Available for contract & full-time remote

Jeremy Tran

Data Engineer & AI Developer

Building production-grade batch and streaming pipelines since 2020 on GCP and AWS. Delivered data systems for Walmart, Macy's, and BlueCross BlueShield. Google Professional Data Engineer certified.

Python SQL Apache Airflow dbt Apache Flink Apache Iceberg Apache Kafka DuckDB GCP BigQuery AWS Databricks Snowflake TypeScript Node.js React Docker LangChain Vector DBs LLM Agents

Projects

Production-oriented pipelines built with modern open-source tooling.

Healthcare Data Engineering Pipeline

Local open lakehouse pipeline processing synthetic Synthea patient records via two ingestion paths — structured CSV and FHIR R4. Airflow orchestrates staging to MinIO, ingestion into Apache Iceberg via Trino and Apache Polaris (REST catalog), and dbt transformations producing a clinical mart. A Streamlit dashboard serves KPIs directly over Trino.

Airflow dbt Apache Iceberg Trino Apache Polaris MinIO FHIR R4 Docker

ScoreChat

Symbolic score analysis and retrieval system for classical piano music in Humdrum format. Parses Beethoven sonatas via music21, stores versioned measure-level analyses (key, harmony, rhythm, texture) in PostgreSQL with pgvector, and uses hybrid vector retrieval to answer natural-language questions with notation-backed evidence. A Verovio WASM renderer displays exact SVG score slices in the browser alongside LLM explanations.

Python PostgreSQL pgvector music21 RAG Verovio OpenAI Docker
More on GitHub

Services

Available for short-term engagements and project-based contracts. Reach out to discuss scope and rates.

Pipeline Engineering

Batch and streaming ETL/ELT pipelines from ingestion to analytics-ready outputs. Airflow, Flink, Kafka, dbt, Apache Iceberg.

Cloud Data Architecture

Data infrastructure design and migration on GCP and AWS. BigQuery optimization, Lakehouse architecture, cost-effective storage strategy.

AI & Agent Workflows

LLM-powered automation using LangChain and multi-agent patterns. RAG pipelines, SQL agents, structured extraction, and workflow orchestration.

Data Quality & Modeling

Schema design, data validation frameworks, and dbt modeling. Audit SQL, fix data integrity issues, and build stakeholder-ready outputs.

AI/ML Data Preparation

Training and evaluation dataset production for ML models — schema consistency, edge case coverage, labeling accuracy, and annotation QA.

Let's Work Together

Have a project in mind? Available for contracts, consulting, and remote full-time roles.

Start a conversation

About

I'm Jeremy Tran, a data engineer and AI developer based in Irvine, CA, building production data systems since 2020. I've worked across the stack at Tredence (Walmart), Infosys (Macy's, BlueCross BlueShield), and as an independent AI/ML data contractor — with a consistent focus on data quality, schema evolution, and analytics-ready outputs.

I hold a BA in Neuroscience from Pomona College and am Google Professional Data Engineer certified. Currently open to remote contract and full-time roles in data engineering and AI.

  • Google Professional Data Engineer Google Cloud Certified
  • AI/ML Data Contractor Freelance, Remote · January 2024 – Present
  • Tredence — Data Engineering Consultant Walmart · GCP BigQuery · Python · PowerBI
  • Infosys — Associate, Data Engineering Macy's · BlueCross BlueShield · Airflow · BigQuery

Languages

Python SQL Java Scala JavaScript/TypeScript

Pipelines & Orchestration

Apache Airflow Apache Flink Apache Kafka dbt Apache Spark

Data & Storage

Apache Iceberg DuckDB BigQuery Databricks Snowflake

Cloud & Infra

GCP AWS Azure Docker Linux

Web

TypeScript Node.js React

AI & Tooling

LangChain Vector DBs RAG Prompt Engineering

Contact

Get in touch.

Open to remote contract work and full-time roles in data engineering and AI.