Big Data & GenAI Engineer · Pharma & Life Sciences

Data platforms
that hold at scale.

I build fault-tolerant pipelines over 200M+ record workloads — Spark, Kafka, and Airflow — and the lakehouses and GenAI agents that run on top of them. Currently engineering pharma data products at ZS Associates.

3B+
claims records processed
200M+
pipeline record scale
3+ yrs
distributed data systems
50RPS
GenAI agent, zero drops
// 01 selected work

Systems I've built

Distributed data platforms and production AI — from self-managed lakehouses to agentic GenAI. A few I can point at.

Personal · infraopen lakehouse

Open Lakehouse Platform

A production-style lakehouse I run myself — the stack Databricks and Fabric abstract away. A Spark 4.0 SQL engine over a Hive Metastore catalog and MinIO S3 storage with Delta ACID tables, every service containerized across isolated namespaces with idempotent deploys and an end-to-end smoke test.

Spark 4.0Delta LakeHive MetastoreKubernetesMinIO
View on GitHub →
Personal · GenAIagentic

AtlasCare — Agentic GenAI Support

An end-to-end GenAI agent that automates Tier-1 support, orchestrating four enterprise backends through LLM tool-calling with a KB-driven RAG policy engine — policy changes ship without a code deploy. Async payment pipeline with semaphore-gated concurrency, auto-retry, and full p50/p95/p99 observability, load-tested at 50 RPS with zero drops.

LLM OrchestrationRAGTool-CallingAsync
View on GitHub →
Personal · distributed systemspayments

FinFlow — Distributed Payments

Event-driven Go microservices over Apache Kafka on Kubernetes, with a double-entry ledger using HOLD-based reservation and idempotent processing to prevent duplicates under failure. Redis-backed stateless services with Prometheus/Grafana monitoring.

GoKafkaKubernetesRedisPrometheus
View on GitHub →
// 02 experience

Where I've worked

2025 — now
Gurugram, IN

Business Technology Solutions Engineer

ZS Associates · Pharma & Life Sciences

  • Payer-performance pipelines (PB & MB) over 3B+ claims — custom PySpark controller/executor on Kubernetes, orchestrated by Airflow, serving Redshift for Tableau.
  • Built a Microsoft Fabric backend + a production data-quality gate that blocks bad refreshes from reaching Power BI dashboards.
  • SLE patient-journey lakehouse: cohort funnel, clinical profiling, and ML-derived journey archetypes.
2024
Chennai, IN

Software Engineer

FourKites

  • Event-driven vessel-metadata ingestion with confidence-scored dedup — SQS + async workers upserting PostgreSQL, DML audit logs to S3.
  • Quantified carrier ETA accuracy from 200K+ daily vessel pings into shipper-facing carrier recommendations.
2021 — 2022
Bangalore, IN

Data Engineer

Tata Consultancy Services · ASML Netherlands

  • End-to-end ML pipeline on semiconductor manufacturing data — unified fragmented sensor logs, engineered 40+ features, automated KPI monitoring.
  • Improved XGBoost accuracy 15% via recursive feature engineering, Bayesian tuning, and NSGA-II feature selection.
// 03 toolkit

What I work with

languages

PythonGoSQLPySparkSpark SQL

big_data

Apache SparkStructured StreamingKafkaAirflowHadoopHive Metastore

lakehouse

Delta LakeMedallionUnity CatalogDatabricksMicrosoft FabricSnowflake

genai

LLM IntegrationAgentic AIRAGLangChainVector EmbeddingsMLflow

cloud_devops

AWS (S3 · SQS · SageMaker · Redshift)KubernetesDockerJenkinsCI/CD

data_tools

PostgreSQLRedisMongoDBPower BITableauGrafana
// let's talk

Building something at scale?
gauravsingh4frnds@gmail.com