Senior Data Engineer
Abhishek Singh
Senior Data Engineer building cloud data architecture on Azure and GCP. I specialize in event-driven ingestion, Delta lakehouse design, and the serving layers that put fast, reliable analytics in front of the people who need them. Strong on SQL, PySpark, and high-volume data at scale.
Toolbox
Skills & stack
Selected work
Things I've built
Event-Driven Lakehouse
A local-first, backend-agnostic medallion lakehouse with queue-driven ingestion. Reliably ingests irregularly-arriving files with exactly-once processing and schema enforcement across raw → aggregate → curated layers. A queue layer (local watcher or Azure Storage Queue) feeds N async workers with in-flight deduplication, while a single-writer commit buffer prevents Delta Lake write contention and enables near-linear worker scaling. YAML-defined schemas mean zero-Python table definitions, and it runs without a JVM or Spark — querying via DuckDB locally or scaling to ADLS Gen2 in the cloud.
Notebook
Latest writing
JAVA_GATEWAY_EXITED: one error, two root causes
Chasing PySpark's most generic startup failure on Windows to two completely unrelated causes.
Getting Spark to read and write files on Windows
winutils is not enough. Native libraries, scratch directories, stale JVMs, and ADLS SAS auth, from a week of Windows Spark debugging.
Contact
Let's build something awesome.
Got a data platform to design, a pipeline to untangle, or a role in mind? My inbox is open.