Big Data Engineer
& Cloud Architect

Building and scaling data platforms that process 300B+ records a day

I lead large-scale pipeline migrations, own data-platform architecture decisions, and deliver measurable gains in reliability, cost, and storage efficiency.

300B+Records/day processed
99.9%Pipeline uptime
3xStorage footprint cut
80%Less manual ETL
Nikhil Kumar

Professional Experience

Delivered high-impact data solutions across cloud platforms and big data technologies for leading global enterprises.

Microsoft

Software Engineer

Aug 2022 - Present
  • Drove 2-3% improvement in load balancing for CosmosDB cluster deployments as part of Azure Cloud Supply Chain Data team, directly impacting resource forecasting accuracy
  • Led migration from legacy BigData Storage Framework to modern Spark-based architecture using ADF, Synapse, Databricks, and Azure Functions
  • Architected and deployed new Data Quality framework, stabilizing Storage telemetry across multiple Azure teams
  • Implemented Erasure Coding techniques, achieving significant reduction in Storage infrastructure footprint
  • Own architecture and design decisions for storage data-platform initiatives, partnering across multiple Azure engineering teams and mentoring engineers on Spark and cloud-native best practices

Epsilon

Software Engineer 2

June 2021 - Aug 2022
  • Engineered framework to migrate MapReduce jobs to Spark, modernizing legacy LinkedIn Camus-based data ingestion pipeline
  • Achieved 3x reduction in HDFS disk footprint through advanced ORC file and Kafka compression optimization, including linger/batch size tuning, sticky partitioning, and high-cardinality column sorting
  • Built comprehensive Kibana Dashboard for Spark Job metrics monitoring using ELK stack, improving observability across data engineering teams
  • Maintained production data pipeline processing 300B+ records daily with 99.9% uptime through on-call support and proactive troubleshooting

HashMap Inc

Cloud Data Engineer

June 2018 - June 2021
  • Waste Management: Integrated on-premise data sources (Netezza, Oracle) and cloud platforms (Google Analytics, S3, REST APIs) with Snowflake using Matillion ETL; automated AWS resource deployment with CloudFormation scripts
  • LAM Research: Developed Spark batch and structured streaming applications for multi-format data parsing (CSV, JSON, Excel); created Python analytics library enabling SQL-like querying of HBase time-series data
  • Murphy Oil: Built automation accelerators for Excel file ingestion using Crealytics library, reducing manual ETL processes by 80%

Case Studies

A closer look at three engagements โ€” the problem, the architecture decisions I drove, and the measurable outcome.

MapReduce → Spark Pipeline Modernization

Epsilon · LinkedIn Camus-based ingestion platform

Problem A legacy MapReduce ingestion pipeline was slow, storage-heavy, and expensive to operate at 300B+ records/day, with limited observability into job health.
Approach Engineered a framework to migrate MapReduce jobs to Spark and tuned Kafka and ORC storage โ€” linger/batch sizing, sticky partitioning, and high-cardinality column sorting โ€” then built an ELK/Kibana dashboard for Spark job metrics.
Impact 3x reduction in HDFS disk footprint and full pipeline observability, sustaining 99.9% uptime on 300B+ records processed daily.

Multi-Cloud Data Integration to Snowflake

HashMap Inc · Waste Management

Problem Data was fragmented across on-premise systems (Netezza, Oracle) and cloud sources (Google Analytics, S3, REST APIs) with no unified analytics layer.
Approach Integrated all sources into Snowflake using Matillion ETL and automated AWS resource provisioning with CloudFormation for repeatable, hands-off deployments.
Impact A single governed analytics platform with automated infrastructure, eliminating manual provisioning and unifying reporting across the business.

Real-Time Streaming & Time-Series Analytics

HashMap Inc · LAM Research

Problem Multi-format sensor and operational data (CSV, JSON, Excel) needed to be parsed and queried as time-series with SQL-like access for analysts.
Approach Built Spark batch and structured-streaming applications for multi-format parsing and authored a Python analytics library enabling SQL-like querying of HBase time-series data.
Impact Analysts gained self-serve, SQL-like access to high-volume time-series data without deep HBase expertise.

Technical Skills

Comprehensive toolkit spanning big data technologies, cloud platforms, and modern data engineering practices.

Languages

Primary: Python, Java & SQL.

Java Python SQL

Big Data & Streaming

Production Spark & Kafka at 300B+ records/day.

Apache Spark Spark SQL Structured Streaming Apache Kafka Delta Lake

Cloud Data Platforms

Databricks, Synapse & ADF running in production.

Databricks Snowflake Azure Synapse Analytics Azure Data Factory

Databases

OLTP, analytical & document stores at scale.

PostgreSQL MySQL MongoDB Azure Cosmos DB Netezza

DevOps & Observability

CI/CD plus observability with the ELK Stack.

Git Jenkins Docker ELK Stack

Education & Certifications

B.Tech in Computer Science

Army Institute of Technology

Information Technology

July 2014 - May 2018

Professional Certifications

๐Ÿ† Google Cloud Professional Data Engineer
๐Ÿ† Snowflake SnowPro Core Certified
๐Ÿ† Applied Data Science - WorldQuant University
๐Ÿ† Mountain Mover Award - Microsoft

Get In Touch

Open to discussing new opportunities, innovative projects, or collaborations in the Big Data and Cloud space. Let's build something great together.

Send Message