Claims you can check
Each card points at something real — open the source and verify for yourself.
- StatedRésumé
CDC Lakehouse Pipeline
Built an end-to-end CDC pipeline streaming PostgreSQL inserts, updates, and deletes through Debezium and Kafka into Apache Iceberg tables with 1.03s P95 latency. Integrated MinIO, Hive Metastore, and Trino for queryable lakehouse analytics; containerized a 7-service stack for reproducible deployment.Check source - StatedRésumé
Databricks Taxi Analytics Pipeline
Built an incremental Bronze–Silver–Gold pipeline processing 355M+ taxi trip records with schema normalization, data quality filtering, and deduplication. Designed partition-aware incremental queries that cut data scanned by 89% and execution time by 33%, producing 3 analytics-ready Gold datasets.Check source - StatedRésumé
Real-Time Crypto Data Engineering Pipeline
Built a real-time pipeline streaming CoinGecko market events through Kafka into Amazon S3 and Snowflake via Snowpipe. Developed incremental dbt transformations with MERGE-based upserts across RAW, STAGING, and GOLD layers, orchestrated with Airflow.Check source - StatedRésuméConverted legacy ADB script outputs from raw text to structured JSON across OTT device/app profiling workflows, reducing storage overhead by 30–40% and cutting downstream parsing time by an estimated ~35%. Optimized memory and CPU profiling workflows for OTT applications and devices, reducing job runtime by an estimated ~20% and improving data collection consistency across device models. Designed and deployed a production ingestion API for crash and profiling data, integrating ADB tooling, custom parsing logic, and database storage into a pipeline used across daily QA test cycles. Built automated collection of device-level metrics (RAM usage, foreground apps, storage stats) using Python and shell scripts, emitting structured JSON for backend ingestion. Debugged data-loss and parsing issues across log parsing, ADB, JSON/TXT handling, and backend database storage, ensuring accurate logging of key fields and improving pipeline reliability.
- StatedRésumé
Software Engineer Intern @ Jio Platforms Limited (JPL)
Worked as a Software Engineering Intern on automation, profiling, and data ingestion workflows for OTT devices and applications. • Worked with Python and ADB to automate device interactions and collect application/device profiling data. • Developed and optimized memory and CPU profiling workflows for OTT applications across different device models and configurations. • Built parsing workflows to transform raw ADB/script outputs into structured JSON for easier storage and downstream processing. • Developed a production ingestion API using Node.js and Express.js for crash and profiling data, integrating ADB tooling, custom parsing logic, and database storage. • Worked with REST APIs, JSON-based data processing, database operations, logging, and automation as part of daily QA testing workflows. • Improved profiling and data-processing workflows, achieving measurable reductions in storage overhead, parsing time, and profiling runtime. - StatedRésumé
AWS re/Start Graduate
Amazon Web Services Training and Certification - StatedRésumé
Introduction to Data Engineering
DeepLearning.AI & Amazon Web Services (Coursera) - StatedRésumé
LeetCode and GeeksforGeeks Coding Achievement
Solved 700+ problems on LeetCode; top 0.5% of users in 2025 (Contest Rating: 1575). GeeksforGeeks Coding Score: 313.
What Abhiraj can prove
Proven Unproven
- ADB: self-declared, 0 evidence items
- Airflow: self-declared, 1 evidence items, linked to Project: Real-Time Crypto Data Engineering Pipeline
- Amazon Redshift: self-declared, 0 evidence items
- Apache Iceberg: self-declared, 1 evidence items, linked to Project: CDC Lakehouse Pipeline
- Apache Spark: self-declared, 0 evidence items
- Athena: self-declared, 0 evidence items
- Auto Loader: self-declared, 0 evidence items
- Automation: self-declared, 0 evidence items
- AWS CLI: self-declared, 0 evidence items
- BigQuery: self-declared, 0 evidence items
- Boto3: self-declared, 0 evidence items
- Business Analysis: self-declared, 0 evidence items
- Business Development: self-declared, 0 evidence items
- CDC: self-declared, 1 evidence items, linked to Project: CDC Lakehouse Pipeline
- CI/CD: self-declared, 0 evidence items
- CloudWatch: self-declared, 0 evidence items
- CSS: self-declared, 0 evidence items
- Data Modeling: self-declared, 0 evidence items
- Data Structures: self-declared, 0 evidence items
- Databricks Workflows: self-declared, 0 evidence items
- dbt: self-declared, 1 evidence items, linked to Project: Real-Time Crypto Data Engineering Pipeline
- Debezium: self-declared, 1 evidence items, linked to Project: CDC Lakehouse Pipeline
- Delta Lake: self-declared, 1 evidence items, linked to Project: Databricks Taxi Analytics Pipeline
- Docker: self-declared, 0 evidence items
- ETL/ELT: self-declared, 0 evidence items
- Express.js: self-declared, 0 evidence items
- Git: self-declared, 0 evidence items
- GitHub: self-declared, 0 evidence items
- Glue: self-declared, 0 evidence items
- Hive Metastore: self-declared, 1 evidence items, linked to Project: CDC Lakehouse Pipeline
- IAM: self-declared, 0 evidence items
- Incremental Processing: self-declared, 0 evidence items
- Java: self-declared, 0 evidence items
- JavaScript: self-declared, 0 evidence items
- JSON: self-declared, 0 evidence items
- Kafka: self-declared, 2 evidence items, linked to Project: CDC Lakehouse Pipeline, Project: Real-Time Crypto Data Engineering Pipeline
- Kotlin: self-declared, 0 evidence items
- Lakebase: self-declared, 0 evidence items
- Lakeflow Declarative Pipelines: self-declared, 0 evidence items
- Lambda: self-declared, 0 evidence items
- LangChain: self-declared, 0 evidence items
- Linux: self-declared, 0 evidence items
- LLM-Assisted Development: self-declared, 0 evidence items
- Medallion Architecture: self-declared, 0 evidence items
- MinIO: self-declared, 1 evidence items, linked to Project: CDC Lakehouse Pipeline
- MongoDB: self-declared, 0 evidence items
- MySQL: self-declared, 0 evidence items
- Node.js: self-declared, 0 evidence items
- OpenAI API: self-declared, 0 evidence items
- PostgreSQL: self-declared, 1 evidence items, linked to Project: CDC Lakehouse Pipeline
- Postman: self-declared, 0 evidence items
- Prompt Engineering: self-declared, 0 evidence items
- PySpark: self-declared, 1 evidence items, linked to Project: Databricks Taxi Analytics Pipeline
- Python: self-declared, 0 evidence items
- RAG: self-declared, 0 evidence items
- REST APIs: self-declared, 0 evidence items
- S3: self-declared, 1 evidence items, linked to Project: Real-Time Crypto Data Engineering Pipeline
- Snowflake: self-declared, 1 evidence items, linked to Project: Real-Time Crypto Data Engineering Pipeline
- Snowpipe: self-declared, 1 evidence items, linked to Project: Real-Time Crypto Data Engineering Pipeline
- SQL: self-declared, 0 evidence items
- TypeScript: self-declared, 0 evidence items
- Unity Catalog: self-declared, 1 evidence items, linked to Project: Databricks Taxi Analytics Pipeline
- Vector Databases (ChromaDB): self-declared, 0 evidence items
Shipped work
See all workFlagship systems this candidate architected will appear here once attached to the dossier. No project has been added yet.
Getting better, on the record
Proven skills over time — today marked in lime.
This trajectory populates once the candidate completes scored interview sessions. No score has been recorded or estimated here.
Career chronology
Gaps shown honestly — only what's on the record.
Software Engineering Intern
Jio Platforms Limited
Software Engineer Intern
Jio Platforms Limited (JPL)
Ajeenkya D.Y. Patil University
B.Tech
Ajeenkya D Y Patil University
JAIN INTERNATIONAL SCHOOL AURANGABAD - India
The Stepping Stone School - India
Tender Care Home
Hiring for Data Engineer?
Request the full evidence dossier — verified interview sessions, cryptographically attested artifacts, and reference outcomes.
Goes straight to Abhiraj. No recruiter spam.