Tanush Aggarwal

Data Engineer

  • 15 proven skills
  • @TanushAgg
Request intro

Goes straight to Tanush. No recruiter spam.

Download résumé

Claims you can check

Each card points at something real — open the source and verify for yourself.

  • StatedRésumé

    Hadoop HDFS to Databricks and Snowflake Migration

    Moved Hadoop HDFS Spark pipelines to Databricks and Snowflake, integrating SQL Server, Excel, and flat files via Spark JDBC/PySpark, orchestrating ETL workflows using Apache Airflow, and optimising financial transaction joins.
  • StatedRésumé

    Modular ADF ELT Pipeline

    Built a modular ADF ELT pipeline with parent-child orchestration across 10+ datasets.
  • StatedRésumé

    Data Migration from Hadoop to Cloud

    Transitioned Spark ETL pipelines from Hadoop HDFS to Databricks and Snowflake, integrating SQL Server, Excel and flat files using Spark JDBC and PySpark transformations. Orchestrated PySpark ETL workflows through Apache Airflow, applying broadcast joins, salting and caching to reduce job execution time by 40%. Optimised financial transaction joins by applying filters and column pruning before the join, reducing the data shuffled and cutting join runtime significantly.
  • StatedRésumé

    End-to-End Automated Data Pipeline

    Created a modular ADF ELT pipeline with parent-child orchestration, reducing pipeline duplication across multiple environments. Automated event-driven ingestion across 10+ datasets using ADLS Gen2 triggers and Mapping Data Flows, processing new files automatically. Implemented parameterised datasets and linked services, enabling reusable pipeline deployment across 3 environments—Dev, QA and Production.
  • StatedRésumé

    Data Engineer - Associate Software Engineer @ Intelibliss

    Engineered PySpark Databricks notebooks across 5+ source formats. Built Databricks, PySpark, Azure Data Factory and ADLS Gen2 solutions with Bronze, Silver and Gold layers. Drove Delta Lake optimisation, reducing query latency by 20%. Tuned Databricks Spark configurations, reducing execution time from 20 to 12 minutes (40%). Migrated Hadoop HDFS Spark pipelines to Databricks and Snowflake. Implemented AIOps Metric Intelligence and Predictive Intelligence Models (routing 35% of tasks automatically). Automated IT and HR workflows, reducing ticket volume by 25%.
  • StatedRésumé

    Data Engineer - Associate Engineer @ Emerson (Contractor)

    Optimised Apache Spark jobs using broadcast joins, salting and caching, improving ETL pipeline performance 30%. Developed Power BI dashboards tracking 25+ HR and payroll KPIs. Analysed 10K+ employee records using SQL. Diagnosed reporting defects through root-cause analysis. Collaborated with cross-functional stakeholders.
  • StatedRésumé

    Data Analytics Intern @ Shree Adisoft Technology Pvt. Ltd

    Developed 10+ Power BI dashboards consolidating 25+ KPIs. Maintained Power BI reporting workflows supporting 30+ monthly report refreshes.
  • StatedRésumé

    Data Engineer - Associate Software Engineer @ Intelibliss

    Engineered PySpark Databricks notebooks across 5+ source formats (cloud, API, CSV, JSON, Parquet). Built Databricks, Azure Data Factory, and ADLS Gen2 solutions with a 3-layer medallion architecture (Bronze, Silver, Gold). Optimized Delta Lake with OPTIMIZE, Z-Ordering, VACUUM, Liquid Clustering, and compaction. Tuned Spark configurations with broadcast joins, salting, caching, and persistence. Migrated Hadoop HDFS pipelines to Databricks and Snowflake with Apache Airflow orchestration. Implemented AIOps Metric Intelligence, Predictive Intelligence, and automated workflows via Virtual Agent and Flow Designer.
  • StatedRésumé

    Data Engineer, Associate Engineer @ Emerson

    Optimised Apache Spark jobs using broadcast joins, salting and caching. Developed Power BI dashboards tracking 25+ HR and payroll KPIs. Analysed 10,000+ employee records using SQL for data validation and recurring reporting. Conducted root-cause analysis on reporting defects and collaborated with 5+ cross-functional stakeholders.
  • StatedRésumé

    Data Analytics Intern @ Shree Adisoft Technology Pvt. Ltd.

    Developed 10+ Power BI dashboards consolidating 25+ KPIs to improve visibility across operational performance and business metrics. Maintained Power BI reporting workflows supporting 30+ monthly report refreshes.
  • StatedRésumé

    Software Engineer Intern @ Water And Power Consultancy Services (Wapcos)

    Updated 10+ official web pages using HTML and CSS, maintaining consistent layouts and supporting smooth website functionality.
  • StatedRésumé

    Rising Star Award - Intelibliss

  • StatedRésumé

    Spot Award - Intelibliss

  • StatedRésumé

    Rising Star Award

  • StatedRésumé

    Spot Award

  • StatedRésumé

    Python(Basic)

  • StatedRésumé

    Data Manipulation In Python

  • StatedRésumé

    Python Foundation with Data Structures and Algorithms

What Tanush can prove

Proven Unproven

See all skills →
INTERVIEW-VERIFIED CAPABILITY (PROVEN)SELF-DECLARED STACK
GRAPH NODES: 87 · CONNECTIVITY FACTOR: 0.17
Skill constellation: 87 skills, 0 interview-verified.ADLS GEN2AIOPS METRIC INTELLIGENCEAPACHE SPARKAUTOMATED REPORTINGAUTOMATED WORKFLOWSAZURE BLOB STORAGEAZURE DATA FACTORYAZURE DATA LAKE STORAGE GEN2AZURE DATABRICKSAZURE SYNAPSE ANALYTICSBATCH PROCESSINGBROADCAST JOINSBRONZE-SILVER-GOLD ARCHITECTUREBUSINESS REPORTINGCACHINGCI/CD CONCEPTSCSSCSVDASHBOARD DEVELOPMENTDATA ACCURACY VALIDATIONDATA ANALYSISDATA INGESTIONDATA MIGRATIONDATA MODELLINGDATA PIPELINE DEVELOPMENTDATA QUALITY CHECKSDATA STRUCTURESDATA TRANSFORMATIONDATA VALIDATIONDATA WAREHOUSINGDATABRICKSDELTA LAKEDELTA TABLESDEV/QA/PRODUCTION DEPLOYMENTELT PIPELINESENGLISHENVIRONMENT MANAGEMENTETL DEVELOPMENTETL/ELT PIPELINESEVENT-DRIVEN PROCESSINGFILE COMPACTIONFILE PROCESSINGFLOW DESIGNERGITGITHUBHADOOP HDFSHINDIHTMLJAVASCRIPTJOB SCHEDULINGJSONKPI REPORTINGLAKEHOUSE ARCHITECTURELINKED SERVICESLIQUID CLUSTERINGMAPPING DATA FLOWSMEDALLION ARCHITECTUREMICROSOFT AZUREOPTIMIZEPARAMETERIZED DATASETSPARENT-CHILD PIPELINESPARQUETPARTITIONINGPERSISTENCEPIPELINE ORCHESTRATIONPLPGSQLPOWER BIPREDICTIVE INTELLIGENCEPYSPARKPYTHONRELATIONAL DATABASESREQUIREMENTS GATHERINGROOT CAUSE ANALYSISSALTINGSHUFFLE OPTIMISATIONSNOWFLAKESPARK PERFORMANCE TUNINGSPARK SQLSQLSQL SERVERSTORAGE EVENT TRIGGERSUATVACUUMVIRTUAL AGENTWORKFLOW ORCHESTRATIONZ-ORDERINGAPACHE AIRFLOW
  • ADLS Gen2: self-declared, 1 evidence items, linked to Project: End-to-End Automated Data Pipeline
  • AIOps Metric Intelligence: self-declared, 0 evidence items
  • Apache Airflow: self-declared, 2 evidence items, linked to Project: Hadoop HDFS to Databricks and Snowflake Migration, Project: Data Migration from Hadoop to Cloud
  • Apache Spark: self-declared, 0 evidence items
  • Automated Reporting: self-declared, 0 evidence items
  • Automated Workflows: self-declared, 0 evidence items
  • Azure Blob Storage: self-declared, 0 evidence items
  • Azure Data Factory: self-declared, 1 evidence items, linked to Project: Modular ADF ELT Pipeline
  • Azure Data Lake Storage Gen2: self-declared, 0 evidence items
  • Azure Databricks: self-declared, 0 evidence items
  • Azure Synapse Analytics: self-declared, 0 evidence items
  • Batch Processing: self-declared, 0 evidence items
  • Broadcast Joins: self-declared, 1 evidence items, linked to Project: Data Migration from Hadoop to Cloud
  • Bronze-Silver-Gold Architecture: self-declared, 0 evidence items
  • Business Reporting: self-declared, 0 evidence items
  • Caching: self-declared, 1 evidence items, linked to Project: Data Migration from Hadoop to Cloud
  • CI/CD Concepts: self-declared, 0 evidence items
  • CSS: self-declared, 0 evidence items
  • CSV: self-declared, 0 evidence items
  • Dashboard Development: self-declared, 0 evidence items
  • Data Accuracy Validation: self-declared, 0 evidence items
  • Data Analysis: self-declared, 0 evidence items
  • Data Ingestion: self-declared, 0 evidence items
  • Data Migration: self-declared, 1 evidence items, linked to Project: Data Migration from Hadoop to Cloud
  • Data Modelling: self-declared, 0 evidence items
  • Data Pipeline Development: self-declared, 0 evidence items
  • Data Quality Checks: self-declared, 0 evidence items
  • Data Structures: self-declared, 0 evidence items
  • Data Transformation: self-declared, 0 evidence items
  • Data Validation: self-declared, 0 evidence items
  • Data Warehousing: self-declared, 0 evidence items
  • Databricks: self-declared, 2 evidence items, linked to Project: Hadoop HDFS to Databricks and Snowflake Migration, Project: Data Migration from Hadoop to Cloud
  • Delta Lake: self-declared, 0 evidence items
  • Delta Tables: self-declared, 0 evidence items
  • Dev/QA/Production Deployment: self-declared, 0 evidence items
  • ELT Pipelines: self-declared, 0 evidence items
  • English: self-declared, 0 evidence items
  • Environment Management: self-declared, 0 evidence items
  • ETL Development: self-declared, 0 evidence items
  • ETL/ELT Pipelines: self-declared, 0 evidence items
  • Event-Driven Processing: self-declared, 0 evidence items
  • File Compaction: self-declared, 0 evidence items
  • File Processing: self-declared, 0 evidence items
  • Flow Designer: self-declared, 0 evidence items
  • Git: self-declared, 0 evidence items
  • GitHub: self-declared, 0 evidence items
  • Hadoop HDFS: self-declared, 2 evidence items, linked to Project: Hadoop HDFS to Databricks and Snowflake Migration, Project: Data Migration from Hadoop to Cloud
  • Hindi: self-declared, 0 evidence items
  • HTML: self-declared, 0 evidence items
  • JavaScript: self-declared, 0 evidence items
  • Job Scheduling: self-declared, 0 evidence items
  • JSON: self-declared, 0 evidence items
  • KPI Reporting: self-declared, 0 evidence items
  • Lakehouse Architecture: self-declared, 0 evidence items
  • Linked Services: self-declared, 1 evidence items, linked to Project: End-to-End Automated Data Pipeline
  • Liquid Clustering: self-declared, 0 evidence items
  • Mapping Data Flows: self-declared, 1 evidence items, linked to Project: End-to-End Automated Data Pipeline
  • Medallion Architecture: self-declared, 0 evidence items
  • Microsoft Azure: self-declared, 0 evidence items
  • OPTIMIZE: self-declared, 0 evidence items
  • Parameterized Datasets: self-declared, 0 evidence items
  • Parent-Child Pipelines: self-declared, 0 evidence items
  • Parquet: self-declared, 0 evidence items
  • Partitioning: self-declared, 0 evidence items
  • Persistence: self-declared, 0 evidence items
  • Pipeline Orchestration: self-declared, 0 evidence items
  • PLpgSQL: self-declared, 0 evidence items
  • Power BI: self-declared, 0 evidence items
  • Predictive Intelligence: self-declared, 0 evidence items
  • PySpark: self-declared, 2 evidence items, linked to Project: Hadoop HDFS to Databricks and Snowflake Migration, Project: Data Migration from Hadoop to Cloud
  • Python: self-declared, 0 evidence items
  • Relational Databases: self-declared, 0 evidence items
  • Requirements Gathering: self-declared, 0 evidence items
  • Root Cause Analysis: self-declared, 0 evidence items
  • Salting: self-declared, 1 evidence items, linked to Project: Data Migration from Hadoop to Cloud
  • Shuffle Optimisation: self-declared, 0 evidence items
  • Snowflake: self-declared, 2 evidence items, linked to Project: Hadoop HDFS to Databricks and Snowflake Migration, Project: Data Migration from Hadoop to Cloud
  • Spark Performance Tuning: self-declared, 0 evidence items
  • Spark SQL: self-declared, 0 evidence items
  • SQL: self-declared, 2 evidence items, linked to Project: Hadoop HDFS to Databricks and Snowflake Migration, Project: Data Migration from Hadoop to Cloud
  • SQL Server: self-declared, 2 evidence items, linked to Project: Hadoop HDFS to Databricks and Snowflake Migration, Project: Data Migration from Hadoop to Cloud
  • Storage Event Triggers: self-declared, 0 evidence items
  • UAT: self-declared, 0 evidence items
  • VACUUM: self-declared, 0 evidence items
  • Virtual Agent: self-declared, 0 evidence items
  • Workflow Orchestration: self-declared, 0 evidence items
  • Z-Ordering: self-declared, 0 evidence items
No architecture case studies yet

Flagship systems this candidate architected will appear here once attached to the dossier. No project has been added yet.

Getting better, on the record

Proven skills over time — today marked in lime.

No scored sessions yet

This trajectory populates once the candidate completes scored interview sessions. No score has been recorded or estimated here.

Career chronology

Gaps shown honestly — only what's on the record.

Trajectory timeline
  1. Spot Award - Intelibliss

  2. Rising Star Award - Intelibliss

  3. Data Engineer - Associate Software Engineer

    Intelibliss

  4. Data Engineer - Associate Software Engineer

    Intelibliss

  5. Data Engineer - Associate Engineer

    Emerson (Contractor)

  6. Data Engineer, Associate Engineer

    Emerson

  7. Data Analytics Intern

    Shree Adisoft Technology Pvt. Ltd

  8. Data Analytics Intern

    Shree Adisoft Technology Pvt. Ltd.

  9. Udemy

    Course Certificate

  10. Software Engineer Intern

    Water And Power Consultancy Services (Wapcos)

  11. Coding Ninjas

    Course Certificate

  12. Guru Gobind Singh Indraprastha University

    Bachelor of Engineering

  13. Guru Gobind Singh Indraprastha University

    Bachelor of Technology - BTech

Hiring for Data Engineer?

Request the full evidence dossier — verified interview sessions, cryptographically attested artifacts, and reference outcomes.

Request intro

Goes straight to Tanush. No recruiter spam.

Download résumé
Verification key —
Request intro