Thanks to visit codestin.com
Credit goes to github.com

Skip to content
View dante381's full-sized avatar
👨‍💼
👨‍💼

Block or report dante381

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
dante381/README.md

Anish Kshirsagar

Data Engineer · Azure · Spark · Python

Portfolio LinkedIn Email


About

Data Engineer at Northern Tool + Equipment building production-scale data infrastructure on Azure. I work across the full data stack — pipelines, data quality, customer data platforms, and lakehouse architecture.

  • Designed and shipped a Customer MDM system resolving 80M+ records to 40M unique identities at 90%+ precision using probabilistic record linkage and Spark batch embeddings
  • Built an Enterprise Data Reliability Platform cutting downstream incidents by 35% via schema enforcement, null validation, and anomaly detection
  • Architected Azure Lakehouse ETL/ELT across 5+ heterogeneous sources, reducing pipeline runtime by 30–40%

Tech Stack

Languages & Query

Python SQL PySpark Java

Azure Data Platform

Azure Data Factory Azure Synapse Microsoft Fabric Databricks Delta Lake

Data Engineering

Apache Spark ETL Data Modeling Lakehouse dbt

Tools

Git Docker Node.js Next.js


Featured Work

Full case studies with architecture diagrams and impact metrics on my portfolio →

Project What Stack Impact
Customer MDM Resolved 80M+ records to 40M unique identities PySpark, embeddings, cosine similarity 90%+ precision, −25% false matches
Data Reliability Platform Schema enforcement + anomaly detection across all pipelines Python, Azure, custom validators −35% downstream incidents
Azure Lakehouse ETL/ELT 5+ heterogeneous sources → centralised lakehouse ADF, Synapse, Delta Lake, Spark −30–40% pipeline runtime

GitHub Stats

GitHub stats Top languages

GitHub streak


View Portfolio

Pinned Loading

  1. career-copilot career-copilot Public

    Forked from RajjjAryan/career-copilot

    AI-powered job search pipeline built on GitHub Copilot CLI

    JavaScript

  2. chatgpt-clone chatgpt-clone Public

    CSS 1

  3. dante381.github.io dante381.github.io Public

  4. Flow-Puzzle Flow-Puzzle Public

    Connect same-colored dots with pipes to fill every cell on the grid. No empty spaces allowed. 10,000 procedurally generated levels.

    TypeScript

  5. PC_Monitoring PC_Monitoring Public

    C++ 2

  6. Smart-Home-Security Smart-Home-Security Public

    Python