Data Engineer at Northern Tool + Equipment building production-scale data infrastructure on Azure. I work across the full data stack — pipelines, data quality, customer data platforms, and lakehouse architecture.
- Designed and shipped a Customer MDM system resolving 80M+ records to 40M unique identities at 90%+ precision using probabilistic record linkage and Spark batch embeddings
- Built an Enterprise Data Reliability Platform cutting downstream incidents by 35% via schema enforcement, null validation, and anomaly detection
- Architected Azure Lakehouse ETL/ELT across 5+ heterogeneous sources, reducing pipeline runtime by 30–40%
Languages & Query
Azure Data Platform
Data Engineering
Tools
Full case studies with architecture diagrams and impact metrics on my portfolio →
| Project | What | Stack | Impact |
|---|---|---|---|
| Customer MDM | Resolved 80M+ records to 40M unique identities | PySpark, embeddings, cosine similarity | 90%+ precision, −25% false matches |
| Data Reliability Platform | Schema enforcement + anomaly detection across all pipelines | Python, Azure, custom validators | −35% downstream incidents |
| Azure Lakehouse ETL/ELT | 5+ heterogeneous sources → centralised lakehouse | ADF, Synapse, Delta Lake, Spark | −30–40% pipeline runtime |


