MS AI @ Northeastern (Silicon Valley) | GPU Kernel & Inference Optimization (CUDA / CuTe DSL) | Silicon Valley Fellow W25
๐ I write articles on FreeCodeCamp ย ยทย ๐ซ Reach me at [email protected]
I make AI models run faster and cheaper by writing GPU kernels close to the metal.
Currently working on GPU kernel optimization with CuTe DSL.
What I'm building
- โก
fast_trimul- drop-in Fused Triangle Multiplicative Update kernel for AlphaFold3-family models (OpenFold-3, Boltz, Chai, Protenix). Up to 4.5โ6.8x faster inference, 2.2โ2.4x less peak VRAM. - ๐ง CUDA / CuTe DSL kernels - softmax attention 62x faster (22.8ms โ 0.36ms) via memory-hierarchy tuning; #2 on LeetGPU's leaderboard.
- ๐ฌ Research in tabular machine learning and synthetic data generation (multiple articles under review / in final writing).
- ๐ Author of The Math Behind AI (160+ โญ) and How to Build Optimal AI Agents - 250k+ readers across my writing.
Foundation: Hardware-level + deep math. In my ECE bachelor's I built a CPU from scratch (transistor up), with a strong math base in linear systems, calculus, differential equations, complex analysis, PDEs, and functional analysis.



