10+ years building data-driven systems. Now architecting production-grade agent platforms — from harness design to governance frameworks. Actively contributing to the open-source AI infrastructure that powers the next wave.
From high-level system design to low-level implementation — the complete stack that turns frontier models into production-grade autonomous systems.
Designing the scaffolding around frontier models — context resets, structured handoffs, generator-evaluator loops, and sprint contracts that keep long-running agents coherent across multi-hour sessions. Inspired by Anthropic's harness design research and GAN-inspired generator-evaluator patterns.
Building self-correcting agent loops where an evaluator agent grades outputs against concrete criteria, feeds critique back to the generator, and iterates until quality thresholds are met. Turning subjective judgments into gradable, testable contracts.
Managing what enters the context window, when, and in what form. Compaction, retrieval-augmented context, and structured artifacts — progress files, feature lists — that let agents pick up where the last session left off without guessing.
Planner → generator → evaluator architectures, A2A and MCP protocols, and frameworks like LangGraph, CrewAI, and the OpenAI Agents SDK. Orchestrating specialized agents that each own a slice of the SDLC — and making sure they communicate without losing context.
AWS Bedrock, Azure OpenAI & Foundry, Vertex AI for model serving. Pinecone, FAISS, ChromaDB, and Azure AI Search for retrieval. Building the retrieval and serving layer that production agents depend on — with observability, fallbacks, and cost controls built in.
Bridging the gap between AI capability and business outcomes. Translating model metrics into ROI, designing KPIs that matter to the C-suite, and building dashboards that make AI performance legible to non-technical stakeholders. The best AI systems are the ones nobody notices — because they just work.
Contributing to the AI infrastructure that the ecosystem runs on — fixing real bugs in production-grade agentic frameworks, LLM platforms, and vector databases. The kind of issues that silently corrupt data, break under concurrency, or surface only at scale.
Agent reliability, profile isolation, TUI stability, and cross-platform compatibility. Actively collaborating with the NousResearch team on core agent loop correctness.
View ContributionsAgent node defaults, MCP client resilience, email validation, file handling, and database migration fixes across the platform's agent and API layers.
View ContributionsSDK correctness and edge case handling — streaming accumulator fixes, NO_PROXY sanitization, and tool-call index coalescing for speculative decoding.
View ContributionsAgent framework reliability and type-safe agent interactions. Collaborating on the Pydantic-powered agent ecosystem.
View ContributionsEmbedding function correctness, FTS5 robustness against NUL byte corruption, and dependency hygiene across the AI-native embedding store.
View ContributionsCode agent reliability and tool-use edge cases — including a critical GIL-deadlock fix for arbitrary-precision integer timeouts and managed-agent summary leak prevention.
View ContributionsMulti-agent execution stability and task coordination — fixing asyncio event loop crashes in async native tool paths.
View ContributionsAgent memory retrieval and persistence correctness — auto-propagating embedding dimensions from embedder to vector store config.
View ContributionsAlso contributing to Firecrawl (web data for AI), Cognee (cognitive graph memory), OpenWorker (agentic workflows), and other foundational AI infrastructure projects.
View AllAn architect who ships without a governance layer is shipping liability. I design AI systems that are not just capable but accountable — built to pass audit, survive regulatory scrutiny, and earn trust.
Designing against the frameworks that matter in 2026: the EU AI Act (risk-tiered obligations, high-risk system conformity assessments), the NIST AI Risk Management Framework (Govern–Map–Measure–Manage lifecycle), and ISO/IEC 42001 (the first certifiable AI management system standard). Compliance is designed in from day one — not bolted on after launch.
Bias detection and mitigation in training data and model outputs. Human-in-the-loop checkpoints for high-stakes decisions. Explainability and audit trails so every automated decision can be traced, justified, and challenged. Model cards and system documentation as living artifacts.
As agents grow more autonomous, blast radius grows with them. I design containment boundaries, permission scoping, and tool-use guardrails so an agent that goes off-script cannot take production systems with it. Inspired by Anthropic's containment research for Claude Code and Cowork.
PII handling, retention policies, and data lineage that satisfy GDPR and sector-specific privacy regimes. Retrieval pipelines that respect access controls — no leaking privileged context across tenants. AI-generated content labelling and chatbot disclosure aligned with the EU AI Act's transparency obligations.
Open to architecture challenges, enterprise AI strategy, freelance engagements, or startup ideas in the GenAI space. If you're building something that matters, let's talk.