Applied AI built with measurement discipline: production LLM products, model-training and evaluation infrastructure, and agentic engineering tooling.
| Project | What it is |
|---|---|
| bufrai-model-harness | File-backed tracking for LLM fine-tuning runs and evals - versioned suites, paired statistical comparison with confidence intervals, build-vs-buy cost modelling. Apache-2.0. |
| model-training-curriculum | Hands-on PyTorch curriculum (Track I public): deliberate-failure teaching arcs, a blind-graded capstone, CI authoring gates. CC BY-NC-SA 4.0. |
| arduino-mcp | MCP server wrapping arduino-cli so AI agents can detect boards, compile, upload, and talk serial. Apache-2.0. |
Copasaidit (copasaidit.com) - production SaaS using the Anthropic Claude API for real-time message moderation and structured data extraction, gated by an AI evaluation harness before every release.
A run whose config you cannot reproduce did not happen. A score without a suite version is not a result. A score difference is a statistical claim. Every price carries provenance.
These are the four rules the harness enforces - and the discipline behind everything else here.
Contact: [email protected]