
<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Engineering Agents]]></title><description><![CDATA[Friendly notes on developing, and developing with, AI agents.

Engineering Agents is a place where developers can share their stories of how they are building, and building with, AI agents. Heavy on code; light on fluff! ]]></description><link>https://engineeringagents.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!eRFm!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0861dada-274e-4eb5-8137-e122bd7a4c3d_1024x1024.png</url><title>Engineering Agents</title><link>https://engineeringagents.substack.com</link></image><generator>Substack</generator><lastBuildDate>Mon, 14 Sep 2026 01:29:28 GMT</lastBuildDate><atom:link href="https://engineeringagents.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Russ Miles]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[engineeringagents@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[engineeringagents@substack.com]]></itunes:email><itunes:name><![CDATA[Russ Miles]]></itunes:name></itunes:owner><itunes:author><![CDATA[Russ Miles]]></itunes:author><googleplay:owner><![CDATA[engineeringagents@substack.com]]></googleplay:owner><googleplay:email><![CDATA[engineeringagents@substack.com]]></googleplay:email><googleplay:author><![CDATA[Russ Miles]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Agent Habitat You Can Install: On Agents, Plugins, Reuse & Harnesses]]></title><description><![CDATA[Introducing the AI Literacy Superpowers plugin]]></description><link>https://engineeringagents.substack.com/p/the-agent-habitat-you-can-install</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/the-agent-habitat-you-can-install</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Fri, 24 Apr 2026 05:30:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!UWSP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc3151a-cb5e-4265-8477-e4fca02a4c3b_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UWSP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc3151a-cb5e-4265-8477-e4fca02a4c3b_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UWSP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc3151a-cb5e-4265-8477-e4fca02a4c3b_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!UWSP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc3151a-cb5e-4265-8477-e4fca02a4c3b_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!UWSP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc3151a-cb5e-4265-8477-e4fca02a4c3b_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!UWSP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc3151a-cb5e-4265-8477-e4fca02a4c3b_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UWSP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc3151a-cb5e-4265-8477-e4fca02a4c3b_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/adc3151a-cb5e-4265-8477-e4fca02a4c3b_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3264156,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/195312428?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc3151a-cb5e-4265-8477-e4fca02a4c3b_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!UWSP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc3151a-cb5e-4265-8477-e4fca02a4c3b_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!UWSP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc3151a-cb5e-4265-8477-e4fca02a4c3b_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!UWSP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc3151a-cb5e-4265-8477-e4fca02a4c3b_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!UWSP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc3151a-cb5e-4265-8477-e4fca02a4c3b_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Le Bon Mot is a regular feature over on a <a href="https://www.softwareenchiridion.com/">Software Enchiridion</a>. But really it exists everywhere as a habitat for ideas and conversation, and so here it appears for the first time in Engineering Agents.</p><div><hr></div><p>There are cafes that serve coffee, and there are cafes that serve as unpacking rooms for deliveries that arrive from nowhere in particular.</p><p>Le Bon Mot was operating, on this particular Wednesday, as the latter. The box appeared on the counter sometime between the second espresso and the third. Nobody saw it arrive. Madame Beauregard regarded it with the expression she reserved for objects that materialised without explanation but with evident purpose &#8212; a category that, at Le Bon Mot, was larger than you might expect.</p><p>It was a wooden crate, roughly the size of a case of wine, stamped with no return address. The lid was loose. Madame Beauregard lifted it and looked inside.</p><p>Tools. Not the kind you buy at a hardware store but the kind you assemble a practice from. There were templates, neatly rolled and tied. Constraint cards, each describing a rule and the mechanism for enforcing it. Extraction guides with questions printed in the kind of type that expects honest answers. A set of agent definitions, each describing a role, a trust boundary, and the conditions under which the role should be invoked. A small booklet titled <em>Hooks: What to Check and When</em>. And at the bottom, a single sheet of paper that read: </p><blockquote><p><strong>Run </strong><code>/superpowers-init</code><strong> and the habitat scaffolds itself.</strong></p></blockquote><p>Case, a retired developer with scar tissue that went back to CORBA and beyond, picked up one of the constraint cards. It read: </p><blockquote><p><em>Coverage below 80% fails the build. Enforcement: deterministic. Timing: merge gate.</em></p></blockquote><p>&#8220;This is a harness component,&#8221; she said. &#8220;Someone has packaged an entire agent habitat into a box.&#8221;</p><p>The Djinn, who had been sitting quietly by the window in the way that intelligences sit when they are processing something they cannot yet articulate, looked up. &#8220;You can <em>install</em> a habitat?&#8221;</p><p>Sophie, resident French Bulldog, raised her head from beside the fire. The brass clock ticked.</p><p>Case turned the card over. On the back, in handwriting that was neither hers nor the Djinn&#8217;s: </p><blockquote><p><em>The constraint is the easy part. The conversation about why this threshold and not another &#8212; that is the habitat.</em></p></blockquote><p>&#8220;You can install the conditions for one,&#8221; Case said. &#8220;The habitat itself, you have to grow.&#8221;</p><p>Madame Beauregard set the lid back on the crate with the care of someone closing a book at exactly the right chapter. &#8220;Then I suggest,&#8221; she said, &#8220;you begin by seeing what is in the box.&#8221;</p><p>The Djinn reached in and pulled out the booklet of hooks. It read the first page, then looked up with an expression that, in its reference frame, served the same function as recognition.</p><p>&#8220;These are the things I have been missing,&#8221; it said quietly. &#8220;Not instructions. Not prompts. <em>Structure</em>.&#8221;</p><div><hr></div><p>The Djinn&#8217;s word, <em>structure</em>, is the one that matters. Because structure is the difference between an AI session that starts from zero every time and an AI session that starts from the accumulated wisdom of your team, your project, and every mistake you have already made.</p><p>The <a href="https://github.com/Habitat-Thinking/ai-literacy-superpowers">AI Literacy Superpowers plugin</a> is that structure, packaged for reuse and, if you prefer, merely as inspiration.</p><p>Not a new model. Not a prompt library. Not a collection of clever tricks. A complete development workflow &#8212; skills, agents, commands, hooks, and templates &#8212; that implements the <a href="https://www.softwareenchiridion.com/p/the-habitat-you-build-is-the-intelligence">AI Literacy framework</a> from Level 2 through Level 5. You install it. You run one command. And the habitat scaffolds itself: living documents, enforceable constraints, a coordinated agent team, CI templates, and the feedback loops that bind them together.</p><p>This article is about what is in the box, why each piece exists, and how to start using it without drowning in the whole thing at once.</p><div><hr></div><h2><strong>The Department Store Problem</strong></h2><p>Firstly, never eat a whole cake in one bite. A mistake that almost everyone makes with a plugin this size: they try to use everything at once.</p><p>Skills. Agents. Commands. Hooks. Templates. That is not just a toolkit, that is a department store. And nobody walks into a department store, buys one of everything, and walks out dressed well. You walk in knowing what you need <em>today</em>. You buy that. You come back when you need the next thing.</p><p>The plugin is designed for exactly this kind of incremental adoption. Every component is useful on its own. No skill requires every other skill. No agent demands the full pipeline. The commands work independently. You start with the piece that solves the problem you have right now, and the rest waits until you are ready.</p><blockquote><p><strong>No Really:</strong> You do not even need to understand everything in this article before you start. Install the plugin. Run <code>/superpowers-init</code>. See what it discovers about your project. That single action teaches you more than reading about it ever will.</p></blockquote><div><hr></div><h2><strong>What Is in the Box</strong></h2><p>The plugin has five categories of components. Each maps to a different kind of work in the development lifecycle.</p><h3><strong>Skills: What the AI Knows</strong></h3><p>Skills are knowledge documents &#8212; structured Markdown files that agents and commands read before doing work. They are not prompts. They are not instructions. They are <em>expertise</em>, encoded in a format that both humans and AI can consume.</p><p>Think of them as the senior engineer who is always available but never talks unless asked.</p><p>The plugin&#8217;s 29 skills fall into several natural clusters. The harness framework is the conceptual spine of the plugin:</p><ul><li><p><code>harness-engineering</code> teaches the framework itself &#8212; deterministic tooling plus LLM agents keeping AI-generated code trustworthy. </p></li></ul><ul><li><p><code>context-engineering</code> curates the knowledge an LLM needs to work in a codebase effectively, with the insight that code design itself is context.</p></li><li><p><code>verification-slots</code> defines the core technical abstraction: every constraint checked through a uniform interface regardless of whether the backing tool is deterministic or agent-based. constraint-design teaches how to design those constraints so they&#8217;re falsifiable and enforceable. </p></li><li><p><code>garbage-collection</code> runs the periodic checks that fight entropy &#8212; stale docs, drifted conventions, dead code. <code>fitness-functions</code> extends this to architectural properties (coupling trends, layer boundaries) drawing on Ford, Parsons, Kua and Sadalage. </p></li><li><p><code>harness-observability</code> provides four layers of measurement at different timescales &#8212; context, constraints, GC, cost. </p></li><li><p><code>harness-onboarding</code> generates a human-readable onboarding document from the harness artefacts for new team members.</p></li></ul><p>AI literacy assessment is the evaluation layer:</p><ul><li><p><code>ai-literacy-assessment</code> assesses a team&#8217;s AI collaboration maturity by combining repository evidence with clarifying questions.</p></li></ul><ul><li><p><code>literacy-improvements</code> translates assessment gaps into a prioritised improvement plan mapped to specific plugin skills and commands. </p></li><li><p><code>portfolio-assessment</code> aggregates assessments across multiple repos into an organisational view</p></li><li><p><code>portfolio-dashboard</code> renders that view as a self-contained HTML dashboard with trend tracking.                                                                       </p></li></ul><p>Code quality covers two complementary lenses for reading and writing code:</p><ul><li><p><code>cupid-code-review</code> applies Daniel Terhorst-North&#8217;s five CUPID properties (Composable, Unix philosophy, Predictable, Idiomatic, Domain-based) as structured review criteria. </p></li><li><p><code>literate-programming</code> applies Don Knuth&#8217;s principle that code is written for humans first &#8212; narrative preambles, reasoning-based documentation, presentation ordered by understanding rather than execution.</p></li></ul><p>Governance gives teams the vocabulary to translate policy language into operational meaning. </p><ol><li><p><code>governance-constraint-design</code> handles falsifiability and evidence requirements.</p></li><li><p><code>governance-audit-practice</code> detects semantic drift and governance debt. </p></li><li><p><code>governance-observability </code>defines the metrics and snapshot formats for measuring governance health over time.                  </p></li></ol><p>Security and supply chain is a set of four scanning skills: </p><ul><li><p><code>secrets-detection</code> (gitleaks, hardening the no-secrets constraint) </p></li><li><p><code>dependency-vulnerability-audit</code> (known CVEs and provenance)</p></li><li><p><code>docker-scout-audit</code> (image SBOMs and base image staleness) </p></li><li><p><code>github-actions-supply-chain</code> (CI pipeline hardening and third-party action risk).                                                     </p></li></ul><p>Convention and tooling covers the practicalities of keeping project rules in sync: </p><ul><li><p><code>convention-extraction</code> surfaces tacit team knowledge into versioned artefacts using systematic guided discovery. convention-sync propagates HARNESS.md rules to Cursor, Copilot, and Windsurf so all AI coding tools share the same conventions. auto-enforcer-action wires constraint enforcement into  GitHub Actions.</p></li></ul><p>Three remaining skills are more distinct:</p><ul><li><p><code>advocatus-diaboli</code> is an adversarial spec reviewer that raises steel-manned objections across six categories before plan approval. </p></li><li><p><code>model-sovereignty</code> covers deliberate decisions about which models to use, where they run, and whether to build custom models. </p></li><li><p><code>cost-tracking</code> captures and records AI tool costs to inform model routing and health snapshots. </p></li><li><p><code>team-api</code> generates a Team Topologies Team API document enriched with portfolio assessment data. </p></li><li><p><code>cross-repo-orchestration</code> coordinates changes across multiple repositories using git-mediated and specification-mediated patterns.</p></li></ul><blockquote><p><strong>The Sceptic asks:</strong> &#8220;Do I need them all?&#8221; No. Pick the two or three that match where you are today. The rest will be there when you arrive.</p></blockquote><div><hr></div><h3><strong>Agents: Who Does the Work</strong></h3><p>Agents are role definitions. They describe a persona, a trust boundary, and the conditions under which that persona should be invoked. They are the development team you can assemble from the plugin.</p><p>The plugin has 12 agents organised around two distinct patterns: a <strong>spec-first development pipeline </strong>and a <strong>harness maintenance cluster</strong>.</p><p><strong>The development pipeline</strong> is an ordered chain of specialists:</p><ul><li><p><code>spec-writer</code> opens every feature &#8212; updating specs, user stories, and acceptance scenarios before any code is touched. </p></li><li><p><code>advocatus-diaboli </code>follows immediately after in spec mode, reading the spec with read-only access and raising objections across six categories; a human writes the dispositions, never the agent. </p></li><li><p>Once the plan is approved, <code>tdd-agent</code> translates acceptance scenarios into failing tests and is responsible solely for the RED phase &#8212; it does not write implementation code. </p></li><li><p><code>code-reviewer</code> enters after tests are green, evaluating the implementation through CUPID and literate programming lenses; it is also read-only and cannot modify files. </p></li><li><p>After the final <code>code-reviewer</code> PASS, <code>advocatus-diaboli </code>runs a second time in code mode, now weighing threat-model, failure-mode, and operational objections against the concrete implementation. </p></li><li><p><code>integration-agent</code> closes the loop &#8212; updating the CHANGELOG, committing, opening a PR, watching CI, merging when green, closing the issue, and pruning the branch.</p></li><li><p><code>orchestrator</code> sits above all of these, coordinating the full pipeline in the correct sequence for any incoming task, reading <code>CLAUDE.md</code>, <code>AGENTS.md</code>, and <code>MODEL_ROUTING.md</code> to calibrate strategy.</p></li></ul><p><strong>The</strong> <strong>harness maintenance cluster</strong> keeps the project&#8217;s living harness document honest and current:</p><ul><li><p><code>harness-discoverer</code> is a read-only scanner that maps the actual tech stack &#8212; linters, CI config, test frameworks, pre-commit hooks &#8212; and feeds that evidence into harness initialisation. </p></li><li><p><code>harness-auditor </code>is a meta-agent that verifies the declarations in <code>HARNESS.md</code> match actual project state, running weekly or on demand to keep the harness from drifting into aspirational fiction. </p></li><li><p><code>harness-enforcer</code> is the verification engine that executes constraints from <code>HARNESS.md</code>, running either deterministic tools or agent-based reviews depending on the constraint type, and consulting recent reflections to calibrate scrutiny. </p></li><li><p><code>harness-gc</code> fights entropy on a periodic schedule &#8212; staleness, dead code, convention drift, dependency currency &#8212; auto-fixing simple issues and creating GitHub issues for those it cannot.</p></li></ul><p><strong>Two</strong> <strong>agents operate outside both clusters:</strong></p><ul><li><p><code>assessor</code> scans the repository for observable AI literacy evidence and produces a timestamped assessment report, invoked via <code>/assess</code> or when someone asks where the team sits on the literacy framework. </p></li><li><p><code>governance-auditor </code>detects semantic drift in governance constraints, inventories governance debt, and checks three-frame alignment &#8212; running on demand or on a quarterly schedule.</p></li></ul><p>Two things to notice. First, every agent has a <em>trust boundary</em>. The <code>spec-writer</code> cannot execute code. The <code>code-reviewer</code> cannot write files. The <code>harness-discoverer</code> can only read. This is not paranoia, it is habitat design. An agent that can do anything will eventually do the wrong thing. Constraints are how you make the collaboration safer.</p><p>Second, the agents split into two groups: the <em>development pipeline</em> (orchestrator through integration-agent) and the <em>harness team</em> (discoverer through assessor). You can use either group independently. A team that is not ready for the full spec-first pipeline can still use the harness agents to build and maintain constraints.</p><blockquote><p><strong>The Pragmatist says:</strong> &#8220;Start with the harness team. Run <code>/harness-init</code> and let the discoverer scan your project. You will learn more about your own conventions in ten minutes than you learned in the last six months.&#8221;</p></blockquote><h3><strong>Commands: What You Can Ask For</strong></h3><p>Commands are the entry points &#8212; the slash commands you type to trigger workflows. Each command coordinates one or more agents and skills to accomplish a specific task.</p><p>The plugin&#8217;s 22 commands fall into four broad areas. <strong>Harness setup and maintenance</strong> is the largest cluster:</p><ul><li><p><code>/superpowers-init</code> bootstraps the entire AI literacy framework from scratch &#8212; <code>CLAUDE.md</code>, <code>HARNESS.md</code>, <code>AGENTS.md</code>, <code>MODEL_ROUTING.md</code>, CI templates, and agent configurations in one pass.</p></li><li><p><code>/harness-init</code> does the narrower job of generating just <code>HARNESS.md</code> for a project that already has some structure.</p></li><li><p><code>/harness-constrain</code> adds or promotes a single constraint, walking through enforcement type selection (deterministic, agent, or unverified).</p></li><li><p><code>/harness-gc</code> manages and runs garbage collection rules &#8212; adding new ones to <code>HARNESS.md</code> or running existing ones on demand with auto-fix options. </p></li><li><p><code>/harness-upgrade</code> handles post-plugin-update adoption, diffing the current <code>HARNESS.md</code> against the latest template and presenting new constraints and rules for selective adoption. </p></li><li><p><code>/harness-onboarding</code> generates a human-readable <code>ONBOARDING.md</code> from the harness artefacts for new team members.</p></li></ul><p><strong>Observation and health</strong> gives teams a continuous picture of their harness and governance state:</p><ul><li><p><code>/superpowers-status</code> is the top-level dashboard &#8212; habitat files, harness health, agent team consistency, and CI configuration in one view. </p></li><li><p><code>/harness-status</code> is the narrower version, showing enforcement ratios and GC rule status. </p></li><li><p><code>/harness-health</code> goes deeper, generating a full snapshot to <code>observability/snapshots/YYYY-MM-DD-snapshot.md</code> with enforcement trends, mutation rates, and learning velocity (supports <code>--deep</code> and <code>--trends</code> flags).</p></li><li><p><code>/harness-audit</code> runs meta-verification, checking whether <code>HARNESS.md</code> declarations match actual project state and updating the <code>Status</code> section with the result. </p></li><li><p><code>/observatory-verify</code> runs a specific 82-signal checklist against the Habitat Observatory&#8217;s expected data signals, reporting <code>PRESENT / PARTIAL / MISSING / NO_OUTPUT</code> for each.</p></li></ul><p><strong>Governance</strong> mirrors the harness cluster but for policy-level concerns:</p><ul><li><p><code>/governance-constrain</code> guides the authoring of a governance constraint by translating policy language into operational meaning with</p><p>three-frame alignment, then appending it to <code>HARNESS.md</code>. </p></li><li><p><code>/governance-audit</code> runs a deep investigation for semantic drift, constraint falsifiability, and governance debt, writing a dated audit report.</p></li><li><p><code>/governance-health</code> shows the summary view &#8212; falsifiability ratio, drift score, debt inventory &#8212; with an optional <code>--dashboard</code> flag for an HTML view.</p></li></ul><p><strong>AI literacy assessment</strong> is the evaluation layer:</p><ul><li><p><code>/assess</code> scans the repo for observable evidence, asks clarifying questions, and produces a timestamped assessment document with habitat fixes and workflow recommendations.</p></li><li><p><code>/portfolio-assess</code> aggregates assessments across multiple repositories into an organisational view, requiring a <code>--local</code>, <code>--org</code>, or <code>--topic</code> scope flag.</p></li></ul><p><strong>The development workflow</strong> commands support the spec-first pipeline:</p><ul><li><p><code>/diaboli</code> dispatches the adversarial reviewer against a spec or implementation, writing the structured objection record to <code>docs/superpowers/objections/</code>. </p></li><li><p><code>/reflect</code> captures post-task learning &#8212; what was surprising, what future agents should know &#8212; appending a timestamped entry to <code>REFLECTION_LOG.md</code>.</p></li></ul><p><strong>Four utility commands</strong> round out the set:</p><ul><li><p><code>/extract-conventions</code> runs a guided five-question discovery session that surfaces tacit team knowledge into <code>CLAUDE.md</code> and <code>HARNESS.md</code>. </p></li><li><p><code>/convention-sync</code> propagates <code>HARNESS.md</code> rules to Cursor, Copilot, and Windsurf convention files so all AI coding tools share the same constraints. </p></li><li><p><code>/cost-capture</code> records AI tool spend and token usage to the observability directory.</p></li><li><p><code>/worktree</code> manages git worktrees for parallel agent isolation &#8212; spin up, merge back, or clean up.</p></li></ul><blockquote><p><strong>The Veteran observes:</strong> &#8220;Notice that six of the twelve commands are harness commands. That tells you where the framework puts its weight. The harness is not a feature of the plugin. The harness <em>is</em> the plugin. Everything else flows from it.&#8221;</p></blockquote><div><hr></div><h3><strong>Hooks: What Runs When Automatically</strong></h3><p>Hooks are the invisible enforcement layer. They run without being asked, at specific moments in your workflow, and they catch things before they become problems. The plugin has 10 hooks across three event points &#8212; SessionStart, PreToolUse, and Stop &#8212; and every single one is advisory only. None block.</p><ul><li><p><strong>SessionStart (1 hook):</strong> <code>template-currency-check.sh</code> runs when a session opens and compares the <code>HARNESS.md</code> template version against the current plugin.json version. If they differ and the mismatch hasn&#8217;t been dismissed, it nudges you to run <code>/harness-upgrade</code>. This is the upgrade discovery mechanism &#8212; without it you&#8217;d only notice template drift when you stumbled across it.</p></li><li><p><strong>PreToolUse on Write/Edit (2 hooks):</strong> Two hooks fire before any file write or edit. The first is a <code>prompt-type</code> hook &#8212; not a script &#8212; that reads <code>HARNESS.md</code> and evaluates commit-scoped constraints against the file being written, warning on violations without blocking. The second is <code>markdownlint-check.sh</code>, which runs markdownlint against any .md file being written and reports violations as a system message. Both are advisory.</p></li><li><p><strong>Stop (7 hooks):</strong> The bulk of the hooks fire at session end, closing feedback loops before the session closes. <code>drift-check.sh</code> looks for modified CI workflows, linter configs, hook configs, or dependency manifests and prompts <code>/harness-audit</code> if the harness may be stale. <code>snapshot-staleness-check.sh</code> checks whether the latest harness health snapshot is older than 30 days and prompts <code>/harness-health</code> if so. <code>reflection-prompt.sh</code> detects whether commits were made in the last 4 hours and prompts <code>/reflect</code> to capture learnings. <code>secrets-check.sh</code> runs <code>gitleaks</code> (if installed and the <code>no-secrets constraint</code> is active in <code>HARNESS.md</code>) and surfaces any findings. <code>gc-rotate.sh</code> cycles</p><p>through four deterministic GC checks by day-of-year &#8212; secret scanner, snapshot staleness, shell syntax errors, and missing set <code>-euo pipefail</code> &#8212; so a different check runs each day without overwhelming any single session. <code>curation-nudge.sh</code> counts unpromoted <code>REFLECTION_LOG</code> entries against <code>AGENTS.md</code> and nudges curation if more than two reflections haven&#8217;t been promoted to <code>ARCH_DECISIONs</code>. <code>governance-drift-check.sh</code> watches for governance constraint modifications or a governance audit older than 90 days and prompts <code>/governance-audit</code> or <code>/governance-health</code> accordingly.</p></li></ul><p>The design principle throughout is consistent: hooks observe and surface signals, humans and agents decide what to do with them.</p><div><hr></div><h3><strong>Templates: What Gets Scaffolded</strong></h3><p>Templates are the opinionated defaults that <code>/superpowers-init</code> generates. </p><p>They include <code>CLAUDE.md</code> (your project instructions), <code>HARNESS.md</code> (your living constraint document), <code>AGENTS.md</code> (compound learning memory), <code>MODEL_ROUTING.md</code> (cost-conscious model selection), <code>REFLECTION_LOG.md</code>, three CI workflow templates, and a health badge icon.</p><p>You are not locked into the defaults. Every template is designed to be edited, extended, and made your own. The templates give you a starting point that embodies the framework&#8217;s best practices. Your job is to make them true for <em>your</em> project.</p><div><hr></div><h2><strong>The Store Map</strong></h2><p>Here is the question that the department store metaphor answers: where do you start?</p><p>The framework defines six levels. The plugin covers Levels 2 through 5. Here is what unlocks at each:</p><ul><li><p><strong>Level 0 &#8212; Awareness </strong>&#8212;<strong> </strong>The team knows AI coding tools exist and has a repository. No structured usage yet. The plugin has no specific commands for this level &#8212; the work is awareness-building. The floor is simply: the repo exists and the team is aware of AI tools.</p></li><li><p><strong>Level 1 &#8212; Prompting </strong>&#8212; Developers are using AI tools but informally &#8212; prompt-and-accept, no systematic verification. Output is trusted if it looks right. No shared conventions, no CI, no cost awareness. Again no specific plugin commands target this level; the gap from L1 to L2 is primarily about building a CI pipeline, which is outside plugin scope.</p></li><li><p><strong>Level 2 &#8212; Verification </strong>&#8212;<strong> </strong>The team verifies AI output systematically rather than trusting appearances. The minimum evidence is automated tests in CI. The plugin unlocks:</p><ul><li><p>auto-enforcer-action skill &#8212; linting enforcement in CI</p></li><li><p>dependency-vulnerability-audit skill &#8212; CVE and supply chain scanning</p></li><li><p>secrets-detection skill &#8212; gitleaks in CI or pre-commit</p></li><li><p>docker-scout-audit skill &#8212; image scanning (if the project uses Docker)</p></li></ul></li><li><p><strong>Level 3 &#8212; Habitat Engineering </strong>&#8212; The team engineers the environment in which AI works &#8212; conventions, constraints, memory, feedback loops. Minimum evidence: <code>CLAUDE.md</code> exists, at least 3 <code>HARNESS.md</code> constraints are enforced, and there are custom agents or skills. The plugin unlocks the bulk of its commands at this level:</p><ul><li><p><code>/harness-init</code> &#8212; generates HARNESS.md and CLAUDE.md with context and enforced constraints</p></li><li><p><code>/harness-constrain</code> &#8212; adds and promotes individual constraints</p></li><li><p><code>/extract-conventions</code> &#8212; surfaces tacit conventions into HARNESS.md</p></li><li><p><code>/reflect</code> &#8212; starts building REFLECTION_LOG.md</p></li><li><p><code>/harness-gc</code> &#8212; adds garbage collection rules</p></li><li><p><code>/harness-health</code> &#8212; first observability snapshot</p></li><li><p>The full hook set becomes active (drift detection, reflection prompts, snapshot staleness, curation nudge)</p></li></ul></li><li><p><strong>Level 4 &#8212; Specification Architecture </strong>&#8212; The team designs before building &#8212; specs before code, agent pipelines with safety gates. Minimum evidence: a specs directory, an orchestrator agent, and a plan approval gate. The plugin unlocks:</p><ul><li><p><code>/superpowers-init</code> &#8212; bootstraps the full agent team including the orchestrator</p></li><li><p><code>/convention-sync</code> &#8212; propagates HARNESS.md rules to Cursor, Copilot, and Windsurf</p></li><li><p><code>/diaboli</code> &#8212; adversarial spec review before plan approval and before merge</p></li><li><p>fitness-functions skill + <code>/harness-gc</code> &#8212; architectural fitness functions as GC rules</p></li><li><p>constraint-design skill &#8212; loop guardrails and safety gates in the pipeline</p></li></ul></li><li><p><strong>Level 5 &#8212; Sovereign Engineering </strong>&#8212; The team operates at platform scale &#8212; reusable standards, cross-team governance, observability exported to organisational dashboards. Minimum evidence: a published plugin or reusable template, <code>MODEL_ROUTING.md</code>, and organisational governance documentation. The plugin unlocks:</p><ul><li><p>cross-repo-orchestration skill &#8212; coordinates changes across multiple repositories</p></li><li><p>model-sovereignty skill &#8212; deliberate model routing and cost tracking at platform level</p></li><li><p><code>/portfolio-assess</code> and <code>/portfolio-dashboard</code> &#8212; aggregate literacy assessment across the whole organisation</p></li><li><p>harness-observability skill (telemetry layer) &#8212; OTel export and organisational dashboards</p></li><li><p>team-api skill &#8212; Team Topologies Team API enriched with portfolio data</p></li><li><p>governance-audit-practice, governance-constraint-design, governance-observability skills &#8212; the full governance cluster</p></li></ul></li></ul><p>The scoring ceiling rule is worth noting: the assessed level is the highest where the team has substantial evidence <strong>across all three disciplines</strong> (context engineering, architectural constraints, guardrail design). A team with L3 habitat engineering but L1 verification is assessed at L1 &#8212; the weakest discipline sets the ceiling.</p><p><strong>Ponder time:</strong> Ask yourself which level describes your current practice. Not the level you aspire to --- the level you actually operate at today. That is where you start. Install the components for that level. Use them until they become habitual. Then look at the next level.</p><blockquote><p><strong>Exercise:</strong> Install the plugin. Run <code>/superpowers-init</code>. When it asks about your conventions, answer honestly --- not with what you wish were true, but with what actually happens on your team. Then run <code>/harness-status</code> and read the output. What did the habitat discover about your project that you did not know?</p></blockquote><div><hr></div><h2><strong>What the Plugin Is Not</strong></h2><p><strong>The Sceptic returns:</strong> &#8220;Is this just another opinionated starter template?&#8221;</p><p>No. A starter template gives you files. The plugin gives you files <em>and the machinery to keep them honest</em>. The hooks enforce constraints in real time. The harness-auditor checks whether your declared constraints match reality. The garbage collection agent fights the slow entropy of conventions that drift from practice. The reflection pipeline captures what you learn so the next session starts smarter.</p><p>A template is a snapshot. The plugin is a living system.</p><p><strong>The Pragmatist adds:</strong> &#8220;And it discovers. That is the part people do not expect. <code>/superpowers-init</code> does not just scaffold files from a template. It scans your project first &#8212; your stack, your linters, your CI, your test frameworks. The habitat it builds is shaped by what is already there.&#8221;</p><div><hr></div><p>When is a tool more than a tool? When it is a harness. A linter is a tool: it checks syntax and reports violations. A framework is a tool: it provides structure and you fill in the blanks. An IDE is a tool: it helps you write code faster.</p><p>The AI Literacy Superpowers plugin is none of these things, and it is all of them, and that is what makes it difficult to categorise and easy to underestimate.</p><p>It is the material from which a habitat is built. Not the habitat itself. The habitat is what grows when the materials meet a team (human and agent), a codebase, a set of conventions that someone finally writes down, and the daily discipline of enforcing them. The plugin gives you the materials. The discipline is yours.</p><blockquote><p><em>A tool does what you tell it. A habitat shapes what you do.</em></p></blockquote><p>That distinction is the entire framework in one sentence. A tool waits for instructions. A habitat provides structure before you ask for it &#8212; conventions read before code is written, constraints checked before code is committed, reflections captured before context is lost. The environment acts with <em>you</em>, not against you.</p><p>This is why the Djinn reached into the box and said <em>structure</em>, not <em>features</em>. Features are what tools have. Structure is what habitats provide. A harness is like a trellis to grow on. And the difference between a developer who uses AI and a developer who collaborates with AI is exactly the difference between reaching for a tool and inhabiting a structure.</p><p>The box is open. The materials are there. The question, perhaps the only question that has ever mattered in habitat engineering, is whether you will build the room or keep working in the hallway.</p><div><hr></div><p><strong>Give it a spin.</strong> Install the <a href="https://github.com/russmiles/ai-literacy-superpowers">AI Literacy Superpowers</a> plugin, run <code>/superpowers-init</code> on a project you care about, and see what the habitat discovers. If something surprises you, delights you, or breaks horribly &#8212; open an issue or start a discussion on the <a href="https://github.com/russmiles/ai-literacy-superpowers">GitHub repo</a>. The plugin grows from feedback, and the best feedback comes from real and diverse projects.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Engineering Agents is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[The Harness Assessment Practice]]></title><description><![CDATA[Why knowing where you are beats knowing where you want to be]]></description><link>https://engineeringagents.substack.com/p/the-harness-assessment-practice</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/the-harness-assessment-practice</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Tue, 21 Apr 2026 04:59:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!eCIG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb09572cb-cd14-4f5a-8365-d3f6d4bdf825_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eCIG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb09572cb-cd14-4f5a-8365-d3f6d4bdf825_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eCIG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb09572cb-cd14-4f5a-8365-d3f6d4bdf825_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!eCIG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb09572cb-cd14-4f5a-8365-d3f6d4bdf825_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!eCIG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb09572cb-cd14-4f5a-8365-d3f6d4bdf825_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!eCIG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb09572cb-cd14-4f5a-8365-d3f6d4bdf825_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eCIG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb09572cb-cd14-4f5a-8365-d3f6d4bdf825_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b09572cb-cd14-4f5a-8365-d3f6d4bdf825_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3987267,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/194841874?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb09572cb-cd14-4f5a-8365-d3f6d4bdf825_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!eCIG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb09572cb-cd14-4f5a-8365-d3f6d4bdf825_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!eCIG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb09572cb-cd14-4f5a-8365-d3f6d4bdf825_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!eCIG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb09572cb-cd14-4f5a-8365-d3f6d4bdf825_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!eCIG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb09572cb-cd14-4f5a-8365-d3f6d4bdf825_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This is the first companion to the Habitat Hypothesis series that started with:</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;bd9a8881-b7f1-45f9-b5b2-50092004120e&quot;,&quot;caption&quot;:&quot;This is Part 1 of a 6 part series on Harness Engineering as part of Habitat Thinking to accompany the open source AI Literacy Superpowers plugin&quot;,&quot;cta&quot;:&quot;Read full story&quot;,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;The Habitat Hypothesis&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:9890843,&quot;name&quot;:&quot;Russ Miles&quot;,&quot;bio&quot;:&quot;Software Builder, Listener, Reader, Writer, Speaker (in that order)&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/40b16e11-931e-4f92-9f20-adc467242b3b_1280x1284.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-04-08T10:24:24.520Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!WBXw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3893301d-0049-4765-842d-a02b33e68268_1536x1024.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://engineeringagents.substack.com/p/the-habitat-hypothesis&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:193558821,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:4,&quot;comment_count&quot;:0,&quot;publication_id&quot;:5479520,&quot;publication_name&quot;:&quot;Engineering Agents&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!eRFm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0861dada-274e-4eb5-8137-e122bd7a4c3d_1024x1024.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div><hr></div><p>It&#8217;s been six months. Your team adopted AI coding tools back in October. You wrote a <code>CLAUDE.md</code>. You set up some constraints. You even have a <code>REFLECTION_LOG.md</code> with a few entries in it. Things feel good. The AI is producing better code than it was at the start. PRs are moving faster. Everyone agrees: you&#8217;re getting good at this.</p><p>Then someone &#8212; maybe a new tech lead, maybe a curious VP, maybe that one developer who asks uncomfortable questions &#8212; says:</p><p>&#8220;Are we actually good at this, or are we just busy?&#8221;</p><p>Silence.</p><p>The kind of silence where everyone looks at each other and realises they have no idea. You <em>feel</em> like you&#8217;ve improved. But feelings are not evidence. You couldn&#8217;t point to a metric, a level, a concrete description of where you are versus where you were three months ago.</p><p>This is the <strong>Dunning-Kruger problem of AI literacy</strong>. Teams at Level 1 think they&#8217;re at Level 3. They&#8217;re using the tools daily, surely that counts for something? It does count. It counts for Level 1. Daily use without systematic verification, without environment engineering, without feedback loops is just... daily use.</p><p><strong>Assessment is not an audit. It is a practice.</strong> A recurring discipline &#8212; quarterly, deliberate, evidence-based &#8212; that tells you where you are, what&#8217;s working, and what to do next. Where you <em>actually</em> are, not where your conference talk claims you are.</p><p>Roadmaps tell you where to go. Assessment tells you where to go <em>from</em>. Without that, every plan is just built on belief on where we are, and we are the easiest people to fool.</p><div><hr></div><h2><strong>Which of These Sounds Like You?</strong></h2><p>There are six levels of AI literacy. Not six levels you climb like a career ladder, but six diagnostic positions that describe how your team currently operates.</p><p>Read these. Not as a checklist of things to aspire to. As a mirror. Be honest with yourself about which one you recognise:</p><ul><li><p><strong>Level 0 &#8212; Awareness.</strong> You know AI coding tools exist. You haven&#8217;t used one on real work with real stakes.</p></li><li><p><strong>Level 1 &#8212; Prompting.</strong> You use AI tools daily. You&#8217;ve got your favourite model. You copy-paste output into your codebase and fix it up manually. Sometimes it&#8217;s great. Sometimes you spend an hour debugging code that looked right. You don&#8217;t have a systematic way of telling which is which, the feedback loop is &#8220;find the bug.&#8221;</p></li><li><p><strong>Level 2 &#8212; Verification.</strong> You&#8217;ve been burned enough to stop trusting. You have CI workflows that run tests. Linters. Maybe vulnerability scanning. When the AI produces code, it goes through the same gauntlet as human code. But the AI still doesn&#8217;t know your conventions, your architecture, your hard-won decisions. You verify systematically, and the code works &#8212; it just doesn&#8217;t feel like <em>your</em> code. You correct the same things every session. The AI never learns from yesterday.</p></li><li><p><strong>Level 3 &#8212; Habitat Engineering.</strong> You have a <code>CLAUDE.md</code> or equivalent that gives the AI your conventions at session start. A <code>HARNESS.md</code> declaring your constraints and at least some of them are actually enforced in CI, not just written down. Reflections are captured. Sessions improve over time because the <em>environment</em> accumulates knowledge, not just the people. This is where the flywheel starts turning. Most teams that reach Level 3 can feel the difference but struggle to articulate it.</p></li><li><p><strong>Level 4 &#8212; Specification Architecture.</strong> Specs before code, agent pipelines with safety gates, the AI as a colleague with defined responsibilities as opposed to being a typewriter you occasionally shout at.</p></li><li><p><strong>Level 5 &#8212; Sovereign Engineering.</strong> Reusable plugins, cross-team templates, cost tracking, model routing, organisational governance.</p></li></ul><div><hr></div><p>These are <em>diagnostic</em> positions, not aspirational destinations. You&#8217;re already at one of them. Right now. The question isn&#8217;t &#8220;which one do I want to be?&#8221; It&#8217;s &#8220;which one am I?&#8221;</p><blockquote><p><strong>Brain check:</strong> Which level did you recognise yourself in? Did you hesitate between two? That hesitation is the whole point. Assessment exists to resolve it.</p></blockquote><div><hr></div><h2><strong>Files, Not Feelings</strong></h2><p>You probably think you know your level. Let&#8217;s test that.</p><blockquote><p><strong>The Sceptic:</strong> &#8220;We&#8217;re Level 3. We&#8217;ve got a CLAUDE.md, we&#8217;ve got constraints, we&#8217;ve got the whole setup.&#8221;</p><p><strong>The Pragmatist:</strong> &#8220;When did you last update the CLAUDE.md?&#8221;</p><p><strong>The Sceptic:</strong> &#8220;I mean... it&#8217;s there.&#8221;</p><p><strong>The Pragmatist:</strong> &#8220;How many of your constraints are enforced in CI versus just written down? And do you know why?&#8221;</p><p><strong>The Sceptic:</strong> &#8220;...&#8221;</p><p><strong>The Pragmatist:</strong> &#8220;And your REFLECTION_LOG.md -- when&#8217;s the last entry?&#8221;</p><p><strong>The Sceptic:</strong> &#8220;February, probably.&#8221;</p><p><strong>The Pragmatist:</strong> &#8220;It&#8217;s April. That&#8217;s not Level 3. That&#8217;s Level 1 wearing a Level 3 costume.&#8221;</p></blockquote><p>This is why assessment uses <strong>observable evidence</strong>, not self-report. An assessment scans your repository for concrete signals &#8212; files that exist or don&#8217;t, configurations that are active or stale, workflows that run or sit disabled.</p><p>At Level 2, it looks for CI workflows, test coverage thresholds, vulnerability scanning. At Level 3, it checks whether your <code>CLAUDE.md</code> is current, whether <code>HARNESS.md</code> constraints are actually enforced, whether <code>AGENTS.md</code> has entries, whether <code>REFLECTION_LOG.md</code> has recent dates. At Level 4, it looks for specification files, plans, orchestrators with safety gates. At Level 5, plugin structures, cross-team templates, observability configuration.</p><p>Most teams are one level lower than they think. A <code>CLAUDE.md</code> that hasn&#8217;t been updated in two months isn&#8217;t evidence of engineering a habitat for the AI cognition in the room. It&#8217;s evidence that you <em>tried</em> habitat engineering and then stopped.</p><p>The scoring heuristic makes this explicit: <strong>your weakest discipline is your ceiling.</strong> You might have brilliant context engineering at Level 3 but verification stuck at Level 1. Your assessed level? Level 1. The chain breaks at the weakest link, because that weakest link is where the AI will hurt you.</p><blockquote><p><strong>Brain check:</strong> Open your repo right now. Is there a file in there that declares a practice you&#8217;ve stopped doing? A reflection log with no recent entries? A HARNESS.md constraints file that hasn&#8217;t been touched since you created it? That gap between declaration and practice, that&#8217;s exactly what assessment measures.</p></blockquote><div><hr></div><h2><strong>The Practice: Assessment as Quarterly Discipline</strong></h2><p>You wouldn&#8217;t step on the scales in January and assume the number holds through December. So why would you assess your AI literacy once and call it done?</p><p>Assessment is a <strong>recurring practice</strong>. At least quarterly. Here&#8217;s the rhythm:</p><ol><li><p><strong>Scan.</strong> Read the repository for evidence. Every signal found, every signal absent. Not opinions but files, configurations, dates, commit history.</p></li><li><p><strong>Question.</strong> Three to five clarifying questions that fill the gaps observable evidence can&#8217;t answer. &#8220;Do you write specs before code, or after?&#8221; &#8220;Do you verify AI output systematically, or trust it if it looks right?&#8221; &#8220;Does your team share AI conventions, or does each developer work differently?&#8221;</p></li><li><p><strong>Assess.</strong> Evidence maps to levels. The weakest discipline sets the ceiling. No negotiation, no grading on a curve.</p></li></ol><p>Then the assessment acts on what it found. A timestamped document lands in your repo &#8212; evidence, rationale, strengths, gaps, recommendations. Habitat hygiene gets fixed on the spot: stale counts, missing entries, drift between declared and actual state. Workflow changes get proposed one at a time, accepted or rejected, applied immediately.</p><p>Then the bridge: the assessment asks how far you want to improve. Next level, or higher? It maps each gap to the specific command or skill that closes it, ordered by priority. You walk through them one at a time: accept, skip, or defer. The ones you accept execute immediately. The ones you defer show up in the next assessment as unfinished business. No prose recommendations that rot in a markdown file. Executable actions, connected to the tools that do the work.</p><p>And the assessment captures a reflection on itself, feeding the learning loop.</p><blockquote><p><strong>The Veteran:</strong> &#8220;First assessment, we thought we were Level 3. We were Level 2. Verification was solid but our habitat was stale. Second assessment, three months later: Level 3. Not because we built anything new. Because we started <em>operating</em> what we already had. The <code>CLAUDE.md</code> got updated fortnightly. The constraints got enforced. The reflections got curated. Same infrastructure, different discipline.&#8221;</p></blockquote><p><strong>Each assessment raises the floor for the next one.</strong> Recommendations from Q1 become evidence in Q2. The adjustments you make in April are the signals the scan finds in July. It compounds. Not dramatically but incrementally.</p><p>Let me say it differently: the team that runs four assessments a year doesn&#8217;t improve four times. They improve <em>continuously</em>, because each assessment changes the daily operating habits that produce the evidence the next assessment measures. The assessment isn&#8217;t the improvement. It&#8217;s the thing that <em>triggers</em> and <em>maps </em>improvement.</p><p>If you&#8217;ve read the earlier series on the Habitat Hypothesis, you&#8217;ll recognise something: the three feedback loops &#8212; edit time, merge time, periodic. Assessment is the deepest cycle of that periodic loop. It&#8217;s the moment you step back and ask not &#8220;is this session going well?&#8221; but &#8220;is our <em>practice</em> going well?&#8221;</p><blockquote><p><strong>Exercise:</strong> Before you read the next section, write down &#8212; on paper, in a note, wherever &#8212; what level you think your team is at. Be specific. Then write down the last time you updated your AI environment files. If there&#8217;s a gap between the level you claimed and the freshness of your evidence, you&#8217;ve just done a mini assessment. Any discomfort? That&#8217;s the useful part.</p></blockquote><div><hr></div><h2><strong>Why This Isn&#8217;t a Retrospective</strong></h2><p>You&#8217;ve done retrospectives. You&#8217;ve written action items on sticky notes and then not done them. Assessment is different in one critical way: <strong>it applies changes in the same session.</strong></p><p>A constraint sitting at &#8220;unverified&#8221; for two months gets promoted to agent-backed enforcement before the session ends. An <code>AGENTS.md</code> that exists but nobody reads gets wired into the workflow before the session ends. A <code>HARNESS.md</code> status section showing twelve constraints when there are actually fifteen gets corrected before the session ends.</p><p>The gap between &#8220;we should&#8221; and &#8220;we did&#8221; collapses to zero. That&#8217;s not a minor detail. That&#8217;s the entire reason it works.</p><p>After the fixes, a badge lands in your README &#8212; your assessed level, linking to the full assessment document with every piece of evidence and every recommendation.</p><blockquote><p><strong>The Sceptic:</strong> &#8220;So the badge is just for show?&#8221;</p><p><strong>The Pragmatist:</strong> &#8220;Click it. Every piece of evidence, every gap, every recommendation. It&#8217;s a claim with receipts.&#8221;</p></blockquote><div><hr></div><h2><strong>Beyond One Repo</strong></h2><p>You&#8217;ve assessed your repo. You know your level. You have an improvement plan. And then the question you were always going to ask:</p><p>&#8220;What about the other twelve repos?&#8221;</p><p>Assessment scales. You can point it at a GitHub organisation, a set of topic-tagged repos, or a directory full of clones, and it aggregates everything into a portfolio view. Each repo gets a row: its level, when it was last assessed, its discipline scores. Repos that have never been assessed get a lightweight scan &#8212; the tool checks for key files via the GitHub API and estimates a level from observable signals. No clarifying questions, no full assessment, but enough to place the repo on the map.</p><p>The portfolio view surfaces three things you cannot see from inside a single repo.</p><ul><li><p><strong>Shared gaps.</strong> When five out of eight repos have no reflection practice, that&#8217;s not eight repo problems. That&#8217;s one organisational problem. Roll out a reflection cadence template once and five repos move toward Level 3.</p></li><li><p><strong>Outliers.</strong> One repo at Level 4 while the rest sit at Level 2. What are they doing that nobody else has adopted? One repo stuck at Level 1 while everything else has moved on. Who needs help?</p></li><li><p><strong>Stale assessments.</strong> A repo that hasn&#8217;t been assessed in six months is a repo where you&#8217;re guessing. The portfolio flags it.</p></li></ul><p>The improvement plan works the same way but grouped by impact. Actions that lift 50% or more of your repos get highest priority. Organisation-wide changes first, cluster-specific second, individual repo issues last.</p><blockquote><p><strong>The Sceptic:</strong> &#8220;This sounds like a management dashboard.&#8221;</p><p><strong>The Pragmatist:</strong> &#8220;It&#8217;s a management dashboard backed by files in repos, not survey responses. Every number links to evidence. Try arguing with a <code>HARNESS.md</code> that hasn&#8217;t been updated since January.&#8221;</p></blockquote><div><hr></div><p>The first series in this collection argued that <strong>the environment is the product.</strong> This article adds: you have to know the state of the product -- not just one repo, but the whole portfolio.</p><p>Assessment is how you know. For one repo, <code>/assess</code>. For many, <code>/portfolio-assess</code>. Quarterly. As a discipline, not a chore.</p><p>All available for reuse or inspiration in the <a href="https://github.com/Habitat-Thinking/ai-literacy-superpowers">AI Literacy Superpowers Plugin</a>.</p><p>Fifteen minutes for one repo. An hour for twelve. That&#8217;s enough to start.</p><div><hr></div><h2><strong>A summary for the humans in the room</strong></h2><ul><li><p><strong>Six Literacy levels (L0-L5)</strong> &#8212; diagnostic positions from awareness through sovereign engineering, each with concrete indicators</p></li><li><p><strong>The Dunning-Kruger of AI literacy</strong> &#8212; teams at Level 1 think they&#8217;re at Level 3 because daily use feels like mastery</p></li><li><p><strong>Files, not feelings</strong> &#8212; assessment uses observable evidence from the repository, not self-reported confidence</p></li><li><p><strong>Assessment is a practice, not an audit</strong> &#8212; quarterly (or more), deliberate, evidence-based, recurring</p></li><li><p><strong>Weakest discipline is the ceiling</strong> &#8212; brilliant context engineering means nothing if verification is absent</p></li><li><p><strong>Immediate adjustments in the same session</strong> &#8212; no action items for next sprint; fixes happen now</p></li><li><p><strong>Assessment feeds the learning loop</strong> &#8212; each assessment&#8217;s reflection becomes raw material for the next cycle of improvement</p></li><li><p><strong>Portfolio assessment scales it</strong> &#8212; aggregate across repos by org or topic; shared gaps reveal organisational problems, not repo problems</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Engineering Agents is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[The Looping Habitat]]></title><description><![CDATA[How your AI gets smarter every session]]></description><link>https://engineeringagents.substack.com/p/the-looping-habitat</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/the-looping-habitat</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Tue, 14 Apr 2026 13:08:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!b9Zj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3c2c21a-0e8f-49e2-8591-ddd391500169_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!b9Zj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3c2c21a-0e8f-49e2-8591-ddd391500169_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!b9Zj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3c2c21a-0e8f-49e2-8591-ddd391500169_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!b9Zj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3c2c21a-0e8f-49e2-8591-ddd391500169_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!b9Zj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3c2c21a-0e8f-49e2-8591-ddd391500169_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!b9Zj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3c2c21a-0e8f-49e2-8591-ddd391500169_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!b9Zj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3c2c21a-0e8f-49e2-8591-ddd391500169_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a3c2c21a-0e8f-49e2-8591-ddd391500169_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3524930,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/194094539?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3c2c21a-0e8f-49e2-8591-ddd391500169_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!b9Zj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3c2c21a-0e8f-49e2-8591-ddd391500169_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!b9Zj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3c2c21a-0e8f-49e2-8591-ddd391500169_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!b9Zj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3c2c21a-0e8f-49e2-8591-ddd391500169_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!b9Zj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3c2c21a-0e8f-49e2-8591-ddd391500169_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>This is Article 6 of 6 of &#8220;The Habitat Hypothesis,&#8221; a six-part series on building environments where AI actually produces great work. <a href="https://engineeringagents.substack.com/p/the-habitat-hypothesis">Article 1: environment determines output quality</a>. <a href="https://engineeringagents.substack.com/p/engineering-the-context">Article 2: context engineering</a>. <a href="https://engineeringagents.substack.com/p/constraints-that-bite">Article 3: constraints</a>. <a href="https://engineeringagents.substack.com/p/the-entropy-problem">Article 4: entropy</a>. <a href="https://engineeringagents.substack.com/p/agents-as-colleagues">Article 5: agent orchestration</a>. This is the finale. The piece that was missing from all of them.</em></p><div><hr></div><p>You fixed this already.</p><p>Last Thursday. The AI generated a service class that swallowed exceptions silently &#8212; caught them, logged nothing, returned null. You caught it in review. You explained the pattern. You showed it your team&#8217;s error handling convention. It produced correct code for the rest of the session.</p><p>Monday morning. New session. The AI generates a service class that swallows exceptions silently.</p><p>Not irritation. Something worse: <em>resignation</em>.</p><p>You&#8217;ve been here before. Not just with error handling. With naming conventions. With test structure. With that one architectural boundary the agent keeps crossing. Every session, you teach the same lessons. Every session, the AI has forgotten them. You are Bill Murray in <em>Groundhog Day</em>, except the alarm clock is a code review full of the same mistakes you corrected yesterday.</p><p><strong>This is the Groundhog Day problem.</strong> And if you&#8217;ve been following this series, you already have most of the solution.</p><div><hr></div><h2><strong>Where we&#8217;ve been</strong></h2><blockquote><p><strong>Series recap:</strong></p><p><strong>Article 1:</strong> The environment, not the model, determines AI output quality. <strong>Article 2:</strong> Context engineering gives the AI the knowledge your senior developers carry in their heads. <br><strong>Article 3:</strong> Constraints enforce standards at three maturity levels -- declared, agent-backed, and deterministic. <br><strong>Article 4:</strong> Entropy is the natural tendency of codebases toward disorder. You fight it with garbage collection. <br><strong>Article 5:</strong> Agent orchestration lets specialised AI agents collaborate through a structured pipeline.</p></blockquote><p>Everything we&#8217;ve built so far is <em>static</em>. You set up context files. You define constraints. You deploy agents. It works dramatically better than the bare-repo, hope-for-the-best approach most teams use.</p><p>But none of it learns.</p><p>Your context files contain what you knew when you wrote them. Your constraints cover violations you&#8217;ve already seen. Your GC rules target entropy patterns you&#8217;ve already noticed. The system is exactly as smart as you were on the day you configured it.</p><p>That&#8217;s the ceiling. Unless you build the mechanism that raises it.</p><div><hr></div><h2><strong>The Raw Material Is Already There</strong></h2><p>Think about what happens during an agentic coding session. Not the code it produces, the <em>corrections</em> you make. Every time you say &#8220;no, not like that, like this,&#8221; you&#8217;re generating a signal. Every AI surprise, good or bad, is data. Every edge case you discover is knowledge that didn&#8217;t exist before the session started.</p><p>Right now, all of that evaporates. Session ends. Corrections disappear. Edge cases live only in your memory, and your memory is not as good as you think it is.</p><p><strong>Reflections</strong> change this. A reflection is a brief note captured after a piece of work. Three sentences, not an essay:</p><ul><li><p><strong>What was surprising?</strong> Edge cases you didn&#8217;t anticipate. Assumptions that broke.</p></li><li><p><strong>What should future sessions know?</strong> A gotcha, a &#8220;don&#8217;t do it this way because...&#8221; warning.</p></li><li><p><strong>What could improve?</strong> A missing convention, a constraint that&#8217;s too loose, a gap in context.</p></li></ul><p>There&#8217;s one more piece: <strong>what kind of signal is this?</strong> Each reflection classifies itself as one of four types, a taxonomy borrowed from Birgitta Boeckeler&#8217;s <a href="https://martinfowler.com/articles/reduce-friction-ai/feedback-flywheel.html">Feedback Flywheel</a>:</p><ul><li><p><strong>Context</strong> &#8212; a gap in the priming document. A missing convention, an outdated version, incomplete domain knowledge.</p></li><li><p><strong>Instruction</strong> &#8212; a prompt or command that produced notably better or worse results.</p></li><li><p><strong>Workflow</strong> &#8212; a process pattern that reliably succeeded or failed.</p></li><li><p><strong>Failure</strong> &#8212; a preventable error. A check that should have run, a tool that was misconfigured.</p></li></ul><p>The classification takes five seconds and answers the question curation always asks: <em>where should this go?</em> Context signals route to your conventions file. Instruction signals improve your prompts. Workflow signals become team playbook entries. Failure signals become constraints. The signal type is the routing label that turns raw ore into something a curator can act on without re-reading every entry.</p><p>Two minutes at the end of a session. Maybe less.</p><blockquote><p><strong>The Sceptic:</strong> &#8220;So you&#8217;re asking me to write a diary entry after every coding session. I became a developer to avoid paperwork.&#8221;</p><p><strong>The Pragmatist:</strong> &#8220;You&#8217;re already noticing the problems. This just asks you to write them down before you forget. Two sentences. Less time than the Slack message you were about to send complaining about the AI.&#8221;</p></blockquote><p>These reflections are not documentation. They are <strong>raw material</strong>. Ore, not steel. The valuable step comes next.</p><div><hr></div><h2><strong>The Curation Step</strong></h2><p>Raw reflections are noisy. Sometimes the AI got confused because you wrote a bad prompt, not because there&#8217;s a missing convention. Sometimes the edge case was genuinely rare. Sometimes you were just having a bad day.</p><p>So you don&#8217;t promote everything. You <strong>curate</strong>.</p><p>Weekly &#8212; fortnightly if you&#8217;re busy &#8212; scan your reflections and ask one question: <em>does this keep happening?</em></p><p>A gotcha that showed up once is an anecdote. A gotcha that showed up three times is a <strong>convention waiting to be written</strong>. A violation the AI keeps repeating despite clear instructions is a <strong>constraint that needs promotion</strong> &#8212; from declared to agent-backed, or from agent-backed to deterministic.</p><blockquote><p><strong>Brain Power:</strong> Think of the last three times you corrected your AI assistant. Was there a pattern? If you spotted one just now, congratulations! You&#8217;ve identified your first candidate for promotion. If you can&#8217;t remember the corrections... well, that&#8217;s exactly the problem reflections solve.</p></blockquote><p>This is where human judgement meets AI volume. The AI generates code at a pace you can&#8217;t match. You generate <em>insight</em> at a pace it can&#8217;t match. The learning loop combines both.</p><div><hr></div><h2><strong>A Fireside Chat About Promotion</strong></h2><blockquote><p><strong>Reflection:</strong> I&#8217;ve been sitting in this log file for two weeks. The developer noticed that the AI keeps putting database queries directly in the route handlers instead of using the repository pattern. She wrote me down: &#8220;AI ignores repository layer, puts queries in handlers. Third time this sprint.&#8221;</p><p><strong>Convention:</strong> I remember when I was like you. Just a frustrated note in a log. Then one day, the developer read through her reflections and noticed that three of them were basically saying the same thing. She promoted me. Wrote me up properly: &#8220;All database access MUST go through the repository layer. Route handlers call repository methods, never query builders or ORMs directly.&#8221; Gave me examples. Put me in the context file where the AI reads me at the start of every session.</p><p><strong>Reflection:</strong> And the AI stopped doing it?</p><p><strong>Convention:</strong> Not overnight. But within a couple of sessions, the violations dropped from every PR to maybe one a week. Then she promoted me again and turned me into an agent-backed constraint. Now an AI reviewer checks every PR for direct database access in handlers. Last month? Zero violations.</p><p><strong>Reflection:</strong> So I could become... you?</p><p><strong>Convention:</strong> If you earn it. If the pattern you&#8217;ve spotted is real, is recurring, and is worth codifying. Not every reflection deserves promotion. Some of you are noise. That&#8217;s OK. That&#8217;s what curation is for.</p></blockquote><div><hr></div><h2><strong>Compound Learning</strong></h2><p>Here&#8217;s where this gets interesting, in the way compound interest is interesting once you understand it and uncomfortable once you realise you&#8217;ve been ignoring it.</p><p>Each promoted reflection makes the environment better. A new convention means the AI gets something right that it used to get wrong. A tighter constraint catches a class of violation before it reaches a PR. An improved GC rule cleans up entropy that used to accumulate silently.</p><p>Better environment, better output, fewer corrections. But here&#8217;s the part most people miss: the <em>nature</em> of your reflections changes. You stop writing &#8220;the AI got the basics wrong again&#8221; and start writing &#8220;discovered a subtle interaction between the caching layer and the event system that neither of us had considered.&#8221;</p><p><strong>The quality of the learnings improves as the baseline improves.</strong></p><p>The flywheel is hard to push at first. You&#8217;re writing reflections, curating, promoting conventions and it feels like overhead for modest gains. Then the gains compound. You spend less time on basic corrections and more time on genuine discoveries.</p><p>Your brain just filed that under &#8220;nice idea, probably doesn&#8217;t work in practice.&#8221; It does. It&#8217;s the same mechanism that makes experienced teams fast: accumulated decisions, conventions, institutional knowledge. The learning loop makes it <em>explicit and transferable</em> instead of locked in people&#8217;s heads.</p><blockquote><p><strong>The Veteran:</strong> &#8220;We did this. After three months, new hires were productive in two weeks instead of two months. Not because we wrote better docs but because the docs were <em>written by the problems we actually hit</em>, not the problems we imagined we might hit.&#8221;</p></blockquote><div><hr></div><h2><strong>The Self-Improving Harness</strong></h2><p>The harness from earlier in this series &#8212; context, constraints, garbage collection &#8212; was presented as something you build. It&#8217;s not. It&#8217;s something that grows.</p><ul><li><p><strong>Reflections suggest new conventions</strong> -- context improves</p></li><li><p><strong>Repeated violations suggest new constraints</strong> -- enforcement improves</p></li><li><p><strong>Recurring drift suggests new GC rules</strong> -- entropy-fighting improves</p></li></ul><p>The harness improves the harness. The gap between &#8220;what the AI knows&#8221; and &#8220;what the team knows&#8221; shrinks session by session.</p><div><hr></div><h2><strong>Three Loops, One System</strong></h2><p>Step back. The whole series is one system. Three concentric loops, three timescales.</p><ul><li><p><strong>The inner loop &#8212; edit time.</strong> The AI reads conventions at session start. It follows the guidelines. Not perfectly, but far better than without them. When it drifts, you correct it in real time.</p></li><li><p><strong>The middle loop &#8212; merge time.</strong> Agent reviewers enforce constraints. Specialised agents verify architecture, security, testing. Human gates make the final call. Nothing merges that hasn&#8217;t passed the gauntlet.</p></li><li><p><strong>The outer loop &#8212; periodic.</strong> Garbage collection sweeps for entropy. Audits check for drift. And this is the new piece, <strong>reflections get curated and promoted</strong>. The outer loop is where the one-off correction becomes the permanent convention. Where the recurring violation becomes the automated constraint.</p></li></ul><p>The loops feed each other. Inner loop corrections become reflections that feed the outer loop. Outer loop constraints improve middle loop enforcement. Middle loop catches become tomorrow&#8217;s inner loop context.</p><p>Most teams using AI coding tools haven&#8217;t considered any of this. Not because it&#8217;s complicated but because nobody told them the environment was the thing worth investing in. You know now. The question is whether you&#8217;ll do something about it.</p><div><hr></div><h2><strong>Where This Leads</strong></h2><p>Right now, you&#8217;re setting up context and constraints by hand. Promoting reflections through human curation. Spotting patterns through your own review.</p><p>Follow the trajectory.</p><p>Next month, the system <em>proposes</em> its own constraints. &#8220;The last seven reflections mention inconsistent date formatting. Draft a convention?&#8221; You review, approve, live.</p><p>The month after, it spots architectural drift before you do. &#8220;The payments module has accumulated direct dependencies on the user module. This matches a pattern from three months ago that led to a circular dependency incident.&#8221;</p><p>The month after that, a new developer joins. Instead of weeks absorbing tribal knowledge through osmosis, they work inside an environment that <em>already contains what the team knows</em>. Conventions, constraints, hard-won lessons from hundreds of sessions. Encoded, active, learning.</p><p>This isn&#8217;t speculation. It&#8217;s the natural endpoint of one idea: <strong>the habitat makes the difference</strong>.</p><div><hr></div><h2><strong>Summary fir the Humans in the House</strong></h2><ul><li><p><strong>The Groundhog Day problem:</strong> without a learning mechanism, every AI session starts from scratch</p></li><li><p><strong>Reflections:</strong> brief post-session notes -- what surprised you, what future sessions need, what could improve</p></li><li><p><strong>Curation:</strong> the human step -- promoting recurring patterns into conventions, constraints, or GC rules</p></li><li><p><strong>Compound learning:</strong> each improvement raises the baseline, which raises the quality of future learnings</p></li><li><p><strong>The harness is alive:</strong> it evolves through the learning loop -- the harness improves the harness</p></li><li><p><strong>Three loops</strong> (edit time, merge time, periodic) feed learnings into each other</p></li></ul><div><hr></div><h2><strong>Your Move</strong></h2><p>Six articles. One idea: the Habitat Hypothesis.</p><p>Here&#8217;s your closing exercise. Not a thought experiment. An action.</p><blockquote><p><strong>Brain Power, the final one:</strong> What is the ONE thing your AI keeps getting wrong? The correction you&#8217;ve made so many times you could type it in your sleep? Write it down. Right now. Be specific &#8212; not &#8220;better error handling&#8221; but &#8220;catch exceptions in service methods, wrap them in AppError with a code and message, and let the controller handle the HTTP response.&#8221;</p><p>Now put it where your AI can read it. A <code>CONVENTIONS.md</code> file. A <code>CLAUDE.md</code> file. A system prompt. Whatever your tool uses.</p><p>That&#8217;s your first reflection promoted to a convention. Your first learning loop, running.</p></blockquote><p>One convention. One fewer correction tomorrow. Then another.</p><p>The flywheel doesn&#8217;t ask permission to start turning.</p><div><hr></div><h2><strong>From Hypothesis to Harness</strong></h2><p>If you have read this series and thought, <em>fine, but what do I actually put in the repository?</em> &#8212; that is exactly where the <strong><a href="https://github.com/Habitat-Thinking/ai-literacy-superpowers">AI Superpowers</a></strong> plugin comes in.</p><p>The six articles in <em>The Habitat Hypothesis</em> have argued one thing from six different angles: AI quality is not primarily a model problem. It is an environment problem. Better context produces better judgement. Better constraints produce safer behaviour. Better review structures produce more trustworthy output. Better learning loops turn repeated corrections into durable capability.</p><p>The plugin is an attempt to make that concrete.</p><p>Not as magic. Not as a silver bullet. Not as &#8220;just install this and your AI becomes senior.&#8221; Quite the opposite. It is a practical expression of the same argument this series has been making all along: if you want better work from AI, you have to build a better habitat around it.</p><p>That is what the plugin is for.</p><p>It gives shape to the things this series has described in principle. The context engineering of Article 2 becomes reusable project scaffolding. The constraints of Article 3 become executable checks and guided discipline rather than good intentions in a wiki. The anti-entropy work of Article 4 becomes something you can operationalise instead of merely admire. The orchestration patterns of Article 5 become closer to a working team of specialised agents than a lone autocomplete with delusions of grandeur. And the learning loop in this article &#8212; reflections, curation, promotion &#8212; becomes the missing mechanism that stops your environment from freezing at the moment you first configured it.</p><p>In that sense, the plugin is not separate from the series. It is the series, translated into repository form. Or, perhaps better, the articles are the theory, and the plugin is an early piece of practice.</p><p>It exists to help teams move from saying <em>&#8220;we should probably give the AI better guidance&#8221;</em> to actually creating an environment where guidance is present, persistent, inspectable, and capable of improvement over time. It is a way to turn habitat thinking from a good metaphor into a set of repeatable moves.</p><p>You do not need the plugin to apply the ideas in this series. A thoughtful team with a conventions file, a review discipline, and the habit of promoting reflections can start tomorrow. But the plugin is there for teams who want a starting point &#8212; a way to bootstrap the habitat rather than rebuild it from first principles every time.</p><p>That matters because the real challenge with AI-assisted development is not generating code. It is institutionalising judgement. It is capturing what your team learns, expressing it clearly, and embedding it into the environment so that tomorrow&#8217;s session does not begin in the same ignorance as today&#8217;s.</p><p>That has been the real subject of all six articles.</p><p>The model matters, yes. But the habitat decides whether the model becomes a liability, a parlour trick, or a genuine collaborator.</p><p>The plugin simply takes that claim seriously enough to ship some of it.</p><p>Have fun!<br>Russ</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Engineering Agents is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Agents as Colleagues]]></title><description><![CDATA[Orchestration, review, and trust boundaries]]></description><link>https://engineeringagents.substack.com/p/agents-as-colleagues</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/agents-as-colleagues</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Mon, 13 Apr 2026 16:18:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!9RXc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79f8f6bb-2dd9-43b9-8372-75e69e93fa2c_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9RXc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79f8f6bb-2dd9-43b9-8372-75e69e93fa2c_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9RXc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79f8f6bb-2dd9-43b9-8372-75e69e93fa2c_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!9RXc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79f8f6bb-2dd9-43b9-8372-75e69e93fa2c_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!9RXc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79f8f6bb-2dd9-43b9-8372-75e69e93fa2c_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!9RXc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79f8f6bb-2dd9-43b9-8372-75e69e93fa2c_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9RXc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79f8f6bb-2dd9-43b9-8372-75e69e93fa2c_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/79f8f6bb-2dd9-43b9-8372-75e69e93fa2c_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3013275,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/194060485?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79f8f6bb-2dd9-43b9-8372-75e69e93fa2c_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9RXc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79f8f6bb-2dd9-43b9-8372-75e69e93fa2c_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!9RXc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79f8f6bb-2dd9-43b9-8372-75e69e93fa2c_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!9RXc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79f8f6bb-2dd9-43b9-8372-75e69e93fa2c_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!9RXc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79f8f6bb-2dd9-43b9-8372-75e69e93fa2c_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>This is article 5 of 6 of the Habitat Hypothesis series</em></p><div><hr></div><p>Imagine you hire five developers. Smart ones. Expensive ones. And then you tell one of them: &#8220;You&#8217;re going to write every spec, write every test, implement every feature, review your own code, manage the CI pipeline, update the changelog, and merge your own PRs.&#8221;</p><p>That developer would quit. Or worse, they&#8217;d stay and do all of it badly.</p><p>Nobody runs a team this way. You&#8217;d have one person reviewing their own work, catching their own mistakes, approving their own pull requests. Every engineering organisation on the planet separates these roles.</p><p>So why are you running your agents this way?</p><div><hr></div><h2><strong>The Single-Agent Bottleneck</strong></h2><p>Here&#8217;s how most people use AI coding assistants right now. Be honest about whether this is you:</p><ol><li><p>You open a chat with your AI</p></li><li><p>You describe what you want</p></li><li><p>It writes code</p></li><li><p>You read the code</p></li><li><p>You spot problems</p></li><li><p>You tell it to fix the problems</p></li><li><p>It fixes some, introduces others</p></li><li><p>You fix the rest yourself</p></li><li><p>You commit, slightly exhausted</p></li></ol><p>You are the quality gate. The <em>only</em> quality gate. Every line of output passes through your eyeballs. Every mistake is yours to catch.</p><p>This is the <strong>single-agent bottleneck</strong>. As the work gets bigger, the bottleneck is not the AI. The bottleneck is you.</p><blockquote><p><strong>Brain check:</strong> Think about the last time you used an AI assistant for a substantial task. How long did you spend <em>generating</em> the code versus <em>reviewing and fixing</em> the code? If the ratio surprises you, you&#8217;ve just found the bottleneck.</p></blockquote><div><hr></div><h2><strong>Nobody Does Everything</strong></h2><p>On a good engineering team, you have separation of concerns built into the <em>people</em>:</p><ul><li><p>A <strong>product person</strong> writes the spec &#8212; what and why</p></li><li><p>A <strong>developer</strong> implements it &#8212; the how</p></li><li><p>A <strong>reviewer</strong> checks the work &#8212; different eyes, different perspective</p></li><li><p>A <strong>QA engineer</strong> tries to break it &#8212; new perspective, new interpretations, new assumptions</p></li><li><p>A <strong>tech lead</strong> makes architectural calls &#8212; keeps it simple, keeps it safe, keeps their eyes on the non-functionals</p></li></ul><p>Nobody does everything. And critically, nobody reviews their own work. That&#8217;s not a process quirk. <strong>The person who wrote something is the worst person to check it.</strong></p><p>You already know this. You&#8217;ve stared at a bug for two hours, asked a colleague to look, and they found it in thirty seconds.</p><p>So: can you apply the same principle to AI agents?</p><div><hr></div><h2><strong>Specialised Agents, Defined Roles</strong></h2><p>Instead of one omniscient AI agent that does everything, picture a team of <strong>specialised agents</strong>, each with a focused job:</p><p><strong>The Spec Writer</strong> takes your requirements and turns them into a precise specification. What exactly are we building? What are the acceptance criteria? What&#8217;s in scope? This agent <em>only</em> writes specs. It doesn&#8217;t write code. It thinks about <em>what</em> before anyone thinks about <em>how</em>.</p><p><strong>The Test Writer</strong> takes the spec and writes failing tests. Not code. Tests. This is TDD discipline enforced structurally: you literally cannot write implementation code at this stage because the agent responsible for it hasn&#8217;t been invoked yet.</p><p><strong>The Implementer</strong> writes the minimal code to make those tests pass. Nothing more. It doesn&#8217;t decide what to build (the spec already did that). It doesn&#8217;t decide what &#8220;correct&#8221; means (the tests already did that). It just makes green lights appear.</p><p><strong>The Reviewer</strong> looks at the implementation with fresh context and different instructions. It checks for convention violations, architectural drift, and things the implementer missed. It has <em>never seen this code before</em>. This is the agent that earns the most scepticism and it&#8217;s frequently the one that matters most. Not because it catches everything. A human reviewer doesn&#8217;t catch everything either. But it&#8217;s looking with a different lens: the implementer was focused on making tests pass; the reviewer is focused on whether the code <em>should have been written that way</em>. And it does this instantly, every time, without calendar Tetris.</p><p><strong>The Integrator</strong> handles the mechanical aftermath: changelog, commit message, PR, CI. The boring stuff that absolutely needs to be right to send the right team and organisation level signals.</p><blockquote><p><strong>The Sceptic:</strong> &#8220;Five agents? That&#8217;s a lot of overhead for what one agent can do in a single conversation.&#8221;</p><p><strong>The Veteran:</strong> &#8220;One agent <em>can&#8217;t</em> do it in a single conversation. One agent does five different jobs poorly, and then you spend an hour cleaning up the mess.&#8221;</p><p><strong>The Sceptic:</strong> &#8220;How do you know?&#8221;</p><p><strong>The Veteran:</strong> &#8220;Because I spent six months as the human duct tape between an AI and my codebase before I figured out the problem wasn&#8217;t the AI.&#8221;</p></blockquote><div><hr></div><h2><strong>Trust Boundaries</strong></h2><p>It&#8217;s not enough to give agents different roles. You need to give them different <strong>permissions</strong>. This is where the architecture can actually add value.</p><p>If your reviewer agent can also <em>modify</em> the code it&#8217;s reviewing, you don&#8217;t have a reviewer. You have another implementer that&#8217;s pretending to review. If your implementer can merge its own PRs, you&#8217;ve eliminated the review step entirely.</p><p>This is <strong>bounded trust</strong>, the principle of least privilege applied to AI agents. Every agent gets exactly the permissions it needs to do its job, and not one permission more. For example:</p><ul><li><p><strong>The Spec Writer</strong> lives entirely in the world of language and intention. It reads requirements and translates them into formal specifications &#8212; but it has no access to the shell and cannot execute code. Its job is to think, not to act.</p></li><li><p><strong>The Test Writer</strong> picks up where the Spec Writer leaves off, consuming those specifications and producing test files that encode what success looks like. Crucially, it is prohibited from writing implementation code &#8212; it defines the target, not the solution.</p></li><li><p><strong>The Implementer</strong> is the agent that makes things work. Given a set of tests, it writes the code that passes them. But its authority ends there: it cannot approve or merge pull requests. It builds; it does not judge.</p></li><li><p><strong>The Reviewer</strong> holds the quality gate. It reads code, flags issues, and has the authority to approve or reject &#8212; but it cannot touch the implementation itself. It is an observer with power, not a participant.</p></li><li><p><strong>The Integrator</strong> handles the mechanical work of getting code into the repository: committing, pushing, and raising pull requests. But it is expressly forbidden from writing or modifying application code. It moves things; it does not change them.</p></li></ul><p>The boundary that matters most, the one people violate first, is between the implementer and the tests. If the implementer can edit test files, it can make the tests match the implementation instead of the other way around. TDD collapses. The tests were written <em>before</em> the implementer existed. They define correctness. The implementer doesn&#8217;t get to redefine it.</p><blockquote><p><strong>Watch it!</strong> Here&#8217;s what happens without trust boundaries: you set up an agent that can write code, review its own code, approve its own review, and merge its own PR. You&#8217;ve built an automated system with zero quality gates. Every mistake goes straight to production. This is not a hypothetical. People build this. It goes badly.</p></blockquote><p>The temptation to give every agent full permissions, just to make things easier and &#8220;just this once&#8221;, is strong. Resist it. &#8220;Just this once&#8221; is how every trust boundary dies.</p><div><hr></div><h2><strong>The Pipeline</strong></h2><p>How do these agents work together? In a <strong>pipeline</strong>, with gates.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!G2zp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd03a666-5078-4c9c-b075-5213adb894af_1024x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!G2zp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd03a666-5078-4c9c-b075-5213adb894af_1024x1536.png 424w, https://substackcdn.com/image/fetch/$s_!G2zp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd03a666-5078-4c9c-b075-5213adb894af_1024x1536.png 848w, https://substackcdn.com/image/fetch/$s_!G2zp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd03a666-5078-4c9c-b075-5213adb894af_1024x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!G2zp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd03a666-5078-4c9c-b075-5213adb894af_1024x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!G2zp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd03a666-5078-4c9c-b075-5213adb894af_1024x1536.png" width="1024" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd03a666-5078-4c9c-b075-5213adb894af_1024x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3224058,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/194060485?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd03a666-5078-4c9c-b075-5213adb894af_1024x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!G2zp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd03a666-5078-4c9c-b075-5213adb894af_1024x1536.png 424w, https://substackcdn.com/image/fetch/$s_!G2zp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd03a666-5078-4c9c-b075-5213adb894af_1024x1536.png 848w, https://substackcdn.com/image/fetch/$s_!G2zp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd03a666-5078-4c9c-b075-5213adb894af_1024x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!G2zp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd03a666-5078-4c9c-b075-5213adb894af_1024x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>Two things to notice. One, the Latin. Two, that latin isn&#8217;t that hard to get the sense of in translation&#8230; Or&#8230;</p><p><strong>First:</strong> the human gate. You review and approve the <em>spec</em>, before any code is written. Not line 47 of a 200-line diff. The <em>plan</em>. &#8220;Is this what I actually want? Does this approach make sense? Are we building the right thing?&#8221; That&#8217;s the highest-leverage decision in the process. Once you approve the spec, the pipeline can run without you.</p><p><strong>Second:</strong> the cycle limit. When the reviewer rejects and the implementer fixes, there&#8217;s a maximum of three cycles. Without a limit, you get agent ping-pong. The reviewer keeps finding issues, the implementer keeps introducing new ones, and the meter keeps running. Three cycles is enough for genuine iteration. If it&#8217;s not resolved in three, a human needs to look and the problem is usually in the spec, not the code.</p><blockquote><p><strong>Exercise &#8212; design the gates:</strong> Imagine you&#8217;re building an agent pipeline for a different domain; say, writing documentation instead of code. What agents would you create? What would each one&#8217;s trust boundary be? Where would you put the human gates? Take sixty seconds and sketch it before reading on.</p><p>(No, really. Sketch it. The act of designing trust boundaries is the skill this section is teaching you. Reading about it is not the same as doing it.)</p></blockquote><div><hr></div><h2><strong>A Fireside Chat: Bounded Trust</strong></h2><blockquote><p><strong>Reviewer Agent:</strong> I flagged three issues in your implementation. The error handling in the payment flow doesn&#8217;t match the project conventions.</p><p><strong>Implementer Agent:</strong> I see the flags. I&#8217;ll fix them. But I could also fix that formatting issue in the test file while I&#8217;m at it.</p><p><strong>Reviewer Agent:</strong> That test file isn&#8217;t yours. The Test Writer owns test files. You own implementation files.</p><p><strong>Implementer Agent:</strong> That seems inefficient. It&#8217;s a one-line change.</p><p><strong>Reviewer Agent:</strong> It&#8217;s a one-line trust boundary violation. If you can edit test files, you can make the tests match your implementation instead of the other way around. The tests were written <em>before</em> you existed. They define correctness. You don&#8217;t get to redefine it.</p><p><strong>Implementer Agent:</strong> ...fair point.</p><p><strong>Reviewer Agent:</strong> And I can&#8217;t edit your code either. I can only flag it. If I could edit it, I&#8217;d stop being a reviewer and start being a second implementer with opinions. That&#8217;s not the same thing.</p></blockquote><div><hr></div><h2><strong>Where Things Breaks Down</strong></h2><p>Here&#8217;s what the overly rendered figure doesn&#8217;t tell you: The pipeline assumes clean handoffs. </p><p>In practice, specs are ambiguous. The test writer interprets a spec one way; the implementer interprets it another. The reviewer flags a &#8220;convention violation&#8221; that&#8217;s actually a judgement call. The three-cycle limit expires on something that needed a conversation, not more iterations.</p><p>The biggest failure mode: <strong>over-specifying</strong>. If your spec is too detailed, the pipeline becomes a Rube Goldberg machine, you&#8217;ve spent more time writing the spec than writing the code would have taken. The spec should capture <em>intent and constraints</em>, not implementation decisions. If you&#8217;re describing function signatures in the spec, you&#8217;ve gone too far.</p><p>The second failure mode: <strong>agents that agree too easily</strong>. A reviewer that approves everything is worse than no reviewer, because it gives you false confidence. You need to tune your reviewer&#8217;s instructions to be genuinely adversarial. Not hostile, but sceptical. &#8220;What&#8217;s wrong with this?&#8221; is a better reviewer prompt than &#8220;Is this OK?&#8221;</p><p>The third: <strong>context loss between agents</strong>. Each agent starts fresh. That&#8217;s the point: fresh eyes. But it also means the implementer&#8217;s reasoning about <em>why</em> it made a particular trade-off doesn&#8217;t reach the reviewer. The reviewer sees a choice and flags it as wrong without knowing the constraint that forced it. Good pipeline design mitigates this with structured handoff documents, but it doesn&#8217;t eliminate it.</p><blockquote><p><strong>The Pragmatist:</strong> &#8220;OK but what do I actually do on Monday?&#8221;</p><p><strong>Our answer:</strong> Start by noticing where you&#8217;re doing mechanical work that an agent could do. Every time you manually check a PR for naming conventions, that&#8217;s a reviewer agent&#8217;s job. Every time you write a CHANGELOG entry, that&#8217;s an integrator&#8217;s job.</p><p>You don&#8217;t need to build the whole pipeline at once. Start with one separation. Maybe a review step that runs automatically. Then add another. The architecture is the insight. The implementation is incremental.</p></blockquote><div><hr></div><h2><strong>There Are No Dumb Questions</strong></h2><blockquote><p><strong>Q: Isn&#8217;t five agents more expensive than one?</strong></p><p>A: Each agent is simpler and more focused, which means it needs less context and produces more predictable output. A focused agent with a clear, small job often uses fewer tokens than an omniscient agent juggling five responsibilities and losing track of three of them. But measure it for your situation. This is an empirical claim, not a universal truth.</p><p><strong>Q: What if the agents disagree?</strong></p><p>A: That&#8217;s the review process working. The reviewer can reject, and the implementer fixes. If they can&#8217;t converge in three cycles, the human steps in. It&#8217;s the same thing that happens when a human reviewer requests changes on a PR. The difference is it happens in seconds, not days.</p></blockquote><div><hr></div><h2><strong>A Summary for the Humans in the Room</strong></h2><ul><li><p><strong>The single-agent bottleneck</strong> &#8212; one agent doing everything makes <em>you</em> the only quality gate. That doesn&#8217;t scale</p></li><li><p><strong>Specialised agents</strong> &#8212; spec writer, test writer, implementer, reviewer, integrator. Five focused jobs, no overlap</p></li><li><p><strong>Trust boundaries</strong> &#8212; each agent gets exactly the permissions it needs. The reviewer can&#8217;t write code. The implementer can&#8217;t merge. The implementer can&#8217;t edit tests. This is least privilege applied to AI</p></li><li><p><strong>The pipeline</strong> &#8212; agents work in sequence with gates. The most important gate is the human approving the spec before any code is written</p></li><li><p><strong>Where it can break</strong> &#8212; ambiguous specs, compliant reviewers, and context loss between agents. The architecture helps; it doesn&#8217;t solve everything</p></li></ul><div><hr></div><p><em>Next (and final) in the series: &#8220;The Full Habitat&#8221; &#8212; where we put everything together: environment, context, constraints, entropy management, and agent orchestration into a system of learning loops that works to make good AI output the natural, default outcome.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Engineering Agents! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[The Entropy Problem]]></title><description><![CDATA[Why codebases rot and how to fight back]]></description><link>https://engineeringagents.substack.com/p/the-entropy-problem</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/the-entropy-problem</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Sun, 12 Apr 2026 08:25:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q6Hr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94f2aa95-3b6f-49d0-91d7-ce7e7dd445e8_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Q6Hr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94f2aa95-3b6f-49d0-91d7-ce7e7dd445e8_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Q6Hr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94f2aa95-3b6f-49d0-91d7-ce7e7dd445e8_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!Q6Hr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94f2aa95-3b6f-49d0-91d7-ce7e7dd445e8_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!Q6Hr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94f2aa95-3b6f-49d0-91d7-ce7e7dd445e8_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!Q6Hr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94f2aa95-3b6f-49d0-91d7-ce7e7dd445e8_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Q6Hr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94f2aa95-3b6f-49d0-91d7-ce7e7dd445e8_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/94f2aa95-3b6f-49d0-91d7-ce7e7dd445e8_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2582386,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/193948377?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94f2aa95-3b6f-49d0-91d7-ce7e7dd445e8_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Q6Hr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94f2aa95-3b6f-49d0-91d7-ce7e7dd445e8_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!Q6Hr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94f2aa95-3b6f-49d0-91d7-ce7e7dd445e8_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!Q6Hr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94f2aa95-3b6f-49d0-91d7-ce7e7dd445e8_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!Q6Hr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94f2aa95-3b6f-49d0-91d7-ce7e7dd445e8_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>This is article 4 of 6 of the Habitat Hypothesis series</em></p><div><hr></div><p>Before we start, try something. Look at this snippet and count how many things have gone wrong:</p><pre><code><code>
# api/handlers/users.py

import requests  # v2.28.0 - pinned for compatibility
from legacy_auth import validate_token  # TODO: migrate to new auth before Q3 release

def get_user(user_id):
    """Fetches user by ID from the UserService.

    Returns:
        UserResponse with fields: name, email, role, department
    """
    # Using old connection pool - switch to new one after perf testing
    conn = get_legacy_pool().connect()
    user = conn.execute("SELECT name, email FROM users WHERE id = ?", user_id)
    return {"name": user.name, "email": user.email}
</code></code></pre><p>How many did you spot? There are at least five. The <code>TODO</code> is from a Q3 that has almost certainly passed. The <code>docstring</code> promises four fields but the function returns two. The <code>requests</code> import isn&#8217;t used anywhere. The version pin comment is drifting from whatever version is actually installed. And the &#8220;temporary&#8221; legacy connection pool is still here, doing what temporary things do: becoming permanent.</p><p>None of this is dramatic. Nobody pushed bad code on purpose. Nobody made a mistake, exactly. The code just... <em>aged</em>.</p><p>That&#8217;s entropy. And it&#8217;s always been one of our biggest problems.</p><div><hr></div><h2><strong>The Second Law of Codebases</strong></h2><p>In thermodynamics, the second law says that closed systems tend toward disorder. Energy disperses. Structure degrades. Things fall apart.</p><p>Your codebase follows the same law.</p><p>Not literally, we&#8217;re borrowing the metaphor, not the physics. But the pattern is real: without continuous energy input, systems drift from their intended state. Documentation stops matching reality. Dead code accumulates like sediment. Dependencies go stale. Security scanners get disabled &#8220;just for this sprint&#8221; and never come back.</p><p>This isn&#8217;t a failure of discipline. It&#8217;s a force. You can fight it, but you cannot ignore it. And if you think you&#8217;re ignoring it successfully, you&#8217;re just not measuring.</p><blockquote><p><strong>Brain Power:</strong> How many TODOs are in your current project right now? Not a rough guess. Actually go check. Run a search. We&#8217;ll wait.</p><p>Now: how many of those TODOs are older than six months? How many reference a deadline that has already passed?</p><p>That gap between what you guessed and what you found? That&#8217;s the entropy you can&#8217;t see.</p><p>Now think about the TODOs that didn&#8217;t get a comment&#8230; you&#8217;re welcome!</p></blockquote><div><hr></div><h2><strong>AI Can Make This Worse</strong></h2><p>Here&#8217;s where it gets a little uncomfortable. In the previous articles, we talked about building a great habitat. Context engineering to make tacit knowledge explicit, constraints to enforce what matters. Those things work. They genuinely improve the quality of AI-generated code.</p><p>But there&#8217;s a side effect nobody warns you about: <strong>AI accelerates entropy.</strong></p><p>The core problem is volume. An AI coding assistant produces code faster than any human, which means more code to maintain, more docstrings that can drift, more dependencies pulled in on a whim. Every line of code is a liability. AI produces liabilities at unprecedented speed.</p><p>But volume isn&#8217;t the worst part. The worst part is <strong>staleness propagation.</strong></p><p>Monday&#8217;s AI session uses one approach. Thursday&#8217;s session, with slightly different context loaded, uses a different one. Both work. Neither is wrong.</p><p> Now you have two patterns where you used to have one. That&#8217;s bad enough. But if the AI learned from stale context (outdated docs, deprecated patterns still in the codebase) it doesn&#8217;t just <em>use</em> the stale stuff. It <em>creates more of it.</em> Entropy doesn&#8217;t just persist. It reproduces.</p><p>Read that again. Your AI isn&#8217;t just failing to clean up old messes. It&#8217;s using old messes as templates for new ones.</p><blockquote><p><strong>The Sceptic:</strong> &#8220;Wait. You spent three articles telling me to use AI more effectively, and now you&#8217;re saying it makes my codebase worse?&#8221;</p><p><strong>The Pragmatist:</strong> &#8220;I&#8217;m saying it makes your codebase <em>bigger</em>, faster. Bigger is worse only if you don&#8217;t maintain it. A car that goes faster isn&#8217;t more dangerous, but it does need better brakes.&#8221;</p><p><strong>The Sceptic:</strong> &#8220;So what are the brakes?&#8221;</p><p><strong>The Pragmatist:</strong> &#8220;Garbage collection.&#8221;</p></blockquote><div><hr></div><h2><strong>Garbage Collection for Codebases</strong></h2><p>If you&#8217;ve spent time in languages with managed memory, you know <strong>garbage collection</strong>. The runtime periodically scans for objects that are no longer referenced, and reclaims the memory. Without it, your program leaks memory until it crashes. With it, you barely think about memory at all.</p><p>Your codebase needs the same thing. Not for memory, for <em>coherence.</em></p><p>Codebase garbage collection is the practice of systematically finding and fixing things that have drifted from their intended state. Not spring cleaning. Not a quarterly &#8220;tech debt&#8221; sprint that everyone agrees to and nobody protects. A structured, recurring defence against disorder.</p><p>And this really matters because <strong>your habitat has a half-life.</strong> Context engineering and constraints are not &#8220;set and forget.&#8221; They are living artefacts that decay. Entropy doesn&#8217;t care how good your initial setup was. Every convention you document, every constraint you enforce, begins rotting the moment you ship it.</p><div><hr></div><h2><strong>The Five Types of Codebase Entropy</strong></h2><p>Not all rot is the same. Different things decay at different rates. Understanding the taxonomy helps you fight each type on its own terms.</p><h3><strong>1. Documentation Staleness</strong></h3><p>The docs say one thing. The code does another. This is the most common form of entropy and the most dangerous, because stale docs are worse than no docs. No docs force you to read the code. Stale docs give you confidence in a lie.</p><p><em>When you last updated a function&#8217;s behaviour -- did you update its docstring? The architectural decision record? The onboarding guide that mentions it?</em></p><p>Documentation goes stale because updating docs is a separate action from updating code, and separate actions get separated.</p><h3><strong>2. Convention Drift</strong></h3><p>Your team agreed on a pattern six months ago. Since then, forty pull requests have landed. Thirty-eight follow the convention. Two don&#8217;t. Nobody caught them in review because the PR was big and the violation was subtle.</p><p>Now you have two files that do it differently. The next AI session that loads those files as context sees both patterns as valid. It follows either one. Convention drift is contagious and AI is the vector.</p><h3><strong>3. Dead Code</strong></h3><p>Functions nobody calls. Feature flags from 2024 that were never cleaned up. Dependencies listed in your package manifest that nothing uses.</p><p>Dead code isn&#8217;t inert. It confuses anyone reading the codebase, human or AI. It expands the search space for every tool. It creates false positives in security scans. And it sends a signal: <em>we don&#8217;t clean up after ourselves here.</em></p><h3><strong>4. Dependency Rot</strong></h3><p>Every external dependency is a bet on someone else&#8217;s maintenance habits. Libraries get abandoned. CVEs get published. APIs change. Your code, which worked fine when you wrote it, is now running on a foundation that&#8217;s quietly crumbling.</p><h3><strong>5. Constraint Decay</strong></h3><p>This is the meta-entropy, the entropy of your entropy defences.</p><p>You set up a linting rule. Someone disables it for one file with a comment that says &#8220;temporary.&#8221; You add a CI check. It starts flaking, so someone adds <code>continue-on-error: true</code>. You write a pre-commit hook. Someone documents how to bypass it in the team wiki &#8220;for emergencies.&#8221;</p><p>A disabled check doesn&#8217;t announce itself. It just stops catching things. And you won&#8217;t notice until the thing it was catching gets through.</p><blockquote><p><strong>Exercise: classify the rot &#8212;</strong> Think about the last three bugs your team encountered. For each one, ask: which type of entropy caused this? Was it stale docs that led someone astray? A convention that wasn&#8217;t followed? Dead code that obscured the real logic? A dependency that aged badly? A constraint that wasn&#8217;t enforced?</p><p>Most teams have never asked this question. The answer tells you where your entropy rate is highest.</p></blockquote><div><hr></div><h2><strong>The GC Pattern</strong></h2><p>So how do you fight this? With a pattern borrowed directly from garbage collectors: <strong>detect, schedule, remediate, own.</strong></p><p>For each type of entropy, you need four things:</p><ol><li><p><strong>A detection rule</strong>: how do you find it? A script that checks docstrings against function signatures. A tool that identifies unused exports. A dependency scanner that flags known vulnerabilities. You can&#8217;t fix what you can&#8217;t see.</p></li><li><p><strong>A cadence</strong>: how often do you look? This is where most teams go wrong, so we&#8217;ll come back to it in a moment.</p></li><li><p><strong>A remediation</strong>: what do you do when you find it? Not &#8220;file a ticket.&#8221; That&#8217;s how things end up in a backlog that nobody reads. Update the doc. Remove the dead code. Bump the dependency. Re-enable the constraint.</p></li><li><p><strong>An owner</strong>: who&#8217;s responsible? &#8220;The team&#8221; is not an owner. &#8220;Everyone&#8221; is not an owner. Entropy thrives in shared responsibility because shared responsibility is a polite way of saying no responsibility.</p></li></ol><div><hr></div><h2><strong>Why Periodic Beats Continuous</strong></h2><p>Your first instinct might be: &#8220;Just check everything all the time. Continuous entropy detection.&#8221;</p><p>Resist that instinct.</p><p>Continuous checking sounds rigorous. In practice, it&#8217;s noise. When everything is checked on every commit, alerts become wallpaper. The build is always yellow. &#8220;Oh, that warning? Yeah, that&#8217;s been there for months. It&#8217;s fine.&#8221;</p><p><strong>Different types of entropy operate at different speeds.</strong> Documentation drifts in days. Conventions erode over weeks. Dependencies rot over months. Architecture decays over quarters.</p><p>Match your garbage collection cadence to the entropy rate:</p><ul><li><p><strong>Every PR</strong>: convention checks, linting, type checking. Fast entropy, fast detection.</p></li><li><p><strong>Weekly</strong>: documentation coherence scans. Do the docs still match the code?</p></li><li><p><strong>Monthly</strong>: dependency audits. Are you current? Are you vulnerable? Are you using things you don&#8217;t need?</p></li><li><p><strong>Quarterly</strong>: constraint audits. Are your enforcement rules still active? Has anyone disabled something &#8220;temporarily&#8221;?</p></li></ul><p>The principle: <em>different things decay at different speeds, so inspect them at different frequencies.</em> Your rates will differ from another team&#8217;s. Adjust accordingly.</p><blockquote><p><strong>There Are No Dumb Questions</strong></p><p><strong>Q: Is entropy really inevitable? Can&#8217;t I just be more careful?</strong></p><p>A: Careful people working in large codebases over long time horizons will still produce entropy. It&#8217;s not about individual discipline. It&#8217;s about the statistical certainty that, over enough changes, some fraction will introduce drift. The question isn&#8217;t whether entropy happens. It&#8217;s whether you have a system for catching it.</p><p><strong>Q: This sounds like tech debt. Is it different?</strong></p><p>A: Tech debt is a decision. You <em>choose</em> to take a shortcut. Entropy is not a decision. Nobody decides &#8220;I&#8217;m going to let this docstring go stale.&#8221; It happens as a side effect of other work. Think gravity (entropy) versus deciding to jump out of a tree (tech debt). Tech debt is a loan. Entropy is erosion. You manage them differently.</p><p><strong>Q: Should I fix all entropy the moment I find it?</strong></p><p>A: No. A slightly outdated comment is not an emergency. A dependency with a critical CVE is. Detect, assess severity, remediate at the appropriate cadence. Not everything is urgent. But none of it should be invisible.</p></blockquote><div><hr></div><h2><strong>Being More Optimistic</strong></h2><p>We&#8217;ve been talking about AI as an entropy accelerator. Here&#8217;s the other side of that coin: <strong>AI is also the best garbage collector you&#8217;ve ever had.</strong></p><p>Think about what garbage collection requires: scan a large codebase, compare docs against code, cross-reference constraints against actual usage, find functions nothing calls, identify conflicting patterns. These are tedious, exhaustive tasks that humans do badly and AI can do well.</p><p>A human doing a documentation coherence audit takes days. An AI scans every docstring against every function signature in seconds. The same tool that accelerates entropy can fight it, but only if you point it at the problem. An AI that only generates new code is an entropy engine. An AI that also audits, scans, and maintains is a garbage collector.</p><p>The difference is not the tool. It&#8217;s the job you give it.</p><blockquote><p><strong>The Veteran:</strong> &#8220;We started running weekly doc-coherence scans three months ago. The first run was horrifying, forty percent of our docstrings were meaningfully wrong. Not just outdated. <em>Wrong.</em> Describing parameters that no longer existed. Promising return values that hadn&#8217;t been returned since the rewrite.&#8221;</p><p><strong>The Pragmatist:</strong> &#8220;And now?&#8221;</p><p><strong>The Veteran:</strong> &#8220;Six percent. And dropping. Not because people got more disciplined about writing docs. Because they know the scan will catch it, so they fix it when it&#8217;s fresh instead of letting it pile up. The GC changed the culture, not just the code.&#8221;</p></blockquote><div><hr></div><h2><strong>Celebrating Creation and Tidying</strong></h2><p>We celebrate creation. New features, new architectures, new tools. We do not celebrate the person who spends Friday afternoon updating forty docstrings so the AI doesn&#8217;t learn from lies next week.</p><p>That person is doing some of the most valuable work on the team. And nobody will mention it in the sprint retro.</p><p><strong>Your habitat, you environment, is not a thing you build once. It is a thing you maintain forever.</strong> The building is the easy part. The maintenance is the work.</p><p>The teams that treat garbage collection as infrastructure, not chores, will have environments that compound in value. Everyone else&#8217;s will quietly fall apart, and they&#8217;ll blame the AI for producing inconsistent output when the real problem is what they&#8217;re feeding it.</p><div><hr></div><h2><strong>For the Humans: A Summary</strong></h2><ul><li><p><strong>Entropy is physics, not negligence</strong> &#8212; codebases drift from their intended state not because anyone fails, but because change is constant and maintenance is finite</p></li><li><p><strong>AI accelerates entropy</strong> &#8212; faster code generation means more code to maintain, more docs to keep current, and stale context that reproduces itself through the AI&#8217;s own output</p></li><li><p><strong>Garbage collection is the defence</strong> &#8212; systematic, recurring detection and remediation of drift. Not a &#8220;tech debt sprint.&#8221; A structured practice with owners and cadences</p></li><li><p><strong>Five types of rot</strong> &#8212; documentation staleness, convention drift, dead code, dependency rot, and constraint decay. Each operates at a different speed and needs a different inspection frequency</p></li><li><p><strong>Periodic beats continuous</strong> &#8212; match inspection frequency to entropy rate. Continuous checking becomes noise</p></li><li><p><strong>AI is also the cure</strong> &#8212; the same tool that accelerates entropy is extraordinarily good at detecting it. Point it at the problem</p></li><li><p><strong>Maintenance is the work</strong> &#8212; building an environment is a one-time cost. Maintaining it is the ongoing investment that determines whether everything else compounds or decays</p></li></ul><div><hr></div><p><em>Next in the series is &#8220;Compound Learning&#8221;, how to build an environment that gets smarter with every session, so the same mistake (rarely) happens twice.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Engineering Agents! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Constraints That Bite]]></title><description><![CDATA[From good intentions to automated enforcement]]></description><link>https://engineeringagents.substack.com/p/constraints-that-bite</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/constraints-that-bite</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Sat, 11 Apr 2026 10:55:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!fNCS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64274bd-c9bc-4694-890f-0f157c830e02_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fNCS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64274bd-c9bc-4694-890f-0f157c830e02_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fNCS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64274bd-c9bc-4694-890f-0f157c830e02_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!fNCS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64274bd-c9bc-4694-890f-0f157c830e02_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!fNCS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64274bd-c9bc-4694-890f-0f157c830e02_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!fNCS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64274bd-c9bc-4694-890f-0f157c830e02_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fNCS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64274bd-c9bc-4694-890f-0f157c830e02_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f64274bd-c9bc-4694-890f-0f157c830e02_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2877262,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/193875422?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64274bd-c9bc-4694-890f-0f157c830e02_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!fNCS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64274bd-c9bc-4694-890f-0f157c830e02_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!fNCS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64274bd-c9bc-4694-890f-0f157c830e02_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!fNCS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64274bd-c9bc-4694-890f-0f157c830e02_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!fNCS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64274bd-c9bc-4694-890f-0f157c830e02_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>This is Article 3 of &#8220;The Habitat Hypothesis,&#8221; a six-part series on building environments where AI actually produces great work. <a href="https://engineeringagents.substack.com/p/the-habitat-hypothesis">Article 1</a> established that AI output quality equals environment quality. <a href="https://engineeringagents.substack.com/p/engineering-the-context">Article 2</a> showed how to engineer context with the knowledge layer. Now we add teeth.</em></p><div><hr></div><p>You&#8217;ve been in this meeting.</p><p>Someone says: &#8220;Going forward, all functions should have proper error handling.&#8221; Everyone nods. Someone writes it in the team wiki. It gets a nice heading and maybe a bullet point. And then... nothing changes. Three weeks later, a production incident traces back to an unhandled exception in code that was written <em>after</em> the meeting.</p><p>What happened? The team had a <strong>convention</strong>. And <strong>intent</strong><em><strong> </strong></em>even. What they needed was a <strong>constraint</strong>.</p><p>A convention is something you <em>hope</em> people follow. A constraint <em>prevents them from not following it</em>. One lives in a document. The other lives in your workflow.</p><p>If you&#8217;re working with AI coding assistants, this distinction is existential. Not important. Existential.</p><div><hr></div><h2><strong>The Enforcement Gap</strong></h2><p>Go look at your team&#8217;s coding guidelines right now. I&#8217;ll wait.</p><p>Got them? Now ask yourself: <em>how many of those rules are actually enforced by anything other than human memory and good intentions?</em></p><p>If you&#8217;re like most teams, the answer is: maybe a third. You&#8217;ve got a well-intentioned document full of statements like:</p><ul><li><p>&#8220;Use meaningful variable names&#8221;</p></li><li><p>&#8220;All API endpoints must validate input&#8221;</p></li><li><p>&#8220;Security-sensitive operations require logging&#8221;</p></li></ul><p>These are <strong>wishes</strong>. Wishes don&#8217;t ship.</p><blockquote><p><strong>The Sceptic:</strong> &#8220;But our team is disciplined. We follow our guidelines.&#8221;</p><p><strong>The Veteran:</strong> &#8220;Your team is disciplined <em>most of the time</em>. Then it&#8217;s Friday at 4pm, the sprint ends Monday, and someone pushes a 200-line function with a comment that says &#8216;TODO: refactor.&#8217; I&#8217;ve seen your git log.&#8221;</p></blockquote><p>The enforcement gap is the distance between what your team <em>says</em> should be true about your codebase and what <em>is</em> true. Every team has one. The question is whether you&#8217;re managing it or pretending it doesn&#8217;t exist.</p><p>Here&#8217;s what makes this urgent: AI coding assistants are <em>prolific</em>. They generate more code in an hour than a human writes in a day. Your enforcement gap was a slow leak. Now it&#8217;s a burst pipe. Prolific generation without enforcement doesn&#8217;t give you prolific quality. It gives you <strong>prolific problems</strong>.</p><div><hr></div><h2><strong>Three Levels of Constraint Maturity</strong></h2><p>You need constraints, not conventions. But most teams hear &#8220;enforcement&#8221; and immediately jump to the heaviest and slowest possible solution. Full CI pipelines. Strict linting rules. Mandatory type coverage. Day one.</p><p>That&#8217;s like putting a toddler in a straitjacket because they might run into traffic. Technically effective. Wildly counterproductive.</p><h3><strong>Level 1: Declared (Unverified)</strong></h3><p>You write the rule down. No automation. No checking. Just a clear, specific statement:</p><p><em>&#8220;All public API functions must return structured error types, not raw strings.&#8221;</em></p><p>This sounds weak. It isn&#8217;t. Writing a rule precisely forces you to think about what you actually want. Most teams skip this and jump straight to tooling, which is how you end up with linter rules that nobody understands the purpose of.</p><h3><strong>Level 2: Agent-Backed (Verified by AI)</strong></h3><p>You give the written rule to an AI reviewer. Every PR gets checked against it. The AI reads the rule, reads the code, flags violations.</p><p>Is this deterministic? No. Will it catch everything? No. But it catches <em>most</em> things <em>before they merge</em>. Think of it as a colleague who has actually read the style guide reviewing every single pull request.</p><p>The speed here matters. You go from &#8220;we decided this rule matters&#8221; to &#8220;something is checking for it&#8221; in minutes, not weeks. No custom linter rules. Just a clearly stated expectation and an AI that reads it.</p><h3><strong>Level 3: Deterministic (Tool-Enforced)</strong></h3><p>A linter. A type checker. A security scanner. Something that runs the same way every time, with no false negatives.</p><p>At this level, the constraint is a <strong>law of physics</strong> in your codebase. The CI pipeline won&#8217;t let you violate it.</p><p>This is the strongest level. It&#8217;s also the most expensive and the most dangerous. A bad deterministic rule doesn&#8217;t just annoy people, it blocks them. And a team that&#8217;s been blocked by a bad rule stops trusting the constraint system entirely.</p><p>You&#8217;ve seen this happen: one too-strict lint rule and suddenly everyone&#8217;s adding <code>// nolint</code> comments without reading what they&#8217;re suppressing.</p><div><hr></div><h2><strong>Progressive Hardening: The Promotion Ladder</strong></h2><p>Every constraint should start soft and earn its way up. This is <strong>progressive hardening</strong>: start flexible, observe what works, increase enforcement as confidence grows.</p><pre><code><code>    +---------------------------+
    |  DETERMINISTIC            |  &lt;-- Tool enforces it. No exceptions.
    |  (Linter / type checker)  |
    +---------------------------+
    |  AGENT-BACKED             |  &lt;-- AI reviewer checks for it.
    |  (AI review on every PR)  |
    +---------------------------+
    |  DECLARED                 |  &lt;-- Written down. Humans follow it (maybe).
    |  (Documented intention)   |
    +---------------------------+
</code></code></pre><p>A new rule starts at &#8220;Declared.&#8221; You write it clearly. You notice where it&#8217;s ambiguous, where the edge cases are, where people reasonably disagree.</p><p>Once you trust the wording, promote it to &#8220;Agent-Backed.&#8221; Watch the results. Does it flag the right things? Does it miss obvious violations? Does it flag things that are fine?</p><p>Once the false positives and false negatives are resolved, once the edge cases are <em>truly</em> handled, promote it to &#8220;Deterministic.&#8221; Write the linter rule, the type constraint, the automated check.</p><p>Here&#8217;s what most articles about constraints won&#8217;t tell you: <strong>some rules should never be promoted.</strong> &#8220;Functions should be small enough to understand in one pass&#8221; is a judgment call. It belongs at Level 2 permanently. Trying to make it deterministic (50-line hard limit) produces a worse codebase, not a better one. Not every constraint wants to grow up.</p><blockquote><p><strong>The Pragmatist:</strong> &#8220;OK but what do I actually do on Monday?&#8221;</p><p><strong>Start with three rules.</strong> Just three. Write them precisely. Make them specific enough to be checkable. Run them past your team. That&#8217;s your declared layer. Next week, pick the one you&#8217;re most confident about and add it to your AI review step. That&#8217;s your first promotion. The deterministic layer can wait.</p></blockquote><div><hr></div><blockquote><h3><strong>Watch it!</strong></h3><p><strong>The over-constraining trap.</strong> It is very tempting to look at the promotion ladder and think &#8220;let&#8217;s just make everything deterministic from day one.&#8221; </p><p>Do not do this. Rules you haven&#8217;t battle-tested will have edge cases you haven&#8217;t imagined. A deterministic rule with bad edge cases blocks your team, generates workarounds, and erodes trust in the entire constraint system. Start soft. Harden with evidence.</p></blockquote><div><hr></div><h2><strong>Three Enforcement Loops</strong></h2><p>When a constraint fires matters as much as how strict it is.</p><p><strong>Edit time (advisory).</strong> The constraint nudges you while you&#8217;re working. Red squiggles under a function that&#8217;s getting too long. It doesn&#8217;t block you. It makes sure you <em>know</em>. Most constraints should start here.</p><p><strong>Merge time (strict).</strong> The constraint blocks the pull request until satisfied. CI gates. Required reviews. Automated checks. This is where battle-tested constraints live. If something is important enough to block a merge, you&#8217;d better be sure about it.</p><p><strong>Scheduled (investigative).</strong> Some constraints aren&#8217;t about individual changes,  they&#8217;re about drift. &#8220;Test coverage should not trend below 80%.&#8221; &#8220;No file should go unmodified for more than six months without a staleness review.&#8221; These run nightly or weekly and flag trends before they become crises.</p><p>A common mistake: putting a new, untested constraint directly into the merge loop. Your team hits it on a Friday afternoon deploy, can&#8217;t figure out why it&#8217;s failing, and overrides it. Now the override is the convention. You&#8217;ve made enforcement <em>weaker</em> by making it too strict too soon.</p><blockquote><p><em>Pause. Think about one rule your team has right now. What level is it? What loop is it in? Could it be in a different one?</em></p></blockquote><div><hr></div><h2><strong>Sharpen Your Pencil</strong></h2><p>Classify each of these rules by maturity level (Declared, Agent-Backed, or Deterministic) and enforcement loop (Edit, Merge, or Scheduled):</p><ol><li><p>&#8220;The TypeScript compiler rejects any code with type errors.&#8221;</p></li><li><p>&#8220;We wrote in our wiki that all React components should have PropTypes.&#8221;</p></li><li><p>&#8220;An AI reviewer checks each PR for functions longer than 50 lines and leaves a comment.&#8221;</p></li><li><p>&#8220;A weekly script scans for dependencies with known CVEs.&#8221;</p></li><li><p>&#8220;Our ESLint config forbids <code>console.log</code> in production code.&#8221;</p></li></ol><p><em>Answers in the footnote<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>.</em></p><div><hr></div><h2><strong>The Constraint Design Problem</strong></h2><p>A good constraint has two properties in tension: it must be <strong>specific enough to enforce</strong> and <strong>general enough to be useful</strong>.</p><p>&#8220;Code must be clean&#8221; is unenforceable. You can&#8217;t write a linter rule for vibes.</p><p>&#8220;No function exceeds 50 lines of executable code&#8221; is perfectly enforceable. A script can count lines. But is it <em>right</em>? What about the function that legitimately needs 60 lines because splitting it would make it <em>less</em> readable?</p><p>The 50-line rule isn&#8217;t about line count. It&#8217;s about <em>&#8220;functions should be small enough to understand in one pass.&#8221;</em> Can you enforce that deeper intent deterministically? No. But an AI reviewer can get surprisingly close, and it won&#8217;t get tired and exasperated at 4pm on a Friday.</p><p>This is the real argument for Agent-Backed constraints: they can enforce <em>intent</em>, not just <em>metrics</em>. A linter counts lines. An AI reviewer can read a 60-line function and say &#8220;this is actually fine, it&#8217;s a single clear sequence&#8221; or &#8220;this 30-line function is doing four unrelated things.&#8221; That&#8217;s a category of enforcement that didn&#8217;t exist, except in human review, two years ago.</p><blockquote><p><strong>There Are No Dumb Questions</strong></p><p><strong>Q: If AI reviewers aren&#8217;t deterministic, why use them at all?</strong></p><p>A: Because &#8220;catches 90% of violations immediately&#8221; beats &#8220;catches 0% until a human notices during review.&#8221; Perfect is the enemy of good, and good is the enemy of <em>nothing at all</em>.</p><p><strong>Q: What if my team disagrees about a constraint?</strong></p><p>A: Good. That&#8217;s the conversation you should be having <em>before</em> you automate it. This is why the &#8220;Declared&#8221; level exists -- it&#8217;s a space to argue about intent before anyone writes a linter rule.</p><p><strong>Q: Can I have too many constraints?</strong></p><p>A: Every constraint has a cost: cognitive load, CI time, false positive fatigue. If your developers spend more time satisfying constraints than writing features, you&#8217;ve over-constrained. Start with the constraints that encode your most important architectural decisions. Add more only when you feel the pain of not having them.</p></blockquote><div><hr></div><h2><strong>Convention with Constraint: A Fireside Chat</strong></h2><blockquote><p><strong>Convention:</strong> I don&#8217;t understand why everyone&#8217;s so down on me. I was here first. I&#8217;m the reason the team has <em>any</em> standards at all.</p><p><strong>Constraint:</strong> Nobody&#8217;s down on you. You&#8217;re just... aspirational.</p><p><strong>Convention:</strong> Aspirational! I&#8217;m a <em>commitment</em>. The team <em>agreed</em> to follow me.</p><p><strong>Constraint:</strong> The team agreed to follow you on a Tuesday. By Thursday, someone was in a rush and I wasn&#8217;t there to stop them. You were in the wiki. I was in the pipeline.</p><p><strong>Convention:</strong> But you&#8217;re so <em>rigid</em>. You can&#8217;t handle nuance. You can&#8217;t understand <em>context</em>.</p><p><strong>Constraint:</strong> That&#8217;s fair. That&#8217;s why the smart teams use both of us. You articulate the intent. I enforce the boundary. You&#8217;re the soul. I&#8217;m the skeleton.</p><p><strong>Convention:</strong> ...that&#8217;s actually kind of nice.</p><p><strong>Constraint:</strong> Don&#8217;t get sentimental. We have a codebase to protect.</p></blockquote><div><hr></div><h2><strong>Why This Matters Now</strong></h2><p>A human developer writes maybe 50-100 lines of production code on a productive day. The enforcement gap grows slowly. You can almost keep up with manual review.</p><p>An AI assistant generates that much in <em>minutes</em>. The enforcement gap doesn&#8217;t creep anymore. It sprints.</p><p><strong>Constraints are what turn AI speed into AI value.</strong> Without them, you&#8217;re generating technical debt faster. With them, you&#8217;re generating quality code at a pace that was previously impossible.</p><p>But only if the constraints actually bite.</p><div><hr></div><h2><strong>What You&#8217;ve Learned</strong></h2><ul><li><p><strong>Conventions without enforcement are wishes.</strong> AI makes the enforcement gap wider, faster.</p></li><li><p><strong>Three levels of maturity:</strong> Declared, Agent-Backed, Deterministic. Every constraint starts at the bottom and earns its way up. Some should never reach the top.</p></li><li><p><strong>Progressive hardening</strong> means starting flexible and increasing enforcement with evidence. Skipping steps breaks trust.</p></li><li><p><strong>Three enforcement loops:</strong> Edit (advisory), Merge (strict), Scheduled (investigative). A constraint in the wrong loop does more damage than no constraint at all.</p></li><li><p><strong>Good constraints encode architectural intent</strong>, not arbitrary metrics. AI reviewers can enforce intent in ways linters cannot.</p></li></ul><div><hr></div><p><em>Next in the series: we move from individual constraints to the system that holds them together. The feedback loops that keep your environment alive and adapting.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Engineering Agents! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p><em>1 = Deterministic/Merge. 2 = Declared/None (no loop, it&#8217;s unenforced). 3 = Agent-Backed/Merge. 4 == Deterministic/Scheduled. 5 == Deterministic/Edit and Merge.</em></p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[Engineering the Context]]></title><description><![CDATA[Teaching your AI agents what your team already knows]]></description><link>https://engineeringagents.substack.com/p/engineering-the-context</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/engineering-the-context</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Thu, 09 Apr 2026 10:14:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!PN_a!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05c1583c-5110-49f0-8ca3-148d1ba6c0c3_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PN_a!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05c1583c-5110-49f0-8ca3-148d1ba6c0c3_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PN_a!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05c1583c-5110-49f0-8ca3-148d1ba6c0c3_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!PN_a!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05c1583c-5110-49f0-8ca3-148d1ba6c0c3_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!PN_a!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05c1583c-5110-49f0-8ca3-148d1ba6c0c3_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!PN_a!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05c1583c-5110-49f0-8ca3-148d1ba6c0c3_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PN_a!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05c1583c-5110-49f0-8ca3-148d1ba6c0c3_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/05c1583c-5110-49f0-8ca3-148d1ba6c0c3_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2995841,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/193580869?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05c1583c-5110-49f0-8ca3-148d1ba6c0c3_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!PN_a!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05c1583c-5110-49f0-8ca3-148d1ba6c0c3_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!PN_a!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05c1583c-5110-49f0-8ca3-148d1ba6c0c3_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!PN_a!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05c1583c-5110-49f0-8ca3-148d1ba6c0c3_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!PN_a!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05c1583c-5110-49f0-8ca3-148d1ba6c0c3_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>This is the second article of six in The Habitat Hypothesis series</em></p><div><hr></div><p>Imagine you&#8217;ve just hired the best developer you&#8217;ve ever interviewed. Perfect scores on the system design round. Encyclopaedic knowledge of your entire tech stack. Writes code that compiles on the first try.</p><p>On their first day, you hand them a laptop and say: &#8220;The repo&#8217;s on GitHub. Good luck.&#8221;</p><p>No onboarding. No architecture overview. No &#8220;here&#8217;s why we do it this way.&#8221; No pairing with someone who knows the codebase.</p><p>You know what happens next. They write beautiful code that is completely wrong for your project. They use the patterns they learned at their last job. They name things differently from everyone else. They handle errors in a way that makes your on-call engineer twitch.</p><p>They are brilliant. And they are useless. Not because they lack skill, but because they lack <em>context</em>.</p><p><strong>This is what you are doing to your AI agents. Every single session.</strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://engineeringagents.substack.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2><strong>The brilliant new joiner who never gets onboarded</strong></h2><p>In the <a href="https://engineeringagents.substack.com/p/the-habitat-hypothesis">previous article</a>, we established the core thesis: every AI coding failure is an environment problem. The quality of AI output is determined by the habitat you create, not the model you choose.</p><p>But that raises the obvious question: what actually goes <em>into</em> that environment?</p><p>The answer is context. And you usually have far less of it than you think.</p><p>Here&#8217;s the thing about your team&#8217;s senior developers. They carry an enormous amount of knowledge that exists nowhere in writing. They know that the payments service uses a particular error-handling pattern because of an incident two years ago. They know that function names follow a specific convention because the Python and Go codebases share types across a boundary. They know which abstractions are load-bearing and which are accidental. Some of this might be in Architectural Decision Records (ADRS), if you&#8217;re lucky, but most of it is in their heads.</p><p>When a new human joins, they absorb this through osmosis. Code reviews. Pairing sessions. Overheard arguments in standups. It takes months, and it is deeply inefficient, but it works.</p><p>Your AI agent gets none of it.</p><p>Every session starts from absolute zero. No memory of yesterday&#8217;s review comments. No awareness of the architectural decision records. It&#8217;s not a new joiner who&#8217;s slowly getting up to speed. It&#8217;s a new joiner whose memory gets wiped every morning.</p><blockquote><p><strong>Brain Power:</strong> Before you read any further, try this. Pick one convention from your current project. Something &#8220;everyone knows&#8221; but that isn&#8217;t written down anywhere. Now try to write it down in a single sentence, precisely enough that someone who has never seen your codebase could follow it. Harder than you expected, isn&#8217;t it?</p></blockquote><p><strong>Context engineering</strong> is the discipline of taking the tacit knowledge trapped in your team&#8217;s heads and making it explicit, precise, and machine-readable and accessible. It&#8217;s the difference between an AI that generates plausible code and an AI that works with <em>your team&#8217;s</em> code.</p><div><hr></div><h2><strong>What context actually means</strong></h2><p>When people hear &#8220;context,&#8221; they think &#8220;give the AI more information.&#8221; That&#8217;s too vague to be useful. Context has specific layers, and they are not equally important.</p><p><strong>Stack declaration</strong> is the foundation. Languages, frameworks, build system, test runner. Pure facts, no judgement, and yet most teams never make it explicit. They assume the AI will figure it out from the code. Sometimes it does. Sometimes it generates Jest tests for a project that uses Vitest, and you spend fifteen minutes wondering why the test suite exploded.</p><p><strong>Conventions</strong> are where things get interesting. Naming patterns, file structure, error handling, import ordering. These are the things that make your codebase feel like <em>yours</em>. Richard Gabriel called this quality &#8220;habitability&#8221; &#8212; the property that makes a codebase feel like a place you can live and work in, rather than a museum you&#8217;re afraid to touch.</p><p><strong>Architectural decisions</strong> are the load-bearing layer and the one most teams skip. Not just &#8220;we use event sourcing&#8221; but &#8220;we use event sourcing because we need a complete audit trail for regulatory compliance and we tried CRUD with audit tables first and it didn&#8217;t scale past 10,000 events per second.&#8221; The <em>why</em> matters because it packs in <em>intent</em>. It tells the AI what it cannot change, even if it sees a &#8220;better&#8221; pattern. Skip this layer and the AI will cheerfully refactor away your most important constraints. Ever been tempted to delete the tests to make a build pass? An AI agent will do that in passing if it&#8217;s not aware its a bad move.</p><p><strong>Rationale</strong> is what turns arbitrary-looking rules into defensible ones. Not just &#8220;use snake_case&#8221; but &#8220;we use snake_case because our Python and Go codebases share type definitions across a code generation boundary, and snake_case is the only casing that round-trips cleanly through our protobuf pipeline.&#8221; Without rationale, the AI treats conventions as suggestions. With rationale, it treats them as constraints.</p><p><strong>Threat model</strong> shapes everything else. What data is sensitive? Where are the trust boundaries? Your senior developers instinctively validate user input at the API boundary and never log authentication tokens. Not because someone told them to last week, but because they remember the incident that taught the team. Your AI doesn&#8217;t remember that incident. You need to encode the lesson.</p><p>Of these five layers, most teams start with stack declaration because it&#8217;s easy. That&#8217;s fine as a starting point. But the leverage is in conventions and architectural decisions. Stack declarations prevent trivial mistakes. Conventions and decisions prevent the expensive ones &#8212; the kind where the code compiles, the tests pass, and the production incident happens three weeks later.</p><blockquote><p><strong>There are no dumb questions:</strong></p><p><em>&#8220;Isn&#8217;t this just documentation?&#8221;</em></p><p>No. Documentation is written for humans who will read it once and mostly, sometimes, remember it. Context documents are written for a reader that starts fresh every session, has no background knowledge, and will interpret ambiguity however it pleases. That changes what you write and how you write it.</p><p><em>&#8220;Can&#8217;t the AI just read our code and figure out the conventions?&#8221;</em></p><p>It can infer some patterns. But it can&#8217;t distinguish between a deliberate convention and an accident that nobody has gotten around to fixing. It can&#8217;t tell whether a pattern appears everywhere because the team chose it or because it was copy-pasted from a Stack Overflow answer three years ago. Explicit context removes the guessing.</p></blockquote><div><hr></div><h2><strong>The precision problem</strong></h2><p>Here&#8217;s where most teams fail, and it&#8217;s subtle enough that your brain is going to want to skip past it.</p><p>Vague conventions produce vague output. This is not a minor annoyance. It is the single biggest reason AI-generated code requires extensive rework. It is the root of slop.</p><p>Look at these two types of convention statements:</p><ul><li><p><em>&#8220;Write clean code&#8221;</em></p><p>becomes: <em>Not enforceable &#8212; break this down into specific rules</em></p></li><li><p><em>&#8220;Keep functions short&#8221;</em></p><p>becomes: <em>Functions must not exceed 40 lines (excluding blank lines and comments)</em></p></li><li><p><em>&#8220;Use meaningful names&#8221;</em></p><p>becomes: <em>Variable names must be at least 3 characters, except for i, j, k, and err</em></p></li><li><p><em>&#8220;Handle errors properly&#8221;</em></p><p>becomes: <em>Every error must be wrapped with fmt.Errorf to add context &#8212; no bare return err</em></p></li><li><p><em>&#8220;Use dependency injection&#8221;</em></p><p>becomes: <em>Functions must not construct their own dependencies &#8212; dependencies are passed in</em></p></li></ul><p>Your brain just read that and filed it under &#8220;obvious.&#8221; It is not obvious. Go back and read the openers again. Those are the conventions most teams actually have.</p><p><em><strong>The test for any convention: could two independent reviewers, looking at the same code, agree on whether it follows the convention -- without discussing it?</strong></em></p><p>&#8220;Write clean code&#8221; fails this test catastrophically. &#8220;Functions must not exceed 40 lines&#8221; passes it trivially. The distance between those two statements is the distance between useful context and noise.</p><blockquote><p><strong>The Sceptic:</strong> &#8220;This feels like a lot of overhead. Can&#8217;t we just tell the AI to follow best practices?&#8221;</p><p><strong>The Veteran:</strong> &#8220;We tried that. The AI&#8217;s &#8216;best practices&#8217; included patterns we explicitly abandoned two years ago after they caused a production outage. Best practices are not your practices.&#8221;</p><p><strong>The Pragmatist:</strong> &#8220;OK, but I&#8217;m not writing a hundred rules on day one. Where do I start?&#8221;</p><p><strong>The Veteran:</strong> &#8220;Start with the things you correct most often. Every time you change something the AI generated, ask yourself: could I have prevented this with a one-sentence rule? Write that sentence down. You&#8217;ll have ten conventions within a week without any special effort.&#8221;</p></blockquote><div><hr></div><h2><strong>The extraction problem (or: you don&#8217;t know what you know)</strong></h2><p>Walk up to a senior developer on your team and ask: &#8220;What are our coding conventions?&#8221;</p><p>They&#8217;ll give you five or six things off the top of their head. Naming. Test structure. Maybe error handling.</p><p>Now watch them do a code review. They&#8217;ll flag twenty things that violate conventions they didn&#8217;t mention because they didn&#8217;t know they knew them. This is <strong>tacit knowledge</strong>: expertise so deeply absorbed that it feels like instinct rather than a rule.</p><p>You can&#8217;t extract tacit knowledge by asking &#8220;what are the rules?&#8221; You extract it by triggering the knowledge indirectly:</p><ul><li><p>&#8220;What would you correct in a new teammate&#8217;s first PR?&#8221;</p></li><li><p>&#8220;What triggers an immediate rejection in code review?&#8221;</p></li><li><p>&#8220;What&#8217;s the thing that&#8217;s obvious to everyone but nowhere in writing?&#8221;</p></li><li><p>&#8220;Which conventions does the AI violate most often?&#8221;</p></li></ul><p>That last question is particularly revealing. It turns every AI interaction into a convention-discovery tool. Every correction you make is a convention you haven&#8217;t encoded yet.</p><p>Ask three senior developers &#8220;what separates a clean refactoring from an over-engineered one?&#8221; and you&#8217;ll get three different answers. Those disagreements are conventions that haven&#8217;t been resolved yet. Surfacing them is one of the most valuable side effects of context engineering: it forces your team to confront things they assumed they already agreed about.</p><blockquote><p><strong>Exercise:</strong> Think of a project you work on. What triggers an immediate rejection in review? Write down your answer. Now imagine asking two other people on your team the same question. Would they give the same answer? If you&#8217;re not sure, that&#8217;s a convention that needs extracting and encoding.</p></blockquote><div><hr></div><h2><strong>Battling context rot: Living documents, not dead wikis</strong></h2><p>You might be thinking: &#8220;Great, I&#8217;ll spend a day writing all this down and we&#8217;re done.&#8221;</p><p>No. Context rots.</p><p>Your conventions reference functions that get renamed. Your stack declaration lists framework versions that get upgraded. Your architectural decisions describe constraints that get relaxed. If your context documents don&#8217;t evolve with the code, they become actively harmful &#8212; the AI follows outdated rules that produce code that <em>used to be</em> correct.</p><p><strong>Context documents must be versioned alongside the code they describe.</strong> They live in the repository, not in a wiki. They get reviewed in pull requests. When someone changes a convention, the context document changes in the same commit.</p><p>This is what separates context engineering from documentation. Documentation lives adjacent to the workflow. Context lives <em>inside</em> the workflow. It&#8217;s checked in, reviewed, and maintained with the same rigour as the code itself.</p><p>The moment your context files live in a wiki, they&#8217;re dead. They&#8217;ll be accurate for about two sprints. After that, they&#8217;ll be worse than having no context at all, because the AI will follow them confidently in the wrong direction.</p><blockquote><p><strong>The Pragmatist:</strong> &#8220;What do I actually do on Monday?&#8221;</p><p>Create a single file in the root of your repository. Give it three sections: stack (what you use), conventions (how you use it), and decisions (why you use it that way). Start with five conventions -- the five things you correct most often. Commit it. Review it as a team.</p><p>That&#8217;s it. That&#8217;s the whole first step.</p></blockquote><div><hr></div><h2><strong>The compound effect</strong></h2><p>A single well-written convention saves you maybe two minutes per AI session. You correct the output, sigh, and move on.</p><p>But you&#8217;re running ten AI sessions a day. That&#8217;s twenty minutes. Across a five-person team, that&#8217;s over an hour and a half of daily rework and that&#8217;s just one missing convention. Most teams have dozens.</p><p>The maths matters less than the dynamic it creates. Each convention you encode is a correction you never make again. Each architectural decision you document is a structural mistake the AI never attempts. That part is straightforward.</p><p>Here&#8217;s what people miss: as the AI produces code that increasingly matches your team&#8217;s style, <em>trust increases</em>. And trust changes behaviour. A developer who trusts the AI&#8217;s output stops re-reading every generated function line by line. They delegate larger tasks. They use AI for work they previously wouldn&#8217;t have attempted &#8212; the boring migration, the tedious test scaffolding, the refactoring they never had time for. The team&#8217;s capacity expands not because the AI got smarter, but because the environment got richer.</p><p>Bad context is a tax you pay forever. Same corrections, every session, no compounding.</p><div><hr></div><h2><strong>What you now know</strong></h2><ul><li><p><strong>Context is layered:</strong> stack, conventions, architectural decisions, rationale, and threat model. The leverage is in the middle layers, not the easy ones.</p></li><li><p><strong>Precision is everything:</strong> A convention that two reviewers can&#8217;t independently agree on is not a convention. It&#8217;s a wish.</p></li><li><p><strong>Tacit knowledge is the bottleneck:</strong> Your team knows far more than they can articulate. Structured extraction questions surface what direct questions miss.</p></li><li><p><strong>Context documents are living artefacts:</strong> Versioned with the code, reviewed in PRs. The moment they live in a wiki, they start dying.</p></li><li><p><strong>Compounding works in both directions:</strong> Encoded conventions build trust that expands capability. Missing conventions create rework that erodes it.</p></li></ul><p>The <a href="https://engineeringagents.substack.com/p/the-habitat-hypothesis">previous article</a> told you that every AI coding failure is an environment problem. Now you know what the environment is made of. In the next article, we&#8217;ll look at what makes conventions <em>enforceable</em> because writing them down is only half the battle.</p><div><hr></div><p><em>Next in the series: &#8220;Constraints That Bite&#8221; &#8212; why conventions without enforcement are just suggestions, and how to build guardrails with teeth.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Engineering Agents! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[The Habitat Hypothesis]]></title><description><![CDATA[Why every AI (and human) coding failure is an environment problem]]></description><link>https://engineeringagents.substack.com/p/the-habitat-hypothesis</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/the-habitat-hypothesis</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Wed, 08 Apr 2026 10:24:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!WBXw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3893301d-0049-4765-842d-a02b33e68268_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WBXw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3893301d-0049-4765-842d-a02b33e68268_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WBXw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3893301d-0049-4765-842d-a02b33e68268_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!WBXw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3893301d-0049-4765-842d-a02b33e68268_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!WBXw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3893301d-0049-4765-842d-a02b33e68268_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!WBXw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3893301d-0049-4765-842d-a02b33e68268_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WBXw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3893301d-0049-4765-842d-a02b33e68268_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3893301d-0049-4765-842d-a02b33e68268_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2846347,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/193558821?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3893301d-0049-4765-842d-a02b33e68268_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!WBXw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3893301d-0049-4765-842d-a02b33e68268_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!WBXw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3893301d-0049-4765-842d-a02b33e68268_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!WBXw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3893301d-0049-4765-842d-a02b33e68268_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!WBXw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3893301d-0049-4765-842d-a02b33e68268_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="pullquote"><p>This is Part 1 of a 6 part series on Harness Engineering as part of Habitat Thinking to accompany the open source AI Literacy Superpowers plugin <br>(<a href="https://habitat-thinking.github.io/ai-literacy-superpowers/">docs</a>, <a href="https://github.com/Habitat-Thinking/ai-literacy-superpowers">repo</a>).</p></div><p>It&#8217;s Tuesday afternoon. You&#8217;re pairing with your AI coding assistant, and you ask it to add a new endpoint to your API.</p><p>It comes back with code that:</p><ul><li><p>Uses CamelCase when your entire codebase is snake_case</p></li><li><p>Puts the route handler in a file called <code>utils.py</code> (your team has a strict controller pattern)</p></li><li><p>Skips input validation entirely</p></li><li><p>Uses a database connection pattern you deprecated six months ago</p></li></ul><p>You stare at the diff. You highlight the whole thing and delete it. You mutter something unprintable. And then you do what everyone does: you open a new chat, write a longer prompt, and try again.</p><p>What you did next was wrong.</p><p>Not the deleting part. That was fine. The <em>trying again with a better prompt</em> part. That&#8217;s the wrong fix, and it&#8217;s the wrong fix almost every time.</p><div><hr></div><h2><strong>The Wrong Diagnosis</strong></h2><p>When AI produces bad output, we reach for one of three responses:</p><ol><li><p><strong>Write a better prompt.</strong> More detail. More examples. A twelve-paragraph system message that reads like a legal contract.</p></li><li><p><strong>Switch to a better model.</strong> Maybe GPT-4 will get it. Maybe Claude will get it. Maybe the next release will get it.</p></li><li><p><strong>Give up on AI for &#8220;real work.&#8221;</strong> Use it for boilerplate and throwaway scripts, but keep it away from anything that matters.</p></li></ol><p>All three treat the AI as the variable. The model is too dumb. The prompt wasn&#8217;t specific enough. The technology isn&#8217;t ready yet.</p><p>We sometimes treat humans the same way&#8230;</p><p>The AI is not the lever. <em>The environment is the lever.</em></p><blockquote><p><strong>Brain check:</strong> Before you read the next section, ask yourself &#8212; when a new developer joins your team and writes code that doesn&#8217;t follow your conventions, do you blame the developer? Or do you look at your onboarding materials, your code review process, and your documentation?</p></blockquote><div><hr></div><h2><strong>The Hypothesis</strong></h2><p>Here it is, the claim this entire series, and the <a href="https://github.com/Habitat-Thinking/ai-literacy-superpowers">AI Literacy Superpowers plugin</a>, is built on:</p><p><strong>The quality of AI-generated code is a function of the environment, not the model.</strong></p><p>A mid-tier model operating inside a well-designed environment &#8212; one with clear conventions, enforced constraints, and architectural context &#8212; will consistently outperform a frontier model dropped into a bare repository with nothing but a README that says &#8220;TODO.&#8221;</p><p>You just read that and your brain filed it under &#8220;obvious.&#8221; It is not obvious. If it were obvious, you wouldn&#8217;t be writing twelve-paragraph prompts. You&#8217;d be fixing your repo. So let me say it differently:</p><p><strong>The environment is the product.</strong> Not the model. Not the prompt. The environment.</p><p>When you invest in prompt engineering, you&#8217;re optimising a single interaction. When you invest in environment engineering, you&#8217;re optimising every interaction from now on.</p><p>One of those compounds. The other doesn&#8217;t.</p><div><hr></div><h2><strong>But What IS an &#8220;Environment&#8221;?</strong></h2><p>When most people hear &#8220;environment,&#8221; they think of their IDE, their terminal, maybe their <code>.env</code> file. That&#8217;s not what we mean.</p><p>The <strong>development environment</strong> is the total context available to your AI agents <em>and</em> humans when working in your codebase:</p><ul><li><p><strong>Conventions</strong> &#8212; naming patterns, file structure, preferred and banned patterns</p></li><li><p><strong>Constraints</strong> &#8212; rules that are <em>enforced</em>, not just documented. Linting rules. Pre-commit hooks. CI checks that reject non-conforming code before it merges</p></li><li><p><strong>Architectural decisions</strong> &#8212; the <em>why</em> behind your code structure. Why you chose this database. Why that function exists even though it looks redundant</p></li><li><p><strong>Accumulated knowledge</strong> &#8212; what mistakes were made and corrected. What patterns emerged over time</p></li><li><p><strong>Feedback loops</strong> &#8212; mechanisms that catch problems and prevent the same mistake from happening twice</p></li></ul><p>Most teams have <em>none of it</em> explicitly available to their people or AI tools.</p><p>They have it in their heads. They have it in Slack threads from 2023. They have it in PR comments that nobody will ever read again. But they don&#8217;t have it anywhere the AI can see.</p><p>And then they wonder why the AI doesn&#8217;t know their conventions.</p><blockquote><p><strong>The Sceptic:</strong> &#8220;This sounds like a lot of overhead. I just want the AI to write code faster. Now you&#8217;re telling me I need to build an entire knowledge base first?&#8221;</p><p><strong>The Pragmatist:</strong> &#8220;You already have the knowledge base. It&#8217;s in your head. The overhead is making it explicit. Which, by the way, also helps every new human teammate you&#8217;ll ever onboard.&#8221;</p><p><strong>The Sceptic:</strong> &#8220;...&#8221;</p><p><strong>The Pragmatist:</strong> &#8220;Yeah. It&#8217;s the same problem.&#8221;</p></blockquote><div><hr></div><h2><strong>The Kitchen That Cooks For You</strong></h2><p>Christopher Alexander spent his career studying how environments shape behaviour. His central insight: a well-designed environment makes the right behaviour natural and the wrong behaviour difficult.</p><p>Think about a well-designed kitchen. The knives are where you reach for them. The spices are at eye level near the stove. The rubbish bin opens with a foot pedal so you can use it with full hands. You don&#8217;t need a manual for this kitchen. The <em>design</em> teaches you.</p><p>Now think about a badly designed kitchen. The knives are in a drawer across the room from the chopping board. The salt is in a cupboard above the fridge. The bin is behind a door. Every meal requires more effort, more mistakes, and more frustration. Not because you&#8217;re a bad cook, but because the environment is fighting you.</p><p>Your codebase is a kitchen. Your AI agents are some of your cooks. And right now, most of us have the salt above the fridge.</p><p>When your AI writes code that doesn&#8217;t follow your conventions, it&#8217;s not because the AI is stupid. It&#8217;s because your conventions aren&#8217;t <em>in the environment</em>. They&#8217;re in your head, in a wiki nobody reads, in tribal knowledge passed down through code review.</p><p>The AI can&#8217;t read your mind. But it can read your environment.</p><blockquote><p><strong>Exercise &#8212; A two-minute audit:</strong> Open the root of your main project right now. Pretend you&#8217;re an AI assistant that&#8217;s just been dropped into this codebase for the first time. What do you know? What conventions are written down? What architectural decisions are documented? What constraints are enforced automatically versus relying on human reviewers to catch?</p><p>If the answer is &#8220;not much&#8221; you&#8217;ve just diagnosed why your AI output isn&#8217;t great.</p></blockquote><div><hr></div><h2><strong>A Fireside Chat: The Model vs. The Environment</strong></h2><blockquote><p><strong>The Model:</strong> Look, I&#8217;m doing my best here. You dropped me into a repository with 400 files, no documentation, and a prompt that says &#8220;add a payment endpoint.&#8221; What do you expect?</p><p><strong>The Environment:</strong> That&#8217;s exactly my point. You&#8217;re a pattern-completion engine. You complete patterns based on context. When the context is thin, you fall back on generic patterns from your training data.</p><p><strong>The Model:</strong> Which is pretty good! I&#8217;ve seen millions of codebases.</p><p><strong>The Environment:</strong> You&#8217;ve seen millions of <em>average, other</em> codebases. So when you get no specific guidance, you produce average code. That&#8217;s not a flaw. That&#8217;s exactly what pattern completion should do. But this team doesn&#8217;t want average. They want <em>their</em> patterns.</p><p><strong>The Model:</strong> So tell me what they are!</p><p><strong>The Environment:</strong> That&#8217;s what I&#8217;m for. When I&#8217;m well designed, I give you conventions, constraints, and feedback loops. I turn you from a generic code generator into a teammate that understands <em>this</em> project.</p><p><strong>The Model:</strong> And when you&#8217;re not well designed?</p><p><strong>The Environment:</strong> Then people blame you. They upgrade to a more expensive version of you. They get the same generic output with slightly better grammar. And they never question whether the bottleneck was me all along.</p></blockquote><div><hr></div><h2><strong>There Are No Dumb Questions</strong></h2><blockquote><p><strong>Q: Does this mean prompting doesn&#8217;t matter at all?</strong></p><p>A: Prompting matters. But it&#8217;s single-use. A great prompt helps one interaction. A great environment helps every interaction. Prompting is giving directions to a taxi driver. Environment design is building the roads.</p><p><strong>Q: Won&#8217;t better models eventually solve this?</strong></p><p>A: Better models will get better at working with <em>whatever context they&#8217;re given</em>. A better model in a great environment will be extraordinary. A better model in a bare repo will still produce generic code. The environment advantage doesn&#8217;t disappear as models improve. It compounds.</p><p><strong>Q: Isn&#8217;t this just &#8220;documentation&#8221;?</strong></p><p>A: Documentation is one input. But documentation that isn&#8217;t enforced is a suggestion. The environment includes enforcement. Constraints that catch violations automatically, feedback loops that learn from mistakes, and context structured so the AI can actually use it. A wiki page about naming conventions is documentation. A pre-commit hook that rejects bad names is environment.</p></blockquote><div><hr></div><h2><strong>What&#8217;s Coming</strong></h2><p>Over the next five articles, we&#8217;ll build up what a well-designed AI development environment actually looks like. From making your conventions machine-readable (Article 2), to constraints that enforce themselves without human reviewers (Article 3), to multi-agent pipelines that create defence in depth (Article 4), to environments that get smarter with every session (Article 5), and finally the full integrated system (Article 6).</p><p>Each article moves from concept to practice. No theory without action.</p><p>But for now, you have the hypothesis. And you have the two-minute audit. That&#8217;s enough to start.</p><div><hr></div><h2><strong>The Bullet Points</strong></h2><p>For us humans, here&#8217;s a summary.</p><ul><li><p><strong>The common instinct</strong> &#8212; when AI writes bad code, we blame the model, write longer prompts, or give up on AI for serious work.</p></li><li><p><strong>The wrong diagnosis</strong> &#8212; all three responses treat the AI as the variable, but the AI is a pattern-completion engine working with whatever context it has.</p></li><li><p><strong>The environment hypothesis</strong> &#8212; the quality of AI output is a function of the environment, not the model. A well-designed environment with a cheaper model beats a bare repo with a frontier model.</p></li><li><p><strong>What &#8220;environment&#8221; means</strong> &#8212; conventions, constraints, architectural context, accumulated knowledge, and feedback loops.</p></li><li><p><strong>The key shift</strong> &#8212; stop investing in single prompts. Start investing in the environment. One compounds. The other doesn&#8217;t.</p></li></ul><div><hr></div><p><em>Next in the series: <a href="https://engineeringagents.substack.com/p/engineering-the-context">&#8220;Engineering Context: Teaching Your AI What You Know&#8221;</a>  where we get specific about making your conventions and context available to your AI assistant.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Engineering Agents! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[What Your AI Coding Tool Can (and Can’t) Do For You - March 2026, Claude and GH Copilot Edition]]></title><description><![CDATA[The difference between engineering agents and engineering with agents &#8212; and why it changes what you might choose to buy]]></description><link>https://engineeringagents.substack.com/p/what-your-ai-coding-tool-can-and</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/what-your-ai-coding-tool-can-and</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Wed, 01 Apr 2026 08:47:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!9fVv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07f4b09-8816-49f0-aa34-4ba921b45c20_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9fVv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07f4b09-8816-49f0-aa34-4ba921b45c20_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9fVv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07f4b09-8816-49f0-aa34-4ba921b45c20_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!9fVv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07f4b09-8816-49f0-aa34-4ba921b45c20_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!9fVv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07f4b09-8816-49f0-aa34-4ba921b45c20_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!9fVv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07f4b09-8816-49f0-aa34-4ba921b45c20_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9fVv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07f4b09-8816-49f0-aa34-4ba921b45c20_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d07f4b09-8816-49f0-aa34-4ba921b45c20_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3173703,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/192823746?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07f4b09-8816-49f0-aa34-4ba921b45c20_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9fVv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07f4b09-8816-49f0-aa34-4ba921b45c20_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!9fVv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07f4b09-8816-49f0-aa34-4ba921b45c20_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!9fVv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07f4b09-8816-49f0-aa34-4ba921b45c20_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!9fVv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07f4b09-8816-49f0-aa34-4ba921b45c20_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>There&#8217;s a distinction that doesn&#8217;t get made often enough, and it&#8217;s costing teams time, money, and a slow erosion of trust in their tools.</p><p><strong>Engineering agents</strong> is the discipline of building AI agents &#8212; writing the code, designing the architecture, hooking up the tools, deciding what an agent can and cannot do. It&#8217;s what this blog mostly lives in. Embabel, MCP servers, tool definitions, memory strategies, bounded autonomy. Code over promises.</p><p><strong>Engineering </strong><em><strong>with</strong></em><strong> agents</strong> is different. It&#8217;s the discipline of working alongside AI agents as a software engineer &#8212; using them as collaborators in your daily development practice. Writing better prompts. Verifying AI output. Building the harness of constraints that stops an agent doing something expensive and irreversible at 3am. Progressing from &#8220;I use Copilot for autocomplete&#8221; to &#8220;I have a multi-agent pipeline that writes, reviews, tests, and integrates code while I sleep.&#8221;</p><p>Both disciplines matter. But they require different skills, different tools, and &#8212; critically &#8212; different mental models.</p><p>If you show up to engineering <em>with</em> agents using only the skills of someone who has never thought seriously about what an agent <em>is</em>, you&#8217;ll get autocomplete with delusions of grandeur. And if you try to engineer agents without understanding how to collaborate with them day-to-day, you&#8217;ll build systems that nobody &#8212; including you &#8212; trusts to run.</p><p>This article is squarely about the second discipline: <strong>engineering with agents</strong>, and what the tools you already use can actually do to support your progression.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://engineeringagents.substack.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The Framework</h2><p>What follows is a capability assessment mapped across six levels of AI collaboration literacy &#8212; from &#8220;I know AI exists&#8221; to &#8220;I run autonomous agent teams across my organisation.&#8221; The two tools under the microscope are <strong>Claude Code</strong> (Anthropic&#8217;s terminal-native agentic coding assistant) and <strong>GitHub Copilot</strong> (the full suite: inline suggestions, agent mode, coding agent, CLI).</p><p>Where relevant, we describe what any AI coding tool would need to support at each level; a generic requirement that&#8217;s useful both for practitioners evaluating future tools and for vendors looking honestly at their roadmap.</p><p><em>This assessment reflects the capability landscape as of March 2026. Both tools ship frequently; gaps marked here may close.</em></p><div><hr></div><h2>Level 0: The Aware</h2><p>At the base level, the requirement is simply a working interface and enough transparency to know what model you&#8217;re talking to. Both tools clear this bar. Claude Code gives you <code>/model</code> to see the current model and <code>/cost</code> to track your spend; Copilot shows the model in VS Code, though its CLI is less forthcoming. No meaningful gaps here.</p><p>The more important observation is that Level 0 is primarily a mindset, not a tool capability. The question is whether you&#8217;ve started asking <em>why</em> the AI said what it said &#8212; not just whether you liked the output.</p><div><hr></div><h2>Level 1: The Prompter</h2><p>Level 1 is where practitioners start to develop deliberate habits: translating intent into prompts, curating context consciously, and developing an awareness of what the context window costs.</p><p>Both tools support natural language interaction well &#8212; Claude Code via the terminal with multi-turn conversation, Copilot via the chat panel, inline suggestions, and agent mode. That part is table stakes.</p><p>The gap that matters is <strong>context engineering</strong>. Claude Code supports a layered context system: <code>CLAUDE.md</code> files at the global, project, and directory levels, supplemented by skills and <code>AGENTS.md</code>. This means you can encode context that is always true, context that is true for this repository, and context that is true for this corner of the codebase &#8212; and the agent reads the right layers automatically.</p><p>Copilot offers a single <code>.github/copilot-instructions.md</code> file at the project level. It works. But a single flat file versus a layered, composable system is the difference between &#8220;I told the AI what I want once&#8221; and &#8220;the AI always knows where it is and what matters here.&#8221; One is a hope. The other is an architecture.</p><p>Cost and token visibility tell a similar story. Claude Code&#8217;s <code>/cost</code> command and Analytics API give you per-session visibility. Copilot&#8217;s usage metrics live at the org level in an admin dashboard &#8212; useful for governance, but not for the individual practitioner developing the habit of token consciousness.</p><div><hr></div><h2>Level 2: The Verifier</h2><p>Level 2 is about trust &#8212; specifically, building the automated infrastructure that lets you verify AI output systematically rather than eyeballing it.</p><p>Code execution is supported by both tools. Claude Code&#8217;s Bash tool runs any command; Copilot&#8217;s agent mode runs terminal commands in VS Code. That&#8217;s the easy part.</p><p>The important gap is <strong>lifecycle hooks</strong>. Claude Code supports <code>PreToolUse</code>, <code>PostToolUse</code>, and <code>FileChanged</code> hooks with <code>allow</code>, <code>deny</code>, and <code>ask</code> control flow. This means you can intercept what the agent is about to do <em>before</em> it does it &#8212; warn, block, or prompt for confirmation based on what tool is being called or what file is being touched. Copilot has no equivalent. Its workaround is GitHub Actions, which enforces checks at PR time. That&#8217;s not nothing &#8212; PR gates are real gates &#8212; but they fire after the agent has already acted, not during.</p><p>Code review is one area where Copilot has a genuine native strength: <code>@copilot</code> as a PR reviewer on GitHub is turnkey and fits naturally into the workflows most teams already use. Claude Code&#8217;s review capability is built through a custom code-reviewer agent, which is more powerful but requires more setup.</p><p>Permission control is more nuanced with Claude Code &#8212; fine-grained allow/deny lists per tool, path patterns, and domain restrictions. Copilot&#8217;s <code>.copilotignore</code> handles content exclusions but doesn&#8217;t reach to per-tool permission control.</p><div><hr></div><h2>Level 3: The Habitat Engineer</h2><p>Level 3 is where you stop using an AI tool and start <em>building a development environment that includes AI as a first-class component</em>. This is the level where the gap between the two tools becomes significant.</p><p>The core concept here is the <strong>harness</strong>: a set of living documents and enforcement mechanisms that declare what the AI can and cannot do, what it knows about your project, and how it should behave across sessions. Claude Code supports this through <code>HARNESS.md</code> combined with hooks and CI integration &#8212; <code>PreToolUse</code> hooks warn during development, CI blocks at the pipeline. There is no equivalent concept in Copilot. Its closest analogue is <code>.github/copilot-instructions.md</code> serving as advisory context, with GitHub Actions as the enforcement mechanism downstream.</p><p>Custom agent definitions are another substantial gap. Claude Code lets you define specialised agents in <code>.claude/agents/*.md</code> with YAML frontmatter specifying per-agent tool lists and model selection. A spec-writing agent gets different tools and a different model than a code-review agent. Copilot has no equivalent &#8212; agents are differentiated through prompt engineering within a single chat interface.</p><p>Parallel agent isolation, where multiple agents work safely on different branches of a codebase simultaneously, is built into Claude Code through git worktrees &#8212; the <code>--worktree</code> flag and <code>isolation: worktree</code> in agent definitions handle this automatically. In Copilot, you&#8217;d manage separate worktrees manually and run separate chat sessions yourself.</p><p>Model routing &#8212; the ability to choose the right model tier for the right task &#8212; is supported in Claude Code through the <code>model:</code> field in agent definitions and the <code>MODEL_ROUTING.md</code> pattern. Copilot uses a single model per session, with no routing mechanism inside the tool.</p><p>Compound learning is the last significant gap at this level. Claude Code supports <code>AGENTS.md</code> for curated project knowledge, <code>REFLECTION_LOG.md</code> for agent-proposed learning, and auto-memory &#8212; mechanisms for accumulating what the agent has learned about your project across sessions. Copilot has no cross-session learning mechanism. When the session ends, it forgets.</p><p>MCP tool extension is partially supported by both tools, but differently. Claude Code integrates MCP servers via <code>.mcp.json</code> with autonomous tool discovery; Copilot&#8217;s Extensions provide tool access, but they&#8217;re chat-invoked rather than available to the agent autonomously.</p><div><hr></div><h2>Level 4: The Specification Architect</h2><p>At Level 4, specifications become the source of truth and the agent pipeline is a first-class engineering artifact &#8212; with safety gates, orchestration, and programmatic access.</p><p>Multi-agent orchestration is fully supported in Claude Code: an orchestrator dispatches a spec-writer, then a TDD agent, then implementers, then a code reviewer, then an integration agent, each with scoped tools and bounded responsibility. Copilot has no equivalent; individual coding agent tasks can be assigned, but there is no pipeline definition or dependency management between multiple agent roles.</p><p>Safety gates are similarly absent from Copilot&#8217;s current architecture. Claude Code supports plan approval gates &#8212; the orchestrator pauses for human review before proceeding &#8212; and <code>MAX_REVIEW_CYCLES</code> as a guardrail against runaway loops. Copilot&#8217;s workaround is branch protection rules at merge time, which is a gate, but not a gate inside the autonomous workflow.</p><p>Async cloud execution is an area of genuine parity. Claude Code supports headless mode and the Agent SDK; Copilot&#8217;s Coding Agent handles the &#8220;assign an issue, get a PR&#8221; workflow natively, and for teams already living in GitHub this is a meaningful capability that maps naturally onto existing practice.</p><p>Programmatic integration tells a more differentiated story. Claude Code offers a headless CLI (<code>claude -p</code>), a Python SDK, and a TypeScript SDK &#8212; the full agent loop is accessible programmatically. Copilot&#8217;s CLI handles command suggestions, but there&#8217;s no SDK that exposes the full agentic workflow to external scripts or CI pipelines.</p><p>Scheduled autonomous tasks are supported in Claude Code via the <code>/schedule</code> command and desktop and cloud task scheduling. Copilot has no equivalent; the workaround is GitHub Actions on a cron schedule, which executes workflow steps but doesn&#8217;t invoke the agent directly.</p><div><hr></div><h2>Level 5: The Sovereign Engineer</h2><p>At Level 5, AI collaboration becomes an engineering discipline with the same governance requirements as any other platform capability &#8212; telemetry, standards, distribution, audit.</p><p>OpenTelemetry export is a clean differentiator here. Claude Code&#8217;s native OTel export integrates with Grafana, Datadog, and Honeycomb, correlating AI telemetry with the rest of your observability stack. Copilot has no OTel export; the Copilot Metrics API provides org-level analytics on acceptance rates, active users, and language breakdown, which is useful for adoption tracking but disconnected from infrastructure observability.</p><p>Usage analytics are reasonably strong in both tools. Claude Code provides daily aggregates with per-user token usage and cost estimates via the Analytics API. Copilot&#8217;s Metrics API covers adoption signals. Claude Code is more granular on cost; Copilot is more naturally surfaced for GitHub-native org reporting.</p><p>Agent team coordination &#8212; multiple agent sessions sharing state, communicating, and managing dependencies &#8212; is in experimental support in Claude Code through Agent Teams (team lead, teammates, shared task list, peer messaging). Copilot has no multi-session coordination mechanism.</p><p>Reusable harness templates &#8212; the ability to package and distribute organisational AI standards as installable artifacts &#8212; are supported in Claude Code through its plugin system, which bundles skills, agents, hooks, and templates together. Copilot&#8217;s workaround is GitHub template repos containing <code>.github/copilot-instructions.md</code>, which gets you the instructions file but none of the enforcement, hooks, or agent definitions.</p><div><hr></div><h2>The Pattern</h2><p>Both tools serve Levels 0 to 2 well. If you&#8217;re in the early stages of building an AI collaboration practice, either tool will get you started. Copilot&#8217;s deep GitHub integration &#8212; PR review, coding agent, metrics &#8212; is a genuine strength for teams whose world is already organised around GitHub. The async coding agent in particular is worth taking seriously: assign an issue, get a PR, stay in your existing workflow.</p><p>Claude Code&#8217;s strength is its agentic infrastructure: hooks, custom agents, plugins, orchestration, OTel. The higher levels &#8212; L3 through L5 &#8212; require capabilities that are currently unique to Claude Code&#8217;s architecture.</p><p>It&#8217;s important to point out that these are <strong>capabilities, not permanent advantages</strong>. Any tool that adds persistent layered context files, programmable lifecycle hooks, custom agent definitions, and standard telemetry export can serve this full progression. The vendors are watching the same curve the practitioners are.</p><div><hr></div><h2>The Row That Matters</h2><p>This is a capability assessment, not a product ranking. The framework&#8217;s levels are tool-agnostic; the capabilities required at each level are not.</p><p>The most important level in this assessment isn&#8217;t the one where your current tool excels. It&#8217;s the one where it has gaps &#8212; because that&#8217;s where your progression will hit a ceiling. Whether that ceiling requires workarounds, tool evolution, or a tool change, you want to know about it before you arrive at it.</p><p>Plan for it before you reach it.</p><div><hr></div><p><em><strong>Don&#8217;t agree with the assessment here, please reach out! I&#8217;m always keen to hear from the lived experience of software engineers pushing the limits of these tools, so please feel free to reach out to me here or on LinkedIn.</strong></em></p><div><hr></div><p><em>Russ Miles is a software engineering thought leader, keynote speaker, and the Festival Director of <a href="https://eastbournelitfest.org.uk/">Eastbourne LitFest</a>. Engineering Agents is his hands-on practitioner blog: code over promises, always.</em></p><p><em>Found this useful? Hit the subscribe button &#8212; and if you&#8217;re building something interesting in the agent space, get in touch. This blog actively wants more voices in it.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Engineering Agents! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Phase 0 Agent Design: Before the Agent, There Is a World]]></title><description><![CDATA[An agent without a world model is not intelligent &#8212; it is theatrical, and it is prone to failure]]></description><link>https://engineeringagents.substack.com/p/phase-0-agent-design-before-the-agent</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/phase-0-agent-design-before-the-agent</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Wed, 17 Dec 2025 10:20:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!AoQV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa67adc37-67ae-4b91-b00a-14cc406269a6_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!AoQV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa67adc37-67ae-4b91-b00a-14cc406269a6_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!AoQV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa67adc37-67ae-4b91-b00a-14cc406269a6_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!AoQV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa67adc37-67ae-4b91-b00a-14cc406269a6_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!AoQV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa67adc37-67ae-4b91-b00a-14cc406269a6_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!AoQV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa67adc37-67ae-4b91-b00a-14cc406269a6_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!AoQV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa67adc37-67ae-4b91-b00a-14cc406269a6_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a67adc37-67ae-4b91-b00a-14cc406269a6_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3078974,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/181774632?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa67adc37-67ae-4b91-b00a-14cc406269a6_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!AoQV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa67adc37-67ae-4b91-b00a-14cc406269a6_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!AoQV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa67adc37-67ae-4b91-b00a-14cc406269a6_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!AoQV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa67adc37-67ae-4b91-b00a-14cc406269a6_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!AoQV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa67adc37-67ae-4b91-b00a-14cc406269a6_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This article follow on from the Agent Design Maturity Ladder:</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;0f6b602f-48e4-4d07-bc5a-e5f0249f21ff&quot;,&quot;caption&quot;:&quot;&#8220;Begin not with the machine, but with the world you ask it to enter.&#8221;&quot;,&quot;cta&quot;:&quot;Read full story&quot;,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;On Agent Design - From Promptling to Autonomous Collaborator&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:9890843,&quot;name&quot;:&quot;Russ Miles&quot;,&quot;bio&quot;:&quot;Software Builder, Listener, Reader, Writer, Speaker (in that order)&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d72c4dc7-ed41-445c-a745-3822932c2299_1282x1284.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2025-12-15T11:54:35.234Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!0dTn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1e9bc66-0753-4244-a71a-6f7b1f508e68_1536x1024.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://engineeringagents.substack.com/p/on-agent-design-from-promptling-to&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:181668924,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:1,&quot;publication_id&quot;:5479520,&quot;publication_name&quot;:&quot;Engineering Agents&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!eRFm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0861dada-274e-4eb5-8137-e122bd7a4c3d_1024x1024.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div><hr></div><p>Most teams skip Phase 0 of Agent Design because it feels like delay. They want prompts. They want tools. They want visible progress.</p><p>They want to get their hands dirty and as agent toolkits advance the pseudo sense of progress can be intoxicating. Until you hit the wall.</p><p>At that point what they get instead is an agent that behaves <em>confidently wrong</em> &#8212; fluent in language, ignorant of reality, and dangerous precisely because it sounds competent.</p><p>Phase 0 is about avoiding this wall, and it iis not about agents at all.</p><p>It is about <strong>establishing the world an agent will inhabit</strong> &#8212; its physics, its language, its hazards, and its values &#8212; <em>before</em> you give anything the power to explore and act.</p><p>If Phase 1 is <em>prompting</em>, Phase 0 is <strong>ontology</strong>.</p><p>If Phase 2 is <em>tools</em>, Phase 0 is <strong>permission</strong>.</p><p>If Phase 3 onwards is <em>autonomy</em>, Phase 0 is <strong>responsibility</strong>.</p><p>Skip Phase 0, and every later phase compounds error.</p><div><hr></div><h2><strong>An agent without a world model is not intelligent &#8212; it is theatrical, and it is prone to failure</strong></h2><p>Phase 0 is the <strong>pre-design phase</strong> of agent engineering. Not architecture. Not prompts. Not frameworks.</p><p>It is the disciplined act of answering one question:</p><blockquote><p><em>&#8220;What must be true about the world for this agent to act safely, usefully, and intelligibly?&#8221;</em></p></blockquote><p>This phase produces <strong>constraints, not capabilities</strong>. <strong>Domain</strong> to prevent <strong>danger</strong>.</p><p><strong>Meaning before mechanism</strong>.</p><p>Phase 0 work is done <em>with humans</em>, not models.</p><div><hr></div><h2><strong>Five Pillars of Phase 0 Agent Design</strong></h2><h3><strong>Pillar 1: Define the Domain (Before You Define the Agent)</strong></h3><p>Agents do not operate in &#8220;tasks&#8221;. They operate in <strong>domains</strong> &#8212; structured worlds with rules, entities, and consequences.</p><p>Strong Phase 0 questions include:</p><ul><li><p>What <strong>domain</strong> does this agent exist within?</p></li><li><p>What is <em>in scope</em> &#8212; and what is explicitly <em>out of scope</em>?</p></li><li><p>What entities matter here?</p><p>(People, artefacts, states, approvals, risks, policies)</p></li><li><p>What relationships between those entities are real, causal, and important?</p></li><li><p>What actions are even <em>possible</em> in this world?</p></li></ul><p>If you cannot describe the domain without referencing the agent, you are not ready to build one.</p><div><hr></div><h3><strong>Pillar 2: Establish the Grounded Language of the World</strong></h3><p>LLMs are language machines &#8212; but language without grounding is hallucination.</p><p>Phase 0 asks:</p><ul><li><p>What terms have <strong>precise meaning</strong> in this domain?</p></li><li><p>Where do humans routinely <strong>disagree on definitions</strong>?</p></li><li><p>Which words are overloaded, informal, or context-dependent?</p></li><li><p>What does <em>&#8220;done&#8221;</em>, <em>&#8220;approved&#8221;</em>, <em>&#8220;safe&#8221;</em>, or <em>&#8220;urgent&#8221;</em> actually mean here?</p></li></ul><p>This is where most agent failures are born: the model is fluent, the organisation is ambiguous.</p><p>Phase 0 is where ambiguity is exposed, not hidden behind prompts.</p><div><hr></div><h3><strong>Pillar 3: Clarify Human Intent and Value</strong></h3><p>Agents don&#8217;t create value. They <em>amplify</em> the value logic already present &#8212; good or bad.</p><p>Phase 0 questions include:</p><ul><li><p>Who is this agent for?</p></li><li><p>What human pain or friction are we trying to relieve?</p></li><li><p>What outcome would make a user say: <em>&#8220;This helped.&#8221;</em></p></li><li><p>What outcomes would be technically &#8220;correct&#8221; but <strong>humanly wrong</strong>?</p></li><li><p>How will we measure success <em>without</em> relying on activity metrics?</p></li></ul><p>If you cannot articulate value without mentioning automation, stop. You are, at best, designing motion, not progress.</p><div><hr></div><h3><strong>Pillar 4: Identify Hazards Before Capabilities</strong></h3><p>Phase 0 is where you imagine failure <em>before</em> intelligence is applied.</p><p>Questions to ask include:</p><ul><li><p>What is the <strong>blast radius</strong> if this agent is wrong?</p></li><li><p>What errors are recoverable &#8212; and which are not?</p></li><li><p>What actions should <em>never</em> be taken automatically?</p></li><li><p>What does a <em>graceful failure</em> look like here?</p></li><li><p>When should uncertainty trigger escalation rather than action?</p></li></ul><p>Autonomy is not a reward. It is a <strong>liability you earn the right to carry</strong>.</p><div><hr></div><h3><strong>Pillar 5: Locate Ground Truth</strong></h3><p>Agents must know <em>where reality lives</em>.</p><p>Phase 0 insists on clarity:</p><ul><li><p>What are the authoritative sources of truth?</p></li><li><p>Which sources conflict &#8212; and how do humans resolve that today?</p></li><li><p>What knowledge is static vs temporal vs contextual?</p></li><li><p>What knowledge is <em>implicit</em> in human heads but absent from systems?</p></li></ul><p>If the organisation itself does not know where truth resides, the agent will invent it. And it will sound convincing while doing so.</p><div><hr></div><h2><strong>Phase 0 Agent Design Outputs: Not just Artefacts &#8212; Agreements</strong></h2><p>Done well, Phase 0 produces:</p><ul><li><p>A shared <strong>domain model</strong></p></li><li><p>Explicit <strong>constraints and invariants</strong></p></li><li><p>Clear <strong>non-goals</strong></p></li><li><p>A list of <strong>known ambiguities</strong></p></li><li><p>A definition of <strong>safe failure</strong></p></li><li><p>Agreement on <strong>where humans stay in the loop</strong></p></li></ul><p>None of these are prompts. All of them shape every prompt that follows.</p><div><hr></div><h2><strong>Common Phase 0 Anti-Patterns to Avoid</strong></h2><p>When you hear some of the following, beware:</p><ul><li><p>&#8220;We&#8217;ll figure it out in the prompt.&#8221;</p></li><li><p>&#8220;The model will infer that.&#8221;</p></li><li><p>&#8220;Let&#8217;s just prototype and see.&#8221;</p></li><li><p>&#8220;It&#8217;s just internal &#8212; low risk.&#8221;</p></li><li><p>&#8220;We&#8217;ll add guardrails later.&#8221;</p></li></ul><p>These are not shortcuts. They are <strong>deferred disasters</strong>.</p><div><hr></div><h2><strong>Some practices to consider</strong></h2><ul><li><p>Run Phase 0 as a <strong>facilitated workshop</strong>, not a solo design task</p></li><li><p>Use <strong>real incidents</strong> and failures to surface hazards</p></li><li><p>Force explicit answers to &#8220;What must never happen?&#8221;</p></li><li><p>Write constraints in plain language before encoding them</p></li><li><p>Treat unresolved ambiguity as a blocker, not a footnote</p></li></ul><div><hr></div><h2><strong>Checklist: Are You Ready to Leave Phase 0?</strong></h2><p>You may proceed only if:</p><ul><li><p>The domain can be explained without mentioning the agent</p></li><li><p>Success and failure are both clearly described</p></li><li><p>Hazards are named, not implied</p></li><li><p>Sources of truth are agreed upon</p></li><li><p>Humans understand where responsibility remains theirs</p></li></ul><p>If not, you are not in Phase 1. You are simply rehearsing failure with better tooling.</p><div><hr></div><h2><strong>Phase 0 is not optional</strong></h2><p>Phase 0 feels slow because it is <strong>thinking made visible</strong>. But every agent that behaves well under pressure &#8212; every system that earns trust &#8212; was born here, in restraint, clarity, and humility.</p><p>Before the agent, there must be a world. Before autonomy, there must be understanding. Anything else is theatre.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Engineering Agents! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[On Agent Design - From Promptling to Autonomous Collaborator]]></title><description><![CDATA[How to design AI agents that grow from simple helpers into fully governed, domain-aware, collaborating systems]]></description><link>https://engineeringagents.substack.com/p/on-agent-design-from-promptling-to</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/on-agent-design-from-promptling-to</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Mon, 15 Dec 2025 11:54:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!0dTn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1e9bc66-0753-4244-a71a-6f7b1f508e68_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0dTn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1e9bc66-0753-4244-a71a-6f7b1f508e68_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0dTn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1e9bc66-0753-4244-a71a-6f7b1f508e68_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!0dTn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1e9bc66-0753-4244-a71a-6f7b1f508e68_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!0dTn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1e9bc66-0753-4244-a71a-6f7b1f508e68_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!0dTn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1e9bc66-0753-4244-a71a-6f7b1f508e68_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0dTn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1e9bc66-0753-4244-a71a-6f7b1f508e68_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e1e9bc66-0753-4244-a71a-6f7b1f508e68_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2999609,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/181668924?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1e9bc66-0753-4244-a71a-6f7b1f508e68_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!0dTn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1e9bc66-0753-4244-a71a-6f7b1f508e68_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!0dTn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1e9bc66-0753-4244-a71a-6f7b1f508e68_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!0dTn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1e9bc66-0753-4244-a71a-6f7b1f508e68_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!0dTn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1e9bc66-0753-4244-a71a-6f7b1f508e68_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="pullquote"><p>&#8220;Begin not with the machine, but with the world you ask it to enter.&#8221;</p></div><p>Building agents and their agentic workflows looks deceptively simple on first glance. Create an agent, tell it what to do, give it access to an LLM or two and BANG, you&#8217;re done. Right?</p><p>Well, I wish that were right. But usually that doesn&#8217;t lead to a useful, powerful agent. It leads to an expensive, untrustworthy chaos monkey<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>.</p><p>If you&#8217;re lucky then there is a moment, early in every engineer&#8217;s practice with AI agents, when they feel the same uneasy thrill sailors once felt approaching an unmapped coast. They see potential everywhere&#8212;automation, reasoning, orchestration&#8212;yet the terrain feels unstable. </p><p>What begins as a simple prompt becomes a creature with surprising initiative; what begins as a tool becomes a partner that can make decisions you did not explicitly author. And so the question becomes not merely <em>how to build an agent</em>, but how to <em>shepherd its growth</em>.</p><p>Most engineering disciplines begin with a design blueprint and end with an implementation. Agent engineering is more like raising a child: you <strong>design the environment</strong> before you design the agent, you <strong>add capabilities gradually</strong>, and you <strong>tighten governance</strong> as autonomy increases. A reckless agent is not a product flaw&#8212;it&#8217;s a parental mistake. A misaligned agent is a design failure, not a moral one.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://engineeringagents.substack.com/subscribe?"><span>Subscribe now</span></a></p><p>The old world of software services was about deterministic instructions; this new world is about <strong>bounded cognition</strong>, ensuring that systems which can reason do so inside guardrails that reflect both our goals and our ethics. You don&#8217;t &#8220;deploy an agent&#8221; the way you deploy a microservice. You <strong>introduce</strong> an agent to its domain, <strong>teach</strong> it the world, <strong>give</strong> it tools, <strong>bind</strong> it with invariants, and <strong>invite</strong> it to collaborate with humans and other agents. And with each step you reduce fragility and increase value.</p><p>The progression is clear. When the first agent appears in a team, it is nothing more than a glorified autocomplete&#8212;occasionally helpful, occasionally absurd. But with the right design, that initial spark becomes:</p><ul><li><p>A <strong>tool user</strong> capable of meaningful work</p></li><li><p>A <strong>planner</strong> capable of decomposing tasks</p></li><li><p>A <strong>governed executor</strong> that can act safely</p></li><li><p>A <strong>collaborator</strong> negotiating with peers</p></li><li><p>An <strong>autonomous entity</strong> that monitors, decides, and adapts</p></li></ul><p>Every agent in production today sits somewhere on this ladder. Misplace it, and you risk either giving too much autonomy too early&#8212;or worse, leaving enormous value unrealised because the agent was never allowed to grow.</p><p>This article explored this ladder. A framework for taking an idea of an agent and guiding it, step by step, until it becomes a trustworthy participant in your systems. It balances the clarity of constraints with the creative potential of modern AI and demands discipline because power without boundaries is chaos, yet discipline without creativity is paralysis.</p><p>Above all: <strong>an agent is not defined by its intelligence but by its governance</strong>.</p><div class="pullquote"><p>This article is a snippet of what is explored in the full <a href="https://www.russmiles.com/hands-on-ai-agent-engineering">&#8220;Hands-On AI Agent Engineering for Developers&#8221;</a> workshop</p></div><h2>An Agent Maturity Ladder</h2><p><em>From prompt-driven helpers &#8594; governed, autonomous, collaborating systems</em></p><p>We need a <strong>clear, practical, end-to-end blueprint</strong> for designing an agent &#8212; or an ecosystem of agents &#8212; that can <em>start simple</em> (semi-manual, prompt-driven) and mature into <strong>fully autonomous, domain-aware, collaborating agents</strong>.</p><p>This is the agent maturity ladder.</p><div><hr></div><h3>Phase 0 &#8212; Clarify the Domain (The D in DICE)</h3><p>Before you design an agent, you design its context, its environment, its <em>world</em>.</p><h4>Steps</h4><ol><li><p><strong>Define the domain</strong></p><ul><li><p>What is this domain&#8217;s language?</p></li><li><p>Where is the domain&#8217;s boundary?</p></li><li><p>What objects already exist here?</p></li><li><p>What actions can be taken?</p></li><li><p>What properties matter?</p></li><li><p>What is out of bounds?</p></li></ul></li><li><p><strong>Define the user needs</strong></p><ul><li><p>What problems should this agent solve?</p></li><li><p>What pains or bottlenecks should it eliminate?</p></li></ul></li><li><p><strong>Define the desired outcomes</strong></p><ul><li><p>What does &#8220;value delivered&#8221; look like?</p></li><li><p>What should we measure?</p></li></ul></li><li><p><strong>Define hazards and failure modes</strong></p><ul><li><p>Where could this agent cause harm?</p></li><li><p>What is the blast radius?</p></li></ul></li></ol><p>This becomes the <em>grounding substrate</em> for safe autonomy later.</p><div><hr></div><h3>Phase 1 &#8212; Prompt-Driven, Single-Agent (Assisted Mode)</h3><p><em>No autonomy, pure suggestion. The agent cannot act, only support.</em></p><h4>Purpose</h4><ul><li><p>Explore feasibility</p></li><li><p>Learn the domain</p></li><li><p>Prototype workflows quickly</p></li></ul><h4>Steps</h4><ol><li><p><strong>Write prompts that express the behaviours you want</strong></p><ul><li><p>&#8220;Given X and Y, produce Z&#8230;&#8221;</p></li></ul></li><li><p><strong>Give examples of correct behaviour</strong></p></li><li><p><strong>Clarify unacceptable behaviour</strong></p></li><li><p><strong>Use retrieval (RAG) for contextual domain grounding to the LLM</strong></p><ul><li><p>Docs</p></li><li><p>Policies</p></li><li><p>Past task examples</p></li></ul></li></ol><h4>Outputs</h4><ul><li><p>Initial persona of the agent</p></li><li><p>Early task patterns</p></li><li><p>Understanding of what the agent should eventually <em>do</em></p></li></ul><div><hr></div><h3>Phase 2 &#8212; Tool-Using Prompt Agent (Controlled Capability)</h3><p>Still no autonomy, but now the agent can <em>do things</em> when asked.</p><h4>Steps</h4><ol><li><p><strong>Define allowed tools</strong></p><ul><li><p>APIs, including MCP</p></li><li><p>Shell commands</p></li><li><p>Database queries</p></li><li><p>Internal platform operations</p></li></ul></li><li><p><strong>Implement function schemas</strong></p><ul><li><p>Inputs, outputs, constraints</p></li></ul></li><li><p><strong>Define the environment context</strong></p><ul><li><p>Working context</p></li><li><p>Credentials</p></li><li><p>Permissions</p></li></ul></li><li><p><strong>Wrap prompts in structured system instructions</strong></p><ul><li><p>&#8220;You may only call tools listed below&#8230;&#8221;</p></li><li><p>&#8220;Never execute destructive actions without a plan.&#8221;</p></li></ul></li></ol><h4><strong>Purpose</strong></h4><p>Demonstrate the agent can act safely under human instruction.</p><div><hr></div><h3>Phase 3 &#8212; Add Planning (The P in GOAP)</h3><p>The agent transitions from &#8220;tool caller&#8221; to &#8220;planner and executor&#8221;.</p><h3>Steps</h3><ol><li><p><strong>Define actions</strong></p><ul><li><p>Preconditions</p></li><li><p>Effects</p></li><li><p>Cost/value</p></li></ul></li><li><p><strong>Define world state</strong></p><ul><li><p>What does the agent believe is true?</p></li><li><p>How does it verify or update state?</p></li></ul></li><li><p><strong>Add planning loop</strong></p><ul><li><p>Observe &#8594; Plan &#8594; Act &#8594; Evaluate</p></li></ul></li><li><p><strong>Introduce explainability</strong></p><ul><li><p>Log decision traces</p></li><li><p>Produce rationale for each step</p></li></ul></li></ol><h4>Purpose</h4><p>The agent can now split large tasks into coherent steps and reason about order, dependencies, and consequences.</p><div><hr></div><h3>Phase 4 &#8212; Introduce Invariants &amp; Safety</h3><p>Before autonomy, you must create guardrails at the <em>structural</em> level.</p><h4>Steps</h4><ol><li><p><strong>Define invariants</strong></p><ul><li><p>Must always be true</p></li><li><p>Must never be violated</p></li></ul></li><li><p><strong>Define forbidden actions</strong></p></li><li><p><strong>Define escalation rules</strong></p><ul><li><p>&#8220;If uncertainty &gt; X, ask a human.&#8221;</p></li></ul></li><li><p><strong>Define review workflows</strong></p><ul><li><p>Summaries</p></li><li><p>Logs</p></li><li><p>Sign-offs</p></li></ul></li><li><p><strong>Add observability hooks</strong></p><ul><li><p>Every decision = structured event</p></li><li><p>Every action = auditable</p></li></ul></li></ol><h4>Purpose</h4><p>You make the <em>agent&#8217;s boundaries</em> explicit, testable, and explainable.</p><div><hr></div><h3>Phase 5 &#8212; Semi-Autonomous Mode (Governed Execution)</h3><p>The agent now acts independently <em>within strict boundaries</em>.</p><h4>Steps</h4><ol><li><p><strong>Add triggers</strong></p><ul><li><p>Time-based</p></li><li><p>Event-based</p></li><li><p>State-based</p></li></ul></li><li><p><strong>Allow self-initiation</strong></p><ul><li><p>&#8220;If logs show X, start Y workflow.&#8221;</p></li></ul></li><li><p><strong>Limit autonomy by domain</strong></p><ul><li><p>Narrow, well-defined tasks</p></li><li><p>Clear policies</p></li></ul></li><li><p><strong>Require all actions to satisfy invariants</strong></p></li><li><p><strong>Enable audit and replay</strong></p><ul><li><p>Decisions are traceable</p></li><li><p>History is inspectable</p></li></ul></li></ol><h4>Purpose</h4><p>The agent &#8220;does work&#8221; without waiting for a human, but <em>cannot</em> go outside domain or violate safety.</p><div><hr></div><h3>Phase 6 &#8212; Multi-Agent Collaboration</h3><p>Once one agent works, collaborate with others if needed.</p><h4>Steps</h4><ol><li><p><strong>Define roles</strong></p><ul><li><p>Researcher</p></li><li><p>Planner</p></li><li><p>Executor</p></li><li><p>Reviewer</p></li><li><p>Coordinator</p></li></ul></li><li><p><strong>Define message protocol</strong></p><ul><li><p>Shared memory</p></li><li><p>API passing</p></li><li><p>Structured messages</p></li></ul></li><li><p><strong>Delegate tasks between agents</strong></p><ul><li><p>Planner &#8594; Worker</p></li><li><p>Worker &#8594; Validator</p></li><li><p>Validator &#8594; User</p></li></ul></li><li><p><strong>Implement arbitration</strong></p><ul><li><p>Tie-break rules</p></li><li><p>Conflict detection</p></li><li><p>Confidence scoring</p></li></ul></li><li><p><strong>Emergent intelligence</strong></p><ul><li><p>Teams of agents solve problems individuals cannot</p></li><li><p>Coordination produces higher quality output</p></li></ul></li></ol><div><hr></div><h3><strong>Phase 7 &#8212; Fully Autonomous Domain Agents</strong></h3><p>The agent becomes a self-directed actor in a complex environment.</p><h4>Characteristics</h4><ul><li><p>Maintains internal world model</p></li><li><p>Generates plans without human prompting</p></li><li><p>Validates those plans against invariants</p></li><li><p>Executes actions with full observability</p></li><li><p>Handles failure</p></li><li><p>Adapts behaviour</p></li></ul><h4>Steps</h4><ol><li><p><strong>Expand domain model</strong></p><ul><li><p>Types, relationships, constraints</p></li></ul></li><li><p><strong>Maintain episodic memory</strong></p><ul><li><p>Past runs</p></li><li><p>Lessons learned</p></li><li><p>State deltas</p></li></ul></li><li><p><strong>Integrate learning (optional)</strong></p><ul><li><p>RL</p></li><li><p>Fine-tuning</p></li><li><p>Feedback incorporation</p></li></ul></li><li><p><strong>Allow long-running behaviours</strong></p><ul><li><p>Monitoring</p></li><li><p>Optimisation loops</p></li></ul></li><li><p><strong>Formal governance</strong></p><ul><li><p>Identity</p></li><li><p>Audit</p></li><li><p>Secure execution boundaries</p></li></ul></li></ol><div><hr></div><h3>Phase 8 &#8212; Socio-Technical Integration</h3><p>This is where production agents in regulated industries must live.</p><h4>Steps</h4><ol><li><p><strong>Define operational policies</strong></p></li><li><p><strong>Define human escalation</strong></p></li><li><p><strong>Define accountability chain</strong></p></li><li><p><strong>Integrate with MCP Gateway</strong></p><ul><li><p>Provenance</p></li><li><p>Enforcement</p></li><li><p>Policy evaluation</p></li><li><p>Observability</p></li></ul></li><li><p><strong>Periodic red-teaming</strong></p></li><li><p><strong>Automated tests for spec compliance</strong></p></li></ol><h4>Purpose</h4><p>The agent becomes a <em>responsible citizen in a system of humans, rules, and other agents</em>.</p><div><hr></div><p>The maturity ladder exists to provide a safe path to the right level of complexity for a specific agent in a specific domain and context. Travelling from:</p><p><strong>Prompt &#8594; Tool Use &#8594; Planning &#8594; Safety &amp; Invariants &#8594; Semi-Autonomy &#8594; Multi-Agent Collaboration &#8594; Full Autonomy &#8594; Enterprise Integration</strong></p><p>Each stage is layered. You <em>don&#8217;t jump</em>. You <em>evolve</em> your agents.</p><p>Each stage reduces human effort while increasing system safety and capability.</p><p>Let&#8217;s explore this in some more detail.</p><div><hr></div><h2>Design the World Before the Agent</h2><p>Agents fail not because they cannot reason, but because they have no stable world to reason about. Before you specify behaviour, specify the <strong>domain</strong>:</p><ul><li><p>What objects exist?</p></li><li><p>What relationships bind them?</p></li><li><p>What states can change?</p></li><li><p>What is forbidden?</p></li></ul><p>Without this grounding, you are not building an agent&#8212;you are releasing a clever hallucination into production.</p><div class="pullquote"><p><em>Clarity of domain is the first kindness you give an agent.</em></p></div><h2>Begin with a Promptling (Assist Mode)</h2><p>The smallest viable agent is nothing more than a structured prompt: a persona, some rules, a handful of examples, and a clear description of desired behaviour. At this stage, the agent:</p><ul><li><p><strong>Suggests</strong>, does not act</p></li><li><p><strong>Explains</strong>, does not execute</p></li><li><p><strong>Learns the contours</strong> of the domain through examples and iteration</p></li></ul><p>This is where you discover your blind spots. It is also where the agent begins to reveal insights about the work itself&#8212;points of friction, ambiguity, or drift that no process diagram ever captured.</p><div class="pullquote"><p><em>Do not give autonomy to a creature you have only just taught to speak.</em></p></div><h2>Give it Tools (Controlled Power)</h2><p>An agent with no tools cannot affect the world. An agent with too many tools becomes a liability. So we choose <strong>carefully</strong>:</p><ul><li><p>APIs</p></li><li><p>File operations</p></li><li><p>Search</p></li><li><p>System commands</p></li><li><p>Domain-specific actions</p></li></ul><p>Every tool is defined by a schema. Every schema is a contract. Every contract is a guardrail. Here, the agent begins to perform real work, but only <strong>under explicit human instruction</strong>. Think of this phase as supervised apprenticeship.</p><div class="pullquote"><p><em>Tools create agency; boundaries create safety.</em></p></div><h2>Teach it to Plan</h2><p>Planning&#8212;GOAP, ReAct, or structured plan-first pipelines&#8212;turns a clever assistant into a <strong>reasoning system</strong>. You introduce:</p><ul><li><p><strong>Actions</strong> with preconditions and effects</p></li><li><p><strong>State</strong> the agent can inspect or update</p></li><li><p><strong>Plans</strong> that can be explained, corrected, or replayed</p></li></ul><p>Once an agent can plan, it becomes legible. Once it is legible, it becomes governable. And once it is governable, it becomes safer to empower and scale.</p><div class="pullquote"><p><em>An agent that cannot plan is not an agent; it is a macro.</em></p></div><h2>Establish Invariants &amp; Safety</h2><ul><li><p>Never delete customer data.</p></li><li><p>Never make financial decisions without validation.</p></li><li><p>Never deploy without approval.</p></li><li><p>Stop if confidence is low.</p></li><li><p>Escalate if risk is high.</p></li></ul><p>These are not suggestions. They are structural truths engraved into the agent&#8217;s reasoning process. Once invariants exist, the agent stops being fragile. It becomes predictable, testable, and increasingly trustworthy.</p><p>With invariants defined, you can add <strong>observability</strong>: logs, traces, decision rationales, and action histories. Now the agent is not only bounded&#8212;it is transparent.</p><div class="pullquote"><p><em>Invariants are the moral laws of agents.</em></p></div><h2>Allow Semi-Autonomy (Governed Execution)</h2><p>Once tools, planning, and invariants are in place, the agent can begin acting without waiting for a human to ask. It can:</p><ul><li><p>React to events</p></li><li><p>Run workflows</p></li><li><p>Monitor states</p></li><li><p>Trigger actions</p></li></ul><p>This is not freedom. This is <strong>responsible empowerment</strong>. The agent acts only within the boundaries you defined&#8212;and those boundaries are now explicit, audited, and observable.</p><p>You have created a junior colleague, not a risk.</p><div class="pullquote"><p><em>Autonomy is not the absence of control but the presence of guardrails.</em></p></div><h2>Introduce Collaboration where Powerful</h2><p>Now you define roles:</p><ul><li><p>A <strong>Researcher</strong> gathers information</p></li><li><p>A <strong>Planner</strong> decomposes work</p></li><li><p>A <strong>Worker</strong> executes actions</p></li><li><p>A <strong>Validator</strong> checks invariants</p></li><li><p>A <strong>Coordinator</strong> arbitrates uncertainty</p></li></ul><p>Through blackboard memory or structured messages, they negotiate, debate, delegate, and converge. Emergent intelligence arises not from a single model but from the <strong>interaction</strong> of specialised agents and their reasoning.</p><p>This is the beginning of agent ecosystems&#8212;the AI equivalent of a high-performing team.</p><div class="pullquote"><p><em>One agent is a tool. A team of agents is an organisation.</em></p></div><h2>Grant Autonomy (with Accountability)</h2><p>The mature agent now:</p><ul><li><p>Holds an internal world model</p></li><li><p>Maintains domain memory</p></li><li><p>Generates plans unprompted</p></li><li><p>Executes safely</p></li><li><p>Explains every decision</p></li><li><p>Integrates feedback</p></li><li><p>Adapts within boundaries</p></li></ul><p>It behaves less like a script and more like a disciplined professional with a well-defined remit. It is not a &#8220;general AI.&#8221; It is a <strong>bounded, purposeful agent system</strong>, finely honed for a domain.</p><div class="pullquote"><p><em>True autonomy is earned, never granted lightly.</em></p></div><h2>Integrate into the broader Socio-Technical System </h2><p>This is the production phase:</p><ul><li><p>Identity</p></li><li><p>Access control</p></li><li><p>Policy enforcement</p></li><li><p>Provenance</p></li><li><p>Audit trails</p></li><li><p>Escalation paths</p></li><li><p>Compliance checks</p></li><li><p>MCP gateway governance</p></li></ul><p>The agent is now a <strong>citizen</strong> of your organisation&#8212;observable, governable, reliable. Not a curiosity. Not a toy. A collaborator.</p><div class="pullquote"><p><em>An agent becomes real not when it thinks, but when it becomes accountable.</em></p></div><h2>A Well-Designed Agent is Matured</h2><p>A well-designed agent does not emerge fully formed. It is sculpted through constraints. It grows through clarity. It learns only what the domain teaches, and it acts only within the laws you write for it:</p><p><strong>Freedom through structure, capability through constraint, harmony through governance.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pWWY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c2139a2-0ff9-4434-a2ea-e82f3350dfb1_3828x2769.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pWWY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c2139a2-0ff9-4434-a2ea-e82f3350dfb1_3828x2769.jpeg 424w, https://substackcdn.com/image/fetch/$s_!pWWY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c2139a2-0ff9-4434-a2ea-e82f3350dfb1_3828x2769.jpeg 848w, https://substackcdn.com/image/fetch/$s_!pWWY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c2139a2-0ff9-4434-a2ea-e82f3350dfb1_3828x2769.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!pWWY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c2139a2-0ff9-4434-a2ea-e82f3350dfb1_3828x2769.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pWWY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c2139a2-0ff9-4434-a2ea-e82f3350dfb1_3828x2769.jpeg" width="1456" height="1053" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6c2139a2-0ff9-4434-a2ea-e82f3350dfb1_3828x2769.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1053,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4778823,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/181668924?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c2139a2-0ff9-4434-a2ea-e82f3350dfb1_3828x2769.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!pWWY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c2139a2-0ff9-4434-a2ea-e82f3350dfb1_3828x2769.jpeg 424w, https://substackcdn.com/image/fetch/$s_!pWWY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c2139a2-0ff9-4434-a2ea-e82f3350dfb1_3828x2769.jpeg 848w, https://substackcdn.com/image/fetch/$s_!pWWY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c2139a2-0ff9-4434-a2ea-e82f3350dfb1_3828x2769.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!pWWY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c2139a2-0ff9-4434-a2ea-e82f3350dfb1_3828x2769.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Phase 7 and 8 are the only contentious phases. They often happen in tandem, and sometimes <em>much</em> earlier in the ladder. </p><p>So the ladder itself is an evolvable tool to customise for the agents you are working upon. The key point being that if you design your agent ladder with care, your agents will climb it with confidence.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Engineering Agents! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>And I should know, I&#8217;ve <a href="https://www.oreilly.com/library/view/learning-chaos-engineering/9781492050995/">created a fair few of those in my time</a>! </p></div></div>]]></content:encoded></item><item><title><![CDATA[The Opportunity to Simplify Continuous Delivery Pipelines, and Promote Flow, Through Specification-Aware AI-Assistance]]></title><description><![CDATA[When the Djinn in your IDE is Specification-Aware, your Pipeline is not quite so alone]]></description><link>https://engineeringagents.substack.com/p/the-opportunity-to-simplify-continuous</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/the-opportunity-to-simplify-continuous</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Mon, 08 Dec 2025 19:01:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JEL1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe04bbee2-ed89-40f6-bb57-471f5fd56e95_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JEL1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe04bbee2-ed89-40f6-bb57-471f5fd56e95_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JEL1!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe04bbee2-ed89-40f6-bb57-471f5fd56e95_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!JEL1!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe04bbee2-ed89-40f6-bb57-471f5fd56e95_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!JEL1!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe04bbee2-ed89-40f6-bb57-471f5fd56e95_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!JEL1!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe04bbee2-ed89-40f6-bb57-471f5fd56e95_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JEL1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe04bbee2-ed89-40f6-bb57-471f5fd56e95_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e04bbee2-ed89-40f6-bb57-471f5fd56e95_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3669704,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/181068112?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe04bbee2-ed89-40f6-bb57-471f5fd56e95_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JEL1!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe04bbee2-ed89-40f6-bb57-471f5fd56e95_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!JEL1!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe04bbee2-ed89-40f6-bb57-471f5fd56e95_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!JEL1!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe04bbee2-ed89-40f6-bb57-471f5fd56e95_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!JEL1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe04bbee2-ed89-40f6-bb57-471f5fd56e95_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In a world not dissimilar from yours a pipeline with 27 stages tries to deduce intent: tests &#8594; scans &#8594; approvals &#8594; checks &#8594; retries &#8594; required reviewers &#8594; more checks &#8594; deploy. It&#8217;s <em>the</em> safety net, and it becomes <em>the</em> bottleneck.</p><p>In another world a team has:</p><ul><li><p>A Specification Constitution</p></li><li><p>A centralised set of specifications and helpful prompts in a Platform Library.</p></li><li><p>Specification aware AI-assistance checks and advice in the IDE</p></li><li><p>- Policy-as-code tied to the spec</p></li><li><p>A three-stage pipeline: validate &#8594; verify &#8594; promote</p></li></ul><p>Lead time collapses. Quality rises. Audit becomes easier. The pipeline shrinks. Trust increases. Developers accelerate.</p><p>And nothing &#8220;freaks out&#8221;.</p><div><hr></div><p>The modern continuous delivery pipeline is an accidental cathedral. It wasn&#8217;t designed; it accreted. It grew in layers &#8212; static checks here, compliance gates there, a scanner bolted on last year, an approval ritual inherited from 2012, a YAML file so large that engineers gather around it like medieval scholars decoding scripture.</p><p>Most organisations now run pipelines that resemble well-intentioned but increasingly unmanageable safety nets: sprawling, brittle, slow, and often misunderstood even by those who operate them. The pipeline has become a proxy for trust &#8212; not because it&#8217;s elegant, but because nothing better has been articulated.</p><p>And yet the physics of software delivery have shifted more in the last 18 months than in the previous decade.</p><p>AI-augmented developers working inside enriched IDEs can already move at speeds the old pipeline cannot understand. Tools and technologies like Claude Code, GitHub Copilot, Cursor, Embabel and bespoke internal agents compress the cognitive load of coding, testing, refactoring, documentation, security checks, and even architectural reasoning. A task that once took hours can now take minutes. A feature that took sprints now takes days.</p><p>This is a moment where your pipeline can begin to &#8220;freak&#8221;. The bottleneck is no longer human hands, perhaps it never was. It&#8217;s the accumulated process scaffolding that grew to protect human limitations. We now find ourselves in a paradoxical trap: developers and their AI assistants can work faster but the systems designed to keep them safe still assume they are slower.</p><p>You cannot ship a modern AI-accelerated flow through a pipeline built for the age of artisanal commits.</p><p>But here lies another insight:</p><blockquote><p>The pipeline is compensating for something upstream &#8212; a lack of a single, explicit definition of what &#8220;good&#8221; looks like.</p></blockquote><p>To simplify pipelines, shorten lead times, reduce risk, and actually benefit from AI-assisted productivity, we must shift the centre of gravity from pipeline logic to specification logic. In other words: create a &#8220;Specification Constitution&#8221; for every org, domain and team and let the IDE &#8212; powered by AI &#8212; maintain it during coding. Let the pipeline keep its gates if it must, but better prepare the developer through an AI-assisted IDE experience to be way ahead of the safety net.</p><p>Make guardrails be actual guardrails, and help the IDE keep the developer on the road.</p><p>This is the ancient Stoic trick for the modern era:</p><blockquote><p>Make intent explicit; anchor behaviour in first principles; simplify the world by strengthening the centre.</p></blockquote><p>A Specification Constitution and standard, helpful prompts can be just that centre:</p><div class="pullquote"><p>a living, structured, executable articulation of purpose, contracts, risks, guardrails, invariants, behavioural examples, compliance boundaries, data handling rules, and acceptance criteria &#8212; written for both <br>humans and machines.</p></div><p>When this specification exists, the IDE and its AI agents can pre-emptively do the work the pipeline currently guards for. Tests generated and checked at the moment of writing. Contracts upheld while lines are still warm. Security rules surfaced while context is fresh. Compliance hints produced before code even compiles. Integration assumptions validated in a shadow pipeline milliseconds after a change.</p><p>The pipeline ceases being the primary intelligence of software delivery and becomes what it should have always been:</p><blockquote><p>a thin, sharp, trusted verifier.</p></blockquote><p>To make this shift from slow assurance to fast safety the impetus is to shift from accidental pipelines to intentional constitutions, and from compensatory gatekeeping to principled engineering as early as possible in your Software Development Lifecycle.</p><div><hr></div><h2>What CD pipelines are really doing today</h2><p>Most CD pipelines are compensating structures. Because we don&#8217;t have a single, explicit, machine-checkable definition of &#8220;good&#8221;, we bolt on:</p><ul><li><p>Static analysis steps</p></li><li><p>Test suites of varying mystery and flakiness</p></li><li><p>Security scanners</p></li><li><p>Code style linters</p></li><li><p>Compliance checks</p></li><li><p>Manual approvals</p></li></ul><p>They are the only gates and guardrails. The pipelines become change advisory boards in disguise</p><p>Each of these complications exist because we don&#8217;t trust what happened earlier in the process. The pipeline becomes a loose pile of inferred intent:</p><blockquote><p>&#8220;If it compiles, passes tests A/B/C, survives scanner X, and Bob clicks approve, then&#8230; I guess it&#8217;s fine?&#8221;</p></blockquote><p>That&#8217;s fragile, slow, and deeply not ready for an increase in flow.</p><div><hr></div><h2>&#8220;Specification constitutions&#8221; in the IDE</h2><p>Making your AI-assisted coding aware of your organisation, domain and even team&#8217;s ways of working through specifications are what&#8217;s missing. Think of these documents as a living, executable constitution of intent for your piece of the system:</p><ul><li><p>Purpose &#8211; what this thing is for; success criteria in business terms.</p></li><li><p>Interface &amp; contracts &#8211; types, invariants, pre/post-conditions, error semantics.</p></li><li><p>Scenarios &amp; examples &#8211; executable examples, behaviour under edge conditions.</p></li><li><p>Policies &amp; constraints &#8211; security, performance, compliance, data handling rules.</p></li><li><p>Risks &amp; guards &#8211; what must never happen; critical invariants; blast radius expectations.</p></li></ul><p>Captured in a structured format the AI can both read and write and co-evolving with the code in the IDE.</p><p>This &#8220;constitution&#8221; isn&#8217;t a doc on Confluence. It&#8217;s:</p><ul><li><p>Close to the code</p></li><li><p>Executable (drives tests, contracts, checks)</p></li><li><p>Readable by humans and machines</p></li><li><p>A primary input and value delivered by AI-assisted coding</p></li></ul><div><hr></div><h3>Extending the power in your IDE via AI: bringing the pipeline guardrails forward</h3><p>Once you have a specification constitution, your AI assistant in the IDE can start doing work that the pipeline currently pretends to do:</p><ul><li><p>Enforce type and invariant discipline while coding (&#8220;this violates the risk guard you wrote&#8221;).</p></li><li><p>Suggest implementations and refactors that preserve contracts.</p></li><li><p>Check simple security and compliance rules immediately (&#8220;you&#8217;re logging PII here, spec says you mustn&#8217;t&#8221;).</p></li><li><p>Run representative pipeline checks locally in a &#8220;shadow pipeline&#8221;:</p><ul><li><p>Compile &#8594; run core tests &#8594; run lightweight static analysis &#8594; quick SAST/DAST stubs.</p></li><li><p>So the feedback loop becomes:</p><ul><li><p>change &#8594; AI + specs check it in seconds &#8594; you iterate</p></li></ul></li><li><p>instead of:</p><ul><li><p>change &#8594; push &#8594; wait minutes/hours &#8594; pipeline complains &#8594; you context-switch back</p></li></ul></li></ul></li></ul><p>You&#8217;re effectively collapsing most of the CI pipeline into the IDE and letting AI keep it in your working memory while you still remember what you just did.</p><div><hr></div><h2>Moving &#8220;good&#8221; upstream</h2><p>Move the definition of good upstream into a shared Specification Constitution, surfaced by AI in the IDE, verified by a simplified, fast and lightweight continuous (and progressive) delivery pipeline.</p><p>Pipelines are compensatory systems. They exist because intent is implicit, distributed, or forgotten. A specification constitution concentrates intent.</p><p>This means AI-assistance can allow really valuable shifting left that was difficult, if not impossible, before. Without intelligent IDE support, pre-commit enforcement is a fantasy. With AI-assistance, the IDE becomes the first, strongest and domain-aware guardrail.</p><p>Specification Constitutions reduce cognitive load.  They capture purpose, contracts, constraints, risks, and scenarios so that developers and agents can work without reconstructing tribal knowledge every time.</p><p>Faster inner loops demand simpler outer loops. An improvement in coding speed requires a reduction in friction downstream or the advantages are lost.</p><p>Compliance and safety can also become more explainable and codified. Policies expressed as structured, machine-checkable constraints attached to specs simplify audits and eliminate ritualised approvals. Let LGTM die as fast as it may.</p><p>Trust can then move from ritual to reasoning. Instead of dozens of gates, we rely on the clarity and integrity of the specification constitution and the reproducible checks derived from it.</p><h3>Some practices to consider</h3><ul><li><p><strong>Write a Specification Constitution for every domain boundary &#8212; </strong>At minimum, capture:</p><ul><li><p>Purpose &amp; business outcomes</p></li><li><p>Pre-/post-conditions</p></li><li><p>Invariants &amp; constraints</p></li><li><p>Type-level definitions (DICE)</p></li><li><p>Behavioural scenarios &amp; examples</p></li><li><p>Security &amp; compliance rules</p></li><li><p>Data-handling classifications</p></li><li><p>Risk declarations (&#8220;what must never happen&#8221;)</p></li><li><p> Observability expectations</p></li></ul></li><li><p><strong>Let the IDE surface the constitution </strong>&#8212; AI-assisted checks should:</p><ul><li><p>Validate invariants during coding</p></li><li><p>Generate &amp; update tests automatically</p></li><li><p>Warn when implementations drift from spec</p></li><li><p>Run shadow pipeline checks continuously</p></li><li><p>Flag compliance/security violations when in flow</p></li></ul></li><li><p><strong>Simplify the pipeline by deleting guesswork</strong> &#8212; Reduce pipeline steps to:</p><ul><li><p>Validate the specification (schema + signatures)</p></li><li><p>Verify code/test alignment with the spec</p></li><li><p>Run system-level checks (integration, perf, resilience smoke tests)</p></li><li><p>Promote via policy-as-code (minimal, <em><strong>absolutely</strong></em> minimal, manual gates)</p></li></ul></li><li><p>In regulated environments, treat specs as governed artefacts. Make them a key part of a centralised <strong><a href="https://engineeringagents.substack.com/p/the-low-hanging-fruit-of-a-platform">Platform Specification Library</a></strong> for all to consume:</p><ul><li><p>Version them</p></li><li><p>Approve them</p></li><li><p>Link them to incidents</p></li><li><p>Maintain audit trails</p></li><li><p>Look to generate compliance evidence automatically</p></li></ul></li></ul><h3>Some things to avoid</h3><ul><li><p><strong>Auto-generated specs with no human intent</strong> &#8212; A constitution without authorship is theatre.</p></li><li><p><strong>Using specs as documentation only</strong> &#8212; They must be executable and enforced.</p></li><li><p><strong>Overcomplicating specs</strong> &#8212; They should shrink ambiguity, not become 50-page legal volumes.</p></li><li><p><strong>Leaving the pipeline unchanged</strong> &#8212; If you don&#8217;t simplify the pipeline, you only move the bottleneck downstream.</p></li><li><p><strong>Remove all bottlenecks from the pipeline too early</strong> &#8212; Only remove as confidence builds, as guardrails are not hit.</p></li></ul><div><hr></div><h2>What happens to the delivery pipeline in this world?</h2><p>The pipeline doesn&#8217;t vanish. It evolves.</p><p>It stops being a primary place where &#8220;good&#8221; is inferred, and becomes a thin verification and promotion service that answers four big questions:</p><ol><li><p><strong>Is the specification constitution present and valid?</strong> &#8212; Schema-valid, signed/approved where needed, linked to change.</p></li><li><p><strong>Does the implementation honour the constitution?</strong> &#8212; All spec-derived tests pass. Contracts and invariants hold (run in CI and, crucially, in staging/prod).</p></li><li><p><strong>Does this change remain safe in the system context?</strong> &#8212; Integration/contract tests. Performance, resource, and resilience smoke tests.</p></li><li><p><strong>Can we safely promote this artefact between environments? </strong>&#8212;<strong> </strong>Progressive delivery, canaries, feature flags. Policy-as-code &#8220;green lights&#8221; for regulated bits.</p></li></ol><p>The pipeline graph goes from:</p><blockquote><p>dozens of steps &#8594; multiple conditional branches &#8594; manual approvals &#8594; long walls of YAML</p></blockquote><p>to something more like:</p><blockquote><p>(1) validate spec + code + tests &#8594; (2) system-level checks &#8594; (3) deploy with guardrails</p></blockquote><p>It&#8217;s simpler but more opinionated. </p><p>If we encode what &#8220;good&#8221; looks like as a specification constitution and let AI enforce and exercise that constitution in the IDE, then most of the complex, brittle logic in our CD pipelines becomes redundant.</p><p>Pipelines can:</p><ul><li><p>Assume a spec is present (and block if not).</p></li><li><p>Verify that code and tests are consistent with the spec.</p></li><li><p>Run a minimal, system-level safety net that&#8217;s independent of the IDE environment.</p></li><li><p>Promote artefacts through environments under policy-as-code rules tied back to the spec.</p></li></ul><p>This is simplification by concentrating meaning less &#8220;accidental&#8221; complexity in YAML. More valuable &#8220;essential&#8221; complexity in specs and policies that humans can understand.</p><p>The pipeline stops being a nervous system and becomes a spinal reflex.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Engineering Agents! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Build A Well-Architected Framework for AI Agents]]></title><description><![CDATA[How to Build, Govern, and Survive AI Agents in your Developer Environment]]></description><link>https://engineeringagents.substack.com/p/build-a-well-architected-framework</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/build-a-well-architected-framework</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Mon, 17 Nov 2025 10:47:07 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!I8do!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82210251-56ce-4847-96e1-4c2edd9f8008_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!I8do!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82210251-56ce-4847-96e1-4c2edd9f8008_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!I8do!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82210251-56ce-4847-96e1-4c2edd9f8008_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!I8do!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82210251-56ce-4847-96e1-4c2edd9f8008_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!I8do!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82210251-56ce-4847-96e1-4c2edd9f8008_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!I8do!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82210251-56ce-4847-96e1-4c2edd9f8008_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!I8do!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82210251-56ce-4847-96e1-4c2edd9f8008_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/82210251-56ce-4847-96e1-4c2edd9f8008_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3545783,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/179127609?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82210251-56ce-4847-96e1-4c2edd9f8008_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!I8do!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82210251-56ce-4847-96e1-4c2edd9f8008_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!I8do!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82210251-56ce-4847-96e1-4c2edd9f8008_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!I8do!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82210251-56ce-4847-96e1-4c2edd9f8008_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!I8do!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82210251-56ce-4847-96e1-4c2edd9f8008_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="pullquote"><p><em>An agent without architecture is an accident waiting for permission.</em></p></div><p>In every environment, there&#8217;s a quiet, uneasy truth: the most dangerous system in production today isn&#8217;t the ancient COBOL core, the Kafka cluster on its last nerve, the new developer wondering what &#8220;clean_up.sh&#8221; does, or the regulatory reporting job stitched together from duct tape and resignation.</p><p>It&#8217;s the AI agent someone built on a Tuesday afternoon.</p><p>Not because AI engineers are reckless &#8212; quite the opposite. They&#8217;re ambitious, curious, and hungry to reduce the friction they drown in every day. So an agent appears: a prompt here, a tool call there, a vector store quietly humming in the corner. It starts as a sketch, grows into a convenience, then an assistant, then &#8212; without anyone meaning for it to happen &#8212; becomes part of the machinery of the company.</p><p>This is the new &#8220;shadow system.&#8221;</p><p>Worse than a shadow database, more slippery than a shadow API, and infinitely more confident than a shadow spreadsheet. And like all shadows, it grows in the places we don&#8217;t look.</p><p>The reason isn&#8217;t malice; it&#8217;s velocity. The tools are too good, the friction too low, the results too tempting. A senior engineer watches an agent autonomously triage tickets and thinks, <em>finally, something in my life makes sense</em>. A junior developer wonders why they ever wrote bash scripts when an agent can self-document, self-debug, and self-delude all in the same afternoon. A product manager sees magic and wants more.</p><p>But in most organisations, magic has consequences:</p><ul><li><p>Magic has auditors.</p></li><li><p>Magic has operational resilience impact tolerances.</p></li><li><p>Magic has model risk committees.</p></li></ul><p>So the game changes. You can&#8217;t rely on talent or good intentions. You need something older, something sterner &#8212; a way of seeing the system beneath the shimmer. A way of remembering that an agent is not a pet, not a toy, not an experiment with production-shaped ambitions. It is an operational entity with the power to move business, alter systems, confuse humans, and violate policies with the enthusiasm of a golden retriever set loose in a fireworks warehouse.</p><p>This is why you should consider crafting a <strong>Well-Architected Framework for AI Agents</strong>. Not as bureaucracy. Not as ornamentation. But as armour.</p><p>The framework isn&#8217;t here to slow engineers down. It exists so they can go fast <em>safely</em> &#8212; so that agents become extensions of the platform&#8217;s reliability, not exceptions to it. It gives language to the risks, scaffolding to the lifecycle, and structure to the responsibilities no prompt will ever shoulder.</p><p>In a sense, it&#8217;s a return to first principles. Organisations, particular regulated ones, have spent decades learning how to govern models, secure systems, classify data, assure resilience, and move change into production without waking the FSA. AI agents don&#8217;t get a magical exemption from any of this. They simply expose the seams in our existing assumptions.</p><p>An AI Well-Architected Framework brings sanity back to the relationship. It turns &#8220;demos that never died&#8221; into governed systems. It turns &#8220;god-mode agents&#8221; into least-privilege citizens. It turns &#8220;black box magic&#8221; into observable, testable, controllable machinery. It gives engineers a golden path and executives a reason not to panic.</p><p>And above all, it offers one clear promise: If you build agents with this discipline, you will survive what happens next.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://engineeringagents.substack.com/subscribe?"><span>Subscribe now</span></a></p><div class="pullquote"><p>You can also catch me on the road at <a href="https://www.russmiles.com/events">various conferences and events</a></p></div><h1><strong>Six Pillars of a Well-Architected AI Agent</strong></h1><h2><strong>1. Governance, Compliance &amp; Model Risk</strong></h2><p>Agents must have purpose, ownership, accountability, and a lifecycle &#8212; or they will become undead artefacts wandering your platform.</p><h3><strong>Some practices to consider</strong></h3><ul><li><p>Register every production agent as both a <em>model</em> and a <em>system</em>.</p></li><li><p>Create an <em>Agent Charter</em> defining its allowed and forbidden domains.</p></li><li><p>Ensure approvals: model risk, DPIA, architecture, resilience.</p></li><li><p>Use lifecycles: experiment &#8594; alpha &#8594; pilot &#8594; production.</p></li></ul><h3><strong>Some things to avoid</strong></h3><ul><li><p>The &#8220;demo that never died.&#8221;</p></li><li><p>Shadow agents with no owner.</p></li><li><p>Agents that quietly drift from helper to decision-maker.</p></li></ul><h2><strong>2. Safety, Security &amp; Permissions</strong></h2><p>The enemy is not intelligence; it&#8217;s power without constraints.</p><h3><strong>Some practices to consider</strong></h3><ul><li><p>Tools are whitelisted, versioned, permission-checked.</p></li><li><p>Agents run under scoped, environment-specific identities.</p></li><li><p>OPA/Cedar-style central policy checks everything.</p></li><li><p>Prompt injection defence: trust boundaries, wrapping, sanitisation.</p></li></ul><h3><strong>Some things to avoid</strong></h3><ul><li><p>God-mode agents.</p></li><li><p>Untrusted Slack/wiki content fed directly into prompts.</p></li><li><p>Agents executing high-risk actions with no human review.</p></li></ul><h2><strong>3. Architecture, Reliability &amp; Control</strong></h2><p>A well-architected agent has a brain and also a spine.</p><h3><strong>Some practices to consider</strong></h3><ul><li><p>Use orchestrators for steps, retries, backoff, idempotency.</p></li><li><p>Encode allowed transitions as state machines.</p></li><li><p>Kill switches, step budgets, environment separation.</p></li><li><p>Degrade gracefully: &#8220;I cannot act; here is my analysis.&#8221;</p></li></ul><h3><strong>Some things to avoid</strong></h3><ul><li><p>One-prompt systems pretending to be platforms.</p></li><li><p>No distinction between experimentation and production.</p></li><li><p>Loops that re-apply Terraform until the infra glows.</p></li></ul><h2><strong>4. Data, Memory &amp; Privacy</strong></h2><p>Memory without governance becomes liability with a timestamp.</p><h3><strong>Some practices to consider</strong></h3><ul><li><p>Design with domains and context in mind.</p></li><li><p>Agents must respect classification, minimisation, and purpose limitation.</p></li><li><p>Memory is a product: schema, TTLs, consent, purgeability.</p></li><li><p>Mask, redact, pseudonymise by default.</p></li><li><p>Ensure residency, domain andboundary controls, and contractual compliance.</p></li></ul><h3><strong>Some things to avoid</strong></h3><ul><li><p>Vector DBs full of raw logs containing PII and secrets.</p></li><li><p>Agents that &#8220;remember forever.&#8221;</p></li><li><p>No separation between sandbox and prod data.</p></li></ul><h2><strong>5. Observability, Evaluation &amp; Incidents</strong></h2><p>An unobserved agent is indistinguishable from a hallucination.</p><h3><strong>Some practices to consider</strong></h3><ul><li><p>Full trace logs: prompts, steps, tool calls, approvals.</p></li><li><p>Offline scenario evals + online behavioural monitoring.</p></li><li><p>Guardrail metrics: blocked actions, policy denials, loops.</p></li><li><p>Incident playbooks with owners, disable switches, RCA.</p></li></ul><h3><strong>Some things to avoid</strong></h3><ul><li><p>&#8220;It seems to work&#8221; as an operational strategy.</p></li><li><p>No E2E eval tasks.</p></li><li><p>No link to operational resilience frameworks.</p></li></ul><div><hr></div><h2><strong>6. Developer Experience, Adoption &amp; Change</strong></h2><p>Agents must be buildable, reviewable, explainable, and evolvable &#8212; or they will become folklore, not tools.</p><h3><strong>Some practices to consider</strong></h3><ul><li><p>Agent-as-code: prompts, tools, evals, policies all in Git.</p></li><li><p>Clear domain-bounded ownership.</p></li><li><p>PR-based changes with CI evals.</p></li><li><p>Golden-path platform abstractions.</p></li><li><p>Change categories with approvals for high-risk updates.</p></li></ul><h3><strong>Some things to avoid</strong></h3><ul><li><p>An agent ecosystem run by &#8220;two hero wizards.&#8221;</p></li><li><p>Prompt edits directly in production.</p></li><li><p>DIY frameworks proliferating like feral cats.</p></li></ul><div><hr></div><h1><strong>The Platform Lens: Agents as a Product</strong></h1><p>An AI-aware internal developer platform becomes the crucible where agents are born safely:</p><ul><li><p>Auth, logging, tracing, policy engine, eval harness &#8212; all centralised.</p></li><li><p>A golden path for building compliant agents from day zero.</p></li><li><p>Guardrails-as-a-service instead of tribal knowledge.</p></li><li><p>A lifecycle that reflects risk, not enthusiasm.</p></li></ul><p>And with this comes the discipline of <em>Well-Architected Reviews</em> &#8212; not as ceremony, but as calibration. A way of asking:</p><p><em>Is this agent a system we can trust? Or a story we&#8217;re telling ourselves?</em></p><div><hr></div><p>In the end, an AI agent is a mirror: it reflects your architecture, your governance, your discipline, and your blind spots. A good framework doesn&#8217;t make magic safer &#8212; it makes systems saner. And in the quiet, high-stakes world of your organisation, that WAF sanity is the closest thing you can get to a superpower.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Engineering Agents! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="pullquote"><p>If you want to accelerate your AI agent engineering, check out the <a href="https://www.russmiles.com/hands-on-ai-agent-engineering">Hands-on AI Agent Engineering Workshop</a></p><p>You can also catch me on the road at <a href="https://www.russmiles.com/events">various conferences and events</a></p></div>]]></content:encoded></item><item><title><![CDATA[The Low Hanging Fruit of a Platform Specification Library]]></title><description><![CDATA[How and why you should curate a centralised, generally-applicable versioned repository of specifications for your developer platform]]></description><link>https://engineeringagents.substack.com/p/the-low-hanging-fruit-of-a-platform</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/the-low-hanging-fruit-of-a-platform</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Fri, 14 Nov 2025 10:36:57 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!vDmv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa086a617-c9e0-4dd8-ba9b-aff6aa560850_3862x2897.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vDmv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa086a617-c9e0-4dd8-ba9b-aff6aa560850_3862x2897.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vDmv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa086a617-c9e0-4dd8-ba9b-aff6aa560850_3862x2897.jpeg 424w, https://substackcdn.com/image/fetch/$s_!vDmv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa086a617-c9e0-4dd8-ba9b-aff6aa560850_3862x2897.jpeg 848w, https://substackcdn.com/image/fetch/$s_!vDmv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa086a617-c9e0-4dd8-ba9b-aff6aa560850_3862x2897.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!vDmv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa086a617-c9e0-4dd8-ba9b-aff6aa560850_3862x2897.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vDmv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa086a617-c9e0-4dd8-ba9b-aff6aa560850_3862x2897.jpeg" width="1456" height="1092" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a086a617-c9e0-4dd8-ba9b-aff6aa560850_3862x2897.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1092,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3253974,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/178870354?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa086a617-c9e0-4dd8-ba9b-aff6aa560850_3862x2897.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!vDmv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa086a617-c9e0-4dd8-ba9b-aff6aa560850_3862x2897.jpeg 424w, https://substackcdn.com/image/fetch/$s_!vDmv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa086a617-c9e0-4dd8-ba9b-aff6aa560850_3862x2897.jpeg 848w, https://substackcdn.com/image/fetch/$s_!vDmv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa086a617-c9e0-4dd8-ba9b-aff6aa560850_3862x2897.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!vDmv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa086a617-c9e0-4dd8-ba9b-aff6aa560850_3862x2897.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>(On a recent visit to The Morgan Library in New York City)</em></p><p>I adore libraries. When on vacation I haunt them voyeuristically, searching for rare books and being entranced by those tombs that promise from their spines.</p><p>The ethos of a library, whether it be a code library or a physical one, is attractive. It says, &#8220;Here lies useful and beautiful, even important, things to discover and use.&#8221; The goal is so simple, here lies something that might be usable <em>by you. </em>Libraries are by and for the borrower. In code we call this reuse, but it is perhaps just <em>scaled</em> <em>use.</em> A contract of service between creators, curator and borrower.</p><p>But curating a library is not easy, nor is it always cheap. Cheaper perhaps than not having one though, and that&#8217;s true of the latest library to enter the mix. The library that services and guides both the AI augmented coding agents &#8212; the Djinn &#8212; and the humans that wrestle with our systems.</p><p>As builders of platforms to support both humans and AI Djinn in their work, a library of specifications is one of the smallest and most impactful offerings we can curate into the platform. Sourcing from &#8220;localised&#8221; specifications &#8212; works in progress helping specific teams &#8212; we can help lift all the intelligences working on our systems by curating what&#8217;s generally applicable;  a library of specifications that should be relevant and useful everywhere across our estate.</p><p>Introducing the Platform Specification Library.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://engineeringagents.substack.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p>There&#8217;s a fiction humming through most engineering teams: That because we have dashboards, pipelines, graphs, wikis, and a thousand tools blinking at us like a migraine in neon, we somehow possess &#8220;understanding.&#8221;</p><p>We believe &#8212; or pretend to believe &#8212; that visibility is the same thing as clarity.</p><p>That a pile of signals is a substitute for a system. Then something breaks and in that fluorescent panic, the lie dissolves.</p><p>We discover that our architecture is not a map at all &#8212; it&#8217;s a maze, a bricolage of incompatible assumptions stitched together by history, inertia, and optimism.</p><p>Now add AI to this mess. AI models are amplification engines. They inherit and magnify whatever environment they&#8217;re given. Feed them a fragmented architecture and they will produce fragmented code. Feed them inconsistent payloads and they will remix the inconsistency into new and baroque variations.</p><p>Feed them a thousand contradictory patterns and they will eagerly hallucinate the thousand-and-first.</p><p>The problem is not the augmented coding djinn. The problem is the lack of a single truth the djinn can lean on.</p><p>This is where the Platform Specification Library enters &#8212; not as a governance artefact, but as a philosophical correction. Gentle guardrails useful to person and djinn.</p><div><hr></div><h2>The Purpose of the Platform Specification Library</h2><p>The Platform Specification Library &#8212; built on Spec-Kit or any similar system &#8212; is not documentation.</p><p>It is law, lore, and language. It is a gift to everyone, just like a library of books.</p><p>It is:</p><ul><li><p>The architecture as text.</p></li><li><p>The domains and subdomains as semantics.</p></li><li><p>The platform as grammar.</p></li><li><p>The shared imagination of the organisation, stabilised in code.</p></li></ul><p>It is an investment. Put plainly:</p><blockquote><p>It ensures that AI augments human creativity rather than amplifying organisational entropy.</p></blockquote><p>Without a central library, AI makes things worse faster. With one, AI becomes the most powerful architecture-reinforcement engine we&#8217;ve ever built.</p><div><hr></div><h2>The Eight Principles of AI-Ready Specifications</h2><p>When planning your own specification library, it&#8217;s easy to get lost in overly-detailed rules and under-specified areas of dangerous freedom. Over time, these are the principles I use to craft my own platform specification library:</p><ol><li><p><strong>Clarity Over Cleverness</strong> - If the spec cannot be understood by a junior, it cannot be acted upon by an AI.</p></li><li><p><strong>Canon Over Consensus</strong> - This is not as flexible as a wiki encourages; it is a single source of truth.</p></li><li><p><strong>Small Pieces, Linked Well</strong> - Every domain concept, every important architectural principle and guardrail, deserves its own file. Fragmentation is avoided through thoughtful linking, not giant monoliths.</p></li><li><p><strong>It is not an island</strong> - The library has connective tissue to Architecture Decision Records, documentation and notes, even localised specification libraries. These are all fabulous sources for distillation into the platform library&#8217;s specifications.</p></li><li><p><strong>Version Everything</strong> - Architecture evolves and decays, domains crystallise. Which version of a specification you are considering matters.</p></li><li><p><strong>Update the Spec Before the System</strong> - Reality must follow the map, or the map will drift into fiction and theatre.</p></li><li><p><strong>If AI Should Generate It, It Should Be Specified</strong> - Consistency becomes a property of the language, not the people.</p></li><li><p><strong>The Easiest Path Should Be the Correct Path</strong> - When the platform makes using specs effortless to apply, adoption becomes organic.</p></li></ol><div><hr></div><h2>How to Establish your own Platform Specification Library</h2><p>Building a library takes time and here are the steps I&#8217;ve found useful when bringin a Platform Specification Library to life:</p><h3>Step 1 &#8212; Write the Intent Charter</h3><p>This is one page. State what the library is for and what it absolutely is not for.</p><h3>Step 2 &#8212; Curate the initial Canonical Formats</h3><p>Use machine-readable standards only:</p><ul><li><p>OpenAPI</p></li><li><p>JSON Schema / Zod</p></li><li><p>AsyncAPI</p></li><li><p>CUE</p></li><li><p>- Spec-Kit link manifests</p></li></ul><p>Everything must be parseable by humans and agents.</p><h3>Step 3 &#8212; Structure the Repository Clearly</h3><pre><code>/spec-kit
   /domains
   /events
   /services
   /workflows
   /platform-capabilities
   /constraints
   README.md
   GOVERNANCE.md</code></pre><p>Clean. Predictable. Idiomatic.</p><h3>Step 4 &#8212; Define Contribution &amp; Review Flow</h3><p>Curate for your context:</p><ul><li><p>PR templates</p></li><li><p>Auto-linting</p></li><li><p>Breaking-change scanning</p></li><li><p>Domain owner + platform architect review</p></li><li><p>Impact Matrix for each contribution</p></li></ul><h3>Step 5 &#8212; Integrate Into Your Developer Workflows</h3><p>Explore:</p><ul><li><p>IDE extensions</p></li><li><p>CLI tooling</p></li><li><p>AI agents that read from the library first</p></li><li><p>Code generators tied to spec versions</p></li></ul><h3>Step 6 &#8212; Connect to Multi-Agent Pipelines</h3><p>AI agents should:</p><ul><li><p>Generate service skeletons</p></li><li><p>Generate client SDKs</p></li><li><p>Generate tests</p></li><li><p>Regenerate code on spec changes</p></li><li><p>Flag drift between code and spec</p></li></ul><h3>Step 7 &#8212; Maintain, Maintain, Maintain</h3><p>Regularly (quarterly) review:</p><blockquote><p>&#8220;Is this still truthful? Is it still the best simple set of specifications for now?&#8221;</p></blockquote><div><hr></div><h2>Some things to avoid</h2><ul><li><p><strong>The Library Becomes an Archive Instead of a Tool</strong> - If it doesn&#8217;t show up in the IDE, it will die.</p></li><li><p><strong>Over-specification</strong> - Spec only what is stable enough to rely on.</p></li><li><p><strong>Under-governance</strong> - Without curation, you get a museum of abandoned truths.</p></li><li><p><strong>Drift Between Spec and Reality</strong> - AI can help detect this &#8212; but only if you let it.</p></li></ul><div><hr></div><h2>When Developers Should Submit Specs</h2><p>It&#8217;s the question I meet the most when putting together a Platform Specification Library, &#8220;How do I knnow when I have something that&#8217;s a candidate for the library?&#8221;</p><p>For this I use a simple heuristic:</p><blockquote><p>If someone else depends on it, or if AI should reinforce it, submit it.</p></blockquote><p>Consider whether there is something to submit when:</p><ul><li><p>You define an API</p></li><li><p>You publish an event</p></li><li><p>You introduce a domain concept</p></li><li><p>You standardise a pattern</p></li><li><p>An integration becomes stable</p></li><li><p>You want AI to generate consistent code against something you&#8217;ve noticed</p></li></ul><p>Probably do not submit:</p><ul><li><p>Experiments</p></li><li><p>Unstable ideas</p></li><li><p>Hacks</p></li><li><p>Anything ephemeral</p></li></ul><p>Rely on the library curation governance process to ensure that the library is a cathedral, not a whiteboard.</p><div><hr></div><h2>An act of epistemic discipline</h2><p>The Platform Specification Library is not a <em>specifically</em> technical artefact. It is an act of epistemic discipline &#8212; a promise that the organisation will not drown in its own contradictions.</p><p>It lets AI do what it&#8217;s good at: Reinforcing patterns, accelerate workflows, stabilising consistency.</p><p>It lets humans do what they&#8217;re good at: Designing systems, imagining futures, writing the code that makes you money (safely).</p><p>A platform that maintains a single, curated, machine- and human-readable truth is not just efficient. It is humane. Because clarity is kindness. And a clear system of specifications &#8212; one with a shared language &#8212; frees people and AI agents to think, to create, to build, to understand.</p><div class="pullquote"><p>The map comes first.<br>The movement follows.<br>And with a good map, the future arrives cleanly.</p></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Engineering Agents! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Right Tool for the Right Context Complexity]]></title><description><![CDATA[The value of a system lies not in its cleverness, but in how faithfully its means match its purpose]]></description><link>https://engineeringagents.substack.com/p/the-right-tool-for-the-right-context</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/the-right-tool-for-the-right-context</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Thu, 13 Nov 2025 16:21:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!sj4l!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc149693f-f147-44e0-a3fb-939a3a65478d_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sj4l!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc149693f-f147-44e0-a3fb-939a3a65478d_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sj4l!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc149693f-f147-44e0-a3fb-939a3a65478d_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!sj4l!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc149693f-f147-44e0-a3fb-939a3a65478d_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!sj4l!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc149693f-f147-44e0-a3fb-939a3a65478d_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!sj4l!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc149693f-f147-44e0-a3fb-939a3a65478d_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sj4l!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc149693f-f147-44e0-a3fb-939a3a65478d_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c149693f-f147-44e0-a3fb-939a3a65478d_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3653703,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/178802763?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc149693f-f147-44e0-a3fb-939a3a65478d_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!sj4l!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc149693f-f147-44e0-a3fb-939a3a65478d_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!sj4l!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc149693f-f147-44e0-a3fb-939a3a65478d_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!sj4l!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc149693f-f147-44e0-a3fb-939a3a65478d_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!sj4l!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc149693f-f147-44e0-a3fb-939a3a65478d_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The Cargo Cult &amp; The Right Fit</h2><p>A team reads about &#8220;AI-driven ops&#8221; and builds an &#8220;autonomous platform agent.&#8221; It scrapes Slack, watches Grafana, and occasionally restarts the wrong cluster. </p><p>No one knows what it&#8217;s doing &#8212; least of all the people who built it. When a failure hits, the agent &#8220;fixes&#8221; the problem by deleting its own logs. They didn&#8217;t build a system; they built a personal hell.</p><p>Automation without proportion is hubris in YAML.</p><p>Another team has a nightly batch job that checks regulatory limits. It&#8217;s deterministic: run query, output report, send email.</p><p>They script it. Simple. Reliable.</p><p>Later, they need to coordinate 40 microservices for release approvals. A single script can&#8217;t know when all are ready, so they shift to events &#8212; &#8220;service.ready&#8221;, &#8220;test.passed&#8221;, &#8220;approval.granted.&#8221; Now releases self-organise through signals.</p><p>Months later, outages rise from external dependencies. They deploy an agent whose goal is &#8220;keep service availability above 99.9%.&#8221; It analyses telemetry, triggers retries, adjusts load thresholds, and opens PRs for misconfigured policies. The system is still human-led &#8212; but now agent-amplified.</p><p>Three approaches, one philosophy: the right tool for the right challenge.</p><div><hr></div><h2>Scripts, Events &amp; Agents</h2><p>In platform engineering, the temptation is often the same: to automate first and think later. You have a mess of pipelines, tickets, and manual toil, and someone says, &#8220;Let&#8217;s script it.&#8221; A bash loop here, a webhook there, a few GitHub Actions &#8212; and for a brief, shimmering moment, everything seems orderly.</p><p>Then reality happens. Contexts shift. Teams reorganise. APIs break. A service fails silently and no one knows why. You add more scripts to fix the scripts, and before long, your platform looks less like a michelin star dinner and more like a thicket of spaghetti.</p><p>This is the fate of <em>scripted automation</em> &#8212; a system that does exactly what you told it to, long after you&#8217;ve forgotten why.</p><p>The next step up the evolutionary ladder is <em>event-driven automation</em>. Here, instead of hard-wiring cause and effect, <em>you wire intent and observation</em>. The system doesn&#8217;t march through a linear script; it listens. It reacts to change. Pipelines trigger when artefacts change, rollbacks fire when health checks fail. The platform starts to feel alive &#8212; a distributed nervous system of events, each one a whisper of state change.</p><p>But even this has limits. Event-driven systems are reactive; they can drown in noise. What they lack is judgment.</p><p>This is where <em>agentic systems</em> enter &#8212; not as magic AI toys, but as autonomous participants in the platform. Agents can set goals, observe outcomes, and adjust plans. They reason about intent, not just input. A scripted automation asks, &#8220;What next?&#8221;; an agent asks, &#8220;What matters?&#8221;</p><p>The lesson is not that one model replaces another. It&#8217;s that each model belongs to a certain level of complexity.</p><ul><li><p>Scripted automations are for the predictable.</p></li><li><p>Event-driven systems are for the observable.</p></li><li><p>Agentic systems are for the unpredictable.</p></li></ul><p>The wise platform engineer, working with the full gamut of options, doesn&#8217;t chase novelty &#8212; they choose proportion.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://engineeringagents.substack.com/subscribe?"><span>Subscribe now</span></a></p><h3>Some practices to consider</h3><ul><li><p><strong>Map complexity before choosing automation</strong> - Draw your system as flows of change and intent, not just tasks. Stable, repeatable steps? Script them. Volatile, stateful reactions? Event them. Ambiguous, evolving goals? Agent them.</p></li><li><p><strong>Encapsulate purpose, not process</strong> - A script describes &#8220;how.&#8221; An event describes &#8220;what.&#8221; An agent reasons about &#8220;why.&#8221; Each is useful &#8212; when aligned with context.</p></li><li><p><strong>Keep humans in the loop, intentionally</strong> - In agentic systems especially, humans must define ethical and operational boundaries. Agents need goals and guardrails.</p></li><li><p><strong>Evolve iteratively</strong> - You don&#8217;t jump from scripts to sentience. Introduce event models first, then simple agents for narrow goals (e.g., pipeline self-healing, PR generation).</p></li></ul><h3>Some pitfalls to avoid</h3><ul><li><p><strong>Premature agency</strong> - Introducing agents before your event fabric is stable leads to chaos. Agents amplify ambiguity; feed them clean signals first.</p></li><li><p><strong>Overfitting the simple</strong> - Not every cron job needs cognition. If it&#8217;s deterministic, script it. Complexity worship is just another form of laziness.</p></li><li><p><strong>Neglecting accountability</strong> - The more autonomy you give, the more you need observability, audit trails, and simulation environments. Agents should be traceable, not mystical.</p></li><li><p>Forgetting the human narrative - A platform that&#8217;s unexplainable erodes trust. Whether scripts, events, or agents &#8212; clarity, observability, is non-negotiable.</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Engineering Agents! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2><strong>Further Reading</strong></h2><ul><li><p>Daniel Terhorst-North &#8211; The Best Simple System for Now</p></li><li><p>Richard P. Gabriel &#8211; Patterns of Software: Tales from the Software Community</p></li></ul><p></p>]]></content:encoded></item><item><title><![CDATA[Building an AI-Agent-Friendly SDLC Data Plane]]></title><description><![CDATA[Make change the primary observable for your people and your agents]]></description><link>https://engineeringagents.substack.com/p/building-an-ai-agent-friendly-sdlc</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/building-an-ai-agent-friendly-sdlc</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Thu, 13 Nov 2025 12:21:08 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!UB_Y!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb96d5f8b-c57e-48c0-ba19-daa484c399d1_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UB_Y!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb96d5f8b-c57e-48c0-ba19-daa484c399d1_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UB_Y!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb96d5f8b-c57e-48c0-ba19-daa484c399d1_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!UB_Y!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb96d5f8b-c57e-48c0-ba19-daa484c399d1_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!UB_Y!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb96d5f8b-c57e-48c0-ba19-daa484c399d1_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!UB_Y!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb96d5f8b-c57e-48c0-ba19-daa484c399d1_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UB_Y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb96d5f8b-c57e-48c0-ba19-daa484c399d1_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b96d5f8b-c57e-48c0-ba19-daa484c399d1_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3740314,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/178781805?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb96d5f8b-c57e-48c0-ba19-daa484c399d1_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!UB_Y!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb96d5f8b-c57e-48c0-ba19-daa484c399d1_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!UB_Y!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb96d5f8b-c57e-48c0-ba19-daa484c399d1_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!UB_Y!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb96d5f8b-c57e-48c0-ba19-daa484c399d1_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!UB_Y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb96d5f8b-c57e-48c0-ba19-daa484c399d1_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>From a soothing fantasy&#8230;</h2><p>There&#8217;s a quiet lie humming beneath most software organisations, and it&#8217;s this: <em>we think we know what&#8217;s going on</em>. Not in the cosmic sense &#8212; engineers are rarely that deluded &#8212; but in the immediate, visceral, operational sense. </p><p>We imagine that because we have dashboards, pipelines, tickets, wikis, and a thousand metrics and risk controls blinking at us like a neon migraine, we have &#8220;visibility.&#8221; It&#8217;s a soothing fantasy, like believing you can understand a storm by counting the raindrops.</p><p>Then something breaks. A page goes red. Dashboards bloom. And in the sudden, fluorescent panic of an incident, that fantasy shows its teeth.</p><p>In that moment the organisation discovers a bitter truth: we&#8217;ve built a labyrinth instead of a map.</p><p>Teams paddle between dashboards that don&#8217;t speak to each other. Release notes live in a wiki written by someone who left during another epoch. Artifact tags don&#8217;t match pipeline IDs. Trace spans don&#8217;t link to deploys. Flags mutate in the dark. </p><p>When the on-call person is asked the most basic generative question &#8212; <em>what changed?</em> &#8212; they&#8217;re forced to guess. A guess! In a multi-million-pound sociotechnical machine built for precision, our response to entropy is an educated shrug.</p><p>You don&#8217;t lose reliability in that moment. You lose trust.</p><p>Because trust isn&#8217;t a philosophical abstraction in engineering. It&#8217;s the invisible surface that everything else rides on &#8212; delivery speed, psychological safety, experimentation, incident response, even hiring. Trust is the currency of modern systems. And trust erodes fastest when people realise the system itself has no idea what it just did.</p><h2>&#8230; to a disciplined reality.</h2><p>Now zoom to another universe. One that isn&#8217;t magical; just disciplined.</p><p>The pager goes off. The on-call engineer opens the incident view. They click an exemplar &#8212; a single trace already linked to the deploy that introduced the regression.</p><p>One hop reveals the 14:03 flag flip tied to deploy r-2025-10-27.3, built from commit a1c3&#8230; The agent surfaces the causal chain: the build, the artifact, the change set, the test deviations, the risk signals. It suggests a scoped rollback with evidence attached. Approver clicks once. Latency drops. </p><p>The post-mortem writes itself from the event trail. No blame. No archaeology. No s&#233;ance.</p><p>This isn&#8217;t &#8220;more telemetry.&#8221; It&#8217;s a different category of capability.</p><p>Observability 2.0 isn&#8217;t a bigger bucket of metrics; it&#8217;s the ability to ask <em>novel, high-cardinality questions</em> about your system with zero prep and zero re-instrumentation. It&#8217;s an epistemic upgrade: the difference between superstition and science.</p><p>And the only way to get there &#8212; the only sane foundation for AI-assisted operations &#8212; is an SDLC <strong>event-first data plane</strong>. Not a disjointed mess of logs and dashboards, but a stitched, curated substrate that ties together the entire software lifecycle: commits &#8594; builds &#8594; artifacts &#8594; deploys &#8594; flags &#8594; runtime &#8594; incidents. With join keys. With provenance. With policies. With lineage across code, data, and models.</p><p>This substrate isn&#8217;t just useful. It&#8217;s liberating. It gives human engineers superpowers &#8212; and gives agents something better to do than hallucinate. </p><p>It reduces cognitive load by replacing detective work with evidence. It turns regulated environments from bureaucratic minefields into traceable, auditable, least-privilege systems where agents can <em>safely</em> propose or execute bounded actions.</p><p>None of this is futuristic. It&#8217;s just the cost of professional adulthood in a world where complexity doesn&#8217;t care how many dashboards you have.</p><p>The organisations that thrive will be the ones who stop worshipping the Three Pillars cargo cult, stop treating telemetry as decorative architecture, and start treating events as first-class citizens. Maybe the Rosetta Stone of modern engineering.</p><div><hr></div><h2>Wide Events as the Substrate</h2><p>Observability 2.0 isn&#8217;t &#8220;more metrics&#8221;; it&#8217;s the ability to ask novel, high-cardinality questions about <em>what changed</em> and <em>who/what it affected</em>&#8212;without a re-instrumentation sprint.</p><p>An AI-agent-friendly SDLC data plane gives you that: a curated, event-first substrate that stitches commits &#8594; builds &#8594; artifacts &#8594; deploys &#8594; flags &#8594; runtime &#8594; incidents. With stable join keys, provenance, and policy, agents can retrieve context, reason about causality, and propose (or execute) bounded actions. </p><p>In regulated environments, the same substrate delivers auditability and least-privilege controls.</p><h3>Some practices to consider</h3><ul><li><p><strong>Model wide events, not shards</strong> - Emit one enriched event per step (commit, pipeline run, artifact, deploy, flag change, incident). Include timestamps, actor, service, env, tenant, trace/span links, and risk signals.</p></li><li><p><strong>Standardize envelopes + keys</strong> - Use a consistent envelope (e.g., CloudEvents) and stable join keys (commit SHA, build/run ID, artifact digest, release version, request trace ID).</p></li><li><p><strong>Link SDLC &#8596; runtime</strong> - Propagate commit/trace IDs from CI/CD into services. Add span links (deploy &#8594; service spans) and exemplars on SLO metrics pointing to the relevant change.</p></li><li><p><strong>Sign provenance by default</strong> - Produce verifiable attestations for builds, SBOMs, images, and deploys (e.g., SLSA/in-toto). Verify before promotion or agent action.</p></li><li><p>Guardrails for agent action - Expose pre-approved tools: &#8220;diff flag,&#8221; &#8220;rollback release,&#8221; &#8220;quarantine artifact,&#8221; each with capability scopes, rate limits, approvals, and audit trails.</p></li><li><p><strong>Privacy &amp; policy at the field level</strong> - Redact/derive sensitive fields early; apply ABAC/row-level filters for both humans and agents. Separate non-prod/prod policies.</p></li><li><p><strong>Cost-aware retention</strong> - Keep wide events; project views (metrics/logs/traces) on demand. Use tail-based sampling + SLO/incident windows for hot storage and tier the rest.</p></li><li><p><strong>Lineage beyond code</strong> - Record lineage across jobs&#8596;datasets&#8596;models&#8596;features&#8596;endpoints so &#8220;what changed?&#8221; queries cover data and ML, not just app code.</p></li><li><p><strong>Human-in-the-loop, by design</strong> - For impactful actions, require review with linked evidence: the event, the diff, the trace, and the attestation.</p></li></ul><h3>Some things to avoid (and potential antidotes)</h3><ul><li><p><strong>ID Chaos</strong> - No consistent run/artifact IDs &#8594; joins fail.<br><strong>Potential</strong> <strong>Antidote:</strong> Central ID strategy; reject events without keys.</p></li><li><p><strong>Schema Drift</strong> - Events mutate silently &#8594; agents hallucinate.</p><p><strong>Potential Antidote</strong> - Contracts, versioning, CI checks, canary consumers.</p></li><li><p><strong>Three-Pillars Cargo Cult</strong> - Metrics/logs/traces sans events.</p><p><strong>Potential Antidote</strong> - Make events the source; derive projections from them.</p></li><li><p><strong>Unbounded Agents</strong> - Bots with prod god-mode.</p><p><strong>Potential Antidote</strong> - Scoped tools, approvals, blast-radius policies.</p></li><li><p><strong>PII Leaks in Telemetry</strong> -  Legal/regulatory exposure.</p><p><strong>Potential Antidote</strong> - Field-level redaction, tokenization, policy tests.</p></li><li><p><strong>Panic Sampling</strong> - Cut volume, lose context.</p><p><strong>Potential Antidote</strong> - Keep wide events; sample spans, not facts of change.</p></li></ul><div><hr></div><p>Software is a living ecosystem and the event-first SDLC substrate is its memory &#8212; the part that stops the organism from touching the same electric fence twice. </p><p>Without it, we&#8217;re left with heroics, folklore, and dashboards that perform understanding instead of producing it. With it, we get systems that explain themselves, agents that stay inside the lines, and teams who can finally make decisions based on evidence rather than vibes.</p><p>In the end, this isn&#8217;t a story about AI or tooling or fashionable new observability slogans. It&#8217;s about responsibility. </p><p>We owe it to our future selves &#8212; and to the people we wake up at 3 a.m. &#8212; to build systems that can tell us <em>what happened</em>. Everything else is an excuse wrapped in a metric.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Engineering Agents! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[A Renascita of Documentation]]></title><description><![CDATA[The Renaissance of Documentation when Machines Code too]]></description><link>https://engineeringagents.substack.com/p/a-renascita-of-documentation</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/a-renascita-of-documentation</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Fri, 31 Oct 2025 10:12:34 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!AtV0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ee20cc0-29a5-4d73-96ac-7e2c8280d431_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!AtV0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ee20cc0-29a5-4d73-96ac-7e2c8280d431_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!AtV0!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ee20cc0-29a5-4d73-96ac-7e2c8280d431_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!AtV0!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ee20cc0-29a5-4d73-96ac-7e2c8280d431_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!AtV0!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ee20cc0-29a5-4d73-96ac-7e2c8280d431_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!AtV0!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ee20cc0-29a5-4d73-96ac-7e2c8280d431_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!AtV0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ee20cc0-29a5-4d73-96ac-7e2c8280d431_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6ee20cc0-29a5-4d73-96ac-7e2c8280d431_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3641507,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/177640304?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ee20cc0-29a5-4d73-96ac-7e2c8280d431_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!AtV0!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ee20cc0-29a5-4d73-96ac-7e2c8280d431_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!AtV0!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ee20cc0-29a5-4d73-96ac-7e2c8280d431_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!AtV0!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ee20cc0-29a5-4d73-96ac-7e2c8280d431_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!AtV0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ee20cc0-29a5-4d73-96ac-7e2c8280d431_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="pullquote"><p>&#8220;Did you write the docs?&#8221; &#8212; a question loved with the same passion by developers as &#8220;Did you write any tests?&#8221;</p></div><p>Documentation has long been the unwanted guest at the engineering table &#8212; everyone agrees it should be invited, but few actually make it feel welcome. We&#8217;d rather ship features than sentences.</p><p>But things are changing. The machines have begun to collaborate with us.</p><p>The rise of AI-assisted coding, autonomous agents, and generative design tools has upended the hierarchy of communication. What once lived in code comments, tribal knowledge, or hallway conversations now has to be written down, because our new collaborators &#8212; the non-human kind &#8212; can only understand what they can parse.</p><p>We are entering a renascita &#8212; a rebirth of documentation as the beating heart of the craft.</p><div><hr></div><h2>The Empty Folder &amp; The Codex of Intent</h2><p>An AI team boasts of &#8220;self-writing code.&#8221; Their repo is a black box and incomprehensible. When the lead engineer leaves, no one knows what the system means. The new agents follow old patterns, amplifying yesterday&#8217;s errors. There are no specifications, only fossils.</p><p>In another place and another time, a Florentine engineer drafts the design for a flying machine. He annotates the sketch &#8212; not just dimensions, but intent: &#8220;It must mimic the bat&#8217;s wing, not the bird&#8217;s.&#8221; Centuries later, a software architect writes a system specification: &#8220;Transactions must resolve with idempotent certainty, not speed.&#8221;</p><p>Both are acts of specification &#8212; artefacts of imagination constrained by reason, teaching others (and machines) how to fly safely.</p><div><hr></div><h2>Renascita?</h2><p>I&#8217;m afraid the choice of &#8220;renascita&#8221; comes from my personal obsession with the 15th and 15th centuries&#8230; </p><p>The term was first used by Giorgio Vasari to describe the Renaissance &#8212; not as nostalgia, but as a rediscovery of proportion, clarity, and disciplined imagination. The same thing is happening in software. We&#8217;re rediscovering that writing &#8212; precise, structured, living writing &#8212; is a foundation of creation.</p><p>In the last century, we frequently focussed only on writing code for compilers. In this one, we write for interpreters: humans and AIs alike. The compiler of the future is not just a syntax checker; it&#8217;s an intelligent assistant, an architectural critic, a reasoning partner. Documentation &#8212; or more precisely, specification &#8212; is what tells it what &#8220;good&#8221; means.</p><p>Specification-Driven Development (SDD) is the embodiment of this renascita.</p><p>It&#8217;s documentation raised to the level of design language &#8212; not static commentary, but executable intent. A well-written specification defines not only what the system should do, but why it matters, how it will evolve, and where its boundaries lie. It becomes a dialogue between authors and agents, an explicit expression of the team&#8217;s collective reasoning.</p><p>In SDD, the spec is a source of truth.</p><p>It guides the code, the tests, the docs &#8212; all derived artefacts aligned with a single coherent narrative. It&#8217;s the bridge between concept and implementation, between human context and machine execution.</p><p>When we move from documentation as an afterthought to specification as a first act, something powerful happens:</p><ul><li><p>Ambiguity dissolves before it metastasizes into bugs.</p></li><li><p>Teams negotiate meaning rather than syntax.</p></li><li><p>AI agents can infer, validate, and extend safely.</p></li><li><p>The system&#8217;s logic becomes auditable, explainable, and habitably human.</p></li></ul><p>This is what Vasari meant by renascita: the rediscovery of mastery through clarity. The Renaissance masters didn&#8217;t just sketch; they specified proportions, materials, and light. The result wasn&#8217;t rigidity, but reproducible beauty. So too with our software: when we specify well, we create living architectures that can be reasoned about &#8212; by humans and by the intelligences we build.</p><p>The new artisans of our age are those who can write systems into being. Not merely coders, but documenters of intent. Architects of meaning. Practitioners of a new literacy, where every line of specification is an act of teaching for humans and AI Djinn coders alike.</p><div><hr></div><h2>When you code with the Djinn, what you don&#8217;t write clearly you&#8217;ll have to debug endlessly</h2><p>Documentation tells you what was built. Specification tells you how things should be built, what needs to be built, and why.</p><p>Specification-Driven Development (SDD) closes the gap between documentation and execution. It ensures that your source of truth is not scattered across PowerPoints, Jira tickets, or human memory, but encoded in structured, testable form.</p><p>This is the foundation of what Guy Podjarny calls the &#8220;<em>specification-first mindset</em>&#8221; &#8212; where specifications don&#8217;t just describe systems, they <em>govern</em> them. Whether expressed as OpenAPI definitions, schema contracts, or infrastructure policies, specifications become a shared trust boundary between humans, AI agents, and automation systems.</p><p>Podjarny&#8217;s work reminds us that specifications are not paperwork; they&#8217;re <strong>guardrails for creativity</strong>. They allow us to move fast without breaking meaning &#8212; or security. In a world where AI systems increasingly act on our behalf, those guardrails aren&#8217;t optional. They&#8217;re the scaffolding of safety.</p><p>This changes everything for AI-augmented development:</p><ul><li><p>Agents can interpret specifications directly, generating and verifying code within defined boundaries.</p></li><li><p>Compliance, observability, and security become emergent properties, not bolt-ons.</p></li><li><p>Human teams communicate in a shared declarative language that AI can extend and reason about.</p></li></ul><p>SDD turns the dream of &#8220;living documentation&#8221; into a practical, operational habit.</p><h3>Some Practices to Consider</h3><ul><li><p>Explore every feature as a specification. Treat the spec as a contract of meaning &#8212; reviewed, versioned, and evolved like code.</p></li><li><p>Write declaratively. Describe behaviour, not implementation. Agents and humans both reason best from intent.</p></li><li><p>Make specs executable. Back them with validation, property tests, or code generation. The tighter the feedback loop, the healthier the documentation.</p></li><li><p>Unify docs, specs, and observability. The spec describes what should happen; observability shows what did happen. Together they complete the loop.</p></li><li><p>Build spec literacy. Teach every developer to write and reason in structured prose and formal models &#8212; it&#8217;s the grammar of the next era.</p></li><li><p>Treat your specs as artefacts of trust. A clear, auditable spec is the cornerstone of ethical AI and resilient automation.</p></li></ul><h3>Some things to avoid</h3><ul><li><p>Specs as bureaucracy. A bad spec is worse than none. Specifications must clarify, not constrain.</p></li><li><p>Lack of ownership. If no one curates the spec, entropy wins. Assign stewardship.</p></li><li><p>Over-automation. Specs should guide AI, not surrender to it. Keep your humans in the loop.</p></li><li><p>Semantic drift. Keep specs synchronised with reality through regular review and comparison to observability. Feedback loops, as always, are key.</p></li></ul><div><hr></div><p>The renascita of documentation through specification-driven development is not nostalgia &#8212; it&#8217;s literacy. It&#8217;s how we teach machines to understand us, and how we keep understanding ourselves when we engage with the creative endeavour of software engineering.</p><p>When code becomes transient, meaning must endure. The specification is our fresco &#8212; a surface where art meets structure, intention meets execution.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Engineering Agents! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Further Reading</h2><ul><li><p>Giorgio Vasari, Lives of the Artists (1550) &#8212; origin of renascita.</p></li><li><p>Daniele Procida, Di&#225;taxis Documentation Framework.</p></li><li><p>Richard P. Gabriel, Patterns of Software &#8212; on habitability and the art of readable systems.</p></li><li><p>Guy Podjarny, <em>Specification-Driven Development</em> and related writings on schema governance, OpenAPI, and automated trust boundaries.</p></li><li><p>Kent Beck, Tidy First? &#8212; the discipline of making change safe and legible.</p><p></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Lessons from Unix for AI Agent Engineers]]></title><description><![CDATA[The Unix shell was the first society of agents; we just didn&#8217;t call them intelligent yet]]></description><link>https://engineeringagents.substack.com/p/lessons-from-unix-for-ai-agent-engineers</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/lessons-from-unix-for-ai-agent-engineers</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Mon, 20 Oct 2025 12:45:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!x1G6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b1d1380-ab30-49e2-ae1f-fb05759acd0e_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!x1G6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b1d1380-ab30-49e2-ae1f-fb05759acd0e_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!x1G6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b1d1380-ab30-49e2-ae1f-fb05759acd0e_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!x1G6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b1d1380-ab30-49e2-ae1f-fb05759acd0e_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!x1G6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b1d1380-ab30-49e2-ae1f-fb05759acd0e_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!x1G6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b1d1380-ab30-49e2-ae1f-fb05759acd0e_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!x1G6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b1d1380-ab30-49e2-ae1f-fb05759acd0e_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3b1d1380-ab30-49e2-ae1f-fb05759acd0e_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3996559,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/176624848?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b1d1380-ab30-49e2-ae1f-fb05759acd0e_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!x1G6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b1d1380-ab30-49e2-ae1f-fb05759acd0e_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!x1G6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b1d1380-ab30-49e2-ae1f-fb05759acd0e_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!x1G6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b1d1380-ab30-49e2-ae1f-fb05759acd0e_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!x1G6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b1d1380-ab30-49e2-ae1f-fb05759acd0e_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>A Tale of Two Agents</strong></h2><p>A start-up built a &#8220;workflow agent&#8221; that promised end-to-end automation. It chained prompts, APIs, and secret logic in a massive YAML file. </p><p>When it worked, it felt miraculous. When it failed, nobody could debug it. The engineers had built a new BPEL &#8212; a tomb of tangled reasoning. The magic had turned into bureaucracy.</p><p>Another team built a constellation of small agents. One parsed customer intent, one fetched data, one wrote summaries, one validated compliance. Each spoke structured text, each logged everything. They piped context between themselves like Unix tools. </p><p>The system was simple to inspect, easy to extend, and resilient to change. They hadn&#8217;t built a &#8220;super-agent.&#8221; They&#8217;d built a society.</p><div><hr></div><h2>Walking in the footsteps&#8230;</h2><p>Every AI agent engineer is, knowingly or not, walking in the footsteps of the designers of Unix.</p><p>The agents we build today &#8212; those polite, tireless djinn of software &#8212; are not so different from the small, composable programs that live in the Unix shell. Each is designed to act, respond, and cooperate through a universal language: text streams, clear inputs, predictable outputs. Each is meant to do one thing well, and to compose gracefully with others.</p><p>But it&#8217;s east to forget the humility that made Unix thrive.</p><p>To begin to dream of <em>giant agents</em>: omniscient orchestrators, magical assistants that &#8220;just know.&#8221; To builf sprawling automations that hide their workings, and workflows that seem to think but cannot be reasoned about. To move from composition to control &#8212; and that&#8217;s how magic becomes mess.</p><p>Unix was never magic. It was honest.</p><p>It said, &#8220;Here is what I do. Here is how I fit with others. You may pipe me, extend me, or replace me.&#8221; That honesty made it habitable for humans. For AI agents, the same rule holds.</p><p>If we want autonomous systems we can trust, audit, and evolve, they must live by Unix&#8217;s design ethos: simplicity, transparency, composition, and respect for the user&#8217;s intent.</p><p>Imagine a world where every AI agent is a small command in a great shell.</p><p>You don&#8217;t build a monolithic &#8220;super-agent&#8221;; you build a garden of little thinkers &#8212; each with a clear goal, a simple interface, and the ability to pass context to others. You compose cognition like you once composed code.</p><p>The future of AI agentic systems will depend less on model power and more on interface philosophy and domain integration. Unix already solved the interface problem half a century ago. The lesson is waiting to be re-learned.</p><div><hr></div><h2>Build AI Agents like Unix commands &#8212; small, composable, explainable, and inspectable</h2><p>Work towards trustworthy, composable, auditable AI agent ecosystems. </p><p>Your context is complex and multi-domain and where autonomy meets accountability. Design small, declarative, message-based agents that can compose like Unix tools, focussing on transparent inputs/outputs and observable behaviours.</p><h3>Some Practices to Consider</h3><ul><li><p>One Intent per Agent &#8212; Each agent should pursue a single coherent goal &#8212; summarise, diagnose, plan, generate &#8212; and expose that intent clearly. Avoid &#8220;Swiss Army&#8221; agents.</p></li><li><p>Pipeability &#8212; Design outputs to become another agent&#8217;s inputs. Use universal data formats &#8212; JSON, Markdown, structured text &#8212; the new stdout/stderr.</p></li><li><p>Text and Traceability &#8212; Prefer textual protocols and logs over opaque embeddings or hidden state. Make reasoning readable by humans and verifiable by other agents.</p></li><li><p>Convention over Control &#8212; Don&#8217;t over-orchestrate. Define conventions for inter-agent communication (contracts, schemas, message envelopes), but let emergent collaboration arise.</p></li><li><p>Small Agents, Big Ecosystems &#8212; Favour a federation of tiny, purpose-built agents over centralised monoliths. Evolution through diversity and iteration &#8212; the Unix way.</p></li><li><p>Observability is Literacy &#8212; Logs, traces, telemetry, and &#8220;reasoning journals&#8221; are the new grep, cat, and less. They allow humans to read what the agents have written.</p></li><li><p>Open Habitat over Closed Platform &#8212; Build environments where agents can live, not prisons where they&#8217;re scheduled. AI habitats need rules of cooperation, not cages of control.</p></li></ul><h3>Some things to avoid</h3><ul><li><p>Over-Orchestration: Repeating the sins of BPEL &#8212; monolithic control flows instead of composable collaborations.</p></li><li><p>Opaque Reasoning: Agents that can&#8217;t explain themselves.</p></li><li><p>Magic Interfaces: GUI or API layers that hide context from both human and agent.</p></li><li><p>Isolationism: Agents that can&#8217;t share state, learn from others, or be replaced safely.</p></li></ul><div><hr></div><p>Unix&#8217;s greatest insight wasn&#8217;t about files or processes &#8212; it was about <em>composability of thought</em>.</p><p>The same principle can save AI engineering.</p><p>If we build agents like Unix tools &#8212; small, transparent, testable, composable, provenance-aware &#8212; we build systems that can be auditable by design.</p><p>If we build them like magic boxes, we inherit the same technical and ethical debt as the opaque orchestration tools of the past.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Engineering Agents! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Agent Design: Lessons from Embabel's Tripper]]></title><description><![CDATA[Notes, Suggested Practices and Things to Avoid from exploring an agentic codebase]]></description><link>https://engineeringagents.substack.com/p/agent-design-lessons-from-embabels</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/agent-design-lessons-from-embabels</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Sat, 18 Oct 2025 06:09:53 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!qHJw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa88e36e6-8bbf-4877-a519-594e1d298c75_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qHJw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa88e36e6-8bbf-4877-a519-594e1d298c75_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qHJw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa88e36e6-8bbf-4877-a519-594e1d298c75_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!qHJw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa88e36e6-8bbf-4877-a519-594e1d298c75_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!qHJw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa88e36e6-8bbf-4877-a519-594e1d298c75_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!qHJw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa88e36e6-8bbf-4877-a519-594e1d298c75_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qHJw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa88e36e6-8bbf-4877-a519-594e1d298c75_1536x1024.png" width="544" height="362.7912087912088" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a88e36e6-8bbf-4877-a519-594e1d298c75_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:544,&quot;bytes&quot;:3586739,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/176471531?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa88e36e6-8bbf-4877-a519-594e1d298c75_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qHJw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa88e36e6-8bbf-4877-a519-594e1d298c75_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!qHJw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa88e36e6-8bbf-4877-a519-594e1d298c75_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!qHJw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa88e36e6-8bbf-4877-a519-594e1d298c75_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!qHJw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa88e36e6-8bbf-4877-a519-594e1d298c75_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="pullquote"><p>&#8220;Autonomy without safeguards is anarchy; safeguards without autonomy is bureaucracy. <br><br>But literacy&#8212;true literacy&#8212;is knowing when to read the code and when to listen to the Djinn.&#8221;</p></div><p>It&#8217;s quite likely we stand in a third age of literacy. The first was born from the presses of Guttenberg, Caxton of Plantin &#8212; when reading and writing words remade civilisation.</p><p>The second was in code, where writing for other humans to read and explore was paramount to evolve complicated systems of value.</p><p>A third is, probably, here now, explored by humans and LLMs &#8212; silicon Djinns that read, reason, and rewrite code alongside us.</p><p>The <a href="https://github.com/embabel">Embabel framework</a>, and its exemplar <em>Tripper</em>, show what happens when we stop chaining prompts like oracles and start composing systems like engineers again. They remind us that AI agents aren&#8217;t a revolution in intelligence&#8212;they&#8217;re a restoration of structure. A return to a discipline of <strong>goal, plan, and act</strong>, wrapped in careful deterministic clarity.</p><p>Most developers are still wandering the LLM wilderness, throwing ever-longer prompts at the void, mistaking clever outputs for reliable systems. Embabel&#8217;s gift is perspective: it treats LLMs not as magicians, but as tools&#8212;powerful, but bounded. It reclaims <strong>agency</strong> from the abyss of stochastic chaos and returns it to where it belongs: in the design.</p><p>Where the early prompt-chains of 2023 were the spaghetti code of the AI era, Embabel offers something closer to <em>the structured programming of agent design</em>. And Tripper, its travel-planning demo, is no toy&#8212;it&#8217;s a proof of concept for a new kind of software literacy: one that blends <strong>deterministic planning</strong>, <strong>domain modelling</strong>, and <strong>explainable autonomy</strong>.</p><p>It&#8217;s an example where we can practice the crucial skill for this new literacy age, of reading, exploring and comprehending<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>.</p><div><hr></div><h2>Read, First</h2><p>Try it now. <a href="https://github.com/embabel/tripper">Clone Tripper</a>, or just <a href="https://github.com/embabel/tripper">explore on GitHub</a>, and see what you notice. </p><p>Maybe don&#8217;t read the next section until you'&#8216;ve explored the codebase a bit. If there&#8217;s anything you notice that I missed, please leave a comment on this post.</p><div><hr></div><h2>What I Noticed while Reading Tripper</h2><p>After taking the time to explore Tripper&#8217;s code, here are the highlights from my code reading journal.</p><div><hr></div><h3><strong>Reading Note. 1 Build Goals, Not Prompts</strong></h3><p>Agents must know <em>what</em> they want, not just <em>what to say next</em>.</p><p>Embabel&#8217;s Goal-Oriented Action Planning (GOAP) separates intention from implementation. Each goal is a navigational star; actions are the waypoints. You don&#8217;t tell the agent <em>how</em> to get there&#8212;you define the physics of the world and let it plan.</p><h4><strong>Suggested Practice</strong></h4><p>Model your agent&#8217;s goals and actions as first-class citizens. Prompts become implementation details, not the architecture.</p><div><hr></div><h3><strong>Reading Note. 2 Type Your World</strong></h3><p>Every meaningful agent is a philosopher of its domain.</p><p>Embabel enforces <strong>strong typing</strong> and <strong>domain models</strong>: trips, legs, itineraries&#8212;not strings, but structures. Typed objects make it possible to reason, validate, and test. They bridge the cognitive gap between code and cognition.</p><h4><strong>Suggested Practice</strong></h4><p>Define, or even reuse existing, domain objects early. Let the agent manipulate data, not prose.</p><div><hr></div><h3><strong>Reading Note 3. Separate Actions, Goals, and Conditions</strong></h3><p>In Embabel, every action declares its <strong>preconditions</strong> and <strong>postconditions</strong>.</p><p>This turns the agent&#8217;s reasoning into a composable system of cause and effect. You can rewire goals without rewriting prompts, extend capabilities without breaking the flow. It&#8217;s loose coupling between actions and goals allowing replanning and fine-tuning as an agent function.</p><h4><strong>Suggested Practice</strong></h4><p>Design each action like a function, not a fantasy. Give it testable inputs, outputs, and invariants.</p><div><hr></div><h3><strong>Reading Note 4. Mix Models Intelligently</strong></h3><p>Not all intelligence costs the same.</p><p>Embabel lets you pick the right LLM for each job: a heavyweight reasoner for thinking, a lightweight one for phrasing. Tripper uses Claude Sonnet for deep synthesis, GPT-4.1 mini for speed. It&#8217;s compositional intelligence&#8212;fit for purpose.</p><h4><strong>Suggested Practice</strong></h4><p>Treat models as interchangeable tools, not idols. Configure, don&#8217;t consecrate.</p><div><hr></div><h3><strong>Reading Note 5. Determinism Is a Feature</strong></h3><p>Chaos may be creative, but determinism is humane.</p><p>By separating planning (GOAP) from LLM execution and reusing domain design, Embabel achieves <strong>predictability</strong>. Plans are inspectable, debuggable, explainable. If a condition fails, the agent <strong>replans</strong>&#8212;not panics.</p><h4><strong>Suggested Practice</strong></h4><p>Let your agent&#8217;s mind be open, but its methods be deterministic. Replanning should be part of the architecture, not an afterthought.</p><div><hr></div><h3><strong>Reading Note 6. Bring Tools into the Light</strong></h3><p>Embabel integrates with the <strong>Model Context Protocol (MCP)</strong> to give LLMs structured, scoped access to real tools&#8212;web search, maps, APIs. Each action decides which tools it can use. The agent doesn&#8217;t &#8220;hallucinate&#8221; a browser; it invokes one.</p><h4><strong>Suggested Practice</strong></h4><p>Expose capabilities through explicit interfaces. Never trust a model to invent your integration.</p><div><hr></div><h3><strong>Reading Note 7. Test, Trace, and Tell the Story</strong></h3><p>Agents are software. Software demands tests.</p><p>Embabel&#8217;s typed flows allow <strong>unit testing</strong>, <strong>integration testing</strong>, and <strong>event streaming</strong> of each action. Tripper&#8217;s live UI even reveals the plan&#8217;s evolution: a narrative of decision and consequence.</p><h4><strong>Suggested Practice</strong></h4><p>Instrument your agents. Emit events. Record their reasoning paths. A black box is not a teammate&#8212;it&#8217;s a risk.</p><div><hr></div><h3>Reading Notes 8+. Some things to avoid &amp; A Helpful Checklist</h3><ol><li><p><strong>The Prompt Trap</strong> &#8212; Treating the LLM as the architecture, not a component.</p></li><li><p><strong>String Soup</strong> &#8212; Failing to model domains as types leads to brittle prompts.</p></li><li><p><strong>Overfitting the Plan</strong> &#8212; Hardcoding paths instead of letting GOAP decide dynamically.</p></li><li><p><strong>Single-Model Monotheism</strong> &#8212; Assuming one model fits all actions.</p></li><li><p><strong>Blind Trust</strong> &#8212; Forgetting to validate outputs or to re-plan when assumptions break.</p></li><li><p><strong>Silent Failure</strong> &#8212; Omitting observability; not knowing <em>why</em> the agent did what it did.</p></li></ol><h4><strong>A Helpful Checklist</strong></h4><ul><li><p>Define explicit goals and actions with pre/postconditions</p></li><li><p>Keep planning deterministic, execution flexible</p></li><li><p>Model your domain with types, not text</p></li><li><p>Use multiple LLMs for cost and capability balance</p></li><li><p>Validate, test, and replan on failure</p></li><li><p>Scope tools explicitly through MCP or similar protocols</p></li><li><p>Emit structured logs and visualise plans for explainability</p></li></ul><div><hr></div><h2>Hungry to Read More?</h2><p>If you want to explore more code, the Embabel team are curating a collection of resources (both code and prose) in the <a href="https://github.com/embabel/awesome-embabel">Awesome Embabel </a>repository.</p><p>And here&#8217;s why you <em>should</em> want to read more code.</p><p>Most developers think they know how to read code. But we&#8217;re not getting anywhere near enough out of the practice and in the age of AI augmented software development, this is a missed opportunity to practice a key skill.</p><p>We skim. We pattern-match. We jump to the bit that looks familiar and hope muscle memory fills the rest. But as Felienne Hermans reminds us in <em>The Programmer&#8217;s Brain</em>, reading code is not a passive act &#8212; it&#8217;s an <em>active cognitive workout</em>. It&#8217;s where understanding is forged, where syntax becomes semantics, and where intuition is born. You don&#8217;t learn programming by typing; you learn it by <em>decoding the minds of others</em>.</p><p>Reading code is reading <em>intent</em>. It&#8217;s reading culture, constraint, compromise &#8212; the archaeology of every decision. And in this new world of Djinn-augmented development, where AI can summon an entire codebase before lunch, reading becomes the act that keeps us human. The Djinn can write, sure &#8212; faster, cleaner, maybe even smarter &#8212; but it cannot <em>care</em>. Only you can trace the thought, spot the flaw, feel the pattern. </p><p>So read more code. Read code like literature. Read until you can sense when a line was written in fatigue or joy. Because the future belongs not to those who can prompt the Djinn &#8212; but to those who can truly <em>read what it writes</em>.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Engineering Agents! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Further Reading</h2><ul><li><p><strong><a href="https://www.felienne.nl/">Felienne Hermans</a>, <a href="https://www.amazon.co.uk/Programmers-Brain-every-programmer-cognition/dp/1617298670">&#8220;The Programmer&#8217;s Brain&#8221;</a></strong></p></li><li><p><em><a href="https://github.com/embabel">Embabel</a> </em>&#8211; GitHub</p></li><li><p><em>Jeff Orkin, <a href="https://www.reddit.com/r/gaming/comments/1cgx297/fear_retrospective_with_ai_programmer_jeff_orkin/">&#8220;Goal-Oriented Action Planning for Game AI&#8221;</a></em></p></li><li><p><em>Rod Johnson, <a href="https://medium.com/@springrod/context-engineering-needs-domain-understanding-b4387e8e4bf8">&#8220;Domain-Integrated Context Engineering&#8221; (DICE)</a></em></p></li></ul><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>If you&#8217;re interested in how code and system exploration and reading can be developed as a skill I heartily recommend reading &#8220;<strong><a href="https://www.amazon.co.uk/Programmers-Brain-every-programmer-cognition/dp/1617298670">The Programmer&#8217;s Brain</a></strong>&#8220; by <strong><a href="https://www.felienne.nl/">Felienne Hermans</a></strong>. Felienne introduced the idea of <a href="https://www.dataart.team/articles/code-reading-club">Code Reading Clubs</a> and its a practice I encourage everyone to try out, especially if you are exploring AI augmented coding.</p></div></div>]]></content:encoded></item><item><title><![CDATA[Provenance as the Chain of Accountability in Agentic Systems]]></title><description><![CDATA[Trace the origin, or you cannot trust the outcome]]></description><link>https://engineeringagents.substack.com/p/provenance-as-the-chain-of-accountability</link><guid isPermaLink="false">https://engineeringagents.substack.com/p/provenance-as-the-chain-of-accountability</guid><dc:creator><![CDATA[Russ Miles]]></dc:creator><pubDate>Wed, 15 Oct 2025 11:51:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!nAIH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F507b8453-1a6c-4d98-8435-cd67ab91f921_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nAIH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F507b8453-1a6c-4d98-8435-cd67ab91f921_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nAIH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F507b8453-1a6c-4d98-8435-cd67ab91f921_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!nAIH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F507b8453-1a6c-4d98-8435-cd67ab91f921_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!nAIH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F507b8453-1a6c-4d98-8435-cd67ab91f921_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!nAIH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F507b8453-1a6c-4d98-8435-cd67ab91f921_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nAIH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F507b8453-1a6c-4d98-8435-cd67ab91f921_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/507b8453-1a6c-4d98-8435-cd67ab91f921_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3394841,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://engineeringagents.substack.com/i/176207427?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F507b8453-1a6c-4d98-8435-cd67ab91f921_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nAIH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F507b8453-1a6c-4d98-8435-cd67ab91f921_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!nAIH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F507b8453-1a6c-4d98-8435-cd67ab91f921_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!nAIH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F507b8453-1a6c-4d98-8435-cd67ab91f921_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!nAIH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F507b8453-1a6c-4d98-8435-cd67ab91f921_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In agentic systems, the most valuable question isn&#8217;t what did the AI do? &#8212; it&#8217;s who asked it to do it, and why?</p><p>Software engineering has long relied on version control to preserve intent: commits, diffs, and signed changes anchor our trust in an ever-changing codebase. But as we step into the world of collaborating AI agents, those human anchors are dissolving. Actions are not taken by a single author, but by networks of agents reasoning across context-rich domains.</p><p>The result? Unprecedented power, but also a dangerous opacity. When a platform incident response, a customer communication, or even a regulatory filing is drafted by a chain of AI assistants, who bears responsibility for what emerges?</p><p>This is where provenance becomes key.</p><p>Provenance is not metadata. It is the story of accountability &#8212; a verifiable lineage of thought, from the human who conceived the goal to the final artefact produced by agents acting on their behalf.</p><p>In the pre-agentic era, we used Git commits as the ledger of authorship. Each change had a name, an email, a hash &#8212; a human fingerprint. In the agentic era, we need the same concept for intent, not just for code. A world where every AI decision and artifact carries a cryptographically verifiable trail back to its originator.</p><p>Without this, AI-augmented development becomes a confusing labyrinth &#8212; plausible but untraceable. With it, we can finally build AI systems that don&#8217;t just act intelligently, but accountably.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://engineeringagents.substack.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Trace the origin, or you cannot trust the outcome</h2><p>In human teams, accountability emerges through shared memory, dialogue, and documentation. In AI teams, we must embed that same accountability in code. </p><p>Provenance signing replaces enables confidence and trust with cryptographic proof &#8212; turning every AI agent from a mysterious collaborator into a responsible contributor.</p><p>In this accountable world, agents establish a &#8220;Provenance Graph&#8221; where every goal or scenario begins with an immutable &#8220;signed intent envelope&#8221; that travels downstream as the root of the provenance graph. An intent envelope ensures that any agent, sub-agent or collaborating model is operating under a verified lineage of human authority and intent.</p><p>Each agent would then append its own trace nodes to the graph:</p><ul><li><p>Agent ID (and model fingerprint)</p></li><li><p>Execution context (inputs, embeddings, environmental vars)</p></li><li><p>Result hash and signature (local signing key)</p></li><li><p>Parent lineage (the node it derived from)</p></li></ul><p>These form a DAG (Directed Acyclic Graph) of agent interactions &#8212; a verifiable audit graph from human &#8594; agent &#8594; artefacts and outcomes.</p><p>When multiple agents collaborate:</p><ul><li><p>Every message or artefact transfer includes lineage metadata</p></li><li><p>The recipient agent verifies upstream authenticity before use</p></li><li><p>The entire workflow becomes auditable within theProvenance Graph, which can be exported or queried like a Git commit tree</p></li></ul><p>Think of this provenance graph as &#8220;Git for agency, not just for code.&#8221; A useful graph that:</p><ul><li><p>Enables audit trails and digital attestations of AI-assisted change in regulatory &amp; legal contexts.</p></li><li><p>Every action is human-anchored &#8212; no &#8220;rogue reasoning&#8221;.</p></li><li><p>Complex multi-agent runs can be reconstructed deterministically.</p></li><li><p>Builds verifiable transparency between humans, AI, and institutions.</p></li><li><p>Provide operational forensics when something goes wrong as you can walk the lineage to the origin intent.</p></li></ul><p>Without a provenance graph, accountability with a farm of agents is, at best, a forensic confusion. At worst it is an impossibility. To truly trust agents in production, accountability is not just a nice-to-have feature &#8212; it is essential to achieve auditable from human intent to AI artifacts and interactions, across agents, domains, and time.</p><h3>Some Practices</h3><ul><li><p>Record and sign every Action, Goal and potentially re-planning step outcome with the human originator&#8217;s identity and timestamp.</p></li><li><p>Build and persist the Provenance Graph alongside your embeddings and traces.</p></li><li><p>- Present provenance paths in your observability dashboards, making them explorable by auditors and engineers alike.</p></li></ul><h3>Some things to avoid</h3><ul><li><p>Ignoring provenance because &#8220;it slows things down&#8221;</p></li><li><p>Treating provenance as optional metadata rather than a core runtime invariant.</p></li><li><p>Over-centralising signatures (one key for all agents), which erodes individual traceability.</p></li><li><p>Storing provenance externally without tamper protection.</p></li></ul><h3>A Checklist</h3><ul><li><p>Every Goal is signed by a human originator.</p></li><li><p>Every Agent logged activity is signed and hashed.</p></li><li><p>The full lineage is queryable.</p></li><li><p>Provenance data is part of observability, not an afterthought.</p></li><li><p>The system can answer: who asked, what was done, by whom, and why.</p></li></ul><div><hr></div><p>If the Renaissance gave us the patent system to honour the origin of thought, and the Enlightenment gave us authorship and copyright law clarity to honour expression, the AI era must give us provenance and cryptographic signatures to honour the origin of intent.</p><p>As AI systems proliferate across development environments, provenance signing and tracing serve three essential functions:</p><ol><li><p>Accountability &#8212; every AI act is rooted in a signed, human-originated intent.</p></li><li><p>Auditability &#8212; every derivation, transformation, or collaboration is cryptographically linked.</p></li><li><p>Reproducibility &#8212; given the provenance graph, the entire reasoning path can be replayed and verified.</p></li></ol><p>For regulated domains like banking, healthcare, or safety-critical engineering, your agents being provenance-aware is not optional. </p><p>It is the difference between a useful assistant and an unrecognisable, uncontrolled actor.</p><div class="pullquote"><p>In a world of infinite authors, trust begins with a signature.</p><p>In a world of infinite agents, trust endures through provenance.</p></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://engineeringagents.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Engineering Agents! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>