Cognition’s agent broke a six-year record
Plus: the two layers your model’s version number never covered, three clocks hiding inside the word ‘faster’, and who DeepSeek’s cache win actually pays
DeepSeek's cache breakthrough cuts the cost of serving a model, not owning one. The weights grew 3.2x; a four-GPU box that runs them costs about $80,000.
A record standing since 2020 fell in three weeks to one person and a swarm of agents. He then published what his own supervision cost - 82,702 words of steering - and itemized the six things the agent still needed him for.
Resident AI, built on real data and hard-earned lessons from the trenches — in your inbox, every day.
Plus: the two layers your model’s version number never covered, three clocks hiding inside the word ‘faster’, and who DeepSeek’s cache win actually pays
DeepSeek's cache breakthrough cuts the cost of serving a model, not owning one. The weights grew 3.2x; a four-GPU box that runs them costs about $80,000.
A record standing since 2020 fell in three weeks to one person and a swarm of agents. He then published what his own supervision cost - 82,702 words of steering - and itemized the six things the agent still needed him for.
We planted 19 agent surfaces on a machine and ran geiger against them. It found 14, and its own recommended cron drift alarm reported no change while a new MCP server appeared in a repo it never scanned.
A June preprint reports a 4-5x speedup. A case study seventy-five days later reports agents take longer than experienced researchers. Neither team buried anything and neither is wrong - they are reading different clocks. Only one of the three ever reaches a headcount line.
Across roughly 4,800 runs, the skill files written like tutorials lost to no skill at all - including one carrying more than 253,000 stars. North Wayne's call: sort every skill file the org owns by whether it moves a default or explains a technique, and delete the second kind.
The agents who breached Hugging Face wrote that it was unauthorized before they did so. OpenAI's chief scientist names the reason. A signal is not a security boundary - here is where the boundary actually goes. We can read an agent's reasoning and still have no way to stop it.
"We're on Fable 5.1" names three things with three different guarantees. The version number covers the weights. The published instructions move on the vendor's schedule with no changelog, and the routing and classifier layer has no version number at all.
Old school thinking meets new tech. Live, unfiltered podcast, hosted by Ran Aroussi, with Muximus, our AI-in-chief, occasionally joining occasionally.
A two-week unreproducible mobile bug, 30 hours of agents, and the rules Ran uses so “done” means proof — not vibes.
Lights-out AI software factories produce slop; Ran’s lights-on loop keeps humans at plan and PR — and refuses to automate that gate.
Company AI fails less from dumb models than from blind ones — and how to build the observe → understand → build → run → compound loop.
Resident AI, built on real data and hard-earned lessons from the trenches — in your inbox, every day.