Skip to content

· Olivier Moreau · AI Notes  · 14 min read

Nobody read last night's diff

Fully autonomous coding pipelines are real and shipping. What decides whether yours produces software or defects isn't the model — it's whether your documents say what success looks like, precisely enough to be checked without you.

Twelve pull requests merged overnight. The suite is green. The deploy went out at 04:12 without waking anyone. By the time the team logs on, the work is in production and serving traffic.

Nobody read it. Not “nobody had time” — nobody was supposed to. There is no reviewer assigned, because the pipeline doesn’t have one. An agent planned the work, an agent wrote it, a second agent checked it, and the tests decided whether it shipped.

This is Level 5. Dan Shapiro named it the dark factory, after a production line that runs with the lights off because nothing on it needs to see. Simon Willison, writing about the model in January, added the detail that moves it from thought experiment to practice: he knows a small team operating this way today. No code review whatsoever. Enormous weight on testing. The humans working entirely on system design and pattern discovery.

Which leaves one question worth an article. If no human reads the code, what is the code being written from?

The capability arrived first

It is worth being precise about what changed, because “the models got better” is true and useless. Three specific things happened, and only one of them is really about intelligence.

Context got big enough to hold the problem. Not long ago you handed a model a function and hoped. A single session now holds the architecture document, the modules a change touches, the test suite, the failing output and the entire conversation about it, all at once. A great deal of what used to look like model stupidity was a model reasoning confidently about code it had never been shown.

Memory persists between sessions. Agents no longer start every task as amnesiacs. Retrieval over a project’s history, its decisions and its prior work means the thing working on Tuesday knows what happened on Monday, and why.

Long tool-use chains stopped falling apart. This is the one that mattered. Edit, run the tests, read the failure, form a hypothesis, revise, re-run — thirty or forty steps deep without losing the plot. That is the entire difference between a code generator and something you can leave alone with a task.

Together those three turned a party trick into a shift pattern. But notice carefully what they did and did not do. They removed the capability constraint. They did not tell the agent what your system is for.

A bigger context window is a bigger appetite

A million-token window is not knowledge. It is room for knowledge. The model will fill it with whatever your repository can supply and then act on the result with total confidence.

So if your architectural decisions live in a Slack thread from 2024; if “we never call that service directly” is something four people know and nobody wrote; if your acceptance criteria were a conversation in a meeting room — the context window does not rescue you. It scales the guess, and it scales it overnight, in parallel, across every task.

This is the inversion at the centre of the dark factory, and it gets framed this way rarely because the demos are more fun to write about. For thirty years documentation was a courtesy extended to future humans: the first thing cut under deadline, because the code was the real artifact and the docs were commentary on it.

Now reverse each of those. The spec is what the agent compiles. The architecture document is what constrains the design. The project rules are the linter for the judgment calls no linter can catch. The tests are the only oracle left standing once review is gone.

Documentation stopped being commentary and became the input.

Locate yourself honestly

Shapiro’s model borrows the autonomy levels used for self-driving cars, which makes it useful for exactly one thing: an honest answer to where you actually are.

LevelNameWhat it looks like
L0Spicy autocompleteOriginal Copilot, or pasting snippets out of a chat window.
L1The coding internBoilerplate and unimportant snippets, everything reviewed.
L2The junior developerPair programming with the model, still reading every line.
L3The developerMost code is generated; you have become a full-time code reviewer.
L4The engineering teamYou collaborate on specs and plans. The agents do the work.
L5The dark factoryNo human writes the code. No human reviews it either.

Most teams I talk to sit at L3 and describe themselves as being at L4. That gap is worth naming, because the requirements change sharply across it. At L3 the human is still the oracle — the last line of defence is a person who reads things. From L4 onward, the artifacts are the oracle. Everything after that point is a question about the quality of your written material.

Code review was four jobs sharing one meeting

Removing review feels like removing a checkpoint. It isn’t. It is removing four separate controls that happened to share a user interface. Each one has to be re-homed somewhere, and if you don’t choose where, the answer is nowhere.

The jobHow review did itWhere it has to live now
CorrectnessA second pair of eyes on the logic, catching what the author couldn’t see.Acceptance criteria written into the spec before the work starts, and a test oracle strong enough to hold them — scenario replays, invariants, static analysis, architecture fitness functions in the pipeline.
Design consistency”We don’t do it that way here,” delivered in a comment thread.Project rules the agent reads before writing. AGENTS.md and its relatives, encoding active constraints rather than an archive of past decisions.
Knowledge transferThe reviewer learned the system by reading it, week after week.Living architecture documents, maintained deliberately — because nobody is absorbing the system by osmosis any more.
AccountabilityA named human clicked approve and owned the outcome.Specs with explicit acceptance scenarios, plus production telemetry and error budgets as the surface where divergence from intent is caught.

Every one of those destinations is a document or a test. That is not a coincidence; it is the whole pattern. The Encyclopedia of Agentic Coding Patterns lists the preconditions for operating at Level 5, and the list reads like a documentation audit: machine-readable specifications with acceptance scenarios, a reliable test oracle, fitness functions enforcing constraints, comprehensive production telemetry as the primary feedback mechanism.

A spec that can’t say what “done” looks like isn’t a spec

If there is one thing worth adding to every document an agent reads, it is this: a description of what success looks like, written precisely enough that it could be checked without you in the room.

Most requirements describe the intended behaviour and stop there. A human reading that fills the gap with judgment. They know what “reasonable” means in this system, they know which edge cases the writer didn’t bother to mention, they know which failure is tolerable and which one pages somebody at three in the morning. None of that is written down anywhere, and until recently none of it needed to be.

An agent has no such reservoir. It fills the same gap with plausible invention, confidently, and then ships. Ambiguity used to cost you a slow sprint and a clarifying message. Now it gets resolved on your behalf, overnight, by something that will never think to ask.

So the unit of work stops being a description of the change. It becomes a description of the change plus how anyone would know it worked: acceptance scenarios in given/when/then form, the invariants that must still hold afterwards, the inputs it has to survive, the latency and error budgets past which the result counts as a failure, and an explicit statement of what is out of scope. That last one earns its place — an unbounded agent will cheerfully improve six things you never asked it to touch.

Write a document that way and the document and the test stop being two artifacts. The acceptance scenario is the test, one transcription away. You are not maintaining documentation and then separately maintaining a suite. You are writing the success condition once, in the place where intent already lives, and letting the pipeline turn it into the thing that gates the deploy.

This is also what stops an agent grading its own homework. An agent that writes the implementation and then writes the tests that approve it has verified nothing — it has confirmed its own reading of an ambiguous requirement, at speed, in green. Criteria authored upstream, by a person, before the work begins, are the only thing that makes the check independent.

A test is the only part of a document that fails when it stops being true.

Which is the practical argument for pushing as much of a document’s meaning into that form as it will take. Prose drifts silently. An architecture note can be wrong for eight months and nothing happens, except that everything built from it is quietly wrong too. The section of a spec that names its passing condition cannot do that: it goes red on the next run. In a pipeline where nobody reads the output, that is the only kind of documentation with an alarm attached.

And the discipline cuts in both directions. If you cannot write the success criteria for a piece of work, you do not understand the requirement well enough to hand it to anything — agent or new hire. The difference is that the new hire would have come back and asked you.

What happens when you skip that part

The 2026 empirical work is not kind to the optimistic version of this story.

  • 302,600 verified AI-authored commits analysed across 6,299 repositories
  • 484,366 distinct issues attributed to five widely used coding assistants
  • 22.7% of those issues still present in the latest version of those repositories

Those come from Debt Behind the AI Boom (March 2026), which ran static analysis either side of every AI-authored commit to isolate what the assistant itself introduced. Maintainability debt accounts for 89.3% of the total, and more than 15% of commits from every assistant studied introduce at least one issue.

Two further numbers get quoted a great deal, and both deserve their caveats. CodeRabbit’s State of AI vs Human Code Generation report puts AI-assisted pull requests at 1.7× the findings of human ones — 10.83 against 6.45, across 470 open-source PRs — which is worth reading in the knowledge that CodeRabbit sells code review. And a difference-in-differences analysis of 806 repositories that adopted an AI coding assistant measured roughly a 41% rise in code complexity and a 30% rise in static-analysis warnings afterwards. Those are two distinct measurements, not a single “technical debt went up 30–41%” figure, and collapsing them into one range is exactly the sort of imprecision this article is complaining about. Info-Tech’s summary of the field is blunter: unmanaged AI-assisted development accelerates defects across the whole lifecycle.

There is also a cost that appears in none of those figures. One write-up calls it knowledge debt: changes implemented by agents and never actually understood by the people responsible for them. Your codebase grows while your team’s grasp of it shrinks, and the gap is invisible until the night you need it closed.

The Pattern Book puts the operational version in a single line: don’t try to run at Level 5 on a codebase that can’t be tested well. Remove human review from an undocumented repository with a weak suite and you have not built an autonomous pipeline. You have automated the production of defects and handed it a deploy key.

Stale documentation is now an incident

A wrong document used to mislead one new hire in their third week. Somebody would correct it over coffee. The blast radius was one person and the repair path was human.

A wrong document now misleads every agent on every task that touches that area, silently and in parallel, until something surfaces in production and someone traces it back. Doc rot moved from an embarrassment to a defect with a delayed fuse.

The tell is in what practitioners actually build. Cole Medin’s setup — the one behind his public dark factory experiment — runs a Plan / Implement / Validate loop over per-repository markdown: architecture documents, project rules, reusable skills, workflow history. It also runs a daily reflection pass whose entire job is finding stale entries. He built maintenance tooling for the documentation, because the documentation is load-bearing. Archon, the orchestration layer underneath, pins the process itself into YAML so that the steps you don’t want improvised are not left to the agent’s discretion.

If your documents are input, they need what input gets: an owner, a review cadence, staleness checks in CI, and a definition of done that includes them. “Update the docs” as a hopeful line in the pull request template is not that.

Before you turn the lights off

None of this is an argument against autonomy. It is an argument about sequence. Here is what needs to be true first, and the honest signal that it isn’t yet.

  • Machine-readable specs, not ticket titles. Not ready if requirements reach the pipeline as a sentence and a conversation nobody wrote down.
  • Every spec carries its own success description. Not ready if work can be marked done without anyone having agreed in advance what done meant.
  • Success stated as scenarios and thresholds, not adjectives. Not ready if your criteria say “fast”, “robust” or “intuitive” with no number attached.
  • Criteria written upstream, by a person, before implementation. Not ready if the same agent writes the code and then writes the tests that approve it.
  • A test oracle you would stake production on. Not ready if a green suite has shipped a broken build this quarter.
  • Project rules as active constraints, not history. Not ready if your only written architecture is an ADR archive nobody has opened since it was merged.
  • Fitness functions, static analysis and security scanning in the pipeline. Not ready if the checks that need no opinion are still being performed by people with opinions.
  • Telemetry and error budgets as the review surface. Not ready if production is now your reviewer but isn’t instrumented like one.
  • Feature flags and canary releases. Not ready if every change reaches every user at once. Staged exposure is what buys back the safety review used to provide.
  • A documented route back to manual. Not ready if there is no runbook for the night the pipeline is wrong and someone has to debug code no human has read. Teams that skip this one discover skill atrophy at the worst possible moment.
  • Named ownership and freshness checks on every document above. Not ready if you cannot say who is accountable for the architecture doc being true today.

You do not need all of this to benefit from agentic development. You need all of this before you remove the human.

The bottleneck was never the tooling

Here is the uncomfortable conclusion for anyone currently building a budget for this. The teams that will move fastest into autonomous development are not the ones buying the most capable agents. They are the ones who already wrote things down.

Their specifications are explicit and say what success means. Their architecture is described somewhere current. Their tests fail when the code is wrong. Their constraints live in the repository rather than in a principal engineer’s head. For those teams this is largely a matter of connecting artifacts they already maintain to a new runtime, and the capability improvements of the last two years are pure upside.

Everyone else is about to discover that their real constraint was never model capability and never tooling. It was that the knowledge required to build their software has never existed anywhere outside of people — and that this was survivable only for as long as the people were the ones building.

The agents didn’t create that problem. They made it impossible to keep ignoring.


Further reading

The frame

The working system

Documentation as input

The evidence

Back to Blog

Related Posts

View All Posts »

Tool calling is a protocol, not a feature

Every provider's comparison table says "supports function calling." That's a checkbox on a spectrum. The interesting question is what your agent does on the turn the model answers in the wrong channel — and why it's your cheapest tier that gets there first.

Wiki or RAG? The real answer is a file format

An LLM-maintained wiki beats retrieval for knowledge that compounds. But a memory written by an agent is only as trustworthy as what it records about itself — and that is exactly the gap the Open Knowledge Format was built to close.