EmailZap is a small team shipping an iOS app, a Chrome extension and a backend on a two-week cadence. For most of the last year, every line of code has been written with an agent in the loop: Claude Code and Codex for implementation, run in parallel worktrees through Conductor, with OpenRouter in front of the models for our own tooling and evals. This post is about the part of that story that went wrong, and what we changed.
The factory we tried to build
The pipeline looked like this. A Linear ticket became a spec and an architecture note. The spec went to a coding agent, which opened a pull request. Instead of a colleague reviewing it, five specialised review agents did: correctness, security, performance, tests, style. If they passed it, it merged. The engineer's job was to write the ticket and watch the dashboard.
The logic was hard to argue with. Peer review was already our bottleneck. Agents had made writing code roughly ten times faster, and a queue of PRs waiting on two humans was where all that speed went to die. Replace the humans in the queue with agents and the queue disappears.
It didn't disappear. It moved.
What actually happened
Three things, in roughly this order.
The reviewers agreed with the author. An agent reviewing an agent's PR reads the same spec, makes the same assumptions and is impressed by the same plausible code. The review agents caught real things: a missing null check, an unbounded query, a test that didn't assert anything. What they never caught was a PR that faithfully implemented the wrong thing. The spec said "add the CTA to the settings screen", the design had quietly moved it, and the agent built the settings screen. An engineer lost a full day to exactly that. No reviewer, human or agent, could have caught it from the PR, because the PR matched the ticket. The error was upstream, and upstream was where we'd stopped spending human time.
Nobody held the model of the system any more. When a person writes a feature, they carry a picture of the surrounding code in their head for a week afterwards, and that picture is what makes the next feature cheap. When an agent writes it and other agents approve it, nobody carries anything. Every ticket started from zero. The tenth feature in a module cost as much as the first. We had a release incident traced to a screen whose states had never been enumerated: not a bug in the code that was written, a bug in the code nobody thought to write, because nobody was thinking about that screen.
Humans stopped reading. This one is on us, not the tools. Once the dashboard was green most of the time, we trusted the green. A PR that had passed five agents felt reviewed. By the time we noticed the pattern, "reviewed" had come to mean "nobody objected", and nobody objected because nobody looked.
The net effect was the paradox that started this: code was cheap and shipping was slow. Slow because of rework, because of incidents, and because the two people who could fix a bad merge were the same two people we'd tried to take out of the loop.
The rule we run on now
Thinking is never offloaded to the model. Typing is. Checking is. Deciding is not.
Lights-out was the wrong goal. The goal is not zero human time. The goal is human time spent only where a human is the only thing that works: deciding what to build, holding the picture of the system, and being the person who says "this is done". Everything else the agents do, and they do more of it than before, not less.
So we didn't roll back to 2023. We rolled back the one decision that broke it, and rebuilt the pipeline around where people are irreplaceable.
The pipeline now
| Stage | Who thinks | Who does the work |
|---|---|---|
| Ticket and grooming | Product and engineering, together, before the sprint | People. Agents draft; a person decides scope and states. |
| Sprint plan | The team, in a fixed planning slot | People |
| Spec | The engineer, using a spec kit that forces the questions | Agent drafts, engineer edits until they can defend every line |
| Implementation | Nobody new: the spec is the thinking | Claude Code or Codex in its own worktree, several in parallel through Conductor |
| Agentic review | Nobody: this is a filter | Review agents run first and clear the mechanical defects |
| Self-review | The engineer, reading their own diff as if they'd written it | Person |
| Peer review | A second engineer, on the design and the fit, not the syntax | Person, with the agentic review already done |
| Dev, then prod | CTO go/no-go on release day | Automated deploy, human decision |
Three things about this table matter more than the rows.
Human review came back, but it changed. Agentic review runs first, so by the time a person opens the PR, the null checks and the unbounded queries are already gone. The human reviewer isn't proofreading. They're asking whether this is the right change, whether it fits the system, and whether the engineer who wrote it can explain it. A review that used to take an hour of nitpicks now takes twenty minutes of judgement, and it's the judgement that was missing.
The spec is where the thinking lives, so the spec is human-owned. An agent drafts it, because drafting is typing. The engineer edits it until they can defend every line, because that's the moment they load the picture of the system into their head. If the spec is right, the agent's implementation is boring. If the agent's implementation is surprising, the spec was wrong, and we fix the spec, not the code.
Parallel agents, serial decisions. Conductor lets one engineer run several worktrees at once, each with its own agent session. That's the throughput win, and we keep it. What we don't parallelise is the deciding: grooming, the spec gate, the release go/no-go. One owner per shared artifact. Decisions in one table, not in chat. When a chat answer and the table disagree, the table wins.
What the agents may and may not do
- An agent may draft anything: tickets, specs, code, tests, release notes, the review comments.
- An agent may not close a decision. Scope, states, architecture and "done" are signed by a person.
- An agent may not merge. Agentic review clears a PR for human review; it doesn't approve it.
- An agent may not be the only reader of a change. Every merged diff has been read by the engineer who owns it and by one peer.
- Tests are written first and must fail first. If an agent writes a test that passes on the first run, it's deleted and rewritten. Mutation testing checks that the tests can fail at all.
- Nothing enters a running sprint except through the on-call "goalie" or a production emergency. Agents make interrupts cheap to act on; that's a reason for more discipline about which ones you accept, not less.
What it looks like in practice
On a greenfield rebuild of our AI agent this month, the workflow was four gates (product, architecture, program design, slice plan), each read and signed by a person before any code, then eleven vertical slices. The first four ran in sequence with one lead agent to establish the core. Then four slices ran as parallel agents on disjoint areas, then two, then the last with the lead. It shipped 591 tests and three end-to-end scripted runs, and it surfaced seventy gaps in a spec that two human review rounds had missed. Every one of those gaps was fixed by a person deciding something. The agents found them; they didn't resolve them.
On the production app, a team of three holds the two-week cadence across iOS, Chrome and backend: code freeze and store submissions on day eight, backend release on day nine, retro and planning on day ten. Six months ago it took twice the people to hold the same schedule, and the difference is not that the agents got better. It's that the humans stopped doing the parts an agent does better and started doing the parts only they can.
Where this is going
The lights-out idea isn't dead. It's just aimed at the wrong layer. Fully autonomous execution of a decided spec is fine and getting better every quarter. Fully autonomous deciding is the thing that failed, and nothing about the models improving changes that, because the failure wasn't model quality. It was that a system with no human holding its shape drifts, and drift doesn't show up in any single PR.
So the next version of our factory is built to make human deciding faster, not rarer: agents with different competencies (a junior that implements, a senior that reviews for fit, an SRE that reviews for operations), voice input for grooming because engineers think out loud better than they write, and every decision captured where the next agent can read it. The person stays in the chair. The chair gets a lot more leverage.
Lessons
- Removing humans from review doesn't remove the bottleneck. It moves it to rework and incidents, where it costs more.
- Agents reviewing agents catch mechanical defects and miss the wrong thing built well. Use them as a filter before human review, not instead of it.
- Somebody has to hold the picture of the system. Make that person's spec the artifact that carries the thinking, and never let an agent close it.
- Parallelise execution, serialise decisions. One owner per shared file. One table of decisions.
- A green dashboard is not a review. "Reviewed" means a named person read it and would defend it.
- Tests first, and they must fail first. An agent will happily write a test that proves nothing.
- Measure shipping velocity, not code velocity. The first number is the only one customers can see.