Zap is the assistant inside EmailZap. It notices email that needs attention, gathers the context, and helps the user take the next step through conversation. It drafts and prepares by default and acts only when the user explicitly asks. Sending an email, archiving a thread or changing a calendar event never happens without a tap on a card that shows exactly what will happen.
When we started, the obvious move was to pick a framework, write some prompts and see what the model did. We had LLM features already: classification, auto-drafts, a search flow on LangGraph. We could have bolted a chat loop onto them in a week.
We didn't. This is the order we went in instead, what each step produced, and what it taught us. It's written for anyone building an agent who suspects that "the model is flaky" is often a spec problem wearing a model costume.
gaps and contradictions found in a spec that had already been through two review rounds, by making the contracts executable with no model attached. Each fix took minutes. Found later, every one of them would have looked like the model misbehaving.
Step 1: understand the job before designing anything
An agent is software that pursues a goal by choosing steps, using tools, checking results, and deciding when to stop or ask. A chat box is not an agent. Neither is a single LLM call that classifies an email. So the first question was not "which framework" but "what job, with what authority, and how would we know it worked?"
Four things came out of that week and shaped everything after:
- One product goal in one sentence. Zap notices, gathers, and helps the user act. Two entry points into one core: proactive ("this needs a reply, here's a draft") and on demand ("find my boarding pass").
- Three authority tiers. Read proceeds. Prepare is shown. Commit (send, archive, calendar change) happens only on an explicit instruction bound to the exact thing that was shown. Never auto-send.
- Memory is three different things, and mixing them is how agents go wrong: task context ("that email"), working state (rebuilt from live data every time), and durable facts (stored with a source and an expiry). An old "my flight is tomorrow" must never become a permanent belief.
- There is no single correctness score. Explicit feedback, opens, draft edits and sends, provider receipts and human-reviewed eval cases each tell you something, and each has a blind spot. The one hard target is zero sends without an explicit instruction.
We also wrote down a shared vocabulary (prompt, context engineering, tool calling, harness, ReAct, RAG, eval, trace, MCP, prompt injection). It sounds like busywork. It stopped a dozen arguments where two people used one word for two things.
Step 2: real data before design
Instead of inventing examples, we pulled thirty real production emails from our tracing tool, where the existing classifier had already captured the context. Random, quota and hard cases. We labelled each one by hand: what it is, the evidence span, what Zap should do, how confident we were.
Five design rules fell straight out of the data:
- Actionable is not the same as replyable. A bill, a signature request or a pickup notice is a task with no reply draft.
- Check the recipient. Real emails in the inbox opened with "Hi Jody" when the inbox belonged to Peter. The request was real; it just wasn't his.
- Conditional actions must show their condition. "If you didn't grant this access, click here" is not an instruction to click.
- Time and thread state change the answer. A later reply or a passed deadline turns a task into noise.
- A matching phrase is evidence, not proof. "Boarding pass" matches last year's flight too.
Our labels agreed with the production classifier on 26 of 27 cases. That is not an accuracy score, because production verdicts aren't ground truth either. Model-proposed labels are proposals until a human has looked at them.
Step 3: design in the order things are hard to change
"Do we start with the framework or the context layer?" got a deliberate answer: neither. Start with whatever is most expensive to change later.
| Order | Layer | Cost to change later |
|---|---|---|
| 1 | Output and evidence contract: what a finished run hands back | Very high. The UI, the evals and the feedback loop all depend on it. |
| 2 | Tool contracts: what Zap can do, with what identity and authority | High |
| 3 | Context layer: what the model sees at each step | Medium |
| 4 | Harness and loop: budgets, approvals, events, traces | Medium |
| 5 | Framework | Low, if 1 to 4 are clean |
A framework is an engine for layer 4. You cannot compare frameworks fairly until you have tools to plug into them and tasks to grade them on. Most teams do it in the opposite order and then discover that the framework has quietly made the output contract for them.
Step 4: write the contracts
A contract is an agreement at a boundary: user and Zap, model and harness, harness and Gmail, harness and UI. Each one answers seven questions: who are the parties, what is the shape, what does it mean, what is guaranteed, what happens on failure, how is it enforced, how does it evolve. People skip meaning, failure and enforcement, and that is where agent bugs live.
Every rule is enforced in exactly one of three ways: in code (guaranteed), in the prompt (usually followed), or by eval (measured). Safety-critical rules go in code. "Never auto-send" cannot live in a prompt that one injected email could talk its way around.
The model proposes. The harness decides. Anything the system can know for itself, the model never gets to assert.
We treat the model the way a backend treats a browser: as an untrusted client. That one idea produced most of the design.
| Design move | What it prevents |
|---|---|
Handles like m3 and t1 instead of real IDs, issued only by the harness | Invented IDs, and access across users or inboxes. Also cheaper in tokens. |
| Citation becomes Evidence: the model quotes, the harness finds the quote in the source and records the verified span | Invented quotes |
| Claim markers that tie each answer sentence to a verified claim | Unbacked statements. An eval flags any sentence without a marker. |
| Result fields split into model-authored proposals and harness-authored facts | "I checked your replies" when it didn't. "Sent!" when nothing was sent. |
| No free-form reasoning shown to users; the model's thinking goes to the trace | Plausible after-the-fact explanations being trusted over evidence |
| No write tools for the model. Approvals come from the validated result and the harness executes after the tap. | The model, or an injected email, triggering a send |
| Computed checks ("is the user a recipient?", "is there a later reply?") run in code whenever an object is read | Relying on the model to remember to check |
| No model calls inside tools | Hidden latency and cost, and steps that never show up in the trace |
The contracts weren't written in one sitting. Each round started with a question, often mine, often a naive one. "Isn't an evidence ref the same as an evidence span?" forced a rename that fixed a real confusion. "What do you mean by one account?" split Account into Connection and User, which is how a primary and a secondary inbox became one user with two connections. "Gmail, Calendar and a CRM should each bring their own tools" turned the design into a core plus plugins, so adding a CRM is a plugin and no core change.
The lesson I keep relearning: when a clear explanation is hard to give, the concept is usually wrong. Confusion is a design signal, not a teaching moment.
Step 5: review the spec like code
We ran the contracts document through a diff review, the same way we'd review a pull request. It found real defects, not wording nits: action kinds with nowhere to be registered, a worked example citing fields that didn't exist, internal IDs leaking into the model's view through check results, two rules that contradicted each other, and a paragraph that silently broke a markdown table so the most important row disappeared.
A long spec contradicts itself in ways nobody notices reading it top to bottom. Structured review catches some of it. Executable code catches far more, which is the next step.
Step 6: make every decision explicit and owned
Every open question became a numbered decision with options, a concrete example and a recommendation. When one wasn't clear, we re-explained it with an actual email rather than more abstraction. Thirteen decisions went in before any code, and a few are worth sharing:
- Approval is a card with one tap. Typed or spoken approval may come later, but approval is never the model interpreting words.
- Reuse infrastructure, not architecture. We kept the existing drafts store and search behind an adapter. Then we checked what that code actually did. "Reuse search" turned out to include a hidden model call for query building, which Zap skips. "Reuse drafts" came with constraints (one live draft per thread, replies only) that went straight into the contract.
- Misaddressed mail surfaces as "Is this yours?" and Zap learns from the answer.
- Attachments go through a managed malware scan inside our own cloud, not a public scanning service, because files uploaded to some of those can be shared with their partners. We checked the current docs rather than trusting memory.
Step 7: make it executable before making it smart
This is the step I'd argue for hardest. Before any model, we built the framework-independent parts as real code: the schemas, the validators, the approval state machine and ledger, the integration registry, fake Gmail and Calendar providers loaded with the hard cases from step 2, and content extraction.
Then we drove it with hand-scripted "golden" runs: the exact tool calls a perfect agent would make. The model is a port. The scripted agent stands in for it now, and a real model plugs into the same slot later.
Three runs work end to end: an invoice follow-up including an injection email and a repair round; today's boarding pass, not last year's; and scheduling a call that an email asks for. All eight misbehaviour cases are refused: an invented handle, a fake quote, a stray URL, a rule-breaking integration, a stale approval, a double tap, a recurring calendar change with no scope, an edit by someone who isn't the organiser. 591 tests, and mutation testing to prove the tests can actually fail.
The findings that earned their keep
- The approval tool was a hole. The model-callable "request approval" tool could put a "Send?" card in front of the user before the result was validated, so a draft with an invented address could reach a tap. The tool was removed. Cards now come only from the validated result.
- Mutation testing found a lie. On a double tap after a send, checking staleness before "already decided" would report a sent email as stale. The first double-tap tests couldn't fail, which is how we found out. The order of checks is now in the contract.
- "User stated" needed proof. The model could declare a memory as something the user had said and skip the confirmation card. It now needs a quote the harness finds in the user's actual message.
- A rule with no runtime check isn't a rule. "No IDs in model context" was written down and enforced nowhere. A leak detector now scans every tool result before the model sees it.
- End-to-end runs find what unit tests don't. The three golden runs surfaced three bugs the 500-odd unit tests had walked past: version binding, recipient checks on outbound mail, and a ledger failure reason.
Seventy findings in total, folded back into the contracts across two versions, plus seventeen new decisions that only building could surface. The findings are the product of this phase, not a failure of it. The success metric was gaps found before a model was involved.
How the build itself was run
Four gates, each approved before code: product, architecture, program design, slice plan. Eleven vertical slices. The first four ran in sequence with one lead coding agent to establish the core. Then four slices ran as parallel agents on disjoint areas, then two more, then the last with the lead. Tests were written first and had to fail first. When I asked "which design patterns are we actually using?", the answer got written down, along with the patterns we deliberately weren't using and why.
Two rules kept parallel work from colliding. One owner per shared file: the build phase owned the contracts document, and other streams logged problems in a findings file instead of editing it. And decisions live in one table, not in chat answers. When a chat answer and the table disagreed, the table won.
What comes next
Only now does the part most teams start with begin. An eval set built from human-reviewed cases and loaded with score configs. Then bake-offs: runtime first (our own loop against Pydantic AI, LangGraph and an agent SDK), then model, then the importance method, all scored against the same contracts and the same evals. Then one real slice end to end with real Gmail and Calendar adapters.
The framework decision comes last on purpose. By the time we make it, we'll have tools to plug in, tasks to grade on, and a harness that already refuses the eight things we most need refused.
Lessons, in one list
- Define the agent before building it: goal, authority, correctness, speed.
- Real data before design. Production traces are the fastest source of hard examples. Production labels are not ground truth.
- Design from hardest to change to easiest: output, tools, context, harness, framework.
- The model proposes, the harness decides. If the system can know it, the model doesn't get to say it.
- Safety rules live in code. Prompts are for judgment. Evals measure both.
- Confusion is a design signal. If a concept is hard to explain clearly, it's probably wrong.
- Review specs like code, then make them executable. The code found 70 gaps the reviews missed.
- Make decisions explicit and owned, with examples and a recommendation, in one place.
- Reuse infrastructure, not architecture, and verify what the reused code actually does.
- Split what must never be lost from what's nice to have: a thin ledger for approvals and receipts, a full trace for everything else.
- One owner per shared file when several people or agents work in parallel.