18 November 2025

Shipping cited answers in six weeks

Field notes from an engagement inside a large, document-heavy enterprise: the bar that matters, the shape of the work, and why citations are not a feature — they are the product.


The brief was unambiguous. A large enterprise with a dense document estate — policies, contracts, controls, operational standards, research — needed an agent that could answer real questions for its own people, drawing exclusively from its own record. The answers had to be cited. Every response had to be replayable. The system had to go live, not into a lab.

Six weeks from kickoff to production. No client naming here — the sector does not matter for this retelling, and the details that do matter are posture, not commercial.

The bar that actually matters

The framing that cost us the most time early in our practice was thinking of the audit bar as a constraint — something the model must not violate. That is wrong. The bar is the product. If you build the agent first and then ask it to behave defensibly, you will fail. You build the evidence trail first, and the agent emerges from it.

Concretely: we do not start with a model choice. We start with a single question. Given a specific answer the agent produces on a specific Tuesday, what is the full record a reviewer would want to see? That record — the retrieved passages, the model version, the prompt, the tool calls, the timestamps, the user — is the specification. The system that produces it is what we build.

The shape of the work

We won't walk the mechanics here. The interesting part is not the recipe; it is the discipline. A few things we insisted on, that buyers should insist on from anyone:

The corpus is treated as a living dataset. Document estates are never as clean as anyone claims. The first pass is not retrieval — it is hygiene, structure, and versioning. Anything less and citations become decoration.

Grounding is enforced, not encouraged. If the evidence isn't there, the agent says so. It does not paper over a gap. This is a contract, not a hope.

Evaluation is co-authored. The people who will have to defend the outputs help define what "good" looks like, including the adversarial cases. A regression suite gates every deploy. This is the difference between a demo and a system.

The trail is a by-product. Every call leaves a record that a reviewer can actually read. Not a log-aggregation screen — a reconstruction of what the agent saw, what it did, and why.

What got skipped

Things we deliberately did not do. A custom model — frontier retrieval-grounded generation was sufficient. A fancy interface — plain prose with inline citation anchors was more than enough, and the reviewers did not want to learn a new UI. A multi-agent orchestration — one well-specified agent outperformed every multi-agent design we prototyped, on every metric that mattered.

Simplicity is underrated. Every additional component is another surface a review has to reason about.

The trail is the deliverable

At handover we produced three things: a live agent in production; a regression suite the firm's team owns; and a reviewable record for every answer it produces. The agent is the visible part. The other two are what make the work repeatable, defensible, and — when the inevitable change lands — adaptable.

This is the thing we wish more buyers understood at the RFP stage. The cost of a bad AI deployment is not the implementation fee. It is the cleanup, the trust debt, the quiet retreat of the people who could have championed it internally. The evidence trail is not a nice-to-have. In a serious enterprise it is the entire point.


— fin —
Book a call