Engineering · 4 min read
Every agent needs a judge.
Why we put a second model in front of every answer, and when it isn’t worth the cost.
A language model is very good at sounding right. That is the problem. An answer with a wrong refund amount reads exactly as confidently as one with the right amount, and a person skimming a queue of drafts will not catch it every time.
So we don’t ask people to catch it. In the marketplace system we built, 90 agents work across eight departments, and every team has a judge: 9 judges in all. A judge is a second model whose only job is to check another agent’s answer before anyone relies on it.
Why a second model
Asking the same model to review its own work helps less than you would think. It tends to agree with itself, and it shares the blind spots that produced the mistake in the first place.
That is why our judges run on a different model family from the agents they check. Different training, different habits, different failure modes. When two unrelated models agree that an answer is grounded and within policy, that means much more than one model agreeing with itself twice.
A judge also sees the answer differently. The agent was trying to be helpful. The judge is only trying to find what is wrong. Giving it a narrow brief, and nothing else to do, is what makes it useful.
Block or grade later
There are two ways to put a judge in the path, and choosing between them is the main design decision.
| BLOCK AND REPAIR | SHIP, THEN GRADE | |
|---|---|---|
| When the judge runs | Before the answer leaves | After it has gone out |
| If it fails | Sent back once with the reasons, then shipped with reservations | Flagged, counted, fed into the next lessons |
| Cost to the user | Some extra seconds | None |
| Use it for | Anything a customer sees, anything with money | High volume, low risk, easy to correct |
In the blocking path, a failed answer is not simply rejected. It goes back to the agent once, with the judge’s reasons, and the agent gets a chance to repair it. If the repaired answer still fails, it goes out marked with the judge’s reservations, or, where the rules require it, to a person. Nothing leaves without a verdict.
In the grading path, the answer goes out and the judge scores it afterwards. The scores show where an agent drifts, and they feed the lessons the system learns overnight. Those lessons are not applied on their own: a person approves each one before agents use it.
What a judge checks
A judge with a vague brief (“is this a good answer?”) is expensive noise. Ours check specific things, and each check can fail on its own:
What a judge checks, every time
- Can every number in the answer be traced to data the agent actually read in this run?
- Does it contradict a policy in the shared memory?
- Does it answer the question that was asked, completely?
- Does it propose an action the approval rules do not allow?
- Is the tone right for who will read it?
The first check is the one that matters most. A refund amount, a delivery date or a stock figure that does not appear in any tool result is treated as invented, however plausible it looks.
What it costs, and when to skip it
A judge is another model call on every answer. It adds tokens, and in the blocking path it adds seconds. For a customer reply that is cheap insurance. For some work, it is waste.
We skip the judge, or move it to grading later, when the output is low stakes, reversible and cheap to check. An internal tag suggestion that a person sees anyway, or a draft that someone always edits before sending, does not need a second model standing in front of it.
We also never ask a model to check what code can check. Whether a price is above its floor, whether a date is in the future or whether an order exists are deterministic questions. They get deterministic answers, which are faster and do not have opinions.
A judge is for judgement. If a rule can be written as code, write it as code.
Fail loudly, not silently
Judges fail too. They time out, or their provider is down for a few minutes. The worst thing a system can do then is pretend the check happened.
When a judge cannot run, the answer is shown with a visible “not validated” mark, and the person reading it decides whether to rely on it. A silent pass is how a system loses the trust of the people who work with it, and that trust is much harder to rebuild than a timeout is to fix.
Got a pilot gathering dust?
In 30 minutes we’ll tell you what it would take to put it live.
André is the founder and CTO of WizardingCode. Eight years building the software companies run on, now putting agents into production.