Skip to main content
Let a judge decide when a review loop stops, instead of a count. Use it when the work is refined over rounds and the right number of rounds is not known in advance: a page, a plan, a shortlist. When a fixed number of rounds is enough, refine: 3 on the stage is the plain way, and a writer and a reviewer shows it. The judge is a small decision model, not the reviewer. It answers typed questions about what the reviewer found and how the rounds have gone, one cheap call per round, and the question it answers is whether the work has reached the point of diminishing returns. A count stays as the safety cap, so the loop always ends.

Shape

The stopping rule

The reader tags each finding [block], [should] or [nit]. The route reads the judge’s answers and decides:
examples/judge-stops-the-loop.ts (excerpt)
A block is a false claim or a sentence the reader would not understand, and it always goes back, whatever the judge says. Otherwise the judge’s chosen reason decides. The cap in maxKickbacks is the last word: a run cannot loop past it.

The questions

The judge sees the use case from the brief, the draft, the latest findings and every round so far with its counts and the size of each change. The default question set asks whether the draft holds, whether the findings are worth acting on for this use case, whether another round is worth it, and why to stop:
examples/judge-stops-the-loop.ts (excerpt)
In workflow() the same loop is one field on the reviewed stage: refine: judge(seat, { cap: 6, questions }), with stopQuestions() from @obversa/runtime as the default set; a dag() takes the same judge() value on maxKickbacks. There a judge that says stop lets the review stand in workflow(), and on a bare dag() it rejects the kickback and leaves the requesting step’s outcome as it was. A custom question set routes on its stop_reason when it has one, and on the cap otherwise. The judge engine takes { state, questions } as its prompt and returns one answer per question; the Jev API engine page has the shape.

What the run did

Run offline, with the judge’s answers replayed from judge.json beside the brief and the person’s answer recorded in approve.json, the example printed:
Output
Two rounds: the first reader pass found two blocks, so the page went back with them; the second found nothing, and the judge chose holds. The record holds each round’s findings and each of the judge’s answers. With JUDGE=jev and a TypeSafe key in the environment, the same file asks Jev instead of replaying.

Next steps

  • Feedback loops: the check, the target and the stopping rule, and the other shapes a loop takes.
  • Jev API engine: the judge’s question types and how it reports its identity.
  • A person decides: the approval step this loop ends on, bound to the exact bytes.