The first principle is that you must not fool yourself—and you are the easiest person to fool.
The first principle is that you must not fool yourself—and you are the easiest person to fool.
You ask AI for a draft, read it over, and it looks fine. Later, a reader points to a vague claim that your review did not flag. That miss gives you a case with a known answer, not a reason to trust a different model by default. A fresh context or another model creates a candidate second opinion. This subchapter shows how to test whether that candidate catches verified shortfalls without manufacturing them.

Ask a model to write something, then ask either that session or a fresh one to review it. Those are different review conditions because one has the drafting exchange and one does not. The separation itself says nothing about accuracy. Run both conditions against the same blinded case pack if you want to know which one makes fewer consequential errors on your task.
Build a small set before you choose the editor. Include representative drafts with errors you already verified, clean drafts that should receive no finding, and at least one close case. Keep the answer key outside the candidate's context, run each candidate with the same criteria, then compare its findings with the hidden labels. Record misses and false alarms separately. The reviewer earns the role when it catches useful failures and leaves clean work alone, not when its answer merely sounds sharper.
Use checks that inspect different evidence. A source comparison asks whether claims match the record. A deterministic test evaluates an exact rule. A candidate editor evaluates the criteria in its prompt. Choose the evidence source for the failure, then measure whether the reviewer recognizes it on blinded cases.
Give the editor explicit criteria, then let each criterion pass, fail, or remain unclear. Requiring a problem in every category rewards the model for inventing criticism. Require a quotation or source comparison for every proposed failure, plus met or unclear when no supported failure is available. That format makes the output auditable; it does not make the verdict correct. The editor should omit any flag it cannot tie to the text or the supplied sources.
The pass runs in three stages: test the editor on blinded cases whose labels you keep separately, review the real draft against named criteria, and verify each finding before revision. If you revise the editor after seeing a result, test that revision on fresh blinded cases. That sequence keeps a sharper-sounding critique from becoming authority by style alone.
Where the morning brief from earlier chapters was a setup you build for yourself, the running example here is a deliverable you send outward: a client proposal, a report, a strategy memo, the kind of document where a stray error reaches someone whose opinion of you it would dent. Watch the same draft handled the do way and the don't way to see the pass earn its cost. The don't way is the loop you may already run: ask one model for the proposal, skim it, ask that same model 'is this good?', get back 'yes, with minor polish,' and send it. The do way drafts the proposal with one model, then hands it to a calibrated editor, which comes back flagging that the budget paragraph states a number the source material never gave and that the recommendation does not follow from the two findings above it. You verify each flag against the budget source and the reasoning in the draft before deciding what to revise.
There is a tempting shortcut that undoes the whole pass: asking the second model to rewrite the draft. A rewrite hands you a different output; an editorial critique hands you a list of fixable problems in the one you have. When the editor rewrites, it buries the original's weaknesses inside its own writing, and you lose the thing you came for, which was a clear view of what was wrong. So the editor's job is to point, specifically, at what is weak, and to leave the fixing to you. Named categories and an evidence rule also make each finding easier for you to score. Whether they improve error detection is a result to measure on blinded cases, not a benefit to assume.
Separating writing from editing gives each context one clear job. answers the next question: whether the editor actually adds useful coverage. Role separation is a design choice; measured misses and false alarms are the evidence that the design works. You remain accountable for what ships because the reviewer proposes findings and the cited draft or source establishes whether each one is true.
A calibrated reviewer costs an extra pass, so it is not free, and not every task earns it. The two-model pass is a tool for work that matters, not an efficiency tool for everything. Run it on the documents where an error would be expensive or embarrassing: client-facing work, research reports, important emails, anything that gets published, the strategic recommendations someone will act on. For lower-consequence internal work, you may choose a lighter review because the extra pass costs time. That is a consequence-and-cost decision, not evidence that one model is accurate enough.
A clean way to sort which work clears the bar is a question rather than a rule. Ask whether you would be embarrassed to have shipped this with an error you could have caught. If yes, the work is worth a second model's different eyes. If no, the extra pass is system you do not need, and the lighter trade-off applies that runs through this whole book: use enough check to protect the work, not so much that the checking becomes the work. The reviewer is there to keep the output honest, not to turn a quick task into a project.
Guneet Kohli · 2026 · Apple Machine Learning Research
Across nine frontier judges from seven model families, correlated mistakes reduced the panel to roughly two independent votes and left its accuracy 8 to 22 percentage points below the independent-voting ideal.
Using another model family creates a candidate reviewer, not independent evidence. Representative labeled cases must establish whether the reviewer adds useful coverage for the task.