There is no best AI judge.
No single judge is best at every check. Hosho sends each check to the one that is, and scores higher than a frontier LLM at a fraction of its latency and cost.
Start with what good looks like
Evaluation has four steps, in order.
- Define what good looks like.
- Define the metrics that make that measurable.
- Pick and design a judge for each metric.
- Calibrate each judge against real answers, and keep tracking it.
A judge is designed, not picked: what it answers well, what it costs and how long it takes all change with the check. This article shows how sensitive that design is, using five public datasets run through every judge that could answer them and scored against the datasets’ own answers. A different judge won each one.
Hosho’s judges
Hosho evaluates with three kinds of judge, plus people, and picks among them check by check. The colours below are used in every chart that follows.
1. Different checks need different judges
No judge came top on more than one of the five datasets.
- Code won where the answer is a count. The strongest LLM failed more than a third of responses that followed every instruction.
- LOGEN won on contracts, at a sixtieth of the frontier LLM’s cost.
- Jev tied the frontier LLMs on which answer is better, at an eightieth of the cost.
- The frontier LLMs won on tables and long grounded answers, the checks that need several things weighed at once.
| Task | Code | LOGEN | Jev | Frontier LLM (lite) | Frontier LLM (flagship) |
|---|---|---|---|---|---|
| Did it follow the instructions?IFEval | 100 | 69 | 70 | 76 | |
| Is the claim true of this table?TabFact | 91 | 59 | 88 | 87 | 94 |
| Is the claim in the source?RAGTruth | 58 | 75 | 75 | 83 | 82 |
| Which of two answers is better?MT-Bench | 69 | 83 | 83 | 82 | |
| Does the contract say this?ContractNLI | 86 | 80 | 81 | 82 |
2. Hosho sends each check to the right judge, and does better for less
Hosho routing scores 88 against 83 for the best frontier LLM, at a third of the cost.
- Every check goes to the judge that wins it: code for instructions and tables, LOGEN and Jev for grounding and contracts, Jev for which answer is better.
- The frontier LLM is called only where the two fast judges disagree.
- Routing does not win every row. On grounding it scores 80 against 83. It wins on the average, and on cost.
3. Only a trained model gives the same verdict twice
A judge that changes its mind cannot gate a release: the same build can pass on Monday and fail on Tuesday.
- LOGEN gave the same verdict on every rerun.
- Every other judge changed some: the flagship frontier LLM 6.4%, Jev 4.4%.
- Small and fast is not what makes a judge stable. Being trained on one question is.
Good judges make better products
Judge design is the hard part of evaluation. It is not picking a model; it is deciding, check by check, what has to be exact, what can be learned, and what needs judgment.
A wrong judge is a wrong signal, and every loop built on it starts working against you.
Get it right and every score is a signal you can trust and build on. That is what we do at Hosho — come and talk to us.
About the evidence. Tests run 1–2 October 2026 on five public datasets: IFEval (Zhou et al., 2023), TabFact (Chen et al., 2020), RAGTruth (Niu et al., 2024), MT-Bench human judgments (Zheng et al., 2023) and ContractNLI (Koreeda and Manning, 2021). 532 to 600 items each; every verdict scored against the dataset’s own answers. The core comparisons were planned and fixed before any result was read. The routing shown sends which-answer-is-better to Jev; the configuration fixed in advance sent it to the frontier LLM and scored the same, 88, at $2.29 per 1,000. Two of our eight planned predictions failed: we expected the flagship model to beat Jev on which answer is better, and to beat our model on contracts. IFEval responses were written by Llama 3.1 8B; its answers come from the benchmark’s official checker, so code is the reference there by definition. The RAGTruth results were first read in an earlier pilot. MT-Bench answers are single expert votes with ties removed. Costs are list prices as metered; LOGEN is GPU time on a cloud T4. Judges: Claude Sonnet 4.6 and Claude Haiku 4.5 at temperature 0; fast judges are LOGEN (Hosho, through its grounding service) and Jev 1.13 (TypeSafe). In the routing, grounding and contracts used both fast judges. Full figures and set-up available on request.
Outside results: MiniCheck (Tang, Laban and Durrett, 2024); LLM judge biases (Zheng et al., 2023).