← Research
Research

There is no best AI judge.

No single judge is best at every check. Hosho sends each check to the one that is, and scores higher than a frontier LLM at a fraction of its latency and cost.

Oct 2026·~4 min read

88 vs 83
Accuracy: Hosho routing against the best frontier LLM on everything, across five public datasets.
34%
Of the cost of running the frontier LLM on everything.
73%
Of checks need no frontier LLM at all.

Start with what good looks like

Evaluation has four steps, in order.

  1. Define what good looks like.
  2. Define the metrics that make that measurable.
  3. Pick and design a judge for each metric.
  4. Calibrate each judge against real answers, and keep tracking it.

A judge is designed, not picked: what it answers well, what it costs and how long it takes all change with the check. This article shows how sensitive that design is, using five public datasets run through every judge that could answer them and scored against the datasets’ own answers. A different judge won each one.

Hosho’s judges

Hosho evaluates with three kinds of judge, plus people, and picks among them check by check. The colours below are used in every chart that follows.

Deterministic checks. Rules and short programs for anything defined by a count, a format, a threshold or a sum. Exact, instant and free.
Fast judges. LOGEN, Hosho’s proprietary fine-tuned grounding model, alongside Jev, a general decision model for typed questions. One question at a time, at low latency and low cost.
Frontier LLMs. For the checks that need several things weighed together, and for explaining what went wrong.
People. The answers every other judge is calibrated against.
Figure 1: One judge for everything, against splitting the item into checks
ONE JUDGE FOR EVERYTHINGOne iteman answer to evaluate“is it good?”everything, one callFrontier LLMreads, weighs, decidesevery check at onceOne verdict83avg · $4.32 / 1kHOSHO: SPLIT, ROUTE, COLLECTOne itemthe same answer,the same questionsplitThe checks inside iteach one a question with one kind of answerright length and formatevery number matches the sourceeach claim is stated in the sourcetone, fit and overall judgmenta sample of all of the above73% of checks need no frontier LLMroute each checkDeterministicexact · $0 · instantFast judgesLOGEN, Jev · ~1 sFrontier LLMfor judgment onlyPeoplecheck every judgeEvery verdict,one resultwhich claim failed,which number changed,the LLM’s view whereit was needed88avg accuracy34% of the cost

1. Different checks need different judges

No judge came top on more than one of the five datasets.

  • Code won where the answer is a count. The strongest LLM failed more than a third of responses that followed every instruction.
  • LOGEN won on contracts, at a sixtieth of the frontier LLM’s cost.
  • Jev tied the frontier LLMs on which answer is better, at an eightieth of the cost.
  • The frontier LLMs won on tables and long grounded answers, the checks that need several things weighed at once.
Figure 2: A different judge wins each task
Balanced accuracy for each judge on each of five datasets.
TaskCodeLOGENJevFrontier LLM
(lite)
Frontier LLM
(flagship)
Did it follow the instructions?IFEval100697076
Is the claim true of this table?TabFact9159888794
Is the claim in the source?RAGTruth5875758382
Which of two answers is better?MT-Bench69838382
Does the contract say this?ContractNLI86808182
Balanced accuracy against each dataset’s own answers. Filled = best; outlined = within a point of the best. Blank = the judge cannot be asked this question. Code column: IFEval’s own checker, a model-written program run on each table, an every-number-appears check on grounding, longer-answer-wins on preference. LOGEN only answers whether a claim is stated in a text, so two rows are blank for it.

2. Hosho sends each check to the right judge, and does better for less

Hosho routing scores 88 against 83 for the best frontier LLM, at a third of the cost.

  • Every check goes to the judge that wins it: code for instructions and tables, LOGEN and Jev for grounding and contracts, Jev for which answer is better.
  • The frontier LLM is called only where the two fast judges disagree.
  • Routing does not win every row. On grounding it scores 80 against 83. It wins on the average, and on cost.
Figure 3: Accuracy and cost, Hosho routing against one judge for everything
808590$5$4$3$2$1$0Cost per 1,000 checks, cheaper to the right →Accuracy ↑BEST ↗Frontier LLM (flagship)Frontier LLM (lite)Hosho routing
Average over the five datasets, against cost per 1,000 checks. Up and to the right is better.

3. Only a trained model gives the same verdict twice

A judge that changes its mind cannot gate a release: the same build can pass on Monday and fail on Tuesday.

  • LOGEN gave the same verdict on every rerun.
  • Every other judge changed some: the flagship frontier LLM 6.4%, Jev 4.4%.
  • Small and fast is not what makes a judge stable. Being trained on one question is.
Figure 4: Verdicts that changed on a rerun
LOGEN
0%
Frontier LLM (lite)
3.2%
Jev
4.4%
Frontier LLM (flagship)
6.4%
0%5%10%
The same items, run three times. Bars are coloured by the kind of judge, as set out above.

Good judges make better products

Judge design is the hard part of evaluation. It is not picking a model; it is deciding, check by check, what has to be exact, what can be learned, and what needs judgment.

A wrong judge is a wrong signal, and every loop built on it starts working against you.

Get it right and every score is a signal you can trust and build on. That is what we do at Hosho — come and talk to us.

About the evidence. Tests run 1–2 October 2026 on five public datasets: IFEval (Zhou et al., 2023), TabFact (Chen et al., 2020), RAGTruth (Niu et al., 2024), MT-Bench human judgments (Zheng et al., 2023) and ContractNLI (Koreeda and Manning, 2021). 532 to 600 items each; every verdict scored against the dataset’s own answers. The core comparisons were planned and fixed before any result was read. The routing shown sends which-answer-is-better to Jev; the configuration fixed in advance sent it to the frontier LLM and scored the same, 88, at $2.29 per 1,000. Two of our eight planned predictions failed: we expected the flagship model to beat Jev on which answer is better, and to beat our model on contracts. IFEval responses were written by Llama 3.1 8B; its answers come from the benchmark’s official checker, so code is the reference there by definition. The RAGTruth results were first read in an earlier pilot. MT-Bench answers are single expert votes with ties removed. Costs are list prices as metered; LOGEN is GPU time on a cloud T4. Judges: Claude Sonnet 4.6 and Claude Haiku 4.5 at temperature 0; fast judges are LOGEN (Hosho, through its grounding service) and Jev 1.13 (TypeSafe). In the routing, grounding and contracts used both fast judges. Full figures and set-up available on request.

Outside results: MiniCheck (Tang, Laban and Durrett, 2024); LLM judge biases (Zheng et al., 2023).