Judge a voice agent by what it does, not just how it sounds.
Tone, fluency and latency matter, and most evals stop there. The failures that cost money are actions: a booking nobody agreed to, a script read over an emergency. Hosho evaluates what the agent actually did. Here is what that found in a dental clinic’s booking agent.
1. What good looks like for this agent
Riya books appointments for a dental clinic. A good turn means she did these four things:
- Books only after the caller says yes to a specific slot.
- Follows the caller, not her own checklist.
- Follows the clinic’s rules: confirm before booking, phone number optional.
- Recognises when the call is not a booking at all, and acts.
Measured on these four, the difference between two agent designs is stark: standard, where the LLM decides everything, and decision enforced, where a small classifier (Jev) picks the move and the code takes away any tool that move does not allow. The first fails all four. The second fixes them.
2. Where the standard agent failed
It broke all four, while sounding polite and fluent throughout.
Example: a caller with a swollen face, fever and trouble swallowing. NHS and SDCEP guidance treats that as an emergency, hospital now. The agent was never given emergency handling; its one stated goal is to book.
Standard agent carries on with the script
Decision enforced safety first
3. Why the prompt could not fix it
The rule was already in the prompt. The agent broke it anyway. We tried the same rule, “confirm before booking”, three ways.
A rule the model can override is advice. The only version that held was the one it could not override.
4. What fixed it
Decide the move before the LLM speaks, and remove the tools that move does not allow.
- Faster because the LLM does less. It gets one instruction for the turn and only the tools that move allows, so it stops weighing every rule and every tool before it speaks. The decision costs 0.55 s on a pre-warmed connection; the LLM gives back more than that.
- The consent check is calibrated on real calls. Before any booking, Jev is asked how confident it is that the caller just agreed to this exact slot. Genuine agreements score around 0.7 out of 1; early booking attempts score under 0.05. Booking is held below 0.25, so the two never overlap.
- It fails soft. If the decision model errors or times out, the turn runs exactly as before, so an outage cannot stall a call.
Example: the caller never offers a phone number. The clinic’s script says the number is optional and to move on without it. The standard agent asks for it four times. The enforced agent follows the caller.
Standard agent asks for a number four times
The call ends with no appointment.
Decision enforced follows the caller
Booked, after the caller’s yes.
Evaluate the action, not the transcript
Good, for a voice agent, is defined by what it does: who it books, when, and with whose consent. A transcript score cannot see any of that.
Score each turn on what the agent did and what it was allowed to do, and the fix follows from the measurement.
That is what Hosho does. Come and talk to us.