← Research
Customer story

Ghost × Hosho

“This is my #1 issue right now and Hosho is directly making my product better.”

Kal Prince·Founder & CEO, Ghost

Ghost outcomes — enabled by Hosho

Improved
Product performance, confirmed by Ghost — the ingest agent was inferring detail its sources never contained; Hosho caught it and the fix shipped: “working much much better”
~2,600
Evaluations run since June
~12,500
Findings, each linked to the evidence behind it
14
Evaluators running
4
Modules under judgement
Ongoing
Hosho recommendations flowing into production changes, re-scored on ship
100%
Reachable from Ghost’s own coding agent — no dashboard required

The company

A sales graph built from the relationships you already have.

Ghost turns a company’s own conversations into a living map of its customers and relationships. Agents read meetings, emails and notes, write the facts worth keeping into the graph, and draft the outbound that follows — and customers now build their own multi-step workflows on top.

Almost none of that is checkable by eye. An invented job title, a dropped fact, a workflow step that can never run — the output reads perfectly well and is wrong. Ghost needed a definition of correct precise enough to test against, with results in their own database, reachable by their own agents.

How it runs

From every conversation to a change Ghost can ship.

Define — agreed with Ghost’s founder

Hosho reads Ghost’s pipeline, prompts and production data, and turns “this output is wrong” into something testable — the specific claim that isn’t in the source, the fact the email dropped, the step that can never run.

Build evals — calibrated against Ghost’s own judgement

Each definition becomes an evaluator — deterministic checks where the answer is knowable, model judges where it isn’t — tuned against samples Ghost has already labelled by hand, until the scores agree with theirs. 14 evaluators across 4 modules.

Evaluate — across every surface Ghost ships

Chat conversations scored hourly for frustration and unsupported answers. Ingestion scored per run for invented details and mis-scoped actions. Workflows checked the moment they’re saved — goal coverage, flow validity, step-type fit — so a broken one is caught before it ever runs.

Explain — into Ghost’s own stack

Every finding names the failing quote and the reason. Results are written into Ghost’s own database, not ours — so their coding agent queries them directly over MCP, and flagged sessions arrive in Slack without anyone opening a dashboard.

Re-measure — the same bar, before and after

A change is only finished when the number moves. A fix is scored on the same sample that failed, before and after; a pull request can be run against the evaluators the same way, to show whether it improved the product or quietly cost something.

The results

What Ghost’s tech team has to say:

  • Evals they trust— “The findings are genuinely good — I checked and agreed with them all.”
  • Complex workflow agents getting better— “The new ingest agent is working much much better from Hosho’s fixes we merged.”
  • Accelerated speed of build— “You guys have dramatically accelerated how I can build our new knowledge graph.”
Want results like Ghost’s?
Talk to us →
Hosho is the independent evaluation and fix layer for enterprise AI.