Ghost × Hosho
“This is my #1 issue right now and Hosho is directly making my product better.”
Ghost outcomes — enabled by Hosho
The company
A sales graph built from the relationships you already have.
Ghost turns a company’s own conversations into a living map of its customers and relationships. Agents read meetings, emails and notes, write the facts worth keeping into the graph, and draft the outbound that follows — and customers now build their own multi-step workflows on top.
Almost none of that is checkable by eye. An invented job title, a dropped fact, a workflow step that can never run — the output reads perfectly well and is wrong. Ghost needed a definition of correct precise enough to test against, with results in their own database, reachable by their own agents.
How it runs
From every conversation to a change Ghost can ship.
Define — agreed with Ghost’s founder
Hosho reads Ghost’s pipeline, prompts and production data, and turns “this output is wrong” into something testable — the specific claim that isn’t in the source, the fact the email dropped, the step that can never run.
Build evals — calibrated against Ghost’s own judgement
Each definition becomes an evaluator — deterministic checks where the answer is knowable, model judges where it isn’t — tuned against samples Ghost has already labelled by hand, until the scores agree with theirs. 14 evaluators across 4 modules.
Evaluate — across every surface Ghost ships
Chat conversations scored hourly for frustration and unsupported answers. Ingestion scored per run for invented details and mis-scoped actions. Workflows checked the moment they’re saved — goal coverage, flow validity, step-type fit — so a broken one is caught before it ever runs.
Explain — into Ghost’s own stack
Every finding names the failing quote and the reason. Results are written into Ghost’s own database, not ours — so their coding agent queries them directly over MCP, and flagged sessions arrive in Slack without anyone opening a dashboard.
Re-measure — the same bar, before and after
A change is only finished when the number moves. A fix is scored on the same sample that failed, before and after; a pull request can be run against the evaluators the same way, to show whether it improved the product or quietly cost something.
The results
What Ghost’s tech team has to say:
- Evals they trust— “The findings are genuinely good — I checked and agreed with them all.”
- Complex workflow agents getting better— “The new ingest agent is working much much better from Hosho’s fixes we merged.”
- Accelerated speed of build— “You guys have dramatically accelerated how I can build our new knowledge graph.”