The Applied Laboratory for AI Quality.
Evaluating and improving agentic systems is an open problem. These systems fail in ways that are hard to observe, hard to attribute, and harder still to fix — and the field’s methods for measuring them are younger than the systems themselves.
Hosho Research works on this problem across its parts, from how evaluators should be built to how their findings should drive repair, and what we learn becomes the platform.
We’re happy to discuss this work, share the harness, or hear what you’re working on: nitish@hoshoai.com
Writing
Essays, customer stories, and research — our point of view on building independent, human-grade evaluation systems for production AI.
Why we exist
The hard part of scaling AI isn’t what you can ship — it’s what you can trust.
AppSmith × Hosho
How a leading low-code platform cut issue incidence by 36%.
Ghost × Hosho
Independent evaluation for AI used by professional publishers.
Research: The Anatomy of a Blind Spot.
A better judge won’t fix your evals. A world model will.