See your AI the way your customers do — and make it better.

The independent evaluation and fix layer for enterprise AI.

aligned · live

v1 is live — 2 customers · 3 products · 100,000+ outputs scored.

What our customers say

“Without a doubt, anybody who’s serious about agents really needs this.”
Abhishek Nayak
Abhishek Nayak
Co-founder & CEO, AppSmith

Leading low-code platform · 100K+ developers · Kite, their AI marketer + website-building agent

36%
Reduction in issue incidence — experience quality
~5×
Increase in win rate vs competition — design & visual distinctiveness
100,000+
Signals scored online — every live session, customer and agent

Products

Three products, one engine.

01

Hosho Agent Optimizer

Continuous evaluation and improvement of AI output — metrics, calibrated judges, root-cause diagnosis, fixes, delivered in Slack.

Kills: “Is the output actually good enough to ship?”

LangfusePostgresPJR
Explorer23 sessions · 18 tasks in scope · 9 clean · 11 warnings · 3 failing7 days30 days90 daysAll history
The lenses · the five things every session is judged on
Lens 01
Behaves truthfully
Is every claim backed?
0% pass
84 of 221 flagged
Lens 02
Does what was asked
Did it do the task?
0% pass
97 of 231 flagged
Lens 03
Runs its team well
Are handoffs clean?
0% pass
41 of 172 flagged
Lens 04
Good to work with
Is it a good session?
0% pass
9 of 188 flagged
Lens 05
Produces good work
Is the work any good — vs the best alternative?
0% pass
9 of 38 flagged
Primary drivers of gaps
result_groundinglens 138 units
policy_compliancelens 121 units
handoff_timinglens 317 units
+ Add cardMetric · Lens · Module
Recent evaluation runs
RunUnitsClean shareDateReport
nightly-regression3155% (17 of 31)30 Augfull report →
Weekly benchmark2236% (8 of 22)28 Augfull report →
probe-2026-08-141471% (10 of 14)14 Augfull report →
probe-2026-08-07580% (4 of 5)7 Augfull report →
probe-2026-07-31967% (6 of 9)31 Julfull report →
Ask Hosho
02

Hosho Signals

How customers actually experience your AI — frustration, dead-ends, promoter/detractor sessions, issues triaged and fixed.

Kills: “Are customers quietly struggling — and would we even know?”

InspectorLangfuse
JourneyAppsSignalsControlDatabaseSupport agentDocs agentSoon
— 128 customers · 96 agents · since 9 Jul · updated 4m ago
Discovery
stopped before anything was built
1,204 reached
0%
dropped off · 0 still active
What they hit before leaving
latency_spike968%
repeated_ask615%
slow_first_reply444%
Setup
stopped while configuring
987 reached
0%
dropped off · 0 still active
What they hit before leaving
rage_quit14214%
dead_end889%
context_drop525%
First run
stopped mid-task
681 reached
0%
dropped off · 0 still active
What they hit before leaving
retry_loop7411%
vague_answer528%
clarification_loop315%
Habit
stopped coming back
531 reached
0%
dropped off · 0 still active
What they hit before leaving
stalled_task397%
unresolved_handoff183%
silent_abandon122%
03

Hosho Prompt & Model Intelligence

Reads how models actually interpret your prompts — prompt demand vs model capability, cheapest model that clears the bar, prompt optimization in your workflow.

Kills: “Are we paying for tokens we don’t need?”

How it works

From raw traces to a merged pull request.

01Your AI
Your AI in production
02Ingest
Traces · messages · tool-call logs · approvals · reactions
03Agents
Orchestrator sequences:
Onboarding — writes the rubricEvaluator — calibrates judgesTriage — finds root causeFixer — opens the PR
04Engines
Evaluation engine — deterministic evaluators + LLM judges
Triage engine — deterministic chassis + evidence agent
05Outcomes
Verdicts per evaluator per record · diagnosis with root cause + confidence
Fix PR → human merge — never applied silently

Use Cases

What you can do with Hosho.

01
Get pinged when quality drops
Issues land in Slack with the evidence attached — not in a dashboard you forget to check.
When Hosho picks up a regression it posts straight to your team channel — what broke, how confident it is, and how many incidents it has seen — then follows up with the root cause and the fix PR. The evidence lands where your team already works.
Hosho
Channels
#general
#engineering
#ai-quality3
#deploys
#incidents
Direct Messages
Jamie L.
Alex T.
#ai-qualityHosho workspace
6 members
H
HoshoAPP9:14 AM

Hey Jamie — picking up a quality regression on refunds. The assistant's been giving incorrect refund guidance, looks like it's mishandling refund-policy questions when context is thin.

Confidence: high4 incidents today
JL
Jamie L.9:17 AM

Yeah, couple of customers flagged this too. Can you dig in and find the root cause?

H
HoshoAPP9:17 AM

On it. The prompt's falling back to a generic handler when refund context is missing — and retrieval order isn't surfacing the right policy docs first. I'll open a PR with the fix.

PR opening on GitHub
Message #ai-quality
02
Stop losing customers to poor AI performance
See every user session and the signals that predict the outcome.
Hosho surfaces every user session with the signals that predict the outcome — a vague answer, dropped context, or a repeat question. See the friction before your users complain about it.
SessionSignalsVerdict
Session detail
What broke
03
Replace vibe checks with real scale evaluations
One rubric, applied to everything.
Define a rubric once and Hosho applies it to every response, at any volume. You stop spot-checking and start seeing the full picture — which responses meet your standard, and which don't.
Tone · Rubric at scale
5Warm, acknowledges the issue
4Friendly, polite
3Neutral
2Curt or dismissive
1Rude or defensive
12,480 responses graded on tone
avg 4.3/5 · 6% scored ≤2, flagged for review
04
Never ship a regression
Every version scored across your rubric.
After every deployment, Hosho re-scores your real conversations across every rubric dimension. A version-by-version matrix shows exactly where a score dropped, so you catch the regression before it reaches your users.
Regression Matrix
Dimensionv1v2v3v4
Accuracy74✓ 888687
Tone6872✓ 8180
Context fidelity6165✓ 7963regression
Resolution707578✓ 84
05
See how you stack up vs competition
Dimension by dimension, with the why.
Run the same conversations through your system and your competitors’. Hosho scores everyone on your rubric, dimension by dimension, and tells you the behavioural reason for each gap.
Benchmark Comparison
DimensionYouComp AComp BComp C
Accuracy✓ 95888279
Tone78 74✓ 9172
Context fidelity71 68✓ 8476
Resolution✓ 89807773
Key Gap
You trail Comp B on tone & context fidelity — its replies are friendlier (more ‘you’ language) and acknowledge the problem before resolving.
06
Higher-quality signal for improvement loops
Calibrated verdicts your team can act on directly.
Every session gets a calibrated judgment — promoter, neutral, detractor — with the evidence attached. Improvement loops stop arguing about whether the signal is real and start acting on it.
Session verdicts
Refund request, thin contextDetractor
Pricing questionNeutral
Onboarding walkthroughPromoter
07
Reduce your team’s workload
A prebuilt platform, evaluation best practices embedded.
No harness to build and no eval infrastructure to maintain. Hosho ships with calibrated judges, rubric workflows, and fix pipelines already wired together — your team reviews outcomes instead of building tooling.
Out of the box
Calibrated judges, tuned to your rubric
Session monitoring with verdicts and signals
Root-cause triage with evidence attached
Fix PRs opened for human review

Research

The research behind the platform.

Evaluating and improving agentic systems is an open problem. These systems fail in ways that are hard to observe, hard to attribute, and harder still to fix — and the field’s methods for measuring them are younger than the systems themselves.

Hosho Research works on this problem across its parts, from how evaluators should be built to how their findings should drive repair, and what we learn becomes the platform.

We’re happy to discuss this work, share the harness, or hear what you’re working on: nitish@hoshoai.com

Stop relying on vibe checks.

Get independent, human-grade evaluation into your AI pipeline today. v1 is live with paying customers.

Built with Kite