← Research
Customer story

AppSmith × Hosho

“Without a doubt, anybody who’s serious about agents really needs this.”

Abhishek Nayak·Co-founder & CEO, AppSmith
Abhishek Nayak on how AppSmith ships Kite with Hosho. Watch on Drive →

Kite outcomes — enabled by Hosho

36%
Reduction in issue incidence — experience issues identified, root-caused and fixed across live build sessions
~5×
Increase in win rate vs. competition — on the primary quality dimension, design and visual distinctiveness
100,000+
Signals scored online — every live session, customer and agent
~10,000
Offline quality judgments, including a weekly competitor benchmark
~15,000
Prompt optimization calls, in CI and via MCP

The company

Kite: an agentic marketing team living in Slack.

AppSmith has launched Kite, an agentic marketing team living in the customer’s Slack. Kite researches the business, sets the strategy and ships the work: websites, new pages, campaigns, outreach and weekly plans — created and edited in chat, and continuously improved.

AppSmith × Hosho: six months embedded in how AppSmith ships Kite. An aligned definition of what good looks like, evaluation signals the team trusts, every session scored online across customer and agent signals, every tool call judged, and websites benchmarked weekly against competitors to keep extending Kite’s lead. Every finding root-caused and shipped as a fix into Kite’s coding agents.

How it runs

From every session to a validated fix — in Slack, Git, the IDE and the Hosho UI.

Define — aligned with AppSmith’s CEO + CPO

Agentic onboarding reads AppSmith’s traces, codebase and business context. Quality becomes measurable: twelve experience dimensions per customer engagement, nine quality dimensions per website.

Evaluate — calibrated with the CTO + AI leads, reviewed in Slack

Deterministic checks + LLM judges for every dimension, tuned until the team trusts the scores. Online: every live session scored. Offline: ~10,000 judgments across ~1,000 websites, including the weekly competitor benchmark. Preflight: pull requests run against the evaluators on a build preview, before they merge.

Triage — agent world model, not guesswork

A failure is carried past the symptom to its true root cause in Kite’s pipeline. The world model links what the customer experienced to the culprit: infra, orchestration, context, prompt or model choice.

Fix — into AppSmith’s engineering flow

Shipped as pull requests to Git and as root-cause insights into AppSmith’s coding agents. Proactive prompt improvements in the IDE via MCP, grounded in Hosho’s model research. Every fix validated by re-testing the cases that failed.

The results

  • Quality keeps compounding:issue incidence across build sessions is down 36%, and Kite’s win rate against competing builders on design is up roughly 5× in six months.
  • Strong release validation: pull requests can be preflighted against the same evaluators before merge, and every Hosho fix is validated by re-testing the cases that failed.
  • The team stays on product:complex evaluation, triage and fix run on the Hosho platform rather than on engineers’ time — 1–2 FTEs freed to keep building Kite.
Want results like AppSmith’s?
Talk to us →
Hosho is the independent evaluation and fix layer for enterprise AI.