Pioneers Insight Method Research Author
No Priors Ep. 97 | With Decagon CEO and Co-Founder Jesse Zhang
Back to Episodes

No Priors Ep. 97 | With Decagon CEO and Co-Founder Jesse Zhang

Summary

  • Customer support is Decagon’s “golden use case” for AI agents because automation and customer outcomes are both directly measurable. Jesse Zhang says buyers track the fraction of conversations handled, CSAT or NPS, and—especially in regulated industries—accuracy. The result can be a “personal concierge” available in any language, 24/7, with potential gains in retention and conversion alongside labor savings.
  • Bilt Rewards stopped scaling its support team within roughly one month of adopting Decagon and has since recorded around 65 agents of headcount saved. Its support volume had been growing with its rapidly expanding user base; automation let Bilt restructure the operation while making responses faster. Zhang calls the case’s ROI “very easy.”
  • The key differentiation sits above foundation models, not in exclusive access to them. Decagon combines multiple models through eval-driven orchestration, molds them around each customer’s business logic, and exposes the data, steps, and knowledge gaps behind responses. “You really don’t want this to feel like a black box.”
  • For customer-service agents, instruction following matters more than the coding and math gains dominating model discourse. Zhang welcomes advances from o1 and Sonnet, but says the decisive capability is whether a model can take a support SOP or workflow and “follow them to a T.” That leaves meaningful application-layer work even as core models improve.
  • The near-term agent winners must support gradual deployment and produce an easily quantified return before reaching perfection. Customer support and coding pass that test; security may not because missing one subtle event can be unacceptable, while text-to-SQL often remains a supervised copilot with murky pricing power. Zhang is therefore “more bearish on a lot of these AI agent use cases in the near term.”
  • Voice, screen context, and agent supervision are key future areas for Decagon. Voice-to-voice models reduce latency, but production workflows may still require data retrieval and multiple calls; Zhang says computer use is probably not production-ready yet. Longer term, he expects more humans to supervise and edit infinitely scalable agents, making observability and control central product priorities.

Deep dive

1. Customer pull selected support as Decagon’s agent wedge

  • Decagon was founded in August 2023. Zhang’s first company was acquired by Niantic; when he and his co-founder Ashwin started Decagon, his biggest lesson was that founders “can’t really overthink things too much.” They began with broad interest in agents, talked to customers, and let those conversations identify customer service as the strongest initial use case.

  • The fit rests on the fact that the use case is tailor-made for what LLMs are good at. Decagon now targets companies with sizable support operations, spanning fast-growing startups and large enterprises.

  • Transparency became the defining requirement. Customers need to inspect which data produced an answer, what steps the agent took, and whether they can provide feedback: “It’s very important for them that the AI agent is not a black box.”

2. Support automation produces unusually legible economics

  • Elad Gil’s benchmark was Klarna: 2.3 million chats in four weeks, satisfaction on par with humans, 25% fewer repeat inquiries, and two-minute resolution versus 11 minutes for a human. AI also enabled 24/7 service across 23 markets and 35 languages while 700 full-time agents shifted to other work.

  • Zhang’s scorecard has two leaders: what fraction of total conversations the agent completes, and whether customers become happier through CSAT or NPS. Regulated buyers add accuracy constraints, but the broader upside includes lower cost, faster service, higher retention, and more conversions.

  • Bilt Rewards supplied Decagon’s clearest case study. Because inquiries rose with its quickly growing user base, Bilt initially feared being overwhelmed; within about a month it stopped scaling the support team. Nearly a year later, the published case study put saved headcount at around 65 agents, alongside a “snappier” customer experience.

3. Application-layer differentiation comes from orchestration, evaluation, and operational control

  • Everyone can access models such as GPT-4o, GPT-4, and Claude Sonnet, so Zhang describes Decagon as a software company using models as tools. He puts the special sauce in orchestration and the surrounding software rather than in access to the models themselves.

  • The orchestration layer can use evals to measure how models perform at particular tasks, combine them, and shape the resulting system around each customer’s business logic. That orchestration will differ by category: a support agent and a coding agent require different structures even when they share underlying models.

  • The surrounding software must also interpret operations at scale. If a customer has one million conversations, no human will read them all; the system should surface major categories, trends, missing knowledge, and other gaps while explaining the data and steps behind individual decisions.

  • Buyers can put an agent into production for 1% of volume, compare results with human performance and other options, and expand deployment from there. Zhang says Decagon’s advantage in those benchmarks comes from “observability, explainability, control,” though he concedes there is “still a long way to go.”

4. Instruction following and latency define the technical frontier

  • Zhang separates quantitative reasoning from instruction following. Recent models have improved markedly at coding and math, but customer support depends more on faithfully executing detailed SOPs—following them “to a T.” That capability, rather than abstract reasoning alone, is what he wants major labs to advance.

  • Voice is another channel for the same underlying customer problem, alongside chat, email, and SMS. Decagon began with text because it was easier for customers to evaluate, but customers that have seen text agents work are now testing voice agents built with companies including ElevenLabs, OpenAI, and Cartesia.

  • The architectural trade-off is latency versus computation. Voice-to-voice models respond quickly, but production cases may need data retrieval and multiple model calls; speech-to-text, text processing, and regenerated voice add delay. One practical bridge is conversational cover—“Give me a second. I’m looking up your data”—while work continues.

  • Beyond voice, Zhang wants agents to use context from a user’s entire screen and interaction history, then help navigate software directly. Anthropic’s computer-use demo illustrates the direction, but he judges it probably not production-ready yet.

5. Viable agents need gradual rollout and measurable ROI

  • Zhang’s retrospective filter has two parts: an agent must deliver value before it is nearly perfect, and the buyer must be able to quantify that value. He believes “the vast majority of use cases right now” still lack commercial readiness under current models.

  • Security shows the perfection problem. Although large log streams look well suited to AI, the job may require catching every subtle event; model nondeterminism makes enterprise buyers reluctant to trust an agentic solution. Adoption there could be “very, very slow”—much slower than impressive demos imply.

  • Text-to-SQL shows the measurement problem. Buyers may like the output yet still require a person to monitor and edit it, turning the system into a copilot. If a claimed AI-agent data scientist probably cannot replace a real one—and teams employ few data scientists anyway—the vendor struggles to justify a large contract.

  • Coding is the contrasting positive example: teams can section off selected tasks, let an agent attempt them, and receive useful output without handing over everything. Better models will unlock more categories, but Zhang’s near-term conclusion remains deliberately cautious.

6. Agent supervision becomes a new form of work

  • As agents proliferate, Zhang expects people to spend more time supervising and editing them. Unlike human workers, agents are “infinitely scalable,” and some behaviors can be hard-coded, creating new possibilities for real-time feedback, monitoring, and control.

  • Decagon’s own product direction follows that shift: human agents and leadership teams should be able to inspect performance, intervene, and make changes. Zhang treats that control layer as the company’s current differentiation and an area for further innovation.

  • Zhang also describes the math- and coding-contest community as an informal founder network. Members angel-invest in one another, exchange operating advice, and socialize, including through a Chinese version of bridge; the background is a useful hiring signal, but he stresses that Decagon’s hiring process is broadly the same and its talent pool extends far beyond contest participants.