
Jesse Zhang
Frontier Insights
Frontier Thesis: Agentic enterprise AI wins not on proprietary foundation models, but on defensible application orchestration—eval-driven execution, deep business logic, and observability that deliver immediate, auditable balance-sheet ROI.
Strategic Playbook: Target high-volume, measurable pain points like customer support first. Decagon drives enterprise lock-in by replacing headcount growth with verifiable resolution workflows, demonstrated by Bilt halting support hiring within a month and saving roughly 65 FTEs.
Risks & Warnings: Multimodal expansion (voice, screen-context) introduces severe latency and reliability penalties. In low-fault-tolerance domains, poorly quantified agent ROI will stall autonomous adoption, trapping solutions as low-margin copilots.
Key Views & Dialogues
How Decagon Runs 90% of Its Agents on Open-Source Models
- 🗓️ Date:
2026-07-31| 🎙️ Show:The a16z Show
Decagon runs 90% of production workflows on open-source models, where task-specific fine-tuning delivers higher accuracy, lower latency, and lower cost than frontier systems on bounded jobs. Frontier models remain the discovery engine, while Decagon’s continuously rebuilt model factory and enterprise infrastructure turn deployment pain into reusable product; migration will remain constrained by proprietary data, security, and governance.
View Dialogue Notes & Key Takeaways
Decagon now runs 90% of its workflow on open-source models because production voice agents reward task-specific accuracy and low latency, not general intelligence. Fine-tuned smaller models can outperform frontier systems on narrow jobs such as topic identification or bad-actor detection while also being cheaper and faster: “We end up getting all three things.” The remaining 10% supports new, experimental, or unusually open-ended work.
The durable model split is frontier for discovery, open source for scaled production. Frontier APIs remain the easiest way to launch an uncertain use case, while a stable workflow creates strong incentives to fine-tune and control smaller models; Decagon Autopilot still uses frontier intelligence to review as many as a million conversations, detect trends, generate variants, and test improvements. Enterprise migration will be slow because proprietary data, custom evals, security, and model-risk governance matter more than model availability.
Decagon treats its research organization as a continuously operating “model factory,” not a one-time infrastructure project. New capabilities create new tasks to automate, while stronger base models make older fine-tunes obsolete; the company therefore trains and retires models continually. It builds tightly coupled system-level evaluation internally, buys commodity labeling and dataset-diversity tools, and optimizes for the customer’s unit of value—a conversation—even as model calls and tokens per conversation rise.
The application-layer thesis is that models do not encode changing business processes, integrations, controls, or systems of record. Decagon fine-tunes for reusable customer-service behavior, but keeps each enterprise’s procedures in context so they can change without retraining; the surrounding product must handle testing, QA, compliance, collaboration, and cases such as rebooking three travelers after a canceled flight. Even with AGI, the founders argue, agents will still need software “to store work and pull information from and reason about things.”
Forward deployment is valuable only when customer pain compounds into reusable product. Early AI workflows require people to “lay out the track as they see which way the train is going,” but known workflows should be productized for the next 10 customers; otherwise, the company becomes “a glorified consulting shop.” Ashwin Sreenivas preserves Palantir’s sharper formulation: forward-deployed engineers “eat pain and excrete product.”
Duo demonstrates Decagon’s compounding product loop: an expensive human deployment task becomes a feature, then Duo Autopilot automates that feature’s ongoing improvement. The larger, slower agent can turn transcripts and documentation into agent operating procedures, integrations, tests, simulations, conversation monitoring, and drafted fixes—work the core conversational agent cannot do. Decagon’s near-term moat, Jesse Zhang argues, is the enterprise infrastructure around that intelligence, though he concedes that if agents eventually generate all such infrastructure on demand, “I don’t know, and we’ll figure out in three years.”
Decagon’s commercial wedge is a “glass box” customers can operate themselves, widening from support into an AI front door for every customer interaction. One customer reportedly left Sierra after producing three journeys in roughly a year, then built seven on Decagon within a month; founder involvement remains heavy, with Jesse Zhang estimating that sales consumes about 80% of his time. Lower service costs can also unlock latent demand rather than translate mechanically into layoffs: Ashwin Sreenivas argues that “AI will kill jobs but not careers.”
🔗 Original source & video: How Decagon Runs 90% of Its Agents on Open-Source Models
The Future of AI Agents | Jesse Zhang Interview
- 🗓️ Date:
2025-10-06| 🎙️ Show:Invest Like the Best
Decagon’s enterprise AI wedge is customer service, where existing resolution rates of 15–20% create a quantifiable path toward 50–80% while escalation to humans de-risks deployment. The moat is continuous improvement: models can review every one of a million monthly conversations, identify the 2% failing, and draft fixes that compound into a brand-level digital concierge. Voice remains the largest unresolved frontier, with voice-to-voice hallucinations probably “like 8x higher” than text, while falling model costs intensify the race for market share and mindshare.
View Dialogue Notes & Key Takeaways
Decagon is an AI customer-service agent, and Jesse Zhang’s explanation of why that use case has the most enterprise gen-AI traction is the tradeable insight: ROI is pre-quantified (existing chatbots/IVRs “resolve” 15–20% of volume; AI could take it to 50–80%, so “I’m going to take the total cost and chop off 60% of it”), and escalation-to-human is already built into call-center infrastructure, de-risking go-live. Enterprises start with 5% of traffic and roll to everything “within weeks” once resolution, CSAT, and accuracy check out.
Zhang’s labor-spectrum framework: AI agents exist to replace human labor, so map use cases by what that labor costs — AI eats the spectrum from both ends. Engineers (the most expensive) are generally augmented rather than simply fired, because “there’s infinite engineering work to do”; outsourced tier-1/2 support (the lower-paid end) can genuinely be replaced, with BPOs unbothered because turnover is high and workers shift to tasks AI can’t do yet, like data labeling.
The founding method is replicable: ask “classic sales-qualification questions but in founder form” — exactly how much would you pay, whose boss approves, how would you present ROI to leadership. Most ideas die at “$100 subscription per month” from a big company; Decagon’s conviction came when tallied willingness-to-pay was “an order of magnitude more than anything else” — even as “very smart older founders” warned the idea was too obvious and incumbents would win.
Voice is the frontier and it isn’t solved: voice-to-voice models hallucinate “probably like 8x higher” than text, so enterprises route voice→text→voice with checks. The prize is large — Fortune 100 support is 90–95% voice — and the bar is high because spoken language is a ~150,000-year-old UI versus ~60 years of keyboards.
The moat claim: an LLM now reads every one of a million monthly conversations (versus a 20-person sampling team), flags the 2% going wrong, and drafts fixes — so after a year “has your agent just continuously gotten better… to the point where it’s very difficult for another agent to come in and perform at the same level?” End state: the agent becomes the brand’s front end — a “digital concierge” replacing the app and website entirely.
On unit economics: at the application layer “your margin doesn’t really matter” right now — costs fall exponentially, so optimize for market share and mindshare — though enterprise deals shouldn’t hemorrhage cash and Decagon’s one rule is never negative margins. His croissant analogy: the last step that actually solves the problem captures the most margin, which is why “wrapper” is a lazy critique and why labs will keep pushing into applications as the API commoditizes (“you just change one line of code”).
Funding-market mania, from the inside: “it just seems way too easy to raise money,” and Decagon has been preempted “almost immediately” after every round — “that alone can’t be right.” His screens: test investors during the window when they want in but haven’t invested (their helpfulness then is a proxy for post-investment behavior), index on “raw intellectual throughput,” and beware expert-call noise — “a lot of people just lie on customer calls,” including claimed Decagon users the company has never heard of.
Contrarian side-calls: bullish Google among mega-caps because consumer reach is where new data comes from (Anthropic’s consumer weakness is a long-term risk); his five-slot private portfolio is Cognition, Cursor, Etched, Pika, and Chai. The forward-deployed-engineer craze is overdone — at likely Palantir, the FDE model serves “$10–25 million” deals and needs roughly $1M in customer size or revenue; 996 “is not that healthy” and works in China in part because employers have leverage.
🔗 Original source & video: The Future of AI Agents | Jesse Zhang Interview
Opendoor CEO, Kaz Nejatian: OpenAI and Oracle, How Can Either Afford to Do This
- 🗓️ Date:
2025-09-19| 🎙️ Show:20VC
Opendoor is being refounded as a software-and-services platform that uses AI to price homes fairly, then monetizes a longer customer relationship rather than relying on the purchase spread. The broader panel highlights a sharper AI market: applications are expanding TAM rapidly, but weak switching costs and two-week competition make liquidity, strategic acquisitions, and IPO windows important catalysts against paper valuations.
View Dialogue Notes & Key Takeaways
Kaz is betting that Opendoor becomes a software-and-services platform, not merely a leveraged home trader. The company should offer fair prices, use software and AI to value properties, and monetize a long relationship through attached services rather than extract its margin at purchase. “Businesses should not exist to make money. Businesses should make money to deliver on a mission.”
The product ambition extends from asset-light transactions to homes that can effectively be returned or guaranteed for life. Kaz argues that Opendoor can eventually find sellers their next home, stand behind properties it sells, and combine asset-heavy and asset-light models. The investor case is correspondingly extreme: “The bull case for Opendoor is obscene,” though the hosts repeatedly stress that houses remain heterogeneous, illiquid assets.
Kaz is aligning compensation and governance for a public-company “refounding,” with execution risk deliberately concentrated. He says he would take less than $1 but is not permitted to; he owns no RSUs and has performance compensation tied to stock-price cliffs. He also says he would not have joined without Keith and Eric providing board-level air cover. After leaving “a few hundred million” at Shopify, his operating test is “hard, valuable, fun,” and his recruiting call is for the “most aggressive, innovative public tech company.”
Oracle’s more-than-$300 billion backlog gave public investors a liquid proxy for OpenAI—and they priced it almost as though every dollar were certain. Oracle rose roughly 36–38% and briefly touched $1 trillion, even though the implied OpenAI commitment is about $60 billion annually over five years against roughly $12 billion of current revenue. The panel doubts the full $300 billion will be paid and questions whether GPU hosting can approach Oracle’s existing 41% operating margins.
The panel’s cycle call is that AI euphoria can be productive while still making liquidity more valuable than another paper mark-up. Private investors cannot sell incrementally as public shareholders can, so “valuation” is not “liquidity”; rejecting a $4 billion offer for an imagined $8 billion, $12 billion, or $24 billion outcome can become “winning the lottery and then losing your ticket.” One panelist’s stated default is to tell every founder with a massive offer to sell—and require the founder’s own conviction to overturn that advice.
OpenAI appears to be escaping Microsoft’s bear hug and converting a strategic dependency into an arm’s-length relationship. The likely end state has Microsoft as a major shareholder, cloud vendor, and model customer, but without exclusivity, while Microsoft simultaneously adopts Anthropic products. A potential 20–35% OpenAI stake could produce a roughly 10x return on Microsoft’s $13 billion investment, yet still barely move a company worth around $3 trillion.
AI applications are creating markets far larger than their pre-AI TAMs, but their revenue durability is radically less proven than SaaS. Higgsfield reportedly raised $50 million and reached $50 million of ARR, while Gamma reached $60 million this year; the mechanism is a 10x or 100x expansion in who can create videos, slides, or software. Yet competition now reacts in “two weeks,” switching costs are tiny, and even Anthropic might lose 30–40% of Claude Code revenue if GPT-5 Codex becomes comparably capable.
Distribution is giving some incumbents a route back into the AI race, while making strong number-two startups unusually valuable acquisition targets. Wix reportedly bought solo-founded Base44 for $80 million, added identity, safety, and its funnel, and could take it toward $50 million of ARR; Workday bought Sana Labs for $1.1 billion at roughly $50 million of ARR. With IPO activity also returning, the trade is increasingly about recognizing when a strategic buyer or open window offers real cash rather than theoretical compounding.
🔗 Original source & video: Opendoor CEO, Kaz Nejatian: OpenAI and Oracle, How Can Either Afford to Do This
No Priors Ep. 97 | With Decagon CEO and Co-Founder Jesse Zhang
- 🗓️ Date:
2025-09-18| 🎙️ Show:No Priors
Decagon’s customer-support wedge offers unusually legible agent economics, with Bilt Rewards stopping team expansion within roughly one month and later reporting around 65 agents of headcount saved. Its differentiation lies in eval-driven orchestration, business-logic customization, observability, and control rather than exclusive model access, while voice and computer-use expansion depend on latency, accuracy, and measurable ROI.
View Dialogue Notes & Key Takeaways
Customer support is Decagon’s “golden use case” for AI agents because automation and customer outcomes are both directly measurable. Jesse Zhang says buyers track the fraction of conversations handled, CSAT or NPS, and—especially in regulated industries—accuracy. The result can be a “personal concierge” available in any language, 24/7, with potential gains in retention and conversion alongside labor savings.
Bilt Rewards stopped scaling its support team within roughly one month of adopting Decagon and has since recorded around 65 agents of headcount saved. Its support volume had been growing with its rapidly expanding user base; automation let Bilt restructure the operation while making responses faster. Zhang calls the case’s ROI “very easy.”
The key differentiation sits above foundation models, not in exclusive access to them. Decagon combines multiple models through eval-driven orchestration, molds them around each customer’s business logic, and exposes the data, steps, and knowledge gaps behind responses. “You really don’t want this to feel like a black box.”
For customer-service agents, instruction following matters more than the coding and math gains dominating model discourse. Zhang welcomes advances from o1 and Sonnet, but says the decisive capability is whether a model can take a support SOP or workflow and “follow them to a T.” That leaves meaningful application-layer work even as core models improve.
The near-term agent winners must support gradual deployment and produce an easily quantified return before reaching perfection. Customer support and coding pass that test; security may not because missing one subtle event can be unacceptable, while text-to-SQL often remains a supervised copilot with murky pricing power. Zhang is therefore “more bearish on a lot of these AI agent use cases in the near term.”
Voice, screen context, and agent supervision are key future areas for Decagon. Voice-to-voice models reduce latency, but production workflows may still require data retrieval and multiple calls; Zhang says computer use is probably not production-ready yet. Longer term, he expects more humans to supervise and edit infinitely scalable agents, making observability and control central product priorities.
🔗 Original source & video: No Priors Ep. 97 | With Decagon CEO and Co-Founder Jesse Zhang
How AI Agents Are Transforming Customer Support, with Decagon’s Jesse Zhang
- 🗓️ Date:
2025-01-16| 🎙️ Show:No Priors
Customer support is emerging as the near-term “golden use case” for AI agents because deployment can start small while savings, satisfaction, NPS, and accuracy remain measurable. At Bilt Rewards, Decagon stopped scaling support within roughly one month and later quantified around 65 agents’ worth of headcount saved, while orchestration above shared models remains the differentiation. Voice expands the workflow opportunity but introduces a persistent latency-versus-computation trade-off worth monitoring.
View Dialogue Notes & Key Takeaways
Customer support is emerging as the near-term “golden use case” for AI agents because deployment can start small while savings and service quality remain directly measurable. Decagon tracks the share of conversations automated alongside customer satisfaction, NPS, and accuracy, effectively offering every customer a “personal concierge in their pocket” across languages and 24/7.
Bilt Rewards provides Decagon’s clearest proof of operating leverage: within roughly one month, it stopped scaling its support team as AI took over much of the automation. Almost a year later, Bilt had restructured the function and quantified around 65 agents’ worth of headcount saved, while customers posted that its support “doesn’t feel like any sort of AI or chatbot system we’ve ever used before.”
Decagon’s claimed differentiation is software above broadly available models, not exclusive access to GPT-4o, GPT-4, or Claude Sonnet. Its orchestration layer evaluates and combines models around customer-specific business logic; its product layer exposes data, decision steps, knowledge gaps, and conversation categories. Zhang’s framing: “Most of the alpha, or most of the special stuff that you build, is on top of models.”
For customer-service agents, instruction following matters more than the coding and mathematical reasoning gains emphasized around o1 and Sonnet. A support agent must execute an SOP “to a T,” particularly in regulated environments; better reasoning helps, but dependable adherence to workflows is the more consequential model-development vector for Decagon.
Voice expands the addressable workflow but introduces a persistent latency-versus-computation trade-off. Voice-to-voice models respond quickly, while production calls may require data retrieval and multiple model calls; speech-to-text-to-speech adds delay but permits that work. Decagon’s pragmatic bridge includes conversational cover such as, “Hey, give me a second. I’m looking up your data.”
The organizational end state is not simply fewer humans but more people “supervising and editing agents.” Zhang expects leaders to monitor, correct, and hard-code behavior across infinitely scalable systems, while screen context and computer use could let agents navigate products—not merely answer questions—though Anthropic’s demonstrated computer use was, in his view, “not production-ready yet.”
Zhang is “more bearish” on most near-term agent categories than the sector’s demos imply. Successful markets need both incremental rollout before near-perfection and easily quantified ROI: security agents struggle with nondeterminism where any small anomaly matters, while text-to-SQL often remains a monitored copilot whose economic value is difficult to price. Better models may unlock those categories later.
🔗 Original source & video: How AI Agents Are Transforming Customer Support, with Decagon’s Jesse Zhang