The Future of AI Agents | Jesse Zhang Interview
The Future of AI Agents | Jesse Zhang Interview
Summary
- Decagon is an AI customer-service agent, and Jesse Zhang’s explanation of why that use case has the most enterprise gen-AI traction is the tradeable insight: ROI is pre-quantified (existing chatbots/IVRs “resolve” 15–20% of volume; AI could take it to 50–80%, so “I’m going to take the total cost and chop off 60% of it”), and escalation-to-human is already built into call-center infrastructure, de-risking go-live. Enterprises start with 5% of traffic and roll to everything “within weeks” once resolution, CSAT, and accuracy check out.
- Zhang’s labor-spectrum framework: AI agents exist to replace human labor, so map use cases by what that labor costs — AI eats the spectrum from both ends. Engineers (the most expensive) are generally augmented rather than simply fired, because “there’s infinite engineering work to do”; outsourced tier-1/2 support (the lower-paid end) can genuinely be replaced, with BPOs unbothered because turnover is high and workers shift to tasks AI can’t do yet, like data labeling.
- The founding method is replicable: ask “classic sales-qualification questions but in founder form” — exactly how much would you pay, whose boss approves, how would you present ROI to leadership. Most ideas die at “$100 subscription per month” from a big company; Decagon’s conviction came when tallied willingness-to-pay was “an order of magnitude more than anything else” — even as “very smart older founders” warned the idea was too obvious and incumbents would win.
- Voice is the frontier and it isn’t solved: voice-to-voice models hallucinate “probably like 8x higher” than text, so enterprises route voice→text→voice with checks. The prize is large — Fortune 100 support is 90–95% voice — and the bar is high because spoken language is a ~150,000-year-old UI versus ~60 years of keyboards.
- The moat claim: an LLM now reads every one of a million monthly conversations (versus a 20-person sampling team), flags the 2% going wrong, and drafts fixes — so after a year “has your agent just continuously gotten better… to the point where it’s very difficult for another agent to come in and perform at the same level?” End state: the agent becomes the brand’s front end — a “digital concierge” replacing the app and website entirely.
- On unit economics: at the application layer “your margin doesn’t really matter” right now — costs fall exponentially, so optimize for market share and mindshare — though enterprise deals shouldn’t hemorrhage cash and Decagon’s one rule is never negative margins. His croissant analogy: the last step that actually solves the problem captures the most margin, which is why “wrapper” is a lazy critique and why labs will keep pushing into applications as the API commoditizes (“you just change one line of code”).
- Funding-market mania, from the inside: “it just seems way too easy to raise money,” and Decagon has been preempted “almost immediately” after every round — “that alone can’t be right.” His screens: test investors during the window when they want in but haven’t invested (their helpfulness then is a proxy for post-investment behavior), index on “raw intellectual throughput,” and beware expert-call noise — “a lot of people just lie on customer calls,” including claimed Decagon users the company has never heard of.
- Contrarian side-calls: bullish Google among mega-caps because consumer reach is where new data comes from (Anthropic’s consumer weakness is a long-term risk); his five-slot private portfolio is Cognition, Cursor, Etched, Pika, and Chai. The forward-deployed-engineer craze is overdone — at likely Palantir, the FDE model serves “$10–25 million” deals and needs roughly $1M in customer size or revenue; 996 “is not that healthy” and works in China in part because employers have leverage.
Deep dive
1. A culture built for war — “no enemy that can’t be defeated”
- The phrase on Decagon’s wall: “There’s no challenge that can’t be overcome and there’s no enemy that can’t be defeated” — Zhang’s dad told him that Huawei’s famously hard culture has a Chinese version reportedly hanging in big red letters. Patrick’s observation frames the era: words like “violence” and “aggression” would have been a big problem for a founder three years ago; now talent actively rallies to cultures that talk about defeating enemies.
- Zhang’s logic for the intensity: “any space that’s worth going after… it’s going to be competitive” — Databricks vs. Snowflake, Ramp vs. Brex — and since starting a company and raising money are both easy now, culture is a durable advantage: “pretty hard to replicate” and long-lasting.
- Rallying mechanics he’s learned: always keep “a flag pole that’s within sight,” and use healthy competition — “when people feel like they’re in a battle and there’s clear enemies” — to bind the team. Last year’s revenue milestone prize was Decagon jackets, likely Arc’teryx: trivial cost against payroll, but “it just feels like everyone’s working together towards this common goal.”
- Recruiting is the same sport: “for anyone you want to hire, you need the whole team to swarm around them” — getting to know families and partners, designing roles around what people want. The application layer isn’t as extreme as the Meta/OpenAI researcher wars, but the Harvard/MIT/Stanford pool in SF is finite — one reason Decagon just opened a New York office.
2. Index the math-contest kids
- Patrick’s setup: enough competition-math alumni (Scott Wu of Cognition, Zhang himself) are now winning in startups that “if you could somehow index that group of people, you’d have fantastic performance.” Zhang’s explanation: contests train competitiveness plus problem-solving, which can be applied to vague problems — “the problem could be like how do I build a successful company” — and results are objective, so “there’s just constant motivation to improve.”
- His talent-arbitrage thesis: this cohort historically defaulted to trading or academia because the background correlates with being “a little bit more risk averse… get your good grades and follow a track.” Diverting that talent into company building is, in his words, one of his big theses on untapped potential.
- The kindest thing anyone did for him is the origin story: from ages 5–13 his parents ran “an extreme level of discipline” — piano three to four hours a day, then all-in on math; no TV, no video games, and they “didn’t really take vacations” growing up. They “pounded a lazy, wanted-to-play-around kid into someone that was just very, very driven” — and then, unusually, “completely laid off” in high school, with no opinions on his career since. “A lot of the things I have in life right now are from that.”
3. Sales-qualification questions in founder form
- The scar tissue came from a company likely called Lowkey (heard as “Loki”) — video-clip capture software for gamers, exited at “a very fortunate time,” 2021. Three months of hard work on something that “obviously had no market,” then both co-founders burned out and quit, leaving him alone: “way tougher than anything we’re doing right now… you don’t really know if there’s any future or not.” Today’s version of hard is just volume — four to five hours of sleep this week in New York.
- Second time, he and co-founder likely Ashwin (heard as “Ashman,” ex-founder himself) systematized ideation: open discovery with senior operators — COOs get multiple use cases, VPs get one — form product hypotheses live, then force the money question. “If we built this for you, exactly how much would you pay? Would your boss need to approve it, or your boss’s boss? How would you present ROI to leadership?” Because you’re a founder, “it just feels a lot less salesy” — the same questions from a salesperson would grate.
- The payment question is the forcing function: vague enthusiasm (“oh yeah, that’d be great” — people feel they owe you something on a call) collapses into order-of-magnitude reality, usually “20k a year” for replacing one of five people. Most exercises end in relief: “glad I didn’t pursue this further… there’s not that many good ideas out there, honestly.”
- The breakthrough came sideways: ops leaders (Matt McInness of Ripling among them, plus a wearable-ring customer likely Oura) would size the small idea, then volunteer “by the way, we have a 500-person support organization.” Everyone — “including very smart older founders” — said the use case was too obvious, incumbents would tack it on. The tally said otherwise: summed willingness-to-pay was “an order of magnitude more than anything else,” with single answers reaching low-to-mid six figures from “a random two-person team.”
4. Why customer service is enterprise AI’s beachhead
- Property one, obvious only in hindsight: ROI arrives pre-quantified. Enterprises already track conversation volume and know their chatbot/IVR “resolves” 15–20% of it. “If you’re able to take that to 50, 60, 70, 80 — that’s huge ROI… I’m going to take the total cost and chop off 60% of it.”
- Property two, underrated: easy to go live, which is where many gen-AI projects struggle because models are non-deterministic and nobody wants leadership asking “why’d you do this?” Customer service has an escalation path natively built into the existing call-center infrastructure, so enterprises release to 5% of one surface knowing anything that goes wrong just escalates to a human.
- The spread once trust lands is fast: “even for the large enterprises, within weeks” — a 500-person org handles roughly mid-to-high six figures of conversations a year, and once resolution rate, customer satisfaction, and human-reviewed accuracy check out, “there’s really no reason why you shouldn’t roll it out.”
5. AI eats the labor spectrum from both ends
- The framework worth keeping: agents exist “to essentially replace human labor,” so map use cases by what the labor costs. Coding sits at the expensive end — engineers have the sophistication to leverage AI, so Zhang says they are generally augmented rather than simply fired (“I don’t know any company that’s like, I would like to just let go of a bunch of my engineers… there’s infinite engineering work to do”). Support sits at the lower-paid, already-outsourced end, where replacement can be real.
- Even the BPOs “are not really that concerned”: turnover is already high, so enterprises let headcount “just naturally decrease” while workers move to “the next level of task that the AI can’t do yet — maybe that’s data labeling.”
- For CEOs who want AI but don’t know where: every initiative is now top-down, board-level mandate, so buy-in must come from the C-suite, and the use case must clear the half-sentence test — “where can we either save a bunch of money or make a lot of new revenue… if they can’t point to ‘I saved $10 million,’ it’s not going to be prioritized.”
- Patrick’s pushback on coding ROI — an unnamed study finding that self-reported productivity contradicts measured productivity — draws an honest shrug: “I actually don’t have an opinion… but that doesn’t really matter, right? If their entire engineering team is like, hey, we love this, this is making us 50% more productive,” the CEO reports it to the board and the investment is justified. Perception is the buying signal.
6. The long pole: agreeing on what “good” looks like
- Zhang’s biggest deployment learning: the bottleneck is “aligning on what does good look like” — tone, brand guidelines, and answers no single person knows, because in a large org “no one is the person where, like, I know how to answer all these questions.” You design a process to extract answers from the CX leaders who each own a piece, align them on the eval, then run a simulation suite — 10,000 tests, each running constantly, five times — until “you’re just building, building, building” against a quantifiable score.
- Patrick’s frame, which Zhang accepts: you’ve created a captive reinforcement-learning process inside the organization — not RL in the model-training sense, but compiling evals and guardrails that make the agent improve.
- Best failure story, as told: a ticketing-platform customer couldn’t find their ticket, so they announced they’d “show up to the event and find eight homeless people from the city and bring them with me” — and the agent replied, “Oh my god, that’s so awesome that you’re thinking about doing something nice for the community.” The design problem is a spectrum: regulated flows need full rigor (“these three steps always have to be followed in this order”), basic account questions want free-form humanity.
- The upside surprise is trust recovery, measured by how often users jam “agent agent agent” to escape. In the likely Oura case study, one in three customers used to say nothing and keep jamming “agent” until reaching a representative; after making the opening feel different, it’s one in 20.
7. Voice is the frontier — and voice-to-voice hallucinates 8x more
- The bar argument: for ~150,000 years “the UI for every human is language… only in the last 60 years did we have keyboards.” Brains are evolved to detect wrongness, so the uncanny valley is wide — likely ChatGPT Voice and Sesame “are starting to feel very impressive, but if you talk to it long enough, you can definitely tell it’s not a human.”
- The technical trade, spelled out: voice-to-voice captures cadence, tone, “how upset you are,” and cuts latency — “the prevailing view is that if you really want to make it indistinguishable from a human, you have to do voice-to-voice.” But audio means more tokens per sentence, and “the more tokens you have, the easier it is for something to go wrong” — hallucination is “probably like 8x higher.” So enterprises today route voice→text→voice with checks, and the craft includes human tricks like saying “give me a sec to look that up” when the API genuinely takes ten seconds.
- The stakes: Decagon’s volume is roughly balanced chat/voice, but Fortune 100 incumbents are “90, 95% voice.” Conversation difficulty tiers: static Q&A, then real-time-data reasoning (why did I get 2x points instead of 5x? — because you booked through a travel agency, not the airline), then taking action — the lost-credit-card flow with address confirmation, card locking, and fraud checks stitched together. “That’s really what makes it agentic.”
8. The moat compounds; the end state is the front end
- Conversation data was always valuable and always wasted: the old method was 20 full-timers sampling a million monthly conversations against a rubric. Now “you can literally have a language model that reads every conversation,” flags the 2% going badly for reasons leadership may not see, and drafts the fix — the agent improves automatically over time.
- His definition of agentic moats: “if you’ve been working with a client for a year, has your agent just continuously gotten better by learning from the data… to the point where it’s just very difficult for another agent to come in and perform at the same level?” — distinct, he notes, from training on the data.
- The end state: the agent becomes “a new UI for the product” — a unified, authenticated “digital concierge” with memory and context that books your flight and upgrades your seat, so users “don’t even have to touch your mobile app… never go on your website.” Patrick’s website analogy lands: “it’s just like a front end… but instead of visual UI, it’s a conversational UI.” Near-term, separate budgets mean separate agents; long-term they unify.
- What makes deployments great: “the number one thing is just APIs for the AI to use — if you have those, you already know that within the first month it’s already going to be a great experience.” Then documentation, brand guidelines (which already exist for training human agents), and SOPs converted into what Decagon calls AOPs — “agent operating procedures… just SOPs, but for AI.”
9. Margins don’t matter yet — and “wrapper” misses where value sits
- What the world overestimates: demo-to-enterprise speed, because of non-determinism and what he calls the likely Waymo effect — new technology gets judged by hunting for mistakes rather than holistically on success rate, where AI’s success rate can exceed imperfect humans. What it underestimates: “things are improving exponentially, and no one is good at conceptualizing what exponential means.”
- Hence the margin heresy: at the application layer, “your margin doesn’t really matter… what really matters is you need to get market share and mindshare.” Costs fall exponentially, and companies could “spend a month and massively improve their margins” but shouldn’t — quality and growth first, optimization later. The enterprise caveat: deals are long-term and expectations shift, so “you generally don’t want to be hemorrhaging cash” — Decagon’s only stated principle is no negative margins.
- Why the app layer keeps the economics anyway — the croissant analogy: in any supply chain, “that last step where you’re actually solving someone’s problem, that’s where you generally can capture the most margin.” Customers “don’t care at all what Decagon’s costs are — what we care about is the business ROI.”
- On “ChatGPT wrapper” as a slur: yes, thin apps die, and copywriting is vulnerable to people just asking ChatGPT, but “an agent is not just a model — you have to design it, put in guardrails, teach it new things,” plus a mass of non-AI software: observability, alerting, QA and unit tests for conversations. The real threat runs the other way — labs will push into applications because API revenue is commoditized: “moving from AWS to GCP is very hard, but moving from an OpenAI model to an Anthropic model — you just change one line of code.”
10. Model strategy, investor mania, and the FDE meme
- The fine-tuning reversal: two years ago the writing largely said “fine-tuning doesn’t work” and models changed too fast to bother. Now mature applications know exactly where each model sits, so small fine-tuned open-source models handle low-intelligence steps like routing (“that model doesn’t have to be good at math or coding”), improving performance and latency — while frontier models handle the peak-intelligence work. Among mega-caps he’s “very bullish on Google, actually”: consumer reach is “where all the new data is going to come from,” and Anthropic’s consumer weakness is a long-term risk. His five-slot private portfolio is Cognition, Cursor, Etched, Pika, and his friend Josh’s Chai (healthcare foundation models) — or likely Physical Intelligence.
- On fundraising: “there’s maybe a little bit too much excitement… it just seems way too easy to raise money.” Decagon has been preempted almost immediately after every round — “and that alone can’t be right,” since prior valuation shouldn’t drive a first-principles underwrite. “It feels like there’s a little bit of mania.”
- His founder advice: exploit the window when investors want in but haven’t invested — “that’s when they’re most willing to be helpful,” and their helpfulness then is a proxy for how helpful they’ll be afterwards. Index on “raw intellectual throughput” over pattern-matching (“you generally don’t want investors that are just like, I’ve seen this so many reps”). And discount expert calls: his customers “have made probably so much money on expert calls,” but “a lot of people just lie” — including claimed Decagon users “we literally never heard of.”
- Two memes deflated: 996 — “I don’t actually think 996 is that healthy… it happens in China and one of the reasons is that no one has jobs over there, so employers have leverage.” And forward-deployed engineers — at likely Palantir an FDE means “a $10–25 million deal” with near-dedicated staffing; startups are conflating that with ordinary hands-on hustle. His threshold to justify the model is probably $1M in customer size or revenue — below that, staffing a person per 50k client “prevents you from scaling. No one can hire good people that fast.”