Pioneers Insight Method Research Author
Back to Pioneers
Div Garg
Entrepreneurs 1 Curated Dialogues

Div Garg

Key Views & Dialogues

The Quest for Autonomous Web Agents with Div Garg, Cofounder and CEO of MultiOn

  • 🗓️ Date2025-05-17 | 🎙️ Show:The Cognitive Revolution

MultiOn argues that agents remain limited by grounding, long-horizon reasoning and environment feedback rather than conversational fluency, with fine-tuning plus online reinforcement learning targeting reliability beyond 90–95%. Short single-site workflows already provide a commercialization wedge, but margins depend on cents-per-step inference, context compression and model routing as the roadmap expands toward recurring jobs, parallel agents and higher safety demands.

View Dialogue Notes & Key Takeaways
  • MultiOn’s core claim is that agents are stalled by reasoning and grounding, not conversational fluency. Div G says GPT-4 can produce “seemingly good content” while “the actual deep work is not there”; websites are dynamic environments models were not trained to represent, so small mistakes compound across long trajectories. The discussion points toward execution data, environment-specific feedback and verification as important layers beyond chat.

  • The proposed path from 90–95% reliability toward nearly 100% is a fine-tuned base plus online reinforcement learning, not RL from scratch. Div G argues sparse-reward RL historically failed because agents began near zero; pretrained models now supply broad knowledge, while supervised tuning can create a roughly 90% starting point from which agents “explore and exploit the environment.” MultiOn was exploring DPO, imitation learning and live-internet experiments on reversible tasks, with explicit care not to “take down the internet.”

  • Commercialization can start before general agency because short, single-site workflows already form a usable wedge. Amazon cart assembly, DoorDash and Instacart orders, NDAs and meeting invitations were cited as current strengths; cross-site state transfer, recurring jobs and long tasks remain harder. At the time, Div G expected MultiOn could soon choose one task and execute it with “a crazy amount of accuracy,” pursuing adoption and research in parallel.

  • MultiOn optimizes for product economics alongside task completion: at least 10× human speed, compact context and cents-per-step inference. A simple task might require fewer than 20 steps and research around 100; Div G put the upper bound near $0.10 per step, with efficient hosted models nearer $0.02–$0.03 and a 100-step job potentially below $2–$3. Its average inference input was “not more than 5,000 tokens,” making compression, caching and model routing central margin levers.

  • The stack is deliberately model-agnostic because GPT-4 remained best for planning but too expensive to anchor a mass-market consumer product. MultiOn mixed fine-tuned open-source models with GPT-4, benchmarked alternatives “plug and play,” and designed its efficiency work to survive a GPT-5 capability jump. Div G attributed GPT-4’s edge to “the quality of the ingredients and the chefs”: experienced researchers, private data and a large human-labeling pipeline.

  • The roadmap moves from one-off actions to long jobs, composition, recurrence and eventually a hidden hierarchy of parallel agents. A user may see one chat surface while an internal scheduler distributes work across subagents, borrowing operating-system concepts such as processes, priorities and failure handling; a Voyager-style skills library would cache reusable procedures. Mobile access and authentication without storing the user’s password aim to deliver “what Siri could never be,” while the API is positioned as a natural-language abstraction over Playwright-like automation.

  • The labor thesis is augmentation first and substitution later, but trust may become the binding constraint on adoption. Div G expects agents to remove “digital chores” and “shitty jobs” as computers displaced typewriter work while creating new roles in teaching, programming and coordinating agents; near-term value comes from taking users from “zero to one” access to assistance. Nathan warned that malicious calling and capability jumps could create systemic risk, while Div G recommended moderation, execution-time verification, prompt-injection defenses and rapid behavior patches.

  • 🔗 Original source & video: The Quest for Autonomous Web Agents with Div Garg, Cofounder and CEO of MultiOn

Listen to full conversation →