Pioneers Insight Method Research Author
Why Every Agent Needs a Box — Aaron Levie, Box
Back to Episodes

Why Every Agent Needs a Box — Aaron Levie, Box

Summary

  • Box’s core thesis is that enterprises will have “10x or 100x” more agents than people, turning dormant corporate files into continuously useful infrastructure. Contracts, research, roadmaps, and customer material become inputs for onboarding, sales, and autonomous work rather than documents humans occasionally reopen. The pitch writes itself: “Every agent needs a Box.”

  • Agent identity and authorization, not raw model intelligence, may determine whether autonomous agents can enter regulated enterprises safely. Today’s “easy mode” makes the agent identical to its human operator; independent agents create harder questions about privacy, liability, oversight, and multi-party access. Levie expects “spectacularly crazy security incidents” unless permissions and governance become agent-native.

  • AI coding’s rapid adoption is a misleading benchmark for the rest of knowledge work because software development enjoys unusually favorable conditions. Code is largely text-in/text-out, engineers commonly access broad repositories, models are heavily trained on code, and the labs’ own developers supply continuous feedback. Bankers, lawyers, and other workers instead face fragmented permissions, undocumented context, mixed media, and information trapped in conversations.

  • Enterprises will have to redesign work around agents rather than wait for agents to assimilate into existing processes. “The agent didn’t really adapt to how we work. We basically adapted to how the agent works.” swyx challenged the consultant-friendly premise and cited OpenAI hiring FDEs and Anthropic embedding at Goldman Sachs as evidence that there is no effortless “come as you are” path. Levie agreed that reaching a well-organized data environment will be difficult and said the opposite extreme—an agent inferring everything from a totally messy environment—is technically impossible.

  • Context engineering is fundamentally a retrieval problem: perhaps 50 million pages of accessible information must be reduced to roughly 60,000 dependable tokens. Larger windows do not remove the need for search, ranking, access control, and judgment about when to stop looking. Better models can detect contradictory or stale documents, but “it still doesn’t work if you just have a total wasteland of data.”

  • Knowledge-work reliability requires private evals because plausible slop can create professional and legal exposure invisible in ordinary software output. Box’s held-out industry benchmark reportedly showed roughly a 15-point jump between model versions in one comparison, while internal tests catch regressions in both models and agent harnesses. Levie expects every enterprise eventually to maintain evals for workflows such as RFP creation, sales collateral, and invoice processing.

  • Box is positioning its governed file system as both an agent data layer and a sandboxed workspace, while organizing a roughly 3,000-person company around an existential agent transition. A core group of a few dozen is supported by search, metadata, infrastructure, security, and compliance teams. Beyond Box, swyx and Alessio argue that software output could increase by 10 to 100 times, making technical workers, deployment, and DevRel more important rather than less.

Deep dive

1. Dormant enterprise files become active agent capital

  • Levie’s starting point is that corporate files already contain contracts, research, marketing material, memos, and roadmaps, but humans use most of them only during an active engagement. Agents turn that archive into “this ongoing source of answers to new questions” and raw material for newly generated work.

  • The concrete applications span the enterprise: an incoming employee can reconstruct a project, a seller can identify what to offer a customer, and a product team can recover the information behind its next feature. The data’s value rises because agents can continuously retrieve and transform it.

  • Some agents will act directly as their human, inheriting the same access. Others will resemble autonomous colleagues with their own machines, tools, and sandboxed environments—closer to the OpenClaw pattern discussed by the hosts. Levie’s shorthand: “Every agent needs a Box.”

2. More agents than employees creates a new infrastructure market

  • Whether the multiplier is “10x or 100x,” Levie regards an order-of-magnitude increase in agents over people as inevitable. That creates demand for governance, permissions, access controls, workflow coordination, and retrieval across multiple enterprise systems.

  • The threat case is concrete: a prompt-injected agent could navigate through a CRM and extract information its user should never see. Levie expects “spectacularly crazy security incidents” because autonomous software combines broad access with the ability to act.

  • Regulation remains unsettled. In financial services, Levie asks whether an agent inherits the same requirements as a human worker or whether responsibility rests entirely with the person who created or directed it; either way, a data-governance layer is still required.

  • When swyx rounded Box’s Fortune 500 customer penetration to 70%-80%, Levie answered “67%” and said the company was projecting to the end of the year. Those relationships expose Box directly to the permission and compliance constraints that decide whether agents can graduate from pilots.

3. An agent cannot simply be treated as another employee account

  • Levie calls today’s pattern “easy mode”: in Claude Code, Cursor, or Codex, “the agent just is you.” It authenticates through the user and can generally do whatever that user can do, avoiding the need for an independent identity or responsibility model.

  • Autonomous agents are different. Their creators will probably retain liability and require oversight, while the agents themselves have neither a human claim to privacy nor legal responsibility. Creating ordinary user accounts would reproduce controls designed for people while obscuring who must inspect and answer for the work.

  • Collaboration makes the boundary harder. If one person creates an agent that later works privately with another employee, the creator needs oversight of the agent but should not automatically see the collaborator’s confidential material. The familiar human Venn diagram of private and shared work no longer maps cleanly.

  • swyx suggested conventional RBAC may be “dead” at this level of granularity; Levie’s more measured answer was that Box’s waterfall permissions create new problems. Agents need selected data, their own workspaces, partial access, and accountable supervision—the “boring problems for 98% of people” that determine whether autonomy leaks data.

4. AI coding is the exceptional case, not the enterprise baseline

  • Coding combines unusually favorable properties: broad repository access, a text-in/text-out medium, extensive model training data, technical users willing to install new tools, and highly networked communities sharing practices. The AI labs also use coding agents daily, producing an unusually tight product-feedback loop.

  • Levie contrasts that with a banker who sees only a fragment of the necessary information, must find whoever controls a deal-room folder, and may need context from another organization. Requirements also arrive through Zoom and in-person conversations that were never captured as authoritative text.

  • Documentation and specifications exist imperfectly in software, but “those things don’t exist for like 80% of work that happens in the enterprise.” Other knowledge domains therefore face six or seven headwinds absent from coding: fragmented data, mixed formats, access controls, tacit knowledge, weaker tooling, and users who need training.

  • The result is a “multi-year march” rather than instant replication of coding-agent adoption. Coding reached escape velocity because its environment was already unusually legible to models; the rest of the economy must first make its workflows and context similarly operable.

5. Companies will adapt their workflows to agents

  • Levie calls coding “the most changed workflow in maybe the history of time” over a two-year interval: developers increasingly describe tasks to agents rather than write every line or even review everything. The decisive shift was organizational—“we basically adapted to how the agent works.”

  • He expects the rest of the economy to follow by redesigning processes, prompts, access, documentation, and review around agent execution. The promised agent that simply “drops in” and automates an existing life has not appeared; early teams nevertheless gain compounding advantages while competitors spend years re-engineering.

  • swyx’s pushback was that this sounds like a consultant’s dream and leaves room for a rival promising, “Come as you are, and we’ll meet you where you are.” He then cited OpenAI hiring FDEs and Anthropic embedding at Goldman Sachs as evidence that even the labs need hands-on workflow transformation. Levie agreed that reaching the “beautiful garden” will be difficult.

  • Levie does not expect a perfectly manicured data garden, but he says the opposite extreme is technically impossible. If context is irretrievably messy, no model can infer absent facts; competitive pressure will force better documentation because the cost of wrong retrieval and lost productivity becomes too high.

6. Better models improve judgment but cannot rescue a data wasteland

  • Box’s internal agents produced bogus answers nine months earlier, sometimes returning five documents that merely “smelled like the right thing.” Levie describes the system as being put “on the clock” to answer despite uncertainty; Alessio summarized the outcome as, “Doesn’t work.”

  • He credits progress from Opus 4.6, Gemini 3.1 Pro, and whatever the latest GPT-5.3 becomes. Where models six months earlier effectively threw darts, newer Opus 4.5 and 4.6 variants can notice contradictory signals, reconsider candidate documents, and rerank results.

  • Box’s agent fans out searches, gathers candidate files, and ranks them before answering. Yet model intelligence has a ceiling: “If a really, really smart human could not do that task in five or 10 minutes” for retrieval, Levie does not expect an agent to overcome the missing or incoherent source material.

7. Context engineering reduces millions of pages to a tiny working set

  • Levie grants that infinite context might become economical around 2035, but it is not a present architecture. Even if a model advertises 200,000 tokens, he estimates perhaps 60,000 remain dependable before substantial degradation—far too little for an enterprise corpus.

  • His scale comparison is the load-bearing point: 10 million documents at an assumed five pages each produce 50 million pages, while the model can reliably inspect only a few hundred pages’ worth of tokens. Search systems, databases, permissions, and ranking must bridge that gap.

  • Box tests this with a request for the addresses of 10 offices when no canonical file contains all 10. Lower-tier models often find six, report four missing, and stop; exhaustive search is expensive, while one requested office might not exist at all.

  • The desired capability is judgment: try alternative queries, check the evidence, and eventually decide that further searching will not resolve the task. “When should it give up?” is a core knowledge-work problem because the answer may be missing rather than waiting somewhere in a repository.

8. Agents need selective forgetting and stricter error standards

  • Alessio observes that humans naturally prune failed approaches, while agents can repeat a mistake merely because it remains prominent in their trace—even when the trace says it failed. The proposed pattern is to remove the distracting attempt while preserving a compact warning not to repeat it; swyx describes this as cutting the mistake without losing the lesson.

  • Software slop can remain invisible behind a working interface. Knowledge-work slop is exposed directly: if a contract is generated 20 times and each version differs by 3%, those variations create organizational risk rather than harmless implementation ugliness.

  • The contrast is professional liability: a software engineer may cause an outage, roll back, and attend a review, but lawyers can be disbarred and medical errors harm patients. Knowledge-work agents therefore need narrower constraints, review responsibilities, and management standards that coding agents did not initially confront.

  • The hosts frame 2025 as the year coding agents rose and 2026 as the knowledge-work turn. Levie agrees with the transferable template—give an agent resources, assign work, then review—but stresses that each domain adds hostile data, access, and liability conditions.

9. Private evals become operating infrastructure for every enterprise

  • Box supported the APEX eval by opening representative data-workspace material for lawyers, investment bankers, and other professions. Its own benchmark uses documents across roughly 10 industries, including public-sector, legal, healthcare, and financial-services scenarios such as data rooms and investment prospectuses.

  • The benchmark evolved from one-shot model testing into an agentic evaluation of both model and Box harness. A rubric scores required facts, and the data is held out from Anthropic and unavailable publicly, preventing a model provider from deliberately training against it.

  • Levie describes “incredible jumps” within model families, citing roughly a 15-point overall gain in one comparison and specifically contrasting Sonnet 4.6 with Sonnet 4.5. The private setup helps distinguish genuine capability gains from leaderboard optimization.

  • Model selection is only half the purpose; Box changes its own agents daily and needs to catch regressions. Levie expects every enterprise to evaluate RFP generation, sales-material creation, invoice processing, and similar pipelines, making agent observability and eval platforms such as Braintrust and LangSmith a “massive space.”

10. The agent becomes a third customer for Box’s entire stack

  • Box historically designed its file system for two customers: human users and applications. The agent is a new kind of user with different workspace and retrieval requirements, including cases where Box may use embedding-based search rather than its typical semantic search.

  • Supporting it touches every layer: data storage, filesystem semantics, metadata, search, permissions, governance, compliance, and infrastructure. Levie describes active experimentation—“testing stuff, throwing things away”—while the agent team continuously generates new requirements for the surrounding organization.

  • The core agent effort is a few dozen people inside a company of roughly 3,000, surrounded by concentric support teams. Levie resists calling it an “innovation center” because innovation must remain company-wide; this group is distinct because getting the agent wave right is “do or die.”

  • The eval effort is led by Ditya and Siddharth, with CTO Ben, AI head Yash, and others involved. Existing security and compliance features are what make Box eligible as an enterprise agent platform, but they are not sufficient to win; the agent roadmap is existential.

11. Box sees a read-write workspace, not merely enterprise search

  • Reading is currently harder than writing because retrieval faces the “10 million to one ratio problem.” Writing can originate in the model and be saved directly, although generated PowerPoint files still fail visible details such as fonts, shapes, and consistent slide updates.

  • Box plans native agents powered by leading models, but Levie sees the larger opportunity in letting any external agent use Box as its filesystem. The agent might store memory, specifications, Markdown, PDFs, intermediate work, or generated deliverables without Box dictating the artifact type.

  • That workspace would be sandboxed yet collaborative: humans could inspect it, contribute to it, or share selected material with others. The combination of private workspace, governed enterprise inputs, persistent output, and controlled collaboration is the product expression of “every agent needs a Box.”

12. Documentation earns a premium, but a company cannot be frozen into skills files

  • Levie’s objection to representing a whole company as Markdown skills is not that documentation lacks value; it is that reality changes a week later. Markets, customers, and internal decisions continuously invalidate instructions, while substantial context remains in conversations that were never digitized.

  • The hosts’ sharper description is that “most companies are practically apprenticeships”: a new employee spends one to three months acquiring tacit knowledge. Agents expose how little of that operating context is written down or maintained authoritatively.

  • swyx argues that better capture could shorten a three-month ramp to roughly two weeks, reduce rework, and make an average employee perform more like the 90th percentile by distributing top workers’ knowledge. That is an immediate productivity argument for agent-ready documentation.

  • swyx also flags the scale problem: at a 10,000-person company, capturing everything does not mean sharing everything; the information must map onto real organizational access boundaries. Digitization without permission design merely creates a more searchable leak.

13. File systems and lightweight wikis may beat graph maximalism

  • Levie views “your company as a filesystem” as a productive metaphor because companies already collaborate through permissioned workspaces. He is less convinced that a formal knowledge graph automatically solves human messiness, recalling earlier cycles when enterprises were expected to run entirely on wikis.

  • His position is deliberately nonreligious: Box can feed somebody else’s graph, consume one, or let an agent query several systems. The durable requirement is governed access to changing information, not winning a debate between graph structure and Markdown simplicity.

  • Alessio argues that the useful graph may emerge dynamically “in the mind of the agent,” as it does for humans. swyx favors persistent agent wikis—linked Markdown as a weak, adaptable knowledge graph—and Alessio points to DeepWiki as evidence that documentation useful to humans can be even more useful to agents.

14. Founder attention follows existential risk, while distribution becomes technical

  • Levie says roughly 90% of Box’s work is delegated; across perhaps 70%-80% of the company, he needs to inspect only about 5% of activity through high-leverage decisions and processes such as quarterly reviews. He sees less distance from Brian Chesky’s founder mode than the hosts initially suggested.

  • AI is different because “two, three, four, five wrong decisions”—in architecture, features, APIs, or platform strategy—could remove Box from the game within a year. That pulls Levie into late-night product work, including an anticipated 11 p.m. Zoom after the recording, while still requiring collaborative leaders rather than dictation.

  • His personal production function links internal problems, public writing, and external feedback. A 20-minute commute and a 7:30-9:00 p.m. scan of AI news become time to distill lessons; public responses then feed back into Box. The instinct predates the company—an internship was once rescinded after he proposed blogging about it.

  • Alessio argues that every company may need to operate as a media company, while DevRel becomes increasingly important because services and APIs must attract agents. swyx adds that software may produce far more features per dollar while companies spend comparable effort getting those features to customers through technical deployment and education.

  • The hosts’ labor-market call is that software output could increase “10 to 100 times,” making technical ability more—not less—valuable. Whether enterprises build their own systems or buy packaged software, engineers will deploy agents, maintain integrations, translate business problems, and support a world in which software reaches every domain.