Pioneers Insight Method Research Author
Building the Cloud for AI Agents | AWS CEO Matt Garman
Back to Episodes

Building the Cloud for AI Agents | AWS CEO Matt Garman

Summary

  • AWS is betting that AI demand is durable enough to justify $220 billion of 2026 capex, with no slowdown anticipated. Matt Garman argues the risk is tempered by AWS’s diversified customer base—individual-customer concentrations are in the single-digit percentages at the highest—and by production workloads already generating positive ROI. “There’s no bubble in which they stop spending on that.”

  • AWS is allocating scarce accelerators across the ecosystem rather than selling them all to frontier labs. Although AWS could sell every available chip to a few large labs, it reserves capacity for startups and eventually says yes in some form to roughly 60% of requests it receives, sometimes in another region or configuration. AWS plans to buy 2 million NVIDIA GPUs over the next couple of years, though Garman concedes, “Who knows if that’s enough?”

  • Agentic workloads push cloud architecture toward fast, potentially disposable, tightly permissioned resources alongside long-lived systems. Agents care about p99.9 latency, three-second database creation, compute sandboxes, gateways and time-boxed permissions rather than inheriting a person’s access. Garman’s core claim is categorical: “Agentic workflows tend to perform better on AWS than anywhere else.”

  • AWS is removing its traditional setup friction without forcing startups onto a simplified dead end. As the new flow rolls out, an account can be opened with Gmail, without a credit card, and become usable in under 30 seconds while VPC and IAM defaults are handled behind the scenes. Crucially, it remains “a real AWS account,” so customers can later expose the full security and organizational controls without migrating.

  • AWS’s custom-silicon path runs from Nitro and Graviton to Trainium. Graviton is described as 20% cheaper with 20% better performance, used by something like 90-plus percent of AWS’s top 100 customers; Trainium 3 capacity is sold out potentially through late next year, while Trainium 4 has been announced but not launched. Garman says Trainium may be “the best inference chip on the market right now” on absolute and cost performance.

  • Enterprise agent adoption is constrained less by model access than by workflow redesign, evaluation and trust. Garman rejects merely copying a human’s five-step process: agents can try 50 approaches in parallel, but enterprises need permissions, guardrails, labeled data, production measurement and drift testing before granting autonomy. AWS’s forward-deployed model aims to teach those skills in 45 days and leave, not create years of consultant dependence.

  • AWS is positioning Bedrock around enterprise data custody while keeping both proprietary and open-weight model paths available. Garman says prompts and data remain inside the customer’s VPC and never return to the model provider; customers with genuinely proprietary data may post-train or fine-tune open-weights models, potentially distilling them, for better performance at lower cost. The prerequisite is an evaluation system proving that advantage rather than assuming it.

  • Inside Amazon, agents are already compressing software, product and business workflows—and beginning to reshape team design. Garman says agent-first “frontier teams” manage agents that write all the code, while HR and finance employees build their own agents for planning, tax and compliance work. A capability that formerly required 10 people might now need three or four, creating a new problem: how to maintain products while moving small teams rapidly between projects.

Deep dive

1. AWS still sees itself near the beginning

  • Raghu Raghuram opens with the scale: AWS has grown from its first dollar of revenue to roughly $169 billion-$170 billion, now growing 37%. Garman nevertheless calls it “the early stages of what the business can be,” because substantial workloads remain on-premises while AI expands the total amount of compute performed each day.

  • Garman’s first AWS assignment was a 2005 business-school internship analyzing who would value the proposed service most. The answer was startups, which became “the lifeblood of the core of what we do” because AWS’s value proposition was especially attractive to them and helped them build architectures that could eventually scale.

  • That relationship has compounded commercially: AWS estimates that 30%-40% of current revenue comes from companies that were startups at some point during AWS’s lifetime. The two-person company matters not only as a future enterprise customer but as an early signal of capabilities that banks, healthcare companies and governments may demand five years later.

  • The startup baseline has changed radically. Where a company might once have raised $10 million to iterate on an app, Garman now sees teams with an idea, $200 million in funding and a $1 billion valuation from day one; their ambition, model-training costs and infrastructure ramp are correspondingly larger.

2. Agents expose a new cloud performance envelope

  • Some startup requirements have not changed: founders still need scalable architecture, security, performance and IAM that survives growth beyond three employees. Garman says this surrounding depth is why many ultimately prefer AWS to a neocloud, even when their immediate request is simply for GPUs.

  • What has changed is the user. AWS now designs for clouds operated by agents as well as people, emphasizing broad-scale API access, fast resource creation and predictable performance. Garman offers a concrete example: “How can you start a database in three seconds?”

  • AgentCore and Bedrock are explicit agent-building services, while “AWS Context,” described as being in preview or beta, is intended to create a context layer across data in S3, Aurora and other AWS stores. An agent can traverse those separate data estates in a way a person typically would not.

  • Agents also magnify tail latency. A human may never notice p99.9 S3 performance, but an agentic workflow can block on it; latency, throughput and the speed of the underlying engine therefore become orchestration constraints rather than abstract infrastructure metrics.

3. AWS is simplifying the on-ramp without removing the controls

  • Raghuram’s pushback is that coding agents now select databases, email systems and deployment targets themselves, potentially making decade-old services feel “legacy.” Garman answers that the core building blocks remain sound; the larger gap has been the usability layer for someone—or some agent—starting without an existing cloud account.

  • In a typical established workflow, a customer tells Kiro, Claude or Codex to build on AWS, supplies credentials and deployment instructions, and the agent handles the rest. An experiment that is not already on a cloud and merely says “deploy,” however, may choose a partner with an easier layer on top—an outcome Garman says AWS welcomes but also wants to address directly.

  • AWS’s new-account flow is being rolled out so users can sign up with Gmail, omit a credit card and avoid manually defining VPCs or IAM roles. Defaults are handled behind the scenes, producing a functioning account in under 30 seconds.

  • The design choice Garman stresses is continuity: this is not a toy account that later requires migration. When the customer needs an organization, custom VPC or fine-grained IAM, “you’re already in a real AWS account” and can progressively expose those controls.

4. Disposable agent infrastructure needs production-grade escape hatches

  • Agents often create a database, perform a small task and discard it, raising a genuine design question: does that database need five nines of durability? AWS’s traditional Aurora posture assumes a production asset requiring durability and availability, which can be “overengineered” for a transient agent task.

  • Garman resists solving that with a casually non-durable option because AWS cannot always know whether a temporary resource will become permanent. The engineering objective is therefore dual-use infrastructure: quick and resource-light enough to discard, yet capable of growing into a durable production database without replacement.

  • Some agent requirements are genuinely new building blocks rather than different uses of old ones: compute sandboxes, gateways and permissions distinct from human or service-role permissions. “You don’t just want to give it Raghu’s permissions”; an agent may need narrowly scoped, time-boxed authority for one task, perhaps without access to an entire tool.

  • Firecracker microVMs provide a ready substrate. Developed roughly a decade ago and now used by many sandbox startups, they start rapidly, impose little traditional-VM overhead and retain a strong security boundary—making them well suited to agent execution.

5. Scarcity makes GPU allocation a strategic portfolio decision

  • Raghuram frames the conflict directly: frontier labs can “gobble up every available GPU,” while smaller companies lack comparable credit but may become tomorrow’s enterprises. Garman agrees AWS could allocate every accelerator to a handful of labs, but says it intentionally reserves supply for startups, enterprises and a broader ecosystem.

  • The buildout is enormous: Garman cites $220 billion of capex for 2026 and says AWS does not anticipate slowing because demand remains “massive.” Constraints rotate among power, data centers, capital, memory, chips and even the construction labor required to erect facilities.

  • AWS says yes in some form to roughly 60% of the requests it eventually receives, though fulfillment may arrive later, in another region or with a different configuration. Every startup still wants more; the company’s announced purchase of 2 million NVIDIA GPUs over the next couple of years may itself prove insufficient.

  • Garman distinguishes AWS from providers whose top one or two customers can represent 30%-60% of capacity. AWS’s individual-customer concentrations are in the single digits at the highest, while most usage is core compute, storage and inference tied to applications; nearly every enterprise he asks says current AI capabilities already produce positive returns.

6. The limiting resource keeps moving down the supply chain

  • Asked to identify the 2027-2028 bottleneck, Garman invokes The Goal: “There’s never one constraint; there’s always just the latest constraint.” Once power eases, the limit might become memory, TSMC capacity, HBM, networking gear, connectors, disk drives or SSDs.

  • Geography compounds the issue because capacity is not wholly fungible: ample power in Indonesia does not cure a shortage in Germany. AWS tracks tens or hundreds of thousands of components and looks four or five tiers into the supply chain for parts that could halt deployment.

  • Planning horizons have expanded from asking a utility for another few megawatts to financing solar, nuclear and other power projects, sometimes behind the meter and sometimes feeding the grid. Power and transmission decisions now extend 20 years, while planning for server, memory and chip needs stretches across 2026, 2027 and 2028.

  • Data-center opposition raises a communication challenge. Garman says AWS needs to communicate renewable-energy, water and employment benefits more clearly, citing one county where people reportedly pay $5,000 less in taxes annually because of the taxes AWS brings—an invisible benefit residents had not been told about.

7. Nitro led AWS from virtualization offload to custom AI silicon

  • AWS’s chip path began with the “virtualization tax.” It first moved network virtualization onto an offload card, then worked with a small team whose Arm-equipped card could absorb storage virtualization and other functions; acquiring that team ultimately enabled Nitro to expose near-bare-metal resources through APIs.

  • The payoff was not only utilization and performance but isolation: Garman says AWS could credibly tell customers it had no access to their running VMs. Because the architecture was not a generalized component others could simply buy, he argues it produced a decade-long lead.

  • Turning those Arm cores into a server yielded an initially underpowered Graviton, followed by a “runaway hit” as Arm performance improved. Garman says Graviton has remained roughly 20% cheaper with 20% better performance for five or six years; something like 90-plus percent of AWS’s top 100 customers use it, and some fleet migrations halved server counts.

  • AWS began Trainium five or six years ago and is now shipping Trainium 3; capacity is sold out potentially toward the end of next year. Most Bedrock traffic runs on Trainium, alongside Anthropic and OpenAI agreements and roughly half a dozen to a dozen smaller startups building on it.

8. Enterprise autonomy depends on redesign, evaluation and data trust

  • Most enterprise agents today remain relatively simple and non-autonomous, though customers already report value. Garman’s first prescription is to stop reproducing “Bob does steps one, two, three, four, five”; an agent can parallelize, try 50 approaches and solve the desired outcome differently from a human workflow.

  • The second barrier is justified nervousness about autonomy. Enterprises need guardrails, sandboxing, data permissions and clarity on when a human stays in the loop; “go nuts” is not an acceptable instruction when an agent could delete a production database.

  • Raghuram lists the evaluation requirements—labeled data, a constant testing loop, goal-seeking criteria, production measurement, back-testing and drift detection—and Garman says enterprises do not know how to solve these today. He doubts anyone is really great at solving them, motivating forward-deployed engineering engagements intended to teach the capability in 45 days and then leave the customer self-sufficient.

  • On model choice, Garman agrees enterprise data is “their most valuable asset.” Bedrock guarantees that data stays within the customer’s VPC and that model providers never see prompts, a foundation AWS prioritized even when critics said it was moving too slowly three years earlier.

9. Open models and machine-speed security broaden the AWS layer

  • Customers with meaningful proprietary data may post-train or fine-tune an open-weight model, distill it and potentially achieve better performance at lower cost. Garman keeps the claim conditional: they need the right data, expertise and evaluations to prove the custom model actually wins.

  • Raghuram calls this “a new lease of life” for SageMaker; Garman says it was always a model-building platform and is now a good fit for enterprises building custom models. He wants AWS to make comparison, tuning and testing across open-weight models progressively easier.

  • CEOs’ overriding question amid existential-risk debate, security vulnerabilities and the Hugging Face attack is still practical: how can they trust agents inside their environments? AWS’s answer spans permissions, sandboxes, guardrails, explicit allowed behaviors and human review where appropriate.

  • Garman describes Continuum as turning powerful models toward defense: it scans environments for vulnerabilities, then prioritizes them using context about permissions and compensating controls. His destination is “security at machine speed,” replacing a workflow where an alarm waits for a human investigation.

10. Amazon’s own agent adoption is beginning to reshape teams

  • Amazon uses AI across security and software development, while Amazon Q has been rolled out to every employee. Garman describes HR staff compressing weeks of team-planning work into hours and finance teams using agents to collect tax rules and ensure compliance.

  • The largest gain is software and product development. AWS’s “frontier teams” work agent-first rather than treating AI as code completion: agents write the code while employees manage teams of agents, producing what Garman calls a “turbo boost” in the release of customer capabilities.

  • Organizational design remains unresolved. Garman expects people to remain important for a long time, but a product capability that once occupied 10 employees may now need three or four—and may be built quickly enough that those people should move to another problem.

  • AWS is experimenting with pods and more fluid staffing while confronting the maintenance question: somebody must continue operating what a small team built. Garman offers no magic structure yet, only the observation that employees like building faster and doing more—and that “there’s real work there.”