Why Asking the Right Questions Is the Most Important Skill in the AI Age
Why Asking the Right Questions Is the Most Important Skill in the AI Age
Summary
- The rumored $20B NVIDIA partnership went from first call to money in the bank in about three weeks. Jonathan Ross went to Jensen Huang asking to buy roughly 100,000 GPUs to deploy a hybrid GPU-LPU system Groq had already built; Jensen instead “thought maybe it would be better to make this available to all of their customers.” Ross corrects the desperation narrative: Groq wasn’t cash-strapped this time; the value of the agreement was only a little over 2x the prior valuation, and Groq could have raised at the amount of the licensing.
- The technical thesis is that LPUs and GPUs are complements, not substitutes. “18-wheelers or vans for last-mile delivery—which one would you pick? The answer is both”: compute-constrained matrix multiplies go to the GPU, memory-throughput-constrained ones to the LPU, because “there is no one perfect architecture” and combining them “defeats the bottlenecks.” What many people get wrong is splitting prefill from generation across hardware—the wrong cut, since “the generation of tokens is the hard part.”
- Speed makes models smarter, not just faster—Ross’s proof is AlphaGo on the TPU he created at Google. Ross gave approximate Elo figures of about 3,200 for AlphaGo on GPUs and about 3,550 for Lee Sedol, while warning he might have a leading digit wrong; the TPU result was described in the exchange as roughly 3,900 or 2,900. The same model found Move 37, a 1-in-10,000 move, only after the hardware enabled deeper search. “Being able to think faster makes you think smarter”—amplified by agentic AI-to-AI traffic, where speed compounds exponentially.
- Fast inference was out of favor as recently as three or four years ago—including inside Groq itself. People were leaving while saying it added no value, customers asked “why do I need an LLM to be faster than I can read,” and Ross twice let his team talk him out of LLM opportunities—including a direct call from GitHub’s CEO wanting chips for code completion. What finally worked was letting people try it: a viral X video of an LLM running on Groq did what years of first-principles argument could not.
- Capital is no longer the winning bet in AI, as Senra frames Ross’s VC story through the Keynesian beauty contest. VC herding was once rational because the most-funded startup could appear advantaged, but “for the first time in history, startups are not starved for cash… putting more money in is not an advantage. But people are still acting as if” it is. The kicker: typical West Coast VCs passed on Groq, while East Coast crossover funds invested in what Senra called NVIDIA’s biggest deal “by almost 3x.”
- The defining AI-age skill is asking questions, not answering them. “Success in the information age was about being able to answer questions; success in the AI age will be about being able to ask the right questions.” Everyone shifts from ICs to “leaders of AI,” and Ross argues that school curricula should center on real community problems that students decompose into questions for AI.
- Code’s marginal cost “is approaching zero,” opening software creation to many more founders. Ross’s EA now builds travel apps without knowing how to code; the literacy analogy says software creation is becoming accessible to people who previously lacked technical ability but may have good taste. He expects individual founders without large teams to create valuable companies.
- Ross’s current manufactured discontent is the world’s compute scarcity. “If it takes us an extra year to cure cancer because we don’t have enough compute, that’s my fault.” The backstory: Groq was once three weeks from running out of money and preserved the team through salary-for-equity “Groq bonds,” in which 80% of employees participated and about half went down to the statutory minimum.
Deep dive
1. Three weeks from phone call to wire: how the rumored $20B NVIDIA deal happened
- Ross’s telling: Groq had already implemented GPU-LPU integration and went to Jensen asking to buy about 100,000 GPUs for its own deployment—Jensen saw the work and thought it might be better to offer it to all NVIDIA customers. “The call where the idea was first floated was about three weeks before money was in the bank.” Senra: “So Jensen moves fast.” Ross: “Of course, that’s how you stay ahead.”
- The mechanism, in Ross’s logistics analogy: 18-wheelers or last-mile vans—“the answer is both.” Processing an LLM token involves matrix multiplies that are variously compute-constrained (GPU) or memory-throughput-constrained (LPU); “the bottlenecks are all over the place… there is no one perfect architecture,” so pairing the two “defeats the bottlenecks.”
- Ross deflates the desperation narrative: the three-weeks-from-death episode was “many years ago”; this time “we were fine.” He says the value of the agreement was only a little over 2x the last valuation, and Groq had the ability to raise at the amount of the licensing.
2. AI using AI: speed, micropayments, and the interrogable daily brief
- Why speed compounds now: “A human can wait a second or two… AI is just sitting there waiting because it produces these tokens so much faster.” Agents kick off jobs to other agents—“AI is really good at using AI. That’s really what agentic is”—producing exponential fan-out where latency matters throughout the chain.
- On agent payments, hedged as early: payments “aren’t really built for this yet, but if you can make micropayments, the number of payments is going to skyrocket.” His example: needing phone numbers for a Signal/WhatsApp agent hobby project meant proving he was human at Twilio; with a delegated budget, the AI could have spent within the budget without his involvement.
- Ross’s hobby projects seed his work practice—including a personalized “presidential daily brief” that evolved from long text to headlines plus follow-up questions: “As I’m learning through AI, I’m not reading a static piece of content. I’m interacting with it,” like the game 20 Questions. Senra’s parallel: Spotify co-CEO Gustav Söderström built something similar as a personal podcast, excluding rage bait and politics while surfacing what people he follows are discussing.
- The thesis Senra pulled from Ross’s tweets: “Success in the information age was about being able to answer questions. Success in the AI age will be about being able to ask the right questions.” Everyone moves from ICs to “leaders of AI”; school trained us to memorize answers, but now “you just ask AI. It knows. You just have to think of the right question. It’s a fundamental shift.”
3. Leadership is having followers—then finding the form true to you
- Ross’s first principle, which he says he got from a Jon Levy book: “You’re not a leader unless you have followers”—and like investing (venture, debt, seed, crossover, PE), there are infinite forms. New founders’ mistake is executing borrowed advice that is “not true to them.”
- His self-knowledge: unlike the control freaks on both of Senra’s podcasts, Ross has not held a driver’s license since 18—“I don’t need to control driving. I want to control thinking”—and hires autonomous people “who would be terrible in most corporate environments.” Corollary for early careers: work somewhere whose leadership style teaches lessons you can actually use.
- The confession: “I was one of the world’s worst leaders when I started”—it cost Groq three to four years. He delegated to people who could not operate autonomously, things ground to a halt, and his belated commands were “so unnatural to me that they didn’t accept it.”
- The fix: a goal simple enough for a challenge coin—every Groq employee carried one reading 25 million tokens per second—plus minimal constraints: “The fewer constraints that you give someone, the more freedom they have to surprise you with the solution.” Senra’s echo from Kelly Johnson of Skunk Works: “Extreme performance often comes from one brutally clear priority.” Ross doubles down: your team can only innovate “if they can surprise you in a good way, which means you must not over-constrain the goal.”
4. Lessons from inside NVIDIA: kill the one-on-one, and the confidence unlock
- NVIDIA is “the least political large organization you will ever see,” and Ross traces it to a mechanism: Jensen never tells one person one thing. One-on-ones produce divergent interpretations and side cliques; big meetings produce one shared message. His rule: copy the person being discussed on the accusatory email—“otherwise you’re allowing politics to happen.”
- Second Jensen lesson: Ross “got way too cute trying to play 3D chess,” while Jensen just asks “what does the customer need? Just build that for them and everything else follows”—including not selling customers things he does not believe they need. Senra adds Jensen’s line on his roughly 60 direct reports, each smarter in their domain, as a template for managing AI agents.
- Ross’s confidence story: at roughly 35 employees he shadowed someone running a 2,000-person organization, silently pre-deciding each call—and matching every one. “I didn’t change my decisions, but it changed my leadership”: people follow decisions delivered with confidence. And on scale: a 450-person creative organization was “more like managing a group of 5,000” in some ways; “the better the people, the harder they are to manage.”
5. Lemmings, East Coast skeptics, and the Keynesian beauty contest
- Ross’s fundraising sociology: “Typical West Coast VCs are more like lemmings”—one pass cascades into universal passes—while typical East Coast VCs “all think they’re smarter than each other” and run their own analysis. Groq ended up funded by East Coast crossover funds; Senra’s gleeful note was that the biggest deal NVIDIA had done, by almost 3x, came after West Coast VCs missed it.
- Senra’s frame, the Keynesian beauty contest: you bet not on the most beautiful model but on the most-bet-on model. He uses it to explain why following other investors could once seem rational. But Ross argues that “for the first time in history, startups are not starved for cash… putting more money in is not an advantage. But people are still acting as if putting more money in gives that startup an advantage.”
6. Faster is smarter: the AlphaGo proof and the deal’s real genesis
- Credit where due: the chip-combination idea came from COO Sunny, an example of Ross’s autonomy principle, and took probably three or four months of work, “maybe a little bit longer.” What many people get wrong is splitting prefill (reading) from generation across hardware, but “the generation of tokens is the hard part. Reading is easier than writing, and that’s true for AI as well.” Groq was not afraid to show NVIDIA the result because it wanted to become a GPU customer.
- Why Jensen acted instantly: LPUs are “like getting broadband instantly on these existing models”—unlike the actual broadband transition, no one has to rebuild websites for the speedup to land.
- The deeper claim, from Ross’s Google TPU days: DeepMind emailed 30 days before the world-champion Go match, AlphaGo lost its GPU test games, and the team ported it to TPU. Ross gave approximate figures of about 3,200 Elo for AlphaGo on GPUs and about 3,550 for Lee Sedol, while warning he might have a leading digit wrong; in the exchange, the TPU result was described as roughly 3,900 or 2,900. The same model had more compute, and Move 37, a 1-in-10,000 move, was not found on GPUs because it was too deep in the chain.
- The concession and conclusion: “As the creator of the TPU, I have to admit GPUs are now better”—ecosystem wins—but an LPU lets a model search deeper faster. “Being able to think faster makes you think smarter.”
7. Reality quotient, the dominant game, and change management as the whole job
- Groq hired for “reality quotient,” distinct from IQ: “There are plenty of really smart people who wouldn’t recognize reality if it tapped them on the shoulder.” Its extreme form is choosing the dominant game—MySpace maximized accounts signed up, Facebook maximized monthly active users, and “if you maximize the monthly active users, you’re going to beat someone who’s maximizing accounts signed up.”
- The 25-million-tokens-per-second goal was that dominant game made legible: chip speed, software, power costs, data centers, fabrication, and supply chain—everyone could connect their work to it.
- Ross’s engineer-to-founder click: “My job was full-time change management. And the first principle of change management is to make it feel like it isn’t a change.” People anchored on the goal experience an approach change as no change at all; the duty is giving enough context that “their job hasn’t really changed.”
8. Return on luck—and the fix of “I intend to”
- From Jim Collins: the best companies do not get more luck; they seize it better. Ross’s counterexample against himself: GitHub’s CEO called needing chips for LLM code completion because GPUs were unavailable; his team said “nope, not going to work,” and he let them convince him—twice, across two opportunities—“even though in my bones I kind of knew that we should.” “How much better off would we have been had we been the original inference engine running LLMs at Microsoft for OpenAI?”
- The third time he did the arithmetic himself, everyone disagreed that it was possible, and Groq “ended up hitting exactly those performance numbers.” Meanwhile fast inference was so out of favor three or four years ago that people were leaving over it; Ross’s diagnosis of the blank stares: “When people don’t understand the first principle of something and they’re getting involved because it’s hype, they don’t understand enough to understand why what you’re doing is different.”
- The marketing unlock mirrored the ChatGPT moment. Ross watched an Anthropic demo about three months before ChatGPT land flat—magic arrives when the answer is specific to you. Groq put its speed online, someone posted a video on X, and Ross discovered during a presentation in Norway that queries felt slow because usage had skyrocketed: “We just went viral.”
- The management fix, from David Marquet’s Turn the Ship Around: intent-based leadership. Asking “Should I do this?” invites pessimism; saying “I intend to do this” generally gets no opinion unless something is truly wrong—the submarine crew that would finally say, “Wait, the hatch is open.” On the third opportunity Ross said “I intend to,” and “rather than people going, ‘We can’t do this,’ they all jumped in and said, ‘This is how we do it.’”
9. Groq bonds: everyone’s hands on the steering wheel
- Groq was three weeks from running out of money while still pre-product and building a never-before-attempted compiler that eliminated the need for humans to write kernels. Ross reviewed the proposed layoff list and concluded, “If we did that layoff, we were dead.” The only answer was cutting burn without cutting people.
- So he held an all-hands with World War II-style war-bond pictures and introduced “Groq bonds,” a salary-for-equity exchange. 80% participated; about half cut their pay to the statutory minimum—engineers earning hundreds of thousands going to “$50, $60,000… real pain.” The program saved more than three weeks of burn, probably closer to two months; Groq raised with three weeks of cash left.
- The expected attrition never came: under 10%, perhaps closer to 5%. Ross’s phrase for why: “Put everyone’s hands on the steering wheel”—passengers fear the windy road; drivers feel more in control and are more willing to take the risk.
10. Hire by negatives, book the win early, and taunt like Jordan
- Groq kept a versioned “people spec”: “If you don’t write down what you’re looking for in people, you’re not going to hire that.” Attributes like return on luck and “poetic design”—“poetry is semantic density… every word matters”—each had a negative twin: squanders luck and maximalist design. “What you’re really hiring for is to avoid those negatives,” because one person can bring a damaging trait into the whole team. His biggest flip: growing talent means showing positives; selecting talent means screening negatives—“very different mental modes,” learned while watching a head of HR who excelled at removing problems.
- One prized trait was loss bias turned productive. In architecture meetings, “someone would say, ‘If we do this, the chip will be twice as fast,’” and nobody moved—they heard “next chip”; Ross heard, “If we don’t do that in this chip, the chip’s going to be half as fast as it could be.” He hires people who “book the win early.”
- On Senra’s Michael Jordan episode: Ross suspects Jordan’s compulsive betting and taunting were intentional stake-raising—a public loss would be humiliating, forcing superhuman performance—the same mechanism as founders who announce success before achieving it. Senra cites Tim Grover’s account: “Once you tell somebody how bad you’re going to [beat] them, you have to actually go and do that.”
11. Manufactured discontent, free code, and teaching kids to ask
- From a room of successful people, Ross’s observation was that some entrepreneurs with hundreds of millions were still unhappy with their wealth, while others were unhappy with their previous work product. Everyone had some discontent that drove them. “You have to have a personality where you are constantly discontent if you’re going to keep pushing things forward.” His current one is compute scarcity: “If it takes us an extra year to cure cancer because we don’t have enough compute, that’s my fault.” Senra’s parallel is Edwin Land’s whiteboard tally of headlight-glare deaths per day of delay.
- The optimistic close: “Code rationing” is ending—“the marginal cost is approaching zero”—shifting software creation from something controlled by technical specialists toward broad accessibility. Ross’s EA now builds live travel apps; “a lot of people are going to get access to creating software… who would never have had the technical capabilities before, but who would have had good taste and known what good is.” He expects individual founders without large teams to create valuable companies.
- His answer to parents: “Stop teaching them to answer questions and start teaching them to ask questions.” Rebuild curricula around real community problems—permitting, local events—where students write useful applications: if they can look up the answer online or ask AI to solve it, “you haven’t taught them what they need… but if you give them a problem where they have to ask the questions and get AI to solve it, then you have.”