Why Tech C.E.O.s Are Blaming A.I. for Mass Layoffs
Summary
Tech’s latest cuts are an early warning, but not yet clean proof that AI has directly replaced workers. Atlassian is eliminating 10%, roughly 1,600 jobs; Block about 40%, or 4,000; and Meta was reportedly preparing cuts of 20% or more, potentially 16,000. Casey Newton’s calibration is the useful one: AI differs in causal importance at each company, yet “sooner or later, I do think we’re going to have to believe them.”
The hosts disclose relevant conflicts: Kevin Roose works for The New York Times, which is suing OpenAI, Microsoft, and Perplexity; Casey Newton’s fiancé works at Anthropic.
For public companies under pressure, AI offers both a productivity thesis and a marketable explanation for old-fashioned restructuring. Block had expanded from roughly 3,800 employees in 2019 to more than 10,000, then spent $68 million flying 8,000 people to an event with Jay-Z five months before its cuts; its shares rose 17% the next day. That makes “AI-washing” difficult to separate from correcting pandemic-era overhiring and management failures.
Kevin argues that Meta’s proposed labor cuts could shift costs from payroll to AI infrastructure rather than reduce aggregate spending. The company plans $135 billion of capital expenditure this year while Zuckerberg argues that projects once requiring large teams can now be completed by “a single very talented person.” Yet the productivity case remains speculative: Meta abandoned Behemoth, reportedly delayed Avocado for missing targets, and apparently achieved only a slight improvement over Gemini 2.5.
Employees are being placed in a no-win adoption test: use AI heavily to demonstrate alignment, or avoid proving that their work can be automated. The hosts argue that repeated layoffs also quiet internal dissent, whether or not workforce discipline is an explicit objective. Kevin Roose consequently revives a prediction he previously got wrong: fear could drive tech workers toward unionization and bargaining over retraining or reassignment.
Chatbots’ mediocre literary voice may be a consequence of product optimization, not an underlying inability to generate surprising prose. Jasmine Sun preferred aspects of GPT-2 and GPT-3 because they were variable, strange, and capable of style matching, whereas post-training and RLHF pushed later systems toward the “helpful assistant”: chirpy, sycophantic, safe, and repetitive. The commercial demand is for excellent corporate emails, while artistic quality lacks the verifiable rewards that accelerated coding.
The defensible near-term writing product is a personalized collaborator, not an autonomous author. Sun says text generation occupies only about 25% of her work; reporting, idea selection, reading, judgment, and lived experience remain harder to reproduce. Her productive Claude workflow instead learns her archive, retrospective notes, audience, and aspirations, then asks questions that push her toward “the best version of myself as a writer.”
Token consumption is becoming a new and potentially enormous component of technical labor costs. OpenAI’s top employee reportedly used 210 billion tokens in seven days—about “33 Wikipedias” of text—while Anthropic’s highest individual Claude Code user spent more than $150,000 in one month. One Swedish engineer said he probably spends more on Claude than his salary, turning unlimited access into both a job perk and a retention mechanism.
Token leaderboards confuse adoption with output and invite classic Goodhart’s-law gaming. Companies are incorporating usage into performance reviews—an employee might be challenged for using “only” 70 million tokens—despite no clear relationship between consumption and value. The investor-relevant question is therefore not who is token maxing, but whether escalating inference spend produces products, revenue, or merely “tasks of uncertain value.”
Deep dive
1. AI has entered the layoff rationale before proving the substitution case
Kevin Roose opens with the scale: Atlassian is cutting about 1,600 jobs, Block roughly 4,000, and Meta was reportedly preparing its largest reduction since the 20,000 jobs eliminated in late 2022 and early 2023. Meta called the report “speculative,” and the cuts remained unconfirmed at recording.
The hosts disclose relevant conflicts: Kevin works for The New York Times, which is suing OpenAI, Microsoft, and Perplexity; Casey Newton’s fiancé works at Anthropic.
Casey resists one explanation for all three companies, but his highest-level conclusion is firmer: executives keep identifying AI as a significant workforce factor, “and sooner or later, I do think we’re going to have to believe them.”
Kevin treats tech as an early-warning market because its workers will be among the first to see jobs change or disappear. Casey’s sharper worker-level point: whether AI caused a dismissal “does it actually matter if the effect on workers is the same?”
2. Atlassian and Block expose two different versions of AI-washing
Atlassian CEO Mike Cannon-Brookes said AI was not simply replacing people, but that it would be “disingenuous” to deny changes in the skills mix or number of roles required. His stated objective is to adapt “thoughtfully, decisively, and quickly” for durable, profitable growth.
Casey’s framing places Atlassian inside the “SaaSpocalypse”: customers may eventually build structured workflows cheaply rather than pay established software vendors as much. With the stock battered, layoffs create a new market story—fewer workers, higher productivity—though Casey gives Cannon-Brookes a pass for acknowledging that AI is only part of the explanation.
Block looks less clean. Headcount rose from about 3,800 in 2019 to more than 10,000, and five months before eliminating roughly 40% of staff, the company spent $68 million flying 8,000 people to an event with Jay-Z. Casey’s verdict: AI may raise remaining-worker productivity, but “you also could just say this company has been mismanaged.”
The market nevertheless rewarded Jack Dorsey’s story: Block shares jumped 17% the day after the announcement. Kevin sees narrative power in presenting cuts as forward-looking AI adaptation; Casey compares it with peak crypto mania, while observing that “the public markets actually can’t just be tricked that easily.”
3. Meta is transferring spending from payroll to compute
Reuters reported that Meta could eliminate 20% or more of its workforce, potentially 16,000 jobs. The proposed cuts sit beside the company’s planned $135 billion of capital expenditure this year—“real money,” even at Meta’s scale—and may reassure investors that its largest-ever bet has some expense discipline.
Zuckerberg supplied the productivity premise: “Projects that used to require big teams now can be accomplished by a single very talented person.” Kevin’s interpretation is not that technology lowers total costs today, but that companies are shifting expenditure “from human labor to AI.”
A venture capitalist told Kevin that some especially AI-native startups already spend more on AI tools than payroll. The destination these companies imagine is one where salaries no longer dominate expenses and businesses instead purchase the models, infrastructure, and tokens on which their work runs.
Casey’s pushback is execution risk: Meta abandoned Behemoth because it “wasn’t very good,” reportedly delayed Avocado after missing performance targets, and apparently barely surpassed Gemini 2.5 before another partial AI reorganization. Frontier developers such as OpenAI and Anthropic are not making comparable mass cuts, although Casey notes that they employ far fewer people.
4. AI adoption has become an internal loyalty test
One big-tech employee described a genuine trap: heavy AI use might prove enthusiasm for the new program, but it might also demonstrate that the employee’s work can be automated. Kevin hears “fear and suspicion and mistrust” because workers know executives are planning reductions.
Casey observes that Meta’s earlier layoffs made employees quieter and reduced internal protests. He stops short of calling periodic cuts an intentional workforce-control mechanism, but says some executives would regard that effect as “a positive byproduct.”
Kevin revisits his failed prediction of sudden mass unionization and wonders whether it could happen in the next year or two. Unlike largely unionized manufacturing workers who bargained over retraining and reassignment, tech employees lack that union bargaining channel; Casey’s advice is pointed: nothing would make Zuckerberg angrier than “a union of software engineers at Meta.”
5. Post-training traded the old models’ weirdness for corporate usefulness
Jasmine Sun distinguishes competent language from literary writing: most human writing is bad, and models outperform many people, but even maximalist AI leaders remain cautious about art. Asked when GPT might write a Neruda poem, Sam Altman offered only the possibility of “a real poet’s okay poem.”
Sun found GPT-2 and GPT-3 more compelling stylistically than current ChatGPT. They lied, wandered, and could be unusable assistants—GPT-2 might answer a tax question with a story about an orphanage—but they were “surprising” and “nutty.” GPT-3 in particular was better at matching voices such as Paul Graham’s than ChatGPT 5.4 Thinking.
The loss, in her account, came through post-training. Example dialogues, prohibited language, and RLHF steered unpredictable base models toward a consistent helpful-assistant persona: excellent for office work, but constrained when creative prose depends on variable tone, surprise, and risk.
6. Literary quality does not fit the rewards that made coding improve
The evaluation machinery can become absurdly reductive. Listings might offer a creative-writing expert $45 an hour while requiring a New York Times bestseller and starred Kirkus review; one Scale AI contractor was instructed to penalize three exclamation marks and assess fan fiction for factuality.
Kevin calls the evaluator problem “the whole story”: “We are taking the entire internet and grading it on factuality.” Sun adds that code can be tested by whether it runs, while experts can debate for decades over what makes Shakespeare or Neruda good; subjective art cannot be reduced to a consistently verifiable reward.
Demand reinforces the technical bias. Most users want “write this email for me,” a task at which models excel, and preference tests reward the bland corporate assistant. Sun therefore thinks both propositions are true: labs face a hard evaluation problem, and the market actively selects the resulting voice.
Sun’s deeper objection is that model language is “not grounded in a life.” Journalists observe scenes and interview people; poets write from emotionally consequential experience. Casey pushes back with models’ evocative writing about music despite never hearing it, while conceding they may simply recombine criticism written by people with ears.
7. AI can generate prose, but writing remains a larger system of judgment
Kevin poses the “cope” objection: programmers once listed everything models could not do, only to watch the gap narrow. Sun says she has spent three years trying to automate herself with Claude and failed, though she explicitly allows that style and literary generation might improve substantially.
Blind tests complicate the claim because readers sometimes prefer AI prose until told its source. Sun’s answer is occupational: text generation takes perhaps 25% of her day; the rest includes finding ideas, interviewing, selecting particular sources, reporting, and deciding what deserves to be written.
Genre fiction shows both capability and constraint. Sudowrite co-founder James Yu and other practitioners described the engineering effort required to undo models’ chirpy, sycophantic, PG-13 post-training. Human authors must keep prompting and “bullying the AI into getting weird,” making the successful arrangement a Centaur rather than autonomous authorship.
Sun’s best workflow makes Claude a personalized editor. She loaded a project with her archive, freelance work, post-publication notes, audience, beat, and goals, then co-developed separate ideation, structure, prose, and fact-checking rubrics while instructing it to evaluate—not write—her drafts.
8. Personalized evaluation turns Claude into a demanding collaborator
The useful feedback is specific to Sun’s aspirations: Claude identifies her “insider anthropologist” position in Silicon Valley and her movement between startup jargon, internet slang, policy, and personal scenes. That is categorically different from counting exclamation marks against a generic standard.
Instead of fabricating an ending, Claude might say her conclusion merely summarizes, recall that another piece ended more powerfully on a scene, and ask: “What were you thinking when the plane took off?” Sun retains judgment over whether to act, describing the objective as becoming “the best version of myself as a writer.”
All three writers feel pressure to preserve odd, colloquial, or blog-native lines as proof of human presence amid “slop.” Sun says AI has made her more comfortable with her loose, irreverent internet voice rather than pushing her toward professionalized newsroom prose.
Her final forecast is conditional: if labs devoted writing-level resources comparable to coding agents, they could plausibly produce strong literary text or turn transcripts into features. Whether that beats “automating 23-year-old software engineers” financially is doubtful; the hosts’ darkly comic alternative is that Grok writes the next great American novel.
9. Token maxing is becoming a costly proxy for modern engineering
Kevin defines a token as “the basic atomic unit of AI labor,” roughly a word fragment. About 10,000 tokens can generate 7,500 words, but agentic coding sessions now consume hundreds of thousands or millions as engineers run longer and more numerous processes.
OpenAI’s highest employee total over one recent seven-day period was reportedly 210 billion tokens—roughly “33 Wikipedias” of text, though some were cached rather than newly generated. Kevin’s conversations focused on this emerging “billion-token club.”
The expense is already salary-scale. Anthropic’s top individual Claude Code user reportedly spent more than $150,000 in one month; other extreme users burn thousands of dollars daily, and a Swedish engineer said he probably spends more on Claude than his salary.
Unlimited internal access becomes a powerful perk: some AI-lab employees use so much that another employer could not afford their habits, effectively making them costly to recruit. Engineering candidates are consequently starting to ask, “What’s my token budget?”
10. Leaderboards turn an imperfect signal into a gameable target
Employers use leaderboards for motivation and tracking, assuming heavier token users are adopting agentic engineering more seriously. Some now incorporate consumption into performance reviews, potentially creating conversations like: “It looks like you only used, you know, 70 million tokens last month. What’s going on?”
Productivity remains unproven. Heavy users may complete far more projects, but they may also generate worthless work; one person speculated that people at the top could be building side companies with their employers’ tokens. Casey invokes Goodhart’s law, and Kevin says he would not create a leaderboard at all.
The historical analogy is lines of code, another proxy engineers learned to game. Casey quotes the old comparison: “Measuring programming progress by lines of code is like measuring aircraft building progress by weight.” Token volume may correlate loosely with output while still failing as an individual target.
The practice is already escaping engineering: a marketer said her performance review now has an AI-use section and that her bonus might be based on how much AI she uses, despite creativity previously being the objective. Kevin rejects calling all token maxing theater, but Casey’s warning stands: “AI use for the sake of AI use” may produce expense, rivalry, and 24/7 agent swarms doing “tasks of uncertain value.”