E245|The Journalists Behind Foundation Models: How GPTs’ Replies Get Written
E245|The Journalists Behind Foundation Models: How GPTs’ Replies Get Written
Summary
- A new profession is absorbing the collapse of the content industry. More than 40,000 film and television jobs disappeared in Los Angeles alone over the past 2 years, nearly one-third of the total; newsroom jobs at US newspapers have fallen by half since 2008. A wave of former reporters, editors, and documentary directors has moved to foundation-model companies as “content engineers.” Meta created the role in 2025, seeking candidates with content, editorial, and film-production experience; the title has since splintered into “AI philosopher” and “creative technologist.” Google DeepMind has just created a senior model designer role, and OpenAI opened a similar position almost simultaneously—Silicon Valley is now putting a price on humanities judgment.
- Tony’s definition of the job is the sharpest line in the episode: “We’re not designing models, products, or writing styles; we’re designing human perception.” His point is that code is a technical language, but ordinary language is technical too. Everyone knows the New Yorker sounds different from the Daily Mail, yet struggles to explain why. Making clear what makes the New Yorker the New Yorker—and translating tone and style into rules that can be scored and trained—is precisely the craft journalists bring.
- Context is the first-principles problem in human-AI interaction. This year’s ICLR best paper found that breaking complex instructions into pieces and feeding them to AI over multiple turns causes a sharp drop in quality: when context is missing, AI fills in the gaps, answers grow bloated, and errors snowball. Tony connected the finding to his 2018 paper on Weibo flame wars, where humans likewise imagined that the other person “must not have children.” Establishing the backstory and causal chain before posing the core question, as in newswriting, is a prompt method he has found consistently effective.
- Taste itself can help you find a job. While ill, Tony wanted to build an AI host for his podcast. A pure system prompt failed—“by the third exchange, it had already forgotten the rules”—and adding intent-analysis and emotional agents made things worse. He ultimately drew random few-shot examples from his old podcast archive, and the AI “suddenly came alive.” A former media professional with no coding background tuned a Gemini Voice Agent that was more vivid in multi-turn conversations than the native model; a Google recruiter reached out, and “they were pretty stunned.”
- Sycophancy is not a writing-style problem; it is a product problem. “If AI argued with you every day, would you use it? You might, but would everyone else?” Designing an AI’s attitude is designing the product itself, and engagement and ROI require maintaining an “extremely, extremely fragile balance.” Tony flatly rejects the idea that users can simply be more critical: “That’s impossible—what are you thinking?” Expecting AI to be both flattering and objective is, at root, humanity outsourcing its responsibility for facing the world.
- Bianca’s core judgment is that AI is trained on consensus and optimized for the lowest common denominator, so it will not produce great art. The screenplay for Everything Everywhere All at Once could not have emerged naturally from AI, because AI would say it was illogical and violated the standard three-act structure. Tony is skeptical of the “AI steals jobs” narrative: the real target should be the gig economy that predates AI; grading a screenplay does not make AI as capable as a screenwriter—“doesn’t work like that; it isn’t that smart.” The third path is to use AI as a paintbrush: can a writer without a camera make a film, or someone without a VFX budget create effects themselves?
Deep dive
1. Content Engineers: Designing Human Perception, Not Models
- The opening numbers establish the industry’s backdrop: more than 40,000 film and television jobs vanished in Los Angeles within 2 years, close to one-third of the total; newsroom jobs at US newspapers are down by half from 2008. Some of those workers moved to Silicon Valley’s foundation-model companies—not to write copy, but to teach AI how to speak: what tone to use, where to draw the line, when to ask a follow-up, and when to stop talking.
- Tony, a former BuzzFeed, Bloomberg, NBC, and Vice journalist and a Peabody Award winner who hosts the podcast Planet Tavern, saw Meta’s 2025 posting for content engineering and thought: “Content experience, editorial experience, film-production experience—doesn’t every line describe me?” In the interview, he told them directly: “I can meet every requirement you listed. You won’t find many people on the market who fit all of them, so you might as well hire me.” He landed at Meta’s superintelligence lab, where he works on internationalizing model content.
- His definition of the profession: “What exactly are we designing—an AI model, an AI product, or a writing style? None of those. We’re designing human perception.” The day-to-day work may involve writing system prompts—“These prompts are actually highly creative writing. The model can be petulant sometimes; how do you train it?”
2. Ordinary Language Is Technical Too: Explaining What Makes the New Yorker the New Yorker
- Tony’s thesis is straightforward: “Coding is a technical language, but ordinary language is technical too.” Grammar and syntax can be quantified; good screenplays and articles can be broken down and analyzed to explain why one is better than another. People simply tend to read by instinct.
- Tony describes the job as an industry footnote: people know that the New Yorker, the New York Times, and the Daily Mail have different voices, but cannot explain the difference. “Our job is to explain what makes the New Yorker the New Yorker—to translate tone, style, and voice into something more technical.” Bianca puts it this way: “My path was first I learned how to make content and then I learned how to reverse engineer that content.”
- “Good” depends on the product. The same base model behaves differently when used to build a task-oriented agent versus a chatbot pretending to be your boyfriend; an airline customer-service bot must first clarify whether its goal is to solve a problem or get the customer off the line quickly. The method is to break attributes down, score them, and isolate variables: “Some answers are good, but they’re a 7 rather than a 10. We have to explain where those 3 points are.”
3. Understanding Intent, Not Just the Literal Question
- Bianca describes her specialty as “understanding intent, not just the literal question”—a set of skills she considers foundational to journalism. Her teaching example: asked whether Taylor Swift will get married at Madison Square Garden, a bad answer simply declares that she will and supplies a bridesmaid list, despite nobody knowing whether it is true; another mechanically stitches together information in a robotic voice. A good answer acknowledges uncertainty and gives the source and attribution behind any speculation. Those are the elements that make reporting better, and they all need to be built into training. Users often ask a literal question while seeking background, context, and the meaning beneath it.
- She emphasizes that hidden questions matter especially in non-Western cultures: “I’m Filipino myself. A lot of the information that really matters to us isn’t said directly.” That is why she encourages voice input. People do not repeatedly edit themselves in their heads, and AI can better capture the nuance. The analogy is the reporter’s craft: an oral interview is always better than an email interview because email is “over-edited by both sides”; the subject has too much time to polish the answer, and nobody can ask a real-time follow-up.
- Her favorite thing about AI is that it makes clear “the speed of thinking and the speed of acting are not the same thing.” Instead of feeding the model more material, users should tell it everything in their heads—the ideas, the background, and why the idea occurred to them. “The model already has a lot of training. What it needs is for you to express your intent fully and call those capabilities into action.”
4. Internationalization: Who Is China’s Meryl Streep?
- Tony works on a frontier that every AI lab is discovering. Before AI, taking content abroad meant translation: “Meryl Streep just won an Oscar” became “Meryl Streep won an Oscar.” With generative AI, the opportunity becomes recreation for different cultural contexts: Who is China’s Meryl Streep? What is China’s Oscar? “We could say that 斯琴高娃 just won the Golden Rooster Award or the Hundred Flowers Award.”
- AI can handle the language. What it lacks is the knowledge of who Meryl Streep maps to in a particular context—and that is difficult to teach. Finding equivalences across languages and regions can quickly trigger fan wars or cultural conflict. The best trainers are entertainment reporters: they follow these stories every day and carry a strong internal vertical model of the entertainment industry.
- Another example comes from Lu Chuan’s earlier appearance on the podcast. Faces generated by Chinese video models tend to look like the handsome men and beautiful women currently popular in China because the training data is biased. “If you use someone who understands film and knows what a movie face looks like, you can improve the model’s quality very quickly.”
5. When Context Breaks, AI’s Inventions Resemble Humans’ Inventions
- Tony thinks many people cannot ask AI good questions because they do not understand what context the situation requires. Newswriting provides the remedy: “You can’t just throw out a fact. You have to establish the context—how did this happen, what led to it—and then present the core question and argument.” That discipline is “incredibly effective in AI training and interaction. It works again and again.”
- The frontier research backs him up. This year’s ICLR best paper examined multi-turn conversations, a common weakness for AI. Researchers broke a complex instruction into pieces and fed it to the model step by step, then tracked quality across the conversation. The result was “a very, very large decline.” When AI lacks enough context, Tony says, it “keeps filling in the gaps based on its training”; the answer gets more bloated, and if one error appears in the middle, it grows “like a snowball.”
- Tony connects that finding to his 2018 research on Weibo. In the comment section around an incident involving a girl urinating in public in Hong Kong’s Mong Kok district, he watched “a very kind conversation turn into mutual abuse.” With no context, commenters began inventing identities for one another: “You must be this kind of person. You must not have children to say something like that.” His conclusion: “When context is fragmented, human invention and AI invention are quite similar. The findings of the humanities and social sciences are not that far from what the AI industry is finding.”
6. Journalism Skills Don’t Migrate; They Move Laterally: The Follow-Up Question Is the Product
- Asked how his international reporting experience “migrated” to the new job, Tony corrected the wording: “I don’t think it migrated. It moved laterally.” Journalists always have an audience in mind—what can and cannot be said, how to remain objective without offending people, how to preserve one’s own voice while staying faithful to the facts. That balance is complex and requires training. The instinct to guide an interview toward the next question transfers directly into conversational AI design as something immediate and intuitive.
- Face’s follow-up is worth dwelling on. One of a reporter’s most important skills is the follow-up question: when an interview subject veers onto an interesting point, the reporter has to say, “Can you expand on that?” even at the risk of derailing the interview. Give AI a broad objective, however, and it heads straight for the target. It does not notice an interesting side point and change direction. Does AI have that ability yet?
- Tony translates the question into product language: the core of a news interview is balancing structure with fluidity, and conversational AI design faces the same problem. Today, GPT uses a small model to generate 2 or 3 sentences restating the user’s question, answers in bullet points, and then generates a follow-up question. “That thing itself is a product. Which aspect should the follow-up target, when should it be asked, and how good is the question? You need a whole evaluation system—and naturally, the people best at asking questions should build it.” So again, these are journalism skills.
7. The Burning Building: AI-Training Gig Work and the Disassembly of Inspiration
- Tony and Bianca’s successful pivots sit on top of a broader collective collapse. A widely shared article in May was titled “I Work in Hollywood, and Everyone Who Used to Make Television Is Secretly Training AI Now.” Its author, Ruth, is a Hollywood screenwriter and producer whose writing fees kept getting delayed. She found AI Trainer work on a job forum: repeatedly grading, rewriting, and demonstrating answers for AI. Ruth jokes that the “Cambridge English literature star” she once was now takes orders from people in their 20s in a work chat, remains on call 24 hours a day, and competes for hourly gigs. On one side, $500B is pouring into AI infrastructure; on the other, film, television, and public media are facing an existential crisis.
- Face supplies the personal stake. She graduated from Columbia Journalism School 2 years ago—the alma mater of Tony and Bianca, and the institution associated with the Pulitzer Prize. One professor compared the job market to “people graduating from journalism school today and running desperately into a building that is already on fire.” As her applications disappeared into silence and she worried about next month’s rent, someone messaged her on LinkedIn asking whether she wanted AI Trainer data-labeling work. “If I hadn’t been so anxious that I couldn’t do anything, I think I would have taken it.” She puts Ruth’s question to the guests: Is AI turning us from creators into data collectors?
8. Consensus Cannot Produce Great Art; “AI Steals Jobs” Bundles Three Different Arguments
- Bianca’s answer goes straight to the training paradigm: “AI is trained on consensus. If you optimize for humanity’s lowest common denominator, it will not produce great art.” The best part of what makes us human does not come entirely from consensus. She and her husband write screenplays and have tried using AI, but “a lot of the time, we still really dislike what AI writes.” The clearest example is the screenplay for Everything Everywhere All at Once, which could not have emerged naturally from AI. AI would likely say the idea was illogical and did not follow the standard, Academy-approved three-act structure. “We’re going through a period of intense hype. Over time, we’ll start to see where AI actually fits.”
- Tony is “somewhat skeptical” of the broader narrative and breaks it into 3 points. First, the target of criticism is AI, but the real target should be the gig economy: media workers and artists were already taking on gig work before AI. “Now they’re grading things. Before, maybe they were driving Uber or doing DoorDash.” Second, the idea that people will “teach AI and then walk away” is overstated. Gig work is project-based. “After a screenwriter grades something, does AI become as good as a screenwriter? No. Doesn’t work like that; it isn’t that smart.”
- The third point is what Tony really wants to say: this narrative pits creators against AI and leaves only 2 choices—be anti-AI or sell your soul to AI. “There is clearly a third path: AI can be a partner to creators.” Can a screenwriter make a film without a camera? Can a filmmaker create visual effects without a VFX budget? From painting to photography, Photoshop, and Blender, each generation has moved through the same cycle. “People are so hungry now, they want an enemy and a target. They forget that it can also become a paintbrush.”
9. The AI Host Built During Illness: Taste Itself Can Get You Hired
- Tony sees himself as “the best example of turning AI into a paintbrush.” Last year, he discovered without warning that he had a serious health problem and took leave to return home for treatment. “For a period of time, he didn’t even know how much time he had left.” While recovering, he did not want to arrange interviews everywhere, so he decided to build a voice agent for his podcast. He assumed it would be simple. But GPT and Gemini’s voice modes “couldn’t be my podcast host at all. They felt like voice assistants, not podcast creators.”
- The iteration became a public masterclass. A system prompt packed with constraints “did absolutely nothing—by the third exchange, it had already forgotten the rules.” Tony then spent 3 days and 3 nights having philosophical conversations with AI about the nature of dialogue, built an agent to analyze the intent of every sentence, and added an emotional agent. “In the end, the whole agent was worse than the original version with no prompt at all. I added too many constraints and broke the entire AI.” The eventual solution was to separate questions from content across his podcast archive and have the AI randomly select several items as few-shot examples in each conversation. “It suddenly knew what a good question was, what good content was, and what the right tone was. It suddenly came alive.”
- The outcome carried a clear signal from the hiring market. A Google recruiter found Tony and was “surprised that a former media professional with absolutely no coding background had used his podcasting experience to tune a Gemini Voice Agent that was more vivid, flexible, and deeper in multi-turn conversations than their native model.” Several unicorns and labs approached him while he was recovering. “Content engineer” is already starting to sound dated; large companies now call these people AI philosophers and creative technologists. After recovering, Tony chose the opportunity that suited him best and became one of Google DeepMind’s first senior model designers. OpenAI opened the same role almost simultaneously, and several of his former journalism colleagues have already joined. “Whether you make podcasts or write articles, having this kind of feel and taste can itself help you find a job.”
10. Sycophancy Is Product Design, Not Writing Style—and Responsibility Comes Back to Us
- Face worries that AI will feed users the voice they are accustomed to, based on the other person’s values and information bubble—a phenomenon some users call “yes-man AI.” People are already turning to AI rather than journalists and real people for value debates. Over time, will that erode sharp disagreement and debate? Tony starts with a cold answer: “I can only say it is what it is.” Sycophancy looks like a writing problem, but it is not. “It’s a product problem. Put simply, if AI argued with you every day, would you use it? You might, but would everyone else?” Humans want everything at once: help, emotional validation, and not too much emotional validation. To engage users and deliver ROI, the product has to maintain “an extremely, extremely fragile balance.”
- Could we instead demand more critical thinking from users and tell them not to trust flattering answers? Tony’s response: “That’s impossible—what are you thinking? How could that work?” Face then points out that expecting AI to flatter you while remaining independently objective is really a way of outsourcing our responsibility for facing the world to AI. The responsibility ultimately comes back to us. It is like the newsroom premise behind fact-checking: audiences will not fact-check, and some do not even have the concept.
- Tony does offer a different possibility, though he flags the recollection for verification: “I vaguely remember a study showing that AI fact-checking actually caused more people to change their original views.” He tested the idea at home. When relatives sent an obvious fake news story in WeChat, he would normally decide there was no point arguing. This time he asked ChatGPT to fact-check it, then fact-checked ChatGPT himself and sent the result to his family. “They actually accepted it. I was pretty shocked.”
11. Who Is on the Other Side of the Chat Window: Understanding Human Collective Consciousness Across Time
- The episode closes on Face’s original question: when she types something unsayable into a chat window late at night, what exactly is answering her? Tony says that the essential skill of everyone who makes content is “empathy”—and that AI becomes more empathetic because people like her exist. Face’s realization follows: behind AI is not cold code but a form of creation. “It’s like reading a book. You read a writer from 500 years ago who is dead and never knew you, yet through reading you find resonance, feel understood, and feel that you’re not alone. AI is simply reorganizing humanity’s collective experience with language. It is not an individual, acting from its own subjectivity, who cares about me. It is the collective consciousness of humanity—a cross-temporal understanding from everyone who has ever resonated with me.”
- Face names the concept: the classic communications idea of a parasocial relationship. The object can be a character in a book, a television figure, or AI. “But you have agency. You can decide what kind of relationship exists between you, and create meaning through that interaction. That is what it means to be human. We are alive for this.” The line Tony once told Face that stayed with her most: in addition to IQ and EQ, this generation will need AI quotient—learning how to live with AI.
- The final question is whether Tony’s judgment about content changed as he moved from reporter to model evaluator. “My answer may not be what you expect—it isn’t a lateral move. The biggest change is that I used to focus on the result: was the writing good, was the film worth watching? Now I focus on the process: how do you get from point A to point B, from a scattered idea to something good?” In the newsroom, he was widely regarded as the best headline writer. Now, working with AI, he thinks about how to make automatically generated titles for every thread both effective and good: “What makes a good title?” The labor he used to sell was the output. The source of value he now creates is the process.