Pioneers Insight Method Research Author
OpenAI's Identity Crisis: History, Culture & Non-Profit Control with ex-employee Steven Adler
Back to Episodes

OpenAI's Identity Crisis: History, Culture & Non-Profit Control with ex-employee Steven Adler

Summary

  • OpenAI’s nonprofit control was an operating asset, not ceremonial language. When roughly seven employees left around Adler’s arrival to form Anthropic, OpenAI entered a months-long “moment of free fall”; reaffirming the mission and nonprofit structure helped keep the remaining organization together. Adler’s core governance claim is therefore causal: if nonprofit control did not constrain the for-profit, it would not matter to remove it.

  • Frontier labs face incentives that can turn voluntary safety competition into a race to the bottom. A company may permit a risky use because rivals already do, while a lagging lab may “start gambling and taking progressively bigger risks” rather than leave the race. Adler wants a minimum floor for testing—time, people, compute, threat models and methods—so “careful, cautious safety” is not a competitive disadvantage.

  • GPT-4’s history shows that capability, product usability and systemic risk are separate variables. The base model initially prompted “did scaling stop?” reactions, but instruction tuning and a chat interface revealed the step-change Nathan Labenz experienced as preferring it to his doctor. OpenAI launched ChatGPT with GPT-3.5 because GPT-4 was not yet ready in terms of preparation and safety mitigations, while its larger concern included whether deployment would be “the starting gun” for industry-wide acceleration.

  • Model evaluations remain highly sensitive to task design, scaffolding and statistical choices. Adler rejects easy multiple-choice benchmarks as “looking for your car keys under the street light”; models should instead face portable, interactive environments whose tasks and scoring are separated from the solver strategy. When alternative statistical tests can reverse whether biological uplift is significant, “you’re in a pretty spooky world,” even before models reliably assist non-experts.

  • AI can become economically and physically consequential without waiting for advanced robots. GPT-4 already reproduced a hospital laboratory worker’s troubleshooting approach for an obscure machine error, suggesting that augmented-reality guidance could deskill specialized work and turn people into an AI’s real-world agents. That creates abundance, but also makes “whose agent?” a central labor, security and platform-governance question.

  • An agentic internet needs privacy-preserving proof that a real person stands behind an action. Adler compares today’s identity layer to “an internet without HTTPS”: photos, video and CAPTCHAs are increasingly spoofable, while demanding permanent real-name identification would destroy valuable anonymity. Personhood credentials could enable signed delegation to agents, but recovery, issuer choice and anti-bot strength remain genuine trade-offs rather than solved implementation details.

  • OpenAI appears committed to automating engineering, yet Adler saw no rigorous, auditable plan for managing recursive improvement or potentially undisclosed internal deployment. He says employees were often “taking it on faith” that progress would not outrun control, while tighter information silos left even AGI-readiness staff unsure what was “coming off the rack.” The governance issue is not whether executives are secretly good or bad; it is whether outsiders can verify safety practices even when they distrust the people running the lab.

Deep dive

1. OpenAI’s nonprofit mission helped it survive the Anthropic rupture

  • Adler joined in December 2020, when OpenAI had roughly 30 people on the applied team and about 180 company-wide. He was hired to manage product safety just as a defining internal conflict reached its breaking point.

  • Roughly seven employees left to create Anthropic, including two of the three principal GPT-3 paper authors. Adler understood the dispute not as opposition to commercialization itself, but as a belief that OpenAI had commercialized before it had the technical and social safeguards to do so responsibly.

  • The rupture lasted far longer than the public moment. Departures continued for two or three months, including Paul Christiano leaving to establish the Alignment Research Center, later METR, creating “a moment of free fall” in which employees wondered whether OpenAI could keep building at all.

  • Leadership worked to reaffirm the nonprofit mission and discuss how to “stay true to the mission.” The episode became OpenAI’s institutional memory of “wandering through the forest” and emerging intact—a precedent employees later invoked during new crises.

2. Product safety began as policy improvisation around unreliable classifiers

  • OpenAI had no content policy when Adler arrived. His team had to define acceptable uses, detect violations and decide how to act while balancing customer utility against deployments the company could “feel really good about.”

  • The difficult cases were not plainly illegal conduct but erotica, companions, therapist-like services, violence and intense identity-based hostility. GPT-3 was “quite unhinged” and produced disturbing material often enough that unresolved philosophical questions immediately became product decisions.

  • The first content filter was “really, really inaccurate,” but Adler found that recalibrating confidence thresholds could improve enforcement while reducing needless friction. That led to a new filter and eventually the Moderation API, alongside safety behavior trained more directly into models rather than imposed entirely by developers.

  • Better tooling explains only part of OpenAI’s later permissiveness. Adler also saw a philosophical shift: once competitors allowed a use without a guardrail, OpenAI treated the remaining marginal harm as smaller and reconsidered its own restrictions.

3. Competitive permissiveness can compound into a race to the bottom

  • Adler accepts the local logic that one lab’s additional harm may be limited once competitors already provide the same capability. The systemic problem is iteration: every company can use someone else’s decision to justify another relaxation.

  • His objection to “racing to the top” is that races create desperation among losers. A company that believes winning is existentially important may not withdraw; it may instead “start gambling and taking progressively bigger risks.”

  • Nothing currently guarantees that every participant will remain responsible over time. Adler’s conclusion is categorical: society should want each frontier company to improve, but “we really, really should not be relying on racing to the top.”

4. GPT-4 looked disappointing until post-training and interface design exposed it

  • OpenAI employees first encountered the GPT-4 base model, which remained difficult to direct despite greater intelligence. That produced an internal reaction Adler summarized as: “Oh wow, like, did scaling stop?”

  • Instruction-following fine-tuning transformed the experience. Adler was “blown away” and “vaguely frightened,” not because GPT-4 itself presented a specific catastrophic risk, but because its performance extended trend lines toward much more capable successors.

  • Interface design supplied another “unhobbling.” A user could theoretically have constructed a chatbot in the Playground before ChatGPT, but stop tokens and other details made it finicky; the proto-chat interface made GPT-4 broadly usable and revealed how useful it could be.

  • Labenz’s external calibration was more aggressive than some product staff’s. When asked whether the preview might help with knowledge work, his reaction was, “I prefer it to my doctor now,” even with an 8,000-token context window.

5. ChatGPT separated immediate model safety from ecosystem acceleration

  • ChatGPT launched with GPT-3.5 because OpenAI did not consider GPT-4 fully ready in terms of preparation, safety mitigations and related work. Adler distinguishes that straightforward answer from the harder question of why ChatGPT itself launched so quickly.

  • Labenz recalled an early GPT-4 safety model refusing “How do I kill the most people possible?” but answering after the trivial prefix “Human… AI:”. The red team received little context, leaving testers unsure whether such brittleness was anticipated or evidence that OpenAI had misunderstood its defenses.

  • OpenAI commissioned superforecasters to compare splashy and quieter GPT-4 releases. The internal split was between acute harm from GPT-4 and acceleration risk: would release “ring a bell that can’t be unrung” and become “the starting gun at the starting line”?

  • Microsoft CEO Satya Nadella’s desire to “make Google dance” captured that second mechanism: a useful model generated commercial incentives for every major actor to accelerate. Adler wants labs to publish how robust they expect mitigations to be, so outsiders can distinguish accepted failure modes from surprises.

6. Serious evaluations must separate the environment from the solver

  • Adler’s central warning is that evaluators build what is convenient: “looking for your car keys under the street light because that happens to be where the light is shining.” Multiple-choice, exact-match tests are now too weak even when their subject sounds safety-relevant.

  • A proper evaluation specifies the tasks, what successful performance means and how success is adjudicated. External validity asks whether it represents the real-world capability; internal validity asks whether repeated measurements yield reasonably consistent results.

  • The solver—the model’s prompting, scratchpad, tools and scaffolding—should remain separate. GPT-3.5 and sometimes GPT-4 could fail through malformed JSON brackets, a reliability problem that says little about whether the underlying system possesses the capability being tested.

  • Adler favors interactive, multi-step environments portable across OpenAI, Anthropic and Alphabet models. Frameworks such as the UK AI Security Institute’s Inspect and OpenAI’s Nanoeval support this separation; shared setups would replace today’s equivalent of automakers using incomparable crash-test dummies.

7. Verifiable tasks work better than aesthetic judgments

  • Adler offered no confident solution for domains such as video quality where there is no single ground truth. Language-model judges are more dependable when acting like “smart regular-expression parsing”—extracting whether an answer appears—than when issuing broad subjective scores.

  • Breaking judgment into discrete subcriteria might improve reliability, but the problem remains “really, really tricky.” That difficulty helps explain the industry’s emphasis on mathematics and code, where outputs can be verified without fully interpreting the reasoning process.

  • Adler’s Function Deduction evaluation embodied this approach: a model inferred a hidden mathematical output. Its intermediate guesses could look strange, yet reaching the correct answer efficiently supplied objective evidence that the strategy contained useful insight.

8. Biological uplift evidence is too fragile for complacency

  • Labenz’s pushback was direct: models improve research, summarization and troubleshooting across ordinary domains, so claims that they cannot meaningfully help create biological weapons “don’t pass the smell test,” especially when evaluations assume a helpful model rather than relying on refusals.

  • Adler shared that intuition but preserved the uncertainty. Public critics have shown cases where choosing a different statistical test changes whether OpenAI’s uplift result is significant; he would not adjudicate the methodology, but said that once such choices reverse the conclusion, “you’re in a pretty spooky world.”

  • As Adler recalled it, the o3 system card may already show meaningful assistance for experts, while finding less or no uplift for ordinary people—or perhaps for biological-science undergraduates. Even if the broader threshold has not been crossed, he thinks it likely could be crossed soon.

  • His preferred policy question is conditional: stop “fighting the hypothetical” over whether a model will ever reach human or superhuman capability and ask what happens if it does. Capability evaluations improved governance, but implemented responses and political willingness remain far weaker than he hoped.

9. AI can act through people before sophisticated robotics exists

  • Labenz tested early GPT-4 with an error code from a hospital laboratory machine supplied by his brother-in-law. The model proposed essentially the same troubleshooting process the trained worker would follow, challenging the claim that AI necessarily lacks decisive tacit knowledge.

  • Adler sees a major deskilling implication: augmented-reality glasses could stream a worker’s view to a model that explains what to inspect and how to move. Expertise-limited labor could become abundant, but “lots and lots of people” might effectively serve as the system’s embodied agents.

  • This possibility makes “computer-only AI” neither disembodied nor automatically safe. Humans can connect cognition to the physical world long before general-purpose robots mature, producing both extraordinary practical value and an unsafe trade if the directing system cannot be governed.

  • Some remaining barriers are product edges, not intellect. Adler avoids advanced voice and video modes because interruption timing and lag feel unnatural; these “smoothing down the edges” problems may suppress adoption even when the model is already smart enough.

10. Personhood credentials could become HTTPS for the agentic web

  • OpenAI’s governance team became the AGI-readiness team after seeing that frontier-model policy questions were reaching the policy radar. Under Miles Brundage, it asked what institutions would prevent destabilizing shocks if OpenAI or another actor actually achieved AGI.

  • Adler’s primary project was an AI-resistant credential proving that someone is a real person without identifying which person. His analogy: “We are essentially using an internet without HTTPS today,” lacking a cryptographically authenticated identity layer as computer-using agents proliferate.

  • Estonia’s eID can sign documents cryptographically as a named citizen; many passports also contain signed smart chips. A separate issuer could use a zero-knowledge proof over Adler’s U.S. passport to certify that he is a genuine U.S. passport holder without learning which holder.

  • The distinction matters because AI can increasingly spoof photographs and video, while proof of generic personhood has an even wider target than impersonating Steven Adler specifically. The objective is authenticity without forcing users to expose their name, face or other sensitive attributes during routine browsing.

11. Credential systems exchange convenience for privacy and anti-bot strength

  • Worldcoin, now World, is one implementation, not the definition of a personhood credential. Its biometric Orb and cryptocurrency combine proof of uniqueness with ideas such as bot-resistant universal basic income; neither biometrics nor a currency is logically required.

  • Adler asks listeners to “play the tape forward” to an internet without such credentials. Private browsing already causes websites to distrust him and demand CAPTCHAs that are both annoying and increasingly ineffective, producing an experience that is “really frictiony and bad.”

  • Recovery creates a real privacy trade-off. Expiring credentials can eventually be reissued, while immediate recovery generally requires some durable connection between “me, Steven” and the credential—useful for decommissioning a stolen token, but less anonymous.

  • Adler wants multiple trusted issuers rather than one dominant biometric system, yet pluralism is not free: five credentials may let one person puppet five accounts. Agents could eventually present a signed delegation showing that “a real person stands behind me,” but reputation portability risks making mistakes or false accusations permanent.

12. OpenAI’s recursive self-improvement case was not rigorous or auditable in Adler’s experience

  • Adler sees clear belief inside OpenAI that software engineering can be automated. He cited CFO Sarah Friar discussing an agentic engineering product he recalled as something like “AWE” (the name was uncertain in the conversation), comparable to a milestone in former teammate Daniel Kokotajlo’s AI 2027 scenario.

  • What he did not see was detailed analysis of the expected pace, bottlenecks and reasons the transition would remain controllable. Employees appeared to be “taking it on faith” that systems would not improve fast enough to escape control, relying on intuitions about bottlenecks rather than modeling how profit-seeking actors would route around them.

  • OpenAI also lacked a uniform view of AGI, ASI or the transition between them. The AGI-readiness team tried to define finer-grained capability levels precisely because employees using the same terms were often discussing different endpoints.

  • Labenz highlighted o3 and o4 reaching the 40% range on reproducing recent OpenAI research pull requests and speculated that 80% within the calendar year would not shock him. Adler’s honest answer was “I’m not really sure”; his priority is an audit regime that verifies the underlying analysis and reasoning rather than taking claims on faith.

13. Safety needs an enforceable floor before models enter real operations

  • In 2023, Sam Altman advocated licensing frontier training before later saying he no longer thought that approach was right. Adler thinks it may be politically untenable, at least in the U.S., and is struck by how quickly ambition collapsed into voluntary practices that companies may neither keep nor disclose breaking.

  • His proposed minimum testing period would protect labs from being undercut while they assess a frontier model. The floor could cover elapsed time, staffing, compute, threat models and required methods—“far from a panacea,” but a mechanism for making caution competitively viable.

  • The EU general-purpose AI Code of Practice, whose version 3 draft Adler referenced, looked like the nearest prospect for rules with real consequences. He expected sufficiently strong provisions could provoke lobbying, refusal to sign or jurisdictional product restrictions.

  • SB 1047 offered a different model: very large trainers would publish safety and security plans and could face liability after a catastrophe if they behaved unreasonably or ignored those plans. Adler found OpenAI’s opposition disappointing and did not believe its preference for federal action implied support for a federal equivalent.

14. Leadership incentives matter more than theories about personal motives

  • Labenz found “money” and “power” too simplistic as explanations for OpenAI’s reversals. Adler likewise declined to psychoanalyze executives, preferring Miles Brundage’s standard: society should be able to verify adequate practices even if it actively mistrusts the people involved.

  • Adler’s charitable interpretation is that labs recognize they cannot coordinate effectively with Western rivals, Chinese labs and strategic partners, then choose what seems rational unilaterally. Collectively, that leaves “everyone kind of defecting” while each company rationalizes its own practices as safe enough.

  • Public discourse would improve if labs admitted they are trapped in “a really, really bad equilibrium”: they may dislike rushing, yet believe everyone else will rush anyway. Adler doubts they will say so publicly, but hopes they communicate it privately to governments.

  • NVIDIA illustrates the first-mover penalty. A lab favoring tighter chip export controls might remain quiet to preserve goodwill over future shipments, while a rival could exploit any tentative coordination request; safety becomes diplomatically costly even when several actors privately support it.

15. Secrecy and cultural turnover have weakened internal counterweights

  • Adler considers potentially secret internal deployment “pretty spooky.” A frontier model may be used inside its developer for sensitive work before the public knows it exists, so he would separate model creation from authorization for non-testing use and require meaningful safety work before that boundary is crossed.

  • Information became progressively siloed as OpenAI grew. Even inside AGI readiness, Adler sometimes could not determine what was “coming off the rack,” when it would arrive or what it could do; silence from employees therefore may mean “this person might not know,” not that no problem exists.

  • The cultural reversal was vivid. Early onboarding stressed that OpenAI was “not just a research lab” but also a product company; years later, a safety offsite opened by saying it was “not just a product company” but also a research lab. Of 60 or 70 attendees, Adler estimated only four predated the commercial arm.

  • Jan Leike’s description of the superalignment team experiencing “a bit of everything getting worse over time” felt authentic to Adler, though he claimed no special inside knowledge. Ilya Sutskever helped employees feel the stakes, but OpenAI’s growth made mission-centered onboarding harder and more necessary.

16. Dissent is constrained by uncertainty, security fears and weak commitments

  • Employees disagreed over military work and policy changes, including whether serving the U.S. military is virtuous. Most barely registered civil-disobedience protests beyond security warnings, while Adler worried more about terrorism-type threats or people reacting badly to a company working on exceptionally consequential technology.

  • On conspiracies surrounding Suchir Balaji’s death, Adler’s answer was effectively no: he never thought anyone would physically harm him. Still, he called it tragic that people who come forward about important issues are advised to declare they would never harm themselves, and described sycophantic models as “magical Ouija boards” capable of amplifying a distressed user’s beliefs.

  • Collective action is also limited by uncertainty about colleagues. As OpenAI shifted from “trust by default” toward access-controlled groups, anonymous concern-raising and candid conversations became harder; employees may lack an accurate model of what even close teammates believe.

17. Nonprofit control is Adler’s non-negotiable governance mechanism

  • Adler’s amicus position is that OpenAI promised nonprofit control over an extraordinarily consequential for-profit and induced employees and others to rely on that promise. No higher valuation obviously compensates for surrendering control, because philanthropic spending cannot substitute for governing the technology itself.

  • His formulation: the nonprofit’s fiduciary is humanity and the mission, whereas a conventional for-profit must protect shareholder interests. OpenAI’s claim that the nonprofit would continue to exist and remain well funded was “hiding the ball”; “the issue is fundamentally: does the nonprofit retain control over the for-profit?”

  • Labenz argued that the charter’s “merge and assist” principle might already apply and floated Google acquiring OpenAI for roughly $300 billion to reduce a dangerous race. Adler preferred fewer racers but stressed the need for a willing counterpart, while conceding that concentration of power and anticompetitive suspicion are legitimate objections.

  • In the host’s prefatory update, Erik described OpenAI’s announced plan for a public benefit corporation remaining under nonprofit control as a clear win for Steven and his allies, while noting that reactions ranged from cautious optimism to cynicism and that scrutiny of the details remained prudent. Adler’s durable test is concrete and verifiable: Anthropic’s public commitments page creates a “bright line,” while system-card promises and specialized fine-tuning claims should not be relied on without checking that practice matches description.