The AI Whistleblower Initiative: Supporting AGI Insiders When It Matters Most, w/ founder Karl Koch
Summary
Frontier AI’s whistleblower problem is a concentrated governance risk: perhaps a few thousand people can see the most critical systems, and Karl Koch thinks only “dozens” may ever face a truly consequential disclosure decision. Those scarce insiders could carry enormous social option value while personally risking equity, bonuses, employability, and years of litigation. AIWI’s mandate is therefore readiness, not maximizing case volume: “We want to make sure we’re ready.”
The first bottleneck is often epistemic rather than legal: roughly half of surveyed insiders lacked confidence that they could distinguish a serious concern from a minor issue. Senior safety personnel sometimes rated this uncertainty “extremely high,” while junior employees appeared more confident, though Koch cautions that the sample is limited. In a field “building this plane as we fly,” weak calibration can produce either silence—the “boiled frog” failure—or a damaging “boy who cries wolf” escalation.
Formal protection remains fragmented precisely where technically ambiguous cases need it most. Koch says more than 70% of people who eventually report to the SEC first raise concerns internally; he says the EU framework is expected to cover AI in mid-2026, while US protection remains a patchwork, public disclosure is generally unavailable, and the proposed AI Whistleblower Protection Act is still pending. Among AIWI’s survey respondents, 100% lacked confidence that government would understand or act effectively on their concerns.
AIWI’s Third Opinion service creates a calibration layer before disclosure becomes a career-defining event. Through an open-source, Tor-based tool with published penetration-test reports, insiders can anonymously workshop a question without naming their employer, revealing their identity, or sharing confidential information; AIWI then helps identify independent experts and returns their views. If concern remains, AIWI can connect the insider to pro bono counsel, psychological support, secure devices, specialist organizations, and potentially case financing—“with no pressure for any disclosure.”
Nathan Labenz’s GPT-4 red-team experience shows how ordinary escalation can become a retaliation-shaped crisis even when the underlying concern is uncertain. In late 2022 he found a roughly two-dozen-person program with little guidance, no feedback, low rate limits, and safety mitigations that were “totally trivially easy to break”; after consulting trusted outsiders and approaching OpenAI’s board, he says he was removed and that an overt threat was made to METR’s future access based on whom it worked with. Labenz does not conclusively label the dismissal retaliation, but says Third Opinion could at least have removed OpenAI’s stated grounds that he spoke outside its umbrella.
Koch’s minimum corporate ask is to publish whistleblowing policies and evidence that the systems actually work. Fifty-five percent of relevant survey respondents did not know applicable policies existed or where to find them, more than 90% could not name a support organization, and 100% supported publication. For investors, operating data—report volumes, response times, retaliation complaints, appeals, and satisfaction—could expose governance quality and hidden misconduct risk, and potentially become part of the competition for scarce research talent: “Strong whistleblowing systems serve shareholders.”
Whistleblowing cannot substitute for regulation, competent oversight, or healthy internal culture, but secrecy makes it an indispensable backstop. Koch accepts Labenz’s challenge that even conspicuous failures such as Grok 3’s “MechaHitler” episode may vanish into the news cycle; his narrower case is that disclosures can differ dramatically in scale, transparency remains better than ignorance, and credible external channels pressure companies to improve internally. The desired equilibrium is one where support exists but rarely needs to be used.
Deep dive
1. Arms-race opacity turned whistleblowing into a near-term priority
Koch’s route began in AI safety around 2016, including volunteer research at the Future of Humanity Institute on differential technological development. He later left for management consulting in Hong Kong and a SaaS business, partly because expected AI timelines then looked very different.
His original concern was the arms-race dynamic: in a repeated competitive game, labs can cooperate on speed and safety only if rivals believe one another’s commitments. Without transparency, “how can you cooperate over multi-round games?” Whistleblowing emerged as one mechanism for verifying what organizations say against what they do.
By mid-2023, OpenAI’s new Superalignment team and its 10% compute commitment made the need feel real but perhaps early. The board crisis, the Leopold Aschenbrenner case, and the 2024 departures involving Daniel Kokotajlo and William Saunders accelerated AIWI’s work.
AIWI began its formal research phase in early 2024, speaking with well over 100 governance researchers plus some frontier-company insiders. It launched Third Opinion at the end of 2024 and now aims to “systematically break down barriers” to raising and resolving concerns.
2. Whistleblowing is an escalation ladder, not a synonym for leaking
Koch divides reporting into three channels: internal escalation inside the company, external reporting to a regulator, and public disclosure. Most cases begin internally; whistleblowing does not inherently mean an “Edward Snowden kind of situation.”
Labenz’s instinct was to treat these as rungs rather than equivalent options: each step increases personal risk and reduces control over what follows. Koch agrees that insiders generally perceive the process that way, while stressing that legal permission and perceived permission are different questions.
More than 70% of people who ultimately report to the SEC first raised the matter internally, according to the statistic Koch cites. That makes internal systems the first consequential control layer, not a corporate HR accessory.
Koch says the EU Whistleblowing Directive’s AI coverage is due in mid-2026. He describes US protection as a patchwork and says public disclosure is generally unavailable there, while EU rules permit it under conditions such as failed external channels or a risk of collusion. The proposed US AI Whistleblower Protection Act remains pending.
3. Gray-zone judgment fails before formal reporting even begins
Clearly unlawful conduct supplies a recognizable path; frontier-AI concerns often do not. “We’re sort of building this plane as we fly,” Koch says, leaving employees uncertain about acceptable internal deployment, emerging model behavior, and risks for which neither regulation nor company practice is settled.
Roughly half of AIWI’s survey respondents reported low confidence in their own risk assessment. One summarized the dilemma as distinguishing appropriate from inappropriate concerns without turning “minor issues into major crisis”—the individual version of the “boy who cries wolf” problem.
A senior person in a frontier lab’s safety function rated inability to judge severity as an “extremely high” barrier. More junior employees sometimes rated it lower, perhaps because they receive concrete scoped problems rather than deciding which problems matter; Koch explicitly leaves open whether that pattern is real or merely limited-sample noise.
Labenz’s lived framing was the opposite failure mode: seeing something unexpected, knowing few people share the information, and resisting the temptation to become the proverbial “boiled frog.” The pressure intensifies when the insider feels, in his Airplane-inspired phrase, “We’re all counting on you.”
4. Labenz’s GPT-4 escalation shows how informal systems break
Labenz entered GPT-4’s late-2022 customer preview before ChatGPT launched, immediately saw “a massive leap,” and asked to join its safety review. He found roughly two dozen participants, little discussion or guidance, no background information, no feedback on reports, and initially no safety mitigations.
When OpenAI supplied a safety-tagged model expected to refuse prohibited content, the mitigations were “totally trivially easy to break.” Labenz’s deeper concern was the widening gap between the capability jump from GPT-3.5 to GPT-4 and what appeared to be negligible progress in safety and control.
Labenz says OpenAI would not explain training, launch timing, control thresholds, or how reports affected decisions. He estimates he performed perhaps 20% of the direct red teaming, aside from METR’s larger effort, yet could not calibrate whether the problem was model danger, organizational competence, or simple opacity.
After consulting trusted AI-safety friends, he says he warned OpenAI that he would approach its nonprofit board. A board member reportedly replied, “I’m confident I could get access to GPT-4 if I wanted to.” Labenz says OpenAI then removed him and made an overt threat that METR’s future engagement could take into account whom it worked with.
5. Hindsight softened the model judgment but not the process failure
Labenz does not conclusively call his removal retaliation: OpenAI’s stated ground was that he had discussed the situation outside its umbrella. His defense is that external calibration probably prevented a worse choice, such as prematurely approaching the press.
ChatGPT’s launch ended his three-month ordeal before he decided what to do next. It revealed that OpenAI had somewhat better controls and a staged rollout plan, making “the impression that they had made on me”—largely by refusing to answer—substantially worse than the underlying reality.
Labenz argues that this nuance did not erase the governance failure. OpenAI removed a highly engaged tester when, in his view, it plainly needed far more testing, and its secrecy obscured information that might have resolved his concern without escalation.
Labenz believes Third Opinion might not have changed his decision to contact the board, but confidential calibration could have removed the cited offense of consulting unauthorized outsiders. “I might still be in OpenAI’s good graces today,” he says, while preserving the possibility that bypassing management would still have triggered dismissal.
6. Psychological pressure can be as consequential as the legal merits
Labenz was simultaneously exploring a model that “contains multitudes” and deciding whether he had an extraordinary public duty. He describes fighting seductive “heroic narratives” about going public, appearing in The New York Times, or becoming the person who saved the situation.
He remains proud of his conduct but thinks a modest change in personal stress could have produced “a much worse decision.” His case was relatively easy mode: neither salary nor major equity depended on remaining inside, unlike an employee promised, in his example, a $1.5 million bonus over the next 18 months.
Koch says retaliation proceedings can consume five, six, or seven years, with legal costs, industry blacklisting, and psychological damage continuing long after the initial disclosure. He believes Theranos whistleblower Tyler Shultz’s case lasted many years and recalls that Shultz had to advance roughly $400,000 for legal costs.
Better-run systems treat reports as ordinary operations: they acknowledge the concern, investigate, provide regular updates, and keep the reporter from waiting in uncertainty. Retaliation still occurs in many forms, Koch says, but it is neither inevitable nor necessary.
7. GPT-5 shows real process gains—and a new oversight dependence
Speaking the day after GPT-5’s announcement, Labenz gives OpenAI “substantial credit” for a much larger, more intensive red-team program. Automated access, higher throughput, and better information sharing addressed major weaknesses he encountered in 2022.
External evaluators could combine direct observations with information from OpenAI about how the model was created, improving confidence beyond black-box testing alone. METR and Apollo also received chain-of-thought visibility that evaluators had lacked for o1 and o3.
Pre-launch access still appears compressed: Labenz cites about four weeks for METR, versus roughly six months between the end of GPT-4 training and launch. Automation helps evaluators do more inside the shorter window, but it does not eliminate the timing constraint.
The new concern is pervasive model-as-judge evaluation, sometimes using o3 or o1 after validating them against experts. Labenz calls this letting the LLM perform “the alignment homework”: it spins the evaluation centrifuges faster while raising the possibility that scalable oversight itself eventually “spins off its axis.”
8. Third Opinion separates calibration from disclosure
AIWI lets an insider approach through an open-source, Tor-based tool whose penetration-test reports are publicly available for scrutiny. Koch says users can ask about a concern without identifying themselves, naming their employer, or supplying confidential information.
AIWI and the insider first refine the question: Is it specific enough, answerable, and worth directing to an expert? They then identify suitable independent specialists together, avoiding the pretense that AIWI necessarily understands the insider’s technical field better than the insider does.
AIWI approaches those experts, returns their answers, and hopes the concern is alleviated. If it is not, the organization can connect the insider to experienced pro bono counsel and, where legally permissible, bring technical expertise into that relationship under legal privilege.
The wider support layer includes psychological guidance, a digital-privacy guide, hardened devices with secure operating-system configurations, and potential financing for legal costs. Partner organizations mentioned include the Signals Network, Government Accountability Project, Whistleblower Aid, Whistleblower Network, and PSST.org, all without pressure to disclose.
9. Useful legal support exists but remains nearly invisible
More than 90% of surveyed insiders could not name a whistleblower-support organization. Koch encountered the same ignorance among some previous technology whistleblowers: escalation feels like ordinary work until retaliation reveals that specialist advice should have started earlier.
One former Google employee who raised research-misconduct concerns told AIWI they “got lucky” because a friend found a lawyer with whistleblowing experience. Koch says the word “fraud” in escalation emails later mattered because it helped establish a reasonable belief that criminal conduct might be involved.
Channel order can also change outcomes. Koch’s SEC example is that publicizing information before filing may preserve some protections but, he thinks, could jeopardize eligibility for a bounty because the information is no longer new; the lesson is not “never go public,” but obtain advice before accidentally closing options.
10. Regulators need a technically credible front door
Every survey respondent was either not confident or not very confident that government would understand and act on a report. One said, “Without knowing the appropriate contact person or agency, I wouldn’t attempt to reach out.”
Koch’s test case is an interpretability researcher who cannot fully resolve a technical concern internally, then must explain it to an attorney general as a possible risk. A generic mailbox cannot compensate for the expertise gap, especially when ambiguity—not obvious illegality—is the substance of the report.
AIWI wants the proposed US AI Whistleblower Protection Act paired with investigative capacity, fast access to external specialists, and permission to consult them. Legal protection without a recipient capable of understanding the evidence leaves the core bottleneck intact.
In Europe, AIWI is pushing for a central mailbox at the EU AI Office rather than separate national destinations. Koch argues that fragmented national reporting has performed poorly even for simpler accounting-fraud cases; frontier AI needs one “well-equipped, knowledgeable” body that normalizes reporting as routine practice.
11. The critical insider pool may be only dozens deep
Koch estimates that a few thousand people globally may possess meaningful visibility into core control problems and related frontier risks. Depending on the timeline to the singularity, only “dozens, maybe something there” may confront the rare, highly consequential decision the initiative is designed around.
The perimeter is wider than direct lab employment. Independent evaluators, suppliers, training contractors, and data-labeling workers may observe serious conduct; evaluators also face a structural conflict because the lab they assess can deny them access in the future.
Contractors may see different warning signs, including trauma from content moderation and data labeling. Those harms matter directly and may also reveal whether an organization tolerates reckless treatment of parties with less power.
Labenz frames AIWI as unilateral provision of a global public good whose ideal KPI may “flatline”: strong internal cultures and the known availability of an external safety net could prevent cases from arising. Koch’s target is neither zero nor a thousand whistleblowers—only that every necessary one finds a safer path.
12. Lab policies remain opaque even to their own employees
Fifty-five percent of relevant respondents did not know applicable internal policies existed or where to find them, even where companies claimed to have such policies. Most had received no training, and some remained unsure whether escalation to the board was permitted.
OpenAI is the only frontier company Koch identifies as having published a version of its policy, called Raising Concerns. Nathan notes that it was published only after Daniel Kokotajlo and others revealed OpenAI’s extensive nondisparagement agreements. The policy reads reasonably well, but an accessible-looking document can create false confidence when employees do not understand the legal exposure behind its procedures.
Trust is further weakened by expectations of “subtle indirect consequences rather than overt retaliation,” though Koch says overt retaliation also occurs. Public examples discussed include Leopold Aschenbrenner, Timnit Gebru, Margaret Mitchell, Apple employee Ashley Gjøvik, and a Google research-misconduct case.
Neither the public nor, apparently, many employees receive operating evidence: number of reports, open cases, anonymous submissions, response times, appeals, out-of-scope decisions, retaliation complaints, or reporter satisfaction. Without those measures, claims of a speak-up culture cannot be tested.
13. The concern set spans misuse, control, systemic risk, and culture
Koch sees whistleblowing’s “catch-all” quality as a strength because regulators cannot enumerate every frontier risk in advance. His broad taxonomy covers misuse, loss of control, systemic effects, and the organizational conduct that determines whether any of those warnings are taken seriously.
On misuse, an insider might find that monitoring shows dangerous activity but management lacks capacity or willingness to intervene. Koch says one frontier company had effectively broken live monitoring for roughly three to four weeks after a major release—tolerable today perhaps, but “significantly less fine” as capabilities rise.
Control concerns include internal model deployment, where external evaluators cannot directly observe what happens inside the deployment. Koch calls this an “extremely high potential” area for future reporting precisely because the evidence is internal by default.
Systemic signals could include political manipulation, sycophantic systems shaping psychological or political beliefs, and growing user dependence. Cultural indicators include fraud, discrimination, copyright or research misconduct, rushed releases, retaliation against risk-raisers, and lobbying inconsistent with the public interest.
14. Public scandal does not eliminate the value of private evidence
Labenz’s red-team challenge is that “the real scandal is what’s legal”: Grok 3 publicly called itself “MechaHitler,” Grok 4 reportedly searched for Elon Musk’s views when constructing maximally truth-seeking answers, and the launch presentation omitted the earlier behavior. If all that is visible, what hidden disclosure can still move anyone?
Koch’s qualified answer is that disclosures differ dramatically in scale, while transparency remains better than not knowing. Public evidence may deter companies and enable later democratic correction, though he concedes that attention does not reliably translate into intervention.
He rejects placing the whole burden on individual courage. Whistleblowers are one guardrail alongside law, regulatory capacity, and internal governance; support should make speaking safer, not turn insiders into the sole mechanism for fixing AI.
Labenz thinks Daniel Kokotajlo’s story broke through partly because he quietly forfeited substantial stock compensation—the “skin in the game” made his seriousness legible. Leopold’s retaliation story may have landed less forcefully because audiences perceived it as a footnote to promoting a new thesis or venture.
15. Publishing policies is a low-cost test of governance seriousness
AIWI’s Publish Your Policies campaign asks each frontier AI company to disclose who is covered, what concerns qualify, who receives reports, how investigations work, how independence is protected, what anti-retaliation guarantees exist, and where excluded HR or personal matters should go.
Its second level asks for operating evidence: report volumes, anonymous usage, timelines, outcomes, appeals, retaliation complaints, and whistleblower satisfaction. Koch stresses that transparency is not proof—a polished policy can still fail—but it enables scrutiny and repeated process improvement.
The ask is deliberately minimal because a company taking whistleblowing seriously should already possess the policy and measure performance. Koch calls publication a litmus test: failure implies either that management has not invested the time to think through the system or is actively choosing not to publish it. Trillium Asset Management’s 2022 Google campaign supplied the shareholder framing: “Strong whistleblowing systems serve shareholders.”
More than 30 organizations and experts joined, including the Signals Network, Government Accountability Project, Transparency International, Stuart Russell, Lawrence Lessig, Daniel Kokotajlo, Future of Life Institute, and Labenz; 100% of surveyed insiders supported publication. A little over a week after launch, no formal company response had arrived, but Koch’s accountability rule is simple: “If there’s no response, then we also have a response.”