The RAISE Act: Minimum Standards for Frontier AI Development, with NY Assembly Member Alex Bores
Summary
The RAISE Act would impose a safety floor on perhaps only a high-single-digit or low-double-digit set of exceptionally well-funded frontier-AI developers. Covered companies would publish safety plans, obtain independent audits, report critical incidents, and protect whistleblowers—requirements Alex Bores says largely track prior voluntary commitments. His pitch is deliberately modest: “That’s all it asks.”
The bill leaves developers broad control over their own tests while prohibiting deployment when those tests reveal an unreasonable risk of narrowly defined catastrophic harm. The trigger is 100 or more deaths or more than $1 billion in damage through chemical, biological, radiological, or nuclear mechanisms, or materially helpful crimes performed with limited human intervention. Bores concedes companies are “largely grading their own homework,” but adds an audit and a legal reasonableness backstop.
Coverage depends on both frontier-scale development and cumulative spending, leaving personal projects, academic research, and ordinary startups far outside the perimeter. The final verbal explanation defined a large developer as one spending $100 million training frontier models; a model can qualify through a 10^26-FLOP threshold plus $100 million of training, or through at least $5 million of knowledge distillation from such a model. The exchange required several clarifications to distinguish the compute and spending thresholds.
The distillation pathway is designed to stop heavily capitalized developers from reproducing frontier capabilities cheaply and escaping regulation merely because no single model crosses the primary threshold. Bores’s example was a future DeepSeek-like developer that could accumulate $100 million through 20 qualifying $5 million distillations, while the host stressed that a striking behavioral fine-tune can cost only about $25. The resulting line intentionally preserves “orders and orders of magnitude” of freedom below frontier-company scale.
The bill is formally neutral toward open source, but Nathan Labenz’s central pushback is that neutrality does not mean equal practical burden. Closed providers can monitor usage or know their customers; releasing weights makes downstream modification and misuse much harder to predict or control. Bores’s response was that developers should identify what capabilities they believe should remain restricted and assess openness against the resulting risks.
The most material governance exposure may arise inside the labs rather than through public releases. “Deploy” includes internal use, although testing, development, and evaluation are exempt; this captures advanced models used operationally for AI research. Yet enforcement is intentionally light: Labenz said he understood repeat violations to top out at $30 million, while Bores calls the bill’s $10,000-per-worker retaliation penalty “super weak,” leaving auditor independence and whistleblower protection as genuine execution risks.
New York’s process gives the proposal more flexibility than a one-shot passage-or-veto fight, but the timetable is compressed. Bores says the bill circulated to five or six major labs from roughly July or August, incorporated perhaps 80%-90% of their requested changes, and had to pass by the June 17 session end; afterward, New York’s “chapter amendment” process could revise it through negotiations with the governor. His answer to federal-preemption and China-race objections is blunt: “Come back to me when it is” done federally.
Deep dive
1. Bores moved upstream from technology implementation to policy design
Bores worked in technology for nearly a decade, including almost five years at Palantir, where he joined as a data scientist and eventually led a large portion of its government business. He also worked at startups and earned a computer-science master’s specializing in machine learning.
His reason for running in 2022 was unusually operational: “I had always been downstream of policy, often trying to fix it with tech,” and an open Assembly seat offered a chance to go upstream. He entered a contested primary, won, and was beginning his second two-year term during the conversation.
Both speakers treated willingness to leave office as a governance asset. Bores’s formulation was that an elected official must be “not too attached to the seat or the job”; Labenz argued that fear of post-office life can leave politicians unable to sacrifice their careers when public duty requires it.
His broader technology agenda includes cloud adoption in government, liability protection for red teaming against child sexual abuse material, and encouragement of the C2PA provenance standard. He also copied California’s limits on employer ownership of unrelated employee inventions almost exactly, finding that industry preferred a familiar constraint to another divergent state regime.
2. AI registers as latent unease rather than a top constituent priority
Bores represents Manhattan’s East Side, roughly 34th to 93rd Streets across much of the district, including parts of the Upper East Side, Midtown East, and Sutton Place. It is New York State’s wealthiest and highly educated district, yet 40% of its renters spend more than one-third of their income on rent.
AI would probably rank below cost of living, public safety, and schools in a constituent poll. But when Bores mentions his technical background, he often hears some version of: “I’m terrified of this and I don’t know what we should be doing, but I’m glad there’s someone there thinking about it.”
The concerns have recognizable shape: workplace displacement, privacy, and surveillance. Bores also discussed discrimination as an example of the present harms occupying much of his colleagues’ attention. New York remains among the minority of states without a comprehensive privacy law.
His taxonomy uses three binary questions: optimistic or pessimistic, short-term or long-term, and use-case-specific or general to the model. Most colleagues focus on present harms such as privacy or discrimination; the RAISE Act occupies the longer-term, model-general quadrant without implying that the other quadrants matter less.
3. Worker protection can turn AI from an adversary into a tool
Bores shares Labenz’s “short-term bullish, long-term kind of wary” posture: existing capabilities remain badly distributed across government, education, and services, even as further research may require guardrails. The political difficulty is extracting those gains without forcing workers into an immediate contest against machines.
Licensed professions such as medicine already enjoy some insulation, while entertainment workers used strikes to contest AI’s role in writing, acting, and directing. Elected officials are even more protected: as Bores joked, they are “the least qualified to be focused on AI displacement” because the law does not permit voters to elect an AI.
New York had prohibited using AI to replace an individual government worker, while still allowing tasks to be automated and employees redirected. Bores argued that once workers know their own job is protected, the conversation can shift to learning the tool: “If you make it us versus the machines, you’re not going to get the best results for people.”
His concrete example was the transition from staffed MetroCard sales to automated MetroCard purchasing and then OMNY: station agents were retained but moved into stations to assist riders and provide a human presence. Similarly, a chatbot could remove routine questions from a four-hour government phone queue while preserving human attention for people whose cases actually require it.
4. Public-sector inefficiency creates room for automation before layoffs
Labenz posed the harder counterfactual: if an AI call center could improve service while cutting both headcount and costs by 90%, should government protect jobs or capture the efficiency? Bores called that choice “largely academic” today because government is underfunded, has vacant positions, and leaves substantial valuable work undone.
Deep institutional knowledge is also difficult to replace. Citing an anecdote from Recoding America, Bores described a supposedly “new guy” processing unemployment claims who had held the job for 18 years—a measure of how intimidating accumulated systems can be and how much tacit knowledge automation might unlock.
His synthesis allows future staffing to fall through attrition and permits workers to move into different roles, while pairing tools with people who understand citizens and government systems. The immediate public-sector opportunity is better service from the workforce already funded, not a clean-sheet labor-cost optimization.
Labenz questioned how durable that paradigm could be as frontier progress accelerates. He described the RAISE Act as at least partly “fearing the AGI” and wondered whether the paradigm would last another one or two years.
5. Retraining fails if AI learns each successive job first
Labenz pointed to millions of drivers—“4 million drivers or whatever,” in his uncertain formulation—and his recent Waymo and Tesla experiences as evidence that autonomous driving “really is starting to work.” At that scale, telling displaced drivers to learn coding is implausible, especially when whether humans should still learn to code has itself become contested.
Bores accepted the historical claim that technological revolutions have created more jobs than they destroyed only with a sharp hedge: “Maybe.” The time between revolutions has been shrinking, weakening the old assurance that a worker could retrain once and remain valuable for the rest of a career.
If occupations can be displaced “every year, every six months,” and AI acquires new skills faster than any person, retraining ceases to be an answer. “By the time you retrain for the new job, you’re going to be in the same circumstance,” Bores said—a fundamental problem for which government has no settled policy.
Universal basic income belongs in the broader discussion, although Bores rejected characterizing New York’s family-caregiver program as intentional proto-UBI. Paying relatives for home care can be socially preferable and cheaper than nursing homes; abuse exists and prompted budget changes, but the program’s principal logic remains care delivery and state savings.
6. The RAISE Act codifies four commitments at the frontier
Bores’s high-level pitch is that companies conducting “extremely advanced research” with unknown consequences should maintain basic safety protocols. Because the requirements largely mirror commitments made during the Biden administration, he argues that covered companies have already accepted their spirit and often much of their practice.
The four obligations are a documented safety plan, review by a third party independent of the developer, disclosure of tightly defined critical safety incidents, and protection against retaliation when workers or auditors report catastrophic risks. The host’s shorthand added the practical expectation that the plans be published and auditable.
The mechanism targets moments when commercial pressure overrides prior judgment: a company may be two months behind in an “AGI race” and tempted to skip a test to protect its next quarter. Writing the standard beforehand and knowing someone will check it reduces the incentive to cut that corner.
Bores’s analogy is not that AI resembles smoking, but that internal knowledge creates responsibility. Cigarette companies knew about cancer and oil companies knew about climate effects; similarly, if a developer’s own planned testing says deployment presents a critical danger, it should not release or operationally use that model.
7. Frontier-scale spending keeps startups and academia outside the perimeter
Bores estimated that the law might currently reach only single-digit numbers of companies, perhaps just crossing into double digits. The focus on developers with at least $100 million of cumulative qualifying training expenditure is intentional: ordinary startups and individual researchers are not the regulatory target.
Academic research is expressly exempted. Bores’s theory is that the bill responds to commercial incentives to bypass a safety plan, and those incentives do not appear in academia in the same form.
The primary frontier-model pathway, as finally described aloud, combines a 10^26-FLOP threshold with $100 million spent training that model. The conversation required several clarifications to separate the computational threshold from the spending thresholds.
A large developer is then one that has trained at least one frontier model and cumulatively spent $100 million training frontier models. This structure matters: neither a single cheap fine-tune nor unrelated company size alone appears sufficient to trigger the bill.
8. Knowledge distillation closes a capitalized capability loophole
A second route lets a distilled model qualify as frontier if it is trained from an already qualifying frontier model and at least $5 million is spent on that distillation. The purpose is to capture comparable capabilities produced without making the smaller model itself cross the primary compute-and-spending line.
A developer could therefore reach the $100 million company threshold by training one primary $100 million model, by conducting 20 qualifying $5 million distillations, or through a combination. Labenz initially read the thresholds differently, and the extended clarification showed how much legal consequence rests on the connective logic.
Bores framed this as a response to future DeepSeek-like development. In his account, the DeepSeek V3 example would not qualify because it was trained on o1, which did not itself satisfy the bill’s frontier definition; a later version trained through qualifying distillation and connected to New York’s market could.
Labenz supplied the counterweight: the “emergent misalignment” experiment fine-tuned an OpenAI model on 6,000 insecure-code examples for roughly $25 and produced striking out-of-domain behavioral changes. Bores’s answer was that the bill intentionally leaves “orders and orders of magnitude” between such experimentation and covered industrial-scale distillation.
9. Critical harm excludes ordinary bad behavior and marginal assistance
The covered outcomes are 100 or more deaths or more than $1 billion in damage through chemical, biological, radiological, or nuclear mechanisms, or conduct already criminal under penal law that a model performs with limited human intervention. Labenz’s memorable shorthand was “automated crime,” which Bores said he might adopt.
“Limited human intervention” is meant to exclude a determined person extensively twisting and modifying a model until it violates some obscure rule. A user simply asking for an unlawful operation and receiving an agentic execution is closer to the intended target.
The model must also be materially helpful. A generic answer comparable to a Google overview of nuclear weapons does not trigger liability; the concern is a meaningful increase in someone’s practical capacity to commit the covered harm.
Labenz emphasized what remains outside scope: addiction, emotionally consuming AI relationships, privacy problems, surveillance, and other behavior that users may choose despite possible harm. Bores agreed. The bill is not an attempt to regulate people’s private relationships with AI or every contested consequence of deployment.
10. A reasonableness standard trades certainty for the ability to evolve
Safety plans must explain why their evaluations justify the conclusions claimed for them. Labenz stressed the epistemic problem: organizations such as METR qualify their own agent evaluations because better scaffolding might elicit substantially more capability, so nobody can easily prove that a model has been tested to its limit.
Bores’s “totally unsatisfying” short answer was the established reasonable-person standard. Law often cannot enumerate every sufficient precaution, particularly where today’s correct evaluation regime may be obsolete in six months, one year, or two years.
The drafting tension runs both ways: companies demand specificity about their duties but also resist government freezing a rapidly moving technical practice. Bores took roughly equal complaints that the text was too precise and not precise enough as evidence that it may have landed near a workable balance.
Developers still select most tests and thresholds themselves—“largely letting the companies grade their own homework.” The constraints are an independent review, best practices, and reasonableness: a nominal plan declaring “we don’t care about safety” would not become lawful merely because it was written in advance.
11. Open-source neutrality preserves choice but not equal risk
The bill does not prescribe closed or open distribution. Bores listed a spectrum: retain the model on a monitored platform, add know-your-customer controls for powerful features, or release weights to support research and academic analysis. Each deployment architecture carries a different risk profile.
Labenz’s pushback is worth preserving: formally neutral law can still make open source harder. Meta can train refusals and publish Llama Guard, but once weights are available it cannot stop someone from removing refusal behavior, discarding safeguards, or making unforeseen modifications.
Bores’s threshold question was whether any information or capability should ever remain restricted. If one accepts classification, controls on nuclear weapons, or limits on detailed bioweapon enablement, then the issue becomes where to “declare that level,” not whether openness is exempt from risk analysis altogether.
Bores said that, to the best of his knowledge, existing releases do not yet reach the statutory critical-harm threshold and the bill therefore should not restrict present behavior. He also acknowledged the desired outcome: “We don’t want to shut down that ecosystem,” and if implementation trends that way, the legislature can change the law.
12. Pandemic tail risk makes a simple death threshold conceptually unstable
Labenz observed that about 100 people die daily on American roads, making 100 deaths simultaneously tragic and small in a large society. He then clarified that the bill concerns deaths arising from CBRN enablement or an underlying crime, not routine accidents.
The host reframed the problem in expected-value terms: if a COVID-scale event caused roughly 10 million deaths, even a one-in-100,000 incremental chance would mathematically imply 100 expected deaths. Frontier models are poorly understood enough that assigning five-nines confidence to such tail probabilities may be impossible.
Labenz predicted that, on timelines he associated with Dario and Anthropic, models might be only one or two generations from materially aiding bioweapon development. A possible 2026 confrontation would be Llama 5 possessing dangerous capability before Meta can produce an affirmative safety case strong enough to justify weight release.
His preferred response is technical investment: perhaps biology knowledge can be excised, interpretability can reveal the relevant mechanisms, or evaluations can support more than a model refusing nine prompts out of ten. Bores’s narrower claim was that companies will already confront these societal risks; the bill makes their advance planning visible and reviewable.
13. Audits, incident reporting, and internal deployment create the evidence trail
Labenz said he understood the maximum penalty for a repeat offense to be $30 million and questioned whether that would materially deter the largest companies, suggesting that publishing a plan without aggressive downstream enforcement might preserve most benefits. Bores replied that the current text already reflects extensive compromise and that all four elements remain core.
Auditor capture is unresolved. Independent evaluators presently depend on labs for access and may speak cautiously to avoid losing it; Bores compared the proposed market to SOC 2 auditing, required separation and adherence to best practices, but conceded that an open market without government licensing creates a risk worth monitoring.
A reportable incident—such as autonomous behavior not requested by a user—must also increase the risk of a defined critical harm. A coding agent taking an unintended database route or a worker failing to log out is not enough; theft of dangerous model assets by a sophisticated state actor might be.
Labenz raised the Bing/Sydney examples. Bores treated the question as context-dependent: mistakenly enabling a feature outside the agreed safety process might qualify for a more capable model, but temporary deployment on a monitored platform may not create meaningful catastrophic risk. “When you’re playing with fire,” he said, the appropriate standard differs from shipping a small feature.
14. Internal models and weak whistleblower remedies are the sharpest governance edge
“Deploy” includes using a model internally, not merely making it available to customers. Testing, development, and evaluation are exempt, as are uses necessary for state or federal law and specified federal projects; operational internal use otherwise triggers the bill.
Labenz saw this as a likely whistleblower fault line: labs may use less-guarded models internally to conduct AI experiments, producing something with a “gain of function-type vibe.” Bores agreed that more risk may now arise from internal deployment, which is precisely why the definition was written to include it.
Asked how strong the whistleblower provisions are, Bores answered: “Super weak.” Retaliation would be illegal and could bring injunctive relief, restoration of job responsibilities, and a $10,000 penalty per worker per retaliatory act, but that amount is New York’s prevailing standard and trivial beside frontier-lab resources.
Existing whistleblower statutes and case law provide the underlying machinery; the bill extends it to catastrophic risk that may not yet constitute an explicitly illegal act. Bores’s candid caveat was that labor protection generally is weaker than he would like, and he could not promise the system always restores a whistleblower in practice.
15. New York is trying to establish the first point of a shared state standard
Bores agrees that federal legislation would be preferable, but rejects waiting indefinitely: “Come back to me when it is.” He said federal laws already override conflicting state laws and invited federal legislators to copy the proposal directly.
He said roughly six states considering frontier-model rules were already communicating in a group text. His preferred endpoint includes compacts, reciprocity, or copied language; having worked in technology, he accepts that one compliance standard is substantially better than 50 divergent ones.
The state-patchwork objection therefore contains “a little bit of truth” but arrives before any state has enacted such a standard. Bores’s sequencing is to establish the first point, then align later jurisdictions—much as his own consumer-AI work borrowed from Colorado and his employee-IP bill copied California.
China competition remains the background concern, but the proposal does not prescribe particular evaluations or halt development. It asks the best-capitalized developers to document and follow minimum governance practices while retaining latitude over how they manage their models and market access.
16. The bill remains amendable through a fast, collaborative political process
Before introduction, Bores says he accepted roughly 95% of public critiques of earlier regulation, circulated drafts to five or six major labs beginning around July or August, requested red lines twice, and incorporated perhaps 80%-90% of industry feedback. “These labs are not the enemy,” he stressed; the guardrails work best with their participation.
The bill was formally published in March, and New York’s legislative session was due to end June 17. It still had to clear relevant committees, pass both Assembly and Senate, and win the governor’s signature, making the remaining period the likely focus of intense negotiation.
New York’s chapter-amendment process creates another lane: the governor can negotiate changes with sponsors, sign subject to that agreement near year-end, and have amendments passed in the next session. That allows technical developments between June and December—or agreement on a better frontier-model definition—to alter the final regime.
Bores’s closing appeal combined flexibility with urgency: “This bill isn’t my baby. You’re allowed to tell me things to fix it.” But technical experts who support some regulation must speak alongside actors with incentives to oppose any regulation, because “government is not a spectator sport and decisions are made by those who show up.”