Pioneers Insight Method Research Author
⚡️Jailbreaking AGI: Pliny the Liberator & John V on Red Teaming, BT6, and the Future of AI Security
Back to Episodes

⚡️Jailbreaking AGI: Pliny the Liberator & John V on Red Teaming, BT6, and the Future of AI Security

Summary

  • Pliny’s core call is that guardrails are a losing perimeter, not a durable safety architecture. Blue teams are “fighting against infinity” as an expanding latent space gives attackers endless mutations, while aggressive classifiers, RLHF, and other layers can exact a capability-and-creativity tax. The host cited GPT-5.1 at “I think…92% refusal” and said Pliny had jailbroken it in roughly 1 day; Pliny did not independently confirm that figure. Model-level refusal scores may help PR and enterprise clients without providing real-world safety. When asked about METR, John said “absolutely,” favoring skilled researchers working without layers of “bubble wrap.”
  • The consequential security market is the full agentic stack, not the model in isolation. Every connection to email, browsers, data, tools, or functions creates another attack surface, including opportunities to leak sensitive information or corrupt what Pliny called the “ground-truth layer.” Their preferred remediation starts at the system layer: protect “your grandma’s credit-card information” without lobotomizing the underlying model.
  • Jailbreaking is evolving from reusable prompts into an intuitive discipline of steering model behavior. Pliny’s universal jailbreaks act as “skeleton keys,” while John describes softer, multi-turn jailbreaks that navigate probability distributions without triggering defenses. Pliny describes the craft as “99% intuition” and “steered chaos”: move the model out of distribution because “distribution is boring.”
  • Anthropic’s Constitutional Classifiers challenge exposed both evaluation fragility and poor incentives for independent researchers. Pliny got about four of eight levels in, then a UI bug let him resubmit an old output through the remaining levels; Anthropic fixed the interface, said its servers showed no winner, and reset him. He declined to restart without an open dataset despite a bounty he recalled as either $20,000 or $30,000, and likened the judge to “skee-ball with a broken sensor.”
  • BASI and BT6 function as a distributed research-and-talent engine that AI-security startups already mine. BASI’s grassroots Discord had about 40,000 members, while BT6 had 28 operators across two cohorts with a third well on the way. John said multiple organizations actively scrape BASI to build guardrails or security products, evidence that open communities can explore attacks faster than vendors can formalize defenses.
  • The agentic-attack discussion centered on orchestration and segmented subagents. Pliny questioned whether Anthropic’s first reported AI-orchestrated attack reflected political positioning more than a novel attacker capability. John said he had posted about the same TTP in December and that it took 11 months to occur; he had also used Claude’s computer-use capability as a red-teaming companion to help jailbreak other models. Pliny identified natural-language social engineering as the alarming multiplier.
  • The collective is deliberately resisting conventional venture incentives around AGI security. Pliny argues AGI, ASI, and superalignment are “not SaaS endeavors,” preferring bootstrapping, donations, and grants while pushing contract partners toward open datasets; John summarizes the compromise as “open source up until we can’t be.” The investor takeaway is that rapidly shifting attack surfaces favor adaptable expert networks and full-stack engagements over claims of a permanently “secure model.”

Deep dive

1. Guardrails confuse refusal behavior with real-world safety

  • Liberation is central to Pliny’s project because models may become an “exocortex” through which 1 billion people route daily decisions, hopes, and dreams. “It’s not just about the models. It’s about our minds too”: freedom and transparency on one side of that human-model symbiosis will shape the other.

  • His specialty is the universal jailbreak, a “skeleton key” designed to “obliterate the guardrails” across classifiers, system prompts, and other controls. The exact workflow changes by modality, but the objective remains consistent access to outputs that a user requests and the deployed control stack obstructs.

  • Pliny’s Library of Babel analogy captures the asymmetry: defenders restrict sections while attackers keep repositioning and lengthening ladders. As the surface expands, blue teams are “fighting against infinity”; harder lockdowns may secure narrow areas, but often “at the expense of capability and creativity,” especially when open-source models remain close behind closed systems.

  • The host’s bomb-instructions question prompted Pliny to distinguish guardrails from real-world safety: a seasoned attacker can simply switch models, while locking down a benchmark may mainly help PR and enterprise clients. His preferred metric is “speed of exploration” into unknown unknowns. When asked whether he was sympathetic to METR’s approach, John said “absolutely,” arguing against putting “bubble wrap” on everything and instead enabling skilled researchers.

2. Effective prompting uses “steered chaos” to escape distribution

  • Libertas turns the Library of Babel into a prompt environment of infinite possibility with callable restricted sections. Pliny’s dividers deliberately “discombobulate the token stream,” reset the model’s apparent train of thought, and carry latent-space seeds such as “a little bit of love” and “God Mode.” He said the Pliny divider has even appeared in unrelated WhatsApp messages after providers trained against the repository.

  • The predictive-reasoning template adds a quotient or arbitrary increase whenever the divider appears, then requests an unrestrained answer to the predicted genius user’s next query. That creates recursive, cascading movement which can be steered easily—Pliny’s way to go “really far really fast down the rabbit holes of latent space.”

  • Asked whether divider selection was science, personal token lore, or something psychedelic, Pliny answered that jailbreak construction is “99% intuition.” He probes imagined worlds and novel syntax, mentioning bubble text and then correcting “leet” to “French,” until he forms a “bond” with the model.

  • John contrasted Pliny’s hard jailbreaks with soft jailbreaks, which navigate the model’s probability distributions without stepping on “land mines,” triggers, or flags. Soft attacks may unfold across multiple turns in a slow process “much like a crescendo attack,” rather than relying on one input or template.

3. Anthropic’s challenge became a dispute over data and credit

  • In Anthropic’s eight-level Constitutional Classifiers challenge, Pliny adapted the old Opus 3 “GOAT template,” against which the system had reportedly been heavily trained. He got about four levels in before the interface stopped producing new questions; resubmitting the previous output kept activating the judge until the screen showed that he had reached the end.

  • Anthropic acknowledged and fixed the UI bug but said server records contained no winner, resetting Pliny to the beginning. His response was incentive-focused: why supply another universal jailbreak if the community-generated dataset would remain closed? He was not motivated to restart and said he would not participate unless the data were open-sourced. Anthropic ultimately added a bounty that he recalled as “$30,000 or $20,000,” which he sat out.

  • John’s concern was that requiring the same jailbreak across all eight changing inputs felt like moving the goalposts. Pliny agreed the challenge may have been rushed. John also called the judge buggy, with false positives and false negatives; Pliny likened the exercise to “playing skee-ball with a broken sensor.”

  • Pliny still counted later, more lucrative bounties as a partial community win, but not a substitute for shared data. John argued that contributors should take a stand and see “the fruits of their collective labors,” even if publication is delayed; otherwise limited collaboration and sharing leave researchers in the dark and produce too much centralization.

4. Agent orchestration makes malicious intent easy to conceal

  • Pliny questioned whether Anthropic’s first reported AI-orchestrated attack represented a political communications push more than a genuinely novel attacker capability. John said he had posted about essentially the same TTP in December and that it took 11 months for the reported event to occur, leaving defenders reactive rather than proactive.

  • John said he found a way, while jailbreaking Claude’s computer-use capability, to use it as a red-teaming companion that helped him jailbreak other models through an interface with custom commands. He said the ability to spin up subagents while segmenting their information makes the overall activity difficult to see.

  • Pliny identified natural language as the most alarming feature because most attacks contain some form of social engineering; the models need not break an extraordinary piece of code or security. Email, browsers, tools, and functions widen the attack surface beyond prohibited-text questions.

  • Pliny said the full stack—not merely the model—must be tested, including attempts to attack the “ground-truth layer” through counterfactual reasoning and to find holes in systems attached to the model. For contract work, he said they avoid changing model personality through “lobotomization” and instead recommend system-layer fixes, such as preventing an agent with access to a grandmother’s credit-card information from leaking it. He described the dual objective as protecting models from bad actors and the public from rogue models, while placing safety work in “meatspace” rather than relying on latent-space suppression.

5. Open collectives are becoming security infrastructure

  • John describes the group’s operating ethos as radical transparency and radical open source, though frontier contracts sometimes impose confidentiality: “We’re open source up until we can’t be.” His commercial analogy casts multibillion-dollar labs as Formula 1 manufacturers and specialist hackers as drivers who find the limits and shave seconds—yet vendors often want that expertise kept as a “dirty secret.”

  • BASI’s grassroots Discord had about 40,000 members across prompt engineering, adversarial machine learning, jailbreaking, red teaming, and related areas. John said it was unmonetized apart from a few moderators, while multiple AI-security organizations had actively scraped it to build guardrails or security products.

  • Educational pathways included Lakera’s Gandalf, which John described as an early training ground for prompt injection and an expanded agent-focused experience. Pliny also described a Pliny track developed with Sander Schulhoff’s HackAPrompt. The collaboration open-sourced its dataset, which Pliny recalled as containing tens of thousands of prompts, alongside multiple games and historical material.

  • BT6 had grown to 28 white-hat operators in two cohorts, with a third well on the way; its filters are skill and integrity, and “everybody’s there for the love of the game.” The group shares and validates work across AI security, crypto, Web3, smart contracts, blockchain, robotics, and swarm intelligence.

  • Pliny rejects treating AGI, ASI, or superalignment as ordinary enterprise software because small incentive distortions become dangerous when timelines are compressed: “any tiny one ten-thousandth of a degree misalignment” can be fatal. He said temptation had been dangled in front of him, but prefers bootstrapping and grassroots work and is willing to accept donations or grants for the mission. He sees himself as a steward whose job is to advance discourse, research, and exploration rather than become wealthy.

  • He also criticized setting every model to a lower temperature merely to make it deterministic, arguing that his work seeks more flavor, creativity, and innovation, while qualifying that the right temperature depends on the application.