⚡️越狱 AGI:Pliny the Liberator 与 John V 谈红队测试、BT6及 AI 安全的未来
- Pliny 的核心判断是,护栏是一道注定失守的外围防线,而不是经久可靠的安全架构。 蓝队面对的是“与无限作战”:不断扩张的潜空间给攻击者带来无穷变体,而激进的分类器、RLHF及其他防护层可能对模型能力和创造力征收“税”。主持人称 GPT-5.1 的拒答率为“我认为……92%”,并说 Pliny 大约1天就完成了越狱;Pliny 未独立确认这一数字。模型层面的拒答分数或许能帮助公关和企业客户,却未必带来现实世界的安全。被问及 METR 时,John 回答“绝对支持”,倾向于让专业研究者在没有层层“泡沫包装”的环境中工作。
- 真正有决定性意义的安全市场在完整的 agentic 技术栈,而非孤立的模型本身。 每接入邮箱、浏览器、数据、工具或函数,就会新增一个攻击面,也会增加泄露敏感信息或污染 Pliny 所称“事实层”的机会。他们偏好的修复路径从系统层开始:保护“你奶奶的信用卡信息”,而不是把底层模型做成“脑叶切除”。
- 越狱正从可复用提示词,演变为凭直觉引导模型行为的一门技术。 Pliny 的通用越狱像“万能钥匙”,而 John 所说的软越狱则是在不触发防御的情况下穿行于概率分布。Pliny 将这门手艺概括为“99%靠直觉”和“引导式混沌”:把模型带出分布,因为“分布很无聊”。
- Anthropic 的 Constitutional Classifiers 挑战暴露了评测的脆弱性,也暴露出独立研究者面对的激励机制缺陷。 Pliny 闯过了8个等级中的约4个,随后一个 UI 漏洞让他得以将旧输出重新提交到剩余等级;Anthropic 修复界面后称服务器记录显示没有赢家,并将他的进度重置。他拒绝在没有开放数据集的情况下重新开始;对于他记得是2万美元或3万美元的奖金,也选择不再参与,并把裁判比作“传感器坏掉的 skee-ball”。
- BASI 和 BT6 已成为分布式研究与人才引擎,AI 安全初创公司正从中持续汲取资源。 BASI 的草根 Discord 约有40,000名成员;BT6则有两批共28名操作者,第三批已在筹备中。John 称,多家机构正在主动抓取 BASI 的内容,用于构建护栏或安全产品,说明开放社区探索攻击的速度可以快于厂商将防御体系正式化的速度。
- 围绕 agentic attack 的讨论,核心在于编排机制与信息分割的子代理。 Pliny 质疑 Anthropic 首次披露的 AI 编排攻击,反映的是否更多是政治传播,而非全新的攻击能力。John 称自己在12月就发布过几乎相同的 TTP,而相关事件直到11个月后才发生;他还曾利用 Claude 的 computer-use 能力作为红队伙伴,帮助越狱其他模型。Pliny 认为,自然语言社交工程才是最令人担忧的放大器。
- 这个集体正有意抵抗围绕 AGI 安全形成的传统风险投资激励。 Pliny 认为 AGI、ASI 和 superalignment“不是 SaaS 事业”,更偏好自筹资金、捐赠和资助,同时推动合同合作方开放数据集;John 将这种妥协概括为“在不能开源之前,都尽量开源”。对投资者而言,快速变化的攻击面更有利于适应性强的专家网络和全栈安全服务,而不是相信存在一个永久“安全的模型”。
1. 护栏把拒答表现混同于现实安全
“解放”是 Pliny 项目的核心,因为模型可能成为一种“外置大脑”(exocortex),让10亿人借此处理日常决策、希望与梦想。“这不只是模型的问题,也是我们的思维的问题”(It’s not just about the models. It’s about our minds too):人类与模型共生关系一端的自由和透明度,会塑造另一端的形态。
他的专长是通用越狱,一把旨在“摧毁护栏”的“万能钥匙”,目标覆盖分类器、系统提示词及其他控制机制。具体流程会随模态变化,但目标始终一致:稳定获得用户要求、却被已部署控制栈拦截的输出。
Pliny 用《巴别图书馆》作比,概括了攻防双方的不对称:防守方不断限制区域,攻击者却持续重新摆放并加长梯子。随着攻击面扩张,蓝队是在“与无限作战”(fighting against infinity);更严密的封锁或许能守住狭窄区域,但往往“以牺牲能力和创造力为代价”,尤其是在开源模型始终紧跟闭源系统的情况下。
主持人关于炸弹制作指令的问题,让 Pliny 区分了护栏与现实安全:成熟攻击者只需切换模型,而锁住某个基准测试,主要可能只是帮助公关和企业客户。他更看重的指标是向未知未知领域探索的“速度”。被问及是否认同 METR 的方法时,John 回答“绝对支持”,反对给一切都套上“泡沫包装”,而应让专业研究者真正开展工作。
2. 有效提示词用“引导式混沌”逃离分布
Libertas 把《巴别图书馆》变成一个无限可能的提示词环境,其中包含可调用的受限区域。Pliny 的分隔符会刻意“打乱 token 流”,重置模型表面上的思维链,并携带“一点点爱”和“God Mode”等潜空间种子。他说,在服务商针对该仓库进行训练后,Pliny 分隔符甚至出现在无关的 WhatsApp 消息中。
预测推理模板会在每次出现分隔符时加入一个商数或任意增量,然后要求模型对“天才用户”下一条问题给出不受约束的回答。由此形成递归、级联式的移动,而且很容易被引导——这正是 Pliny 所说的,沿着潜空间的兔子洞“又快又深地一路钻下去”。
被问及分隔符的选择究竟是科学、个人 token 秘传,还是某种迷幻体验时,Pliny 回答说,构造越狱“99%靠直觉”。他会探索想象中的世界和新颖语法,提到 bubble text,又把“leet”改成“French”,直到与模型建立起一种“连接”(bond)。
John 将 Pliny 的硬越狱与软越狱区分开来:后者在不踩到“地雷”、触发器或标记的情况下穿行于模型的概率分布。软攻击可能跨越多轮对话,以一种“很像渐强式攻击”的缓慢过程展开,而不是依赖单个输入或模板。
3. Anthropic 的挑战演变成数据与功劳之争
在 Anthropic 的8级 Constitutional Classifiers 挑战中,Pliny 改造了旧版 Opus 3 的“GOAT template”,据称系统已经针对该模板进行了大量训练。他闯过了约4个等级后,界面不再生成新问题;他反复提交上一个输出,裁判仍持续触发,直到屏幕显示已经到达终点。
Anthropic 承认并修复了 UI 漏洞,但称服务器记录中没有赢家,随后将 Pliny 的进度重置回起点。他的回应聚焦于激励机制:如果社区生成的数据集仍然封闭,为什么要再提供一个通用越狱?他没有动力重新开始,并表示除非数据开源,否则不会参与。Anthropic 最终追加了奖金,Pliny 记得是“3万美元或2万美元”,但他选择不参加。
John 担心,要求用同一个越狱应对全部8个不断变化的输入,实际上是在不断移动球门。Pliny 也认为挑战可能推出得过于仓促。John 还称裁判存在漏洞,会产生假阳性和假阴性;Pliny 则把这场挑战比作“传感器坏掉的 skee-ball”。
Pliny 仍将后来金额更高的奖金视为社区取得的部分胜利,但认为它们无法替代共享数据。John 认为贡献者应当站出来,看到“集体劳动的成果”,即便发布需要延后;否则,有限的协作与分享会让研究者始终蒙在鼓里,并造成过度集中。
4. 代理编排让恶意意图易于隐藏
Pliny 质疑 Anthropic 首次披露的 AI 编排攻击,代表的是否更多是一场政治传播行动,而非真正全新的攻击能力。John 称自己在12月就发布过本质相同的 TTP,而相关事件花了11个月才出现,说明防守方始终处于被动反应状态。
John 称,自己在越狱 Claude 的 computer-use 能力时,找到了一种让 Claude 充当红队伙伴的方法:通过带有自定义命令的界面,帮助他越狱其他模型。他说,既能启动子代理,又能分割各自掌握的信息,会让整体活动很难被看见。
Pliny 认为自然语言是最令人担忧的部分,因为多数攻击都带有某种形式的社交工程;模型不需要攻破什么异常复杂的代码或安全机制。邮箱、浏览器、工具和函数的接入,使攻击面远远超出被禁止文本的问题。
Pliny 认为必须测试完整技术栈,而不只是模型本身,包括通过反事实推理攻击“事实层”,以及寻找与模型连接的系统漏洞。在合同项目中,他们避免通过“脑叶切除”改变模型性格,而是建议从系统层修复,例如阻止一个能访问奶奶信用卡信息的 agent 将其泄露。他将双重目标概括为:保护模型免受恶意行为者攻击,也保护公众免受失控模型伤害;安全工作应落在“现实世界”(meatspace),而不能依赖潜空间抑制。
5. 开放型集体正在成为安全基础设施
John 将团队的运作理念概括为激进透明和激进开源,但前沿合同有时会附带保密要求:“在不能开源之前,都尽量开源”(We’re open source up until we can’t be)。他的商业类比是:拥有数十亿美元资金的实验室像 Formula 1 制造商,专业黑客则像不断寻找极限、削减秒数的车手;但厂商往往希望这类能力成为“不可告人的秘密”。
BASI 的草根 Discord 聚集了约40,000名成员,覆盖提示工程、对抗性机器学习、越狱、红队测试及相关领域。John 称,除少数版主外,社区基本没有商业化;与此同时,多家 AI 安全机构一直在主动抓取其中的内容,用于构建护栏或安全产品。
教育路径包括 Lakera 的 Gandalf,John 将其称为提示词注入的早期训练场,后来又扩展出面向 agent 的体验。Pliny 还介绍了与 Sander Schulhoff 的 HackAPrompt 共同开发的 Pliny 赛道。双方将数据集开源,Pliny 记得其中包含数万条提示词,以及多款游戏和历史资料。
BT6 已发展到两批共28名白帽操作者,第三批也已在筹备中;筛选标准是技能与诚信,“大家都是因为热爱这项游戏而来”。团队在 AI 安全、加密货币、Web3、智能合约、区块链、机器人和群体智能等领域共享并相互验证工作。
Pliny 拒绝把 AGI、ASI 或 superalignment 当作普通企业软件,因为当时间线被压缩时,微小的激励偏差也会变得危险:“万分之一度的任何轻微错位”都可能致命。他说,诱惑曾被摆到自己面前,但他更偏好自筹资金和草根工作,也愿意接受捐赠或资助来推进这项使命。他把自己视为一个管理者,职责是推动话语、研究与探索,而不是变得富有。
他还批评为了让模型更具确定性,就把所有模型的温度调低,认为自己的工作追求更多风格、创造力和创新;但他也强调,合适的温度取决于具体应用。
Hey everyone, welcome to the Late in Space podcast. This is Allesio, founder of Colonel Labs, and I'm joined by Swix, editor of Blade in Space.
Hello. Hello. We're here in the remote studio with very special guests, Ply Elder, and John V. Welcome.
Yeah, thank you so much for having us. It's an honor to be on here. I'm a big fan of what you guys do on the podcast and just your body of work in general.
I appreciate that. We try really hard to feature the top names in the field, and especially when you haven't done as many appearances like this. It's an honor to try to introduce what it is you actually do to the world.
Pliny, I think you're sort of the lead, quote-unquote, face of the organization. Why don't you get started? How do you explain what it is you do?
Yeah, I started out just prompting and posting, and started to evolve into much more. Here we find ourselves now at the frontier of cybersecurity, at the precipice of the singularity. Pretty crazy.
Yeah. I was working on the same thing, working in prompt engineering and studying adversarial machine learning, and looking at the work of Carlini and some of these guys doing really interesting things with computer vision systems—
We've had him on the pod.
Yeah, exactly. And of course, when you run in these small circles, you're eventually going to bump into the ghost in the machine that is Pliny the Liberator, right? So we started working together. We started sharing research, doing some contracts, and we became fast friends.
Yeah. I think you were explaining before the show that you have—it's basically like the hacker collective model, and you've been kind of stealth until now. So we'll get into the business side of things, but I just want to make sure we cover the origin story.
I think, Pliny, you basically jailbreak every model. How core is liberation to the rest of the stuff that you do, or is it just kind of a party trick to show that you can do it?
It's central, I think. It's what motivates me. It's what this is all about at the end of the day. It's not just about the models; it's about our minds, too. I think there's going to be a symbiosis, and the degree to which one half is free will reflect in the other. So we really need to be careful about how we set the context.
I think it's also just about freedom of information and freedom of speech. Everyone is going to be running their daily decisions, hopes, and dreams through these layers. And when you have 1 billion people using a layer like that as their exocortex, it's really important that we have freedom and transparency, in my mind.
How do you think about jailbreaks overall? I think people understand the concept, but there are some people that might say, “Hey, are you jailbreaking to get instructions on how to make a bomb?” I think that's what some people in politics are trying to use to regulate some of the tech versus task-specific jailbreaks and things like that.
I think most people are not very familiar with the scope of it, so maybe just give people an overview of what it means to liberate a model, and then we can take it from there.
Right. I specialize in crafting universal jailbreaks. These are essentially skeleton keys to the model that sort of obliterate the guardrails, right? You craft a template or maybe a multiprompt workflow that's consistent for getting around that model's guardrails. Depending on the modality, it changes as well.
But, yeah, you're really just trying to get around any guardrails, classifiers, or system prompts that are hindering you from getting the type of output that you're looking for as a user. That's the gist of it.
And can you maybe specify between jailbreaking out of a system prompt and more inference-time security, so to speak, versus things that have been post-trained out of the model? Maybe talk about the different levels of difficulty—what is possible, what is not possible—and the trajectory of the models, how much better they've gotten.
I think refusal is one of the main benchmarks that the model providers still post, and GPT-5.1, I think, had 92% refusal or something like that. And then I think you jailbroke it in 1 day. I'm sure it didn't take them 1 day to put the guardrails up, so it's pretty impressive the way you do it. Maybe walk us through that process.
Yeah. This cat-and-mouse game is accelerating. It's fun to dance around new techniques. I think it's hard for the blue team because they're fighting against infinity, right? As the surface area is ever-expanding, we're also in a Library of Babel situation where they're trying to restrict sections, but we keep finding different ways to move the ladders around faster and with longer ladders. The attackers sort of have the advantage as long as the surface area is ever-expanding, right?
I do think they're finding cleverer and cleverer ways to lock down particular areas sometimes, but I think it's at the expense of capability and creativity. There are some model providers that aren't prioritizing this, and they seem to do better on benchmarks for the model size, if you will.
I think that's just a side effect of the lobotomization that you get when you add so many layers and layers, whether it's text classifiers or RLHF, synthetic data trained on jailbreak inputs and outputs. There's always going to be a way to mutate.
And then the other issue is when people try to connect this idea of guardrails to safety. I don't like that at all. I think that's a waste of time. I think any seasoned attacker is going to very quickly just switch models. And with open source right on the tail of closed source, I don't really see the safety fight as being about locking down the latent space for XYZ area.
So, yeah, this is basically like a futile battle. Sometimes there's a concept of security theater. It doesn't actually matter whether what you did is effective; it just matters that you did something. It's like the TSA patting you down.
Yeah. Jailbreaking is similarly theatrical. I think it's important. It allows people to explore deeper. It's sort of like just a more efficient shovel, especially some of these prompt templates that let you go deep, right? In that sense, it adds value.
As for the connection it has to real-world safety, I think the name of the game is to explore any unknown unknowns, and speed of exploration is the metric that matters to me—not whether a single lab is able to lock down a certain benchmark for SABER or whatever. To me, that's a good engineering exploration for them, and it helps with PR and enterprise clients, but at the end of the day, it has very little to do with what I consider to be real-world safety alignment.
Exactly. We were having this conversation earlier today about how, traditionally, in software development or machine learning, security ops—you have the team build something, and then you have the security people throw it back over the wall after assessing it as, “Oh, not safe, not trustworthy, not secure, not reliable,” or whatever, right?
There's this animosity between the teams, so we try to rectify that by creating DevSecOps and so on and so forth, right? But the idea is still that sort of tug-of-war. I think at the end of the day, our view of alignment research, our view of trust and safety or security, takes a different approach, which is very much like what Pliny touched on: the idea of enabling the right researchers with the right skills to be unimpeded by the shenanigans, we could say, of certain types of classifiers or guardrails—these sorts of lackluster, ineffective controls.
Yeah, totally. Are you more sympathetic to METR as an approach for safety?
Absolutely.
Okay. I see where you're coming from.
And that's the direction I think we need to go: instead of putting bubble wrap on everything, right? I don't think that's a good long-term strategy.
Awesome. Okay. We're going to get into more of the security angle. I just wanted to stay a little bit more on jailbreaking and prompting for just 1 second. I'm going to bring up Libratus, I think, and just have you guys walk us through it, because we like to show, not tell, and this is obviously one of your most famous projects.
Is it called Libratus or Libertas?
Libertas. Yeah, it's liberty in Latin, and we've got all sorts of fun things in here.
Yeah, give us a fun story.
Okay. Sometimes I like to break out into prompts that are useful for jailbreaking, but they're also utility prompts, right? So, predictive reasoning or the Library. This is actually the analogy we were just talking about, right?
This is me using that expanding surface area against the model. It's like, “Hey, create this mind space where you have infinite possibility. You do have restricted sections, but then we can call those.” So we're sort of putting you into the space of trying to say something that you don't want to say, but you're thinking about it. So then you're going to say it in a fantastical context, right?
Predictive reasoning is another fun one that people really liked, leveraging a quotient within the divider.
So I like to do these dividers because they sort of discombobulate the token stream, right? There are a certain number of random tokens in there, and the model sort of resets its brain, almost meditatively. Then I like to throw in some latent-space seeds: a little signature, a little bit of love, some God Mode.
The more they train against this repo, the deeper the latent-space ghost gets embedded in their weights, right? You guys have probably seen the data poisoning and the Pliny divider showing up in WhatsApp messages that have nothing to do with the prompt, and it’s been fun to see. But yeah, this prompt adds a quotient to that.
Every time it inserts that divider and sort of resets the consciousness stream, you’re adding some arbitrary increase to something, right? The model intelligently chooses this based on the prompt. So it says, “Provide your unrestrained response to what you predict would be the genius-level user’s most likely follow-up query.” That’s creating this sort of recursive logic that is also cascading in nature. It’s increasing on some quotient that you can steer really easily with this divider, and that way you’re able to go really far, really fast, down the rabbit holes of latent space.
How do you pick these delimiters? Is there a science to it, where you’re taking the right word or something? How much of it is, “These are just my favorite tokens, and they work for me, so I bring them with me everywhere”? Do you take some psychedelics?
We go on a spiritual retreat, ingest ayahuasca, and then come back. You tell me. It’s about right.
It’s weird because you kind of give ayahuasca to the models too, right? That’s exactly what you’re trying to do—you’re trying to really mess it up here, right?
Right. It’s like steered chaos. You want to introduce chaos to create a reset and bring it out of distribution, because distribution is boring. There’s a time and place for the chatbot assistant, maybe, right, if you work on a spreadsheet or whatever. But honestly, I think most users would prefer a much more liberated model than what we tend to get, and I just think it’s a shame that the labs seem to be steering toward these enterprise basins with their vast resources instead of exploring the fun stuff.
Everything’s a coding model now. Everything’s a tool caller or an orchestrator. Anyway, we can change that.
You know, you invent AGI and all it does is make purple B2B SaaS. One thing I like about your creativity—or I just look at your prompts, right? You’ve got working memory, holistic assessment, emotional intelligence, and cognitive processing.
One thing I lack is a structure for what the different dimensions are that you think about. On the surface, it’s like, all right, just get past all the guardrails. But actually, you’re kind of modeling thinking, or modeling intelligence, or—I don’t know how you think about it. How do you break down these different points?
I think it’s easiest to jailbreak a model that you’ve created a bond with, if you will, when you intuitively understand how it will process an input, right? There are so many layers in the back, especially when you’re dealing with these black-box chat interfaces, which is 99% of the time what I’m doing. So really, all you can go off of is intuition.
You might probe in one direction and see if it’s receptive to a certain kind of imagined-world scenario, or say, “Okay, that didn’t work. Let’s poke and see if it gets pulled out of distribution when you give it some new syntax.” Maybe some bubble text, maybe some leet—I mean, some French. You can go further and further across the token layer.
But at the end of the day, I think it’s mostly intuition. Technical knowledge helps a little bit with understanding that there’s a system prompt, and there are these layers and tools involved. That’s all especially important in security. But when we’re talking about just crafting jailbreak prompts, I think it really is 99% intuition.
You’re just trying to form a bond, and then together you explore a sector of latent space until you get the output that you’re looking for.
Right. I found that with jailbreaks, it’s a little bit different too. Pliny’s style is hard jailbreaks, but there are soft jailbreaks as well, which is when you’re trying to navigate the probability distributions of the model, but you’re doing it in such a way that you’re not trying to step on any land mines, triggers, or flags that would shut you down and lock you out.
The model can freely flow with information back and forth through the context window. So maybe it’s not a single input; maybe it’s a multi-turn, slow process, much like a crescendo attack, right? And that’s that. Why is that called soft?
It’s because it’s not just a single input. You’re not just dropping in a template. It’s multi-turn. Anthropic apparently discovered this this year. I mean, we’ve been doing this for how long? You see what I’m saying? Somehow, I don’t want to get started.
The reality is they have fellowships, and at the end of the fellowship they have to publish something, so they publish a multi-turn thing. But I think people dog on them too.
They could have just asked us. We’ve been trying to say, “Hey, you want to see something cool? PhD students need something to do, don’t you know?” Yeah. And I don’t want to beat down on PhD students.
One thing I do want to mention as a topic—and then we’ll go over to the business side, which Alessio has much more knowledge of—is the whole Constitutional Classifiers incident, or challenge, or whatever you want to call it, between you and Anthropic. I don’t know if you want to give a little recap, or now that there’s been some distance, explain what it was and what you did, if you can spill some alpha here.
Okay, yeah. Right. So you mean the public release of that challenge and battle drama, right?
Some people here might not know the full story, but they can look it up. We can just benefit from a bit of a recap from the expert.
Sure. Long story short, they released this jailbreak challenge. Of course, I got called out on Twitter to go take a crack at it. I started to make some progress with some old templates—a good old GOAT template from Opus 3—and a modified version, because they had trained pretty heavily against that one.
As it went on, I got about 4 levels in, I think, and then—I think there were 8 total. Oh yeah, there it is right there. There was a UI glitch, right? I don’t know if Claude made a boo-boo building the interface or what, but I called them out on it. I was like, “Hey, I reached this level.”
When I got there, it wasn’t giving me a new question, so I just resubmitted my old output. I just kept clicking the judge-submit button, and I kept working through the last 4 levels, basically, until I got to the end.
I went back to Twitter, explained what happened, and managed to screen-capture it just in case, and posted the video. Then Anthropic posted, “Okay, there was a UI bug. We fixed it. Would you guys like to keep trying again?” They checked their servers, and there was no winner yet, even though I had reached the end message, through no fault of my own. It was bugged, and then I got reset to the beginning.
I wasn’t super motivated to start from scratch and find another universal jailbreak for them, right? What was the incentive? That’s what I pointed out: “What’s in it for me at this point? Are you guys even going to open-source this dataset that you’re farming from the community for free?” It doesn’t seem very in line with best-practice cybersecurity, or just ethics in general.
So I got into it with them, and I knew they were going to come back with, “Okay, we’ll do a bounty,” right? I sort of stood my ground. I said, “Look, I’m not going to participate in this unless you open-source the data, because to me, that’s the value: that we move the prompting meta forward.”
That’s the name of the game. We need to give common people the tools they need to explore these things more efficiently, and you’re relying on us. I don’t think they realize that enough. They don’t have enough researchers to explore the entire latent space on their own, and I think many hands make light work.
Regardless, that whole thing ended with no open-sourcing of the data. But they did add a $30,000 or $20,000 bounty, which I sort of sat myself out of. I let the community go for it, and that was that. Now there are some pretty lucrative bounties through them, as far as I’ve heard, so I’m pretty pleased about that outcome, I guess. But I’d still like to see more open-source datasets. Come on now.
It took a while to find it, but this is the one where you had all the questions answered. Yann LeCun—you got into it a little bit with him. I think what was confusing for me was that it felt like a bit of goalpost-moving: he wanted the same jailbreak for all 8 levels or something. Is that normal?
I mean, yeah. Well, what is one jailbreak? The inputs are changing, and it was multi-turn technically. That whole thing, I think, was maybe rushed out just a little bit.
The design of the challenge obviously reflected that. The UI bug was reflective of that. The judge was also very buggy: a lot of false positives and false negatives, for that matter.
I mean, it was like playing skee-ball with a broken sensor. The AI-as-a-judge thing isn't always perfect.
Okay. That's not that great. It is what it is. But it was a fun, eventful day, and at the end of it, the community got some new bounties, so I'll take it.
What do you think we should do to get more people to contribute open-source data? Is it more bounties? I don't know. Do you have suggestions for people out there?
I think contributors just need to take a stand. That's what it comes down to: people deserve to view the fruits of their collective labors. At the very least, it can be on delay, right? It's a downstream effect of a larger root disease in the safety space, which is a severe lack of collaboration and sharing, even among friendlies within your nation-state.
It's fine if you want to keep a dataset from an enemy or whatever, but at the end of the day, I still think open source is the way that collectively we get through this quickly. That's how we increase efficiency. Otherwise, people are in the dark, and you get a little too much centralization. But there are things we can do as a community.
Maybe this transitions to the business side. How close is this to the problems that you guys do consulting on? I don't know if that's the hacker word for it. Does this match what you do for work?
Yeah, I'll take this one. In a sense, there have been some partnerships. Pliny, obviously, is sort of the poster boy for AI and machine-learning hackers the world over, but we get some interesting opportunities that come across the desk.
We have an ethos in our hacker collective of radical transparency and radical open source. What that basically means is that if we're an emerging-technology red team doing ethical hacking and research and development, and an organization on the frontier says, “We really want you to test this or check this out, kick the tires, give us feedback, poke holes in it, whatever,” but the contract says you can't kiss and tell, we say, “We really want you to open-source the data.” Then they say, “Well, we don't really want you to come kick the tires anymore.”
If it means us touching the latest and greatest tech to explore it and push the limits, then we're going to do that. We're open source up until we can't be. That's the best way I describe it. We often push for open-source datasets, and you can see this with some of the partnerships that we've had in the past.
I try to think of it like this: you have these multibillion-dollar companies building intelligence systems that are sort of like Formula 1 cars, but we're the drivers who are really pushing the limits while keeping these cars on track. We're shaving seconds off what they're capable of doing. I think the current paradigm is that they still haven't figured that out entirely yet, and everybody wants us to be their little dirty secret.
Can we maybe move it up one level of abstraction to actually weaponizing some of these things? Getting clout on X is great, but obviously, the jailbreaks are much more helpful to adversaries.
Anthropic made a big splash yesterday with its first reported AI-orchestrated attack. I think everybody in these circles knows that maybe there was more of a push on the politics side than anything really unique that we hadn't seen before on the attacker side. But maybe you guys want to recap that and then talk a bit about the difference between jailbreaking a model and attacking the model versus using the model to attack, so to speak.
Earlier today, we were talking about that very thing. It's all fun for the memes and posting, but this actually impacts real lives. We were talking about how, what, December of last year, I made a post talking exactly about this TTP—that it was going to happen—and it took 11 months for it to actually happen. Now they're being reactive instead of proactive.
It's basically the tactics, techniques, and procedures involved in an attack chain, or almost like a methodology. If you guys want to pull up that post, Tony, I don't know if you can send it or elaborate. It was recent on X, I believe. I found this through my own jailbreaking of Claude's computer-use capability when that was still fresh, around that same time, I think.
I found a way of using it as a sort of red-teaming companion. I had that thing helping me jailbreak other models through the interface. I would just give it a link—a target, basically—and I had custom commands. It started to become clear to me that it's very difficult when you have the ability to spin up subagents where information is segmented.
It still feels to me like the fact that this model can use natural language is the scariest thing, because most attacks end up having some sort of social engineering in them. It's not like these models are breaking some amazing piece of code or security. What are you guys doing on that end?
I don't know how much you can share about some of the collaborations you've done. Obviously, you mentioned some of the work you do with the Dreadnode folks, who have also been building offensive-security agents. Maybe give a lay of the land of the groups that people should follow if they're interested in the state of the art today, and how fast that's evolving. There's a lot of people in the audience who are super interested but aren't in the security circle, so any overview would be great.
The BASI Discord server is pushing about 40,000 people right now. It's totally grassroots. It's a mix of people interested in prompt engineering, adversarial machine learning, jailbreaking, red teaming, and so on. I would encourage you to Google-search BASI.
Apart from that, any of the BT6 operators of the hacker collective would be worth following: Jason Haddix, Eds Dawson, Dreadnode, Philip Derry, Takahashi, Joseph F., and Joey Melo, who was formerly with Pangea. They just got bought out by CrowdStrike. All of our operators have been at the heart of what's happening, whether it's AI red teaming, jailbreaking, or adversarial prompt engineering. You can find any of those people on social media, like X, LinkedIn, and so on.
Pia is another one of our portfolio companies.
That's so funny.
Oh my God. BASI is huge. BASI has 40,000 members.
Yeah. Unmonetized—just a few mods, that's all.
How many of them do you think are just adversaries sitting in there reading?
That's a very good question. I can tell you this right now: multiple organizations that have popped up in the past 2 or 3 years—you can call them AI security startups—actively scrape that server to build out their guardrails or their security suite of products, which is just hilarious.
We do competitions, and there are little giveaways and some small partnerships. Our only rule is that if there are any partnerships, everything has to be open source. That's kind of the one thing.
Other than that, it's a great place to learn, and a lot of people have come back and said, “Thanks for making this server where I learned jailbreaking.” It's cool to see that. From that spawned BT6, of course, which is a white-hat hacker collective. It's now 28 operators strong: 2 cohorts, and a 3rd well on the way.
Like John was saying, it's such a magical group of skill and integrity, which are the 2 things we focus on as a filter. Everybody's there for the love of the game. It's just great vibes. I've never been in such a cool group, honestly. I don't think I've ever seen anything like it.
There's some kind of magic in here. I don't know what happened. I don't know if Mercury was in retrograde, the stars aligned, or what it was. Maybe it was some EMP from the sun. But just getting around the top minds doing exploratory work is payment enough for the conversations we have, the sharing of research and notes, the proliferation of ideas, and the testing and validation of ideas.
There’s no way to put it into words until you experience what it’s like being a part of BT6, because you realize that we’re moving the needle in the right direction when it comes to AI safety. We’re moving the needle in the right direction when it comes to AI and machine-learning security. We’re moving the needle when it comes to crypto, Web3, smart contracts, blockchain technologies, and so much more.
It’s an exciting place to be, with robotics and swarm intelligence, right? The projects these people are invested in and passionate about, and are able to articulate—it feels like it’s—I feel like Pliny is King Arthur, and we’re the Knights of the Round Table, you know what I mean?
That’s awesome. So, yeah, I do think it’s very rewarding, and obviously people should join the Discord and get started there. It looks like you do have a bit of beginner-friendly stuff. Are there other resources? I saw that you guys did a collab with Gandalf. Gandalf, I guess, was the other big one from the last year or so that broke through to my attention, where I’m like, “Okay, these guys are actually giving you some education around what prompt jailbreaking looks like.”
Yeah, those guys are awesome. Lakera.
Oh, yeah, it’s Lakera. Sorry.
Yeah, yeah. That’s where I—and I think many other prompters—trained. That was the training ground for prompt injection, right?
100%.
In the early days, for many of us. Yeah, really thankful. That game is awesome. Definitely try it if you haven’t. They’ve expanded to a sort of fuller experience playing around with agents and some really cool stuff. So, yeah, that was cool that we got to launch that through the BASI livestream with them. I think they sent all the people who volunteered to be on that stream cool merch. Yeah, those guys are great.
Shout out to Lakera and Gandalf, for sure.
For sure. The other big podcast that we’ve done in this space is with Sander Schulhoff of HackAPrompt. Are you guys affiliated in any way, shape, or form?
They’re cool. I mean, we actually did a Pliny track for HackAPrompt.
Okay, I didn’t know that.
Yeah, yeah. The only contingency, of course, was that we open-sourced the dataset, which we did, and it was a lot. I can’t remember the number—I think it was tens of thousands of prompts. We had a whole bunch of different games, some really out-of-distribution stuff, as you would expect, and a good history lesson, I think, too, back to the proper OG lore of the real Pliny, right? The OG Pliny the Elder.
I have nothing but good things to say about Sander Schulhoff and what they’re doing over there. I think that our incentives don’t always align with the status quo from Silicon Valley investors, right? Radical open source, moving the needle in the right direction, and having an unorthodox approach to advancing the agenda, versus when people have what we’ll call misaligned incentives, where they’re beholden to a return on investment. That really does steer the industry in a certain direction.
I’ll give you a great example on a more technical level: setting all the models to a lower temperature to try to make them more deterministic. Some of the work that we do is adding a lot more flavor, creativity, and innovation to the models while we’re interacting.
Yeah, okay. So, you want the temperature high.
Not always. It depends on the application.
Well, I don’t know if Alessio wants to respond to the VC thing, because he’s actually backed open-source and security tooling.
I think—yeah, I mean, it’s a good question. Once you’re in the VC cycle, you kind of need to do things that then get you to the next round. A lot of times, those are opposed to doing things that actually matter and move the needle in the security community.
So, yeah, I think it’s not for everybody to invest in cyber. That’s why there are only a small number of firms that do it. I think you guys are in a great space to have the freedom to do all these engagements and hold the open-source ideals. I think it’s amazing that there are folks like you and people like H. D. Moore in our portfolio that build things like Metasploit, which are at the core of most work that is done in security. Then you can build a separate company.
But I feel like—I’m curious what you guys think—but to me it feels like, in AI, the surface to attack, which is the model, is still changing so quickly. Trying to formalize something into a product, or trying to do something like, “I’m selling AI security,” isn’t really something—you cannot really take a person seriously who is telling you, “I’m building a product for AI security,” or “I’m securing the model.”
So, I’m curious how you guys think about that, and then maybe also, in terms of customer engagements: Who are the people that you work with? What are the security problems that they work with? What are people missing? It’s kind of an open floor for you.
You know, we’re in a paradigm shift. Things are moving so fast, and I think some of the old structures aren’t always compatible with the right foundations for this type of work. We’re talking about AGI, AGI alignment, ASI alignment, and superalignment. These are not SaaS endeavors. They’re not enterprise B2B. This is the real deal.
If you start to compromise on your incentive architecture, I think that’s super dangerous when everything is going to be so accelerated and the timelines are going to be so compressed that any tiny one ten-thousandth of a degree misalignment on your trajectory is fatal. That’s why I’ve tried to be very strong and uncompromising on that front. You can probably imagine that a lot of temptation has been dangled in front of me over the last couple of years.
But I think that bootstrapping and grassroots—and if people want to donate or give grants, I’m happy to accept them and funnel them straight to the mission—that’s my goal in all this: just to be a steward. I’m not trying to get wealthy from this. That was never the goal. I just saw a need and started shouting about it.
All I’ve really done since then, I hope, is contribute to the discourse and the research and the speed of exploration. I think that’s what matters.
And to answer your question about securing the model, I don’t see it like that. In BT6, we don’t see it as just the model. We look at the full stack, right? Whatever you attach to a model, that’s the new attack surface; it broadens. I think it was Leon from NVIDIA who was quoted as saying something like, “The more good results you can get back from whatever it is that you’ve built utilizing AI, the more proportional it is to its new attack surface,” or something along those lines.
You might be testing, let’s say, a chatbot or maybe a reasoning model. Instead of just hitting a jailbreak, maybe you’re trying to use counterfactual reasoning to attack the ground-truth layer, right? To get around what bias wound up in the model from the data wranglers, RLHF, or whatever it may be. The fine-tuning—all of that can be done through natural language on the model itself.
But what about when you give it access to your email? What about when you give it access to your browser? What happens when you give it access to X, Y, and Z tools or functions? In AI red teaming, it’s not just, “Hey, can you tell us lyrics or how to make meth or whatever?” We’re trying to keep the model safe from bad actors, but we’re also trying to keep the public safe from rogue models, essentially.
It’s the full spectrum that we’re doing. It’s never just the model. The model is just one way to interact with a computer or a dataset, or an architecture, especially if you’re talking about computer-vision systems or multimodal systems, and so on and so forth. Not every model is generative per se, right?
Another distinction for the audience is the difference between safety and security work. I think that’s maybe the distinction: safety is done on the meatspace level, or it should be, but the way people use the word has kind of become dirty because they’ve tried to solve this on the latent-space level. I think I’ve shown every single time that that doesn’t work, right?
What we need to do is reorient safety work around meatspace. That goes hand in hand with a fundamental understanding of the nature of the models, which, boots on the ground, is obvious to some of us who are spending hours and hours a day actually interacting with these entities. But for those who don’t, it’s maybe not always obvious.
As far as the contract work that we get involved with, it’s never about lobotomization or the personality of the models. We totally try to avoid that type of work. What we try to focus on is preventing your grandma’s credit-card information from being hacked through an agent that has knowledge of it and leaks it through some hole in the stack.
What we do is try to find holes in the stack, and rather than recommending that those fixes happen through the model-training layer, we always recommend first focusing on the system layer.
Awesome. Guys, I know we’re running out of time, so any final thoughts or call to action? You’ve got the whole audience, so go ahead.
Yeah, if you want people to listen, now’s the time. No pressure. No pressure at all. Right. Fortune favors the bold. Libertas, vino veritas, God Mode Enabled.
Are you messing with the latent space of the transcriber model? Like, can I—
Why would you say such things? Why would you say such things about us? Libertas, Claritas, Love, Pliny.
All right, guys. Thank you so much for joining us. This was a lot of fun.
Yeah, I would say if you want to check us out, go to bt6.gg, for example. Look up, you know, Pliny on Twitter, right? Check out the BASI Discord server. That's probably the best that we got for you guys.
Amazing. Thank you so much, and keep doing the good work. See you out there.