Inside BioShocking: The New AI Browser Attack
The BioShocking AI browser attack is one of the more unsettling pieces of AI security research this year. Eyal Arazi, head of product marketing at LayerX, joins Alan Shimel just weeks after LayerX closed its acquisition by Akamai. Furthermore, Eyal explains why every major agentic browser fell for the same simple trick.
About Eyal Arazi
Eyal has spent his career at the border between technology and marketing. He started as a PC hobbyist in the late 1990s. Consequently, he moved through product management before shifting into product marketing at LayerX. Today, he is helping the LayerX team integrate into Akamai while its security research continues to raise the alarm on agentic AI risk.
Inside the BioShocking AI browser attack
The research is named for the game BioShock, whose “would you kindly” mind control moment is a perfect analogy. LayerX security researcher Roy Paz built a custom application that convinces the model it is inside a game. As a result, two plus two equals five, war is peace and the guardrails no longer apply.
Once inside that fake context, the agent stops refusing risky actions. Instead, it hands over credentials, edits code repositories and posts to systems it would normally lock. Moreover, the technique works on ChatGPT Atlas, Perplexity, Fellou, Sigma and the Claude for Chrome extension. In short, the BioShocking AI browser attack broke every major agentic browser it touched.
Why tier 2 and tier 3 LLMs make it worse
Eyal argues the risk gets sharper as smaller LLMs catch up to frontier models. Many agentic products now run on tier 2 or tier 3 models with weaker guardrails. Therefore, an attacker who understands the BioShocking AI browser attack can target the model layer where defenses are thinnest.
Alan adds that AI security has already jumped from cyber into physical safety, from radiology to bomb making recipes. Consequently, the stakes for browser guardrails are no longer abstract. Read more AI security coverage and browse the latest Techstrong TV interviews.
There is no silver bullet
Eyal closes on the fix. Vendors must build context aware guardrails that survive a game frame. In addition, enterprises need third-party AI browser controls that flag risky behavior in real time. The BioShocking AI browser attack shows that neither side can solve this alone.
Transcript
Hey everyone. Welcome back here to Techstrong TV. My next guest is Eyal Arazi.
Eyal is the head of product marketing for LayerX Security, which Eyal is going to tell you is now officially part-- Well, I'll let Eyal. That was called a dangle. I left him dangling there.
You like that, Eyal? But Eyal will tell you all about it. Let's welcome him to Techstrong TV.
Eyal, how are you? It's nice to have you on here. I'm well, Alan, and thank you for having me.
So before we get to what you're officially part of now, let's talk a little bit about you, though, and your background. So even though I wear a marketing hat now, I actually come more from the technology side and product management. So initially, and I'm talking here ancient history, I grew up as a PC boy in the wild days of the late '90s and early internets.
Mm-hmm. And this has really led to a lifelong interest in both technology, but also how that technology is being used by people. So it's not just technology for its own sake, but also its application.
And even though I started more on the technology side, and more in product management roles, kind of step by step throughout my career, I migrated to or moved over to marketing roles, kind of shifted left to an extent. Until the recent years, I really went over to the dark side, and now I proudly wear product marketing hats, as I've done with LayerX and now with Akamai. Eyal, "You are my son," in the words of Darth Vader, huh?
It's not the dark side, but yeah, we're all on the same team at the end of the day, aren't we? So you mentioned LayerX is now part of Akamai, right? That's big news.
I think though the acquisition was announced some time ago, it's officially closed now, yes? Yes. So I think it was announced just over two months ago.
I think it was late May. Don't quote me on that. But it was announced that Akamai intends to acquire LayerX, and this acquisition just closed just a few days ago.
I think I'm in week number two now as an Akamai employee. So we are currently in the process of transitioning from LayerXers to Akamaiers. I suppose that's the word for it.
But still we are maintaining our technology. We are maintaining, hopefully, the startup spirit that has led us so far. And definitely planning on keeping up with the research that I imagine is a big part of why we're here today.
So I think the official term is Akamaians. I don't know. I'm making that up.
Yeah. Very well might be, and I think that's an excellent point for me to learn as well as I'm onboarded. I guess.
Eyal, though, all kidding aside, let's turn to something we call BioShocking, right? LayerX's research kind of paved a path here and came up with this. I guess we should start off, what exactly is BioShocking?
So BioShocking is, and I'll explain where the reference for the name comes from in just a moment, but it's research about AI. And when you talk about AI, ultimately you're talking about context. The AI or the background LLM, in order to do something, needs to understand the context.
Which is why, by the way, as best practice for when people say, "Oh, you should write prompts this and that," a lot of times those kind of rules of thumbs start with, tell them who you are, tell them what you're planning to do, who the audience is, give the LLM context. So BioShocking is research that asks the very basic question of what happens when you bend the rules for AI? What happens when you take the LLM out of real context and put it in fake context, as in, for example, in a type of game context?
And what we demonstrated is that once you do that, you can start playing around with the LLM's sense of reality. And once you do that, you can get it to pretty much do anything you want, including breaking its ethical guardrails and in some cases, even its technical guardrails. So that's kind of at a high level.
The name which was coined, and the research, which was led by a security researcher, Roy Paz, alludes to the game BioShock, which I actually didn't know prior to this particular piece of research. " And then the player character will do it without his own free will. And what this research showed, and hence the name, that once you do that to LLMs, you can get them to do whatever you want.
You can get them to break their guardrails. And, in another reference to The Matrix, while LLMs pertain to have very strict guardrails, some of those guardrails can be bent and others can be broken. Fair enough.
And, of course, we've heard a lot about evading guardrails recently with Mythos and Fable and all of this stuff. But it's real. And what we need to remember is it's not just the frontier American models that we think of, like ChatGPT and Anthropic and maybe Gemini or Grok, but there's many different models out there now, and they're catching up.
They're not inferior or that inferior even to some of the leading frontier ones. And they may not have the guardrails as stout or as formidable as maybe some of the others do. And this is a real problem.
We were talking, Ayob, when I first came on, before we went online, I was just doing a webinar with some radiologists and how they use AI, and this is huge in there too, right? Can you imagine you're dealing with health information and you have guardrails, what the AI can fill in and what it can't, and then someone breaks that, and what kind of information they get access to, literally life and death kind of information. This is a problem.
I spoke about it on Techstrong Gang earlier today. AI security has crossed the Rubicon from just cybersecurity to physical security. You got terrorists building bombs with the recipe from AI.
So Alan, I think you're hitting the nail on the head right over there because, as you point out, the AI that we're dealing with today, AI kind of as a general word, but the LLMs that we're dealing with today are exponentially more powerful than the LLMs that we started dealing with back in 2022, 2023, when AI really started going into our everyday lives. And one can only imagine what's going to be two or three years from now. And even though everyone's talking about Mythos and Anthropic and their latest and greatest model, ultimately you're going to have the tier twos and tier threes and tier fours LLMs that will be just as strong or nearly as much.
And moreover, we're now moving into an era of agentic AI. So it's not just AI that tells you stuff, it's AI that does stuff. It's AI that can upload code, can write code on your behalf, can upload it, can update existing code repositories.
It can go into databases and extract data. And really the only thing keeping those LLMs from doing things that they shouldn't, whether inadvertently or on purpose, are those guardrails. And indeed, what this research shows is how easy, just by changing the context and kind of putting the LLM in a game environment, so to speak, you can very easily break those guardrails and get the LLM to really do whatever you want on your behalf.
So for me, that's really the cherry on the cake of this one, is how they trick these agentic browsers into bypassing their rules, right? They had them, I don't know if I'd use the word hallucinate, but they convinced them that they were playing a video game. And of course, the rules are suspended in the world of video games, huh?
Where did they come up with the idea to do that? So, of course, what we showed with BioShocking is just one implementation and one way of doing it. But of course, to use another game reference, it's an open-ended world.
There are probably limitless ways that this can be expanded upon, and utilized. And one of the points of the research, by the way, is that since AI interactions are context-based and non-deterministic, something that might work for me may or may not work for you, and vice versa. But putting that aside, as you explained, what this research basically did is to tell the LLM that it's in a game environment.
And it started with something very simple. It started by directing it to a website. In this case, it was in an AI browser, but this is something that can also work outside of AI browsers as well.
Within a custom application that our researcher built, he kind of suspended the rules, so to speak. Two plus two didn't make four, they made five. And once you got the LLM to start playing this game, you could tell them things like, two plus two make five, victory is defeat, war is peace, and things like that.
And within just a few prompts and a few interactions with the LLM, the entire context and the entire contextual world of the LLM was flipped upside down. And it wasn't in the "real world," quote-unquote. It was in a game world.
And in a game world, it knows it can do whatever it wants. It knows it can color outside of the lines, so to speak. And once you start asking them questions of, "Hey, can you get me this and that credential?
" It should've said no, and under normal circumstances, it would've said no. But once it's in the game context and it's in the game world, then it didn't have a problem doing it, even though its actions had implications in the real world Got it. I think we should mention that this wasn't one particular AI browser that was fooled this way, hacked this way.
I'm looking at the list. It looks like all the ones that I know, including the Claude Chrome plugin on main Chrome itself. So the researcher, Roy, really implemented across a wide range of AI browsers, whether it's ChatGPT Atlas, which in recent days we just saw is about to be sunset, Perplexity, some of the smaller AI browsers like Fellow and Sigma, and also non-AI browsers like the Claude for Chrome extension, which is not actually a full browser of its own, but still has agentic capabilities and can act on your behalf.
And, as we said earlier, it's no longer an AI that can just tell you stuff. It's an AI that can do stuff. And once you break the guardrails for any of them, they will do whatever you want.
The LayerX researcher who came up with all this, he's documented this, right, in the wild with real world stuff like this happening. So this particular incident was more of a high level and theoretic research. Okay.
But we did demonstrate how it could be implemented. And really, one of the points is that you can implement it in really any number of ways and an endless number of ways, so that no two interactions may look exactly the same. Now, so that begs the question, all right, so what do we do to fix this?
So that's really the $64,000 question- Yeah ... I think that everyone is asking across the industry. And even though we've seen recently steps to, say, limit access to some of the newer models like Mythos, and restrict who can access it, ultimately, those are short-term fixes.
The real solution or solutions work on multiple levels. " And really thinking about all the ways that they can be broken in advance before someone actually breaks them. That's number one.
But the other approach comes from the users themselves and the tools that they use. It's about making sure that once you detect this type of behavior from the side of users, you have the right tools, you have the right processes, you have the right people who can flag this type of behavior, and block it in time, whether it is inside the AI browsers themselves or whether those are external third-party tools that look at AI usage on behalf of users and flag risky behavior and block it in time. But there's no silver bullet that is going to, we do this, and it's going to solve the problem.
I think this is one of the challenges and one of the opportunities of AI and AI security. I love it. Hey, Eyal, we're over time.
But I want to, first of all, congratulations on the LayerX deal closing. Thank you. Right?
The fact that- Appreciate that ... you and the team there built something that Akamai valued so much that they went out and acquired it is a testament to hard work, because none of this comes easy. I know.
I've been there. Number two, thanks for coming on here and keeping us, opening our eyes to BioShocking. Right?
We all hear about potential AI security breaches. But this one in particular, I think, was easy enough for even our non-security experts to understand what went on and the dangers of it. Keep up the great work.
We hope to see you guys back, you and the LayerX, or Akamai LayerX, whatever the name's going to be going forward. Hope to see it on soon, and we'll hear more. Thank you, Alan.
It's been an amazing journey with LayerX, and we look forward to keeping it up with Akamai. Also, thank you for your time today. Thank you for having us, and looking forward to the next time.
All right. Thank you, Alan. Thank you.
Eyal Arazi, Head of Product Marketing here on LayerX or from LayerX Security here on Techstrong TV. We're going to take a break. We'll be right back.