Inkling vs. GPT-Red: AI’s Open-Weight War Is Here
AI open-weight models are colliding head-on with AI security in today’s episode of Techstrong Gang. Hosts Alan Shimel and Jon Swartz are joined by Mitch Ashley, Yvette Schmitter, and Liz Safran. The panel breaks down a new Futurum report on enterprise AI adoption. It shows 56% of enterprises already run multiple foundation-model providers — a sign the model layer itself may already be commoditizing.
The panel then turns to the open-weight race defining 2026. Mira Murati’s Thinking Machines Lab has released Inkling, its first AI open-weight model. It’s positioned as a direct answer to Chinese open-source dominance from DeepSeek, Kimi, and Qwen. But as Techstrong.ai reports, Murati’s real bet may not be the model at all. It’s Tinker, her company’s fine-tuning platform.
Why AI Open-Weight Models Are Becoming a Geopolitical Battle
China isn’t standing still. At the World Artificial Intelligence Conference in Shanghai, Xi Jinping positioned China as the technical champion of AI openness. He backed a new 29-country World AI Cooperation Organization, framed as a counterweight to the US-led Pax Silica initiative. The centerpiece is Moonshot AI’s Kimi K3, a 2.8-trillion-parameter model. Full weights release July 27 — potentially the largest open-weight drop in AI history. Read the full breakdown at Digital CxO.
Mitch Ashley’s Futurum research maps the AI market into eight layers, from silicon to applications. The panel discusses why every vendor racing to own a layer is chasing the same trap. Making a layer indispensable does not guarantee lasting economic control. Full analysis is available at Techstrong.ai.
OpenAI’s GPT-Red and the Rise of AI-vs-AI Security Testing
Closing out the show, the panel examines OpenAI’s newly unveiled GPT-Red. It’s an internal red-teaming model trained via self-play reinforcement learning. GPT-Red attacks frontier models before real-world deployment. It beat human red-teamers 84% to 13% on novel attack scenarios. Training against it made GPT-5.6 Sol six times more resistant to prompt injection than OpenAI’s best model just four months earlier. Full details are in OpenAI’s announcement.
Watch today’s Techstrong Gang panel unpack AI open-weight models, the US-China open-source race, and the new era of AI-on-AI security testing. Today’s panel: Alan Shimel, Jon Swartz, Mitch Ashley, Yvette Schmitter, and Liz Safran. Techstrong Gang streams live every weekday at 12:00 PM ET.
Transcript
Hey everyone. Happy Monday to you, and welcome to Techstrong, gang. I was doing an interview this morning with a friend of mine, a Techstrong TV interview.
We both mentioned that you come in Monday so full, so excited to talk about, because so much happens over the weekends even now, right? If you weren't catching what's going on Friday afternoon and then carrying it through the weekend, you come in Monday and it's like a new world. It's an exciting time to be doing these, to be in tech, and to have an outlet.
Yvette was talking off camera about, it's nice to have a place like this to even just get this stuff off your chest. And have a bunch of pundits to shoot it around with. We hope you enjoy it, and we'll keep doing more of it.
But let me introduce you to our gang for today. We've got, I guess, the way it looks on my screen anyway, left to right, Yvette Schmitter. Yes, Yvette, it's good to have you back here.
It's our Monday. We got a real Monday thing going on here. A Monday thing.
Joining Yvette tonight, yeah, is Liz Safran, the Monday Club. That's right. And joining us on Monday is my friend Mitch Ashley.
Always good to be here. The man out West. Happy to have you.
Good to be here. We've got Jon Swartz. Hi.
Hey, John. I should mention we've got no Mike Vizard today. I don't know exactly what it was, but he mentioned he was doing something.
I think he's in Montreal, or he's in Canada. That's right, Montreal and Quebec and all of that. Trying to escape the wildfire.
Well, he probably went right into it. Well, that's Mike's sense of direction. Don't ever ask me about Mitchell's sense of direction.
But anyway, it's Monday, guys. We got a lot to go over, gang. I want to jump right into it.
Mitch, you published a piece up on the Futurum site on Monday. It's really a big piece of research you've been working on that leads into a theme. Really, it's the arc of your life, right?
Of your career. But it's about defining this new AI stack, right? That we're all dealing with, and interesting players and everything else.
But why don't you share a little bit about what your paper's about? Well, thanks, Alan. Happy to.
Kind of going through multiple of these cycles, right? With technology, whether it's cloud, client-server, cloud native, et cetera. You kind of see this pattern emerging where vendors are making a lot of announcements, and we're not always sure what they mean.
In this case, with AI, every vendor's making every announcement all across the board, and sometimes it's pretty dizzying trying to understand what exactly they're doing, what's their strategy, because they don't lay out the whole strategy for you. So now I'm officially in an analyst role, so I get to kind of think about it that way. So I created the agent control plane framework before I created cloud, and, excuse me, observability native.
There's a couple of frameworks for understanding those two aspects of what's happening in the market. But stepping back and take a bigger picture at it, is I created my version of the AI stack. Now, it's not going to be shocking what's in it.
What's important about it is not what any individual vendor is announcing, but what they're announcing where in the stack. What are they doing around governance? Do they have a runtime element?
Are they playing more in the kind of OpenShift operating system, infrastructure side of this? Are they up in the builder tools, the application surface? Where are they at?
And you can see, for example, a company like Google in their announcements at IO really have something all the way up and down the stack, or at least they've announced in each one of these areas. In Microsoft, you can see where they're concentrating, right? In the applications, in the builder, some of the runtime, and also the substrate.
But at the bottom, of course, we have silicon and the SDKs and frameworks that go with that. So my whole intent in this is to try to establish a model that we can use, A, so we have a frame of reference, a language to use, but even more importantly, buyers can use it to try to assess, hopefully with some analyst guidance, what the vendors are doing, what their strategies really are. And vendors can use it as a diagnostic tool to make sure they've got their messaging, how they're approaching the market makes sense to buyers and people like us that are talking about what they're doing.
So Mitch, I got to ask you though, do you got a tail back there? Is there some kind of animal behind you? If you want to know, I have animals to the left of me.
All right. Well, I saw a tail wagging. I didn't know if something was going on or what.
But okay. There'll be a cat, Nate the cyber cat, he'll be coming by soon. He always stops by and makes a visit.
So yeah. All good. Big animals.
But no, it's a great paper, Mitch. And I actually did a follow-up article, and the article got a little long on me as I kept digging into it because it was a meaty paper. The whole concept is meaty with a lot to dig into.
" And that's why I think I got so into it. Here's the bottom line. Why are these people doing this?
Because historically, when we look at waves of technology, no one wants to be commoditized out of the high margin. No one wants to be stuck Just being the electric company, not GE. No one wants to just be the fiber player, not the cloud hypervisor.
And so Mitch, I took your diagrams and I gave them an up and a down arrow where the value's rising, where it's going down. " Everybody wants to move upstack. However, those who already have upstack real estate, they want to shore up their foundation to make sure no one can pull the rug out from under them.
Mm-hmm. And so they're looking at moving downstack, if you will, not to invest the trillions of dollars we're seeing invested, but just to shore up that no one can pull the rug out from them. It's an interesting time.
John, you got your hand up? Yeah. So the concept, it was an excellent paper, Mitch.
I learned a lot from it, and it opens my mind up to what's actually happening and gives me a kind of an overall narrative, and it overlays with what Alan's been talking about with his book that he's working on. So my takeaway, and I think Alan has mentioned this. So to avoid becoming trapped in this commoditized layer, every major vendor's got to expand beyond their point of origin.
So the chip makers are shifting to software platforms. The model providers are bundling or building user-facing apps. I think the SaaS companies are positioning themselves as agent control planes.
In other words, vendors are kind of racing to establish dominance in layers that create unique sticky business value, like data access, memory, governance. And I think, as Alan said, they are doing this before their foundational products are reduced to mere utilities. So it was very interesting.
And it shows me where this is going and why all the major players are shifting in the way they are. So, that was my takeaway. Sure.
Yvette, you've got experience with big companies like this and Upstack and so forth. What's your take? So, I guess my mic drop on this is basically I'm like, in this market, indispensable isn't a moat.
In my opinion, I believe it's a countdown clock because the only really defensible layer is the one where your customer's data and your workflows live. And most vendors on that stack and that chart don't own either. So yes, they're moving up, but I just want to say this, remember what Microsoft did.
I do think that there's going to be an in where who's going to own the logins? Who's going to own your new Active Directory layer for agents? Just think how this is going to go.
It's going to go- No. It's funny. There are a bunch of companies, some I've interviewed, that actually do want to own identity and access control for agents, which is sort of what Active Directory was doing for people 100 years ago.
Yeah. But listening to it and picturing it, you guys don't remember this. Liz may because she's a New Yorker.
But when we used to play stickball in the streets in New York as a kid, how do you choose your teams? Well, you would take a stickball bat, and you'd throw the bat at the other captain. He had to grab it.
And the person who was at the top of-- You're a New Yorker. You probably have seen this. Mm-hmm.
The person who got to the top of the bat, who had the grip on the very top of the bat, he got first pick. Yep. In the beginning, you would take up a lot of space, and then there'd be the fight to just get those fingertips at the top of the bat.
This way, you could pick the best kid on the team. Yeah. That's what we're playing here with the AI stack.
I want to be at the top of that one who owns the login. Who- That's a great analogy Yeah. I love that you're bringing back the be home before the streetlights come on memories that I'm having.
But here's the one thing that I want also on the gang as well as everyone that's listening is that think about this: 56% of the enterprise today, the builders already run on multiple model providers in production. Yes. Only like 15% are a single vendor.
And that's not the trend line. That's the verdict. " So I do think the fight is going to continue to move up, and in my opinion, it's going to be the orchestration layer, like who decides which model gets called, what data it sees, what tools it touches, and how the output becomes a business process.
And that's not a technology decision. That's ultimately going to be a governance decision wearing, cosplaying a tech costume. So that's just my thoughts right there.
I like how you think, Yvette. Alan and I have this unique experience. We, in kind of different capacities, started a company called Lattice Networks in 1999.
It was all about the data center market. Everybody was co-lo-ing, putting their stuff offsite before cloud, if you will. And we set up a company basically to arbitrage reselling all of this overbuild.
And what's interesting is all the focus is all around the cloud, meaning the AI cloud, the data centers. And of course, it causes a lot of public attention, but that is and will be commoditized, per Alan's book. We've seen it happen over and over.
The value moves up. It's going to get overbuilt. There are companies that will go out of business because they tried to do me too, the data center, the AI data center business.
" Well, there are a lot of things that aren't, but there are some that are different. One is we're seeing an announcement at every layer of the stack all at once. You go to any vendors conference, and Google announces something in the application layer, the new agentic work surface.
AWS has Quick at the application layer. Who would have guessed? Why would AWS come out with that?
Why would IBM come out with an IDE? What sense does that make? How does that fit into their strategy when they're more of a buy companies integration, sort of, infrastructure player, if you will?
Yes, some software as well and tools also. So you have to look at it, and you can't just throw it against the wall and say, "What does this look like? " You have to have some way of analyzing it and just kind of categorizing and understand.
And when you start to fill in the gaps, when you start to fill in what each one is doing, I created these archetypes. So like a harness-coupled archetype, a cross substrate for the model companies like the Anthropics and the OpenAIs. What are they doing?
They're trying to move into, yes, Anthropic already does development, but also move into cowork, the work surface, as well as what OpenAI is trying to do in the enterprise surface. So point being, I hope this is a useful both diagnostic and just kind of getting your arms around what's happening, and I appreciate your article, Alan. I think you did a fantastic job of taking another lens to it, which is, yes, commoditize, but also where's the value going?
Can I ask you a quick question, Mitch? Mm-hmm. And tell me if I'm totally off base, because is there kind of a paradox in play with the AI market?
And here's the point I guess that I was trying to-- I'm thinking of, is that the more indispensable a technology's capability becomes, the faster the market forces capital and alternatives to act- That's the premise of the book. That's the trap ... Yeah.
That is Alan's book. The faster the market forces capital and alternatives to act to turn it into a cheap, standardized utility- Yeah ... which I think is what Alan's book is about.
Yeah, that's exactly the book. But what happens is the amount of money invested to get there, right? So in the dot com era, it was a trillion dollars.
And back then, a trillion dollars was a trillion dollars, not like now. It was a lot of money. And it was a trillion dollars invested in fiber, of which about 5% to 10% was lit.
The rest was dark. And that money got wiped out in bankruptcies by WorldCom and Global Crossing and all those good... I'm not saying that's necessarily going to happen here, but here's the fact: when you go back and look at that fiber boondoggle, what was built on top of it, what went upstack?
The cloud. AWS, Google, Microsoft's clouds built on the bones of those fibers. And what's built on top of that?
Cloud native. Right? That allowed you to scale out the way these things scale.
" That gives the road for AI to come in. These bottom layers that are in the bottom of Mitchell's chart or diagram are the bones of what came before it, right? And that is the game.
And there's different ways of getting cap, though. You could have the government, you could have the competition do it. There's all different...
Once you're indispensable, you got a big target on your head. That's always the way. Are you- Right.
Well, that made me think of way back in March, so many internet years ago when- ... Anthropic was declared a supply chain risk, right? Claude was kind of generally accepted to be the best model, but that designation from the US government- Can set back ...
made them from indispensable to frozen out, so... Yeah. Mm-hmm.
There's one other factor that I want to point out. Maybe you were kind of hinting at this too, John, is what we have now is there's always the what is your uniquely identifiable, defensible, and intellectual property, right? Everybody asks that about their start.
You can have something that differentiates yourself, but is it defensible? There's something that someone else can't replicate easily. Mm-hmm.
And in this case, parts of that are out the door because I can recreate what you did in software yesterday, today. I can do it that quickly. So that's why we see the model, the Anthropics and OpenAI and Google going back and forth announcing the same things over and over because there is no defensibility in that.
So that, and I think I- No. And it's not just those three, Mitch, right? Or not those three.
I think it's everybody. That brings us to the- Right ... that brings us to our next segment- Yeah ...
which I want to jump into. But before we do, I do want to mention that one of the things Futurum's a little different, Mitch's report and article is not behind some paywall that you got to be a subscriber for even. You can go to the Futurum Group, look at the research notes, and grab Mitchell's paper complete there, as well as my response to it, I believe, on Techstrong AI.
But speaking of what everyone can jump into, right? Look, we had a big week last week. " But that's the world we find ourselves in.
And then on top of that, we got some announcements this week of the latest open weight models from China being roughly on par with the three you mentioned, Mitch. John, this was a story you had helped start, right? Yes.
So I started off with an emphasis on Inkling, and then you went with Tinker, which you think is more significant, but I'll let you talk about that later. I'll just do it really quickly. So Thinking Machines, which is the startup that was founded by the former CTO of OpenAI, which is Mira Murati, they announced the release of Inkling, which is this massive open-weight AI model.
And in a sense, it's basically this viable alternative to the dominant open source models currently coming out of Chinese AI labs. It's interesting because Inkling enters this market where we've lagged behind Alibaba's Qwen. By the way, we talk about how fast this news cycle is moving.
8 at the World Artificial Intelligence Conference in Shanghai, claiming it's second only to Claude Fable 5 among the frontier models. So we have this kind of arms race or open source race which is going on. In a sense, my big takeaway is Inkling is this calculated effort to revitalize the Western open-weight AI ecosystem.
But I know, Allen, you think that the actual bet that is more significant is Tinker, and I wanted to know more because that intrigued me as well. Yeah, no, it is. I'm going to jump on that, John, but forgive me a little bit.
Not just the Qwen model. There's also the Kimi. Was it Kimi?
Yes. Is... I forgot.
AI. AI. That also claims it's just below Fable, but ahead of- That also happened, and it's really hard.
Yeah. As you said at the beginning of the show, it's so hard to keep track of this, but things are moving. In the last two days, both of them claim to be between OpenAI and Fable, and they're open weight.
So, they're not as proprietary or as closed off as, let's say, some of the US frontier models are. This is a serious concern, right? And I'll throw it out to the rest of the panel.
The numbers I see is at least 30% of the market, including the US, are using these open-weight models because they're cheaper. They're cheaper to use. There's no if, ands, or buts, or tokens about it.
What does this mean? Yvette, what do you think? So, you touched on it, Allen.
So I'm going to go back to the sharks and the jets. I'm dating myself, but the- No, there was a remake. You probably saw the remake by the time you reached there.
Inkling is, for me, Inkling is the razor, and Tinker is the blades, and let me give you my reasoning why. Because once the weights are public, nobody who downloads them owes Thinking Machines a dime. So the model literally cannot be- Mm-hmm ...
the business. The business is the fine-tuning platform, and I think that's her bet. She's betting that the future isn't going to be one god model.
It's 10,000 specialized little ones, and she wants to sell the factory, not the product. And if you go back to what we were talking about with "The Indispensable," right? Mm-hmm.
If 56% of enterprises is multi-provider stats, that's her best marketing. If the models are now commoditized, the money's in the customization infrastructure. Right.
The models- Right ... they became indispensable, but commodity. Yeah.
And then one last thing. No one's saying this, but you literally need two terabytes of GPU memory, and that ain't open in any way where it matters most to companies, right? I think with BF16, that checkpoint needs roughly two terabytes of aggregated VRAM, and even when the quantized version needs about 600 gigabytes.
So that's not democratization at all. That's like a velvet freaking rope. You all want to line up and you can't get in, right?
Because not everybody can use that. So, yeah, that's free, and I got a Brooklyn Bridge to sell you that I can sell you today for no money down. Absolutely, if that's free.
Liz, you were going to say something? And some Florida property to go with it. Yeah.
Well, I also think that she's leaving it in the hands of the customer. If you talk about the models are being commoditized, then the end user has control over that governance and control plane layer, right? Yes.
Which is what I think is mm-hmm. And Liz, spot on, because now open-weight transfers the safety liability as well. So TechCrunch kind of peeped this.
So when customers fine-tune that stuff, customers own all the safety outcomes, and the closed labs, they sell you a model, and they absorb the reputational blast radius. But Thinking Machines is selling you the model and handing you the blast radius. Well, the interesting- So I'm going to make a statement here, Allen.
Go ahead. I'm going to make a statement. Multi-model won, period, is already won.
That war, that decision is over. Oh, there's no doubt that multi-model is won. And I would contend that local models as well as cloud-based models have won also.
And the reason being is tip of the spear for AI, you've heard me say this for a long time now, Allen and John, is development. It's what happens in development is what escalates and grows into the other parts of the life cycle. What does every developer do?
They download and run their own models. Yvette, your point is spot on. I can only run so much on a unified RAM and GPU, CPU on a Mac or on a NVIDIA card, but I can run some pretty good models in an eight gig, 12 gig card.
Then I can do a lot of things locally. You can do more in the data center. Mitchell, what you're talking about, though, is Do I need these latest, greatest models to do most of my work?
Probably not. Well, I think no, you don't. You absolutely don't.
Right. You don't want a fable for everything. You don't need the latest, greatest model for everything.
And in a way, this is not only the razor blades and the razor, it's sort of like we're going to replace the razor and the razor blades every month. Mm-hmm. It just all transforms all the time.
That's the definition of modeling. What a consumer economy that is, huh? Yeah.
But let me come back to John's question to me. Why did I think that it wasn't Inkling, right, that that was the real deal here, but that it was Tinker. Right?
That's the name of it. Because Tinker's an upstack move. Tinker says, "I don't care what model you run on.
" As a matter of fact, the example they cited was Tinker running on top of... John, which is the one that just came out from Alibaba Cloud? Qwen.
Qwen. 8, I guess. Yeah.
Well, it was running, I think, the prior version of Qwen. Right. Previous version.
But really, and Mitch, to your point about local, that's what Tinker does. It says, "I don't care what universal world model," that's a special term here, "but I don't care what indispensable model you use here. " And it doesn't get to that world model.
You don't lose your data. You get to keep your data, right? And you can run that locally.
You don't need to run it right back at the big cloud or the hypervisor, or with this model maker who's sucking your data in to train the next version of their model. And so that, to me, is what makes Tinker so special, is- Yeah ... it runs on everyone's model.
So no matter what that indispensable layer of commodity becomes, Tinker runs on top of it. So, in the simplest, and maybe I'm oversimplifying this, but what Thinking Machines is doing is rather than competing purely for the title of "I'm the strongest general chatbot," whatever you want to call it, via traditional metered API fees, they're betting on customizability and infrastructure control. Is that- That's basically it, and the local aspect.
Like Mitch said, local one, you could run that local. Okay. But let me come back to the bigger issue here, guys.
No one here sees the irony of China saying the US is a closed system and we should be open? Well, my thought is that Anthropic is accusing Moonshot of espionage, basically, or stealing their IP. Well, yeah, I know.
Just like the 2020 election. Or distillation maybe. Let's not go there.
They accused them of distillation. They've accused every Chinese- Every, yes ... model maker of stealing their stuff or distillation from the get go, right?
At some point, and that's pretty ironic too, considering that OpenAI and Anthropic stole the whole internet- Right ... to do their own models. Pretty rich.
Yeah. My dad's actually part of that class action case. Yeah.
Yeah. Expecting a check for three grand at some point. But what does it mean?
Are we living in bizarro world where China is the beacon of freedom and choice and openness? Or are they just doing that because that's the cards they were dealt in this poker game, and as soon as they get a stronger hand, they're going to clamp that down like they do everything else in that society? Think about the theory of constraints, right?
What was China constrained on, and that was access to GPUs, though they got some, right? So what did they emphasize? Efficiency, speed, memory, things that they could control, and that's what got them into the game.
And now they're a player, and they can claim open if they want, if they're doing open, doing open weights at least. But that's also sort of a thumb in the eye of the US companies from a geopolitical standpoint, right? Yeah, that's what I thought too.
That's just you guys, right? This is about all geopolitics. Mm-hmm.
Well, no, it is. It always comes back to that. Yvette, go ahead.
What do you think? So my take is that China's not trying to build the best AI. They're trying to become the AI everyone else builds on, and those are two completely wars, and we keep scoring the wrong one.
So at the end of the day, US, why we have no guardrails, no guidelines, we're competing to build the best AI, and China is competing to be the best substrate. 8 trillion parameters. Weights are promised July 27th.
It would be the largest open weight release in history, and history is a day. Okay? Right?
So China doesn't need the crown. It just needs to be close enough to it so that the other 150 countries on the planet pick the free option, and that's not charity. That's the Android playbook run at nation state scale.
That's exactly it. You hit it on the head. Well said.
Wow. It is. This is what they were dealt, right?
They weren't going to beat the US by having more GPUs. Yeah. They weren't going to beat the US by maybe investing more trillions of dollars.
So they took the open source route, which is a well-known route in technology. We've seen it, right? How did Linux win?
How did Apache win? How did Android win? Pick any open source.
Yeah. OpenTelemetry, any of these. And they played it.
But now I think you're coming to the point where they're no longer forced to play that catch-up game. I think the world knows good enough Trump's best when we put dollars and cents next to it for most things. So do you foresee them pulling back?
Because Xi, the Chinese premier, president, whatever, he's intimated as much that- Mm ... these models may be becoming so powerful and dangerous, they may not continue to keep them open. Yeah.
And open means not open source. That would be an ultimate bait and switch. Somebody in the "New York Times," it was kind of an interesting comparison.
They talked about the global AI race in general, and they said that the US approach is always like a nuclear weapons race, and the Chinese think of it more of like a nuclear energy type of approach. It's just two totally different approaches. And it's- That's an interesting way of looking at it, too.
You're right, John Think about what happened with the A-bomb, Oppenheimer, right? The reason why, whoever builds it first is going to use it first. " And that's exactly what happened, and I do think that Xi's speech was a split screen, and the Global South was his audience- Yeah ...
if you think about it. " And if you think about the timing, it wasn't subtle because it landed right after whose primetime broadcast? Yeah.
Accusing people of compromising and election data. So Washington performed security, Beijing performed inclusion, and they are both auditioning for the same 150 countries. So think about it that way.
We, here in the US, I don't know what allies we have left. But everyone is looking to unhitch their wagon to US tech. But that's my point exactly.
I don't understand how that's not geopolitical. What am I missing? Oh, it is.
Oh, it's totally geopolitical. No, I understand. I don't think I- But it's business geopolitical.
It's about dollars and cents, and look, they're playing it well. But we've got to move on. I want to turn to our third segment, to we always like to do a little heavier on the cyber here.
And we've got OpenAI. AI war games. So it wasn't enough that the AI, Mythos and all these things are finding vulnerabilities.
Now we're using AIs as red teams. Liz, you've got this story. What have we got here?
Yeah. Well, I actually find this a little bit heartening in the never a glimmer of good news world of security, that they're red teaming models before they're getting out, and that the red teaming seems to be bearing some fruits. Like the GPT human red teamers, 84 to 13% on novel attack scenarios.
That sounds pretty good to me. And even just from the vendor community, autonomous pen testing is now the new normal. That's what everybody's doing.
That's what Mythos brought. Is that the response to that is autonomous pen testing. And so that we're doing it before the models get released I think is smart, because it says here that ChatGPT's own models were particularly vulnerable to GPT-Red's prompt injection attacks.
That's kind of scary, but it feels to me like we're moving more towards the right direction of getting these things more secure. Yvette, I can't tell if you agree or if you're dying to disagree with me. Your assessment is spot on, I'll give you that, but let's zoom out.
All right? So this announcement is also a confession, Liz, because over 90% of GPT red team's attacks worked against GPT-5, the model the world is running in production since last fall. So the security posture of the agents everyone deployed in 2025 was, by their own admission and measurement, tissue paper.
Well, but that's the way of new technology, right? But upgrade. I've said this probably before.
What's the answer? Upgrade. But honestly, we should know this.
And you know what? I don't know why we give them so much trust in the first place. I mean, like- Because we always do.
" I was writing an open source column. And I did an interview with the CEO of MongoDB and Couchbase, two very, at the time, open- Early ... SQL database.
Early, yeah. And I asked them, Mogul was on it with me, Rich Mogul. And I asked the two CEOs, "A lot of people say no SQL stands for no security.
" The fact that all these people used this AI and didn't demand it tells you all you need to know. Well, that's my whole point. That's why it's buyer beware, right?
I also want to step back and look at the system of what's happening. So step back for a moment. Alan's heard me say, John's heard me say, the only place you solve security is at the point of origin, is at the place where software's created, where prompts are created, where they're executed.
Everything else is forensics. Everything else is, it's already happened, and it's subject to those injections or whatever it is. So when you look at the system, what's happening is the market is forcing pressure on the model makers because of Fable and all the things that led up to it, but other reasons to say, "We need these to be more secure," right?
They're not announcing these things, OpenAI and Anthropic and others, out of the goodness of their hearts. They're responding to the market, and they're maybe glossing over GPT-5 of like, "Oh, don't look over here. " But the point being is, this is an example of solving it at the point of origin.
Because if I can solve it at the point that the prompts could get injected but now can't, all right, that's another security issue I don't have to fix later. And when you operate at a very high scale, say 10X from the speed that we do today, no manual process will ever keep up. So we have to have these kind of solutions for AI to be successful.
And not only that, but I'm also hoping that if you do that red teaming there, then the architecture of any harness or anything else will be more secure. Because I deal, I work with startups all the time. I can't tell you how much threat research I see across the board where it's not just a vulnerability in a model, it's a vulnerability in the architecture that allows something to get prompt injected or hijacked, right?
And so if you do this sooner, if you shift all this left, then maybe that won't be so easy for all these researchers to do. Here was my take on, and none of you have mentioned it. This is an internal-only tool from OpenAI that they're using to test their own model and other models.
Obviously, the thought being that we can't let this into the wild because people will be able to break the models and do these things. So my question, John, I'll throw it to you. You know what they say in "Jurassic Park," right?
Yes. Life will find a way. Yeah.
What happens if you- Models will find a way. Yeah. You don't think this will get out, or someone won't be able to use this to build the dream crate?
I know, exactly. What could possibly go wrong, Alan? Yeah.
When I see OpenAI and I see anything related to security, I always cringe. I'm sorry. I have such a- ...
wall about this company. I don't trust... Is it okay to say this?
Because I mentioned, and Alan- Sure, you can say whatever you want, John. You have insurance ... probably said it during this.
It's a Pavlovian response for you, John. It's from all the training they've done with you. Yeah.
Just make sure you caveat, it's not the official response of Techstrong, but go ahead. Anybody with a... I'm overstating it, but Mira left for a reason.
They all left for one reason or another. And then when I see OpenAI trying to play the white hat, I always just smirk when I see this, and it's a market play, it's a defensive play. " The stock market play.
And so I always think it's always calculated, in my opinion. Well, look, I don't want to give away tomorrow's show, but another one of the big AI models got caught with his hand in the cookie jar this past week, sucking customers' information in when they swore they didn't, and it was none other than Slippery Sam or Scam or whatever you want to call him over. Sham.
Or Sham the Sham. Good one. Right?
Sam the Sham, who brought it out and said, "Oh, this is concerning," to his favorite foil. Yeah. I'm reading "Regime Change," and there is some really interesting insight into AI politics and AI foundation, and just cutting to the chase, Greg Brockman met with Trump early on.
And Trump, to his credit, sees the problem with what was happening, and he looked at this, and I'm thinking as kind of a manufacturing boom, right? This is before all the... What was that project, Stargate?
Whatever happened to that? Well, it's still there. It's still there.
It's worth the 25 jobs it said it would. Okay. It's always hard to keep track of these promises.
But just going back to this, it's just interesting the dynamic, and Altman always is squarely, I think with Zuckerberg, just squarely on the wrong side of history. That's just all I'm... I can't think of any other way.
I was talking to a friend of mine in the PR business. Not Liz. I have other friends in the PR business.
You do? Yeah. But Liz, not as good as you.
You shouldn't disclose those things in an open show. This is an MS, is it? No, but they said, I'd want to be whoever Anthropic uses for PR.
They've done a tremendous job of capturing the good guy. Right? They're the people with a heart, the people who seem to care.
With a conscience. Yeah, with a conscience. Yes.
That's knowing. That's like the Apple approach, or was. Yeah.
It was. Well, I mean- It was. Right ...
a conscience, but they also have a crisis of conscience when they have to make money, right? Yeah. They of course have that.
I didn't say they had a conscience. That's just me. Well.
Yeah. So- They're able to project the appearance of one. Yeah.
Yeah. Yeah. So- That's the key.
Yeah ... here's the one thing. So Alan, you mentioned what happens, this is internal only.
Well, how internal really? So the model may be internal only, but didn't they just publish a paper describing the training methodology? Isn't that public?
Yeah. So, you can self-play it out for attack generation. It's now a published recipe.
So every capable state actor and well-funded criminal organization just got the blueprint minus the weights. So the moat is compute and time, and both just eroded. So it's not a matter of if, it's going to be when.
Fair enough. I think we're going to end it right there because we're about out of time. Yvette, Liz, Mitch, John, thank you for joining us today.
Thank you for watching. We hope you enjoyed this conversation. I'll be very honest with you, I think we'd have this conversation whether you watched it or not, but we appreciate you watching it, and we hope you find it interesting.
If you do like it, we do this every weekday, Monday to Friday at noon, on the Techstrong network. If you can't catch it at noon because you got a job or something, or you don't want to ruin lunch- ... tv on demand, on the Techstrong TV YouTube channel, the Techstrong TV OTT channel, available just about on any screen, whether it's iOS or Android or Apple, Roku, Amazon.
It's available there. And of course, we have Techstrong TV that plays throughout most of the rest of the day, so you can check out good videos there. Or if reading is your thing, we do have about a dozen different websites you can go read.
A lot of these stories are based on articles that come from there. We will be back tomorrow with more Techstrong gang. But for now, on behalf of our esteemed panel today, we're out.
Have a great day, everyone. Bye.


