Veracode’s 2025 GenAI Code Security Report: GPT-5 Models Lead in Secure Coding Performance
Veracode’s CTO Jens Wessling discusses the 2025 GenAI Code Security Report which reveals major differences in LLM code security. Testing 20 leading models, OpenAI’s GPT-5 Mini and GPT-5 achieved the highest secure-code scores, outperforming Claude Sonnet 4.5 and Gemini 2.5 Pro. The findings highlight new risks for AI-driven development
Transcript
Hey everyone, it's Alan Shimel and welcome back here to Techstrong tv. Excuse me. My next guest is he's frequent guest of ours, Jens Wessling at, uh, Veracode Ys.
I was just checking to make sure I got your name right, but I did, and, uh, every once in a while though, you get that moment of panic. Right. Did I get his name right?
Ys, of course, is, uh, with Veracode, as I mentioned, and he's been with Veracode a while. It will let him introduce himself. Ys, it's great to see you again.
I hope all is well. Yeah, it's wonderful to see you again too, Alan. It's, uh, For those Interesting times.
Absolutely. Well, may you live in interesting times, is what they say. Right.
For those who, um, maybe have not seen you on here before though, why don't, if you don't mind, give us a little background. Yeah. Um, I'm the, I run up all sort of, all architecture and research for Veracode and I've been there for a while now and I've sort of kicked off the project, uh, started evaluating the security of large language models, Which is a hot topic, certainly, and of course wear code being a leader in the AppSec space.
This is something kind of near and dear to them. Um, yeah, it's just before we get into the findings of, of this particular survey and research, um, any idea if I had to ask you what percentage of code being generated today is being generated by ai? Oh, I think it's so hard to tell anymore, but it's certainly climbing by leaps and bounds day over day.
I think, uh, probably rare to find a lot of developers these days that aren't using AI in some way, shape or form in their development. Agreed. Agreed.
I mean, I, I saw something that said 90% of developers actually are using AI before Yeah. At though 40% don't trust it. I'm surprised it's not a little higher than that, but Yeah.
Yeah. Well, 65, 60 5% absolutely believe it. It, it, it creates instabilities.
It's still, it's a, it is early days yet we'll have to see where things go. I think it's leaps and bounds ahead of where it was six months ago. It mostly wasn't a thing a year ago, so who knows where it'll be in six months, but I think there's gonna be an upper limit to how effective it can do what it needs to do.
And I don't see a world in which you won't need a human verifying that what it's doing is what you want it to do. So the human in the loop. Well, I mean, and, and that's kind of the subject we're gonna talk about today, but it seems with every new release and between the frontier models, there's new releases, it seems coming out every week or so.
Um, it seems with every new release, the the accuracy, the quality of the code that these things are producing is, is improving. So interesting there. But before we jump into that, kind of the specifics there, Jens can set the table for us.
You, you guys, I mean, Veracode is known for some great reports, a lot of based upon their own customer data that, you know, anonymously anonymized, they've used for, you know, tremendous data sets. Yeah. What about, so Yeah, about a year ago we started looking over the research on how secure LM generated code was.
And what we discovered is sort of depending on where you looked and who you asked, you can get a lot of different answers. And even more than that, they'd give you an answer that's a snapshot in time. And as you said, every time these models change, the results differ.
And what we really needed is something that didn't exist, which is a report that gave a, a very candid like description of how well they were performing with respect to security that got updated regularly. So you could actually see whether or not things are moving in a given direction or not. And we did our first report last spring and we did our most recent update in October with the latest frontier models and, uh, interesting results.
I mean, I think the, the chat GPT reasoning models showed a significant uptick, roughly 10% better than they did prior up into say the low 70 percentile. And basically every other frontier model that had been released over the prior six months didn't actually show any improvement from a security standpoint. Kind of interesting.
So I think Chatt PT did a nice job of sort of focusing on that as one of their key points, and they managed to deliver a probably the largest jump we've seen in the history of the, the reporting data we pulled. Really? Yeah.
Which this is five one, This is five one. The reasoning models, the, the chat model, the non reasoning model didn't do any better than their prior version really, but the, yeah, the, the five and the five mini both showed a significant uptick roughly 20% better than their prior model had done. What do you, what do you attribute the, the, the regular non reasoning model, if we call it that, not having any improvement versus the big improvement of the reasoning model?
I think, so I'll offer some guesses, but I, you'd really have to be in-house at, uh, OpenAI to know the answer. But reasoning models will take multiple passes through a problem considering sort of different perspectives on how to accomplish their goals. And if you develop a code generating reasoning model that actually has a past that's concerned with security, it could actually make a significant improvement in how well it does in that respect.
And I know they've actually spoken at, at length about how they've invested in security and they put some adversarial models out there and they've, they've tried to improve the quality of what they've done and I think it shows that that work has paid off. And the reasoning model is considering these things now when it's putting together code samples. Good.
Yeah. Let me ask the next question, which is, uh, as I, I think I alluded to earlier in our discussion, you know, there's a new model coming or a new version of the models, you know, collectively coming every weeks, every few weeks or whatever. And it seems, well, except for this non reasoning open model, each one gets a, at least a little bit better.
Is it really just a question of whoever, whoever went last is the best at this point? Do you see that leveling off at some point? Well, no, I, I mean I don't think that's the case with respect to application security.
I think when we look at the history of these over the last year, the overall improvement in the average has been a few, two, 3% over the last year. And the open a i, there are, there are more recent frontier models than the OpenAI ones than aren't performing as well as theirs have. So I think it's more than just, we hope it gets better 'cause it's not getting better fast enough.
I think you need to make an investment in that as you're producing these models, if you actually want to see improvement and even with all the investment they put in, they're hitting 70% and while that's much better than 50%, which a lot of them are running at, it's still not good enough that you trust it to go to production without actually testing it to make sure that it was secure. Just, just for reference, what percentage is let's say human generated code usually? So that's tough.
I think it's normally around 60 to 70% humans are also not great and it makes sense 'cause the models are training in human data. Well Then, you know, you're always as good as what you needed. Right.
Um, so but what you then that but that in itself is interesting. You're saying at this point yeah, they're about as good as a human coder. Yeah.
And I think there's a long history of human coders not writing secure code. And I think the same thing is gonna be true of LLMs and that's why even if they're, you know, modestly better than humans, it still doesn't mean you get a blank check to release code without being concerned about security. You know, I mean in all honesty, I don't think anyone sits here and says, we're just gonna release this without first checking it, run it through some security tests and everything.
I think we recognize that. I think the question is how diligent are we with this code and with our testing, how rigorous is the testing? I I've seen a lot of organizations where they're using one AI to generate the code and another AI to test the code and there's no human in the loop in that equation.
Right. And, and though I get the, the logic of it right, this AI looks at it differently than that AI or what have you. It still just doesn't sit well with me to tell you the truth.
I don't, I think the final validation of security should not be left to anything that routinely hallucinates. Mm-hmm. I I get it, I get it.
But you know, to be fair, the flip side of this is, um, the pressure from above. Even use ai, experiment with ai, use ai, get your code out, keep up, move more, publish more, right. Push more.
It's, it, it makes it harder and harder to kinda do the right thing, right? It can be, I mean I think AI is a useful tool that makes sense for developers to leverage, but like any tool like not indiscriminately, they're, they're situations where it's the right tool and situations where it's the wrong tool and understanding the difference is what sort of separates the novice from the master. Fair enough.
Jen Jens, if you don't mind, I'd like to go back to, to the report though. 'cause we, we went off you and I talking about what, you know, we find it interesting. What else in the report though might be of interest to our listeners watchers View?
This is the first time we've actually broken out the security sort of by model. The first one was just sort of the aggregate and we thought it was important to do this time is we did actually for the first time see a pretty significant difference in a few of the, the models. So if security is your top priority, I think the open A models AI reasoning models are probably the, the best in the business right now.
And I think if that's where your advertisement, if you want to have an LM that's reviewing code for security, that would probably be a good choice as opposed to some of the other models that don't even hit 50%. Hmm, absolutely. Um, the models that don't even hit 50%, do you think that's just a way point on the way to getting better for them or do you think those are just inferior models when it comes to security?
I mean, for instance, one of the models that hit 49% was Anthropics Claude Opus four one, which just came out, which Considered a good a good coding Machine. It's a, it's an excellent model for writing code but not necessarily for writing the most secure code. And I think when we look at some of these things over time, we don't necessarily see the direction is consistently linear in a positive direction.
Um, I think like Opus four was 50%, four, one was 49%, four five is back to 50%. It's sort of flats. And I think if you don't invest in making that better, it's probably not just going to improve because you've built another model.
You know, a lesson I learned early on in my security career was they'll, they'll make it better when customers demand it to be better. And I, I quite frankly, again, I think that's part of the problem here. Customers want to use these things to generate code.
They know the code isn't of the highest security quality, they don't trust it, a good chip percentage of them, however they still use it and they're still doing it. So they're not demanding better security. I mean, it's not necessarily worse than what they get from their human developers now.
So I, I understand why they feel the way they do, but I also, and I'm, I'm hopeful because I think the first focus for a lot of these large language models with co-generation is getting it to generate the right code correctly and consistently. And I think we're just getting to that point now and I think once you get to that point, then you can start asking for more. 'cause having code that's more secure but still doesn't function isn't really a win.
Agreed. I I, I would have to agree with you. Yeah.
Gens anything in the report or findings that you kinda say, geez, that wasn't on the bingo card. Um, I think that the GPT model just being way out of band and the highest growths we've ever seen is probably the biggest thing. net seem to be steadily improving with respect to security and less so with some of the other languages that we've looked at.
And then I think this model, we actually did a breakdown of reasoning models versus non reasoning models. And the reasoning models seem to on average do about five points better than the non reasoning models. That's significant.
That's significant. It is significant, yes. And I think a lot of the AI coding assistant are leaning more towards reasoning models over time.
I I, I do agree with you on that as well. Um, correlation between how good it is on code versus how good it is on making secure code. Mm-hmm.
Right. Like you, we brought up the anthropic model, right? A lot of developers swear by anthropic for code according to this.
They may swear by it, but that doesn't mean it's very secure code. Yeah. And, and, and again, I mean the, the, the delta here between the most secure code being, uh, generated in, you know, middle of the pack is is what I mean the, the big delta is 70 to 50, but I imagine Yeah, a lot, a lot of 'em are a lot closer.
So one of the, the evaluations is we looked at how good a job they did at producing syntactically correct code, like does what they produce compile and October of 2024, it was roughly 50% of the code they generated would compile successfully and the rest of it wouldn't. 5% for creating syntactically correct code, which they can now do extremely consistently, but only four or 5% better with respect to security. So I I I've heard this before, you know, and my, my take on it is that narrow use case, whatever you want to call it, of making syntax correct, is something that they cracked the milk, pun intended, they cracked the code on and we and figured it out pretty well, hence the, they're really good.
But that, I think now that we're at almost a hundred percent, I think people are gonna start asking questions like, okay, now it's intact the correct, it's doing what we want, but is it secure? And I think now they got their first ask, which is, does it work at all? And now it's like, okay, now how do we make it work better?
So I, I'm hopeful that we're gonna see them investing more in the security side of things. So the code they produce is a little more trustworthy. When will you be testing it?
Uh, it's a bit of wait and see. We sort of wait till another whole swath of frontier models come out again and I think the latest batch is planning to come out early next year. So as soon as they do, like we run the, the report regularly, we just generally don't publish the results until we have enough additional data that actually says something new.
'cause there's not much fun for me to get on here with an interview and say it looks exactly the same as last time, The same speed ahead. You know what, speaking of the report though, Jens, where can people get that? com, we've got a link to the report on the front page and uh, it's a very interesting read.
I think, uh, the original report and the extension with the October updates are important reading for anybody that's serious about security and you know, amongst the 90% that are using a to help them generate their code. I love it. Alright.
Yeah, we're about outta time here. I want to thank you for coming on as always. I'm sure we'll be seeing you soon.
Maybe ours RSA is only like three, four months away now. Yeah. Always a pleasure.
Alan. Hopefully I'll run into you at RSA. Yes.
Well, we'll, you know, don't leave it to chance. We'll make it happen. I someone who Sounds good With us Yeah.
And say hello to everyone and Veracode, keep up the great work on this. Hey, we are making progress. I'll, I'll say that.
Yes. And that Things are getting better. I like it.
All right, thanks Alan. And take care. Okay.
Jens Wessling Veracode here on text tv. We're gonna take a break. We'll be right back.