Exploring AI’s Impact on Application Security with Veracode’s Jens Wessling
Jens Wessling, Chief Technology Officer of Veracode discuss the impact on application security and a recent report on AI’s influence in the field. Jens highlight advancements in AI code generation, including models like ChatGPT 5.0 and Claude, while addressing concerns about the security of AI-generated code. The importance of human oversight in identifying vulnerabilities is emphasized, alongside the potential of AI to enhance security systems and the need for secure coding practices.
Transcript
Hey everyone. Welcome back here to Techstrong tv. You know, I was just a black hat last week and I, I didn't get a chance to see this gentleman.
I'm glad we getting to catch up so soon after. Let me introduce you to, uh, Jens Wessling. Jens is of course, chief Technology Officer at Veracode.
He's been on here before, but not often enough. But Jens, it's great to see you here. How are you?
I'm doing great. It's good to see you too. It's been a minute.
Yes, it has. I think the last time we had you on, you had just become CTO. Yep.
It was about a year ago, so, so yeah, it's been a, Yeah. How, how's it, how's it been going? It's been going fantastic.
It's, uh, been doing a lot of research and innovation stuff and as AI and the market changes, there's been a lot of opportunities and a lot of discoveries, and it's kind of part of what we're here to talk about today. God knows. God knows.
Yeah. It's, most people in our audience are familiar with Veracode. We've been covering Veracode, I think, since I started Security Boulevard.
com. But in case someone out here is not familiar, give them a little Veracode background. Uh, so Veracode, uh, specialize in, in securing software.
We're an application security specialty company, and we do static scanning and dynamic scanning and security component analysis. And we're getting more and more into ai. And we had the, the pers automated security remediation product in the market, and we sort continue to press on with innovation and doing new and interesting things.
Excellent. All right. com, find it all there.
Absolutely. Yep. You know, yeah.
It's Veracode's research team has always been a, uh, just a stellar part of the company, a stellar part of the industry. Uh, absolutely. The annual security report from Veracode is, is kind of a, you know, a benchmark.
You guys did a new report for this year though, and, uh, why don't you tell us a little about it? Yeah. So as AI has taken a bigger and bigger role in our industry, it became clear and clear that we didn't necessarily have the answers to all the questions we wanted.
And if you look around at all the reports people do on AI and security, depending on who produces the report and what you're asking, you'll get a lot of different answers. And a lot of the answers were sort of snapshot in time answers and not really we're looking for. So we really wanted to sit down and put together a report that would look at the state of ai, the impact it has on security and software and application development, and put the data together that we needed to make the, the important decisions that we needed strategically to move forward.
And as we put it together, we said, this seems like it'd be a direct resource for a lot of people. So we actually put a report together and released our results. And we sort of, I think we shared that at Black Hat and we're here talking about that today.
Absolutely. AB and look, this is, I think, a very timely report, right? 0, which is supposedly a little better with coding than its predecessor.
Of course, Claude has a, a, a, well, it's not new anymore because if it's over a week old, it's not new anymore in ai. Yeah. I came out on the fifth, I mean, that was whole six, well, six days ago.
We can call it new for another day. It's crazy, isn't it? Yeah.
It's, but uh, but Claude has a new version out Gemini keeps pumping out new stuff. We've got, you know, the promise of, of, uh, AI browsers coming out. But clearly the race is on in, in terms of what AI is best for coding.
Right. Everybody's talking about vibe, coding and coding and, you know, which AI is the best for it. Um, as we were talking off camera, I don't know if we're ever gonna have a definitive answer to that, but what we do know is there's a lot of people generating a lot of code using ai.
Yes. Mm-hmm. And, and while you know, I'm, I'm never one to say, oh, we need less code.
Right. More code's good. I guess, um, secure code would be nice.
And I think that's, It would be, That's where, that's where rubber meets the road here on this report. Yeah. That's one the big questions we wanted to answer.
It's like, is the code generated by AI Secure as a whole? And we look back through a couple years and clearly it started out a little chaotic and it's gotten a lot more consistent over the last year or so. But what we've discovered is that they've sort of, there's very clear trends.
So their ability to produce syntactically correct code has increased by leaps and bounds. Yes, you should expect it to be in the high 90 percentiles, but as far as secure code goes, sort of entered in around 60% Correct. Code, which is not really where you'd like it.
Right. Especially you're producing a lot more volume of code. If it's only right.
Six times in 10 and producing something that's secure, you're producing twice as much code, you have a lot more problems than you did a month ago. Right. I mean, look, the only analogy I could tell you with that is when you, you know, forever and ever, it's been my dream that I don't have to type Yes.
And I could dictate, right? Yep. And transcription and, and, you know, even nine outta 10, which sounds great.
Well go back in and change every 10th word. It's, It's a pain. It's a lot pain in the butt.
Yeah. Six outta 10 don't even bother. I'm doing that more than I am co.
You know what I mean? I, I just, it's not useful at all. Um, is it really that load, Jens, In some cases it's worse.
It, it really is that low. And we also evaluated across different programming languages and across some of the most important cws that we run into and over time with sort of different models and different sizes and the size and the, the recency of the model doesn't seem to have a significant impact on its ability to produce secure code, which was a bit of a surprise to us. 'cause we expect it to move similar to how Syntex did, where it just keeps getting better.
But I think a lot of the code that's out there isn't by default secure. And that's what it's training on. So it's not really getting more secure than what you see in the wild.
Huh. Like, programmers aren't gonna check in code that doesn't compile. They can tell that right away, but they'll check in code that's not secure and maybe never know.
Yeah. You know, it's interesting, the last week while I was outta black hat, everybody there was touting some new study or another and there was such a discrepancy in dichotomy between them. Right.
You know, I, I saw one, I saw one, uh, study that said, oh, by having AI do the code we're increasing code production. I forgot if it was 60% or some crazy number, 40%, you know, not only increasing pro pro code production, but making developers more productive. Yep.
Right. Then I saw another study that said, no, the fact of the matter is using AI to generate code makes developers 19 or 20% less Productive. Yeah, I saw, I saw the same study and the part of the reason we did our own report, because if you look at it, if you can find whatever results you want out there, and usually it's someone that well Yeah.
Interest in a particular result. Yeah. Yeah.
But, but you gotta ask yourself, look, if I have to correct six out of every 10, there's no way I'm more productive. Excuse me, four out of every 10. Yeah.
It, it's tough. I mean, what we don't have is like, if the same situation, if you're given the same tasks to a developer, what would their accuracy have been? Well, that, that is interesting.
Right. And, you know, 'cause for instance, Microsoft did a a, a study or they released some numbers about, I guess was over a month ago that this new product they have for, for diagnosing medical conditions was four x more accurate than human doctors. So I think it was like about 80% accuracy on diagnoses, you know, for things out of like the New England Journal of Medicine, but real life doc, you know, human doctors were only 20% or 21%, something like that.
So, you know, you had a good comparison. Humans here, machine there, we don't have that with the 40 60 here. Right.
We don't, is that in par for the course? I don't know. But what we do know is that whether humans are 10% or 90% accurate, you don't want 40% of your code going with security implications going out with a, a vulnerability in it.
Right? Yep. Agreed.
So Agree it human, It's gonna do something. You gotta make sure that it's secure. And that's sort of, AI is not fixing that problem for us.
We still need to make sure that we're addressing the security issues and they're getting remediated. So here's something that I've seen though, in speaking to people with AI in security. If you use AI to generate your code, you should have a human check it for vulnerabilities.
If you use a human to generate your code, you should have an AI check it for vulnerabilities. What do you think about that? I don't trust anything to scan for vulnerabilities that might occasionally hallucinate on me.
Okay. Well, humans have been known to hallucinate, right? I mean, usually drug induced, but nevertheless, That's why we had security scanners that are trying to keep us all honest, that aren't gonna hallucinate, are gonna make sure they find all the issues and none of the false positives and sort of help humans do that job effectively.
AG agentic AI could run scanners. Yes, that's already going out. And we have products in the works for doing exactly that kind of thing.
And I think that provides a lot of value, but you need the security expertise to train the agents to make sure that they're looking for the right things and doing the right things and returning the right results. Fair. Um, yeah, it's just reading kind of the headlines here is you, you did this by testing over 100 LLMs.
Mm-hmm. So that sounds like it goes well beyond the frontier models that most of us are familiar with. What other, you know, LLM models, if you can talk about were, were part of this study.
I mean also keep in mind that we're testing all of the models from open a AP OpenAI for instance, and breach release of chat GPT. There might be three or four different models they release as part of that. They'll have their, their mini model or their nano model and their chat model and sort of, that sort of adds up to being quite a few of them.
And we sort of grabbed a lot of the bigger open source models, um, deep seek. So it is mostly the pro the frontier models, if we can call it that. Those are certainly the most important ones in there.
And we made sure that all of those were covered and we few others in there. But yes, we weren't searching out edge case models that were hardly used to do the evaluation. This was, we made sure all the mainstream models were covered and evaluated.
So in that 40 60 number, is that including like chat GPT 3 0 1 oh like older models where we knew Oh yeah. All the way back to, waited All the way back to one and you actually see, when you look at the chart, they started out quite terrible at syntax. And then over time you see that they sort of asked em to, to giving really good syntax numbers and roughly 60%.
So we sort of charted over time. But then the important bits is we wanna know where the industry's going. Like is it on track to solve this problem itself or is it not?
And like obviously we need to know that right? As this application security company. So we did the evaluation and they, we hope they get better over time.
Like I think they're starting to focus more on security now, but we haven't seen it yet. And I think certainly with the volume of code we're producing, we're producing twice as much code, even if it's half as vulnerable as, uh, a software engineer would be still the same amount of security data going in. Yeah.
Yeah. net and Python and curiously, uh, three of the four did almost identically and one did significantly worse than all the rest. Give you one guess who it was?
JavaScript? Java. Java itself.
Java was only accurate about 30% of the time. 30 Oh yy y That's pretty bad. That's pretty mess.
It was pretty bad. Why, why do you, why do you think that? So we have hypothesis, but I wouldn't say we know for sure, but I, I have observed that like Java is older than injection attacks and cross site scripting.
It predates a lot of things and I think there's code out there that existed prior to us even knowing that were security issues and that's gets trained on just like everything else. Right. Well that, that I was just gonna say, that's really the issue, right?
Yeah. These, you know, these models are only as good as the code they're trained on. And so yes, If we only train them on secure code, they probably get much higher scores, but they train on all the code.
Why isn't someone doing that? Why isn't someone doing that then? Well, I would say there's two reasons.
Uh, one, it's relatively expensive to do that. So you'd have to basically fix every security vulnerability in open source code. You're scanning, and I'm imagining that's a, a tall order to do.
So they use what they have available to them and they're not, you know, evaluating it on security. If it goes through the system, that might be something they do in the future. I mean, I would, I'm sure they've asked the question, can we at least identify all the security vulnerabilities before we scan it so we know what's secure and what's not, and use that to train the models models, but models based on the data, they haven't accomplished that yet.
You know, it's interesting, and I've been saying this for a while, lesson I've learned in security in 30 years, is vendors, non-security vendors build in security into their products, you know, a trustworthy computing kind of initiative when customers demand it. Yeah, absolutely. And, And it could be, we're at a stage where, where quite frankly the customers aren't demanding it just yet.
I think we're starting to see that happen. I think there's more and more pressure on LLM, you know, producers and companies are using LMS to do it to a way that's producing secure code because producing more insecure code doesn't ultimately solve the problems we have of making sure we have a, a safe computing environment and we're operating our business in a reasonable way. Absolutely.
On the other hand, as you said, we're churning out twice as Mitch Code and There's some people, I guess we're turning point to that. Yep. Crazy.
It is, it's crazy. So I think things are gonna need to change. Yep.
Yep. I, I know you guys were talking, or not you guys, but, uh, Veracode was talking about it at blackout last week, but the reports available on the Veracode site. Yes, It is out there and find it.
com, code com linkable right off of that. Yep. If you're able to find it from there.
Um, as you mentioned, this is the first year you did this report. I'm assuming there's, you said you're gonna try to do it every six months or so? Hopefully, yeah, a couple times a year we can provide updates, um, probably when new models come out and people wanna see what the trend is and if things are moving, if there's something very interesting to report and we may add more data to the report over time and uh, sort of flesh it out from there.
But this answered sort of the core questions we had and it's been very useful in how we sort of target our research and our AI innovation work to make sure we're addressing the problems that are actually being produced by sort of modern AI based software coding principles. Let me ask you, the $64 billion question for developers and development shops out out there who see this report, you think it slows 'em down? No, I don't think you slowed down.
I was afraid you were gonna say that. And isn't isn't that the, the security pros dilemma right there, man. Yeah, yeah, it is.
Um, I know how it is. I know how it is. It's, it, it could be depressing and discouraging, but it's unfortunately, you know, so let, let me, let me try to put a good spin on this.
Knowing, knowing that Jens, what do we do as security pros to, you know, sort of sticking our finger in the d**e here, what could we do to make it better? Uh, I think I'm gonna paraphrase Einstein when he said something along the lines of we generally aren't qualified to solve the problems we create at the time we create them, right? But I think we need to, to figure out a way to solve the problem.
And I do think AI can also help contribute to that. And I think we're looking at how to leverage it, not just to give you a fancy chat bot to explain all your problems, but to actually help remediate them, understand them and get them addressed. Make sure they never current your code to begin with and to make sure the systems are more secure fundamentally from the LLM out.
And I think there's potential in AI and agent to actually do things that actually enhance the security of, of our systems. But as you said, I don't know that the demand has quite been there yet. Yeah, it will be, you know, it takes time to catch up and until then we keep fighting the good fight, don't we?
We do. We're trying, we're trying our best to hold the line against the tide. Absolutely.
Hey Yas, thanks for coming on. It better not be a year till I see you again on here. I hope so.
I hope not. Yeah, no, we'll make that happen. Um, sorry we didn't catch up a black hat.
As you had mentioned, you weren't there and I misfired in seeing my Veracode friends, but we will be in touch. Congratulations and good work on this report. This's a report that I think people need to see.
You know, thank you very much in light of what's going on. I appreciate it and we'll speak to you soon. Always a pleasure.
Thanks Alan. Thank you. Yas Weslake, chief Technology Officer Veracode, you're on Tech Trunk.
We're just take a break. We'll be back in a moment.