AI Code Security Demands More Than an LLM
More Code, More Pressure on Security
AI code security is becoming a boardroom concern, not just a developer task. Coding assistants are accelerating software production. Meanwhile, attackers can use powerful models to discover weaknesses and develop exploits.
Checkmarx CEO Sandeep Johri examines that collision with Mike Vizard. Johri describes rising anxiety among security leaders as new code compounds existing vulnerability backlogs. Finding problems is no longer enough. Teams also need practical ways to prioritize and fix them.
Combining AI with Deterministic Testing
Johri argues that large language models bring valuable security capabilities, but cannot provide comprehensive coverage alone. Their findings can vary across models and prompts. False positives and operating costs also complicate enterprise adoption.
The discussion explores a hybrid approach to AI code security. Johri describes Checkmarx Fusion, which combines deterministic scanning with multiple language models. The goal is to broaden detection while filtering duplicate findings and false positives.
Rather than choosing between established testing and AI, he makes the case for using both. Deterministic engines provide repeatable analysis. Models can help uncover additional weaknesses and explore complex attack paths. Johri positions that combination as a way to improve coverage without relying entirely on probabilistic results.
Prevent New Flaws and Reduce Existing Debt
The conversation then shifts from discovery to developer workflows. Johri advocates checking generated code before it enters a repository. He describes Checkmarx DevAssist as a security plugin that works within AI coding environments.
Existing vulnerabilities require a different response. Johri explains how agents can prioritize findings, prepare fixes and return proposed changes through pull requests. Developers remain responsible for reviewing those changes. This approach aims to reduce manual effort without separating security from everyday development.
Workflow integration also matters for collaboration. Johri sees an opportunity to ease friction between teams when security tools help developers fix issues earlier. He argues that useful controls should fit established processes rather than force developers into another interface.
Security remains an ongoing effort as models, applications and threats evolve. The interview closes with a practical direction: prevent avoidable flaws, tackle accumulated debt and keep improving defenses.
Transcript
Hey guys, thanks for the throw. We're here with Sandeep Johri, CEO of Checkmarx, and we're having a little chat about what's going on with all these vulnerabilities that AI tools are starting to discover, and worse yet, exploits are not far behind. Sandeep, welcome to the show.
Hey, thanks, Mike. Good to meet with you again. Unless you've been hiding under a rock, you probably have heard something about the fact that these more advanced AI models are discovering vulnerabilities, and we're very concerned about not only the number of them, but the fact that they seem to be easily exploitable, and we're all wondering, are we looking at some sort of application security crisis that's about to descend?
Sandeep, what's your assessment of what's going on here, and are we prepared, and if we're not, well, what should we be doing about it? Yeah. So what's going on is, first of all, even before we think about the exploitability of LLMs, we need to think about what's happening in the dev environment.
Every developer I know is adopting coding assistance, and the volume of code is increasing rapidly. We have access to all our customers' repos, and we see the volume increasing dramatically. That's number one.
Number two, most of this new code is LLM generated, and not therefore, but happens to have a significantly higher vulnerability density. Why? Because LLMs don't generate clean code.
There's plenty of benchmarks in the industry that show that auto-generated code has a higher density of vulnerabilities. So on the supply side, you have this geometric increase in vulnerabilities, and then on the other hand, like you pointed out, you have LLMs that can be weaponized. They become more and more powerful.
They can be weaponized to discover new things that previously were not known and make it a whole lot easier to exploit the ones that are even there. Even things that were very difficult to exploit before are now what needed a state actor level of expertise to exploit now can be done very simply with LLMs, and you don't need the latest model to be able to do those exploits. So is there a crisis in AppSec?
From a-- I would say there's no crisis in AppSec, but there's certainly a crisis in code security, and we are hearing that from every CISO. I've probably talked to a hundred CISO in the last six months, and everyone is super anxious about code security, about the fact that they have vulnerabilities. They have known that they've had vulnerabilities, but because they were difficult to exploit, the number of vulnerabilities remediated in the past was relatively low.
And now they feel a lot more anxiety about having to, one, be ahead of the bad actors to make sure that they discover everything, and two, not just discover, but quickly remediate. So that's what we're hearing from customers. Everyone, every single CISO I'm talking to is talking about that.
Code security is a top three issue for them. Not only that, Mike, they're required now. It's become a board-level issue.
The boards have read stories about Glasswing and Mythos and other models, Astral, and they're asking the CIO and the CISO, "What are we doing about vulnerabilities? Are we doing anything? Is our risk posture improving?
" So that's what's going on. Some folks would say, well, won't we use AI to validate and fix the code generated by other AI agents, and that's how this whole process is going to work. And yet when I talk to other folks, they'll say that the false positives generated by a lot of these AI tools is really high, and so it's not really helping, and we need some sort of method to provide some additional context if that's going to work.
But how does that come about in your mind? Yeah. So LLMs are amazing engines as we see them generating code, doing all the functions that they're doing.
They're amazing, and they're pretty amazing at security as well, meaning they do discover things that previously could not be discovered. They do allow you to test things by stringing together simple things into a very complex exploit. So there's real benefit in LLM capability.
However, there's a few things. They are probabilistic engines. Most of them, in their own literature, will tell you that we are not comprehensive, meaning while they discover things, they don't claim to discover all known vulnerabilities.
So they might discover a hundred, but they might leave several hundred because they are probabilistic. Not only that, they're inconsistent in that if I have it run with some prompt versus you have it run with a different prompt on the same code base, you'll get slightly different results. Different models will give you different results because their logic is different, right?
It's all probabilistic. And so those are a couple of challenges. The third thing is what you called out, which is they are probabilistic, and therefore, by definition, there's a lot of false positives.
So they give you a stack of hay. There definitely are needles in there, but you've got to now sort through the hay and the needle, right? You have to find the needle in the haystack.
And then lastly, I would say the challenge with LLMs alone is that they're super expensive, right? So if you run all your code through an LLM, you're talking about taking 10 times as much time, and the budget literally will bust the budget for every enterprise. So those are some of the challenges, and that is where we have a very unique perspective, which is why I said I don't think there's panic in AppSec, at least not in us, because we think we have a great offering that actually addresses this.
So let me talk about that. Does that make-- Would it be useful for me to explain that? Yeah, I think a lot of people are trying to figure out, well, what is the right balance between a set of probabilistic tools and a bunch of deterministic tools, and when do I use what for what?
Yeah. So let's just lay out the options. One option is, so first of all, we all have to agree that the LLMs are phenomenal engines, right?
For all the reasons I stated. They do find new things. They can be weaponized, so you better be using that to make sure someone else is not weaponizing it, or you are ahead of others weaponizing it.
But one option is to just write your own harness and go directly at the tool-- go directly at the LLM. Most companies don't know, most enterprise teams don't know how to optimize that harness, and it'll end up being super expensive, and you're left with uncertainty of what you get. The second option is to use some of their solutions, like Code Security, which is essentially a security harness around their models from Anthropic.
Codex has an equivalent security product, but they still only do the probabilistic part. And for all the reasons I stated, they miss stuff, they're inconsistent, blah, blah, blah. What we have come up with is what we believe to be the best of both worlds.
The product is called Fusion. We just launched it a few weeks ago. We put it in production yesterday.
What it is, is it's a hybrid. So this exact dilemma that you were talking about is what we address. It's a hybrid solution that uses deterministic engines.
We have 20 years of expertise in that, and we believe we have the best deterministic engines out there. And we also run LLMs. So we have optimized and harness to run against multiple models.
We then take those results, we dedupe it, we identify the false positives, and what we are left with is the cleanest, most highest fidelity results that combine the best of both the deterministic and the probabilistic. And I'll throw some numbers because we have started running benchmarks. Just yesterday, we ran a benchmark against two million lines of code.
It was five projects, roughly two million lines of code. To run it on Cloud Code, it cost us about $1,400. Okay?
We got about 108 results out of that, true positives. Actually, 84 true positives out of that. 7 and the deterministic engine, our engine, we found 508, I believe was the number.
So we're finding five times more vulnerabilities. It's more comprehensive, and you can run that again and again, and you'll get the same set. And we suppressed close to 100 false positives that the LLMs discovered.
So that's what it does. And we can do this run, you can integrate it into your workflow, so you can run it daily, you can run it weekly, and we can do it at less than half the cost of what it would be to run it against your Cloud Code Security. So it is a super, super compelling value proposition.
As we go down this path, it almost seems like we're trying to do two things simultaneously. One is reduce this massive amount of technical debt that we've allowed to accrue over the decades. At the same time, though, we're creating more new code than ever, and the new code is of somewhat unknown provenance and quality.
At what point does this just break and people become overwhelmed, and how do I reinvent my entire SecOps workflow? Yeah, actually, I don't think it breaks. And I'll tell you why, Mike.
For two reasons. One, there are things you can do at this code generation stage itself. So one, for new code coming in, you can actually test for security vulnerabilities.
Even though LLMs generate not very clean code, you can cleanse it right there before it comes into your repo. So there is a prevention element to it, which is important. And there are tools.
We have a tool, we have a solution there called DevAssist. It's a security plugin that sits inside of Cursor or of Hero or Devin. And just as the code is being generated by the LLM, we can rinse it right there for vulnerabilities.
Now, that allows you to prevent new code coming in with a lot of vulnerabilities. So you're trying to catch it as left as possible, right? And many customers are beginning to put that in place so that they, as the volume of code is increasing, at least the new code coming in, you try to catch as many things as possible as early as possible.
That's one. " Even with an LLM, it'll take them an hour or two to fix it. Why?
Because they need to understand what it is. They need to go back and forth with the LLM to understand what it is, to get a fix, to try it out. We do all of that.
We have agents that can do both the prioritization of the vulnerabilities and the remediation. We do all the back and forth. We've come up with a remediation.
We put it into a PR and send it back to the developer, fully integrated into the workflow. So now, what would have taken, let's say you have 10,000 vulnerabilities. If you take even two hours each vulnerability, you're talking about 20,000 hours, man-hours.
We can do that in minutes. Literally, this agent can automate all of that. We run it through the process back to the developer as a PR.
The developer reviews. In two minutes, they accept it, and you're on your way. So the point I'm making is, one, you should be preventing.
Two, you should be running through your tech debt rapidly with agentic capability to remediate it faster. Mm. And you can do that.
" And they're making great progress on that because they're using these agents. As this evolves, how will the relationship between the cybersecurity teams and the app dev teams mature, change? And I'm asking the question because historically, there's not always been a lot of love lost between those teams.
Mike, that is true. There's always that tension between dev teams wanting to ship fast and security teams wanting to make sure. But even as everybody shifted from a DevOps-only mentality to a SecOps mentality, that is integrating security into the workflow was the theme even before Mythos.
So for that matter, our Checkmarx product, one of the core value propositions is that it is a very developer-friendly product. What does that mean? We don't expect the developers to ever come into Checkmarx One.
The Checkmarx One capability is integrated into their workflow. So in the AI context, I don't see that relationship worsening. I actually see acceleration of integrating security tools into the dev process because the dev teams also realize that they have this problem, that this new code is being generated that has a lot of vulnerabilities.
So they are welcoming tools like ours, which are integrated into their workflow, are frictionless, and help them fix things earlier and easier than before. So I don't see it worsening. The anxiety level from CISO is higher than it was before, I do agree with that, where before they would say, "Okay, well, the dev team doesn't listen to me" or something to that effect.
" So I think CISO have a little more power to influence the dev team, but I don't see too much resistance from the dev team, as long as the tools they pick can be integrated into the workflow. I think the one question everybody seems to have is how long will it take us to work through these issues because, I think we're kind of expecting that there will be some period of pain in the short term, but eventually we'll be better off. The question is, is how long will it take to get from here to there?
Well, it's going to be a journey. That's the case. The debt can be taken off relatively quickly, and then you need to have, like with anything else in security, with every new technology come new security threats.
And there are new tools to deal with that. For example, as people are adopting AI into building applications, forget dev, just other aspects of the company, people want an inventory of all the AI assets. They want to be able to scan all the skills.
So we already have solutions in the works on that front where I don't think we should think of this as, okay, it's going to take me six months, and then I'll be done. You're never done with security. But you have to keep up with the time.
And I think that solutions like Fusion, our triage remediation agent, allow you to deal with the debt very easily. And you move to a new world where all of this capability is integrated. 7.
But we're planning on offering Mythos as soon as we are allowed by Anthropic. We also plan to offer other LLM models. So the customers will have a choice to really be able to use this harness to test against multiple models and really be sure that they are secure.
And the solutions today, like Fusion, actually allow you to do that. So I don't see it as, okay, it's going to take us six months, and we'll be done. Nor do I look at it as a doomsday, oh my God, we'll never be able to get over it kind of a thing.
But the models will keep evolving and getting better, and there will be new security threats as well. There's skills, there's LLM security, there is MCP needs to be secure. So there's new threat vectors we need to deal with as well.
All right. Well, folks, you heard him here. There are clearly challenges ahead, but there's hope.
We have tools and the tech, and we know what we're doing. Hey, Sandeep, thanks for being on the show. Thank you, Mike.
And back to you guys in studio.