StackHawk Wingman Gives Coding Agents a Security Partner
Security testing for coding agents is where AppSec must go when AI writes most of the code. Scott Gerlach, co-founder and CSO of StackHawk, joins Alan Shimel on Techstrong TV. Furthermore, he explains why fixing, not finding, is now the real constraint.
About Scott Gerlach
Scott ran security teams at GoDaddy for about 10 years. In addition, he was CSO at SendGrid right before Twilio bought the company. Consequently, he and CEO Joni Klippert launched StackHawk to put application security inside development.
Why fixing matters more than finding
Frontier models have made finding vulnerabilities cheap for defenders and attackers alike. Meanwhile, AI generated code has blown past tickets waiting at the end of a queue. Therefore, remediation must become a core part of every step of software delivery.
Scott notes that open weight models give threat actors the same power with fewer safeguards. Furthermore, human reviewers cannot keep up with thousands of lines of generated code. As a result, security testing for coding agents has to run at AI scale.
Inside StackHawk Wingman
Wingman adds a deterministic security partner to your coding agent. In addition, it works like linters and unit tests, acting as a gate before code moves forward. Consequently, vulnerabilities go straight back into the agent loop to be fixed.
It ships as a single binary plus agent skills that teach any plugin capable agent to use StackHawk. Meanwhile, it runs in the IDE, on the command line or inside a software factory. Therefore, security testing for coding agents fits every stage of agentic maturity.
Results and availability
Since a June soft launch, Wingman has fixed more than 10,000 vulnerabilities. Furthermore, it is generally available with a 14 day free trial and then 10 dollars per user per month.
Explore more DevSecOps coverage and the latest Techstrong TV interviews.
For more information please visit stackhawk.com
Transcript
Hey everyone, welcome back here to Sh strong TV. I'm really happy to have my friend Scott Gerlach back. Scott, of course, is CSO co-founder over at StackHawk, and it's always great to talk to Scott because he reminds me of Boulder, and we have some pleasant memories of my times in Boulder.
Scott, how are you, man? I'm great, Alan. Always a pleasure to be here.
Thanks for having me. Absolutely. How is Boulder doing?
The Buffs are doing what they're doing up there, and college is back in, and it's just Boulder, and it's college town central. Yeah. It's Boulder.
" But we never made it out there. We were so busy doing videos and stuff. Anyway, Scott, I'm familiar with your background, but I'm sure people out here aren't.
So why don't you give them a little bit of your journey- Yeah ... to putting together StackHawk. Happy to.
Thanks, Alan. As you mentioned, Scott Gerlach, co-founder, chief security officer here at StackHawk. My background was actually running different security teams, so I ran security teams at GoDaddy for about 10 years.
I was the CSO at SendGrid right before Twilio bought them, and I've been involved in running different aspects of the security program, including application security. When I got done with SendGrid, and our CEO, Joni, was done with VictorOps at the same time, we collaborated on how do we kind of upset the apple cart of AppSec? We started StackHawk to try to figure out how to get more application security information into the development life cycle so people can actually find and fix some of the security problems that exist in API code and web app code, those kinds of things.
If only you had founded an AI model that could find untold vulnerabilities, right? We wouldn't- If only. If only.
We did the next best thing. How's that? Right.
Couldn't do that, at least you did this. That's right. All kidding aside, though, of course, Scott, mythos and post-mythos, the cheese has moved a little bit in AppSec, right?
Yeah. Where so much of our focus was on finding vulnerabilities. And this goes back to when I was doing security at Still Secure 25 years ago.
Quite frankly, we always found more vulnerabilities than we could fix anyway. Sure. Right?
That was part of the issue. But now, in a post-mythos world, that's a little bit on steroids, right? We're finding so many vulnerabilities, or we can if we desire.
Right? Or someone else will. And the issue is one of governance, one of prioritization, one of, well, that's no longer the constraint.
Mm-hmm. The constraint now is how do I remediate? How do I- Yeah ...
bring these vulnerabilities down? And- How do you fix? Yeah.
It was interesting. We were talking about it on the gang today. Like a fix was in, but it wasn't a patch.
It was more of a temporary block. Right. And now it turns out people can get around the temporary block, and so it's no fix at all.
It's a different world in AppSec, Scott. Very different than when you and Joni first looked at it, what was it, maybe seven years ago now? Six years ago, something like that.
Yeah, seven years ago. I think the fundamental problem is the same. How do we incorporate that fix into the development life cycle so that software comes out as secure as it can be.
The game changer is the amount of code that can be created now. Yeah. So the ability to create code at such a rapid pace has just completely blown out the buffers of how traditional AppSec is done, where you've got an AppSec team that's waiting at the end of a queue for checking the security of an application or an API.
They find some stuff, they build a ticket, they put that into the ticketing system, and then they hope and pray that someone does something with that ticket. That is extremely broken with the amount of code that can happen. " And my whole premise was fixing is actually hard, and I got 50/50 results on that particular thing, which was interesting at the time.
But fixing is so important to what we're doing because of the power of some of these new frontier class models, whether they're gated like the Anthropic models and the OpenAI models, or open weight, like some of the Qwen models or some of the other things that don't have quite the amount of restrictions and safeguards in place. Threat actors have access to these models just like defenders do, and being able to change the process and tooling around how application security works so that fixing is a core part of every part of software delivery is so important now because of the power of those models. Scott, too bad you can't get in front of those same people today and ask them.
Yeah. Fixing or finding, as I'm sure would be different. But you're right.
So I call this the AI scale problem. Anything where AI is getting involved in a meaningful way is just blowing our human-based scale out of the water. Mm-hmm.
Whether it's how much code we're generating, how many vulnerabilities we're finding. It operates on a scale that we can't match human to machine. The only way to really kind of do it is AI to AI, if you will, and we're seeing a lot of that.
What's interesting with the progression is two years ago, a year and a half ago, look, the code it was turning out was, yeah, they could turn out a lot of code, but it was slop. Mm-hmm. And it was not only not secure, it wasn't really very good code, to tell you the truth.
It didn't do what it was supposed to do. It was buggy as hell. That, of course, has changed now, right?
A lot more code out here is AI-generated, and of course, we're using AI to test it and do everything else. StackHawk has now launched something called Wingman. Mm-hmm.
Well, before we even get into what Wingman does- Okay ... I got to ask you, only because I know you and I know Joni, and I know Casey, and I know the StackHawk team. Yeah.
Who came up with Wingman? Who came up with the name Wingman? Yeah.
Got to give that one to Joni. Okay. She was ideating on different names for what it is, and we were talking about what it does, and it really is like kind of got the security back of your coding agent, and we may or may not have been watching some "Top Gun" at the time.
I don't know. That's what I was thinking. Was this a Tom Cruise thing?
Or, in which case, I guess I'll go with Joni, or was it the Boulder college kid thing, and I need my wingman going out on Pearl Street with me at night. Yeah. It's less of the Boulder college kid thing, but the simile is really kind of the thing, right?
In "Top Gun," they're talking about not leaving your wingman, the whole movie, and even in two movies. And the whole thing is how do you add a deterministic security partner to your coding agent so that it knows when it's making mistakes, and kind of has your agent's six, for lack of a better term. But that's how we came up with the name.
It's supposed to be the coding agent's deterministic partner on is there security problems in this code or this application that we're building. It's funny you described it that way, right? You developed a deterministic partner using non-deterministic tools, right?
AI is famously non-deterministic. How do you close that gap? Yeah.
So the thing that's really great about how Wingman works, and when you look at agentic software creation loops, the best way to inform the agent or a coding session whether or not it's making mistakes is by having tools that run that give deterministic output. And so if you think about linters and unit tests and integration tests, and now security tests, those are deterministic gates. Am I good to proceed?
Does this have vulnerabilities? Does the linting pass? Does the unit test pass?
" Excellent. Where does it live in the IDE? Is it outside of that?
How- Yeah ... talk to me about, as a product, how do I run this thing? As a product, it works with your coding agent.
So if you run your coding agent in an IDE, it will absolutely work there. So it's a single binary you install that does all the testing and feedback to the agent, and then some agent skills that kind of, if you think about... My favorite analogy here is "The Matrix," when they plug Neo into the Matrix, and they teach him kung fu.
Mm-hmm. That's basically what a StackHawk plugin is doing there. Teach your agent how to StackHawk, and then it starts working.
Just starts working by itself. You can kick it off manually if you'd like, but naturally, it just works along with the coding agent as it's building features. It will then spin up the application, test those features, feed the vulnerabilities back to the agent as a, "Hey, you still have things to fix," so the agent can actually fix things.
And the really exciting thing, Alan, is things are actually being fixed. So we soft launched Wingman in June, and since June, we've fixed... I just looked at the number this morning, and we've fixed over 10,000 vulnerabilities so far from people either trying Wingman or current customers.
And that is a really exciting number because I don't know if I've ever fixed 10,000 vulnerabilities in my 25-year career. No. And again, that's AI scale, right?
Mm-hmm. That is AI scale. Crazy.
It's just crazy. Now, when you say it talks to the agent- Mm-hmm ... first of all, what agents?
And secondly, how? Is it sort of an MCP kind of connection? Is it...
What are we talking here? Yeah, so any agent that supports the plugin framework, so that's pretty much every agent at this point, plugin and skill framework. Pretty much every agent supports that today.
The way that the tooling actually work is a single command line install, so the agent can actually drive the command line tool, to either run scans, run tests, interact with the StackHawk platform and get information back, as well as how to get vulnerabilities back to the agent in a language that they understand and just is very well described. So the agent can actually drive the command line utility. That command line utility gives it output.
And now Alan's got a fire at his place apparently. Do you hear that? I was hoping you didn't.
Well, is this a fire or a fire alarm? I don't hear anything. Oh, I hear it for sure, and I- Oh, I hear it.
Yeah, it's a fire alarm. It's a test. They're testing it.
You know for a fact they're testing it? It says something. Excuse me.
You can hear it. Okay. Will we be able to take this out?
Yeah. Okay. All right, Scott, we're good.
All right, cool. If you see... Scott, if you see smoke behind me or anything- Yeah, let me know ...
let me know, okay? Anyway, do you want to pick up where you were? Yeah, where was I?
Let's see. How the agent works. Right.
So the- So you could use any agent using the plugin. Any agent. Yep.
And the plugin- Any agent using the plugin, whether you're using it in an IDE, a ton of people that I talk to right now are writing codes just on the command line, not even using an IDE, so it works that way as well, and it's really built for no human interaction. So can you oversee it in kind of human over the loop mode or human in the loop mode? Absolutely.
But it's built for kind of the future where people have software factories where you're giving an input to a software factory, it's building all the code, testing it, making sure it's right, then security testing it, and making sure what comes out at the end is the best code, the most secure code it can be. But it works kind of all along the different software development maturity models, as you were, to where if I'm just getting started with an agent and I'm instructing it how to build code, StackHawk will work along. If you're doing software factories, it will work there as well.
Okay. So Scott, that same AI scale rears its head here. Does Wingman work at AI scale?
Yeah. Wingman does work at AI scale because it works with your AI. So that's the really cool thing about how Wingman works is we're not just generating PRs for someone to pick up or something else to pick up.
We're not just generating code fixes. We're actually working with your coding agent, which has the context of how you build, where you build, how other services interact with this code. So being able to utilize your coding agent in your environment with its rules really makes Wingman crazy powerful.
So as fast as your agent can go, Wingman's just right alongside there, kicking in the AI afterburner. All right. I can see we're going to have to do some "Top Gun" kind of quotes in here.
But, so let's nuts and bolt business. Wingman's generally available right now? com.
There's a 14-day free trial, so if you're doing agentic coding, you can give it a try, and then it's $10 per user per month after that. It's an easy thing. Scott, one of the things that I think has kind of changed the game a little bit is the rise of platform engineering.
Mm-hmm. Right? A lot of developers, you said some don't even use an IDE anymore.
They're doing it right in command line, but the whole CI/CD workflow has moved into a platform. And developers use it, testers use it, DevOps uses it, SREs use it, right? A common platform that allows us all to go faster within the rails.
How does Wingman fit into your platform engineering strategy? Yeah. In a very similar way.
If you think about what's going on with that kind of platform engineering and that AI scale problem, one of the most precious resources is brain, like human brain power, human brain attention, human brain context, and we have to start protecting that a little bit more. And one of the things that you can do to protect it is, in the evolution of code development and how much code can be created, you can't just pass all this code over to a person and go, "Hey, make sure you do a thorough code review," without any other information and context, because when we're talking thousands of lines of code that AI can generate, it turns into either rubber stamp or it just doesn't get reviewed. So that's why you're looking at things like having an agent do a first review of it to get the human reviewer to look at the code smells.
This thing is a little unsure. This particular thing needs review because it's an anti-pattern in our environment. Those things really help that scale.
But the other thing that helps is not making people be domain experts in every single thing, and security is one of those things, and that's a challenge that has always existed, and that's why security teams are what they are. They're domain experts in security. " You can be a champion, you don't have to be an expert.
Yeah, exactly. You don't have to be an expert. You have tooling that helps you understand, where do I need to pay attention?
" Or even as just a regular developer or an agent, I should be fixing these problems. Got it. I love it, Scott.
I like the direction for StackHawk. It's good to see. I've been kind of collecting AppSec companies' responses, because Mythos has changed this game, and I think for the better, because I think it's gotten us out of that 50/50 finding vulnerabilities versus fixing them that you described, was it a year ago, two years ago?
And clearly placing us in the new front, right, which is, hey, we got to use AI, we got to do it at AI scale, but we've got to fix these things. com, right? com.
I think there's a really awesome... It's scary and awesome at the same time right now because of all of this ability and tooling. I think we keep talking internally about the defender advantage.
If you can figure out how to change your process, change your tools to integrate AI at scale and at the same scale that the engineering team is going on, because you have all of the source of truth, you can get to defender advantage and really tip the very lopsided defender/attacker scale today to your side. And that's super exciting for me, and we hear it all the time from Anthropic and OpenAI when they give speeches. " And that's what we're out here trying to help support security teams' ability to do.
I love it. Fix and fix and fix. Fantastic.
Hey, Scott, say hello to Joni and Kasey and the gang. Will do. It sounds like a band that I've heard.
No, that's a band. It's a strong three-piece, that's for sure. Yeah.
We might need somebody else. Kasey and the Sunshine Band. But anyway, best of luck with StackHawk.
Keep us posted. It's been too long since I had you on here, Scott. Fair enough.
Alan, thanks for your time today. I really appreciate it. Thank you.
Scott Gerlach, StackHawk, CSO co-founder. Go check out Wingman. You're watching Techstrong TV.