Andy Ellis – What Are the Metrics that Matter in AppSec?
Building out an AppSec program is already a challenge, and learning how to measure the efficacy of your program even more so. Listen in as Andy Ellis walks through the philosophy of measuring your AppSec program — both qualitatively and quantitatively — and communicating its value to your stakeholders. There will be an open question period, so bring your metrics and get Andy’s opinion on them, or any other question you bring!
Transcript
Good morning, or good afternoon, maybe good evening or good night. If you're coming in from somewhere else today. I'm going to talk about how to build your security program and how to focus a little on absec.
This is a derivative of a talk I gave at RSA. So if you didn't see that you can always go check out that longer form version. I included all the slides but I'm gonna slide through a lot of them pretty quickly today.
com and click on talks and you'll see that I think it's the most recent talk I've added there. So let's talk a little bit about how do you build and measure a security program and a lot of our programs start with this concept of Defense in depth. This is the Mage no line just very quickly important thing to remember is the defense in depth was about trading space to buy time and in the app SEC world that makes absolutely no sense developers, you're of attacks design them to Pivot so quickly that even if it takes two or three steps to get through So those steps are happening within milliseconds, you know, their adversaries are basically coming around whatever defense is you build in.
So it's important to know that you have defenses that work together to stop adversaries. You know this basic model, we've always had of the perimeter is a challenge. This goes back to Roman Camp times you'd show up at the edge.
You need to have a password to get in. So we have built a moat to make it harder to climb over the walls, you know, and hopefully we put Defenders on them to shoot you. Unfortunately, we don't get to shoot the people who are attacking our applications, so we should think carefully about whether this model actually works for us.
So that's sort of gets us to where we are. Today is we have Security Programs and we have measures for them. But are they really effective?
So let's look at one metric and this applies to applications to which is vulnerability management. Right, and I have a big problem with vulnerability management the way it's often measured because what we do is we talk sometimes about how many open vulnerabilities you have as a company and what is the average age of open vulnerabilities? It's really almost a defect measurement, but it's not even a really good defect measurement.
So let's look at what we're doing. So first question is anytime you're measure something. You should ask like what isn't being measured like what systems aren't included in here.
Right are we actually covering all of our systems? You know, I have run into cases with vulnerability Management Programs where you're one organization is reporting up really good data. And so their data is what gets used to cover the whole company.
And so the people who don't have great data on their programs, we're not reporting on that and often that is the application teams probably have less coherent and clean data compared to say your Windows Server admins who only have to deal with the small set of things and they're all coming from one vendor that big Microsoft so it's what vulnerabilities aren't counted in here and maybe sometimes what less relevant vulnerabilities are being counted. So these are things that we should always challenge ourselves and ask when we look at some Metric but let's look at some ways that this metric is bad. We're gonna sort of war game it and this is always fun.
I love breaking things. So let's break. Metrics to get ourselves started and sort of woken up for the day.
So let's say that a vulnerability comes out to new one. Let's go back in time log4 shell but a lot of us are very familiar with the log Forge and the multiple, you know iterations of that vulnerability. And so let's take the first what if like the vulnerability happens and let's assume we didn't patch it and first to give some context on like what's the data behind this chart is we assume that you had I think 10 new vulnerabilities every month only it was that small but you didn't fix any of them.
Okay, so maybe just 10 new ones that didn't get fixed and you have a million that you fixed. You have a great program either way. And so you get 10 new vulnerabilities come in in December and what happens?
Well, the total number of open vulnerabilities goes up a little bit like you can barely even see it on a chart but what's fascinating is the average age of open vulnerabilities. This defect measurement actually goes down. So if you're arguing that we're a Bad Company because our average age keeps going up because we don't fix anything.
Well when we get extra vulnerabilities in arguably, we're in a more dangerous place the measurement that we would be showing to people actually says the opposite it says we just got better. That's a bad metric and if you have to put these two together and explain it to a lay person that's a sign of a dangerous metric that you should be very careful about using So let's just start with that one is that's that one bad. What if we didn't even patch anything we looked at.
Well, what if we just patch it after a month? So it looks the exact same and then a month later you'll get 10 new vulnerabilities come in but we've dealt with these 10. So the the Blue Line sort of you ends up back where it would have been had the vulnerability never happened, but it does have this little uptick and the red line.
Boom. Look at that. It jumps.
It's way worse. in fact This is kind of interesting is that like we'd have been better off going and patching whatever these irrelevant vulnerabilities were that had been around forever to keep our numbers low. And not patch the most relevant vulnerability that just hit us.
And so that's the thing that we should worry about. Now what if we just patch the moment it came out and in between reporting Windows, right? Basically doesn't even look like it happened.
Our metric does not even show. That we had a serious event and we dealt with it. And you might say well that's kind of arbitrary except many Security Programs suffer from this intentionally.
If the people who are reporting data to a security team know when the security team is going to checkpoint that data to show to someone else that's an incentive to them to clean things up right before the checkpoint and then all of a sudden they end up in a cycle where all they're trying to do is clean up for the checkpoint rather than for the security reasons themselves. So this is a bad metric and so we should explore ways to improve this metric as well as ways to apply this logic to other aspects of our security program, especially as we dive into appsec. So here's the new metric you might consider for vulnerability management, right?
It's challenge our definition and say like here's this issue like our goal is we have vulnerabilities that matter differently based on their criticality. Let's assume that our criticality is accurate. It might not be but that's a conversation for another day and we have some SLA that we have established and we believe right they are wrongly that that SLA is tied to outcomes, you know, critical things need to be fixed within one week before there's Mass exploitation arguably.
That's really about 36 hours. Maybe a little bit less, you know High maybe you're giving yourself a month. Maybe you give yourself 90 days for medium and 180 days for love reality is I don't think anybody ever has an SLA for patching love vulnerabilities.
But let's hypothesize. It's there. Well now let's measure what the actual sla's are that we're hitting.
Right. So if you have a seven day critical vulnerability, you know, how often are we patching that 85% That would be fantastic. In fact arguably.
These numbers should look like this the more important it is the more likely we are to reach the SLA. Now important caveat that you need to make sure you have in the definitions which is like what happens if somebody knows they can't fix it in time. Where do you score?
I'll come back to that a little bit later, but always be careful thinking about expectations and exceptions and how people might game one of these even unintentionally. So an important thing to realize is that I started with this sort of linear defense in depth model, but we should recognize that defense is not linear, you know, our adversaries can attack Us in breadth in height or in depth. And so let me just quickly explore each of those three, but our Defenders need to match our attackers where the attackers are going to come and that's what we're going to have to meet that challenge.
So let's talk about the First Dimension, which is this breadth or width which is, you know, if you had a castle and you have two gates the attacker gets to pick which gate or climb over a wall. They don't have to come in the way that you want them to come in. And so this is often about a tax surface, which is another way of fancy way of talking about inventory.
Do you know all of the places that somebody could attack you this is not yet about the weakness. This is just about the places you have you collected them. What is your inventory?
So for me the first part of tackling Dimension one begins and ends with the most important. Important security tool ever invented which is the spreadsheet. Maybe you've progressed from Excel sheets to Google Sheets, or maybe you're in Office 365 but I have found that creating a spreadsheet that just lists type of assets you say.
Okay. I'm going to list all of my types of assets. Been the reason why you're going to do this is because you have to tackle them in different ways different people own them.
You maybe have Enterprise systems owned by the CIO. Maybe you have Dev environment owned by the developer productivity or the devops team, you know production owned by your actual devops public Cloud who the heck knows who owns that and for each of those count. Like how many do I have how many says services do I have and then ask yourself one simple question?
How hard was it to collect that data? If this was something that I was able to get in an easy automated fashion It is Well lit that's fantastic. But maybe this is really Shadow it like I don't even know how many Devin build servers I have.
And so it's like living in the dark. And some of these numbers might be variable like public Cloud that number is changing every minute and that's okay. The important thing is do you have an easy API to quickly collect that data?
And do then know what's running on all those systems and what controls are around them. This is just a starting point. And now often people leave out one of the most important asset classes people have which are your applications and your apis.
Do you know all of the ways that somebody can interact with your environment at the application stack because all too often Security Programs are network-centric so they're still relying on protocol vulnerabilities. So many vulnerability detection Suites are focused on like are you on the latest version of open SSL? Very important thing to be like let's make sure you're patched and up to date make sure you've all taken your Chrome updates and your iOS updates for today.
But some of the most serious vulnerabilities we're going to have are going to be in our apis and in the application stack, we wrote they're often bespoke vulnerabilities. They're not always just like log4j imported into our software supply chain. So do you even know how many of those API endpoints are out there?
And do you have any systematic way to collect them? So that's how you're going to start on your app set journey is by doing this and I say start and I don't mean you don't get to move forward until you have finished this. One of the biggest challenges many Security Programs have is they start at inventory?
And they refuse to proceed until they have finished the inventory inventory is a parallel track to security. And what's going to happen is you implement security controls you're going to implement it against the inventory as you know it and as you learn more things you then apply more controls to them. That's just how it's going to work think of it as working in two dimensions on a spreadsheet and as you're adding systems and you're adding controls, you're just trying to fill out and grow more of this spreadsheet.
Next thing about height so height is really what people thought about with defense in depth. It's really do our defenses stack. Right is this just seven speed bumps in a row, they don't stack or is it a wall that somebody has to climb over now seven speed bump high wall isn't very big unless you're trying to drive over it in which case it might be a little bit challenging for most Vehicles.
So we need to understand like from the adversaries perspective what controls we have and how well do they actually work together? And so we're going to think about those controls, you know from this sort of Define them perspective. Let's start outside the app set world just to give a high level view of this right you have some controls you might say inventory vulnerability management, and these are tied together inventories.
Do I know what software is on the machines so that I can patch those machines. So that I can configure them correctly. Do I have authentication tied in all the appropriate places on those systems, you know do I have access control that works with that?
So I'm limiting who has access and for all of these you're going to want to define a measurement and a metric that you're going to look at and understand if that is actually an effective measure and I put a mix of effective and not effective measures here. So yeah, I'm just using the one for vulnerability management that we just defined but then think about configuration hygiene here and all I've done is said, you know, there's high medium and low findings from whatever is my configuration hygiene detection system. You're my cspm or whatever and this is how many I get to and that's a good start but it's a measure of activity not really effectiveness.
You know, but think about authentication and you might say well here it's about user MFA and managed identities for my machine systems. And how many have moved over into a control regime where I'm more comfortable that we actually authenticated user and we're not suffering from an account takeover. You know for Access Control, maybe we're looking at what I call Grant utilization rate.
And think of that as applying the principle of least privilege which everybody has talked about for longer than I've been in the industry. But actually thinking about it logically. If you say you have a hundred identities you whether the users or systems and you have a hundred different rights that they could be granted.
Maybe it's access to systems. Well that gives me a matrix of 10,000 possible rights and in many systems like that's an end to any great. We give everybody access to everything.
Now you when you look at it you say what is actually being used. You know, it's maybe every one of those hundred users is interacting with maybe up to three of the systems. So we're using three percent of our grants.
Well, what if we could narrow down those user groups and maybe group every user into a group of 10 users Each of which has access to 10 systems, maybe choosers still only use accessing three machines, but now our utilization rate is up to 30% So we have carved out an awful lot of access that wasn't necessary did we've eliminated 90% of the access in our environment. And so Grant utilization is sort of this good measurement of how close are we to the principle of least privilege? You know, you might ask what's an ideal number.
It is not 100% Unless you're very good on just in time access Rights Management. That is user friendly, but those are really let's let's not talk about with you at sort of the pipe dreams here. I think that a number between 10 and 80 percent is your ideal original starting point.
That's where you want to get to you want to be able to say, oh like users have a little more access than they need but that's their provisioned because we have assisted mint Team that manages 10 systems and while they do some partitioning who normally manages things. We know that in a crisis, they need to Pivot very quickly. And so we've just provisioned them a little bit greater.
But you can decide what that number might be. And that's an example of a measurement. Nobody really does today.
I've included exploit monitoring here. Sometimes called threat hunting one of my least favorite controls It's actually an anti-control. Because it was a control you'd actually have it somewhere up here.
This is did somebody break every one of our controls and then did we find them later? So really you're measuring dwell time here. And I would actually functionally argue that you want dwell time to go up and not down in your environment.
And the reason for that is you actually want to make it the people don't break in. So all of the people you would find very quickly you should instead prevent them from actually dwelling in your environment by moving those fine times down to the moment of compromise kicking them out instantly and then not really counting that as dwelling, you know, they broke into a system and they were evicted before anything could happen. And then you should think about data protection.
That's still I think the new and novel space in the cloud World especially is people getting more and more systems. So now let's think a little bit about how to do that into the application world. Like how does that really apply for us?
com produced by enso security and what this does is it basically lets you take for any one of your applications you look at what are the components of it and I'll show you a few of that in a moment, but then it talks about what are all of the controls that you might want to put in place around that. And here's what that might look like as you might say. Here's a set of controls for an application.
Do I have inventory? Do I have source code analysis? Do I have any of the application security testing tools?
Do I have code review? Do I have a pen test or a bug Bounty? Do I have training?
Do I have laugh and Bot mitigation do I have a cicd security and you could break these out if I listed these all on separate lines. It would just be much harder to read so I didn't. And now when you think about defenses here.
The important thing to think about is how well do these run in the absence of executive oversight? Because once you have a program that is working fine and doesn't require a vice president to check in, you know, every two weeks, that's great. But some things require that executive oversight like nothing happens less than executive is yelling at people and then you slowly that might move away and that's important in the app SEC program, especially where these are often new and novel ideas you're going to transition from you have no process to it requires some executive oversight to the executives.
Just trust that it happens. And so let's think a little about what measurements you might use for these, you know, inventory was easy, but we think about source code analysis and application security testing. We're probably going to want to think about code coverage, you know, if you're doing your dast you need to make sure that you're actually exercising all of your lines of code and not just exercising a given API.
But you want to really understand like what is being analyzed here, right? If it's SCA, you know, is this SCA? That's just doing open source detection, or is it SCA that includes hygiene detection that's going to find vulnerabilities for you.
And so think about that a little bit as that code coverage from a code review perspective, you know here you really want to think about how effective is it many code review programs are in my opinion worthless actually of negative value because what they end up being is a second set of eyes that is incentivized to look very quickly and never find problems. And to accept that for a moment that if you have a code review process that is someone just has to check a box that says that they reviewed this. Well, that's probably most of what's going to happen.
The problem is that most code review Systems are designed to detect adversarial and actually weak adversarial failures that this is let's just assume that the first person who wrote it decides to write in this obvious back door. We want somebody to code review it to look and see that there was no obvious back doors. But is this really instead should this really that functional analysis that the second person you know, maybe is reading the requirements specification sketching out pseudocode and then checking the code against their pseudocode to make sure this does what they wanted to do.
Do they have a list of safety checks that should be in any piece of code that they're verifying and validating that they exist. So when you're building your code review program, that's a really important thing to focus on which is how do you make sure that the code review is providing value because it's not here's the dangerous point a developer who knows code review happens, but has it really thought adversarial about whether the adversarily about whether the code review program is effective will actually take more risks. That's what's known as the pelce manufact or risk homeostasis in the presence of a safety control humans will behave more riskally because the safety control provides them with protection.
That's the point of most safety controls. We have seat belts and breaks in cars so that we can drive really fast. Well, if you think about code review as a seatbelt, but if it's only a seatbelt that the code the code writer believes exist, but doesn't provide safety.
Well, they might you know, take shortcuts knowing that the code Reviewer is going to find it even if the code reviewer does it so really pay attention on that when you think about your pen testing and Bug bounties you hear you really want to look at? What is your true positive rate? How often are you really finding things that are interesting and novel there that you couldn't have found internally because otherwise they're not really augmenting they're so they're a compensating control for a different failure that if things you probably should have found in your application security testing.
For training this one really should tie to sort of the defect classes that you want to make sure get detected or get fixed. And are you detecting those in other parts of your application security testing? Waffenbot mitigation.
I think here one metric is really, you know, the virtual patching. How often does your WAFF provide you defenses against an attack while you go fix the application. That's a really important one.
Obviously you want your applications to be safe, but we should recognize when something like, you know new vulnerability comes out that it's sometimes faster to have the laugh detect it. Do you have bought mitigation and in what ways has bought mitigation protecting you And lastly for cicd you should think about all of the gates in your software development lifecycle all the ways the code could be altered. And how many of those are actually being protected by your ci/cd system and that's how you should think about your defenses in this standpoint.
And then last you should document your process maturity, you know for each of these I didn't write in an actual metric but let's first start with what is the maturity of those processes and until you have a bright light on them and you have a measurement for them. You shouldn't rest and say oh, I think I've got it but that's how you build the apps that map you might have something qualitative like this. You can go back to absec math and I love playing with this one because you can just put in like who your vendors are for each of these programs to give you a visual view for any application what your application security program looks like now, this is important because you have to be able to tell stories Right and let's think about the depth problem, which is when an adversary breaks into a system.
They rarely want what's on the edge system they want what's behind it? Now advertise don't often know what's behind it until they show up applications are one place where they might already have a hint but sometimes they attack one application find keys to a database that's you more interesting than the application they had. And that's where they're going to go ahead to.
So we need to be able to tell ourselves stories about how the attackers operate and then def how our defenses are actually stopping them. Some of a big fan of writing up an attack scenario. So foreign attack type you're defined your defenses and you're gonna tell a story including your instant response.
So let's talk about account Takeover in the apps that context, you know, we might talk about a persistent cross-site scripting attack that is present and available in some user page account management and attacker payloads that you know, either copy the session authenticator or make a call to an API to get their own token that they can then reuse and then they're gonna extract, you know information out of this application, right? And this is the thing we see all the time, you know, whether it's just reusing credentials that we're breached in the wild or something like Mage cart that's you know, running in browser and doing something on the cross site. Now we know their ways to defend against this, you know, you can use a WAFF, you know with a page Integrity module that's going to look for some of these attacks going on or look for interesting velocity things that don't look like Accusers we know there's ways to this encoding so it's a training control to implement, you know, some form of, you know, dual confirmation similar to csrf or maybe crsrf would be sufficient for you.
And then yeah, we also know that it'd be really nice if we had better identity and account management so that we could an authorization so that even if somebody's identity was stolen which you know, it's gonna be harder in this in a two-factor world that it's much harder for someone to go get things that they're not supposed to and maybe there's additional mitigation mitigating controls here with that source code analysis and Pen testing and a bug Bounty like place a bug bounty on this type of attack so that if all of your controls fail somebody in the outside will show you what it is and you can pay them directly in cash rather than paying through having a bad incident happen to you. And so now you can sort of tell a story and these are important to sell the value of your absec program. But if you want more money you have to be able to demonstrate value.
This is a great way, you know, we train our developers specifically this kind of attacks. We're reducing our user privilege so that we don't have to worry as much about this, you know, we have implemented page Integrity on the Waf our bug Bounty program specifically pays a bonus for defects found in this area. And maybe here's the vulnerabilities we've found in the last three years and we have patched in mitigated.
And this is going to be one of the most important things is narrative for selling your program to make sure that people understand that you've built an effective program and you've tied this to business value, right? They used to be able to extract lots of data and now they can't straight business value right there. And then lastly very important for us to think about time.
How do our controls get better or worse over time? And they're really important question to ask and I'm going to go back to vulnerability management is you might have a way for people to get Exempted from this metric. Right.
It's a seven day SLA. Well what happens if you can't fix it in seven days because it's impossible. You usually have some way for someone to get an approved exception and now the seven days doesn't apply to them and if so, they don't show up in this 85% number for good or bad, right?
They're not in the 15% bad. They're not in the 85% good. They just got removed.
And what will after a Time happen is your developers will recognize this and really not maliciously. They just don't want to get in trouble. They're gonna show up at six days and say we couldn't fix this it's going to take us at least another four days and the real answer is it would have taken them four days if they'd started the first day they waited until day 6 to say they couldn't And so you should start measuring how many violations are escalated in time enough to make a real decision.
If it's a question between four days to fix it and three weeks to fix it. You should make sure we're aware of that by day three so that we can make a reasonable decision as an executive leadership team about whether this one needs to be fixed at high priority or not. So remember defend with coverage comprehensiveness context and have continuity over time and that's the important way to think about your security program and specifically your application security program.
And with that I'd like to say thank you all for your time today and a very shortly you're going to get to talk to Alan in honor of his Pittsburgh Steelers actually winning on the last second. I decided to wear a few colors for him today. So I hope you appreciates that not out to him.
So have a fantastic day everyone.





