Why Cybersecurity Risk Scores Often Mean Nothing
Security dashboards are designed to simplify complexity, but in cybersecurity, that simplification can quickly become distortion.
On this episode of Security Boulevard, Tom Hollingsworth, Fernando Montenegro and Jay Cuthrell take aim at one of the industry’s most persistent bad habits: reducing cyber risk to a single score. The conversation starts with a deceptively simple number and quickly turns into a larger critique of how security teams, vendors and executives often rely on metrics that look precise but lack the context needed to support real decision-making.
The panel explores why many cybersecurity risk scores can be misleading on their own. A number on a dashboard may appear authoritative, but if the methodology behind it is unclear, the output becomes difficult to trust. Without visibility into the assumptions, weightings and business context shaping that score, organizations may end up reacting to optics instead of actual exposure.
Tom, Fernando and Jay also discuss familiar examples such as CVSS, noting that even widely accepted scoring systems break down when they are separated from environmental and operational realities. A severe vulnerability affecting a small, isolated asset may deserve less urgency than a lower-rated issue tied to a critical business process. That is where raw severity scores stop being useful and context has to take over.
The episode also examines the challenge of communicating cyber risk to leadership teams that want a clean, digestible answer. While simplified dashboards may make reporting easier, they often create false confidence. More mature approaches to risk quantification require a clearer understanding of business impact, likelihood and exposure, rather than a single abstract number presented as definitive truth.
For security leaders, the takeaway is straightforward: Metrics can support better decisions, but only when they are grounded in context, transparency and business relevance.
Transcript
Welcome to Security Boulevard, the cybersecurity podcast from the Futurum Group. Each episode explores a variety of topics within cybersecurity and the technologies that drive it. com, the Security Boulevard YouTube channel, Techstrong TV, and all of your favorite podcast platforms.
Let's meet today's panel before we jump into the episode, starting with Fernando. Fernando, it's good to see you back from RSAC. Thank you very much.
We're recording this shortly after the conference, and I think I have a little bit of con flu, but it was so nice to see so many people in San Francisco last week, and I'm sorry for the ones we couldn't meet. But in this industry, we have tons of events, and there's always Black Hat starting to show up on our radars already. And before that, for the European crew, you have InfoSec Europe, which is also very nice of it.
And joining us for the first time on "Security Boulevard" is my good friend, Jay Cutrill. Jay, tell everybody who you are. I am Jay Cutrill.
I'm the chief product officer at Nexus Tech. I deal with the product marketing and analyst relations part of the business. And last week, I had a chance to be in Raleigh Durham, North Carolina, for all things AI.
So while everyone else was soaking up RSAC, and Bsides, and the wonders of security on the West Coast, I was holding down here on the East Coast. It actually turned a little chilly, but that was an amazing moment, seeing 500 people lined up outside the door trying to get into the Durham Convention Center and learn about all things AI. This was also, of course, during the...
If you follow the stock market for the cybersecurity stocks, you may have noticed that little zip when someone said Claude, or Anthropic, or other mythos model out loud, and the market did what the market does, sometimes overreact. So, great to be here. Appreciate being in the invite list today.
Thank you so much. Market's overreacting? Surely not.
I've never seen that happen before. It's almost like numbers matter or something. Ooh.
Let's jump into today's episode, because it came out of a conversation I was having with some people at RSAC, and it kind of cracked me up. So I'm just going to throw this out there, gentlemen. 76.
Does that number mean- You should have said 67. 67. I could've said anything that wouldn't trigger my kids or any other combination of numbers.
You're probably wondering to yourself, 76 what? "76 Trombones" is what you hear when you march down the street? Big parade.
Spirit of 1776? No, it's a score on my dashboard, and it says a thing, and that 76 is important, or maybe it's not. " And I want to talk about the fact that all of these scores are hot garbage.
They don't mean anything to anyone. " And I'm like, "Such as? " So have you guys ever run into these things?
I'm sure you have. It's a feature that everybody touts. It's a big number.
Oh my God, yes. And we run into these in many different ways, this quantification, gamification aspect of cybersecurity. And there is a lot wrong with it, but there is also a lot that goes right with it.
So I'm dying to get more into the topic. But I'll say this, yes, we've seen this forever. Those of us who have been around risk management in any sort, we've been looking at those scores as a component of risk management decisions, and with varying degrees of success.
And there's so much to talk about. Jay, as we get to know each other, I won't monopolize the call too much, and I'll yield. But I'm very excited about the topic.
I just want to know what flavor of lifesaver that is. I'm used to the magic lollipops. Sometimes you have the red lollipop, which I guess would be the cherry flavor, the yellow, which apparently is sugar plus yellow coloring, and green, maybe, I don't know, is that a light pine tar taste, but ever so sweet?
I think there's a false precision in math of saying, what is 76? Is 76 on a scale of one to 100, which I think we would all kind of assume, but I also have access to other tools that have a 200-point scale, just to prevent someone from arbitrarily thinking 76 is a good number. And you have to do math in your head to double that, and then there's 140, 150...
Yeah, so you can tell I'm already bad at this. 152. Is that good, bad, relative, indifferent?
So that's the other part of this is, is it or isn't it good? Is it or isn't it bad? Or is it or isn't it really telling us something that it's implied, but you really have to go into that second or third click?
And if so, is the color misrepresenting? Is the number, by virtue of it being a number, effectively misrepresenting? So I also am looking forward to a robust discussion around why each of us is right and wrong at the same time.
And I think it's important to bring up, because you guys have already talked about some of thevagueness in this, you have to know the scale. " But also, I love the stoplight system because to me, the stoplight system is one step beyond making it completely useless. Ooh, light red, bad.
" And I feel like these stoplight systems or these risk score systems, at least initially, the idea was egalitarian, right? There are so many inputs into all of these things that we have to figure out how to make it more digestible for the non-technical people. We have to find a way to take the various inputs, is this important?
Is this critical? Is this something or other? And make it so that someone at a glance can go positive, negative, good or bad.
But I think that we've overstressed that system now that everything has to be boiled down to a number. Well, is 76 good or bad? Well, that depends.
Is 75 and below red? Is 76 and above yellow? It's like when my kids get scored on tests.
Because it used to be that 90 to 100 was an A and 80 to 90 was a B, but now it's 93 to 100 is an A and 85 to 92 is a B, and I need to know how they break down because a 92 could be both good and bad depending on how I look at it. And then the problem is, is that I know that as someone who deals with data all the time and as someone who loves to label the axes on the graphs, but I'm not the one answering those questions. " And so there's a lot that we can parse through this.
I think let's take a many, not one, not two, but many steps back, right? I think that fundamentally we want to do these things as a combination of we're doing decision support. Right?
Let's get this out of the way. We want a number, we want not even a number, we want the information because we want to make a decision, right? Is this something that we need to act on?
Is this something that we're not going to act on? What's our comfort level with acting on this particular thing? Right?
And that is a perfectly valid need from anybody doing cybersecurity or IT or anything else, right? I joke that economics, right, here we go, is the study of scarcity, right? And the first rule of economics is that there's always scarcity.
You never have enough resources to do everything you want. How do you prioritize what you want, right? And then the joke is, so first rule of economics, there's always scarcity.
First rule of politics, ignore first rule of economics, promise everything to everybody, and then you're good. Digression. But I think that the idea of having scoring, it's an element, like I said, I see it as an element of decision support.
And I'm going to be optimistic in that I see this use of numerical risk scores as an attempt to elevate the conversation, which we need, right? We need to move the conversation at the board level, beyond red, yellow, and green, right? Beyond risk is high.
Now, my CISSP is back from many years ago, and I remember that we were doing calculations on risk score. Like five plus three is eight, and that means it's a medium. That didn't make any sense.
You don't calculate on ordinals, right? And so I see this scoring as an attempt to do that. Now, there are two scenarios here, right?
And there are the scores that we see, there is a value to those scores as a progression over time. I've dealt with many vendors who have, "Oh, okay, we have a risk score in your dashboard," and you can see the nice squiggly line, it's trending up or it's trending down. And that may be valuable or not, depending on what goes into that score.
Right? " I'm going to pick on the numbers. Right?
Okay, fine. Does that mean it got worse? Why did it change to 67?
Did it change to 67 because the aggregate health of the devices dropped for some reason, right? So did the reason drop because the devices we have are weaker, are materially worse? Or did we add new devices to the environment that haven't yet been patched up to the level of controls that we have, and therefore they are weaker?
So again, meaningful intentBut the execution is kind of weird. But yeah, I'll say this, I remain positive, in the use of more precise scoring. I'll yield the floor to Jay.
I'll come back later. Yeah. So I would say that there's things like a defense readiness kind of condition, the so-called DEFCONs.
And so you may be at DEFCON Cookie Monster, which is the five, blue, or you may be all the way up to DEFCON Elmo, DEFCON 1. And there was even a meme, if you remember, back in the old Twitter days, where it was when maybe it was Department of Homeland Security or the TSA apparently had these different color levels of situational awareness for travelers. And it always seemed like we were constantly at a state of Elmo.
And if you're always at a state of Elmo, are you really at a state of Elmo, or are you just over Elmo'd out? And I'm not picking on Elmo specifically, but it's these conditional views that are entirely subjective at some points. " Well, how long can you sustain that effort?
Is that sustainable over a long period of time? And so I see some of these other interfaces, especially in security reporting, where you talk about if it's a numerical quantitative score, relative to what? If it is a quantitative score, relative to what?
If it is a qualitative score, do we all agree on the definition, what it implies? By the way, if we're culturally expanding beyond just, say, the United States or a specific country, is your sovereign country definition of how that color or story or iconography should be interpreted consistent globally and accepted? And one of the other things I was thinking through is, in the early days of the web, there was this common concern around the many dashboards that were being created, because the web was moving so fast, you had infinite ways to think about visualization of data, and multivariate data, even more complex.
So if you're not familiar with Chernoff faces, basically happy face, smiley face, versus flat, versus frowny face, and possibly moving their eyes or their eyebrows up in the cartoon to show or convey different types of qualitative and quantitative measure. But a Chernoff face is interesting. It's a provocative way to think about it.
" Is the server happy? And not a server in your restaurant, I mean a server probably in your data center. And if the server's not happy, should the front of the server have a Chernoff face?
And this was in the world of the physical, kind of one-to-one, if you remember the pets versus cattle, this is a pet type server. " Is that enough? Or is it like, no, there's four or five of these that are really unhappy Chernoff faces.
How do you find those five and determine whether or not that's in the blast radius of whether you care or not? Or is there an architecture that allows it, or to permit it to run and then get remediated later? So I do see challenges both in the qualitative and in the quantitative that we spend a lot of time on, but I think even the qualitative measures are still challenging, like whether it's lollipop colors or what you think is the flavor of whatever is on...
Actually, you mentioned the traffic lights. I've never climbed up to a traffic light to try to lick the cherry red light. So maybe that's on my bucket list now, but I do know that red, for me, means stop, but it might mean something else in a different culture.
And I'm glad you brought up the fact that the scales that we've been using forever are always off, right? The DHS one that was released was comical because at no point were we ever lower than grade three, which was heightened readiness. It's almost like the other things existed just to make people feel that- We were at Big Bird for a long time.
Big Bird was a problem, right? Yeah. Big Bird yellow.
Yeah. But, and go ahead, Fernando. No, I'm just going to say, I love to throw quotes at people, right?
" Right? And Tom, I have a visceral reaction to you mentioning it never went down from something because the incentives in the system are misaligned in that it's easier for you as the creator of the risk score to report a higher risk, right? Because it's safe, right?
And again, I go back to decision support, right? One of the best books I ever read on decision support is "Thinking Bets" by Annie Duke. I highly recommend it.
Right? And one of the first things the book talks about is the notion of resulting. Humans are bad at conflating the result of an action.
We conflate the quality of the decision with the quality of the outcome, right? We don't control outcomes, we control decisions, right? And if you make a good enough decision, you maximize your chances of the outcome going your way, but you can't guarantee it, right?
And she uses that famous example of the Super Bowl where the, I think it was the Seahawks tried to throw at the one-yard line against the Patriots a few years ago, and the Patriots intercepted it and eventually won the Super Bowl. And the outrage that people had against the coach. Well, but it turns out that that was actually the right decision given the circumstances.
The outcome didn't turn out what they wanted, right? But it's the same thing here. Like if the outcomeIf people measure the outcome of a decision by tying that to the quality of decision, they're making a mistake.
And I'm sorry, I went off complete different tangent as usual. But Tom, you're right about the level. It never went below something because it didn't make sense for somebody.
Somebody would be safe reporting that it was a high enough level. And we deal with this all the time in a lot of things, right? One of the things that I remember when I first started working with these monitoring systems was the fact that we had to turn off memory monitoring.
" And it's always going to be red because, ooh, it's above 90%, therefore it has to be bad. No, I would rather it be 90% because then that means it's being used efficiently or things like that. And one of the problems that we always run into when we're doing that is, again, we know context behind things.
Another way to look at this is, I think one of the original things that made me think about it was CVSS scores, right? 5? And you're like, "Oh my God, that is absolutely horrible.
" Wait, this only happens on a really random product that was sold in, like, a lot of 100 to one company in Antarctica. But if they could get into it, it's really bad, therefore it's a 10. I'm like, it's a 10, but the likelihood of it happening is, like, a two.
There's no context behind that. Likewise, oh, well, that's a six and a half. That's not a big problem, except it's a problem on every device that's been sold for the last 85 years.
You got to patch that. It's giving people this false sense of security when you publish a score and you're like... Because let's be fair, if you went to RSAC, you walked through the booth, I should have just taken a picture of every, like, score says 74, you're okay.
Score says 94, you've got a problem. And just the usual, like Jay said, is that a scale on 100, 200, 1,000? Is it another mathematical calculation?
Is that graph a logarithmic graph? We have enough problems with people who don't understand math on a daily basis. Thank God none of them work in finance.
They all seem to work in the executive, though. Because the only math they understand is, did bottom line number go up? And when you get to security scores, they really don't care.
What they want to hear is, "Everything is fine and we're not going to get hacked. " And if that does happen, they're going to want to know why, and they're going to need details, which is why every one of those numbers, when you click on it, should produce a report that says, "You're not at 100 because your CEO's password is bad. " If you can't provide even just baseline context for people, then you're doing them a disservice because again, number must be good because number green.
God, yes. And I go back to conversations I've had with vendors in the past. The composition of that score is absolutely essential, right?
I remember being particularly peeved at a vendor whose score composition was the number of endpoints that had their product installed. Sure. Not the number of endpoint that had endpoint security installed, not the number of security, the number of endpoints that had endpoint configured.
The number of modules of their offering was a component of that risk score, right? I'm sorry, that's a pre-sales thing at best, right? Mm-hmm.
And so those scores are extremely confusing, right? And when I work with vendors, I go, "Look, we are trying to help this industry overall reach a higher level of maturity on these things. That kind of score that you're doing does not work toward that.
" Not only because it is the right thing to do, but because we are seeing people, there is a movement, there is an entire movement in this industry of getting better, right? And I would be remiss if I didn't call out the phenomenal work that SIRA, the Society for Information Risk Analysts, SIRA does. I was a volunteer at SIRA for a long time, and I absolutely love the organization, right?
And with that comes significant knowledge that is being shared in terms of what does good risk management look like, right? So anyway, I'll refrain from- But that's a good point, is going back to what you said. These people deal in risk all day long, and so sometimes they have to quantify that risk.
And that's- And they do amazing work at it. But what does the quantification mean? It means that they've plugged a bunch of information into an algorithm or to an equation, and the number at the end is effectively normalized, and that's important within their context.
My problem with all of this risk score stuff is that you have effectively dumbed it down to the point where nobody has any context for anything, and you like it that way because you need to sell training or professional services to help them understand why the number is the way that it is. " Okay. You had mentioned earlier that we tend to be bad at math, and if risk is another way of saying that you can quantify or do the math, then it also means we're kind of bad at quantifying risk.
Mm-hmm. And so even if you go back to the-- It's an acronym, FAIR? Yes.
Right. That means that there's some kind of a factor that you have done analysis on. And from that point of view, the information you're talking about, you've determined risk.
" But on the other hand, it's $10 million of risk. And I guarantee you here that someone's going to understand the $10 million of risk. They won't care about the 100 whatever Linux things.
It just, p**f, just go right over their head, right past them. The question then becomes is, are the vendors supporting and driving towards something more like FAIR, or is a proprietary view and visualization within their product experience going to perpetuate a renewal? Are they really just thinking about like, "If I can get you to adopt this thing, I know you can't get off me now.
" And because everyone's learned what 76 means, 76 means exactly what it means to you in your company because you have used my product for so long. So is the 76 also a kind of an entrenchment, digging in? Or if you could get away from 76, going back to the 76 example at the beginning of it, does that mean that maybe that there's interoperability or shared understanding of how relative risk can be summarized amongst different disparate vendors that ultimately make up what might be that dashboard of doom for the CISO?
And Tom, may I say that if I'm ever not available for economic commentary, Jay here has brought up switching costs. So thank you very much, Jay. That is spot on.
Right. But I'll go back to something. I think that this is positive, like that the overall movement is positive to move us from a heat map, to move us from a color, right, into something more defensible from a statistics perspective, right?
And I think that we are seeing more risk management practitioners get their word out there. I'm going to do a shout-out to both Richard Syerson, who wrote the "Metrics Manifesto" back in 2022, I believe, right, as somebody who's been doing some phenomenal work on this, as well as Tony Martin Veiga, who has a new book called "From Heat Maps to Histograms" that should be coming out this week, actually. So I'm eagerly waiting to take a look.
And both Tony and Richard have done phenomenal work in the context of CERA, of doing precisely that, of helping to find common ground between environments. At least from where I stand right now, what I see as the best approach we've come up so far at that highest executive level, right, is something that Richard has advocated, which is something similar to value at risk, right? Which is the notion that you build your security controls using FAIR or other mechanisms, right, or other methodologies, and you come up not with the number 76, which I still think we should be using 67 as an example, but that's okay.
Right? And the number you come up with is that, look, given our current risk profile, right, there is a 5% chance that this organization will see a loss greater than $8 million in the next year. 5%, $8 million, there's the curve.
Is this number acceptable to you, dear board? Oh, no, this number is too high. Okay.
In order for us to bring this number down, right, what number is it? Is it the 8 million, or is it the 5%, or is it both? Okay.
Based on, oh, we want you to bring this down to 2%. Okay. If we're going to bring this down to 2%, here is the set of controls that I need to put in place, and not only the controls in place, but the operational practices that we have to have in order for this number to come down.
But that is a much more sophisticated risk conversation than 76 or 67, right? When you mentioned that, it took me back to the '90s. There's a TV show, "Saturday Night Live," here in the US.
Oh, wow. And now there's an export of that, I guess, to the UK. They're going to do "Saturday Night Live UK" now.
Yep. " Pre-tapes are when you have a figurehead, speaker, whoever is the anchor, and they're going to go on vacation. And if they're going to go on vacation, there might be some things that are very likely to happen where you want to have something in the can just in case, like key legislation passes- Yeah, yeah ...
someone passes away that's on their deathbed, very morbid thing sometimes. But you do the pre-tape so it appears as though you're still there and available even if you're on vacation. So in the skit, when you go through that, you watch this, I'm going to give a spoiler alert, it's hilarious.
There's also some not safe for work things, so don't watch this at work if you're in a sensitive environment. ButIt does highlight that if you have these systems in place, what happens if you do have people that go on vacation? Is there only one person that really intuits or understands?
Or again, is this more organizational? If it's organizational and we have common understanding, I think that's great. But I also think as you go through different parts of the organization, you may have different levels of, this is very high risk from my perspective, but from my perspective in this other part of the organization, that might, for the same exact event or concern, might be very low risk by comparison.
And so that's another important context is, how are we surfacing these numbers? And I like 67 now that you've convinced me of that. Is 67 the same 67 awareness in different parts of the organization, in the topology, or is it really only ever going to make sense, is the 67 really only 67 just for the teams in security?
I think we could spend another half an hour just debating all of this stuff. I think what needs to happen is if you're out there and you're listening to this episode, go to the vendor and ask them where the number comes from and why the number is the way that it is. And if they can't justify the number, then they're probably just making it up out of thin air, and maybe it's time for you to do a little bit more investigative work.
I know that we just got back from RSAC. You're probably listening to this episode a couple of weeks out from RSAC, but just know that we've been putting a lot out since then. Fernando, what are some things that you've got coming out that people should check out from maybe from RSAC or maybe more in general?
Yes. By the time this comes out, we'll have our main note out on the conference. I also use the conference to inform a lot of our research, of course, and related to this topic.
So the whole topic of cyber risk quantification, very close to this is third-party risk management, right? That is something I'm going to be digging into over the next few months, right? I love how we are maturing these disciplines.
One of the things, for example, that we're doing is we're starting to take, related to this course, we're not talking so much about what is my outside in view. " It's like, okay, what do your internals look like? The inside out view, and then start to put those together, right?
But sorry, I digress. I think that coming over the next little while is, Tom, you and I have a report out on secure access service edge that should be out momentarily or next few days, weeks. Outside of that, we have a refresh of our security operations signal report that's going to come out in the early summer.
And then, as I said, a few months from now, I'm diving deep into CRQ and third-party risk management and other areas. So I'm really looking forward to that. And J, I know you've been producing a lot of content as well.
If people want to check out some of the stuff you're doing, where can they go to learn that? org. org, my blog there.
I also have newsletter content as well as a podcast. I'm on my third episode of that. I'll also be at the Click Connect for 2026.
I'll be with Brian, Frederick, Gina, and others there again. This will be my third year. This year's going to be exciting because the element of AI reproducibility, security, validation, trust in the actual AI experience will be top of mind.
So we're well past the training wheels phase. People are producing this stuff now in a production environment. So I'm looking forward to those stories as well.
com because we have great videos that we've published regarding our Techfield Day Extra at RSAC presentations from companies like Veem, ObjectFirst, and Commvault. com because we'll be highlighting all of the coverage from our delegates that were there. They have some interesting thoughts.
" If you enjoyed the conversation, do us a favor, subscribe on YouTube or in your favorite podcast application. We don't want you to miss any episodes. Also, leave us a rating, and a review, and a comment.
That always helps the show grow and reach new audiences. com and the Futurum Group. com is the place to go.
You can also head over to the Techstrong TV website or check us out in the Techstrong TV app that's available on Apple TV, Roku, and other smart devices. Follow Security Boulevard on X, Twitter, and LinkedIn. They are @securityblvd on all those platforms.
Lots more content for you to enjoy. Thank you very much for tuning in, and we'll see you all next week.