AI Agents Corrupt Workflows, Coding Gets Rules and Security Starts Negotiating
Mike Vizard, Jack Poller, Jeff Reich, Jon Swartz and Tracy Ragan break down Microsoft’s warning that AI agents can corrupt long workflows, the rise of spec-driven coding and the security consequences of the exception economy.
Transcript
Hey everybody, happy Friday. We're here with the Techstrong gang having a chat about all our favorite topics, many of which include AI as usual. But let me introduce the gang today.
We have, of course, Jon Swartz is joining us from California. How you doing, buddy? I'm doing well.
Happy Thursday. Happy Friday. Okay, John, you're ready to work one more day this week, are you?
No, I'm not. I'm done. All right.
Also joining us today is Jeff. Jeff, how are you? Good to see you again.
Good to see you as well. I'm looking at my calendar confirming it's Friday, and I'm happy. All right.
Jack Poller is back home, and I saw Jack earlier this week here in New York. You paid us a visit. Jack, how you doing?
Well, after that trip, it is definitely Friday, even if it's not Friday. Well, there you go. Did it take you a long time to get back or what happened?
No, it's just conference season. It's busy. It is true.
The information overload in the brain. My brain just... I'm done.
All right. And of course, Tracy Ragan is in... Looking in the background, that is New Mexico, right?
Not a fake New Mexico. Yes. I am in New Mexico, enjoying a lovely day, and I plan to go for a horseback ride very shortly.
Ooh, that's very nice. All right. Awesome.
Well, let's just jump into this, because there's a fascinating bit of research that the folks at Microsoft did where they were looking at agents that were being assigned tasks that were running for a very long time. And, well, lo and behold, they discovered that there's a tendency for these agents to corrupt the data that they encounter over time, and that does not sound like a good thing at the end of the day. John, I know you wrote this story, but are we kind of starting to see where the limits of these AI agents are?
Yeah. That's just been a theme it seems, hasn't it, for the last several weeks, last several shows, based on studies and discussions we've had and some of the topics, is that there's a limit to what these agents can do. But also, I think, to echo what Alan and I think Fred Wilmot said yesterday, we're pretty early in all the process, so it's a learning process.
It's touch and go, trial and error. But yeah, you're right, Mike, this was a really fascinating study, I think. So the Microsoft researchers came up with a benchmark called Delegate 52, and they contend that the frontier models, and they're talking about GPT, Gemini, Claude, they can silently, they said, damage work artifacts when they carry out these multi-step workflows.
And their conclusion was pretty stark. " So I'll give you an example. They're talking about these models and others losing 25% of a document's content over 20 interactions and degradations reaching 50% across all models tested.
They're talking about catastrophic bursts of failures where a single step could wipe out 10 to 30% of a document's integrity. And they mentioned that adding agentic harnesses made outcomes even worse by an additional 6% on average. So, there was a bit of good news.
They reported that only Python programming met the bar, which I guess would imply that document-centered knowledge work is not ready for hands-off agent autonomy. But what it does underscore is despite all that we hear during this conference season, and we're hearing a lot, there are some definite flaws and shortcomings that are being discovered in the field and on the ground among the companies and this is something that has to be addressed. And I'm glad the researchers came up with it and as we move forward, because we're moving forward at an incredibly fast speed, and it's a little bit overwhelming to, I think what Jack would say, this whole thing is overwhelming, this information overload.
Jeff, I'd love to get your opinion on this because I'm trying to figure out if this is a moment in time issue where the AI will continue to get smarter and will just get better at these things because some people will say that the worst AI we have is the one we have today. Or does this give you cause for pause? Well, there's certainly going to be some of both.
Anyone that says, "AI is great, I'm comfortable with it, just let it run," is foolish today, because we need to remember that AI is in its infancy. " So treat it like an eight-year-old. You can give it tasks, you can watch what it does, but if you just leave it alone for too long, it's going to wander off and do something it shouldn't.
And I think that's a good way to look at AI right now. I'm not saying it's bad. I'm saying it's immature, and we have to help it grow.
Oh. Tracy, you going to be comfortable turning loose AI agents on software development life cycles? There's a tendency among those folks where they love to let things run for a very long time, regardless of how much it costs, but is that a good idea?
I want to. I would very much love to, but the thing that I read in this article that kind of pinged my brain is that these are not failures that are easy to find. These things just aren't breaking.
I would call them non-dramatic failures, which we are not used to. We're used to things breaking, right? And it's not that they're breaking, they're just continuing when they have these multi-step workflows.
So I think it's concerning for any enterprise that's trying to adopt AI agents for software development, which many of us are. Certainly trying to automate cybersecurity or anything around operations, because the corruption may not be noticed really, and it may not get discovered until it impacts somebody downstream trying to make a business decisionAnd you know how we are as people. AI-generated corruption isn't especially dangerous because employees tend not to notice it if it's polished output, right?
Mm-hmm. If it looks good and it feels right, we tend to believe what it says. So this is a new world for us in trying to understand the data that it's generating, and the results.
Is it correct? It's like talking to somebody who acts like they're an expert, and suddenly you realize that maybe they're not actually an expert . But boy, they carry all the credentials to be an expert, and then you just wanted to ask them to stop talking .
Mm-hmm. And that's the problem I see with it. But I really, really want to start using these tools to automate what we do, in particular around cybersecurity.
It would be super helpful if we could trust it. Mm-hmm. And maybe the challenge here is that the implication is that we'll eventually have to learn how to understand these long chain behaviors.
And how do we do that? I don't know. Jack, by definition, aren't all cybersecurity workflows and tasks essentially long-running, as in never ending?
So, what's going to- Yeah, but that's not the real problem here, and I'm going to take a contrarian view on this, which is, I think the study was completely flawed because it doesn't put it in perspective. They failed to take into account what humans do and what's the probability that a human makes subtle or not subtle errors in the same long-flowing process. Now, I work on the marketing side of the world these days.
I used to be an engineer, like Tracy. And in software development, it's very obvious when somebody makes a typo. Things break.
Either the compiler breaks or the program goes haywire very quickly. NoWeReallyMeanItFinal, because so many people's hands get on it, and when you actually go and review what you get back, it's completely different than what you wrote two days ago. Yeah.
So, this happens all the time on the human side as well, and I think it's not fair to put out a study like this and just put it out there and say, "Oh, the world is coming to an end because AI- Oh ... " I think there's a parallel, like that telephone game. Remember that game where we- Yeah ...
would say something, this says- Exactly. That's inherent. And Mike and I can vouch for working in communications companies, they're among the worst at communicating, I think.
So much- Right? But if they were like this innate ability to kind of screw things up when we talked among ourselves, especially when I worked in a newsroom, it was always just... I'm thinking newspapers, just chaos, and by the time something got to the third person, you had no idea what you were supposed to do, or you did it, and you were wrong, or you were just slightly off.
So what I'm wondering, though, is given all that and the fact that there is definitely what's happening in workflows that these guys discovered is not that uncommon, should some of these tech companies that are pushing their autonomous agents or kind of really hyping them, should they I'm not saying provide a warning label, but publish some sort of reliability scores along their marketing claims about these long-form workflows? I think here's the issue. We're definitely holding AI agents to a higher standard than we do humans.
And a lot of that has to do with the marketing messages that comes with it. So there's an assumption that these things are going to be machines and going to do things consistently the same way. As I look at it, though, I start to wonder if it's not going to be similar to what Jack just described as a workflow involving humans, except it'll be a workflow involving multiple AI agents that are assigned the same tasks and the same reviews, and they will surface inconsistencies as they go along, and then there'll be a markup of a document going back and forth between the AI agents till we get some- Mm-hmm ...
in here. And it's not just going to be one AI agent that's assigned a thing. It's going to be essentially a swarm for every little thing.
Now, as I work that through in my head, it gets very costly to have swarms of AI. So at some point, does the cost of the thing exceed the probability of just hiring somebody at 60 grand a year to go check the data? Kind of like a delegator of some sort, maybe.
Almost- Mike- Yeah, sorry. I was going to say, you remind me, though, that even this is not a new problem. Even in old days of computers in mission-critical environments, we actually do do multiple computers that check work on the space shuttle or an aircraft.
You actually have two or four, two, three, or five units that actually vote on every decision, and if all three agree, then the decision takes place. If one of them disagrees with the other two, that computer gets kicked out of the environment. That's how the space shuttle runs.
That's how a lot of jet fighters work when you have computer control of flight surfaces and other things. So this is not a new problem. It's just- Consensus.
Yeah, right. Consensus LLMs, right? Right.
Maybe that's the way we have to go. But we do do this in the computer industry and other environments. I want to- So I don't want to throw the baby out with the bathwater here.
Like Jeff pointed out, we have to treat AI right now like it's an eight-year-old. And the study did point out that AI performs better in kind of a highly structured environment like coding, but the natural language and that kind of loosely structured workflows really degrade more quickly. So it's just immature, and I believe we'll figure out how to fix it.
I have faith. I do think... I'm sorry, Tracy.
No, I have faith. We'll get there. I agree with you.
I do believe we're going to get there, and I don't think it's going to be done by agents watching agents watching agents either. I think there's going to be more of a distributed intelligence, rather than a centralized within each agent. Which allows you to, and sorry for getting geeky, if we perform differential equations on what we're producing out of AI agents, we can start seeing where does the variance start.
Because where we get the corruption is especially because of long-running processes, a small variance which expands once it reaches its target. And it's being able to identify those, and I think distributed intelligence may be a way to do it because one, it's going to address the cost issue, Mike, that you brought up. Do you really need a swarm for everything?
And it also allows the different perspectives from different agents to be able to create, for lack of a better term, a consensus. So there's my prediction for the day. I've been talking about consensus for a while.
It's like we have to have a way to have a consensus. There's also a movement afoot to shift the context and the maintaining of the relationship between data and the integrity outside of the LLM and outside of the AI agent. And there are folks talking about everything from graphs to indexes to newfangled relational databases that, in addition to keeping track of the relationships between things, are also coordinating when there is something that is out of sync.
And so, I guess the question sometimes I'm starting to have is, are the guys who are building the LLMs are basically saying, "Yeah, we can do everything inside the LLM," and there are some other things that maybe they shouldn't do inside the LLM, and we should have some other thing that acts as an overlay to check the veracity of the LLM. I don't know. Jeff, what do you think?
Oh, I'd like to think Tracy's going to think along the same lines. A developer shouldn't be validating data, and I think that's what you just said. You have to be able to have some separation.
Whoever owns the data needs to validate the data, not the person writing the code. There should be an expectation of what the code produces, but they can't truly validate it because they don't own that process. And to have different assignments for different processes, I think makes sense, and I'm just going to say it again, I think that lends itself towards distributed intelligence.
Mm-hmm. And at the risk of grossly oversimplifying some things, I feel like what happened here is we created a bunch of copilots and people were like, "I don't want to become a prompt engineer," and so there wasn't going to be this whole massive, immediate return on value for all the people who built those AI things. So we came up with these things called AI agents.
But as I look at them, I'm like, well, at the end of the day, as an AI agent, kind of feels to me like it's a bot. It's doing something in particular. But we dressed it up now as an AI agent in its new agentic era of doing some things, but I don't know.
Jack, is this really net new, or is this kind of like at the end of the day, we're slapped a little lipstick on something? I feel like I'm a historian here. You keep bringing up things that we've done in the past.
But like you say, everything old is new again, right? We've been here before in the computer world. We've been here before in the manual world.
There was a time when you went to a bank to go get a loan to buy a house. You get a mortgage, and you'd have 10 people who would process the loan documents- Mm-hmm ... and it would take a month because it was hand-typed, and there was not enough forms, and everything got double and triple-checked and carbons and everything like that.
When the computer industry automated that process, they automated the human process rather than redoing it for computers. And so it still took a month or two months to process that loan application because it would still have to go passed around to the 10 people who had to do a sign-off. But the computer does the math correctly.
Nobody had to check the math on the mortgage app. So eventually, we figured out how to build a system that didn't require all the checks and balances. And I think we forget that because it's moving so quickly.
This thing with LLMs is only a couple years old. It's not 20 years old now, right? We are in the very early infancy.
Tracy said we're operating like they're eight-year-olds. Well, they're operating like they're two-year-olds that are thinking like eight-year-olds, right? Not 20-year-olds thinking like they're eight-year-olds.
We're really early in this, so we got a long way to go, and to expect that we've solved every single problem today or tomorrow, I think is not realistic. Yeah, it's just healthy. I don't think of these reports as alarmist.
I think of them as a constructive criticism. We're trying to find flaws in the system and addressing them as you move forward. And I think that's healthy.
That's a good thing. And we're just quickly learning that scaling AI safely is less about making the models larger, which is what we've been focused on, and starting to become more about building these reliable systems around them, which have been part of our conversation that we've had in the last couple of weeks. We've had many topics around that exact problem, and the companies who are stepping in to build more reliable systems around AI, and that's when we're going to mature.
Mm-hmm. Yeah, and to that point, and John, I'll ask you this, but do we just need to collectively reset expectations? I don't think anybody is saying AI should be tossed.
I think what we are saying to folks is that, hey, there are some things that aren't quite fully baked yet, and there are things that we're going to work through over the next year or so. But I don't think CEOs should be standing around going, "I'm afraid we're going to miss out because XY"- Yeah. I kind of see this evolution in the way some of these presentations are going during these conferences.
Before, it used to be pie in the sky. It was a utopian vision, and now they seem to have kind of scaled back a little bit or at least added some sort of nuance or some sort of proceed with caution type of messaging. And in a sense, because maybe they're seeing it, and they're trying to be honest with the crowd, butI do sense kind of a more realistic vision or more realistic messaging because before it was just ridiculous.
It was so over the top, and I think few of us believed what they were saying. And I think now given some of this research, some of the studies, and some of the real-life examples that they're contending with, they have to acknowledge what's out there. Mm-hmm.
So ultimately, as you look at all of this, and I'll toss this one to Jeff, but what's your best advice to folks about how to look at all these studies and how to think about agentic AI at the moment? Because I think I feel like everybody's kind of backing off the Kool-Aid a little bit and trying to figure out, well, what's real here? And what you just said, I think makes perfect sense.
Just to use the Kool-Aid example, you shouldn't crash through the wall and say, "Here I come. " You need to have some caution. How many big, significant business decisions-- Actually, I'm thinking about this, and I realize it may not be true, but how many big business decisions are really made on a whim?
Unfortunately, more than we'd like to think. But this can't be one of them because it drives everything that you're doing, your operations, your data, and everything associated with it. " I think we're all comfortable making that statement.
So it's not quite ready for full autonomy yet, but it really does make sense to say, "Let's find where it makes sense. " And we guide along. We help it grow.
All right. And don't forget to suspend your disbelief because once, if that happens then you're going to find yourself trying to do something that'll probably end up wrong, but we'll see. Anyway.
Probably will. Yeah. I do want to shift a gear here, but the next topic is somewhat similar in the same levels of concern.
So there's a movement afoot among the application development community, at least, to create something that they're calling specs, specifications. Other people probably have other terms for it, but GitHub is describing this as kind of a new constitution that will be set up that defines the parameters of what an AI coding agent can do and what its relationship is going to be to the rest of the environment. And you hear AWS picking up the same verbiage now around their Kuro coding tool, and it's one of the things that they're saying differentiates it because it provides not just the guardrails for the application development and the agents, but it also offloads certain things in a way that makes it reusable at scale and maybe helps us achieve that goal of building more software faster, but safely.
But Tracy, I don't know, is this net new in your mind, or is this just best practices are being coded in a slightly different way? Oh, it's kind of both, right? It really is both.
So first of all, I just want to point to the conversation we just had where we talked about AI being pretty decent at coding, but maybe not so good at natural language and kind of structured workflows, right? Those we see that they're degrading more quickly. So now we get to GitHub's new spec kit.
I feel like it's a response to the kind of the chaos of vibe coding. When I was at VolCon a couple of weeks ago, everybody was talking about vibe coding, where developers give AI agents these broad prompts and hope it generates the code that aligns with the business intent. So the core behind this spectrum and development is specifications, right, not code, become the source of truth.
So instead of generating code first, and as a developer, I know this is a mistake we all make. I can't tell you, we always run and start coding as soon as somebody comes up with an idea. You do the documentation first.
We've been told this for years. You define the architecture, you look at the task, and look at your constraints all up front before you start coding. So I feel like this is an attempt to pull back and put these standard best practices back into the coding process and put some guardrails around this vibe coding that we're talking about.
So I feel like there's a bigger industry trend here is that AI coding tools were kind of evolving into kind of orchestrated software engineering systems, and AWS is kind of following the same kind of process, extending Kuro with its spec analysis and its planning and its autonomous agents. And all of it's aimed at improving the quality of code and delivering excellent software. And the best practices are things that developers have been told for years and years and years, and this is just, I feel like it's putting it into maybe a natural language, loosely structured workflow that may not do so well.
So we'll see, because we're kind of pulling those concepts together. I think it's going to be done in a way where it will be much more difficult for them to ignore, at the very least. So, we should have some hope that they will follow the rules because we will be more gently guiding them, perhaps.
Sure. And one of those rules, Jeff, is it possible in your mind that we could get more secure applications out of this because essentially the best practices for application security become a spec that's built into the workflow, and people won't just continue to do things that we see every time, time and time again because that OWASP list of vulnerabilities still doesn't change. Yeah.
That's on us, not the code. But I do think there's a way to get good security out of it, and it goes back to the principles Tracy talked about. For two years, a couple of years ago, I did a tour speaking at different conferences.
They really talked about bad security comes from crappy code, and I still believe that. Bad code creates an insecure environment, and bad code is most likely to happen when you don't really know what architecture and specifications to which you're supposed to be coding. If you're just coding to accomplish something, you're going to miss a whole bunch.
So when's the last time you heard about having a good enterprise architect involved with an AI agent? I haven't met one yet. And because of that, there's that gap in how do we work, what are our principles, what's our environment like, what are the specifications we use for CPU usage and data size and everything else around that that you can use in your specs to start coding.
And I think the organizations now are recognizing that we're missing that almost entirely, and we need to start kind of at least nudging, if not really shoving people into it. Mm-hmm. Jack, you've been around security for a long time.
There's never been a lot of love lost between the application development teams and the cybersecurity folks. Do you think things might get better if we all embrace specs and people start doing the right thing, or is this just a pipe dream? Oh, no, no, no.
This is absolute 100% reality, and when Tracy used the term, "We've been talking about this for years and years," she wasn't joking. When I was at the Software Engineering Institute in the late 1980s, which is quite a while ago, this was what we were talking about, is how do you design software before you ever put hands on the keyboard, pen to paper. It's designing the system, architecting it, figuring out what the specs are.
What are your inputs, what are acceptable inputs, what are acceptable outputs? What are unacceptable inputs? And when I was doing networking stuff, the rule was always accept all inputs, even if they're bad, so don't crash on any bad input, but cause only good output.
And that was sort of a basic security principle is no matter what somebody feeds you, it shouldn't cause you a problem. You should be able to deal with that. And that was all based on specifications, not code, and I think it will go a long way towards improving our security.
The other thing that will improve security, however, is if we compensate developers on considering security as part of the design of the product. Until we do that, we will still have all of the issues we have today. Well, we don't compensate them on that because we're busily compensating them on how fast they can actually generate the code, but that's a core issue.
Sorry, it's a different argument, but still. But- And it's part of the culture, too. " I sat down, and I wanted to create a game.
That was my first program. So many developers, really the strong developers, they didn't learn in school, and even universities don't go through. If you look at a university course, it doesn't teach developers to start coding based on an architect document or a requirements document even.
Maybe the requirements are really high level, like a prompt, right? Mm-hmm. The professor gives them a prompt, and that's why this vibe coding has become so popular, because it allows developers to do what they like to do faster.
Mm-hmm. So there's a cultural piece in this, but if something like the Spec Kit encourages developers to include things like documentation and requirements and the architecture to start generating code, then we begin making it easy for developers to pursue those best practices. Mm-hmm.
Best practices are a pain in the a*s, to be quite honest. I don't like following them in anything we do. It feels like it slows everything down, but we try as best as we can to follow some standard best practices.
And this conversation has been going on for such a long time, I think it's great that- Mm-hmm ... both GitHub and AWS is thinking about how to incorporate this into generating code, because that's what we've been asking developers to do for 30 years. I think- Can I ask a quick question?
And again, so Spec Kit tools, the agreement they're going to become standard, right? So is there any concern about what impact that has on junior developers and their ability to learn how to code versus kind of merely approving what the AI has planned for them to do? Is it- I would hope that junior developers would start learning how to write requirements, right?
Mm-hmm. Yes. And be involved in the architecture discussion before they're asked to do any coding.
But junior developers are the ones who have been doing the vibe coding. They're the ones who have been tinkering with computers since probably their teen years, and they just want to jump on and start coding. So junior developers should be asked to start working with all of the different teams that you have to write requirements for, from the end user all the way to operations, and help build that architecture document so that they understand from the beginning the process.
When I was doing development work, and this is '70s, '80s, I couldn't even begin by my first line of code without having a full-scale set of requirements that was approved. So yeah, Tracy, we are skipping that. Something else Jack mentioned, which resonates with me very strongly.
And- In ... and yeah, and we ended up paying them for writing crappy code. This is where the whole crappy code came from.
" It took six months, but boy, it turned right around. To Jack's point, when you compensate people for doing it right, eventually they'll start doing it right. Well, here's my- We've gone through some phases, though.
Remember rapid application development? Mm-hmm. I was coding during that time, and literally, I'd be sitting there on a trader's workstation, and they'd be saying, "I want you to move this over here and this over here, turn this a different color, make this really, really dark green," and that would be my requirements document.
So they were using me as a prompt, and they were doing coding through me. All right. Well, theoretically, that requirements document is going to get automated and dynamically adjusted.
But here's my pet theory, and I tested this out earlier this week, and I didn't get too far with it, but I'm sticking with it. I'm going to go with a Winston Churchill phrase on this, but you shouldn't let a good crisis go to waste. And if we're going to have this vulnerability issue where AI is discovering all these vulnerabilities, and there's going to be all this stuff coming back in through the SDLC to get fixed, well, now's a good time to rethink how that software's being built in the first place to maybe not have these issues.
And maybe we'll build stuff that's more secure the first time around so that we're not being held hostage by a bunch of LLMs that are pointing out the flaws in our code bases that we all know are far too prevalent to begin with. I think we'll just have fatter stacks with more vulnerabilities, to be quite honest. Okay.
Well, my hope spring is eternal, but we'll see. So much for that faith Tracy had. I know.
I just think AI's going to pull in a lot more packages. It's going to be long-running, as we talked about in the first block. All right.
Well, it's Friday, and maybe we'll all be more hopeful next week, but right now we're just not buying into it. We'll see. But let's shift the gear here up because, well, we had another one of these topics that is also maybe in the eye of the beholder, but Jack has a story up on Security Boulevard.
Basically, it's a Andy Rooney-esque type rant, and one thing I would say about it is it points out a fallacy that says if we have something where there are rules, but there are more exceptions to the rules, then the thing is fundamentally broken. And so Jack, cybersecurity is fundamentally broken because, well, there's an exception for everything. Well, yes, but let's just say start out with this is fun because I don't have to explain the complex technicalities of some obscure bug or vulnerability.
Replica Cyber came out with a study. They surveyed 200 companies, and they found that 100% of the people they talked to made exceptions to their security policy in order to get things done. Now, sort of seems obvious, but that it's 100% is a little bit concerning.
" That's the good news. " And that's really concerning. And the other part of that is that leads to what they're calling both the exception economy, which is this is all standard practice now, and there's a visibility gap.
So when they talked to the C-suite leaders of the company, a quarter to a half of the leaders felt that even though they had exceptions, that when they wanted to go do another project, their company was ready to implement that project securely, whether it was an M&A activity or something else. And a lot of this is like how do you do clean rooms for M&A activity or those types of high-level strategic projects. So a quarter of the executives thought they were ready, but only 5% of the CIOs believed that they could satisfy security for these projects.
So there's this big gap between what the CIO C-suites think the company can do and what the executive leadership thinks the company can do as far as implementing security and building secure projects. I would posit there's another gap. In my mind, it goes something like this.
The more senior the person is, the more exceptions they seem to have. So you know- Mm-hmm ... that sales has more exceptions than the rank and file.
" And everybody else looks at what they do, and then they do that themselves. So Jeff, does this kind of create the cultural problem where we don't take the security rules seriously because everybody knows there's an out? Absolutely.
And Jack, I appreciate your column on this. I will pose that if you wrote this and posted it 35 years ago, it would've been somewhere else, not online. Everything would've still been holding true.
I've been in security for close to 50 years, and I learned early on that becoming the department of no means that my successor will negotiate, and I'll have to go find something else to do. So I learned pretty early on everything has to be a negotiation, but what you have to do, just like the spec hit, is come up with how do weum, approve and handle exceptions. Because an exception can't be, "Oh, we need it now, just do it.
Oh, it's an emergency. " How many people have I talked to that have had direct disastrous M&As? There was one I was involved in where I was in an acquiring company, and the CEO said to me specifically, "I want a firewall between everything.
" Well- Right. Well, that disconnect between executives' perception and this, I guess, operational reality has been around for a very long time. Even when we started talking about separation of duties, and we talked about separation of duties often in the concept of, or in the context of deployments.
And we had Jenkins running, right? And Jenkins did everything. " But Jenkins doesn't really do anything.
It's just an orchestration engine. And developers were writing scripts, and developers were controlling the build, and developers were controlling the product releases. There was no separation of duties, even though the executive's perception was that.
And so I feel like this is just a continuation of that discussion that we've been having. And the takeaway is that security can't operate purely as a control function because it can become a blocker. And the velocity of business is outpacing security architecture, let's just put it that way.
Mm-hmm. Especially around AI. Mm-hmm.
So, this is going to continue, and I feel like right now, it's for security in particular, they should be looking at these exceptions and understanding why they have to have them. Um, I had- And here I thought we were going to get through this entire conversation without mentioning AI. But AI is one of the major implications for this problem, right?
Yes. Mm-hmm. Absolutely.
They're switching forward with the- Everything's driven by AI ... and AI's becoming a big exception. Right.
So, this is an ongoing battle between what executives really understand's happening and what's really happening. Here's the report I want to see going forward, and John, you and I have been subjected to any number, one, of these security reports, and they're all kind of very similar. But the one I've never seen is the one that says how many of these breaches were the result of an exception, that policy that somebody had actually put in place and that resulted in this breach occurring because, well, we had a policy, but we had an exception.
And I swear, every time these people do any kind of report here, they all make it feel like the company was the victim, but I don't know. Right. Blame the technology, don't blame the executive leadership.
Yeah. If I'm walking through Times Square in the middle of the night, and my wallet's in my back pocket, and I get pick-pocketed, yes, I am a victim, but I'm also kind of stupid. Mm-hmm.
" Mm-hmm. What do you think, Jeff? Same thing, right?
Oh, yeah, and when you do a root cause analysis, you can usually predict the answer. Yeah. And here's a weird thing, the thing about the...
Sorry, Tracy. The thing that always, my takeaway of the exception economy, and I love that phrase, it just creates this dangerous normalization of deviance. That's my takeaway from this.
But again, this is not new. This has been the norm, and as Mike pointed out, these studies, they do tend to point the finger at an inanimate object rather than human error or the people in charge, and that's going to continue. I think Jeff said something about that, about job longevity as part of the issue there.
Yeah. Well, if you have- If you blame your boss on the issue, you're not going to be employed very long. Yep.
If you're having to deal with exceptions, and they're happening often, and in this article, 100% of companies are doing exceptions. It just means that we haven't built security directly into the workflows. Because if they were built in, they wouldn't have to get a policy exception because the policy would be built into the workflow.
And we are not seeing money put into that right now. We, in fact, in terms of maybe in the platform engineering space, they have many more discussions on secure by design. But if we look at the DevOps space, we have struggled to get to DevSecOps, and we don't want new plug-ins.
We don't want to touch the workflows. So then everything becomes a policy. We have so much work to do just in DevSecOps, and that's not even bringing in building AI applications that have brand-new implications on security.
So, to your point, and then let's stitch this all together, do we need to have a bigger conversation where we're talking to application developers about how to build code more securely that we're going to use to reinvent all these workflows that are going to be AI-driven? That's the B and the A block stitched together. Mm-hmm.
Then we're hopefully going to be more secure now because, well, in the C block, we're going to actually not have a million exceptions built into these workflows. Is that achievable? Or is this just the nature of the condition, and we're always going to have the things, and we'll never change?
Jeff? Oh, if I had an answer to that, I might be talking to Elon Musk instead of Mike Vizard. I don't know.
I don't know. That's the sad truth of that. Does that just adds up or down?
You would have a logical discussion with Mike. This is- Mm-hmm ... a resolution there.
Sorry. Right. All right.
I do believe that we have to address this from the, we have to go back to our CI/CD workflows, and we have to start evolving them because we can build policies into those, and we're just not doing it. And partly because the CI/CD platforms are way long in the tooth, and I love the idea of MCP servers. I know that at CDCon next week, you're going to hear about CD events to try to solve this problem.
But secure by design is the only way to get out of this exception economy, and it's not a good economy to be in. And we have the capability now to do it. " And that was really the only control we had.
Now, we can automate that. We can automate, we can build in those policies, as Tracy said, and automate it. And so we're going to be writing secure code by design, and until we bite that bullet and pay that cost, we're accruing interest.
This is tech debt that we're not going to be able to pay off if we don't do it. All right. I'm going to leave it here, folks, because, well, I think it's pretty clear to everybody who's watching this just who the enemy really is these days, and we'll figure it out from there.
Hey, everybody on the panel, thanks for sharing your insights and your thoughts. As always, have a great happy Friday. Stay tuned for the rest of the Techstrong TV replay right behind us.
There's some awesome conversations, just as equally brilliant as they are here, and we will talk to you all on Monday.