Enhancing Resilience and Efficiency with Robin Yeman | DevOps Experience 2024
Robin Yeman shares insights from Her experience in industrial DevOps, focusing on the challenges of implementing Agile and DevOps in safety-critical systems. Robin highlights issues like long lead times and regulatory requirements, and propose solutions such as organizing teams for value delivery, modular designs, and iterative planning. Emphasizing early integration and a growth mindset, she encourages learning from failures and provide resources for further exploration.
Transcript
All right. Hello. Um, my name is Robin Yeman, and I am going to be talking about enhancing resiliency and efficiency building systems in highly regulated environments.
Uh, so, so this is me. I, uh, spent over, uh, 28 years in this field, and 26 of those years I spent at Lockheed Martin as a senior fellow, but just recently published a book on industrial DevOps, which is the application of agile and DevOps, again, to large scale safety critical cyber physical systems. Um, and so that's what I'm gonna walk you guys through today.
Thank you. So let's begin with the challenges. And there is a lot, right?
So, um, when we were looking at, you know, applying Agile and DevOps, um, you had a number within software, but when we get to hardware, when we get to things that are, are cyber, physical and safety critical, and a couple others, um, one of the, the big ones that are long lead times, right? Takes a long time to procure hardware and, and get it in place. There's some pretty expensive test requirements.
We have to deal with multiple dependencies, right? So if I'm building a missile, I have a variety of dependencies. Um, I have things like complex risk management, and then a very large attack surface.
Once we, we connected everything together, the attack surface went up exponentially. We have safety and reliability concerns, and then regulatory requirements. I'm gonna walk you through some of those.
Here. You can see, for example, if upon manufacturing a satellite, I have, uh, procurement, I have multiple components and sub-assemblies that I'm gonna get from a variety of suppliers. And this takes time.
In this particular case, thermal vacuum changer chambers cost hundreds of millions of dollars, right? And the test can take months long, right? So the whole rapid feedback is, is, is much more difficult in this domain.
I have multiple dependencies. So can you imagine all the dependencies that a satellite constellation has? There are hundreds, right?
So multiple dependencies, and we need to link those together. I have complex risk management. So when dealing with GPS, for example, things that can happen to me are solar interference, uh, frequency crawling, signal degradation, jamming, spoofing, and, and that's, you know, some of those are environmental as well as, um, IT related, right?
So there's, there's a variety of different options there that, that I can deal with. And extensive pack surface, you can see here, um, all of the places that a satellite could, could be vulnerable, right? So everything across the supply chain from my onboard computer, and then even things like, you know, uh, the ground system and things that are happening outside, we're gonna talk about high stakes, reliability and safety.
For example, here, you can see that NASA has a variety of standards under 8,000 that, that space vehicles have to comply with. Um, and one of the problems that you're gonna see is that a lot of safety standards and things are occurring in areas where typically you'd see a phase gate. So what does that look like from an agile perspective, regulatory and compliance hurdles.
So things like, you know, API for automotive or ISO 27,001, right? For information security management, there's a lot of compliance that we actually have to build into our pipeline. Now, is everything lost?
Hopefully not. All right? Um, again, most of my experience has been building large scale safety critical cyber physical systems.
And myself and Dr. Suzette Johnson put together a set of patterns that we believe can actually support and overcome these challenges, right? The goal is that I need to be able to get the same benefits that I did in software, such as adapting and rapid delivery and rapid feedback within cyber physical systems.
Um, and so how do I do that, right? Because the market is still requiring rapid delivery. So we used a variety of bodies and knowledge to, to come up with this work.
Um, everything from Agile and lean and DevOps, but also things like systems thinking, design thinking, um, and, you know, uh, different, uh, highly regulated, you know, things like FICA analysis, um, and formal methods. So the first pattern that we came up with is organizing around the flow of value. And I know that many of you probably believe that, hey, I, I organized for value, but is it organized for value in order to deliver a system?
So here you can see, say I'm building a rocket, um, I've got program management, systems engineering, design, hardware engineering, software tests and operations. And in many large companies, these are individual areas of excellence. They're, they're silos.
So once they complete their piece, they're handing it off. Um, and, and that requires a lot of extra time. It also impacts the quality.
So program management, they're, they're typically looking at cost and schedule, and they might be looking at lean management systems. Engineering will talk a lot about systems thinking, system design. They're talking about design thinking in the hardware world, we're talking about rapid prototyping.
In software engineering, you're gonna hear them talk a lot about agile. In te you hear things like shift left. And in operations we will refer to ITIL or IT infrastructure library.
And here's the thing. Each one of these processes in all of these environments are actually looking to deliver at the speed of relevance, right? They're trying to deliver, but inherently because we have handoffs and a different language, we can't even talk together, right?
Organizing around the flow of value is about organizing across this chain, not within the silos. The next pattern we use is multiple horizons of planning, right? So typically agile, you've got people, um, maybe planning in two weeks sprints, and that's okay, uh, except that it doesn't actually, you know, uh, work well when I have to integrate over multiple years to build things, right?
So, um, so in waterfall, you're seeing integrated master schedules that are multiple years. That doesn't really work well either, because a lot of that's, um, made up or assumptions, and the moment anything changes, which it always does, um, it is no longer valid. So we said, Hey, we have to connect those together.
And really we want the, the, uh, multi-year plan that breaks down into the annual plan. It breaks down in our quarterly plan, and then breaks down into a sprint plan. Maybe every two weeks breaks down into a, uh, a daily plan, right?
And each of those horizons of planning, I get empirical data from. And one of the things I'm gonna do with that empirical data is I'm gonna use that to inform the next horizon of planning. So here, for example, you can see Artemis, right?
So, you know, Artemis Untr launched in 2022, and we were supposed to see Artemis two crude test play actually launch in 2024. But based upon the horizons, based upon empirical data, actually had to move that right to 2025. So, um, they're using empirical data to further inform the next plan.
And when you're building large scale systems, they do take multiple years. So second, second principle, bring these things together, multiple horizons of planning. The next pattern we found is implementing data-driven decisions, right?
And currently, uh, the technology has improved so much that, that with cyber physical systems, there's a number of ways we can get that empirical data regarding what we're building. So things like simulators, emulators, digital shadows, digital twins, which is basically the intersection of, you know, the model that's, uh, has full telemetry and the working system with telemetry, and we're feeding that back and forth. And 3D printing.
So while these are not all new, the technology now is currently at the state where I can get real time data-driven decisions. And this can be seen, um, for example, um, a star. Uh, so Aari just was awarded a contract to actually digitally certify the X-Wing fighter.
That's amazing. And the amount of time reduction that's gonna take when building cyber physical systems is huge. The next principle we're gonna talk about is architecting for change in speed.
Um, again, not a completely new principle, but we wanna look at that modularity with standard interfaces. This allows us to make change later in development, right? This, this impacts that whole monolith and be able to change pieces in and out.
Now, this isn't typically how things were built in the past, right? So a lot of legacy has to undergo things like, you know, uh, refactoring to get to this modular state that's critical in order to optimize for change in the future. There's a new approach that, um, we've seen a lot of lately, which is software defined X.
So software defined X is we architect based on the software. And so you're looking at software defined satellites, software defined weapons, software defined vehicles, um, and it has the maximum amount of flexibility. The next principle we're gonna talk about is iterating and managing cues.
Right here you can see the software acquisition pathway, so DOD 5,087 and DOD 5,087 put together to allow much smaller batches. Now, while it began as the software acquisition pathway, people are leveraging it for larger weapon systems as well. And you can see here that they're iteratively planning and giving real capability to the stakeholders, which then reduces the amount of rework, right?
In the past, we'd go off, we're gonna build, let's say, you know, the F 16, we're gonna build the whole thing, and then we're gonna say, is this what you wanted? And if it wasn't, it's pretty much too late to make any kind of significant changes here. We're going to actually have that continuous bean back, which is going to allow us to have much smaller batch sizes, and it's gonna allow us to adapt to change, which right now things are changing every day, the speed of light, right?
So it's, it's, it's really important to be able to adjust and adapt. The next pattern we're gonna talk about is applying cadence and synchronization. So you can see here I have a series of MVPs and minimum viable product and MVPs, which is the next viable product.
And I am putting all of my teams that are building this rocket on the same cadence with tight integration points. So you can see the avionics team, for example, they're integrating every single, um, every single one of these. But the environmental control team, well, they, they're, they are not on the exact same, you know, time box, but they have those intersex right?
Now, it's critical that each of your teams when building anything, had this common cadence and synchronization so that the system will actually come together, right? If, if everybody's on their own timeline, it extends it a lot more. The other thing is, if you've ever read principles of product development flow, um, Don Renson goes on to explain that in manufacturing, the number one priority is to reduce all variability in the system, everything.
But in product development, we want to reduce bad variability, and we want to exploit good variability in the way of innovation. That's really difficult. So we wanna make sure we remove as much noise from the system.
And that's another benefit that cadence and synchronization brings to the table, is it allows me to see those areas and make a better determination of bad variation and, and good variation, right? Because it's, it's a little different than manufacturing. Next, we want to integrate early and often.
And here you can see a missile program, right? And they're integrating every single sprint. Many times people will tell you, Hey, I need much longer sprints for hardware.
I'm not saying that that's, you know, not true in, in every single case, but in many cases, um, they can actually be on the exact same sprint as as software. Um, the, the difference is, I'm assuming that I'm not 100% done. I'm not deploying something, but I am learning something.
So here you can see, I, they've, they've got their breadboard set up, they've got their net chassis that's linked to it, and they can validate at this stage in the program, this integrates, it works. We can send messages through. It's not deployable, but I can learn when I wait to integrate in every case, I keep all of my risk till later.
All right? Um, you know, my very first program when I was, um, you know, just starting out in engineering was working on a submarine. And I can tell you that it took, you know, multiple years to build the submarine.
Cool. Um, but then we did integration and tests afterwards, and it took even more years to integrate it and test it as opposed to integrating as we went, right? And buying down that risk.
It's a different approach, but it's, it's required in order to reduce our timelines and be able to adapt to change. If we believe that integration at the till the end and things magically come together, they almost never do. Next principle we're gonna talk about is shifting left, right?
And I wanna begin with test. Um, people are aware that we need to test and validate. In many cases, they'll tell you, Hey, I need somebody completely different to test.
And the reason that is, is because whoever built the product is gonna know how it works. So they're not gonna be the best tester. 'cause inherently they're gonna use it the way it was meant to be used.
But if you begin with test, you kind of remove that problem. The other thing is, you build exactly what's needed, not extra, not less, exactly what meets the test case. So while many people will tell you it takes longer to do test driven development, it just feels like longer in their state, but actually it's much shorter to get to the customer.
Uh, here you can see McLaren, and they do a lot in the way of digital engineering and shifting lots in every single case. They basically make an adjustment to their system, and they make an adjustment and, and validate in their, their digital environment. Um, I was told, and I thought this was really interesting, that if you use the exact same car that you started the race with in Formula One, and you were the fastest car that you, and you didn't make any changes during the race, you would no longer be the fastest car, right?
So they're constantly having a hypothesis testing in that digital environment and pushing to the system even during the race. So I, I think it's just fantastic. The last principle I'm gonna discuss is applying a growth mindset, right?
And, and in our culture, especially in safety critical over time, right? We were like, we can't fail. Well, we need to rethink that.
We can't fail. We can't fail in a place where I have human impact, right? I definitely don't wanna fail there, but all of the other steps and prototyping and evaluation along the way, that's gonna allow me to push past my boundaries and learn, right?
I'm gonna learn the fastest. So here you can see SpaceX, um, they, they put out a, a tweet that said, Hey, you know what? Look at how successful we were.
We've got all the data we need. Um, but right from a traditional, uh, you know, contractor, I can tell you that, um, many traditional folks would've not considered this blowing up, uh, rapid, unscheduled disassembly like they call or rud. Um, they wouldn't consider that a, a success.
But basically SpaceX is trying to learn, and we can see over the last decade that they have reduced the cost launch so much that most people are leveraging their vehicles, right, to send satellites into space, um, or to bring astronauts back from space because they've been so successful. And the reason they've been so successful is because they're constantly learning, right? So, and the best way to learn, I hate to tell you, it's, it's, it's when we fail.
So myself and, uh, Dr. Suzette Johnson and others have been on a journey since 2018, and we have written multiple papers, uh, through it Revolution and, and Jenkins, uh, publishing group on industrial DevOps. The latest one we've done is on digital twins.
And up next will be coming out AI enabled digital twins. If you're interested here, you can get a free first chapter of industrial DevOps. So determine if it's interesting to you, if it makes sense to you.
Um, and if you reach out to me on LinkedIn, I can probably even get you a free digital copy of the entire book 'cause I still have some. Um, but you've gotta reach out and you've gotta gimme feedback. And with that, I am completing my presentation on how to apply Agile and DevOps to large scale safety critical cyber physical systems.
I.