What Would You Send a Cloud Scout to Fix with SOUTHWORKS
This segment grounds the idea in practice. We’ll examine how embedded engineers have helped product teams go beyond reactive fixes — from automating post-mortems to co-designing self-healing infrastructure and predictive testing frameworks. The focus is on what changes when teams own reliability together: faster iteration, fewer handoffs, and more precise success metrics. We’ll close with an open discussion on how organizations can experiment with the Cloud Scout model — and what it signals for the next evolution of DevOps.
The presentation addresses the challenge of organizations needing to adopt new technologies, such as AI, but facing uncertainty and risk. The Cloud Scout model is presented as a way to mitigate these risks by embedding engineers to assess the current state, identify opportunities, and demonstrate the value of new tools and practices. The goal is to de-risk innovation and empower teams to embrace change, particularly concerning AI adoption, which is driven by business mandates but often faces resistance due to security concerns or a lack of clear implementation strategies.
A key aspect of the Cloud Scout approach is its focus on practical application and measurable business outcomes. The scouts aim to demonstrate, not just tell, how AI can be utilized to achieve specific goals, such as reducing alert fatigue or enhancing efficiency. While the initial engagement is typically a 40-hour-a-week commitment for three months to understand the problem and prototype a solution, it can evolve into a fractional engagement with a specialist or lead to a separate project for building out the solution. This approach emphasizes the importance of senior expertise in navigating uncertainty and mitigating risk associated with new technology adoption, ultimately enabling organizations to become more mature and effectively embrace innovation.
Presented by Johnny Halife, Chief Technology Officer, SOUTHWORKS. Recorded live at KubeCon North America in Atlanta, Georgia, on November 11th, 2025. Watch the entire presentation at https://techfieldday.com/appearance/southworks-presents-at-tech-field-day-at-kubecon-north-america-2025/ or visit https://techfieldday.com/event/kubecon25/ or https://southworks.com/ for more information.
Transcript
I saved some time at the end for some questions. If you have. Um, I've already touched on it a little bit.
My, my question or my answer to that would be, you know, analytics, observability, alert fatigue, that kind of thing. I think it'd be super useful for that kind of thing for administrators and help 'em focus on what really matters in their day-to-day operations and not the little things, and take that operational efficiency and, and send it over here and kind of free 'em up a little bit. Good.
The question on the, um, the pre-release predictive assessment of the product, like what guarantees do you put around that for the, the customer? In, in which regard? Uh, it seems there's such a big world of, and you're thinking about software, we don't have the same kind of qualifications and testing that, like standard engineering practices go through, like people who build bridges and analyze materials.
But you're saying you've got a test that can detect, you know, pre potential pre-release failures or risk modes. That would be, I think, I feel like very hard to detect. I wonder like, so how do you guarantee that you're gonna, or what, what percentage correctness are you looking at for these kinds of things?
So it's, it's a moving target because some of the data exist, Some we need to help create and we use some strong predictors like this coverage. What types of tests do they have? If there is like, um, the change rate of that file, the correlation with issues, the trace stack, four hour, you know, production issues, logs and those sort of things to figure out like the risk of touching.
Like I, I use a fancy name, but it would be what's the risk of touching this fire? That is kind of the scoring that we are working after. Um, of course that's information rather than action.
So at the end of the day, what we are after is this is a risk score, how can we help you reduce that risk? And there, there the, you know, testing comes into play, the regression testing and uh, the edge case scenarios for the things that we found. Even the times in which we have seen a bug or a production issue hitting that file is a tested, is it covered?
Like, to mimic those scenarios to reduce the likelihood of failing in production. But it's kind of a north like that that as that score assessment or that predictability, it's kind of a north star. Mm-hmm.
When it comes to like, what each are the things that you need to work on or why is it so high for this file? Yeah. And in more legacy, like traditional legacy settings, we have seen the do not touch this, you know, annotated commented thing, like, just don't touch this.
Well, that shouldn't be happening in 2025 because with all the tools and all the, uh, opportunities we have for improvement, but it's a good way to, you know, craft a score. Yeah. Based on all that data that we have that we are capable of reasoning on top and predicting, like, again, I call it prediction because we try to just close as the reality as possible, but still there is some, some failure.
What we try there is to educate the teams on what are the missing links when it comes to quality control. Yeah. You gave the example of a building a bridge, like where you have test, we type of test against what, like, are these production cases every time you have a bag, like the bag that we saw, is that a test case that passes effectively?
So, so do you integrate this into the CI pipeline and then have a threshold score that would block a release? It would be more on the, yeah, I'll call it the, the CD pipeline or their GitHub pull request type of thing as an additional controls that we want to stack. Because that's the other thing that, that we've learned that originally, you know, you will deploy a pre-production environment that you want out of the, you know, uh, just deal for that pool request.
So nothing weird happens. I need revision against production and all those sorts of things. But You will stack without stacking this and a missing link on that will fail the whole thing.
So we treat these as independent checks and each team will decide how confident they are. We'll give our advice, but we also understand that we cannot show up, lock the door and say like, until everything is a hundred percent covered, until you have tests for all your production issues, there is nothing that is going to go into production, uh, anymore. Like, that will be unrealistic.
And, and the idea of this is to help the organization become mature enough to drive change forward, uh, to lose to the risk change and to embrace innovation. But there is nothing a hundred percent guaranteed, and we don't want to be, you know, unrealistic about like, sure, there is a way to get to a hundred percent, but it will require you not to work on the product for the next 12 months or so. Yeah.
What, but you, um, you described the problem statement at the beginning, I guess more in terms of the function the Clouds Scout provides, which is to provide communication. You describe the problem statement as, um, the, I guess, you know, largely independent agendas and operations of these different groups, but I, it seems to me that the, the outcomes of this are really more business outcomes of speed, um, responsiveness, that sort of thing. But what would, what would you say they are?
I mean, it could be efficiency, it could be smoother operation. So, and you've said innovation a lot. Yes.
Because What, what, I mean, And I would take, I'll put it another way. What do you think is the right scenario for a a, you know, for a customer or potential customer if they're sitting there, say, if, if you were to say, are you, are you experiencing this as an organization, then you should look at this. Like, what scenarios?
So I I think that they, it, it's, it's a great framing, right? I would start by saying like, are you looking to, um, increase efficiency and to tap into all these tools, or they're risking the adoption of all the tools that are becoming a business mandate for you to adopt? Because that's something that we hear a lot, right?
Like I met with a lot of c CTOs and CIOs in the last year or so, when, you know, we go for, uh, we sit down, I pitch South works and we work together, they know me. Then we go for lunch, and it's like my CEO and the board is really pushing me to say like, how are we using AI and what is the, you know, the AI adjunct story for our organization when it comes to all these practices and we feel that we are falling behind? Well, then there is, again, the technology is almost there.
There is a pressure for the business to go into one direction. And if you want to derisk that and to do it De-risk, that was another thing that's, that came to mind is, is a lot of, um, The, A lot of management is much easier to, to, to do things And, and, you know, Easier to make decisions. When we were thinking about the name scout became the thing of, or the name of the thing, because for us it was a thing of like, Hey, we'll send somebody in to figure out the ground and to understand the risk or to articulate the risk.
It won't come back with a report. It will show, don't tell about the opportunities that you have. And that's why it's scout.
Because we understand that there is a lot of discovery that needs to happen in a high uncertainty context of like, Hey, I know I need to be using AI using all these practices, you know, em empowering my team to do more with less and driving the agentic wave and all these trends, facts, you know, fashion like that. There are different things, uh, that you can say about this, but at the end of the day, it's a reality. AI is here.
These tools work. If you put them to work, there is some associated risk to it as we were talking in terms of security exposure, like who build this and why, and what information do they have access to. So the scout is that practical way of going in that is the situation.
And show don't tell, say, Hey, you can use an agent to fix your bug and reduce the alert fatigue, or turn those alerts into actual fixes in no time. And here is a zero risk or the risk scenario in which I prove to you, let's figure out how we take this company wide, right? Like how we turn this into a practice that we can adopt into a company capability.
But it goes hand in hand because innovation has to do with doing thing different things or doing things different. It always comes with risk. And with ai, we see a lot of, you know, fans or fanatics.
I would say that no, sure, I will share everything I have with OpenAI and I will drop my bank statement into chat. GPT free, free version, uh, unregistered, right? Like for the world to see.
And then you have the people that are like, yeah, unless I can run the model on my own, I won't be using it. So it takes a little bit of both, right? Like to the risk those who are, you know, all in.
And to empower those that are afraid of the potential, um, of the potential implications and help them, their risk, the whole AI story. Do the scouts get called into customers? Does South works get called into the customers by the teams who feel they're not achieving enough and need, need more assistance?
Or do they come down from above as the managers start saying, well, you have been given this opportunity. You should be doing all of these things and you are not, here's some assistance. So that is a great question because scouts is not what people ask for.
No. It's the recipe that we found to increase our likelihood of succeeding with a customer in these type of projects. Why?
Because these requests usually come top down and it's peel me a project in which all my teams will be using AI and teach them, teach them how, and show me nice dashboards of adoption and change management. Simple, Simple stuff. Yeah.
And it's like, well, that's what you want. Not really what you need. What you need is somebody that, you know, goes, boots on the ground, figures out what's the real opportunity to bring this into the field, to remove the risk out of it, and to show them that it's possible.
Because on all these, and we impart part in different stages of this change management initiatives in which we will adopt the cloud, or we will adopt AI or we will adopt agile. It's a political battle. And when it comes to technology, you have funds in both corners, and there is no reason to believe until you can show instead of telling.
And nowadays, we believe that scouts is not what people will ask is what people need. Hmm. So are most of the scouts working on AI initiative?
AI is a tool. AI is at the, the adoption of AI or the streamlining of the processes or increasing the efficiency or achieving some goal through ai? Yes.
Not on AI per se. We see AI as means to an end. It has to be tied to a KPI, whether it's like reducing the time it takes or the number of s or increasing the efficiency or something like that.
It has to be tied to a business goal. It's super fun on the technical side, it has a lot of intricacies on the technical side, but it has to prove business value, and it's a means to an end. Are, are your engagements fractional or do you literally plant, uh, a re two resources in 24?
Or in other words, it sounded like you had some fractional metrics up there, like weekly, daily standups, but are the engagements you, you listed two? The, the, the engagements are 40 hours a week, three months? Yes.
Okay, so it's not fractional. No. Okay.
No. We, what what we have found, and, and if I go back into the, the, the, the, how I say the fixture of our services is that once we understand the problem and we go through the three months and we figure out, and we have a prototype, it can become a fractional thing through a specialist because we understand the problem. And it could be, Oh, it is fractional.
Yeah. Okay. It could be 10 hours a week for something really specific on a, on a place we call that specialist.
The specialist is a traditional consultant with the SME knowledge and can look at the problem. But we really, But that's the selling point though, right? The selling point is to my earliest question, if I go into a bank, I need a banking expert.
Yes. That can develop Java code and understand now AI and understand SRE and troubleshooting. Yes.
That's gotta be a specialist. Yes. It's a combination of both the process for the scout, Just like, this is my world, so Yeah.
Yeah. I know how hard it is to hire people with those skills. Almost impossible.
So It, it's hard. It has to be fractional. There's, yes, there's no way you could scale.
So It has to, I, I would say it's not meant for the long run to be running 40 hours a week, uh, for five years. It's, it doesn't make sense for us. It doesn't make sense for the customer.
We understand there is a period at the beginning of the three month figuring out what's the actual problem I install and all that. And then it can become fractional. It can become a project on its own.
Like, okay, I have an SME that figured out and created a roadmap. We prove the value, we prove the point with a risk, the whole thing. Now, there are things that we need to go build that could be either through a fire team or a dev crew, but it doesn't make sense to have something, somebody that senior with all that valuable knowledge on a technical, on a technical problem forever, right?
It's more about the uncertainty and making sense out of things that you need at the beginning, and then it needs to move on. Because once you understand, then, then it's our bread and butter, then it's technology. That's what we leave and breathe.
But it doesn't make sense to stick forever. As you know, the scout that lives with us, but happens to be an expert on media and entertainment, TPGs cloud and Java code. It doesn't make sense for the customer.
It doesn't make sense for us.