Cloud Scouts are Embedded (Human) Builders for the Cloud Native Frontier with SOUTHWORKS
Most SRE teams were designed to ensure uptime, not to evolve products. They monitor systems but often lack the context to influence architecture or design. This segment examines why proximity — being part of the product team — is what transforms reliability into progress. When engineers operate as embedded partners, they surface deeper insights, close the gap between observation and action, and help the team own the outcome end to end.
Johnny Halife from SOUTHWORKS introduced Cloud Scouts, a new service designed to bridge the gap between SRE teams and product development teams. He explained how traditional SRE practices, while valuable for maintaining uptime and monitoring applications, often operate in silos, leading to “ticket battles” and a lack of context when issues arise. This SOUTHWORKS service addresses the evolving need for engineers who understand both the application and the platform, and can work directly with development teams to identify the root cause of problems, not just surface-level errors.
Cloud Scouts are senior software engineers who are embedded within customer teams, possessing domain expertise and the ability to quickly prototype solutions. They actively engage with software engineers, platform engineers, and architects to foster better communication and collaboration. These scouts also use AI-powered “companions” to analyze telemetry data, identify patterns, and propose fixes, while always maintaining human oversight to ensure accuracy and alignment with business goals. It is intended to provide hands-on support, rather than a consulting position.
The goal of Cloud Scouts is not to replace existing SRE or development teams, but to enhance their effectiveness by facilitating knowledge sharing, promoting end-to-end ownership, and accelerating the adoption of new technologies, such as AI. The engagement begins with a three-month assessment to evaluate the current state and establish a baseline, to achieve measurable improvements in areas such as alert fatigue, time to resolution, and overall system reliability. SOUTHWORKS emphasizes transparency and a collaborative approach, empowering customers to mature their practices and become more self-sufficient over time.
Presented by Johnny Halife, Chief Technology Officer, SOUTHWORKS. Recorded live at KubeCon North America in Atlanta, Georgia, on November 11th, 2025. Watch the entire presentation at https://techfieldday.com/appearance/southworks-presents-at-tech-field-day-at-kubecon-north-america-2025/ or visit https://techfieldday.com/event/kubecon25/ or https://southworks.com/ for more information.
Transcript
My goal today is to walk you through this concept of Cloud Scouts, uh, a new service from South works that we've been working on over the last year or so. And the idea is to formally introduce it today to you for the first time, right? So I will start with some, some history.
Like in the beginning there was nothing and so on, but you know, uh, when we started writing code in the early days, uh, the bagging and logging and all those practices fall into the developer, right? Like as a developer, you were the one responsible for maintaining the software that you built. And the idea was that, you know, you needed better telemetry, better things for you to debug the code that you have written.
Over time, we came to this concept of SRE books has been written about it. A lot of products and things happened, and it became its own practice, right? We have a lot of tools, a lot of, uh, data collection points, a lot of services, and a lot of things that create these as it as a standalone discipline.
So we have this team that it's now responsible for creating alerts and, you know, monitoring how applications are doing and informing the rest of the organization, why do we have a downtime and, you know, get somebody to fix it. But with this evolution and all this tooling and all this data that we have nowadays, um, it became, uh, its own silo, right? Like a silo in which people would look at the metrics, would look at their own dashboards, and then something happens and the ticket battle begins.
What do we mean by this? It's basically, you get an alert. That alert is something is broken.
You look at the elevated error rate, you say like, but all services show healthy, and probably it's a development thing. So why will open up a ticket, send it somewhere, somebody will prioritize it and hopefully it will get resolved by somebody that knows what's going on. It's not, you know, the containers, it's not the cluster, it's not the network.
It's something within the app. So the famous, like, you know, it's a bag, it's a feature, whatever, and this, you know, whatever is running inside the infrastructure, it's a black box. It's, you know, I just, you know, take care of containers.
And, and I found myself saying that many times where I was like, whatever, you're running as long it's on a container. We can make it scale, we can deploy it everywhere. We can build those things.
But then there is a reality, right? Then there is like, why the app is not working, why the system is not responding, why are we getting an elevated error rate? And our conclusion, right?
Was that it, it isn't one or the other. It's not like, let's take back everything that developers used to do out office and bring it back. And, but to come up with this concept of somebody that knows both the how the app is built, what's the goal, what's the business, and try to close the gap between that context.
So we are not exchanging tickets. We created this concept of cloud scouts that it's heavily inspired in, you know, the latest and greatest trend, even though Palantir has been doing it for a couple of years now, an open AI of the forward deployed engineer, right? The idea is that we have somebody from our team that joins our customers and works the way we used to work with our customers, embedded on their teams, looking at the problems and trying to close the gap by being fluent in both the product.
Like what is that we are building, what are the business requirements and why things are built the way they are? And the platform, like we are running es we are using, um, containers. We are, uh, you know, syndicating logs here, there, these are the alerts and how the way it works.
The idea is that these cloud scout, it's a hands-on support. It's not a consultant. It's more like hands-on support that knows both worlds and tries to close a gap in a constructive way by not saying like, you should be doing this.
Go ahead and do it right? Go ahead and close that gap. Uh, it's not meant to replace this at, it's not meant to replace, um, plat, uh, the the product engineering or the platform.
It's a way to have somebody that talks to both. So with, with that introduction, I will go to what it is, and I use the forward deploy seniors of software engineer embedded inside the product team to continuously observe, build, and improve the reliability layer of the stack. Which doesn't mean it's SRE, it's, it's in between SRE and the product.
Sharing that context, we understood as with AI that the most important thing when it comes to running these systems and running these platforms is understanding the full context. Because there are a lot of things that we do as developers by design to create these systems. We put business rules, we put things that oftentimes, you know, come later to bite us when we are running the thing.
But this idea sits in between all of them to work with all of them, right? To be actively engaged in a daily basis with the software engineers writing the code, the platform engineering, the platform engineers running the actual infrastructure and the software architect that it's like, you know, balancing in between, uh, this world. So what, what we are after and what we have achieved by introducing this to some of our first pilot customers is going from these, you know, dashboards.
We build dashboards, we have numbers, we have all of it to have something that you can really, by using AI and other tools, and I'll show an example later on, to have the real diagnostics, not just like a bunch of screens that shows numbers and things, so we can proactively detect and fix issues as they happen. The difference between being on call, and this has been something that, that we have experience, uh, running systems for our customers. When you get an issue, that issue, it's an elevated error rate on those, you know, errors that we are all familiar with.
And all of the sudden it's like, no, no, but that's a bag on the source code and we need to wake up another developer or somebody else to go look into the code to go fix that code because it's not a platform issue. And the idea of by leveraging AI and, you know, uh, all these concept of agents and all the tools that are available, uh, going from monitoring to building actual companions, that monitoring that do the monitoring for you, Johnny, yeah. Guy Courier, Futurum Group.
Um, is this a shift in South work's thinking that led to, um, the scout or, uh, is it something that the scout leads organizations to do? So it's, it's an interesting thing. I think it's a combination of our own experience doing this and a, uh, a gap that we observed while working with our customers in all the different silos at the same time.
So we have a concept that we call the dev crew that goes, sits with the SREs and work and do services on that front. We have our concept of fire teams that work on migrations, on building these apps and, you know, we do consulting in the traditional sense, like this is how the architecture should look like, but when, when all was running, we found this, like, this connection even with our own team style. So we were operating something and eventually somebody will go and say, Hey, this is a bag on the product.
This is, so we understood that there, there has to be somebody there doing this, you know, communication and, and pro and fast prototyping and showing the art of possible hands-on, not a consulting like, Hey, here is my PowerPoint deck, this is what needs to happen. But somebody that is actually capable of writing that code and say like, Hey, this is what we can do. Some of the teams will take it and evolve it, but, you know, lead the prototyping and the, I would say innovation from the ground no longer like a thing.
Because there is another thing that, that we have observed with ai, the number of lines of code that we write every day. As, you know, actually running writing code has decreased. We rely on all these AI tools and stuff like that.
So we need somebody that's capable of prototyping fast. And it was kind of a, a, a, um, mind change, if you will, that there is a lot of, I would say you treat the AI as, as something really you are against it or with it, right? So you have people that says like, no, all of it, nothing works like cha GPTs great, but when you try to take it to reality, it won't work.
And then you have people that says like, it's a solution to world peace. It can do everything. It will do everything.
Like we have a bag, we have ai, and we were like, eh, the technology's almost there. I wouldn't say like it's, it's done, but it's capable of helping a human in that loop to really accelerate the prototyping process. Therefore we said like, Hey, we spend a lot of time working with our customers from the consulting side trying to craft these roadmaps.
This is what's going to happen. And as with everything, you know, there is a big quota of trust and a leap of faith that you need to take when you say like, okay, I'm going to embark myself on this project. So we are trying to change that by bringing someone on the ground that works with the team and says, Hey, here's the prototype I built.
It's working. We can now chat how we want to iterate, iterate this, but it's no longer like, trust me, I know what I'm talking about. Gimme a million bucks 12 months and I will come back with the solution for all your problems.
I was gonna wait a little longer, but he opened Pandora's box. Um, so you had a Venn diagram of engineering, SRE and Cloud Scout would be in the middle. But isn't that the point of SO and SLIs the contract because by the, the proper way to run SRE is you take the business, you take the sort of the SRE, you build a contract around SLOs and SLIs, and so it, I don't know, it's not a clear picture if I understand SRE the way I understand it.
So That's a great question and, and you are right. But one of the things that we have seen from working with these organizations that have ma mature enough to say, Hey, we have contracts and we have somebody that is orchestrating those agreements in between the teams, it's that the, the, their position, like the internal position of these groups, whether it's like the product teams or the platform teams is to build defensive barriers to say, it wasn't me, right? So when you are trying to say, Hey, we are noticing something, no, it wasn't me.
Like my logs are clean, my alarm. So the role of the scout ideally is to look at the root of the problems and try to create a fix. And I will show using AI and MCP servers and it gets a really, really interesting in such way that it won't happen again, right?
It's a, it's a, an attack rather than a defense of like, okay, we'll create this barrier here. It's like, how do we fix the problem from the root rather than creating new policies and new agreements in between us to say, whose fault was that? But I mean, not to belabor it, but that's the whole point of an SLI is that it, there is no fault in the contract.
It is, it it is abstracted away. I By definition, by definition. Well, that's the thing.
Ask another question. So it seems to me like I, you're probably not focused on a legacy software, right? How so?
Well, like a banking application, it seemed to me it would be very hard to sort of contract in somebody who has to fit into something like Capital One, which has, you know, at probably now 5,000, but used to have 20,000 Java developers and then say, okay, we're gonna be that intermediate, be like, we're, our day job is gonna be coding, banking, transaction tier zero applications, but we're also gonna figure out the infrastructure. Like, I don't think you can find those kind of people. So I would say that the scouts are scarce because what, who do we put in that role are people with specific domain expertise?
And you will see from one of the examples that I brought that you need to know the industry. So if it's banking is somebody that has to have the banking experience and has grown their career in both ways, but we prioritize the, the domain, uh, the, the subject matter expertise when it comes to a domain, because otherwise it will be really hard to play and play. Because one key component to that will be the fact that you need to also understand the business you in that, in those discussions, when we are doing this, you have a product manager that comes with, Hey, this is what we are trying to achieve for our customers, right?
Then you have the yellower that it's focused on how do we make these transactions? And then you have the 3 9, 4 9 promise that somebody else is taking care of. So dealing with that requires a, uh, higher seniority and domain expertise because you cannot parachute into a bank if your whole experience has been, I don't know, writing, retail, whatever, retail or media or something else.
So it's, it's a really interesting senior profile and you know, it's meant to wear many hats and should have that experience of wearing those hats. Otherwise it would be impossible to really capitalize on the business knowledge and knowing what it takes to work on a tier zero transaction application. Good.
Okay. So what do we use? We use MCP zeros, we use copilots, we, you know, uh, create custom agents to find patterns and, and try to surface the decision to a human in the loop, right?
Uh, we don't, based on our experience, and hopefully the technology will evolve over time, we don't trust a hundred percent the solution to the ai. It's not like, Hey, now I have my agent, it's auto healing. We look into MCP server and fix everything and anything.
For me, it's more, these are the patterns that I'm seeing. How about we take a look at this, this, this, and that, and it becomes an interactive conversation with the agents and so on. So the idea is that these scout turns telemetry into action.
Can I ask a question? Yes. This might be a really dumb question, but are scouts human?
Yes. Okay. No, no, no.
Yes, yes. And it's a great question because everybody talks about it's agents, and it might have been like the name I chose for, for agents, but it's not, it's a human. All right.
Just making sure, just making sure That works with a, that work. It's a senior engineer and an automation companion. Okay.
That was good. And the idea is that, again, it's embedded with the product and the platform teams on these daily and weekly things. It becomes part of the rituals of the teams.
It's really embedded and spends time with our customers. It, again, as a, as I was saying before, it's not trust me, it's like I'm looking at what you're doing, these are the ideas that come to mind, and this is what I could prototype that solves part of that problem. It's KPI driven.
We try to figure out, you know, where are the issues, whether it's like a signal versus noise, like I'm getting tons of outlets, so nothing it's important anymore. Or whether it's the reliability or the time to resolution and, and how that works. And you know, we, we believe that it takes up to three months to get embed into the customer.
And after that, if our customers don't see value, we open up to say like, okay, that's fine. No, we cancel and, and move on. It's our take on software, uh, on the forward deploy engineering for SRE.
As you know, to your question, it's a human that sits at the keyboard writing code that talks with an agent that ri that drives or writes the MCP settlers to find patterns, draft and proposes fixes, and then becomes the final decision maker, right? This, this was what I meant by not delegating a hundred percent to the ai, but one of the things that's interesting about this is that we try to have that, that idea of end-to-end look right? So there might be issues, there might be drifts, there might be bugs, but what's the actual problem takes more than an alert on Slack on a log?
And, and I'll show that in a minute, but the idea is This is Tony with tecton. Real quick on that last bit, you mentioned that you have the human agent in there, then you have the AI that's analyzing and kind of assisting you, uh, and then it's making determinations. Is there some second level of human interaction so you can verify things are being done the way you want 'em before it actually applies these fixes?
Yes. Okay. Definitely.
The idea is that we use the AI companions to figure out the root cost to communicate the root cost to go deeper than the actual error. Like, hey, I don't know, an HTTP request payload is bigger than supportive. Yeah.
Whatever the scenario, it's Whatever the scenario. So you go down, you figure out, you do the RCA with ai, and then we craft the pull request and then the, uh, the final approval and maybe the tweaking on how the issue is fixed. It's run by a, by a human.
Okay. So that closes that loop. Yes.
And the idea is that we generate, and, and I'll show this on my example, sorry, we generate knowledge as we go. This means that we want to move away from things are happening, happening magically. Like, hey, it is a fix for a problem you didn't know you had.
Right? Or where this fix is coming from or who changed that? So we also rely on the ai, a lot of the, um, I would say documentation process, why things are changed, where it come from, like, you know, build me an RSA will review it if it makes sense how it's working and so on and so forth.
But have that process with a proper human in the loop that, you know, works as a fail switch that can say, hold the list, start over. And I'll, and I'll show an example in a minute, but the idea is to document issues that in, in such way that both sides, whether it's SRE or product, can understand that a tweak on the source code, you know, impacted beyond what you thought, right? Like you are missing a validation on a payload size when uploading a file.
Well that trigger a bunch of errors and a bunch of alerts and things like that and all that because of a business requirement that wasn't met. And those sort of deeper dive analysis that you can do when you have the compute power to. So, and I assume there's gonna be some sort of an audit trail logging Yes.
Et cetera. You can deep dive if you Need to. Exactly.
Now South Works typically embeds someone in the product team, the client product team, isn't that right? Yeah. Yes.
So when I see product team there, I ask myself, who, is that the embed or is that No, the client product team. It's the client product team. So we want to surface, like we, uh, our product teams, we call it fire teams, the ones we do the migrations and, you know, re-architecting a modernization of the applications are built to do things and move on.
We don't wanna put our customers on a endless retainer for doing that. We might do managed services to, you know, evolve the platform, but all the work that we do on the product side and the way our customers come to us, it's because our ability to, you know, build document and move on. So the idea is that we share that knowledge back to their, the customer product team so they know where it's coming from and, you know, closes the loop and, and reaches the, the, the ultimate owner that are they?
Well, that's, that's the top part. That's the knowledge. I'm looking down here and saying, seeing that the, the automation is built Yeah.
And then the product team adopts Yeah. But maybe they don't wanna adopt. Yeah.
That's how the knowledge comes back to the scout. Because there might be things that they say like, no, we don't want to, because I don't know, there is something that goes against the, our policies or there is a feature we are working on that goes against these and things that we might not know. That's how the knowledge, why the knowledge, it's a circle and it doesn't go one way only from us to the customer.
Right. One of the principles of Agilent of and of, um, of uh, uh, something else who's, yeah, the DevOps, Well, DevOps, but I was thinking more, um, pilots pilot more generally. One of the principles is that there, there's no such thing as failure because should it not be adopted, incorporated, or should it regress or any of those other sort of things, it's not a failure because you've learned something.
Yes. As long as you learn Something. So yeah.
So part of the value proposition here is a maturation of the product. Yes. Right?
Even if things don't get adopted. Yes. And, uh, And you know, what, You are touching in something that, it's really, really interesting that we have seen with customers that it's the lack of documentation or the lack of knowledge properly.
Like, Hey, we haven't done this, or why we present something and it gets discarded. If they don't have that practice, it will come back and, and we, we see that as maturity or, or helping them material. It's like teams change the industry, you know, people move around.
So you would switch and somebody would come back and say, Hey, I have an idea. And it's like, Hey, we, we, we try that before and this was the reason why maybe the business wasn't ready and now it's ready. But it's a way to reval and have that context while producing these iterations of the software product that you have enough context to don't go into battles that you already fought, if you will.
Yeah. And the, the best consulting companies are the ones that actually teach a client how to fish. And they, their goal is to get outta that business completely.
Mm-hmm. Now it sounds like you are built not a model to permanently embed, right? Sorry, when she asked The question.
Yeah, that's, that's, that's, that's part of the, that South Works model. All right. But, But yeah, but there's questions on that, which is in sort of DevOps or in general, there's no such thing as done.
Yes. But there's no done. So the real question then is do you have a setting goal for an organization to be basically work their way out of working, which outwork And Yes.
And um, yeah. And then what, how do you set those kind of goals to begin with? Because it, it doesn't sound like, you know, you talk about the embedding and then we don't know much at first.
So how do you, how do you set sort of the expectation that we're gonna be here for till you get to here and what does here look like Part? And and that's why we always start with like, okay, we need three months to work with you and start producing stuff to set the baseline. And that, and what we've seen in practical experience can go two ways.
One is like, this is our baseline. There are things that we would love to happen and we can say, okay, it would take six months, 12 months to get there. And then we shake hands and move on.
And there are others in which we, our customers don't like, know that something is wrong and it's really hard to explain what's, what's broken, right? So we prototype the things and build the knowledge around to say like, okay, here you have an example of how to deal with this situation. This is the way you should evolve it.
We can help you evolve it. Or here is take it. And they might say, that's great, but now I have this problem, right?
Or that's great, move on. Uh, we try to be pretty transparent on the assessment because that's our definition of success. So I wouldn't go in and say, Hey, we will, you know, reduce the number of issues that you have if I don't know how many issues you have.
It's not the same thing if you are being bombarded every day with a thousand notifications to somebody that gets five. So I cannot, I, I don't wanna make promises that I cannot keep. So our idea is to explore from the ground up, like not by, you know, coming up with an idea, this is a solution to all of your problems.
And then drive a sort of roadmap and say like, okay, here you have something that it's working. That it's not like, Hey, in the traditional consulting world, that will be an assessment or an evaluation that would turn into a PowerPoint that will be, sure, now gimme a million bucks 12 months and come back later. Our idea is that from the ML is to build something that's actionable, that it's hands on, that it's working, and then offer like, Hey, we can help you evolve this, or here is the instruction manual, go ahead and, you know, keep it yourself.
That's, that's the goal, to maximize that idea of independence. And as guy said, like The role of the scout is to help organizations mature and put the technology to work to them. It's a maturity thing that oftentimes, uh, in the traditional consulting will be like, sure, I have a thousand slides PowerPoint, I feel smarter, but it feels years away, like all these horizon three, five, whatever.
And on the other side it's like, we want to be tactical enough. So you feel that evolving and truly an agile way of working. So, and we believe that AI has opened up the game to really, uh, prototype fast and show value.
And if it doesn't get adopted, it's like, sure, we've learned something. We won't do this again. We'll find the solution.
Yeah. This sounds very similar to the TAM program. Is this something that is built in or is it an add-on to your service?
It's a, um, it's a great question. We, We started piloting this with customers by saying like, Hey, we will show you what we are capable of if we embed and figure out the problems we have had. And, and mostly when implementing ai, and, and I'm not talking about building chatbots, I'm talking about like putting AI to work for infrastructure work or software development practices in general.
That it was this deadlock when it comes to decisions. 5 or it's code better and everybody pitching thousands of lights and stuff like that. And that was the angle in which we started building this as maybe the entry point.
There might be a longer term project, something that needs to be built. There might be a managed service that, you know, it was that you lack the capability for maintaining or doing the upkeep of your system. And what you really needed was a smaller team looking into the constant evolution modernization because as you said, there is no done in DevOps or maybe it's that you just need, you know, somebody to break the analysis paralysis for the decision.
So it's kind of an entry point that we started to turn into a practical activation of the technology. So it's hands on. And oftentimes we recommend this to our customers because a trend that we are seeing in their, their organization or when they pitch a problem and they come to us with like, Hey, let's build a two year roadmap for AI adoption.
We have all these thousand tools and we never try them. Uh, I read an article that it might be good if, and I, and I'm like, let's try three months with something that it's more flexible, more agile and nimbler and see how your organization adapts to that. And when we can, we can figure out how the next 12 months would look like.
So it's A limited term engagement and you can figure it out from there. Yeah. Based on customer adaptability and visibility.
Yeah. So I like the wedge idea, like the co companies like trying to figure out all the tools and should I use MCP, should whatever, and, you know, and that idea of like an embed personal, let me try it out. But then the counter to that is my concern is organization for these things being dangerous and unsafe.
Like I haven't thought through an MCP architecture, particularly in a, in a bank or a place where, you know, there there're, there's zero trust, you know, OCC regulated, zero trust compliance. So it seems like a little bit of a dichotomy there in that I'm gonna help you get started 'cause you're all over the map. Nobody really knows what to do.
But same token, I'm not sure I want these outside people implementing MCP or you know, or sort of agentic type processes in a world where we still haven't figured out what that looks like in our world. And does that Make sense? Yes, definitely.
Because that's, that's the key To why it's embedded, what it's embedded inside of the organization. Because one of the patterns that we observe, even when pitching our own AI projects to be at enterprises, it's that concern that, you know, you design, you build a castle in the air where all these MCP servers will exist and everything will be connected to the internet. And you have all these fantastic models looking at your source code and then all of the sudden comes the InfoSec guy, the CIO, the ciso, whatever name you choose to it, and it's like, yeah, 90% of it won't happen, right?
We haven't seen that, you know, level of pushback when it comes to working integrated. We have seen it from like the outside pitch. So this is also an answer to that, that it's like, sure, we have seen companies saying like, Hey, sure you can use GitHub copilot and all these models or bring codex and all, and then it, it's, they don't arrive right.
Like you, they won't happen. So that's why for us, for this model to work and to become the real fire starter for the evolution of the organization, it needs to be embedded. It needs to be sitting on that table where, you know, open AI is a no-no or deep seek no way.
We are going to rely on a Chinese model or you know, we wanna operate with bedrock or we want our own private deployment and I can get a thousand blackwells if I want. So those conversations that you are not exposed unless you're sitting at the table, are crucial for pushing the envelope forward. Right.
But it sounds like you're offering this as a service. So are you bringing some kind of pre-programmed, like building blocks to the table and are those getting reviewed by these clients or are you kind of greenfield building as you go into a new client each time? So bring the knowledge, the experience, and the expertise of doing it.
We found that with AI there is little to no value to bring actual source code because for the InfoSec process and all those things, even though if it's AI generating and such, they feel more comfortable on a green greenfield. Okay. And seeing, you know, how the building is coming together floor after floor gives them more confidence than, sure I'm bringing an agent that will take care of Yeah.
Solving Everyone loves a black box. Everybody loves a black box until you meet with, uh, CSO or somebody that says like, sure, open the kimono. And it's like, yeah, that's not possible.
They Get pod Everybody loves a back box too.