Demonstrating AI-Assisted Development for Leading European Streaming Service with SOUTHWORKS
A Cloud Scou is a forward-deployed engineer who joins the product team to co-own reliability, scalability, and evolution. Drawing from the Forward-Deployed Engineer for SR and AI-Managed DevCrew models, Scouts act as both architectural advisors and implementers — blending human judgment with AI-driven companions to build, test, and tune cloud-native systems. We walk through how this embedded approach fosters continuous improvement, strengthens technical decision-making, and creates a shared sense of accountability between Dev, Ops, and AI.
Johnny Halife from SOUTHWORKS presented an example of their work with a European streaming service facing issues with their electronic program guide (EPG). The EPG, built on Node.js, Lambda, S3, BigQuery, and XML, was experiencing blank displays due to ingestion problems. The issue was traced to an unexpected 413 error indicating that the request entity was too large, specifically related to image transformation failures. This problem was impacting viewers, who were seeing blank screens.
To address this, SOUTHWORKS employed a Cloud Scout, leveraging tools such as GitHub Copilot and their own MCP servers, which are connected to AWS CloudWatch. The process began with the scout prompting GitHub Copilot to create a Jira ticket, which was then assigned. The agent analyzed the error by running CloudWatch MCP, finding related logs, and contextualizing them within the solution codebase. This analysis revealed a missing validation and a data conflict between files, providing evidence-backed insights. The agent then proposed solutions, including code changes, which were compiled into a pull request.
The final step involved a code review by the Scout, along with standard organizational pre- and post-requisites, including SonarQube and linting. This process, previously taking days, was reduced to a few hours. By implementing this AI-assisted approach, the streaming service experienced faster issue resolution, fewer noisy alerts, and predictive scoring for deployments, resulting in a significant reduction in recovery time. This approach enabled them to transition from a defensive strategy of increased monitoring and tooling to a proactive approach, aimed at preventing issues before they arise by analyzing past incidents and identifying potential risks.
Presented by Johnny Halife, Chief Technology Officer, SOUTHWORKS. Recorded live at KubeCon North America in Atlanta, Georgia, on November 11th, 2025. Watch the entire presentation at https://techfieldday.com/appearance/southworks-presents-at-tech-field-day-at-kubecon-north-america-2025/ or visit https://techfieldday.com/event/kubecon25/ or https://southworks.com/ for more information.
Transcript
I'll show a, an actual example of something that, that we have done and, and the plumbing and the tooling, how it came together. Uh, we were working or we are actually working, uh, with a European streaming service that distributes content, uh, all over the world. And we were running their streaming app in general.
Uh, they had an issue with the electronic program guide, or as the media guys would call it, with the a PG, the TV schedule that you show was dead. And it was showing nothing. And I'm sharing some, like it's running on no JS Lambda S3.
BigQuery, XML ingestion of the programming was blank. Everybody was, you know, complaining. Uh, view Is my experience every time I push that button to try And it's like, I don't Think I'm alone at the tape.
No, No. So it was like, okay, something is broken. If there's one thing South Works can do for the world, it's fix this problem everywhere.
Continue. It could. So this is how it works.
So the customer puts data on an EPG that files get monitored, goes into a queue, it goes into a database. There is lambda processing, sends another queue, and then the output of the A PG, something was broken. And as guy said, I'm trying to watch TV and I'm getting like a blank thing and nobody knows what's happening.
Yeah. From the traditional side, we notice like, hey, there is an, an issue with the a PG ingestion. The team, the SRE or the monitoring team also received this, and our scout received this.
So it was like, hey, event, whatever, let's keep the ingestion, the remote server return an, an unexpected for 13. That means that the request entity is too large following what error was found in the image transformation. And there was a, an issue that was the root or the potential cause for these EPG to go black.
So how do we react? We have, you know, our scout has the tools and the MCP servers in this case connected to the AWS CloudWatch. We were using, uh, GitHub copilot.
We prompt GitHub copilot, first of all to create a ticket to say like, Hey, we have reviewed the attach alarming production, create a Jira ticket in the board to track the issue and assign it to me like I'm going to fix this. The scout continuous interaction with the GitHub copilot asking to assign the ticket. And, and the agent starts reasoning, right?
I will help you create the ticket for the A PG and the ingestion and it will run some queries against adian MCP server to figure out whether Jira issue is already present for that or not. Then it will check the, the resources list, the projects, and it create the bag issue. It will show me, hey, or it will show our scout in this case, Hey, I created a ticket.
This a bag. Um, it's an issue. It's been assigned to you.
There is a link, the critical error is a data conflicts between events. Uh, it was, you know, eight instances of the image creation fail because of being too large. And the next step is investigating what's happening.
The ticket is now assigned to you for, to work on this is how a ticket looks like. It, it's like a damp of like who are the effect, which are the affected files and everything. And we keep going, right?
So we continue the interaction with the copilot that is connected to all these things, and we now ask to analyze the error. The agent that we have starts analyzing the, the, the error running. The, the, the CloudWatch MCP find the related logs goes into a context of the solution code base.
So it figure out the resources, the local groups, it goes into AWS as a human would, you know, troubleshoot this issue and it will find the issue. Sorry. And say like, Hey, here is the analysis.
This is what I found. I found two issues. One is that you are missing a validation.
And the idea is that there is a data conflicts in between the files that have been dropped into the packets for ingestion. And there is a conflict there. I like that it's got the evidence from code right there so you can see what it's actually finding of you and not what is it guessing at?
That I'm not sure is really Accurate. No, and for, and, and you touch into something that it's really interesting that I show there that it's, it's super important to have the reasoning process and the, and the resources that you are looking. You can See it's thinking, It's weird to use thinking, but Yes.
And you can see it reasoning. There you go. Now the technical term, Inferring, Inferring and against which files.
And that's the other thing because again, technology almost there, but hallucinations are real, right? Like, it, it, it could happen. No, but this brings up an interesting point.
I was waiting, but just like you, you broke, you know, broke us into this. So, so are this, I mean, how productized is this for you guys? Does the scout have runbooks?
Yes. Okay. And and the idea is it's two.
So we have the runbooks and we create, the outcome of this process is a runbook for the ops team. Like we, we end up building the Ram book and it's like, hey, I feel like a little bit of the, a little bit of the, the, the secret strength, um, um, maybe not so secret for South works is, is documentation just generally all, all yes. All around aboard service, service definitions, uh, at All.
Yeah. And, and Yeah. And it's not only, and I would say it's a big part of our success, but in this world where you have agent agents reasoning on top of things though, who, those who have written things are meant to thrive because all these context, all these decisions, why we took the decisions we took, Yeah.
Providence, prominence, transparency. Yeah. At some point when we've been in business for 20 years and communications and writing things down has been front and center always.
It was part of our promise of being transparent, of sending every day what we work on, how we work on links to will request or change logs even before that. But now with ai, it opens up a new opportunity that it's like there is something with enough compute power to reason on top of all these things that we have written. And that it's a unique advantage to create the most accurate runbooks, if you will, with all the decisions of those things that hadn't been adopted.
And also to tap into our own internal knowledge like, hey, I run into this problem. Have we seen this before? How we fix it?
And, and it becomes a real, like everybody said, documentation is a living and breathing organism. But truth to be told, nobody went back to read it, right? Like we wrote it and we don't know if our customers look at it.
Now, if you have something that can reason on top of that, and it's, you know, with enough card rails to look at the proper things, whether it's source code or documentations that we have written and not going out to the internet to figure out what Jane Doe wrote on a ready thread somewhere, and, you know, be grounded in the actual context that we have, that's crucial. So I, I've built a fair amount of ag agentic systems and I wanna follow up on two questions. Your question about is this human?
It doesn't look human, it looks ent. And then the second, his question was, do you bring tools with you? And it looks like you're bringing tools with you because I, I saw about 30 MCP servers on that list, and I don't think you've developed those every time you go to a client.
No. And, and, and some of them are existing. You follow up on those two questions and one, 'cause it doesn't, it seems more non-human, especially in the reasoning than it does human and everybody does human in the middle.
So that's, and then to his point, um, it sounds like you are bringing in a bunch. It, it looks like a tool. It it, It looks like a tool, but it isn't.
It's a service, right? It is human. It it, it has an AI companion that it's working on it, but the orchestration, it's with the human in the loop.
The human in the loop is not kick this process, go to sleep and figure out the human is actually interacting with the agents. And when it comes to the cps, we might have some missing links for things that are not yet available from first party providers. But what we know is we have identified which are the first party MCP servers that might exist and when to use them and if they are trustworthy enough.
So if I were to write an MCP server, which we have done in the past, we will make it open source. We'll put it out there, we'll have all the community, you know, bash it, kill it, or do whatever on top of the source code. And we are interacting with it because as you said, it runs with a big trust and responsibility.
And I think, as I said before, the, the big part on the knowledge is like knowing which ones to use, which are the ones to use, and who to ask whether these are trustworthy or not. But if we were to bring our own pieces of circle, we make those open source. And the real reason is because we have nothing to hide.
And we believe that there is a lot of, you know, Bad actors that might be exposing data to models that you don't want exposed. So that it's our approach. So it is a service, it's human driven, it has a big component of AI and that knowledge and that, you know, documentation of the do and don'ts is way more valuable than the actual source code of an MCP seller.
So if we were to write one that it's not yet available by the first party, it's there in the open source, have the community kill it, pass it, send, put request to it and review it as many times as needed because we believe on the being on the, on the room, like sitting at the table, having the discussion, that's where we see the actual value rather than the source code that we are writing. This is fun. So, ah, we, we do the proper inference or the reasoning on top of the files.
We figure out, we connect to the source code, the solution, we identify the issues and the agent will provide recommendations. But one of the things that we say is like, hey, before jumping into conclusions and creating source code and stuff like that, open up a pull request, open up a pull request, you know, and or create the r the RSA that then will be tied to a pool request explaining what is going on. Because we don't want magic, we want like, hey, we look here, we saw this on the log, we understood this, this and that.
We might edit the outcome of this with our own knowledge of like, hey, yeah, that's not completely accurate, or there is something that the agent might have missed and we will document the root cause of the issue. As you find an issue on the logs, and this is what it translates to, we will ask now to add the, uh, the image size check. If a query has a file size data available, let, let me look at the images to understand the image selection logic better.
So we will troubleshoot using the agent and then we'll start applying the fixes, right? Okay. So yes, we'll use string pattern package to check and we will touch this and that file here and there and the outcome of that is a pool request.
It will never go into production. It will go into, hey, create a pool request, sorry, a outside is a, I need translate into the pool request, site limits and versioning for file processing. This triggers a couple of things.
One is the actual scout review of the source code. The second thing is the, all the set of pre and post requisites that the organization might have in terms of, you know, running sonar queue and compliance and LinkedIn and all the things. So it's a conformant, the they don't conform and test as if it was written by a human.
The other thing is like, we also write the approval chain as it was set by a customer. So if you have, like depending on the files that you touch, you need an approval or sign off from two people, three people, somebody on this team, on that team, it would still happen. But what won't happen is that this whole process have took two hours and I'm being generous.
In the previous world, it would have been an issue. Somebody would look at the law, say, Hey, this is about, we create a ticket, we'll assign it to somebody else to look at it tomorrow in two days it would take another two, three days to fix it would take another day to approve. So by doing this, the actual value is that the recovery time of a an issue that we all hate, that it's like going into the a PG and being blank, it's two hours, three hours.
We went from five days to five hours by looking at the right thing. And, and this is the value of being a may of being actually looking at the infrastructure, having access to the logs that were pre cryptic that we're saying like, Hey, sure there is a four 13 everything that starts with a four because these are the things that I like. It's not me saying it is what I heard.
Everything that starts with a four, it's a new, a user initiated issue, so it's not my problem. Right? 4 0 4, 4 0 3, like you don't have access and those sort of things.
So the way that we help, you know, connect all these tools, build a pipeline, build the, the, I would say build the, the discipline to run these tools and show the actual value of using AI and being there. It's not that the code, It's written by ai, it's not the technology itself, it's the time it will have taken to fix this in a traditional organizational setting where tickets will start flowing from one to the other. Like, Hey, have a look at this, figure out what's the problem?
And so on and so forth. So that's the real, you know, value when we have the discussion with this customer, it was like, hey, we went from zero to fixed in a couple of hours and it was like, this is great. And there was no override to any of their own processes.
Meaning that if they have quality aids, quality controls, controls on their put request environments in which they run regression testing and those sort of things, that would still be valid. It is not that the AI is taking over and fearing out, it's, you know, uh, To me it's, you're taking advantages of the efficiencies where AI can help still keeping a human in the loop, but it makes sense, you know, back in the day when I was an administrator and I had whatever observability tool I had and I had a hundred alerts, which of them really matter? Being able to filter that out and focus my time where it matters, that's where it ops comes into efficiency and gaining a lot there.
And, and that's, you know, when I talk about the signal versus noise issue and those sort of things and, and people often ask me like, okay, so this is going to replace SRE function, and I'm like, no, you need somebody to implement the telemetry, to have the logs to increase the complexity. It's a hybrid, it's a hybrid approach. Yeah.
And at, at the end of the day, the How the industry matured over time, how, you know, the tooling available and the the things that we can look at, the things that we can observe and act upon have grown. It's why we can build the solutions because we can ground it on all these data that didn't exist before or, or wasn't available at the time. So this is capitalizing on top of all of that work.
And, and, and it's a compliment to that or an acceleration, you know, benefiting from the power of being able to reason, influence, compute, whatever name we want to use of large amounts of data with a purpose that it's like, hey, uh, the human is there to give the AI purpose. Like yeah, find me a pattern, all log files start with a Z or they are start, like, those are the things that if you ask point blank ai, like, hey, what's the pattern? Sure.
Everything start and, and giving a purpose is the role of the human. The upside is that if you don't know what you're doing and you understand the problem, the compute power lets you crunch data at a pace that wasn't possible before. So with our customer, we go to the, you know, faster issue resolution up to 40% in 90 days, 25%, uh, fewer noisy alerts.
One of the things that we have done is that Slack channel goes through a, uh, pipe of prioritizing, Hey, this is coming, it's repeated and doing the ization based on the logs and some other things. So to your point, yeah, if I get a thousand alerts of the same thing, I would rather just get one as complete as possible so I can go look at it. It's not replacing the actual logs, it's, you know, summarizing or doing the, Hey, this is something we think you should look at, type of thing.
And we have helped them with what we call the predictive scoring before the deployment that it's like by looking at the things that we have changed and the issues that we have had in the past, how likely or how risky is to send this into production? And I assume this is adaptable to different tools that the customers are using. It's not any kind of a cookie cutter.
It has to be this has to be this. No, and it's multi-cloud. So you could be on Azure, you could be on, um, AWS, you can be on GCP, uh, you could be using multiple languages, you could be using bms or you can be using Kubernetes.
And there is a minimum set of telemetry that we need that it's not prescriptive on the tool, but it's like we need logs. Like if you don't have logs, well, let's start there by, you know, aggregating and rotating and keeping buckets with your logs and those Yeah, or even opening a ticket. You used Jira as your example.
What if they're using ServiceNow? It's still a, It could be ServiceNow. It could be GitHub, it could be, uh, we have seen a really interesting workflow in which you can convert GitHub issues into pull request using the command line.
So we'll create the issue directly on the repo and it will turn the RCA and all that into an actual pull request. So you have even more context in the same thing. It varies by like we, we understood that the ultimate goal of this is to accelerate the cycles to really empower the human that can make the proper decisions to get to resolutions faster And to adapt to what the customers are using in terms of tooling and in terms of cloud and instead and their own maturity levels, right?
Uh, if I was to say, Hey, in order for you to start using something like this, you need to be on the edge of Kubernetes. Running on AWS using it won't make sense because those who have this problem, these problems are usually trying to catch up. And this is a good way to ensuring that, that, that that process of, of catching up one come later to bite you or generate an issue as you start creating these contentions, uh, nets or these analytics tools to predict the risk of the things that happen.
Why now? Because the complexity has increased like never before, right? So, uh, you have people operating specific tools, so it became so complex that you have the Kubernetes cluster expert that just look at the co.
So we need somebody to make sense out of all that data across the board and to have a seat at the table and to be able to discuss and to have that, you know, ownership and, and that stewardship of pushing the envelope forward. We have the, we have our, you know, 20 years helping customers at the edge trying to modernize their platforms. And this is a compliment to our existing stack of services where we have our fire teams for product development or the dev crew for sustaining engineering or reliability or the tactical precision of saying like, Hey, I need an expert on cornes or something like that.
This comes to, you know, push the envelope to bring innovation, to be the real hands-on keyword support that you need to foster that innovation process. And not just to come back with a thousand slides PowerPoint deck. The way The, the way we, we thought about this and, and we have our service of dev crew to do that, you know, reliability work to work as an extension of SRE teams and oftentimes to take on SRE for customers that don't have that, but that is our defensive play, right?
Like we are all the time increasing like the alerts, the monitoring, the tooling and all that to keep the system alive. Uh, the cloud scout is our, you know, our attack play, figuring out how the future is not to wait for issues to actually happen to understand what are the opportunities to prevent this from actually happening, uh, in the first place. That's why we came up, or how we came up with the risk scoring for pull request and those sort of things by crunching all that data and so on.
Um, we, and, and this was, you know, how we work with customers or how to engage with customers. We ask them to nominate their toughest reliability challenge with sending the, to send in your scout where we'll have it and limit the pilot back to your point on, you know, how we know when, when it's done to limit the pilot to start trialing, to start figuring out and seeing the actual result to convey into a proper business case in which it will make sense. And then if SRE was about watching the cloud scouts are about billing daily when incidents happen and make sure that this won't happen again.