Transforming IT Operations: The Power of AI Agents, RCA, and Operational Twins
In this session, learn how AI is transforming the world of network operations. The shift in the industry is leading towards AI-drive infrastructure management. Learn how Selector.ai is combining copilots, RCA, and Operational Twins together to help streamline your operations and ensure teams are working at peak efficiency.
Presented by John Capobianco, Product Marketing Evangelist. Recorded live at Networking Field Day 37 in San Francisco, CA on March 20, 2025. Watch the entire presentation at https://techfieldday.com/appearance/selector-ai-presents-at-networking-field-day-37/ or visit https://techfieldday.com/event/nfd37/ or https://www.Selector.ai/ for more information.
Transcript
Alright, well thank you so much for having me here. We're gonna be talking truly about transforming IT operations with the power of AI agents, root cause analysis, and operational twins. Uh, before we get started, I'm not gonna kill you with PowerPoint.
I do have some slides to tell the story, but we're gonna really focus on the videos and dissecting some videos and hopefully be able to pull off a bit of a live demonstration. So today's agenda, we're gonna introduce selector in case we're new to you. We are a new startup.
As Tom mentioned about a year ago. We were, you know, unveiling ourselves to the world through network field day. Um, when I joined Selector in August, I actually used that video from Nitin at Tech Field Day to ramp myself up on the technology.
I probably watched it 15, 20 times. So watch that video. We're gonna be talking to our infrastructure with a co-pilot for it.
We're gonna be looking at root cause analysis and AI agents and closing the loop on automation. And we're gonna talk about a operational twins and I like to say looking at the past, present and future of your network and then a bit of a summary. So just an incredible introduction.
Thank you Tom. I'm really humbled by that. Uh, but I'm a product marketing evangelist.
I joined in August of last year. Um, I've written a couple books about network automation. I've got all kinds of GitHub code for you to use YouTube videos.
Um, I've been all in on artificial intelligence for about two and a half years now. And let's talk a little bit about selector. So it really is about providing you with effective answers, giving you root cause analysis and cutting the noise out of the signal at any time, past, present, or future where you need it.
I could get an alert right here on my phone in this room with all the information I need and giving you meaningful answers. Not just giving you alerts, but giving you smart alerts and insights into what's going on with your network. So what customer verticals are we in?
I, I think we really started with large ISPs and large enterprise, but in the past year we've exploded into retail broadcast and video entertainment and sports finance, healthcare. It really is a solution for large scale enterprise and ISP networks, but, but can fit into a lot of different verticals. We have different target environments, primarily your data centers, your backbones, your SD WANs and wide area networks, but even your Kubernetes architectures and microservices and applications.
Um, I think what differentiates us, and maybe this is important to highlight, is that it really is observability as a service if you get selector. It's not like you need to have a team of data scientists in your organization or experts on artificial intelligence. We're there to help guide you, we're there to help set up the system, set up the platform and the infrastructure so that you can just start using it for observability.
And we've, we have dramatic ROI our customers are experiencing 97, 90 8% in ticket reduction alone. Root cause analysis leads to dramatic improvements in meantime to innocence, meantime to recovery. And we're seeing a lot of tool consolidation.
Uh, I think the number one question that I get about selector is where can we run? What type of infrastructure does an organization need to have to adopt selector because it is a built on Kubernetes and microservices, we can run on any environment. The most common environment we see is hybrid cloud, right?
So a mix of on-prem, a mix of in the cloud, a mix of multi-cloud. But that's not to say that we can't deploy in an air gapped on-prem only environment. We can at the other end of the scale, if you're a pure cloud organization, we can operate in your cloud or in our cloud in terms of the pricing and the dollars.
We build our pricing off of use cases and nodes. It's not about the volume of data. So it's a little bit of a different focus than some of the other platforms that, uh, your prices explode as you start to see more and more data.
We want more data in our platform. It makes our outcomes better because of the machine learning and the correlations that we can make. So we actually encourage more data.
John, John, quick, quick, quick question. Yes, Scott Roon, good morning. Um, we had an interesting discussion on the term air gap yesterday.
Okay, just want some clarification. You can deploy the solution without any access to any cloud resources, like a completely private network. The caveat being that you have GPU capability and are able to host an LLM.
Okay, But like in a, a totally cloistered, you know, No need to nothing, no heartbeats, no cloud dashboard. That could be alternate internal on-prem solution. Again, with the caveat being that you've got some GPU and the ability to host an LLMA large language model.
Thank You for the clarification Ron Westfall Futurum group. And uh, this is, uh, I think definitely provides a, a valuable level set on, you know, why selector AI and use case naturally AI comes to mind. Can you share any comments or observations about what's going on with AI workloads?
How is this impacting selector's ability to sell into, you know, the organizations? Yeah, so the ability, so for that same scenario of having, having GPU and being able to monitor your, our own platform, right? So selector actually monitors itself by looking at the Kubernetes nodes.
And a big focus in our roadmap is going to be AI observability. We see things like sharding workloads across multiple GPUs coming from Cisco and other vendors where you might have multi rack with GPUs spread across them all using IP for the connectivity. You need observability there, you need visibility there.
So we can provide visibility down into that layer. Very good. So just a little bit of a story here.
How did we get to AIOps? Now? AIOps, right?
Just to, you know, I know it's a loaded term. Why not ML ops? Why not ML slash AIOps?
It's just the name, it's just what we've landed on as AI operations, um, as an evolutionary step. So traditional operations, right? Some people are still doing things traditionally, not everyone is at automation, not everyone is at ai Syslog s and p manual effort, reactive DevOps comes around around 2000, 2001 with agile building off of lean purpose practices from the manufacturing space.
And we get things like continuous integration, continuous delivery in the software space, and developers are starting to spin up their own infrastructure now, right? We're breaking down silos in this continuous release, deploy, operate, monitor, plan code, build test that just continues forever. Networks, what, 15 years later in 2015?
2014. Finally start to adopt network DevOps infrastructure as code Ansible, Python, all the new tools that the network engineers have to use with a focus on automation. I strongly believe that AI is just an extension, an augmentation.
The next evolutionary step beyond network automation. So rapid change doesn't come in normal times. 5 released their GUI and everyone started to democratize artificial intelligence.
We, I can't move over here we are, you know, a year, two years past the inflection point, but as you can see, we're going up and things are changing rapidly. But not only are we going up older ways of doing things are starting to fall off, people are realizing there's a whole new tool. You know why if I can't lift this boulder, I can put a lever under it and move the boulder, right?
How many people are gonna still move the boulder by hand now that we have a lever? And that I think is, is a decent analogy that we have this word calculator in the term, in the form of artificial intelligence. Now looking at the variety of ages here, some people didn't live through the computing boom, right?
They were born in the two thousands, but you know, in 1984 I got a 2 86 or something, world changes, internet comes out, world changes again, mobility comes out, world changes, more cloud services and streaming now into AI and it's just gonna keep going. Narrowing ai, generative ai, A GI right? Now, what does this have to do with networks and network operations?
I believe that network network operators today are looking for needles and stacks of needles, at least in a needle, in a stack of haystack. You can use a magnet or a tool to get the needle out of the haystack, right? But a needle in a stack of needles where we have corporate offices with traditional core access distribution and wireless on top of it and a data center or two multi-vendor, multiple sources of information and telemetry.
Then we move into our ISPs, multi ISP links, diverse ISPs, if you are an ISP, the complexity of operating a network at that scale, then we overlay clouds on top of this Multi clouds and Azure dashboard or Google dashboard, an Amazon dashboard, different tools to monitor those dashboards. Then we have even more stacks of needles. The actual tooling that we're supposed to all know, every single one of these platforms.
All their syntax, all their ins and outs. How do we do it? How do we maintain this infrastructure?
Especially when there's a problem? Where do we look right now? This is the core proposition value of selector.
We're gonna take all that noise that you saw here, topology events, configuration logs, metrics, metadata, and we're gonna bring it into a single unified platform that has really three core values. The event correlation, root cause analysis to drive ticket reduction and to close the loop and automate remediation. Our FIR industry first network language model.
So similar to a large language model, we have a network language model. I'm gonna peel that out in a little bit in a bit. And a copilot.
So tool and vendor consolidation, a self-serve portal. This really democratizes things. Your CTO level people, your senior network engineer people, your frontline help desk people.
Everyone can ask questions in their own words and get context back. They don't have to have commands memorized, they don't have to be ccis, they don't have to know the ins and outs of rest. API calls and JSON and all of that.
It's natural language. And finally the operational twin where we can model your services in business and have what ifs scenarios, what hap what would happen to my traffic if link one goes down, but also replay the he the past and have a DVR of your network before I go into the network language model. Are there any questions?
Am I going too fast? Question? Yeah.
Can I ask one question here, John? Yeah, Please. Bruno Wallman.
Um, you were talking about digital twin and modeling your network is, uh, what is selector using for that? Like there, you know, that's been talked across, um, you know, different vendors can provide that. Does um, selector have tools that you built that do that or you covered doing something else?
Yeah, so, uh, the, the third portion, Bruno, if you could just hold that thought. I'm gonna really do a deeper dive into the operational twin. Okay, perfect.
Yeah. And, uh, Question on market progress. Is there a way to characterize how many customers are really on board with NLP and You know, um, I think we're approaching 50 or 60 customers.
Okay. And what I'm finding very interesting and I'll, I'll make, uh, I'll allude to this a little bit later. On this point of the root cause analysis and the closing the loop, what we find our existing customers are, they have so much trust and confidence that we've earned over the past few months or years with them in the root cause analysis.
It's become routine to them and they rely on our platform. So now they're saying, well now that you've identified the problem perfectly with the root cause, can we close the loop on that? Can you kick off a playbook?
Can you hand off to say potential? Can you run an Ansible job? So there's so much confidence that the root cause is, is so highly accurate and, and, and valid that they wanna now extend that trust to maybe doing configuration management, shutting ports, bouncing interfaces.
So we've made a lot of traction with the root cause in the copilot. Alright, So on the network language model, we, this is, this is just, um, a visual representation, right? Um, we start with a core model, a base model, and that's where you see the llamas and you got the llamas from in that air gap scenario.
3 or or Gemma. If you're Google, if you have cloud, this could be Gemini from Google, we, we see customers gravitating, which makes sense. Your emails in Google, your calendars in Google, why not have your LLM in Google?
We allow for that. The name selector comes from an SQL select statement, SQL select star, right? So that first band of fine tuning is natural language to sql.
So I don't need to know SQL queries, I can put it in a natural language phrase. And we fine tune the base model to so-called understand sql. Now every customer gets this and then we add a layer of per customer fine tuning.
Now this is known as raft. If you wanna look up the, how the technique is used. Retrieval, augmented fine tuning raft, where we're retrieving generation, augmented generation, uh, excuse me, retrieving augmented data from the telemetry and then fine tuning the model further.
Now to clarify a few things, selector does not have a so-called parent super network language model in the cloud. And we're learning from your telemetry and data, it's per customer and customer A is separate from customer B, there's no, they're not learning from each other. Truly is an enterprise deployed network language model specific for the customer.
Now this is where the machine learning comes in. Telemetry comes in, which is typically metrics, numbers that we can automatically baseline and log mining of SIS log named entity recognition. Um, you know, the THESYS log number, the event IP addresses as numbers, whatever we can mine out of the SIS log.
This lets us do anomaly detection. Anything above the natural baselines is an anomaly. Certain keywords in the name entity recognition is an anomaly, which lets us cluster these things together, which leads to correlations, which then the artificial intelligence can detect a root cause, which then interfaces with that network language model.
Now the network language model, right? We have the co-pilot interface, which you'll see in a few minutes. Show me the circuit errors at San Diego.
We translate that from natural language to sequel. I'm doing it on my, um, which goes into the storage and pulls the data that it needs, which augments the generation of the output from the nat, the network language model.