Selector AI RCA AI Agents: Closing the Loop on Automated Remediation
AI can help network operations teams with root cause analysis (RCA). Gone are the days of reactive troubleshooting. Instead, see how Selector.ai uses AI to assist with detection, diagnosis, and remediation of network issues automatically to create a self-healing network. Shown in the video is a demo of AI detecting an outage and automatically applying fixes to bring it back online.
Presented by John Capobianco, Product Marketing Evangelist. Recorded live at Networking Field Day 37 in San Francisco, CA on March 20, 2025. Watch the entire presentation at https://techfieldday.com/appearance/selector-ai-presents-at-networking-field-day-37/ or visit https://techfieldday.com/event/nfd37/ or https://www.Selector.ai/ for more information.
Transcript
So a little bit on, well, a significant amount about the root cause analysis and AI agents. I don't wanna keep hammering the point about needles and stacks of needles, but these are some, this is some research we did. The average network engineers looking at 38 dashboards, the average infrastructure engineers looking at 15 dashboards and applications are up almost 50 dashboards.
And if you think of a full stack developer, now you're looking at close to 75 dashboards, right? Too many tools. Too many dashboards, too many alerts.
I call these panes of glass. No, I call them glasses of pain, right? These are glasses of pain.
The other thing is, is the incident response challenge. 46 alerts from a fiber cut going to 12 different teams, all with different messages, all with different dashboards, all with different silos. So we have the P one war room or the WebEx call the zoom meeting, bring everyone together and what's going on?
What's the root cause? What is, what are your alerts? Say, here's what my alerts say.
How do we do it? Right? These are the three critical challenges, bringing in all the heterogeneous data into one data lake.
Once we have a standardized data lake, we can then do normalization and named entity recognition to, to find anomalies. This is exactly what a human does, by the way. All the data, throw out the good stuff.
'cause I don't care what's working. I care about what's not working. Try to find the anomalies and come up with the correlations and ultimately a root cause.
That root cause then filters over to the collaboration. Now this is bi-directional, as you're gonna see. This is what a smart alert looks like that you would receive on your phone.
So you've seen me on my phone talking to the network in one direction. This is the network talking to us in the other direction. Smart alert right away.
BGP down established A down on these five devices and we've detected a config change on this device by this user. Click on this link to go to the portal. Here's the 58 related events.
One ticket, not 60 tickets, and it's in metro, Kansas City and America, right? So this is the idea of the smart alert context, actionable items, real insights. And so, you know, not to troubleshoot OSPF.
This is a BGP problem, even though it caused OSPF downstream right now. You can click on the portal, which would take you to this view. Those are the devices involved in this outage.
We hop into the copilot and start asking questions. What's the probable root cause of the event at NI at 3:28 PM Now I just want to, where did I go? Sorry.
Um, I just wanna highlight that when we talk about operational twin and the past, present and future, this is looking back at 3 28, right? And I don't have to go into the SIS log dump and find the log from that 10 minute window. I just ask, what is the probable cause at this time?
So there's the two answers, the visual view and the text view. And since we know it's this change by Martin, maybe we should get Martin on the phone, right? And find out what he changed.
No, we don't need to show me all the config changes by Martin, right? Or in the last hour. Boom.
There you go. Here's what Martin did based on the TAC acts, and you're gonna see we're getting shut down. Command shut down command was executed multiple times.
Martin shut down a port, it led to BGP, which led to 60 events downstream, no noise. Identify probable root cause, accelerate incident resolution through natural language. Now about the trust in moving on, closing the loop.
What you see here is a ticket filed with ServiceNow, fully ally. And you wouldn't know if you read that ticket that a human didn't open the ticket and fill in the details. We have the root cause.
Go ahead and open up the ticket and ITSM start, start managing this incident properly with, with tracing and logs and visibility. So we get the full ticket, the symptoms, the notes, related records, all completely done by an AI agent. What about sending emails automatically?
So this is an agent that's saying details from provider maintenance, email planned outage notification in order to maintain the highest levels of availability, blah, blah, blah. You get this agent sending emails, letting people know your customers know, or your organization know about potential disruptions. This one I thought was really cool.
What if I put a calendar invite right into your calendar about the outage on Friday night coming up? AI agent can do that, right? Right in your calendar.
Hey, look, there's an invite Friday night, there's a maintenance window. Um, going back to the example you were doing where you're using the L LM feature and asking, showing me all devices, what happens if someone's, you know, not really confident in the output and query again, like reassured, did you miss anything? What, how does the system respond, Um, To show me its Work?
Yeah, so we do have a show you the work, uh, capability to actually take the query and go to the query dashboard. So we have ways of interrogating and, and, uh, showing your work or enforcing the trust with almost like cookies throughout the system. So when I give you the natural language answer, you're actually able to go to the SQL query and validate the metrics and the logs that it used.
The raw data is available behind the scenes, and that really helps reinforce trust because when a network engineer does this the first time, they can't believe it and they want to be shown the evidence of that truth over time. They need less and less evidence and they start to build more and more confidence in the answers. What about other tech domains?
So security alerting information, optical layer, compute, you know, Datadog, will you talk about that at all this morning? Um, no, but it's a good question though. To the point of I have seen root cause analysis suggests that it's memory on a host.
Mm-hmm. And you see the whole link through the show, your work, right? Of the knowledge graph, right?
I have experiencing traffic drops and you, it almost eliminates the network. It says actually, but it's a, it's a resource issue on this node, in this rack, right? Optical layer, we can monitor that in terms of predictability of the future of your network.
We can see trends and temperatures, for example, on optics and say in eight days you're gonna be in trouble on optic four because it's getting hot, right? Yeah. Um, in terms of security, I would say tertiary.
I I wouldn't say that we're a security focused company, that would be misleading. However, there, there are natural correlations that can come out of that, uh, natural language that are semi-related to security. I, I, you know, on the one hand I think there's great value in product focus.
I also think those guardrails can be artificial. And if you think about things as a whole system and start to take all the telemetry reporting, um, alerting that I can get, you might be able to do much better at seeing the big picture of what's happening with an application running on all sorts of infrastructure. Right.
No, good point. Good point. And we know that there's firewalls in the path, there's load balancers in the path, right?
So we have to get the visibility from those devices as well. Sure. Yep.
So what about, um, using external sources for, uh, some of your troubleshooting? For instance, a very common, uh, thing that a network engineer would look at is how is my route being propagated through the internet? Um, I, you know, is, and a lot of that is interrogating looking glasses and things like that, right?
Um, can this handle going out and fetching external source data to Yes. So, so I'll use an example of let's say NetBox or Autobot as an external source. We bring in all that data to augment the metadata around.
It's not just an IP address, but it's an IP address in San Diego, right? So we bring that metadata in, we are able to ingest from rest API systems and external systems. I, I don't know of anyone who's integrated looking glass, for example, with our tool.
Um, but our tool's flexible enough that we can ingest, you know, any data source at all into the system and enhance and augment the correlations. But I don't know that that would help with maybe root cause. When I get to the, the path tracing in a little bit, uh, you'll see how we can pa uh, trace paths through the network out to the internet.
So like I have a vendor service like DDoS mitigation, right? Like DDoS protection provider. We'll, tell me if I'm, if I'm under mitigation, right?
Right. Um, but that is not an an internal, I mean, usually a network engineer would have to go to the portal, right? Log in, look at the mitigation right status.
Now if that, that service you mentioned has arrest API, if they did, then we could tap into the rest API and ingest that data and let you know that something's been blacklisted or not whitelisted or whatever. As long as there's a rest API, we we're happy to incorporate the data. And could it use that to predict a or tell me what my traffic, why I'm seeing a traffic drop?
Or is my traffic It, it would To another edge For it? It's a little theoretical, but in theory it should actually make those correlations to say, I'm bringing this data in. This is why we're dropping traffic because of the DDoS mitigation.
Okay. It, it's, it's an exercise on paper to try. It would be an interesting use case.
Okay.