A Deep Dive into AIOps with Selector AI
Selector AI’s presentation at Cloud Field Day 22 offered a technical deep dive into the architecture and functionality of its platform. The core of the platform relies on machine learning (ML) techniques for data ingestion and analysis, ingesting raw data such as metrics, events, and syslogs to understand network behavior without generating content. This ML-driven process focuses on accuracy, utilizing methods like regressions, clustering, and cosine similarities to identify patterns and correlations within the data, avoiding the hallucinations often associated with Large Language Models (LLMs).
The platform’s unique strength is its ability to handle a wide variety of data sources, leveraging a declarative ETL and compiler to easily ingest new data types. This flexibility is showcased by the system’s ability to process data from diverse sources, ranging from legacy network devices and modern cloud services to custom CSV files. The system’s architecture is built on Kubernetes, ensuring horizontal scalability to handle the volume and velocity of data ingested, with a strong focus on creating context through the integration of CMDB data and metadata to give meaning to the raw data. This data integration is a critical component, bringing together disparate data silos to provide a holistic view of network operations.
Generative AI plays a supporting role in Selector AI’s platform, primarily enhancing user experience. It translates natural language queries into the platform’s query language and converts the resulting JSON output back into human-readable English. This use of generative AI is carefully managed to ensure accuracy, acting as a complementary tool to the underlying ML engine and not replacing it. The platform is designed to be flexible, allowing customers to utilize various LLMs and ensuring that the system learns and adapts to the specific details of each customer’s network infrastructure over time. The company emphasizes its commitment to customer support, providing ongoing platform maintenance and support throughout the entire lifecycle of their product implementation.
Presented by Sachin Natu, VP of Product and Technical Marketing, Selector AI. Recorded live in Santa Clara, California, on February 19, 2025, as part of Cloud Field Day 22. Watch the entire presentation at https://techfieldday.com/appearance/selector-ai-presents-at-cloud-field-day-22/, https://TechFieldDay.com/event/cfd22/ or visit selector.ai for more information.
Transcript
Now let's, the first big question that first came in, right? Is this all LLM or this is all ml? And the real answer is use the right tool for the right job, right?
And what you see here on the left hand side, this is where we are ingesting lots of raw data, whether it's metrics, whether it's events, whether it's, uh, cis logs, stuff like that. And that is we are, we are using ML techniques to figure out exactly what is happening, right? So there is no generative part on this side, right?
Why about the accuracy, right? There is no solution on this side. This is all well understood math, right?
I mean, even this is understood math, but this is all math, like your, like the regressions and clustering and uh, co-sign similarities and stuff like that, right? Because we are getting all this data, analyzing it and putting together things without generating anything, okay? Now, when does we, when do we start to use the generative ai?
Is, is really to offer that easy experience for an operator, right? So rather than somebody typing a query saying slash select, choose a stable and show me the interface errors, right? A lot of those graphs are accurate representation of the data that exists because there is a query that happens.
But translation from that English language to query is a job of generative ai. Or once you get an output, you get a JSON output. Translating that into an English language is a job of generative ai, right?
And that's a very important part, right? We are using generative AI as a combin real tools not a replacement, right? And again, right, as these technologies evolve, we will basically evolve with whatever is best outcome, right?
So again, this is where there's no hallucination. There's a question. This, the, the portion you're showing, you know, three different potential, um, services that have models.
Are you leveraging all of those services or do you have your own internal service? Yeah. So like you can use whatever service that is allowed by your corporation, right?
So internally, we can even use LAMA three, right? So that, that is, again, I'll go into the details. It's an API call.
So it can be replaced with whatever engine that is allowed in your organization. Okay? So, Pete, I was gonna say, you talked about the evolution.
So I assume that, uh, ML was built into this before Gen AI came out before, yeah, yeah. Got GPT. So as you evolve, you took, now you've got copilot, you've got Gini, and you've got got, I assume this disrupt Those, those are just representative, right?
You can use what deep seek if you want. You can use whatever you want. When did you put, uh, AI at the end of your name?
I'm just trying to figure out the, the path. So the, the starting company was based on ai, right? Because Ml, right?
Ml, Sure. Yeah. Well, because I, you're, I'm trying to make, figure out which one.
You don't want to get people confused. You're saying that we've been legitimate? Yeah.
AI in the form of ml. We're not just Johnny come lately to this. We've been doing this for real, correct?
And again, so now we'll go into the actual details of the data flow, right? Because that's super important. So again, right, we talked about, hey, how do we collect data, right?
So this is basically a distributed process. So your, some of your data is going to come from these remote engines, which is your on-prem stuff, understands SMPs and Kafkas and all that. Some of this data is going to come from your cloud.
And this is where using ServiceNow and PagerDuty, and these are where we are logging tickets. So this particular part, each of these collecting devices are also based on Kubernetes. So they are horizontally scalable.
And I also answered the question earlier, even the top cloud is based on Kubernetes, horizontally scalable, right? So that takes care of the volume and the velocity part of the data. Another big problem is the variety.
How do we solve for variety? Over last five years, we have built a whole lot of connectors of the box, right? So there was a question about, hey, S-N-M-P-G-N-M-I and GRPC, right?
SNMP also has like bunch of vendor specific nibs. Standard nips, right? And there are event stuff like Kafkas and Rabbit MQs.
There are standard, uh, like the data stores like your Prometheus and Splunk and all that, right? So these are all over the fires of exper, uh, uh, existence already built in. But what is more critical is how we make this easy to, easy to extend.
So the one realization is there is no universal standard telemetry format. So we built a declarative ETL and we built a compiler, right? To understand your schema and right code to write the parcel to understand your schema, right?
So if you bring me into a new environment and suddenly there is a different type of data source, we can ingest that data source using this compiler in a matter of days. All we are to understand is your, what the syntax is and the declarative a TL. Make sure that the data goes into the system in the standard format.
And if you want to invent a new type, we'll do it right? Very, very flexible, right? The again, right, the fundamental thing is we do not assume anything.
We are built with flexibility. We could stop right there for five hours. This is fantastic.
So I, I think in this one, I, I think there was a question before. Also, I would say I can challenge anyone to give us a type of data source that we cannot ingest. Because so far we have in that we have ingested like legacy network devices, different kinds of, even things like folks have their own CSV files with different kind of data schema.
So I would say any data we throw at us, we'll take it that confidently. I can say that, uh, like a hundred percent ha Has anybody heard of TL one? Right?
So TL one, when I started working TL one as a no protocol, right? So, and it tells you how, how old that is, but we support that because a customer wanted it. Hey, it's a schema.
Like the compiler builds this, builds the password. We can start to ingest it. Yeah, well, so there's two things.
One is of course the, the data that you're getting, and there's also the way in which you're getting the data. So leveraging things like TEL and EBGP where you can get really low overhead, but high, you know, detail. And of course the, the forever battle between SMP and gm.
I like, you know, we hope that the world will get to gm I faster, but it took a long time. S-N-P-S-N-P-G-N-M-I are, those are obvious ones, right? But even like rib, even like BMP, right?
I mean, these are our data sources, right? I mean, that's the real time routing data, right? Real time data from like your cloud providers.
Is there any way for us to gather, like are you tapping also into like the backend provider? So if I'm connecting to Azure through Verizon, where you can actually see like, not just my as boundary, but the actual, yeah. 'cause otherwise it's a black box from as to as until I get to the provider.
And I know that's always a tough gap to cross see Again, right? Whatever you can see, and you can provide me the data, we can take it, right? But this is where, in the example that I showed that Verizon or at t was a black box, right?
And so there we figured out there was something wrong either based on interface gas, which I understand, or even the latencies which are going over the system, right? So like whatever data you can give, we can handle it, right? So now we talked about a lot about data, right?
But data makes sense only in the form of context, right? So what does the context gives you, right? So the context gives you, hey, what does this even mean, right?
So what is the inventory? What is the configuration? What is the location, right?
What is, so like a, a real example, when Hurricane Helene came in, one of our customers wanted to essentially prepare for that thing. And they said, you know what? Tell me all my data centers where Hurricane Helene is going to pass through.
And so we just ingested that data from the Weather channel, right? And we know the light long, we ted the data from Weather Channel, we could tell the customer that, Hey, these are the properties that you should be watching out for. And we are taking this data from everywhere, NetBox and ServiceNow.
Anyone as de said, Hey, if you don't have anything, we'll take your CSV file. And then in addition to everything, this, this, the metadata makes it easy for people to interact with it, right? So there's a circuit id, there is an interface name that does not mean anything to a end user.
So when you are giving that conversational interface, hey, the connection to Verizon in San Diego makes a sense and not that particular what that IP address is or what that interface name is. So I'm going a little bit faster because this is really the money slide, if you will. And that gives you the, the accuracy of the data and how it, how we process it.
So the data we talked about, it's coming from all these different places. It got a lot of context with CMDB and metadata. So this is the first time all data from all different silos come together in one place, right?
And so this is where we really start to get to work, right? So far it is all about data ingestion. This, this data flow is what, where we get to what is the first thing we do, right?
The metrics are a bunch of numbers, right? So 25% CPU utilization, 23%, 29%, just the numbers. They become interesting when suddenly whatever is 20, 25 to 30% becomes an 80%.
That's an event That is something interesting that you need to understand, right? So we baseline these numbers to understand what is normal and what is not normal. And we do this for every metric we get like the millions of metrics baseline number one, so that we immediately know what is not normal.
Similarly, the other big concept is log mining. What is log? Log is a fragmented English statement written by the developers for their own use.
What do they have? They have IP addresses, they have process names, they have interface names, and they have bunch of things that only a developer understands. And so nobody can really use that to, like, nobody can even parse it.
So this is where we start to use another machine learning stuff called named recognition, right? Once you feed system thousands of logs, it knows what all different types of IP addresses are or what all different types of interface names are, right? So now you've got the named entity recognition.
Then you run other algorithms like clustering algorithms, right? So using clustering algorithms, you know what logs kind of believe behave together and what causes an event. And that is all enriched with the metadata, right?
So at the end of these two big processes, and again, right? These processes are not generative. These processes are working on real data to find real correlations.
Okay? That's an important part. Accuracy, You mentioned normal, which is always one of the hardest words to use when we describe this.
Is it normal in the sense of like 95th percentile, 99th percentile? Is it relative to like, hey, it's bad, but it's okay to be bad That that's the slider we even showed you, right? Right.
So you can decide what percentile you want. So this is again, right? Once you do these big, uh, operations like key operations, log mining and uh, baselining, you start to generate those smart events.
As you can see, a lot of these events, the blue ones are perfectly fine events, right? There's nothing to worry about. They are okay events, but then these are the anomalies.
The red ones are the anomalies. This is what we do by losing first step is temporal, like look at the particular period of data. And this period of data is configurable, but particular period of data where you say, you know, Hey, these are all the anomalies that I saw, okay, in the previous example, the VM and the link down, it's likely it it happened in the same 10 minutes, right?
How do I know they're different? So this is where we do another actual statistical machine learning technique called all your cosign similarities and clustering, right? And so with those algorithms, we figure out what are this correlations?
A lot of times the first correlations are undirected based on again, right? This event depends on this event. And once you learn the patterns, you can even find the causality On this.
Do you have a chance to actually tell it to stop watching things like somebody do you with like zero events planned maintenance? Yes. Like, and also even when there's live and there's some network outage, how do you do there is do you deal with zero data that's feeding in but not affecting the inference?
The the beauty is we can even get an get an, we can process an email that is coming from your provider saying there is going to be a maintenance outage. And in that time we suppress those alerts. So we do not produce fault alerts where somebody is actually doing maintenance, right?
We do that. Excellent. And then it goes to the correlation a little bit thing about event correlation, right?
This is pretty standard stuff in the interest of time. This is exactly what say Netflix or what your Instagram works, right? Ev every event has a bunch of properties which are annotated and all that.
And any event that has more properties together is strongly correlated. And any event that has less properties together is weekly correlated. And to answer your question, based on the confidence interval that you want, and based on the type of correlations you want to want, there is actually a slider which allows you to see different types of events And then the manual context.
And like, you could basically give it a sort of a bayesian push towards a, a potential thing that will then be defined where's the data stored about the customer specific enrichment. All this data stored in either MongoDB or, and again, that in terms of where it is stored, it's a single tenant system. The store data is stored wherever that, uh, it could be on-prem or it could be in your cloud instance, or it could be in our cloud instance.
Nice. And this is a single tenant, so this is your data. Nobody else sees it.
Fantastic. Thank you. Now again, right?
The big question in the, in the room, where is ML and where is ai? So this again, right? We are worried like this is AI ops, this is ac has to be accurate.
So this, everything that you see is that ML part which gives you the output, which is a smart engineer can generate. So this is a human level intelligence generated from raw data using this machine learning algorithms. Now, how do we make it easy for people to work with the system?
This is where that generative AI comes in, right? Some of that, you already saw that, hey, the root cause analysis, right? Like changing the DNS, a lot of that is not coming from this correlation that is coming from the other places on top of it.
Like now you take any base LLM, which understands English, we add our specific query language details on that. That's the blue circle that translates that English sentence into a query to make it sure it's accurate. And then the last circle that you see is your specific data.
So the, the black box could be any engine that you want, right? It's like, it could be any LLM that is approved for you to use. On top of it, we come up with a blue circle, which is tuned for fine tuned for our application, the networking application.
And then the right circle that you're seeing is your actual locations. It's your actual IP addresses and learning about your system. So let's say I'm taking a base, uh, model, right?
Whether it's GT four, whether it's lama, whatever, what I'm doing with this essentially is I'm all my data's being collected on the backend. You are creating a rag or fine tuning for forecasting or whatever you're doing. Uh, and then that data's coming out in, in real time, right?
But I I, I'm sorry I keep getting hung up on this real time thing because, uh, ingesting data, creating a rag around that data, depending on where that all taking place, it takes a little bit of time. Fine tuning takes even longer. That's why it's typically used for forecasting.
But I, I think I'm just getting hung up on the real time thing. Sure. Let me show you this, right?
This data and the blue circle happens in our, our systems. It's not for the inference time. Mm-hmm.
Right? So the fine tuning does not happen in your network to begin with. The only, that final red part, which is about your data that's learning your, your IP addresses and your systems, right?
So that would happen in first few hours of the deployment. Hmm. Right?
And that's a very thin layer, right? The English and the selector language already understands the notion of an IP address. All they need to understand what is your IP address?
That's a very thin layer, But, but it would differ based on the base model that you're using. No, Let, let me, let me, we should just have a, we can table it, but this is an important part. Yeah.
So what this here is doing, you asked me the question in this plain English, okay, I found based on LM three entities here, right? There is a circuit, it could be a port, it could be an interface, it could be a circuit, you know, this is an error I have error table we need to look at, right? And I know where the condition where I'm looking at is not in Chicago, not in Denver, but in San Diego.
So that becomes a specific query into our database and that runs through this SQL and provides that answer. Okay? So this part of LLM is English translation is not generating data.
And again, right, this is where this could be real time because he is doing very little work. He's just doing the entity recognition. It's always tricky 'cause real time, especially when we think of database.
So we think of like subsecond response as real time and it's tricky. I know the tricky real time's a as a good marketing term and it's, you get, you get run over the coals every time about it. One other question about dynamic changes.
So an application, we'll call it sunrise, it's a containerized application, seven pods, whole lot of craziness going on, lots of containers, but it's constantly rebuilding itself, right? We're doing fresh deploys every three days I wanna monitor the application, but all of the entities underneath it are constantly getting swapped out. How do you track in the DVR format?
Let you know the POD may be the same label, but everything underneath it is completely different every three days. And how do you maintain that along with the correlation? So can we take that?
Absolutely. Yeah, yeah. Already it's a tough one.
No, I take that answer like again, right? It's just a matter of going with the last light and again, right? So now talking about the real time, right?
The system figured out all these correlations, it found a JSO format. We have already trained our systems to understand, uh, ml and now it is also fine tuned with understanding your IP addresses and your locations and newer VMs. So now we again invoke LLM to, to convert that into a human readable format, right?
So it is an inference level interaction. That training is not happening. At least LLM training is not happening on a every basis.
It's lightweight. Understanding the entities when that question is coming from English to the query. And once the response comes back, understanding the JSON and converting into an human readable English, that's where the LLM is An odd question.
Maybe, uh, I have my ops team are in Brazil, can they speak to it in Portuguese? Yes. Yes.
Right? So that, uh, we always, we do naturally always lend, move towards English. It's a North American, you know, arrogance, but obviously I wanna make sure that the rest of the world is represented.
Yeah. Yes, yes. It again, like we, we showed in one of the customers they wanted to say, Hey, talk to me in, in the pirate language.
And it started to talk in pirate language, right? So those are the properties of the LLN, right? The core platform is in English.
Okay. So this is the main last slide. And again, hopefully it answered all the questions.
And because again, right, the LLM is doing relatively small work. It's really real time as you saw in the demo. So again, right, this is a summary.
This is a cloud native platform meant for the infrastructure and network operations. So we are using AI or ML as a horizontal technology, but with our expertise, we are focusing on these two verticals and providing you the accurate results. Right?
I think just, I'll stop At that. Just to close it, that last point is the most important one. So our product comes with where we help the customer build the solution.
It's not that we just sell the product and we don't touch it. We onboard the PLA platform, we maintain the platform, and we stay with the complete lifecycle of the platform. That's the key difference for which we are winning a lot of business.