A Day in the Life of the Enterprise Administrator with Selector AI
Selector AI’s presentation at Cloud Field Day 22, delivered by VP of Product Sachin Natu, focused on illustrating a typical day for a cloud engineer managing Fortune 500 company environments. Natu highlighted the significant challenges these engineers face, emphasizing the complexity arising from the multitude of technologies, administrative domains, and siloed views across on-premise networks, multiple ISPs, and hybrid cloud deployments. The core problem showcased was the difficulty in troubleshooting even a seemingly simple issue like a malfunctioning chatbot, requiring investigation across numerous interconnected systems and data sources.
The presentation centered around Selector AI’s platform, which aims to address these complexities by providing a Slack-native AI-powered interface. Using a scenario involving a chatbot outage, Natu demonstrated how the platform proactively identifies problems, correlates events across various systems (including network devices, ISPs, and cloud services), and provides actionable recommendations within the Slack workspace itself. The platform doesn’t simply identify issues; it also offers contextual information, such as historical data and user-added notes, to build a clear picture of the situation, allowing for more efficient and informed decision-making.
Crucially, Natu differentiated Selector AI’s approach from generative AI, emphasizing the platform’s reliance on machine learning techniques for accurate data analysis and correlation, rather than generating hypothetical solutions. While leveraging Large Language Models (LLMs) for natural language interaction, the core of the platform uses precise data analysis to ensure accurate representations of the systems being monitored. This approach addresses concerns about the reliability of AI-driven insights in critical infrastructure management and ensures that the recommendations provided are grounded in factual data and historical context. The presentation concluded with a planned deeper dive into the platform’s architecture and data processing methods.
Presented by Sachin Natu, VP of Product and Technical Marketing, Selector AI. Recorded live in Santa Clara, California on February 19, 2025 as part of Cloud Field Day 22. Watch the entire presentation at https://techfieldday.com/appearance/selector-ai-presents-at-cloud-field-day-22/, https://TechFieldDay.com/event/cfd22/ or visit selector.ai for more information.
Transcript
My name is Sachin Na. I'm the VP of product management, and in this particular section we'll be going over what does a day in the life looks like when somebody has a select AI platform, okay? And what you'll understand is what, how do we understand various types of data?
How do we correlate it and how do we present it to the end users to make it very easy for them to understand? And then later on after the demo, we'll do an actual reveal of what the technology underneath that, right? So it's going to be a two stage presentation where I'll be presenting twice.
So if you look at today, right? Today I'm going to latch on to a simple question where somebody says, my chatbot is not working. And a pretty, pretty standard application in an enterprise space, right?
What are you looking at? You are looking at bunch of your offices, remote offices, head offices, whatever they could be. Then you are looking at one or two ISPs that you are essentially working with.
Then you're going to the cloud vendors. And the cloud vendors. Essentially, you always have multiple clouds, right?
For two reasons, right? One, you don't want the vendor lockin. And the second one is a lot of cloud vendors offer very unique capabilities that if you want to build a best in class product, you actually, you need to use multiple clouds, right?
So in this particular scenario or the use case that you'll walk through, the question is very simple. Hey, my chat chatbot is not working. The, the journey of the cloud or journey of the packet is going to start from your enterprise network is going to go over your cloud on-ramp, whether it could be any ISP.
The front end of this application is in AWS, and the back backend of this application is in GCP very standard type of network deployment. This, this example is very, very simple, but now what makes it so complicated, it's really finding needle in the stack of needles, right? It's not your needle in the haystack where needle at least shines here.
There's a large bunch of data and which looks very similar. So let's look at, at the top level, if I want to justify what is driving this complexity is you are dealing with different types of multitude of different technologies, which are built across generations in last 30 years and multiple administrative domains, and each of them has their own siloed view. So the problem statement that was saying, what, Hey, my chat bot is not working.
Who would you even give it the give the problem to? Would you give it to your networking team? Would you give it to your ISP or would you give it to the cloud guys?
And this is where the complexity comes in. Let's just start with your corporate office. If you're an enterprise admin, this is the best place you have control over.
What you're dealing with here is you're dealing with whole bunch of, uh, servers and switches and, uh, access points. You are dealing with vendors like your traditional vendors like your Ciscos and Arubas and Juniper, right? And what are the technologies that you're dealing with them?
You are talking to them or SNMP gets and get next, or SNMP tracks or bunch of CIS logs or nets. So these are the technologies that you're being using to drive a network. Now, this traffic or the packet goes into, say, an ISP, what you're dealing with here, right?
In the ISP, you don't now now control the control the packet. The packet is in somebody else's admin domain. So what does, what does that entail to you?
You need to know what is the path, what is the length of the path? Because that figures out what the latency is and what is the est point of the path based on whatever the routing table is set up. And a lot of operators give you that type of a view.
Then again, right? You are dealing with different types of entities, whether it's say at t or Verizon, Comcast, anybody. Then what are the technologies that you're looking at?
You're looking at things like looking glass. Who's my, what is my egress point? Or you're looking at things like service ping.
You're looking at service trace route. And more critically, right? A lot of times service ping and service race route just tell you whether the IP connectivity exists.
It doesn't tell you how your application behaves. So these were the synthetics, and actually we have our own synthetics, which run at HTTP or your application level. So we can tell you, hey, this is, this is a particular nuance of this chat bot application, which is being mimicked on the synthetics, right?
And that's a very key thing that how is my application running over, uh, ISP where I have no control over? So I'm really looking into the performance of ISP. Now, what's the last part?
Right now the packet went all the way into the cloud. Now what are the entities that you're looking at, right? You're looking at entities like Cloud Router or the VPCs, or Hey, what is my, uh, AI engine, or what is my vm?
What is my database? What is my, uh, A-P-I-A-P-I gateway, right? There?
Again, you are dealing with all different types of vendors and you're dealing with different types of technologies like your CloudWatch, or these are all API based technologies. But even if you look at the API, there could be rest, API, there could be webhooks, there could be web sockets, right? And depending on what information that you're pulling in, you are talking different method.
And so now that one innocent question that, hey, my chatbot is not working, there could be a problem anywhere in this place, and the end user bot end user sees is what my chatbot is not working. And again, right? Our goal really as Dev was saying, is to be proactive.
So we want to be filing a ticket and solving the problem or helping our customers to solve the problem before somebody even notices that, right? And so for the, for the demo as well as for the, like the actual day in the life, I'm going to take two scenarios where in one case, in the first case, the issue is on your cloud, on ramp site where the packet is dropped because of a flaky link on your way to the, on the, on the way to the cloud, okay? And in both these cases, I have chosen the gray failures where the links are flaky or the CPU is high utilized, and why the gray failures are even more difficult to debug, right?
If something is down, it's really cut and dry. If the link is down, you know what to do. But if the link is dropping a lot of traffic, that's where the challenge becomes kind of difficult, right?
I mean, for this person application is working fine, or this person application is not working fine, right? For this location, things are great. In this location, things are not good.
So we'll cover that. So again, right, as Deborah said, we are, uh, we are a Slack native ai, slack native interface. What that means is our, and this day and age, right?
It is 2025, everybody is on Slack, everybody's on teams, and you see a problem when you generally least expect it, right? So I'm going to, in most of the times, are going to see the first alert on your mobile, say, when you're having a dinner with your family. And so what do you need to know?
In this case, you need to see, hey, there is a alert from my application. This alert says the ISP. Verizon is experiencing high number of packet drops, by the way, right?
We don't want to be picking up on Verizon. Verizon and at t both are, are the names of the customers I'm going to use just to make it more relatable. Uh, both are our customers and like both, both work extremely well.
But this is just, we picked a fine cost to figure out where the, where the challenge challenges. So again, right? Coming back, so the, the link, the link going towards Verizon is experiencing high packet drops.
The alert also told you the specific affected circuit, okay? The alert told you the affected site and the region. And alert also told you what is the alternate remediation, right?
So for an admin, you got a proactive message, it identified what the problem was, it suggested a recommended solution. And on top of it, a lot of times these type of alerts show up into multiple different events. Say for example, hey, if a link is dropping traffic, a lot of times, hey, it also affects the BGP, right?
So BGP also starts to send events. So we are essentially to, uh, giving you an event summary of all the events that are correlated with this particular flappy link. Okay?
In your first message, in your first Slack message, you almost, you have the entire information that you need to act and fix a problem, ideally before anybody else sees it. Now again, right? What is a curious person would do, you would type on the slack using plain English language that hey, you know, show me what's happening in my network and maybe I'll give a specific, specific site, tell me what's happening in the San Diego site.
So then now the chat bot and the system came back with a very specific circuit diagram almost, where it's showing, hey, you have two edge, two CP routers connected to Verizon and at t and the link towards Verizon is flaky. And you can see that. And whereas link towards at t is, is working fine, and you can continue this interrogation on Slack on your mobile forever, right?
But now I'm just going to go to the dashboard view just because I can use the real estate little bit better. I wanna, yeah. Can I ask a question real Quick?
Yes, please. And I may have missed something, but I, I want to go back to what you, you, you guys said in the beginning was, our goal is to come in and integrate with existing tools. Yes.
Ultimately, at some point, taking over those tools, unless I'm missing something. This is very network focused, which is a big gap today as far as observability in many different ways. Will we get to the point where we actually see something?
I know Michael asked like a PM, we don't have traces and logs, so we know there's a gap there. No, no, no. We do actually have traces and logs, and I'll cover that.
Okay? Okay. And I'll, I'll save some, I'll save.
So let's finish through this and then I'll probably have more. Yeah. And again, right?
I chose two different failures. One is in the access side towards the cloud and one is in the cloud. And you'll see how these all come together, okay?
So again, I told you that there is a problem, but if you want to go and fix it, I also want to prove it to you that there indeed is a problem. So this is where you can interact with, again, same plain English and say, you know, show me the packet drops on that particular link. And as you can see, it's actually showing the packet drops.
And again, right? This is where that gray failure matters. The packet drops are inconsistent, right?
And how does an inconsistent packet drops look like? So this is a, uh, idea at the link level, how does an application behave, right? Suddenly all your TCP applications start to become flaky because the packet drops, right?
So latencies and all start to increase, right? So now based on this deep drive, and we can go even a lot further based on this deep drive, you are convinced that there indeed is a problem. And I want to take the solution Question about the, the graphics that we're seeing in this and the previous slide.
Are these being dynamically generated by They're real time? Yes. Okay.
How accurate is the information in the graph? A hundred percent accurate. Mm, That's, that is a old claim and has lied to me many times about things.
I was, this is a fantastic go ahead, fantastic question. Let's go into the deep dive on that. So again, um, I, I'll cover that in the deep dive, but again, right, what you'll see here, right, what Deb was presenting, I literally loving this question, right there is a, we are looking at the real data, right?
We are not generating, this is not a generative ai, this is the accurate AI based on standard machine learning techniques like, like sampling with simple linear regressions, log mining, entity recognitions or clustering, right? And since you are acting on this data, this is accurate, and this is where we actually show you. And in an AI world or AI ops world, we cannot be hallucinating responses, right?
Right. Otherwise nobody will use us. I I think it's very important that you pointed out, out that this is not generative ai.
Yes. That seems to be the base assumption. Now, when someone says AI is that it's generative, I will cover that.
There is a, again, right? You have to use the right technology for the right problem statement. Mm-hmm.
And we do use generative AI mainly to interact with the customer, right? That's, so think about this. Go to any gen model, ask them to create a picture, they cannot, right?
Whereas this system is going to look at specific circuit errors and is going to actually look into the database, say the Prometheus database, where all the metrics are kept, and it's going to show data from here, Okay? Right? So again, this is not gen i, no hallucination here.
This data is created by looking at the actual data in the system. And when we get to the demo, uh, we'll show you the, how the data is generated. Maybe once, um, in another, uh, few minutes, we'll show the exact demo that we have to show how it is.
Okay? I, I'm loving this question, right? We'll go into the deep dive, right?
Because hey, these are standard stuff, right? Can generative AI solve all my problems? We use generative AI for the right place in the right, in the tool, right?
So then similarly that there is a circuit delay. Now we made a recommendation, right? Saying, you know, you should be moving to at TI want to prove it to you that this even recommendation is accurate.
And so through our ana, through our synthetics or even simple packet drops and circuit things, we actually know that in last 24 hours, Verizon has been not, uh, it has been, uh, has some issues about the availability, whereas at and t has fantastic availability. So moving there is safe, right? Again, this data is not generated, it's imputed, it's calculated based on actual packet loss, actual pings and trace routes and all that.
Yeah. Even basically temperature set to zero. Like there can be no, we're not guessing on Anything.
There is, again, this is machine learning, not generative. Yeah. Please understand this.
So To that point, there's obviously, there's a lot of confusion because as you say, it's synonymous these days in most people's mind that AI is gen ai. And then you also talk about AI ops. How familiar is that to your customers as you have SEC ops, DevSecOps, you have GI ops, you've got ML ops, so why AI ops, not ML ops.
And I'm just curious how you communicate what you do and, and what's made of it. I'm glad you asked that question because I was going there next. So blame Gartner that they're the ones that came with the bloody term.
So again, Right, if you look at AI and ML are pretty much coming together, and as I said, right in our technology, we use ML for a lot of core functionality and we use the generative ai, the LLS for communicating in English with the rest of the users. So if I can take that question, uh, quickly, uh, the ML ops and AI ops differentiate here is that in case of ML ops, we are looking at, they're looking at the pipeline, which is being used for machine learning. But here, if when we are monitoring the infrastructure, which is used, uh, the CPUs, GPUs, uh, network infrastructure, there comes the AI ops.
So they kind of both coexist in, uh, AI ops and ML ops. I think by larger questions, just general customer confusion. If you're selling to enterprise admins and you talk about ML ops or AIOps, does that resonate with them or do they get kind of a, so Again, for us, ML ops or AIOps is a method, right?
What are we monitoring? We are monitoring your cloud infrastructure and uh, network infrastructure and together, right? And we are using the tools of ai, whether it's machine learning tools or the AI tools to, to give you the results that you're looking for, Which makes sense.
I think that's a great way of explaining. I think when you get into AIOps and ML ops, it'll confuse more than it'll, yeah, It's the one thing when we talk a lot of customers to them, AI is generative ai, it's chat GBT and a lot in the last two years, right? And so, um, that's why we leverage that obviously that technology, but AI ops isn't just chat GBT because chat, bt what's the status of the network in San Diego?
They have no idea. Um, so it's a combination, right? But I think for us, because we have AI in our, in our name and all that, we, we've obviously kind of distinguished between, you know, what we're doing in the pipeline side versus like the piece, but people kind of think of as AI today, to your point.
Well, I, I want to, I want to just hit on this just real quick 'cause I'm glad that that question was asked because I deal with this stuff all the time. Getting these conversations, AI ops is completely blown up to confusion. Most customers are not even doing automation, so they can't even get to that next layer of AI ops.
So I guess my bigger question, and I think we could move on, is, can I take the data that you currently have? And I think Eric actually hit on this a bit ago. Can I put it on a Kafka bus and take action On it?
Absolutely, yes. Okay. Yes, yes.
Okay. Yes. Carry on.
Yeah. And Al, when I do my demo, I'm gonna, this is a funny, our chief data scientist, you know, he, he presents a lot to like, what is the ml AIOP selector? And his last slide said, this is my ai.
And it's like I have a thousand alerts over here and I have 10 correlated incidents here. Like this is AIOps. 'cause I'm taking all the human, um, manual effort of correlating events and the needles and saying these needles together are an event that to him is AIOps.
And that's how I think about it When I present to customers, ultimately it's it's saving your teams and time, money or time team money and time. So again, right, he has like our CCDO has a fantastic slide which talks about, you know, what, what does AI mean to you, right? And it's really I of the beholder, if you will, right?
And typically of the hierarchy, you go, AI only means gen ai, right? But for a lot of network operators, as you are asking, hey, you need to gimme the accurate questions or accurate answers, alright, you wanna just top The demo, so maybe we can do the demo and then if No, just please. Okay, so again, right now I'm taking a very different again, right?
We are talking about cloud. So the first thing that we talked about is the circuit failed. Now we are going to deal with a VM failing inside a cloud.
So same here, right? Similar alert notification, completely different set of like obviously right? The right relevant set of details here, but which location the VM failed, what is the primary, uh, suspect?
What is the secondary thing that we are seeing, right? The likely cause is high utilization. And here the recommended solution is essentially pointing your users in affected regions here in EU towards a new, new place like say North America.
Okay, hold onto this, right? Because this is not something that somebody can just understand based on an event. There is a higher level of intelligence that is required here, right?
And I'll show how we draw that higher level of intelligence. Okay? So, and again, similar there are latency events.
When I ask a question on a chat bot or the Slack chat bot on Slack, this is the answer it gives. It says, you know what this CPU is, right? As well as this DCI link the application API that are happening over at DCI, they have gone red.
Now again, similar thing while that interaction is happening, the built-in agent, it's really an API called agent is a fancy name for the API call. So the builtin agent files an ITSM ticket, right? What does that ITSM ticket talk include, right?
Deva talked about reducing the number of incidents that one ITSM ticket has all the related events in issue about all the related events. That means all of a sudden there is one ticket filed with all the related events. You are not like handling hundreds of incidents.
So there is a summarization of bunch of events into one incident that you need to act on. Okay? This obviously becomes a system of records and the last bullet is super important on this system of records.
You can start to put your resolution notes, comments and stuff like that. And that data is further used to provide gen provide you the recommendation for the next time, right? In this example, what happened, right?
The CPU is down, the DCI link is not, not reactive, right? It's flaky. Nobody where it says that you're not, you should be pointing DNS somewhere else.
But that intelligence is derived based on the previous notes from, from the operator that you know, the first action is to point point users to the working place. And then you can figure out exactly what happened with the CPU. Because changing A CPU and all that stuff might take some time.
You don't want that downtime changing DNS is a matter of one minute, right? So that intelligence did not come from pure events, that intelligence come from this incident intelligence. Can you actually add manual context to something like this where you could say like, well it looks like it was an ISP, it turned out it was somebody drove into the, the CO and, and that wiped out my line.
Because going back then over previous incidents, like does that have an opportunity to impact the correlative effect so that you can say, Hey, we've keep adding manual notes or some kind of No, you can, you can keep on adding notes as much as you want and we will use those notes to learn from the incident perspective, right? The correlation part is like a lot of times it's just based on what you're seeing in the network, right? But then there are almost two levels of intelligence here, right?
One that I see in the network and the one that I see from what users input and we, we, we process both, Right? So now, now just do a simple again, right? Very similar deep dive here.
I found that the server in the, the server in Ireland airline, the EU west has gone like is not the CP is not handling traffic. The CPO has become saturated and because of which our synthetics, and this is what I talked about, right? Our synthetics can mimic how a particular API looks like it's not just ping and trace route, right?
So that will figure out, hey, this particular API is not going to work because of the, the CPU is saturated. It's not handling, uh, calls fast enough. Yes.
You, you had me at uh, synthetics, uh, is this like a rum type of synthetic you're actually dealing with like application level interactions or is it protocol level that you have to define See To some it's a, it's really a small Python program. So it could be protocol, it could be application. So whatever we can write.
Okay, great. Thanks. One thing maybe to add is, you know, there are leaders, you know, there's, there are people selling to the synthetics out their thousand I that, you know, we compete with what we also will integrate with.
Yeah. But what they don't do is like, you can't go to Cisco and say I want a synthetic to do these five things that are just for my organization, right? So we've kind of, our way to kinda compete is to build, take that, that synthetics we have but layer very custom integrations on for your organization or yours that might be unique just to you.
And then go and kind of render that with using ml. So we compete, but we also take any synthetic somebody has because it's a source of data and it's, you know, it's great okay to hop in there. So again, right?
Don't know to go, but again, right here also building confidence in the recommendation. So if I'm saying, hey, you know, you should point your EMEA users to the America America data center, that latency has to be correct. And then this is where we are showing that a different graph of how these events are correlated, right?
Different events happened in this entire failure case and we are showing all the different event correlated, and I'll talk into the deed session. There is a event and there is an event association. So based on that, we know what is the causation of bunch of events, right?
What event causes other events and what is that relationship, right? Typically a VM down causes an application down, application down does not call VM down, right? So we understand that type of, uh, correlations.
Just a a question on that. Sorry, I I I don't mean to pull this, I know we'll get to demo. When you have that happen, does it give you more than just the single like estimated top causation or do you have a chance to look at like, here's like the top, there's the 99, the 98, and the 45 and you can say, I may know that's actually the 45 because I know something just happened A a absolutely right.
And so again, right, even in that correlation, we actually have a slide bar, which shows me, right? You know how much level of correlations you want, right? And so that slide bar lets you that confidence interval.
Fantastic. Thanks. And to quickly answer your question, yes, we do print out all the possible, uh, causes and then if there are top two of them, then at that point, a human has to look at that okay.
Saying that, okay, these are the top two and I have to look at this particular site or this device. They get the enough context to work on that. And then the additional stuff you can do with notes and can actually enrich that for future stuff.
Yes. Got it, Got it. Correct.
And again, right, so now the last aspect, right? You, you solve the problem immediately by pointing the DNS, but you also want to show, hey, why is that CPU became saturated? Then you can go back here we are going back two days and really finding that, hey, you know what, somebody changed the, the spec of the VM from a 64 core to a 32, right?
And that may not show you the effect right away. It may show you the effect after two days, right? Maybe this, this change was made over the weekend, it worked perfectly fine and then suddenly there was a, like the school starting day where suddenly you start to see more number of events and so you can go back and solve the issue properly.
The long-term fix event, Because I keep seeing the, the slides where, you know, it's working inside of the Slack channels. Um, one thing I, I have in mind 'cause I deal, deal with monitoring all the time inside of Slack using chat bot and stuff like that. What is like the timing when you're interacting with, uh, selector AI for it to respond back to you inside of the, It's the, the Slack channel.
The slack queries are real time. And you'll see it in the, yeah, we'll see it in the demo. Okay.