Customer Use Cases and Product Demonstration with Selector AI
Selector AI’s presentation at Cloud Field Day 22 showcased real customer use cases demonstrating how its platform aids operators and management teams in making critical cloud-based business decisions. John Heintz, Global Systems Engineering Director, highlighted the platform’s ability to resolve critical issues rapidly, emphasizing its unique position in meeting the daily needs of operators and management. The presentation featured a recorded demo showcasing the platform’s functionality, including the creation of “smart tickets” that automatically summarize events, provide context, and suggest remediation actions. The demo further illustrated how these tickets integrate with collaboration tools like Slack, offering a streamlined workflow for incident management.
A key aspect of the demo involved the platform’s dynamic dashboarding capabilities. These dashboards are contextually driven, automatically generated based on the details of an alert, presenting relevant topology renderings, color-coded KPIs, and drill-down capabilities for deeper investigation. The presenter addressed audience questions regarding the dashboards’ dynamic generation, highlighting the utilization of JSON data from alerts to build visualizations on the fly, while emphasizing the possibility of customer customization. He also explained how the platform’s chat ops functionality allows users to interact with the system using natural language, eliminating the need for complex queries and streamlining the investigation process.
The demo also showcased Selector AI’s capacity to integrate various data sources, including cloud providers’ native monitoring services, and its ability to correlate events across different systems to pinpoint root causes. The presenter highlighted features like correlation graphs and a time-series DVR function, which allow users to visualize the sequence of events leading to an incident. Finally, the discussion addressed the platform’s architecture, emphasizing its scalability built on Kubernetes and the use of techniques like Kafka for efficient data processing, and the flexibility to deploy agents to collect data from various environments. This ability to ingest data from diverse sources, whether existing monitoring tools or directly from the target systems, represents a core strength of the Selector AI platform.
Presented by John Heintz, Global Systems Engineering Director, Selector AI. Recorded live in Santa Clara, California, on February 19, 2025, as part of Cloud Field Day 22. Watch the entire presentation at https://techfieldday.com/appearance/selector-ai-presents-at-cloud-field-day-22/, https://TechFieldDay.com/event/cfd22/ or visit selector.ai for more information.
Transcript
I'll introduce myself. I'm John Heinz. I run the SE team over here at, uh, selector.
So, um, I'm gonna kind of show a demo with some of the stuff you saw on the screen, but you kinda see how a user will interact, uh, with the product. And we kind of picked some scenarios based on real customer, uh, experiences. So, um, we're, we didn't do this live this time, we're gonna get the recording, but I'll kind of narrate and kind of tell you what's going on here.
So I'll start the play button. You know, there was a question about, uh, AIOps at the beginning, right? Um, this will go away in, just can you drag it?
You can drag it out there. Drag. There we go.
To me, this is AIOps in some ways, right? I now have, I don't have enough down ticket. I have a smart ticket.
I have a ticket summarizing event. So this summarization can be done using our LM for example. I get a bunch of context, like what's affected and like he was, uh, sat was showing I have a recommended remediation.
I have event summary, so I can actually see all the details that are part of this event. And, um, and I can now have context to go and take an action. I'm gonna pause the video real quick here.
So I also have things like preview mute knowledge incident. So again, when we talk to our customers, they have different workflows they want. Not everybody wants the same thing.
Some love teams, some hate teams, and so on. So we've kind of built, like, we can make these buttons do different things depending on the customer. This could create incident in ServiceNow.
This could create an acknowledgement saying, John Heinz has this ticket, uh, Bob stand down. And so it just kind of creates a nice, uh, collaboration way to kind of do this. Now I'm showing Slack 'cause that's easiest one to kind of demo.
This works across all collab, not all, but most collaboration tools that we've had to integrate with. Um, so now as a user, when I get this, uh, in this case, it's the event that Satchin described earlier. Uh, I have basically a decision to make.
I can choose to go and start asking questions via the ChatOps, or I can actually pivot into the selector UI and start to do an investigation there. And so in the first demo, we're gonna go to the UI to start the investigation, then we'll go into ChatOps. And the second one, we're gonna do the opposite.
We're gonna go into ChatOps first and then go into the ui. So let's let this, uh, let, let catch up here a little bit. All right.
So what I did here is I clicked on the alert, it's gonna go to the portal. And the first thing I wanna point out is, um, everything you see here is relevant to the alert that was just presented. This is not their homepage.
This is not any kinda dashboard. It's, it's basically contextual driven, alert, uh, view. And so, as you saw, we took screenshots from this that we put in the slides.
But basically we can do topology renderings, not just physical topology. It could be like, you know, cloud-based apologies, uh, routing, protocol based apologies. Um, lots of ways we render this stuff visually.
You'll also notice we do a lot of color coding. So we think of things as, you know, is it good or bad? Yeah, that's green and red, but we also have oranges and shades of orange or yellow depending on how you look at it.
The goal here is say like in the case of Verizon, it's not down. It's degraded. So it's easy if it's down, because I can just say, the circuit's out this case, we have to make a call.
Are we going to adjust this thing over at at t or not? And so we use color coding a lot, uh, in our examples. The other thing I'll point out is we think of these honeycombs, and you'll see these different kind of renderings as KPIs at and t.
Is it healthy at this site? Yes or no, green or red. Um, but things make up that health circuit, air stability, uh, BGP flaps, et cetera.
And so we take that context into saying, I think is good or bad. Um, so I'll keep playing the video, um, over here on the circuit. SLA, this is interesting too.
We actually, we do this for a number of customers where if they have multiple providers they use and they wanna know what's my worst and best one? Am I meeting, are they meeting their contractual obligations? And so we again, can take a composite metrics to say, is this SLA, what is the value of it?
Is, is, am I in compliance with the, uh, provider? What they've told me that I'm paying for? That can be done with anything.
It could be your cloud infrastructure, it could be your instances, it could be your connectivity on your private, your direct connects or private links going to Amazon. Like we can me, if it can be measured, it can be quantified into an SLA or to a health value. But what's we've kind of found is that everybody might have a different thing they care about at our ization, they care about X and y debit cares about something else.
And so we, our solutions seem to build this for customers based on their requirements. Now you see here, I clicked on at and t and I get what I call a drill down. Uh, and I'll pause here again.
Basically it says is I can take like that at TI said it was a rendering of the health of a circuit, but there's lots of details behind it. The safe, itt, good or bad. And so in our software, we call these query chains or drill downs, but basically it allows me to get more context about what is good or bad about a thing, uh, and bring up more detail.
So in this case, I'm looking at the admin status, the interface, uh, you'll see the opera status below errors and so on. If there was, uh, a ology there, we would see that here. Yeah.
There's a Question. I have a question about the, the dashboard that's being shared. Yeah, yeah, Please.
Typically you would have to build a dashboard. Yes. And so that would require someone knowing ahead of time, these are the core metrics that I'm interested in.
If I click in on at and t Yeah. Is this dashboard being dynamically created or is it something that it's drawing on a template? Like what's Yeah.
Process Two answers. Um, this dashboard right here, if you go back to, I'm at the three 13 mark. If you go back to this ticket or incident, you know, behind this, there's a bunch of JSON details, like device name, timestamp, uh, data, you know, the, the database that we use for the generate the data and so on.
So we reference those details when we produce this visualization. Like, so this is a dynamic dashboard built on the events that happened. Now, in the case of me clicking on at t, which I did, and you saw that I, I got this little, uh, drill down here.
This was defined by us. We said, Hey, if we click on a circuit, what are the things we care about from a circuit? KPI level?
I click on a device, I wanna look at things like the system uptime, um, you know, maybe the neighbors inventory of all the different hardware components of it, et cetera. So we control that in a, in a schema. Basically, customers can influence that dramatically.
They may say, I want these 10 widgets in a dashboard, and we build those 10 widgets. Um, the trick though is when you have a problem, I don't have to go pull five different widgets together to go investigate it. Selector should be smart enough where all that data's presented to me upfront, which is how we do it for, for customers, answer it.
Okay. Maybe, maybe not. So in short, we have dashboards in the product Okay.
That everybody gets, and it can be customized, but what happens when you click on something can be tailored to you. Mm-hmm. If you want it to be, you know, I want these 10 things in the event that an alert happens, those dashboards are based on the events and the alert.
And so that's kinda dynamically rendered on the fly. Okay. And that alert data is in that JSON blob I'm talking about.
Yeah. And you'll, I think you have an example of that in the slides, kind of how that works. Yeah.
Okay. And if you did wanna create your additional, you know, common dashboards, is it something you can, can you chat to generate that? Yeah.
Is it So you have to know GraphQL. 'cause lord knows I've had to try to explain GraphQL to a lot of people that use sql and it, it's, it's an adventure. Yeah.
The, the whole point of chat ops is to not have to have, you know, any kind of SQL language. Yeah. Um, so behind, like behind every widget here is a query, right?
It's, you know, show me this data as a honeycomb or as a table or as a line plot. But realizes people don't wanna know. They just wanna say, I wanna see, you know, Verizon for the last 12 hours in a certain site.
Yeah. Quick question. Sure.
On dashboarding, sorry, Eric f ql, gummy thinking. Um, how many customers do you see that are going, like with Grafana and tapping into this, using Grafana as an example to tap into this give visualization whole more holistically? Does that make sense?
That, so the question is, um, are our customers tapping in a selector Yes. With Grafana as as a visual representation? Yes.
A visualization layer that they can get? What Money? No, I think not many.
Because in case of, uh, issue with the Grafana is that now you have to prebuilt all those dashboards. Dashboards, yeah. In one of our customers, very initially we had worked with them, uh, for every device they would build a dashboard, they would create the folder.
And finally when something happens, they have to now search for that dashboard in which folder it is. So it's very cumbersome. Now with this, when you can get the dashboards and widgets on the fly for the place where the issue is happening, that's like two clicks away rather than searching for the dashboard in Grafana.
Okay. How are you? Yeah, I'm just thinking like those integrations like more Further, Eric, go ahead.
How do you handle performance and that real time nature? 'cause you have a timescale database as well as a hierarchical representation with a lot of potential edges and nodes. Yeah, like the DVR stuff is tough.
I know this 'cause I used to literal, I worked at, uh, with SEV one and we had to build a really gnarly system. We learned the hard way that outta the box, like just timescale db and a lot of those influx DB couldn't handle and it was rendering was really brutal. So I know it's tough when you're dealing with real time, okay, I'll cover this in the deep dive session, but at the top level, this platform is built on Kubernetes.
And so inherently, uh, like scale out model, right? Additionally, right? Even when they talk about data, the various microservices essentially take a snapshot of data and talk over a Kafka bus.
And so using those type of scale out techniques, we handle the performance, right? And again, if you say, if your network grows, these were, you can just add additional nodes and like scale out even the Kubernetes infrastructures And are using like sampling, uh, rather than full retention. Because that's always the challenge too, is the amount of data you have to hold onto to and how long you hold onto it.
For Right now, there is no sampling here. This is actual full data. But obviously, right.
In many cases, we don't need to necessarily understand, uh, like the data in the raw format, we draw insights from it, and then you can work on those insights. So there is a little bit of filtering happens as we even understand and process data. So I I I can, there are two answers to that.
Yes. Uh, for example, the pro says time series data logs, uh, logs, we keep it depending on the customer, they want to keep it seven days, 30 days. That's the raw data we keep.
But the sampling piece of it, if you want to keep the data for one year, all the events that has happened in the last one year, what are the anomalies happen? Those are the, those are sample, those are kept separately because then you can go back. But definitely when an issue is happening, what we have seen seven days to 30 days is good enough to go back.
We don't have to store all whatever. Yeah. Like high granularity, high cardinality is brutal, but if you, yeah, like as long as you know, MinMax average and the major, you know, ranges you need over time, you can go back and Yeah, as we, as we, you know, some of the use cases around predictive analytics, which you need longer, longer.
And so then you probably start to do average by, you know, one hour, then you might go to one day because you're looking at, you know, six months of data to say what will happen two months from now. Um, so I would try to get through the most of them out here. I'm gonna fast forward some of this stuff 'cause a little bit repetitive.
But the point here is, you know, I'm starting to ask questions. And this is, I think the lady's, uh, question about how fast it, like I'm saying what is the circuit SLA for, um, San Diego. And so slash select is our way to engage our agent.
It's gonna hit the UI or it's gonna run that query in selector and it comes back with an image which I can then blow up and I have links and it's gonna summarize it. 23, like that's, some of the gen comes in, it's taken that data that's in this widget and it's summarized that in English language and we'll talk more about that, um, later on about how that kind of works. But that's, that's kinda how we're doing.
It's a combination of visual and, and text renderings based on the question that's been asked. And so the other thing I wanna point out, this is kind of a, you know, it's a demo for this conference, but you know, we actually have customers that have built, and MLB was a great example of this. They took our platform, they built a whole ChatOps interface in front of it so they can disable interfaces, they can reboot phones, all, you know, using a combination of selector plus NetBox in their case.
And so you can actually drive automations from Slack through our platform to things like potential or other like cloud native kind of automation, Terraform to go and make differences to make changes. And so, um, obvious is more to that and how that works and the plumbing involved, but that's, we're seeing more and more interest in that of, of taking repetitive common things and saying, let's find a way to automate this, um, via a tooling. Now, uh, this is the second scenario where you have, you know, we're saying we have latency detected and a certain area we have a, a ServiceNow incident.
We have a likely root cause, uh, which is the high sea utilization on instance, uh, uws three, and we have a recommended remediation. So now in this case, when I kind was thinking about building a demo, I go, well, I'm at dinner right now and so I can't just click on the link because it's Valentine's Day and you know, it's been a busy week, so I'm gonna start to qualify this alert a little bit. So I'm gonna say things like, you know, show me the topologist application.
'cause maybe I don't, it's a tryouts application. I don't know the topology. Maybe it changed yesterday because it's about native, right?
And so as I ask these questions, um, you'll start to get this, these answers coming back VR copilot, you know that here's the, here's the, the, uh, the topology. Now we're using a combination of like LA long information, uh, metadata from the cloud instances like tags to, to render this, the topology, like how are things connected? Where are they physically sitting at?
Like an what data center are they at? Uh, as well as well as like even like iconic icons to help you understand like what's GCP versus AWS. So the goal here is I see I have my topology, I see I have this link in red, which kind of, you know, makes, makes sense.
But I wanna see the actual latency for everything for ChatOps right now because this is a, uh, a multi-cloud application. And so, um, I'll pause this real quick. And you can see here I have, I have two areas that are in red.
The rest are pretty low millisecond latency. And so I've kind of confirmed my fear that, all right, this is just between these two regions. I don't know if my cause yet just based on the ChatOps in uh, interaction, but I incident other parts of my cloud network.
And so in this case, the user wants to ask a couple more questions like, you know, maybe I'll get lucky, right? Was there any config events and config changes in like user led or like machine led changes are a big awesome source of data for us. Like if de logged in and just did a thing on 10 minutes ago and it broke something, I wanna know about that as quick as I can.
'cause he is not gonna tell me he is gonna go hide. Um, but the point is like this has become a key solution for us. It's pulling in config events because that kind of tells you the secret of maybe what happened that you wouldn't have only known about.
In this case we're saying that there was an instance type change where it went from a 64 core to a 32 core. In this case, we're pulling it from CloudWatch. Now alternatively, I can do all this through our portal, which you're gonna see now.
Um, so that was kind of like me asking questions while at dinner. This is basically showing me a similar dashboard you saw before, but all contextualized based on this ChatOps kind of thing. And so now I can interact with the dashboards, I can see the link, I can see the why the link's red, uh, the latency display.
I have my, the devices that had the violations that are, that you saw in the, in the latency widget there. That's why those two are there. I have my server health looking at different utilizations for disc CPU and memory that we're monitoring.
Um, and so on and so forth. And this will kind of show through, go through each of these things. The one I'm gonna call attention to here is just a second is the correlation graph.
And again, sachin's gonna, I'm trying to give 'em some time, so I'm not gonna go through in too much detail, but here we actually have, like, this is cool, we can actually show diffs, people love diffs when looking at running configs and stuff. So we can show diffs, we can show raw configs. Those are all native capabilities of the platform, um, that we can do.
And these are more cloud native things. I think there's question about the cloud side. Now how do we, how are we connecting all this together?
This is important. Um, and so you see this, you see four different events, app timeout, config, high CPU, high latency. Then you have these lines connecting in some kind of pyramid, right?
But when I start to hover over an event, you'll see context. I'm gonna pause this and what we're saying is this, this event came from CloudWatch. Um, it has an application associated with like a tag.
We're taking CloudWatch or taking tag information from Amazon. This case we have a data center where it's involved and we have a device name or instance name, a local IFP address in a region. Now that's for the config event.
Now we also have like an app timeout event, for example, and I'll fast forward here a little bit. Oops. Um, these other events have similar context.
So, so this is the one for high CPU or high latency. We look at the metadata associated with these events. We look at time, we look at topology, and that's how we start to build a correlation graph of how things are related.
Then we use our causal model to say, all right, did high CPU cause high latency or did high latency cause high cpu, right? And high CPU generally causes latency issues, not vice versa. And that's how we get that causation snippet you saw on the alert.
But the key is getting all this data aggregated together, finding events like config, finding anomalies like high latency, getting the metadata associated with that. So we can actually get a, a real smart integrated alert that goes to the a platform of your choice. Um, the last thing I'll show is it's a new feature we added the other day, uh, actually not too long ago.
It's a DVR. So actually what I did is I took that event and I'm playing it back over time. It's saying, all right, the event happened at say 1800 and you'll see a difference between UT UTC UTC time and local time here, but what actually happened two hours ago, so I can in selectors say, show me the last four hours of this page so I can kind of see how something played out.
So you'll see the convict event happening earlier, like right now, and then it kinda, everything's still green though, right? My memory and my disappear are green. My link is green up here, but as I go to the next spot, it's still green again, which we see happen.
Sometimes it may take a minute for a thing to actually take the effect that you didn't expect it to happen. But now as I progress in time again, you see that now I have this violation happening where the link has gone red, my CPU's gone hot. Very useful seeing how things propagate across the network.
Um, but it's just a native capability of the platform that we wanted to show today because, um, it's neat. So this is like an, is this an agent installed on each system? Is this, uh, a service in, in the AWS and Azure marketplace that you're installing to scan your entire cloud environment?
It's, uh, yeah, I get what you're, No. So, uh, as I was co covering at that time, so it can be installed on customer, on on-prem, or it could be in a public cloud or in our cloud. So we installed where the data is available.
Uh, so it's is closer to the data in just into the platform and provides this analysis. So now in some cases, uh, we have a remote agent, uh, which we install. For example, if you want to do A-S-N-M-P walk and collect the data, if you want to ingest data from the, where the log data is available, the remote agent will get the data and send it to the selector analytics, uh, platform.
So the control pane or the way all the analysis is done, that can be in the public cloud or on-prem where the collection of data can be anywhere else. You as a customer own that. So you say, Hey, I want you to, I'm gonna give you a, um, a roll that can read my CloudWatch logs.
That's when we have doing it. Or you may say it's in Kafka already, just hook to this topic and get the data there. Or you may say, I don't have a option.
You tell me how you want to get it right and we'll go work with you to find a solution. The key is if it's stackdriver logging or cloud wash or like, you know, NSG logs and Azure, like the data's out there. It's a challenge of getting it all pulled together, which is what we're really good at doing.
Like that is our superpower. So, So, so there were two pieces that, that y'all were referring to earlier, either pulling data from existing tools, which is what you just described. Yeah.
Uh, and then having selector AI as just the monitoring and observability tool, which is what you just described around, you have remote agents or you have stuff installed in the cloud or whatever. Okay, so these are, these are two separate things. There's two separate choices that customers have.
No, it can be separate choice or a combination. So for example, uh, if I look at the selector, uh, the stack, mm. So the data that is connected, it's, as John was saying, it could be through a Kafka bus, it's could be a rest API or through a web hook, or we have remote SNMP engine or Syslog engine where it is collecting the data.
All is streaming into the platform. So it can be a combination or it is all available on-prem where the data is next to it. We get the, we can get the data from Splunk or any other log store.
That's also possible. So what what I mean by that is that it could be all at one place or it could be remote engines and selector outside. Got it.
So let's say I, I don't have any monitoring or, and observability tools right now. I'm not using CloudWatch, I'm not using Azure Monitor, nothing. Uh, and I have a bunch of Kubernetes clusters running.
I can go and I can install the agent that's running in a pod to collect the data. To send the data. Got it.
Okay. Got it. Yep.
Yeah, we could hook into your Kubernetes orchestrator and start to pull that stuff back, POD status, et cetera. I Have final question about the demo before we leave. Um, you know, as you were walking through the data, when we look at the realtime data capture, are you utilizing AI in that or is that data analytics and are you going to clarify where AI is?
Clarify. Clarify. Perfect, thank you.
Just one quick question about the ITSM side, that's previous to, to the demo. When you create a ticket, when you make changes and you interact with the system, does that continue to add context to the existing tickets so you can keep adding knowledge to that? Yep.
Yeah. So we have customers where we update tickets. They update tickets, we pull those updates back to us or we keep updating or can close 'em out.
Even it's, again, each customer's kind of different in how they want their ITSM to work. So our job is to be flexible and kind of follow the guidelines they give us. But typically, yeah, updates is a huge value.
It's a lot of systems they have. Like if it's five minutes later, like it just treats as a new ticket. Yeah.
You Oh then Oh yeah. Nice. Thanks.
Okay. Alright, me, I've got one more question and then I promise I won't ask anymore. You mentioned SMP, um, from net, let's just talk network devices.
Um, I assume, but we haven't seen it. Does it also support like netcom, netcom, GNMI? Yes.
All those things? Yes. Because I know, you know where I'm going with this.
Like going into a customer, like don't talk to me about net, uh, you know, CIS log and pulling SNMP data 'cause that's an old school way of doing things and it doesn't scale. So I.