Enhanced Observability and Correlation for Hybrid Networks with AIOps from Selector AI
Selector AI, presented at Cloud Field Day, showing an AIOps solution providing comprehensive visibility and intelligence into the complex networks, cloud infrastructure, and applications of large enterprises. Unlike existing monitoring tools focusing on single data sources (metrics, logs, events), Selector ingests diverse data types from various sources including existing databases and monitoring tools, config events, alerts, topology information, and even CSV files. This unified approach allows for advanced event correlation, drastically reducing alert volume and the associated workload on operations teams.
Selector’s natural language interface is a key differentiator, enabling users to query the platform using plain English rather than complex SQL queries. This, coupled with a digital twin capability for operational insights and “what-if” analysis, provides a significantly more user-friendly and accessible experience. The platform integrates with various communication channels like Slack and Microsoft Teams and ITSM tools, enabling proactive alerting and reactive querying across different teams and workflows, breaking down the typical siloed approach of NOC centres.
Selector’s deployment model offers flexibility, supporting public and private cloud environments, on-premise installations, and even integration directly into existing Google Cloud or AWS instances. Importantly, their pricing model is not tied to data volume or user count, instead focusing on a predictable cost structure based on monitored devices and use cases. This allows for more straightforward budgeting and encourages the ingestion of larger datasets, which in turn improves the accuracy and effectiveness of Selector’s insights. While Selector integrates with existing tools, its ultimate aim is to provide a single source of truth, eventually replacing the need for multiple, disparate monitoring systems as customers realize its efficiency and comprehensive capabilities.
Presented by Deba Mohanty, VP of Solutions Engineering and Customer Delivery, Selector AI. Recorded live in Santa Clara, California on February 19, 2025 as part of Cloud Field Day 22. Watch the entire presentation at https://techfieldday.com/appearance/selector-ai-presents-at-cloud-field-day-22/, https://TechFieldDay.com/event/cfd22/ or visit selector.ai for more information.
Transcript
Uh, today, uh, we will be presenting first time at the cloud field day. We had presented previously at the network field day and primarily our focus today will be on the hybrid cloud networks and how we use platform observability and provide visibility to the network and infrastructure, not just, uh, cloud. So to introduce myself, I am Deba, I manage the solutions team at, uh, selector and co-presenting with me will be Sachin, who is leads the product management team.
And John will be doing the demo. He's our lead for the global systems engineer. And we'll go through this.
Uh, today primarily I will do a brief introduction about what selector is because this is our first presentation for the cloud field day. Then we'll go into a description of the product and the platform and do a demo also. So just to set the stage, what selector is and what this platform provides it.
If you look at the existing platforms or existing monitoring tools, they're primarily focused on a single type of data source, whether it's metrics, logs, events. So primarily when someone is looking at the whole infrastructure, what's happening, they have to look at multiple different data sources, multiple different devices to figure out what's happening. So where selector sits in the, uh, in the infrastructure where you are providing the observability on top of these data sources.
The data sources could be from your existing databases, existing monitoring tools, and you can see the list of items that we can ingest. We can ingest, configure events, metrics, alerts, logs, topology information from even from a CSV file or from LLDP or CDP, or any other information that, uh, the customer may have. Once this information is ingested into the platform, what we provide is primarily the first and the foremost thing that is the event correlation and reducing the number of tickets.
Because anytime there is a issue happening, you'll see hundreds of alerts coming your way and they have to all, they have to deal with all the alerts, figure out what's happening. The first and the foremost we provide is that event correlation and reduced the number of tickets so that operators don't have to spend time looking at, uh, all these, uh, alerts that are coming their way. That's the first, uh, benefit.
The second one is we provide a way to communicate with the platform in natural language, because if you look, you have to very, you have to know the platform to get information out of that. But what we have provided is a interface where folks can ask natural language questions. For example, let's say I want to know what are the top five events that happened in last 24 hours in a particular site.
I can get that information from there. This is one, just one example and we'll do the demo. The third is we provide a digital twin, which provides a operational information about the system.
For example, if you want to do a what if analysis and simulate some of the, uh, failures want to know what is the usage you'll see at a certain time of the day. Can you go back in the time and look at what has happened? We provide that also.
Now, just looking at the problem statement, if you, this is primarily if you look at a current NOC center, they will have all these silos. And the silos are not just about tools, devices, it's about teams. Also that will be a network team looking at their own setup tools.
There is a infrastructure team looking at their own setup tools and similarly application team. But when a application failure happens or there is issue, let's, let me pick something like, let's say there is a high latency that we see for IPTV transmission. It is not just application that could be correlation and that could be events that is happening in the infrastructure could be happening in the network where it is happening.
So we correlate all this data and provide a summarized view of why, what is happening and why it is happening. So now what are the core key things that we primarily focus on? So as I said, we provide a interface where you can ask questions in natural language and also we provide an interface where if needed, we provide a SQL-like interface to ask queries and get the information.
Both are available, as I said for the digital twin. We can go back in the time and see what happened and also use that operational model to see do what if analysis that's also is available. Now, there are multiple ways you can interact with the platform.
You can use our portal, uh, the web portal and go and see the dashboards that's available. But the primary mode of interaction that is very useful is through ChatOps like Slack, Microsoft teams providing this information to A-I-T-S-M tool like ServiceNow, PagerDuty. So now, now what that has done is that now the number of folks who can come to the platform and start using has expanded.
It is not now not limited to only the NOC team or the folks with access to that particular tool. You can ask a question in a Slack. The same question that I said, okay, what are the top five issues that happened in the last 24 hours for a given site?
You can ask that in a, in the portal in Slack teams or any other place. Now asking the question is much more reactive. The proactive things that we do is that when we correlate the information and file a actionable insight and a ticket, that ticket can go to a ServiceNow or PagerDuty or ITSM tool or that alert can go into Slack and teams.
So there are two ways of interacting with the platform. When there is an issue, you go and ask in a reactive way or proactively we look at what's happening and inform the user. Now what are the top verticals that we have been successful with?
Uh, definitely service providers. Um, all the large, uh, telcos we have in us, Canada and other places. That's primarily and now we have expanding to other geographies.
Also, another big segment is retail. Primarily. If you look at any retail or any hotel chain, they have multiple different locations and they are looking at what is the health of my particular store?
What is the health of my particular, uh, hotel, uh, location, health, I mean connectivity, reachability to the cloud application that is running and serving those locations. What is the health of that? That single pane of glass was missing in a lot of these places and we are deployed there.
And two other, uh, segments that we have also been successful with is, for example, transmission of uh, games sports. Uh, like, uh, NBC sports was one of our first customer. It's also on our website, uh, their testimonial and also for finance in financial segment.
The key here is that if some of the applications are misbehaving because they are doing a lot of financial transactions, why is it happening? Is it happening because of the issue on-prem because of the application that is being serviced from the cloud or it is because of the network? That's the first question they ask when they come to our uh, platform.
So overall, to summarize what are, what is the value that we provide to our customer, as I said, proactively detect and inform them what's the issue happening while those tickets consolidate the number of alerts so that they have, don't have to deal with thousands of alerts. They're just looking at five or 10 with which a human can deal with. And it has enough context and summary available so that they can go and take the next action.
I'm seeing Kubernetes and application infrastructure there. Does that mean you're doing anything in the a PM realm as well or just at the system and network layer? We are doing, uh, so I I I'll, we are not doing in terms of going inside an application and doing tracing, but we are monitoring the CPU memory.
Got it. Usage of, for example, we use selected to monitor our own applications. So for example, there are multiple different types of databases, Prometheus, low key, lot of others.
What, what are the number of requests these, uh, uh, we are receiving? Is the number of requests going up? Is the usage going down?
Those are the metrics that we measure in terms of the, uh, application and then correlate that is there any other event that is happening in the network and join them together. Makes sense. So now, uh, where is this deployed?
This is the question that comes up every time. So the way the platform is built, it can be deployed in a public cloud, in a private cloud, in a on-prem customer, uh, location. We just need VMs based on the volume of data, velocity of data, uh, where you have a spec available, it can be deployed anywhere.
The reason we have this model, because most many of our customers, when you go to them, they don't want data to go to a public cloud. They want to keep that data OnPrem and there are customers who want to have, in a way, they want to have that is that our deployment is in their, uh, Google project or AWS instance. That's also PO possible.
So any combination is possible and also we provide a SaaS service. Uh, if we get the data directly to our public cloud instance, we can, uh, provide the service from there. And a key innovation that we have done is in the pricing model.
The price of our platform is not dependent on amount of data or number of user. This is the first thing we realize that if we price based on number of data, because the more data we get, our insights get better. So we don't want to charge based on number, amount of data and number of user.
It's primarily based on use case and number of devices or number of entities that we are monitoring, which is very predictable. You say I'm monitoring thousands, 1000 devices in a one location. If that doubles, yes, it's a predictable pricing growth.
We don't want to price based on amount of data that we are ingesting. That's kind of the difference between existing tools out there and our pricing model. So before I hand it over to such to do a deep dive, any questions or anything just from an intro point of view that we can answer or we can take the questions at the end of the session?
Also Just a, sorry, just a logo call out. 'cause I noticed that you had sort of the enterprise a PM side and you listed AppD is the first one rather than Datadog, obviously AppD is a very good enterprise, uh, play and they're already well in there. They're also sort of sliding underneath the Splunk world.
So I was curious why I would think that Datadog would be a primary vendor you'd see in the wild. Uh, There's a reason behind that. Uh, when we have gone to customers and try to get data events from, uh, Datadog, they limit their API calls and they don't export the data.
So therefore, uh, we can take it, uh, uh, and um, build the solution. But we have hit roadblocks in the number of requests that we can go do and the type of data they can export to us. So for the Datadog people are watching, this is for you learn how your customers use your products.
That's that's really cool. Thank you for catching that. Did You say you price per device?
So for different use cases, that could be per device or let's say, uh, number of entities and depends on also if it's a, uh, edge use case or there are thousands of CP devices, but there are two core devices. So we have modules for different kinds of devices and use case. But my point was that it'll be predictable.
It is, we just don't increase the price just because the number of events or alerts have gone up. So You have things out at the edge and in the, in the data centers and all over, I've got about 35 other licensing questions that I'm not gonna ask 'cause it's not relative. But when we're offline, let's, let's talk later.
So to be clear, you're, you're positioning your product, you're integrating with other observability tools, you're not replacing. So In majority of the case, we go in in a way that we take data from existing tools and provide the information. We have a remote engine to take the data from devices or from existing tools.
But yeah, over time customers have realized why do you have need one more monitoring tool under selected and they've replaced it. But to point, when we go to, You slide in interfacing with the tools and then your goal is to displace them eventually. Got it.
If they determine they don't need them anymore, correct it. Okay. And yeah, it's a journey.
Uh, it does not happen overnight. And they see that, okay, why have the same thing that we can collect, go through another set of tools and provide the information. Yeah.
So Would your competitors also be your, your partners of the people you're integrating or who would be your biggest competitors? So right now our biggest competitor if you look at it, is primarily the folks who want to do it yourself. Mm-hmm.
And they want to build a, they have a team, they tried it out. Those are the places where we have been more successful because they understand the problem, they understand the complexity. And when we go to and present a proposal and they have their existing monitoring tools, it kind of works best for us.