The Operational Twin: Past Present and Future of Your Network with Selector.ai
Learn all about operational twins in this video. See how past historical data and present insights from your network in real time can be combined in the operational twin to predict future performance and issues. In addition, find out how the operational twin leverages copilots and agents to provide full lifecycle management of your network infrastructure.
Presented by John Capobianco, Product Marketing Evangelist. Recorded live at Networking Field Day 37 in San Francisco, CA on March 20, 2025. Watch the entire presentation at https://techfieldday.com/appearance/selector-ai-presents-at-networking-field-day-37/ or visit https://techfieldday.com/event/nfd37/ or https://www.Selector.ai/ for more information.
Transcript
Let me just start with what we're not, and, and we're not offering a Cisco modeling lab or a container lab or a rural environment. We're not so focused on the iOS version of the Cisco device. And if you upgrade it with the new features, how is that gonna impact the network?
We call it an operational twin to differentiate a little bit from the digital twin in that that massive amount of data that we have is the state of your network in the past and with predictability into the future, and obviously the present, right? So again, more needles, internet of things globally distributed network. Hybrid clouds, downtime impacts revenue in SLAs, right?
Downtime. I read some articles, something like $600 billion worldwide from network outages. 75 to 80% of those are human induced.
Um, and the complexity of these networks, the needles in the stack of needles. So what is the operational twin? It's a real time digital representation of your network with live and historical data for deep visibility.
And it enables safe experimentation without impacting the network through natural language interrogation. If I were to change the path on this, if I were to put this BGP in, if I was to remove this network, if I was to add this subnet safe thought experiments using real data and real predictability, now how do we build this? How does this happen?
It's not easy. This is a very complicated problem to solve. But we have, again, optical services device layer one information all the way up to layer seven, offering real-time state configuration differentials at any point in time between any device, um, control plane and data plane visibility and, and the physical topology.
Uh, I'm gonna focus on the DVR capabilities in the video. Um, but here is sort of the dashboard of the DVR. So this is what the path looked like, you know, right now that is what the path looked like yesterday, it looked like two days ago.
A full Google like experience overlaying your traffic with geolocation metadata. So let's watch the video and I'm gonna pause this one a few times 'cause there's some, um, some things to highlight here to set the stage. So what are we looking at here?
This is an internet path tab that we specify an ingress, pe and then the internet address of 8, 8, 8 1. So a Google address. So this is how that PE reached the internet from one side of the country to the other, or across state lines, right?
So what we're gonna do is let the video catch up a little bit. What we're waiting for that, John? Yes.
Quick question. How, how does it determine that? Is it based on like the routing table?
Is it based on ICMP? How does it actually Determine the path? Yeah, it is based on the routing tables and the forwarding plane information.
Okay. Yep. So what we have here is a couple of widgets showing the path time, which is kind of a neat concept, right?
The time of the path, here's, you know, 0 3, 0 5, 0 4, 0 4 at the exact same timestamp. So at the same time of day across different days. And when he clicks on the path, we're gonna see that we can drill down.
So we can click on the topology view, we can click on any of these links. We can drill down and see that that device has three edges. We can drill into the edge nodes and drill into the topology view.
It's fully interactive view based on the metadata and the forwarding plane. So now he's gonna change it to another date and time. And we can see that this part of the path is different.
It got highlighted. So we now, we have a different start and end for a path to the internet, which is using different devices and we can interrogate those devices. The next thing we're gonna do is change to infrastructure paths.
And infrastructure paths is between an ingress and an egress router. Just two random ingress egress routers. This is how the traffic got all the way across the country.
Your providers, your links, your IP addresses, everything you'd wanna know about it. And you can long press for details. You can, um, see the number of edges.
Here's the edge nodes. You can actually click and see what interface is facing, what other interface. Really remarkable stuff.
Great to compare against intent because we'll give you the full path as a text string and you can interrogate that against your intended path. So can we do this through the co-pilot, right? I could use the classic web way and click and click and find and sort and filter and blah, blah, blah.
Or I could just say, show me the path from this PE to this PE in natural language. There it is. And we have two answers.
We have the, actually this doesn't have the co-pilot enabled, um, but we have the path here just from the natural language prompt with a fully interactive dashboard. We can see the, we can click on show and see the different PEs involved and different devices involved. And one of these buttons is actually gonna show the entire path.
This one. So this here you actually see from one end of the country to the other end of the country, every interface that's involved, every bundle that's involved, core routers, PE routers, the whole thing. And we could change this timescale, like a DVR show me yesterday, show me right now, show me a month ago.
Now the last one we're gonna show is very similar, but it's gonna be L three VPN paths. And in this case we, we can put the VPN and the customer and the route. Um, so if you are running multi customer environments or multi MPLS or whatever, you just can adjust your customer corp to get the path through for, uh, L three V-P-N-I-I Have done.
And it's regarding, um, you know, sizing on, you know, if you've got hundreds or thousands of devices in your network, um, this could potentially, uh, be a huge amount of information to select for electric. Yes. And so what are the methods selector use, ingest that information and you know, what method using to kind of guarantee that you're not dropping that information, that you are actually getting a true DVR like experience and, you know, uh, getting information around all the information around and show routing changes and things like that, that, you know, that that could take time.
How do you, uh, mitigate that? That's a great question. So in terms of scale, one of our largest customers, I believe the 24 hour data lake is a terabyte.
Can you believe that? That's just logs and metrics. So we can handle very extreme, large scale networks simply by scaling our nodes.
And we ingest everything into a Kafka bus. So we're monitoring the bus, we're monitoring the nodes. We have Kubernetes that will autoscale the pods.
So if there's a, let's say there's a fiber cut and suddenly there's a huge reconvergence, and now suddenly I'm advertising routes differently, right? That's a, that's almost like a, a mini catastrophe when that happens, right? So our pods in Kubernetes, and it actually leads me to this next slide, so I'll put the slide up, the pods autoscale.
So when there was a hurricane, um, that, that caused a massive increase for one of our customers, um, within five minutes, the, the selector Kubernetes had auto corrected and auto scaled the pods to handle the massive influx of information. So, so Bruno, to answer your question, I think it's, it's baked into our architecture as a microservices architecture that can scale horizontally as it needs to, or vertically as it needs to. And in terms of scale of our customer base, right?
It could be hundreds, it could be tens of thousands of devices, and we simply just continue to scale the number of pods. So we have collector points on-prem, um, that feed the cloud. So, so there are certain data collect collection points, but if we look at the bottom, this is the selector stack in four layers.
The collection service, the data hypervisor, the knowledge service, and the collaboration service. So we've really been heavily focused today on the top of this stack, the collaboration services through Slack or teams through a web goi through a rest API, through gen ai soon through MCP, through a Juniper notebook, Jupyter Notebook, excuse me. And, and these are all individual microservices.
So if, if Slack has a huge demand, we just scale those pods for that particular micro. Now, when we move into the knowledge service, this is, this is really that network language model. You can see it in the center here, and it comes with its own query engine alert engine notifications, and it's fed, right?
This line here is us fine tuning this network language model with the telemetry from the bottom parts of the stack, right? There is a hypervisor. Now the hypervisors just like a, like a VM hypervisor except it's for data.
So when data comes in in one form, another form, J-S-O-N-X-M-L custom text, we wanna normalize all that and make it homogenous through the data hypervisor layer. And we have different transformers that do that. Finally, we have that ingestion engine, Bruno, to your point where we can take things in from rest, from Splunk, from influx, we can feed from your own customer, um, message bus.
If a customer is running their own Kafka, we can take it in from, um, routers and switches to SNP and Syslog obviously, and also through GNMI and streaming telemetry. Um, Quick question for you. Yes, please.
Uh, Jason Ner, um, is this, where do you guys host all of this? Is it, is it like some of it in public cloud, uh, some of it like on your own, your own private infrastructure or is it, you know, It's kind of a mix. So for a customer, there's going to be some collection engines on-prem to get the data out of their on-prem infrastructure into our cloud.
Sure. Um, and then the upper layers of that stack are hosted in either our cloud or the customer's cloud or the customer's on-prem. It's totally a flexible deployment model.
Um, but like I said, this is, think of this as a per customer silo. Okay? Right.
So there's no cross pollution between customers from ingestion layers or, or sharing data or anything. Is Your stuff hosted in a specific public cloud? Like are you guys Google?
Amazon is, Yeah. I, I believe we are using Google. Okay.
If for the select, if you wanted to host it in Selectors cloud. Okay. I would believe, I believe that's in the Google Cloud.
Okay. Yeah. Thank you.
The other thing is, um, Reja Rerum, who is a founding engineer at Selector, a brilliant architect and engineer. Her and I are doing a podcast at the end of the month, blowing up this architecture and really doing a deep dive. So if there's any questions about, if you're curious about how Selector builds this or our architecture or our Kubernetes, or how we scale, uh, tune into that podcast at the end of the month.