Forward AI Demo Troubleshooting Network Operations with Forward Networks
Nikhil Handigol’s presentation showcases how Forward AI revolutionizes network operations troubleshooting. Before diving into the AI capabilities, Handigol provided a concise tour of the foundational Forward Enterprise platform. This robust, deterministic software connects to all network devices across hybrid multi-cloud and multi-vendor environments to create detailed, point-in-time snapshots of network configuration and behavior. The platform offers various analytical views, including graphical topology, inventory dashboards, vulnerability assessments, blast radius analysis, and precise path tracing. A key component is the Network Query Engine (NQE), which transforms raw configuration data into a normalized, hierarchical data model and supports queries via a SQL-like language, enabling users to extract specific network insights and verify compliance against predefined checks, triggering alerts when discrepancies arise.
The core demonstration focused on how Forward AI, operating as a conversational interface, streamlines the resolution of common network connectivity issues. By ingesting a service ticket describing a host’s inability to reach a database server over SSH, the AI agent dynamically constructs and executes a diagnostic plan. This plan involves gathering context about the involved hosts and performing a precise path trace through the network’s digital twin. In the scenario presented, Forward AI swiftly identified the issue: SSH traffic was blocked by a specific firewall due to an explicit Access Control List (ACL) deny rule. Crucially, the system provides a clear, “bottom line up front” diagnosis, supported by detailed explanations of the blocking device, the rule, and the full traffic path, all substantiated with direct links to the relevant “evidence” views within the Forward application, enhancing transparency and user trust.
Extending its utility, Forward AI can also generate proposed Command Line Interface (CLI) commands as a starting point for resolving identified issues, such as creating a new firewall security policy. Nikhil strongly emphasized that these generated fixes are for planning purposes only and require human validation and adherence to established operational change procedures, underscoring that the system does not autonomously execute changes. Discussions highlighted essential guardrails, including the AI’s ability to reject unanswerable requests and the enforcement of Role-Based Access Control (RBAC) to restrict data access and command generation based on user permissions. While a feedback mechanism (thumbs up/down) is in place to gather user input for continuous improvement, future iterations may incorporate business policies into AI recommendations and develop simulation capabilities within the digital twin before deploying changes to production, further building trust and enhancing automation.
Presented by Nikhil Handigol, Chief AI Officer, Forward Networks. Recorded live at AI Infrastructure Field Day in Santa Clara on January 29th, 2026. Watch the entire presentation at https://techfieldday.com/appearance/forward-networks-presents-at-ai-infrastructure-field-day/ or visit https://techfieldday.com/event/aiifd4/ or https://www.forwardnetworks.com/ for more information.
Transcript
Let me show you a demo of what we have made so far. In fact, I'll show you four demos today. So, uh, today I'll show you four demos, uh, each covering a different scenario and showcasing a different aspect of our ai.
Uh, so one thing I wanna mention, uh, this is a live demo and, uh, live AI demo. Uh, so I'll probably be seeing some of this output for the first time. So I'm going to read this, interpret this on the fly, live with you.
It should be fun. Very brave. The first demo, uh, covers, uh, is around troubleshooting, which is a network operations use case.
Uh, before I show you forward ai, let me give you a brief tour of the forward enterprise. Can you introduce yourself so we can make this, uh, okay. Okay.
Hello everyone. I'm Nickel Al, uh, co-founder and Chief ai. Hello everyone.
I Amel Al. I'm co-founder and chief AI officer of our networks, and here I'm showing you forward ai. Before we jump into forward ai, let me give you a brief tour of the forward enterprise product itself.
The base product, it's a software platform that connects to all the devices in the network and creates snapshots of the network. A snapshot is a point in time image of what the network looks like and how it behaves. You can look at the latest snapshot to look at the current behavior of the network, or you can go back to any historical snapshot to understand how the network behaved at any point in the past.
And in order to understand the behavior of the network, it presents various views to the users. There are some graphical exploratory views. For example, this view here shows the topology of the network.
What we are looking at here is a typical enterprise network with, uh, with site spread around the world. As I zoom in, this is a hybrid multi-cloud network with with on-prem locations as well as locations in the cloud. This is Google Cloud, AWS, so on.
I can also see what is inside the network with a dashboard view. This is showing what's present in the network. Uh, this is a multi-vendor environment.
There are devices, uh, many devices from Cisco. There are devices from Arista, uh, there are Palo Alto devices, Fortinet, uh, Juniper and so on. So multi-vendor environment, and there are different types of devices.
There are switches, routers, firewalls, load balancers, access points, and so on. That's the on-prem part. There is also presence in the cloud, and the cloud presence spans AWS Azure and Google Cloud.
There are also tabular views that, that I can explore. Here's an inventory view that shows everything that's present in the network, an inventory view of, of this network in a tabular fashion. This is a table showing all the hardware that is present in the environment, uh, including the parts part IDs and serial numbers.
So from a snapshot perspective, are you taking this on a periodic basis or is it something like a change driven snapshot or, or how, how do you determine when to take a snapshot or not? Yeah, Uh, both actually. So, uh, there is a regular cadence with which customers take snapshots.
This can be a few times a day, once every hour. There's a, there's a standard cadence, that's one, but there's also on demand snapshots that are taken around changes. This is Marian Newsom.
I have a quick question. How do you enforce the boundary between the math, the AI interpretation, and how do you enforce the source of truth? Is that through inferred data or how are you doing that?
Yeah, what you're seeing here is all based on this is deterministic system. I will show you forward AI that sits on top of this foundational platform that is entirely deterministic. The AI part is, is, is the conversational part that comes on top, which you'll see very soon.
Okay. So for every AI answer, is it traceable with a timestamp or something like that? Exactly.
And I will show you that in the demo. Okay, thank you. These are some tabular views.
Here is one interesting tabular view that shows all the vulnerabilities that are present in the network. These are all the CVEs that are present in the network based on the operating system version that these devices are running, and also the configuration that is present on these devices. Uh, it also shows not just the CVEs that are present, but also whether these are exposed to the outside world based on the behavior of the network.
There are also ex uh, there are also queryable views where I can come in and ask questions and get answers relevant to that query. Here's one example. This is called blast radius analysis.
Here, I can come in and enter the name of a host or an IP address. And what this view shows me is if this host were to get compromised, then what are all the other assets in my environment that can be reached from this host? This is an exhaustive list of all the targets that are reachable from this host.
Based on the connectivity that is provided by the network network, I can also query and understand the exact path any particular traffic of interest can take through the network. Here I'm querying a path from this SSH, so one post to this other destination, and here it shows me the exact path that this traffic takes through the network. It goes from an on-prem location through switches and firewalls out into AWS uh, VPCs as well.
Another interesting queryable, uh, view that I want to show you is called N-Q-E-N-Q-E stands for network query Engine. I had mentioned before that we take all of this raw configuration and state and turn it into a normalized data model. Mm-hmm.
This is that data model. This is a hierarchical data model that shows everything that is in the network and how it's configured. It's, it starts with the network at the top.
It shows details about all the devices and the details span information like system information, platform information, VLANs interfaces, routing, information spanning, tree acls, nas, uh, security zones, and so on. So this is, this is a very comprehensive data model, and it comes with a query language query. The user can write a simple SQL like query to extract any part of this data, any, any data from this data model, run logic on top of it, compare against other pieces of information if they need to, and have the system run a report.
When I run this, the system extracts all of this data and produces a report with just the data that I care about. Now, there's a small, uh, bit of friction here, uh, is that our users need to write a query. Mm-hmm.
Uh, and, uh, and you will see very soon with forward AI that the users like we are going to remove that friction point of friction as well. Yeah. Another, uh, capability of the forward platform is the, uh, ability to verify.
It's called, uh, it's called verify verification. The ability to verify that the actual configuration and behavior of the network matches intent. I can go ahead and define checks in terms of, uh, what it means for the network configuration to be hygienic.
Uh, I can also write checks in terms of what sort of end-to-end connectivity should or should not exist in the network. And every time power takes a snapshot of the network, it validates that those checks, like the expectations and the behavior match. And if, if, if they don't, then it flags, uh, flags an alert.
So that's, that's a quick, uh, tour of our, uh, platform as it stands today. But let's jump back to our demo. Right.
So this is the, the first demo is, uh, is a scenario, it's a network operations. We are all in the network operations team, and we are gonna be troubleshooting a problem. So for that, uh, I am going to go to forward ai, which is a simple conversational interface, and I'm gonna collapse everything else because everything else is completely relevant for this demo.
A simple conversational interface. Imagine, uh, we are in the network team, uh, in the network operations team, and we get a ticket that looks like this. We got a ticket, uh, from an employee, and the ticket is about a connectivity issue.
The, the ticket says there is a host, SJC DC one DB host 12, that can't reach another database server in Atlanta with i some IP address over SSHA pretty routine kind of ticket that network operations teams have to deal with on a day, in day, day-to-day basis. Mm-hmm. Now let's look at what is, uh, a traditional way of dealing with a ticket like this.
And so as an operator, if I get a ticket like this, I'll start by first reading the ticket and understanding what this ticket is about. And then I start gathering some context about the various things that are described in the ticket. This ticket talks about this SJC DC one DB host.
Well, so I need to understand what is this? What is this host? Where is it located?
What is it connected to? Some basic context about that source also talks about this other IP address. So I need to understand more about this IP address.
Again, who does this IP address belong to, and where does it live in the network? What does it connected to? And so on.
Once I gather that basic context, I have to start diagnosing the problem. If I have, if I'm able to log into this host, I may run some pings, some trace routes just to get, to get to understand like what's going on. Or, uh, more often I will start logging into devices.
I'll start logging into devices, starting with the device that this source horse is directly connected to, run some show commands, try to understand what that device could be doing for that particular packet, move on to the next device and do this hop by hop analysis. What our customers, uh, tell us is routine tickets like this can easily take hours to, to resolve. Mm-hmm.
And, and it's also, uh, not uncommon that while I'm in the middle of resolving this ticket, I may get interrupted with a, with a higher priority item. So if that happens, then I have to put this back in the queue, jump to my higher priority item, and then this ticket sits in the queue. So what would take hours can easily stretch two days.
The problem is, while this queue is sitting unresolved, the end user who needs this connectivity is blocked and inefficiencies like this cause businesses to suffer power. AI changes that. Let me show you what the workflow looks like with forward a with forward ai.
It's a conversational interface. So I can come in and express my goal directly in plain English triage the ServiceNow incident with the, with the ticket id, and I hit run under the hood. It's an agent tick system based on the intent, it dynamically puts together a plan and starts executing through it.
And as it builds the plan and executes through it, you can watch that happening live on the right side panel First. This is consulting. Do, do you validate that the request is, is is truthful, meaning maybe it's somebody saying, I want request, but the request was never authorized in the first place.
Yeah. So there, there are various scenarios where we have to put guardrails in the place. For example, this could be a request that may not even be answerable by this AI agent, right?
Right. So we have this, uh, gating system that will reject anything that the agent is not capable of handling. Okay.
Yeah. So what, what, uh, let's look at what this, uh, agent just did it, uh, it, it got details about this ServiceNow incident. It read the contents of the ServiceNow incident, and it started getting some context exactly what I would've done as a human being, but automatically started getting details about SJ CTC one, DB host 12, and then it did a path trace using the digital twin of the network.
Remember the hierarchical data stack I mentioned there is a behavioral knowledge that is available to the AI agent. It's using that to trace packet, trace the path from the source IP address to the destination IP address over 4 22, because I meant the ticket mentioned. The problem was with SSH traffic, it does a path trace, and then it goes ahead and gets some details about this, uh, other device, SJCD DC one panel SFW one, I had not even mentioned this in the ticket.
Maybe it found something as a result of, uh, tracing that path. And maybe it's like it's getting more detail. So this is, this is what I mean by putting together a dynamic plan and executing through it.
Mm-hmm. And once it's done with all of this, it comes back with a response. I'm, I'm, I'm gonna read this response live with you.
Let's, let's look at what, uh, the system, uh, just said. It says, SSH connectivity from this SJC DB host to this other Atlanta host IP address is blocked at the SJC data center firewall. This particular firewall due to an explicit ACL deny rule on the egress interface, it's giving me bottom line upfront.
It's giving me a direct answer to the problem that I was encountering. Then it goes into more details, uh, like it digs into more details, uh, about the blocking device and the rule that is blocking it. Uh, it tells me more about the path that this traffic, uh, would take through the network.
Uh, and then it gives me more, uh, other more context. So it's also giving me some host context, uh, just in case that's useful. It's, uh, showing me that, uh, this host actually runs a bunch of services.
It has six critical vulnerabilities and 34 severe vulnerabilities in case this is, this is relevant to the problem. It's giving me this additional context as well. More importantly, this is, this is the interesting section.
This is, this section is evidence. These is the data, these are the views of the data that it extracted from the digital twin based on which it came up with the answer when it asked for host details for that S-J-C-D-B host 12, this is what it received from the digital twin. This host, uh, is connect.
This is the connected interface of the host. This is the IP address, vlan, mac information, gateway information, et cetera. So this is all the context it gathered about the host, and it did a path trace.
When it did a path trace. This is the path that it saw the traffic taking through the network based on the digital twin. And this is a hub by hub view of the path that the traffic would take from the social destination.
And here's a firewall along the path. The Pan OSF W one that is blocking the traffic. It actually goes beyond that.
It actually models what the path would be beyond this firewall as well. So it goes ahead through the firewall because all of this is modeling. This is based on the behavior of the network.
This is not ping or trace route. It is able to push through beyond the firewall and see what is, what else is going on, uh, in the network. Is the routing ready, for example, to deliver the traffic all the way to the end?
And all of this evidence at the bottom comes with a link that I can click here to see that view directly in the forward application. I can use this as a starting point, dig in deeper if I need to, but the point is I don't need to take the answer that's coming from an AI agent on its face value. There is always evidence.
This is transparency. This is how we can build trust in agent tech operations. So Al Jack Fuller with, uh, paradigm Technica, so I know networks don't change very frequently, but they do change.
How are you picking up the change and what's the time lag? Yeah, uh, excellent question. Uh, as I said this, every, all the, everything that we saw here, the behavior that we are analyzing is based off of a snapshot of the network, right?
So that the data here is as, uh, as fresh or as old as the last snapshot of the network. But we have mechanisms in the platform where you can quickly refresh just pieces of the data that you care about. For example, this is a path, uh, that we saw in the four platform.
Mm-hmm. That is an ability to real time refresh this path just to see, just to confirm that right now, at this moment, this path is behaving exactly like what we saw in the last snapshot. Okay.
So, so the snapshots are manually generated, or is there like automatic and routine generated? These Are periodically taken snapshots. These are automatically generated, but you can take manual snapshots also if you need to.
You want, yeah. Okay. So let's, let's go on And il, in that particular case, you were taking a snapshot just of the change path, uh, just a path, not, not necessarily the whole network configuration.
Yeah. You have the ability to take just, just a snapshot of just that path. Uh, if you want a quick refresh of just that section of the network.
Okay. So now we have, uh, we have a diagnosis from forward ai. Let's have it update the ServiceNow incident with the diagnosis information.
And I hit run and the agent, AI agent kicks in. And what it's doing now here is it's adding a ServiceNow comment. It's taking all of the diagnosis that it just performed, it summarizes it, and it's gonna update the ServiceNow, uh, incident with, with all of that diagnosis information automatically.
And this is the, this is the summary information that, uh, it is going to update the ServiceNow incident with and it succeeded with that. So if I go to the ServiceNow incident, I scroll here, it's updated with the diagnosis information in the form of a work note is the root cause that's identified and, and all of the investigation details. That's great.
Uh, so, but, uh, what would be the next step if I am dealing with a trouble ticket like this? So after getting the context, after diagnosing the problem, my next step would be to figure out how do I resolve this issue. Mm-hmm.
Uh, we can, we can take forward AI's help for that as well and ask forward AI to come up with a fix to resolve the issue. And again, the AI agent is gonna kick in and it's gonna take all of the necessary context that it needs to come up with a resolution. It takes into account the original problem.
Mm-hmm. The diagnosis that it performed, and it, during the diagnosis, it found that there was a firewall that was blocking the traffic. Mm-hmm.
Right? And based on the digital twins, uh, twin digital twins knowledge, it knows exactly what the device is, what platform, what operating system that device is running. Mm-hmm.
It also knows what is the base configuration on that device. So using all of that information, it is now generating CLI commands. Did it ask for authorization or was that implied when you typed in, please fix this?
Uh, great question. So just to be clear, this is not making the fix. Gotcha.
This is just planning the fix. This Is just giving you command line, a headline, like yeah, a starting point, a head start into fixing this. Like what would be a good fix for this problem and what did it do?
It's creating, again, it's being very open and transparent about the work that it is doing here. Mm-hmm. Create a firewall security policy to permit SSH traffic from the source IP address to this destination IP address.
Uh, and this rule should be applied before any deny rules and should allow traffic to pass through the firewall. Like that's what it's trying to do. Right.
And these are the CLI commands that it is producing as a result of that. Again, just to be clear, this is just a starting point, right? So earlier, uh, we saw that the digital twin, uh, includes business level context or can include metadata.
Uh, if there was metadata or if there were, uh, access policy security policies in that metadata, would the AI have called that out and said, yes, you can't do that because you're not allowed to do that? Absolutely. I think that that would be, that would be an excellent way and like, uh, very appropriate way to enrich the, the agent forward AI's capabilities to take into account the, the business data.
Mm-hmm. If they have, if I have business policies on what should or should not be allowed, it can take into account and it may, uh, and it would come, come back and say, if that was the case, it would come back and say, yeah, these would, uh, like this is how we would do it. But you're not supposed to because of your business policies And forward ai, um, in this sort of fixed mode, can it actually say you need to purchase new, new equipment, new switches that nature to be able to, to do what you want or maybe to optimize, uh, SSH access, you may need to put another, uh, I don't know, router in this, in this environment, something like that.
Or is, is that not, is that part, is that outside the scope of Well, I have not tried it myself, so I, I don't know what it would say like one way or the other. I have not seen it say anything of that nature. It's, it's, it's strictly confined A new Cisco 9,000 foot, something like that.
They charge the vendors for that feature. Oh, I see. That's the free version.
It comes with ads. Okay. I just, just, just a question.
So I have a, I have a question related that to that though. So you talked early on about the fact that this is a natural language interface, which it is. And the first part was definitely, well, okay, we're giving you a natural language interface to our environment, and the first set was basically using that to translate to that, to commands that would be given to just like a human would give them the, the same commands to the system and it would do that diagnosis that I get.
Now we're back into the, we're actually doing generation here, right? The true generat AI generation. What guardrails are in place and how are we controlling that, that that's actually generating valid rules and that they not only just valid commands and valid rules, but that they do actually what you want 'em to do.
Yeah. Like, that's, that's an excellent question. And, and like this is super important, right?
Like this is like, it also explicitly stays says here, validate before you execute. And so what this is helping you is giving you a head start in terms of how you could resolve it. Mm-hmm.
But this is absolutely no substitute for all of the validation work that you would have to do to actually execute the changes. Like you, you need to, uh, enterprises need to follow that normal operational procedure when it comes to validating and executing the changes. Like that doesn't change at all here.
Question, follow up to that, since we're talking about the validation and that's still necessary. Are you collecting any statistics from NetOps or SecOps about whether these changes are accurate, right? Are they, are you collecting some sort of I see the little thumbs up there, right?
Yeah. Yeah. Is that important in your dataset saying we are creating those validation rules, you're getting that feedback loop from the individuals using that.
I was just curious if that's a super important thing. Maybe the, maybe the thumbs up needs to be a lot bigger, so say, please click this if it's right, you know? Yeah, yeah.
So I know that this is accurate. Yeah. Like that, that's, that's an excellent point and absolutely there's a reason why we have this thumbs up, thumbs down because it's really important for us to see what forward AI is doing and whether it's actually solving the problems correctly or not.
In fact, uh, what's behind this thumbs up thumbs down is actually a more descriptive, uh, input section where the user can input more specific details. It's not just a, a binary Yes, no, good, bad. Okay, gotcha.
Feedback. But like, they can give more elaborate feedback and that would be taken into account by us to improve forward Ai. Right.
And you're building the trust. Right, exactly. That's the, that's the whole thing.
You gotta have the human in the loop, build the trust so that I can come back to this more and more. Yeah. Get to that 99 percentile of things are looking good.
Yeah. Wouldn't you apply this to the, the digital twin before you actually applied it to production? Yeah.
So like that, that would be like, uh, one of the like, I mean that's like absolutely an important way in which the digital twin needs to evolve. Uh, I don't have any announcements on that today. I have a question.
Is this our back base? So like, um, will, will the execution of commands be restricted to whatever permissions that current user has? Yeah.
Just to be very clear, again, we are not executing any commands right now. This is just generating commands, but even access to the data, uh, is subject to subject to r back roads. Our platform has, uh, very fine-grained RAC uh, policies that allows users access to just certain kinds of data.
For example, there may be, uh, network teams that may, there may be classes of users who don't have access to firewall configurations. And when that happens, they won't even be able to generate this because yeah, they, they won't, they won't know what's there in the net. That's really good.
Then another very kind of really maybe silly question, is this, does this only work in English or can you have this conversation in other languages? Uh, yeah, good question. Uh, so we have, uh, uh, I, we have mostly tested this with English, but there was one instance where we, uh, some of our, uh, Japanese customers just typed it out in Japanese without us prompting them and actually produce really valid like answers.
Oh, awesome. Yeah. That's good.
So Josh's question a little bit, uh, I wanna dig in a little bit more. I mean, obviously because you are, you've already got this digital twin, and then like, I, I'm, I'm, I guess what my question is around like logging and, and change management and how that feeds back into the system. What I mean is, so like in this example, right?
I've now asked a question of the system. So I've basically told the system that there was a problem. The system has given me a potential answer for that solution, and then if I go and load that config into a switch or router or whatever it is, I've now changed the system that's gonna pull a new copy of a digital twin.
Yes. Is there any like, validation of that, of like, does, is there a learning happening there where, okay, a problem was identified by a user, this problem is now gone, this code changed, so I know that code was good and so I can suggest that again in the future or the code that was applied was different and so I should suggest different code in the future. Is there any learning that's happening there, I guess is the Question?
Uh, yeah. I mean, that would be one way to evolve the system over time. Right now it is not doing that.
Okay. Like in, in this current version? Yeah.
Of Fair enough. Okay. I think one thing Hil that, um, you know, we talked a little bit about this yesterday.
It, at the end of the day, I think we'd all like to get to remediation, but, but there's a trust before we get there. So whatever we can do to collect metrics to provide people, you know, you, you can trust us now, go ahead and do the remediation. So I would like to see the thumbs up, thumbs down be massive, right?
Yeah. Yeah. Because I think, and I think it feeds the broader community, right?
Uh, all these tools are are people of the day, again, they wanna get to remediation. Yeah. So whatever all of us can do collectively as a community to say, yes, you can trust AI to make these changes in your network, regardless of whose Google is doing it.
You know, it'd be great if we all fed that into the beast, right? To make sure Exactly. People were comfortable doing it.
Agree. I'll talk a little bit about trust, uh, in like after the demos as well, so we'll talk about that. Okay.
Perfect. Perfect. Okay.
So, uh, with that, I mean, uh, if, if we are actually good with all of these changes, I can even ask it to go ahead and start a ServiceNow change request with the proposed fix. I can, uh, yeah, I can, I can have it, uh, churn on that as well. But, uh, and then go ahead and create a change request with all of the details, all of the diagnosis and all of the proposed fix.