Nile’s Architectural Approach with Shiv Mehra
Go behind the scenes on what makes up the Nile Service Block. The Nile service block is comprised high end enterprise grade access points, switches, aggregation switches, physical sensors. Get a peek into how the Nile’s Layer 3 based secure network fabric autonomously forms itself with zero config.
Presented by Shiv Mehra, VP, Service and Solution. Recorded live at Mobility Field Day 13 in Santa Clara, CA on May 9, 2025. Watch the entire presentation at https://techfieldday.com/appearance/nile-presents-at-mobility-field-day-13/ or visit https://techfieldday.com/event/mfd13/ or https://www.NileSecure.com for more information.
Transcript
Hello everyone. I'm Shiv. Uh, let's go deep into the, you know, NSP if you will, uh, the NI Nile service block.
But it all starts with the purpose with hardware. I think ESH spoke to that. Uh, our goal really here is not to sell hardware, right?
We are not selling you an ap. So if you look at our aps, they're all four by four, calling four today, we have three types of aps. We've got a wifi, six indoor ap, a wifi, six outdoor ap, and a wifi six E indoor ap, and we have a wifi seven direction antenna ruggedized AP that's gonna be coming out.
And we also have wifi seven indoor AP coming out. But the key thing is it's one sku, meaning that it's just four by four code four. You don't have a two by two version or one by one version.
The idea really here is to give the Cadillacs of Cadillac to, to all our customers. They all up to 5G PS uplinks. We've got a Bluetooth built into it.
So we make sure that this is the top-notch hardware that's out there on access switches, which is where the APS plug into, right? We've got a 48 port switch. We basically look at, got multi gigs a hundred all the way to 5G PS on every single wide port.
We've got P OE POE plus plus on every single white port as well, and they get uplink into a distribution switch. You guys are still very young and, and I'm sort of curious what your plans are for lifecycle managing out in particular wifi aps, wifi six to 60 to seven. Are you plan if somebody needs a new AP and you're no longer making wifi six aps and you send them a wifi seven ap, it's gonna create clumping and all of that, right?
How do customers make sure that they keep the latest hardware as you make new products? Absolutely. You wanna take this fresh?
Sure, Sam. So this goes back into, you know, like, uh, our contacts. So when we talk to the customers, we look at, you know, probably three to five years contact terms.
So when we continue to make sure that if you have wifi six or 60, we continue to make sure if you have to replace one or two pc. We know like yesterday there was a clear article saying that you cannot make some of these things and we understand. So we continue to keep some of that, but we move on from six to 60 to seven.
We do not try to continue to sell wifi six when there is a wifi seven. It's only needed to support the existing customers. And then when that, uh, uh, contract expires or when we work with the customers and figure out how many devices are capable of wifi seven, what's the capacity on the network, what kind of constraints you have on the network.
So even during the contract time, if needed, we will upgrade the hardware too. And those are one-off conversations, or are those planned through the, the life of the Contract? So we have a customer success manager who works with a customer on a regular basis.
It's an ongoing conversations with the customers. Thank You. Awesome.
Thank you. Sosh. So the, on the access switches, what you'll see is you can, you know, kind of stack them, they look like a stack, they're not a stack, they're ring.
And we'll talk a bit more about those details. And you can actually easily scale up and down. And these connect into a distribution switch.
And these distribution switches are aggregating all those access switches. Now, the key thing to note here is, you know, up here, if you see the distribution switch is connected to the, you know, upstream router and firewall. And as said, we have single architecture, but we can have multiple topologies.
So if you have a small office suite, you don't need a distribution switch, you can just have an access switch that connects to the upstream router firewall. It could even be that internet gateway that he was talking about. The edge service.
What the idea really here is you basically make it as simple turnkey as possible. So whether it's the access switch or the distribution switch that are identical, the code is the same. You can actually plug in your 10 gig servers into that distribution switch if you want client access.
Uh, and then finally that distribution switch also has something called as the head end, which is basically the brains of the NSB. The head end can be virtualized in an access switch. If it's a small deployment, it can be virtualized inside a distribution switch or it can be physical appliance where you talk that 30 K capacity that we're going up into.
The key thing here is, uh, we also have physical sensors that are used to monitor the NSV, right? So with 24 7 monitoring the NSV, think of them as dedicated clients that are, you know, plugged into the wall or the APS third radio, which is dedicated sensor, that's monitoring network. And the beauty is it's an outside in view, which means, you know, you can do that simple test, that ping test from your laptop connected the NSV and validate whether what we are showing you minute by minute is, is, is the availability, coverage and capacity of the network.
Can we talk more about the sensors? What are they doing? Is it synthetic testing?
Is it Absolutely. So we've got three types of sensors. We've got a physical sensor, the one that you see there.
Uh, we've got a third radio sensor because you cannot plug these sensors everywhere. So this sensor and the third radio sensor are identical. All they're doing is a ping every 12 seconds to validate the network.
They're looking at the beacon reports sending to the cloud so that we can do the coverage and capacity planning. Uh, and then the APS third sensor also does widths and whips. So that's all fully automated.
And then the third sensor that we have is built into the switches. And this sensor is doing some, some of those synthetic tests, right? So we can actually do a radius transaction on UDP 1 8 1 2, uh, with the credentials that we provide every single minute we can do A-D-H-C-P transaction, we can do a DNS transaction, and then there are top 10 apps for the site, which we do an HGT PS transaction with.
And is, is your sensor in the ap, is it actually doing a do 11 association to another ap or is it doing that test from within inside the APS network perspective To the neighboring ap? Does the neighboring ap yes. It does not have the same ap.
So it Associates gets an IP address and does Exactly. Okay. And you'll see that in the next slide.
On on what, you know, why we have different subnets Are, are all three sensors deployed in an environment? Do you make a decision or does the end user make that decision? Right.
So the aps obviously have a dedicated sensor. They're turned on automatically. Uh, the sensor in the switch is obviously always turned on because it's, the switch is always there, whether it's an access switch or distribution switch, it does not matter.
Uh, and the physical sensor is where there is some determination. We have some guidance trying to put in the perimeter of the building in conference rooms in locations where, you know, you don't wanna put in a dorm room, for example, right? Someone's gonna kick it off.
Yeah. So yes. Do you ever see or foresee not having physical sensors and just absolutely completely leveraging AI to do everything?
Yes. So we do, we do have installations where there's no physical sensor at all, just because you can't past them. But then the aps, uh, built-in radio, uh, does that physical sensing.
And what do you think about putting a sensor like on a device? On a, Yeah, that's something you, Sure. Well, I think that question comes up, frankly, I, we talked a lot of IT teams.
It's another nightmare. There's so many agents on your laptops this another nightmare for one is operation side of it. Second is that could be your security gap.
Frankly, so many CISOs they hate, uh, adding another agent for that. But instead, we are able to capture a lot of information using the sensor data and the data that we receive from the clients. We are able to access a lot of that information.
So, and also, you know, even though we may not have at a particular customer site, but we have the sensors across all our customer sites, we are able to gather that data to understand the client behavior and continue to fine tune it. Perfect. Awesome.
Thank you. Sesh, Are you aggregating that sensor data across your clients then and able to do things like regional, regional information? Hey, we know these, these offices from these different clients are in these areas.
There's a problem upstream, it potentially regionally. Are you doing that aggregation across multiple clients right now? Or is it just single client, single site?
So Few things, right? One is the data the sensor collects that input can go into multiple places. Like channel planning is one thing where we do, uh, if something is down, those are all real tests that are happening in every single environment.
So for example, this, you know, if you have a D HCP server that's global, if you cannot reach to that, you cannot reach that DP DH CCP server. One site does not mean the DCP is down, right? So we actually have that data coming from every single site, uh, to kind of show you that, hey, that DCP server is accessible or not.
So there are two tests we do. One is, is that DH CCP server itself available or not, right? From transaction perspective.
And the second is to show you the view. Is it accessible from the site? Right?
So, but, but we do real time tests every single minute from every single site and show you the data from every single site. Sure. I think to answer your question about aggregating across the tenants, right?
So the idea is yes, you know, we, that's how we do a lot of channel planning simulations in the background. So, you know, like universities may have certain type of enrollment, distribution centers may have different type of enrollments, right? As a warehouse, as the stock goes up.
And so we do simulate those kind of inform, you know, per, uh, deployment type, for lack of a better word, right? And then apply those logic based on the deployment type. So one question that we always get, Hey, you're a cloud company.
What happens if the internet goes down? Does everything go down? The answer is no.
Right? Once you set the network up, uh, if you have an on-prem radius, on-prem DHCP, everything just works fine. New clients come in, they connect, they can pass traffic.
So nothing goes down. You only lose that control and visibility from the cloud. But, but the industry itself stays up.
The next thing is wireless redundancy. I mean, we do not want escape eau, right? Or haina.
We wanna do a proper, uh, design. We wanna make sure we can, you know, pull all that design into Nile. Uh, and that's where, you know, we were talk about the digital twins because we have all that information.
Uh, and then in an automated fashion we try and figure out a salt and pepper design, meaning two neighboring aps should be on two separate switches, right? We don't want them to be the same switch. That way when you do upgrades, we at least have that coverage.
And if something goes wrong, you know, that AP can at least cover, uh, the entire area. And then wide redundancy, again, that's the whole ring. So if you look at our, our switches, you have a deck cable in the back might look like a stack.
It's not, it's just a ring running OSPF. That way if one switch goes down, we can easily work around that issue and have multiple paths out of a closet, if you will. We also run OS VF on the uplinks.
So the uplinks can go up to a hundred gigabits per second, uh, uh, on our distribution switches. So every distribution switch can do two uplinks. So we can have a total of four uplinks from an NSP.
It's an active active mode. And the beauty is you don't have to set OSPF on the N site. You set it on your router and your firewall, uh, through how the messages we learn, we self convi configure ourselves.
And that's how the network comes up. And we'll double click on, on that whole setup in a minute. So very quickly, I wanna touch base on that L three fabric we talk about, right?
Why did we go about doing the L three fabric? So as you know, VLANs were already designed to control broadcast storms, but they got morphed into segmentation, right? And then the whole policy downloadable user roles, all these things came along with it.
But if you look at it by default, if you have a VLAN on a switch, it's already bridged, right? Unless you start doing private VLANs or doing IP access lists, which very, very few customers already do. So it's pretty much wide open and that basically enables lateral movement of traffic.
But if you really wanna campus zero trust, you need to start, you know, locking all that stuff down. So loops for example, we don't even have STP 'cause we don't have that traffic that goes through, uh, you know, within the network by default, every port is secured. And we get into more details on how we do that.
If you look at, you know, identity, there's no identity really tied to vlan. N yes, you can say this, VLAN n is for my employees, but once you get on that network, we just go anywhere and everywhere, right? If you can get on that network by just plugging into wire port.
But again, we bring identity on every single, uh, wireless land as well as on every single wire port. And we, we chat about that in more detail. The more important thing is once you're on the network and you're authenticated, how do you apply policies?
Do you apply it based on subnets or can you do that based on true identity? And again, we touch base on what we do on micro segmentation, uh, in a few slides. So very quickly, whether it's, you know, a two access switches connect to 10 aps in a small office suite, or whether it's a large campus where you have distribution switches, multiple rings of access, switches coming to distribution switch, the effort it takes to really bring up the network is the same, right?
So we really don't have any config on our access point and switches. They're truly beep and scan. The way we do that is, you know, we have a mobile app, uh, and the installers are empowered to use that mobile app.
Customers like yourself could also use that. But what we are really asking you is for an uplink ip, uh, one for this distribution switch, uplink IP two for that distribution switch, and a couple of subnets. That's all.
Once they give us that information, the entire network, whether it's again, two switches and 10 aps, or you know, 500 aps with, you know, a hundred switches and distribution switches, it's all the same. So what happens is that config gets pushed through Bluetooth. Bluetooth is only turned on in two scenarios.
One is factory default state, and the second is if something crashes on that device, right? Because we don't have any console ports, we don't have any SSH into these devices. So that's the only time Bluetooth is turned on.
You need to have a job, you need to have a username password. So it's completely locked down from that perspective that no one can just log in and and grab the details. So sorry, is There a way to force that Bluetooth on Yes.
Manually on the box That is yes. If need be, we can, we can certainly, assuming it's online, we can force that. And not only that, the beauty is sometimes, you know, remote site and you don't even have the mobile app available.
So we can actually turn the Bluetooth on, on a P one to connect to AP two if AP two crashed, right? And then pull that last breath of logs if need be. But if there's a stuck state of some kind, um, Bluetooth will automatically turn on.
But it, it has to recognize that there is a stuck state. Yes. The Bluetooth, yes.
So that's how it's been designed. Anytime there's a crash or anything that's stuck, uh, and we can't get to it, it is designed to basically turn Bluetooth on. So for example, now if, if the internet goes down and, and it does not come back up in, in a certain time, we might turn on Bluetooth on the gateways, right?
And obviously that is physical access that's needed to a closet. Great. So once that is done, what we do is, uh, we have the NSB gateways, gateways, nothing but a role in this case, the distribution switch is talking to that edge.
So it is a gateway. It could be the access switch that could be acting as a gateway, but gateway is a device that connects to the upstream router or firewall. They're running an active active mode.
Uh, we basically have them being the default gateway for all the N elements. So all the traffic is gonna, you know, go to them and they're gonna route it out. And since OSPF is con is configured, you know, all the traffic is advertised to the Epstein router and firewall and that configuration is coming from the firewall, router, whatever you set up, uh, from a config standpoint in a firewall router, we learn that through hello messages and configure ourselves.
Yes. And, and this model, are you guys also providing that edge and firewall? We have built something called the edge service, which is, which will basically be that, uh, edge.
And that would be become the antenna, be connecting to the internet. So you could have a couple of up links and a disaster recovery. Okay, like a rad point that would connect because You made it sound like it was something that the end user providing.
And So both options are available. So if it's a large campus, we would connect to your existing router firewall, right? You may not want something from us.
If you're a small, you know, mid market customer segment, you probably just wanna replace the entire stack with Nile. And that is something that we have as well. Now the key thing here is, uh, once we set all this stuff up and we pushed that config, there were two subnets we asked you for.
We asked you for NSV subnet, which is for the APS and the switches, and we asked you for a census subnet, we automatically start running a DH CCP server in that distribution switch or access switch if it's, if it's connected upstream. And this DH CCP server is only used for elements. So you are not really using this for your client devices.
We have the DHCP service in the cloud that you could leverage if you wanna use Nile or you could use your own DHCP. But this one is turned on automatically. We do not configure it.
You do not configure it. All you do is give us subnets. And once you give us those subnets, we basically start advertising those, uh, uh, those IP addresses.
So as in AP or switch plugs in, it basically gets an IP and the census subnet is different. So that AP actually has two ips, one that it gets on its physical wide interface and the sensor also gets an IP address which it can connect to, you know, to a separate ap. And that's the reason we have a separate census subnet because we like to treat it like a client.
Alright, so this is getting into a bit more details. How does this all come together? I think we are very familiar with what the controller based architecture is from, from an AP perspective, you never, ever configure VLANs on on those aps.
You don't configure SID on those aps, right? You literally just plug them to the network, they discover the controller. The controller then you know, pushes the config, whatever's needed.
They advertise the beacon. When the client traffic comes in, all they do is just send it through that tunnel cap app or IP SEC to the controller, right? That's how it is.
So obviously we do the exact same thing with the aps, but we've extended that to the access switch. So every wired port is creating a tunnel to the head end, just like the ap. And that is why you do not need to have a VLAN on that switch.
That is why you do not need to have any configuration on that switch because all the traffic from that port directly gets tunneled into the head end, right? So if I have two devices on the same subnet, that traffic is all gonna go to the head end and the head end can't decide what to do. And that's how we can implement zero trust because you don't need to configure that port.
Not sure. In some cases you want a lock to port because you have a phone and it's tied to E 9 1 1. You have the ability to do a port level config, but by default there is no port level config, there is no segment, there is no vlan that is no information that the switch knows.
If that switch goes bad, you just chuck it out, put a new switch and you're good to go because it's not doing nothing. But just creating those tunnels directly into that net, into that ds. So whether it's a wireless client that hits the ap, AP pushes into the tunnel, the NVGR tunnel, it's a wire client, it goes to the switch, it puts it to the tunnel and it goes through and at the head end you decide what should I do?
I'm gonna talk to Radius and identify the identity, or I'm gonna basically go and look at Macau and look at the identity and put it on the right segment. So that is why no VLANs exists within the Nile network. And how are you creating that tunnel?
Is it proprietary or is it a standard? It's automatic. N-V-G-R-E.
We just, as soon as the cord comes up, discovers the controller gets an ip, it'll build that tunnel. And, and how much of capacity can nest head ends in terms of throughput? So throughput, it depends, uh, whether it's a virtual headend or whether it's a physical head end.
I don't know the numbers too. Does anyone know the numbers for the capacity? Sure.
So we go up to 25 gigabits per second Rupert. So we have the controllers depending, so this is where when we do the sets survey, we understand the requirements, scale improvements, and based on that, the network is automatically designed with the appropriate capacity. So you, so you'll put in multiple head ends with you if you need more than 25?
No. So Typically you have only two aheads as you showed the primary active and standby. And each one can handle up to 25 gigabytes per second throught.
If you need more than that, you can have multiple Nile service blocks. These are like Lego blocks and you can interconnect them as needed using a next level. But you can have multiple Nile service blocks to support larger scale.
So today we can go up to 30,000 clients and we gonna increase it, but if you need beyond that, you can deploy multiple of that. Yeah, you'd Okay. Deploy multiple had ends would you?
To get, Uh, multiple night service blocks. Oh, okay. Yeah, that's right.
And what about the other scenario where you would only need like a couple of aps per site? Do you have to deploy the entire, So as you said, you know, we have a built things, uh, head. This headings are typically built into our switches, so you don't even realize that you have a separate head.
Yeah, so Like a switch CT right here, the virtual head, it can be inside the access switch. Okay. And if you are even doing the internet gateway functionality, that has also been the access switch.
So you might just have two, you know, physical hardware appliances, right? Which is what you'll see here. Or a 16 port switch.
We have a 16 port variant of this. Uh, and you might have two aps. That's it.
Yeah. Like a P OE switch. And then in that everything enabled in one, I think I, I missed it, but the tunnel for the ap, does it just get set up with one head end?
Is there a, a backup tunnel that goes to the other head end? Or is it a stateful type of failure? So, so right now this is, this is an active standby, right?
And the aps do have a standby tunnel. And yes, they would do that From that endpoint perspective. 'cause I know other manufacturers that not, I'm not gonna say they do what you do, but the other manufacturers that are doing kind of land-based microsegmentation throw a slash 32 at the, at the end point because they don't own the full network stack, right?
So they can't control it otherwise what is the client thing right? During This? So, so client this, this can be on a slash 16 or slash 32, we don't care, but we treat every client as it is a slash 32.
Okay? Right? So it is isolated no matter whether you're in the same subnet or different subnets, we just do not allow any traffic by default.
You have to put a policy to enable that. So if you want employ an employee to talk to each other, a policy has to be created. Yes.
And, uh, quickly to add on to royal's common question. So it's a state full switchover if you ever need to do it, it's a stateful Switchover. And then what were the limitations again for the, So we can go up to 30,000 clients and we can go up to 25 gigabits per second.
Okay. And how do you interconnect? So if, if I need to deploy multiple Nile service blocks to get the capacity, do you interconnect at the head end side or are you expecting it to go up through the client's edge and then back in edge?
Yeah, so I'm gonna have to build capacity on my firewalls to port capacity and all the rest or create, uh, a core block outside of your core block. That's right. That I manage to run that then interconnects to my, okay.
Yeah. Alright. What we spoke about was the setup for Nile, right?
Which is almost zero touch. That's not zero touch, zero config, right? We talk about zero touch in a minute.
That is what it takes to bring the network up. It's like that Uber car that's waiting for you, right? But you still have, you know, control over, you know, how you are going to, what passengers are you gonna bring along?
How many suitcases, what a destination is. And that's what we like to call the context, you know, setup. This is all done through Nile portal, so you will set it up.
What we're trying to get to is we wanna make it as intuitive as possible without the need for groups and profiles. And then to figure out which AP gets which group, which profile, as you know, that's still cumbersome, uh, licensing that's required. You have to claim a license, you have to, you know, doing RMA, you have to figure out how to, you know, unclaim that license.
So with Nile, its, it starts with the customer telling us what service areas they want, right? Service areas starts with a site, a logical construct. So Santa Clara could be a, a site, if you will, uh, in site you could have multiple buildings.
So this particular building could be an actual building in there with an address, a physical address, a building has a floor, and then on the floor you can actually draw what we call a zone, right? So here, I'm, I'm drawing a zone, I'm falling the guest zone. Uh, and this is what service areas are.
So it's very straightforward. Just create a simple site building floor and a zone on, on that. Now, once you do that, all the other constructs from a setup perspective are tied to that service area.
So I can create a DH CCP server and say this DH CCP server caters site one, right? Very logical, right? I have a D HCP server in Bangalore, I have a DHCP server here in San Jose, I've got, you know, one in London.
So you just create your, your resources if you will, and map them to a geo. You can easily create one single DH CCP server and map it to all geos if you have a global DSCP server. Uh, same thing with Sid.
You decide where do you want those? Sid I want all my employee, I want my employee societies at all my sites, right? I want iot only in, in site A in this case.
Now, once you basically do that, when you plug an AP in, depending on where that AP is or headend or switch, uh, because we know the X xy, when we do the install, we know exactly where this AP is going. So if you install this AP on floor one, which is outside that, that guest zone, it'll inherit the employee society and the IOT society automatically. The beauty is this AP dies.
You just go there, you click on the floor plan in the mobile app, you deactivate that ap, take a new ap, plug it in, and you're good to go. There's no claiming licenses, there is no pushing profiles, right? There is no group information that you need.
It is as simple as just knowing where you wanna install that, which is already prebuilt as part of the mobile app. And that's what the digital twin is very critical because we need to know exactly where every single device is, whether it's a switch head in distribution switch or an ap. So in this example, the AP two is, and a P four, uh, are gonna have the guesses Society, because they are in that zone.
The XY falls in the zone that you had created on that floor plan, and it automatically inherit stack effect based on that. So we basically land out with, you know, no profiles, no groups, uh, and it's, it's very, very straightforward. I'll quickly touch on key capabilities and then I'll hand it over to Avinash.
No, who's going next? Uh, uh, ish is going next. Uh, and we can spend some more time talking about this, but we wanted to highlight, uh, some of the automation that's already built into Nile, right?
So one is we obviously have Deepak inspection, multiple ports protocols, study multiple protocols and applications is what we can figure out. Not only that, based on that information we can do automated quos, right? So we know this is FaceTime audio, we know this is FaceTime video, we know this is Zoom audio, zoom video, and we can essentially mark the packets with the right DSCP.
So when it hits it outer or the firewall, those markings are are in place, right? Is that Deepak inspection full stack? Uh, it's on the AP or switch or, or at your gateway?
It's all on the head end. All on the head, yeah, because all the traffic just comes to the head end and that's where we apply the policies. Okay?
Yes. Uh, a lot of people ask us, Hey, you are layer three, what about layer two? Right?
Uh, what about that passive client? What are multicast? What about broadcast?
By default, everything is blocked. And in the portal you'll actually see a list of every broadcast traffic on a network. You'll see every single multicast on a network, uh, and you can allow them or disallow that, right?
When it comes to, you know, IGMP snooping, we can actually completely in an automated fashion, you know, snoop those packets and try and connect all the devices without needing, you know, a thousand pages of config. Yes, sir. Oh, I was just gonna see if you'd talk a bit about M-D-S-M-D-N-S.
Yes. So MDNS discovery is automated across segments. Uh, and it's also location aware, which means that if there's an airplane device in building a, that airplane device is not gonna show up in, in building B.
So it'll be context aware based on where you are, and the advertisement would be available on every segment, uh, within that network. But when you connect to it, it, it is basically a, uh, you know, an L three connection, right? That's where the policy comes into play.
So even though you can see, it does not mean you can necessarily connect to it if your uh, you know, admin has not permitted that. Yeah. Well, quick starting about the deep packet inspection, is that something that the customers have access to?
So, couple of things, but first of all, you know, we have a lot of content. So can we hold off on some of the questions till the end of the session and then we'll definitely go into the deep packet inspection, the capabilities and like if they have private applications, how would they get access to that information so that you can do the pro quality of service.