Engineering Campus-Wide Mobility: Arista’s Scalable Wi-Fi Roaming Design
Designing for large Wi-Fi roaming domain especially in environments like university campus has many challenges. In this video, Arista’s own founder and CTO, Ken Duda will talk about how we applied some of our learnings from large scale data center & AI cluster deployments to solve Wi-Fi roaming and thus unifying our wired and WLAN data plane fabric.
Presented by Kenneth Duda, Founder and CTO. Recorded live at Mobility Field Day 13 in Santa Clara, CA on May 8, 2025. Watch the entire presentation at https://techfieldday.com/appearance/arista-presents-at-mobility-field-day-13/ or visit https://techfieldday.com/event/mfd13/ or https://arista.my.site.com/AristaCommunity/s/article/Roaming-in-Wi-Fi-Networks for more information.
Transcript
My name is Ken Duda, one of the original founders of Arista Networks. I run the software team at the company. What I'd like to talk to you about today is layer two mobility scaling layer two, say it, it is so fascinating.
We're doing this, you know, because I feel like for 50 years, layer two and layer three have been duking it out, and it's always the pendulum swings back and forth. And I, I thought we had this nailed. It's all gonna be layer three from here on out.
And then more reasons for layer two show up. And we are now looking at scaling layer two to the next level here. So one scenario here is you've got a, a v switch on some server somewhere with a bunch of VMs, probably have more than one of them, and you connect them up to some top of rack switch, and you probably have a whole bunch of those and connect that whole thing to some kind of network interconnect to got all these VMs running workloads.
But, you know, applications have layer two requirements, right? Mac broadcasts, or they, uh, have hardcoded IP addresses in them or th apples or are, are fixed. Whatever reason you end up with these layer two kind of requirements on your applications, but you also wanna be able to put any application anywhere in your data center.
So you need layer two mobility across this entire thing. This is mobility field day, right? I'm supposed to be talking, oh shoot, this is the wrong mobility.
Lemme try again. Okay, so you got an access point here with a bunch of wifi clients and, uh, and you probably have more than one access point, and there's clients connected to them. They connect up to some kind of a POE switch somewhere, and then, uh, those switches all connected into a network look familiar.
And you need layer two mobility across this whole thing. Why do you need layer two mobility there? We'll get to that in a sec.
But, um, first let's notice how similar these problems are. And this is probably, I think, one of the best cases for Arista's approach of having one e os one software architecture, one feature set, one binary image running across all your switches. Because while data centers and campuses are different in many ways, they're also similar in many ways.
And the having all the diversity, you know, diversity may be the spice of life. It's not the spice of your infrastructure, okay? Or at least not the kind of spice you want.
And, uh, having the consistency of EOS here is just a tremendous advantage. The scale targets we're after here, we're looking at, you know, environments that wanna have 300,000 wifi clients with layer two mobility because with one layer, two domains stretched across the entire thing, and potentially a very large fraction of that 300,000 all in that one layer two domain with tens of thousands of access points. And so what, so like, haven't we learned this lesson?
Like do we need to relive the broadcast storms of the 1980s, right? Like, how many of you were working on networks in 1980s? All right, um, yeah, well here we are.
And it's, it's 2025 and it's still there too because, um, in the data center, of course you need this for the reasons I already said, but in the campus, you also need this because you need rapid roaming and clients and client applications just don't do well with client devices changing their IP address. It causes sessions to drop voice calls to drop need to reestablish, and customers don't like it. They want to be able to take any device anywhere and get seamless connectivity across the whole thing.
So this is the, the, the, the key motivation for providing this very large layer two domain mobile across a very large number of access points. Ken, a quick question. Are you finding in the sales cycle arista's credentials in the layer two data center is helping the, uh, the campus sales, the mobility?
Oh, a hundred percent. We, we, we've had, we've had, we had customers who wanted our reliability before we had POE switches who were using POE injectors on our data center switches in their campuses, right? Um, so we've, we've very much get pull from, uh, the data center side, but we're also seeing some, uh, campus first customers now, which is really exciting for, for, for me because, uh, uh, I feel like, um, that we're at the inflection point we're really gonna become a major player in, in campus.
Alright, thank you for that. Let's see. So the, so the, the, the question is then, well, how do you provide this large mobile layer two edge unknown DA flooding?
Really scary. You're gonna flood across 30,000 access points. Um, you are gonna solve this with a proprietary control plane in a proprietary data path and lock the, try to lock the customer in.
I mean, we, we, I really objected all the proprietary single vendor nonsense. Um, are you gonna turn every access point into an EVPN speaker? You imagine having 30,000 EVPN speakers, how many BGP sessions, route reflector mesh you need?
I mean, my God, how do you, how do you do this? Well, our solution here is two pieces. The first piece is Vespa, virtual ES with path aliasing.
So the Vespa picture is you start with some kind of a, a gateway device. This is a kind of a high-end router, um, pole line rate, a hundred gig, 400 gig kind of, uh, device. You have, of course, more than one of these, uh, in a, maybe they're in a data center somewhere.
And, uh, these things all coordinate with EVPN. There's probably a smallish number four to eight, that kind of a, that kind of a range. We're not talking, you know, thousands or even even hundreds.
Uh, it is probably less than 10 of these for a very, very large environment. And we assign to each pair of these gateways, we assign a vip. And in this picture here, it's 10 0 0 and read my own slide, 10 0 0 5 and 10 0 0 6.
Uh, and this is a V that terminates VXLAN by advertising the V into the IGP. And then sort of, sort of basically any cast and any cast VXLAN termination point that connects to your IP fabric. That's the little cloud thing there is the, is an IP fabric, and then connected to that are all of your campus P OE switches and all of your access points and all of the clients.
Now on the access points, we're not gonna run EVPN, but we are gonna use vxlan. So every access and IP address and serves as a very basic data path learning v tep, um, and, uh, it's configured to direct all traffic to one of those two vs. The 10 0 0 5 or 10 0 0 6.
Um, and so, sorry, I I I don't have a slide, my slide doesn't capture this quite right, but the, the way to think about this is picture the VXLAN tunnel, sorry, I'm probably off the camera. 1 to the top switch and one to the bottom switch on the left. And those two VXLAN tunnels, you can think of them as a logical link.
Each one is a, is a logical link terminating on two EVPN speakers. What can EVPN do with two links? It can turn them into an ES and ethernet segment, right?
And now traditionally this is only done with wired ethernet. So you've got a, a couple of ports coming outta your routers, you put them together into an es and then you connect that into some other switch network. And now you can do, uh, multi active active multihoming treating that that pair of links as a, as a lag, right?
So it looks like a lag. On the other side, we're doing the same thing, but with VXLAN tunnels as the link. So these guys are using VXLAN with EVPN, but this from here to here, these links are outside of the EVPN domain.
They are links in EVPN and together form an es. And so that means that when you, the packet comes into an access point, it takes the packet encapsulates in vxlan sends it to the configured address, 10 0 0 5, say, it then gets any cast routed, maybe it hits the lower gateway. The lower gateway then does some key things.
If this is the first packet it's seen from that access point, it creates a type six ESI for the access points IP address. This is a dynamic ESI, you don't have to configure your ess, we learn the ess let us sink in for a second. These are the interfaces in EVPN being dynamically learned based on incoming VXLAN traffic from VTEPs that don't even have to be configured.
And so we created a new, uh, type six, uh, ESI for this. Um, and then we learn the, the, the, the gateway. The gateway switch learns the client Mac on that ES then advertises a type two route into EVPN, uh, and with where the, the MAC address is assigned to that ES with a next desktop being self.
And now all of the switches know how to get there. Um, and the, so that means that if traffic comes in on the right side, it will get VXLAN tunneled because of EVPN to the left because that's what the next hop is. And then from there out to the access point, you say, well, that seems like two hops.
Why do you, you could go straight, couldn't you? Well, there's table scale limits and physical switches. They can only deal with so many tunnel endpoints and so many encapsulations.
One of our platforms has a limit of 8,000. And the customer's like, but I've got 24,000 access points. Well, three repairs and you're done.
So, uh, this, so this gives you really good scalability. Uh, then when traffic comes, say from the internet or something and goes into the upper gateway, the upper gateway looks up the MAC address and EVPN sees that the ES is an ES that it itself carries as well. So it just immediately turns around and tunnels that traffic straight to the access point.
This is EVPN, active, active multihoming right here. So this is a way of using standards based technology to solve a scale layer two problem. This gives you this enormous L two edge, a plain simple layer three underlay, um, get transparent mobility across the, the whole thing.
There is zero unknown MAC flooding in this environment because you're using EVPN to exchange MAC address information. The switch is simply drop. Uh, unknown das whole thing is standard space.
The only EVPN edition here is the treating of VXLAN tunnels as virtual ess, but this interoperates as it is with any standard EVPN speaker. Now, the, obviously they aren't, if they're not doing the virtual ES thing, they cannot themselves tunnel directly to and from access points. But if you have an extended, uh, uh, EVPN environment that includes layer two stretch to other wired devices, this interoperates with that completely transparently, right?
'cause it's standard EVPN, the other switch doesn't need to know what the virtual ES means. Um, and also there's, uh, no changes on the access point here. This works with essentially any access point this capable of serving as a, like the world's most basic vap, like send all traffic to one fixed, um, VT a ip, and then take any, any vx uh, any VXLAN encapsulated traffic, you get dcap and bridge.
That's all you need outta the access point. So it works with any access point that provides that basic, uh, VXLAN functionality. Any questions, comments so far?
Alright, so now, um, some remaining limitations we're not done yet because we have some, some remaining limitations. We do not like flooding in these environments. Okay, boy, layer two loops in an, in an environment with, we've got, uh, 30,000 access points wholly moly, and, uh, um, so we don't like art flooding.
We also have a, a limitation here that if the, the client needs to speak when it's roaming because there's nothing else triggering the EVPN type two route to get updated. Um, there are still table size limits in the EVPN gateways in the hardware switches. We have limited mac table capacity, but worse, the, the killer is this one.
Um, some of our products have a L three f**k limitation that when you're routing into a subnet, we only have so many rewrite table entries to rewrite the MAC address to the client's destination Mac. And if you have, you're trying to support 300,000 clients, you're gonna run out of table capacity there real quick. So, um, we have a solution for this, which requires no changes at all to the switches, is entirely on the access point side.
So the beautiful thing about these two in innovations Vespa and MRO is they're both, they're completely orthogonal. So they, they work with standards based on the other side. Um, MRO is MAC rewrite offload.
And the benefit here is the is is, as I said, is only an AP behavior change. There will be no more flooding period. We're done with MRO and switches will only need a rewrite per ap, not per client 30,000.
We can do 300,000. We can't, okay, so how does the MAC rewrite offload work? So same picture as before.
I'll put the diagram here in the lower right and just show the packet flows. So here's part one. This is what happens with MRO when the, the client associates, the client associates, and then the first thing a client typically does is sends an ARP for its default gateway.
The access point does not forward the arp. The access point responds immediately with a configured gateway mac. There's a, there's essentially a virtual Mac configured as that, that, that represents the routing function out of this, uh, entire stretch layer two domain.
Instead the access point sends a garp, a gratuitous ARP response that contains the client's ip, the APS own Mac. And the switch responds to this by learning the aps Mac, not the client's Mac. The whole point here is the switch never sees the client's Mac.
We're gonna offload all the per client Mac processing onto the access point. And so the, so we only have to invest to learn the APS Mac and then update the switches ARP cache because of the GARP recording the client's IP getting rewritten to the gateways Mac. And so that means that all clients on that AP all have the same rewrite.
So I can share that switch rewrite entry across all of those different client aps because client ips, because my problem, remember my problem wasn't the, I have plenty of lookup capacity on my switches. I can put all those client ips into my, uh, you know, my, my exact match all three table. It's the rewrites I can't do, but now I can share the rewrite entry across all clients attached to that ap.
It sounds like MAC routing almost, Um, like the ap it's, it's sort of like art proxy. Yeah. It's because you're using one one Mac address to represent a whole bunch of different servers, but it's sort of, it's sort of like art proxy, I think.
Yeah. Um, so now the, the, uh, uh, the last three things shouldn't be there. So then to, to illustrate the data path, now that the client has successfully AED towards Gateway Mac and sends a, um, now the access point rewrites, this is the part of the MAC rewrite offload, it rewrites the Mac source of this packet.
Instead of using the client's Mac, the AP substitutes its own Mac so that the switch doesn't need to learn any of the client Macs. And then the, uh, access point via VXI end caps as I described before, sends it across the layer three underlay the, uh, switch dec caps and routes to back it out to the internet. And then the server re some server on the internet responds with a syn.
The, um, switch now has to IP forward that to the client's ip and it hits the RP entry with the client with the APS Mac. So it gets rewritten to the, the switch in hardware. VX line encapsulates the packet, does the Mac rewrite to the APS Mac, and, uh, and sends a packet across the layer three underlay.
The, the, uh, the access point. Obviously Decapitates terminates vxlan Deencapsulate, uh, finds his own Mac as a destination, looks at the, uh, IP destination and says, oh, that's not for me, that's for this other guy. And I think that's what you're talking about, right?
Is sort of, sort of a routing function, kind of. It's a very primitive, very, very tiny part of IP routing is actually taking place in the access point here because we're, we're doing essentially the MAC rewrite step of the IP forward action in the access point. The ip, the access point looks at the IP destination, knows which client that is, puts the client's own Mac into the packet and then sends it out over wireless.
There you go. Any questions, comments? It makes sense.
Yes, sir. I'm wondering how do you handle roaming? Aha, that was the next, guess what part three is?
That's what I've been waiting for right here. Part three is rapid roaming. So when the, uh, our access points, uh, run a control protocol, uh, it's a control, it's a controller list, peer-to-peer control plane protocol for exchanging which client is au is authenticated, signed into which, uh, um, uh, which, uh, SSID with what IP address, all that information is being exchanged.
It's very normal for to support rapid roaming for somehow for that information to be distributed. Maybe it's through a controller in a lot of our competitors' architectures, but in our case it's peer to peer among the access points. So the new access point when the, when the, when the rapid roaming client associates with the new access point, the new access point already knows the client Mac and the client IP Make sense so far?
How does, how does it know it? Yeah. Through Our control plane protocol.
So we have a control plane protocol that runs between the access points. Yeah. That exchanges information about authentication about, uh, client Mac and client ip, uh, VLAN assignment, uh, SSID assignment.
So that's already distributed across the access points, but how, by the existing, How does that scale to like 300,000 clients across 20,000 aps in a single mobility domain? So we, uh, learn about RF neighbors over the air, mostly using the multifunction radio. Um, I say that because some aps don't have a multifunction radio.
So we learn about our RF neighbors and we share, uh, on the wire in a uncast fashion. All the, uh, the data that we need to have for seamless roaming and, and other things we only Share with people who are RF Reachable. That's right.
That's Very clever. It's not see scale. Great answer.
So that's how you're passing your R ones around is based on that. That's right. That's right.
So speaking of scale, how is M-D-M-D-N-S handled, which tends to destroy the airtime on network? Another question I don't have an answer for. Okay.
I imagine the answer would be we would proxy it, but I don't actually know the Backpack on. Um, so, uh, when the, when the client associates with the new ap, the new AP already has the, the IP related information. So part of that process, the AP sends a garp, it's the same GARP it sent when the initial association took place, but this time it doesn't require the ARC request because there's no reason for the client rapid roaming to re arc for its own default gateway.
And the access point already has the IP information, so it, it send, sends the same exact GARP with the client's ip, the APS Mac address sends that off to the, uh, to, to the VX VXLAN tunnel. And then the switch does the same thing. It vesper learns the APS mac if needed.
I mean, probably already knows about the AP updates, this ARC cache for the client and it, uh, advertises that Mac IP route, uh, into EVPN so that everybody else knows that the, the IP to MAC binding just changed. It's interesting from the switches point of view, this is not a MAC move, this is a arps a change of which IP address map maps, which Mac address, and then the data path just works exactly as before. So you may be a little scared of this.
I was a little scared when I first heard about this, about the exact order in which these messages take place. And can you get into a situation where there's an inconsistency between the IP to AP mapping in the, in the switches arc tables versus where the client's actually associated to get into the situation where the client's sending traffic and going in this direction, the client comes from returning traffic's going to the wrong ap, and so you're, you're in, in the, in for some pain. And so the way this works is we have an arc refresh mechanism where the access point notices if the client sends a packet but doesn't get anything back, and after a small number of seconds, it responds by sending, you guessed it, the same exact garp.
So this is a unicast a response. And I sort of think of this as like, this is basically like rip for post routes, right? So like rather than building a BGP stack and running that on our access point, we're essentially using GARP like RIP and just sort of saying, here's the, here's the routes I have, here's the route I have, and just keep 'em alive through refresh.
We only need to do it if we're seeing traffic flow and not seeing any response, send the GARP message. And then if, if the, in some cases maybe the client's just sending into a black hole and nothing's supposed to come back, in which case the switch looks at the GARP message and says, yeah, I already knew that and just discards it. But if in fact this things had become inconsistent, then the, the switch will update.
Its a cache based on the GARP and re advertise the route into EVPN and you're back on the air and connected. So it's, uh, resilient against race conditions, failures. Um, I'm, uh, running out of time here.
Uh, I think we are of time for your session. Okay. If you want continue, Uh, this is just a summary slide and I think I've already mentioned all these benefits, so I will, uh, I will leave it there.
Uh, any last questions? Yes, sir. You, You'd mentioned earlier that, that you could have different aps, different switches.
Yes. Are, is Arista the only AP manufacturer that supports MRO? Um, as far as I know.
And so my, my my point was that Vespa and MRO are independent. Mm-hmm. So MRO will work with anybody's switches and Vespa works with anybody's aps.