The X Series Architecting for High Performance Scale with Xsight Labs
This section provides a high-level overview of the product roadmap, specifically introducing the X-Series and E-Series lineups. It identifies the six critical chips required to build modern AI Factories and explains the concept of a “Truly Software Defined” stack that operates at full line-rate across layers L1-7. This serves as the technical foundation for the subsequent specialized deep dives.
The X-Series, an Ethernet switch, distinguishes itself through a “truly software-defined” programmable architecture, utilizing 3072 Harvard architecture cores operating on a “run-to-complete” model, unlike competitors’ fixed pipelines. This provides unparalleled flexibility, enabling parallel packet operations, recursion, and extensive header processing, including 11 layers of MPLS and various encapsulations. This design is particularly well-suited for emerging AI-centric protocols such as Ultra Ethernet (UEC) and ESON, enabling customizable congestion management and efficient in-flight packet handling. The X-Series boasts significantly lower latency, achieving 450 nanoseconds compared to the typical 800 nanoseconds, and demonstrates exceptional buffer utilization, consistently above 86% even under heavy load.
The X-Series also stands out with its low power consumption, operating at under 200 watts for a 12.8T switch, which is described as disruptive. Its software-defined physical layer supports diverse SERDES speeds (10G to 200G) and modulation schemes, enabling mixed-and-matched configurations that facilitate connections between new and legacy interfaces. The programming model, though initially assembler-based with Python wrappers and libraries, has seen customers such as Oxide develop P4 compilers, with Xsight Labs planning to develop their own. This powerful, flexible, and low-power solution is specifically designed for edge deployments, including half-rack to two-rack configurations, satellites, and base stations, delivering significant reductions in power, rack space, and cost. The X-Series product was generally available in November 2022 and has been in mass production since the summer of 2023.
Presented by Ted Weatherford, Vice President of Business Development, Xsight Labs, and John Carney, Distinguished Engineer, Software Architecture, Xsight Labs. Recorded live at AI Infrastructure Field Day in Santa Clara on January 29th, 2026. Watch the entire presentation at https://techfieldday.com/appearance/xsight-labs-presents-at-ai-infrastructure-field-day/ or visit https://techfieldday.com/event/aiifd4/ or https://xsightlabs.com/ for more information.
Transcript
I'm Ted Weatherford, I'm VP of Business Development. I've joined with John, who's our distinguished engineer and architect on our E-Series products. Nice to be here.
Yeah. And I'm gonna cover the X-ER products now. Uh, the X-er product is, uh, ethernet switch.
And what distinguishes it from all the other ethernet switches literally is that it's programmable and its programming model is a load store type model. You have 3072 little Harvard architecture cores inside this ethernet switch chip. Um, all the competition is approaching the problem with a fixed pipeline.
They may call it programmable, but at the end of the day, it's a fixed VLIW pipeline with fixed elements and programmable elements, and it has a fixed latency and you don't have the flexibility to do tasks. In parallel. What we've done is run a complete model with this 3072 little Harvard architecture cores.
And it gives us two things. It gives us more flexible than any ethernet switch chip that's available or has ever been designed with an actual architecture I'm showing here that allows us to scale down or up really gracefully to call it an FPGA. 'cause it kind of looks like an FPGA floor plan would be a complete misnomer, but it is symmetric.
And if you look at all the little squares, there's 64 of them and each of those 64 squares has 48 processors. Okay. And they're run to complete.
So when you go to program this, you're not forced into a pipeline, okay. Of sequential steps where you write the code, you turn the code 90 degrees and you drop it into the pipeline. You literally can do recursion, you can do whatever you do within your instruction space and the time you have for the minimum packet size.
We've sized all the caches, we've sized all, all the clocking so that you have this amazing low power device with tons and tons of header processing. You can, you could do 11 layers of MPLS, for instance. So all kinds of IP, inside ip, any kind of encapsulation, whether it's already a standard or something that's developed.
I'm gonna, um, so yeah, thank you. Uh, alter ethernet, things like that. These sorts of things that are starting to come outta the woodwork UA link, I'm guessing that's UA link here, but, um, how does that play in this space?
It's exactly where this shines And for alter ethernet or eon or some of the ethernet centric extension protocols that are coming out for high. We're, we're, we're, we're there. We can ship this and some of the physical layer mechanisms like the layer two retries, there's things in there that we didn't catch with this particular tape out.
But, uh, protocol wise, uh, congestion management wise, um, the PA in flight packets for the servo loops for all your congestion management, all that's dialed in and completely flexible and build your own server for any scheme. Um, so it, it's, uh, it's ready for that, uh, better than all the other products 'cause they're fixed in function. Um, and that was the one exception is to retry.
And then on the, uh, the UA link reference, this is not a UA link switch. That's correct. It's not, um, I'll just offer the latency of this switch while we're at it.
Um, a UA UA link is a very low latency switch. Uh, it's backend scale up centric. This is a front end scale out or a backend scale out a a product.
Um, and we can get down to, you know, 400, uh, and 50 nanoseconds, which is screaming for a traditional scale out switch. I mean your broadcom's rate 800 when you really measure 'em. So the, one of the advantages of having a run to complete architecture is you, you can dial in the latency to be as good as it gets.
So Ted and Jack Poller with Paradigm Technica, can you talk a little bit more about run to complete and what that means For you? Yeah, thank you. So, um, there's really three different architectures.
There's a Von Noman of Harvard and of data flow, uh, and run to complete, uh, speaks to, um, an architecture where you have an instruction set, uh, that you're running, right? Uh, and you'll run to, you complete the overall operation of the packet frame coming and being modified and being sent out. So the frame shows up, ethernet, it gets modified, it gets shipped on its way.
Uh, and what run to complete allows you to do is have complete flexibility. You don't have to, uh, handle the packet one piece at a time sequentially. You can move around anywhere in that packet header or that packet within a window.
So the run to complete is a amount of time or clock cycles. You have to do work, but that work does not have to be only sequentially. That's my level of understanding.
Yeah. It's more flexibility and it allows you to do things in parallel. And there are many packet operations that can be done in parallel.
So you, you end up saving, um, time and getting latency from it, uh, and not being caught off guard if protocols change or you have something interesting where it really matters. 'cause all of us use standards, right? Mm-hmm.
It matters With all of the instrumentation in telemetry people, the observation and the troubleshooting of these network is getting increasingly complex. And this kind of model just allows you to build the best instrumentation in telemetry, um, with, with really less constraints. Yeah.
So it's symmetric and it'll scale either way. So each one of those blocks is a 400 gig block. So if I want to go build a 400 gig, a tiny little switch, I just do one block.
If I wanna scale up to 400 terabit and, uh, and I shouldn't have said that number, um, then, then you increase the block size. Um, so it, it does have that sort of, um, geometric scaling up and down in this graceful, symmetric way. Um, which you won't see when you look inside the, the fixed function data center switches, you'll see pipelines and they're fixed in nature and there's a number of them.
I'm not gonna cover this, it's an eye chart, but I want to point out that we have large amount of packet buffer. We're very low power. 8 T in derivatives.
It's the same dye, but we down clock it. You can turn the ser down. The sies can be all different speeds.
So we're software defined even at the physical layer. You can have er coming in at a hundred and going out at 25 with different modulation schemes. You can mix and match.
And this is actually really, really, really sexy because you've got these new fabrics up there with a hundred gig er that have just been deployed last summer and they'll be around for a while. And then downward, you've got all the legacy. So you can run a 10 gig, 25 gig, 50 gig, a hundred gig, 200 gig with whatever series and whatever modulations you want.
So it really gives you this building block for looking backward as well as forward. 8 is 'cause we're going after the edge, and I'll cover that later. We're going after half rack, full rack, two rack.
We're going after satellites, we're going after base stations. Okay? Um, the performance is there, the chips are here, the boxes are out by our, our Taiwanese friends Act.
In an edge core, you can just measure everything, but you gotta show people some performance if you're claiming to be programmable and lower power than than a Broadcom or an Nvidia or a Cisco or a Marvell. So we do that, and that's what this is. It's the watts on the left and it's breaking down the certis, the core, you know, and it's showing what your max and and mins are over, over that.
8 T switch that's programmable at under 200 watts. It's disruptive. Um, this is showing you efficiency of the most important thing about a switch, besides its rad switches or connectivity at the end of the day and how much bandwidth and how many different connection points or ports you can have.
But the other thing that matters, and it really matters is the shared memory or not buffer the frame buffer. The packets come in and do they come in fair? And can you utilize, in times of rustiness, can you utilize the whole buffer that you paid for without overrunning it?
So in our example here, we take 127 ports running it at a hundred gig each, and we ram it out 100 gig port, and we find out how long till the packet buffer fills up, does it overrun. And then we do it, you know, at the different, over subscription rates and different packet sizes. And we show that the utilization never drops below like 86 or something.
And in real world tests. That's amazing. Okay.
So if you're a switch head like I am, then you're like, oh, wow, that's amazing. The, the, the tomahawk products that are dominating the market, Ted. Yeah.
Regular ese silver. Okay. Still trying to get a handle on this.
You're not actually generate, you're not actually manufacturing dus, you're actually manufacturing the chips that would go into dus or ships that would go into switch. We have two products, they're both chip products. I'm covering the switch first.
It's the separate tape out and it's just an ethernet switch chip with 120 800 gig er on, it's a, it's a switch chip. Next we're gonna cover our EERs, which is A DPU. So we have two products.
It is amazing. 200 engineers doing two products of this complexity as it's a lot. So we have two chips.
They go together nicely, they have the same ER ip, they play really well together. 'cause the highest volume opportunity is the top rack and the front end interface or the backend interface into the server or the GPU server. And so we book in that with these two products, especially for the edge, um, root chip company, two chips.
Um, I want to, uh, I'm gonna pick up the pace a little bit. This is showing, um, the fixed pipeline and the map pipeline approach, which is not ours against this run to complete full SDN. We can imitate the other architectural approaches.
Um, they can't imitate us. So we can make the trade-offs between latency or the amount of work you're doing and the amount of power you're spending. Um, so this is more of a deep dive for somebody that wants to compare these products to data flow architectures or fixed function stuff.
But suffice to say, we're trying to say that we have, you can have multiple pipelines, you can have branches, you can have, um, a physical connection and a logical connection that's flexible. Uh, Challenge with something, yeah, like this in the past has been latency. I mean, to do a, to do something that's not pipeline and on map pipeline and achieve the line speeds has always been impossible before it's Pr we're proud of it.
Yeah. Uh, I'll give you a clue. If you're designing for your worst case and you're building a pipeline, you've got a lot of stages you don't need.
If you build with us and you put our 3000 processors in, in a line, if you want to pretend it's a pipeline, uh, you can and you'll just have less instructions. Or you could have one processor handle a whole packet. You can build a pipeline like they do.
Or you can have one processor handle one flow. You have this whole range. I guess the question is what's the clock speed?
And you know how Oh, sure. How fast are you being able, are you able to maintain Yeah. Line speed across, you know, however, 128 ports I guess.
8. It's, it's mind bending. Um, and I can just say that the team, uh, is basically on their eighth or ninth generation network processor or switch when you combine the people.
We have people from Motorola, Freescale, uh, you know, Nvidia, Broadcom, um, Juniper, Cisco. It's a pretty senior strong team that's been doing network processing, DPU and switches for, you know, 30 plus years. Yeah.
Now it, it is, you know, if you, before we had the products, it can be a lot more, uh, interesting debate. Now we just have the products so you can just put 'em on the test and test. In fact, we have built in self test that we don't advertise, but you, that's a really nice feature we have too.
8 terabyte of, of line rate. Every port has, its built in Xia tester. Um, so you could even test the device in its own print circuit board without an expensive $3 million tester.
So programming model is always the challenge for programmable products. We have an assembler that we've wrapped in Python and we provide libraries. You've gotta configure the tables for forwarding the tables for security, the tables for quality of service, the meters, the counters.
You have to set all that up and they, their structures. And we have libraries and you have to program this thing in our assembler. However, we opened up the instruction set and our first customer, which we did a PR on, uh, called oxide, uh, they're a local, uh, cloud as you know, in field.
Yeah, they, they, they wrote a P four compiler on it. So this is a simple risk instruction set. I shouldn't call it risk technically a Harvard architecture, but it's a small little instruction set that you'd be familiar with if you're a, a programmer, especially somebody that really programs deep and they just wrote a compiler on top.
So our whole ethos is it's open, do what you want. And we also have plans of putting out a P four compiler early next year also. 'cause we have a lot of customers that really want that, that higher level or what I would call fourth generation language.
It's kind of dated terminology. But, um, so today we give you courses, we've got all the examples, um, and we haven't had a customer we had to write all the code for. They've all taken the classes and written the, written the code and uh, and uh, done quite well.
And one of them is, uh, all foreshadow is SpaceX. So we're really excited about that. Um, the architecture and the normal stack of what you target to on the very bottom, I'll start there.
You've got the switch device, that's what you really care about. But you could also do simulators and you can also do, um, hardware emulators. We built our own hardware emulator.
We have a room full of FPGAs of design. We did ourselves, we do all our own hardware emulation. We don't buy hardware emulators.
Uh, we have our eval boards, uh, which are just, uh, rack systems, uh, you know, pizzas of boxes with front panel ports like you'd see in a top rack switch or a fabric. And then this just shows the network operating system down and what we provide, um, and we put all this, you know, on open, anybody can get to it. Uh, and the network operating system of choice for all of us now is Sonic.
Uh, and we do our own sonic distribution and then we have two partners that provide hardened sonic as well. Um, so that's what your, your normal stack looks like. And this is what a box looks like.
I have it right over here. I just wanna say it's real, it's available. This particular one, um, uh, is what we call the universal switch because the MPA connectors are Q SFPs and you can put whatever you want in them.
Each little rectangle is four C days and those SER can run at whatever speed you want and there's common media for four by 25, 4 by 50, two by 50, et cetera, et cetera. So that you can build a top rack here, which we call a, a top rack upgrade or a TOR upgrade so that again, you can connect to any kind of speed up and any kind of speed down. You can migrate from older network interfaces to newer ones or maybe one storage box has a certain kind of thing and it's in the same rack with a server or a GPU server.
Got a got a question. Yeah. Specifically about the, the, to here.
Uh, do you think, and you can theorize here a little bit, do you think once 2 24 UR comes around, you'll be able to stick with the same form factor, low power and not go liquid cooling? Depends straight up on how much bandwidth you want. Straight up bandwidth.
6 T, you could stick with this. Yeah. 2 T would be harder.
Yeah. Uh, it'd be harder. Um, that's, I that I'd have to defer to a switch expert.
Um, but that's the answer. Okay. No, that's, yeah.
May be maybe 51 too, but Yeah. But, And, and our architecture, um, is really on par at a geometry level with the others. So if we do a hundred terabit switch right.
It, it's gonna, it's gonna be a thousand watts. Yep. So what we've done to just be so disruptive on power Now, full disclosure Sure is we went to five nanometer when everybody else was Still in the Old place, was going forward with very large rated switches.
And we did it to capture the edge. Yeah. The economics of the edge, the power of the edge, and the right amount of connectivity for the edge.
Um, I got an example on that coming. Um, I'm gonna just move faster. We built this for a large, uh, hyperscale for this exact thing.
Their particular format is just a different cage. This is a 16 by 800. So that's the tour they happen to use.
Uh, and this gives them a benefit of half the power, uh, half the Rackspace, um, uh, a, a quarter of the cost of, of, so that product was JA 24. Yeah. So how long has this puppy been out there?
We first sampled this April of 2024 and we called it generally available. Both our products have been first spin, no metal spins to market. We called it generally available in November of, of 2024.
And it's been in mass production since summer of 25. Yeah. No, this is out there.