The E-Series Delivering Cloud-on-a-Chip from Xsight Labs
The E-Series session explores the convergence of storage, networking, compute, and security into a single, cohesive silicon platform. This “Cloud-on-a-chip” approach is dissected through its architecture and programming model to show how it simplifies complex data center environments. We will highlight our partnership with the Hammerspace solution, demonstrating how E-Series silicon powers a global data environment. Xsight Labs presents its E-Series, a System-on-a-Chip (SOC) designed to deliver “Cloud-on-a-chip” capabilities by integrating essential cloud elements. This includes Ethernet connectivity, robust security features, virtualized storage, and powerful processing via 64 ARM Neoverse N2 cores. The E-Series chip, which has been generally available for about four months, is offered in various form factors, including a server, an add-in card, and a COMEX module, targeting applications ranging from embedded systems to full servers.
Xsight Labs differentiates its E-Series architecture from traditional DPUs, which typically evolve from a NIC with a constrained CPU cluster. The E-Series began with a server-class compute system featuring 64 ARM Neoverse N2 cores, specifically optimized and sized for data-plane applications. This allows all packets and PCIe transactions to be terminated and processed in software using standard programming models like Linux, DPDK, or SPDK, eliminating proprietary code. The chip integrates an E-unit for Ethernet connectivity, offering inline encryption and stateless offloads, and a P-unit for PCIe Gen 5, providing up to 40 lanes and 800Gb bandwidth. This PCIe unit can software-emulate various devices (storage, networking, RDMA), offering immense flexibility. With a typical power consumption of 50-75W (up to 120W TDP) and a SpecInt rating of 170, the E-Series offers significant compute power efficiently. Beyond network and memory encryption, the roadmap for the follow-on E2 product includes CXL support, targeting 1.6T bandwidth.
The E-Series supports a broad range of use cases, from front-end DPUs in public cloud and AI clusters (offloading the host, providing virtualization and isolation) to back-end DPUs in AI inference clusters for KV cache offload. It also extends to local storage, bump-in-the-wire network appliances for security and load balancing, smart switches for stateful processing, edge servers, and storage target appliances. Xsight Labs provides a comprehensive software development kit, ensuring compatibility with standard ARM server operating systems such as Ubuntu, as well as Linux and DPDK drivers. A key demonstration of the E-Series’ capability is its performance on the Sonic Dash “Hero Benchmark,” a highly intensive SDN workload. This test requires processing millions of routes, prefixes, and mappings, which largely depends on off-chip DRAM due to poor cache locality. The E1 exceeded the benchmark requirement of 12 million new connections per second with 120 million background connections, without packet drops, by almost 20%, while still retaining CPU capacity for control-plane operations, making it the only DPU to pass this test at 800Gb with a single device.
Presented by Ted Weatherford, Vice President of Business Development, Xsight Labs, and John Carney, Distinguished Engineer, Software Architecture, Xsight Labs. Recorded live at AI Infrastructure Field Day in Santa Clara on January 29th, 2026. Watch the entire presentation at https://techfieldday.com/appearance/xsight-labs-presents-at-ai-infrastructure-field-day/ or visit https://techfieldday.com/event/aiifd4/ or https://xsightlabs.com/ for more information.
Transcript
I'm Ted Weatherford, VP of, uh, business development at Xite Labs, and my distinguished colleague. Yep. John Carney.
I'm a software architect, architect on the E-Series, which is what we'll talk about. And I'm just gonna kick it off by saying the E-Series is a separate chip, it's a sock, and we call it cloud on a chip because it's got all the elements of cloud. It's got the ethernet connectivity.
It's got security. Okay. It's got the virtualized, uh, storage, and of course it's got the processing with the arm cores, 64 neo versus two arm cores.
Uh, it's in production, uh, in about four months, and it's been general available for four months. Um, and these are just the different form factors you could get it in. Okay.
We sell a server. We have one here, uh, to show you guys. Uh, and we also have the add in card, uh, and you can add it into the server, so the server has it add in card slaughter too in it, so we can put our own chip as a server and then a smart nick together.
And then we have, uh, the comex and the COMEX is really targeting the control plane applications, but it's actually just a really nice embedded format that, well, you could have a 64 or core product on. John, take it away. Yeah.
Thanks. So the first thing I'm gonna do is I'll talk about, um, what the architectural differentiation is between our DPU, um, and, um, what traditionally has been, uh, the DPU offerings from our, our competitors. So, um, the, the traditional DPU really started and evolved from a nic.
Uh, so what has happened in the past is, um, the NIC would be taken, there would be some flexibility added within the pipeline of the nic. Um, and then, uh, to add even more flexibility, a CPU cluster would was added to the nic. And this is what, uh, our co competitors architectures look like as A DPU.
And this architecture's constrained. And the reason why it's constrained is because the CPU cluster is not really sized. Uh, and the connectivity of the CPU cluster to the offload NIC is not sized so that all of the packets, all of the processing of the data plane could go to the arm cores or go to the CPU cluster, uh, for processing.
It really relies on being able to process most of the packets, uh, in the hardware pipeline and only exception packets, or maybe the first packet of a flow, uh, goes to the, the CPU cluster. And what we found with this architecture is that, um, what we found with this architecture is that for, um, for many workloads, um, there's, there's flexibility that you need, uh, that doesn't exist in that hardware pipeline. And you often are trying to get all of the packets to that CPU cluster, uh, and then your performance is, is limited.
What we did differently is we started from that CPU cluster and we said, let's take a very energy efficient, uh, and optimize for data plane applications, server, class, compute system, and let's size it so that, um, all of the packets, uh, and, and in fact, all of the PCIE transactions when you're attached to a host can all be terminated and processed in software using standard, uh, programming models. Um, not proprietary programming, um, but just a standard Linux, uh, or user space programming. Uh, and so that's what we started with.
And then we added to that, uh, an ethernet unit, which provides, uh, the ethernet connectivity, uh, and provides, um, some processing that's done on sort of the, the throughput or the, the bits of the packet like encryption. Uh, and then we added a a p unit, which is the PCIE connectivity, which when you connect to a host or you can connect devices to the DPU, um, they go through the P unit. But again, all of the data plane processing can be done, uh, in the cores, uh, and not in this model on the left where, uh, you really have to get, do most of your processing in the nic the NIC portion of this architecture.
Um, this, this on the right, um, allows us, um, scalable performance, we can scale performance with cores. And again, it's a very standard development model. Um, and, uh, I think, uh, I don't think we can go to the next slide.
Yeah. The competition does, does it differently. They start at the offload nick, and they add arm cores, and they traditionally, for three generations haven't added enough arm cores and haven't added the more powerful arm cores, like the neo verse twos.
Yeah. So, and it we're really vindicated that with the blue field, right? The blue field's come out and said 64 cores, and it's, it's come out and said, okay, let's put a real compute in there.
So we, uh, we see them as our principal competition. The Intels and the pin sando AMDs are, are still kind of sticking to their guns with the, the other approach. And, and your programming model On that solution is Linux.
Yeah. So, uh, we run Linux. We run and, and we'll get into this more on some of the other slides, but, uh, we run Linux.
Um, we, we are, um, we will be certified as an ARM system ready. Uh, actually we're certified as a server class, not an embedded class, uh, meaning that any, uh, operating system distribution that can run on an arm server, uh, uh, can run on our DPU. And so we see, uh, and we'll get into use cases as well, but we see use cases where customers want to do everything in the Linux kernel, um, you know, all of the, uh, you know, networking stack of the Linux kernel.
Uh, and we see a lot of use cases where customers want to use user space frameworks like DPDK or SPDK and do all of the data plane programming. Uh, in, in user space. One of the big advantages that we have is that, uh, a lot of these data planes already exist.
The customers have been running them on servers, um, and now what they can do is they can take those data planes that they're all running, running today on servers, and they can retarget them unmodified, uh, onto the DPU, and they get a huge, uh, gain in performance per watt, uh, in, in, in density, um, uh, to be able to put multiple of these e ones in a small form factor. Uh, and so that's really an advantage we have is that all of those existing data planes that people have run on servers, um, can just be, uh, placed onto the, the E one Go. Sorry, go ahead.
Start, go ahead, Jack. I'm sorry. I was just gonna ask, so on the traditional side, it's all custom code, right?
It's, yeah, so it's a combination and it actually, the combination is what actually makes it difficult because, um, there's sort of custom code that runs in that offload nick portion of it. It's usually proprietary. Um, it's usually very resource constrained.
Uh, and then there may be more sort of standard programming model on the CPU cluster. But to create, uh, an application, uh, you kind of have to split the application between these two very two very different programming models. Mm-hmm.
Uh, and that's what makes that like, kind of that environment, like very difficult to, to program. Yeah. I think it was also a problem in the security hardware market too, right?
When I've gotta do packet processing between CPU and my offloading nick. Right? It's, it's a huge issue.
'cause got to have like almost a one core application to do that kind of processing. So if I can put it closer, then makes a lot of sense. Right.
So in that case, then, because you are traditional Linux based, you can run into Linux things. So are people doing EBPF on this? Yep.
You can run eeb PF um, so we see that there's a camp that's sort of the Linux, EBPF. There's another camp that's the user space, DPDK. Okay.
And, uh, and we support, uh, both and I think we'll show that a little bit more in, in some of the future slides. Um, so it's really, again, like, um, it's just if you've programmed data planes on a server, um, this is just a very familiar envir environment to you. Um, okay.
I think we can go to the next slide. So this is a closer look at the, at the actual chip architecture. So, um, in the, the heart of the architecture is that compute fabric.
Um, that's where our 64 arm neo verse N two cores are. Um, they're connected with a high speed scalable c coherent fabric. Um, we support 32 megabytes of system level cache that's shared by all of the cores.
The cores also have their own private L one and L two caches. Um, we have a memory subsystem where we have four channels of DDR five memory. Uh, there, those can operate up to 5,200.
Um, and, uh, this is a little bit unique. Most, most dpu, uh, don't have, uh, that much, uh, memory io. Uh, and depending on the form factor, Ted showed those form factors earlier.
Um, in some form factors where you're not as space constrained, um, you can populate all four of those memory channels, uh, in some other form factors where you may be more space constrained, you might only populate, uh, two of those, those memory channels. But we give you that ability to use, uh, uh, more, uh, memory io, uh, depending on your, your use case. Um, a very unique aspect of this chip is our PCIE connectivity.
So we have PCIE connectivity that could be host attached where you connect the DPU to a host. That's a traditional kind of DPU application where we're acting as a nic, uh, to a host. We did something unique where we in hardware, um, created PCIE hardware that presents all of the services of PCIE to the software so that the software can emulate any kind of PCIE device.
So to a host, we can look like a storage device, we can look like a networking device. Uh, we can look like an RDMA fabric device, and it's all software defined, again, using just standard, uh, pro programming models. Uh, we support up to 30, sorry, sorry, 40 lanes of PCIE Gen five.
Um, some of those lanes are typically used for, you know, like an MVME boot, uh, or peripherals. Uh, but then there's 32 lanes, which would typically be used in your data path. Uh, so, um, that data path can allow up to 800 gig of P-C-I-E-E bandwidth towards a host, uh, or towards storage.
We'll talk more about storage, uh, later, um, uh, or towards other, uh, devices. On the bottom of the diagram, we show our ethernet, or we call it a e unit. Um, and this, um, appears to, um, the programmer, um, as your nick.
Um, it appears like a nick. It appears sort of like an enterprise class NIC, with the features that you would expect in a nic. It also has several offloads.
Uh, we support inline, uh, encryption. Uh, we have a lot of flexibility and what kind of formats we can support. Um, uh, and, uh, it also supports all of the normal stateless offloads that you would find in n nick.
Things like CRCs and check sums. Um, it also supports, um, lookups and you can, you can do some flow offloads, uh, in the NIC portion of this chip. And the way we think about it, and the way our, our philosophy is, is that that ethernet hardware is really there to help, um, steer traffic to the cores to enable you to scale the processing of those packets, uh, on the cores.
So we're not trying to do all of the processing and all of the, um, flexibility in the NIC portion, like our competitors are. We're trying to do the primitive operations that you need there to enable you to use those arm cores, uh, effectively. Um, little bit on, on power.
Uh, so the chip itself, uh, you know, typical use cases would be, you know, for typical would be 50 to 75 watts. Um, and, you know, for thermal design power, um, it's up to 120 watts. Um, this is rated with a spec end of one 70, which is a, you know, significant amount of compute power in, in such a energy efficient, uh, chip And John, uh, Jack Poller with Paradigm Technica, I noticed that you have a memory encryption block in front of the memory.
Yeah. Do you support, Uh, so is it, uh, you can support, uh, secure encryption Data in use? So we have, we have like multi multiple different places where we do encryption in this chip and I'll, I'll talk about them.
So, um, one of them is inline, you know, as data comes in and out of the ethernet interface. Mm-hmm. Um, and so that's like your IP sec or Google has something called PSP.
Um, also ultra ethernet has their own security. So all of those can be done in line, and if you can do them in line, that's really the best place to do it because it doesn't require you to bounce the, the data through your memory system multiple times. Sorry about that.
We also have, um, uh, a look aside crypto engine. Uh, and so there's things like TLS, uh, you know, uh, TLS where you wanna terminate, um, like TCP traffic that has TLS, and it's a very staple, um, and it's not something that's easily done in line. So we have a look aside engine where the arm cores can kind of feed that look aside crypto engine, um, and do crip encryption there.
Then we have the memory encryption. The memory encryption is really there so that, um, data that's stored in the drams on the chip is, can actually be encrypted. Okay.
Uh, so that there's not any way you can kind of snoop on those memory interfaces and be able to, you know, access any kind of sensitive data. It also ultimately can allow you to have, um, multiple tenants sharing, uh, uh, you know, cores on this chip and be able to isolate, um, you know, their data, uh, in a way where it's, where they're each encrypted from each other. That's sort of where I was going.
I was wondering if you could do a secure uncla in this. Yeah. So I'm not really the right person to like, get into the details of that.
We have security architects that could really, uh, an answer all, all of that. Yeah. Uh, very Quick question on your PCI side.
Uh, is that capable of doing CXL interfaces as well? So we did not implement CXL, uh, in this chip. Uh, we're not gonna really talk about road roadmap too much for this chip, but that's something that, uh, is, is very much, uh, um, in mind for us, uh, uh, on our roadmap.
Okay. Because it's, I mean, there's, there is some interest in the industry Yeah. In the ability to actually, uh, have essentially Nick to CXL Yeah.
Interfaces. Yeah. And yeah, and this is something that we're, that we're following closely again, and, and seeing that Yeah, we're Seeing that, yeah.
Yeah. Memory for sure. Yeah.
This chip, this chip does not, uh, uh, support CXL, but, um, but stay tuned. Yeah. The follow on's called the E two.
We don't have a roadmap slide on it, but stand up. Yeah. Sorry.
Yeah, sorry. Yeah. The, the follow on product is the E two, uh, and it's roughly, um, double the bandwidth of this one.
6, and that's where you'll, you'll, you'll see the CXL support come in. Um, yeah. Okay.
I think we can go next slide. Um, so we'll talk about some of the applications. You know, starting from this server class kind of architecture, you can imagine that there's just a very broad set of use cases, and we'll talk about some of 'em.
We'll talk about the ones that we've really been focused on. Um, CHIP was really architected for this use case on the left, uh, we'll call it the front end DPU. Um, and this use case, um, could either be, um, you know, just a traditional public cloud compute node, uh, in that node.
Um, the tenants need to get access to the infrastructure. Um, the DPU is what provides, um, the virtualization of the infrastructure, the securing of the infrastructure, um, and for, for both, you know, networking and storage. Um, and it provides isolation amongst the tenants that are running on the host.
Um, but the front end also, you know, applies to, uh, an AI cluster, um, where essentially the same kinds of things have to be done for the AI cluster to provide access to the, the infrastructure, access to the storage. Um, and in the, uh, you know, for the, the AI use case and, and both for the traditional cloud compute use case, the DPU is really offloading the host and, and offloading the DPU from having to do those functions and offloading it in a way where it's done on a, you know, very optimized architecture for those workloads that are running, uh, on the DPU, um, the backend DPU. So in the, typically in the sort of training clusters, these high scale training clusters, this backend nick function as traditionally and, and predominantly been done with the performance nicks, uh, those nicks allow you to get, you know, RDMA, uh, performance at very, uh, high rates and low power.
Um, and, you know, the, there are opportunities in those large clusters for dpu more around innovations innovating around congestion control or, um, congestion avoidance, uh, packet spraying telemetry. Um, there are many, many use cases where even in a training, um, uh, cluster, the DPU makes sense, but really in the inference cluster is where the DPU, um, can shine. Uh, because, uh, in the inference cluster, that's where we start to see, um, uh, uh, the offload of, uh, KV cash offload, uh, from the GPU to something else.
And the DPU is a natural place to do that because it has direct access to the GPU's memory system. Uh, but it's also, you know, directly connected to the scale out network, uh, of those inference clusters. So as we see, you know, a push to building special purpose clusters for inference, there's a lot more opportunity for the DPU, uh, in, in those.
Um, we also have use cases for local storage, um, and in the local storage use cases to the host, we can virtualize storage devices, but those can actually be, um, the, the E one and the DPU can actually be managing a local set of, of discs. Um, there's a bump in the wire use case. And, and this is one diagram.
There's many ways to draw this bump in the wire, but you can think of the bumper in the wire as just simply a network attached appliance where packets come in, maybe come in on one port, get processed, go out another port, many applications for bump in the wire, um, both, um, in security, um, and, you know, software load balancing. Um, and, and we see many use cases there. Smart switch, um, is another use case that I think we'll we'll touch on more later, where we actually combine the switch with the dpu.
Um, and, um, what this allows, um, is it allows you to use the switch, um, to have, you know, a large bandwidth of traffic coming into the switch and to selectively choose which flows you want to go to the DPU for more processing. Typically the switch is doing your stateless processing, um, and the DPU could do, uh, very state stateful processing. I think we can go to the next slide, Ted.
And then we have, um, server use cases. So, uh, we have edge server use cases. We see a lot of applicability for DPU for things like CDN, um, for, um, gaming.
We, we see use cases. Um, and, uh, and then, uh, on the right we see, um, the E one and sort of an appliance form factor, uh, for storage target kinds of use cases where the storage is disaggregated from, uh, from the host. Uh, this is sort of a look at our software developer kit.
I've said multiple times, I think so far that we're standard programming model. Um, but we do provide, we do provide software, um, in a few forms. Uh, we provide the drivers and the necessary infrastructure pieces of software to enable our customers to use the DPU.
Uh, and we also provide reference software for many of these applications and use cases. So at the, at the lowest level, um, we provide the secure boot, the, the bios, um, and uh, uh, uh, the sort of low level software that gets the chip booted and, and operating. Um, and then as far as the operating system, as I said earlier, um, any, um, operating system distribution that can run on an arm server, uh, can run on the E one internally, we've been using Ubuntu.
That's kind of our sort of, uh, default. Um, and then we provide the DPDK Linux drivers on top of that for our networking. Um, and then as I talked about the PCIE where we can emulate any kind of PCIE device, we implement what we call backends.
So we implement a backend for networking. Uh, we have a couple flavors of that vert io and XNA. We implement a backend for N-N-V-M-E emulation and as well as backends for RDMA and, and Rocky.
Um, and then as you go above that, you start to get into the, the applications. And typically our customers will use our chip as a platform, uh, and then they'll put their own applications on top of, uh, uh, the E one. Um, there is an application that we've developed called Sonic Dash.
Um, we'll talk about that more, um, in, in a couple of slides. Um, and that's where we'll really get into some of the performance, uh, in a real world, uh, use case SDN use case for, for the E one. Okay.
Right now, so the reason why, you know, we kind of selected Sonic Dash to talk about is because this is a very heavy workload, um, SDN workload. And, um, so what is Sonic Dash? Sonic Dash is a Linux Foundation project that was, uh, started in, in 2021.
Uh, it was really a project pushed by Microsoft. Um, and the goal of the project was to take the stateful services that run on the host in the public cloud and define them and define the APIs and the object models in the, um, in, in the, those services. And the goal of that was to enable a broad set of technology providers to really create performance and power and cost optimized implementations of these services.
Traditionally, um, that cloud, SDN had been built on, um, frameworks that had primitive operations and a service was sort of defined as stringing these primitive operations together, um, in order to implement the service. And over time, as the cloud matured, um, it got to the point where we no longer need, um, this kind of low level way of defining the services in the cloud that we could do it by. Actually, we know now after decades of cloud computing that these are the seven services SDN services that you need in the cloud.
So they've been defined and we've implemented, been implementing them. One of the services is called vnet to vnet. Um, what does this do?
So, um, packets come in. Uh, you apply lookups transformations to be able to map the overlay to the underlay. Um, does state tracking for TCP and UDP connections, um, dash defines that you have to do five ACL operations.
ACL algorithmic. ACL is a very intensive, um, especially at high scale on memory accesses, um, accounting enforcement of rate limits, low scale limits, um, and also, uh, you need very high availability. Um, because if tracking millions of connections, you need to make sure that that state is synchronized on another data plane.
So if there's any kind of failure in the network, all of that traffic can be taken over by, um, a peer data plane. And so there's no loss of connections to the users of the network. Um, in order to run Sonic Dash at an 800 gig scale, um, it requires, um, millions of routes, um, millions of prefixes in your ACL tables, millions of mappings.
The the point, um, of showing the scale is that this kind of scale is not something that you can fit on, chip on, on chip memory. All of this has to live in dram, uh, off of the chip. And there's also no locality as packets come into the chip.
A packet might be for one flow, next packet is for a different flow, and you can't rely on those caches for locality. Everything has to go to dram, uh, in order to, uh, be able to process the packets for Sonic Dash, it doesn't matter what architecture you have, it doesn't matter if what you have like the e architecture, uh, or one of our competitor architectures. It's a very dram intensive, uh, use case.
Um, so if we go to the next slide, we can show that, um, there's a test dev defined for Dash called the Hero Benchmark. It's the highest scale, most intensive test that you can do. Uh, the test requires that you have, um, over 120 million background connections.
So this background traffic is running, and then while that background traffic is running, you have to be able to support 12 million new connections per second. Um, and you have to be able to do that for a hundred seconds without a single packet drop. We, on the E one, uh, have been able to exceed that performance.
Um, we were able to get, um, almost 20%, um, uh, excess on that performance requirement. We're able to do it where we still have cores left over. So you can run those, uh, those remaining cores for your Sonic Control plane.
Um, and even with the performance we're able to achieve, we still see, uh, the potential to achieve even more performance gain, uh, or uh, be able to do it with fewer cores. Um, and, uh, and we see like the potential for another 25 to 40% performance improvement over what we've already tested. So this is just to give you an example of a Cloud SDN use case.
It's really the use case that we designed and architected the chip for. Um, and to show that we're actually the only DPU, um, that's able to pass this test at 800 gig with a single device.