AI Unbound, Your Data Center Your Way with Juniper Networks
Praful Lalchandani, VP of Product, Data Center Platforms and AI Solutions at Juniper Networks, opened the presentation by highlighting the rapid growth of the AI data center space and its unique challenges. He noted that Juniper Networks, with its 25 years of experience in networking and security, is uniquely positioned to address these challenges and help customers meet the demands of AI. Juniper is experiencing maximum momentum in the data center space, with revenues exceeding $1 billion in networking alone in 2024. The presentation then dove into the increasing distribution of AI workloads across hyperscalers, Neo Cloud providers, private clouds, and the edge, emphasizing Juniper’s comprehensive portfolio of solutions spanning data center fabrics, interconnectivity, and security.
Lalchandani focused on the critical role of networking in the AI lifecycle, particularly for training and inference. High bandwidth, low latency, and congestion-free networking are essential to optimizing job completion time for training and throughput and minimizing latency for inferencing. The discussion highlighted Juniper’s innovations in this space, including developing AI load balancing capabilities such as Dynamic Load Balancing, Global Load Balancing, and RDMA-aware load balancing. Juniper was the first vendor in the industry to ship a 64-port 800-gig switch, showcasing Juniper’s commitment to providing the bandwidth needed for AI workloads and achieving a leading 800-gig market share.
The presentation also emphasizes AI clusters’ operational complexity and security challenges. Juniper’s Apstra solution offers lifecycle management, from design to deployment to assurance, providing end-to-end congestion visibility and automated remediation recommendations. Security is paramount, and Juniper advocates a defense-in-depth approach with its SRX portfolio, protecting the edge, east-west security within the fabric, encrypted data center interconnects, and security for public cloud applications. The presentation concluded by addressing the dilemma customers face between open, best-of-breed technologies and proprietary, tightly coupled ecosystems, and that Juniper offers validated designs with AI labs to show customers they don’t have to make a trade-off and can get the best of both worlds.
Presented by Praful Lalchandani, VP of Product, Data Center Platforms and AI Solutions. Recorded live in Santa Clara, California, on April 23, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/juniper-networks-presents-at-ai-infrastructure-field-day-2/or https://techfieldday.com/event/aiifd2/ for more information.
Transcript
My name is PR Ani, and I lead product management at Juniper for our data center business unit. So the delegates in the room, and for everybody else who's tuning in, uh, welcome to Juniper and thank you for being here. We do appreciate the valuable amount of time that you spend with us today.
Now, if you've been in the networking domain for a while, I would suspect you know a little bit about Juniper, but those of you tuning in and wondering who the hell is Juniper about? This is a one slide primer on our history, right? So we've been around for 25 years.
Uh, we have deep roots in networking and security, a global customer base, and we serve enterprise, cloud and service provider verticals, uh, customer segments. We have products and solutions that power networking and security in multiple places in the network, starting with wired and wireless in campus, and branch van routing. Number two.
And number three, which is the focus of today's discussion, is we have networking and security solutions for data center as well. In fact, data center is where Juniper is seeing the maximum momentum these days, these years in the last two years with our revenues in data center crossing $1 billion in networking revenues alone. That's not even including security in 2024.
And we have more than 10,000 happy customers who have deployed Juniper in the data center. Some of those logos, in fact, just a fraction of those logos are on the screen, uh, in front of you, who are some public references that we have. But today, uh, the focus of today's session, uh, we are gonna dive into the realm of AI data centers.
And from our vantage point in the industry as an infrastructure provider for AI clusters, we see one very clear trend that AI is increasingly distributed. Let's explain how, why. So, right?
So if you're an enterprise, you're gonna have some of your applications that you're gonna host with your favorite hyperscaler. But in the new world of ai, there's a breed of GPU as a service providers and AI as a service providers that are coming out there with niche offerings, and we call them neo cloud providers. So your application or your AI workload may be hosted with your favorite hyperscaler, but very likely is also hosted in some other cloud provider like some, some of these neo cloud providers.
Now, if you're an enterprise for various business reasons, whether it is cost, whether it is data privacy reasons, or whether it is regulatory reasons, you may also decide that some of these applications, you wanna host them in your own private cloud or on-prem. And then that's also edge ai because for some inferencing applications, latency matters and you wanna be as close to as possible to the end user, uh, for those inferencing use cases. Now, at Juniper, we believe that in order to support this distributed architecture, we have, we are uniquely positioned with a comprehensive portfolio that spans data center fabrics, it spans data center interconnect to connect all of these various clouds together and security natively built in.
So wherever our AI is happening, whether it is in the cloud, whether it's in private cloud or at the edge, Juniper is right there to provide solutions for you. Now, this is what really makes us excited about data centers in general, because this new wave of ai, if I take an example of a general purpose server, right, for general purpose computing reasons, even today, two ports of 25 gig for a total of 50 gig bandwidth per server is considered more than sufficient. If somebody maybe going to a hundred gig, but predominantly 50 gig per server, you look at an Nvidia node, even the last generation of hopper series with eight GPUs, you're looking at 400 gig connectivity per GPU.
You look at multiply that by eight GPUs, front end networking, storage, networking, that's a total of four terabytes of te terabytes per second of networking per server. That's 40 x or sorry, 80 x more than what you need for a typical general purpose server, right? So this is what really makes us exciting because AI is lifting all boards and networking is no exception to the rule over here.
And that doesn't stop, in fact, it goes on fire in the next few years. This is a slide straight from NVIDIA's GTC, uh, jenssen's presentation. And what Jensen showed is that in 2025, when the Blackwell 300 series of GPUs comes out, it's gonna be powered with 800 giga networking with the CX eight nix and doesn't slow down there.
6 terabytes per second in just the next few years. And so now I wanna kind of impress upon why networking is important in your entire AI lifecycle and no rocket science in this slide so far over here. I think people understand the lifecycle of an AI model, but just to baseline it a list a little bit, right?
So you start with what you call is foundational model training. This is the purview of large cloud providers or large enterprises who have resources of thousands and thousands of GPUs to train a model from scratch. But most enterprises are essentially gonna be doing fine tuning where you take a generalized model, which has been trained on a generalized data set, and fine tune it for your, with your proprietary data so that it can be customized for your own specific use case.
And then the rubber hits the road in the case of inferencing, where you take that trained model and dis deploy it in a distributed way across many, many AI clusters so that it can serve end users. Now, we know that training was always a multi GPU problem, but what I wanna impress upon today is that with the rise of things like reasoning models and agentic ai, uh, even influencing is becoming a multi GPU communication problem. Uh, while networking is important for all of them, what the, the metrics tend to change between training and, uh, inferencing training.
The most important metric is job completion time. How much time did it take to train my model? And with inferencing, the metrics are slightly different.
You're looking at throughput, uh, which is my, I have an AI factory that is doing inferencing. How many tokens per second, or what's my throughput that I can generate for all of my users? And a second metric is because there's a real time user sitting behind that LLM waiting for a response, the latency or the time to first look.
Token is an important metric for inferencing as well. And like I said, networking plays a critical role in maximizing all of these metrics. And we'll, you know, today we'll show you how.
Now, diving a little bit into training, uh, you know, the training cluster, the, the, the green box in the middle, think of it as a cluster of GPUs, typically has three different networks involved. There's a front-end network, there's the backend storage network. But the important one that we wanna focus on today is that backend GPU training network that provides GPU to GPU connectivity across that cluster.
And this is where you wanna provide high speed, low latency, congestion free networking, because any congestion in these, in this infrastructure can lead to traffic drops. And anything that slows down GPUs due to retransmission of packets or any other reason, is really criminal in a, in the case of AI infrastructure, because these ex GPUs are expensive, you're spending millions and millions of dollars for these GPUs, uh, for your AI infrastructure. And having them sit idle is not something that you want.
So the networking plays a critical role in improving the job completion time for training jobs. You look at inferencing, like I said, it's is no different. You start with the left hand side.
The most simplest case is single node inferencing, right? Single client request comes in a GPU response. Sometimes it may be even be a CPU, right?
But as I said, with the rise of new types of models, for example, reasoning models where there's a internal dialogue happening in the model before even the first token is sent out to the user, there's a lot more EastWest traffic happening as the model tries to reason its output before the first token or the first response goes out to the user. Uh, again, there the you need. So that becomes a multi GPU problem.
The same thing with agent AI frameworks. So what we are seeing is that increasingly, uh, inferencing is becoming a multi GPU problem as well, and requires a need for a backend inferencing network for east west communication. And la lastly, we have RAG or retrieval augmented generation where idea behind retrieval augmented generation is that you're supplementing an L LMS output to give it, to enable it to give more precise responses by giving, by supplementing proprietary data that it can use to, to generate a response.
Uh, again, in the case of rag, you're basically looking at an environment where you need to deliver very low latency because again, there's an end user sitting at the end of the response, and you need to have, uh, you know, those shared storage nodes that can be accessed at very, very high speed. So networking, again, is, plays a very important role over there. Now, you know, stepping one step about, uh, you know, what are the challenges that anyone who's deploying an AI cluster faces?
I hopefully by now, I have stressed upon the fact that performance is important because GPUs are, are critical asset, but these AI clusters are increasingly more operationally complex to manage as well. Think about it, I told you that it was four terabytes per second of networking per server. That's so many more optics that can fail.
So many more links that can bounce so many more BGP sessions that can flap. And on top of that, you have the need and AI clusters to have real time visibility of congestion events and the ability to manage and remediate those congestion events as well. So operationally, way more complex than typical data center networks.
Security was always important in an AI cluster. Now, in the world where your proprietary data is being used to train these models, you face new rest, new threat vector around model exfiltration, data exfiltration, et cetera. And we're gonna spend some time on that.
And then finally, in the, in the world of AI clusters, what we are seeing is that customers are facing this new dilemma. Should I pick open best of breed technologies based on standard ethernet, or am I locked into proprietary technologies just because they offer, uh, you know, a closed ecosystem or a tightly coupled ecosystem? And we are gonna spend some little bit time on that aspect as well.
How much are you seeing the shift these days from scale up mega clusters to more scale out, uh, more efficient ones that only use resources as needed? Is that something that, that you all are seeing and you're planning for? Or do you already have that?
So we are in the scale out domain right now, and, uh, the scale up domain today with Nvidia is a pretty locked in and we switched architecture, but as a MD GPUs become more relevant, that scale up, um, scale up opportunity opens up to other vendors as well. So today we are mostly scale out, but we are also looking at scale up with non Nvidia GPUs. So that seems to be opposite.
We have the dense models that people started with and the that are vertically, but you've started out smaller with the more efficient LLMs. So I mean, look, until NVL 72 came out, uh, most of the hopper class of GPUs were sold as discrete at DGX and at GX systems, right? It's only with the Blackwell where the NV L 72 is coming out, right?
So most of our customer deployments, even even there's a customer who's deployed us for 200,000 GPUs, this was all scale out. It was not deployed with scale up. Scale up is in fact is emerging recently since NV 72 became more, uh, popular starting in 2024 GTC.
Do you have at some point a slide, and maybe I missed it earlier, that shows the whole topology of a, of a AI network? Yeah, We're gonna spend some time on that later. Okay, Great.
Thank you. So, uh, starting with performance, uh, no shame in saying the plus thing that we did was to throw bandwidth at the problem, right? And we were the first vendor in the industry to ship a 64 port, 800 gig switch.
We were the first OEM vendor in the industry to ship a 64 port, 800 gig switch that can, uh, in a two RO compact form factor, that can actually have two ports of 400 gig connectivity to the hopper class of GPUs. And it's future proof for the Blackwell class of GPUs as well. So you can actually build a pretty large size cluster eight K GPUs cluster in a three stage design with just the QFX 50 to 14 a leaf and spine roll.
But if you want larger, we did not only stop there, we also introduced a PTX 10,008 chassis that has 288 ports of 800 gig. So you can build clusters again in a three stage design going up to 32,000, uh, GPUs. And now, because the PTX is also a deep buffer system, like I said, there are multiple AI clouds.
They need to be interconnected together using data center interconnect. This deep buffer platform is a, is a superior choice even for your data con data center interconnect use cases. In fact, we were the first vendor in the industry to have a deployment of 800 gig using coherent optics for ZR led.
So we were first with 800 gig on the fixed switches. We were first with 800 gig on, on the PTX with a deep buffer switches both for data center fabric as well as data center interconnect, uh, use cases. Now, uh, now we were obviously first to the market, like I said, with 800 gig, and that's hopefully a testament to our engineering innovation and engineer, but doesn't hurt that we got some market recognition out there as well.
This is something that cust, you know, nobody expects to see. But as reported by six 50 group, just a few months ago in the year 2024, Juniper had a leading 800 gig market share. So the reason for that is we had the products first to market, uh, but obviously this AI wave was right picking up right at that time.
So it was a good product market fit, but we had a market share that was larger than Nvidia, larger than Arista, larger than Cisco. In this data right here, are we looking at any specific verticals or regions or would that be a, a slide talking about global market share? This is global market share, but data center alone, it's not outside of data center, it just focused on data center.
Now, like I said, first thing we did was throw bandwidth at the problem, but we all know that standard ethernet was not designed to be lossless. Uh, it is not standard ethernet without any advancements. Advancements that we are gonna talk about today is not suitable for AI workload.
And the reason for that is that ethernet switches make a random hashing decision on which path of flow should take. So if you look at a picture on the top right, two, two flows, two 400 gig flows are coming in from the leaf, from GPUs hitting the leaf. And even if I have built up enough capacity between my leaf and spine, it's a possible that that ethernet switch makes a decision that both those flows go to the same path to the spine, and that link is gonna get congested.
So you have a situation where some links are underutilized, others are overutilized and need, and you start hitting congestion and failures out there. Similar situation happens on the picture on the, on the top left as well. So we started looking at this problem, you know, way back in 20 22, 20 23.
And the first thing that we did was introduced a capability called dynamic load Balancing. What Dynamic Load balancing does is that it makes a path forwarding decision, not just based on what available parts that routing tells me, but based on the quality of those links in real time, it's making a real time decision on the quality of those links and making a forwarding decision. You'll see in the next session that that gave us tremendous improvement in performance over standard ethernet.
But again, we did not want stop there. Uh, we introduced this fee feature in 2024 called Global Load Balancing. The idea behind Global Load Balancing is that it's like a Google Maps for the network.
When I'm driving, I see the congestion in front of me, but I only Google Maps can tell me the end-to-end path quality. So I make a decision now to turn left versus right. The same way global Load balancing.
What it does, it gives this LEAF device that it needs to make a path decision end-to-end quality information of the entire path, all the way to the destination leaf. And using that information, it can make more educated decisions to avoid, uh, to avoid, uh, the, both the local length as well as the remote congested link. But did we stop there?
No. And today what we are announcing is something called RDMA Aware Load Balancing. This is a capability that enable the switch to see RDMA subflows, be aware of RDMA subflows and map them to specific parts through the fabric.
It is gives us the highest performance, as you'll see shortly. It gives you the lowest variability in performance. If you run that AI job 10 times, you're gonna get the same performance each and every time.
DLB and GLB still give you some, some amount of variability around a mean, and it requires at least amount of tuning to optimize the performance of your infrastructure. So today we are announcing AI load balancing, uh, which is these umbrella term that we are using for all of these load balancing capabilities that deliver a 50% or greater than 50% improvement over standard ethernet. Now, at this point, if you're thinking that's, you know, you made me curious, but did not satisfy my curiosities, uh, completely.
I need to learn more. I need to know more. Uh, this is about what I was designed to do.
Uh, Ali Rasing, who's joining me, uh, joining the next session is actually gonna cover these in a lot more detail. So stay tuned for, uh, his session. But I wanna put a little bit plug over here.
Um, and this is an important one. You will hear our competitors who have proprietary networking topologies for AI clusters constantly say things like this proprietary technology is 60% better than standard ethernet, 70% better than standard ethernet. What they're comparing against is unoptimized ethernet.
That was not designed for AI clusters. You know, with the capabilities and the improvements that we are bringing to the table, we feel very, very confident that customers deploying these customers will get the performance as like comparable to any proprietary technology out there, including infin bear. So in this case, what does open mean as far as proprietary?
Does this mean that people have, so anyone has access to the protocols? Is it, is It the, it's based on BGP, it's based on ethernet. I mean, all and know it can connect to an A-M-D-G-P, it can connect to CEUs, anova.
It's not, it's agnostic to the GPU. Okay. It's not locked in.
Yeah. Industry standard. Industry Standard.
Yeah. I have a question about, um, um, uh, on the previous slide you said no tuning, I would expect that mm-hmm. Uh, that you, uh, need to do some tuning definitely with RDMA.
If you're gonna go all the way into the stack of the operating system or the, the servers as such, tuning On day zero, like, uh, you know, you don't need to tune, uh, like you, it's not like you have to tune constantly. Like if I'm running a different model, model A versus model B, I don't need to tune the network for that. DLB and GLB do require some amount of tuning.
There's some timeouts that you have to configure appropriately to get the maximum performance. Uh, but you know, this technique, it will not. And I think CRU session will make, make that a whole lot more clearer.
Okay, thanks. So, and some of the things that you're talking about are, are these part of the Ultra Ethernet Consortium work? These, so, uh, so part some of them are, uh, some of them will become standardized as part of the UEC.
Uh, so for example, our switches already support packet spraying, right? What, what is UEC standardizing their standardizing the ability of switches to packet spray and the ability of the nicks to be able to, to handle those out of order packets. That's something we already support today.
Okay. Right, because I knew that was part of it when you were talking about the, the load balancing. I just wasn't sure if this was tied into to what's actually gonna be standardized or not.
Yeah, some of these things, so honestly, from a, from a UEC perspective, they're standardizing on package spring when it comes to congestion avoidance, right? And that's something we already support in our platforms. Now, uh, we can't really have an AI data center discussion without actually talking about power.
And, uh, if any of you attended jenssen's, uh, keynote this year in GTC, he eloquently stated the problem statement, especially around optics, right? He said, a hundred thousand GPU cluster has 400,000 optics associated with it. You know, if we think that those optics consume 12 watts of power, depending on whether what the reach is, uh, that's total of close to five megawatts of power that is consumed just by optics alone, right?
And what Jensen said that, you know, he's introducing technologies, uh, like CPO that will be available in 2026. We are working on CPO as well. But what I wanna say the operative word is that we have technologies today in something called linear pluggable optics that takes some of the DSP functionality away from the regular transceivers into the switch, the tomahawk five switch itself on the chip itself, so that we can actually deliver that reduction in power consumption, like 67% in the case of VR optics.
Again, this is available for our customers to take advantage of even today while we are working on other technologies like CPO, AIM for 2026. Now, I did mention operational complexity, uh, as the one of the bigger challenges in AI clusters. And, uh, we have a whole session talking about abstract and how, how it, uh, is the simple antidote to that complexity.
But just a little bit of an introduction on what abstract is. Abstract is our day zero to day and full lifecycle management, uh, solution for fabric management and assures so right from design to deploy to day to assurance, it can do all of those things for regular data centers. Now, obviously in the last year, year and a half, uh, we have been spending time enhancing the capabilities of abstract to be ready for AI clusters.
And we'll see a lot of that today. But I just wanna point out two capabilities that I've, you know, close to my heart, because in talking to our customers, one of the things that we realized was congestion management was the hardest part of operating an AI cluster. So we have delivered some capabilities that provide, not just switch visibility, but switch to nick visibility, end-to-end visibility from all the way from the switches and the nicks, to give you that view of where is the congestion in my network, in my, in a, in a singular view.
But again, we did not stop there seeing congestion, hotspot is one part of the problem. Operators could sometimes spend days in remediating those congestion hotspots. And what we do with RA is this intelligent mechanism that is constantly monitoring for congestion and in making recommendations to operators that here's the hotspot, this is our recommendation to tune the network or change the con congestion handling profiles so that you can ele remediate that congestion as you switch from one model to the next.
So the third pillar of our solution is around security. And, you know, again, QAR will join us later to talk about the security portfolio. But like I said, security was always important in any data center.
In this new world of ai, uh, where your proprietary data is being used to train a model, you wanna make sure that you are protected against those new threat vectors around model exfiltration, around data exfiltration, et cetera. Now, let me be very clear over here that when your proprietary IP is at risk, this is not just a technical issue. This is, you know, a hit to your business value if your entire model gets exfiltrated out, and that's what your and what that model was trained on, your proprietary ip, uh, more more of that to come.
But what I wanna point out over here is that this is not just theoretical risks. These things are happening in the industry as we speak. And I just wanna touch upon a couple of cases just from 2024 in an example on the top, uh, in early 2024, there was a case where hugging face models were infected with malware.
You know, data scientists or users who downloaded that malware of that model essentially had that model model running some, you know, some suspicious code that opened a back door to a bad actor in the cloud. And from there, you know, that bad actor can now do anything nefarious with your models, your data, et cetera. The second one actually blows my mind completely is because, you know, this is a case where an AI as a service provider was, uh, had a vulnerability where a tenant in that, in that, uh, in, in that infrastructure could be a bad actor, and they run a malicious model.
Once that malicious model is run, that model now laterally moves across to other tenants in that infrastructure. So it's not just you who are, you know, maybe did something wrong, uh, maybe, you know, maybe by intention or, or not. But now you are affecting every other con customer that was residing in that AI as a service or GP as a service, uh, infrastructure.
So I think that was pretty crazy. Uh, so that's why at Juniper we recommend a defense in depth approach. You protect at the edge with our secure se, with our SRX portfolio.
You do not want that threat to enter your infrastructure. But you know, sometimes even with the best efficacy in the industry, day, zero threats do penetrate. And what you wanna do is that then to provide that east west security within your fabric itself.
So you assume that, you know, you have some day zero vulnerability, some day zero threats, you know, uh, did enter your infrastructure. You want to prevent the lateral movement of those threats across your infrastructure with our, both our QFX multi-tenancy capabilities that provide isolation for tenants in the, with the QFX as well as our SRX portfolio. Then you have data center interconnect.
Uh, you want to encrypt every communication between your, between your clouds to prevent from snooping style of attacks. So we have max sec IPSec and, uh, encryption capabilities available on both our srx as well as our MX series of data center internet portfolio. And then finally, if your applications are distributed and are running on public cloud, then the SRX portfolio, and more specifically the virtualized SRX or the containerized SRX provides the same level of security for your public cloud applications as you would see, uh, for your, uh, private cloud, uh, infrastructure.
And QAR will spend a whole session, you know, talking a little bit more in depth about this. So I said, I'll touch upon the topic of open versus proprietary a little bit more. And, um, and what we see is that, you know, in every engagement I, we have with our customers, they're facing this dilemma, right?
You know, ethernet had become the open defacto status standard for pretty much every place in the network, whether it was campus, whether it was van, whether it was data center, and now comes this new world of ai, and now they have to, they're dealing with this new dilemma. Should I go with best of breed technologies? But then I have the, have the challenge of kind of trying, having to stitch it together and making sure it performs and works well.
Or I go to this single vendor out there that provides the GPU also provides a network. And I have this allure that going with this tightly coupled sys, you know, ecosystem is the only thing that, you know, it gives me that performance. So I have to make this trade off between lower cost with a best of breed versus proprietary solutions that are more expensive, but have that allure of of performance.
So we really decided to take this challenge head on, and the way we did this was built an AI lab where we bought in all of these best of breed conference, right? We spent a lot of money, uh, honestly, uh, but we brought in all of these best of breed components, a GPUs from NVIDIA and a MD storage from our, uh, partners like Becca and West Nicks from Broadcom, a MD pen sand, uh, sorry, polar nick, as well as NVIDIA's Connect text seven, six, and seven Nicks. And we brought them together, stitched it together into an AI cluster regularly, rigorously tested it for, for a few months, uh, against, you know, real world training and in fencing workloads.
And we're constantly doing that. That kind of exercise never really stops as technology keeps evolving. And what we produce, what these validated designs, like the Juniper validated designs.
So if you do a Google search, even now for Juniper AI validated designs, what you'll hit are two assets. One is a validated design that is anchored on AMD's, GPUs, another validated design that is anchored on Nvidia GPUs. So what we are, what we, what we are hoping to offer our customers is something very simple, that they do not have to make a trade off anymore between breast of breed flexibility and performance.
On the other hand, they can get the best of both worlds. So I'm a little confused about, okay, so you're doing a validated design based upon this stack here, Kimberly Bates with HR Group, and I'm not a network person, so just really not Yeah. Um, are, you're doing a validated design using those, that criteria as opposed to what?
As opposed to, uh, as opposed to Nvidia saying that, Hey, if you buy my GPU, you have to buy Infinity brand or my buy my proprietary networking technology. So Nvidia, a Full stack, it's like an Apple Wasn't because some of those things are also in the stacks. Yeah.
That is in with, so, so what you're basically saying it's NVIDIA's, um, networking and NVIDIA's, GPUs, which you have Nvidia. So it's basically you've taken NVIDIA's networking out, put yours in Correct. Built that stack out, have a validated design.
So I've got confidence in a ship is are you seeing Well, that's fine. No, that's enough. I'll get Into trouble if I Exactly that.
Yeah. Uh, and, and in the case of Nvidia, we've actually followed their superpower design, right? Okay.
The only thing that we have replaced, we even use Sloan for orchestration. Everything is the same. We use, uh, we only thing we replaced is Infinity B with ethernet.
That, that's pretty much the only thing that we changed over there. So with as validated designs, you're saying, Hey, we've tested all these, you can mix and match. Mm-hmm.
Do you have reference architectures that say this is exactly how you can connect these and Yeah. And sizing guides and all sort? Absolutely.
That, that validated design actually is our reference guide. So it has all the best practices down to which knob to configure, how to tune your DLB dynamic load balancing everything right is, is all out there. Uh, and not only that, we've actually codified that into blueprints and abstract.
So if you don't want to, you're not a CLI junkie and you want abstract to deploy that, all of those best practices that we have learned and deploy them in one click reliably, repeatably abstract can do that as well. The abstract can deploy that entire reference design. So what part of Broadcom did you use?
What, what's Broadcom element in that? The Broadcom obviously the, the, the switch is powered based on Broadcom's, uh, switching silicone. That's what I figured.
Okay. But Broadcom also has Nicks. They have the tall two 400 gig.
Nick, was That in there as well? Yes. Yeah.
Okay. Mm-hmm. Okay.
So, um, you know, here at Juniper, we are our engineering led company and we do enjoy solving some of the most pressing needs for our customers. But once in a while, it feels nice to get recognized. And this is something that was, uh, recently announced by Gartner just three or four weeks ago.
Gartner reintroduced the Gartner Magic Quadrant for data center in 2025 after a five year gap. And we were squarely in that leader's quadrant along with the other giants in that space, right? So the Gartner Magic Quadrant being in the leader's quadrant, I think speaks for itself.
So I won't linger on this slide, uh, anymore, uh, but is powerful. What I do enjoy, uh, and I like even more, is that there is a companion report to the Gartner Magic Quadrant called the Critical Capabilities Report. And here is where Gartner goes into details of a vendor's capabilities to assess that capabilities against a specific use case.
And we were number one in enterprise data center buildouts, like I was talking about those 50 gig servers for enterprise data center buildouts. We were number one, uh, out there, uh, ahead of Arista, ahead of, you know, the other, you know, typical vendors you would expect to be out there. And, uh, uh, this is a testament mostly to our fabric management capabilities with abstract that you're gonna see in a few minutes.
And the topic of today's discussion was AI data centers. Uh, I did say we have a leading market share, but Gartner recognized that as well. Uh, we are only number two to Nvidia, right?
02 behind NVIDIA for AI networking, right? So again, great testament to a great validation of our technology and the momentum that we are seeing with AI data centers out there with our networking and security solutions. And it's not just analysts who are, uh, you know, raving about us.
Uh, we have customers as well. Uh, this is just a small fraction of the customers who are actually public references. And all of these, whether you're looking at some GPUs or service providers like Digital Ocean, we ion stream xai who's building this massive cluster in, uh, you know, called Colossus in me, in Memphis or some other, uh, enterprises out there like Sam, Nova, Wyoming, PayPal, uh, they all have trusted, uh, Juniper to build their AI cluster, whether it's for storage, networking, backend, GPU networking or frontend networking.
So, um, again, that was my final slide on, uh, the traction that you're getting, uh, with customers. And I'll take any few, few questions before turning it over to Bikram. Going back to your, uh, slide where you showed the, the, uh, hybrid AI cloud, how are you defining that compared to multi-cloud?
Are you seeing more of multi-cloud environments? So I think of hybrid, I think of on-prem and on the cloud, and then multi-cloud, of course being, you've got multiple clouds within that, not on-prem stuff. Yeah.
So I, so we are seeing, so I think hybrid cloud, maybe having loosely used term, it is multi-cloud as well. Uh, we are definitely seeing a lot of repatriation of workloads happening due to the AI wave, right? Because, you know, you, you need, the models are being trained on your proprietary data, right?
So whether, so we are seeing that applications moving closer to the data rather than data kind of moving to the application in public cloud. So we are seeing a lot of private cloud buildouts happening, uh, with, uh, the no whole AI wave. Uh, and yes, it is actually, you could say it's a, it's multi-cloud in a, in a sense that some applications are running in private cloud, some applications are running in, in public cloud in that sense, Yeah.
So data gravity and latency And data gravity kind of pulling applications towards private cloud. Yeah, I would assume security too, for particular industries. Absolutely.
Like privacy data. Privacy security reasons as well. Thanks.
Another question, um, until now, what I heard most about Juniper was missed. How does it fit in this big picture here? Mm-hmm.
So, mist, uh, started out as our, uh, you know, portfolio for managing campus and branch right data in data center. We always so had abstract for managing and operating, uh, data centers, right? But over the last few years we have, uh, integrated some of these technologies together.
So Mist is basically a leading AI ops engine. Like for, you know, today we're talking about networking for ai, but that's more of AI for networking, right? And abstract started out as an on-prem instance management instance, but now RA also has a cloud presence and their cloud presence is powered Bym AI ops engine to provide more predictive capabilities, uh, for data center as well.
So, simple answer, mist was for campus, abstract was for data center. But still going forwards, we are seeing more and more of that integration where we are bringing this AIOps capabilities from mist into abstract in a cloud instance as well. Okay?
Mm-hmm. Thanks. So Jim, spring, CDC, um, your customers, you know, you've laid out your three or four pillars is if the security aspect that most people are interested in, I mean, I'm just interested if you kind of did a pie chart, You know.
Yeah, I would say that it, when it comes to AI clusters, I talked about four pillars. Uh, you know, let's say performance, operational complexity, security, and open versus proprietary. If you don't have performance, the rest of the pillars don't, are not, nobody's gonna consider you, right?
So you need to first show that, you know, I'm neutralizing every other proprietary technology out there with an ethernet low cost solution, lower cost solution that delivers the performance. So I would say everybody's in there for the performance. Otherwise, you know, the remaining pillars do not even matter.
But yes, security becomes a, a differentiator for us. Our operational capabilities with RA and our security do become the differentiator. Once we have neutralized that performance is, is good.
Okay.