GPYOU: Building and Operating your AI Infrastructure with Juniper Networks
AI infrastructure is a critical but complex domain, and IT organizations face the pressure to deliver results quickly. Juniper Networks shows Juniper Apstra as a solution to streamline the management of AI data centers, providing proven designs. Kyle Baxter emphasizes the necessity of a robust network foundation for AI and ML workloads and highlights the challenges of traditional network management tools, which often overwhelm users with data, making it challenging to pinpoint root causes and resolve issues efficiently.
Juniper addresses these challenges by offering a comprehensive solution built on the Apstra platform. This platform features a contextual graph database, intent-based networking, and a vendor-agnostic design approach. Combined with Mist AI and the Marvis Virtual Network Assistant, Juniper aims to provide a holistic view of the data center, moving away from managing individual switches to focusing on delivering desired outcomes. This approach simplifies the complex network, allowing for precise identification of root causes, related symptoms, and impacted applications or training jobs.
The presentation focuses on managing the training side of AI and ML clusters. It highlights Apstra’s global capabilities to manage various data center networks, including back-end, storage, and inference networks, for large enterprises. Juniper offers designs and flexibility to manage any network design using a single tool. The key takeaways are the ability to design, deploy, and assure network operations, utilizing Juniper’s leading switching portfolio and security solutions. This aims to provide a streamlined, efficient, and reliable AI infrastructure management solution.
Presented by Kyle Baxter, Head of Apstra Product Management, Juniper Networks. Recorded live in Santa Clara, California, on April 23, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/juniper-networks-presents-at-ai-infrastructure-field-day-2/ or https://techfieldday.com/event/aiifd2/ for more information.
Transcript
So my name is Kyle Baxter. I head up the abstra product team. Today we're gonna be talking about GPU on how you are gonna be able to manage and deploy your AI infrastructure.
So we're gonna walk you through your AI journey. We're gonna give you a quick intro on what is abstra for those that don't know, and then we'll get into how you use abstra to design your AI data center, how you manage your data center at scale for deploying and how you operate your data center. And so let's get into a little intro on what is abstra.
So I think everybody will agree here that you can't build a great building without a great foundation. So if you have a weak foundation, you're going to have problems. And the network is the foundation in AI and ML training and inference jobs.
The problem is the network is the problem. There's always lack of insight. How do you know why you're training jobs slow?
Why are you getting congestion? How do you deploy faster when there's business needs to go deploy, you know, new fabric or expand faster? How do you get that speed and get that reliably?
Um, and then how do you prevent outages and make sure your network is operating reliable? This is is where Juniper comes in, um, because the network is complex. There's so many elements that make up a network When you look at, you know, spines, leaves and all the links and we're amplitude and amplifying that with ai when we're talking about 800 gig networks.
6 and beyond, they're gonna get some crazy speeds that if anything goes wrong, it's gonna go wrong fast. And the problem with most network management tools out there is they flood you with data is, which is nice, but then you're looking for the needle in the haystack. You're just looking through all this data points trying to figure out what means what.
And that's what we've done differently at Juniper. We've brought together our technology together to build something different and something better. And so brought Abstra, which brings the industry's only contextual graph database with intent based networking and designs with a vendor agnostic approach.
We've combined that with mist in the AI there to bring with the industry's only AI native platform and Marvis virtual network assistant to the data center. And so we wanna look at the data center a little differently rather than as a collection of individual switches that you're configuring one by one. We wanna look at it holistically as a solution on how we can deliver the right outcomes.
And we do that by cutting through the complexity. Instead of seeing all these random events and dots and trying to figure out what's going on, abstract cuts through that complexity by bringing context to that data, we can pinpoint exactly what's the root cause that you need to address and fix, but what is, what are related symptoms that you can ignore that you don't need to worry about, that you can shift all that aside. And then what are the impacted applications or training jobs because of that issue that we have pinpointed?
We can do all of that for you and we'll walk through a lot of that today. And what this delivers is a complete solution from Design Deploy assure that can run on our industry leading, um, switches and switching portfolio with the integrated solution from security that we talked about in another session and be able to bring that to you. And we're gonna focus a lot about on the training side of AI and ML clusters today, but I wanna make sure it's clear that apps can manage more than that.
Thank you. We've been managing data centers with all over the grow globe for some of the largest enterprises out there that we can do. Not only the backend networks that we're talk about today, but storage networks, inference networks, or any kind of other data center network.
We have the designs, we have the flexibility to manage any design and you can use one tool and we'll see that here in, uh, in a later session. We talk about operating on how we can use one tool to manage multiple different networks.