Securing AI Clusters, Juniper’s Approach to Threat Protection with Juniper Networks
AI clusters are high-value targets for cyber threats, requiring a defense-in-depth strategy to safeguard data, workloads, and infrastructure. Kedar Dhuru highlighted how Juniper’s security portfolio provides end-to-end protection for AI clusters, including secure multitenant environments, without compromising performance. The presentation addressed the challenges of securing AI data centers, focusing on securing WAN and data center interconnect links and preventing data loss, which are amplified by the increased scale, performance, and multi-tenancy requirements of these environments.
Juniper’s approach to securing AI data centers involves several key use cases. These include protecting traffic between data centers, preventing DDoS attacks to maintain uptime, securing north-south traffic at the data center edge with application and network-level threat protection, and implementing segmentation within the data center to prevent lateral movement of threats. These security measures can be applied to traditional data centers and public cloud environments with the same functionalities and adapted for cloud-specific architectures. Juniper focuses on high-speed IPsec connections to ensure data encryption without creating performance bottlenecks.
Juniper uses threat detection capabilities to identify indicators of compromise, including inspecting downloaded software and models for tampering and detecting malicious communication from compromised models. Their solution employs multiple firewalls, machine learning algorithms, threat feeds to detect and block malicious activity, and a domain generation algorithm (DGA) and DNS security to protect against threats. The presentation also highlighted Juniper’s new one-rack-unit firewall with high network throughput and MACSec capabilities, along with multi-tenancy and scale-out firewall architecture features.
Presented by Kedar Dhuru, Sr. Director, Product Management, Juniper Networks. Recorded live in Santa Clara, California, on April 23, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/juniper-networks-presents-at-ai-infrastructure-field-day-2/ or https://techfieldday.com/event/aiifd2/ for more information.
Transcript
Good morning, everybody. My name is Karu, and, um, I'm here to talk to you, uh, through some of our security use cases for, uh, AI data centers. Um, I lead product management for, uh, Juniper security, um, both product development as well as, um, helping with, um, you know, building security strategy for Juniper security, as well as how we embed security across, um, all of our different product lines to ensure that we can deliver things like secure data centers, um, to our customers.
Um, part of the agenda is to look into some of the use cases when it comes to AI data centers. Uh, you know what, um, uh, covered a few things. This is about how do we help secure the van and data center interconnect links.
Uh, we'll talk about, um, uh, you know, using the example of a couple of, um, uh, threats. How do we, how can we help focus on the data and prevent data loss? Uh, and then one of the challenges we have when it comes to, uh, AI data centers is the fact that, um, you know, your, everything gets amplified, got more scale, more performance, um, more, uh, tenancy, um, requirements.
And so how can, what have we done to kind of help, uh, address some of these challenges? Mm-hmm. So that's kind of the structure of what, um, um, you know, we will, we'll follow today.
Um, I, I think that's key. Can you say that again? More scale, more latency, and what was the Third one?
Uh, multi-tenancy. Okay. Um, and I think this is not new.
I think both Bikram and Profu have covered these to a large extent in, in the first two sections. So, um, they, they apply to security as well. And then we've done a, uh, we've, we've implemented, um, higher performance, higher scale and multi-tenancy capabilities into, um, our security use cases too.
Um, so when it comes to, uh, AI data centers, it's about, um, being able to deliver or, or run data centers closer to, uh, where your data, where your users are, where your, uh, applications are running. You want to have, um, uh, data centers be, uh, local and regional. Sometimes you, you want to be able to train your data in one location, move it closer to where you can, um, you know, uh, deliver, um, inferencing may, maybe it's through a public cloud or it's in a co-location facility, but essentially you have, uh, different locations where your data's located, different locations where your data centers are going to run, um, and, and, um, um, access this data.
Um, from a protection and a security standpoint, um, essentially you have three or four different, um, use cases. Many of these apply to standard, um, ai, uh, data centers. And when we look at AI data centers, they, uh, don't necessarily, uh, I mean, they, they evolve, uh, but we break them down into four different areas.
You've got data centers that need to connect to each other, and so you want to be able to protect traffic, um, between those different data centers. You want to ensure that the data centers are up and running. Um, you know, somebody's looking to, um, bang on your door and take that data center service down.
You want to, uh, ensure that that doesn't happen. So things like DDoS protection become important. Uh, the next use case is, um, as you are allowing traffic users, uh, or other, um, um, kinds of traffic to get in and look at your applications, you want to be able to secure that KnotOut traffic.
So you've got the data center edge, or data center perimeter firewall, which is where you wanna focus on the application. You wanna focus on network level threats. You want to be able to weed out the bad stuff with things like, um, you know, uh, threat feeds.
You want to be look, um, able to identify, uh, zero day, uh, attacks or zero day threats in, in some way. And so, how do we help do that? Um, the third one is, once you get inside the data center, you want to be able to, um, uh, ensure that you can do, um, uh, segmentation.
If you've got something compromised, it is not moving laterally within your organization and, um, um, you know, compromising other applications. Um, you know, or, or I think one of the examples Profo used earlier was, um, uh, going from one tenant to another tenant from one user's data compromising, somebody accessing somebody else's data. So you've, you've got things like the data center core firewall, uh, addressing your inter DC use cases.
And then as we, um, you know, when we looked at the architecture or the, um, the high level diagram for, uh, data center, um, and how they're disputed, you may also want to take some of your data into the public cloud. And then how do we bring these same use cases into the public cloud? I don't think it's anything, um, new.
It is still the same data center parameter, firewall, data center, core firewall, uh, functionality. Um, it just needs to, uh, it, it needs to be adapted to operate in a public cloud environment. Um, breakup between using the AWS terminology, breakup between different VPCs, um, ensure that you've got secure connectivity into the cloud and so on.
Same thing applies to Azure, GCP or wherever else you're running your data. Um, and workloads. Um, let's look at, uh, two things.
One is from a van and data center interconnect standpoint, you want to ensure that every, um, data center that is connected or all of the data centers that you have connected, um, have a link that is in some way protected. So the two things we do here, and what we have done is, um, we have ensured that, um, if you are doing some kind of a, a connection between, let's say, a firewall, the SRX is our product, uh, our firewall product line, the MX is our data center kind of gateway and routing, uh, product line. If you do have a connection, how do we ensure that that, uh, data center's up and running?
So we, we integrated, uh, tdo s mitigation, uh, capabilities into our, uh, data center gateways, whether that's MX or the PTX product line. Uh, we can detect, uh, somebody doing things like volumetric DDoS attacks meant to take down your applications or your, um, um, data center links as a whole, um, identify those and then, um, start blocking against those DDoS attacks there. The second thing is, when you do have these connections, many times those connections may not have a dedicated link.
They may be going out over a high performance public link, um, the public internet. And so you want to ensure that when you're moving data, your customers are bringing data to a, uh, Kenar. Yes.
Before you go onto that Jack Poller paradigm technical, so you, you're making a distinction there on public links. Is there a reason you're doing that or she behind a firewall? Are you looking at zero trust at all in this Environment?
So we are looking at zero trust, and I'll come to that. Okay. Um, but in this particular case, we are just looking at the van link.
Mm-hmm. Uh, you're looking at an open public link, uh, where you, if you're transferring data, you may want to bring in a level of, um, encryption. Mm-hmm.
That's where the IPSec piece comes in, right? PPN IPSec. And, um, many times, and, you know, when it comes to, uh, AI data centers, it's large volumes of data, right?
Uh, some of that needs to happen at high performance, but if your IP SEC links are few gigs, you're going to act as a, a choke point. Mm-hmm. So some of the things we did was, um, whether it's the SRX, which is running as a firewall, or things like the mx, which are your data center gateway, we can do, uh, high high speed IP set connection, um, you know, all the way up to, um, single flow being 300 gigs.
Right? Um, so whether you have a public link or even if it's a, a private link, if you do have a need for encryption, you do get, um, that high speed capability. Alright.
Does that answer your question? Yeah. I just, you know, in the context of ai, is it performance that's driving the need for not using encrypted links versus encrypting your data all the time?
Right. So from a general data security point of view, we want to encrypt all links all the time. You Wanna link, uh, encrypt links all the time.
All the time, Right? Yes, Yes. And, and we can help do that.
Right. Um, I was just being, you know, it, it's up to the, it's up to the, uh, customer Okay. Um, the, the operator, how they want to do the encryption.
Right. I, I would recommend doing it all the time. Mm-hmm.
But that's me, you know? Yeah. Might have a different design.
Okay. So it's flexible. I just wanna make sure that, that, that there wasn't some specific reason we were doing No, No.
Okay. No, it's, uh, it's a choice. Yep.
But, uh, we made it easy and we built it in the products. Mm-hmm. So, um, if you had to turn it on, it's, it's, it's, it's a couple of lines code.
Yep. Right. So it's, it's there.
And, um, I'll, I'll come to this in the last slide, but essentially, um, if, if you're doing a few gigs of encryption, most products and most firewalls can do that. Yep. Um, if you need, um, you know, what, what happens with, um, AI data centers, you're moving large volumes of data.
You, your, your strain on your, um, IPSec connection is also not a couple of kinks. Could be tens, could be, you know, multi tens of gigs. Yeah.
And so you need higher performance that's available as well. Gotcha. Thank you.
So, like I said, single flow, I think because we are using, um, uh, we built this into the Juniper ASEC to be able to do, um, IPSec, that's looking at 300 gigs of, um, uh, uh, single tunnel performance. Mm-hmm. So, gotcha.
You've got that wide range. Um, when it comes to, um, security for AI data centers, uh, I think the, um, um, you know, I'm just using this slide, uh, to level set. You've essentially got your backend networks.
You've got your private frontend networks and, and things like your public frontend networks. Um, when we break these down and we look at, uh, some of the different, um, security use cases that show up there, you've got things like data loss, data contamination in the backend networks. You, you may have, um, you know, you don't want, uh, intermingling of data.
That's where the multi-tenancy piece comes in on the private front end networks. You've got, obviously, you know, if you've got something compromised, if you've got a compromised machine learning model, which is what the example we'll see, you have, um, uh, IP theft or, you know, data loss, that that can potentially happen. Uh, and, and you've got tampered models that you want to be able, uh, to detect.
And then finally, on the public front end, you want to allow, you know, the right kind of user authentication you want to do, uh, intrusion prevention. You don't want somebody to come in and, you know, um, uh, um, uh, waste your resources. Uh, so these are the different use cases that we, we, uh, were looking at, uh, from a security standpoint for the different kinds of networks, um, you wanna use Explaining these at a higher level.
So backend, those could be hosted, they could be in your data center, They could be in your data center if you want to run training models, they, you could be using somebody else's data center where you move data run, uh, or, or your running your training models. Right. So you could be using in Google Cloud or, um, Uh, AI data center as a service Or nvidia.
And then, so private fronted, what do you mean by private fronted? Public fronted, Okay. Uh, so private, uh, front end is essentially, uh, your data centers where you bring, you've, you've already got, um, you've, you've, you are using models, you're bringing them, uh, down and LLM model, and you are running it against your data set to, um, you know, um, um, whether it's within your organization to be able to, um, uh, do analysis.
Uh, for example, on the threat research side, uh, we may, we may be using a, uh, LLM model for, you know, doing threat analysis. Um, we, we may bring different kinds of data. We may bring data from, uh, our honeypot information.
We may bring data from customer, um, submitted information. We may bring publicly available information. So that's where we are running all of our, uh, inferencing and, and, um, um, to get the results.
And then public front end is, if you're running, um, uh, an application that you want to, um, have as a front end app, this could be just any other kind of public app that's connecting to, um, uh, an l LM based service on the backend. Uh, and you allow a large number of users or, um, even, um, publicly registered users access to, uh, those applications, I guess. So either Coming in or going out to the public.
Okay. Like a chat frontend interface. And private is, uh, OnPrem intracompany Intra company within your application.
Okay. Thanks. So other than tra Kimberly Bates, other than movement of data mm-hmm.
Um, how else are you detecting the data? Correct. Um, Um, especially when you talk about things like model 10.
Correct. So that's the example I was going to use. Okay.
Uh, in the next slide. So next Slide. Okay.
Well, that Set it off. Well done. Way to go camera.
Okay. We, we, we talked about this earlier. Yeah.
Um, but, um, essentially, I mean, the example we, we've, I think Pfu brought this up, we've probably heard about, you've probably heard about this in the news, is, um, you know, things like publicly available marketplaces like hugging face, um, um, are, are, um, exist to be able to share and, and, you know, um, different kinds of models. What you, um, are doing is a user's probably downloading some kind of model to their, um, private front end in this particular case to be able to run against their data sets. Um, and let's say the model is compromised, I think that's one of the use cases.
That's one of the things you were asking, right? You've Got Well, yes, I'm, I'm more saying, okay, so that's, I'm bringing a model that's already been compromised, correct. In, correct.
Yes. That's a totally different thing than One of the things that I'm very worried about is that once I have created the model and I've deployed the model, somebody goes in and alters that code. And how would you know, how would you know that?
That's the one that scares the crap outta me is because I'm like, okay, so I, I'm putting out these models that are running, let's say, at the edge, correct. You know, water monitoring, water monitoring electricity or whatever, and it's at that edge that somebody goes in there and does something, I'm gonna think about somebody. Right.
Whatever. Um, and that's kind of the place that, So when somebody goes in and tries to tamper with the model or any application, there's, there's potentially a few things happening. One of those is, um, Oh, I'm sorry.
I wasn't even tapping all that brilliance away. You can still here a little bit. Wanna hear, go ahead.
Do you wanna re-ask the question? No, no, you can go. Okay.
So the, the question is, you've got a model, you've trained the model, you've deployed the model, and somebody goes in and alters the model. And, um, how do you protect against that? Yeah.
Right. So, um, from, from a network standpoint, that basically means, uh, a couple of things. When you've got a model, uh, that is, um, uh, tampered on that look, that's looking at, um, your, your, your data, and if it's doing nothing else, one of the challenges there is, um, is, is that, uh, giving you the right outcomes for, um, the data that you had trained the model for.
The second one is if you have tampered and played around with the model, somebody's looking to do something with it, potentially mm-hmm. Exfiltrate data or, or, uh, communicate with and give access to somebody outside the network, um, uh, backdoor into the network. And so those are things that we can look at, um, and we can identify to cease malicious activity or malicious behavior, we can either block those or we can report on those that allow you to go in and investigate further.
Um, so that's what we are looking at from both of these use cases. In this particular case, I'm bringing in a compromise model. Mm-hmm.
I bring that onto my, uh, network. And essentially when I install that, that, uh, model is, um, looking to open a back door to somebody that's giving them access to my network. So how do we detect some of those different things, uh, so that we can stop you from bringing in a pad model, um, identify a bad model if it has already been brought in, and be able to, um, um, minimize, uh, the impact of that, um, keep it contained.
So those are the two different use cases. And essentially the first thing is, if you're bringing in a bad model itself, uh, we can have, so we have two firewalls in this particular example, you've got firewall A. Is it, is it really, is it fair to say that what you're really doing is you're detecting that your, your threat detection, what threat protection is looking for specific IOCs indicators of compromise?
In this case, you're looking at things that are network related Yes. As indicators of compromise. So you're looking for, if you see somebody coming in to look at the model that isn't doing it, where the, that shouldn't be looking at the model coming across the network.
So in this case, you're downloading a model model. Right. And, um, just to set this up, uh, we are using two of those four use cases I mentioned earlier, right?
You're using a data center parameter firewall, and you're using a data center core firewall. Right? Right.
And if you're downloading a model that is compromised, um, that firewall is capable of detecting, uh, a compromised model, you know, we, we've built in, um, a machine learning algorithm onto the firewalls itself that is capable of looking at, um, um, indicators in the file that is being transferred. Okay. So if you've got a model that, uh, somebody has tampered with mm-hmm.
Um, or whether it's an AI model or it is an application that has been tampered with, um, there are certain signatures in the file that is being downloaded itself that, um, indicate, um, that a, uh, somebody has tampered with, uh, with it. For example, if they're using a, um, old unsupported version of Visual Studio to make some of the changes, that's one example, right? And then you've got other markers that we can look for.
So you're, when you aggregate all of those, that gives, you know, there's less chance that somebody good is using bad software to create Right. You know, a product. So you're doing more than just looking at network.
Yes. Intrusion. We're looking at the actual, So more, more than just network effects, network intrusion detection, some of those types of things.
So you're actually doing an inspection of the data being downloaded and saying, This is not of model being downloaded in This What? I'm sorry. Of the software.
The Software, sorry, yeah. Software being downloaded and saying, we expect a model to look like this, and this model has signatures in it that it shouldn't have of one form or another, that indicators that it's been tampered with. Yes.
And that can give you, um, a lower level of probability, which you may want to alert on, and a higher level of probability that you may want a threshold that you may wanna block it. Right? So that's one use case.
Yep. Bringing in something. Now, let's say you brought it in anyhow, before you had security in place or through, through another mechanism, and you've now got, um, that model that is compromised.
It's trying to connect and commit. That's it. I have five minutes.
Okay. Let's speed up then on the Solution. Uh, so that model that's trying to communicate with, um, uh, a threat actors infrastructure, this typically happens.
I mean, why is somebody compromising your model? They want your data. They're not doing it for fun.
They, they, they want your data. They want to be able to get into your network so that they can move laterally so they can get to something else. Uh, if it is a critical infrastructure, for example, they may want to, you know, take over more and more applications in your critical infrastructure.
So if that's going out, then um, that same data center parameter firewall, um, um, has, um, several things. It's got direct signatures like command and control feeds that we know where there are known threat actors, IP addresses, or infrastructure ips that we can block against. Um, typically, you know, these models are looking at different ips, looking at com, connecting to different kinds of, um, uh, domains.
Uh, so things like domain generational algorithms can, or DNS security can be used to help protect that. We built another machine learning, um, um, algorithm called Encrypted Traffic Insights. Even if your connection to, um, uh, the, the, the threat infrastructure is, uh, encrypted.
There are a lot of markers in that encrypted tunnel. Um, the IP address that's connecting to the handshake, the certificates, all of these, the beaconing behavior that is there, all of this can be used to save the high level of confidence that that model is compromised is talking to some kind of a threat, um, uh, actor's infrastructure. And we need to block that or we need to report on that.
And then finally, um, one of the things, um, uh, that happened with the hugging face, um, model, um, corruption was a lot of the models put in, um, uh, river, um, like a reverse shell connection. And so being able to detect those, uh, can help you identify a compromised model in place or compromise application in place and help, uh, block that. The third thing, everything's come in, everything's running.
Now you've got, uh, traffic flowing from this model, not just to your data, but you know, trying to go somewhere else. That's where the data center core firewall comes into, uh, play. And, uh, it is able to, um, um, you know, uh, identify traffic moving from one, uh, model to a different data set, uh, that movement across, um, multiple tenants, uh, of data, and you're able to help bring that in.
Uh, if we have detected that my model is compromised, we have a feed that we, um, that we can share directly with the QFX switches in the data center called the infected host feed. So if I've detect, um, um, my firewall here has detected my, uh, IP address to be malicious, I can share that with the QFX itself. And, uh, you can help contain and block, um, the, the compromised, uh, ips or servers, uh, within that environment.
So that's more about, um, you know, reducing the spread there. Uh, and finally, um, you know, uh, we have introduced, uh, things like, uh, multi-tenancy capabilities on the firewall itself. So, uh, we can help prevent, um, uh, or, or block an unauthorized application from talking to somebody else's, uh, data.
So again, this is more about containment. How would this protect against, um, like the people who want to tamper with your models, not to exfiltrate data, not to go other places, but to poison your business decisions and that kind of stuff internally. And then that sort of also relates to the second question about what about insider threats doing all the things.
So, um, the insider threats, again, that's where you have the data center. And so I'll just take that one. That's where you have the data center core firewall, which essentially is you have access to this application, not to this application.
So, um, that's more about where you have access to, that's implementing things like a zero trust, um, security approach. Uh, so I'm the insider, I have access here, but I'm trying to go in here, but I don't have a, uh, access to an application X. Um, that's where, um, a zero trust based policies would, uh, apply.
So things like, uh, like a universal zero trust architecture, uh, integrating with your, um, active directory so you can set the right kind of permissions. That's where it would apply. And I think the other one is I do have a compromise model.
I'm giving you bad data. I'm giving a bad response. Mm-hmm.
Um, that one's, uh, a little bit harder from a network and firewalling point of view. Yeah. Uh, but we are looking at ways where, um, um, you know, um, to, to, we are kind of, that's in the research phase.
We, uh, we understand that's a challenge. Yeah. But we, you know, um, a network based approach necessary can't solve that.
That's more, it's not based between the model and the data. Yeah. So you need to look at the data level.
You need to understand model context. So those, those are a little bit different. Mm-hmm.
So Do you, to get the hook, uh, we'll be, have time to get more que ask more Questions. It, it's, I I did have a question following up on, uh, Karen's question. Okay.
Which is, uh, how do you handle, um, what would be unprotected ethernet traffic, like our DMA traffic? Um, can we take that later? I don't have, I don't have right now.
Okay. Okay. Um, can I get two more minutes to finish up?
You need to finish up in the next 30 seconds. Okay. 30 seconds.
Um, so everything gets amplified in an AI data center. We, we talked a little bit about this. It's much higher performance, it's much higher scale, and obviously multi-tenancy.
Um, so we have been thinking about these challenges for a while. We did, um, uh, take some steps to get there. 4 terabytes of, uh, network throughput with multiple 400 gig, a hundred gig links to allow you to connect into your data center fabric itself, or AI data center fabric.
All of those ports can do ec. Um, and then the other thing we did was, uh, we, we, uh, introduced machine learning and, uh, algorithms, uh, into the software code itself to be able to detect, um, um, you know, AI based, um, threats. Things like using AI to be able to take malicious, um, um, or compromise models.
Uh, we introduced, uh, integration directly into an IP fabric like EVP and vxlan. So we can make it easy to integrate security, uh, into your data center fabric that allows for multi-tenancy, that allows for multi-tenancy across a data center stretch. And then finally, I heard a lot of questions earlier about scale up versus scale, uh, out, uh, that is, um, so we introduced a concept called scale out firewall, where you can connect multiple of these firewalls together.
Um, they act and function as one logical firewall rather than having to, um, uh, treat them individually, separately from an operations point of view, from a, um, CLI management point of view and everything.