Building AI infrastructure for a Global Network | The Six Five Summit
With the launch of Workers AI, Cloudflare has deployed GPUs at scale around the world. In doing so it confronted many challenges that would be familiar to any other enterprise looking to deploy accelerated computing at scale, as well as some unique challenges. Syona Sarma, head of hardware engineering at Cloudflare, will share lessons they learned in the process. Audience members who’re responsible for enabling AI and machine learning in their own organizations will leave with new ideas and tools to support their companies’ AI transformation.
Transcript
Hello, Dave Nicholson here. Welcome back to six five Summit. We've got a very exciting guest lined up, SMA, senior Director, hardware engineering at CloudFlare.
It's important that you remember what I just said. Hardware engineering. We're gonna have an awesome discussion about infrastructure in the cloud and AI space.
Welcome. Thank you. How are you?
I'm good and excited to be here, um, and look forward to sharing more about the CloudFlare perspective on building out AI infrastructure. Absolutely. Well, tell us more about CloudFlare.
For, for people who aren't familiar with CloudFlare, a little bit of its history, um, give us, fill us in a little bit. Who is CloudFlare? What do you do?
CloudFlare started about 13 years ago as a CDN provider, and as time evolved, started supporting different types of services that included network and security, specifically things like DDoS mitigation, VPNs bot management, et cetera. And finally, we are in a place where we're able to support a developer platform that we call workers in order to build out a developer ecosystem. As of last year, we've start ventured into the AI insurance space with a product called Workers ai and a couple of other AI platform solutions.
And given that we are an Edge network with a presence across the world, we think we have an inherent advantage with latency, which serves as value entrance types of application. Okay. So, so first of all, CDN the TLA three letter acronym Content Delivery network.
Correct? Yeah, that's, so you, so you have a history of having this infrastructure. Gimme a little bit more of a sense of the scale of the network that you built out.
What does that mean exactly? Like what's what's your reach? It's, it means a presence in all of the different regions of the world.
Um, and we separate out from multi colo points of presence to small data centers in the most remote corners, um, that we think is the specialty of CloudFlare in terms of reach. Um, another thing I should mention is our architecture is built such that every service, irrespective of what, uh, bucket it lands in, runs at the same level of performance, is as secure and reliable as any other place in the world. Okay.
So, so, so on the, on the subject of services, you, you have, you mentioned workers, workers, ai. What, what do, what do you, how do you, how do you, what do you call it? Yeah, so, uh, we built this AI entrance is a service solution called Workers ai.
It's built on our developer platform, which is called Workers. Um, and what we are attempting to do here is to support a developer ecosystem. It, it's more than just an inference product, though.
We are attempting to provide an AI platform solution that includes inference as a service, but also gives you a place to store and, uh, use and, uh, train your models pre-train, fine tune your models. I wanna be clear, we are not in the full training space yet, and don't think we've ever been given. We are an edge solution.
So, um, so, so far we've talked about, uh, some, we, we, we've loosely talked about, uh, the hardware infrastructure and the sets of the scale of your network, but let's get down to it. What about GPUs? Uh, and I I I, is it fair to assume that that, uh, inference as a service includes the deployment of GPUs in your network?
Oh, yes. Um, so we started out, like I said, as a CDN providers. So we already had an existing compute and storage infrastructure, um, and we use that and leverage that to add GPU attached to our servers in order to be able to support inference.
Um, I do wanna mention, in addition to inference, we also have in our two solution, which is our object stored solution, because inference is really in the context of a larger system. It's not a workload on its own. Um, and along with a couple of other tools that help you rate limit and manage your workloads if you're running training elsewhere, and one will run cloud inference on CloudFlare.
We have a platform solution available, um, and we are in the midst of rolling out GPUs, um, to be in milliseconds of eyeballing latency. Um, so it's going pretty well so far and we have a growing list of models that we're supporting in partnership with, uh, hugging face in other places like that. So when you, so you have this network, uh, and now you're going to do inference as a service and other things in the AI space, is it as simple as opening up the servers that are in these various points of presence and, uh, sticking a GPU into A-P-C-I-E slide?
Is that all you have to do? No, it's much more complicated than that. Um, and particularly from a hardware side, there are a bunch of considerations and challenges that we faced that I, I would like to talk through, um, in case it's helpful to someone who's either bending out their own infrastructure or choosing with solution to run that AI applications on.
What are certain things that you should be thinking about, um, to start starting out with what is the type of accelerator you want to provide? Um, so it starts by considering what is the space in the target market segment that you're looking at. So with CloudFlare, we're focused on entrance and entrance in the context of the larger system, not training for a start.
The second thing is to understand what types of workloads, even within inference you wanna support, because the characteristics of these workloads differ widely and from a hardware perspective and given the long hardware product life cycles, it's really important to understand if you would want a one size fits all solution or a custom built accelerator solution. Yeah, you know, there's been, I I think we, a bit of a theme has developed, um, during six five Summit, uh, when we're talking about infrastructure and cloud and ai. And that theme is that, um, there isn't a single tool to do all jobs and the thousand watt or 2000 watt GPU is not the only way that you can accelerate AI operations.
I, I just wanna kind of double click on your mention of choosing the right inference accelerator. Can you, can you gimme an example of what some of those choices might look like? And I imagine that sometimes the, you know, the considerations have to do with power and efficiency for a given job, but what, what's a version of an accelerator that you offer?
Um, a lot of this comes down to, in the case of CloudFlare, what we can support in our infrastructure, like I mentioned, are, um, racks and servers have different system constraints that, that they operate in. But our, our, um, goal is to provide the same level of performance across the board, which means we have to do a lot in terms of system design constraints and optimizing to them. Um, so things like power thermal, top of the list, right?
What, what is the system level current node power that you can accommodate? What is the rat level power that you can accommodate? What are the bottlenecks that you're most likely to see, whether it's memory, capacity or bandwidth or network.
Um, so kind of having a view of these bottlenecks and monitoring them as they change is really important. Um, as we build out a, a roadmap and I wanna just focus on, on the roadmap because a single solution is not going to meet all your needs at the pace that things are changing. So, um, we have more of a phased approach where we deploy a certain type of an accelerator.
Right now we've been focused on n and m's most, mostly because that's what our customers would like to see. Um, but it's smaller to mid-size models with higher throughput and latency requirements right now, um, in the future we want choose something that's more fine tuning, more custom build your own type of use case, which at which point we might want to transition to a different type of vaccinator. So it's really becoming a multi-pronged approach where you have different solutions for different types of solutions.
And our, our attempt is to make that I visible to the customer in terms of what they see in terms of the KPIs that they need. Interesting point on making it invisible to the customer because I would, I would argue that we're coming out of an age where the mantra was, uh, cloud first. Uh, I don't care about the infrastructure, infrastructure doesn't matter.
In fact, in fact we're running serverless applications, which some people started to believe didn't include servers on the backend. So I think the point that, uh, that we're both making here is it's really important to pay attention to the infrastructure and deploying it correctly. Now the customer doesn't have to because cloud floor CloudFlare or some infrastructure person will on the backend.
Can you, can you give me an example of, um, of, of how this kind of works in the real world? Maybe customer examples or, or at least industry use cases for the, for the inference as a service and you, you know, the, the, the new services you're delivering? Yeah, so we have a couple of different buckets of insurance workloads that we started out characterizing.
Um, the first was insurance, the second n and m and the third recommender, there was not as much interest in the recommendation system workload, primarily because it's on the edge and needs a whole lot of memory capacity, which our systems are not built to provide. So we decided to focus in on the other two buckets. And, um, in terms of the exact requirements, they vary between high memory bound workloads to high compute bound workloads.
So we started out by saying, um, we would like to support smaller models, say 7 billion parameters with the throughput requirements that would be satisfying to customers, but also have the capability in our systems to, to support larger inference models, say up to 50 billion parameters, and still provide the latency benefit at the edge from a customer standpoint. Well, very, very exciting times for the world of AI And CloudFlare. Si is there something that we missed here?
Um, what, uh, what else should people understand about the latest and greatest from CloudFlare? We, We are looking at an AI platform solution that includes workers ai, which is serving inference, but we also have, uh, two other products that I wanna mention. One is AI gateway product, and this is meant for customers who are not running, um, training on CloudFlare and would like to move to inference on CloudFlare.
We have tools that allow you to cache and rate limit and run data analytics and give you a seamless way to transition to CloudFlare to be able to run inference. The other product that I wanna mention is our R two, which is an object store, which will help you if you are trying to do any sort of retraining or fine tuning, which is a customer use case that is becoming dominant with, um, data sets that are more customized to the use case that you are looking at. So, um, it, uh, wanna emphasize we are, we are trying to become the platform for edge inference, not necessarily just the place to be able to run inference.
Fantastic. Si, thank you so much for spending time with us here at six five Summit. Lots of exciting stuff in the infrastructure space coming from CloudFlare.
Stay tuned for more from six five Summit coming right up.


