Scaling Up The Cloud Computer with Oxide Computer
Cloud computing has been the most significant platform shift in computing history, allowing companies to modernize and grow their businesses. While cloud computing has accelerated businesses, it has begun to hit its limits. Companies need to extend their operations beyond the public cloud for reasons like locality, security, sovereignty, and regulatory compliance. However, operating infrastructure outside the public cloud often feels like a step back in time, relying on traditional rack-and-stack approaches that are inefficient and lack utility.
Oxide Computer Company aims to address this by bringing true cloud computing on-premises. To achieve this, they’ve built a completely different type of computer, a rack-scale system designed holistically from the printed circuit board to the APIs. This approach delivers improvements in density, energy efficiency, and operability, all plumbed with software for operator automation. The goal is to provide businesses with elastic, on-premises scalable computing that mirrors the efficiencies enjoyed by hyperscalers.
The Oxide system features a modular sled design for easy component upgrades, DC power, and a comprehensive software stack, including firmware, an operating system, a hypervisor, and a cloud control plane. This design enables elastic services like compute, storage, and networking, with multi-tenancy and security built in. The company has seen a surge in demand, expanding manufacturing operations and targeting various verticals, including federal, financial services, life sciences, energy, and manufacturing. Oxide focuses on providing modern API-driven services, improved utilization, enhanced energy efficiency, and a trusted product to alleviate public cloud costs and enterprise software challenges.
Presented by Steve Tuck, Co-Founder and CEO, Oxide Computer. Recorded live at Cloud Field Day in Emeryville on October 21, 2025. Watch the entire presentation at https://techfieldday.com/event/cfd24/ or visit https://oxide.computer/ for more information.
Transcript
It is Oxide computer company. We, we kind of put the computer company part in the name, uh, to remind folks that while we are a company board born out of the cloud computing era, uh, we very much make computers. Uh, it's just that we make computers that look nothing like what one can purchase and deploy on premises today.
And a lot more like the kinds of computers that the cloud hyperscalers have built for themselves to deliver cloud computing. Now, cloud computing has been, uh, the biggest platform shift in computing history. It, it is the foundation that allowed companies who believe in investing in technology to drive revenue, to be able to modernize and grow their businesses.
And, and this is not a contrarian viewpoint. Uh, you know, starting in 2010, 2011, the cloud computing industry really ushered in this ability for companies big and small to be able to build the digital experiences that they need at the end of an API, that they did not have to wait two months to be able to launch a new project on new infrastructure unboxing systems and deploying things. They could have a team that could hit an elastic compute or storage or network service and begin building software.
And over the next 15 years, we have seen the fastest growing sector again in technology being that cloud computing set of services. That has been wonderful for that the, the being able to accelerate what these companies need. But it has started to hit some limitations.
And those limitations are positive ones. They're not because the, the public cloud computing providers are, uh, are, are providing limited services. It's that businesses need to extend the kinds of things that they are doing, storing data, doing analysis on that data arming teams that need to operate things in areas where they can dictate locality security, dealing with sovereignty and regulatory compliance reasons.
Um, things that have to sit outside of that rental only public service model. And those problems are only getting worse with each passing month. When you look at how can I operate infrastructure outside the public cloud, things take you back about 10 to 20 years.
They take you back to when I was at Dell, first 10 years of my career, uh, and we were selling people servers, one U or two U2 U servers, and then someone would go buy a switch and then maybe they would buy some storage and then they would find a software provider to run on top of that, and they would try to integrate that into their own kind of snowflake of infrastructure eventually call to private cloud. And if you fast forward 15 years later, after I spent 10 years at a cloud computing company joint, the state-of-the-art on premises does not move forward very much. You know, the, the, the, the case is still to go and, and buy one U2 U servers and then add another four vendors worth of kit into a rack that you go deploy.
And the efficiency and the utility of that infrastructure truly has not modernized much. Um, whereas if you talk about what have the hyperscalers done for themselves about that same timeframe, 15 years ago, they took a different path. The hyperscalers said, we need to reimagine computing for a cloud computing future.
And that meant that we have to start building holistic systems that deliver on improvements in density, energy efficiency, operability, plumbed with software to deliver operator automation so that I can have far fewer people that need to manage my infrastructure that can stack together into rows and data centers and regions that can deliver energy efficiency. That is two to four times better than that rack and stack approach. And that is API driven for the end users.
And this was to enable the teams at Facebook to go fast and Amazon to go fast and Google to go fast. Um, but as we were leaving that cloud computing era ourselves, and we're thinking about the next company, um, that we wanted to go start, um, we were deep believers in cloud computing. We believe that absolutely is the future of computing, but we also believed as such that it needs to be ubiquitous.
It ha the, the capabilities that operational efficiency and the developer capabilities have to be accessible everywhere businesses need to run. And for that to be the case that meant you had to have true cloud computing on premises, you shouldn't be living in this false dichotomy of do I want kind of modernity and elastic infrastructure that moves the business forward faster? Or do I want control, control over costs, control over locality, control over, uh, what I do, where I do it?
And, and so this was, uh, that was sort of the backdrop behind us starting oxide. And, uh, with oxide, it meant if we were gonna truly bring cloud computing on premises that we had to build, like the hyperscalers had done for themselves a completely different type of computer. And that meant instead of rack and stack rack scale, that meant instead of AC power, DC power, it meant instead of a bunch of miscellaneous parts that were pulled off shelves and kind of woven together a holistic design that started from the printed circuit board itself, all the way up to the APIs that the end users would interact with.
And that if we could pull that off, we would be able to deliver true elastic scale computing on premises with the kinds of efficiencies that the hyperscalers get for themselves. Efficiencies that mean you now can do, uh, energy efficient computing 12 times more efficient than the rack and stack approach that you can get density that is two x per rack compute per watt than was possible even today in that kind of rack and stack approach. And that the experience could be one where rather than all of the boxes landing on that shipping dock, uh, on premises and it taking two months to get to usable services for developers, it being done in one hour.
And we, uh, I wanted to just share the very earliest days. This was a, uh, a napkin drawing of sorts from one of our earliest engineers, Anne ler. And, um, I think the thing that is so that we appreciate so deeply now is that it was so formative when we were thinking about what to go do now.
Yes, there was prior art because the hyperscalers had published some of the things they had done for themselves. And so we had this kind of idea of the shape of what a cloud computer should look like. Um, but the remarkable bit is just how close to that first drawing the actual system itself would land.
You can see, uh, the sled design, uh, a a kind of middle of rack set of switches, uh, with, with the connectivity running front to front and back. Um, and this was, you know, this is five years ago and it would take us years because we did the entire design ourselves. We did not use a reference design off the shelf, which is there are like literally two of them that every OEM uses, and I don't have enough time to go through the reasons why we absolutely were not gonna use that same reference design and break from the mold of kind of 30 years of, of OEM server computing.
Um, but then doing our own switch, ripping out the bios, this like legacy blob of firmware that causes unending issues around security and performance and observability in systems, and then taking that all the way up. Um, but this was a shot of us shipping our very first rack and, uh, it was a, uh, a a an incredible day a couple of years ago to prove that we could go build this thing. And, uh, when we had folks last out here for Cloud Field day, we were talking about not only do we prove we could build it, but now we've started to build a bunch of them when we started deploy them in customer use cases and getting production software running on top of them.
Um, and what we're excited about today is we have really seen, uh, the motion shift from us evangelizing and trying to communicate to folks that we exist and why you should care and the benefits this can bring to your business, to things flipping just in the last six months and being dragged now into environments where we can't build them fast enough. Hmm. And so this was a, a great shot that just came out.
Uh, we've been massively expanding our manufacturing operations. Uh, those are the PCBs that that's a, a copper layer exposed of what is a, a complex 20 layer board. Um, and again, we can't fab them fast enough and, and we're very, very grateful for our customers.
Uh, and, and we'll talk a little bit about how we scale this business up to meet, um, a, a surge in, in enterprise demand. Um, wanna just spend a quick minute on what the system itself is, because it is truly a system. It is not hardware, it is not software, it is hardware and software co-designed.
It is, uh, software engineering thinking about what does the hardware engineering team need? Is hardware engineering thinking about what the software engineering team needs to, to develop a holistic system. I mentioned there is a, a bunch of innovation in the rack scale design that drives the kinds of density and energy efficiency that is critical today as enterprises are constrained, try to do more with less out of space and power in their data centers, uh, and things, uh, the parts coming out of the chip providers are not getting any cooler.
Uh, they're not getting any less, uh, power hungry. And rather than having to pour new concrete or figure out the next colo location, um, actual true hyperscaler computing affords an opportunity to consolidate what you're doing, almost two to one, almost opening up 50% of data center capacity as it exists today, while being able to deliver a better substrate to the operations teams and the developer teams. But to de to deliver that experience to the developer teams, that means that you cannot expose them to traditional infrastructure and virtual machines.
What they want is they want APIs that allow them the access to elastic services, just like they're getting in the public cloud today. Elastic compute that gives them self, self-service, elastic storage, elastic networking and security services. So a project lead can grab their team and just go, they're not dealing with ticket based development for on-prem and moving very fast for the use cases where they can use in the public cloud.
But that meant that not only did we have to do our own firmware and operating system and hypervisor, but we had to do the whole cloud control plane. That distributed system that deals with multi-tenancy, that deals with the security and isolation required for, uh, being able to have tens or hundreds or thousands of tenants in a shared pool of resources and be able to treat them all as individual companies. Most multinational enterprises, they've, they will have a hundred different divisions in the company and every one of those divisions has to be treated as its own customer and it has to be able to meet regulatory compliance and the things that, that you can't just spray on to traditional enterprise infrastructure.
It has to be designed in from the start, just like Amazon did, just like Google did, just like Microsoft did. Um, and, and then at the top of the stack that, that, that intersection point, uh, allows you to then snap in the kinds of tools that you're familiar with running in the public cloud, be it a container orchestration substrate that you wanna run, um, be it a, uh, a set of plugins for, you know, uh, the way that you deploy your code, be it Terraform plume or, or Kubernetes. Um, but API native.
And, uh, and that is, uh, the important part of being able to take it from hardware all the way up to software and a bunch of benefits that come therein. Because if we could pull this off, which we were, we were excited to be able to announce having done two years ago, it would now be an experience where as you are doing planning on site, you can snap in one or 10 or 50 of these cloud computers that will operate as a pool of resources and then be able to start to, to segment your infrastructure in availability zones and regions and start to manage your on-prem resources. Like your system administrators are managing your off-premises resources in the public cloud.
Um, with one big exception, maybe two, you now own the depreciation of this stuff. There was kind of an interesting announcement. One big AWS announcement, I don't know, maybe a year ago, they had changed their depreciation schedule from five to six years returning a billion dollars to the bottom line of the company.
And, you know, I think every AWS customer is like, okay, that's nice. Is my bill gonna go down? Um, but one of the big benefits of owning part of your computing is that you control cost, you control depreciation, uh, the rental model is great.
Staying in a hotel is phenomenal, gives you the flexibility of going to any city you need and having services like someone accessing your room to clean it and food showing up. Um, but when you're staying in the same city for nine straight months, uh, you want an apartment, maybe a house, and not spending $27,000 a month for three, 300 square feet, and having a model by which you can get the modernity of cloud computing in your own data center, um, has been really exciting for enterprises that are thinking about how they modernize the last 10 years have been focused on how much we push to the public cloud. And that was the right energy, that was absolutely the right focus.
And now that businesses have realized those advantages, the next era is how do we get those same sorts of modern capabilities on-prem and why we really believe that on-prem is gonna have its cloud moment. Hey Steve, um, Jack Poller with Paradigm Technica, have you seen an impact from the recent change in tax laws here that now allow us to, uh, not take depreciation, but right off entire hardware investments? Uh, we have seen, I think, I think there are opportunities for being able to improve how the financials are managed Yeah.
In companies. I think just speaking to the ones that we have, the, the, the customers that we work with, um, it is still early mm-hmm. In being able to take advantage of some of those opportunities and like looking for any opportunity to be able to get an asset that's gonna last 5, 6, 7 years.
Um, and with the slowing down of Moore's Law, you, you have this condition where you can get much more useful life out of your computing. With flash, you know, you, you have very long useful life of this stuff, and being able to get to, uh, depreciation schedules has meant that you had to be your own private cloud integrator and operator with a bunch of different enterprise partners. So being able to kind of bring the best of both worlds where you can get true cloud computing, but in a domain in which you can get some of those, those financial, uh, benefits has been a boon.
Yep. Real quick, uh, since you're talking about policy, have you seen any increased because, especially outside the United States of, um, increased sovereignty? Yes.
Uh, so just what's your experience been with that so far? Oh, I, I mean, we are, uh, we are trying to move as fast as we can to open up serving additional markets. We started in the us uh, by just nature of needing to focus in the first couple of years, um, which has meant we've had to, we have had to defer a bunch of, uh, conversations with folks across all industry verticals in Europe, eu, the uk, uh, in the Asia Pacific region.
Um, I mean where, where US data centers are power constrained, it is downright crushing in, uh, the EU and, uh, and, and certainly in the Asia Pacific region. So, um, very much so. Um, part of the reason that we raised our recent a hundred million dollars series B was to make sure that we were well capitalized to not only expand manufacturing operations to, to meet the demand, uh, that we have seen surging in the last several months, but, um, so that we can begin laying the footprint to service these additional markets in additional regions.
Because in the limit, we are a global technology provider that is able to service companies in, uh, every major market. And, um, we're excited to be able to do so. Uh, just to take a quick, uh, minute or two before, uh, I hand things off.
Um, like all companies, two years ago we started in one vertical. It's actually a very surprising vertical. It was, it was the federal vertical, uh, which we kind of almost tried to talk them out of.
We said, oh no, we're, we're, we're too early for you and we're not prepared for you. Uh, and they, um, once I was done getting in my own way, they said, Steve, do you mind if we tell you our use case? Uh, and then quickly educated me about how much of in this particular national lab their uses were kind of traditional rack and stack, and how much time their critical engineering resources were being spent on doing things like switch, firmware update, uh, and, and, and VMware license management and not actually delivering critical capabilities like doing security audits for the types of things they needed to do for agencies in the federal government.
Um, we apologized and then very quickly, uh, started working with them to deploy capabilities so they could truly have air gapped, cloud computing, a need that we have only heard get louder and louder and louder across the federal space. Uh, and one that we're very, very excited to be able to support. Um, but, but very soon thereafter, our second vertical is financial services, also a regulated industry.
This is an industry where, uh, of course data privacy is critical. Now, I, I think there is a misnomer of like, oh, cloud computing is not secure. Cloud computing is very secure, but there are data regulations and data requirements where a subset of data, and oftentimes a lot of the data needs to reside under the control of the bank or the financial service provider.
And they have had two sets of capabilities in those conditions where they could put data in the public cloud. They were going much faster, getting better insights, delivering more value to their customers where they had to keep that data on premises. It was, it was a much slower and less innovative environment for them to be able to deliver those critical capabilities.
Um, and we were excited to start working with these financial services companies to give them a closer, getting them closer to having both, to being able to serve data in both domains, but be able to do so with the kind of velocity and the capabilities that cloud computing could afford. From there, life sciences, energy and manufacturing sector, um, we, we've kind of been opening up to a lot more of industry verticals and there have been some things that are unique to each. Uh, I think in the kind of tech AI vertical, it is all about being able to go at speed at scale and deliver these high performance cloud computing services on premises in a better economic footprint, um, and do so in a way that is wholly differentiated from a security perspective, from the traditional rack and stack approach.
Um, if any of you have have, uh, come across the, the, the rand modeling for scoring of AI companies from a security perspective, you'll see in there there's a couple sections that talk about beware of firmware exploits in traditional rack and stack infrastructure. Not only beware like, um, you, you will have a hard time moving up the stack if you cannot go and, and, uh, and, and, and address that kind of concern. So security has been really relevant in some energy efficiency in others, but the unifying force has been being able to have modern API driven elastic services where I no longer have that operational burden of two months of integration, install, deploy.
I can move my utilization of my deployed infrastructure from 25% utilization in north of 50% utilization. I can get an energy efficiency profile that's 12 times better than I have had before, and I can have a trusted product that enterprises depend on that I can lean on as I am being pressured by either public cloud costs or my current enterprise software companies acquisitions leading to exponentially increasing costs OnPrem for virtualization. Steve?
Yep. Um, can I ask a short and a long question? Please.
You can skip the long question. So sovereignty, that seems easy. I get it.
I can hug my rack. It's right here. Um, I control it, I control where it is, I control its gap.
No one has access to the control. Right. Invite you.
Yep. Okay. That's, that seems easy.
The short question is speed. Uh, I, I want this to be clear to everyone, including me. You're talking about speed to value because you're practically rolling in a rack and plugging it in and good to go software, right?
Yep. That's the, that's what you're talking about about when it comes to speed, This crate comes fully assembled. Yeah.
You you wheel it. Yeah. That crate, you wheel it onto the data center floor, you plug in power plugin networking, and under 60 minutes you have development teams that are operating against cloud computing Services.
Right? So that's speed. So then the long question is reliability.
Yep. Um, two backs that has a dimension in the, the system design. Mm-hmm.
And this integration of software and hardware development, I think is, is, is a really strong ox point for oxide, but there are other dimensions to reliability as well. There's operations, um, there's services, there's your, the services or support that you or partners provide. Can you let's, I'm, yeah, I'm I'll I on it briefly and then I'm gonna, uh, co-founder CTO, Brian will, will definitely, um, can pick this up in part of, of how he's describing what we have built.
Um, but you know, a big part of oxide came out of us as operators and the scar tissue that we faced trying to run global data centers with commodity hardware on in our cloud control plane. And so I I love you're hitting on exactly the kinds of issues that we faced for 10 long years trying to run a global public cloud. And that included a fundamental lack of observability into the lower level system.
We could never know what power our systems were drawing. We could never know what the firmware was doing when things were not working well. We did not have any sort of automation in that lower level system software to tell us the health of systems, the usage of the underlying system.
And to your point on reliability, you know, the, the, if you, if you kind of take a holistic design, first and foremost, the rack, the, the rack is the computer, the rack, and that back plane should last you seven to 10 years. And it, and, and the switching infrastructure should last at least five to seven years. And so how do you have a reliable system that allows you to modernize as new capabilities come into the market?
And that is why we went with this sled design that allows you to kind of snap in modular components to take advantage of the new CPU stack from say, a MD without replatforming. Um, and then plumb the whole thing with software from telemetry. So you can understand whether it's power utilization, health of system, you can get an advanced warning on the kinds of things that you need to do to ensure that you are able to keep your software systems running operationally online, um, with, with all of the substrate to be able to do updates without disruption to the software that is running on there for your customers.
Um, so kind of at every layer of the stack, we have designed this thing with operability and reliability, um, baked in and it, but it, it starts with being able to get new insights into what the systems both software and hardware are doing so that you can make decisions against that. Um, and these are, these are things you just can't do with kind of a four vendor rack and stack approach.