Evolving Storage Management with Qumulo’s Douglas Gourlay
Newly appointed Qumulo CEO Douglas Gourlay dives into how storage management is evolving in an era when more compute than ever needs to be brought to the data.
Transcript
This is Textron tv. Hey guys, D Wither. We're here with Doug Gole, who is newly appointed CEO and president of Qumulo.
And we're gonna be talking about storage cloud native and all kinds of cloud stuff like that. Douglas, welcome the show. Thanks for having me.
It's good to be here. So what attracted you to this company? There's clearly a lot of activity in the cloud native space, there's a lot more stateful applications than ever, but from your perspective, what's going on?
Well, I think the really, the main interesting thing is when you look at the mainstream enterprise and you look at primary storage for mainstream enterprise, there's been this incredible challenge that customers have had over the years, which has been, do I keep things on premises? Do I move them to the cloud? We've heard customers talk about or consider repatriation, we put it in the cloud, we found that it was really expensive, now let's bring it back.
And what we've built is a system that allows an enterprise's primary storage to run anywhere. First off, it can on premise X 86 or arm architecture in a disaggregated model on whatever vendors storage platform they want. Then they can run the exact same file system in the public cloud, Azure or Amazon natively at a cost point.
From an economics perspective, that is basically equivalent. So this financial challenge, the economic challenge of making a determination cloud, non-cloud, what's best for my business, have eliminated the technology as a barrier to that coupling that with thirdly, a thing we call like the global namespace, which, uh, to use a movie analogy allows you to basically look at your data as everything is everywhere, all at once, globally replicated and making it so that if I'm building a movie and I have a follow the sun model of editors and artists and special effects technicians working on it, everybody can be working in their workflow harmoniously as the data is still very rapidly accessible, globally synchronized and extremely durable and protected. The, the combination opens up tremendous opportunities for new solutions for enterprises and how to protect their data, how to utilize their data, how to have their data feed into the applied AI systems and so on.
We've been spending arguably the last 10 years trying to move data into the cloud. And more recently it feels like there's a more conversation about, uh, bringing compute to where the data actually is. And it seems like we've come full circle on this conversation that started 20 years ago when I was much younger.
It was always about bringing the compute to the data. And then somewhere along the line we shifted gears from your perspective, what is happening here these days? Are we trying to find some middle ground?
You know, they always, as you said, we come full circle or does the pendulum swing back, right? I mean, insource, outsource, go cloud, go OnPrem. I think those cycles always happen.
And the job of technology is not to lock somebody into a particular business decision. I think the job of the technology is to unlock the business decision so that the business can say what is best for us. There are scenarios on, to run a big on-prem data center, you gotta have people in human capital.
You gotta have expertise, you gotta have reasonable scale. If you don't have human capital expertise and scale, don't build a data center, don't run one, go to the cloud if you need something that is elastic. If I need to forward project things and stand it up rapidly in Singapore right now and then Switzerland in two hours from now, the cloud is a perfect, perfect place to run it.
So we've seen enterprises say, if I don't have the people, if I don't have the expertise, if I don't have the scale, I'm gonna go on the cloud. We've seen some enterprises lift workloads and shift them over to the cloud and go, whoa, this is a lot more expensive in reality than I thought it'd be. Maybe we should come back.
And we've seen people refactored applications for the cloud and then realized they locked themselves in. Now our view is our customers should have choice and there shouldn't be a technology inhibitor to moving your workloads into the cloud if that's the best business decision. And to be able to do that without a negative economic impact.
And similarly, if a customer has workloads in the cloud and they determine though that would be better served running on premises, they should not have a barrier to exit either. They should better bring those exact same applications, the exact same data and the exact same file systems back on premises and not have to refactor an application or again, have to change the way the data's accessed. So in fact, being able to use both concurrently is what it becomes incredibly powerful.
I can have my data exist in the cloud in multiple clouds and on-prem concurrently, I can then choose to migrate an application if I need to. Are we kinda, it's not really a binary choice, I feel, 'cause sometimes I might create an application that sits in the cloud and it's perfectly comfortable there, but then the dynamics change and a year later I may wanna move it to on-premise because the nature of the thing has changed. So do people kind of get that idea that applications are a little more fluid than we sometimes think?
I think it depends by market, uh, and by the age of the application, to be honest. Um, I'll give you a good example. One of the, one of the first earliest, best examples I I've had of this and this this was a decade ago, was a video game company.
And they said, you know what we're gonna do? We don't believe our product managers forecast, so we're going to launch all new video games in the cloud. We're gonna let the cloud take the elasticity.
And if we need to rapidly scale up, we're gonna rapidly scale up because the last thing they want is a poor user experience. You know, that is like you're bringing call of duty out. If all of a sudden it's slow and it's lagging and it's going down, players are complaining, Reddit blows up, you know, and disrupt blows up and next thing you know, or discord blows up.
And next thing you know, everybody's like this game's horrible, don't play it. So they used the cloud to absorb that elastic capacity until they had predictable demand. And once they understood and they had very predictable demand curves and utilization, they were then able to say, let's make a business decision.
Should we repatriate this to on-premises data centers? Or should we leave it in the cloud? At that time, it was generally more cost effective for that business to read because of some of the storage costs and compute costs and so on.
It was cheaper for them to repatriate it to an on-prem facility. But then when they would relaunch that game in a new market, we're gonna go launch it in Korea, now we're gonna launch it in Australia. They didn't have data center facilities there.
So they'd relaunch it in the public cloud in those facility, in those areas. And then again, make a business determination as to whether to move it to a hosting site or it was, was it worth building a data center for, um, in every one of those scenarios, you know, your your point is exactly spot on. It's not a binary decision.
It's enabling the business to make the best decision possible. And when you're locked in because you know, I can't move my storage or I can't move my applications, or my applications are only designed to work with a certain type of storage system or using a certain protocol, if I am designed to be ultra cloud native and locked myself into one cloud provider, or you know, you lose that portability, you lose that choice in the business and that that's not a good thing. Right?
Normalizing the economics and removing the technology barriers to migration make it incredibly simple for customers to now have the choice to use the data center, the edge, the cloud, a hosting facility, whatever's best for them. Are the applications themselves becoming more hybrid? 'cause you know, back in the day we used to say nothing good happens when you move data.
So, or we build applications that are able to invoke data sets regardless of where they're, Some are. Uh, definitely if you look at some of the more modernized applications they used, um, more HTDP and web style principles for how to call data, allowing for things like redirections allowing for it to be, you know, uh, be far more tolerant of different latencies. And you know, I remember back in the day it was like if a scuzzy right had to happen at a certain time, had to synchronize over to another data center, the acknowledgement had to come back and then the right acknowledgement had to happen back to the database.
And all of this had to happen in like under a millisecond or your database would time out and a file lock would have an issue. And, you know, increasingly, uh, newer applications or using protocols that are more tolerant of a longer latency in that allowing the data to exist in across town or two states away and not have to be co-resident completely. At the same time, there are applications that are, have incredible performance demands.
You see this in AI workloads. You see this in media and entertainment. You see this in, um, medical imaging, right?
Where distance does matter and what you may want are large centralized repositories of the data. And then distributed elements that when I walk into a hospital and I log in and I register at the hospital, imagine it starts prefetching, my health record, prefetching my imaging, my, my most recent imaging system. So the oncologist, whoever I'm working with instantly has that data on hand.
And I'm not sitting there waiting for an hour for a date for the data to come down of these, you know, very large, you know, medical imaging files, Right? Most patients walking through the door think that that's what's actually happening anyway when it's really not. Um, you mentioned AI is AI kind of creating something of a renaissance period for data and storage management that we kind of ignored for many years and, and now suddenly, uh, everybody's kind of waking up going, Hey, uh, data management is kind of crucial.
You know, uh, two things there. Um, we try to break AI into, you know, two aspects. One are the training and learning model aspects where, you know, you're feeding this model lots of data and you're seeing, you know, g PT 3G, PT three, 5G, PT four in incredible evolutions and in advancements in what's happening with generative AI and these lar extremely large clusters that are being built.
Um, I mean there's probably 75 to a hundred organizations in the world that have, again, the human capital, the scale, the need, the ability to build these very large clusters for these types of learning models and applications. Um, but when you look broadly across the mainstream enterprise, you start finding organizations that say, how do I use GPT-4 with our data without our data becoming part of GPT-4, right? I want to write an email that says, congratulate my top three salespeople for the last quarter and include their numbers and put that in as a prompt and have it give me accurate data back that I can cut and paste into an email.
And had said just that would be an incredibly impersonal way to generate a thank you email. But let's just assert I actually did that. I don't want my competitor coming in going, tell me a story about Kumu Low's top three sales reps and how they performed last quarter and having my data now part of open AI's model and you know, somebody being able to get access to that data through that type of environment.
So there's some new models and you know, I broadly think of it as applied ai, which is, you know, there's things like, um, this model called rag, which uses your data, runs it through a vector database and interconnects it with an o open AI style model, but doesn't allow my data to become part of the model. I think mainstream enterprise will be using far more inference, cloud-based delivery and you know, these types of applications of using public models using open models with their data. And I think what it will result in, and kind of you, the, the headline is there's probably gonna be more data generated in the next three to five years than in the entire previous history of humankind at the rate of data generation.
We're seeing from videos and photos and sound and text data all coming through using and synthesizing what we have and what we've written. And then using AI systems to accelerate the creation of more data, Right? Data management and storage is roughly equivalent to the Levi's and shovels of the gold rush, right?
Who made all the money. I I think that that's a, that's a pretty fair, fair statement. Um, you know, and our view is that I wanna work with the companies who are gonna maximize the value, you know, out of that, which are the, you know, the thousands of mainstream enterprises who have decades of data that is file data.
That is, you know, how the formula for Coca-Cola, right? The, you know, what choices and trials did they do for every drug they ever tried to make at a particular pharmaceutical? And then be able to utilize that data with AI models to say, here's what worked before.
Here's what we've tried, what haven't we tried, what new models should we do? And allow them to again, utilize public models, utilize their data, whether it's in the cloud, whether it's OnPrem to maximize the value of that data. You mentioned security in passing, but I can't help but wonder if, uh, storage data management and data security are about to converge.
I think there's a tremendous point of convergence there. And if you look at one of the biggest risks of having, you know, large scale data is we see large scale ransomware attacks now, right? And it's one of the, uh, you know, one of, again, one of the bigger attack vectors, uh, and more common threats right now.
And we're seeing healthcare systems get shut down and go back to paper charting, right? And that had a tangible impact on billing. It had a tangible impact on patient care.
And you know, I mean, I literally talked to some of the doctors and nurses at, you know, a healthcare system that for eight weeks plus was operating on paper charts. And the question becomes, how can the storage system, the data management system, the applications, is there a way they can work together to detect when someone's trying to, for instance, encrypt the entirety of the volume of electronic healthcare records? And can that be detected?
Can that be prevented? Can that be rapidly recovered and restored from can you capture enough data about the attack that's happening to regarded as prosecutorial evidence? These all become, you know, again, parts of how the systems have to work together to create, you know, and I, I don't think there's any, you know, every security company in the world will tell you we are the answer.
You know, use our box. We are the answer to this. I, to me, I believe it's a systems level problem.
There's things we will need to continue to do in enterprise data management, uh, that will work with DLP companies, that will work with other security companies, that will work with SIM systems, that will work with other systems vendors and other E-D-R-N-D-R-X-D-R plays. And I think it's that ability to work within the ecosystem at a systems level that creates the maximum opportunity to defend and defeat a malefactor who's trying to do something, you know, that horrible to, especially to, you know, a hospital system that provides patient care. So who's in charge of data and storage these days?
'cause back in the day we had storage admins and then we went through this, uh, hyperconverged cycle, and now it seems like we're kind of decoupling things again and making them scale up and independently of each other. But, um, who do you see playing the lead role in managing data and storage these days? You know, in the larger enterprises, there's absolutely storage, engineering, storage operations, data management, you know, people who look after the entirety of the data plant of the enterprise, because it is, you know, it was a quote, data's the new oil, right?
It's, it, it needs curation, right? It needs someone who cares about it, who looks after it, whether it's in the cloud, whether it's OnPrem, whether it's in a branch, an edge site, whatever. It's the aggregate of this.
And these storage admins, storage operators, storage engineering, um, probably have gone through some title shifts in some organizations, especially as they, you know, went into the cloud or then used a hybrid approach. But in the end, the responsibility has stayed the same. Now, what I have seen is that you go up one level, the infrastructure management, right?
The, uh, VP of infrastructure at a financial firm or whatever is looking at data network, platform compute, oftentimes cloud infrastructure and asking the question, how do I make all of these systems work better together? Because traditionally, and they, they would often use the term, my storage tower versus my network tower. And I'm like, why do you call them towers?
Well, they're like, well, we're in New York, so we call them towers, we would call them silos if we were in Chicago. And I, I got a good laugh out of, you know, a CIO telling me that, but they think of them as these vertically integrated technology components that aren't interconnected well, but one of the key priorities I hear from every one of them is, can you make them integrate and make them work better together? And, you know, my background has largely been networking and systems, and right now I'm sitting in a data management company that stores exabytes of data for our customers and manages exabytes of data across a thousand plus enterprise customers.
And I look at that and I'm like, wait a second, what's the networking guy doing here? Well, it's simple when we have a global namespace and we start distributing data everywhere, that's a networking and storage problem. When it's in the cloud and it's OnPrem, that's a networking, storage systems and cloud problem.
How we interconnect these things and how we transport the data between them and how we allow that data to be synchronized and be readily available and be durable, right? That that is a systems level problem. And I think there's a tremendous area of innovation that we can unlock there by allowing these systems to work better together than they ever have before by default, as opposed to through, you know, archaic configuration.
All right, folks, you heard in here? Well, if data's the new oil, I guess the issue is we don't exactly know how to optimally refine it just yet, but we're working on it. Douglas, thanks for being on the chat.
Thank you for, thank you for the time today. Appreciate it. All Right.
And back to you guys.