Storage East AI
Keith Townsend, Camberley Bates and Steven Dickens review the AI Infrastructure Field Day and the vendors that took part. They also discuss the Bank of America outage that impacted more than 60 million customers. Lastly, the team covered VAST Data’s Cosmos announcement and the new OCI region in Malaysia.
Transcript
Hello and welcome to episode 57 of Infrastructure Matters. It's three of what's gonna be ag angle of four. On the call with you today.
We've got Keith Townsend and Kimberly Bates. Hey guys, welcome to the show. Good Morning, Uh, morning, afternoon, evening, wherever you're at.
I think everyone's, It's five o'clock somewhere, right? Keith? Somewhere.
And it's Friday. It's Friday and it's beer time somewhere. We are missing Diane Hinch Cliff.
He's gonna be a new regular on the show, but he called be with us today. So you just just got the three of us. Yeah.
And we just decided for everybody listening in that Diane is bringing in the class. Yeah, we, we, we, we, we are rough around the edges. We, we will admit that, but, you know, we own it.
You know, we're all at a age where we can own that we're a little rough at the edges. And Diane definitely up levels the conversation. Totally, Totally.
So this is the scrappy trio of the Gang of Four, and Diane's gonna bring the class going forward. Okay. That's good.
I like it. So busy docket This week Kimberly will go to you first take field day. You've been out there with AI Infrastructure Field Day.
I got to sit in on a few of the sessions. Tell us a little bit about who was there, what you saw, what was interesting, and, and sort of get us aware. Well, we had a, uh, it was, this is, um, AI data infrastructure, tech field day.
Um, usually we've had isn't That just fancy words for storage And for, for those who not really Now. Okay. Kimberly really does have a really mean scowl going on.
Oh, yeah. And I just was on the end of it. Yeah, there we go.
Okay. Data, data infrastructure. So, and as if we've talked about, it's like the, the world has moved.
We, we we're still in the drives and that kind of stuff, but the real meaningful stuff is stuff. There we go. Technical term, as most people understand, is in this software capability, it's also into the, I mean, it's, I mean, we had pure storage talking very, very much so in detail about their direct flash modules, which is they've, you know, got lots of smarts in these drives and that kind of, and the, that area.
And there's other firms that have invested heavily in the tech as we know, like Nvidia kind of important. The hardware is, but the software is very important too. So as we look at the infrastructure and it's also gets beyond storage, it gets, it's beyond that idea as I'm just sticking it into the drive kind of thing.
Um, I'm really working on this entire infrastructure and how the key thing on AI and data infrastructure, it's more, it's as much about the speed and performance of that technology and the software that's on there as it is, is about the integration and the understanding of what we're doing with this. I mean, for instance, we had soy who is a solid state drive vendor, who did a phenomenal job there talking about the entire data pipeline breaking down every piece of it. I mean, their charts were just amazing.
Um, in terms of a learning process. I mean, the guys that we was a lot of fun. We had two kinds of folks there.
We had traditional infrastructure people like Keith and I, and then we had the data engineers there. So we had this really interesting conversation that was going on. And let me tell you, the data engineers absolutely loved what soy a drive vendor was presenting.
If you can imagine, well, I mean, the Google session that I got to sit in, you know, the Jux position of the HBE session, the, the soy session, the Google session, we were kind of going all the way up and down the stack. The Google piece was kind of coming at it from the AI kind of down, um, solid. I sound like they were coming at it from the drives up.
I mean, really holistic. They Came from here down, they started from, what are the things that are going on? What does this look like?
Because training imprints, rag archive, every one of those have got different, different requirements and, and different size amount of data that we're doing. And it's very, very different. So you gotta break up that entire thing to say, how do I build this?
Because not one thing fits. Yeah. And this a whole lot of kind of kinds of technologies here.
And I think the, this field day highlights where infrastructure is going. Like you had this super abstracted layer that Google's talking about in this infrastructure, this, that was a infrastructure field day. And then you're going all the way down.
I was looking at the, uh, some of the soine, uh, not just the, the front channel, the, the actual presentation, but some of the back channel conversation and the data infrastructure folks, the data engineers were amazed at like how much soy has given thought to these higher level abstractions. And I think that's representative of where infrastructure's going, where, you know, this term infrastructure has become really go, uh, you might talk something as low as, you know, drive speed and fees, and that's something that's high as what services, what abstracted level services is infrastructure teams managing. And into Google's thing, they presented their very broad range of options for storage, if you wanna use that term.
So everything from a parallel file system, they're involved with DAOs, which came out of Intel in terms of an open system or open source, uh, parallel file system, you know, just regular file system, you know, block where all this fits. And so they, they've gotten this broad range of saying, and, and one of the things that we were looking for is like, okay, so, you know, when you go in an engineer or even a, you know, a data infrastructure dude has gotta be able to parse what pieces do I use where, um, you know, HPE was a little bit more subscript, uh, um, prescriptive in terms of what they're doing is bringing a, a stack to the world, which is a huge value. If someone just wants to roll something in here, here, here you got, we've curated this open source, um, work.
We've curated the technology under that. We've, we've, we've made sure it's, you know, it's, it's in, it's working together. Um, we we're allowing you to pick in pieces, whatever you need to in, from the open source space.
So, so they've really, you know, they've done a great job of that kind of offering that's out there. And of course it's a GreenLake offering if you wanna do that. Um, so it was, and then we had Pure, which was fascinating because it was very technical.
They got to into some really technical items. And what we were trying to do is understand what the things that they have built into their direct flash modules and their networking behind it and how that gives them an advantage in terms of speed, in terms of being able to feed the beast, the, the GPU and in that area. So that's kind of where they went in, in talking about it, but it's just some really broad range.
And I'm, I know I'm missing somebody here. Who did I miss In There? We go in Finn Fitting.
Let me maybe come to you, Keith, you were there. What did you, what was your I was not there, so I'm I'm gonna defer to and I get to have it There. Oh, I thought you were there.
No, I did. I was going to sneak by and go there. But we're in travel season and this is my opportunity to take a week off from travel.
Oh, Travel. Oh, cool. Go do that.
Yeah. Yeah. Bill NEIS was there and he was presenting on where they fit both into talking about where they were working in terms of providing, so their, their big focus has been on, you know, the, the traditional transactional high-end systems.
And one of the things that we're thinking about is saying, okay, so most of what we think about here is file, file an object, unstructured data, but we have to marry somewhere along the line. We have to bring in and marry the transactional data that's gonna go through this, uh, vast for who we're gonna talk about a little bit later is he, they talk about time series coming into that data to, and think about it this way as if I have a, I've got an infra, I've got an engine that I'm doing and I'm making some decisions. It's already been trained and everything else, but I've got new data coming in, or I wanna use new transactional data that's going on, whatever that is.
Maybe it's the banking system or maybe it's retail or whatever, but I need to feed that in, in real time, more than likely to augment my knowledge base that I've created. And that's that rag. I'm augmenting my knowledge base with current absolute new data.
How do we do that? And, and those are challenges. Those are really big challenges because we got data in different places.
Um, and how does that feed into this process? So it's, there's, you know, we're just at the, just the cusp of this industry. And I think as delegates, we walked around, walked away from there going, wow, there's a lot to be done.
Mm-Hmm. There's a huge amount of collaboration between the engineer data engineers and the infrastructure people to communicate, because we speak very different languages. Yeah.
Kimberly th this reminds me, I don't know how we missed this because we did talk about it last week. I think it was in between shows the ML PERF results came out. Oh, yeah.
0 how the industry is starting to look at storage performance and its impact on ai. 0 came out last week, and I'm going say upfront, I have not spent the time on it. Although our guys over in Signal six five have buried themselves into it and looked at it.
Um, and I have not, I've been on the road, so I have not had a chance to kind of get into what the, the pieces that are first out there. 0 that they, they brought out. Um, and I'm not exactly sure what this is actually reporting, because there's a lot of claims out there.
I'm number one, I'm number one, I'm number one. So there's a lot of slices And still is Origin. So, so it's kind of like, okay, I'm gonna take a look at it, but I think I be, I'm sure what they're doing is they're looking at everything from your training to the rag, to the whatever.
And how do you, as I said, this is not one type of workload. This is, you know, five or six workloads that are in that, and then there's more that we don't know that's gonna be needed in order to work through that. And now we haven't even started talking about classification and PII stuff.
So that's the, that's an entirely different, but that's kind of what ML PERF is trying to do from a storage standpoint. Yeah. I'm looking forward to, I haven't dived into it.
The, I I still need to read the Signal 65 report, but this is something that we've been trying to get our storage customers to pay attention to and help us and collaborate with us to see the impact, especially on rag, because I think most people are not doing training, obviously, but almost everyone is going to do some level of infra and rag, and what is the impact of storage performance in subsystems and services on rag, which is probably a good segue to, you know, what Vast is doing. Yep. Yeah, let's go there.
Keeping us rolling along Vast made some big announcements. I wasn't involved in those, but I know you guys were, what did they announce? Was it significant?
Tell me more. Yep. Yeah, the storage is still the Kimberly's you, if you ask me if you ask a virtualization question, I'm all over it.
People identify me as a, okay. Virtualized Lights and Keith. Yes.
Uh, yes. So they, what they, what they were doing, first of all, um, conference, conference called Cosmos. But there's also, what Cosmos is, is what they created with Cosmos is a community, um, of, uh, technologists, technology companies to come together and understanding that this is not a one system kind of environment, but it is a, it's a will take, take a village, if you will, on bringing this to market.
So that's what Cosmos is all about. And, um, and the, the usual suspects were very involved with that. You know, everything from the vendors soft, uh, hardware, software and, um, partners like UWT.
Um, but the big piece of this is coming out with the data engine that they're doing, um, which is what they've already, they've already done is they've got this, um, you know, they've got the data storage, the base level data storage. They've been working on, um, the vector data vector database capability across the entire data of storage, which is, is used for, um, rag. And now what they've added in there is the ability to bring in nims, um, which is capability coming from, um, Nvidia to be able to iterate on, you know, the, the rag systems, et cetera, and push, push that out there.
There's not a huge amount of detail on it because they're not shipping until they say they're ga. Um, and, uh, what they say they, they, they're shipping in first half and they've also talked about first quarter. So, you know, somewhere along in there that we're gonna see that.
And at that time is when we're gonna get a whole lot more of the details of what they're doing. Um, but essentially their vision there is to be able to, and, and the other piece they're doing is that's bringing in time series data, which is significant as I was just talking about. How do you bring time series data, um, real time, um, into the, the rag system to update and be able to produce out.
So that entire platform, their vision is that here is one platform you can use that will streamline the entire process of what you have. So the question is, you know, you've got everything there. You've got the ability to do block, you've got the ability to do file, you've got the ability to do object, bring all that stuff into your, um, your vector database or the environment that you need to do, or your data warehouse environment, process all that, and spit out the answer on the backside.
So that's, that's where they're at. And all within, you know, privacy governance that you need to have well integrated into Nvidia, who's their key partner here. Um, the question then, the big questions that we've had already is, will customers and their data is already on all these other systems, are they gonna bring in a brand new system and do that?
And the answer coming back from many people we talked to already is maybe, and the reason being is because this is a new application or a new environment. And so they are setting up new areas. You know, for instance, we hear about Dell talking about this AI factory.
What they're talking about is that, you know, what we've done in transactions is only is this much what you're gonna do on AI analysis and data is gonna be, you know, 10 times bigger, 20 times bigger. So yes, 10 To 20 times more storage required. I don't Know.
I wonder where that, I wonder. No, no, no, no. Processors.
Processors. Oh, okay. Primarily processors.
I was gonna Say that sounds like a campaign to sell more disk and splash. And if we look, if we look back at kind of, you know, what we've done traditionally in enterprise, it, there's opportunity here for vendors. Because when I was at, uh, my last assignment as a, uh, enterprise architect, I had what was called Vmax back then, which is now, you know, whatever the, uh, power Power max, which is, you know, the, the mainframe as we like to call it, the mainframe of storage.
We had three of them supporting one app, SAP. So SAP was important enough to support, uh, uh, three, uh, power maxes, two for production, one for dr. And the rest of the infrastructure ran on Dell's lower level storage.
So if the application brings enough business value and there is, uh, exactly up time and performance, and there's a return on that, investment enterprises will spend the money on it. What's very uncertain it came across in, in, in the Cosmo, uh, partnership, is that we really don't know what these applications are gonna look like in the future. We're trying to hedge our bets and, you know, there's, you know, uh, partners in this Cosmo from super micro to Equinix to WT Deloitte, uh, what's missing is like the big cloud providers, right?
Uh, and uh, there's this argument, especially for ai, that this is going to be primarily a on-premises workload. That the QT xs, the Equinix of the worlds will be hosting these workloads and we'll have to do a lot of Lego block putting together initially. And it's important for these types of partnerships to basically show enterprises the way, because this is, this is, this is PhD level stuff right now, especially, uh, you mentioned the time series database, uh, the time series data stuff.
I've just worked, uh, read a, a huge research paper on whether or not you should try and rapidly retrain versus use rack. We just don't know. So fantastic discussion.
Lot of storage. Kimberly, you've been talking a lot. Oh, should I say data infrastructure?
Sorry, I shouldn't say storage. Um, couple of big things for me this week that I saw system of record outage for Bank of America on Wednesday hit the press. Very, very few details.
I went digging, spent a couple of hours going down the rabbit hole, um, really hard to find what this was. Spent some time with a couple of the IBM guys last night for, and you know, there's a massive mainframe infrastructure that is the system of record at Bank of America. And I kind of said, did you guys break things on Wednesday?
And there was no kind of reaction from those guys. I think that was a cyber attack and that's not why the, and why there was no, um, no sort of details that I could find. Typically, if a bank's having, uh, cyber issues, they won't say anything.
If they're having system uptime and availability issues 'cause they broke something, they're a bit more transparent. So I don't wanna fuel any conspiracies, but I spend a fair bunch of time looking at this. I kind of track outages and track downtime and this type of stuff.
It's kind of my bailiwick. Really Not very many details out there. I don't know, Keith, whether you saw this one come across you already all this week.
I did see it come across my radar and it is exceptionally quiet for a bank outage. It's, yeah, we won the world's largest banks. Mm-Hmm.
You typically get some pretty transparent things. People get nervous when bank, when the banking system has problems. And someone, uh, that's frankly too big to fail the size of this bank.
We do have some pretty transparent announcements. You know, there's, you know, some system failure or whatever the, some high level explanation. I expect that one way or another we will find out what happened because this isn't, you know, some mid-size bank in the, I'm in the Midwest, so I can pick on us on much Midwestern.
It's not a mid-size bank in the Midwest that only a handful of folks use. You know, it's not i'll, I'll use a defunct name. It's not Harris Bank of of Chicago.
This is, you know, a proper global player. Something like 70 million customers. I mean, there was literally no, I must have read enough, 15 different articles then went digging from there, didn't find any detail.
So there was one, and the only article I saw was a report that some people had their balances zeroed out. Was there any kinda, Yeah, so I mean, people going in by their mobile app looking at their balances, their credit card balance was intact. Interestingly enough, it's always a trade run.
That's none system. Nobody ever catches a break in these outages. I mean, all of this is FDIC insured, you know, your money's not gone.
Don't create a run on the bank. Don't go to the teller and get all your cash out 'cause you think your money's disappeared. That's not what happened.
This will just be some type of sys they'll recover, as we said, they'll be, you know, they'll be, um, protected copies of the data. They'll be getting, your money's still there, but it was just really interesting that some of those balances were zeroed out when people were looking at their accounts, which is obviously stressful, right? Yeah.
That, that would, that would mean, that would mean that me $800. Yeah, I mean my, my 50 bucks wouldn't be there and I, you know, I'd be really stressed about it. I mean, Kimberly May be, has got a bit more money than you and I, Keith, but I'm, I'm wanting to go up and look at my account right now because I happen to be with Merrill Lynch in Bank of America.
So there we go. Oh, okay. So that, so that wasn't broke.
I, I'm gonna continue to dig, but it doesn't to me sound like a system of record outage. It sounds to me like there was a cyber attack. I'm putting two and two together there.
I could be making four, I could be making 48. But that's what I'm saying, based on what I was able to find out. So I'll be really interested to hear if it was and how they caught it and what did it and what, yeah, what systems did, did, did do this.
I'm sure we may may not be able to report on it someday, but, um, it could be interesting. We, I'll, I'll keep digging. Uh, the other big news for was, um, an OCI region and ament in Malaysia.
So Oracle cloud infrastructure, I've been tracking, I've got a research note that's coming out on this that's in the system that should publish either today or, or early next week. Another big region from Oracle. I think I see a really good trajectory from those guys around sovereign cloud.
I think they get that data sovereignty is really important in various markets. Obviously with the database background that they've got. Those are typically Infrastructure matters or is it storage that matters?
Well, I think based on this podcast, it's storage matters, right? We should just rename the podcast. Um, but no, I mean, I think given their background around database where they are, they get that those system of records are gonna be very often in Oracle and they're gonna need to be OnPrem.
So lots of details here. 8 billion, which is significant. Normally these are in the bi the billion dollar range.
So this is a biggie. This is like the Amazon one in the UK that I reported recently. So we're seeing some very big investments that these hyperscalers are having to make into software cloud deployments, uh, I think kind of continues on a trend I've been tracking all this year.
And Oracle is a unique, with the cloud providers, they're, you know, kind of differentiators that when they bring on a region, they, uh, they make the marketing claim that every service is available. So it's not, oh, this is a region with, you know, these core set of services. Uh, if you're using Oracle Cloud VMware service, it should be available day one in this new region in, in, in Malaysia.
So it is that $7 billion number is a big deal, and it's a big deal that they're bringing up a new region and a new part of the world for them and, uh, the services that comes along with it. The other thing for me, and I've think Oracle's able to make their regions and their cloud instances with what they're doing with Alloy and cloud customer, they can go a lot smaller with some of these regions. 8 billion, you know, that that's, that's up there.
That's up there as big as any of the ones I've seen. But they're also go, I think I took away from cloud world that it's down to a smallest three racks. You know, so if you've got, and you want to connect to your, the cloud customer, the alloy, the Exadata cloud customer, if you want to connect and create that sort of small region, you can connect that infrastructure back in what they did and announced at cloud world.
We put in OCI into AWS to complete the set and now have AWS Googling Azure. There's some really good stuff that Oracle's doing here with, its, its sort of, I don't want to call it private cloud, but sort of data sovereignty and how that connects into other clouds. So I think, Yeah, I think the, the, the new term is exactly that sovereign cloud, this ability to take these three racks, put 'em in an Equinix data center next to your vast data storage, blah, blah, blah, and the whole, uh, the whole, uh, uh, market architecture picture and have this data locality capability, uh, with these cloud services interacting with your own data within your own region.
Yeah. And that's vital for enterprises that operate locally, government agencies, defense organizations. There's a whole raft of people that are really paranoid about where their data sits and the regulatory and legal framework that sits around it.
8 billion. Again, big amount of money at the top end of where I've seen these. So, fascinating announcement.
I've got more details coming out in the research note. So we have three of us this week. We're getting up on to half an hour.
Anything else from you, Bo before we wrap? No, this has been a, uh, I've learned a lot because now I don't have to watch, go back and watch all the A IDI videos. Actually, I'm going to feed it into, no, I'm, I'm gonna do an experiment.
I'm gonna feed it into Notebook, LM and see if Notebook LM could just replace the three of us. We'll replace ion, but we placed the three of us on giving news. We're gonna create a avatars and just completely get the AI to do it.
Keith, is that what the planet Exactly. So, so our compatriot, David Nicholson is feeding all his sarcasm into something. He's calling Dave bought 9,000.
Oh geez. And he's gonna have it start writing his, Well do, I, I, I don't know if anyone can really, any AI can quite capture Dave's snark yet, so We'll, we'll need to feed an awful lot of data. Snark is a service.
Yes. Is that what we're at? RK is a surface.
There we go. Yeah. It's success.
Okay. Well, fantastic. You've been listening to episode 57 of Infrastructure Matters.
We are here every week to bring you the breaking news, what's been going on in the week of infrastructure and how that matters and why you should care. Please click and subscribe and do all those things, share with all your friends so that we can grow the audience here, and we'll see you next week. Thank you very much for watching.