Too Much AI, Not Enough Chips! Plus Amazon AI and Smartsheet
Camberley, Keith and Dion review the latest conferences and announcements — Amazon Generative AI, Smartsheet ENGAGE, and AMD — while also discussing what happens when you can’t get hold of AI chips. How do the cloud and on-premises environments factor into this? And what are the options and implications beyond Nvidia?
Transcript
Good morning everyone, and welcome to Infrastructure Matters number 58. And we have some, some of the new faces. We usually gonna see them on a regular basis because there's four of us now here, Diane Hinchcliffe, and he is up in Seattle at the airport.
So hello, early morning to you. Yes, it is, but uh, always fun to be here. Thanks.
And of course, we have the infamous Keith Townsend. Where are you? I'm in Seattle as well.
I'm just still at the hotel. We're gonna be talking about that. So, and I'm stuck here in beautiful Colorado, and Steve is someplace flying on a plane, so he can't join us today, but here we go.
So we are, a big theme today is gonna be on AI because, um, the guys were at this thing called the Amazon Generative AI Summit, and we also had a MD announcing a bunch of things. So we're gonna be talking about a, but the first thing we wanna do is one other conference that was going on this week, um, which was the Smartsheet gauge. And I, Diane, you were there.
So what is Smartsheet and why should people even care about that? Well, A Smartsheet is the, the leading solution in what's called the collaborative work management space. So, uh, unlike project management, they actually help work get done.
And they use the sheet model, that's what they called Smartsheets, like a spreadsheet, but it's collaborative and many people can use it at once. Um, they're, but they're really kinda entering a new chapter, uh, in their, in their journey. Uh, they were public, um, um, for six years, uh, uh, on the, on, uh, wall Street and nasdaq.
Uh, they're now being taken private, um, by, um, uh, Blackstone and, uh, Vista Equity Partners. Uh, and, and it's allowed 'em to really focus on, you know, now, now that they're categories being really commoditized, everyone's entering the space. Atlassian, uh, in ServiceNow on the IT side have just added, uh, the same type of capability to their platform because it's how, how, uh, you know, things like Jira tickets actually get turned into work is, uh, you, you structure it in Smartsheet, everyone has their name assigned to it.
You go out and do it. Uh, but they've, and really kind of are now not moving away from the sheet model, but they're really focusing their credit, a new workspace approach that allows you to, um, to structure all your assets from out a project and putting all the sheets you use for a project, um, in, in one place. That was one of the bigger announcements.
And of course, they're also rolling out generative AI that was, um, very well received. Uh, they're all, they're, if you haven't, if you haven't tracked Smartsheet, they're, they have a cult following. They're extremely popular with their customers.
Um, the announcements were received, I would say very well. Uh, the, the customers don't seem to be concerned about private equity, uh, taking 'em off the market. Um, uh, you know, and they're, uh, I talked to their, their CEO, uh, mark Mader who says that the, their investment partners are aligned with their kind of, their very collaborative open culture.
So we'll see what happens. It's gonna be interesting. I, I don't know, Keith, what, what, what, what did you see about Yeah, when you were there?
So, as you mentioned, the very cult following, I mean, the, they announced the collection and, and there was a round of roaring applause, about 4,000 people was really, uh, shocking to see. But couple of notes, one, these SaaS conferences are usually attended by people who are using the product. Uh, and that's refreshing.
And the nature of this project are people who are working in PMOs, operation teams, uh, disaster recovery teams, et cetera. And they were really one into the announcements, but two very sober about AI and how they use ai. I talked to maybe five or six customers, and each one of them had either a plan for AI or using ai, and this was outside of it, which I thought was an extremely interesting, uh, piece.
And as a reformed project manager, uh, I really appreciate the collections view because the amount of time that I spent at PWC, taking data out of sheets and et cetera, and putting that into PowerPoint just to provide a dash an executive dashboard, I can see why people are excited about it. You know, I kind of, you know, I haven't dived into these guys at all because there's a whole, whole, you know, series of them that are, are producing this. But it's almost like, um, you know, how we had Dropbox and share files, right?
You know, we created a system for sharing files, and now, you know, that those were just word docs or sharing the files, and now we're sharing them in a much broader way, um, and using more intelligence. Um, it's not sure if it's AI or whatever, but more capability of collaborating. And I think this is just another movement forward in this very, very hybrid multi, you know, geographically dispersed kind of world that we're dealing with as, as well as making things simpler, as you said, you know, if I can just drop it into my PowerPoint and that kinda stuff and not have to prepare for that.
Cool. Okay. So let's go into the big topic of today, uh, which is the AI situation.
ai, and we, we get bored about talking about ai. This, I think we got some big stuff to talk about today. So I know this title, like cybersecurity, ai, what are we, so here we are, we, there were, Amazon held a generative AI summit, and I understand they brought analysts in from all over the world for this to attend, um, including YouTube.
Um, so why don't we talk about what are the top things that came out of there? And then we're gonna kind of drop into some, you know, what are the, what are the things that are going on that are really pressing the CIOs and the IT directors and you know, the guys that are trying to implement these things. So, um, Dion, you wanna take that first piece?
Uh, the top? Yeah, yeah, sure. Uh, well love to have Keith to weigh in as well.
Uh, but, uh, Amazon really kind of brought us into to ca calibrate us for what is gonna be coming out, uh, at reinvent at the end of the year. All of us will be there again, of course. Um, they, uh, they didn't have a lot of new announcements, but they really wanted to give us their point of view on how businesses can be successful with ai.
What did it actually take? So they had a former CIO and now a, a, a cloud evangelist, um, to my stage, really walk us through step by step, you know, about, you know, data management, data strategy, being foundational for, for executing on ai, how they look about, uh, look at it. Um, and then the second session that, that was, that was the main session, the second session really focused on basically a lot of NDA topics.
Um, uh, that I, I can't get into it in detail, but I can give you the, the gist of it without, without violating that NDA and, uh, Amazon really is about, uh, uh, model choice. So their all bedrock, uh, capability allows you to use one API to access all the models they offer and all the other public models. They have anthropic and meta and, and cohere and all the other popular AI models.
And you can also put your own model in Bedrock and still use it via a single API. And they also have capabilities coming that allow you to dispatch, uh, AI requests to the, the most cost efficient model. They'll still give you the results you're looking for.
That's a hot topic. Uh, CIOs are very concerned about the cost of ai, not just from the actual bottom line cost, but also from the sustainability, uh, perspective. The AI models consume a lot of energy.
We train and run them, and there's lots of concerns about that. So from an ESG perspective as well as a budget perspective, they, people want cost efficient models, and they, and Amazon's trying to deliver that automatically. We use Bedrock, they'll be able to route the requests to the most efficient model.
And of course, they are now venturing into the hot topic where everyone, um, really focused on, which is AI agents. So not just AI generating content, but AI think needs action for you. So action models, uh, uh, is a hot topic.
Uh, Amazon have any specific announcements, but, uh, they gave indication that we're gonna see things reinvent around that, and they have to, in order to stay competitive. Companies like Salesforce are, uh, are, uh, using, uh, action models, uh, and agents as a, uh, competitive differentiator. And, uh, and Amazon's always been playing defense on ai.
Very interesting. Um, really Microsoft and Google have, have taken the lead in the conversation, and I think, uh, the summit was a, a attempt to kind of get out in front of that, uh, conversation, uh, and, and start, and start owning that, that that topic, at least in the analyst community. That's, I dunno what you, yeah.
Let me comment on your agent thing. Yeah, sure. And know we'll pass over to Keith.
The agent idea is an interesting idea because we, I was on a, uh, conference, um, sat, which is a global file system. It's a, um, where, where people will share data all over the, all over the world and be able to rapidly bring it back, um, heavily focused on the gov, uh, government institutions. They've got really, things are locked down, um, being that's kind of a risky kind of area to be in.
But one of the things they rolled out with on that is they have an AI initiative, so the data intelligence piece of it, but they also rolled out with agents assistant. Interesting. Yeah.
They have assistant personalities there, and they've, you know, they've got a lawyer personality, they've got a document personality, they have a marketing personality, they have a customer service personality. So the, so it's, I guess what they're doing is they're creating ways or training these personalities in certain ways to pull the data through or something like that. Is, that's what's going on with these things, is that what the assistants are?
They, They ares uh, action models, which is, um, uh, what is the largest language model you use for, for agents? They actually work on the APIs for you. So they'll call all the APIs to get the work done.
So you want want, if you want to create a marketing campaign, it'll write the campaign for you, go to all your tools, create the campaign, and then send it out. I mean, it's amazing. So these agents actually take action directly on the IT systems that you make the request for.
Where do trust? Yeah, it's really gonna, it's the security and trust implications are enormous, and I think there's gonna be very difficult, um, um, for them to, um, uh, for organizations to really sort through all the issues. Uh, but the promise is enormous.
The opportunity is great. So yeah, I'm, I'm super excited about them, but I have a lot of concerns, so we'll see how it out. Well, Well, I think it's like anything else that's automated, initially, it presents you with the information, and I, as an expert, go through and look at the expert and say, that works.
That works, that works. That doesn't work through it away and train it. And then we go, go and then initially get to the point that we can let it automated, but maybe it just streamlines the process.
So, Well, I, and that's way Amazon talked about the humans in the loop, uh, will be key for, um, to make yeah, agents successful early on. And of course, eventually you can let get humans outta the loop once you trust it, so. Yeah, exactly right.
Okay. So Keith, we've dominated the conversation in onto you about what happened over there at Amazon. Yeah, so I, I, I spent some time with their, uh, accelerated compute team, their compute team, um, the, uh, storage team as well as, uh, some of the Amazon queue team.
And one of the themes that was clear is that Amazon has given a lot of thought about optionality when it comes to accelerate compute and ai, uh, seeing for the first time in person in CIA and, uh, uh, gaton for, they, they had both chips out for display for us to take pictures of. It wasn't even in the NDA area. I was surprised.
A and, and, and people don't even know how unusual that is. You cannot buy those chips. They, they, and I've never seen 'em before, and even the Amazon people there said they'd never seen them on the top sauce before.
So, so that was, yeah, only in the data center. That was pretty cool. But one of the things that I really pushed back what on them about or, and around not just optionality, is this idea that Amazon is being threatened by the private data center, again, over at, in the eu, the EU Competition Committee.
Oh, yeah. Amazon announced that, you know, uh, or at least reported that data center has put pressures on their business and they're seeing competition. You know, it's, I is is that the stop from, you know, getting, uh, sued by the EU for Monopoly?
Or is it real? I think our data is showing that it is real, that customers are considering private options over the public cloud. So I was asking a lot about lock-in, uh, you know, when we think about these chips, when we think about what a MD, uh, Intel and Nvidia are trying to do at the software level to lock customers into software platforms is tied to the hardware.
Amazon is probably the, you know, poster child for Walt Garden. And these solutions are extremely, extremely powerful and addictive. So when I all in and bedrock, what's my options for basically leaving?
And Amazon, you know, they were pretty, uh, they were pretty for forward in saying they, they wanna make it so that you don't want to leave that, uh, from a cost performance perspective, they expect to compete with any solution and, uh, that they would, uh, and, and they, they would keep pace of change and allow you to modernize in a way that's, uh, hopefully much easier than what we're seeing on the traditional is side and move, uh, to platforms and, uh, processors and accelerators that fit your needs, whether you're talking about insing or, uh, a lot about training on inferential. So that's, okay. So let's get to this other piece of it.
Um, we talked about, I mean, a MD had their, their announcements this week. They had their big conference, um, Patrick, um, and Daniel were both down there. So there's a lot of, you know, there'll be a lot of publication coming out on the very details, but maybe, um, Keith, you were tracking what they were rolling out.
So maybe you can touch base on that, because what then, what we wanna kind of drill into is this issue, um, that Diane you raised, which is, you know, right, we've had, Nvidia has dominated the market, and now we're starting to see other companies bringing out other pieces, et cetera. You know, what are the options for getting hold of Blackwell? You know, what are the options for doing, you know, is there really gonna be competition here?
And how does that competition look in terms of the, the GP tip and ships and such? So maybe, uh, Keith, would you go through, talk a little bit about what you saw from the A MD release and what was going on? Yeah, so I was following mainly Patrick and Daniel on X and trying to keep pace with some of the releases.
Some of the highlights will, I think we can dive deep into it maybe next week once we actually read the release and get some more detail and talk to Daniel and get the debrief. But couple of really interesting pieces, uh, stat that will help Diane and his conversation. 5 million epic CPUs in support of their MI 300 x, uh, GPU accelerator or, uh, AI accelerator that is exclusively used, or, uh, 4 0 5 Lama 4 0 5 B.
Their 405 billion, billion parameter model exclusively runs on that. So that's an interesting stat for Diane to chew on a little bit later. And then the other big announcement coming out was that they have announced a new DPU, and they're all in alter ether ethernet.
This isn't a surprise for those of us that have watched the company. They bought a DPU company a few years ago. Uh, I'm surprised that it's taken this long to come out with an actual DPU, but this is in competition with, there's a consortium, the Alter e ethernet consortium has come together to basically battle Nvidia and this, uh, ability to use ethernet over, uh, the proprietary, eh, somewhat proprietary, so, uh, solution that Nvidia used for high speed networking between GPUs and, uh, I think, uh, I think, I think there's going to be a lot to watch, especially for us ultra data center geeks around ultra e ethernet.
Yeah. And it's, that is the, we we're seeing the battle coming out with that, with a whole bunch of those elements. We can also probably see eventually, you know, the Broadcom kind of bring some of those pieces out as the, the battle continues for the ethernet connections for the AI environment.
Um, I know there's still more, we're talking about whether or not InfiniBand is gonna exist in that, but probably not. It's probably gonna go, you know, 'cause we're doing, you know, the east west kind of d design of these architectures as the north south, and they lend themselves more to the ethernet technology from what I understand. So, um, so one of the things we talked about is, um, if, and you mentioned, um, Dion at the market estimate coming out of a MD for AI chips, um, was $1 trillion by 2030.
Um, you know, that can't all be, hopefully, God, that's not all gonna be interview. Maybe it would be, you know, we have a lot of other people rushing to bring these things to bear into market, but, you know, we've been in a, um, a supply issue because we've got one supplier pretty much. But we're also seeing some other people come out.
You talk about, you know, what are you looking at? How are you advising clients? You know, what if you can't get your hands on the Blackwell?
Mm-Hmm, yep. Yeah. So, uh, for those who are, who aren't tracking, Blackwell is the next generation AI processor from Nvidia, uh, is highly anticipated.
There have been production problems and they've had to redo the mask lately. Uh, those are not, not for pin signs, but it seems like they've gone through the, uh, they've managed to get through that and are now shipping because Microsoft just announced they had the first Blackwell chips in the public cloud, uh, this week. So that, that, that was major news.
Um, but the, the reality is NVIDIA's got 90% of the GPU market, uh, around AI chips. And the, and the reason that is, is because they own the entire stack. They have, um, uh, a, uh, an API called Cuda, which, uh, is a developer, darling.
Everyone likes it. Everyone's built their model around it, and it's very hard. It's the lock in, um, for ai, and it's very hard to move your code over to a entirely different stack.
And so if you wanna go with Intel, uh, chips, a, uh, a MD, uh, you want to go with, uh, Amazon Differe, uh, and Traum, there are two chips, one's for inference, uh, and one's for, for training AI models. They use entirely different APIs. And so, uh, there's a lot of reticence around moving your AI code.
Um, if you're building foundation models, uh, uh, moving away from the proven, um, very popular, uh, kuda, uh, API, uh, for your, your ai, we don't know how well, uh, Intel and a MD and even to some extent Amazon, where they, they probably had the most, I think, the most notable success in the market other than Nvidia in terms of actual uptake. But you can't buy the ships, you can only use it to do, do Amazon services. Uh, so it's very interesting.
So what do you do? And you have to be contrarian, and you have to be willing to bet against the market if you want to use a different, uh, AI chips. And so blackwells are gonna be constrained and whatever's accurate and be constrained for the foreseeable future with these, if these forecasts are correct.
Uh, and we believe they are. So 1 trillion annual spend on AI chips is, uh, where the market is headed. That means that Nvidia is gonna sell everything they have and more, um, and they're gonna be very constrained for, for the rest of the decade.
This if, if this pace continues, uh, and all satisfied that it will. Uh, and so, uh, there are alternatives. Um, and I would, right now, I it's, I I was head do my best.
I would, I tell organizations, uh, don't get yourself locked in the Cuda. Um, other chips are out there. Ceres has the world's largest AI chip.
It's, it's an entire wafer. It's amazing. It's incredible.
And, and it's actually kind of putting Moore's Law back, uh, on, on track, but it uses its own API too. So, um, the key is I, uh, I think if you wanna be an AI long term, they'll get locked into a particular vendor. Uh, start hedging your bets, uh, build a, a, uh, an independence layer, uh, between your model and the API I that you're using.
Uh, and that will put you a good stead because Nvidia is not likely to remain the, the, the, the only option, uh, uh, given that how constrained they're gonna be in terms of popularity. So that's, that's my take on it. I don't know, Keith, that'd be very interesting, uh, to hear what you say.
Uh, the same with you, Angela. Yeah. So this is where middleware has a huge opportunity, this ability to sit in between the AI chip, low level programming language like Cuda, Intel, a MDA little bit more open with their and more, uh, progressive in their licensing around those, uh, languages or APIs, you know, so if you've messed around with AI at all, whether you're talking about within one of the cloud providers or with one of the AI studio platforms like PyTorch, you don't see Kuda, you don't see the, the PyTorch does that for you, that translation for you.
So you can go from one model to the next one, uh, GPU, uh, iteration to another. If you're developing at that AI level, at that PyTorch level, if you're doing those types of projects, those are the tools that's really impacted. Where are you gonna see the difference?
You're gonna see the difference in performance. You can run the same models across d different GPUs and see a difference in performance. A big question is, what are you gonna use your AI for?
Are you training models? If you're training models, you're probably not taking advice from the three of us. If you're, if you're doing ai interesting, which is where the vast majority of the industry is going to go, I wouldn't be fixated on which GPU, which accelerator to use.
You're gonna have plenty of options. Don't wait on Nvidia or the availability, the Nvidia, if you see a market opportunity that's going to accelerate your, you Know, I think we were talking about, you know, how do you get your hands on, you know, fill in the blank GPU kind of situation. And since I work in the data world, one of the problems is is that your data is here and the GPU is here.
And, okay, so, which is one of the reasons why, you know, you, you know, earlier talking about whether or not we're gonna go on prem or in the cloud, it's because data has gravity and we don't particularly like to move it. And then it's get, and, and it can get very costly if you then take it out of the cloud, um, on the Mac side. So, um, what we're seeing, some of the implementations that we're seeing from the data storage people is enabling, um, through the snapshots or through the mirroring, or through whatever it is, is constantly keeping updated.
Maybe a system that sits right next to the Amazon environment or the Microsoft environment, but the data is actually sitting in the Equinix, um, co because you hold onto your data, then I don't lose my data, but I'm able to have that speed of connection back to the, the processors that are over there. So those are some, you know, that is one of the strategic kind of things that people are looking at, is they're looking at privacy and how do you also, how do you access when you, you know, your primary data center is someplace else and where the transaction is going. So I look at other tech, other strategies and technologies to get it over there.
Um, so Beth is, we'll see, see if that one, that that kind of strategy works with any of the guys and, and where they're going. Um, and, and Keith, what I really liked what you said is that it's, it's this issue of Will we, is, is the gp is a single GPU really driving all the decisions about what you're implementing and how you're implementing and the issues on training, the issues on rag. You know, having data ready to feed into the, uh, inference engine, et cetera, is a huge issue.
Um, and probably has it almost a stickier problem than just getting a hold, than getting a hold of the, the GPUs systems. So, um, it's, we have a rack of problems here. Stop there for any comments.
Well, I think it's just interesting, um, you know, uh, organizations, uh, will, will want to get their, their hands on their own GPUs if they wanna run their models locally. And there's, of course, a lot of reason to do that is controlling and protecting your data, uh, and ensure privacy, you know, as in, um, you know, taking and preserving, uh, PII, you know, um, first I, uh, uh, information. We don't necessarily, most organizations, especially large companies, don't want to release this information into their, into the public AI models, not knowing for sure what's gonna happen to them.
There can be accidents and data spills mm-Hmm. Um, they'd rather not. They have that information get out there, uh, into the AI ecosystem where they can control it.
So that's, you know, driving a lot of the enterprise demand is saying, I, we want ai, but for many of our most important AI tasks, and we are not gonna use public models when we can use private models of our own data and running our own inference on our own models. Uh, and that's gonna drive, I think a lot of the, uh, trying to figure out how you're gonna accelerate the AI and, and make it cost effective. Yeah, I mean, and, and then, uh, we're seeing AI models bring shrunken down to run on a phone.
It really does matter what you're trying to accomplish. We're, look, you know, we're fixated as the industry on GPUs and accelerators, uh, that are separate from the CPU, but Intel a MD uh, the cloud providers have really done a really good job of putting accelerators built into the instruction set of the GP or the of the CPU itself. So if, especially if we're looking for a model that's probably gonna be closer to what we're going to implement in enterprise, and most use cases for most agents around 7 billion, uh, parameters, depending on the scale, your projects CPUs might just be fine.
And, you know, when you start thinking about, and, and this is where I'm going with it, is as we have inference engines that are doing realtime ingestion of data, ra, the RA realtime rag information that goes in there that has to, um, generate a response based upon, you know, a, a user hitting the website or you know, asking, asking a query, et cetera, um, as opposed to more of a recommendation engine or something like that line, you know, yeah. The speed matters. But is there a point somewhere along the speed of the GPU is enough by the enterprise?
I mean, is that a possibility? You know, absolutely. This is where tokens per second matter, this is easy math.
The human, the average human, I think reads about 33 words permitted. Um, that is much faster. That is much slower than a CPU can serve up a 7 billion, uh, parameter model.
You can probably get about 15 users per second in an average CPU for that average model. So unless you're talking about massive scale of ai, there is a point where enough GPU matters, you know, we're fixated on Blue Field and a bunch of, of newer GPUs, but, uh, the L four, the LL forties, these lower levels, uh, GPUs from NVIDIA are much less capable than a H 200 or Blackwell chip. They're perfectly fine position to the enterprise.
So gaudy, uh, from Intel, uh, 300 series from a MD ROC has come up with some inference focused, uh, AI chips. We're gonna see plenty of options for the enterprise, especially when it comes to ersing. Right.
Well thank you very much gentlemen for, um, joining me this morning up for infrastructure matters. Um, and thank you all for listening in. We appreciate it.
Make sure to click follow, share, do all those things and we will see you next week. Yeah, everyone.