92. AI Needs Resource Efficiency – Tech Field Day Podcast
As we build out AI infrastructure and applications we need resource efficiency, continuously buying more horsepower cannot go on forever. This episode of the Tech Field Day podcast features Pete Welcher, Gina Rosenthal, Andy Banta, and Alastair Cooke hoping for a more efficient AI future. Large language models are trained using massive farms of GPUs and massive amounts of Internet data, so we expect to use large farms of GPUs and unstructured data to run those LLMs. Those large farms have led to scarcity of GPUs, and now RAM price increases that are impeding businesses building their own large AI infrastructure. Task-specific AIs, that use more efficient, task-specific models should be the future of Agentic AI and AI embedded in applications. More efficient and targeted AI may be the only way to get business value from the investment, especially in resource constrained edge environments. Does every AI problem need a twenty billion parameter model? More mature use of LLMs and AI will focus on reducing the cost of delivering inference to applications, your staff, and your customers.
Transcript
More horsepower, more power, more get those horses moving faster. Your ai, it's huge. It needs more power.
Or does it, can you actually do great things for your business with less horsepower and do things more efficiently? Join me on the tech Fields, a podcast as we drive this stage, coach through the ideas of efficiency in AI infrastructure. Welcome to the Tech Field Day podcast, where we bring together a group of IT technical experts to discuss a single idea about key concepts in the industry.
This podcast features a variety of perspectives from members of the Tech Field Day delegate, and it's often recorded in association with one of our events, in this case, AI Infrastructure Field, day four Tech Field day's, part of the Futureum Group. And this podcast is also published, uh, on our sister company site Techstrong tv. On this episode, we'll be discussing that AI needs resource efficiency and that more horsepower isn't the only direction.
But before the discussion, let's meet who's on the panel today. Hi, I'm Pete Welcher. Uh, I've been around networking for a long time now.
Retired, still doing it because I think it's fun. Hey, I'm Gina Rosenthal. I'm a product marketing manager.
I've been doing that for a while, been doing product marketing for AI since about 2017. And I'm Andy Banta. I've been hanging around the infrastructure industry for quite a while, um, mostly from the technical aspect of it.
And I'm Alistair Cook. I'm the event lead for AI infrastructure Field Day here at Tick Field Day. And one of the things that we see is that AI needs to kind of grow up from this land grab and this idea that massive amounts of infrastructure, ridiculously powerful GPUs stacked up in, uh, rack after rack with hopefully a very strong flaw and liquid cooling to get to ridiculous numbers of, uh, kilowatts, hundreds of kilowatts of power being delivered through a single rack.
Um, I've long felt that there needs to be a much more efficient way of delivering ai. I was hoping there'd be some revolution where this more horsepower would cease to be the issue because all of this horsepower and all of this compute power, that energy has a cost. It's really expensive to generate these things.
It's really expensive to build these data centers. And as we've covered in the tech fields and news rundown, there are contracts for more data centers to be built to host AI than have ever existed in the past. Exist now.
So this is a huge build out for ai, uh, using vast amounts of resources. And even if we set aside the difficulty of actually producing those, the cost to people producing them, is that even a sensible thing for us to be doing with the planet? Should we be generating AI by heating the entire planet up with these massive GPUs?
And it's kind of a a point of, of difficulty along here. Um, Andy, you've been around this, this block a few times before you've seen massive infrastructure get built out and maybe get more efficient. Can you see that happening here with ai?
Well, I, I think it has to. Uh, one of, I, I've actually been, um, talking about the power consumption used for data centers and AI data centers in general for the past couple years. But one of the things that I would be really interested in hearing from, from the AI infrastructure companies out there is how they're doing things more efficiently, not how they're doing things faster or more densely or, uh, or various different ways of adding more and more.
But it's, you know, tell us something about some algorithms that you're using to make more efficient use of the hardware that you have available or the speeds that you have available, or talk to us about a way, uh, a technique that you come up with to do more with less and, and instead of just telling us how you can do more, that's really what I'm interested in hearing more of from infrastructure companies. I completely agree, and I've got a slightly, uh, different perspective on the whole market. I think the companies that develop the very large models needed, um, vast funding because of the huge costs that go into those LLMs, and they couldn't afford to specialize, uh, because they really needed something that would have a vast potential market.
Unfortunately, what I think has happened is that they've assumed that's the future. I'm not sure could, it's certainly got a role particularly for conversational, um, interactions, I guess to put it. But if you are trying to model something, be it network troubleshooting or, um, some biological process or something, maybe a smaller dedicated model that knows about, say, computer networks, um, might be more suited to the task and a whole lot more efficient to run.
So that's kind of a question in the back of my head. I don't have an answer, I'm just kind of watching the space, uh, with intense interest. I agree with both of y'all.
And Andy, I agree with the whole idea around power. When you think about, if you read any type of description of what, uh, the data centers are being built for, it's described in gigawatts. So, um, it's not a back to the future reference, but it's actually how much power they need to run, um, these servers and to run the different components of the servers.
They never talk about the data. They never talk about how much data that will serve us. They never talk about how much data will be created.
They never talk about, you know, the actual part of ai, it's the end result of it. They only talk about what it takes to run it. But when you do an infrastructure, when you do any kind of architecture, you're always looking at what is the workload that's going to be run on it.
And it's not at top speed because not everybody can afford to run everything at gigawatt speed all the time. Not only that, but if you think about what is ai, this is common, I'm gonna always come back back to this. What is ai?
AI is what we used to call right now. What what is available from AI is what we used to call high performance computing. So machine learning and deep learning, and you don't need gigawatts of power to run those types of workloads.
So why aren't we talking about this as a true architecture instance, right? Like, what needs to happen, how much data you, you need, what do you need for each section of the pipeline that we're gonna be doing to transform this data? Um, that, those are kind of different questions, but I think those are important.
So we kind of skip over the whole, we can design for this, it's a workload, this is what we know how to do, and if it's a workload, we know how to make it run more efficiently. So I I'm with you. I I wanna hear more about that.
Well, you touched on something that I've been focusing on, which is, uh, where's the data? And I think we're gonna see that during the ai, uh, infrastructure field day, uh, presentations, because if the data's not local or if you've got, uh, such a huge LLM that you need to run your, um, training in multiple sites, then just accessing large amount of data is yet another barrier as well as latency and bandwidth and power. Right?
And I mean, the, the other piece of this that I, I really wanna try to address on some of our discussion here is that the, um, in addition to the data, the ability to actually create the data isn't being done efficiently. Uh, and Fabrica presented it one of these sessions a long time ago before they got consumed, and part of their presentation was talking about the fact that the LAMA model is only, uh, theoretically capable of using about 70% of the GPUs available to it, and in practice is only using about 30% of the GPUs available to it. And that is clearly showing that more is done.
What you need in here, what you need is to actually do the good old fashioned software engineering to make these models more efficient, to make use of resources that are available to them. And that is really part of the efficiency that I would love to hear people talk about. Uh, Gina made this point a little bit earlier, uh, when we were in the pre discussion for this about why aren't people looking at solutions similar to VMware of, uh, of sharing resources among various different processes.
Uh, and you know, I I think Phoenix can offer a little bit more information on what she was thinking, but, uh, that's like an excellent question that we don't hear from Ethan VMware these days. I think VMware has a hard time getting the message out, but, um, they have a lot, they've had a lot of tools for a long time where you could do things like virtualize the GPU and split it into, I wanna say 16, but it might have been eight to deliver that, that that vir that, uh, power to different virtual machines. So the same thing we did with Oracle back in the day, remember nobody thought you could virtualize Oracle.
That, that it would be, you would cost so much because you needed those servers to be bare. You need to, uh, uh, put that Oracle OS right onto the server, and that's how you had to run it. But once you started looking at, no, these are workloads, you can virtualize all of it, we can virtualize every component of it from a VMware perspective, and it's virtualization in general.
I'm sure it's all of the virtual, um, virtualization hypervisors that can do that for you. But, but it, we're not thinking about it that way. We're, I think it's, it's, it's definitely the hype that's driving that.
Like you have to have ai, you get ai, you have to have this many servers and this many GPUs that that's the formula, the blocks, and that's kind of how they're being sold. I don't see anybody pushing back and saying, nah, man, we need, what we're needing to do is we're needing to take this historical data, do this with it, it's gonna cost this much for us to crunch through it, and this is what we actually need to install in our data center. But people were doing this before the days of ai, they were doing it when you went to super compute, you would see people doing cool things on bare battle.
And I know this because we would, there would be more saying, and you could also do this with virtualization. So it, it, it's just workloads. And I think that the architects and the engineers need to get back to, to that piece, which would also force people back into the what is the use case?
Why are we going to do ai? What's on this side? What's on the other side, and what do we expect to report back?
You know, what, what's this gonna gain the company if we, Yeah. And it's, um, and VMware certainly did do virtualization of GPUs, and they, they did it both, they did it both in software, and they also were able to split out the hardware functions of various GPUs and spread those to various different VMs. And through some of their later technology, they were actually able to share those across the network.
Um, I, one of the other concepts that has been brought up that I really haven't seen a whole lot of traction on is there are CXL vendors out there who are attempting to say, harvest the memory from your, um, systems that you were shutting down and put it into a cxl rack that you can actually use on a new system. The idea is that DRAM doesn't go bad. DRAM can last for a very long time, and even if it's not the newest, most, most fast DRAM that's out there, you can certainly make use it again.
So I mean, the, in addition to efficiency, we also don't hear too many people talking about reuse or sharing of resources. We're gonna have to, right? Because what happens now when all of a sudden the prices of everything is going through the roof, you can't get rid of your memory.
You can't, you really have to hold onto it because if you don't hold onto it, you, you're not gonna buy any new, unless you three times what you're already paid, what you paid for it yesterday. So we're in this memory crunch because they can't make it fast enough. The same way you can't get ahold of GPUs because they can't make them fast enough, can't manufacture them.
We're in a place where the chips are getting smaller, but getting, doubling in, in capacity size. So they haven't really, uh, perfected that, um, that, that in the, in, in the manufacturing plants to get those rolling and get those out to customers. So many of them are already pre-sold And they're pre-sold, so you can't even get them once you, once you get ready to go deploy this.
So you're gonna have to get creative and figure out how do you architect for this workload with what I got or what I can maybe find on eBay for a decent price. Right. Well, Jane, you hit on another, um, thing that I think is key, which is shorthand would be ROI, but what's the company benefit from what people are doing?
And that's kind of the challenge with everybody does ai, uh, which seems to be the, a meme in some co companies, and it really is much more helpful to start thinking about ROI. But another factor in it is that maybe not everybody, just like not everybody is cut out to be a programmer. People have their job they wanna do, and they may have a idea that AI can help them, but right now, uh, how to prompt A LLM and, uh, get good results and sort of cycle that is, uh, something that's not for everybody.
And we're starting to see some mechanisms coming, uh, gate and some of the other things that help with automating that sometimes even in massive ways, which could be a problem. But, uh, when is this, when is AI sort of experimentation gonna be easier for, uh, Joe Schmo and marketing or whatever to use, uh, as opposed to somebody who's a computer science, uh, specialist And, and then Pete, is that the right thing to have happen? Because as we're talking about this, as, as you doing AI through resource efficiency, what are the tools that that abstract away the intelligence of the operator, right?
That, that take away the specialization of being able to work with the AI tool? Well, that puts more load on the ai. That means that we need more horsepower in that AI to handle the fact that we're asking somebody who, who is not specifically skilled or oriented towards understanding the tool to just throw simple language stuff at this tool.
This is how we're getting this idea of needing more and more horsepower. I think coming back to one of Gina's points was that having an AI tool that is specific to a task rather than a general purpose ai. So building into your application instead of using, uh, a full 7 billion or 20 billion parameter model, that is a general purpose model.
Having a model that's specifically set up to handle maybe, uh, analyzing the productivity of your, um, factory floor. Having an AI that is specifically trained for that function is gonna be more efficient than feeding the all of the information about your production floor into a general purpose model. And that's where we need to look at that resource efficiency versus simply building more horsepower.
There isn't a single solution here, there isn't a single answer to this. Sometimes in order to get value out of ai, we're gonna need to a very specialized AI that does a very specialized task. And I come back to some of the early discussions around this, that a lot of the time it was still the, uh, predictive ai, the old fashioned machine learning that, uh, was looking for trends and data and uh, and extrapolating those trends and data a lot of the time, that's what the production AI was.
And strapping a chat bot, large language model on top of a huge amount of resources wasn't necessarily the right way to solve this. Another of the things that we saw across our AI field days, uh, as, as though were maturing, was the idea that chat bots aren't the answer. The, the AI needs to be invisible inside the application doing whatever it's doing.
Uh, it, it needs to not be the, the primary interface. It needs to just be the thing that makes the thing better to quote Bobby. So AI should disappear and, and should be task specific.
Well, that's where maybe the LLM gets used in a slightly different, uh, form of agentic, uh, use, where it recognizes what field a question, what type of question is being asked, or what type of inferencing is, is being requested. And then it calls on a dedicated model that, uh, for instance, like what Cisco's done with an in-house network trained, uh, model to actually answer the question. So it calls an AI agent rather than a software code agent.
That's a classic program. Yeah. Yeah.
And I mean that's, I think that's part of where Alistair was going. I I think part of it, uh, the idea of the efficiency could be that you, you don't need to train something on the entire world of knowledge to, for a specific thing. And I mean, if, if you go in with the idea that we only need to train something for a specific piece of knowledge, then you, you probably can do it more efficiently.
And it, it's, uh, you know, you talk about Cisco, a a friend of mine who worked at a MD said that one on one of their early models of their, uh, virtual assistant that was based on AI model, that he, the first thing he did was ask him, what is Epic? Which is the name of, uh, desktop chip? And the virtual assistant couldn't answer it for them.
Fascinating. Uh, one of the things that, you know, you brought up, I, I think it's unfair to say that Joe Schmo in marketing, this is like LLMs are the perfect tool for marketing, right? And I know people aren't giving the right prompts yet, but, um, we had a conversation about this this morning on the Textron game.
What happens when, um, what happened with email, when email first came out, everybody couldn't use it. I mean, I had to support post drops who were astrophysicists that really had a hard time with email. So, you know, because it was so new, it was something brand new that nobody knew how to do.
Um, and they didn't, they lost, they didn't understand the rules. But now you can't imagine having an organization that doesn't have a proper email system running one way or the other. And I think that's how it's gonna be with ai.
AI is gonna be how we communicate with computers in the future, right? This is how, and now going forward, and it's gonna get more and more, we'll, we'll figure out the latency problem, we'll figure out the power problem, but people dunno how to use it now because nobody knows how to use it now. Like the few people that do are, I would say the very, very technical people don't understand how important those marketing people are in the organization and how they can pick the information out and distribute it back.
Again, I'm defending marketing because I'm product marketers, so have to do it. But, but I, I think, you know, the, the very technical people forget that that data that we're moving around and trying to transform things into, one of the way it gets transformed is by the, the knowledge and the point of view that the people that aren't technical that are gonna take a little longer to train them on the prompts and, and how to do things properly, they understand the way to pull the information out because they have the vocabulary, everything else that nobody else wants to learn about marketing. If you're not a marketer, your technologist, you don't care.
You don't wanna learn that. But the marketers will know the right words to pull that out and pull the information out of the data. And that's really what a lot of those LLM, um, workflows are all about, is to help pull information outta data.
So I think some of this is just, we are in such early stages and we were bombarded absolutely bombarded with hype about what this is and how you don't get on board, your company is gonna be sunk and all the rest of it, you know? So I think it's really important to remember that all of us are suffering from hype overload and we're gonna get there because this is how we communicate with computers now. Yeah, but this goes back to the there point, or yes, I, I agree with exactly what you're saying, but I, I think, uh, some of the, some of what's being lost here is that, um, the efficiency that you actually would get, lemme back up that this is, this is making things more inefficient rather than, uh, more efficient.
And I think one of the things that we, we lose sight of is actually paying attention to the genuine computer science necessary. And instead, we, we keep adding layers that are, uh, wasting cycles to get the same job done. Thousand percent agree.
Have you ever looked at stuff and it looks like front page like AI stuff looks like front page from Microsoft way back in the day. Yes, I agree. I agree With you.
So will, will it optimize when it looks like, um, MySpace and see the front page? I think that'll be a little bit optimized. But that's like when they're just, that's probably like when they're just to the point where it's too much and then they're like, holy crap, we gotta pull this back and make this really work for people and we've gotta make it good.
Now, One other challenge that I'm, I think is lurking here with the idea that everybody does AI is, uh, it takes me back probably several decades to, uh, ro it, uh, the prospect of people just learning spreadsheets and doing corporate financials on a spreadsheet and screwing up a formula with consequences. And that's actually good news from the AI perspective is we've been here before, so people learning and making mistakes, uh, hopefully you can fence them and be aware that that's a potential problem. But, um, I think assisting people with frameworks that, uh, minimize the mistakes is probably a useful thing to think about In terms of resource efficiency.
Um, I, I really don't expect that an unskilled user should be handed here. Here's an AI that is a general purpose. Ai, you can do anything you want with it, and we're not gonna tell you any limitations around that, right?
That's, I don't think that's a good way of using ai. AI should be invisibly embedded into the thing you are already doing. So as a, as a marketing person general is perpetually writing documents, writing, um, presentations.
And if you're going outside of that, using AI to generate some things and then come back into it, the workflow's wrong as, let's say there's just another layer being added that adds no value. The AI should, should actually be embedded, however frustrating. It's to have the, uh, Google sheet say, would you like me to generate everything for you?
Right? That's where it should be. Uh, the AI should be a feature that makes the application you are using makes the tool you are using function better.
It shouldn't be a separate place that you have to go. Um, some of the evidence that we're misusing the separate place is, is absolutely ask the ai what is epic when you can go to Google and say, what is epic? And be taken to the documentation that describes exactly what Epic is.
People are using, um, chat GPT as a replacement for search and finding that chat. GPT isn't necessarily up to date on the latest news. It certainly doesn't know, uh, about the ice storm that's coming to the eastern US uh, this week, or about the landslips that happened close to me here in New Zealand over the last couple of days due to the weather.
Um, but I can find that information through Google search, right? Using AI and particularly chat interfaces for the wrong things is gonna lead to poor results. And I think, think hiding them away.
Don't you think that if you, I'll go back to my front page, you know, that's why front page was so awful because it had everything in one, one place. So don't you think that we're gonna end up having more bloat problem and more resource problem if, um, if we put everything in the place? Well, data sovereignty is the other thing I'm sitting here thinking, because if you have, uh, people just asking chat GPT, uh, well, it's already happening that, uh, very private and sensitive information is getting exposed because it's sloshing out into the ai.
And given some recent results where people have been able to reproduce entire books or major chunks of books by feeding, by iteratively prompting the LLM, uh, they could probably get at that secret data, uh, confidential data as well, which is not a good thing. Returning to, to one of the things that Gina said that, so the sort of bloat to building AI into everything, I think you've got a really vital point how rebuilding, putting AI into every single application independently as we're seeing at the moment isn't efficient. Is it?
It's the same thing again, again, again, uh, but we're seeing some movement towards some standards about how you inject ai, uh, or specific knowledge in AI into an application technologies like MCP, um, standardizing the way you attach to an ai. So I don't think there's necessarily the need for that, that complete rebuild. Uh, and there's other standards like the agent to agent standards because as Pete was saying, a lot of the stuff is, is actually gonna be done by agents rather than directly by human interaction.
So I think there are developments in that space of not having to rebuild, redeploy, uh, do again, exactly the same thing we've done in the past. So I think there is hope for us that the AI delivered into a specific applications doesn't have to be a completely from scratch or really an efficient reimplementation of some other ai. Uh, of course that's probably what we're gonna see first because that's how you get to a minimum viable product.
Uh, I think Pete mentioned that, that idea of, uh, we built huge large language models and, and so that's the way we view all of AI needs to be this huge amount of resource for these huge large language models. Again, I I view that a little bit as a minimum viable product. I really hope as maturity turns comes, we will see more efficiency, more reuse of technology, and more use of tools that are specialized to a single function and that do that a function more efficiently.
The thing is, we've gotta stop the pace of going for more and larger, which was coming back to our original premise, more and larger isn't necessarily the right way to go. Um, so yeah, I, I think I've rambled all over the place on, on a bunch of stuff we've talked about. Uh, and I, I, I really do hope we do see more efficient use more tasks, specific use of large language model AI and, um, some of that maturity of working out how we're getting business value rather than just the, um, arms race to acquire as much resources we can.
Because if we don't, somebody else will acquire that resource and, and we'll lose, I think in the end, if all you do is acquire resource and never use it, you're gonna lose anyway. You may as well, uh, look at efficiently using those resources that you acquire. Well, there's competition for the resources.
There's also time to market. So if you're using massive resort costly, uh, resources and training a model, you have to build this Mongo data center and everything just to get rolling, that's a long lead time. And so anything you can do to make it lighter weight, it affects, uh, improves your time to market, Right?
I, and you know, I, I think one of the other places where we could have potentially see some efficiency is, uh, being able to filter out AI generated content when we're actually building our models. Uh, it's one of the things that that's, we, I mean, we, we end up with, uh, hallucination amplification when we do that. And it's, um, it also means that we're, we, we end up with AI that's seeded with imaginary information to begin with.
And that's, that was one of the terrifying things about deep sake was, was when I built deep sake, one of the things they did was use a lot of AI generated content to teach the ai. And yeah, that does seem like fantasy land being used to build another fantasy land. And that doesn't translate well to being useful in the real world generally.
I think one of the other struggles is right, you know, of course in the marketer I use, I use, um, Microsoft. And, um, one of the things that's gotten way better is their copilot, especially if your documents are on the platform, that's gotten way better. But you talk about, um, uh, the bloat is there and the bloat is there because the systems underneath, uh, office doc, you know, online office are not all the same.
They've come from mismatches of what they have and put up on, you know, the different, different sites and we use them, but you can see it when you try to start to run this whole pipeline, you know, writing and producing and editing and adding things and that kind of thing, because you can't do it all in one place. So I think that that's a problem too, is like your businesses are going into AI and they might be going out and buying whatever they can to be the most powerful and the fastest and blah, blah, blah, but they're reaching data that they're using data, they're organizational data that's on all sorts of stuff, maybe even on tape. There's a lot, that whole piece they're saying.
That's what I've heard is that whole piece of preparing the data to actually be used by a model and to be used to make some business battle that take value. That takes a lot of time and a lot of doing. And you think about rehydrating things from tape and getting things into an object ready, if that's the way you built things, how, how that looks for everything to be consumable by the model.
So it's, it's, we we're gonna have a lot more bloat before we can get to the trim down thing, even if we're trying to, so like, we've got this problem of not being able to get the components, this whole problem with getting the data ready, you know, and then making it, um, palatable for the business users to use. So it's a, there's a lot to be done. Did you just say garbage in, garbage out?
Yeah, that's what I said. I think that that idea of bringing data together from different sources leads to a, a sort of disconnect where, as you say, it's not all the same format or naming of things. And this is the semantic layer that Brad Shiman from Futureum Research has been writing about recently.
I'll be presenting, uh, a little bit of his research at, uh, infrastructure Field Day tomorrow. And, uh, I think there's gonna be a continuing conversation at I Infrastructure Field Day about how you build this infrastructure, how you make it valuable, and we're probably gonna come in with a lens of how you make it efficient. Andy Ton join Gina, uh, here in, in, uh, Santa Clara with us, and I have a, a collection more awesome delegates.
Uh, so before we fill out hours and hours of this conversation, I think I'm gonna put a line under it here. And thank you all for joining us today at the Tech Field Day podcast. So before we close out this podcast, where can people connect with each of you and maybe continue this conversation or maybe decide they wanna buy you a beer at your local, uh, bar?
I can be found on LinkedIn. Um, my initials PJ Welcher should find me. You can find me on LinkedIn too, and I am so excited that my podcast partner will be with me at Tech at AI Day, and you can find us at tech aunties com.
Okay. And you can find me on LinkedIn as well, Andy Banter, uh, as well as on Blue Sky Social, and I, uh, my, my occasional, um, blog posting on andy banter substack com. And of course, you could find me, Alistair Cook on LinkedIn and all of your favorite social media, as well as at AI Infrastructure Field Day this week.
Looking forward to having some great, uh, conversations with people there. So thank you so much for listening to this episode of the Tech Field Day podcast. And of course, if you enjoyed this discussion, subscribe on whatever your favorite podcast application is or on YouTube so you don't miss a single episode.
Uh, give us a nice review as well, a high rating because you love Listen podcast when we have such entertaining delegates joining us that will help other people to find the awesome content. This podcast is brought to you by Tick Field Day, the home of it experts from across the enterprise and a part of Therum Group for upcoming events and more episodes here to tick field day com slash podcast or viewers on Tick Strong tv. Thanks for listening, and of course, we'll as always, see you next week on the podcast.