70. Early Adoption of Generative AI Helps Control Costs with Signal65 – Tech Field Day Podcast
If you haven’t already, start working with Generative AI now and make sure to control your ongoing costs. This episode of the Tech Field Day podcast features Russ Fellows, Mitch Lewis, and Brian Martin, all from Signal65, and is hosted by Alastair Cooke. Generative AI is delivering value to businesses of all sizes, but significant evolution in models and technologies remains before maturity is achieved. Experimentation is essential to understand the value of new technologies, starting with cloud resources or small-scale on-premises servers. Business value is derived from the inference stage, where AI tools generate actionable information for users. Generative AI is like a knowledgeable and well-intentioned intern; someone more senior must ensure AI is given good instructions and check their work. In production, grounding and guard rails are vital to keep your AI an asset, not a liability.
Transcript
Generative AI is everywhere. Everyone's talking about it. Many people are doing it.
You should probably get started and do something with generative ai, but make sure that you're getting value for what you're doing, that you are doing sensible things. Join me on this episode of the Tech Field Day podcast as we talk with the awesome experts from Signal six five. Welcome to the Tech Field Day podcast, where we bring together a group of IT technical experts to discuss a single idea about a key concept in the industry.
This podcast features a variety of perspectives from members of the Tech Field Day delegate community, and is often recorded in association with one of our events. Tech Fail Day is part of the TUM Group, and this podcast is also published on our sister company Site Techstrong tv. And this is a very special episode for me because it has some three guests from the Signal six five part of, uh, Futurum as well.
And we're gonna be discussing the idea that generative AI is evolving. You should start now and control your cost. But before we start the discussion, let's meet who's on the panel.
Hello, I'm Russ Fellows from Signal six five, and I'm one of the team members that, uh, concentrates on looking at developing proof points and like many in the, uh, industry today, a lot of focus on ai. Hi everyone. I'm Mitch Lewis.
Uh, I, I'm also part of the Signal 65 team, uh, working with Russ and Brian as well. Um, also currently very focused on, uh, AI that we're gonna be talking today. And, uh, excited to be here.
Hi, I'm Brian Martin, signal 65, uh, heading up AI and data center performance, uh, working with Russ and Mitch, uh, and in our scaling lab. And of course, I'm Alistair Cook. I'm an event lead at Tick Field Dice, specifically AI infrastructure.
Field Day is one of my events and one of the things we have seen across the last few episodes of AI Infrastructure Field Day is the idea that you need quite a lot of infrastructure in order to run large generative AI applications, to build models to particularly fine tune models and then run them for inference in production production. And I think we also have seen that there's a lot of uncertainty about how you go about building an application and get started. And so we wanted to look in this, this particular episode, thinking about generative ai, how it's evolving, how newer techniques, newer ways of using generative AI are really bringing better and stronger answers, possibly faster answers, but also that the business value for AI comes when you're doing the inference phase, when these AI tools are generating information that you can hand on and, and make actionable from your users.
And so managing that in that intersection between the cost and benefit is gonna be vital as we're starting to build out applications on top of ai. And Brian, I know you've recently been building out a data center for, uh, running AI projects. Have you got some thoughts about design decisions in that process of building the data center that might be valuable here?
That's a great question, Alistair. Yeah. Down, uh, Colorado Springs been working on our AI scaling lab and, uh, are fortunate to be working with partners, uh, and have two eight node clusters there, 64 GPUs each, which gives us the ability to test from small scale single GPU up through much larger clusters.
Um, looking specifically at generative right now and inferencing as that workload appears to be growing rapidly, uh, among the user base and customers, uh, we're seeing that as, um, we're seeing that we can start small and grow. Um, some of the newer models, especially mixture of experts, uh, while they take a fair amount of BAM, um, can be very fast for their size because they're activating only a fraction of those parameters in runtime. Uh, I was recently able to run the new moonshot can be K two, uh, inferencing that required 16 GPUs.
So this was 16 H two hundreds for inferencing. That's a trillion parameter model with 32 billion active and seeing token rates in the tens of thousands on a trillion parameter model. Now this is scale, uh, toward the upper side, uh, for enterprises, but that just shows what's available in on-prem usage, uh, for these types of models.
At the complete opposite end of the spectrum, I've got, uh, an eight gig Jetson Oren, um, doing inferencing on two, uh, billion parameter models, uh, for an IOT project. Uh, so inferencing is happening at all those scale points. Um, get in, try something and, and see what fits your use case.
And last time I was in Colorado, uh, sitting in Russ's office, uh, Russ, you, you had a workstation with A GPU and it's sitting on your desk that you're doing some experimentation with, and I think you've, your experimentation's gotten a little more sophisticated than it was early last year when I was with you. Uh, how do people get started? What do you see as being a good on-ramp to using generative ai?
Right. Yeah, I think getting started is the important thing. And yeah, I used to be proud that I had one of the better setups, but then, you know, Bri, Brian, uh, uh, shown us all with his, uh, a 6,000 pro, but I, I have a system with a, uh, NVIDIA 40 90, which is, you know, today, at the time it was the, the newest, highest end.
That was before the 50 90 came out. But, uh, you don't need to have the, the highest end or the latest, you know, GPU to get started. Almost anything can get you started.
In fact, you can even get started on, um, CPU based systems. You don't even need A GPU, of course. You're just gonna have to be a little bit more patient.
Um, things are gonna run a bit more slowly, uh, so don't be looking for high throughput rates, but in terms of experimentation, you can do a lot even without a GPU. And you know, even a low end one, if you don't have the funds to purchase one, you can rent those in the cloud at less than a dollar an hour as well. So there's a lot of ways that you can get started pretty inexpensively.
I've even been working recently on a, on a project here at Tech Field Day with, uh, click and, uh, was amazed at the how rapidly I could build a rag solution retrieval augmented generation solution just using their cloud. Uh, currently I'm not paying for that service because the project is on. I'd be interested to see how it compares with, uh, other platforms for running the same kind of, um, kind of tools and development.
I know Mitch, you were working on projects and we, we covered this at Cloud Field day 23, uh, a project around cost effectiveness versus necessarily straight up raw performance. And as we get into production inference, um, cost effectiveness comes to be kind of vital. Yeah, exactly.
We've done, you know, a lot of, so we've done a lot of AI testing, but what goes hand in hand with that often is, um, you know, how much are you paying for it? Um, and I think there's, you know, lots of levels to it. Like we're seeing, um, things that are maybe not super expensive where you can just, uh, get started.
Like Russ is saying, like, we've done some testing on, uh, you know, AI inferencing on CPUs and, uh, as the performance, as good as, you know, Brian's big GP clusters. No, uh, but you can do things. Um, so that's, you know, one way to get started and that's gonna drive the price down.
Um, but you know, we've also done testing looking at different hardware solutions, um, like I talked about, uh, cloud Field Day. Um, so, you know, uh, is Nvidia your only, uh, GP vendor? No, there's other solutions, right?
So, uh, you know, we've, we've looked at, um, A MD and we've looked at Intel, um, and there's different price points. So I think people need to be kind of mindful of, um, you know, what you're doing, um, and what resources you can use, uh, to kind of bring those costs down. Whether you just need to do something quicker than cloud, uh, can you just, you know, leverage an API endpoint because you're just playing around with something and you can throw 10 bucks into open AI and, you know, build something cool real quick just to play around.
Uh, or do you need to go, you know, procure, uh, a thousand GPUs because you are, uh, building, building something massive? So I think there's, you know, people, people get kind of scared off by AI because they think it's just, um, you know, Tesla and open AI and Google with thousands of GPUs, uh, training these models. But, uh, there's, there's a wide range going from, I can run something on CPU to, uh, you know, I have a whole data center with a, a nuclear power plant.
Yeah, I, I'd add onto that quickly, just that, um, we, we have, uh, some upcoming research coming out, hasn't been published yet. We'll be shortly, um, evaluating, you know, three different CPU vendors. So there are three, well, there's more than three, but, uh, believe it or not, people think of kind of two, but there's more than that.
Um, but looking at doing different inferencing and some AI ML workloads as well. And, um, they've all improved significantly over the past several years. Um, you know, there was no focus on the CPU side dedicated to, you know, giving circuitry to optimizing a lot of these matrix multiplication tasks that are behind a lot of AI models.
But with more focus comes, obviously more funding in that area, more research in that area. So the, all the vendors, including CPU vendors are significantly improving, you know, their performance in that area. So just to add on, yes, there's choices among GPU and there's choices among CPU and they're all rapidly improving.
I think there's, you know, and and beyond just, you know, the, the CPU or GPU, there's some other creative approaches that are, um, kind of emerging from, from different vendors, uh, that we've done some testing projects on, like, how can you reduce the storage or how, um, uh, can you actually leverage the storage to offload some of the memory? Uh, so we just did a project, uh, testing on a solution for Fon, um, that takes that kind of approach. Um, and pretty cool.
We were able to run, uh, both a 70 a LAMA 70 B model and a clean 72, uh, B model on a single, uh, workstation, GPU, um, that, you know, normally would not be able to do that. Um, so it did take a while, but there's kind of that time versus cost and performance, uh, trade off. One of the other areas I've noticed that, uh, we talked about generative AI evolving and being price conscious is looking at the code support environment, code development, whether it's Cloud code or Cursor or Windsurf, uh, or Google Fire Base Studio.
Um, they've all recently been starting to change their pricing models as well to the, to the dismay of developers and I'm sure to the, uh, delight of their accountants. Um, but I had a, a recent challenge over the weekend, I was tracking down two pesky bugs and jumped into Claude Code and flipped it into Opus four mode. And I'm like, okay, here, here goes the most expensive approach I can to solve these problems.
And, you know, within an hour and a half or so, at a cost of $19 and 88 cents, I'd knocked down two pretty good size bugs. So, you know, it's both, uh, sobering, uh, to see the effectiveness of it. Uh, and also I like the transparency of getting a sense so we can watch what it's costing now, see how that evolves over time, but really helps put a price on some of these features.
Yeah, watching the meter run is fun. Um, I, you know, it's funny, I've been using AI coding more and more over the past six months, and I use it differently than a lot of people suggest. I, I use it differently than like Brian just outlined.
And I am, I'm aware of that method, and I think it's probably more productive, but I'm still a bit hesitant. So I, I still go with the chat GPT model and I change the model depending on what I'm doing. So I always start with an O model, what I'm doing, the initial outline and, and you know, project, you know, what things do I wanna focus on?
What libraries do I want to use? How do I wanna design this? I, I use one of the thinking reasoning models, it makes a huge difference.
But then I, I'll, I'll take the code and I'll copy and paste and, and so I just do the old copy and paste. Um, it's not as effective. It is more cost effective, though I probably only burned about two to $3, I'm guessing, I don't know, over the weekend.
Um, and pumped out 800 lines of working rest code. So I think that's one of the things we see is that using the right tool has always been important, but there's a huge difference in scale and cost for different tools. So I think that it's crucial that you start optimizing for the value that you're gonna get.
Now, if Russ gets value out spending $2 on, uh, on some, some code generation during the day, that may be, uh, all that's, that's important for him, but Brian spending, you know, $30 to, to knock out some significant bugs that maybe might have meant that your platform was unavailable for a while, right? That's, that's pretty cost effective too. I love that we can choose between running things ourselves and sending them to a different API.
So one of the tools I use is, is whisper to do transcription on my Mac and my, um, uh, ARM-based Mac does a really good job of running through that transcription, but when I want to do something more like extract what were the key points in this conversation, then I hand it off to a cloud-based service. And so I'm minimizing cost by choosing the right tool for it. And again, like as Ross was saying, different model for different phases of your investigation.
Yeah, it's, it's like having, uh, a handful of experts available for what's needed. Uh, I have a friend who's completely vibe coding a project and has figured out a system where he uses chat GPT, like Russ for high level concepts, but when chat GPT comes up with an answer, he asks it to write it in the form of specific directions to hand to his junior developer, and then he takes that text and pastes it into his code IDE to do the edits. Yeah, that, that's good.
And actually, I've done some things back and forth between having, uh, chat check Gemini and Gemini, check, uh, chat and, you know, go back and forth, change models. And I also use Gemini just because, you know, we, we get our email through Google, so we have paid accounts. So to me, Google is a, a free service Gemini pro.
So it, yeah, it, it's a useful checking. So you, you don't have to be tied to one model, you can use two or three. So that one that you, you said though, Brian, that that's pretty creative.
I'll, I'll have to try that. Yeah, it's, I, there's a couple, two, two articles I read recently. This reminds me of one of them, which is, uh, the writer had an insight about prompting generative ai.
And rather than trying to be, uh, thorough and meticulous in what he was asking for, uh, the writer instead relaxed that part and spend more time talking about how like, you know, write up something as if you are a developer at the end of a 12 hour shift chasing down bugs, you're exhausted and you're explaining to management for the fourth time why this is a problem. And what he found was it, it invites the models to lean into what they've been trained on. They're not really good at factory called, they're not really good at specificity, but they are incredibly creative and they understand tone and presentation, I think more than we give them permission to exercise.
Creative prompting, as you're doing ad hoc work, seems to be absolutely vital. Uh, have, have any of you come across any of the coverage of challenges around particularly vibe coding, um, and hallucinations from, uh, from these AI coding tools? Because that was something I covered on the rundown recently.
Oh my goodness. Uh, I would, you know, hallucinations looping, um, over enthusiastic extra work that wasn't asked for. You know, there's probably a list of a half a dozen bad behaviors that AI coding tools have evolved as they've gotten better.
It, it's almost like they want to prove how good they are. It's like watching a junior coder overachieve when you asked them to do one thing or and them to do one thing and they come back with like, oh, look, I did this and this and this, and I compiled it this way. Like, okay, thanks, but I asked for this.
I didn't need any of that. Yeah, it depends on what you call hallucination. I, I just call it flat out being wrong.
Um, I mean, sometimes, yeah, it depends on the model and the way you ask the questions. It'll, so I'm doing a lot of development. It's funny in a language I don't really know, which is rust, almost all my development now is in rust because I love the results.
When, when you get the code to work, it works like flawlessly and really fast, you know, no runtime memory dumps like you get with C which I've been dealing with for 30 years, so no garbage collection with Java or, or go. So I love the results. So I code almost exclusively in rust, and there's a ton of libraries to choose from.
And sometimes it'll just like dream up interfaces that don't even exist, right? And the compiler knows, it's like, this thing doesn't exist. It's like, Hey, what are you talking about?
This, this library doesn't exist. Why, why don't you check this again? So, you know, you, you could call that hallucination or call it being wrong, but yeah, it happens all the time.
You just have to understand that it's gonna come and how to deal with it. I see a Vibe coding podcast coming online soon called Russ on Rust. So if we're concerned about the hallucinations that are happening in these AI coding applications, how do you feel about handling hallucinations in production applications, running inference to, to drive your, your business, to drive your interaction with your customers, Um, on a spectrum from terrified to curious?
Um, uh, I'd say the whole spectrum. There's, there's an interesting, you know, speaking of evolution, there's interesting changes I've seen recently. Uh, I first saw it in deep research mode with Google Gemini Pro, uh, and that is, uh, listing citations.
So when the, the research comes back, when the answer comes back pointing to not a fictitious URL that they made up, but actually making sure that it's really real, and it is in fact the one they referenced. Um, I was speaking with another company earlier this week that specializes in metadata enrichment of source material. Uh, and one of the things that they do for AI platforms is offer citations down through the data.
So whether you're building a chatbot or a rag system or some other AI enabled data-driven solution, you can either in real time or after the fact audit down to citations of the actual source data. Yeah, that's, that's an important fact. And I, I've noticed some models give those citations and some don't, uh, but you need to check them because I found when I do check them, uh, I wouldn't even care to guess, but at least 25% of them are either wrong or made up.
Um, so yeah, it'll, it'll give you citations, but uh, maybe it's an outdated one or maybe it's different than what it's assuming, or sometimes they just don't even exist. So check the citations, uh, yeah, those can be helpful. But check references, Trust, but verify or don't trust and verify.
Exactly. And then to get back to your question, Alistair, about, you know, how much do you tell, um, how much do you trust AI coding assistance? Um, uh, the way I think of it is kinda like Brian described as sort of like a, a junior, you know, fresh outta college, very eager to please person, and how much would you trust them?
Hmm. You, you're gonna verify, uh, you'll let them get started, but that doesn't mean you're not gonna test the heck out of it and have some more senior people review things. So yeah, you can get a lot of code and get things working, but that doesn't mean that you shouldn't review things and test as much or maybe even more than before, because now you can automate the, the writing of test code too, right?
So why not to verify, have another agent a a different model, write the test cases. And there's a key thing Russ said there, which is, have a different agent write the tests. I think in general, you know, whether it's coding assistance or just general language models, I think it's just kind of reflective of where we're at with AI right now, or it's really useful.
Um, but you do kind of need to know what you're doing. It's gonna hallucinate it's gonna be wrong. Sometimes your rest code isn't going to work.
Sometimes your CI citation is gonna, uh, you know, bring you to a nonexisting link. Um, but I think, you know, it is getting better, uh, for a couple reasons. So, I mean, models are generally getting better.
Uh, we found better prompting, uh, techniques. Um, and you know, there, there's some other things you can do like build and rag or, or tune or model for specific tasks to make it a little bit better and not just, um, invent things. Um, but I think, you know, AI's kind of going out, but we need to, you know, make sure we're being intelligent about how we use it, right?
So, um, if you're using AI to code, maybe you should at least sort of know how to code as well, uh, so you can kind of fact check it as you go. I think it's crucial that as you're building production AI systems, you are doing thorough testing and there's techniques of grounding and, and rag, uh, is the, one of the, the classic ones for grounding give information out of this corpus of data that I gave you. And the, the, the original large language model is really just providing a user interface where you can ask things in natural language.
It's the, the primary, um, large language model isn't the thing that is giving you the, the source of truth. It's coming from your own internal, um, information as well. What we could spend a long time talking about ai, in fact, we'll, we'll spend many more podcasts and many more tech fields a events both AI infrastructure field days and AI field days.
Uh, thank you very much for joining us today on the Tech Field Day podcast. But before we go, where can people connect with you and carry this conversation on, Uh, get ahold of me? Probably the best place is, uh, find me on LinkedIn, Russ Fellows.
There's only a couple of them there. There's another guy who makes, um, mufflers in the UK for Volkswagen, Beatles. I'm not him, I'm the other Russ fellows.
com. Yeah, pretty similar story for me. I'm on uh LinkedIn, I am on Twitter, um, and then I'm on signal com as well.
com or LinkedIn as Mr. Brian J. Martin.
And of course you can find me Alister Cook on LinkedIn or many of your other types of social media. You can find things that I've written on Tick Field, Diane, on the Tick Strong side, as well as occasionally on the futurum do com site. So thank you so much for listening to this episode of the Tech Field Day podcast.
If you enjoyed the discussion, please subscribe on YouTube. If you like to see our smiling faces and in your favorite podcast application as well, so you don't miss an episode, do consider giving us a review and, uh, maybe a nice positive rating. This podcast was brought to you by Tech Field Day, the home of IT experts from across the enterprise and a part of the Future Group.
For upcoming events and more episodes, head to tech field day com slash podcasts or view us on Techstrong tv. Thanks for listening, and we'll see you next week.