Techstrong Gang – November 11, 2024
Alan, Mike, Mitch and special guests John Willis and Tracy Ragan dive into the degree to which artificial intelligence (AI) agents can take on software engineering tasks.
Then, they discuss if AI models are more or less secure than any other type of software artifact. Finally, the gang turns its attention to what needs to be done to rein in the cost of AI.
Transcript
Hey, everyone. Happy Monday to you. Also.
Happy Veterans Day and to our vets out there, thank you for your service. We've got a great show today. You know, in my best, uh, Hal 9,000 voice will AI agents evolve.
You're watching Textron Gang. Hi everyone, it's Alan Shimel here at Techstrong, and, uh, happy Monday to you. Happy Veterans Day to you.
You know, it used to be everyone was off of Veterans Day. At least I remember it that way. Schools were closed, banks were closed, offices would close.
What a great day to play Stickball in the street. Actually, by this time of year, we might have been playing a little too end touch in the street. We played something in the street, though.
I don't know how many of your kids are out playing in the street, but, uh, we're here a text drug gang today, though, and like most of you were working. Um, but nevertheless, veterans deserve a day and, and so happy Veterans Day to them. Uh, got a great bunch of stuff to go over on the, on the Gang today.
But first, let me introduce you to our gang. We've got some real core gang members at. One of my favorite people still sporting is he, we attract Yankee fans to this show for whatever reason.
But, um, let me introduce you, first of all, joining us in Colorado in his red, the band in red from the Guitar Factory. It's our CTO and Fu of VP Mitch Ashley. Hey, Mitchell, how are you?
I'm doing really well. Happy Monday. Happy Veterans Day.
Okay. Is there any significance to the red? It's just a red, red day.
That's all. I just wanted to wear red. I like you're entitled.
Looks good. I'll wear that today. All right.
Happy Reds. Um, joining us from New Mexico, one of our favorite states and one of our favorite people. She's the CEO of Deploy hub, Tracy Reagan.
Hey, Tracy. How are you? I'm great.
We have, I got in the last weekend or couple of days. We've had almost two feet of snow, so I'm a bit snowed in today. It's kind of weird.
That is kind of weird. Yeah. Boy, I thought it was always warm by you.
No, we're didn't work Santa Fe's up in the, we're we're, we are higher than, um, Denver, actually. We're really almost 8,000 feet. Yeah, We're not, you may be having at a higher elevation, but you don't get higher than Denver.
Yeah, the altitude is a lot too. Yeah, exactly. Exactly.
Um, hey, we'd be smoking here in Florida, but No. Well, no. I promise we're not gonna discuss elections anymore.
Let me go next over to Alabama. He's, he's probably the biggest Yankee fan in Alabama. Could be.
No doubt. My Son Daniel, maybe he's got Yeah. But he's not long for Alabama.
That's true. Yeah. That's a good point.
Yep. Bacher Gloo. Hey guys.
Hey everybody. John Willis. Hey, John.
Good to have you. I haven't been on with you on the gang and me and you at the same time. Yeah, I missed the big guy.
I've been coming. Yeah, I know. I always there why I come, You know, I, I get it.
I get it. But I'm glad to see your smiling face. And then finally he's, well, I don't know if he's the biggest Yankee fan in Harrison, New York.
'cause there's a lot of like sky breads may live there for all I know. Um, there's, there's a, There's a lot of Yankee fans here. It's a safe place for Yankee fans.
Yeah, It's, well, no politics, Mike. Remember you told me. Yeah.
Uh, but anyway, he's our chief content officer, Mike Ard. Hey, Mike. How are you?
I'm doing great. How are you? Good.
All right. So guys, on this fine Veteran's Day, we're gonna start off with, uh, how will software engineering evolve in the age of AI agents? Agentic ai?
Mike, why don't you set the table here for us and we'll bring our cleanup hitter in keeping with that baseball theme. There you go. So, i, BM research has been showing off some agents that they built, and they're not alone.
Seems like just about everybody's playing around with AI agents these days and they're gonna change the way we think about software engineering. But Mitch, I know you've been tracking this space and it seems to me at least we've gone from, you know, everybody wants to be a prompt engineer to, maybe I don't need to be a prompt engineer 'cause I just need to orchestrate a bunch of agents. But How do you think this is all gonna play out?
Well, we are, it, it's such a fascinating time to be part of software and John Will is doing some great work with, uh, incorporating AI into your development and your applications. This is, this is about applying AI to our code, and we're starting to see companies issue some early agents, maybe releasing a few here and there that are examining code bases and, uh, fixing specific problems or recommending fixes. Now, keep in mind, you know, there are, there are academic projects that work on really large code bases, um, you know, like that just taking, uh, 12 la 12, uh, repositories at a GitHub to be able to, um, to process those and see how much they could kind of tackle it.
com. Um, it's actually from a paper that was written by some folks center that were partially affiliated with Carnegie Mellon. And their whole idea is, let's do research by taking these 12 repositories, they've built, um, their own code on top of, um, LLMs that they're using and turning it loose to see what kind of problems they can solve.
Here we have, you know, IBM using Watson with their own LLM targeted at seeing what kind of problems that they can fix. And the announcement, I don't believe went into specifics about which kinds of bugs are they going after. I, I think my opinion is we will see verdict kind of narrow focused agents that I'm really good at fixing, you know, memory protection issues.
I'm really good at fixing, uh, buffer over flow issues or whatever it might be. The, the challenge is that when you get into, you really get into code, not just generating code with an LLM today is software is very interrelated, right? There's libraries, there's code.
You write, there's calls going between lots of different elements of code. There's the environment that it's operating in and using some of those service services from your application. So to get to the point where we're really fixing some of the real vexing ones, um, is a little bit off, I think, a little bit farther down the road.
But we will get there. I think 2025 will be the year of the code fixing agents for in, in dev AI and DevOps. That's not more other things too.
John Tracy. Yeah, go ahead. Yeah, no, I've been tracking this space pretty hard, uh, because, um, you know, um, I, I think this whole thing got exposed.
They called SW bench, the SW bench paper and, and the paper, it's, it was actually just, I, I think one of the guys was CMU, but mostly Princeton guys, and now they've done a bunch of startups. But there was a, I, you know, we don't have to go too far into it, but Devin, back in early this year, like, um, uh, March. So, so that the whole suite bench is, it's mostly Python code.
You know, it, it's a lot of problems from existing bugs. And the SWE bench people wrote a paper and said, you know, could we use sort of some type of ag agentic processing to solve, you know, well-known bugs, bugs that already had this sort of issue, the fix. And then they went back and got the original code and then you benchmark against it.
And the original SWE agent from the Princeton, uh, students or um, researchers was basically like somewhere about 9%, 8% accuracy. And then, um, and then this Devin came out, which was really sort of a marketing ploy. It was a bunch of competition coders.
They actually got, uh, you know, funded by what's his face, um, uh, Peter Thiel, right? And they got like $20 million just kids that win competition coders. And they, and they made up a bunch of s**t.
Um, and, and they, who Else? Peter Thiel's behind it. What do you expect?
Yeah, yeah, Yeah, yeah. We said we don't wanna get throw code, but, but they showed up with 13%. It woke everybody up within a week.
There were like, there were a bunch of people that were saying, I'm doing this too. And then this open Devon came out. And then, you know, over the time there's been a bunch of solutions, uh, ada, and I don't know if you probably people have heard of Cursor.
It's all the rage right now. So this, this thing is pretty heated as open source is proprietary. The place I'm more interested in is what the big people are doing, right?
So I saw the IBM thing, uh, that you guys had posted, and that was pretty cool. I've been putting IBM unfortunately on the back burner, you know, to, to see if they really flesh out to be competitive with, you know, with, you know, some of Amazon, Microsoft, uh, OpenAI, all those guys, right? So, um, and so my focus has been looking like, you know, so my, I got early in on, uh, Microsoft copilot workspace.
So you got copilot, which everybody knows they created a workspace, um, I don't know, probably about five, six months ago. And I got early access to it. And it's sort of interesting, you know, I, I've gone in and I've tried to get, it's very buggy, a lot of memory issues, like you said, Mitch, there's, uh, you know, it, it, these things are still very toyish, but, but they do pretty phenomenal things.
Like the workspace, you literally grab an issue in GitHub and it analyzes the whole repo. It then comes up with explaining what the code actually does. Then it gives you, say if you wanna go next, it actually generates a code to change it.
Um, and then you can sort of, you change it, update it gives you, you know, ability to edit that. And then if you, you wanna go forward, you, it generates a pull request, right? And it's, it's pretty crazy for basic things, you know.
Um, and then, you know, the, uh, Amazon has Amazon queue, um, developer, um, I'm, I've been trying to get to see what Google's doing with, um, their vertex. But the interesting thing is, I think they're all going down the wrong path, in my opinion, because they got, it's a hammer nail problem is all these benchmarks and everybody's, and I, I saw, I, I looked into the actual what IBM was doing. They, they're on the SW bench.
There's a SW bench, uh, benchmark, and they're sort of listed. And so all the big guys are doing about 25, 20 6% efficiency in percent resolved. Um, but the problem that is, is hammer nail problem is, this is the real solution here for me.
It's the enterprise, which is Java code, which is banking code, which is retail code. And guess where none that code is, it's not in GitHub. And so, uh, so there's a couple of us guys from Cortex and a couple of sort of core Java people.
We're trying to see if we can't take the SW bench model and use, you know, get some organizations to see what, how this sort of agentic process works when a lot of the coding structures are not well known. Like, 'cause the copilot knows Python really well based on what's on GitHub and Python code. The argument is it might not know, you know what JP Morgan or, uh, capital One who had 20,000 Java developers at one point, that code base is not available to copilot.
Right? So it, so there's an interesting, I I think this whole thing is absolutely fascinating. I think it's real.
It's baby steps. I think the copilot workspace is, in my opinion, the most usable tool, even though it's still in beta. But it is crazy buggy.
I mean, crazy buggy. But it solves, it does solve some wicked problem, you know? So, yeah.
You know, Mike, you said something sent a chill through my heart, Yikes. Through my heart. So we wasted our time learning how to do prompts.
And it is just as agents, agents is the way to go. I'm not gonna be prompting. I just tell my agent to go do something.
This is the debate that's, people are starting to ask the question. I don't know the answer for sure. But it seems like, you know, as we keep going down this path towards agents, the agents are just gonna be told what to do in plain English.
And I don't need to know more than that. I absolutely think that's where we're headed. We're headed to natural language being the primary UI for instructing these things, right?
That's what John's talking about, right? Is you're, you're, you're giving it something, telling it, go look at this. Right?
Or, and now you can, you know, if you're a user of Grammarly, it has, you know, AI built into it, but it comes up, you don't prompt anything. It comes up and says, would you like some recommendations? And it picks out sections of something you've written, and you can choose to adapt those things or not.
So as, as AI gets embedded, whether it's agents or in, in the ui, you know, now we know it's ai, but I think writing the prompts we're not very far from the point where you just say natural language, what I want. Yeah. And they go, write me the prompt.
Here's a point I want to get in. This is something I realized. I, I was showing a a a, a development team, you know, these guys have a startup.
It's the X rancher guys, right? And I was showing them the, this workspace, and they're like, okay, that's cool. Let's try this issue.
And it failed miserably, right? And then I, I said, well, wait a minute. And I looked at the issue description, and the issue description was not well formed.
So like this, that you're actually right. The language that we used, you know, you look at that IBM paper and they're going through, if the issues are not, I had a conversation at, at Gene's conference with the guy who, one of the problem per managers of GitHub next. I'm like, what we really need is these tools to create well-formed issues, descriptions.
So then when the tool goes in and tries to understand what that is in the repo. Yeah. And so it is, it's, it's always going to be, whether it's agent or sort of general, just prompt it is gonna, to your point, Mitch, it's natural language.
And, and so we need to make sure that we're, we're we, we've gotta get better and crystal clear about our questions, our intent, and our goals. Well, lemme just make a note here. Cancel the prompt engineering training, but Okay.
Say Corporate training. Yeah. And there's, there's another topic that we have that we're not kind of diving into here, and that is the culture of these tools.
Uh, you know, I've worked with developers for a very long time, and they really do see what they do as a craft. Um, I believe that the pull request is the way to go to say it's such a, basically suggesting what they need to do. But really and truly it needs to be more like Grammarly.
It just needs to be, if, if we're going to make it really work, it has to be fed directly to them. They have to know it as soon as they're coding, it's gotta be done right in their face. Now, for maintaining a mature software, let's just call it that legacy software, that may be helpful, right?
So if somebody doesn't have to look at that code, 'cause I've been on projects where you're having to maintain legacy software and it's not very fun. And if we can make that more efficient, that would be important. But for developing new code and new applications, um, I think we have a way to go before the technology's there that we can trust it.
We always talk about trust. And before developers are willing to adopt it and give away their, you know, their creativity skills. I had one developer tell me once about builds scripts.
She was like, well, I like making my own shoes. My response was, well, is it true that the shoe cobbler's children never has shoes? Maybe we have to be doing, you have to be, you know, a little more efficient in your work.
And some of these tools are about efficiency, but developers are not always about efficiency. So we have to see how developers work. And I think serving it to him is a, is is a more direct method than to create the poll request, even though I think that's a good way too.
But, uh, it, I think they'll be ignored. Yeah. Really good.
Important point. I wanna also raise, you know, you remember James Wickett that we, we've talked to before? He is now, He's still good stuff with this Yeah.
Driver and security. They're doing some really interesting things. Have, have some product in market now, but, uh, analyzing code, I mean, the whole repository looking for where might there be security exposures or security compliance.
You know, somebody started using a different library than we normally use for, you know, um, the shopping cart. Did we know that was happening? And is that in compliance with what we're looking for?
Or, you know, here's a code that allows command injection or SQL injection kind of thing. So people are doing, I mean, it's, John and I probably could go have dinner and, you know, but about noon the next day we'd stop talking 'cause we just couldn't talk anymore. There's so much going on in this space.
Absolutely. And it, and that's why we talk about it. And that's why people come to text and gang to hear about stuff like that.
Mm-Hmm. But someone's gotta pay the bills. We need to run a little, uh, commercial here.
We'll take a break and we're gonna come back and talk about, are these AI models any more or less secure? We used to have this conversation, open source, closed source, open source was more secure, less secure. Now the same, same question for AI models.
We'll talk about that and more. You're watching Textron game. We're back.
And as Alan said, we're gonna talk a little bit about, well, just how secure are these AI models that we're about to inject into all our software? Because there's a conversation out there that says that they are subject to the same kind of vulnerabilities that all other software is. And there's a list now from the folks who do, uh, oasp.
And it's not clear to me though, that these are more serious vulnerabilities than the ones we know already, or if they're just basically the same ones dressed up for an AI model. And an AI model is just another type of software artifact. But John, I know you were looking into this.
What's going on here? Yeah. You know, the answer's both Mike.
I mean, that's the scary part, right? Um, you know, my, I I do this presentation dear. CIO I've been doing it.
And, and really my sort of passion now is this shadow ai, this technical debt tsunami that's coming down the pike. And the, the fear that, and I know I've said this before in a show that CEOs are hiring chief AI officers, chief AI harvests are hiring kids from Stanford and, and CMU. And they, and they tell me things, and I've had these conversations where, you know, what's so hard about installing Kubernetes?
I did it in 15 minutes. Like, yeah, do you ever hear the plumber joke? You know, you know where they hit the hammer?
You know, I mean, they, they just don't know what they don't know. And, and CIOs are being told, and I get confirmation of this over and over the stay that this Chief AI officers are told to stay away from the C Ccio O and ciso, um, because this steward go fast. And it, it's, it's, and what, what we're building is scaffolding of infrastructure.
And so the, so the answer is both, right? You have all those vulnerabilities. Now, ai, the, the problem in, in a couple years from now, maybe in a year from now, we wouldn't even call it ai.
We're just gonna call it services. We're building stacks on Kubernetes and Kafka and Redis and all this stuff. And 90 85, and this is just my napkin math, 85 to 90% of those AI solutions, those chatbots are just plain old cloud native infrastructure.
And if the people who normally protect the fort are not getting involved in building that infrastructure, that's ground zero for problems. But then the scarier part is the people who protect, don't know how to protect in this new world, right? And so that, this is where I get into what, you know, I, I think, you know, like, and you know, I haven't done my sort of exhaustive check lately, but NIST has done a terrible job so far, in my opinion, of trying to protect us in ai.
I mean, when you read, in my opinion, you read the NIST discussions, it's, you replace Gen AI or an LLM with Google search, and it's like the same thing, right? Like, oh, all the scary things that you could get from a, you know, and, but, but, um, but oas, which is the sort of the article I pointed out, like oas, I think is oas, what OAS is doing right now is I think they're doing the best job now. You know, our good friend, um, Tracy Bannon, I think Mire is doing incredible research.
So, you know, they, they definitely have a gold star. But OAS is, to me, when I, you know, was getting involved in this stuff in the eighties, like ITO was not great, but it was the only game in town, right? It at least gave you a way to figure out how to manage all this chaos that was happening in the, you know, mid to late eighties, early nineties.
And I think to me, I was thinking about oasp. I mean, they're just doing a really good job of rolling up their sleeves, you know, an O osp, uh, LLM top, top 10, LLM is a, you know, is a great starting point. Um, and then this paper that, you know, this thing, this sort of update to the LLM, um, top 10 is, um, they've got, um, a COE guide, guidance paper.
And I, you, you know, I like glass half full, glass half empty, glass half full, is they're, they're laying out like how you should, what teams should be involved? What is the roles and responsibility? What is explainability and ai, um, prompt injections.
Um, you know, there's these, uh, these, uh, jailbreaks are like really scary, you know, I can get into that where people are poisoning public documents knowing that a corporation that talked about their new GPT is using that public document, and they're putting bite code or jailbreak, you know, bite code exception in the text that they know will show up in a rag, and that'll get put into the context window. Yikes. Yeah.
I mean, like, it, it, it like, so like, dear CIO, like, you can't fall asleep on this one. You gotta be the hero. So anyway, so, but the o again, I, my sort of meta pointed out is their papers.
I think OAS is doing the best job. The only thing I, I think all of them sort of over rotate on the fundamentals of security, which have to be done necessary, but not sufficient. We need to teach the protectors how to talk this new language.
Like, I'll give you one last example, which is, you know, they, in both of the papers, which just came out like last month, um, they don't talk about observability and observability in AI as a completely different animal. It's about evaluations, it's about hallucination management, it's about correctness, it's bias. There's no mention that I can see of any of that.
And that's about as important a subject that I want a CIO to know right now is how important it is to have that tool as a tier one portion of your arsenal. So yeah, that's, that's sort of my take on it. It's worth reading the updated papers from oasp and, you know, so Look, I, I applaud Oasp for recognizing this as a security issue and shining it in the light of day.
My question is, five years from now, is that OAS top 10 list gonna be the same, because they don't tend, you know, one of the knocks on no WA quite frankly, is, you know, the, the look, they were the ones who kind of put AppSec on map, right? But their top 20 AppSec list hasn't changed very much in 17 years or whatever it is. Um, yeah.
But, but you know, the thing is that, like in this, we know this is an accelerated timeline. I mean, the O Ops top 10 12 is already, it's sort of like, it, it is like, you know, I'm not a security person by trade, right? So when I started doing DevSecOps, I navigated quickly the os, right?
And then I started, oh, okay, now what is that? You know, an sq I mean, I knew what an SQ injection was, but, but, but some of that stuff is still good for people to understand how attackers attack, right? Mm-Hmm.
And, and so I think, you know, you're right, the oosh TAP 10, which came out like, what? Last year, you know, or earlier this year. Earlier this year, I, it's already sort of like mutated, like the, like that example I just gave.
Like, that's not it. That's something that just showed up in the last month or so, you know? So, but I think it's, it's a good foundation.
It's a good place for people to start. So, you know, like is no original top 10 good or bad? In the grand scheme of things, I Think I don's good.
I think, I don't Think the fact, I don't think the fact that the list doesn't change is an indictment of oas, but I think it's an indictment of the development community. 'cause we keep making the same mistakes over and over again. But, you know, is what it is.
Well, Here's what I like about the OAS top 10 for LLMs. It's, is it's very approachable. You don't have to be a John Willis to the depths of all the things that are happening behind the scenes.
It's things like prompt injection, uh, insecure handling of prompt, uh, prompts and output handling. It's data poisoning, it's denial service. It's, you know, there are things that people are gonna recognize, kind of bigger macro issues.
And yes, there's a lot of detailed things behind it that maybe it doesn't cover yet. I don't think we'll see more evolution of kind of getting into the bits and bites of how we protect this. But right now, you know, to your point, John, about, you know, what, what do you want on the CIO's list?
I would take the O top 10 and sit down and say, these are all the things we need to start working on. Absolutely. Absolutely.
It's, it's not the, it's not the turn the knob, it's the, you know, get the train rolling. Let's get started. Yeah.
Mm-Hmm. Pay attention. Right?
And maybe we should start with the nist, uh, software development framework, you know, because right now most companies haven't even, you know, a achieved levels of security in traditional development, you know, just what we're delivering today, much less all the problems that we'll experience with our LLMs and this, you know, you know, everybody knows by now, I'm not a fan of agents. And the reason why I'm not a fan of agents is because it complicates the process from a security, uh, standpoint drastically when it comes to managing the thousands of agents that we will be seeing. It is a, a death star that is going to be almost impossible to manage with the tooling that we currently have.
And I get frustrated because I feel, feel, I feel a huge hole in the force. Yeah. Thousand agents, all of stop, thousands of agents sudden.
It is a very complex dependency graph. It's extremely complex. Mm-Hmm.
And the stack that we're building is massive. Tracy and I look Tracy at it every day and go, why? This is catch 22 here, though.
This, I, I, I, I, I, you're, you're so right. And spot on. Like, I, I talked to somebody who's really big for one of the big four, big five, and, and he was bragging about how they've built tens of thousands of bots.
And I was like, the first thing I think about was RPA and what a disaster that was. Right? Like the, the screen scraping across the whole enterprise, right?
Um, so there, there's something in the middle here about agents, you know, and I know haven't quite figured out, but you're dead, right? If, if an organization thinks they're just gonna write 10 thou, or they're gonna bring in some big four big five consulting organization to create, you know, 10,000 bots to just automate everything, it's gonna be a dependency. Absolutely.
And just, just trying to manage the drift. Oh, on those, no doubt. It's that that will be a challenge that people will be struggling with.
And, you know, and from a security standpoint, you don't have to hit all of those bots or all those agents. You only have to hit a few. Sure.
Yeah. Yeah. Right.
And so how do you know those few have been hit? So the more, the more agents we put out there, the more difficult, uh, securing and managing and the platform engineers are, the, the, their, their time is going to be spent doing a lot of very gory detail work. Well, the problem is, Agents are, and right now we just don't have tools.
Agents are, it, it's like chatbots. Agents are easy to produce, and that's why we're gonna see a plethora of agents being created. The real game changer is what you were talking about, Tracy, is when you really, when you really ingrain AI usefulness in the workflow of people doing the work, you know, the Grammarly example, right?
And what, what is it doing behind the scenes? We're, we're using agents to kind of do analysis on things and, uh, do some basic workflow kind of steps for us that can benefit from things like context graphs and, uh, that inform the agent a little bit more about how to actually understand the data that it's, well understand relatively speaking, the information that it's, you know, that it's analyzing. So I think we're, yeah, we're gonna be smacking a lot of those agents for a long time.
They're gonna be around for a while. Not that that's what what we want to have, but I think we, we need, We need some, we need something that feels like an orchestration framework for all these agents, but then the agents are gonna conflict with each other. And how are we gonna sort that?
And that seems to Tracy point a lot of complex software. And I'm looking forward to continuing coverage of this subject all the way through 2027, maybe 20, 28 old Bar. Oh yeah.
Once the agents are out There, we're not bringing 'em home. Yeah, exactly. There'll be sworn of them.
Attack has, do you guys know, Are there any Agent Orchestrator projects out there right now? I haven't heard any. Well, This is why I, I think that the, that's why I want focus more, even though like, things like Cursor and ADA and Open Devon are very fascinating for like, sort of a hacker mentality.
Um, I, I think the promise has gotta be with Microsoft. Um, Google. Yeah.
And, and so that's why I've been really tracking, you know, Microsoft co-part workspace, um, what, what Google's doing with Vertex. Um, you know, and, and because, you know, they, they, they're gonna have a grip. 'cause thing is they use these tools to manage their own software base, right?
So like, and they're pretty good at that stuff, you know? Um, you know, so, so I think, you know, I think there's, I think we're gonna find really at, you know, hopefully I, I, I use RPAA lot as, as like the, the bad story. Like, you know, everybody, you know, five, seven years ago, every CIO wouldn't talk to you about anything other than RPA, right?
And RPA turned out to be a disaster, except it solved those real narrowing, not any problems, right? So, like, you know, I had a guy one time, and I, I was dragging on RPA and, and he comes up to me after session and says, John, you know, during the Covid SBA loan thing, we had one week to, and well, we didn't have APIs, we had a doc, A PDF document, and we had to have that thing done in a week or else we weren't gonna get, you know, our customers were gonna go to B of A, right? And RPA saved our lives.
So I think the, the way we have to think about is the way we always try to somehow fight our way through this battle with the intelligent and how we use these things. Mm-Hmm. And just, this is this, Alan, to your point.
So IBM has the skunk work project where they're talking about orchestration of agents. And that's one of the things they showed. Every vendor who has anything to do in this space has a similar project going on.
And this is gonna be the next big fight in the, the IT sector. It's gonna be control over who's gonna dominate on this whole orchestration framework thing. 'cause they all smell it as a huge opportunity.
I agree. It's the next coob. Alright, let's take a break.
We're gonna come back. There's a topic, you know, do we need finops for ai? com is the number one online destination for DevOps education and community building.
com covers all aspects of DevOps, including DevOps, best practices and tools, DevOps culture, DevSecOps, business impact, continuous testing, continuous delivery, and more. com has the largest collection of original DevOps content featuring breaking news, blog posts, podcasts, and more. com to learn more.
com where the world meets DevOps. All right, we're back. And yeah, we're talking about finops for ai.
Turns out that people are starting to figure out that AI is not like butterflies. It's not free. It's rather expensive when you sort it all out.
And the cost of the infrastructure in particular seems to be super expensive. And maybe we'll get smarter about how we build these things. But Tracy, I don't know, are you talking to folks out there on the IT side?
Are they starting to ask questions about how come this stuff costs so much and what do we do to reign it in? Well, I think this is an age old question. Um, how do you reign in cost in software development?
We've always had this discussion, and every time we have a new technology, we have this discussion. When relational databases first hit the, you know, hit the road, everybody ran with it. And then they started realizing that they have higher cost.
And how do you get control of that? And we talk about, you know, chargebacks and not being able to see them. But I believe that companies are starting to now recognize that they have a potential, uh, to, to start reigning back some of those costs.
The, the, the charge forward in AI has been a little crazy. It's like when fools rush in, everybody just starts throwing code at, and everybody tries to start building these new, uh, services as if we're gonna call 'em services, building new agents. And this does have a cost.
It's not free. It's, it's expensive. And we, you know, we, we understood that Kubernetes had the potential to allow us to have software that could self-heal and could grow when we need it.
But, and, and that we saw an immediate benefit. I think the problem here is we're not seeing immediate benefits from AI yet. So we're spending a lot of money and we're not getting a return.
And while we are doing the best we can to manage these costs, there are still areas that we don't know about. We really are, we're struggling. Companies are struggling with trying to have some, a little bit of transparency to understand if the extra cost has improved de efficiency.
What is the ROI and what we've already, uh, we've already delivered many of these AI tools are not meeting the demands that people had originally thought that they were going to create this massive new efficiency around many tasks that we do. But it just hasn't gotten there yet. So we have to go back to, you know, basic, um, management of, you know, one of the articles we talked about, technology business management.
I haven't heard TBM in such a long time, but I think right now we need to start talking about it. We need to understand if our enhanced cost has improved our efficiency. We really need to understand, um, this, the, the, the increase in this, this financial expense and have some transparency in it.
'cause most companies are struggling with it and they don't do a good job of sort of data-driven decision making right now. I feel it's very emotional. And we're just getting around to saying, this emotion is starting to cost us quite a bit of money.
And while we may not know exactly which team is costing us money, we have one big bill. Um, we know we gotta do something about it. And scaling back and really understanding why we're developing, um, ai, why we are spending this money and understanding what benefits we're bringing to our end users and our our customers is where we have to be.
We have to do that first. I mean, don't just tell me I have a great AI tool that doesn't really make my life better. I, you know, I have some friends and we decided that every week we would say, what's your favorite AI tool that you've come up with?
I've only come up with two Grammarly and this really cool tool that, uh, now I don't remember the name of it, but you can take a picture of your, your food at a restaurant and it'll give you all of the, the calories and the protein content of it that actually makes my life, uh, really happy. But what, what are the tools that are being delivered and why are we spending so much money? And is AI delivering?
I don't think it is yet. And maybe we have to spend a lot more money and have a lot more fools rush again, uh, before it's making money. I, I sort of disagree on a couple points, but respectfully.
But, but I totally agree with the cost rating in there. People are not putting any sort of mythology. I mean it like, like I love just looking at history and saying, oh, there we go again, right?
Uh, cloud, private cloud, back to cloud, back to, to some semblance of, we did little of both, right? We went all into a cloud, cost got high, and a bunch of organizations tried to pull everything back. They realized they didn't get, you know, all that story, right?
We're going through that right now. Um, on a personal level, things like copilot and coding agents have changed my life. I've, I've coded more in the last year and a half, two years than I've coded in my whole 40 years of career.
Um, but it's, you know, I've been tracking some of these organizations. I, I, you know, I'm a consultant for, uh, PU and at Margo db, um, I do a lot of that work at hackathons. They, they, there are real, there's real stuff happening.
You know, and even at Gene's conference at ET TLS Adidas, you could watch their presentation when they talk about performance gains and they're real, there's, these aren't just, uh, Fernando co I, I can never pronounce his last name. He's a real deal guy. He's been in the DevOps community forever.
Uh, he runs digital transformation there. Adidas, when he gets up and talks about the, the productive gains that they're getting, they're not like Spotify blog articles. Um, Adobe, you know, uh, Brian Scott over at Adobe who used to be at Disney.
Like, he talks about what they're doing. Vanguard, uh, they've created a, an audit GPT tool that they're giving their field auditors that are just, and they're, they're, they're showing. And then, you know, and then the, uh, John Deere, you know, I've, I had some great conversations with John Deere.
So there is, I, I totally agree with you on the, um, like I don't think these organizations are, they're, they're starting to hit the brick wall on token costs, but, but they are, um, they have gone into it like the fools rush in. But, but there are, I mean, these, these companies are not, you know, they're not, you know, they're, they're pretty bright about their business and I, I think they're tracking the invest, the ones I've been following. And the only other thing I wanted to add is there are ways to sort of save on costs that you need to sort of understand, which is, you know, things like, uh, you know, we've, we've talked about carbonization on this show.
Um, you know, just, you know, rewrite re-changing the model, instead of having, you know, uh, a billion, uh, weights and parameters all as 32 bit float, you can make 'em four by inch. Right? And you get the same performance for most problems.
Um, crump compression. And then, um, you know, this, uh, this really interesting thing, I should write an article for you guys on this. It's called LM to vec.
And it's basically, it, it's, it's fascinating. I think what the opportunities of this is, is basically it's, um, just quickly, you know, when you go into chat GP and ask a question, a prompt, it's called a decoder model. It's like, here's my sort of sentence, go it, you know, sort of attention based mechanism.
Go find out what the next, this is what they call bi-directional attention. So it can actually do what the training models do when you create the embeddings so that it has this ability to bidirectional. So you can literally create amazing efficiencies if you start taking things that you would normally do in a prompt with a decoder like your G PT four.
Anyway, Hey, I'm getting very technical now, but the, these are three right off the bat that I'm recommending to some of my clients right now. So, And you know, I don't disagree that in, in the, in copilot and chat GBT chat g bt I use every day. But what, what software are we delivering to non-technical people?
Let's talk about everybody else, because that's where the big money is, where it's going to impact the lives of everyday individuals who may not be in tech. And I don't see AI delivering a whole lot Yeah. Benefit in that.
And that's what's puzzling to me, because that's where the unicorn is. No, I mean, you gotta look at JP Morgan Chase and what they're doing. Look at, uh, there's a, one of the largest workforce automation companies that they're literally revamping their product portfolio for workforce automation.
I mean, these are, I mean, John Deere is doing sort of supply chain logistics, like just changing the game. So there, there's real stuff happening. Uh, we're not seeing it.
That's the other thing. A lot of these companies right now, I can't mention the company, the workforce automation company. 'cause they actually, not only was I under end NDA for the client, I was working through them, they made me fill out another NDA.
They're so secretive. So a lot of these companies are not giving away how they're using this information. Come on, John.
No one's watching. You could tell us. Dang, you Just a friends.
I would like to know, I would like to start seeing real products come out that make our lives. I would, I it's a, it's a double secret. Indeed.
It was a double. I was, I was, it cancels the other one out. So it's not secret.
The second day ETLS the whole morning is Adidas, Adobe Vanguard and John Deere later in the day. And it was, you know, I, I told Jean that I thought it was the 10 deploys a day at Flicker day. Um, because these were real companies, um, showing how they're doing like real stuff with this technology and not this sort of, you know, sort of nonsense of just, you know, it, John, you said something about, Hey, you like watching history.
The fact of the matter is, what you're watching here is the creative process at work in normal times. This is how stuff works, especially in technology, but it works like this. I think through all creative processes, first you try to get something.
You, you have the belief that this thing could do something, then you get it to do something. Then you figure out how to do something more efficiently, right? To the, the efficiency piece of it comes in and unfortunately, somewhere near the caboose of the train or behind someone says, what about security?
And, and, and we, and then we say, oh, we gotta do it efficiently and secure. Done Why? My head blows up.
Yeah. Yeah. And we over, but that's normal.
That's the normal process. Can we do it over and over? Why can't we just, but my dear CIO is why can't we just like, at least this time, learn a little bit.
We've got, so maybe It's human nature, maybe it's human nature. Maybe it's tied directly into the creative process. I think the momentum working against us too is, you know, every, every CIO, every CEO has been told, if you don't have an AI strategy, now you are gonna, you are a dinosaur, you're gonna die with the dinosaurs.
It's Stuart dies what I hear. And so it's like over rotate. But at 9,000 RPM, you know, we're like way over, over pushing this.
Well, But think about that though. I'm Not saying it isn't the process. I think we're just the, the mount and the speed is just, I don't see we're delivering anything better.
I mean, I, John, I'm glad you're having a different experience, but from where I'm sitting and I'm not getting a view into John Deere or jp, I don't see it. So I just see a lot of money being spent on AI and not a lot being delivered. Oh, I think there's absolutely a lot being delivered.
Talk to my sales team. I just ripped them about not using AI on our sales things. We should be, it's being delivered.
It's being delivered in software. It's being delivered in marketing. It's being delivered all over.
But just let, let's think about, I'm gonna ask you something. Substitute nuclear fusion for ai, right? Right now we're just trying to get nuclear fusion to work in a lab or an at to car or whatever they call those things that they, they do it in, right?
Then once they get a sustained fusion reaction, we're then gonna work to make it commercially feasible. In other words, figure out the cost involved, right? 'cause a sustained f reaction gives off more energy than it takes to create.
Then we're gonna figure out how to make that efficient so that we can do it at scalable, you know, scalable ways and, and we can start really powering our world on, on this stuff. Then we'll figure out how to make it so that no one can come in here and do something catastrophic with it. But that's, it's science.
This is science. This is how science works. I, I don't, I don't think, well, first of all, I was gonna say, are you in the Colorado high or the Santa Fe High?
Right? Look, I got a whole universe on my thumbnail right here, man. Wow.
Um, its a long, I seem to remember having that nuclear fusion conversation in 1978. So I don't think it's moving along that very far. What in the next five years we're gonna have commercial level nuclear CLE fusion.
Well, I do think that we quantum Computers, I, I do think that speeding up this process and making, um, AI more a, a more real, an easier software to learn and to develop in and to deliver software to end users, it's gonna make our lives better. I think that these managed, uh, platform services and cloud infrastructure services that provides everything in one place for us, um, will be helpful. I don't think it will LA they will last for the long run, but I think it gets us started and I'm, uh, I'm, you know, I don't know if many companies are going to a single platform, um, like a, a neas, but, uh, I believe it will be hel it will be useful for new companies getting into this and cutting costs down.
Fair enough. Tracy, henceforth, your Delta House name is AI skeptic I's so skeptic. I'm just, I just wanna push and push and say, if we're gonna spend the money, you better be delivering.
Okay. I thought Mike got all the skeptic titles. It's not fair.
No, from now on, Mike's known as Pinto. Pinto. Oh, oh, okay.
Alright. Hey, we're about outta time. We got to call a rap on this wonderful Veteran's Day Monday version of Textron Gang.
Tracy Tra stay, uh, safe and dry and, you know, warm out there in, in, uh, Santa Fe. John, it's always great seeing you, my friend. Give, give hugs to the boys and Vicki.
Thank you Mike. And, and Mitch, thanks for joining us guys this week, Mike, Mitch and I are in Salt Lake City, happening Salt Lake City, where we will be covering the, uh, CubeCon event and we'll be, we'll be doing a couple of Textron gangs live right here in Salt Lake. I don't know if they'll be live, but we'll be doing, we'll be recording some Textron gangs.
We'll be broadcasting live Wednesday, Thursday, Friday from Q Hey, I, one quick thing, which is my, you can call me Delta House. Uh, the Cynic, isn't that the conference where 5,000 people try to learn how to wear to put semicolons and commas? Something like that?
Yeah. Yeah. It's 15,000 this year.
Replaced by ai. There's an agent for that. John, we we're gonna hang out at YAML house.
Is That what you saying? There's multiple agents, Yael House, Louis, And if you're, if you're in Salt Lake, Oh, okay. Enough.
There's an Toga parties, there's an NHL game. There's a good game in, in Salt Lake City with The Utah, what they called Utah Hockey Club. Hockey Club and the Golden Ices.
We're seeing. There's Also also App App Dev Field Day happening next week that's gonna be broadcast from Salt Lake. And I think the Utah Jazz might be playing there next week as well.
Who would think that Utah would have a basketball team called the Jazz? Right? Salt Lake City.
The birthplace of the blues. Ho Ho. Anyway, have a great day everyone.
Check us out at CubeCon. This week is for now, though. We're out.



