Navigating AI Challenges in Platform Engineering – EP 8
Transcript
Hey everyone, it's Alan Hummel. Welcome to another edition of the Platform Engineering Show. We've got a special edition of the Platform Engineering Show for you today.
In addition to Mike Compadre, who's in the Alps today of all places. org community. Is that fair?
Contributor. Contributor to the community contributor. Okay.
He's humble too. So my co-host, Luca Gallenti is here with me and our very special guest today is my friend Keith Townsend. You may know Keith, he has a long career advising some of the biggest organizations in the world.
Keith's company is the Advisor Bench, right? That's TAB, the Advisor Bench. You could check that out.
You could follow 'em on LinkedIn. Keith, thanks for joining us on the Platform Engineering Show today. It's great to have you on, Man.
As always fun talking to you, Alan. We always have a good time. Appreciate it.
All right, so today's show is really stems out of a LinkedIn post Keith made that kind of hit home with me 'cause it's the kind of thing, you know, we were just a platform con, what was it, two weeks ago, Luca? Yeah. And we had people there from Google and Google Cloud and AWS and the hyperscalers.
And of course everyone's talking ai, AI platform engineering. And Keith wrote a, a kind of provocative post, the missing link in enterprise ai, why platforms are failing to empower Dev teams. And you know, what hit home for me was, it's not the platform engineer's fault.
I don't know if it's the AI company's fault, but the, somewhere here there's a disconnect. Keith, you could frame the the problem better than I. Why don't, why don't you frame it and then Luca jump in.
Yeah, so if we look at what the cloud providers, let's focus on the cloud providers and Luca, I would imagine you know this extremely well. 'cause this is your area. They've done a really great job of standardizing, at least serving, not, standardizing, might not be the best word, serving the platform engineering audience.
You know, we, we've gotten to a point where platform engineering as a discipline has evolved from trying to make every cloud look like a single cloud to really getting developers what they need. They need observability developers and operators. They need observability, they need, uh, standard patterns.
They need these tools to develop and maintain applications. So if we look at the AI services that most of the cloud providers and the big OEMs like Dell, HPE and Lenovo are providing, is still very much focused on that early stage of ai. Uh, the ability to just give me raw power, the ability to train models, the ability to consume APIs for the most part.
But then it starts to fall off when we look at this kind of solved problem. And I don't want to treat platform engineering as a completely solved problem 'cause it's the always evolving. But for the most part, you know, some of these big problems like observability, et cetera, they're solved.
We've, we've learned how to do that. And ai, you know, we just saw it with, with, uh, Xi Xxi the other day, X ai the other day with, uh, when it went off the rails. I don't have tools internally to prevent something like that from a app platform engineering perspective.
Every team has to recreate their guardrails. They have to do their model observability, they have to do their pa the development pattern, uh, process over and over and over again. And that is where we're at today in the maturity level, uh, uh, for providing AI as a platform to develop on.
Yeah, I I, I totally agree with you, Keith. Um, I think, you know, what you're hinting at is, is essentially the space is just not enterprise grade, right? It's not enterprise ready.
Um, they, I think like a lot of the use cases, and this is across like development, but also just like general, like enterprise users, right? Like, you see, I think like, you know, the only applications that work in, in production, let's say, are really like, basically like basic customer success stuff right now. Um, you know, and some, some kind of like o gem stuff on the individual contributor level.
The problem is to your point, is like when you're trying to like, tie those things together, like across different individual contributors, and especially across different teams or business units, different departments, then everything breaks, right? Because there's no sort of like enterprise wide guardrails and enterprise wide system thinking even, right? Um, now I think like one thing that I would like slightly challenge is, you know, when you mention like, you know, observability and I guess like, you know, security and like all these things, right?
Like, I totally agree with you that they are like, for the most part solved problems. But I would, I would, you know, slightly reframe it in the sense that it is not really, like, that's not really what platform engineering is about for me. Like, that's kind of like the difference, sort of like infrastructural silos, right?
And then platform engineering is really about like, okay, how do I tie this, this stuff together into, you know, a self serviceable layer that I, that I serve to developers as a product, right? And I think that's exactly, and so this is why, you know, you mentioned this, right? Like it's, it's never evolving problem, right?
And this is why I actually, I'm very excited about, you know, kind of like intersection of AI and platform engineering and like, Alan, you mentioned platform con, you know, we had it like two, three weeks ago. It was definitely like the hottest, you know, very unsurprisingly the hottest kind of like topic that everybody was trying to tackle from different angles is like, you know, what does the Venn diagram between AI and PE looks like look like? And you know, how, you know.
And then I think like, we know, we talked last week or like two weeks ago in the show about like this dichotomy between sort of like AI enabled platforms and platforms for ai. And I think like really what, what you're attaching on Keith is like the latter, which I also would describe as actually where the money is and like the important thing to solve, not necessarily how an LLM makes your interaction with a platform better, but like, okay, what's the underlying platform for all this, like exploding AI workload workloads and workflows? Um, you know, how does that help?
And this is where this like product mindset, right? Um, is essential because it's like, okay, how do we take these new tools and, you know, and these new capabilities that we wanna provide to developers and package them into something that actually works and that actually works. And the enterprise means that it's like secure that is, you know, that has like governance built in that is, you know, um, that drives standardization to your point and automation by design.
And I think like there's too much focus right now in this space on like, okay, how do you know this thing can like, automate all this other stuff, you know? And, and this can like speed us up, you know, but the, you know, speed is, you know, there, there's, I think there's this like, um, I don't dunno why it just popped in my head. There's like this commercial I think from like Bridgestone, you know, the, the tires, they're like, power is nothing without control.
And it's like, it's the same thing, right? It's like this, you know, you need the, you need like really solid wrapper, uh, for, for this like AI thing to go fast. Otherwise, otherwise speed kills.
Yeah. Right? Speed Kills.
But, But guys, I'm sorry. Go ahead, Keith. You're making a really great point.
The, and I don't think, and I think I accept your pushback on the premise of what the problem is. And if we look at some of the things that platform engineering has solved for enterprises on the traditional application side, there's very much a pattern that develops that we're, you know, that's very high risk when it comes to ai. So most, most enterprises don't have product groups.
Like, and if it's, it's a failing, right? This is why Enterprise cloud keeps failing. You need someone to actually manage it as if it's a product.
And there's no product management group. But there's this at this side effect, I don't think we expected with platform engineering that we kinda get in built in lifecycle management with platform engineering. Not exactly everything that we need, but some attributes of it.
And one of those attributes is, you know, API maintenance, uh, and when I upgrade the what, what one of the things that we're going to learn with ai, even the models that we consume that we built are going to keep moving. 'cause we're going to reuse those models over and over again long past when the applications are past their prime. And we're no longer developing those applications, actually developing those applications.
That's where we see, you know, we see it today when, you know, we change your API and the old application breaks. Someone has to go and fix it. Well, what happens when the model changes in a way that we can't predict and we're no longer monitoring the application in the ways that we, you know, we're not putting a human in the loop to monitor the application.
How do we do that? That's the platform engineering problem. That's not necessarily a developer problem.
That is a platform engineering problem. Absolutely. And I think it's really comes down to, you know, how do you design this like pipelines, right?
And, and of course, like all of this stuff, I think is gonna be radically different in like two or three years from now in ways that we can't really predict. Um, but like if I look at the kind of the situation on the field today, um, I was like, some of the most inspiring conversations that I've had beyond platform con of course were, um, at, um, at, at Google, uh, cloud next this year. Um, 'cause you could really feel like there were like every, you know, everything was ai, everybody was coming in with like a lot of energy and, and around this.
And there were, and, and it was clearly like the range was so broad, right? Like, it was like the vast majority of people had no idea. They were just like there to like listen and try and figure out, there were a few people that were like experimental some stuff.
And then I was really, really impressed because like going into that conference, my, my gut feeling was just like, okay, everybody's in to do groups. But then I found, well, there's actually like a third group of people that are already pushing stuff to production. Um, now of course, like to your point, Keith, it is like they can't really like reduce the, the error rate to zero.
But it was like, you know, like a few standard deviations already, like in a way where it was actually like, you know, upper and, and, and it was, and it was basically done in a very, you know, simple way actually in a, in a sense where we're just like chaining like model after model after model, right? And like constantly like, you know, so that even if like one, like massively hallucinated or like if anything went wrong in the process, you know, there's just so many checks and balances essentially, right? That, um, that you would, you know, that you would basically end up with like an output that is actually to some extent almost production ready in a way, or, or, or another.
And this actually, and you know, this might be a segue to another interesting thing to explore, Alan, I'm not sure, but like, I was talking to somebody that was framing, that was reframing, I think how, you know, he was like, look, like in the last like 10, 15, 20 years, like the, you know, if you look at like developers, like, you know, the, the more, you know, the, the, the, you know, QA wasn't necessarily always like the most kind of like, Hey, that's where you start, right? Like, it was all more like a, like, okay, what's the necessary thing that we need to have is a, is a check, you know, that we need to have and so on in place. But actually the people like really innovating, creating new code, creating new features are somewhere or in another team.
And, and I think what's interesting is actually when you think about where we're going in the setup, like QA or whatever it's gonna be called, right? 'cause it's gonna be, you know, like AI engineering or whatever, it's actually like in the enter, you know, on the individual level is prompt engineering on the enterprise level, it's really like, you know, like, how do I make sure that like all these things like don't hallucinate, you know, and, and they're actually usable in production, right? And so it's very interesting because all of a sudden this like QA that was, you know, I'm not saying like a secondary figure, but not the primary figure, I think becomes in a, you know, non-deterministic, more probabilistic world actually, the figure of reference, right?
For how you make this stuff enterprise ready, The big dog. So I, I gotta jump in guys 'cause I haven't gotten worded. So I, I think to a certain extent we're ignoring the elephant in the room here.
And that elephant is, is that a lot of these AI platforms, if we can use that word by calling you, you know me a gentleman, but you know, a lot of these AI platforms, they weren't really designed for the developer, for the platform engineer to use. They were designed for the data scientists, for the model trainers, for the, for the people who were, you know, the initial workers or the initial audience for these ai uh, applications. org community.
Yeah, there's a lot of people who have platform engineer is their title, but there's a lot of people who have data scientists in their title. org community. Why?
Well, the it, it's obvious why, right? They, they understand. And so, you know, the a lesson I learned in a lot of the startups i, I helped start was un you know, understand who your customer is, understand who your personas are, understand what your building for who.
And so you can't say, I built something for these people and now I'm going to co-opt it for these people. Sometimes you can, but it usually takes a lot of re-engineering, a lot of rejigging, sometimes just redesigning. And, and I think that's the cycle we're in right now, maybe is trying to take something that was built for these people and make it work for this crowd.
Keith, what do you think? Yeah, yeah. So you, Alan, you're, you, uh, you're hearing on one of the first questions I ask anytime I create content is who is this for the, at the end of the day, who is this for?
So as we're, you know, as we, as I put on my CTO advisor hat and Luca hit a super key point that I, I want to go down this rabbit hole a little bit, that AI as an enterprise ready, and he used, uh, uh, AI assisted cold as an, as an example. com com on this Yeah. Topic, which is, uh, where, how do you scale this?
Like the, I've talked to individual contributors at AWS, they had a really great conversation with a principal engineer at AWS, that's a big time title. And this is, you know, when you're talking about the most senior of engineers, this is, this is one of the big boys. This is a fame, uh, developer who's getting paid, you know, probably a million and a half dollars a year to be productive.
And they were exceptionally capable with, uh, uh, uh, uh, AI assistant to 10 x their productivity. And the question I asked them was, how do I spread that around? How do I expand that into the enterprise?
How do I scale that? And that is the, uh, essential problem that we're seeing, and this is why these personas are joining organizations like this, because they're naturally coming to, is similar to when, uh, uh, cloud native first came around and they said, oh, we're gonna show the enterprise how to scale applications. Oh, okay.
Thank you very much. We've never, we've never scaled applications before. We don't know what we're doing.
But they soon discovered that they were solving the same problems that we had already solved, uh, time in and time in again. So, uh, Luca, I think you really hit on a keynote, uh, here that scale breaks everything. It breaks, uh, platform things that should go to platform engineering.
Yeah, yeah, absolutely. And, and, and, um, and I think like where we, I, I don't think it's the answer, but I think like where we need to start here is to your point, Alan, is like, how do we bridge this gap? Because yes, like the fastest growing segment in the community is data engineers and then, you know, security people.
And now of course there's gonna be all sorts of like AI titles coming in. Um, but the, you know, people are trying to figure out, okay, how do I, you know, I need to interface myself increasingly. So with the, with the platform engineering or with the platform team, you know, there is this like platform, so how do I leverage it for data?
And to your point, Alan, right? Like the, the, the, the, the issue thing is, you know, right now these two, you know, these are two completely separate siloed worlds, right? There is this kinda like SDLC platform engineers, how do I deploy workloads with their dependencies and so on.
And then on the other side, there is, you know, like datas and like data, data, data like data lakes and models and like, you know, yeah. Like the, and, and, and you know, the user over here is application developers. The user over there is like, you know, ops, uh, uh, ML ops engineers and like, uh, data scientists and, and whatever, right?
Um, and so like one of the, I think like interesting exercises that we've started doing in the, in the community is actually, uh, trying to map out what does a joint reference architecture for a data platform, um, kind of look like, right? So like a platform that can servee both these users. And, you know, to be honest with you there, I don't think there is an answer there yet.
There are a lot of questions, but I think we're starting to at least pin down, okay, what are the key things that we need to answer to figure this out? Um, and, and, and kind of put in the different, you know, uh, puzzle pieces together. Um, because, uh, otherwise, um, you know, we're gonna, and, and, and I think like, you know, I've, I've always, um, used as a guiding principle, you know, you mentioned Keith, like, you know, where do you start when you create content?
Like, the way I think about one of the, one of the, uh, key things that I think about when creating content is, is is this sort of like, okay, what can I learn, um, from sort of like top performing organizations that I see is, is gonna trickle down to the rest over time, right? Um, because it's, you know, it's, it's, it's that quote of like, the, you know, the future is already here. It's just unable and distributed.
It's like, it's, that's always true across, you know, any sector, any like, new technology. Like it just gets earlier somewhere, um, and you just need to like go and listen there, right? And so, and so I think like if you look at top performing organizations, they're already thinking of this as just one unified layer because of course, why wouldn't you, right?
Like, I mean, apart from engineering, I think it's just like, yeah, another permutation of what we've been doing in software engineering for all the last decades, which is just like progressively abstract, like one level higher, right? And the thing is right now, you know, you have this like two levels, they're kind of like next to each other, but, and so the next level is just like one thing that sort of like pulls them both into, into, into one platform layer, which again, is, is I what I wanted to say earlier. And I, I, you know, I i, I always end up in tangents, but what I wanted to say earlier is, um, you know, I'm really excited about this whole AI thing because I think like while as a trend, it is overshadowing and overpowering a lot of outer stuff in the industry and, you know, industry meaning like cloud native, whatever, right?
Like enterprise, like I actually think it's, it's really powering platform engineering because you need this, you need to figure out this platform, um, thing, uh, to to, to make everything else possible. Fair, fair guys, we got maybe time for one more round of comments and we gotta close it out. Keith, what, what's your, what's your summation here?
Yeah, so we need to give practical advice, right? The, at the end of the day, the people who listen to this need to understand, well, okay, I hear that the big boys aren't taking care of my needs. What should I be doing?
And I think that's where the last piece of the stuff at is one, you know, our approach, this like any other engineering solution, measure what I can measure, understand where the gaps exist, and this is going to be unique for every, or organization. What do you need as you're building AI ops? So you need to understand the problem in the gap in which you have, then you start going down the, uh, solutioning of talking to your vendors and, and making sure, validating what we are saying, you know, from a high level we talk to the vendors, but, uh, I'm not a practitioner anymore, so I, you, you really do need to talk to the vendors and understand where they're offering potential solutions or understanding what their roadmaps are.
'cause we haven't talked about the vendor's roadmap when it comes to serving the platform engineering audits, there's going to be gaps because this is not where they're focused on. They're focused right now on this building the biggest, uh, clusters, the most through throughput. They can, the best models that they can build and that they're cap, they're focused on building the capability around the infrastructure and the compute, and not necessarily operations and deployment and management, et cetera.
So there's going to be gaps, understand those gaps and then reuse some of the patterns that you've already have. You might already have solutions for the most basic of things. Why predict that you're going to find the most problem is around this whole lifecycle management piece, because we haven't gotten to that, right?
We, we haven't gotten to what happens when an AI app has passed this useful development lifecycle. Uh, we're pretty, we're very much too early. I can't predict that, and I don't think any of us can predict that, but we need to be thinking about the problem.
Fair enough. I'm gonna give you the last word and I Just, yeah, I would just gobble down on what Keith said. I think like the, the TLG is like, nobody's coming to save you.
Like, You know, you, you need to, and which I think goes nicely like full circle with where we started. Like, you know, this big, uh, cloud providers, the OEMs, like nobody's is building a solution for you. Nobody ever had.
Um, it's like platform engineering is about like taking whatever it's out there. So don't go and reinvent a wheel from scratch so you don't have to build entire stack leverage, whatever this providers give you. But then you need to put in the work to like really make it sensible for what your use cases are.
Um, nobody's gonna do that for you. And if you don't do it, you know, to Keith, to Keith's point like this, this stuff is not gonna be usable. It's not gonna be enterprise ready.
Um, and you're gonna fall behind. Excellent. Gentlemen, thank you both for coming on.
This was a great platform engineering show podcast, Luca, I have, I'll see you in two weeks. Keith, for people who wanna follow you and, and grab more, what's their, what, where, where do we point 'em? com is the content machine.
That's where, that's where I'm creating content. Go back to the blog. It is where I always create content.
The advisor bench. If you want the that next level help, that's where the hit the contact us on the c advisor, uh, dot com. I love it.
All right, Luca, Keith, thank you. Thank you. By the way, we didn't mention thank you very much for Check Marks sponsoring our podcast engineering show.
So kudos to them and thank them. Uh, we'll be back in two weeks with another show. I think we probably have a live round table coming up soon too, on this.
But until then, this is Alan Shimmel for the podcast Engineering show. Thanks everyone. Have a great day.