79. Your Edge Projects will Fail Without Fleet Lifecycle Management with ZEDEDA – Tech Field Day Podcast Spotlight
Projects to deliver applications to edge locations will fail without comprehensive fleet lifecycle management. This episode of the Tech Field Day podcast features Sachin Vasudeva from Zededa discussing the importance of long-term edge management with Guy Currier and Alastair Cooke. There are unique challenges of managing edge deployments compared to cloud or on-premises environments. Focusing on business logic and application outputs while leveraging infrastructure providers to handle the complexities of packaging, deploying, and monitoring AI models enables diverse edge environments. Edge locations might have different hardware deployed, intermittent connectivity, requiring a balance between standardization and flexibility in managing edge devices and applications. Teams with rapid responsiveness and adaptation will better enable their business to respond to changing conditions, especially with the rapid pace of AI innovation.
Transcript
Your Edge Fleet, your Edge project, you need to make this successful because you're investing a lot of money in it, not just today, but tomorrow, next year, and probably five years from now. Join us on this special episode of the Tech Field Day podcast as we talk with, uh, session from Zaida and also Guy Coer. Welcome to the Tech Field Day podcast, where we bring together a group of IT technical experts to discuss a single idea around key concepts in the industry.
This podcast features a variety of perspectives from members of the Tech Field Day delegate community, and it's often recorded in association with one of our events. Tech Field Day is part of the Futurum Group, and this podcast is also published on our sister company site Text Hong tv. On this episode, presented by Zaida, we'll be discussing the premise that without Fleet Lifecycle Management, your edge projects will fail.
Before we start the discussion, let's meet who's on the panel today. Uh, hi, I'm Guy, er, I'm an analyst at the Futureum Group. Um, ai, AI infrastructure edge, um, as well as, uh, some of the things that fuel, fuel, those like chip sets, containers, and so forth.
Those are my areas of coverage. Hey guys, I'm Sachin Vasudeva. I am the VP of product here at Zita, and I am responsible for driving the product strategy pricing as well as driving the portfolio around Edge AI as well as management and orchestration solutions.
And of course, I'm Alistair Cook. I'm an event lead here at Tech Field today. And what to look at the kind of context of edge deployments and some of the things that make edge deployments quite different to how we might sort of view things.
Sometimes we view, uh, edge deployments as an, an edge projects as being a, a cross between being a, a cloud platform and being, uh, like managing a fleet of laptops. And while it has elements of both of these, its own quite unique complexion, there's projects to deploy, uh, intelligence and, and applications out to the edge are very different to giving a a, a group of traveling salespeople their laptops or using cloud services. We have near infinite resources available in front of us.
Um, one of the other elements on this is that that edge deployments tend to be characterized by a huge number of very small deployments. We end up with devices that every branch or attached to every oil well or to every, uh, ship that we have traveling around the world. And that, that, that's, um, lots of small deployments is very different from dealing with clouds that are, that are massive.
And so things that sort of make sense in the public cloud and makes sense if you, you have a, a small number of laptops that are in the hands of skilled and experienced staff don't locations with potentially staff who really don't care about those devices at all. And, uh, but makes things quite difficult as we scale and deploy these applications, and particularly as we're running these edge locations for longer periods of time, as day two, operations start to dominate. What we're actually doing, uh, moves, ads, changes to the application setting.
These things that are out at the edge, uh, section is you are working with customers, particularly those who are adding functionality like the AI that you are focusing on. Um, what are the sort of challenges as they, they hitting with managing their fleet of locations as changes have come, come in, Right? Um, there's actually quite a few challenges that we are seeing, uh, especially with, uh, AI coming more and more, uh, apparent and, and, and more organized and more structured.
First, uh, and foremost, you know, so fleet management is about managing total cost of ownership of your devices, the lifecycle. And, uh, that includes also the, uh, not just updating the software stack, but also updating the applications that run on these devices. So the, with ai, what's going on is the, the applications are iterating faster.
They are AI models that are coming around, uh, with, with newer technology, uh, you know, interfaces, um, without even talking about gen ai, if we take the classic machine learning models, they are, uh, now starting to be fully commoditized. And so we are seeing the, uh, challenge, uh, of taking and packaging these models in a way that they're lightweight and they're easily deployable. Second, uh, the version management of these models, uh, become super important across the fleet.
Uh, you might have different versions deployed for different application types because one version sort of exists and continues to operate as it was trained, but another version might be required because, um, you know, uh, the data has changed for a particular region, for a particular site. So within the fleet, there are, uh, you know, there are, um, parts of the fleet that will require, uh, you know, the, the sort of, uh, model management, uh, to be different versions. Uh, so packaging will change accordingly.
Uh, uh, the, the ability to then coming back to like back to lifecycle is automation, right? So how do we easily drive, uh, uh, changes that are maybe driven by the IT team for security purposes, but the OT team, uh, does not want to apply the change right away. So there's a, a, a, a management, uh, for the edge when it comes to feed management to actually hybridize this IT OT bridge and make sure that, uh, there's a right level of approval that's built into this fleet orchestration solution as well.
I think that, um, when I think about, I mean, this is really good insight, uh, sien because when I think about, um, let's take the word fleet out for just a second. When you think about managing, um, edge applications to start with, there are at least three basic dimensions to this because there's the management of the physical device and, and its physical environment. Um, there is management of, uh, the payload, um, and then there's, uh, a management of, uh, uh, communications, uh, whether it's, uh, it could be a data collection device or it could be a data utilization device, um, uh, you know, a drone or something like that, a data utilization just meaning, uh, the application performing, uh, functionality of some sort in field.
But when you add AI into this, um, that creates an added level of complexity and, and variation. The variation of, I mean, you talked about machine learning. There's a certain amount that can and must happen for disconnected devices in the field.
There's data collection and training back in the core that's required. You might wanna do pre-processing for that. So, so, um, the, uh, the relationship of the ultimate application to the ai, um, is an added area of management.
And it was helpful to me for you to open up this idea of the provisioning and the payload and the application and versioning position in the lifecycle. That also all needs to be considered part of fleet management as well. And I, and I suppose, and this might be my question back to you, Sachin.
Um, I suppose that, you know, part of the game here, uh, for customers is to understand where they need to focus. Because you, you, you can't focus on six things. You, you need to focus on, on the most important elements of that, that that management part of it that comes after development when you're on day one and beyond.
Is that, is that right? And what sort of shape does that focus take? How do you advise your customers where to focus?
That's a, uh, that's a good segue, right? So back to sort of, um, let's, you know, when, when talk about fleet and lifecycle, right? All these, uh, um, words are sort of interlinked.
Let's talk about the lifecycle and the anatomy of an AI application, right? Let's start there and then, and then we can talk about the focus that we need customers to sort of really put energy on. So the first and foremost, um, you mentioned, right, there's a, there's a training that happens.
Uh, there's data that's involved, right? All of that can happen in a centralized location where all of the centralized data lake, uh, is connecting all this info. Um, and once this training has occurred, however, there is a, a need for the customer to take this package, it in a way that it's lightweight, so it fits the edge form factor.
We might be deploying this on a small, sort of, uh, Nvidia Jetson based platform, which has, uh, you know, limited, uh, resources, or we might be deploying this on a high-end GPU AI server, which has a, you know, lot of resources. So, so there has a, there has to be a deterministic or there's a determination on the resource at the edge. And, and ideally you don't want to think about it at the time of, you know, thinking of deploying this.
So ideally you want to sort of push that decision to somebody else, like where does it go in a consistent way? And if the resource is not available, what happens then? Uh, then if you continue on the lifecycle, right?
Once you move into a deployment, then there's monitoring, which is, uh, observability, you know, how is my application performing? How is this AI model accuracy? You know, is it up or down?
Is it data drift? Uh, to your point, right? Data collection has to occur.
Uh, typically that, um, ecosystem is driven by, uh, data collection, not at the edge, but data collection all the way making data, making its way all the way back into the cloud. So there's an office cost to pushing data all the way back and then figuring out if the model drift occurred or not. So, uh, let's, uh, take this apart, right?
So on the deployment side, uh, the, um, identifying the target, identifying the right environment, and then packaging it to the right form factor, and then monitoring means, um, maybe collecting the data at the edge and then having, you know, something to do with retraining the model at the edge. Uh, those become sort of really sort of big drivers for driving this. So now, um, what we are starting to talk about with customers is their focus needs to be on the model itself.
I mean, they have the domain expertise on the business logic that gets attached to the model. The model is doing its bit, which is the inference part, but once the output of the inference comes in, right, they have to bring, bring some post-processing logic that helps, uh, you know, conclude to certain business decisions that happen either at the edge or happen maybe centrally. And so, let's, uh, help you, Mr.
Customer focus on solving those problems within the application and take the rest of the application challenges around packaging, deploying, fitting it to the right form factor, observing or collecting data, uh, leave it to, you know, an infrastructure provider to help you get all that info, uh, so that you can actually come back and then decide whether your model actually worked properly, whether you are able to glean the right information to make the right business decision, whether you can automate around it deterministically, or whether you need to continue to sort of iterate and, uh, and continue to, you know, uh, what we call test, uh, and, and deploy and it trade, right? So, and, and we know with ai, uh, iteration and testing, and, you know, what we call AB testing is, is the norm nowadays, right? Because a single model doesn't fit all use cases, whether it's overfitted or unfitted.
Uh, but, and so, so you know, these, um, so focus on the business logic, focus on the out outputs of these models and let us take care of the rest. I mean, it starts to get really interesting when you're thinking about variation, um, based on local conditions and, and learning and, and, and maybe some sort of, you know, automated or, or, or not fully automated iteration inversing. I mean, that's like the mother of all forks.
I mean, uh, that's not necessarily something that, uh, that, that it's, it's the sort of thing that can give a customer pause. But it sounds like Aida's point of view here is that there are certain elements of that that are more technical than domain related. If you're an ag or if you are in energy, um, you wanna focus on, um, you know, uh, applying your, what you call the business business, I think, business rules.
But it's really about how you're operating your company, what your go-to-market model is in energy, in ag, um, and, um, certainly focus AI development, data management and all that sort of thing on that domain. And Aida's point of view is there's, there are technical elements of this that can and should be handled by provider tool software. Yeah.
Is that right? That's absolutely correct. I think one of the things I wanted to hit on was quite a lot of what you've talked about s is true for people who are doing AI deployments that are not edge deployments, that are doing this in, in their own data centers as well.
A lot of that, that same process goes on for a deployment in your own data center, but that's quite different from where the actual business logic is being applied. Hundreds of locations spread around the world, and that decisions need to be made locally inside each of those locations for speed and cost reasons. Uh, and particularly one of the things we see is the intermittent connection of some of these locations.
And that adds another dimension that maybe the, the model that we believe is our, our best, most accurate model is only in 90 of our a hundred locations, because 10 of them have been offline during that update, and maybe two of those were offline in the last update as well. And this is that, that perspective on fleet management and wanting to have policy-based management across it rather than having to manage every single site individually. And I think that's really where you're talking about handing off this responsibility for the, essentially providing a, a platform and an infrastructure layer in a consistent and predictable way.
I think that that difference between doing this in your own data center to doing it across lots of potentially intermittently connected low power locations is what's really different in edge deployments versus cloud or on premises. Yeah, that's a, that's a great point. And, um, two words come to mind when, you know, we talk about these two items where you start off in this centralized mode of operations, and then you're trying to really deploy something at the edge, which has different constraints.
And you may not have tested for it, you may not have thought through all the implications of it. Um, so quantization and optimization, right? So, so we believe that, um, that is a significant effort, uh, in, especially as, as the edge as you're aware, there are a variety and diversity of platforms out there, right?
So, and, and we know there's a constant sort of, um, uptick on the number of, uh, systems or platforms or chips we are seeing now in the market. So, you know, for example, Nvidia is pushing, uh, the Jetson infrastructure and the Jetson chip, uh, platform, and it's actually revving, its super fast. So, you know, within, uh, a couple of years we've already seen, like within this year, we've already seen two versions, two variations come up.
And, and the last one, the Jetson Thor was based on the Blackwell GPU, but that's not alone, right? I mean, you've got Qualcomm with, with, uh, with their NPUs, and then you've got custom chip makers, um, um, like Halo and you know, with TPUs, and then you've got the hyperscalers bringing their own flavors off the chips that they're also building, which they might wanna, and also push towards the edge. So we are seeing a plethora of all these systems.
And, and so when you start thinking about quantization and optimization, you have to think about the common, common, you know, elements that exist from an infrastructure point of view to help you deploy this in a consistent way without worrying about, um, uh, you know, getting, uh, a hit on your accuracy or hit on your performance. Now, um, it is a tough problem. It's not an easy problem, but, you know, we love to like try to like address this in a, in a way that we can address it, uh, one step at a time and, you know, ensure that we can actually move towards some level of standardization.
Uh, there's a lot of open source efforts going on towards taking the work that, uh, is being done, um, for one specific chip set maker, and then bringing it as a generic, uh, you know, approach to inference, for example. And inference engines are, are sort of, you know, um, are taking on that the open source community is taking on that effort as well. I wonder though, uh, if I, I rather feel like this is an issue that's not talked about much in edge.
Um, and maybe it's because it's a non-issue, but that's diversity of, of platform, uh, meaning hardware platform, meaning edge device platform. My, my general sense from talking with customers and, and especially uh, application managers over the years, um, 'cause this has been going on for a long time, including AI at the edge going on for quite a long time, is that, uh, any one application actually can have several different types of edge devices. Um, it's not just a question of architecture, it's a question of size, rugged, non rugged.
Is it a, does it have a human interface? Does it not? Is it collecting, is it, in other words, is it, is it mobile or is it fixed?
You know, what kind of connectivity that single applications have? Diverse hosting environments, and, you know, I'm sorry, but you know, I don't, I don't see K two s as really solving this. I don't see anything solving this other than, uh, a way to, I, I'm gonna use the word centralized, but I just mean sort of in a virtual sense, centralizing developments so that there is a way to develop one application, but that can have its various functionalities deployed into diverse, um, you know, onto diverse hosts.
So, so I, I just wonder if, you know, from, from the analyst perch that I sit on, if that's just me speculating or if that's actually a reality, and if so, why isn't it talked about? Isn't, so it's the same idea that, you know, um, a lot of companies in the networking space talk about when they built the hardware, ab abstraction layers, right? The hal, so, so you're right.
Uh, at, at some point, you know, when the rubber meets road and you're actually trying to bring a HAL that actually works on a specific platform or device, right? Uh, you've gotta make sure that it actually works, end, works end to end. So it's the same concept for us, right?
We, we will handpick a few that are most predominant in the industry and start with them. And then as, as adoption occurs across, you know, other, uh, players and emergent or existing players coming in, uh, and, and so customer driven, right? Adoption, then, you know, we'll start to address it as part of, uh, this, uh, this notion of a hell or this notion of an abstraction interface, which, uh, starts to abstract it.
What we are, what you're correct about is that, uh, we would have to pick and choose. Sometimes, you know, sometimes it's not feasible, you know, the portability is a, is, is a, is a desired sort of, and consistency of workflow is a desired end state. But, you know, for some applications, we mean never achieve that.
So especially like where we, we are seeing a very diverse, uh, set of, uh, applications, but I think our, the vendors we're working with are also trying to address that. So HMI related, like where you have touch screens, uh, the apps are gonna be very different. There's gonna be a lot of sort of, you know, um, um, interfaces, uh, that built in around the touchscreen and the actions, and then versus, you know, an, an industrial pc, which is not, uh, you know, looking for that interactive mode of operation.
So, uh, we do, you know, it's, it comes back to the anatomy of, uh, of the application itself and what you're trying to do. Um, you know, and if you break it down into a series of microservices with accompanying models and accompanying data for managing drift, uh, we can argue, yeah, we, you know, how we then package this, uh, whether there's a UI or not UI component or not becomes an interesting conversation to your point. Um, but yeah, I mean, the, the, the model will certainly have a dependency and the runtime of the model will have a dependency on the type of hardware you deploy.
That that is, Yeah, that's, that's the real, I mean, Alistair, you, you, you've lived that, that it's the, these, these reference platforms and everything, they just sound really great when you're in buy mode. But then when you're in use mode, that's when you know you have the company breathing down your neck and, and yeah, know setting off your beeper in the middle of the night, right? There's, there's also the element that your edge, each deployment is not gonna be a written replace everything when there is an update that you'll end up as, as you see with the diversity of hardware, even if you have multiple sites that have exactly the same requirements as the same application set, you may have one site that has hardware that's six months old, another site that has hardware that's four and a half years old, and managing that diversity, making sure that we're not pushing down a model to that old hardware that is so, uh, heavily quantized that it'll fit on the hardware, but no longer produces a useful result.
There's absolutely some challenges around fleet management across that diversity and, and then flagging back to the hardware lifecycle as well as to the application. And the, the infrastructure underneath the lifecycle of these things are all tied together and driven changes driven by the business need. But there's a very high change for, uh, high cost for changing the hardware at all of these edge locations compared to trying to shoehorn whatever application we can get into the hardware that's already there.
Uh, this is what leads us to four and a half year old hardware sites, because it's sitting out in a, in Alaskan mine, uh, and it's just monitoring as, as one of the, the examples I've seen is monitoring a conveyor belt to see if one of the staff has, uh, gotten onto the conveyor belt because they're maybe, hmm, not in good working condition, uh, the staff that is not the conveyor belt. So yeah, that, that diversity is a challenge That that is exactly right. So, uh, you know, there's a promise of, of an abstraction and a reference, uh, you know, uh, platform.
But the reality is there's diversity to manage, and it becomes half process and half technology. So, you know, so yeah, one of our customers, uh, you know, has more than 18 vendors. I, I don't think, I don't think customers are asking for this problem to be eliminated, but that's what vendors are offering.
We have this magic potion that's gonna eliminate this problem. I think customers are looking for a way to rapidly and relatively simply respond, be responsive to these scenarios and situations. A node goes out, you need to be able to do something about it, um, in a reasonable way because, you know, as they say, stuff happens.
Yeah. So, so I would argue that, um, for ai, right? It's wild, wild west.
So, you know, and so there is some level of standardization that can happen. Uh, for example, you know, the inference engines, uh, out there like open vio, infra, you know, uh, the, uh, your, uh, tensor RT or, you know, like the new Dynamo or, um, owner nx, right? So, so, so we will see some level of settling and, and our customers asking us, well, tell us upfront if you are comfortable driving on nx, then we'll align to on NX as a format for packaging.
And so there is that level of conversation also happening at the same time. So, so I, you know, while, while there is no magical cure for, for the diversity of hardware, there's definitely a, a a, a runtime that can be, uh, you know, that the format, uh, for example, could be agreed upon. It could be an agreed upon.
It doesn't mean it's, uh, it's an industry-wide standard that's adopted, but, you know, adoption and sort of this goes hand in hand. So, so I think if you can agree upon one or two or a couple versus, uh, bring a lot of flexibility into your fleet management, then, you know, then, then you can really sort of own and manage it. I think it really frees the, um, the, the, the, um, let's say the device management team too, to respond to, uh, more quickly to, uh, new opportunities, new, not just new applications.
I'm thinking, you know, uh, um, you know, entering new fields, new operational zones and areas, whatever it is, I'm, I'm losing my lingo here because I don't have that domain expertise necessarily, but it allows them to, um, to be more creative. A as always on the Tech Field Day podcast, our guests could continue to discuss and learn from one another for hours, possibly days. In fact, that's probably why we have longer events at Tech Field Day.
But I'd like to thank you all for joining us today at the Tech Field Day podcast. com. Um, and, uh, as I attend events and conferences or do, um, uh, shows like this one I tend to post in LinkedIn, that's a good place.
And I'm also at Blue Sky at Guy Courier, blue Sky, whatever. It's, You can find me at, uh, you know, on the, uh, za do com website, um, where you can reach out to the zaa team as well as on LinkedIn. Um, so yeah, happy to provide more, uh, more insights as needed and looking to engage.
Yeah, And of course, you can find me a cook. Uh, do make sure to check out the previous presentations that we've seen from zaida. You'll find them all on the Tech Field Day website.
If you just, uh, go to Tech Field Day slash company slash zaida, you'll find all of the great, uh, presentations from, from Zaida. And so thank you for listening to this episode of the Tech Field Day podcast, uh, showcasing Adidas Edge expertise. If you've enjoyed this discussion, subscribe on YouTube or your favorite podcast application, so don't miss a single episode.
Uh, do also remember to give us a rating, a very positive rating, and a nice review. This podcast was brought to you by zaida and Tech Field Day, the home of IT experts from across the enterprise and a part of the RUM Group. For upcoming events and more episodes, head to Tech podcast or view us on text on tv.
Thanks for listening, and we'll see you next week.