ZEDEDA Edge AI – Object Recognition Use Case
In this ZEDEDA Edge Field Day Showcase, Sergio Santos, Account Solutions Architect shows how ZEDEDA manages edge AI for a practical object recognition use case, specifically for computer vision. His presentation shows how to deploy a stack of three applications—an AI inference container, a Prometheus database, and a Grafana dashboard—using the Docker Compose runtime across a fleet of three devices, one equipped with a GPU and two without. The demo highlights the ability to deploy and manage applications at scale from a single control plane, leveraging ZEDEDA’s automated deployment policies. The process starts from a clean slate, moves through provisioning the edge nodes, and automatically pushes the application stack based on predefined policies, including GPU-specific logic.
A key part of the demonstration is the live update and rollback process. Santos shows how to remotely update the inference container to a new version and then roll it back to the original without restarting the runtime. This highlights ZEDEDA’s lightweight, efficient updates and the use of its Zix infrastructure to push configuration changes. The demo also shows the ability to monitor application logs and device metrics (CPU, memory, network traffic) from the central ZEDEDA controller, proving the platform’s comprehensive management capabilities. The session concludes by demonstrating how to easily wipe the entire application stack by simply moving the edge nodes to a different project.
Presenter:
Sérgio Santos: https://www.linkedin.com/in/cshari/
Delegates:
Alastair Cooke: x.com/DemitasseNZ
Ned Bellavance: x.com/Ned1313
Josh Warcop: x.com/Warcop
Guy Currier: x.com/GuyCurriersFeed
Transcript
Welcome. So my name is SE Santo. I'm, um, a solution architect, uh, a solution architect, uh, at, so I'm based in, uh, Southwest Europe in Liman.
And I'm here to introduce your packet as a practical example of, uh, HI for object recognition. Use, use case that you can use in different market market, uh, verticals, retail shops who count number of people in the cashier, or you can use like in a, in the, in the, in the oil and gas industry or in maritime, we can detect protection equipment. So this is basically a generic, um, object punition use case that we are going to show how to deploy in this case, using the docker composed, uh, infra hand time infrastructure, and also, uh, deployed at scale in a fleet of, uh, three devices.
So the demo workflow that we, we prepared, uh, here was essentially composed by, uh, three applications. So the AI piece, so the container that runs the AI France pipeline together with the business logic. So together with the, so we're just doing the single container, so we don't, we didn't split this in, uh, in two.
So we have this one single, uh, container inference that in this case, this container can run on A GPU or without a GPU. So it's, it basically, we, we made the, the container available in a hand time in the docker, uh, compose hand time. And in that hand time, we also make the available the NPDF could the, uh, drivers.
So that's in case those drivers are there, in case there's the GPU in that device, the container can actually consume, uh, those who occurs, right? So in this case, the, the inference pipeline is actually, uh, using, uh, open source model from, uh, from meta from Facebook. The Snet, uh, model that, uh, is used genetically to detect different, uh, kinds of objects in the pictures, and is leveraging also the PyTorch, uh, framework for image transformation calculations and actually outputting them the result of the inference.
And then the, the container itself will expose that, uh, that information to business logic. And in this case, locally, you, you want to also to keep some, uh, some that data persistent in the device. So we can imagine that this could be a, a ship that doesn't have internet access or actually information is only relevant local in case of protection equipment.
You need locally dashboard to alert captain of the ship that there's a, there's an author or there's the problem. So we somehow need to keep database on this case. We brought, uh, promeus at database that will keep that inference data, uh, locally in a time series.
And you also have a local visualization to, to actually present that data in a, in a dashboard. So that's why you see these three applications, uh, and they are actually available in this marketplace. Uh, and then from that marketplace, basically we can, uh, deploy in a flip of devices, we have three node devices, one with the GP we kit, the other one, the other two they don't have.
And the flow of them is essentially basically go from zero from this zero to day two plus. So where we start with a, uh, without any edge node created in our platforms, we'll deploy the edge node, uh, and then we'll trigger the deployment of this application stack, right on those, uh, three edge nodes using, uh, automated deployment policy, right? And, and then at the end, we'll, we'll update basically this interest container so that you may, we will, we'll push the new version of that container, then we'll do it, uh, on the 10 time without actually starting the, the hand time, right?
So, so technically, um, when this is looking, doing a zoom in, in the, in the edge node. So we, we pick that first edge node that we have a GP, so a Discre GU, um, available. So when we deploy this, uh, ai, uh, inference from container, our, our inference pipeline running on, um, on the docker hand time with the NP with the drivers.
So it'll basically expose, it'll actually connect to the camera feed. So it's actually image processing, computer vision. So you'll basically connect to the camera feed.
So the pointer together there is the configuration perimeter in our docker compose yamo. And then the user can also get the web interface that is exposed, or the business logic or the client's logic is also exposed in that, that same container. So we'll get a picture like this, so we can see in the video stream, in this case from New York City, we can see the, the objects, car person, buses, and you can have some statistics.
And then we are going, these statistics are actually live in the, in the, in the webpage. So we want basically to keep these statistics historically over time. So that's why, uh, the container also the business, the client logic of container also exposes the permit to scrapper.
So our Prometheus, uh, container running next to the 10 time is scrap lab metrics and keep it in the database, right? And then the, the refine database, we can also be used, uh, to present those, those metrics, the time series and user can also connect to our G Grafana and see g those statistics over time. Okay.
So essentially this is the, the demo how this with, how, how do we deploy it? So we have this concept of, um, deployment project or deployment policy. So we saw already te have form in action, uh, and you can actually declare all these objects part of the terraform, or you can reduce the Terraform state so that we can just declare our, declare our devices.
And you have the policy definition in the controller where is say that for this device, in this project, I want this application stack to be deployed. So we bring the notes, the notes are part of the deployment project, and automatically the controller will create all the objects that we need in that project. So if you'll basically create these three applications that we see in the picture, the, the doc, the ai, f front, the AI infancy, uh, running on the docker, uh, runtime, the Prometheus, that device and the refiner.
And next to that there are other elements. So you might need a network, uh, a network segment internally to connect to those devices. You might need other, other policy that you need to, what should remote manage those?
Device remote connected. So you see here h few policies not at the station. So all these are actually, um, hardware level or uh, or maintenance level of the, of policies for our devices, but you also bring part of this deployment policy.
Yeah, so that's underwood what is happening. So, uh, the docker, the dock hand time, that was, uh, we already explored. So here, uh, we present a bit more detail how it works.
So we have actually, uh, curated, uh, docker hand time, and we have an agent running in inside a dock docker hand time that basically can grab and pull and pushed the updates from the controller. So we have test the data, uh, agents running there, which is basically getting confusion, updates and pushing also of your state updates back to the controller, right? So, and our container basically is heart of our album charts, okay?
So we call it this ZZ so we call it the inventory configuration services. It's a generic framework that it is made, it, it's available for any application in the marketplace. Uh, you, you actually use it, uh, with a, with a specific, um, communication channel here between the, the docker hand time and it.
So here you see basically the, the docker composed yamo file, and you see in this case basically two variables pointing to our, um, to our camera feed, right? So in this case, I have two camera feed, and you also see the access, access to the kuda cars. So we are using the NNPD Puda, uh, cars available in our GPU.
Okay? And then we also inject some secrets so that we can access the, we can access, um, the hash history basically with where our container is hanging, and you have some variables here so that we can dynamically, uh, push updates in this container image. Okay?
Does, does, does Dix, um, uh, comprehend rollback as well? Is that he's Not sure? Yeah.
So basically w with this, with this, we can, we can, we can basically push updates into the pushup updates into, into the hand time. So we can just change tech or actually get saved. So we can actually update and hold back the text or the versions of these containers handing in hand time.
So it's basically just a mechanism that we have. It is a generic mechanism to propagate configuration updates into the hand time and actually push generic state right from the hand time into the control. Okay?
So you can even develop your count, uh, your own agent, right? That will follow your own, uh, your own, um, your own data structure, right, that you can use to export information back to the controller. And we have a single pane of thrust where you can see, still hold your head.
Okay, so let me jump into the, the video. So, we'll, as I mentioned, we'll start from the, from the zero. So we don't have any no edge nodes deployed, uh, in the first moment.
So we'll just go. And so this is basically default project, so where not the list of notes is empty, and then we're going to extort marketplace. So now we are going to see these three applications that, uh, that we have in the marketplace.
So we are going to see the docker compost hand time. Uh, we are going basically to, to check that, uh, we have, uh, here the, the docker compose yamo that I showed you, um, with the two camera things. Um, and then the next, and here you see the GPU, right?
So you see basically in the marketplace, you, you, you tell that this application requires a GP attachment and switch GQ attachment, and then you have another version of this app without a GPU, uh, attachment requirement. Um, and then we also have the defender, uh, container. So again, uh, we, in the, in the marketplace, we expose which ports are available, what is, what is the image pointing to, so you can just point to any registry.
And here we are going to define basically the, the deployment policy, right? So that's where you were, you're saying that, okay, we have a deployment project. There are two policies part of this deployment, one for GP nodes, another one for the A GP nodes.
And as part of this policy, they, you are saying that we need to deploy these three apps for every single node that is part of the deployment project. So this is the version for GPU attachment. So you see the applications basically for marketplace that's are, uh, part of this deployment policy, okay?
The three applications that I showed you in the, in the marketplace. So you see here that, uh, quite a GT attachment for this person of the policy, It's this unified policy in the Z to control panel and in z that that is really the, the secret source here is having that, that sort of scaling of a tool that's designed for dealing with just one site to being able to apply policies that, that use that same tool across potentially thousands of sites. Yes.
So basically we just define your policy and then we can use that across thousands of sites. So we just bring your notes right on that policy, make sure that this constant across all those sites, so in this case right now, I'm now deploying the, the devices that I have for them. So we just need to collect the devices and make them our Right with those policies.
Can you do like a phase rollout to different locations? So you don't, you're not pulling, uh, uh, you know, the same container a thousand times, uh, from all these different locations, or if something goes wrong, you can kind of do a rollback before it hits all thousand of your locations. Well, what this, this policy is applied basically to the edge locations, right?
So this will, this will be executed by the, the edge nos, uh, directly, right? So we'll monitor the policy deployment on the flip of these devices. So if there's the failure in one of those nodes, maybe we can move that node out of the, of the project, or you can move that node into a different version of that deployment policy, right?
So as you can hold back, so as you saw in deployment project, there are multiple policies, right? That they are they three, they are triggered based on the association of the nodes to the project and the tech. So if it failed in one of those nodes, we can basically move that node to the previous version of that, um, of that deployment policy, correct?
And that will hold back to the previous version. Okay? So I mentioned that you have changed with application stack.
So right now we have least three applications. You can add a new one, we'll create, you'll add a new policy with a new application, right? And then you bring, you move the notes comm policy one to policy two question, right?
So right now, what I'm doing here is just essentially, uh, declaring the notes, right? Using FF form again, so I'm just putting the nodes in the default project. Uh, so this is the variable.
We just have three nodes. Uh, so these three devices, uh, will be created under default project. So I have some flex if device GPU enabled or non GPU enabled.
I also define IP address statically. And this is basically the Terraform plan for node creation where I just essentially, I just declared the nodes. I don't declare any applications of application stack.
So all that will be brought in automatically by the controller, uh, based on this policy definition that I just showed you. Okay? So you see the tech, so this tag that we see in the node is actually matching, uh, uh, responding, uh, policy.
Okay? So when you see some, some logic in the terraform, basically if node as a GP or non GPU, it go basically attach the notes to the GP policy or non GPU policy, okay? So I just create the notes.
So the notes will show up now as, uh, proficient in the, in the controller. So what I'm going to do now, I'm, so you see I have one no of GPU in this case, and the other two, they don't have any GP attached. So what I'm doing next is just I'm using ox, so I'm using pre nodes in this case, so I'm going to bring up the nodes.
So, uh, before going that actually now going to change from the default project into the, into the deployment project. So even before ring up the nodes. So at this, we basically triggered the creation of how objects by the controller of those, in those devices, in those nodes, right?
So, so even if the nodes are not yet there, right in the field. So we don't have any, any hardware yet, uh, running. So we just basically administratively move the device from the default project into Dow our deployment project.
So they are gone from my default project, and now we, we go to the deployment project and you see the nodes out there. Um, and then if we inspect the node, you'll see now in the edge app list, we see now three applications already proficient there. So as part as, as, as we defining the policy, right?
So we see our AI inference, h gq, and now when you bring the node online, so it the register to the controller and, and then, uh, all these, the application stack that we, um, we define deployment policy will be automatically, uh, installed, uh, in these three nodes. So right now, uh, so the node is, is just doing the normal onboarding process, so using uni remote at station TLS, right? So, and now that nodes are online, then it'll start, uh, the deployment of our application stack.
So I, I go to list, so you see, uh, over these three nodes, right, we have total, uh, nine applications, so instances, triple node. So we see here that our applications are now coming online. So yeah, they're getting implemented.
So we can monitor the state of all these application instances across the, the feed of three devices. So now we have it online. So we see, we can see that the application is running.
We see that the service is exposed on 48,000. So this is one of the, the client, uh, the client service is a web service that we can connect. And uh, we are going basically to just figure out what is like PRS of that device.
So 30, and they do connect on 40,000, and that's what we see, right? So that is basically a video stream, uh, and you can see basically the statistics of the inference on the objects. So here we see it's a bit faster in the detecting this is having a GPU.
So, and, and then the next thing is how with how we, how we keep that updated, right? So we can actually remote access that node. So we can, we can just check locally, uh, if in the hand time if this process to training in A GPU.
So right now I'm just connecting, uh, locally. So you see the Nvidia, uh, utility that is showing that, uh, our Python process from our container is running on gq, okay? And also we see the container running with our tech.
It's right now structured one. So later on we update this, uh, version of the, of the inference pipeline container. So is it the same thing on the, on the second node?
So we have two nodes. So we'll have, we have this co uh, deployed across two nodes. In this case there's no GPU.
So you see, you'll see the image a bit slower in this case. So the, 'cause it's not GQ compared to the other, compared to the other device. Okay?
So, and the same for no three, right? So no three is also running without, uh, without the GT U. And we have the same watch.
We can also access over the IP address, And you'll see that each is also slow compared to the first one. Conceivably, there's no practical limit to the number of, or the variety of number of variety of images you're using for this. You, you, you have two that it demonstrates the, uh, the, the platform.
But you could have like 10, 12, Yes. So there's no limit. So you can have, in this case, I have two without chief view or you, you mean, you mean the notes, right?
That you are, we're talking here? Yeah, I mean, one of the, one of the, I I, to, to my sense, one of the, um, one of the big advantages of SIDA presents, and I think it's one of the reasons why, um, so much of your, um, customer based use cases are on like that classic edge, you know, physical or demanding environments or what have you, um, is an ability to support a, a real wide variety of devices for what amounts to one application, um, or, you know, one application set. And so, um, it's, it's really, I mean this is, this is a really good example of it where, uh, you need to target maybe similar functionality to a lot of different types of edge devices or even conversely, um, have more uniform edge devices, but um, different packages that you're sending out and you can do it in a centralized and managed way, Correct?
So in your schedule, these call, so that you can actually just read the device through the policy and actually the policy based on the tag, which you quantify, okay, this device doesn't have a GPU, so that's, this is that application stack, that LP and you'll ask, do that version of the deployment policy, right? So that's what I'm showing here. So there are two versions of that deployment and the, with this we can actually do, of course the hand time there are slightly difference to, to it, but we have those two versions in the, in the marketplace.
But onboarding process, because you're talking about the same space, right? So right now, the next next uh, step in this demo is just, um, I already show the Grafana, uh, piece. So I'm going basically to show how we can update that version of the docker compose using this zis infrastructure that I mentioned.
So I'm just going to push, um, so I'm going to show that I am, I'm running basically this, uh, version draft zero one of our inference, uh, pipeline container. So that draft zero one and is what I will going, what I'm going to do now is basically the select list of our three hand times in these three devices. I will purchase this update to change a container image graph two.
And, uh, if I go now to, so this actually opex status you see here, so let me just pause the video. So opex status you saw there in that previous frame. So this one, uh, is basically the, the update that, uh, the hand time is pushing to the control.
So you can just say, okay, I'm running, okay, this is my personal, my on my container. So this right now is encoded, but what I did, I just used the API client on the side of it. So you see here the three nodes and basically just helps that, okay, I'm handing draft zero to already and this is, uh, what the hand time is publishing back to us.
Okay? So as I said, it is generic infrastructure, so you can publish any, uh, any j uh, formatted, uh, content. So, so now I'm going to remote access again from the, so you can imagine this, these nodes are remote, so there can be identifi off behind that.
So we can use this mechanism just to remote access the nodes, and then you can see that the device is already hand the draft zero two. And if I data, see now the UI as a, as a new look and fill. So there's a a, it means that you actually updated our, our image, uh, into, in those three devices, right?
So we see as well, that was updated. So it, uh, it's a slightly, we didn't, they actually start hand time. We didn't redeploy everything.
It just was a, a lightweight operation that we did in the hand time. And they use this Dix infrastructure to push the configuration updates into the, into the hand time. Okay, so now basically I'm just going to remove it.
I'm just going to put it back to addition. Its initial plate. So I'm just going to touch again the draft one and uh, that will, we'll see again, the state already been published by our, uh, agents running in the, in the hand time.
So it's this publishing back to the control, but now I'm running the version draft zero one and uh, I should basically remote access again and confirm, and that is now draft zero one as as well. So that, uh, it's back to this initial update. Okay?
So, and then the NVIDIA is also then n video util is also reporting the, so you c here DY spec, that, that initial state, um, the next, so you see this from three notes as well. So we did a bulk update, right on those three applications, uh, from the controller. So you can also extract logs.
So all these 10 times the container can be expect for logs. And logs are also visible in the, in the controllers. So you can see logs about the inference.
So you can see first persons or, so these are basically the logs of our, um, of our promeus proper that is exposing locally. And you can also see metrics of the, of our like N-C-C-P-O memory metrics, ization, uh, and we can also see network, uh, network utilization, network traffic. So we can also see here, uh, for instance the traffic towards the camera, so traffic towards the, the, the NVD GI for instance and things like that to get the plates.
So that's also visible in the, in our statistics. So I'm just going back to the, the catch all, the patch lop, all the plates are done. So now I don't want to use this, this split of devices anymore for HII don't want to deploy anything else.
I will like everything, so I'll just move the device from this a HI project into default project again. So that's what I'm doing by, just in the TE forum by exchanging the project assignments. So I apply now my terraform, so the stake will be changes for the, for those three devices.
So devices are gone for my HI and now they are back to default and then it default, there's no application. So basically just by all that, by just newing the devices from one project or another. So those three are now empty.
So that's, I think essentially what I wanted to show you. I think from my perspective, all of that work you showed is really hard to do manually. And that is what we're really driving at right here is the orchestration and the ability to have those apps to remove them without quite significant amount of work in our GI ops pipeline.
So, uh, yeah, great demo. Appreciate that.