Couchbase’s Rahul Pradhan on Edge Databases
Rahul Pradhan, vice president of product and strategy for Couchbase, explains how databases running at the network edge are making it possible to process and analyze data at the point where it is created and consumed.
Transcript
This is Textron tv. Hey guys, thanks for the throw. We're here with Raul Pradan, who's vice president of Product and Strategy for Couchbase, and we're talking about the rise of edge computing and what types of applications we're actually gonna see at the edge because, well, there's a lot of data out there and it's gonna need some form of persistence, and we'll see what gets involved in that.
Raul, welcome to the show. Hey, Mike. Hey, thanks for having me on, and glad to be back on your show.
I feel like there's a certain amount of irony in this conversation we're having, 'cause we've spent the last 10 years trying to shove everything up into the cloud, and yet here we are now talking about the rise of edge computing, and it feels like we are trying to process and analyze data closer to the point where it's being created and consumed. But from your perspective, what is driving this next generation of edge computing? No, I I think you kind of, um, uh, captured that correctly.
I think we've spent all this time going from, uh, going to the cloud, and I think right now with what's happening is, uh, with, with the, with the confluence of events that have happened in terms of the technology evolution, we are seeing the, the, the emergence of edge computing and overall, uh, edge AI in general. Uh, so before we get into that, let me kind of step back and, um, kinda define what we think about as edge computing and what's generally accepted as edge computing. Um, edge computing typically means processing or, or performing competition as close to, uh, the generator or the recipient of the data itself.
Uh, and the, the goal of the, uh, of, of edge computing is to really enable the real timeliness of that, of response to get some, uh, analysis done in real time so that the, the, the, the, the latency that the end user sees is as minimal as possible. Um, now depending upon the, the level of latency that your application can tolerate and the nature of the, of your use case might be, there are various different levels of, uh, edge, uh, computing itself. So we talked about cloud and which is typically your hundreds of milliseconds of latency for your application.
But there's also the, what, what we kind of define as a near edge, which is, which is mainly the, um, the, uh, the local regions within a cloud, um, within a cloud platform. So think about this as like AWS local zones or wavelengths or some of the emerging technologies with, with 5G, where now you have, um, 5G mec, uh, regions that, uh, cloud providers have created. So those are the, the, what we think about as near edge, where the connectivity is still, uh, uh, on the backend of that, uh, uh, region is pretty significant as it gets connected back into a cloud.
But from an application perspective, the, the latency is anywhere, uh, up to 50 milliseconds or so. And then you get into more about, uh, well, what we call kind of the far edge. Uh, and then the device edge, the far edge is really more about your typical branch office location.
It could be also be, um, be, be, uh, an a connected platform itself that is either, uh, out on an oil rig or in a, or in an ocean liner or a cruise ship, uh, where it, it may, uh, end up losing connectivity to the, um, to, to the, uh, to the internet itself. And then finally you have the device edge, which is on the device, and that's the one that will have the lowest latency, um, uh, for the application and for the end user. So if you think of that spectrum of edge, uh, the, the use cases and how, um, the data is processed and consumed depends upon, uh, where that data is going to reside along that spectrum.
So given that spectrum, are these mainly going to be new greenfield applications, or are we gonna start re-architecting our existing applications more to incorporate stuff at the edge? And this is part of a larger kinda modernization movement. I think it's a, it, it is, it's partly both, but I think there's also, um, uh, there's also a modernization moment, uh, that's happening to incorporate edge computing because one of the key, uh, uh, key key reasons that you want to develop an edge computing application is really to reduce that late latency for the application.
It reduces cost. One of the, one of the key things, um, with the cloud today that people are discovering is the, is the cost it takes to actually move data across, uh, regions into the cloud and, uh, back and forth from the cloud and to the edge. So there, there is that, there's also the complexity of the overall application architecture.
So this is an opportunity for, um, an application developer and an organization to actually simplify and be at, at, and at the same time be more, uh, uh, receptive to the, uh, to the, uh, to the use to the response time that their, their customers and their users are going to get. Um, the other key thing that, uh, edge actually is unique, uh, about, uh, about what's unique about Edge, uh, is the fact that, um, the edge can itself be offline or, uh, or online. So having that disconnection tolerance, being able to develop application in an with an offline first mindset, uh, is, is really critical for the success of edge applications.
And that's where having a highly performant, globally distributed data layer comes in, where the data layer needs to be able to, uh, be, be, um, uh, take into account the disconnected nature of the network itself, still be able to persist the data, and then being able to resolve conflicts when the, when the, when the connectivity is resolved and sync the data back to a central location so that you do not lose the, lose the transaction or the context that happened when there was, um, there was no connectivity. Is this a variation of what we typically used to refer to as batch? Or are we moving to more of a event driven processing architecture where some things are online and offline and we can remember persistence, but ultimately we're not just holding on to data and shipping it every six hours and hoping everything reconciles?
No, I think that's, that's a, uh, great, uh, way to think about it is there is a distinction. This is not, this is not batch, this is more of the real time nature of it, and that's why the sinking becomes real important eventually. But the same time, getting you access to a data in the moment is crucially important.
And we'll talk more, a little bit more about AI and how that comes in with regards to inferencing on the edge. But the key thing there is as applications, uh, we know in general the applications that we all love today and how applications are evolving, they're becoming more hyper contextualized and and personalized. And what that means is you, uh, you as a user are expecting the application to exhibit real time now.
So you, you want responsive applications that are, um, that as you use them on your phone or your devices are able to react in real time, uh, in the moment that you are in, which means that they be, they will, should be able to process the data locally whether there's connectivity or, or, uh, or no connectivity. Uh, and the same thing applies to, uh, we haven't, uh, talked yet about, um, devices in general because we all have several devices now that are connected and are in various modes of connected and disconnected that are also iot devices, which are continuously gathering information that need to eventually, uh, get connected to a central store where they can be processed and analyzed further. So there's a, there's a whole spectrum of, um, of, uh, use cases and devices, uh, that need this level of connectivity and processing power in order to give, um, their users the experience that they need.
And to your point about the user experience, have we reached a point where the end user of the application expects the data to be current at all times? And I usually cite this example. I mean, sometimes you get a, uh, they'll tell you a certain item is available in a store and by the time you get there, it's not there.
'cause it turns out that they haven't updated the system in 24 hours. I think our tolerance for that kind of stuff is dropping a near zero. So is this just becoming new table stakes for expectations for how current data's supposed to be?
Yeah, and and that's a, that's a great example, right? And this is, this is kind of what we all, uh, uh, live every day and, and the frustration and, uh, uh, and the, the, the experience that the end user gets if they have to go and make a trip to a store or in their, and it does not meet their expectations, that is, that is yet another, an op, uh, of an opportunity for a retailer or a, or a or a business to actually lose that customer. This is where, um, it becomes, um, uh, it becomes extremely complex from a data architecture perspective too, to make sure that you're syncing the, uh, the, in, in that example, uh, sink syncing the, the, the SKUs that are in stock in the store, uh, to, to the application and back to a central location so that such that when the user is, is trying to make the competition of where to go and pick that specific, um, item, uh, it is actually in stock at the time that they're looking at.
Um, and, and this is, uh, this is where, um, AI becomes even more interesting because now you are, you are in this, uh, in this world where you have, you have the data, you have real timeliness of that data and, uh, to know exactly which specific location has, how many units of a specific item, um, and as a user or as the unit moves off of the shelf, to be able to reconcile that locally within the store as well as back in a central location is absolutely critical and important. Uh, this is where the, the, the, the, the concept around offline and online connectivity that I was talking about comes in. And now if you marry that with ai, uh, you are now, uh, you, you can envision a world where, um, uh, with, uh, especially in a retail case, it might be an omni-channel retail experience that a customer is all, uh, used to.
They might be looking for a specific item, uh, online in a store, um, and then, um, they would want to know where that item is or what's the fastest way to get that item. Sometimes it might be shipping, sometimes it might be just driving and doing a pickup for their order to be able to actually know that, get that information out, um, it is vitally important for that data to be as accurate as possible. And then if you think about what, uh, generative AI is also able to bring, uh, to this is not just the capability of analyzing that, but let's say that, uh, that that specific user or the customer is driving down a specific location, um, and then there are a few miles away being able to entice the user to come into the store by giving them, uh, specific offers or giving them reminders.
There is specific item that they were looking for is in stock a few minutes away, uh, is a very compelling, um, use case and a compelling value prop for a, for a retailer in order to, um, uh, to build customer loyalty. How will this play out in the context of ai? And I'm gonna do some gross oversimplification here, but we train the AI model in the cloud and then we deploy the inference engine at the edge, but over time we get more data and we wanna update the AI model.
So how do we get the data that we just collected at the edge back to some form of an AI model that can be updated? Can that all take place at the edge or do I gotta do a round robin circle to the cloud? How do you perceive that this is all gonna play out?
No, that, that's a great question. And that's one of the, one of the key, um, the advancements that are going on in this area is, is this, is this whole notion of model quantization is how do we simplify and make the model small enough, uh, in order for them to actually run on, uh, lower or constrained, uh, computing environments. Um, so think of it as, um, as there are various techniques that go in, in, in order to actually quantize a specific model, uh, including, uh, reducing the, um, reducing some of the, the parameters and the weights, um, making sure you can actually fine tune the model.
Uh, so model can be trained, but if you, for example, in a, in a, in a generative AI model, uh, you have, uh, techniques like rag or retrieval, augmented generation, uh, that you can inform the model with newer set of data for it to make a decision for you or, or generate some content for you. Um, same thing, uh, goes along with, um, uh, adapting, uh, models which are, uh, which which are quantized for example, to, uh, to to to smaller or resource con constraint environments where they can actually run on this edge location itself, can be fed new data and they can generate the, um, the insights for you or recommendations for you. Uh, that is a key, um, key, key focus area for a lot of, uh, researchers and a lot of, uh, effort that is going in this space in order to enable exactly the, the kind of use case you're talking about.
It used to require a significant amount of horsepower in a data center to run those types of applications. Is that level of horsepower and compute power moving out to the edge? Is it becoming more affordable to do this stuff at the edge?
Because we do have the compute and storage processing capabilities now has advanced that far? Um, not quite though. I mean, if you look at, um, again, I think there's a spectrum of edge, um, like we talked about earlier, and there are going to be areas where there is going to be a higher power compute available where you could potentially do it, especially in the near edge situation, on the device edge or on the, on the, on the far edge.
That's where, uh, you get into, uh, the, um, the, the, the smaller models or the quantized models, uh, that will, uh, be more applicable for. And in a lot of these use cases, the, the results that you get of the, of, out of those models are good enough for the use case that they're trying to solve. So you are, you are not technically trading off much of the, um, much of the accuracy of the model in order to get the benefits out of it.
So it's, it's a situation where based on where you want the model to be deployed, there will be different, uh, computational resources available to you. Uh, the other other point I'll make is, um, even with, from a, uh, computational resourcing perspective, uh, what is available today even in our, in our devices on our phones, is significantly more capable than what was what used to be available, just say five, 10 years back. Uh, so the, the computing power has definitely increased, but so have, have our computing requirements increased too, and ba and there's a, there's a, there's a match of those that is, um, that is happening, uh, based on the location of where you are, uh, trying to run that model in.
So ultimately, what's your best advice for folks about how to get started down this path? 'cause I think a lot of people, they kind of intuitively understand what's occurring, but they're not quite sure where to begin because in some ways it can be overwhelming. So where's my first step in the journey?
Yeah, it's, it's, I think the first step of the journey, um, as with everything, uh, really starts with from your user perspective, what, what are you trying to solve? So really understanding the problem that you're trying to solve, and then what's the best solution to get to that problem, um, to solve that problem. Uh, and a lot of times it's, it's going to come down to, uh, you, you are going to, especially in the s situ, uh, scenarios, you want to have, uh, or give the user the best experience from a latency contextualizing and personalization perspective, in which case you are need, you are going to need to have data that is, uh, locally available in that specific edge region.
Um, to be able to do that, you need to be able to have the, the, the, the data layer, uh, that you're storing and needs to have the, um, and needs to be able to offer the disconnection tolerance so that if the connectivity does go away, your data is still persisted and can eventually be, um, be, uh, synced back into a central location that becomes a key, uh, aspect of it. So understanding what, what is going to be, uh, which data layer you pick, uh, that gives you that capability to have that, uh, online, offline dis uh, connectivity mode as well as the ability to, uh, uh, to replicate back into a central location becomes extremely important. And then if you, as you think about with any data, um, there are going to be numerous, uh, devices or instances or something, uh, of, of a data that is being concurrently accessed.
Uh, is that data going to be, uh, are those conflicts getting resolved or not, or who, uh, or how do they get resolved? That becomes an important consideration too. Otherwise developing an application that sits on top of it that actually is going to do that for you is going to be a big, uh, challenge from a developer perspective, which frankly is, should be a concern of the data layer and not of the, uh, of the application developer.
All right, folks. Well, you heard it here. There may come a day soon when we're processing as much data at the edge as we are in the cloud today.
And if you asked me that question five or six years ago, I would've said unlikely. But here we are, Raul, thanks for being on the show. Thank you so much, Mike.
All right. And thank you. And back to the studio.