Modernizing Data Protection and Migration with AI – Elizabeth Nammour, Teleskope
Teleskope CEO Elizabeth Nammour dives into how artificial intelligence (AI) will modernize data protection and migration.
Transcript
This is Techstrong tv. Hey guys, thanks for the throw. 2 million for applying AI to data protection.
So we're gonna dive into how the world, the data protection is changing. Elizabeth, welcome to the show. Thank you.
Thank you for having me. I think, to be honest, one of the reasons that we're kind of bad at this whole data protection thing is it's scut work. Nobody wants to do it.
It's difficult, it's a pain. You might not get it right most of the time, and there's about a thousand other more fun things to be done. So you guys are about to apply AI to this.
So is the whole process about to change and how So? Data security and compliance is just definitely a lot of grunt work. I think right now, the way most companies are doing it.
And what I did as well, uh, at Airbnb was doing a lot of manual assessments, manual reviews, manual, uh, inventorying of all the data that exists, which is impossible and gets deprecated as soon as it's done. So at Telescope, we hope to automate that work to make data security, privacy, uh, and compliance achievable at scale and reduce the manual work on security and engineering teams. Mm-hmm.
And what's involved in that exactly? I mean, how did you go about doing that? Because we've been trying to find the magic silver bullet for data protection for years.
Yeah, I think there's two things. One is understanding the data that they've already had, that they already have and have collected. So automatically, continuously inventorying all the data that a company has accumulated automatically using ai, pinpointing where personal sensitive data lives and how is it being stored?
Is it stored securely or not? And doing this at scale and with no false positives has been a really hard problem. Uh, and this is what we're hoping to solve.
So flexible, uh, scalable and accurate, uh, understandings on the data. Now, what type of AI are you applying to this? Is it machine learning algorithms or is it more generative ai?
What, what exactly is the path here? Yeah, so we're using, uh, large language models, not generative ai. So it's kind of a large language model that gives you the best accuracy, but the lang language model has to be also small enough to give you speed to be able to classify, uh, petabytes of data, uh, at scale and actually make it work.
Um, so it's a combination of getting a very large model, using that to train a smaller model that can actually handle, uh, the amount of data that even small companies have. And how do I take advantage of that without necessarily losing control of my data? Cuz a lot of folks are a little bit sensitive these days about where their data's gonna wind up.
So how do I know that my data's not gonna wind up in somebody else's large language model Model? So the large language model is, can actually be deployed on the customer's premises if they're, uh, paranoid about the data ever leaving their environment. So it would be a large language model that's actually hosted in the customer's environment.
O or it's a single tenant SaaS. So we deploy a large language model for each of our customers. And so the data never gets commingled or shared with any other customers.
We're not calling like chat G B T or like open AI's APIs to and send sending. We're not doing that. Yeah.
There is no shortage of vendors out there selling data protection platforms these days. I mean, here you are as a startup taking on the big guys. What makes you think that they won't just turn around and use the same kind of technologies?
I think they haven't been really, they, they haven't been operators, so they don't understand like what it's truly like working with those tools on the other end. And so they come with a few problems. One is they can't scale.
They talk about things in gigabytes, but when it comes to even scanning terabytes, uh, they break, you'll hear that a lot. Uh, and then also the false positive rates. Like a lot of our, uh, competitors, the large guys, they'll flag you for things that are country that are like, they say our first names, but our country names or they say are an address, but it's actually like a public restaurant address.
So they don't give you any like, contextual understanding around the data on top of just like what element is it? And so that's where we think we can have an edge, uh, of giving good, very accurate contextual understanding as well as something that's flexible and scalable and integratable into developer pipelines. Have we reached a point where it's just not possible for the average administrator to keep pace with the volume of data that they're supposed to protect?
I mean, have we just reached a point where we're choking on the data we have? There's also no longer, in a lot of companies, there's no longer these types of administrators, especially with the cloud. Like you have infrastructure teams, site reliability teams, but they are, there's no longer that like physical gatekeepers.
So I, a lot of companies like engineers can spin up their own instances, their own databases. So no one really knows what's going on behind the hood. So will it become simpler to make data protection therefore part of a larger DevOps workflow?
Because I'm gonna have the machines, the alerts will be more accurate and I can act on this stuff. Exactly. You can automate on top of it.
You can integrate like the findings and automate security and privacy controls on top of those automated findings. What do you think ultimately the relationship between data protection and cybersecurity is gonna be? We have been talking about making these things a little more tightly joined at the hip for years, but it seems like it's problematic.
Can we get past that? I think eventually it will have to. Uh, I think it's more because like sometimes you talk about cybersecurity and then engineering and data teams just don't want the additional friction.
But I think if you can build a tool that has no friction and actually works really well with your development pipelines, then I think that eventually they can all merge together. I Think one of the issues that we don't like to talk about in the land of data protection is that we don't often test whether or not we can actually recover the data in the first place. And, uh, much to our surprise, the attack comes and data's encrypted or the stuff was corrupted in the first place.
So will AI make it easier for us to figure out, you know, what actually is a good backup versus what is something that is problematic? Yeah, and then also like letting you know, okay, you have sensitive data here. You need to make sure you have all the right correction, like protections in place so that you don't get like ransomware attacks or like anything that can corrupt the data and make it unusable.
Uh, and you can make sure you have as many backups as you need in different regions, different replicas, just to make sure the availability is there. So as we go along, where do we go from here? I mean, once you kinda start throwing AI at data protection, you know, what's your roadmap look like?
Where, what's kind of on the top of mind for you? From here? Obviously making the AI as good as possible.
It's not just something you can throw on. It's like a continuous learning process, uh, continuously making it more accurate, making it smarter in terms of contextual understanding. So right now we can tell you whether we found an email address, uh, that's related to your employee or if we found an address that's related to your customers, but eventually we wanna be able to tell you, okay, this is a publicly related address, like it's a restaurant address.
You don't need to even look at this alert because it's not sensitive information or telling you, okay, this is if you're a healthcare company, this is like a doctor's address versus a patient address. So being able to hone in on like, we, you have this data sensitive, it's about your doctors or about your patients, uh, and then allowing you to automa automatically remediate way things, um, that's not using it yet. A lot of folks are also using their data protection tools to migrate data, right?
I mean, the backup allows the ship and then they use that up in the cloud. Why buy a separate tool? How do you think AI will help with that whole process beyond just data protection, but data migration?
Yeah, so one thing, uh, that's really helpful is to know, okay, you have 10 data sets that contain the same type of data. Why don't you just merge it into one or migrate all of it into one? Even if you talk about mergers and acquisitions, if they're trying to integrate like a new company's data with yours, uh, it can also help with that.
Cause now you know, okay, these are all the customers first names. Let's merge them all together into one place. All right.
There is, um, a lot of demand for folks out there that have titles like data engineers. Do you think that that function becomes part of a larger DevOps workflow or we always need a data engineering specialist? I think you, you need specialists in terms of data engineering.
Cause it's very complicated to build those robust like pipelines. But eventually all those roles will have similar responsibilities. You're already seeing data engineers have data governance responsibilities.
So that includes like security and privacy and compliance. The other phrase we hear a lot lately is data ops. That's becoming a set of best practices for managing data.
We have been trying to manage data for decades, and I would say most of us aren't very good at it. So, um, are we about to get better at this because we're gonna have a set of guidelines and best practices that people will implement. I think the issue with those is like, obviously the guidelines can be implemented by the data engineers, but when you talk about guidelines that all employees have to follow, I feel like sometimes that ends up not really working.
Cause no one really reads those guidelines. No one really applies those guidelines or remembers them after they read it when they first onboarded. So I think things need to be like enforced, uh, every time they're trying to do something.
Yeah. You of course spend some time doing this stuff yourself or Airbnb and some other folks. What's your best advice?
I mean, you've been in the trenches. What do you kinda see or wish that your organizations had done that others might learn from? I think obviously it's trying to automate, like when I first joined Airbnb, uh, they gave me a list of like 500, 400,000 columns and we're like manually label this, uh, to pinpoint where personal insensitive data is stored, uh, because you have to do this manually to start.
Uh, but that exercise is so boring. Uh, no one likes to do it, uh, it's error prone. And so we ended up automating a lot of these things at Airbnb to make lives better.
Not just for me, but for all the teams that were labeling their own data, uh, afterwards. So try to automate as much as possible. I, I think no team, no engineer likes operational work and they'll be much happier spending a week like building a script to help them automate a manual task than doing that manual task over and over again.
So kind of prioritizing those like, uh, optimizations are great. All right. Well folks, you heard it here.
If it's boring, automated, that's the way to go. Elizabeth, thanks for being on the show. Thank you so much.
All right. Back to you guys in the studio.