AI Security for the Health Care Industry – Techstrong AI Podcast EP56
Amanda Razani speaks with Yannik Schrade, CEO and co-founder of Arcium, about privacy and data protection in the health care industry, and leveraging confidential AI to safeguard sensitive patient information.
Transcript
Hello, and welcome to the Techstrong AI Podcast. I'm Amanda Ani, and I'm excited to be here today with Yanik Schrade. He's the CEO of and co-founder of Akim.
How are you doing today? Doing great. Happy to be here, Amanda.
Happy to have you on the show. Can you share a little bit about AUM and what services do you provide there? Yeah, so with a, we are building what we call unencrypted supercomputer.
Um, and what it means is that we are building a global decentralized network that anyone can use, um, in order to run any type of computation on top of any type of data in a fully encrypted way. So basically, our mission is to eliminate any types of single points of failure in our current internet infrastructure and enable full privacy for computations and the processing of highly sensitive data, um, and to, to enable that for, for any type of industry. So that, um, ranges from data analysis, computations, financial computations, but also the training and interference of AI models in a fully encrypted way.
And so all of that we are offering within our permissionless encrypted supercomputer. Oh, wow. That sounds very interesting.
So our topic for the day is generative AI for healthcare. Um, and we know we're, we all get concerned about our data, um, when it comes to the healthcare industry. Um, so how can we leverage, um, these tools that you're working on and leverage AI for, um, safer storage of our healthcare data?
Yeah, so, um, the, the current healthcare system, um, basically how it functions in any jurisdiction globally, um, is that there is sensitive patient data and most of that sensitive patient data sits in isolated data silos. Um, mostly for good reasons, because you don't want that data to be exploited. Um, and that at the same time greatly hinders what we can do with the data, right?
Because it's just this highly sensitive data and that that sits somewhere. And how the, how the current internet functions really is that, um, there's different, um, yeah, types of phases data can be in. Um, and one of those phases is data addressed data that just sits somewhere.
And usually, um, this kind of data is actually sitting somewhere in a protected way. Um, and so you have some server in a hospital, let's say, where your data is being stored in a secure way. Um, most of the time at least, but once the data should be processed, let's say we want to train an AI model on top of all of the data of all of the patients in the us, that becomes a problem because what we need now is we need direct access to that data.
So we need to decrypt the data and need to see the data in order to run computations. That's how computations traditionally function. Um, and that's a huge problem that we need to overcome because if we, if we can't solve this problem, we can't really use this type of data because there will always be some, some actors that will be able to exploit the data and exfiltrate data, right?
We've seen many data leaks over the years. Basically every week there's some, some major data leak. Um, and so what, what we set out to do with AQM is to enable exactly that, to be able to compute over this kind of distributed sensitive, encrypted data without ever having to see it, which is a very strange con concept.
Um, how, how, how can you run computations? How can you train AI models over data without seeing it? Um, but it's actually, um, on the most fundamental level, relatively straightforward mathematics.
And then I guess down the line, it, it gets more complex. But what we are doing really just boils down to mathematical formulas and cryptography, and that's really, um, the way we, we think about, um, implementing these kinds of systems where we cannot have any trusted actors in between, because trusted actors always mean there's a single point of failure. And so our approach at R is to go down, um, and study all of the mathematics.
We are a team of a lot of, um, PhDs in, in mathematics and build these systems from the ground up that, um, function entirely, um, on, on top of cryptography. If that becomes possible, then is to, to have this distributed sensitive data, keep it stored in a secure, encrypted way where you as the patient or the hospital, um, remain the data owner because only you are the, the party that can decrypt this data. Um, and then collectively be able to train on top of all of this data, um, and train models on top of this data or improve foundational models, fine tune models using this data, um, where funnily enough, the models that are being generated also can remain encrypted.
So it's, it's not just we are, we are using all of this encrypted data and then we get some model, and someone, again, could take ownership over this model, but it's possible to build what we call end-to-end encrypted ai, where these models that are the results of those encrypted computations remain encrypted again. And then if a doctor, let's say, wants to use one of those more powerful models, because those models have been trained on way more valuable, sensitive, um, and, and hence powerful data, um, they again have sensitive private information, um, that they want to input into those inference, uh, workflows. And so they can use those encrypted models with their encrypted data, get an output back, nobody sees anything.
And so that's end strand encrypted ai. Um, what we are enabling with rqm and what we, um, actually started, um, working with, um, different hospitals on, on building exactly these kinds of platforms of training models over sensitive patient data without it ever getting exposed. Um, yeah, and I think it even goes beyond that.
Um, if you look at, at hospitals, I think that's one aspect, but, um, 23 and me, I think is a excellent example of what can go wrong when, when we don't think about data privacy because, um, I, I read an article yesterday about 23 and me, um, I think they, they went bankrupt. And so, um, you can basically bid on all of this highly sensitive personal information and DNA of individuals, which I think, um, concerning humans might be the most sensitive information you could, you could have access to, right? And so, um, the way something like this would function with RQM is that of course you need someone that, that that can take your DNA and can actually sequence it.
Um, but then once you have this sequenced DNA, you can process it in a secure way. You can check your ancestry, you can check relatives, whatever, in a fully encrypted way where there's no party in between that has access to your data. And so I think we, we can move towards a future where we, we have hospitals where we have, um, these kinds of, I guess, um, private ways of, of providing DNA sampling and doing things like, um, taking all of the data from health trackers from, from whoops, right?
Um, taking all of this data, um, and being able to securely, you know, fully encrypted, way aggregated, train more powerful models on top of it, and then benefit, um, for everyone. Absolutely. It's funny that you mentioned 23 and me because literally before our interview, I just finished deleting all my data from there and closing the accounts.
I used it for years, and it was very useful and, and helpful information, but, um, just concerns there with, uh, data privacy and where it's gonna go, what's gonna happen to it. So my question though is this data, if it's so secure that nobody can see it, how do you guarantee that it's reliable and accurate? Yeah, yeah, that's, that's, that's an excellent question.
So, um, there's this, this mathematical concept of verifiability where, um, I can give you a proof, um, that I have executed some computation. Um, and you just have to verify this proof. Basically, you verify the output and input and this very compact mathematical proof, and then you can be convinced of the correctness of the output I've given you, um, in relation to those inputs because you have this mathematical proof, and I can't cheat in this process.
And so this mathematical concept of computational verifiability is what, what we have within rqm as well. And so you can really think of it as a very verification chain where you have some, some initial data and on top of this initial data that you put into the computation, you can perform checks. And you, you, you, you can validate, um, that, that data and in this process of validating that data, you can generate those proofs as well where you say, okay, this data is valid, and so now I encrypted the encryption that you have performed as valid.
So it's not just some random data, but instead you, every step of the way you verify those, those computations that take place. And so now you end up with those verifiable, encrypted data points that sit in different locations are being owned by different parties, and we, we run those training or interference workflows with rqm in a fully encrypted way, which again, also yield those verifiable outputs where you can be mathematically convinced that those values are correct. And so that's very interesting because nobody ever sees anything, but still you can get an output and be convinced, okay, this depends on the model that has been trained on all of this data that, that I've never seen, but it has been verified using these kinds of computations.
Um, and so this system of course also lies, um, to, to some level on, um, on I guess transparency. I, um, I, I guess, right? So, so you can actually, um, understand this audits trail, if you will, of, of, of verifications taking place.
And so that's why we've, we've also designed RQ m not in a way where we are just some company that has patented technology and then sells that to the world, but instead we open source all of our technology, right? That's, that's a key aspect to, um, to, to how all of this functions that you could go into the code base and look at every mathematical formula and be convinced of the correctness of the computations and all of those mathematical proofs. Um, and another aspect, um, and, and we recently, uh, published a, a paper about this is explainable privacy preserving ai.
Um, and so in the grand scheme of things, um, the way I consider what I call responsible AI is two main pillars. One being privacy, preserving ai, and the other aspect being explainable ai. Um, because I guess AI models even still nowadays sort of are black boxes now you more or less square this black box because we make it encrypted, right?
And so, um, there's, there's different ways of, of achieving explainability. Um, one of the ways we've, we've been doing that is using so-called chappy values. Um, and, um, what's interesting is that you can have an AI model, and as I said earlier in this concept of end-to-end encrypted ai, this model is unknown to everyone.
Nobody knows this model because it has been trained in this encrypted way and is being stored in this encrypted way. So nobody knows this model yet we can get explainable outputs where the, the, the outputs that are being produced can explain to us what input parameters, what values, um, were the reason why we got some specific output, again, without revealing underlying information about the model. So, um, it's, I think it's, uh, very hard to grasp concept, but on a high level, I think, um, it's, it's easy to see why it makes sense because once we achieve those two ends, explainability and privacy, we can actually, um, yeah, make all of those AI systems more secure and we can implement them into especially medical workflows where, um, we, we, we don't want to rely on, on some randomness, right?
We, we we want to be, be certain that the outputs that are being produced are correct. And, um, I think another aspect to this is, um, multi-modality. Um, and whereas, um, I think we've, we've seen this trend of everyone just wanting to use, let's say chat GPT for every single application, those those generative AI models.
Um, what, what, what I'm most interested in really is, um, having a large set of different kinds of models, um, really fine tuned for very specific tasks. So it's not just this kind of generative ai, but also different machine learning models which can do pattern recognition, right? And so, and, and, and different aspects.
So you're really constructing this, this virtual brain of different specialized, um, models that, that take on different tasks. And then I think the, the role of, of, of, of such, um, generative AI really is being the orchestrator, um, amongst all of these models, being able to make sense of the outputs that are being produced. But all of those different models can be highly specialized.
And again, since we have access to way more sensitive data without having to see it, we can get those very specialized models that then can be coordinated over. So you said you were working with, um, some hospitals, uh, for, uh, you know, a medical facility, for example, as an example, if they wanted to, to, to switch over this type of technology and this method. What is step one and what are maybe some, um, some issues you've seen during this process?
Um, uh, some tips or advice you can give to avoid any roadblocks during the process? So, um, we've been, we've been building this, this technology for three years now, and last year we, we actually acquired our largest competitor. Um, we, we really in the past have focused more on, on the financial side of things of, of enabling, um, yeah, financial transactions trading and the, the corresponding compliance that, that goes, um, in hand with that.
So being able to trade privately in general, um, in a, in a compliant way using this technology. Um, and they have really focused on, on healthcare and, and AI for healthcare. And so, um, we, we had a shared vision of, of bringing this technology to every industry.
Um, and at that point we, we then acquired them, um, and Dave spent nearly, um, eight years, um, developing that, that technology. Whereas we are our relatively young startup. Um, and they did a lot of work, as I said, with, with healthcare.
So, um, they in the past, um, I think have struggled a bit, um, because, um, they had a different, different approach to the topic where it really was there this company, this private company that keeps patents and has all of this technology sort of locked up. And then you, you need to very manually integrate their technology where it was really also big sales process on their end of, of actually, um, getting those partnerships. Whereas with RQM, we are really building this open platform that, that anyone can simply use.
Um, and that makes things way easier for, for, for companies, hospitals to use our technology because it's just this open computing stack they can use to simply compute over distributed, um, data. And that makes it way easier because, um, you don't as heavily rely on multiple parties deciding to collaboratively now do this thing where they have data, right? Where it's multiple hospitals 'cause that, that that's the, the most logical application.
Um, whereas with, with with our infrastructure, um, you don't need that. You don't need partners to, to get started because it's this, this open infrastructure where you can collaborate with anyone. So I think, um, that's something that we, that we recognized with, with, um, what, what this team had been doing in the past and um, that, um, yeah, gave us con confidence in, in our approach where, where we, um, yeah, make it way easier to use and, and more open.
Um, and I think one of the, one of the interesting problems as, as always is, um, I guess then data privacy laws actually, and, and compliance in that regard. Um, and what, what we've seen, um, and um, yeah, we've, we've done in the past is we can actually have a, um, solution with rqm that, um, is very interesting in that regard because you don't have traditional data processors anymore. Um, so where previously you had some data processor, which is something that most of those regulatory frameworks define as a term where yeah, it's just some party that has access to the data, right?
And so they have some, some compliance requirements, um, in, in handling that data, but we don't have that anymore because there's no one processing the data. It's only this encrypted data that you collaboratively compute over, um, each, each ca um, yeah, remaining in, in ownership over, over their data. So, um, that in the past has been one of the big hurdles for, for collaboration, and we can very easily, at least from what we've been been doing in the European Union, overcome all of those issues.
Um, yeah. All right. Well, if there was one key takeaway you could leave our audience with today, what would that be?
Um, that there is still privacy out there. Um, privacy isn't that. And um, I think privacy will see a big comeback, um, where we will see things like your data at 23 and me and, and other companies being leaked as a red of the past because everything in the future can run on this type of encrypted infrastructure where you can really see it as encrypted fiber, optic cables, um, where everything is connected, yet everything can remain encrypted, but still everyone can get shared insights.
Wonderful. Well, thank you for coming on the show and sharing your insights and about this new technology. Thanks for having me, Amanda.
Um, it was a pleasure And thanks to our audience. Stay tuned. There's more.