SPLX CEO Kristian Kamber on Why AI Model Accuracy and Security Aren’t Improving
In this Techstrong.ai Leadership Insights interview, SPLX CEO Kristian Kamber dives into the state of artificial intelligence (AI) model accuracy and security that, despite the arrival of new offerings, isn’t improving.
Transcript
Hello and welcome to the latest edition of the Techstrong AI Leadership Insight series. I'm your host, Mike Bazaar. Today we're with Christian Can as the C-E-O-S-P-L-X.
And we're talking about, well, just how much progress are we making in securing AI models because, well, the theory at least was as we go along, they'll get more secure, but that may not be the case. Christian, welcome to show. Thank you, Mike.
Pleasure to be here. Alright. What's your assessment of what's going on here?
And some folks are saying that tests would show that chat GT five, for example, is not more secure than chat GT four. And is this just the nature of the particular beast, or are we just not paying attention to security? Again, I think security is not the main topic.
Uh, probably open ai. Other folks are, um, currently focused on, it's just more data, it's just more capacity, it's multimodality and other features. So I guess it's coming back to the topic, why security matters and why hallucination matters and why accuracy matters.
I think it's security is a pretty wide topic, and I think what we found out through our research is that, um, the company was not paying attention to hallucinations and I would say fake URLs and, uh, any type of other, um, fake news and bias, uh, biases, uh, which can happen and predominantly, uh, also fell into the security bucket as well. So, and uh, obviously, um, Sam Melman took it, uh, from a production, I think over the weekend it was apologizing that the team was tired. So I think we're now going back big time to, um, security and security testing, I guess.
And it's not only models of course, I think, um, we see, uh, a pattern in, uh, the main specific application when big companies are building those agents. Uh, the attack surfaces way bigger than just testing the model itself. The model's just the underlying model.
It's just the base you start with. Mm-hmm. Are you perceiving that people are kind of encountering these issues during testing and this is what's holding them up from deploying in production?
Or are they just kind of ignoring it and hoping for the best? I Think a lot of companies are ignoring that. Uh, we don't, we, we don't hear a lot of things in the public, but I guess there are some, some, uh, examples with custom, uh, customer facing applications.
But I think what we see currently is like big corporations use it internally for AI workflows and their agents. And I think, uh, they haven't been paying that much of, um, attention to testing. But as we do more, uh, work more with their own data and uh, basically fine tune their own data and incorporate into those models together, it's gonna be a massive thing and it's already a massive thing.
And I think, um, yeah, we're just at the beginning. So we work, for example, with one company. They have 48 hour SLA to deploy to use a model in production.
The pressure is huge from the business to deploy and use models at scale because it's not a winner takes it all market. And, um, I guess, um, we need more people and we need more skills, um, predominantly more skills, um, beside any type, any kind of automation currently. Um, is it always gonna be the responsibility of the people using the models to figure out how to secure them?
It doesn't seem like the providers of the models themselves are normally gonna be too focused on that. They might provide some maybe rudimentary guardrails, but I feel like we're basically saying the responsibility for securing them lies with the organization using them. That's a good question.
I think the responsibility for the plane model should be definitely with the biller of the model, but, um, the guardrails are not really knowing exactly what the user is supposed to use the model for. So if you take, I don't know, uh, logistics healthcare, you're gonna train the model, you're gonna fine tune a bit, you're gonna have your own rag. So it's, um, it's not easy to use any type of default, um, guardrail.
And that's why we always say there, no matter doesn't matter who does the guardrail on the lightweight guardrail to be custom built, custom policies for, for example, bias and off topic and I don't know, competitive checks and business alignments even, I don't know, profanity, those kind of things as that can be only done with insights into the domain of the, of the user. So that's why the user in this case, is the company who's putting those models of wrap models somewhere in the, in production towards the customers, the users. Um, and this is definitely the liability of the end user in this, in this, in this topic.
Yeah. Aren't we making a conscious decision about the risk or are we just kind of rushing in headlong? Because I can see some business users saying, well, it's good enough, it's right, you know, 95 out of a hundred times, and that other five times is just part of the risk factor that we're taking into account for the benefit of increased productivity.
It depends on the business. Um, we have been seeing use cases in healthcare where models were hallucinating and recommending you to take the pill before the breakfast or after the breakfast, or clean the needle twice, reuse it with alcohol and stuff. This can have cause, uh, this can cause massive, uh, massive, uh, issues.
Uh, not only health issues, but legal issues. So it depends which kind of business you are. If you have a, if you are a car seller and you don't care if people are just buying a car for, uh, for a buck, then it's a different topic.
So, um, you can put always a disclaimer and you just basically can say, Hey, you know what? I'm not liable for anything. The model spins out and you can build guides, but those are basically then limiting your user, uh, to a certain number of use cases.
And then you're just going back to the machine learning model itself. It has been a few years, so I think we should enhance gen AI and we should let genai, uh, be the automation tool and help help us to, as, as you said, have more productivity. But I think security testing and specifically AI security and safety testing becomes essential while, while you're building more AI models and more more use cases.
It all started with customer pacing chapels. But meanwhile, as I said before, we have AI workflows. We have, um, ai, uh, AI agents internally, communication tools, wrappers any type of ai uh, usage in the company.
We're just going above a hundred or a thousand in some fortune fifties. Yeah. Mm-hmm.
So are we therefore just waiting on some sort of cataclysmic event to occur before everybody kind of wakes up and takes a look at all this stuff? Or is there a way of thinking about this maybe in a more proactive manner? I think, uh, the legislation or I think compliance framework should definitely make a step.
Uh, a lot of companies don't really know, A lot of businesses don't really know what to comply with. There is some standards like os and this framework, but it's not really, um, giving or it's not really a mandate for CISO to secure the business. Right?
They can use it, but they're not like liable for anything in this case. Like they can secure themselves and saying, Hey, we're not liable for a PC. We can put a disclaimer.
I think, uh, once we're getting to the step, like having heavy usage of AI anywhere, uh, there should be a legal framework or at least a kind of, uh, industry-wide framework, which gives a mandate to the, to the CSO and their businesses to be liable, uh, to be, um, to be compliant to comply with. Mm-hmm. And this will, this will trigger the usage because in, in, in Europe we've seen AI and it was a mess.
People are just scared of deploying ai. That's totally a totally different story. I think people are just scared.
But AI does not built for, for, for AI security. It's just a legislation. It's just a, it's just a thing saying, Hey, you should secure your ai, but how do you do it?
There's many ways from offensive security to defensive security to governance controls, compliance controls. I think those measurements, uh, measures should be triple implemented anywhere. And then on domain, and I think this is the, this is the real future, I think, where we will have several frameworks for domain specific applications and industries.
And then also, um, yeah, secur, uh, variations on how to secure those. I can't help but wonder if we're also just ignoring some of the existing regulations we have that should apply, whether it's HIPAA and the United States or GDPR in Europe. Don't these things assume that there's some level of privacy and security being applied in the first place and it should apply to ai?
Yeah, absolutely. And they should. Uh, but they have, they should have a, at least a section which is, uh, covering AI security, let's say, um, data exfiltration, rec poisoning, data poisoning, so many new attacks which are coming out.
It's just not only proper objection on jailbreaks, right? Every attacker goes first with the contact leakage or the jailbreak. But if you don't succeed what you're gonna do, you're gonna do some patent code execution, enter the source code, uh, enter the, the, the server files.
Like a lot of things can happen. And I think this should be written down somewhere and this should be secured somewhere, or they have at least a guideline. Because what happens is security teams are now entering the space, but a lot of people in cybersecurity don't really understand the non-ad administered behavior of LLMs and AI models.
So what needs to be done in big corporations is like a lot of AI engineers and data science need to help them out. 'cause it's gonna be a whole totally different new animal. And what I would say smart businesses do is they create standalone AI security teams, which are reporting directly to the CSUN in this case and just taking care of this manner.
'cause it's a totally different animal that's uncomparable to cloud security. We've, we've seen before, and I would say from timeline perspective, we're like in 2007 or th 2008, uh, when cloud security just came out, and this is the infancy situation we have right now with the United Security. We've seen this debate before, but people will say, well, there's not enough cybersecurity folks who know anything about ai and therefore we need to train the data science teams more about security.
But you don't seem particularly interested either. Yeah, I mean, we have 30 people and I think, um, 25 people are just data scientists. Never.
They've never done cybersecurity before, never. We have a CISO and she's out of the cybersecurity space and I think she's helping us out with those acronyms and stuff. But at the end of the day, why we do continue, uh, AI security testing is we've seen on our, like we saw the first manual pen testing, like, and we saw, we saw, oh wow, this can't be fixed, fixed.
Uh, and then just, and, and one, uh, this can't be fixed in one time and let it go like for six months. Like we did that for w or whatever. It's impossible.
You need, you need do it continuously because the day you have text to text capabilities, tomorrow you have text to speech, you will document upload CSP pipes, whatever, and enhance new system prompts. So I think, uh, every two, three weeks and looking at model and, uh, specifically at the domain specific application with a model and several models, if we talk about agent, agent to agent communication, it's massive. It just can't do it one half and just let it go.
So in addition to the data science teams, we also need the auditors to be aware of what the security issues might be as it relates to ai. And yet they too are not trained. So how would they even recognize what an issue is?
Uh, they have the biggest business right now. I always say AI is a services business. It's not a product business.
It can be a product business and the B2C face and where our kids and teenagers and my grandma is using Chad GPT, but the use case productization and big firms and companies and corporations is still not there. We're pre-production. That's why we work with a lot of auditors and as they are at the core, they're at the core and they see what's happening and they actually switch to a lot of, I would say, penetration testing services and AI testing services and have their own own AI frameworks.
I think we have KPMG with their own AI framework, which is already going very deep into, towards how to secure AI pre-production and in runtime. So, um, you're totally right. The auditors are super important in this case.
Hmm. I also wonder if the bad guys, AKA cyber criminals might be a little more AI literate than the rest of us at this point. 'cause they're like, Hey, this is the greatest thing since sliced bread, or are they still trying to figure it out themselves?
A lot of things can happen by accidents and I think we, uh, we just had a, a a an example on the west coast, but a student, uh, guy, like a student who was um, doing an internship in the company and was like doing like, um, text to text communication with another employee. And we've seen some day exfiltration of Google workspace to g run Confluence out of this chat. So anyone who has access to this kind of chat, internal company chat and you work with customer data, internal data, any type of PII data, it's, it's, it, the attack surface is huge.
Uh, we're talking about 95% un un undetected, uh, uh, attack surface still. Mm-hmm. Long term.
I mean, I feel like we've been talking about the need to secure our data forever, but might the rise of AI just finally shine a spotlight on this issue as a, both a data management and a data security crisis that people will finally address? Yeah, I think we all, what we all need is visibility. I think currently we're having a huge, uh, ton of AI usage, which we can't, uh, assess and can't address.
I think it's, uh, from the user perspective, but from the application perspective. And I think what companies are currently trying to do is like uncover the AI assets and try to understand what's happening where have at least a kind of an overview and then they can build policies around that. I wouldn't be too strict about it.
I think what I like about the US is that people are just super positive about AI usage and trying to use that as much as possible. But I would just say when we're dealing with sensitive data, we should be always sensitive. And this starts with healthcare, going back to banking and insurance.
Um, it, we just need to educate the market and the users as well. What can happen. And this can't be just sold with a disclaimer or just, I don't know, quick questionnaire or, or whatever.
So this is the thing, but I wouldn't be that much of the pessimistic or a negative. Uh, it's just part of a wider ai, uh, cybersecurity space, uh, which AI security will play. What's your best advice to folks about how to go after that particular task or issue?
Because I think a lot of them will look at it and just say, wow, this is just so overwhelming. I have no idea where to get started. I think, as I said before, I think AI repository, ai, AI code analysis, um, scanning analysis, I think the shift left approach makes sense.
There's a lot of companies out there, um, we're going that direction. Even, um, classic SaaS and desk companies. Um, we call it AI bomb AI developed material.
I think having the first visibility of your AI workflows, um, as a CSO and as an organization helps you actually to uncover, oh, uh, is this in production? Is this in pre-production where it's connected to what MCP server is needs to be whitelisted and where this is the first step I would do. And then I would think about guard and then I would think about security testing.
That's so, because anything which is anything which is not uncovered, you can't even know what's behind. And then later on you would understand and you would need to assess is this of a high risk or if this is something in pre-production, I'm just playing with around if this is a high risk, then of course we need to do some, uh, governance controlled security testing, control, uh, security testing measurements and s measurements. But I would, I would go first with the discovery, um, with the ai, uh, AI inventory discovery.
I think part of the issue that people are trying to determine as well is you mentioned that most AI projects today are led by what I might call a TIGER team, and it should have some sort of security person involved in that. But folks are also saying, well, maybe only a matter of time before AI is pervasive and maybe I just have to figure out how to extend my existing security workflows and processes to include ai. Is that where, is that the curve that we're on and how long might it take to get there?
Or is that just never gonna be feasible? Yeah, I think both ways are, um, realistic. Uh, it depends what the company does.
If we're talking about a bank which is completely transforming its business and having a kind of a low touch business with the customer and client and user, um, having a chat interface or having some, uh, AI automation and, and talking and discover and talking with the customer and, and, and, and companies, I think, uh, a team is needed, right? A specific team is needed. Call it tiger team or call the AI security team, whatever.
But if your business is not that I would say, uh, that much of, um, affected by any type of cybersecurity risks and specifically I security risks and a lot of small business do that enterprise, midlevel enterprise do that. I think, um, I would go always first with, uh, with full risk or with a lot of risk to try to see what's happen, what's happening. 'cause we never protected insecurity something which is unknown.
We're always try to fix something or protect if we know the threat, right? Sometimes you just don't know the threat, right? And you can't anticipate it, uh, because you don't know how users will reactive how, how they gonna use it in that way.
But if we're talking about highly sensitive businesses, um, going to healthcare and biomedicine and, and, and, and drug production, whenever it's, it's necessary we have this steam, I would say Ultimately we talked about that skill shortage, but might now we one day see AI models that are trained to help protect AI models and therefore, you know, that's how we're gonna kind of narrow that skills gap. Yeah, we talk about that a lot. Agent to agent communication or uh, agent tech, uh, LLM based creative text.
I can, I can only say that the number of false positives still high 'cause those models hallucinate and stuff. So how we do it, we have a lot of 11 AI researchers. Every test we see out there and every prompt injection we see out there, we retest them with human and lube before embedding it anywhere.
It just just doesn't make sense. Because if you have a security team, let's say, and you mentioned in a question before with the cybersecurity teams, a lot of AI security components being right now embedded in cybersecurity controls and SAS and das in firewalls and DLP functionalities. And if you are just bringing some kind of software, which should secure this model and use LMS for that only, um, the companies will not use it that much.
'cause the number of fault spots is still high, right? So I would say a kind of number of expertise retesting human in the loop, adding a kind of some controls in plays already is something which, which I see as realistic as of current. Yeah.
And not a hundred percent protection is, is is not guaranteed. 5%. All right, well folks, you heard it here, AI security, it's not easy.
But the one thing was for certain is the longer it takes you to address it, the harder it will become. Christian, thanks for being on the show. Thank you, Mike.
Pleasure to be here. All right. Thank you all for watching the latest episode of the Techstrong AI Leadership Inside series.
You can find this episode and others on our website. We invite you to check those out. Until then, we'll see you next time.