Reinventing ML Model Management and Elevating Software Supply Chain Security – Yoav Landman, JFrog
JFrog’s new ML Model Management capabilities, an industry-first set of functionalities, are designed to streamline the management and security of machine learning models. The new capabilities bring ML model delivery in line with an organization’s existing DevOps practices to accelerate and govern the release of ML components. JFrog also unveiled new capabilities for its software supply chain platform that set the standard for quality, security and integrity of software releases. From creation to production, the new capabilities infuse security at the binary level in every stage of the SDLC to ensure applications are traceable, reliable, compliant and secure.
Transcript
This is Textron tv. Hey everyone. Welcome back here to techron tv.
You know, I, I spent most of this week out in San Jose for the annual Jfr Swamp Up event. And as usual, it was a great event. You know, the folks at Swamp Up, I've said this for years, they, they've created, not at Swamp Up, but Jfr, they've created a great culture, not only of the frogs themselves who work there, but the partners, the, the, the users, the, the entire community really comes together at Swamp Up.
And you can kind of feel the love. It's, it's tangible. And, and I think it's a testament really, to the three co-founders at Jfr.
Of course, you know, Shlomi, Shlomi, Beka, C E O, Fred Prince, and y Laman. Um, I'm very happy to have, yo, we, we snagged them from swamp up here with us. Yo, of course is co-founder and c t O at Jfr.
Yv, always a pleasure to have you on Techstrong. How are you? Great.
Good to be here again, Ellen. It's good to have you. So, yobo was another fantastic event.
Swamp up this year. It's good to be, you know, not that Covid iss totally gone, but people aren't as worried. It's in the rear view mirror.
We've got a full kind of turnout again. Um, lot of interesting discussions, a lot of interesting announcements from Jfr. If you don't mind, why don't you, I'm gonna ask, I'm gonna put it on you.
Uh, what, what, what do you think were the big takeaways for people, you know, that they can listen to here now about, uh, was announced that, uh, swamp up this week? Yeah, so first it was a great event. Uh, I think we're out of the woods, or it feels like it's out of the woods with Covid.
So, great attendance, uh, many partners, uh, many customers, and many future customers. So we made exciting announcements around, uh, the Jeffo platform as a full secure software supply chain, end-to-end software supply chain in three major areas. The one, uh, the first one is really lifecycle management.
Uh, second one is around security. We're extending our security portfolio. I'll touch around it about it, uh, in a second.
And the last one is about, uh, machine learning and the support of J Fog as a machine learning, uh, uh, supply chain platform. Uh, so maybe, maybe we should start with, uh, with the release lifecycle, uh, management. Love it.
So, Jeffo always had, uh, binaries in mind. And as you know, a release, uh, at the end of the day, it's a binary that's going to end up in your runtime. That's everything you care about.
Like, you have the pipeline and the pipeline is significant, but it's just a vehicle to get a release to run somewhere. Otherwise, why, why bother developer, uh, this, uh, release? And what we are introducing is, uh, with this release first approach is the ability to, to attach evidence on top of the release.
So you can think about the release as a binary that leaves behind the trail of evidence. It could be, can be security scanning, it can be approval for promotion, maybe from staging to production. It can be, uh, some something else like, uh, even a general document like, uh, maybe a code coverage document or, or something like that, that q you can attach to your release.
Mm-hmm. Uh, we already have, uh, very strong building blocks in the j o platform that are highly adopted by customers, such as release bundle, which is, um, uh, which, which is the way to, uh, basically bundle an application. And we introduced the next version of release bundle or release bundle, uh, version two, uh, with this new feature of attaching, uh, extensible metadata, signed metadata evidence on top of it, and get this lineage and get the, get this full traceability, uh, about your release, the build that created it, uh, all the artifacts that were part of it, and any external evidence, like you can go back in time and say why this release is even in production.
It has a very critical security vulnerability. And you can show the evidence and say, but look, when, when I scanned it a week ago before I put it in production, it was not a known, uh, vulnerability. Here's the manifest, the signed manifested, or the evidence that, uh, that is attached to this release and proves that, uh, this was the situation.
So that's the release lifecycle management. Now on security, we announced a couple of, uh, exciting things about, uh, shifting left. Uh, so the first thing is that we're expanding our, uh, ID integration with sass.
So, uh, source codes, uh, scanning SaaS, and, uh, uh, so, so that's, uh, the jfo approach to, to SaaS scanning. SaaS scanning. We have our own, we devised our own, uh, mechanism, uh, of doing that, that can work across multiple files and support, uh, various, uh, uh, ecosystems.
Uh, and another very big announcement that we made is, uh, is curation and catalog. This combination of, uh, a curation engine, uh, which allows any security personnel to really focus on, uh, on, um, quality control even before quality assurance. So if you know this term from the industry, uh, you can do quality assurance on your, on your materials that are already in the pipeline.
But in order to be, uh, to save, uh, uh, a lot of time and money, you want to kind of, uh, eradicate on the external perimeter, anything that you want to bring, you don't want to bring in from the first place. And sometimes it's very intuitive, like if there's a malicious package, no way. I want it in my organization from the get go.
If there is a release that is too new, maybe it's two weeks old, so I don't want to be the new Guinea pig of, of the software industry and get, uh, uh, hammered by, uh, faker J or color js, if you remember those, uh, incidents, not, not so long ago. I can set a policy that say, okay, don't allow this, uh, two new release to, to get in. Uh, and so on.
So this is ge o curation. The nice way about it that it's seamless. You just, uh, point your artifactory at the curation engine and you get immediately all the good research, uh, uh, and metadata that the ge o research team has injected into catalog.
Uh, so for instance, you can set policies, uh, which will be also expanded in the future around the security, the, the risk level of a certain project, because the number of, uh, commits or, or the, the cadence of releases and so on. So this is catalog. This is something, it's like the database that curation consults with in order to create those policies.
And when you try to download something for multifactor as a developer, Artifactory will reach out to the curation engine, make, uh, the, the question and get a thumbs up thumbs, thumbs down decision of whether this dependency is allowed. We also know how to, uh, deal with the range of, like the graph of dependencies. Uh, and as a developer, you can, you will get an email.
If you get rejected, you can start a workflow of waivers and you can see, uh, exactly why, uh, sometimes in your ID or in your, uh, c L I tools, uh, why this, uh, uh, specific download request was rejected. And this comes on top of our, uh, of our current security offering. So if there is a situation where something which is vulnerable, uh, which wasn't vulnerable at the time of request, and maybe a week later the situation changed, so xray and advanced security will find out about it.
So we have Excel for, uh, component scanning and advanced security that is doing secret detection and, uh, zero day, uh, uh, discoveries and contextual analysis like, uh, applicability, uh, analysis, um, and yeah, and infrastructure code scanning and, and all that, that, uh, goodness. Um, so that comes on top of, uh, of our current security offering. Uh, so I'll, I'll pause here for if you, if you have any question.
Sure. I mean, there's a lot to digest there. You know, you have, we, uh, next month, October 16th I think is, uh, we're doing our annual virtual event DevOps experience.
I think actually Shmi maybe on a panel of CEOs from CloudBees and digital AI and a few others. Um, the theme this year that we picked is we're calling it Achieving balance in DevOps, right? com in 2013, 14, right?
And before even, we certainly have now seen the shift left happen. We've seen DevSecOps. I mean, from, look, just listening to you over the last few minutes, we could see how important security and software quality has become to that DevOps for the, not only to DevOps to the whole software development life cycle, right?
Are you afraid that we we're shifting left too far, we're putting too much emphasis on security? It it, or is it, is this now balanced and it was unbalanced before? What, what are your feelings?
Yeah, it's, it's a very good question. So at the end of the day, it boils down to a question of trust. Like everything that, uh, that we're doing is around, uh, really, uh, like, uh, emphasizing the, the trust or, or, uh, um, accelerating the trust that you have in your releases.
Because everything is automated with every, uh, pipeline is automated. At the end of the day, if you cannot instill trust in the process, uh, uh, you, it, it hampers your automation. So the, the one thing that you don't want is to overload your developers with too much information, uh, and too many findings.
Or some vendors, they may consolidate different open source tools that, uh, at the end of the day, they will reflect to them many findings. Some of them may be conflicting, and some of them may not even be applicable, right? So you may get a security vulnerability, vulnerability about something that your code isn't even making a call to.
So that can be reduced. Uh, so sometimes you have a network, uh, uh, exposed vulnerability, but your dock container is not allowing any network connection. So you, you are not exposed.
So the whole idea is not just to, it's actually twofold. It's first of all to consolidate everything around, uh, one, uh, single pane of, uh, of information. Uh, otherwise you're just going to be bombarded with, uh, with more and more, uh, information.
Um, and some, some security guys like that, they, they kind of, it looks like you have multiple insurances and, and, uh, it's not necessarily a bet thing, but when it comes to the developer, it slows you down. So you want to have everything consolidated, uh, and maybe you, you want to challenge the tool that you're using from time to time by comparing it to other tools. And the other thing is that you want to have only the applicable findings.
That's even more important. You want to reduce the noise. You want to, uh, avoid this, uh, vulnerability fatigue that developers are getting.
And at the end of the day, you have a huge depth, uh, of, uh, security issues that you already are not fixing or you are just, uh, waiving them. Uh, we, we also saw these kind of situations happening. Um, so increasing the quality of results, making sure that you have, uh, everything consolidated.
Uh, sure. So Let me bring up another topic. You can't, you can't walk three steps without tripping over something with AI these days, right?
Generative AI and how it's helping. We had a hackathon, uh, here two weeks ago with John Willis and Damon Edwards, Patrick Dubas, Shannon, Lisa, a whole bunch of DevOps people operationalizing ai. And, uh, the, the, the potential to use this to help make better code, more secure code, when I say better quality, you know, it's, it's not just fiction anymore.
It's not pipe dreams. It, it, I saw for myself, it's real. I haven't seen a, I didn't see a lot in, in, in your description of new functionality, but I'm sure jfr is, is looking at this, what kind of effect or impact is that gonna have, do you think, over the next six months, 12 months, 18 months?
So we're not just looking at it, we actually announce some, uh, Yes, Big, uh, uh, announcements, uh, and features around that. Uh, but you know, Alan, the reason that people are using J Rog, uh, initially Artifactory, and now also the, the security solution, uh, is, like I said, it's trust is, is to be able to trust the releases, manage them in one place, get good access control, using the, uh, checks and based storage of Artifactory as a mean to, uh, find out about tempering and, uh, and basically have one single source of truth for all the input of your builds and all the output of your, uh, well, the release, the, the final output. Uh, AI is not different.
It's just, uh, if you think about AI models, uh, so the model itself is a binary. Uh, normally it's, uh, it's a binary plus plus, uh, y plus plus because there are other, uh, binaries that surround it and are not less critical for, um, for the holistic view of the model, like the training data that you use to train the model, the artifacts like the dependencies, most of the models are python packages are, are, are using Python packages. I mean, at the end of the day, it's some sort of a neural network, uh, embedded with dependencies in Python packages.
So those packages may impact the, the, like, the stability and the quality of your model. Uh, the model itself is not runnable by itself. So at the end of the day, like most cases that we see, you put some sort of a docker container with some convention for the a p I endpoint that will invoke your prediction.
And then there's the results of your model that you're using for retraining and finding out whether your predictions are still accurate or whether you need to, uh, deploy new model. So there is a whole new workflow there, uh, that needs to be managed. And for, for the J F O customers, when we spoke with customers, their models, many times they're the crown jewel of Aries because they directly impact the, the, the, the revenues.
Sometimes, you know, they're making critical business decisions based on this model, not just, uh, identifying, uh, whether something, uh, is an animal or a human being in a picture. A lot of these models, they, they really drive the business, uh, so you need to manage them. And what, uh, what we found out is that, uh, there is, with l l m, you know, there is huge adoption of machine learning.
Gono says like, uh, 90% of applications in 2027 are going to to be, uh, and I to, to embed machine learning. And I think it's real because it's very easy. The barrier to embed it is, is very easy.
Uh, but we found out that the situation is that, uh, it's very much like the early days of DevOps. Like, uh, you have the data scientists, uh mm-hmm. Or the, the researchers, and they walk in their own small world on, on their desktop a lot of times, or maybe on a, on a remote development environment, uh, in the Jupyter Notebook, uh, doing the, uh, the scientist, uh, stuff and creating the models and training them and using all those, uh, nice Python libraries.
At the end of the day, they are not productive unless they get a DevOps to hold their hand and move them to production. You know what, even before moving them to production, the DevOps have to streamline the data from other systems for them in order for them to clean it up and, and train the model on it. So they need the, the, the operational end, um, in order to, uh, create trusted workflows.
And today, they, they, whenever we spoke to a customer that has machine learning for a couple of years, we found out that a lot of customers are already using the Artifactory as the, um, registry for their models because of the access control, because of the trust, 'cause of the checks and based storage and so on. There's always a DevOp team that, uh, uh, helps the, the, the scientist. Sure.
And what we announced, uh, is we, we said, okay, we have to, uh, we, we have to offer our customers a much better and much more mature and much more trusted solution around managing those models. Uh, so we, we did a couple of things. First of all, we introduced a dedicated machine learning type of registry.
In Artifactory, we are using hugging face as the format to, to begin with. Mm-hmm. Uh, so we are, you can host hugging face compliant models, but, but what's more nice is that you can bring on all the foundation models from hugging face to your own organization.
Uh, and those models, they can be used. Sometimes they are, uh, I mean, if you take lama uh, too, it's uh, it's around 10 gig, so it can take a lot of time to download. You don't want, uh, your organization to redownload it, uh, every time.
Uh, so we bring, we, we proxy the, the, the models from hugging face, uh, and more other, we scan them. So we scan them for, uh, vulnerabilities. And we found out there are, um, malicious models already in Hackfest using, uh, mainly Python to, uh, to allow a takeover.
Uh, I mean, at the end of the day, it's like a Python library that runs something. So you can, uh, you can use that, uh, for, um, for not KO server purposes. And the other thing Got it, is we're scanning for compliance, for license compliance.
So if you're using a model which is not friendly in terms of licensing, we will also also allow to about that. And, um, and we, we, we actually see a lot of customers already using these foundation models for logging phase, uh, retraining those models, uh, adding, uh, lower land, lower LE layers on top of the, the existing foundation models. Uh, so that's, that's a big announcement we made around trusted, uh, model management, uh, with fro.
So Y is all fantastic deep information. Unfortunately, people watching this, if they were not swamp up, they missed it. Uh, but we can't help that, but we can help them get on the ramp to, to jfr and, and to take advantage of these new things.
What's the best way for them to engage and, and to maybe, you know, try out some of these new features and capabilities and everything that was announced at Swamp Up. Yeah. So the nice thing with what we announced is that everything then we, that we announced is either, uh, ready for you to use or, uh, in, uh, in, uh, last, uh, beta phases.
Uh, so you can, you can use it today. com and you start downloading Yeah. The trial and, and, uh, and start experimenting yourself or reach out to us.
We'll know how to give you more information if you need. Absolutely. And we should mention that at the same time, swamp Up was going on, JFR was also, also had some folks over in DC right at a, an important cybersecurity conference there, uh, in, in conjunction with the government.
com as well for, for those who are interested. But if you want to give them maybe of just a 30 seconds what was going on there? Yeah.
So in, in the same spirit of, uh, of fostering software trust and trusted releases, the White House is, uh, uh, gathering again the group of experts to discuss what the net next steps should be with securing open source software and, uh, releases in general. So, uh, our CSO Azi is there together with one of our, uh, product, uh, uh, leaders. Uh, we were invited to these discussions under the, uh, open source, uh, foundation.
Fantastic. All, we gotta pull the plug on this one. Yo, thank you so much for coming on.
I know between traveling from Israel to California and back and everything else, it's a lot. And I appreciate you taking time out to come talk with us today. Be well, hopefully I'll see you soon again, but it was great seeing you, and it was a great swamp up.
Thank you very much. Thank You. My pleasure.
All right. Yo Laman, co-founder, C t o for Jfr here on Tech Drunk tv. We're gonna take a break.
We'll be back in a minute.