The Hero Project with DeployHub’s Steve Taylor at OSS Seattle 2024
Steve Taylor, CTO of DeployHub, discusses the Hero project at the Open Source Summit 2024, which aims to integrate security seamlessly into DevOps pipelines. By leveraging CD events and Ortelius, Hero passively gathers data from millions of workflows, enabling intelligent decision-making to enhance software supply chain security. Taylor emphasizes the project’s potential to revolutionize security practices by utilizing existing infrastructure and AI techniques without requiring developers to make significant workflow changes.
Transcript
This is Textron tv Everybody. We are here at Open Source Summit 2024 in Seattle, Washington, talking to the amazing people who're all part of contributing to open source projects, as well as their own companies and activities they're involved in. And another great person we're talking to is Steve Taylor.
Welcome Steve. Welcome. Thank You.
Good to have you here. Steve is CTO with Deploy Hub, and also contributor, involve and Engager and lots of open source stuff. I know you're doing a talk today Yep.
Later today, right? Correct. And it tied into the, in essence, in the keynote about, uh, your project is called, uh, hero, just the ones talking about this afternoon.
Tell us more about it. So, um, what we're doing is, um, I'm part of the arterius, uh, open source project that's under the CD Foundation. And one of the things that we've recognized is, uh, in order to bring SEC into the DevSecOps world Mm-Hmm.
Um, we need to bridge the gap between what's happening in development and the, the security side of things. So when in order to do that, we have to build out the Devs SEC pipeline. So we have to somehow get the security tools into the pipeline, say.
So I'm curious about that, and I'll let you finish. Yep. You said DevSecOps Pipeline is, is, do you think there should be a Dev DevSecOps or maybe there already is pipeline or multiple in?
Is that how the right way to think about it? Well, I think the way we're approaching it with the Hero project is, um, we need to achieve, uh, a DevSecOps pipeline. Okay.
Okay. But in order to do that, we're not gonna be able to convince all the developers and all the, the DevOps engineers to go modify every single Right. Jenkins file that's out there.
Yeah. Just, just not that gonna happen. Yeah.
So part of the people Around to do it, aren't they? Or probably in many Chases Right? Or people doing other stuff.
Yeah. And it, it's, we're talking millions, millions and millions of, of, uh, workflows. And then all the testing has to go behind it once you've, so we, we recognize that it's not going to, um, happen that way.
So in order to achieve the DevSecOps Pipeline, uh, part of the hero project we're launching is u using CD events out of the, uh, CD F1 of the CDF projects to start listening for, um, things that happen in the DevOps pipeline. Mm-Hmm. So, for example, we're gonna listen for an artifact to be published, um, an SBO to be generated, things like that.
And once that happens, uh, Orus, uh, would be out there listening for those events and goes out and starts gathering all the gory details about what happened in the pipeline. And with that data, we can now associate the security stance of a project or a repo, um, cross-referencing the CVEs vulnerabilities. Um, those type of things at that level.
Interesting. So is, is Hero sort of enhancement toward Totius to kind of create, I'll call 'em listeners to go out and gather this data? Is that a good way to think About it?
Yeah, that's part of it, uh, is to be able to, um, the CD events project did some small POCs, um, to prove out the, you know, the viability of, of the event project. Mm-Hmm. Um, so we're gonna take that to the next step.
So, um, we're going to, you know, we're basically, as part of the hero project, we need a, a, a big end user, somebody that's out there generating, you know, running thousands of workflows, um, whether they're gonna be Jenkins, Spinnaker, GitLab, doesn't really matter. Mm-Hmm. We're gonna plug into the, the, the CI part of that through the events.
Um, and then Arterius is gonna be there, uh, listening to the message queue of those things coming in. And we will start gathering the security data for those for that. So it's gonna, it is kinda like a joint effort Mm-Hmm.
That we need to, we're basically gonna try to bring all the different players from the DevOps side, well, from the end user side to the dev, develop DevOps, uh, all the way to the, the security side. So we'll need to player from each one of those silos. So.
Interesting. Then, so there's the gathering, the data, the events that are happening, uh, probably some form of correlation about how you put that information together and then producing it in whatever usable outcomes Yeah. For that.
Is this is hero about all of those things or tackling parts of that process? It's, it's about all that. It's about all that.
So all the pieces are there. You know, we have, um, we have the CI/CD side, so we have like Jenkins, we have GitHub actions. So we have that side.
We have the Orillia side, which is good at gathering SBOs and CCBs and Coral doing the correlation. Um, we have other tools that help produce the SBOs, like sift, um, you know, some of the native tools and built into, uh, like Docker, um, has a native tool to generate SBO m Um, and then we have the CD events that is, um, created a standard for us to basically produce the cloud event in certain schema that Ortel would be able to consume and then act upon. So the hero project really is, is taking all these pieces that have been do, done, you know, that are out there and combining 'em into a working solution.
It kind of sounds, this is maybe a bad analogy, but it sounds kind of like you're doing what OpenTelemetry does for event data and in kind of operational world, you're kind of doing that for the DevOps pipeline in a way. E Exactly. That Sort of roughly.
Yeah, exactly. Ballpark. You know, the, the OpenTelemetry idea is, you know, no matter what type of log file or, um, you know, runtime event that's out there Mm-Hmm.
I'm going to get that into a certain schema and then be able to then broadcast and persist it so other people can consume it and make decisions about what's happening. Because that's one of the big things that's missing in the, the DevOps and DevSecOps and the security side, is we leave a lot of the data on the floor Mm-Hmm. The exhaust, right?
Yeah. We, we just, we do the build, we walk away and, and then it's data that we can now use. If we start capturing it, uh, organizing it, we can then take the next step, which is apply obviously AI and ML to it.
Mm-Hmm mm-Hmm. You know, so that's, that's where we're, we're headed, but we first have to get the pieces connected together. Um, and that's where we're, we're, we're headed right now.
Okay. Good. So you said you're looking for someone who does a lot of, you know, software builds and and development.
If, if you were gonna talk to somebody out there that describing who you're looking for, who would be a good candidate to work with you on this? Um, what kind of company? What, yeah.
Uh, like some of the companies that are in the CDF, um, like Fidelity, um, has their own in-house version of this that they've written. Um, the Fidelity hasn't publicly immediate open source yet. So something like a Fidelity, um, that would be in the financial world insurance.
Uh, and then obviously you get into the government world, like, um, SpaceX, um, so that's SpaceX, space Force, space Force, uh, A FRL, which is Air Force Research Labs, um, the defense contractors, those type of things that really get into having that, uh, SBO m that's, you know, part of the presidential, uh, mandate. Mm-Hmm. Uh, that's the next connection I was gonna ask you about.
Seems like if you can do enough of the right things and position this well, this could be a good solution to software supply chain security issues that the government is not mandated yet. But you all know we're down ahead down. Yeah.
It's, it's coming that way really quick. Yeah. And you know, a lot of the pushback when you've talked to developers is they don't want to be bothered with generating SBOs if nobody's gonna look at 'em.
Right. Exactly. So why do we need to have them do more work?
And even on the DevOps side, the DevOps engineers going in, in tweaking all these workflows if nobody's gonna use the data. Exactly. And, but the data's critical.
The data's gonna allow us to make, uh, intelligence decisions down the road. I mean, we didn't even get into like open policy agent and those type of things that you could do with the data to see, um, if something should be blocked. You know, I got a bad package that I'm consuming, um, that has a vulnerability.
Uh, other things that when you look at the data, you could actually do, uh, threat mitigation through Mitre attack, um, match up your packages with the different techniques, um, the different groups that are exploiting the different techniques and really bring the data to the security officers to say, this is where we stand, and this is the CVEs and the vulnerabilities we know that are in our software. What is our overall risk? Where should I tell the, you know, the security officers, you know, them trying to convince a developer do something.
It, it's, it's very hard for them to go through it and Sure. And get that on the project schedules. But if they go, if you do this one thing, we're gonna solve 90% of our security problems Mm-Hmm.
And they can then take that to the, the C-I-O-C-T-O and they could push it down that way. Um, as part of that, that way to mitigate, you know, potential attacks Seems like a way too of kind of implementing, you wanna call 'em secure, secure security guard rails, which essentially policy, right? Yeah.
It doesn't have to be an offer and on kind of thing you can say, we want to kind of, we wanna start to shift doing it this way or doing some things this way. And rather than going to all the developers and saying, you know, you've been doing it wrong, this is the right way now. Yeah.
That's a good way to get people to change. Yeah. Right.
Yeah. But, but you could do that. But the interesting thing, it seems like you could take this data with some interesting models and almost simulate, okay, we can simulate CD pipeline DevOps pipeline all day long.
Let's introduce this kind of a security issue into that and see what the model produces. Oh, yeah. Right.
Definitely. Now you've got real data that you're operating scenarios. You're not just table topping it.
Yeah. And that's, that's a key, you know, with any AI or ml, is you need the data. Mm-Hmm.
And you need a good amount of it. And, you know, um, the, the cloud beads, Jenkins, that cloud beads runs on their SaaS platform Mm-Hmm. Those 90 million workflows a month.
Mm-Hmm. So the data's there, it's just being ignored. Like I said, it's being dropped on the floor and nobody's going around and, and acting upon it.
Well, and that's, and that's one drip in a very big bucket, you know? Yeah. There of all the, the things going Chen is massive, but Yeah, you look at GitHub 350 million repos in GitHub Mm-Hmm.
And all of them have some sort of action. Yeah. I mean, just with their aelius.
'cause Aelius is microservice based, uh, it's the, the dashboard piece of it, uh, is microservice based. So we have, uh, 50 different microservices. Um, we will run, uh, uh, basically, um, renovate from mend, uh, to bump our dependencies Mm mm-Hmm.
Um, every day we do about, uh, two dozen updates to our repos for every single repo, you know, so we're, we're constantly Another great open source tool. Yeah. We, we, we constantly are kicking out new versions of our software, um, all the time.
So when you look at, uh, not only the number of repositories, but the number of times those repositories get acted upon, you're up into the billions Mm-Hmm. You know, so that's, that's, that's stuff we can definitely do. AI and ML against, um, auto remediation of vulnerabilities.
Um, you can take and look at, uh, pipelines and say, oh, your pipeline isn't running the open SCF scorecard and you're running a GitHub action. Let me make a pull request for generative AI to insert that into your pipeline can create that. It's a pr they can approve it whenever they want.
Now we're collecting that data without the developers. All they're doing is Overview. That makes a ton of sense.
Right. It's already happening. Yeah.
Let's either use the data that we were already gathering or start to capture some additional things that will help Us. Exactly. Exactly.
So your talk today is this afternoon. What are you, what are you hoping to come out of the talk today? Our goal is to, uh, bring this awareness, uh, to the folks in the room and let 'em know that, you know, this project, the Heroes Project started, um, what we're looking for and what we're looking to achieve and what the possibilities are.
And the whole point of it is to do this all passively. Mm-Hmm. You know, again, the goal is not to get the developers to go and change their, their workflows, but to passively do all this, there may need be, may need to be a little bit of, uh, work on the DevOps, you know, Jenkins administrator to enable a plugin Mm-Hmm.
And set up A-A-U-R-L for us to capture the events to, but we're trying to do this as minimal as possible and at a global scale. So, um, you know, Jenkins, it would be for all of the workflows running in Jenkins. You don't have to go and install the plugin for every single one.
Same thing with Spinnaker, GitLab, all those things. GitHub Seems like it's CloudBees GitLab, Microsoft. Yeah.
With GitHub all be great kind of potential partners doesn't Yeah, exactly. And, and you know, like I said, it, it's all the pieces are there and we just gotta, and it's been proven, um, at the smaller level, we just wanna bring it up to the, the next higher level. Mm-Hmm.
Oh, excellent. Well, good luck with the talk. Good luck, especially with getting the right folks.
It's gonna be Folks to help you. Yeah. They're asking me, the tricky part is getting the right f folks in the room.
Uh, and once we are, we start, uh, getting this data collected, we will be back. 'cause we'll wanna brag about what we're doing. Exactly.
You know, it seems like if you, if you can kind of get across that idea of we want to do this in a passive way, it does require to developers to make changes can be, you know, leverage on 90% plus of what you're already doing and create, start to create these outcomes you would use in these ways. That would be, okay, sign me up. Right?
Yep. Yep. Exactly.
That's the goal. And you know, like I said, it's, it's gonna be, the benefit of this is, is just astronomical when you start looking at, uh, the data that you now have that you can act upon and, and make, make intelligence decisions based on that data. Wonderful.
Well, good luck with the talk today. Good luck with getting the right folks involved and Thank you. Hope that happens.
Steve Taylor, CTO with Deploy Hub and uh, it's gonna be talking about the, uh, hero project, uh, in addition to his all his other work with Alius and other projects. So thanks Steve. Thank you.