Optimizing DevOps Pipelines with RAG and GraphRAG AI Retrieval Models with Stephen Chin | swampUP 2025
AI models are getting tasked to do increasingly complex and industry specific tasks, where different retrieval approaches provide distinct advantages in accuracy, explainability, and cost to execute. GraphRAG retrieval models have become a powerful tool to solve domain-specific problems where answers require logical reasoning and correlation that can be aided by graph relationships and proximity algorithms. Stephen will demonstrate how the agentic combination of RAG and GraphRAG retrieval patterns can improve analysis of dependency and security information to optimize a DevOps pipeline.
Transcript
Hi everyone. We're back here in Napa Valley at, uh, swamp Up Jfr Swamp Up event. It's been a great two days.
We're kind of winding down, but we saved some of the best for last. If you take a look at this guy, you've seen him on Tech Drunk TV before. Um, we've been working with Steven Chin for four or five years, maybe more.
We've seen him when he first came to Jfr, coming from the Java community, leaving Jfr, Neo four J and now back. Well, you're not, you're not a frog, but you're, You're Well presenting. I'm bringing the best of both worlds.
'cause now we can use graph intelligence, uhhuh plus DevOps and solve some real security challenges, supply chain challenges. So I can I call that dev graph Intel ops? Yeah.
Yeah. Let's call that, let's call that, yeah. That'll be off Columbia next year.
Next year. Dev graph, Intel ops. Yeah.
If you can get him to say that you really are a magician. Okay. Um, but seriously, Steven, it's great to have you on, you presented this year here at, at, uh, swamp Up.
But before we get into your presentation, you are a swamp up veteran as much as I am or more even. What'd you think? It's nice to be back in Napa.
Yeah. So the first swamp up was actually here at the Meritage Right resort in 2015. Um, and I remember like super casual, but all of the thought leaders and what probably didn't wasn't even called DevOps at the time, but like, like people actually doing real world deployments and infrastructure, dealing with security issues, dealing with challenges, getting to deployment.
Now you fast forward 11 years from now, we're back at the Meritage humongous event sold out. Yeah. Full audience.
And now we're looking at the same problems, but with the lens of how we use ai right. To solve and to automate and to really like, streamline your DevOps processes, because that's the biggest challenge with AI that nobody's talking about is getting I outta production, releasing ai. Everybody has amazing prototypes, amazing applications.
They have this humongous value, but the quality's not, there are Humongous investment, let's investment Investment. The quality is not there. It's, it is not providing business stakeholder value.
It's not production ready. It's not secured. No.
I mean, this came out, there was a recent study you probably saw from MIT, there was another one I think from, I wanna say from Deloitte, though it might have been Accenture. One said 95% of, uh, AI apps have not affected the bottom line or been classified as not successful. Another one said 80%.
So depending who you wanna believe, neither one of a good picture. But, but you know what, I remember when they said the same thing about DevOps. I remember when they said the same sort of things about cloud.
This is, you know, it take, the thing about AI is it's so much a victim of its own success that to hype me know, went so sky high, you know, Um, Defying gravity to Well, well you said, you said that in past tense. It's still going up. You Think it's still going up the, the height meter?
I, no, I, you know, I'm starting to see a lot of people say, you know, already going down to that trial of disillusionment, right? Y you know, like, so, so when you look at AI in aggregate mm-hmm. Like, like a lot of the early things like LMS and, and models and inferencing, like those, those are more stable.
Like the, the progression is more incremental. But when you look at what people are doing with AI on top of that, where they're building agent systems, they're incorporating like different knowledge sources in their organization, um, they're taking advantage of, of data science and different like layer techniques. Those are still new buzzwords every six months.
New techniques for doing it, like new advancements and what the capabilities you're able to bring. And, um, what start out with chatbots, which are kind of, you know, passe now is turning into business intelligence and, and dashboards and like insights into customers. And there's a whole bunch of real use cases which, um, if if they were production ready, if they were accurate, if they could provide explainable auditable results, it would be amazing.
But there's just a gap in getting to production and, um, I think swamp up and like a conference like this, which is so focused on releasing and DevOps is a really good, um, litmus ground for the latest in getting to production because it's, it's just focused on professionals who do this for a living and they're, they're the ones where all the apps and the companies feed into and they have to make sure they meet the quality bar. They're actually like at a level where you can release them and maintain them and support them. Agreed.
Agreed. All right, let's pivot. Let's talk about your talk here at Swamp Up.
Yeah. So what, what I did here, because it's, it's an audience of people dealing with security issues with software bill materials, with like releases is I used Artifactory as the system of record exported a bunch of SBO M and VEX files and Cyclone DX format. Um, you could also do SBDX.
Mm-hmm. And I fed those into a knowledge graph system where now you're using an LM to take all of that information and construct a knowledge graph, pull in a bunch of insights from the data, and then create these connections. And so when you, when you ask like an lm let's say you're like a security scenario, you're like, well, you know, in this library, what, what sec secure security vulnerabilities are they, how exposed am I, yada yada.
It will, it will tell you a wonderful story, very, very long-winded. And like, like it'll find similar things, similar security exploits, like similar production issues, but it's not very relevant to your system. No.
Like, is it a library you use? Do you even call the API which matters for it? Um, so what knowledge graphs are really good at is grounding.
And so they take that information, they encode it into a, a knowledge graph and a knowledge system. And then what you do is you tell the LM answer from this knowledge graph. And I, I did it two different ways.
One is, um, uh, it's called, um, doing a a, a graph vector search with graph enhancement. Mm-hmm. And you first acts, uh, vector database, um, NEO four J also has a vector store for the answer, and it does similarity searches.
So it gives you back related information, but not very pertinent at all times. And then you then pull some of those nodes out and you say, what nodes are similar to this? So like, if you find the particular security exploit by the first search, now you'll pull in all the libraries, it's nested in the authors of those libraries, the systems it's deployed in.
You pass that as context to the lm and it does a ver a more precise job of answering. And it's also very fast because the vector search returns immediately, the graph look ups quick and you get back a very quick response. The second thing I did, and this was, um, new this year, and I think this is the future and why people are investing in some of the new AI technologies, is I stood up an MCP server.
Um, so I used cloud desktop. I deployed the MCP server as a DXT file to the local desktop. So super easy.
We, we have an open source, um, cipher. Cipher is a query language for graph to, um, um, text decipher MCP agent. And I gave it the same database, the same knowledge graph, but then asked it to solve the question.
And this time it wasn't doing a vector lookup at all, but the agent was using the tool to help answer the question. So first it retrieved the schema of the database and I saw, okay, well, you know, you're asking about a library now I see like that's a package. Now I'm gonna ask about all the vulnerabilities, which related to that package.
So I did a query. I was like, okay, well you asked about how the vulnerability applies to my application. So then it dug in deeper which applications were deployed that used it.
And after a couple round trips where it was querying the knowledge graph and building it out and without any, I didn't need to write queries, I didn't need to optimize the flow. It it navigated this. So yeah, it came up with basically a very detailed report on, you know, this is your application, this is the risk areas, this is the things you should investigate, here's all the information I know about it.
And it's, it's the same sort of research like we would do as security researchers. Yep. So let me ask a question though.
'cause one of the beauties of SAM is that they're not static as the release you are using or the, the component that's in this piece of software as that component we find out about vulnerabilities in it. The, there's a new version of the component, whatever SBOs are supposed to be telling us all that. Does that kind of, um, not portability, but automated updating, does that live in the graph as well, Doug?
Yeah, so the, the nice thing about graphs and, and like graph database technology is, it's been around for a long time. So like doing updates of the graph, like doing transformation of the graph is all a pretty solved problem. Yep.
Um, so you can continually update the graph, you can use. Does it continually update itself? I guess?
Well, my, my question, my demo didn't, well, This is just a demo. I mean, yeah, Yeah, yeah. But you, you, you can basically set up an automated system where as you make changes, it'll update and it'll pull, pull the entities out, update the graph, and then the other thing which, um, graphs are very commonly used for is removing data silos across the organization.
So my, my use case was, um, all the data was an artifactory. So theoretically, like this could be a product capability artifactory, but what if you also need to cross reference those security vulnerabilities and the, the applications against like another database, which is our, you know, our known mitigations for different vulnerabilities. Maybe you have like a another application list, which is our, where all the applications are deployed to different environments, what hardware they're running on.
And now you can use the graph to pull all this information together, have it be the system of record which gets, you know, updated and managed from all these different systems. And then you can directly derive value by building dashboards, by building query interfaces and things on top of the graph database. Absolutely.
Um, exciting times, huh? Exciting times. How do people, the regular people, and keep in mind our audience are not regular people.
Our audience are our people, right? They're, they're techie people, they're developers, they're DevOps engineers, they're platform engineers and security folk and, and so forth. Is this beyond them, Steven, can they do this themselves?
Is someone gonna come along and wrap this into a product or SaaS or something? Yeah. So that's, that's interesting.
Now, I think the point we're at in AI evolution is anybody who tries to sell you a, like a platform or a package solution, it's, it's never gonna provide the business value you need for, for your system and your use case. Now, on, on the flip side, it's never been easier to roll your own. And I'm not even talking about vibe coding.
So literally my demo was, um, I created a knowledge graph using an LLM and I used a prototype web application, our knowledge graph builder, it's open source, it's free. I threw the documents in it, connected to a free or a database that's our cloud database. And I have my knowledge graph built without running a single line of code.
That's what people want to hear. And then for the MCP server, so of course you could, you could do it, you know, a docker deployment. Yeah.
It's like, do all the infrastructure and do all the port configuration, yada yada. I didn't even bother with that. I downloaded Claude desktop, took the open source MCP server, which is, you know, in a packaged format and I added it as an MCP tool to Claude configured a few like URLs and usernames and passwords for the database, and then I could query it.
So this is something anybody can do, and it's really lowered the bar for, um, people who are technically skilled. Mm-hmm. Right?
They, they understand the business domain, they understand the requirements, they understand even like, like how to architect systems, but you just don't have the time to, to build and maintain a, a large code base to do a specialized application. The, the toolings got to the point where you can take off the shelf MCB tools, you can take some technologies like graph databases and you can compose a very custom tailored and productive system for use cases and scenarios you have internally with, you know, in, in a, in a quick hackathon project, you know, 24 hours a few days with a team and you have like a, like working system. Fantastic.
That you gotta love. I mean it's an in, you know, they say you may, you live in interesting times. It's crazy times to be living in crazy times.
Hey man, we're about outta time we're gonna bring on, I think it's gonna be our last, uh, swamp up. Steven, it's a pleasure seeing you. Say hello to Cassandra.
Jen got Steven's just, he's almost the precursor, but you know, chin two oh is Cassandra. Check her out. She's doing amazing things too.
We're here at Swamp Up. I think we've got one more great one for you and we'll be back.