Gremlin Foresight AI Tests and Fixes Reliability Risks
AI reliability testing is moving from a pipe dream to a closed loop. Alan Shimel welcomes back Kolton Andrus, Founder and CEO of Gremlin. Furthermore, Kolton explains how the new Foresight AI finds, tests and fixes reliability risks with a human at the helm.
From chaos engineering to reliability management
Kolton built reliability tools at Amazon and Netflix before founding Gremlin. In addition, Gremlin started as an expert level SRE tool for chaos engineering. Over the past decade, it became easier for everyday engineers and leaders to use. Consequently, the platform now drives both technical and cultural change across organizations.
How Foresight AI works
Foresight AI analyzes a system and identifies what could go wrong. Meanwhile, it runs the tests, interprets the results and explains why something failed. It then hands engineers a diff or code snippet that fixes the issue. Therefore, teams can rerun the test and prove the system is more reliable.
Foresight draws on millions of experiments across Fortune 100 and Fortune 1000 systems. As a result, Kolton wants the recommendations to be as dependable as the systems themselves.
Why verification matters in AI reliability testing
Kolton expects more code, more bugs and more outages as AI speeds up development. In addition, he says many remediation tools guess at a fix without checking the root cause. Consequently, closed loop verification will separate real fixes from knock on effects.
Alan compares today’s caution to the early days of automated intrusion blocking. Furthermore, both agree that confidence grows only when results can be proven.
Available now
Foresight AI is a standalone product built on Gremlin platform data. Meanwhile, it spent six to eight weeks in beta with existing customers. It is priced per team with generous usage limits.
Teams exploring AI reliability testing can start a trial and ask the built in agent where to begin. Explore more DevOps coverage and the latest Techstrong TV interviews.
For more information please visit gremlin.com
Transcript
Hey everyone, welcome back here to Textron TV. Our next guest is my old friend, Kolton Andrus. Colton, for those of you who don't know, is the CEO and founder of Gremlin, one of the leaders going way back into chaos engineering and stuff.
But so much has happened and there's so much water under that bridge since then. Colton, it's great to have you back on Textron TV. How are you?
Doing well. Thanks for having me, Alan. It's always a pleasure.
Absolutely. We're going to jump into a lot of the new things that you guys are doing, but not everyone watches every Textron TV interview. I know hard to believe.
But for those who aren't familiar with your story or the Gremlin story, why don't you give them a quick update? Yeah. So I'm an engineer by trade.
I've been focused on reliability of distributed systems for the last 15, 20 years. I did some great work at Amazon, some great work at Netflix. We took that work and founded Gremlin, and our goal is to help build a more reliable internet.
And we've done it by going out and proactively identifying and testing the failure scenarios that bring systems down in order to prevent outages. And we built a kind of expert level SRE tool in the beginning, had all the knobs and dials, could do everything you need it to do. And then we really evolved it over the last decade to make it more approachable for your average engineer to have that visibility and tracking the leadership needs to be able to measure, improve, and improve reliability, and drive not just the technical change, but the social change across organizations.
I love it. I like the way you did that one. com?
com. Yeah. ai, so either one of those will take you to the right place.
It's the world we live in today, right? That's a testament. Colton, I want to jump right in.
I think we gave people enough background. They can go check it out for themselves. But I wanted to jump in and spend the bulk of our time today talking about what's new with Gremlin.
You guys are rolling out something here called Foresight AI, and I don't know if it's curing cancer as Dario Amodei says AI may do, but it does just about everything else, right? Tell us about it. Yeah.
Well, one of our visions in the beginning was make it easy to do the right thing, and how much can we just do for people? And one of the things we found when it comes to building distributed systems is there's a lot of little details you've got to understand and figure out in order to accomplish it. People are busy.
They're shipping product. They've got to manage their AI agents now. The world's changing, and things are going faster.
And so you'd think we'd have more time to go take care of some of these important pillars, but we don't. And so Foresight's really about how much can we do for you? And the name of the game is, let's show up.
Let's analyze your system. Let's tell you what could go wrong. Let's tell you what risks exist.
Let's go run the tests that validate those, and let's see how the system responds. Let's interpret the results of those tests. Let's tell you why it failed, and then let's tell you how to fix it, and let's hand you a diff or a code snippet that actually corrects the issue that was found.
And then finally, we can close the loop. We can go back and run it again, and we can prove that we found the issue within the system, and we fixed it, and the system's in a more reliable state. And all of this is just being done human in the loop, human at the helm, but not human having to do it.
Yeah. " If we're going to go patch your Kubernetes config or push a change to your source code repo, we want to make sure a customer feels comfortable with that, and that's where people are at today. People are like, "That sounds great.
" Maybe down the road, I'm sure, people want, we'll give them the, "Here you go, go fast" button. But yeah, what it takes is, we've had all these SREs that really understand this distributed system testing we've been doing for the last decade, but it's hard to scale that across the whole organization. And reliability is not a one-team problem.
It's an everybody problem. And so when you go to a team, an engineering team that's built a new product, or maybe they're having their agents build a new product, have they done the testing they should to make sure that it's really in a solid place? And they often haven't, and they're often not sure where to begin.
They're not sure exactly what to test. They ran the test. They're not exactly sure if it behaved correctly.
Okay, it didn't behave correctly. They're not exactly sure what they should do next. And we're able to leverage this last decade of knowledge we've built.
We have millions of experiments we've run. We've seen tens of thousands of these Fortune 100, Fortune 1000 systems, and we've distilled that down more than just on the LLM side, but on a machine learning side, into a set of real provenance that we baked into the system. Because one of our goals is having it be actionable and credible.
And the way I think about it is, we want nines of availability in our system. I want nines of availability in the accuracy of what the systems are recommending and suggesting we do. Agreed.
Let me ask you a serious question, just between you and I. No one else is listening. When you were first working back in the day, even before Gremlin, chaos engineering was a new thing.
This whole idea of testing this kind of thing was new. Did you ever envision a time where it would be so automated using AI like this that, once the human signed off, it just ran and did it, and not only just ran the testing to see what broke, but also fixed it as well, just automated the heck out of it like this? Well, not from an AI perspective.
Where AI was 20 years ago, that felt like a pipe dream. When we'd originally envisioned this idea of Gremlin, we thought it would be kind of an autonomous system in the background that just went and ran and probed without a whole lot of oversight and found issues. And so that's the way we built it originally.
And a bunch of people were like, "Whoa, don't do that. " I think 10 years ago, when I founded the company, I was like, "Hey, you've got to push for a large vision. " And at the time, that was really hard to do.
That was something that felt very far away. " Mm-hmm. And I think that's been super exciting for me, super exciting for my CTO.
We're skeptical engineers. " And you know what? There is a way, and it's pretty awesome to see it come that far so quickly.
It is. I mean, quickly. You know what?
Ten years, Colton. But you're right. It does feel like if I looked at that 10-year period, a lot of the, especially when it comes to the automation progress and everything else, is in the last 25%, the last two and a half years.
And it's continuing to accelerate. Yeah. Which kind of brings me to my next question.
Where do you go next with this? Yeah, I think that's what's interesting about where we're at today. For many teams and companies, we're just trying to keep pace with the innovation that keeps blasting forward.
That's part of where I think there's a lot of good for Gremlin in the future. We see people shipping more code. We see people shipping more bugs.
We see more outages happening. I think we're going to have a year in the next year or two where we kind of fall out of the enlightened era, and we see a lot of the trough of disillusionment, and there's a lot of bugs and issues and things we've got to figure out and fix. But really it's really around that automatic remediation, and how do we get to the point where we take the human fully out of the loop?
Can we get to the place where we do all the analysis, we do all the testing, and we do all the fixing in the background so people can focus on whatever's most important to them, and the system can fade into the background as just a very reliable, robust platform that they operate on, and that their applications behave the way they expect. Agreed. I think this is an important thing, though, too, Colton, is we haven't reached the end of time, we haven't reached the end of evolution here.
It's still the beginning of the beginning, not even the end of the beginning of this journey that we're going on. And so it is going to be interesting to see where we go from here. As you mentioned, people are still a little hesitant to go full speed ahead with remediations and automated remediations, automated fixing, automated changes.
This is not an AI thing. I saw this when I started Still Secure in 2001, 2002. We were trying to get people just to automatically block god and variety intrusions, and they were scared to death about doing it.
They would turn on blocking for a small fraction of the things we could find and block. It was the same thing with automated vulnerability remediation. Right?
They had to test, make sure those patches work, and all that good stuff. So it doesn't surprise me that we're having a similar sort of journey with AI automation here. But at some level, you got to just let the chips fall.
You're going to go ahead. I think it's a question of building confidence in what you're doing and how you're doing it. Yeah.
That's a big one for me. I'm a pedantic engineer. I'll take some risk.
I don't want to take unnecessary risk. I want to have a high degree of confidence in what I'm doing. And I think that's a great angle for us as well, because a lot of the remediation tools out there, first of all, my CTO thinks they're more like 20/80 solutions, not 80/20 solutions.
So we got a little bit of room to go before we can really fix all the things. But a lot of them look at what went wrong, look at the metrics, look at the logs, make a guess, and then fix something. But it's not always clear, did they fix the root cause or did they fix a knock-on effect?
Yep. And our approach is really this closed loop where at the end we verify that the thing is working the way we expect. And I think that verification in these auto-generated systems is going to become massively important.
That's what's going to allow us to move from the Wild West, moving fast and breaking things, into the kind of production systems and digital infrastructure that we're going to need to rely upon as a society in the coming 10, 20, 30 years. Agreed. Hey, I'm going to turn us another way here, if you don't mind, and get down and dirty in the weeds.
Is it available now? How do people get their hands on it? Yeah.
com, and you'll have access to the new capabilities. So you can jump in and start playing with it. ai.
That'll give you some intel, some details about how it works and how it fits together. But one of the great parts about this approach is if you're not sure what to do, just ask the bot. Just ask the agent, and it'll tell you what to do, and it'll guide you in the journey.
And I think that's what's so exciting to me is a lot of people, "Hey, I'm not really sure what to do. " Well, hop in, ask it some questions, ask it what you should do, and we've put a lot of time and effort in the ways we build our harness, our tools, and our skills so that it's really effective. It's really useful in how to guide you to the right answers, guide you to the problems, and help you fix them.
I love it. Would it, existing Gremlin customers, this is part of their package? Or is this like an add-on or a new product?
It's really, we built it as a new product because we wanted it to be- Different ... standalone. It's a new set of capabilities.
We rely on a lot of the data within the Gremlin platform, so it obviously fits very well with the systems we built. We've had it in beta for the last six to eight weeks with our existing customers. Some of my favorite quotes, we have a few customers who say, "I can't live without this now.
This changes the way I do this. This makes it so much easier. " You have got to- So we've got it in their hands.
We had them battle test it. Our team, we drink our own champagne, as Jeff Bezos said when I was at Amazon, and so we've been using it ourselves for the last couple of months. And so we got a high bar, and we wanted to meet that high bar, and now we're ready to put it in the hands of people and let them put it through the ringer.
Beautiful. How's it priced, Colton? We're pricing it per team.
So a team will come in, and we'll turn it on, and they'll able to use it. The price of tokens is changing on an almost daily basis, and so we're really structuring it in a way that we want you to have enough capacity to go do all the things you need to do. So we're setting very high limits, very high usage, and we're taking a page from the Anthropic and OpenAI books to give people just enough capacity to go do all the work they need to do, hopefully without hitting any kind of ceiling or limits.
I love it, man. Hey, Colton, we're about out of time, but I want to thank you for coming on here. I wish you all the best of luck with this.
I'll tell you, man, I think back to all the times you and I have spoken. Right? It's got to be close to 10 years now.
And in some ways, we've been sitting here watching this whole industry kind of just keep evolving, keep moving, keep changing. And you've managed to stay right in the thick of it and keep Gremlin right on the edge, which is what you should be doing, and good work, man. Keep it up.
Oh, I appreciate it. Always a pleasure, Alan. All right.
Kolton Andrus, CEO founder, Gremlin, here on Techstrong TV. We're going to take a break. We'll be back with more.