Ciroos.AI CEO on AI Teammates Assisting SREs in DevOps After $21 Million Funding Round
Ciroos.AI CEO Ronak Desai following the raising of $21 million in funding explains how artificial intelligence teammates will increasingly handle DevOps tasks that previously required the expertise of site reliability engineers (SREs).
Transcript
Hey guys, thanks for the throwaway here with Ron Desai, who's the CEO for Sierra Rus, and they've just picked up $21 million in funding to develop and I team made for SRE and DevOps engineers. And we're gonna get into what that means and how maybe DevOps is evolving. Ron, welcome to the show.
Thank you very much, Mike. Great to be here. First time on, thanks, uh, tech Strong TV here.
Everybody, of course, by now, unless you're living under a rock, has heard about something about AI and AI agents and how that might be, uh, changing the way we build and develop software. But, um, how do you see all this coming together and, and where are we on this journey? Yeah, so Mike, um, if we just take a sort of step back and look at the, where the industry is, right?
Especially looking at the IT operations, uh, in a bigger enterprise today. The investigations when you have outage on applications takes a long time. Hundreds of experts has to get involved.
You know, you are to go through a whole bunch of dashboards. It takes hours on an average, uh, to resolve the incidents. And so that's basically where we are coming from, right?
So SROs ends all of that. We want to mitigate and really build a teammate for our SRE and DevOps teams. It's a multi-agent tech system, uh, built with the model of experts.
So we can really target different, um, domains, which your application depends on. So if you think about like today's modern applications, right? It depends upon a lot of different components.
It's not typically people would talk about saying, oh, it's a Kubernetes based applications, but it's deployed on-prem or deployed in the cloud. Uh, it has dependency on internet, it has dependency on your infrastructure, right? So now if you have outages on the application, it depends upon all of those components.
So that's basically where we are building this capability, which is allowing you to sort of reason as a human would reason across all of those domains and really find out those insights. That's basically where we are focused on it. Um, as an industry, I think every domain is getting disrupted and as you know, right, every, uh, I think there is a talk and a lot of CEOs, uh, the big public companies have talked about it, that we would have 10 to hundred agents helping us do our regular day-to-day mundane task, which we should delegate it.
Those are tasks which nobody wants to do it. So that's basically how we sort of view it, and that's where we got started There. Of course, there's a lot to unpack there, but one of the core questions people seem to have about these AI agents or teammates or whatever we wind up calling 'em, um, is do I need to have that come as part of the platform that I'm using, or will there be more of a horizontal approach where it's an an overlay for lack of a better phrase of, and I don't have to change any of my underlying platforms to take advantage of it.
Yeah, I think great insight there, Mike. Like, so the way we kind of look at this is you enterprises has already deployed a whole bunch of ity tools. So the way approach which we are taking it is you don't need to replace any of your observate tool, any your, your cloud monitoring tools or your, um, uh, ticketing systems or your incident management system.
What we do it is think about your human SREs and DevOps teams. What do they do? They use all of those tools including collaboration tool like Slack, right?
So when you have teammate, think about it that you've got this experts who has expertise across all of those domains, works with all of those tools, extracts the right set of informations reasons, like the way your human SREs would reason it and kind of get to the root cause. That's basically how we kind of think about it, right? So you're, you're sort of think, think about it.
You're sitting at just into all of those tools, which you already have it, and really trying to sort of get to the really the important aspect of how do you reduce meantime to repair, how do you, uh, help your SRE team to really reduce that to, uh, because there's a tons of work they do today, which is just tedious and manual, and nobody should be doing that work, right? So that's basically how we kind of think about it. Do not replace the tools, sit next to it.
Is there one agent that's kind of an expert in all these things or is there really a bunch agents that each knows a particular tool? And then there's something that feels like a planner that kind of manages all these other agents. Only Cyros has built this capability where we have expert in all of those domains, which we talked about.
One agent for Kubernetes, one agent for cloud, one agent for your network, one agent for security, and then all of these agents are sort of working together to really figure it out. Because anytime when you have an application outage, the biggest problem talking to hundreds of enterprises, what I've heard is they may have expertise in one particular domain, but if somebody made a change in another domain, they're completely blindsided. They don't even know where to start, right?
So this is where that cross domain ability to go across all of that, investigate that vast amount of data, uh, and really figure it out. Where the problem is, is where the crux of the, uh, problem here is. And so you, to your point, right, it's multiple agents working together, uh, for this cross domain and multi-domain is something which is so critical, uh, for, uh, bigger IT organizations.
Will these agents wind up talking to other agents outside of your platform? And will there be some sort of collaboration and interoperability across all these different agents? Yeah, so, uh, you, you I'm sure heard about a to a agent to agent and MCP.
Um, and so what I've seen, Mike, and this is a very interesting phenomena, which is happening in every company out there, um, which is building agent for their own set of capabilities, they have it, right? So take take an example of if you have a network vendor, they're building agent for networking. If you have security vendor, they're building agent for security.
If you have a storage vendor, they're building agent for that, right? So we are literally creating agent silos, right? And this is where we come into the picture.
We believe that we need to be building this ecosystem where if somebody has built a, a agent, uh, for a specific domain, we should be able to easily integrate with a to a internally, of course we use MCP to talk amongst our own agents, right? So we build agents for our own, uh, set of capabilities. We have built it across all the domains, but we will easily integrate with a two, a interface if somebody has built a much more deeper capability for that particular domain, right?
So that's how we kind of think about it. Um, it's agents working together to really solve the problem, uh, for our operations teams. So will this become the primary mechanism with which we engage various DevOps platforms?
And I'm asking the question because today, you know, there's all these pipelines and they kind of span a bunch of different tools, and each of the tools has a slightly different user interface or sometimes a CLI and am I gonna get to like just one consistent natural language interface now and all those other things become, you know, backend services? Yeah, no, and I think the, um, uh, the whole, uh, we used to talk about UI and ux right now. People talk about ax Yeah.
AI interfaces, right? So this is where, you know, in reality, uh, we as a human best communicate, uh, by talking or by entering what our exact questions are in a natural language, right? What has happened in the industry because of the lack of technology, we developed this bunch of dashboards, well, very pretty looking dashboards, but we put the burden on the humans to really click through those dashboards, right?
So really one is to sort of build and the interface, which is very easy for humans, but also sort of elevate it instead of giving them very fragmented view of your infrastructure, your application landscape, really ask higher layer questions to get to where you need to be, right? And so we look at it from the whole, uh, AI usage perspective. There's a part of augmenting the functionality of, um, uh, SREs and DevOps, and then there's a part about autopilot, right?
So because SREs do a lot of work where they're doing proactively investigating it and just making sure that those low signal alerts don't get, get unnoticed. So there's a part of augmentations AI can help by a natural language, but then there's a part of autopilot. And I think I kinda really like this Tesla analogy, uh, because if you think about it, uh, 2015, I was one of those early, uh, adopter of Tesla, uh, autopilot.
And then the vision was it's going to be a full autopilot, but it was a mere, uh, auto cruise control and a lane change, right? Uh, still human in loop, still human has it control, right? After 10 years, we've got to this full self-driving, right?
Where human is still in charge. You still have to be in driver's seat, but it takes care of a lot of those things where if you're super tired, it's really helpful, right? There's a places where way more like, uh, opera operations make sense?
So the way, if you look at it from a, uh, way, uh, enterprises are thinking about, um, really adopting the IT tools. One is the interface, but then also how you sort of place them in, in terms of augmenting the functionality. That's why we really love this teammate concept.
It is really somebody you sort of onboarded it and think about, uh, expert who has a million hours of training on some of these domains. So like having that kind of expertise, uh, is going to be super helpful, right? So there's a, um, human in the loop, there is a full self driving, and then there are segments which is low impact environment.
You would do a full autopilot, right? So that's basically how taking that pragmatic approach is very critical for the AI adoption. How will those AI agents continue to be trained?
'cause the environments will continue to evolve. So how do I keep them up to speed? Yeah.
So, uh, the way we look at it is, um, outta the box when we deploy SRE teammate in our customer's environment, it starts to sort of behave like as if you onboarded a SRE teammate, it has a access to your collaboration environment. So if the investigation fires, it's actually picks up that investigation, starts investigating it across this domain. So there is no learning loop per se.
So at the time of inferencing, it already has the capability, but then human in the loop, right? So human can say, look, I think you investigated this correctly, you eliminated those domains, which was perfect, uh, because it helped me avoid bringing another 10 set of people on the war room call. But this is the area where I would love to investigate, have you investigate in this particular direction or this particular domain, right?
So getting that hu human feedback is something which we will continue to incorporate it as we go forward. Uh, and as we build the capability, but out of the box, it'll start to deliver the value which customers are adopting us for, uh, in order of minutes, right? So there's no learning loop from that perspective.
Are we kind of filing, moving beyond scripts and plugins here to where the AI agents are gonna take care of a lot of these functions that historically a DevOps engineers spent time writing these scripts and then the carrying and the feeding of said scripts? And of course, uh, you know, this is one of the issues that people have with DevOps, and then it doesn't scale because it's dependent on all these brittle scripts. So are, are we counting entering a new phase here?
Yeah, so, uh, runbook and static runbooks, i, I kind of call runbooks and preface it with static because, you know, once it's written, it's outdated because your infrastructure continues to evolve, you add more capabilities in your applications, and nobody goes back and reviews those runbook. Like not many organizations have that discipline. Um, but what if, what if, if you had a system which was thinking like the human had understanding of your environment, you wouldn't need those static grant books, right?
And this is where you come up with this patented technology, what we call it behavior pattern, right? This is the behavior pattern is think of it like if you all got the best experts for all of the component your application depends upon, and it has the ability to sort of eliminate and reason like the way your all the experts will do it, you don't need those runbooks right? Now.
You've got the system which is able to reason like the humans, right? So in this hue, uh, systems can watch your environment 24 by seven by 365. So think of it, you got a superhuman SRE who's watching your system and able to do what your experts are able to do it.
Why would you need static runbooks, right? So I think my take is those era of building those static runbooks era of keeping those up to date things that have been gone, right? I mean, you've seen this across coding agents as well, right?
People are using a coding agents to write some of the software documentations and keeping them up to date. So that's how I think we would evolve, uh, even in the ops era. So You raised the 21 million, what's the priority for that?
What are you guys thinking? What needs to be done next? So, uh, 21 million.
So super exciting. Uh, it was over subscribed round. Uh, absolutely we are, uh, inviting all the talent across all the groups.
Uh, whether it's, uh, talented engineers go to market, uh, product managers, we are inviting them to join us on the journey here. And then of course, we wanna really double down on our GTM accelerate inviting customers to join, uh, us on this journey. Um, early feedback has been amazing.
Uh, so the goal is to accelerate, uh, the journey here. So when you show this to people, what's kind of the reaction you get? I mean, because, you know, on the one hand I can imagine that there's a fair amount of excitement, and then there's also a certain amount of, well, you know, who's moving my cheese, right?
Yes. So, um, and then this is where the teammate concept is super critical, right? So let me just take a step back and kind of describe, right, what all, uh, has happened.
The pace of innovation on the application perspective from a developer's perspective is going through the roof, right? We ourselves get 90% plus coding using some of the AI assistance, right? Uh, when we are building the product.
So that pace has started to go up. So things coming at the SRE are a lot more than what it used to be because now our developers are more productive. Before this phenomena happened, first of all, environment was very complex because the very distributed nature of applications.
But when we pulled enterprises, we noticed that there were, for every one SRE, you had 15 to 20 SREs uh, developers, right? So one SRE, 15 to 20 developers. So there was an imbalance in terms of the amount of work, amount of alerts which were getting generated, and this team's ability to kind of take care of it, right?
So they were already sort of falling behind and we, we were talking to one of our enterprise, uh, financial enterprise customers. 8 million alerts. We are a team of 10 SREs.
There's just no way we can deal with that, right? That's today, now just apply that exploration, which is happening. So they're falling behind.
So we really need to sort of focus on really bridging that gap and give them the same set of capability, which our development team has it so that they can start to really focus on the critical work in terms of architecting, building, the reliable system, thinking about and proactive so that they can add the bottom line to the business, right? So we really believe that this is a part which is augmentation and help, uh, the SRE teams where they feel comfortable to let it run in autopilot environment. This is not about productivity gain.
There is already things which are so manual and tedious tasks, which we want actually want to take it away from their plate and let the system do that work. Um, who wants to sort of query logs, metrics, traces, try to figure it out across four different domains. I mean, it's just too painful.
And to your point, we're already struggling, but as far as I can tell, the amount of code that's being created is also accelerating. So we may be looking at a tsunami of applications that are coming that, um, and we're not adding more SRE bodies to the equation. So the only thing to do is to lean more into ai, it would seem.
Exactly, exactly. No, I think, I think it's, it's all of those things coming together. We really need to build this teammate, uh, for the SRE.
And I think you raise a very interesting point, which is in 2025, if anybody was going to build SaaS application, it's going to be based on agentic ai, right? So there is a aspect of modernizing the practice of monitoring even those applications, right? Which are native agentic applications.
So monitoring them needs a different set of capabilities. And I think building that out of the box from day one is super critical, right? So that's something which absolutely something which we would attack, um, as, as, as we, uh, sort of get into this journey.
All right, folks, while you're hear in ear, hey, DevOps at scale equals a gente, k, that's one way of thinking about it. Ron, thanks for being on the show. Thank you very much, Mike.
All, and back to you guys in the studio.