Yunhao Jiao on the DevOps Bottlenecks Emerging in the Age of AI
TestSprite CEO Yunhao Jiao dives into the bottlenecks that DevOps teams are going to encounter in the age of artificial intelligence (AI)
Transcript
Hey guys, thanks. The throw, we're here with Y How Chow, who's the CEO for Testrite. And we're having to chat about how AI may be finally breaking up some of those DevOps bottlenecks that have been plaguing us all these years.
You and Hal, welcome to the show. Thank you, Mike. Yeah, nice to meet you.
Yeah. So walk us through your thinking here. I mean, we all know that AI agents are coming, or in some places they've already arrived, but we have all these bottlenecks and DevOps that we've been struggling with for all these years.
Can we finally break them using AI agents? Uh, yes, for sure. Um, so, but I, I will say it's still kinda like a progress, um, at this moment.
So, and everything actually is, is changing so rapidly in the past couple years, as we know. Um, agents like Cursor coding, agent like Cursor, GitHub, copilot, uh, tray, and all those, uh, coding AI help developers actually write code very fast, and they can sometimes write thousands lines of code within just um, minutes. Uh, but this also create actually new problems.
Sometimes we'll find bugs are flying everywhere. Sometimes we also find AI probably just riding too fast so that you cannot really keep the pace with it, review all those content manually by yourself. Uh, when you generate tens of thousand lines of code, million lines of code, it's almost impossible to automatic to review by human beings.
That's new, kinda like a bottlenecks or what do we call new issues or problems raised by those new agents. And test by actually was created to solve those new bottlenecks and to make programming smoother or easy Again, for developers, uh, we are trying to use our agent, uh, AI testing AI to help to automatically validate software or code autogenerated by those AI coding agent like Cursor or GitHub copilots. We can be directly installed into programmers, IDE, so they can install our MCP directly into their IDE with just a one click within minutes, and then they can just use natural language prompts in their cursor chat bot or maybe AI coding IDs chat bot to run test Bryan, and something like, Hey, can you test this software or project using test by MCP?
And then it's all done. So our AI will automatically run all those steps, including analyzing their code base and learn from their software, read their existing, um, product requirement documentation, generate a test plan, generate test the code, run test cases, do test result analysis, and share the final results with customers. So all these, so, So I, so I get that we have new bottlenecks that are emerging in the age of ai.
I guess, you know what I was curious about and can we use AI to replace or fix the legacy bottlenecks that we've been struggling with? Or are we just gonna start piling more code through those bottlenecks and essentially hit the wall faster? Got it.
Got it. So, um, it's actually, um, so I will say it's actually, uh, interesting progress that AI agents is like, actually lots of AI agents is actually already fixing legacy bottlenecks in the past, I believe, five years, especially for programmers. I, I used to work in Amazon for more than five years.
The biggest bottleneck when I was a programmer probably is the coding efficiency like this, how whether we can write code that fast, whether we can actually ship things, um, way faster, and meanwhile keeping the quality, because previously, if it's a human being, you can probably only write 200 lines of code, 300 lines of code in a day. So at that time, the main bottleneck is how fast you can understand a problem, how fast you can propose a solution, implement it, how fast you can end to end test it, make sure it's really good running successfully, and then deliver it. So at that time, everything's probably about speed.
And, and right now I think, yes, existing agents and, and all those, um, like AI indeed solved that problem. The pace is great, everything was solved within minutes. They can generate hundred thousands of lines, like we mentioned before, within just 10 minutes.
So speed is really fast today. This bottleneck in the past of three, five years is successfully solved by agent. But meanwhile, when they're solving ex legacy problem on bottlenecks, they're creating new ones.
For example, uh, the code quality, for example, those hallucinations introduced by found foundational models because I, I know OpenAI cloud code, they are doing fantastic jobs. They're releasing like models almost every year, GPT-4, GPT, five, cloud three, cloud four. But, uh, those models still are actually making some hallucination issues or arrows.
Uh, if you look at their performance on SWE benchmark, you will notice that their accuracy is only about 60%, which meaning that 40%, they make mistakes. Uh, and that is already the state of art per like, like the best model among all those existing AI models. So 40% basically meaning you if you have like, um, 10 features that you want AI to implement, probably six got successfully implemented, but four maybe partially still have some issues or didn't fully meet your requirements.
And that is the new bottleneck created by these agents. So especially for startup like us, what we do or what we focused to do is solving the new bottlenecks probably doesn't exist before five years ago. So this is actually a quite interesting change which is happening in the industry In a certain perspective.
It seems like the bottlenecks are just shifting and we can write code faster, but the problem is if I get to the back end of the workflow and then that a lot of that code gets rejected because it's too verbose, there's too many vulnerabilities, and the overall amount of technical debt starts to increase. Doesn't that defeat the purpose of kinda having these AI coding tools in the first place? Because now I got all this code that I'm just sending back to have redone anyway, probably by human.
You are right, you are right. Um, be, this is actually what's happening right now in the industry. Um, because we work with, um, almost 50,000, uh, enterprise customers in the past one year, and we observe that lots of things are happening to them, uh, especially when they're adopting the, these kind of like AI coding agent at the first, they were surprised by their speed.
Uh, they find that one AI to some degree, their speed can replace maybe 10 engineers when it's generating code. So they were, they were surprised. They would say, wow, it changed everything.
So they purchased lots of AI tools. They start to use all those tools, uh, to write a code, and programmers actually have more time to, to drink coffee or, you know, just, uh, free their hands, but probably focus on designing, uh, customer facing problems, which is good. But later they find out, seems like they cannot fully count on ai.
I'm not saying AI is not good, so that we, we, we cannot use it at all, just the degree, the balance. So you can for sure let cursor, let GitHub copilot and all these tools or cloud code to help you to create the first draft version, but that version definitely cannot be directly released to a customer if it's enterprise level feature or enterprise level software, because as you can see, 40% chances they will, they will probably create some hidden issues, unseen bugs or hallucinations somewhere. So right now, existing programmers job is on how to identify those issues, how to manually fix those issues.
So I would say statist, uh, speaking, their job efficiency right now indeed got increased. So whatever takes them, like for example, a week before to ship to finish today with the help of cursor, GitHub, copilot, all these popular coding tools, they probably would get it done for two, three days. So half of their time and half of their efforts probably is already saved today, although they're still fighting with those ai solving those hallucination issues, solving those bugs, but indeed it saved them some time.
Uh, due to the efficiency of the ai, our tool is basically helping them to say, can we save them some extra time? Helping them to even be better to faster shift the confident ship, the software with more confidence, make it even like what used to take seven days, five years ago, and right now only take half a day so that they can be even 10 times faster with the same quality than before. Y yeah, that's kinda like, um, what I'm trying to, to share.
Yeah. So yes, basically, uh, in short we can say, um, AI is indeed helping people or already, if you look at how much time, how much effort, how easy it is right now compared with five years ago, but it just like, it's still not perfect, uh, still not probably what people thought, uh, it is today. Yeah.
So I understand that we can clearly write more code in a day than we did before. And, um, but we have to be smart about this, I think because I cannot ask the AI that created the code to review that code, and I need a different model to kind of look at that, and then I need some sort of ability to reason across it and judge it. But so do we need an an an entirely new way of thinking about our DevOps workflows because there's gonna be multiple AI agents that need to play off each other?
Yes. This is actually a, uh, this is actually something I really want to share. Yes, we are actually, uh, sharing with actually developers worldwide, lots of startups, enterprises that the DevOps and also the whole development habit is probably changing today.
So especially on the testing side, we are encouraging something called left shifting, left the shifting. Uh, if we wanted to easily understand it, I probably wanted to share something about my background when I was working at Amazon, uh, five years ago. If in those big tech companies, the typical, um, DevOps kinda like, uh, cycle is you write a code first and as the programmer, you do some unit test to making sure that you don't make mistakes among the code.
So all the logic is fine, and then you create a new PR code review, submit to your team to do some code review so that other team members help you to quickly review it, making sure that your design, your logic, your implementation on the human review also makes sense and doesn't make too many obvious mistakes. And then you merge the code to, um, your dev endpoint, to your, uh, kinda like beta endpoint and then let the QA team, the testing team to do some manual testing. They'll play with the feature mouse clicking around record the whole scenario and trying to tell you whether this code under the more kind of like production endpoint is working correctly or not.
So they will let basically bunch of manual resources to use your feature, um, mimic the customer's real behavior, trying to tell whether everything's all fine. This is what we call integration testing or end-to-end testing, and that including sometimes ui, front-end UI testing, sometimes backend API testing and all these kind of testing. So we do have different kinds of testing even when we talk about software testing.
And today when we talk about left shifting, actually this is even, um, uh, recommending from injury, uh, from deep learning, uh, deep mind and left shifting, basically meaning we don't have to wait for the development to be fully finished and then do testing. We can even do testing while we are doing development. And to some degree we can even do test driven development, basically, meaning you can first design those test cases ahead.
Uh, something like if I want to create a feature a, then basically I need to have these five test cases for feature A and if the AI correctly implemented the feature a meaning the AI has to at least pass all those test cases, otherwise it doesn't mean it successfully did the job or finished my requirement. So basically they have a high level design doc and then they immediately design some testing associated with those features, and then they do implementation. The implementation will be fully finished or marked as finished only when all those tested cases are green or passed.
So that at that time, human beings and also ais both have confidence, I did great job, this software at least satisfy all these requirements and it can be probably moved to the next stage or maybe for beta users to give a try or something like that. Yeah, so this is what we call after shifting do testing earlier that for example, when you generate the sum code using cursor, using GitHub copilot, you can already give it a test wrong to see because we know they are making mistakes somewhere, we just don't know where. So we can make a testing wrong or execution at that time.
And when we notice, okay, so it, so 70% are right, but here, there this place we have some issues and then we let a cursor iterate again, fix all those issue, and then we run test again, and then we find remaining issues and that keep in a loop until everything is great. This is exactly what test Bright MCP is doing right now. So we don't engage like traditional QA tools at the end of the day when everything from the developer side is finished, because that basically meaning developer has already fixed all those, um, cursor issues.
Uh, they manually spotted them and fixed already using prompts, but we wanted to engage earlier when they're coding. So every time when they use Cursor to generally something they can immediately run test Bri to spot or to validate whether there are issues and they can keep running it, keep those two AI agents working together until all test green and then maybe create a PR for other team team members to review or maybe deploy it to their dev and point, uh, for some beta users or internal members to give a try. Things like that.
I feel like though there's something wrong with our little human condition in all of this, and people will see that they can write more code and then they'll just continue to write more code and they won't think about spending more of the time they freed up on testing and improving the quality of that code. Yeah. Yeah.
So, uh, if it's, uh, actually a junior level of engineer, or sometimes if, if you, if you first time you using, uh, tools, uh, you'll, you'll feel that way for sure. Uh, but with your software become more complex, with more of your customer complaining, Hey, I have this issue, oh, I encountered that issue, they will immediately realize that the quality of the software is not enterprise level and they cannot afford to maintain that kind of software to their customers because they will, they will lose customer. Interesting thing is today with the rising of tools like Cursor, GitHub, copilot, and all these coding agents, the barrier to to, to create a software is very low.
So almost everybody, when they have an idea, uh, this is why we call it a vibe code. So when you have a vibe, when you have an idea, you can already create something, um, by yourself or with a very LinkedIn small team. So the resources needed to create a software is is not that big, big or huge.
Everybody can almost create a software when they have, uh, I idea that means people with the same idea probably a lot. And there are probably a lot of similar products in the market because you, you, there are definitely lots of people with similar idea, like you, they'll also create a software, they will also release it. These, all these changes also lead to the fierce competition.
Right now in the AI sector, lots of, um, industries, lots of, um, like different attracts, we can all see d similar products right now in the market because of, because creating a software is so easy today. So the main, I would say the main factor for you to win a competition among your competitors is naturally right now becoming whose quality is better, whose AI is more accurate, whose software is more user friendly, who whose AI is faster and, and more, maybe more affordable. So everything is about the, the, the software is user experience inequality.
If you can become the best software among your competitor, you naturally just wing it because idea, vibe is cheap right now. So every, everybody with an idea can, can somehow generate something. So people will immediately realize that.
And what we can see is people already realizing that and they're trying to pay more attention as well as money, uh, into the software testing world. Uh, this is more driven by their revenue, I think by the competition, by their customer's, feedback by the market is, um, current situation. So I believe in the future, people will just pay more attention, more and more attention in the future due to the natural, um, uh, I, I would say the competition, uh, in the market so different, doesn't matter what what the industry is, doesn't matter what things they're building, as long as they want to win the competition in the future, they definitely want to focus on quality.
And, and, and right now, the only way to do that for sure is, um, is do more testing, making sure your, your, your software is robust and, and, and your platform is stable, reliable, under all kinds of conditions for any of your customers. Yeah, Folks, you heard it here, the choice is clear. We can either write more bad software faster or we can take some time and focus on writing better quality software that people actually use and enjoy.
Hey, YHA, thanks for being on the show. Thank you so much, Mike. Yeah.
All Right. And back to you guys in the studio.