Performance at scale: Keep up with cloud-native complexity
Tuning performance in cloud-native environments while balancing cost, speed, and reliability.
Scaling cloud-native apps is a balancing act—performance, reliability, and cost all pull in different directions. So how do you keep things fast, efficient, and budget-friendly? In this roundtable, we cut through the complexity to share real strategies for tuning performance at scale without the headaches.
Transcript
Hey, everyone. My name's Alan Humel from Techstrong and welcome to Control Alt Deploy. You're not familiar.
Control. Alt Deploy is a, uh, podcast webcast that we do in conjunction with our partners at OpenText, great company to work with, and we enjoy working with them. This is Epi episode three of Control Alt Deploy.
If you haven't seen the first two, I highly recommend you go and do it. Let me introduce you quickly to our guest on today's episode of Control Alt Deploy. First of all, he's a gentleman that had the pleasure of meeting at Platform Con, and we hit it off and I said, I need more of you on Textron.
Need more of you to talking to our audience. I'm gonna love you. I wanna introduce you to Ricky Zachary.
Ricky, welcome to Control Alt Deploy. Ricky, if you wouldn't mind, just a quick, I don't know. Yeah, of course.
Background on you. 30, 60 seconds. Yeah.
Uh, thanks for having me, Alan. Uh, Ricky Zachery. I'm the global lead of platform engineering at ThoughtWorks.
ThoughtWorks, a global consultancy. We've probably written a book that you're reading or have on your shelf, either about infrastructure as code or platform engineering, or Agile or any of the other kind of like software development practices. Um, I've been there for four years, and, uh, prior to that I was in the defense and federal contracting space.
So if you follow me on LinkedIn, you can ask me about COBOL and JCL programming at the different federal agencies. And again, it's great to be on the, on the podcast. Alan.
Thank you, Ricky. Enjoy having you here. Next up we have Julio Artiega, Julio with Open OpenText.
Julio, you tell us a little bit about yourself. Sure. Thank you, Alan.
Thank you for, uh, having me on the podcast. So I'm, uh, a principal solutions consultant for OpenText specializing in the, uh, application delivery portfolio, uh, for DevOps. Uh, basically covering your performance testing, uh, functional testing, uh, test management, and all the associated, uh, methodologies like Waterfall and Agile as well.
Excellent. So, guys, let me start off today, right? Today's, today's episode is about performance at scale.
Keep up with the cloud native complexity. But let me tell you a quick story. I guess it's around 20 16, 20 15, something like that.
I go down to Austin, Texas for DockerCon. We're thinking about launching a cloud native. We didn't even call it cloud native.
We're thinking about launching a, a new site around containers and, you know, development involving containerized, kind of, you know, uh, patterns. I guess Cloud native was starting to become a word. But anyway, I go down there and I, you know, I did some interviews, I listened in on some sessions.
I came home and I said, guys, we shouldn't use anything with the word docker and the domain for the new website. A, because Docker probably will sue us, but BI don't know if Docker's gonna be what we thought it was. The, the, the Google guys are down there.
They got something called, they're calling Kubernetes. 8. It's not even officially released yet.
And that's all anybody could talk about. That's all it, it is the, the shizzle, right? Everybody wants to do something with Kubernetes.
I said, the only thing is I sat in on a bunch of sessions, I saw work in it. I don't think it's ever gonna catch on. It is hard as hell.
And this whole cloud native thing, combining that with Docker, with the Kubernetes, with, with the, uh, microservice architecture and everything else, you think this is gonna replace what we do now? I don't think so. Well, that's why I still have to work for a living.
I should have got into that early. But, you know, I'm gonna put forth the premise that cloud native was complex from the get go, and it hasn't changed. And in, in many cases, it's even gotten more complex, especially when we talk about testing and deployment and observability and stuff like that.
Julio, you sound like you spend a lot of your time around that. Yes. Am I, am I making it up or is there any truth there?
No, there's a lot of truth. Um, I mean, we, we see it and we live it every day at OpenText, right? I mean, it's, it's the classic development and, and evolution of an application from, uh, end tier or client server.
And then we added web services, microservices, virtualization. And today, anybody that says cloud is simple, is really not seeing the big picture, you know, from where applications live to how they're configured, to how we interact. There's just so many moving pieces that it becomes extremely challenging.
So the, the applications have gotten better, but the complexity still remains the same or, or maybe even gotten a little bit worse. Ricky, isn't this the reason for platform engineering? That's the, that's the goal, right?
Is to hide and decrease some of that complexity, particularly when it comes to the, you know, developers that are, are building applications, right? Do, do they care if they're using Istio for sidecars or Istio for service mesh? Not really.
What they do care about is, okay, I can create and write an application, deploy it, and it's going to be performing and I can test it, right? So the ultimate goal of platform engineering is to hide some of that complexity away from the developers, but in hiding complexity, it becomes more complex somewhere. So, you know, the platform engineers, the DevOps engineers, the people managing the infrastructure now are taking on a lot of that complexity.
That, and that is a theme, right? Like, Hey, we, we, all right, we learned a lesson during the DevOps rollouts, right? We can't just throw everything on the developer.
Our testers need to do the testing. Our security people need to do security platform engineers. We need to build the platform.
We can't tell the developer to build his old plow or build their own platform. Julio, how, where does the, where's the rubber meet the road on that? Well, I, I, I think we're seeing, we're definitely seeing, uh, a lot more shift left speaking of developers, right?
They have a bigger role. Um, traditionally when I started in my testing career, you, you had a siloed approach. Developers never spoke to testers, and it was two separate rooms, two separate worlds.
And today, it really doesn't work. If you have that siloed approach, there has to be communication. There has to be that beginning all, all the way shifted, left.
Now developers have more responsibility. Um, and, you know, that's where we spend a lot of our time. I spend a lot of my time working with developers, making sure they're comfortable, and making sure that they understand that we're not asking them to re-engineer their process.
We're simply asking them to use tools that, that are available to them today to become part of the solution for performance testing and, and, uh, resilience testing and so on. So we're seeing more of a shift left without, of course, discounting the shift, right? Or the middle layer as well.
So we're seeing a shift left in terms of breaking the silos and communications. Yes, but we're not telling the developer, Hey, you gotta do the testing. Uh, no.
Well, we're telling them to contribute to the testing, right? They're gonna be a contributor, not responsible for the entire testing, but we see it more as a, uh, more as a process, right? So developers begin the process, then they hand it off to the performance engineers, the folks that do the large scale, uh, testing.
And part of the reason for that is because it's expensive. Performance testing can be expensive. So you don't want your developers debugging a script, taking up resources, load generators and controllers and so on.
So they're able to work locally, kick off the process, do early testing, you know, concurrency and, and some quick smoke testing, and then hand that over to the next phase in the process where the larger platform is consumed. Got it. Ricky, how's that jive with, with the platform engineering though, right?
It sounds, sounds a little bit like that. Very, I I think it drives quite a bit. Um, very, very much so.
It's about giving the developers the right amount of feedback at the right time. And that could be through a platform abstraction, right? Um, I'm not going to, I'm going to give you the developer, the infrastructure or the interface to perform a performance test.
But you don't have to go in, write all of the harness and write all of that infrastructure to perform the performance test. And, and we like to talk about the testing pyramid a lot at ThoughtWorks, which is this idea that there, you know, there's a pyramid, there's unit tests at the bottom. We definitely want developers to do that, and we want to accelerate that through different platform extractions.
Hey, here's a testing ski, uh, template for our J unit. We're gonna give that to you, uh, from a platform engineering perspective, just gonna build to be built in a platform. But as you move up the testing pyramid to Julio's point, it becomes way more expensive to do a performance test or to even have, you know, 50 developer, I mean, 50 testers doing manual testing.
So you wanna do that less often, and you want to drive the developers to being able to get most of their feedback through unit and integration testing, which is why, you know, Alan, when you were telling the story about, you know, Docker Khan 2020 to 20 15, 20 16, that resonated with me because that was kind of like the rise of the containers. And now the container is something that I can be running on my local machine and getting feedback from. I can be, I can run it in a GitHub action or, or, or, or some other, you know, runner, I can get feedback in a lot of different ways, but it does jive with, you know, platform abstractions, platform engineering.
How do I get developers to shift left? I build that into the platform so they don't have to do a lot of work. Yep.
And, and I think it also makes them feel more comfortable, um, because I, you know, in the early days, earlier in my career, I mean, I still have the scars, um, you know, approaching developers and saying, Hey, you have a new responsibility you're gonna test. And they looked at me like, wow, what, you know what, that's what you think. No way, Julio, I'm definitely not doing, not doing that.
So, so it's about making them comfortable, right? Um, and, and make, and helping them fail. Now, when I say that, you know, most people look at me like, wait, you're crazy.
But no, it's about failing early, because once you fail early, you have a lot runway, uh, in order to, to, uh, uh, take care of problems, defects, and so on down the line. So it's, the developers become very, very important part of this equation. Agreed.
Agreed. So guys, I think we've defined the problem, right? Or we've defined the fact pattern.
Yeah. Julio, you're a solutions consultant, right? Ricky, you, you, you, you know, you're in the same, I mean, you're your job there is director of platforms.
Where are people watching this? Both developers as well as testers, right? People doing that.
What can we do to make their lives easier? What can we do to let them go faster? What can we do to be more efficient?
I'm not saying cheaper, but more efficient. So one of the things that we, we talk about at, at ThoughtWorks, um, that I talk about in particular is the, and, and this is gonna sound very cliche and maybe even kind of like consulty, but, but it's really about identification of friction and waste within the system, right? The old lean and six Sigma type stuff is making a comeback, um, in, in platform engineering, which is, if I'm a developer and I wake up every day and I have to configure a local environment, and it takes me six hours to get my local environment set up before I can even start writing code and testing that code locally, that's six hours is a, a lot of friction and waste in the system.
So the platform six hours, you're not getting back six, six hours. I'm not getting back, right? Yeah.
So as a platform engineer, I am always listening and talking to the developers that are gonna be using those systems that I'm building. Um, and I'm, I'm working with clients to build, and I'm gonna, I'm going to ask them questions like, what's your biggest pain point in the day? And it could be something super mundane.
It could be, you know, I have to log into my Okta five times every time I, you know, leave my browser. I've gotta go and log back in. The frequency of that happening, the friction that that causes is something that we should be trying to extract out of the system and giving the developers and testers the tools so that they can do the things that Julio was talking about, getting feedback as quickly as possible, getting developed software as quickly as possible.
So it's really about identification and extraction of that friction, not just, hey, throwing the 200 plus CNCF projects that are out there at, at some of the developers, right? Who cares if I'm using, you know, um, um, vena or Sty a for OPA, how does that make the developers' lives easier or the tester lives easier? Should be our starting point when we're having conversations about platforms and platform engineering.
Thanks, Phil, Julio, and a couple, and a couple of points on, on the, uh, and Ricky, you said it perfectly feedback, right? Feedback is, is, is huge. And, you know, going back to my siloed example, right?
We, we traditionally thought of testing, performance testing, you tested the application, you certified it, everything's great, you threw it over the wall. Now it's somebody else's problem. But I always, always trust to all the performance testing teams that the process is cyclical, never ending feedback, right?
And that includes operations, that includes the folks monitoring the application once the application is in production, if there's a defect, which hopefully there isn't, but if there is, it happens, right? You wanna be able to feed that back into the testing loop, because now it is up to the performance team to determine how can we approach a more accurate, uh, way of testing so it reflects the reality of the application in production. The other thing too that we're seeing more and more in the industry is chaos testing, right?
Throwing a wrench into the works. Let's, let's break the system. And, you know, I always like to say it's, you're successful when you fail because you found that needle in the haystack, right?
And ideally, your user should not feel the failure. They should never know that the application failed, right? It failed, sure.
But then things took over in the backend provisioning new, uh, uh, machine instances and, and so on. And then the other thing is, you know, look at performance from a multi-layered approach. And that includes not forgetting the, the, the client side of things, you know, making sure that, you know, we, we talk about the backend, we talk about, uh, database servers and application servers and response times and things of that nature.
But now let's look at the, the front end. Let's look at the other layers and, you know, make sure that the user's having a positive experience as well. So a lot of things happening.
And, and, and, and Julio, really quickly, one thing that you mentioned that I like to talk to, um, developers and and clients about is what is the reason why you're doing all of these tests? The performance tests and all of those things? It's really about release confidence, right?
It's, it's giving the developers the confidence that, um, I, I just wrote some code. I know that things have happened that are outside of my purview, but I have confidence that I'm going to be, that that's not going to cause a ma major outage, right? That whatever I just did.
So it's a series of all of these different tests and the feedback that increases the, the, the, the confidence of the developer, right? So it's like, Hey, I, I, I wrote this code and I know that Julio's team has a performance testing harness that's set up that if it finds something, I'm going to get direct feedback about how to fix it. And I have confidence that that fix is actually going to to matter.
So that, that, that idea of release confidence is something that's coming up more and more. It's not just testing for testing sake, testing for compliance, or testing for, you know, okay, we've got the check marks and everything passed. It's about really release confidence.
Yep. And it changes the nature of the conversation too, because one of the things that we bring back to developers, we say, look, we, we, we've taken care of the backend for years. We've, we've seen the response times from databases, the errors with, with application servers.
Now let's change the context of the conversation to application design. How can you make your application more efficient? How can you design a better experience for the end user?
You know, to put it frankly, how can we make you a hero, uh, to your, to your team, to your team, right? So, so I think it drives a much broader conversation than, than we've ever seen before. Guys, I'd love to have a much broader conversation, but these are only supposed to be 15, 20 minutes.
So we're about outta time. But this was a great, great discussion, Julio, Ricky, I, we will continue it, hopefully in a future event. Thank you both for, you know, bringing your years of experience knowhow to our control of deployed discussion today.
Um, I hope you've enjoyed this. You know, we, as I mentioned this episode three of this podcast series, there's also six on demand webcast that you can go check out. So check those out.
But until next time, on behalf of OpenText and Text Strong as well as ThoughtWorks, thanks for joining us here on Control Alt Deploy. Bye-bye.