28. Network Automation Is More Than Just Tooling with Mike Bushong of Nokia – Tech Field Day Podcast Spotlight Series
The modern enterprise network automation strategy is failing. This is due in part to a collection of tools masquerading as an automation solution. In this episode, Tom Hollingsworth is joined by Scott Robohn, Bruno Wollmann, and special guest Mike Bushong of Nokia to discuss the current state of automation in the data center. They discuss how tools are often improperly incorporated as well as why organizations shouldn’t rely on just a single person or team to affect change. They also explore ideas around Nokia Event-Driven Automation (EDA), a new operations platform dedicated to solving these issues.
Transcript
The world of network automation is complicated enough as it is, and we don't need to make it any worse by creating our own custom tooling and failing to document everything. In this episode of the Tech Field, a podcast, we explore the ideas behind these custom automation solutions and whether or not they'll survive you leaving your organization. Welcome to the Tech Field Day podcast, where we bring together a group of IT technical experts to discuss a single idea about key concepts in the industry.
This podcast features a variety of perspectives from members of the Tech Field Day delegate community, and is often recorded in association with one of our events. Tech Field Day is a part of the future and group in this podcast is also published on our sister company site Techstrong tv. On this episode, it was brought to you by Nokia.
We will be discussing whether or not your custom-based tooling will allow future generations to succeed in your automation efforts. Before we jump into a little bit more detail on that though, I'd like to take a moment for our guests to introduce themselves, starting with Scott. Hey, Scott Roben, uh, host of total network operations and co-founder of Network Automation Forum.
Hello, uh, my name is Bruno Wallman. com. And lastly, joining us as a special guest from Nokia is Mr.
Michael Bouchon. Uh, hi. Thanks.
I'm Mike Boong. I do data center things at Nokia, having previously done data center things at other companies that have data center ambitions. Alright, well thank you all very much for joining us.
Let's jump into this episode As I teased, we're gonna be talking a little bit about data center automation, uh, thanks to having Mike here with us, but also we're gonna talk a little bit about some of the challenges that people face in it. Specifically what happens when your company has made it a concerted effort to create some kind of an automation strategy. But you found it kind of difficult, uh, you know, to get it going.
So you kind of built this custom tooling and then you have to train somebody how to use it and what's gonna happen. And more importantly, how can we prevent the friction that usually comes along with this? Because I know a lot of people just take one look at this pile of, of, uh, tools and weird bash scripts that you've created to make 'em all talk to each other and they just throw up their hands in the air and you're like, you know what?
I'm not gonna work on this. I'll go do something else. Like, I don't know Kubernetes.
So I wanna jump in and kind of, uh, open the floor to our experts here. We've spent a lot of time over the last couple of decades kind of coming up with these solutions and lord knows there's enough of them out there. What is it that seems to make people gravitate toward this idea of I'm gonna take best of breed tools from all over the place and then just kind of drown them in crazy glue to make them all stick together and hopefully accomplish what I want to do?
Yeah, I I think the easiest answer to that question, which may not be the most comprehensive, is they weren't aware of the right tool set that would work together for them. So they were forced to either go out and pick piece parts and assemble their own, you know, tools in their toolbox and or create pieces that they couldn't find other solutions for open source or commercial. That's, that's the obvious answer, but it goes deeper than that.
I think that makes a lot of sense. I think we're sort of an incrementalist, uh, industry, so we'll make incremental changes. And so you get individuals who wanna do something that makes their jobs a little bit easier, and so they bolt on and they bolt on and they bolt on.
Um, I don't think you're always looking at kind of the intersection of jobs. You're looking at how do I reduce effort? And if you look at really just an effort thing, then it's like, how do I do something that makes my job, you know, marginally easier?
And that leads to, I guess, a bunch of point decisions. Um, the number of network automation architects, people who kind of sit back and say, lemme look at the entire thing. Let me look at everyone's jobs.
Let me look at all of the workflows. The number of companies that have those roles is relatively few. And I think the people capable of putting that system level, you know, I guess, uh, thinking and design together.
I think I just think that's a, a pretty sparse skillset, um, and a and a thinly populated, uh, environment for us. Yeah, I would agree with, um, both your points. I think, um, especially at the start of any network, uh, automation efforts at a, at a company, usually it's, uh, somebody just trying to reduce their, um, you know, reduce their workload or try to get through a, a, a huge stack of to-do lists.
Um, and so they, uh, implement point solutions. And in, in that case, it's usually, um, a network engineer who may not, um, necessarily have, uh, great coding skills, but good enough to get the job done. And then yeah, lacking that systems thinking to, um, build a start with a foundation and may not have the support to even build that foundation.
So starts off with, uh, the simplest thing that, that that person can, can use. So I want to toss this idea out here because it was something we were kind of discussing as we were prepping for this podcast. And, and it's my take on it that a lot of times what ends up happening with these kind of organizational efforts is that you find a, a really, um, intelligent person in the organization who has maybe a little bit of bandwidth to take something on like this on, and you kind of toss this over the fence, right?
You're like, Hey, we want you to automate this. Um, you know, the, this is kind of the other track from people who are just like, I'm trying to make my life easier. But I think that people kind of fall into two different buckets.
Generally, there are creators and users, creators look at this problem and go, oh, I'm gonna build the perfect automation solution that marries these tools together and tries to make it all work. And, and it's gonna work out great. I mean, maybe you'll have to kind of think like I do to make this work properly, but I promise you it's gonna be the best thing you've ever seen.
And on the other side of the fence, you have users who are like, I don't wanna build anything. I wanna go out and find the perfect solution, the perfect platform, the perfect organizational structure to make all of this work. And I'm just gonna operate that in perpetuity until I need to make some changes.
Now, users are great for companies that sell tools because, hey, we'll sell you whatever you wanna make. Creators are great for people who are kind of like bush whacking their way through the jungle of, I don't know what I'm doing. But the problem is, is that when you turn around and see the path that they've taken, it's got lots of snakes through it, and they left a stump somewhere, and you're like, but how am I supposed to drive down this road that you've created?
Because it doesn't really work. And and oftentimes what happens is, is that when you bring the users into the equation to work with the creators, they turn around and go, this is unusable and I'm not gonna bother. And so they either throw their hands up or they go out and buy something that isn't effective simply because, well, it does most of what I want and I don't care what it costs because I'm not gonna sit here and waste any more my time trying to figure out how to use this thing.
Well, the, the discipline of documentation goes right along with that, right? How to, and now I can't get this picture of you with a huge machete cutting your way through. Pick your favorite, uh, jungle.
Um, we'll put you in a Jumanji sequel Tom. How's that? Um, but uh, yeah, you know, there are people who love to go solve problems, do it creatively and effectively, but don't necessarily memorialize what the next people downstream are gonna need, um, to follow, to follow the roadmap, to drive through the jungle.
To misappropriate your analogy there, Uh, it's worse than that, right? The creators will go through the path like one time, but anyone who's, uh, who's blazed a trail and knows that, that that trail stays blazed for, you know, a short amount of time and then it grows back, you need people to go back and, and the carrying cost of some of these tools is, is frequently not really considered. You know, so you build one tool and you do some great thing, and now you have to maintain that tool.
You built it for yourself, now you've gotta maintain it for 10 users. Okay, like, maybe I can still do that. You add a second tool, right?
Okay, that's good. And maybe you add another 10 users. What's the maintenance burden of two tools for 20 people?
And are you gonna keep up with that? And then what happens when you want to add a third tool or another 10 users on top of that? Like, I just think the carrying costs it, it's not, the creators don't often consider the carrying cost.
And then the challenge is that the creators are frequently motivated by interest. And so when their interest in maintaining the thing goes away, like what happens next? I just, I see a lot of companies start and then they sort of, they, they realize they can't scale.
And some of our breakthrough moments, they don't happen because you merely start, they happen when things become codified and somewhat routine. And at that point, I, I just think it's a different model for, for maintaining it. Mike, let me, let me ask you about that.
Like, you talk to lots of customers and end users and, you know, calculating total cost of ownership is something that no one in our industry has ever been good at, ever that continues to this day. I love your concept of carrying cost. Have you ever, ever come across an organization that actually thinks of it that way or runs those numbers?
Uh, almost never. I'm trying to think if I even have one example of someone that, that, that did. Sometimes you'll see people, I, I guess the companies that do this and actually the cloud companies or you know, I'll say cloud majors, they like, think SaaS companies, you know, clearly like an Amazon or a Google or a Microsoft.
They have, you know, relatively robust teams that are, you know, building out software for expressly this purpose. But again, that's part of their product. So it's, I think it's a different use case when you get into the SaaS companies.
Um, some of the SaaS companies will have like their network engineering teams paired with their DevOps teams, their DevOps folks, approximate software developers. Um, maybe they're networking adjacent, um, but their primary role is to go and, and build the tools and then to maintain them. So the carrying cost is built into the org structure.
I don't think they talk about it that way in terms of carrying costs, but I think they're actually organized that way because you have a, a set budget for a set number of people that are, um, expressly allocated to support these types of tools. But if you get into like enterprises, I don't think that's necessarily the case. I think there, the teams tend to be merged together.
I think, um, Bruno's comment earlier about, you know, people start down that path and it's like an individual who has interest in a problem and maybe some time, you know, they go when they do it. I think there, I just, I don't think the carrying cost concept is even, is even present. And then not only is it not present, it's, it puts their organizations in this like really, really, really perilous position where if that individual leaves you are absolutely like behind the eight ball.
Um, so I, I, so, so no, I mean, I guess to answer your question, no, I don't think that many people think about it that way, but I, I think if you wanna scale anything, you gotta start thinking about like how do you, how do you do it beyond just the tool or the user? Yeah, I would agree. And I actually am gonna refer back to something we saw at an early cloud field event here.
Uh, Ben Siegelman at LightStep, I'm paraphrasing a little bit, but he, he came from Google and he said that in Google we would write tools to solve problems and then we would release them on the internet and then people would immediately download our tools and start using them for things that they weren't designed for. He's like, don't do that. Like, Google makes tools for Google's problems and your problems are not Google's problems.
And Google's problems definitely aren't your problems in a lot of cases, but people think that if it's a tool that was released by Google, it has to be good. And I think that one of the problems is, is that people will download a tool that looks like it's gonna do exactly what they want it to do, and then what ultimately ends up happening is they end up transforming their problem set so that the tool can solve it. And they didn't actually fix their problem, they fixed a problem they didn't have because they wanted to see the tool be successful.
And like you said, Mike, that kind of leads to a problem of, yeah, don't make Brent very mad because if he leaves, we're in a lot of trouble, or Brent can never be promoted, Brent can never go on vacation. It's the Phoenix problem, uh, project problem that Gene Kim wrote about, right? Like, when all the work has to go through him because he's the only one that knows how to do anything with it, then your company is effectively reduced to the bandwidth of a single person.
And when, you know, automation scripts are constantly breaking or encountering new things that they don't understand how to operate, like how can you scale that up or out when you are utterly reliant on a single person to do it? Yeah, I think these are are lessons that, um, have been learned in other disciplines inside of networking and have yet to be learned for automation. Like, um, I, I have customers that went through some of the same, uh, problems with cloud, like trying to build tools to make the cloud more consumable.
And all of my customers are enterprise level, so that's all I can speak to. But, um, you know, they didn't, they were, it was recommended for them to create a cloud specific team to deal with that. And, and they didn't do that.
It was, that functionality was munged in with other network engineers. Um, so the same, I see the same thing happening with network automation. It's just they're not, it's kind of a systemic problem where it's, um, I don't know if you wanna call it lack of training or lack of vision or, or what the reasons are, but you know, those automation specific teams aren't, aren't being created yet.
And, you know, I'm not sure the reason, maybe that's 'cause it's, there isn't enough, uh, momentum to do that yet where organizations have created cloud specific teams after years of failing and flailing away and, you know, with their machete through the wrong jungle and things like that. Well, you know, there's an interesting distinction there. Where do I want network automation to be a job title or a skillset, right?
I have basic carpentry skills, I can fix certain things around my house, not very well and nobody would ever pay me for it. Um, but you know, in my role as a homeowner, I've learned how to deal with certain things around the house. I think that's what we want from most network engineers, right?
Not to turn them into exclusive automation experts, but to have enough smarts and skills and network automation to make it a part of what they do, make it a part of their workflows and so forth. The, the challenge with some of that though, I think I, I guess if I were to characterize the, the network automation movement, there's a distinction between what individuals want, right? Which is really a measure of effort.
I do this thing and I would like that thing to be easier to do. And then there's what organizations want, you know, they, they measure things in terms of time, right? And I think those are two different, and in some cases, competing disciplines.
If you go in and you take something that was 178 steps and you reduce it to one, but you were the only person that executed those 178 steps has huge value to you, I would argue it's marginal value to the company. Um, whereas if you look at where, where time accumulates in a, in a system, um, it's not usually an effort thing. It's the handoff between people.
It's at the intersection of either workflows or in some cases different tools or systems where they come together. And there, when it's, if it's an individual, if it's a, if to your analogy, if it's a, um, you said is it a job title or is it a skillset? If, if we treat it as merely a skillset, whose job is it to handle the handoffs between people or organizations or whatever?
Like, it's not obvious to me who picks that up. And so I think those things end up falling through the cracks or never being addressed. And then that leaves you with this disconnect where it's like, we put all this effort into network automation objectively individuals will probably toil less because they've solved some of their local problems, but the organization doesn't feel the benefits.
And so now it's, you know, year two and you're looking for where's the next, you know, network automation, investment dollar, and the organization will go back and conclude, you know, it's nowhere because we're not seeing organizational benefit. Whereas the individuals would be like, oh my gosh, we desperately need it. We feel it.
And so there's this tension which is like, we, we really need it and yet we can't invest in it. I feel like, like the industry is kinda locked there. I don't, I don't know how you unlock that, by the way, but I I, I think this effort versus time thing is, is important.
Um, I just don't know that people see it that way. Yeah. So one of the, one of the things I would propose that can help with this problem is you've got somebody that functions like an operations architect who is responsible for how all these operations focus tools work together.
And, and that doesn't need, doesn't need to be a job title either. It can be just a responsibility. Um, but if it's a big enough organization, right?
Maybe there is a person who does just that, but somebody's gotta be keeping an eye on, you know, what are these 173 scripts in the scriptorium? What do they do other, you know, am I, am I contributing things that have common standards within my organization? And to reduce all that entropy, it takes work, right?
And it's gotta be amortized across the, the, the, uh, time and effort saved by the actual tooling. And I think what's important for people to understand, kind of like what Mike said is that it's real easy for me to figure out how to reduce the amount of time that I'm spending on something. But because anything that's not time I'm spending is outside of my purview.
I have no idea about things I don't know about. Um, it tends to, that's like you said, that's where time accumulates because taking an extra three seconds to do something doesn't sound like a lot. Taking an extra three seconds to do it over 500 devices does.
And that's where people need to get better at tracking these kinds of things. It's almost like we need some kind of a, I don't know, like a database. 'cause I know that the one time that I tried to do anything that involved me speeding up a process and, and adding more tooling into it to make it go faster because it was an ISO certified process, I actually had to go through document how quickly it took for me to do the, the original method document, how quickly it took me to do the new method, compare how much time was saved per deployment, and then explain, you know, all of the like skilling up on the tools and stuff like that.
And it actually took me like three times longer to do that, to save in the end what amounted to about 10 minutes. And so I think that one of the problems that we have is that people have gotten so good at just coming up with these solutions and saying, oh, well, you know, I saved a little bit of time, so it's gotta be good for everybody, right? 'cause I'm not spending as much time working on this and, and not doing a good job of communicating how this helps down the road.
And I think that that's one of those things that, you know, 'cause when we talk about this with, with our colleagues, you know, provisioning a switch is easy, right? Like I, I, I put the, the code in and I tell 'em what to do. It's when do I get the order to provision the switch?
Is it an hour after it was needed or is it a week after somebody requested it? I, I think, um, there's two things I want to bring up with what you just said, Tom. Uh, the first is, um, I'll address the second one first, and that's the, uh, provisioning a switch is easy.
Um, I I think if you have, um, without automation or without templates, without um, you know, standardized configurations that, um, you know, network resources can draw from, um, you know, if you've got five resources, uh, deploying five different switches are, are they gonna be identical? So there's that, um, configuration drift, um, with a lack of tooling as well. The other one I wanted to address, the other point I wanted to address was the, uh, sharing.
And that kind of goes to one of Scott's points too, and that's documentation. And I think anytime you, um, create a tool, it saves you time. You know, you gotta upsell it to the rest of your team, to your manager and things like that and, and share it, document it, put it on a local wiki or whatever, and make sure that's available to e everybody else on your team and they know how to run it and are capable of running it.
Um, e even though it might be a a point solution that may be hard to grow into something, um, more foundational to the organization, but be willing to share it and document it and upsell it to whoever's gonna listen so that automation can be that, uh, workforce multiplier that it should be. Yeah, that's a corporate discipline that applies to a lot more than just automation, right? Who, who, who here loves writing documentation?
Come on, Mike. Admit it. I know.
Yeah. Um, you know, and then, and there's, there are very few tools that are gonna help that happen automatically, but man, you give me a good tool that automatically documents everything I do with minimal need for cleanup. There's probably some revenue there.
I in some cases, the tools are the documentation by the way. Um, and so I actually think the automation that, one of the benefits that I don't think gets, gets talked about enough when things are automated, you take what was in somebody's head and you, you force 'em to type it out. I think that there's like huge value in that.
Um, I've always thought, like just on the vendor side, you know, having tools that do discovery, if you could just print out Vizio from there, like just export a lot of that stuff. I mean, like, there's huge value in that kind of thing. Um, people don't buy it because nobody identifies their problem as I have a documentation problem.
But I promise you if you built it, people would be like, oh my gosh, that is difference making. You see it in the wireless side. I think it was, is it hamina?
I don't know all the wireless guys that well, but they have like a whole tool around, you know, how do you map out, um, you know, AP placement and some of the stuff they've done there, which document. Like, I just think that's fascinating. And so that's a free shout out to them, by the way.
So you can send me the swag later. Um, but I just like, like to me it's like focus if when you have a maniacal focus on like the user experience, things like documentation end up being, you know, more important than some of the features and functionality. I think people who are, who care about the, the user experience where the user is, the network engineer should have that same inclination.
And I think that network automation is one, it will automate stuff, but two, I think it, it could be that this is the, the defacto documentation tool for a lot of these workflows. And, and I think that you're, you're onto something there, Mike, because one of the things that network engineers are notorious for not doing is documenting anything, right? Because it just takes too much time.
So this goes back to that whole where are you spending your time doing things, but I think a platform like say Nokia EA, it's effectively self-documenting, right? Like it creates a database of all the changes that we've made. It has a system that you can query to figure out what's going on.
And in a way, I don't have to forget to do something because it's built into the process as I do it. And that's kind of the way that the tools like Kubernetes were effectively built, right? Is I have the ability to track all of this stuff.
Why am I not doing it? I think back to the days of using things like TAC acts to authenticate and to basically say, okay, I'm gonna check all of my configuration changes in through this TAC Act server so that I know that if something went wrong, who typed it in, when did they do it? And, and how can I roll back from it?
Although I will say that compared to the, the, uh, the good old days, something like edic is a lot more powerful. I was just thinking about it the other day, like, you know, how many times have we found ourselves, like we, we started deploying something and then like three or four hours into it we're like, oh no, there's a problem and we need to figure out how to fix it, but we gotta roll everything back. Oh, we're 400 switches into this software deployment.
How can I, where, where did I stop? Can I, can I figure out how to back this, roll this all the way back? And if you have a platform, not not just a collection of tools, but an actual cohesive configured platform to just go back and go, okay, here's where I started the deployment roll back to here.
And then when you do that, you go back to the system and you query and you say, can you show me any devices that are still running this code version? And if it returns zero rows, guess what? You fix the problem until you can go back and do that.
And, uh, anybody will tell you that that halfway through your change window is when you start figuring out if things are actually busted or are gonna work, because that's the point of no return. Once you get past the halfway point, you're kind of committed and you have to finish. I think there's a whole model for this in the Marvel cinematic universe with the time variance authority and every change, instantiates another branch of a timeline, and those TVA cops are always going back trying to make sure we're all on the same timeline.
Sorry, I won't go any further with that, I promise. I I think that's right. The, uh, the, in in that timeline view though, I mean the, the, the question is how do you version everything?
It just becomes like a, like it actually approximates, um, you know, development. We talk about DevOps and treating your network as code. Um, you know, if you think about network as code, you have like an an IDE, you know, some way of manipulating things, but you also have like a, a source code management tool on the backend.
And so, you know, the question is how do you handle branching? Um, you know, we've chosen to do it using Git because we thought that the tool is familiar and has a bunch of capability in there. Whether that's the right decision on the tooling or not.
I just, I would say that people need to, in this space, you need to be thinking through, you know, how do you make changes, how do you validate changes? How do you store those changes? And then what are the operations that work around that?
That becomes then that underlying SCM, like for, again, for us, it's git that becomes the documentation of what is good. And then from there, every change you make should be relative to some point of good, which gives you the safety net in case something goes wrong. Um, but that's, i I, I didn't think of it this way until you mentioned something Tom, but like the, I, I agree with the self-documenting piece, but it's, it's almost like document based or documentation based automation because you're, you're starting from a point of documented good before you move.
And I think that's, I mean, that's something that, that, you know, way back in the day with the whole commit model, I mean, Juniper was early in that space, kind of the same idea. Um, if you took that and said, let's extend that same idea to the network at large, so we make atomic sets of changes, but at a network level, I, I think it's a, it's an important change in how people operate and I think it would meaningfully affect how you plan around, um, you know, maintenance windows. So another, I, I think we're putting a lot of eggs in the GI basket as an industry and I wonder what that looks like when Git gets disrupted.
'cause it's, it's not an, it's not an if it's a when, right? But if I have a wrapper, if I have a tool like EDA or other tools that uses it on the backend and I could swap something in for the next thing that actually mitigates against other tool changes in the industry and makes it more accessible to network engineers who haven't touched GI even though they should. And we know that, you know, we need to preach skilling up to folks, but, uh, there's a lot of value in having something on the front end that hides that from me and makes it easier to use.
Well, thank you all very much for joining us on this episode of the Tech Field Day podcast. Before we go, uh, where can people connect and learn more about what you're doing? Starting with Scott, Please reach down on LinkedIn.
That's the easiest place to find me. Yeah, me as well. I'm also on LinkedIn.
That, um, gets me to my web, gets you to my webpage as well. com. Uh, you can find me, uh, on LinkedIn as well.
Also buan on Twitter. com. Awesome.
Well, thank you very much for listening to this episode of the Tech Field Day podcast. If you enjoyed this discussion, please subscribe to our YouTube channel or in your favorite podcast application so you don't miss an episode. Also, consider giving us a rating and a nice review.
This podcast is brought to you by Nokia, as well as Tech Field Day, the home of IT experts from across the enterprise, which is a part of the Futurum Group. com/podcast. Review us on Techstrong tv.
com. Thanks for listening to this episode, and we will see you next week.