35. Snowflake Networks are Built to Break – Tech Field Day Podcast
Network engineers are notorious for doing whatever it takes to keep their customers and users happy. No reference architecture is safe from modification. However, these unique designs, commonly referred to as “snowflakes”, create challenges when unforeseen consequences occur. In this episode of the Tech Field Day podcast, Tom Hollingsworth is joined by Dakota Snow, Steve Puluka, and Bob McCouch as they discuss the challenges behind snowflake design and operations. They talk about the best way to build better systems and prevent the challenges caused by uniqueness.
Transcript
We love to build systems. We love to make them important and do amazing things, but unfortunately, in the process of doing that, we sometimes build custom things that don't work the way that they're supposed to. We, uh, build snowflakes.
We need to stop that in this episode of the Tech Field Day podcast. Welcome to the Tech Field Day podcast, where each episode we bring together a group of IT experts from across the enterprise IT space to discuss a premise or a topic usually related to one of the things that we talk about during the Tech Field Day event series. We're here at Networking Field Day this week, and we brought together a group of experts to discuss exciting trends in the networking space.
I'd like to take a moment for our guests to introduce themselves before we jump into today's topic, starting with Dakota. Hi everyone. My name's Dakota Snow.
I am a YouTube host of the Bearded IT dad and also the IT career podcast and my day job that actually pays the bills. I'm a director of network operations for Fiber Optic ISP. Hi, I'm Steve Plucka.
I'm a network, uh, architect with DQE communications out of Pittsburgh. Also a fiber ISP and a METRONET provider. And I'm Bob Mcco.
Uh, a 25 year networking veteran, uh, who works for a, uh, national, uh, solutions integrator called ahead. Alright, thank you all very much for joining us. Let's jump into today's premise.
We have spent a lot of our careers building bespoke custom artisanal networks and systems. Everything is special and different most of the time, but in reality, we're iterating on the same designs that we've been doing over and over again. But this time it's different and maybe it's only about 5% different until it comes time for us to be able to do things like automation or observability.
And then we find out that this beautiful, unique snowflake is causing us a massive amount of headache. And if we'd have just built it a little bit differently to start with, we probably wouldn't be dealing with all this hassle. So the premise for this episode is that we really need to stop building snowflakes.
And I say that in my heart of hearts. I know I've done this. If, if I'd have just come up with a standard way of doing things, it would've been so much easier, and yet I find myself taking the same shortcuts time and time again.
So I'm gonna toss this out to my expert friends here. Why do we spend so much time building these snowflakes? Well, I'll, I'll jump in and kick it off.
And I, I, you know, coming from the, uh, metro ethernet provider, a lot of our snowflake, uh, operations come from customer requests that you just really, you don't want to say no. I mean, they're customers are like your kids. If they want something, you want to give it to them.
And it, it's only a little bit different than the standard offering. So we should be able to pull this off and the equipment will do it. The, uh, the parameters are all there.
So why not? Oh, I know. Why not?
Because you're gonna forget why it's different and special when you go back in six months to try to troubleshoot something. And then you're sitting there going, what kind of an idiot would build this? Oh, wait, I'm that idiot.
Well, and one thing I've noticed is when the organization, an organization is small or in their, they're in their startup phase, you kind of want to bend a little bit more. You have that feeling like you have to go off the norm, off the being path. But just like you said, as soon as you do, and it's not either properly documented, because let's be honest, none of us has time to document things properly.
Speak For yourself, But yeah, you'll, you'll forget why it's like that. Or when stuff hits the fan and you're trying to troubleshoot this and have no recollection, recollection on how it's built, how are you gonna fix it at that point? Yeah, I, you know, I think the reality is most networks undergo, uh, largely organic growth, right?
Mm-hmm. You know, they're, they're designed for a certain scale, a certain, um, feature set. And then, uh, typically the organization, hopefully the organization grows where it adapts in some way, right?
Maybe it shrinks, but, uh, you know, it grows, uh, maybe there's mergers, acquisitions that take place. And so there's this continual sort of organic growth that usually happens, uh, with a, a smaller set of operators who may be learning, you know, little tweaks, little optimizations along the way that they wanna work into the next and the next and the next. And of course, what we don't typically do is go back and, and update something that's working and make changes to it in order to try to, you know, have it conformed to a new, uh, what we think is a new standard, what we've told ourselves, a new standard, which is only the standard until we discover the next little tweak to make.
So I, I think that that organic growth tends to happen, um, and that, that organic change tends to cause a lot of those snowflakes, um, very infrequently, I think does, uh, a a team of network engineers sit down and, and, uh, have the opportunity to develop a completely thorough set of requirements that will remain fixed and not have any deviations required, and then be able to implement that time after time after time, right? So the question is, how do you now go back and start to, uh, consolidate and unify, right? I think to me that's the, the real, the real key.
I would argue that that's, that's true. Like we, we see this kind of organic growth, but even in it, we have boundaries. Like there's only certain amounts of things that we can do.
Like obviously, you know, oh, I want 800 gigabit per second performance from this switch that I bought from, you know, new Egg. Well, I'm sorry. It just doesn't support that.
But more importantly, we also are working within a certain set of bounds. And one of the things that I always, I think about this when I'm giving people an analogy is we, we talk about cars, right? Um, I, I want you to order a car from a car dealer, and you have a certain set of options that you can pick from, right?
Like, you can get cloth seats or leather seats. You can get a moonroof or no moonroof, but you don't get total like customizability. Like you can't put the Corvette seats in the trailblazer.
You can't, uh, you know, add additional things that are not there. You can't put like 35 inch wheels on a Corvette. That's just not an option that's available.
And when you give people a limited set of deployment options, a lot of the Snowflake units goes away because they're like, oh, well, you mean I can't configure this? No, that's not an option that you can do. Oh, okay, fine.
So should we, should we stop bending to the customer's will and saying, no, we just can't do that. I think it's not even the customer we're bending to. Sometimes it's the management team, the leadership team that really doesn't have an understanding of what's practical and what they saw on the latest, you know, blog article that's wholly impractical.
They're like, Hey, you know, this new technology is a new thing put into production next week. And to do that properly and not make this snowflake scenario, you have to forklift your entire environment in at all actuality. So I think we find these shortcuts to sometimes try to meet in the middle and try to halfway do it.
And in that instance, we're creating these snowflake networks. Yeah, that's, that's definitely true. Is that if the, the advancing forward in, uh, in increments be ba is not something they do just to do, they're doing from demand from something.
Mm-hmm. There's a requirement to do something, and so you have to do it right away. And, and then what happens is if, as you said, the the network grows, and when it gets too big, um, or when it gets big enough, I should say, then it's a, it's a requirement that we have these, um, these tools available to both maintain it and continue to grow it.
And those tools become exponentially harder to create and use when the de when the design is, is completely snowflake. Yeah. I think there's, there's a level of discipline within, you know, the, the, the network engineering space as well that we can, that we have some level of control over, right?
It's, and I think the key there is recognizing that one standard one, one specific, you know, golden configuration or golden design is probably not a fit for an entire enterprise, right? So you need to t-shirt size or you need to have, you know, the high, medium and low or something like that. But you need to have enough options that, uh, you know, can meet the, the, the needs in terms of flexibility for the business that again, might be budgetary based.
They might be, uh, you know, scale based. There could be, you know, any number of, of sort of influences there. Um, but then within, once you've defined those couple of, of, of options, right?
Maybe it's a couple of, maybe it's six, then you need to stick to those, right? And, and the challenging part there can be that everybody, again, from a budgetary or managerial perspective, whatever always wants, you know, well, we want the medium, but we want just a little bit more resiliency. Or we, you know, we want the large, but we can't afford that, so why don't we do this?
That's actually one of the reasons why I feel like we're getting to a point where there are more snowflake type things, is that a lot of it comes from doing the minimum. Like yeah. When, when someone says, well, I want you to configure this, but I don't, I can't afford these other pieces of it.
And so you're sitting there thinking to yourself, well, how can I accomplish that goal while preserving as much cost as possible? And so then you kind of overreach yourself, you know, it's like, oh, well, yeah, I can totally turn this feature on. That gives you a little bit of that, except nobody would, in their right mind would use that feature in a particular scenario like that.
Yeah. You know, like, uh, think of the number of times that you've had to turn on like some kind of weird, like layer two data center interconnect. Or my other favorite one from my career was extension mobility for phones.
I'm like, by the way, don't use that ever, that idea, but it's like I, I built a, a snowflake that I never got a chance to fix because no one in their right mind should ever do anything like that. Yet the customer was like, well, we want to do this little feature, but we don't wanna pay for the license for that feature. And I'm like, well, I think I can make it work like this.
So are we, are we creating our, our own worst enemy by doing a little bit more than we should? Yeah. Well, I, I think, I think that that's a, that's a trend that I've seen or a, uh, uh, an attribute that I've seen across my entire career, which is, I think by our nature, I think most it engineers, and I think network engineers in particular are resourceful.
We're natural problem solvers. We're people, right? People pleasers.
Yeah. And we're people pleasers, right? But, you know, so, so when those challenges come up or when those requests come up, we're very apt to say, well, yeah, I mean, I would prefer not to do it this way, but there is a way we can do that.
Right? Um, and, and I think that that, uh, tends to lead to, you know, as you call it, kind of kind of doing this, this bare minimum or, or just, you know, finding a way to make a little compromise to make someone happy that that now has broken from the standard, right? Mm-hmm.
And, and so it's one of those times where you need to have enough flexibility to, again, have some variety in your options, but then have enough discipline to say, Nope, these, these are our options. Because now you can templatize those, you can automate around those, you can audit against those, right? Right.
You can, you can validate those configurations, those designs, those environments and, and actually be able to demonstrate business value through that standardization, you know, and be able to show the business, Hey, look, it is more efficient to operate a standardized network, even if that standard involves four different site designs or six different scale options, or whatever it is. Go ahead. I, I think that's one of the things that I, I, I took away, um, when I, when I made a brief transition into programming, so sorry, a couple decade, decades ago, is this modularization?
Mm-hmm. And I think that's something that we have been slow to come to Absolutely. In the network engineering space.
And one of the analogies I've, I've been using recently because it it seems to click with, um, network engineers is Legos. You've gotta think of your, your click Your, Yeah. Your features and your options in the network.
You pick the ones that, that you want to support, and you can put them together in different ways, but you always use these same, uh mm-hmm. Bricks and you decide as a team, these are the best ones for the types of services we have to deploy. These are the best ones for our gear and, and architecture that we have, and these are what we use.
And then once you've made that decision, then when you actually have to get back to being a programmer, which you actually have to be now as a network engineer and write those, those scripts and those automations, it becomes a very practical thing to be able to do because you've taken that, that big picture step and, and made those, those modules. Yeah. Oh, and one biggest problem I keep on coming back to is a lot of times you might not be the one who initially built the snowflake.
You inherited the snowflake. Mm-hmm. And it's hard to get buy-in when the snowflake is technically working.
And we could just continue to bandaid, perma fixx the problem instead of, Hey, this is the way it should be done. This is it. It's hard to see and communicate that return on investment to project stakeholders.
Yeah. And there's one big, well, I'd say two big drivers that are causing us to have to validate this again. And one of them is network automation.
Because as we've seen, as we've started implementing these network automation solutions, uh, the enemy of a good automation solution is an exception. Yes. It's like, oh, well, I can run this across 20 switches.
It's the 21st switch that becomes a problem because it has a different firmware version on it, or it has a different configuration because it was the first one that I deployed, or something like that. And so network automation projects are being halted or, you know, put it on an indefinite suspension because, well, we just can't make it work the right way across the whole enterprise, when in fact, the reason why we can't is because we don't have a homogenous enterprise. Right?
We have a multi-vendor cross connected nightmare. Yeah. Of, well, these switches were on sale this week.
So, you know, is automation going to cause us to reevaluate how we build these things to make them more Lego like, or are we just gonna have to give up on the dream of having robots run our network? Well, I, so I can tell you that the, like the approach that, that, uh, my company actually takes with our customers is looking at a, like a holistic network modernization strategy, right? And, and, um, so many individuals, engineers, partners, vendors, you know, when they say network modernization, what they mean, it's just a tech turn, right?
Just, just take the old gear out and put new gear. That's not end of life anymore in place. Yeah.
Yeah. Just upgrade. 'cause that's the, well, yeah.
'cause that's the easy way, right? That's easy. If it works, just replace the, the gear so that it's still under support and move on.
Um, and, you know, we look at network modernization in a much more holistic manner, right? And, and that's really like, like applying large pillars of a modern, uh, uh, envisioning of your network, right? So that includes, you know, lifecycle management, right?
Making sure that you actually have a plan. And that, that even that goes to not just doing, uh, you know, like, oh, hey, everything's end of life this year. Like, let's replace it all right?
But instead actually having a strategy and saying, Hey, we turn over 20% of our gear every year, or, you know, whatever it may be. Patching. Patching, right?
Things like that. Right? So overall lifecycle management, right?
Uh, automation being a pillar of looking at a network modernization program, right? And building to that. So as you're rethinking your network, rethinking how you can automate it in the process, right?
Zero trust, how do you, as you're doing this process and you're looking at automation, how do you build the network with security and observability in mind from the get go, right? How do you extend it into, you know, a hybrid cloud and multi-cloud type environment? So, um, so I mean, that, that's our particular strategy when we're working with customers is, is you, you do have to take a holistic look at it.
And it's not a project, it's a program, and it's not something that happens overnight. It's not something that happens for free. It's a, it's gotta be, there's gotta be some organizational, uh, you know, backing behind it and some, and some, you know, buy-in from leadership.
Um, but it, it can result in moving toward that standardized, uh, that, that automatable observable network. Mm-hmm. Mm-hmm.
But it takes time and it's gotta be something that's got, you know, that's beyond just the individual network engineers saying, boy, I wanna make this better tho those, those people, those guys and gals have the best intentions in the world that I'm just gonna start automating things. Right? But if they don't have the backing of the organization from a leadership and management and executive level, it, it, it's a real uphill battle.
Well, and previously I've talked on this podcast about how no one wants to be a network engineer, and the industry as a whole is facing the problem of these growing networks and these larger and larger network deployments, while the people who are actually supporting these networks is shrinking and shrinking the teams. And it's, it's becoming a necessity. You have to have these systems in place, this network automation in place because there's becoming less and less people aware of how everything should be working together.
And if you don't have these processes and you go to leave for another organization, you're leaving that snowflake out for someone else to hopefully be able to figure out and fix. And it's, it's not just that the network is growing, it's the, the demands for, uh, change day to day is also growing too. Mm-hmm.
The networks used to be static. You know, you built it and then you replaced it in seven years. It's not static anymore.
We're getting daily or, or regular demands for small changes that are necessary for various equipment. More and more things have to be plugged into the network one way or another, right. In order to work stuff we never thought of before as being networked devices.
Well, the, yeah. I mean, the pa the pace of change within the business environment overall is just doing nothing but accelerating. And of course, the dependence on connectivity and moving data from one place to another, right?
And, and, you know, uh, uh, internet enabling all the things iot and OT and, uh, you know, process control. All of these, you know, that's done nothing but just hockey stick in the last, you know, 10 years or so, 10, 15 years. So, so yeah, you're absolutely right.
Like the pace of demand, the only way, it's the only way, right? The pace of the, the pace of change and the demand around that from the business and, and frankly, the importance of the network to the business, whether they recognize it or not, has done nothing but accelerate. And so that's recogniz, that's what drives It, goes down, they Cognize it when it goes down, you know, the, the network is truly what's generating revenue for the most organizations.
'cause if the network goes down, work's not getting done, they're not making money. So that uptime is crucial. And by having that standardization in the network, you're hopefully preventing that downtime from happening.
Yeah. That the resiliency. Resiliency, yes.
Absolutely. Yeah. And the biggest thing a network operator can do to prove that value to the business is actually, you know, get those processes and tools and automation tool chains and everything in place to be able to actually keep up with that face.
Because when you no longer look like the laggard when it's no longer while we're waiting on the network team again, right? Right. To help enable this new application or bring up the new remote office or whatever it is, when that can be ideally even pipelined into a, a provisioning process or, or an overall automation job.
But even if not, just something that can happen quickly, uh, proactively, right? I i, that's where you start to drive genuine business value. And that's when the business starts to recognize the importance of investing in its technology.
The technology has to be an asset and an advantage, not a drain on, you know, the, the, uh, you know, on the p and l, right? And when you have the mature automation system, self-service then too. So the actual customer that needs to bring on the, that new service or new devices can actually just go in and say, here I'm in this, uh, this location or this room, I need this service for, for this device.
And boom, the automation runs in the background and within a reasonable amount of turnarounds and, and approvals, the, the, the thing is happening, right? With, with the user initiating it on demand for when the user needs it. Right?
But that brings you back full circle, right? If you, if if everything is a snowflake that you will never, you'll never get there, you'll never get to that, to that place, right? Because every single time there will be some exception, right.
That, you know, kills the automation. Because One of the things we, we forget, uh, now in this, this growing cloud eras, I know the number one reason AWS entered into organizations when it was, when it was in its infancy, was, I can, I'm a IT person in a company somewhere. I want to do something today.
And my IT department told me it would take two weeks, two months or whatever to get it done. So what do I do? I swipe out my company credit card and I go to AWS and it's spun up today, Right?
Well, I mean that, that's the reason behind pretty much any shadow it right? Is it's faster, it's, you know, more agile. Um, occasionally it's cheaper, but it's, it, it usually has to do with keeping up with the pace of business, right?
So Absolutely. You know, you need to get the, uh, uh, you know, the, the intended, you know, IT environment to be able to keep internal services, the internal services to be able to be agile, that available, that you know, that, that flexible, and that's how you get rid of the shadow it. But you're absolutely right.
I mean, that's, that, that was a big, that's a big reason that we saw the explosion of cloud services, right? So one of the things I like to do when I record these podcasts, I like to give a little bit of a takeaway for people at home. So I'm gonna ask you professionals, what's one thing that someone can do today to reduce their snowflake, to melt it a little bit, if you will.
One thing that you can implement, one thing you can change, one thing that you can do that will prevent this from being a stopper of work in the future. Throw money at it, Throw money. You are so American man, But Dakota, I feel like that's a solution to a lot of problems.
Yeah. It, it's really, I think having an understanding of the problem, you know, you need to fully grasp what's going on. And again, with automation and stuff, you know, there is no one solution fits all.
There's several different ways you can perform this task. Um, it can be as simple of, and you can go with third party solutions. You can try to make something in-house, or you can go with a baked in solution, such as like, something like Cisco Meraki is, in a sense, one way to get rid of the snowflake because you no longer have that granular control built in.
Um, platforms like Cisco, Meraki and others out there really kind of give you a set of rails you can kind of stick to, and you really can't go outside this box. And it enables smaller teams to, you know, be able to pivot quicker and spin things up before they're even hit the site and stuff like that. So if you're in, in the network now that's, that's in this, this state.
So you're in the brownfield, the number one thing you can do is get your documentation, update your documentation, and then study it with an eye to this modularization. What are the key components of the services your customers need? Not what your equipment can do.
Mm-hmm. But your end to end services that a customer is asking you for in, in your organization. Break those down into their modular formats.
And once you, at that point, now you're ready to go out to the, uh, example designs from the various vendors you're dealing with and see what of these modules, what of these example designs are the, the right ones for the services I need to deliver. And now you're in a position to build yourself a path to get there. Maybe starting by just saying, new services are built in this new way, and eventually as gear is replaced, I move the whole network into this mold.
Yeah, Yeah. I would, I would agree with that. Right?
It's, it, it's, it's exceedingly difficult to go back to everything that's already working or mostly working or working, but maybe inefficiently Right. And make a lot of change there. It's far easier to just draw a line in the stand and say, Hey, from here on we're gonna follow these standards.
Yeah. Right. Um, and, and, and what I would say there in terms of, of, you know, the, the suggestion, the recommendation is, you know, be bold, right?
Be, you know, be a little bit rigid in the fact that there's reasons we need to follow these standards because they will pay dividends down the road in our ability to mm-hmm. Uh, you know, uh, uh, reduce mean time to resolution, um, increase availability, uh, reduce cycle time for new deployments, right? You're Making an investment, Right?
You've gotta kind of build that business case a little bit with your leadership because it does require a bit of an investment. Maybe it's monetary investment, maybe it's a time investment, maybe it's a people investment. Yep.
But there does need to be some investment and the payoff is there. Um, but you need to articulate that This isn't an overnight plan. You're building, this is a decade long probably goal could be depending on the size of your organization, but, you know, um, this, you're not gonna make the change overnight.
But by doing the legwork now and having a plan, like you said, draw that line in the sand from this point forward, you're just making it easier on yourself, making it easier on the organization, and, um, it's just gonna pay back dividends in the long run. Mm-hmm. Well, as you can see, um, even the best intentions in the world, uh, sometimes end up creating something unique and difficult to manage.
Usually it comes down to uncertainty and speed. Um, we don't know quite what we're doing, but we know we need to do it as quickly as possible. And so we get the minimum amount of work done to make it work, and then years later we go back and curse ourselves from making it happen.
So, as you can tell from these, uh, group of experts, you know, you need to have good documentation. You need to be ready to say no when someone tells you they need something custom and different. And you need to be ready to understand your requirements so that it will do exactly what you want it to do in a way that is repeatable, maintainable, and automatable in the future.
That will just about do it for this episode of The Tech Field Day podcast. I wanna thank each and every one of our guests for joining us. I also wanna thank you for listening.
Remember that you can find us on our YouTube channel. com/podcast, and you can follow us as a podcast in your favorite pod catcher. Uh, just search for the Tech Field Day podcast.
Leave us a rating and a review so that everybody knows kind of what we're all about here. com. We look forward to the next opportunity to share an interesting premise with you on the Tech Field Day podcast.
Until then, stay tuned.