57. Managing Hybrid Cloud Networks Complexity with Infoblox – Tech Field Day Podcast Spotlight Series
Managing hybrid-cloud networks is complex due to differing architectures and naming between on-premises and the multiple public cloud platforms. This Tech Field Day Podcast episode features Glenn Sullivan, Senior Director of Product Management at Infoblox, Eric Wright, and Alastair Cooke. Each public cloud has a unique management console and network management paradigm; none provides deep integration with each other or with on-premises networking. It is left to individual customers to assemble a jigsaw of pieces into a coherent whole. Customers may not plan to use multiple public clouds, but through different project requirements or mergers and acquisitions, most large organizations find themselves in a hybrid multi-cloud environment. Combining fast-changing public cloud applications with on-premises applications further complicates network management, requiring an automation-based approach. Infoblox UDDI (Universal DNS, DHCP, and IPAM) provides a consistent, automatable interface to manage and operate basic network infrastructure across all enterprise locations. UDDI includes bi-directional operation where changes using cloud-native consoles are visible in UDDI and vice versa.
Transcript
Hybrid multi-cloud networkings. They are so complex and yet so fragmented. It feels like a jigsaw puzzle with pieces that don't quite fit or maybe pieces from multiple puzzles.
In this episode of the Tech Field Day podcast, I'm joined by Eric Wright and Glen Sullivan. Glen is the general manager of a product at Info Info present most recently at day 20. But also join us for all of the insights that we have around the complexity of managing hybrid multicloud networks.
Welcome to the Tech Field Day podcast, where we bring together a group of it technical experts to discuss a single idea about key concepts in the industry. This podcast features a variety of perspectives from members of the Tech Field Day delegate community, and is often recorded in association with one of our events, tech Field Day as part of the FU and group. And this podcast is also pub published on our sister company Site Techstrong tv.
On this episode presented by Infoblox, we'll be discussing hybrid multi-cloud fact that it's a jigsaw puzzle full of pieces that don't fit together. Lots of complexity out there. But before we dive into today's conversation, let's start with some introductions.
Uh, first up on our panel today is Mr. Eric Wright, my friend. How are you today?
Very good, Alistair, thank you for letting me join. Uh, happy to be here today. And, uh, for those that are brand new to me, my name is Eric Wright.
I am otherwise known as at Disco Posse. I'm the co-founder of GTM Delta, and, uh, I create content and, uh, noise for a living. And joining us from Infoblox is Glenn Sullivan, the Senior Director of Product Management.
Hi Glenn, how are you doing today? Hello, Alistair. I'm good.
Thank you very much for having me. And of course, I'm Alistair Cook. I'm an event lead here at Tick Field Day.
And I think one of the themes we saw through Cloud Field Day, and particularly through the real world deployment of applications on the cloud, is that gluing together a hybrid multi-cloud environment does feel like you're assembling jigsaw pieces and often building solutions on the cloud is about assembling together jigsaw pieces. The pieces from a single cloud are designed to go together, but assembling pieces across multiple clouds and on-premises environments feels like you're trying to assemble three or four or five different jigsaw pieces into a single picture. Glen, are you seeing customers really struggling with this?
Is this am am I making stuff up or is this real that it's hard? No, this is absolutely something our customers, uh, are seeing as a challenge on a regular basis. And, um, it's, it's even just in a single cloud environment too, right?
Multi-cloud makes it worse. But even, even if you're sitting there thinking, well, I have everything in, you know, AWS or Azure or GCP, you probably have dozens if not hundreds, maybe even. Uh, I think the most I've ever heard is like 3000 accounts in AWS.
Um, so even in those environments, everything seems like a very different, uh, environment. 'cause you have different regions in different, you know, AZs and different, you know, different zone types. Um, so what our customers are struggling with, you know, kind of out of the box is getting consistency between all of those different accounts and all of those different cloud environments and rectifying it and reconciling it with what they have on-prem.
Uh, because it's not the same tools, it's not the same expertise. And, uh, like I was talking about in the cloud field day presentation, it's, it's not even the same terms. So you have to have a cheat sheet to, to determine like, what is this thing called in, you know, on-prem or Azure AWS or GCP?
'cause even the terms are different, right? It just, the very fact that they keep changing the names of products on us between the clouds. And then on top of that, like I often say that multi-cloud is not a strategic choice, but adopting it safely is.
And one of the problems we have is that we acquire clouds. You often, you don't say, I wanna build one app across four clouds. What you have is four apps.
Each is in a single cloud and maybe you acquire another company or you run a SaaS product and it is tied to another cloud. And so now you've got all these things that need to work together, but the way you manage them does not in the outta the box way. And then on top of that, the human aspect, right?
I've got team members that, like I said, just is it a Route 53? There's no route 53 over here on, on Azure. You're like, oh, okay, what do we call it over there is Azure DNS, it might be today.
And then, you know, the, the elements and the artifacts are different. It's not, not like there's one JSON, like bind was happy, bind was a good thing. Everybody had the same format.
And we've, we've steered far from that. The core database is the same, but the way we manage it is not, Yeah. And, and people often ask me like, what is the optimal, um, you know, organizational structure?
Because I get to talk, you know, the luxury of my job is I get to talk to, you know, hundreds of customers about how their different cloud environments are set up. And, um, I see everything from like cloud centers of excellence, where you have one team that has, you know, governance, architecture, network, uh, cloud teams, cloud ops, all on one team. And that works well in some environments.
And then in other environments, it's siloed down to we have an AWS, you know, development team that, that has their own governance and, and an Azure, you know, development team that has their own governance and the network team is, is completely separate. And there's no right answer. It kind of just depends on your, your needs and how you organically or non-organically, you know, developed your cloud infrastructure, right?
And like you said, with m and a, right? If I acquire, if I'm an AWS shop and I acquire somebody who's primarily in Azure, you know, I'm, I don't even know what I'm getting half the time, right? Um, I was talking, you know, with, with cricket, uh, the other day on a webinar about this.
And, uh, you know, the m and a teams, the tech team is sometimes the last to find out that the m and a is happening, right? So then it's like, hey, the, the contract's signed, the ink's drying, uh, now you need to go figure out what you're getting. And, uh, you know, the first thing you figure out is, hey, they're using 10 slash eight for all of their IP addressing.
We're using 10 slash eight for all our IP addressing. So re iPing is inevitable. So then it's just a matter of how do we make the re iPing, you know, uh, uh, less impactful than, uh, you know, than it would be otherwise.
It's not, it's not a matter of, Hey, we're just gonna find some copacetic, you know, IP space that in between that works. It's, it's re iPing is inevitable. And no matter how much we've loved the idea of, of variable length subnet masks, we've had them for decades at this point.
We love 8, 16, 24. You know, like it's, we just love those natural ending points. A hundred percent, a hundred percent.
I mean, even, even getting, um, you know, someone to change the default cluster size from a slash 24 to a slash 27 is a challenge. Um, and this is an, this is like an exact, um, scenario that I've seen time and time again happen between cloud teams and network teams, right? Is you'll have a, you'll have a cloud team who's using Kubernetes, you know, A-K-S-E-K-S-G ke, doesn't matter what flavor, and maybe the default pod size is slash 24, okay?
And get 255 addresses, you know, minus two for network and broadcast. But, um, maybe I only need 30. And it's like, well, what are you doing with the other 200 plus ips?
Well, you're wasting them, and it might not be a big deal because maybe you don't need that many clusters, but if you need more than, you know, a handful of clusters, you're gonna want to reduce that size from a 24 to a 27 so that you can, you know, reclaim all that wasted IP space. And if you think about the complexity of that situation, right? Um, the, the cloud engineer, the, the, the, the Kubernetes admin, they, you know, they don't really speak IP much because most of that is abstracted underneath the, the coverage for them.
So they've gotta get on with a network team, you know, who speaks siters and who really understands, you know, IPM you know, IP address management, and they have to kind of like look at both consoles, right? And say, okay, I am looking at Kubernetes, it's using this, I'm looking at, you know, ipam, this is what the toss 24 versus slash 27 gives me. It's actually really, really complex to figure that problem out if you don't have a holistic view on both sides.
And just one thing that you brought up earlier, uh, that will lead probably into something, I know Alistair has had a lot of experience with the idea of Conway's law in that organizations build communication patterns that mirror the systems they build. And so if you've got an Azure team and AWS team, a VMware team, a network, you know, team, a field network team, they will act and operate as if they work for different companies and they expect their local toolkit to be the toolkit of record for everybody else. And it's often like, I know it's works fine on my machine.
And, and that is an unfortunate pattern. And you know, Al I know you've, you and I have spent a lot of time in a lot of environments. You know, you, you know this too well.
Yeah. The other thing I wanted to reflect on was a couple of elements of how this has been done badly in the past. And we've all seen environments where the IP address management was a big spreadsheet, and it wasn't even a spreadsheet full of formulae.
It was search and replace within each tab for the subnet that you have. Um, I'm looking at a particular network team I work with in the not terribly distant past who, who had a, a large campus network that was exactly that, and it was so painful. The other element I wanted to hit on, uh, is Glen, you, you talked about the, you know, even within a single, um, provider, single AWS environment with, with lots of, um, lots of accounts and therefore lots of VPCs probably in there.
Uh, it's still a complex thing because often we've used scripting and declarative tools, uh, scripting and, and, and aligned declarative tools and just made managing the network part of our fragile scripting or hopefully less fragile declarative tools. But it's been a, a home baked solution that we understand. And this kind of hits a little on Eric's point of, uh, one team who's building the automation on AWS will build it in a way that makes sense on AWS and, uh, will have a whole set of structures around that.
And then the team that's building on premises will have a whole different set of structures. And so there's, there's elements where we've built our own tooling to suit our own requirements, but it tends to be very fragile and not very cohesive. And I think that was really a place where we start seeing pain, right?
We move from that idea of automated all of the things that I know Eric and I were on, that that, uh, rollercoaster as, um, PowerShell was, was new as the, the thing to automate all the things with. And we very quickly realized that really we're automating things because there isn't a good product to do the things for us. And really that automation that everybody has to build, well, that's, that's actually an opportunity for a really good product that takes away the need to build the automation.
Yeah. That's the way we see it with, you know, what we're doing with Universal di because it, it's, it's actually twofold, right? And sort of the obvious thing is, okay, I've got a bunch of humans causing errors.
Um, I'm gonna automate everything because that's gonna prevent conflicts. That's gonna prevent from exhausting ips, that's gonna prevent from, you know, this bad scenario to that bad scenario. So let me automate everything, but there's, there's two, there's two outcomes that are not super obvious from, from that decision, right?
One is, I am signing up for managing the API structures of all the different, you know, third party systems that I integrate with, right? So I mean, you could have a relatively simple environment blink, and now you're managing APIs in AWS Azure, um, maybe you're not even at GCP yet. Microsoft ad, um, you know, Infoblox, maybe some generic bind out there.
You know, this is really, really fast. You've got 5, 6, 7, 8 different APIs that you're managing. Um, that's kind of one bucket of, of, you know, tech debt or tech, I won't call it tech debt, but like, you know, something that needs care and feeding.
The second thing that you have to handle is the person who figures all this out. There's probably one or two people in your organization that have the skillset to do all of this. They could be using Terraform, they could be using Ansible, they could be using a combination of all of that.
Um, but the kind of universal answer that I've received from most people I talk to is there's one, maybe two people at the entire organization who really understand how that 50, 60, 70 lines of Python really work. And as soon as they, and it's a hot market right now, as soon as they get a different offer or as soon as they decide they're, you know, they wanna move on, that knowledge is going out the door. So it's not just the managing four or five different API structures, it's also the institutional knowledge that you bake into the automation.
So native automation, you know, is always gonna be better interface with a third party, pull the data in, get next available, you know, make, make, make your single IAM source, do the hard work for you, hard Work when what we find too is like, I used to always see, like even we have some IAM platform in place, their read versions, like they're generally not really, they're not really management. It's more like it's a nicer version of an Excel spreadsheet with a bit more searching. Like there's some dynamic updates, but they're, they, if at best they were unidirectional, right?
Like it's reading the data, getting you one place to look at, you know, two or three sources, but in the end it was just a, a very slightly, you know, elevated abstraction from just managing right in the final interface. And then the biggest problem is then you have to then go elsewhere to actively manage that. And that's, that's the problem I think that we used to always bump into, There's two things, right, um, that are critical when we're doing anything within a universal, you know, IAM perspective.
One is we can't eliminate a swim lane, right? So if, if the cloud team does VPC provisioning and maybe they use AWS ipam, great, I don't take that away from them because as soon as I tell a cloud team, they're not allowed to use their native tools, I might as well tell 'em they need to deploy an appliance, right? They have antibodies and, and everything else and heartburn against doing that.
So that's critical. I can't eliminate a swim lane. The second thing is, like you said, it has to be bidirectional, right?
So, um, anything I do in one has to show up in the other, and that's part and parcel for that swim lane, right? As as long as I make an update in a WSI pick it up in, in my UDDI console, my UDDI portal, and the same, same as vice versa, I make a change in UDDI, it pops up in AWS, so it has to be, it has to be bi-directional otherwise. Um, the third thing that it has to do, I know you said there's only two, but there's really three, um, is it has to do the get next available.
Like if it's just a pre, like you said, if it's just a spreadsheet, right? Then it's not doing me a whole lot of good, sure I can visually look and see, hmm, okay, what's the next slash 24 inside of the slash 16, but it needs to have an API structure that says, Hey, given, given everything I know about, you know, the environment, here's the metadata I need to be able to pull out the next available ip, right? So it could be, Hey, give me a slash 24 in US East region that is PCI compliant.
Here you go, API, right? Not looking at a spreadsheet, not looking at the ui, it has to unify the data structure so that when I do get next available, it gives me the real next available. That's a critical piece of ipam.
And I love that you highlighted the idea before of like the, you can have incredibly talented people, uh, for fans of the Phoenix project. We call them Brent, right? So shout out to Jean Kim and, and the crew there, but that is the one person that they're incredibly valuable, but they're also incredibly dangerous in that the amount of risk you carry by having them be such a critical bottleneck to knowledge.
And, you know, it's not that we shouldn't celebrate that automation, that knowledge, but like, I would rather automate a tool that already adapts to changes in outbound APIs. Like, like you said, I've, I'm a huge fan of automation, Alistair, and I've spent tons of time buried in PowerShell stuff till, until I couldn't even see. And what we end up finding is that you build incredible solutions that are bespoke and that's great, but then yeah, new version power show comes out or the endpoint API gets to deprecate some function and all of a sudden I'm seeing new elements come in and now I have to go back relearn the code that I barely knew the first time and then I've got a context switch for that.
I mean, deity bless cursor and all these amazing things that can help me get there faster now. But still, I'd rather just say, I know for a fact that this, I can buy a product that I know will continuously keep up with the API changes so that I don't have to create braking changes. Everything I do is like I've got one API that I need to care about.
So if I wanna automate, I'm automating against UDDI, right? I'm automating, I'm automating one place much safer. And like a anybody who's into compliance, they should know.
Like the more things you try and work against, it's, it's a general IT risk. Yeah. I mean, fundamentally what you don't want to be authoring the, the, the automation that interacts with the API, you want to be authoring policies about what it should achieve for you by interacting with that API, you want to be able to say, here are my standards, just go and implement my standards because that's much closer to business knowledge that is far less likely to be locked into that one Brent, who then gets hit by a bus and, uh, you find yourself out of business six months later because the network falls apart and you have to redesign from the beginning.
That's right. There's also a security element too, right? Like, um, with, with such a, you know, such a hot market right now, you know, there, I know one of the challenges I hear on a constant basis is onboarding and offboarding.
Think about if I've got eight different nine different systems that I have to update the API keys. I mean, we don't even who, who, who updates their quarterly passwords as often as they should, right? Much less, you know, revokes all the API tokens every time someone leaves their company, right?
So It's like to tell everybody that says they got a TLS cert, the first thing is like there's only one time that you ever update the cert, and that is about 22 minutes after it expires, because that's when you find out that it expired. We're not good at where it people have A DHD by nature, we are not good at following long-term rules. That's right.
This is why alert and alarms are your, your coping factor here for not just for A DHD, but also for certificate expiry and all of those kinds of routine maintenance things, uh, on active notifications. I love, And, and just actually a question you gave an Infoblox is in large environment and small environment, you know, like you've got a good variety of types of industries and like, what does multi-cloud really look like? I know my, I'm sort of biased in how, what my view of it was, but like multi-cloud used to be, occasionally I would bump into people that had a second cloud and they kind of grudgingly took it on.
But are we actually seeing in practice that it's just cloud, cloud is just a ubiquitous thing and we just have at least two clouds in, in a lot of environments? Am I wrong in thinking that that's more common? It's like, it's like 90 plus percent of organizations have, are in multiple clouds and they end up there one of two reasons.
One is 'cause it organically happened. Um, they had a different dev team who needed something that was really, you know, specific and solid in GCP. So they stood that up.
You know, you could call it shadow it, you could call it, you know, swiping your credit card to enable, you know, cloud accounts. But in reality it's just the way business gets done, right? So you either find yourself that or you find yourself in a mandate.
Um, multi-vendor mandate has always been a thing. It kind of goes in and out a phase for what it means. Does it mean that I have two switch vendors for top of rack or does it mean that I have, you know, some maur so that I can play AWS in the penalty box when I need to?
Um, I don't see that as often. Like you would think that would be it. Like, oh, there'd be some, you know, ex executive C-suite level mandate to be multi-cloud.
That's not what I'm seeing. I'm seeing a mandate to be cloud first, a mandate to be SaaS first, but then the multi-cloud part is just, it's just organically happens and, um, the smart organizations don't fight it. The smart organizations realize that this is what we're, this is what we have to deal with, so we have to be adapt to the business needs.
It's about as realistic as me, you know, getting a second girlfriend to keep the first one in line. Like, it's just not like that. These are not realities of like, we can't, you can't just say AWS yeah, you bugger something up.
And so I'm, we're swinging over to Azure, you know, like I can probably barely even log into Azure in the amount of time that they'll be laughing at me, that it's just, you can't possibly move your whole estate like that. And G ccp, like look at GCP with the cross cloud networking, the thing that they just, they just announced to it Google next, right? With the Google Cloud wan, right?
Which does happen to use our stuff as well for DDI under the covers. Um, but the thing is, is that, you know, even the cloud teams are recognizing, the cloud providers are recognizing that you're, they're not gonna win a hundred percent of your business and that's not in their best interest to insist that. Um, so now they're building tools.
Um, and then MCN is a thing too, right? You know, your aas and your aviatrix's and things like that, like it's, it's a real thing. Uh, so it's not going away.
We are running ourselves towards the end of time and as usual, this would be a great conversation to carry on over a, a full lunch or a full week. In fact, it would be great to just hang out with you all. Uh, but I'd like to, uh, give people the opportunity to follow up with you and continue the conversation with you once moved on maybe from listening to this podcast.
Uh, Eric, where can people find you and carry on to this discussion around multi-cloud and hybrid cloud? Well, uh, yeah, thanks for letting me take a part of the conversation. Always fantastic.
Uh, I'm at Disco Posse all over the place. You can find me on LinkedIn. Uh, you could search for Eric Wright.
There's probably a bunch of them, but there's only one Disco Posse. Uh, and of course on XI may even be on Blue Sky, I can't remember. com, you know, uh, likely you are a customer.
We have lots of customers, all big and small. So you likely already have an account team. Uh, me specifically, I am, I'm on LinkedIn.
I'm fairly active on LinkedIn. Uh, so reach out via Glen Sullivan on LinkedIn. Um, there's a few of us out there, but probably only one guy that works at Infoblox who used to work at Snap Route and other companies before that.
So, uh, Glenn Sullivan on LinkedIn. Oh, and of course I'm Alistair Cook. You can find me on LinkedIn and as well as across a whole bunch of the, uh, Textron and uh, Futureum group properties as well.
Uh, I think we will be hearing lots more about the complexity of hybrid cloud networking. It's certainly something that I think is important as we're doing more Cloud Field day events, eye out for those in the future. Thank you for listening to this episode of the Tech Field Day podcast.
If you enjoyed this discussion, please dis subscribe on YouTube or your favorite podcast applications so you don't miss an episode. Do consider giving us a rating and a nice review. This podcast was brought to you by Infoblox and Tech Field Day, part of the Futurum Group for upcoming episodes and more events and all of the fun things that we do it at fu uh, tech Field Day on rum.
Head to the tech field day com slash podcast or view us on Techstrong tv. Thanks for listening and we'll see you next week.