Cloud Discipline for On-Prem Infrastructure: Nokia IT’s Requirements for NetOps
Ahmed Abutaleb explains how Nokia IT defined requirements for a next-gen, global data-center network and why they pursued a NetOps approach. The goal was to operate on-prem infrastructure with cloud-like consistency: simple intent expression, repeatable automation, and predictable outcomes. Their first attempt (on a different platform) fell short, but moving to an open, Kubernetes-based platform—and surrounding it with open-source tooling—let a small team (3–4 engineers) learn fast, customize, and automate deeply.
Architecturally, Nokia shifted from inherited “flat” designs to a modular, flexible topology that can stitch in remote sites and support dual, geo-redundant data centers in each region (Europe, U.S., Asia) for disaster resilience. The tooling had to understand and enforce that flexibility at scale. This modernization, part of Nokia IT’s broader NetOps journey with Nokia SR Linux and Nokia Event-Driven Automation (EDA), contributed to major operational gains, including roughly an 80% reduction in trouble tickets.
This video is number 3 in a series of 6. To see the other posts, visit: https://techstrong.tv/videos/modernizing-the-data-center-nokia-its-netops-playbook
Transcript
So as you, you know, made these decisions, you had certain things, you know, and with your view from an architecture perspective, real requirements to drive real network outcomes and performance outcomes, what are some of the things that ended up on your list of things the new network infrastructure must do? What were some of your key issues that you were addressing through your requirements gathering? First, you know, we were getting also exposed to, uh, um, public cloud, you know mm-hmm.
The automation there. We wanted to drive our on-prem, like we do, like in public cloud, you know? Sure.
Uh, make it more, more consistent automated, and also, uh, simpler to, to manage, you know, simpler, to specify what you wanna do and get consistent results. Um, so we wanted on-prem to look like cloud. Mm-hmm.
But we, we ended up with, even on-prem, I think in a bad little bit better state than cloud because Wow. We had more control there. Sure.
We had more control there. We were able to, to control the underlying, um, management platform, uh, where we couldn't do it in the cloud. Really.
Um, so, um, would You, would you call that, was that, was that an objective or was that a happy side effect? We had it as an objective. We didn't achieve it with our first iteration.
Our first iteration was a few years ago with a different platform. Sure. And we, we didn't really achieve that part.
Okay. But with, with our newest platform, it gave us the, the, uh, it's op, you know, being open source. Sure.
And with lots of tooling, it enabled us to put around it the, the, the, the solution that will achieve this objective. Sure. Yeah.
You, I, I know from previous conversations, the ability to use more open source tooling really gave you more options and give you more ability to customize. Yes. And, and you, I mean, what we did, I mean, we did a lot of innovative work really.
And, and we started as we were a small team. We were maybe three or four engineers at the beginning. Wow.
And when we started, we really didn't know any of the modern networking techniques. But having an open source tool, you know, then you can, you can put somewhere, you know, you can look on the internet really, or, uh, you know, other people would have faced what you are facing. It might be not be not related to networking, it might be related to something else.
Sure. But having an open source tool that people have, have, have touched this open source, which is Kubernetes, where our tool is based on Kubernetes and, uh, you know, millions of people are using Kubernetes and they've built tools that enable it to communicate with this Kubernetes cluster and do interesting things that we have used in our solution. Got it.
I don't want to, um, under count, you know, the, the, the global scope of the network, yet it wasn't just like one data center, a pair of data centers. Talk, talk a little more about global topology and, uh, what data center resources you had to bring all in line together. Okay.
So, um, I mean we've got I data centers in all content, um, Europe, uh, Europe U us, Asia. Uh, they're very complex data centers. They're dual data centers.
And, um, got applica, you know, one is active standby of, of the other. Um, it started with a very flat, flat architecture. You know, we've inherited a flat architecture and we continued thinking flat, but then we made it a much more interesting, uh, modular architecture mm-hmm.
You know, of the data centers where we can add remote data centers together. And we wanted, you know, uh, uh, this flexible architecture, but we wanted the tool that will understand this flexibility Sure. And enable us to have this modular flexible architecture.
So, um, that is, that is what we had in the architecture side. And when you say dual data centers, you mean dual data centers in each geographic location? Like, you know, like within Yes, we have it's dual geo-redundant data center.
Yeah. So, so they are, they are in the same region, but, uh, geographically apart to back up each other. Yeah.
In case of a disaster, Separate power, um, volcanoes Separately, et cetera. Yes, yes. Yeah.
Yes. Not to imply that anything is actually in Iceland. Um, but, uh, I know people have had routers taken out in Iceland because of lava.
But anyway.