The Age of Operations: Third Party Views of DCF Innovation
Scott Robohn of the consulting firm Solutional provided a third-party, operational perspective on data center innovation, based on a collaboration with The Futurum Group and Nokia. He introduced his background as a former network engineer and Tech Field Day delegate, now focused on NetOps and AI adoption. Robohn’s central thesis is that with AI driving new infrastructure builds and hardware becoming normalized, the next major performance gains in data centers will come from a relentless focus on improving operations. The joint project between Solutional, Futurum, and Nokia aimed to validate this by looking at AI’s dual role: both as a driver for new network infrastructure and as a set of tools to be used for network operations.
Robohn detailed the market trends shaping this new age of operations, starting with AI as a durable technology driving massive fabric build-outs by hyperscalers and a new class of NeoCloud providers. He defined NeoClouds as specialized providers focused on renting expensive, complex GPU-interconnect infrastructure for workloads like model training and video rendering. He then argued that as data center hardware has normalized around merchant silicon and stable architectures like Spine-Leaf, the hardware itself is no longer the key differentiator. This consistency makes automation more achievable and shifts the entire industry’s focus to operations as the critical area for innovation and reliability.
To validate this operational focus, the project stood on four legs. First, a Futurum reliability survey which found that network reliability is the number one purchasing criterion for data center infrastructure, far outweighing cost, and that human error remains a major issue. Second, a collaboration with Bell Labs on a reliability model, which concluded the most significant reduction in downtime comes from fixing operations-related issues. Third, interviews with Nokia’s own 80,000-employee internal enterprise IT team, who are successfully migrating their complex, high-stakes manufacturing and office networks to Nokia’s SR Linux and EDA platform. Finally, general market engagement confirming a palpable, industry-wide appetite for a new class of automation and AIOps tools.
Presented by Scott Robohn, CEO and Co-Founder, Solutional. Recorded live at Networking Field Day 39 in Silicon Valley on November 5, 2025. Watch the entire presentation at https://techfieldday.com/appearance/nokia-presents-at-networking-field-day-39/ or visit https://techfieldday.com/event/nfd39/ or https://Nokia.com for more information.
Transcript
It's great to see you all. I'm Scott Roan. Um, I'm with Al a consulting firm, uh, that's been collaborating with Futurum and Nokia on some really interesting third party validation that we're gonna talk about for a bit here.
So I think I'm the last up correct, Andy, so I get to go long as I want to, uh, until you all get up and walk out on me, which, you know, won't be the first time. Um, I wanna bring a very operational, um, feed on the street view of what we see with, um, the emphasis on data center networking and that push towards, um, operations as being where it's at to quote, um, a good friend of mine who was mentioned earlier on a podcast, uh, last year, um, we really think there's room for focusing on operations to get that next level of performance out of data center infrastructure. That's kind of my perspective.
Um, quick background on me. I've been an internet plumber. Um, I spent a long time on the vendor side, and now I'm a recovering SE leader, I guess.
Um, running my own consulting business, uh, had the privilege of being on your side of the table as a tech field day delegate, and this is a little different, uh, gig for me today. So it's really fun to be here. Um, this work is done in collaboration, obviously with Nokia, but also the Futurum group and a hat tip to Mitch.
Um, Ashley, um, cheer. Mitch and I have been collaborating along with many other people on the Nokia side, on the futurum side, on the solution side to pull this analysis together. Um, so if you wanna jump in and say anything as we go, you should feel free, um, to do that.
Okay. Um, but the Futurum Group, um, is huge in industry analysis across all IT disciplines. Um, you know, Mitch is one of many, many, um, who are out there watching what's happening in the IT environment.
Um, and they bring very sharp analysis, um, report generation and amplification to things that are happening, again, across a different it, uh, silos. Solutional is a very small consultancy that is exists to help, uh, customers rationalize adoption of AI and operations. So NetOps is one of our, our, our, uh, key focus areas, but we are looking at adoption of ai, uh, across many different IT disciplines as well.
So together we have this very synergistic view of what's happening. Um, one of the things I get to do is talk to a lot of operators, and I'll, uh, expand on some of that as we go. Um, part of what I do is a podcast called Total Network Operations.
I'm not much of a marketer. You know, the, the, uh, title of the show is pretty obvious. And, uh, thanks to Ethan and Packet pushers after, um, engaging, uh, irritating Ethan for a long time, we finally got it off the ground.
So we're, uh, we're having a lot of fun with that. You were very persistent, Was Thank you. Thank you for stating that so nicely.
Okay. So the outsider perspective here, um, just to set the stage for the conversation, this is all under the assumption that AI is a durable technology. I hope nobody needs to be convinced of that at this point, right?
Um, are we early in the game and its development and application? Absolutely. Right.
Um, but you see the investment, it's real. I'm not here to debate, you know, the B word bubble. Um, but we've seen this movie before and we have seen periods of extreme investment and we've seen overbuild of capacity and the capacity get absorbed eventually, right?
So that's a scenario that we could see here. Um, but we do see the, the continued push for the development of, uh, training facilities for models driving, um, fabric build out, and driving innovation in, uh, fabric networking technologies, right? Um, inference is not just here.
Um, it's going to keep going up into the right and that's gonna dominate our decisions as we go. I mean, think about it. How many, how many real AI users are there today versus how many there will be next year and the next year and the next year.
So, we'll see inference and infrastructure for inference play more and more of a role as we go here. But this is all building a foundation for, for what's to come. And, uh, you know, there are applications and problem sets that we don't even know yet that are gonna be addressed by AI technologies, right?
It's gonna be, you know, what's the killer app for the internet, right? From 30 years ago, you know, was it video? I don't know.
We can debate that over, uh, happy hour later. But, uh, we're gonna be discovering what's useful here in AI for some time to come. Um, it's obviously driving network infrastructure for ai.
We see this concentrated in some key places. You know, obviously the hyperscalers have, uh, have engaged and are building out their own infrastructure. The foundation model companies, you know, have some of their own infrastructure or they're renting it from hyperscalers.
And now we have this new class of providers called the Neo Cloud providers. And the emergence of Neo Cloud has been fascinating to me, and I think it builds on two really important things that you all around this table have helped build, build over the last 30 years. We had the internet build out, which provided all the connectivity that we needed that allowed the emergence of cloud, which brought us aggregation of resources and new business models that both feed into what's happening with Neo Cloud now.
So if you think about it, it's kind of no surprise. Denise. Hi Denise.
Don, could you please, uh, explain Neo Cloud one was using that term before too, and we were just kind of discussing. So what's the neo cloud, Um, cloud providers focused on GPU interconnect for GPU workloads. I didn't just say AI workloads.
Um, model training is one special application that works really well on GPUs, but, uh, video rendering and uh, computer graphics, right? That's where GPUs came from. Mm-hmm.
And Hollywood and the special effects industry are an important, um, source of business for these aggregations of clusters. Um, I used to have two other really good application ideas, but does that make sense? Crypto Mining?
I, uh, yes. That was, that was one of them. And that's also still real and people are willing to pay for those services.
Um, But okay, that, yeah, that makes sense. Thank You. So that's what they are.
Lemme tell, say a little more about the why. So GPUs are really expensive, um, and we see an, an increasing number of specialized technologies in terms of cooling, delivering power, getting energy to the facility and deliver, distributing the power within a facility and getting energy wholesale to, to a building. Those are actually two different things.
Yeah. Um, all these things come together that create special engineers and operators that your average enterprises really can't afford to hire and maintain. So just like we have the cloud emerge, now everybody, you know, with a credit card could have their own IT department.
Mm-hmm. Right now, anybody with a credit card or maybe a larger PO amount, uh, can rent time, you know, rent time in a neo cloud or get a dedicated section that they don't have to maintain. Got it.
Thank you. Great question. Um, there you go.
Much better than My ex explan. Yeah. It was better than his ex.
Much better than my ex-wife. Well, I, I doubt that, but, uh, um, I don't wanna leave out select enterprises, right? There are enterprises that are gonna grow, grow, build their own, um, model training infrastructure.
They're gonna have special security and privacy requirements, and they're gonna have funding to be able to do this. But I think you're gonna see more and more model building done in neo cloud type environments. So, um, that's all on the network infrastructure side.
Um, then there's the use of AI tooling for network operations. And that's really where the focus has been on EA today, right? The tooling that's embedded here.
Um, so that's really gonna be the focus of the rest of the discussion today. Um, you'll see in this, um, survey, we're gonna reference here that reliability, network reliability for data centers is really the number one buying concern of anybody investing in new data center infrastructure. I see some heads nodding.
Okay. So maybe we're onto something here, right? Um, other things to set the stage here, hardware is essentially normalizing.
Scott, what do you mean? Great question. I'd love to tell you.
Um, it's merchant silicon everywhere and merchant is often a, you know, a synonym for Broadcom, you know, different flavors of Broadcom. There are other merchant providers that are emerging too, are getting more traction, right? But, uh, we see different vendors onboarding merchant silicon and using that as the basis for their, for their data center infrastructure, right?
And no key, no exception. Noki has been doing a great job packaging, uh, merchant silicon in the IXR product line for quite some time. So, um, we do see a shift in stabilization of architectures.
I'm saying this really carefully. I picked the word stabilization, um, uh, with intent. Um, it's not all the same everywhere.
You know, you have spine leaf, which is really common, but sometimes you have super spine, sometimes you have funky butterfly designs, but there's great modularity in build out of data centers, and that's great for automation. Like, one of the problems we have in automating most of the rest of the network is there's too many snowflakes. And that's not a comment on any one of you in this room, right?
But special cases that I have to do something with this script to do automation in this piece of the network, the more uniform I can get, the more I can templatize configs, the easier it is to automate in the long term. Agreed. You wanna argue with me about that?
Hmm. Okay, later. That's fine.
Um, so consistency is good for automation. Um, lets me have fewer exception cases to deal with, right? Um, I think these all set the stage for operations mattering even more than it has, you know, in our lifetimes, right?
And there's an opportunity here for how, what can I do to make operations better? Get rid of the fat finger scenarios, you know, error introduced by us when we've been a little sleepy. Like Andy's first slide at the beginning of the, uh, the talk today.
You know, we do bad things when we're just not well rested or we're not paying it to squirrel attend. So here's what futurum and solution have done with Nokia on this front. I'm not supposed to only have three legs of the stool, but I have four legs for this stool.
I, I, I hope you'll forgive me. We just make sure we put it on a nice level spot so we can sit on it. We looked at, uh, constructed and executed a reliability survey.
The one thing you should take away from that is network reliability is job one, right? We engaged with Bell Labs on their development of a data center fabric reliability model. And I'm going wax eloquent on the capabilities of Bell Labs there for a little bit.
We spoke with, interviewed and continue to work with the internal Nokia enterprise IT team that is drinking their own champagne. Can't say eat your own dog food, right? Um, where they created an RFP, laid out the requirements for what they wanted, migrated from legacy infrastructure and multiple types of legacy, legacy infrastructure toward, um, SR.
Linux with IDA on I xrs. And we'll talk about what, uh, what they're finding. And then just our general market engagement.
You know, what we're having in conversations, you know, across the board, through podcasts, through events and things like that. So the reliability survey. Have you ever been involved in constructing like an industrial strength survey?
It's not, it's not easy, it's not trivial. Um, this was my first experience going through that and, you know, 17 revisions of this question, right? And you know what, it's worth it in the long run.
And I think, you know, the, uh, application of heat and pressure can result in a diamond. I think we have a really good result. Um, we futurum, uh, has this motion down.
I just got to help with advising on questions. And uh, you know, we found the, the number one result on reliability. People are making purchasing decisions for data center network infrastructure.
The first thing they ask about is reliability. And then they ask about cost reliability. And then after that they ask about reliability.
Sorry, drama for, uh, for emphasis here. Um, and then the manager asks how much it costs then something. But, you know, so thank you.
It's a great point. And uh, I'm gonna encourage you to go get the report and look at it cost. I just dumped it into the Slack channel.
Thank, but he's got read my mind. Excellent. Um, cost was way down in terms of criteria, partly because of the normalization of the hardware.
'cause those per port costs, they're really competitive across vendors, right? And so how do vendors differentiate, right? And it's not per port cost, at least not yet.
So unless you're buying Netgear for your data center, I don't wanna make anybody upset at net gear. I apologize. So, um, human error is still an issue.
It still stings, right? We still, you know, I I, my Tone Network operations podcast, I ask every guest the following closing question, what's the worst outage you've ever caused? What?
And only one guest has said, you know, I've never really caused an outage, and I'm not sure I believe that individual. Um, so, you know, these things are out there. Then there is strong interest in automation and AI ops.
Adoption is still lagging interest, of course, right? 'cause we're early in the game on this. But, but it's, it's on the top of mind for, uh, for anybody making these decisions, um, on that point.
So the 36% number, um, 36% of respondents, uh, reported dedicated AIOps tooling or interest in it. So that 36% is not, 36% have AIOps in production. It was in production or interest.
Wow. Right? And I'll let you speculate as to what Slice, where's the slider bar within that, you know, roughly one third.
I'm gonna guess it's more interest than actually adoption at this point, but I can't say that from the results of the survey. Next. Um, bell Labs, um, were you familiar with Bell Labs before hearing about them today?
Yeah, just about everybody. Okay. Um, I had the, um, great privilege about six years ago of going to the Bell Labs, Murray Hill facility, and in the lobby of Murray Hill or two statues or, uh, busts, um, one of, um, Thomas Edison and one of Claude Shannon, the father of Information Theory.
And guess who got, um, his picture with each one of them right away? You're standing on the shoulders of giants, right? And we, we think ethernet everywhere today, but Claude Shannon is where we get 64 kilobits per second for an, you know, a raw uncompressed, digitized voice conversation.
Sorry, I won't di dive deep on that. The Bell Labs people are steeped in this. Um, and they brought their expertise to looking at reliability for data center fabrics.
Um, I won't be able to do it justice here, even if I took 20 more minutes. There's a report, you can go grab it, Mitch, if it's available. The, the, the Nokia app note is out on it now, I think just the other day.
Great. I'll grab it. So that can go out there too close up here.
Then. Um, they did a great job of poking in on to all the different factors involved in, you know, what's going on with hardware, what are different operations tasks, you know, where are the opportunities for improvement there? And basically, one of the major outcomes of their analysis is the most significant reduction in downtime comes from operations related issues.
You know, initial config, um, troubleshooting and more over, over the course of time. But the, the model shows there's a clear way to get from old Legacy PMO present mode operation fabrics to over five nines with a new FMO that really mimics what's available. Um, with Sr.
Linux and ida, this third leg of the stool may be my favorite. Um, as part of this project, we have had the opportunity to talk to the enterprise IT people at Nokia. Um, do you have any idea how many employees?
Nokia has over 80,000 worldwide. They make stuff, they have manufacturing operations, they have extranet connections for partners. Um, they have all sorts of different types of office buildings all over the world, right?
It's a complex enterprise IT environment, right? So it's, it's a great, um, terrarium of network migration and implementation. I could, I could use a better word than that, but, you know, it's a hot house environment, right?
Where there's real pressure, um, outages that impact manufacturing could cause anywhere from 500 K to a million per incident, right? And that's just one example of the cost associated with downtime in an enterprise like that. Um, we are in the process of putting out more material on, um, the journey of this team, and it's still ongoing.
Like on Monday, I got an update from the team that talked about their last phase of their most recent phase of here's what we just cut over. And like, they're, they're stoked, they're excited about it. They just want their network to be better.
Um, and it's really cool to see them engaged on this. The last thing I'll throw on the table for you is, you know, people like Mitch and myself making our job to just be out there, right? Talking to people, hearing, hearing what they hear, you know, many of you are in the same situation, right?
Um, I am never the smartest guy in the room, but I get to talk to the smartest people in the room. Um, and I love that through, you know, what I've been able to do in network automation forum, what I'm doing in my consulting work, what I'm doing through the podcast, um, showing up at, you know, OG nano, lots of other operator centric events. Um, we definitely see across the board this increasing appetite, um, for automation and AI ops, right?
It's real, it's palpable. Um, everyone's starting to ask the questions of how, where do I start, right? And we've found a way to help some people along in that from a, from a vendor independent perspective.
So what I would recommend to you as next steps, um, you can engage by taking a look at the reliability study. Mitch has already distributed that. Um, you can also get a hold of the Bell Labs model report and you can actually work with Nokia folks to tweak the model for your environment.
If that's something of interest to you, you can kind of see how it works. They won't, they won't hand you a package and leave you, you know, good luck. Play with the model yourself.
Um, you'll get guidance from them. And I think that's super appropriate. Um, watch as we put out more material in the next couple months on the enterprise IT migration at Nokia.
Um, and you, uh, you should reach out to the Nokia folks who are here today, you know, hear what they're saying, test their claims. I think Andy might have some fun things, uh, to propose to you later that will help with that. Think about it in your business context, in your environment, um, and figure out how to go learn more.
And of course, Mitch and I are always happy to talk about what, um, FUO and solution are doing together and in our different pieces of the world. And that Just wanna add seconds over. It's been a great pleasure working with you on this.
Really enjoyed it. Learned a lot from you. So Right back at you.
Yeah. And, uh, I want to give a hat tip to Kathleen Barron. Kathleen, uh, Kathleen is, um, the master reducer of entropy.
She's the project manager and like, uh, I don't know how she does, it's the Locomotive on the train. Okay, there you go.