Amazon Web Services Recovers After Major Global Outage | Tech Field Day Rundown: October 22, 2025
AWS is restoring operations after a massive outage disrupted internet access worldwide, affecting major platforms like Snapchat, Facebook, Fortnite, Delta, Coinbase, and several banks. The issue stemmed from a DNS failure that temporarily prevented access to data stored in AWS systems, causing widespread service interruptions and “Error 404” messages. Experts say the outage, which exposed how heavily global internet infrastructure depends on AWS, may have cost hundreds of billions of dollars. Amazon says it has “fully mitigated” the issue and continues investigating the root cause. This and more on the Tech Field Day News Rundown with Tom Hollingsworth and guest host Kate Scarcella.
Time Stamps:
0:00 – Cold Open
0:27 – Welcome to the Tech Field Day News Rundown
1:21 – MLCommons Unveils New Standard for AI Security
3:47 – Nation-State Hackers Breach F5, Endangering Thousands of Networks
6:49 – Hackers Breach U.S. Nuclear Weapons Plant via SharePoint Flaws
11:20 – Broadcom Unveils Thor Ultra: 800G Open Ethernet NIC for AI
15:12 – Veeam Acquires Securiti AI for $1.7 Billion
17:56 – ShengShu Launches Vidu Q2 to Compete with OpenAI and Google
21:48 – Amazon Web Services Recovers After Major Global Outage
29:48 – The Weeks Ahead: Upcoming Tech Field Day Events
31:52 – Thanks for Watching the Tech Field Day News Rundown
Transcript
AI security in this market. F five Blown away by Hackers, SharePoint went nuclear. Rocom has the power of four Ultra Veeam backs up for security.
Ai Shang Shu Make Voodoo video Voo. And we pay the AWS bill in this closer look on the Tech Field Day rundown. Hello everyone, it's Tom Hollingsworth back once again with the weekly Tech Field Day rundown.
It is October 22nd. We are that much closer to Halloween, and I'm very happy to have you all joining me here along with a new co-host joining me today, Kate Scar. Kate, welcome to the Rundown.
Hello. And thank you for having me on. So I, I'm very thrilled to have you here.
Uh, you were recently a delegate at Security Field Day and, uh, look, look, luckily for us, there were a lot of security related stories this week. I think. Uh, you're gonna have a lot to say, aren't you?
I, and I'm looking forward to it. I do. I have so much to say.
I absolutely, yes. Awesome. Well make sure you sit down and enjoy something good for lunch.
'cause it is National Tavern Style Pizza Day, probably brought to you by people who make tavern style pizza. Um, but, you know, whatever the case, there's a lot that we're gonna be talking about in, as I say, tipped a little bit. It is gonna be security focused.
So, uh, put on your 10 foil hats and make sure you change your passwords. Folks. 5 jailbreak Benchmark.
The first framework to measure how AI systems resists safety bypass attempts. It's new resilience gap metric reveals how much safety performance drops under attack, showing vulnerabilities across major AI models. 0 release.
That should be happening sometime next year. Kate, is it invaluable for people to have a benchmark that shows them just how resistant AI models are to attempts to get under the covers and see things and know things they're not supposed to? I do think it's important, uh, basically tracking how AI safety drops as you know, once it's attacked, how stable, um, or fragile its guardrails are, really are and look at and know how it's supposed to be.
Um, you know, I hate the word resilience and we use it a lot. There's a lot of reason I feel like it's like one of these buzzwords that, um, that has come into our vernacular for the last, I don't know, five plus years. And for me, I like to look at more the word of behavior.
Um, and that's where I think that's a buzzword that I think we need to look at, especially when it comes to AI and, and those models. You know, because resilience to me just basically means that, you know, we're able to, to stand up and, and to keep going, but just because we can stand up and keep going, is that really what we need to do? Is that really resilient?
Um, I don't, I don't necessarily think so. I I think we have to look more of adaptive, um, integrity, like systems that just don't survive stress, but they learn from it. That that's gonna become key, especially as we go into this new AI generation to keep, um, unsafe, uh, choices from being attacked.
Um, we go away from resilient. We start to look at predictable, predictable behavior. And that's what I think becomes these real benchmarks.
Um, not how it resists, but how it can revalidate trust through its actions. F five networks has revealed a major breach by a nation state hacking group that gained long term access to its systems, used to build updates for big ip, a product used by most Fortune 500 companies and US government agencies. The attackers stole source code unpatched vulnerability data and customer configurations, raising risk of supply chain attacks and credential thefts.
While Investig found no signs of tampered updates, US and UK cybersecurity agencies warned of an imminent threat and ordered organizations to patch systems and follow F five security guidance immediately. Boy, that's a terrible thing that someone managed to hack into F five, which is a company that does a lot of pass through data and was able to get ahold of some code and some unpatched vulnerability code for zero day exploits. I hope there's no big customer out there that uses a lot of F five products that would be a prime target for a nation state.
Boy, I really hope that the United States federal government isn't a big customer of F five. Don't you? I mean, how, how crazy would that be?
Uh, you're not running out to look at the F five customer list, are you? I hope you're not, 'cause you're not gonna like what you find. This is one of the biggest worries that we have in modern hacking, right?
Like we, we know that there are people out there that are doing this for a business, whether they're ransomware crews or organized, uh, groups running out of well unfriendly countries. We'll put it that way. Uh, but this is the next step in that this is nation states that are using this as a form of intelligence gathering.
So in the case of whomever did this, they managed to jump in, gain persistence and get customer information big whoop. Uh, they were able to get code, they were able to find out what F five was working on, uh, possibly some unreleased security vulnerabilities, and they're gonna go try to exploit them. The problem is they're not doing this for notoriety and they're not doing this to get paid.
They're doing this to gain a foothold and persistence and hang around. And that is problematic for the kind of customers that five works with because as I mentioned, a lot of governments use F five. Remember when we talked about Saul Typhoon last year and the fact that people were able to gain a foothold in persistence in switches and do things like monitor political candidates phone calls?
Remember SolarWinds, this has all the fingerprints of the SolarWinds hack because what they did is they were able to get into the supply chain, introduce compromise tooling into the system and gain persistence. I'm not saying that that's exactly what's happening here. I am saying that if you are a customer of F five and you haven't already patched your devices, I would now, and if you have patched your devices, I'd be on the phone with F five asking how they're going to prevent this and what kind of indemnity you have against people being able to get into the system after the fact.
Because this is gonna be one of those things where I bet you we're gonna be seeing fallout from this for months to come, which of course, you know, we're gonna be talking about here on the rundown. Speaking of foreign hackers, uh, they were able to exploit unpatched Microsoft SharePoint vulnerabilities, and they breached the National Nuclear Security Administration, Kansas City National Security Campus. You may recall that if you read Wikipedia as much as I do, because it's a key site for us nuclear weapons components, the attack happened back in July.
It has been attributed to Chinese or possibly Russian actors. It of course, exposed weaknesses in federal cybersecurity and the gap between IT and operational technology protections, no classified data has been confirmed to be stolen. However, experts warn that even minor technical leaks could reveal sensitive details about US weapons manufacturing.
And of course, this highlights the need for much stronger, more unified defenses in the federal government. Kate, do you think that the people who were doing this were looking for something they could leverage? Or do you think they were just looking to see what they could scoop up wherever they could get in With that?
I'm gonna say yes. I, I think actually both. Um, and should we just highlight that?
I think NSA actually just laid off some people nice riff. Um, so let's just combine that with, uh, a riff and what voila. So what we're seeing here isn't just another exploit of collaboration software.
It is governance failure masquerading as a technical breach when the organization manufactures manufacturing. 80% of nonnuclear parts for our deterrent is accessed through a SharePoint flaw. We're past perimeter failures.
I we're into identity process, trust, value, validation, failures, and oh my goodness, SharePoint, I, you know, as a person who has so much talked about critical infrastructure and IT and OT and, and how we definitely need to, to harden this critical infrastructure, not only with nuclear, but throughout our critical infrastructure, um, systems. This just highlights something that is, is coming. It's, it's here, but it's also coming.
So the attacker didn't blast through any hardened weapons control system. They stepped in through what many considered and what we consider mundane business tools. It this is a wake up call, people defense architects, threat surface, you know, business collaboration.
It just, you know, CADA and ICS treat it accordingly. I mean, this, this is a serious, serious breach. We keep talking about architecture isn't verified, but by nothing bad hasn't happened yet.
It's verified by how did we behave when bad happened? And for us, you know, for us to still be throwing out, we don't know if it's Chinese, we don't know if it's Russia. You know, at the end of the day, it's, it seems that we don't have a lot of visibility into what happened.
Uh, if the system defaulted to alert, isolate, we validate. That's trust architecture. And you know, at this point, you know, we really need to be thinking about how the footprint may still be in the system here.
So when we're looking at, um, systems, what happened with Kansas City, the physical consequence might be weeks or months away. We have no idea of knowing this. Um, but the risk of reconnaissance fingerprinting, future hold points become immediate.
You know, to, to your question and beginning here, this slow burn threat model, I mean this, you have asked me before what keeps me up, iot, IIOT, this is what keeps me up because we treat it, we treat OT like it, and I don't know what it's gonna take for us to learn that this is, that this is serious. And I, I, and I, I hope it doesn't get to that point. I hope at one point that we actually say enough is enough.
So we shall see Broadcom unveils Thor Ultra 800, um, gig ethernet Nick built for large AI clusters, emphasizing open standards and high efficiency, fully compliant with ultra ethernet consortium specs. It boasts, um, it boosts task performance by up to 15% while using just 50 watts. Thanks to smarter congestion control and selective retransmission with flexible deployment options and strong security Thor ultra positions ethernet as a leading interconnect for hype hyperscale AI workloads.
Tom, what do we think about this? I'm excited that we finally have an 800 gigabit nick on the market that isn't an Nvidia nick, because one of the things that we've seen is that those Nvidia nicks are effectively running NV link, which, okay, I get that. Like, that's what I would expect from the, the giant who's building all of this vertically integrated stack.
This is an 800 gigabit ultra ethernet nick from Broadcom, a company that has been introducing 800 gigabit ports on their switches. So now if you wanna run it in flat out mode, go for it. And that's what we're expecting.
Now, as I've said on a number of occasions, remember ultra ethernet is not ethernet. I know it's weird. The ethernet's in the name, it does run over traditional ethernet type signaling.
However, it is not ethernet like the way you know it, you have to plug this nick into a switch that is alteration that capable and it does all this weird stuff. You know, as mentioned in the read-in, uh, you know, task performance re optimization and it uses different kinds of congestion notification. Um, this really is more like a fabric, but that's what you want, right?
I think the big thing is the fact that they managed to do this in 50 watts of power. It is' an SFP. Now, don't, don't get me wrong here, this is a nick that has one port on it.
It's the fastest port you're gonna need right now because I, I believe 800 gigabit ethernet is, is the, the fastest you can get in a, in a network facing port. Um, but I mean 50 watts of power for an 800 gigabit port. Like I can remember hearing Andy Bettelheim talking about how we're probably gonna cap out pretty soon because the amount of heat that's being generated by these things is enormous.
By the way, if you wanna see a picture of the card head over to serve the home, because Patrick actually has a picture of the heat sink on this thing, I'm pretty sure that I could mount a laser on it and probably cool that thing off pretty quickly. Uh, this, this to me is kind of the next step, though. We're seeing Broadcom basically aligning against Nvidia saying, you know, we're gonna match you step for step.
And remember that the current thinking is that the fastest InfiniBand chips that they're using are gonna match up very closely to where we are at with 800 gigabit ethernet because of the signaling and things like that. But with the way that Nvidia has been using Spectrum X technology and all the things that they've been announcing recently, we know that InfiniBand kind of has a horizon, right? Like we, we know that that's eventually gonna have to go away, and the real development is gonna be happening inside of spectrum X ethernet.
So eventually we're gonna have spectrum X versus ultra ethernet, and that's gonna be a real interesting showdown. 'cause on the one side, you've got Nvidia who's really pushing their solution and on the other, but they're partnering with companies like Cisco and Meta and Oracle and, and companies like that. But then on the other side, you have Cisco and Broadcom and other companies who are heavily involved in the ultra ethernet consortium trying to build a standard that not only competes, but can provide similar performance to infinite band.
So good times are headed our way. And if you wanna buy me one of those next tests out, please let me know. Um, I I have a wish list.
And I, I think though, if you have to ask how much it costs, you're not in the market for it. 725 billion in cash and stock. Uh, they're aiming to combine veeam's data, backup and recovery expertise with security AI's, data security, posture management, and governance tools.
The merger will create an integrated platform that unifies data protection, privacy, and AI across multiple multi-cloud and hybrid environments. Security ai, CEO, rehan, Jaleel will join Veeam as president of security and AI after the deal closes in Q4 of 2025. Hey, wait, that's this quarter.
Veeam says that the acquisition will help customers better secure govern and recover data while enabling transparent AI driven innovation. Kate, is data security posture management something that Veeam needed to add to their portfolio today? I think so.
I think it actually makes sense. Um, it marks a real shift of moving from just data protection to data trust architecture. Backup alone isn't enough anymore.
We need to know what data is, uh, how to govern it, how to recover it in ways that align with both risk and AI driven decision making. Uh, when VM talks about enabling customers to understand secure recover and roll back data, that's essential for the trust pipeline I've been describing, discover, validate, protect, restore for a, for a long time now, and integrating this into that stack is smart because too many backup systems are still blind. They can restore everything, but not necessarily the right things.
Uh, two, trust architecture isn't just this checkbox, it's this lifecycle. The maturity leap here is end to end from backup media through governance to responsible, a I use, but behavior still matters. Will this new system surface risk, um, show lineage, uh, validate integrity during recovery?
Because without that, it's just a bigger stack. What we're seeing is a convergence, governance, backup, and AI readiness coming together. So yes, it's, I think it's a great idea.
And the question every enterprise, um, is the same data stack reacting, architecting for trust before the next breach or AI misuse hits? You know, one of the things that, um, that we always have the problem with being, being data hoarders that we are, we just have this idea of, let's, let's just protect all this data and that's not the best solution. Um, so I think by being able to see, see what you're backing up, I I believe that this will absolutely help.
1, voodoo Q2 improves facial expressions body movement and see seeing consistently while generating content faster and cheaper. Sheng SHU also released a global API, so businesses can integrate the technology into their workflows. Wow, if this doesn't resonate, some security concerns, Tom, what do you think?
Ah, it won't be a problem, right? I can just upload whatever images I want. I can make all these really cool videos of fighter planes flying through the sky and, uh, Dan babies dancing with dogs and all kinds of other stuff.
And yeah, I'd be a little bit worried if I was Google and open AI right now, because remember when Deep Seek came out and everybody rushed to do it because it was as good as the competition, it was a lot faster. And yeah, privacy concerns, security concerns, we don't care. It's faster and better, and it's not open AI and worst deep seek today.
I, I don't, I don't really see it. Now granted, we know that there was a move to, to ban it and other things like that. I think what we're seeing here is the way that this is going to work for a while, you know, we've seen a lot of reports from Open AI saying that they've poured millions upon millions upon billions of dollars into the research, and they created SOA and it does all this crazy stuff.
And then about a month later, somebody comes along and goes, well, we, we built it too. It's a little bit faster and, and it it does things a little bit better. Are they building on some of the intellectual property that, that OpenAI has put together?
No, they probably just opened a asked open AI's check, GPT, Hey, build me one of these platforms and vibe, code the whole thing, and, and it'll work just fine. This, this is the battle. Even if Xhu, uh, gets their app banned from the app store, which I highly, I highly anticipate is going to be a conversation we're gonna have pretty soon.
The, the, if you wanna say the damage is done, sounds cliche, but it is because what you're basically saying in the Chinese market is, we don't have a need for Sora, we don't have a need for Veo, we don't have a need for anything else because we've got our own at home and we're gonna use it and we're gonna encourage you to use it. And if it gets banned from the app store, so what that means that we'll just have more time to do what we wanna do. Like making sure it can't generate any videos that concern a certain fluffy bear from the a hundred acre woods.
That is a banned term in, um, you know, in, in certain countries of the world for various reasons. But when you think about it like, this is the way that life is going to be for a while, we are going to be generating all of these crazy new functionalities in AI that quite honestly, nobody is asking for. And then about six to eight weeks later, we're going to see the next generation being rapidly developed, because as it turns out, one of the things that AI is really bad at is thinking up interesting new novel ideas.
One of the things that AI is really good at is copying interesting novel new ideas quickly so that, that that gap of first mover advantage is gonna be shrinking before you know it. And I, I think that, you know, the folks over at Open AI are, are gonna have some soul searching to do, maybe they should ask chat GPT to make them a video about what that looks like, or they can just use shsu. It might be cheaper in the long run.
Alright, we had a story that we wanted to take a closer look at, and there's no denying who it's gonna be this week, folks. Uh, if you tried to order your Starbucks yesterday morning, you know what happened? AWS is restoring operations after a massive worldwide outage that disrupted internet access and disrupted major platforms, including Snapchat, Facebook, Fortnite, Delta, Coinbase, a few banks, and possibly the reservation booking system at Costco.
Who knew, uh, the issue of course, wait for it. DNS. Yeah, it was a DNS failure that temporarily prevented access to data stored in AWS systems that caused widespread service interruptions and everybody's favorite numerical error message 4 0 4 baby, uh, experts say that the outage, which exposed how heavily global internet infrastructure now depends on AWS, may have cost upwards of hundreds of billions of dollars.
Amazon says that it has fully mitigated the issue and it is continuing to investigate the root cause. Now, I know for a fact that even though they had fully patched the vulnerability and, and, uh, fixed the problem as of like 6:00 AM central time, yesterday morning, we were seeing rolling concerns going on all day long. And of course, as soon as it went down, everything that happened that went wrong yesterday was blamed on that just like it was blamed on CrowdStrike or the last time that something happened.
So, Kate, I, I kind of wanna dive into this a little bit other than don't put all your stuff in US East one, like how are we going to be able to reduce our reliance on Amazon, because it really does feel like most everything runs on it now. So could you ask me like a simpler question? I mean, come on, what the heck really?
Um, so what I'll say though is that I believe that, that we need to go, you know, more to a distributed model. You know, I, for a long time we're putting all of our eggs in one basket. 0 technology, it really brings in, um, the case for distributed.
0 technology, distributed technology become more important as we become more connected. And as we see these cloud, you know, these single system failures, we have to, we really have to think about, um, how we, how we've we're using an archaic type of, uh, system. 0.
Uh, when we look at network and, and, hey, let me just go back to, to Thor. I mean, I, I, I do think that we're gonna be able to get there to this distributed more a distributed technology. Um, instead of putting all of our eggs in one big AWS basket, I, I love your optimism, Kate, and I would totally agree with you, except it's really hard to stand up a cloud computing instance.
And did you know that Amazon actually has a service now that will just take care of all that for you? All you gotta do is just, uh, sign a little more on the bottom line here, and we're only gonna charge you a few pennies per hour and everything will work out just fine. I love the fact that this exposed for a lot of people, the law of unintended consequences.
For example, did you know that there was a company that made a, uh, bed that was completely cloud enabled and allowed you to set all kinds of fun positions, you know, like sitting up to read in bed and it would cool itself and all these other things? Do you know what happened at 6:00 AM on Monday morning when that bed suddenly couldn't contact Amazon? It flew forward and people who were in it were actually getting thrown out of bed because it turns out that's its default behavior.
I, it reminds me of a few years ago, you know, one of the last times that AWS had a huge outage, uh, when someone cratered a router in US East one. Uh, and it went down for a couple of hours. And one of the things that had people very, very annoyed was the fact that at the time Amazon was hosting the stat page for US East one on US East one, which meant that when it went out, the lights were stuck on green because nobody could get in to update them.
And so everyone's like, well, the, the status page says that it's up. I don't understand why everything is acting so haywire. There were was there were isolated incidences yesterday of people realizing that while they think that they're not running on a WSA service that they rely on might be, so, like for example, I could post new things on Reddit, but I could not vote the comments or, you know, this thing was working, but this little other fractional piece over here wasn't.
And I think that that people need to get more control over their environment that way and, and woe be to the people who said, oh, well the cloud never goes down, so I don't have to worry about that, that stuff. There are things called availability zones. There are other regions of Amazon that are not based in Northern Virginia, and I think people really have to understand that unless you have a very good understanding of the way that your system is supposed to work or you are a site reliability engineer, you are effectively doing the same thing that you've always done in the past.
You're just doing it in somebody else's data center, right? Like we talk all the time about Netflix being an example of a very survivable system. No one outage can take them out unless it's the the Tyson Paul fight where they just had too many people trying to watch it all at once.
But part of the reason for that is because they were forcing themselves to build a survivable system over the years by purposefully breaking things along the way in small ways to see exactly what happened. That's why, for example, not all their infrastructure runs in one availability zone. They spread it across the nation, across the world.
And that has huge impacts for people all around the world because someone out there is saying, oh, well this will just be so much easier if we migrate it to the cloud. And the other thing that I need to make sure that people understand very, very much, if your name is Microsoft, Oracle, Google, IBM, Alibaba or any other cloud provider, you keep your mouth shut because this could have very easily happened to you and Amazon would be the one laughing all the way to the bank today while everybody moves all of their data off of AWS onto somebody else's cloud. Because you know what?
Yeah, I just don't know that how this is, it's the same thing no matter where you're running it. If you're running everything in one group of, uh, one server cluster and you haven't built it to be resilient, this will happen again because we're sitting there saying this, this is not the first time that this has happened to Amazon. I actually had a mockup of a t-shirt on Cafe Press one time and it was one of those really simple black t-shirts with white writing that says, don't install on us East one like that.
I know it's the default. Amazon, if you're listening, uh, Bezos doesn't pay attention to me anymore. Andy Jassy.
Andy, you're a friend of mine and I say friend of mine, meaning I'm assuming that you know that we exist on this podcast. Do me a favor, create a round robin algorithm that forces people to pick a different region of AWS. It will take about 30 seconds to code 45.
If you ask chat Chi pd, divide code it for you no matter what. This will fix most of your problems. Get people out of Northern Virginia, trust me on this.
Alright, enough about that, but not enough about cloud because guess what? Cloud Field Day is happening right now. com, you can tune in to see all of the great things that are happening at Cloud Field Day.
Alistair Cook braved the seasons he flew out of spring into fall and he will be joining us at Cloud Field Day. He's got a great lineup of delegates and a wonderful group of presenters and you're not gonna wanna miss them. com, then come back next week because we move from cloud to ai, AI field.
Day seven is going to be happening on October 29th and 30th. Stephen is gonna be dressing up as the scariest thing that you can imagine and AI generated Stephen Foskett. I know the horror, but thankfully he's gonna have a lot of delegates there that are going to be able to, um, decode his strange structural syntax and, uh, possibly give him a better writing guide.
And we are gonna have a great group of presenters there as well. Then the next week, which of course will be the first week of November, November 5th and sixth, guess who's back in Silicon Valley? That's right.
It's your boy Tom. I'm gonna be out there for networking Field day and we have a wonderful lineup of presenters of people that are building the infrastructure that run AI in the cloud. You know, the plumbers that we always keep forgetting about.
com of the presenting companies and of the delegates that are gonna be there. Stay tuned to that page because we might be adding some stuff very, very soon. Kate, if people wanna check out some of the stuff that you add very soon, where can they go to read some of your writing or see some of your musings?
So the CD Foundation, I put out, uh, a couple articles. Um, so I'm the chair of the C-D-F-C-I-C-D cybersecurity sig. Also, if you look on LinkedIn, uh, you'll see some articles that came out of Security Field Day from since I was out there as well.
And so I'll miss you guys this time. But, uh, definitely look, look on either LinkedIn or uh, the CD Foundation. Awesome.
And we want to thank each and every one of you for watching the Tech Field Day rundown. Don't forget we post new episodes every Wednesday on YouTube or in your favorite podcast application of choice. If you wanna use Apple Podcasts, that's great.
There's a few out there that you can check out as well. The rundown is also being streamed on Techstrong tv and if you've got one of those cool set top boxes like an Apple TV or a Roku, make sure you check out the Textron TV app there as well. You can also catch myself and many of the other folks here on Other Techstrong Future Group, uh, programs such as the New Security Boulevard podcast.
Uh, Kate was a guest of our on the second episode. You can also see me, uh, somewhat frequently on, uh, you know, other things. So make sure that you tune in there.
Uh, we will be back next Wednesday to talk about all the IT news in the week. That was, but until then, for myself, Tom Hollingsworth, for Kate Scar, and for everybody else here, ad Tech Fuel Day, and the Futurum Group, we hope that you have a great week. And remember, don't install things on US East.
One.