Veeam Understand – Know Your Data
In this session, Field CTOs Michael Cade and Emily Tellez dive into the practical application of Veeam’s four-pillar strategy, focusing heavily on the Understand phase. Central to this approach is the recent acquisition of a Data Security Posture Management (DSPM) solution, now integrated as the Data Command Center. This tool acts as a “social network of data,” utilizing a connector framework of over 350 integrations to inventory data systems across platforms like Microsoft 365, Kubernetes, and various cloud environments. By building a comprehensive map of data lineage and access, Veeam helps organizations identify sensitive information, uncover “God mode” privileges, and conduct ROT analysis to eliminate redundant, obsolete, and trivial data, thereby reducing the attack surface and storage costs.
Beyond visibility, the presentation highlights how this intelligence informs smarter backup and recovery workflows. The speakers emphasize that understanding data is the prerequisite for securing it, particularly in the face of agentic AI risks where data might be overshared or mismanaged by automated models. Veeam’s orchestration capabilities, which have evolved since 2018, allow for dynamic documentation and automated readiness checks to ensure compliance with Recovery Time Objectives (RTO) and Recovery Point Objectives (RPO). This ensures that disaster recovery plans are not just static documents but living, tested processes that can transition workloads, such as moving VMware backups to Hyper-V or Azure, at scale while maintaining a clear audit trail for cyber insurance and regulatory requirements like GDPR.
The discussion concludes with a focus on clean recovery, addressing the critical need to prevent the re-infection of environments during restoration. Veeam integrates multiple layers of defense, including inline scanning for anomalies, indicator of compromise (IOC) detection, and the use of YARA rules or antivirus signatures. This process can occur at rest, during backup, or before restoration into isolated sandbox environments for forensic testing. By partnering with an ecosystem of over 60 security providers, such as CrowdStrike, Veeam ensures that if a threat is detected in production, the backup system is immediately informed. This holistic approach transforms backup from a black box into a proactive security asset that validates data integrity and operational resilience in a post-AI world.
Presented by Michael Cade, Field CTO, and Emilee Tellez, Field CTO. Recorded live at Tech Field Day Extra at RSAC 2026 in San Francisco on March 23, 2026. Watch the entire presentation at https://techfieldday.com/appearance/veeam-presents-at-tech-field-day-extra-at-rsac-2026/ or visit https://techfieldday.com/event/rsac2026/ or https://Veeam.com for more information.
Transcript
Hey everyone, I'm Michael Cade from Veeam. I'm a field CTO, and I'm joined with Emily. Hello, everybody.
My name's Emily Tayas. I'm a field CTO, also with Veeam. So in the previous segment, Rick touched on the four pillars and a bit more of the business outcomes of where we are, what we're doing at the moment from a Veeam point of view.
So now want to go into those four pillars, right, as to why, and probably answer some of Tom's question there as well, around why are we getting into this space around helping our customers understand that data. So just to reiterate these, I put some text along these to make my simple British mind understand it. But really about that understanding is helping our customers and to Tom's point, so Tom's question for the previous section was how are our customers relating to this?
How are they using this? And the good thing about this, and I was going to come off mute and say, but didn't want to disrupt, but ultimately, these are all modular. A customer doesn't need to take all of these pillars.
They might be just completely focused on the resilient and the secure side. They might be way down the line of unleashing that data. They've got a good hygienic understanding of their data, and they're going for it.
But the story is how these interlink with each other as well, which is where that acquisition comes in from a security perspective, is being able to help our customers understand their data. Who has access to it, where all the critical or sensitive data is within their estate, ultimately then, so they can secure it better. What does that flow look like?
So you've got a mind map, a social network of where that data is being shared, OneDrive links being shared publicly, SharePoint links being shared publicly, ROT-based data, et cetera, redundant, obsolete, and trivial data. And then that leads into how can we be smarter when it comes to protecting that as well, right? Just as a very big picture.
Equally, I mentioned ROT, and we're going to touch on ROT a little bit later, but we're seeing that massively resonate with our customers. We're only 100 odd days into this acquisition, and we're already seeing a synergy of conversation around reducing that. Not only is it reducing cost on enterprise storage in the data center, where they've got data that's been stored for X amount of years that could be sensitive, or even if it's not sensitive, it's still taking up expensive storage space.
Let's tier that off into cheaper, deeper storage somewhere else. So it's reducing the cost, but also it's mitigating risk. Reduce that attack surface so you've got a good lean understanding of that data so you can use it, or at least just protect it a little bit better.
So we're going to go through with the sections that we're doing. We're going to go through each of these, but this one is really focused on the understand. So knowing where all your critical data is and who and what uses it to protect it smarter.
So I want to spend a bit more time on this, is that, so you might have seen on Rick's original slide, it talked about the Data Command Graph. So it's a graph database, the social network of data. This is security as an acquisition or as a product, the Data Command Center, has the ability to inventory all of your data systems.
So where Veeam focuses on protecting the platform, like whether it's VMware, whether it's another hypervisor, whether it's M365, whether it's the cloud, whether it's Kubernetes, the platform that the data system lives on. So if it's a Postgres database or a Postgres cluster, security has a connector framework that enables them to go into that and build a map around who has access to it, what user, what kind of data is in that structured or unstructured data, and starts to build this map out, and you see just a snapshot of the connector framework down below. But there is around 350 plus connectors that plug into these databases, NAS-based services, data warehouses, data lake houses, et cetera, like big data.
And it basically builds this map of everything that's going on in that space. Now, the outcomes of that could be, I need to know what's the sensitive data that we have in our estate, ROT analysis. You'll hear me say ROT quite a bit, redundant, obsolete, and trivial data, because I think that resonates with the backup admin, the IT staff that we generally speak to.
That's where we're seeing it fit from a story perspective. Equally, because now we've got an understanding of what that flow of data looks like, what the lineage of that data looks like, how it's moving from A to B to C. Well, when bad things happen, like an exfiltration of data, for example, in a cyber event, if you've got a good timeframe and know when something's happened, we can work back and understand where that data has been and what data has been exfiltrated.
We're not a data loss prevention tool, because we don't have that capability. We partner with a load of DLP type systems, but this is going to give you a mind map of where has that data been? What was the data that was exfiltrated?
Because many of the customers that we've spoken to, not just customers, but a broader survey, is they don't know what data's been exfiltrated, and the cybercriminal, they're not writing down what they've taken, they're holding you for ransom for a reason. What this will do is give you a lens into what that dataset was, so you can make an informed decision on, well, was it actually sensitive, or was it just some old data that is irrelevant? The other thing is around compliance, especially from me.
So being able to build those frameworks into understanding that data. So not just understanding what your Social Security numbers and where they are, and credit card details, but also aligning that to your regulatory requirements within your business. So out of the box, it can marry those up.
You tick and choose what ones you want to marry that up to. Who has access to it? I guarantee when you implement Security AI or the Data Command Center, and you add all your data systems, you're going to find some God mode privileges out there still in 2026.
You're going to find something that you didn't know about your environment. So this gives you the ability to revoke some of that access, again, reducing that attack surface and the vulnerabilities of that. And then AI governance.
So I know we're here at RSA, we know that everything is going to be about AI, and AI resilience, as Rick touched on as well, is the next wave of resilience. None of the other things have gone away, like fire, flood, blood, and accidental deletion is still going to be high up there on people's minds, as well as cyber resilience. AI resilience is now we're introducing new vectors into touching our data, using our data.
So how do we get a good grasp on that data system? So we'll dynamically also discover those data systems, those models that are being used, AWS Bedrock, Microsoft Copilot, et cetera. What's being put in there?
What's being shared? Without getting too into the weeds, a concept of an LLM firewall. Let's make sure that we're not sharing sensitive data up into a public model so that you've got control of that.
There's three slides on this, but I'm going to just touch on one. So PII, just one variant of sensitive information. And it also depends on where you are in the world as well as to what regulation, what compliance, what governance you have to follow.
But generally speaking, this framework is out of the box included inside of Security AI and the Data Command Center to be able to say, that's PII. Equally, that could be PHI, it could be PCI, lots of different acronyms. But ultimately, it's got a good grasp on what that data is and gives a flag of what that data is so that you can see that in this, I keep calling it the social network of data.
But it gives you an idea of where that is, and I think... Yeah. So just to paint that into the data element.
So you've got all of these regulations. You as a company might also have a regulation that you've stipulated. It might be a bit of another regulation or a compliance rule that you want to bring into your company, as well as all the industry regulations.
But we have this concept of data elements. How do we pull those together and build this content profile? What does that look like from a business perspective?
So you can build a set of rules around that, because that might be the most sensitive data. That might be the gold data in your business that you need to protect, or you just need to have a good understanding of. So this could be security before we acquired them.
We're obviously a standalone company, DSPM type tool that started off in the privacy, governance, and compliance space. So I've been calling it DSPM Plus. And that's a standalone tool that you can go and procure on your own.
Let's say you've got another backup system that is already backing up your data, you're happy with that, but you want to have a good understanding, you want to be more hygienic when it comes to data. This can come in and help you derive what that data set looks like from all of the different data systems that you have. So just one other thing we'll add on this, right?
The biggest portion of why did we do this acquisition. Yeah. Right?
So I think for a lot of our customers, it's number one, just making sure that they actually understand the data that actually exists out there, right? And that's probably one of the biggest things that being at Veeam for 12 years, I don't think I've ever went into a customer account or went into a POC with them, and they didn't share-- There wasn't something that we found that they didn't know about. Meaning we identified or we uncovered something, whether it was, hey, you have virtual machines that have orphan snapshots, or hey, you have virtual machines that should've been protected, or data that should've been protected that wasn't.
And then going back to the earlier question as to like, well, how are your customers reacting to this news? For me, especially someone that speaks to customers and speaks to partners on a day-to-day basis, the news is actually positive, right? If we think about how backup has existed prior to most recent years, it has been a black box.
It has been something that we're just told that we have to do, and hopefully when something goes wrong, we have something that we can go and we can recover, too. But now with cyber incidents, now with AI, it's shined a light on a lot of the ways that we go about handling risk, and what is our current risk posture, right? And so you see a lot from cyber insurance company getting involved and validating with a customer, well, do you have backups?
Do you have a workflow? Do you know what your SLA is? Do you know what your regulation looks like in the time in which you have to go and actually report a breach?
So all of those items kind of help us to feel that understanding of, well, how can we help you, Mr. and Mrs. Customer, be ready for an incident that can occur and feed that information early versus you having to do it post-process and having to juggle a lot of different hats.
So a lot of the understanding that Michael kind of led up to, that's the biggest portion, right? How can we help you understand your data better so that way we can make better decisions in terms of overall protection? So we've already protected data for a long time, and we've already added capabilities like being able to orchestrate the data set itself.
I don't know why that keeps on moving forward. Tom, do you need some help switching? No, we're good.
No, it looks good. So essentially, when we think about how we helped organizations for a long period of time is let's think about how can we actually restore your data from backups or from replicas. How can we help you to build those bigger workflows and essentially make sure that we are hitting your compliance minimum, so whatever your recovery time objective is, whatever your recovery point objective is, and that we have that fully aligned to the SLA standards.
How do we make sure that this is documented? That's probably one of the biggest issues that I see with a lot of organizations. They have a separate tool.
Maybe they're using Word, maybe they're just using something to capture, like Notepad, their entire workflow strategy for how you go about restoring your most critical data. But then how often does that get updated? Where is that being stored?
Who has access to it? What happens if the one person that is in charge of orchestrating those workflows for that particular incident is out, is not available, they can't be reached? Do you have somebody that's trained that could go through and can walk through step by step and know exactly that it is going to fail over to the right process that you have set up?
So that dynamic documentation and compliance is really key, and this is something that Veeam we've been doing since 2018 and helping customers understand that overall orchestrated recovery. And then we get to the other two, clean recovery. I think this is probably one of the most overused terms that we've seen in the last three years.
What does it mean to have clean data? Well, technically from a security standpoint, your version of clean is very different than maybe an IT ops person's version of clean. We have data, we can restore it.
That's their version of clean. From a security perspective, it's, well, no. Has that data been messed with?
Is there any type of malicious or anomalous activity that has happened inside of it? Is there backdoors that maybe a threat actor had put and now we are going and restoring into a separate environment, we are now reinfecting that new environment with this data that we are pulling from backups because it was malicious and we didn't do the right process to actually go through and validate to make sure that we removed anything before we do the restore. So with Veeam, we think about that and we add in some of those capabilities for malware analysis and being able to scan depths.
And I'm going to show some of that. I have a quick question- Sure ... about the recovery process.
Is that something that has to be done before the data is restored in flight, or can it be done at rest? It can be done at rest. So effectively, I can store that data, and when I know, just say, for example, there was a zero-day that came out.
Once I've detected that, I can go back through my previous catalog and say, "Oh, yeah, we have a signature capability to remove that," and we can go ahead and disinfect. Absolutely. Yes.
So there's actually a few different ways. You could do it during the recovery process. We could also do it not before the recovery process, so just take our version of different backups.
You could sit down with your security team. They could provide you YAR rules, or they could provide you with that AV signature that you want to be able to go ahead and scan those backups with. And we could sit there, and we could scan multiple iterations of it before it actually goes into a recovery process.
And that can all be done from the backups leveraging our core technology, which was instant recovery, of just mounting and going through and scanning those workloads. The other thing I'll add onto that, Tom, as well, is being able to provide a sandbox environment so that your security team can go and test what does that patch even look like before I even roll it out to production. So put it into an isolated environment, inject your update, whatever that may be, to fix the vulnerability, make sure everything works together in this sandbox isolated environment, and then tick, drop that down, and then go and do it in production.
So you're not infecting, or you're not causing a problem in production straight off the bat. Yeah. So I'll run through what this actually looks like too because we didn't cover the cloud.
Orchestrator covered the cloud, but I'll cover that off as well. So you can see in here, let's say, for example, we need to create a recovery plan. And here, I could actually choose the plan type, depending on what it is I'd like to recover to, whether it's going to be cloud, recovering to a new hypervisor platform or restoring from a replica.
I could select my different backup types. And then from there, I could actually go ahead and create a new plan. And so for this example, we're going to say this is our recovery plan.
It's going to be for Tech Field Day 2026. And for my recovery objectives, I could actually set in here what my RPOs and what my RTOs are. And those are going to be something that'll get tested and validated for me.
Now, I'm actually going to be recovering this to a Hyper-V environment, so I'm taking VMware backups and restoring these over into Microsoft Hyper-V. And the great thing about that is when we start talking to customers that are going through the woes of Broadcom, or maybe they're just reimagining what their environment could look like from a migration strategy or from a migration process. The benefit here is now you actually get an opportunity to test your migration strategy at scale.
I could select a whole bunch of hosts of different backups from here. I can choose how I want to go ahead and have those be migrated over to this new hypervisor platform or even sending them over to Microsoft Azure. So you can see all of my virtual machines that I have associated, and I could pick and choose and select the areas in which I want to have these recover in what order.
So you also have that opportunity to do so here. Now, once I create this plan, the one beauty of this product is that, number one, I could just automatically do a readiness check, a lightweight check. " Meaning, do we see any warnings?
Now, I did choose a workload in here to show you what do those warnings actually look like. So this is going to be the full set of documentation that a user will see just from creating that plan. You get all of this documentation of showing these are the workloads, this is the area that we are going to be recovering to, these are the groups in which we've actually set them up, and of course, we're seeing that you already have some warnings in here in terms of the actual recovery that possibly might not work or is already missing a recovery point objective, meaning I set my RTO or RPO for 24 hours, and so that backup job in particular is past that 24-hour period, so it's not going to make that RPO that I have set up.
Now, another way that you can leverage this is, again, recovering to a cloud. So we can leverage Microsoft Azure as a target, but going back to Tom's point, I can actually scan before I restore into Microsoft, leveraging YARA rules, leveraging AV signature scans, and I can plug those in here to do that scan before it actually restores up into Microsoft Azure. " So that way we could do additional tests.
I could have my security team go through and actually take a look at it. Can you do the scan as you're backing up the data so that you can then identify an infected machine, not waiting until it's recovery time? Yes.
So in the next session, we're going to cover all the different ways that we can scan the data. So with Veeam, we could do it before we actually started a backup, so we can inform you of current risk behavior that is happening within the production environment. Then during the backup, we could scan inline and we can check to see for any type of anomalous activity.
We could look for indicators of compromise. We can search for tools that threat actors use to perform exfiltrations of data, and then also scanning with our own AV signature-based detection. And then after the fact, we can go through those different iterations of AV signature plus YARA rule scanning.
So three different versions or variations of which we're scanning data. And then on top of that, we here at RSAC, large ecosystem of providers and partners that are here. We integrate with over 60 of them.
So if CrowdStrike is finding something from an ADR perspective, they could send that information to us and they can inform us of that potential information that's happening in production. We'll talk a little bit about some of the other capabilities as well. But yeah, from this standpoint, we could do multiple versions of those scanning capabilities before we actually perform a recovery.
So the last part of this is the why it matters. So for us, it's all about making smarter decisions with that data. So when we started this acquisition, it was for that specific use case.
How can we make sure that we are understanding the data that currently exists within the organization? How can we leverage that to make smarter decisions around our RTOs and our RPOs versus maybe somebody just guessing what it is? The beauty of the product that Veeam has built is that the fact that we can actually go through and test and validate those SLAs for our customer and inform them, "Hey, if you have an SLA of 24 hours RPO, clearly you have backup jobs that aren't hitting those RPO periods.
" And then on top of that, this gives us a capability of being able to precisely recover what is necessary, depending on whoever made the change, whether it was accidental deletion, disaster recovery initiatives, or even agentic. And then, of course, on top of this, this just helps us to validate SLAs and compliance more. All right.
Firstly, are there any questions on the understand? And then we'll move on to the secure side. I think we answered the questions as we were going.
Okay. So we'll wrap up there. That's the end of our session for the understand data.
I'm Emily Tayas. I'm Michael Cade. Thank you.