Microsoft Security Copilot Conditional Access Optimization Agent
This session explores the evolution and capabilities of Microsoft Security Copilot, focusing on how it’s transforming security operations. Microsoft Security Copilot operates as a unified platform, providing a consistent user experience across its various agents and underlying products. Key features like transparent decision trees, identity and RBAC management, and human-in-the-loop design principles are common across all agents, ensuring that users retain control and can audit AI-driven actions. The Conditional Access Agent, for instance, autonomously scans policies and recommends changes to ensure they align with the current state of the business, enabling rapid updates to security posture and reducing the risk window from months to minutes or hours.
The system incorporates robust guardrails, allowing organizations to control agent operations, particularly concerning new users and applications, and to apply custom natural language instructions to tailor agent behavior. This ensures that AI-generated policy recommendations are balanced with human oversight and business context. Users can also provide feedback to the agents, which directly influences their future reasoning and decision-making, akin to training a new human employee. This continuous learning mechanism is crucial for the AI to adapt to an organization’s specific nuances and improve its effectiveness over time.
While agents are designed to handle resource-intensive tasks like triaging user-submitted phishing emails, the generative AI component is not intended for real-time, high-volume inline processing due to its computational demands. Instead, Microsoft focuses on applying AI where it can most significantly augment human efforts, such as automating time-consuming and low-value tasks. The platform aims to provide clear metrics like resolution rates and time to triage, allowing organizations to assess the economic value of deploying these agents. Furthermore, Microsoft is committed to expanding integrations with third-party data sources and partners, empowering agents to leverage a broader ecosystem of security tools and data, and ultimately enabling customers to build more comprehensive and adaptive security workflows.
Presented by Nick Goodman, Product Manager, Microsoft Security Copilot. Recorded live at Security Field Day 13 in Santa Clara, CA on May 29, 2025. Watch the entire presentation at https://techfieldday.com/appearance/microsoft-security-presents-at-security-field-day-13/ or visit https://techfieldday.com/event/xfd13/ or https://techcommunity.microsoft.com/category/security-copilot/blog/securitycopilotblog for more information.
Transcript
I'm Nick Goodman. I'm a product manager here at Microsoft. I'm also one of the, uh, here's the conditional access agent.
So a couple things to pay attention. I'm just gonna highlight 'em. A lot of the screens are gonna look really similar, and that's because security CO's a platform.
It's behind all of these products. It's behind all those third party products that, uh, I mentioned earlier. And so the things like the decision tree that helps the human understand how the agent made his decision.
That's common. The concepts of identity and art back management, that's common. So there's some of these that will be somewhat repetitive and I may fast forward through them in the interest of time.
Uh, but the idea with the agents is primarily to focus on the business use case. But I wanna make sure to show you this one too. Okay.
So here we are in intra, I'm gonna fast forward a little bit. Again, same concepts. I can see how it's triggered, what permissions it requires, what licenses it requires, it runs in the background and when it gets running, starts running, um, we start to see it scanning policies.
There are guardrails on this one that I'll show you here in a minute. Um, and it starts making recommendations about what we can do. And so these recommendations are things that effectively are policies that it's recommending, uh, that we add.
Um, and so we might see things like MFA policy or or device compliance policy, um, just like all of the other things that we've been showing you, these are, these are designed to be human in the loop. They're designed where you are making the decision. You are still in charge.
Uh, the AI is not changing state in your environment without your knowledge. Um, again, I see a, a summary of what's going on and I can drill in to see the details. Again, really similar UI here.
The idea is this decision trees, it shows you the steps that the agent took. It's got tags on it so that you can understand what's important. I can click on them for details.
Um, and it really covers the full flow of the agents. So it's creating this transparent path which allows you to debug it. Also, if you've got an organization that's, um, maybe highly regulated, needs to be able to say that, Hey, my operations that use ai, um, are responsible.
Um, these are, these approaches are part of how we deliver on our responsibility. And at the end of it there's po again, there's policy recommendation that it's making and or multiple in this case actually. And I can drill in and take a look at it.
And so the policy is gonna show me, here's the policy changes. You know, if I didn't know this was AI powered, I might not realize that it was AI powered. And so one of the ideas again, is to have the process.
The person is used to the product, the person is used to have the AI built directly into it, and we still have warnings to help you know that, that it's an AI process that you're seeing. Now, I also have the ability, uh, uh, to drill in here. I'm gonna have it drill in and take a look at the actual data.
So policy is just a JSO document. So I can also get the diff of what's actually changing. I don't have to necessarily read over the the page.
So when I apply the policy, this is what will happen. And again, we're now down to button click deploy of the updated policy. So, uh, Laura, Sara, question.
Yeah, Because configuring conditional access can certainly prevent many users from getting access, like you said earlier. Mm-hmm. Is it possible to get on statistics before you apply such a policy that, uh, users within the past week will now no longer be able to log on because it's changed?
Very good question. Um, that's absolutely, uh, the type of thing that people have concerns about with this agent. And you couldn't have asked a better question for my next couple of slides here.
Um, because these are the guardrails. So, um, I can the guard and, uh, I'll answer it. And if I didn't answer it here, like within a slide or two, holler at me.
So your guardrails are pretty much your, your basics, things like triggers, but you also get to control what it can operate on. So what we've heard from our design partners and customers is that the most important guardrail, um, is the new users and new applications. And they often don't want the agent making policy recommendations for new users and new applications within a 24 hour span of, and that giving you the choice to apply it to those seems to help a lot with preventing some of the worst types of problems that can happen.
And um, again, also the, these policy choices in the check boxes down here about MFA and off strength, these are based upon actual, um, learnings from customers about where this agent can go off the rails and what type of guardrails you might want. You can also apply custom natural language instructions upfront here so that you could say, I really want you out of this area, don't do this. Or, or other sorts of things.
I think I lemme see if, I think I might have another one on here to show on this. Sorry, let me speed it up slightly there. Yeah, again, you might wanna always exclude users with a break glass account from an MFA policy adjustment.
Um, sorry, I gotta speed it up a little bit here. Do I have, yeah, so those are the pri these are the primary guardrails that we have implemented based upon customer feedback. I'm curious how well does that answer your question?
Yeah. At least you have song guard rails there, but it will be nice to have some, uh, and also more than 24 hours I think will be great. Okay.
Good feedback. Thank you. Uh, okay, so schedule the guardrails.
Um, let's see here. Uh, it also, like the other ones shows KPIs. Um, and again, you'll see there's a, each of these agents has a little bit of its own flavor.
The data is the same, the concepts are the same. Um, it's just designed to fit within the workflow of the user. So you will obviously, so see some difference between what an IT administrator, identity administrator's looking for and what the security admin is looking for.
Fast forward it through time here is protects lots of people. Cool. So, uh, that's the scope of my demos.
What other questions can I answer for you all? So one of the things that Nick, that you mentioned was that, um, some, somebody might not even be able to tell that it's co-pilot behind there, which I think is great, but how do we go about telling extended team members that co-pilot's part of this, but it's also human interaction or human in the loop with it? Yeah, it's a, it is a great question.
So, um, and there's a couple of of things that you wanna probably care about there. Um, probably first and foremost is AI could make mistakes. Mm-hmm.
We all know this and you want humans in the loop, humans to review. And so I'm gonna back up on my slides and show you some examples of the, the things that you'll see in the product to help. Okay.
Lemme go back. So we've tried to design the user experiences for it to be feasible for hu easy for humans to figure out that AI is involved. Nick, I I don't wanna interrupt.
Are you currently sharing your slides? 'cause we don't see them. Um, I'm almost there.
Okay. Just double checking. Okay.
So here's an example. Continue to share. Yeah.
So here's an example. So you'll notice there's an agent tag on each row, uh, that the, uh, AI operated on. And the idea there is, even in this big list view, I should be able to pretty quickly find it.
Um, and because tag, I can also filter for it. So I could build and save filters that help me easily find what AI has done. And then you'll also see on here, when I look at the incident, it has, it says that it's labeled as being operated on by an agent.
Uh, you'll see labels like right here, AI generated content may be incorrect. That's often paired right next to the thumbs up and thumbs down where you can provide, um, indication back to Microsoft that this agent is not working properly. Um, it's also why I'm gonna fast forward to this design right here, why we have a common decision tree that you'll see in here.
It's not just about engineering efficiency, it's also about if people are gonna learn new ways of checking on ai, it's really helpful if the way they do it in defender is the same as the way they do it in intra is the same as the way they do it in purview. We can't have perfect harmony across all products and all aspects of the world, but we tried to make sure that inside the security domain, at least all of the user experiences are the same. So that once the, once the security team goes, oh, I know if I wanna see how the agent reasoned about it, I go find the decision tree, I can drill in, I can look for things that are red, which tell me it thought there was something fishy going on on, I can click on a note in the tree and it'll give me details treating the user to be able to do the same thing over and over rather than have them do it.
Uh, learn different processes for each product. So that's our kind of core approach to this. And how do you get like the anti-spam versus the threat management to play well together?
Is it, is there a copilot for that as part of defender? And how do they, how do you deal with both things at the same time? Um, elaborate a little bit more on, on what you mean.
Well, Just regular anti-spam that we know that most email providers have to worry about. And if I'm using 365 as my email or exchange online is my email predecessor, how are these things gonna not stomp on each other or get in the way of each other? Gotcha.
At the moment, our approach on the AI side has been to be very targeted about where we apply ai. And so that's allowed us to think as we're developing it about questions like how does this fit into the bigger workflow. Mm-hmm.
Um, and hopefully land the right design patterns that scale as we move AI into more of the, like I said, the threat protection piece and the anti-spam piece. Um, so that's, that's really our approach is, is be very targeted with it right now and learn how to do it right so that as we scale up, we can continue to do it effectively. Thanks.
One question on scale. Yeah. Uh, since, uh, how should we, how should organizations estimate cost for agents?
I know you have SCUs that cost four bucks an hour or something like that. Uh, what, how much did this particular run cost and Yeah, I'm, I'm, uh, actually surprised it took us till the 25 work to get to cost. Um, right now these agents are in preview and for customers that already use Security co-pilot, um, we've made that preview free.
Okay. And that's partially so that we can learn how they perform in the real world from an underlying cost perspective so that we can make sure that when we move them to being generally available, the costs make sense for the business cases that they do. Um, so obviously I'm, I'm not giving you a terribly direct answer about how you would do this, but you picked, you did Intuit correctly.
Um, it's designed to be on top of the SCU. Um, our intention is absolutely to make it, uh, cost effective for customers, uh, to get the, get the value out of these use cases. Um, and it is a learning process and we're in the preview stage in parts so that we can learn that element so that it lands correctly.
Uh, when we get to ga What's the, uh, interactivity with the agents? How, how interactive can you get with the agents right now? You sort of said, okay, well the agents are running, for instance on the, on the email phish the email phish detection here.
Right. So, you know, it shows you that, okay, here's the list of alerts and 75% of those were tagged with the agent. Right.
Is there any way you can go beyond that and say, okay, I want to ask the agent questions, show me all the other emails that are like this that you tagged or like this that you missed or all those types of types of things. It's not something that we have built today. I'm sorry, say that again please.
We have not built that today. Thank you. Hey Nick, what, looking at this from a, a temperature perspective of these agents, um, is there a possibility that we could see variance between what an agent recommends on one run versus another?
Um, specifically example would be, you know, conditional access policy. Um, one day it recommends that these, you know, 14 user accounts get multifactor applied and then the next it's go use passwordless. Oh, uh, also a great question.
Um, when you think, yeah, and that's, it's a natural, it's a, it is a natural question. Uh, like you said, like you've got temperature settings and things like this and, and we've all seen that, uh, with just prompting there, there's an both an art and a science to getting, uh, generative AI application to return consistent results. And so like we think about that int conditional access policy, some of what drives consistency is our use of code terministic logic mixed with ai.
So for example, uh, oh and state management. So for example, um, if say I don't apply the policy recommendations and I, I'm just, I leave them in, I don't dismiss, but I leave, it says, Hey, add multifactor to these 16 accounts that's still there the next day and then the next day. So it's, it, it knows that it has made the recommendation and that you've got it there.
And so it's not gonna recommend that again. It's also not gonna like automatically remove it because it was still a valid policy recommendation To the question though of like, what if I just keep rerunning it? Well, the agent is designed to recommend the most valuable policy enhancements.
And so, and I believe it's, I think the parameter we have it set up now to do is to give you three. And so it's always giving you the most valuable three policy enhancements. And so that is a stable list.
So if you were to say, um, uh, hypothetically run it and get it like reset and run it again, it would give you the same three policy statements over and over. Um, and if you say apply number two and you open up a slot, well then the fourth most valuable would then the next time it ran pop into that. And that's been a lot of the development work on the agent has been to ensure the consistency of the policy recommendations.
Uh, 'cause you're right, because otherwise you would, you could have a, a confusing system if it was constantly switching back and forth between the types of recommendations. The other thing it does, I didn't show at all, um, it tries to make a minimum set of policy recommendations, which is also an interesting art and science question because I could have said, apply MFA to to A apply MFA to user B to user C to user DI could create 16 policy recommendations to add MFA to 16 users. But the one that I have an example correctly lumped them together 'cause it's the same policy.
Um, and the, a lot of the tuning of the agent has been to drive it into doing that sort of behavior, uh, just against the, the, the human and the, the business process involved have the simplest, uh, path forward. Awesome. So follow up question.
You said that it, it recommends the most valuable. How do you understand and how does it interpret what's valuable to the enterprise? Is it going off of trusted advisor score or, or security recommendations?
Yeah, It's, it's a good question and there's obviously judgment involved in how we code that, how we design it, and, um, and reasonable people could disagree about some of those judgements we've made. Um, the, it's effectively based upon our understanding of how to use these policies as effectively as possible, um, and based upon what we've learned interacting with customers. And the override piece is the guardrails I showed where you can turn pieces on and off as well as the natural language instruction.
So if you think Microsoft is particularly wrong about one element of, of a policy, of a type of policy, you can easily tell, Hey, I don't, don't do it this way. Um, or alternatively, if you think something is way more important than our policies, sorry, kind of default ranking puts it, you could say, I really prefer this type of policy. Gotcha.
Is is, is it also looking for things where, um, I, I liked your example where number four is on the list, but it's not shown yet. Um, could I end up in a scenario where number two is block x, y, Z country, for example, um, and then it comes back and says, no, go and block this region now that you've implemented that? Or is it just gonna come out and say block the region?
Uh, uh, so like, uh, policy four was a subset of policy two. Yeah. Uh, should not happen the way this is designed.
Okay. Perfect. Nick, just question on my side.
Uh, we expect of course that the, the use cases that you are launching, uh, are working pretty good on, uh, on the Microsoft products boundaries, right? Mm-hmm. Are, are you planning any future also to interact with, uh, Microsoft ISVs and technology partners?
Yeah, uh, yes. Um, some of that happens today. So some of the, some of the, uh, agents have abilities to pull in third party data.
Uh, threat intelligence, for example, is a space where many customers use multiple threat intelligence providers. Um, and I believe our threat intelligence agent already has integrations with two, I'm not gonna name them 'cause I'm not 100% sure, uh, off the top of my head if I would name them properly. And I don't wanna do disservice to our partners.
Um, but I believe we've got two, uh, third party integrations that if you've got a license to those TI vendors, they show up in the agent. This is an area we're exploring of how do we do this right? Um, uh, to incorporate the, the right intel into that agent.
Uh, there's other obvious places where it's applicable. Um, the defender, uh, user submitted Phish agent, uh, could operate on emails from outside Microsoft, I mean like email servers outside of Microsoft and, uh, email protection systems outside Microsoft. Um, I don't believe we built that yet, but it's the type of thing that we would like to do.
Um, our general strategy here is to work with the customer's ecosystem. We've started with Microsoft because we have the, the resources on those teams that have been committed to go build the capability and the integration. We want the Microsoft version of it to be a great example and to be highly valuable.
And we want, uh, third party integrations both into our agents and third party agents built on our system. So like the seven agents built on our system, we worked with those teams, educate them and train them use when we're using that to inform like a more scalable model where we can bring in many more partners to help them build agents on top of our system. Okay.
And, um, okay, so basically mo mostly on gathering data from other other partners, but do you think that in the future a customer that maybe doesn't have a mainstream technology, uh, that can integrate it, uh, to the, the copilot to the API and then, uh, of course program it to, to interact with them in a specific Yes. The, the short answer is yes. Um, the, there's complexity in doing that because, uh, like if the data source doesn't contain the information, it may be that the agent is incapable of doing the type of, like the agent has to have the data to do the type of reasoning it's designed to do.
Um, but the, but the intention on, uh, on these is to be able to operate on data wherever it lives, um, and, and to be able to trigger action, eventually be able to trigger action in the systems where the action needs to be triggered. Um, and we do have agents, I don't want it to sound like it's just about data. Um, many of our third party agents are not feeding data to Microsoft.
They are using Microsoft's like the security copilot AI reasoning engine to do Angen operation on data in their system that drives workflow in their system. I think of, um, OneTrust is a, a really good example of this. They've got a, a cool privacy workflow, um, and there's really, it's Microsoft is providing data to them, uh, for shared customers, and we're providing the ai, uh, brain for the orchestration, um, of their agent.
Um, and it's a more, it's a very OneTrust centric process that they've built following up The, the privacy topic. And then that's a link me to another question, you know, uh, will be a moment on which, uh, compiled will integrate with different technologies or from the private perspective and data sovereign is important to have a clear map where is working with the, and uh, when are you Yeah, exactly. And we've, we've taken steps in that direction.
So, uh, security copilot is built to be audited not by us, but by our customers. Um, and so we emit audit logs and audit records into the audit system. Um, the identity, we think identity is very important both from a security and a privacy and an audit perspective.
And so agents have an identity, uh, that allows, uh, really empowers the customer to control what's going on and to govern it. Um, same thing with like each of the agents has RA permissions on it so that you can govern what data they have access to. Um, uh, so yeah, it's, it's, it's similar in many ways the mental model of I'm bringing on a new, a new human workforce, how do I govern what they have access to?
How do I audit what they did access? Um, and it's our intention to make this work smoothly for enterprises that have those needs. Just following up a little bit on what you mentioned about third party and OneTrust and the data flows that happen, uh, is the reasoning capability of pilots, uh, implement.
I'm not sure if you can comment as, uh, with within your typical gen AI model because you, you refer to before that you have code as well controlling some of these things. Is the code that the non gen AI code that you've written something that outside of the reasoning model, or is that something that's inside the reasoning model? Um, kind of an answer is both.
Um, if you think of it, so if a, if a, if a third party, so a non-Microsoft developer we're building on top of our system, yes, they get the full benefits of everything we've written. Um, so, and, and the both in my answer was the AI platform itself has deterministic code and AI in it. All of that is available to everybody who builds on security copilot regardless of what company they work for.
Okay, thank you. Obviously there's licensing and authorization and security checks to be in the ecosystem, but once somebody is in the ecosystem, they have access to all of this. The part of the reason I said both is a lot of times security copay is OP is interacting with existing SaaS products.
Most of those existing SaaS products are primarily non AI products. And so there's a lot of non deter, or sorry, deterministic code in the things that we invoke and trigger and call. What, uh, what type of metrics are available for understanding your use of the agents and their success and those types of things?
Yeah, so, uh, there's a, a few different metrics. So you've got KPIs that are specific to the agents themselves. That could be things like resolution rate and time to triage.
Um, you also have common ones like success rates, um, uh, latency, uh, what are the, what are the other ones that are visible? Um, I think those are the two common ones off the top of my head, uh, that are visible. But yeah, the idea is again, with the frame, the agent framework exposes a certain set of core metrics back out to the, the products that use the agents.
How does, oh No, go ahead. Sorry. How does this, how does security copilot deal with like the problems of data exfiltration or data loss prevention, any of those types of things?
Um, and when you, when you say that, you mean from IT or in assisting those workflows In detecting those things? Like, I don't know if it's supported by security copilot much like a fish. Yeah.
Yeah. So, um, it can help on the triage piece. Um, we don't currently have an agent that is actively doing detection in that space.
Um, it's definitely the type of thing our, our purview products group is looking into. Okay. Um, they targeted, you know, it's an interesting question.
Like when you look at what we targeted, you'll see three of the agents are about triage, and three of them are proactive. Two of 'em are proactive policy effectively, and one is a proactive briefing. The three on triage, we picked them because that triage space is really great for AI because it's very human intensive work.
Um, are time con time intensive work that humans are doing and often see no value in doing. Um, which just makes it like a really great space for ai. And then the three proactive, we love those because they are, they're improving posture, the reducing risk, and they save time and we really like that and in the equation.
Mm-hmm. Yeah. So I'm thinking of things like, you know, dealing with the false positives, the positive negatives, whatever, is that, you know, every time I send around a test database that has credit card things in it, I want it to be not detected because I can say that this is not real data, it's test data, that kind of stuff.
Hmm. Interesting. Yeah, so we've not built that sort of thing into the products, but, um, uh, I'll definitely, I'm ha especially if you wanna shoot me details of what you're, what interested in happy to share with the product group working.
Yep. Hey Nick, this is Tyler Christensen. What are you thinking about or doing, um, in terms of detecting rag or model poisoning, uh, from like a customer has an insider threat problem?
Yes. So let's take the p the parts of that separately. com page.
I know a lot of people who ask this question have to send the answer along to their compliance team. So, um, we do have a pretty deep, uh, deeply detailed section on our documentation for this type of question. Um, the short answer though is, uh, for the models, we control the models.
There is not a bring your own model aspect. Uh, we're controlling the supply chain of the model, the hosting of the model, the execution, and um, uh, and so like that's Microsoft's responsibility and our responsibility matrix to, to control that. Um, on the rag poisoning element, again, I'm sorry, I'm wanna make sure I'm within the bounds of a non NDA conversation.
Um, on the rag poisoning element, our system is designed to be resistant to that type of attack. And, um, obviously we don't control the downstream in every instance. We don't control the downstream.
You want to connect security copilot to, uh, NHP based server that you run, you're responsible for the data in that server. So I don't, we don't control that, but our design is meant to be resistant to that type of attack. And if anybody is interested in the room online in that, the more deeper details of that question, if you're a customer and have that type of request to get through your compliance process, uh, we can definitely give you more details than I'm able to give on a, on a public recording.