Mike Miller with AWS on How Automated Reasoning Makes AI Systems Trustworthy
In this Techstrong.ai Leadership Insights interview, Mike Miller, director of product management at AWS, explains how automated reasoning—grounded in centuries-old logic and mathematics—can improve the trustworthiness, reliability, and correctness of modern artificial intelligence systems and applications.
Transcript
Hello, and welcome to the latest edition of the Techstrong AI Leadership Insight series. Today we're with Mike Miller, who's director of AI product Management for Amazon Web Services, and we're talking about well trustworthiness when it comes to ai, which can be automated with some math apparently. Hey, Mike, welcome to show.
Hey, Mike, thanks for having me. Glad to be back with you. All right.
So explain if you would, I mean, we're all talking about trustworthiness when it comes to ai and everybody's concerned, especially with AI agents, but I think we have a thought in our heads that somehow rather, this can be done with some other external magic thing, but maybe it just requires good programming and some old fashioned math and reasoning. So can this be done and how can it be done? Yeah, I, I would say kind of all of the above and, uh, you know, I I, I'm super excited about this because as, as you or your listeners may know, uh, yesterday was World Logic Day.
And, uh, logic is really the foundation of, uh, a lot of the kind of tools and techniques that we're gonna be chatting about today, uh, about how you can, uh, improve trust awareness of AI and of AI agents. Uh, and, and if you, you're kind of thinking like logic, like, wait, like high school geometry? What is, what, what do we mean by logic?
And, uh, logic actually like, kind of imbues our life every day. And we, we use logical, uh, deductions, um, and logical inferences all the time. So if you think about like, uh, it rained yesterday and overnight, it's gonna drop below freezing, therefore I need to watch out for, you know, ice on the sidewalk.
Right? That's a deductive reasoning. It's kind of taking these sort of logical statements and putting them together and then reaching conclusions, uh, that, you know, are true because of the sort of inference capability.
And so that's kind of the same approach that we take, um, with this, uh, technology called automated reasoning. It's this new, it's, it's actually not a new field. It's actually been around for, for many years.
Um, and it attempts to use mathematical logic to provide assurance about what a system or a computer program will do. And that assurance is based on mathematical proof. Hmm.
I get the concept, but, um, how much skill does it require to implement this? And do I have to be a rocket scientist or is this something that a me developer or your average data scientist can wrap their head around? Yeah.
Well that's actually, uh, one of the really, uh, interesting, um, implications of AI and machine learning that we've seen, and that in the past, um, you know, AWS has embraced, um, automated reasoning for, you know, the past decade. And we've had to bring on board, uh, PhD scientists who, you know, spent their life sort of diving into this mathematical logic and reasoning and understand how we can represent computer programs in this mathematical format. And one of the things that we've seen is with the advent of, uh, LLMs and sort of generative AI, is that that task has now become a lot easier putting this technology within reach, uh, of your common developer.
And so what we're working on are capabilities that start to imbue, um, our products with automated reasoning. And this combination of, uh, generative AI machine learning and automated reasoning we call neuro symbolic ai. 'cause it combines the two sort of elements that statistical sort of prediction kind of elements of AI and these mathematical sort of symbolic reasoning elements of AI and, uh, and coming together to kind of enhance, um, the capabilities and, and drive truthfulness, right?
Because if you think about the world is moving to AI agents, um, you know, and these AI agents are becoming more complex. They're handling, uh, more independent tasks, they're taking actions. And so we're gonna need to feel a high level of confidence that these agents are operating according to, like policies or rules or guidelines and making the right decisions or preventing them from making sort of, uh, the wrong decisions.
And this is where neuros symbolic AI can kind of come into play. Did we just simply forget about this AI discipline? It sounds like symbolic ai, I, I seem to remember hearing about this before, and I, I certainly know that reasoning's been around for a while and mathematical proof.
So did we just overlook all this and are now rediscovering it, or where Well, no, it's, you know, it's actually been used for quite a while, but because of the sort of specific detailed skills needed through the science application, it's really only been useful for the sort of most high risk kind of, uh, tasks like, you know, constructing software for the space station or managing rail networks, things like that, where the cost of a, you know, mistake was very high, where it was worth investing all this time, um, bringing these scientists on board to build these mathematical representations and make sure that we could prove the correctness of these systems. But generative AI has allowed us to really massively accelerate that, um, and create these sort of mathematical models in a much easier way. Um, and, and that's kind of what we're doing here at AWS now is taking that decade of experience, um, determining where the right places are, where we can apply this automated reasoning and then getting those into the hands of our customers.
Mm-hmm. So how do we know or validate something? I mean, I get that we can do it with the reasoning and, and programming, but at some point, will a third party need to validate the math or how does that kind of work?
Yeah, like, let me give you maybe a, a a quick analogy and then we can talk a little bit about how, how you can sort of audit these things, right? So if I think about, um, you know, back to geometry, your right triangles, right? Euclid 2000 years ago, sort of reasoned about right triangles, and what he did was he proved, uh, the Pythagorean theorem, A squared plus B squared equals C squared, right?
The lengths, the, the, the lengths, uh, the two shorter sides of your right triangle when, you know, squared and added together equal the length of that longer side. Um, and what he did was he proved it in a way that allowed us to, uh, understand that that was applicable to the infinite number of right triangles. Uh, because you could have gone about this and said, well, let me examine, you know, hundreds or thousands or tens of thousands of right triangles and kind of estimate what the relationship is between these sides.
And that's kind of how machine learning works today. You give it a bunch of training data and it sort of starts to generate, um, you know, predictions about, about sort of, uh, the questions that you're asking, whereas reasoning in mathematical logic basically says, Hey, over the infinite number of sort of scenarios here, this output is gonna be guaranteed the same time. And so when it comes back to sort of auditing, how do we know?
Well, automated reasoning is really interesting because it doesn't operate like a black box. It's, it's open when we, when it performs a proof, uh, we can sort of print out that proof and use that as sort of an audit trail to validate the behavior of the system that it's being proved correct. So there is a built in sort of visibility, um, and explainability that comes into play when we use automated reasoning to validate programs.
Will this solve the following issue? I mean, I talk to people about this and they all like generative ai, but they realize that the outcomes are probabilistic and a lot of the workflows that we're trying to apply it to, or deterministic in the sense that we want them done the same way every time. So does this give us a mechanism to kind of marry the two in a way that is reliable?
Yeah, you, you hit the nail on the head, Mike, and that's why we're so excited about using automated reasoning, uh, to help customers, you know, use generative AI in these trustworthy applications, uh, and know that when they're building agentic systems, um, you know, we're validating or constraining or sort of making sure that the behavior of those age agentic systems, um, is, uh, according to our policies. And in fact, we kind of do that today at AWS already. Um, you know, we have a couple of products.
So one product that we just recently announced is called Policy in our Bedrock Agent Core Product. So Agent Core is a set of tools to allow customers to build these agentic solutions and policy. And combined with another one of our products called, uh, gateway allows you to define in natural language policies that you want to restrict the behavior, uh, or constrain the actions that your agents can take.
Um, and these policies are then applied and validated through, uh, formal reasoning and automated reasoning to make sure that, uh, you know, that generative AI isn't gonna sort of slip through the slip through a crack and kind of perform an action that wasn't, uh, according to, uh, the policy that you've defined. Mm-hmm. How will this manifest itself at the end?
Will there be some sort of output with a little check mark next to it that says that this has been validated, certified? Or how do you envision us understanding when this has actually been completed? Yeah, I think we, I think we work with our customers, um, you know, to help them kind of find the right way.
If they're building an innovation and they're using these techniques to, uh, provide higher assurances to their customers, you know, we'll work with them on sort of how we do that. One of the ways that we did it is through a couple of our tools. So when customers are building solutions on, on AWS, they have to define policies and kind of access controls.
And so we actually have a tool called the IAM Access Analyzer that allows customers, uh, to, uh, plug in what kind of policies and access controls or details they've got, and we use automated reasoning to validate them. For instance, if I make this change, does it become more or less restrictive, right? Because let's say I only want my admins to be able to modify this database.
If I create a policy, I can kind of validate whether that policy is more or less restrictive than needed, and we use automated reasoning to do that. So it kind of depends on the particular application, um, and the type of assurance that you wanna provide, you know, the end users, whether there's like a check mark or a validation or, you know, uh, sometimes we've talked about customers about just providing like a, Hey, explain this to me, or sort of like, you know, show me where this answer came from. And that's where we get back to our auditability and explainability and we can kind of detail, uh, the proof or sort of the reasons that were used for this, uh, for this kind of answer.
In fact, we have one product, uh, called Automated Reasoning Checks for Bedrock Guardrails, which kind of does exactly this. It's a product that helps to detect and remediate hallucinations when you're interacting with a chatbot. So where there's like a defined policy, let's take for example, like an airline ticket refund policy.
What we can do is we can ingest that and then build out sort of a, a formal set of mathematical logic that defines the terms and conditions, if you will, of that refund policy. And so, when I'm chatting with a chatbot about getting a refund for my airline ticket, uh, bedrock Guardrails can be used as a guardrail to validate the LMS output. And so we can be certain then that when it sells me, like, Hey, yes, you can get a refund on that ticket because you bought it within the last 30 days, and it's unused and you live in, you know, this US state, uh, we can know that that's a valid answer and the, the customers can see that in the, in the, uh, chat bot experience based on what the customer wants to do.
Yeah. The curiosity for a technology that's been around, you would think I would've seen or heard more, uh, providers offering similar capabilities. But, um, is there something here that's unique about the way AWS went after this or, um, yeah.
How come I just don't see this as a broad-based capability? Yeah, these techniques? Well, so there's a couple AWS products, um, that already leverage automated reasoning, and we didn't necessarily make a big deal about the fact that automated reasoning powers these things.
So I mentioned the, you know, access, the, the Policy Analyzer, the IAM Access analyzer. We have, uh, network reachability analysis tools. There's a tool, which you might have heard of, called S3 Block Public Access, uh, which we actually use automated reasoning.
So we mathematically validate that when you turn that on, uh, there's no way that somebody from the public internet can access that S3 bucket. So a lot of times these capabilities have just been ingredients in, um, you know, sort of security, um, and related sort of high risk kind of capabilities or, or products that we've released. Um, I think one of the reasons why you don't see this, uh, proliferating across a lot of other providers is that, uh, AWS really has the deepest bench of this automated reasoning science experience, right?
We've been building it up over the last 10 years, and really right now, uh, we're starting to see that come to fruition with the Bedrock guardrails, the Agent core policy. Um, you know, we have a number of other products related to, uh, internal validation. So, um, you know, uh, so in our, in the work that we do in building our own chips, right?
For Graviton five, for instance, automated reasoning played a key role in some of the, what's called the Nitro Isolation engine, to validate that when multiple customers or you know, multiple tenants are using the same server, that there's no possible way that they can access each other's data. So we've been using the automated reasoning, but a lot of times it's kind of behind the scenes and it's just an ingredient to drive, uh, you know, the security and the reliability and the trust in, in AWS products. So, the way I, maybe I should start thinking about this, I mean, we've all been obsessed with generative AI since it came out, but, um, maybe the future of AI is using multiple types of AI models alongside each other, and they'll be predictive and generative and causal, and now symbolic.
And it's all about mixing and matching these things when needed. That's right. There's a whole, um, kind of universe of techniques that companies can use to drive trustworthiness of their AI solutions.
Uh, and, you know, you need to apply them, uh, sort of gingerly, if you will, right? So there's automated reasoning, uh, there's machine learning, there's guardrails, there's fine tuning, uh, there's a whole range of techniques that all can kind of be put into play to improve the trustworthiness, um, you know, of the solutions. Uh, and I think in 2026, we're really gonna see, uh, that start to come to fruition and kind of trustworthiness, especially as we get more into age agentic, um, development and agentic tools are gonna be operate on your behalf.
Uh, that trust, that trustworthiness is gonna be, you know, a key element of these solutions for, you know, uh, people to adopt them. All right, folks, what you heard in here, they used to say back in the day, you know, don't trust anyone over. Uh, try that again.
Lemme try that again. We write it down. Well, folks, you heard in here, they used to say back in the day, don't trust anyone under 30.
Now we're saying don't trust any AI agent that hasn't been validated. Hey, Mike, thanks for being on the show. Absolutely.
That's a great one, Mike. I appreciate your time. All right, and thank you all for watching the latest episode of the Techstrong AI Leadership Insight series.
You can find this episodes and others on our website. We invite you to check them all out. Until then, we'll see you next time.