Salesforce’s Muralidhar Krishnaprasad on Measuring ROI for AI Agents
In this Techstrong.ai Leadership Insights interview, Muralidhar Krishnaprasad, president & CTO of 360 Platform, Apps, Industries and Agentforce at Salesforce, dives into how organizations should be thinking about justifying the return on investment (ROI) in artificial intelligence (AI) agents.
Transcript
Hello and welcome to the latest edition of the Techstrong AI Leadership Insight series. Today we're with Mi lid Hard k Prasad, who's president and CTO for engineering over at Salesforce. And we're talking about agentic AI and the rise of it in the enterprise.
Mk, welcome to the show. Thank you, Mike. Great to be here.
Uh, and good, good morning, good evening to everybody listening. I think at this point, everybody's generally familiar with the concept of an AI agent, but that's not quite the same thing as understanding of how to build, deploy, maintain, and secure these things at scale. So where are we on this journey right now and what are folks gonna need to know about building the agent AI enterprise?
Wow, that's, that's a big question. So what, maybe we should parse it down, right? What is an agenting enterprise?
An agenting enterprise, really where we envision agents and humans working side by side, right? To like, solve complex problems. Now, the problems could be, Hey, how do I get more people to buy my product?
Or the problem could be I have a product and there's a lot of issues with it, and how do I make it easier for them to solve it? Or simply just background processes where you want workflows to be solved for your employees and, uh, HR representatives and so on. So there are like different aspects, sort of for an agentic enterprise now, but the key thing in all of this is you could break it down into a few things.
One is you certainly need the right AI foundations. When we say AI foundation, that starts with data and APIs. You need to extra, have all your data to be able to go make the decisions.
You need to have access to the APIs to go take the actions on it. And then you need a good sort of planner system, the agentic system if you may, which can then orchestrate all of these data and actions and then surrounding it, you need to create that agentic enterprise. You need all the tools around it to make it easy for you to then observe, create those agents, observe what the agents are doing, and talk to different agents, orchestrate across all of them.
And then finally, you also need like an underlying, uh, security and metadata that brings it all together so that you don't want the agents, you want rather you want the agents saying the right things to the right people and not, uh, not like leaking your sensitive information to customers or employees and so on. So that really forms what we call as a bedrock of an agent. Take enterprise.
And of all this, of course, you need to make sure your employees or your customers, this agent take thing, is able to go reach them in all the channels of their choice, be it slack for your employees, whether it's your web, SMS, WhatsApp, whatever it may be. It should be able to go answer in all of those channels as well. So that's kind of how we are looking at it as an agent enterprise.
It's a sort of a multi-part system with a foundation of data APIs with a strong planner that augments with the tooling system, the observe, observe, uh, observability and analytics around it, the governance on top of it, and of course the security and metadata around it. Mm-hmm. How will I govern all of that?
Because as I look at this issue, each of these AI agents will be trained on handling a specific task or a set of tasks they will probably work in with each other, but at some point I need to understand what they're doing and I may even need another AI agent to validate what the first one did. That's a good point, right? Uh, in fact, even before we go to multi-agent, this is a problem even for a single agent, which is an agent is answering a lot of questions.
How do you know it's doing the right thing? And so part of what we call as Judge LLMs that we have, uh, so if you use, uh, our thing called the testing centers, we call it our agent for studio, it allows you to GoCreate tests or evals if you may. And then we have a judge, LLM that's running on the site to see if the agent actually answered in the right way.
Uh, and if it didn't, we, it might either replan or at least tell you the tests are failed and so on. So I, it's definitely, uh, an important thing because it's non-deterministic, right? It's a non-deterministic output.
We're gonna be getting. So to your sort of question around how do we really do the governance, I think there is two levels of governance here. One is there's a governance at the data layer, which you call as a data governance, making sure the right agent is give it, getting the right data so that it can actually go answer the questions correctly.
And that starts with an AI based classification, uh, and then tagging around it and making sure the agent is running in the context of the that particular user or that particular action. And so it has access to the right data. That's number one.
And again, it's structured or unstructured data shouldn't matter. We need that strong data governance layer. Then when you go up to the level of cross agent communication, you wanna make sure there is governance there as well.
And if you go back to the worlds of APIs, we had done that before, right? With API governance, making sure who has access to what APIs, a p, Salesforce, and so on. And the same thing we are trying to bring, uh, in, in the world of multi-agent, what we call this the MuleSoft agent fabric, where you can actually add the similar governance across agents to know, okay, which agents can talk to what agents, how, who should orchestrate what, uh, and putting the throttles, right throttles and security, uh, things there as well.
Mm-hmm. How will agents negotiate with each other? We humans negotiate all the time when we wanna, uh, have something accomplished.
I'll do something new, you do something for me. Right? Some agents may even have competing agendas and priorities.
So how will they negotiate with each other and come up with some sort of resolution, or are they just gonna call us for an answer? Good question. I think certainly it starts again with a few things, right?
One is, uh, things like MCP and eight two, A eight, two A in particular, the protocol's evolving, but it's giving us a good foundation. Just like, uh, you had the Ram Ls and other kind of standards before you now have the emerging protocols that's allowing for things like basic communication. It's like, okay, which agent, what can I call?
What are the error conditions to handle all of those things? But you asked a very important question, which is negotiation. How can agents negotiate with each other, uh, just like humans do?
Uh, it, it would be few things in my opinion where it'll, it's, it's really one, first of all, we need to have the right context pass between the agents. That's like crucial, right? If this agent has a lot of context and the context doesn't go to the other agent, it's gonna just make up its own thing.
That's number one. Uh, and then, so that's part of the thing that we are working on to make sure cross agent communication includes a lot of the context passing either through sort of what we have is through profiles and other things or through the protocol itself. But that's the second thing around this as well.
In a typical enterprise setup where you have very clear agent's, mark, you may have this agent doing your workflow process. I mean, for your sort of workday like thing, this agent might be doing your sales and everything else. There's less negotiation going on.
It's more around figuring out the right agents to call for your particular work. Uh, and that's where a planner comes in handy. Um, and you can, you need to make some of these things deterministic as well because you want to make sure there are very set of steps you had to follow, okay, you need to do an order first before you can go close the quote and so on.
Uh, and so we have the mechanisms sort of built in to say, how can I bring in a sense of determinism in this non-deterministic orchestration? The second aspect is when you get into the negotiation is where you have a plethora of agents available. Maybe it's like a supplier dealer like model where you have multiple agents running and you may wanna actually negotiate with the agent, say, which agent is actually going to answer my task?
And then maybe give it the, give it the job to do. com, if you remember, we had all the B2B, uh, at that time it were agents, but you had all of these companies and protocols pop up for cross company negotiations to happen. I think we'll end up creating more of these protocols with these, uh, agent exchanges, if you may, where agents can advertise what they're wanting to do.
And then as an Uber agent, you can go pass the information to all the agents to say, which one will get me a better deal? Um, and I think the more interesting things also is as we are learning, uh, this is all a learning for all of us that agents can lie to. So then we need to figure out, okay, how do we figure out is this agent really gonna tell the truth or not?
And I think this is kind of where we can go back to some of our older, uh, things in terms of grabbing the data, it's outputs, doing trends, doing predictive sort of scoring to say, if this agent is like really the last time it said yes, uh, it didn't really, right? Like the action really didn't happen. And so we'll probably wall all those methods as well.
And that's kind of why the observability is kind of critical in this whole thing like observability and also closing the chain, if you may. Uh, in the human case, we did that before, right? When you actually sign up with a, let's say a dealer, they don't deliver that is there, record it in your CRM or other systems and you can figure out, okay, this person is unreliable and then maybe we won't give the next deal to them.
We will evolve certain systems like that for agent things as well, Right? I think those are employees are usually our relatives, so that's why we know not to trust them. Right?
True, true. Um, I'm trying to figure out though, how loosely or tightly coupled are agents and LLMs gonna be? Am I gonna have a scenario where I'm gonna have an AI agent that will invoke different LLMs based on the task and it will, uh, shop the LLMs as it were?
Or am I gonna have a scenario where there's gonna be, I don't know, four, uh, agents that are all capable of doing the same task and one might do it better or less expensively than another and they'll compete for the privilege? That's a great point. I think my, my, the way I'm seeing happen is that I think we are finding that certain LLMs are better for certain tasks.
We've seen that, right? And so agents that are specializing in certain tasks is gonna go optimized for that LLM because it's easier to say, oh, I'll just switch LLMs. But really the results are very different.
We have to go make sure the same prompts will get you different answers across the L LMS and so on. So my, what I'm seeing happen in my gut field tells me we will end up specializing agents based on the L lms and then that way they'll compete to say, okay, this LLM has all this functionality meets cheaper. Maybe it doesn't give you everything.
Whereas this LLM gives you more functionality, but it's more expensive and you'll have fine tuned agents, which may cost more. Okay. This agent might give you the best answer possible for your code writing scenario.
It may cost you more. This LLM may be good enough for simple code writing that you may not care, right? Uh, so I think that's probably where we will evolve.
But I think an LLM switching, uh, thinks at an action level, yes, but at a planner level, it's very hard, uh, because you gotta fine tune it to make sure it works for that particular LLM. Certainly at an action level, yes, it can say, okay, this action can be done by Claude, this action can be done by open ai. This action can be done by Gemini.
That's possible. But dynamically flipping at a planner level, I think it's a little more farfetched. Uh, we will end up more with specialized agents, is my gut feel.
Hmm. All these agents will need to be integrated. Does that require some sort of new dedicated platform to achieve that?
Or are we just gonna extend our existing integration platforms that we're already using to access data and APIs? And it's just really a matter of making sure that the right data shows up at the right place at the right time. That's right.
I think my, what we will end up happening is expanding this. So think about, think about this, right? Every agent, just like every human needs the right data.
And what is the data? It's all about your customers. Whether it's your sales data, who I talked to, who I didn't talk to, who came on your website, or it's service data.
Like, okay, what cases I've had and did I solve your problem or not, or early on in the pipe, it's your marketing funnel to say, okay, which person, uh, that I should be targeting or not, right? And so all that information is critical. So I know it's the same Mike who came to my website, didn't purchase or made a purchase, didn't like put it in their cart, et cetera, or had the sales call, had the service incident unhappy or happy.
All of those things is really what forms the context for what you think about a customer is. And then you have the memory as an agent to like all the conversations you might have had had with the agents. And this is an important distinction between an agent, a human, because you may talk to different humans, it really depends on what the human then types into your CR mothers to know, what's the context of what's the conversation you had?
And most of us don't want to type in everything, whereas agents can actually preserve all that history very easily across agents too. So bringing those two together with your profile, with your memory of the agent, I think you can really understand what the customer is doing with your business. That I think is a very profound thing.
And I think that's gonna where we as I think as Salesforce feel very proud that we have probably some of the best data on there to be able to go represent. And so it's an extension, if you may, of the platform that used to serve humans to be able to go serve agents. And part of this is also connecting to the rest of the enterprise, which is why I said APIs and other connectivities can critical to be able to go take actions, uh, onto the different enterprise, uh, scenarios.
So those two is going to be an extension of what we've already done for humans into the agent era. But there is gonna be new things coming in and new things coming in is where I, like I said, we need different planners. We need more, uh, determinism within non-determinism.
What do I mean by that? Like, if you just give questions that people are asking to an LLM, we found there's a lot of issues. Meaning if you give more than eight instructions to an LLM, it starts hallucinating.
Or if you give more than a hundred topics to an LLM, it won't work. It doesn't know what to do. So, uh, you may wanna have that determinist team to say, okay, somebody should have given you an order ID before we can actually go call this method.
Or if somebody is like an important VIP customer, make sure you don't keep bugging them with more questions. Go straight to the human escalation as the case may be. All of these needs a little bit of determinism, and that's kind of where some of these innovations are happening, where we can bring in determinism within the non-deterministic sort of LLM sphere, but grounded in the right context and memory so that you can actually give the right answers.
Mm-hmm. We start to hear the phrase context engineering more and more. And it seems to me that this is what that art form is.
It's getting the data, which is, and the context that's wrapped around that data, right. And the metadata to the right place so that there's enough of that information Yeah. For the AI agent to come up with some sort of reasonable output.
But on the other end of it too, I think maybe you don't want to give it too much data because then you wind up getting a lot of extraneous output. So That's right. Is there an art to this thing?
It's a very good question, and I'll give the example with our own, um, customer success story, right? 5 million questions from customers have been answered by our agent and it looks simple like, Joe, just feed all the documents to it, you'll answer. Uh, but this is kind of where the context engineering becomes important, because we did that same thing too.
First we just said, okay, let's give it all the documents. Turns out we found something interesting, which is you need to constantly also look at what is happening with your agent. What are people asking?
And is the agent answering correctly? Like it started off first, I'll give one simple example where somebody asked us to compare our agents with a competitor's agents, and the agent actually answered it, right? Like saying, oh yeah, we can do this, we can do that.
Uh, the other company can't do this. And so on. We were like, oh no, that we should not be doing it.
And so we put a rule that said, Hey, don't, uh, don't talk anything about, uh, other companies. Sounds simple, except next week we found out people are asking, saying, Hey, how do I integrate your agent or your system Salesforce with another customer, like another company like X, Y, Z? The system said, sorry, I cannot talk about it, right?
I'd be like, no, no, no, no. That's an important one to actually support. So then you had to fine tune the instruction to say, you know what?
You should be talking about integrations with other companies, but don't try to answer competitive kind of questions, right? Things like that. This is one simple example, but uh, it's a powerful thing to kind of say that you need to be looking at what your agent is doing and evolve.
The evolve can be you need to fine tune the instructions or fine tune your data, because a lot of times it could be missing data, they're asking a question, you don't know the answer to it, that's where you go add more sources or it may not be answering it correctly. You go fine tune the instructions to it. And that's kinda where we are.
We have added a lot of tools, including tableau sort of tools, et cetera, to go analyze it across all of your agents to say, okay, what are the topics it's answering correctly or not answering correctly? Uh, and what the remedies that you can do, um, at a fine-grain level. So that's is really what context engineering.
Context engineering is really about like making sure your rag is right, uh, making sure your pipelines are right, making sure your in instructions are correct to answer the right thing, and also making sure you are bringing all that right data together. Like you said, not garbage data, but all the right data together, structured and unstructured, uh, for the agents to work correctly. Hmm.
Most battle plans are excellent until first contact with the enemy. So I'm gonna give you this scenario and see how it plays out. So Salesforce widely used by salespeople, and they will use that to come up with offers and things like that, that they will send out to folks.
But on the other end grid, won't there be a purchasing platform somewhere that has AI agents that will be acting on behalf of the buyers, and the two of these sets of AI agents are gonna somehow or other come together in in some external ether somewhere and mm-hmm. What to each other won't. 'cause isn't there a chance they'll just cancel each other out?
Wow. But in some ways, if you really step back and think if the AI agents act just like humans, their goal should be to go maximize, right? To maximize whatever they're bid for.
Whether as a customer you're trying to go talk and get your tasks done on the other side, trying to maximize the dollar potential from from that other side. I feel we will evolve on that one. When you have these AI agents, they will all be goal-based AI agents.
And we've seen that already. Like if you look at our SDR agent, by the way, we now run SDR agents on our website. They're actually pretty good now.
Their, their goal is like, how do you make sure the customer will set up a meeting, uh, or the customer will actually go to your property and go buy it? Uh, I mean go go to the uh, uh, digital online store and go buy it. And what we are seeing is pretty fascinating already.
Like these are, um, prospects that we would've never had, nobody would've picked up the phone to call them because these are like long tail prospects that just visit our website. But now letting our agents, because a goal is to go do this thing, it's able to actually talk to them, convince them, can you do the same? Now you're correct.
On the other side, you might actually start to, to get agents, right, instead of humans, in which case I, my, the thing that it'll evolve is their both will start negotiating understanding. And actually we've seen some, uh, recent, I think, uh, there was some recent things where when both sides understood those agents actually downshifted from talking English to their own language super fast too. Uh, maybe we'll get that too.
Say, Hey, you know what? I know what you're doing. I, you know what you're doing, do this.
But I think in the end, I would say it still comes back to the goals that the agents are trained to. Uh, and they will go evolve and be flexible and they'll also be, all the ambient data will be available for them as well to be able to make the right calls. Um, I think that's kinda where we will end up with And, and all the code will be written in Assembler because they'll figure out that that's the most appeal.
Exactly. Maybe each will try to hack the other, who knows. Um, when you put all that together, what is your best advice to folks?
'cause I think everybody's out there trying to experiment with various things and, and, but getting that over the goal line into something that feels like it, I can run it in production, feels like it's, it's a maybe a, a little heavier lift than there ready for, but what do we do to get there? Yeah, great question. See, that's why if you look at the MIT study that came out right where 95% of the agent, uh, AI agent thinks fail because I think people are just looking at demos and saying, oh, I'll just so throw some things and things will just work.
Uh, that's great for a demo. But then when you put into real life practice, you need to have the solid foundation. Uh, and like I said earlier, the three foundations are one, it starts with making sure you know what data you need first.
Start with the scenario. I think that's the most important thing. Don't try to boil all the ocean.
Start with the scenario. Scenario could be something simple. It could be for your employees, your service, your sales, whatever it may be.
Um, once you know that, then figure out what data you need to go accomplish the task, what APIs or um, agents or APIs that you need access to, to go make that task and then create that agent with the right planning, with the right instructions to say what to do. But that's just step one. The step two is make sure once you put that agent in pilot or in in beta or or even production, start monitoring it with all the right tools, with the observe observability so that you can keep fine tuning it.
It's an art. Um, and then once you have that success, then it's easy for you to build upon it and then go create the next scenario and the third scenario and so on. Uh, never assume that just because you've created one agent, it's all fine.
I think, I think it, that is the part where I think people, people are missing their thing to say, you need to continuously observe it and fine tune it, uh, as you learn from it. All right, folks, you heard it here. Hey, has that great AI leader, Benjamin Franklin once said, you know, failing the plan is planning to fail.
Still true in the age of ai. Yeah, that's great. Thanks for being on the show.
Thanks, Mike. That was great. All right.
And thank you all for watching the latest episode of the Techstrong AI Leadership Insight series. Can find this episode and others on our website. We invite you to check all those out.
Until then, we'll see you next time.