Tracking and Ranking AI Agents with Galileo’s Atin Sanyal
In this Techstrong.ai video, Galileo CTO Atin Sanyal dives into why the capabilities of artificial intelligence (AI) agents will need to be continuously tracked and ranked.
Transcript
Hello and welcome to the latest edition of the Textron AI video series. I'm your host, Mike Zern. Today we're with a tins who's CTO for Galileo.
And we're talking about a leaderboard for AI agents that they've created. And well, a lot of these technologies are innovating so fast that everything seems to leapfrog each other on the leaderboard, but so no one's quite sure who's in best in what at any given moment, but at 10 you guys are trying to figure that out. How do you determine that?
How do you track that? I mean, and, and for that matter, how did you guys get in this business? Yeah, so, um, so for those who don't know, uh, I'm Atan, I'm one of the founders of Galileo.
And the reason we got into this space 'cause was 'cause, uh, from my personal experience, uh, I kind of spent much of my career in the, in, during the growth phase of Uber and very early AI at Apple. And we saw evaluations being a very cumbersome task and kind of the primary reason why AI systems fail in production is because of lack of robust evaluations. So we kind of went down this path and, uh, ended up, uh, choosing language modeling as our sort of strength and, uh, forward down down that path.
And, uh, for us, uh, of course, you know, the whole industry is kind of caught on and, uh, started using LLMs for industry use cases. And one of our goals was always to focus on benchmarks, uh, for a lot of these LLMs. 'cause we saw an explosion of, uh, of LLMs being a real thing in the, you know, first couple of years of, of, um, you know, the, the chat g pity moment.
Um, and, uh, we knew that, uh, often in ai people focus on academic benchmarks, you know, just to show strengths and weaknesses of models and they don't really do good justice to industry use cases. And, uh, that was one of the reasons why we built, uh, what's called the hallucination index. And we published a couple of versions of those in, uh, the months prior.
And more recently because agents and agent workflows have become, uh, very standard and, uh, it's still early days, but people have started adopting building agent applications. We sought to build an agent leaderboard, which kind of uses a lot of the evaluation techniques we used for the hallucination workflows and rag workflows when we published our hallucination index in that we essentially curate, uh, high quality datasets across multiple industry domains. In the case of agentic leaderboard, we looked at over 25 common industry use cases and domains, and then really used our own metrics and scorers, which we publish as part of the Galileo platform, kind of like dog feeding the Galileo platform to see which models, uh, perform well for certain kinds of agent tasks and don't for others.
'cause it's always this game of, oh, which model should I use for my application? Uh, one of the first questions any practitioner would ask is that, and our goal has always been to kind of publish industry benchmarks and kind of show how these models perform, but in the light of more practical industrial use cases, and a lot of the insights which are, uh, pretty surprising in the agent leaderboard, they kind of point to, uh, the fact that, you know, a model can do great in an academic benchmark, but the, the rankings can be completely different when it comes to the industry. How much difference is there in capability for these various AI agents versus what it might cost to implement them?
It seems like if I read the report, um, there's not a huge gap between them, but the pricing is kind of fundamentally different. Actually, that was one of the most surprising, uh, if I were to rank the three top three surprises from the leaderboard. One was this, uh, the this, uh, uh, sort of tussle between cost and performance.
'cause amongst the top three ranked models, which included Gemini and uh, as well as OpenAI models, uh, the performance difference between the best and the worst in the top three was 4%, which is minimal marginal, uh, and the cost difference is 20 x. So, you know, as an industry practitioner, if you really care about cost, you're kind of in luck. 'cause you might find a model that performs equally well for your use case, uh, as any other at one 20th, the cost.
How much are people kinda evaluating an AI agent today initially and going with one? And then are they gonna swap those out over time based on pricing or is this kinda like, I make a decision once and I stick with it? It Is more the former where you would kind of expect, uh, a developer or you know, a maintainer of an agent system to really go with what's best.
'cause on one hand there's this massive reduction of the cost of intelligence where a small model, uh, 100th the size of say a super large, very intelligent reasoning model can do as well, if not better for their use case. And uh, that combined with the fact that every day there's newer models, newer tools, newer frameworks, which are emerging, it's only natural for you to kind of do a, you know, a cost benefit analysis. And, um, uh, that's why a lot of, uh, on the tooling side of things, uh, many companies are kind of building these AB testing tools internally as well as using from open source, which kind of allows them to almost swap out models, uh, in their production applications, uh, and quickly do a test of, you know, if I swap out andro with something else, you know, what's the variation in performance will my system incur?
Um, and and that's kind of part of the story that Galileo also intends to solve where we wanna provide a platform and tooling to be able to do that very efficiently and do that in real time. Yeah, that was my next question. How dynamic can this be?
'cause can I have a scenario where LLMA is what I use in the morning and LLMB is what I'm invoking in the afternoon because some of the pricing structures changed on you in on that LLM per se. I mean, just how, uh, disposable are these models gonna wind up being? That's a great question.
Uh, I do think that there's a lot of variance, um, kind of in, in the era in which we are in, um, uh, but that, that variance will sort of go down as the tooling, uh, matures. And there's um, there's, there's minimal difference between, uh, different LLMs but also in terms of cost. Um, but also I think, uh, one prediction that I have is, um, in probably the next 12 to 15 months, what we'll see is this clear, uh, segregation between, uh, hey, here's the class of reasoning models, which, uh, are expensive, but they solve one particular kind of task very well, whether it's, you know, deep research or, uh, you know, things that require you to think a lot and reason about a lot.
And they are very expensive for a reason because behind the scenes there's, you know, if you go to the technicalities of it, there's multiple forward passes which are happening. And uh, the amount of GPU operations, it's really like the cost is kind of driven by the GPU operations. It's a lot more for those kind of models versus the second class is, hey, you're building some sort of an industry application, whether it's a chat bot or a summarizer or q and a system.
Um, there's a, a list of maybe a dozen or two practical use cases and there's a swath of much cheaper models which help you achieve that. Um, uh, that, that's kind of on the LLM side. Uh, one of the other things is with agents, it's not just the LLM, it's different functions are in the mix now.
There's different tools and uh, the, the ecosystem has become a little more complex and spirited with a myriad of different operations. So that also kind of causes additional variance where, you know, you might have stability in the LLMs and you're like, alright, this particular model Gemini, for example, is great for my agent application, but today I'm using auto gen for my tool orchestration and tomorrow maybe land graph or something else might, you know, just perform more deterministically for my use case. I, I do see for the foreseeable future that as the tooling and tool ecosystem matures, there will be this, uh, challenge of a lot of variability in the different components.
'cause everything's now predictive and, uh, sometimes something may work for one use case and may not work for another use case. So what you really want to build is like this system that allows you to multiplex between different models, between different frameworks and uh, that's kind of your, that would be my advice to anyone who's in the early days of building agent systems. Do you think at some point we may even have, um, I don't know, an AI agent or AI model that is optimized to figure out all that multiplexing logic?
Absolutely. In fact, we are already seeing, um, this at least in the private, uh, private models ecosystem with, with OpenAI Swarm and various others who have built in function calling as well as tool selection capabilities. Uh, the, the truth is that there will be innovation on both fronts.
Uh, like the large model providers are also in the game of not only building great systems for academic purposes, but also, uh, they're big businesses now and they, they're trying to cater to industry, but there will also be open source models and, uh, open source tooling which will kind of, you know, race one-to-one with, uh, a lot of the innovation that OpenAI does. Uh, speaking of open source, one of the other great findings from the agent leaderboard was that, uh, Mistral is probably the top open source model, which came out in our rankings with a very high tool selection score. So if you're a fan of open source, Mistral is certainly a great starting point for you.
Do you think we'll see some segmentation along the lines of, I may use open source for most of my functions, but then I'm gonna use proprietary models for specialized purposes that maybe they're better able to reason across? I mean, will there be segmentation kind of like the difference between, I don't know, buying a Honda and a Ferrari? That's a good point.
Um, I would say that the jury is still out on, uh, whether that would be the sort of segregation factor between open source and proprietary Deep seek is a great example of that where it was able to compete and outcompete on certain benchmarks. Oh one a level models by, uh, by OpenAI. So you might as well, you know, not use OpenAI and use deeps seek or try out deeps seek flavor of models for your use case.
And that is open source. And uh, uh, so, so there is this push towards having open source reasoning models, which are equally powerful, if not more. The race, I think is in optimizing the costs.
'cause the lower the costs go, the more available it is for, for usage and people will only pay for, uh, usage, whether it's open source or whether it's, uh, using an OpenAI model or any proprietary model. Um, it'll only be worth the dime if they're able to solve industry specific applications, uh, for the rest of us who are tinkering around and just doing, you know, simple q and a. Of course there is that, you know, uh, reason to pay for just directly using these models, but real like that, the real monetary value kind of comes from industry.
What's that one thing you see people doing over and over again that kind of makes you shake your head a little bit and go, folks, we need to be a little smarter about this than that. It's really how people are still playing whack-a-mole with the different models and different frameworks. Uh, there's almost tool fatigue and tool noise at this point around which model or which framework to use for your specific application.
Uh, I, I think the right way to go about in, in an environment where there's so much options to choose from is to, uh, really build a system, whether it's an internal tool or, you know, go for a, some kind of a, a platform which allows you to multiplex between the different choices and quickly get to the decision of that this is the combination that works for my summarization use case or some other, uh, q and a system that I'm building. Uh, Galileo, just to throw a shameless plug in there, is kind of one of those systems which allows you to make these decisions pretty quickly. Uh, but then that, that's half the story.
Uh, then because once you build your application with say, you know, the ideal combination of LMS and prompts and frameworks that is bound to change, uh, based on the realtime performance that you would get in production. So you need another system that allows you to monitor that and, uh, make those decisions on the production side of things. Uh, and there I think the real differentiator is your production application will constantly have newer forms of data to deal with, whether it's new questions or new kinds of ways of querying your application that will likely determine the performance.
So you need a way to measure it and then kind of come back to the drawing board and, um, redo your entire experiment that, Hey, here's a new kind of data that my chat bot received clearly, you know, frameworks A, B, and C are not working for me. Here's a quick test that I'm gonna do with d, e, and F and throw that back into the production. And you need a system that allows you to do this in with minimal downtime.
Folks, you heard in here to quote one great philosopher, don't worry, be happy because well, the cost of switching between AI agents and LLMs is not nearly as high as you thought. Hey, Adam, thanks for being on the show. Thank you so much, Michael.
It was a pleasure. All right, and thank you for all watching the latest episode of the Techstrong AI video series. You can find this episode and others on our website.
We invite you to check them all out. Until then, we'll see you next time.