09×04: Bringing Agentic AI to Production with Articul8
Most generative AI work has been fairly general purpose to date, but it is far more effective to develop expert models focused on specific industries. This episode of Utilizing Tech features Arun Subramaniyan, founder and CEO of Articul8 AI, in conversation with Guy Currier and Stephen Foskett. According to Arun, an agent is a model with a set of tools plus data that has agency to be called and interact with other agents. If these agents are domain-specific they can perform tasks more effectively than general purpose agents at certain points in this chain. Agentic AI is able to accomplish tasks previously thought impossible, and these systems keep improving. But people remain responsible for using and managing these systems. Costs can rise significantly if AI is used improper, but it is possible to deploy it profitably. Companies that can combine domain expertise with a novel AI-powered application are breaking free from the pack.
Transcript
Most generative AI work has been fairly general purpose to date, but it's far more effective to develop expert models focused on specific industries. This episode of utilizing tech features Arun Subin, founder and CEO of articulate in conversation with Guy Courier and myself as we talk about specific models and building AI applications in production. Welcome to Utilizing Tech, the podcast about emerging technology from Tech Field Day.
Part of the Futurum Group. This season focuses on practical applications of ag, agentic, ai, and other related innovations in artificial intelligence. I'm your host, Steven Foskett, organizer of the Tech Field Day event series, including AI Field Day, which is next week.
Joining me this week as my co-host is Mr. Guy Courier Guy. Welcome to the show.
Thanks, Steven. It's great to be back. So, guy, um, it's been fun having you here.
Uh, you bring a different perspective. Frederick and I have done quite a few seasons and, uh, episodes. Uh, this episode, uh, we're gonna be really zooming in on something that's kind of near and dear to I think the topic of the episode, which is, uh, or the season, which is practical applications for specific industry verticals and basically making sure that these things actually go and get into production.
Yeah, into production. So, um, that is where the proverbial rubber hits the road in application work. And I think a AI has AI development, AI capabilities, um, have accelerated so rapidly, um, even previous, uh, you know, cycles transform transformative cycles.
I think, you know, we sort of measured in years as, uh, the latest one, I guess being cloud as it ramped up and, you know, people got used to it, started doing useful things. Um, really in ai it's months because arguably agen AI is a sub transformation within the overall transformation. Um, the point being that, uh, there is 90% experimentation in 10% production.
I think at this, I'm probably even overstating it, it might be 99 and 1%, and some of that is due to all the experimentation and innovation going on in enterprises and industries all over, what can you use this for? Part of it is the sub transformation within the transformation of ai. Um, all of that represents change, new paradigms in, um, processing, um, data ingest, uh, uh, data storage in distribution, orchestration, all that other sort of stuff, the classic resource elements, uh, should have mentioned, networking, security, everything else.
Um, the classic, uh, uh, resource needs for, uh, applications going back since forever, but now whole new patterns in a rapidly changing, um, technology landscape that is AI and this case adjunct ai. So that is a major issue in taking something to production. And what does taking it into production means?
It means you need a certain certain level of reliability, security, availability, um, regional coverage, all of which are really important in industry, not to mention the enterprise. Um, so I, I, I love, uh, having this topic around agent ai, which is not the art of the possible stuff anymore, as much as let's get something out of it and figure out what's getting in the way of it and do something about that. Yeah.
And, and one of the things that I think is also holding it back is that a lot of the AI solutions out there are really generic. It's sort of like, you know, we made this foundational model, you know, it's a chat bot or it's a, you know, speech recognition or whatever it is. And it's useful for anything and everything, but that's not necessarily the right approach.
And that came out of a conversation that I had with our guest today, Arun Subramanian from Articulate. Arun and I were speaking about what they're doing, and he was quick to point out that sometimes the best way to bring something to market is to develop a focused solution for that specific industry. So, Arun, welcome to the show, Steven.
Gary, thank you so much for having me. Fantastic to be here. Um, as you mentioned, I'm, uh, Arun Subramanian, founder and CEO of Articulate.
We are building a domain specific genea platform focused on industries that typically have a very high barrier to entry, as you said, where genetic solutions, first of all don't work, and second can be pretty dangerous to be implemented there. I'll give you several examples of those as well, but it's not just about solutions, but also having platforms that cut across use cases that are specific to an industry, but then they're widely applicable across the industry. So you're not developing sort of hyper-focused solutions so much as industry specific solutions.
And as you said, I mean, I guess why bother, uh, training a, um, I don't know, a, a liquid flow analysis, ai, how to make HAI coups or how to cook, um, steak, you know, I, I guess it's, you know, you you, maybe it makes more sense to have something that's a lot more specific. Is that what you mean? Um, yes.
So like think of, uh, say fields like say manufacturing or, uh, aerospace or energy or even places like say cybersecurity. You may be able to understand some of those concepts. If you know language, like of course at the highest level, you can understand that.
But then in order to actually go implement anything there, you need experts. And anytime you need human experts, you should also think about the notion of would you hire somebody whose only knowledge is what they learned from the internet, or do they actually have some hands-on training? Have they gone to school for doing something?
Have they actually gotten battle tested in that particular field? The answer is pretty obvious. Most places you would want somebody who has the context, who's an expert in the field, us going to these general purpose models no matter how big they are or how, how proficient they seem to be in a lot of different aspects.
It's like going to somebody who's only learned from the internet and is able to only understand what is already on the internet, right? So that's the, that's the comparison bar that probably nobody would wanna go for a, a specialist who considers themselves a specialist from internet knowledge only. What's really interesting, um, when you, uh, apply that to AG agentic AI is, um, that agentic ai, I mean, it can, it can be based on just a single, the operations of a single model inference.
But, but as a general rule, the idea is for agenda guide to interact with numerous resources and other ai, um, and even, even grow in that regard, um, and create ad hoc connections, ad hoc connections to data, ad hoc connections, uh, to other, um, agents. And so that kind of specificity is not necessary, uh, let's say for every element of a workflow or that specificity can change over the course of a workflow. I'm just gonna focus on workflows, um, because I think that's a really general purpose concept across the gen ai.
So I think that that we, we talk a lot about tuning or context, um, whether for foundational models or, or for, or for more tuned or more specific models in agen ai, the problems associated, um, with, uh, AI that's not, uh, appropriately focused, I think would just multiply. So I think that makes, um, the application of the, of, of the technique you're talking about run even more important. Absolutely.
Right. So let me perhaps like, define what I mean by an agent, just so that we can leverage it, because there's slightly different, uh, definitions of that. Really for me, an agent is a combination of a model.
I'm making it even generic. It's a model, not necessarily LLM combined with a set of tools and with a set of optional data sets, let's put it that way. Whether the data is something that you go out to the internet to go do a web search, or it might be your own, uh, data sets in your enterprise.
And the tools are what are important because those not only specify what an LLM or a model cannot do, but you don't want the model to do, even if it doesn't, right. An example would be something as simple as find the average of a column of data. Like some LLMs can do it, some LLMs can do it better than others, but why?
Like, it's a, it, it's, it's things like that that you'd rather do with tools that are battle tested that you know, will get you the right answer every single time. And then combining that with a model actually gets you to do something. And the reason why it's an agent is it actually has agency meaning it can call things, it can decide, and most importantly, act on your behalf, right?
That's at least my simplistic definition of what an agent is. And then you make the agents work with each other or work with other tools around them, really what makes it into an agent system. And, um, if you think about it, that's pretty much how we solve all our problems.
Every system is a human driven system that has multiple people who have agency. Now you're augmenting that with systems that also have agency. Yeah, and I, I think that, well, I, I, I agree with your definition.
I think that it's important that people don't, I guess too tightly constrain what they think this is. But as you say, I think the important aspect is the fact that they have in a way specific uses for each agent in the chain, and that it's a chain of agents. And so they, they know how to look up, uh, other agents that can do things for them, and they know what the capabilities of those are and how to call them and how to specify things to them.
And that those agents then can take the output and, and, and kind of go to the next one in the chain. Uh, that's what you're saying. That's exactly what I'm saying.
And what makes these things domain specific are a combination, right? So you could, for example, make the model a domain specific model, meaning a model that understands energy really well, and lakes energy production and transmission can become an energy specific agent, which has access to tools that are both general as well as specific to that particular domain. For example, if you have a tool that predicts, uh, power demand, or how do you optimize the grid, and you mix that with a model that is domain specific, that becomes a highly domain specific agent.
But look, look at us. We're, we're, we're, we're talking again about, you know, sort of these concepts, because I think that's the nature of this market right now. Um, I wonder like if, if, if we focus on bringing, uh, agents sugen AI into production, though, um, can you maybe around from your perspective, um, put into, uh, light, um, the, the pers the perspective that I provided earlier, which is, uh, how critical it is, um, to, um, or what, what kind of challenges I should say around the complexity of these systems?
Um, what, what challenges there are to putting, uh, uh, agent AI into production? I can see integration and connectivity being part of, I can see data avail availability being part of, I could see a lot of things, but what are you seeing? I mean, the first step that you need to beat is secure, right?
So that's really why people get to agent systems because if a model could give you an answer, you've stop there. That itself is reasonably complex because you need to deploy GPUs. Your, uh, infrastructure becomes a little bit more complex.
Then you have to combine models with agents, like models with tools. So that becomes an agent. So you have to manage tools, you have to manage your models.
Now you have a group of agents working together, so you have to manage a group of things that themselves need management. So from a compute system perspective, it's a reasonably complicated environment. Uh, like there's a few, uh, systems out there that like give you the ability to manage agents, but they're all rudimentary today.
Nothing at enterprise scale, nothing that, uh, meets your enterprise quality standards. But even if you did, the real challenges that you face are things of scale. Meaning if, just imagine a company with 10,000 people, if you let agent take systems lose, you are talking about hundreds of thousands to millions of instances of agents running daily.
That's not a small amount of scale. And some of it might be doing silly things like, uh, managing email systems. Some of them might be actually doing production quality systems and, uh, things that are critical.
So you'll have to be able to manage that. But even if you did say, how do you manage the fact that this is reproducible? How do you manage the fact that this needs to be auditable?
How do you manage the fact that you have security at the data level? You have security at an application level. But here the question is just like how, um, I'll make a tongue, uh, cheek statement, just like when there is no spoon here, there is no apps, like the app becomes available on the fly.
It instantiates itself when you need it for whatever you need to do. If you wanna keep it along, you keep it along as an app, otherwise it disappears. So like when you think about security, you had data security, you have application security, you had enterprise security, the what if your applications themselves had FM metal?
So you have to think about security very differently. So the same way when people went from on-prem to cloud, there's a massive transition of security protocols and then shared security practices and all of that stuff. The security practice around agent systems also have to evolve.
So I assume that you have a suggestion, a proposal on how to handle this. Uh, absolutely problem. Uh, let's cut right into it.
Um, how do you think these issues should be tackled? So first and foremost, I'm an optimist, right? So all of these problems tell me that there are so many things to go invent and so many things to go solve.
That's one. But all of these things start because there is a real use case and there is a real value story at the end of it. Enterprises can legitimately do things today that they could not even dream about six months ago.
So that is the value that we need to cross, otherwise there's no point in putting all these additional costs, right? So that's number one. But the second thing that we, uh, come across, across every enterprise is the, the, the traditional question of build versus buy.
In my opinion, we no longer live in the world of build versus buy. We live in a world of build faster or build slower. There is no build versus buy.
Because first and foremost, everybody not just developers, every single person in every single company is now potentially possibly becoming a developer. Like you can ask something a natural language and it builds you an app. Now, it may not be production quality, it might be flaky, but if the app lives only for two hours, do you really care because you were able to build it, it was useful for you and you moved on.
And if you are living in that kinda world, you want to be able to enable everybody to build faster, but at the same time, you need to put guardrails, but you also need people to be able to build things that are reliable. And so it's the build slower versus build faster world that I'm much more interested in thinking about. And the build faster is really whether you partner with somebody like Articulate or you partner with, like you go to an AWS or you, uh, work with any of the other vendors, building with AI tools and AI vendors should be a no brainer.
And from a cost perspective, I think that's a easily a three to five to 10 times cheaper exercise when you think about a total cost of ownership perspective and time to market perspective, right? Because time today is money. And if you can get out there with your solutions grounded in your own enterprises data, your own enterprises know-how, add it with ai, that's really what is going to differentiate your business to somebody else's.
And that's what we call hyper-personalization. And that's, um, something that we're seeing like again and again play out in many industries. And I'll give you one more example and stop, uh, answering the, if you think about AI enablement across enterprises today, there might be a gradation, some companies that are ahead, some companies are behind, they're just getting started.
But AI for productivity in my opinion, is going to become table stakes. Just like how email is table stakes today, having internet or an intranet, uh, I mean, maybe I'll date myself, but uh, when I went into, into the workforce, having an intranet used to be an advantage, like no company today would consider that an advantage. Same way, having access to AI tools are becoming remarkably similar.
Whether you are with Google or whether you're with Anthropic or you're with Open AI or Microsoft, your capabilities for productivity are more or less going to be the same. So if that is the case, the tide has risen, how do you stand out? You have to hyper-personalized.
And to do that is really where you're building faster with the capabilities and partners, uh, that can bring to bear, um, I would say solutions and platforms that are domain specific in your own domain. I'll give you some examples as well, but I'll pause to see if you have any questions or comment. It sounds like that puts more of the, um, production of readiness burden on the platform though, Arun, because if the, you're, if you're saying the application is ephemeral, that means the, the, the stack is ephemeral in a sense, at least the upper part of the stack.
So, um, how, how do you ensure the question that comes to mind then is how you ensure that the production enables the That's right. The platform enables production readiness for an app that lasts two hours, two days or two years, that scales or doesn't scale, that's very individualistic or more, let's say more broadly individualistic, uh, you know, good for a department and organization and industry that that puts more pressure on, on the platform because the, the need for production readiness exists however long. The, I actually, it also lowers the barriers for of require or lowers the not barriers, lowers the requirement level, like you said, depending on the nature of the app, that's a pretty big change.
It is a big change, right? But then think of all of the changes we've had going all the way back to the printing process. So like, it's pretty much from an information standpoint, it's not that different, even though it may feel very different for us.
So think of it's much faster, but the, the, the unlock of capabilities that were not possible before is very, very similar, right? That's one. But the, the responsibility as you put it, is not just with the platform.
The responsibility is now shared and it's even more shared between the creators and the providers than it used to be before. 'cause before you could buy something off the shelf and then say you'll own the application stack, and then you'll build the application on top. To a large extent, the stack is not changing.
Like, I'll give you a specific example. When you're building an application with React, for example, you have, uh, a bunch of databases behind that you have to go connect your data with, maybe you even have a data warehouse. The difference is taking two years to standardize your data, having your data engineers and data scientists bring you materialized views and your business analysts and, uh, your application developers working together with your materialized views and data scientists to come up with an application that goes through review cycles and tightening that to having your first version of your application.
With all the things I mentioned in your first week, is it production quality? No. Is it, uh, something that you can put out there to test internally?
Absolutely, yes. I, I don't know if you can say though that any longer that it's not production quality, because production quality refers to adherence to a set of specifications. And yeah, sure.
Two years ago, even as recently as two years ago, those requirements could be really quite strict. Um, now I feel like it's really the scope of the application, the scope of of an agentic AI or use of agentic AI can, uh, reduce those requirements a good deal potentially even in a highly regulated enterprise environment. No.
Am I going too far? No, not at all. Because before, if you had to spend all of the time and resources to build an application, your standards were high and you wanted to make sure that they actually lasted.
So the return on your investment you're making is really what the calculation is today. If you're taking two weeks to do the whole thing, and if it lasts only for two months, you got more than the bank for a buck, right? And that's really where the delta comes in.
However, the production readiness and getting to, uh, say quality of use is important because the other side of the whole thing is what we have to acknowledge, which is AI swap. Like how many times you read an email and you go, huh, that was generated. Like, it, it just starting to, to get on everybody's nerves to a point where you're going to get to a state where you not only want productivity, you actually want quality.
I'll give you my own example. Before, if I had to write a one pager or a two pager for my team, I would've taken half an hour or one hour. I've researched, I've written something up, I go to late today, I still take an hour.
Sometimes it hour turns in two hours, not because I don't use ai, I actually use more than one tool all over the place, and my research goes much deeper. My analysis of how this piece of writing is going to get absorbed is much, much deeper because I had to guess before now I can actually do what if analysis. And sometimes depending on the time of the day, my what if analysis gets out ahead and it like one hour used to what it used to take, ends up taking two hours.
However, the quality of the content I produce is significantly more, right? And there's not a word that I would put on a piece of paper with my name on it that I haven't read and agreed to, right? So that's a quality level that individuals have to hold themselves accountable for.
Whether you're writing a piece of paper, you're writing code, you are generating a PowerPoint that needs to come through. And I think all of us will hold ourselves to that level of quality, at least I hope so, because like nobody wants to read an email that somebody didn't even write, right? Um, and, and those are kinds of things that, uh, as a society, I think we have to get better, but even as enterprises, everybody's going to demand more quality, right?
That's really where the, the rubber really hits the road. Because if you look at the kinds of applications we'd solve, it is not the give me 70% or 80%, or even 90% accuracy. Our customers are, like, for example, I'll, uh, name some of our public customers like Franklin Templeton, their accuracy starting point was 92%, and understanding tables above a 95% accuracy level was their criteria to meet.
So those are not kinds of things that you would do if you didn't have a model that understands financial records, that understands tables. These are what I would call the non cool aspects of putting something into production, but that's really what moves the needle in the top last 10% or even the last month. You know, I've been having a lot of conversations with companies in the ENT space about exactly what you're describing here.
And one of the things that was, um, articulated, if you'll forgive me, uh, for the pun, uh, to me about that was the sort of compounding nature of, of AI errors. Effectively, if you have a chain of agents, and if each of those agents are 90% accurate, yes, by the time you get to the bottom of that chain, you could be much, much, much less accurate. You're less than 10% accurate.
At that point, if you like change just six of those agents each with 90% accuracy, your final accuracy is less than 10%. Yeah. And and they also pointed out that, uh, this sort of paradox that even as, um, systems are getting more efficient and hardware is improving and the models are improving, and we have distilled models and tiny models and tuned models, um, we're using more and more and more tokens.
We're doing more and more processing because of these agents. Because once again, if you have, uh, a dozen different, um, you know, elements that are all talking one to the next, to the next, to the next, yes, uh, by the time you get to the end of that chain, uh, you have, uh, used way more than you might have if you had one big super AI thing. You know, I don't wanna say, I'm not gonna say general intelligence, if you had a big, a big model, you, you know, a a bunch of small ones.
So how do you answer that problem, the cascading errors and also the sort of compounding effort required? So Two ways to answer that, right? So one is to make sure that every response you do is grounded on actual data that has to be like a non-negotiable, uh, condition.
And if you're answering something that does not have a support with data, the user has to know at that point in time, right? And it's a, it's a decision that they have to consciously make. The second one is, after you've made a prediction, by the way, like this is something that, uh, I think even the, the practitioners of AI have forgotten that every prediction is uncertain.
Every prediction has an error bar around it. So we need to know what is the relevance of that prediction? And I'm using the word prediction very carefully here because all the outputs, just because it comes from an application or an agent, we have forgotten that is actually coming from a model, which means it's just a prediction.
It could be just as wrong as any other prediction. And having the ability to evaluate how good or bad that prediction is at every step helps you. And you also need to have an overall evaluation that's independent of the systems that are actually causing the prediction to happen in the first place, right?
And it's, it's those things where it's not just the amount of tokens that are getting generated. You need to have efficient compute to judge the tokens that are getting computed as. And so for example, if you take, um, take our platform, every prediction we make comes with the relevant score.
Now, we cannot make an accuracy score because we don't know the truth from the falsities. All we can say is, here's your dataset, here is the grounding on the dataset, and given your dataset, and given your question, is this relevant or not? Right?
The accuracy question is the next question that has to come from like say, validation data sets, right? That also we provide, if you're talking about models, but having the ability to know that I'm getting this answer and this answer has this relevance scope, even from the system that is actually doing the analysis is important. Because many times we forget that the tone of the response might suggest that it's very confident, but actually the prediction might be highly not confident.
Also, the nature of gen generative AI in particular to be confident because sort of appear confident. 'cause that's the training point, that's the training value point. You challenge it, and then it is, uh, it seems to agree to everything you say the number of times, um, like, um, I get told that I'm such a deep thinker and, uh, the, I found the mistake like gets on my nerves at least, Or just, well, I'm, I'm actually the deep thinker, not you, the, I dunno what You guys are talking about.
I'm the deep thinker, you know? Oh goodness. Well, so, so I guess what's the prognosis for all this?
Um, you know, what is your prescription, Dr. Arun, for companies that really want to deploy a system that is, uh, productive and accurate and reliable, uh, what should they do apart from, I don't know, talk to articulate, but what, what else should they do? So, uh, the first thing I would say is please be cautiously optimistic.
And, uh, it's important to be optimistic because these systems can legitimately do things that are impossible even six months ago. So questioning whether this is real or not is really past us. It's behind us.
These, these systems can actually help irrespective of which platform you use, which, so that's number one. But you also need to be cautious because these are very, very early days. The tech is evolving super fast.
The costs are changing significantly, and the caution really needs to be just like how you would drive, uh, an autonomous car. It's still level three, meaning you are still responsible for whatever happens in the car. You are still responsible for, uh, like, uh, any kind of mishaps, meaning you need to still pay attention, but not using it would be a significant handicap for you, right?
That's number one. The second thing is, if you're not careful, costs can significantly go outta hand. But if you're careful, your actual total cost would be much, much lower.
And you don't necessarily have to be a significant expert to go do that. Just need to be able to partner, uh, smartly with, uh, the right partners to do that. And that's the second one I would leave people with.
And the third one is, you are going to see companies break out and you're already seeing companies breaking out. Like we work with some of the thought leaders in the world. They have thought through the problems, they've hit their walls, they've hit their, like say boundaries, gotten help and gotten past them, and you can very clearly see them breaking away.
And the breakaway is almost always when they combine their own domain expertise with an application that nobody in their own domain could do. So the top of the line use cases versus bottom of the line use cases, right? My opinion, every enterprise is going to get more efficient.
That's really bottom of the line. If you wanna stand out, you have to go after use cases that increase your top level. And if you do, you are going to significantly break away from your competition.
And we've seen that in many industries. It's a fascinating discussion and, uh, eye-opening in, in the sense of what production means, what an application is as it relates to what an agent is. So I'm gonna have to get a little extra time afterwards to think about all of this.
I, I think that's a great summary because it shows that we, despite the sort of naysayers about ai, there are practical applications, there are ways that companies can leverage this technology, and as you say, that ways that these companies can break away from the pack. So before we go, um, tell us a little bit more Arun, um, where can we, uh, learn more about what you're doing? And, um, I guess I'll start by saying, Hey, how about tuning in at, uh, AI Field Day next week, uh, when y'all are presenting?
Absolutely, Steven. I look forward to seeing you there, uh, next Thursday. And, uh, that's a fantastic place to actually look at all the different actual practical applications.
We have some exciting demos there as well. Um, and, uh, we just had a, a huge refresher of our website. Uh, lots of new information, lots of new case studies we focused on, uh, like what people can actually do today to improve their, uh, enterprise, uh, focus and sharpen their skillsets.
So please check that out as well. And please hit us up on any of the social channels. Love to engage And Guy, uh, how about yourself?
Well, one place you'll find me is definitely at, uh, uh, online, um, watching and commenting on, uh, AI Field Day next week. Everyone looking forward to seeing you there. Um, and, uh, also actually I'm a delegated cloud Field day, which is this week.
So you'll see me, uh, on screen. com/in/guy coer, you can follow me there. com.
Thanks guy. And as for me, yes, uh, I will be, uh, virtually joining you with, with Cloud Field Day and, and joining in person for AI Field Day. So sort of the, the opposite of you.
Uh, and of course, I would encourage you all to tune in to AI Field Day as well. We're, we're gonna be launching a brand new AI podcast for the RUM Group, so that's gonna happen on Thursday with a special live session. So thank you very much for, uh, joining us and all of you listening, thank you for listening to this episode of the Utilizing Tech podcast series.
com. You can also, as guy mentioned, find us on socials LinkedIn X Twitter Boost, guy Mastodon at Utilizing Tech. Thanks for listening, and we will catch you next week.