The Rise of Domain-Specific AI with Digitate’s Efrain Ruh
Domain-specific large language models are reshaping AI by offering more control and cost-efficiency than general models. Specialized agents work in distributed systems to tackle challenges like LLM deployment skill gaps. The focus is on evaluating automation needs, setting KPIs, and understanding costs to ensure strong ROI.
Transcript
ai video series. I'm your host, Mike Bazaar today with Efrin Ru, who is CTO for Europe for Digitate. And we're gonna be talking about, well, the rise of domain specific large language models, which are a little bit different than your general purpose ones.
Re welcome to the show. Thank you. Uh, uh, nice, uh, nice to be here with you, Mike.
Pretty excited for today. I think we've all seen or experienced by now in some form a large language model. Um, but there appear to be general purpose.
They're jacks of all trade and maybe not optimized for specific tasks. And now we're talking about the rise of Agen ai. So my question to you is, are we about to see a whole new generation of these things that are really optimized for specific domains?
Yes. Um, and definitely Mark. Um, we see a big trend, uh, as you know, they, these generic models are being basically, they're getting better and better over time, but there are specific constraints.
Um, then, um, like, um, enterprises looking a way of embedded their own basic classic knowledge of their industry. They also want to have a little more control about these models. And there's another important factor is to make it more cost effective, uh, to make the business case.
So those two, three parameters are actually driving these, basically the evolution of these domain specific LLMs. So to your point about the cost, are they gonna be smaller and more efficient and maybe less costly to train? 'cause I don't need as much data.
I mean, how do they become something more affordable? Well, um, if you think about the very large language models, I mean the biggest more parameters you have, basically, you have either or you host it. So, which basically would have a pretty constant cost around hosting or basically you taking care of you pro getting it provided by, you know, um, uh, companies like OpenAI.
And then, uh, basic, every time you make a call, you get in charge. So ideally it is to make this call, I mean, utilize, depending on the use case, you may be able to utilize a smaller model, which is more cost effective, has a smaller footprint, you're able control more than you're able to run by yourself and not depending on, on a third party vendor for run it for you. Is there any trade off to be made in terms of the reasoning engines that are embedded in those things?
'cause every time you hear from the large language model folks, they talk about their next generation model and it has more memory and consumes more power. Uh, and they say, well, that's needed to drive these reasoning capabilities, but can I still have the advanced reasoning capabilities with smaller models? Yeah, sure, sure, sure.
I mean, there's, uh, there are, uh, there's, uh, so many models available on the market today, and if you look at, uh, what the, the cloud vendors so today are providing, uh, companies like Microsoft Azure or AWS are giving you the possibility of selecting the model that better fits for your use case. And then, uh, so you're able to do, do a specific training and then, um, and then areas where you have to basically differentiate whether you're gonna go with more, let's say, a specific TAT training or you're gonna basically, um, bring on some, some rag retrieval augmented generation into your model. So, um, um, what we have right now, uh, in the market is it's a very large variety of options that, uh, companies can leverage.
And instead of basically going to this very large and very complex, um, um, um, LLMs, which are being, um, you know, offered in the market, yet We also hear about the rise of AI agents. So, um, will these agents essentially be able to navigate multiple domain specific LLMs to drive a process? Or will it be more like multiple agents talking to each other and each one is trained on a specific domain and Yeah, so Correct.
So we've been working on this AI business for quite some time. I mean, we, we've been in the market for the last 10 years. We've been started investing, uh, a lot on traditional AI like, uh, machine learning, and we move into gen ai, especially the it, we are working, especially with these AI agents to deliver value on our, already our expertise, which is the automation of ideal operations.
And basically, and it's more or less the same approach that, um, that the industry is taking. So you select certain agents to give them a very specific task. And for example, you will have an agent, which is very good on, on understanding data, then, then you might have another agent, which would basically be in tuned for more reasoning and decision making.
And then you would have another agents which are more focused on the execution, basically taking, uh, remote connections, collecting, uh, data, or taking even, even actions. And then all together in an architecture where you have a kind, an agent that helps orchestrate everything. So it is, it is like, um, like you were mentioning, the trend is to, uh, have multiple agents, which each a specific agent being like very good other specific domain or a specific task that at hand How disposable are these domain specific large language models is gonna become.
Because I can imagine a world where they're constantly being, um, updated. So will my agent kind of be loosely coupled to that LLM and then when something else comes along, I'll just swap out the API? Yeah, that's, that's true.
That's one of the things that, uh, uh, when we talking about this very distributed architecture where we have like, say multi-agent would allow us to basically swap sim certain agent, maybe in an area where they have an an, uh, AI agent, which is very good, for example, of, um, applying changes or fixes on an inter IT infrastructure, maybe have an another agent which is more efficient or more, more, uh, faster or even has more knowledge, then I should be able to swap them, basically maintaining that, uh, we are keeping the same API, um, communication between the agents so that, uh, that is already happening, is happens on, on these very distributed architectures where we have established, um, um, very standardized communication protocol within these agents. So then we able to make those changes when we, when required. 'cause remember these agents then some of these agents we may have to, uh, keep updating them or, or, or, or doing some fine tuning on, on, on retraining.
And the whole ecosystem is just evolving so fast that the agents are getting more and more efficient, they're being, uh, requiring less computing power, uh, and even reducing costs. And then, uh, we have to keep these architectures flexible enough to be able to, to do this swap. Yeah.
How will we orchestrate all these AI agents then, and how will we kind of connect them to these various LLMs that they may be using? 'cause sometimes it may, may even need to access something that's a little less domain specific, but how fluid is all this gonna be? Well, um, it, it, it, it already depends on your use case or where you're gonna make this from.
And, and, and currently there are, um, a, a very large amount of different patents that been been deployed from everything. It could be a single engine that, uh, that is pretty good to basically take care of what you need to do or a complete multi-agent approach. Um, so it, it really depends of, of what kinda kind of use case, uh, you're gonna make.
If, if we're gonna make, we make, for example, we, like I was talking about in, in the IT environment where we trying to make automation across operations, not making agents autonom only, uh, take care of it faults or, uh, outages and downtime. We need to have those, those agents which are basically, um, able to take, um, these remote connections, uh, and take actions in the IT systems that we have deployed in the environment. Mm-hmm.
So as we kind of put all this together, um, who has the understanding of the processes that's gonna be required to stitch all these agents together to automate something on an end-to-end basis? Well, that's, that's a very good question. So one of the cha big challenges, uh, right now in the industry is, uh, the lack of skills.
Uh, we have, um, lms as you know, they're being out in market, uh, I think less than two years. We think about chat. GBT was launched, uh, just by the end of 2022.
Um, so there is a lack of, let's say, expertise, uh, in, in, in, in our IT teams, the people that have to, to deploy them. So we have to make, uh, very strong investments on our, our teams. Um, we need to let the people, um, let's say try and experiment with, with different models.
And that's where the, the, the hypervisors, like I said, you know, uh, like, uh, Azure and, and AWS already provide, um, you know, like a safe, uh, sandboxes to, uh, to start, um, um, improvising or, or start, uh, innovating with this, with these lms. So it's, it's definitely something that, uh, the companies need to invest internally. Of course, they can get help from, um, external providers, you know, very large, uh, companies or, you know, the typical sits to help them accelerate this, this, uh, this process.
But, um, um, definitely it's a skill that needs to be grown up, uh, within the, the enterprise. Um, will, I need multiple AI agents to be trained on the same domain, specific LLM? Because one, I need an AI agent to kind of check the work of the other AI agent.
'cause all these things are somewhat probabilistic and I'm trying to maybe use them in a process that's more deterministic. So I need it to be done the same way every time, but these LMS don't do anything the same way every time. Yeah.
So that's, that's what, um, we need to, um, bring, um, another capabilities. I mean, one is of course, uh, the, the correct guardrails and to make sure that, um, you know, our answers are within, um, the, the scope. Uh, that's also very important.
That is also the idea of bringing, um, things like, uh, uh, I mean, it's very standard, uh, reasoning into the system. Um, and then, um, and, and the knowledge, you know, think one of the big challenges companies has is said, how do, how do I collect all this knowledge? How to make it normalize it?
How do I, I treat this data, how do I clean it, make sure the data is correct for, for training, um, do my my model training, and then, uh, then what happens when this data becomes obsolete? How do I get it updated? Et cetera, et cetera, et cetera.
So that's, that's I think one of the biggest challenges on, on, on tuning this, these, uh, models for, uh, specific, specific domain tasks. And that's what I, I kind of make emphasis. I think we have to keep these models relatively small and simple, uh, not a lot of complexity.
And, uh, to divide, you know, it's very, when we talk about a specific domain, um, we need to basically constrain it so that it doesn't become a very complex problem, uh, you know, this process of training and, and, and keeping up to date. So I'm with you. Um, so what's your best advice to folks about how to go about doing all this?
'cause it does feel like it, it's gonna take a journey, it's gonna require some expertise that has to be aggregated from all across the organization, but I think a lot of folks kind of look at this and it's a little overwhelming. So how do I get started? Yeah, yeah, definitely.
So especially on, on, on automation. So we need to make a first, a good assessment when it makes sense to automate, because, uh, um, it's like you said, it's, it's a lot of effort. And then, uh, you know, the cost of implementations can, can really explode.
Uh, sometimes it's very hard to, to estimate the overall, you know, TCO of this kinda implementation. So one thing is just to make sure that we solving a real problem and we are able to, to, um, estimate very clearly what is gonna be the impact. I mean, how much, how much time my team is best spending on this, how important it's for me to reduce those effort.
How important for me, for example, to accelerate, I mean, there's, there's two things. One is the effort I'm able to reduce, which will be money, um, but also how fast I'm actually getting a response. So that actually give me an edge in the business area.
For example, when I'm, um, I'm quicker on giving you, uh, the best quote, or I'm quicker in giving you the best solution because I have a complete, let's say, a objective solution for, you know, responding, for example, from A to an RSP. Um, so that will give me a, a good ad advantage. So I think it's best to make sure that we pick up the, the correct use cases.
Uh, we define, um, very clear KPIs that we wanna track, uh, to make sure that, uh, we're heating, we're hitting those numbers. Then of course, uh, keep an eye on, on, you know, the, you know, the, the, the overall distance case around it. Well, folks, you heard it here.
There's a lot of math involved in AI and LLMs, and it doesn't all have to do with science. It has to do with the return on investment. So look twice before you jump.
Hey, ent, thanks for being on the show. Thank you. Thanks.
Uh, it was, uh, a real pleasure to be with you here, Mike. Uh, thanks very much and I hope to we see you again. All right.
Thank you all for watching the latest episode of the Techstrong AI series. You can find this and others on our website. We invite you to check them all out.
Until then, we'll see you next time.