Voice AI Needs Specialized Models, Not Just Bigger Ones
A Third Path Between Building and Buying AI
Voice AI agents need more than an impressive demo to deliver business value. They need fast responses, relevant context and ongoing training. In this Techstrong AI Leadership discussion, Phonely co-founder and CEO Will Bodewes explores those requirements with host Mike Vizard. Their conversation examines the choices behind practical enterprise AI adoption.
Open-weight models offer greater control over model behavior and customization. However, deploying and maintaining them requires expertise that many businesses struggle to recruit. Bodewes describes a third path between building everything internally and relying on general-purpose frontier models. Companies can work with providers that fine-tune specialized models for specific business tasks.
Response Speed Changes the Model Decision
Voice conversations create demands that differ from text-based interactions. A caller cannot comfortably wait while a model works through a lengthy reasoning process. Bodewes explains why latency makes fine-tuning and post-training particularly important for voice applications.
Smaller, task-focused models can also change the cost equation. Yet specialization does not eliminate the need to connect models, tools and business systems. The discussion considers orchestration and the role of Model Context Protocol connections in accessing other agents and enterprise data.
Bodewes also distinguishes conversational AI from traditional interactive voice response systems. Rather than forcing callers through fixed menus, these systems can use context to respond dynamically. That flexibility still requires careful training, testing and improvement. Businesses should not expect a new agent to handle every situation immediately.
Measure Growth Alongside Cost Savings
The strongest business case may extend beyond reducing the cost of existing calls. Bodewes highlights opportunities to answer customers around the clock, capture more inquiries and turn leads into booked appointments. Voice AI agents can support growth when teams evaluate them against clear business outcomes.
He encourages leaders to treat AI adoption as an ongoing investment rather than a one-time purchase. A polished demonstration is not a substitute for measuring performance against the metrics that matter.
The conversation also explores responsible deployment and the difference between inbound assistance and unsolicited outbound calls. Bodewes favors useful, customer-oriented applications while cautioning against AI cold calling. The takeaway is to match the model, workflow and safeguards to the task, then keep improving the system.
Transcript
ai Leadership Insight Series. I'm your host, Mike Bizant. And we're having a little chat about, well, there's a lot of interest all of a sudden in open source, open weight AI models, but, well, what comes with that is the interesting question.
Will, welcome to the show. Yeah. Super excited to be on, and this is one of my favorite things to talk about, so can't wait to get in.
So everybody's kind of trying to put their arms around the cost of AI, and then that naturally pushes people towards open source and open weight. But the trouble with open source has always been the same. It's kind of like getting a free puppy, then you've got to raise the thing and manage it, and keep it.
So, what's your sense of what's real here, what's not real, and maybe is there some sort of alternative between the two extremes that we're talking about today, open source and very expensive frontier models? Yeah, it's a great question, and I think that there's some argument to be made that even if you use the latest frontier models, it's still a bit like raising a puppy, if you've ever used ChatGPT or anything like that. But even still, the difference between the value of open source, open weight is that you have much more control than you would with a frontier lab.
So you can actually tweak the model and the underlying components of it, rather than having to rely on just prompt engineering as your only source of truth. So the pros are it can be cheaper, it's good for scope of business units, and you can use it to do things that weren't possible with the larger models. But the cons are you have to have somebody that does it.
And to your point about that, where do I get those people? Because it seems like they're all working for very large AI-oriented companies, and they're not exactly lined up to come work for me in my average little enterprise. Yeah.
It's definitely a tough thing. And I think also part of the hard part is it's constantly evolving. So that's one of the challenges of building any company is finding the right people.
The folks that we feel like are qualified, our team is just PhDs in machine learning and artificial intelligence, so they're hard to find, they're hard to come by. But if they're there, and the model is able to save you tens, hundreds, maybe even millions of dollars a month by specifically using an open source, open weight model and post-training it on your use case, it may be worth it to pay the salary. In the absence of the ability to deploy open source, open weight models and train them and maintain them, but I don't want to necessarily pay a small fortune to the frontier people.
Is there a third path? I guess everybody wants to know if there's some other way of thinking about this. Yeah.
Well, the third path is to use somebody that's already done that. And so what we've done is we've post-trained these models, actually, it's a series of models together, specifically for voice. And we offer that as a service for some of the businesses that want to use that.
And so we'll actually take your data and we'll run it through our fine-tuning pipeline and process, and then you can use your own model without you having to train it. And so the third option is, well, look for a smart group of people that have already trained this. And there are companies that are popping up that are doing this outside of just us.
They're training these models, and they're allowing folks on the enterprise side or businesses to solve their business problems without having to use these large frontier models. Yeah. Voice is one aspect of my application.
So am I going to wind up kind of stringing together a bunch of different models for different tasks, and that becomes my application experience? But then I got to orchestrate all this stuff. So how do I wrap my head around all that?
There's probably no other way, even if you're using voice in the traditional sense, out of orchestrating everything. The way that you'll save money on open source, open weight models is because you're training a smaller model. Which means that inherently, it's not as comprehensive as one of those larger models.
And also, when you're talking about voice versus other, voice has a different modality. And so regardless of what you're using on the back end, you will still have to have a completely separate instance for voice. Now, that being said, that voice agent can call the MCP server that is orchestrated with the rest of your data and the rest of your agents.
But there's no way around not having a separate instance specifically for voice. Now, when it comes to other use cases and other processes, that really depends on what you have set up. You've been at this for a fair amount now.
" Gosh, where do I start? There's I think a lot of probably an overestimation of how complicated a problem may be. And so I think a lot from a business perspective is certain businesses will take in a look at the economics of using AI and not evaluate it on the right tools and the right metrics.
" So that's what really good businesses do. A lot of businesses, they will evaluate it based on some demo, they'll evaluate it based on the way that some AI handles some response, and they don't understand that there's the ability and the capability to train and to improve agents. And so I think the biggest thing is when people are adopting AI in general, whether that's voice or other, you need to think about it as an investment, as just like you're training a person.
So people, if you want to augment what humans are doing, you still need to train and treat like a person, and understand that you're going to still have to make that investment over time, regardless of you're using a frontier lab or if you're using your own post-trained, fine-tuned model. My trouble with the whole ROI conversation is a lot of the things that people are building are nice, but they're not sustainable, competitive advantages. And I cannot help but wonder if we need to start thinking about AI as essentially table stakes.
And if I don't do it, well, the competitor down the road is going to do it anyway, and so we both kind of have to do this, but it may not be something that I can actually say that I am going to get a ROI from. Does that make sense? I think it really depends on your use case.
So there's two sides of this coin. I think what everyone has traditionally viewed AI as is it's a cost-saving tool, and it's designed to save businesses money. But there's certain businesses out there, and we're starting to see this as a new movement that's happening, is businesses that are using AI for growth and for getting more customers.
" And they will deploy voice agents to save X percentage. " Or, "I want to enable myself to get more phone calls so I can convert more of those leads into booked appointments, closed insurance sales," whatever that might be. And I think the difference is not whether or not you're adopting AI, but it's how you're adopting AI and how you're thinking about it.
And I think of the businesses that are doing it really well are thinking about how they can use AI to grow their business rather than just to save money on what they're already doing. Of course, anything with voice and everybody immediately goes back to their worst IVR experience you could possibly imagine. So, in the age of AI, how is that whole voice interaction different, better, and what can we actually expect?
Yeah. So first things first, you probably have already interacted with a voice AI agent, whether you know it or you haven't. And so the difference in an IVR versus a voice agent is voice agents use LLMs, which means that they can respond dynamically, and they have a lot more context.
And a lot of what the problems that we're solving with AI is just giving it the right context to be able to handle the situation that you might have. And so the difference between an IVR and an AI is that it can actually do things, and it can respond like a person, and it can react to your emotions, and it can respond in real time. And so I think that's the biggest difference between an IVR is the old way of the past, and the voice AI agent is the new way of the future.
Now, it can't do everything that a person can do, and sometimes there's a lot of businesses that will adopt them thinking that they can do everything right from the bat. But really what they do is they take training, and they take optimization, and they take coaching, but they can get there, and they can do a lot of really impressive stuff. At least that's what we're seeing with a lot of our businesses that use us.
There also seems to be something of a debate emerging around, well, what should I have the LLM do versus what should I maybe rely on some sort of context engine for? And I ask the question because it seems like the LLMs are getting better at breaking down a task into a bunch of subtasks that a bunch of AI agents can complete. Yeah.
But other folks would say, well, maybe more of that should be leaned into an AI skill that is executed by some context engine somewhere, because that will be cheaper, less expensive, and more accurate. Are those two things diametrically opposed, or is there a middle? I think there's a middle.
It's one of the things that is yet to be seen. I can talk more about for voice, the challenge is that you don't have the amount of time that it takes for us to think in real time. Got you.
And so that's the whole challenge that we face with voice. And part of the reason why fine-tuning and post-training a model is so critical is because you can't wait for these models to think through 10 different steps before they give you an answer. You have to respond just like that.
And so a big component of why would you even be fine-tuning is because latency is such a critical use case. Everywhere you look, there's people talking about regulation now for AI. What's reasonable?
What's maybe wishful political thinking? And how concerned should we be that swarms of AI agents are going to make a billion calls one day, and we won't know what to do with it all? Yeah.
I think it's a good question. My opinion for the voice space specifically is anything that's inbound, a customer is calling inbound, that should be allowed, so long as it's legal and whatever, because it's helping a business out, it's helping a customer out. And so from the consumer perspective, that should be hand-in-hand.
Now, the other side is outbound calls. And outbound calls, in my opinion, shouldn't really even be done by AI. It shouldn't be something that's allowed by artificial intelligence unless it's very specific use cases.
So when I'm talking about is I don't want people cold calling with AI because that's a terrible experience for everybody. And so what it should be is, oh, I have an appointment reminder that I want to get called back on, or specific use cases like that where they can be beneficial. Now, the fact around AI regulation with some of these bigger models as they get more and more advanced into really complicated stuff that people don't have control over, that's a tough question.
I wish I had all the answers to that, and I think that AI is evolving so fast that even the people in San Francisco, in Silicon Valley, are watching it unfold right in front of them and are reacting pretty much in real time along with the rest of the public. So I wish I had a better answer. Let me ask you a question about that.
Do people actually take calls from numbers that they don't know? It's been years since I take a call from a number I didn't know. So, what is the right balance there on the outbound anyway, because I think most people have been conditioned to if it's serious, they'll text me.
Yeah. And that's pretty much true anyways. But even still, it's annoying to get your phone buzzing right when you're in the middle of a meeting, no matter what's going on.
This is true. As you look at all this, voice is one interface, but will that become the dominant interface, do you think, as we go forward here? Because most people today, when they're interacting with a model, they're typing something into a chat window.
But at some point, are we all going to be using voice just to talk to any type of model in any type of workflow? There will always probably be a hybrid approach, but I think as AI models get better and better, especially voice, we'll lean more and more towards voice. And the simple proof in the pudding comes from, we've been talking for millions of years, or thousands of, I don't know, however long humans have been talking.
We've not been typing, right? Typing is a concept that came up in the last 100 years or so, and there's a reason why we're not emailing this and we're having a conversation. It's because we're people, and we're designed to listen, and we're designed to talk.
And I think that as means become more and more efficient for people to do so, we will lean harder and harder on using it. I believe it was the Secretary of the Department of Commerce said we needed more open-weight, open source models here in the US. If you could ask the United States government to do anything in particular, what would be at the top of your list?
What should be the role of the government in all of this? And, from your perspective, how can they help? " That's a very rare question.
I think there's a lot of opinions on this and a lot of people that are doing a ton of research on what's the right way to go about this and what's not. I think the important part is that we all have a very good understanding and alignment of what we're building towards. And I think the thing that sometimes is a bit scary to me is we have a couple of big competitors that are competing against each other in a speed race.
And while potentially everything ends up okay, there's also the risk that one person tries to get better than the other person, then they cut regulations, and they do things that they shouldn't do simply because they're trying to get ahead. And so I think the concerning part is, hey, there is an AI race that's happening right now, and how do we make that race-- It's okay if we slow down a little bit on it. And so I think that the government should be, when we talk to some of these more advanced use cases, should be viewing this as a level of understanding and a level of scrutiny to make sure that this is, like, the last thing that humanity will create is voice AI.
Not voice AI, it's AI in general. It's the biggest contribution to what will be next in our society. And I think it's very important that we do it correctly.
I wish I had all the answers for how that is, but I think that it is in somewhat everybody's responsibility to ensure that this thing that we introduce into the world is introduced in the correct way. All right. Well, folks, you heard it here.
There's clearly a lot at stake, and there's a lot of choices. And by the way, it's never going to be a black and white choice. It's going to be a lot of gray and a lot of balance and probably more of all of the above.
Hey, Will, thanks for being on the show. Yeah, thanks so much. Pleasure to be on.
All right. AI Leadership Insight series. You can find this episode and others on our website.
We invite you to check all those out. Until then, we'll see you next time.