AI Token Costs Fall While Enterprise AI Spending Climbs
### AI Cost Management Moves Into Focus
AI token costs are falling, but enterprise AI spending is still rising. In this Techstrong AI Leadership Insights conversation, Mike Vizard talks with Jon Knisley, Director of AI Value Management at ABBYY, about why lower token prices do not always translate into lower operating costs.
The issue is not just the price of a single model call. It is the larger system around AI. Agents, workflows, tools, integrations and repeated reasoning steps can consume far more tokens than a simple chatbot response. That changes the way teams need to measure value.
### AI Agents Change the Cost Equation
Knisley explains that AI agents can create a version of Jevons Paradox for enterprise technology teams. As each unit of AI becomes cheaper, organizations often use more of it. That can drive total spending higher, even when the cost per token keeps dropping.
This makes AI cost management more important for leaders who are trying to move beyond experiments. Teams need to understand when an AI agent is useful, when a simpler model is enough and when automation adds too much complexity. The goal is not to avoid AI. The goal is to apply it where the return is clear.
### Workflow Value Matters More Than Model Choice
The conversation also explores why many AI projects struggle to show business value. Knisley points to workflow design as a major factor. A powerful model may not improve outcomes if the process, data and measurement strategy are weak.
One example involved contract lease extraction. The organization needed hundreds of data points from thousands of complex documents. Accuracy problems were not solved by simply sending the work to a large language model. Better structure, process understanding and domain-specific design mattered more.
### Practical Takeaways for Enterprise AI Leaders
For technology leaders, the message is direct. AI token costs should be tracked, but they are only one part of the bigger picture. Teams should also measure utilization, accuracy, workflow impact and the cost of missed business opportunities.
AI cost management can help organizations decide where agents make sense and where lighter automation will do the job. As AI becomes part of everyday operations, the winners will be the teams that connect technical choices to measurable value.
Transcript
AI Leadership Insight Series. I'm your host, Mike Bizos. Today, we're with John Nicely, who is Director of AI Value Management for Epic.
And we're having a little chat about, well, AI cost, because there's an odd thing happening in the world. The actual cost of a token is dropping, depending on what the token is, and at the same time, however, the total cost of using AI is rising. It's a little bit of a paradox.
John, what's going on here? It's an interesting time, Mike, and thanks very much for having me on today. A lot of interesting sort of updates, news going on in the industry, but you're exactly right.
The cost of tokens is going way down. I think I saw the other day, Gartner predicts it's going to go down like 95% in the next two years, which is pretty remarkable. But at the same time, the costs are skyrocketing.
And it really just has to do with the tooling that's built around the AI and the models. There's so much more talk today about agents, and we're moving away from chatbots, things like that, that all that complication and complexity just adds to the cost and the number of tokens that get used. So I think it's an old economic principle from the sort of the steam age days, Jevons Paradox, they call it sometimes.
But the usage is going down, but cost and utilization ends up going higher. Another number I saw was for an agent to do a task, it's about 150 times more tokens get used than if the chatbot responds to essentially that same task. So just shows you as we get more of this technology, yes, the cost may go down, but ultimately, utilization and cost end up rising as well.
So we've got a lot more attention on this topic around token cost. So is this giving people cause for pause because they're kind of looking at all this stuff and saying, "Well, maybe I won't use $2 million worth of tokens to eliminate a $45,000-a-year job, and we're going to be smarter about what we're using this stuff for"? Yeah, I think there's a lot of movement in a whole bunch of different areas.
You now see occasionally the reports that the agent developer costs more than your human developer, which is remarkable. But I think it's a wider sort of macro event that's happening. You've got all this sort of interest and excitement around AI.
Primarily, you've also got this shift of moving from sort of personal productivity into more operating redesign. And, where before, sort of the bulk of your AI activity was all focused around, okay, how can we increase personal productivity? That's an easy one to do.
You added Claude Code, and you added some of these, sort of the co-pilots of the situation. But everybody agrees at this point to really drive the key benefits around AI and maximize the impact it can have, you really got to touch on the sort of larger workflow and that whole operational model redesign. Much more complex, many more changes, obviously much bigger impact on the organization, so you're going to have much more attention to costs.
And I think back to a decade ago when we were in the digital transformation phase, very similar story. You had 80% of projects driving almost no value, but everyone still fought forward and kept trying to push the technology because missing out was such a massive concern and a potential problem for the organization. So even though chance of success was suspect sometimes, again, it was that sort of 80/20% rule, and none of those projects really ever expected to-- You never saw more than about a 20% success rate on them.
And I think you're seeing sort of a similar story with just a different technology today with AI. Are people looking at alternative models? Because some of these token costs are attributed to the cost of a proprietary model, but there are all these open weight models that people are talking about, and so maybe there's an alternative way of executing some of these workflows.
Yeah, I think as people get more experienced and more comfortable with the technology, obviously the cost question that we're talking about becomes front and center. And as a result, there's a lot of different strategies that people are looking at. There's a lot of talk about using for sort of those higher level areas around reasoning and sort of deep thinking.
You go and use your frontier model for that, but then on your sort of more day-to-day production issues, you might drop down to a cheaper model or even open source type model. And you're seeing companies also start to create their own models as well. I saw just the other day that Travelers had built their own domain specific model based on the millions and millions of documents that they've had over the years to really drive, A, cost savings, and B, when you get that more specific context involved, typically you get better performance as well.
So the models perform in their target areas much better when they're more tightly aligned with the use case that they're tackling, and that's exactly what Travelers saw and many other companies are seeing as well. The other aspect of this that people are talking about is the actual cost of training, and you're seeing some, to your point, organizations training their own AI models. But I think I read somewhere that the effect of creating ChatGPT-2 costs about as much as maybe a used car, and the latest models cost about as much as an aircraft carrier.
So as that method and those costs start to rise for training, how do we kind of think about that in a way that maybe allows us to build, or maybe more companies at least, to build more models in a way that doesn't break the bank? Yeah, I think one of the challenges you're seeing is a lot, and you're seeing in the market, a lot of the talk around the frontier models is that that is becoming a commodity, I think, faster than a lot of people thought it would. Which just goes to illustrate how far advanced the technology is moving and how quickly it's moving.
So now you're able to see these organizations, like Travelers, build their own LLM, relatively, I don't want to say easily, but it's cost effective for them to do that. On the other hand, in terms of managing the costs around it, yes, you're seeing the open source models. I think there is still a bit of a subsidy, which is amazing to think about, with sort of the top frontier models.
But again, that's going to have to disappear, and how we address that is to be determined a little bit. But organizations implementing this technology have different options. There's different standards that can be used.
We've got a new program that's going on that we've done with NVIDIA and IBM and Red Hat and a couple other folks for a new document format called DocLang. And what we're seeing on that is the cost to process, the tokens used to process a DocLang formatted document that's made for a machine to read and not a human to read it is 500, 600 tokens. That same document just sent to the same LLM, but in a different format, may be 6,000, 7,000 tokens.
Over 50,000, 100,000, millions of documents, real material savings just based on the input format that you use. So there's a lot of tools and techniques and different applications that can be used to manage that cost. Really comes down to planning and how you design your program to move forward and get it into production.
The other challenge is, well, just the fundamental scarcity of GPUs, and there's a lot of folks who are now wondering, well, do I need GPUs for everything? Can I use something else out there that might be a little more cost effective, either for training or for inference? So are we going to see a lot more maybe diversity of the underlying processor technologies that we're using to run these AI models?
Exactly. There's a finite source of GPUs that are available, and they're not cheap, as you indicated. Some of these technologies can run on CPUs, and perform close to parity, with the higher powered, larger investment that's required for GPUs.
So again, that ability to really think through and design for the use case and figure out, okay, what am I trying to get at the end of the day? That's where you can really drive more efficient planning and reduce your development costs significantly. We also hear a lot about, well, world models, and there's different types of AI models that people are playing around with.
Some folks have more of these small language models. So is that also going to change the way we think about the economics of AI? Because there's just going to be a lot of other things to play with, and maybe we will mix and match them as needed.
What do you say? Yeah, exactly. We've always advocated, use the right technology, use the right model for the use case that makes the most sense.
There's a lot of interest in the past 12, 24 months around domain-specific language models, sort of a version of LLMs that are specifically tuned for a specific business application. Again, much more cost efficient. That's exactly what Travelers saw when they used their own data to build their own LLM.
But you're also seeing some incredible things happen with these larger models that go on. I saw the other day that, I think it came out of MIT, that they've actually simulated the entire population, every sort of individually, globally, and they can run around and look at specific new product introductions, because they've essentially modeled the entire world population in a lot of respects. And it used maybe 40% of actual data on people.
And then, the other 60% or so was synthetic. But again, an incredible model that really illustrates the entire globe, and how you can use these things. There's just so many incredible advances going on these days.
But what's your use case that you're trying to address? And there's probably a more efficient way to tackle it than you might be thinking about. How dynamic will this swapping in and out of models get?
Because today, I think people are, well, I built it and it runs on this model, and here's my inference. But there's these routers that are starting to emerge, and I'm starting to wonder, down to even an individual prompt, will the prompt be evaluated for which model should service this best, and that will be done at machine speed, and this will be a way that we keep swapping in and out the most cost-effective model? Yeah.
Routing the work to the right model is becoming more and more common today. You're almost seeing it, as you indicated, sort of near real-time. But that's going to be, I think, one of the key drivers to increasing efficiency and reducing costs for a lot of this work.
You need to have the right controls in place. But that ability to, even with a single API in some platforms, allows you to test and validate and route to the most economic model that's available to deliver what you're trying to achieve. What are people not thinking through enough?
" It's the million-dollar question. So much of this value that can be driven by AI is not based on the model, and it's really built around the workflow. And I think that's where probably the biggest impact can happen.
8, and then they'll come and sort of tout the achievement without really thinking through what's that cost in terms of dollar value to deliver, but also just lost business opportunity. So again, it's never really the model challenge. These models are so incredibly powerful these days, and as I said, they're becoming more of a commodity.
It's typically a workflow problem. So your ability to rethink that workflow and make it sort of work and function in an AI native manner, and maybe eliminate some of those unnecessary steps and figure out where the human in the loop needs to stay involved, all those different elements, that's where you can drive, I think, a ton of your value. That's from the design perspective.
I think from the sort of the more global area, a lot of it's around AI fluency. And a number of these new regulations that are coming out, the EU AI Act and similar guidelines and frameworks that are there really call for increasing and mandate AI fluency within your workforce. Again, the amount of money that gets spent on people asking OpenAI what the weather is today is incredible.
And if people get smarter about that, again, there's significant savings that can be achieved. As they say, good enough always triumphs over perfect. I know.
It's interesting times we're living in and a lot of fun, and that's why I enjoy working so much in the industry. So what are some of the use cases that you've seen that you go, wow, that just kind of rocked my world a little bit? Because I think a lot of folks are trying to figure out the ROI, and I ask this question because increasingly it seems to me that it's going to be harder and harder to differentiate in the age of AI.
If you build something interesting, I'll probably replicate it in short order, and so is AI rapidly becoming just the cost of doing business? Yeah, it's a great observation, and don't disagree with you. Everybody's got access to the same technologies and the same models, so you really got to differentiate around the user experience and what exactly you're trying to capture.
I had a project about a year ago that looked at contract lease extraction. And this organization, they needed to report to the government up to about 300 different data points within each lease that they received, and they received up to about 50,000 every year. Obviously, very complex documents.
The group that they'd outsourced it to in the Philippines ended up being only 58% accurate, which we didn't know about at the time, just because they were, oh, a human did it, it must be accurate. But when we really dug in and tried to figure out, okay, why aren't these results as good as we're seeing, a lot of the training data was off. We tried to send that just to an LLM to see what that would do, and we got a little bit better.
I think we got to 67, 68%. But when we paired sort of more traditional machine learning techniques with the generative techniques and sort of that hybrid AI approach, bringing the right technology to the use case, we were out of the box at 87, 88% with no reinforcement training, nothing else. So we went from 58% with a human doing the work to 88%, I don't want to say overnight, but very quickly, just by, again, pairing the right technology to the use case and bringing multiple technologies together versus just, oh, let me just send it to an LLM, let me just send it to this tool.
And that's where I think that sort of architecture question comes in, like, okay, how do we bring the different pieces and different elements together, whether it's computer vision or traditional machine learning or generative technologies, and mix and match those to deliver the right value for the work that we're trying to accomplish. All right. Well, folks, you heard it here.
No matter how advanced AI gets, the fundamentals of economics are still going to apply. Hey, John, thanks for being on the show. Thanks very much, Mike.
Have a great day. All right. AI Leadership Insight series.
You can find this episode and others on our website. We invite you to check all those out. Until then, we'll see you next time.