Agent-Optimized Inference Cuts AI Token Costs
AI token costs keep climbing even as the price of a single token falls. Mike Vizard talks with Jack O’Brien, co-founder and CEO of Subconscious, an MIT-born startup that raised $5.1 million. Together, they explore why long-running AI agents drive so much spending and how smarter inference can bring it down.
Why agents drive AI token costs
Token consumption is doubling roughly every 11 weeks, Jack says. Furthermore, most of those tokens now go to agentic systems rather than chatbots. A chatbot session might run 50 or 100 messages. An agent, however, often runs hundreds of messages and hundreds of thousands of tokens on one task.
Routing is a red herring
Routing requests to cheaper models works for short tasks. For long-running agents, though, switching models mid-task breaks the reasoning chain and invalidates cached tokens. As a result, the system slows down and AI token costs go up. Instead, Subconscious built an inference engine around runtime memory compression and efficient caching.
More users per GPU
GPU scarcity is only part of the problem; utilization is the other part. According to Jack, the compression approach supports two to three times as many concurrent users on the same hardware. In addition, the engine runs on GPUs, CPUs and NPUs, and customers can use an API or deploy on their own GPUs.
Lowering AI token costs with good enough models
Many companies still run models that are a generation or two old. Consequently, moving to newer open models and serving them efficiently can cut AI token costs while improving accuracy. Looking ahead, Jack expects agents to run much of the work behind every company, and he sees standards such as MCP sticking around.
Smarter inference is becoming as important as smarter models. Explore more AI coverage and the latest Techstrong TV interviews.
For more information please visit subconscious.dev
Transcript
Hey guys, thanks for the intro. We're here with Jack O'Brien, who's the CEO of Subconscious. 1 million in funding to help us figure out how to cut the cost of AI token consumption.
Jack, welcome to the show. Yeah, thank you so much for having me on. There's a lot of people focused on this issue right now, and to be honest, though, I'm not quite seeing if we're making a whole lot of progress.
The cost of an individual token seems to have declined, but we're consuming more of it than ever, and people are actually cutting investments elsewhere to fund AI token consumption. So what's to be done about all this, Jack? Yeah.
Our take on the world is a couple different mega trends are happening right now. Number one, token consumption is doubling every 11 weeks, and has been for the past three years. So there's a lot more companies consuming a lot more tokens.
The second thing is we see open source models are really at parity with the closed source models, so that's been a nice source of the drop in price. Why use the premium version of a system when you can use something that's a lot cheaper? And then finally, we see that just the vast majority of these tokens are being consumed by agentic systems instead of humans interacting with chatbots.
And the difference between an agent and a chatbot is, a chatbot, let's say I have a very long session. I'm writing a message, it sends me a message back. Maybe it runs 50 messages, 100 messages.
An agent, on average, is doing hundreds of messages and hundreds of thousands of tokens, and that's really the source of where all of the consumption is coming from. And then if you talk to any software engineer, and it's really bleeding into other professions like legal and marketing and HR and finance, people are just becoming addicted to using these agents to work. It's very hard to imagine life beforehand, even a couple of years in.
So what we've done is we've built an inference system, top to bottom, so how we serve models, that is really optimized for these agent workloads specifically. And when you build a system top to bottom that treats agent workloads and not chatbots or these other ways that you can use language models as the primary workload, we save significantly on cost. We can generate tokens a lot faster, we actually improve the accuracy of the systems overall.
So we're taking an approach to this that is fixing the very core of why it's very expensive to run agents over long periods of time on lots of tokens, which is, let's find a way to utilize GPUs better in a way that's optimized for agent workloads specifically. Are you routing the AI agent requests therefore to different models based on the task at hand? Because not all tasks are of equal value or equal complexity for that matter.
And if so, how do you know which task is of what nature? Yeah. Routing's an interesting topic that I think is a bit of a red herring.
The point is, it's really useful if you're doing very short tasks, if you're doing a one shot here, maybe five or six messages back and forth there. But for these problems that require really the most valuable things that you can solve with AI systems, things that are, "Hey, go do this task and come back to me when it's done," those require lots of steps, lots of tokens over long periods of time, and are not a good fit for routing because you don't want to switch back and forth between models during those tasks. The two main reasons are, number one, a system that reasons early on, you want it to be the same system that's using that reasoning as it's executing the task.
These models are really trained to understand their own forms of thinking. The second thing is if you route models in between requests, you invalidate a lot of the cache tokens and it ends up significantly slowing down the system and also making it much more expensive. So at a very high level, we are not routing.
You can route to us. But what we are doing is we built an inference engine, so think the software that runs on a GPU to turn a model into a system that actually serves you tokens. We built that top to bottom to support lots of steps, lots of tokens as really the primary use case.
We do that with some newly developed techniques around compression and very efficient caching, making the assumption that you're going to be interacting with the system back and forth over time instead of just trying to generate one response really well. So I can pick one model versus another, but once I make that choice, I still need to optimize the interactions with that model and improve the overall utilization rates of the underlying infrastructure. Which kind of brings us to our next question.
Yeah. It does seem like the cost of tokens in part is high because we don't have enough GPUs that are out there, but at the same time, the utilization rates for those GPUs are, shall we say, highly suboptimal. So can we get better at using those, and is that part of the whole equation here?
Yeah. That's exactly right. Part of the reason that those systems really get bogged down is, let's take a really specific example.
Super high-end GPU, the V200. Very expensive to own or lease as a company. It can really only support a handful of concurrent users, though, who are using it really thoroughly.
Part of what we've built, and a side effect of this compression that we're doing, is you can support many more concurrent users. So somewhere between two and three X as many concurrent users, and that drives cost down. That makes the infrastructure that we already have, it's essentially a force multiplier that you can download.
How hard is it to set this up? Because I think a lot of folks out there think that at least deploying anything in AI requires a few rocket scientists. But I have to wonder, at some point, is this going to become something that mere mortal IT folks can run?
Yeah, absolutely. I think the core of what we built is a lot of deep research. My co-founder who really developed the core technology was a research scientist, PhD postdoc at MIT for a decade, and this is the output of a lot of his research from 2016 to now.
But we package it up in a way that's like using any other type of software. I think the big difference is you have to have this very high value hardware in order to really utilize it. So that means running in a data center.
But there's more and more powerful versions that you can download on your own machine if you're running a powerful enough laptop or edge device. Moral of the story is how we work with customers today, we have an API. You can get an API key and start using.
We've helped enterprises deploy our system on their GPUs in their data centers or on-prem locations. And then we're working more and more on something that could live locally. There's definitely some trends headed in that direction.
Are we too obsessed with GPUs for inference, and will there be other classes of processors that we're going to use at some point? And is that something you're going to factor into your thinking? Yeah.
What we've built is relatively hardware agnostic. We're at a level where the kind of key optimizations that we've built are around runtime memory compression. So as an agentic system builds up all of this context, we found really effective ways of compressing less relevant information in a near lossless way.
We've ported that to a whole bunch of inference engines, and that can run on GPUs, CPUs, NPUs. We've experimented up and down the stack. Where we found the most success right now is deploying our system on NVIDIA GPUs and getting tokens to our users that are much cheaper.
But as the underlying infrastructure changes, we're going to be another additive step to utilize models better. " I will never talk bad about any of our customers, but it's amazing. It's almost surprising that there's a lot of companies out there using Second, third gen models and treating that like it's just fine.
Um, and I think what that points to is a couple of things. Number one, if you're using something that's a couple of years older or even six, 12 months older, it's probably not as capable. It's also not as efficient.
Um, so you can make big gains just by using something that's closer to the cutting edge and then using technology like ours in order to serve that takes you another stepwise gain in performance. So there's a huge gap delta in the cost and accuracy benefits that you get. Um, the other thing that that points to, though, is we're hitting a threshold where the models are good enough for most tasks.
" You want to use the latest, greatest, hottest thing. Um, but the truth is, as OpenAI, Anthropic push the frontier of what's possible, it's not going to be necessary for the vast majority of businesses. What they're going to care about is customizing it for their company, making sure cost and latency are as low as possible.
And we're helping a lot of companies do that. As you look down the road a little bit, what are you guys thinking about next? Or where does this research project kind of go from here and how does that manifest itself down the line?
Yeah, the platform is live today. dev, get an API key, start experimenting. Um, for bigger enterprises, we've done these enterprise deployments.
Um, yeah, we think the vast majority of tokens are going to be consumed by agentic systems, and we think that our infrastructure is absolutely necessary if these systems are actually going to proliferate with the hardware that we have available today. So our vision for the company is that we'll be a core piece of the stack to power trillions of agents that are going to be embedded everywhere in the economy. I mean, we kind of envision a world and AI-native companies already operate this way, like us, where agents are just an integral part of every business.
And it kind of looks something like this, where the consumers of companies will interact with their personal agents. This is already kind of how I interact with the website and our company data and even definitely our platform and code. Um, and that personal agent actually ends up interacting with companies via their APIs or MCPs or something like that.
And then you think about what's actually on the other side of that company's API. It's agents that are maintaining the API, and it's agents that are filling up the database and making sure things are working properly. And then behind that level of agents is actually the humans that are running the company, right?
But this middle ground between my personal agent and their API and what's actually running it, it's all run by agents, top to bottom. It's going to be a lot more efficient. It just works really well.
That's how us as a now less than 20 person team are able to do so, so much. Um, and we think virtually every company is going to look like that. Um, what does that actually-- What's the grand vision?
So our grand vision is we're going to power trillions of agents all over the world. What does that actually get us? We think it gets us to a point where your life as a person is you get to show up to work, solve meaningful problems, and then go enjoy your life.
Um, where all of the meaning, monotonous b******t in the middle is hopefully swept up by these systems. So to your point about that, how will these agents kind of negotiate with each other? Because in some cases, um, my personal agent or the agent that I'm invoking from you guys on my behalf will have to negotiate with agents from cloud service providers and other folks, and they may have, um, shall we say, diametrically opposed missions.
Yeah, I think that's true. Um, I'll say it's just not our core business, how that's going to work. What we want to enable is for your agent to run as efficiently as possible and to remain accurate and reliable, no matter how much work it's doing.
As you think about all this and where we are at the moment, do we need more standards in this space, or do you think that at some point, the tools we have kind of will be sufficient? But I wonder, um, do we need to make it easier for startups like yours to kind of create some value on top of things? Because right now I feel like every AI platform, to a certain degree, is very proprietary.
Yeah, I would agree with that. It's one of the more most competitive, fast-moving waves of all time. Standards are going to keep popping up.
Companies, some will use them, some won't. Um, but I think the winning standards will win. And what actually ends up winning, it's something that people like, um, it's something that the systems can understand kind of natively, and it's something that the popular companies adopt quickly.
So people are going to keep trying to create standards. Um, some of them will win, some won't. MCP is a good example of something that came out.
People weren't so sure. There was a big surge, there was a bit of a retraction, and now it's becoming clear that that's kind of here to stay. Um, it's pretty embedded into the systems.
And the interesting thing is, when you create a standard like that, that survives two or three generations of language models, that becomes pretty embedded in the training data, and then there's really no shaking it. Um, so, I'm not exactly sure what a standard looks like, but we're going to keep experimenting with the ones that come out, and I'm sure several of these are going to stick around for a long time. All right, well, folks, you heard it here.
AI agent optimization is definitely on the way, and based on what we've seen so far, it's going to be sorely needed. Jack, thanks for being on the show. Yeah, thank you very much.
All right, and back to you guys in the studio.