Tokenomics Brings FinOps Discipline to AI Cost Management
AI Cost Management Needs a Shared Language
AI cost management is becoming a board-level concern as organizations move from experiments into production use cases. In this Techstrong TV interview, Mike Vizard talks with Amit Kinha, Field CTO at DoiT, about the rise of tokenomics and the work of the Tokenomics Foundation.
Kinha explains that tokenomics is not only about lowering the price of tokens. It is about understanding the full economics of AI. That includes model usage, GPU cycles, infrastructure, application design, business value and the outcomes that AI systems produce.
Tokenomics Extends the FinOps Model
FinOps gave cloud teams a common way to discuss visibility, governance and optimization. AI introduces a faster and more complex version of that problem. Teams now need to understand what a specific AI task costs and whether that cost is justified by the value created.
That shift makes AI cost management more complicated than traditional cloud cost tracking. A single task can involve input tokens, reasoning tokens, output tokens, database calls, vector searches and agentic workflows. Each step can add cost. Each step also needs to be tied to business impact.
Value Matters More Than Token Price
The conversation highlights why price per task may be more useful than token price alone. A model that appears cheap can become expensive if it needs more steps to finish the work. A more expensive model may deliver better results if it completes the task faster or with fewer failures.
Kinha also explains why AI adoption depends on proving value. Engineering teams can often show faster delivery with coding tools. Other teams, such as marketing, sales, research and healthcare, may need different measures. The goal is to connect AI spending to real outcomes.
Context Engineering Becomes Critical
The discussion also turns to context engineering. AI systems need the right information to produce useful work. Too little context leads to weak answers. Too much context can create noise, risk and unnecessary cost.
For technology leaders, the takeaway is practical. AI cost management requires more than usage reports. Organizations need standards, measurement practices and governance models that show what AI costs, where that spend goes and what the business receives in return.
Transcript
Hey guys, Thanks for Throwing. We're here with Amit Kinra, who is field CTO for DoiT, and we're having a little chat about, well, tokenomics, specifically this new Tokenomics Foundation, of which he is also a member of the governing board, and he's also a member of the governing board for FinOps. So I think our man here probably has some answers to some questions about, well, just how do we measure this AI stuff we're all tracking because it's getting expensive.
Amit, welcome to the show. Thank you so much for having me. Yeah.
So what is the fundamental mission here of the Tokenomics Foundation, and what makes it different than what we were already trying to do with the FinOps Foundation? Yeah, great question. I think when I think about FinOps Foundation, the mandate really was a lot of companies are going to public cloud, and the complexity of running cloud efficiently is a very different paradigm than it was when you're on data centers.
And FinOps Foundation really is born out of necessity that, including me, all of us are trying to solve the same problem. We have this data set that exists, and we know it's valuable. We know cloud is really important, and yet we couldn't really shake the feeling that is this the right approach, from the visibility side, from the optimization side, from governance, right?
What I call preventing the finops moment that runaway costs and everything else. I think tokenomics is different than that. And tokenomics I think really centers on two things.
One, the mandate for tokenomics that we call it, it's not just about the price of a token and being efficient there, right? It's really about how do you see value coming out of AI. And the mandate here is much more fast-paced than it was for public cloud adoption, which it's still a lot of companies still getting on away from data centers when needed onto public cloud.
With AI, the board, the CEO level, have a mandate to now use AI not only for their customers to build a better experience, but also for operational efficiency in their company as well. And with the fast pace that's happening right now with AI especially, I think a lot of enterprises are trying to solve this problem of the age of POCs, that was last year. Token maxing, which lasted about five days, felt like this year.
And away from those things towards how do we prove that we want to invest more money on AI for certain areas, certain areas we want to be more efficient with. And finally, it expands away from only consumption and value, which we're describing, which is similar to FinOps. But we also think about then the creation of tokens as well, right?
" And well, it's not really free. You still have to run the hardware somewhere, right? The Nvidia chips aren't any cheaper if you're running open source versus GPUs that you were using for running your own fine-tune LLM there.
So I think the mandate really for the Tokenomics Foundation is not only with all the stuff that FinOps Foundation was really doing, but to really help companies and organizations really figure out how do we actually use AI and use it in a way that we have confidence that we're not just shoveling money into a fire pit there. Mm-hmm. I think there's not a lot of clarity on exactly what a token is from one provider versus another.
It doesn't seem like there's a whole lot of easy ways to do an apple-to-apple comparison. So what needs to happen here on how we define the tokens, how we track them to kind of get to some ability to actually compare them? 20 a gallon.
And on one hand, that's important to know, but that's only part of the information, right? You'd want to know what kind of car you're driving, how many stops it'll have, how many miles are going to go. And so for me, token price is important.
But with the rise of two things, agentic AI and then also reasoning models, it's not an apples-to-apples comparison that might've been valid last year. You can compare, say, a Haiku versus Gemini Flash versus I think the sole model or whatever. Now you may have a situation where the token pricing may be cheaper, right, for one model, and yet the overall cost for the task that we're describing might be a lot higher because it's doing a lot of iterations on reasoning, for example.
So, I like to think of the way to really solve that is abstract away from only this point-in-time pricing of a token, and we'll think about, well, what was the outcome we wanted in the first place, right? And that really depends on each team, each product line. An example that I'm working through right now is say, what is the price per task of when we're marketing teams creating a pitch deck?
If I looked at only the token pricing, that might give me something, but what I really care about is the output that I wanted was this pitch deck that is the right format, the right version, has the right data in it. And if it costs $5 for that task, well, I don't really care if there was a lot less input tokens on one model and a lot more reasoning tokens on another model. Unlike the task is what I care about, right?
So I think the shift that we have to really do is start thinking about AI value, not in terms of the unit price that we had, say, for example, in FinOps, right? " And we keep seeing that in the industry, at least some of our customers that, let's take another example of a call center. " That's useful, but it's not really useful when you think about should we invest more or less there.
What's more useful is saying the call per successful call where a human was not in the loop, which is perfect, is about $7 per successful call. And anytime we have to have a human in the loop, it's $4. Because now you can think about where do you want to either optimize, where you want to invest, or where you want to change out the model provider to figure out, maybe we don't need to use this model because the unit price for a call is going down.
Right? So again, I know it sounds a lot very similar to FinOps. We think about unit economics, we think about the output.
So in a way, the mandate really right now in the first, to be fair, the Tokenomics Foundation was created about, what, two months ago? The governing board was announced a month ago. A lot of the work is happening in the working groups within the foundation around three pillars of creation of tokens, consumption of tokens, and then finally the value of tokens.
Right? On the creation side, we never really thought about in FinOps about, for example, from atoms to tokens or atoms to VMs. But in this space, we really do care about that, right?
We care because energy is a constraint, the GPUs are a constraint, and all these different indexes that we care about, we didn't really have to consider before. And then on the value side, in FinOps, we talk about unit economics a lot. But here the need for proving the value of AI is way more important upfront than maybe a lot of companies where we would say crawl, walk, run.
We'll get to unit economics later. Right now let's build some dashboards. We'll start optimizing our cost, and then we'll figure out this part later because it's kind of constrained, right?
The cloud cost is a budget that we have, and we don't want to exceed it. With AI, it's opposite, right? " Companies who have pressure to use AI in their companies, and yet they're finding this area where they want to use AI responsibly.
They want to make sure that it's getting some value back there. So, it's almost like we're speed running whatever we did in FinOps, with expanded scope here, with more pressure, more visibility than ever. Value is in the eye of the beholder, is it not?
Because if I've got a research team that's working on, I don't know, some project related to cancer, the cost of a token there is considerably different, or at least its value is considerably different than the meme that somebody in the social media department created. So how do I kind of think about this? Because at the end of the day, tokens are a reflection of the scarcity of GPUs.
So do I need to get smarter about who's using what for what task when? Yeah, that's a great question. I like to think of value as saying it's an investment we're making in order to either do a couple of things.
Either be more efficient with our time, be able to go to market faster, be able to augment or do higher quality work without having to spend an inordinate amount of time. So for me, the value of a token is not really the right question. Because a token, again, is just a price that will keep collapsing, but the same model, the same capability is probably 60, 70% cheaper this year, and it'll continue to happen.
But there's a concept of Jevons paradox that everyone has been talking about, which is as the price starts to collapse of intelligence, well, there'll be a lot more use cases that we previously would not have used AI for. " We don't want our sales team using AI to do one-pagers because it might cost $25 per one-pager creation. If that's down to say four cents now, well, we should only be using that.
So for me, value is, I completely agree with you, value is in the eye of the beholder. But that value part is the reason we'll be able to accelerate in companies the adoption of AI, because now we can prove this is what we're getting in return. And I think a lot of the work around where its clarity is where I think AI is being used very definitively today, which is on coding.
We have Codex, we have Cloud Code, and there we can see value very quickly. We can say we're going to production faster, we're able to build features faster, we're able to remediate things when there's issues. So the idea of a 1 to 10x engineer, which we all used to hire for, well, now we're really shifted completely.
When we hire people in engineering, we don't just think, can you go on a whiteboard and draw an algorithm that you know. We want to see, can you use the tools like AI to be more efficient use of your time to build faster? And I think outside of engineering, say things like marketing, sales, research, medical, all these different places, there's the value that will be created, yes, absolutely, we'll have to define per domain, per expertise.
But at the same time, value is simple in the way that there's an amount of cost that will go into it. It might be training before, it might be actual operational using it, but what are you getting back in return quantifying that so you can actually make better investment choices, make better comparisons between models, make comparisons between partners, all that stuff. You can't really do that if you don't have the data like, is this good or bad?
The other question I have, and I think when I read the announcement, there was some discussion about this, but are tokens even the right construct for figuring out pricing? And is there something else we could consider, and what might that be? Yeah, that's a great question.
We published a blog about a week ago about exactly this topic, and my view is that it's not about the token price. It's really about the price per task. And the task itself encompasses, yes, the token price.
The token price of input tokens, reasoning token, output tokens. But it also might, it includes, well, what if you had, for example, a lot of calls made to a database in order to get the information when you were working on a task, for example, and that database cost is, say, $5. When you think about just token pricing, you're just trying to index on the cheapest token.
But if you're using a lot more of them, reasoning tokens or an output token, you're not getting the full picture. So for me, where I've landed on this at least, and obviously things are shifting very quickly, is really at the task level. And the task is that thing that's the output that you really want in the first place.
Because then you're able to figure out, we spent 10 hours on this task, like a research report, and the output of it is what we needed, and that task might have cost $800. Now you can actually compare and what would that have cost without AI? Is the quality higher?
Should we now try to do this replicator across? Should we try to optimize everything else there? So to your earlier point, we started out trying to get everybody to use AI, and then we created this lovely term called token maxing.
Right. And everybody went out and used AI for all kinds of crazy things, and a lot of it may not have been, shall we say, useful, or some cases, it was probably cheaper to do it some other way. And now we're, for lack of a better phrase, sending many of those people to Tokens Anonymous because we're trying to wean them off of using so many tokens.
What is the process of kind of teaching people how to optimally consume tokens look like? And we hear a lot about context engineering and all these other things. Mm-hmm.
How does all that come together? Yeah, we keep iterating on words. Like first it was prompt engineering, then it was context engineering.
Now then we have loop engineering. Now I think graph engineering. All these concepts are evolving to basically say what is the best way to use AI in the most efficient manner?
It's a funny thing. I truly think the only way to really learn AI, unlike, say, cloud, we can get a certification where all the different things that we had before, the best way to learning how to start using AI. So every company that I've seen successfully implement change, where they're not force-feeding people to use something where you have unintended side effects of token maxing leaderboards and everything else, really is to have a budget to say, "This money is not to be able to now quantify value.
" Because lessons learned from that will be we can build something where, take a skill file, for example. I know we keep talking about pitch decks, been working on them all day. If creating a PowerPoint deck, the first time I try to use AI, a lot of lessons learned.
I create a skill file out of it. And it turns out everyone on my team is also creating their own skill files to do the same exact thing again and again and again. Well, that's when you recognize a company itself should probably have one thing that people can use, and we can continue to try to optimize that one skill versus having every single individual starting from ground zero.
But I don't think there's a shortcut to having the people who have to be using AI in your company being able to experiment and learn. Learn not only how to use it effectively, but also how to use it efficiently. There's obviously some things that as you get more and more advanced in your organization that you want to start thinking about, for example, context cache.
It's not easy to, or rather not easy, but it's counterintuitive to kind of be able to say, well, do you want to use the cheapest model for some tasks and the bigger models for complex reasoning? But if you keep switching in the same context, in the same chat window, for example, well, you have to reload that entire cache. It doesn't show up automatically on the bill when you're working right there.
So I think being able to not only let people experiment and have a budget for that, but also give them the data so they understand the cost of what they're doing. So then you can actually start to think about budgets and constraints and everything else. And I know it's not an answer that a lot of people would want to hear that, well, let them experiment, and of course, it's within reason.
You want to say unlimited budget, but at the same time, it might be that the person who used the most tokens in your company, that individual should be teaching people, what are they getting back in return? How did they actually leverage, say, $5,000 in a month? And what have they learned from it?
And what is the things that we should be implementing across the company? Within the company, it's really funny. If you think about a bell curve, the most extreme examples are where the biggest concern is.
Someone who's not using AI whatsoever, someone who's using too much AI. Both sides are a problem for an organization to now deal with, where there's no right or wrong answer, but at the same time, you kind of want to keep iterating and not having individual loops of people learning. You want to be able to now take those lessons and try to consolidate them, and then be able to build for the company as well.
So ultimately, what is that one thing that you see people doing over and over again at this point that just makes you shake your head a little bit and go, "I wish folks were just a tad smarter"? It's funny. My answer probably would've been different a month ago.
The biggest unlock I see for AI adoption in a company is not so much about choosing the right model or choosing the right token. It honestly is that, at the end of the day, if I use a ChatGPT, for example, on my personal account versus using ChatGPT in my corporate account, on one hand, if I do nothing, it should be the same answer. But the value really is being able to load the right context, and not overloading it where we say we have 100 MCP calls connected to everything I'm doing, but being smart about it.
What is the context I really need for this task that I have where I don't have to paste everything? Maybe I should have, there's a concept of like the AI brain, a company brain that we all talk about. How do we actually get the right amount of information for the task at hand?
" And now the entire literature of the last 15 years of anything we built on Confluence is available. Well, that's probably not the smartest thing to do. " It doesn't have your company proprietary information loaded into it.
" I think that's what we're struggling with, where people are either overloading too much context into it or are unhappy with it because they didn't load enough context there either. There you go. Folks, that's kind of where we are at the moment.
AI is a precious, limited resource. Use it wisely. Amit, thanks for being on the show.
Thanks so much for having me. All right. And back to you guys in the studio.