Steadcast
The AI Daily Brief cover art
The AI Daily Brief

Everything You Need to Know About AI Tokens

August 2, 202650 min · 9,519 words

Show notes

In this Operator's edition, Nufar Gaspar explains what AI tokens actually are, why costs can spiral in agentic workflows, and how to distinguish valuable usage from waste. Learn how to measure cost per successful task, eliminate “tokens that spin,” choose the right models and protect the experimentation that creates real value.

Highlighted moments

the most expensive token is the one that your best person is afraid to spend.
8:00
even though it was documented, de facto, we paid more for the same intelligence. And this is like a shrinkflation, right? The same sticker price, but a smaller candy bar.
17:54
Sonnet cost around $2 per task or $2.09 per task versus $1.94 for Opus. So because Sonnet needed more iterations and more reasoning, had to spend way more tokens to get to the same results. Overall, Opus, which is significantly on paper, more expensive model, it was cheaper to operate,
24:50
I spent in two weeks $1,500 on an agent that I was not using. So I opened the dashboard, like double clicked, and I realized that I have almost 400 million tokens in and almost zero tokens out.
31:01

Transcript

Introduction to AI Tokens

0:00Today on the AI Daily Brief, an operator's cut episode with Nufar, everything you need to know about AI tokens. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, Rackspace, Blitzy, Section, and Airtable. To get an ad-free version

0:30of the show, go to patreon.com slash AI, or you can subscribe on Apple Podcasts. And to learn more about sponsoring the show, send us a note at sponsors at aidailybrief.ai. All right, friends,

Nufar Gaspar Introduction

0:40well, Nufar Gaspar is back today, and Nufar and I have been cooking up a lot recently. A whole slew of you have done our most recent program, have explored our most recent program, the Choose Your Own Adventure style AI summer adventure. Plus, we've been cooking up an expanded set of educational resources that we'll be telling you about soon. But one of the realities that both Nufar and I have been living in is every company we interact with dealing with the same questions of AI tokens and token economics. We are now firmly in the agentic era of AI, where companies have to

1:11think not only about how to get adoption and how to maximize AI's value, but how to do so in a way that doesn't just totally break the bank and where the right intelligence is being used for the right problems. Anyone who's ever built an open clock can tell you that getting the right models to do what you want them to without going off into endless cycles of spin takes some real consideration. Today's episode is designed to be the ultimate primer on AI tokens, what we're

Defining AI Tokens

1:33talking about when we say that term, what the new challenges are, and some of the key pitfalls to avoid as well as strategies to maximize the way that you and your company use AI tokens. All right, Nufar back with another Operator's Cut talking about the topic du jour, the topic on everyone's minds. We are talking tokens. How are you doing? I'm good. Very psyched to talk about tokens. Yeah, I think it's, I love this period in a discourse where we've gone from sort of pulling hair out, freaking out about the new change to actually settling into new tactics, new strategy,

2:07and I think this is a perfect fit with that. So tell us a little bit about what we're going to be talking about and let's dive in.

Eras of Token Consumption

2:13Good. So the reason why I wanted to do this episode is because every room that I walk into these days, like literally every room, have some version of the same token conversation. Some practitioners feel like they are being watched when they use an expensive model. The regular users wonder whether one ambitious prompt will eat their weekly allowance and the leadership teams, they see a bill growing faster than expected and then they start asking a lot of questions on whether all of these tokens produced anything useful. So I don't know where you guys are sitting, but there is a new anxiety around using too much

2:46intelligence. And I actually want to flip the conversation, first of all, to make sure that everybody understands what tokens are and what the bill actually means and then how to spend them wisely rather than sparingly. So that's why I'm here and what I'm planning to do today. One of the places that I've found myself with this conversation is there's been such a visceral reaction now as the cost has gone up. I have a bigger concern around people retreating, back to known ROI biases and not boring, but ultimately low stakes use cases, let's say,

3:19as compared to what AI can actually do, that I've found myself in the position of having to defend things like token maxing and token leaderboards just relative to where the tone has shifted. I think obviously we'll get into today the smarter version of that conversation. So I'm excited for it. Exactly. All right. Because there is a growing conversation around that, I think that there is a better language around the feeling of tokens shouldn't be just a financial thing. And just recently, the OpenAI CFO proposed a scorecard and it was called Useful Intelligence Per Dollar.

3:52And that's built around the one question, what does each successful task actually cost? And that's the conversation that I think people should have. And I want to help you read the whole story.

Understanding Token Costs

4:01What are tokens, what your work costs, and where usage creates value and where it quietly leaks value rather than adding. So to kick us off and how we got in here, I want to walk you through four eras of token consumptions. And probably you will recognize where you are. And we all started by being basically token oblivious. That was the all-inclusive era where model companies subsidized the usage and the flat subscriptions hid the meter. And many individual users, and many of them are still there, just see the ceiling, not a pair token price. So that's where

4:35we started. And then, like you said, we got into the era of token maximizing. That was the leaderboard era where usage became the badge of high maturity. And we all remember some of the conversations around Meta who tracked employee AI usage on an internal leaderboard. They used to call it tech, and they used roughly between 60 to 74 trillion tokens in a single month. And according to the data that was published, the top individual user used 280 billion tokens. So to give a sense of how

5:06many this is, that is roughly 2.3 million books worth of text. So if you want to try and imagine that, that's about 50 books every minute continuously for a month. So that was the meta story. And then Uber launched also an adoption leaderboard and burned through the entire 2026 AI coding budget in about four months. And there was another company unnamed, but according to TechCrunch, they ran up 500 million dollars of Claude Bill with no usage limits in place. So then in that era, usage became the metric,

5:38and Dashboard measured activity while claiming to measure value. Obviously, this was unsustainable, and I know you have some opinions on that, but I'll be curious to hear your points. But I also want to say that it actually got us to be, as always, the pendulum took us way too far to the era that

Token Interest Era

5:54I call token interest, which is where we are. But if you have anything to say in defense of leaderboards, I'm here to listen. Yeah. Look, the defense is less. Leaderboards are a great concept. They come with, I think, a set of very predictable challenges. In fact, so predictable that I would say that the hand wringing around the idea of people gaming them has always struck me as a little absurd. Of course, people are going to game systems if you put real stakes around them, but that's pretty predictable and also fairly, from first principles, you could figure out a lot of ways to deal with

6:27that. So I think, one, it's overwrought all the sort of people who are freaking out about that. Secondly, my bigger point was that a company that wildly overspends right now via a token leaderboard or anything else, I will bet any amount of money that they will be farther ahead than a company that underspends because they're overly concerned with proving out ROI or whatever it is on a sort of year time scale. Now, the Goldilocks scenario, which I think is what you're going to get into, is being able to experiment, being able to learn, being able to build, and being able to actually

7:00understand consumption while also not being afraid of it. But I agree. I think that the pendulum swung so aggressively, far too aggressively, back from token maxing and excitement to token anxious. And that's the paradigm that we've been living in, in the recent few weeks, couple months, whatever it is. Yeah. Good. So I think you love my model of how to use tokens wisely. But with regards to being token anxious, what I'm seeing in many companies now is that many employees are self-censoring themselves, basically trying to avoid costs. And even if we look at the same

7:32companies that were token maxing, so Meta went from the leaderboard to sending a memo that constrained the AI usage. And now the press is calling it token minimizing instead of token maxing. And Uber caps employees at $1,500. So even the same companies who are token maxing are now significantly shortening that. And when employees self-censor, it gets them to feel like every prompt is an ROI conversation. And that's not something that we want to have. And I think that this is a very bad error to stay in because I believe that the most expensive token is the one that your best person

8:06is afraid to spend. So this is where I want to direct all of us to be in. And I call it the token smart error, meaning that you need to spend wisely and not sparingly and understand what creates value and where usage quietly leaks. That's the entire kind of theme of this episode. And what I wanted to go by in this era, in this course, is to talk about the four elements of what a token actually is,

Auditing Token Usage

8:30why tokens were not born equal, how to audit your own usage, and how to govern or do it better within your company. And I'm trying to make it relevant to anybody, whether you're the practitioner that needs to apply some cost engineering playbook and be smart about that, or the executives and admins that need to be proactive and avoid having a difficult conversation with the CFO without a proper response to, let's just cut the bill without talking about the business implications, as you just said. So that's the plan for us today. So I want to start with introducing you to the token, because I feel

9:06that even though it's the most used term in AI, not too many people truly understand what it is, because that's in the root of every bill, quota, and rate limit with AI in tokens. So in a simple word, token is a chunk of a text that the model reads and writes. It's typically bigger than one character, and it's usually smaller than a word. And if you have never ever seen a tokenizer or how tokens look in action, OpenAI has a very good page that is open to everyone that you can just take a look at how tokens actually look. So it looks something like that. You can paste the text and then you will

9:41see how words are being chunked. So you can see that some words are staying like as one token, while others might be separated into multiple tokens. And interestingly, numbers often are being chopped in the middle and so on. So we'll put it in the show notes, but a very interesting experiment, if you have never seen how your text looks. And by the way, if you paste a non-English or a non-Latin language, you will see that typically the amount of tokens is much larger than an English language. So that's the OpenAI tokenizer. And a few things just to lend it home. In general, like the ratio in

10:18English is around three quarters of a word to token, meaning that if you have a page of text, it's roughly 1,000 tokens. And some languages that are like Hindi, Thai, Greek, other languages like that might get two to five X more tokens for the same content. And because billing is per token, then some questions, if you ask them in other languages, might cost you much more. And that's sometimes referred to as language tax. With AI, with code also, it's different and it has its own way. Indentation and brackets and white spaces, they all become tokens.

10:49There are some newer ways to tokenize text that are more code friendly in order to do that. But

Token Layers and Pricing

10:55still numbers is a huge problem. So you've seen the one, two, three, four, five being chopped in the middle. And by the way, that's also why whenever everybody's doing like the strawberry test for AI and it very badly fails in trying to count how many R's are in the word strawberry. In many cases, that's just a tokenization feature rather than a failure. And the model just has never ever seen the individual letters. It just saw the straw and the berry as separate words. And that's why it's counting it off. So modern models have various workarounds, but many of the like AI is so dumb

11:30memes are literally just tokenizer issues. So that's tokens. In terms of what everyday work costs, I think that's a good kind of mental model to have. So for example, drafting an email is around 500 to 700 tokens. A page of text as noted is about 1000 tokens. You can see a longer text. It can be more than that. If you send a model or a tool to do like a AI web search, often it will add a few thousand more tokens for the results, sometimes much more. Images, interestingly, are in many cases not that large

12:05in terms of how many tokens. They're roughly around slightly more than 1000 tokens. Interestingly, deep research can very easily be 70,000 or hundreds of thousands of tokens. But just the other day, one of our learners in one of our courses had a yes, no question. And accidentally, instead of asking for a web search for the problem, he was asking for the agentic tool to do a deep research. The tool spawned about 100 sub-agents to do the research. And then his yes, no question cost over 4 million tokens

12:36just to answer this question. So it can very easily amount to much more than that. Specifically, a few additional places where you can find very token-heavy workloads will be data analysis that can easily get to 1 million or more tokens per task. And heavy coding can also be very aggressive, similarly with many agentic working flows. So just to give you a sense of where it is, and to be a little bit more concrete here, the everyday stuff, as you've seen, like the emails and so on is almost free. So it's around half a cent and nobody should ration emails. It's not where

13:12the money goes. Search and research can multiply very quietly. So that can be a place to look for efficiency. And the top of the ladder, that's a completely different sport. So if you compare like email to agentic coding task, it can be a factor of a thousand or even more. And another thing that you need to pay attention is that every conversation compounds. So the model doesn't remember your previous messages. And as such, it sends all of the previous conversations within the same session back to the model. So by, let's say, turn number 10, it may be processing so much of the earlier

13:44exchange alongside your new message that the total grows much faster than the number of turns suggests. That even happened before the system prompt. And we'll talk about strategies later on. But this is one of the things that can very easily, just having very long sessions, can very easily amount to a ton of tokens being consumed. I think this is one of the reasons why this is such an important conversation is another way to put this is that the more advanced and ultimately higher value use cases consume more tokens, which is intuitive that more intelligence is

14:17required for bigger challenges. But the direction of use cases is proceeding this way. And so the reason that the token anxiety is going to create problems, if not addressed, is that it will incentivize people to stay swimming around less sophisticated use cases. So this is, the trajectory is clear in terms of less token consumption. You want, as a leader group, your people to be doing more advanced, more useful things with AI. It's just how they do it well. So you want them to do deep research where deep research is required, but you don't want them to

14:50accidentally do a deep research on a yes or no question that they can Google in a second. Good. So speaking of the agentic or the more advanced capabilities, those can significantly grow the amount of tokens because agents work autonomously in loops. And as such, they consume by very widely cited industry estimates five to 30 times the tokens of a simple chat and poorly designed agentic loops or agentic harnesses can be even worse than that because a typical task involves between 10 to 20

15:23model calls carrying instructions and history and tool definition and previous results. And I think according to McKinsey, they estimate that the roughly around 60% of an agentic tasks cost is tied to the checking and refining and the regeneration of the answers after the first response. So the expensive part is often getting from the answer to the accepted results. So that's an interesting one. And now it gets even more complex because tokens were not born equal. So by the way, the point here is not to not use the agentic tool just to know that, as you said, intelligence cost, but it gets even

15:56more complex because tokens were not born equal. And every model lab has its own tokenizer. You've just seen the OpenAI, but different model labs have different tokenizers. So for example, the OpenAI current tokenizer has a vocabulary of about 200,000 tokens. Gemini has around 256,000. And Lama by Meta has about half of that. And Claude is unpublished. And the reason why we all should care is that the price per million tokens is denominated in each lab's own tokens. And often we don't know them. And the same document can be 10 to 20% more tokens on one provider than another,

16:32and even more so for a code and non-English text. So the model behavior widens the gap. And one model may answer in a single pass while another reasons longer and writes more and takes more agentic steps or needs retries. And the tool around the model adds its own system and context and the loop design. So the two stacks doing the same task can have different token counts and different completion rates. And as a result, it's completely different builds. So the per token price is kind of the sticker, but the cost per accepted task is the operating metric because otherwise there is no way

17:04for you to compare between different providers and different tools. All right. An important story that also illustrates that what happened when Opus 4.7 came on board, the tokenizer basically under the hood changed and it was this April and the price sheet was identical to the previous model, the same dollar per million token. But the model was shipped with a new tokenizer that produced by Entropics on documentation. They didn't hide it. Roughly 30% more tokens for the same text. So there were quite a few

17:34independent analysis of over a million requests that found native tokens. They grew and the count grew by about 32% all the way to 45%. And the real world bills grew by 12% to 27% because some of the difference was absorbed by caching. Even Simon Wilson, he measured one of his own prompts at around almost one and a half X more tokens. So even though it was documented, de facto, we paid more for the same intelligence. And this is like a shrinkflation, right? The same sticker price, but a smaller candy

18:06bar. So nobody prints now 30% for your words per dollar, which is the case that happened there. So that's something that is constantly changing. Every lab tunes the tokenizer and often for good reasons. But the operator lessons here is that we have to talk about dollars per task and not dollar per token because the budget is like a moving denominator and it's not the way for you to try and understand how much it's going to cost. Let's talk about what tokens are used for by the AI tools.

Token Usage and Cost

18:33And you have to understand that every AI request has three token layers and they are priced very differently. We have the input tokens. Those will be the prompts and the conversation history and the files and the tools definition and everything that is part of the input. This is what the model reads and this is the cheapest per token, but can accumulate fast because if the history is being recent or if a lot of context is being read, that can cost quite a lot. Then we have the reasoning tokens. That's the second layer. These are the tokens being used for the model internal thinking before answering. For the

19:05most part, it's going to be invisible to you, but it's billed at the output rates, meaning at the high rate of per token cost. And those can add between 4 to 20x cost per request. And finally, we have the output. That's the answer that you actually see. And this is typically 3 to 5x more expensive than the input price per token. And I think the reasoning layer is the one layer that catches everybody by surprise because you might have a 400 token answer, but under the hood, it carried like, I don't know, 4,000 thinking tokens underneath because the model was having an internal monologue and doing a lot of

19:40thinking in order to give you the answer. And if you want the analogy, it's like thinking about the part of the restaurant bill that is labeled the kitchen time. So you don't get to see it. It's not part of the dish, but you still have to pay a lot of it for that. And the models with the high reasoning effort are often the one with the 20x amount of tokens being consumed versus the lower reasoning effort. It can be the same question with a significantly different price tag. And sometimes spending higher reasoning does not get you better results. So some metrics even show that for simple questions, it's better to use lower reasoning

20:13because the overall cost per task will be significantly lower and the quality will be improved without having the model overthink everything. So it's not always that smarter or spending more time thinking gets you better results.

20:29One of the more interesting shifts in enterprise AI right now is how quickly the conversation is moving towards infrastructure and operations. As AI moves into core workflows, regulated data environments, and agentic systems, enterprises need governed infrastructure and inference that can operate reliably day to day with clear operational accountability built in from the start. As those systems scale, the operating model increasingly becomes part of the AI strategy itself. Rackspace technology is the operator of the full enterprise AI stack from agents to infrastructure across private

20:59cloud, hybrid cloud, and edge environments. Rackspace builds and operates governed AI infrastructure, inference, and production AI systems for organizations where sovereignty, compliance, and uptime are non-negotiable. Therefore, deployed engineers stay embedded beyond deployment to help operationalize and run AI in live environments. To learn more about where enterprise AI runs and outcome scale, go to rackspace.com. Every AI coding tool on the market does the same thing first. It starts writing code. Blitzy does the opposite. Before writing a single line, Blitzy spends days reverse engineering your entire code base. Thousands of agents ingest millions of lines mapping every

21:34dependency, every undocumented constraint, every architectural decision made over the last decade. The result is a dynamic knowledge graph that understands your software the way a principal engineer would after 30 years in the building. Other tools guess at context with grep searches and markdown files. Blitzy never guesses. It builds true understanding first, then delivers over 80% of entire software epics autonomously. Validated, end-to-end tested, production-grade pull requests. That's why Fortune 500 engineering teams trust Blitzy with the code bases that matter most. See for yourself at Blitzy.com. That's B-L-I-T-Z-Y.com.

22:07Here's a harsh truth. Your company is probably spending thousands or millions of dollars on AI tools that are being massively underutilized. Half of companies have AI tools, but only 12% use them for business value. Most employees are still using AI to summarize meeting notes. If you're the one responsible for AI adoption at your company, you need Section. Section is a platform that helps you manage AI transformation across your entire organization. It coaches employees on real use cases, tracks who's using AI for business impact, and shows you exactly where AI is and isn't creating value. The result? You go from rolling

22:39out tools to driving measurable AI value. Your employees move from meeting summaries to solving actual business problems. And you can prove the ROI. Stop guessing if your AI investment is working. Check out Section at sectionai.com. That's S-E-C-T-I-O-N-A-I.com. This episode of the AI Daily Brief is brought to you by HyperAgent, where you run fleets of agents your team can manage together. New users get $1,000 in inference. Forget local agents and chat workflows waiting on your laptop to be prompted. HyperAgent deploys always-on agents in the cloud,

23:11doing real work across the tools your team already uses. Marketing's agent turns competitor moves into landing pages. Sales's agent enriches leads, drafts emails, and updates the CRM. Ops agent chases the paperwork and tracks the budget. Every agent has access to shared context and follows your rules about scope and approvals. It's time you add agents that feel like teammates. Hire yours at HyperAgent, built by the team at Airtable. Claim your $1,000 in inference at hyperagent.com slash AI Daily Brief.

Optimizing Token Consumption

23:40All right. And I think that the gap between input and output pricing keeps widening at the frontier. If you look at the Fable 5, it's about 10, not about, it's $10 per million input tokens and $50 per million. With the GPT 5.6 solid, it's a 6x ratio. So we're seeing the gap even widening. And those effort levels, that's also something that highly adds the complexity because these frontier models increasingly letting you dial the reasoning effort. With higher efforts, the reasoning tokens are

24:11significantly higher. And that's probably the dial that you should even be more mindful of, even beyond the models, because those can easily cost you 10 to 12x token increase between high or extra high effort to the lower medium. All right. One last thing here, price tag that experiment that came from Databricks, because a smarter model might not always be more expensive than a less expensive model. What Databricks did is they tested coding agents on real engineering tasks from its own code base, and they were using Sonnet 5. It was 1.7 times cheaper per token than Opus 4.8. However,

24:50Sonnet cost around $2 per task or $2.09 per task versus $1.94 for Opus. So because Sonnet needed more iterations and more reasoning, had to spend way more tokens to get to the same results. Overall, Opus, which is significantly on paper, more expensive model, it was cheaper to operate, which means that we shouldn't just reach to the cheapest model possible. We need to reach to the right model for the task. And that's not easy to get, but something to be mindful. The other thing that

25:23matters to the bill is the tool itself. So Databricks, in their same experiment, ran the same model at the same thinking effort through different agent harnesses. And they saw that more than a 2x difference in cost per task with the same quality using different harnesses, just because primarily one tool was feeding the model roughly three times less context than the others, and thereby the overall cost was lower. So very difficult bill to read and very difficult bill to navigate, and I'll try to help you as best I can. So bottom line, we're dealing with cost per task and not tokens, because otherwise

25:57we will not be able to actually compare apples to apples. And the cost will include the rate rise, the review, every correction that needs, and every additional iteration. And then you need to divide by the number of accepted results. That's your cost per accepted task. That's the metric that you should aim for and optimize for. And this also brings the conversation much more into return on investment and business value, rather than just having a conversation around tokens that is very hard, as hopefully binary understand, to meter. Good. The practical task that you can do, you need to take between five to ten

26:31representative tasks of what you do, run them through two model or tool options in order to have a good understanding of, in your option space, what you should do. Hold the input and the quality bar very constant and compare first pass success, attempts, human correction, and elapsed time, and the total cost. The winner is the stack that gets your actual work done reliably. I know it sounds like a lot, but if you have a good taxonomy that was optimized for yourself, and you know for the task that you do, which models overall

27:01get you better results, potentially with fewer tokens, or if you can do that for your team or your company, and you will need to do that recurrently because things change quickly, then at least you can teach folks that if you're doing that type of research, the recommended model to get you to the overall best quality and the value per task is the following, and so on. That's the current reality that we live in. Okay, now I want to give you a language on how to look at your own tokens, and hopefully using that you will be able to distinguish between the tokens that add value to the ones that

27:35not so much. Every token that you or your organization spends, in my opinion, is one of three kinds. There are tokens that I call tokens that teach, and this is running in both directions, meaning that you're teaching yourself, as Nathaniel said before, we don't want to stop the experimentation. So the tokens that include the experimentation, the failed workflows, and let me try these three different ways so I will learn. And so these are tokens that are worth spending because you get learning out of

28:05them, and you can look at them as tuition. And by the way, those also include what you are teaching AI about yourself. So those will be the identity files, and the curated context, and the knowledge packs, the memory can be counted as those, because teaching your AI who you are, what is your context, and learning what works for you in AI is very critical for you to continue moving forward. They look a little bit like a waste on a dashboard, or a lot like a waste on a dashboard, because no deliverable ship. But I claim that these are tokens that you need to defend fearlessly,

28:38because if you want to defend those, you will very quickly go back to just getting AI's help to draft emails and translate between languages rather than moving towards the workflows that matter. And especially if your company and yourself has a lot of catch-up to do on where AI is currently at. So I want the tokens that teach to be defended, because those are the things that will move the needle beyond the next category, which I call them the tokens that produce. So obviously, those are the most defensible ones, because those are the tokens that use to

29:09create work that ships. It can be the final proposal, or the research, or the code. Obviously, that's the thing that is much easier to show the ROI. But lastly, we also have the tokens that should be eliminated. And those are tokens that I call tokens that spin. Those can be machines talking to themselves, or automations that nobody is looking at their output, or automations that are running too infrequently. Idle agents, bloated context, misused tools and context, using Fable to write

29:40an email, so using the wrong model, and optimized flows, and so on. Those will be activity without sufficient output. So the token smart move, if I need to summarize, is to kill the tokens that spin, to tune the production, to make sure that it is cost-effective, and protect the teaching. And that's the order. First, go and do the audit on your spin tokens, and then do the rest. I have a very embarrassing tokens that spin story, which I will share in a minute. But I do want to also note, with regards to tokens that teach, that a failed experiment is as important as a successful

30:14experiment. So you should definitely encourage your employees to fail, to try, because otherwise, the tokens that produce will not yield as much value as possible. So it's embarrassing, as an AI expert, to talk about it. But my open claw was a chief of staff, was because it's currently disabled, a chief of staff that I called Chloe. And it was using the Entropic API. And because it was using an API, it was like auto-renewing all the time. And the bills were sent to a secondary inbox. I wasn't really monitoring them. And I was seeing that the charges

30:46seemed quite high, but because I was getting a ton of value, and because I was not paying attention to how frequently I'm getting a new bill, I wasn't noticing. And then early June, I was traveling. So I was not using my open claw at all. And still, I see that the bill kept coming. So I was saying, like, why am I still getting some bills? So I opened the dashboard, only to realize that I spent in two weeks $1,500 on an agent that I was not using. So I opened the dashboard, like double clicked, and I realized that I have almost 400 million tokens in and almost zero tokens out. So it was

31:19a ratio of almost 3,000 to 1 from input to output. And that's literally the definition of a machine talking to itself and billing me for like an internal monologue that it was running with itself. Looking further, there was a bunch of corn jobs that the open claw created for itself. And it was like a compaction job, the trend every 30 minutes on empty sessions. And even worse, like the trend was going up. So I, first of all, closed my open claw and only to optimize it differently. But if it happens to me in this setup, it can happen to literally everybody. And especially

31:52when the credit card is owned by your company and not by yourself, often you will not pay attention because you're not sitting on the billing. Yeah. And also the fact that you have a bunch of other things that are working well, that you might assume it's those things that are amounting for the cost. So one thing that I wanted to mention with spin is that I think a lot of the framing of spin, if people were to pick this up, they might assume that it's only mistakes or errors that produce that spin, but that's not always going to be the case. Like sure, this is sort of an in-between example where it wasn't exactly an error because it was doing something that it was meant to,

32:26but you weren't really paying attention. So it was doing more of it than it needed to. But I think a lot of times spin will also be just ill defining the parameters for a job that you actually do want. An example of this that I had is I had an open claw going for a while that was perpetually researching new data sources in AI that could help us figure out where the state of certain adoption metrics was, right? Every day there's new studies that come out that measure this or measure

32:56that. And that tell you about data readiness or systems integration or use cases or whatever. And it's too much to monitor for humans, but agents are really good at it. And so this open claw agent was a researcher that its only job was to, on a set schedule based on its heartbeat, go out and check for new things. And it was never meant to stop. It was always, it was on a specific schedule, but it basically was this continuous research process that was crawling to the ends of the internet every day. And it ended up just not being valuable enough for the cost, but it was doing what

33:30it was supposed to. And so I think part of the auditing spin is also just figuring out what things have accidentally become spin, even if they started in the right area. And I think that's why this idea of auditing, I think is a good framework because sometimes it's going to be about just updating or changing a process that was valuable as well as catching mistakes. I agree. And even more, I see many automations that people created because they think they will be useful. Like, oh, I can't read. I have too many Slack messages. Let me just create like a Slack miner that runs every hour and reads my entire set of channels. And that can easily become $1,000 in

34:06tokens that literally do a job that moves the needle for nobody. Or the morning brief that you created wholeheartedly with the intention to read it every morning, but for some reason you don't find value and you don't read it. So these are the things that you should definitely audit and kill. And my rule of thumb is if you created an automation and for one or two weeks, you have never, ever used the output, you should definitely kill because it's a definition of a spin. Or if you are using this automation, but there is a very bad proportion between the value of summarizing all of your Slack channels to the bill at the end of the month, that's also something that I consider to be a spin.

34:40And I think that now that everybody gets co-work or GPT work, that even more and more within companies, because it's so easy to build these automations and without sufficient literacy about how to effectively use the tokens, people create a ton of these automations that look good on paper, but don't look so great on the paper of the bill at the end of the month. All right. So let me give you a list for the suspects for silent token spenders. First of all, it's going to be your idle agents and the overfrequent jobs. These are going to be the things that run without any meaningful output or way, way too frequently. We also, in many cases,

35:14see automations that nobody uses. So it can be like the weekly report or the dashboard that nobody ever goes to read. Additional thing can be what we refer to often as the pre-prompt tax. So anything the model runs and reads before the very first prompt. So those include the always-on rules or instructions, the skill definition, the tool definition, and so on. And those can very easily, if not properly organized amount to many thousands of tokens each run without you typing a single word. So those amount significantly. Many folks also hold the immortal conversation, meaning that they

35:49will continue an endless session that keeps carrying old history and old context, often if even creating a poorer quality. We also, in many cases, see users never filtered data retrieval. So instead of just getting 20 rows from a database, they will pull 500 rows or they will process the entire inbox to look for a specific mail that they know what was the subject line and so on. Many other folks will have their context all over the place. So the agent will have to read through a ton of documentation just to understand

36:19what are they talking about and what's the truth here, as well as rework loops. So anytime that your agent or your skill or your just day-to-day usage gets you to do more iterations just to get the same result, this is like just more tokens being spent on nothing. So these are the like immediate suspects. And I want to show you how you can try to potentially identify whether your system is in a spin situation or that the spin to production ratio of tokens is not well formulated. And the thing here is that

36:53not everybody can detect in the same way. Some folks have concrete meter. Those will be people who are using the API version of the models. They have the API console that they can use, or if they are cloud code or cursor users, they have usage view. And of course, people with admin privileges, they have an admin dashboard. So if you are one of those, you can do the following things. One thing that you should definitely do is the weekend test, meaning that if you didn't do anything with AI, but you look at your bill and you see that your bill keeps compounding, you know that there are

37:27things that are adding to your value without, to your bill without any value. That was what happening to me. Also look for very extreme input to output ratio. So agentic work legitimately runs with high ratio, but if you get to a point where it's many thousands to one between input and output, in many cases, that's empty loops. In my open claw case, it was 2,600 to one, which is ridiculous. And if you see that your spend keep rising while the work or the value that you do stays flat, that's also

37:58potentially an indication that you're in a scenario of spin and you need to go and further understand what's the case. However, there are many folks that don't have direct meter because they're not using one of these tools or they don't have the admin privileges, which is probably most of the regular users. For them, you should probably use proxies. So just go directly to list all of your automation and the scheduled jobs that you own and ask which one of them added business value last week. If you don't know, that's a suspect. Then also watch your quota. And if you're burning through your weekly

38:31quota extremely fast, especially if you compare it to other people in your setup or in similar roles, that might be that you're doing something wrong there. And if you are in an enterprise plan, your admin do have the view at least of how much you're consuming and also typically the input-output that you can just ask them. And there are many places that you can look. There are specific like slash context and slash usage in CloudCode. There is also in application visualization now both in CloudCode and in cursor that you can just click on the usage meter and try to understand that.

39:02So regardless of what and how visible it is for you, you should definitely put some caps on how much you spend rather than letting the bill just extend all the time and put some alerts. If there is some kind of a significant jump in how much you consume, that can be an indication that something is up in your system. So that's for identifying spin. And now the habits that we should all adopt to mind our tokens. These are several things that anybody can do immediately that typically improves

39:32the token consumption without reducing the business value. New task is a new session. This one is an

Best Practices for Token Management

39:38interesting one because we will talk in a minute about also model routers, but at least for now, for the most part, be intentional about which model you use for what task. Sometimes it's actually going up to like an opus or even fable class models because they will get the job done in one iteration and overall reduce the spend. In some other cases, it's not doing a web search with fable, but rather going to the haiku or the lower cost of models. Right-size your context. Tell your AI what it needs to know. This is a classical Goldilocks, not too much, not too little, but

40:11sufficient such that it will not go into endless internal reasoning token loops just to try and understand what you're talking about. Build reusable capabilities often when we're just vibing with our model and trying to use it ad hoc rather than sitting down and creating the skills, creating the proper automation, creating the proper agents. We're just wasting a ton of tokens to re-ask the tools to do something again and again. So seeing and building proper systems often is one of the best levers that you have to use their tokens wisely. And filter everything that you can. Tell it in which rows of the table

40:46the data exists, in which parts of the project board the data accounts for, which select channels and so on. The more you point the model to the right place, the better the results that you will get. And lastly, in many cases, we start doing the work. We realize that the model is completely off. Maybe it's the wrong model. Maybe it's missing something. Don't let it spin. Just kill the job early and start again while understanding what you do. And this is one of the cases where looking at the model reasoning will go a long way to understanding that it's completely off in the wrong direction. So I would recommend whenever you send the model to start doing something, especially if it's a

41:19significant portion of work, open the thinking to understand what the model is and understanding from the task that you gave it. And if it seems to be off, stop and improve the instructions rather than letting it air. So that's the habits for everyone. Two additional levers that you should consider, and some of them are very new. So if you are a Cloud Code user, you can use the slash doctor command. This will basically check not only how much like a past installations take on your machine, but also how are your token divided? Whether you have stale skills, stale tool

41:50configuration, whether your overall instructions are overly long or overlapping. So it's a very good command that will be created for us that you can go and execute if you are a Cloud user. If you're not a Cloud user, you can just have your AI tool investigate what the slash doctor command does and basically recreate it for your own tool because it's not like a very complex thing to do. It just audits all of your system for you and gives you a structured report with concrete recommendations of things that you can kill because you haven't run them for a while or things that are duplicated or

42:24stale or contradictory that you can potentially reduce significantly. And with regards to routing, a lot of the industry conversations sit right now around the model routing. And you were just talking, I think today or the other day around some interesting M&A around model routing. Picking the right model is still one of the highest return things that you can do, even if you are able to use like the cursor automated router or some of the other solutions that are coming our way, because it's not always going to be, even if you have like a router in the background, it's not always

42:59going to be as precise as you knowing which model to use. And in many cases, you still don't have in your existing tool, a good enough or even an existing router. Yeah, I think that we are very early in figuring out the right patterns around routing. Obviously, there are a million solutions coming to market. They're all taking slightly different approaches. You have independent experiments from enterprises who are building their own systems that route between, you know, custom models that they've trained as well as, you know, the premier model. Like it is,

43:29there's no one clear approach yet. And even when there do start to be clear use cases and patterns, they may not fit everyone in every use case. I think it would be entirely unsurprising to me, or I expect that routing norms around certain types of software engineering get solved first, because it's more deterministic and clear. And you can kind of actually have, you know, more sort of verified success or not. I think when it comes to knowledge work tasks more broadly, it's going to be immensely more complicated, especially considering how much of our personal

44:05model routing that we do right now is about not what the benchmarks would say on a test, but how we like the particular nature of one type of response versus another for a particular context. So I continue to believe that understanding different model capabilities and having model preferences is still a very high leverage activity and is going to be for quite some time. I agree. And I think the ultimate test was when GPT-5 was automatically routing us and all super users or just like more than occasional users, we were all very frustrated by what we got from the

44:36auto mode. I think that's the original test that we want control. And we will probably even with a great router for many things will continue to be opinionated and rightfully so. Just for people who are also building their own, obviously they have additional levers, like you can create more caching and so on, but still for them, it's much the same physics. Like the more control you have, the more you are able to be smart about the way you use the models. That's the additional levers. We talked about tokens that teach. And I think that up until now, we were very much focused on

45:12things that we can do to reduce the bill. But here I want to fight a good fight and say that we want to protect those tokens because those are in many cases, the tokens that you spend in order to get much better return. And it's not just about optimizing the bill to go downwards, but rather to improve also the return that we're getting. And often to improve the return, we need to improve the tokens that teach. And we're talking about two ways, whether it's you teaching yourself, meaning that you run the same task using three different models in order to get to this taste

45:43of which models you like for each task, or you try the same task in three different ways until you learn which one works best, or you experiment with a new tool, or you try a new skill or a new automation and it doesn't work and you try something else. So all of these typically gets you overall to much better results from AI. So those should be protected firstly, and also the other side of you teaching AI who you are, building the systems, adding more context, such that you will get much more personalized results or much more organizational aware results. Those are almost always with direct

46:17correlation to how much value you get from AI. And so does data. I've seen a study of 20K developers that found that the heaviest AI users were roughly twice as productive in terms of the amount of production code that was shift. So in many cases, it's actually becoming much more like a smart exploratory user will get you to better results. So to summarize, what you need to do in two sides. So for the individual users, these are the things that you should definitely do. Go and see whether you

46:50have tokens that spin. I'm sure that all of us have those idle automations, or maybe some of us have like an even worse scenarios of the amount of tokens being spinned without any business value. Practice those six habits. You can even put them on a post-it and just get yourself to work more effectively with the tokens that you have. Do spend the time to invest in reusable capabilities and improved context that the model can be much more selective and discover the relevant context where it matters. I also want you to audit the things on a schedule, meaning regularly go back

47:23to the system and see what is now stale, or maybe something that was working well has become a stale automation. Maybe you need to improve the context, the instructions. Maybe you can remove some of the instructions per the new advice coming from Entropic that the modern models need fewer instructions, not more. And make sure that you protect the learning budget. And as needed, go and negotiate that with the people responsible for the budget to make sure that you're not now being reduced to the amount of tokens that leaves you with very little room for exploration. For the organizational side,

47:57make the usage visible and then teach the people. Because when managers and employees see their own, they are much smarter about how they use, but make sure that they are not being encouraged to spend as little as possible, but to spend smartly. And also make sure that the budget is by workload and by individuals. If someone is building skills and context and usable capabilities for their entire team, they need to get significantly higher budget than the person that just uses the tool as a extended Google. And all the time, we need to make sure that it's by that you tear it up in some

48:29organizations and some individuals get significantly higher, while others potentially less and not just one size fits all for the entire organization. And make sure that everybody listens to something like that or that you do an internal training that teaches people on how to be smart about tokens, but not how to spend as little as possible, but also how to be mindful about the ROI and aiming to use tokens for the things that move the needle for the company. So that's the concrete actions for you and the team. And if you want to be even more token smart, so beyond the audit of your own usage,

49:04we created for you a token gym that you can go and learn and flex your token smart muscles. And if you want to go even further and to learn how to build and work with AI and agents properly, we do have our existing trainings and the next cohort start on early September. So we'd love to have you there in the executive catch-up or the executive agent leadership that will bring you all the way to be very smart about AI or very smart about agents, depending where you are. That's it. Awesome. You look, I think that this, we're always at the beginning when we're talking about things on

49:39this show, but this one is, I think, particularly inflection pointy, let's say, to use a word that doesn't exist. We are so clearly just at the beginning of figuring out how to organize the relationship between people and the compute and intelligence that they're going to consume. And it is going to be iterative and messy, which is why I think so many of these ideas that you presented are shared as frameworks, you know, patterns to explore, right? It's a set of steps that you can take to try to get

50:12a handle on these problems, but every organization at the beginning is going to solve them or not in different ways. So thank you for sharing some starting points and, you know, we'll continue to evolve this conversation as the tools around us change too.

More from The AI Daily Brief

41 Stats That Tell the Story of AI Right Now

Aug 8, 202622 min

The Right Way to Worry About AI

Aug 7, 202628 min

Google’s AI Leadership Shakeup: Disaster or Exactly What It Needs?

Aug 6, 202633 min

Why the Data Center Fight Has Little to Do With AI

Aug 5, 202635 min

Why AI Washing Won’t Work Much Longer

Aug 4, 202624 min