TokenMaxxing was always the wrong question. Here's the right one.
Most organisations are failing at AI because they're measuring the wrong things entirely.
We have all been there:
A student trying to turn a 2-page essay into a mandatory 10-page paper by changing the font size, widening the margins, and writing bloated sentences like “Due to the fact that the historical context implies...” instead of just saying “Because...”.
There’s a concept circulating in tech and data circles right now called TokenMaxxing. The idea is essentially: measure how much your teams are using AI by the volume of tokens consumed. More tokens in, more tokens out, more productivity. Sounds simple, right?
I will explain why this is one of the more dangerous ideas I’ve seen associated with AI and why it’s only leading to wrong things.
Measuring tokens is measuring effort, not outcomes
Good product building practices define one simple yet difficult to absorb fundamental: measure outcomes, not outputs.
TokenMaxxing reminds me of something that should have died in software engineering decades ago: measuring developer productivity by lines of code written. Funnily, many top companies used this metric for the longest time. It sounds logical until you think about it for five seconds. A junior engineer might write two hundred lines to do what a senior engineer does in twenty. More code, less value.
Tokens work the same way. A team generating enormous volumes of AI output is not necessarily doing good work. They might be doing the opposite as some examples below show.
The consultant at Deloitte or EY who fed a client deliverable through an AI and got back a report full of fabricated citations, non-existent books, and hallucinated statistics must have generated plenty of tokens.
The Air Canada chatbot that promised a grieving passenger a bereavement fare that didn’t exist was very productive by token standards.
McDonald’s AI order taker that famously couldn’t process a simple modification had no shortage of output.
Token volume is a measure of activity. It tells you nothing about whether anything useful happened. Strangely, many of the sharpest minds seem to be missing it!
The lazy patterns are everywhere
The most insidious thing about TokenMaxxing is that lazy AI use doesn’t announce itself. Rather, it just quietly embeds itself into quality work while looking productive on a dashboard.
I’ve watched it happen across functions:
A Marketing team produces blog posts that technically cover the right topics but sound like they could be about any company in any industry.
A Product Manager pastes a one-line feature request into an AI and gets back a ten-page PRD that feels complete until you notice it has no real persona, no edge cases, no actual technical constraint.
A Sales team uses an AI insight tool, trusts the competitive analysis without touching the competitor’s product themselves, and walks into a demo with a fundamentally wrong read of the market.
The worst version I’ve seen: performance reviews written by AI from generic project metrics. No real mentorship in them. No honest feedback. Even scarier: seems to be becoming a norm.
In all these cases, the token count was great. The value was close to zilch.
What should actually be measured
The reason organisations fall into the token trap is that real AI outcomes are harder to count. But hard to count is not the same as impossible.
In Product, track whether AI-assisted work is actually shipping features that users engage with. Not features that shipped, features that landed (i.e., more engaged users, better feedback loops, better retention - BETTER, BETTER, BETTER) . If AI is helping your team move faster but you’re shipping to dormant users, you haven’t found productivity but acceleration toward the mature end of your product cycle.
In Sales, track whether teams using AI close faster or convert at a higher rate. Not whether they used the tool. Whether it moved the number. Simple.
In Customer Support, track First Response Time, Average Handle Time, and CSAT. These aren’t novel or glamorous metrics, but they’re honestly tied to business outcomes. They tell you whether customers are getting better help, not whether the team is generating more text.
The framing should not be “did we use AI?”
It should be “did using AI make the outcome better than it would have been otherwise?”
That’s a harder question. But it’s also the only one worth asking.
Shift in AI thinking - a sparring partner, not a ghostwriter
Here’s the shift in thinking that separates teams using AI well from teams running up token counts.
The good use cases I’ve seen treat AI as adversarial by default. One pattern that genuinely works imho: upload a finished product strategy or feature proposal, then prompt the AI to act as a deeply sceptical principal engineer and a cynical CFO simultaneously. Ask it to find the hidden technical risks. Ask it to identify where you’re wasting money. Ask it where the plan assumes things that aren’t true.
That use of AI generates maybe the same amount of tokens. But tt’s also worth enormously more.
The same logic applies to persona work. Using AI to process large volumes of qualitative research, find patterns, and surface themes is genuinely powerful. But then you need a human to sit with a real customer, verify the persona, and notice the things that don’t fit the pattern. AI gives you the hypothesis. The PM has to stress-test it in the field.
One without the other produces either a beautiful fiction or an overwhelming pile of raw material. You need both.
Who will actually get value from AI
The companies that will actually get value from AI aren’t the ones with the highest token counts. I think that’s probably more than clear already from what we are seeing right now.
Some Fortune 500 companies have burned through the entire AI budget in about a quarter and haven’t shipped anything meaningful to business.
The ones that will actually get value out of AI are the ones that have thought clearly about what good looks like in their context, and built habits around catching where AI is wrong.
Measuring token consumption and calling it an AI strategy is the equivalent of measuring how many hours your team stared at spreadsheets and calling it a finance strategy.
The question isn’t how much AI your organisation is using. It’s whether it’s making anything better.
If you’ve seen this play out in either direction - AI that genuinely improved an outcome, or AI that made something worse while looking productive, I’d love to hear about it too!


