There is a number showing up in AI reviews that everyone has quietly agreed to be impressed by: tokens consumed. Usage is up and to the right. The team ran millions of tokens through the model this quarter, more than last. Someone puts it on a slide next to the productivity story. Hours saved. Tasks automated. Reports written faster. The implication is that heavy usage is heavy value, that a busier model is a more useful one.
It is the wrong metric. More tokens does not imply more ROI, and treating consumption as a proxy for value is how the bill climbs while the return stays flat.
We've been measuring enterprise AI on a copilot's terms. How fast it types. How many emails it drafts. How quickly it spits out a summary. Those are useful metrics if you sell a copilot. They're a category error if you're trying to run a business.
A great account manager isn't valued for how fast they answer questions. A great inventory analyst isn't valued for how quickly they pull a report. A great CFO isn't valued for how many forecast reports they crank out. They're valued for spotting the thing nobody else saw, raising the flag before anyone asked, getting ahead of a problem that hadn't happened yet. That's the work that moves a business.
So measuring AI by productivity is like measuring a CFO by sheer volume of forecast reports. You'll get a number. You won't be measuring what matters.
Measuring AI by productivity is like measuring a CFO by how many forecast reports they crank out.
The consumption trap
Here is how the wrong metric takes hold. Tokens are easy to count and they always go up, so they become the number people report. High usage gets read as high adoption, and high adoption gets read as value. The chain sounds reasonable, but usage is not adoption, and adoption is not value. Consumption measures how much work the model did, not whether any of that work was worth doing.
Customer success is where this shows up most clearly. Point an agent at your CS org and it will happily burn through millions of tokens: summarizing every ticket, drafting a reply to every message, writing a health note on every account. Work that used to take a rep hours now takes minutes. The usage chart looks incredible. But most of that work did not need doing. Summarizing a ticket nobody was going to act on is fast, cheap per token, and worth nothing. Speed on the wrong task is not ROI. It is a bigger bill.
The rep who matters is not the one who answers the most tickets. It is the one who spots the account about to churn and does something before renewal. Same idea for the agent. The question is never how many tokens it spent. It is which accounts it caught, and what that was worth.
The trap gets worse when the metric itself is the wrong one. Point AI at a customer success team and count how many emails each CSM sends, and you will get more emails. You will not get more renewals. Emails sent is a vanity number. The questions that decide the quarter are whether the renewal rate moved, and whether the customer is getting more measurable value out of what they already pay you for. Track those, and most of the email volume turns out to be motion, not progress.
| Same CS week | Consumption play | Right task |
|---|---|---|
| Tokens consumed | 4.1M | 0.2M |
| What the model did | 12,000 tickets summarized | 9 churn-risk accounts caught |
| Time per task | hours to minutes | hours to minutes |
| ROI attributed | None traceable | $84k renewals saved |
This is the test. If you cannot tie a block of tokens to an outcome it produced, in dollars, then you are not measuring ROI, you are measuring your own spend. Attribute revenue per token or you are doing it wrong. Every action Cimba takes is logged against the outcome it produced, so the question “what did this consumption buy us?” always has an answer.
The metric that does matter
The question worth asking isn't “how much time is your AI saving people?” It's “how much is your AI seeing before people see it themselves?”
Call this proactivity. The willingness and ability of an AI system to surface what you need to know before you even have to ask.
In a managed marketplace, that's the agent that flags the merchant whose orders have been slipping, before the AM's quarterly review notices. It's the system that sees a stock-out coming a week out, while the inventory team is still working on yesterday's restock. It's the workflow that notices delivery SLAs softening in one zone before the morning standup, instead of after.
In finance ops, it's the controller's agent that catches the variance trending the wrong direction in mid-month, not at close. It's the FP&A workflow that surfaces the assumption that's about to break the model, before the CFO asks why the forecast is off.
None of these are productivity wins. They're judgment wins. The AI didn't help someone do their job faster. It helped the business avoid a problem or catch an opportunity that would have been missed entirely.
That's the work enterprise AI should be measured on.
Why the market got this wrong
The productivity framing isn't an accident. Most of the AI tools that landed in enterprises over the last two years were copilots, and copilots are productivity instruments by design. Copilots wait to be asked. They sit at someone's elbow and respond. The metric that fits that posture is “how much faster did the asking and answering go.” Which is fine, as long as the question being asked is the right one.
This same dynamic is sometimes framed as reactive AI versus proactive AI. That terminology is accurate, but the cause runs deeper.
Reactive isn't a design flaw. It's what productivity metrics produce.
Reactive
The tool waits to be asked
Proactive
Cimba surfaces the Next Best Action
The problem comes when copilot metrics get applied to systems that should be doing something different. When an AM has fifty merchants, the question isn't “can I answer queries about each one faster.” It's “which of these fifty should I be paying attention to right now, and why.” A copilot can't answer that. Nobody asked. A proactive system can.
Same thing in delivery ops. A copilot can write a faster report on yesterday's SLA misses. A proactive agent can tell you which zone is about to miss tomorrow's, while there's still time to do something about it.
The copilot model isn't broken. It's just one chapter in what enterprise AI is supposed to be.
What changes when you measure proactivity
Two things, mostly.
The buying criteria shift. You stop asking vendors how fast their AI is, and start asking what their AI is monitoring on your behalf. What signals it's watching for. What it does when it sees one. Whether it surfaces the thing without being prompted, or just answers faster once you finally ask.
Four questions to ask any AI vendor
- 1Monitor. What is your AI watching for on my behalf?
- 2Detect. What signals does it surface as soon as they appear?
- 3Act. What does it do when it sees one?
- 4Initiate. Does it surface things without being prompted, or just answer faster once asked?
The operating model shifts too. Teams stop using AI as an on-demand tool and start treating it as a continuous service. Instead of asking the agent a question every morning, the agent has already pushed three things to your queue overnight, ranked by what you need to act on first. Your morning isn't “what should I look at.” It's “what action do I take on what already surfaced.”
We call that surfaced item the Next Best Action. Not just an alert, but a specific recommended move with its projected impact attached, routed to the person who can approve it. And because the action is named, the outcome it produces can be measured against it. Proactivity stops being a feeling and starts showing up in the top line. Here is one, end to end:
08:1208:1208:1914:30One action, one measured outcome. Compounded across hundreds of actions a week, that is how ops teams reach a real number, on the order of +20% revenue per customer. Not hours saved. Revenue you can attribute.
That's not faster. That's different. And it's what business operations and finance teams actually need.
Productive is the floor, not the ceiling
None of this is a knock on productivity. Faster is better than slower. Saving hours is better than wasting them. If your AI is making your team productive, good. Take the win.
But productivity is table stakes now. Anyone can put a copilot in front of a team or automate a slice of a workflow in an afternoon, so speed on its own stopped being a differentiator. A generic copilot is built to make a faster horse: the same job, done quicker. That is a real convenience, and it is not the thing an operations team is short on. Ops does not need a faster way to answer questions nobody prioritized. It needs the system to decide which questions are worth answering, act on them, and show the result in the top line. A faster horse never gets you there.
Just don't confuse productive with proactive. They aren't the same thing, and the gap between them is where most of the value lives. Ask the five questions you'd ask of any governed AI workflow, then ask one more: what does this system tell me before I ask?
Proactive is just one of four qualities we think enterprise AI should be measured on. The full set is PACT: proactive, auditable, consistent, trusted. We'll come back to the other three in future posts.
The best ops and finance teams aren't the ones who answer questions the fastest. They're the ones who see things first.
Your AI should be measured the same way.
Cimba is proactive AI for enterprise business and finance operations. If you're tired of paying for AI that waits to be asked, book a demo.
Evaluating enterprise AI?
Ask us the hard questions.
