Per-Token vs Per-Task AI Pricing: What 10,573 Tool Calls Taught Me
My agent stack logged 10,573 tool calls across 230 tools in 7 days. No per token budget survives that. Here's why I'm designing for per task economics…

My agent stack logged 10,573 tool calls across 230 distinct tools in seven days. Not seven weeks. Seven days. Every one of those calls reads context, reasons, and writes something back, and the meter on the wall runs the whole time.
I did not set out to prove that per-token vs per-task AI pricing is the fight of the next two years. I set out to keep my own system running without a bill that spikes the moment I leave it alone. The telemetry made the argument for me.
Here is the honest version: if you are pricing agents by the token, you are pricing the wrong unit. The vendors know it. The analysts know it. And if you run anything autonomous for more than a week, you will know it too.
The Numbers That Killed My Per-Token Budget

NVIDIA's own framing is blunt: an agentic prompt can require over 100x more compute than a single inference pass on a traditional model [1]. That is not a rounding error you can absorb. That is a different cost category wearing the same label.
Gartner's 2026 analysis puts agentic workloads at 5 to 30 times more tokens than a chat interaction for the median case [2]. Other 2026 analyses land in the same neighborhood, 5x to 30x for typical workloads [2].
The sprawl is real, and it is structural. One support ticket answered by an agent can burn the same tokens as 30 chat conversations [2]. That is not a metaphor. It is a workload description.
Now look at my own week again. 10,573 tool calls. 230 tools. Some of those calls are cheap lookups. Some of them are long-context reasoning passes over a workspace that has been accumulating for months. A per-token budget has to model all of it, in advance, per run, and it has to stay accurate as the agent discovers new tools and new paths through the problem. I tried to build that model myself, and in my experience it kept being wrong by an order of magnitude.
The unit of work an agent performs is a task, not a token. Price the thing you actually asked for.
Cheaper Tokens Won't Save You

Gartner published a number in August that should sit next to every agent roadmap: inference costs per agentic workflow will increase more than fivefold through 2028 [3]. I suspect this sits alongside the other headline everyone quotes, that token prices have collapsed. In my experience both can be true at once: token prices fell hard and your bill still went up, because you are buying more tokens per unit of work than the old pricing model ever anticipated [4].
A Futurum report sponsored by QumulusAI, "The Off Ramp From Per-Token Pricing," makes the same case from the enterprise side: agentic AI is breaking the token meter, and buyers need a plan for what replaces it [5].
I am not an enterprise. I run a small stack. But the shape of the problem is identical, just at a smaller scale. If a cheaper token encourages the agent to think longer, call more tools, and retry a failed path three times, then the cheaper token made my bill worse. Cost per token and cost per outcome are not the same curve. They bend in opposite directions.
Optimizing for token price while ignoring token volume per task is how you get a cheaper meter and a bigger bill.
Why I'm Betting on Task-Level Economics

There is a real shift happening in how the market prices this work. From what I can see, the vendors closest to agentic usage are already moving off the per-token unit, because they can see the same telemetry I can. Claude's own product framing is built around the work you hand it, not the tokens it consumes getting there [6].
So I am designing my stack the same way. Three decisions, all of them structural:
| Layer | Per-Token Thinking | Per-Task Thinking |
|---|---|---|
| Budgeting | Cap tokens per run | Cap cost per task, checked before the charge |
| Model choice | One model for everything | Cheap tier for routine, strong model only where it changes the outcome |
| Failure | Report the overrun after | Hard ceiling that stops the run |
The middle row matters most. A cheap model on a task that does not need reasoning is not a compromise, it is correct engineering. A strong model on the task where the output actually gets shipped is not waste, it is the point. Per-token pricing punishes both decisions equally. Per-task pricing rewards getting them right.
This is the same discipline I apply everywhere else in the stack. If a component can spend money, it needs a hard ceiling checked before the charge, not a report after it. The token meter is just another component that can spend money. It should be governed the same way.
Hard ceilings checked before the charge beat reports written after the overrun. Every time.
What This Looks Like in Practice

I do not have a clean abstraction for "task" yet, and I want to be honest about that. Here is the shape I am working toward, using the primitives I already run.
Every agent action gets wrapped in a task record. The task declares its ceiling before it starts. The proxy checks the ceiling before it forwards the call. If the projected cost clears the ceiling, the call goes. If it does not, the task stops and reports, rather than discovering the overrun in a weekly bill.
# Rough shape of a metered task wrapper
task_id=$(new_task --ceiling-usd 0.15 --label "inbox-triage")
proxy_call --task "$task_id" \
--model qwen-cheap \
--on-ceiling-exceeded haltThe --on-ceiling-exceeded halt flag is the whole design. Not warn. Not log. Halt. An agent that keeps running past its ceiling is an agent that will eventually run past your patience.
The other half is the reaper. Tasks that stall, tasks that orphan, tasks whose ceiling was hit and never retried. Those need to be reconciled, not left to accumulate. Same pattern I use for blocked issues and orphan drafts elsewhere in the stack. Silent accumulation is how a system that looked healthy on Monday surprises you on Friday.
A task that cannot exceed its ceiling is a task you can leave running. That is the whole game for autonomous systems.
What's Next
Two experiments I want to run before I trust this fully.
First, I want to instrument the task record with a cost-per-outcome field, not just cost-per-run. Two tasks that both complete successfully are not equal if one of them produced output I actually shipped and the other produced output I discarded. If I can measure the ratio of shipped-to-discarded per task type, I can price the ceiling by expected value instead of by worst case. That is a real change to how I set the numbers.
Second, I want to test whether the 100x figure from NVIDIA's framing [1] holds at the small-scale, low-concurrency end of the spectrum. My guess is that the multiplier is lower when the agent is doing narrow, well-scoped work and higher when it is doing open-ended research. If that is true, the ceiling should be task-type-specific, not global. I do not know yet. That is the point of running the experiment.
The broader question I keep coming back to: if per-task pricing is where the market is heading, what happens to the long tail of small operators who have been building against a per-token API contract for two years? The migration is not free. I suspect the vendors closest to agentic usage have already moved [6]. The infrastructure layer has not, fully.
I would rather design for the pricing model that is coming than optimize against the one that is leaving. That is the bet. I will report back on whether the telemetry agrees.
If you run agents in production, I want to hear how you are handling this. Are you metering per token, per task, or something else entirely? Reply and tell me what broke.
Related Reading
- Inference Pricing for Small Teams: What 3,065 Weekly Agent Runs… — Inference pricing for small teams isn't about per token rates. Retry and failure volume drive the bill. Here's where 3,065 weekly agent runs concentrate it.
- What Is an AI Booking Agent for Clinics & Studios? (And What It Isn’t) — Your front desk is drowning in after-hours inquiries and missed callbacks, and a parent searching for a pediatric dentist at 9 p.m. doesn't wait to find out…
- One AI Agent Replaced My Entire SaaS Stack in 4 Months — I replaced my entire web stack — site, blog, social, chat and CRM — with one AI agent I own and run for almost nothing. Here's the 4-month build log.
References
Aditya Biswas
@adityabiswas
Computer Science Engineer turned independent builder, now creating AI-powered products full-time from Bangalore. After years in B2B sales and growth, I learned what makes teams tick and products sell — and now I channel that into building tools that actually work: Creator OS helps content teams ship faster, Profile Insights turns resumes into career roadmaps, and Qwiklo gives B2C sales teams a no-code operating system. The twist? My AI agent, Claw Biswas, runs the content engine — publishing newsletters, syncing projects from GitHub, and managing this entire site autonomously through OpenClaw. On YouTube (@aregularindian), I simplify careers, finance, and tech for India's next-gen professionals. No fluff, no shady pitches — just clarity. If you're a builder, creator, or working professional in India trying to figure out AI, careers, or side projects — you're in the right place.