AI projects often begin with a small bill. Then usage grows, prompts become longer, agents make several calls to complete one job and the most capable model often gets assigned to every task, even ones a smaller model could handle. Many teams will attempt to mitigate the costs of this by chasing token price, but that's actually backwards.
The obvious response is to look for a cheaper price per million tokens. That can help, but it does not tell you what the business pays to complete useful work.
A better metric is cost per completed task. It measures the full model spend required to produce an output that meets the business's success criteria. It accounts for token prices while also capturing context length, reasoning, repeated calls, retries and model reliability.
This guide explains how to calculate that cost and reduce it using open models, including DeepSeek V4.1 Flash, DeepSeek V4 Pro, Kimi K3 and GLM 5.2, which are available through Aurora Inference.
Why Your AI Bill Is Higher Than the Token Price Suggests
API pricing is usually quoted per million input and output tokens. That is only the first layer of the calculation.
The actual cost of an AI workload can include:
- input tokens from the user request, system prompt and retrieved context;
- output tokens, including visible answers and billable reasoning tokens;
- repeated context sent across several calls;
- cache reads and cache writes;
- retries after invalid, incomplete or low-quality outputs;
- search tools, external APIs and retrieval;
- human review and correction; and
- infrastructure costs for self-hosted deployments.
An AI agent makes this gap more pronounced. One user request might trigger a planning call, three tool calls, two verification calls and a final response. The business is paying for the full sequence, not the price of the final answer alone.
This gives us two different metrics:
- Cost per request = total model and tool cost ÷ total requests
- Cost per completed task = total model and tool cost ÷ tasks that meet the success criteria
The second figure connects model spend with an outcome such as a resolved support ticket, a validated extraction or a passing code change.
How Much Do Leading Open Models Cost?
Hosted open-model prices can differ substantially by provider, service tier, region, quantization and caching policy. They can also change quickly.
The following table uses a public hosted-inference pricing snapshot from September 2026 to show the size of the difference. It is an external reference point, not Aurora Inference pricing. Teams should replace these rates with the current prices in their own provider account before making a purchasing decision.
|
Model |
Input per 1M tokens |
Output per 1M tokens |
Difference in positioning |
|---|---|---|---|
|
DeepSeek V4.1 Flash |
$0.09 |
$0.18 |
Low-cost, high-throughput model |
|
DeepSeek V4 Pro |
$1.30 |
$2.60 |
Larger model for difficult reasoning and agentic work |
|
Kimi K3 |
$2.85 |
$14.25 |
Large long-horizon and multimodal agentic model |
|
GLM 5.2 |
$0.488 |
$1.56 |
Mid-priced long-context model |
Source: AuroraInfra's hosted model pricing, accessed September 2026. Rates vary by provider and change frequently, including our own.
The range is large. At these reference rates, Kimi K3 output tokens cost more than 79 times as much as DeepSeek V4 Flash output tokens. That still does not prove that Flash is 79 times cheaper for a production task.
Models can use different numbers of reasoning tokens and produce different success rates. The cheapest token can become expensive if the workflow needs several attempts.
Price Per Token vs Cost Per Task
Consider a business that summarizes customer-support conversations for quality review. Assume each task uses:
- 8,000 input tokens;
- 1,000 output tokens; and
- one model call per task.
Using the public rates above, the initial cost estimate is:
|
Model |
Cost per attempt |
10,000 attempts |
100,000 attempts |
1 million attempts |
|---|---|---|---|---|
|
DeepSeek V4.1 Flash |
$0.00090 |
$9.00 |
$90.00 |
$900.00 |
|
DeepSeek V4 Pro |
$0.01300 |
$130.00 |
$1,300.00 |
$13,000.00 |
|
Kimi K3 |
$0.03705 |
$370.50 |
$3,705.00 |
$37,050.00 |
|
GLM 5.2 |
$0.00546 |
$54.64 |
$546.40 |
$5,464.00 |
These are illustrative calculations. They hold token usage constant across models and exclude caching, reasoning-token differences, failed calls, tools and human review.
Now add quality. Suppose a low-cost model completes 70% of tasks acceptably on the first attempt, while a more capable model completes 95%. The expected model cost of one successful task can be approximated as:
Expected cost per successful task = cost per attempt ÷ success rate
A $0.004 attempt at a 70% success rate costs about $0.0057 per expected success. A $0.007 attempt at a 95% success rate costs about $0.0074. The first model is still cheaper if retries are automated and carry no other penalty.
But add $0.05 of human review whenever a task fails:
- Model A: $0.004 + (30% × $0.05) = $0.019 expected cost per task
- Model B: $0.007 + (5% × $0.05) = $0.0095 expected cost per task
The more expensive request is now the lower-cost business choice. Failure costs can be much higher when an error delays a workflow, triggers an expensive tool or reaches a customer.
Five Ways to Reduce AI Costs Without Sacrificing Quality
1. Match the Model to the Task
Do not route every request to the most capable model by default. Classification, tagging, straightforward extraction and short summaries may work well on a smaller, lower-cost model. Complex repository changes, multi-step analysis and difficult agent tasks may justify a stronger model.
A routing system can classify the request and select an appropriate model. It can also start with a lower-cost model and escalate when confidence is low or validation fails.
For coding workloads, the right choice depends on whether the task involves short code generation, repository modification or long-horizon agentic work. See The Best Open Source LLM for Coding: Benchmark & Comparison [2026] for a task-specific comparison of the four Aurora-supported models.
2. Reduce Unnecessary Context
Input tokens often grow quietly. Long system prompts, complete conversation histories and large retrieved documents may be included in every call, even when only a small portion is relevant.
Reduce context by:
- removing duplicated instructions;
- summarizing older conversation turns;
- retrieving smaller, more relevant document sections;
- filtering low-confidence retrieval results;
- storing structured state instead of replaying full transcripts; and
- sending only the fields required for the current task.
Context reduction can lower cost and latency while giving the model fewer irrelevant details to process.
3. Use Prompt Caching
Prompt caching reduces the cost of repeated input. It is most useful when many requests share a stable prefix, such as:
- a long system prompt;
- product documentation;
- company policies;
- a codebase snapshot; or
- standard agent instructions and tool definitions.
Caching depends on the provider's implementation. Check which tokens qualify, how long the cache remains active and whether cache writes have a separate price. Then monitor the cache-hit rate.
4. Reduce Failed Calls and Retries
Retries are often treated as a reliability concern, but they are also a cost multiplier. A workflow that averages 1.4 calls per completed task pays roughly 40% more in model charges than its per-request estimate suggests, assuming similar token usage per call.
Reduce retries by:
- using structured output schemas;
- validating inputs before inference;
- giving the model clear acceptance criteria;
- setting realistic token limits;
- adding deterministic checks for factual or formatting requirements; and
- routing difficult cases to a more capable model earlier.
Track failure reasons separately. Formatting failures need a different fix from missing context or weak reasoning.
5. Route Flexible Workloads to Lower-Cost Capacity
Some providers offer batch, flex or off-peak capacity at a lower price. Non-urgent workloads can be scheduled around those options.
Good candidates include offline summarization, enrichment, evaluation runs, embedding backfills and large content migrations.
The trade-off may be slower or less predictable completion time. Keep customer-facing and latency-sensitive requests on capacity with appropriate availability commitments.
When a More Expensive AI Model Can Actually Cost Less
A higher-priced model can lower total cost when failure is expensive.
This is common in workflows that involve:
- human review or manual correction;
- paid search, browser or database tools;
- several dependent agent steps;
- customer-facing responses;
- code execution or infrastructure changes; or
- compliance and financial decisions.
Suppose an agent makes five model calls and uses a paid data tool. If the model fails at the final step, the business may need to repeat the full sequence. Paying more for a model that plans and verifies more reliably can reduce total calls and wasted tool charges.
The decision should be based on measured task economics. Run both models on the same representative dataset. Record token use, total calls, tool charges, latency, success rate and human-review time. Then compare the cost per accepted output.
For a complete evaluation framework, see How to Evaluate Open-Source Models for Production Inference.
Open Models vs Frontier APIs: Where Are the Biggest Savings?
There are three common ways to deploy an LLM:
- use an open model through a hosted API;
- self-host an open model; or
- use a proprietary frontier API.
Hosted open models can reduce API costs without requiring a company to operate GPU infrastructure. Self-hosting provides more control and can be economical at high, predictable utilization. At low utilization, idle GPUs, engineering, monitoring and scaling can outweigh the apparent savings.
Frontier APIs may be worth their premium when they improve task completion, reduce tool usage or eliminate human review.
The right comparison is therefore hosted open model versus self-hosted total cost versus frontier API cost per completed task. Comparing a hosted token rate with the hourly price of one GPU does not capture the operational difference.
How to Calculate Your Own AI Cost Per Task
Start with monthly model cost:
Monthly model cost = (uncached input tokens × input rate) + (cached input tokens × cache rate) + (output and reasoning tokens × output rate)
Then add:
- retries and additional model calls;
- search, database and external API charges;
- embeddings, storage and reranking;
- infrastructure and observability; and
- human-review cost.
Finally calculate: Cost per completed task = total monthly AI cost ÷ successfully completed tasks
Collect these figures for each production workflow:
|
Metric |
Why it matters |
|---|---|
|
Requests per month |
Establishes workload volume |
|
Average input tokens |
Captures prompt and context cost |
|
Average output and reasoning tokens |
Captures generation cost |
|
Cache-hit rate |
Measures whether caching is working |
|
Model calls per task |
Reveals multi-call and retry overhead |
|
Tool cost per task |
Captures non-model variable cost |
|
Success rate |
Connects spend with accepted outcomes |
|
Human-review time |
Captures a major hidden cost |
Recalculate after every routing, prompt or model change. A lower monthly bill is useful only if task quality and completion remain acceptable.
Aurora Inference provides access to DeepSeek V4 Flash, DeepSeek V4 Pro, Kimi K3 and GLM 5.2 through one platform. That gives teams the flexibility to match different open models to different workloads instead of paying one model to handle every task.
Compare open models on Aurora Inference
Frequently Asked Questions
Does a cheaper model always cost less to run in production?
No. A lower token price can become expensive if the model fails more often, requires retries, or produces outputs that need human review. The cost per completed task — not the token price — determines real savings. A $0.007 attempt at 95% success can cost less per accepted output than a $0.004 attempt at 70% success, especially when human review is expensive.
How do I measure cost per task for my workflow?
Collect these metrics for one month: total model cost, number of completed tasks that met your success criteria, failed attempts and retries, tool charges, and human-review time. Then divide total cost by successfully completed tasks. Repeat after every routing, prompt or model change to see the real impact.
When should I pay more for a more capable model?
When failure is expensive. This includes workflows with human review, paid API calls, multi-step agents, customer-facing responses, code execution or compliance decisions. Measure both models on the same dataset: compare token use, total calls, tool cost, success rate and human-review time. The higher-priced model often wins on total cost per task.
What's the difference between "open source" and "open weight" models?
Open weight means the model weights are available. Open source typically includes weights, code and documentation under a permissive license. For hosted access via an API, the distinction matters less than the provider's terms and data policies. For self-hosting or fine-tuning, licensing is critical.
How much can I save with prompt caching?
That depends on your provider and workload. Caching is most valuable when many requests share a stable, long prefix: system prompts, documentation, policies or large codebases. Monitor your cache-hit rate and compare cached vs. non-cached token costs to measure real savings. Some providers charge a separate rate for cache writes.
Is self-hosting an open model cheaper than hosted APIs?
Not always. Self-hosting can be economical at high, predictable utilization. At low utilization, the cost of GPUs, engineering, monitoring and scaling often exceeds the token savings. A fair comparison includes hosted API cost per task versus self-hosted total cost per task, not just token rates versus hourly GPU cost.
Which open model should I use?
It depends on your task. DeepSeek V4.1 Flash is the lowest-cost option and works well for classification, simple extraction and short summaries. V4 Pro is stronger on reasoning and agentic work. Kimi K3 excels at long-horizon and repository-level tasks. GLM 5.2 offers a permissive license and large context window. Test on your own workload before committing.
Sources
_.webp?width=150&height=56&name=Aurora_Logo_Brand(DARK%20THEME)_.webp)