Open-weight models through
an OpenAI-compatible API.
Aurora Inference runs your open-weight models for less, on reserved or metered capacity.
Economics you can plan around, without hiring the team that usually comes attached.
Not Every Call Needs Your Biggest Model
Aurora's catalog spans that range on a single OpenAI-compatible endpoint, from a fast instruction-tuned model to a long-horizon reasoning model, with more than an order of magnitude between what their calls cost. Same key, same API, one bill. Moving a call to a cheaper model is a string change, not an integration.
Cached input is priced well below list, with the exact rates on the model list below. So the two levers that move your bill are both in your hands: how much context you re-send, and which model you send it to.
Start metered. Move to reserved when your usage has a shape worth committing to.
Same Models. Same API.
The Difference Is What You Can See.
Cached input, priced down.
Agent loops re-send the same context every turn. Cached input is priced well below list across the catalog, so the tokens you repeat cost a fraction of the ones you send once.
What you see is what you get.
Rates and context windows are published on the model list, not behind a sales call.
Your boundary, by construction.
On reserved capacity, prompts, outputs and weights stay in your region. Private AI as a property of the deployment, not a clause in a contract.
Metered now, reserved when it pays.
Start on the metered endpoint with no commitment. Move to reserved when your volume has a shape worth committing to, without changing your integration.
What You Get on the Endpoint.
| Models | DeepSeek V4 Flash, DeepSeek V4 Pro, Kimi K3, and GLM 5.2. You pick the model per call, so simple work stops running on reasoning-priced models. |
| Integration | OpenAI-compatible endpoints: base URL and key swap. Test in staging against your own evals, with no re-architecture and no trace handover. |
| Pricing | Metered per token from the first call. No commitment and no minimum, with rates published per million per model in the console. |
| Caching | Agentic workloads repeat context heavily. Cached input is priced well below list, so the context you re-send costs a fraction of what you pay to send it the first time. |
| Context | Long context across the catalog, up to roughly one million tokens. Per-model windows are published alongside the rates. |
| Reserved | Available when your volume has a predictable shape, including in-region deployment. Sales-led, and it does not change your integration. |
Frequently-Asked Questions
Sign up, generate a key, change your base URL, and send one call. Then point a real workload at it and compare against your own evals.
Capacity That Learns Your Workload.
What it does for the engineer
Continuously evaluates placement, parallelism strategy, batching, cache policy, and precision against your live traffic patterns. Aurora Inference reconfigures as the workload drifts, instead of serving yesterday’s tuning to today’s traffic.
What it means for everyone else
You’re not buying a machine that ages. Every month on Aurora Inference, your capacity does more useful work per dollar than the month before without anyone on your team touching it.
Your Inference Bill Isn’t a Price Problem. It’s a Tuning Problem.
Inference is becoming the largest variable line in AI COGS, and the market's answer is a discount on the meter. That treats a control problem like a procurement problem, and we've laid out why cheaper tokens don't fix the economics with the sourced numbers.
The real lever is how much value each dollar of capacity produces through routing, caching, and utilization. That lever is operated, not negotiated. Aurora operates it for you, on infrastructure whose costs we control, so your spend becomes something you can plan and your margins something you can commit to.
Metrics that move: cost per request, gross margin %, budget variance vs. plan.

Built for the Teams Feeling It First.
Agentic AI in Production
When token bills climb faster than the value your agents produce, the fix is cheaper calls where the work is simple and cached context where the loop repeats. Both are yours to set on Aurora Inference.
Migrating Off Frontier APIs
Open-weight models now barely trail the frontier by months and they do it at a fraction of the cost. It's where the builders have already landed. Aurora runs multiple models behind an OpenAI-compatible endpoint, so moving a workload over is a base URL and a key.
Voice and Real-Time Workloads
Inference deployed near the traffic, so the latency budget goes to the model with per-minute economics you can plan.
Private AI and Data Residency
Regulated traffic runs in-region: reserved processing stays in the jurisdiction, not just storage. If you're navigating data residency under GDPR and the EU AI Act, that's a requirement Aurora meets by construction.
Start It. Test It. Scale It.
No migration project, no trace handover. You judge everything with your own evals.
Start
Sign up, generate a key, change your base URL. Metered from the first call, with no commitment and no minimum.
Test
Point staging at the OpenAI-compatible endpoints and judge with your own evals. Start with one workload, in shadow mode.
Scale
Move calls to cheaper models where the work is simple. When your volume has a predictable shape, talk to us about reserved capacity in your region.