Pricing
Pay for the
tokens you use.
Every request is metered to the token. We charge the provider's per-million rate plus 25% (30% on xAI and DeepSeek) — and we show you both numbers, every time. No subscriptions. No seat fees. No minimum.
Top up
Buy credits to fund your account ($1 = 1 credit)
Make requests
Each call charged: tokens × provider rate × 1.25 (×1.30 on xAI/DeepSeek)
See the cost
Real provider cost + our margin, on every response
Same per-token rate regardless of pack size. The packs are just top-up amounts — your real cost is per-call. Credits never expire.
Per model
What does $25 get you?
All prices below include our margin — 25%, or 30% on xAI and DeepSeek. The "$25 buys" column shows input tokens you can send for $25 at that model's rate.
| Model | In $/M | Out $/M | $25 buys (input) | Best for |
|---|---|---|---|---|
Claude Fable 5 flagship 1M context | $12.50 | $62.50 | 2M input tokens | The frontier. Use it when it has to be right. |
Claude Opus 4.8 1M context | $6.25 | $31.25 | 4M input tokens | Frontier quality at half of Fable's rate. |
Claude Sonnet 5 new 1M context | $3.75 | $18.75 | 7M input tokens | The workhorse. Best balance of speed + quality. |
Claude Haiku 4.5 200K context | $1.25 | $6.25 | 20M input tokens | Reviews, formatting, quick lookups. |
GPT-5.5 1M context | $6.25 | $37.50 | 4M input tokens | OpenAI's frontier. Try it next to Fable 5. |
GPT-5.4 1M context | $3.13 | $18.75 | 8M input tokens | Solid GPT-5 default. |
GPT-5.4 Mini 400K context | $0.94 | $5.63 | 27M input tokens | Budget GPT-5 — small tasks, sub-agents. |
Grok 4.5 new 500K context | $2.60 | $7.80 | 10M input tokens | xAI's new coding model — agentic software work. |
Grok 4.3 1M context | $1.63 | $3.25 | 15M input tokens | Fast general reasoning + structured outputs. |
Grok 4.1 Fast 2M context | $0.26 | $0.65 | 96M input tokens | Massive-context (2M) work at budget rates. |
Gemini 3.5 Flash value 1M context | $1.88 | $11.25 | 13M input tokens | Google's newest — near-frontier at half the price. |
Gemini 3.1 Pro 1M context | $2.50 | $15.00 | 10M input tokens | Google's multimodal frontier. 1M context. |
Gemini 3.1 Flash-Lite 1M context | $0.31 | $1.88 | 81M input tokens | High-volume pipelines at minimal cost. |
DeepSeek V4 Pro value 1M context | $0.57 | $1.13 | 44M input tokens | Top open-weight reasoning under $1.20/M out. |
DeepSeek V4 Flash value 1M context | $0.18 | $0.36 | 139M input tokens | Cheapest model on the platform. Period. |
Cost = (input tokens × in $/M) + (output tokens × out $/M). Output tokens are typically 2–6× more expensive than input on most providers.
Plus image, video, voice, and TTS models priced per asset. Full table in the docs.
Compute
Run code, not just prompts
The assistant can build and test in an isolated cloud sandbox — full toolchain, destroyed after each run — or on your own machine, where your code never leaves your computer. Compute is metered by the second only while a command runs.
Your machine
Connect the bridge and the assistant builds on your own hardware — your data never leaves your computer. You only pay for the models it calls.
Cloud sandbox
An isolated container (2 vCPU · 4 GB), prompt-injection-safe. Stays warm while you work so incremental builds are fast; reaped after an idle hour. Billed per second of run time — a typical build + test is a fraction of a cent.
Coding session
A full autonomous coding agent in a bigger container (up to 8 vCPU · 32 GB) that works a whole task end-to-end. Priced by container tier + model usage.
All compute runs on Quantum Encoding's own multi-region infrastructure. Credits cover models and compute alike — one balance, one bill.
Ready to build?
Sign up, get free starter credits, and see the real per-call cost on your first response.
Get startedFAQ
Frequently asked questions
How am I actually charged?
Per request, per token. Every API call is metered: input tokens × the per-million input rate, output tokens × the per-million output rate. We charge the provider rate plus our margin — 25% on most providers, 30% on xAI and DeepSeek. The exact cost shows in the response headers and on every entry in your dashboard. No rounding up, no per-task fudge factors.
Why credits then?
Credits are just a top-up balance — $1 = 1 credit. Your balance ticks down as actual API costs accrue. Nothing more clever than that. We don't pretend a 'credit' equals a fixed amount of compute, because it doesn't — different models cost wildly different amounts per token.
Do credits expire?
No. Your money is your money. Use it whenever.
Can I choose which model runs each call?
Yes. Pick your model per call, per mission, per task. Mix providers — Fable 5 for planning, Sonnet 5 for execution, V4 Flash for cheap throwaway calls. Your call. We just route the request and bill you the real cost plus margin.
What happens if I run out of credits mid-mission?
The mission pauses. Top up and resume exactly where you left off. No work lost, no restarts.
Why the 25–30% margin?
Because the orchestration engine, multi-provider routing, sessions, billing, observability, and the live workspace aren't free to build or run. Every API response shows the raw provider cost and our margin separately — no obfuscation. If you're a heavy user (>$50/mo) you can apply for the Developer tier (12.5%) or Lifetime (0%). See the developers page.
Do I need a credit card to try it?
No. Sign up, get free starter credits. Top up only when you want to keep going.