Pricing

Pay for the
tokens you use.

Every request is metered to the token. We charge the provider's per-million rate plus 25% (30% on xAI and DeepSeek) — and we show you both numbers, every time. No subscriptions. No seat fees. No minimum.

1

Top up

Buy credits to fund your account ($1 = 1 credit)

2

Make requests

Each call charged: tokens × provider rate × 1.25 (×1.30 on xAI/DeepSeek)

3

See the cost

Real provider cost + our margin, on every response

Starter

$25
25 credits
≈ 7M Sonnet 5 input
≈ 4M Opus 4.8 input
Buy credits
Most popular

Builder

$100
100 credits
≈ 27M Sonnet 5 input
≈ 16M Opus 4.8 input
Buy credits

Professional

$250
250 credits
≈ 67M Sonnet 5 input
≈ 40M Opus 4.8 input
Buy credits

Team

$1000
1000 credits
≈ 267M Sonnet 5 input
≈ 160M Opus 4.8 input
Buy credits

Same per-token rate regardless of pack size. The packs are just top-up amounts — your real cost is per-call. Credits never expire.

Per model

What does $25 get you?

All prices below include our margin — 25%, or 30% on xAI and DeepSeek. The "$25 buys" column shows input tokens you can send for $25 at that model's rate.

ModelIn $/MOut $/M$25 buys (input)Best for
Claude Fable 5 flagship
1M context
$12.50$62.502M input tokensThe frontier. Use it when it has to be right.
Claude Opus 4.8
1M context
$6.25$31.254M input tokensFrontier quality at half of Fable's rate.
Claude Sonnet 5 new
1M context
$3.75$18.757M input tokensThe workhorse. Best balance of speed + quality.
Claude Haiku 4.5
200K context
$1.25$6.2520M input tokensReviews, formatting, quick lookups.
GPT-5.5
1M context
$6.25$37.504M input tokensOpenAI's frontier. Try it next to Fable 5.
GPT-5.4
1M context
$3.13$18.758M input tokensSolid GPT-5 default.
GPT-5.4 Mini
400K context
$0.94$5.6327M input tokensBudget GPT-5 — small tasks, sub-agents.
Grok 4.5 new
500K context
$2.60$7.8010M input tokensxAI's new coding model — agentic software work.
Grok 4.3
1M context
$1.63$3.2515M input tokensFast general reasoning + structured outputs.
Grok 4.1 Fast
2M context
$0.26$0.6596M input tokensMassive-context (2M) work at budget rates.
Gemini 3.5 Flash value
1M context
$1.88$11.2513M input tokensGoogle's newest — near-frontier at half the price.
Gemini 3.1 Pro
1M context
$2.50$15.0010M input tokensGoogle's multimodal frontier. 1M context.
Gemini 3.1 Flash-Lite
1M context
$0.31$1.8881M input tokensHigh-volume pipelines at minimal cost.
DeepSeek V4 Pro value
1M context
$0.57$1.1344M input tokensTop open-weight reasoning under $1.20/M out.
DeepSeek V4 Flash value
1M context
$0.18$0.36139M input tokensCheapest model on the platform. Period.

Cost = (input tokens × in $/M) + (output tokens × out $/M). Output tokens are typically 2–6× more expensive than input on most providers.

Plus image, video, voice, and TTS models priced per asset. Full table in the docs.

Compute

Run code, not just prompts

The assistant can build and test in an isolated cloud sandbox — full toolchain, destroyed after each run — or on your own machine, where your code never leaves your computer. Compute is metered by the second only while a command runs.

Your machine

Free

Connect the bridge and the assistant builds on your own hardware — your data never leaves your computer. You only pay for the models it calls.

Cloud sandbox

metered / sec

An isolated container (2 vCPU · 4 GB), prompt-injection-safe. Stays warm while you work so incremental builds are fast; reaped after an idle hour. Billed per second of run time — a typical build + test is a fraction of a cent.

Coding session

by tier

A full autonomous coding agent in a bigger container (up to 8 vCPU · 32 GB) that works a whole task end-to-end. Priced by container tier + model usage.

All compute runs on Quantum Encoding's own multi-region infrastructure. Credits cover models and compute alike — one balance, one bill.

Ready to build?

Sign up, get free starter credits, and see the real per-call cost on your first response.

Get started

FAQ

Frequently asked questions

How am I actually charged?

Per request, per token. Every API call is metered: input tokens × the per-million input rate, output tokens × the per-million output rate. We charge the provider rate plus our margin — 25% on most providers, 30% on xAI and DeepSeek. The exact cost shows in the response headers and on every entry in your dashboard. No rounding up, no per-task fudge factors.

Why credits then?

Credits are just a top-up balance — $1 = 1 credit. Your balance ticks down as actual API costs accrue. Nothing more clever than that. We don't pretend a 'credit' equals a fixed amount of compute, because it doesn't — different models cost wildly different amounts per token.

Do credits expire?

No. Your money is your money. Use it whenever.

Can I choose which model runs each call?

Yes. Pick your model per call, per mission, per task. Mix providers — Fable 5 for planning, Sonnet 5 for execution, V4 Flash for cheap throwaway calls. Your call. We just route the request and bill you the real cost plus margin.

What happens if I run out of credits mid-mission?

The mission pauses. Top up and resume exactly where you left off. No work lost, no restarts.

Why the 25–30% margin?

Because the orchestration engine, multi-provider routing, sessions, billing, observability, and the live workspace aren't free to build or run. Every API response shows the raw provider cost and our margin separately — no obfuscation. If you're a heavy user (>$50/mo) you can apply for the Developer tier (12.5%) or Lifetime (0%). See the developers page.

Do I need a credit card to try it?

No. Sign up, get free starter credits. Top up only when you want to keep going.