Three new models, and up to 80% off two you already use
Cheaper, starting today
We moved two models onto a new upstream provider and passed the saving straight through. No action needed — the price change is already live on your existing key.
| Model |
Input / 1M |
Output / 1M |
Change |
deepseek-v4-flash |
$0.28 $0.056 |
$0.56 $0.112 |
80% lower |
glm-5.2 |
$2.10 $0.56 |
$6.60 $1.76 |
73% lower |
glm-5.2 is our flagship coding agent with a 1M-token context window. It now costs less per token than most models charge for their small tier.
Three new models
| Model |
Input / 1M |
Output / 1M |
Context |
Good for |
glm-5.1 |
$0.56 |
$1.76 |
200K |
Coding agent, same price as 5.2 |
kimi-k2.6 |
$0.38 |
$1.60 |
256K |
Reasoning and images |
kimi-k2.5 |
$0.24 |
$1.20 |
256K |
Cheapest of the reasoning tier |
Both Kimi models read images natively — no separate vision model, no reroute. Point them at a screenshot, a diagram, or a failing UI and ask.
Cached input is cheaper still
Long agent sessions resend the same context every turn, so cached reads are billed at a fraction of the input rate:
glm-5.2 and glm-5.1 — $0.10 per 1M cached tokens
kimi-k2.6 — $0.064
kimi-k2.5 — $0.04
deepseek-v4-flash — $0.00112
That is where most of the real-world saving lands on a long coding run.
Nothing else changed
- Same key, same endpoint — just name the new model in your config.
- Still pay-as-you-go, still no subscription.
- The free tier (
glm-4.7-flash, glm-4.6v-flash) is untouched.
- Requests containing images still route automatically to a vision-capable model, and the small per-image surcharge is unchanged.
One thing worth knowing: when our cheapest upstream is overloaded, we automatically retry your request against the original provider so it still succeeds rather than failing. Those retried calls are billed at that provider's higher rate. It only applies to requests that would otherwise have errored.
The full price list is always on the welcome page.
New here? Top up from $2 and a key is provisioned instantly.