DeepSeek V4 Flash
deepseek-v4-flash DeepSeek's frontier-quality flagship, served fast on our stack.
- Input
- $0.15
- Cached input
- $0.03
- Output
- $0.35
per M tokens
deepseek-v4-flash DeepSeek's frontier-quality flagship, served fast on our stack.
per M tokens
llama-3.3-70b-fast A 70B target with exact speculative verification and 1.61× faster measured decoding.
per M tokens
qwen3.6-27b Mid-size workhorse for assistants, retrieval, and tool use.
per M tokens
qwen3-4b-fast Ultra-light tier for high-volume pipelines at the lowest input cost.
per M tokens
qwen3-8b-safety-fast Our fastest safety classifier for continuous SAFE or UNSAFE checks.
per M tokens
qwen3-32b-awq-safety-smart A larger safety classifier for workloads where accuracy matters most.
per M tokens
Vision & action
Use one All models API key across text, video, and action-policy inference.
Video generation
Generate 5-second, 1280×704 videos from a text prompt on our dedicated GH200 runtime.
wan2.2-ti2v-5bRobot action policy
Turn paired robot-camera observations into stateful LIBERO action predictions.
lingbot-va-liberoRates
Pay only for tokens used. No subscription.
Immediate processing · 1.00× base rate
| Model | Input / 1M tokens | Cached input / 1M tokens | Output / 1M tokens |
|---|---|---|---|
| DeepSeek V4 Flash | $0.15 | $0.03 | $0.35 |
| Llama 3.3 70B Fast | $0.024 | $0.024 | $6.43 |
| Qwen3.6 27B | $0.08 | $0.08 | $0.722 |
| Qwen3 4B Fast | $0.021 | $0.021 | $2.273 |
| Qwen3 8B Safety Fast | $0.021 | $0.005 | $2.273 |
| Qwen3 32B Safety Smart | $0.08 | $0.02 | $0.722 |
Priority is the immediate market rate. Cached input pricing applies to tokens served from cache. Flex / Delayed is 40% lower for delivery within 10 hours. Llama 70B rates are rounded GPU-only breakeven at 100% productive measured throughput; idle capacity and operating overhead are excluded.