Lightweight, fast
and cheap inference

DeepSeek V4 Flash

deepseek-v4-flash

DeepSeek's frontier-quality flagship, served fast on our stack.

Input
$0.15
Cached input
$0.03
Output
$0.35

per M tokens

Llama 3.3 70B Fast

llama-3.3-70b-fast

A 70B target with exact speculative verification and 1.61× faster measured decoding.

Input
$0.024
Cached input
$0.024
Output
$6.43

per M tokens

Qwen3.6 27B

qwen3.6-27b

Mid-size workhorse for assistants, retrieval, and tool use.

Input
$0.08
Cached input
$0.08
Output
$0.722

per M tokens

Qwen3 4B Fast

qwen3-4b-fast

Ultra-light tier for high-volume pipelines at the lowest input cost.

Input
$0.021
Cached input
$0.021
Output
$2.273

per M tokens

Qwen3 8B Safety Fast

qwen3-8b-safety-fast

Our fastest safety classifier for continuous SAFE or UNSAFE checks.

Input
$0.021
Cached input
$0.005
Output
$2.273

per M tokens

Qwen3 32B Safety Smart

qwen3-32b-awq-safety-smart

A larger safety classifier for workloads where accuracy matters most.

Input
$0.08
Cached input
$0.02
Output
$0.722

per M tokens

Benchmarks & stats

Model benchmarks and statistics

Vision & action

More models, hosted on our stack.

Use one All models API key across text, video, and action-policy inference.

Video generation

Wan 2.2 TI2V-5B

Limited beta

Generate 5-second, 1280×704 videos from a text prompt on our dedicated GH200 runtime.

Model ID
wan2.2-ti2v-5b
Pricing
No charge during beta

Robot action policy

LingBot-VA 14B

Limited beta

Turn paired robot-camera observations into stateful LIBERO action predictions.

Model ID
lingbot-va-libero
Pricing
No charge during beta

Rates

Usage pricing

Pay only for tokens used. No subscription.

Immediate processing · 1.00× base rate

Model Input / 1M tokens Cached input / 1M tokens Output / 1M tokens
DeepSeek V4 Flash $0.15 $0.03 $0.35
Llama 3.3 70B Fast $0.024 $0.024 $6.43
Qwen3.6 27B $0.08 $0.08 $0.722
Qwen3 4B Fast $0.021 $0.021 $2.273
Qwen3 8B Safety Fast $0.021 $0.005 $2.273
Qwen3 32B Safety Smart $0.08 $0.02 $0.722

Priority is the immediate market rate. Cached input pricing applies to tokens served from cache. Flex / Delayed is 40% lower for delivery within 10 hours. Llama 70B rates are rounded GPU-only breakeven at 100% productive measured throughput; idle capacity and operating overhead are excluded.