Join Us

A 70B model built the context. An 8B model answered at near-8B speed.

Research note 011

Ultravox 70B built a richer audio KV cache and transferred it to an 8B model for near-native-speed generation. The handoff reached 81.7% accuracy at 59.6 tokens per second, recovering 12.7 points over native 8B.

Read blog

A 32B read handed off to 8B-speed generation.

Research note 010

Qwen3-32B prefilling followed by mapped Qwen3-8B decoding reached 1.94 times native-32B decode throughput at 16K while using 14% less active GPU compute, with promising paired quality results.

Read blog

A faster serial topology made 128-way TTFT 65× worse.

Research note 009

A topology that improved single-request latency collapsed under 128 concurrent decode requests, showing why production inference tuning has to measure load, visible answers, and recovery together.

Read blog