Test one input across three inference paths.

Compare a native small model, OneTriangle KV cache transfer, and a native large model. Live runs show outputs and latency; sealed tables report fixed benchmark evidence.

Text comparison

Text comparison

Send one prompt through all three paths.

Checking GPU

Live mode runs the same prompt through native 4B, our transferred 4B-to-32B path, and native 32B on the local two-H100 host. Use the outputs and live timing to compare this request. The sealed table below is the benchmark evidence.