About OneTriangle

OneTriangle works on optimizing prefilling computation, providing companies with the cheapest and fastest inference.

FocusPortable context

Built forMulti-model inference

OneTriangle is built by a group of five MIT engineers and researchers with experience at Google DeepMind, Jane Street, SpaceX, MIT Lincoln Laboratory, and CSAIL. The team includes International Physics and Astronomy Olympiad medalists and has published at NeurIPS and ICML.

Advised by vLLM lead at Red Hat AI, Robert Shaw, and a Google Distinguished Engineer in AI Infrastructure.

Any model should be able to pick up where another left off. We envision AI systems that move context across models without recomputing it, making inference faster, cheaper, and more capable.

Make computed context transferable across the open-weight ecosystem. We build infrastructure that transfers state between models, so teams can use the best model for each task without paying the prefill cost twice. This way, we can provide inference at a fraction of the cost and TTFT.