Technology

Keep precision
where it matters.

AI models store billions of numbers. ShapeLearn learns which need more precision and which can use fewer bits, shaping the model around your hardware.

ShapeLearn / learned precision

Different parts of the model. Different needs. Different precision.
Illustration of learned bit allocation.

See how ShapeLearn learns (opens in a new tab)
01 / DEFINE

Your deployment sets the target.

The model, hardware, memory budget, and tasks you need it to perform.

02 / LEARN

Precision follows the model.

ShapeLearn learns bit lengths and numeric formats across the model.

03 / MEASURE

The result runs on real hardware.

We measure task quality, speed, and memory, then deliver the selected build.

The allocation makes a difference

Better use
of every bit.

54.7% higher normalized benchmark score

Devstral Small 2 24B, at nearly identical generation speed on the same RTX 4080 (16 GB).

Mean score across the release's seven tasks, normalized to this model's BF16 baseline. Others is the selected public comparison build. Full results and setup (opens in a new tab).

Optimize the execution, too.

Multi-token prediction takes Qwen3.6-35B-A3B from 214.54 to 285.53 tok/s on the RTX 4090 (24 GB), a 33.1% gain at the same benchmark score. GPU-5 and MTP-GPU-5, llama.cpp.

Explore MTP (opens in a new tab)

From optimization to deployment

Built for the
tools you use.

From model files to the hardware beneath them.

Our longer-term objective is a common layer between AI software and hardware: optimizing how tensors are stored and moved, independently of the formats used for computation.

What could your model do?