35B on a 24 GB GPU
285.53 tok/s
Qwen3.6 on RTX 4090 with multi-token prediction.
Inspect the results (opens in a new tab)Custom optimization
ShapeLearn is ByteShape's automated optimization library and SDK. It learns the precision your model needs to meet your hardware and quality targets.
Independent SDK use is our planned product path. Talk to us about bringing ShapeLearn into your own model pipeline.
Apply ShapeLearn with our team, from integration to optimized model files, benchmark results, and runtime settings.
285.53 tok/s
Qwen3.6 on RTX 4090 with multi-token prediction.
Inspect the results (opens in a new tab)8.03 tok/s
Qwen3 on Raspberry Pi 5 (16 GB), running on the CPU with llama.cpp.
Inspect the results (opens in a new tab)5.26× smaller
Qwen-Image-2512: 40.86 GB to 7.77 GB of diffusion weights.
Compare the images (opens in a new tab)Tell us what you want to optimize and how you'd like to work together.