A 35B model on a single 24 GB graphics card, with multi-token prediction enabled.
On a GPU that fits both: 2.6×the original BF16, 343.11 vs 131 tokens per second on the same RTX Pro 6000 (96 GB).
Original BF16: 131 tokens per second on an RTX Pro 6000 (96 GB), the only tested GPU with enough memory for the BF16 model.
35B total parameters, 3B active · mixture of experts. Model file 18.61 GB. Build MTP-GPU-5. Download this build ↗
Original BF16: 16 bits / weight. ByteShape: 4.19 bits / weight. 3.79× intelligence per bit vs BF16.
Published results for the selected build and device. Benchmark scores use that model’s own original-precision (BF16) baseline and release-specific evaluation suite.
Intelligence per bit is our summary of retained benchmark score per bit of weight precision, relative to the same model’s BF16 baseline (1×). We calculate it as (benchmark-score retention / 100) × (16 / average bits per weight).
The precision bars show weight precision relative to BF16. The speed bars show token generation on the RTX Pro 6000. The GPU examples use multi-token prediction (MTP). Speed measures token generation; model file size covers stored weights. Runtime memory and prompt-processing results are detailed in each release.
From a University of Toronto computer architecture research group.
BACKED BY
ACROSS MODEL FAMILIES
One approach. Any kind of AI.
Our method applies to any AI model and any use case. We showcase it on language models, coding agents, vision models and image generation.
CODING AGENTS
Two coding agents. One game.
Qwen3 Coder 30B A3B on an RTX Pro 6000 workstation. The original model on the left, the ByteShape build on the right.
Each agent builds a Flappy Bird-style game. The recording shows its own elapsed time and ends with both games on screen.
Explore the iris texture, pupil outline, and reflected window in this close-up.
Qwen-Image-2512 in the GGUF pipeline: original BF16 and ByteShape’s Q3_K_S build at 3.04 bits per weight. Each published pair uses the same prompt and seed, at 1024 × 1024, 50 steps, and CFG 4.
The 40.86 GB and 7.77 GB figures cover diffusion weight files, excluding the text encoder, VAE, and runtime memory. Explore the gallery for more prompts; the release includes separate timing and memory tests with their own settings.