The full ShapeLearn release of Qwen 3.8 27B: five GGUF models that
all sit on the measured quality-speed frontier across six GPUs.
GPU-5 is the default wherever it fits, at 99.63% of BF16, and
GPU-4 stays very competitive when memory is tight. ShapeLearn-Lite held up
better than its KLD ranking suggested, and we compare MTP against
DFlash2 speculative decoding.
An interactive explorer for the Krea-2-Turbo release: 24 curated
prompts rendered by the GGUF and Humming builds, each family next
to its own BF16 baseline. Compare any two variants side by side or
with a slider, with synced pan and zoom.
Our first image-generation release: quantized GGUF builds for
ComfyUI and Humming builds for vLLM-Omni, generating a
1024×1024 image in roughly 8–9 seconds on a high-end
GPU. Includes measured VRAM and speed for every model, setup guides
for both runtimes, and 24 curated prompts compared across every
variant of both families.
A three-part series on quantized-model evaluation: does the model
fit, how well does it perform, and how fast does it run on the
target hardware? Why KLD measures displacement rather than
direction, why BPW measures storage cost rather than realized
speed, and why the quant that wins the proxies is often not the
quant that wins in deployment.
A Canada Day release: ByteShape-compressed GGUF models for Cohere
North Mini Code 1.0, a Canadian coding model compressed by a
Canadian team. See the best quality/speed trade-offs across RTX
4090, 5090, 4080, and 5060 Ti — pick Model 3 on 24GB+ GPUs, and
Model 1 or 2 on 16GB GPUs.
Compare standard and multi-token-prediction builds of Qwen 3.6
35B. Published tests show how model size, benchmark score, and
token-generation speed vary across desktop and edge hardware.
ByteShape's ShapeLearn-quantized release of Qwen 3.5 35B A3B. An
MoE model where CPUs are surprisingly consistent but GPUs are much
pickier. See the best quality/speed trade-offs across RTX 4090,
4080, 5090, RTX Pro 6000 Blackwell, Intel i7, Ryzen 9, Ultra 7, and
Raspberry Pi.
Step-by-step guide to running a fully local, fully free AI coding agent
on your own hardware. Covers LM Studio, llama.cpp, and Ollama with
ByteShape GGUF models, from installation to building Flappy Bird in one prompt.
ByteShape's ShapeLearn-quantized release of Qwen 3.5 9B. GPUs agree
on the best models, CPUs have strong opinions. See the best
quality/speed trade-offs across RTX 5090, 4080, 3090, 5060 Ti, Intel
i7, Ryzen 9, Ultra 7, and Raspberry Pi.
Published comparisons of optimized Devstral and Qwen3-Coder builds
on Raspberry Pi and RTX GPUs. Explore measured memory, speed, and
benchmark-score tradeoffs, with download and setup links.
A large language model running on Raspberry Pi 5, with published
tests on Intel CPUs and NVIDIA GPUs too. Explore device-specific
speed, memory, and benchmark-score tradeoffs, then find the build
and setup instructions for your hardware.
We're excited to announce ByteShape's first public release of
ShapeLearn-quantized models. Learn how our datatype learning
technology delivers better quality at lower sizes, with benchmarks
across Qwen3 4B and Llama 3.1 8B models showing superior
performance on GPUs, CPUs, and Raspberry Pi.