The model catalog

Find your next model.

Optimized language, vision and image models, with downloads and published evidence in one place.

Choose for my hardware
  • 12 base models
  • 15 build families
  • 138 released artifacts

ByteShape models and optimized builds, with best published benchmark score shown as a percentage of each model's own baseline
Model Params Builds Smallest file Best score % of own BF16 Backends Links
Qwen 3.8 27BVLM Alibaba / Qwen 27B dense 5 2.56–3.84 bpw 8.84 GB 99.63% GPU-5 llama.cpp Benchmarks (opens in a new tab)Hugging Face
Qwen 3.6 35B A3BVLM Alibaba / Qwen 35B MoE, 3B active 10 2.17–4.22 bpw 9.42 GB 99.27% GPU-5 llama.cpp · LM Studio Benchmarks (opens in a new tab)Hugging Face
Qwen 3.6 35B A3B MTPVLM Alibaba / Qwen 35B MoE, 3B active 5 2.25–4.19 bpw 10.02 GB 99.27% MTP-GPU-5 llama.cpp Benchmarks (opens in a new tab)Hugging Face
Qwen 3.5 35B A3BVLM Alibaba / Qwen 35B MoE, 3B active 10 2.17–4.12 bpw 9.41 GB 99.81% GPU-8 llama.cpp · LM Studio Benchmarks (opens in a new tab)Hugging Face
Qwen3 30B A3B Instruct 2507LLM Alibaba / Qwen 30B MoE, 3B active 15 2.66–4.67 bpw 10.18 GB 99.75% GPU-8 llama.cpp · LM Studio Benchmarks (opens in a new tab)Hugging Face
Qwen 3.5 9BVLM Alibaba / Qwen 9B dense 11 2.81–5.10 bpw 3.15 GB 99.63% GPU-7 llama.cpp · LM Studio Benchmarks (opens in a new tab)Hugging Face
Llama 3.1 8B InstructLLM Meta 8B dense 18 2.54–4.31 bpw 2.56 GB 99.21% CPU-9 llama.cpp · LM Studio Benchmarks (opens in a new tab)Hugging Face
Qwen3 4B Instruct 2507LLM Alibaba / Qwen 4B dense 17 2.55–4.74 bpw 1.29 GB 99.80% GPU-8 llama.cpp · LM Studio Benchmarks (opens in a new tab)Hugging Face
Qwen3 Coder 30B A3B InstructLLM Alibaba / Qwen 30B MoE, 3B active 13 2.65–4.20 bpw 10.11 GB 98.98% GPU-6 llama.cpp · LM Studio · Ollama Benchmarks (opens in a new tab)Hugging Face
North Mini Code 1.0LLM Cohere 30B MoE, 3B active 4 3.17–5.64 bpw 12.10 GB 100.00% GPU-4 llama.cpp · LM Studio Benchmarks (opens in a new tab)Hugging Face
Devstral Small 2 24B Instruct 2512VLM Mistral AI 24B dense 8 2.34–4.04 bpw 6.90 GB 99.38% GPU-8 llama.cpp · LM Studio Benchmarks (opens in a new tab)Hugging Face
Qwen-Image-2512 (GGUF)Image 20B dense 6 3.04–6.61 bpw 7.77 GB Compare images (opens in a new tab) ComfyUI (ComfyUI-GGUF) · stable-diffusion.cpp
Windows, macOS, Linux · NVIDIA, AMD, Apple Silicon, CPU
Benchmarks (opens in a new tab)Hugging Face
Qwen-Image-2512 (Humming)Image 20B dense 6 3.07–6.77 bpw 7.83 GB Compare images (opens in a new tab) vLLM-Omni + Humming
Linux + NVIDIA only
Benchmarks (opens in a new tab)Hugging Face
Krea-2-Turbo (GGUF)Image Krea 13B dense 5 3.91–8.93 bpw 6.26 GB Compare images (opens in a new tab) ComfyUI (ComfyUI-GGUF)
Windows, macOS, Linux · NVIDIA, AMD, Apple Silicon, CPU
Comparison (opens in a new tab)Hugging Face
Krea-2-Turbo (Humming)Image Krea 13B dense 5 3.83–8.92 bpw 6.14 GB Compare images (opens in a new tab) vLLM-Omni + Humming
Linux + NVIDIA only
Comparison (opens in a new tab)Hugging Face

The memory filter uses measured peak VRAM for Qwen-Image at 1024×1024 with CPU offload. Language models use weight-file size plus a 2 GB allowance. Krea publishes no peak VRAM, so it uses the sizes its cards document: the GGUF diffusion file plus its text encoder and VAE, or the Humming transformer plus its shell, again plus 2 GB. This allowance is a starting point; runtime memory depends on your workflow.

Don't see the model you need?

Apply ShapeLearn to your model with our team, or explore planned SDK integration for your own workflow.