The model catalog
Find your next model.
Optimized language, vision and image models, with downloads and published evidence in one place.
Type
Family
| Model | Params | Builds | Smallest file | Best score % of own BF16 | Backends | Links |
|---|---|---|---|---|---|---|
| Qwen 3.8 27BVLM Alibaba / Qwen | 27B dense | 5 2.56–3.84 bpw | 8.84 GB | 99.63% GPU-5 | llama.cpp | Benchmarks (opens in a new tab)Hugging Face |
| Qwen 3.6 35B A3BVLM Alibaba / Qwen | 35B MoE, 3B active | 10 2.17–4.22 bpw | 9.42 GB | 99.27% GPU-5 | llama.cpp · LM Studio | Benchmarks (opens in a new tab)Hugging Face |
| Qwen 3.6 35B A3B MTPVLM Alibaba / Qwen | 35B MoE, 3B active | 5 2.25–4.19 bpw | 10.02 GB | 99.27% MTP-GPU-5 | llama.cpp | Benchmarks (opens in a new tab)Hugging Face |
| Qwen 3.5 35B A3BVLM Alibaba / Qwen | 35B MoE, 3B active | 10 2.17–4.12 bpw | 9.41 GB | 99.81% GPU-8 | llama.cpp · LM Studio | Benchmarks (opens in a new tab)Hugging Face |
| Qwen3 30B A3B Instruct 2507LLM Alibaba / Qwen | 30B MoE, 3B active | 15 2.66–4.67 bpw | 10.18 GB | 99.75% GPU-8 | llama.cpp · LM Studio | Benchmarks (opens in a new tab)Hugging Face |
| Qwen 3.5 9BVLM Alibaba / Qwen | 9B dense | 11 2.81–5.10 bpw | 3.15 GB | 99.63% GPU-7 | llama.cpp · LM Studio | Benchmarks (opens in a new tab)Hugging Face |
| Llama 3.1 8B InstructLLM Meta | 8B dense | 18 2.54–4.31 bpw | 2.56 GB | 99.21% CPU-9 | llama.cpp · LM Studio | Benchmarks (opens in a new tab)Hugging Face |
| Qwen3 4B Instruct 2507LLM Alibaba / Qwen | 4B dense | 17 2.55–4.74 bpw | 1.29 GB | 99.80% GPU-8 | llama.cpp · LM Studio | Benchmarks (opens in a new tab)Hugging Face |
| Qwen3 Coder 30B A3B InstructLLM Alibaba / Qwen | 30B MoE, 3B active | 13 2.65–4.20 bpw | 10.11 GB | 98.98% GPU-6 | llama.cpp · LM Studio · Ollama | Benchmarks (opens in a new tab)Hugging Face |
| North Mini Code 1.0LLM Cohere | 30B MoE, 3B active | 4 3.17–5.64 bpw | 12.10 GB | 100.00% GPU-4 | llama.cpp · LM Studio | Benchmarks (opens in a new tab)Hugging Face |
| Devstral Small 2 24B Instruct 2512VLM Mistral AI | 24B dense | 8 2.34–4.04 bpw | 6.90 GB | 99.38% GPU-8 | llama.cpp · LM Studio | Benchmarks (opens in a new tab)Hugging Face |
| Qwen-Image-2512 (GGUF)Image | 20B dense | 6 3.04–6.61 bpw | 7.77 GB | Compare images (opens in a new tab) | ComfyUI (ComfyUI-GGUF) · stable-diffusion.cpp Windows, macOS, Linux · NVIDIA, AMD, Apple Silicon, CPU |
Benchmarks (opens in a new tab)Hugging Face |
| Qwen-Image-2512 (Humming)Image | 20B dense | 6 3.07–6.77 bpw | 7.83 GB | Compare images (opens in a new tab) | vLLM-Omni + Humming Linux + NVIDIA only |
Benchmarks (opens in a new tab)Hugging Face |
| Krea-2-Turbo (GGUF)Image Krea | 13B dense | 5 3.91–8.93 bpw | 6.26 GB | Compare images (opens in a new tab) | ComfyUI (ComfyUI-GGUF) Windows, macOS, Linux · NVIDIA, AMD, Apple Silicon, CPU |
Comparison (opens in a new tab)Hugging Face |
| Krea-2-Turbo (Humming)Image Krea | 13B dense | 5 3.83–8.92 bpw | 6.14 GB | Compare images (opens in a new tab) | vLLM-Omni + Humming Linux + NVIDIA only |
Comparison (opens in a new tab)Hugging Face |
The memory filter uses measured peak VRAM for Qwen-Image at 1024×1024 with CPU offload. Language models use weight-file size plus a 2 GB allowance. Krea publishes no peak VRAM, so it uses the sizes its cards document: the GGUF diffusion file plus its text encoder and VAE, or the Humming transformer plus its shell, again plus 2 GB. This allowance is a starting point; runtime memory depends on your workflow.
Don't see the model you need?
Apply ShapeLearn to your model with our team, or explore planned SDK integration for your own workflow.