Supports high-speed download of cutting-edge local models within the software
Supports dual-source local models from HuggingFace and ModelScope, as well as numerous domestic and international cloud model service providers
Qwen 3.5 0.8B (Chinese is better)
unsloth/Qwen3.5-0.8B-GGUF:UD-Q5_K_XL
Extremely fast and lightweight text model, suitable for low-end and CPU environments, with balanced Chinese analysis performance.
Qwen 3.5 0.8B (Lightweight image recognition)
unsloth/Qwen3.5-0.8B-GGUF:UD-Q6_K_XL
Extreme running speed, suitable for extremely low-configuration environments, and supports image analysis, which is less accurate.
LFM2.5 2.6B(Prison Break)
Abiray/LFM2.5-2.6B-Heretic-Abliterated-GGUF:Q4_K_M
LFM2.5 2.6B is a delimited uncensored version, balancing speed and quality.
Gemma 4 E4B-it(Support audio)
unsloth/gemma-4-E4B-it-GGUF:Q4_K_S
Google's original quantified version, supporting text, image and audio analysis.
Qwen 3.5 4B(Balance)
mradermacher/Qwen3.5-4B_Abliterated-GGUF:Q4_K_M
Abliterated version balances speed and quality.
Qwen 3.5 2B(Better)
mradermacher/Huihui-Qwen3.5-2B-abliterated-GGUF:Q8_0
4G video memory is the first choice.
Qwen 3.5 2B(Japanese optimization)
tatsuyaaaaaaa/Qwen3.5-2B-gguf:Q4_0
Text analysis model optimized for Japanese data sets, only supports text analysis.
Qwen 3.5 2B(Jailbreak text)
jordanwoodson/Qwen3.5-2B-heretic-GGUF:Q4_K_M
A jailbreak model that has a certain understanding of the content.
MiniCPM-V 4.6 (Picture-reading small steel cannon)
ggml-org/MiniCPM-V-4.6-GGUF:Q4_K_M
The top end-side multi-modal model performs extremely well in OCR, object recognition, and complex scene understanding.
Gemma 4 E2B-it QAT MTP(All-round - high speed)
unsloth/gemma-4-E2B-it-qat-MTP-GGUF:UD-Q4_K_XL
With full capabilities, plus MTP technology for speed, it is by far the best model to choose.
MiniCPM5 1B Claude微调(Thought chain)
GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF:Q4_K_M
MiniCPM5-1B, which is fine-tuned based on the Claude Opus Fable5 data set, supports chain-of-thinking reasoning and is good at programming and instruction following.
Nanbeige 4.2 3B (beyond 9B)
owao/Nanbeige4.2-3B-GGUF:Q4_K_M
3B body hard steel 9B, highly recommended, but only supports text.
Nanbeige 4.2 3B(Jailbreak-Optimization)
mradermacher/Nanbeige4.2-3B-heretic-i1-GGUF:i1-Q5_K_M
3B body hard steel 9B, the jailbroken version of i1 has higher quantization accuracy, but only supports text.
Gemma 3 1B Instruct(基础)
unsloth/gemma-3-1b-it-GGUF:UD-Q4_K_XL
English model leader
LFM2.5 8B A1B(Prison Break)
gaston-parravicini/LFM2.5-8B-A1B-Uncensored-Gaston-GGUF:Q4_K_M
LFM2.5 8B A1B MoE Uncensored version, only 1B activation parameters, balanced speed and quality.
Bonsai 27B 1-bit(Super compression)
dealignai/Bonsai-27b-1bit-CRACK-GGUF:Q1_0
Bonsai 27B one-bit quantization jailbreak version, de-aligned and uncensored, supports text and image analysis.
Qwen 3.8 27B(The Strongest-Jailbreak)
JonathanColetti/Qwen3.8-27B-Uncensored-GGUF:Q4_K_M
Qwen3.8-27B is an uncensored quantitative version that retains MTP long predictions and supports text and image analysis.
Qwen 3.8 2B(Distillation)
empero-ai/Qwen3.8-2B-Distill-GGUF:Q4_K_M
Qwen3.8 Distilled Edition, ultra-fast performance, text analysis only.
Qwen 3.8 9B(Distillation)
empero-ai/Qwen3.8-9B-Distill-GGUF:Q4_K_M
Qwen3.8 Distilled Edition, inheriting deep reasoning capabilities from large models, text analysis only.
Ornith 1.5 9B(Top-tier)
ornith-ai/Ornith-1.5-9B-GGUF:Q4_K_M
Ornith 1.5 Multimodal model, supports text and image analysis.
Spark X2.5 1.7B(Lightweight)
XHToken/Spark-X2.5-1.7B-GGUF
Spark X2.5 1.7B 文本模型,轻量高效,仅支持文本分析。
Ready to experience more powerful local AI file management?
No complex configuration needed. Firefly AI Folder supports a built-in engine and one-click model download, beginning your journey to intelligent and efficient organization.