Supports high-speed download of cutting-edge local models within the software
Supports dual-source local models from HuggingFace and ModelScope, as well as numerous domestic and international cloud model service providers
Qwen 3.5 0.8B (Basic)
ggml-org/Qwen3.5-0.8B-GGUF:Q4_0
It can run without a separate graphics card, and it has extremely fast running speed. It only supports text analysis and is less accurate.
Qwen 3.5 0.8B (lightweight image recognition)
unsloth/Qwen3.5-0.8B-GGUF:UD-Q6_K_XL
Extreme running speed, suitable for extremely low-configuration environments, and supports image analysis, which is less accurate.
Qwen 3.5 9B (Claude fine-tuning)
DavidAU/Qwen3.5-9B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKING-MAX-NEOCODE-Imatrix-GGUF:D_AU-IQ3_M-imat
Fine-tuned based on the Claude 4.6 dataset and performs well in following instructions.
Gemma 4 E4B-it (audio supported)
unsloth/gemma-4-E4B-it-GGUF:Q4_K_S
Google's original quantified version, supporting text, image and audio analysis.
Qwen 3.5 4B (balanced)
mradermacher/Qwen3.5-4B_Abliterated-GGUF:Q4_K_M
Abliterated version balances speed and quality.
Qwen 3.5 2B (better)
mradermacher/Huihui-Qwen3.5-2B-abliterated-GGUF:Q8_0
4G video memory is the first choice.
Qwen 3.5 27B (strongest)
unsloth/Qwen3.5-27B-GGUF:Q4_0
Very large parameter model requires high-performance graphics card to support text and image analysis.
Qwen 3.5 2B (Japanese optimization)
tatsuyaaaaaaa/Qwen3.5-2B-gguf:Q4_0
Text analysis model optimized for Japanese data sets, only supports text analysis.
Qwen 3.5 2B (jailbreak text)
jordanwoodson/Qwen3.5-2B-heretic-GGUF:Q4_K_M
A jailbreak model that has a certain understanding of the content.
Qwen 3.5 2B Instruct (high IQ)
mradermacher/Qwen3.5-2B-Polaris-HighIQ-INSTRUCT-GGUF:IQ4_XS
A highly intelligent, fine-tuned version of Polaris that supports text and image analysis.
MiniCPM-V 4.6 (Picture-reading small steel cannon)
ggml-org/MiniCPM-V-4.6-GGUF:Q4_K_M
The top end-side multi-modal model performs extremely well in OCR, object recognition, and complex scene understanding.
Qwen 3.6 27B MTP (newer)
unsloth/Qwen3.6-27B-MTP-GGUF:UD-IQ2_XXS
Version 3.6, which has explosive IQ, needs to be upgraded to 2.1+ to support it.
Gemma 4 E2B-it QAT MTP (all-around-high speed)
unsloth/gemma-4-E2B-it-qat-MTP-GGUF:UD-Q4_K_XL
With full capabilities, plus MTP technology for speed, it is by far the best model to choose.
MiniCPM5 1B Claude fine-tuning (thinking chain)
GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF:Q4_K_M
MiniCPM5-1B, which is fine-tuned based on the Claude Opus Fable5 data set, supports chain-of-thinking reasoning and is good at programming and instruction following.
Nanbeige 4.2 3B (beyond 9B)
owao/Nanbeige4.2-3B-GGUF:Q4_K_M
3B body hard steel 9B, highly recommended, but only supports text.
Nanbeige 4.2 3B (beyond 9B-high accuracy)
owao/Nanbeige4.2-3B-GGUF:Q8_0
3B body hard steel 9B, Q8_0 highest precision quantification, highly recommended, but only supports text.
Nanbeige 4.2 3B (jailbreak-optimized)
mradermacher/Nanbeige4.2-3B-heretic-i1-GGUF:i1-Q5_K_M
3B body hard steel 9B, the jailbroken version of i1 has higher quantification accuracy, but only supports text.
Ready to experience more powerful local AI file management?
No complex configuration needed. Firefly AI Folder supports a built-in engine and one-click model download, beginning your journey to intelligent and efficient organization.