AI Model Ecosystem Center

Supports high-speed download of cutting-edge local models within the software

Supports dual-source local models from HuggingFace and ModelScope, as well as numerous domestic and international cloud model service providers

30+
Compatible with local lightweight models
100%
Privacy guaranteed, offline operation
35+
Mainstream cloud API integration
能力筛选:
ggml-orgHuggingFace

Qwen 3.5 0.8B (Basic)

ggml-org/Qwen3.5-0.8B-GGUF:Q4_0

内置
文本
参数:0.8B
大小:563MB
量化:Q4_0
上下文:128k
速度极速
质量

It can run without a separate graphics card, and it has extremely fast running speed. It only supports text analysis and is less accurate.

#official#lightweight#Support CPU operation#text only
unslothHuggingFace

Qwen 3.5 0.8B (lightweight image recognition)

unsloth/Qwen3.5-0.8B-GGUF:UD-Q6_K_XL

文本图像
参数:0.8B
大小:976MB
上下文:128k
速度极速
质量

Extreme running speed, suitable for extremely low-configuration environments, and supports image analysis, which is less accurate.

#GPU extremely fast#Support CPU operation#multimodal#mini
DavidAUHuggingFace

Qwen 3.5 9B (Claude fine-tuning)

DavidAU/Qwen3.5-9B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKING-MAX-NEOCODE-Imatrix-GGUF:D_AU-IQ3_M-imat

文本图像
参数:9B
大小:6.5GB
上下文:128k
速度极速
质量

Fine-tuned based on the Claude 4.6 dataset and performs well in following instructions.

#Fine-tuning optimization#Claude dataset#multimodal#NSFW#No censorship#prison Break
UnslothHuggingFace

Gemma 4 E4B-it (audio supported)

unsloth/gemma-4-E4B-it-GGUF:Q4_K_S

文本图像音频
参数:4B
大小:5.83GB
量化:Q4_K_S
上下文:128k
速度快速
质量

Google's original quantified version, supporting text, image and audio analysis.

#multimodal#fast#Audio#English is better
mradermacherHuggingFace

Qwen 3.5 4B (balanced)

mradermacher/Qwen3.5-4B_Abliterated-GGUF:Q4_K_M

文本图像
参数:4B
大小:3.08GB
量化:Q4_K_M
上下文:256k
速度快速
质量

Abliterated version balances speed and quality.

#to limit#No censorship#NSFW#prison Break#multimodal
mradermacherHuggingFace

Qwen 3.5 2B (better)

mradermacher/Huihui-Qwen3.5-2B-abliterated-GGUF:Q8_0

文本图像
参数:2B
大小:2.68GB
量化:Q8_0
上下文:256k
速度很快
质量极高

4G video memory is the first choice.

#to limit#No censorship#NSFW#prison Break#multimodal
UnslothHuggingFace

Qwen 3.5 27B (strongest)

unsloth/Qwen3.5-27B-GGUF:Q4_0

文本图像
参数:27B
大小:16.63GB
量化:Q4_0
上下文:128k
速度较慢
质量极高

Very large parameter model requires high-performance graphics card to support text and image analysis.

#Very large parameters#multimodal#High performance GPU
tatsuyaaaaaaaHuggingFace

Qwen 3.5 2B (Japanese optimization)

tatsuyaaaaaaa/Qwen3.5-2B-gguf:Q4_0

文本
参数:2B
大小:1.2GB
量化:Q4_0
上下文:128k
速度快速
质量

Text analysis model optimized for Japanese data sets, only supports text analysis.

#Japanese optimization#text only#Low video memory
jordanwoodsonHuggingFace

Qwen 3.5 2B (jailbreak text)

jordanwoodson/Qwen3.5-2B-heretic-GGUF:Q4_K_M

文本
参数:2B
大小:1.27GB
量化:Q4_K_M
上下文:128k
速度快速
质量

A jailbreak model that has a certain understanding of the content.

#prison Break#No censorship#NSFW#text only
mradermacherHuggingFace

Qwen 3.5 2B Instruct (high IQ)

mradermacher/Qwen3.5-2B-Polaris-HighIQ-INSTRUCT-GGUF:IQ4_XS

文本图像
参数:2B
大小:1.54GB
量化:IQ4_XS
上下文:128k
速度快速
质量

A highly intelligent, fine-tuned version of Polaris that supports text and image analysis.

#Fine-tuning optimization#High IQ#multimodal
OpenBMBHuggingFace

MiniCPM-V 4.6 (Picture-reading small steel cannon)

ggml-org/MiniCPM-V-4.6-GGUF:Q4_K_M

文本图像
参数:0.8B
大小:1.17GB
量化:Q4_K_M
上下文:32k
速度快速
质量极高

The top end-side multi-modal model performs extremely well in OCR, object recognition, and complex scene understanding.

#multimodal#OCR#Scene understanding#Top image recognition
unslothHuggingFace

Qwen 3.6 27B MTP (newer)

unsloth/Qwen3.6-27B-MTP-GGUF:UD-IQ2_XXS

文本图像
参数:27B
大小:9.77GB
量化:UD-IQ2_XXS
上下文:128k
速度较慢
质量极高

Version 3.6, which has explosive IQ, needs to be upgraded to 2.1+ to support it.

#multimodal#MTP#Very large parameters#High performance GPU
unslothHuggingFace

Gemma 4 E2B-it QAT MTP (all-around-high speed)

unsloth/gemma-4-E2B-it-qat-MTP-GGUF:UD-Q4_K_XL

文本图像音频
参数:2B
大小:3.71GB
量化:UD-Q4_K_XL
上下文:128k
速度极速
质量

With full capabilities, plus MTP technology for speed, it is by far the best model to choose.

#multimodal#Audio#QAT quantification#MoE#Very small video memory#English is better
GnLOLotHuggingFace

MiniCPM5 1B Claude fine-tuning (thinking chain)

GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF:Q4_K_M

文本
参数:1B
大小:688MB
量化:Q4_K_M
上下文:128k
速度极速
质量极高

MiniCPM5-1B, which is fine-tuned based on the Claude Opus Fable5 data set, supports chain-of-thinking reasoning and is good at programming and instruction following.

#Fine-tuning optimization#Claude dataset#Thought chain#Programming optimization#lightweight#Support CPU operation#text only
owaoHuggingFace

Nanbeige 4.2 3B (beyond 9B)

owao/Nanbeige4.2-3B-GGUF:Q4_K_M

文本
参数:3B
大小:2.40GB
量化:Q4_K_M
上下文:128k
速度快速
质量

3B body hard steel 9B, highly recommended, but only supports text.

#prison Break#No censorship#NSFW#text only
owaoHuggingFace

Nanbeige 4.2 3B (beyond 9B-high accuracy)

owao/Nanbeige4.2-3B-GGUF:Q8_0

文本
参数:3B
大小:4.13GB
量化:Q8_0
上下文:128k
速度快速
质量

3B body hard steel 9B, Q8_0 highest precision quantification, highly recommended, but only supports text.

#prison Break#No censorship#NSFW#text only#High precision
mradermacherHuggingFace

Nanbeige 4.2 3B (jailbreak-optimized)

mradermacher/Nanbeige4.2-3B-heretic-i1-GGUF:i1-Q5_K_M

文本
参数:3B
大小:2.78GB
量化:i1-Q5_K_M
上下文:128k
速度快速
质量

3B body hard steel 9B, the jailbroken version of i1 has higher quantification accuracy, but only supports text.

#prison Break#No censorship#NSFW#text only#High precision

Ready to experience more powerful local AI file management?

No complex configuration needed. Firefly AI Folder supports a built-in engine and one-click model download, beginning your journey to intelligent and efficient organization.