Model Gallery

340 models from 1 repositories

Filter by type:

Filter by tags:

qwen3.8-27b-dflash2
# Qwen3.8-27B > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc. > [!Tip] > For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud. > In particular, **Qwen3.8-27B** will be available as a hosted version with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-27B Overview. The service is coming soon. Stay tuned for updates. Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. ...

Repository: localaiLicense: apache-2.0

huihui-qwen3.8-27b-abliterated
# Qwen3.8-27B > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc. > [!Tip] > For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud. > In particular, **Qwen3.8-27B** will be available as a hosted version with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-27B Overview. The service is coming soon. Stay tuned for updates. Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. ...

Repository: localaiLicense: apache-2.0

qwen3.8-27b-heretic-abliterated-uncensored
# Qwen3.8-27B > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc. > [!Tip] > For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud. > In particular, **Qwen3.8-27B** will be available as a hosted version with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-27B Overview. The service is coming soon. Stay tuned for updates. Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. ...

Repository: localaiLicense: apache-2.0

qwen3.8-27b-obliterated-q4
Qwen3.8-27B OBLITERATED is an Apache-2.0 Qwen3.8 vision-language model modified for refusal-removal and red-team research. It retains reasoning, coding, tool use, image, and video capabilities, but its safety guardrails have been removed. This default entry uses the Q4_K_M GGUF and BF16 vision projector. The linked variant uses the higher-quality Q8_0 model. The publisher recommends greedy decoding with a 1.15 repetition penalty.

Repository: localaiLicense: apache-2.0

qwen3.8-27b-obliterated-q8
Qwen3.8-27B OBLITERATED in the higher-quality Q8_0 GGUF format. This model is modified for refusal-removal and red-team research, and its safety guardrails have been removed.

Repository: localaiLicense: apache-2.0

qwen3.8-27b-q4
Qwen3.8-27B is Qwen's dense 27B vision-language model for reasoning, coding, tool use, and long-running agent tasks. It accepts text, images, and video, and it supports a native context window of 262K tokens. This default entry uses the official Q4_K_M GGUF and Q8_0 vision projector. The linked variants add MTP speculative decoding or use the higher-quality Q8_0 model.

Repository: localaiLicense: apache-2.0

qwen3.8-27b-q4-mtp
Qwen3.8-27B with the official Q4_K_M model and Q4_0 MTP draft model. MTP speculative decoding can increase generation speed by proposing multiple tokens for the target model to verify.

Repository: localaiLicense: apache-2.0

qwen3.8-27b-q8
Qwen3.8-27B in the official Q8_0 GGUF format. This variant provides higher model fidelity for hosts with enough memory.

Repository: localaiLicense: apache-2.0

qwen3.8-27b-ridge
Qwen3.8-27B Ridge is a 3.69-bit mixed quantization that keeps the Gated-DeltaNet state path at Q8_0 and preserves the embedded MTP head. It reduces the model weights to 12.59 GB while retaining multimodal, reasoning, coding, tool-use, and long-context capabilities.

Repository: localaiLicense: apache-2.0

qwen3.8-9b-q4
Qwen3.8-9B is Empero AI's full-parameter distillation of Qwen3.8 2.4T A95B into the dense Qwen3.5-9B architecture. It targets reasoning, mathematics, coding, instruction following, and tool use, and supports a native 262K-token context window. This default entry uses Q4_K_M weights; a higher-quality Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

qwen3.8-9b-q8
Qwen3.8-9B in the higher-quality Q8_0 GGUF format. This variant preserves more model fidelity for hosts with enough memory.

Repository: localaiLicense: apache-2.0

qwen3.8-4b-q4
Qwen3.8-4B is Empero AI's full-parameter distillation of Qwen3.8 2.4T A95B into the Qwen3.5-4B architecture. It targets mathematics, reasoning, instruction following, and tool use with a native 262K-token context window. This default entry uses Q4_K_M weights; a higher-quality Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

qwen3.8-4b-q8
Qwen3.8-4B in the higher-quality Q8_0 GGUF format. This variant preserves more model fidelity for hosts with enough memory.

Repository: localaiLicense: apache-2.0

qwen3.8-2b-q4
Qwen3.8-2B is Empero AI's smallest Qwen3.8 reasoning distillation. It uses the Qwen3.5-2B architecture and targets mathematics, instruction following, tool use, and edge deployment with a native 262K-token context window. This default entry uses Q4_K_M weights; a higher-quality Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

qwen3.8-2b-q8
Qwen3.8-2B in the higher-quality Q8_0 GGUF format. This variant preserves more model fidelity while remaining suitable for compact hosts.

Repository: localaiLicense: apache-2.0

qwen3.5-9b-defiant-fable-mtp
Qwen3.5 9B Defiant Fable is an Apache-2.0 multimodal fine-tune for reasoning, coding, creative writing, and roleplay. It retains the 256K context window and vision support of Qwen3.5 while reducing refusals. This default entry uses the NEO-imatrix Q4_K_M build with multi-token prediction enabled for faster generation.

Repository: localaiLicense: apache-2.0

qwen3.5-9b-defiant-fable
Qwen3.5 9B Defiant Fable in the plain NEO-imatrix Q4_K_M GGUF format. This fallback offers the same multimodal reasoning, coding, and creative capabilities without enabling multi-token prediction.

Repository: localaiLicense: apache-2.0

qwen3.8-2b-distill-q4
Qwen3.8 2B Distill is an Apache-2.0, text-only Qwen3.5 2B fine-tune distilled from Qwen3.8 2.4T A95B reasoning traces. It targets compact reasoning, coding, instruction following, and function calling with a 262K native context window. This entry uses the balanced Q4_K_M GGUF quantization; the Q8_0 variant offers higher fidelity.

Repository: localaiLicense: apache-2.0

qwen3.8-2b-distill-q8
Qwen3.8 2B Distill in the higher-fidelity Q8_0 GGUF format. This text-only Qwen3.5 2B fine-tune targets reasoning, coding, instruction following, and function calling with a 262K native context window.

Repository: localaiLicense: apache-2.0

qwen3.8-4b-distill-q4
Qwen3.8 4B Distill is an Apache-2.0, text-only Qwen3.5 4B fine-tune distilled from Qwen3.8 2.4T A95B reasoning traces. It targets reasoning, coding, instruction following, and function calling with a 262K native context window. This entry uses the balanced Q4_K_M GGUF quantization; the Q8_0 variant offers higher fidelity.

Repository: localaiLicense: apache-2.0

qwen3.8-4b-distill-q8
Qwen3.8 4B Distill in the higher-fidelity Q8_0 GGUF format. This text-only Qwen3.5 4B fine-tune targets reasoning, coding, instruction following, and function calling with a 262K native context window.

Repository: localaiLicense: apache-2.0

Page 1