Model Gallery

127 models from 1 repositories

Filter by type:

Filter by tags:

qwen3.8-27b-dflash2
# Qwen3.8-27B > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc. > [!Tip] > For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud. > In particular, **Qwen3.8-27B** will be available as a hosted version with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-27B Overview. The service is coming soon. Stay tuned for updates. Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. ...

Repository: localaiLicense: apache-2.0

huihui-qwen3.8-27b-abliterated
# Qwen3.8-27B > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc. > [!Tip] > For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud. > In particular, **Qwen3.8-27B** will be available as a hosted version with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-27B Overview. The service is coming soon. Stay tuned for updates. Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. ...

Repository: localaiLicense: apache-2.0

qwen3.8-27b-heretic-abliterated-uncensored
# Qwen3.8-27B > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc. > [!Tip] > For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud. > In particular, **Qwen3.8-27B** will be available as a hosted version with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-27B Overview. The service is coming soon. Stay tuned for updates. Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. ...

Repository: localaiLicense: apache-2.0

ornith-1.5-397b-q4
Ornith-1.5-397B is Ornith AI's MIT-licensed flagship mixture-of-experts model for agentic coding, reasoning, repository-level tasks, and tool use. It supports text and image input with a context window of 262K tokens. This default entry uses the Q4_K_M GGUF and BF16 vision projector. A higher-quality Q8_0 model is available as a variant.

Repository: localaiLicense: mit

ornith-1.5-397b-q8
Ornith-1.5-397B in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

qwen3.8-27b-obliterated-q4
Qwen3.8-27B OBLITERATED is an Apache-2.0 Qwen3.8 vision-language model modified for refusal-removal and red-team research. It retains reasoning, coding, tool use, image, and video capabilities, but its safety guardrails have been removed. This default entry uses the Q4_K_M GGUF and BF16 vision projector. The linked variant uses the higher-quality Q8_0 model. The publisher recommends greedy decoding with a 1.15 repetition penalty.

Repository: localaiLicense: apache-2.0

qwen3.8-27b-obliterated-q8
Qwen3.8-27B OBLITERATED in the higher-quality Q8_0 GGUF format. This model is modified for refusal-removal and red-team research, and its safety guardrails have been removed.

Repository: localaiLicense: apache-2.0

qwen3.8-27b-q4
Qwen3.8-27B is Qwen's dense 27B vision-language model for reasoning, coding, tool use, and long-running agent tasks. It accepts text, images, and video, and it supports a native context window of 262K tokens. This default entry uses the official Q4_K_M GGUF and Q8_0 vision projector. The linked variants add MTP speculative decoding or use the higher-quality Q8_0 model.

Repository: localaiLicense: apache-2.0

qwen3.8-27b-q4-mtp
Qwen3.8-27B with the official Q4_K_M model and Q4_0 MTP draft model. MTP speculative decoding can increase generation speed by proposing multiple tokens for the target model to verify.

Repository: localaiLicense: apache-2.0

qwen3.8-27b-q8
Qwen3.8-27B in the official Q8_0 GGUF format. This variant provides higher model fidelity for hosts with enough memory.

Repository: localaiLicense: apache-2.0

qwen3.8-27b-ridge
Qwen3.8-27B Ridge is a 3.69-bit mixed quantization that keeps the Gated-DeltaNet state path at Q8_0 and preserves the embedded MTP head. It reduces the model weights to 12.59 GB while retaining multimodal, reasoning, coding, tool-use, and long-context capabilities.

Repository: localaiLicense: apache-2.0

grug-27b
Grug 27B is a multimodal Qwen3.5-derived model for chat, reasoning, vision, and tool use. This entry uses the QAT Q4_K_M GGUF build.

Repository: localaiLicense: apache-2.0

grug-27b-q8
Grug 27B Q8 is the higher-precision Q8_0 GGUF build for multimodal chat, reasoning, vision, and tool use.

Repository: localaiLicense: apache-2.0

grug-27b-mtp
Grug 27B MTP is the Q4_K_M GGUF build with multi-token prediction enabled for speculative decoding, plus the shared vision projector.

Repository: localaiLicense: apache-2.0

qwythos-27b-v1
Qwythos-27B-v1 is an Apache-2.0 dense 27B reasoning and agentic model derived from Qwen3.5-27B. It supports tool use, vision through the included projector, and a one-million-token context window. This entry uses the recommended Q4_K_M GGUF quantization; an MTP-enabled build is available as a variant for hosts with recent llama.cpp support.

Repository: localaiLicense: apache-2.0

qwythos-27b-v1-mtp
Qwythos-27B-v1 MTP is the Q4_K_M build with its native multi-token prediction head enabled for faster speculative decoding. It also includes the shared vision projector and supports tool use and long-context reasoning.

Repository: localaiLicense: apache-2.0

qwopus3.6-27b-fusion
Qwopus3.6-27B Fusion is an experimental Qwen3.6-27B merge that combines reasoning and code-execution fine-tunes. It targets agentic coding, mathematics, tool use, and long-context work while retaining image input. This default entry uses the Q4_K_M GGUF quantization and the shared Q8_0 vision projector.

Repository: localaiLicense: qwen

qwopus3.6-27b-fusion-q8
Qwopus3.6-27B Fusion is an experimental Qwen3.6-27B reasoning and coding merge. This entry uses the near-lossless Q8_0 GGUF quantization and the shared Q8_0 vision projector.

Repository: localaiLicense: qwen

tess-4-27b
Tess-4-27B is an Apache-2.0 agentic and reasoning model built on Qwen3.6-27B. It scales its thinking depth to the task and supports tool use, long-context work, and image input. This default entry uses the Q4_K_M GGUF quantization and the shared F16 vision projector.

Repository: localaiLicense: apache-2.0

tess-4-27b-q8
Tess-4-27B is an Apache-2.0 agentic and reasoning model built on Qwen3.6-27B. This entry uses the near-lossless Q8_0 GGUF quantization and the shared F16 vision projector.

Repository: localaiLicense: apache-2.0

tess-4-27b-mtp
Tess-4-27B with its Q4_K_M multi-token prediction draft enabled for speculative decoding. The main model verifies every proposed token, and the entry also includes the shared F16 vision projector.

Repository: localaiLicense: apache-2.0

Page 1