Model Gallery

110 models from 1 repositories

Filter by type:

Filter by tags:

ling-3.0-flash-iq1
Ling-3.0-flash is InclusionAI's MIT-licensed hybrid reasoning MoE model with 124B total parameters and 5.5B active parameters per token. It targets coding, deep research, instruction following, and agentic workflows with a native 256K-token context window. This default entry uses the 36.5 GB AD-IQ1_M GGUF. A higher-quality 44.7 GB AD-IQ2_XS model is available as a variant.

Repository: localaiLicense: mit

ling-3.0-flash-iq2
Ling-3.0-flash in the higher-quality 44.7 GB AD-IQ2_XS GGUF format. This variant preserves more model fidelity for hosts with enough memory.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-q4
Ornith-1.5-35B-A3B is an MIT-licensed Qwen3.5 mixture-of-experts model from Ornith AI for agentic coding, reasoning, repository-level software tasks, and tool use. It activates about 3B parameters per token and supports text and image input with a context window of 262K tokens. This default entry uses the Q4_K_M GGUF and BF16 vision projector. A higher-quality Q8_0 model is available as a variant.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-q8
Ornith-1.5-35B-A3B in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q4
Tiel-Coder-35B-A3B is a 35B-parameter mixture-of-experts model for coding, reasoning, tool use, and vision tasks. This default entry uses the Q4_K_XL GGUF and BF16 vision projector.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q4-mtp
Tiel-Coder-35B-A3B in Q4_K_XL format with MTP speculative decoding and a BF16 vision projector.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q8
Tiel-Coder-35B-A3B in the higher-quality Q8_K_XL GGUF format, with the BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

ornith-1.5-397b-q4
Ornith-1.5-397B is Ornith AI's MIT-licensed flagship mixture-of-experts model for agentic coding, reasoning, repository-level tasks, and tool use. It supports text and image input with a context window of 262K tokens. This default entry uses the Q4_K_M GGUF and BF16 vision projector. A higher-quality Q8_0 model is available as a variant.

Repository: localaiLicense: mit

ornith-1.5-397b-q8
Ornith-1.5-397B in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

btl-4-compact
BTL-4 Compact is Bad Theory Labs' text-only 35B mixture-of-experts model compressed into a single 9.96 GB IQ2_XXS GGUF. Around 2.1B parameters are active per token, and the model is tuned for agentic work, tool use, coding, and reasoning. The compact build omits the vision tower and disables the source model's MTP layer for compatibility with stock llama.cpp.

Repository: localaiLicense: apache-2.0

deepseek-v4-pro-0813
DeepSeek V4 Pro 0813 is DeepSeek's MIT-licensed flagship mixture-of-experts model for agentic coding, reasoning, and long-horizon tool use. This entry uses Unsloth's UD-Q4_K_XL GGUF build, split into 20 shards for llama.cpp.

Repository: localaiLicense: mit

instella-moe-16b-a3b-think
AMD Instella-MoE-16B-A3B-Think is a reasoning and instruction-following mixture-of-experts model with 16 billion total parameters and 3 billion active parameters. It supports long-form reasoning, chat, coding, and tool use. This entry uses the Q4_K_M GGUF quantization.

Repository: localaiLicense: other

instella-moe-16b-a3b-think-q8
AMD Instella-MoE-16B-A3B-Think is a reasoning and instruction-following mixture-of-experts model with 16 billion total parameters and 3 billion active parameters. It supports long-form reasoning, chat, coding, and tool use. This entry uses the near-lossless Q8_0 GGUF quantization.

Repository: localaiLicense: other

north-mini-code-1.0
North Mini Code 1.0 is Cohere Labs' Apache-2.0 sparse mixture-of-experts coding model with 30B total parameters and 3B active parameters. It targets code generation, agentic software engineering, terminal tasks, tool use, and interleaved reasoning with a 256K-token context window. This entry uses the UD-Q4_K_M GGUF quantization.

Repository: localaiLicense: apache-2.0

north-mini-code-1.0-q8
North Mini Code 1.0 is Cohere Labs' Apache-2.0 sparse mixture-of-experts coding model with 30B total parameters and 3B active parameters. It targets code generation, agentic software engineering, terminal tasks, tool use, and interleaved reasoning with a 256K-token context window. This entry uses the Q8_0 GGUF quantization.

Repository: localaiLicense: apache-2.0

pocket-35b
POCKET-35B is an Apache-2.0 Qwen3.5-family mixture-of-experts model from FINAL-Bench/VIDRAFT, derived from Darwin-36B-Opus and packaged for stock llama.cpp. This entry uses the quality-oriented Q4_K_M GGUF quantization.

Repository: localaiLicense: apache-2.0

pocket-35b-q3
POCKET-35B is an Apache-2.0 Qwen3.5-family mixture-of-experts model from FINAL-Bench/VIDRAFT, derived from Darwin-36B-Opus and packaged for stock llama.cpp. This entry uses the balanced Q3_K_M GGUF quantization.

Repository: localaiLicense: apache-2.0

pocket-35b-q2
POCKET-35B is an Apache-2.0 Qwen3.5-family mixture-of-experts model from FINAL-Bench/VIDRAFT, derived from Darwin-36B-Opus and packaged for stock llama.cpp. This entry uses the smaller Q2_K GGUF quantization.

Repository: localaiLicense: apache-2.0

pocket-35b-iq1
POCKET-35B is an Apache-2.0 Qwen3.5-family mixture-of-experts model from FINAL-Bench/VIDRAFT, derived from Darwin-36B-Opus and packaged for stock llama.cpp. This entry uses the most compact IQ1_M GGUF quantization.

Repository: localaiLicense: apache-2.0

pocket-26b
POCKET-26B is an Apache-2.0 Gemma 4 26B-A4B mixture-of-experts model from FINAL-Bench/VIDRAFT, tuned for Korean and packaged for stock llama.cpp. This entry uses the quality-oriented Q4_K_M GGUF quantization.

Repository: localaiLicense: apache-2.0

pocket-26b-q2
POCKET-26B is an Apache-2.0 Gemma 4 26B-A4B mixture-of-experts model from FINAL-Bench/VIDRAFT, tuned for Korean and packaged for stock llama.cpp. This entry uses the smaller Q2_K GGUF quantization.

Repository: localaiLicense: apache-2.0

Page 1