Model Gallery

570 models from 1 repositories

Filter by type:

Filter by tags:

hy-mt2-1.8b-q4
Hy-MT2-1.8B is Tencent's compact multilingual translation model. It follows translation instructions across 33 languages and supports tasks such as terminology control, style transfer, and structure-preserving translation. This default entry uses the 1.1 GB Q4_K_M GGUF. A higher-quality Q8_0 model is available as a variant.

Repository: localaiLicense: apache-2.0

ling-3.0-flash-iq1
Ling-3.0-flash is InclusionAI's MIT-licensed hybrid reasoning MoE model with 124B total parameters and 5.5B active parameters per token. It targets coding, deep research, instruction following, and agentic workflows with a native 256K-token context window. This default entry uses the 36.5 GB AD-IQ1_M GGUF. A higher-quality 44.7 GB AD-IQ2_XS model is available as a variant.

Repository: localaiLicense: mit

qwen3.8-9b-q4
Qwen3.8-9B is Empero AI's full-parameter distillation of Qwen3.8 2.4T A95B into the dense Qwen3.5-9B architecture. It targets reasoning, mathematics, coding, instruction following, and tool use, and supports a native 262K-token context window. This default entry uses Q4_K_M weights; a higher-quality Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

qwen3.8-4b-q4
Qwen3.8-4B is Empero AI's full-parameter distillation of Qwen3.8 2.4T A95B into the Qwen3.5-4B architecture. It targets mathematics, reasoning, instruction following, and tool use with a native 262K-token context window. This default entry uses Q4_K_M weights; a higher-quality Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

qwen3.8-2b-q4
Qwen3.8-2B is Empero AI's smallest Qwen3.8 reasoning distillation. It uses the Qwen3.5-2B architecture and targets mathematics, instruction following, tool use, and edge deployment with a native 262K-token context window. This default entry uses Q4_K_M weights; a higher-quality Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

homura-30b-q4
Homura 30B is an English, agent-focused fine-tune of Muse Glimmer 30B. It targets autonomous tool use and direct instruction following. This entry uses the publisher's 16.9 GB Q4_K_M GGUF and supports a 131K-token context window.

Repository: localaiLicense: apache-2.0

qwen3.8-2b-distill-q4
Qwen3.8 2B Distill is an Apache-2.0, text-only Qwen3.5 2B fine-tune distilled from Qwen3.8 2.4T A95B reasoning traces. It targets compact reasoning, coding, instruction following, and function calling with a 262K native context window. This entry uses the balanced Q4_K_M GGUF quantization; the Q8_0 variant offers higher fidelity.

Repository: localaiLicense: apache-2.0

qwen3.8-2b-distill-q8
Qwen3.8 2B Distill in the higher-fidelity Q8_0 GGUF format. This text-only Qwen3.5 2B fine-tune targets reasoning, coding, instruction following, and function calling with a 262K native context window.

Repository: localaiLicense: apache-2.0

qwen3.8-4b-distill-q4
Qwen3.8 4B Distill is an Apache-2.0, text-only Qwen3.5 4B fine-tune distilled from Qwen3.8 2.4T A95B reasoning traces. It targets reasoning, coding, instruction following, and function calling with a 262K native context window. This entry uses the balanced Q4_K_M GGUF quantization; the Q8_0 variant offers higher fidelity.

Repository: localaiLicense: apache-2.0

qwen3.8-4b-distill-q8
Qwen3.8 4B Distill in the higher-fidelity Q8_0 GGUF format. This text-only Qwen3.5 4B fine-tune targets reasoning, coding, instruction following, and function calling with a 262K native context window.

Repository: localaiLicense: apache-2.0

instella-moe-16b-a3b-think
AMD Instella-MoE-16B-A3B-Think is a reasoning and instruction-following mixture-of-experts model with 16 billion total parameters and 3 billion active parameters. It supports long-form reasoning, chat, coding, and tool use. This entry uses the Q4_K_M GGUF quantization.

Repository: localaiLicense: other

instella-moe-16b-a3b-think-q8
AMD Instella-MoE-16B-A3B-Think is a reasoning and instruction-following mixture-of-experts model with 16 billion total parameters and 3 billion active parameters. It supports long-form reasoning, chat, coding, and tool use. This entry uses the near-lossless Q8_0 GGUF quantization.

Repository: localaiLicense: other

mellum2-12b-a2.5b-instruct
Mellum2-12B-A2.5B-Instruct is an Apache-2.0 mixture-of-experts model from JetBrains with 12 billion total parameters, 2.5 billion activated per token, and a 131,072-token context window. This entry uses the Q4_K_M GGUF quantization.

Repository: localaiLicense: apache-2.0

mellum2-12b-a2.5b-instruct-q8
Mellum2-12B-A2.5B-Instruct is an Apache-2.0 mixture-of-experts model from JetBrains with 12 billion total parameters, 2.5 billion activated per token, and a 131,072-token context window. This entry uses the higher-quality Q8_0 GGUF quantization.

Repository: localaiLicense: apache-2.0

inkling
# Inkling BF16 | NVFP4 | Playground | Tinker Cookbook | Acceptable Use ## 1. General Information Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs. It is intended for use in English and other languages, and across multiple coding languages. The model is designed to be used by developers building AI-powered applications, including agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems, and is suitable for general-purpose conversational use, instruction-following, and other natural language and multimodal tasks. It is released with open weights to support research, fine-tuning and integration into third-party products by downstream developers. **Languages:** English, with general multilingual capabilities across other languages. ## 2. Getting Started Try Inkling on the Tinker Playground or access via API using the Tinker Cookbook. Inkling supports local deployment using the following open-source libraries: * SGLang (recipe, PR) * vLLM (recipe, PR) * TokenSpeed (recipe, PR) * Unsloth (recipe, PR) * Huggingface (recipe, PR) ...

Repository: localaiLicense: apache-2.0

inkling-small
Inkling Small is a 276B-parameter mixture-of-experts multimodal model with 12B active parameters for text, image, and audio understanding, instruction following, coding, and tool use. This entry uses the Q4_K_M GGUF quantization, whose five language-model shards total approximately 162.5 GB.

Repository: localaiLicense: apache-2.0

inkling-small-iq2-m
Inkling Small is a 276B-parameter mixture-of-experts multimodal model with 12B active parameters for text, image, and audio understanding, instruction following, coding, and tool use. This entry uses the IQ2_M GGUF quantization, whose three language-model shards total approximately 82.4 GB.

Repository: localaiLicense: apache-2.0

minicpm5-1b-claude-opus-fable5-v2-thinking
# MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking GGUF quantizations for local deployment: **MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF** 中文说明 **MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking** is a compact 1B **Thinking** language model built on openbmb/MiniCPM5-1B. Compared with V1, this V2 release is further fine-tuned on **Fable 5** data with a stronger focus on **tool calling / function calling**, while also improving **coding** and **instruction-following**. It keeps MiniCPM5's native Thinking chat template and XML tool-call format. Previous version: **MiniCPM5-1B-Claude-Opus-Fable5-Thinking** (V1) For llama.cpp / Ollama / LM Studio deployment, see the **GGUF repository**. ## Overview ## Capabilities - **Tool calling (enhanced in V2)** — more reliable XML / function-calling style tool use on top of MiniCPM5's native format - **Coding** — code generation, debugging, and software-engineering-style tasks - **Instruction following** — more reliable adherence to user prompts and structured constraints - **Thinking mode** — chain-of-thought reasoning via the MiniCPM5 chat template - **Long context** — up to **128K tokens** (131,072 tokens per `config.json`) ...

Repository: localaiLicense: apache-2.0

bonsai-8b-1bit
Bonsai 8B (PrismML) is an end-to-end 1-bit language model built on the Qwen3-8B dense architecture (GQA, SwiGLU, RoPE, RMSNorm, 36 layers, 65,536 context). Every weight is a single sign bit (`-scale` / `+scale`) with one FP16 scale per group of 128 weights, for an effective 1.125 bits/weight and a ~1.15 GB footprint (14.2x smaller than FP16) while matching full-precision 8B instruct models at ~70.5 average across 6 benchmark categories. The Q1_0 quantization is only decodable by the PrismML llama.cpp fork, so this entry runs on LocalAI's `bonsai` backend (that fork), not the stock `llama-cpp` backend. License: Apache 2.0.

Repository: localaiLicense: apache-2.0

minicpm5-1b-claude-opus-fable5-thinking
# MiniCPM5-1B-Claude-Opus-Fable5-Thinking GGUF quantizations for local deployment: **MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF** 中文说明 **MiniCPM5-1B-Claude-Opus-Fable5-Thinking** is a compact 1B **Thinking** language model built on openbmb/MiniCPM5-1B. It is further fine-tuned on **Fable 5** data to improve **coding** and **instruction-following** while keeping MiniCPM5's native Thinking chat template and tool-call format. For llama.cpp / Ollama / LM Studio deployment, see the **GGUF repository**. ## Overview ## Capabilities - **Coding** — code generation, debugging, and software-engineering-style tasks - **Instruction following** — more reliable adherence to user prompts and structured constraints - **Thinking mode** — chain-of-thought reasoning via the MiniCPM5 chat template - **Tool calling** — inherits MiniCPM5's XML tool-call format - **Long context** — up to **128K tokens** (131,072 tokens per `config.json`) ## Quick start ```python from transformers import AutoModelForCausalLM, AutoTokenizer import torch model_id = "GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-Thinking" ...

Repository: localaiLicense: apache-2.0

agents-a1-4b
Agents-A1-4B is InternScience's Apache-2.0 dense 4B agentic model, based on Qwen3.5. It is trained for long-horizon search, engineering and scientific research, instruction following, tool use, and multimodal tasks. This entry uses the official Q4_K_M GGUF quantization and vision projector.

Repository: localaiLicense: apache-2.0

Page 1