Local LLM releases

Dense73B— ctxJun 17, 2025

MIT-licensed coding LLM trained with repository-level reinforcement learning for software issue resolution.

Magistral Small

Dense24B40K ctxJun 10, 2025

Mistral AIOpen weights

Open-weight 24B reasoning model from Mistral's Magistral family, popular for local reasoning experiments.

Phi-4 Reasoning

Dense14B— ctxApr 30, 2025

Phi-4 reasoning-specialized model family for math, science, and chain-of-thought style tasks.

Granite 3.3 8B

Dense8B128K ctxApr 30, 2025

Granite 3.3 text update for enterprise chat, RAG, and instruction-following workflows.

Kimi-Audio-7B-Instruct

Hybrid10B— ctxApr 25, 2025

Open audio foundation model for audio understanding, generation, speech recognition, audio QA, captioning, and speech conversation.

Kimi-VL-A3B-Instruct

MoE16B128K ctxApr 17, 2025

Efficient MIT-licensed vision-language MoE for OCR, image/video understanding, long documents, and OS-style agent tasks.

Llama-3.3-Nemotron-Super-49B

Dense49B128K ctxApr 2, 2025

NVIDIAOpen weights

Open Llama Nemotron reasoning model from NVIDIA's 2025 Nemotron family.

Qwen2.5-Omni-7B

Local omni-modal Qwen model that supports text, image, audio, video, and speech generation in a 7B package.

Dense7B— ctxMar 26, 2025

Mistral Small 3.1

Dense24B128K ctxMar 17, 2025

Apache-licensed Small update adding vision and a 128K context window to the efficient 24B line.

OLMo 2 32B

Allen Institute for AI (Ai2)Open source

A fully open model — weights, data, and training code all public — and the first such to beat GPT-3.5 / GPT-4o mini.

Dense32B4K ctxMar 13, 2025

Granite 3.2 8B

Dense8B128K ctxFeb 26, 2025

Granite 3.2 update with reasoning controls and multimodal/document-oriented Granite variants.

Moonlight-16B-A3B-Instruct

MIT-licensed 16B/3B-active MoE trained with Moonshot's scalable Muon optimizer experiments.

MoE16B8K ctxFeb 24, 2025

DeepHermes 3 Llama 3 8B

Nous ResearchOpen weights

Nous reasoning-oriented Hermes model trained to combine concise answers with optional deep reasoning traces.

Dense8B8K ctxFeb 18, 2025

Dolphin 3.0 Llama 3.1 8B

Cognitive ComputationsOpen weights

Popular local assistant model tuned for coding, math, function calling, and agentic workflows.

Dense8B128K ctxFeb 2, 2025

Mistral Small 3

Dense24B32K ctxJan 30, 2025

A latency-optimized 24B dense model under Apache-2.0 — a popular local-deployment workhorse.

Qwen2.5-VL-72B

Dense72B128K ctxJan 26, 2025

Vision-language Qwen2.5 model for image, document, video, and agentic visual grounding tasks.

Granite 3.1 8B

Dense8B128K ctxDec 18, 2024

IBM's enterprise-focused open model with a 128k context, Apache-2.0 licensed.

Falcon 3 10B

Technology Innovation InstituteOpen weights

UAE's TII open model designed to run on light infrastructure, including laptops.

Dense10B32K ctxDec 17, 2024

Command R7B

Dense8B128K ctxDec 13, 2024

CohereOpen weights

Cohere's smallest, fastest R-series model, tuned for RAG and tool use on modest hardware.

Phi-4

Dense14B16K ctxDec 12, 2024

MicrosoftOpen source

A 14B dense model that rivals far larger ones on math and reasoning, under a permissive MIT license.

EXAONE 3.5 32B

LG AI ResearchOpen weights

EXAONE 3.5 32B open-weight model for bilingual reasoning, coding, and long-context tasks.

Dense32B32K ctxDec 9, 2024

Llama 3.3 70B

Dense70B128K ctxDec 6, 2024

Late-2024 70B Llama update delivering much of the 405B instruction-following quality at lower serving cost.

QwQ-32B-Preview

Alibaba (Qwen)Open source

Qwen's first public reasoning-preview model, aimed at math, coding, and deliberate problem solving.

Dense32B32K ctxNov 28, 2024

Qwen2.5-Coder-32B

Alibaba (Qwen)Open source

Code-specialized Qwen2.5 model family, with the 32B checkpoint as the flagship open coding model.

Dense32B128K ctxNov 12, 2024

SmolLM2 1.7B

Dense1.7B— ctxNov 4, 2024

Hugging FaceOpen source

Compact on-device model family trained on 11T tokens, popular for lightweight local chat and experimentation.

Sarvam-1

Sarvam AIOpen weights

Sarvam's 2B open model trained for ten major Indian languages.

Dense2B— ctxOct 22, 2024

Granite 3.0 8B

Dense8B4K ctxOct 21, 2024

Apache-licensed Granite 3.0 text model, part of IBM's push toward enterprise-friendly open models.

Llama-3.1-Nemotron-70B

Dense70B128K ctxOct 15, 2024

NVIDIAOpen weights

NVIDIA-tuned Llama 3.1 70B instruction model optimized with Nemotron reward and alignment recipes.

Molmo 72B

Allen Institute for AI (Ai2)Open weights

Open multimodal model family trained for strong image understanding, pointing, and visual grounding.

Dense72B— ctxSep 25, 2024

Qwen2.5-72B

Dense72B128K ctxSep 19, 2024

Broad Qwen2.5 foundation-model update spanning general, coding, math, and multimodal descendants.

Pixtral 12B

Dense12B128K ctxSep 17, 2024

Mistral's first open multimodal model, adding image understanding to a Mistral text backbone.

Yi-Coder-9B

Dense9B128K ctxSep 5, 2024

01.AIOpen weights

01.AI's compact code model trained for repository-scale programming and code completion tasks.

OLMoE 1B-7B

Allen Institute for AI (Ai2)Open source

Fully open sparse MoE model with 7B total and about 1B active parameters.

MoE7B— ctxSep 3, 2024

Phi-3.5 MoE

MoE42B128K ctxAug 20, 2024

Phi-3.5 mixture-of-experts model, scaling Microsoft's small-model line while preserving efficient active parameters.

EXAONE 3.0 7.8B

LG AI ResearchOpen weights

LG's first open-weight EXAONE model, a compact bilingual instruction model for Korean and English.

Dense7.8B— ctxAug 7, 2024

MiniCPM-V 2.6

OpenBMBOpen weights

8B vision-language model for local image, multi-image, OCR, and video understanding, with llama.cpp and Ollama support.

Dense8B— ctxAug 2, 2024

Mistral NeMo

Dense12B128K ctxJul 18, 2024

Apache-licensed 12B model co-developed with NVIDIA, including a 128K context window and strong multilingual tokenization.

Gemma 2 27B

Google DeepMindOpen weights

Second-generation Gemma model, improving open-weight quality and efficiency at 9B and 27B sizes.

Dense27B8K ctxJun 27, 2024

Qwen2-72B

Dense72B128K ctxJun 7, 2024

Qwen2's largest dense model, introducing stronger multilingual support, coding/math gains, and long-context variants.

GLM-4-9B

Z.ai (Zhipu AI)Open weights

Open GLM-4 9B model family, covering chat, long-context, and code-oriented variants.

Dense9B128K ctxJun 5, 2024

Codestral 22B

Dense22B32K ctxMay 29, 2024

Mistral AIOpen weights

Mistral's first code-specialized model, trained for code generation, fill-in-the-middle, and multi-language programming tasks.

Aya 23 35B

Dense35B— ctxMay 23, 2024

CohereOpen weights

Open multilingual research model covering 23 languages, released by Cohere For AI.

Yi-1.5-34B

Dense34B4K ctxMay 13, 2024

01.AIOpen weights

Yi 1.5 update with stronger instruction following, coding, math, and multilingual performance.

Falcon 2 11B

Technology Innovation InstituteOpen weights

Falcon 2 generation, including text and vision-language 11B models under a permissive TII license.

Dense11B8K ctxMay 13, 2024

Granite Code 34B

Dense34B8K ctxMay 6, 2024

Apache-2.0 code model from IBM's Granite Code family, used for local code generation and enterprise coding assistants.

Phi-3 Mini

Dense3.8B128K ctxApr 23, 2024

3.8B-parameter Phi-3 model released as a phone-capable small model with 4K and 128K variants.

Llama 3 70B

Dense70B8K ctxApr 18, 2024

First Llama 3 release, with 8B and 70B open models and a stronger tokenizer, data mix, and post-training stack.

CodeGemma 7B

Google DeepMindOpen weights

Open code-specialized Gemma model for local code completion, generation, and instruction-following.

Dense7B8K ctxApr 9, 2024

Jamba

Hybrid52B256K ctxMar 28, 2024

AI21 LabsOpen weights

First Jamba hybrid Transformer-Mamba MoE model with open weights and a 256K context length.

StarCoder2 15B

Dense16B16K ctxFeb 28, 2024

BigCodeOpen weights

Next-generation BigCode code model trained on 4T+ tokens and 600+ programming languages, with 16K context.

Gemma 7B

Google DeepMindOpen weights

First Gemma open-weight text model family, derived from the same research lineage as Gemini.

Dense7B8K ctxFeb 21, 2024

OLMo 7B

Allen Institute for AI (Ai2)Open source

Ai2's first fully open language model release, including weights, training data, code, logs, and intermediate checkpoints.

Dense7B4K ctxFeb 1, 2024

Stable LM 2 1.6B

Dense1.6B— ctxJan 19, 2024

Stability AIOpen weights

Small multilingual Stable LM release built for low hardware barriers and local experimentation.

DeepSeekMoE 16B

DeepSeekOpen source

Early DeepSeek sparse MoE research model that foreshadowed the later V2/V3 architecture direction.

MoE16B4K ctxJan 11, 2024

Nous Hermes 2 Mixtral

MoE47B32K ctxJan 11, 2024

Nous ResearchOpen source

Nous instruction-tuned Mixtral model with strong open-chat and tool-use adoption.

OpenChat 3.5

OpenChatOpen source

Compact Mistral-based local chat model trained with C-RLFT, popular in early 2024 local leaderboards.

Dense7B— ctxJan 6, 2024

TinyLlama 1.1B Chat

Dense1.1B— ctxJan 1, 2024

TinyLlamaOpen source

Compact Llama-style 1.1B chat model trained for local experimentation and low-memory deployments.

Phi-2

Dense2.7B— ctxDec 12, 2023

2.7B-parameter Phi model showing strong reasoning and language understanding at small scale.

OpenHathi-7B

Sarvam AIOpen weights

Sarvam AI's first open Indic language model, adapted from Llama 2 for Hindi and Indian-language work.

Dense7B— ctxDec 12, 2023

Mixtral 8x7B

MoE47B32K ctxDec 11, 2023

The open sparse Mixture-of-Experts that brought MoE efficiency to the open ecosystem.

Qwen-72B

Dense72B32K ctxNov 30, 2023

Alibaba's first major open Qwen model and the start of a prolific open-weight line.

DeepSeek LLM 67B

Dense67B4K ctxNov 29, 2023

DeepSeekOpen source

First general DeepSeek language model family, with 7B and 67B base/chat checkpoints.

Yi-34B

Dense34B200K ctxNov 6, 2023

01.AIOpen weights

01.AI's strong bilingual open model, with a 200k-context variant.

DeepSeek Coder 33B

Dense33B16K ctxNov 2, 2023

DeepSeekOpen source

DeepSeek's first public code-model family, released before the general DeepSeek LLM line.

LLaVA 1.5 13B

Hybrid13B— ctxSep 30, 2023

LLaVAOpen weights

Open vision-language assistant and one of the most widely run early local multimodal models.

Mistral 7B

Dense7B8K ctxSep 27, 2023

The 7B that punched far above its weight and put Mistral on the map.

Qwen-14B

Dense14B8K ctxSep 25, 2023

Second open Qwen size, expanding the first-generation Qwen language-model lineup.

Granite 13B

IBMOpen weights

IBM's early Granite foundation model family for enterprise language and code tasks.

Dense13B— ctxSep 7, 2023

Code Llama 34B

Dense34B16K ctxAug 24, 2023

Meta's first code-specialized Llama model family, released in base, Python, and instruction-tuned variants.

Qwen-7B

Dense7B32K ctxAug 3, 2023

Alibaba's first open Qwen checkpoint and the start of the Qwen open-model line.

Nous-Hermes-Llama2-13B

Nous ResearchOpen weights

Early Nous Hermes instruction model on Llama 2, widely used in the open-model fine-tuning ecosystem.

Dense13B4K ctxJul 24, 2023

Llama 2 70B

Dense70B4K ctxJul 18, 2023

The release that made capable open-weight models genuinely usable for production.

ChatGLM2-6B

Z.ai (Zhipu AI)Open weights

Second open ChatGLM generation, improving long context, inference efficiency, and bilingual chat quality.

Dense6B32K ctxJun 25, 2023

Phi-1

Dense1.3B— ctxJun 21, 2023

Microsoft's first Phi small-language-model release, demonstrating strong code performance from textbook-quality synthetic data.

Falcon 40B

Technology Innovation InstituteOpen weights

TII's breakout open Falcon model, released before Falcon 180B and trained on the RefinedWeb corpus.

Dense40B2K ctxMay 25, 2023

MPT-7B

Databricks / MosaicMLOpen source

MosaicML's permissively licensed 7B model, an early favorite for commercial local fine-tuning and long-context variants.

Dense7B2K ctxMay 5, 2023

Vicuna 13B

LMSYS / SkyLabOpen weights

LMSYS instruction-tuned LLaMA model that became a landmark early local ChatGPT-style assistant.

Dense13B— ctxMar 30, 2023

ChatGLM-6B

Z.ai (Zhipu AI)Open weights

Zhipu AI and Tsinghua KEG's first widely used open bilingual ChatGLM checkpoint.

Dense6B2K ctxMar 14, 2023

LLaMA

Dense65B2K ctxFeb 24, 2023

Meta's first LLaMA, released to researchers; its leak catalyzed the open-weight movement.

GPT-2

Dense1.5B1K ctxNov 5, 2019

OpenAIOpen source

Initially withheld over misuse fears, then fully released in Nov 2019 — an early 'limited release' debate.

BERT