Local LLM releases push coding and reasoning gains
As of July 2026, local LLM releases are led by Poolside Laguna XS 2.1, Z.ai GLM-5.2, Moonshot AI Kimi K2.7 Code and Kimi K2.6, Alibaba Qwen3.6, Google Gemma 4, DeepSeek V4 Pro/Flash, OpenAI gpt-oss:20b/120b, Meta Llama 3.3 70B, DeepSeek-R1, and the Qwen3 and Qwen3-Coder families. GLM-5.2 is described as the top open-weights model on the Artificial Analysis Intelligence Index v4.1 with 51 pts, #1 open, 4th overall, while Qwen3.6 27B is positioned as the best overall local model with 84% MMLU, 262K context, and 201 languages.
Coding and reasoning are the main battlegrounds. Laguna XS 2.1 targets agentic, long-horizon coding with 33B total / 3B active MoE, 256K context, and SWE-bench Verified 70.9%. Kimi K2.6 reaches SWE-Bench Pro 58.6 and SWE-bench Verified 80.2%, while DeepSeek V4 Pro posts 93.5% LiveCodeBench. OpenAI’s gpt-oss:20b runs in 16 GB at roughly o3-mini level, with gpt-oss:120b requiring 80 GB.
Ollama availability remains uneven but broad. Laguna XS 2.1, Kimi K2.7 Code, Qwen3.6 27B, Gemma 4 26B-A4B, Gemma 4 E2B, DeepSeek V4 Pro, Kimi K2.6, GLM-5.1, and gpt-oss:20b have pull commands listed, while GLM-5.2 is not yet in the Ollama library and requires a hosted API or community GGUF quantization.
Local model quality has advanced sharply since early 2024. 7B-class models improved from 64% MMLU to 74%, and 70B-class models rose from 75% to 82-84%, narrowing the gap with frontier cloud models to roughly 18-24 months of equivalent capability.