Z.AI details GLM-5.1 for long-horizon agents
Z.AI has introduced GLM-5.1 as its latest flagship model for long-horizon work, with text input, text output, a 200K context length, and 128K maximum output tokens. The model is positioned for autonomous agents and long-horizon coding agents, with capabilities including thinking modes, streaming output, function calling, context caching, structured output, and MCP tool integration.
GLM-5.1 is described as broadly aligned with Claude Opus 4.6 in general and coding capability. On SWE-Bench Pro, it records a score of 58.4 and is presented as outperforming GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro, while showing balanced results across 12 benchmarks spanning reasoning, coding, agents, tool use, and browsing.
The model’s main emphasis is sustained execution. Z.AI says GLM-5.1 can work autonomously on a single task for up to 8 hours, moving through planning, execution, testing, fixing, optimization, and delivery while maintaining goal alignment over extended workflows.
In engineering examples, GLM-5.1 can build a complete Linux desktop system from scratch within 8 hours and carry out 655 iterations to improve vector database query throughput to 6.9× the initial production version. On KernelBench Level 3, it achieved a 3.6× geometric mean speedup, compared with 1.49× for torch.compile in max-autotune mode.