NVDA 220.78 ▲1.48%GOOGL 339.35 ▼2.09%MSFT 507.29 ▼1.22%AMD 470.72 ▲1.10%INTC 89.51 ▲0.04%TSMC 415.32 ▼0.53%AMZN 259.77 ▼2.50%META 572.34 ▼0.98%AAPL 316.85 ▼0.89%PLTR 186.38 ▲0.05%
Markets at last close

Models

Z.AI details GLM-5.1 for long-horizon agents

·1 min read

Z.AI has introduced GLM-5.1 as its latest flagship model for long-horizon work, with text input, text output, a 200K context length, and 128K maximum output tokens. The model is positioned for autonomous agents and long-horizon coding agents, with capabilities including thinking modes, streaming output, function calling, context caching, structured output, and MCP tool integration.

GLM-5.1 is described as broadly aligned with Claude Opus 4.6 in general and coding capability. On SWE-Bench Pro, it records a score of 58.4 and is presented as outperforming GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro, while showing balanced results across 12 benchmarks spanning reasoning, coding, agents, tool use, and browsing.

The model’s main emphasis is sustained execution. Z.AI says GLM-5.1 can work autonomously on a single task for up to 8 hours, moving through planning, execution, testing, fixing, optimization, and delivery while maintaining goal alignment over extended workflows.

In engineering examples, GLM-5.1 can build a complete Linux desktop system from scratch within 8 hours and carry out 655 iterations to improve vector database query throughput to 6.9× the initial production version. On KernelBench Level 3, it achieved a 3.6× geometric mean speedup, compared with 1.49× for torch.compile in max-autotune mode.

Originally reported by docs.z.aiRead the source →
Related coverage