NVDA 217.55 ▼2.86%GOOGL 357.52 ▲0.91%MSFT 506.06 ▲1.21%AMD 469.56 ▼2.86%INTC 97.52 ▼4.06%TSMC 418.47 ▼0.37%AMZN 278.09 ▲1.32%META 594.92 ▲0.48%AAPL 308.26 ▼1.62%PLTR 175.23 ▲1.87%
Markets at last close

Meta · Models

Meta returns to open weights with Muse Glimmer

·1 min read

Meta has launched Muse Glimmer, a 30 billion-parameter LLM distilled from its larger proprietary Muse Spark model, marking its first open weights release in more than a year. The move follows criticism that Meta had drifted from the open approach that helped establish Llama in 2023, especially after Llama 4 underperformed and the company restructured its AI group.

Glimmer is aimed at local inference workloads such as agents, code assistants, multimodal tool use and function calling. Meta released it under the Apache 2.0 license, allowing enterprises to deploy, modify and use the model, and early support is appearing in Llama.cpp, Ollama and Unsloth.

The model is positioned for small-to-medium sized enterprises and enthusiasts rather than as a challenger to larger Chinese open weights models. At BF16 precision it can fit in a single Nvidia RTX Pro 6000 or AMD MI350P, while 4-bit quantization reduces weights from around 60 GB to just under 16 GB, enough for many 20 to 24 GB consumer graphics cards.

Meta says RTX 5090 users can see between 75 and 233 tok/s with sufficient memory bandwidth and speculative decoding, while an M5 Max MacBook Pro is estimated at 26.2 to 57.8 tok/s. Alexandr Wang also said an open weights version of Muse Spark 1.2 will arrive soon.

Originally reported by theregister.comRead the source →
Related coverage
All Meta news →