NVDA 195.04 ▲2.65%GOOGL 333.66 ▼0.91%MSFT 451.10 ▲15.51%AMD 485.39 ▲13.00%INTC 91.13 ▲11.30%TSMC 403.31 ▲7.64%AMZN 235.50 ▲3.90%META 539.03 ▼7.95%AAPL 333.43 ▼1.41%PLTR 122.26 ▼0.60%
Markets at last close

AMD · Models

AMD posts Instella-MoE thinking model on Hugging Face

·1 min read

AMD has posted Instella-MoE-16B-A3B-Think on Hugging Face, describing it as the final RL-refined checkpoint in a fully open Mixture-of-Experts language model family. The model has 16 billion total parameters and 2.8 billion active parameters per token, and was trained from scratch on AMD Instinct™ MI300X and MI325X GPUs using AMD’s Primus framework.

The release covers checkpoints across pre-training, mid-training, long-context extension, SFT, DPO and RL, along with the complete training recipe, including data mixtures, hyperparameters, frameworks and inference code. The architecture uses Gated Multi-head Latent Attention and FarSkip-Collective, with 27 decoder layers, hidden size 2048, 16 attention heads, 64 experts, 2 shared experts, 6 activated experts per token and a vocabulary size of 128,896.

The model is available for text generation through Transformers with custom code and can also be served through vLLM, SGLang or Docker Model Runner. AMD says the checkpoints are licensed for academic and research purposes under ResearchRAIL, with warnings that they are not intended for safety-critical, medical or high-factual-accuracy use cases and may produce inaccurate, harmful, biased or otherwise objectionable output.

Originally reported by huggingface.coRead the source →
Related coverage
All AMD news →