Open Source · MarkTechPost ·
AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs
AMD released Instella-MoE-16B-A3B, an open mixture-of-experts language model with 16B total parameters and 2.8B active per token. Trained on Instinct MI300X and MI325X GPUs, it includes weights from each training stage, data mixtures, configurations, and inference code.