Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Scope of Hacks on U.S. Water Supply Widens as Evidence Points to Iran

    August 1, 2026

    Inside the London hacker house taking a stand against founder burnout

    August 1, 2026

    Accelerating Transformer Training with NVIDIA Transformer Engine, Fused Kernels, BF16, FP8, and GPU Benchmarking

    August 1, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Scope of Hacks on U.S. Water Supply Widens as Evidence Points to Iran
    • Inside the London hacker house taking a stand against founder burnout
    • Accelerating Transformer Training with NVIDIA Transformer Engine, Fused Kernels, BF16, FP8, and GPU Benchmarking
    • Coldcard Bitcoin Thief Likely Used Top Blockchain Services Provider
    • California announces minimum wage increase as governor taunts Trump | California
    • EU calls emergency meeting to discuss Ceuta migrant crossings
    • Trump Calls for $1.8 Billion Fund Over Blanche Attorney General Fight
    • Alienware 15 vs. LOQ Essential 15: I compared both budget gaming laptops – it was close
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Saturday, August 1
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Artificial Intelligence

    AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKAugust 1, 2026 Artificial Intelligence No Comments5 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    AMD released Instella-MoE-16B-A3B, a fully open Mixture-of-Experts language model trained from scratch on Instinct MI300X and MI325X GPUs. The model holds 16B total parameters but activates only 2.8B per token. AMD is publishing weights from every training stage, along with data mixtures, training configs, and inference code. Two systems-level choices carry the release: Gated Multi-head Latent Attention and FarSkip-Collective connectivity.

    Is it deployable?

    Partly. The weights ship under a ResearchRAIL license for academic and research purposes only, so this is not a drop-in commercial model. The training codebase is MIT licensed, and that is the more reusable asset here.

    • Company level: AI research labs, university groups, and enterprise R&D teams with data-center GPU capacity. Not a fit for lean startups wanting a hosted commercial endpoint.
    • Industries: semiconductor and cloud infrastructure, AI tooling vendors, and academic research.
    • Applications: reproducing an end-to-end MoE recipe, studying expert-parallel serving, evaluating 64K long-context behavior, and running RL post-training experiments.
    • Serving cost: 16B parameters in BF16 need roughly 32 GB of weight memory, so one high-memory accelerator suffices. AMD ships SGLang inference code.
    https://rocm.blogs.amd.com/artificial-intelligence/instella-moe/README.html

    Architecture

    Instella-MoE is a decoder-only MoE with 27 layers, hidden size 2048, 16 attention heads, and a 128,896-token vocabulary. Each MoE layer uses 2 shared experts plus 6 routed experts selected from 64. That yields 2.8B active parameters against 16B total. A Multi-Token Prediction objective is used during pre-training and mid-training.

    There are two structural choices that are important to know. Gated MLA adds a lightweight learned output gate to Multi-head Latent Attention. A dedicated linear projection derives an input-conditioned gate, applied multiplicatively before the output projection. FarSkip-Collective passes outdated and partial activations into the MoE and attention layers, overlapping expert-parallel communication with computation. AMD reports a 12.7% pre-training speedup and up to a 39.2% reduction in time to first token when serving with expert parallelism.

    Training pipeline

    Pre-training covers 7.1T tokens from open corpora including Nemotron-CC-v2, MegaMath, FineMath, RefineCode, and TxT360. Mid-training uses Dolma3 Dolmino 100B across three data variants, merged by weight averaging. A long-context stage extends the window from 4K to 64K using YaRN, an increased RoPE theta, and document masking.

    Post-training runs SFT on Dolci-Think-SFT-7B plus Nemotron mixtures, ending on a feedback-driven 512K-example set targeting measured weaknesses. DPO follows, with router bias updates and the auxiliary load-balancing loss disabled to prevent degradation. RL runs in the Miles framework: 1,400 steps of instruction-following RLVR, then Multi-Teacher On-Policy Distillation to fold that gain back without losing math or code.

    Results

    The base checkpoint averages 76.7, the strongest among fully open models, ahead of Moonlight-16B-A3B (76.2), SmolLM3-3B-Base (70.5), OLMo-3-7B (70.1), and OLMoE-1B-7B (61.9). It trails Qwen3.5-4B-Base (79.5). It leads on WinoGrande (86.5) and scores 65.7 on HumanEval+. Long-context averages are 41.5 on HELMET and 79.4 on RULER.

    Post-training climbs from SFT (71.58) to DPO (72.67) to Think (73.22), above Olmo3-7B-Think (71.97), Gemma-4-E4B think (70.47), and Qwen3.5-4B (69.73). IFEval rises from 77.08 to 83.70.

    Interactive explainer

    Key Takeaways

    • 16B total parameters, 2.8B active per token: 2 shared plus 6 of 64 routed experts.
    • Gated MLA and FarSkip-Collective give a 12.7% training speedup and 39.2% lower TTFT.
    • Trained end-to-end on AMD Instinct MI300X and MI325X with ROCm, Primus, and Miles.
    • Base averages 76.7 and Think averages 73.22, both leading fully open peers.
    • ResearchRAIL weights limit commercial use; the MIT-licensed training code does not.

    Check out the ROCm blog, Hugging Face collection and GitHub. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

    Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us

    Sources: ROCm blog · Hugging Face collection · GitHub


    Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

    2.8B active AMD Fully GPUs InstellaMoE16BA3B Instinct LLM MixtureofExperts open Parameters Releases Trained
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Accelerating Transformer Training with NVIDIA Transformer Engine, Fused Kernels, BF16, FP8, and GPU Benchmarking

    Anthropic Says Claude Mistook the Open Internet for a CTF and Breached Three Organizations

    Rails patches critical Active Storage flaw with RCE potential

    MiniMax Releases MiniMax H3: An Omni-Modal Video Model That Generates 15-Second 2K Clips With Native Stereo Audio

    Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks

    Three Recent Chrome Releases Fix 1,442 Flaws, More Than Prior 23 Updates Combined

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Scope of Hacks on U.S. Water Supply Widens as Evidence Points to Iran

    August 1, 2026

    Inside the London hacker house taking a stand against founder burnout

    August 1, 2026

    Accelerating Transformer Training with NVIDIA Transformer Engine, Fused Kernels, BF16, FP8, and GPU Benchmarking

    August 1, 2026

    Coldcard Bitcoin Thief Likely Used Top Blockchain Services Provider

    August 1, 2026
    Latest Posts

    Wildfires ravage Spain, France and Italy, killing three firefighters

    July 23, 2026

    Three speeches on a single day signaled a dying American democracy | Robert B Shpiner

    July 23, 2026

    Keystone clashes: Millions pour into three Pennsylvania races that could decide control of the House • OpenSecrets

    July 23, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Scope of Hacks on U.S. Water Supply Widens as Evidence Points to Iran

    August 1, 2026

    Inside the London hacker house taking a stand against founder burnout

    August 1, 2026

    Accelerating Transformer Training with NVIDIA Transformer Engine, Fused Kernels, BF16, FP8, and GPU Benchmarking

    August 1, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.