Close Menu
NCIJ Network NCIJ Network
    What's Hot

    How 60,000 people swam to Spanish territory

    July 31, 2026

    Amazon is splitting its $600 million tariff refund with customers – here’s who’s eligible

    July 31, 2026

    CareCloud Data Breach Impacts Over 350,000

    July 31, 2026
    Facebook X (Twitter) Instagram
    Trending
    • How 60,000 people swam to Spanish territory
    • Amazon is splitting its $600 million tariff refund with customers – here’s who’s eligible
    • CareCloud Data Breach Impacts Over 350,000
    • Bitcoin (BTC) price’s July gain survives hawkish Fed, AI meltdown and Coldcard fallout
    • Indigenous Parakanã fear election U-turn may spark new invasions in Brazil
    • France’s Wildfires Are Swallowing Its Politics
    • The NHS isn’t offering Premier League match tickets as a treatment for depression – Full Fact
    • Body of US climber Mallory Geis found after Broad Peak avalanche
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Friday, July 31
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Artificial Intelligence

    DeepSeek Upgrades DeepSeek-V4-Flash-0731 with Major Agentic and Coding Gains

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKJuly 31, 2026 Artificial Intelligence No Comments4 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    DeepSeek published DeepSeek-V4-Flash-0731 on Hugging Face and moved the official V4-Flash API into public beta on July 31, 2026. The model card is explicit that this is the official release superseding the preview, and that the architecture and size are unchanged. The gains come from re-post-training, not a new design.

    The checkpoint ships with the DSpark speculative decoding module attached, matching the structure of DeepSeek-V4-Flash-DSpark. Hugging Face reports 304B parameters for the repo, which includes that draft module on top of the 284B base.

    On the API side, deepseek-v4-flash now natively supports the Responses API format and is adapted for Codex. The V4-Pro API and the app and web models were not updated.

    Is it deployable?

    Yes, in two very different ways.

    Via API, it is deployable by almost anyone: DeepSeek’s pricing page lists deepseek-v4-flash at $0.14 per 1M input tokens on a cache miss, $0.0028 on a cache hit, and $0.28 per 1M output tokens, with a 2,500 concurrency limit. That is roughly a third of deepseek-v4-pro output pricing ($0.87). Seed-stage startups, indie developers, and internal platform teams can run agent loops at this price without a GPU budget.

    Via self-hosting, the bar is much higher: The weights are MIT-licensed and ungated, but every expert stays resident in memory even though only 13B activate per token. DeepSeek’s vLLM example serves it on a single 4×GB300 node. Unsloth’s dynamic GGUFs put the lossless 8-bit build at 162 GB and a 3-bit build at 103 GB, needing roughly 110 GB of combined RAM plus VRAM. Self-hosting suits mid-size and large enterprises with a serving cluster, or one well-specced workstation at aggressive quantization.

    Architecture

    Per the DeepSeek-V4 technical report, V4-Flash is a 284B-parameter MoE with 13B activated per token and a 1M-token context window. Each MoE layer holds 1 shared expert and 256 routed experts with an intermediate dimension of 2048, and 6 routed experts fire per token. The first three MoE layers use hash routing. Multi-token prediction depth is 1.

    Attention is hybrid, combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA). Manifold-Constrained Hyper-Connections (mHC) replace conventional residual connections, with expansion factor 4 and 20 Sinkhorn-Knopp iterations. Pre-training used more than 32T tokens and the Muon optimizer. The paper’s headline efficiency figure — 27% of single-token inference FLOPs and 10% of KV cache versus DeepSeek-V3.2 at 1M context — is stated for V4-Pro, not Flash.

    agentic Coding DeepSeek DeepSeekV4Flash0731 Gains major upgrades
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    LingBot-Map Tutorial: GPU-Aware Inference and Point Cloud Export

    OpenAI aligns safety practices with EU AI Act’s GPAI Code

    JetBrains Open-Sources KotlinLLM: Smart Macros That Generate Kotlin Source Code at Runtime and Hot-Reload It Through JDI

    Nous Research Ships Three Integration Paths for Hermes Agent and Buzz, Block’s Open Source Nostr Workspace for Humans and Agents

    PolyAI Releases Dialog-RSN-1: An Audio-Native Dialog Model That Fuses Turn-Taking, Speech Recognition, Function Calling, And Response

    Building a Policy-Governed Multi-Agent Financial Research Workflow with Omnigent

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    How 60,000 people swam to Spanish territory

    July 31, 2026

    Amazon is splitting its $600 million tariff refund with customers – here’s who’s eligible

    July 31, 2026

    CareCloud Data Breach Impacts Over 350,000

    July 31, 2026

    Bitcoin (BTC) price’s July gain survives hawkish Fed, AI meltdown and Coldcard fallout

    July 31, 2026
    Latest Posts

    New to Linux? This 10-day checklist will help you settle in nice and easy

    July 22, 2026

    Tories ask HMRC to investigate whether Nigel Farage owes tax on £5m gift | Nigel Farage

    July 22, 2026

    Greece derails EU’s Russia sanctions plan – POLITICO

    July 22, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    How 60,000 people swam to Spanish territory

    July 31, 2026

    Amazon is splitting its $600 million tariff refund with customers – here’s who’s eligible

    July 31, 2026

    CareCloud Data Breach Impacts Over 350,000

    July 31, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.