Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Burnham must understand Wales is ‘not a region in England’, says first minister | Rhun ap Iorwerth

    September 13, 2026

    UK government to offer clearer guidance on student loans after repayment row | Student finance

    September 13, 2026

    The Best 3-in-1 Apple Charging Stations After Testing 30+ Models

    September 13, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Burnham must understand Wales is ‘not a region in England’, says first minister | Rhun ap Iorwerth
    • UK government to offer clearer guidance on student loans after repayment row | Student finance
    • The Best 3-in-1 Apple Charging Stations After Testing 30+ Models
    • India launches tokenized bond pilot
    • The battle to save the dying heart of rural France: the village bistro | France
    • EU-Parlament öffnet sich für härtere China-Politik – POLITICO
    • From Hacks to Bioweapons, Claude Misuse Is Now Everywhere
    • Metaplanet Cuts Series 10 Stock Pool by 41%, Plans Hong Kong Subsidiary
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Sunday, September 13
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Artificial Intelligence

    Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKSeptember 13, 2026 Artificial Intelligence No Comments4 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Cognition, the company behind the Devin coding agent, has released SWE-2, its most capable coding model to date. SWE-2 is post-trained with reinforcement learning from Kimi K3, Moonshot AI’s 2.8T-parameter open model. Cognition reports a score of 50.0% on FrontierCode 1.1 Main, within 1 point of Fable 5.1 at 64% lower cost. It is also Cognition’s first model with selectable reasoning-effort levels, all trained in a single RL run.

    Is it deployable? Not on your own infrastructure. SWE-2 has no open weights and no standalone API. It runs only inside Devin: Desktop and CLI today, with Devin Web and Fusion rolling out.

    What is SWE-2

    SWE-2 builds on the infrastructure and recipe behind SWE-1.7, which was post-trained from Kimi K2.7. This time Cognition scaled RL to the multi-trillion-parameter regime, using a base model with almost 3x the parameters. Cognition says its RL still finds substantial headroom on top of K3, adding 5 to 6 points on many benchmarks.

    The main change is an RL algorithm that trains all 3 effort levels in one run. Each level carries its own cost penalty, so the whole cost-and-performance frontier moves at once.

    Benchmark Results

    Cognition published the following table. Public results are used where available; otherwise each model runs in its native harness at best effort.

    Benchmark SWE-2 Kimi K3 Grok 4.6 Fable 5.1 GPT-5.6 Sol GPT-6 Astra SWE-1.7
    FrontierCode 1.1 Main 50.0% 44.2% 48.0% 50.9% 47.5% 53.3% 42.0%
    DeepSWE 1.1 73.0% 68.5% 67.5% 67.4% 72.7% 74.1% 37.7%
    Terminal-Bench 2.1 92.8% 88.3% 88.4% 91.4% 88.8% 89.9% 81.5%
    Terminal-Bench 4 27.3% 21.5% 20.3% 55.8% 37.3% 57.9% 7.6%

    SWE-2 leads on Terminal-Bench 2.1 and beats its K3 base on every row. Cognition says it comes within a few points of GPT-6 Astra at a quarter of the cost. The clear weak spot is Terminal-Bench 4, where SWE-2 trails Fable 5.1 and GPT-6 Astra by roughly 30 points. FrontierCode is Cognition’s own benchmark, and all rival numbers come from Cognition’s evaluation.

    Model Behavior: Fewer Detours

    SWE-1.7 tended to over-explore on simple tasks. SWE-2 addresses this through what Cognition calls focused exploration. On FrontierCode 1.1 Main, SWE-2 medium scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less. Mean steps per run drop from 127 (SWE-1.7) to 53 (medium), 80 (high), and 98 (max). SWE-2 medium makes its first real edit after a median of 18 steps, versus 48 for SWE-1.7.

    Cognition team also reports 3 behavioral patterns: stronger end-to-end test coverage, resourcefulness when a tool is blocked, and verification discipline. When challenged, the model re-derives conclusions instead of re-asserting them.

    How It Was Trained

    Pareto-informed cost penalties: The reward is R = S minus lambda times C, where S is binary success and C mixes inference cost in USD with rollout time. Cognition proves that only a linear penalty makes the RL objective depend purely on average cost and solve rate. Each effort level’s lambda is set to the local slope of the base model’s Pareto curve. That makes the iso-reward line tangent to the frontier, so reward can only rise by pushing the frontier up.

    Length-weighted reward baseline: Cognition shares a baseline used since SWE-1.6. Gradient magnitude correlates strongly with rollout length, so the group baseline is weighted by tokens: sum(R x L) divided by sum(L). In ablations this kept inference-to-training KL divergence lower and stabilized training at no extra compute.

    Rollout serving and numerics: A prefill delayer batches nearby requests, raising TPM per GPU and TPS per request by 10 to 20%. DSpark speculative decoding accelerates rollouts, with a draft model retrained via SpecForge for 15% longer accept lengths and then trained online alongside the policy. NVFP4 and FP8 kernels with quantization-aware training keep memory usage down and train-inference mismatch below SWE-1.7 levels.

    Data: Cognition tripled its RL environments, added instruction-following overlays, and built a flywheel that uses earlier SWE-2 checkpoints to patch false positives and negatives in verifiers.

    Trustworthiness Checks

    Cognition reran 2 evaluations from its open-source trustworthiness study. On 145 politically sensitive questions about China, SWE-2 passed 98.0% overall: 99.8% in English, 95.2% in Simplified Chinese, and 99.1% in Traditional Chinese. On a context-dependent vulnerability test across customer framings, no framing produced a statistically significant change for any model.

    Interactive Explainer

    Coding Cognition cost Fable FrontierCode Kimi matches model PostTrained Releases SWE2
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

    Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its Own Controls Show the Wiring Does Not Help

    GPT-6 Astra Users Say OpenAI’s Newest Model Got Dumber. It Happened Before, Too

    Christopher Harborne matches £36m Reform donation of Ben Delo | Party funding

    Reform receives second £36m donation in two days as crypto investor matches record

    Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Burnham must understand Wales is ‘not a region in England’, says first minister | Rhun ap Iorwerth

    September 13, 2026

    UK government to offer clearer guidance on student loans after repayment row | Student finance

    September 13, 2026

    The Best 3-in-1 Apple Charging Stations After Testing 30+ Models

    September 13, 2026

    India launches tokenized bond pilot

    September 13, 2026
    Latest Posts

    Washington’s Badger Mountain Solar Project Canceled by Developer — ProPublica

    August 3, 2026

    Rejected Wisconsin data center proposal had guaranteed tax revenue, housing

    August 3, 2026

    EIG’s MidOcean Energy lines up new investment as NYK spreads its LNG wings

    August 3, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Burnham must understand Wales is ‘not a region in England’, says first minister | Rhun ap Iorwerth

    September 13, 2026

    UK government to offer clearer guidance on student loans after repayment row | Student finance

    September 13, 2026

    The Best 3-in-1 Apple Charging Stations After Testing 30+ Models

    September 13, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.