Close Menu
NCIJ Network NCIJ Network
    What's Hot

    UN Security Council Will Get Advice on AI Risks From Tech Giants Building It

    September 22, 2026

    Renewables: Grid shortages are forcing India to produce and waste green energy

    September 22, 2026

    Australia news live: Joyce says One Nation talking to Coalition about power-sharing agreement; fog causes Sydney airport delays | Australia news

    September 22, 2026
    Facebook X (Twitter) Instagram
    Trending
    • UN Security Council Will Get Advice on AI Risks From Tech Giants Building It
    • Renewables: Grid shortages are forcing India to produce and waste green energy
    • Australia news live: Joyce says One Nation talking to Coalition about power-sharing agreement; fog causes Sydney airport delays | Australia news
    • Chinese president arrives in US with new energy leverage as war rattles global markets
    • The Guardian view on Keynes in the West End: James Graham’s latest play captures the zeitgeist | Editorial
    • Far-right activist Daniel Thomas slashes dinghy in Channel with blade while rescue worker onboard | The far right
    • Paramount will need to release way more movies to make this merger work
    • Check Point Warns of Management Server Zero-Day Exploited in Targeted Attacks
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Tuesday, September 22
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Artificial Intelligence

    Microsoft’s SkillOpt Shows Optimized Agent Skill Artifacts Transfer Across Model Scales and Between Codex and Claude Code Harnesses

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKAugust 6, 2026 Artificial Intelligence No Comments5 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    SkillOpt is a text-space optimizer developed by a team of researchers from Microsoft, Shanghai Jiao Tong University, Tongji University, and Fudan University.

    SkillOpt trains a single natural-language skill document while the target model stays frozen. An optimizer model reads scored rollouts and proposes bounded add/delete/replace edits. A held-out selection split accepts an edit only when the score strictly improves. The exported artifact is one file, best_skill.md.

    The transfer tables report three columns. Baseline is the target’s no-skill score. Direct is SkillOpt trained in-domain on that exact target. Transferred applies a skill trained elsewhere, with no further optimization.

    The useful comparison is not transferred versus direct. It is how much of the in-domain gain survives the move.

    Cross-model transfer: within-family, mixed retention

    Skills were trained on GPT-5.4 and deployed on smaller variants.

    SpreadsheetBench GPT-5.4-mini 36.1 47.5 45.5 +9.4 82%
    SpreadsheetBench GPT-5.4-nano 23.5 42.5 26.5 +3.0 16%
    LiveMath GPT-5.4-mini 14.7 32.8 19.2 +4.5 25%
    LiveMath GPT-5.4-nano 23.2 27.2 28.8 +5.6 140%

    Two rows deserve attention. SpreadsheetBench on GPT-5.4-mini keeps 82% of the in-domain gain. That is close to free reuse. The LiveMath row on GPT-5.4-nano is stranger: the transferred skill scores 28.8 against an in-domain SkillOpt result of 27.2. The paper reads this as evidence that some learned procedures are target-model agnostic.

    The GPT-5.4-nano SpreadsheetBench row is the weak one at 16%. Retention is not uniform, and the paper does not claim it is. Its stated bound is narrower: no row falls below the target’s no-skill baseline.

    Note the scope. All four rows stay inside one GPT family. Cross-family transfer, such as GPT to Qwen, is not tested.

    Cross-harness transfer: the strongest result

    This is the section that matters most for deployment. All rows use GPT-5.5.

    Benchmark Source → Target Baseline Direct Transferred Gain Share of in-domain gain
    SpreadsheetBench Codex → Claude Code 22.1 80.4 81.8 +59.7 102%
    SpreadsheetBench Claude Code → Codex 27.5 85.0 71.1 +43.6 76%
    LiveMath Claude Code → Codex 35.2 78.4 48.0 +12.8 30%
    LiveMath Codex → Claude Code 40.8 56.5 42.4 +1.6 10%

    The first row is the headline. A skill optimized inside Codex lifted Claude Code from 22.1 to 81.8. That slightly exceeds the 80.4 Claude Code reached by training its own skill from scratch.

    The two harnesses expose different tool and file APIs and different command surfaces. A skill that survives that shift is not encoding command recipes. The research paper attributes SpreadsheetBench’s portability to workbook-level procedures: structure-first inspection, formula-aware verification, and static-value materialization. Those hold regardless of which CLI runs the Python.

    LiveMath tells the opposite story. Codex → Claude Code retains only 10% of the in-domain gain. The asymmetry is worth sitting with. Procedural skills — how to inspect, verify, and format — appear to be the portable class. Reasoning-heavy skills appear more tied to their training environment.

    Cross-benchmark transfer: real but small

    Source → Target Model Baseline Transferred Gain
    OlympiadBench → Omni-MATH GPT-5.4 56.6 60.3 +3.7
    OlympiadBench → Omni-MATH GPT-5.4-mini 34.8 36.6 +1.8
    OlympiadBench → Omni-MATH GPT-5.4-nano 38.8 40.1 +1.3

    There is no Direct column here. No in-domain SkillOpt run on Omni-MATH is reported, so the comparison is against no-skill only. Gains are positive across all three model scales but small. The research paper’s reading is that the skill retained reusable mathematical procedure after both the test instances and the answer-format conventions changed.

    Why the artifact moves at all

    The mechanism is stated plainly in the research paper. All three execution modes: direct chat, Codex, Claude Code – consume the same best_skill.md file format. That shared contract is what makes the cross-harness experiment possible in the first place.

    The Codex harness renders the current skill to a per-task SKILL.md alongside task files, then reads back a compact execution trace. The Claude Code harness mirrors the same workspace contract through the claude CLI. Neither harness gets a bespoke skill format.

    The artifact’s shape supports portability too. Final skills run 379 to 1,995 tokens across the six benchmarks, with a median near 920. They are assembled from 1 to 4 accepted edits. The paper’s Figure 4 samples one learned rule per benchmark, and all are procedural rather than instance-specific. The SpreadsheetBench rule, verbatim: inspect workbook structure and formulas, then write evaluated static values across the full requested target range instead of relying on Excel recalculation.

    What this implies for portability

    Training cost is paid once, offline, and measured. The research paper reports 0.6M to 46.4M training tokens per absolute test point, depending on benchmark. SpreadsheetBench sits at 0.6M per point; DocVQA at 46.4M. The optimizer model runs only during training and adds zero inference-time calls at deployment.

    If a skill trained in one harness holds up in another, that one-time cost spreads across environments. The Codex → Claude Code SpreadsheetBench result is the existence proof. It also implies you can optimize where tooling is cheapest and deploy where the product lives.

    The audit angle is separate and underrated. The deployed artifact is a text file a domain practitioner can read in minutes. Every change to it is traceable: each step records an edit_apply_report.json with per-edit accept and skip status. Portability plus inspectability is a different operational posture than shipping fine-tuned weights.

    Key Takeaways

    • Evidence covers one GPT family and two benchmarks per axis, so portability is demonstrated, not yet generalized.
    • A Codex-trained SpreadsheetBench skill scored 81.8 inside Claude Code, above that harness’s own 80.4 in-domain result.
    • All 4 cross-model, 4 cross-harness, and 3 cross-benchmark transfer rows land above the target’s no-skill baseline.
    • Transfer strength tracks task type: procedural spreadsheet skills move well, math-reasoning skills move weakly.
    • The portable unit is one best_skill.md of 379 to 1,995 tokens, built from 1 to 4 accepted edits.

    Resources: Paper, GitHub, Project page, Docs, PyPI and Demo video

    Baselines referenced: GEPA, TextGrad, EvoSkill and Trace2Skill

    Benchmarks referenced: SearchQA, SpreadsheetBench, DocVQA, LiveMathematicianBench and ALFWorld


    Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

    agent Artifacts Claude Code Codex Harnesses Microsofts model Optimized Scales shows Skill SkillOpt transfer
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Anthropic Releases Claude Opus 5.5: Fable 5.1-Level Performance at 40% Lower Running Cost Than Opus 5

    Claude Opus 5.5 delivers Fable 5.1 performance – and costs 40% less

    Rabbit Is Back, This Time With an AI Agent App

    Can AI reason without words? A small model puts the idea to the test

    Did Trump’s kindness ‘ruin’ Secret Service agent? Don’t fall for tall tale

    AutoScheduler launches warehouse app builder for logistics teams

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    UN Security Council Will Get Advice on AI Risks From Tech Giants Building It

    September 22, 2026

    Renewables: Grid shortages are forcing India to produce and waste green energy

    September 22, 2026

    Australia news live: Joyce says One Nation talking to Coalition about power-sharing agreement; fog causes Sydney airport delays | Australia news

    September 22, 2026

    Chinese president arrives in US with new energy leverage as war rattles global markets

    September 22, 2026
    Latest Posts

    COLDCARD security audit phishing attack installs remote access tool

    August 5, 2026

    Reddit aims to make ‘karma’ less important for first-time posters with shift to AI moderation tools

    August 5, 2026

    Right turn on green: is the Telegraph changing its tune on the climate? | Daily Telegraph

    August 5, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    UN Security Council Will Get Advice on AI Risks From Tech Giants Building It

    September 22, 2026

    Renewables: Grid shortages are forcing India to produce and waste green energy

    September 22, 2026

    Australia news live: Joyce says One Nation talking to Coalition about power-sharing agreement; fog causes Sydney airport delays | Australia news

    September 22, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.