Close Menu
NCIJ Network |NCIJ Network |
    What's Hot

    Chevron’s Cypriot gas project edges closer toward binding supply deal with Egypt

    July 23, 2026

    Keystone clashes: Millions pour into three Pennsylvania races that could decide control of the House • OpenSecrets

    July 23, 2026

    Three speeches on a single day signaled a dying American democracy | Robert B Shpiner

    July 23, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Chevron’s Cypriot gas project edges closer toward binding supply deal with Egypt
    • Keystone clashes: Millions pour into three Pennsylvania races that could decide control of the House • OpenSecrets
    • Three speeches on a single day signaled a dying American democracy | Robert B Shpiner
    • Wildfires ravage Spain, France and Italy, killing three firefighters
    • ‘Falling apart’: Trump’s Boeing deal hits turbulence with Beijing
    • Clarity Act Mired In Debate Over Whether to Bar President From Selling Crypto
    • Funding bus fare cap from aid budget will hit world’s poorest, Burnham told | Transport policy
    • Teenager drops social media addiction lawsuit against Meta
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network |NCIJ Network |
    Thursday, July 23
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network |NCIJ Network |
    Home»Artificial Intelligence

    Research-Grade EdgeBench Analysis: AI Agent Benchmarking, Leaderboard Analytics, Scaling Laws, and Evaluation Metrics

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKJuly 22, 2026 Artificial Intelligence No Comments8 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    In this tutorial, we explore EdgeBench as a practical benchmark for evaluating advanced AI agents across diverse task categories, runtime environments, and interaction-time budgets. We begin by downloading the dataset snapshot from Hugging Face, parsing the released task specifications, and examining the benchmark taxonomy, execution settings, internet requirements, judging logic, and scoring metadata. We then extract the leaderboard data directly from the repository README, standardize model names, reshape task-level results into an analysis-ready format, and compare performance across multiple time budgets. Finally, we fit log-sigmoid scaling curves, measure category-level score improvements, inspect the tasks with the largest gains, and study how SForge rescale functions transform raw evaluation outputs into normalized benchmark scores.

    !pip -q install "huggingface_hub>=0.23" pandas numpy scipy matplotlib pyyaml
    import os, glob, json, textwrap, warnings
    import numpy as np, pandas as pd
    import matplotlib.pyplot as plt
    from scipy.optimize import curve_fit
    from huggingface_hub import snapshot_download
    warnings.filterwarnings("ignore")
    pd.set_option("display.max_colwidth", 90)
    pd.set_option("display.width", 160)
    REPO_ID = "ByteDance-Seed/EdgeBench"
    TIME_BUDGETS = [2, 4, 6, 8, 10, 12]
    MODELS = ["Claude Opus 4.8", "GPT-5.5", "GPT-5.4", "GLM-5.1", "DS-V4-Pro"]
    def banner(t):
       print("n" + "=" * 78 + f"n  {t}n" + "=" * 78)
    def canon_model(name):
       n = name.lower()
       if "opus" in n:
           return "Claude Opus 4.8"
       if "5.5" in n:
           return "GPT-5.5"
       if "5.4" in n:
           return "GPT-5.4"
       if "glm" in n:
           return "GLM-5.1"
       if "ds-v4" in n or "deepseek" in n:
           return "DS-V4-Pro"
       return name
    banner("1. DOWNLOADING DATASET SNAPSHOT")
    local_dir = snapshot_download(repo_id=REPO_ID, repo_type="dataset")
    print("Snapshot cached at:", local_dir)
    

    We install the required Python libraries, import the analytical tools, and configure the notebook display settings. We define the dataset repository, interaction-time budgets, model names, and helper functions to format output and standardize model labels. We then download the complete EdgeBench dataset snapshot from Hugging Face and store its local cache path for the remainder of the workflow.

    banner("2. LOADING TASK SPECIFICATIONS")
    def flatten_task(d):
       work, judge = d.get("work", {}) or {}, d.get("judge", {}) or {}
       rescale = judge.get("rescale", {}) or {}
       return {
           "task_id": d.get("task_id"),
           "name": d.get("name"),
           "category": d.get("category"),
           "base_image": d.get("base_image"),
           "internet": d.get("internet"),
           "game_mode": d.get("game_mode", False),
           "cwd": d.get("cwd"),
           "n_submit": len(d.get("submit_paths", []) or []),
           "submit_paths": ", ".join(d.get("submit_paths", []) or []) or "(interactive)",
           "parser": judge.get("parser") or "(game)",
           "score_dir": judge.get("score_direction", "n/a"),
           "selection": judge.get("selection"),
           "eval_timeout": judge.get("eval_timeout"),
           "rescale_kind": rescale.get("kind"),
           "rescale": rescale,
           "agent_query": work.get("agent_query", "")
       }
    records = []
    for fp in sorted(glob.glob(os.path.join(local_dir, "*.json"))):
       try:
           with open(fp) as f:
               records.append(flatten_task(json.load(f)))
       except Exception as e:
           print("  ! skipped", os.path.basename(fp), "->", e)
    df = pd.DataFrame(records).dropna(subset=["task_id"]).reset_index(drop=True)
    print(f"Loaded {len(df)} task specifications.n")
    print(df[["task_id", "category", "base_image", "internet", "rescale_kind"]].head(10).to_string(index=False))
    banner("3. BENCHMARK TAXONOMY")
    print("Tasks per category:n", df["category"].value_counts().to_string(), "n")
    print("Runtime (base_image):n", df["base_image"].value_counts().to_string(), "n")
    print("Judge rescale kinds:n", df["rescale_kind"].value_counts(dropna=False).to_string(), "n")
    print(f"Tasks needing internet: {int(df['internet'].sum())} | game_mode tasks: {int(df['game_mode'].sum())}")
    fig, ax = plt.subplots(1, 2, figsize=(13, 4.2))
    df["category"].value_counts().plot.barh(ax=ax[0], color="#4C78A8")
    ax[0].set_title("Released tasks per category (51)")
    ax[0].invert_yaxis()
    df["base_image"].value_counts().plot.bar(ax=ax[1], color="#F58518")
    ax[1].set_title("Runtime environment")
    ax[1].tick_params(axis="x", rotation=45)
    plt.tight_layout()
    plt.show()
    banner("3b. ANATOMY OF ONE TASK")
    s = df.iloc[0]
    print(f"task_id: {s.task_id} | category: {s.category} | base_image: {s.base_image}")
    print(f"judge parser: {s.parser} | rescale: {s.rescale_kind} -> {s.rescale}")
    print("n--- agent_query (truncated) ---")
    print(textwrap.fill(s.agent_query[:800], width=96))
    

    We parse every task specification into a structured table containing its category, runtime image, internet access, submission paths, judge configuration, and agent query. We summarize the benchmark taxonomy by counting tasks across categories, execution environments, rescaling methods, and game modes. We also visualize these distributions and inspect one representative task to understand how an EdgeBench evaluation is defined.

    banner("4. PARSING THE LEADERBOARD")
    readme = open(os.path.join(local_dir, "README.md"), encoding="utf-8").read()
    def unescape(x):
       return x.replace("\_", "_").replace("\", "").replace("*", "").strip()
    def to_float(x):
       x = x.replace("*", "").strip()
       return np.nan if x in ("", "—", "-") else float(x)
    def extract_md_tables(md):
       tables, cur = [], []
       for ln in md.splitlines():
           s = ln.strip()
           if s.startswith("|"):
               cur.append([unescape(c) for c in s.strip("|").split("|")])
           elif cur:
               tables.append(cur)
               cur = []
       if cur:
           tables.append(cur)
       return [[r for r in t if not all(set(c) <= set("-: ") for c in r)] for t in tables if t]
    tables = extract_md_tables(readme)
    def parse_series(cell):
       parts = cell.split("/")
       if len(parts) != len(TIME_BUDGETS):
           return None
       try:
           return [to_float(p) for p in parts]
       except ValueError:
           return None
    long_rows = []
    for tbl in tables:
       head = tbl[0]
       if head and head[0].lower() == "task" and any("categ" in h.lower() for h in head):
           model_cols = [canon_model(m) for m in head[2:]]
           for row in tbl[1:]:
               if len(row) != len(head):
                   continue
               for mname, cell in zip(model_cols, row[2:]):
                   series = parse_series(cell)
                   if series is None:
                       continue
                   for t, sc in zip(TIME_BUDGETS, series):
                       long_rows.append({
                           "task": row[0],
                           "category": row[1],
                           "model": mname,
                           "hours": t,
                           "score": sc
                       })
    scores = pd.DataFrame(long_rows)
    print(
       f"Parsed {scores['task'].nunique()} tasks x "
       f"{scores['model'].nunique()} models x "
       f"{len(TIME_BUDGETS)} budgets = {len(scores)} cells."
    )
    agg_time, groups, cur = [], [], []
    for tbl in tables:
       head = tbl[0]
       if head and "model" in head[0].lower() and any("@2h" in h for h in head):
           cols = head[1:]
           for row in tbl[1:]:
               if len(row) == len(head):
                   rec = {"model": canon_model(row[0])}
                   rec.update({c: to_float(v) for c, v in zip(cols, row[1:])})
                   cur.append(rec)
           groups.append(cur)
           cur = []
    agg51 = pd.DataFrame(groups[1] if len(groups) > 1 else (groups[0] if groups else []))
    if not agg51.empty:
       print("nREADME aggregate (51-task subset):")
       print(agg51.to_string(index=False))
    

    We read the repository README and extract its Markdown tables into structured Python records. We convert the task-level leaderboard into a tidy dataset containing task, category, model, interaction time, and score values. We also parse the aggregate 51-task leaderboard table to compare the README summary with our task-level calculations.

    banner("5. LOG-SIGMOID SCALING LAW (fit on per-task means -> robust)")
    def log_sigmoid(t, lo, hi, k, t0):
       return lo + (hi - lo) / (1.0 + np.exp(-k * (np.log(t) - np.log(t0))))
    def r2(y, yhat):
       ssr = np.nansum((y - yhat) ** 2)
       sst = np.nansum((y - np.nanmean(y)) ** 2)
       return 1 - ssr / sst if sst > 0 else np.nan
    agg = (
       scores.groupby(["model", "hours"])["score"]
       .mean()
       .unstack("hours")
       .reindex(index=MODELS)[TIME_BUDGETS]
    )
    print("Per-task mean by model & hour:n", agg.round(2).to_string(), "n")
    t_h = np.array(TIME_BUDGETS, float)
    t_dense = np.linspace(2, 12, 200)
    fig, ax = plt.subplots(figsize=(9, 5.5))
    for color, model in zip(plt.cm.tab10(np.linspace(0, 1, len(MODELS))), MODELS):
       if model not in agg.index:
           continue
       y = agg.loc[model].values.astype(float)
       try:
           popt, _ = curve_fit(
               log_sigmoid,
               t_h,
               y,
               p0=[y[0], y[-1] + 1, 2.0, 4.0],
               bounds=([-50, -50, 1e-2, 0.5], [200, 200, 50, 50]),
               maxfev=200000,
           )
           rr = r2(y, log_sigmoid(t_h, *popt))
           ax.plot(
               t_dense,
               log_sigmoid(t_dense, *popt),
               color=color,
               lw=2,
               label=f"{model} (R²={rr:.3f})",
           )
       except Exception as e:
           print("  fit failed for", model, e)
       ax.scatter(t_h, y, color=color, s=45, zorder=3)
    ax.set_xlabel("Interaction time (hours)")
    ax.set_ylabel("EdgeBench score (51-task mean)")
    ax.set_title("Log-sigmoid scaling of score vs. interaction time")
    ax.legend(fontsize=8)
    ax.grid(alpha=0.3)
    plt.tight_layout()
    plt.show()
    piv = scores.pivot_table(
       index=["task", "category"],
       columns=["model", "hours"],
       values="score",
       aggfunc="mean",
    )
    top = MODELS[0]
    uplift = (
       piv[(top, 12)] - piv[(top, 2)]
    ).groupby("category").mean().sort_values(ascending=False)
    print(f"nMean 2h→12h uplift for {top}, by category:n", uplift.round(2).to_string())
    focus = (
       piv[(top, 12)] - piv[(top, 2)]
    ).sort_values(ascending=False).head(5).index
    fig, ax = plt.subplots(figsize=(9, 5))
    for task, cat in focus:
       ax.plot(
           TIME_BUDGETS,
           [piv.loc[(task, cat)][(top, h)] for h in TIME_BUDGETS],
           marker="o",
           label=task[:32],
       )
    ax.set_title(f"{top}: per-task learning trajectories (largest gains)")
    ax.set_xlabel("Interaction time (hours)")
    ax.set_ylabel("Task score")
    ax.legend(fontsize=8)
    ax.grid(alpha=0.3)
    plt.tight_layout()
    plt.show()
    

    We calculate the mean score for each model at every interaction-time budget and fit a log-sigmoid scaling curve to the resulting trajectories. We evaluate the goodness of each fit using the coefficient of determination and visualize both the observed scores and fitted curves. We then measure category-level score gains and plot the individual tasks that benefit most from longer interaction time.

    banner("6. SForge SCORING 'RESCALE' FUNCTIONS")
    def rescale_linear(raw, lower, upper):
       return float(np.clip((raw - lower) / (upper - lower) * 100.0, 0, 100))
    for kind in ["linear", "piecewise_max"]:
       ex = df[df["rescale_kind"] == kind]
       if ex.empty:
           continue
       row = ex.iloc[0]
       rs = row["rescale"]
       print(f"n[{kind}] example: {row.task_id}n  params: {rs}")
       if kind == "linear" and "lower" in rs and "upper" in rs:
           for raw in [rs["lower"], (rs["lower"] + rs["upper"]) / 2, rs["upper"]]:
               print(
                   f"    raw={raw:>12.2f} -> "
                   f"{rescale_linear(raw, rs['lower'], rs['upper']):6.2f}"
               )
       else:
           xs = [rs["baseline"], rs["rank30"], rs["rank1"], rs["super_anchor"]]
           ys = [0.0, 70.0, 99.0, 100.0]
           for raw in xs:
               print(
                   f"    raw={raw:>16.1f} -> "
                   f"{float(np.interp(raw, xs, ys)):6.2f} (illustrative)"
               )
    banner("DONE")
    print("Real evaluation runs via the SForge two-container harness:")
    print("  https://github.com/ByteDance-Seed/EdgeBench  |  https://bytedance-seed.github.io/EdgeBench/")
    

    We examine the scoring rescale configurations used by the SForge judging system and implement the linear normalization formula. We demonstrate how raw scores map to normalized benchmark scores for both linear and piecewise-style configurations. We finish the tutorial by printing the official EdgeBench and SForge resources required for running complete evaluations.

    In conclusion, we built a complete analytical workflow for understanding both the structure and the reported results of EdgeBench. We moved beyond simply viewing a leaderboard by connecting task specifications, runtime configurations, scoring rules, model performance, and interaction-time scaling within a single reproducible Colab pipeline. We also identified where additional interaction time yields the greatest performance improvements and visualized how different models scale across the released benchmark tasks. It provides a technically grounded foundation for interpreting EdgeBench results, comparing agent capabilities, and preparing for deeper evaluation using the full SForge execution harness.


    Check out the Full Code here. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

    Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us


    Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.

    agent Analysis Analytics Benchmarking EdgeBench evaluation Laws Leaderboard metrics ResearchGrade Scaling
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Cursor Releases Cursor Router: A Request-Level Classifier Delivering Frontier Coding Quality at 30–50% Lower Cost

    Google Releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber: A Cheaper, More Token-Efficient Flash Tier Built for Agentic Workloads

    Poolside Releases Laguna S 2.1, an Open-Weight Agentic Coding Model Punching Above Its Weight Class on SWE-Bench Multilingual

    David Vélez and Robin Vince join the boards of the OpenAI Foundation and OpenAI Group PBC

    Cisco Foundation AI Releases Antares: 350M and 1B Open-Weight Models That Localize Known Vulnerabilities Inside Real Codebases

    Advancing the next era of national science

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Chevron’s Cypriot gas project edges closer toward binding supply deal with Egypt

    July 23, 2026

    Keystone clashes: Millions pour into three Pennsylvania races that could decide control of the House • OpenSecrets

    July 23, 2026

    Three speeches on a single day signaled a dying American democracy | Robert B Shpiner

    July 23, 2026

    Wildfires ravage Spain, France and Italy, killing three firefighters

    July 23, 2026
    Latest Posts

    Trump slaps 50% tariffs on Canada and Carney vows to ‘intensify’ trade talks

    July 21, 2026

    How Two Brothers Dug for Dead Relatives: With a Shovel and a Kitchen Knife

    July 21, 2026

    Chile floods: Towns evacuated following heavy rain in Coquimbo

    July 21, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Chevron’s Cypriot gas project edges closer toward binding supply deal with Egypt

    July 23, 2026

    Keystone clashes: Millions pour into three Pennsylvania races that could decide control of the House • OpenSecrets

    July 23, 2026

    Three speeches on a single day signaled a dying American democracy | Robert B Shpiner

    July 23, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.