Close Menu
NCIJ Network NCIJ Network
    What's Hot

    A mine polluted a Zambian river in 2025: Residents continue to live with the impacts

    August 14, 2026

    North Korea fumes over upcoming US-South Korea military drills | Military News

    August 14, 2026

    The job Zelenskyy can’t fill: Ukraine’s envoy to Trump’s Washington  – POLITICO

    August 14, 2026
    Facebook X (Twitter) Instagram
    Trending
    • A mine polluted a Zambian river in 2025: Residents continue to live with the impacts
    • North Korea fumes over upcoming US-South Korea military drills | Military News
    • The job Zelenskyy can’t fill: Ukraine’s envoy to Trump’s Washington  – POLITICO
    • Farage’s by-election victory won’t stop questions about finances
    • JPMorgan debanked Polymarket over regulatory concerns
    • Babbel Promo Code: Up to 65% Off in August 2026
    • Meet Needle 2: An Open 45M-Parameter Tool-Calling Model That Ships as a 14MB Binary and Runs a Full Session in 28MB of RAM
    • Hackers breach govt webmail while running parallel crypto fraud
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Friday, August 14
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Artificial Intelligence

    Create a Reasoning-Focused LLM: A Practical Guide to Streaming, Curating, and Fine-Tuning the SupraLabs Reasoning Corpus

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKAugust 14, 2026 Artificial Intelligence No Comments7 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    In this tutorial, we build an end-to-end workflow for working with the SupraLabs reasoning corpus. We stream a representative subset directly from the Hugging Face Hub, inspect its source distribution, token-length patterns, task composition, and reasoning-to-answer ratios, and then apply a series of quality filters to remove unsuitable training examples. We transform the retained samples into a chat-based supervised fine-tuning format with explicit reasoning tags and use them to adapt SmolLM2-135M-Instruct with LoRA through TRL’s SFTTrainer. By combining scalable data access, exploratory analysis, dataset curation, parameter-efficient fine-tuning, structured inference, and Parquet export, we create a complete Google Colab pipeline for turning a large multi-model reasoning corpus into a compact reasoning-focused language model.

    import subprocess, sys
    def pip_install(pkgs):
       subprocess.check_call([sys.executable, "-m", "pip", "install", "-q", *pkgs])
    subprocess.call([sys.executable, "-m", "pip", "uninstall", "-y", "-q", "torchao"])
    pip_install([
       "datasets>=3.0.0",
       "transformers>=4.46.0",
       "trl>=0.12.0",
       "peft>=0.13.0",
       "accelerate>=1.0.0",
       "bitsandbytes",
       "matplotlib",
       "pandas",
    ])
    import os, re, json, math, random, itertools, warnings
    import pandas as pd
    import matplotlib.pyplot as plt
    import torch
    from collections import Counter
    from datasets import load_dataset, Dataset
    warnings.filterwarnings("ignore")
    random.seed(42)
    torch.manual_seed(42)
    DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
    print(f"Device: {DEVICE}")
    if DEVICE == "cuda":
       print(f"GPU: {torch.cuda.get_device_name(0)}")
    DATASET_ID = "SupraLabs/reasoning-corpus-4K-5M-v1"
    SAMPLE_SIZE = 8_000
    print(f"nStreaming {DATASET_ID} ...")
    stream = load_dataset(DATASET_ID, split="train", streaming=True)
    stream = stream.shuffle(seed=42, buffer_size=30_000)
    rows = list(itertools.islice(stream, SAMPLE_SIZE))
    ds = Dataset.from_list(rows)
    print(f"Materialized sample: {len(ds):,} rows")
    print(f"Columns: {ds.column_names}")
    ex = ds[0]
    print("n" + "=" * 70)
    print("EXAMPLE ROW")
    print("=" * 70)
    print(f"repo_id : {ex['repo_id']}")
    print(f"tok_len : {ex['tok_len']}")
    print(f"user            : {ex['user'][:300]} ...")
    print(f"thought_trace   : {ex['thought_trace'][:300]} ...")
    print(f"assistant       : {ex['assistant'][:300]} ...")
    

    We configure the Colab environment, install the required machine learning libraries, and remove the incompatible torchao package. We detect the available compute device, connect to the SupraLabs reasoning corpus through Hugging Face streaming, and avoid downloading the complete dataset. We shuffle the streamed records, materialize a representative sample, and inspect the structure and contents of an example row.

    df = ds.to_pandas()
    print("nTop 15 source repos in sample:")
    src_counts = df["repo_id"].value_counts()
    print(src_counts.head(15).to_string())
    fig, axes = plt.subplots(2, 2, figsize=(14, 10))
    axes[0, 0].hist(df["tok_len"], bins=60, color="#4C72B0", edgecolor="white")
    axes[0, 0].set_title("Token length distribution")
    axes[0, 0].set_xlabel("tok_len"); axes[0, 0].set_ylabel("rows")
    src_counts.head(12).plot(kind="barh", ax=axes[0, 1], color="#55A868")
    axes[0, 1].invert_yaxis()
    axes[0, 1].set_title("Top-12 source repos (sample)")
    df["think_chars"] = df["thought_trace"].str.len()
    df["answer_chars"] = df["assistant"].str.len()
    df["reason_ratio"] = df["think_chars"] / (df["think_chars"] + df["answer_chars"] + 1)
    axes[1, 0].hist(df["reason_ratio"], bins=50, color="#C44E52", edgecolor="white")
    axes[1, 0].set_title("Reasoning ratio  (think / (think + answer))")
    axes[1, 0].set_xlabel("ratio")
    axes[1, 1].scatter(df["tok_len"], df["reason_ratio"], s=4, alpha=0.25, color="#8172B2")
    axes[1, 1].set_title("tok_len vs reasoning ratio")
    axes[1, 1].set_xlabel("tok_len"); axes[1, 1].set_ylabel("ratio")
    plt.tight_layout()
    plt.show()
    print("nSummary stats:")
    print(df[["tok_len", "think_chars", "answer_chars", "reason_ratio"]]
         .describe().round(2).to_string())
    def tag_task(row):
       u = row["user"].lower()
       a = row["assistant"]
       if "```" in a or re.search(r"b(def |class |import |function|#include)", a):
           return "code"
       if re.search(r"(prove|equation|integral|theorem|\frac|\int|solve for)", u):
           return "math"
       if re.search(r"b(patient|diagnosis|symptom|treatment|clinical)b", u):
           return "medical"
       if re.search(r"b(which of the following|options?:|(a)|(b))", u):
           return "mcq/logic"
       return "general"
    df["task"] = df.apply(tag_task, axis=1)
    print("nHeuristic task mix:")
    print(df["task"].value_counts(normalize=True).round(3).to_string())
    

    We convert the sampled dataset into a pandas DataFrame and analyze the distribution of source repositories and token lengths. We calculate reasoning and answer character counts, measure the reasoning-to-response ratio, and visualize the relationships across the dataset. We also apply lightweight heuristic rules to classify each record as a code, mathematics, medical, multiple-choice, or general task.

    def filter_length(row, min_tok=200, max_tok=3000):
       """Keep samples within a training-friendly token budget."""
       return min_tok <= row["tok_len"] <= max_tok
    def filter_degenerate(row):
       """Drop empty/near-empty thoughts or answers."""
       return len(row["thought_trace"]) > 100 and len(row["assistant"]) > 20
    def filter_repetition(row, max_line_repeat=0.30):
       """Drop traces where one line repeats too often (looping models)."""
       lines = [l.strip() for l in row["thought_trace"].split("n") if l.strip()]
       if len(lines) < 5:
           return True
       most_common = Counter(lines).most_common(1)[0][1]
       return (most_common / len(lines)) <= max_line_repeat
    def filter_reason_ratio(row, lo=0.15, hi=0.97):
       """Keep samples that actually reason but don't ONLY reason."""
       t, a = len(row["thought_trace"]), len(row["assistant"])
       r = t / (t + a + 1)
       return lo <= r <= hi
    n0 = len(ds)
    ds_f = ds.filter(filter_length)
    ds_f = ds_f.filter(filter_degenerate)
    ds_f = ds_f.filter(filter_repetition)
    ds_f = ds_f.filter(filter_reason_ratio)
    print(f"nFiltering: {n0:,} -> {len(ds_f):,} rows "
         f"({100 * len(ds_f) / n0:.1f}% retained)")
    MODEL_ID = "HuggingFaceTB/SmolLM2-135M-Instruct"
    from transformers import AutoTokenizer, AutoModelForCausalLM
    tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
    if tokenizer.pad_token is None:
       tokenizer.pad_token = tokenizer.eos_token
    SYSTEM_PROMPT = (
       "You are a careful reasoning assistant. Think step by step inside "
       "... tags, then give your final answer."
    )
    def to_chat(row):
       return {
           "messages": [
               {"role": "system", "content": SYSTEM_PROMPT},
               {"role": "user", "content": row["user"]},
               {"role": "assistant",
                "content": f"n{row['thought_trace']}nnn{row['assistant']}"},
           ]
       }
    train_ds = ds_f.map(to_chat, remove_columns=ds_f.column_names)
    train_ds = train_ds.shuffle(seed=42)
    N_TRAIN, N_EVAL = 1_500, 100
    eval_ds = train_ds.select(range(N_TRAIN, min(N_TRAIN + N_EVAL, len(train_ds))))
    train_ds = train_ds.select(range(min(N_TRAIN, len(train_ds))))
    print(f"nTrain: {len(train_ds):,}  |  Eval: {len(eval_ds):,}")
    print("nRendered training sample (truncated):")
    print(tokenizer.apply_chat_template(train_ds[0]["messages"], tokenize=False)[:800])
    

    We construct a quality-filtering pipeline that removes samples with unsuitable token lengths, incomplete responses, excessive repetition, or unbalanced reasoning content. We load the SmolLM2 tokenizer and transform each retained record into a structured conversation containing a system prompt, user message, and reasoning-enhanced assistant response. We then shuffle the formatted data, create training and evaluation subsets, and inspect the final chat template used for supervised fine-tuning.

    from trl import SFTTrainer, SFTConfig
    from peft import LoraConfig
    try:
       import peft.import_utils as _piu
       import peft.tuners.lora.torchao as _plt
       _piu.is_torchao_available = lambda: False
       _plt.is_torchao_available = lambda: False
    except Exception:
       pass
    model = AutoModelForCausalLM.from_pretrained(
       MODEL_ID,
       dtype=torch.bfloat16 if DEVICE == "cuda" else torch.float32,
    ).to(DEVICE)
    peft_config = LoraConfig(
       r=16,
       lora_alpha=32,
       lora_dropout=0.05,
       bias="none",
       task_type="CAUSAL_LM"
    sft_config = SFTConfig(
       output_dir="smollm2-reasoning-demo",
       max_length=2048,
       per_device_train_batch_size=2,
       gradient_accumulation_steps=8,
       num_train_epochs=1,
       learning_rate=2e-4,
       lr_scheduler_type="cosine",
       warmup_steps=10,
       logging_steps=10,
       eval_strategy="steps",
       eval_steps=50,
       save_strategy="no",
       bf16=(DEVICE == "cuda"),
       gradient_checkpointing=True,
       report_to="none",
    )
    trainer = SFTTrainer(
       model=model,
       args=sft_config,
       train_dataset=train_ds,
       eval_dataset=eval_ds,
       peft_config=peft_config,
       processing_class=tokenizer,
    )
    print("nStarting fine-tune (≈10–20 min on a T4 with these settings)...")
    trainer.train()
    print("Done. Final eval loss:", trainer.evaluate().get("eval_loss"))
    

    We load the SmolLM2 causal language model and configure LoRA adapters for parameter-efficient training. We define the optimization, batching, evaluation, precision, and gradient-checkpointing settings through TRL’s SFTConfig. We initialize the SFTTrainer, fine-tune the model on the curated reasoning conversations, and evaluate its final training performance.

    def generate(question, max_new_tokens=512, temperature=0.7):
       msgs = [
           {"role": "system", "content": SYSTEM_PROMPT},
           {"role": "user", "content": question},
       ]
       prompt = tokenizer.apply_chat_template(
           msgs, tokenize=False, add_generation_prompt=True
       )
       inputs = tokenizer(prompt, return_tensors="pt").to(DEVICE)
       with torch.no_grad():
           out = trainer.model.generate(
               **inputs,
               max_new_tokens=max_new_tokens,
               temperature=temperature,
               top_p=0.9,
               do_sample=True,
               pad_token_id=tokenizer.pad_token_id,
           )
       text = tokenizer.decode(out[0][inputs["input_ids"].shape[1]:],
                               skip_special_tokens=True)
       m = re.search(r"(.*?)(.*)", text, re.DOTALL)
       if m:
           print("─" * 60, "nTHINKING:n", m.group(1).strip()[:1500])
           print("─" * 60, "nANSWER:n", m.group(2).strip())
       else:
           print(text)
    print("nn### TEST 1: logic puzzle")
    generate("If all bloops are razzies and all razzies are lazzies, "
            "are all bloops definitely lazzies? Explain briefly.")
    train_ds.to_parquet("reasoning_subset_train.parquet")
    eval_ds.to_parquet("reasoning_subset_eval.parquet")
    print("nSaved: reasoning_subset_train.parquet / reasoning_subset_eval.parquet")
    

    We create an inference function that formats new questions with the same system prompt and generates responses from the fine-tuned model. We separate the generated section from the final answer and test the model on logic and arithmetic problems. We finally export the processed training and evaluation datasets as Parquet files for reuse in larger experiments.

    In conclusion, we developed a practical pipeline that connects large-scale reasoning-data exploration with small-language-model training. We streamed the corpus efficiently, analyzed its internal composition, filtered examples using token, repetition, completeness, and reasoning-balance criteria, and converted the resulting data into a consistent conversational training structure. We then fine-tuned SmolLM2 with LoRA, evaluated the adapted model, inspected its generated reasoning and answers, and exported the curated datasets for future experiments. This workflow provides a reusable foundation for source-aware data mixing, curriculum learning, larger student models, longer-context training, and production-scale reasoning model development without requiring the entire dataset to reside in Colab memory.


    Check out the FULL CODES here. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

    Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us


    Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.

    Corpus create Curating FineTuning Guide LLM Practical Reasoning ReasoningFocused Streaming SupraLabs
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Meet Needle 2: An Open 45M-Parameter Tool-Calling Model That Ships as a 14MB Binary and Runs a Full Session in 28MB of RAM

    Dyna Robotics Introduces Dyna-2: A World-Action Model Pre-Trained on 1 Million Hours of Human Video

    Liquid AI Releases LFM2.5-VL-3B: A 3B Vision-Language Model That Reads Screens, Grounds Objects, and Calls Tools On-Device

    Google AI Just Released Gemini 3.7 Flash: A Coding and Agent Model at $0.75/1M Input Tokens

    How RingCentral builds AI-native work from engineering to ops

    Listening shaped our homelessness resource guide for northeast Wisconsin

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    A mine polluted a Zambian river in 2025: Residents continue to live with the impacts

    August 14, 2026

    North Korea fumes over upcoming US-South Korea military drills | Military News

    August 14, 2026

    The job Zelenskyy can’t fill: Ukraine’s envoy to Trump’s Washington  – POLITICO

    August 14, 2026

    Farage’s by-election victory won’t stop questions about finances

    August 14, 2026
    Latest Posts

    Evacuated villagers in Cairngorms allowed home after wildfire threat lifts | Wildfires

    July 25, 2026

    The Fraternal Order Of Police Supports The Clarity Act.

    July 25, 2026

    How Synthetic Identity Fraud is Coming for Machine Identities

    July 25, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    A mine polluted a Zambian river in 2025: Residents continue to live with the impacts

    August 14, 2026

    North Korea fumes over upcoming US-South Korea military drills | Military News

    August 14, 2026

    The job Zelenskyy can’t fill: Ukraine’s envoy to Trump’s Washington  – POLITICO

    August 14, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.