• About Us
  • Disclaimer
  • Contact Us
  • Privacy Policy
Thursday, October 1, 2026
mGrowTech
No Result
View All Result
  • Technology And Software
    • Account Based Marketing
    • Channel Marketing
    • Marketing Automation
      • Al, Analytics and Automation
      • Ad Management
  • Digital Marketing
    • Social Media Management
    • Google Marketing
  • Direct Marketing
    • Brand Management
    • Marketing Attribution and Consulting
  • Mobile Marketing
  • Event Management
  • PR Solutions
  • Technology And Software
    • Account Based Marketing
    • Channel Marketing
    • Marketing Automation
      • Al, Analytics and Automation
      • Ad Management
  • Digital Marketing
    • Social Media Management
    • Google Marketing
  • Direct Marketing
    • Brand Management
    • Marketing Attribution and Consulting
  • Mobile Marketing
  • Event Management
  • PR Solutions
No Result
View All Result
mGrowTech
No Result
View All Result
Home Al, Analytics and Automation

Create a Reasoning-Focused LLM: A Practical Guide to Streaming, Curating, and Fine-Tuning the SupraLabs Reasoning Corpus

Josh by Josh
August 16, 2026
in Al, Analytics and Automation
0
Create a Reasoning-Focused LLM: A Practical Guide to Streaming, Curating, and Fine-Tuning the SupraLabs Reasoning Corpus

[ad_1]

In this tutorial, we build an end-to-end workflow for working with the SupraLabs reasoning corpus. We stream a representative subset directly from the Hugging Face Hub, inspect its source distribution, token-length patterns, task composition, and reasoning-to-answer ratios, and then apply a series of quality filters to remove unsuitable training examples. We transform the retained samples into a chat-based supervised fine-tuning format with explicit <think> reasoning tags and use them to adapt SmolLM2-135M-Instruct with LoRA through TRL’s SFTTrainer. By combining scalable data access, exploratory analysis, dataset curation, parameter-efficient fine-tuning, structured inference, and Parquet export, we create a complete Google Colab pipeline for turning a large multi-model reasoning corpus into a compact reasoning-focused language model.

import subprocess, sys
def pip_install(pkgs):
   subprocess.check_call([sys.executable, "-m", "pip", "install", "-q", *pkgs])
subprocess.call([sys.executable, "-m", "pip", "uninstall", "-y", "-q", "torchao"])
pip_install([
   "datasets>=3.0.0",
   "transformers>=4.46.0",
   "trl>=0.12.0",
   "peft>=0.13.0",
   "accelerate>=1.0.0",
   "bitsandbytes",
   "matplotlib",
   "pandas",
])
import os, re, json, math, random, itertools, warnings
import pandas as pd
import matplotlib.pyplot as plt
import torch
from collections import Counter
from datasets import load_dataset, Dataset
warnings.filterwarnings("ignore")
random.seed(42)
torch.manual_seed(42)
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
print(f"Device: {DEVICE}")
if DEVICE == "cuda":
   print(f"GPU: {torch.cuda.get_device_name(0)}")
DATASET_ID = "SupraLabs/reasoning-corpus-4K-5M-v1"
SAMPLE_SIZE = 8_000
print(f"\nStreaming {DATASET_ID} ...")
stream = load_dataset(DATASET_ID, split="train", streaming=True)
stream = stream.shuffle(seed=42, buffer_size=30_000)
rows = list(itertools.islice(stream, SAMPLE_SIZE))
ds = Dataset.from_list(rows)
print(f"Materialized sample: {len(ds):,} rows")
print(f"Columns: {ds.column_names}")
ex = ds[0]
print("\n" + "=" * 70)
print("EXAMPLE ROW")
print("=" * 70)
print(f"repo_id : {ex['repo_id']}")
print(f"tok_len : {ex['tok_len']}")
print(f"user            : {ex['user'][:300]} ...")
print(f"thought_trace   : {ex['thought_trace'][:300]} ...")
print(f"assistant       : {ex['assistant'][:300]} ...")

We configure the Colab environment, install the required machine learning libraries, and remove the incompatible torchao package. We detect the available compute device, connect to the SupraLabs reasoning corpus through Hugging Face streaming, and avoid downloading the complete dataset. We shuffle the streamed records, materialize a representative sample, and inspect the structure and contents of an example row.

READ ALSO

XPENG Commissions Humanoid Robot Lines as IRON Walks Off Production – Unite.AI

Matt Clifford Steps Down as ARIA Chair After Anthropic Move – Unite.AI

df = ds.to_pandas()
print("\nTop 15 source repos in sample:")
src_counts = df["repo_id"].value_counts()
print(src_counts.head(15).to_string())
fig, axes = plt.subplots(2, 2, figsize=(14, 10))
axes[0, 0].hist(df["tok_len"], bins=60, color="#4C72B0", edgecolor="white")
axes[0, 0].set_title("Token length distribution")
axes[0, 0].set_xlabel("tok_len"); axes[0, 0].set_ylabel("rows")
src_counts.head(12).plot(kind="barh", ax=axes[0, 1], color="#55A868")
axes[0, 1].invert_yaxis()
axes[0, 1].set_title("Top-12 source repos (sample)")
df["think_chars"] = df["thought_trace"].str.len()
df["answer_chars"] = df["assistant"].str.len()
df["reason_ratio"] = df["think_chars"] / (df["think_chars"] + df["answer_chars"] + 1)
axes[1, 0].hist(df["reason_ratio"], bins=50, color="#C44E52", edgecolor="white")
axes[1, 0].set_title("Reasoning ratio  (think / (think + answer))")
axes[1, 0].set_xlabel("ratio")
axes[1, 1].scatter(df["tok_len"], df["reason_ratio"], s=4, alpha=0.25, color="#8172B2")
axes[1, 1].set_title("tok_len vs reasoning ratio")
axes[1, 1].set_xlabel("tok_len"); axes[1, 1].set_ylabel("ratio")
plt.tight_layout()
plt.show()
print("\nSummary stats:")
print(df[["tok_len", "think_chars", "answer_chars", "reason_ratio"]]
     .describe().round(2).to_string())
def tag_task(row):
   u = row["user"].lower()
   a = row["assistant"]
   if "```" in a or re.search(r"\b(def |class |import |function|#include)", a):
       return "code"
   if re.search(r"(prove|equation|integral|theorem|\\frac|\\int|solve for)", u):
       return "math"
   if re.search(r"\b(patient|diagnosis|symptom|treatment|clinical)\b", u):
       return "medical"
   if re.search(r"\b(which of the following|options?:|\(a\)|\(b\))", u):
       return "mcq/logic"
   return "general"
df["task"] = df.apply(tag_task, axis=1)
print("\nHeuristic task mix:")
print(df["task"].value_counts(normalize=True).round(3).to_string())

We convert the sampled dataset into a pandas DataFrame and analyze the distribution of source repositories and token lengths. We calculate reasoning and answer character counts, measure the reasoning-to-response ratio, and visualize the relationships across the dataset. We also apply lightweight heuristic rules to classify each record as a code, mathematics, medical, multiple-choice, or general task.

def filter_length(row, min_tok=200, max_tok=3000):
   """Keep samples within a training-friendly token budget."""
   return min_tok <= row["tok_len"] <= max_tok
def filter_degenerate(row):
   """Drop empty/near-empty thoughts or answers."""
   return len(row["thought_trace"]) > 100 and len(row["assistant"]) > 20
def filter_repetition(row, max_line_repeat=0.30):
   """Drop traces where one line repeats too often (looping models)."""
   lines = [l.strip() for l in row["thought_trace"].split("\n") if l.strip()]
   if len(lines) < 5:
       return True
   most_common = Counter(lines).most_common(1)[0][1]
   return (most_common / len(lines)) <= max_line_repeat
def filter_reason_ratio(row, lo=0.15, hi=0.97):
   """Keep samples that actually reason but don't ONLY reason."""
   t, a = len(row["thought_trace"]), len(row["assistant"])
   r = t / (t + a + 1)
   return lo <= r <= hi
n0 = len(ds)
ds_f = ds.filter(filter_length)
ds_f = ds_f.filter(filter_degenerate)
ds_f = ds_f.filter(filter_repetition)
ds_f = ds_f.filter(filter_reason_ratio)
print(f"\nFiltering: {n0:,} -> {len(ds_f):,} rows "
     f"({100 * len(ds_f) / n0:.1f}% retained)")
MODEL_ID = "HuggingFaceTB/SmolLM2-135M-Instruct"
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
if tokenizer.pad_token is None:
   tokenizer.pad_token = tokenizer.eos_token
SYSTEM_PROMPT = (
   "You are a careful reasoning assistant. Think step by step inside "
   "<think>...</think> tags, then give your final answer."
)
def to_chat(row):
   return {
       "messages": [
           {"role": "system", "content": SYSTEM_PROMPT},
           {"role": "user", "content": row["user"]},
           {"role": "assistant",
            "content": f"<think>\n{row['thought_trace']}\n</think>\n\n{row['assistant']}"},
       ]
   }
train_ds = ds_f.map(to_chat, remove_columns=ds_f.column_names)
train_ds = train_ds.shuffle(seed=42)
N_TRAIN, N_EVAL = 1_500, 100
eval_ds = train_ds.select(range(N_TRAIN, min(N_TRAIN + N_EVAL, len(train_ds))))
train_ds = train_ds.select(range(min(N_TRAIN, len(train_ds))))
print(f"\nTrain: {len(train_ds):,}  |  Eval: {len(eval_ds):,}")
print("\nRendered training sample (truncated):")
print(tokenizer.apply_chat_template(train_ds[0]["messages"], tokenize=False)[:800])

We construct a quality-filtering pipeline that removes samples with unsuitable token lengths, incomplete responses, excessive repetition, or unbalanced reasoning content. We load the SmolLM2 tokenizer and transform each retained record into a structured conversation containing a system prompt, user message, and reasoning-enhanced assistant response. We then shuffle the formatted data, create training and evaluation subsets, and inspect the final chat template used for supervised fine-tuning.

from trl import SFTTrainer, SFTConfig
from peft import LoraConfig
try:
   import peft.import_utils as _piu
   import peft.tuners.lora.torchao as _plt
   _piu.is_torchao_available = lambda: False
   _plt.is_torchao_available = lambda: False
except Exception:
   pass
model = AutoModelForCausalLM.from_pretrained(
   MODEL_ID,
   dtype=torch.bfloat16 if DEVICE == "cuda" else torch.float32,
).to(DEVICE)
peft_config = LoraConfig(
   r=16,
   lora_alpha=32,
   lora_dropout=0.05,
   bias="none",
   task_type="CAUSAL_LM"
sft_config = SFTConfig(
   output_dir="smollm2-reasoning-demo",
   max_length=2048,
   per_device_train_batch_size=2,
   gradient_accumulation_steps=8,
   num_train_epochs=1,
   learning_rate=2e-4,
   lr_scheduler_type="cosine",
   warmup_steps=10,
   logging_steps=10,
   eval_strategy="steps",
   eval_steps=50,
   save_strategy="no",
   bf16=(DEVICE == "cuda"),
   gradient_checkpointing=True,
   report_to="none",
)
trainer = SFTTrainer(
   model=model,
   args=sft_config,
   train_dataset=train_ds,
   eval_dataset=eval_ds,
   peft_config=peft_config,
   processing_class=tokenizer,
)
print("\nStarting fine-tune (≈10–20 min on a T4 with these settings)...")
trainer.train()
print("Done. Final eval loss:", trainer.evaluate().get("eval_loss"))

We load the SmolLM2 causal language model and configure LoRA adapters for parameter-efficient training. We define the optimization, batching, evaluation, precision, and gradient-checkpointing settings through TRL’s SFTConfig. We initialize the SFTTrainer, fine-tune the model on the curated reasoning conversations, and evaluate its final training performance.

def generate(question, max_new_tokens=512, temperature=0.7):
   msgs = [
       {"role": "system", "content": SYSTEM_PROMPT},
       {"role": "user", "content": question},
   ]
   prompt = tokenizer.apply_chat_template(
       msgs, tokenize=False, add_generation_prompt=True
   )
   inputs = tokenizer(prompt, return_tensors="pt").to(DEVICE)
   with torch.no_grad():
       out = trainer.model.generate(
           **inputs,
           max_new_tokens=max_new_tokens,
           temperature=temperature,
           top_p=0.9,
           do_sample=True,
           pad_token_id=tokenizer.pad_token_id,
       )
   text = tokenizer.decode(out[0][inputs["input_ids"].shape[1]:],
                           skip_special_tokens=True)
   m = re.search(r"<think>(.*?)</think>(.*)", text, re.DOTALL)
   if m:
       print("─" * 60, "\nTHINKING:\n", m.group(1).strip()[:1500])
       print("─" * 60, "\nANSWER:\n", m.group(2).strip())
   else:
       print(text)
print("\n\n### TEST 1: logic puzzle")
generate("If all bloops are razzies and all razzies are lazzies, "
        "are all bloops definitely lazzies? Explain briefly.")
train_ds.to_parquet("reasoning_subset_train.parquet")
eval_ds.to_parquet("reasoning_subset_eval.parquet")
print("\nSaved: reasoning_subset_train.parquet / reasoning_subset_eval.parquet")

We create an inference function that formats new questions with the same system prompt and generates responses from the fine-tuned model. We separate the generated <think> section from the final answer and test the model on logic and arithmetic problems. We finally export the processed training and evaluation datasets as Parquet files for reuse in larger experiments.

In conclusion, we developed a practical pipeline that connects large-scale reasoning-data exploration with small-language-model training. We streamed the corpus efficiently, analyzed its internal composition, filtered examples using token, repetition, completeness, and reasoning-balance criteria, and converted the resulting data into a consistent conversational training structure. We then fine-tuned SmolLM2 with LoRA, evaluated the adapted model, inspected its generated reasoning and answers, and exported the curated datasets for future experiments. This workflow provides a reusable foundation for source-aware data mixing, curriculum learning, larger student models, longer-context training, and production-scale reasoning model development without requiring the entire dataset to reside in Colab memory.


Check out the FULL CODES here. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us


Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.

[ad_2]

Source_link

Related Posts

XPENG Commissions Humanoid Robot Lines as IRON Walks Off Production – Unite.AI
Al, Analytics and Automation

XPENG Commissions Humanoid Robot Lines as IRON Walks Off Production – Unite.AI

September 8, 2026
Matt Clifford Steps Down as ARIA Chair After Anthropic Move – Unite.AI
Al, Analytics and Automation

Matt Clifford Steps Down as ARIA Chair After Anthropic Move – Unite.AI

September 7, 2026
In “An Alien Mind,” OpenAI’s Jakub Pachocki Urges Shared Safety Bars – Unite.AI
Al, Analytics and Automation

In “An Alien Mind,” OpenAI’s Jakub Pachocki Urges Shared Safety Bars – Unite.AI

September 7, 2026
H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder
Al, Analytics and Automation

H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder

September 7, 2026
The Model, Tools, Memory, and Control Loop – Unite.AI
Al, Analytics and Automation

The Model, Tools, Memory, and Control Loop – Unite.AI

September 6, 2026
UC Berkeley Researchers Release CUA-Lite, an Open Platform Unifying Sandboxes, Data, Evaluation and RL for Computer-Use Agents
Al, Analytics and Automation

UC Berkeley Researchers Release CUA-Lite, an Open Platform Unifying Sandboxes, Data, Evaluation and RL for Computer-Use Agents

September 6, 2026
Next Post
What Features Set Them Apart

What Features Set Them Apart

POPULAR NEWS

Trump ends trade talks with Canada over a digital services tax

Trump ends trade talks with Canada over a digital services tax

June 28, 2025
15 Trending Songs on TikTok in 2025 (+ How to Use Them)

15 Trending Songs on TikTok in 2025 (+ How to Use Them)

June 18, 2025
Communication Effectiveness Skills For Business Leaders

Communication Effectiveness Skills For Business Leaders

June 10, 2025
Comparing the Top 7 Large Language Models LLMs/Systems for Coding in 2025

Comparing the Top 7 Large Language Models LLMs/Systems for Coding in 2025

November 4, 2025
App Development Cost in Singapore: Pricing Breakdown & Insights

App Development Cost in Singapore: Pricing Breakdown & Insights

June 22, 2025

EDITOR'S PICK

Major League Soccer on its Experiential Strategy, World Cup Plans

Major League Soccer on its Experiential Strategy, World Cup Plans

March 11, 2026
OpenAI’s Unreleased AGI Paper Could Complicate Microsoft Negotiations

OpenAI’s Unreleased AGI Paper Could Complicate Microsoft Negotiations

June 27, 2025
Email Campaigns and Loyalty Programs: The Ultimate Power Couple

Email Campaigns and Loyalty Programs: The Ultimate Power Couple

September 4, 2026
The 6 Best Latte Machines for Automatic Espresso Drinks (2025)

The 6 Best Latte Machines for Automatic Espresso Drinks (2025)

June 14, 2025

About

We bring you the best Premium WordPress Themes that perfect for news, magazine, personal blog, etc. Check our landing page for details.

Follow us

Categories

  • Account Based Marketing
  • Ad Management
  • Al, Analytics and Automation
  • Brand Management
  • Channel Marketing
  • Digital Marketing
  • Direct Marketing
  • Event Management
  • Google Marketing
  • Marketing Attribution and Consulting
  • Marketing Automation
  • Mobile Marketing
  • PR Solutions
  • Social Media Management
  • Technology And Software
  • Uncategorized

Recent Posts

  • The AI visibility gap: Why great brands disappear from AI answers
  • XPENG Commissions Humanoid Robot Lines as IRON Walks Off Production – Unite.AI
  • Google partners on West Virginia energy storage project
  • Gemini Live and Find Hub
  • About Us
  • Disclaimer
  • Contact Us
  • Privacy Policy
No Result
View All Result
  • Technology And Software
    • Account Based Marketing
    • Channel Marketing
    • Marketing Automation
      • Al, Analytics and Automation
      • Ad Management
  • Digital Marketing
    • Social Media Management
    • Google Marketing
  • Direct Marketing
    • Brand Management
    • Marketing Attribution and Consulting
  • Mobile Marketing
  • Event Management
  • PR Solutions