• About Us
  • Disclaimer
  • Contact Us
  • Privacy Policy
Saturday, September 5, 2026
mGrowTech
No Result
View All Result
  • Technology And Software
    • Account Based Marketing
    • Channel Marketing
    • Marketing Automation
      • Al, Analytics and Automation
      • Ad Management
  • Digital Marketing
    • Social Media Management
    • Google Marketing
  • Direct Marketing
    • Brand Management
    • Marketing Attribution and Consulting
  • Mobile Marketing
  • Event Management
  • PR Solutions
  • Technology And Software
    • Account Based Marketing
    • Channel Marketing
    • Marketing Automation
      • Al, Analytics and Automation
      • Ad Management
  • Digital Marketing
    • Social Media Management
    • Google Marketing
  • Direct Marketing
    • Brand Management
    • Marketing Attribution and Consulting
  • Mobile Marketing
  • Event Management
  • PR Solutions
No Result
View All Result
mGrowTech
No Result
View All Result
Home Al, Analytics and Automation

Poolside Releases Laguna S 2.1, an Open-Weight Agentic Coding Model Punching Above Its Weight Class on SWE-Bench Multilingual

Josh by Josh
July 22, 2026
in Al, Analytics and Automation
0
Poolside Releases Laguna S 2.1, an Open-Weight Agentic Coding Model Punching Above Its Weight Class on SWE-Bench Multilingual


Poolside has released Laguna S 2.1, a 118B-parameter open-weight model built for agentic coding. It is a Mixture-of-Experts (MoE) model with 8B activated parameters per token. It supports a context window of up to 1M tokens in both thinking and no-thinking modes. The weights are on Hugging Face under an OpenMDW-1.1 license, and the model is small enough to run on a single NVIDIA DGX Spark.

On long-horizon coding benchmarks, Laguna S 2.1 holds its own against models several times its size, including DeepSeek-V4-Pro-Max, NVIDIA’s Nemotron 3 Ultra, and Thinking Machines’ Inkling. Laguna S 2.1 is a scale-up of the Laguna XS family, trained on the same pre-training data as XS 2.1.

READ ALSO

Adaption Labs Introduces ‘Invent a Dataset’: Training Data Generated From a Task Description, Not a Seed Corpus

Retrieval vs. Memory in Agentic AI System

What is Laguna S 2.1

The model activates roughly 6.8% of its parameters on any given token. All 118B parameters remain resident in memory, but only ~8B route through the network per step. That sparsity is why a mid-size model can behave like a larger one while staying cheap to serve.

Poolside team publishes weights in BF16, FP8, INT4, and NVFP4, along with official GGUF and MLX conversions and DFlash draft models. It went from the start of training to launch in under nine weeks. Pre-training began on 22 May 2026 on 4,096 NVIDIA H200 GPUs. It is the first Poolside model where reinforcement learning ran in FP8 precision.

Performance

Laguna S 2.1 scores 70.2% on Terminal-Bench 2.1 with thinking enabled. That places it first among open, disclosed-size models on Poolside’s compiled leaderboard, behind only larger or closed systems. On SWE-Bench Multilingual it scores 78.5%, topping the published table outright. The full comparison Poolside released is below.

Benchmark Laguna S 2.1 (118B-A8B) Tencent Hy3 (295B-A21B) Inkling (975B-A41B) Nemotron 3 Ultra (550B-A55B) DeepSeek-V4-Pro-Max (1.6T-A49B) Kimi K3 (2.8T-A50B) Qwen 3.7 Max Muse Spark 1.1 Claude Fable 5
Terminal-Bench 2.1 70.2 71.7 63.8 56.4 64.0 88.3 74.5 80.0 88.0
SWE-Bench Multilingual 78.5 75.8 – 67.7 76.2 – 78.3 – –
SWE-Bench Pro (Public) 59.4 57.9 54.3 – 55.4 – 60.6 61.5 80.3
DeepSWE v1.1 40.4 – – – 9.0 69.0 – 53.3 70.0
SWE Atlas (Codebase QnA) 46.2 – – – 27.2 – – 42.2 –
Toolathlon Verified 49.7 – 45.5 34.3 55.9 – – 75.6 –

The clearest signal is DeepSWE v1.1, which still has real headroom. There, Laguna S 2.1 scores 40.4% against DeepSeek-V4-Pro-Max’s 9.0%, with roughly one-sixth the active parameters. Closed frontier models such as Claude Fable 5 and Kimi K3 still lead on several benchmarks. Poolside’s claim is about the weight class, not the outright top. Trajectories from the final evaluation set is published at trajectories.poolside.ai.

Two thinking modes, and where the score comes from

Laguna S 2.1 has two modes: off and max, with max enabled by default. In max mode the model sets its own test-time compute budget. Poolside is shipping without user-configurable low/medium/high effort control for now.

Max thinking lifts Terminal-Bench 2.1 from 60.4% to 70.2%. It lifts DeepSWE from 16.5% to 40.4%. Those gains cost tokens: DeepSWE trajectories run about 249k completion tokens in thinking mode against 99k without. Poolside team reports coherent, productive reasoning running for hours and hundreds of thousands of tokens.

Three published runs

Poolside team shared three unedited trajectories to show behavior rather than scores. In one, the model built a working HTML/CSS browser engine from an empty folder. That run took 181 steps across a 50-minute session, with no human intervention, and it validated its own output against headless Chromium.

In a second, the model optimized Poolside’s own agent harness. It made the harness 5.2% faster with roughly 71% lower memory allocation, replacing O(n²) string concatenation with buffers. In a third, it re-derived Erdős Problem #397 offline in Perl over 68 minutes, since the sandbox had no Python. That result is an independent rediscovery; GPT-5.2 Pro solved the same problem in January 2026, and Laguna’s knowledge cutoff is November 2025.

Can you actually deploy it?

Sizing uses the full 118B parameters, not the 8B active count, because every expert stays in memory. At 4-bit (NVFP4 or INT4) the weights need about 59 GB, which fits comfortably on a single DGX Spark’s 128 GB of unified memory. At FP8 they need about 118 GB, still within a single Spark or a single H200. At BF16 they need about 236 GB, which calls for two linked Sparks or a multi-GPU datacenter node.

Poolside worked with NVIDIA to optimize inference from TRT-LLM serving to NVFP4 on Blackwell, down to a single DGX Spark. It shipped day-one support for vLLM, SGLang, and Ollama. Hosted access runs through OpenRouter, free at 256K context and paid at the full 1M context for $0.10 / $0.20 / $0.01 per 1M input / output / cache-read tokens. The model is also on Baseten, Kilo, Prime Intellect Prime Lab, and ZML.

Key Takeaways

  • Laguna S 2.1 is a 118B-total / 8B-active MoE coding model with a 1M-token context, open under OpenMDW-1.1.
  • It scores 70.2% on Terminal-Bench 2.1 and 78.5% on SWE-Bench Multilingual, leading open disclosed-size models.
  • At 4-bit it runs on a single NVIDIA DGX Spark; FP8 fits one Spark or H200, BF16 needs two.
  • Default ‘max thinking’ drives most of the score, at a real token cost (DeepSWE: 16.5% → 40.4%).
  • Trained in under nine weeks on 4,096 H200 GPUs, it is Poolside’s third model in under three months.

Sources: Poolside Technical— Introducing Laguna S 2.1, Robert McHardy on X, Poolside on X and Hugging Face model card.


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.



Source_link

Related Posts

Adaption Labs Introduces ‘Invent a Dataset’: Training Data Generated From a Task Description, Not a Seed Corpus
Al, Analytics and Automation

Adaption Labs Introduces ‘Invent a Dataset’: Training Data Generated From a Task Description, Not a Seed Corpus

September 5, 2026
Al, Analytics and Automation

Retrieval vs. Memory in Agentic AI System

September 5, 2026
OpenAI Commits $1B to Frontline Cyber Defense, Launches MS-ISAC Pilot – Unite.AI
Al, Analytics and Automation

OpenAI Commits $1B to Frontline Cyber Defense, Launches MS-ISAC Pilot – Unite.AI

September 5, 2026
MIT Quantum Initiative launches postdoctoral fellowship program | MIT News
Al, Analytics and Automation

MIT Quantum Initiative launches postdoctoral fellowship program | MIT News

September 4, 2026
Google DeepMind’s WeatherNext 3 Trains on Weather Station Observations to Deliver 5 km Global Forecasts, Refreshed Every Hour
Al, Analytics and Automation

Google DeepMind’s WeatherNext 3 Trains on Weather Station Observations to Deliver 5 km Global Forecasts, Refreshed Every Hour

September 4, 2026
Al, Analytics and Automation

Understanding the Role of Latent Space in Machine Learning Models

September 4, 2026
Next Post
France Will Ban Social Media For Children Under 15

France Will Ban Social Media For Children Under 15

POPULAR NEWS

Trump ends trade talks with Canada over a digital services tax

Trump ends trade talks with Canada over a digital services tax

June 28, 2025
15 Trending Songs on TikTok in 2025 (+ How to Use Them)

15 Trending Songs on TikTok in 2025 (+ How to Use Them)

June 18, 2025
Communication Effectiveness Skills For Business Leaders

Communication Effectiveness Skills For Business Leaders

June 10, 2025
Comparing the Top 7 Large Language Models LLMs/Systems for Coding in 2025

Comparing the Top 7 Large Language Models LLMs/Systems for Coding in 2025

November 4, 2025
App Development Cost in Singapore: Pricing Breakdown & Insights

App Development Cost in Singapore: Pricing Breakdown & Insights

June 22, 2025

EDITOR'S PICK

DeepSeek AI Releases DeepSeek Harness in Developer Preview: An MIT-Licensed Agent Harness Where Everything is a Plugin

DeepSeek AI Releases DeepSeek Harness in Developer Preview: An MIT-Licensed Agent Harness Where Everything is a Plugin

August 17, 2026
What You Need to Know

What You Need to Know

October 7, 2025
Define effective digital marketing objectivs and KPIs to achieve your goals in 2025

Define effective digital marketing objectivs and KPIs to achieve your goals in 2025

June 6, 2025
Our Top 5 Blog Posts of 2025 (And What Made Them Work)

Our Top 5 Blog Posts of 2025 (And What Made Them Work)

January 19, 2026

About

We bring you the best Premium WordPress Themes that perfect for news, magazine, personal blog, etc. Check our landing page for details.

Follow us

Categories

  • Account Based Marketing
  • Ad Management
  • Al, Analytics and Automation
  • Brand Management
  • Channel Marketing
  • Digital Marketing
  • Direct Marketing
  • Event Management
  • Google Marketing
  • Marketing Attribution and Consulting
  • Marketing Automation
  • Mobile Marketing
  • PR Solutions
  • Social Media Management
  • Technology And Software
  • Uncategorized

Recent Posts

  • The Wine AI Visibility Index 2026: Named vs. Cited
  • GeoGuessr Daily Challenge Answer Today for September 5, 2026
  • Adaption Labs Introduces ‘Invent a Dataset’: Training Data Generated From a Task Description, Not a Seed Corpus
  • 7 Best Virtual Desktop Infrastructure (VDI) Software (2026): My Picks
  • About Us
  • Disclaimer
  • Contact Us
  • Privacy Policy
No Result
View All Result
  • Technology And Software
    • Account Based Marketing
    • Channel Marketing
    • Marketing Automation
      • Al, Analytics and Automation
      • Ad Management
  • Digital Marketing
    • Social Media Management
    • Google Marketing
  • Direct Marketing
    • Brand Management
    • Marketing Attribution and Consulting
  • Mobile Marketing
  • Event Management
  • PR Solutions