Laguna S 2.1: Open-Weight Agentic Coding Model Beats Larger Models

Poolside's Laguna S 2.1 is an open-weight MoE model (118B/8B active) scoring 78.5% on SWE-Bench Multilingual. It runs on a single DGX Spark and redefines

viernes, 24 de julio de 2026 • 3 min read • Q2BSTUDIO Team

Pequeño en parámetros activos, grande en rendimiento: Laguna S 2.1

The landscape of artificial intelligence applied to software development has taken a significant leap with the arrival of Laguna S 2.1, an open-weight model released by Poolside. With 118 billion total parameters (118B) but only 8 billion activated per token (8B), this Mixture-of-Experts (MoE) model achieves performance rivalling much larger systems, reaching 78.5% on SWE-Bench Multilingual and 70.2% on Terminal-Bench 2.1. Its ability to handle up to 1 million tokens of context, along with 'max' thinking modes that boost results, makes it an exceptional tool for agentic coding tasks.

From a technical perspective, Laguna S 2.1 strikes a remarkable balance between efficiency and power. By activating only 6.8% of its parameters at each step, it behaves like models with hundreds of billions of active parameters, but at a much lower inference cost. This is achieved through the MoE architecture, which keeps all experts in memory but routes only a fraction per token. The model was trained in under nine weeks using 4,096 NVIDIA H200 GPUs, starting on May 22, 2026. Poolside has published weights in BF16, FP8, INT4, and NVFP4 formats, along with official conversions for GGUF, MLX, and DFlash draft models, easing deployment in diverse environments.

In benchmark evaluations, Laguna S 2.1 leads among open models with disclosed size. On SWE-Bench Multilingual it surpasses competitors like Tencent Hy3 (75.8%) and DeepSeek-V4-Pro-Max (76.2%), while on DeepSWE v1.1 it achieves 40.4% versus DeepSeek's 9.0%, despite having one-sixth the active parameters. Although closed models like Claude Fable 5 or Kimi K3 still lead on some tests, Laguna S 2.1's achievement is impressive for its weight class. The 'max' thinking mode (enabled by default) is crucial: it lifts Terminal-Bench from 60.4% to 70.2% and DeepSWE from 16.5% to 40.4%, though at a token cost (e.g., 249k completion tokens versus 99k without thinking).

For companies involved in custom software development, the arrival of open models like Laguna S 2.1 opens immense opportunities. Integrating AI agents capable of writing, debugging, and optimizing code autonomously can accelerate development cycles, reduce costs, and free human teams for higher-value tasks. In this context, companies like Q2BSTUDIO, specialized in custom multiplatform software development, can leverage these capabilities to deliver smarter solutions to their clients. The combination of open-source models with expertise in artificial intelligence, cybersecurity, cloud AWS/Azure, and Business Intelligence with Power BI enables the creation of AI-assisted development platforms that boost productivity and code quality.

Deploying Laguna S 2.1 is feasible even on modest hardware. With 4-bit quantization (NVFP4 or INT4), the weights occupy about 59 GB, fitting comfortably in a single NVIDIA DGX Spark with 128 GB of unified memory. In FP8, it requires 118 GB, still within a Spark or an H200; in BF16 it needs 236 GB, requiring two Sparks or a multi-GPU node. Poolside has optimized inference with TRT-LLM, vLLM, SGLang, and Ollama, and hosted access is available through OpenRouter, Baseten, and other platforms, with prices starting at $0.10 per million input tokens. This flexibility allows developers and companies to integrate the model into their workflows without relying on massive infrastructure.

Beyond the numbers, Poolside has shared real trajectories that demonstrate the model's ability to autonomously solve complex problems: from building an HTML/CSS browser engine from scratch (181 steps in 50 minutes), to optimizing its own agent harness reducing memory allocation by 71%, to independently rediscovering an Erdős mathematical problem. These examples illustrate how an open and efficient model can become a strategic ally for innovation in software engineering.

In summary, Laguna S 2.1 marks a milestone in democratizing AI for agentic coding. Its performance on SWE-Bench Multilingual (78.5%) and its ability to run on a single device make it an attractive choice for startups and companies seeking competitive advantages without investing in massive clusters. For Q2BSTUDIO, such advances reinforce the importance of staying at the technological forefront, integrating cutting-edge models into custom software development, artificial intelligence, cybersecurity, and process automation services. The future of AI-assisted programming is already here, and models like Laguna S 2.1 are the spearhead of a new era of productivity and creativity in software development.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.