The artificial intelligence ecosystem has gained a new member that promises to change the rules of the game in multimodal model development. Thinking Machines Lab has introduced Inkling, a state-of-the-art language model with 975 billion total parameters and 41 billion active, based on a Mixture-of-Experts (MoE) architecture. But beyond the numbers, what truly sets it apart is its ability to control 'reasoning effort' during inference — a feature that allows adjusting computational cost and latency per request. In this article, we dive into the architecture, business applications, and how Q2BSTUDIO, as a software and technology development company, can help organizations get the most out of this model and others like it.
Inkling is not just another large model. Its MoE design with 256 routed experts and 2 shared experts per layer, along with an attention mechanism that interleaves sliding windows and global layers at a 5:1 ratio, allows it to handle contexts of up to 1 million tokens. The ability to simultaneously process text, images, and audio (though outputting only UTF-8 text) makes it an ideal tool for complex multimodal applications. But perhaps the most innovative aspect is the reasoning effort control: a parameter ranging from 0.2 to 0.99 that lets you decide how many tokens to dedicate to each response. This translates into direct cost optimization, critical for companies that need to scale AI solutions without skyrocketing bills.
From a technical perspective, Inkling was trained on 45 trillion tokens of multimodal data and employed techniques like Muon for large matrices and Adam for the rest of the parameters. The result is a model that, at maximum effort (0.99), achieves competitive scores on benchmarks such as AIME 2026 (97.1%), GPQA Diamond (87.2%), or FORTRESS Adversarial (78.0%), surpassing other open-weights models in adversarial robustness. However, it is not perfect: in tasks like SimpleQA Verified (43.9%) or Terminal Bench 2.1 (63.8%) it lags behind competitors like DeepSeek V4 Pro or GLM 5.2. Still, its efficiency is remarkable: it spends one third of the tokens that Nemotron 3 Ultra uses for equal performance on Terminal Bench.
For businesses, this level of control and efficiency opens the door to very specific use cases. For example, in developing custom software applications that integrate voice and vision — such as support assistants that process calls and screenshots to generate structured tickets — the ability to adjust effort allows using low-cost mode for routine tasks and maximum mode only when deep analysis is needed. This fits perfectly with Q2BSTUDIO's philosophy, which offers personalized AI solutions for clients in sectors like finance, logistics, or healthcare. The possibility of fine-tuning the model with proprietary data using platforms like Tinker (64K or 256K context) or through APIs on TogetherAI, Fireworks, or Databricks facilitates creating AI agent systems tailored to specific needs.
Moreover, cloud integration is a key factor. Inkling can be deployed in cloud environments like AWS or Azure, and Q2BSTUDIO has experience with AWS/Azure cloud services to help companies orchestrate scalable infrastructures. Inference with Inkling requires at least 2 TB of aggregated VRAM in BF16 (e.g., 8x NVIDIA B300 or 16x H200), but with NVFP4 quantization (W4A4) it drops to 600 GB, allowing execution on 4x B300. This makes the model viable for companies that already have investments in cutting-edge hardware, but also opens the door to lighter solutions like Inkling-Small (276B parameters, 12B active) which, though not yet available, promises to match or exceed the large model on many benchmarks.
Another area where Inkling can make a difference is cybersecurity. Its score of 78.0% on FORTRESS Adversarial indicates superior robustness against malicious inputs. This makes it an ideal candidate for anomaly detection systems or chatbots that must resist prompt injection attacks. Q2BSTUDIO, with its cybersecurity division, can integrate models like Inkling into offensive and defensive security platforms, helping companies protect their digital assets through intelligent agents capable of identifying attack patterns and suggesting countermeasures in real time.
We cannot overlook the value of business intelligence. Inkling can process multimodal data — such as financial reports, charts, and meeting recordings — and generate summaries or extract key metrics. This aligns with the BI / Power BI solutions offered by Q2BSTUDIO, where generative AI can automate dashboard generation or interpret natural language queries. The model's ability to handle long contexts (1M tokens) allows analyzing extensive documents without losing coherence, ideal for audits or compliance.
Finally, the combination of Inkling with enterprise automation strategies is promising. For example, in customer service pipelines, a low-effort agent can classify incidents, while a high-effort one resolves the more complex ones. All managed from the same infrastructure. The availability of a multi-token prediction drafter (MTP) for speculative decoding further speeds up inference, reducing latency in real-time applications.
In summary, Inkling represents a significant advance in democratizing multimodal AI: it is open (Apache 2.0 license), customizable, and offers granular control of computational cost. However, its effective adoption requires deep knowledge of the architecture and best deployment practices. This is where companies like Q2BSTUDIO, specialized in custom software applications, AI, and cloud, can bring their expertise so that organizations not only integrate the model but optimize it for their real workflows. From fine-tuning with proprietary data to integration with BI and cybersecurity systems, the potential is enormous. What is clear is that reasoning effort control marks a before and after: AI is no longer a fixed resource, but a service that adapts to the budget and urgency of each task.





