The landscape of open-source artificial intelligence models has taken a significant turn with the arrival of Inkling, the largest model ever created by an American company outside Asian giants. Presented by Thinking Machines Lab, the startup founded by former OpenAI CTO Mira Murati, this 975-billion-parameter model promises to redefine what businesses can achieve with open AI. Under an Apache 2.0 license, Inkling is not only available for download and use but also allows deep customization, making it a key tool for developers seeking to build custom software with advanced reasoning capabilities.
The model employs a mixture-of-experts (MoE) architecture with 256 routed experts and two shared ones, activating about 41 billion parameters per generated token. This allows Inkling to offer competitive performance on modern hardware like NVIDIA B300 or H200 GPUs despite its massive size. Thinking Machines trained the model from scratch using 45 trillion tokens of text, images, audio, and video, running on NVIDIA GB300 NVL72 systems. The company claims Inkling matches or exceeds models like DeepSeek V4 or Nemotron 3 Ultra on benchmarks such as Terminal Bench 2.1, while using thinking tokens more efficiently, reducing operational costs for businesses.
One of Inkling's standout features is its ability to handle contexts of up to one million tokens, equivalent to exceptional short-term memory. This makes it ideal for complex tasks like managing large codebases, needle-in-a-haystack searches, or analyzing extensive documents. Additionally, as a reasoning model trained with reinforcement learning, it can 'think' before responding, reducing hallucinations and improving accuracy in logic-driven tasks. For companies looking to integrate these capabilities into their workflows, having technology partners like Q2BSTUDIO can make a difference: their expertise in artificial intelligence, cybersecurity, and AWS/Azure cloud allows organizations to adapt models like Inkling to their specific needs, whether for virtual assistants, process automation, or predictive analytics.
Inkling's flexibility extends to its tool ecosystem. Through Thinking Machines' Tinker platform, developers can fine-tune the model with scripts generated by the AI itself, accelerating customization. This is particularly useful for companies requiring process automation or AI-powered Business Intelligence solutions. For instance, a BI team using Power BI could integrate Inkling to generate natural language queries and uncover hidden patterns in data, while the cybersecurity department could use it to analyze logs and detect threats in real time. The model's ability to write its own fine-tuning scripts makes it an attractive option for developers aiming to reduce deployment time.
Inkling competes not only in size but also in efficiency. Thinking Machines claims that for reasoning tasks, the model uses about a third of the tokens that Nemotron 3 Ultra would need to achieve equivalent results. This translates to lower API costs for end users, a critical factor in enterprise projects where inference budgets can skyrocket. The company has also released an NVFP4 quantized version that halves memory requirements, allowing the model to run on more modest setups without sacrificing too much quality. For businesses that haven't fully migrated to the cloud, Q2BSTUDIO offers AWS/Azure cloud services that facilitate deploying large-scale models while ensuring scalability and security.
Alongside Inkling, Thinking Machines has previewed Inkling-Small, a 276-billion-parameter model with 12 billion active parameters, designed for latency-sensitive applications. Although not yet publicly available, its existence indicates a clear strategy: to cover everything from intensive workloads to real-time responses. This approach echoes the trend of creating specialized AI agents, a field where customization is key. Companies wishing to explore this path can benefit from Q2BSTUDIO's expertise in developing custom software, integrating AI models with existing systems to create intelligent workflows.
The arrival of Inkling marks a milestone in the democratization of advanced AI. By offering a top-tier model under a permissive license, Thinking Machines Lab opens the door for startups, SMEs, and large corporations to compete in the AI arena without depending on closed providers. However, the real advantage lies not only in the model but in how it is implemented. This is where services like those from Q2BSTUDIO become essential: from AI consulting to integration with BI/Power BI tools or adopting cybersecurity measures, the right expertise can turn a generic model into a unique business solution. With Inkling, the future of open AI looks brighter than ever, and companies that act quickly will be better positioned to leverage its capabilities.





