Inkling’s open-weight AI model shakes up global competition

The gist
Thinking Machines Lab’s new 975-billion-parameter open-weight AI, Inkling, is shaking up the global AI race with radical transparency, developer freedom, and benchmark-beating performance.
What to know
- Inkling fuses text, images, audio, and video in a massive Mixture-of-Experts architecture, offering a one-million-token context window and dynamic compute scaling.
- Released under Apache 2.0 and hosted on Hugging Face, Inkling is fully customizable and deployable on proprietary infrastructure—no closed APIs required.
- Already outperforming Nvidia’s proprietary AI on agentic tasks, Inkling is reigniting U.S. competitiveness and challenging China’s dominance in the 2026 AI arms race.
Dynamic MoE Powers Efficiency
Inkling’s 975B-parameter Mixture-of-Experts design activates just 4.2% of its parameters per token, letting developers fine-tune computational cost and reasoning depth on demand.
Inkling’s architecture is anchored by a massive 975-billion-parameter sparse Mixture-of-Experts (MoE) design that activates only about 4.2% (41 billion) of its parameters per token, striking a balance between scale and efficiency. Each of its 66 decoder-only Transformer layers contains 256 routed experts plus two shared experts, with a novel sigmoid-based routing mechanism that employs a learned bias for load balancing, eschewing auxiliary-loss training instability common in other MoEs. This dynamic expert selection enables scalable inference and controllable reasoning effort, allowing developers to dial computational cost and latency per request—a feature refined through over 30 million asynchronous rollouts to optimize token generation versus compute trade-offs.
Inkling pushes the boundaries of long-context modeling with an unprecedented one-million-token open-weight context window and API contexts supporting up to 256K tokens. This capability is powered by a hybrid attention architecture that interleaves five sliding-window (local) attention layers for every global attention layer, combined with learned relative positional encodings that outperform traditional Rotary Positional Embeddings in extrapolating to longer sequences. These innovations ensure efficient local computation while maintaining global information flow, crucial for reasoning over extended text and multimodal inputs.
Unlike many multimodal models that rely on separate encoders, Inkling achieves genuine token-level multimodal fusion by projecting small 40x40 pixel patches directly into tokens, enabling native reasoning across text, images, audio, and video within a unified Transformer framework. Pretrained on a staggering 45 trillion tokens spanning these modalities, Inkling outputs text but integrates diverse input types seamlessly, broadening its versatility in applications ranging from vision-language tasks to audio understanding without the overhead of distinct modality-specific encoders.
Inkling’s architectural innovations extend beyond scale and modality to include efficiency and training optimizations that enhance both speed and robustness. It incorporates short convolution layers within Transformer blocks to explicitly mix local sequence information beyond attention mechanisms, employs DeepSeek-style auxiliary-loss-free load balancing for expert utilization, and uses parameter-specific optimizers—Muon for large matrices and Adam variants with adaptive weight decay schedules—to improve training stability. This design philosophy, influenced by deep SEQ models and a 'fast path' approach, results in a relatively fast, open-weight model from Thinking Machines Lab that prioritizes architectural novelty and practical accessibility over benchmark supremacy.
True Open-Weight Customization
Inkling’s Apache 2.0 licensing and modular build empower organizations to fully own, fine-tune, and deploy the model on private infrastructure—sidestepping closed APIs and vendor lock-in.
Thinking Machines Lab’s Inkling model, a mammoth 975-billion-parameter open-weight AI, breaks new ground by offering developers unprecedented customization through fine-tuning and modular adjustments. Released under the permissive Apache 2.0 license and hosted on Hugging Face, Inkling empowers organizations to deploy the full model on their own infrastructure using proprietary data, sidestepping reliance on closed APIs and preserving domain expertise. This approach aligns with the lab’s business model, which monetizes customization services via its Tinker platform rather than restricting access to the model itself, fostering a decentralized AI ecosystem.
Inkling’s technical innovations extend beyond open access to include a massive one-million-token context window and a unique 'reasoning effort' dial that lets developers dynamically balance inference depth, latency, and cost per request. This controllable compute feature allows users to tailor the model’s reasoning intensity from rapid, cost-efficient responses to deep, high-quality analysis, making Inkling adaptable to diverse workloads and use cases. Such flexibility enhances developer control and operational efficiency, a critical advantage in real-world applications.
By making Inkling’s full weights publicly available on platforms like Hugging Face, Thinking Machines Lab catalyzes a shift toward decentralized AI development, inviting a broad community of developers to innovate without gatekeepers. This open-weight accessibility not only reduces costs—Bridgewater reportedly outperformed proprietary models at roughly one-fourteenth the expense—but also mitigates data-sharing risks by enabling businesses to run customized AI on their own data and infrastructure. This democratization challenges the dominant closed-model paradigm and signals a new era of AI sovereignty.
U.S. Reclaims AI Innovation
Inkling’s open-weight release signals a strategic pivot in the global AI race, emphasizing reproducibility and user agency as the U.S. counters China’s lead with transparent, customizable architectures.
Thinking Machines Lab’s Inkling model has emerged as a pivotal force in the 2026 AI arms race by delivering a 975-billion-parameter open-weight multimodal architecture under the Apache 2.0 license, reigniting global competition beyond the dominant Chinese models and hyperscalers. Its modular, tunable design emphasizes cost-efficient, transparent, and adaptable AI solutions, positioning the U.S. as a leader in open-source AI innovation amid escalating geopolitical tensions and soaring hyperscaler expenditures. This strategic push not only challenges proprietary giants but also accelerates enterprise-grade AI customization and agentic workflows, marking a significant milestone in the race for scalable, efficient AI architectures.
While Inkling leads the U.S. open-source AI charge alongside contenders like Neotron and Reflection, it currently ranks tenth globally, trailing nine Chinese models, underscoring the fierce international competition and the imperative for continued development. Despite internal restructuring at Thinking Machines Lab, the rapid six-month development cycle of Inkling highlights resilience within the U.S. AI sector. Moreover, having a domestically developed open-source model is critical for American enterprises aiming to mitigate geopolitical risks associated with reliance on foreign AI technologies, as one expert emphasized, “If we don't have an American open source model, businesses will simply be incurring greater risk.”
Inkling exemplifies a strategic shift in AI innovation from merely scaling model size to empowering users with customizable, fine-tunable intelligence, reflecting a broader U.S. philosophy in open-source AI development. Its reproducible and deterministic training approach provides a controlled experimentation environment that could enable rapid evolution into a frontier model, signaling a methodical and sustainable U.S.-led path forward. This challenges the prevailing narrative that open-source AI breakthroughs are primarily driven by Chinese labs, showcasing American strength in transparent, adaptable AI architectures that prioritize user agency and customization.
Agentic Performance, Real-World Edge
Outperforming Nvidia’s proprietary models on agentic tasks, Inkling delivers enterprise-ready multimodal AI that’s accessible, hardware-flexible, and tailored for sectors demanding domestic control.
Thinking Machines Lab’s Inkling, a 975-billion-parameter open-weight model, has made a striking impression by outperforming Nvidia’s proprietary AI on agentic tasks and securing a strong position on the Design Arena benchmark, where it ranks between GPT-5.6 Soul and Cloud Opus 4.6. This performance underscores Inkling’s robust multimodal design capabilities, particularly in creating high-quality, human-judged websites, even as it shows mixed results on other software benchmarks, highlighting its specialized strengths in creative and agentic domains.
Despite having fewer active parameters—41 billion compared to the much larger Quinn 3.8 and Quinn 3 models—Inkling delivers competitive results, demonstrating that sheer size isn’t the sole determinant of AI efficacy. Developed from scratch by a US-based company, Inkling offers a compelling alternative to proprietary giants like GPT-5, blending efficiency with high performance in a model that is less than half the size of some competitors yet still holds its own in key benchmarks.
Inkling’s open-weight release marks a significant shift in AI accessibility and enterprise adoption, particularly in North America. By allowing users to deploy the model on their own or rented hardware—requiring about 1.9 terabytes of disk space and compatible with high-end machines like Apple’s Mac Studio—it circumvents the restrictive usage policies often associated with Chinese AI models. This flexibility, combined with its US origin, makes Inkling especially attractive to sectors such as banking that demand domestic, cost-effective, and adaptable AI solutions.






