Introducing Inkling: The Future Of AI Innovation

📊 Full opportunity report: Introducing Inkling: The Future Of AI Innovation on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Thinking Machines has launched Inkling, a massive 975-billion-parameter multimodal AI model, available on Hugging Face. Its open access offers new possibilities for cross-modal reasoning, but hardware requirements and evaluation details remain uncertain.

Thinking Machines has released Inkling on Hugging Face, a 975-billion-parameter multimodal model designed to process text, images, and audio within a large context window. The release marks a significant step in open AI model availability, though hardware requirements and independent evaluations are still emerging. This development is relevant for researchers and developers seeking advanced cross-modal reasoning capabilities.

Inkling is described as a decoder-only Mixture-of-Experts model with 975 billion total parameters and 41 billion active during processing, trained on 45 trillion tokens across multiple data types. It employs a sparse architecture with 256 experts, selecting six routed experts per input, and combines various attention mechanisms to handle multimodal data. The model supports a one-million-token context window, enabling long-range reasoning across text, images, and audio.

Access is provided via Hugging Face, with support in popular inference frameworks such as Transformers, SGLang, vLLM, and llama.cpp. Two main checkpoints are available: a BF16 version requiring approximately 2 TB of VRAM and an NVFP4 version needing around 600 GB, making deployment beyond typical consumer hardware. The release includes a hierarchical image patching module and audio converted into mel-spectrograms for processing.

While the release emphasizes the model’s potential for domain-specific fine-tuning in scientific, media, and enterprise contexts, no independent benchmark results, safety evaluations, or licensing details have been disclosed. The model’s performance, safety, and practical deployment considerations remain under evaluation by the AI community.

At a glance
announcementWhen: announced July 2026
The developmentThinking Machines released Inkling on Hugging Face, providing access to a large-scale multimodal AI model with significant hardware demands and limited independent testing.
At a glance
announcementWhen: announced on Hugging Face; the source m…
The developmentThinking Machines has made its Inkling multimodal model available through Hugging Face with day-one support from several major inference frameworks.

Implications of Inkling’s Open Multimodal Scale

The release of Inkling introduces a new level of scale and multimodal integration in AI models, potentially enabling more sophisticated applications across scientific research, media analysis, and enterprise workflows. Its open availability could accelerate innovation, but the high hardware demands and lack of independent validation mean widespread adoption may be limited initially. The model’s capacity to reason across text, images, and audio within a single framework marks a notable advancement in AI capabilities, though practical deployment remains challenging for most users.

Amazon

high VRAM external GPU for AI training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Large Multimodal Models and Open Releases

Recent years have seen rapid growth in large language models and multimodal AI systems, with companies releasing increasingly massive models for research and commercial use. Prior to Inkling, models like GPT-4 and PaLM have demonstrated multimodal capabilities, but often with limited open access and substantial hardware requirements. The trend toward open models aims to democratize AI development, though scalability and safety remain concerns.

Thinking Machines, a relatively new player in the AI space, has now introduced Inkling, positioning it as an open alternative that combines scale with multimodal reasoning. The release follows a pattern of major AI labs sharing large models, but with limited independent validation and unclear licensing terms, raising questions about safety, bias, and practical usability.

“This model is huge.”

— Hugging Face

MX3 M.2 AI Accelerator

MX3 M.2 AI Accelerator

High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional…

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unverified Aspects and Performance Unknowns

Independent benchmark results, safety assessments, and real-world performance data are not yet available. It remains unclear how well Inkling performs on practical workloads, especially with video processing, or how its speed and accuracy compare to existing models. Licensing terms and restrictions are also not specified, leaving questions about usage rights and open-source status.

ASUS Dual NVIDIA GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Graphics Card, (PCIe 5.0, DLSS 4, HDMI 2.1b, DisplayPort 2.1b, 2.5-Slot, Axial-tech Fan, 0dB Technology), 3 Year Warranty

ASUS Dual NVIDIA GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Graphics Card, (PCIe 5.0, DLSS 4, HDMI 2.1b, DisplayPort 2.1b, 2.5-Slot, Axial-tech Fan, 0dB Technology), 3 Year Warranty

AI Performance: 767 AI TOPS

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Evaluation and Deployment

Developers and researchers are expected to begin testing Inkling through supported inference frameworks, with early assessments focusing on latency, memory use, and accuracy across modalities. Independent evaluations, safety testing, and domain-specific fine-tuning are anticipated in the coming months. Clarifications on licensing, deployment costs, and model safety will be critical for broader adoption and responsible use.

ARCHITECTING RELIABLE INDUSTRIAL AI: EDGE DEPLOYMENT, MULTIMODAL AGENT, AND VERIFICATION: Building Safe, Low-Latency LLM and Vision Systems for Manufacturing, Infrastructure, and Mission-Critical

ARCHITECTING RELIABLE INDUSTRIAL AI: EDGE DEPLOYMENT, MULTIMODAL AGENT, AND VERIFICATION: Building Safe, Low-Latency LLM and Vision Systems for Manufacturing, Infrastructure, and Mission-Critical

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Inkling?

Inkling is a large-scale, 975-billion-parameter multimodal AI model from Thinking Machines, capable of processing text, images, and audio within a long context window, designed for domain-specific fine-tuning.

Can Inkling process videos?

While the architecture supports inputs with a temporal dimension, native video performance has not been evaluated, and no specific claims about video processing capabilities are currently available.

Can I run Inkling on my personal computer?

Given its hardware demands—requiring around 2 TB of VRAM for the BF16 checkpoint—full deployment on typical consumer systems is unlikely. Access through hosted inference services is recommended for most users.

Is Inkling open source?

The release describes it as an open model, but licensing details, restrictions, and availability of training code or complete data are not yet specified.

How does Inkling compare to other multimodal models?

Independent benchmark results are not yet available, so performance comparisons with models like GPT-4 or PaLM remain unclear. Evaluation is ongoing.

Source: ThorstenMeyerAI.com

You May Also Like

Sony seemingly really serious about eliminating PS5 shovelware, as one such publisher gets hit with a new set of “stricter guidelines”

Sony has introduced new, stricter publishing guidelines for PS5 games, targeting low-quality titles and shovelware, as part of its efforts to improve game quality.

QAtrial: Compliance That Shows Its Work

QAtrial introduces an open-source, provenance-first compliance platform designed for regulated life sciences, ensuring AI assistance meets strict validation standards.

The Kill Switch: What the Anthropic Export Ban Really Costs the AI Industry

Anthropic’s models were abruptly shut down by US export controls, raising concerns over industry reliance on AI and regulatory risks.

Cloud’s Hidden Memory Bill

Memory shortages are increasing cloud costs through hidden surcharges, impacting prices for cloud services and prompting reconsideration of on-premise solutions.