📊 Full opportunity report: Introducing Inkling: The Future Of AI Innovation on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Thinking Machines has launched Inkling, a massive 975-billion-parameter multimodal AI model, available on Hugging Face. Its open access offers new possibilities for cross-modal reasoning, but hardware requirements and evaluation details remain uncertain.
Thinking Machines has released Inkling on Hugging Face, a 975-billion-parameter multimodal model designed to process text, images, and audio within a large context window. The release marks a significant step in open AI model availability, though hardware requirements and independent evaluations are still emerging. This development is relevant for researchers and developers seeking advanced cross-modal reasoning capabilities.
Inkling is described as a decoder-only Mixture-of-Experts model with 975 billion total parameters and 41 billion active during processing, trained on 45 trillion tokens across multiple data types. It employs a sparse architecture with 256 experts, selecting six routed experts per input, and combines various attention mechanisms to handle multimodal data. The model supports a one-million-token context window, enabling long-range reasoning across text, images, and audio.
Access is provided via Hugging Face, with support in popular inference frameworks such as Transformers, SGLang, vLLM, and llama.cpp. Two main checkpoints are available: a BF16 version requiring approximately 2 TB of VRAM and an NVFP4 version needing around 600 GB, making deployment beyond typical consumer hardware. The release includes a hierarchical image patching module and audio converted into mel-spectrograms for processing.
While the release emphasizes the model’s potential for domain-specific fine-tuning in scientific, media, and enterprise contexts, no independent benchmark results, safety evaluations, or licensing details have been disclosed. The model’s performance, safety, and practical deployment considerations remain under evaluation by the AI community.
Implications of Inkling’s Open Multimodal Scale
The release of Inkling introduces a new level of scale and multimodal integration in AI models, potentially enabling more sophisticated applications across scientific research, media analysis, and enterprise workflows. Its open availability could accelerate innovation, but the high hardware demands and lack of independent validation mean widespread adoption may be limited initially. The model’s capacity to reason across text, images, and audio within a single framework marks a notable advancement in AI capabilities, though practical deployment remains challenging for most users.
high VRAM external GPU for AI training
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on Large Multimodal Models and Open Releases
Recent years have seen rapid growth in large language models and multimodal AI systems, with companies releasing increasingly massive models for research and commercial use. Prior to Inkling, models like GPT-4 and PaLM have demonstrated multimodal capabilities, but often with limited open access and substantial hardware requirements. The trend toward open models aims to democratize AI development, though scalability and safety remain concerns.
Thinking Machines, a relatively new player in the AI space, has now introduced Inkling, positioning it as an open alternative that combines scale with multimodal reasoning. The release follows a pattern of major AI labs sharing large models, but with limited independent validation and unclear licensing terms, raising questions about safety, bias, and practical usability.
“This model is huge.”
— Hugging Face
professional AI inference hardware
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unverified Aspects and Performance Unknowns
Independent benchmark results, safety assessments, and real-world performance data are not yet available. It remains unclear how well Inkling performs on practical workloads, especially with video processing, or how its speed and accuracy compare to existing models. Licensing terms and restrictions are also not specified, leaving questions about usage rights and open-source status.
large capacity VRAM graphics card
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for Evaluation and Deployment
Developers and researchers are expected to begin testing Inkling through supported inference frameworks, with early assessments focusing on latency, memory use, and accuracy across modalities. Independent evaluations, safety testing, and domain-specific fine-tuning are anticipated in the coming months. Clarifications on licensing, deployment costs, and model safety will be critical for broader adoption and responsible use.
multimodal AI model deployment hardware
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Inkling?
Inkling is a large-scale, 975-billion-parameter multimodal AI model from Thinking Machines, capable of processing text, images, and audio within a long context window, designed for domain-specific fine-tuning.
Can Inkling process videos?
While the architecture supports inputs with a temporal dimension, native video performance has not been evaluated, and no specific claims about video processing capabilities are currently available.
Can I run Inkling on my personal computer?
Given its hardware demands—requiring around 2 TB of VRAM for the BF16 checkpoint—full deployment on typical consumer systems is unlikely. Access through hosted inference services is recommended for most users.
Is Inkling open source?
The release describes it as an open model, but licensing details, restrictions, and availability of training code or complete data are not yet specified.
How does Inkling compare to other multimodal models?
Independent benchmark results are not yet available, so performance comparisons with models like GPT-4 or PaLM remain unclear. Evaluation is ongoing.
Source: ThorstenMeyerAI.com