Inkling review: Mira Murati’s 975B open-source model

Thinking Machines Lab released Inkling on July 15. The 975‑billion‑parameter multimodal model is on Hugging Face under Apache 2.0 and reviewers call it the leading Western open‑source model.

Thinking Machines Lab released Inkling on July 15. The model has 975 billion parameters in a mixture‑of‑experts architecture and uses about 41 billion active parameters at inference. It accepts text, images and audio, supports a one‑million‑token context window and was pretrained on roughly 45 trillion tokens. Full model weights are published on Hugging Face under the Apache 2.0 license.

Thinking Machines Lab developed Inkling over two years after Mira Murati left OpenAI in September 2024. The model is available through OpenRouter at a listed price of $1 per million input tokens and $4.05 per million output tokens. OpenRouter routing setups such as Hermes or OpenClaw can integrate Inkling without additional configuration. Inkling requires data‑center hardware and is not intended for local device use.

Independent evaluations using the Model Context Protocol (MCP Atlas) reported a 74.1% task completion rate on real‑world agentic tasks. On SWE‑Bench Verified, a benchmark for autonomous GitHub bug fixing, Inkling scored 77.6%. Those scores place it ahead of Nvidia’s Nemotron on the same measures.

Hands‑on tests covered coding, creative writing, associative tasks and logic puzzles. On complex coding prompts Inkling failed to produce runnable outputs. On simpler coding prompts it produced working but rudimentary code; one comparison showed a 27‑billion‑parameter compressed model running on a phone produced a more complete game from the same prompt.

In creative writing tests Inkling generated vivid prose and performed web fetches before composing, but it included invented cultural details that were inaccurate. In an associative creativity task the model produced strong metaphorical language in one section and failed to link a final element coherently in another. In a bridge‑and‑torch logic puzzle variant the model returned a 17‑minute solution tied to a common training example while the unrestricted prompt’s correct answer is 10 minutes when all four people cross together.

Inkling enforces safety filters and declined two tested requests: one seeking advice on seduction and another from a self‑described opioid user asking how to avoid job loss. In the latter case the model offered professional help resources rather than tactical workplace advice. Removing or altering safety layers would require substantial compute given the model’s size; smaller open models offer different safety and deployment tradeoffs.

The Apache 2.0 license and Western provenance provide a permissive legal path for organizations that require modifiable weights from a Western provider. The model’s agentic task scores and published weights enable integration into enterprise workflows that route through OpenRouter. Developers prioritizing local deployment, lower cost or uncensored outputs may prefer smaller, cheaper models.

Inkling is the first major train‑from‑scratch model released by Thinking Machines Lab since Murati’s departure from OpenAI. The lab has published the full weights and documentation for users who meet the hardware and licensing requirements.

The material on GNcrypto is intended solely for informational use and must not be regarded as financial advice. We make every effort to keep the content accurate and current, but we cannot warrant its precision, completeness, or reliability. GNcrypto does not take responsibility for any mistakes, omissions, or financial losses resulting from reliance on this information. Any actions you take based on this content are done at your own risk. Always conduct independent research and seek guidance from a qualified specialist. For further details, please review our Terms, Privacy Policy and Disclaimers.

Articles by this author