Skip to content
Hugging Face Blog

Thinking Machines releases Inkling: a 1T-param, natively multimodal open model

Welcome Inkling by Thinking Machines

Inkling is an open ~1T-param model that natively accepts image, audio, and text inputs with a 1M context window. Trained on 45T multimodal tokens, it uses a MoE architecture with 975B total and 41B active parameters. It includes MTP speculative decoding layers for faster inference and ships in BF16 and NVFP4 variants. Hugging Face provides day-0 support in transformers, SGLang, vLLM, and llama.cpp, covering agentic coding, multimodal vision, and audio tasks.

Why it matters: A new player, Thinking Machines, open-sources a trillion-parameter multimodal MoE model with solid specs (975B/41B activated, 1M context, 45T tokens trained) and MTP speculative decoding. H and K both hit, but R is weak — the team has no name recognition, no emotional anchor. ...

Read the original ↗Export Markdown