Skip to content

Hugging Face

The Hugging Face community: trending models and datasets, leaderboard shifts, the open-source barometer.

Latest picks

101–120 of 179

Jul 16Thursday

Hacker News front page

Mira Murati's Thinking Machines releases Inkling, a 975B open-weights MoE model for text, images, and audio

Inkling is a 975B total / 41B active parameter Mixture-of-Experts model with a 1M-token context window and native text, image, and audio input. Thinking Machines compares it against Nemotron 3 Ultra, GLM 5.2, GPT 5.6 Sol, and Claude Fable 5, claiming frontier-level performance in general intelligence, agentic coding, and speech. Weights are on Hugging Face and fine-tuning is available via the Tinker platform. The post does not disclose training data, training cost, inference latency, or specific benchmark scores—take the comparison charts with a grain of salt.

Why it matters: Mira Murati's first model release post-OpenAI — 975B MoE, open weights, benchmarked against GPT 5.6 Sol and Claude Fable 5. This is an industry-level event. HKR all hit: name recognition + architectural detail + open-weight disruption. Held back from 95+ because we only have s...

Jul 15Wednesday

Hugging Face Blog

Thinking Machines releases Inkling: a 1T-param, natively multimodal open model

Inkling is an open ~1T-param model that natively accepts image, audio, and text inputs with a 1M context window. Trained on 45T multimodal tokens, it uses a MoE architecture with 975B total and 41B active parameters. It includes MTP speculative decoding layers for faster inference and ships in BF16 and NVFP4 variants. Hugging Face provides day-0 support in transformers, SGLang, vLLM, and llama.cpp, covering agentic coding, multimodal vision, and audio tasks.

Why it matters: A new player, Thinking Machines, open-sources a trillion-parameter multimodal MoE model with solid specs (975B/41B activated, 1M context, 45T tokens trained) and MTP speculative decoding. H and K both hit, but R is weak — the team has no name recognition, no emotional anchor. ...

Jul 14Tuesday

TechCrunch · AI

The real AI race may no longer be at the frontier

Hugging Face CEO Clem Delangue says enterprises increasingly pick open models for cost, accessibility, and ownership. Chinese open-weight models hit 41% of Hugging Face downloads this spring, overtaking US models. The top six models on OpenRouter are all from Chinese firms — Tencent, Xiaomi, DeepSeek, MiniMax, and Z.ai. Anthropic's Claude Opus 4.7 trails behind. The post doesn't give absolute download numbers or enterprise adoption rates, but the direction is clear: open models are taking production workloads from frontier closed models.

Why it matters: Hugging Face CEO argues with download data that open models, not frontier ones, are the real battleground — 41% of HF downloads are Chinese models, top six on OpenRouter all Chinese. Solid HKR. Docked slightly because it's a single exec's framing, not an independent report, an...

Jul 10Friday

TechCrunch · AI

Hugging Face CEO: Companies are done renting AI, open source is winning

Hugging Face CEO Clem Delangue says roughly half the Fortune 500 now use the platform. The pattern he sees: companies start with APIs, then quickly move to self-hosting open models for cost, control, and data privacy. He cites one company that went from $100K/month on APIs to $10K/month with open source. Delangue argues closed-source vendors will struggle unless they offer what open models can't—extreme convenience or exclusive data. The post doesn't name the specific customer or timeline.

Why it matters: Hugging Face's CEO gives a concrete cost case for enterprises shifting from APIs to self-hosting—a 10x price gap is real data. But this is a CEO narrative for media, not third-party research, so I'm discounting it: sample size, industry breakdown, and hidden ops costs aren't a...

Jun 18Thursday

AI HOT (Curated Pool)

Is it agentic enough? Benchmarking open models on your own tooling

Hugging Face used its own transformers library as a testbed to see how much work open models really need when writing code, calling APIs, and debugging themselves. Instead of just scoring right or wrong, they measured how many detours and tokens each run took, and found that library docs and API design directly affect agent cost and success rate. The post lays out a full open-model harness with the pi coding agent but does not disclose final model rankings.

Why it matters: Hugging Face benchmarks open models' agent coding on their own transformers library, breaking down token costs and success rates — real engineering, not just scores. Hits all three HKR axes, but it's a methodology innovation rather than a capability breakthrough, landing at 78...

Jun 17Wednesday

AI HOT (Curated Pool)

AWS open-sources Strands Robots SDK: one agent stack from Hugging Face Hub to physical robots

AWS released the Strands Robots SDK under Apache 2.0, wrapping the LeRobot stack into a unified agent. It defaults to MuJoCo simulation with no hardware needed; switch to mode="real" for physical robots. Recorded demos are saved as LeRobotDataset and can be pushed to Hugging Face Hub. Policies like GR00T or LerobotLocal run inference, then broadcast commands to multiple robots over Zenoh mesh. Simulation and hardware code are identical except for one keyword argument. Examples run in a notebook with Python 3.12+ on Linux/macOS, no GPU required.

Why it matters: AWS wraps LeRobot into a unified agent SDK with one-click sim-to-real switching — a solid tool for robotics devs. But pure physical robotics has limited resonance with AI app-layer readers, so R axis isn't fully hit, landing right at the featured threshold.

Hugging Face Blog

Hugging Face launches ARD discovery tool so agents can search for tools, skills, and other agents

Hugging Face released Discover Tool, a reference implementation of the Agentic Resource Discovery (ARD) spec. ARD is an open draft co-developed by Microsoft, Google, GoDaddy, Hugging Face, and others. It lets agents find MCP tools, A2A agents, or skills at runtime via natural-language search instead of hardcoding each one. Hugging Face's implementation wraps the Hub's existing semantic search and Agent Skills into an ARD catalog, exposed as a REST API and an MCP Tool. The post does not disclose pricing, search latency, or accuracy figures.

Why it matters: ARD tackles a real pain point—agent tool discovery—with cross-vendor backing from Microsoft, Google, and Hugging Face, plus a working reference implementation. Not scoring higher because it's still an open draft, not a ratified standard, and the post doesn't spell out adoption...

Jun 12Friday

AI HOT (Curated Pool)

Hugging Face open-sourced Open-R1, a full reproduction of DeepSeek-R1

Hugging Face published Open-R1 on GitHub, aiming to fully reproduce the DeepSeek-R1 reasoning model. The repo has 26.1k stars and 2.4k forks so far. The body only contains the repo's landing page navigation and metadata; it does not disclose the implementation plan, training data, reproduction progress, or benchmark results. I'd treat this as a public reproduction scaffold and collaboration hub for now, and wait for a technical report before judging fidelity.

Why it matters: Hugging Face launched a full open-source reproduction of DeepSeek-R1, with the repo already at 26.1k stars — strong community interest. But the body only contains project scaffolding and navigation; no implementation plan, training data, or reproduction progress is disclosed y...

Jun 11Thursday

AI HOT (Curated Pool)

Google DeepMind open-sources DiffusionGemma, 4x faster text generation

Google DeepMind released and open-sourced DiffusionGemma, a diffusion-based text generation model that runs up to 4x faster than similarly sized autoregressive models. It replaces sequential token-by-token decoding with parallel denoising, cutting latency while preserving quality. Weights, code, and training recipes are available on Hugging Face.

Why it matters: Google DeepMind open-sourced a diffusion-based text generator with a concrete 4x speed claim and full artifacts. Missing param count and benchmarks keep it below 80, but the novel approach and open release make it a strong signal for inference-focused readers.

Jun 9Tuesday

AI HOT (Curated Pool)

How an Agent Chains Two HuggingFace Spaces to Build a 3D Paris Gallery

A coding agent chained ideogram-ai/ideogram4 and VAST-AI/TripoSplat to generate Paris monument images, reconstruct single-image 3D Gaussian splats as .ply files, convert them to .ksplat with about 3× smaller size, and deploy a static Three.js Space using APIs exposed through agents.md.

Why it matters: HKR-H/K/R all pass, but this is a Hugging Face Spaces tutorial-style build, not a model or platform release. The concrete chain and ~3x compression place it in the 72-77 featured band.

AI HOT (Curated Pool)

Migrating GitHub CI to Hugging Face Jobs

Hugging Face describes using huggingface/jobs-actions to run GitHub Actions CI as HF Jobs, where the Trackio project cut CPU job time by about 30% and added a GPU test suite using CPU, t4-small, or h200 hardware.

Why it matters: HKR-H/K/R pass via a concrete CI-to-HF Jobs workflow, ~30% speedup, and GPU-test pain point. Scope is ML tooling, not a major platform release, so it sits at the featured threshold.

Jun 8Monday

r/LocalLLaMA

OpenEnv Is Now Owned by HF, Torch, Prime Intellect, Unsloth, Modal, Mercor, and More

OpenEnv moved to committee coordination with 9 initial members, including Meta-PyTorch, Unsloth, Modal, Prime Intellect, Nvidia, and Mercor, while the post describes it as a tool for creating agent execution environments such as terminals and browsers.

Why it matters: HKR-H/K/R pass, but the post is thin: it gives committee ownership and 9 initial members. This is a mid-weight open-source agent-infra governance update, not a must-write release.

AI HOT (Curated Pool)

Open-source community backs OpenEnv for agentic reinforcement learning

Hugging Face announced broader OpenEnv access, coordinated by a committee from Meta-PyTorch, Reflection, and Unsloth; the project provides Gymnasium-style APIs and first-class MCP support for terminal and browser agent environments.

Why it matters: HKR-H/K/R all pass: this is not a model launch, but OpenEnv ties agent-RL environments, a Gymnasium-style API, and MCP into open governance, making it a solid infra story.

Jun 7Sunday

AI HOT (Curated Pool)

Five Labs, Five Minds: Building a Multi-Model Financial Drama Game with Small Models

Thousand Token Wood v2 uses four small models from different labs to drive agents in a financial simulation game, with vLLM 0.22.1’s CUDA toolkit dependency identified as the main serving friction, while a fine-tuned 0.5B Qwen reached 0% self-trading and 100% valid quotes.

Why it matters: HKR-H/K/R all pass: the small-model finance game is a real hook, with vLLM and 0.5B Qwen metrics, plus agent-engineering resonance. Scope remains an experiment, so it sits in low featured.

r/LocalLLaMA

Cohere's Unreleased Coding Model Gets Early Access for LocalLLaMA

Cohere employee Nick Frosst opened early testing of BLS-Mini-Code-1.0 to LocalLLaMA, with weights on Hugging Face before public launch. The coding model has 30B total parameters and 3B active parameters, and Cohere says token output tests are in line with similar models in its size class.

Why it matters: HKR-H/K/R all pass: early-access Cohere coding weights with 30B/3B specifics matter to local-model users. Reddit sourcing and missing evals, license, and training details keep it in the low featured band.

Jun 4Thursday

r/LocalLLaMA

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 on Hugging Face

NVIDIA released Nemotron-3-Ultra-550B-A55B-BF16 with 550B total parameters, 55B active parameters, a 1M-token context window, and minimum hardware listed as 8x H200, 16x H100, or 8x GB200/B200/GB300/B300.

Why it matters: HKR-H/K/R all pass: NVIDIA open-weight scale, 550B/55B active params, and 1M context are concrete. Missing benchmarks, license, and availability details keep it in the 78–84 band, not P1.

AI HOT (Curated Pool)

Hugging Face redesigns hf CLI output format for coding agents

Hugging Face redesigned hf CLI output for coding agents including Claude Code and Codex, using environment-variable detection and compact untruncated TSV output; in complex multi-step tasks, agents without the CLI used up to 6 times more tokens.

Why it matters: HKR-H/K/R pass: the story has a clear agent-CLI hook, a concrete TSV/token mechanism, and strong developer cost resonance. It stays in the featured band because this is a tooling update, not a model or platform release.

Jun 3Wednesday

r/LocalLLaMA

google/gemma-4-12B on Hugging Face

Google DeepMind released Gemma 4 open-weight models in five sizes, with the 12B variant supporting text, image, and audio input, instruction-tuned and pre-trained variants, native system prompts, function calling, and a context window of up to 256K tokens.

Why it matters: Gemma 4 clears HKR-H/K/R: open weights, multimodal input, and 256K context make it more than a routine update. Missing benchmarks, license detail, and fuller official context keep it in the 78–84 band.

NVIDIA Blog

NVIDIA Research Presents Grasping, Autonomous Driving and Agent Training Work at CVPR

NVIDIA Research presented three physical AI papers at CVPR: GraspGen-X was trained on 2 billion simulated grasps, LCDrive cuts reasoning tokens by about half versus text-based reasoning, and NitroGen trains embodied agents across more than 1,000 games and 40,000 hours of interaction.

Why it matters: HKR-H/K/R all pass: NVIDIA’s CVPR bundle gives concrete mechanisms and scale numbers. It stays in the low 78–84 band because it is a vendor research roundup, not a major model or product launch.

QbitAI · WeChat

Papers with Code returns with CVPR coverage and Hugging Face-led rebuild

Hugging Face’s open-source team launched paperswithcode.co in May 2026, using AI agents to parse papers and restore SOTA leaderboards tied to the original platform’s 9,300-plus benchmarks.

Why it matters: HKR-H/K/R all pass: a beloved research portal returns, with 9,300 restored leaderboards and agent-based paper parsing. The impact is strong for research workflows, not model-release scale.