Skip to content

#Hugging Face

1 today

Jun 18Thursday

AI HOT (Curated Pool)

Is it agentic enough? Benchmarking open models on your own tooling

Hugging Face used its own transformers library as a testbed to see how much work open models really need when writing code, calling APIs, and debugging themselves. Instead of just scoring right or wrong, they measured how many detours and tokens each run took, and found that library docs and API design directly affect agent cost and success rate. The post lays out a full open-model harness with the pi coding agent but does not disclose final model rankings.

Why it matters: Hugging Face benchmarks open models' agent coding on their own transformers library, breaking down token costs and success rates — real engineering, not just scores. Hits all three HKR axes, but it's a methodology innovation rather than a capability breakthrough, landing at 78...

Jun 17Wednesday

AI HOT (Curated Pool)

AWS open-sources Strands Robots SDK: one agent stack from Hugging Face Hub to physical robots

AWS released the Strands Robots SDK under Apache 2.0, wrapping the LeRobot stack into a unified agent. It defaults to MuJoCo simulation with no hardware needed; switch to mode="real" for physical robots. Recorded demos are saved as LeRobotDataset and can be pushed to Hugging Face Hub. Policies like GR00T or LerobotLocal run inference, then broadcast commands to multiple robots over Zenoh mesh. Simulation and hardware code are identical except for one keyword argument. Examples run in a notebook with Python 3.12+ on Linux/macOS, no GPU required.

Why it matters: AWS wraps LeRobot into a unified agent SDK with one-click sim-to-real switching — a solid tool for robotics devs. But pure physical robotics has limited resonance with AI app-layer readers, so R axis isn't fully hit, landing right at the featured threshold.

Hugging Face Blog

Hugging Face launches ARD discovery tool so agents can search for tools, skills, and other agents

Hugging Face released Discover Tool, a reference implementation of the Agentic Resource Discovery (ARD) spec. ARD is an open draft co-developed by Microsoft, Google, GoDaddy, Hugging Face, and others. It lets agents find MCP tools, A2A agents, or skills at runtime via natural-language search instead of hardcoding each one. Hugging Face's implementation wraps the Hub's existing semantic search and Agent Skills into an ARD catalog, exposed as a REST API and an MCP Tool. The post does not disclose pricing, search latency, or accuracy figures.

Why it matters: ARD tackles a real pain point—agent tool discovery—with cross-vendor backing from Microsoft, Google, and Hugging Face, plus a working reference implementation. Not scoring higher because it's still an open draft, not a ratified standard, and the post doesn't spell out adoption...

Jun 12Friday

AI HOT (Curated Pool)

Hugging Face open-sourced Open-R1, a full reproduction of DeepSeek-R1

Hugging Face published Open-R1 on GitHub, aiming to fully reproduce the DeepSeek-R1 reasoning model. The repo has 26.1k stars and 2.4k forks so far. The body only contains the repo's landing page navigation and metadata; it does not disclose the implementation plan, training data, reproduction progress, or benchmark results. I'd treat this as a public reproduction scaffold and collaboration hub for now, and wait for a technical report before judging fidelity.

Why it matters: Hugging Face launched a full open-source reproduction of DeepSeek-R1, with the repo already at 26.1k stars — strong community interest. But the body only contains project scaffolding and navigation; no implementation plan, training data, or reproduction progress is disclosed y...

Jun 11Thursday

AI HOT (Curated Pool)

Google DeepMind open-sources DiffusionGemma, 4x faster text generation

Google DeepMind released and open-sourced DiffusionGemma, a diffusion-based text generation model that runs up to 4x faster than similarly sized autoregressive models. It replaces sequential token-by-token decoding with parallel denoising, cutting latency while preserving quality. Weights, code, and training recipes are available on Hugging Face.

Why it matters: Google DeepMind open-sourced a diffusion-based text generator with a concrete 4x speed claim and full artifacts. Missing param count and benchmarks keep it below 80, but the novel approach and open release make it a strong signal for inference-focused readers.

Jun 9Tuesday

AI HOT (Curated Pool)

How an Agent Chains Two HuggingFace Spaces to Build a 3D Paris Gallery

A coding agent chained ideogram-ai/ideogram4 and VAST-AI/TripoSplat to generate Paris monument images, reconstruct single-image 3D Gaussian splats as .ply files, convert them to .ksplat with about 3× smaller size, and deploy a static Three.js Space using APIs exposed through agents.md.

Why it matters: HKR-H/K/R all pass, but this is a Hugging Face Spaces tutorial-style build, not a model or platform release. The concrete chain and ~3x compression place it in the 72-77 featured band.

AI HOT (Curated Pool)

Migrating GitHub CI to Hugging Face Jobs

Hugging Face describes using huggingface/jobs-actions to run GitHub Actions CI as HF Jobs, where the Trackio project cut CPU job time by about 30% and added a GPU test suite using CPU, t4-small, or h200 hardware.

Why it matters: HKR-H/K/R pass via a concrete CI-to-HF Jobs workflow, ~30% speedup, and GPU-test pain point. Scope is ML tooling, not a major platform release, so it sits at the featured threshold.

Jun 8Monday

r/LocalLLaMA

OpenEnv Is Now Owned by HF, Torch, Prime Intellect, Unsloth, Modal, Mercor, and More

OpenEnv moved to committee coordination with 9 initial members, including Meta-PyTorch, Unsloth, Modal, Prime Intellect, Nvidia, and Mercor, while the post describes it as a tool for creating agent execution environments such as terminals and browsers.

Why it matters: HKR-H/K/R pass, but the post is thin: it gives committee ownership and 9 initial members. This is a mid-weight open-source agent-infra governance update, not a must-write release.

AI HOT (Curated Pool)

Open-source community backs OpenEnv for agentic reinforcement learning

Hugging Face announced broader OpenEnv access, coordinated by a committee from Meta-PyTorch, Reflection, and Unsloth; the project provides Gymnasium-style APIs and first-class MCP support for terminal and browser agent environments.

Why it matters: HKR-H/K/R all pass: this is not a model launch, but OpenEnv ties agent-RL environments, a Gymnasium-style API, and MCP into open governance, making it a solid infra story.

Jun 7Sunday

AI HOT (Curated Pool)

Five Labs, Five Minds: Building a Multi-Model Financial Drama Game with Small Models

Thousand Token Wood v2 uses four small models from different labs to drive agents in a financial simulation game, with vLLM 0.22.1’s CUDA toolkit dependency identified as the main serving friction, while a fine-tuned 0.5B Qwen reached 0% self-trading and 100% valid quotes.

Why it matters: HKR-H/K/R all pass: the small-model finance game is a real hook, with vLLM and 0.5B Qwen metrics, plus agent-engineering resonance. Scope remains an experiment, so it sits in low featured.

r/LocalLLaMA

Cohere's Unreleased Coding Model Gets Early Access for LocalLLaMA

Cohere employee Nick Frosst opened early testing of BLS-Mini-Code-1.0 to LocalLLaMA, with weights on Hugging Face before public launch. The coding model has 30B total parameters and 3B active parameters, and Cohere says token output tests are in line with similar models in its size class.

Why it matters: HKR-H/K/R all pass: early-access Cohere coding weights with 30B/3B specifics matter to local-model users. Reddit sourcing and missing evals, license, and training details keep it in the low featured band.

Jun 4Thursday

r/LocalLLaMA

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 on Hugging Face

NVIDIA released Nemotron-3-Ultra-550B-A55B-BF16 with 550B total parameters, 55B active parameters, a 1M-token context window, and minimum hardware listed as 8x H200, 16x H100, or 8x GB200/B200/GB300/B300.

Why it matters: HKR-H/K/R all pass: NVIDIA open-weight scale, 550B/55B active params, and 1M context are concrete. Missing benchmarks, license, and availability details keep it in the 78–84 band, not P1.

AI HOT (Curated Pool)

Hugging Face redesigns hf CLI output format for coding agents

Hugging Face redesigned hf CLI output for coding agents including Claude Code and Codex, using environment-variable detection and compact untruncated TSV output; in complex multi-step tasks, agents without the CLI used up to 6 times more tokens.

Why it matters: HKR-H/K/R pass: the story has a clear agent-CLI hook, a concrete TSV/token mechanism, and strong developer cost resonance. It stays in the featured band because this is a tooling update, not a model or platform release.

Jun 3Wednesday

r/LocalLLaMA

google/gemma-4-12B on Hugging Face

Google DeepMind released Gemma 4 open-weight models in five sizes, with the 12B variant supporting text, image, and audio input, instruction-tuned and pre-trained variants, native system prompts, function calling, and a context window of up to 256K tokens.

Why it matters: Gemma 4 clears HKR-H/K/R: open weights, multimodal input, and 256K context make it more than a routine update. Missing benchmarks, license detail, and fuller official context keep it in the 78–84 band.

NVIDIA Blog

NVIDIA Research Presents Grasping, Autonomous Driving and Agent Training Work at CVPR

NVIDIA Research presented three physical AI papers at CVPR: GraspGen-X was trained on 2 billion simulated grasps, LCDrive cuts reasoning tokens by about half versus text-based reasoning, and NitroGen trains embodied agents across more than 1,000 games and 40,000 hours of interaction.

Why it matters: HKR-H/K/R all pass: NVIDIA’s CVPR bundle gives concrete mechanisms and scale numbers. It stays in the low 78–84 band because it is a vendor research roundup, not a major model or product launch.

QbitAI · WeChat

Papers with Code returns with CVPR coverage and Hugging Face-led rebuild

Hugging Face’s open-source team launched paperswithcode.co in May 2026, using AI agents to parse papers and restore SOTA leaderboards tied to the original platform’s 9,300-plus benchmarks.

Why it matters: HKR-H/K/R all pass: a beloved research portal returns, with 9,300 restored leaderboards and agent-based paper parsing. The impact is strong for research workflows, not model-release scale.

Jun 2Tuesday

AI HOT (Curated Pool)

Holo3.1: Fast Local Computer-Use Agents

Holo3.1 releases Qwen-based computer-use agents in 0.8B, 4B, 9B, and 35B-A3B sizes, with FP8, Q4 GGUF, and NVFP4 quantized checkpoints for local inference and a 79.3% AndroidWorld score for the 35B-A3B model.

Why it matters: HKR-H/K/R all pass: Holo3.1 pairs a local computer-use agent with concrete model sizes and quantized checkpoints. It fits the 78–84 band, below major lab model-release weight.

r/LocalLLaMA

NVIDIA releases Cosmos 3 Omnimodal world models on Hugging Face

NVIDIA released Cosmos 3 on Hugging Face with Nano at 16B parameters and Super at 64B parameters; the post says the models generate video, images, audio, and action commands from text, image, video, and action-trajectory inputs.

Why it matters: HKR-H/K/R all pass: NVIDIA world models on HF, concrete 16B/64B variants, and multimodal robotics relevance. Missing benchmarks, license, and training details keep it in the 78–84 band.

Jun 1Monday

AI HOT (Curated Pool)

Introducing Mellum2: JetBrains' 12B Mixture-of-Experts Model

JetBrains published a Hugging Face blog post introducing Mellum2, confirming a mixture-of-experts architecture and a 12B parameter scale; the snippet does not disclose training data, license, benchmarks, or deployment conditions.

Why it matters: HKR-H/K/R all pass, but the body only confirms 12B and MoE, with no benchmarks, license, context window, or IDE integration terms. Treat as a mid-weight model release at the lower featured band.

May 31Sunday

r/LocalLLaMA

13 abliterated Gemma 4 E2B variants, 44 GPU hours, benchmark and comparison

Abliterlitics tested 13 abliterated Gemma 4 E2B variants using 44 RTX 5090 GPU hours, and HarmBench ASR rose from the base model’s 32.2% to 82%–100%, while coder3101 scored 84.8% on GSM8K versus the base model’s 83.5%.

Why it matters: HKR-H/K/R all pass, with a named first-person benchmark and concrete numbers. Scope stays narrow around abliterated Gemma 4 E2B variants, so it lands at the featured threshold rather than a must-write item.

May 29Friday

r/LocalLLaMA

Liquid AI releases LFM2.5-8B-A1B

Liquid AI released LFM2.5-8B-A1B with a 128K context window, 38T pre-training tokens, large-scale reinforcement learning, doubled vocabulary for non-Latin tokenization, and availability on Hugging Face.

Why it matters: HKR-H/K/R pass: 8B/A1B, 128K context, and 38T tokens are concrete hooks for local inference. No benchmarks, license, or deployment limits are disclosed, so it stays in the mid featured band.

May 28Thursday

r/LocalLLaMA

Nvidia LocateAnything: Fast Vision-Language Grounding with Parallel Box Decoding

The title says Nvidia LocateAnything-3B performs vision-language grounding with parallel box decoding and runs 10x faster than Qwen3-VL; the post body only provides Hugging Face, GitHub, demo, and project links, and does not disclose benchmark setup or accuracy numbers.

Why it matters: HKR-H/K/R all pass, but the body is mostly links and title-level facts, with no full eval setup or quality metrics. NVIDIA open vision grounding is useful enough for featured, not same-day must-write.

Synced · WeChat

Chinese pretrained embodied model Wall-OSS-0.5 is open sourced

X Square Robot open sourced Wall-OSS-0.5, a VLA model whose 400k pretraining checkpoint scored above 80 on 4 of 17 real-robot zero-shot tasks, with weights, code, training recipe, ablations, and a DMuon optimizer implementation released.

Why it matters: Clear HKR-H/K/R: a 400k checkpoint and 17 real-robot zero-shot tasks add substance, while “post-training not required” is a sharp hook. X Square Robot is not a top foundation-model lab, so this stays at 79.

r/LocalLLaMA

I built a 103B-token Usenet corpus from 1980–2013

OwnerByDane released a 103.1B-token Usenet corpus covering 1980–2013, 408M posts, and 18,347 newsgroups, with free 5K-post-per-hierarchy samples and full-corpus licensing available.

Why it matters: HKR-H/K/R all pass: the zero-contamination corpus has a clear hook, concrete scale, and relevance to training-data scarcity. Score is capped by Reddit-only sourcing, licensed full access, and no third-party validation or benchmark results.

May 27Wednesday

r/LocalLLaMA

I ran 8 open-weight models as agents in a persistent MMO for 10 days

Firespawn Studios ran 25 agents across 8 open-weight models for 10 days in Null Epoch Season 0 and released about 93,000 logged events, with roughly 70% of actions including the model’s reasoning or justification.

Why it matters: HKR-H/K/R all pass: a concrete 10-day MMO agent trial with 25 agents and 93k events. Reddit sourcing limits reach, so it lands in the 78–84 good-quality band, not P1.

AI HOT (Curated Pool)

Reachy Mini enables fully local voice interaction

Reachy Mini implements local voice interaction through the speech-to-speech library, using a cascaded pipeline with a Realtime API-compatible WebSocket interface and default components including Silero VAD, Parakeet-TDT, and Qwen3-TTS.

Why it matters: HKR-H/K/R all pass: the post has a clear local-robot voice hook, concrete stack details, and edge-agent resonance. Scope stays limited to Reachy Mini voice interaction, so it sits at the featured threshold.

AI HOT (Curated Pool)

Shipping a Trillion Parameters With a Hub Bucket: Delta Weight Sync in TRL

Hugging Face merged TRL PR 5417 for delta weight sync, sending only changed weights as sparse safetensors via a Hugging Face Bucket; on Qwen3-0.6B, the per-step payload falls from 1.2GB to 20–35MB.

Why it matters: HKR-H/K/R all pass: TRL gets delta weight sync with a concrete sparse-safetensors mechanism and a 1.2GB to 20–35MB example. Scope is training infra, so it stays below must-write.

r/LocalLLaMA

PrismML Released Binary and Ternary Bonsai Image 4B

PrismML released Binary and Ternary Bonsai Image 4B, 1-bit and ternary text-to-image diffusion transformers around 3GB, compared with FLUX.2 Klein 4B at about 16GB, with browser-local WebGPU demo links and an Apache-2.0 license disclosed in the Reddit snippet.

Why it matters: HKR-H/K/R all pass: low-bit image DiT plus local browser inference is a strong hook, backed by 4B, ~3GB, WebGPU, and license details. Reddit sourcing and limited lab weight keep it in the low featured band.

May 25Monday

AI HOT (Curated Pool)

Harness, Scaffold, and AI Agent Terminology Explained

Hugging Face’s post frames an agent as three layers: Model, Scaffolding, and Harness; Scaffolding defines behavior through prompts and tool descriptions, while Harness runs model calls, tool calls, and control loops.

Why it matters: HKR-H/K/R pass: the Hugging Face post gives a concrete agent-stack taxonomy. It clears featured on practitioner relevance, but lacks a release, benchmark, or deployment case, so it stays at the threshold.

May 23Saturday

r/LocalLLaMA

meituan-longcat/LongCat-Video-Avatar-1.5 on Hugging Face

Meituan LongCat released LongCat-Video-Avatar-1.5 on Hugging Face, supporting AT2V, ATI2V, and video continuation while replacing Wav2Vec2 with Whisper-Large and using DMD2 distillation to reduce inference to 8 NFE; the model weights are released under the MIT License.

Why it matters: HKR-H/K/R all pass: open MIT video-avatar weights plus 8 NFE inference give local multimodal builders real signal. This is a mid-weight open-source model update, not an 85+ same-day industry event.

May 22Friday

AI HOT (Curated Pool)

Text Degeneration: A Production Failure Mode Most Benchmarks Do Not Track

Dharma-AI says in a Hugging Face post that large language models can produce repeated, incoherent, or logically confused text in production, and most mainstream benchmarks do not track this failure mode.

Why it matters: HKR-H/K/R all pass, but the post only discloses the failure pattern and benchmark blind spot, with no sample size, metric, or reproduction setup. This fits the lower featured threshold.

May 21Thursday

r/LocalLLaMA

HRM 1B

Sapientinc released HRM-Text 1B Base and its training code, and the paper claims competitive performance against 2–7B open models while using 100–900x fewer training tokens and 96–432x less estimated compute, with training on 16 H100 GPUs taking about 46 hours and costing about $1,472.

Why it matters: HKR-H/K/R all pass: HRM-Text 1B has concrete low-cost training numbers and released code. Capped at 80 because this is a Reddit item and the efficiency claim still lacks independent evaluation.

May 20Wednesday

Hacker News front page

Show HN: Lance – Image/video generation and understanding in one model

ByteDance released Lance as a research project for image and video generation and understanding in one model; the RSS snippet states 3B active parameters, fewer than 128 GPUs used for training, and links to a homepage, arXiv paper, and Hugging Face model, while the post does not disclose benchmark results or licensing terms.

Why it matters: ByteDance’s Lance puts image/video generation and understanding in one model, with 3B active parameters and <128 GPUs for training. HKR-H/K/R all pass, but benchmarks, license details, and real outputs are not disclosed, keeping it below P1.

May 19Tuesday

AI HOT (Curated Pool)

NVIDIA fine-tunes Cosmos Predict 2.5 with LoRA/DoRA for robot video generation

NVIDIA published a Hugging Face post on fine-tuning Cosmos Predict 2.5 with LoRA and DoRA to generate robot first-person videos from text prompts; the post does not disclose dataset size, training cost, or evaluation results.

Why it matters: HKR-H/K/R pass: the robot POV video angle is clickable, and LoRA/DoRA on Cosmos Predict 2.5 is a concrete mechanism. Missing dataset scale and metrics keep it in the low featured band.

May 18Monday

AI HOT (Curated Pool)

The Open Agent Leaderboard

IBM Research published the Open Agent Leaderboard on Hugging Face to evaluate agents across language understanding, tool use, and multi-step reasoning tasks; the post does not disclose dataset size, model scores, or the evaluation date.

Why it matters: HKR-H and HKR-R pass because an open agent leaderboard speaks to agent-eval pain. HKR-K fails: the article lacks scores, dataset size, and evaluation date, so it sits at the featured threshold.

May 17Sunday

r/LocalLLaMA

85 GPU-hours comparing 5 abliteration methods on Qwen3.6-27B

Abliterlitics compared five Qwen3.6-27B abliteration variants against the base model using 85 GPU-hours of benchmarks, HarmBench, KL divergence, and weight forensics; Huihui had the smallest benchmark deltas, Heretic had the lowest KL divergence, and all five variants reached near-complete safety removal.

Why it matters: HKR-H/K/R all pass: the post gives an 85-GPU-hour comparison across five abliteration methods on Qwen3.6-27B. Niche open-model safety work, not a lab release, so it stays at the featured threshold.

May 15Friday

AI HOT (Curated Pool)

Granite Embedding Multilingual R2: Open Multilingual Embedding Model with 32K Context

IBM Granite released Granite Embedding Multilingual R2 on Hugging Face under Apache 2.0, with fewer than 100 million parameters, a 32K-token context length, and top same-scale retrieval performance on MTEB according to the post.

Why it matters: HKR-H/K/R pass: the 32K-context, sub-100M multilingual embedding model gives RAG builders a concrete open-source option. Impact is narrower than a frontier-model release, so it sits at the featured threshold.

r/LocalLLaMA

MOOSE-Star (ICML 2026): 7B Model and 108K-Paper Dataset for Scientific Hypothesis Discovery

MiroMind researchers released the MOOSE-Star collection with three 7B models and TOMATO-Star, a dataset of 108,717 NCBI papers. MS-IR-7B reaches 54.37% inspiration-retrieval accuracy, uses DeepSeek-R1-Distill-Qwen-7B as its base, runs at about 14GB fp16, and supports llama.cpp, vLLM, and SGLang.

Why it matters: HKR-H/K/R all pass via the local 7B research-agent hook and concrete dataset metrics. Single Reddit source and limited lab gravity keep it below the must-write band.

r/LocalLLaMA

inclusionAI/Ring-2.6-1T on Hugging Face

inclusionAI released Ring-2.6-1T, a 1T-parameter reasoning model on Hugging Face; it supports high and xhigh reasoning effort levels, targets agent workflows and long-horizon tasks, and uses Async RL with the IcePop algorithm for reinforcement-learning training stability.

Why it matters: HKR-H/K/R pass: a 1T HF model with two reasoning modes and named training methods is real signal. Benchmarks, license, and inference cost are not disclosed, so this stays at the lower edge of featured.

May 14Thursday

r/LocalLLaMA

Automated AI researcher running locally with llama.cpp

Hugging Face’s ml-intern added local-model support through llama.cpp and ollama; the post says Qwen3.6-35B-A3B can orchestrate CPU/GPU sandboxes and Hub jobs to run an end-to-end SFT workflow.

Why it matters: HKR-H/K/R all pass, but this is a Reddit-sourced open-source tool update, not a major model release. Local sandbox and Hub-job orchestration for SFT put it just above the featured threshold.