Skip to content

Hugging Face

The Hugging Face community: trending models and datasets, leaderboard shifts, the open-source barometer.

Latest picks

121–140 of 179

Jun 2Tuesday

AI HOT (Curated Pool)

Holo3.1: Fast Local Computer-Use Agents

Holo3.1 releases Qwen-based computer-use agents in 0.8B, 4B, 9B, and 35B-A3B sizes, with FP8, Q4 GGUF, and NVFP4 quantized checkpoints for local inference and a 79.3% AndroidWorld score for the 35B-A3B model.

Why it matters: HKR-H/K/R all pass: Holo3.1 pairs a local computer-use agent with concrete model sizes and quantized checkpoints. It fits the 78–84 band, below major lab model-release weight.

r/LocalLLaMA

NVIDIA releases Cosmos 3 Omnimodal world models on Hugging Face

NVIDIA released Cosmos 3 on Hugging Face with Nano at 16B parameters and Super at 64B parameters; the post says the models generate video, images, audio, and action commands from text, image, video, and action-trajectory inputs.

Why it matters: HKR-H/K/R all pass: NVIDIA world models on HF, concrete 16B/64B variants, and multimodal robotics relevance. Missing benchmarks, license, and training details keep it in the 78–84 band.

Jun 1Monday

AI HOT (Curated Pool)

Introducing Mellum2: JetBrains' 12B Mixture-of-Experts Model

JetBrains published a Hugging Face blog post introducing Mellum2, confirming a mixture-of-experts architecture and a 12B parameter scale; the snippet does not disclose training data, license, benchmarks, or deployment conditions.

Why it matters: HKR-H/K/R all pass, but the body only confirms 12B and MoE, with no benchmarks, license, context window, or IDE integration terms. Treat as a mid-weight model release at the lower featured band.

May 31Sunday

r/LocalLLaMA

13 abliterated Gemma 4 E2B variants, 44 GPU hours, benchmark and comparison

Abliterlitics tested 13 abliterated Gemma 4 E2B variants using 44 RTX 5090 GPU hours, and HarmBench ASR rose from the base model’s 32.2% to 82%–100%, while coder3101 scored 84.8% on GSM8K versus the base model’s 83.5%.

Why it matters: HKR-H/K/R all pass, with a named first-person benchmark and concrete numbers. Scope stays narrow around abliterated Gemma 4 E2B variants, so it lands at the featured threshold rather than a must-write item.

May 29Friday

r/LocalLLaMA

Liquid AI releases LFM2.5-8B-A1B

Liquid AI released LFM2.5-8B-A1B with a 128K context window, 38T pre-training tokens, large-scale reinforcement learning, doubled vocabulary for non-Latin tokenization, and availability on Hugging Face.

Why it matters: HKR-H/K/R pass: 8B/A1B, 128K context, and 38T tokens are concrete hooks for local inference. No benchmarks, license, or deployment limits are disclosed, so it stays in the mid featured band.

May 28Thursday

r/LocalLLaMA

Nvidia LocateAnything: Fast Vision-Language Grounding with Parallel Box Decoding

The title says Nvidia LocateAnything-3B performs vision-language grounding with parallel box decoding and runs 10x faster than Qwen3-VL; the post body only provides Hugging Face, GitHub, demo, and project links, and does not disclose benchmark setup or accuracy numbers.

Why it matters: HKR-H/K/R all pass, but the body is mostly links and title-level facts, with no full eval setup or quality metrics. NVIDIA open vision grounding is useful enough for featured, not same-day must-write.

Synced · WeChat

Chinese pretrained embodied model Wall-OSS-0.5 is open sourced

X Square Robot open sourced Wall-OSS-0.5, a VLA model whose 400k pretraining checkpoint scored above 80 on 4 of 17 real-robot zero-shot tasks, with weights, code, training recipe, ablations, and a DMuon optimizer implementation released.

Why it matters: Clear HKR-H/K/R: a 400k checkpoint and 17 real-robot zero-shot tasks add substance, while “post-training not required” is a sharp hook. X Square Robot is not a top foundation-model lab, so this stays at 79.

r/LocalLLaMA

I built a 103B-token Usenet corpus from 1980–2013

OwnerByDane released a 103.1B-token Usenet corpus covering 1980–2013, 408M posts, and 18,347 newsgroups, with free 5K-post-per-hierarchy samples and full-corpus licensing available.

Why it matters: HKR-H/K/R all pass: the zero-contamination corpus has a clear hook, concrete scale, and relevance to training-data scarcity. Score is capped by Reddit-only sourcing, licensed full access, and no third-party validation or benchmark results.

May 27Wednesday

r/LocalLLaMA

I ran 8 open-weight models as agents in a persistent MMO for 10 days

Firespawn Studios ran 25 agents across 8 open-weight models for 10 days in Null Epoch Season 0 and released about 93,000 logged events, with roughly 70% of actions including the model’s reasoning or justification.

Why it matters: HKR-H/K/R all pass: a concrete 10-day MMO agent trial with 25 agents and 93k events. Reddit sourcing limits reach, so it lands in the 78–84 good-quality band, not P1.

AI HOT (Curated Pool)

Reachy Mini enables fully local voice interaction

Reachy Mini implements local voice interaction through the speech-to-speech library, using a cascaded pipeline with a Realtime API-compatible WebSocket interface and default components including Silero VAD, Parakeet-TDT, and Qwen3-TTS.

Why it matters: HKR-H/K/R all pass: the post has a clear local-robot voice hook, concrete stack details, and edge-agent resonance. Scope stays limited to Reachy Mini voice interaction, so it sits at the featured threshold.

AI HOT (Curated Pool)

Shipping a Trillion Parameters With a Hub Bucket: Delta Weight Sync in TRL

Hugging Face merged TRL PR 5417 for delta weight sync, sending only changed weights as sparse safetensors via a Hugging Face Bucket; on Qwen3-0.6B, the per-step payload falls from 1.2GB to 20–35MB.

Why it matters: HKR-H/K/R all pass: TRL gets delta weight sync with a concrete sparse-safetensors mechanism and a 1.2GB to 20–35MB example. Scope is training infra, so it stays below must-write.

r/LocalLLaMA

PrismML Released Binary and Ternary Bonsai Image 4B

PrismML released Binary and Ternary Bonsai Image 4B, 1-bit and ternary text-to-image diffusion transformers around 3GB, compared with FLUX.2 Klein 4B at about 16GB, with browser-local WebGPU demo links and an Apache-2.0 license disclosed in the Reddit snippet.

Why it matters: HKR-H/K/R all pass: low-bit image DiT plus local browser inference is a strong hook, backed by 4B, ~3GB, WebGPU, and license details. Reddit sourcing and limited lab weight keep it in the low featured band.

May 25Monday

AI HOT (Curated Pool)

Harness, Scaffold, and AI Agent Terminology Explained

Hugging Face’s post frames an agent as three layers: Model, Scaffolding, and Harness; Scaffolding defines behavior through prompts and tool descriptions, while Harness runs model calls, tool calls, and control loops.

Why it matters: HKR-H/K/R pass: the Hugging Face post gives a concrete agent-stack taxonomy. It clears featured on practitioner relevance, but lacks a release, benchmark, or deployment case, so it stays at the threshold.

May 23Saturday

r/LocalLLaMA

meituan-longcat/LongCat-Video-Avatar-1.5 on Hugging Face

Meituan LongCat released LongCat-Video-Avatar-1.5 on Hugging Face, supporting AT2V, ATI2V, and video continuation while replacing Wav2Vec2 with Whisper-Large and using DMD2 distillation to reduce inference to 8 NFE; the model weights are released under the MIT License.

Why it matters: HKR-H/K/R all pass: open MIT video-avatar weights plus 8 NFE inference give local multimodal builders real signal. This is a mid-weight open-source model update, not an 85+ same-day industry event.

May 22Friday

AI HOT (Curated Pool)

Text Degeneration: A Production Failure Mode Most Benchmarks Do Not Track

Dharma-AI says in a Hugging Face post that large language models can produce repeated, incoherent, or logically confused text in production, and most mainstream benchmarks do not track this failure mode.

Why it matters: HKR-H/K/R all pass, but the post only discloses the failure pattern and benchmark blind spot, with no sample size, metric, or reproduction setup. This fits the lower featured threshold.

May 21Thursday

r/LocalLLaMA

HRM 1B

Sapientinc released HRM-Text 1B Base and its training code, and the paper claims competitive performance against 2–7B open models while using 100–900x fewer training tokens and 96–432x less estimated compute, with training on 16 H100 GPUs taking about 46 hours and costing about $1,472.

Why it matters: HKR-H/K/R all pass: HRM-Text 1B has concrete low-cost training numbers and released code. Capped at 80 because this is a Reddit item and the efficiency claim still lacks independent evaluation.

May 20Wednesday

Hacker News front page

Show HN: Lance – Image/video generation and understanding in one model

ByteDance released Lance as a research project for image and video generation and understanding in one model; the RSS snippet states 3B active parameters, fewer than 128 GPUs used for training, and links to a homepage, arXiv paper, and Hugging Face model, while the post does not disclose benchmark results or licensing terms.

Why it matters: ByteDance’s Lance puts image/video generation and understanding in one model, with 3B active parameters and <128 GPUs for training. HKR-H/K/R all pass, but benchmarks, license details, and real outputs are not disclosed, keeping it below P1.

May 19Tuesday

AI HOT (Curated Pool)

NVIDIA fine-tunes Cosmos Predict 2.5 with LoRA/DoRA for robot video generation

NVIDIA published a Hugging Face post on fine-tuning Cosmos Predict 2.5 with LoRA and DoRA to generate robot first-person videos from text prompts; the post does not disclose dataset size, training cost, or evaluation results.

Why it matters: HKR-H/K/R pass: the robot POV video angle is clickable, and LoRA/DoRA on Cosmos Predict 2.5 is a concrete mechanism. Missing dataset scale and metrics keep it in the low featured band.

May 18Monday

AI HOT (Curated Pool)

The Open Agent Leaderboard

IBM Research published the Open Agent Leaderboard on Hugging Face to evaluate agents across language understanding, tool use, and multi-step reasoning tasks; the post does not disclose dataset size, model scores, or the evaluation date.

Why it matters: HKR-H and HKR-R pass because an open agent leaderboard speaks to agent-eval pain. HKR-K fails: the article lacks scores, dataset size, and evaluation date, so it sits at the featured threshold.

May 17Sunday

r/LocalLLaMA

85 GPU-hours comparing 5 abliteration methods on Qwen3.6-27B

Abliterlitics compared five Qwen3.6-27B abliteration variants against the base model using 85 GPU-hours of benchmarks, HarmBench, KL divergence, and weight forensics; Huihui had the smallest benchmark deltas, Heretic had the lowest KL divergence, and all five variants reached near-complete safety removal.

Why it matters: HKR-H/K/R all pass: the post gives an 85-GPU-hour comparison across five abliteration methods on Qwen3.6-27B. Niche open-model safety work, not a lab release, so it stays at the featured threshold.