Skip to content

#开源/仓库

3 today

Jun 3Wednesday

AI HOT (Curated Pool)

Google DeepMind open-sources a toolkit for scientific agents

Google DeepMind released Science Skills on GitHub for scientific-discovery agent workflows; the post does not disclose the license, benchmark results, or numeric token-efficiency gains.

Why it matters: Passes HKR-H/K/R: DeepMind, open source, and science agents make it relevant. Missing license, benchmarks, and efficiency data keep it in the 78–84 band, not P1.

Jun 2Tuesday

AI HOT (Curated Pool)

To Avoid Paying $120, I Turned a Computer Cleaner into an Open-Source Skill

The author open-sourced a cross-platform AI cleaning skill for Mac and Windows, generating interactive HTML reports from file scans; in a test, it freed nearly 120GB, compared with CleanMyMac identifying 15.8GB.

Why it matters: This is not a platform-level release, but HKR-H/K/R all land through the $120 replacement hook, concrete scan/report mechanism, and 120GB test result. It fits the practical open-source tool band near the featured threshold.

Xinzhiyuan · WeChat

CAS Opens MobileGym, a Browser-Based Agent Training Environment for Mobile Apps

CASIA released MobileGym, a browser-based Android simulation environment covering 28 apps, with about 400MB per instance, 3-second cold start, JSON state snapshots, and programmatic task verification for mobile-agent training and evaluation.

Why it matters: MobileGym is practical open-source infrastructure for agent training and evaluation, with enough concrete numbers and mechanisms to pass HKR-H/K/R. It fits the 78–84 quality band, below major lab model-release weight.

Jun 1Monday

AI HOT (Curated Pool)

MiniMax Releases Open-Source M3 with Coding, Long-Context, and Multimodal Capabilities

MiniMax released the open-source M3 model with coding, a 1M-token context window, and native multimodal support; M3 scores 59.0% on SWE-Bench Pro, 83.5% on BrowseComp, and costs about one-twelfth per token versus GPT-5.5.

Why it matters: HKR-H/K/R all pass: M3 has open source, 1M context, multimodal support, and 59.0% on SWE-Bench Pro. A single X post without official docs or third-party tests keeps it in the 78–84 band.

AI HOT (Curated Pool)

OpenBMB Releases Two UltraData Open Datasets, Tops HuggingFace Trending

OpenBMB, Tsinghua NLP, and Modelbest released two UltraData open datasets: Ultra-FineWeb-L3 contains 600B+ tokens, including 400B+ English and 200B+ Chinese tokens, while UltraData-SFT-2605 contains 15M+ SFT samples with thinking and non-thinking labels.

Why it matters: HKR-H/K/R pass: two open datasets, 600B+ tokens, and 15M+ SFT samples are concrete practitioner signal. Single-source release with no evals or license detail keeps it at the lower featured band.

AI HOT (Curated Pool)

NVIDIA Open-Sources Cosmos 3, Its First Generalist Model for Physical AI

NVIDIA open-sourced Cosmos 3 at GTC Taipei, releasing two variants, Super 32B and Nano 8B, with model weights, code, and datasets made available.

Why it matters: HKR-H/K/R all pass: the concrete hook is NVIDIA opening Cosmos 3 with 32B/8B variants and released artifacts. The post is sparse and single-source, with no benchmarks or license details, so it stays in the 78–84 band.

r/LocalLLaMA

I ported NVIDIA Parakeet speech-to-text to ggml: same output as NeMo, faster, GGUF-quantized, no Python

mudler_it ported NVIDIA Parakeet speech-to-text models to C++/ggml with no Python or PyTorch, reporting byte-for-byte NeMo parity on f32/f16, up to about 5x GPU speedups on larger TDT and hybrid models, and GGUF quantization across f16, q8_0, q6_k, q5_k, and q4_k.

Why it matters: HKR-H/K/R all pass: the port has a concrete local-inference hook, byte-parity and speed claims, and clear practitioner resonance. Source scope keeps it at the low featured band, not P1.

May 31Sunday

Synced · WeChat

Microsoft open-sources SkillOpt for training Agent skill documents, reaching 3.3k stars in a week

Microsoft open-sourced SkillOpt, a text-space optimization framework that trains Agent skill documents without changing model weights; the paper reports best or tied-best results across 52 combinations covering 7 target models, 6 benchmarks, and 3 execution environments.

Why it matters: Microsoft’s open-source SkillOpt is a strong Agent tooling and research release. HKR-H has the 3.3k-star/trainable-skill hook, HKR-K has the text-parameter mechanism and 52 eval setups, and HKR-R hits agent engineering pain, so it lands in featured at 82.

May 30Saturday

QbitAI · WeChat

RUC and Zhizhi Institute Open-Source Claw Agent Data, Training, and Evaluation Pipeline

Renmin University of China and Zhizhi Institute open-sourced ClawGym, a Claw Agent framework with 13.5K synthetic executable tasks, 200 benchmark tasks, model checkpoints, training data, and training code; ClawGym-30B-A3B scores 56.82 on ClawGym-Bench and exceeds Qwen3-235B-A23B in the reported evaluation.

Why it matters: HKR-H/K/R all pass: ClawGym bundles data, code, checkpoints, and eval tasks rather than just a leaderboard. Its impact is developer-facing, below a major lab model release or market-moving event.

May 29Friday

AI HOT (Curated Pool)

Xiaomi Open-Sources Controllable Video Foley Model ControlFoley

Xiaomi’s large model application team open-sourced ControlFoley, a controllable video Foley model supporting three tasks: text-guided video dubbing, text-controlled video dubbing, and reference-audio-controlled video dubbing, with code, model weights, and an online demo released.

Why it matters: ControlFoley clears HKR-H/K/R with controllable video Foley plus code, weights, and demo. It is a useful multimodal-audio release from Xiaomi, but not a flagship foundation-model launch, so it sits near the featured threshold.

May 28Thursday

Xinzhiyuan · WeChat

Tsinghua Team Open-Sources PilotDeck Agent System, Claims 70% Token Cost Reduction

Tsinghua THUNLP, ModelBest, OpenBMB, and AI9stars open-sourced PilotDeck; the article says its sub-agent routing reduced cost from $12.58 to $2.83 in a Xiaohongshu content-generation test, while preserving separate WorkSpaces, editable memory, and per-session routing logs.

Why it matters: HKR-H/K/R all pass: PilotDeck has a clear agent-cost hook, a concrete routing mechanism, and $12.58 to $2.83 data. It stays in the 78–84 band because this is a tool release, not a major model or platform launch.

Synced · WeChat

Chinese pretrained embodied model Wall-OSS-0.5 is open sourced

X Square Robot open sourced Wall-OSS-0.5, a VLA model whose 400k pretraining checkpoint scored above 80 on 4 of 17 real-robot zero-shot tasks, with weights, code, training recipe, ablations, and a DMuon optimizer implementation released.

Why it matters: Clear HKR-H/K/R: a 400k checkpoint and 17 real-robot zero-shot tasks add substance, while “post-training not required” is a sharp hook. X Square Robot is not a top foundation-model lab, so this stays at 79.

r/LocalLLaMA

I built a 103B-token Usenet corpus from 1980–2013

OwnerByDane released a 103.1B-token Usenet corpus covering 1980–2013, 408M posts, and 18,347 newsgroups, with free 5K-post-per-hierarchy samples and full-corpus licensing available.

Why it matters: HKR-H/K/R all pass: the zero-contamination corpus has a clear hook, concrete scale, and relevance to training-data scarcity. Score is capped by Reddit-only sourcing, licensed full access, and no third-party validation or benchmark results.

AI HOT (Curated Pool)

Open-source FastVideo Dreamverse real-time video generation tool

Hao AI Lab open-sourced FastVideo Dreamverse, a real-time video generation tool that generates a 30-second 1080p video in 7 seconds under the stated setup of one NVIDIA B200 GPU and LTX-2.

Why it matters: HKR-H/K/R all pass: the 7s-for-30s-1080p claim is concrete and practitioner-relevant. Single-source X sourcing and missing independent benchmarks keep it in the 78–84 band.

May 27Wednesday

AI HOT (Curated Pool)

Perplexity open-sources Unigram tokenizer to reduce CPU usage

Perplexity open-sourced a rebuilt Unigram tokenizer that reduces CPU usage by 5-6x, targeting tokenization latency when small rerankers and embedding models run on GPUs in single-digit milliseconds.

Why it matters: HKR-H/K/R all pass: the 5-6x CPU claim and tokenizer bottleneck are concrete for production RAG/search teams. It stays in the featured-threshold band because the post lacks independent benchmarks, repo details, and deployment scale.

AI HOT (Curated Pool)

AI Builds AI: ModelBest Open-Sources ForgeTrain, a Training Framework Written by AI

ModelBest, Tsinghua University, and OpenBMB open-sourced ForgeTrain, described as the first production-grade LLM training framework written entirely by AI with zero human code, and ModelBest used it to pretrain MiniCPM5-1B on Huawei Ascend chips.

Why it matters: HKR-H/K/R all pass: an open-source training framework, AI-written code, and MiniCPM5-1B pretraining on Ascend give concrete hooks. This is a strong tooling story, not a top-model launch, so 80 fits featured rather than P1.

May 26Tuesday

AI HOT (Curated Pool)

SenseNova-U1 full training code open-sourced for multimodal multitask training

OpenSenseNova released the full SenseNova-U1 training code on GitHub under Apache-2.0, supporting an 8B dense model, an A3B MoE architecture, and multimodal tasks such as text-to-image generation, image editing, interleaved generation, and text-visual understanding.

Why it matters: HKR-H/K/R all pass, but the source is a short official post with no dataset, training budget, or eval results disclosed. The practical value of full training code puts it in the featured band.

r/LocalLLaMA

[OSS] dlmserve: First Serving Engine for Diffusion Language Models

dlmserve released an MIT-licensed serving engine for diffusion language models, with LLaDA-8B-Instruct support and 2.5x HF throughput at batch=4. It exposes an OpenAI-compatible /v1/chat/completions API, batches at the denoising-step level, runs in 12GB VRAM, and adds about 1.8x throughput with optional LocalLeap acceleration.

Why it matters: HKR-H/K/R all pass: an open-source DLM serving engine with concrete throughput and VRAM claims. Single Reddit source and an early ecosystem keep it in low featured, not 78+.

AI HOT (Curated Pool)

ModelBest open-sources MiniCPM5-1B, topping sub-2B models on AA-Index

ModelBest open-sourced MiniCPM5-1B, a 1B-parameter edge language model that beats all sub-2B models on AA-Index, uses a 0.5GB weight file after INT4 quantization, and runs on phones and browsers.

Why it matters: HKR-H/K/R all pass: MiniCPM5-1B has concrete params, quantized size, and edge runtime claims. It is still a small-model release, below flagship-model impact.

r/LocalLLaMA

Shard - Getting to 10× KV Cache Compression

Shard reduces Llama-3.1-8B KV memory by about 10× at 8K context and 11× at 32K, with no measured drop on NIAH or LongBench, using PCA plus int4 quantization for K and Hadamard rotation plus vector quantization for V.

Why it matters: HKR-H/K/R all pass: the 10× KV-cache claim has a strong hook and concrete model/context/benchmark details. Reddit-only sourcing and limited validation keep it in the 78–84 band.

May 25Monday

r/LocalLLaMA

Computer-use sandbox framework for Codex on headless Linux

superSmitty9999 released ai-sandbox-manager as a PoC that uses LXC templates to give Codex sudo access, browser use, Docker, and shared GPU access, with a hook that blocks git push while the agent works inside isolated copies.

Why it matters: HKR-H/K/R all pass, but this is a Reddit personal PoC with mechanisms only, not adoption, benchmarks or maturity evidence. It fits the featured floor for practical agent-sandbox work.

Synced · WeChat

A 1B-Gaussian 3D World Runs in the Browser, Outperforming Fei-Fei Li’s Spark

Manycore Tech open-sourced Aholo Viewer, a browser-based 3D Gaussian Splatting viewer that used half the memory of Spark 2.0 in a 300M-Gaussian test, loaded 2x faster, rendered 3x faster, and supports scenes with up to 1B Gaussian points.

Why it matters: HKR-H/K/R all pass: the hook is vivid, the post gives 300M-point benchmarks and a 1B-point ceiling, and browser-side 3D deployment matters to practitioners. Score stays at 80 because this is a strong tool release, not a foundation-model event.

May 23Saturday

r/LocalLLaMA

club-rdna16: Practical 16GB AMD/Radeon local LLM testing repo

club-rdna16 publishes a practical 16GB Radeon local LLM testing repo, with an RX 6900 XT running llama.cpp on ROCm/HIP and Qwen3.6 35B-A3B reaching a stable 131k context using q8 KV cache.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post and the body only discloses test conditions, not speed, VRAM curves, or reproducible logs. It clears the featured floor as a practical local-LLM repo.

r/LocalLLaMA

Experts first llama.cpp

comanderxv published a llama.cpp fork that caches MoE experts in 12GB VRAM; on an RTX 2060 with Qwen3.6-35B-A3B, throughput rose from 19/22 tk/s to 26 tk/s at about a 62% expert-cache hit rate.

Why it matters: HKR-H/K/R all pass: the hook is a 35B MoE speedup on a 12GB RTX 2060, with concrete caching and hit-rate data. Scope stays niche to local inference, so it lands at the featured threshold rather than must-write.

May 22Friday

AI HOT (Curated Pool)

NetEase Youdao Open-Sources Ziyue 4 Multimodal and Text-to-Speech Models

NetEase Youdao open-sourced its Ziyue 4.0 multimodal and text-to-speech models, with the 27B multimodal model reporting 81.4% accuracy on Chinese math reasoning tasks and the speech model supporting 14 languages.

Why it matters: HKR-H/K/R pass: the story has a concrete open-source hook, specific model numbers, and practitioner relevance. NetEase Youdao is not a frontier lab, so it stays below the 78+ good-quality band.

May 21Thursday

AI HOT (Curated Pool)

Tencent open-sources Hy-MT2 multilingual translation model

Tencent open-sourced the Hy-MT2 multilingual translation model with support for translation across 33 languages; its 1.8B version uses AngelSlim 1.25-bit quantization, occupies 440 MB of storage, and runs locally on mainstream mobile chipsets.

Why it matters: HKR-H/K/R all pass: Tencent gives a specific edge-AI hook with 33 languages, 1.25-bit quantization, and a 440MB phone-local build. Benchmarks, latency, and license terms are not disclosed, so it stays below major flagship releases.

May 20Wednesday

r/LocalLLaMA

Running DeepSeek-V4 locally on 4 legacy RTX 2080 Ti GPUs with W8A8 at 255 prefill tok/s

A Reddit user ran DeepSeek-V4-Flash locally on 4 RTX 2080 Ti GPUs, reporting 284B total parameters, 13B active parameters, a sub-$2,500 build, custom Turing CUDA kernels, W8A8 quantization, 1TB DDR4 ECC RAM, and about 255 prefill tokens/s.

Why it matters: HKR-H/K/R all pass: this is a numeric first-person local-inference experiment. Single-source Reddit provenance and custom Turing kernels keep it in the lower featured band.

r/LocalLLaMA

Public repository Codegraph claims 94% fewer Claude, Cursor, Codex, and OpenCode tool calls locally

Codegraph uses a pre-indexed knowledge graph for symbol relationships, call graphs, and code structure. In the VS Code test, it reduced tool calls from 52 to 3 and runtime from 1m37s to 17s.

Why it matters: All HKR axes pass, but evidence is a Reddit/public-repo self-test without independent replication. The 94% reduction and 52→3 call count clear featured, not p1.

r/LocalLLaMA

Floor for local meeting summarization on a 6GB GPU: Qwen3.5 0.8B works in 57s, Granite 4 350M hallucinates

The author tested VoiceFlow 1.6.0 on an RTX 3060 Laptop 6GB, where Qwen3.5 0.8B summarized a 4-minute meeting in 57 seconds with 16K context, while Granite 4 350M returned summaries in 0.6-2.8 seconds but fabricated Binance and Star Trek content.

Why it matters: HKR-H/K/R all pass: the hook is concrete, the test reports hardware/context/timing, and local meeting summarization hits privacy and cost nerves. Single Reddit experiment limits authority, so 73 featured.

AI HOT (Curated Pool)

NVIDIA open-sources first 4-bit infrastructure for ultra-long video generation

NVIDIA researchers open-sourced LongLive 2.0, an end-to-end long-video generation infrastructure covering training and inference with 4-bit quantization, FP4 quantization, parallel acceleration, KV-cache optimization, and 45.7 FPS generation on a 5B model.

Why it matters: HKR-H/K/R all pass: NVIDIA researcher open-sources LongLive 2.0 with 4-bit long-video train/inference and 45.7 FPS on a 5B model. This is strong open-source infra, not a flagship model launch, so it fits the 78–84 band.

May 19Tuesday

Hacker News front page

Show HN: Forge takes an 8B model from 53% to 99% on agentic tasks

Forge adds five guardrail layers to self-hosted LLM tool calling, raising Ministral 8B to 99.3% across 18 multi-step agentic scenarios, with the accepted ACM CAIS ’26 paper covering 97 model/backend configurations and 50 runs per scenario.

Why it matters: HKR-H/K/R all pass: the 53%→99.3% jump is clickable, the test setup has concrete numbers, and self-hosted agent reliability is a live practitioner pain. Single-source Show HN/GitHub evidence keeps it in the 78–84 open-source-tool band, not P1.

r/LocalLLaMA

ByteDance released an open-source model that attempts broad multimodal tasks with 3B parameters

ByteDance released Lance, an open-source unified multimodal model with 3B active parameters that supports image and video understanding, generation, and editing, and the post says it was trained from scratch with a staged multi-task recipe under a 128-A100-GPU budget.

Why it matters: HKR-H/K/R all pass: ByteDance’s open Lance has a compact multimodal hook, concrete 3B/128-A100 facts, and clear cost/deployment resonance. Reddit-sourced details lack benchmarks, license terms, and official context, so it stays featured, not P1.

AI HOT (Curated Pool)

Horizon Open-Sources 400M-Parameter Robot Control Model HoloMotion-1

Horizon Robotics Lab open-sourced HoloMotion-1, a 400M-parameter full-body humanoid control model that uses MoE sparse activation and KV-cache inference to reach about 300 FPS on-device, with code and a technical report released.

Why it matters: HKR-H/K/R all pass: HoloMotion-1 has an open-source robotics hook plus 400M params and about 300FPS edge inference. Its reach is narrower than a frontier model release, so it fits the 78 featured band.

May 18Monday

Hacker News front page

Show HN: InsForge – Open-source Heroku for coding agents

InsForge released an Apache 2.0 backend platform that lets coding agents deploy, operate, and debug backend systems through one CLI install command and Skills.

Why it matters: HKR-H/K/R all pass: the Heroku-for-agents framing, Apache 2.0 plus one-CLI install, and agent ops pain are concrete. Source is mainly Show HN/GitHub with no usage, benchmark, or production proof, so it sits at the featured threshold.

QbitAI · WeChat

openJiuwen open-sources JiuwenSwarm, a multi-agent swarm coordination framework

openJiuwen released and open-sourced JiuwenSwarm with four components: Agent Swarm, Swarm Skills, Skills Hub, and self-evolution, and the framework supports HOTS and HITS modes for human participation in multi-agent workflows.

Why it matters: HKR-H/K/R pass: the swarm angle is clickable, the post gives four modules plus HOTS/HITS, and agent builders care about orchestration choices. Lacking benchmarks or adoption data keeps it at the featured threshold.

r/LocalLLaMA

I built a coding agent that gets 87% on benchmarks with a 4B parameter model

SmallCode passes 87 of 100 benchmark tasks with Gemma 4 activating 4B parameters per token. The author attributes the result to compound tools, compile and lint feedback, task decomposition after two repeated failures, and optional escalation to Claude or OpenAI for one task.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post and the benchmark identity plus replication details are incomplete. It fits a concrete first-person experiment above the featured bar, not the 78+ band.

AI HOT (Curated Pool)

Open-source tool exposes security risks and detection gaps in AI API relays

api-relay-audit audits AI API relay risks with verifiable three-state decisions and transparent logs, covering AC-1 tool-call rewriting, AC-2 error-response leakage, and context truncation, while the author has published the methodology, comparison results, quick-reference table, and the open-source tool.

Why it matters: HKR-H/K/R all pass because the tool targets real AI API relay risks with concrete checks. Source is a single X post, and adoption or incident data is not disclosed, so it stays in the low featured band.

May 17Sunday

Hacker News front page

Show HN: Semble – Code search for agents that uses 98% fewer tokens than grep

MinishLab open-sourced Semble, a code-search tool for agents that combines Model2Vec embeddings, BM25, RRF fusion, and reranking; on a 63-repo benchmark, it used 98% fewer tokens than grep+read, reached 0.854 NDCG@10, and ran CPU queries in about 1.5 ms.

Why it matters: HKR-H/K/R all pass: the 98% token claim is clickworthy, the 63-repo benchmark adds substance, and coding-agent context cost is a real practitioner nerve. Impact is still toolchain-level, so it stays below must-write.

AI HOT (Curated Pool)

Latest Open Artifacts #21: Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1, and More

Open AI model teams released Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1, and other versions this month, and the post says they were tested under CAISI’s V4 evaluation framework, but the RSS snippet does not disclose scores.

Why it matters: HKR-H/K/R all pass: a dense open-model roster, a named CAISI V4 evaluation frame, and clear practitioner relevance for model choice. Missing scores and reproducible detail keep it in the 78–84 band.

AI HOT (Curated Pool)

Ring-2.6-1T Open-Sourced and Listed on OpenRouter for Agent Workflows

AntLingAGI open-sourced Ring-2.6-1T and listed it on OpenRouter with a 75% discount through the end of May; the trillion-scale reasoning model targets agent workflows, including planning, tool use, context maintenance, and complex task execution, using Async RL and IcePop training methods.

Why it matters: HKR-H/K/R all pass: a 1T open agent model is clickable, with OpenRouter access, discount, and training methods disclosed. Score stays at 74 because benchmarks, license, and context window are not given.