Skip to content

Open source

Open models, frameworks and repositories: open weights, community hits and the balance between open and closed.

Latest picks

141–160 of 329

Jun 19Friday

AI HOT (Curated Pool)

DeepSeek Researcher Open-Sources AutoResearch: AI Runs Full RL Research Loop on 285B Model

DeepSeek researcher Deli Chen open-sourced AutoResearch, a protocol where an AI agent independently ran a full RL research loop on a 285B model—designing experiments, writing code, submitting GPU jobs, debugging, and summarizing results with zero human intervention. The system used GRPO. The post doesn't disclose the specific task, training duration, or success rate, so I'd hold off on getting too excited until there are reproductions.

Why it matters: A DeepSeek researcher released an experimental protocol where an agent independently ran a full RL research loop on a 285B model—strong premise. But the post doesn't disclose the specific task, training duration, or success rate; all key metrics are missing, so the score stays...

AI HOT (Curated Pool)

Steve Yegge: Fable’s shutdown signals frontier AI will be locked down like nukes

Steve Yegge argues Fable’s brief USG shutdown marks the moment model intelligence became dangerous. He predicts frontier models will be controlled like nuclear weapons within 2–3 generations, with most Fortune 500 companies locked out. Open-source can reach Fable-class but won’t blow past it due to compute walls and supply-chain lockdowns. The capability curve will appear flat to most people—not because progress stops, but because the smartest models will be kept out of public hands.

Why it matters: Steve Yegge's deep analysis of the Fable takedown argues the AI capability curve is about to be flattened by government regulation. Sharp thesis with concrete predictions, but it's commentary, not primary reporting — docked for lacking verifiable new facts.

Jun 18Thursday

AI HOT (Curated Pool)

Alibaba open-sources LOGOS, a 1B-param science model that beats Microsoft NatureLM on multiple tasks

Alibaba's ATH-Token Foundry and Renmin University's Gaoling School of AI open-sourced LOGOS, a generative model that uses a unified 'science grammar' to handle seven modalities including proteins, small molecules, and materials. It encodes 3D pocket-ligand contacts as discrete tokens, predicting spatial interactions without explicit 3D coordinates. LOGOS-1B uses only 1/56 the parameters of Microsoft NatureLM (8×7B) and matches or beats domain-specific methods across six science tasks. Pretrained on 44.87B tokens, it shares the same sequence format and next-token prediction objective for both pretraining and downstream tasks, eliminating heavy adaptation. Weights, inference code, and the tech report are fully open on HuggingFace and GitHub.

Why it matters: Alibaba open-sourced LOGOS, a 1B-param scientific model that unifies seven data types into token sequences and beats Microsoft's 56x-larger NatureLM on multiple tasks. Concrete numbers and open code give it strong knowledge value, but the niche domain limits resonance — lands ...

Computing Life · Share · Yage

Vercel open-sources eve: an agent is a directory, built as standalone software

Vercel open-sourced eve under Apache 2.0 at its London Ship conference. The core claim: an agent is a directory. File names auto-register as tools, the Git repo is the agent itself, every instruction change gets a diff and a preview deploy. It ships with durable execution (zero compute during approval waits), sandboxed microVMs, and multi-channel support for Slack, Discord, Teams, and HTTP. This is a different path from LangChain's assemble-it-yourself parts and Claude Managed Agents' cloud-config approach. Eve handles runtime and deployment; it does not write your agent's judgment—instructions.md and skills/ are loading slots, and you bring the content. Multi-platform support is promised but not yet scheduled.

Why it matters: Vercel open-sourced eve under Apache 2.0, with the core claim that an agent is a directory, including durable execution, sandboxed microVMs, and multi-channel support. The article positions eve between LangChain and Claude Managed Agents with concrete mechanism details — not a...

Jun 17Wednesday

AI HOT (Curated Pool)

AWS open-sources Strands Robots SDK: one agent stack from Hugging Face Hub to physical robots

AWS released the Strands Robots SDK under Apache 2.0, wrapping the LeRobot stack into a unified agent. It defaults to MuJoCo simulation with no hardware needed; switch to mode="real" for physical robots. Recorded demos are saved as LeRobotDataset and can be pushed to Hugging Face Hub. Policies like GR00T or LerobotLocal run inference, then broadcast commands to multiple robots over Zenoh mesh. Simulation and hardware code are identical except for one keyword argument. Examples run in a notebook with Python 3.12+ on Linux/macOS, no GPU required.

Why it matters: AWS wraps LeRobot into a unified agent SDK with one-click sim-to-real switching — a solid tool for robotics devs. But pure physical robotics has limited resonance with AI app-layer readers, so R axis isn't fully hit, landing right at the featured threshold.

AI HOT (Curated Pool)

GLM-5.2 open-sourced: tops Code Arena, built for long-horizon tasks

Zhipu released GLM-5.2 under MIT license. It hit #1 among publicly available models on Code Arena, a front-end dev blind benchmark. Coding ability lands between Claude Opus 4.7 and 4.8. The model handles 1M-token lossless context and multi-day tasks. On FrontierSWE it trails Opus 4.8 by only 1%, beating GPT-5.5 and Opus 4.7; on Terminal-Bench 2.1 it's 4% behind Opus 4.8 but up 17.5% over GLM-5.1. A new thinking-budget control lets users dial reasoning depth. IndexShare architecture cuts unit FLOPs to 2.9×, and improved MTP layers boost acceptance length by 20%. It's already adapted to domestic hardware like Huawei Ascend, with the API live under the GLM Coding Plan.

Why it matters: Zhipu open-sourced GLM-5.2 under MIT, with Code Arena frontend blind test ranking #1 among public models, slotting between Claude Opus 4.7 and 4.8, and FrontierSWE only 1% behind Opus 4.8. 1M-token lossless context adds practical weight. Not scoring higher because only the hea...

Hugging Face Blog

Hugging Face launches ARD discovery tool so agents can search for tools, skills, and other agents

Hugging Face released Discover Tool, a reference implementation of the Agentic Resource Discovery (ARD) spec. ARD is an open draft co-developed by Microsoft, Google, GoDaddy, Hugging Face, and others. It lets agents find MCP tools, A2A agents, or skills at runtime via natural-language search instead of hardcoding each one. Hugging Face's implementation wraps the Hub's existing semantic search and Agent Skills into an ARD catalog, exposed as a REST API and an MCP Tool. The post does not disclose pricing, search latency, or accuracy figures.

Why it matters: ARD tackles a real pain point—agent tool discovery—with cross-vendor backing from Microsoft, Google, and Hugging Face, plus a working reference implementation. Not scoring higher because it's still an open draft, not a ratified standard, and the post doesn't spell out adoption...

AI HOT (Curated Pool)

Zhipu releases open-source GLM-5.2, focused on coding and long-horizon tasks

Zhipu released and open-sourced GLM-5.2, scoring 51 on the Artificial Analysis composite leaderboard—top three alongside Anthropic and OpenAI. It ranked first among globally available models in the Code Arena front-end dev blind test. The headline upgrade is solid 1M lossless context for long-horizon tasks: the model handled an 880K-token multi-platform app pipeline in one go and scored only 1% below Claude Opus 4.8 on FrontierSWE. Developers report more stable project-level context and fewer derailments on complex tasks. It runs on domestic hardware including Huawei Ascend and Cambricon, and is released under the MIT license for commercial use.

Why it matters: Zhipu released GLM-5.2 as open-source under MIT license, scoring 51 on Artificial Analysis alongside Anthropic and OpenAI, and #1 on Code Arena for frontend dev. The core upgrade is solid 1M lossless context, with long-horizon benchmarks landing between Claude Opus 4.7 and 4.8...

Jun 16Tuesday

AI HOT (Curated Pool)

Ant Group BaiLing releases Ling & Ring 2.6 tech report, all three models open-sourced

Ant Group BaiLing published full architecture, pretraining, post-training, and agent RL details for Ling-2.6-flash, Ling-2.6-1T, and Ring-2.6-1T. All three use a Hybrid Linear Attention that mixes Lightning Attention and MLA at a 7:1 ratio. Ling-2.6-flash hits 340 tokens/s decoding on 4×H20 hardware. Ling-2.6-1T shows roughly 4× token efficiency gain over its predecessor on the Artificial Analysis Intelligence Index. Ring-2.6-1T high scores 87.60 on PinchBench and 63.82 on ClawEval. Code and weights are open.

Why it matters: Ant Group's BaiLing team open-sourced three models with a Hybrid Linear Attention design blending Lightning Attention and MLA at 7:1, backed by concrete long-context efficiency data. Code and weights are public, making this a verifiable release. Not scoring higher because Ant'...

Jun 15Monday

AI HOT (Curated Pool)

MiniMax open-sources M3 model weights (428B total, 23B active) with lower long-context cost

MiniMax open-sourced M3 model weights last Friday—428B total parameters, 23B active—along with the MSA sparse attention paper that cuts long-context inference cost. M3 is the first open-source model trained with interleaved text and image data from the pre-training stage. Two weeks post-release, it ranked #1 among open-source models on the Artificial Analysis Intelligence Index and GDPval-AA, reached Pareto-optimal on Code Arena WebDev, and topped Chinese models on Vals.AI. Output speed improved from ~30 TPS to ~80 TPS, with another 30–40% planned. A usage dashboard was added to the Token Plan backend.

Why it matters: MiniMax open-sourced a 428B MoE model with interleaved image-text pretraining and two #1 open-source rankings in two weeks — enough signal for featured. Held back from p1 because the post is a first-party announcement without third-party benchmarks or concrete MSA cost numbers...

Jun 13Saturday

AI HOT (Curated Pool)

Zhipu launches GLM-5.2 flagship model with 1M context, open-sourcing next week under MIT license

Zhipu's new flagship GLM-5.2 is live for Coding Plan subscribers, emphasizing coding strength and 1M context. API and chatbot access arrive next week, alongside an MIT-licensed open-source release. The post doesn't disclose benchmark scores or pricing details.

Why it matters: Zhipu drops GLM-5.2 with 1M context and MIT open-source next week, coding-focused. No benchmarks or pricing disclosed, so real capability is unverified — hence below 85. But a domestic flagship update plus open-source is strong signal, worth featuring.

AI HOT (Curated Pool)

Zhipu GLM-5.2 fully released with 1M context window, open-source next week

Zhipu released GLM-5.2, its strongest open-source model yet, available tonight to all GLM Coding Plan users. It supports a genuinely usable 1M context window, leads in long-range tasks, and is called the strongest domestic coding model by Zhipu. API access arrives next week, and the model goes open-source under MIT license next week.

Why it matters: Zhipu rolls out GLM-5.2 to all paid tiers with a 1M context window and a concrete open-source timeline under MIT license. This is a domestic flagship release, scored on par with equivalent US lab launches. The self-claimed strongest coding performance and the open-source date ...

AI HOT (Curated Pool)

MiniMax open-sources M3 weights, takes a swipe at Anthropic's export control ban

MiniMax released M3 model weights on HuggingFace. The post says 'M3 would never,' a jab at Anthropic's Fable 5 and Mythos 5 being forcibly disabled under US export controls, blocking all foreign nationals. The post doesn't disclose M3's parameter count, benchmarks, or license.

Why it matters: MiniMax open-sources M3 weights as a direct response to Anthropic's export controls — strong conflict and topicality, but the post lacks parameter count, benchmarks, and license details, capping the score.

Jun 12Friday

r/LocalLLaMA

MiniMax-M3 open-sourced: a 428B MoE model with 23B activated parameters

MiniMax released MiniMax-M3 weights on Hugging Face. It's a mixture-of-experts model with ~428B total parameters and ~23B activated per inference. The post doesn't disclose training data, benchmarks, or minimum VRAM for local runs.

Why it matters: MiniMax dropped full weights for M3 on Hugging Face — a 428B MoE model activating only 23B per forward pass, putting it in the top tier of open-weight efficiency plays. No benchmarks or hardware requirements disclosed yet, which caps the score, but the weight release alone is ...

AI HOT (Curated Pool)

Kimi releases and open-sources Kimi-K2.7-Code

Kimi open-sourced K2.7-Code, scoring 11%–31.5% higher than K2.6 on three in-house benchmarks. Inference token usage dropped 30%, and long-coding-task instruction-following and end-to-end success rate both improved. A 6x speed mode is coming; the model is available now via Kimi API and Kimi Code. The post doesn't disclose parameter count, training data, or the open-source license.

Why it matters: Moonshot open-sourced a code model with solid gains on three in-house benchmarks and a 30% inference efficiency improvement — a real cost signal. No external benchmarks (LiveCodeBench, SWE-bench) or parameter count disclosed, so capped below 85. Still, a major Chinese lab open...

r/LocalLLaMA

Huawei launches openPangu 2.0, open-sourcing June 30; Pro version has 505B total params but only 18B active

Huawei announced openPangu 2.0 at HDC 2026. Two sparse models: Pro at 505B total / 18B active, Flash at 92B total / 6B active, hitting a 28:1 sparsity ratio. 512K context window, heavily optimized for Ascend chips with claimed 2x single-card throughput vs mainstream open-source models. Richard Yu said the large total param count reflects limited compute left for Huawei after supporting other Chinese enterprises, so the focus is on latency and throughput gains. Open-sourcing starts June 30, covering weights, inference code, training code, and training operators. I'd hold off until we see actual benchmarks—the post only gives relative improvement percentages, no absolute scores.

Why it matters: Huawei announced openPangu 2.0 at HDC: two sparse variants, Pro 505B/18B active and Flash 92B/6B active, 512K context, open-sourcing June 30. The 28:1 sparsity ratio is a technical hook, and the 2x Ascend throughput claim needs independent verification. Score stays below 80 be...

Ruan YiFeng's Weblog

rsync maintainer's use of Claude to write code sparks heated community debate

rsync v3.4.3 was found to be generated by Claude, raising community concerns about vulnerabilities. Maintainer Andrew Tridgell responded that AI-driven attacks are coming, and he lacks the energy to patch AI-discovered bugs manually, so he shifted to 'AI writes code, humans write tests.' The thread has over 300 comments, mostly critical.

Why it matters: rsync's maintainer openly admitted using Claude to write code and proposed a 'humans write tests, AI writes implementation' model — this isn't a routine product update but a public clash over open-source maintenance methodology. The 300+ comment thread is itself a signal. Not ...

AI HOT (Curated Pool)

Hugging Face open-sourced Open-R1, a full reproduction of DeepSeek-R1

Hugging Face published Open-R1 on GitHub, aiming to fully reproduce the DeepSeek-R1 reasoning model. The repo has 26.1k stars and 2.4k forks so far. The body only contains the repo's landing page navigation and metadata; it does not disclose the implementation plan, training data, reproduction progress, or benchmark results. I'd treat this as a public reproduction scaffold and collaboration hub for now, and wait for a technical report before judging fidelity.

Why it matters: Hugging Face launched a full open-source reproduction of DeepSeek-R1, with the repo already at 26.1k stars — strong community interest. But the body only contains project scaffolding and navigation; no implementation plan, training data, or reproduction progress is disclosed y...

Jun 11Thursday

Synced · WeChat

Google open-sources 26B text-diffusion MoE; Pichai: generation speed like a racehorse

Google open-sourced DiffusionGemma, a 26B MoE model that activates only 3.8B parameters at inference. Instead of generating tokens one by one, it drafts 256-token blocks in parallel, hitting 1,000+ tokens/sec on an H100—up to 4× faster than autoregressive models. Output quality is lower than standard Gemma 4, so Google still recommends the autoregressive version for production. It ships under Apache 2.0, fits quantized on consumer GPUs with 18GB VRAM, and targets latency-sensitive nonlinear tasks like inline editing and code completion.

Why it matters: Google open-sourced a 26B text diffusion model that skips autoregressive decoding, activating only 3.8B params at inference and hitting 1,000+ tok/s on a single H100. Apache 2.0, with concrete speed comparisons and mechanism details — directly useful for inference folks. Not s...

QbitAI · WeChat

Google releases DiffusionGemma, a diffusion-based text model that generates 4× faster than autoregressive models

Google open-sourced DiffusionGemma, a 26B MoE diffusion text model that activates only 3.8B parameters at inference and fits in 18GB VRAM after quantization. It denoises 256 tokens in parallel—like a printing press instead of a typewriter—hitting 1,000+ tokens/s on an H100 and 700+ on an RTX 5090, roughly 4× faster than a comparable autoregressive model. Bidirectional attention enables real-time self-correction; after fine-tuning, Sudoku accuracy jumped from 0% to 80%. Quality still trails Gemma 4, and Google positions it as an experimental “racehorse” for speed-sensitive local use. Released under Apache 2.0, weights available on Hugging Face.

Why it matters: Google open-sourced DiffusionGemma, applying diffusion models to text generation with 256 tokens denoised simultaneously, roughly 4x faster than comparable autoregressive models. Score isn't higher because only speed numbers are out—generation quality and downstream task perfo...