Skip to content

Open source

Open models, frameworks and repositories: open weights, community hits and the balance between open and closed.

Latest picks

121–140 of 329

Jul 8Wednesday

AI HOT (Curated Pool)

Liquid AI open-sources Antidoom, a final-token preference optimization method that fixes reasoning model doom loops

Reasoning models can get stuck in doom loops, repeating useless tokens until the context window fills up. Liquid AI open-sourced Antidoom, which uses Final Token Preference Optimization (FTPO) to fix this. The method trains the model on 1,040 preference pairs to learn when to stop at the end of reasoning. On DeepSeek V4 Pro, the doom-loop rate dropped from 3.2% to 0.3% without hurting math or coding scores. The post doesn't disclose training cost or how well it transfers to non-DeepSeek models.

Why it matters: Liquid AI open-sourced a practical fix for reasoning model doom loops, dropping the rate from 3.2% to 0.3% on DeepSeek V4 Pro — solid numbers. Not scoring higher because it's a single blog post with no paper or third-party validation yet; 78 for a strong single-source piece.

Jul 5Sunday

AI HOT (Curated Pool)

Meituan LongCat-2.0 fully open-sourced under MIT license, releasing 1.6T MoE weights and inference code

Meituan fully open-sourced LongCat-2.0 under MIT license, releasing both weights and inference code. It's a 1.6T-parameter MoE model activating ~48B per token, with 1M-token context. LongCat Sparse Attention handles long sequences, Zero-Compute Experts dynamically activate 33B–56B to avoid wasted compute, and MOPD routes tasks across Agent, Reasoning, and Interaction expert groups. On benchmarks: SWE-bench Pro hits 59.5, edging out GPT-5.5's 58.6; Terminal-Bench 2.1 scores 70.8; multilingual SWE-bench reaches 77.3. It natively integrates with Claude Code, OpenClaw, and Hermes Agent, supports GPU and NPU deployment, and has been validated on large-scale domestic clusters.

Why it matters: Meituan fully open-sources LongCat-2.0, a 1.6T MoE model, under MIT license with weights and inference code — a rare move from a major Chinese tech company. The 1M-token context window and sparse attention design are concrete technical hooks, not just marketing. Score held at ...

Jul 2Thursday

Hacker News front page

git-annex maintainer spent 100 hours removing LLM-generated code from dependencies

Joey Hess audited git-annex's entire dependency tree to exclude LLM-generated code. He found an incoherent 1,489-line commit message with 10,000 lines of changes, and an LLM prompt that copied code from another project—avoiding infringement only by luck. Hess says the only upside of this 100-hour effort is better dependency quality data for future decisions. He notes the Software Freedom Conservancy has already backed off on this issue, and he is reconsidering his own participation in these communities.

Why it matters: Joey Hess personally spent 100 hours auditing git-annex's dependency tree for AI-generated code, surfaced two concrete horror stories, and noted SFC already punted. HKR all hit, but this is a personal practice report, not an industry-level event — 78 featured.

AI HOT (Curated Pool)

Tencent Hy3 released: matches much larger flagship models with 2-5x fewer parameters

Tencent officially released Hy3 under Apache 2.0. With 1/5 to 1/2 the parameters of competing flagship models, Hy3 matches or beats them on reasoning, agent, and long-context benchmarks. In a 270-person internal blind test, Hy3 scored 2.67/4 vs GLM5.1's 2.51/4. Hallucination rate dropped from 12.5% to 5.4%, multi-turn error rate from 17.4% to 7.9%. WorkBuddy task completion jumped from 72% to 90%, with 34% less time. API pricing: ¥1/M input tokens, ¥4/M output, ¥0.25 cached. The post does not disclose exact parameter count or training details.

Why it matters: Tencent Hunyuan releases Hy3, open-source under Apache 2.0, with parameter counts 1/5 to 1/2 of competitors yet matching or beating them on reasoning, agent, and long-context benchmarks. A 270-person blind test shows it beating GLM5.1. Domestic flagship open-source release is ...

Computing Life · Share · Yage

Fable 5's 18-day ban: Anthropic's share went to GLM

Anthropic's Fable 5 was taken offline by US export controls three days after launch, for 18 days. OpenRouter daily token data shows total volume grew from 24T to 32T, but Anthropic's share dropped from 20.7% to 17.6% and its absolute volume shrank. GLM was the biggest winner, jumping from 1.8% to 7.4% share. GLM-5.1 saw a spike on the day GLM-5.2 launched, then collapsed 48 hours later as GLM-5.2 took over. Community sentiment shifted from sympathy for Anthropic to mocking its business strategy. The post notes this only captures API-layer data, not first-party subscriptions, and token volume comparisons overstate GLM's share due to a 10x price gap.

Why it matters: A data-driven attribution of the 18-day Fable 5 ban's competitive impact using OpenRouter daily token data: Anthropic lost 3.1pp share, GLM gained 6.7pp, plus a weird anomaly where GLM-5.1 traffic spiked 4-5x after GLM-5.2 launched. Counterintuitive findings that directly touc...

Jul 1Wednesday

Hacker News front page

Open-source game engine Godot will no longer accept AI-authored code contributions

Godot maintainers will reject AI-authored code contributions, arguing that heavy AI users often don't understand their own code well enough to fix it. The project worries that AI-generated patches look correct but hide bugs, undermining long-term maintenance. The post doesn't specify the effective date or which detection tools are used.

Why it matters: Godot is a major open-source project in the game engine space. Its maintainers publicly rejecting AI-authored code contributions, with a concrete reason (contributors don't understand their own code and can't fix bugs), is directly relevant to the AI-assisted coding debate. Sc...

AI HOT (Curated Pool)

Meituan releases LongCat-2.0: a 1.6T-parameter model trained on 50,000 domestic GPUs, now open source

Meituan open-sourced LongCat-2.0, a 1.6T total-parameter model with ~48B activated per inference and native 1M context. It was trained and served entirely on a 50,000-card domestic GPU cluster. The architecture combines LSA sparse attention, zero-compute experts, ScMoE, and MOPD multi-expert fusion that blends Agent, Reasoning, and Interaction expert groups. SWE-bench Pro hits 59.5, Multilingual 77.3. A preview is live on OpenRouter and longcat.ai, already ranking top three globally in monthly calls on OpenRouter. The post doesn't disclose training cost, inference latency, or the specific domestic chip model, so I'd hold off on those details.

Why it matters: Meituan's trillion-param model trained end-to-end on domestic GPUs is the headline; code benchmark scores are solid. Not scoring higher because Meituan isn't a tier-1 model lab yet, and real-world usability depends on the open-source release.

Jun 30Tuesday

Hacker News front page

Meituan open-sources LongCat-2.0, a 1.6T MoE model with 48B active params, trained entirely on AI ASIC superpods

Meituan open-sourced LongCat-2.0, a 1.6-trillion-parameter MoE model with ~48B active parameters per token. It was pretrained on over 35 trillion tokens using 50K+ in-house AI ASICs with no rollbacks or irrecoverable loss spikes, showing frontier-scale training is viable on non-GPU hardware. The model targets long-context and agentic workloads: it introduces LongCat Sparse Attention to speed up 1M-token processing and was trained on hundreds of billions of 1M-context tokens. Official charts place it alongside Gemini 3.1 Pro, GPT-5.5, and Opus 4.8 on Terminal-Bench 2.1, SWE-bench Pro, and other coding/agent benchmarks, though the post does not provide exact numeric comparisons. An N-gram Embedding module with 135B parameters expands the embedding space roughly 100×, which the team claims outperforms scaling standard MoE experts by the same amount. The model is integrated with Claude Code, OpenClaw, and Hermes; code and weights are available on GitHub and HuggingFace.

Why it matters: Meituan open-sources a 1.6T MoE model trained entirely on in-house AI ASICs across 50k+ cards with zero rollbacks, plus dedicated long-context and agent optimizations. Score held at 82 rather than higher because we only have the official blog post — no third-party evals or rea...

Jun 26Friday

New York Times Chinese

Chinese AI Models Narrow Performance Gap with Anthropic and OpenAI

Zhipu's GLM-5.2 surged in popularity after Anthropic restricted access to Fable and Mythos, entering OpenRouter's top ten. It costs about one-eighth of Claude Opus 4.8 for certain tasks and is fully open-source. Experts estimate China's lag behind US firms has shrunk to six months or less. The post notes Zhipu's compute spending exceeded 7x its revenue in H1 2025, but does not disclose whether GLM-5.2's training involved distillation.

Why it matters: Zhipu's GLM-5.2 quickly filled the gap after Anthropic restricted access, costs one-eighth of Claude Opus 4.8, is fully open source, and the US-China gap estimate has shrunk to six months — three signals stacking up, worth recommending. Not scoring higher because the post does...

AI HOT (Curated Pool)

Xiaohu open-sources 'Xiaohu IP Studio' with 31 original characters and an auto-illustration pipeline

Blogger Xiaohu released an open-source tool called 'Xiaohu IP Studio' that auto-generates illustrations for articles. It ships with 31 original characters—15 hand-drawn line-art figures and 16 pun-based meme images. The agent reads the article, decides on an illustration type (mood image, diagram, or four-panel comic), generates the image, and self-checks with rework if needed. The default style is hand-drawn line art with light color; five alternative skins are available, including 3D blind-box and black-and-white line art. Setup requires only Python 3, works with Claude Code or Codex, and needs an OpenAI-compatible image API key (defaults to GPT-image-2). You can also output prompts only and generate images manually.

Why it matters: A practical open-source tool release with a concrete workflow design and 31 original characters, directly valuable for AI content creators. But it's a personal project open-sourcing, not an industry-level event, so it stays at the featured threshold.

Jun 25Thursday

AI HOT (Curated Pool)

Meituan LongCat open-sources VitaBench 2.0, a long-horizon dynamic agent benchmark

Meituan's LongCat team open-sourced VitaBench 2.0, a benchmark for testing how well agents model users over long, dynamic real-life scenarios. It includes 56 simulated users, 819 complex tasks, over 2,000 shifting preferences, and 66 executable tools—averaging 2,093 interaction events per user across roughly 1,580 days. Even the top model, Claude-Opus-4.6, barely scored above 0.5 in open-book mode. Thinking mode didn't consistently help on personalization tasks, and all models saw a sharp drop on tasks requiring proactive questions. The benchmark and tools are open-sourced.

Why it matters: Meituan LongCat open-sourced a large-scale long-horizon agent benchmark with concrete numbers on data volume, task design, and results — not a vague leaderboard. Score isn't higher because there's only one WeChat post so far, no cross-source confirmation yet, and the benchmark...

Jun 24Wednesday

Hacker News front page

Greptile's OpenClaw PR study shows AI-generated spam PRs now resemble early-2000s email spam

Greptile analyzed PR data from the OpenClaw repo. Weekly PRs jumped from 2 last December to 3,400 by February, with merge rates dropping from 48% to under 9.3%. One contributor submitted 106 PRs in a day at a median interval of 3 seconds. Three takeaways: PRs will need sender reputation like email spam filters—Mitchell Hashimoto's Vouch project already tackles this. More contributors using the same AI coding tools leads to convergent thinking: 4 people submitted identical SearXNG feature PRs, and 6 independently fixed the same Brave Search locale bug. Refactors merge at 35% vs. 9% for features, showing that deep codebase understanding still wins.

Why it matters: Greptile quantifies the AI-generated PR noise problem with real data from the OpenClaw repo — the numbers are striking. Downside: single-repo case study, and Greptile sells a code-review product, so there's a vested interest, but the data and methodology are transparent enough...

Jun 23Tuesday

Hacker News front page

Krea releases Krea 2 technical report: open-source text-to-image models built for aesthetic diversity and creative control

Krea 2 is a series of open-source text-to-image foundation models released under a permissive license. Instead of optimizing for a single polished default look, it aims to cover a broad range of visual styles and give users ways to explore them via text or reference images. The pretraining data deliberately excludes AI-generated images and avoids aesthetic-score filters—only duplicates, samples VLMs can't describe well, harmful biases, and overly complex images are removed. The architecture is a diffusion transformer (DiT) with iREPA, improved VAEs, Qwen3-VL text encoder, and components like GQA and sigmoid-gated attention to speed up convergence. Training runs through pretraining, midtraining, SFT, preference optimization, and RL. To bridge the gap between short user prompts and the model's rich conditioning space, Krea 2 adds a prompt expander (two-stage SFT+RL on open-source LLMs) and a style-reference system that lets users control style and mood from uploaded images, with adjustable strength and weighted mixing. It ranks in the top 10 on the Artificial Analysis text-to-image leaderboard and second among independent labs.

Why it matters: Krea 2 ships open-source with a detailed technical report and a clear data curation stance (no AI-generated images, no aesthetic scoring). Useful for model builders, but the image-gen space is crowded and Krea isn't a tier-1 lab, so it lands at the 78 featured threshold.

AI HOT (Curated Pool)

IBM open-sources CUGA, a lightweight agent framework with 20+ single-file example apps

CUGA bundles planning, execution, reflection, and tool calling into a configurable agent—just supply a tool list and a prompt. It ranked first on both AppWorld (Jul 2025–Feb 2026) and WebArena (Feb–Sep 2025) benchmarks. Three inference modes (Fast / Balanced / Accurate) are available, and code can run locally, in Docker, or inside an E2B sandbox. The tool layer supports OpenAPI, MCP, and LangChain functions; switching between OpenAI, watsonx, Ollama, and other providers is done via environment variables. Over 20 single-file example apps ship with the framework—movie recommendations, an IBM Cloud architecture advisor, and more—each requiring only one FastAPI file.

Why it matters: IBM open-sourced CUGA with concrete benchmark wins and reproducible examples, giving it solid knowledge density. But the agent-framework space is crowded and the post lacks a distinctive hook for practitioners to debate, so resonance is weak. Defaulting to the lower band per p...

TechCrunch · AI

SpaceX inks $150M/month compute deal with open source AI lab Reflection AI

SpaceX's Colossus 2 data center near Memphis landed its third major AI compute customer. Open source lab Reflection AI will pay $150 million per month starting July 1, 2026 through 2029 for immediate access to Nvidia GB300 chips. The deal follows earlier contracts with Anthropic ($1.25B/month) and Google ($920M/month). The post doesn't disclose what models Reflection AI plans to train or its funding sources.

Why it matters: SpaceX lands a $150M/month compute deal with open-source lab Reflection AI through 2029. The story has novelty, hard numbers, and industry buzz, but Reflection AI's low profile and undisclosed funding keep it at the 78 featured threshold.

Jun 22Monday

Hacker News front page

Git is forever, but Zach Geier built Oak anyway—a version control system for AI agents

Zach Geier spent four years building a VCS called Jam, sold it, and watched the acquiring company shut down within a year. Now he's using AI to build Oak, getting more done in four months than in the previous four years. Oak is a version control system designed for AI agents: virtual mounts let agents work without cloning full repos, and parallel tasks don't require worktrees. No Windows build yet, no CI, issues, or comments—but the team has been fully bootstrapped on Oak with no Git backup for months. Core and CLI are open-source; you can self-host and export to Git anytime. First 100 paid users get a custom e-ink display.

Why it matters: Strong founder narrative and a concrete technical hook (virtual mounts for agent workflows) that addresses real Git friction. Downside: this is an announcement blog with no public product, no user reports, no benchmarks — can't verify the claimed experience yet. 72 is the righ...

AI HOT (Curated Pool)

OpenAI Launches Daybreak: Codex Security and GPT-5.5-Cyber for Patch Automation

OpenAI shifts its security focus from finding bugs to automating patches. Codex Security has scanned 30M commits and flagged over 500K fixed findings. The full GPT-5.5-Cyber hits 85.6% on CyberGym, up from GPT-5.5's 81.8%. The Patch the Planet initiative, co-founded with Trail of Bits and HackerOne, brings 30+ open-source projects like cURL and Python into the fix pipeline. The post doesn't disclose Codex Security pricing or the exact scope of GPT-5.5-Cyber's limited release.

Why it matters: OpenAI launches Daybreak, shifting from vuln discovery to automated patching. Codex Security backs it with 30M scans and 500K flagged findings; GPT-5.5-Cyber posts 85.6% on CyberGym. Score held below 90 because the post doesn't disclose fix accuracy or false-positive rates — r...

Jun 21Sunday

Product Hunt · AI

Conduit: a local gateway that cuts AI agent tool-list overhead by ~90%

Adding more MCP servers slows AI agents because every server dumps its full tool list into context on each request—3 servers cost ~24k tokens before any prompt. Conduit sits as a local gateway between the agent and MCP servers, exposing 3 meta-tools the agent searches on demand instead of loading every tool. Measured: 97% less tool overhead per request, ~90% fewer total tokens, same task success rate. API keys stay in the OS keychain, no phoning home. Works with 17 clients across Windows, macOS, and Linux. Free and open source. The post doesn't list which MCP servers or clients are supported, nor the full benchmark setup.

Why it matters: A practical fix for MCP tool-list bloat with measured 97% overhead reduction, directly useful for developers building MCP agents. Score capped because it's a Product Hunt launch rather than a formal product release, and the post doesn't disclose the gateway's own compute overh...

Jun 20Saturday

Hacker News front page

GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2, challenging the bigger-model dogma

The author benchmarks GPT-5.5, DeepSeek V4 Pro, and GLM-5.2 on the AA-Omniscience hallucination metric and a Python async coding prompt. GPT-5.5 hits 86% hallucination, DeepSeek V4 Pro 94%, while GLM-5.2 scores 28%. DeepSeek V4 Pro spent nearly 4 minutes and 7.7k reasoning tokens producing a confidently wrong solution; GLM-5.2 needed 12 seconds and ~800 tokens to flag the prompt as technically impossible under single-threaded, no-polling constraints. GLM-5.2 trails GPT-5.5 by only 4 points on the AA Intelligence Index and Claude Fable 5 by 9 points—Fable 5 was restricted by the US government three days post-launch over a single jailbreak. The post argues that scaling parameters and data makes models worse at saying “I don’t know,” and frames an unsolved trilemma: raw capability, hallucination calibration, and compute efficiency. The article does not disclose GLM-5.2’s training data size or exact release date.

Why it matters: First-person benchmark with concrete, counterintuitive numbers; hits all three HKR axes. Score held at 78 rather than 85+ because it's a personal blog, sample size and methodology aren't fully detailed, limiting authority.

Jun 19Friday

AI HOT (Curated Pool)

Banning Open Source AI Would Be A Mistake

Nathan Lambert and Kevin Xu argue that Washington's recent AI regulatory moves—including an executive order, a congressional proposal, and a ban on foreign nationals accessing Anthropic's top models—could inadvertently harm open source. They frame Anthropic and OpenAI as a consolidating duopoly, noting Anthropic was caught reducing its model's capability when used to improve competitors. Open source, which underpins over 90% of global software and $8 trillion in economic value, is the only counterweight. The post is a general-audience op-ed; it does not propose specific policy fixes.

Why it matters: Co-authored commentary by Nathan Lambert and Kevin Xu directly responds to recent DC regulatory moves and discloses that Anthropic actively degrades model capabilities when used to improve competitors. The piece has conflict hook, new concrete info, and hits the open-source co...