Skip to content

#开源/仓库

3 today

Jul 23Thursday

Latent Space

Poolside co-CEO on how a 70-person team ships a 118B MoE model in 8 weeks

Poolside co-CEO Eiso Kant walked through their 'Model Factory' on the Latent Space podcast. A team of fewer than 70 researchers runs 10,000–20,000 experiments per month, cutting model cycles from six months to five to eight weeks. Their new Laguna S 2.1 is a 118B-total, 8B-active MoE model with a 1M context window and dual thinking/no-thinking modes, beating Thinking Machines' ~1T open-weights model. Eiso argued 95% of model building comes down to better data or compute efficiency, called MCP and traditional tool calls 'stupid,' predicted RL will move earlier into pre-training, and said he'd rather live in a world with 100 foundation model companies than five.

Why it matters: Poolside opens up its model factory internals for the first time, with real numbers on the 118B MoE architecture and 5-8 week iteration cadence — useful for anyone doing model training or code tooling. Not scoring higher because Poolside's audience is still code-niche, and thi...

Jul 22Wednesday

Hacker News front page

Codeberg bans vibe coded projects via ToU amendment

Codeberg members voted to amend the Terms of Use, banning projects that mostly consist of LLM-generated code without human review. The proposal argues such projects have unclear copyright and lack safeguards. The post doesn't define 'mostly' or specify an enforcement timeline.

Why it matters: Codeberg membership voted to ban unreviewed AI-generated code via ToU amendment — a substantive governance move with conflict, new information, and emotional resonance. Score held back by vague enforcement details and scope limited to Codeberg ecosystem, not industry-wide.

Jul 21Tuesday

Ben's Bites

Kimi K3 tops Fable on frontend coding leaderboard, but token inefficiency cancels cost edge

Moonshot AI's Kimi K3 beat Fable and GPT-5.6-Sol on Arena's frontend coding leaderboard and came close on other benchmarks. It's a 2.8T-parameter model with a 1M-token context window and image support; weights will be open-sourced by July 27. Token inefficiency cancels its per-token price advantage: half the cost per token but twice the tokens used. New subscriptions are paused due to GPU shortages. Fable 5 is now a permanent part of Claude Max/Team plans, with Pro users getting a one-time $100 credit. Fable also found a counterexample disproving the 87-year-old Jacobian conjecture. Sierra launched Horizon, outcome-priced long-running agents. NotebookLM rebranded to Gemini Notebook and added Collections.

Why it matters: Moonshot drops Kimi K3, topping Fable and GPT-5.6-Sol on Arena's frontend coding board. 2.8T params, 1M context, open-source on July 27 — all hard signals. The token-efficiency gap is a real weakness but makes the story more substantive. Held at 82 rather than 85+ because only...

Jul 20Monday

Hacker News front page

Kimi K3 and Qwen 3.8 go open, squeezing Anthropic from both sides

Moonshot's Kimi K3 and Alibaba's Qwen 3.8 launched this week, both near Anthropic Fable 5 in performance and set to release weights publicly. The piece runs the numbers: Anthropic leases data centers and buys electricity, so inference costs scale with usage. Fable 5 costs nearly 3× per completed task vs. competitors. Open models catching up makes a premium-pricing strategy fragile. Anthropic bets on regulation and recursive self-improvement, but its product moat is thin—open-source harness startups are flooding in. The post doesn't spell out a clear countermove.

Why it matters: The K3 and Qwen 3.8 releases are notable, but the real value is the cost analysis: Fable 5 inference costs 3x competitors, and Anthropic's lack of owned infrastructure means costs scale linearly with usage. This is a concrete economic argument for open-source catching up, not ...

AI HOT (Curated Pool)

Xiaohongshu and Peking University open-source UltraEP for real-time MoE load balancing

Xiaohongshu and Peking University open-sourced UltraEP, a real-time load balancing method for large MoE models. It dynamically replicates hot experts per microbatch and per layer using exact routing info, hitting 94.6% of ideal training throughput. On Qwen3-235B, training throughput is 42% higher than Megatron-LM and prefill throughput is 1.56× SGLang. The post doesn't disclose the license or deployment requirements.

Why it matters: Joint open-source release from Xiaohongshu and Peking University tackles real MoE idle-GPU pain with strong numbers (94.6% ideal training throughput, 1.56x SGLang inference). Missing license and deployment requirements keep it from scoring higher—those determine real-world ado...

Hacker News front page

Xiaomi drops XR-1, a robot foundation model pre-trained on 100K hours of embodiment-free data

Xiaomi Robotics open-sourced XR-1, a ready-to-use robot foundation model. It pre-trains on 100K hours of embodiment-free manipulation videos across 1,700+ scenarios, then post-trains on 7,200 hours of real-robot data for embodiment and instruction alignment. Pre-training shows clean scaling laws—lower action error with more data and larger models—and those gains transfer directly to real-robot success rates with no sign of saturation yet. After post-training, XR-1 picks up new tasks like phone packing and printer refilling from under 10 hours of demos on average, hitting 75% overall success (nearly 2× π 0.5); with under 40 hours it reaches 85%. It also achieves SOTA on four sim benchmarks. Code, weights, and paper are public.

Why it matters: Xiaomi Robotics open-sourced XR-1, a robot foundation model with code, weights, and paper. The two-stage recipe (100K hrs embodiment-free pretraining + 7,200 hrs real-robot post-training) and the pretraining scaling law are hard signals, directly comparable to π 0.5. Scored as...

Jul 19Sunday

Computing Life · Share · Yage

Grok Build open-sourced its client harness, not the model or cloud

xAI released the Rust client harness that handles local files, commands, and permissions for Grok Build under Apache-2.0. The Grok model, cloud services, and the official binary build chain remain closed. The repo doesn't accept external PRs. The commit from the earlier upload controversy isn't in the public history, so the current code can't close that case. The real win: you can now pin a public commit, build it yourself, and compare its behavior against the official binary.

Why it matters: xAI open-sourcing Grok Build's client harness is substantive—Apache-2.0, headless mode, and ACP support go beyond signaling. But the model and build chain remain closed, and the repo rejects PRs, capping it below 85. All three HKR axes hit, so featured.

TechCrunch · AI

Moonshot AI open-sources Kimi K3, competitive with GPT 5.6 and Claude Fable 5

Moonshot AI open-sourced its Kimi K3 model this week. The company says it still trails Claude Fable 5 and GPT 5.6 Sol, but independent evals from Arena.ai and Vals AI place it near flagship closed models. The release coincided with Xi Jinping's speech at the World AI Conference in Shanghai; the Nasdaq dropped about 1% on Friday as chip stocks like Nvidia sold off. The discourse echoes the DeepSeek R1 moment from early 2025, now amplified by the Trump administration's tariff war with China, Anthropic's national-security scrutiny, and major AI firms preparing to go public. The post does not disclose K3's parameter count, training cost, or open-source license details.

Why it matters: Moonshot open-sourcing Kimi K3 with third-party evals showing it can compete against GPT-5.6 Sol and Claude Fable 5 is a significant signal from China's flagship model ecosystem. Score capped at 78 because this is a TechCrunch commentary piece, not the original release — key t...

Jul 17Friday

Hacker News front page

Moonshot AI launches 2.8T-parameter Kimi K3, calling it the first open 3T-class model

Moonshot AI released Kimi K3, a 2.8T-parameter model and the most expensive from a Chinese lab so far at $3/$15 per million input/output tokens—matching Claude Sonnet pricing. Self-reported benchmarks mostly beat Claude Opus 4.8 and GPT-5.5 but lose to Claude Fable 5 and GPT-5.6 Sol. On Artificial Analysis's private long-horizon knowledge eval, K3 hit an Elo of 1547, +732 over K2.6, behind only Fable 5. Cost per task is $0.94, close to GPT-5.6 Sol's $1.04 and roughly half of Opus 4.8. Output tokens dropped 21% vs K2.6. The model only offers a 'max' reasoning effort; Simon's pelican-on-a-bike SVG cost 25 cents and burned 13,241 reasoning tokens. Input token count suggests an ~85-token hidden system prompt. Vision works well. Open weights promised by July 27.

Why it matters: Moonshot AI released Kimi K3, a 2.8T-param model priced at $3/$15 — matching Claude Sonnet and making it the most expensive Chinese lab model. Self-reported benchmarks mostly beat Claude Opus 4.8 and GPT-5.5 high, but lose to Claude Fable 5 and GPT-5.6 Sol. Simon Willison's ta...

AI Chat-Group Daily (群聊日报)

Kimi K3 tops Frontend Code Arena, weights to open-source, early tests show brilliance and burnout

Kimi K3 hit #1 on Frontend Code Arena with 1679 points, beating Claude Fable 5's 1631 and taking six of seven frontend domains. It packs 2.8T params, 1M context, $3/$15 per million tokens, with full weights opening by July 27. Early testers got mixed results: one user's 199-yuan monthly plan produced stunning particle VJ effects from chat history, while another burned through a $40 coding plan in five hours as the model looped on a domain spelling error. Benchmark trust is shaky—GLM-5.2 scored well on paper but felt worse than 5.5 in practice. Writing style drew split reactions: less AI flavor but forced casual tone, nowhere near the natural Chinese of the old Opus 4.6. Same day, GPT-5.6's frontend taste was called 'very Claude-like,' Sol traced a deadlock only reproducible on Ubuntu, Linus told kernel devs AI is here to stay, and Schema harness pushed ARC-AGI-3 efficiency to 98.98% by making models think like physicists.

Why it matters: Kimi K3 tops Frontend Code Arena, winning 6 of 7 frontend categories with weights opening July 27 — a major domestic flagship release. The chat digest provides scores, params, pricing, and hands-on user feedback. Not scoring higher because the source is a community digest rath...

Jul 16Thursday

Hacker News front page

Sentinel: an open-source QA agent that reads your code before it clicks

SimbaStack open-sourced Sentinel under MIT, a QA agent that reads the codebase first, derives business flows on its own, then tests them end-to-end across frontend and backend. They pointed it at their own hotel PMS with only the repo and admin credentials, no test plan. Sentinel read the code, concluded it was a boutique hotel system, and auto-derived nine critical flows including the full reservation lifecycle, group bookings, and night audit. It ran the top two flows twice each and caught three bugs invisible to UI-only checks: a backend NO_AVAILABILITY error on a reservation that already held the room, a calendar showing a room as available when the API said it was booked, and a check-in returning 200 but leaving the guest registration status unchanged. The pipeline: a deterministic grep/find recon pass extracts code structure, Xiaomi's Mimo model derives business flows, Playwright drives the browser, and an api_request tool checks server state. Each flow runs twice by default, findings are unioned, and a 90-call cap bounds each attempt. A final vision pass scores visual hierarchy, spacing, and contrast on visited screens. It currently supports common JS stacks like Next.js, Express, Fastify, and Prisma; other stacks need a recon patch.

Why it matters: A new entrant in the open-source QA agent space with a real end-to-end experiment on a hotel PMS — not a toy demo. Score stays below 80 because there's only one blog post so far, no third-party reproduction or head-to-head comparison yet.

Hacker News front page

The LLM Critics Are Right. I Use LLMs Anyway

At Local-First Conf in Berlin, the author noticed a shared dissonance: speakers criticized LLMs while the audience applauded with Claude Code open. He concedes every critique—slop, trust erosion in OSS, broken junior-senior teaching loops, geopolitical supply risks—yet still uses LLMs heavily. The post doesn't resolve the tension; it lays out the contradiction and asks others to share their usage patterns so the community can better understand this collective unease.

Why it matters: An honest personal observation that lays out the collective dissonance devs feel about LLMs, with a concrete on-stage anecdote (Armin Ronacher's reply). Strong resonance, but lacks hard data or actionable takeaways, so the score sits right at the featured threshold.

AI HOT (Curated Pool)

xAI open-sources Grok Build coding agent and terminal UI

xAI released the full Grok Build codebase on GitHub, covering the agent loop, tool dispatch, terminal UI, and extension system. You can read the source to see how context assembly and tool calls work, or compile it yourself and point it at a local inference setup.

Why it matters: xAI open-sourced Grok Build's full codebase — agent loop, TUI, extension system, local-first support. Hits all three HKR axes for the dev audience. Score stays at the featured threshold because we only have the official announcement so far; no third-party benchmarks or hands-o...

Jul 15Wednesday

Hacker News front page

StyleSeed: A design-rules engine so AI coding agents stop shipping generic-looking UI

bitjaru open-sourced StyleSeed, a design-rules engine for AI coding tools like Claude Code, Codex, and Cursor. It teaches design judgment rather than just generating code: 74 rules, 48 components, 7 brand skins (Toss, Stripe, Linear, Notion, Raycast, Arc, Vercel), a named motion system, and 15 /ss-* skills. MIT licensed, currently at 731 stars. The post doesn't detail how rules are enforced or how the motion system works in practice, but the structure aims to suppress the 'AI-generated' look in shipped UI.

Why it matters: Adding design constraints to AI coding tools addresses a real need, and 74 rules plus brand skins give this substance beyond a concept demo. Score capped because it's a fresh Show HN launch with no user feedback or real-world results yet — graded on tool completeness alone.

Computing Life · Share · Yage

Codex stays open source, but parent-to-sub-agent task messages are now encrypted

On June 5, OpenAI merged PR #26210, encrypting task messages that Codex's parent agent sends to sub-agents. Previously, local session logs showed plaintext instructions like 'Review the authentication changes'; now only <ciphertext> remains. Sub-agent tool calls, commands, and outputs are still visible, but debugging can't tell whether the parent gave a wrong task or the sub-agent misunderstood. Encryption happens server-side in the Responses API; the local client only forwards ciphertext. This differs from earlier hidden reasoning and compaction—what's now hidden is content that directs another agent to act, not internal model thinking. The post doesn't spell out OpenAI's rationale; speculation includes prompt protection or unified cloud multi-agent services.

Why it matters: A product-change report with concrete technical details, not marketing fluff. PR numbers, issue links, and before/after comparisons are all provided. The deduction is because this is a feature adjustment rather than a new capability launch, and its impact is limited to Codex u...

Jul 14Tuesday

AI HOT (Curated Pool)

Tencent Hunyuan releases 1-bit and 4-bit quantized Hy3, a 295B MoE that runs on a single GPU

Tencent Hunyuan quantized its flagship Hy3 (295B MoE) into 1-bit and 4-bit versions that run on a single GPU. Hy3 is claimed to be best-in-class at this scale and competitive with trillion-parameter models for most agent scenarios. The quantized versions work via llama.cpp with MTP support, drastically lowering hardware requirements. Apache 2.0 license, commercial use allowed, plus two weeks of free API through OpenRouter. The post doesn't disclose quantization accuracy loss or the specific GPU memory needed.

Why it matters: Tencent Hunyuan's quantized Hy3 puts a 295B MoE model on a single GPU — immediately actionable for local deployment and agent builders. Apache 2.0 license plus a two-week free API window lowers the barrier to test. Held below 85 because the post doesn't disclose quantization a...

TechCrunch · AI

Nous Research is raising at least $75M at a $1.5B valuation, led by Robot Ventures

Nous Research, the startup behind the open-source Hermes agent, is finalizing a round at a $1.5B valuation, raising at least $75M. Robot Ventures is leading, with USV joining significantly. Three sources confirmed the deal; Nous declined to comment, and the investors didn't respond. Founded in 2023, the company previously raised $70M from Paradigm, OSS Capital, Balaji Srinivasan, and others. The post doesn't spell out how the new capital will be used or give recent Hermes updates.

Why it matters: Nous Research's Hermes agent has real traction in open-source circles, and both the numbers and investor lineup are solid. The ding is that this is 'in talks' not closed, and neither Nous nor the investors have commented — everything comes from sources.

Jul 13Monday

AI HOT (Curated Pool)

Tencent Hunyuan open-sources HyOCR-1.5: a 1B end-to-end OCR model with 6.37× faster inference

Tencent Hunyuan fully open-sourced HyOCR-1.5—training, inference, and model weights—a first for end-to-end OCR large models. The 1B-parameter model handles 8+ text-centric tasks and scores 94.74 on OmniDocBench v1.6, ranking first end-to-end. DFlash speculative decoding speeds up inference 6.37× under Transformers and 2.14× under vLLM, hitting 1.408s per page. It supports 4K resolution and a 128K context window, and uses Agentic Data Flow to extend low-resource OCR to 331 languages, ancient script recognition, and multi-image QA.

Why it matters: Tencent Hunyuan fully open-sourced an end-to-end OCR model — training code, inference code, weights. 1B params, 94.74 on OmniDocBench v1.6 (#1 among end-to-end models), 6.37x inference speedup via DFlash. A genuine open-source move from a major Chinese lab, not weights-only. S...

Jul 11Saturday

Hacker News front page

Cloudflare blocks AI agents by default, and your agent can't tell

Since July 1, 2025, Cloudflare blocks AI crawlers by default on new domains. Worse, blocked requests return a 403 with a full HTML body—the 'Just a moment' challenge page—which language models read as real content and summarize confidently with fabricated answers. The author tested eight major anti-bot vendors: naive fetches got real content zero times, got a summarizable block page eight times, and got zero signals that the fetch failed. Their open-source Fortress stealth browser clears five of eight, returning live job listings from Indeed, 878 Zillow listings, and StockX's GraphQL pricing API. DataDome and Amazon click-walls remain unsolved in the open-source build; those require the hosted Tilion Cloud layer. The fix: detect block-page signatures and fail loud so agents stop before hallucinating from a challenge screen.

Why it matters: All three HKR axes hit. The counterintuitive trap (agent reads block page as fact) drives strong click intent; concrete data on 12 sites plus the 403-with-body mechanism is real new knowledge; anyone building agents or RAG will resonate immediately. Comes with open-source tool...

Jul 9Thursday

AI HOT (Curated Pool)

Ant Lingbo open-sources LingBot-Video, a MoE video base model for embodied AI

Ant Lingbo open-sourced LingBot-Video, the first MoE-based video generation model built for embodied AI. It has 30B total parameters but activates only ~3B during inference, roughly 3× faster than a dense model of similar size. Training used 70,000 hours of robot-related video—dexterous manipulation, navigation, egocentric interaction. On the RBench benchmark for robot manipulation videos it scored 0.620, ahead of Wan2.6 (0.607) and Seedance 1.5 Pro (0.584). Internal tests also place it above NVIDIA Cosmos 3 and Hunyuan Video 1.5 on physical plausibility and motion consistency. The model targets robot action prediction, simulation data generation, and world-model research. Code is public.

Why it matters: Ant Lingbo open-sourced the first MoE video foundation model for embodied AI — 30B total params, ~3B activated during inference, 3x faster than dense models of similar scale, trained on 70k hours of real robot video. HKR all hit, but it's a fresh release with no external repro...

Jul 8Wednesday

TechCrunch · AI

French AI startup ZML releases free inference accelerator for multiple chip types

ZML released ZML/LLMD, open-source software that speeds up inference for models like Llama and DeepSeek across Nvidia, AMD, Google TPU, Apple Metal, and Intel Arc chips. Founder Steeve Morin says the goal is cheaper inference, and it's free for now. Turing Award winner Yann LeCun previously endorsed the startup. The post doesn't disclose funding details or benchmark comparisons—I'd wait for real-world numbers before getting excited.

Why it matters: Cross-chip inference acceleration is a real need, and ZML/LLMD is free, open-source, and endorsed by LeCun. But the post lacks performance benchmarks and funding details, capping the score at 72.

AI HOT (Curated Pool)

Liquid AI open-sources Antidoom, a final-token preference optimization method that fixes reasoning model doom loops

Reasoning models can get stuck in doom loops, repeating useless tokens until the context window fills up. Liquid AI open-sourced Antidoom, which uses Final Token Preference Optimization (FTPO) to fix this. The method trains the model on 1,040 preference pairs to learn when to stop at the end of reasoning. On DeepSeek V4 Pro, the doom-loop rate dropped from 3.2% to 0.3% without hurting math or coding scores. The post doesn't disclose training cost or how well it transfers to non-DeepSeek models.

Why it matters: Liquid AI open-sourced a practical fix for reasoning model doom loops, dropping the rate from 3.2% to 0.3% on DeepSeek V4 Pro — solid numbers. Not scoring higher because it's a single blog post with no paper or third-party validation yet; 78 for a strong single-source piece.

Jul 5Sunday

AI HOT (Curated Pool)

Meituan LongCat-2.0 fully open-sourced under MIT license, releasing 1.6T MoE weights and inference code

Meituan fully open-sourced LongCat-2.0 under MIT license, releasing both weights and inference code. It's a 1.6T-parameter MoE model activating ~48B per token, with 1M-token context. LongCat Sparse Attention handles long sequences, Zero-Compute Experts dynamically activate 33B–56B to avoid wasted compute, and MOPD routes tasks across Agent, Reasoning, and Interaction expert groups. On benchmarks: SWE-bench Pro hits 59.5, edging out GPT-5.5's 58.6; Terminal-Bench 2.1 scores 70.8; multilingual SWE-bench reaches 77.3. It natively integrates with Claude Code, OpenClaw, and Hermes Agent, supports GPU and NPU deployment, and has been validated on large-scale domestic clusters.

Why it matters: Meituan fully open-sources LongCat-2.0, a 1.6T MoE model, under MIT license with weights and inference code — a rare move from a major Chinese tech company. The 1M-token context window and sparse attention design are concrete technical hooks, not just marketing. Score held at ...

Jul 2Thursday

Hacker News front page

git-annex maintainer spent 100 hours removing LLM-generated code from dependencies

Joey Hess audited git-annex's entire dependency tree to exclude LLM-generated code. He found an incoherent 1,489-line commit message with 10,000 lines of changes, and an LLM prompt that copied code from another project—avoiding infringement only by luck. Hess says the only upside of this 100-hour effort is better dependency quality data for future decisions. He notes the Software Freedom Conservancy has already backed off on this issue, and he is reconsidering his own participation in these communities.

Why it matters: Joey Hess personally spent 100 hours auditing git-annex's dependency tree for AI-generated code, surfaced two concrete horror stories, and noted SFC already punted. HKR all hit, but this is a personal practice report, not an industry-level event — 78 featured.

AI HOT (Curated Pool)

Tencent Hy3 released: matches much larger flagship models with 2-5x fewer parameters

Tencent officially released Hy3 under Apache 2.0. With 1/5 to 1/2 the parameters of competing flagship models, Hy3 matches or beats them on reasoning, agent, and long-context benchmarks. In a 270-person internal blind test, Hy3 scored 2.67/4 vs GLM5.1's 2.51/4. Hallucination rate dropped from 12.5% to 5.4%, multi-turn error rate from 17.4% to 7.9%. WorkBuddy task completion jumped from 72% to 90%, with 34% less time. API pricing: ¥1/M input tokens, ¥4/M output, ¥0.25 cached. The post does not disclose exact parameter count or training details.

Why it matters: Tencent Hunyuan releases Hy3, open-source under Apache 2.0, with parameter counts 1/5 to 1/2 of competitors yet matching or beating them on reasoning, agent, and long-context benchmarks. A 270-person blind test shows it beating GLM5.1. Domestic flagship open-source release is ...

Computing Life · Share · Yage

Fable 5's 18-day ban: Anthropic's share went to GLM

Anthropic's Fable 5 was taken offline by US export controls three days after launch, for 18 days. OpenRouter daily token data shows total volume grew from 24T to 32T, but Anthropic's share dropped from 20.7% to 17.6% and its absolute volume shrank. GLM was the biggest winner, jumping from 1.8% to 7.4% share. GLM-5.1 saw a spike on the day GLM-5.2 launched, then collapsed 48 hours later as GLM-5.2 took over. Community sentiment shifted from sympathy for Anthropic to mocking its business strategy. The post notes this only captures API-layer data, not first-party subscriptions, and token volume comparisons overstate GLM's share due to a 10x price gap.

Why it matters: A data-driven attribution of the 18-day Fable 5 ban's competitive impact using OpenRouter daily token data: Anthropic lost 3.1pp share, GLM gained 6.7pp, plus a weird anomaly where GLM-5.1 traffic spiked 4-5x after GLM-5.2 launched. Counterintuitive findings that directly touc...

Jul 1Wednesday

Hacker News front page

Open-source game engine Godot will no longer accept AI-authored code contributions

Godot maintainers will reject AI-authored code contributions, arguing that heavy AI users often don't understand their own code well enough to fix it. The project worries that AI-generated patches look correct but hide bugs, undermining long-term maintenance. The post doesn't specify the effective date or which detection tools are used.

Why it matters: Godot is a major open-source project in the game engine space. Its maintainers publicly rejecting AI-authored code contributions, with a concrete reason (contributors don't understand their own code and can't fix bugs), is directly relevant to the AI-assisted coding debate. Sc...

AI HOT (Curated Pool)

Meituan releases LongCat-2.0: a 1.6T-parameter model trained on 50,000 domestic GPUs, now open source

Meituan open-sourced LongCat-2.0, a 1.6T total-parameter model with ~48B activated per inference and native 1M context. It was trained and served entirely on a 50,000-card domestic GPU cluster. The architecture combines LSA sparse attention, zero-compute experts, ScMoE, and MOPD multi-expert fusion that blends Agent, Reasoning, and Interaction expert groups. SWE-bench Pro hits 59.5, Multilingual 77.3. A preview is live on OpenRouter and longcat.ai, already ranking top three globally in monthly calls on OpenRouter. The post doesn't disclose training cost, inference latency, or the specific domestic chip model, so I'd hold off on those details.

Why it matters: Meituan's trillion-param model trained end-to-end on domestic GPUs is the headline; code benchmark scores are solid. Not scoring higher because Meituan isn't a tier-1 model lab yet, and real-world usability depends on the open-source release.

Jun 30Tuesday

Hacker News front page

Meituan open-sources LongCat-2.0, a 1.6T MoE model with 48B active params, trained entirely on AI ASIC superpods

Meituan open-sourced LongCat-2.0, a 1.6-trillion-parameter MoE model with ~48B active parameters per token. It was pretrained on over 35 trillion tokens using 50K+ in-house AI ASICs with no rollbacks or irrecoverable loss spikes, showing frontier-scale training is viable on non-GPU hardware. The model targets long-context and agentic workloads: it introduces LongCat Sparse Attention to speed up 1M-token processing and was trained on hundreds of billions of 1M-context tokens. Official charts place it alongside Gemini 3.1 Pro, GPT-5.5, and Opus 4.8 on Terminal-Bench 2.1, SWE-bench Pro, and other coding/agent benchmarks, though the post does not provide exact numeric comparisons. An N-gram Embedding module with 135B parameters expands the embedding space roughly 100×, which the team claims outperforms scaling standard MoE experts by the same amount. The model is integrated with Claude Code, OpenClaw, and Hermes; code and weights are available on GitHub and HuggingFace.

Why it matters: Meituan open-sources a 1.6T MoE model trained entirely on in-house AI ASICs across 50k+ cards with zero rollbacks, plus dedicated long-context and agent optimizations. Score held at 82 rather than higher because we only have the official blog post — no third-party evals or rea...

Jun 26Friday

New York Times Chinese

Chinese AI Models Narrow Performance Gap with Anthropic and OpenAI

Zhipu's GLM-5.2 surged in popularity after Anthropic restricted access to Fable and Mythos, entering OpenRouter's top ten. It costs about one-eighth of Claude Opus 4.8 for certain tasks and is fully open-source. Experts estimate China's lag behind US firms has shrunk to six months or less. The post notes Zhipu's compute spending exceeded 7x its revenue in H1 2025, but does not disclose whether GLM-5.2's training involved distillation.

Why it matters: Zhipu's GLM-5.2 quickly filled the gap after Anthropic restricted access, costs one-eighth of Claude Opus 4.8, is fully open source, and the US-China gap estimate has shrunk to six months — three signals stacking up, worth recommending. Not scoring higher because the post does...

AI HOT (Curated Pool)

Xiaohu open-sources 'Xiaohu IP Studio' with 31 original characters and an auto-illustration pipeline

Blogger Xiaohu released an open-source tool called 'Xiaohu IP Studio' that auto-generates illustrations for articles. It ships with 31 original characters—15 hand-drawn line-art figures and 16 pun-based meme images. The agent reads the article, decides on an illustration type (mood image, diagram, or four-panel comic), generates the image, and self-checks with rework if needed. The default style is hand-drawn line art with light color; five alternative skins are available, including 3D blind-box and black-and-white line art. Setup requires only Python 3, works with Claude Code or Codex, and needs an OpenAI-compatible image API key (defaults to GPT-image-2). You can also output prompts only and generate images manually.

Why it matters: A practical open-source tool release with a concrete workflow design and 31 original characters, directly valuable for AI content creators. But it's a personal project open-sourcing, not an industry-level event, so it stays at the featured threshold.

Jun 25Thursday

AI HOT (Curated Pool)

Meituan LongCat open-sources VitaBench 2.0, a long-horizon dynamic agent benchmark

Meituan's LongCat team open-sourced VitaBench 2.0, a benchmark for testing how well agents model users over long, dynamic real-life scenarios. It includes 56 simulated users, 819 complex tasks, over 2,000 shifting preferences, and 66 executable tools—averaging 2,093 interaction events per user across roughly 1,580 days. Even the top model, Claude-Opus-4.6, barely scored above 0.5 in open-book mode. Thinking mode didn't consistently help on personalization tasks, and all models saw a sharp drop on tasks requiring proactive questions. The benchmark and tools are open-sourced.

Why it matters: Meituan LongCat open-sourced a large-scale long-horizon agent benchmark with concrete numbers on data volume, task design, and results — not a vague leaderboard. Score isn't higher because there's only one WeChat post so far, no cross-source confirmation yet, and the benchmark...

Jun 24Wednesday

Hacker News front page

Greptile's OpenClaw PR study shows AI-generated spam PRs now resemble early-2000s email spam

Greptile analyzed PR data from the OpenClaw repo. Weekly PRs jumped from 2 last December to 3,400 by February, with merge rates dropping from 48% to under 9.3%. One contributor submitted 106 PRs in a day at a median interval of 3 seconds. Three takeaways: PRs will need sender reputation like email spam filters—Mitchell Hashimoto's Vouch project already tackles this. More contributors using the same AI coding tools leads to convergent thinking: 4 people submitted identical SearXNG feature PRs, and 6 independently fixed the same Brave Search locale bug. Refactors merge at 35% vs. 9% for features, showing that deep codebase understanding still wins.

Why it matters: Greptile quantifies the AI-generated PR noise problem with real data from the OpenClaw repo — the numbers are striking. Downside: single-repo case study, and Greptile sells a code-review product, so there's a vested interest, but the data and methodology are transparent enough...

Jun 23Tuesday

Hacker News front page

Krea releases Krea 2 technical report: open-source text-to-image models built for aesthetic diversity and creative control

Krea 2 is a series of open-source text-to-image foundation models released under a permissive license. Instead of optimizing for a single polished default look, it aims to cover a broad range of visual styles and give users ways to explore them via text or reference images. The pretraining data deliberately excludes AI-generated images and avoids aesthetic-score filters—only duplicates, samples VLMs can't describe well, harmful biases, and overly complex images are removed. The architecture is a diffusion transformer (DiT) with iREPA, improved VAEs, Qwen3-VL text encoder, and components like GQA and sigmoid-gated attention to speed up convergence. Training runs through pretraining, midtraining, SFT, preference optimization, and RL. To bridge the gap between short user prompts and the model's rich conditioning space, Krea 2 adds a prompt expander (two-stage SFT+RL on open-source LLMs) and a style-reference system that lets users control style and mood from uploaded images, with adjustable strength and weighted mixing. It ranks in the top 10 on the Artificial Analysis text-to-image leaderboard and second among independent labs.

Why it matters: Krea 2 ships open-source with a detailed technical report and a clear data curation stance (no AI-generated images, no aesthetic scoring). Useful for model builders, but the image-gen space is crowded and Krea isn't a tier-1 lab, so it lands at the 78 featured threshold.

AI HOT (Curated Pool)

IBM open-sources CUGA, a lightweight agent framework with 20+ single-file example apps

CUGA bundles planning, execution, reflection, and tool calling into a configurable agent—just supply a tool list and a prompt. It ranked first on both AppWorld (Jul 2025–Feb 2026) and WebArena (Feb–Sep 2025) benchmarks. Three inference modes (Fast / Balanced / Accurate) are available, and code can run locally, in Docker, or inside an E2B sandbox. The tool layer supports OpenAPI, MCP, and LangChain functions; switching between OpenAI, watsonx, Ollama, and other providers is done via environment variables. Over 20 single-file example apps ship with the framework—movie recommendations, an IBM Cloud architecture advisor, and more—each requiring only one FastAPI file.

Why it matters: IBM open-sourced CUGA with concrete benchmark wins and reproducible examples, giving it solid knowledge density. But the agent-framework space is crowded and the post lacks a distinctive hook for practitioners to debate, so resonance is weak. Defaulting to the lower band per p...

TechCrunch · AI

SpaceX inks $150M/month compute deal with open source AI lab Reflection AI

SpaceX's Colossus 2 data center near Memphis landed its third major AI compute customer. Open source lab Reflection AI will pay $150 million per month starting July 1, 2026 through 2029 for immediate access to Nvidia GB300 chips. The deal follows earlier contracts with Anthropic ($1.25B/month) and Google ($920M/month). The post doesn't disclose what models Reflection AI plans to train or its funding sources.

Why it matters: SpaceX lands a $150M/month compute deal with open-source lab Reflection AI through 2029. The story has novelty, hard numbers, and industry buzz, but Reflection AI's low profile and undisclosed funding keep it at the 78 featured threshold.

Jun 22Monday

Hacker News front page

Git is forever, but Zach Geier built Oak anyway—a version control system for AI agents

Zach Geier spent four years building a VCS called Jam, sold it, and watched the acquiring company shut down within a year. Now he's using AI to build Oak, getting more done in four months than in the previous four years. Oak is a version control system designed for AI agents: virtual mounts let agents work without cloning full repos, and parallel tasks don't require worktrees. No Windows build yet, no CI, issues, or comments—but the team has been fully bootstrapped on Oak with no Git backup for months. Core and CLI are open-source; you can self-host and export to Git anytime. First 100 paid users get a custom e-ink display.

Why it matters: Strong founder narrative and a concrete technical hook (virtual mounts for agent workflows) that addresses real Git friction. Downside: this is an announcement blog with no public product, no user reports, no benchmarks — can't verify the claimed experience yet. 72 is the righ...

AI HOT (Curated Pool)

OpenAI Launches Daybreak: Codex Security and GPT-5.5-Cyber for Patch Automation

OpenAI shifts its security focus from finding bugs to automating patches. Codex Security has scanned 30M commits and flagged over 500K fixed findings. The full GPT-5.5-Cyber hits 85.6% on CyberGym, up from GPT-5.5's 81.8%. The Patch the Planet initiative, co-founded with Trail of Bits and HackerOne, brings 30+ open-source projects like cURL and Python into the fix pipeline. The post doesn't disclose Codex Security pricing or the exact scope of GPT-5.5-Cyber's limited release.

Why it matters: OpenAI launches Daybreak, shifting from vuln discovery to automated patching. Codex Security backs it with 30M scans and 500K flagged findings; GPT-5.5-Cyber posts 85.6% on CyberGym. Score held below 90 because the post doesn't disclose fix accuracy or false-positive rates — r...

Jun 21Sunday

Product Hunt · AI

Conduit: a local gateway that cuts AI agent tool-list overhead by ~90%

Adding more MCP servers slows AI agents because every server dumps its full tool list into context on each request—3 servers cost ~24k tokens before any prompt. Conduit sits as a local gateway between the agent and MCP servers, exposing 3 meta-tools the agent searches on demand instead of loading every tool. Measured: 97% less tool overhead per request, ~90% fewer total tokens, same task success rate. API keys stay in the OS keychain, no phoning home. Works with 17 clients across Windows, macOS, and Linux. Free and open source. The post doesn't list which MCP servers or clients are supported, nor the full benchmark setup.

Why it matters: A practical fix for MCP tool-list bloat with measured 97% overhead reduction, directly useful for developers building MCP agents. Score capped because it's a Product Hunt launch rather than a formal product release, and the post doesn't disclose the gateway's own compute overh...

Jun 20Saturday

Hacker News front page

GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2, challenging the bigger-model dogma

The author benchmarks GPT-5.5, DeepSeek V4 Pro, and GLM-5.2 on the AA-Omniscience hallucination metric and a Python async coding prompt. GPT-5.5 hits 86% hallucination, DeepSeek V4 Pro 94%, while GLM-5.2 scores 28%. DeepSeek V4 Pro spent nearly 4 minutes and 7.7k reasoning tokens producing a confidently wrong solution; GLM-5.2 needed 12 seconds and ~800 tokens to flag the prompt as technically impossible under single-threaded, no-polling constraints. GLM-5.2 trails GPT-5.5 by only 4 points on the AA Intelligence Index and Claude Fable 5 by 9 points—Fable 5 was restricted by the US government three days post-launch over a single jailbreak. The post argues that scaling parameters and data makes models worse at saying “I don’t know,” and frames an unsolved trilemma: raw capability, hallucination calibration, and compute efficiency. The article does not disclose GLM-5.2’s training data size or exact release date.

Why it matters: First-person benchmark with concrete, counterintuitive numbers; hits all three HKR axes. Score held at 78 rather than 85+ because it's a personal blog, sample size and methodology aren't fully detailed, limiting authority.