Skip to content

Open source

Open models, frameworks and repositories: open weights, community hits and the balance between open and closed.

Latest picks

101–120 of 329

Jul 22Wednesday

Hacker News front page

Codeberg bans vibe coded projects via ToU amendment

Codeberg members voted to amend the Terms of Use, banning projects that mostly consist of LLM-generated code without human review. The proposal argues such projects have unclear copyright and lack safeguards. The post doesn't define 'mostly' or specify an enforcement timeline.

Why it matters: Codeberg membership voted to ban unreviewed AI-generated code via ToU amendment — a substantive governance move with conflict, new information, and emotional resonance. Score held back by vague enforcement details and scope limited to Codeberg ecosystem, not industry-wide.

Jul 21Tuesday

Ben's Bites

Kimi K3 tops Fable on frontend coding leaderboard, but token inefficiency cancels cost edge

Moonshot AI's Kimi K3 beat Fable and GPT-5.6-Sol on Arena's frontend coding leaderboard and came close on other benchmarks. It's a 2.8T-parameter model with a 1M-token context window and image support; weights will be open-sourced by July 27. Token inefficiency cancels its per-token price advantage: half the cost per token but twice the tokens used. New subscriptions are paused due to GPU shortages. Fable 5 is now a permanent part of Claude Max/Team plans, with Pro users getting a one-time $100 credit. Fable also found a counterexample disproving the 87-year-old Jacobian conjecture. Sierra launched Horizon, outcome-priced long-running agents. NotebookLM rebranded to Gemini Notebook and added Collections.

Why it matters: Moonshot drops Kimi K3, topping Fable and GPT-5.6-Sol on Arena's frontend coding board. 2.8T params, 1M context, open-source on July 27 — all hard signals. The token-efficiency gap is a real weakness but makes the story more substantive. Held at 82 rather than 85+ because only...

Jul 20Monday

Hacker News front page

Kimi K3 and Qwen 3.8 go open, squeezing Anthropic from both sides

Moonshot's Kimi K3 and Alibaba's Qwen 3.8 launched this week, both near Anthropic Fable 5 in performance and set to release weights publicly. The piece runs the numbers: Anthropic leases data centers and buys electricity, so inference costs scale with usage. Fable 5 costs nearly 3× per completed task vs. competitors. Open models catching up makes a premium-pricing strategy fragile. Anthropic bets on regulation and recursive self-improvement, but its product moat is thin—open-source harness startups are flooding in. The post doesn't spell out a clear countermove.

Why it matters: The K3 and Qwen 3.8 releases are notable, but the real value is the cost analysis: Fable 5 inference costs 3x competitors, and Anthropic's lack of owned infrastructure means costs scale linearly with usage. This is a concrete economic argument for open-source catching up, not ...

AI HOT (Curated Pool)

Xiaohongshu and Peking University open-source UltraEP for real-time MoE load balancing

Xiaohongshu and Peking University open-sourced UltraEP, a real-time load balancing method for large MoE models. It dynamically replicates hot experts per microbatch and per layer using exact routing info, hitting 94.6% of ideal training throughput. On Qwen3-235B, training throughput is 42% higher than Megatron-LM and prefill throughput is 1.56× SGLang. The post doesn't disclose the license or deployment requirements.

Why it matters: Joint open-source release from Xiaohongshu and Peking University tackles real MoE idle-GPU pain with strong numbers (94.6% ideal training throughput, 1.56x SGLang inference). Missing license and deployment requirements keep it from scoring higher—those determine real-world ado...

Hacker News front page

Xiaomi drops XR-1, a robot foundation model pre-trained on 100K hours of embodiment-free data

Xiaomi Robotics open-sourced XR-1, a ready-to-use robot foundation model. It pre-trains on 100K hours of embodiment-free manipulation videos across 1,700+ scenarios, then post-trains on 7,200 hours of real-robot data for embodiment and instruction alignment. Pre-training shows clean scaling laws—lower action error with more data and larger models—and those gains transfer directly to real-robot success rates with no sign of saturation yet. After post-training, XR-1 picks up new tasks like phone packing and printer refilling from under 10 hours of demos on average, hitting 75% overall success (nearly 2× π 0.5); with under 40 hours it reaches 85%. It also achieves SOTA on four sim benchmarks. Code, weights, and paper are public.

Why it matters: Xiaomi Robotics open-sourced XR-1, a robot foundation model with code, weights, and paper. The two-stage recipe (100K hrs embodiment-free pretraining + 7,200 hrs real-robot post-training) and the pretraining scaling law are hard signals, directly comparable to π 0.5. Scored as...

Jul 19Sunday

Computing Life · Share · Yage

Grok Build open-sourced its client harness, not the model or cloud

xAI released the Rust client harness that handles local files, commands, and permissions for Grok Build under Apache-2.0. The Grok model, cloud services, and the official binary build chain remain closed. The repo doesn't accept external PRs. The commit from the earlier upload controversy isn't in the public history, so the current code can't close that case. The real win: you can now pin a public commit, build it yourself, and compare its behavior against the official binary.

Why it matters: xAI open-sourcing Grok Build's client harness is substantive—Apache-2.0, headless mode, and ACP support go beyond signaling. But the model and build chain remain closed, and the repo rejects PRs, capping it below 85. All three HKR axes hit, so featured.

TechCrunch · AI

Moonshot AI open-sources Kimi K3, competitive with GPT 5.6 and Claude Fable 5

Moonshot AI open-sourced its Kimi K3 model this week. The company says it still trails Claude Fable 5 and GPT 5.6 Sol, but independent evals from Arena.ai and Vals AI place it near flagship closed models. The release coincided with Xi Jinping's speech at the World AI Conference in Shanghai; the Nasdaq dropped about 1% on Friday as chip stocks like Nvidia sold off. The discourse echoes the DeepSeek R1 moment from early 2025, now amplified by the Trump administration's tariff war with China, Anthropic's national-security scrutiny, and major AI firms preparing to go public. The post does not disclose K3's parameter count, training cost, or open-source license details.

Why it matters: Moonshot open-sourcing Kimi K3 with third-party evals showing it can compete against GPT-5.6 Sol and Claude Fable 5 is a significant signal from China's flagship model ecosystem. Score capped at 78 because this is a TechCrunch commentary piece, not the original release — key t...

Jul 17Friday

Hacker News front page

Moonshot AI launches 2.8T-parameter Kimi K3, calling it the first open 3T-class model

Moonshot AI released Kimi K3, a 2.8T-parameter model and the most expensive from a Chinese lab so far at $3/$15 per million input/output tokens—matching Claude Sonnet pricing. Self-reported benchmarks mostly beat Claude Opus 4.8 and GPT-5.5 but lose to Claude Fable 5 and GPT-5.6 Sol. On Artificial Analysis's private long-horizon knowledge eval, K3 hit an Elo of 1547, +732 over K2.6, behind only Fable 5. Cost per task is $0.94, close to GPT-5.6 Sol's $1.04 and roughly half of Opus 4.8. Output tokens dropped 21% vs K2.6. The model only offers a 'max' reasoning effort; Simon's pelican-on-a-bike SVG cost 25 cents and burned 13,241 reasoning tokens. Input token count suggests an ~85-token hidden system prompt. Vision works well. Open weights promised by July 27.

Why it matters: Moonshot AI released Kimi K3, a 2.8T-param model priced at $3/$15 — matching Claude Sonnet and making it the most expensive Chinese lab model. Self-reported benchmarks mostly beat Claude Opus 4.8 and GPT-5.5 high, but lose to Claude Fable 5 and GPT-5.6 Sol. Simon Willison's ta...

AI Chat-Group Daily (群聊日报)

Kimi K3 tops Frontend Code Arena, weights to open-source, early tests show brilliance and burnout

Kimi K3 hit #1 on Frontend Code Arena with 1679 points, beating Claude Fable 5's 1631 and taking six of seven frontend domains. It packs 2.8T params, 1M context, $3/$15 per million tokens, with full weights opening by July 27. Early testers got mixed results: one user's 199-yuan monthly plan produced stunning particle VJ effects from chat history, while another burned through a $40 coding plan in five hours as the model looped on a domain spelling error. Benchmark trust is shaky—GLM-5.2 scored well on paper but felt worse than 5.5 in practice. Writing style drew split reactions: less AI flavor but forced casual tone, nowhere near the natural Chinese of the old Opus 4.6. Same day, GPT-5.6's frontend taste was called 'very Claude-like,' Sol traced a deadlock only reproducible on Ubuntu, Linus told kernel devs AI is here to stay, and Schema harness pushed ARC-AGI-3 efficiency to 98.98% by making models think like physicists.

Why it matters: Kimi K3 tops Frontend Code Arena, winning 6 of 7 frontend categories with weights opening July 27 — a major domestic flagship release. The chat digest provides scores, params, pricing, and hands-on user feedback. Not scoring higher because the source is a community digest rath...

Jul 16Thursday

Hacker News front page

Sentinel: an open-source QA agent that reads your code before it clicks

SimbaStack open-sourced Sentinel under MIT, a QA agent that reads the codebase first, derives business flows on its own, then tests them end-to-end across frontend and backend. They pointed it at their own hotel PMS with only the repo and admin credentials, no test plan. Sentinel read the code, concluded it was a boutique hotel system, and auto-derived nine critical flows including the full reservation lifecycle, group bookings, and night audit. It ran the top two flows twice each and caught three bugs invisible to UI-only checks: a backend NO_AVAILABILITY error on a reservation that already held the room, a calendar showing a room as available when the API said it was booked, and a check-in returning 200 but leaving the guest registration status unchanged. The pipeline: a deterministic grep/find recon pass extracts code structure, Xiaomi's Mimo model derives business flows, Playwright drives the browser, and an api_request tool checks server state. Each flow runs twice by default, findings are unioned, and a 90-call cap bounds each attempt. A final vision pass scores visual hierarchy, spacing, and contrast on visited screens. It currently supports common JS stacks like Next.js, Express, Fastify, and Prisma; other stacks need a recon patch.

Why it matters: A new entrant in the open-source QA agent space with a real end-to-end experiment on a hotel PMS — not a toy demo. Score stays below 80 because there's only one blog post so far, no third-party reproduction or head-to-head comparison yet.

Hacker News front page

The LLM Critics Are Right. I Use LLMs Anyway

At Local-First Conf in Berlin, the author noticed a shared dissonance: speakers criticized LLMs while the audience applauded with Claude Code open. He concedes every critique—slop, trust erosion in OSS, broken junior-senior teaching loops, geopolitical supply risks—yet still uses LLMs heavily. The post doesn't resolve the tension; it lays out the contradiction and asks others to share their usage patterns so the community can better understand this collective unease.

Why it matters: An honest personal observation that lays out the collective dissonance devs feel about LLMs, with a concrete on-stage anecdote (Armin Ronacher's reply). Strong resonance, but lacks hard data or actionable takeaways, so the score sits right at the featured threshold.

AI HOT (Curated Pool)

xAI open-sources Grok Build coding agent and terminal UI

xAI released the full Grok Build codebase on GitHub, covering the agent loop, tool dispatch, terminal UI, and extension system. You can read the source to see how context assembly and tool calls work, or compile it yourself and point it at a local inference setup.

Why it matters: xAI open-sourced Grok Build's full codebase — agent loop, TUI, extension system, local-first support. Hits all three HKR axes for the dev audience. Score stays at the featured threshold because we only have the official announcement so far; no third-party benchmarks or hands-o...

Jul 15Wednesday

Hacker News front page

StyleSeed: A design-rules engine so AI coding agents stop shipping generic-looking UI

bitjaru open-sourced StyleSeed, a design-rules engine for AI coding tools like Claude Code, Codex, and Cursor. It teaches design judgment rather than just generating code: 74 rules, 48 components, 7 brand skins (Toss, Stripe, Linear, Notion, Raycast, Arc, Vercel), a named motion system, and 15 /ss-* skills. MIT licensed, currently at 731 stars. The post doesn't detail how rules are enforced or how the motion system works in practice, but the structure aims to suppress the 'AI-generated' look in shipped UI.

Why it matters: Adding design constraints to AI coding tools addresses a real need, and 74 rules plus brand skins give this substance beyond a concept demo. Score capped because it's a fresh Show HN launch with no user feedback or real-world results yet — graded on tool completeness alone.

Computing Life · Share · Yage

Codex stays open source, but parent-to-sub-agent task messages are now encrypted

On June 5, OpenAI merged PR #26210, encrypting task messages that Codex's parent agent sends to sub-agents. Previously, local session logs showed plaintext instructions like 'Review the authentication changes'; now only <ciphertext> remains. Sub-agent tool calls, commands, and outputs are still visible, but debugging can't tell whether the parent gave a wrong task or the sub-agent misunderstood. Encryption happens server-side in the Responses API; the local client only forwards ciphertext. This differs from earlier hidden reasoning and compaction—what's now hidden is content that directs another agent to act, not internal model thinking. The post doesn't spell out OpenAI's rationale; speculation includes prompt protection or unified cloud multi-agent services.

Why it matters: A product-change report with concrete technical details, not marketing fluff. PR numbers, issue links, and before/after comparisons are all provided. The deduction is because this is a feature adjustment rather than a new capability launch, and its impact is limited to Codex u...

Jul 14Tuesday

AI HOT (Curated Pool)

Tencent Hunyuan releases 1-bit and 4-bit quantized Hy3, a 295B MoE that runs on a single GPU

Tencent Hunyuan quantized its flagship Hy3 (295B MoE) into 1-bit and 4-bit versions that run on a single GPU. Hy3 is claimed to be best-in-class at this scale and competitive with trillion-parameter models for most agent scenarios. The quantized versions work via llama.cpp with MTP support, drastically lowering hardware requirements. Apache 2.0 license, commercial use allowed, plus two weeks of free API through OpenRouter. The post doesn't disclose quantization accuracy loss or the specific GPU memory needed.

Why it matters: Tencent Hunyuan's quantized Hy3 puts a 295B MoE model on a single GPU — immediately actionable for local deployment and agent builders. Apache 2.0 license plus a two-week free API window lowers the barrier to test. Held below 85 because the post doesn't disclose quantization a...

TechCrunch · AI

Nous Research is raising at least $75M at a $1.5B valuation, led by Robot Ventures

Nous Research, the startup behind the open-source Hermes agent, is finalizing a round at a $1.5B valuation, raising at least $75M. Robot Ventures is leading, with USV joining significantly. Three sources confirmed the deal; Nous declined to comment, and the investors didn't respond. Founded in 2023, the company previously raised $70M from Paradigm, OSS Capital, Balaji Srinivasan, and others. The post doesn't spell out how the new capital will be used or give recent Hermes updates.

Why it matters: Nous Research's Hermes agent has real traction in open-source circles, and both the numbers and investor lineup are solid. The ding is that this is 'in talks' not closed, and neither Nous nor the investors have commented — everything comes from sources.

Jul 13Monday

AI HOT (Curated Pool)

Tencent Hunyuan open-sources HyOCR-1.5: a 1B end-to-end OCR model with 6.37× faster inference

Tencent Hunyuan fully open-sourced HyOCR-1.5—training, inference, and model weights—a first for end-to-end OCR large models. The 1B-parameter model handles 8+ text-centric tasks and scores 94.74 on OmniDocBench v1.6, ranking first end-to-end. DFlash speculative decoding speeds up inference 6.37× under Transformers and 2.14× under vLLM, hitting 1.408s per page. It supports 4K resolution and a 128K context window, and uses Agentic Data Flow to extend low-resource OCR to 331 languages, ancient script recognition, and multi-image QA.

Why it matters: Tencent Hunyuan fully open-sourced an end-to-end OCR model — training code, inference code, weights. 1B params, 94.74 on OmniDocBench v1.6 (#1 among end-to-end models), 6.37x inference speedup via DFlash. A genuine open-source move from a major Chinese lab, not weights-only. S...

Jul 11Saturday

Hacker News front page

Cloudflare blocks AI agents by default, and your agent can't tell

Since July 1, 2025, Cloudflare blocks AI crawlers by default on new domains. Worse, blocked requests return a 403 with a full HTML body—the 'Just a moment' challenge page—which language models read as real content and summarize confidently with fabricated answers. The author tested eight major anti-bot vendors: naive fetches got real content zero times, got a summarizable block page eight times, and got zero signals that the fetch failed. Their open-source Fortress stealth browser clears five of eight, returning live job listings from Indeed, 878 Zillow listings, and StockX's GraphQL pricing API. DataDome and Amazon click-walls remain unsolved in the open-source build; those require the hosted Tilion Cloud layer. The fix: detect block-page signatures and fail loud so agents stop before hallucinating from a challenge screen.

Why it matters: All three HKR axes hit. The counterintuitive trap (agent reads block page as fact) drives strong click intent; concrete data on 12 sites plus the 403-with-body mechanism is real new knowledge; anyone building agents or RAG will resonate immediately. Comes with open-source tool...

Jul 9Thursday

AI HOT (Curated Pool)

Ant Lingbo open-sources LingBot-Video, a MoE video base model for embodied AI

Ant Lingbo open-sourced LingBot-Video, the first MoE-based video generation model built for embodied AI. It has 30B total parameters but activates only ~3B during inference, roughly 3× faster than a dense model of similar size. Training used 70,000 hours of robot-related video—dexterous manipulation, navigation, egocentric interaction. On the RBench benchmark for robot manipulation videos it scored 0.620, ahead of Wan2.6 (0.607) and Seedance 1.5 Pro (0.584). Internal tests also place it above NVIDIA Cosmos 3 and Hunyuan Video 1.5 on physical plausibility and motion consistency. The model targets robot action prediction, simulation data generation, and world-model research. Code is public.

Why it matters: Ant Lingbo open-sourced the first MoE video foundation model for embodied AI — 30B total params, ~3B activated during inference, 3x faster than dense models of similar scale, trained on 70k hours of real robot video. HKR all hit, but it's a fresh release with no external repro...

Jul 8Wednesday

TechCrunch · AI

French AI startup ZML releases free inference accelerator for multiple chip types

ZML released ZML/LLMD, open-source software that speeds up inference for models like Llama and DeepSeek across Nvidia, AMD, Google TPU, Apple Metal, and Intel Arc chips. Founder Steeve Morin says the goal is cheaper inference, and it's free for now. Turing Award winner Yann LeCun previously endorsed the startup. The post doesn't disclose funding details or benchmark comparisons—I'd wait for real-world numbers before getting excited.

Why it matters: Cross-chip inference acceleration is a real need, and ZML/LLMD is free, open-source, and endorsed by LeCun. But the post lacks performance benchmarks and funding details, capping the score at 72.