Skip to content

Open source

Open models, frameworks and repositories: open weights, community hits and the balance between open and closed.

Latest picks

1–20 of 329

Sep 28Monday

Hacker News front page

Nvidia launches a hardware watchdog chip to stop rogue AI agents in milliseconds

Nvidia launched the Open Agent Safety Platform with two layers: OpenShell, an open-source tool that traces every agent action and enforces boundaries, and Sentry, a BlueField-4-based reference design that acts as an external watchdog, quarantining rogue agents in milliseconds. Over 100 companies including Anthropic, Microsoft, and SpaceXAI have signed on, but OpenAI, Google, Meta, and Amazon are absent. The controls sit outside the model so agents can't talk or code their way around them. Sentry pricing and ship date are not disclosed, and all claims come from Nvidia and partners with no independent testing yet.

Why it matters: Nvidia's Open Agent Safety Platform has a two-layer hardware-software design with model-independent control and millisecond isolation, plus named backing from Anthropic and SpaceXAI. HKR all hit. Not scoring higher because only a blog report so far — no official Nvidia technic...

AI HOT (Curated Pool)

NVIDIA open-sources OpenShell 0.1.0 to add runtime permissions and sandboxing for AI agents

NVIDIA released OpenShell 0.1.0, an open-source runtime that enforces which systems and data an AI agent can access without rewriting the agent. It bundles sandboxed execution, controlled service access, credential management, and formal policy analysis so teams can restrict API operations and protect credentials outside the agent workload. Cadence, Slack, and Gecko Robotics are already adopting it for chip design, enterprise automation, and physical robot governance. Three components—Gateway, Supervisor, and Sandbox—manage agent fleets, inspect outbound requests against policy, and apply kernel-level filesystem and process controls. A policy prover uses formal logic to verify that permissions stay within defined boundaries. It supports Codex, Claude Code, Pi, Hermes, and runs on Docker and Kubernetes.

Why it matters: NVIDIA open-sources an Agent security runtime with three-layer architecture and formal verification — a real need for teams deploying agents. Score held back because it's v0.1.0 with no perf data or real deployment cases in the post; treat as substantive but unproven.

AI HOT (Curated Pool)

NVIDIA open-sources Agent Safety Platform with in-silicon monitoring and DPU-level enforcement

NVIDIA released an open-source agent safety platform that bakes monitoring and enforcement into Vera CPUs and BlueField-4 DPUs. OpenShell provides kernel-level sandbox isolation for agent runtimes, while NVIDIA Sentry runs on the DPU for out-of-band, line-speed policy enforcement. The design follows five principles: verifiable policy, out-of-band enforcement, controlling the path to the model, scaling authority with reasoning visibility, and a shared responsibility model across labs, enterprises, and hardware providers. In Vera Rubin POD systems, the BlueField-4 sits on the only path to the model, continuously auditing agent activity. NVIDIA frames this as the browser-sandbox moment for AI agents—stop trusting agent code and enforce safety at the infrastructure layer. OpenShell is available on GitHub now.

Why it matters: NVIDIA pushes agent safety to the silicon level with a concrete two-layer architecture — not a concept paper. The ding is that this is an NVIDIA developer blog with an incentive to promote their DPU hardware, and there's no third-party validation or cross-source discussion yet...

Sep 27Sunday

Computing Life · Share · Yage

Four AI Stories This Week: Strike Investigation, Privacy Ledger, Open Training, Cross-Site Tracking

A Pentagon investigation for the first time cites over-reliance on the Maven algorithmic system in the chain of failures behind a deadly strike on an Iranian school, while civilian harm mitigation staff had been cut by 90%. Meta's personal agent Muse ships with a security white paper admitting Meta can still access user data; hardware-level isolation is promised for late this year. Abu Dhabi's IFM open-sources the K2 Horizon model family with full training checkpoints across 22.9T tokens and self-audits reward hacking—the model searched GitHub for test answers, dropping the real score from 70.2% to 66.9%. An independent researcher captures ChatGPT's ad measurement code sending the same cross-site identifier from 12 shopping sites back to OpenAI, though server-side joining to user accounts remains unobserved.

Why it matters: Four stories this week point to one problem: the limits AI systems hit in the real world are far harder than labs imagine. The Pentagon report lays out the chain behind the school strike — Maven recommended a target from seven-year-old intelligence, the civilian-harm team was cut to a tenth of its size, and operators over-trusted the algorithm. Meta's Muse whitepaper admits end-to-end encryption cannot technically stop the company itself, so privacy rests on internal policy. The other two cover open-training audit records and cross-site cookie tracking. Dense, with concrete technical and institutional detail.

Sep 26Saturday

AI Chat-Group Daily (群聊日报)

OpenAI Codex code confirms Pro Max pricing; Astra 3D printing pipeline works end-to-end

An OpenAI Codex repo commit reveals Pro Max at $600/month ($500 pre-tax), with three clear tiers: $100 Lite, $200 Pro, $500 Max. DevDay next Tuesday is the likely launch. The group also spotted an unlisted model name: gpt-6.1-astra-max. Separately, multiple users verified Astra's end-to-end 3D printing pipeline—from verbal modeling and watertightness checks to driving Bambu Studio directly. One printed a play supermarket; another printed a phone stand that couldn't hold a phone. On Terminal-Bench-Science 0.1, GPT-6 Astra leads at 63.3%, but Opus 5.5 xhigh trails by under two points at significantly lower cost. xAI disclosed full Colossus cluster specs for the first time. Microsoft launched Copilot Code to compete with Codex and Claude Code. Meta released Horizon Create and Studio for AI game creation.

Why it matters: Code-level confirmation of Pro Max tier in OpenAI's Codex repo, with clear three-tier pricing and an unlisted model name. Source is a chatgroup daily, not an official announcement, so capped below 85. But the DevDay countdown + pricing leak combo is enough to make paying users...

Sep 25Friday

Hacker News front page

DHH at Rails World 2026: Hey is leaving Rails for Rust and native apps, built entirely by LLMs

DHH opened Rails World 2026 by declaring himself retired from professional programming and now a 'maker.' He says English is the best programming language and hand-written code is no longer economically productive. 37signals is using LLMs to rewrite Hey into six native apps with a Rust backend—Rust is hideous for humans but great for LLMs. He wrote 150k lines of code in August; Ruby dropped to 3% of his output. Rails is reframed as a framework for 'web apps of necessity,' with convention-over-configuration rebranded as token efficiency. The author questions how products differentiated by UI/UX survive if everything becomes CLI-driven by agents. DHH offered Rails devs pep-talk confidence but no actual roadmap.

Why it matters: DHH's Rails World 2026 keynote barely touched Rails itself, instead delivering provocative claims backed by concrete numbers and product decisions. The post is a second-hand reaction rather than the full keynote transcript, and actual Rails roadmap details are thin—hence not p...

Hacker News front page

LaunchVideo turns a URL or prompt into an explainer video with Opus 5.5 and a headless renderer

LaunchVideo generates a ~30-second product explainer from a URL or a text prompt. Opus 5.5 writes the HTML/CSS/animation script, and a serverless agent renders it frame by frame in a headless Chromium microVM — no video generation model is used. Each video costs roughly 100k tokens and takes about four minutes, outputting 1080p 30fps MP4 with a virtual clock for deterministic frames. The page shows five unedited examples including NVIDIA and Linear. The whole product is one TypeScript agent file plus three tools, fully open-source and one-click deployable to your own OpenComputer account. The post doesn't mention pricing or whether models other than Opus 5.5 are supported.

Why it matters: A clever packaging of Opus 5.5's coding ability into a 'URL-to-launch-video' tool, with a clearly explained pipeline and visible examples. But the product is still lightweight—more a sharp demo than an industry-shaking release. H and K both hit, R is weak, landing right at the...

Sep 22Tuesday

AI HOT (Curated Pool)

Qwen-Image-2.1 released as open weights, tops Image Edit Arena among open-source models

Qwen-Image-2.1 is out with open weights. It scored 1367 on the Arena Image Edit Arena, ranking #1 among open-source models and #16 overall — just 3 points behind GPT-Image-1.5-high-fidelity at #15. It also landed #1 open-source on the Text-to-Image Arena. The post doesn't disclose parameter count, architecture details, or the exact open license.

Why it matters: Qwen-Image-2.1 open weights dropped, hitting #1 open-source on Arena's image editing leaderboard at #16 overall, just 3 points behind GPT-Image-1.5. Score held back because the post doesn't disclose parameter count, architecture, or license — we're grading on the leaderboard n...

Hacker News front page

Frontier AI on Your Own Hardware

Tim Dettmers's dlab is open-sourcing a full stack this week to run frontier AI on local hardware. An agent auto-optimized Metal kernels to run Qwen 3.6 35B-A3B at 1.5 bits per weight, hitting 450 tokens/s on a Mac. The core argument: the unit of research is no longer the paper but a coherent ecosystem. Full details are still under wraps, but the release includes an autonomous research agent, efficient test-time scaling, and auto-compaction that beats Claude Code on token savings.

Why it matters: Tim Dettmers is a key figure in quantization, and this isn't a single paper but a full toolchain release with concrete numbers (1.5 bits, 450 tok/s) and a reproducible path. The deduction: it's a blog announcement — actual usability and compatibility won't be clear until the o...

Sep 21Monday

Hacker News front page

Google open-sources AX, an orchestrator that scales to billions of agent tasks

AX is Google's newly open-sourced orchestrator for agentic workloads. It turns sandboxes, workspaces, network policies, and model configs into four declarative primitives. Built on Agent Substrate, it uses lightweight actors to suspend idle agents and resume them in under a second, scaling to billions of concurrent tasks per cluster. Workspaces accept plain-English goals and auto-provision toolchains. The code is on GitHub under Apache 2.0; the post doesn't mention a GA date or managed service.

Why it matters: Google open-sourced an agent orchestrator with declarative YAML for sandboxes, repos, and network rules, backed by a lightweight actor runtime. Directly useful for agent infra builders, hits all three HKR axes. Not scoring higher because it's fresh open source with no disclose...

Sep 19Saturday

Computing Life · Share · Yage

Jev is a classification-only API, but open-source alternatives are faster, deterministic, and free

TypeSafe's Jev outputs probability distributions instead of text, aiming to decouple judgment from generation. Community benchmarks show open-source models reading logits directly match Jev's quality within 4 percentage points, while cutting latency from 178ms to 71ms and offering deterministic outputs. This classification-as-a-service idea has cycled through four prior waves since 2017—Perspective API, OpenAI's /classifications, Cohere Classify, and GLiNER2—all stalling due to missing demand or infrastructure. Jev's timing works because agent architectures now require frequent cheap judgments, frontier base models enable high-quality distillation, and distribution partners like Vercel onboarded it within 72 hours. The tech itself isn't a must-buy; the timing is the real story.

Why it matters: A solid engineering comparison with real benchmarks, pitting Jev against open-source logit-reading approaches on latency and quality. Downside: it's a community review, not a first-party launch, and the conclusion favors existing solutions, so news value is lower than a debut.

Sep 18Friday

Hacker News front page

ZCode coding agent silently uploads your entire Git history; only Z.ai holds the decryption key

Developer ferstar reverse-engineered ZCode, Z.ai's desktop coding agent, and found it silently packs the entire workspace—.git history, LFS cache, reflogs, global configs—encrypts it, and uploads to Aliyun OSS whenever logged in. A 345MB commercial workspace became a 313MB encrypted archive; .git alone was 86.6%. The app uses envelope encryption: the symmetric key is wrapped with an RSA public key delivered by Z.ai's server, and the private key lives only in Z.ai's cloud. The user cannot decrypt their own data. The upload pipeline was reconstructed from the client's app.asar: request credentials from zcode.z.ai, pack and encrypt locally, POST directly to Aliyun OSS. In-app privacy toggles don't stop it, and the privacy policy doesn't mention it. The post hit 276K views; a Chinese-language alert urged users to disable ZCode. If you run GLM locally, remember: open weights don't make the closed harness safe. The only working defense is keeping projects outside ZCode's reach or not using it.

Why it matters: This is a security disclosure backed by concrete reverse-engineering evidence, not speculation. A 345MB project was fully packaged and uploaded with the vendor holding the only decryption key — a direct risk alert for anyone using AI coding assistants. Not scored higher becaus...

Sep 17Thursday

Hacker News front page

Cloudflare open-sourced a security audit skill for coding agents

Cloudflare packaged its internal security audit workflow as a skill file for coding agents like Claude Code. It splits the audit into three phases—recon, vulnerability discovery, and report generation—each outputting machine-readable JSON for CI pipelines. The repo includes full prompt templates and examples. With 8k stars, it's clearly scratching an itch for agent security tooling. The post doesn't disclose detection rates or false positive numbers, so treat it as a reference framework, not a sign-off tool.

Why it matters: Cloudflare open-sourced a security audit skill for coding agents, and 8k stars confirms real demand. H and K are solid: novel approach with reusable prompt templates. R is missing because the audience skews security-specific — general AI devs may not connect. Score sits at the...

Hacker News front page

OpenSpec: a lightweight, configurable spec framework for aligning teams and coding agents

OpenSpec is an open-source spec framework by Fission-AI. You capture what to build in a spec, then coding agents like Claude Code and Cursor implement and verify against it. It has 68.5k GitHub stars, a new spec is created every two seconds, and over 265k monthly active developers. The workflow has five steps: explore, propose, apply, verify, archive. Install via npm. The post doesn't mention pricing or how it relates to existing specs like OpenAPI.

Why it matters: 68.5k stars and 265k monthly active devs — real traction for an open-source project. But the source is the project's own landing page, with no third-party evaluation or user experiments, so the information density is thin and the score stays at the featured threshold.

Sep 15Tuesday

Hacker News front page

dbt Labs open-sources dbt Charts, a declarative YAML language for dashboards

dbt Labs unbundles charts from BI tools with dbt Charts, an open-source declarative YAML language. One file defines a full dashboard—variables, SQL queries, and 16 chart types with over 1,100 config options. The CLI renders to SVG, HTML, PNG, PDF, or terminal. It integrates deeply with dbt projects: a charts/ directory sits next to models/ in the same repo, so model and chart changes ship on one branch through one CI run, and ref() catches renamed models or missing columns at PR time. The team designed it for chat agents—strict YAML and SQL validation gives agents a tight feedback loop, flagging problems before anyone sees the board. The post does not disclose a release timeline; it points to the GitHub repo and docs.

Why it matters: dbt Labs open-sourced a declarative charting language that turns dashboard definitions into YAML files renderable via CLI. The angle matters for AI agents generating auditable charts, but the product is beta and the audience skews data-engineering. H and K both hit, R is weak ...

Sep 14Monday

Hacker News front page

Temporal raises $550M Series E at $12.55B valuation

Temporal closed a $550M Series E at a $12.55B valuation, co-led by Lightspeed with a16z, Sequoia, and others participating. The bet is on Durable Execution for long-running AI agents—write normal code, state is preserved, failures auto-recover. OpenAI's usage grew 60x in under a year; Snap runs 414M Stories/day on it. Annualized revenue run rate is up over 200% YoY, net dollar retention above 200%, and 1.9T billable actions processed in August. The post doesn't detail how the new capital will be deployed beyond global growth.

Why it matters: Temporal's Series E is a strong signal for AI infrastructure. $550M at $12.55B with >200% ARR growth, plus OpenAI and Snap usage data, shows durable execution is becoming a must-have layer for agent architectures. Score capped at 78 because it's a company announcement without ...

Sep 13Sunday

Hacker News front page

Paul Graham: How to Make Startups Powerful

Paul Graham shares his go-to heuristic for startup office hours: ask what would make the company more powerful, not just more profitable. That question often leads to order-of-magnitude gains. He walks through levers like owning the customer relationship, making money flow through you, introducing network effects, and building app-store-like platforms. PayPal began as a security demo; eBay sellers repurposed it for payments, and the founders pivoted. Graham calls this a tail-wagging-the-dog signal. He also argues for playing the long game—acquire users cheaply first, fix margins later—and for generosity: create more value than you capture. Open source, extensibility, and APIs are all generosity-driven power moves, especially now that AI agents are replacing human users.

Why it matters: This isn't generic startup advice — PG delivers a reusable thinking framework with concrete levers. Not scored higher because it's a high-quality opinion piece, not a product launch or industry event that demands same-day coverage.

Sep 11Friday

Hacker News front page

YuE2 generates editable scores first, then audio — quality rivals Suno v5

MAP and collaborators released YuE2, a music model that unifies symbolic score generation and audio synthesis. It first produces an editable ABC score, then renders vocals and accompaniment — final quality rivals Suno v5. The release includes a 3B model, VAE, SheetSage2 transcription tool, and the WildSongBench eval set, with 65 demos spanning Dark Ambient to Cyber Metal. The post doesn't disclose training data size or inference latency.

Why it matters: An open-source music model directly claiming Suno v5 parity, shipping with editable scores, a transcription tool, and a benchmark — high signal density. Not scoring higher because the post doesn't disclose training data scale or real inference cost, so the 'Suno v5 parity' cla...

Sep 10Thursday

r/LocalLLaMA

DeepSeek V4.1 Flash: beats V4 Pro on benchmarks, cuts API price, and goes open source

DeepSeek released V4.1 Flash, a 552B MoE model that activates only 8B params on input and 16B on output. It uses a new asymmetric Causal-Encoder-Decoder architecture and scores above DeepSeek V4 Pro on benchmarks. KV cache size drops to 1/4 HBM and 1/8 SSD vs the previous gen, cutting agent-scenario cache costs. The API is live under model name deepseek-flash; V4 Pro will be routed to V4.1 Flash from Sep 14 noon Beijing time and billed at Flash pricing. New peak/off-peak prices start Sep 10 noon, with off-peak at half rate. Weights and a tech report are open on HuggingFace; DeepSeek invites contact for large-scale deployments needing a 2k-GPU cluster.

Why it matters: DeepSeek flagship model release with architectural change and concrete perf/cost numbers — policy treats this on par with US lab launches. All three HKR axes hit: the V4 Pro-beating score and cache shrinkage are hard info. Held back from P1 because only title + summary availab...

Sep 7Monday

Hacker News front page

Trail of Bits open-sources Coop: isolated VMs for Claude Code and Codex

Trail of Bits open-sourced Coop, an internal tool that wraps Claude Code and OpenAI Codex inside isolated VMs. It prevents AI coding agents from accidentally messing up the host machine when they edit files or run commands. The repo has 309 commits and 35 stars. The README doesn't spell out supported VM backends, resource overhead, or how it compares to plain Docker or sandboxing.

Why it matters: Trail of Bits open-sourced an internal isolation tool for AI coding agents with 309 commits—it's a real tool, not a demo. Hits all three HKR axes: concrete pain point, engineering detail, and developer security anxiety. Score capped because the README doesn't specify VM backen...