Skip to content

Open source

Open models, frameworks and repositories: open weights, community hits and the balance between open and closed.

Latest picks

61–80 of 329

Aug 14Friday

Hacker News front page

GLM-5.3: Post-training-only gains push open-weight coding and exploit capability to the top

Z.ai released GLM-5.3 with the same base model as 5.2 — every gain is from post-training. Coding jumped 50% on their internal Z.ai Code Bench, and Terminal Bench 3.0 went from 4.6 to 28.3. The bigger surprise: exploit capability grew far faster than expected. ExploitGym 2h score rose from 29 to 105, 6h from 39 to 130. The team credits training environments that mirror real expert workflows, pushing the model to chain full exploit sequences. Weights will be open-sourced in two weeks after safety hardening.

Why it matters: Zhipu releases GLM-5.3 — same base model as 5.2, all gains from post-training. Code bench up 50%, Terminal Bench from 4.6 to 28.3, 2-hour exploit score from 29 to 105. The lab admits cyber capability emerged faster than expected. Domestic flagship model launch with concrete nu...

Aug 13Thursday

Latent Space

xAI drops Grok 4.6 and Grok Bot, a strong new entrant in the AI teammate race

xAI launched Grok 4.6 and the Grok Bot early beta. Grok Bot logs into your tools, operates them like a human, and returns finished work—positioned as an AI teammate. The 1.5T-parameter Grok 4.6 scores near GPT-5.6 Sol Max on the AA-Briefcase knowledge-work benchmark but costs far less: $2/M input tokens, $6/M output. Training reused Grok 4.5 to regenerate SFT traces and added agentic RL across coding, web, CAD, and kernel optimization. Elon says Grok 4.7 is already training. The same day, Qwen3.8-Max dropped as open weights: a 2.4T total / 95B active MoE.

Why it matters: Grok 4.6 matches GPT-5.6 Sol Max on a knowledge-work benchmark at an order-of-magnitude lower price, while the simultaneously launched Grok Bot enters the AI teammate race built by the ex-Cursor team with positive early feedback. Score isn't higher because the Bot is still in ...

Computing Life · Share · Yage

DeepSeek open-sources DSH: agent loop as a hot-swappable plugin, paving the way for self-evolving agents

DeepSeek released its first agent harness, DSH, as open source on August 13. Unlike Codex or Claude Code, DSH treats the agent loop itself as a plugin that can be swapped at runtime. The Cordis runtime handles hot reloads, dependency notifications, and transactional rollbacks. For everyday coding, declarative plugins plus a quick restart are enough—DSH's imperative model adds complexity. But if you want an agent that can generate new tools or replace its own control flow mid-run, DSH is the only option with the infrastructure in place. The post does not disclose performance benchmarks or production-scale data.

Why it matters: DSH makes the agent loop itself a hot-swappable plugin — a real architectural difference, not marketing. But this is a third-party analysis, not an official launch, and DSH has zero production track record yet. Defaulted to the lower band per policy.

AI HOT (Curated Pool)

Alibaba open-sources Qwen3.8-2.4T-A95B: 2.4T MoE, 95B active, native 256K context

Alibaba's Qwen team open-sourced its first Qwen-Max-level weights. Qwen3.8-2.4T-A95B has 2.4T total parameters with 95B active per token, native 262K context expandable to 1.01M tokens. It uses a 512-expert MoE, routing 10 experts plus one shared expert per token, and includes multi-token prediction training. The model targets coding, office tasks, research, and long-horizon agent workflows. Benchmarks against Opus 4.8, Fable 5, and GPT 5.6 Sol show mixed results, with top scores on PaperBench and IFBench among listed models. Post-training combines combinatorial environment scaling, a unified reward system, and an online data balancer to reduce gradient variance. The post does not disclose the open-source license or inference hardware requirements.

Why it matters: Alibaba's first full open release of a cloud-grade flagship — 2.4T total params, 95B activated, native 256K context — puts it in the top tier. Hits all three HKR axes and triggers the domestic flagship model positive signal. Held back from 90+ because we only have the announce...

Aug 12Wednesday

Hugging Face Blog

Liquid AI releases LFM2.5-VL-3B, a vision-language model for edge devices

Liquid AI open-sourced LFM2.5-VL-3B, a 3B-param vision-language model that runs on local hardware. It skips long reasoning chains and answers directly, targeting real-time and on-device use. Four main upgrades: screen/UI understanding, natural-language object grounding, multi-image reasoning, and stronger function calling. The post includes benchmark comparisons and CPU/GPU inference speed, but doesn't give exact latency numbers.

Why it matters: Liquid AI ships a 3B vision model tuned for local inference with four concrete capability upgrades and benchmarks. H and K both hit, but Liquid AI lacks brand pull in the vision space so R is absent — score lands right at the featured threshold.

Hacker News front page

NVIDIA ships Nemotron 3.5 Lightning and NeMo Switchyard for faster, smarter agent routing

NVIDIA added a 30B-parameter MoE model, Nemotron 3.5 Lightning, to its Nemotron 3 family. It targets specialized tasks inside multi-agent systems, delivering 4x faster output and 30% faster agentic task completion than peers. It runs locally on RTX PCs, DGX workstations, and Jetson. The company also open-sourced NeMo Switchyard, a routing library that directs requests to the best model for each job without app rewrites. CrowdStrike, Harvey, and CodeRabbit are already using customized versions. The post does not disclose pricing or a release timeline.

Why it matters: Nvidia released a 30B MoE model positioned as a specialized worker in multi-agent systems, not a general-purpose model. The 4x output speed and 30% task acceleration claims are useful references, and Switchyard is open-sourced. But this is Nvidia's own blog with no third-party...

Aug 11Tuesday

Mistral AI

Mistral launches regional inference endpoints and a Priority Tier, adding third-party open models like GLM-5.2

Mistral announced general availability of Mistral Regional Endpoints, letting customers choose whether inference runs in Europe or the US. Mistral Priority Tier also entered public preview, offering custom rate limits and an availability commitment backed by an SLA.

Why it matters: Mistral puts regional inference endpoints, an SLA service tier and third-party open models on one infrastructure stack, a read on how European sovereign AI is being delivered.

Aug 10Monday

Hacker News front page

Zuckerberg attacks closed AI rivals as Meta returns to open models

Zuckerberg called out OpenAI and Google by name in an internal meeting, arguing open models will win long-term. He confirmed Meta's next Llama generation will stay fully open and said AI teams are merging into product units to speed up shipping. No release date or specs were disclosed.

Why it matters: Zuckerberg's internal talk calls out OpenAI and Google by name, confirms Llama stays fully open-source, and reveals AI teams are being merged into product groups. Conflict, org change, and a clear stance hit all three HKR axes. No timeline or specs disclosed, so it lands at 78...

AI HOT (Curated Pool)

SGLang adds Day-0 inference support for Meta's local agent model Muse Glimmer

Meta released Muse Glimmer, a 30B multimodal model built for local agentic workflows. SGLang ships Day-0 support with dedicated optimizations: on a single RTX 5090 with NVFP4 quantization and DFlash speculative decoding, per-user decode hits 236 tok/s and total throughput reaches 1,452 tok/s. The model uses a hybrid of sliding-window and full-sequence attention with a 128k+ context window. Apple Silicon is supported via the MLX backend, though speculative decoding isn't available there yet.

Why it matters: Meta shipping a new model is an industry event, but this post centers on SGLang's inference optimization, not the model itself. Concrete perf numbers (236 tok/s, 1452 tok/s total throughput) give it enough knowledge density to clear the featured bar, though the narrow audience...

Hacker News front page

Meta open-sources Muse Glimmer, a 30B agentic model that runs locally on a single GPU

Meta released Muse Glimmer weights under Apache 2.0. It's a 30B model built for always-on local agent workflows, small enough to run on a Mac or PC with a single consumer GPU. 4-bit quantization shrinks it below 20 GB, leaving room for KV cache and the vision encoder within a 24 GB or 32 GB envelope. Training used logit distillation from a larger Muse Spark teacher, followed by mid-training on long-context agent data and post-training with SFT, on-policy distillation, and RL. Meta's benchmarks show it outperforming Gemma4-31B and Qwen3.6-27B on agentic, coding, multimodal, and safety evals. The post doesn't disclose specific latency numbers, only that inference optimizations were applied to keep it responsive.

Why it matters: Meta drops a 30B local agent model under Apache 2.0, quantized under 20 GB for consumer GPUs. Clear positioning — not a general chatbot but purpose-built for always-on agent workflows. Score held back from higher bands because we only have the launch blog; third-party benchmar...

AI HOT (Curated Pool)

NVIDIA Releases NemotronLabs VoiceChat 11B: An Open Full-Duplex Speech-to-Speech Model with ~450 ms Turn-Taking and Live Tool Calling

NVIDIA open-sourced an 11B end-to-end speech-to-speech model that handles streaming understanding and generation in one network, skipping the usual ASR-LLM-TTS pipeline. Measured turn-taking latency is 448 ms, and it supports live tool calling during conversation. Weights and code are public.

Why it matters: NVIDIA open-sourced an 11B end-to-end speech model with 448 ms interruption latency and full-duplex turn-taking — real engineering progress, not a benchmark flex. But the MarkTechPost piece reads like a product announcement with no third-party benchmarks or head-to-head compar...

Aug 8Saturday

AI HOT (Curated Pool)

Tencent Hunyuan open-sources production inference kernels and integrates them into SGLang

Tencent Hunyuan open-sourced its production-proven Attention, Router GEMM, and MoE kernels as HPC-Ops and merged them into SGLang's main branch. On H20, Hy3-FP8 TPOT dropped by up to 48.8%; the Attention kernel averaged 2.25× faster than FlashInfer/FlashAttention, and Router GEMM ran 1.30–3.22× faster than FP32 cuBLAS with lower error. H200 validation passed, with the MoE kernel hitting 4.21× over Triton on Qwen3.

Why it matters: Tencent Hunyuan open-sourced production-hardened Attention, Router GEMM, and MoE kernels and merged them into SGLang main. Three concrete optimizations with reproducible detail — not vague 'we made it faster' claims. High practical value for inference engineers, but audience i...

Hacker News front page

Oracle bans AI-generated code from OpenJDK while Ellison says AI writes Oracle's own code

Oracle told OpenJDK contributors: use LLMs privately for debugging and review, but don't submit AI-generated code to repos or pull requests, citing safety, security, and IP risks. That stance clashes with Larry Ellison's recent claim that AI models now write Oracle's code, and co-CEO Mike Sicilia's praise of AI tools for letting smaller teams ship faster. Oracle is spending $70 billion this year on datacenter expansion, which prompted S&P to downgrade its rating to BBB-, one notch above junk, over uncertain returns.

Why it matters: A sharp public contradiction between Oracle's internal messaging and its open-source community policy, with concrete details. Not pushed to 85+ because only a single Dealroom source so far — no direct statement from Oracle or OpenJDK maintainers yet. Settled at 78.

Aug 7Friday

Hacker News front page

Herdr joins Y Combinator F26, open-source agent runtime stays free

Herdr is a runtime and TUI for CLI coding agents, letting agents run persistently across machines. Solo developer Can Celik built it to 25k GitHub stars and 340k downloads, and is now joining Y Combinator's F26 batch to turn it into a company. The runtime stays free under Apache-2.0, the core remains small, and everything else goes through plugins—over 500 already exist, plus community-built Raycast, Stream Deck, and iOS clients. The post doesn't disclose funding amount or team size.

Why it matters: Solo dev takes a CLI agent runtime to 25K stars and 340K downloads, then joins YC F26 while keeping the core Apache-2.0 open. The story has contrast, numbers, and an open-source commitment that resonates with agent builders. Downside: it's a single blog post, not a product lau...

AI HOT (Curated Pool)

Agent Plugins 1.0.0: Google, Amazon, Microsoft, and others ship a unified agent plugin spec

Agent Plugins 1.0.0 is an open, vendor-neutral spec that packages Agent Skills and MCP servers into a portable directory. Google joins Amazon, Cursor, Microsoft, OpenAI, and Vercel as a core maintainer. The format is deliberately minimal: plugin.json declares only a name and schema, skills live in skills/, and MCP servers go in mcp.json with explicit transport types. v1 intentionally omits install mechanisms, permission models, and sandboxing—those are left to each client. The post also notes that a single skill or single MCP server doesn't need a plugin; the format earns its keep when components must travel together.

Why it matters: Five major players jointly shipping a unified agent plugin spec — strong cross-source signal with real ecosystem impact. Capped at 78 because it's a spec release, not a runnable product; adoption remains to be seen.

Aug 6Thursday

Hacker News front page

Prime Intellect open-sources Prime Agent, a coding harness that lets models manage their own context, tools, and sub-agents

Prime Intellect released Prime Agent, an open-source coding harness where models treat context as variables and sub-agent calls as functions inside a persistent IPython kernel. Two core abstractions drive it: RLM gives the model programmatic access to its own history and tools, while Continual Harness lets the agent create, update, and delete its own prompts, skills, and memory at runtime. A background daemon manages all sessions with attach/detach, crash recovery, and agent-to-agent messaging. The repo is public on GitHub and installs with a single curl command.

Why it matters: Prime Intellect open-sourced a code agent framework with a clear architectural hook: models managing their own memory and prompts inside a persistent environment. H and K both hit, but R is weak — Prime Intellect isn't a tier-1 lab, so the identity resonance is limited. Meets ...

Aug 5Wednesday

Hacker News front page

A char-level transformer trained from scratch on an $8 ESP32-S3

Carloscodix open-sourced qapla, a char-level transformer trained entirely on an $8 ESP32-S3 chip. The chip runs the full training loop, not just inference, with backprop hand-written in C. The repo doesn't disclose parameter count, dataset size, or training time, but the code is public with 6 commits. Treat this as an extreme engineering demo rather than a practical tool—fitting training onto a microcontroller is the wild part.

Why it matters: Running a full training loop on an $8 microcontroller is a hardcore engineering demo — H and K both hit. But the repo has only 6 commits, no param count, dataset size, or training time disclosed, so practical value is limited and R misses. Scored 72 at the featured threshold, ...

Hacker News front page

Flowise is shutting down; repo to be archived in August

Flowise is winding down operations, citing a shift toward coding agents like Claude Code that handle complexity better than rigid low-code workflows. Active development stops July 29, the GitHub repo will be archived on August 10, and official support ends August 31. The Apache 2.0 code remains available for anyone to fork and maintain.

Why it matters: Flowise is a flagship open-source project in the low-code AI agent space. Its shutdown announcement directly attributes the cause to the rise of coding agents like Claude Code — essentially using its own death as a footnote for an industry trend. HKR all hit, but this is ecosy...

AI HOT (Curated Pool)

LLM 0.32 adds reasoning traces, OpenAI Responses, server-side tools, and smarter logging

Simon Willison shipped LLM 0.32, the biggest update since launch. Reasoning traces now stream to stderr so you can pipe clean output elsewhere. The new default model is GPT-5.6 Luna. Server-side tools like OpenAI's code interpreter and web search are supported, and the Anthropic plugin adds matching tools plus an MCP connector. The Python API drops the forced conversation abstraction—you pass a messages list directly and use stream_events() to separate reasoning, text, and tool calls. Logging switches to a Git-like content-addressable store to avoid duplicating long contexts.

Why it matters: LLM 0.32 is a substantial release with developer-facing improvements that actually matter — reasoning trace isolation and content-addressable logging are real quality-of-life upgrades. Not scored higher because it's a tooling-layer update, not a model capability or industry sh...

Aug 4Tuesday

Hacker News front page

Soup: fine-tune an 8B model on a 4 GB laptop GPU with one config and one command

Soup wraps LLM fine-tuning into a single-command workflow that runs on as little as 4 GB of VRAM. It uses QLoRA with Unsloth acceleration and supports mainstream 8B models like Llama, Mistral, and Qwen. You prepare a JSONL dataset, write a YAML config, and run `soup run`. A built-in Streamlit UI lets you test the model in the browser. The README doesn't disclose exact training speed or memory numbers, but it emphasizes that everything works on a regular laptop GPU.

Why it matters: A practical tool that lowers QLoRA fine-tuning to 4 GB VRAM and one command. Hits H and K, but it's a solo dev's fresh project with no community validation, so R is absent — lands right at the featured threshold of 72. If real-world benchmarks or community feedback appear late...