Skip to content

Open source

Open models, frameworks and repositories: open weights, community hits and the balance between open and closed.

Latest picks

241–260 of 329

May 10Sunday

r/LocalLLaMA

We tried vectors, ASTs, and brute-force context stuffing for code retrieval; LLM semantic graphs worked best

ByteBell open-sourced a code indexing system that stores per-file LLM-generated purpose, summary, business context, entities, classes, functions, keywords, and imports in a Neo4j graph, then uses full-text search instead of vector similarity, with SHA-256 diffing to reindex only changed files and keep LLM calls proportional to churn.

Why it matters: HKR-H/K/R all pass: the hook is counterintuitive, and the post gives a concrete Neo4j semantic-graph mechanism with SHA-256 incremental rebuilds. Reddit sourcing and missing metrics keep it at the 72–77 featured threshold.

r/LocalLLaMA

I have DeepSeek V4 Pro at home

Reddit user fairydreaming ran DeepSeek V4 Pro Q4_K_M with a modified llama.cpp CUDA repo on one RTX PRO 6000 Blackwell Max-Q workstation GPU, using an 859GB model file; the shared log reports a 1M context window and 8.6 tokens per second generation speed.

Why it matters: HKR-H/K/R all pass: the hook is single-GPU local inference, with concrete file size, context, speed, and runtime path. Reddit single-source sourcing keeps it below must-write model-release territory.

r/LocalLLaMA

BeeLlama.cpp: DFlash and TurboQuant with reasoning and vision support

Anbeeld released BeeLlama.cpp, a llama.cpp fork that runs Qwen 3.6 27B Q5 with 200k context and vision on a single RTX 3090 or 4090; the title claims 2–3x faster than baseline and a 135 tps peak.

Why it matters: HKR-H/K/R all pass, but the claims come from a Reddit title and summary without independent reproduction. Treat as a mid-weight open-source inference update, so it lands in the low featured band.

May 9Saturday

AI HOT (Curated Pool)

YC CEO Open-Sources Personal AI OS GBrain for a Compounding Second Brain

Y Combinator CEO Garry Tan open-sourced GBrain, a personal AI operating system that processed more than 20 books in five months and manages over 100,000 pages of structured knowledge.

Why it matters: HKR-H/K/R pass: Garry Tan’s open-source personal knowledge system has a notable-user hook and three concrete usage numbers. Missing repo activity, architecture detail, and tests keep it at the featured threshold.

AI HOT (Curated Pool)

Redis founder uses a C inference engine to run a large model on a personal computer

Antirez open-sourced ds4, a native inference engine for DeepSeek V4 Flash that uses a few thousand lines of C to run a 1M-context model on a 128GB MacBook Pro at a reported 27 tok/s.

Why it matters: HKR-H/K/R all pass: Antirez open-sourced a native C inference engine with hardware, model, context, and speed numbers. Single-source X provenance keeps it below P1, but it is strong open-source inference signal.

Xinzhiyuan · WeChat

CUHK Open-Sources ArbiterOS Agent Governance Kernel With 92.95% High-Risk Interception

CUHK CURE Lab open-sourced ArbiterOS, an agent runtime governance kernel that intercepts, parses, governs, and observes actions before execution, raising high-risk step interception on OpenClaw tasks from 6.17% to 92.95%.

Why it matters: HKR-H/K/R all pass: the story has a sharp execution-control hook, a concrete 6.17%→92.95% result, and clear agent-safety resonance. It is a strong open-source research tool, not a top-lab model release, so it stays in the 78–84 band.

Synced · WeChat

StarVLA Open-Sources a Unified VLA Framework from HKUST and the Community

HKUST and the open-source community released StarVLA, a unified Vision-Language-Action framework that integrates backbones, action heads, training strategies, and evaluation interfaces; the repository has 2.2k GitHub stars and supports benchmarks including LIBERO, SimplerEnv, RoboTwin 2.0, RoboCasa-GR1, and BEHAVIOR-1K.

Why it matters: HKR-H/K/R all pass: StarVLA ships a concrete open-source VLA framework with unified interfaces, 2.2k stars, and named robotics benchmarks. The robotics scope keeps it in the 78–84 band, below model-release weight.

r/LocalLLaMA

MTP + TurboQuant Running: Qwen3.6-27B Hits 80+ t/s on a Single RTX 4090

indrasmirror ran Qwen3.6-27B-Heretic-v2 on a single RTX 4090 with 262K context, TBQ4_0 KV cache, and MTP draft 3, improving throughput from about 43 t/s to 80-87 t/s with roughly 73% MTP draft acceptance.

Why it matters: HKR-H/K/R all pass, backed by a numbered first-person experiment. The Reddit-only source and niche local-inference focus keep it below the 78–84 band for broader industry releases.

May 8Friday

Hacker News front page

Show HN: Git for AI Agents

regent-vcs released the open-source re_gent project for AI-agent version control, currently supporting Claude Code, with workflows for tracking why an agent changed files, rewinding sessions, and bisecting agent actions; the post does not disclose the license, storage format, or installation details.

Why it matters: HKR-H/K/R all pass: the Git analogy is clicky, the mechanism is concrete, and Claude Code rollback pain is real. The post lacks license, storage format, and install details, so it stays at the featured threshold.

AI HOT (Curated Pool)

Donating the Open-Source Alignment Tool Petri

Anthropic transferred the open-source alignment testing tool Petri to Meridian Labs to preserve independence and credibility. Petri 3.0 separates auditor and target models, adds Dish for real prompts and deployment settings, and integrates Bloom.

Why it matters: HKR-H/K/R all pass: the independent donation is a real hook, Petri 3.0 and Dish add testable mechanisms, and audit credibility resonates. Anthropic open-source safety tooling is strong, but below a model-release-level event.

AI HOT (Curated Pool)

DeepSeek 4: Flash Local Inference Engine for Metal

DeepSeek 4 Flash is open-sourced on GitHub for offline inference on Apple Silicon Macs. The post says it uses Metal Performance Shaders to reduce latency and memory use, but discloses no benchmark numbers. The key item is the Metal local inference stack, not another model wrapper.

Why it matters: HKR-H/K/R pass: the hook is offline Apple Silicon inference, with GitHub OSS, MPS, and a clear run target. No latency or memory benchmarks, and not an official DeepSeek model launch, so it stays near the featured floor.

May 7Thursday

AI HOT (Curated Pool)

SenseNova-U1 Open-Sources 8-Step Distilled LoRA, Speeds Diffusion Inference by 11x

SenseNova-U1 open-sourced an 8-step distilled LoRA that cuts diffusion generation from 100 steps to 8. GPU inference time drops from 23 seconds to 2 seconds, with ComfyUI workflows for text-to-image, image editing, and interleaved generation. The key signal is distillation for latency, not parameter scale.

Why it matters: HKR-H/K/R all pass: the 11x speedup hooks attention, the post gives step and latency numbers, and open LoRA affects diffusion deployment cost. Scope stays within image generation, so this is featured, not P1.

r/LocalLLaMA

Running Qwen3.5/Qwen3.6 with NextN MTP in llama.cpp on one RTX 3090 Ti

A Reddit user posted a llama.cpp guide for Qwen3.5/3.6 with NextN MTP on one RTX 3090 Ti. It requires two unmerged PRs, #22400 and #22673; Qwen3.6-35B-A3B-MTP reaches 157 tok/s at 350W, 1700MHz, with q8 KV. The key reproducible detail is nextn=q8_0 quant override; missing it yields “////” output.

Why it matters: HKR-H/K/R all pass: single-GPU 157 tok/s is a strong hook, and the PR/power settings make it testable. Scope stays narrow because it is a Reddit guide using unmerged PRs.

AI HOT (Curated Pool)

Open Slide lets AI write PPT code

Open Slide builds PPTs with React, using a workflow designed for AI agents. It integrates SVGL with 1,500+ brand logos, supports manual edits, and lets AI read user comments for revisions.

Why it matters: HKR-H/K/R pass: the programmable-slide angle is clickable, with concrete React and 1500+ logo details, and deck work is a real practitioner pain. No usage metrics or hands-on test keeps it at the featured threshold.

r/LocalLLaMA

GB10 inference engine Atlas is open source, with Qwen3.6-35B-FP8 over 100 tok/s

Avarok open-sourced Atlas, an inference engine running Qwen3.5-35B at ~111 tok/s sustained on one DGX Spark. It uses Rust+CUDA, a ~2.5GB image, and sub-2-minute cold start; the author claims 3.0–3.3x vLLM in tests. The key details are Blackwell SM120/121 kernels, NVFP4/FP8, and MTP decoding.

Why it matters: HKR-H/K/R pass: open-source inference engine, 35B FP8 at 111 tok/s, and a direct vLLM comparison. Single Reddit sourcing and unreproduced benchmarks keep it at the lower featured band.

May 6Wednesday

r/LocalLLaMA

CopilotKit (MIT): Open-source building blocks for agent apps and generative UI

CopilotKit offers MIT-licensed React components and claims 30k GitHub stars. It covers chat, streaming, tool calls, HITL, and generative UI, with AG-UI support for LangGraph, CrewAI, LlamaIndex, and other backends. The key point is decoupling the UI layer from agent frameworks.

Why it matters: HKR-H/K/R all pass: MIT open source, 30k stars, and AG-UI links to major agent backends. Kept in 72–77 because the post lacks a new version, benchmark, or named production adopter.

r/LocalLLaMA

2.5x Faster Inference with Qwen 3.6 27B Using MTP on 48GB

A llama.cpp PR adds MTP support for Qwen 3.6 27B, with a reported 2.5x inference speedup. The author measured 28 tok/s on a Mac M2 Max 96GB and shared GGUF builds, compile steps, and a 262144-context server command. The key detail is turbo4 4.25-bit KV cache: a 48GB Mac runs Q5_K_M at 262K context.

Why it matters: HKR-H/K/R all pass: the hook is concrete, the post names mechanisms and numbers, and local coding-agent cost resonates. Single Reddit source and setup complexity keep it in the low featured band.

Synced · WeChat

Alibaba open-sources PromptEcho for T2I rewards using frozen VLMs

Alibaba open-sourced PromptEcho, which uses one frozen Qwen3-VL-32B forward pass to score T2I training rewards. It computes token-level cross-entropy for the original prompt under teacher forcing, then uses the negative value as a continuous reward. In 5,000 poster tests, text accuracy rose from 68% to 75%.

Why it matters: HKR-K is strong: the post gives a concrete reward mechanism and a 68%→75% text-accuracy result. HKR-H/R pass, but this is a training-side research release, not a flagship model or major product update.

Synced · WeChat

DeepSeek Version of Claude Code Tops Trending Chart With 8,700 Stars

DeepSeek TUI topped GitHub trending with over 8,700 stars. Hunter Bown built it in Rust for local terminal use with DeepSeek V4, supporting chat, file edits, shell commands, and task management. The key detail is RLM mode: up to 16 V4 Flash subtasks, plus a 1M-token context window and approval gates.

Why it matters: HKR-H/K/R all pass: the 8,700-star hook is strong, RLM adds concrete mechanisms, and coding-agent competition resonates. It is a third-party open-source tool, not an official DeepSeek model release, so it stays in the 78–84 band.

Synced · WeChat

Two Chinese open-source projects turn Mac into a private AI workstation

Mininglamp open-sourced Cider and Mano-P 1.0 for Apple Silicon local inference and GUI agents. Cider speeds Qwen3-VL-2B prefill by 57%–61% on M5 Pro; Mano-P 1.0-72B scores 58.2% on OSWorld. The key constraint is W8A8 memory: on 16GB devices accuracy falls from 58.0% to 54.0%, so 32GB+ is recommended.

Why it matters: HKR-H/K/R all pass: the Mac-local workstation angle is clickable, and Cider/Mano-P include testable numbers. Score stays at 80 because the source entity is not a top-tier model lab.