Skip to content

AI coding

Everything about AI writing code: coding assistants, vibe coding, code model evals and new developer workflows.

1,196 picksRelated topicsAgentsCursorTutorials

Latest picks

941–960 of 1,196

May 9Saturday

r/LocalLLaMA

80 tok/sec and 128K context on 12GB VRAM with Qwen3.6 35B A3B and llama.cpp MTP

Reddit user janvitos ran Qwen3.6-35B-A3B-MTP-GGUF with a llama.cpp MTP PR on an RTX 4070 Super. The posted benchmark shows 69.2-81.9 tok/s, 0.694-0.947 draft acceptance, 131072 context, and a -fitt 1536 setting that reserves 1536 MB for the draft model and KV cache.

Why it matters: HKR-H/K/R all pass with concrete single-user benchmark data and reproducible settings. Source is one Reddit post, so verification is thin; this lands above featured threshold, not in must-write range.

AI HOT (Curated Pool)

Using Codex to debug and verify fixes in parallel

The author uses Codex in temporary crabbox environments to recreate bug states, verify failures, apply fixes, and re-verify them, while running 10 sessions in parallel to avoid local state pollution and speed loss.

Why it matters: HKR-H/K/R all pass, but this is a single first-person workflow note, not a product release or benchmark. The 10-session Codex/crabbox setup earns featured-level practical signal, near the lower band.

QbitAI · WeChat

Why Perfect AI Agents Do Not Exist: Five Design Philosophies and Trade-offs Behind Claude Code

MBZUAI VILA Lab and UCL analyze Claude Code v2.1.88 source code and identify 5 design philosophies, 13 design principles, 7 permission layers, and 5 context-compaction layers behind its production-agent architecture.

Why it matters: All HKR axes pass: the contrarian Claude Code angle is clickable, the v2.1.88 permission/context mechanisms add substance, and agent tradeoffs resonate with builders. It is third-party analysis, not an Anthropic release, so it stays below must-write.

Synced · WeChat

OpenAI's Jiayi Weng: Is the Next AI Training Paradigm Beyond Gradients?

OpenAI researcher Jiayi Weng proposes Heuristic Learning: codex gpt-5.4 reached a perfect 864 score on Breakout and generated 342 search trajectories across Atari 57, with updates applied to code, tests, replays, and memory rather than neural-network weights.

Why it matters: HKR-H/K/R all pass: an OpenAI researcher proposes Heuristic Learning with concrete hooks like Breakout 864 and 342 Atari 57 trajectories. This is strong research/commentary signal, not an official model or product release, so it stays in the 78–84 band.

Latent Space

Anthropic growing 10x/year while others lay off over 10% of staff

Anthropic is described as growing 10x annually and being valued at $1T-$1.2T, while the post cites layoffs of 40% at Block, 14% at Coinbase, and 20% at Cloudflare under AI-readiness framing.

Why it matters: HKR-H/K/R all pass: the title has contrast, the post gives growth, valuation, and layoff figures, and it hits jobs plus AI-capital concentration. It is high-signal industry commentary, not an official funding or product event, so 78-84 fits.

r/LocalLLaMA

MTP + TurboQuant Running: Qwen3.6-27B Hits 80+ t/s on a Single RTX 4090

indrasmirror ran Qwen3.6-27B-Heretic-v2 on a single RTX 4090 with 262K context, TBQ4_0 KV cache, and MTP draft 3, improving throughput from about 43 t/s to 80-87 t/s with roughly 73% MTP draft acceptance.

Why it matters: HKR-H/K/R all pass, backed by a numbered first-person experiment. The Reddit-only source and niche local-inference focus keep it below the 78–84 band for broader industry releases.

AI HOT (Curated Pool)

Claude Code Practice: The Effectiveness of HTML Output

Thariq Shihipar recommends requesting HTML output from Claude, and the post cites GPT-5.5 generating an interactive Linux vulnerability page with SVG diagrams, interactive components, and in-page navigation.

Why it matters: HKR-H/K/R all pass, but this is a workflow tip rather than a Claude release. As a quality Claude Code tutorial, it sits in the 72–77 band, with Simon Willison’s source authority clearing featured.

May 8Friday

Hacker News front page

Show HN: Git for AI Agents

regent-vcs released the open-source re_gent project for AI-agent version control, currently supporting Claude Code, with workflows for tracking why an agent changed files, rewinding sessions, and bisecting agent actions; the post does not disclose the license, storage format, or installation details.

Why it matters: HKR-H/K/R all pass: the Git analogy is clicky, the mechanism is concrete, and Claude Code rollback pain is real. The post lacks license, storage format, and install details, so it stays at the featured threshold.

AI HOT (Curated Pool)

Running Codex Safely at OpenAI

OpenAI runs Codex with four safeguards: sandbox isolation, human approval, strict network policies, and native agent telemetry; the post does not disclose evaluation metrics, incident rates, or enterprise deployment requirements.

Why it matters: HKR-H/K/R all pass: the OpenAI Codex post gives concrete safety mechanisms for code agents. I keep it at 74 because it lacks eval data, incident rates, or enterprise rollout details.

Synced · WeChat

OpenAI launches official CLI for terminal-based model access

OpenAI released the open-source openai-cli, letting developers call Responses, cloud tools, image generation and editing, speech transcription, and TTS from a single terminal command.

Why it matters: HKR-H/K/R all pass: an official OpenAI CLI, open-source packaging, and terminal access to multimodal APIs. This is a useful developer workflow update, not a major model capability release, so it sits in low featured.

AI HOT (Curated Pool)

Adaptive Parallel Reasoning: A New Paradigm for Efficient Reasoning Scaling

BAIR’s post describes adaptive parallel reasoning, where ThreadWeaver and Multiverse dynamically control parallel threads for math and code reasoning; the RSS snippet does not disclose benchmark scores, latency reductions, or reproducible settings.

Why it matters: BAIR authority supports the 72+ band, and HKR-H/K/R all pass. The post names mechanisms and dynamic thread control, but lacks scores, latency gains, and reproducible conditions, so it stays below 78.

Ruan YiFeng's Weblog

Technology Enthusiast Weekly Issue 395: The Third Way of Software Development

Ruanyifeng Weekly issue 395 frames AI-assisted coding as a “mystery house” style of software development and cites HN SOTA, which ranks model popularity by scanning 200 top Hacker News topics each day and their programming or AI discussions.

Why it matters: HKR-H/K/R pass: the “third way/mystery house” framing, HN SOTA’s 200 daily HN topics, and developer workflow anxiety all land. It is commentary, not a model or product release, so it stays at 72.

AI HOT (Curated Pool)

Codex Plugin Now Supports Parallel Runs Across Chrome Tabs

OpenAI says Codex now runs in Chrome on macOS and Windows. The plugin works across tabs in the background without taking browser control; the post does not disclose version, concurrency limits, or enterprise policy.

Why it matters: HKR-H/K/R all pass, but the post gives platform and execution mechanics only; version, concurrency limits, and enterprise controls are not disclosed. Score: 76 as a practical OpenAI Codex product update.

AI HOT (Curated Pool)

Agent Pull Requests Are Everywhere: How to Review Them

GitHub published a guide for reviewing pull requests generated by AI agents. The snippet lists 3 focus areas: code changes, logic or security bugs, and pre-merge technical debt. The key issue is a review process before automated commits reach production.

Why it matters: HKR-H/K/R all pass: GitHub gives a practical checklist for agent-generated PRs with 3 review areas. It is guidance, not a product or model release, so it stays at the featured threshold.

May 7Thursday

AI HOT (Curated Pool)

Trillion-parameter instruction model Ling-2.6-1T released

inclusionAI says Ling-2.6-1T is now live on OpenRouter. The trillion-parameter instruction model uses “fast thinking” and claims top AIME26 and SWE-bench Verified results with about 75% lower cost. The post does not disclose pricing, context length, or full benchmark scores.

Why it matters: HKR-H/K/R all pass: a 1T instruction model on OpenRouter with fast thinking, AIME26/SWE-bench claims, and ~75% cost reduction. Missing price, context window, and full scores keep it in the 78–84 band.

Hacker News front page

AlphaEvolve: Gemini-powered coding agent scaling impact across fields

Google DeepMind describes AlphaEvolve as a Gemini-powered coding agent; the body is only an RSS snippet. The title discloses coding-agent scope and cross-field impact, but the post does not disclose model version, benchmarks, or deployments.

Why it matters: HKR-H and HKR-R pass on a DeepMind Gemini coding-agent announcement, but HKR-K fails: only title-level facts are disclosed. This reaches featured threshold, not 78+, because evals, model version, and deployments are absent.

OpenAI News

Scaling Trusted Access for Cyber with GPT-5.5 and GPT-5.5-Cyber

OpenAI expanded Trusted Access for Cyber to GPT-5.5 and GPT-5.5-Cyber. The RSS snippet says access is for verified defenders; the post does not disclose criteria, pricing, or benchmark data.

Why it matters: HKR-H/K/R all pass: OpenAI expands trusted cyber access to GPT-5.5 and GPT-5.5-Cyber. Kept below 85 because admission rules, pricing, evals, and reproducible tests are not disclosed.

r/LocalLLaMA

Running Qwen3.5/Qwen3.6 with NextN MTP in llama.cpp on one RTX 3090 Ti

A Reddit user posted a llama.cpp guide for Qwen3.5/3.6 with NextN MTP on one RTX 3090 Ti. It requires two unmerged PRs, #22400 and #22673; Qwen3.6-35B-A3B-MTP reaches 157 tok/s at 350W, 1700MHz, with q8 KV. The key reproducible detail is nextn=q8_0 quant override; missing it yields “////” output.

Why it matters: HKR-H/K/R all pass: single-GPU 157 tok/s is a strong hook, and the PR/power settings make it testable. Scope stays narrow because it is a Reddit guide using unmerged PRs.

Latent Space

Anthropic-SpaceXAI's 300MW/$5B/yr Deal for Colossus I, ARR Growth Is 8000% Annualized

Anthropic announced a SpaceX compute partnership, doubled Claude Code’s 5-hour limits for Pro, Max, Team, and seat-based Enterprise, raised Opus API limits, and said Claude inference would ramp on Colossus within days; the post treats the 300MW and $5B-per-year figures as widely circulated but not canonized in Anthropic’s own announcement.

Why it matters: HKR-H/K/R all pass: the compute-deal numbers and Claude Code limit changes are concrete and practitioner-relevant. The 300MW/$5B/year claim is unofficial, so it stays below P1.

AI HOT (Curated Pool)

Amp releases Neo CLI as coding agents shift toward long-horizon workflows

Amp released Neo, a CLI tool covering remote orchestration, automatic context compression, and a Plugin API. Neo lets local threads be controlled remotely, allows all operations by default, and moves safety control to plugins; the post does not disclose version, pricing, or performance gains.

Why it matters: HKR-H/K/R all pass: Neo adds remote orchestration, context compression, Plugin API, and default-allow permissions. Amp’s reach and missing price/version/perf data keep it in the 72–77 band.