Skip to content

AI coding

Everything about AI writing code: coding assistants, vibe coding, code model evals and new developer workflows.

1,196 picksRelated topicsAgentsCursorTutorials

Latest picks

761–780 of 1,196

May 28Thursday

AI HOT (Curated Pool)

NVIDIA Releases AI Framework Polar, Raising Codex Benchmark Score by 594.74%

NVIDIA’s research team open-sourced Polar, an agent reinforcement learning framework that connects GRPO training at the model API boundary without rewriting Codex CLI, Claude Code, Qwen Code, or Pi; on Qwen3.5-4B, Polar raised Codex pass@1 on SWE-Bench Verified from 3.8% to 26.4%, while prefix_merging cut training steps from 1,185 to 218.

Why it matters: HKR-H/K/R all pass: NVIDIA open-sourced Polar with a concrete GRPO mechanism and SWE-Bench Verified numbers. This is a strong research/open-source item, not a major model or product release, so it stays in the 78–84 band.

Computing Life · Share · Yage

After SWE-Bench Pro Saturation, Someone Built a New Benchmark

DeepSWE says SWE-Bench Pro lost discrimination because of data contamination and verifier flaws; the same model set showed a 62-point spread on the benchmark, while the snippet does not disclose the audited models or test protocol.

Why it matters: HKR-H/K/R all pass: the “new ruler” hook, contamination/verifier claims, and 62-point spread give this real signal for code-agent evaluation. Source reach and impact are below the 85+ same-day tier.

AI HOT (Curated Pool)

Grok Build 0.1 on API

xAI released Grok Build 0.1 in public beta through the xAI API for agentic coding tasks, with throughput above 100 tokens per second and pricing at $1 per million input tokens and $2 per million output tokens.

Why it matters: HKR-H/K/R all pass, but this is a 0.1 public-beta API and pricing launch; benchmarks, context window, and task success rates are not disclosed. It fits a solid mid-weight product update at 78, featured not p1.

AI HOT (Curated Pool)

Security Changes in the AI Agent Era

Lemonade CISO Jonathan Jaffe says a single endpoint can run 200 to 10,000 agents, so security teams need to assign identity to each agent and enforce policies at the point of action, beyond current identity and access management systems.

Why it matters: HKR-H/K/R all pass, but this is an event-recap commentary rather than a product or research release. The concrete signal is the endpoint agent count and identity-control model, placing it at the 72-77 featured threshold.

AI HOT (Curated Pool)

Cognition AI raises over $1B, targets 10x software engineering productivity

Cognition AI raised over $1 billion at a $26 billion pre-money valuation, while annualized revenue grew from $37 million to about $492 million in one year, and Devin is positioned as an autonomous junior engineer that can plan, test, and deploy through multi-step workflows.

Why it matters: HKR-H/K/R all pass: the story has hard numbers on funding, valuation, and ARR, plus a direct junior-engineer automation angle. Single-post sourcing keeps it below the 95+ industry-shaking band.

AI HOT (Curated Pool)

Using LLMs to secure source code

Anthropic describes a six-step Claude Opus workflow for source-code security: threat modeling, sandboxing, vulnerability discovery, validation, triage, and remediation; in its open-source scanning work, it disclosed 1,596 vulnerabilities by May 22, 2026, with 97 already fixed.

Why it matters: HKR-H/K/R all pass: Anthropic gives a Claude Opus security-audit workflow plus 1,596/97 outcome numbers. It stays below 85 because this is not a new model or platform-level capability release.

AI HOT (Curated Pool)

Cognition becomes the world’s largest independent agent lab

Cognition announced over $1 billion in funding at a $26 billion valuation, with enterprise usage up more than 10x this year and annualized revenue reaching $492 million.

Why it matters: HKR-H/K/R all pass: Cognition’s >$1B raise, $26B valuation, and $492M ARR are hard numbers tied to coding-agent competition and developer workflows. This fits the 85–94 same-day band, with no cross-source bump or deal details.

AI HOT (Curated Pool)

I Think Anthropic and OpenAI Found Product-Market Fit

Anthropic and OpenAI changed enterprise pricing around April 2026, moving coding agents from heavily discounted seat plans to API-usage billing, with Anthropic Enterprise at $20 per seat per month plus API fees and OpenAI Codex billed by API token usage.

Why it matters: HKR-H/K/R all pass: the piece ties OpenAI and Anthropic PMF to a concrete billing shift for coding agents. It is influential commentary, not an official launch, so it fits the 78–84 band.

TechCrunch · AI

AI coding startup Cognition raises $1B at $25B pre-money valuation

Cognition raised $1 billion at a $25 billion pre-money valuation, with annualized revenue run rate reaching $492 million, and the company says its valuation more than doubled in eight months.

Why it matters: HKR-H/K/R all pass: the story has a sharp valuation hook, concrete revenue and round data, and strong resonance around AI coding economics and developer displacement.

May 27Wednesday

AI HOT (Curated Pool)

AI Builds AI: ModelBest Open-Sources ForgeTrain, a Training Framework Written by AI

ModelBest, Tsinghua University, and OpenBMB open-sourced ForgeTrain, described as the first production-grade LLM training framework written entirely by AI with zero human code, and ModelBest used it to pretrain MiniCPM5-1B on Huawei Ascend chips.

Why it matters: HKR-H/K/R all pass: an open-source training framework, AI-written code, and MiniCPM5-1B pretraining on Ascend give concrete hooks. This is a strong tooling story, not a top-model launch, so 80 fits featured rather than P1.

Latent Space

[AINews] New AI Infra Decacorns: Fireworks, Baseten, with OpenRouter on the Way

Latent Space says Fireworks is in talks for a $15 billion valuation round, Baseten is raising at an $11 billion valuation, and OpenRouter closed a $113 million Series C after volume grew 5x in six months.

Why it matters: HKR-H/K/R all pass: the decacorn hook is clickable, the post gives valuation, round, and usage figures, and the topic speaks to inference economics. Fireworks and Baseten are still reported as in talks or raising, so this stays in the 78–84 band.

AI HOT (Curated Pool)

Code w/ Claude London event: Rethinking the developer experience

Anthropic announced two Claude Managed Agents capabilities at Code w/ Claude London: self-hosted sandboxes in public beta and MCP tunnels in research preview, with Spotify, Base44, and Legora already using them.

Why it matters: Official Anthropic product update with two concrete Claude Managed Agents capabilities. HKR-H/K/R pass, but this is a developer-tooling update rather than a major model release, so it lands at 78.

May 26Tuesday

The Verge · AI

Uber president says AI spending is getting harder to justify

Uber president Andrew Macdonald said the company exhausted its 2026 AI budget in four months, while rising Claude Code token consumption has not been tied to a measurable increase in useful consumer features delivered.

Why it matters: HKR-H/K/R all pass: a senior Uber exec gives a contrarian AI-spend quote, the story has a 4-month budget-burn number, and it hits Claude Code ROI anxiety. Strong industry signal, not a model or major product launch, so it sits in 78–84.

r/LocalLLaMA

SkillOpt treats markdown skill files as trainable parameters with proper optimization machinery

SkillOpt uses a frontier model to propose add, delete, and replace edits to markdown skill files, then accepts only strict gains on a held-out validation set; the best skills usually converge after 1 to 4 accepted edits.

Why it matters: HKR-H/K/R all pass: the hook is trainable markdown skills, with held-out validation and 1-4 accepted edits. Single Reddit/project source and no broad adoption data keep it at 78, featured not p1.

AI HOT (Curated Pool)

Qwen3.7-Max Becomes the World’s No. 2 AI Coding Model

Qwen3.7-Max scored 1541 on Code Arena and ranked behind Claude; the post says it can run 35-hour tasks and perform more than 1,000 tool calls.

Why it matters: HKR-H/K/R all pass, but the source is a single Alibaba Cloud post and the evidence is benchmark plus vendor claims. This fits a strong product/benchmark update, not P1 without independent validation.

QbitAI · WeChat

Chinese AI-Written Pretraining Framework ForgeTrain Trains MiniCPM5-1B

ModelBest released ForgeTrain and MiniCPM5-1B, saying ForgeTrain was written by AI and trains 10% faster than NVIDIA Megatron under the same hardware conditions. MiniCPM5-1B is a 1B-parameter edge model with about 2GB FP16 weights and about 0.5GB INT4/Q4 weights.

Why it matters: HKR-H/K/R all pass: an AI-written trainer, a 10% same-hardware Megatron speed claim, and a 0.5GB 1B edge model are concrete hooks. Score stays at 80 because the first-ever claim and benchmark lack third-party reproduction.

Synced · WeChat

Grok keeps updating after xAI disbandment as Musk announces a new model

Elon Musk said the 1.5T-parameter Grok V9-Medium has finished training, will enter reinforcement learning in a few days, and is planned for release in two to three weeks. Grok Build supports up to 8 parallel sub-agents, a 256K-token context window, Plan Mode, Arena Mode, MCP, and ACP.

Why it matters: HKR-H/K/R all pass, but this is a Grok V9-Medium preview before RL and release, with no benchmarked capability yet. That fits a strong model-race/product update at 82, featured but not p1.

Synced · WeChat

AI-written training framework trains 1B edge model MiniCPM5-1B

ModelBest open-sourced MiniCPM5-1B and ForgeTrain; the 1B edge model scores 17.9 on AA-Index, while the AI-written ForgeTrain framework matches Megatron’s training results and runs 10% faster on Nvidia H100 under the article’s reported setup.

Why it matters: HKR-H/K/R all pass: the AI-written training framework hook is strong, with concrete AA-Index and H100 speed claims. It is not a flagship model release, so it stays in the 78–84 band.

Xinzhiyuan · WeChat

Chinese agent SkyClaw targets Opus 4.6-level performance with free trial

Kunlun Tech released SkyClaw-v1.0 and SkyClaw-v1.0-lite with a 2-4 week free trial, claiming SkyClaw-v1.0 input costs are 1/24 of DeepSeek V4 Pro and about 1/43 of Sonnet 4.6.

Why it matters: HKR-H/K/R all pass: SkyClaw-v1.0 has a sharp cost hook, concrete trial and pricing ratios, and budget resonance. Source facts remain vendor claims, so it stays at the low featured band.

AI HOT (Curated Pool)

OpenAI GPT-5.6 Reportedly Set for Next Month With 1.5M-Token Context

Developers found an unannounced OpenAI GPT-5.6 entry in Codex backend logs under the codename iris-alpha, with a 1.5 million-token context window, about 43% higher than GPT-5.5’s 1.05 million-token limit.

Why it matters: HKR-H/K/R all pass: the Codex-log leak, 1.5M-token window, and 43% increase are concrete and practitioner-relevant. It stays below 85 because this is not an official GPT-5.6 launch.