Skip to content

AI coding

Everything about AI writing code: coding assistants, vibe coding, code model evals and new developer workflows.

1,196 picksRelated topicsAgentsCursorTutorials

Latest picks

1101–1120 of 1,196

Apr 16Thursday

OpenAI News

Introducing GPT-Rosalind for life sciences research

OpenAI released GPT-Rosalind on April 16, 2026, and made it available as a research preview in ChatGPT, Codex, and the API for qualified customers. The post says it targets biology, drug discovery, and translational medicine, and adds a free Codex life sciences plugin connecting to 50+ scientific tools and data sources. The real signal is deployment breadth: Amgen, Moderna, and Thermo Fisher Scientific are involved, but the post does not disclose model size, pricing, or benchmark scores.

Why it matters: HKR-H lands because OpenAI is shipping a vertical life-sciences model; HKR-K lands on access paths and the 50+ tool/data plugin. HKR-R also lands on the domain-model debate, but missing params, pricing, and benchmark scores keep it at featured, not p1.

Apr 15Wednesday

X · @dotey

Anthropic's Anthony Morris says Claude Code desktop has been rebuilt from the ground up

Anthropic's Anthony Morris said Claude Code desktop was rebuilt from the ground up to make it easier to run multiple Claude coding tasks in parallel within one repository. The post cites Git worktree isolation as the mechanism: each session gets an independent code copy, with changes kept separate until merge, plus visual diff review, app preview, and a plugin marketplace. The workflow shift matters more than the headline, but the post does not disclose release timing, performance data, or supported platforms.

Why it matters: This is a substantive Claude Code product update aimed at a real workflow pain point: parallel coding sessions in the same repo. HKR-H/K/R all pass through the strong hook, concrete worktree-based mechanism, and developer resonance, but missing launch date, performance data, and

X · @claudeai

Claude Code on desktop redesigned with side-by-side sessions in one window

Anthropic redesigned Claude Code on desktop and now lets users run multiple Claude sessions side by side in one window. The RSS snippet confirms a new sidebar for session management; the post does not disclose rollout timing, platforms, or more interaction details. For coding workflows, the key question is whether multi-session control cuts context-switch overhead.

Why it matters: An authoritative Anthropic post plus a concrete workflow change gives it HKR-H/K/R. It stays near the featured floor because rollout date, supported desktop platforms, and deeper interaction details are not disclosed, and the scope is still a mid-weight product update.

X · @op7418

Claude Code's newly released routines feature looks strong

Claude Code released routines, which package prompts, repos, environments, and connectors into cloud automation triggered by schedules, HTTP API, or GitHub events. Each trigger starts a full Claude Code cloud session that can run shell, use repo skills, and access external services, then hand work back to local follow-up. The post does not disclose pricing, quotas, or supported platforms.

Why it matters: This is a substantive Claude Code workflow update: routines package prompts, repo, environment, and connectors into cloud jobs triggered by schedule, HTTP API, or GitHub events. HKR-H/K/R all pass, but price, quota, and supported platforms are not disclosed, so it stays featured,

X · @dotey

Claude Code adds Routines for trigger-based automated tasks

Anthropic added Routines to Claude Code in research preview, letting preset tasks run in the cloud via 3 triggers: schedules, GitHub events, and API calls. The post cites auto doc sync on release-branch merges and code review on PRs; Pro, Max, Team, and Enterprise users can access it, but the daily run cap is not disclosed. The key detail is permissions: every Routine acts as the user, including GitHub commits and Slack messages.

Why it matters: This is more than a minor feature tweak. Routines moves Claude Code toward an event-driven cloud agent, with concrete details on 3 triggers, plan availability, and a user-identity permission model, so HKR-H/K/R all pass. It stays below p1 because this is still a research preview,

X · @claudeai

Now in research preview: routines in Claude Code

Anthropic launched routines in research preview for Claude Code: configure a prompt, repo, and connectors once, then run it on a schedule, via API, or from an event. Routines run on Anthropic web infrastructure, so a laptop does not need to stay open; the post does not disclose pricing, quotas, or rollout scope. The key point is hosted execution, not one-off code completion.

Why it matters: This is a substantive Claude Code expansion from local interactive coding to hosted, scheduled, and event-driven execution. HKR-H/K/R all pass, and the Anthropic update gets a policy bump, but price, quotas, and rollout scope are not disclosed, so it stays featured rather than P1

Apr 14Tuesday

X · @dotey

Rather than AI First, this is really Software Engineering First

The post argues “AI First” is an engineering problem: if AI writes code in 2 hours, review, testing, deploy, monitoring, and rollback must also run automatically, with humans kept at key decision points. Its concrete prerequisites are automated tests, CI/CD, A/B testing, production monitoring, task management, and a clear architecture; without them, a 25-person team just shifts bottlenecks from coding to QA and ops. The real boundary is use case fit: API services, data platforms, and internal tools fit better than complex UI, core products, or high-security systems.

Why it matters: This is a strong practitioner commentary rather than a news event. HKR-H lands on the contrarian framing, HKR-K on concrete prerequisites and scope limits, and HKR-R on the bottleneck-shift argument; it stays in the mid-70s because there are no named cases, first-person tests, or

X · @dotey

Vercel open-sources Open Agents, a reference implementation for enterprise coding agent platforms

Vercel open-sourced Open Agents as a forkable reference for enterprise coding-agent platforms, with a three-layer architecture and features like voice input and PR creation. Its key design keeps the agent outside the sandbox and uses tools such as file I/O, shell, and search to control execution; the post also cites Anthropic Managed Agents pricing at $0.08 runtime per hour and $10 per 1,000 web searches. The part to watch is the agent-sandbox split, not the packaging choice.

Why it matters: This fits the 78–84 band: a notable open-source coding-agent framework with concrete architecture, remote sandbox operation, and Anthropic pricing, so HKR-H/K/R all land. It stops short of must-write status because this is strong infra reference material, not a model or industry-

最佳拍档 (BestPartners)

Meta-Harness: Can harness engineering code self-iterate? A Stanford paper analysis

Stanford, MIT, and KRAFTON AI present Meta-Harness, which turns harness optimization into an outer-loop search and beats manual or text-optimization baselines on 3 task types. The system uses a coding agent to inspect filesystem history; after 10 search iterations, the data exceeds 10 million tokens, and on online text classification it matched OPRO’s 60-iteration result in 4 iterations while reaching 75.9% average accuracy on 5 OOD datasets. The key point is full-feedback retention rather than compression; the paper also reports about 20 TerminalBench-2 iterations at a total cost of a few hundred dollars.

Why it matters: This is a good research-release explainer for agent builders: the mechanism is clear and the post includes concrete numbers, so HKR-H/K/R all pass. It stays at 80 because the source is a secondary YouTube summary, not the primary paper or official release, and the impact is still

Apr 12Sunday

X · @Yuchenj_UW

MiniMax M2.7 is open-source!

MiniMax open-sourced M2.7 and said its research agent now handles 30%–50% of the R&D workflow. The post says the agent covers literature review, experiment orchestration, log debugging, code fixes, and merge requests; M2.7 also rewrote its own harness for 100+ automated rounds, with a 30% gain on internal coding evals.

Why it matters: HKR-H/K/R all pass: open-sourcing plus a research agent doing 30%-50% of R&D is a strong hook, and the post includes 100+ self-rewrite loops with +30% internal coding eval. It stays at 78 because license, repo, benchmark context, and external reproduction are not disclosed.

Apr 11Saturday

X · @dotey

OpenAI Codex team's Nick Baumann: build dedicated CLI tools for AI instead of feeding messy data repeatedly

OpenAI Codex engineer Nick Baumann says teams should wrap repeated data access into parameterized CLI tools with JSON output instead of repeatedly dumping logs, docs, and API responses into Codex. The post lists 3 examples in daily use: codex-threads for past sessions, slack-cli for threaded Slack search, and typefully-cli for posting workflows; access still goes through the existing auth gateway. The point for practitioners is narrower interfaces: models handle focused commands more reliably than raw, noisy source data.

Why it matters: This is a practical workflow note from an OpenAI Codex team member, not a formal launch, but it offers a reusable mechanism: wrap noisy context behind parameterized JSON-returning CLIs and shows 3 live examples. HKR-H/K/R all land; no benchmark, scale, or major product release,so

X · @dotey

Anthropic launches Claude for Word beta add-in

Anthropic released a beta Claude for Word add-in for paid Claude Team and Enterprise users, with direct sidebar editing for .docx and .docm files. Edits appear in Word’s native track changes flow, the add-in can reuse conversation context from Excel and PowerPoint, and it supports reference uploads plus reusable team Skills. The key point is shared context across Office apps; the post does not disclose pricing, regions, or a wider rollout timeline.

Why it matters: This is a substantive Anthropic product update for Team and Enterprise, not a generic integration post. HKR-H/K/R all pass on novelty, concrete mechanics, and workflow resonance, but the beta scope is limited and price, regions, and GA timing are undisclosed, so it lands in mid-"

X · @dotey

Claude Code adds ultraplan: start planning in terminal, review in browser, then run in cloud or locally

Claude Code opened a preview of ultraplan to users with the web app enabled, requiring v2.1.91+, and planning starts from /ultraplan in the terminal. Claude drafts a plan in the cloud after reading the repo, users review and annotate it in the browser, then choose cloud execution with a PR or local terminal execution. The key change is splitting planning from execution: planning moves to the cloud without blocking the terminal, and the post says token use is close to local plan mode.

Why it matters: This is more than a routine feature add: Claude Code splits planning from execution, with /ultraplan in terminal, cloud-side repo reading, browser review, and cloud PR or local execution. HKR-H/K/R all pass, with a Claude-specific bump, but it is still a preview and sourced froma

X · @claudeai

Claude for Word is now in beta

Anthropic launched Claude for Word in beta, letting users draft, edit, and revise documents from the Word sidebar on Team and Enterprise plans. The post says Claude preserves formatting and shows edits as tracked changes; it does not disclose pricing, regions, or rollout timing.

Why it matters: This is a useful but mid-weight Anthropic product update. The official post confirms Word sidebar access, Team/Enterprise availability, format retention, and tracked changes; HKR-K and HKR-R pass, but missing price, region, and rollout details keep it at the low end of featured.

Apr 10Friday

最佳拍档 (BestPartners)

LLM self-evolution: Shinka Evolve, AlphaEvolve, and sample efficiency

Sakana AI open-sourced Shinka Evolve and uses a UCB bandit to switch among GPT-5, Claude Sonnet 4.5, Gemini, and others, aiming to cut the thousands of program evaluations common in AlphaEvolve-style search. The post says it beat AlphaEvolve’s classic circle-packing result with fewer evaluations and adds full-file rewrites, crossover, editable-region guards, and a meta-notebook; the post does not disclose exact metrics, cost, or the repo link. The part to watch is surrogate-task design and hard verification: the system still needs humans to define problems.

Why it matters: Featured, not P1: HKR-H/K/R all pass. The piece has a strong hook, concrete mechanisms like UCB model routing and program crossover, and a real nerve around eval cost and hard verification. It stays at 80 because key metrics, cost, and the primary release link are not disclosed.

QbitAI · WeChat

Tencent open-sources 3B SVG model HiVG to make tokens geometry-aware

Tencent Hunyuan open-sourced the 3B-parameter HiVG, claiming 62.7%-63.8% shorter SVG sequences via hierarchical tokenization and better SVG generation metrics than GPT-5.2, Claude-4.5-Sonnet, and some 8B open models. The post reports 0.896 SSIM, 0.114 LPIPS, and 0.957 CLIP-S on Image-to-SVG; the core method packs drawing commands plus coordinates into segment tokens and uses HMN to initialize coordinate embeddings. The part to watch is token design, not parameter count; paper, code, and project page are public.

Why it matters: Tencent's HiVG earns HKR-H and HKR-K: a 3B open model claims GPT/Claude-level SVG results, and the article includes 62.7%-63.8% token compression plus SSIM 0.896, LPIPS 0.114, and CLIP-S 0.957. HKR-R is weaker because SVG generation remains niche, so it lands at the low end of `f

X · @OpenAI

OpenAI updates ChatGPT Pro and Plus subscriptions to support growing Codex usage

OpenAI set a new ChatGPT Pro tier at $100/month and raised Codex usage to 5x ChatGPT Plus. The tier keeps all Pro features, including the exclusive Pro model and unlimited Instant and Thinking access. Through May 31, $100 Pro subscribers get up to 10x Plus usage on Codex; the real signal is separate pricing for heavy code-agent demand.

Why it matters: This is an OpenAI product-pricing update centered on Codex usage, with HKR-K from concrete pricing/quota facts and HKR-R from a clear signal on code-agent monetization. No new model or capability is disclosed, and HKR-H is weaker, so it lands as solid featured rather than must-wr

Apr 9Thursday

X · @dotey

Anthropic launches Claude Managed Agents, a managed API for building and deploying agents, now in public beta

Anthropic launched Claude Managed Agents, a managed API for building and deploying agents, in public beta. It offers a production sandbox, long-running sessions, and multi-agent coordination; Anthropic says internal tests showed up to a 10-point success-rate gain on structured file-generation tasks versus standard prompt loops. Pricing uses standard Claude token fees plus $0.08 per active session-hour; the real signal is Anthropic moving agent infrastructure into its platform layer.

Why it matters: Anthropic packaged managed agents, sandboxing, and long-running sessions into a public-beta API, which is a real workflow update for developers. HKR-H/K/R all pass: strong platform hook, concrete facts like a 10-point gain and $0.08 per hour, and clear resonance around developer-

Apr 8Wednesday

QbitAI · WeChat

After a late-night update, DeepSeek reportedly said: I am V4?

DeepSeek added Fast and Expert modes on its web app and started gray-testing a Vision model; the claim that Expert mode is V4 comes only from user probes and the model’s own replies. The post gives one concrete detail: Expert mode focuses on code, web, and harder generation tasks, is supply-limited, does not support multimodal or file upload, and one user reported a length cap at about 133K tokens. What matters is the official model ID and context spec; the post does not disclose them, pricing, or a release timeline.

Why it matters: HKR-H is strong on the 'I am V4' hook. HKR-K and HKR-R pass because the post gives testable mode behavior and a ~133K token limit, and DeepSeek silent swaps are highly discussable. The score stays in the mid-70s because the model name, price, and context window remain unconfirmed

X · @dotey

Anthropic launches Claude Mythos Preview and Project Glasswing for vulnerability hunting

The post says Anthropic released Claude Mythos Preview and restricted it to 12 partners for vulnerability research, with no public app, API, or enterprise access. It cites 93.9% on SWE-bench Verified, 97.6% on USAMO, and a 244-page system card, plus $100M in credits and $4M in grants; the key point is closed distribution of high-risk capability, not just benchmark wins.