Skip to content

AI coding

Everything about AI writing code: coding assistants, vibe coding, code model evals and new developer workflows.

1,196 picksRelated topicsAgentsCursorTutorials

Latest picks

1141–1160 of 1,196

Mar 5Thursday

OpenAI News

GPT-5.4 Thinking System Card

OpenAI published the GPT-5.4 Thinking System Card on March 5, 2026 and says it is the latest GPT-5 reasoning model and the first general-purpose model with mitigations for high-capability cybersecurity. The post confirms the safety approach follows prior GPT-5 models and builds on measures used for GPT-5.3 Codex, but it does not disclose benchmark scores, mitigation details, or deployment conditions. The key signal is the risk threshold change: OpenAI has extended high-cyber mitigations to a general reasoning model.

Why it matters: This clears HKR-H/K/R: a new GPT-5 reasoning model and the first general-purpose model with high-capability cyber mitigations. It stays below p1 because the disclosed text does not provide eval scores, mitigation details, or deployment conditions.

Feb 26Thursday

OpenAI News

OpenAI Codex and Figma launch code-to-design roundtrip workflow

OpenAI and Figma launched a Codex integration on Feb. 26, 2026 that turns code into editable Figma designs and brings Figma Design, Figma Make, and FigJam content back into code. The workflow uses MCP via the Figma MCP Server in the Codex desktop app; OpenAI says Codex has 1M+ weekly users and usage is up 400%+ since the start of the year. The key issue is whether roundtrip context stays intact; the post does not disclose supported models, permission boundaries, or pricing.

Why it matters: This is a solid OpenAI/Figma workflow update with clear HKR-H/K/R: a bidirectional code↔design loop via MCP and Figma MCP Server. It stays below 85 because the post does not disclose model support, permission boundaries, pricing, or roundtrip reliability.

Feb 20Friday

Hugging Face Blog

GGML and llama.cpp join Hugging Face to support the long-term progress of Local AI

Hugging Face said the GGML and llama.cpp team is joining the company, while Georgi Gerganov’s team will still spend 100% of its time maintaining llama.cpp. The post says the project remains 100% open source and community driven, with full technical and community autonomy. The key angle is tighter delivery from transformers model definitions into llama.cpp, aiming for near “single-click” shipping; the post does not disclose timeline, team size, or deal terms.

Why it matters: This is a meaningful local-AI infrastructure move: HF brings in the GGML/llama.cpp team, so HKR-H/K/R all pass. I kept it at 78 because the post confirms staffing and integration direction, but not a ship date, team size, or deal terms.

Feb 14Saturday

Ruan YiFeng's Weblog

Using ByteDance's Seed 2.0 and TRAE with Skills for app building and deployment

Ruanyifeng used ByteDance's Seed 2.0 Code and TRAE to generate one ASCII-to-Excalidraw web app and preview it at localhost:8080. The post says Seed 2.0 includes Pro, Lite, Mini, and Code models, and shows Skills as YAML-headed Markdown files, including Anthropic's frontend-design and Vercel deploy examples.

Why it matters: HKR-H and HKR-K land because the post turns Seed 2.0 Code + TRAE into a runnable mini app and explains the Skill mechanism with concrete setup details. HKR-R also lands for coding-agent workflow reuse, but this is a strong tutorial, not a major ByteDance launch, so it sits at the

Dwarkesh Patel

Dario Amodei: “We are near the end of the exponential”

Anthropic CEO Dario Amodei said in a long interview that model capability gains are still tracking an exponential, but are near its end, with the timeline off by only 1-2 years. He attributes progress to compute, data, training duration, and scalable objectives, and says RL shows log-linear gains on math and coding tasks; the post does not disclose exact curves, model versions, or reproducible parameters. The key claim is that pretraining and RL follow one scaling story, not two separate ones.

Why it matters: A top-lab CEO is making a direct claim on scaling, RL returns, and a 1-2 year timeline, so HKR-H/K/R all pass. I stop at 85 because this is thesis-level signal, not a product or research artifact: no curves, model IDs, or reproducible conditions are disclosed.

Feb 13Friday

OpenAI News

Beyond rate limits: scaling access to Codex and Sora

OpenAI says in the headline it will scale access to Codex and Sora beyond current rate limits. The body is empty and does not disclose quota changes, eligible users, pricing, or rollout timing. The key missing fact is the access mechanism, not the headline claim.

Why it matters: This is an official OpenAI product update, so HKR-H and HKR-R pass: the rate-limit angle is clickable and quota pain resonates with users. HKR-K fails because the body discloses no quota delta, eligible tiers, pricing, or rollout date, so it stays at the featured floor.

Feb 12Thursday

MIT Technology Review · AI

AI is already making online crimes easier. It could get much worse.

Microsoft said it blocked $4 billion in scams and fraudulent transactions in the year to April 2025, with many likely aided by AI-generated content. The article cites research estimating at least half of spam email is now LLM-generated, and LLM use in targeted email attacks rose from 7.6% in April 2024 to 14% in April 2025. Don’t overread “fully automated AI hackers”: the immediate issue is AI scaling phishing, deepfakes, and malware support, while the post does not disclose total attack growth.

Why it matters: HKR-H/K/R all pass: the swindle angle is strong, and the article adds concrete abuse metrics ($4B blocked, half of spam, 7.6%→14%). Featured, not p1, because this is a solid trend report on AI-enabled fraud, not a same-day industry-moving release or incident.

MIT Technology Review · AI

What’s next for Chinese open-source AI

MIT Technology Review says that after DeepSeek released R1 in January 2025, Chinese firms kept shipping open-weight models near top Western systems; Moonshot AI’s Kimi K2.5 was close to Anthropic Claude Opus on early benchmarks at about one-seventh the price. The post also says Qwen took over 30% of Hugging Face downloads in 2024 and surpassed Meta Llama in cumulative downloads by 2025–2026; the key shift is from a few general models to many fine-tunable, distillable variants.

Why it matters: All three HKR axes pass. This is not a launch, but it offers concrete market signals—~1/7 pricing, Hugging Face download share, and a clear thesis that Chinese open source is moving toward specialized, distillable variants—so it merits featured, not p1.

OpenAI News

Introducing GPT-5.3-Codex-Spark

OpenAI posted an item titled “Introducing GPT-5.3-Codex-Spark,” confirming the model name GPT-5.3-Codex-Spark. The body is empty in the RSS snippet, so pricing, context window, launch scope, and code-specific details are not disclosed.

Why it matters: An official OpenAI post confirms a new model name, so HKR-H and HKR-R pass on novelty and developer attention. HKR-K fails because the body discloses no specs, pricing, context window, benchmarks, or product scope, keeping this at the featured floor.

Ruan YiFeng's Weblog

Hands-on with Zhipu's flagship GLM-5: compared with Claude Opus 4.6 and GPT-5.3-Codex

Ruan Yifeng compared GLM-5, Claude Opus 4.6, and GPT-5.3-Codex on 4 coding tasks, and judged GLM-5 competitive with the two closed models overall. The post covers web redesign, a 3D sandbox, an Angry Birds clone, and Laravel-to-Next.js migration; in the migration task, GLM-5 and GPT-5.3 took about 5 minutes, while Opus 4.6 took about 20. The key point: this is a single-author hands-on comparison, not a standardized benchmark.

Why it matters: This clears HKR-H/K/R because it is a named first-person test with 4 tasks, video evidence, and a 5-minute versus ~20-minute gap. I did not score it higher because it is one author's evaluation, not a standardized benchmark or a broad multi-source release event.

Feb 6Friday

TechCrunch · AI

OpenAI launches new agentic coding model minutes after Anthropic releases its own

OpenAI launched an agentic coding model minutes after Anthropic released a similar one, and the model is meant to accelerate Codex, which OpenAI launched earlier this week. The RSS snippet gives only the timing and purpose; the post does not disclose the model name, benchmarks, pricing, context length, or availability. The signal is direct competition in agentic coding, not a substantiated performance claim.

Why it matters: Major-lab product news plus a minutes-apart Anthropic clash gives this HKR-H and HKR-R. The score stays in the low featured band because HKR-K is weak: the post lacks the model name, benchmarks, price, context window, and availability.

Feb 5Thursday

MIT Technology Review · AI

This is the most misunderstood graph in AI

MIT Technology Review says METR’s plot shows frontier models’ software-task time horizon doubling about every seven months; Claude Opus 4.5 was estimated at about five hours in December 2025. The post stresses that five hours means human time for comparable tasks, not five autonomous model hours; METR gave Opus 4.5 a roughly 2-to-20-hour range. The key caveat: the plot mainly measures coding tasks and defines time horizon at 50% task success, not general AI ability.

Why it matters: HKR-H/K/R all land: the piece has a strong hook and clarifies the METR chart with concrete, testable details. It stays in the low featured band because this is authoritative explanatory commentary, not a new model, product, or research release.

Feb 3Tuesday

Computing Life · Yage

Beyond Tutorial Thinking: Why AI Education Should Add Engineering Infrastructure, Not Just Content

The team says it ran 4 courses over 2 years for 2,500+ learners, yet only a minority shipped usable products; drop-off centered on setup, experimentation, deployment, and context handling friction. The post says AI Builder Space gives students a no-card unified API, one-click deployment to <name>.ai-builders.space free for 1 year, and MCP access for Cursor and Claude Code via one command. The point is productized teaching infra, not more tutorials; retention, conversion, and cost are not disclosed.

Why it matters: The piece turns a familiar complaint into operational detail: 2500+ learners, 4 failure points, and a concrete platform response with API, deployment, and MCP access. HKR-H/K/R all pass, but missing conversion, retention, and cost data keeps it at the low end of featured.

Feb 1Sunday

Lex Fridman (YouTube RSS)

State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI | Lex Fridman Podcast #490

Lex Fridman, Sebastian Raschka, and Nathan Lambert discuss the 2026 AI race in podcast #490 and frame DeepSeek R1’s January 2025 release as a key inflection point. The episode names Claude Opus 4.5, Gemini 3, Z.ai GLM, Minimax, and Kimi Moonshot, but the post does not disclose a shared benchmark, cost table, or reproducible eval. The useful takeaway is the lens: gaps look more like compute, budget, and org culture than secret ideas.

Why it matters: High-quality commentary, not a news break. HKR-H and HKR-R pass because Lex Fridman, Sebastian Raschka, and Nathan Lambert frame China, agents, GPUs, and AGI for practitioners. HKR-K misses: the post names models and DeepSeek R1 but provides no shared benchmarks, cost table, or a

Jan 30Friday

Ruan YiFeng's Weblog

Technology Enthusiast Weekly #383: What Level of AI Programming Are You?

Steve Yegge frames AI coding into 8 levels and says he is at level 8, where an orchestrator manages parallel AI coding sessions. The post lays out a path from IDE copilots to YOLO acceptance, 3-5 windows, 10+ windows, then orchestration; it also says his AI-built tool Gas Town has 225,000 lines of Go code, which he has never read, and had 6,000 stars as of last week. The real signal is black-box programming as a workflow choice, with cost and failure risk stated plainly.

Why it matters: Strong HKR-H/K/R: the 8-level framing is sticky, and the post carries concrete workflow and project numbers. The score stays below 78 because this is secondary commentary, not a primary model, product, or research release.

MIT Technology Review · AI

DHS is using Google and Adobe AI to make videos

A DHS document says the agency uses Google Veo 3, Google Flow, and Adobe Firefly for public-facing content, with an estimated 100 to 1,000 licenses. It also says DHS uses Microsoft Copilot Chat for drafting and summarization and Poolside for coding; the post does not disclose which specific videos used which tool. The key point for practitioners is that commercial video generators are now inside a federal public-communications workflow, while watermark retention and attribution remain unverifiable across platforms.

Jan 29Thursday

Ruan YiFeng's Weblog

Kimi’s integrated stack vs. Manus’s layered approach

Kimi released the K2.5 model and K2.5 Agent together, with an agent mode already available on its website. The post cites 1,500-step long-horizon actions, up to 100 agents in parallel, and visual coding from design files or web videos; pricing, context window, and API terms are not disclosed. The key point is product shape: not just a model launch, but a bundled model-plus-agent release.

Why it matters: HKR-H lands on the integrated release angle; HKR-K lands on the 1,500-step, 100-agent, visual-programming details; HKR-R lands on the stack-design debate. Missing price, context window, and API terms, plus a commentary source, keep it below p1.

Jan 28Wednesday

Mistral AI

Mistral releases terminal coding agent Mistral Vibe 2.0

Mistral released Mistral Vibe 2.0, a terminal coding agent powered by the Devstral 2 model family. It adds custom subagents, multi-option clarification, slash-command skills, a unified agent mode and automatic updates.

Why it matters: The post lists Vibe 2.0's custom subagents, slash-command skills and subscription entry point, enough to judge how terminal coding agent workflows change.

Jan 5Monday

Import AI (Jack Clark)

Import AI 439: AI kernels; decentralized training; and universal representations

Meta says KernelEvolve cut kernel development from weeks to hours and delivered up to 17x over PyTorch baselines in production tests. The system uses Llama, GPT, and Claude to generate kernels, validates them, and feeds results into a knowledge base across NVIDIA, AMD, and MTIA; the post also says decentralized training is growing 20x per year but still uses about 1000x less compute than frontier runs. The real signal is continuous self-optimizing infra in production, while decentralized training matters if that 1000x gap keeps shrinking.

Why it matters: HKR-H/K/R all pass: the kernel-writing angle is novel, the post includes concrete numbers and mechanism, and the decentralization thread hits cost and power-concentration nerves. I stop at 80 because this is a newsletter synthesis of technical work, not a single industry-defining

Jan 1Thursday

36Kr (direct RSS)

Escaping the user-acquisition nightmare: Moonshot AI's 10 billion yuan cash reserve and Yang Zhilin's confidence

Moonshot AI raised $500 million at a $4.3 billion post-money valuation; Yang Zhilin said the company holds over 10 billion yuan in cash and is not rushing to IPO. Named backers include IDG, with Alibaba, Tencent, Gaorong Ventures, and Capital Today reportedly taking super pro rata; the memo also says paid users grew over 170% MoM on average and overseas API revenue rose 4x from September to November. The signal that matters is the shift from paid traffic to open source, model capability, and agents: the post says K2 reached No. 2 on OpenRouter's global trending list within a week of open-sourcing.

Why it matters: Moonshot is a Chinese frontier-model company, so a fresh $500M round plus operating metrics matters. HKR-H/K/R all pass on the strategic pivot and hard numbers, but this is still funding and business reporting, not a major model or product launch, so it stays featured rather than