Skip to content

#编码

0 today

Yesterday · Sep 29Tuesday

OpenAI News

OpenAI releases GPT-6.1 Sol model

OpenAI released GPT-6.1 Sol, positioned as near-Astra-level intelligence for coding, computer use and professional work. Standard API input and output tokens cost one-fifth of Astra's price.

Why it matters: OpenAI's GPT-6.1 Sol launch shows the capability target for coding and computer use, plus the pricing shift.

Sep 22Tuesday

OpenAI News

Parallel cuts research time and cost in half with GPT‑6 Astra

Parallel, an AI agent infrastructure startup, used GPT‑6 Astra to research labor-market data across six states over six months. The model cut both time and code cost by 50% by issuing more targeted searches and delegating sub-tasks to parallel agents. The post doesn't specify which prior models were used for comparison.

Sep 18Friday

GitHub Blog · AI & ML

Should you read the code, is RAG dead, and did Skills kill MCP?

GitHub Podcast 最新一期拆解了五个 AI 热门观点:AI 生成的代码仍需阅读和负责,但审查力度应按风险分级;Skills 与 MCP 解决不同问题,前者是打包的团队经验,后者是连接工具与数据的标准,可组合使用;RAG 并未死亡,它为模型提供训练数据之外的相关信息,减少 token 浪费并让回答更有依据。

Sep 16Wednesday

OpenAI News

Hex turns complex analysis into visual reports with GPT‑6 Astra

Data platform Hex uses GPT‑6 Astra to turn complex analysis into interactive visual reports. Co-founder Caitlin Colgrove says models have long struggled with visualization, but Astra handles underlying libraries and geospatial transformations to produce functional and beautiful outputs. It also applies “analytical judgment”—checking whether answers make sense, match the user’s question, and serve the business goal. The post doesn’t disclose Astra’s pricing or latency.

Sep 14Monday

OpenAI News

Perplexity trusts GPT-6 Astra with end-to-end systems, checking in far less often

Perplexity co-founder Johnny Ho says GPT-6 Astra can now draft communications, edit live systems, and monitor production—tasks earlier models couldn't handle. They let Astra write its own test harnesses that simulate external API responses, running full end-to-end workflows. Because the model is more reliable, the team checks in much less often. The post gives only qualitative statements; no specific performance metrics or latency figures are disclosed.

Why it matters: Perplexity's cofounder describes GPT-6 Astra in production with concrete scenarios — more substance than a typical customer story. But the post only gives qualitative claims, no perf numbers or latency data, so the score sits right at the featured threshold.

Sep 12Saturday

OpenAI News

Cognition uses GPT‑6 Astra to let Devin test its own code and ship faster

Cognition plugged GPT‑6 Astra into Devin so the AI coding agent can test its own work and return recordings plus reports. One example shows Astra driving Devin to test an iPhone game called Otter Run, returning a simulator recording and a checklist of passed and untested areas. The team also feeds customer bug screenshots to Devin, which fixes the issue and sends back a result screenshot, cutting response time. Co-founder Walden Yan says the goal is less manual code review and more shipping over time. The post doesn't disclose specific performance numbers or latency figures.

Why it matters: GPT‑6 Astra integrated into Devin for self-testing is a concrete workflow landing, not a concept demo. The post provides three scenarios—screen recording, checklist generation, customer bug fixing—with enough detail. Score held below 85 because this is an OpenAI customer story...

Sep 11Friday

GitHub Blog · AI & ML

GitHub Copilot app for Beginners: Using the diff, terminal, and browser

GitHub Copilot 应用内置 diff、终端和浏览器三个面板,让用户无需离开应用即可审查、运行和预览 AI 智能体生成的代码变更。diff 面板以绿色和红色高亮显示代码的增删改,终端面板支持直接运行项目命令并可通过 Run 按钮配置脚本,浏览器面板则提供 Pick & Polish 工具来选取页面元素并让智能体调整。

OpenAI News

Using ChatGPT and Codex to search genomes for new antibiotics

César de la Fuente's lab uses AI to scan genomes of living and extinct organisms for antimicrobial molecules. Their deep-learning models cut candidate search from years to hours; ChatGPT and Codex help write code, process data, and bridge disciplines. About 5 million deaths in 2021 were linked to bacterial antimicrobial resistance, projected to double by 2050. The post doesn't disclose specific candidates found or clinical progress.

Sep 9Wednesday

OpenAI News

OpenAI launches GPT-6 Astra, built for computer use, document work, and cost efficiency

OpenAI launched GPT-6 Astra, a model designed for complex enterprise work. It can directly operate everyday apps like Excel and Figma without APIs. In an Excel modeling challenge, it was about four times faster than the winning human. On Terminal-Bench 4.0, it scored 57.9%, compared to 37.3% for GPT-5.6 Sol and 55.8% for Claude Fable 5.1, with roughly 9% and 63% lower estimated API cost per task. Pricing starts at $10/1M input tokens and $50/1M output tokens. Early customers praised its judgment, deck-building fidelity, and lower hallucination rate. Internally, OpenAI used it to edit a multi-camera video and fix a memory bottleneck, cutting latency by 25x.

Why it matters: OpenAI's GPT-6 Astra release, with direct computer use and speed surpassing human champions, is an industry-shaking event. HKR all hit, score near ceiling. The post doesn't disclose pricing or exact rollout scope — that's the only info gap right now.

OpenAI News

GPT-5.6 Sol runs quantum chip calibrations, freeing MIT grad student from routine lab work

OpenAI published a case study: MIT grad student Beatriz Yankelevich connected GPT-5.6 Sol to lab software to autonomously run calibration measurements on superconducting qubits. The model handled standard sequences—finding frequencies, calibrating pulses, measuring coherence—with little intervention when signals were clean. Weak or noisy signals still required researcher guidance. EQuS now routinely runs agents overnight; researchers check results from their phones. The post doesn't specify hours saved but says a chip previously took days to characterize.

Sep 8Tuesday

OpenAI News

OpenAI CFO: GPT‑6 Astra is here, and consumer + enterprise reinforce each other

OpenAI CFO Sarah Friar published a blog framing GPT‑6 Astra as the world's most capable and aligned model. ChatGPT now has over 1B weekly active users and 2.5M business customers. Internally, the research org uses 3.1 agent-workdays per human workday. The post also claims an internal model solved the Navier–Stokes Millennium Prize Problem, but gives no technical detail. I'd treat this as a strategy narrative, not a technical report.

Why it matters: OpenAI CFO publishes a strategic framing piece for GPT-6 Astra with two concrete numbers: 1B weekly users and a 3.1x agent-workday ratio. Hits all three HKR axes. No technical details — this is narrative, not a product launch — so it stays below 85.

Sep 4Friday

GitHub Blog · AI & ML

GitHub Copilot app for Beginners: Run several agents at once

GitHub Copilot 应用支持同时运行多个智能体会话,每个会话运行在独立的 Git worktree 上,互不干扰且各自保留上下文,可随时切换并从中断处继续。用户可在会话视图中查看各任务标题与进度,例如在同一项目上并行执行 funded sort 开发、无障碍审查和测试运行。

Sep 3Thursday

OpenAI News

Playco cuts manual fixes 50% prototyping games with GPT-6 Astra

Playco built Playbot, an AI-powered IDE for game dev, using GPT-6 Astra. From one grey box prototype, the model generated three themed game worlds in one go, most working on first take. Manual fixes dropped 50% vs the previous model. Spatial reasoning, UI responsiveness, and game feel all improved. The model also plays the game to find bugs itself.

Aug 31Monday

OpenAI News

Polimill builds Japan's next-gen public AI infrastructure with OpenAI, serving 1,050 municipalities

Japanese startup Polimill built QommonsAI, a public-sector AI platform using OpenAI's GPT models and Codex. About 1,050 municipalities and 550,000 public employees now use it. The platform standardizes fragmented administrative data—assembly minutes, welfare records, legal documents—into a cross-municipality searchable knowledge base. Development speed increased 3-5x. Polimill's CAIO says GPT's broad familiarity lowers adoption barriers for government staff. The platform includes audit logs and model access controls for security. Polimill aims to evolve QommonsAI into a shared public OS for all Japanese municipalities.

Aug 25Tuesday

Hugging Face Blog

Quantization-Aware Healing: a 4-bit model that beats its full-precision original

Multiverse Computing introduces Quantization-Aware Healing (QAH), a recovery step for models that have been both structurally compressed and quantized. Applied to a GPT-OSS 120B pruned to 60B and quantized to MXFP4, the 4-bit model beats its bfloat16 original on 7 of 9 benchmarks, including reasoning and math. QAH also outperforms standard QAT on compressed models. The post doesn't disclose latency or throughput numbers, so real-world savings are still TBD.

Why it matters: Counterintuitive compression result: a 4-bit model beats its bfloat16 original on most benchmarks. Method is concrete, numbers are clear, directly useful for deployment and inference folks. Not scoring higher because Multiverse Computing isn't a tier-1 lab, and the post doesn'...

Aug 19Wednesday

OpenAI News

Replit launches Free Mode powered by GPT-5.6 Luna, removing token costs for software creation

Replit introduced Free Mode running on GPT-5.6 Luna, so users can plan, ideate, and explore projects without tracking token spend. CEO Amjad Masad credits recent OpenAI price cuts for making the free tier viable at millions-of-users scale. Complex reasoning tasks get routed to GPT-5.6 Sol, then return to Luna while preserving project context. Sam Altman frames it as a step toward anyone with internet building a product or startup. The post does not disclose Free Mode quotas, concurrency limits, or the exact launch date.

Why it matters: Replit's free tier running GPT-5.6 Luna is a concrete product update with a real mechanism (dual-model handoff) and a direct CEO quote on cost economics — enough signal for featured. But it's an OpenAI customer story, not a model release, so the score stays at 72.

Aug 18Tuesday

OpenAI News

Asana cleared 5 years of engineering work in 2 weeks with Codex

Asana used OpenAI Codex to fully remove Enzyme, an outdated testing framework, from its codebase. The work was originally estimated at five years and roughly $6M; it took two calendar weeks and $12K in model and infrastructure costs. Engineers wrote a five-sentence prompt, ran up to four coding agents in parallel, and reviewed every proposed change twice a day. Asana's CTO noted that not every multi-year project will collapse into weeks, but agents make once-impossible engineering work worth attempting.

Why it matters: Asana used Codex to rip out the Enzyme testing framework — 5 years of estimated work done in 2 weeks, cost dropped from ~$6M to $12K. The numbers carry the story. The post gives a reproducible method, not just PR fluff. Dings: it's an OpenAI official case study, so there's a m...

Aug 6Thursday

OpenAI News

OpenAI publishes first country-by-country ChatGPT usage data: from asking to doing

On Aug 6, OpenAI released its first country-level ChatGPT usage data covering over 1B users. At work, people are more than twice as likely to use ChatGPT to produce output or complete tasks—coding and analysis are typical—compared to outside work. Multimedia is the fastest-growing use case at 7.8% of messages, exceeding 10% in Brazil and Colombia. Latin America, Oceania, and Africa are closing the per-capita adoption gap; Peru, Uruguay, and Costa Rica gained the most in Q2 rankings. Usage among people over 35 rose in nearly every country, with France and Czechia up over 10 percentage points in the past year. Data comes from OpenAI Signals and covers Free, Go, Plus, and Pro individual accounts only.

Why it matters: OpenAI published country-level usage data covering over 1 billion users — 'doing' is twice as likely as 'asking' at work, multimedia messages hit 7.8%, and Latin America is catching up. The data is substantive, but it's an official blog post without third-party verification or...

Jul 31Friday

OpenAI News

OpenAI lays out its “abundant intelligence” playbook: price cuts, efficiency gains, and a full-stack flywheel

OpenAI published a strategy post on July 31 explaining its “abundant intelligence” approach. The core loop: more capable and cheaper models drive broader adoption, which generates revenue and feedback to fund the next round of R&D and infrastructure. Concrete numbers: GPT-5.6 Luna input/output prices dropped 80% to $0.20/$1.20 per million tokens; GPT-5.6 Terra dropped 20%. GPT-5.6 Sol Fast mode delivers 2.5x speed at 2x price with no intelligence change. On the engineering side, Sol helped cut end-to-end serving costs by 20% and improved speculative-decoding efficiency by over 15%. On the public ARC-AGI-3 benchmark, better retained reasoning and context management lifted Sol’s score from 13.3% to 38.3% while using 6x fewer output tokens. Product stats: ChatGPT has over 1B active users and 2M businesses; six months after signup, daily messages rise ~50% and use-case breadth roughly doubles. Agentic work via Codex now accounts for 99.8% of OpenAI’s weekly output tokens. No new model was announced—this is a strategy piece.

Why it matters: OpenAI's official blog lays out its 'abundant intelligence' strategy with concrete pricing data (GPT-5.6 Luna down 80%). Not a product launch, so it doesn't hit 85, but as a strategic signal it's worth featuring.

Jul 29Wednesday

OpenAI News

OpenAI launches ChatGPT for Academic Researchers, giving 100,000 scientists free access to GPT‑5.6

OpenAI is giving 10,000 researchers free access to GPT‑5.6 Sol Pro and Codex this summer, scaling to 100,000 through 2027. Each participant can invite up to four collaborators; data is not used for training by default. The program includes training and hands-on support, and is part of a $250M+ commitment to external research. GPT‑5.6 Sol scores 83% on FrontierMath Tier 4 vs. 72.5% for GPT‑5.5. The post does not spell out eligibility criteria or selection process.

Why it matters: A large-scale free academic rollout with concrete model names and cohort numbers. Capped below 85 because it's a distribution play, not a capability release, and the impact is concentrated in the research community.

Jul 27Monday

OpenAI News

OpenAI study: 43.5% of occupation-specific ChatGPT use crosses job boundaries

OpenAI Economic Research analyzed 800,000+ ChatGPT messages from US users. 16.8% of work messages and 43.5% of occupation-specific messages involve tasks from another occupation—a pattern they call 'task crossover.' Customer experience (77%), design (75%), and HR (69%) workers borrow the most. Marketing and engineering tasks travel farthest across fields. Crossover is more common in small businesses. The report also notes AI is creating new tasks like prompt engineering and output review that don't fit standard job classifications. This is the first paper in the 'Work at the Frontier' series; the full PDF is available.

Why it matters: OpenAI's own research with 800k conversations as the dataset—credible scale. The 43.5% crossover rate is a fresh signal, far more concrete than generic 'AI changes work' narratives. Not an 85 because it's a report, not a product launch or model release—impact is more diffuse.

Jul 17Friday

Google DeepMind

Google DeepMind releases Gemini 3.5 Flash Cyber security model

Google DeepMind released Gemini 3.5 Flash Cyber, fine-tuned from 3.5 Flash to find, verify and patch vulnerabilities quickly. With multiple calls, it approaches larger models on benchmarks such as CyberGym.

Why it matters: It reports how a lightweight security model performs on several benchmarks and inside Google's own codebase, so readers can judge the cost-benefit for vulnerability discovery.

Jul 8Wednesday

OpenAI News

OpenAI audits SWE-Bench Pro, finds ~30% of tasks are broken

OpenAI audited SWE-Bench Pro and estimates ~30% of its tasks are broken. An automated pipeline flagged 286 suspicious tasks; Codex-based investigator agents and five experienced engineers then reviewed them. Engineers identified 249 (34.1%) flawed tasks, mostly due to overly strict tests, underspecified prompts, low-coverage tests, and misleading prompts. OpenAI advises model developers to scrutinize results rather than trust leaderboard scores. The post does not disclose a fix timeline or a revised dataset release.

Why it matters: OpenAI audited SWE-Bench Pro and found ~34% of tasks defective — a ratio that forces the industry to re-examine coding benchmark reliability. The post provides concrete defect categories and a human review pipeline. Not scored higher because this is a benchmark quality report,...

Jun 25Thursday

OpenAI News

OpenAI publishes economic research paper on how Codex is reshaping work

OpenAI released an economic research paper on June 25, using internal and external usage data to track Codex adoption over the past year. By May 2026, 80.6% of sampled individual users had run at least one Codex task estimated to exceed 30 minutes of human work, and 25.6% had run tasks exceeding eight hours. Inside OpenAI, Codex now accounts for 99.8% of weekly output tokens; Legal and Recruiting switched their primary AI tool from ChatGPT to Codex around April 2026. Non-developer users grew fastest—137x for individuals, 189x for organizations. The paper does not disclose Codex pricing or external enterprise conversion rates.

Why it matters: OpenAI's economic research team published a paper quantifying Codex's shift from chat to long-horizon agent tasks, with 80.6% and 25.6% penetration as the core hooks. It's a self-published promotional study, not independent research, so the score stays below 85.

Jun 22Monday

OpenAI News

Samsung Electronics rolls out ChatGPT and Codex to employees in one of OpenAI's largest enterprise deals

Samsung Electronics is deploying ChatGPT Enterprise and Codex to all employees in Korea and its DX division worldwide, covering R&D, manufacturing, marketing, and more. OpenAI calls it one of its largest enterprise launches ever. Codex now has over 5 million weekly active users; weekly actives in Korea grew nearly 800% since Feb 1, 2026. The post does not disclose deal value or rollout timeline.

Why it matters: One of OpenAI's largest enterprise deployments ever, with a concrete 800% Codex WAU spike in Korea. No deal size or timeline disclosed, so it stays at 78 rather than the 85+ band.

Jun 17Wednesday

Hugging Face Blog

Z.AI releases GLM-5.2: first open-source model with solid 1M-token context, built for long-horizon coding tasks

Z.AI open-sourced GLM-5.2, a model built for long-horizon coding tasks. It delivers a genuinely usable 1M-token context—not just accepting more tokens, but maintaining quality across long agent trajectories. IndexShare reuses one indexer across every four sparse attention layers, cutting per-token FLOPs by 2.9× at 1M context; MTP acceptance length improved by up to 20%. On FrontierSWE it beats GPT-5.5 by 1%, and on PostTrainBench it outranks both GPT-5.5 and Opus 4.7, placing second. It's the top open-source model across all three long-horizon coding benchmarks. MIT license, no regional restrictions.

Why it matters: Z.AI open-sources GLM-5.2 with a 1M-token context window and two new architectural components, explicitly targeting long-horizon agent tasks. Domestic flagship model release gets full weight per policy, but the body excerpt lacks full benchmarks, capping it below 85.

Jun 3Wednesday

Alibaba Technology · WeChat

Rethinking R&D Infrastructure When Agents Become First-Class Citizens

Xu Xiaobin argues that agent-based development compresses the intent-to-code loop from weeks or months to minutes, using a weekly-report system, a multi-role agent development setup, and image-repository provisioning as examples; the article identifies mismatches in Git, CI, code review, release flows, permissions, harness setup, and dry-run validation.

Why it matters: HKR-H/K/R all pass, but this is infrastructure commentary rather than a model or product launch. The named cases and week/month-to-minutes claim put it in the 72–77 featured band.

Jun 1Monday

OpenAI News

OpenAI banned a likely PRC-origin cluster using ChatGPT to generate anti-US-data-center social media content

OpenAI's June threat report details a banned cluster of ChatGPT accounts likely originating in China. The operators used Simplified Chinese prompts to generate English posts and images on X, posing as ordinary Americans and claiming data centers and AI are driving up electricity costs for households. They also used ChatGPT for image editing, automation scripts, and harassing overseas dissidents. An internal work report they uploaded outlined tactics for building credible personas on Facebook and evading platform detection.

Why it matters: OpenAI's official threat report details a likely PRC-linked AI influence op with concrete tradecraft and a topic — AI driving up living costs — that's already a public flashpoint. Hits all three HKR axes, but as a security incident report rather than a product or research brea...

May 28Thursday

Mistral AI

Mistral upgrades Le Chat into unified agent Vibe, covering office work and coding

Mistral upgraded Le Chat into a unified AI agent called Vibe, with one license covering both office work and coding. Existing chats, settings and plans all carry over. Work Mode supports enterprise knowledge search, structured data analysis, document and report generation, scheduled multi-step tasks and reusable skills, and connects to Google Workspace, Outlook, SharePoint, Slack, GitHub and more.

Why it matters: It discloses Vibe's Work Mode, coding mode and CLI updates in full, so readers can judge how it plugs into existing workflows.

Alibaba Technology · WeChat

AI-Native Project Management: Two Git Repos Replace Weekly Updates, Insights, and Metrics Reports

Zhou Zhiwei describes a project-management setup that uses two Git repositories, an AI coding assistant, Shell, and Python to replace at least 80% of manual weekly-update chasing, data moving, chart generation, and engineering-metrics reporting.

Why it matters: HKR-H/K/R all pass: the hook is counterintuitive, the post gives an 80% replacement claim and a two-repo mechanism, and it hits engineering-management toil. This is a strong practical workflow piece, not a model or platform launch.

May 16Saturday

Google DeepMind

Google DeepMind releases Gemini 3.5 Flash

Google DeepMind released the Gemini 3.5 model family, with the first model, Gemini 3.5 Flash, available the same day in the Gemini app, Google Search AI Mode, Google Antigravity, the Gemini API and Gemini Enterprise.

Why it matters: Google published 3.5 Flash's coding and agent benchmark scores and where it is available, enough to judge its place in long-horizon workflows.

May 15Friday

Alibaba Technology · WeChat

Qoder 1.0 launches as an agentic development workspace beyond AI IDE

Alibaba released Qoder 1.0 with downloads for Windows, macOS, and Linux, adding a standalone Quest workspace, cross-project parallel agent tasks, a team knowledge engine, and Experts mode with five roles for planning, research, coding, review, and testing.

Why it matters: Alibaba’s Qoder 1.0 is a mid-weight AI coding product release with concrete agent-workflow features and developer resonance. No pricing, benchmark, or task-success data is disclosed, so it stays near the featured threshold.

May 13Wednesday

OpenAI News

Building a Safe, Effective Sandbox for Codex on Windows

OpenAI built a secure sandbox for Codex on Windows. The RSS snippet discloses controlled file access and network restrictions, but the post does not disclose implementation details, performance data, or rollout conditions.

Why it matters: OpenAI details a Windows sandbox for Codex with file-access and network controls. It is not a major model release, but HKR-H/K/R all pass because the safety boundary matters for coding-agent adoption.

May 7Thursday

OpenAI News

Scaling Trusted Access for Cyber with GPT-5.5 and GPT-5.5-Cyber

OpenAI expanded Trusted Access for Cyber to GPT-5.5 and GPT-5.5-Cyber. The RSS snippet says access is for verified defenders; the post does not disclose criteria, pricing, or benchmark data.

Why it matters: HKR-H/K/R all pass: OpenAI expands trusted cyber access to GPT-5.5 and GPT-5.5-Cyber. Kept below 85 because admission rules, pricing, evals, and reproducible tests are not disclosed.

Apr 27Monday

OpenAI News

An Open-Source Spec for Orchestration: Symphony

OpenAI released Symphony, an open-source spec for Codex orchestration. The RSS snippet says it turns issue trackers into always-on agent systems; the post does not disclose spec details, license, APIs, or benchmarks.

Why it matters: HKR-H and HKR-R pass: an OpenAI open-source Codex orchestration spec is relevant to agent workflows. HKR-K is weak because license, interfaces, and reproducible mechanics are not disclosed.

Apr 23Thursday

OpenAI News

Introducing GPT-5.5

OpenAI introduced GPT-5.5 and says it targets complex cross-tool tasks such as coding, research, and data analysis. The RSS snippet only confirms “faster” and “more capable”; the post does not disclose benchmarks, context window, pricing, release timing, or availability, which are the details practitioners should watch.

Why it matters: An OpenAI flagship-model release is same-day news, so HKR-H and HKR-R are clear. HKR-K fails because the post discloses the name and use cases but not benchmarks, context window, price, or availability, so this stays featured rather than p1.

Apr 22Wednesday

OpenAI News

Introducing workspace agents in ChatGPT

OpenAI introduced workspace agents in ChatGPT, describing them as Codex-powered agents that automate complex workflows in the cloud. The RSS snippet confirms secure work across tools for teams, but the post does not disclose pricing, availability, supported tools, or performance metrics.

Why it matters: This is a substantive OpenAI product update inside ChatGPT. HKR-H lands on the jump from chat to workspace agents, HKR-K on Codex-powered cloud execution across tools, and HKR-R on team workflow automation; the score stops at 86 because pricing, rollout, tool support, and metrics

Apr 17Friday

Tencent Technology · WeChat

From Vibe Coding to Agentic Engineering: Rebuilding the Full Backend Development Workflow

Tencent engineers report a one-week practice that used Claude Code plus custom Skills, Commands, and MCP servers to run an 11-stage backend workflow in one terminal session. The post gives reproducible details: one requirement-exploration step used 20 tool calls, 93.8k tokens, and 56 seconds; execution was split into 4 tasks and produced 3 commits. The real point is workflow orchestration, not raw code generation; human review remains at plan, deploy, and review gates.

Why it matters: HKR-H/K/R all pass: the story turns agentic engineering into a measured backend workflow test, with tool-call, token, timing, plan-length, task, and commit data. Stronger than generic coding hype, but still a practitioner case study rather than a major product or model release.

Apr 16Thursday

OpenAI News

Introducing GPT-Rosalind for life sciences research

OpenAI released GPT-Rosalind on April 16, 2026, and made it available as a research preview in ChatGPT, Codex, and the API for qualified customers. The post says it targets biology, drug discovery, and translational medicine, and adds a free Codex life sciences plugin connecting to 50+ scientific tools and data sources. The real signal is deployment breadth: Amgen, Moderna, and Thermo Fisher Scientific are involved, but the post does not disclose model size, pricing, or benchmark scores.

Why it matters: HKR-H lands because OpenAI is shipping a vertical life-sciences model; HKR-K lands on access paths and the 50+ tool/data plugin. HKR-R also lands on the domain-model debate, but missing params, pricing, and benchmark scores keep it at featured, not p1.

Apr 15Wednesday

X · @claudeai

Claude Code on desktop redesigned with side-by-side sessions in one window

Anthropic redesigned Claude Code on desktop and now lets users run multiple Claude sessions side by side in one window. The RSS snippet confirms a new sidebar for session management; the post does not disclose rollout timing, platforms, or more interaction details. For coding workflows, the key question is whether multi-session control cuts context-switch overhead.

Why it matters: An authoritative Anthropic post plus a concrete workflow change gives it HKR-H/K/R. It stays near the featured floor because rollout date, supported desktop platforms, and deeper interaction details are not disclosed, and the scope is still a mid-weight product update.