Skip to content

#OpenAI

47 today

May 29Friday

Xinzhiyuan · WeChat

Claude Opus 4.8 tests split users: strong at high effort, costly under rate limits

The article says Claude Opus 4.8 scores 63 on an Extra-High senior engineering benchmark, 30 points above Opus 4.7, but drops to 42 at High effort, while $200/month Max users report hitting rate limits within hours on complex agent tasks.

Why it matters: Anthropic/Claude relevance plus concrete test numbers clears HKR-H/K/R: the hook is strength versus cost, K has benchmark and quota details, and R hits agent-budget anxiety. Source is a media test rather than an official release, so this lands at low P1.

AI HOT (Curated Pool)

Strengthening Societal Resilience with Rosalind Biodefense

OpenAI launched Rosalind Biodefense and provides trusted GPT-Rosalind access to vetted developers and U.S. government partners; the post does not disclose model parameters or pricing.

Why it matters: HKR-H/K/R all pass: OpenAI launched GPT-Rosalind access for vetted developers and US government partners. Missing parameters, pricing, and eval results keep it below a major capability release.

Latent Space

Anthropic raises $65B Series H, releases Opus 4.8 and Dynamic Workflows

Anthropic announced a $65B Series H at a $965B post-money valuation, disclosed a $47B revenue run rate, and released Claude Opus 4.8 plus Claude Code Dynamic Workflows as a research preview for parallel subagent orchestration.

Why it matters: HKR-H/K/R all pass: this combines a frontier-lab financing event with an Anthropic model and Claude Code workflow release. I score using the summary’s $65B raise and $965B post-money valuation because the title’s dollar figure conflicts with it.

Ruan YiFeng's Weblog

Technology Enthusiasts Weekly Issue 398: Token Costs Are Hard to Afford

Peter Steinberger posted one month of usage showing 7.6 million requests and 603 billion tokens, with CodexBar estimating a $1.3 million value under preset rates rather than his actual spend as an OpenAI employee.

Why it matters: HKR-H/K/R all pass: the CodexBar case turns token economics into concrete usage and cost. This is strong practitioner commentary, not a model or platform release, so it fits the 72–77 featured band.

AI HOT (Curated Pool)

Skill distillation

Skill distillation has Opus 4.7, GPT-5.1, and Gemini 3 Pro write standardized SKILL.md procedure files, while local Qwen 35B and Gemma 26B models execute those files step by step.

Why it matters: HKR-H/K/R pass: the agent-skill distillation pattern is concrete and practitioner-relevant. The summary lacks success rates, cost data, or task outcomes, so it sits at the featured threshold, not must-write.

Financial Times · Technology

Anthropic finalises $65bn funding deal to surpass OpenAI’s valuation

Anthropic finalised a $65bn funding deal that values the Claude AI maker at $965bn including the new money, taking its valuation above OpenAI’s, according to the RSS snippet.

Why it matters: HKR-H/K/R all pass: FT reports Anthropic finalized a $65bn funding deal at a $965bn valuation, overtaking OpenAI. That is a top-model-lab capital reset, above the must-write band.

Bloomberg Technology

Anthropic Raises $65 Billion in Funding Round, Eclipses OpenAI

Anthropic raised $65 billion at a $965 billion post-money valuation, and Bloomberg says the round put its value above OpenAI for the first time.

Why it matters: HKR-H/K/R all pass: $65B raised and a $965B valuation would reorder frontier-lab capital rankings. The item stays below 95 because investor names and terms are not disclosed.

Bloomberg Technology

Anthropic Eclipses OpenAI With Valuation of $965 Billion

Anthropic raised $65 billion at a $965 billion post-money valuation, surpassing OpenAI’s valuation for the first time; the RSS snippet does not disclose investors, deal terms, or a timeline.

Why it matters: HKR-H/K/R all pass: Bloomberg reports a $65B raise at a $965B valuation, putting Anthropic above OpenAI. Investors and terms are not disclosed, but the scale makes it same-day top AI funding news.

May 28Thursday

AI HOT (Curated Pool)

OpenAI Frontier Governance Framework

OpenAI published its Frontier Governance Framework to align its AI safety, security, and risk management practices with new EU and California regulations; the post does not disclose specific evaluation metrics, implementation timelines, or the list of covered frontier models.

Why it matters: HKR-K/R pass: an official OpenAI frontier-governance framework carries safety and regulatory signal, but metrics, timeline, and covered models are not disclosed, so it stays at the lower featured band.

AI HOT (Curated Pool)

Cognition AI raises over $1B, targets 10x software engineering productivity

Cognition AI raised over $1 billion at a $26 billion pre-money valuation, while annualized revenue grew from $37 million to about $492 million in one year, and Devin is positioned as an autonomous junior engineer that can plan, test, and deploy through multi-step workflows.

Why it matters: HKR-H/K/R all pass: the story has hard numbers on funding, valuation, and ARR, plus a direct junior-engineer automation angle. Single-post sourcing keeps it below the 95+ industry-shaking band.

AI HOT (Curated Pool)

OpenAI Products Support Secure Connections to Private MCP Servers

OpenAI supports ChatGPT, Codex, and the Responses API connecting to internal MCP servers through outbound-only HTTPS, while teams keep those servers inside private networks.

Why it matters: HKR-H/K/R pass: OpenAI adds private MCP server support with outbound-only HTTPS, a concrete enterprise agent integration mechanism. Missing permission model, pricing, and rollout details keep it in the lower featured band.

AI HOT (Curated Pool)

I Think Anthropic and OpenAI Found Product-Market Fit

Anthropic and OpenAI changed enterprise pricing around April 2026, moving coding agents from heavily discounted seat plans to API-usage billing, with Anthropic Enterprise at $20 per seat per month plus API fees and OpenAI Codex billed by API token usage.

Why it matters: HKR-H/K/R all pass: the piece ties OpenAI and Anthropic PMF to a concrete billing shift for coding agents. It is influential commentary, not an official launch, so it fits the 78–84 band.

May 27Wednesday

The Verge · AI

AI tried to bury this politician — now people have actually heard of him

Leading the Future, a super PAC funded by OpenAI, Palantir, and a16z executives, has spent millions against NY-12 candidate Alex Bores since late 2025; the snippet says Anthropic and OpenAI will spend millions before the June Democratic primary over who regulates AI and who faces political costs for trying.

Why it matters: HKR-H comes from the backlash angle, HKR-K from named PAC spending millions, and HKR-R from AI lobbying over regulation. This is a strong policy feature, not a same-day industry shock.

AI HOT (Curated Pool)

Runway launches Model Context Protocol server

Runway launched an MCP server that lets compatible agents such as Claude, ChatGPT, and Cursor generate images and videos inside chat interfaces, with access to Gen-4.5, Seedance 2.0, GPT Image 2, Kling 3.0, and Nano Banana Pro.

Why it matters: HKR-H/K/R all pass, but this is a Runway product integration, not an MCP protocol change or model release. It clears featured, with the score kept in the 72–77 band.

QbitAI · WeChat

7B Medical AI Agent Beats o3 and GPT-5 by Learning Where and How to Look

Shanghai Innovation Institute’s LeapQuest and three universities released Ophiuchus and MedScope, applying Think with Images and Think with Videos to medical AI; Ophiuchus-7B scored 68.0 on eight VQA benchmarks, above OpenAI-o3 at 62.2, Gemini 2.5 Pro at 61.8, and GPT-5 at 59.9.

Why it matters: HKR-H/K/R all pass: a 7B model beating o3/GPT-5 is a strong hook, 8 VQA benchmarks with 68.0 vs 62.2 add a testable claim, and medical specialist evaluation will trigger debate. Not a frontier-lab general model release, so it stays in 78–84.

New York Times Chinese

How Google Rebounded and Started Winning the AI Race

Google said regular Gemini users more than doubled in one year to 900 million, while ad revenue rose 16% to $77 billion last quarter, and its Siri partnership with Apple will place Gemini inside future iPhone assistant features.

Why it matters: HKR-H/K/R all pass: NYT ties Google’s comeback narrative to 900M Gemini users, ad growth, and a Siri distribution deal. This is strong industry analysis, not a model launch, so it fits the 78–84 band.

Computing Life · Yage

Using AI Better, Step Two: Write the Skill Before Execution

The author proposes writing a Skill before asking AI to execute a task; each Skill should include three elements—success criteria, observed pitfalls, and deterministic tools—and can be organized through index.md plus AGENTS.md or CLAUDE.md for reuse.

Why it matters: HKR-H/K/R pass via a concrete Skill-first workflow and reusable agent practice. No model release, product capability, or experiment numbers, so it sits at the featured threshold.

AI HOT (Curated Pool)

Claude Mythos reportedly solves OpenAI’s landmark Erdős problem with a “cute simple proof”

Anthropic engineer Sholto Douglas said Claude Mythos solved OpenAI’s Erdős unit distance conjecture problem over the weekend and produced a “cute simple proof”; the RSS snippet does not disclose the proof, verification process, or benchmark setup.

Why it matters: HKR-H/K/R all pass: the claim is clickable, specific, and tied to frontier reasoning rivalry. The post does not disclose the proof, validation process, or Mythos release status, so it stays featured rather than P1.

May 26Tuesday

AI HOT (Curated Pool)

SynthID watermarking expands partnerships, covering over 100 billion content items

Google DeepMind says SynthID has watermarked more than 100 billion content items and is being integrated into models from OpenAI, ElevenLabs, and Kakao, extending prior industry work with NVIDIA.

Why it matters: HKR-H/K/R all pass: the story has a >100B usage number and named integrations with OpenAI, ElevenLabs, and Kakao. It is strong provenance infrastructure news, but still a partnership expansion rather than an 85+ must-write release.

Xinzhiyuan · WeChat

OpenAI Nearly Collapsed? President Says He Resigned the Day Altman Was Ousted

Greg Brockman recounted OpenAI’s 72-hour crisis: on November 17, 2023, the board removed Sam Altman as CEO and took Brockman off the board, after which Brockman resigned the same day and said he initially put the chance of taking the company back at 10%.

Why it matters: HKR-H/K/R all pass via an insider crisis hook, a 10% recovery-odds detail, and OpenAI governance resonance. It is still a retrospective on a heavily covered 2023 event, so it stays in the 72–77 band.

New York Times Chinese

Pope Leo XIV Challenges Silicon Valley and Warns of AI Risks

Pope Leo XIV issued the 42,300-word encyclical Magnifica Humanitas, warning that AI amplifies the power of people with economic resources, expertise, and data access, and calling for regulation and transparency.

Why it matters: HKR-H/K/R all pass: the hook is unusual, the article gives a 42,300-word encyclical and a concrete power-concentration claim, and the topic hits regulation and safety accountability. Not a model, product, or company-moving event, so 78 featured.

AI HOT (Curated Pool)

OpenAI GPT-5.6 Reportedly Set for Next Month With 1.5M-Token Context

Developers found an unannounced OpenAI GPT-5.6 entry in Codex backend logs under the codename iris-alpha, with a 1.5 million-token context window, about 43% higher than GPT-5.5’s 1.05 million-token limit.

Why it matters: HKR-H/K/R all pass: the Codex-log leak, 1.5M-token window, and 43% increase are concrete and practitioner-relevant. It stays below 85 because this is not an official GPT-5.6 launch.

May 24Sunday

AI HOT (Curated Pool)

Greg Brockman: The 72 Hours That Nearly Destroyed OpenAI

The title says Greg Brockman discusses the 72 hours that nearly destroyed OpenAI, but the post does not disclose the timeline, participants, or specific mechanisms behind the crisis.

Why it matters: HKR-H and HKR-R pass: Brockman’s insider account of OpenAI’s near-collapse is clickable and resonant. HKR-K fails because no timeline, actors, or mechanism are disclosed, so it sits at the featured floor.

May 23Saturday

AI HOT (Curated Pool)

Anthropic reportedly nears over $30B funding round, with valuation set to overtake OpenAI

Bloomberg reports that Anthropic is nearing a funding round of over $30 billion, expected to close as soon as next week, pushing its valuation above $900 billion, while the company projects second-quarter revenue of $10.9 billion and its first profitable quarter.

Why it matters: HKR-H/K/R all pass: Bloomberg’s reported $30B+ round, $900B+ valuation and $10.9B Q2 revenue make this a same-day Anthropic capital-race story. It is still not officially closed, so it stays in the lower 85-94 band.

Bloomberg Technology

Anthropic to Close Over $30 Billion Round as Soon as Next Week

Anthropic plans to close a funding round of over $30 billion as soon as next week at a valuation above $900 billion, Bloomberg reported, citing people familiar with the matter, which would put it ahead of OpenAI as the world’s most valuable AI startup.

Why it matters: Bloomberg reports Anthropic may close a $30B-plus round next week at a $900B-plus post-money valuation, a frontier-lab capital-structure story. HKR-H/K/R all pass; the deal is not closed, so it stays below the 95 band.

May 22Friday

MIT Technology Review · AI

Google I/O showed how the path for AI-driven science is shifting

MIT Technology Review says Google used I/O to shift its scientific AI framing toward Gemini for Science, a package that groups AI Co-Scientist and AlphaEvolve, while researchers can now apply for access and older specialized systems like AlphaFold and WeatherNext remain active.

Why it matters: HKR-H and HKR-K pass: MIT Technology Review frames a real Google science-AI product shift with named components and access conditions. HKR-R is weak because the impact is mostly research-facing, not practitioner-wide.

AI HOT (Curated Pool)

OpenAI Codex /goal Feature Officially Launches with Usage Guide

OpenAI moved Codex /goal mode from experiment to stable release, letting users set milestones in the Codex app, IDE extension, or CLI and keep tasks running for hours or days with progress checks, direction changes, and pause controls.

Why it matters: HKR-H/K/R all pass: OpenAI Codex /goal is now stable, with milestones across app, IDE extension, and CLI. The article is thin on permissions, safety limits, and tier access, so it stays in the lower featured band.

Computing Life · Share · Yage

A general-purpose AI model refutes an 80-year-old conjecture

An OpenAI general-purpose reasoning model refuted Erdős’s 1946 unit distance conjecture in the plane; the post says the model was not specially trained for mathematics, and Tim Gowers said he would recommend it to Annals of Mathematics.

Why it matters: HKR-H/K/R all pass: an OpenAI general reasoning model allegedly refuting Erdős’s 1946 conjecture with Tim Gowers approval is same-day material. The summary lacks paper link, proof details, and reproduction conditions, so it stays below 95.

AI HOT (Curated Pool)

Codex Enables Secure Cross-Device Mac Control Around the Clock

OpenAI Devs says Codex can use apps on a Mac from a phone while the Mac remains locked and the screen is off; the post does not disclose permission boundaries, pricing, or a release timeline.

Why it matters: HKR-H/K/R all pass: OpenAI Devs disclosed a concrete Codex Mac-control condition. Missing permission boundaries, pricing, and launch timing keep it below the 85+ band.

Bloomberg Technology

SpaceX Files for Nasdaq IPO | Bloomberg Tech 5/21/2026

Bloomberg says SpaceX filed for a Nasdaq IPO and pitched a $28.5 trillion opportunity spanning AI to Mars; the snippet also says OpenAI is preparing an IPO filing that could arrive as soon as Friday.

Why it matters: HKR-H/K/R all pass: an OpenAI IPO filing as soon as Friday is a high-impact finance node from Bloomberg. The lead is still SpaceX, and OpenAI valuation, deal size, and filing link are not disclosed, so this lands at 88.

May 21Thursday

MIT Technology Review · AI

Anthropic’s Code with Claude Showed Off Coding’s Future—Whether You Like It or Not

Anthropic used its two-day Code with Claude event in London to show Claude Code automation, with nearly half the room saying they shipped a pull request fully written by Claude in the past week, and many keeping their hands raised when asked whether they had shipped it without reading the code.

Why it matters: HKR-H/K/R all pass: the MIT Tech Review piece has a strong Claude Code hook, a concrete developer-behavior number, and clear resonance for programmers. It is not a model release or major product launch, so it stays in the 78–84 band.

The Verge · AI

Musk v. Altman: Much Ado About Nothing

The jury found Elon Musk’s lawsuit was filed after the statute of limitations had expired, while the case centered on OpenAI’s shift from nonprofit to for-profit status and whether Musk lost money; the post does not disclose damages awarded or next legal steps.

Why it matters: HKR-H/K/R all pass: the OpenAI-Musk-Altman feud has drama, a concrete time-bar ruling, and governance resonance. Thin facts on damages or next legal moves keep it in low featured, not 78+.

Xinzhiyuan · WeChat

Anthropic Acquires SDK Toolmaker Stainless, Leaving OpenAI and Others to Maintain SDKs

Anthropic has completed its acquisition of Stainless, an SDK generation company used by OpenAI, Anthropic, Meta, Cloudflare, and other infrastructure vendors; Stainless says prior SDK ownership remains with customers, but it will shut down hosted products including SDK generator and stop providing ongoing support.

Why it matters: HKR-H/K/R all pass: the deal targets API SDK generation, names OpenAI/Meta/Cloudflare as customers, and says hosted products will shut down. Anthropic bump applies, but this is not a model or core capability release, so it fits 78–84.

Latent Space

OpenAI GPT-next Disproves 80-Year-Old Erdős Planar Unit Distance Problem for Under $1000

OpenAI said an internal general-purpose reasoning model disproved the 1946 Erdős planar unit distance problem by finding a new family of constructions; the reasoning summary reportedly spans about 125 pages, while outside observers speculate the run used under 32 hours or under $1,000.

Why it matters: HKR-H/K/R all pass: an OpenAI internal reasoning model allegedly refuting the 1946 Erdős problem with ~125 pages is a major capability signal. Cost and runtime are still external estimates, keeping it below 95.

Synced · WeChat

Zhipu deploys ZCube, raising inference throughput 15% on the same GPUs

Zhipu deployed ZCube in a thousand-GPU GLM-5.1 production inference cluster, replacing ROFT while keeping GPUs, software stack, and business code unchanged; throughput rose by over 15%, TTFT P99 fell 40.6%, and switch plus optical module costs dropped by one third.

Why it matters: HKR-H/K/R all pass: Zhipu reports ZCube in a GLM-5.1 1k-GPU production inference cluster with +15% throughput and 40.6% lower TTFT P99. Single-source infra optimization keeps it below major model-release weight.

Hacker News front page

OpenAI to Confidentially File for IPO as Soon as Friday

The title says OpenAI will confidentially file for an IPO as soon as Friday; the RSS body only includes the CNBC URL, Hacker News with 41 points and 2 comments, and does not disclose valuation, offering size, or listing timetable.

Why it matters: HKR-H/K/R all pass: an imminent OpenAI confidential IPO filing is a top-band foundation-model-company IPO event. The post is thin and lacks valuation or raise size, but the stated timing keeps it P1.

AI HOT (Curated Pool)

OpenAI Model Independently Solves 80-Year-Old Math Problem

An OpenAI AI model solved the plane unit distance problem proposed in 1946, using Golod-Shafarevich theory to produce a family of more efficient constructions.

Why it matters: HKR-H/K/R all pass, but the item is only an X summary and lacks model name, paper link, reproducibility, and third-party verification. Strong OpenAI reasoning-research signal, kept below P1.

Financial Times · Technology

Anthropic on Track for First Profitable Quarter

Anthropic is on track to record its first profitable quarter ahead of OpenAI and xAI; the RSS snippet does not disclose the quarter, revenue, profit figure, or accounting basis.

Why it matters: HKR-H/K/R all pass: the FT claim reframes Anthropic’s business race against OpenAI and xAI. Missing quarter, revenue, and profit figures keeps it below P1.

AI HOT (Curated Pool)

OpenAI may file draft IPO prospectus as soon as Friday, targeting September 2026 listing

OpenAI expects to file a draft IPO prospectus as soon as Friday, with Sam Altman setting a September 2026 listing target and the company carrying a private valuation above $850 billion.

Why it matters: HKR-H/K/R all pass: an OpenAI draft IPO filing with a $850B+ valuation and September target is a must-write capital-markets story. It stays below the top band because this is still a single-source expected filing, not the filed prospectus.

r/LocalLLaMA

HalBench: Custom sycophancy and hallucination benchmark tests 4 frontier models

HalBench tested 4 frontier models on 3,200 false-premise prompts, with Sonnet 4.6 ranking first at a 0.565 mean score and Gemini 3.1 Pro last at 0.339; higher scores mean the model more often named the false premise and pushed back instead of complying.

Why it matters: HKR-H/K/R all pass: HalBench has a clear custom-eval hook, 3,200 prompts with scores, and a live trust/safety angle. Single Reddit sourcing and an unvalidated benchmark keep it at the low featured band.