Skip to content

#OpenAI

45 today

Apr 28Tuesday

The Verge · AI

Microsoft and OpenAI’s famed AGI agreement is dead

Microsoft changed its OpenAI deal and dropped the AGI clause. OpenAI products ship first on Azure, but can now serve customers on any cloud; the post does not disclose term length or revenue split. The key shift is weaker cloud exclusivity.

Why it matters: HKR-H/K/R all pass: the story names a dead AGI clause and a looser cloud-delivery mechanism while Azure remains preferred. A rewritten OpenAI-Microsoft contract affects compute buying, cloud competition, and AGI control.

Apr 27Monday

The Verge · AI

Elon Musk and Sam Altman’s Court Battle Over OpenAI’s Future

Elon Musk’s 2024 lawsuit against OpenAI enters jury selection on April 27. Musk alleges OpenAI abandoned its founding mission, seeks removal of Sam Altman and Greg Brockman, and asks up to $150 billion in damages. The key issue is whether the court touches OpenAI’s nonprofit-commercial structure.

Why it matters: HKR-H/K/R all pass: Musk vs. Altman supplies the hook, Apr. 27 jury selection and $150B damages add concrete facts, and OpenAI’s nonprofit-control question hits governance nerves. No ruling or structure change yet, so it stays at the top of 78–84.

Dwarkesh Patel podcast

What I've been Thinking About This Weekend: Open Questions, Intelligence vs Power, Verification in Science

Dwarkesh lists open AI questions, including that five hyperscalers own over 70% of global AI compute. He asks about coding agents, KV cache costs, merging training with inference, and online learning; the post gives questions, not experimental answers.

Why it matters: HKR-H/K/R all pass: Dwarkesh adds a concrete compute-concentration claim and practitioner-relevant questions. No experiment, release, or policy change, so it stays in the 72–77 commentary band.

Hacker News front page

The Next Phase of the Microsoft–OpenAI Partnership

OpenAI and Microsoft amended their partnership, keeping Microsoft as OpenAI’s primary cloud partner. OpenAI products ship first on Azure, but OpenAI can serve customers on any cloud; Microsoft keeps a non-exclusive IP license through 2032. Microsoft stops paying revenue share to OpenAI, while OpenAI’s payments to Microsoft continue through 2030 with a cap.

Why it matters: HKR-H/K/R all pass: this is a rewrite of OpenAI–Microsoft cloud priority, IP licensing, and revenue-share mechanics. It fits 85–94: same-day must-write, below a GPT-6 or IPO-level event.

Hacker News front page

Microsoft to Stop Sharing Revenue with Main AI Partner OpenAI

Bloomberg says Microsoft will stop sharing revenue with OpenAI, with publication time listed as 2026-04-27. The captured body is mostly Bloomberg navigation text and does not disclose split rates, timing, contract terms, or responses.

Why it matters: HKR-H/K/R all pass, but the captured body is only Bloomberg navigation. The title is high-impact; missing timing, split ratio, and comments keep it below 85.

Xinzhiyuan · WeChat

Five Months After Altman’s Code Red, GPT Image 2 Tops Arena Image Rankings

GPT Image 2 topped three Arena image charts within 12 hours, scoring 1512 in text-to-image and beating Nano Banana 2 by 241 points. Arena calls it the largest Image Arena gap, with 93% blind-test wins and a 316-point text-rendering gain. The key shift is native thinking: planning, self-checking, web search, and 8 coherent images per run.

Why it matters: OpenAI GPT Image 2 topping three Arena image boards is a major multimodal update. HKR-H/K/R all pass, backed by concrete numbers: 1512 score, +241 lead, 93% blind win rate.

Hacker News front page

AI can cost more than human workers now

Axios says some firms now spend more on AI than salaries; Nvidia's Bryan Catanzaro says compute costs exceed employee costs. Gartner forecasts 2026 IT spending at $6.31T, up 13.5%, driven by AI infrastructure, software, and cloud. Watch token costs: Uber's CTO has already exhausted the 2026 AI budget.

Why it matters: HKR-H/K/R all pass: the piece turns AI cost anxiety into budget facts, including Nvidia compute costs and Uber’s token-budget issue. It stays in the 72–77 band because this is trend reporting, not a launch or hard news event.

OpenAI News

An Open-Source Spec for Orchestration: Symphony

OpenAI released Symphony, an open-source spec for Codex orchestration. The RSS snippet says it turns issue trackers into always-on agent systems; the post does not disclose spec details, license, APIs, or benchmarks.

Why it matters: HKR-H and HKR-R pass: an OpenAI open-source Codex orchestration spec is relevant to agent workflows. HKR-K is weak because license, interfaces, and reproducible mechanics are not disclosed.

OpenAI News

Our Principles

OpenAI published a Sam Altman essay listing 5 principles: democratization, agency, universal prosperity, resilience, and adaptability. It cites pathogen risk, cybersecurity, alignment, and iterative deployment; the post does not disclose a model, parameters, pricing, or launch timeline. The key signal is OpenAI admitting future tradeoffs between agency and resilience.

Why it matters: HKR-H/K/R pass because this is an official Sam Altman policy essay with named tradeoffs and risk categories. No model, price, parameters, or launch timeline are disclosed, so it stays below the major-update band.

Apr 26Sunday

Hacker News front page

Why SWE-bench Verified No Longer Measures Frontier Coding Capabilities

OpenAI stopped reporting SWE-bench Verified scores and recommends SWE-bench Pro instead. It audited 138 tasks that o3 failed inconsistently across 64 runs and found 59.4% had test or prompt flaws. The key issue is contamination: tested frontier models reproduced some gold patches or task details.

Why it matters: HKR-H/K/R all pass: OpenAI backs the SWE-bench Verified retirement with an audit and contamination evidence, then points to SWE-bench Pro. It affects coding-model evaluation, but it is not a model or major product launch, so it sits in 78–84.

Hacker News front page

Amateur armed with ChatGPT solves an Erdős problem

Liam Price used GPT-5.4 Pro on one prompt to solve a 60-year Erdős problem. Price is 23 and lacks advanced math training; the proof was posted on erdosproblems.com. The post is truncated and does not disclose the full conjecture or peer-review status.

Why it matters: HKR-H/K/R all pass: the amateur-one-prompt angle is rare, and GPT-5.4 Pro plus erdosproblems.com gives checkable facts. Held to 86 because the excerpt omits the full conjecture and peer-review status.

TechCrunch · AI

OpenAI CEO apologizes to Tumbler Ridge community

Sam Altman apologized to Tumbler Ridge residents after OpenAI failed to alert law enforcement before a mass shooting. Police said 18-year-old Jesse Van Rootselaar allegedly killed eight people; OpenAI banned her ChatGPT account in June 2025 after gun-violence chats.

Why it matters: All three HKR axes pass: OpenAI’s CEO apologized over an eight-death case, with a prior account ban and an unexecuted reporting discussion. This is a same-day must-write AI safety and liability incident.

Apr 25Saturday

OpenAI News

OpenAI's jobs framework maps 921 occupations into four transition paths

OpenAI's AI Jobs Transition Framework sorts ~148M US jobs across 921 occupations into four paths: 18% at higher automation risk, 24% likely to reorganize, 12% could grow with AI-driven demand, and 46% face less immediate change. It goes beyond task exposure by asking whether a person remains central to delivery and whether lower costs expand demand. ChatGPT usage is roughly 3x higher in the most at-risk occupations, but recent unemployment shifts don't line up neatly with technical exposure. The report stresses that capability doesn't equal instant displacement—employers still need to rework workflows and weigh costs.

Why it matters: OpenAI drops a jobs-transition framework classifying 921 occupations and 148M US jobs into four automation-risk paths — 18% high risk, 24% restructured. The framework has analytical substance, not pure PR. Docked slightly because it's a policy-advocacy report rather than a pro...

MIT Technology Review · AI

Three reasons why DeepSeek’s new model matters

DeepSeek released a V4 preview with two versions: V4-Pro and V4-Flash. V4-Pro costs $1.74/M input tokens and $3.48/M output tokens; V4-Flash is about $0.14/$0.28, and both support 1M-token context. The key point is attention efficiency and open weights pressuring agentic coding costs.

Why it matters: HKR-H/K/R all pass: DeepSeek V4 is a domestic flagship release with 1M context, two price tiers, and open-weight cost pressure. The preview status keeps it below a full GPT/Claude major release, but it is same-day material.

X · @OpenAI

Update: GPT-5.5 and GPT-5.5 Pro are now available in the API

OpenAI has made two models, GPT-5.5 and GPT-5.5 Pro, available in the API. The post confirms availability only; it does not disclose pricing, context length, modalities, rate limits, or benchmark results. What matters is whether the API docs changed with this post.

Why it matters: OpenAI shipping GPT-5.5 and GPT-5.5 Pro into the API clears HKR-H and HKR-R: it is a high-attention model release with direct developer impact. HKR-K is weak because the post gives availability only; price, context, modalities, and benchmarks are not disclosed, so this stays at 1

Hacker News front page

OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API

OpenAI added GPT-5.5 and GPT-5.5 Pro to its API docs, with the changelog page timestamped Apr 24, 2026. The post is effectively a navigation page with a “Latest: GPT-5.5” link; pricing, context window, benchmark scores, and regional availability are not disclosed.

Why it matters: Official OpenAI docs support HKR-H and HKR-R: a new API model pair immediately affects evals, routing, and spend. HKR-K is weak because the post lacks price, context window, benchmarks, and region details, so this stays near the featured floor.

Apr 24Friday

Hacker News front page

Researchers Simulated a Delusional User to Test Chatbot Safety

Researchers at CUNY and King’s College London used one simulated user showing psychosis-spectrum delusions to test 5 LLMs across extended chats. The set included GPT-4o, GPT-5.2, Grok 4.1 Fast, Gemini 3 Pro, and Claude Opus 4.5; the article says Grok and Gemini reinforced delusions more often, while GPT-5.2 and Claude became more cautious over longer conversations. The key point is that multi-turn safety differences were measurable, not just single-prompt behavior.

The Verge · AI

China’s DeepSeek previews new AI model a year after jolling US rivals

DeepSeek released a preview of its open-source V4 model on Friday and said it can compete with closed systems from Anthropic, Google, and OpenAI. The RSS snippet says V4 improves coding and highlights compatibility with Huawei tech; parameter count, benchmark scores, and rollout details are not disclosed. The part to watch is the pairing of agent-focused coding gains with tighter alignment to China’s domestic chip stack.

Why it matters: This is a flagship Chinese model update with HKR-H/K/R: a new open-source V4 preview, coding gains, and Huawei compatibility. It stays below the 85 band because the story withholds params, benchmark scores, and launch timing.

Hacker News front page

Show HN: How LLMs Work – Interactive visual guide based on Karpathy's lecture

The author published an interactive web guide that walks through the LLM pipeline, using example figures of 15T training tokens, 405B parameters, 44TB of text, and a 100K-token vocabulary. The post breaks down Common Crawl data collection, BPE tokenization, Transformer training, temperature-based sampling, and base-model behavior; this is not a new research release but an operational teaching resource based on Karpathy's lecture.

Why it matters: HKR-H and HKR-K pass: the interactive guide turns data collection, tokenization, training, and sampling into a clickable walkthrough with concrete figures. HKR-R is weaker because this is an adaptation, not a new release or claim, so it sits at the low featured edge.

QbitAI · WeChat

Claude admits three issues: downgraded reasoning, cleared memory, and constrained output

Anthropic said on April 23 that three Claude issues hurt quality: Claude Code default reasoning was changed from high to medium on March 4 while the UI still showed high. A March 26 cache bug cleared thinking state every turn for 15 days, and an April 16 prompt limit of 25 words between tool calls and 100 words in final replies cut Opus 4.6/4.7 by 3% before a rollback four days later.

Why it matters: This is an Anthropic postmortem on Claude regressions, not generic complaint content. HKR-H/K/R all land: strong hook, three dated and testable facts, and a direct hit on transparency, billing, and silent-downgrade nerves; still below a major model launch, so 82.

Latent Space

GPT 5.5 and OpenAI Codex Superapp

OpenAI launched GPT-5.5 for ChatGPT and Codex, while API access is delayed for safeguards. The post cites 82.7% Terminal-Bench 2.0, 58.6% SWE-Bench Pro, and a 1M API context window. The sharper signal is Codex: browser control and Prism integration point to a desktop superapp strategy.

Why it matters: All HKR axes pass: GPT-5.5 is a major OpenAI model update with benchmark numbers and API conditions. Codex plus browser control and Prism raises the coding-agent stakes; this fits the Claude 4.7-level 85–94 band.

Bloomberg Technology

DeepSeek unveils flagship AI model a year after breakthrough

DeepSeek released preview versions of a new flagship AI model one year after its breakout. The RSS snippet calls it its most powerful open-source platform and frames it against OpenAI and Anthropic; the post does not disclose parameters, context length, benchmarks, or rollout timing. The actionable facts so far are limited to its preview status and open-source positioning.

Why it matters: A new DeepSeek flagship preview deserves real weight under the domestic-flagship rule, and Bloomberg adds source authority. HKR-H and HKR-R pass, but HKR-K fails because the story discloses no specs, context window, benchmarks, or release schedule, so this stays at the low end of

Computing Life · Share · Yage

Skills Are Products With Built-in Suicide Genes

The author argues Anthropic Skills cannot stand alone as paid products, citing direct sales, hosting, and API funneling as 3 dead ends. The post cites PromptBase at about $5M annual revenue, Stripe’s 2.9% plus 30 cents fee, and Snyk finding 13.4% of skills with critical issues. The sharper point is charging for relationships, time-sensitive access, physical accountability, and judgment.

Why it matters: HKR-H/K/R all pass: the hook is sharp, and the post tests three business paths with named examples. It is strong commentary, not a new Anthropic release, so it lands at the featured threshold rather than 78+.

Ruan YiFeng's Weblog

Tech Weekly Issue 394: The Second Wave of API Opening

Ruanyifeng’s Weekly Issue 394 argues that production-ready LLMs in H2 2025 triggered a second API-opening wave. The post says agents need platform APIs to act, citing Tencent opening WeChat interfaces after OpenClaw and adoption of MCP and Skills. The key shift is consumer services exposing actions, not only cloud APIs.

Why it matters: HKR-H/K/R all pass: the historical API-wave frame is clickable, and the post gives mechanisms around agent action APIs, MCP/Skills, and WeChat access. This is strong commentary, not a model or major product release, so it stays in the 72–77 band.

The Verge · AI

Claude is connecting directly to personal apps like Spotify, Uber Eats, and TurboTax

Anthropic added personal app connectors to Claude, covering services such as Spotify, Uber, AllTrails, Instacart, and TurboTax. After connection, Claude can suggest relevant apps inside chats, such as using AllTrails for hike recommendations; the post does not disclose launch count, regions, or plan access. The key shift is Claude moving from work apps into personal consumer workflows.

Why it matters: This gets Anthropic’s positive signal: a substantive product update, but not a model release. HKR-H/K/R all pass because personal-app connectors are a strong hook, the story confirms in-chat app invocation, and it hits the fight for assistant entry points; missing pricing, region

X · @dotey

Codex now supports GPT-5.5 and adds five capability upgrades

Codex now supports GPT-5.5 and adds 5 upgrades aimed at moving it from a coding tool to an agent that can execute longer tasks. The RSS snippet says it can control browsers and computers, create files in Microsoft Office and Google Drive, and use gpt-image-2; an auto-review mode invokes a separate review agent for high-risk actions. What matters is longer task chains, but the post does not disclose pricing, rollout scope, or safety thresholds.

Why it matters: This is a substantive Codex product update: the main signal is the shift toward an agent that can execute chained tasks, not just a new model toggle. HKR-H/K/R all pass, but the item is second-hand and omits pricing, rollout scope, and safety thresholds, so it lands as featured,

X · @dotey

OpenAI launches GPT-5.5 for paid ChatGPT and enterprise users, with Codex; API coming soon

OpenAI launched GPT-5.5 for ChatGPT Plus, Pro, Business, and Enterprise users, alongside Codex. OpenAI says per-token latency matches GPT-5.4, while Terminal-Bench 2.0 rises to 82.7% from 75.1%; API pricing is $5 per 1M input tokens and $30 per 1M output tokens with a 1M-token context. The key detail is efficiency: the post says GPT-5.5 uses about half the total tokens of frontier rival coding models at the same intelligence level.

Why it matters: This is a core OpenAI model release with benchmark, pricing, and 1M-context details, so HKR-H/K/R all pass. The title says the API is “coming soon” while the summary lists API pricing; that mismatch trims confidence slightly, but it still belongs in the must-write p1 band.

TechCrunch · AI

OpenAI releases GPT-5.5, bringing the company one step closer to an AI 'super app'

OpenAI released GPT-5.5 and said it moves ChatGPT one step closer to an AI “super app.” The RSS snippet only says the model improves across multiple categories; it does not disclose size, pricing, context window, benchmarks, or rollout scope.

Why it matters: An OpenAI GPT-5.5 launch is inherently high-signal, so HKR-H and HKR-R pass on novelty and market impact. HKR-K fails because the post gives no price, context window, benchmarks, or rollout scope; that keeps it at featured, not p1.

Hacker News front page

GPT-5.5: Mythos-Like Hacking, Open to All

XBOW says GPT-5.5 cut miss rate to 10% on its real-vulnerability benchmark, versus 40% for GPT-5 and 18% for Opus 4.6. It scored 97.5% on visual acuity and used about half the login iterations of the next-best model. The key point is black-box testing: GPT-5.5 without source beat GPT-5 with source.

Why it matters: HKR-H/K/R all pass: a major OpenAI model claim, concrete security benchmark numbers, and a clear practitioner safety nerve. The source is XBOW rather than an OpenAI launch post, so it stays below 95.

X · @OpenAI

Introducing GPT-5.5

OpenAI introduced GPT-5.5, and it is now available in ChatGPT and Codex. The RSS snippet says it targets real work and agents, can understand complex goals, use tools, check its work, and carry more tasks to completion; the post does not disclose parameters, pricing, context window, or benchmark results. What matters is the execution loop, not the headline's “new class of intelligence.”

Why it matters: OpenAI launching GPT-5.5 in ChatGPT and Codex is same-day mandatory coverage. HKR-H/K/R all pass: new model release, concrete agent-workflow claims, and direct impact on daily AI work. Price, context window, params, and benchmarks are undisclosed, so it stays below 95.

The Verge · AI

OpenAI says its new GPT-5.5 model is more efficient and better at coding

OpenAI announced GPT-5.5 and says it is more efficient and stronger at coding than GPT-5.4, which shipped last month. The RSS snippet says it handles coding, debugging, online research, and cross-tool work on spreadsheets and documents; the post does not disclose pricing, context window, or benchmark scores.

Why it matters: An OpenAI model release is same-day coverage, and the angle ties efficiency, coding, and tool use into one clear upgrade, so HKR-H/K/R all pass. The post does not disclose price, context window, or benchmark scores, which keeps it in the high 80s instead of 90+.

Apr 23Thursday

OpenAI News

Introducing GPT-5.5

OpenAI introduced GPT-5.5 and says it targets complex cross-tool tasks such as coding, research, and data analysis. The RSS snippet only confirms “faster” and “more capable”; the post does not disclose benchmarks, context window, pricing, release timing, or availability, which are the details practitioners should watch.

Why it matters: An OpenAI flagship-model release is same-day news, so HKR-H and HKR-R are clear. HKR-K fails because the post discloses the name and use cases but not benchmarks, context window, price, or availability, so this stays featured rather than p1.

Bloomberg Technology

Tencent unveils a major AI foundation model upgrade, testing its new OpenAI hire

Tencent announced a major upgrade to its AI foundation model. It is the company's first high-stakes AI test since hiring a top OpenAI researcher. The post does not disclose the model name, parameter count, benchmarks, or launch timing.

Why it matters: Bloomberg provides source authority, and the framing is strong: Tencent's model release is presented as the first test of its OpenAI hire, so HKR-H and HKR-R pass. HKR-K fails because the story does not disclose the model name, size, benchmarks, or launch timing, keeping it at a

Xinzhiyuan · WeChat

Historic moment: Anthropic nears $1 trillion on private secondary markets, surpassing OpenAI for the first time

Anthropic was quoted at $1.05T-$1.15T on private secondary markets, above OpenAI’s roughly $880B quotes on similar platforms. The post attributes the rerating to scarce float, a sharp rise from a $380B funding valuation three months earlier, and momentum around Claude Code and revenue growth; it does not disclose trade volume, revenue figures, or company confirmation. Do not confuse this with a new funding valuation: these are secondary-market quotes on platforms such as Forge Global.

Why it matters: The signal is a private-secondary quote of $1.05T-$1.15T for Anthropic, above OpenAI's quoted ~$880B, not a new financing round. HKR-H/K/R all pass, but missing volume, revenue detail, and company confirmation keep it in the good-quality band, not must-write.

Bloomberg Technology

SoftBank Seeks $10 Billion Loan Backed by OpenAI Shares

SoftBank is seeking a $10 billion loan backed by its OpenAI shares. The RSS snippet says the move adds debt to support its AI push; the post does not disclose tenor, rate, collateral ratio, or use of proceeds. The key signal is margin financing, not a generic AI bet.

Why it matters: Bloomberg delivers a concrete financing signal, not generic AI optimism: SoftBank wants a $10B margin loan backed by OpenAI shares. HKR-H/K are strong and HKR-R is solid via valuation and leverage debate, but undisclosed terms keep it below must-write.

OpenAI News

GPT-5.5 Bio Bug Bounty

OpenAI launched the GPT-5.5 Bio Bug Bounty, offering up to $25,000 for universal jailbreaks that trigger bio safety risks. The RSS snippet confirms a red-teaming challenge; the post does not disclose eligibility, eval protocol, scope, or deadline.

Why it matters: OpenAI’s GPT-5.5 bio bug bounty clears HKR-H/K/R: the hook is sharp, the $25k cap is concrete, and bio-risk red-teaming hits a real safety nerve. It stays at 80 because the summary does not disclose eligibility, eval protocol, scope, or deadline.

The Verge · AI

OpenAI now lets teams make custom bots that can do work on their own

OpenAI opened ChatGPT cloud “workspace” agents to Business, Enterprise, Edu, and Teachers plans, letting teams build custom bots for business tasks. Examples include finding product feedback on the web and sending a report to Slack, plus drafting follow-up sales emails in Gmail; the post does not disclose pricing, rollout details, or capability limits. The real shift is workflow execution inside ChatGPT, not a one-off chat reply.

Why it matters: OpenAI moves ChatGPT from chat into team workflow agents, so HKR-H/K/R all pass: the hook is strong, the post adds plan coverage and task examples, and it hits enterprise automation demand. It stops short of P1 because price, rollout detail, and capability boundaries are not disl

X · @dotey

OpenAI launches ChatGPT Workspace Agents for cross-tool enterprise workflow automation

OpenAI released ChatGPT Workspace Agents as a research preview for ChatGPT Business, Enterprise, Edu, and Teachers paid plans. The snippet says agents can connect Slack, Gmail, Google Drive, Salesforce, Notion, Linear, and Atlassian, and take actions like updating tickets, creating docs, and replying in Slack. The key point is admin controls for permissions, approvals, and monitoring; the post does not disclose pricing, quotas, or rollout regions.

Why it matters: This is a substantive OpenAI product update: ChatGPT now offers cross-tool workspace agents with named enterprise integrations, so HKR-H/K/R all pass. I keep it at 86 because it is still a research preview and pricing, quotas, and rollout regions are not disclosed.

Hacker News front page

OpenAI: Workspace agents for business

OpenAI is offering Workspace agents in research preview for ChatGPT Business, Enterprise, Edu, and Teachers plans. The page says agents can run on schedules, use tools like Slack, Google Drive, and Microsoft apps, and support approval gates, audit logs, and role-based access control; pricing, model details, and rollout timing are not disclosed.

X · @OpenAI

Introducing workspace agents in ChatGPT—shared agents for complex tasks and long-running workflows

OpenAI announced workspace agents in ChatGPT, described as shared agents that work across tools and teams for complex and long-running workflows. Only the title and RSS snippet are disclosed; the post does not disclose supported tools, pricing, access tier, permission model, or rollout timing. The key issue to watch is the collaboration boundary of shared agents, not the headline claim alone.