Skip to content

OpenAI / ChatGPT

Everything OpenAI: the GPT models, ChatGPT and Sora, company strategy and people moves.

Latest picks

1121–1140 of 1,549

May 6Wednesday

The Verge · AI

OpenAI claims ChatGPT’s new default model hallucinates way less

OpenAI says ChatGPT’s default GPT-5.5 Instant reduced hallucinations in internal evaluations. Versus GPT-5.3 Instant, hallucinated claims fell 52.5% on high-stakes prompts. Inaccurate claims fell 37.3% on flagged hard chats; the post does not disclose full eval size.

Why it matters: OpenAI changed ChatGPT’s default model and gave two hallucination-reduction figures, satisfying HKR-H/K/R. Internal evals lack set size and reproduction details, but a default ChatGPT model change is same-day material.

TechCrunch · AI

OpenAI releases GPT-5.5 Instant, a new default model for ChatGPT

OpenAI released GPT-5.5 Instant as ChatGPT’s new default model. The company says it reduces hallucinations in law, medicine, and finance while keeping prior low latency; the post does not disclose benchmarks, rollout scope, or pricing.

Why it matters: HKR-H/K/R all pass: a new ChatGPT default model, testable reliability claims, and direct workflow impact. Missing eval numbers, rollout scope, and pricing keep it in the mid 85–94 band.

May 5Tuesday

The Verge · AI

OpenAI is reportedly launching a phone for ChatGPT

Ming-Chi Kuo says OpenAI is fast-tracking a ChatGPT phone for mass production in early 2027. It reportedly uses a customized MediaTek Dimensity 9600 with enhanced-HDR ISP; the post does not disclose price, design, or OS details.

Why it matters: HKR-H/K/R all pass, but this is a Kuo report rather than an OpenAI launch. Missing price, form factor, and OS details keep it below must-write territory.

OpenAI News

OpenAI Introduces MRC for Large-Scale AI Training Networks

OpenAI introduced MRC for large-scale AI training cluster networks. MRC stands for Multipath Reliable Connection and is released via OCP to improve resilience and performance; the post does not disclose throughput, latency, or cluster size.

Why it matters: HKR-H/K/R pass: OpenAI shared MRC via OCP, with a concrete multipath reliability mechanism. No throughput, latency, or cluster scale is disclosed, so this stays in the 72–77 featured band.

OpenAI News

GPT-5.5 Instant: smarter, clearer, and more personalized

OpenAI updated ChatGPT’s default model to GPT-5.5 Instant for default chat use. The RSS snippet says answers are more accurate, hallucinations are reduced, and personalization controls improved; the post does not disclose metrics, pricing, or context window.

Why it matters: HKR-H/K/R all pass: OpenAI changed ChatGPT’s default model to GPT-5.5 Instant. The post lacks evals, pricing, and context window details, so it stays at the low end of the 85–94 band.

OpenAI News

GPT-5.5 Instant System Card

OpenAI published a GPT-5.5 Instant system card; the title confirms one model version. The post body is empty and does not disclose eval scores, safety limits, context window, or release date.

Why it matters: HKR-H and HKR-R pass because an official GPT-5.5 Instant card is a strong OpenAI hook. HKR-K fails: the body has no evals, safety limits, context window, or release details, so this stays at the featured floor.

r/LocalLLaMA

DeepSeek V4 Pro matches GPT-5.2 on FoodTruck Bench, 10 weeks later and about 17x cheaper

DeepSeek V4 Pro ranked No. 4 on FoodTruck Bench. The 30-day agentic benchmark uses 34 tools, persistent memory, and daily reflection; its median is within 3% of GPT-5.2 at about 17x lower workload cost. Xiaomi MiMo v2.5 Pro also ranked No. 6, with 5/5 survival, 1,019% median ROI, and $2.41 per run.

Why it matters: HKR-H/K/R all pass: the cost gap is clickable, and the post gives a 30-day, 34-tool setup plus a 17× cost delta. Single-source Reddit benchmark with no cross-validation keeps it in the 78–84 band.

Xinzhiyuan · WeChat

OpenAI President Admits in Court He Got Up to $30B Equity for No Cash

Greg Brockman testified that he paid no cash for equity in OpenAI’s for-profit arm worth over $20B and near $30B. The hearing also covered Brockman and Sam Altman’s Cerebras stakes, a $10B OpenAI order, a $1B loan, and a later $20B order. The key issue is nonprofit asset conversion.

Why it matters: HKR-H/K/R all pass: the court disclosure gives concrete equity and supplier-conflict numbers tied to OpenAI governance. Single-source sourcing and sensational framing keep it at the low end of the 85 band.

OpenAI News

New Ways to Buy ChatGPT Ads

OpenAI expanded ChatGPT ad buying with a beta self-serve Ads Manager, CPC bidding, and enhanced measurement tools. The post says ads protect privacy and keep chats separate; it does not disclose pricing, rollout scope, or timing.

Why it matters: HKR-H/K/R all pass: OpenAI is turning ChatGPT ads into buyable tooling. Price, placement scope, and rollout timing are not disclosed, so this stays a mid-weight business product update.

r/LocalLLaMA

Benching Local Qwen as a Codex Validator, Co-agent, and Challenger

robert896r1 tested Qwen3.6 27B GGUF beside Codex as a coding validator and released a reproducible eval suite. The runs covered Bartowski, Unsloth, 65k/128k context, and q8/f16 KV cache; three 128k profiles tied for best, with no measured q8 KV accuracy loss in this suite. The useful signal is the sidecar eval: missed directives, overbuilding, UI judgment, and long-context misses, not a universal leaderboard.

Why it matters: HKR-H/K/R all pass: a reproducible sidecar eval with concrete Qwen/Codex conditions beats a normal Reddit tip. Source authority and event scale keep it in the 72–77 band, not a same-day must-write.

TechCrunch · AI

OpenAI’s cozy partner Cerebras is on track for a blockbuster IPO

Cerebras is moving toward an IPO at a valuation of $26.6 billion or more. The snippet says its OpenAI relationship is deep, but does not disclose ownership, revenue, or timing. The key signal is OpenAI-linked supply-chain valuation, not just AI chips.

Why it matters: HKR-H/K/R all pass: OpenAI partner, $26.6B valuation, and an IPO angle tied to AI compute supply. Lack of revenue, ownership, and timetable keeps it below must-write model-release territory.

Financial Times · Technology

OpenAI president defends motives in for-profit restructuring as he reveals $30bn stake

OpenAI’s president defended its for-profit restructuring and disclosed a $30bn stake. Elon Musk’s lawsuit says executives sold out the charity mission for personal gain. The post does not disclose the president’s name, equity structure, or restructuring terms.

Why it matters: All three HKR axes pass: OpenAI’s for-profit shift, a $30bn stake, and Musk’s lawsuit make it same-day material. Missing name, equity structure, and restructuring terms keep it below the 95+ band.

Bloomberg Technology

Musk’s Lawyer Pushes OpenAI’s Brockman to Give Back $29 Billion

Greg Brockman testified his OpenAI stake is worth almost $30 billion, and Musk’s lawyer asked why he had not donated most earnings to OpenAI’s nonprofit foundation. The post does not disclose stake size, valuation basis, or case procedure details.

Why it matters: HKR-H/K/R all pass: a $29B courtroom confrontation, a new valuation claim, and OpenAI governance tension. Missing stake percentage, valuation basis, and procedural detail keep it below must-write.

May 4Monday

TechCrunch · AI

Anthropic and OpenAI Are Both Launching Joint Ventures for Enterprise AI Services

Anthropic and OpenAI will each launch joint ventures for enterprise AI services. Both partnered with asset managers to market enterprise AI products more aggressively. The RSS snippet does not disclose partner names, equity terms, pricing, or launch dates.

Why it matters: HKR-H and HKR-R are strong because two frontier labs mirror the same enterprise JV move. HKR-K is limited to the sales-vehicle mechanism; names, equity, pricing, and launch timing are not disclosed.

Xinzhiyuan · WeChat

Top AI wrote dozens of pages of derivation before reviewers found the problem was wrong

Xinzhiyuan says Google DeepMind used Aletheia on 700 Erdős problems and got 13 original answers. The pipeline had Gemini Deep Think produce 200 candidates, then a verifier reduced them to 63. The post says Erdős-75 had a wrong premise, yet Aletheia wrote dozens of proof pages.

Why it matters: HKR-H/K/R all pass: the mistaken Erdős-75 setup gives a sharp hook, while the 700/13/200/63 pipeline adds substance. This is strong research coverage, not a GPT-scale product release, so it fits 78–84.

May 3Sunday

r/LocalLLaMA

LLM proxy that lets Claude Code talk to any model

DataNebula released open-source rosetta-llm, letting Claude Code call multiple providers through one gateway. It translates Anthropic Messages, OpenAI Chat, and OpenAI Responses, and round-trips encrypted reasoning via the signature field. The key detail is thinking-block fidelity for multi-turn agent prompt-cache hits.

Why it matters: HKR-H/K/R all pass, but this is a Reddit open-source tool post with no adoption, stars, or benchmark data disclosed. Score stays in the mid-weight tooling band, not 78+.

r/LocalLLaMA

Upskill: skill registry your agent consults before it starts, with 10k+ indexed skills

Autoloops released Upskill, an open-source skill registry with 10k+ indexed skills for agents. Search combines Postgres full-text search, 1024-dim embeddings, and reranking by stars, installs, and feedback. LLM adversarial review blocked hundreds of skills at index time.

Why it matters: HKR-H/K/R pass: a useful open-source agent registry with concrete retrieval and safety mechanics. Source authority is low and adoption is unproven, so it stays in the 72–77 featured band.

Synced · WeChat

Why CTOs at Billion-Dollar Companies Are Joining Anthropic as Engineers

Jiqizhixin lists at least six CTOs who joined Anthropic as individual contributors. Cases include Workday, You.com, Box, Super.com, and Adept AI from Jan 2025 to Apr 2026. The key issue is career leverage, not just AGI mission talk.

Why it matters: HKR-H/K/R all pass: the career-status reversal is clickable, the post gives 6 cases, and it touches AI talent competition. No hard exclusion, but it is commentary, not a model or product release.

Hacker News front page

OpenAI's o1 correctly diagnosed 67% of ER patients vs. 50–55% by triage doctors

OpenAI o1 correctly diagnosed 67% of ER triage patients, versus 50–55% for doctors. The title cites a Harvard trial, but the RSS post does not disclose sample size, case mix, or evaluation protocol. Practitioners should track the test setup, not only the accuracy gap.

Why it matters: HKR-H/K/R all pass: a high-risk ER comparison gives the hook, 67% vs 50–55% gives a testable number, and clinical trust/safety creates resonance. Missing sample size and protocol keep it in 78–84, not P1.

May 2Saturday

r/LocalLLaMA

A Dark-Money Campaign Is Paying Influencers to Frame Chinese AI as a Threat

Build American AI is funding an influencer campaign to spread pro-AI messaging and fear about China. The post links it to a super PAC backed by OpenAI and Andreessen Horowitz executives; amounts, influencer names, and targeting mechanics are not disclosed.

Why it matters: HKR-H/K/R all pass: the post names an organization, funding link, and campaign condition; amounts and influencer lists are missing. This fits the featured edge, not a model-release-level event.