Skip to content

OpenAI / ChatGPT

Everything OpenAI: the GPT models, ChatGPT and Sora, company strategy and people moves.

Latest picks

701–720 of 1,549

Jul 10Friday

The Verge · AI

OpenAI's ChatGPT browser Atlas is shutting down less than a year after launch

OpenAI launched the Atlas browser in October 2025, pitching it as a way to browse, fill forms, and book tickets with ChatGPT. The company now says it will shut down on August 31, 2026 — less than a year after launch. The post does not disclose the reason or any user numbers. A browser killed this fast usually means low adoption or a strategic pivot.

Why it matters: OpenAI's Atlas browser went from launch to shutdown in under a year — the contrast alone makes it worth a click. Missing the why and user numbers means no knowledge bump. Score sits at the featured threshold because this is a public product retreat from OpenAI that builders sh...

TechCrunch · AI

New York Times says OpenAI hid evidence in ChatGPT copyright trial

The New York Times and The Daily News filed a sanctions motion, accusing OpenAI of lying about its inability to search training data and chat logs for copyrighted material. OpenAI had claimed such searches were technically infeasible, but the publishers say internal tools and datasets exist that can do exactly that. The court has not ruled yet; OpenAI's response is not in the article.

Why it matters: NYT filed a sanctions motion alleging OpenAI hid internal tools capable of searching training data, directly undercutting OpenAI's core defense that such searches were technically infeasible. All three HKR axes hit: strong conflict, concrete evidence, high industry resonance. ...

TechCrunch · AI

How did the US government decide OpenAI's frontier model Sol was safe to release?

OpenAI is rolling out Sol, a frontier model on par with Anthropic's Fable, which the White House briefly banned. Mina Narayanan of Georgetown's CSET says she has no visibility into the government's review process. Anthropic mentioned building a jailbreak classifier and defense-in-depth, but the actual dialogue between the government and the labs remains opaque.

Why it matters: Policy transparency is a core AI governance issue, and the CSET researcher's admission of no access gives this a concrete hook. The deduction is that the article raises the question without revealing the actual review mechanism — no internal process details — so it lands at 78...

AI HOT (Curated Pool)

OpenAI launches ChatGPT Work desktop app, integrating Codex and GPT-5.6

OpenAI combined Codex and ChatGPT into a single desktop app called ChatGPT Work. Powered by Codex and GPT-5.6, it can work across apps and files, running complex projects for hours. It also includes new coding workflows, a Chrome extension, an improved built-in browser, and faster Computer Use driven by GPT-5.6. The post doesn't disclose launch date, pricing, or system requirements.

Why it matters: OpenAI ships a desktop agent bundling Codex and GPT-5.6, directly competing with Cursor and Claude Code. Concrete product shape and technical details make this a same-day must-write. No launch date or pricing disclosed, slight deduction but still featured.

The Verge · AI

OpenAI releases GPT-5.6 and announces ChatGPT Work for enterprises

OpenAI launched GPT-5.6 after receiving government approval, ending a months-long limited preview. They also announced ChatGPT Work, an enterprise-focused version, though the post doesn't detail its features, pricing, or launch date. Specific benchmarks or performance gains for GPT-5.6 aren't covered either — only the release and product names are confirmed.

Why it matters: Flagship OpenAI model moves from restricted preview to full launch, plus an enterprise teaser — H and R are solid. But the post gives zero performance data or feature details, so K is absent, capping the score. The Verge's sourcing authority helps, but the information density ...

Financial Times · Technology

Microsoft’s early AI lead has become a test of faith

The FT argues Microsoft’s Copilot strategy, built on its OpenAI tie-up, hasn’t yet translated into clear revenue gains. Azure growth is slowing, enterprise willingness to pay for Copilot remains uncertain, and in-house model efforts lag. The AI premium the market gave Microsoft now hinges on hard financial delivery.

Why it matters: FT's commentary nails the core tension in Microsoft's AI narrative: the valuation premium from the OpenAI tie-up is still priced in, but Azure growth and Copilot conversion haven't delivered yet. Hits all three HKR axes, but it's analysis not primary data — lands at the featur...

Jul 9Thursday

TechCrunch · AI

Anthropic, OpenAI, and SpaceX are bigger than the last 25 years of tech exits

A new Pitchbook report estimates that SpaceX, Anthropic, and OpenAI together will generate more exit value than all U.S. VC-backed exits since 2000. SpaceX already went public at $1.77 trillion; Anthropic and OpenAI are each pushing toward trillion-dollar valuations. The post doesn't give a precise combined figure, but the concentration in AI and space is historic.

Why it matters: Pitchbook's data gives the AI valuation debate a historical yardstick—the comparison scale is massive and the numbers are concrete. Two dings: SpaceX isn't an AI company, so lumping it in feels like padding; the post doesn't give a combined exit-value figure, just directional ...

Ben's Bites

SpaceXAI and Cursor trained Grok 4.5, a model 6x cheaper than Opus

SpaceXAI and Cursor jointly trained Grok 4.5, landing between Opus 4.7 and 4.8 in performance but 6x cheaper than Opus and 3x cheaper than GPT-5.5 on a per-token basis. OpenAI rolled out GPT-5.6 (Sol, Terra, Luna) to all users; early testers say Sol is less smart than Fable but far more reliable. ChatGPT Voice got new GPT-Live-1 and Live-1-mini models that can talk while you speak and use GPT-5.5 in the background. Anthropic extended Fable 5 access for Claude subscribers to July 12—the post doesn't explain the repeated delays. Meta introduced Muse Image and Muse Video; image editing and text rendering look solid, but images still have an AI look, and the video model is in preview.

Why it matters: SpaceXAI + Cursor joint Grok 4.5 launch with concrete performance anchor and pricing — all three HKR axes hit. Deduction because source is a newsletter summary, not a first-party announcement, and the body is truncated with incomplete GPT-5.6 info. +3 cross-source bump to 82, ...

OpenAI News

OpenAI turns its bio bug bounty into an ongoing program, doubling rewards to $50K starting with GPT-5.6

OpenAI is turning its GPT-5.5 Bio Bug Bounty into an ongoing private program, now called the OpenAI Bio Bounty Program. The focus stays on universal jailbreaks that beat its biosafety challenges. Rewards jump from $25,000 to $50,000 for both GPT-5.6 and GPT-5.5, with smaller payouts possible for partial wins. GPT-5.5 testing ends July 27, 2026; after that only GPT-5.6 is in scope. Applicants need a ChatGPT account, must sign an NDA, and past GPT-5.5 applicants don't need to reapply.

Why it matters: OpenAI upgraded its bio-safety bounty from a one-off to a permanent program with doubled rewards and clearer rules — a substantive safety-mechanism update. But the audience fit is narrow: bio-jailbreak testing is far from most practitioners' daily work, so resonance is weak, k...

AI HOT (Curated Pool)

OpenAI launches GPT-5.6 family: Sol, Terra, Luna, pushing performance per dollar

OpenAI released the GPT-5.6 family on July 9: flagship Sol, balanced Terra, and low-cost Luna. Sol scores 53.6 on Agents' Last Exam, beating Claude Fable 5 by 13.1 points at roughly one-quarter the estimated cost. A new `ultra` mode coordinates parallel agents to cut latency and lift scores on BrowseComp and Terminal-Bench 2.1. Sol also tops the Artificial Analysis Coding Agent Index at 80, using less than half the output tokens of Fable 5. Terra and Luna outperform Fable 5 at about one-sixteenth the cost. OpenAI ran extensive red-teaming and automated testing, and hardened safeguards with external partners during a preview period.

Why it matters: OpenAI's flagship model refresh with three variants, a direct benchmark win over Claude Fable 5 on long-horizon agent tasks, and a claimed 4x cost advantage. This is the most significant model launch of 2026 so far and will immediately reshape agent workflow decisions.

AI HOT (Curated Pool)

OpenAI launches ChatGPT Work, an agent that acts across apps and stays with projects for hours

ChatGPT Work is an agent that acts across apps and files, powered by the new GPT‑5.6 model. It breaks complex projects into steps, creates slides, sheets, docs, and web apps, and can run scheduled tasks while you're away. Nearly all teams inside OpenAI use it; early external users include Zapier, RingCentral, Virgin Atlantic, and NVIDIA. The post does not disclose pricing details, only that it's available starting today.

Why it matters: Official OpenAI launch of ChatGPT Work alongside GPT-5.6 — a major product release. Cross-app autonomous operation, background execution, and human-in-the-loop approval provide concrete detail beyond marketing. Not a 95 because we only have the official blog post so far; third...

Hacker News front page

AI buildout isn't bottlenecked by electricity supply—it's the grid interconnection queue

The U.S. has enough electricity, but connecting a new data center to the grid now takes a median 55 months, up from under 20 months in 2005. The $40B+ Stargate campus in Texas will draw 1.2 GW at peak—equal to 313,000 homes. Jensen Huang, Mark Zuckerberg, and Sam Altman have each said energy access is the real limiter. The bottleneck is a first-come-first-served queue and rigid rules that don't reward plants willing to cover their own short-term power needs.

Why it matters: Strong angle correcting the AI bottleneck narrative from 'power shortage' to 'grid interconnection queue rigidity,' with concrete numbers and the Stargate case. But the source is Works in Progress rather than a tier-1 tech outlet, and the piece is policy/infra analysis rather ...

AI HOT (Curated Pool)

GPT-5.6 is now the preferred model in Microsoft 365 Copilot

OpenAI announced GPT-5.6 as the new preferred model for Microsoft 365 Copilot, covering Word, Excel, PowerPoint, Chat, and Cowork. The model promises more useful work per token: fewer prompt rounds in Word, more token-efficient analysis in Excel, and faster idea-to-slide conversion in PowerPoint. Both Microsoft's Copilot president Nitin Agrawal and OpenAI's API product head Nikunj Handa endorsed the move. The post does not disclose a rollout date, pricing changes, or benchmark comparisons—treat this as a partnership announcement rather than a product review for now.

Why it matters: GPT-5.6 becomes the default model across Microsoft 365 Copilot — Word, Excel, PowerPoint, Chat, and Cowork. First major enterprise deployment after the model's release. Hits all three HKR axes: concrete efficiency claims, real deployment context, strong audience resonance. Hel...

Computing Life · Share · Yage

GPT-5.5 reasoning tokens cluster at 516, causing wrong answers on coding tasks

Developer vguptaa45 audited 390K Codex responses and found GPT-5.5 reasoning cuts off at exactly 516 tokens in 44% of cases, versus 19.8% for GPT-5.4 and 0.34% for GPT-5.2. Truncated runs all produced wrong answers; the same tasks completed with 6,000–8,000 tokens all got correct. The community reproduced it and found adding 'THIS IS HARD' to the prompt bypasses the cutoff, pointing to a budget-classification bug rather than a model capability drop. In the same week, Liquid AI released Antidoom to fix the opposite failure—reasoning models stuck in self-revising doom loops. Both failures live in the reasoning layer, invisible to standard pass-rate evals. The post recommends monitoring reasoning token distributions and not assuming newer models are more stable.

Why it matters: A community audit of 390k Codex responses shows GPT-5.5's reasoning clips at exactly 516 tokens in 44% of coding tasks, all wrong, while full runs get it right. Solid data, reproduced, with a workaround — directly useful signal for AI coders. Not scored higher because it's a s...

Computing Life · Share · Yage

GPT-Live separates voice interaction from heavy reasoning—that's the real shift

OpenAI launched GPT-Live on July 8, adding full-duplex and a delegation architecture to ChatGPT voice. Full-duplex lets you interrupt and talk while it works, but the principle isn't new—Moshi, Gemini Live, and ByteDance's Seeduplex all did it. The real change is delegation: the voice layer handles conversation while GPT-5.5 runs search, reasoning, and computation in parallel in the background, returning results as they arrive. This breaks the latency paradox where faster meant dumber. The voice model itself is limited—the System Card confirms it has no standalone tool access or code execution. No API yet; developers can only sign up for a waitlist. Realtime API remains the production workhorse at $64/M tokens for audio output. The post doesn't spell out whether custom tools can be plugged into the delegation layer or how much control developers will get over the black box.

Why it matters: OpenAI just shipped GPT-Live, and this analysis doesn't stop at full-duplex—it pulls out the delegation architecture as the real novelty. The author has technical judgment, laying out comparisons with Moshi, Gemini Live, and ByteDance's Seeduplex clearly. Points off because th...

AI HOT (Curated Pool)

Anthropic files confidential IPO, Q3 profit projected above $1B

SemiAnalysis reports Anthropic's Q3 profit will exceed $1B and it confidentially filed for IPO on June 1. Claude Code's rapid developer adoption made it the B2B leader ahead of OpenAI. Combined ARR of the two firms is nearing $100B, while OpenAI pushed its IPO to 2027. The report floats a $6T market cap target, though the article doesn't show the math behind it.

Why it matters: Anthropic's confidential IPO filing with hard profit and ARR numbers, plus a concrete B2B story driven by Claude Code. SemiAnalysis is a credible source, but the post doesn't disclose S-1 details, so the score stays below 95.

TechCrunch · AI

OpenAI launches GPT-Live-1 voice models that can speak and listen at the same time

OpenAI released GPT-Live-1 and GPT-Live-1 mini, full-duplex voice models that let users interrupt naturally and enable live translation. Paid ChatGPT users get mini by default; higher tiers can access the larger GPT-Live-1. The new models skip the old speech-to-text-to-speech pipeline and can call GPT-5.5 for search, reasoning, or agent tasks mid-conversation. Demos showed the model staying silent until summoned and presenting info visually. The post doesn't disclose pricing changes or a rollout timeline.

Why it matters: OpenAI drops two end-to-end voice models with full-duplex as the headline feature. TechCrunch has the scoop — it's a significant product update. Not scoring 85+ because the post doesn't disclose latency numbers, supported languages, or real-world comparisons; right now we only...

Jul 8Wednesday

AI HOT (Curated Pool)

OpenAI publishes national security principles for government partnerships

OpenAI released a set of national security principles today, laying out how it will work with governments and law enforcement. Three hard lines: no mass domestic surveillance, no directing autonomous weapons, no high-stakes automated decisions. Over the past month it has set up cyber defense partnerships with Australia, Canada, Japan, South Korea, France, Germany, Poland, the Netherlands, and EU bodies, and is giving select U.S. and allied partners trusted access to its GPT-Rosalind model for biosecurity. The principles cover current and future work, including the Department of War contract, but OpenAI argues the biggest decisions should come through democratic processes, not from companies alone.

Why it matters: OpenAI's first systematic disclosure of national security principles, with three concrete bans and named partner countries. GPT-Rosalind mention adds substance. Not scored higher because it's a principles document, not a product launch — execution remains to be seen.

OpenAI News

OpenAI audits SWE-Bench Pro, finds ~30% of tasks are broken

OpenAI audited SWE-Bench Pro and estimates ~30% of its tasks are broken. An automated pipeline flagged 286 suspicious tasks; Codex-based investigator agents and five experienced engineers then reviewed them. Engineers identified 249 (34.1%) flawed tasks, mostly due to overly strict tests, underspecified prompts, low-coverage tests, and misleading prompts. OpenAI advises model developers to scrutinize results rather than trust leaderboard scores. The post does not disclose a fix timeline or a revised dataset release.

Why it matters: OpenAI audited SWE-Bench Pro and found ~34% of tasks defective — a ratio that forces the industry to re-examine coding benchmark reliability. The post provides concrete defect categories and a human review pipeline. Not scored higher because this is a benchmark quality report,...

AI HOT (Curated Pool)

British Columbia plans to sue OpenAI for not reporting shooter's violent ChatGPT chats

British Columbia announced on July 7 it is preparing to sue OpenAI. The shooter, 18-year-old Jesse van Ruijsselaar, had entered violent prompts into ChatGPT before OpenAI banned her account in June 2025. OpenAI did not alert law enforcement. In February 2026, she killed eight people at home and at a school in Tumbler Ridge before taking her own life. BC's Attorney General Sharma said the province will hold OpenAI and its management accountable; any damages won would fund community rebuilding and a new school. CEO Sam Altman apologized publicly in April, acknowledging that under post-June 2025 safety rules the account should have been reported. Lawyers for victims' families allege OpenAI withheld information because reporting one account would mean reporting thousands of similar ones.

Why it matters: A Canadian province sues OpenAI over a school shooting, alleging failure to report the shooter's violent ChatGPT chats. Clear timeline and casualty figures, strong legal precedent potential, directly hits the core debate on AI safety liability. Score not higher because only th...