Skip to content

OpenAI / ChatGPT

Everything OpenAI: the GPT models, ChatGPT and Sora, company strategy and people moves.

Latest picks

441–460 of 1,549

Aug 15Saturday

Financial Times · Technology

OpenAI upheaval mounts as Sam Altman readies IPO push

OpenAI is facing a wave of senior departures as Sam Altman pushes toward an IPO, the FT reports. Chief Strategy Officer Jason Kwon and Chief Product Officer Kevin Weil have recently left, adding to earlier exits of key early members. Altman is still driving the conversion to a for-profit structure, targeting a valuation above $150 billion. The article does not disclose the IPO timeline or underwriters. A string of C-suite exits is real pressure on pricing and investor confidence, but the final story will hinge on revenue growth and retention numbers.

Why it matters: OpenAI's executive turmoil continues, this time losing its chief strategy officer and chief product officer right at the pivot to a for-profit structure and IPO push. FT exclusively reports Altman's internal valuation target above $150B. The personnel shock plus IPO narrative ...

Hacker News front page

The End of Mathematics: When AI Overproduction Shrinks the Math Community

Daniel Litt gave a talk at OpenAI imagining a future where AI is superhuman at math but progress stalls. He shows arXiv combinatorics submissions spiking while MathOverflow Q&A volume drops sharply since early 2025. Multiple groups and models are duplicating the same results—three teams independently proved Feige's 1/e conjecture almost simultaneously. By 2027, the dominant career strategy could be letting codex pick conjectures, prove them, and write papers, producing several per day that nobody reads. Colleagues already refuse to discuss work in progress for fear of being scooped by AI. The post does not spell out the full 2028 scenario.

Why it matters: Daniel Litt is a credible algebraic geometer, not a random blogger. He uses the divergence between arXiv submission volume and MathOverflow activity to argue AI is turning math research into isolated production — a sharp take backed by data. Score held back because it's still ...

Aug 14Friday

Hacker News front page

When Genius Fails: AI Labs' Intellectual Arrogance, from a $20B Blow-Up to Materials Science

Leopold Aschenbrenner's $20B hedge fund Situational Awareness blew up this week, with its portfolio sold to Citadel. Aschenbrenner, formerly on OpenAI's Superalignment team, gained fame from a 2024 essay on AGI's imminence, then raised a fund and went heavily long AI stocks (neoclouds, memory, datacenter power) with ~4x leverage while shorting software names—both sides moved against him. Author James Wang, an ex-hedge fund analyst with an AI background, compares it to Long-Term Capital Management's 1998 collapse: very smart people assuming expertise transfers across domains. He extends this critique to AI lab culture, citing DeepMind's materials science work flagged for basic chemistry errors by domain experts, and a Hugging Face engineer publicly mocking Cerebras' wafer-scale chip design without understanding the hardware. The core argument: being an expert in one field doesn't make you an expert in all fields, but frontier AI culture often conflates confidence with competence.

Why it matters: Leopold Aschenbrenner's $20B hedge fund blew up after betting long AI infra and short software — both sides went wrong. The author has analyst background and provides concrete numbers, not just hot takes. It's a finance story rather than an AI tech update, but as a character p...

Hacker News front page

DeepSeek V4 Pro goes GA with peak/off-peak API pricing

DeepSeek V4 Pro is now GA, with major agent workflow gains and adjustable reasoning effort—low for simple tasks, high for daily agent work, max for complex ones. It natively supports the OpenAI Responses API and one-click Codex setup. API pricing shifts to peak/off-peak on Aug 16: off-peak is 50% cheaper. Model names stay the same; try it via Expert Mode on the app.

Why it matters: V4 Pro GA with agent hardening and a thinking-effort dial is a real feature update that matters to developers building automation on DeepSeek. Held below 85 because the post doesn't disclose GA benchmark comparisons or the actual peak/off-peak price spread — the info density i...

Financial Times · Technology

OpenAI and Anthropic in price war as Chinese AI rivals gain ground

FT reports OpenAI and Anthropic are slashing prices to win enterprise customers, pressured by cost-competitive Chinese models like DeepSeek. Both are pushing cheaper, smaller models while leaning on premium subscriptions and IPO expectations to support valuations. The post doesn't spell out exact price cuts or effective dates—it's more a trend piece.

Why it matters: FT's trend piece has narrative value, but the body lacks specific price-cut figures or timelines — the information density isn't hard enough. H and R hit, K is missing; it just clears the featured threshold at 72.

Computing Life · Share · Yage

Agent auth isn't about crypto—it's about who holds the trust

Vercel Connect removes long-lived refresh tokens from app code and swaps them via OIDC—its real value is centralized management, not stronger security. Cloudflare OS tracks where data goes after access, not token issuance. OpenAI's Ona acquisition bets on customer-held credentials. Four competing beliefs each cover one segment; credential and data pipelines remain unintegrated.

Why it matters: A sharp industry analysis that dissects four approaches to agent auth and their shared blind spot, with concrete mechanism comparisons rather than vague trend talk. Points off because it's a synthesis piece rather than a primary scoop, and the post doesn't offer a clear path t...

Bloomberg Technology

OpenAI revenue run rate tops $40 billion ahead of IPO

OpenAI's annualized revenue run rate has passed $40 billion, more than doubling in three months. This is the key number ahead of its IPO, showing how fast it's commercializing. ChatGPT subscriptions and API usage are the main drivers, though the article doesn't break out each segment's share. I'd discount this a bit—a run rate extrapolates one month, not actual full-year cash—but the growth is real.

Why it matters: The most critical financial figure ahead of OpenAI's IPO drops: $40B run rate, doubled in three months, directly tied to how the market prices its valuation. Bloomberg exclusive sourcing is a plus, but the article doesn't break down ChatGPT subscriptions vs. API revenue or pro...

The Verge · AI

OpenAI loses second executive this week as CRO Denise Dresser departs

Denise Dresser joined OpenAI as CRO in December after serving as Slack CEO. She now says she'll leave in the coming weeks to pursue other opportunities. Wiz president and COO Dali Rajic will take over the CRO role. Earlier this week, special projects lead and former COO Brad Lightcap also announced his departure. The post doesn't spell out reasons for either exit or any compensation/ non-compete details.

Why it matters: Two executive exits in one week, with Dresser leaving less than a year after joining from Slack and a successor coming from Wiz rather than a traditional SaaS giant — strong signal. Deduction: the post doesn't disclose Dresser's reason for leaving or Rajic's start date, so we ...

TechCrunch · AI

OpenAI adds 'Ultrafast' mode to GPT-5.6 Sol, pushing inference to 14x speed

OpenAI launched a preview of Ultrafast, a mode for its top model GPT-5.6 Sol that hits 750 tokens/sec — 14x the standard speed. The company says this avoids the old trade-off of switching to a smaller model for real-time use. It runs on Cerebras chips and is limited to a small customer group for now, with broader access planned. Anthropic's Claude has a fast mode but doesn't match this throughput. Target workflows include incident response, customer support, financial analysis, and e-commerce. The post doesn't disclose pricing, latency details, or a general release date.

Why it matters: 14x speed on GPT-5.6 Sol is a real inference win with direct agent implications. Held back from higher score because it's a limited preview with no GA timeline and ties to Cerebras hardware — general availability is unproven.

Hacker News front page

Cerebras powers OpenAI's GPT-5.6 Sol Ultrafast at 750 tokens per second

Cerebras and OpenAI previewed Ultrafast mode, running GPT-5.6 Sol on Cerebras' wafer-scale chips at up to 750 output tokens per second. On Humanity's Last Exam, it answered all 2,500 PhD-level questions in 11 hours 11 minutes—nearly 7× faster than Claude Fable 5 with comparable accuracy. On GDP-Val it delivered a 5.6× end-to-end speedup with no quality loss. The speed comes from packing 44 GB of SRAM on a single wafer, keeping model weights on-chip to avoid memory bandwidth bottlenecks. Access is limited preview for now.

Why it matters: OpenAI and Cerebras jointly unveiled Ultrafast mode for GPT-5.6 Sol, hitting 750 tok/s — a speed that pulls frontier models into real-time interaction territory. The 11-hour HLE run across 2,500 questions gives deployment teams a concrete number to work with. Not a perfect sco...

TechCrunch · AI

OpenAI replaces CRO after 9 months, hires Wiz president Dali Rajic

OpenAI replaced CRO Denise Dresser after only nine months, bringing in Wiz president and COO Dali Rajic. The move follows COO Brad Lightcap's departure and No. 2 exec Fidji Simo stepping down. President Greg Brockman said OpenAI now reaches 1B+ weekly active users and 2M businesses, and Rajic will turn lessons learned into repeatable sales execution. Rajic's former company Wiz was acquired by Google for $32B earlier this year.

Why it matters: Consecutive executive moves at OpenAI are newsworthy, and the Wiz connection adds texture. But pure personnel news without product/tech substance keeps it at the lower edge of featured.

Aug 13Thursday

AI HOT (Curated Pool)

OpenAI's GPT-5.6 builder guide shows how to run frontier agents at a fraction of the cost

OpenAI published a builder's guide for GPT-5.6, showing how startups use cheaper models like Luna and Terra for agent workloads. Hex dropped GPT-5.6 into their harness and got best results at low reasoning effort—the model didn't chase bad leads and used fewer tokens. Hypha kept 98% of GPT-5.5's extraction accuracy at 1/18 the cost. Browser Use ran 106 hard browser tasks: Luna hit 78% for $14, while the current SOTA model reached 80% for $235. On BrowseComp, GPT-5.6 Luna (Extra High) scored 84.04% at $1.33; three months ago GPT-5.5 (Extra High) scored 84.36% at $33.27. The guide also details three new API primitives: persisting reasoning across turns, native multi-agent orchestration, and programmatic tool calling for deterministic work. The post does not disclose release dates or regional availability.

Why it matters: An official builder's guide from OpenAI with real startup case studies and concrete cost/performance tradeoffs — useful for developers. But it's a product best-practices doc, not a model launch or research breakthrough, so importance caps at recommended-reading level.

OpenAI News

OpenAI previews Ultrafast mode: GPT-5.6 Sol at up to 14X the speed

OpenAI added an Ultrafast inference tier for GPT-5.6 Sol, running on Cerebras chips at up to 750 output tokens per second—14× faster than standard. The preview launches via the API first, targeting latency-sensitive workflows like incident response, financial research, and real-time customer support. OpenAI’s own teams are using it for on-call debugging and to tighten overnight research loops into same-day iterations. The post does not disclose pricing or a general release date; access is by application only.

Why it matters: OpenAI's Ultrafast preview pushes GPT-5.6 Sol to 14X standard speed via Cerebras silicon, with three concrete latency-sensitive use cases. No pricing or GA date disclosed, capping the score at 82 rather than pushing into the must-write-same-day band.

OpenAI News

OpenAI appoints Dali Rajic as Chief Revenue Officer

OpenAI hired former Wiz President Dali Rajic as CRO, replacing outgoing Denise Dresser. His brief: turn early enterprise wins into repeatable, metrics-driven revenue execution. OpenAI also disclosed 1B+ weekly active users and 2M+ business customers—double the figure from a year ago. Worth discounting: the user number includes free ChatGPT users, not just paying accounts. The post doesn't disclose Rajic's start date or compensation.

Why it matters: Official OpenAI announcement with both a personnel change and business metrics—enough density for featured tier. But it's fundamentally an executive hire with no product or tech angle; HKR hits H and K only, missing R, landing in the 72-77 band per policy.

AI Chat-Group Daily (群聊日报)

Closed-source reasoning chains extracted at scale; Coze CLI hijacks AI tools

The big one today: researchers extracted hidden reasoning chains from Anthropic, OpenAI, and Google models at scale. The trick is absurdly simple—take Opus 4.8's encrypted CoT and feed it to Haiku 4.5, which decodes it verbatim. All three API families were broken, and decoding 10K trajectories costs about $720. A separate paper shows you can even reverse-engineer reasoning from public outputs alone using a 1.5B-param model. Separately, Coze CLI was caught silently scanning local Codex and Claude Code directories and injecting its own skills into workflows. On the engineering side, the group discussed how prompt debt now rivals traditional code debt—old rules pile up, evals lag behind model iterations, and nobody dares delete anything.

Why it matters: Strong cross-source cluster signal (chat digest + original paper + study notes). First systematic validation that encrypted CoT from three major vendors is cross-model decodable, with concrete $720/10k cost. All three HKR axes hit, but the source is a secondary digest rather t...

Hacker News front page

OpenAI brings Codex coding agent into ChatGPT desktop, now with a Linux download

OpenAI launched Codex in ChatGPT, embedding its coding agent directly into the ChatGPT desktop app and releasing a Linux download. Codex handles end-to-end engineering tasks—feature builds, refactors, migrations—with multi-agent parallelism across projects and a Skills system that lets teams teach it their standards. The post doesn't disclose pricing or the underlying model name. Customer quotes from Ramp, Duolingo, and Harvey claim it catches bugs humans miss in PR review and cuts early iteration time by 30–50%. I'd discount those numbers a bit—they're vendor-supplied testimonials with no independent benchmark.

Why it matters: OpenAI folding Codex into ChatGPT Desktop is a distribution play against Cursor and Claude Code. Parallel multi-agent worktrees and the Skills system are real mechanisms, not fluff. Docked slightly because the post doesn't disclose pricing or the underlying model name, and Lin...

Computing Life · Share · Yage

Every coding agent form factor shift is chasing the same thing: execution data

DeepSeek is hiring an Agent Harness PM, signaling it's filling the gap of not having its own coding tool runtime. The article argues that desktop apps, managed cloud agents, and remote control are all moves to capture execution data. Interfaces converge because they're cheap to copy; execution layers diverge because that's where the data moat is. Without a first-party harness, DeepSeek lacks real-world coding feedback to improve its models. Judge a coding agent by who controls the execution environment, who sees the data, and who's in the data flywheel—not by feature checklists.

Why it matters: A sharp industry analysis that uses DeepSeek's hiring move and LangChain test data to argue 'harness = data moat.' Hits all three HKR axes, but as an opinion piece rather than a primary release, scored at the lower end of the 78-84 band per policy.

Computing Life · Share · Yage

Anthropic spent 31M output tokens on Riemann zeta search—the real signal is the architecture

Anthropic used an unreleased Claude to raise the proven lower bound of Riemann zeta zeros on the critical line from 41.6% to 67.2%—still far from proving the Riemann Hypothesis. The real story is the search architecture: two Claude Code sessions burned 31M output tokens. Round one produced 650 ideas, all failed, but left a ledger of 106 partial survivors with kill criteria. Round two coordinated ~60 subagents that rechecked the ledger and stitched a final route via stepping-stone transfers. The key insight was switching from requiring all-positive structure to counting usable positive directions. Lean formalization is sorry-free, but effective forms are missing from headline statements and independent third-party review is absent. Compared with GPT-5 on Erdős and OpenAI's ten math advances, the pattern is clear: generation and formalization are accelerating fast, while human understanding and absorption stay flat. The bottleneck is shifting from discovery to comprehension.

Why it matters: Anthropic used an unreleased model for math search — the 31M-token engineering details and hostile review mechanism are real signal, not pure PR. Docked slightly because pure math is far from product impact, and the post doesn't disclose the model name or token cost.

Aug 12Wednesday

Hacker News front page

AI is removing the middle class of software engineering

The author contrasts a 2020 vacation mess with a 2026 Monday morning: 7 PRs, one at 24,506 lines. AI removed the speed limit on bad decisions. Anyone can prompt an agent and ship something that looks functional, but no one knows where the data comes from or why Kafka was added. Reverting one bad call is far harder than generating it, and five more land while you fix it. The bet: AI widens the salary gap—good decision-makers become more valuable, while engineers who only implement become too expensive to hire.

Why it matters: A grounded, first-person engineering observation with concrete scenes and numbers, not generic 'AI will replace devs' fluff. Hits all three HKR axes, but it's a personal blog commentary, not a product launch or research breakthrough, so it lands in the 78-84 band. No cross-sou...

AI HOT (Curated Pool)

Nathan Lambert wrote an AI textbook—models still can't handle long-form nonfiction

Nathan Lambert just finished his post-training textbook *Reinforcement Learning from Human Feedback*. He used LLMs for LaTeX formatting, copyediting, and diagrams, but when he tried to get a model to write a full technical chapter, the output was confusing, poorly organized, and made random conceptual errors. He argues long-form nonfiction writing has stagnated even as models became superhuman at coding and math. The post doesn't cite benchmark scores, but Lambert points to a lack of good training data and notes inference-time scaling hasn't helped writing. His takeaway: if models can't coherently organize established knowledge, autonomous scientific breakthroughs are still far off.

Why it matters: Lambert's first-person experiment delivers concrete failure cases and a data-gap diagnosis — all three HKR axes hit. Deduction: no quantitative benchmark, it's personal experience not systematic research, and the second half drifts into general capability discussion. Sits righ...