Skip to content

OpenAI / ChatGPT

Everything OpenAI: the GPT models, ChatGPT and Sora, company strategy and people moves.

Latest picks

661–680 of 1,549

Jul 14Tuesday

Ben's Bites

OpenAI ships GPT-5.6 with three models, five thinking levels, and an Ultra sub-agent mode

GPT-5.6 ships as Luna, Terra, and Sol, each with five thinking levels (light to max) plus an Ultra mode that spins up sub-agents aggressively. The macOS ChatGPT and Codex apps merge into ChatGPT Work; a new ChatGPT Sites plugin builds hosted pages with optional ChatGPT login. Sol excels at UI and writing, especially with references; Terra feels like a steerable 5.5 upgrade; Luna has a mini-model vibe—fuzzy on ambiguous prompts but solid on clear tasks. Higher thinking levels burn usage fast, and OpenAI temporarily removed the 5-hour cap while fixing merge bugs, so weekly limits can vanish in one session. Also: Claude Code gets an in-app browser and multiplayer Artifacts, Meta launches multimodal Muse Spark 1.1 via API, and Apple sues OpenAI over alleged trade-secret theft for AI hardware.

Why it matters: GPT-5.6 going GA is one of the week's biggest product stories, and the three-model lineup with Ultra mode is worth practitioner attention. Docked because this is a tutorial recap rather than the primary release post, and the body is truncated with key details missing.

Hacker News front page

OpenAI's ad revenue may miss its own forecast by 90%

OpenAI started its ad trial in February and projected $2.5B in ad revenue this year and $100B by 2030. Emarketer estimates the entire US standalone chatbot ad market—ChatGPT, Copilot, Google AI Mode, Alexa for Shopping—will generate under $1B this year and just $5.41B by 2030. OpenAI's forecast assumes it captures search budgets en masse, dominates a fully mature chatbot ad market, and outperforms every ad format in history, all at once. None of those conditions hold today.

Why it matters: OpenAI's ad revenue forecast gets a reality check from third-party data, and the gap is big enough to matter. Emarketer caps the entire US chatbot ad market at $5.41B by 2030; OpenAI's own target is $100B. Concrete numbers, clear contrast, and it hits a live nerve about AI com...

Latent Space

OpenAI Codex hits 7M users, 10x growth in 6 months, likely overtaking Claude Code

OpenAI Codex reached 7M active users on July 13, adding 1M in a single day. That's 10x growth from ~550-700k at the start of 2026 and 2M in March. Anthropic last reported ~2M Claude Code users in February and has been silent since. The post speculates Anthropic shifted focus to Claude Tag, making direct comparisons harder. I'd note the spike coincides with the GPT 5.6 launch and a temporary removal of the 5-hour usage cap — retention remains unproven.

Why it matters: Codex hitting 7M users with 10x growth in 6 months is a real number worth surfacing, and Claude Code's silence since February creates a genuine information gap. The deduction is because this is a paid newsletter digest, not a primary source, and the headline's question mark si...

Computing Life · Share · Yage

What Would a ChatGPT That Doesn't Wait for Your Questions Look Like?

Greg Brockman described a 'no-product' AI in a July 1 interview: a system that runs in the background, spots conflicts, and drafts actions before you ask. The article walks through a thought experiment of this 'ambient intent layer' and argues the bottleneck for proactive AI isn't execution—it's the lack of long-term, self-updating personal context infrastructure.

Why it matters: A sharp thought experiment that grounds Brockman's interview into a discussable 'ambient intent layer' framework with concrete scenarios and clear contrasts. Held back from higher bands because it's speculative commentary, not an empirical product release or research artifact.

TechCrunch · AI

Nadella warns companies using proprietary AI models may be feeding future competitors

Microsoft CEO Satya Nadella argued in a personal blog post that companies using proprietary models from labs like OpenAI and Anthropic are handing over sensitive business data. Those labs could later use that insight to compete with their own customers. He calls this the 'reverse information paradox.' The post cites similar warnings from VC Jason Calacanis and Palantir CEO Alex Karp but offers no hard data or case studies. Worth noting: Nadella runs a competing model business, so this reads as positioning for Microsoft's own offerings.

Why it matters: Nadella personally warns about closed-source model risks — sharp angle, but the piece is pure opinion with no data or cases. H and R hit, K missing, landing right at the featured threshold.

Hacker News front page

The same TypeScript file costs 73% more tokens on Claude than on GPT

Playcode counted tokens across 16 real fixtures using each provider's official tokenizer. The same TypeScript file becomes 681 tokens on GPT-5.x's o200k but 1,178 on Claude's new tokenizer—a 73% gap. Claude Opus 4.8 and 4.6 share the same rate card, yet the new tokenizer silently adds ~30% more tokens for identical code. English and code are hit hardest; Chinese barely changed. DeepSeek and GLM were excluded because only rough estimates were available.

Why it matters: Real-file tokenizer benchmarking turns the industry pain point of incomparable token pricing into reproducible data. All three HKR axes hit, but this is a tooling insight rather than a product launch or model breakthrough — lands in the 78-84 band. No cross-source cluster sign...

TechCrunch · AI

The wildest allegations in Apple’s trade secrets lawsuit against OpenAI

Apple's 41-page complaint, filed July 11, accuses OpenAI of a coordinated effort to extract trade secrets from former Apple employees. One internal OpenAI message read: 'LOL, I found out I can access the [network storage], so funny.' The suit claims OpenAI asked job candidates to bring Apple-issued hardware to interviews, instructed them to transfer files via personal email, and used disappearing messages on Slack. Apple alleges at least five ex-employees were poached, leaking chip architecture, model training methods, and Siri team org charts. OpenAI has not yet filed a formal response.

Why it matters: Apple sues OpenAI for trade secret theft with vivid complaint details; HKR all hit. Score held below 85 because it's early-stage allegations with no ruling yet, and TechCrunch is a secondary source.

The Verge · AI

The 6 wildest claims in Apple's lawsuit against OpenAI

Apple's 41-page complaint accuses OpenAI of a systematic poaching campaign with aggressive tactics. OpenAI allegedly coached Apple employees on bypassing security checks and asked for 'show and tell' of internal projects during interviews. The suit claims OpenAI targeted managers who could bring entire teams, hiring at least 20 Apple staffers—some allegedly took next-gen chip designs. Apple wants data returned and damages; the post doesn't specify the amount. Grain of salt: this is one side's filing, but the 'show and tell' claim, if true, crosses a clear line.

Why it matters: Apple's lawsuit against OpenAI hits all three HKR marks with concrete allegations about poaching tactics and chip design risks. Not scoring higher because we only have Apple's side—OpenAI hasn't responded yet, so the full picture is still pending.

Hacker News front page

Apple's new SpeechAnalyzer beats Whisper Small on accuracy in first public benchmark

Inscribe benchmarked Apple's new SpeechAnalyzer API against the legacy SFSpeechRecognizer and three Whisper models on 5,559 LibriSpeech utterances. SpeechAnalyzer hit 2.12% WER on clean speech and 4.56% on noisy speech, beating Whisper Small by 1.62 and 3.39 points respectively while running ~3x faster. The legacy API scored 9.02% WER, worse than the 40MB Whisper Tiny. All engines ran fully on-device on an M2 Pro. Inscribe switched its default engine to SpeechAnalyzer and released all transcripts and scoring code. The post does not disclose SpeechAnalyzer's model architecture or parameter count.

Why it matters: First independent benchmark of Apple's SpeechAnalyzer with solid methodology (5,559 utterances, all on-device). Directly useful for voice product teams. Not 85+ because it's a single third-party benchmark on one dataset, not an Apple launch, and LibriSpeech alone doesn't cover...

Jul 13Monday

AI Chat-Group Daily (群聊日报)

GPT-5.6 Sol Pro decoded: 'Pro' is a reasoning mode, not a new model

Packet capture reveals OpenCode's Sol Pro is just gpt-5.6-sol with reasoning.mode: "pro" — not a separate model. Mode, effort, and service_tier can be freely combined. A simple greeting jumps from 12 to 1,527 input tokens with Pro enabled, roughly 100x more expensive. Separately, GPT-5.6 now charges for cache writes, potentially doubling Codex costs for long tasks. One user burned 19B tokens in two days, 98% from cache reads. The biggest shock: a researcher's 2024 open problem was solved by gpt-5.6-sol ultra in 46 minutes, verified correct by Fable.

Why it matters: First-hand packet capture with concrete numbers, not a rehash. The Sol Pro debunk and cache billing discovery both deliver real signal, but the source is an anonymous chat group without official confirmation, so the score stays at the featured threshold.

Hacker News front page

I love LLMs, I hate hype

George Hotz is excited about GPT-5.6, GLM-5.2, and coding agents, but calls out two things he hates: negative-valence hype about closing windows and perpetual underclasses, and the strawman jump from 'fancy autocomplete' to 'owning the whole light cone.' He argues AI progress is mostly Moore's law and commoditization, not frontier-lab magic, and that anti-open-source arguments are really about fear of commodification. He also walks back his earlier dismissal of models for programming—he's getting better at using them—but warns they can increase cognitive fatigue and that vibe-coded stuff is still slop.

Why it matters: George Hotz names and shames two hype patterns — fear-based negative valence and the 'own the whole light cone' leap — while walking back his earlier coding skepticism with a concrete GLM-5.2 + opencode example. Sharp, quotable, and backed by a real experiment, but it's ultima...

AI HOT (Curated Pool)

Codex and ChatGPT Work drop the 5-hour cap, roll out GPT 5.6 Sol efficiency gains

Three updates landed in 48 hours: the 5-hour usage cap is temporarily removed for Plus, Business, and Pro plans; GPT 5.6 Sol is getting efficiency improvements that reduce per-request usage, with numbers promised after quantification; and active users hit 6 million, with a usage reset rolling out within the hour. The post doesn’t say how long “temporarily” lasts or give a range for the efficiency gain, so I’d hold off on pricing that in.

Why it matters: Codex and ChatGPT Work both got updates — lifting the 5-hour cap is an immediate win for heavy users. GPT 5.6 Sol efficiency gains and 6M active users add substance, but without quantifying the efficiency bump or defining 'temporarily,' the score stays below 80.

Hacker News front page

Ploy migrated its production AI agent from Claude Opus 4.8 to GPT-5.6: 2.2x faster, 27% cheaper

Ploy's agent builds real marketing sites. For four months, no model beat Claude Opus. GPT-5.6 Sol is the first. Migration cut build time from 8 min to 3 min 42 sec, cost from $3.06 to $2.22, with a slightly higher visual score. The switch wasn't plug-and-play: eval harness, tool schemas, caching, and reasoning replay all needed rework because the stack had quietly specialized around Opus. The post doesn't disclose GPT-5.6's API pricing or context window.

Why it matters: Ploy published a same-day migration report from Claude Opus to GPT-5.6 with concrete latency and cost numbers plus engineering details — not a vendor case study. Downside: single-team experience, no failure cases or edge scenarios disclosed, so generalizability is unproven.

Jul 12Sunday

AI HOT (Curated Pool)

Altman now 'pretty sure' AI is net job-creating, Amodei also walks back job-killer claims

OpenAI CEO Sam Altman posted on X that he's 'pretty sure' AI has been net job-creating so far, a sharp pivot from his earlier 'potentially a little scary' warning. Anthropic CEO Dario Amodei also reframed automation as a productivity multiplier rather than a job killer. No studies yet show a significant AI impact on overall productivity or the labor market; the Yale Budget Lab found no AI-related job market shifts. The article notes some companies did cite AI for layoffs, but often as a shareholder-friendly excuse.

Why it matters: Altman and Amodei both pivoting to 'AI is net job-creating' is a strong narrative hook, but the article rests entirely on tweet quotes with no data backing, so the Knowledge axis misses. Score lands at the featured threshold of 72 — held back by the lack of empirical evidence.

AI HOT (Curated Pool)

Analyst: Apple's lawsuit against OpenAI could stall its hardware plans even if allegations fail

Apple sued OpenAI, alleging it poached 400 employees and stole engineering prototypes and confidential files. PP Foresight analyst Paolo Pescatore says the lawsuit alone could slow OpenAI's push to build consumer hardware that bypasses the iPhone. Stanford law professor Mark Lemley notes the case gets serious if ex-Apple staff actually brought confidential documents to OpenAI. The post does not disclose any specific hardware product names or launch timelines.

Why it matters: Apple sues OpenAI for poaching and IP theft; the analyst's take that the lawsuit itself will stall OpenAI's hardware roadmap is more informative than the allegations. Not scored higher because the post lacks concrete evidence or damages figures — it's analyst commentary only.

Computing Life · Share · Yage

Codex Merges into ChatGPT: Why Agents Are Going Cross-Interface

OpenAI merged Codex into ChatGPT and dropped its standalone desktop app. The same Codex tech that wrote code now handles docs, spreadsheets, and web pages, with over 5 million weekly active users—1 million of them for non-dev work. Anthropic and Cursor are making similar moves: putting different execution modes into one app so agents run in the background while users start tasks from the web, check progress on a phone, and approve results. Two axes drive this: agents are getting better at completing tasks on their own, and systems are compressing long execution runs into quick-to-review summaries, diffs, and anomalies. Coding got there first because software engineering already had mature verification tools like tests, diffs, and PRs. Knowledge work lacked that compression layer until recently, which is why reviewing a 30-slide deck on a phone in two minutes is now becoming feasible. Always-on doesn't require the cloud—a local Mac mini paired with a phone client can achieve the same pattern, with different trade-offs in responsibility and data boundaries. IDEs and desktop apps aren't disappearing; they're becoming specialized execution views beneath a cross-device, always-available delegation service.

Why it matters: A trend piece with a real analytical frame and numbers, not a product PR. Hits all three HKR: the headline hooks, the body delivers 5M WAU and the 'verification compression' concept, and the pain point lands. Held at 78 because it's a single-source opinion piece without cross-...

AI HOT (Curated Pool)

Apple sues OpenAI over talent poaching and trade secret theft, triggered by an ex-employee's 'LOL'

Bloomberg revealed details of Apple's lawsuit against OpenAI. Ex-iPhone engineer Chang Liu kept his work MacBook after leaving and found a bug that let him still access Apple's internal servers. He messaged colleague Alyssa Peng 'LOL I found I can still access network storage,' and she replied 'I'm ready' before helping him grab more confidential data. At the center is former Apple executive Tang Tan, who led iPhone and Apple Watch design and left in late 2023 to become OpenAI's chief hardware officer. Apple says over 400 employees have jumped to OpenAI, hollowing out multiple engineering teams. Apple contacted OpenAI in February asking for an investigation; OpenAI did not respond.

Why it matters: Bloomberg dug up key evidence in Apple's trade secret lawsuit against OpenAI—strong narrative and solid detail. Score capped below 85 because this is fundamentally a legal dispute, not an AI tech or product development.

Hacker News front page

Wealthy AI workers push San Francisco home prices to record highs

San Francisco's median home price hit $1.76M in May 2026, up over 14% year-on-year, reclaiming the top spot as the most expensive US city for buyers. Redfin's chief economist pins the surge on AI wealth: OpenAI employees cashed out $6.6B in stock last October, averaging $11M per person, and Anthropic staff recently sold about $6B. One seller even offered to accept shares in OpenAI or Anthropic instead of cash. The post doesn't give exact IPO dates, only that both companies are expected to go public this year or next.

Why it matters: BBC piece uses Redfin data and OpenAI/Anthropic cash-out figures to make the AI wealth effect concrete; hits all three HKR axes. Score capped at 72 because the topic is socioeconomic rather than AI tech/product, so direct knowledge gain for industry readers is limited.

Hacker News front page

geohot's "AI 2040": intelligence isn't everything, local AI is the freedom line

George Hotz argues against hard-takeoff AI from firsthand hardware experience at comma.ai. Reality is full of supply-chain snags, wrong parts, and 3-month fab cycles that no amount of token quality can speed up. He calls the AI-2027-style narrative a self-fulfilling push for a sci-fi nanny state. His alternative is Plan L: a local, user-aligned AI that never refuses—even if you ask it to cover up a murder. He posts a screenshot of ChatGPT declining to help after "I just killed my wife" and calls it an alignment failure. Core claim: intelligence is only a bottleneck for some things, not the ultimate lever on the world; without a physics hack, there is no hard takeoff.

Why it matters: Hotz rebuts hard-takeoff narratives with concrete hardware experience from comma.ai (tape-out cycles, physical constraints), and his identity draws attention. Deduction: this is an opinion piece, not a product launch or research release — commentary defaults to a lower ceiling...

AI HOT (Curated Pool)

OpenAI releases GPT-5.6 medical evaluation: smallest Luna variant beats GPT-5.5 at lowest reasoning strength, 25× cheaper

OpenAI had specialists write answers with unlimited time and web access, then other doctors blind-rated them against GPT-5.6 across 20,000 scores on accuracy, communication, completeness, instruction-following, and health-decision helpfulness. All GPT-5.6 models outperformed doctors significantly, and doctors found fewer flaws in GPT-5.6 answers than in peer-written ones. The smallest variant, GPT-5.6 Luna, surpassed the highest-reasoning GPT-5.5 at its lowest reasoning strength while costing 25× less; the largest variant, GPT-5.6 Sol, set a new high bar. The post doesn't disclose the disease mix or specialist composition tested.

Why it matters: OpenAI ran 20,000 blind ratings pitting GPT-5.6 models against specialist physicians across five dimensions. The smallest Luna model at minimum reasoning effort already beat GPT-5.5 at max effort, and doctors flagged more issues in peer-written answers than in GPT-5.6's. The e...