Skip to content

OpenAI / ChatGPT

Everything OpenAI: the GPT models, ChatGPT and Sora, company strategy and people moves.

Latest picks

301–320 of 1,549

Sep 4Friday

Hacker News front page

Which tools Claude Code, Codex, and Cursor pick in 16,893 real coding sessions

Armature ran nearly 17k experiments across 75 repos and 1,163 prompt variants to see which services Claude Code, Codex, and Cursor actually install. They simulated four personas—vibe coder, junior, senior, and enterprise engineer—and had agents go from analysis to implementation. The post discloses partial findings: in object storage, Cloudflare R2 started beating Amazon S3 once a simulated human was added to the loop; in databases, Neon was repeatedly recommended. Full leaderboards and raw traces are published, but the article body cuts off before covering more categories.

Why it matters: Armature ran 16,893 simulated sessions to surface tool selection preferences across Claude Code, Codex, and Cursor—solid sample size, useful signal. The caveat: Armature sells growth services to dev-tool companies, and while they disclose it upfront, that stake puts a question...

AI HOT (Curated Pool)

Perplexity to integrate GPT-6 Astra, CEO says it tops WANDR benchmark

Perplexity CEO Aravind Srinivas says the company will integrate OpenAI's newly released GPT-6 Astra, claiming it far outperforms other models on deep and broad research tasks at lower cost. It will roll out to Perplexity Computer Pro and Max users first. The post does not disclose a launch date, WANDR scores, or cost figures.

Why it matters: A top AI search product quickly adopting the latest flagship model is newsworthy. But the post lacks WANDR scores, cost figures, and a launch timeline — the actual improvement is still unclear, so it doesn't push past 85.

Latent Space

GPT-6 Astra: an automated AI Engineer you can hire for <$6 an hour

Latent.Space got early access to GPT-6 Astra and burned over 20B tokens on real-world tasks. The biggest surprise: it works as a fully capable AI engineer—choosing models, labeling data, monitoring pipelines, reading logs, deploying and debugging systems, and managing 20–50 sub-agents in parallel. At 33 tokens/sec and a max rate of $50 per million tokens, that comes out to under $6 an hour. Over a month the team built a dozen internal tools, including a GitHub+Vercel replacement prototype and a game AI for a board with 10,000x more legal moves than Go. Astra scored 97.6% on FrontierMath and 99.9% on ARC-AGI-3, though the post doesn't specify benchmark versions or evaluation conditions. I'd discount this a bit: these are preview latency numbers, and GA speeds may differ.

Why it matters: GPT-6 Astra is OpenAI's first Stargate supermodel, and Latent.Space got early access with a 20B-token real-world test, quantifying it as a sub-$6/hour AI engineer. This is an industry-level event with dense cross-source coverage and all three HKR axes hit. Not 95+ yet because ...

AI HOT (Curated Pool)

OpenAI GPT-6 Astra system card: model's control over its own chain-of-thought jumps from 16% to 61%

Rohan Paul pulls one key shift from Astra's 117-page system card: the model's ability to control its own chain-of-thought rose from 16.1% in GPT-5.6 Sol to 60.9%, with monitorability dropping accordingly. The post doesn't detail the evaluation method or risk scenarios—I'd discount the number until the full system card is out.

Why it matters: A safety finding from GPT-6 Astra's system card with concrete numbers and a counterintuitive tradeoff hits all three HKR axes. Score held below 85 because this is a secondhand interpretation, not the original card, and the measurement methodology isn't disclosed.

AI HOT (Curated Pool)

OpenAI launches GPT-6 Astra, hitting SOTA on multiple benchmarks

OpenAI dropped GPT-6 Astra, claiming SOTA on FrontierMath Tier 4, ARC-AGI 3, and TerminalBench-4.0, plus leading scores on Terminal-Bench Science 0.1 and HealthBench Pro. The post is a headline with benchmark names only—no params, architecture, release date, or raw scores, so I'd hold for more details.

Why it matters: The GPT-6 Astra codename and SOTA claims are newsworthy on their own, but the post contains only benchmark names with zero concrete numbers, architecture details, or timeline. Per policy, default to the lower band when info is thin — 82 within the 78-84 range.

AI HOT (Curated Pool)

OpenAI launches GPT-6 Astra, hits 99.9% on ARC-AGI 3 — but that score comes with a big asterisk

OpenAI released GPT-6 Astra, rolling out today to select orgs and soon to all ChatGPT Plus, Pro, Business, Enterprise, and API users. API pricing matches Claude Fable 5/5.1 at $10/M input and $50/M output. The headline 99.9% on ARC-AGI 3 is real but inflated: it used OpenAI's custom Provider Adapter harness at $19K, while the default harness scored 62.7% at $26K. The custom harness preserves reasoning state across requests and compacts long conversations, letting the model reuse prior work. Security scores are genuinely strong — 100% on ExploitBench, 42.4% on ExploitGym, 99.2% on SRE-Bench reverse engineering. Long-context needle retrieval hit 100% at 256K–512K and 96.3% at 512K–1M. On Artificial Analysis's Intelligence Index, Astra ties GPT-5.6 Sol at 61, 5 points below Claude Fable 5.1 and behind Meta's Muse Spark 1.3. It leads the Coding Agent Index cost-efficiency frontier: same cost as Sol at max effort but 2 points higher, and less than half the per-task cost of Fable 5 for the same score. Simon hasn't tried it yet; the API label will be gpt-6-astra.

Why it matters: GPT-6 Astra is OpenAI's direct Fable competitor, priced identically and claiming higher benchmarks. The 99.9% ARC-AGI 3 score required a custom harness — default harness hit 62.7% — which is the key caveat. ExploitBench went from 78.5% to 100%, a concrete security jump. Simon ...

AI HOT (Curated Pool)

Sam Altman announces GPT-6 Astra, calling it the world's best model across multiple domains

Sam Altman announced GPT-6 Astra, positioning it as the world's best model for computer use, professional work, science, coding, and cybersecurity. He said the team took extra time to meet the safety and alignment standards required for this capability level. Three benchmark scores were shared: FrontierMath Tier 4 at 98%, ARC-AGI 3 at 99.9%, and ExploitBench at 100%. The post does not disclose parameter count, pricing, access method, or a concrete launch date—only the title and these scores are available so far.

Why it matters: A flagship model generation drop from OpenAI, announced by Sam Altman himself, is an industry-shaking event. Three benchmark scores are new SOTA, explicitly targeting hardcore use cases like computer use, coding, and security. The post doesn't disclose parameter count or archi...

Hacker News front page

OpenAI GPT-6 Astra hits 99.9% on ARC-AGI-3 for $19K

GPT-6 Astra scored 62.7% for $26K on ARC-AGI-3 Semi-Private with the Standard harness, and 99.9% for $19K with the Provider Adapter harness, which preserves opaque reasoning state and uses compaction. Astra beat the median human in action efficiency on 96% of levels. It built compact symbolic world models from unfamiliar environments and invented its own shorthand to track state and plan. The post does not disclose parameter count, architecture, or release date.

Why it matters: GPT-6 Astra hits 99.9% on ARC-AGI-3, the first flagship model near-perfect on this benchmark, with cost dropping from $26K to $19K. Cross-source coverage is guaranteed. Not 95+ because this is the ARC Prize blog, not an OpenAI release, and the post doesn't disclose Astra's arc...

AI HOT (Curated Pool)

OpenAI launches GPT-6 Astra, targeting Computer Use and agent alignment

OpenAI Chief Research Officer Mark Chen announced GPT-6 Astra, calling it the result of years of pretraining, RL, and post-training work—the most capable and best-aligned model yet. The post is a single sentence; it doesn't detail what Computer Use can do, how agent alignment was achieved, or provide any performance numbers or timeline.

Why it matters: OpenAI's Chief Research Officer announces GPT-6 Astra with Computer Use and agent alignment — an industry-shaking event. But the post is a single sentence with no performance numbers, safety mechanisms, or gen-over-gen gains, so the K axis is a complete miss. Per policy, flags...

AI HOT (Curated Pool)

OpenAI starts rolling out GPT-6 Astra to all Plus users

OpenAI announced the rollout of GPT-6 Astra, prioritizing all Plus users rather than limiting it to Pro, Business, and Enterprise plans. The release will take a few days, with multiple new systems running at scale for the first time and significant compute being brought online. The post does not disclose model parameters, pricing changes, or specific capability benchmarks.

Why it matters: GPT-6 rolling out to all Plus users at once is OpenAI's largest model launch to date. The post doesn't disclose parameters, pricing, or benchmarks — real-world performance remains to be seen — but the launch itself is an industry-level event.

AI HOT (Curated Pool)

Artificial Analysis benchmarks GPT-6 Astra: coding agent score matches Fable 5 at 2.5× the price

Artificial Analysis ran its Coding Agent Index on GPT-6 Astra. Score 67, on par with Claude Opus 5 and Fable 5. Cost is under half of Fable 5 but roughly 2.5× GPT-5.6 Sol (max). Token efficiency improved ~70% over GPT-5.6 Sol. The post doesn't disclose latency or task completion rates, so hold off on real-world expectations.

Why it matters: Artificial Analysis's Coding Agent Index is a widely-cited independent benchmark. GPT-6 Astra scores 67, tying Claude Opus 5 and Fable 5, with ~70% better token efficiency but at 2.5x the price of GPT-5.6 Sol. The price-performance reversal is newsworthy, but this is a third-p...

AI HOT (Curated Pool)

OpenAI launches GPT-6 Astra, the first model it classifies as critical-risk under its own cybersecurity framework

OpenAI shipped GPT-6 Astra, and president Greg Brockman says it may already qualify as AGI under OpenAI's own definition—outperforming humans at most economically valuable work. Astra scores 99.9% on ARC-AGI-3, 97.6% on FrontierMath Tier 4 v2, and a perfect 100% on ExploitBench. It is the first model OpenAI has rated as a critical cybersecurity risk in its Preparedness Framework. Token prices are 2.5× higher than predecessor Sol and on par with Anthropic's Fable 5.1, though OpenAI argues per-task cost is lower. Pretraining ran on over 100,000 GPUs at the Stargate facility in Texas—OpenAI's largest training run ever. The post says paying ChatGPT customers and cloud platforms will get access in the coming days, but does not give a specific date.

Why it matters: GPT-6 Astra launch with OpenAI's first self-declared AGI-era framing and Critical-level cybersecurity classification under its Preparedness Framework. Brockman's direct AGI claim is backed by concrete ARC-AGI-3 and FrontierMath scores. Cross-source cluster confirmed; this is a...

AI HOT (Curated Pool)

OpenAI launches GPT-6 Astra, benchmarks fully surpass Claude Fable 5.1

OpenAI published official benchmarks for GPT-6 Astra: 99.9% saturated ARC-AGI-3, 100% on ExploitBench, fully beating Claude Fable 5.1 which held SOTA for just two days, and at a lower price. The post only gives headline numbers—no pricing details, parameter count, or release date, so I'd wait for third-party evals.

Why it matters: OpenAI officially posted GPT-6 Astra benchmarks, beating Claude Fable 5.1 on ARC-AGI-3 and ExploitBench — an industry-shaking release. Pricing, param count, and launch date are missing from the post, so I'm holding at 92 until third-party evals land.

AI HOT (Curated Pool)

OpenAI launches GPT-6 Astra, starting with vetted Daybreak cybersecurity clients

OpenAI released GPT-6 Astra, rolling it out first to vetted Daybreak cybersecurity clients. Plus, Pro, Business, Enterprise, API, and AWS access will follow within days. The post doesn't disclose model specs, pricing, or capabilities—hold off on conclusions until more details land.

Why it matters: GPT-6 launching exclusively through vetted Daybreak security customers is a featured-worthy rollout strategy on its own. But the post gives zero model specs, pricing, or capability details, so the K axis misses and the score caps at 82. Will raise it once concrete numbers land.

AI HOT (Curated Pool)

OpenAI launches GPT-6 Astra, rolling first to Daybreak Access orgs

OpenAI released GPT-6 Astra, currently limited to organizations in the Daybreak Access program. The post doesn't spell out what Daybreak Access is, nor any model specs or benchmarks. Plus, Pro, Business, and Enterprise users will get it in the coming days.

Why it matters: A GPT-6 launch is an industry-level event — featured tier is warranted even with only a title and access-tier info. Score stays below 90 because the post lacks any specs or benchmarks (K axis missed); can bump once concrete details surface.

Hacker News front page

OpenAI starts rolling out GPT-6 Astra after flagging its advanced cyber capabilities

OpenAI is rolling out GPT-6 Astra in phases, starting with companies in its application-based cybersecurity program. ChatGPT Plus, Pro, Business, and Enterprise users will get access later. OpenAI itself just warned about Astra's advanced cyber capabilities, but the post doesn't detail safeguards or restrictions.

Why it matters: First public rollout of GPT-6 Astra, coming right after OpenAI's own warning about its advanced cyber capabilities — the 'warn first, ship later' rhythm is itself the story. CNBC exclusive, industry-shaking tier. Minus 3 points because the article doesn't detail the safety gua...

Hacker News front page

OpenAI launches GPT-6 Astra; Brockman says 'Welcome to the AGI era'

OpenAI released GPT-6 Astra on Thursday, with president Greg Brockman calling it a potential arrival of AGI. Trained on over 100,000 GPUs at the Texas Stargate site, it is OpenAI's first model to use other models heavily in training supervision. Astra works directly inside software: it formatted a legal contract, built a 3D game, laid out a circuit board, and filled a tax draft, while setting new marks on math and science evals. OpenAI admits Astra is harder to monitor—it showed declines in oversight-evasion tests—and chief scientist Jakub Pachocki said improving monitorability is a research priority. The model rolls out first to a limited set of orgs via the Daybreak Access program, then to paid users and API developers in coming days. I'd temper expectations: Astra's cyber capabilities hit OpenAI's 'critical' threshold, meaning it can find and exploit unknown vulnerabilities autonomously, so the strongest cyber features stay restricted to trusted testers.

Why it matters: GPT-6 launch with OpenAI's president calling it the start of the AGI era — an industry-shaking event. 100K+ GPU training, multi-model supervision, and direct software operation are all first disclosures with solid detail. Hits all three HKR axes, importance near ceiling.

AI HOT (Curated Pool)

OpenAI launches Astra, a model for computer and browser use that's drawing fire over opaque recurrence

OpenAI released Astra on Thursday, pitching it as a new high for speed, accuracy, and safety in computer and browser tasks. President Greg Brockman called it the company's most intelligent and aligned model yet. Astra rolls out first to Daybreak cybersecurity customers, then to paid plans and the API within a week. The controversy stems from an earlier OpenAI blog that mentioned an opaque recurrence mechanism—the post doesn't explain how it works or what risks it introduces. I'd hold off on the hype: the capability claims are big, but transparency and safety details are still missing.

Why it matters: OpenAI's new flagship model Astra, focused on computer use, is a same-day must-cover. The 'opaque recurrence' controversy is flagged but not explained in the body — otherwise this would be a 92.

Sep 3Thursday

The Verge · AI

ChatGPT, Grok, and Claude all went down at the same time on Thursday

Around 11AM ET Thursday, ChatGPT, Grok, and Claude all started having issues at roughly the same time. ChatGPT returned errors across chat, login, file uploads, voice, search, deep research, and image generation; its status page cited elevated errors for ChatGPT and Codex. Anthropic's Claude chatbot and Claude Code were also affected. The post doesn't detail Grok's specific symptoms, the recovery timeline for each service, or whether the outages share a root cause.

Why it matters: A simultaneous outage across ChatGPT, Grok, and Claude is a rare event that directly disrupts workflows for a huge user base. Missing root cause and recovery timeline keeps it from 95+, but the topic is strong enough for featured.

Hacker News front page

OpenAI, Claude, and Grok all went down at once—users suspect a Cloudflare cascade

A Hacker News thread noted that OpenAI, Claude, and Grok all went down around the same time. Users pointed to Downdetector spikes for Cloudflare, Azure, AWS, and Google Cloud near 7:30, suspecting a cascade from Cloudflare or another shared dependency. Other guesses include user migration overload and deliberate attack, but the post is community speculation—no official root cause is confirmed.

Why it matters: Simultaneous outage across OpenAI, Claude, and Grok with high HN engagement. Downdetector data points to Cloudflare or shared infra as a possible common cause. The event is conversation-worthy but lacks a confirmed root cause, so it lands at the 78 featured threshold rather th...