Skip to content

#推理

0 today

Sep 28Monday

OpenAI News

Basis cuts tax workbook time in half with GPT-6 Astra

Accounting AI startup Basis tested GPT-6 Astra against GPT-5.6 Sol on a 50-tab tax workbook. Astra finished 50% faster. Basis says Astra understands user intent better, picks a more direct path from the start, and wastes fewer tokens. The model also adjusts reasoning effort per step—more compute for hard parts, less for easy ones—while keeping its cache intact. Internal eval scores improved ~20%, driven by Astra knowing when to ask questions, flag assumptions, or follow templates without explicit rules. The post doesn't disclose exact latency or cost figures, only says it's "more economical."

Sep 22Tuesday

OpenAI News

Parallel cuts research time and cost in half with GPT‑6 Astra

Parallel, an AI agent infrastructure startup, used GPT‑6 Astra to research labor-market data across six states over six months. The model cut both time and code cost by 50% by issuing more targeted searches and delegating sub-tasks to parallel agents. The post doesn't specify which prior models were used for comparison.

Sep 10Thursday

OpenAI News

OpenAI launches ChatGPT for Financial Services with built-in financial data and GPT-6 Astra

OpenAI introduced ChatGPT for Financial Services, a tailored Work experience that pairs GPT-6 Astra's reasoning with built-in premium data from Daloopa, PitchBook, LSEG News, and Crunchbase. Designed with Morgan Stanley and Evercore, it targets investment banking and equity research workflows: value analysis, LBO modeling, buyer screening, earnings analysis, and pitchbook prep. OpenAI indexes and hosts the data to improve accuracy and provide granular citations. The post does not disclose pricing or a launch date; it notes that firms can centrally manage access and data connections under ChatGPT's enterprise governance.

Why it matters: OpenAI's first vertical-specific product, directly integrating four premium financial data sources and co-designed with Morgan Stanley and Evercore — not a generic wrapper. But the post doesn't disclose pricing, data latency, or compliance certifications, which are hard gates ...

Sep 2Wednesday

Hugging Face Blog

Allen AI's BenchMIRT uses psychometric IRT to reveal what LLM benchmarks actually measure

Allen AI open-sourced BenchMIRT, a method that audits LLM benchmarks using multidimensional item response theory. It analyzed 100 models across 16 benchmarks and 34K+ questions, automatically recovering two dominant capability dimensions: safety and general reasoning. A BBQ question about a grandson and grandfather booking an Uber tests age bias but also requires reasoning. WildJailbreak's harmful and benign prompts map to safety and reasoning respectively—averaging them into one score hides that split. BenchMIRT identifies which questions best separate strong from weak models, enabling cleaner evaluation with fewer items. Code, data, and the tech report are public.

Why it matters: Allen AI open-sourced a method that uses item response theory to audit benchmarks, backed by 100 models, 16 benchmarks, and 34k questions. Score stays below 80 because it's a methodology tool rather than a shippable product update, but it hits all three HKR axes and is genuine...

Aug 25Tuesday

Hugging Face Blog

IBM details the full pipeline behind Granite 4.2, from pre-training to agentic RL

IBM published a technical walkthrough of the Granite 4.2 model family on the Hugging Face blog. It covers architecture, pre-training, SFT data quality control, and a multi-stage RL pipeline. The RL curriculum has three phases: foundational skills, agentic RL for tool use on the 8B and 30B models, and RLHF alignment. The post also mentions FP8, FP4, and GGUF quantization. Specific benchmark scores and hardware details are not included in the provided excerpt.

Why it matters: A solid training pipeline breakdown with strong H and K, but Granite's limited community pull drags down R. The post doesn't disclose pretraining data or hardware specs, so it can't push past 78. Featured because the engineering detail is real — model trainers will bookmark this.

Hugging Face Blog

Quantization-Aware Healing: a 4-bit model that beats its full-precision original

Multiverse Computing introduces Quantization-Aware Healing (QAH), a recovery step for models that have been both structurally compressed and quantized. Applied to a GPT-OSS 120B pruned to 60B and quantized to MXFP4, the 4-bit model beats its bfloat16 original on 7 of 9 benchmarks, including reasoning and math. QAH also outperforms standard QAT on compressed models. The post doesn't disclose latency or throughput numbers, so real-world savings are still TBD.

Why it matters: Counterintuitive compression result: a 4-bit model beats its bfloat16 original on most benchmarks. Method is concrete, numbers are clear, directly useful for deployment and inference folks. Not scoring higher because Multiverse Computing isn't a tier-1 lab, and the post doesn'...

Aug 1Saturday

OpenAI News

OpenAI's internal model Astra solved ten open math problems untouched for over a decade

OpenAI published ten new results in math and theoretical CS produced by its internal model Astra. The problems—untouched for at least a decade—include high-dimensional sphere packing, existence of non-sofic groups, a disproof of Connes's rigidity conjecture, and polynomial-factor hardness for the closest vector problem. All arguments were formalized in Lean, and the model's reasoning traces are released. Total token cost was roughly $2,000 at Sol API rates. OpenAI states the mathematical arguments were generated by the system; humans only prepared manuscripts and formalized proofs, and authorship should reflect that.

Why it matters: OpenAI's Astra model produced verifiable advances on ten decade-old math problems, all formalized in Lean. A landmark for AI in hard science, but pure theory is distant from product/agent impact — policy deducts 10–15, landing at 78.

Jul 31Friday

OpenAI News

OpenAI lays out its “abundant intelligence” playbook: price cuts, efficiency gains, and a full-stack flywheel

OpenAI published a strategy post on July 31 explaining its “abundant intelligence” approach. The core loop: more capable and cheaper models drive broader adoption, which generates revenue and feedback to fund the next round of R&D and infrastructure. Concrete numbers: GPT-5.6 Luna input/output prices dropped 80% to $0.20/$1.20 per million tokens; GPT-5.6 Terra dropped 20%. GPT-5.6 Sol Fast mode delivers 2.5x speed at 2x price with no intelligence change. On the engineering side, Sol helped cut end-to-end serving costs by 20% and improved speculative-decoding efficiency by over 15%. On the public ARC-AGI-3 benchmark, better retained reasoning and context management lifted Sol’s score from 13.3% to 38.3% while using 6x fewer output tokens. Product stats: ChatGPT has over 1B active users and 2M businesses; six months after signup, daily messages rise ~50% and use-case breadth roughly doubles. Agentic work via Codex now accounts for 99.8% of OpenAI’s weekly output tokens. No new model was announced—this is a strategy piece.

Why it matters: OpenAI's official blog lays out its 'abundant intelligence' strategy with concrete pricing data (GPT-5.6 Luna down 80%). Not a product launch, so it doesn't hit 85, but as a strategic signal it's worth featuring.

Jul 16Thursday

Hugging Face Blog

Model routing is simple—until you measure real cost, not sticker price

IBM Research found that routing by model sticker price backfired in agent workloads. Across 417 AppWorld tasks, Claude Sonnet 4.6 cost $79 total vs. GPT-4.1's $155—nearly double—because Sonnet's lower cache-read pricing exploited high context reuse across steps. The post argues real cost, latency, and complexity all depend on workload-infrastructure interaction, making routing a systems optimization problem, not a classification one.

Why it matters: IBM ran 417 AppWorld tasks and found that routing by list price alone fails—Sonnet 4.6 cost $79 total while GPT-4.1 cost $155, nearly double. The core insight: when agents reuse the same context repeatedly, cache-read pricing dominates the total bill. Concrete numbers, counter...

Jun 25Thursday

Google Research Blog

How reasoning unlocks parametric knowledge in LLMs

Google Research shows that letting models think before answering sharply improves their ability to recall facts from training data. On Natural Questions, Gemini 2.5 Pro jumps from ~40% accuracy without reasoning to over 70% with it. The gain comes from the model connecting fuzzy memories into verifiable chains, not from external retrieval. The reasoning traces often include self-questioning and fact-checking steps. The post only covers QA tasks so far.

Why it matters: Google Research published a mechanism study with concrete numbers showing how reasoning helps models retrieve parametric knowledge, with a clear 40%→70% jump. Missing generalization evidence beyond Natural Questions keeps the score from going higher. Useful for RAG and eval pr...

Jun 18Thursday

OpenAI News

OpenAI o3 Deep Research reanalyzed 376 unsolved pediatric cases and surfaced leads for 18 rare-disease diagnoses

Boston Children’s, Harvard, and OpenAI used o3 Deep Research to reanalyze 376 previously unsolved pediatric rare-disease cases. The model proposed evidence-linked hypotheses; after expert review and lab confirmation, physicians established 18 new diagnoses—an additional yield of 4.8%. The model never made clinical decisions. All confirmed diagnoses went through CLIA-certified lab validation. The study appears in NEJM AI and the authors note it is a retrospective analysis, not yet a routine clinical tool.

Why it matters: NEJM AI-published study: o3 deep research reanalyzed 376 unsolved pediatric rare-disease cases and surfaced 18 new diagnoses (4.8%). Has a paper, concrete numbers, and a CLIA validation pipeline — not a PR fluff piece. Held at 78 rather than 85+ because it's a single study, no...

Jun 3Wednesday

NVIDIA Blog

NVIDIA Research Presents Grasping, Autonomous Driving and Agent Training Work at CVPR

NVIDIA Research presented three physical AI papers at CVPR: GraspGen-X was trained on 2 billion simulated grasps, LCDrive cuts reasoning tokens by about half versus text-based reasoning, and NitroGen trains embodied agents across more than 1,000 games and 40,000 hours of interaction.

Why it matters: HKR-H/K/R all pass: NVIDIA’s CVPR bundle gives concrete mechanisms and scale numbers. It stays in the low 78–84 band because it is a vendor research roundup, not a major model or product launch.

OpenAI News

Introducing new capabilities to GPT-Rosalind

OpenAI says GPT-Rosalind adds biological reasoning, medicinal chemistry, genomics analysis, and experimental workflow capabilities; the RSS snippet does not disclose model parameters, benchmark results, pricing, or access conditions.

Why it matters: OpenAI’s vertical model update clears HKR-H and HKR-R, but HKR-K fails because evals, parameters, and access terms are missing. That keeps it at the featured floor.

May 22Friday

NVIDIA Blog

NVIDIA GTC Taipei at COMPUTEX: Live Updates on What’s Next in AI

NVIDIA won four COMPUTEX 2026 Best Choice Awards for Vera Rubin NVL72, Jetson Thor, and Alpamayo; Vera Rubin NVL72 connects 36 Vera CPUs and 72 Rubin GPUs, and NVIDIA says it delivers up to 10x higher inference performance per watt and 10x lower cost per token.

Why it matters: HKR-H/K/R all pass: NVIDIA gives concrete Vera Rubin NVL72 specs and a 10x inference-efficiency claim, directly tied to AI compute costs. The source is NVIDIA’s event blog, so this stays below the 85 same-day must-write band.

May 20Wednesday

OpenAI News

An OpenAI model has disproved a central conjecture in discrete geometry

An OpenAI model solved the 80-year-old unit distance problem and disproved a major conjecture in discrete geometry; the post does not disclose the model name, proof mechanism, or reproducibility conditions.

Why it matters: HKR-H/K/R all pass: the OpenAI math result is novel, concrete, and debate-starting. Missing model name, proof mechanism, and reproducibility keep it at 85, not a higher P1.

May 7Thursday

OpenAI News

Advancing Voice Intelligence with New Models in the API

OpenAI introduced new realtime voice models in its API for voice intelligence. The RSS snippet says they reason, translate, and transcribe speech; the post does not disclose counts, pricing, or limits.

Why it matters: OpenAI’s official voice API update hits HKR-H/K/R, but the available body gives capability direction only. Model count, pricing, latency, and context limits are not disclosed, so it stays at the top of 78–84.

May 5Tuesday

OpenAI News

GPT-5.5 Instant: smarter, clearer, and more personalized

OpenAI updated ChatGPT’s default model to GPT-5.5 Instant for default chat use. The RSS snippet says answers are more accurate, hallucinations are reduced, and personalization controls improved; the post does not disclose metrics, pricing, or context window.

Why it matters: HKR-H/K/R all pass: OpenAI changed ChatGPT’s default model to GPT-5.5 Instant. The post lacks evals, pricing, and context window details, so it stays at the low end of the 85–94 band.

Apr 29Wednesday

X · @OpenAI

A 60-Year-Open Erdős Problem Was Solved With Help From GPT-5.4 Pro

OpenAI says GPT-5.4 Pro helped solve an Erdős problem open for 60 years. The post names Sebastien Bubeck, Ernest Ryu, and Andrew Mayne, but does not disclose the problem name, proof details, or reproducible conditions.

Why it matters: HKR-H and HKR-R pass because an OpenAI model aiding a 60-year Erdős problem is a strong AI-research hook. HKR-K fails: no problem name, proof details, or reproduction conditions are disclosed.

Apr 25Saturday

X · @AnthropicAI

New Anthropic research: Project Deal

Anthropic announced Project Deal and had Claude buy, sell, and negotiate for employees in a San Francisco office marketplace. The setup is confirmed as an internal marketplace; the post does not disclose scale, model version, or outcome metrics.

Why it matters: This clears featured on HKR-H and HKR-R: Anthropic has attention weight, and an agent negotiating office deals is inherently discussable. It stays mid-band because HKR-K is weak; the post gives the setup, but not sample size, model version, success metrics, or controls.

Apr 24Friday

X · @OpenAI

Introducing GPT-5.5

OpenAI introduced GPT-5.5, and it is now available in ChatGPT and Codex. The RSS snippet says it targets real work and agents, can understand complex goals, use tools, check its work, and carry more tasks to completion; the post does not disclose parameters, pricing, context window, or benchmark results. What matters is the execution loop, not the headline's “new class of intelligence.”

Why it matters: OpenAI launching GPT-5.5 in ChatGPT and Codex is same-day mandatory coverage. HKR-H/K/R all pass: new model release, concrete agent-workflow claims, and direct impact on daily AI work. Price, context window, params, and benchmarks are undisclosed, so it stays below 95.

Apr 23Thursday

OpenAI News

Introducing GPT-5.5

OpenAI introduced GPT-5.5 and says it targets complex cross-tool tasks such as coding, research, and data analysis. The RSS snippet only confirms “faster” and “more capable”; the post does not disclose benchmarks, context window, pricing, release timing, or availability, which are the details practitioners should watch.

Why it matters: An OpenAI flagship-model release is same-day news, so HKR-H and HKR-R are clear. HKR-K fails because the post discloses the name and use cases but not benchmarks, context window, price, or availability, so this stays featured rather than p1.

Apr 21Tuesday

OpenAI News

Introducing ChatGPT Images 2.0

OpenAI introduced ChatGPT Images 2.0 as a new image generation model, highlighting better text rendering, multilingual support, and visual reasoning. The RSS snippet names only these three upgrades; the post does not disclose architecture, resolution, pricing, latency, or availability. What matters is whether text fidelity and multilingual consistency improve in real use; for now, only headline-level details are disclosed.

Why it matters: A primary-source OpenAI image update clears HKR-H and HKR-R: the 2.0 label and text-rendering claim hit real workflows. HKR-K is weak because the post discloses only three upgrade areas; resolution, price, latency, architecture, and rollout are absent, so it stays just above the

Apr 17Friday

X · @OpenAI

Introducing GPT-Rosalind, OpenAI's frontier reasoning model for biology, drug discovery, and translational medicine

OpenAI introduced GPT-Rosalind as a reasoning model for biology, drug discovery, and translational medicine research. The title and snippet disclose its intended domains; the post does not disclose size, benchmarks, availability, pricing, or launch timing. The key point is research targeting, but reproducible details are absent so far.

Why it matters: An official OpenAI announcement plus the unusual biology/drug-discovery positioning gives this HKR-H and HKR-R. HKR-K is weak because the post discloses only the model name and target domains; benchmarks, params, access, and launch timing are not disclosed, so it stays at the low

Apr 16Thursday

X · @claudeai

Introducing Claude Opus 4.7, our most capable Opus model yet.

Claude introduced Opus 4.7 and describes it as its most capable Opus model so far. The RSS snippet gives three claims: better rigor on long-running tasks, more precise instruction following, and self-verification before replying; the post does not disclose benchmarks, context window, pricing, or rollout scope. What matters is whether those claims show up in public evals, not the tagline.

Why it matters: This is a substantive Anthropic model release and clears HKR-H/K/R: a new Opus, three testable behavior claims, and strong resonance with Claude-heavy practitioners. The score stays in the high 80s because benchmarks, pricing, context window, and rollout scope are not disclosed.

OpenAI News

Introducing GPT-Rosalind for life sciences research

OpenAI released GPT-Rosalind on April 16, 2026, and made it available as a research preview in ChatGPT, Codex, and the API for qualified customers. The post says it targets biology, drug discovery, and translational medicine, and adds a free Codex life sciences plugin connecting to 50+ scientific tools and data sources. The real signal is deployment breadth: Amgen, Moderna, and Thermo Fisher Scientific are involved, but the post does not disclose model size, pricing, or benchmark scores.

Why it matters: HKR-H lands because OpenAI is shipping a vertical life-sciences model; HKR-K lands on access paths and the 50+ tool/data plugin. HKR-R also lands on the domain-model debate, but missing params, pricing, and benchmark scores keep it at featured, not p1.

Apr 15Wednesday

X · @AnthropicAI

New Anthropic Fellows research: developing an Automated Alignment Researcher

Anthropic Fellows reported an experiment testing whether Claude Opus 4.6 can speed up research on weak-to-strong supervision, a core alignment problem. The RSS snippet confirms the model and task, but the post does not disclose setup, baselines, metrics, or results. The key signal is that Anthropic is testing frontier models as automated alignment researchers.

Why it matters: A credible Anthropic-source research teaser plus a novel safety angle clears HKR-H and HKR-R. HKR-K fails because the post discloses the direction and model only; setup, baselines, metrics, and results are not disclosed, so this sits near the featured threshold.

Apr 10Friday

X · @claudeai

We're bringing the advisor strategy to the Claude Platform.

Claude is adding the advisor strategy to Claude Platform, with Opus as the advisor and Sonnet or Haiku as the executor. The RSS snippet says this yields near-Opus-level agent intelligence at lower cost; the post does not disclose pricing, benchmark scores, or rollout timing.

Why it matters: Anthropic ships a substantive Claude Platform update, and HKR-H/K/R all pass: the Opus-advisor plus Sonnet/Haiku-executor setup is novel, concrete, and directly relevant to agent builders. The score stays below P1 because price, benchmarks, and rollout timing are not disclosed.

Mar 17Tuesday

Mistral AI

Mistral releases Mistral Small 4, unifying reasoning, multimodal and coding

Mistral AI released Mistral Small 4, the first Mistral model to unify Magistral reasoning, Pixtral multimodal and Devstral coding-agent abilities in a single model. It ships under the Apache 2.0 license.

Why it matters: Merging reasoning, multimodal and coding agents into one open model is a direct test of what unified models do to deployment cost.

Mar 12Thursday

NVIDIA Blog

NVIDIA Nemotron 3 Super delivers 5x higher throughput for agentic AI

NVIDIA launched Nemotron 3 Super, a 120B open model with 12B active parameters, and says it delivers up to 5x higher throughput for agentic AI. It has a 1M-token context window and uses hybrid MoE, latent MoE, and multi-token prediction; the post says Blackwell NVFP4 gives up to 4x faster inference than Hopper FP8, with over 10T training tokens disclosed. What matters is that NVIDIA is releasing open weights, training recipes, and RL environments for reproduction and fine-tuning.

Why it matters: This is a solid model-release story with all three HKR signals, led by strong HKR-K: parameter counts, active params, context length, training scale, and Blackwell/Hopper comparison are all concrete. It stays below 85 because the key performance claims come from NVIDIA's own blog

Mar 10Tuesday

OpenAI News

New ways to learn math and science in ChatGPT

OpenAI launched interactive math and science visualizations in ChatGPT on March 10, 2026, covering 70+ core concepts and rolling out globally across all plans. Users can adjust variables, manipulate formulas, and see graphs update in real time; OpenAI says 140 million people use ChatGPT weekly for math and science learning. The key point is productized interactivity, while the post does not disclose the underlying model, evaluation method, or outcome data.

Why it matters: HKR-H lands on the interactive-visual hook, HKR-K on 140M weekly learners plus 70+ concepts and live manipulation, and HKR-R on the product and edtech nerve. It is still a mid-weight product update; model details and learning-outcome evaluation are not disclosed, so it stays in a

Mar 5Thursday

OpenAI News

Reasoning models struggle to control their chains of thought, and that’s good

OpenAI frames an article around the claim that reasoning models struggle to control their chains of thought, and that this is a good thing. Only the title is available here, with no body text, so there are no verifiable numbers, methods, or mechanisms to summarize. The claim relates to reasoning and safety discussions, but any interpretation should stay limited to the headline.

Why it matters: OpenAI presents a contrarian but testable safety claim, so HKR-H/K/R all pass. The excerpt shows the thesis, section headers, and paper link, but not the key numbers, setup, or limits, so this stays high featured rather than P1.

OpenAI News

GPT-5.4 Thinking System Card

OpenAI published the GPT-5.4 Thinking System Card on March 5, 2026 and says it is the latest GPT-5 reasoning model and the first general-purpose model with mitigations for high-capability cybersecurity. The post confirms the safety approach follows prior GPT-5 models and builds on measures used for GPT-5.3 Codex, but it does not disclose benchmark scores, mitigation details, or deployment conditions. The key signal is the risk threshold change: OpenAI has extended high-cyber mitigations to a general reasoning model.

Why it matters: This clears HKR-H/K/R: a new GPT-5 reasoning model and the first general-purpose model with high-capability cyber mitigations. It stays below p1 because the disclosed text does not provide eval scores, mitigation details, or deployment conditions.

OpenAI News

Introducing ChatGPT for Excel and new financial data integrations

OpenAI launched ChatGPT for Excel beta on March 5, 2026, bringing GPT-5.4 into Excel workbooks and finance workflows. The post says it can build and update models, trace changes to cells, and is off by default for Enterprise and Edu admins; OpenAI's internal banking benchmark rose from 43.7% with GPT-5 to 87.3% with GPT-5.4 Thinking. The key move is data access: Moody’s, Dow Jones Factiva, MSCI, Third Bridge, and MT Newswires are live, while FactSet is listed as coming soon.

Why it matters: This is more than a routine add-on: OpenAI puts ChatGPT into Excel, names major finance data feeds, and cites a 43.7%→87.3% internal banking benchmark gain. HKR-H/K/R all pass; importance lands at 82 because this is a strong vertical workflow move, not a market-wide model release

Mar 3Tuesday

OpenAI News

GPT-5.3 Instant: Smoother, more useful everyday conversations

OpenAI released GPT-5.3 Instant on March 3, 2026 as an update to ChatGPT’s most-used model, aiming for fewer unnecessary refusals, fewer disclaimers, and more accurate everyday answers. The post shows one concrete contrast: GPT-5.2 Instant refused long-range archery trajectory help, while GPT-5.3 Instant requested parameters and gave a no-drag example at 300 fps (about 91 m/s), 45°, and 845 m; the key issue is the safety-boundary shift, while the post does not disclose benchmark scores, system card details, or API pricing.

Why it matters: OpenAI updated a core ChatGPT everyday model, and the story clears HKR-H/K/R because the refusal-boundary shift is concrete and widely relevant. The post includes a specific 5.2 vs 5.3 behavior example, but no system card, benchmark table, or API pricing, so it lands below the 85

Feb 26Thursday

OpenAI News

Pacific Northwest National Laboratory and OpenAI partner to accelerate federal permitting

OpenAI and Pacific Northwest National Laboratory evaluated coding agents on NEPA drafting tasks from 18 federal agencies, finding 1-5 hours saved per subsection, or about 15% less drafting time. The DraftNEPABench benchmark was designed with 19 experts and covers 102 tasks, using Codex CLI with GPT-5 for long-document synthesis, cross-checking, and structured writing. The key limit is explicit: this measures well-scoped drafting work, not full real-world permitting decisions.

Why it matters: HKR-H/K/R pass: federal permitting is an unusual hook; the post gives 19 experts, 102 tasks, and 1–5 hours saved; the debate is agents entering regulated workflows. Score stays below major product news because this is a scoped benchmark, not a shipped capability.

Jan 6Tuesday

NVIDIA Blog

NVIDIA presents Rubin platform, open models and autonomous driving roadmap at CES

At CES 2026, NVIDIA said its six-chip Rubin AI platform is now in full production and cuts token generation cost to about one-tenth of the prior platform. The post cites 50 petaflops NVFP4 inference for Rubin GPUs, 5x gains from its KV-cache storage tier, and the new open autonomous-driving model family Alpamayo; the key signal is production status and cost curve, not the “AI everywhere” framing.

Why it matters: HKR-H lands because Rubin is in production, not just on a roadmap. HKR-K is strong with ~1/10 token cost, 50 PFLOPS NVFP4, and 5x long-context throughput; HKR-R lands because NVIDIA still sets the tone on inference economics, though the company-blog framing keeps it below 90.

NVIDIA Blog

NVIDIA DGX SuperPOD Sets the Stage for Rubin-Based Systems

NVIDIA introduced Rubin-based DGX SuperPOD systems, with DGX Vera Rubin NVL72 and DGX Rubin NVL8 slated for the second half of this year. One DGX SuperPOD can combine eight NVL72 systems for 576 Rubin GPUs, 28.8 exaflops FP4, and 600TB memory; NVIDIA says inference token cost drops by up to 10x versus the prior generation. The key detail is rack-scale design: 260TB/s NVLink per rack, which the post says removes model partitioning.

Why it matters: This is a substantive NVIDIA infra roadmap with hard numbers: 576 Rubin GPUs, 28.8 exaflops FP4, 600TB memory, 260TB/s NVLink, and up to 10x lower token cost. HKR-H/K/R all pass, but it is still a vendor roadmap post rather than a shipping model or broad product release, so it is

Sep 2, 2025Tuesday

OpenAI News

Building more helpful ChatGPT experiences for everyone

OpenAI said it will ship ChatGPT safety changes over the next 120 days and roll out Parental Controls within a month. Disclosed steps include routing conversations with signs of acute distress to reasoning models such as GPT-5-thinking, and letting parents link accounts for teens 13+, disable memory and chat history. The post does not disclose router trigger thresholds or alert false-positive rates.

Why it matters: This changes core ChatGPT behavior, so HKR-H/K/R all pass: the routing hook is novel, the post gives concrete controls, and teen safety is a live industry topic. I keep it below 85 because trigger criteria, false-positive rate, and rollout scope are not disclosed.

Aug 7, 2025Thursday

OpenAI News

GPT-5 and the new era of work

OpenAI launched GPT-5 on August 7, 2025, started rollout to Team users the same day, said Enterprise and Edu access would follow next week, and made it available in the API immediately. The post gives two hard numbers: 5 million paid ChatGPT business users and nearly 700 million weekly ChatGPT users; it does not disclose benchmark scores, pricing, or context length.

Why it matters: An OpenAI GPT-5 launch is a market-wide event, so HKR-H/K/R all pass. The post gives rollout timing and a 5M paid-business-user datapoint, but it omits benchmark scores, pricing, and context length, so this lands at the low end of the top band.