Skip to content

#推理

1 today

Apr 24Friday

QbitAI · WeChat

Claude admits three issues: downgraded reasoning, cleared memory, and constrained output

Anthropic said on April 23 that three Claude issues hurt quality: Claude Code default reasoning was changed from high to medium on March 4 while the UI still showed high. A March 26 cache bug cleared thinking state every turn for 15 days, and an April 16 prompt limit of 25 words between tool calls and 100 words in final replies cut Opus 4.6/4.7 by 3% before a rollback four days later.

Why it matters: This is an Anthropic postmortem on Claude regressions, not generic complaint content. HKR-H/K/R all land: strong hook, three dated and testable facts, and a direct hit on transparency, billing, and silent-downgrade nerves; still below a major model launch, so 82.

X · @dotey

DeepSeek releases and open-sources V4 preview; 1M context is standard across all services

DeepSeek released and open-sourced the V4 preview, making 1M context standard across all official services with no tier or price split. The post says V4-Pro and V4-Flash use token compression plus DSA sparse attention to cut compute and memory costs for 1M context; legacy APIs remain for 3 months and stop after July 24.

Why it matters: DeepSeek is a flagship Chinese model vendor, and this V4 preview is a substantive release with open source and 1M context made standard across official services. HKR-H/K/R all pass: the post includes mechanisms and a migration deadline, and the tier reset makes it a same-day P1.

X · @dotey

OpenAI launches GPT-5.5 for paid ChatGPT and enterprise users, with Codex; API coming soon

OpenAI launched GPT-5.5 for ChatGPT Plus, Pro, Business, and Enterprise users, alongside Codex. OpenAI says per-token latency matches GPT-5.4, while Terminal-Bench 2.0 rises to 82.7% from 75.1%; API pricing is $5 per 1M input tokens and $30 per 1M output tokens with a 1M-token context. The key detail is efficiency: the post says GPT-5.5 uses about half the total tokens of frontier rival coding models at the same intelligence level.

Why it matters: This is a core OpenAI model release with benchmark, pricing, and 1M-context details, so HKR-H/K/R all pass. The title says the API is “coming soon” while the summary lists API pricing; that mismatch trims confidence slightly, but it still belongs in the must-write p1 band.

X · @OpenAI

Introducing GPT-5.5

OpenAI introduced GPT-5.5, and it is now available in ChatGPT and Codex. The RSS snippet says it targets real work and agents, can understand complex goals, use tools, check its work, and carry more tasks to completion; the post does not disclose parameters, pricing, context window, or benchmark results. What matters is the execution loop, not the headline's “new class of intelligence.”

Why it matters: OpenAI launching GPT-5.5 in ChatGPT and Codex is same-day mandatory coverage. HKR-H/K/R all pass: new model release, concrete agent-workflow claims, and direct impact on daily AI work. Price, context window, params, and benchmarks are undisclosed, so it stays below 95.

Apr 23Thursday

OpenAI News

Introducing GPT-5.5

OpenAI introduced GPT-5.5 and says it targets complex cross-tool tasks such as coding, research, and data analysis. The RSS snippet only confirms “faster” and “more capable”; the post does not disclose benchmarks, context window, pricing, release timing, or availability, which are the details practitioners should watch.

Why it matters: An OpenAI flagship-model release is same-day news, so HKR-H and HKR-R are clear. HKR-K fails because the post discloses the name and use cases but not benchmarks, context window, price, or availability, so this stays featured rather than p1.

X · @dotey

Microsoft makes Copilot Agent Mode the default in Word, Excel, and PowerPoint

Microsoft made Copilot Agent Mode the default experience in Word, Excel, and PowerPoint, available now to Microsoft 365 Copilot and Premium subscribers, including Personal and Family plans. Microsoft reports internal test gains: Excel engagement up 67% and approval up 65%, Word engagement up 52%, and PowerPoint new-user retention up 36%. What matters is that multi-step in-document execution is now the default, with preview and rollback controls.

Why it matters: This clears HKR-H/K/R: the default flip is the hook, the post includes rollout scope, concrete pilot metrics, and preview/keep/revert controls, and the Office distribution angle will get discussed. I keep it at 84 because this is a large product update, not a new model release or

Apr 22Wednesday

r/LocalLLaMA

ServiceNow-AI/SuperApriel-15B-Instruct · Hugging Face

ServiceNow released SuperApriel-15B-Instruct, a single-checkpoint 15B model with 8 deployment presets spanning 1.0× to 10.7× decode throughput at 32K sequence length. It has 48 decoder layers with 4 mixer variants per layer and up to 262K context positions depending on runtime; the key point is that speed-quality tradeoffs and speculative decoding are exposed from the same weights.

Why it matters: A single checkpoint spanning 8 deployment presets with 1.0x-10.7x decode throughput gives strong HKR-H and HKR-K, and the serving tradeoff gives HKR-R. The blast radius is narrower: this is a 15B inference-focused release, not a frontier-lab flagship update, so 76 and featured.

Synced · WeChat

Transformer can be converted into Mamba: Apple uses cross-architecture distillation to make inference cost linear

Apple presents a two-stage cross-architecture distillation path that converts Pythia-1B Transformer into a 1B HedgeMamba, reaching 14.11 perplexity with 10B tokens, about 2.7% of the teacher data. The teacher scores 13.86 PPL, while direct Transformer-to-Mamba distillation jumps above 100; the method first aligns with Hedgehog linear attention, then maps into Mamba initialization and fine-tunes. The key point is the path, not one trick: long-context inference shifts from quadratic to linear cost, and the post says downstream results on ARC, PIQA, BoolQ, RACE, and LogiQA approach the teacher.

Apr 21Tuesday

Synced · WeChat

Monet: Enabling multimodal LLMs to reason in latent visual space

Monet trains Qwen2.5-VL-7B into Monet-7B to reason with continuous latent visual embeddings instead of external tools; the work is accepted by CVPR 2026 and releases paper, code, model, and a 125K SFT dataset. The method uses three-stage SFT plus VLPO reinforcement learning; the post reports 3% to 9.75% gains on in-distribution tasks and 2.31% on out-of-distribution abstract visual reasoning versus the base model. The key detail is the VLPO mechanism and dataset construction; the post does not disclose one unified table of absolute headline scores.

Why it matters: This hits HKR-H and HKR-K: the angle is abstract visual reasoning, and the post includes 125K SFT data, a 3-stage SFT setup, VLPO, and 3%–9.75% / 2.31% gains. HKR-R is weaker because full absolute leaderboard scores and real deployment evidence are not disclosed, so it lands as a

OpenAI News

Introducing ChatGPT Images 2.0

OpenAI introduced ChatGPT Images 2.0 as a new image generation model, highlighting better text rendering, multilingual support, and visual reasoning. The RSS snippet names only these three upgrades; the post does not disclose architecture, resolution, pricing, latency, or availability. What matters is whether text fidelity and multilingual consistency improve in real use; for now, only headline-level details are disclosed.

Why it matters: A primary-source OpenAI image update clears HKR-H and HKR-R: the 2.0 label and text-rendering claim hit real workflows. HKR-K is weak because the post discloses only three upgrade areas; resolution, price, latency, architecture, and rollout are absent, so it stays just above the

Apr 20Monday

r/LocalLLaMA

Compared some models for feature planning

A Reddit user tested 9 models on planning a “load tracking” feature for a Go budgeting app, then used Claude Code to rank the generated specs, with Claude Opus 4.6 placed first. The table shows Opus 4.6 produced a 19 KB spec with 44 code reads at $2.47; GLM 5.1 ranked second and Qwen 3.6 35B fp8+vLLM ranked third. Do not treat this as a benchmark: the author says it is not representative, and the post does not disclose any manual quality review yet.

Why it matters: A named first-person test gives real workflow data, so HKR-H/K/R all pass. The ceiling stays low: one task only, ranked by Claude Code itself, and no human acceptance result is disclosed, so this lands at the low end of featured.

Apr 18Saturday

QbitAI · WeChat

RAG retrieves the right docs but still answers wrong? Saarland University team diagnoses why | ACL 2026

A Saarland University-led team introduced Disco-RAG, adding a 3-step “reading” layer between retrieval and generation, and says the paper was accepted as an ACL 2026 main-conference long paper. The post says it uses RST-based argument trees, cross-passage relation graphs, and outline generation with zero training; it reports gains on Loong, ASQA, and SciNews, but does not fully disclose the exact scores. The key claim is that many RAG failures come from reading and discourse understanding, not retrieval recall.

Why it matters: This is a solid research release with HKR-H, HKR-K, and HKR-R: a strong practical hook, a concrete mechanism, and a pain point RAG builders know well. I keep it at 80, not higher, because the post does not fully disclose benchmark numbers and external replication is still missing

Synced · WeChat

What is OpenAI prioritizing under compute limits?

Greg Brockman said OpenAI narrowed priorities under hard compute limits to two bets: a personal assistant and AI workers that solve hard user problems, and current compute cannot fully support both. The snippet says Sora resources were reduced while focus shifted to reasoning models, a unified AI layer, and the next base model Spud; it does not disclose the claimed compute budget, timeline, or model specs. The key point is not a B2B retreat but a compute-driven reprioritization.

Why it matters: HKR-H/K/R all pass: the compute-ceiling angle is strong, the piece adds concrete priority shifts, and OpenAI roadmap triage hits cost and dependency nerves. It stays at 80 because this is secondary reporting; spend, timing, and technical details are not disclosed.

Xinzhiyuan · WeChat

Claude Opus 4.7 splits users 48 hours after launch: benchmark lead, reasoning tests drop

Anthropic's Claude Opus 4.7 drew split reactions within 48 hours: Artificial Analysis scored it at 57, tied for No.1, while NYT Connections Extended fell from 94.7% on 4.6 to 41.0%. The post says a new tokenizer raises token usage to 1.0-1.35x on the same text, and old thinking parameters can return 400 errors; Anthropic also cites a 1753 Elo GDPval-AA score, 79 points above No.2. The real issue is migration cost and capability trade-offs, not a single leaderboard.

Why it matters: The signal is not the “backlash” framing but the four concrete shifts: benchmark lead, reasoning drop, higher token use, and API breakage. HKR-H/K/R all land, but this is secondary analysis 48 hours after launch, not the primary Anthropic release, so it stays below p1.

Financial Times · Technology

Months-old start-up Recursive raises $500mn for self-teaching AI

Recursive raised $500mn, and the headline says the company is building “self-teaching AI.” The body is empty, so beyond the firm being months old and the $500mn amount, the post does not disclose investors, valuation, or technical method. Those missing details matter more than the label.

Why it matters: This clears HKR-H, HKR-K, and HKR-R on one strong fact: a months-old AI startup raised $500mn. The score stays near the featured floor because the body does not disclose investors, valuation, or the mechanism behind the 'self-teaching AI' claim.

Apr 17Friday

Latent Space

[AINews] Anthropic Claude Opus 4.7 - one step better than 4.6 in every dimension

Anthropic launched Claude Opus 4.7 at the same $5/$25 per million input/output tokens; the post says 4.7-low through 4.7-high each outperform the matching higher 4.6 tiers. Reported changes include a new xhigh reasoning tier, Claude Code defaulting to xhigh, an 11-point gain on SWE-Bench Pro, and image input up to 2,576 px on the long edge (~3.75 MP). Do not overread the tokenizer change: the same input can use up to 35% more tokens, but the post says total token use still falls by up to 50% from prior equivalents.

Why it matters: Anthropic's flagship-model release fits the policy's 85–94 band. HKR-H/K/R all pass because the post gives concrete pricing, benchmark, image-limit, and token-accounting changes that hit Claude users' core coding and cost concerns.

X · @OpenAI

Introducing GPT-Rosalind, OpenAI's frontier reasoning model for biology, drug discovery, and translational medicine

OpenAI introduced GPT-Rosalind as a reasoning model for biology, drug discovery, and translational medicine research. The title and snippet disclose its intended domains; the post does not disclose size, benchmarks, availability, pricing, or launch timing. The key point is research targeting, but reproducible details are absent so far.

Why it matters: An official OpenAI announcement plus the unusual biology/drug-discovery positioning gives this HKR-H and HKR-R. HKR-K is weak because the post discloses only the model name and target domains; benchmarks, params, access, and launch timing are not disclosed, so it stays at the low

X · @dotey

Claude Opus 4.7 uses more thinking tokens, so Anthropic permanently raised rate limits for paid users

Anthropic permanently raised rate limits for all paid subscribers because Claude Opus 4.7 uses more thinking tokens than its predecessor. The post confirms the affected group but does not disclose the increase size, pricing rules, or rollout timing; users who do not see the change should verify they are on Opus 4.7 and have updated Claude Code.

Why it matters: This is a substantive Anthropic quota update for paying users, with all three HKR axes present: a strong surprise hook, a concrete operational fact, and direct resonance on usage limits. It stays at featured, not P1, because the post does not disclose the size of the increase, pr

X · @dotey

Official best practices for using Claude Opus 4.7 with Claude Code

Anthropic shared guidance for Claude Opus 4.7 in Claude Code: the default Effort level is now xhigh, and users should provide goals, constraints, and acceptance criteria upfront. The post lists five Effort tiers—low, medium, high, xhigh, and max—with xhigh recommended for most coding, API design, migration, and code review tasks. The key shift is behavior: adaptive thinking is built in, while tool use and SubAgent spawning are less frequent by default, so prompts should state those needs explicitly.

Why it matters: This is not a model launch, but an official Anthropic workflow note that changes day-to-day Claude Code usage: default effort, 5 levels, and fewer tool/SubAgent calls unless asked. HKR-H/K/R all pass, but the scope is narrower than a major product release.

Apr 16Thursday

X · @op7418

Claude Code now supports Claude Opus 4.7

Claude Code now supports Claude Opus 4.7, and the RSS snippet confirms X-HIGH as the default reasoning level. The only concrete detail disclosed is that users must switch manually to Max if X-HIGH is insufficient. The post does not disclose pricing, rate limits, or launch timing.

Why it matters: This is a substantive Claude product update with all three HKR signals: a new model in Claude Code plus one concrete operational detail, X-HIGH vs. manual Max. I kept it below the top band because price, rate limits, and formal release timing are not disclosed.

Hacker News front page

Claude Opus 4.7 System Card

Anthropic published a 232-page system card for Claude Opus 4.7 on April 16, 2026, saying it outperforms Opus 4.6 but remains below the limited-release Claude Mythos Preview. The card says Opus 4.7 does not advance Anthropic’s capability frontier, catastrophic risk remains low, cyber capability is roughly similar to Opus 4.6, and it does not cross the threshold for automated AI R&D. The excerpt does not disclose benchmark scores or the new cybersecurity safeguard details.

Why it matters: This is not a flashy launch post, but it is a substantive Anthropic system card update. HKR-K is strong: Opus 4.7 beats 4.6, stays below automated AI R&D thresholds, and is roughly similar to 4.6 on cyber evals; HKR-R lands because Claude users track general-access model ceilings

X · @claudeai

Introducing Claude Opus 4.7, our most capable Opus model yet.

Claude introduced Opus 4.7 and describes it as its most capable Opus model so far. The RSS snippet gives three claims: better rigor on long-running tasks, more precise instruction following, and self-verification before replying; the post does not disclose benchmarks, context window, pricing, or rollout scope. What matters is whether those claims show up in public evals, not the tagline.

Why it matters: This is a substantive Anthropic model release and clears HKR-H/K/R: a new Opus, three testable behavior claims, and strong resonance with Claude-heavy practitioners. The score stays in the high 80s because benchmarks, pricing, context window, and rollout scope are not disclosed.

r/LocalLLaMA

Qwen3.6-35B-A3B released

Qwen released Qwen3.6-35B-A3B as open source under Apache 2.0; it is a sparse MoE with 35B total parameters and 3B active. The post also claims agentic coding, strong multimodal perception and reasoning, plus thinking and non-thinking modes; the post does not disclose benchmarks, context length, or latency.

Why it matters: HKR-H/K/R all pass: a new open Qwen model is timely, and the post confirms 35B total, 3B active, and Apache 2.0. The score stays at 82 because this is still a launch post; benchmarks, context window, latency, and multimodal details are not disclosed here.

Hacker News front page

AI cybersecurity is not proof of work

antirez argues AI bug finding is bounded by model intelligence level I, not by brute-force sampling alone; for the same code, execution paths eventually saturate. His concrete example is the OpenBSD SACK bug: weaker models fail even with unlimited tokens because they do not connect window validation, integer overflow, and the NULL branch. The key variable is model quality and access speed, not just more GPU.

Why it matters: High-quality commentary with HKR-H from the contrarian headline, HKR-K from the OpenBSD SACK mechanism and firsthand test, and HKR-R because it hits the 'more sampling vs better models' debate in AI security. Not a product, research release, or multi-source event, so it stays mid

OpenAI News

Introducing GPT-Rosalind for life sciences research

OpenAI released GPT-Rosalind on April 16, 2026, and made it available as a research preview in ChatGPT, Codex, and the API for qualified customers. The post says it targets biology, drug discovery, and translational medicine, and adds a free Codex life sciences plugin connecting to 50+ scientific tools and data sources. The real signal is deployment breadth: Amgen, Moderna, and Thermo Fisher Scientific are involved, but the post does not disclose model size, pricing, or benchmark scores.

Why it matters: HKR-H lands because OpenAI is shipping a vertical life-sciences model; HKR-K lands on access paths and the 50+ tool/data plugin. HKR-R also lands on the domain-model debate, but missing params, pricing, and benchmark scores keep it at featured, not p1.

最佳拍档 (BestPartners)

Post-AGI may arrive within 50 years: Demis Hassabis on AlphaFold, three AI risk classes, and human value

Demis Hassabis said in a 1-hour interview that post-AGI scenarios can arrive within 50 years, while AGI should stay in labs for another 10-20 years. He cited concrete numbers: AlphaFold has been used by 3M+ scientists, Isomorphic Labs is running 18-19 drug programs, and the most urgent risks in the next 2-4 years are misuse and agent misalignment.

Apr 15Wednesday

X · @AnthropicAI

New Anthropic Fellows research: developing an Automated Alignment Researcher

Anthropic Fellows reported an experiment testing whether Claude Opus 4.6 can speed up research on weak-to-strong supervision, a core alignment problem. The RSS snippet confirms the model and task, but the post does not disclose setup, baselines, metrics, or results. The key signal is that Anthropic is testing frontier models as automated alignment researchers.

Why it matters: A credible Anthropic-source research teaser plus a novel safety angle clears HKR-H and HKR-R. HKR-K fails because the post discloses the direction and model only; setup, baselines, metrics, and results are not disclosed, so this sits near the featured threshold.

Apr 12Sunday

最佳拍档 (BestPartners)

Breaking RLHF scaling bottlenecks: DeepMind raises data efficiency 10x with information-directed exploration

A Google DeepMind team reports that online RLHF plus information-directed exploration on Gemma 9B reaches about 55% win rate with under 20k preference labels, versus about 200k for offline RLHF. The post describes four algorithms—offline, periodic, online, and information-directed exploration; online training uses batches of 64 prompts and 16 sampled responses per prompt, while the ENN head adds under 5% parameters. The key point is methodological, not that RLHF failed; the post also says results use Gemini 1.5 Pro simulated feedback, and the 1000x gain is an extrapolation toward 1M labels.

Why it matters: HKR-H/K/R all pass: the 10x label-efficiency claim is a strong hook, and the post includes concrete setup details. I kept it at 77 because this is a secondary video summary, feedback is simulated with Gemini 1.5 Pro, and the 1000x figure is an extrapolation.

Apr 11Saturday

QbitAI · WeChat

Liu Zhuang and Danqi Chen team open-source Vero, a general visual reasoning RL framework, reaching SOTA with zero thinking data

Princeton researchers including Liu Zhuang and Danqi Chen open-sourced Vero, an RL framework for visual reasoning, and report beating Qwen3-VL-8B-Thinking on 23 of 30 benchmarks. The post says Vero uses 600K samples filtered from 59 datasets, task-routed rewards, and single-stage RL across six task groups. The key point is the mechanism mix: no private thinking data, but the post does not disclose training cost or base model configuration.

Why it matters: Featured on HKR-H/K/R: the zero-thinking-data claim is a strong hook, and the post includes concrete benchmark and method details. I keep it in the low 80s because training cost, base model choice, and full reproduction conditions are not disclosed.

Apr 10Friday

X · @claudeai

We're bringing the advisor strategy to the Claude Platform.

Claude is adding the advisor strategy to Claude Platform, with Opus as the advisor and Sonnet or Haiku as the executor. The RSS snippet says this yields near-Opus-level agent intelligence at lower cost; the post does not disclose pricing, benchmark scores, or rollout timing.

Why it matters: Anthropic ships a substantive Claude Platform update, and HKR-H/K/R all pass: the Opus-advisor plus Sonnet/Haiku-executor setup is novel, concrete, and directly relevant to agent builders. The score stays below P1 because price, benchmarks, and rollout timing are not disclosed.

Apr 9Thursday

X · @op7418

Meta releases Muse Spark model

Meta released the Muse Spark model with native multimodal reasoning, tool use, visual chain-of-thought, and multi-agent orchestration, but it is only available in the Meta AI app and is not open source for now. The snippet says its Contemplating mode coordinates multiple parallel agents for reasoning, and its Artificial Analysis score is below Gemini 3.1 Pro, GPT-5.4, and Claude Opus 4.6. The post does not disclose model size, pricing, or rollout timing.

Why it matters: A major-lab model launch plus the “poached team’s first output” angle lands HKR-H/K/R. The score stays near the featured floor because the post offers capability claims and relative benchmark placement only; params, pricing, rollout timing, and access scope are not disclosed.

Apr 4Saturday

Latent Space

Marc Andreessen introspects on The Death of the Browser, Pi + OpenClaw, and Why “This Time Is Different”

Marc Andreessen argues in a 76-minute interview that this AI cycle differs from 2016 because of reasoning, coding, agents, and recursive self-improvement. The post gives one concrete mechanism: Pi/OpenClaw as LLM + shell + filesystem + markdown + cron loop; it mentions “death of the browser,” but does not disclose a verifiable timeline or product plan. The sharper point is his Unix-like framing of file-backed agent state and portability.

Why it matters: This is a strong commentary piece, not a market-moving event. HKR-H comes from the browser-death hook, HKR-K from the Pi+OpenClaw mechanism, and HKR-R from the interface/distribution nerve; lack of roadmap, metrics, or launch details keeps it at the low end of featured.

Mar 31Tuesday

MIT Technology Review · AI

There are more AI health tools than ever—but how well do they work?

Microsoft launched Copilot Health this month, and Amazon expanded Health AI beyond One Medical; the piece also cites OpenAI’s ChatGPT Health and Anthropic’s Claude, showing consumer health chatbots are becoming a trend. Microsoft says Copilot gets 50 million health questions per day, but all six academics interviewed raised safety concerns over the lack of independent evaluation; the post cites a Mount Sinai study saying ChatGPT Health can over-recommend care for mild cases and miss emergencies. The key issue is external validation, not vendor-run benchmarks.

Why it matters: Strong HKR-K and HKR-R: it combines concrete scale, named critics, and Mount Sinai error modes around a high-risk AI vertical. HKR-H also lands through the 'more tools, but do they work?' tension, but this is trend reporting rather than a market-moving launch or breakthrough, so

Mar 20Friday

MIT Technology Review · AI

The Download: OpenAI is building a fully automated researcher, and a psychedelic trial blind spot

OpenAI says it plans to build an autonomous AI research intern by September 2026 for a small set of research problems, ahead of a multi-agent automated researcher targeted for 2028. The RSS snippet gives the timeline and staged plan, but the post does not disclose evals, compute budget, or research scope. The real question is whether the agent can produce verifiable research output.

Why it matters: HKR-H lands on the “fully automated researcher” hook, HKR-K on the two roadmap dates, and HKR-R on research-job substitution plus lab rivalry. It stays below must-write because the post does not disclose benchmarks, compute budget, or scope, so this is a strong roadmap signal, no

MIT Technology Review · AI

OpenAI is making a fully automated researcher its North Star

OpenAI made a “fully automated researcher” its multi-year North Star and plans an autonomous “AI research intern” by September for a small number of specific problems. The post says this roadmap combines reasoning, agents, and interpretability, with a multi-agent research system targeted for 2028; it does not disclose pricing, compute, or evaluation criteria. The real thing to watch is long-horizon execution and task decomposition, not the slogan.

Why it matters: This lands on HKR-H/K/R: the roadmap has a strong hook, new timelines, and a direct job-and-competition nerve. Kept at 84, not p1, because this is a reported strategy piece rather than a shipped product, and price, compute, and evals are not disclosed.

Mar 17Tuesday

Mistral AI

Mistral releases Mistral Small 4, unifying reasoning, multimodal and coding

Mistral AI released Mistral Small 4, the first Mistral model to unify Magistral reasoning, Pixtral multimodal and Devstral coding-agent abilities in a single model. It ships under the Apache 2.0 license.

Why it matters: Merging reasoning, multimodal and coding agents into one open model is a direct test of what unified models do to deployment cost.

Mar 13Friday

MIT Technology Review · AI

The Download: how AI is used for military targeting, and the Pentagon's war on Claude

A US Defense Department official said the military can feed target lists into a classified generative AI system to analyze and rank strike priority, with humans reviewing the output. The title also says the Pentagon CTO called Claude a risk to the defense supply chain because of a built-in “policy preference”; the post does not disclose the exact model, timeline, or control mechanism. The key point is that generative AI is entering high-stakes decision loops while audit details remain undisclosed.

Why it matters: HKR-H/K/R all land: the post links genAI directly to target-priority ranking and frames a Pentagon pushback against Claude over embedded policy preferences. Key facts—the model used, deployment timing, and audit controls—are not disclosed, so it stays in the low featured band.

Mar 12Thursday

NVIDIA Blog

NVIDIA Nemotron 3 Super delivers 5x higher throughput for agentic AI

NVIDIA launched Nemotron 3 Super, a 120B open model with 12B active parameters, and says it delivers up to 5x higher throughput for agentic AI. It has a 1M-token context window and uses hybrid MoE, latent MoE, and multi-token prediction; the post says Blackwell NVFP4 gives up to 4x faster inference than Hopper FP8, with over 10T training tokens disclosed. What matters is that NVIDIA is releasing open weights, training recipes, and RL environments for reproduction and fine-tuning.

Why it matters: This is a solid model-release story with all three HKR signals, led by strong HKR-K: parameter counts, active params, context length, training scale, and Blackwell/Hopper comparison are all concrete. It stays below 85 because the key performance claims come from NVIDIA's own blog

Mar 10Tuesday

OpenAI News

New ways to learn math and science in ChatGPT

OpenAI launched interactive math and science visualizations in ChatGPT on March 10, 2026, covering 70+ core concepts and rolling out globally across all plans. Users can adjust variables, manipulate formulas, and see graphs update in real time; OpenAI says 140 million people use ChatGPT weekly for math and science learning. The key point is productized interactivity, while the post does not disclose the underlying model, evaluation method, or outcome data.

Why it matters: HKR-H lands on the interactive-visual hook, HKR-K on 140M weekly learners plus 70+ concepts and live manipulation, and HKR-R on the product and edtech nerve. It is still a mid-weight product update; model details and learning-outcome evaluation are not disclosed, so it stays in a