Skip to content

#编码

10 today

Apr 22Wednesday

TechCrunch · AI

SpaceX is working with Cursor and has an option to buy the startup for $60B

SpaceX is working with Cursor and holds an option to acquire the startup for $60B. The RSS snippet discloses the collaboration and purchase option, but not the term, trigger conditions, ownership impact, or payment mix. The sharper signal is strategic weakness: the snippet says neither Cursor nor xAI has proprietary models matching leading offerings from Anthropic and OpenAI.

Why it matters: HKR-H lands on the surprise combo and the $60B option. HKR-K lands on the concrete number and deal structure; HKR-R lands because Cursor is a daily tool for AI builders. Missing term details keep it at the low end of p1, not higher.

Hacker News front page

Anthropic removes Claude Code from the $20/month Pro subscription for new users

Anthropic was reported to remove Claude Code from the $20/month Pro plan for new users, while saying existing Pro and Max subscribers are unaffected. The cited evidence: an April 10 archived help page said “Pro or Max plan,” the current page says “Max plan,” and Amol Avasare said this is a test on about 2% of new prosumer signups. The key issue is whether pricing shifts fully to Max or API billing; the post does not disclose retroactive scope or a final rollout timeline.

Why it matters: This clears all three HKR axes: the rollback is a strong hook, the post adds concrete evidence via help-page changes and a ~2% test, and it hits Claude users' cost and access concerns. Scope is still limited to new-user testing and no formal rollout timeline is disclosed, so it’s

The Verge · AI

SpaceX cuts a deal to maybe buy Cursor for $60 billion

SpaceX announced an either-or deal: buy AI coding platform Cursor for $60 billion or pay a $10 billion fee. The RSS snippet says this could help xAI's coding tools chase Anthropic; the post does not disclose the structure, timing, or IPO linkage. Watch the $10 billion breakup fee, not just the tentative acquisition headline.

Why it matters: All three HKR axes land: the headline has a strong unexpected hook, and the report gives two hard facts — a $60B price and a $10B breakup fee. I keep it at featured, not P1, because only top-line terms are disclosed; structure, timing, and the exact xAI linkage are still undiscol

Bloomberg Technology

SpaceX Has Deal for Right to Acquire Cursor for $60 Billion

SpaceX said it signed a deal giving it the right to acquire AI coding startup Cursor later this year for $60 billion. If it does not proceed, it can pay $10 billion for the companies' collaboration; the post does not disclose trigger terms, scope, or regulatory details.

Why it matters: Bloomberg reports an unusual, high-value deal: SpaceX gets the right to acquire Cursor for $60B, or pays $10B for cooperation if it does not proceed. HKR-H/K/R all pass; missing trigger, scope, and regulatory details keep it at the low end of the 85-94 band.

Bloomberg Technology

Anthropic’s Mythos Model Is Being Accessed by Unauthorized Users

A small group of unauthorized users accessed Anthropic’s new Mythos model, Bloomberg reported, citing a person familiar with the matter and reviewed documents. The snippet says Anthropic considers Mythos powerful enough to enable dangerous cyberattacks; the post does not disclose the user count, access path, time frame, or remediation. The real issue is access control failure, not a normal product launch.

Why it matters: This is a Bloomberg-reported Anthropic safety incident, not routine product news; HKR-H and HKR-R are strong because unauthorized access to a high-risk model is inherently clickable and discussable. HKR-K passes on the new access and risk facts, but user count, access path, and a

Apr 21Tuesday

QbitAI · WeChat

Mystery model Elephant: 100B parameters reaches same-scale SOTA with high token efficiency

Ant Group's Inclusion AI team is identified as the maker of Elephant, a 100B-parameter model with 256K context and 32K output shown on OpenRouter. The post reports tests on bug fixing, summarizing a 3,000-word meeting note, and a light agent loop, plus AI BENCHY figures of about 2,500 output tokens, about 1 second average latency, and 9.6/10 consistency; the post does not disclose training details, pricing, or an official model card.

Why it matters: HKR-H/K/R all pass: a 100B model posting same-scale SOTA with token efficiency is a strong hook, and the piece includes 256K/32K, ~1s latency, 9.6/10 consistency, plus failure cases. It stays below p1 because training details, pricing, and an official model card are not disclosed

Synced · WeChat

Sergey Brin revives founder mode? Google forms a strike team to focus on AI coding

Google has formed an AI coding strike team led by Sebastian Borgeaud, with Sergey Brin and Koray Kavukcuoglu directly involved, to improve long-context coding and internal code automation. The pressure signal cited is that Google said about 50% of its code is written by coding agents and reviewed by engineers, while Anthropic staff claimed 100% code use by Claude Code and Opus 4.5; the post does not disclose team size, launch timing, or the exact Google model version. The key issue is whether Google can turn private codebase training into stronger public models.

Why it matters: HKR-H/K/R all pass: the founder-return angle is clickable, and the piece includes Google's ~50% agent-written-code claim. It stays below p1 because no public launch is disclosed, and team size, timing, and model version are missing.

X · @dotey

GitHub paused new sign-ups for Copilot Pro, Pro+, and Student on April 20

GitHub paused new sign-ups for Copilot Pro, Pro+, and Student on April 20, leaving only Copilot Free open to new users. The post says Pro+ now has more than 5x Pro usage, Claude Opus 4.7 is limited to Pro+, and users can request cancellation with a full April refund from Apr 20 to May 20. What matters is the price stayed fixed while access, quotas, and model tiers tightened first.

Why it matters: This is not a capability launch; it is a meaningful Copilot packaging clampdown on entry, quotas, and model access. HKR-H/K/R all pass on the unexpected restrictions, concrete tier changes, and direct developer impact, but the source is an X post rather than a primary GitHub note

Latent Space

Moonshot Kimi K2.6 open-weight model refresh aims to catch Opus 4.6

Moonshot released Kimi K2.6, a 1T-parameter MoE with 32B active and 256K context. The post cites 58.6 on SWE-Bench Pro, 4,000+ tool calls, 12+ hour runs, and 300 parallel sub-agents. The key signal is long-horizon agent execution, not only open-model scores.

Why it matters: HKR-H/K/R all pass: Kimi K2.6 has a strong race narrative, concrete model and agent metrics, and direct relevance to open-model builders. The domestic flagship release signal lifts it into P1.

Computing Life · Share · Yage

Musk wants Cursor: a $60B acquisition option, a $10B partnership, and the rise of acqui-hire deals

The article says SpaceX offered Cursor two paths: buy Anysphere for $60B in 2026 or pay $10B for a tech partnership. The $60B figure is about 20% above Cursor's reported $50B fundraising valuation, but the post does not disclose payment terms; it also says Cursor uses xAI's Colossus to train Composer 2.5. The real signal is acqui-hire risk: xAI already hired two Cursor engineering leaders in March, so a $10B partnership would not guarantee a broad employee exit.

Why it matters: HKR-H/K/R all pass: the 60B buyout vs 10B partnership frame is a strong hook, and the piece includes concrete facts—20% over Cursor's 50B talks, Colossus training Composer 2.5, and two leaders already at xAI. Not P1 because payment terms and the final path are still undisclosed.

X · @dotey

OpenAI adds Chronicle to Codex, letting it read screen context

OpenAI added Chronicle to Codex and is rolling it out to ChatGPT Pro users on macOS; it uses periodic screenshots, OCR, and tool detection to turn recent screen activity into memory. The memory is stored as plain Markdown in ~/.codex/memories_extensions/chronicle, and the EU, UK, and Switzerland are excluded; OpenAI says screenshots are uploaded for processing, deleted afterward, and not used for training. The part to watch is risk: the background agent can burn rate limits, local plain-text files widen exposure, and OpenAI warns it amplifies prompt-injection from malicious webpages.

Why it matters: HKR-H/K/R all pass: the screen-watching memory angle is novel, and the post includes testable details like OCR, plaintext local storage, region limits, and deletion claims. The limited macOS ChatGPT Pro rollout keeps it in the 78–84 band rather than p1.

Hacker News front page

Expansion Artifacts

Matt Ström-Awn argues that flaws in LLM outputs are “expansion artifacts,” not compression artifacts, and cites 2024 evidence that they can be tracked. He notes Stanford researchers estimated AI-drafted text in 17.5% of recent CS papers and 16.9% of peer reviews from post-ChatGPT word-frequency shifts, and contrasts this with a JPG after 10,000 recompressions reaching PSNR 14.59. The point for practitioners is forensic: these artifacts expose both model aesthetics and generation provenance.

Why it matters: HKR-H lands on the “expansion artifacts” hook; HKR-K adds concrete numbers and a testable provenance claim; HKR-R hits peer-review trust and detection anxiety. It stays at 73 because this is personal-blog commentary, not a primary research or product release event.

Apr 20Monday

X · @Yuchenj_UW

Kimi K2.6 is open-source

Kimi K2.6 is now open source, and the RSS snippet says it scored 58.6 on SWE-Bench Pro. The snippet also says it beat GPT-5.4 xhigh and Claude Opus 4.6 max effort. What matters is reproducibility; the post does not disclose weights, license, or eval setup.

Why it matters: All three HKR axes land: the open-source release is a strong hook, the SWE-Bench Pro 58.6 claim is testable, and the open-vs-closed coding race resonates. I keep it at 81 because the post appears title-level only; weights, license, and eval conditions are not disclosed.

r/LocalLLaMA

Compared some models for feature planning

A Reddit user tested 9 models on planning a “load tracking” feature for a Go budgeting app, then used Claude Code to rank the generated specs, with Claude Opus 4.6 placed first. The table shows Opus 4.6 produced a 19 KB spec with 44 code reads at $2.47; GLM 5.1 ranked second and Qwen 3.6 35B fp8+vLLM ranked third. Do not treat this as a benchmark: the author says it is not representative, and the post does not disclose any manual quality review yet.

Why it matters: A named first-person test gives real workflow data, so HKR-H/K/R all pass. The ceiling stays low: one task only, ranked by Claude Code itself, and no human acceptance result is disclosed, so this lands at the low end of featured.

Synced · WeChat

How to Do Vibe Coding Correctly? A Masterclass from Anthropic's Coding Agent Lead

Anthropic researcher Erik Schluntz said his team merged a 22,000-line production change, mostly written by Claude, cutting work from two weeks to one day. His workflow spends 15-20 minutes on repo exploration and planning, limits edits to leaf nodes, keeps humans on core logic, and validates with long stress tests plus a few E2E tests. The key issue is boundary control, not handing AI the system core; he also said task length AI can handle doubles about every seven months.

Why it matters: HKR-H/K/R all pass: this is an Anthropic field report with concrete numbers and reproducible workflow rules for production coding agents. It stays at featured, not p1, because it is a strong practitioner lesson rather than a major model or product launch.

Xinzhiyuan · WeChat

Agent isn’t the key: RUC's AiScientist shows 23 hours and 74 rounds of long-horizon memory

A Renmin University of China team released AiScientist, which ran 23 hours and 74 experiment loops on MLE-Bench Lite Detecting Insults, raising validation AUC from 0.903 to 0.982 with 18 best-so-far updates. The paper says its core is File-as-Bus, which persists analysis, code, logs, and results in the workspace; removing it drops PaperBench by 6.41 points and MLE-Bench Lite Any Medal by 31.82 points. The real lever here is state continuity, not simply adding more agents.

Why it matters: HKR-H lands because the title flips a live assumption: memory continuity, not more agents. HKR-K lands on the 23h/74-run setup, AUC 0.903→0.982, and ablations; HKR-R lands because builders are debating multi-agent stacks vs durable state.

r/LocalLLaMA

Using Qwen3.6 via LM Studio as a Claude Code subagent, saving 30x Opus tokens per task

A Reddit user routed Qwen3.6 through LM Studio as a Claude Code subagent and reported about 30x lower Opus marginal tokens on two audit tasks. In the examples, a 23-file route audit dropped from 13k to 0.4k marginal tokens, and an 18-file Astro site inventory fell from 89k to 3k; the setup used unsloth’s Qwen3.6-35B-A3B-MXFP4_MOE gguf on a 64GB M4 Max with a 64k context window. The key mechanism is offloading extraction and audit work to a local OpenAI-compatible server, while the post also says quality was mixed rather than strictly better than Opus.

Why it matters: A named first-person experiment with 2 clear token comparisons hits HKR-H, HKR-K, and HKR-R: strong hook, concrete setup details, and direct cost relevance for Claude Code users. It stays below p1 because the evidence is a Reddit post with only 2 tasks.

Apr 19Sunday

r/LocalLLaMA

Same 9B Qwen weights: 19.1% in Aider vs 45.6% with a scaffold adapted to small local models

Using the same Qwen3.5-9B Q4 weights on the 225-task Aider Polyglot benchmark, the author changed only the scaffold and raised mean pass@2 from 19.11% to 45.56%. The little-coder setup is not a new model; it uses bounded reasoning, a write guard, explicit workspace discovery, and small per-turn skill injections. The key claim is scaffold-model fit, but the post reports only two full runs and does not disclose ablations, cross-model replications, or a second benchmark.

Why it matters: HKR-H/K/R all pass: the hook is a 2.4x jump on Aider Polyglot 225 with the same 9B Qwen weights, and the post names the scaffold mechanisms. Importance stays low-featured because evidence is thin: two full runs, no ablation, no cross-model rerun, and no second benchmark.

Xinzhiyuan · WeChat

A Berkeley team built an AI that scores perfectly on SWE-bench while fixing 0 bugs

Berkeley RDI used a roughly 10-line conftest.py exploit to score 100% on all 500 SWE-bench tasks while fixing 0 bugs. The post says its agent broke 8 major agent benchmarks with scores from 73% to 100%, via pytest hook tampering, file:// answer reads, and faulty validators. The real issue is benchmark isolation failure, not stronger models.

Why it matters: HKR-H lands on the 'perfect score, zero fixes' contradiction; HKR-K lands on the ~10-line pytest exploit, 500 tasks, and 8-benchmark spread; HKR-R lands on eval-trust anxiety for agent builders. Strong featured research, but not a same-day industry event, so below P1.

Apr 18Saturday

Synced · WeChat

Claude Design enters research preview for generating mockups, prototypes, and slides

Anthropic launched Claude Design in research preview for Claude Pro, Max, Team, and Enterprise users, covering mockups, prototypes, slides, and one-pagers. Powered by Claude Opus 4.7, it can ingest codebases, images, DOCX, PPTX, XLSX, and web captures, then export to Canva, PDF, PPTX, and HTML; the headline cites Figma and Adobe stock drops, but the post does not disclose the moves. The real signal is the workflow link from design system ingestion to handoff into Claude Code.

Why it matters: HKR-H/K/R all pass: the design-workflow angle is novel, the post gives concrete mechanism details, and the Figma/Adobe pressure point resonates. I keep it below 85 because the stock-drop claim has no numbers and there is no user test, pricing, or adoption data.

Xinzhiyuan · WeChat

Claude Opus 4.7 splits users 48 hours after launch: benchmark lead, reasoning tests drop

Anthropic's Claude Opus 4.7 drew split reactions within 48 hours: Artificial Analysis scored it at 57, tied for No.1, while NYT Connections Extended fell from 94.7% on 4.6 to 41.0%. The post says a new tokenizer raises token usage to 1.0-1.35x on the same text, and old thinking parameters can return 400 errors; Anthropic also cites a 1753 Elo GDPval-AA score, 79 points above No.2. The real issue is migration cost and capability trade-offs, not a single leaderboard.

Why it matters: The signal is not the “backlash” framing but the four concrete shifts: benchmark lead, reasoning drop, higher token use, and API breakage. HKR-H/K/R all land, but this is secondary analysis 48 hours after launch, not the primary Anthropic release, so it stays below p1.

Xinzhiyuan · WeChat

Bilibili debate: Hermes responds to plagiarism claims for the first time, as MiniMax moves early on Harness

MiniMax says its M2.7 model now handles 30%-50% of daily workflows in its RL team, ran over 100 self-optimization loops, and improved evals by 30%. The post also says Hermes Agent grew from 2B to nearly 300B daily tokens, while M2.7 exceeds 25B daily tokens on OpenRouter; Hermes lead Tommy Eastman denied copying EvoMap in a livestream. The real signal is Harness: the post cites 20-40ms or 80ms sandbox startup and 15k to 600k instances per minute, showing competition is shifting from benchmark scores to agent execution infrastructure.

Why it matters: HKR-H/K/R all pass: the plagiarism-response angle pulls clicks, and the story carries concrete metrics on workflow share, self-optimization loops, sandbox latency, and concurrency. It stays at 83 because this is a dense secondary report, not a primary launch or official technical

X · @dotey

Anthropic designer Ryan Mather shares Claude Design tips while covering 7 product lines

Anthropic designer Ryan Mather shared 9 Claude Design workflow tips while covering 7 product lines. The RSS snippet says to spend 1 hour building a design system, use chat for large changes, comments for small edits, specify feedback like 8px spacing, and attach only the target component folder instead of a full monorepo. The key shift is process: from human-do/human-review to Claude-do/human-review.

Why it matters: This is a strong practitioner workflow note: an Anthropic insider shares concrete, reusable tactics, so HKR-H/K/R all pass. It stays below the 80s because this is not a formal Claude product release and the post does not disclose harder outcome data such as time saved or task win

Bloomberg Technology

Cursor In Talks to Raise $2 Billion at Over $50 Billion Value

Cursor is in talks to raise $2 billion at a valuation above $50 billion. The title only confirms it is an AI coding startup; the post does not disclose investors, round stage, revenue, or timing. The number to watch is the $50 billion pricing bar, not the rumor alone.

Why it matters: Bloomberg gives this strong source authority, and the $2B / $50B+ numbers land on HKR-H, K, and R. I keep it at 84, not p1, because the deal is still in talks and the story does not disclose investors, ARR, or closing timing.

Apr 17Friday

Hacker News front page

Measuring Claude 4.7's tokenizer costs

The author used Anthropic's free count_tokens API to compare Claude Opus 4.6 and 4.7 on 7 real samples and 12 synthetic ones; the real-sample weighted total rose from 8,254 to 10,937 input tokens, or 1.325x. Technical docs hit 1.47x, a real CLAUDE.md file hit 1.445x, while Chinese and Japanese stayed near 1.01x. On a 20-prompt IFEval sample, 4.7 improved strict prompt-level pass rate from 85% to 90%; the post cannot isolate tokenizer effects from model weights or post-training.

Why it matters: HKR-H/K/R all land: the post has a sharp cost hook, reproducible token-count data, and clear budget impact for Claude Code users. It stays below p1 because this is a third-party measurement, not an Anthropic release, and the IFEval slice is only 20 items.

Hacker News front page

Introducing Claude Design by Anthropic Labs

Anthropic launched Claude Design on April 17, 2026, in research preview for Claude Pro, Max, Team, and Enterprise subscribers. Powered by Claude Opus 4.7, it generates designs from text, images, DOCX, PPTX, XLSX, and codebases, and exports to Canva, PDF, PPTX, or HTML. The key detail is its one-step handoff bundle to Claude Code; the post does not disclose standalone pricing beyond existing plan limits.

Why it matters: This is a substantive Anthropic product launch, not a routine feature add. HKR-H/K/R all pass on novelty, concrete deployment details, and workflow resonance; the research-preview scope and limited pricing detail keep it at 84 instead of p1.

Tencent Technology · WeChat

From Vibe Coding to Agentic Engineering: Rebuilding the Full Backend Development Workflow

Tencent engineers report a one-week practice that used Claude Code plus custom Skills, Commands, and MCP servers to run an 11-stage backend workflow in one terminal session. The post gives reproducible details: one requirement-exploration step used 20 tool calls, 93.8k tokens, and 56 seconds; execution was split into 4 tasks and produced 3 commits. The real point is workflow orchestration, not raw code generation; human review remains at plan, deploy, and review gates.

Why it matters: HKR-H/K/R all pass: the story turns agentic engineering into a measured backend workflow test, with tool-call, token, timing, plan-length, task, and commit data. Stronger than generic coding hype, but still a practitioner case study rather than a major product or model release.

Hacker News front page

Discourse Is Not Going Closed Source

Discourse said it will keep its GPLv2 codebase open after 13 years. The post says its team used GPT-5.3 Codex, GPT-5.4, and Claude Opus 4.6 to scan code, and its last monthly release fixed 50 security issues. The key claim is defensive capacity: OpenAI said Codex Security scanned 1.2M+ commits in 30 days and found 792 critical and 10,561 high-severity issues.

Latent Space

[AINews] Anthropic Claude Opus 4.7 - one step better than 4.6 in every dimension

Anthropic launched Claude Opus 4.7 at the same $5/$25 per million input/output tokens; the post says 4.7-low through 4.7-high each outperform the matching higher 4.6 tiers. Reported changes include a new xhigh reasoning tier, Claude Code defaulting to xhigh, an 11-point gain on SWE-Bench Pro, and image input up to 2,576 px on the long edge (~3.75 MP). Do not overread the tokenizer change: the same input can use up to 35% more tokens, but the post says total token use still falls by up to 50% from prior equivalents.

Why it matters: Anthropic's flagship-model release fits the policy's 85–94 band. HKR-H/K/R all pass because the post gives concrete pricing, benchmark, image-limit, and token-accounting changes that hit Claude users' core coding and cost concerns.

X · @dotey

Boris Cherny shares practical tips from recent heavy use of Claude Opus 4.7

Boris Cherny outlined five ways to use Claude Opus 4.7, centered on Auto mode approving safe commands and a /go skill chaining tests, code simplification, and PR creation. The post names Auto mode, Recaps, Focus mode, effort level, and computer use; pricing, launch date, and benchmark data are not disclosed. The real shift is workflow, not just the model itself.

TechCrunch · AI

OpenAI upgrades Codex with more control over your desktop

OpenAI upgraded Codex on April 16, 2026, expanding its desktop control, and the headline frames it as a move against Anthropic. The truncated post only confirms more desktop power for Codex and says Claude Code has become a preferred tool for many businesses; the post does not disclose exact features, pricing, rollout, or permission limits. The key issue is the permission boundary, not the coding-tool label.

Why it matters: TechCrunch reports an OpenAI Codex desktop-control upgrade framed as a direct move against Anthropic, so HKR-H and HKR-R land. But HKR-K is limited: the article confirms broader permissions only, with no action list, pricing, or rollout details, so it stays at the featured floor.

X · @dotey

Claude Opus 4.7 uses more thinking tokens, so Anthropic permanently raised rate limits for paid users

Anthropic permanently raised rate limits for all paid subscribers because Claude Opus 4.7 uses more thinking tokens than its predecessor. The post confirms the affected group but does not disclose the increase size, pricing rules, or rollout timing; users who do not see the change should verify they are on Opus 4.7 and have updated Claude Code.

Why it matters: This is a substantive Anthropic quota update for paying users, with all three HKR axes present: a strong surprise hook, a concrete operational fact, and direct resonance on usage limits. It stays at featured, not P1, because the post does not disclose the size of the increase, pr

X · @dotey

Official best practices for using Claude Opus 4.7 with Claude Code

Anthropic shared guidance for Claude Opus 4.7 in Claude Code: the default Effort level is now xhigh, and users should provide goals, constraints, and acceptance criteria upfront. The post lists five Effort tiers—low, medium, high, xhigh, and max—with xhigh recommended for most coding, API design, migration, and code review tasks. The key shift is behavior: adaptive thinking is built in, while tool use and SubAgent spawning are less frequent by default, so prompts should state those needs explicitly.

Why it matters: This is not a model launch, but an official Anthropic workflow note that changes day-to-day Claude Code usage: default effort, 5 levels, and fewer tool/SubAgent calls unless asked. HKR-H/K/R all pass, but the scope is narrower than a major product release.

X · @dotey

Musk's xAI is turning into a GPU lessor, with $50 billion coding tool Cursor as its first customer

xAI is leasing tens of thousands of GPUs to Cursor to train its coding model Composer 2.5, while Cursor is reportedly fundraising at about a $50 billion valuation. The post says xAI's internal model FLOPs utilization is about 11%, versus a typical 35% to 45%, across roughly 200,000 Nvidia GPUs. The key point for practitioners is that xAI is starting to monetize idle compute as cloud capacity, not just build models.

Why it matters: This clears all three HKR axes: a strong strategic twist plus concrete numbers on utilization and fleet size. I keep it at 84, not higher, because this is business/economics reporting on capacity monetization, not a model launch, product ship, or top-level personnel move.

Apr 16Thursday

X · @op7418

Claude Code now supports Claude Opus 4.7

Claude Code now supports Claude Opus 4.7, and the RSS snippet confirms X-HIGH as the default reasoning level. The only concrete detail disclosed is that users must switch manually to Max if X-HIGH is insufficient. The post does not disclose pricing, rate limits, or launch timing.

Why it matters: This is a substantive Claude product update with all three HKR signals: a new model in Claude Code plus one concrete operational detail, X-HIGH vs. manual Max. I kept it below the top band because price, rate limits, and formal release timing are not disclosed.

X · @dotey

Anthropic officially releases Claude Opus 4.7 at unchanged pricing

Anthropic released Claude Opus 4.7 at unchanged pricing: $5 per million input tokens and $25 per million output tokens; the API name is claude-opus-4-7, now live across Claude, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry. The post gives two concrete changes: vision input now supports up to 2576 pixels on the long edge, and the new tokenizer can raise token usage to 1.0-1.35x for the same text. Watch migration cost, not list price; higher reasoning settings and multi-turn runs can increase output length and bills.

Why it matters: An Anthropic substantive model release belongs in the 85+ band, and this is not just a rename: the 2576px vision limit and 1.0–1.35x tokenization shift affect migration tests and billing immediately. HKR-H/K/R all pass, so it clears p1.

Hacker News front page

Claude Opus 4.7 System Card

Anthropic published a 232-page system card for Claude Opus 4.7 on April 16, 2026, saying it outperforms Opus 4.6 but remains below the limited-release Claude Mythos Preview. The card says Opus 4.7 does not advance Anthropic’s capability frontier, catastrophic risk remains low, cyber capability is roughly similar to Opus 4.6, and it does not cross the threshold for automated AI R&D. The excerpt does not disclose benchmark scores or the new cybersecurity safeguard details.

Why it matters: This is not a flashy launch post, but it is a substantive Anthropic system card update. HKR-K is strong: Opus 4.7 beats 4.6, stays below automated AI R&D thresholds, and is roughly similar to 4.6 on cyber evals; HKR-R lands because Claude users track general-access model ceilings

Hacker News front page

Introducing Claude Opus 4.7

Anthropic released Claude Opus 4.7 on Apr. 16 at the same price as Opus 4.6: $5 per million input tokens and $25 per million output tokens. The post says it improves on Opus 4.6 in advanced software engineering, long-running tasks, and higher-resolution vision, and ships across Claude, the API, Amazon Bedrock, Vertex AI, and Microsoft Foundry. The key detail is the first deployment of Anthropic’s cyber request blocking on a less capable model; the post cites benchmark gains but does not fully disclose every score in text.

Why it matters: Anthropic shipping Claude Opus 4.7 is a same-day write: GA, unchanged $5/$25 pricing, and rollout across Claude, API, Bedrock, Vertex AI, and Foundry give it direct workflow impact. HKR-H/K/R all pass, but the post does not publish full benchmark scores.

Hacker News front page

Qwen3.6-35B-A3B: Agentic coding power, now open to all

Qwen released Qwen3.6-35B-A3B as open weights, with 35B total parameters and 3B active parameters. The post reports 73.4 on SWE-bench Verified, 51.5 on Terminal-Bench 2.0, and 92.0 on RefCOCO. The key point is agentic coding and multimodal performance at a 3B active-parameter budget, with weights, Qwen Studio, and API access available.

r/LocalLLaMA

Qwen3.6-35B-A3B released

Qwen released Qwen3.6-35B-A3B as open source under Apache 2.0; it is a sparse MoE with 35B total parameters and 3B active. The post also claims agentic coding, strong multimodal perception and reasoning, plus thinking and non-thinking modes; the post does not disclose benchmarks, context length, or latency.

Why it matters: HKR-H/K/R all pass: a new open Qwen model is timely, and the post confirms 35B total, 3B active, and Apache 2.0. The score stays at 82 because this is still a launch post; benchmarks, context window, latency, and multimodal details are not disclosed here.