Skip to content

#其他

3 today

Sep 9Wednesday

OpenAI News

OpenAI launches GPT-6 Astra, built for computer use, document work, and cost efficiency

OpenAI launched GPT-6 Astra, a model designed for complex enterprise work. It can directly operate everyday apps like Excel and Figma without APIs. In an Excel modeling challenge, it was about four times faster than the winning human. On Terminal-Bench 4.0, it scored 57.9%, compared to 37.3% for GPT-5.6 Sol and 55.8% for Claude Fable 5.1, with roughly 9% and 63% lower estimated API cost per task. Pricing starts at $10/1M input tokens and $50/1M output tokens. Early customers praised its judgment, deck-building fidelity, and lower hallucination rate. Internally, OpenAI used it to edit a multi-camera video and fix a memory bottleneck, cutting latency by 25x.

Why it matters: OpenAI's GPT-6 Astra release, with direct computer use and speed surpassing human champions, is an industry-shaking event. HKR all hit, score near ceiling. The post doesn't disclose pricing or exact rollout scope — that's the only info gap right now.

AI HOT (Curated Pool)

OpenAI's millennium proof dispute raises the question of whether researchers can trust AI labs

Mathematician Tristan Buckmaster accuses OpenAI of training on drafts he uploaded to Codex and pressuring him to drop his Anthropic-employed co-author. OpenAI admits it mobilized resources after hearing rumors that Anthropic had solved a Millennium Problem, denies plagiarism, but says it 'cannot rule out' that de-identified data helped its models. Altman backs his team; Alpöge disputes Altman's account of his willingness to cooperate. Terence Tao warns this sets a precedent where labs can overtake original research based on rumors alone.

Why it matters: The dispute has strong topical pull — a Millennium Prize problem, a named accuser with a concrete timeline, and two top AI labs involved. The deduction is because the excerpt only gives Buckmaster's side; OpenAI's response and Codex's position aren't fleshed out, so the full p...

AI HOT (Curated Pool)

Thomas Wolf: AI math isn't solved yet—Navier-Stokes result looks more like counterexample search

Hugging Face co-founder Thomas Wolf responded to OpenAI's claim that a swarm of next-gen model agents proved the Navier-Stokes Millennium Problem false. He called the result impressive but sees it as counterexample search rather than a full proof—AI math isn't solved yet. The post doesn't disclose the model name, proof details, or verification status.

Why it matters: Thomas Wolf's public pushback against OpenAI carries inherent news value, and his distinction between counterexample search and full proof adds real insight. Score capped at 72 because the post lacks model names, proof details, and external verification — the information densi...

AI HOT (Curated Pool)

DeepSeek reportedly hires CITIC Securities for STAR Market IPO, aims to file this year

Reuters reports DeepSeek has hired CITIC Securities to prepare for a STAR Market IPO, aiming to file this year and list next year. Fundraising size and target valuation are not yet set. The company plans to use IPO proceeds to expand computing infrastructure, boost model R&D and chip development, and strengthen talent incentives. Revenue in the first seven months of 2026 reached about 475 million yuan, roughly 10x its full-year 2025 revenue. DeepSeek is also raising a new funding round targeting a ~500 billion yuan valuation, following a June round that valued it at over $50 billion post-money. Investors include Tencent, CATL-linked entities, JD.com, and NetEase. Intense demand has spawned multi-layered SPVs reselling access with front-end fees exceeding 15%. Founder Liang Wenfeng is personally vetting final investor lists to block unknown entities ahead of the IPO.

Why it matters: DeepSeek's STAR board IPO push is an industry-level event, with a 10x revenue jump and in-house chip plans adding real substance. Not scoring higher because the fundraising amount and final valuation are still undecided — it's early in the process.

New York Times Chinese

OpenAI claims solving Navier–Stokes Millennium Problem with 10,000 AI agents

OpenAI says an unreleased model solved the Navier–Stokes existence and smoothness problem in 88 hours, using up to 10,000 AI agents working together. Researchers acted as 'bumblebees' cross-pollinating ideas across agent groups, and the final proof was written in the Lean language. OpenAI researcher Noam Brown called it 'a very expensive process,' likely costing millions of dollars. Terence Tao compared it to a guide showing one path to a hidden waterfall—people will take that path and stop searching for others. OpenAI published a paper and the Lean proof for external review. The post does not name the specific unreleased model.

Why it matters: OpenAI claims to have solved Navier-Stokes with an unreleased model — an industry-shaking event. The 88-hour run, 10,000+ agent swarm, and Lean proof submission are dense with signal. The deduction: the proof hasn't passed external review yet, so this stays below 95 until veri...

Product Hunt · AI

Perplexity Hybrid Compute: cloud for research, Mac for privacy

Perplexity splits AI workloads: cloud handles research-heavy tasks, while privacy-sensitive processing stays on your Mac. The post is a one-liner with no details on switching logic, supported Mac models, latency, or cost. Treat this as a concept announcement until real-world tests show whether privacy and UX can coexist.

AI Chat-Group Daily (群聊日报)

OpenAI solves Navier-Stokes with 10K agents, but Codex data privacy debate steals the show

OpenAI deployed ~10K concurrent agents to solve the Navier-Stokes Millennium Problem in 88 hours, consuming 130B output tokens. But NYU mathematician Buckmaster publicly alleged OpenAI may have accessed his and collaborator Alpöge's unpublished drafts via Codex—their technical approaches overlapped heavily. OpenAI hasn't directly denied accessing Codex data, only stating they 'cannot rule out that de-identified data helped improve models.' The group debated whether personal subscriptions offer true zero data retention: only Team/Enterprise plans do. On the practical side, third-party benchmarks show Astra's xHigh effort costs more than High but scores slightly lower—High is the daily sweet spot. DeepSeek V4.1 Flash internal test model hits 340–450 tok/s with impressive SVG morphing quality, expiring Sept 10. GPT Image 2.5 launched with doodle canvas and native transparency. Codex's new experimental context management replaces compression with note-taking, cutting window-switch time from 27s to 1.8s.

Why it matters: A claimed Millennium Prize solution is already industry-shaking; the Buckmaster plagiarism accusation and OpenAI's non-denial push it into must-cover territory. Source is a curated group-chat digest, but it cites the official OpenAI post and a named mathematician's public alle...

New York Times Chinese

Calif built a zero-click WeChat worm that jumps between iOS and Android using AI

Calif researchers used a mix of AI models to build WeWorm in just over a week—a zero-click worm that spreads across iOS and Android through WeChat. An incoming call from a compromised contact is enough; the worm hijacks the account and then dials the victim's own contacts. Tencent confirmed the flaw and patched it, stating no users were affected. The attack exploits WeChat's trust model for saved contacts and, combined with other bugs, can give full phone control. Calif says human expertise was still needed end-to-end, but AI dramatically accelerated vulnerability discovery and exploit construction. The disclosure lands right after 100+ companies warned of an incoming wave of AI-powered cyberattacks.

Why it matters: NYT exclusive with Tencent confirmation and patch. AI-assisted vulnerability discovery moves from theory to real-world demo. Slight discount because the bug is already fixed with no known user impact, but the warning value clears the featured bar.

Latent Space

OpenAI claims Navier-Stokes singularity find with ~10k agents and 88 hours of compute

OpenAI posted that a swarm of ~10k agents powered by a next-gen model (Astra-next) produced a Navier-Stokes finite-time singularity result in 88 hours, consuming 130B tokens at an estimated cost over $40M. If verified by the math community, it would be the second solved Millennium Prize problem. No preprint, proof sketch, or peer review is public yet. The 88-hour figure comes from a satirical post, not an official OpenAI statement, and the human-vs-model division of labor isn't spelled out.

Why it matters: A Millennium Prize-level math breakthrough would be historic if verified. But there's no preprint, no proof sketch, no peer review — just a paid newsletter recounting the claim. The post doesn't link to OpenAI's original announcement or any verifiable source. I'm discounting t...

Product Hunt · AI

Type.com: A shared workspace for Claude, Codex, and your team

Type.com launched a shared workspace that supports both Claude and Codex. Teams can use different AI models in one interface without switching tools. The post doesn't spell out specific collaboration features, model versions, or pricing.

Financial Times · Technology

DeepSeek fundraising frenzy spawns a shadow market with steep access fees

DeepSeek is raising a new round, and getting in is tough. The FT reports that some funds and family offices are buying access through intermediaries, paying a 2% management fee plus 20% performance carry—far steeper than standard terms. One investor, chasing a $5–10M allocation, also had to promise the middleman a bigger cut on future deals. DeepSeek says it hasn't authorized anyone to sell stakes, but the post doesn't disclose the round's total size or valuation. I'd take these off-market quotes with a grain of salt, but they show how hot demand is.

Why it matters: FT exclusive on a shadow market for DeepSeek fundraising, with concrete fee numbers and an investor anecdote—high signal. Score held at 78 because the round's total size and valuation are missing, and DeepSeek only gave a denial without further detail. HKR all hit, but the inf...

AI HOT (Curated Pool)

How to Choose GPT-6 Astra Inference Levels to Save Tokens

The article body is blocked by WeChat, only the title remains. It mentions GPT-6 Astra has multiple inference levels and choosing the right one saves tokens. But the post discloses no details on levels, selection criteria, or savings.

AI HOT (Curated Pool)

OpenRouter Tutorial: Edit Images with Nano Banana 2 in Code

OpenRouter published a tutorial showing developers how to call Gemini's image editing model with one API. The default model is Nano Banana 2 (google/gemini-3.1-flash-image). You send a source image and a text instruction, and the API returns the edited image. The tutorial includes full Python and TypeScript examples: encode local files as base64 or pass a hosted URL. Edit in small steps—send each result back as the next input to stack changes. Switch models by changing one field. The post doesn't spell out pricing or latency for Nano Banana 2, only that the family includes a cheaper Lite and a pricier Pro version.

AI HOT (Curated Pool)

OpenRouter reviews Seedance 2.5: strong at long takes and editing, but no 1080p

OpenRouter published a hands-on review of Seedance 2.5 on Sep 9. The model went live Aug 7, 2026, and excels at 30-second single takes and editing from existing footage. Cost is ~$0.103/sec at 480p and ~$0.231/sec at 720p; using a video reference cuts the token price by ~40%. Audio generation adds no extra charge. The clear trade-off: it caps at 720p. For 1080p or 4K, you still need Seedance 2.0 or Veo 3.1. Frame-exact reproducibility is also not guaranteed.

Why it matters: OpenRouter's hands-on review of ByteDance's Seedance 2.5 delivers concrete pricing ($0.103/sec at 480p, $0.231/sec at 720p, 40% off for video-reference tokens, free audio) and a clear resolution cap at 720p. Solid product intel but narrow audience fit, landing right at the fea...

AI HOT (Curated Pool)

OpenRouter launches US in-region routing, keeping decryption and inference inside the country

OpenRouter added a US in-region routing endpoint (us.openrouter.ai) for Business and Enterprise plans, matching the EU routing launched last October. Requests are decrypted and served only by US-based providers; if no in-region endpoint exists, the call fails with a 404 instead of falling back outside the region. Chinese open-weight models like DeepSeek V4 Pro, Kimi K3, and GLM 5.2 are available through US routing because Baseten, Fireworks, and Azure host them in US data centers. The post also flags that many gateways only pin inference to a region while decrypting traffic elsewhere—OpenRouter locks the full path inside the chosen region.

Computing Life · Share · Yage

Three Small Things Last Week: Agent Ledgers, Interfaces, and Rooms

Several AI engineering efforts last week converged on the same bottleneck: agents forget, collide, and can't touch the physical world. Security researcher Jordy Zomer open-sourced Lemmalog, splitting agent memory into a probabilistic front-end for fact extraction and a deterministic Datalog engine for causal reasoning and cascading retraction. Anthropic and Janelia unveiled the Model Hardware Standard, letting LLMs control lab instruments via natural-language labels while hard-coding safety limits in driver firmware—a lesson learned after Claude mistook liquid foaming for a software error at Genentech. Startup Raft blamed multi-agent chaos on the room, not the models, identifying a reasoning-commit gap where agents act on stale snapshots; their fix includes draft holding, pull-based inboxes, and silence as a valid action. Anthropic also formalized Fermat's Last Theorem in 11 days, but early multi-agent attempts collapsed until the team moved to the Prove2Me dependency-graph platform. All four stories share one pattern: wrapping probabilistic models in deterministic engineering scaffolds. Lemmalog scored just 0.128 on preference tasks, MHS figures are all self-reported with no public spec, and Raft's scale claims lack third-party discussion—discount these numbers for now.

Why it matters: Three stories converge on the same agent engineering bottleneck; Lemmalog's causal ledger has concrete implementation and open-source code. But this is a personal blog roundup, not a first-party release, and only Lemmalog gets detailed treatment — MHS and multi-agent parts are...

Financial Times · Technology

OpenAI faces competing claims around maths breakthrough

FT reports OpenAI achieved a math reasoning breakthrough, but at least two teams claim they independently produced similar results. The full article is behind a paywall and does not disclose technical details, model names, or benchmark scores. Only the existence of a priority dispute over math reasoning progress is confirmed.

AI HOT (Curated Pool)

Meta's AI agent Muse opens for trials, team responds to early feedback

Meta's AI agent Muse is now open for more users to try. Chief AI officer Alexandr Wang thanked the community for feedback. Users praise its design, speed, and browser agent workflows, plus deep integration with Instagram. However, naming like 'soul.md' confuses non-tech users, and feed relevance needs improvement.

Product Hunt · AI

Frigade Assist API: Let your AI agent show users where to click

Frigade launches Assist API, letting AI agents generate on-screen guides that tell users exactly which button to click next. The post doesn't specify supported models or latency, but the pitch is clear: agents write their own tutorials, no manual recording needed. For teams building AI customer support or automation flows, this cuts guide maintenance overhead.

TechCrunch · AI

Hackers are stealing Claude tokens from subscribers

A UK-based AI consultant noticed his Claude Max 20x account kept burning tokens while idle, climbing from 45% to 55%. Anthropic confirmed the anomaly, suspended his paid account, and refunded £44.49, but didn't provide an itemized usage log. The suspension disrupted his business, which relies on Claude for client agent workflows. The post doesn't explain how attackers obtained the tokens or how many users are affected.

Why it matters: Anthropic security incident with a named victim, concrete numbers, and official response — hits all three HKR axes. Held back from higher bands because it's a single case reported by TechCrunch, not an Anthropic disclosure, and we only have the user's side. 78, low featured.

AI HOT (Curated Pool)

OpenAI rolls out Astra to all paid-tier users

OpenAI pushed Astra to Plus, Pro, Business, and Enterprise users across Codex and ChatGPT Work. The post doesn't explain what Astra does or list any specs or pricing changes, but it links to a live demo.

Why it matters: A full-tier rollout is a signal, but the post doesn't explain what Astra is — the info gap is too large, so the score sits right at the featured threshold. If the demo shows concrete capabilities or numbers, it can go higher.

TechCrunch · AI

Cognition hits $48B valuation, signaling AI coding is far from a winner-take-all market

Cognition raised a new round at a $48B valuation with ~$250M in annualized revenue. The multiple is higher than Cursor's before its sale to SpaceX, showing investors don't see AI coding as winner-take-all. Devin's 'AI software engineer' pitch is landing enterprise deals, though the post doesn't disclose the exact funding amount or investors. Worth a discount: revenue is less than 1/200 of the valuation—the market is still voting with its feet.

Why it matters: Cognition's $48B valuation and $250M ARR are concrete, and the non-winner-take-all thesis is contrarian. But the post doesn't disclose the round size or investors — missing key facts keeps it at 78, not 85.

TechCrunch · AI

Meta launches Muse, a personal AI agent that wants access to your email, calendars, and payments

Meta just launched Muse, a personal AI agent currently available only in the US. It connects to a user's email, calendars, payments, health apps, smart home devices, and more to handle everyday tasks. This is Meta's biggest consumer AI bet yet, but the article points out it comes just two weeks after Meta's $18 billion settlement over social media harms to children—making user trust a major open question. The post doesn't disclose technical details, pricing, or a rollout timeline.

Why it matters: Meta betting big on a consumer AI agent is significant, and Muse's permission scope is genuinely more aggressive than existing assistants. But the post is a product announcement with no technical details or pricing — K axis missed. The trust angle resonates, but the informatio...

AI HOT (Curated Pool)

NYU mathematician accuses OpenAI of dirty tactics in millennium problem race

NYU math professor Tristan Buckmaster announced three proofs with a preliminary finding on the Navier-Stokes existence and smoothness problem, which carries a $1 million prize. In his statement, Buckmaster accused OpenAI of dirty tactics in a parallel effort to solve the same problem. The work was done with Anthropic mathematician Levent Alpöge using Codex and Claude models. OpenAI's Sébastien Bubeck denied the allegations. The post does not spell out the specific misconduct.

Product Hunt · AI

DuckFightClub: Train your MicroDuck and win the Golden Beak Belt

DuckFightClub is an AI robot fighting league where teams train reinforcement-learning policies for Pollen's open-source MicroDuck, then battle in a simulator livestreamed to viewers. You don't build the robot—you train its brain. Teams can register, host local showdowns, and eventually fight IRL when real MicroDucks ship.

OpenAI News

GPT-5.6 Sol runs quantum chip calibrations, freeing MIT grad student from routine lab work

OpenAI published a case study: MIT grad student Beatriz Yankelevich connected GPT-5.6 Sol to lab software to autonomously run calibration measurements on superconducting qubits. The model handled standard sequences—finding frequencies, calibrating pulses, measuring coherence—with little intervention when signals were clean. Weak or noisy signals still required researcher guidance. EQuS now routinely runs agents overnight; researchers check results from their phones. The post doesn't specify hours saved but says a chip previously took days to characterize.

Dwarkesh Patel podcast

Pretraining progress is mostly coming from data

Dwarkesh Patel and Jerry Han trained models using year-specific open recipes and data corpora from 2019–2025. At a 1e19 FLOPs budget, data improvements delivered a 12x compute-efficiency gain versus 3.7x from model improvements, with additive effects. The authors argue model research's main value was enabling larger-scale training—MoE, FlashAttention, stability fixes—not just saving FLOPs. Small models benefit more from data quality; big models may prefer quantity over aggressive filtering, though the post doesn't confirm this at frontier scale.

Why it matters: Dwarkesh and Jerry Han ran a controlled experiment across six years of public recipes, decomposing pretraining progress into 12x from data and 3.7x from model improvements—clean conclusion backed by numbers. Not scoring higher because this is a blog post, not a peer-reviewed p...

AI HOT (Curated Pool)

Mistral's post-mortem on migrating 40k lines of Fortran 77 to C++ with AI agents

Mistral published an engineering post-mortem on using their own AI agents to migrate a 40k-line Fortran 77 codebase to C++. The piece focuses on three hard parts: getting agents to understand undocumented legacy logic, preserving numerical precision after translation, and designing a verification pipeline to catch bugs. The post doesn't disclose which model was used, total time spent, or how many manual fixes were needed. Treat this as a methodology reference, not a product launch.

Why it matters: Mistral published a real engineering retrospective on using their own AI agents for a legacy migration, with concrete breakdowns of three hard problems — not a product launch fluff piece. Score held back because the post doesn't disclose which model, total time spent, or how m...

Sep 8Tuesday

TechCrunch · AI

Chrome ships updates every 2 weeks as AI reshapes security

Google has cut Chrome's release cycle from four weeks to two, starting with Chrome 153 on desktop, iOS, and Android. The move is driven by AI-powered attack tools and a surge in community bug reports, which demand faster patch turnaround.

Hugging Face Blog

Safety alignment should refuse the harmful subset of a topic, not the whole topic

Multiverse Computing's new paper argues that current safety alignment treats entire topics as refusal units—LlamaGuard-3, for instance, labels elections as 'factually incorrect information,' causing models to refuse even benign queries. They propose 'narrow-boundary safety': within a single topic like politics, refuse only the harmful subset (e.g., writing targeted manipulation) while still answering benign questions (e.g., election facts). The method uses deployment-specific boundary labels to self-distill a model that respects per-setting splits instead of topic-level blocks. Experiments focus on politics; the post doesn't disclose generalization results for other topics.

Why it matters: H and K both hit: the angle is sharp and the LlamaGuard-3 mislabeling case is concrete. R is weak — this is a safety-alignment niche topic that won't resonate broadly. Landed at the featured threshold of 72; didn't go higher because the body excerpt is partial, with full exper...

TechCrunch · AI

Mistral raises €3B Series D at €21B+ valuation

French AI lab Mistral AI closed a €3B (~$3.58B) Series D at a post-money valuation above €21B. Samsung Electronics led the round, with EQT's Scaleup Europe Fund and PSG Equity as co-leads. The company says it's the largest equity raise by a European tech firm. Funds go to compute, infrastructure, commercial growth, and international expansion. Mistral insists it isn't building a European ChatGPT and still sees itself as an AI lab.

Why it matters: Mistral lands the largest equity round in European tech history, Samsung-led, valuation above €21B — all three HKR axes are solid. Not scoring higher because competitive positioning details are thin; the post doesn't disclose current revenue or market share, so it stays in the...

OpenAI News

OpenAI CFO: GPT‑6 Astra is here, and consumer + enterprise reinforce each other

OpenAI CFO Sarah Friar published a blog framing GPT‑6 Astra as the world's most capable and aligned model. ChatGPT now has over 1B weekly active users and 2.5M business customers. Internally, the research org uses 3.1 agent-workdays per human workday. The post also claims an internal model solved the Navier–Stokes Millennium Prize Problem, but gives no technical detail. I'd treat this as a strategy narrative, not a technical report.

Why it matters: OpenAI CFO publishes a strategic framing piece for GPT-6 Astra with two concrete numbers: 1B weekly users and a 3.1x agent-workday ratio. Hits all three HKR axes. No technical details — this is narrative, not a product launch — so it stays below 85.

OpenAI News

OpenAI launches ChatGPT Images 2.5 with faster generation and sharper editing

OpenAI released Images 2.5, a new image model that cuts generation latency by up to 50% and improves lighting, textures, and multi-turn editing consistency. Over 3 billion images are already created weekly across ChatGPT and the API. A new Sketch feature lets users draw directly in ChatGPT as a reference. API availability is confirmed, but the post does not disclose pricing details.

Why it matters: OpenAI officially released Images 2.5 with 50% lower latency, quality improvements, a new Sketch feature, and 3B images/week volume. It's a substantive update to a core ChatGPT capability, hitting all three HKR axes. Not scored higher because this is an iterative upgrade rathe...

OpenAI News

OpenAI commits $5M to study how generative AI affects teens

OpenAI is funding $5 million in independent research on how generative AI affects teens aged 13–17. The program covers emotional development, social relationships, demographic variance, and safety design. Applications are open globally, with priority for countries with high AI adoption. Research involving minors must detail ethics review, consent, privacy, and data security.

AI HOT (Curated Pool)

Mathematician Buckmaster announces PDE blowup results aided by LLMs, details OpenAI communication

NYU mathematician Tristan Buckmaster and collaborator Levent Alpöge announced three finite-time blowup results for incompressible porous media, Boussinesq, and 3D incompressible Euler equations, all with smooth forcing. They relied heavily on LLMs (Claude, Codex, GPT-5.6 Sol, Astra) and verified proofs in Lean. Buckmaster called the Euler writeup "AI slop" and detailed his communication with OpenAI: an internal OpenAI model claimed a forced Navier-Stokes blowup proof, but Buckmaster believes the team used extensive human effort and compute, contrary to claims of "very little human input." The post does not disclose the details or verification status of OpenAI's proof.

AI HOT (Curated Pool)

Mistral raises €3B Series D at a valuation above €21B

Mistral closed a €3B Series D round at a valuation north of €21B. The company says the money will push sovereign, open-weight AI to the frontier. The post does not name the investors or the closing timeline.

Why it matters: Largest single round for a European AI lab, with a clear differentiated positioning around sovereign + open-weight AI. Held back from 85 because the announcement doesn't name investors or specify how the €3B will be allocated — the post is light on concrete detail.