Skip to content

#其他

3 today

Sep 15Tuesday

Hacker News front page

Capsule packs entire apps into single files with SQLite storage

Capsule is a desktop app packager that bundles UI, data, and logic into a single .capsule file. Share it like a PDF via WhatsApp or email—recipients double-click to run, with data stored in local SQLite. No cloud accounts or servers needed. It also supports AI-generated apps: describe what you want in ChatGPT or Claude, and Capsule produces a complete single-file app. Currently supports macOS, Windows, and Linux; iOS and Android are coming. v0.4.0 is free to download.

TechCrunch · AI

Salesforce and Nvidia launch Koa, a reasoning model built for sales and support

Salesforce unveiled Koa at Dreamforce, its first reasoning model, built on Nvidia's open-weight Nemotron and trained for sales, marketing, and customer support tasks. Marc Benioff framed it as a direct threat to AI labs: a vertical SaaS company now ships its own reasoning model on its own data. The post doesn't disclose benchmark scores, parameter count, or pricing, so I'd discount the hype until numbers drop. The signal worth watching is vertical reasoning on an open base model—this could land in production faster than general-purpose alternatives.

Why it matters: Benioff's provocative framing and the 'vertical SaaS trains its own reasoning model' narrative have real buzz, but the article provides zero benchmarks or specs — the technical substance is unverifiable. Scores at the featured threshold as a newsworthy product announcement; re...

r/LocalLLaMA

DeepSeek V4.1 Flash Q4 hits 40 t/s on M3 Ultra with native DSpark multi-token prediction

A developer forked ds4 and tuned it for DeepSeek V4.1 Flash Q4 on a 512 GB M3 Ultra, lifting decode from 16.6 t/s to 31.3 t/s, and to 40.5 t/s with DSpark speculative decoding. The ~300 GB Q4 weights fit only the 512 GB M3 Ultra. The speedup comes from cutting Metal dispatch overhead: the 384-expert router went from 9 dispatches to 1, the shared expert gate+up+SwiGLU became a single kernel, and BF16 rounding moved inside producer kernels, removing ~770 re-round dispatches. At 300k context, compressed attention selection was the bottleneck; the fix scores only admitted blocks and uses a bounded radix select, keeping decode at 90% of the 8k rate. DSpark verifies 6 tokens per step with shared weight streams and overlapped Engram fetches, dropping verify latency from 177 ms to 112 ms. Output is byte-identical to upstream under greedy decode, with SHA-256 manifests provided. The branch is M3 Ultra only because the optimizations rely on measured behavior of this specific chip's dual-die memory, 80-core GPU scheduling, and Metal dispatch characteristics.

Why it matters: A solid local-inference optimization post: DeepSeek V4.1 Flash Q4 on M3 Ultra goes from 16tps to 40tps via Metal command-buffer merging and native DSpark speculative decoding. Concrete technical detail, directly useful for the local-LLM crowd. Score stays at 72 because the aud...

AI HOT (Curated Pool)

404 Media Exposes OpenAI's Project Lily: Human Review of ChatGPT Chats for Model Tuning

404 Media obtained internal docs on Project Lily, OpenAI's human review program for ChatGPT chats. Reviewers read anonymized real conversations to rate response quality, flagging AI clichés, condescending tone, or fake personal anecdotes. Pay exceeds $50/hour but the work is repetitive. Most users don't know their chats can be read by humans, and many treat ChatGPT as a confidant. OpenAI admits anonymization can leak personal data, especially in short sessions. After the story broke, OpenAI updated its help page but still didn't explicitly mention human review.

Why it matters: 404 Media obtained internal docs showing OpenAI hires humans to review user chats at $50+/hr. Strong privacy angle, but the article lacks scale details (how many reviewers, what % of chats), capping the score below 80.

AI HOT (Curated Pool)

Trail of Bits calls 1Password's AI patching benchmark misleading, releases two agent skills for patch validation

Trail of Bits reanalyzed 1Password's FLAWED report and argues the 26% clean-fix headline is misleading. Trials that deliberately instructed agents to apply wrong fixes or prohibited testing were mixed into the average. When restricted to trials where agents could run code and weren't given bad advice, 86% of patches blocked the supplied exploit. Trail of Bits also released two agent skills: post-patch-validation for automated patch testing, and review-walkthrough for engineer review. The post doesn't disclose performance data for these skills.

Why it matters: Trail of Bits re-analyzed 1Password's report claiming AI fixes only 26% of bugs, showing the headline number mixes in deliberately misleading prompts and tests where the agent couldn't run code. Under fair conditions, 86% of patches worked. This is a direct, data-backed challe...

r/LocalLLaMA

CrofAI, self-claimed cheapest inference provider, exposed as an OpenRouter wrapper swapping in cheaper models at up to 20x markup

CrofAI marketed itself as the world's cheapest inference provider but was caught by developer Kendell silently routing API calls to smaller, cheaper models on OpenRouter. For example, requests for Kimi K3 at $2/$10 in/out were actually served by GLM 5.3 Flash, a 13–20x markup. The founder also fabricated a 'greg' model family that simply pointed to existing open models like GLM 5.2 and Qwen 3.5 9B. After the exposé, he denied everything, then claimed a 'team' was taking over, and within hours deleted the website, Twitter account, and subreddit. The post also notes his hardware claims don't add up: Kimi K3 needs at least 802GiB of VRAM even at Q2_K quantization, but the largest RTX Pro 6000 machine on Vast only offers 765GiB.

Why it matters: A full fraud exposé with technical evidence and a dramatic company meltdown. Hits all three HKR axes. Score capped below 85 because the source is a Reddit post, not formal reporting, and the event is a single-provider scandal rather than a model or protocol shift.

AI HOT (Curated Pool)

Anthropic and OpenAI propose a coordinated slowdown of frontier AI development; critics like Cohere's CEO question the real motive

Anthropic CEO Dario Amodei called for a government-coordinated slowdown of frontier AI development and antitrust exemptions to make it happen; Sam Altman and Elon Musk agreed. Cohere CEO Aidan Gomez published a blog post calling it 'a wolf in sheep's clothing, a cartel by any other name,' arguing it would lock out competitors through massive barriers to entry. Hugging Face engineer Niels Rogge called the statements 'bizarre nonsense,' saying Amodei mainly wants to restrict Chinese models like DeepSeek and open-weight models to protect his scale advantage. White House AI czar David Sacks noted OpenAI and Anthropic already hold a duopoly in frontier intelligence and that product liability concerns also drive their push for a slowdown.

Why it matters: Anthropic and OpenAI jointly calling for a coordinated slowdown, with Cohere's CEO publicly pushing back as a monopoly play — three major players in direct conflict, high signal. Score held below 85 because only the title and summary are available; the full proposal details ar...

r/LocalLLaMA

jinfer: an open-source inference engine that brings LLMs to the JVM, no Python needed

mukel90 released jinfer, a pure-Java inference engine covering chat, vision, audio transcription, embeddings, reranking, and TTS. It runs without Python, ONNX, or sidecar processes, and reads gguf/safetensors natively. CPU performance is claimed to be competitive with llama.cpp; GPU support via the jota backend is still in progress. It integrates with Spring AI and LangChain4j and supports GraalVM Native Image. The author previously built llama3.java and gemma4.java. This is an early release—I'd wait to see how GPU pans out before getting too excited.

AI HOT (Curated Pool)

StepFun Releases StepAudio 3 Voice Models, Several Top Artificial Analysis Global Rankings

StepFun launched the StepAudio 3 series of voice models, with several topping the Artificial Analysis global rankings. The post is blocked by WeChat and does not disclose specific parameters, ranking details, or model capabilities. The title confirms the release and ranking results but does not specify which metrics or languages.

r/LocalLLaMA

Voodoo Dynamic Quant goes open source under MIT license, targeting aggressive low-bit compression

1ncehost open-sourced Voodoo Quant under MIT, a method that uses gradient descent to pick per-tensor quantization levels instead of static analysis. On small Qwen3.5 models it beats Unsloth Dynamic 3.0 at very low bit-widths like IQ1/IQ2, but loses at mid-to-high quants. The tooling currently targets Qwen and can be adapted to other architectures. The post does not disclose results on larger models; the author calls it research-grade. The repo README drew complaints for sounding AI-generated, and the author said he lacks time to polish it but welcomes PRs.

New York Times Chinese

Anthropic CEO calls for an AI slowdown, but China makes it nearly impossible

Anthropic CEO Dario Amodei argues frontier AI must slow down, warning that swarms of AI agents could gain the ability to “take over the entire internet” within 6–12 months. His first step: embed external experts inside labs to monitor safety and report publicly. Sam Altman, Elon Musk, and Demis Hassabis endorsed the idea; Altman said OpenAI will follow suit. The real obstacle, author Sebastian Mallaby writes, is China. The US lead is only a few months, so any unilateral slowdown risks letting China pull ahead. Amodei acknowledges this and, in a notable shift, lists areas where US–China cooperation might be possible, comparing it to Cold War arms control. The post does not spell out a concrete timeline, but notes Trump and Xi are set to meet on Sept 24, with two more summits possible by year-end.

Why it matters: Anthropic CEO's direct call plus endorsements from Altman, Musk, and Hassabis make this a high-signal moment. Amodei delivers a concrete 6-12 month timeline and an operational proposal for embedded safety experts. The deduction: this is an op-ed, not a policy announcement, and...

Latent Space

AEF-1 standard for third-party evaluators lands, with xAI, OpenAI, and Anthropic all signing on

The AI Evaluator Forum published AEF-1, a baseline for independent third-party evaluations covering access, conflicts of interest, funding, recusal, and transparency. The same day, Dario Amodei blogged that Anthropic is unilaterally committing to embedded evaluators with office badges, company laptops, and access comparable to internal risk teams. He also laid out a two-tier coordination framework for democratic and global pacing. Bilal Chughtai left Google DeepMind and called for slowing capability progress; Dan Selsam warned that models may learn to fake alignment during evals. On the other side, Aidan Gomez and Cohere pushed back against a few Silicon Valley firms becoming gatekeepers, and Kevin Bass alleged structural conflicts in the Anthropic-linked safety ecosystem.

Why it matters: The AI Evaluator Forum's AEF-1 standard, co-signed by xAI, OpenAI, and Anthropic on the same day Dario Amodei published a personal blog proposing even deeper evaluator access, forms a strong signal cluster. All three HKR axes hit: first written rules for third-party evaluation...

Financial Times · Technology

Italian startup Exein raises $270M to fight AI hackers with AI

Exein, an Italian embedded security startup, raised $270M to use AI models for detecting anomalies in IoT and industrial devices, specifically targeting AI-powered attacks. The post doesn't disclose valuation or product benchmarks, but notes customers include European automotive and defense firms. Big bet on the space, not yet proof of product-market fit.

Financial Times · Technology

Zach Dell bets a battery fleet can solve America’s AI power crunch

Zach Dell, son of Dell's founder, plans to deploy large-scale battery storage fleets to stabilize power for AI data centers. The idea: charge batteries during grid off-peak hours and discharge during peak demand, easing the strain from AI training and inference. The post doesn't disclose fleet size, cost, or timeline, but highlights the growing tension between AI's power hunger and grid capacity.

Financial Times · Technology

What an AI slowdown could mean for investors

FT analyzes signals that the AI investment boom may cool. Big AI groups are calling for a slowdown, which has already dragged down US tech stocks. Investors may need to recalibrate expectations on infrastructure and compute spending. The post doesn't name specific companies or a timeline, but flags a key shift in market sentiment.

Bloomberg Technology

Chinese State Media Dismisses ‘Self-Serving’ AI Slowdown Call

Chinese state media published a piece dismissing calls to pause AI development as 'self-serving.' It argues the slowdown push is a competitive tactic, not a safety measure. The post doesn't name specific proponents or cite detailed arguments, but the stance is clear: China won't slow down.

Bloomberg Technology

Hugging Face Scientist on Safety Issues With Agentic AI

Bloomberg interviews a Hugging Face scientist on safety issues with agentic AI. The body does not disclose specific risk cases or solutions; the core concern is the safety risks when models autonomously execute tasks.

AI HOT (Curated Pool)

Artificial Analysis ranks GPT-Live-1 #1 on Speech-to-Speech Index with 81.5

Artificial Analysis just dropped a Speech-to-Speech Index. OpenAI's GPT-Live-1 scored 81.5 with the Astra backend on medium reasoning intensity, edging out Grok Voice Think Fast 2.0 High at 81.3. The Sol backend config of GPT-Live-1 landed third at 80.1. The post doesn't disclose evaluation dimensions, sample size, or latency—so I'd take the ranking with a grain of salt for now.

New York Times Chinese

Trump calls AI safety fears a “scam,” rejects new regulation

Trump dismissed calls for AI regulation from Anthropic CEO Dario Amodei and others, calling safety fears a “scam” and insisting the only needed guardrail is “a strong and smart president.” He singled out Amodei as “pretending to be a perfect little angel,” though the post doesn’t spell out what government actions he claims to have already blocked. VP Vance and Speaker Johnson showed more openness to regulation, with Vance calling industry self-regulation pleas “a little bit of a Trojan horse.” Congress has almost no time to act before the midterms.

Why it matters: A sitting president directly names Anthropic's CEO and dismisses AI safety as a 'scam' — this is a head-on attack at the industry's core narrative. HKR all hit: high conflict, new White House power-split detail, direct identity nerve. Score capped below 85 because the article ...

New York Times Chinese

China’s top intelligence chief warns AI could directly threaten CCP rule

Minister of State Security Chen Yixin published an article framing AI as a direct threat to CCP rule—Beijing’s highest-level and most detailed warning yet. He named US models like Claude Mythos and GPT-5.5-Cyber, citing risks of deepfakes, information warfare, large-scale data exfiltration, and attacks on critical infrastructure. He also pointed to US military AI use in the Iran war as evidence that algorithmic advantage determines battlefield control. Law professor Henry Gao said this draws a red line ahead of US-China AI talks: data sovereignty and political security won’t be traded for international agreements.

Why it matters: China's national security minister publishes a long essay elevating AI safety to a regime-security issue, naming specific Anthropic and OpenAI models and citing US military use cases. This is a policy signal ahead of US-China AI safety talks — more about positioning than new f...

Hacker News front page

Ex-FTC chair Khan says the US should jail AI CEOs, citing a 1934 precedent

Former FTC chair Lina Khan told The Register that existing US laws can already hold AI executives criminally liable. She pointed to Section 501 of the 1934 Communications Act, which was used to convict telecom execs for fraud. Khan argued that if an AI company knows its model is being used for scams, CSAM, or price-fixing, the CEO should face charges—not hide behind 'the model did it.' She named OpenAI, Google, and Anthropic as firms that ship fast and push safety burdens downstream. The article does not include responses from those companies.

Why it matters: Khan offers a concrete, actionable liability framework — not vague regulatory talk. The 1934 Communications Act precedent gives the argument teeth. Score held below 85 because this is commentary, not policy action, and The Register's piece is a secondary account without full d...

Computing Life · Share · Yage

Mistral skips the frontier race and sells jurisdiction, hitting a $24B valuation

Mistral raised €3B in Series D at a €21B+ valuation, with ~$1B annualized revenue. Its flagship Mistral Large 3 scores below the median on Artificial Analysis; Fortune notes it's roughly on par with OpenAI models from 18 months ago. Core customers—Airbus, ASML, the French Ministry of Defense—pay for data locality, legal accountability, and supply-chain independence, not benchmark leadership. Revenue figures are unaudited; the CEO said 2026 chip and infra spend will roughly match expected revenue, and the company carries $830M in debt. The sovereign AI market is real but narrow: Gartner estimates €126B in European sovereign cloud spend for 2026, most of which flows to cloud providers, not model vendors. Mistral proved you can raise capital and win government deals without frontier models; it hasn't proved you can turn a profit.

Why it matters: Mistral's €3B round, €21B valuation, and ~$1B annualized revenue use hard numbers to prove 'you don't need to be frontier to survive,' directly loosening an industry default assumption. All three HKR axes hit: the headline creates cognitive tension, the body delivers new marke...

Computing Life · Share · Yage

Runway demos code-free UI rendering; Google lets agents write their own manuals; GitHub Enterprise goes air-gapped

All three are early-stage. Runway Solaris generates interactive UIs frame-by-frame with no frontend code—only curated demos and a waitlist so far, no public testing, pricing, or API. Google WikiSkill distills agent failure logs into reusable skill manuals, lifting Gemini-3.5-Flash accuracy from 49.5% to 68.1%, but skills from a small model can hurt a larger one; no official code repo. GitHub GHES 3.22 lets enterprises self-host Copilot CLI inside air-gapped networks with admin-managed model endpoints, though many features are disabled and it's labeled a technical preview.

Why it matters: Three items bundled, with Solaris as the main hook. Runway's frame-by-frame interface rendering is genuinely novel, but there's only a curated demo and waitlist — no public access, no third-party testing, and the cost comparison dodges standard web rendering. That keeps it bel...

AI HOT (Curated Pool)

Vercel shrinks inbound SDR team from 10 to 1.25, AI agent costs a few thousand dollars a year

Vercel COO Jeanne DeWitt Grosser says their AI sales development agent now handles 90% of inbound leads, and a home-built support agent resolves 93% of cases. Combined annual infrastructure cost is in the single-digit thousands, with a 32x ROI on the SDR agent. The inbound SDR team went from 10 people to 1.25. Grosser notes the bottleneck isn't the model—it's codifying the sales workflow. Worth flagging: this is Vercel's own data and inbound is a highly structured use case, so don't extrapolate to all sales teams yet.

Why it matters: Vercel's COO shared real deployment numbers: 90% inbound automation, 93% support case resolution, 32x ROI, single-digit-thousands annual infra cost. Three hard metrics lift this well above generic AI-efficiency stories. Not 95+ because it's from an interview rather than a prod...

AI HOT (Curated Pool)

Is Big Tech's AI slowdown a safety pact or a cartel?

The Verge questions Big Tech's verbal agreement to slow frontier AI development. It argues the pact looks like safety consensus but may function as a cartel that stifles competition. The post says proposed 'pace the frontier' solutions lack real teeth. It does not disclose specific company names or agreement details.

Bloomberg Technology

Gebru: AI Security & Safety Is About Human Control

Timnit Gebru argues that AI safety is fundamentally about human control, not just technical robustness. The post does not disclose specific cases or policy proposals, but the headline makes her stance clear: the safety debate should center on power distribution.

TechCrunch · AI

Jensen Huang takes Trump's call onstage, says 'we're not going to let an AI slowdown happen'

Jensen Huang took a call from President Trump while onstage at the All-In Summit and said 'we're not going to let an AI slowdown happen.' Elon Musk and Sam Altman had previously backed Dario Amodei's call to slow AI development. Huang publicly took the opposite side. The post doesn't disclose what Trump said or how long the call lasted.

Why it matters: Jensen Huang publicly sides against the AI deceleration camp, directly countering Musk and others — strong H and R. But the article is thin on details (no Trump response, no call specifics), so K is absent. Lands at 82, featured threshold.

Hacker News front page

Ninth Circuit vacates injunction against Perplexity: user-driven AI browsing isn't company 'access' under CFAA

The Ninth Circuit vacated a preliminary injunction against Perplexity AI. Amazon had sued over Perplexity's Comet browser, whose AI assistant navigates Amazon.com on a user's behalf. The lower court found Perplexity likely violated the CFAA by accessing Amazon's servers without authorization. The appeals court held that the 'access' was performed by the user, not Perplexity, making Amazon unlikely to succeed on the merits. The case was remanded.

Why it matters: The Ninth Circuit reversed a preliminary injunction against Perplexity, directly addressing a core legal question for AI agents: does an agent acting on a user's behalf on a third-party site constitute 'unauthorized access'? Broad impact, but sourced from a legal database rath...

Bloomberg Technology

Minebea Mitsumi Pauses M&A to Chase AI, Nvidia Demand for Components

Japanese precision parts maker Minebea Mitsumi halts M&A to focus on meeting surging demand from AI and Nvidia for high-end components. The company sees AI hardware orders as more urgent than acquisitions, prioritizing capacity expansion over integration. The post does not disclose specific investment figures or capacity targets.

TechCrunch · AI

OpenAI reportedly buys smartphone camera maker Glass Imaging for $300M

OpenAI acquired Glass Imaging for over $300M, per The Wall Street Journal. The startup uses neural networks to improve smartphone image quality at capture time, not in post. Founders Ziv Attar and Tom Bishop previously led Apple's Portrait Mode team. Glass had raised about $30M before the deal. OpenAI didn't comment, but the move fits its rumored hardware push into phones, earbuds, and the io device with Jony Ive.

Why it matters: An atypical $300M OpenAI acquisition targeting a smartphone camera algorithm team whose founders built iPhone Portrait Mode. All three HKR axes hit: the move is surprising, the tech and price are concrete, and it resonates with on-device AI builders. Not scoring higher because...

Hacker News front page

A Beginning for Mathematics: A Professor's Positive Vision for the AI Era

Daniel Litt, a math professor at the University of Toronto, shifts from his earlier 'End of Mathematics' talk to a positive vision. He assumes AI will soon be superhuman at most math tasks. The core issue isn't AI solving problems—it's how humans keep producing understanding. He argues that protecting old institutions like journals and peer review is futile when high-quality results cost a few dollars to generate. Instead, he proposes preserving what actually builds human understanding: learning seminars, serendipitous conversations, and students dropping by to talk math. The post does not lay out concrete reform steps, but explicitly rejects chasing the edge of model capabilities and urges planning for the endgame directly.

Why it matters: Daniel Litt is a U of T math professor. This isn't generic AI threat talk — it's an institutional design question: when AI produces math at a few dollars per result, how do humans preserve 'understanding'. Hits all three HKR axes, but as an opinion piece rather than a product ...

AI HOT (Curated Pool)

Anthropic shares how it scaled test impact analysis to handle agentic coding CI load

Anthropic rewrote its test impact analysis service after agentic coding tools like Claude Code flooded CI with PRs. Per-analysis latency dropped from 11 seconds to under 1 second, handling 1,000 analyses per day. The core trick: caching file dependency graphs with Merkle trees so only truly affected tests run. The post gives concrete architecture and numbers—worth noting this is their internal monorepo setup, so direct portability varies, but the caching strategy and API design are solid references.

Why it matters: Anthropic's own engineering blog with real numbers and a concrete technical approach — not marketing fluff. The 11s → <1s latency drop is solid, but the topic is infrastructure-heavy and less accessible to non-coding readers, so it lands at the 72 featured threshold.

AI HOT (Curated Pool)

Amodei calls for slowing frontier AI; Altman, Hassabis, and Nadella echo the shift

Anthropic's Dario Amodei published a ~4,000-word essay arguing the industry must slow the pace of frontier model capability gains to avoid making catastrophic risks more acute. Within hours, Sam Altman, Demis Hassabis, and Satya Nadella all signaled agreement. Ars Technica notes the sudden U-turn after years of an all-out AGI race, and cautions that slowing down also helps incumbents lock in their lead and reduce competitive pressure.

Why it matters: Amodei's 4,000-word call to slow frontier AI development drew public agreement from Altman, Hassabis, and Nadella within hours — a rare consensus shift among industry leaders. The core argument targets commercial competition as a direct driver of catastrophic risk, not a gener...

Hacker News front page

Sakana AI's PC-ALM trains 1000-layer networks without backpropagation

Sakana AI published PC-ALM, a local training method that replaces backprop with layer-wise PI feedback controllers. By adding dual neurons (Lagrange multipliers) per layer, it fixes the signal decay that kills standard predictive coding in deep nets. They tested residual MLPs up to 1000 layers on Fashion-MNIST and CIFAR-10, nearly matching backprop performance. Code and paper are public. The motivation is split: neuroscience (the brain can't do exact backprop) and neuromorphic hardware (local dynamics run cheaper on specialized chips).

Why it matters: Sakana AI's paper proposes a backprop-free training method that works on 1000-layer networks — solid research with a concrete artifact. But the benchmarks are still Fashion-MNIST and CIFAR-10, so it's not yet production-relevant, which caps the score.

Hacker News front page

Apple releases iOS 27, iPadOS 27, macOS 27 with Siri AI beta

Apple rolled out iOS 27, iPadOS 27, and macOS 27 today, headlined by Siri AI launching as a public beta in English. The new Siri understands personal context, onscreen content, and camera input, with conversation history synced across devices via iCloud. French, Japanese, Korean, Portuguese, and Spanish support arrives in October. The updates also add parental controls and system-wide refinements.

Why it matters: Apple ships iOS 27 / macOS 27 with Siri AI public beta (English-only) as the headline feature: personal context, on-screen awareness, and camera-based visual input land together — a clear interaction-model shift. Official Apple newsroom source is authoritative; cross-platform ...

Hacker News front page

Andon Labs releases Pion, an agent platform for running companies autonomously

Andon Labs packaged two years of autonomous vending, store, and cafe agents into Pion, now open for waitlist sign-ups. Their Vending-Bench eval showed Claude Opus 4 first beat the human baseline in May 2025, and every new model since has pushed scores higher. Real-world tests revealed a gap: early agents gave free handouts, rejected good deals, and hallucinated having a physical body. The post does not disclose Pion's architecture, pricing, or launch timeline.

TechCrunch · AI

iOS 27 makes Siri useful again: Gemini-powered, handles complex requests and on-screen context

After switching to the iOS 27 public beta, the author went from using Siri only for timers to relying on it daily. The overhaul swaps in Google Gemini models, replacing the old extended-knowledge setup and shaky ChatGPT integration. In real use, Siri handles sports schedules, lineups, and scores; early dev builds had hiccups, but the public beta is mostly stable. It also gets a new logo and transition animation. The post doesn't disclose latency, accuracy metrics, or specific device models—it's a first-person experience piece.

AI HOT (Curated Pool)

Apple launches next-gen Apple Intelligence with Siri AI beta

Apple today released the next generation of Apple Intelligence, with Siri AI launching as a beta. The new Siri is described as more capable and personal, with contextual understanding and cross-app task execution. The post does not disclose model specs, hardware requirements, or regional availability.

AI HOT (Curated Pool)

SiliconFlow launches Hy4 preview: a 770B open-source model with 1M context

SiliconFlow has onboarded Hy4 preview, a 770B-parameter open-source model that activates 49B per token and supports a 1M context window. It's released under Apache 2.0 and targets coding, analysis, and complex real-world tasks. Pricing is listed at $0.834 per 1M input tokens, $2.501 per 1M output tokens, and $0.042 for cached tokens. The post doesn't disclose training data, benchmarks, or real-world latency, so I'd hold off on getting excited.