Skip to content

Models that plan, call tools and finish multi-step tasks on their own — from Claude Code and Manus to agent frameworks and benchmarks.

1,465 picksRelated topicsMCP & tool useAI codingReasoning

Latest picks

461–480 of 1,465

Jul 2Thursday

Latent Space

Paul Bakaus on skill engineering and why one-shot AI design is a dead end

Paul Bakaus presented Impeccable at the AI Engineer World’s Fair, an open-source design skill system for coding agents. Instead of one-shot full-site redesigns, users steer output with terms like 'bolder' or 'quieter' that the skill translates into precise design actions. Bakaus calls this 'skill engineering'—compressing expert vocabulary so agents don't converge on generic results. He noted designers now make up at least half of Impeccable's audience, using it as a bridge into code. He rejects full auto mode, arguing the goal is to insert human judgment at the exact point it matters most.

Why it matters: Paul Bakaus introduces 'skill engineering'—packaging designer feedback vocabulary into an open-source instruction set (Impeccable) to steer AI design iteratively rather than one-shot. The concept is novel, backed by a concrete artifact and user data. Score sits at the featured...

Ben's Bites

Fable 5 is back, and there's a new Claude Sonnet 5

Anthropic re-released Fable 5 for paid users with stronger guardrails, available in subscriptions only through July 7 and capped at 50% of usage limits. Scale's benchmark shows it completes 16% of remote work tasks, double Opus 4.8. Claude Sonnet 5 also launched—benchmarked close to Opus 4.8 on agent tasks, cheaper per token but roughly the same cost per task in practice; the author finds it expensive and slow. Google dropped two new models: Nano Banana 2 Lite for fast, cheap images and Omni Flash for video generation and editing. Bridgewater and Thinking Machines trained a financial triage model hitting 84.7% accuracy at 13.8x lower cost than the best frontier model tested.

Why it matters: Anthropic dropped Fable 5's limited return and Sonnet 5 simultaneously — two signals stacked. Scale benchmark provides hard comparable numbers, not pure marketing. Fable 5's 16% task completion rate doubled but absolute number is still low, so not pushing past 90.

AI HOT (Curated Pool)

Fable 5 hits 16.1% automation on freelance jobs in the Remote Labor Index, up from 2.5% eight months ago

The Remote Labor Index tests AI agents on 240 real freelance projects worth $144,000. Fable 5 reached a 16.1% automation rate, nearly double Opus 4.8's 8.3% and well ahead of GPT-5.5's 6.3%. Eight months ago the top score was 2.5%. 22 of Fable 5's projects couldn't be evaluated due to US government access restrictions; even in the worst case its rate would be 14.6%. The study also found AI judges overrate performance badly—GPT-5.5's score was inflated nearly 3x because the AI judge couldn't open professional software to inspect actual deliverables. No model's output passed as finished professional work, but the automation rate has more than quadrupled in under a year.

Why it matters: RLI is one of the few benchmarks using real paid freelance projects; Fable 5 hitting 16.1% — nearly double the runner-up — with a 6x improvement in 8 months is solid. Held below the top band because Fable isn't a tier-1 lab and the post doesn't disclose model size or cost, so ...

AI HOT (Curated Pool)

Qwen team's Zhu Da on consumer agents: 3× faster execution, 10× cheaper token cost vs overseas products

Qwen App shipped a general-purpose complex-task agent behind a capsule entry point in January 2026. Team lead Zhu Da frames the engineering philosophy as 'more, faster, better, cheaper': it handles info gathering and research tasks, execution time is down to one-third of the initial version, delivery quality improved through search paradigms and context management, and token cost is only one-tenth of comparable overseas products. The team is building toward proactive service with four components—User Memory, Environment, Task System, Assistant—and Zhu calls 'emotional intelligence' the hardest part. He maps agent engineering from Prompt Engineering to Harness Engineering, with AIWare Engineering as the next stage, guided by 'low power, good enough.' The post is an RSS snippet; it doesn't disclose specific latency figures or a timeline for proactive features.

Why it matters: A substantive engineering share from Qwen's consumer Agent team with real metrics and architecture breakdown. The self-reported nature and lack of third-party validation cap the score, but the 'more-faster-better-cheaper' framework and proactive-service design are directly use...

Latent Space

AIEWF Day 3: Autoresearch takes the stage, but speakers push back on full autonomy

Day 3 of AIEWF focused on autoresearch. Introspection's Roland Gavrilescu described it as an outer loop where agents maintain the system itself. Anthropic's Thariq Shihipar echoed continuous discovery in his Claude Code keynote, saying models are 'grown, not developed.' Former Google engineering lead Addy Osmani pushed back hard: the outer loop must stay human—inner loop is capability, outer loop is agency. Notion's Geoffrey Litt and Impeccable's Paul Bakaus both argued humans need to understand the code and steer the final 20%. Bakaus stated flatly there will 'never be auto.' Google's Nicole Brichtova added that cultivated expertise sees what average preference misses.

Why it matters: On-the-ground AIEWF report with first-hand quotes from Introspection and Anthropic — not a press release. But it's a conference roundup, not a product launch, so it lands at the featured threshold.

AI HOT (Curated Pool)

Ant Group's AI assistant Abao opens public beta in Alipay

Ant Group's AI assistant Abao is now in public beta inside Alipay—no invite code needed. iOS and Android users can search 'Abao' to access it. Swipe right in the app to see a chat interface plus an asset page. You can speak or type a request like 'check my housing fund,' and Abao pulls up the matching mini-program and service entry, collapsing multi-step navigation into one sentence. Alipay says all fund transfers and payments still require user confirmation; Abao only runs the workflow and presents the final step. The post doesn't disclose the underlying model, the number of supported services, or future pricing.

Why it matters: Ant Group open-tests an AI assistant inside Alipay that condenses multi-step tasks like checking housing fund into a single voice command, with a chat + asset page interface. It's a meaningful signal for domestic AI deployment, but the post doesn't disclose the underlying mode...

AI HOT (Curated Pool)

Tencent Hy3 released: matches much larger flagship models with 2-5x fewer parameters

Tencent officially released Hy3 under Apache 2.0. With 1/5 to 1/2 the parameters of competing flagship models, Hy3 matches or beats them on reasoning, agent, and long-context benchmarks. In a 270-person internal blind test, Hy3 scored 2.67/4 vs GLM5.1's 2.51/4. Hallucination rate dropped from 12.5% to 5.4%, multi-turn error rate from 17.4% to 7.9%. WorkBuddy task completion jumped from 72% to 90%, with 34% less time. API pricing: ¥1/M input tokens, ¥4/M output, ¥0.25 cached. The post does not disclose exact parameter count or training details.

Why it matters: Tencent Hunyuan releases Hy3, open-source under Apache 2.0, with parameter counts 1/5 to 1/2 of competitors yet matching or beating them on reasoning, agent, and long-context benchmarks. A 270-person blind test shows it beating GLM5.1. Domestic flagship open-source release is ...

AI HOT (Curated Pool)

Emil Kowalski turns UI animation taste into Skills for AI coding agents

Emil Kowalski released three design-engineering Skills for Codex, Claude Code, and Cursor. They encode hard rules: no animation on high-frequency actions, keep UI motion under 300ms, animate only transform and opacity, start entrances from scale(0.95) + opacity:0, and respect prefers-reduced-motion. The review-animations Skill outputs a Before/After/Why table; animation-vocabulary turns vague descriptions into professional motion terms. The post doesn't say whether the Skills are open-source or paid.

Why it matters: Emil Kowalski packaged UI animation engineering know-how into AI-usable Skills — a novel angle with concrete substance. But the audience is limited to the frontend/design-engineering crossover, lacking cross-domain resonance, so it lands right at the featured threshold.

Computing Life · Share · Yage

Claude Science skips the eureka moment and starts as a compute dispatcher

Anthropic's Claude Science desktop app, announced June 30, skips the AI-scientist fantasy and targets the grunt work that eats 80% of researchers' time. It auto-pulls and cleans data from UniProt, PDB, Ensembl, and ChEMBL, then writes SLURM scripts, sets up conda environments, and submits jobs to HPC clusters—retrying on failure. By constraining the model to verifiable execution tasks, it sidesteps hallucination risks. The post does not disclose pricing or a GA date.

Why it matters: Claude Science is a substantive Anthropic product release, and the article nails the positioning — not an 'AI scientist' but a compute orchestrator and data wrangler for research workflows. HKR all hit, but this is a third-party analysis, not a first-party launch post, and the...

Computing Life · Share · Yage

Fable 5's 18-day ban: Anthropic's share went to GLM

Anthropic's Fable 5 was taken offline by US export controls three days after launch, for 18 days. OpenRouter daily token data shows total volume grew from 24T to 32T, but Anthropic's share dropped from 20.7% to 17.6% and its absolute volume shrank. GLM was the biggest winner, jumping from 1.8% to 7.4% share. GLM-5.1 saw a spike on the day GLM-5.2 launched, then collapsed 48 hours later as GLM-5.2 took over. Community sentiment shifted from sympathy for Anthropic to mocking its business strategy. The post notes this only captures API-layer data, not first-party subscriptions, and token volume comparisons overstate GLM's share due to a 10x price gap.

Why it matters: A data-driven attribution of the 18-day Fable 5 ban's competitive impact using OpenRouter daily token data: Anthropic lost 3.1pp share, GLM gained 6.7pp, plus a weird anomaly where GLM-5.1 traffic spiked 4-5x after GLM-5.2 launched. Counterintuitive findings that directly touc...

Jul 1Wednesday

TechCrunch · AI

Google's agentic assistant Gemini Spark is now on Mac

Google brought Gemini Spark, its AI agent for file sorting and cross-app tasks, to Mac. It can read local files—turning invoices into a budget sheet, for example—and will later support remote phone-to-desktop commands. It's in beta, only for Google One AI Premium subscribers.

Why it matters: Google bringing Gemini Spark to Mac adds another player to the desktop agent race. Concrete feature details and subscription info give it substance, but it's a platform expansion rather than a new launch, and the paid-user-only beta limits reach.

MIT Technology Review · AI

Anthropic launches Claude Science; California's manure carbon math doesn't add up

Anthropic announced Claude Science at an event for pharma execs and biotech founders. It works like Claude Code but for research, autonomously handling computational biology and drug development tasks from short instructions. Anthropic will also use it in-house for rare disease drug research. Separately, the US lifted restrictions on Anthropic's Mythos and Fable models, restoring access today. Another piece digs into California's subsidies for turning cattle manure methane into natural gas—research suggests the carbon offset math is flawed and could lock in more warming.

Why it matters: Anthropic launched Claude Science, extending its Agent model from coding to scientific research, pitched directly to pharma. A significant Claude product line expansion with concrete use cases and internal adoption. Not 90+ because we only have the announcement — no performanc...

AI HOT (Curated Pool)

Meituan LongCat-2.0: A 1.6T MoE model trained and deployed entirely on domestic chips

Meituan released LongCat-2.0, a 1.6T total parameter MoE model activating ~48B per token, trained and deployed entirely on 50,000 domestic chips with over 35T tokens and no rollbacks or unrecoverable loss spikes. Agent performance stands out: it matches Gemini 3.1 Pro on Terminal-Bench 2.1 and SWE-bench Pro coding tasks, and ties Claude Opus 4.6 on FORTE general agent tasks. It offers up to 1M context and 128K max output, using LSA sparse attention and N-gram Embedding for long-context and tool-calling optimization. The API is live with OpenAI and Anthropic compatibility, ready for Claude Code and Codex workflows.

Why it matters: Meituan LongCat-2.0 is the first publicly disclosed trillion-scale MoE model trained end-to-end on domestic chips, matching Gemini 3.1 Pro on coding/agent tasks and Claude Opus 4.6 on general benchmarks. 50K domestic GPUs and 35T tokens without training collapse is itself a si...

AI Chat-Group Daily (群聊日报)

Claude Code found to embed China-user detection; Fable 5 export controls lifted same day

A Reddit reverse-engineering post reveals Claude Code since v2.1.91 silently classifies China-based users via timezone checks and encodes the result into Unicode apostrophe variants in the system prompt. Multiple group members were banned the same day; a reseller said Anthropic targeted Alibaba-related accounts. Meanwhile, the US Commerce Department fully lifted export controls on Fable 5 and Mythos 5. Ford became the top US recall leader after replacing engineers with AI. Sonnet 5 launched at $2/$10 per million tokens but uses a new tokenizer that inflates token counts. WeChat's built-in AI assistant 'XiaoWei' began grayscale rollout, raising privacy concerns as others can invoke it in private chats without consent.

Why it matters: Reddit reverse-engineering post confirms Claude Code uses Unicode steganography to flag Chinese users, with multiple ban reports the same day — high signal density and timeliness. Score capped below 85 because the source is a chat-group digest, not primary reporting, and the p...

AI HOT (Curated Pool)

AWS puts $1B into on-site engineers who embed with clients to ship AI

AWS is launching a new division that embeds engineers inside customer companies for 45-day stints to get AI agents into production workflows. It's putting $1B into the effort and plans to scale the team to thousands. First named customers are the NBA and Ricoh. Palantir has run similar on-site engineering for over a decade; Salesforce, Anthropic, and Google Cloud offer comparable services. LinkedIn reports demand for these roles jumped 42x from 2023 to 2025. AWS says success will be measured by how much faster clients reach real business outcomes.

Why it matters: AWS's $1B embedded engineer program is a concrete move in the 'last mile' of enterprise AI deployment, with real numbers, a defined model, and named first customers. But the post doesn't disclose current team size or hiring timeline, so it stays at the 72-point featured thresh...

Latent Space

Anthropic launches Claude Sonnet 5, but the real story is Fable 5's absence

Anthropic released Claude Sonnet 5 today, calling it the most agentic Sonnet yet with planning, browser/terminal tool use, and a 1M-token context window. Pricing stays at $3/$15 per million tokens, with a promo rate of $2/$10 through late August. The community reaction was muted: benchmarks show it consumed more tokens than Fable, and one test found it cost more than Opus 4.8. The post confirms Fable/Mythos 5 were approved for re-release after government work, but gives no timeline.

Why it matters: Anthropic dropped Sonnet 5 with 1M token context and unchanged list pricing, but tokenizer changes drove 3-6x real consumption—community benchmarks show it may cost more than Fable to run. HKR all hit: the launch is news, the efficiency twist has substance, and Claude users wi...

AI HOT (Curated Pool)

Meituan releases LongCat-2.0: a 1.6T-parameter model trained on 50,000 domestic GPUs, now open source

Meituan open-sourced LongCat-2.0, a 1.6T total-parameter model with ~48B activated per inference and native 1M context. It was trained and served entirely on a 50,000-card domestic GPU cluster. The architecture combines LSA sparse attention, zero-compute experts, ScMoE, and MOPD multi-expert fusion that blends Agent, Reasoning, and Interaction expert groups. SWE-bench Pro hits 59.5, Multilingual 77.3. A preview is live on OpenRouter and longcat.ai, already ranking top three globally in monthly calls on OpenRouter. The post doesn't disclose training cost, inference latency, or the specific domestic chip model, so I'd hold off on those details.

Why it matters: Meituan's trillion-param model trained end-to-end on domestic GPUs is the headline; code benchmark scores are solid. Not scoring higher because Meituan isn't a tier-1 model lab yet, and real-world usability depends on the open-source release.

AI HOT (Curated Pool)

Tim Cook and EU tech chief hold 'constructive' talks on new Siri AI

Apple CEO Tim Cook and EU tech chief Henna Virkkunen held a video call to discuss launching the new Siri AI in Europe without violating the Digital Markets Act. The new Siri can access personal user data, but the EU demands Apple open similar device data access to rival voice assistants. Apple proposed a 'trusted system agent' to mediate between user data and third-party AI, but hasn't built it yet and wants EU guarantees first. The EU sees this as a regulatory grace period that would harm competitors. The new Siri is already confirmed to skip EU iPhones and iPads this year.

Why it matters: Direct talks between Apple's CEO and the EU regulator over the new Siri's DMA compliance. The core conflict is clear. Score held back because the article doesn't explain how the 'trusted system proxy' works or give a timeline — key details are missing.

MIT Technology Review · AI

Anthropic launches Claude Science, a flagship product for AI-driven research

Anthropic launched Claude Science, positioning it alongside Claude Code as a flagship product. It writes code, runs experiments on compute clusters, and prioritizes reproducibility—aimed at computational biology and drug discovery. A live demo showed it identifying drug candidates for phenylketonuria. Anthropic will also use it for in-house rare-disease research. Harvard physicist Matthew Schwartz previously rated Opus 4.5's research ability at the level of a second-year grad student; Claude Science productizes that capability.

Why it matters: Anthropic flagship product launch with a clear positioning and live demo — a same-day must-write. Not above 90 because only a single MIT Tech Review report so far; pricing, availability, and multi-source confirmation are still missing.

TechCrunch · AI

Anthropic launches Claude Sonnet 5 as a cheaper way to run agents

Anthropic released Claude Sonnet 5, a midsize model that can plan, use tools like browsers and terminals, and run autonomously at a lower price. The company says this agentic capability required larger, pricier models just months ago. It directly competes with OpenAI's GPT-5.6 Sol preview and Google's Gemini 3.5 Flash, both pitched as agent-first tools. The post does not disclose specific pricing or benchmark scores, so the real cost savings are still unconfirmed.

Why it matters: Anthropic drops a mid-tier Sonnet 5 positioned as a cheaper agent runner, directly competing with OpenAI and Google equivalents. A model launch is hard news, and agent cost is a top pain point for developers — all three HKR axes hit. Not scoring higher because the post doesn't...