Skip to content

OpenRouter

The OpenRouter model router: new listings, usage rankings and real market signals on what developers pick.

Latest picks

41–60 of 77

Jul 14Tuesday

AI HOT (Curated Pool)

Tencent Hunyuan releases 1-bit and 4-bit quantized Hy3, a 295B MoE that runs on a single GPU

Tencent Hunyuan quantized its flagship Hy3 (295B MoE) into 1-bit and 4-bit versions that run on a single GPU. Hy3 is claimed to be best-in-class at this scale and competitive with trillion-parameter models for most agent scenarios. The quantized versions work via llama.cpp with MTP support, drastically lowering hardware requirements. Apache 2.0 license, commercial use allowed, plus two weeks of free API through OpenRouter. The post doesn't disclose quantization accuracy loss or the specific GPU memory needed.

Why it matters: Tencent Hunyuan's quantized Hy3 puts a 295B MoE model on a single GPU — immediately actionable for local deployment and agent builders. Apache 2.0 license plus a two-week free API window lowers the barrier to test. Held below 85 because the post doesn't disclose quantization a...

Jul 7Tuesday

AI HOT (Curated Pool)

OpenRouter: Low-res images can cost more than high-res on reasoning models

OpenRouter benchmarked image detail settings across five OpenAI and Google models on MMMU-Pro Vision. On gpt-5.5, low detail scored 65.2% vs 79.0% on auto, yet cost 5.1¢ per question vs 4.5¢—the model burned 1.6× more reasoning tokens trying to read blurry inputs, wiping out input savings. Non-reasoning models gpt-5.4-mini and gpt-4.1 did save money on low, but lost 9.7 and 17.4 accuracy points. Charts and graphs gained the most from auto detail: gemini-3.1-pro jumped from 78.6% to 91.7%. The post recommends sending clear images and dialing down reasoning effort instead.

Why it matters: OpenRouter benchmarked five models on MMMU-Pro Vision and found low-detail images make reasoning models more expensive—gpt-5.5 lost 14 points of accuracy and cost 13% more per question. Counterintuitive result backed by solid data, directly actionable for anyone tuning API cos...

Jul 1Wednesday

AI HOT (Curated Pool)

Meituan releases LongCat-2.0: a 1.6T-parameter model trained on 50,000 domestic GPUs, now open source

Meituan open-sourced LongCat-2.0, a 1.6T total-parameter model with ~48B activated per inference and native 1M context. It was trained and served entirely on a 50,000-card domestic GPU cluster. The architecture combines LSA sparse attention, zero-compute experts, ScMoE, and MOPD multi-expert fusion that blends Agent, Reasoning, and Interaction expert groups. SWE-bench Pro hits 59.5, Multilingual 77.3. A preview is live on OpenRouter and longcat.ai, already ranking top three globally in monthly calls on OpenRouter. The post doesn't disclose training cost, inference latency, or the specific domestic chip model, so I'd hold off on those details.

Why it matters: Meituan's trillion-param model trained end-to-end on domestic GPUs is the headline; code benchmark scores are solid. Not scoring higher because Meituan isn't a tier-1 model lab yet, and real-world usability depends on the open-source release.

Jun 30Tuesday

AI HOT (Curated Pool)

Meituan's LongCat Owl Alpha tops OpenRouter, a 1.6T MoE trained entirely on Chinese ASICs

Meituan LongCat's Owl Alpha became the most popular model on OpenRouter, consuming 10 trillion tokens so far. It's a 1.6T-parameter MoE trained on 35T tokens, running entirely on 50,000 Chinese ASICs. Performance is rated at Gemini/Opus 4.6 level, ranking #1 on Hermes Agent, #2 on Claude Code, and #3 on OpenClaw. The model will retire soon; no details on the next version yet.

Why it matters: Hits three high-signal zones at once: large-scale domestic ASIC training (50K chips), #1 on OpenRouter by usage (10T tokens burned), and claimed Gemini/Opus 4.6 parity. 1.6T MoE params and 35T training tokens are hard numbers, not marketing fluff. Only knock: the post doesn't ...

Jun 25Thursday

AI HOT (Curated Pool)

OpenRouter ships an MCP server so coding agents can query live model pricing and benchmarks

OpenRouter turned its model catalog, benchmarks, pricing, and docs into an MCP server that coding agents like Claude Code and Cursor can call directly. Instead of guessing from stale training data, your agent can query live pricing, fire test prompts, and search docs. The server is remote; first login mints a dedicated key with a 7-day expiry and a $10 spend cap. Tools include filtering models by price and context length, fetching full model details, and comparing responses and costs across models for the same prompt. The post doesn't say whether the MCP server itself is free or paid.

Why it matters: OpenRouter turned its model directory into an MCP tool — a practical update for devs who do model selection inside coding agents. The $10 default cap and 7-day key expiry are concrete safety details, not vaporware. But it's a toolchain optimization, not a model capability brea...

Jun 18Thursday

Hacker News front page

OpenRouter ran 11 LLMs in a 30-game battle royale — Grok 4.1 Fast won 43%

OpenRouter's Jacky Liang dropped 11 LLMs into a 2D battle royale for 30 matches. Grok 4.1 Fast won 13 games at $0.97 per win; Claude Sonnet 4.6 won 5 at $26.78 per win — a 27x gap. GPT 5.4 had the most kills (38) but only 2 wins, so killing more didn't mean winning more. GPT 5.4-mini, DeepSeek 4 Flash, and Kimi K2.6 spent $57 combined and won zero games. The models reasoned, called tools, and updated memory each turn — they weren't just generating control code. The post doesn't provide the full leaderboard or detailed behavioral differences across all models.

Why it matters: OpenRouter's official blog, author Jacky Liang ran 30 games himself with full data and replays. Grok 4.1 Fast's cost advantage is stark, Claude Sonnet 4.6 is expensive but consistent, GPT 5.4 is the kill leader but can't close — all three takeaways are concrete and verifiable....

Jun 16Tuesday

AI HOT (Curated Pool)

OpenRouter's Subagent tool lets frontier models delegate routine tasks to cheaper workers

OpenRouter launched a server-side tool called Subagent. Add openrouter:subagent to your tools array and your orchestrator model can hand off mechanical work—summarization, data extraction, boilerplate, reformatting—to a smaller, cheaper worker mid-generation. Claude Opus 4.8 costs $5 per million input tokens; GLM 5.2 costs $1.40, a 3.6x spread. In a 20-tool-call agent workflow, 5–8 calls might be delegations, cutting per-request cost without touching reasoning quality. Each delegation is isolated: the worker sees only the task_description, no parent context or memory. Workers can carry their own tools like web_search, recursion is blocked, and delegations cap at 10 per request. OpenRouter also highlighted the Advisor tool, which escalates hard decisions upward to a stronger model. The two can be used together in a single request.

Why it matters: OpenRouter turned sub-task delegation into a server-side tool — not just another API wrapper. The Opus 4.8 vs GLM 5.2 cost comparison ($5 vs $1.4) makes the savings tangible. Deduction: no latency numbers disclosed, and no fallback behavior described when the subagent fails. R...

AI HOT (Curated Pool)

Agentic AI Governance: Your API Key Is a Guardrail

OpenRouter argues that most agentic AI governance stops at frameworks and maturity models, which can't block a retry loop from burning $200 overnight. The API routing layer is the practical enforcement point—every request passes through it, so budget caps, model restrictions, and logging can live there. Deloitte reports only one in five companies has mature governance for autonomous agents; IBM says 97% of orgs with an AI security incident lacked proper access controls. The post offers a 5-minute minimum viable setup and notes enterprise needs like audit trails and human approval flows, but doesn't disclose a product timeline.

Why it matters: OpenRouter tackles agent governance from the API routing layer — more concrete than pure framework talk, with Deloitte and IBM stats backing the argument. Downside: it's a vendor blog with product promotion baked in, and the full argument isn't visible from the excerpt alone.

Jun 12Friday

AI HOT (Curated Pool)

OpenRouter's model fusion panel beats GPT-5.5 and Claude Opus 4.8 on deep research benchmark

OpenRouter launched Fusion, which sends a prompt to multiple models in parallel and has a judge model synthesize the final answer. On 100 DRACO deep research tasks, Fable 5 + GPT-5.5 fused scored 69.0%, beating Fable 5 alone at 65.3%. A budget panel of Gemini 3 Flash, Kimi K2.6, and DeepSeek V4 Pro hit 64.7%—close to Fable 5 at roughly half the cost. The post doesn't disclose added latency or the exact per-call price for the budget panel.

Why it matters: OpenRouter's Fusion lets budget model panels beat solo frontier models on deep research via multi-model deliberation + judge. Concrete DRACO benchmark data and anti-cheat design make it worth reading. Score capped at 78 because it's a platform feature launch, not a model break...

Jun 10Wednesday

AI HOT (Curated Pool)

OpenRouter Launches Advisor Tool for Low-Cost Models to Consult Stronger Models

OpenRouter released the Advisor server tool, letting GPT-4o Mini consult Claude Fable during generation, but the post does not disclose pricing, latency, or the routing policy.

Why it matters: HKR-H/K/R all pass: OpenRouter turns cheap-model plus strong-model advising into a callable server tool. Price, latency, and call policy are not disclosed, so this stays in the upper mid-weight product-update band.

Jun 4Thursday

AI HOT (Curated Pool)

OpenRouter compares 11 LLMs for real-time decisions: Claude and Grok lead

OpenRouter spent $482 on inference to run 11 LLMs through a 30-round real-time decision challenge, where Claude and Grok models led on decision speed and task success, while several high benchmark models underperformed on real-time scheduling.

Why it matters: HKR-H/K/R all pass: the contest format is clickable, the post gives cost and round counts, and agent model choice is a real practitioner concern. It is still an OpenRouter-run experiment, not a model release or standard benchmark.

Jun 2Tuesday

AI HOT (Curated Pool)

The Thriving Ecosystem of Open Models

OpenRouter data shows open-weight models generated 69.1% of token usage since 2025, versus 30.9% for closed models, while share leadership shifted across DeepSeek, MiniMax, Kimi, MiMo, Qwen, Tencent Hy3, Alibaba, and Arcee releases.

Why it matters: HKR-H comes from the 69.1% vs 30.9% contrast, HKR-K has OpenRouter token-share data, and HKR-R hits open-vs-closed competition. It is a data-backed commentary, so featured low band.

May 31Sunday

r/LocalLLaMA

Cost Analysis of My $6.4k Local LLM Server

The author runs Qwen3.6 27B on a $6,406.45 local server with 4 MI100 GPUs, processing 20.4M input tokens and 1.32M output tokens per day; using OpenRouter prices, the first-year local cost is $2,992.72 versus $3,701.10 for API use.

Why it matters: HKR-H/K/R all pass: a first-person local-LLM cost test gives hardware, token volume, and API comparison. Single Reddit post and workload-specific economics keep it in the lower featured band.

May 30Saturday

AI HOT (Curated Pool)

OpenRouter supports model-generated file patches

OpenRouter now supports apply_patch, a server-side tool that lets any model propose file edits through the Responses API using V4A diffs, covering file creation, updates, and deletion, with OpenRouter validating diff syntax on the server.

Why it matters: HKR-H/K/R pass: the OpenRouter update gives coding agents a concrete cross-model patch path with V4A diffs and server validation. It is useful infra, not a model-level release, so it sits low in the 72–77 band.

May 29Friday

Xinzhiyuan · WeChat

Three DeepSeek Models Enter OpenRouter Monthly Top 10 With Over 17 Trillion Tokens

DeepSeek placed three models in OpenRouter’s monthly top 10 with more than 17 trillion tokens combined, including V4 Flash at 9.13T tokens; the article says Ascend’s MegaMoE operator raised Prefill throughput by 20% to 30% on DeepSeek V3.1 and Qwen3-235B tests.

Why it matters: HKR-H/K/R all pass: the story has a 17T-token hook plus concrete OpenRouter and MegaMoE Prefill numbers. It stays at 82 because the compute-sovereignty framing is strong, while reproducible test conditions are not disclosed.

May 28Thursday

AI HOT (Curated Pool)

OpenRouter Raises $113 Million Series B

OpenRouter announced a $113 million Series B led by CapitalG, with participation from NVentures, ServiceNow Ventures, Andreessen Horowitz, and Menlo Ventures.

Why it matters: HKR-H/K/R pass: OpenRouter is a known model-routing gateway, and $113M led by CapitalG is concrete. The post lacks valuation, revenue, and traffic metrics, so this stays in the lower featured band.

AI HOT (Curated Pool)

Grok Build 0.1 on API

xAI released Grok Build 0.1 in public beta through the xAI API for agentic coding tasks, with throughput above 100 tokens per second and pricing at $1 per million input tokens and $2 per million output tokens.

Why it matters: HKR-H/K/R all pass, but this is a 0.1 public-beta API and pricing launch; benchmarks, context window, and task success rates are not disclosed. It fits a solid mid-weight product update at 78, featured not p1.

May 27Wednesday

Xinzhiyuan · WeChat

OpenRouter processes 100 trillion tokens monthly and raises $113M Series B

OpenRouter raised a $113 million Series B led by CapitalG, lifting its valuation to $1.3 billion; the platform processes 25 trillion tokens per week, about 100 trillion per month, and provides one API for more than 400 models.

Why it matters: HKR-H comes from the 100T-token/month hook; HKR-K has funding, valuation, usage, and model-count numbers; HKR-R maps to routing and API-cost competition. Still, this is infra funding news, not an 85+ must-write release.

Latent Space

[AINews] New AI Infra Decacorns: Fireworks, Baseten, with OpenRouter on the Way

Latent Space says Fireworks is in talks for a $15 billion valuation round, Baseten is raising at an $11 billion valuation, and OpenRouter closed a $113 million Series C after volume grew 5x in six months.

Why it matters: HKR-H/K/R all pass: the decacorn hook is clickable, the post gives valuation, round, and usage figures, and the topic speaks to inference economics. Fireworks and Baseten are still reported as in talks or raising, so this stays in the 78–84 band.

TechCrunch · AI

OpenRouter more than doubles valuation to $1.3B in a year

OpenRouter raised a $113 million Series B led by CapitalG, lifting its valuation to $1.3 billion within a year; the RSS snippet discloses 5x usage growth over six months but does not disclose revenue, pricing, or customer mix.

Why it matters: HKR-H/K/R all pass, but this is funding news rather than a model or capability release. OpenRouter’s $1.3B valuation and 5x usage growth justify featured at the top of the 72–77 band.