Skip to content

OpenRouter

The OpenRouter model router: new listings, usage rankings and real market signals on what developers pick.

Latest picks

61–77 of 77

May 26Tuesday

AI HOT (Curated Pool)

OpenRouter Raises $113M Series B

OpenRouter raised a $113 million Series B led by CapitalG; its weekly volume rose from 5 trillion to 25 trillion tokens over the past 6 months.

Why it matters: HKR-H/K/R all pass: OpenRouter is a common model-routing layer, and $113M plus 25T weekly tokens gives hard scale. Kept below 85 because this is funding plus growth data, not a new model or capability release.

May 19Tuesday

AI HOT (Curated Pool)

OpenRouter Tool-Calling Models Can Now Run Web Search Autonomously

OpenRouter now lets any tool-calling model on its platform autonomously invoke web search and webpage scraping, with the model deciding when to search, what to query, and how many searches to run; OpenRouter also added @p0 as a web search provider.

Why it matters: HKR-H/K/R pass: OpenRouter lets tool-calling models decide search timing, queries, and frequency. The source is tweet-thin and lacks pricing, limits, or evals, so it lands near the featured threshold.

May 17Sunday

AI HOT (Curated Pool)

Ring-2.6-1T Open-Sourced and Listed on OpenRouter for Agent Workflows

AntLingAGI open-sourced Ring-2.6-1T and listed it on OpenRouter with a 75% discount through the end of May; the trillion-scale reasoning model targets agent workflows, including planning, tool use, context maintenance, and complex task execution, using Async RL and IcePop training methods.

Why it matters: HKR-H/K/R all pass: a 1T open agent model is clickable, with OpenRouter access, discount, and training methods disclosed. Score stays at 74 because benchmarks, license, and context window are not given.

May 15Friday

r/LocalLLaMA

Evaluated a RAG Chatbot: The Most Expensive Model Was the Worst Performer

The author evaluated a customer-support RAG bot and raised the quality score from 6.62 to 7.88 while cutting per-session cost from $0.002420 to $0.000509, using retrieval logging, LLM-as-judge scoring, chunk deduplication, stricter grounding, and a five-model sweep.

Why it matters: HKR-H/K/R all pass: counterintuitive model ranking, concrete quality and cost deltas, and direct RAG production relevance. Reddit source authority keeps it near the featured floor despite the first-person experiment signal.

May 12Tuesday

r/LocalLLaMA

I catalogued every way local models break JSON output and built a repair library across 288 model calls

Reddit user kexxty ran 288 structured-output calls through OpenRouter models, including Llama 3, Mistral, Command R, DeepSeek, and Qwen, and found similar JSON failure categories across local and API-only models. The MIT-licensed Python library outputguard validates against JSON Schema, applies 15 ordered repair strategies, includes 2,001 tests, and has no LLM provider dependency.

Why it matters: HKR-H/K/R all pass: 288 tests, the outputguard library, and a 15-step repair chain give practitioners reusable detail. Source is a single Reddit post, so it stays in the 72–77 featured band, not 78+.

May 11Monday

AI HOT (Curated Pool)

Pareto Code Reorders Model Selection Using Market Demand

OpenRouter says Pareto Code observes the Pareto frontier using real market demand; DeepSeek V4 Pro ranks first, followed by GPT 5.4 Mini and Gemini 3.1 Pro, while the post does not disclose the scoring formula or evaluation sample size.

Why it matters: HKR-H/K/R all pass, but the source is a single OpenRouter post with no sample size, time window, or pricing basis disclosed. It clears featured as a model-selection benchmark, not the 78+ band.

AI HOT (Curated Pool)

AntLingAGI Releases Trillion-Parameter Ring-2.6-1T Model

AntLingAGI released Ring-2.6-1T, a trillion-parameter thinking model available for free on OpenRouter until May 15, with adjustable thinking intensity, agent-oriented multi-step execution, tool calling, and tasks covering math logic and scientific research.

Why it matters: HKR-H/K/R all pass, but the post is thin: no benchmarks, pricing, architecture, or training details. Treat as a mid-weight model launch on OpenRouter, not a same-day must-write.

May 7Thursday

AI HOT (Curated Pool)

Trillion-parameter instruction model Ling-2.6-1T released

inclusionAI says Ling-2.6-1T is now live on OpenRouter. The trillion-parameter instruction model uses “fast thinking” and claims top AIME26 and SWE-bench Verified results with about 75% lower cost. The post does not disclose pricing, context length, or full benchmark scores.

Why it matters: HKR-H/K/R all pass: a 1T instruction model on OpenRouter with fast thinking, AIME26/SWE-bench claims, and ~75% cost reduction. Missing price, context window, and full scores keep it in the 78–84 band.

AI HOT (Curated Pool)

Consistent web search and scraping for all models

OpenRouter released tools for tool-calling models to run web search and page scraping. The post says multiple search and scraping engines are supported, but does not disclose names, pricing, or limits. The key item is cross-model tool interface consistency.

Why it matters: HKR-H/K/R all pass, but engines, pricing, and limits are not disclosed. This is a mid-weight Product update: useful for model-agnostic agent stacks, not a major model or capability release.

r/LocalLLaMA

Analysis of 922 Agentic Task Traces Finds DeepSeek v4’s Cost Edge in Caching

A Reddit user analyzed 922 agentic task traces and reported $0.01 per task for DeepSeek v4 Flash versus $1.52 for Opus 4.7. Both used about 960K tokens per task, but DeepSeek showed a 97% cache hit rate versus 87%, with a 0.02 cache read/write price ratio versus 0.08. The key issue is caching, not headline pricing.

Why it matters: HKR-H/K/R all pass: 922 agent traces tie a large cost gap to cache hit rate and cache read/write pricing. Reddit single-source data and incomplete method detail keep it in the 78–84 band.

Apr 30Thursday

r/LocalLLaMA

Actual comparison between locally run Qwen-3.6-27B and proprietary models

The author compared 5 model setups on an autoresearch-loop task; only Qwen-3.6-27B via OpenRouter nearly solved it. The local q4_k_m run took about 8 hours and used 39k/45k tokens; full-quality Qwen used 4.4M tokens and cost $0.939. The useful signal is failure quality: both Qwen runs needed small fixes, while Gemma, Codex-Spark, and Claude Haiku 4.5 missed tests or key logic.

Why it matters: HKR-H/K/R all pass: the post has a concrete agent-test surprise, token and cost data, and local-vs-proprietary tension. Single Reddit run limits source authority, so it stays in the lower featured band.

Apr 29Wednesday

NVIDIA Blog

NVIDIA Launches Nemotron 3 Nano Omni for Vision, Audio, and Language Agents

NVIDIA launched Nemotron 3 Nano Omni, claiming up to 9x higher throughput at the same interactivity. It uses a 30B-A3B hybrid MoE with Conv3D, EVS, and 256K context, taking text, images, audio, video, documents, charts, and GUIs as input. Open weights, datasets, and training methods arrive April 28, 2026 on Hugging Face, OpenRouter, build.nvidia.com, and 25+ platforms.

Why it matters: HKR-H/K/R all pass: NVIDIA’s open multimodal model has a 9x efficiency claim, 30B-A3B MoE, and 256K context. Single-vendor sourcing keeps it in the good-quality band, below must-write.

Apr 21Tuesday

QbitAI · WeChat

Mystery model Elephant: 100B parameters reaches same-scale SOTA with high token efficiency

Ant Group's Inclusion AI team is identified as the maker of Elephant, a 100B-parameter model with 256K context and 32K output shown on OpenRouter. The post reports tests on bug fixing, summarizing a 3,000-word meeting note, and a light agent loop, plus AI BENCHY figures of about 2,500 output tokens, about 1 second average latency, and 9.6/10 consistency; the post does not disclose training details, pricing, or an official model card.

Why it matters: HKR-H/K/R all pass: a 100B model posting same-scale SOTA with token efficiency is a strong hook, and the piece includes 256K/32K, ~1s latency, 9.6/10 consistency, plus failure cases. It stays below p1 because training details, pricing, and an official model card are not disclosed

Apr 20Monday

r/LocalLLaMA

Compared some models for feature planning

A Reddit user tested 9 models on planning a “load tracking” feature for a Go budgeting app, then used Claude Code to rank the generated specs, with Claude Opus 4.6 placed first. The table shows Opus 4.6 produced a 19 KB spec with 44 code reads at $2.47; GLM 5.1 ranked second and Qwen 3.6 35B fp8+vLLM ranked third. Do not treat this as a benchmark: the author says it is not representative, and the post does not disclose any manual quality review yet.

Why it matters: A named first-person test gives real workflow data, so HKR-H/K/R all pass. The ceiling stays low: one task only, ranked by Claude Code itself, and no human acceptance result is disclosed, so this lands at the low end of featured.

Apr 19Sunday

r/LocalLLaMA

I tested 8 LLMs as tabletop GMs: a 27B model beat the 405B on narrative quality

The author tested 8 LLMs on 6 fixed tabletop-GM scenarios, and google/gemma-3-27b-it ranked first in narrative quality with a 4.33 overall score. The probe used 8 auto metrics plus 3 LLM-judge scores, and the full run cost about $0.02; the title says a 27B beat a 405B, but the snippet does not disclose the 405B model name or full rankings.

Why it matters: A named first-person benchmark with a strong surprise hook clears HKR-H, HKR-K, and HKR-R. I kept it at featured, not higher: the source is Reddit, the post is truncated, and the 405B model name plus full ranking are not disclosed.

Apr 16Thursday

Hacker News front page

Darkbloom – Private inference on idle Macs

Eigen Labs launched Darkbloom, linking 100M+ Apple Silicon Macs into a decentralized inference network. It offers an OpenAI-compatible API, claims end-to-end encryption plus hardware attestation, and lists prices up to 70% below OpenRouter comps. The real point is the trust model: hardware keys, hardened runtime, and signed outputs are disclosed, but enterprise audit scope still needs the paper.

Why it matters: HKR-H/K/R all pass: the idle-Mac inference angle is novel, and the post includes concrete scale, API, encryption, and price claims. I keep it at 80 because this is still a self-published research preview; audit scope, network reliability, and attack boundaries are not yet third-p

Apr 3Friday

X · @op7418

Alibaba released the Qwen 3.6 Plus model

Alibaba released Qwen 3.6 Plus with a 1M context window, 64K input, and nearly 991K max output. The RSS snippet says it improves over Qwen 3.5 on agents, coding, image, and document understanding, priced at RMB 2 per 1M input tokens and RMB 12 per 1M output tokens; benchmark scores and test conditions are not disclosed.

Why it matters: Alibaba shipping Qwen 3.6 Plus is a substantive domestic model update. HKR-H/K/R all pass on the 1M-context plus pricing combo, but it stays below P1 because benchmark scores, baselines, and test conditions are not disclosed in the body.