Skip to content

OpenRouter

The OpenRouter model router: new listings, usage rankings and real market signals on what developers pick.

Latest picks

1–20 of 77

Sep 26Saturday

Latent Space

OpenRouter: from Seed to Stripe — with Alex Atallah & Anjney Midha

Stripe acquired model routing platform OpenRouter for $7B. In this episode, co-founder Alex Atallah and investor Anjney Midha trace its path from being dismissed by VCs as 'just a wrapper' to handling over 10 trillion tokens per day. They discuss why model labs spend billions on training yet fail at distribution, how Mistral's price war proved the inference marketplace, and OpenRouter's early failed model fusion experiments that were revived years later. Anjney also explains why Stripe's fraud infrastructure matters strategically for OpenRouter and warns that the next wave of fraud will come from autonomous agents attacking token flows.

Why it matters: Stripe's $7B acquisition of OpenRouter is one of the largest AI infra exits this year. The podcast discloses 10T+ daily tokens and the acquisition price for the first time, with Alex Atallah walking through the full seed-to-exit arc. Score capped slightly because it's a retros...

Sep 24Thursday

AI HOT (Curated Pool)

Kimi K3 is open-weight, not open-source: license, checkpoint, and how to call it

Moonshot AI released Kimi K3 weights on Hugging Face under a custom license that isn't OSI-approved, so it's open-weight, not open-source. The checkpoint is a 2.8T-parameter MoE with 104B active parameters per token, stored in MXFP4. The license allows commercial use, modification, and distribution, but adds two conditions: if you run a Model-as-a-Service business with over $20M annual revenue, you need a separate agreement with Moonshot; if your product exceeds 100M MAU or $20M monthly revenue, you must display 'Kimi K3' on the UI. Internal use and access via official partners are exempt. On OpenRouter the model ID is moonshotai/kimi-k3, accepting text, image, and video input with a 1,048,576-token context window. No free tier.

Why it matters: OpenRouter's license breakdown for Kimi K3 is more substantive than the official announcement, clearly distinguishing 'open-weight' from 'open-source' and flagging the commercial API revenue threshold. But without the actual revenue figure or any hands-on benchmarks, it stays ...

Sep 23Wednesday

AI HOT (Curated Pool)

OpenAI GPT-6 Sol and Luna land on OpenRouter at half the price

OpenRouter just listed two new OpenAI models: GPT-6 Sol and GPT-6 Luna. Pricing is half that of the previous GPT-5.6—Sol at $2/M input and $10/M output, Luna at $0.10/M input and $0.50/M output. On AutomationBench, both beat the prior best score while costing less per task. The post doesn't disclose exact scores or latency figures.

Why it matters: OpenAI's next-gen flagship launch with dual variants and halved pricing is an industry-level event. The post doesn't disclose full benchmarks or context window, but the pricing and AutomationBench leap alone justify featured.

AI HOT (Curated Pool)

Claude Opus 5.5 lands on OpenRouter with better agentic coding and a 20% price cut vs Opus 5

Anthropic released Claude Opus 5.5 on OpenRouter, the first model in the Claude 5.5 series. It beats Opus 5 and Fable 5.1 on agentic coding, knowledge work, and computer use, with a 1M context window. Pricing is $4 per million input tokens and $20 per million output tokens, 20% cheaper than Opus 5. The post doesn't include benchmark scores or latency figures.

Why it matters: Anthropic's flagship Claude Opus 5.5 lands on OpenRouter as the first 5.5-series model, with explicit gains in agentic coding and computer use, plus clear pricing. Hits all three HKR axes — a same-day must-write. Not scoring higher because only the platform announcement is ava...

Sep 22Tuesday

AI HOT (Curated Pool)

NVIDIA Nemotron 3.5 Lightning: a 30B sparse model built for high-frequency agent execution

NVIDIA positions Nemotron 3.5 Lightning as the execution layer in agent workflows—handling frequent tool calls, file reads, and result checks rather than heavy planning. It's a 30B MoE model that activates only ~3B parameters per token, keeping latency and cost low for high-volume calls. It complements, not replaces, Nemotron 3 Ultra. Weights are open, with tool calling and structured output support. Context goes up to 1M tokens, though OpenRouter's standard tier caps at 262K. Worth a look if your agent makes many model calls per run.

Why it matters: OpenRouter's breakdown of NVIDIA's new model is substantive, clearly explaining high-frequency agent calls and MoE architecture choices, but the topic is engineering-focused and lacks an emotional hook — R missed.

Sep 18Friday

AI HOT (Curated Pool)

OpenRouter tested 20 image gen models: cheapest at $0.006, priciest at $0.134

OpenRouter sent the same prompt to 20 image models and read the actual billed cost. GPT Image 2 was cheapest at $0.006 per 1024×1024 PNG; Gemini 3 Pro Image was priciest at $0.134—a 22x spread. Pricing units differ across providers (tokens, megapixels, per image), so side-by-side list prices mislead; generate once and check usage.cost. Five of six models rendered text correctly, including the cheapest. Recraft V4.1 Vector outputs editable SVG at $0.08. The post also details formats, resolution caps, and seed support per model.

Why it matters: OpenRouter ran one prompt through 20 image models and posted the actual bills — the kind of real cost data pricing pages never show. All three HKR axes hit: the headline pulls you in, the billing breakdown is genuinely new info, and it nails a daily pain point for builders. No...

Sep 17Thursday

Hacker News front page

I had Gemini train its own replacement for $9

The author paid Gemini 3.1 Pro $9 to label 4,290 Reddit comments with knife brands, models, and steels, then fine-tuned GLiNER large v2.5 on those labels. The resulting model runs locally and hits 0.83 F1 against Gemini's labels after 24 minutes on a Tesla T4. Zero-shot GLiNER scored roughly 0.65 F1. The hardest bug was words_mask: the docs suggest a binary mask, but it's actually a word index; filling it with ones kept loss flat at 70. Five of ten runs produced no usable model—three config failures, two from words_mask. The post doesn't report human accuracy on Gemini's labels, so 0.83 is measured against Gemini, not ground truth.

Why it matters: A hands-on fine-tuning case with real numbers: $9 to distill Gemini labels into a local GLiNER model, lifting F1 from 0.65 to 0.83. Score capped because the domain (knife NER) is narrow, but the method transfers.

Sep 15Tuesday

r/LocalLLaMA

CrofAI, self-claimed cheapest inference provider, exposed as an OpenRouter wrapper swapping in cheaper models at up to 20x markup

CrofAI marketed itself as the world's cheapest inference provider but was caught by developer Kendell silently routing API calls to smaller, cheaper models on OpenRouter. For example, requests for Kimi K3 at $2/$10 in/out were actually served by GLM 5.3 Flash, a 13–20x markup. The founder also fabricated a 'greg' model family that simply pointed to existing open models like GLM 5.2 and Qwen 3.5 9B. After the exposé, he denied everything, then claimed a 'team' was taking over, and within hours deleted the website, Twitter account, and subreddit. The post also notes his hardware claims don't add up: Kimi K3 needs at least 802GiB of VRAM even at Q2_K quantization, but the largest RTX Pro 6000 machine on Vast only offers 765GiB.

Why it matters: A full fraud exposé with technical evidence and a dramatic company meltdown. Hits all three HKR axes. Score capped below 85 because the source is a Reddit post, not formal reporting, and the event is a single-provider scandal rather than a model or protocol shift.

Sep 10Thursday

AI HOT (Curated Pool)

OpenRouter launches Fusion: a compound model that debates across models before synthesizing a final answer

OpenRouter Fusion is a compound inference pipeline, not a new model. It fans out one prompt to up to 8 panelist models in parallel, has a judge compare their answers for consensus and blind spots, then lets the calling model write a final synthesis. A default three-model panel costs roughly 4–5× a single completion and takes 2–3× longer. On the DRACO deep-research benchmark, a budget panel of Gemini 3 Flash, Kimi K2.6, and DeepSeek V4 Pro scored ~64.7%, close to Claude Fable 5’s solo 65.3%. OpenRouter’s own test paired two Claude Opus 4.8 runs and saw a 6.7-point gain over a single run. The team positions Fusion as an escalation path for complex research and high-stakes decisions, not for simple chat. The post does not disclose per-token pricing, only the cost multiplier.

Why it matters: Fusion is a multi-model debate-and-synthesize workflow, not a new model. Concrete cost/latency numbers and DRACO benchmark data give it substance beyond marketing. But it's a routing-layer product update, not a foundation-model breakthrough — capped at the low end of featured,...

Sep 9Wednesday

AI HOT (Curated Pool)

OpenRouter reviews Seedance 2.5: strong at long takes and editing, but no 1080p

OpenRouter published a hands-on review of Seedance 2.5 on Sep 9. The model went live Aug 7, 2026, and excels at 30-second single takes and editing from existing footage. Cost is ~$0.103/sec at 480p and ~$0.231/sec at 720p; using a video reference cuts the token price by ~40%. Audio generation adds no extra charge. The clear trade-off: it caps at 720p. For 1080p or 4K, you still need Seedance 2.0 or Veo 3.1. Frame-exact reproducibility is also not guaranteed.

Why it matters: OpenRouter's hands-on review of ByteDance's Seedance 2.5 delivers concrete pricing ($0.103/sec at 480p, $0.231/sec at 720p, 40% off for video-reference tokens, free audio) and a clear resolution cap at 720p. Solid product intel but narrow audience fit, landing right at the fea...

Sep 8Tuesday

AI HOT (Curated Pool)

OpenRouter launches shell sandbox and Files API so any model can run commands in a hosted Linux container

OpenRouter added a server-side shell tool and Files API so any model can run commands inside a hosted Linux container. Sandbox time costs $0.0001 per second, billed with the request. Network is off by default; you can enable it with an allowlist. The Files API handles uploading inputs and downloading outputs. The shell tool supports both OpenAI and Anthropic tool specs—set engine: openrouter to force server-side execution. The post doesn't disclose container resource limits or max runtime per invocation.

Why it matters: OpenRouter added a managed shell sandbox and Files API for all models, letting them execute commands, read errors, and retry scripts autonomously. Per-second billing and network-off-by-default make it credible in the agent toolchain. Not scoring higher because this is a platfo...

Sep 5Saturday

Hacker News front page

OpenAI GPT-6 Astra lands on OpenRouter, built for long-horizon agentic work

OpenAI's new flagship GPT-6 Astra is now listed on OpenRouter, released Sep 4, 2026. It's positioned for demanding end-to-end work: advanced analysis, software engineering, deep research, science, and document creation, with a stated strength in long-horizon agentic tasks involving computer and browser use. Pricing is $10/$50 per 1M tokens, 1M context window. The fastest provider on OpenRouter is OpenAI Fast at 2.10s latency but $20/$100; the best value is OpenAI Flex at $5/$25 with 2.72s latency and 56 tps throughput. The post does not disclose benchmark scores or comparisons to other models.

Why it matters: OpenAI's flagship GPT-6 silently landing on OpenRouter is an industry-shaking event. Clear positioning for long-running agent tasks, with concrete pricing and context window numbers — high information density. Deduct 4 points because only the OpenRouter page is available so fa...

Sep 3Thursday

Computing Life · Share · Yage

Agent token usage 5× human, but caching discounts cut the real bill to ~2×

OpenRouter data shows agents consume 7.3T tokens weekly, nominally 5.2× human usage. But 70–85% are cached reads; with ~90% discount, the real bill is roughly 2×. GitHub's Knowledge Compressor prototype halves doc length and claims breakeven at 2,000 reuses, but factoring in caching pushes the median to 5,000+. OpenAI's Jalapeño chip beats Nvidia GB200/GB300 on fixed-length benchmarks, yet lacks AgentX scores for real agent workloads. All three stories share one distortion: prompt caching inflates headline numbers.

Why it matters: Three stories bundled, but the core value is the first: someone finally separated nominal agent token consumption from the caching-discounted real cost, landing at ~2x. The OpenAI chip benchmark and GitHub compression prototype are bonuses but less dense. Cross-source cluster ...

Sep 2Wednesday

AI HOT (Curated Pool)

Claude Fable 5.1 is live on OpenRouter, targeting agentic coding and long-running workflows

Anthropic released Claude Fable 5.1 on OpenRouter as a direct upgrade to Fable 5. The focus areas are agentic coding, long-running workflows, visual code generation, finance, and analytics. The post doesn't disclose benchmark numbers or pricing changes, so I'd wait for third-party evals.

Why it matters: Anthropic model update with four clear focus areas, directly relevant to Claude developers. But no benchmarks, no pricing, no third-party evals in the post — stays at 78, the featured threshold, pending real-world testing.

Aug 27Thursday

AI HOT (Curated Pool)

Tang Jie announces GLM-5.3 Flash AA tops OpenRouter, running on domestic chips

Tang Jie posted that GLM-5.3 Flash AA (codename Ox Alpha) scored 57 on OpenRouter at 1/100th the price of frontier models. It runs entirely on domestic Chinese chips and captured nearly 20% of weekly token share, ranking first. The post doesn't disclose the chip model, benchmark details, or comparison targets.

Why it matters: Zhipu's GLM-5.3 Flash AA hit #1 on OpenRouter, with Tang Jie posting three hard numbers: score 57, ~20% weekly token share, 1% cost, plus a claim of running on domestic chips. HKR all hit, but the post doesn't name the benchmark, comparison models, or chip model — those gaps k...

Aug 26Wednesday

TechCrunch · AI

Z.ai confirms it built Ox Alpha, the anonymous model topping leaderboards

Z.ai confirmed it is the lab behind Ox Alpha, the open-weight model that appeared anonymously on OpenRouter and immediately topped rankings. The company calls it the newest GLM iteration, built for coding, sustained agentic work, and multimodal reasoning. Weights drop Wednesday for developers to build on. Earlier GLM-5.3 already matched Anthropic's Fable 5 on some benchmarks. Ox Alpha adds more pressure on frontier pricing from OpenAI and Anthropic.

Why it matters: Revealing the identity of a chart-topping anonymous model is inherently newsworthy; Z.ai also commits to open-sourcing weights on Wednesday and clearly positions the model for code, agents, and multimodal reasoning. The score is held back because the article provides no benchm...

Computing Life · Share · Yage

The term 'local LLM' conflates two separate markets

Yage breaks down 'local LLM' into two markets: a cost market buying 5–20× price gaps, and a control market buying 25–33-year certainty. Using a four-quadrant framework (open/closed weights × time/token billing), the piece explains why surging open-weight model usage on OpenRouter doesn't mean local deployment is winning. Self-hosting payback depends entirely on which cloud billing mode you replace—decades for subscriptions, months for high-cache-hit agent API calls. In July–August 2026, Anthropic and others made four moves at the inference layer: silently remapping parameters, repeatedly extending usage boosts, adding watermarks, and redefining self-hosting as 'your harness plus my inference.' But simultaneous deep price cuts mean the misalignment is real but direction is unresolved.

Why it matters: Splits 'local LLM' into cost vs control markets with OpenRouter data and hardware payback math — directly useful for infra decision-makers. Not scored higher because it's commentary rather than a product launch or research breakthrough, but hits all three HKR axes and earns a ...

Aug 25Tuesday

Hacker News front page

CTGT behaviorally fingerprints Ox Alpha to GLM-5.x lineage, finds censorship is a switch not a tilt

CTGT traced the anonymous model Ox Alpha, which appeared on OpenRouter on Aug 20, to Zhipu's GLM-5.x family via an 11-of-11 tokenizer match and parameter checks. On sensitive topics, Ox Alpha answers like an American model on Xinjiang and Taiwan but flatly refuses 7 domestic-risk topics including Xi Jinping personally—a blacklist-style censorship distinct from DeepSeek's pervasive softening. Three refusals still billed completion tokens; one later returned a full answer on retry.

Why it matters: Anonymous model provenance + censorship split analysis hits all three HKR axes. The 11/11 tokenizer match is hard evidence, and the split censorship pattern is a genuinely new observation. Slight discount because it's third-party research, not an official release, and the mode...

AI HOT (Curated Pool)

GPT 5.6 discounts drove Terra/Luna token usage up to 13.8x, with ~32% user retention

OpenRouter data shows that during OpenAI's July 27–Aug 14 discount on Terra and Luna, daily Terra tokens rose 5.6x and Luna 13.8x, while the undiscounted Sol model saw only a 1.1x bump. Most of the share gain came from competitors: the OpenAI family's token share grew from 7.1% to 12.4%, with roughly three-quarters taken from outside labs. After the discounts ended, about 32% of the 100K+ users who tried Terra/Luna kept using them, and 18% ran at or above their discount-period pace. Sol later reproduced the same spike when it got its own 50% discount on Aug 17.

Why it matters: First-party OpenRouter data showing market displacement after GPT 5.6 price cuts, with concrete multipliers and share shifts. Not an 85+ because it's platform analytics rather than a model capability update, but solid enough as a market signal for featured.

Aug 24Monday

TechCrunch · AI

Mysterious reasoning model Ox Alpha sparks frenzy over who built it

A free reasoning model called Ox Alpha appeared on OpenRouter Thursday, described as built for coding and sustained agentic work. Stripe CEO Patrick Collison called it 'very impressive' on X. The listing says it's a 'stealth model' from an anonymous third-party provider. Speculation centers on two theories: an unreleased GLM model from Chinese company Zhipu, or a hidden version of Microsoft's MAI. Reddit and X are split, but the article offers no hard evidence—only community guesses.

Why it matters: Anonymous reasoning model lands with a Patrick Collison endorsement and a clear code/agent focus. Speculation points to Zhipu or DeepSeek — enough signal and mystery to matter. Held at 78 because all info is external guesswork; the post didn't confirm the developer.