Skip to content

Kimi / Moonshot AI

Moonshot AI's Kimi models and products: the open K series, long context and product changes.

39 picksRelated topicsQwenDeepSeekMiniMax

Latest picks

21–39 of 39

Jul 17Friday

AI Chat-Group Daily (群聊日报)

Kimi K3 tops Frontend Code Arena, weights to open-source, early tests show brilliance and burnout

Kimi K3 hit #1 on Frontend Code Arena with 1679 points, beating Claude Fable 5's 1631 and taking six of seven frontend domains. It packs 2.8T params, 1M context, $3/$15 per million tokens, with full weights opening by July 27. Early testers got mixed results: one user's 199-yuan monthly plan produced stunning particle VJ effects from chat history, while another burned through a $40 coding plan in five hours as the model looped on a domain spelling error. Benchmark trust is shaky—GLM-5.2 scored well on paper but felt worse than 5.5 in practice. Writing style drew split reactions: less AI flavor but forced casual tone, nowhere near the natural Chinese of the old Opus 4.6. Same day, GPT-5.6's frontend taste was called 'very Claude-like,' Sol traced a deadlock only reproducible on Ubuntu, Linus told kernel devs AI is here to stay, and Schema harness pushed ARC-AGI-3 efficiency to 98.98% by making models think like physicists.

Why it matters: Kimi K3 tops Frontend Code Arena, winning 6 of 7 frontend categories with weights opening July 27 — a major domestic flagship release. The chat digest provides scores, params, pricing, and hands-on user feedback. Not scoring higher because the source is a community digest rath...

Bloomberg Technology

China's Moonshot unveils new Kimi model that rivals top US AI on benchmarks, triggering a tech selloff

Moonshot AI released its next-gen Kimi model on July 17, matching or nearing OpenAI o3 and Anthropic Claude Sonnet 4.5 on benchmarks like MATH and HumanEval. The model is available for testing via the Kimi chatbot, with an API already live. Nvidia dropped over 3% premarket on the news, as markets worry about pricing pressure on US AI firms. The post doesn't disclose parameter count, training cost, or inference latency, so I'd hold off on those practical metrics.

Why it matters: Moonshot's new Kimi model benchmarks against o3 and Claude Sonnet 4.5—a domestic flagship release that policy says should be weighted equally with US labs. Bloomberg coverage plus immediate market reaction (Nvidia down >3%) form a cross-source signal. HKR all hit, but the arti...

AI HOT (Curated Pool)

Kimi K3 tops frontend coding leaderboard, open weights coming July 27

Kimi K3 scored 1679 on Frontend Code Arena, taking first in 6 of 7 frontend sub-tasks and beating Claude Fable 5 and GPT-5.6 Sol. It's a 2.8-trillion-parameter MoE model with a 1M context window, and open weights are promised for July 27. API pricing is $15 per million tokens—no low-cost play here, it's priced against top closed-source models and aimed at long-context coding and agent workflows.

Why it matters: Moonshot AI's Kimi K3 tops Frontend Code Arena at 1679, winning 6 of 7 subtasks against Claude Fable 5 and GPT-5.6 Sol. 2.8T MoE params, 1M context window, weights opening July 27. A domestic flagship model directly challenging the closed-source duopoly on a concrete coding be...

Jul 2Thursday

Hacker News front page

Kimi K2.7 Code is now generally available in GitHub Copilot as the first open-weight model option

GitHub Copilot added Kimi K2.7 Code to its model picker—the first open-weight model available. It runs on Microsoft Azure and is billed at provider list pricing under usage-based billing. Rollout starts with Pro/Pro+/Max plans; Business and Enterprise admins must enable it manually. The post doesn't disclose benchmark scores, only that GitHub will monitor quality and performance.

Why it matters: First open-weight model in Copilot's model picker — a real change to developer toolchains. Deployment details (Azure-hosted, usage-based billing at provider pricing) are new info. Score held back because the announcement provides zero benchmark data, so we can't assess code ca...

Jun 15Monday

AI HOT (Curated Pool)

Kimi K2.7 Code high-speed edition live: 5–6× faster output, 2× API price

Kimi released a high-speed variant of K2.7 Code. Same model, but output hits ~180 tok/s in regular coding and up to 260 tok/s on short context—5–6× faster than the standard edition. API price doubles; Kimi Code Plan users pay 3× token consumption. Thinking mode must be on, or it errors out or falls back to K2.6. Compared to K2.6, K2.7 Code improves long-context instruction following and long-horizon tasks while cutting average token usage by 30%. For non-coding work, K2.6 is still recommended. A three-week API top-up promo gives 20–30% vouchers on deposits of ¥500+.

Why it matters: Kimi K2.7 Code Turbo is a substantive product update from Moonshot AI — 5-6x speed boost on the same model, at 2x the price. Hits H and K, misses R. Score stays at the featured threshold because this is an inference acceleration channel, not a new model release, and the mandat...

Jun 9Tuesday

Product Hunt · AI

Kimi launches Kimi Work, a desktop agent that can run up to 300 agents in parallel

Kimi launched Kimi Work on Product Hunt, a desktop agent for knowledge work. It reads local files, automates browsers via WebBridge, runs scheduled tasks, and can spin up to 300 agents in parallel for heavy jobs, outputting PPT, Excel, Word, or PDF. The post doesn't disclose K2.6 model specs or pricing, only mentions a free option.

Why it matters: Moonshot added a desktop agent tool to Kimi. The WebBridge plugin and 300-agent cluster are real new mechanisms, not a wrapper. But the info comes from a Product Hunt page with no hands-on data, and the post doesn't explain task coordination or error handling in cluster mode, ...

May 19Tuesday

AI HOT (Curated Pool)

Kimi's Latest Funding Adds State Capital and Central SOEs, Valuation Quadruples in Six Months

Moonshot AI’s Kimi is raising $2 billion, with Guozhitou and China Mobile added to the shareholder list; in January and February, Kimi completed three funding rounds totaling more than $3.9 billion.

Why it matters: HKR-H/K/R all pass: Kimi is a top Chinese model player, with a reported $2B raise, 4x valuation jump, and Guozhitou/China Mobile entering. Because the round is still in progress, it stays below a completed major launch or IPO.

AI HOT (Curated Pool)

Cursor releases Composer 2.5, calling it its strongest model yet

Cursor released Composer 2.5, claiming a 10x efficiency gain at comparable capability, with larger training scale, more complex reinforcement-learning environments, and a text-feedback mechanism.

Why it matters: Cursor Composer 2.5 is a substantive model update for a front-line AI coding tool, with HKR-H/K/R from the 10x efficiency and RL-training details. The single social-source summary lacks benchmarks, pricing, and reproducible tests, keeping it in the 78–84 band.

AI HOT (Curated Pool)

Cursor releases Composer 2.5 coding model

Cursor released Composer 2.5, claiming up to 10x higher efficiency on long coding tasks; the model is further trained on Moonshot’s Kimi K2.5 and uses text feedback for 100k-token-scale trajectories.

Why it matters: HKR-H/K/R all pass: Cursor is a core AI coding tool, and Composer 2.5 adds concrete claims around 10x long-task gains and Kimi K2.5 tuning. Limited sourcing and no independent eval keep it in the 78–84 band.

May 18Monday

AI HOT (Curated Pool)

Composer 2.5 release and technical analysis

Cursor released Composer 2.5, built on a Moonshot open-source checkpoint, trained with synthetic data from real codebases at 25 times the previous scale, and updated with text-feedback reinforcement learning and a sharded Muon optimizer.

Why it matters: HKR-H/K/R all pass: Cursor is a core coding-agent surface, and the post gives concrete training details around Moonshot, 25x data, RL, and Muon. It lacks benchmarks, pricing, or user-facing capability limits, so it stays in the 78–84 band.

May 17Sunday

AI HOT (Curated Pool)

Latest Open Artifacts #21: Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1, and More

Open AI model teams released Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1, and other versions this month, and the post says they were tested under CAISI’s V4 evaluation framework, but the RSS snippet does not disclose scores.

Why it matters: HKR-H/K/R all pass: a dense open-model roster, a named CAISI V4 evaluation frame, and clear practitioner relevance for model choice. Missing scores and reproducible detail keep it in the 78–84 band.

May 14Thursday

AI HOT (Curated Pool)

Kimi launches Web Bridge browser extension for multi-platform interaction

Kimi launched the Web Bridge browser extension, which lets agents search, scroll, click, type, and complete website tasks, with support for Kimi Code CLI, Claude Code, Cursor, Codex, and Hermes.

Why it matters: HKR-H/K/R all pass: the product hook, action list, and workflow relevance are clear. Kept in the 72–77 band because this is a mid-weight tool update, not a model release, and safety or performance details are not disclosed.

May 13Wednesday

AI HOT (Curated Pool)

90% of People Are Wasting Tokens

Andrej Karpathy says 90% of AI coding bills is wasted on unnecessary context, including repeated full-repository sends, expensive models for simple tasks, and missing prompt caching.

Why it matters: HKR-H/K/R all pass via the 90% claim, named waste mechanisms, and practitioner cost pain. It reaches featured, but stays at 72 because the post gives no billing sample or reproducible test.

May 7Thursday

Bloomberg Technology

Kimi Chatbot Maker Moonshot AI Valued at $20 Billion in Meituan-Led Round

Moonshot AI raised about $2 billion, reaching a $20 billion valuation. The title says Meituan led the round; the post does not disclose investors, stake size, or use of funds. It signals strong demand for Chinese AI startups.

Why it matters: Bloomberg reports Moonshot AI raised about $2B at a $20B valuation, a major capital event for a Chinese model lab. HKR-H/K/R all pass; investor details and use of funds are not disclosed, so this sits in the lower 85–94 band.

May 1Friday

r/LocalLLaMA

MiMo-V2.5-Pro: the actual best open-weights model

Reddit user cjami benchmarked Xiaomi MiMo-V2.5-Pro in autonomous Blood on the Clocktower games. It scored 88% as Good and 48% as Evil, with 183,639 output tokens per game, $0.99 cost, and a 0.4% tool-call error rate. The key comparison is Kimi K2.6: 580,000 tokens, $2.65, and 10–15 hours per game.

Why it matters: Single Reddit benchmark limits authority, so this is not a model-release story. HKR-H/K/R all pass via a named test with win rates, token counts, cost, and tool-error data, placing it in the 78–84 featured band.

r/LocalLLaMA

32x AMD MI50 32GB runs Kimi K2.6 at 9.7 t/s TG and 264 t/s PP

Reddit user ai-infos ran Kimi K2.6 int4 on 32 AMD MI50 32GB GPUs, reaching 9.7 tok/s TG on 136 output tokens. PP hit 263 tok/s on 14,564 input tokens using vllm-gfx906-mobydick across two 16-GPU nodes over 10G Ethernet. Power was about 640W idle and 4,800W peak inference; PCIe bandwidth and the vLLM distributed stack are the real bottlenecks.

Why it matters: HKR-H/K/R all pass via an unusual 32x MI50 build with concrete throughput, power, and network conditions. It stays in the 72–77 band because it is a niche Reddit benchmark, not a broader product or model release.

Apr 21Tuesday

Latent Space

Moonshot Kimi K2.6 open-weight model refresh aims to catch Opus 4.6

Moonshot released Kimi K2.6, a 1T-parameter MoE with 32B active and 256K context. The post cites 58.6 on SWE-Bench Pro, 4,000+ tool calls, 12+ hour runs, and 300 parallel sub-agents. The key signal is long-horizon agent execution, not only open-model scores.

Why it matters: HKR-H/K/R all pass: Kimi K2.6 has a strong race narrative, concrete model and agent metrics, and direct relevance to open-model builders. The domestic flagship release signal lifts it into P1.

Apr 19Sunday

r/LocalLLaMA

Prefill-as-a-Service: KV Cache of Next-Generation Models Could Go Cross-Datacenter

Moonshot says Kimi Linear makes KV cache transfer practical across datacenters, with a 20x scaled-up model showing 1.54x throughput and 64% lower P90 TTFT. The post describes prefill/decode disaggregation across datacenters and heterogeneous hardware; the cost metric and reproducibility details still require the linked arXiv paper.

Jan 29Thursday

Ruan YiFeng's Weblog

Kimi’s integrated stack vs. Manus’s layered approach

Kimi released the K2.5 model and K2.5 Agent together, with an agent mode already available on its website. The post cites 1,500-step long-horizon actions, up to 100 agents in parallel, and visual coding from design files or web videos; pricing, context window, and API terms are not disclosed. The key point is product shape: not just a model launch, but a bundled model-plus-agent release.

Why it matters: HKR-H lands on the integrated release angle; HKR-K lands on the 1,500-step, 100-agent, visual-programming details; HKR-R lands on the stack-design debate. Missing price, context window, and API terms, plus a commentary source, keep it below p1.