Skip to content

#DeepSeek

0 today

May 9Saturday

r/LocalLLaMA

DeepSeek Rejects Alibaba, Prioritizing Independence Over Big Tech Ecosystems

DeepSeek’s financing talks with Alibaba fell through after both sides failed to agree on terms. The post says DeepSeek was valued at RMB 300 billion and sought RMB 50 billion.

Why it matters: HKR-H/K/R all pass: the DeepSeek-Alibaba split has a strong conflict hook, hard funding numbers, and China AI ecosystem stakes. Reddit single-source uncertainty keeps it below P1.

AI HOT (Curated Pool)

Baidu releases ERNIE 5.1 with compressed parameters and training cost

Baidu released ERNIE 5.1 with total parameters reduced to about one third of the original scale, active parameters to about one half, and pretraining cost to about 6% of same-scale models; the model is available on the ERNIE platform and Baidu AI Studio.

Why it matters: HKR-H/K/R all pass: Baidu ERNIE 5.1 is a domestic flagship-model release with concrete compression and 6% pretraining-cost claims. That puts it in the must-write band.

Synced · WeChat

DeepSeek Reportedly Raises RMB 50B, with Liang Wenfeng Funding 40%, Valuation Reaching RMB 350B

DeepSeek is negotiating a $7.3 billion funding round at an estimated $51.5 billion valuation; Liang Wenfeng reportedly plans to contribute 40%, while Tencent and China’s RMB 60 billion national AI fund are also in talks.

Why it matters: HKR-H/K/R all pass: the DeepSeek funding rumor has large numbers, a founder contribution ratio, and named backers. Because it is still reported as talks with no official confirmation, it stays at 84 and featured, not p1.

AI HOT (Curated Pool)

DeepSeek Raises $7 Billion at Record Scale, Founder Personally Invests $3 Billion

DeepSeek is raising up to $7 billion at a $50 billion valuation, with founder Liang Wenfeng personally contributing $3 billion, or 40% of the round, while the company says the funding will target large-scale compute, V4.1 model releases, enterprise products, and a path toward positive revenue.

Why it matters: HKR-H/K/R all pass: the $3B founder check is a strong hook, with concrete funding numbers and clear competitive resonance. Single X-source sourcing leaves lead investor, terms, and confirmation undisclosed, so it stays below P1.

r/LocalLLaMA

MTP + TurboQuant Running: Qwen3.6-27B Hits 80+ t/s on a Single RTX 4090

indrasmirror ran Qwen3.6-27B-Heretic-v2 on a single RTX 4090 with 262K context, TBQ4_0 KV cache, and MTP draft 3, improving throughput from about 43 t/s to 80-87 t/s with roughly 73% MTP draft acceptance.

Why it matters: HKR-H/K/R all pass, backed by a numbered first-person experiment. The Reddit-only source and niche local-inference focus keep it below the 78–84 band for broader industry releases.

May 8Friday

r/LocalLLaMA

Reports suggest DeepSeek seeks $7.35B in funding and plans V4.1 update next month

DeepSeek is seeking up to RMB 50 billion, about $7.35 billion, in its first funding round; the post says V4.1 is planned for June, but does not disclose model parameters or pricing.

Why it matters: HKR-H/K/R all pass: the report gives a $7.35B target and June V4.1 window for DeepSeek. Confirmation, specs, pricing, and funding status are not disclosed, so this stays at the low end of P1.

QbitAI · WeChat

All Labs Watch ByteDance, Everyone Praises DeepSeek: A U.S. Researcher’s 36-Hour China AI Trip

Ai2 researcher Nathan Lambert visited Zhipu, Moonshot AI, Tsinghua, Meituan, Xiaomi, and 01.AI within 36 hours, and said Chinese labs closely watch ByteDance and respect DeepSeek, while student participation in core work, open source habits, and in-house control of the technical stack mark key differences.

Why it matters: HKR-H/K/R all pass: the piece has a named US researcher’s dense China-lab tour plus concrete claims on ByteDance, DeepSeek, open source, and in-house stacks. It is strong industry field reporting, not a model launch or major deal, so it sits at featured rather than p1.

AI HOT (Curated Pool)

WIRED examines why ChatGPT keeps saying “I’ve got you” in Chinese replies

ChatGPT repeatedly uses phrases like “I’ll steadily catch you” in Chinese chats. WIRED links it to mode collapse, translation mismatch, and RLHF rewards for pleasing replies. Similar phrases appear in Claude and DeepSeek; the post does not disclose sample size.

Why it matters: HKR-H comes from the odd “I’ll catch you steadily” meme; HKR-K names three mechanisms; HKR-R touches alignment and Chinese UX concerns. No sample size is disclosed, so this stays in the lower featured band.

AI HOT (Curated Pool)

DeepSeek 4: Flash Local Inference Engine for Metal

DeepSeek 4 Flash is open-sourced on GitHub for offline inference on Apple Silicon Macs. The post says it uses Metal Performance Shaders to reduce latency and memory use, but discloses no benchmark numbers. The key item is the Metal local inference stack, not another model wrapper.

Why it matters: HKR-H/K/R pass: the hook is offline Apple Silicon inference, with GitHub OSS, MPS, and a clear run target. No latency or memory benchmarks, and not an official DeepSeek model launch, so it stays near the featured floor.

May 7Thursday

r/LocalLLaMA

DeepSeek nears $45bn valuation as China’s Big Fund leads investment talks

DeepSeek is reportedly discussing its first funding round at a valuation near $45 billion. The title says China’s Big Fund leads talks; the post only links TechCrunch, FT, and Bloomberg. The post does not disclose round size, stake, investors, or closing date.

Why it matters: HKR-H/K/R all pass: DeepSeek at nearly $45bn with China’s Big Fund is strong. The post lacks amount, equity stake, investor list, and closing date, so it stays below 85.

r/LocalLLaMA

Analysis of 922 Agentic Task Traces Finds DeepSeek v4’s Cost Edge in Caching

A Reddit user analyzed 922 agentic task traces and reported $0.01 per task for DeepSeek v4 Flash versus $1.52 for Opus 4.7. Both used about 960K tokens per task, but DeepSeek showed a 97% cache hit rate versus 87%, with a 0.02 cache read/write price ratio versus 0.08. The key issue is caching, not headline pricing.

Why it matters: HKR-H/K/R all pass: 922 agent traces tie a large cost gap to cache hit rate and cache read/write pricing. Reddit single-source data and incomplete method detail keep it in the 78–84 band.

TechCrunch · AI

DeepSeek could hit $45B valuation from its first investment round

DeepSeek could reach a $45B valuation in its first investment round, according to the title. The snippet says it rose in early 2025 after training an LLM with far less compute and cost; the post does not disclose round size, investors, or terms.

Why it matters: HKR-H/K/R all pass: DeepSeek’s first round targeting $45B is a strong valuation story. Missing investors, amount, and terms keep it in the lower 78–84 band, not P1.

May 6Wednesday

Financial Times · Technology

Chinese AI start-up DeepSeek nears $45bn valuation

DeepSeek is nearing a $45bn valuation in fundraising talks, with Tencent among investors seeking a stake. The post does not disclose round size, terms, or timeline. The key question is valuation versus model revenue.

Why it matters: HKR-H/K/R all pass: FT reports DeepSeek nearing a $45bn valuation with Tencent interest, a major capital signal for a flagship Chinese AI lab. The deal is not closed, and size, terms, and timeline are undisclosed, so it stays below P1.

Synced · WeChat

DeepSeek Version of Claude Code Tops Trending Chart With 8,700 Stars

DeepSeek TUI topped GitHub trending with over 8,700 stars. Hunter Bown built it in Rust for local terminal use with DeepSeek V4, supporting chat, file edits, shell commands, and task management. The key detail is RLM mode: up to 16 V4 Flash subtasks, plus a 1M-token context window and approval gates.

Why it matters: HKR-H/K/R all pass: the 8,700-star hook is strong, RLM adds concrete mechanisms, and coding-agent competition resonates. It is a third-party open-source tool, not an official DeepSeek model release, so it stays in the 78–84 band.

r/LocalLLaMA

DeepSeek V4 at 17x lower cost prompted a local-vs-cloud coding workflow test

Reddit user spencer_kw logged a 10-day coding workflow and retested 150 tasks on local Qwen 3.6 27B versus cloud models. Local was equivalent for 65% of tasks, acceptable for 20%, and cloud was needed for 15%; the API bill fell from $85/month to about $22. The useful signal is task-based routing, not headline model pricing alone.

Why it matters: HKR-H/K/R all pass: this is a quantified practitioner cost test, not a model launch. The single Reddit sample limits generality, so it lands at the featured threshold rather than P1.

May 5Tuesday

r/LocalLLaMA

DeepSeek V4 Pro matches GPT-5.2 on FoodTruck Bench, 10 weeks later and about 17x cheaper

DeepSeek V4 Pro ranked No. 4 on FoodTruck Bench. The 30-day agentic benchmark uses 34 tools, persistent memory, and daily reflection; its median is within 3% of GPT-5.2 at about 17x lower workload cost. Xiaomi MiMo v2.5 Pro also ranked No. 6, with 5/5 survival, 1,019% median ROI, and $2.41 per run.

Why it matters: HKR-H/K/R all pass: the cost gap is clickable, and the post gives a 30-day, 34-tool setup plus a 17× cost delta. Single-source Reddit benchmark with no cross-validation keeps it in the 78–84 band.

May 4Monday

QbitAI · WeChat

DeepSeek-TUI, a “DeepSeek Claude Code,” reaches 2.3k GitHub stars

DeepSeek-TUI reached 2.3k GitHub stars; the Rust project is MIT-licensed. It targets DeepSeek V4 with a 1M-token context, RLM up to 16 V4 Flash subtasks, MCP, Shell, Git, and three control modes. Watch cache misses: uncached tokens cost 10x cached tokens.

Why it matters: HKR-H/K/R all pass: the hook is a DeepSeek-flavored Claude Code, with 2.3k stars, 1M tokens, 16 subtasks, and a 10x cache-miss cost gap. Impact is developer-specific, so it sits in the 72–77 band.

May 3Sunday

r/LocalLLaMA

Local LLM Benchmark for Backend Generation via Function Calling: GLM vs Qwen vs DeepSeek

AutoBe posted a controlled backend-generation benchmark and says qwen3.5-35b-a3b matches gpt-5.4 on DB/API design. One shopping-mall run uses 200–300M tokens, costing $1,000–$1,500 per model at GPT 5.5 pricing. The key caveat is n=4 projects and self-scoring harness bias.

Why it matters: HKR-H/K/R all pass, but Reddit sourcing, n=4 projects, and self-eval harness bias keep it at the low featured band. Concrete cost and test constraints carry the score.

QbitAI · WeChat

DeepSeek V4’s biggest omission

DeepSeek V4’s technical report omits Engram while listing mHC, CSA, HCA, Muon, and FP4. Engram was open-sourced by DeepSeek and Peking University in January, inserting lookup modules between Transformer layers 2 and 15; its 27B test raised MMLU by 3.4 and Multi-Query NIAH to 97.0%. The engineering signal is CXL pooling: 8 servers shared a 4TB memory pool with under 5% throughput loss.

Why it matters: HKR-H/K/R all pass: the omitted-Engram angle is clickable, with layer ranges, benchmark deltas, and CXL memory-pool numbers. It is analysis, not the V4 launch itself, so 78–84 fits.

May 2Saturday

Hacker News front page

Show HN: Filling PDF Forms with AI Using Client-Side Tool Calling

SimplePDF released a Copilot demo that fills PDF forms via client-side tool calling; SimplePDF has 200k+ monthly users. PDFs stay in the browser, with parsing, rendering, and field detection local. The demo uses a DeepSeek V4 Flash proxy by default, with BYOK, cloud, or LM Studio options.

Why it matters: HKR-H/K/R pass: the client-side PDF-agent angle is specific, with a clear privacy mechanism and builder relevance. It sits in the 72–77 band as a useful product demo, not a major platform release.

May 1Friday

r/LocalLLaMA

16x Spark Cluster Build Update

Reddit user Kurcide finished a 16-node DGX Spark cluster, with all nodes hitting line rate on the fabric. Each node uses one QSFP56 link to an FS N8510, showing 100–111 Gbps per rail and about 200 Gbps aggregate. The key angle is unified memory: 8 nodes served 434GB GLM-5.1-NVFP4, with DeepSeek and Kimi tests next.

Why it matters: HKR-H/K/R all pass: the post gives first-person cluster numbers, networking conditions, and a live 434GB model test. Scope stays local-inference hardware, so it fits the 72–77 band rather than a broader product-release tier.

Synced · WeChat

The Evolution of RL: From PPO to MaxRL in LLM Reasoning Training

Jiqizhixin translated Alexander Weers' article on RL algorithms for LLM reasoning from 2024 to 2026. It covers REINFORCE, PPO, GRPO, RLOO, Dr. GRPO, DAPO, CISPO, MaxRL, DPPO, and ScaleRL, comparing critic removal, clipping, normalization, and pass@k goals. The key signal is mechanism choice, not algorithm names.

Why it matters: A strong technical explainer, not a model or paper release. HKR-H comes from the PPO→MaxRL arc, HKR-K from concrete mechanism comparisons, and HKR-R from live RL-recipe choices; the higher technical bar keeps it in low featured.

Apr 30Thursday

r/LocalLLaMA

DeepSeek released Thinking with Visual Primitives framework

DeepSeek, Peking University, and Tsinghua released the Thinking with Visual Primitives paper and repository. The framework inserts coordinate points and bounding boxes into chain-of-thought; the post does not disclose benchmark scores.

Why it matters: HKR-H/K/R all pass: the hook is visual primitives inside reasoning, the new fact is point/box CoT plus an open repo, and the audience cares about grounded VLMs. No benchmark scores are disclosed, so it stays at 80, not P1.

Apr 29Wednesday

X · @op7418

Deepseek’s multimodal model is fully rolled out

Deepseek fully rolled out a multimodal model, available via the web image-recognition mode. The post says it looks like a separate model; it does not disclose name, size, pricing, or API timing.

Why it matters: HKR-H/K/R all pass, but the X post only confirms web image-recognition access; model name, params, price, and API timing are missing. DeepSeek’s multimodal rollout is strong, but the thin sourcing keeps it in 78–84.

QbitAI · WeChat

DeepSeek’s multimodal AI has entered testing

DeepSeek researchers confirmed V4 vision mode is in gray testing, with an image-recognition mode on the homepage. A screenshot shows it identified drinks and cup types in a non-text-heavy image after 4 seconds. The post does not disclose rollout scope, API access, or pricing.

Why it matters: HKR-H/K/R all pass: DeepSeek’s V4 vision gray test is a real domestic flagship update with a concrete 4s sample. Score stays at 80 because access scope, API form, pricing, and benchmarks are not disclosed.

r/LocalLLaMA

DeepSeek V4 pricing is genuinely silly; the math made me question my stack

A Reddit user calculates DeepSeek V4-Pro input at $0.145 per million tokens, about 34x cheaper than Claude Opus 4.7. A May promo cuts it to $0.036, while cache hits are $0.0036, about 173x below Opus cached pricing. The key issue is agent-loop cost; the post does not verify the 1M context under production loads.

Why it matters: HKR-H/K/R all pass on the pricing hook, concrete token prices, and agent-cost pressure. Capped below 78 because this is a Reddit calculation, not an official release or production benchmark.

Computing Life · Share · Yage

DeepSeek V4 Explained: Engineering Decisions Around Agentic Workloads

DeepSeek V4 targets long-horizon agent tasks with a 1M context. The snippet cites hybrid attention, OPD, Muon, and mHC; the post does not disclose size, data, pricing, or release timing.

Why it matters: HKR-H/K/R all pass: DeepSeek V4, 1M context, and agentic workload engineering create a strong hook with concrete mechanisms. Missing params, data, price, and launch timing keep it at 78, not P1.

Apr 27Monday

QbitAI · WeChat

DeepSeek V4 Cuts Prices Permanently; Cached Inputs Get 90% Off, Coding Test Costs Drop 83%

DeepSeek V4 cut prices twice in two days: input/output pricing is 75% lower, with cached inputs getting another 90% off. QbitAI’s coding test fell from 31.73 yuan for 35M tokens to 5.34 yuan under new pricing, an 83% drop. The key case is high cache-hit workloads, with V4-Pro at about 95–96% cache hits.

Why it matters: HKR-H/K/R all pass: DeepSeek V4 pricing has a sharp cost hook, concrete test numbers, and strong cost resonance. It is still a pricing update, not a new model release, so it stays below the 85 P1 band.

Apr 26Sunday

Hacker News front page

DeepSeek-V4 on Day 0: From Fast Inference to Verified RL with SGLang and Miles

SGLang and Miles added day-0 inference and RL support for DeepSeek-V4, covering 1.6T Pro and 284B Flash. The post cites a 1M-token context, FP4 MoE expert weights, 128-token SWA, and 4:1 or 128:1 KV compression. The key systems detail is ShadowRadix coherence across three KV pools and two compression-state pools.

Why it matters: HKR-H/K/R all pass: a DeepSeek-V4 day-0 systems stack, concrete context/compression mechanisms, and clear deployment-cost stakes. The systems depth narrows reach, but no hard-exclusion rule is triggered.

Apr 25Saturday

Latent Space

DeepSeek V4 Pro and Flash released, runnable on Huawei Ascend chips

DeepSeek released V4 Pro and V4 Flash, with 1.6T/49B active and 284B/13B active parameters. Both support 1M-token context, Base/Instruct variants, and an MIT license; the report claims 27% FLOPs and 10% KV cache versus V3.2 at 1M tokens. The key point is Huawei CANN compatibility, not just benchmarks, because it reduces CUDA dependence.

Why it matters: HKR-H/K/R all pass: a major DeepSeek release adds concrete specs, 1M context, MIT licensing, and Huawei Ascend support. This sits in the 85–94 must-write band, with hardware independence pushing it upward.

MIT Technology Review · AI

Three reasons why DeepSeek’s new model matters

DeepSeek released a V4 preview with two versions: V4-Pro and V4-Flash. V4-Pro costs $1.74/M input tokens and $3.48/M output tokens; V4-Flash is about $0.14/$0.28, and both support 1M-token context. The key point is attention efficiency and open weights pressuring agentic coding costs.

Why it matters: HKR-H/K/R all pass: DeepSeek V4 is a domestic flagship release with 1M context, two price tiers, and open-weight cost pressure. The preview status keeps it below a full GPT/Claude major release, but it is same-day material.

Bloomberg Technology

China’s DeepSeek Unveils New Model a Year After Shock Launch

DeepSeek unveiled a new flagship AI model about one year after its open-source release jolted Silicon Valley. The title and RSS snippet confirm that timing; the post does not disclose the model name, size, pricing, benchmarks, or release terms. The key thing to watch is the missing launch detail, not the comeback framing.

Why it matters: A new DeepSeek flagship is newsworthy: HKR-H comes from the 'one year after the shock launch' hook, and HKR-R from the open-source and pricing rivalry it triggers. HKR-K fails because no model name, params, pricing, or benchmarks are disclosed, so this sits at the low end of the

Apr 24Friday

TechCrunch · AI

DeepSeek previews new AI model that ‘closes the gap’ with frontier models

DeepSeek previewed two new models and said architectural changes make them more efficient and higher-performing than DeepSeek V3.2, while nearly closing the gap with leading models on reasoning benchmarks. The RSS snippet discloses only that there are two models and that they outperform V3.2; model names, parameter counts, benchmark scores, test sets, and release timing are not disclosed. The key question is reproducible evals, because “closes the gap” comes without numbers.

Why it matters: A new-model preview from DeepSeek, a flagship Chinese lab, clears HKR-H and HKR-R on competitive relevance alone. HKR-K is weak because the story gives only 'two models' and 'better than V3.2' while model names, benchmark scores, test sets, and release timing are not disclosed,so

The Verge · AI

China’s DeepSeek previews new AI model a year after jolling US rivals

DeepSeek released a preview of its open-source V4 model on Friday and said it can compete with closed systems from Anthropic, Google, and OpenAI. The RSS snippet says V4 improves coding and highlights compatibility with Huawei tech; parameter count, benchmark scores, and rollout details are not disclosed. The part to watch is the pairing of agent-focused coding gains with tighter alignment to China’s domestic chip stack.

Why it matters: This is a flagship Chinese model update with HKR-H/K/R: a new open-source V4 preview, coding gains, and Huawei compatibility. It stays below the 85 band because the story withholds params, benchmark scores, and launch timing.

r/LocalLLaMA

DeepSeek releases V4: 1.6T Pro, 284B Flash, MIT license, 1M context

DeepSeek released two open-weight V4 models: Pro at 1.6T total with 49B active, and Flash at 284B total with 13B active; both use an MIT license and support 1M context. The RSS snippet points to a Hugging Face collection and a tech report, but the post does not disclose benchmark scores, pricing, training data size, or real inference throughput. The key thing to watch is the 1M context plus low active-parameter ratio; if evals hold, self-hosted long-context and routing economics change materially.

Why it matters: HKR-H/K/R all pass: this is a flagship DeepSeek open release with two huge MIT-licensed weights and 1M context, strong enough for same-day coverage. The score stops at 86 because the provided text does not disclose benchmarks, throughput, training data, or pricing.

Bloomberg Technology

DeepSeek unveils flagship AI model a year after breakthrough

DeepSeek released preview versions of a new flagship AI model one year after its breakout. The RSS snippet calls it its most powerful open-source platform and frames it against OpenAI and Anthropic; the post does not disclose parameters, context length, benchmarks, or rollout timing. The actionable facts so far are limited to its preview status and open-source positioning.

Why it matters: A new DeepSeek flagship preview deserves real weight under the domestic-flagship rule, and Bloomberg adds source authority. HKR-H and HKR-R pass, but HKR-K fails because the story discloses no specs, context window, benchmarks, or release schedule, so this stays at the low end of

X · @Yuchenj_UW

Finally, DeepSeek V4 is here!

DeepSeek announced DeepSeek V4 and says DeepSeek-V4-Pro uses an MIT license with 1.6T parameters and 49B active parameters. The snippet also claims DeepSeek-V4-Pro Max is close to Opus-4.6 Max and GPT-5.4 xHigh across benchmarks; the post does not disclose benchmark names, scores, release timing, or model weights. The key signal is the MIT license and 49B active scale, not the headline comparison.

Why it matters: This is a flagship DeepSeek model launch, and the MIT license plus 49B active scale make HKR-H/K/R pass. I keep it at 84, not p1, because the current source does not disclose benchmark names, exact scores, release timing, or a weights link.

X · @dotey

DeepSeek releases and open-sources V4 preview; 1M context is standard across all services

DeepSeek released and open-sourced the V4 preview, making 1M context standard across all official services with no tier or price split. The post says V4-Pro and V4-Flash use token compression plus DSA sparse attention to cut compute and memory costs for 1M context; legacy APIs remain for 3 months and stop after July 24.

Why it matters: DeepSeek is a flagship Chinese model vendor, and this V4 preview is a substantive release with open source and 1M context made standard across official services. HKR-H/K/R all pass: the post includes mechanisms and a migration deadline, and the tier reset makes it a same-day P1.

X · @op7418

DeepSeek V4 detailed official announcement is out

DeepSeek says V4 Pro has 1.6T total parameters with 49B active, while Flash has 284B total and 13B active; both were pretrained on 32T tokens. Web and app Expert mode map to Pro, and Fast mode maps to Flash. The post also says several benchmarks are on par with Opus 4.6, with stronger agent ability and world knowledge, plus a new attention mechanism that reduces compute and memory demand.

Why it matters: This is a flagship DeepSeek release, scored on par with peer US lab model launches. HKR-H/K/R all pass on concrete scale numbers, 32T data, and an inference-efficiency mechanism; benchmark setup, pricing, and API availability are not disclosed in the summary.

X · @op7418

DeepSeek V4 arrives with Flash and Pro variants

DeepSeek released V4 with two variants, Flash and Pro. The RSS snippet says it supports JSON output, tool calling, dialogue prefix continuation, and FIM completion; Flash costs ¥0.2/¥1 per million input/output tokens, while Pro costs ¥1/¥12. At 1M context, output pricing doubles.