Skip to content

#DeepSeek

0 today

Jun 12Friday

AI HOT (Curated Pool)

Hugging Face open-sourced Open-R1, a full reproduction of DeepSeek-R1

Hugging Face published Open-R1 on GitHub, aiming to fully reproduce the DeepSeek-R1 reasoning model. The repo has 26.1k stars and 2.4k forks so far. The body only contains the repo's landing page navigation and metadata; it does not disclose the implementation plan, training data, reproduction progress, or benchmark results. I'd treat this as a public reproduction scaffold and collaboration hub for now, and wait for a technical report before judging fidelity.

Why it matters: Hugging Face launched a full open-source reproduction of DeepSeek-R1, with the repo already at 26.1k stars — strong community interest. But the body only contains project scaffolding and navigation; no implementation plan, training data, or reproduction progress is disclosed y...

Jun 7Sunday

AI HOT (Curated Pool)

AI Substitution Wave: Three Forces Reshape Cost Structures

Coinbase, Lindy, Harvey, and Cursor shifted workloads to cheaper models; Harvey reported Kimi 2.6 reached a 15% all-pass rate on Legal Agent Benchmark, versus Opus at 14%, with 100 tasks costing $84 versus $954.

Why it matters: HKR-H/K/R all pass: the $84 vs $954 cost delta and named cases from Coinbase, Lindy, Harvey, and Cursor give it concrete signal. It is a strong cost-structure commentary, not a major model or product release, so it fits the 72-77 band.

Jun 6Saturday

AI Chat-Group Daily (群聊日报)

Chat Group Weekly Vol. 2: The AI Tricks You Learned This Year May Be Wasted

The author retired an OpenClaw AI assistant after more than one month of use; the post says it required self-hosting, API setup, and keeping one home computer running 24 hours a day.

Why it matters: HKR-H/K/R all pass, but this is a personal weekly write-up, not a model or platform release. The month-long OpenClaw use and 24/7 PC requirement make it just clear the featured threshold.

Synced · WeChat

DeepSeek V4 Proves Math with 500x Cost Advantage as Agent System Sets Records

Princeton researchers released Goedel-Architect, an agent framework for Lean formal theorem proving. Using DeepSeek-V4-Flash, it reached 75.6% pass@1 on PutnamBench, with $294 in API cost for 672 problems, compared with Hilbert’s 70.0% and about $170,000 cost.

Why it matters: HKR-H/K/R all pass: Goedel-Architect pairs a 75.6% PutnamBench score with $294 for 672 problems, versus Hilbert at about $170k. It is still research-heavy, so it stays in the 78–84 band rather than P1.

Jun 5Friday

Ruan YiFeng's Weblog

Tech Enthusiasts Weekly Issue 399: Visits to China’s AI Majors

Ruan Yifeng excerpts observations from U.S. analysts who visited 14 Chinese AI and robotics companies in early May: the article estimates U.S. AI compute at about 8 times China’s by the end of 2025, while Chinese firms’ intelligence output per unit of compute is estimated at 4-7 times naive scaling.

Why it matters: All three HKR axes pass: many named visit targets, concrete compute ratios, and a China-US AI competition nerve. It is still a secondary commentary post, not a primary release or major product event, so it sits just above the featured threshold.

Jun 3Wednesday

Synced · WeChat

Understanding SFT Mechanisms in LLMs: Resolving Practice Disputes and Avoiding Wasted Compute

Junpeng Zhang and coauthors argue that SFT on highly homogeneous data has an effective window of only hundreds to about 1,000 training steps, and their interaction-based warning signal detects overfitting before loss gaps appear, saving roughly 30%–50% of training compute.

Why it matters: HKR-H/K/R all pass: the paper gives testable SFT windows, earlier overfitting warnings, and 30%-50% compute savings. It is strong research, not a major model or product release, so it stays below 85.

AI HOT (Curated Pool)

DeepSeek Reportedly Seeks RMB 50 Billion in First Funding Round with Tencent and CATL

DeepSeek plans to raise about RMB 50 billion in its first funding round, with post-money valuation expected at RMB 350 billion to RMB 400 billion; Liang Wenfeng, Tencent, and CATL plan to invest RMB 20 billion, RMB 10 billion, and RMB 5 billion respectively.

Why it matters: HKR-H/K/R all pass: DeepSeek's rumored RMB 50B first round includes a RMB 350B-400B valuation and named checks from Tencent and CATL. The rumor status keeps it at 88, below confirmed industry-shaking funding news.

Computing Life · Share · Yage

Microsoft AI's MAI-Thinking-1: Getting Models to Think Is Easy, Sustained Thinking Is Hard

Microsoft AI says MAI-Thinking-1 uses three mechanisms—thermostat, circuit breaker, and self-distillation—to keep RL training stable for several thousand steps; the RSS snippet contrasts MAI’s discipline with DeepSeek’s efficiency and GLM’s endurance.

Why it matters: HKR-H/K/R all pass: the hook is training persistence, the new facts are three stability mechanisms and thousand-step RL runs, and the audience cares about reasoning-model stability. Not a major model launch, so it stays below 85.

Jun 2Tuesday

AI HOT (Curated Pool)

StepFun releases Step 3.7 Flash for efficient inference

StepFun released Step 3.7 Flash with a 196B MoE architecture, using multi-matrix factorized attention to cut KV-cache cost to about 22% of DeepSeek models.

Why it matters: HKR-H/K/R all pass: Step 3.7 Flash has concrete specs, not just launch copy, with 196B MoE and ~22% KV-cache cost versus DeepSeek. It is below top-lab flagship weight, so 78 featured.

AI HOT (Curated Pool)

The Thriving Ecosystem of Open Models

OpenRouter data shows open-weight models generated 69.1% of token usage since 2025, versus 30.9% for closed models, while share leadership shifted across DeepSeek, MiniMax, Kimi, MiMo, Qwen, Tencent Hy3, Alibaba, and Arcee releases.

Why it matters: HKR-H comes from the 69.1% vs 30.9% contrast, HKR-K has OpenRouter token-share data, and HKR-R hits open-vs-closed competition. It is a data-backed commentary, so featured low band.

Jun 1Monday

Xinzhiyuan · WeChat

400 tokens/s: StepFun Step 3.7 Flash cuts Agent task costs

StepFun released Step 3.7 Flash, a sparse MoE model with 196B parameters plus a 1.8B ViT, activating 11B parameters per inference and reaching up to 400 tokens per second.

Why it matters: HKR-H/K/R all pass with concrete speed and parameter numbers. The feed does not disclose pricing, benchmark setup, or open-source terms, so this stays in the 78–84 quality update band.

r/LocalLLaMA

Deepseek V4 Flash performance on DGX Spark

A Reddit user ran DeepSeek-V4-Flash with vLLM on two ASUS GX10 DGX Spark nodes and reported 1,680 prefill tokens/s plus 39.8 decode tokens/s at a 256K context with MTP=2; the setup uses TP=2 over RoCE, fp8 KV cache, and fits about 1M tokens safely in KV cache.

Why it matters: This is not broad industry news, but it is a first-person benchmark with reproducible details: TP=2, RoCE, fp8 KV cache, 256K context, and ~1M KV. HKR-H/K/R all pass, so it lands at low featured.

r/LocalLLaMA

I bolted an 8-arm reasoning MoE onto a frozen 1.4B Mamba backbone on a single RTX 3060

The author trained Mamba-Titan-1.4B-Reasoning on a 12GB RTX 3060: a frozen 1.4B Mamba-1 backbone with 8 trainable MoE arms, 2.54B total parameters, Top-2 routing at layers 24/25, and about 50% math accuracy.

Why it matters: HKR-H/K/R all pass via a numbered first-person experiment, but it is a single Reddit post with no independent replication and a fairly technical setup, so it stays in the low featured band.

May 30Saturday

Bloomberg Technology

MiniMax Eyes China Listing, Takes on AI Rivals Like DeepSeek

MiniMax Group has begun preparations for a domestic China listing, according to a regulatory filing, and the post identifies DeepSeek as a local AI rival; the RSS snippet does not disclose valuation, listing timeline, exchange venue, or fundraising size.

Why it matters: Bloomberg reports MiniMax has begun domestic listing prep via regulatory filings, clearing HKR-H/K/R. Missing valuation, timing, and raise size keep it in the 78–84 band, not P1.

May 29Friday

Xinzhiyuan · WeChat

Three DeepSeek Models Enter OpenRouter Monthly Top 10 With Over 17 Trillion Tokens

DeepSeek placed three models in OpenRouter’s monthly top 10 with more than 17 trillion tokens combined, including V4 Flash at 9.13T tokens; the article says Ascend’s MegaMoE operator raised Prefill throughput by 20% to 30% on DeepSeek V3.1 and Qwen3-235B tests.

Why it matters: HKR-H/K/R all pass: the story has a 17T-token hook plus concrete OpenRouter and MegaMoE Prefill numbers. It stays at 82 because the compute-sovereignty framing is strong, while reproducible test conditions are not disclosed.

Synced · WeChat

The Ma Jiaqi Failure Exposed an LLM Issue He Spotted in the Shower a Year Earlier

FaceMind links low-frequency token degradation to two papers: SLoW appeared at EMNLP 2025, Adam's Law was accepted as an ACL 2026 Oral, and high-frequency rewriting raised DeepSeek-V3 math accuracy from 63.55% to 71.54%.

Why it matters: HKR-H/K/R all pass: the odd celebrity-token hook is clickable, and the post gives a mechanism plus a 63.55%→71.54% DeepSeek-V3 result. Practical research signal, but not a major model launch.

r/LocalLLaMA

StepFun 3.7 Flash

StepFun released Step 3.7 Flash with 196B total parameters, 11B active MoE, a built-in 1.8B ViT, and local execution on 128GB RAM.

Why it matters: HKR-H/K/R pass via the 196B/11B MoE specs and 128GB local-run claim. Sparse Reddit sourcing leaves license, eval method, and access conditions undisclosed, so it stays in the lower featured band.

May 28Thursday

QbitAI · WeChat

Behind DeepSeek V4's Chip-Model Co-Design, China's Compute Ecosystem Gains Speed

QbitAI says DeepSeek V4 validated Ascend chip-model co-design, with CANN open-sourcing 65 repositories and supporting day-zero adaptation for more than 70 mainstream models, while AIGCode reported 65% MFU in MoE pretraining on Ascend.

Why it matters: HKR-H/K/R all pass, but this is mainly a compute-ecosystem progress story, not a DeepSeek V4 capability release. Concrete repo, adaptation, and MFU numbers lift it into featured, below must-write.

AI HOT (Curated Pool)

DeepSeek plans STAR Market IPO after completing roughly $50B funding round

DeepSeek plans to apply for a STAR Market IPO after completing a roughly $50 billion funding round, according to a large fund manager participating in the round; the post does not disclose valuation, timetable, filing documents, or company confirmation.

Why it matters: HKR-H/K/R all pass: a DeepSeek STAR Market IPO after a $50B round would put a Chinese foundation-model lab into public-market pricing. Single X sourcing and no formal filing keep it at the low end of the must-write band.

May 27Wednesday

AI HOT (Curated Pool)

MiMo 2.5 Pro Gets Major Price Cut, Matching DeepSeek V4 Pro

Xiaomi permanently cut MiMo-V2.5 API prices by up to 99%, matched DeepSeek V4 Pro pricing, increased same-price token allowances by 5–8x, reset existing user quotas in full, and set the new pricing to take effect on May 26.

Why it matters: HKR-H/K/R all pass: the 99% cut creates a price-war hook, the post gives 5-8x token economics, and API cost pressure resonates. It remains a pricing update, not a model or capability release, so it stays below the 78+ band.

May 26Tuesday

New York Times Chinese

The Shared U.S.-China AI Anxiety: Being Harvested by the Future

Yi-Ling Liu compares U.S. and Chinese AI anxiety through labor, companionship, and agency: over 70% of U.S. teenagers report using chatbots as companions, while China is projected to reach 200 million single-person households by 2030.

Why it matters: HKR-H/K/R all pass, but this is commentary rather than a model, product, or policy release. Its signal comes from two social data points and a US-China framing, so it fits the featured threshold for an insightful opinion piece.

May 25Monday

r/LocalLLaMA

The reason small-model agent stacks aren't the default is not whether they work

A Reddit post argues small-model agent stacks are not default for business reasons, not capability limits: Gemma 4 31B reaches 86.4% on tau2-bench, and DeepSeek V4-Flash output tokens are priced about 89x below Claude Opus 4.6. The operational risk is verification, because 7–9B models produced broken reasoning for roughly half to two-thirds of correct answers in a cited audit.

Why it matters: HKR-H/K/R all pass: the angle is contrarian, with benchmark, cost, and verifier-failure numbers. Reddit-source uncertainty keeps it in the 78–84 recommendation band, not P1.

QbitAI · WeChat

Reasonix for DeepSeek V4 reaches 99.82% cache hit rate and cuts costs to 20%

Reasonix uses an append-only loop for DeepSeek V4 and reports a 99.82% cache hit rate in long coding sessions, cutting an example 400M-token bill from $61 to $12.

Why it matters: HKR-H/K/R all pass, but this is a third-party cost tool around DeepSeek V4, not a model launch or platform update. Concrete mechanism and billing numbers put it in the 72–77 featured band.

May 23Saturday

Bloomberg Technology

DeepSeek To Make Permanent 75% Discount on Flagship AI Model

DeepSeek will make a 75% discount on its flagship AI model permanent, but the post does not disclose the specific model name, original price, discounted price, or effective date.

Why it matters: HKR-H/K/R pass on a concrete 75% permanent discount from DeepSeek, a cost and price-war story. Sparse extracted body lacks model name, list price, discounted price, and timing, so it stays in low featured.

QbitAI · WeChat

DeepSeek V4 cuts prices as CATL, JD.com and NetEase discuss investment; Liang Wenfeng targets AGI

DeepSeek-V4-Pro API will keep its promotional pricing from June 1, with cached input at RMB 0.025 per million tokens, while Bloomberg says DeepSeek is pursuing a RMB 70 billion round at a USD 45 billion pre-money valuation.

Why it matters: HKR-H/K/R all pass: DeepSeek V4 API price cuts and Bloomberg’s RMB 70B raise at a $45B pre-money valuation are same-day material. The cost and capital angles directly affect China model competition.

May 22Friday

Hacker News front page

DeepSeek Makes the V4 Pro Price Discount Permanent

DeepSeek will set deepseek-v4-pro API pricing to one quarter of the original price after the 75% promotion ends on 2026-05-31 at 15:59 UTC; the post does not disclose the exact per-token price.

Why it matters: HKR-H/K/R all pass: the hook is a permanent DeepSeek price cut, the new fact is 1/4 pricing after a stated UTC time, and the nerve is API cost. Missing unit pricing keeps it at the featured floor, not a must-write release.

r/LocalLLaMA

DeepSeek Advances $10.29B Financing as Liang Wenfeng Commits to Open-Source AI Models

The title says DeepSeek is advancing a $10.29 billion financing round and Liang Wenfeng commits to continued open-source AI model development; the body only links to Bloomberg and does not disclose round terms, investors, or a commercialization timeline.

Why it matters: HKR-H/K/R all pass: the $10.29B DeepSeek financing and open-source pledge are high-signal. The thin Reddit body only links Bloomberg and omits investors, valuation, and terms, so it stays in the low 85–94 band.

AI HOT (Curated Pool)

DeepSeek Advances RMB 70B Funding as Liang Wenfeng Commits to Open-Source AI Models

DeepSeek is pursuing RMB 70 billion in funding at an estimated valuation of about $45 billion, with Tencent and IDG Capital close to participating and founder Liang Wenfeng potentially investing RMB 20 billion personally.

Why it matters: HKR-H/K/R all pass: a DeepSeek RMB 70B financing at a $45B valuation is a major China-model capital story with open-source stakes. It stays below 95 because the deal is still in progress and final terms are not disclosed.

Bloomberg Technology

DeepSeek Founder Declares AGI Goal as $10 Billion Round Advances

The title says DeepSeek’s founder declared an AGI goal and that a $10 billion funding round is advancing; the post does not disclose the founder’s statement, financing terms, investors, or timeline.

Why it matters: HKR-H/K/R all pass: DeepSeek plus a $10B round and AGI goal is same-day AI-business news. The scrape provides title-level facts only, with no investors, terms, or timeline, so the score stays at the low end of the 85+ band.

Computing Life · Share · Yage

How to Run DeepSeek V4 Flash Locally on Mac: DS4 Engine Explained

DS4 provides a macOS local runtime path for DeepSeek V4 Flash; the post only discloses three mechanisms—multi-agent integration, KV cache disk persistence, and activation steering—and does not disclose performance numbers, hardware requirements, or pricing.

Why it matters: HKR-H/K/R all pass, but the body only names DS4 mechanisms and omits performance, model size, Mac support, and reproducible tests; this fits the featured threshold for a local-inference tutorial.

May 20Wednesday

r/LocalLLaMA

Running DeepSeek-V4 locally on 4 legacy RTX 2080 Ti GPUs with W8A8 at 255 prefill tok/s

A Reddit user ran DeepSeek-V4-Flash locally on 4 RTX 2080 Ti GPUs, reporting 284B total parameters, 13B active parameters, a sub-$2,500 build, custom Turing CUDA kernels, W8A8 quantization, 1TB DDR4 ECC RAM, and about 255 prefill tokens/s.

Why it matters: HKR-H/K/R all pass: this is a numeric first-person local-inference experiment. Single-source Reddit provenance and custom Turing kernels keep it in the lower featured band.

May 17Sunday

r/LocalLLaMA

DeepSeek V4's 1M Context Window: The Breaking Point

A Reddit user tested DeepSeek V4 on 45k, 180k, and 520k-token codebases and found 150k-250k tokens best for coding work. Past 300k tokens, line-number precision degraded; at 520k, outputs shifted toward architecture summaries and skipped implementation details.

Why it matters: A single Reddit post limits authority, but HKR-H/K/R all pass: it is a numbered first-person test with a concrete long-context failure pattern. The right band is featured, not 78+, because replication and model details are thin.

AI HOT (Curated Pool)

Latest Open Artifacts #21: Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1, and More

Open AI model teams released Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1, and other versions this month, and the post says they were tested under CAISI’s V4 evaluation framework, but the RSS snippet does not disclose scores.

Why it matters: HKR-H/K/R all pass: a dense open-model roster, a named CAISI V4 evaluation frame, and clear practitioner relevance for model choice. Missing scores and reproducible detail keep it in the 78–84 band.

May 15Friday

r/LocalLLaMA

MOOSE-Star (ICML 2026): 7B Model and 108K-Paper Dataset for Scientific Hypothesis Discovery

MiroMind researchers released the MOOSE-Star collection with three 7B models and TOMATO-Star, a dataset of 108,717 NCBI papers. MS-IR-7B reaches 54.37% inspiration-retrieval accuracy, uses DeepSeek-R1-Distill-Qwen-7B as its base, runs at about 14GB fp16, and supports llama.cpp, vLLM, and SGLang.

Why it matters: HKR-H/K/R all pass via the local 7B research-agent hook and concrete dataset metrics. Single Reddit source and limited lab gravity keep it below the must-write band.

May 14Thursday

Synced · WeChat

ACL 2026: Alibaba DAMO I²B-LPO Improves RLVR Exploration

Alibaba DAMO Academy introduced I²B-LPO, an RLVR post-training framework that branches rollouts at high-entropy nodes and filters them with an information-bottleneck self-reward, reporting up to 5.3% accuracy gains and 7.4% semantic-diversity gains on math benchmarks using Qwen2.5-7B and Qwen3-14B.

Why it matters: HKR-H/K/R all pass: the ACL 2026 DAMO paper has a clear RLVR exploration hook, concrete I²B-LPO mechanics, and benchmark gains. It is still a training-method paper, not a major model or product release, so 78 fits the lower good-quality band.

May 13Wednesday

New York Times Chinese

China Seeks AI Technology Self-Reliance, Weakening Washington’s Leverage Over Beijing

DeepSeek optimized its latest model for inference on Huawei chips for the first time, while two semiconductor sources said training still relies on Nvidia chips; Huawei says it plans to release a training chip this year, but matching current Nvidia performance will take another year.

Why it matters: HKR-H/K/R all pass: NYT ties DeepSeek-Huawei chip optimization and Huawei's training-chip timeline to US export-control leverage. It is not a model launch and lacks benchmark results, so it stays in the 78–84 band.

May 12Tuesday

r/LocalLLaMA

I catalogued every way local models break JSON output and built a repair library across 288 model calls

Reddit user kexxty ran 288 structured-output calls through OpenRouter models, including Llama 3, Mistral, Command R, DeepSeek, and Qwen, and found similar JSON failure categories across local and API-only models. The MIT-licensed Python library outputguard validates against JSON Schema, applies 15 ordered repair strategies, includes 2,001 tests, and has no LLM provider dependency.

Why it matters: HKR-H/K/R all pass: 288 tests, the outputguard library, and a 15-step repair chain give practitioners reusable detail. Source is a single Reddit post, so it stays in the 72–77 featured band, not 78+.

May 11Monday

AI HOT (Curated Pool)

Pareto Code Reorders Model Selection Using Market Demand

OpenRouter says Pareto Code observes the Pareto frontier using real market demand; DeepSeek V4 Pro ranks first, followed by GPT 5.4 Mini and Gemini 3.1 Pro, while the post does not disclose the scoring formula or evaluation sample size.

Why it matters: HKR-H/K/R all pass, but the source is a single OpenRouter post with no sample size, time window, or pricing basis disclosed. It clears featured as a model-selection benchmark, not the 78+ band.

May 10Sunday

r/LocalLLaMA

I have DeepSeek V4 Pro at home

Reddit user fairydreaming ran DeepSeek V4 Pro Q4_K_M with a modified llama.cpp CUDA repo on one RTX PRO 6000 Blackwell Max-Q workstation GPU, using an 859GB model file; the shared log reports a 1M context window and 8.6 tokens per second generation speed.

Why it matters: HKR-H/K/R all pass: the hook is single-GPU local inference, with concrete file size, context, speed, and runtime path. Reddit single-source sourcing keeps it below must-write model-release territory.

May 9Saturday

AI HOT (Curated Pool)

Redis founder uses a C inference engine to run a large model on a personal computer

Antirez open-sourced ds4, a native inference engine for DeepSeek V4 Flash that uses a few thousand lines of C to run a 1M-context model on a 128GB MacBook Pro at a reported 27 tok/s.

Why it matters: HKR-H/K/R all pass: Antirez open-sourced a native C inference engine with hardware, model, context, and speed numbers. Single-source X provenance keeps it below P1, but it is strong open-source inference signal.