Skip to content

#DeepSeek

0 today

Jun 6Saturday

Synced · WeChat

DeepSeek V4 Proves Math with 500x Cost Advantage as Agent System Sets Records

Princeton researchers released Goedel-Architect, an agent framework for Lean formal theorem proving. Using DeepSeek-V4-Flash, it reached 75.6% pass@1 on PutnamBench, with $294 in API cost for 672 problems, compared with Hilbert’s 70.0% and about $170,000 cost.

Why it matters: HKR-H/K/R all pass: Goedel-Architect pairs a 75.6% PutnamBench score with $294 for 672 problems, versus Hilbert at about $170k. It is still research-heavy, so it stays in the 78–84 band rather than P1.

Jun 3Wednesday

Synced · WeChat

Understanding SFT Mechanisms in LLMs: Resolving Practice Disputes and Avoiding Wasted Compute

Junpeng Zhang and coauthors argue that SFT on highly homogeneous data has an effective window of only hundreds to about 1,000 training steps, and their interaction-based warning signal detects overfitting before loss gaps appear, saving roughly 30%–50% of training compute.

Why it matters: HKR-H/K/R all pass: the paper gives testable SFT windows, earlier overfitting warnings, and 30%-50% compute savings. It is strong research, not a major model or product release, so it stays below 85.

Computing Life · Share · Yage

Microsoft AI's MAI-Thinking-1: Getting Models to Think Is Easy, Sustained Thinking Is Hard

Microsoft AI says MAI-Thinking-1 uses three mechanisms—thermostat, circuit breaker, and self-distillation—to keep RL training stable for several thousand steps; the RSS snippet contrasts MAI’s discipline with DeepSeek’s efficiency and GLM’s endurance.

Why it matters: HKR-H/K/R all pass: the hook is training persistence, the new facts are three stability mechanisms and thousand-step RL runs, and the audience cares about reasoning-model stability. Not a major model launch, so it stays below 85.

Jun 1Monday

r/LocalLLaMA

I bolted an 8-arm reasoning MoE onto a frozen 1.4B Mamba backbone on a single RTX 3060

The author trained Mamba-Titan-1.4B-Reasoning on a 12GB RTX 3060: a frozen 1.4B Mamba-1 backbone with 8 trainable MoE arms, 2.54B total parameters, Top-2 routing at layers 24/25, and about 50% math accuracy.

Why it matters: HKR-H/K/R all pass via a numbered first-person experiment, but it is a single Reddit post with no independent replication and a fairly technical setup, so it stays in the low featured band.

May 29Friday

Synced · WeChat

The Ma Jiaqi Failure Exposed an LLM Issue He Spotted in the Shower a Year Earlier

FaceMind links low-frequency token degradation to two papers: SLoW appeared at EMNLP 2025, Adam's Law was accepted as an ACL 2026 Oral, and high-frequency rewriting raised DeepSeek-V3 math accuracy from 63.55% to 71.54%.

Why it matters: HKR-H/K/R all pass: the odd celebrity-token hook is clickable, and the post gives a mechanism plus a 63.55%→71.54% DeepSeek-V3 result. Practical research signal, but not a major model launch.

May 15Friday

r/LocalLLaMA

MOOSE-Star (ICML 2026): 7B Model and 108K-Paper Dataset for Scientific Hypothesis Discovery

MiroMind researchers released the MOOSE-Star collection with three 7B models and TOMATO-Star, a dataset of 108,717 NCBI papers. MS-IR-7B reaches 54.37% inspiration-retrieval accuracy, uses DeepSeek-R1-Distill-Qwen-7B as its base, runs at about 14GB fp16, and supports llama.cpp, vLLM, and SGLang.

Why it matters: HKR-H/K/R all pass via the local 7B research-agent hook and concrete dataset metrics. Single Reddit source and limited lab gravity keep it below the must-write band.

May 14Thursday

Synced · WeChat

ACL 2026: Alibaba DAMO I²B-LPO Improves RLVR Exploration

Alibaba DAMO Academy introduced I²B-LPO, an RLVR post-training framework that branches rollouts at high-entropy nodes and filters them with an information-bottleneck self-reward, reporting up to 5.3% accuracy gains and 7.4% semantic-diversity gains on math benchmarks using Qwen2.5-7B and Qwen3-14B.

Why it matters: HKR-H/K/R all pass: the ACL 2026 DAMO paper has a clear RLVR exploration hook, concrete I²B-LPO mechanics, and benchmark gains. It is still a training-method paper, not a major model or product release, so 78 fits the lower good-quality band.

May 3Sunday

QbitAI · WeChat

DeepSeek V4’s biggest omission

DeepSeek V4’s technical report omits Engram while listing mHC, CSA, HCA, Muon, and FP4. Engram was open-sourced by DeepSeek and Peking University in January, inserting lookup modules between Transformer layers 2 and 15; its 27B test raised MMLU by 3.4 and Multi-Query NIAH to 97.0%. The engineering signal is CXL pooling: 8 servers shared a 4TB memory pool with under 5% throughput loss.

Why it matters: HKR-H/K/R all pass: the omitted-Engram angle is clickable, with layer ranges, benchmark deltas, and CXL memory-pool numbers. It is analysis, not the V4 launch itself, so 78–84 fits.

May 1Friday

Synced · WeChat

The Evolution of RL: From PPO to MaxRL in LLM Reasoning Training

Jiqizhixin translated Alexander Weers' article on RL algorithms for LLM reasoning from 2024 to 2026. It covers REINFORCE, PPO, GRPO, RLOO, Dr. GRPO, DAPO, CISPO, MaxRL, DPPO, and ScaleRL, comparing critic removal, clipping, normalization, and pass@k goals. The key signal is mechanism choice, not algorithm names.

Why it matters: A strong technical explainer, not a model or paper release. HKR-H comes from the PPO→MaxRL arc, HKR-K from concrete mechanism comparisons, and HKR-R from live RL-recipe choices; the higher technical bar keeps it in low featured.

Apr 30Thursday

r/LocalLLaMA

DeepSeek released Thinking with Visual Primitives framework

DeepSeek, Peking University, and Tsinghua released the Thinking with Visual Primitives paper and repository. The framework inserts coordinate points and bounding boxes into chain-of-thought; the post does not disclose benchmark scores.

Why it matters: HKR-H/K/R all pass: the hook is visual primitives inside reasoning, the new fact is point/box CoT plus an open repo, and the audience cares about grounded VLMs. No benchmark scores are disclosed, so it stays at 80, not P1.

Apr 29Wednesday

Computing Life · Share · Yage

DeepSeek V4 Explained: Engineering Decisions Around Agentic Workloads

DeepSeek V4 targets long-horizon agent tasks with a 1M context. The snippet cites hybrid attention, OPD, Muon, and mHC; the post does not disclose size, data, pricing, or release timing.

Why it matters: HKR-H/K/R all pass: DeepSeek V4, 1M context, and agentic workload engineering create a strong hook with concrete mechanisms. Missing params, data, price, and launch timing keep it at 78, not P1.

Apr 25Saturday

Latent Space

DeepSeek V4 Pro and Flash released, runnable on Huawei Ascend chips

DeepSeek released V4 Pro and V4 Flash, with 1.6T/49B active and 284B/13B active parameters. Both support 1M-token context, Base/Instruct variants, and an MIT license; the report claims 27% FLOPs and 10% KV cache versus V3.2 at 1M tokens. The key point is Huawei CANN compatibility, not just benchmarks, because it reduces CUDA dependence.

Why it matters: HKR-H/K/R all pass: a major DeepSeek release adds concrete specs, 1M context, MIT licensing, and Huawei Ascend support. This sits in the 85–94 must-write band, with hardware independence pushing it upward.