Skip to content

DeepSeek

DeepSeek's model releases, open weights and technical reports — the bellwether for open-model price and performance.

194 picksRelated topicsQwenOpen sourceModel releases

Latest picks

181–194 of 194

Apr 24Friday

X · @Yuchenj_UW

Finally, DeepSeek V4 is here!

DeepSeek announced DeepSeek V4 and says DeepSeek-V4-Pro uses an MIT license with 1.6T parameters and 49B active parameters. The snippet also claims DeepSeek-V4-Pro Max is close to Opus-4.6 Max and GPT-5.4 xHigh across benchmarks; the post does not disclose benchmark names, scores, release timing, or model weights. The key signal is the MIT license and 49B active scale, not the headline comparison.

Why it matters: This is a flagship DeepSeek model launch, and the MIT license plus 49B active scale make HKR-H/K/R pass. I keep it at 84, not p1, because the current source does not disclose benchmark names, exact scores, release timing, or a weights link.

X · @dotey

DeepSeek releases and open-sources V4 preview; 1M context is standard across all services

DeepSeek released and open-sourced the V4 preview, making 1M context standard across all official services with no tier or price split. The post says V4-Pro and V4-Flash use token compression plus DSA sparse attention to cut compute and memory costs for 1M context; legacy APIs remain for 3 months and stop after July 24.

Why it matters: DeepSeek is a flagship Chinese model vendor, and this V4 preview is a substantive release with open source and 1M context made standard across official services. HKR-H/K/R all pass: the post includes mechanisms and a migration deadline, and the tier reset makes it a same-day P1.

X · @op7418

DeepSeek V4 detailed official announcement is out

DeepSeek says V4 Pro has 1.6T total parameters with 49B active, while Flash has 284B total and 13B active; both were pretrained on 32T tokens. Web and app Expert mode map to Pro, and Fast mode maps to Flash. The post also says several benchmarks are on par with Opus 4.6, with stronger agent ability and world knowledge, plus a new attention mechanism that reduces compute and memory demand.

Why it matters: This is a flagship DeepSeek release, scored on par with peer US lab model launches. HKR-H/K/R all pass on concrete scale numbers, 32T data, and an inference-efficiency mechanism; benchmark setup, pricing, and API availability are not disclosed in the summary.

X · @op7418

DeepSeek V4 arrives with Flash and Pro variants

DeepSeek released V4 with two variants, Flash and Pro. The RSS snippet says it supports JSON output, tool calling, dialogue prefix continuation, and FIM completion; Flash costs ¥0.2/¥1 per million input/output tokens, while Pro costs ¥1/¥12. At 1M context, output pricing doubles.

Hugging Face Blog

DeepSeek-V4: a million-token context that agents can actually use

DeepSeek released V4 with two MoE checkpoints, Pro and Flash, both supporting a 1M-token context. Pro has 1.6T total and 49B active parameters; Flash has 284B total and 13B active. The key detail is KV cost: Pro uses 27% of V3.2 single-token FLOPs and 10% of its KV cache; Flash uses 10% and 7%.

Why it matters: DeepSeek-V4 is a flagship Chinese model release with 1M-token context and KV cache at 7%–10% of V3.2. HKR-H/K/R all pass, placing it in the 85–94 same-day band.

Apr 23Thursday

Financial Times · Technology

DeepSeek targets a $20bn valuation to stop poaching of staff

DeepSeek is seeking its first funding round at a $20bn valuation to reduce rival poaching of researchers. The RSS snippet discloses prior defections and that this is its first raise, but the post does not disclose round size, investors, or headcount lost. The real signal is talent retention, not the headline valuation.

Why it matters: HKR-H lands because the title ties a $20bn valuation to stopping staff poaching. HKR-K and HKR-R also pass: FT adds first-fundraise and talent-war facts, but deal size, investors, and exit counts are undisclosed, so this is featured rather than p1.

Apr 22Wednesday

Bloomberg Technology

Tencent, Alibaba in Talks to Join DeepSeek’s First Funding Round

Tencent and Alibaba are in talks to join DeepSeek’s first funding round, and the snippet confirms this is DeepSeek’s maiden financing. The RSS text discloses only the talks and the first-round status; it does not disclose the round size, valuation, lead investor, or timing. What matters is whether strategic capital from two Chinese internet giants also brings compute or distribution terms, but the post does not disclose them.

Why it matters: Bloomberg adds one real datapoint: DeepSeek is pursuing its first funding round, with Tencent and Alibaba in talks. Amount, valuation, lead investor, and timing are still undisclosed, so it stays below P1; HKR-H/K/R all pass because the capital-and-cloud implications are strong.

Apr 8Wednesday

QbitAI · WeChat

After a late-night update, DeepSeek reportedly said: I am V4?

DeepSeek added Fast and Expert modes on its web app and started gray-testing a Vision model; the claim that Expert mode is V4 comes only from user probes and the model’s own replies. The post gives one concrete detail: Expert mode focuses on code, web, and harder generation tasks, is supply-limited, does not support multimodal or file upload, and one user reported a length cap at about 133K tokens. What matters is the official model ID and context spec; the post does not disclose them, pricing, or a release timeline.

Why it matters: HKR-H is strong on the 'I am V4' hook. HKR-K and HKR-R pass because the post gives testable mode behavior and a ~133K token limit, and DeepSeek silent swaps are highly discussable. The score stays in the mid-70s because the model name, price, and context window remain unconfirmed

Apr 4Saturday

X · @dotey

DeepSeek's next-generation V4 model will run on Huawei chips

DeepSeek delayed V4 for months and rewrote some low-level modules with Huawei and Cambricon so it runs on Huawei's Ascend 950PR, with launch expected in weeks, per The Information. The post cites 112GB memory, 1.4TB/s bandwidth, 600W power, and FP4 inference support; it does not disclose V4 size, pricing, or measured performance.

Why it matters: This clears HKR-H/K/R: Huawei-chip deployment is a strong hook, the report includes concrete module and chip details, and the China compute-stack angle will travel. It stays below 85 because this is pre-release reporting; model size, price, and real benchmarks are undisclosed.

Apr 3Friday

X · @dotey

LatePost on DeepSeek before V4: traits, organization, and Liang Wenfeng's goals

LatePost says DeepSeek has confirmed 4 core departures, and V4's large model slipped from around Lunar New Year to April; the report says it will likely remain open source. The snippet cites 2x-3x recruiting offers, some 8-digit packages, a 100-plus research team, and a shift from CUDA/Triton to TileLang for domestic GPU adaptation. The real signal is strategy: DeepSeek had spent less on agents and coding, but now names an agent product role; the post does not disclose V4's size, price, or benchmarks.

Why it matters: This is not the V4 launch, but it carries real signal: four confirmed departures, an April delay, a 100+ research team, and partial migration from CUDA/Triton to TileLang. HKR-H/K/R all pass; missing V4 specs, price, and benchmarks keeps it below launch-tier or p1.

Feb 15Sunday

Computing Life · Yage

OpenClaw deep dive: why it suddenly took off, and what it means for us

OpenClaw surged in late January 2026 because it plugged local coding agents into Slack, WhatsApp, and Feishu, giving non-technical users file access, command execution, and persistent memory in a chat UI. The article also names the costs: 12% of third-party skills contained malicious code, and the $CLAWD token scam took $16 million; the chat interface remains linear, low-density, and hard to observe. The real takeaway is not to copy OpenClaw blindly, but to reuse its unified context, file-based memory, and composable skills in a controllable stack like OpenCode.

Why it matters: This is more than a recap: it breaks down OpenClaw's adoption mechanism, downside, and reusable design pattern. HKR-H/K/R all pass with two hard facts—12% malicious skills and a $16M scam—but as a personal analysis rather than an official release or industry event, it lands in `+

Computing Life · Yage

OpenClaw Deep Dive: Why It Went Viral and What It Means for You

The post says OpenClaw went viral in late January 2026, changed names 3 times in one week, and a $CLAWD scam token took $16 million. It cites two concrete risks: 12% of third-party skills had malicious code, and some users exposed consoles to the public internet without passwords. The excerpt is truncated, but the core claim is distribution: OpenClaw put agentic AI into WhatsApp, Slack, and Lark for non-technical users.

Why it matters: HKR-H/K/R all pass: the viral arc is dramatic, the post includes a 12% malicious-skills figure and a specific exposed-console risk, and the distribution angle matters to agent builders. It is still a secondary deep-dive, not a primary launch or official research, so 78 and tiered

Feb 12Thursday

MIT Technology Review · AI

What’s next for Chinese open-source AI

MIT Technology Review says that after DeepSeek released R1 in January 2025, Chinese firms kept shipping open-weight models near top Western systems; Moonshot AI’s Kimi K2.5 was close to Anthropic Claude Opus on early benchmarks at about one-seventh the price. The post also says Qwen took over 30% of Hugging Face downloads in 2024 and surpassed Meta Llama in cumulative downloads by 2025–2026; the key shift is from a few general models to many fine-tunable, distillable variants.

Why it matters: All three HKR axes pass. This is not a launch, but it offers concrete market signals—~1/7 pricing, Hugging Face download share, and a clear thesis that Chinese open source is moving toward specialized, distillable variants—so it merits featured, not p1.

Jan 30Friday

Bloomberg Technology

US Lawmaker Says Nvidia Worked to 'Co-Design' DeepSeek Model

The Republican chair of the House China committee said Nvidia gave DeepSeek technical support that helped improve a breakthrough AI model despite US export controls on high-end chips to China. The RSS snippet discloses the allegation and the regulatory context, but not the model name, support mechanism, timeline, or evidence. The key issue is not chip shipment alone, but whether technical collaboration undermined the controls' intent.

Why it matters: Bloomberg gives this source-authority, and HKR-H / HKR-R land because the allegation is surprising and hits export-control compliance. HKR-K misses: the summary discloses no model name, mechanism, timeline, or evidence, so this stays low-featured rather than must-write.