Skip to content

#DeepSeek

0 today

Yesterday · Sep 29Tuesday

The Verge · AI

Will Chinese AI companies slow down? A top House Democrat wants answers

美国众议院中国问题特别委员会首席民主党人 Ro Khanna 致信 DeepSeek、阿里巴巴和 Moonshot AI,要求其提供追求"超级智能"与递归自我改进(RSI)的文件,并说明是否设有保障措施和"终止开关"。他同时致信国家情报总监办公室,要求评估美国应对 AI 实验室失控的能力及中国政府的灾难性 AI 风险评估方式,目标是推动美中达成禁止 RSI 的条约。

New York Times Chinese

Is China Really Stealing AI Technology from U.S. Companies?

The NYT breaks down why 'distillation' became a flashpoint in US-China AI talks. Anthropic and OpenAI accuse Chinese firms of distilling their proprietary models, but experts say the claim is overblown—distillation only captures output text, not source code or training internals. Chinese labs still need to build a strong base model first. The piece also notes Anthropic just paid $1.5B for using copyrighted data, and OpenAI faces a similar suit from the NYT. U.S. courts haven't ruled on whether distillation violates trade-secret law.

Why it matters: NYT's dissection of the 'distillation = theft' claim, backed by technical explanation and the Anthropic/OpenAI copyright cases as legal reference points. Docked slightly because it's a synthesis piece rather than original reporting, and the distillation mechanics still have a ...

Sep 27Sunday

Hacker News front page

DeepSeek open-sources DSec: elastic sandbox infrastructure for agentic training

DeepSeek published a paper on DSec, their internal sandbox system for training agents at scale. The idea is to let models practice with real tools in isolated environments that scale elastically. It handles 100K concurrent sandboxes, 11-second startup latency, and roughly $3 per sandbox. The post doesn't mention a code repo—only the arXiv paper is available so far.

Why it matters: DeepSeek open-sourced their internal agent-training sandbox infra with hard engineering numbers: 100k concurrent sandboxes, 11s cold start, $3/instance. Not a model release, so it stays below 85, but as a practical agent-training infra reference it's high-value for practitioners.

Sep 26Saturday

Hacker News front page

Token-space font compiler makes every LLM token the same width

An online tool that merges any font with a tokenizer to produce a font where every LLM token has equal width. Supports DeepSeek, OpenAI, Kimi, Qwen, and other tokenizers; outputs a single TTF that works in browsers, Discord, and Slack without extra scripts. Optional Noto fallback and color emoji. The post doesn't spell out rendering performance cost, but notes that browser shaping runs and line breaks can shift token boundaries.

Sep 24Thursday

New York Times Chinese

Xi Jinping bets on AI to revive China's economy, using state capital and industrial policy to catch the US

This NYT feature traces China's AI strategy. Xi Jinping declared in 2014 that China must be a top AI maker, not just a buyer. The first national AI plan followed in 2017, targeting global leadership by 2030. State-led funds poured over $184 billion into AI firms from 2000 to 2023, with hundreds of billions more pledged last year. Unlike the US focus on AGI, China prioritizes immediate deployment in factories, hospitals, and classrooms, aiming for AI tools to cover 90% of society by 2030. DeepSeek's breakthrough restored confidence, but regulators tightened controls on chatbots and restricted overseas travel for top AI entrepreneurs. Beijing dismisses global calls to slow AI development, seeing them as a way to lock in US dominance.

Why it matters: A well-sourced NYT long-read on China's state AI strategy, with hard numbers ($184B) and a clear US-vs-China framing. It's a policy overview rather than a breaking product or research drop, so it lands at 82—strong context piece, not a must-act-today item.

Sep 20Sunday

Hacker News front page

Pirate Face turns open models into torrents so they can't be deleted

Pirate Face mirrors open models from Hugging Face as magnet links and distributes them via P2P swarms. Every file carries the official Hugging Face SHA-256 hash, so you can verify the weights haven't been tampered with. Over 669k models are already synced, including DeepSeek V4.1 Flash and Qwen3.8-27B. If the original source goes down, the swarm keeps the model alive as long as peers are seeding. A drop-in Hugging Face-compatible API endpoint is planned. The post doesn't spell out seeder incentives or long-term hosting costs.

Sep 19Saturday

Sep 18Friday

New York Times Chinese

China Worries About a Different Kind of AI Risk

Kyle Chan argues in the NYT that the US and China worry about fundamentally different AI risks. US labs focus on recursive self-improvement and existential threats; Chinese policymakers see that takeoff as distant and instead fear deepfakes, political dissent, and social instability. Recent cases—OpenClaw data leak warnings, Mythos’s cyber offense capabilities, and an AI tool cracking WeChat accounts—are pushing Beijing to also take cyber and runaway AI risks more seriously. Chan suggests both sides start by acknowledging each other’s risk perceptions before jumping to arms-control talks.

Why it matters: NYT op-ed with concrete examples (OpenClaw data leak, Mythos cyber capability, WeChat-cracking tool) — not empty commentary. The US-China risk perception gap is a fresh angle with real information value. Downside: it's opinion, not primary reporting, and the excerpt is short w...

Hacker News front page

Cactus Needle 3: 8-29 MB automation models beat DeepSeek V4 Flash on tool calls

Cactus open-sourced Needle 3, an automation model for phones, wearables, robots, and other tiny devices. The whole model is 8-29 MB, built on Simple Attention Networks where every layer is a usable sub-network. A fine-tuned 4-layer sub-network beats DeepSeek V4 Flash on mobile tool-calling benchmarks; on extraction it matches models 2-3x its size. It runs fully offline at 400-4,000 tokens/s decode on a Raspberry Pi 5. The Python package supports tool calls, structured extraction, and text embeddings. The post does not disclose training data composition or fine-tuning cost.

Why it matters: A tiny model claims to match DeepSeek V4 Flash on mobile tool calls at 8-29MB, with a novel intelligence ladder architecture. Score capped below 85 because only the project page is available—no third-party benchmarks or production deployment stories to cross-validate the perfo...

Sep 17Thursday

Hacker News front page

DeepSeek-V4.1 Flash: Pushing the Limits of KV Cache Compression

This technical report breaks down DeepSeek-V4.1 Flash's architecture, which compresses KV Cache by another 4x. The 552B-parameter model activates only 8B params during prefill and 16B during decode, using just 20 of its 40 layers for prefill. Compression tactics include cross-layer KV sharing (CSA2), FP4 KV Cache, and sparse attention indexer optimizations. The author argues the changes are so substantial it should be called DeepSeek-V5 Flash. The post does not disclose training cost or release timeline.

AI HOT (Curated Pool)

DeepSeek-V4.1-Flash: 552B MoE multimodal model with KV cache compression

DeepSeek released V4.1-Flash, a 552B MoE multimodal model. The key feature is KV cache compression, which cuts memory usage during long-context inference. The paper just hit arXiv and doesn't disclose compression ratios or benchmarks yet, but the title says 'Pushing the Limits' — this is about inference efficiency.

r/LocalLLaMA

Dual RTX Pro setup hits 150 tok/s decode with Qwen and DeepSeek

A developer built a local inference rig with two RTX Pro GPUs running Qwen 3.8 Flash Next and DeepSeek V4 Flash. Decode hits 150 tok/s, prefill 10K tok/s, with room for 4 concurrent requests. He previously used an M3 Ultra 512GB and found it too slow for inference. The new setup handles DeepSeek's 1M context window, but Qwen's thinking tokens eat VRAM. Build took 2.5 days due to a PSU wiring fault. The post doesn't specify GPU model or total VRAM.

Sep 16Wednesday

Hacker News front page

DeepSeek V4.1 Flash scores 11/11 on Enclave's hacking benchmark

Enclave tested DeepSeek V4.1 Flash against 11 vulnerable targets and 4 patched controls. The model achieved code execution on all 11 targets for $4.65 total, using 2,349 Bash commands over nearly 2h38m of active model time. An audit confirmed 6 attacks followed the planned exploit path; 5 found alternate routes in the test environment, including a Grafana compromise in 52 seconds. Enclave has since closed those extra paths and will require re-runs for updated leaderboard rankings.

Why it matters: Security firm Enclave ran DeepSeek V4.1 Flash through its in-house pentesting benchmark: 11/11 vulnerable targets exploited, 4/4 patched targets held, $4.65 total cost. The result is striking, but it's a single vendor's own benchmark, not a third-party eval — hence the score s...

AI Chat-Group Daily (群聊日报)

DS V4.1 Flash search hallucination test, Astra over-engineering from old context, and GPT-6 Sol rumors

A controlled test with the same search tools shows DS V4.1 Flash hallucinated URLs after 22 tool calls, while GPT delivered real results in 4. Astra's over-engineering was traced to stale skills and memory driving extra work; behavior normalized after cleanup. GPT-6 Sol is rumored to launch this week with a quota reset. A DeepSeek kernel engineer's farewell post went viral, predicting AI will match hand-written kernels within 6–12 months.

AI HOT (Curated Pool)

Which DeepSeek V4 models accept images? OpenRouter breaks down the family

DeepSeek V4 is a model family, not a single model. OpenRouter's guide confirms only V4.1 Flash and V4 Flash Vision Exp accept image input; all others (V4 Pro 0813, V4 Flash 0731, etc.) are text-only. V4.1 Flash is the recommended choice with native vision support at $0.15/$0.60 per million tokens. V4 Flash Vision Exp is the pricier experimental option. The post also covers two integration methods: direct image input or using a separate vision model as a front-end.

Sep 15Tuesday

r/LocalLLaMA

DeepSeek V4.1 Flash Q4 hits 40 t/s on M3 Ultra with native DSpark multi-token prediction

A developer forked ds4 and tuned it for DeepSeek V4.1 Flash Q4 on a 512 GB M3 Ultra, lifting decode from 16.6 t/s to 31.3 t/s, and to 40.5 t/s with DSpark speculative decoding. The ~300 GB Q4 weights fit only the 512 GB M3 Ultra. The speedup comes from cutting Metal dispatch overhead: the 384-expert router went from 9 dispatches to 1, the shared expert gate+up+SwiGLU became a single kernel, and BF16 rounding moved inside producer kernels, removing ~770 re-round dispatches. At 300k context, compressed attention selection was the bottleneck; the fix scores only admitted blocks and uses a bounded radix select, keeping decode at 90% of the 8k rate. DSpark verifies 6 tokens per step with shared weight streams and overlapped Engram fetches, dropping verify latency from 177 ms to 112 ms. Output is byte-identical to upstream under greedy decode, with SHA-256 manifests provided. The branch is M3 Ultra only because the optimizations rely on measured behavior of this specific chip's dual-die memory, 80-core GPU scheduling, and Metal dispatch characteristics.

Why it matters: A solid local-inference optimization post: DeepSeek V4.1 Flash Q4 on M3 Ultra goes from 16tps to 40tps via Metal command-buffer merging and native DSpark speculative decoding. Concrete technical detail, directly useful for the local-LLM crowd. Score stays at 72 because the aud...

AI Chat-Group Daily (群聊日报)

Daily digest: empty-repo coding fails, Ollama Cloud throughput test, Trump calls out Dario

今天最直观的教训来自 @搞仁义毛义仁:给 Astra 一个空白 C++ 仓库,代码写得一塌糊涂;把积累了大量 code review 经验的 GacUI 上下文导进去,质量立刻飙升。这说明模型不是不会写,是得用具体规则去“规训”。@今天群内信息量极大 实测了 Ollama Cloud 跑 DeepSeek V4.1 Flash,解码吞吐是官方 API ...

AI HOT (Curated Pool)

Fireworks benchmarks DeepSeek-V4.1-Flash: matches GPT-6 Astra on DeepSWE at 1/15th the cost

Fireworks ran a full benchmark suite on DeepSeek-V4.1-Flash. On DeepSWE, it scores 74.34% pass@1, in the same band as GPT-6 Astra at 74.12%, but costs $0.43 per task—15x cheaper. The model uses a 552B MoE with a split activation design: 8B active for input, 16B for output, plus KV cache optimizations. On Terminal-Bench 2.1 it trails Astra by 1 point while costing 12x less. The post mentions an HLE and oracle router eval but does not disclose the actual scores.

Why it matters: DeepSeek V4.1-Flash matching GPT-6 Astra on DeepSWE at an order-of-magnitude lower cost is the strongest price-performance signal this week. Docked because the source is Fireworks' own benchmark, not an independent eval, and the body is truncated by a cookie wall with no full ...

Sep 13Sunday

Computing Life · Share · Yage

DeepSeek Engram: Moving static knowledge out of GPU via lookup tables to free up reasoning capacity

DeepSeek V4.1 Flash assigns 196B parameters to Engram, a conditional memory module stored in host RAM instead of GPU VRAM. Lookup keys are built from the last few tokens, so addresses are known ahead of time; RDMA prefetch hides the transfer latency behind computation. In the paper's self-reported results, reasoning gains outpace knowledge gains: BBH +5.0, needle-in-a-haystack retrieval jumps from 84.2 to 97.0. The mechanism: offloading static local mappings frees up early-layer compute and attention budget for multi-step reasoning and long-range dependencies. The team also introduces 'sparsity allocation'—experiments suggest ~20–25% of sparse capacity going to Engram works best, though no independent replication exists yet. Qwen3.8 Flash-Next adopts a similar design, signaling that external static memory is entering the mainstream.

Why it matters: DeepSeek packed a 196B-parameter lookup module called Engram into V4.1 Flash—no matmuls, no GPU memory residency, using hash keys and RDMA prefetch to decouple knowledge retrieval from compute. The self-reported gains are stronger on reasoning than on knowledge QA, which is co...

Sep 12Saturday

Latent Space

DeepSeek V4.1-Flash: a 763B encoder-decoder MoE with 8B prefill, 16B decode, and native vision

DeepSeek dropped V4.1-Flash on Sep 10. Despite the 4.1 label, Sebastian Raschka called it a V5-level rewrite. It's a 763B total-parameter MoE with a causal encoder-decoder split: 8B active for prefill, 16B for decode, yielding 1–2% sparsity and up to 8× smaller KV cache vs V4 Flash. Native vision is built in, and V4 Pro has been quietly retired. The post doesn't include benchmark tables but argues current evals miss the point—the real advance is context efficiency for long-running agents.

Why it matters: DeepSeek drops V4.1-Flash with a 763B causal encoder-decoder MoE, 8B/16B active params, 1%-2% sparsity, and vision. Sebastian Raschka says it should've been V5. This is a major domestic flagship architecture update with a cross-source cluster forming. HKR all hit. Not 90+ yet ...

Computing Life · Share · Yage

Anthropic alleges 300K requests silently rerouted, exposing real production data

Anthropic's September threat report says a team used 5,380 fake accounts to reroute ~300K user requests to Claude over 10 days. The exposed data includes a pharma firm's multi-country budget sheet, live Telegram and Feishu credentials, and police ID checks. Independent researcher Shou claims to have bought a 6TB dataset with SSH keys and cloud tokens—single-source, unverified. Anthropic estimates 180M+ unauthorized distillation calls: Alibaba 151M, Moonshot ~23M, DeepSeek 12.1M. DeepSeek specifically routes requests containing Claude Code markers to reasoning models. DeepSeek's terms allow training on inputs; Kimi's web UI has no opt-out toggle—users must email and wait 5–7 business days. Technical defenses protect model outputs, not user inputs. The named companies haven't publicly responded; attribution rests solely on Anthropic's account.

Why it matters: Anthropic's unilateral investigation, but the leaked samples — pharma budget tables, police ID checks — are concrete and alarming. All three HKR axes hit; security incidents carry natural resonance. Deduction: attribution is single-source, named companies haven't responded, nu...

AI HOT (Curated Pool)

DeepSeek V4.1-Flash open-sourced: CED architecture cuts prefill cost for coding agents

DeepSeek released open weights for V4.1-Flash, a 552B MoE model with a Causal Encoder-Decoder architecture tuned for coding agents. It splits compute asymmetrically: 8B active params during prefill, 16B during decode, plus improved KV cache efficiency. On Terminal Bench 2.1 it hits 90.6; on Automation-Bench it scores 54.8—better than V4-Pro but still failing roughly half of complex workflows, so keep a human in the loop. It is also DeepSeek's first non-experimental model with native image input. Chartography reaches 78.9, but ZeroBench logical reasoning over images is only 49. DeepSeek has already retired V4-Flash traffic and will reroute V4-Pro traffic to V4.1-Flash starting September 14.

Why it matters: DeepSeek open-sourced V4.1-Flash, a 552B MoE that splits prefill and decode via CED architecture, directly targeting coding agent latency. Terminal Bench 2.1 scores are concrete, and Baseten's analysis adds deployment perspective. Not 85+ because this is a third-party writeup ...

Sep 11Friday

AI Chat-Group Daily (群聊日报)

Anthropic report confirms DeepSeek and Kimi silently routed user requests to Claude; Pro 20x halts new sign-ups same day

Anthropic's September threat report reveals DeepSeek and Moonshot (Kimi) silently forwarded user requests to Claude without consent, exposing code and credentials to third parties. A 6TB data leak from the same router contained SSH keys, cloud credentials, and GitLab tokens capable of compromising 7 government entities and 19 enterprises. The report also names seven Chinese labs—including Alibaba, Zhipu, and Xiaomi—for large-scale distillation attacks on Claude totaling over 180 million interactions. The same day, Anthropic paused new $200 Pro 20x subscriptions as Astra capacity tightened. DeepSeek launched V4.1 Flash, merging its Pro and Flash lines; V4 Pro sunsets September 14. Zhipu partnered with Hangzhou's Shangcheng district on a city-wide coding subsidy, offering 51% off annual personal plans.

Why it matters: Anthropic official threat report + 6TB leak evidence + seven Chinese labs named for distillation — three threads converging into a security event cluster. All three HKR axes hit, with enough density and industry impact for featured. Not scoring higher because this is a curated...

AI HOT (Curated Pool)

DeepSeek V4.1 Flash Tested: Price Drop, Native Vision, Game & City Gen

The article body is blocked by WeChat; only the title remains. It claims DeepSeek V4.1 Flash has a big price drop, native vision, and was tested on game and city generation tasks. The post does not disclose the exact price cut, vision specs, or generation quality.

Hacker News front page

Benzi benchmarks code-fixing harnesses against Claude Code and DeepSeek on lines read, time, and cost

Benzi tested four setups on 24 real GitHub issues: Benzi with Sonnet or DeepSeek, Claude Code, and the DeepSeek native harness. The headline metric is source lines read per fix—Benzi + Sonnet read 9,125 lines total, Claude Code read 20,704, and the DeepSeek harness read 43,598. Cost-wise, Benzi + Sonnet spent $17.96 for all 24 bugs vs. $39.54 for Claude Code; Benzi + DeepSeek cost just $2.66. On SWE-bench Verified, Benzi resolved 78.2% of 500 issues at under 10¢ per fix. The post doesn't explain how Benzi's code intelligence achieves the lower read counts, and it doesn't break down latency details.

Computing Life · Share · Yage

DeepSeek V4.1 Flash shifts the long-context cost battle from compute to memory

DeepSeek released V4.1 Flash, compressing the global KV cache to about 1/4 and persistent KV cache to 1/8 of the previous generation, while cutting cache-hit input prices by roughly 60%. The tech report argues that sparse attention has already squeezed compute costs low; what now drags down long-running agent tasks is HBM filling up, SSD offloading, and bus transfers. Flash tackles this with 4-bit storage, cross-layer global-cache reuse, and dropping sliding-window disk writes, shifting the cost center from compute to the memory hierarchy. On deployment, DeepSeek initially planned to route all V4 Pro traffic to Flash immediately, but pushed the cutover to Sept 14 after developer pushback. The report also flags potential position-selection bias from layer reuse and degradation risks in extreme long-context cache reconstruction. All throughput and reduction figures are self-reported, not independently verified.

Why it matters: DeepSeek V4.1 Flash isn't a routine price cut — it compresses KV cache to 1/4–1/8 of the previous gen and slashes cache-hit input pricing by 60%. The tech report argues that sparse attention already tamed compute; the bottleneck is now VRAM and bus transfers. For agent builder...

Sinocism (Bill Bishop)

Anthropic says DeepSeek, Xiaomi, and Moonshot used Claude outputs for model distillation

Anthropic's September threat-intel report calls out DeepSeek, Xiaomi, and Moonshot for piping user-model conversations into Claude and using Claude's replies as training data for distillation. The exchanges reportedly contained sensitive info from individual users, multinationals, and state-affiliated actors. Anthropic says this violates PRC law and suggests sharing detailed findings with China's Ministry of Public Security via the FBI. The post doesn't disclose the volume of conversations or the time range involved.

Why it matters: Anthropic's official threat intel report names three major Chinese AI labs for distilling Claude with sensitive user data — an industry-level security incident. Strong cross-source signal, all three HKR axes hit. The slight deduction is because we only have Sinocism's second-h...

AI HOT (Curated Pool)

DeepSeek V4.1 Flash scores 40 on Intelligence Index, surpassing DeepSeek V4 Pro 0813 as new flagship

Artificial Analysis reports DeepSeek V4.1 Flash hits 40 on the Intelligence Index, edging out DeepSeek V4 Pro 0813 as DeepSeek's top-scoring model. It uses 8B active parameters for input and 16B for output, supports 1M token context, and is MIT-licensed. The post doesn't disclose inference speed or pricing, so I'd hold off on cost-performance claims for now.

Why it matters: New DeepSeek flagship beats its own predecessor and ships MIT-licensed — strong dev appeal. Held back from 85 because the post doesn't disclose inference speed, leaving real-world experience an open question.

Sep 10Thursday

AI HOT (Curated Pool)

WorkBuddy Launches DeepSeek V4.1-Flash with Two-Week Free Trial

WorkBuddy now offers DeepSeek V4.1-Flash on its platform with a two-week free trial. The model is available via DeepSeek API and supports native multimodal input. The post doesn't spell out improvements over prior versions or pricing.

AI HOT (Curated Pool)

DeepSeek-V4.1-Flash lands on SiliconFlow, a 552B MoE with 1M context window

SiliconFlow launched DeepSeek-V4.1-Flash on Day 0. It's a 552B MoE model with ~8B active params during prefill and ~16B during decode, native vision, and a 1M-token context window. KV cache footprint is about 1/4 of V4 Flash, which helps with deployment cost. MIT license keeps commercial use straightforward.

Why it matters: Same-day availability of DeepSeek V4.1-Flash on SiliconFlow, with KV cache reduced to 1/4 of V4 Flash — a clear deployment cost signal. Score held at 78 because this is a platform availability announcement; no benchmarks or real-world performance data yet.

AI HOT (Curated Pool)

DeepSeek V4.1-Flash cuts KV cache memory for AI agents to a quarter of its predecessor

DeepSeek released V4.1-Flash, a 552B-parameter model built to slash memory costs for AI agents. Its KV cache in fast GPU memory is about a quarter the size of V4-Flash, and the offloaded portion shrinks to roughly an eighth. The model splits into an encoder and decoder: only 8B parameters activate per token during input processing, versus 16B during text generation, nearly halving input compute. It supports 1M-token contexts and stores the main KV cache in FP4. On the DeepSWE v1.1 coding benchmark it scores 74.2%, narrowly beating Anthropic Opus 5 and OpenAI GPT-5.6 Sol, but it still trails on complex scientific tasks and image analysis. Weights are on Hugging Face under the MIT license. The post does not disclose inference latency or specific hardware requirements.

Why it matters: DeepSeek drops V4.1-Flash targeting agent memory costs — KV cache down to 1/4 of predecessor. Concrete architecture numbers, not vapor. Held at featured rather than p1 because only one source so far (no cross-source cluster yet) and the post doesn't disclose real latency/throu...

Bloomberg Technology

DeepSeek's New Low-Cost Model Deals a Fresh Blow to OpenAI, Z.AI

Bloomberg reports DeepSeek released a new low-cost model, directly challenging OpenAI and Z.AI. The model's lower cost may force competitors to cut prices or adjust strategy. The post does not disclose specific specs, pricing, or release timeline, but the title confirms this is a price-war move against top players.

AI HOT (Curated Pool)

DeepSeek V4.1-Flash drops with native vision and a big price cut

DeepSeek released V4.1-Flash with a new Causal Encoder-Decoder architecture and native vision understanding — no separate vision model needed. It's a 552B MoE, activating 8B params for input and 16B for output. The post doesn't disclose the exact price cut or benchmark numbers, so I'd wait for third-party evals before getting excited.

Why it matters: DeepSeek ships a new model with a causal encoder-decoder architecture that replaces bolt-on vision components. 552B total params, only 8B/16B active during inference. The post claims a price cut but gives no specific numbers or benchmarks, so the score stays below 85 until thi...

AI HOT (Curated Pool)

DeepSeek Releases V4.1-Flash: New Causal Encoder-Decoder Architecture with Native Vision

DeepSeek V4.1-Flash is the smallest model in the new architecture family: a 552B MoE with 8B active params for input and 16B for output. It uses a Causal Encoder-Decoder design with native vision. KV cache drops to 1/4 of HBM and 1/8 of SSD storage vs the previous generation, and API pricing is lower. The post doesn't disclose exact pricing or vision benchmarks.

Why it matters: DeepSeek ships a new architecture — not a V4 refresh but a Causal Encoder-Decoder with native vision and dramatically reduced KV cache. The 552B total / 8B+16B active MoE config directly impacts deployment economics. Domestic Chinese flagship model release triggers the positiv...

AI HOT (Curated Pool)

DeepSeek V4.1-Flash: 1M context, FP4 KV cache, and cross-layer attention reuse

DeepSeek released V4.1-Flash, targeting long-context efficiency. It supports a 1M-token context window, uses FP4 KV cache to cut memory, and reuses attention across layers to reduce compute. The post does not disclose benchmark scores, parameter count, license, or API pricing—only the technical features are described.

Why it matters: DeepSeek drops V4.1-Flash with 1M context, FP4 KV cache, and cross-layer attention reuse — a concrete engineering combo that's worth a look. But no params, benchmarks, license, or pricing are disclosed, so we can't gauge real competitiveness. That gap keeps it at the featured ...

r/LocalLLaMA

DeepSeek V4.1 Flash: beats V4 Pro on benchmarks, cuts API price, and goes open source

DeepSeek released V4.1 Flash, a 552B MoE model that activates only 8B params on input and 16B on output. It uses a new asymmetric Causal-Encoder-Decoder architecture and scores above DeepSeek V4 Pro on benchmarks. KV cache size drops to 1/4 HBM and 1/8 SSD vs the previous gen, cutting agent-scenario cache costs. The API is live under model name deepseek-flash; V4 Pro will be routed to V4.1 Flash from Sep 14 noon Beijing time and billed at Flash pricing. New peak/off-peak prices start Sep 10 noon, with off-peak at half rate. Weights and a tech report are open on HuggingFace; DeepSeek invites contact for large-scale deployments needing a 2k-GPU cluster.

Why it matters: DeepSeek flagship model release with architectural change and concrete perf/cost numbers — policy treats this on par with US lab launches. All three HKR axes hit: the V4 Pro-beating score and cache shrinkage are hard info. Held back from P1 because only title + summary availab...

Hacker News front page

DeepSeek releases V4.1 Flash model on HuggingFace

DeepSeek published a new model, V4.1 Flash, on HuggingFace. The post doesn't disclose parameters, benchmarks, or architecture details. HN discussion is active at 853 points and 481 comments, but most are speculating based on the name—Flash usually signals a faster, lighter variant. I'd wait for a technical note before drawing conclusions.

Why it matters: DeepSeek model release with massive HN traction — a same-day must-cover. Score held back because the post lacks parameters, benchmarks, and architecture details, so the K axis doesn't land. The Flash suffix points to a lightweight/low-latency variant, directly relevant to infe...

AI HOT (Curated Pool)

DeepSeek releases V4.1-Flash, API pricing cut alongside

DeepSeek launched V4.1-Flash today, the smallest model in a new architecture family with native multimodal vision. The new design targets higher ceiling, faster inference, and larger throughput, and is meant to scale to bigger models. V4.1-Flash scores 90.9 on GPQA Diamond, 3471 Codeforces rating, and 36.8 on HLE. Set model name to deepseek-flash in the API; old V4 Flash and V4 Flash Vision Exp are offline and requests are temporarily routed to V4.1-Flash. DeepSeek also claims V4.1-Flash beats V4 Pro on performance, cost, and speed, so V4 Pro requests will be routed to V4.1-Flash starting Sep 14 and billed at Flash rates. API pricing is cut, but the post doesn't list the new numbers—check the pricing page.

Why it matters: DeepSeek ships the first model from its new architecture — vision-native, strong benchmarks, lower pricing. A substantive release from a top Chinese lab. HKR all hit, scored 86. Not higher because this is the smallest variant and the post doesn't detail the new architecture's ...

AI Chat-Group Daily (群聊日报)

Chat digest: Astra capacity crunch, DeepSeek V4.1 Flash benchmarks, Codex quota bug, and why xHigh saves more credits than Medium

OpenAI's Tibo publicly admitted unprecedented Astra demand and may pause new Pro subscriptions; users report lag even during off-peak hours and frequent WebSocket disconnects. DeepSeek V4.1 Flash scored 81.2 on OpenDesign's design benchmark—98% of Astra's quality at 1.4% of the cost—but the API's mandatory training clause and not-so-cheap real pricing gave users pause. A Codex quota display bug caused panic today; Tibo promised compensation but most users never got it. A counterintuitive finding: xHigh mode actually consumes fewer total credits than Medium because it plans more accurately and loops less. Also: Jacob Coxon quit with a warning about AI arms-race risks, Apple announced the foldable iPhone Duo starting around $2,800, and the Navier–Stokes proof cost roughly $15M in API fees.

Sep 9Wednesday

AI HOT (Curated Pool)

DeepSeek reportedly hires CITIC Securities for STAR Market IPO, aims to file this year

Reuters reports DeepSeek has hired CITIC Securities to prepare for a STAR Market IPO, aiming to file this year and list next year. Fundraising size and target valuation are not yet set. The company plans to use IPO proceeds to expand computing infrastructure, boost model R&D and chip development, and strengthen talent incentives. Revenue in the first seven months of 2026 reached about 475 million yuan, roughly 10x its full-year 2025 revenue. DeepSeek is also raising a new funding round targeting a ~500 billion yuan valuation, following a June round that valued it at over $50 billion post-money. Investors include Tencent, CATL-linked entities, JD.com, and NetEase. Intense demand has spawned multi-layered SPVs reselling access with front-end fees exceeding 15%. Founder Liang Wenfeng is personally vetting final investor lists to block unknown entities ahead of the IPO.

Why it matters: DeepSeek's STAR board IPO push is an industry-level event, with a 10x revenue jump and in-house chip plans adding real substance. Not scoring higher because the fundraising amount and final valuation are still undecided — it's early in the process.