Skip to content

DeepSeek

DeepSeek's model releases, open weights and technical reports — the bellwether for open-model price and performance.

194 picksRelated topicsQwenOpen sourceModel releases

Latest picks

21–40 of 194

Sep 10Thursday

AI HOT (Curated Pool)

DeepSeek Releases V4.1-Flash: New Causal Encoder-Decoder Architecture with Native Vision

DeepSeek V4.1-Flash is the smallest model in the new architecture family: a 552B MoE with 8B active params for input and 16B for output. It uses a Causal Encoder-Decoder design with native vision. KV cache drops to 1/4 of HBM and 1/8 of SSD storage vs the previous generation, and API pricing is lower. The post doesn't disclose exact pricing or vision benchmarks.

Why it matters: DeepSeek ships a new architecture — not a V4 refresh but a Causal Encoder-Decoder with native vision and dramatically reduced KV cache. The 552B total / 8B+16B active MoE config directly impacts deployment economics. Domestic Chinese flagship model release triggers the positiv...

AI HOT (Curated Pool)

DeepSeek V4.1-Flash: 1M context, FP4 KV cache, and cross-layer attention reuse

DeepSeek released V4.1-Flash, targeting long-context efficiency. It supports a 1M-token context window, uses FP4 KV cache to cut memory, and reuses attention across layers to reduce compute. The post does not disclose benchmark scores, parameter count, license, or API pricing—only the technical features are described.

Why it matters: DeepSeek drops V4.1-Flash with 1M context, FP4 KV cache, and cross-layer attention reuse — a concrete engineering combo that's worth a look. But no params, benchmarks, license, or pricing are disclosed, so we can't gauge real competitiveness. That gap keeps it at the featured ...

r/LocalLLaMA

DeepSeek V4.1 Flash: beats V4 Pro on benchmarks, cuts API price, and goes open source

DeepSeek released V4.1 Flash, a 552B MoE model that activates only 8B params on input and 16B on output. It uses a new asymmetric Causal-Encoder-Decoder architecture and scores above DeepSeek V4 Pro on benchmarks. KV cache size drops to 1/4 HBM and 1/8 SSD vs the previous gen, cutting agent-scenario cache costs. The API is live under model name deepseek-flash; V4 Pro will be routed to V4.1 Flash from Sep 14 noon Beijing time and billed at Flash pricing. New peak/off-peak prices start Sep 10 noon, with off-peak at half rate. Weights and a tech report are open on HuggingFace; DeepSeek invites contact for large-scale deployments needing a 2k-GPU cluster.

Why it matters: DeepSeek flagship model release with architectural change and concrete perf/cost numbers — policy treats this on par with US lab launches. All three HKR axes hit: the V4 Pro-beating score and cache shrinkage are hard info. Held back from P1 because only title + summary availab...

Hacker News front page

DeepSeek releases V4.1 Flash model on HuggingFace

DeepSeek published a new model, V4.1 Flash, on HuggingFace. The post doesn't disclose parameters, benchmarks, or architecture details. HN discussion is active at 853 points and 481 comments, but most are speculating based on the name—Flash usually signals a faster, lighter variant. I'd wait for a technical note before drawing conclusions.

Why it matters: DeepSeek model release with massive HN traction — a same-day must-cover. Score held back because the post lacks parameters, benchmarks, and architecture details, so the K axis doesn't land. The Flash suffix points to a lightweight/low-latency variant, directly relevant to infe...

AI HOT (Curated Pool)

DeepSeek releases V4.1-Flash, API pricing cut alongside

DeepSeek launched V4.1-Flash today, the smallest model in a new architecture family with native multimodal vision. The new design targets higher ceiling, faster inference, and larger throughput, and is meant to scale to bigger models. V4.1-Flash scores 90.9 on GPQA Diamond, 3471 Codeforces rating, and 36.8 on HLE. Set model name to deepseek-flash in the API; old V4 Flash and V4 Flash Vision Exp are offline and requests are temporarily routed to V4.1-Flash. DeepSeek also claims V4.1-Flash beats V4 Pro on performance, cost, and speed, so V4 Pro requests will be routed to V4.1-Flash starting Sep 14 and billed at Flash rates. API pricing is cut, but the post doesn't list the new numbers—check the pricing page.

Why it matters: DeepSeek ships the first model from its new architecture — vision-native, strong benchmarks, lower pricing. A substantive release from a top Chinese lab. HKR all hit, scored 86. Not higher because this is the smallest variant and the post doesn't detail the new architecture's ...

Sep 9Wednesday

AI HOT (Curated Pool)

DeepSeek reportedly hires CITIC Securities for STAR Market IPO, aims to file this year

Reuters reports DeepSeek has hired CITIC Securities to prepare for a STAR Market IPO, aiming to file this year and list next year. Fundraising size and target valuation are not yet set. The company plans to use IPO proceeds to expand computing infrastructure, boost model R&D and chip development, and strengthen talent incentives. Revenue in the first seven months of 2026 reached about 475 million yuan, roughly 10x its full-year 2025 revenue. DeepSeek is also raising a new funding round targeting a ~500 billion yuan valuation, following a June round that valued it at over $50 billion post-money. Investors include Tencent, CATL-linked entities, JD.com, and NetEase. Intense demand has spawned multi-layered SPVs reselling access with front-end fees exceeding 15%. Founder Liang Wenfeng is personally vetting final investor lists to block unknown entities ahead of the IPO.

Why it matters: DeepSeek's STAR board IPO push is an industry-level event, with a 10x revenue jump and in-house chip plans adding real substance. Not scoring higher because the fundraising amount and final valuation are still undecided — it's early in the process.

AI Chat-Group Daily (群聊日报)

OpenAI solves Navier-Stokes with 10K agents, but Codex data privacy debate steals the show

OpenAI deployed ~10K concurrent agents to solve the Navier-Stokes Millennium Problem in 88 hours, consuming 130B output tokens. But NYU mathematician Buckmaster publicly alleged OpenAI may have accessed his and collaborator Alpöge's unpublished drafts via Codex—their technical approaches overlapped heavily. OpenAI hasn't directly denied accessing Codex data, only stating they 'cannot rule out that de-identified data helped improve models.' The group debated whether personal subscriptions offer true zero data retention: only Team/Enterprise plans do. On the practical side, third-party benchmarks show Astra's xHigh effort costs more than High but scores slightly lower—High is the daily sweet spot. DeepSeek V4.1 Flash internal test model hits 340–450 tok/s with impressive SVG morphing quality, expiring Sept 10. GPT Image 2.5 launched with doodle canvas and native transparency. Codex's new experimental context management replaces compression with note-taking, cutting window-switch time from 27s to 1.8s.

Why it matters: A claimed Millennium Prize solution is already industry-shaking; the Buckmaster plagiarism accusation and OpenAI's non-denial push it into must-cover territory. Source is a curated group-chat digest, but it cites the official OpenAI post and a named mathematician's public alle...

Financial Times · Technology

DeepSeek fundraising frenzy spawns a shadow market with steep access fees

DeepSeek is raising a new round, and getting in is tough. The FT reports that some funds and family offices are buying access through intermediaries, paying a 2% management fee plus 20% performance carry—far steeper than standard terms. One investor, chasing a $5–10M allocation, also had to promise the middleman a bigger cut on future deals. DeepSeek says it hasn't authorized anyone to sell stakes, but the post doesn't disclose the round's total size or valuation. I'd take these off-market quotes with a grain of salt, but they show how hot demand is.

Why it matters: FT exclusive on a shadow market for DeepSeek fundraising, with concrete fee numbers and an investor anecdote—high signal. Score held at 78 because the round's total size and valuation are missing, and DeepSeek only gave a denial without further detail. HKR all hit, but the inf...

Sep 4Friday

r/LocalLLaMA

Qwen3.8-27b called the first local model users can 'blindly trust'

A Reddit user reports that Qwen3.8-27b ran 8+ hours of continuous agentic work without a single mistake, making it the first local model they trust like a frontier model. Another user confirmed 20-hour sessions with sub-agents and commit gates, and said the INT8 quant even solved a coding problem that DeepSeek V4 Flash couldn't fix. The post doesn't disclose specific task types or failure rates, but the community feedback points to noticeably better reliability in long-chain agent workflows. Take it as personal experience, not a systematic eval.

Why it matters: Two independent users report Qwen3.8-27b's stability in multi-hour agent tasks, one with a direct comparison to DeepSeek V4 Flash. But the post doesn't specify task types or failure criteria — this is community word-of-mouth, not a reproducible eval. Score 72 at the featured t...

Sep 3Thursday

Hacker News front page

Chen Danian returns with a 27B local model that trails DeepSeek-V4-Pro by only 1.3 points in CAICT's MCP benchmark

Chen Danian is back with StartLux, a company betting on local models. Its first release, StartLux-V1.0-27B-Preview, scored 39.25% in CAICT's MCP benchmark—second place, just 1.3 points behind the 1.6-trillion-parameter DeepSeek-V4-Pro. The 27B model runs on consumer PCs without the cloud and ranked first in location navigation, financial analysis, and browser automation. Two case studies: when calculating a two-year Microsoft stock return, Claude Sonnet 4.6 misidentified a trading day due to missing raw data; StartLux backtracked and got it right. Asked to search flights in a browser, Claude said it couldn't open a browser. Chen has publicly claimed local models will catch up with Claude in three years and take 80% of the market—StartLux is his bet on that thesis.

Why it matters: Chen Danian's first model lands second in CAICT's MCP benchmark, with a 27B parameter count that runs on consumer hardware and three first-place sub-scores — a concrete signal for the Agent space. Score capped at 82 because only benchmark results are available; the model isn't...

Sep 2Wednesday

AI Chat-Group Daily (群聊日报)

DeepSeek V4 Flash beats Sol in real-world use; Anthropic drops Fable 5.1

Community members ran two-month SBS comparisons and a week-long 5.1B-token workload on DSH + DeepSeek V4 Flash, concluding it feels better than GPT-5.6 Sol in real tasks. Sol overthinks and produces bloated output; V4 Flash is fast (2.3s first token) and cost ¥362.84 total. A 'subscription gym paradox' theory argues subscription-based harnesses quietly throttle usage while pay-per-token models don't. Anthropic launched Fable 5.1 with 75% cheaper cache reads, but Fable 5 scored below Opus 5. Also: Astra hits Critical cybersecurity tier, Anthropic's $35B compute deal, Qwen 3.8-Max-0902 benchmark run, Microsoft AI secretary setup, and Grok Bot hands-on.

Why it matters: The side-by-side data is solid — 5.1B tokens, ¥362.84 total spend, 2.3s first-token latency — but the source is an anonymized chat log, not an official release or reproducible benchmark. That caps the authority. HKR all hit, so featured is the right tier.

Aug 31Monday

AI HOT (Curated Pool)

DeepSeek open-sources V4-Flash-Vision-Exp, its first vision model, with multimodal agent performance near Opus-4.8

DeepSeek released V4-Flash-Vision-Exp on Hugging Face under MIT License—the first V4 model that accepts image inputs. The repo includes a minimal PyTorch inference implementation covering the vision encoder, MoE, DFlash Attention, and other core modules. It handles JPEG, PNG, GIF, and WebP for tasks like image captioning, screenshot OCR, and chart reading. Text-only performance matches the stable V4-Flash; multimodal agent benchmarks show a big jump, nearing Opus-4.8. This is an experimental version—it hit the API on Aug 21 and now has open weights.

Why it matters: DeepSeek's first multimodal V4 model, MIT-licensed, directly targeting Claude Opus-4.8 on agent tasks — a significant update from a major Chinese lab. Score held back because it's an experimental release and the post doesn't disclose specific benchmark numbers or comparison de...

Aug 27Thursday

AI Chat-Group Daily (群聊日报)

GLM-5.3-Flash and Qwen 3.8-Flash-Next debut on the same day, both drop global attention

GLM-5.3-Flash matches Claude Opus 4.8 across six benchmarks at $0.045 per task, but testers report slow speed and hallucinations. Qwen 3.8-Flash-Next opens weights, hitting 64.7 tok/s single-stream decode on DGX Spark and beating DeepSeek V4 Flash across the board. Both models adopt MoE plus sparse attention hybrids, ditching global attention. NVIDIA acquires Hugging Face for $12.9B, roughly 86x its annualized revenue, to control the open model distribution channel. Anthropic preps IPO at a ~$2T valuation target, with ~$559M adjusted operating profit in Q2, while OpenAI posted ~$12.3B operating loss in the same period. Altman admits on a podcast that OpenAI hasn't had its iPhone moment and has scrapped Sora and Atlas. RTX 30 series GPUs resume production using Samsung 8nm to avoid TSMC bottlenecks. Shopify's CEO complains Claude Code ignores AGENTS.md, causing split brain in teams. QUASAR-QAT quantizes all 496 linear layers of Qwen 3.8-27B to NVFP4, saving another 1.8GB VRAM. The group also discusses Sol's context bloat and the limits of fully automated PR merges.

Why it matters: Two domestic Flash models launched the same day — GLM-5.3-Flash posts strong benchmarks but slow real-world speed and hallucinations, while Qwen 3.8-Flash-Next is open-weight with measured inference speed beating DeepSeek V4 Flash. Concrete numbers, real-user feedback, archite...

Aug 26Wednesday

Hacker News front page

Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

Z.ai has claimed the previously anonymous Ox Alpha model, confirmed it belongs to the GLM series, and announced plans to open-source its weights. Ox Alpha scored close to DeepSeek on several benchmarks, but the company hasn't disclosed parameter count, training data, or a release date. The post doesn't spell out technical details or the license yet.

Why it matters: Z.ai claims the stealth Ox Alpha model, confirms it's GLM-series and will open weights. Bloomberg exclusive adds authority. Downside: no param count, training data, or timeline — still a teaser.

Aug 25Tuesday

Computing Life · Share · Yage

High Fidelity Nearby, Lossy at a Distance: The Shared Intuition Behind Three Long-Context Approaches

YaRN, DeepSeek V4, and DeepSeek-OCR tackle long-context bottlenecks at the coordinate, information-pathway, and input-representation layers respectively, all converging on the same intuition: keep nearby tokens high-fidelity, compress distant ones. YaRN applies frequency-partitioned interpolation to RoPE, letting Llama 2 7B reach 128K context with ~384 A100 GPU hours. DeepSeek V4 uses full-attention within a 128-token window and heavy compression plus sparse selection beyond it, cutting V4-Pro's per-token FLOPs to 27% of V3.2 at 1M context. DeepSeek-OCR compresses full pages into dense visual tokens, hitting 97% text accuracy at 10x compression. The three were developed independently by different teams. The post also flags a new challenge—maintaining positional awareness after compression—and outlines three solutions: dual-track position scales, document-wise coordinate resets, and Kimi K3's removal of positional encoding entirely.

Why it matters: Unifies three long-context approaches across different tech stack layers under one sharp intuition, backed by concrete numbers. Not a primary research release, and the excerpt cuts off mid-argument, which caps the score.

Aug 19Wednesday

Hacker News front page

modelmap: paste a HuggingFace model ID, get an interactive architecture map

modelmap.cc turns any HuggingFace model ID into an interactive architecture diagram—no weights downloaded. It instantiates the model on a meta device to get the structure, then runs a traced fake forward pass to infer tensor shapes. The homepage shows trending models like Qwen3.8-27B, DeepSeek-V4-Pro-0813, and Kimi-K3, plus classic reference architectures such as GPT-2, BERT, and DeepSeek-V3.1. Public repos work out of the box; gated ones work after adding a token. The post doesn't disclose whether it's open source, backend costs, or concurrency limits.

Why it matters: A practical Show HN tool that maps model architectures without downloading weights, featuring DeepSeek-V4-Pro and Kimi-K3 on the landing page. Hits H and K but lacks the discussion hook R requires — more of a bookmark than a conversation piece. Scores 72 at the featured thresh...

Aug 18Tuesday

Bloomberg Technology

DeepSeek, Qwen, and Moonshot worry US AI rivals not by leading in tech, but by being dramatically cheaper

Bloomberg argues the real threat from Chinese AI firms isn't superior model capability—it's cost. DeepSeek, Alibaba's Qwen, and Moonshot are delivering comparable performance at a fraction of the price, forcing US rivals to rethink their economics. The article credits more efficient training methods and hardware utilization, but doesn't provide specific pricing comparisons or recent benchmark figures. Treat this as an industry trend piece rather than a technical deep dive.

Why it matters: Bloomberg's trend piece reframes the China AI threat from capability catch-up to cost undercutting, which is a sharp angle. But without concrete pricing data or recent benchmarks, the information density isn't high enough to push the score further.

Aug 17Monday

AI HOT (Curated Pool)

Unitree to list on Shanghai STAR Market Aug 19, becoming A-shares' first humanoid robot stock

Unitree will list on the STAR Market Aug 19 at 150.80 yuan/share, implying a ~60.99 billion yuan market cap. The 219.23x P/E ratio far exceeds the industry average of 38.56x. It raised about 6.1 billion yuan, nearly half earmarked for robot model R&D. 2025 revenue hit 1.699 billion yuan with 278 million yuan net profit—one of the few profitable general-purpose robot firms globally. Q1 2026 revenue grew 68.49% YoY to 423 million yuan, though higher R&D and selling expenses dragged down adjusted net profit. Strategic investors include China's social security fund, DeepSeek, and CNPC.

Why it matters: Unitree's STAR Market IPO is a milestone—one of the few companies globally making a profit on general-purpose humanoid robots. The 219x P/E ratio, 5x the industry average, signals serious valuation debate. Score stays at 82 rather than higher because we only have the offering ...

New York Times Chinese

China pushes state-aligned datasets to shape global AI narratives

China’s National Data Administration released a blueprint this year aiming to make the country a data powerhouse by end of 2028, with plans to create “high-quality” datasets across 20+ strategic fields and share them globally. The Shanghai AI Laboratory has already published large multilingual datasets like “WanJuan” on GitHub and Hugging Face, covering history, law, and medicine, while requiring alignment with “mainstream Chinese values.” Analysts say the push serves two goals: pulling developing nations into China’s AI orbit and closing the gap in Chinese-language training data, which is fragmented across domestic silos and has forced labs to rely on distillation from stronger models. A Princeton study also found that Chinese state-media narratives have seeped into ChatGPT and Claude, making their Chinese-language responses more favorable toward Beijing.

Why it matters: NYT deep-dive on China's National Data Administration AI data blueprint, with a clear timeline and named projects — not a press release. Hits all three HKR axes, but it's a policy/ecosystem story rather than a product launch, so it lands in the 78-84 band. Not higher because i...

Financial Times · Technology

The next China shock will come from open-source AI

An FT op-ed argues that China's open-source LLMs are repeating the playbook of its manufacturing boom—turning tech into a commodity at ultra-low cost and eroding Western pricing power. It names DeepSeek and Alibaba's Qwen series as key examples, noting their open-source strategy builds ecosystems fast while US firms stay closed-source and capex-heavy. The post doesn't cite specific market share or enterprise adoption figures, so treat this as a directional argument.

Why it matters: FT op-ed frames Chinese open-source AI as a replay of the manufacturing shock — a catchy angle, but the body lacks hard data, making it more of a directional warning. H and R hit, K is missing evidence, landing right at the featured threshold.