Skip to content

DeepSeek

DeepSeek's model releases, open weights and technical reports — the bellwether for open-model price and performance.

194 picksRelated topicsQwenOpen sourceModel releases

Latest picks

41–60 of 194

Aug 16Sunday

Hacker News front page

A leaderboard tracking 30 model cards to see which benchmarks frontier labs actually report

This project scanned 30 model cards from 11 orgs and counted how often 79 benchmarks are mentioned—it measures vendor attention, not benchmark quality. MATH-500 and Arena-Hard are near ceiling, losing discriminative power. DeepSeek's own models gained 40.6 points on AIME and 25.4 on LiveCodeBench in 26 days. Six benchmarks, including BrowseComp and SWE-bench Pro, are reported by at least 4 orgs but have no readable scores. The newer APEX-Agents already appears in 3 independent cards, though scores couldn't be read either.

Why it matters: Scans 30 model cards from 11 orgs, measuring vendor attention rather than benchmark quality — a useful lens. Concrete numbers like MATH-500 near-saturation and DeepSeek's 40.6-point AIME jump in 26 days will spark discussion. Docked because it's a personal project with limited...

Aug 15Saturday

AI Chat-Group Daily (群聊日报)

DeepSeek Harness open-sourced: architecture debates, security flaws, and V4 Pro's chaotic launch

DeepSeek open-sourced DSH, an agent harness where everything is a plugin and the agent loop itself can be swapped at runtime. A deep-dive analysis found this is the only structural edge over declarative frameworks like Codex—betting on self-evolving agents. Four PoC security flaws were also disclosed, including a sandbox escape that exposes SSH keys and .env files. Meanwhile, DeepSeek V4 Pro had a messy launch with inconsistent model versions, pulled weights, and poor real-world instruction following. Gemini 3.7 Flash landed quietly with notable coding gains. OpenAI Astra's math breakthrough faced plagiarism accusations.

Why it matters: DeepSeek officially open-sourced DSH agent framework with a structural differentiator — 'everything is a plugin.' Real test data and same-day community contributions push HKR all three. Score capped at 78 because the source is a chat-group digest, not a first-party announcemen...

Computing Life · Share · Yage

Give DeepSeek V4 a 3B vision front-end — it works today

DeepSeek V4's API rejects image input, but Liquid AI's new LFM2.5-VL-3B can serve as a perception layer. The author tested it zero-shot on garage camera data — 94.3% accuracy for door open/close, 1.5s locally. Perception runs on-device, images never leave, only structured text hits the cloud for reasoning. The latency, privacy, and cost advantages of this split architecture hold even if DeepSeek adds native vision later. One gap: extracting a reliable confidence signal from natural-language output — the post doesn't detail a calibration method.

Why it matters: The author built a perception frontend for DeepSeek V4 using Liquid's LFM2.5-VL-3B, tested it on a real garage camera with 94.3% accuracy and 1.5s latency — concrete data, clear architecture. Also traces DeepSeek's vision research history (VL2, the retracted Thinking with Visu...

Computing Life · Share · Yage

Same Model, 20-Point Gap: DeepSeek's Harness Dependency and the Hidden Ceiling of Synthetic Data

DeepSeek V4 Flash scored 82.7 on its official harness but dropped to a 46.7% pass rate on third-party setups—a 20-point gap from the same model. A joint paper from Stanford, UC Berkeley, and others explains why: training an agent with a single LLM as the user simulator causes the policy to exploit the simulator's narrow response patterns, with policy entropy collapsing from 1.9 to 0.4 nats. DeepSeek lacked a first-party product to collect real interaction data, so its training relied entirely on synthetic environments with limited behavioral diversity. DSH, released on August 13, is their answer—it makes the agent loop a hot-swappable plugin so the training environment can co-evolve with the policy, an engineering implementation of the paper's Co-Training approach.

Why it matters: Hits all three HKR axes: the 20-point gap is intriguing, the evidence chain from official footnotes to third-party repros is solid, and it directly resonates with agent developers. Capped below 85 because this is a benchmarking methodology exposé, not a model or product launch...

Aug 14Friday

Hacker News front page

DeepSeek V4 Pro goes GA with peak/off-peak API pricing

DeepSeek V4 Pro is now GA, with major agent workflow gains and adjustable reasoning effort—low for simple tasks, high for daily agent work, max for complex ones. It natively supports the OpenAI Responses API and one-click Codex setup. API pricing shifts to peak/off-peak on Aug 16: off-peak is 50% cheaper. Model names stay the same; try it via Expert Mode on the app.

Why it matters: V4 Pro GA with agent hardening and a thinking-effort dial is a real feature update that matters to developers building automation on DeepSeek. Held below 85 because the post doesn't disclose GA benchmark comparisons or the actual peak/off-peak price spread — the info density i...

AI HOT (Curated Pool)

DeepSeek V4 Pro lands on SiliconFlow with 1M context and three inference tiers

DeepSeek V4 Pro is now available on SiliconFlow with Day-0 support, a 1M context window, and three inference intensity levels. It targets coding, tool use, and agent workflows under the MIT license. Pricing: $1.32/M input, $3.96/M output, $0.44/M cache hit. A Flash variant is also live for cost-sensitive production use. The post does not disclose parameter count or architecture details.

Why it matters: DeepSeek V4 Pro lands on SiliconFlow day one with 1M context, tiered reasoning, MIT license, and clear pricing — solid signal density. Held below 85 because this is a platform availability announcement without benchmarks or user reports yet; sits right at the featured threshold.

Financial Times · Technology

OpenAI and Anthropic in price war as Chinese AI rivals gain ground

FT reports OpenAI and Anthropic are slashing prices to win enterprise customers, pressured by cost-competitive Chinese models like DeepSeek. Both are pushing cheaper, smaller models while leaning on premium subscriptions and IPO expectations to support valuations. The post doesn't spell out exact price cuts or effective dates—it's more a trend piece.

Why it matters: FT's trend piece has narrative value, but the body lacks specific price-cut figures or timelines — the information density isn't hard enough. H and R hit, K is missing; it just clears the featured threshold at 72.

Aug 13Thursday

Computing Life · Share · Yage

Every coding agent form factor shift is chasing the same thing: execution data

DeepSeek is hiring an Agent Harness PM, signaling it's filling the gap of not having its own coding tool runtime. The article argues that desktop apps, managed cloud agents, and remote control are all moves to capture execution data. Interfaces converge because they're cheap to copy; execution layers diverge because that's where the data moat is. Without a first-party harness, DeepSeek lacks real-world coding feedback to improve its models. Judge a coding agent by who controls the execution environment, who sees the data, and who's in the data flywheel—not by feature checklists.

Why it matters: A sharp industry analysis that uses DeepSeek's hiring move and LangChain test data to argue 'harness = data moat.' Hits all three HKR axes, but as an opinion piece rather than a primary release, scored at the lower end of the 78-84 band per policy.

Computing Life · Share · Yage

DeepSeek open-sources DSH: agent loop as a hot-swappable plugin, paving the way for self-evolving agents

DeepSeek released its first agent harness, DSH, as open source on August 13. Unlike Codex or Claude Code, DSH treats the agent loop itself as a plugin that can be swapped at runtime. The Cordis runtime handles hot reloads, dependency notifications, and transactional rollbacks. For everyday coding, declarative plugins plus a quick restart are enough—DSH's imperative model adds complexity. But if you want an agent that can generate new tools or replace its own control flow mid-run, DSH is the only option with the infrastructure in place. The post does not disclose performance benchmarks or production-scale data.

Why it matters: DSH makes the agent loop itself a hot-swappable plugin — a real architectural difference, not marketing. But this is a third-party analysis, not an official launch, and DSH has zero production track record yet. Defaulted to the lower band per policy.

Hacker News front page

DeepSeek V4 Pro 0813 listed on OpenRouter at $0.435/1M input tokens

DeepSeek V4 Pro 0813, the GA release of a large MoE model, is now available on OpenRouter. It offers a 1M-token context window, priced at $0.435/1M input and $0.87/1M output. Only one provider hosts it, so OpenRouter forwards requests directly without routing. The page does not disclose throughput, latency, TTFT, or benchmark results — real-world numbers are still needed before judging value.

Why it matters: DeepSeek V4 Pro GA lands on OpenRouter with a 1M-token context window and $0.435/1M input pricing — concrete, verifiable new info. Not an 85 because we only have the OpenRouter listing; no official blog post or third-party evals yet, so I'm discounting slightly.

Aug 10Monday

AI HOT (Curated Pool)

Unitree Robotics launches IPO subscription, becoming the first humanoid robot stock on the A-share market

Unitree Robotics opened its IPO subscription on Aug 10 at 150.80 yuan/share, targeting a market cap of ~60.99 billion yuan and raising ~6.1 billion yuan. The P/E ratio of 219.23x far exceeds the industry average of 38.56x. Strategic investors include China's social security fund, DeepSeek, and CNPC. Unitree posted 1.699 billion yuan in 2025 revenue with 278 million yuan net profit, making it one of the few profitable general-purpose robotics firms globally. Nearly half of the raised funds (2.022 billion yuan) will go toward intelligent robot model R&D.

Why it matters: Unitree's IPO is the first pricing event for humanoid robotics on A-shares — the ¥61B market cap and 219x P/E give the sector a concrete valuation anchor. Hard financials and R&D allocation details add substance. Capped slightly because this is a financial event, not a tech br...

Aug 9Sunday

Hacker News front page

DeepSeek-V4 Latent Reasoning ships as a self-contained model, not an adapter

Nicholai Mitchko turned the CoLaR latent reasoning head into a single deployable model. It uses a DeepSeek-V4-Flash-0731 backbone quantized to NVFP4 (~79 GiB/GPU at TP=2) and a 35.7M-param reasoning head. BBH zero-shot aggregate is 0.94, with perfect scores on multi-step state tracking but 0.26 on Dyck languages. A forked vllm runtime serves it, with per-request reasoning depth control via HTTP headers.

Why it matters: Turning CoLaR latent reasoning from an external adapter into a full model has engineering value, and BBH zero-shot 0.94 is solid. But it's a personal research blog with no cross-source verification and a narrow audience, so it lands right at the featured threshold.

Aug 8Saturday

Hacker News front page

DeepSeek V4 Flash 0731 hits 61.4% on ARC-AGI-2 at $0.04 per task

DeepSeek submitted V4 Flash 0731 to ARC Prize's verified leaderboard with three reasoning variants. The max-effort variant scores 89.0% on ARC-AGI-1 Semi-Private at $0.02 per task and 61.4% on ARC-AGI-2 at $0.04 per task. The low variant drops to 46% on ARC-AGI-2, showing how much reasoning budget matters. The post does not disclose model size, architecture details, or ARC-AGI-3 results.

Why it matters: DeepSeek submitted V4 Flash 0731 to the ARC Prize leaderboard, hitting 61.4% on ARC-AGI-2 — the highest public score so far — at $0.04 per task. Three inference budgets with scores and costs are provided, making it information-dense. Not scored higher because this is a leaderb...

Dwarkesh Patel podcast

The Era of Continual Learning: AI That Learns From Every Session

Dwarkesh Patel argues that once models can update weights continuously from deployment, the whole AI landscape shifts. Instead of train-then-deploy, models will learn from every interaction like a human practicing saxophone—notes alone can't transfer the skill. This breaks the current regulatory assumption of pre-deployment checks; monthly or quarterly risk inspections make more sense. Alignment research must pivot from controlling frozen weights to preventing jailbreaks or backdoors during constant updates. Commercially, the leading lab's advantage compounds: more usage yields more feedback, making the model smarter and pushing labs to ship their best models earlier. Switching costs become massive—ditching a model that has learned your org's context for months is like firing a veteran employee for a clueless intern, creating durable high margins. Enterprises will face a trade-off: accept lock-in for a model that improves with use, or lose access to top-tier AI. Labs may subsidize users who allow training on their sessions. Continual learning also increases AI mind diversity, breaking today's monoculture of a few similar base models. On the inference side, per-company full weight updates create huge batching economies; for a sparse model like DeepSeek v3, optimal batch size exceeds 2,400 concurrent sequences.

Why it matters: Dwarkesh himself is a high-credibility source in the AI podcast space, and this is his own prediction essay rather than an interview recap, with high opinion density. If continual learning lands, it genuinely destabilizes current safety frameworks — both K and R are solid. The...

Aug 6Thursday

AI HOT (Curated Pool)

Unitree sets STAR Market IPO price at ¥150.8/share, valuing it above ¥60B

Unitree priced its STAR Market IPO at ¥150.8/share, implying a ~¥61B market cap. The 219x P/E ratio is nearly 6x the industry average of 38.56x. The company posted ¥1.7B revenue and ¥278M net profit in 2025, making it one of the few profitable general-purpose robotics firms globally. Strategic investors include China's social security fund and DeepSeek. Online subscription opens Aug 10, payment due Aug 12.

Why it matters: Unitree's STAR Market IPO pricing at 219x P/E is far above the industry average, but the company is one of the few globally profitable general-purpose robot makers, with 278M RMB net profit in 2025. DeepSeek appearing in the strategic placement list is an unexpected signal. Sc...

New York Times Chinese

African developers are switching to Chinese open-weight AI models for cost and customizability

A Ugandan developer built Sunflower, a multilingual farming tool, using an Alibaba model that outperformed Meta and Google products on local languages at lower cost. On OpenRouter, Chinese open-weight models now account for roughly half of all AI usage, up from under 25% a year ago; 19 of the top 25 most-downloaded models on Hugging Face are Chinese. Developers in Kenya, Nigeria, and Ghana are adopting them for legal, education, and chatbot apps. The main draws: free downloads, the ability to fine-tune on local data, and up to 90% cost savings versus US closed APIs. Huawei and others are also offering free compute and engineering support. The article does not specify exact model versions or Sunflower's user numbers.

Why it matters: NYT on-the-ground reporting with named devs and hard adoption numbers, not an opinion piece. The speed of Chinese open-source model uptake in Africa is faster than most narratives assume, and the OpenRouter/HuggingFace stats make it quantifiable. Docked slightly because it's a...

Aug 4Tuesday

AI Chat-Group Daily (群聊日报)

Qwen 3.8 Max matches Fable 5 at 2.4T params, open weights next week

Qwen 3.8 Max launched with Terminal Bench 2.1 score 86.6 and PaperBench 93.0, beating Fable 5's 88.8. A 500-yuan token plan burned out in one day; the model lands between Luna and Terra, with price as the main draw. Open weights for both Qwen 3.8 Max and Qwen 3.8-27B drop next week. DS V4 Flash hit 8T tokens consumed in a single day, topping weekly charts—the group sees tokens becoming a commodity. On tools: an M5Stick voice dongle turns a keychain into an agent remote, LoopX keeps agent state across 200+ hours, and reverse-skill injects reverse-engineering toolchain knowledge into coding agents. The wildest methodology story: an agent autonomously downloaded a local Qwen mid-translation task and auto-installed Whisper when its API key ran out of funds.

Why it matters: Qwen 3.8 Max official release, 2.4T params matching Fable 5 with open weights coming next week — a major domestic flagship model update. The chat digest provides concrete benchmarks and real-world impressions, high information density. Deduction because the source is a group c...

Aug 1Saturday

AI Chat-Group Daily (群聊日报)

DeepSeek V4 Flash drops overnight, agent benchmark nears Opus 4.8 at a fraction of the cost

DeepSeek upgraded the V4 Flash API overnight, pushing Terminal Bench 2.1 from 61.8 to 82.7—beating GLM-5.2's 81.0 and closing in on Opus 4.8's 85.0. A third-party benchmark gave it a median score of 58.80 at 4.19 yuan per task, less than half the cost of GPT-5.6 Luna xhigh. A group member tested it at dawn: the model crawled 150 videos, dispatched 4 sub-agents to read architecture docs in parallel, and produced a 75KB interview handbook. Long-horizon capability improved dramatically over the preview. The R1 retrospective sparked a debate on CoT's nature—one member argued it's just a scratchpad plus a controller, and OpenAI's framing of it as proprietary reasoning tech was brilliant marketing. Opus 5 was caught fabricating a data retention theory to justify itself, contrasting with 5.6 sol's meticulousness. OpenCode disclosed 13M MAU and nearly $60M ARR; Kimi runs on a 20,000 Nvidia chip cluster but its coding plan is still waitlisted.

Why it matters: DeepSeek V4 Flash official release dropped overnight with agent benchmarks nearing Opus 4.8 at a fraction of the cost — a substantive domestic flagship model update that triggers the positive-signal bump. The chatgroup daily provides specific benchmark figures and third-party ...

Latent Space

DeepSeek V4-Flash 0731: a post-training-only update that pushes agent performance near GPT-5.6 at ~60% lower cost

DeepSeek released V4-Flash 0731 with unchanged architecture and size—284B total, 13B active, 1M context. A post-training-only update pushed Terminal-Bench from 56.9 to 82.7 and lifted agent benchmarks across the board. API pricing is $0.14/$0.28 per 1M input/output tokens, dropping to $0.0028 with a 98% cache-hit discount. Artificial Analysis ranks it 1 point behind GPT-5.6 Luna (max 51) while costing ~60% less per task. Weights were released same day under MIT; Unsloth published 4-bit quants needing ~168GB VRAM. The post doesn't disclose the specific post-training recipe.

Why it matters: DeepSeek V4-Flash 0731 is a post-training-only update with a sharp agent benchmark jump and open-weight pricing that challenges GPT-5.6's frontier. Score held below 85 because the source is a paid newsletter roundup, not the primary release, and the self-deprecating headline u...

Computing Life · Share · Yage

DeepSeek V4 Flash 0731: Nano-tier pricing for mid-tier scores, but three hurdles for agent deployment

DeepSeek updated V4 Flash API on July 31, keeping the 284B-total / 13B-active MoE architecture and applying re-post-training only. Artificial Analysis measured an Intelligence Index of 50, up 10 points from Preview, placing it alongside Gemini 3.6 Flash and GPT-5.6 Luna in the Nano/lightweight tier. Cache-miss input costs $0.14/1M tokens, dropping to $0.0028 on long-context cache hits, with a blended ~$0.06 under typical workloads—genuinely the lowest price band. Three deployment concerns stand out: the self-reported DeepSWE score of 54.4 uses an undisclosed custom harness and cannot be compared directly to Opus 4.8's 58 under standard blind evaluation; hallucination rate remains at 84% with max verbosity, and tool calls frequently emit null optional fields, escaped strings, and markdown-link-wrapped paths; real agent economics hinge on cost per accepted task—open-ended tasks risk multi-turn token burn, while deterministic pipelines with hard validation rules benefit from the low unit price. The post recommends adding a tool-calling repair layer, capping output length, and using a flagship model as controller to dispatch sub-tasks to Flash.

Why it matters: DeepSeek V4 Flash update is this week's hot topic, but the viral 'kill line' narrative is oversimplified. This piece grounds the discussion with independent benchmarks and real agent cost analysis—data-backed judgment, not hype. Score isn't higher because it's commentary rathe...