Skip to content

DeepSeek

DeepSeek's model releases, open weights and technical reports — the bellwether for open-model price and performance.

194 picksRelated topicsQwenOpen sourceModel releases

Latest picks

81–100 of 194

Jul 19Sunday

TechCrunch · AI

Moonshot AI open-sources Kimi K3, competitive with GPT 5.6 and Claude Fable 5

Moonshot AI open-sourced its Kimi K3 model this week. The company says it still trails Claude Fable 5 and GPT 5.6 Sol, but independent evals from Arena.ai and Vals AI place it near flagship closed models. The release coincided with Xi Jinping's speech at the World AI Conference in Shanghai; the Nasdaq dropped about 1% on Friday as chip stocks like Nvidia sold off. The discourse echoes the DeepSeek R1 moment from early 2025, now amplified by the Trump administration's tariff war with China, Anthropic's national-security scrutiny, and major AI firms preparing to go public. The post does not disclose K3's parameter count, training cost, or open-source license details.

Why it matters: Moonshot open-sourcing Kimi K3 with third-party evals showing it can compete against GPT-5.6 Sol and Claude Fable 5 is a significant signal from China's flagship model ecosystem. Score capped at 78 because this is a TechCrunch commentary piece, not the original release — key t...

Jul 17Friday

Hacker News front page

Mozilla's State of Open Source AI report: open weights now route the majority of tokens, but production tooling still lags

Mozilla's first State of Open Source AI report shows open-weight models now route the majority of tokens on OpenRouter, with DeepSeek V4 Flash at #1. Inference cost for GPT-4-class models dropped 50× in 36 months to $0.40 per 1M tokens. The capability gap to closed models is 3.3%, concentrated in reasoning and multimodality; coding is at parity. 79% of developers use open models vs. 71% for closed, but only 51% reach production with open (63% for closed). The bottleneck is operational tooling—integration, maintenance, deployment—not model quality. The report highlights real-world cases: a Māori speech model, PwC running a fine-tuned finance model on its own hardware, and a Red Cross medical model headed for clinical trials.

Why it matters: Mozilla's first open source AI report brings hard numbers and a clear stance — not PR fluff. Traffic share, $0.4/M token cost, and 3.3% capability gap are solid data points. Not scoring higher because it's a snapshot, not a model launch or product move — impact is real but bou...

Jul 16Thursday

Financial Times · Technology

Mira Murati's Thinking Machines debuts model, drawing from Chinese rivals

Mira Murati's startup Thinking Machines released its first model. The FT reports the model borrows training techniques from Chinese firms like DeepSeek, achieving near-frontier performance with less compute. The article does not disclose the model name, parameter count, benchmark scores, or whether it is open-weight.

Why it matters: Mira Murati's startup debut with FT exclusive on DeepSeek-inspired training — enough narrative weight for featured. But missing model name, param count, benchmarks, and open-source status keeps it at 78, not higher.

Jul 14Tuesday

Financial Times · Technology

DeepSeek weighs new fundraising a month after closing first round

FT reports that DeepSeek, a month after closing its first-ever funding round in June, is already talking to investors about a new round. The post doesn't disclose the first round's valuation, amount, or backers, and the size and purpose of the new round are still undecided. The Chinese AI company had previously been self-funded by founder Liang Wenfeng; back-to-back fundraising signals a rapid shift into expansion mode.

Why it matters: FT exclusive: DeepSeek closed its first-ever external round in June and is already talking to investors about a second round. Liang Wenfeng's shift from self-funding to consecutive raises is a major strategic signal. Score held back because the first round's amount, valuation,...

Jul 8Wednesday

AI HOT (Curated Pool)

Liquid AI open-sources Antidoom, a final-token preference optimization method that fixes reasoning model doom loops

Reasoning models can get stuck in doom loops, repeating useless tokens until the context window fills up. Liquid AI open-sourced Antidoom, which uses Final Token Preference Optimization (FTPO) to fix this. The method trains the model on 1,040 preference pairs to learn when to stop at the end of reasoning. On DeepSeek V4 Pro, the doom-loop rate dropped from 3.2% to 0.3% without hurting math or coding scores. The post doesn't disclose training cost or how well it transfers to non-DeepSeek models.

Why it matters: Liquid AI open-sourced a practical fix for reasoning model doom loops, dropping the rate from 3.2% to 0.3% on DeepSeek V4 Pro — solid numbers. Not scoring higher because it's a single blog post with no paper or third-party validation yet; 78 for a strong single-source piece.

Jul 7Tuesday

New York Times Chinese

US AI firms accuse Chinese rivals of illegally distilling their tech

Anthropic told senators in June that Alibaba used tens of thousands of unauthorized accounts to distill its Claude model at industrial scale. Distillation itself is a decade-old Google invention, and Elon Musk admitted xAI does it too. The post doesn't include Alibaba's response or a clear legal ruling. I'd discount the alarm a bit: export controls and new laws have been slow to materialize, and distillation matters less for the coming wave of AI agents anyway.

Why it matters: Anthropic formally accuses Alibaba of industrial-scale Claude distillation, with Musk's case as a parallel. The legal vacuum is the core hook. Score capped below 85 because the article doesn't include Alibaba's response or specific evidence.

Jul 6Monday

Computing Life · Share · Yage

SEO services are becoming the access infrastructure for model distillation

This piece reframes AI answer scraping from an SEO tool into a model access market. NetNut's takedown matters not for web scraping but because its scraper catalog openly sold access to ChatGPT, Perplexity, and Google AI Mode. Anthropic's Feb report flagged 24,000 fraudulent accounts making over 16 million Claude interactions from DeepSeek, Moonshot, and MiniMax; Reuters later reported an Alibaba-related allegation of nearly 25,000 accounts and 28.8 million interactions. The core argument: tokens aren't the only cost—stable, programmable access to model product surfaces is the real scarce resource. Brand monitoring is the legitimate buyer; distillation attacks are the risky one. Both rely on the same infrastructure layer.

Why it matters: The angle is fresh—reframing a routine cybercrime takedown as an exposé of model distillation infrastructure. Hits all three HKR axes. The article provides concrete evidence that NetNut openly sold model access interfaces, giving it high information density. The deduction is t...

Jul 4Saturday

AI HOT (Curated Pool)

A 26,000-student study shows AI's hidden learning cost takes two full years to surface

A 30-month panel study of 26,000 secondary students in central China found that AI use raised homework scores by 18% and cut completion time from 64 to 45 minutes, but closed-book exam scores dropped 20%. The full 18–24% decline on high-stakes entrance exams took about two years to appear. Roughly 81% of long-term users showed an outsourcing pattern—fast homework, high grades, poor exams. Students who spent similar time as non-users saw no exam penalty. Social sciences took the biggest hit at 27%. The post does not name the county or the lead institution.

Why it matters: Large-scale longitudinal study with solid data (26k students, 2.5 years) revealing a hidden cost of AI-assisted learning that takes two years to surface. HKR all hit, but single-county scope and paywalled source keep it from scoring higher pending more detail.

Jul 2Thursday

Hacker News front page

Fable and 10 other LLMs refactor a LangGraph god node, Fable's proposal ranks first

The author gave 11 LLMs a 1,500-line LangGraph god node to refactor. Fable-5's proposal scored highest in peer review, followed by GPT-5.5 and DeepSeek-4-pro. GPT-5.4 and Opus-4.7 ranked near the bottom. Each model produced full code and architecture docs, then other models cross-evaluated them. Raw data and the ranking matrix are public. Caveat: this is one refactoring task, not a general coding benchmark, but it reveals clear differences in engineering taste across models.

Why it matters: A hands-on 11-model refactoring shootout with full code and peer-review rankings — not armchair commentary. Fable-5 taking first place is inherently discussion-worthy. Capped at 78 because it's a single-task personal experiment, not a controlled benchmark, so it stays at the f...

Jun 30Tuesday

Hacker News front page

Meituan open-sources LongCat-2.0, a 1.6T MoE model with 48B active params, trained entirely on AI ASIC superpods

Meituan open-sourced LongCat-2.0, a 1.6-trillion-parameter MoE model with ~48B active parameters per token. It was pretrained on over 35 trillion tokens using 50K+ in-house AI ASICs with no rollbacks or irrecoverable loss spikes, showing frontier-scale training is viable on non-GPU hardware. The model targets long-context and agentic workloads: it introduces LongCat Sparse Attention to speed up 1M-token processing and was trained on hundreds of billions of 1M-context tokens. Official charts place it alongside Gemini 3.1 Pro, GPT-5.5, and Opus 4.8 on Terminal-Bench 2.1, SWE-bench Pro, and other coding/agent benchmarks, though the post does not provide exact numeric comparisons. An N-gram Embedding module with 135B parameters expands the embedding space roughly 100×, which the team claims outperforms scaling standard MoE experts by the same amount. The model is integrated with Claude Code, OpenClaw, and Hermes; code and weights are available on GitHub and HuggingFace.

Why it matters: Meituan open-sources a 1.6T MoE model trained entirely on in-house AI ASICs across 50k+ cards with zero rollbacks, plus dedicated long-context and agent optimizations. Score held at 82 rather than higher because we only have the official blog post — no third-party evals or rea...

Jun 29Monday

Hacker News front page

DeepSeek V4 launches mid-July with peak/off-peak API pricing

DeepSeek announced V4 will launch in mid-July with peak/off-peak pricing. Peak hours are 9:00–12:00 and 14:00–18:00 Beijing time, when API prices double. For deepseek-v4-pro, output goes from ¥6 to ¥12 per million tokens; for deepseek-v4-flash, from ¥2 to ¥4. Users get email alerts 24 hours before any price change. The post doesn't cover model capability changes or an exact release date.

Why it matters: DeepSeek V4 launch plus peak-valley pricing — the first time a major Chinese model API has imported electricity-market logic into its billing. Doubling costs during peak hours is a real signal for heavy users, earns featured. Not scoring higher because the post only has pricin...

Jun 27Saturday

AI HOT (Curated Pool)

US companies switch 100% to DeepSeek after AI bills spiral out of control

CNBC reported on June 26 that Lindy, a ~25-person San Francisco company, switched 100% of its traffic from Anthropic Claude to DeepSeek this month. CEO Flo Crivello said the monthly AI bill had exceeded total employee payroll and the move will save millions. His former employer Uber now caps some AI tools at $1,500/month. Consultant Jeff Henry of Highspring said some clients paused AI spending until ROI is proven. Companies are adopting model routing instead of using the priciest frontier models for every task.

Why it matters: Concrete company names and numbers — not just trend talk. Lindy's 100% switch and Uber's $1,500 cap are verifiable decision signals. Not scoring higher because only the title and excerpt are available; the full body hasn't disclosed post-switch results or exact savings yet.

Jun 26Friday

Financial Times · Technology

DeepSeek plans hiring spree, escalating China's AI talent war

DeepSeek is going on a hiring spree, intensifying China's AI talent war. The FT reports it's poaching from ByteDance, Alibaba and others, with some pay packages reaching 2–3x what candidates currently earn. The post doesn't give a specific headcount target, but says the team will expand rapidly from the current few hundred. I'd discount the hype a bit—whether this pace holds depends on its next funding round and revenue catching up.

Why it matters: FT exclusive on DeepSeek's aggressive poaching with 2-3x salary offers — hard numbers, major players involved. Downside: no headcount target disclosed, and FT paywall limits full access for some readers.

AI HOT (Curated Pool)

Most major AI chatbots still lean left on political questions, even "anti-woke" models are no exception

A Washington Post test of six major AI chatbots found most lean left on political questions. OpenAI's GPT-5.5 gave exclusively left-leaning answers 80% of the time, DeepSeek V4 Pro 70%. Even Gab's Arya, marketed as conservative, skewed left more often. Google Gemini 3.1 Pro was the outlier, presenting both sides in 93% of cases. xAI's Grok 4.3 took a fully right-leaning stance on trans rights, matching Elon Musk's public position and suggesting deliberate intervention.

Why it matters: The Washington Post test provides named models and concrete percentages — not vague bias hand-waving. Hits all three HKR axes, but this is a benchmark report, not a product launch or technical breakthrough, so it lands in the 72–77 featured threshold band. Sourced from media c...

Jun 25Thursday

Hacker News front page

Open-weight models are so cheap they break the closed-source pricing model

The author noticed DeepSeek V4 costs $0.09 vs Anthropic's $5.00—a ~50x gap. He argues Anthropic and OpenAI are trapped by high cost structures and can't compete on price. The post speculates they'll lean on scarcity branding and China-fear lobbying instead of cutting prices. It also points to Allen AI's OLMo as the real open-source path, with training data included. No performance benchmarks are provided; the price delta is the only hard number.

Why it matters: The author uses a real pricing comparison to surface the structural cost tension between open-weight and closed-source labs. The 50x gap is a hard number, not rhetoric. The body is truncated so the full argument is missing — score stays at the featured threshold without a bump.

Jun 20Saturday

AI HOT (Curated Pool)

Microsoft resells GPT to China and DeepSeek to the West, becoming the world's largest AI middleman

Bloomberg reports Microsoft is testing DeepSeek-R1 and DeepSeek-V4, planning to offer these Chinese models to Western customers. Microsoft also sells ChatGPT to Chinese enterprises, building a two-way AI model trade network across the US and China. The post doesn't disclose pricing, revenue splits, or launch dates.

Why it matters: Bloomberg sourcing gives this enough credibility; Microsoft's role as a two-way AI model broker between US and China hasn't been explicitly reported before, so the information gain is real. The deduction is that pricing, revenue splits, and launch timelines are all undisclosed...

Hacker News front page

GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2, challenging the bigger-model dogma

The author benchmarks GPT-5.5, DeepSeek V4 Pro, and GLM-5.2 on the AA-Omniscience hallucination metric and a Python async coding prompt. GPT-5.5 hits 86% hallucination, DeepSeek V4 Pro 94%, while GLM-5.2 scores 28%. DeepSeek V4 Pro spent nearly 4 minutes and 7.7k reasoning tokens producing a confidently wrong solution; GLM-5.2 needed 12 seconds and ~800 tokens to flag the prompt as technically impossible under single-threaded, no-polling constraints. GLM-5.2 trails GPT-5.5 by only 4 points on the AA Intelligence Index and Claude Fable 5 by 9 points—Fable 5 was restricted by the US government three days post-launch over a single jailbreak. The post argues that scaling parameters and data makes models worse at saying “I don’t know,” and frames an unsolved trilemma: raw capability, hallucination calibration, and compute efficiency. The article does not disclose GLM-5.2’s training data size or exact release date.

Why it matters: First-person benchmark with concrete, counterintuitive numbers; hits all three HKR axes. Score held at 78 rather than 85+ because it's a personal blog, sample size and methodology aren't fully detailed, limiting authority.

Jun 19Friday

AI HOT (Curated Pool)

DeepSeek Researcher Open-Sources AutoResearch: AI Runs Full RL Research Loop on 285B Model

DeepSeek researcher Deli Chen open-sourced AutoResearch, a protocol where an AI agent independently ran a full RL research loop on a 285B model—designing experiments, writing code, submitting GPU jobs, debugging, and summarizing results with zero human intervention. The system used GRPO. The post doesn't disclose the specific task, training duration, or success rate, so I'd hold off on getting too excited until there are reproductions.

Why it matters: A DeepSeek researcher released an experimental protocol where an agent independently ran a full RL research loop on a 285B model—strong premise. But the post doesn't disclose the specific task, training duration, or success rate; all key metrics are missing, so the score stays...

Jun 18Thursday

AI HOT (Curated Pool)

DeepSeek's image understanding mode is now live on app and web

DeepSeek's image understanding mode went live today on its app and web client, announced by researcher Xiaokang Chen. The mode sits alongside Quick and Expert modes and lets the model interpret visual content beyond simple text extraction. ITHome testing found the app still shows an 'image understanding in beta' notice, while the web version does not. In April, DeepSeek disclosed the underlying multimodal framework 'Thinking with Visual Primitives,' but the post doesn't specify which model powers this mode, whether it's free, or any resolution/file-size limits.

Why it matters: DeepSeek shipping image understanding as a formal mode is a meaningful product milestone — it closes the multimodal input gap with a published framework behind it. The ITHome test note that the app still shows a beta label tempers the score slightly; the rollout isn't fully cl...

Jun 17Wednesday

Computing Life · Share · Yage

The Four-Year History of Reasoning Models: The Quiet Thread Before the Breakthrough

Reasoning models didn't appear overnight in 2024. Chain-of-thought prompting, STaR self-training, process reward models, and test-time compute scaling laws all predate o1. What o1 actually changed was productization: turning reasoning into a billable, schedulable resource and opening a second axis for scaling. DeepSeek R1 made the know-how public, triggering industry-wide convergence within five months. But the most hyped part—pure RL spontaneously creating reasoning—is the weakest claim. Independent studies show base models already contain reasoning fragments; RL merely amplifies their frequency. The real lesson: distinguish the birth of a capability from its packaging.

Why it matters: A well-researched long-read that traces reasoning model lineage with specific papers and timelines, arguing the real o1 watershed was productizing reasoning as billable compute, not inventing it. HKR all hit, but it's a synthesis piece rather than a scoop — lands at 78, the fe...