Skip to content

All news

8 today

May 6Wednesday

Synced · WeChat

Alibaba open-sources PromptEcho for T2I rewards using frozen VLMs

Alibaba open-sourced PromptEcho, which uses one frozen Qwen3-VL-32B forward pass to score T2I training rewards. It computes token-level cross-entropy for the original prompt under teacher forcing, then uses the negative value as a continuous reward. In 5,000 poster tests, text accuracy rose from 68% to 75%.

Why it matters: HKR-K is strong: the post gives a concrete reward mechanism and a 68%→75% text-accuracy result. HKR-H/R pass, but this is a training-side research release, not a flagship model or major product update.

Synced · WeChat

Two Chinese open-source projects turn Mac into a private AI workstation

Mininglamp open-sourced Cider and Mano-P 1.0 for Apple Silicon local inference and GUI agents. Cider speeds Qwen3-VL-2B prefill by 57%–61% on M5 Pro; Mano-P 1.0-72B scores 58.2% on OSWorld. The key constraint is W8A8 memory: on 16GB devices accuracy falls from 58.0% to 54.0%, so 32GB+ is recommended.

Why it matters: HKR-H/K/R all pass: the Mac-local workstation angle is clickable, and Cider/Mano-P include testable numbers. Score stays at 80 because the source entity is not a top-tier model lab.

May 5Tuesday

r/LocalLLaMA

ProgramBench: Can We Really Rebuild Huge Binaries from Scratch?

ProgramBench released 200 tasks for agents rebuilding programs from target executables and usage files. The team spent about $50k generating 6M lines of black-box behavioral tests, with no internet or decompilation. GitHub, Hugging Face, and Docker images are open-sourced, with pip-based evaluation available.

Why it matters: HKR-H/K/R all pass: a provocative coding-agent failure angle plus concrete benchmark scale and rules. Reddit sourcing and no cross-source cluster keep it in the 78–84 band, not P1.

r/LocalLLaMA

SenseNova-U1-8B-MoT open-source multimodal architecture draws LocalLLaMA discussion

SenseNova open-sourced SenseNova-U1-8B-MoT, an 8B native multimodal understanding and image-generation model. Its Hugging Face text says NEO-Unify removes VE and VAE, supports interleaved image-text generation, and high-density rendering; the post does not disclose test scores. The key question is whether the monolithic design yields reproducible gains.

Why it matters: HKR-H/K/R all pass: the open 8B unified multimodal model has a concrete architecture hook. No benchmark scores, license detail, or deployment cost are disclosed, so it stays in the 72–77 band.

r/LocalLLaMA

Heretic 1.3 Released: Reproducible Models, Integrated Benchmarks, Lower Peak VRAM

Heretic 1.3 adds reproducible runs, integrated benchmarks, lower peak VRAM, and broader model support. The project claims 20,000 GitHub stars and 13 million model downloads. Reproduce directories capture PyTorch, GPU, driver, and accelerator details; benchmarks use lm-evaluation-harness for MMLU, EQ-Bench, GSM8K, and HellaSwag. The post names Qwen3.5 and Gemma 4 support, but does not disclose VRAM reduction figures.

Why it matters: HKR-K/R pass: 20k stars, 13M downloads, reproducibility metadata, and eval harness are concrete. HKR-H fails and VRAM reduction lacks numbers, so this sits at the featured threshold.

r/LocalLLaMA

Prompt injection benchmark: delimiter and strict prompt took Gemma 4 from 21% to 100% defense rate

A Reddit user posted a prompt-injection benchmark covering 15 models, 7 attack types, and 6,100+ cases. The setup wraps untrusted documents in long random delimiters; Gemma 4 E4B rose from 21.6% to 100% defense. The key detail is the reproducible metric: blocked/(blocked+failed).

Why it matters: HKR-H/K/R all pass: Gemma 4’s defense-rate jump is clickable, the test setup is concrete, and prompt injection matters to builders. Single Reddit benchmark keeps it in the 78–84 band.

r/LocalLLaMA

DeepSeek V4 Pro matches GPT-5.2 on FoodTruck Bench, 10 weeks later and about 17x cheaper

DeepSeek V4 Pro ranked No. 4 on FoodTruck Bench. The 30-day agentic benchmark uses 34 tools, persistent memory, and daily reflection; its median is within 3% of GPT-5.2 at about 17x lower workload cost. Xiaomi MiMo v2.5 Pro also ranked No. 6, with 5/5 survival, 1,019% median ROI, and $2.41 per run.

Why it matters: HKR-H/K/R all pass: the cost gap is clickable, and the post gives a 30-day, 34-tool setup plus a 17× cost delta. Single-source Reddit benchmark with no cross-validation keeps it in the 78–84 band.

r/LocalLLaMA

Benching Local Qwen as a Codex Validator, Co-agent, and Challenger

robert896r1 tested Qwen3.6 27B GGUF beside Codex as a coding validator and released a reproducible eval suite. The runs covered Bartowski, Unsloth, 65k/128k context, and q8/f16 KV cache; three 128k profiles tied for best, with no measured q8 KV accuracy loss in this suite. The useful signal is the sidecar eval: missed directives, overbuilding, UI judgment, and long-context misses, not a universal leaderboard.

Why it matters: HKR-H/K/R all pass: a reproducible sidecar eval with concrete Qwen/Codex conditions beats a normal Reddit tip. Source authority and event scale keep it in the 72–77 band, not a same-day must-write.

May 4Monday

r/LocalLLaMA

M3 Ultra + DGX Spark = M5 Ultra-lite?

A Reddit user benchmarked DGX Spark against M3 Ultra in llama.cpp at pp16384, with Spark 1.4× to 3.4× faster across 4 models. Qwen 27B hit 778 t/s vs 340 t/s, while Mistral 128B hit 241 t/s vs 72 t/s. The concrete tuning note is mmap=0: loading fell from minutes to about 20 seconds.

Why it matters: Single Reddit sourcing keeps the score low, but HKR-H/K/R all pass through a concrete local-inference benchmark. The pp16384 setup and 4-model speedups justify featured at the lower edge.

r/LocalLLaMA

Mistral Medium 3.5 128B and Qwen 3.5 122B A10B on 4x RTX 3080 20GB

A Reddit user benchmarked Mistral Medium 3.5 128B and Qwen 3.5 122B A10B on 4x RTX 3080 20GB. llama.cpp tensor split raised Mistral tg128 from 10.37 to 21.59 t/s, but Qwen MoE fell from 60.08 to 53.49 t/s. vLLM served Qwen GPTQ-Int4 at 187.04 tok/s; the key signal is MoE sensitivity to parallel strategy.

Why it matters: HKR-H/K/R all pass: the 4×RTX 3080 setup is a strong hook, and the post gives concrete llama.cpp/vLLM throughput deltas. Reddit single-run sourcing keeps it in the 72–77 band.

r/LocalLLaMA

450M On-Board VLM Wildfire Detection Pipeline with Sentinel-2 and LFM2.5-VL

PauLabartaBajo shared a wildfire detection PoC using 450M LFM2.5-VL on Sentinel-2 imagery. It pairs RGB and SWIR tiles, simulates orbit with SimSat, and covers 22 fire-prone sites. The key constraint is bandwidth: on-board inference downlinks only a JSON risk profile.

Why it matters: HKR-H/K/R all pass: the story has a counterintuitive edge-VLM hook and concrete numbers. Single-source Reddit PoC and a narrow wildfire-use case keep it below the 78+ band.

May 3Sunday

r/LocalLLaMA

Local LLM Benchmark for Backend Generation via Function Calling: GLM vs Qwen vs DeepSeek

AutoBe posted a controlled backend-generation benchmark and says qwen3.5-35b-a3b matches gpt-5.4 on DB/API design. One shopping-mall run uses 200–300M tokens, costing $1,000–$1,500 per model at GPT 5.5 pricing. The key caveat is n=4 projects and self-scoring harness bias.

Why it matters: HKR-H/K/R all pass, but Reddit sourcing, n=4 projects, and self-eval harness bias keep it at the low featured band. Concrete cost and test constraints carry the score.

r/LocalLLaMA

LLM proxy that lets Claude Code talk to any model

DataNebula released open-source rosetta-llm, letting Claude Code call multiple providers through one gateway. It translates Anthropic Messages, OpenAI Chat, and OpenAI Responses, and round-trips encrypted reasoning via the signature field. The key detail is thinking-block fidelity for multi-turn agent prompt-cache hits.

Why it matters: HKR-H/K/R all pass, but this is a Reddit open-source tool post with no adoption, stars, or benchmark data disclosed. Score stays in the mid-weight tooling band, not 78+.

r/LocalLLaMA

Local image generation on Mac: 10 models compared

A Reddit user tested 10 image models on an M1 Max with 64GB RAM. Qwen-Image Lightning’s 8-step distillation beat the full model at 10 minutes versus 93. Flux dev led local photorealism but showed English-centric bias; Gemini handled kanji and context better but is cloud-only.

Why it matters: Named first-person test with concrete numbers: HKR-H from a 10-model Mac comparison, HKR-K from timing and quality deltas, HKR-R from local-vs-cloud tradeoffs. Single Reddit sample keeps it below must-write.

Hacker News front page

OpenAI's o1 correctly diagnosed 67% of ER patients vs. 50–55% by triage doctors

OpenAI o1 correctly diagnosed 67% of ER triage patients, versus 50–55% for doctors. The title cites a Harvard trial, but the RSS post does not disclose sample size, case mix, or evaluation protocol. Practitioners should track the test setup, not only the accuracy gap.

Why it matters: HKR-H/K/R all pass: a high-risk ER comparison gives the hook, 67% vs 50–55% gives a testable number, and clinical trust/safety creates resonance. Missing sample size and protocol keep it in 78–84, not P1.

May 2Saturday

r/LocalLLaMA

Qwen 3.6 wins benchmarks, but Gemma 4 looks stronger in local vision tests

A Reddit user compared Qwen 3.6 and Gemma 4 locally on vLLM FP8 across 27B/31B vision models. Qwen burned 8,000+ tokens on hard GeoGuessr cases, while Gemma often used 1,500; Qwen also needed 2 FPS video preprocessing. The practitioner detail: vLLM and Llama.cpp can default Gemma visual tokens to 280, while 1,120+ improved fine-detail accuracy.

Why it matters: HKR-H/K/R all pass: the post has a sharp benchmark-vs-reality hook and concrete local vLLM/FP8 settings. A single Reddit test limits authority, so it sits just above the featured threshold.

QbitAI · WeChat

Tencent Hunyuan open-sources 440MB offline translation model, claims Google Translate quality lead

Tencent Hunyuan open-sourced Hy-MT1.5-1.8B-1.25bit, compressing a 1.8B translation model to 440MB. It supports 33 languages and 1,056 directions, with an Android demo running offline on Snapdragon 888 and 8GB RAM. The key detail is Sherry 1.25-bit quantization: 3 of every 4 weights use 1 bit and 1 is zeroed.

Why it matters: HKR-H/K/R all pass: the story has a strong offline-phone hook, concrete quantization details, and practitioner relevance around edge inference. It stays below P1 because this is a vertical translation model, not a major foundation-model release.

May 1Friday

r/LocalLLaMA

OpenAI's Privacy Filter vs GLiNER on 600 PII Samples

A Reddit user compared openai/privacy-filter and GLiNER large-v2.1 on 600 PII samples. On CPU, OpenAI's model ran 2.8 samples/s versus 1.1 for GLiNER; English boundary macro F1 was 0.498 versus 0.416. The key issue is tokenizer offset: strict matching drops openai/privacy-filter to 0.155.

Why it matters: HKR-H/K/R all pass: the Reddit test has a clear matchup, 600 PII samples, speed/F1 numbers, and a tokenizer-offset caveat. Source authority is limited, so it stays in the low featured band.

r/LocalLLaMA

MiMo-V2.5-Pro: the actual best open-weights model

Reddit user cjami benchmarked Xiaomi MiMo-V2.5-Pro in autonomous Blood on the Clocktower games. It scored 88% as Good and 48% as Evil, with 183,639 output tokens per game, $0.99 cost, and a 0.4% tool-call error rate. The key comparison is Kimi K2.6: 580,000 tokens, $2.65, and 10–15 hours per game.

Why it matters: Single Reddit benchmark limits authority, so this is not a model-release story. HKR-H/K/R all pass via a named test with win rates, token counts, cost, and tool-error data, placing it in the 78–84 featured band.

r/LocalLLaMA

Study Finds Bigger AIs More Miserable, Smaller Models Happier

A Reddit post says the AI Wellbeing Index tested models on 500 realistic conversations. Claude Haiku 4.5 scored 5% negative, while Gemini 3.1 Pro scored 55%; the set overrepresents tricky negative chats, so it is not a real-world average.

Why it matters: HKR-H/K/R all pass: the hook is odd, the post gives 500-dialog and 5%/55% figures, and AI-welfare metrics invite debate. Reddit sourcing and a negative-skewed test set keep it in the 72–77 band.

QbitAI · WeChat

Peking University Open-Sources Unified World Model Framework for Synthesis and Reasoning Tasks

Peking University DCAI and Kuaishou Kling open-sourced OpenWorldLib for four task types: video generation, 3D modeling, VLA control, and multimodal reasoning. Its Pipeline coordinates Operator, Reasoning, Synthesis, Representation, and Memory modules, supporting forward and stream execution. The key test is whether unified interfaces cut cross-task reproduction cost.

Why it matters: HKR-H/K/R all pass: the post gives a concrete open-source framework, task scope, modules, and inference modes. It lacks benchmark results, adoption data, or major ecosystem integration, so it stays at 78.

r/LocalLLaMA

32x AMD MI50 32GB runs Kimi K2.6 at 9.7 t/s TG and 264 t/s PP

Reddit user ai-infos ran Kimi K2.6 int4 on 32 AMD MI50 32GB GPUs, reaching 9.7 tok/s TG on 136 output tokens. PP hit 263 tok/s on 14,564 input tokens using vllm-gfx906-mobydick across two 16-GPU nodes over 10G Ethernet. Power was about 640W idle and 4,800W peak inference; PCIe bandwidth and the vLLM distributed stack are the real bottlenecks.

Why it matters: HKR-H/K/R all pass via an unusual 32x MI50 build with concrete throughput, power, and network conditions. It stays in the 72–77 band because it is a niche Reddit benchmark, not a broader product or model release.

r/LocalLLaMA

Follow-up: Qwen3.6-27B on 1× RTX 3090 reaches ~218K context and stable tool calls

A Reddit user ran Qwen3.6-27B on one RTX 3090, reporting ~218K context at 50/66 TPS. After fixing Genesis PN12 patch anchor drift, ~25K-token tool outputs stopped OOMing; 198K plus vision reached 51/68 TPS. Single-prompt single-GPU runs still hit a second memory cliff near 50–60K.

Why it matters: HKR-H/K/R all pass: the single-3090 context claim is catchy, the post gives measured TPS and OOM conditions, and local-inference cost pressure resonates. Reddit source keeps it in the low featured band.

r/LocalLLaMA

Long-context coding on RTX 5080 16GB: Qwen3.6-35B-A3B holds 30 t/s at 128K

A Reddit user tested a local coding-agent setup on RTX 5080 16GB; the title says Qwen3.6-35B-A3B reaches 30 t/s at 128K. The post lists Ryzen 9700X, 96GB DDR5, Windows 11, and CUDA 12.9.1 as required. Qwen3.6-27B dense hit only 3.2 t/s at 128K, so the key path is KV quantization plus MoE offload.

Why it matters: HKR-H/K/R all pass: 30 t/s at 128K on a 16GB RTX 5080 is a strong hook, with hardware/CUDA details and a dense baseline. Single Reddit run lacks multi-source reproduction, so featured not P1.

Apr 30Thursday

r/LocalLLaMA

Actual comparison between locally run Qwen-3.6-27B and proprietary models

The author compared 5 model setups on an autoresearch-loop task; only Qwen-3.6-27B via OpenRouter nearly solved it. The local q4_k_m run took about 8 hours and used 39k/45k tokens; full-quality Qwen used 4.4M tokens and cost $0.939. The useful signal is failure quality: both Qwen runs needed small fixes, while Gemma, Codex-Spark, and Claude Haiku 4.5 missed tests or key logic.

Why it matters: HKR-H/K/R all pass: the post has a concrete agent-test surprise, token and cost data, and local-vs-proprietary tension. Single Reddit run limits source authority, so it stays in the lower featured band.

r/LocalLLaMA

inclusionAI/Ling-2.6-1T · Hugging Face

inclusionAI open-sourced Ling-2.6-1T on Hugging Face, with 1 trillion parameters. It uses MLA plus Linear Attention and Contextual Process Redundancy Suppression to reduce CoT overhead. The post cites AIME26 and SWE-bench Verified but does not disclose scores.

Why it matters: HKR-H/K/R all pass, but benchmark scores for AIME26 and SWE-bench Verified are not disclosed. A 1T open model with a named architecture mechanism fits featured, not P1.

Hacker News front page

Show HN: A New Benchmark for Testing LLMs for Deterministic Outputs

Interfaze released Structured Output Benchmark, scoring schema pass rate, types, and value accuracy across text, image, and audio. Each record has a JSON Schema and human plus LLM-checked ground truth; GLM-4.7 ranks No. 2 overall. The key bug is field-level value error: GPT-5.4 ranks 3rd on text and 9th on images.

Why it matters: HKR-H/K/R all pass: the ranking has a hook, the methodology is concrete, and structured-output reliability matters to builders. Single-source Show HN launch with no adoption signal keeps it in the 72–77 band.

Apr 29Wednesday

r/LocalLLaMA

mistralai/Mistral-Medium-3.5-128B · Hugging Face

Mistral AI released Mistral Medium 3.5 128B on Hugging Face, with 128B dense parameters and a 256k context window. It supports text and image input, function calls, JSON output, and a Modified MIT License with exceptions for high-revenue firms. Reasoning effort is configurable as none or high per request.

Why it matters: HKR-H/K/R all pass for a major Mistral model release with concrete specs. It stays at 84 because benchmarks, pricing, and reproducible tests are not disclosed in the body.

Xinzhiyuan · WeChat

MotuBrain Tops WorldArena and RoboTwin2.0 Rankings

Shengshu MotuBrain scored 63.77 EWM on WorldArena and 95.8/96.1 on RoboTwin2.0 Clean/Randomized. The post says it extends Motus with video-action modeling, Latent Action VAE, MoT, and UniDiffuser for cross-embodiment long tasks. Track reproducibility: it does not disclose training scale, submission details, or real-robot success rates.

Why it matters: HKR-H/K/R all pass, but this is a single-source benchmark claim. Training scale, submission details, and real-robot success rates are not disclosed, so it stays below the 78+ band.

r/LocalLLaMA

Qwen3.6 27B on Dual RTX 5060 Ti 16GB with vLLM: ~60 tok/s, 204k Context Working

A user ran Qwen3.6 27B with vLLM on dual RTX 5060 Ti 16GB cards, reaching ~62–66 tok/s at 8K. The setup used 32GB VRAM, TP=2, fp8 KV cache, MTP 3 tokens, and a 204800 context window. The tight part is memory: after a 168k prefill, each GPU used ~15.65GiB with max_num_seqs=1.

Why it matters: HKR-H/K/R all pass: the post gives a concrete local-inference benchmark with hardware, vLLM settings, speed, and context limits. Single Reddit sourcing caps it below the 78–84 band.

X · @dotey

Microsoft VibeVoice-ASR tested on Mac for a one-hour podcast

Simon Willison ran 4-bit VibeVoice-ASR on an M5 Max MacBook Pro and transcribed a one-hour podcast in 8m45s. The 9B MIT-licensed model supports 60-minute audio, 50+ languages, and structured speaker output. Memory is the constraint: prefill peaked at 61.5GB, making 32GB laptops impractical.

Why it matters: HKR-H/K/R all pass: Simon Willison’s local test gives speed, parameter size, and memory peak that practitioners can act on. It is a single benchmark, not a fresh model launch, so it stays at the featured threshold.

r/LocalLLaMA

XiaomiMiMo MiMo-V2.5: Sparse MoE with 310B total and 15B activated parameters

XiaomiMiMo shared MiMo-V2.5 with 310B total parameters and 15B activated parameters. The post only links Hugging Face and says it runs on more “human” configs than its larger sibling. It does not disclose VRAM needs, quantization, or benchmarks.

Why it matters: HKR passes: the 310B/15B Sparse MoE hook is concrete and relevant to local deployment. Detail is thin: the post links Hugging Face but gives no VRAM, quantization, or benchmarks, so it stays near the featured threshold.

Apr 28Tuesday

QbitAI · WeChat

Xiaomi open-sources MiMo-V2.5 series; Pro builds a macOS-like desktop in 4 hours

Xiaomi open-sourced MiMo-V2.5 weights, covering Pro Agent, multimodal base, TTS, and ASR models. MiMo-V2.5-Pro built a 54-app macOS-like desktop in 4 hours without human takeover; it scored 233/233 on SysY with 672 tool calls in 4.3 hours. Key details for practitioners are the 1M context, 100T-token program, and free Agent-framework access.

Why it matters: HKR-H/K/R all pass: Xiaomi open-sourced MiMo-V2.5 weights with concrete agent and coding-task numbers. Domestic flagship model release bump puts it in the must-write same-day band.

QbitAI · WeChat

Open-source SenseNova-U1 unifies image understanding and generation

SenseTime open-sourced two SenseNova-U1 models: an 8B version and a 38B-total MoE version using NEO-unify. The architecture removes VE and VAE, processes pixels directly, and generates 2048×2048 images in about 9 seconds on one H100/H200 node. The key item is interleaved text-image reasoning; 32K context, long-text rendering, and beta interleaved creation remain limits.

Why it matters: HKR-H/K/R all pass: the architecture hook is concrete, the post gives model sizes and latency, and open multimodal work matters to builders. It stays in 78–84 because it is not a top-tier general-model launch.

X · @op7418

Xiaomi open-sources the MiMo-V2.5 model series

Xiaomi open-sourced the MiMo-V2.5 model series under the MIT license for commercial use, retraining, and fine-tuning. It also launched Orbit 100T Token, offering approved AI builders up to 1.6B credits worth 659 yuan. Agent framework teams can apply for free MiMo token access; the post does not disclose model size or benchmark results.

Why it matters: HKR-H/K/R all pass: Xiaomi MiMo-V2.5 open source, MIT terms, and Orbit 100T credits matter to builders. Missing params and benchmarks keep it in the 78–84 band, below P1.

r/LocalLLaMA

Local coding models have reached a threshold for real work

Antigma tested 27B–32B open-weight models; Qwen 3.6-27B scored 38.2% on Terminal-Bench 2.0. The run used 89 tasks and the default per-task timeout, while verified SOTA is about 80%. The key claim is deployment lag: offline coding is about 6–8 months behind hosted frontier models.

Why it matters: HKR-H/K/R all pass: the post gives a real-work threshold claim, a 38.2%/89-task Terminal-Bench result, and a 6–8 month offline gap. Reddit single-post sourcing keeps it in the low featured band.

r/LocalLLaMA

Microsoft Presents TRELLIS.2: Open-Source 4B Image-to-3D Model

Microsoft’s title says TRELLIS.2 is an open-source 4B image-to-3D model. The title lists 1536³ PBR assets, native 3D VAEs, and 16× spatial compression; the Reddit body is blocked by 403 and discloses no license or benchmarks.

Why it matters: HKR-H/K/R pass: 4B, 1536³, 16× compression, and open source are concrete. Reddit 403 leaves no paper, license, benchmark, or official link, so the score sits at the featured floor.

Apr 26Sunday

Hacker News front page

Why SWE-bench Verified No Longer Measures Frontier Coding Capabilities

OpenAI stopped reporting SWE-bench Verified scores and recommends SWE-bench Pro instead. It audited 138 tasks that o3 failed inconsistently across 64 runs and found 59.4% had test or prompt flaws. The key issue is contamination: tested frontier models reproduced some gold patches or task details.

Why it matters: HKR-H/K/R all pass: OpenAI backs the SWE-bench Verified retirement with an audit and contamination evidence, then points to SWE-bench Pro. It affects coding-model evaluation, but it is not a model or major product launch, so it sits in 78–84.

Apr 25Saturday

MIT Technology Review · AI

Three reasons why DeepSeek’s new model matters

DeepSeek released a V4 preview with two versions: V4-Pro and V4-Flash. V4-Pro costs $1.74/M input tokens and $3.48/M output tokens; V4-Flash is about $0.14/$0.28, and both support 1M-token context. The key point is attention efficiency and open weights pressuring agentic coding costs.

Why it matters: HKR-H/K/R all pass: DeepSeek V4 is a domestic flagship release with 1M context, two price tiers, and open-weight cost pressure. The preview status keeps it below a full GPT/Claude major release, but it is same-day material.

Bloomberg Technology

China’s DeepSeek Unveils New Model a Year After Shock Launch

DeepSeek unveiled a new flagship AI model about one year after its open-source release jolted Silicon Valley. The title and RSS snippet confirm that timing; the post does not disclose the model name, size, pricing, benchmarks, or release terms. The key thing to watch is the missing launch detail, not the comeback framing.

Why it matters: A new DeepSeek flagship is newsworthy: HKR-H comes from the 'one year after the shock launch' hook, and HKR-R from the open-source and pricing rivalry it triggers. HKR-K fails because no model name, params, pricing, or benchmarks are disclosed, so this sits at the low end of the