Skip to content

All news

8 today

Apr 24Friday

Bloomberg Technology

DeepSeek unveils flagship AI model a year after breakthrough

DeepSeek released preview versions of a new flagship AI model one year after its breakout. The RSS snippet calls it its most powerful open-source platform and frames it against OpenAI and Anthropic; the post does not disclose parameters, context length, benchmarks, or rollout timing. The actionable facts so far are limited to its preview status and open-source positioning.

Why it matters: A new DeepSeek flagship preview deserves real weight under the domestic-flagship rule, and Bloomberg adds source authority. HKR-H and HKR-R pass, but HKR-K fails because the story discloses no specs, context window, benchmarks, or release schedule, so this stays at the low end of

X · @dotey

DeepSeek releases and open-sources V4 preview; 1M context is standard across all services

DeepSeek released and open-sourced the V4 preview, making 1M context standard across all official services with no tier or price split. The post says V4-Pro and V4-Flash use token compression plus DSA sparse attention to cut compute and memory costs for 1M context; legacy APIs remain for 3 months and stop after July 24.

Why it matters: DeepSeek is a flagship Chinese model vendor, and this V4 preview is a substantive release with open source and 1M context made standard across official services. HKR-H/K/R all pass: the post includes mechanisms and a migration deadline, and the tier reset makes it a same-day P1.

Hacker News front page

GPT-5.5: Mythos-Like Hacking, Open to All

XBOW says GPT-5.5 cut miss rate to 10% on its real-vulnerability benchmark, versus 40% for GPT-5 and 18% for Opus 4.6. It scored 97.5% on visual acuity and used about half the login iterations of the next-best model. The key point is black-box testing: GPT-5.5 without source beat GPT-5 with source.

Why it matters: HKR-H/K/R all pass: a major OpenAI model claim, concrete security benchmark numbers, and a clear practitioner safety nerve. The source is XBOW rather than an OpenAI launch post, so it stays below 95.

Apr 23Thursday

QbitAI · WeChat

Qwen3.6-27B open-weights, beats its 397B flagship predecessor on agentic coding

Qwen released Qwen3.6-27B and says it beats Qwen3.5-397B on 4 agentic coding benchmarks with about 1/15 the parameters. The post cites SkillsBench rising from 30.0 to 48.2, GPQA Diamond at 87.8, and AIME26 at 94.1; it uses a dense architecture, Thinking Preservation, and Gated DeltaNet, with weights on Hugging Face and ModelScope.

Why it matters: This is a substantive Qwen open-source model release with concrete agent-coding and reasoning scores, so HKR-H/K/R all pass. I keep it at 84, not higher, because the post gives strong benchmarks but no pricing, context window, or independent reproduction yet.

Bloomberg Technology

Tencent unveils a major AI foundation model upgrade, testing its new OpenAI hire

Tencent announced a major upgrade to its AI foundation model. It is the company's first high-stakes AI test since hiring a top OpenAI researcher. The post does not disclose the model name, parameter count, benchmarks, or launch timing.

Why it matters: Bloomberg provides source authority, and the framing is strong: Tencent's model release is presented as the first test of its OpenAI hire, so HKR-H and HKR-R pass. HKR-K fails because the story does not disclose the model name, size, benchmarks, or launch timing, keeping it at a

Hacker News front page

Coding Models Are Doing Too Much

The author programmatically corrupts 400 BigCodeBench problems with single-point bugs to test whether coding models over-edit code during fixes. The post defines the minimal fix as exactly reversing the corruption and measures excess changes with token-level Python Levenshtein distance. The provided body does not disclose final results, model rankings, or training gains.

Why it matters: Strong HKR-K from a concrete 400-task bug-injection eval and a clear minimal-patch metric. HKR-R also lands because over-editing is a daily pain point for Copilot/Cursor/Claude Code users, but the excerpt omits results, model rankings, and effect sizes, so this sits near the low

Apr 22Wednesday

Hacker News front page

Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

Qwen released the open-weight 27B dense model Qwen3.6-27B and made it available in Qwen Studio. It scores 77.2 on SWE-bench Verified vs. 76.2 for Qwen3.5-397B-A17B, and 59.3 on Terminal-Bench 2.0 under a 256K context and 3-hour timeout. The real takeaway is deployment: this is not a larger MoE, but a denser 27B model with stronger coding results.

Why it matters: Qwen3.6-27B is a substantive flagship-model release with open weights, concrete coding benchmarks, and a practical dense-deployment angle. HKR-H/K/R all pass, and per policy a major Chinese model launch should score on par with an equivalent US-lab release.

Apr 21Tuesday

Synced · WeChat

Anonymous world model MotuBrain tops WorldArena and RoboTwin2.0

MotuBrain ranked first on both WorldArena and RoboTwin2.0, with a 63.77 EWM Score on WorldArena and 95.8/96.1 in RoboTwin Clean and Randomized settings. The post says it also leads Motion Quality, Flow Score, and Motion Smoothness, and averages 96.0 across 50 RoboTwin tasks versus 92.3 for second place; the post does not disclose its owner, model size, or training setup. The result matters because it supports a single-model path that combines world prediction with robot action, at least on benchmarks.

Why it matters: HKR-H lands on the anonymous double-#1 hook; HKR-K lands on concrete scores across WorldArena and RoboTwin; HKR-R lands on the embodied-AI nerve around one model doing prediction and action. I kept it in the low 80s because ownership, scale, training data, and reproducibility are

Latent Space

Moonshot Kimi K2.6 open-weight model refresh aims to catch Opus 4.6

Moonshot released Kimi K2.6, a 1T-parameter MoE with 32B active and 256K context. The post cites 58.6 on SWE-Bench Pro, 4,000+ tool calls, 12+ hour runs, and 300 parallel sub-agents. The key signal is long-horizon agent execution, not only open-model scores.

Why it matters: HKR-H/K/R all pass: Kimi K2.6 has a strong race narrative, concrete model and agent metrics, and direct relevance to open-model builders. The domestic flagship release signal lifts it into P1.

Apr 20Monday

r/LocalLLaMA

Training LoRA adapters for Apple's on-device 3B model on a free Colab T4 and a Mac

The author built a QLoRA pipeline for Apple’s on-device 3B model, cutting training needs from about 24GB to about 1GB RAM and 5GB GPU, enough for a free Colab T4 or a 24GB Mac. The post says A100 LoRA, T4 QLoRA, and Mac QLoRA adapters perform about the same, raising accuracy from about 40% to 75%, or 86% with retrieval; it also reports a confirmed Apple bug that writes a hidden ~160MB cache copy per CLI call, reaching 269GB over ~300 runs.

Why it matters: A named first-person experiment with reproducible memory and accuracy numbers clears HKR-H/K/R and beats routine tutorial posts. The score stays below the 85 band because this is a single Reddit post with limited source authority and a narrow benchmark scope.

r/LocalLLaMA

Compared some models for feature planning

A Reddit user tested 9 models on planning a “load tracking” feature for a Go budgeting app, then used Claude Code to rank the generated specs, with Claude Opus 4.6 placed first. The table shows Opus 4.6 produced a 19 KB spec with 44 code reads at $2.47; GLM 5.1 ranked second and Qwen 3.6 35B fp8+vLLM ranked third. Do not treat this as a benchmark: the author says it is not representative, and the post does not disclose any manual quality review yet.

Why it matters: A named first-person test gives real workflow data, so HKR-H/K/R all pass. The ceiling stays low: one task only, ranked by Claude Code itself, and no human acceptance result is disclosed, so this lands at the low end of featured.

r/LocalLLaMA

TRELLIS.2 image-to-3D now runs on Mac (Apple Silicon) with no NVIDIA GPU required

A developer ported Microsoft's TRELLIS.2 to Apple Silicon and reports generating ~400K-vertex meshes from one photo in about 3.5 minutes on an M4 Pro with 24GB. The port replaces five CUDA-only extensions with PyTorch MPS and custom backends; texture baking takes about 18 seconds, removing the NVIDIA and cloud requirement.

Why it matters: This is a community port, not an official release, but HKR-H/K/R all pass: the hook is NVIDIA-free image-to-3D on Apple Silicon, and the post includes testable details (M4 Pro 24GB, ~400k vertices, 3.5 minutes, 5 CUDA-extension rewrites). Reddit-level source authority keeps it in

QbitAI · WeChat

Sudo, valued above $2 billion, unveils embodied model Sudo R1 with zero real-robot data and ~98% first-try grasp success

Sudo unveiled embodied model Sudo R1 and says it achieved about 98% first-try grasp success in 200+ zero-shot tests with zero real-robot training data, nearing 100% within two attempts. The post says the 60-minute run covered 100+ unseen objects, including transparent, metallic, soft, and reflective items, using integrated world-model and reinforcement-learning training on a high-fidelity simulator. It also says Sudo is valued above $2 billion and is working with CATL, but the post does not disclose round size, benchmark protocol, or third-party validation.

Why it matters: Strong HKR-H/K/R: the zero-real-data, zero-shot, 98% claim is novel and concrete, and it hits robotics' data-cost nerve. Kept below 85 because the metrics are self-reported; funding amount, benchmark definition, and third-party validation are not disclosed.

New York Times Chinese

Chinese humanoid robot 'Shandian' finishes a half marathon in 50:26, faster than the human world record

Honor’s humanoid robot Shandian finished a Beijing half marathon in 50:26, faster than Jacob Kiplimo’s 57:20 human world record. The 1.65-meter robot fell after hitting a barrier, resumed with human help, and far beat last year’s best robot time of 2:40:42. The key signal is stronger robotics engineering, not a disclosed AI leap.

Why it matters: This clears HKR-H/K/R: strong headline contrast plus concrete numbers and conditions. It stays below the top bands because this is a benchmark event, not a directly reusable model or product release, and the control stack and race-rule details are not disclosed.

Hacker News front page

Show HN: TRELLIS.2 image-to-3D running on Apple Silicon, no Nvidia GPU needed

Developer shivampkumar ported Microsoft's 4B-parameter TRELLIS.2 to Apple Silicon with PyTorch MPS for single-image 3D generation. He replaced flash_attn, nvdiffrast, and custom sparse conv kernels with pure PyTorch sparse 3D conv, SDPA attention, and Python mesh extraction. On an M4 Pro with 24GB, it generates ~400K-vertex meshes in about 3.5 minutes; slower than H100 seconds, but fully offline.

Why it matters: Strong on all HKR axes: a clear hook, concrete implementation details, and benchmark-like numbers. This is not a Microsoft model launch, but a reproducible local port with real practitioner relevance, so it lands in featured rather than p1.

Apr 19Sunday

r/LocalLLaMA

Same 9B Qwen weights: 19.1% in Aider vs 45.6% with a scaffold adapted to small local models

Using the same Qwen3.5-9B Q4 weights on the 225-task Aider Polyglot benchmark, the author changed only the scaffold and raised mean pass@2 from 19.11% to 45.56%. The little-coder setup is not a new model; it uses bounded reasoning, a write guard, explicit workspace discovery, and small per-turn skill injections. The key claim is scaffold-model fit, but the post reports only two full runs and does not disclose ablations, cross-model replications, or a second benchmark.

Why it matters: HKR-H/K/R all pass: the hook is a 2.4x jump on Aider Polyglot 225 with the same 9B Qwen weights, and the post names the scaffold mechanisms. Importance stays low-featured because evidence is thin: two full runs, no ablation, no cross-model rerun, and no second benchmark.

r/LocalLLaMA

I tested 8 LLMs as tabletop GMs: a 27B model beat the 405B on narrative quality

The author tested 8 LLMs on 6 fixed tabletop-GM scenarios, and google/gemma-3-27b-it ranked first in narrative quality with a 4.33 overall score. The probe used 8 auto metrics plus 3 LLM-judge scores, and the full run cost about $0.02; the title says a 27B beat a 405B, but the snippet does not disclose the 405B model name or full rankings.

Why it matters: A named first-person benchmark with a strong surprise hook clears HKR-H, HKR-K, and HKR-R. I kept it at featured, not higher: the source is Reddit, the post is truncated, and the 405B model name plus full ranking are not disclosed.

Apr 17Friday

Hacker News front page

Measuring Claude 4.7's tokenizer costs

The author used Anthropic's free count_tokens API to compare Claude Opus 4.6 and 4.7 on 7 real samples and 12 synthetic ones; the real-sample weighted total rose from 8,254 to 10,937 input tokens, or 1.325x. Technical docs hit 1.47x, a real CLAUDE.md file hit 1.445x, while Chinese and Japanese stayed near 1.01x. On a 20-prompt IFEval sample, 4.7 improved strict prompt-level pass rate from 85% to 90%; the post cannot isolate tokenizer effects from model weights or post-training.

Why it matters: HKR-H/K/R all land: the post has a sharp cost hook, reproducible token-count data, and clear budget impact for Claude Code users. It stays below p1 because this is a third-party measurement, not an Anthropic release, and the IFEval slice is only 20 items.

Xinzhiyuan · WeChat

Yixin says its finance Agent harness runs single tasks for 16 hours and plans an H2 open-source release

Yixin says its finance Agent harness can run a single task for 16 hours across 12 sessions, with 65% autonomous delivery. The post adds a 50k-token cap per case, projected approval speedups above 150%, and projected unit cost at one-fifth of human work; it says an open-source release is planned for H2 2026, but does not disclose the repo, license, or reproducible evals. The key signal is governance design, not the “smarter over time” framing.

Why it matters: This clears HKR-H/K/R with a rare production claim: a finance agent runs 16 hours, spans 12 sessions, hits 65% autonomous delivery, and stays under a 50k-token cap. It stays below 85 because the evidence is self-reported and the post does not disclose a repo, license, or reproduc

Hacker News front page

Qwen3.6-35B-A3B on my laptop drew me a better pelican than Claude Opus 4.7

Simon Willison ran a 20.9GB quantized Qwen3.6-35B-A3B on a MacBook Pro M5 and judged its SVG pelican output better than Claude Opus 4.7. He used LM Studio with an Unsloth Q4_K_S GGUF, then repeated the test with “a flamingo riding a unicycle” and again scored Qwen higher. This is not a general capability result; the author says this joke benchmark no longer tracks overall model usefulness in this comparison.

Why it matters: A named first-person experiment with reproducible setup gives this strong HKR-H/K/R: the headline has a sharp contrast, the post includes a 20.9GB GGUF on an M5 MacBook Pro via LM Studio, and it hits the open-local-vs-closed-frontier debate. It stays in featured, not higher, لأن/

Apr 16Thursday

36Kr (direct RSS)

Anthropic plans to release its Mythos model to UK banking institutions next week

Anthropic PBC plans to grant UK financial institutions early access to its Mythos model within the next week. The mechanism is the “Glass Wing” program for selected institutions; Anthropic says the model can identify and potentially exploit cybersecurity flaws, while the post does not disclose specs, pricing, or customer count. The key signal is controlled access, not a broad launch.

Google DeepMind

Google DeepMind releases Gemini 3.1 Flash TTS

Google DeepMind released Gemini 3.1 Flash TTS, a text-to-speech model built around controllability and expressiveness. It is in preview on the Gemini API, Google AI Studio, Vertex AI and Google Vids.

Why it matters: The post covers the new model's audio-tag controls, Elo scores and preview entry points, so you can judge how controllable speech generation has become.

Apr 13Monday

Google DeepMind

Google DeepMind releases Gemini Robotics-ER 1.6

Google DeepMind released Gemini Robotics-ER 1.6, an upgrade to its reasoning-first robotics model. It strengthens spatial reasoning and multi-view understanding, and adds gauge-reading ability.

Why it matters: The post details the new model's changes in spatial reasoning, multi-view understanding and gauge reading, plus where it is available, so you can judge progress in high-level robot reasoning.

Apr 12Sunday

X · @Yuchenj_UW

MiniMax M2.7 is open-source!

MiniMax open-sourced M2.7 and said its research agent now handles 30%–50% of the R&D workflow. The post says the agent covers literature review, experiment orchestration, log debugging, code fixes, and merge requests; M2.7 also rewrote its own harness for 100+ automated rounds, with a 30% gain on internal coding evals.

Why it matters: HKR-H/K/R all pass: open-sourcing plus a research agent doing 30%-50% of R&D is a strong hook, and the post includes 100+ self-rewrite loops with +30% internal coding eval. It stays at 78 because license, repo, benchmark context, and external reproduction are not disclosed.

Apr 10Friday

最佳拍档 (BestPartners)

LLM self-evolution: Shinka Evolve, AlphaEvolve, and sample efficiency

Sakana AI open-sourced Shinka Evolve and uses a UCB bandit to switch among GPT-5, Claude Sonnet 4.5, Gemini, and others, aiming to cut the thousands of program evaluations common in AlphaEvolve-style search. The post says it beat AlphaEvolve’s classic circle-packing result with fewer evaluations and adds full-file rewrites, crossover, editable-region guards, and a meta-notebook; the post does not disclose exact metrics, cost, or the repo link. The part to watch is surrogate-task design and hard verification: the system still needs humans to define problems.

Why it matters: Featured, not P1: HKR-H/K/R all pass. The piece has a strong hook, concrete mechanisms like UCB model routing and program crossover, and a real nerve around eval cost and hard verification. It stays at 80 because key metrics, cost, and the primary release link are not disclosed.

QbitAI · WeChat

Tencent open-sources 3B SVG model HiVG to make tokens geometry-aware

Tencent Hunyuan open-sourced the 3B-parameter HiVG, claiming 62.7%-63.8% shorter SVG sequences via hierarchical tokenization and better SVG generation metrics than GPT-5.2, Claude-4.5-Sonnet, and some 8B open models. The post reports 0.896 SSIM, 0.114 LPIPS, and 0.957 CLIP-S on Image-to-SVG; the core method packs drawing commands plus coordinates into segment tokens and uses HMN to initialize coordinate embeddings. The part to watch is token design, not parameter count; paper, code, and project page are public.

Why it matters: Tencent's HiVG earns HKR-H and HKR-K: a 3B open model claims GPT/Claude-level SVG results, and the article includes 62.7%-63.8% token compression plus SSIM 0.896, LPIPS 0.114, and CLIP-S 0.957. HKR-R is weaker because SVG generation remains niche, so it lands at the low end of `f

Apr 9Thursday

QbitAI · WeChat

Beyond MoE, Tencent introduces MoT: a 2B embodied model ranks first in 16 of 22 evaluations

Tencent Hunyuan and Robotics X released HY-Embodied-0.5; its MoT-2B uses 4B total params with 2B active and ranks first in 16 of 22 embodied evaluations. The post says it uses 100M+ embodied data, 600B+ pretraining tokens, 30M+ mid-training samples, plus visual latent tokens, bidirectional attention, RFT, RL, and online distillation. The key point is a rebuilt edge-oriented embodied stack, not a simple VLM fine-tune.

Why it matters: Strong on HKR-H/K/R: the headline has a real hook, the body includes concrete numbers and training mechanisms, and the edge-robotics angle lands with practitioners. I keep it at 83, not 85+, because this is a high-quality embodied-model release, not a broad same-day industry-def

X · @op7418

Meta releases Muse Spark model

Meta released the Muse Spark model with native multimodal reasoning, tool use, visual chain-of-thought, and multi-agent orchestration, but it is only available in the Meta AI app and is not open source for now. The snippet says its Contemplating mode coordinates multiple parallel agents for reasoning, and its Artificial Analysis score is below Gemini 3.1 Pro, GPT-5.4, and Claude Opus 4.6. The post does not disclose model size, pricing, or rollout timing.

Why it matters: A major-lab model launch plus the “poached team’s first output” angle lands HKR-H/K/R. The score stays near the featured floor because the post offers capability claims and relative benchmark placement only; params, pricing, rollout timing, and access scope are not disclosed.

Apr 8Wednesday

QbitAI · WeChat

Free open-source 2B Chinese speech model reproduces Mangzhuang Ren with high-speed tonguetwisters

ModelBest, OpenBMB, and Tsinghua University released VoxCPM 2, a 2B open speech model that supports 9 Chinese dialects, 30 foreign languages, and 48kHz audio. The post says generation often finishes within 1 second, recommends reference audio of at least 5 seconds, and supports denoising, LoRA, and full fine-tuning; the key detail is its tokenizer-free diffusion autoregressive continuous representation design.

Why it matters: This is a substantive open-source speech release, not a thin demo: the post gives 2B, 48kHz, 9 Chinese dialects, 30 languages, ref audio ≥5s, and a tokenizer-free route. HKR-H/K/R all pass, but the event is not large enough for a must-write P1.

X · @Yuchenj_UW

GLM-5.1 beat Opus 4.6, GPT-5.4, and Gemini 3.1 Pro on SWE-Bench Pro

GLM-5.1 scored 58.4 on SWE-Bench Pro, ahead of Opus 4.6 at 57.3, GPT-5.4 at 57.7, and Gemini 3.1 Pro at 54.2. The post also says it is an MIT-licensed open-weight model; the post does not disclose eval setup, cost, or whether all models were tested under identical conditions. Watch reproducibility, not a single leaderboard snapshot.

Why it matters: Open-weight GLM-5.1 beating closed leaders on SWE-Bench Pro is a real hook, and the score deltas are concrete. Source authority is weak: this is a single X post with no disclosed eval setup, cost, or equal-condition proof, so it stays low-featured rather than higher.

Apr 7Tuesday

Latent Space

[AINews] Gemma 4 crosses 2 million downloads

Google’s Gemma 4 reached about 2 million downloads in its first week. The post compares that with Gemma 3 at 6.7 million over the past year, Gemma 2 at 1.4 million since June 2024, and Qwen 3.5 at about 27 million in roughly 1.5 months. The signal for practitioners is local deployment: one iPhone 17 Pro demo ran Gemma 4 E2B at about 40 tok/s via MLX, with support across Hugging Face, vLLM, llama.cpp, Ollama, and NVIDIA.

Why it matters: HKR-H/K/R all pass: the story has a clean hook, concrete comparative download data, and a real open-model adoption nerve. It stays low-featured because this is a secondary-source uptake snapshot, not a primary Google release or a substantive capability update.

Apr 3Friday

X · @dotey

Google releases the Gemma 4 open model family under Apache 2.0

Google released the Gemma 4 family and switched the full line to Apache 2.0. The post says it includes 31B Dense, 26B MoE, E4B, and E2B; 31B and 26B support 256K context, and 31B fits on one 80GB H100. The key change is distribution terms: fewer limits on commercial use, modification, and redistribution, plus native function calling and structured JSON for agent workflows.

Why it matters: This is a substantive Google model release, with the Apache 2.0 switch carrying as much weight as the model specs. HKR-H/K/R all pass on novelty, concrete deploy details, and commercial relevance; it stays below P1 because the post lacks formal eval links and direct head-to-heads

Google DeepMind

Google DeepMind releases the Gemma 4 open model family

Google DeepMind released Gemma 4, which it calls its most intelligent open model yet, aimed at advanced reasoning and agentic workflows under an Apache 2.0 license. The family comes in four sizes: E2B, E4B, 26B MoE and 31B Dense. The 31B ranks 3rd among open models on the Arena AI text leaderboard, and the 26B ranks 6th.

Why it matters: Gemma 4 is Apache 2.0 and spans four sizes from on-device to workstation, so you can weigh deployment and fine-tuning options for open models.

Mar 31Tuesday

MIT Technology Review · AI

AI benchmarks are broken. Here’s what we need instead.

The author proposes HAIC benchmarks that evaluate AI over longer periods inside teams and workflows, not on isolated tasks alone. The post lists four shifts and cites a UK hospital study from 2021–2024 plus an 18-month humanitarian case; the key signal is coordination, error detectability, and downstream effects, not a 98% accuracy headline.

Why it matters: This hits all three HKR axes: a contrarian headline, a concrete 4-part framework with two field cases, and a strong resonance with the industry's eval-vs-production debate. It is a strong commentary piece, not a model release, benchmark launch, or research drop, so it lands in `f

Mar 24Tuesday

Mistral AI

Mistral AI releases Voxtral TTS, a 4B-parameter speech model

Mistral AI released Voxtral TTS, its first text-to-speech model. It has 4B parameters and supports nine languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi and Arabic. It handles emotional expression and zero-shot cross-lingual voice adaptation.

Why it matters: The 4B size, nine languages, 70ms latency and pricing give readers a basis for judging cost and model choice in enterprise voice agents.

Mar 17Tuesday

Mistral AI

Mistral releases Mistral Small 4, unifying reasoning, multimodal and coding

Mistral AI released Mistral Small 4, the first Mistral model to unify Magistral reasoning, Pixtral multimodal and Devstral coding-agent abilities in a single model. It ships under the Apache 2.0 license.

Why it matters: Merging reasoning, multimodal and coding agents into one open model is a direct test of what unified models do to deployment cost.

Feb 12Thursday

Ruan YiFeng's Weblog

Hands-on with Zhipu's flagship GLM-5: compared with Claude Opus 4.6 and GPT-5.3-Codex

Ruan Yifeng compared GLM-5, Claude Opus 4.6, and GPT-5.3-Codex on 4 coding tasks, and judged GLM-5 competitive with the two closed models overall. The post covers web redesign, a 3D sandbox, an Angry Birds clone, and Laravel-to-Next.js migration; in the migration task, GLM-5 and GPT-5.3 took about 5 minutes, while Opus 4.6 took about 20. The key point: this is a single-author hands-on comparison, not a standardized benchmark.

Why it matters: This clears HKR-H/K/R because it is a named first-person test with 4 tasks, video evidence, and a 5-minute versus ~20-minute gap. I did not score it higher because it is one author's evaluation, not a standardized benchmark or a broad multi-source release event.

Feb 10Tuesday

36Kr (direct RSS)

Alibaba Qwen launches the new image generation foundation model Qwen-Image-2.0

Alibaba Qwen announced Qwen-Image-2.0, an image generation foundation model, and opened API invite testing on Alibaba Cloud Bailian. Developers can also try it for free in Qwen Chat; the post does not disclose model size, pricing, eval results, or a general release date.

Feb 6Friday

TechCrunch · AI

OpenAI launches new agentic coding model minutes after Anthropic releases its own

OpenAI launched an agentic coding model minutes after Anthropic released a similar one, and the model is meant to accelerate Codex, which OpenAI launched earlier this week. The RSS snippet gives only the timing and purpose; the post does not disclose the model name, benchmarks, pricing, context length, or availability. The signal is direct competition in agentic coding, not a substantiated performance claim.

Why it matters: Major-lab product news plus a minutes-apart Anthropic clash gives this HKR-H and HKR-R. The score stays in the low featured band because HKR-K is weak: the post lacks the model name, benchmarks, price, context window, and availability.