Skip to content

#多模态

1 today

Apr 21Tuesday

Latent Space

Training Transformers to Address the 95% Failure Rate in Cancer Trials — Noetik

Noetik uses TARIO-2 to predict tumor spatial transcriptomics, targeting a 95% cancer-trial failure rate. GSK signed a $50M technology deal, and TARIO-2 predicts a ~19,000-gene spatial map from routine H&E assays. The key issue is patient-tumor-treatment matching, not the claim that AI cures cancer.

Why it matters: HKR-H/K/R pass: the hook ties 95% cancer-trial failure to transformer matching, with TARIO-2 predicting ~19k spatial genes from H&E and a $50M GSK deal. Vertical AI productization, not a general model release, keeps it at featured threshold.

Apr 20Monday

r/LocalLLaMA

TRELLIS.2 image-to-3D now runs on Mac (Apple Silicon) with no NVIDIA GPU required

A developer ported Microsoft's TRELLIS.2 to Apple Silicon and reports generating ~400K-vertex meshes from one photo in about 3.5 minutes on an M4 Pro with 24GB. The port replaces five CUDA-only extensions with PyTorch MPS and custom backends; texture baking takes about 18 seconds, removing the NVIDIA and cloud requirement.

Why it matters: This is a community port, not an official release, but HKR-H/K/R all pass: the hook is NVIDIA-free image-to-3D on Apple Silicon, and the post includes testable details (M4 Pro 24GB, ~400k vertices, 3.5 minutes, 5 CUDA-extension rewrites). Reddit-level source authority keeps it in

QbitAI · WeChat

Sudo, valued above $2 billion, unveils embodied model Sudo R1 with zero real-robot data and ~98% first-try grasp success

Sudo unveiled embodied model Sudo R1 and says it achieved about 98% first-try grasp success in 200+ zero-shot tests with zero real-robot training data, nearing 100% within two attempts. The post says the 60-minute run covered 100+ unseen objects, including transparent, metallic, soft, and reflective items, using integrated world-model and reinforcement-learning training on a high-fidelity simulator. It also says Sudo is valued above $2 billion and is working with CATL, but the post does not disclose round size, benchmark protocol, or third-party validation.

Why it matters: Strong HKR-H/K/R: the zero-real-data, zero-shot, 98% claim is novel and concrete, and it hits robotics' data-cost nerve. Kept below 85 because the metrics are self-reported; funding amount, benchmark definition, and third-party validation are not disclosed.

Synced · WeChat

In the first year of “deployment mode,” AgiBot expanded its rollout plans to seven solutions

AgiBot said at its April 17 Shanghai event that it released 4 robots, 6 AI models, and 7 standardized deployment solutions, and framed 2026 as the first year of embodied AI “deployment mode.” The post cites concrete metrics: Expedition A3 runs 8-10 hours, WITA Omni 1.0 targets sub-500ms interaction latency, and BFM was trained on 100 million-plus frames and 700 hours of motion-capture data; it also claims 5,100-plus shipments and 39% share in 2025, with the 10,000th robot rolling off in March 2026. The real point for practitioners is repeatable delivery rather than launch volume: the post lists 7 scenarios from 3C line loading to patrol, but independent validation details are not disclosed.

Why it matters: HKR-H/K/R all pass: the story leads with seven deployment playbooks and backs it with shipment, share, latency, and training figures. It stays at 76 because key outcome claims are company-sourced; customer impact and independent validation are not disclosed.

Hacker News front page

Show HN: TRELLIS.2 image-to-3D running on Apple Silicon, no Nvidia GPU needed

Developer shivampkumar ported Microsoft's 4B-parameter TRELLIS.2 to Apple Silicon with PyTorch MPS for single-image 3D generation. He replaced flash_attn, nvdiffrast, and custom sparse conv kernels with pure PyTorch sparse 3D conv, SDPA attention, and Python mesh extraction. On an M4 Pro with 24GB, it generates ~400K-vertex meshes in about 3.5 minutes; slower than H100 seconds, but fully offline.

Why it matters: Strong on all HKR axes: a clear hook, concrete implementation details, and benchmark-like numbers. This is not a Microsoft model launch, but a reproducible local port with real practitioner relevance, so it lands in featured rather than p1.

Apr 19Sunday

QbitAI · WeChat

Did Musk Really Sell Lao Gan Ma on Douyin?

QbitAI says the shown “Musk selling Lao Gan Ma on Douyin” and “GTA-6 crossover” images were generated by OpenAI GPT Image 2; the claimed 100K+ live viewers were part of fake visuals. The post argues Image 2 can render realistic posters, game screenshots, and readable long text, and links that to Codex-style UI workflows; the post does not disclose pricing, rollout scope, or launch timing. The real issue is verification: image realism is eroding “photo as evidence.”

Why it matters: HKR-H/K/R all pass: the hook is novel, the article shows a concrete capability jump, and the trust/verification angle resonates with practitioners. It stops short of p1 because the body does not disclose rollout, pricing, or an official launch scope.

QbitAI · WeChat

Amap unveiled ABot, its first full-stack embodied AI stack for AGI, and claimed 15 SOTA results

Amap unveiled embodied AI stack ABot and claimed SOTA on 15 metrics. The post says ABot-3DGS builds 10k-scale 3D scenes from centimeter-level map data, while ABot-PhysWorld uses a 14B DiT and 3M real manipulation videos. What matters is the interactive world model and VLA loop; the post does not disclose the 15 benchmarks, exact metrics, or the open-source timeline and scope.

Why it matters: HKR-H/K/R all pass: the angle is surprising, and the post includes concrete mechanisms and numbers. It stays below the 80s because the claimed 15 SOTAs lack benchmark names, and the open-source scope and timeline are not disclosed.

Apr 18Saturday

Synced · WeChat

Claude Design enters research preview for generating mockups, prototypes, and slides

Anthropic launched Claude Design in research preview for Claude Pro, Max, Team, and Enterprise users, covering mockups, prototypes, slides, and one-pagers. Powered by Claude Opus 4.7, it can ingest codebases, images, DOCX, PPTX, XLSX, and web captures, then export to Canva, PDF, PPTX, and HTML; the headline cites Figma and Adobe stock drops, but the post does not disclose the moves. The real signal is the workflow link from design system ingestion to handoff into Claude Code.

Why it matters: HKR-H/K/R all pass: the design-workflow angle is novel, the post gives concrete mechanism details, and the Figma/Adobe pressure point resonates. I keep it below 85 because the stock-drop claim has no numbers and there is no user test, pricing, or adoption data.

Bloomberg Technology

OpenAI’s Former Product Chief and Sora Head Leave Company

OpenAI is losing two leaders: its former product chief and the head of Sora; the title confirms the count is two. The post does not disclose timing, reasons, successors, or names; the key watchpoint is whether the Sora org changes as well.

Why it matters: A Bloomberg personnel report on OpenAI and the Sora line clears HKR-H/K/R: surprise, a concrete new fact, and direct relevance to org stability and roadmap risk. The body gives roles only; names, reasons, and succession are missing, so it stays below the 95+ industry-shaking band

X · @dotey

Anthropic launches Claude Design, a conversational design generation product

Anthropic released Claude Design in research preview and is rolling it out to Pro, Max, Team, and Enterprise subscribers. Powered by Claude Opus 4.7, it can start from text, images, docs, or web clips, then iterate via chat, comments, direct edits, and sliders. On first use it reads a team's codebase and design files to build a design system; outputs export to Canva, PDF, PPTX, or standalone HTML, with one-click handoff to Claude Code.

Why it matters: Anthropic pushes Claude into a new design workflow, so HKR-H/K/R all pass. The post includes rollout tiers, model name, first-run design-system ingest, and export paths; strong featured story, but still a gradual research preview rather than a top-tier model release.

Apr 17Friday

Hacker News front page

Introducing Claude Design by Anthropic Labs

Anthropic launched Claude Design on April 17, 2026, in research preview for Claude Pro, Max, Team, and Enterprise subscribers. Powered by Claude Opus 4.7, it generates designs from text, images, DOCX, PPTX, XLSX, and codebases, and exports to Canva, PDF, PPTX, or HTML. The key detail is its one-step handoff bundle to Claude Code; the post does not disclose standalone pricing beyond existing plan limits.

Why it matters: This is a substantive Anthropic product launch, not a routine feature add. HKR-H/K/R all pass on novelty, concrete deployment details, and workflow resonance; the research-preview scope and limited pricing detail keep it at 84 instead of p1.

X · @claudeai

Introducing Claude Design by Anthropic Labs: make prototypes, slides, and one-pagers by talking to Claude

Anthropic Labs launched Claude Design in research preview for Pro, Max, Team, and Enterprise plans, letting users create prototypes, slides, and one-pagers by talking to Claude. The post says it runs on Claude Opus 4.7, Anthropic’s most capable vision model; the post does not disclose pricing, output constraints, or a detailed rollout schedule. The thing to watch is the interactive design workflow, not just another writing surface.

Why it matters: This is a first-party Anthropic capability launch, and HKR-H/K/R all pass: Claude expands from chat into prototypes, slides, and one-pagers, with paid tiers and Opus 4.7 named. It stays below p1 because price, export limits, and rollout timing are not disclosed.

Xinzhiyuan · WeChat

AgiBot says robots have entered the deployment phase with 8-hour continuous factory work

At APC 2026 on April 17, AgiBot defined 2026 as year one of the “deployment phase” and said its robots had run for 8 hours on a real production line. The clearest case in the post is Genie G2 at Longcheer’s Nanchang factory: 2,283 loading tasks, over 99.5% success, and 18-20 seconds per cycle; these figures are company disclosures, and the post does not disclose independent audit results. The real signal is scale and line integration: AgiBot said it shipped over 5,100 units in 2025 and reached 10,000 cumulative units by March 2026, while Longcheer plans nearly 1,000 deployments.

Why it matters: HKR-H/K/R all land: the 'demo is over' angle is clickable, and the post gives testable factory data—8 hours, 2,283 runs, >99.5% success, 18-20s cycle. Not P1 because the evidence is company-reported and the article shows no independent audit or cross-site replication.

MIT Technology Review · AI

How robots learn: A brief, contemporary history

Companies and investors put $6.1 billion into humanoid robots in 2025, 4x 2024, and MIT Technology Review attributes the surge to a shift in how robots learn. The piece highlights two mechanisms: around 2015, simulation plus reward signals enabled millions of trial-and-error runs; after ChatGPT in 2022, robotics models took images, sensors, and joint states to predict dozens of motor commands per second. The key change is data-driven learning over hand-written rules; the provided text is truncated, so later examples are not fully disclosed.

Why it matters: HKR-H/K/R all pass: the $6.1B and 4x funding jump provide the hook, and the piece maps the shift from sim+RL to multimodal action models. It stays in the lower featured band because this is commentary rather than a new release, and the excerpt is truncated on company-level detail

X · @dotey

Seedance 2.0 API is now available on Volcano Engine and BytePlus

Volcano Engine has released the Seedance 2.0 API for enterprises, individual developers, and overseas users via BytePlus; China pricing is RMB 46 per million tokens, or about RMB 1 per second for pure video generation. The post says it supports text, image, audio, and video inputs, plus face verification, portrait authorization, and 10,000+ preset avatars for workflow automation; overseas pricing is not disclosed here. The part to watch is orchestration: the post cites up to 10x efficiency gains, but does not disclose a common benchmark or model specs.

Why it matters: HKR-H/K/R all pass: the overseas rollout is a real hook, and the post includes usable pricing and modality details for practitioners. It stays at 74 because this is an API availability update, not a major model launch, and the post does not disclose model params, benchmark method

X · @op7418

Seedance 2.0 API is now fully open

Volcano Engine has opened the Seedance 2.0 API to domestic users, while BytePlus serves overseas access; the API currently accepts 4 input modalities: text, image, audio, and video. The post also confirms face registration, portrait authorization, and preset virtual avatars, but does not disclose pricing, rate limits, model variants, or regional availability. The real watchpoint is whether video-agent workflows can be wired through Skills and MCP, not the ecosystem rhetoric.

Why it matters: This is a real product update from ByteDance’s stack: HKR-H on full API availability, HKR-K on 4-modal input and consent mechanics, and HKR-R on builder demand for deployable video APIs. I keep it at 75 because pricing, rate limits, regional rollout details, and quality evidence

Latent Space

[AINews] Anthropic Claude Opus 4.7 - one step better than 4.6 in every dimension

Anthropic launched Claude Opus 4.7 at the same $5/$25 per million input/output tokens; the post says 4.7-low through 4.7-high each outperform the matching higher 4.6 tiers. Reported changes include a new xhigh reasoning tier, Claude Code defaulting to xhigh, an 11-point gain on SWE-Bench Pro, and image input up to 2,576 px on the long edge (~3.75 MP). Do not overread the tokenizer change: the same input can use up to 35% more tokens, but the post says total token use still falls by up to 50% from prior equivalents.

Why it matters: Anthropic's flagship-model release fits the policy's 85–94 band. HKR-H/K/R all pass because the post gives concrete pricing, benchmark, image-limit, and token-accounting changes that hit Claude users' core coding and cost concerns.

Hacker News front page

Qwen3.6-35B-A3B on my laptop drew me a better pelican than Claude Opus 4.7

Simon Willison ran a 20.9GB quantized Qwen3.6-35B-A3B on a MacBook Pro M5 and judged its SVG pelican output better than Claude Opus 4.7. He used LM Studio with an Unsloth Q4_K_S GGUF, then repeated the test with “a flamingo riding a unicycle” and again scored Qwen higher. This is not a general capability result; the author says this joke benchmark no longer tracks overall model usefulness in this comparison.

Why it matters: A named first-person experiment with reproducible setup gives this strong HKR-H/K/R: the headline has a sharp contrast, the post includes a 20.9GB GGUF on an M5 MacBook Pro via LM Studio, and it hits the open-local-vs-closed-frontier debate. It stays in featured, not higher, لأن/

X · @dotey

browser-use open-sources video-use, a Claude Code skill that turns raw camera footage into edited videos

browser-use released video-use, a Claude Code skill that turns raw footage into a final.mp4 automatically. It converts footage into ElevenLabs word-level timestamp transcripts, shrinking one asset to about 12KB; the post says feeding frames directly would cost about 45 million tokens. The key detail is the structured editing pipeline: the model mostly reads text, uses timeline images only at uncertain cuts, and runs up to 3 self-check repair passes after rendering.

Why it matters: Strong HKR-H/K/R: the result is instantly clickable, and the post includes a concrete text-first editing architecture with 12KB vs about 45M-token economics. Kept below higher bands because this is a builder-facing Claude Code skill, not a platform-level release.

Apr 16Thursday

X · @dotey

Anthropic officially releases Claude Opus 4.7 at unchanged pricing

Anthropic released Claude Opus 4.7 at unchanged pricing: $5 per million input tokens and $25 per million output tokens; the API name is claude-opus-4-7, now live across Claude, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry. The post gives two concrete changes: vision input now supports up to 2576 pixels on the long edge, and the new tokenizer can raise token usage to 1.0-1.35x for the same text. Watch migration cost, not list price; higher reasoning settings and multi-turn runs can increase output length and bills.

Why it matters: An Anthropic substantive model release belongs in the 85+ band, and this is not just a rename: the 2576px vision limit and 1.0–1.35x tokenization shift affect migration tests and billing immediately. HKR-H/K/R all pass, so it clears p1.

X · @op7418

Anthropic releases Claude Opus 4.7 with the following main updates

Anthropic has rolled out Claude Opus 4.7 across all Claude products and the API, with pricing unchanged from Opus 4.6. The post lists better long-horizon task handling, more precise instruction following, self-verification before reporting, vision support up to 2,576-pixel long-edge images, plus Claude Code Ultra Review, an xhigh thinking level, and auto-approval for Max users.

Why it matters: This is a substantive Anthropic model release across Claude and the API, with testable details: unchanged pricing, a 2,576px vision limit, self-checking outputs, and Claude Code workflow changes. HKR-H/K/R all pass; it fits the same-day must-write band, so p1.

Hacker News front page

Introducing Claude Opus 4.7

Anthropic released Claude Opus 4.7 on Apr. 16 at the same price as Opus 4.6: $5 per million input tokens and $25 per million output tokens. The post says it improves on Opus 4.6 in advanced software engineering, long-running tasks, and higher-resolution vision, and ships across Claude, the API, Amazon Bedrock, Vertex AI, and Microsoft Foundry. The key detail is the first deployment of Anthropic’s cyber request blocking on a less capable model; the post cites benchmark gains but does not fully disclose every score in text.

Why it matters: Anthropic shipping Claude Opus 4.7 is a same-day write: GA, unchanged $5/$25 pricing, and rollout across Claude, API, Bedrock, Vertex AI, and Foundry give it direct workflow impact. HKR-H/K/R all pass, but the post does not publish full benchmark scores.

Hacker News front page

Qwen3.6-35B-A3B: Agentic coding power, now open to all

Qwen released Qwen3.6-35B-A3B as open weights, with 35B total parameters and 3B active parameters. The post reports 73.4 on SWE-bench Verified, 51.5 on Terminal-Bench 2.0, and 92.0 on RefCOCO. The key point is agentic coding and multimodal performance at a 3B active-parameter budget, with weights, Qwen Studio, and API access available.

r/LocalLLaMA

Qwen3.6-35B-A3B released

Qwen released Qwen3.6-35B-A3B as open source under Apache 2.0; it is a sparse MoE with 35B total parameters and 3B active. The post also claims agentic coding, strong multimodal perception and reasoning, plus thinking and non-thinking modes; the post does not disclose benchmarks, context length, or latency.

Why it matters: HKR-H/K/R all pass: a new open Qwen model is timely, and the post confirms 35B total, 3B active, and Apache 2.0. The score stays at 82 because this is still a launch post; benchmarks, context window, latency, and multimodal details are not disclosed here.

Hacker News front page

Darkbloom – Private inference on idle Macs

Eigen Labs launched Darkbloom, linking 100M+ Apple Silicon Macs into a decentralized inference network. It offers an OpenAI-compatible API, claims end-to-end encryption plus hardware attestation, and lists prices up to 70% below OpenRouter comps. The real point is the trust model: hardware keys, hardened runtime, and signed outputs are disclosed, but enterprise audit scope still needs the paper.

Why it matters: HKR-H/K/R all pass: the idle-Mac inference angle is novel, and the post includes concrete scale, API, encryption, and price claims. I keep it at 80 because this is still a self-published research preview; audit scope, network reliability, and attack boundaries are not yet third-p

TechCrunch · AI

Google rolls out a native Gemini app for Mac

Google launched a native Gemini app for Mac on April 15 for all users worldwide on macOS 15 and later, with Option + Space as the summon shortcut. Users can share their screen or local files with Gemini, and the app also supports image generation with Nano Banana and video generation with Veo. The key shift is desktop access plus live context sharing, not just another client.

Why it matters: Google shipping a native Gemini app for Mac clears HKR-H/K/R: the hook is desktop entry, the new facts are hotkey and context sharing, and the resonance is the desktop assistant race. Still a mid-weight product update, not a model leap, so it sits at the low end of featured.

Apr 14Tuesday

最佳拍档 (BestPartners)

Global GPU shortage worsens: H100 rental prices rose nearly 40% in five months

SemiAnalysis says Nvidia H100 one-year rental pricing rose from $1.70 to $2.35 per GPU-hour between Oct 2025 and Mar 2026, up nearly 40% in five months. The post attributes this to Anthropic-driven demand, multi-agent and media generation workloads, and memory cost spikes, with LPDDR5 and DDR5 contract prices up about 4x and 5x year over year; much new capacity is already prebooked. The key variable is the supply gap, not Blackwell refreshes alone.

Why it matters: Strong HKR-H/K/R: the story has a sharp price-shock hook, concrete market data, and clear resonance with compute-cost anxiety. It stays below P1 because this is a secondary video synthesis of a SemiAnalysis report, not a primary company or product announcement.

Apr 11Saturday

QbitAI · WeChat

OpenClaw-style methods reach multimodal generation, with a 6B model beating Nano Banana 2 on some tasks

A team led by Shanghai AI Laboratory introduced GEMS, adding Agent Loop, Memory, and Skills to multimodal generation, and reports that 6B Z-Image-Turbo beats Nano Banana 2 on some tasks. The post reports +14.22 average gains on 5 mainstream tasks and +8.92 over the best baseline on 4 downstream tasks; the paper and code are public, but the post does not disclose Nano Banana 2's full setup.

Why it matters: Strong HKR-H/K/R: the hook is a 6B multimodal model beating Nano Banana 2, and the post includes mechanism plus testable deltas (+14.22 / +8.92) with paper and code. It stays below P1 because the article does not disclose the full Nano Banana 2 comparison setup.

QbitAI · WeChat

Liu Zhuang and Danqi Chen team open-source Vero, a general visual reasoning RL framework, reaching SOTA with zero thinking data

Princeton researchers including Liu Zhuang and Danqi Chen open-sourced Vero, an RL framework for visual reasoning, and report beating Qwen3-VL-8B-Thinking on 23 of 30 benchmarks. The post says Vero uses 600K samples filtered from 59 datasets, task-routed rewards, and single-stage RL across six task groups. The key point is the mechanism mix: no private thinking data, but the post does not disclose training cost or base model configuration.

Why it matters: Featured on HKR-H/K/R: the zero-thinking-data claim is a strong hook, and the post includes concrete benchmark and method details. I keep it in the low 80s because training cost, base model choice, and full reproduction conditions are not disclosed.

QbitAI · WeChat

A Chinese embodied model reached global No.1 as a 100,000-hour human dataset for robots was released

Psibot says it released a 100,889-hour human-plus-robot manipulation dataset, and that Psi-R2 ranked first on AllenAI’s MolmoSpace benchmark. The post lists 95,472 hours of human data, 5,417 hours of robot data, 1,000 open-sourced hours, 294 scenes, 4,821 tasks, and 1,382 objects; Psi-W0 adds 30% failure samples, and Psi-R2 latency drops from 2.2s to under 100ms. The key point is the data loop and benchmark framing: the post claims nearly 10x higher success, but does not disclose task setup, full baselines, or statistics.

Why it matters: HKR-H/K/R all pass: the data scale, failure-sample mix, and latency cut are concrete and discussable. I keep it at 80 because the No.1 ranking and near-10x success claim lack task setup, full baselines, and statistical detail in the body.

Apr 10Friday

QbitAI · WeChat

Tencent open-sources 3B SVG model HiVG to make tokens geometry-aware

Tencent Hunyuan open-sourced the 3B-parameter HiVG, claiming 62.7%-63.8% shorter SVG sequences via hierarchical tokenization and better SVG generation metrics than GPT-5.2, Claude-4.5-Sonnet, and some 8B open models. The post reports 0.896 SSIM, 0.114 LPIPS, and 0.957 CLIP-S on Image-to-SVG; the core method packs drawing commands plus coordinates into segment tokens and uses HMN to initialize coordinate embeddings. The part to watch is token design, not parameter count; paper, code, and project page are public.

Why it matters: Tencent's HiVG earns HKR-H and HKR-K: a 3B open model claims GPT/Claude-level SVG results, and the article includes 62.7%-63.8% token compression plus SSIM 0.896, LPIPS 0.114, and CLIP-S 0.957. HKR-R is weaker because SVG generation remains niche, so it lands at the low end of `f

Apr 9Thursday

QbitAI · WeChat

Beyond MoE, Tencent introduces MoT: a 2B embodied model ranks first in 16 of 22 evaluations

Tencent Hunyuan and Robotics X released HY-Embodied-0.5; its MoT-2B uses 4B total params with 2B active and ranks first in 16 of 22 embodied evaluations. The post says it uses 100M+ embodied data, 600B+ pretraining tokens, 30M+ mid-training samples, plus visual latent tokens, bidirectional attention, RFT, RL, and online distillation. The key point is a rebuilt edge-oriented embodied stack, not a simple VLM fine-tune.

Why it matters: Strong on HKR-H/K/R: the headline has a real hook, the body includes concrete numbers and training mechanisms, and the edge-robotics angle lands with practitioners. I keep it at 83, not 85+, because this is a high-quality embodied-model release, not a broad same-day industry-def

X · @op7418

Meta releases Muse Spark model

Meta released the Muse Spark model with native multimodal reasoning, tool use, visual chain-of-thought, and multi-agent orchestration, but it is only available in the Meta AI app and is not open source for now. The snippet says its Contemplating mode coordinates multiple parallel agents for reasoning, and its Artificial Analysis score is below Gemini 3.1 Pro, GPT-5.4, and Claude Opus 4.6. The post does not disclose model size, pricing, or rollout timing.

Why it matters: A major-lab model launch plus the “poached team’s first output” angle lands HKR-H/K/R. The score stays near the featured floor because the post offers capability claims and relative benchmark placement only; params, pricing, rollout timing, and access scope are not disclosed.

Apr 8Wednesday

QbitAI · WeChat

Xiaomi unveils two AI audio frameworks: Any2Speech and Midasheng-audio-generate

Xiaomi's large-model application team introduced Xiaomi Any2Speech and Midasheng-audio-generate. Any2Speech generates up to about 10 minutes per inference, while the other model turns one text prompt into mixed audio with speech, music, and ambient sound. The post names GST labeling, dual-path planning with dimension dropout, Flow Matching, and five-field structured labels; benchmark scores, training scale, and commercial terms are not disclosed.

Why it matters: Xiaomi released two audio-generation frameworks with a clear hook and concrete mechanisms, so HKR-H and HKR-K pass. HKR-R is weaker because benchmark results, training data scale, open-source status, and commercial terms are not disclosed, so this sits at the low end of featured.

QbitAI · WeChat

After a late-night update, DeepSeek reportedly said: I am V4?

DeepSeek added Fast and Expert modes on its web app and started gray-testing a Vision model; the claim that Expert mode is V4 comes only from user probes and the model’s own replies. The post gives one concrete detail: Expert mode focuses on code, web, and harder generation tasks, is supply-limited, does not support multimodal or file upload, and one user reported a length cap at about 133K tokens. What matters is the official model ID and context spec; the post does not disclose them, pricing, or a release timeline.

Why it matters: HKR-H is strong on the 'I am V4' hook. HKR-K and HKR-R pass because the post gives testable mode behavior and a ~133K token limit, and DeepSeek silent swaps are highly discussable. The score stays in the mid-70s because the model name, price, and context window remain unconfirmed

Apr 7Tuesday

Latent Space

[AINews] Gemma 4 crosses 2 million downloads

Google’s Gemma 4 reached about 2 million downloads in its first week. The post compares that with Gemma 3 at 6.7 million over the past year, Gemma 2 at 1.4 million since June 2024, and Qwen 3.5 at about 27 million in roughly 1.5 months. The signal for practitioners is local deployment: one iPhone 17 Pro demo ran Gemma 4 E2B at about 40 tok/s via MLX, with support across Hugging Face, vLLM, llama.cpp, Ollama, and NVIDIA.

Why it matters: HKR-H/K/R all pass: the story has a clean hook, concrete comparative download data, and a real open-model adoption nerve. It stays low-featured because this is a secondary-source uptake snapshot, not a primary Google release or a substantive capability update.

Apr 3Friday

X · @op7418

Alibaba released the Qwen 3.6 Plus model

Alibaba released Qwen 3.6 Plus with a 1M context window, 64K input, and nearly 991K max output. The RSS snippet says it improves over Qwen 3.5 on agents, coding, image, and document understanding, priced at RMB 2 per 1M input tokens and RMB 12 per 1M output tokens; benchmark scores and test conditions are not disclosed.

Why it matters: Alibaba shipping Qwen 3.6 Plus is a substantive domestic model update. HKR-H/K/R all pass on the 1M-context plus pricing combo, but it stays below P1 because benchmark scores, baselines, and test conditions are not disclosed in the body.

X · @op7418

Google releases Gemma 4 for on-device use under Apache 2.0

Google released Gemma 4 in four variants—E2B, E4B, 26B MoE, and 31B Dense—targeting phones, edge devices, and up to single-H100 workstations. The RSS snippet says the 26B MoE activates 3.8B parameters and adds native function calling, JSON output, multimodal I/O, speech-to-text, and Apache 2.0 licensing; the post does not disclose benchmarks, context length, or rollout details.

Why it matters: Google releasing Gemma 4 is a substantive open-model update. HKR-H/K/R all pass on the size spread, 3.8B-active MoE detail, and deployment-cost relevance; it stays at 81 because benchmarks, context window, and test conditions are not disclosed here.

X · @dotey

LatePost on DeepSeek before V4: traits, organization, and Liang Wenfeng's goals

LatePost says DeepSeek has confirmed 4 core departures, and V4's large model slipped from around Lunar New Year to April; the report says it will likely remain open source. The snippet cites 2x-3x recruiting offers, some 8-digit packages, a 100-plus research team, and a shift from CUDA/Triton to TileLang for domestic GPU adaptation. The real signal is strategy: DeepSeek had spent less on agents and coding, but now names an agent product role; the post does not disclose V4's size, price, or benchmarks.

Why it matters: This is not the V4 launch, but it carries real signal: four confirmed departures, an April delay, a 100+ research team, and partial migration from CUDA/Triton to TileLang. HKR-H/K/R all pass; missing V4 specs, price, and benchmarks keeps it below launch-tier or p1.

X · @dotey

Google releases the Gemma 4 open model family under Apache 2.0

Google released the Gemma 4 family and switched the full line to Apache 2.0. The post says it includes 31B Dense, 26B MoE, E4B, and E2B; 31B and 26B support 256K context, and 31B fits on one 80GB H100. The key change is distribution terms: fewer limits on commercial use, modification, and redistribution, plus native function calling and structured JSON for agent workflows.

Why it matters: This is a substantive Google model release, with the Apache 2.0 switch carrying as much weight as the model specs. HKR-H/K/R all pass on novelty, concrete deploy details, and commercial relevance; it stays below P1 because the post lacks formal eval links and direct head-to-heads