Skip to content

#语音

0 today

Jun 8Monday

AI HOT (Curated Pool)

VoxCPM2 technical report released

OpenBMB released the VoxCPM2 technical report, covering a 2B-parameter speech generation model trained on more than 2 million hours of multilingual speech data, with support for 30 languages and 9 Chinese dialects.

Why it matters: HKR-H/K/R pass via the 2B size, 2M+ training hours, and dialect coverage; the score stays at the low end of 78–84 because the post lacks benchmarks, license terms, and adoption data.

May 12Tuesday

Latent Space

Thinking Machines' Native Interaction Models: TML-Interaction-Small 276B-A12B Advances Realtime Voice

Thinking Machines released TML-Interaction-Small, a 276B-parameter MoE model with 12B active parameters, and the post says it advances realtime voice through 200ms time-aligned microturns, encoder-free early fusion for audio and images under 200ms, and benchmark wins over GPT-Realtime-2 and Gemini 3.1-Flash.

Why it matters: HKR-H/K/R all pass: TML-Interaction-Small gives architecture, active parameters, 200ms interaction, and named rivals. Benchmarks still need replication, but a real-time voice SOTA claim is same-day material.

Apr 30Thursday

Google DeepMind

Google DeepMind announces AI co-clinician medical research program

Google DeepMind announced an AI co-clinician research program, exploring how AI agents can assist patient care under a doctor's clinical supervision. In a blinded evaluation of 98 real primary care queries, the system made no critical errors in 97 cases, and doctors preferred its answers over mainstream evidence synthesis tools. On 140 consultation skills, it matched or beat primary care physicians on 68, but expert physicians were still better overall at spotting red flags and key physical exams.

Why it matters: Google DeepMind published its AI co-clinician research program and a multimodal consultation evaluation, showing where medical agents' abilities currently end.

Apr 28Tuesday

QbitAI · WeChat

ModelBest Releases MiniCPM-o 4.5 Technical Report for Consumer-GPU Deployment

ModelBest, OpenBMB, Tsinghua THUNLP and THUMAI released the MiniCPM-o 4.5 technical report, covering a roughly 9B-parameter model. It supports video, audio and text streams; a 12GB RTX 5070 runs full-duplex mode at RTF 0.4. The key mechanism is Omni-Flow: a unified timeline with time-division multiplexing, without external VAD.

Why it matters: HKR-H/K/R all pass: a 9B omni model runs full-duplex on a 12GB RTX 5070 with RTF 0.4, using Omni-Flow timeline alignment. It is below a frontier-lab flagship release, so 78–84 fits.

Mar 21, 2025Friday

OpenAI News

Early methods for studying affective use and emotional well-being on ChatGPT

OpenAI and MIT Media Lab studied affective use on ChatGPT with two tracks: nearly 40 million interactions in an observational analysis and a 4-week RCT with nearly 1,000 participants. The post says emotional engagement is rare overall and concentrated in a small subset of heavy Advanced Voice Mode users; the provided body does not fully disclose all quantitative well-being results. Watch subgroup effects, not platform averages.

Aug 8, 2024Thursday

OpenAI News

GPT-4o System Card

OpenAI published the GPT-4o System Card on August 8, 2024, reporting 3 of 4 Preparedness categories as low risk and persuasion as borderline medium. The post says GPT-4o accepts text, audio, image, and video inputs, responds to audio in as little as 232 ms with a 320 ms average, and is 50% cheaper than GPT-4 Turbo in the API. The key issue for practitioners is voice safety: the card names unauthorized voice generation, speaker identification, and sensitive trait attribution, and says only models with post-mitigation scores at medium or below can be deployed.

Why it matters: This is not a routine post: it adds concrete preparedness ratings, 232ms voice latency, and a clear deployment threshold. HKR-H/K/R all pass, but it is a safety disclosure rather than a new model or major launch, so it lands as featured, not p1.