Skip to content

#数据/训练

0 today

Sep 8Tuesday

Google DeepMind

Google DeepMind releases AlphaGenome Atlas, predicting every single-base variant in the human genome

Google DeepMind released AlphaGenome Atlas, a platform holding effect predictions for 9 billion single-nucleotide variants across the human genome. It spans 1PB, more than 30 times the size of the AlphaFold Database.

Why it matters: The post gives the 9 billion-variant prediction dataset and its AVI scoring, showing what a new tool for interpreting genomic variants looks like.

Jun 10Wednesday

r/LocalLLaMA

Fine-tuned Qwen2.5-7B to 96% of Claude Haiku using about $3 of API calls

A Reddit user fine-tuned Qwen2.5-7B with 1,040 DV-DPO preference pairs, costing about $3 at Claude Haiku rates, and reported 96% composite performance versus Claude Haiku on a domain task, with 11-second latency on a 4-bit T4 setup.

Why it matters: HKR-H/K/R all pass: the hook is cheap fine-tuning, with concrete DV-DPO and latency numbers. Kept at 74 because it is a single Reddit post; task scope and eval details are thin.

Jun 8Monday

AI HOT (Curated Pool)

VoxCPM2 technical report released

OpenBMB released the VoxCPM2 technical report, covering a 2B-parameter speech generation model trained on more than 2 million hours of multilingual speech data, with support for 30 languages and 9 Chinese dialects.

Why it matters: HKR-H/K/R pass via the 2B size, 2M+ training hours, and dialect coverage; the score stays at the low end of 78–84 because the post lacks benchmarks, license terms, and adoption data.

Google DeepMind

Google DeepMind publishes Sierra Leone AI tutoring trial results

Google DeepMind published results from a pre-registered randomized controlled trial in Sierra Leone. Students using Guided Learning gained 0.258 standard deviations in math over the control group, equal to roughly 1.2 to 1.7 years of normal learning progress in eight weeks.

Why it matters: It gives quantified RCT results and interaction data from a real classroom, showing where AI tutoring helps and where it does not.

Jun 7Sunday

QbitAI · WeChat

Kuaishou Kling Proposes VLM-as-Teacher for Test-Time Video Reasoning Optimization

City University and Kuaishou Kling proposed VLM-as-Teacher, which uses VLM feedback to optimize VGM LoRA at test time, reporting a 16.7-point average gain and raising VBVR-Bench from 0.666 to 0.781.

Why it matters: HKR-H/K/R all pass: the story has a novel test-time VLM-teacher hook, concrete VBVR-Bench gains from 0.666 to 0.781, and clear resonance around controllable video generation. It is a strong research release, not a must-write model launch.

AI HOT (Curated Pool)

Five Labs, Five Minds: Building a Multi-Model Financial Drama Game with Small Models

Thousand Token Wood v2 uses four small models from different labs to drive agents in a financial simulation game, with vLLM 0.22.1’s CUDA toolkit dependency identified as the main serving friction, while a fine-tuned 0.5B Qwen reached 0% self-trading and 100% valid quotes.

Why it matters: HKR-H/K/R all pass: the small-model finance game is a real hook, with vLLM and 0.5B Qwen metrics, plus agent-engineering resonance. Scope remains an experiment, so it sits in low featured.

Jun 3Wednesday

Synced · WeChat

Understanding SFT Mechanisms in LLMs: Resolving Practice Disputes and Avoiding Wasted Compute

Junpeng Zhang and coauthors argue that SFT on highly homogeneous data has an effective window of only hundreds to about 1,000 training steps, and their interaction-based warning signal detects overfitting before loss gaps appear, saving roughly 30%–50% of training compute.

Why it matters: HKR-H/K/R all pass: the paper gives testable SFT windows, earlier overfitting warnings, and 30%-50% compute savings. It is strong research, not a major model or product release, so it stays below 85.

Jun 1Monday

r/LocalLLaMA

I bolted an 8-arm reasoning MoE onto a frozen 1.4B Mamba backbone on a single RTX 3060

The author trained Mamba-Titan-1.4B-Reasoning on a 12GB RTX 3060: a frozen 1.4B Mamba-1 backbone with 8 trainable MoE arms, 2.54B total parameters, Top-2 routing at layers 24/25, and about 50% math accuracy.

Why it matters: HKR-H/K/R all pass via a numbered first-person experiment, but it is a single Reddit post with no independent replication and a fairly technical setup, so it stays in the low featured band.

May 31Sunday

QbitAI · WeChat

Fudan and Tongyi introduce ToolCUA for GUI-Tool path selection in agents

Fudan University and Tongyi Lab introduced ToolCUA-8B, which reaches 46.85% accuracy on OSWorld-MCP after training with about 4k synthetic tools and 180k interleaved GUI-Tool trajectory steps.

Why it matters: HKR-H/K/R all pass: the tool-selection failure hook is concrete, with OSWorld-MCP 46.85% and 180k steps. It stays in the 78–84 band because this is a research release, not a major model or product launch.

May 30Saturday

Synced · WeChat

CUHK Pion optimizer updates LLMs on iso-spectral manifolds to address AdamW and Muon instability

CUHK and collaborators introduced Pion, an optimizer that preserves weight singular values through orthogonal equivalence transformations, and reported that it kept a 60M normalization-free LLaMA-like model stable for 9.6B training tokens while AdamW and Muon collapsed with NaNs.

Why it matters: HKR-H/K/R pass: the hook is AdamW/Muon NaN instability, with a concrete isospectral update and 9.6B-token run. Niche optimizer math keeps it in 78–84, not same-day product news.

May 29Friday

AI HOT (Curated Pool)

Adam's Law: Prompts Written with High-Frequency Words Work Better

FaceMind tested 100 languages and four core tasks, finding that, with semantics unchanged, prompts or fine-tuning text using higher-frequency expressions from pretraining data improves large language model performance.

Why it matters: HKR-H/K/R all pass: the claim is counterintuitive and backed by 100 languages and four task types. Missing models, datasets, and effect sizes keep it in the low featured band.

Synced · WeChat

The Ma Jiaqi Failure Exposed an LLM Issue He Spotted in the Shower a Year Earlier

FaceMind links low-frequency token degradation to two papers: SLoW appeared at EMNLP 2025, Adam's Law was accepted as an ACL 2026 Oral, and high-frequency rewriting raised DeepSeek-V3 math accuracy from 63.55% to 71.54%.

Why it matters: HKR-H/K/R all pass: the odd celebrity-token hook is clickable, and the post gives a mechanism plus a 63.55%→71.54% DeepSeek-V3 result. Practical research signal, but not a major model launch.

May 28Thursday

AI HOT (Curated Pool)

NVIDIA Releases AI Framework Polar, Raising Codex Benchmark Score by 594.74%

NVIDIA’s research team open-sourced Polar, an agent reinforcement learning framework that connects GRPO training at the model API boundary without rewriting Codex CLI, Claude Code, Qwen Code, or Pi; on Qwen3.5-4B, Polar raised Codex pass@1 on SWE-Bench Verified from 3.8% to 26.4%, while prefix_merging cut training steps from 1,185 to 218.

Why it matters: HKR-H/K/R all pass: NVIDIA open-sourced Polar with a concrete GRPO mechanism and SWE-Bench Verified numbers. This is a strong research/open-source item, not a major model or product release, so it stays in the 78–84 band.

May 27Wednesday

Synced · WeChat

AMD paper: FP4 training instability is not caused by insufficient randomness

AMD and Penn State pretrained Llama 3.1-8B with MXFP4 on MI355X native FP4 hardware, achieving 9-10% end-to-end speedup over an FP8 baseline, while the paper identifies Wgrad quantization as the bottleneck that raises token overhead to 26-27% without deterministic Hadamard stabilization.

Why it matters: HKR-H/K/R all pass: a counterintuitive FP4 claim, concrete Llama 3.1-8B numbers, and a cost/hardware nerve. The topic is narrower training-infra research, so it stays in the 78-84 band.

May 24Sunday

r/LocalLLaMA

BitCPM-CANN: Native 1.58-Bit Large Language Model Training on Ascend NPU

OpenBMB released BitCPM-CANN, a 1.58-bit QAT training stack on Ascend NPU with 0.5B, 1B, 3B, and 8B models trained from scratch, where the 1B to 8B variants retain 95.7%–97.2% of full-precision MiniCPM4 performance across 11 benchmarks.

Why it matters: HKR-H/K/R pass: low-bit native training on Ascend is novel, and the summary gives sizes plus retention rates. Reddit-only sourcing and no throughput or reproduction details keep it at the featured floor.

May 21Thursday

Xinzhiyuan · WeChat

USTC Papers Study Lifelong Learning for LMMs via Multimodal Knowledge Injection

USTC researchers released MMEVOKE and KORE: MMEVOKE contains 9,422 samples across 159 subcategories, while KORE uses knowledge-tree augmentation and null-space constrained fine-tuning to reduce catastrophic forgetting during multimodal knowledge injection.

Why it matters: HKR-K and HKR-R are solid: the post gives dataset size plus a concrete fine-tuning mechanism. It stays in the low featured band because this is paper-level knowledge injection without production evidence or full reproducibility details.

r/LocalLLaMA

HRM 1B

Sapientinc released HRM-Text 1B Base and its training code, and the paper claims competitive performance against 2–7B open models while using 100–900x fewer training tokens and 96–432x less estimated compute, with training on 16 H100 GPUs taking about 46 hours and costing about $1,472.

Why it matters: HKR-H/K/R all pass: HRM-Text 1B has concrete low-cost training numbers and released code. Capped at 80 because this is a Reddit item and the efficiency claim still lacks independent evaluation.

May 19Tuesday

QbitAI · WeChat

JD and CAS IIE Publish Three Papers Defining Self-Taught RLVR

JD and CAS IIE released three Self-Taught RLVR papers covering RLSD, NPO, and CoPD; RLSD reports that 200 training steps on Qwen3-VL-8B-Instruct exceed GRPO at 400 steps across 8 benchmarks.

Why it matters: HKR-H/K/R pass: self-taught RLVR is a clear hook; RLSD reports 8 benchmarks and a 200-vs-400-step GRPO comparison; it hits reasoning fine-tuning cost. Not a top-lab model launch and replication heat is undisclosed, so it stays low featured.

Google DeepMind

Fast-tracking genetic leads to reverse cellular aging

Google DeepMind 的 Co-Scientist 正被用于加速细胞衰老研究,它扫描数万篇论文后提出 20 多个可测试的新遗传因子,其中数个经实验室验证能驱动细胞进入更年轻状态并改善整体功能。它还能将原本需长达六个月的筛选数据分析缩短至几天。

May 18Monday

Synced · WeChat

ICML 2026: Huawei GTS proposes EDCO for dynamic curriculum fine-tuning

Huawei GTS proposed EDCO, a dynamic curriculum method that selects fine-tuning samples by inference entropy; prefix entropy estimation cuts per-sample scoring time from 2.24 seconds to 0.37 seconds.

Why it matters: HKR-H/K/R pass: the story has a lab-race hook, a concrete entropy-based mechanism, and a 2.24s→0.37s efficiency claim. It stays below 78 because it is still a training-method paper, not a major model or product release.

r/LocalLLaMA

I trained TIME: short context-triggered thinking on Qwen instead of overthinking

An independent author trained TIME with QLoRA on Qwen3 4B/8B/14B/32B to trigger short mid-response reasoning when context changes; the post says datasets, notebooks, scripts, curriculum, and TIMEBench are public, with 24GB VRAM enough for training up to 14B.

Why it matters: HKR-H/K/R all pass: the post has a clear tuning hook, concrete reproducible details, and strong local-LLM resonance. Reddit single-post sourcing keeps it in the 72-77 featured band, below lab-level releases.

May 17Sunday

QbitAI · WeChat

TGO Aligns Visual Generative Models with Scalar Feedback Without Preference Pairs | ICML 2026

NUS proposed Threshold-Guided Optimization, which converts scalar feedback into positive or negative updates through a score-distribution threshold and was accepted by ICML 2026; experiments cover Stable Diffusion v1.5, FLUX, Wan 1.3B, and Meissonic across image and video generation settings.

Why it matters: HKR-H/K/R pass: the paper has a concrete mechanism and tests across SD v1.5, FLUX, Wan 1.3B, and Meissonic. Impact is research-heavy, so it lands in featured, not must-write.

May 16Saturday

Google DeepMind

Finding the molecular switches behind new infectious diseases

剑桥大学 Clare Bryant 教授利用 Google Co-Scientist 研究流感等病原体跨物种传播时引发脓毒症等重症的分子开关。Co-Scientist 生成并排序假设,优先锁定一个她此前未关注的蛋白,并逐步将假设细化到具体氨基酸。Bryant 团队正构建含氨基酸突变的细胞系验证,原本需两到三年的工作预计六个月完成。

Google DeepMind

Opening new paths in aging research

Calico Life Sciences 的 Matt Onsum 与 Katherine Labbé 正使用 Google DeepMind 的 Co-Scientist 整合衰老生物学中零散的研究发现,将其转化为可验证的假设。在整合应激反应(ISR)研究中,该工具帮助团队生成了一条关于代谢如何调控 ISR 的新假设,并协助优化实验设计。相关实验已产生新发现,团队计划发表这些结果。

Google DeepMind

Accelerating discovery of liver disease mechanisms

爱丁堡大学团队用 Google Co-Scientist 研究 MASH 肝病,系统整合肝生物学与药理学证据,锁定值得关注的机制并筛选出候选组合疗法。针对 resmetirom 仅对少数合格患者有效的问题,Co-Scientist 提出 NLRP3 炎症小体是连接炎症与代谢的分子桥梁,该假说随后经实验验证,有望推动靶向双重疗法。

Google DeepMind

Uncovering repurposed medicines to fight liver fibrosis

斯坦福大学医学院遗传学家 Gary Peltz 团队在《Advanced Science》发表研究,用 Google DeepMind 的 Co-Scientist 从现有药物文献中筛选可重定位治疗肝纤维化的候选药。

May 15Friday

r/LocalLLaMA

I Let a Small Model Train on Its Own Mistakes; It Reached 80% on HumanEval and Beat GPT-3.5 on Math

The author fine-tuned Qwen 2.5 7B base on self-mined mistake-correction pairs, raising HumanEval from 25/164 to 112/164; Qwen 2.5 14B used 100 pairs and a 95-minute H100 run costing $3.50.

Why it matters: HKR-H/K/R pass: the hook is strong and the post gives samples, H100 time, cost, and HumanEval deltas. Kept at 78 because it is a single Reddit post and the 80% claim differs from 112/164.

r/LocalLLaMA

MOOSE-Star (ICML 2026): 7B Model and 108K-Paper Dataset for Scientific Hypothesis Discovery

MiroMind researchers released the MOOSE-Star collection with three 7B models and TOMATO-Star, a dataset of 108,717 NCBI papers. MS-IR-7B reaches 54.37% inspiration-retrieval accuracy, uses DeepSeek-R1-Distill-Qwen-7B as its base, runs at about 14GB fp16, and supports llama.cpp, vLLM, and SGLang.

Why it matters: HKR-H/K/R all pass via the local 7B research-agent hook and concrete dataset metrics. Single Reddit source and limited lab gravity keep it below the must-write band.

May 14Thursday

Synced · WeChat

China in Focus: PsiBot Uses 100,000 Hours of Human Data for Embodied AI

PsiBot says it uses 100,000 hours of human operation data to train robot policies, with the W0 world model acting only as a training-time transfer module while deployment runs R2 alone.

Why it matters: HKR-H/K/R all pass, but the facts come mainly from company framing and lack an artifact link, benchmark, or third-party replication. This fits a solid robotics research/product story, not the 78+ band.

Synced · WeChat

ACL 2026: Alibaba DAMO I²B-LPO Improves RLVR Exploration

Alibaba DAMO Academy introduced I²B-LPO, an RLVR post-training framework that branches rollouts at high-entropy nodes and filters them with an information-bottleneck self-reward, reporting up to 5.3% accuracy gains and 7.4% semantic-diversity gains on math benchmarks using Qwen2.5-7B and Qwen3-14B.

Why it matters: HKR-H/K/R all pass: the ACL 2026 DAMO paper has a clear RLVR exploration hook, concrete I²B-LPO mechanics, and benchmark gains. It is still a training-method paper, not a major model or product release, so 78 fits the lower good-quality band.

May 13Wednesday

AI HOT (Curated Pool)

SenseNova-U1 Technical Report Released: Guide to Native Multimodal Model Building

SenseTime released the SenseNova-U1 technical report, covering six-stage training, RL post-training, and distillation; the open-source SenseNova-U1-A3B-MoT uses an MoE architecture and activates only 3 billion parameters.

Why it matters: HKR-H/K/R all pass: A3B-MoT’s 3B active parameters and six-stage training recipe give concrete signal. The score stays near the featured floor because this is a vendor post with no benchmarks, license terms, or reproduction details disclosed.

May 12Tuesday

Google DeepMind

Google DeepMind publishes Co-Scientist multi-agent research system

Google DeepMind published Co-Scientist research in Nature, introducing a Gemini-based multi-agent AI system that iteratively generates, debates and evolves new hypotheses for complex scientific problems.

Why it matters: The post discloses the system's three-stage collaboration mechanism and deployment cases at several labs, showing how AI takes part in scientific hypothesis generation.

QbitAI · WeChat

Shanghai AI Lab Study: SFT Generalizes Under Three Conditions

Shanghai AI Lab, Shanghai Jiao Tong University, and USTC tested Long-CoT SFT on Qwen3-14B-Base and found that cross-domain performance recovered and improved after 8 epochs, with generalization conditioned on optimization depth, data quality and structure, and base-model capability.

Why it matters: HKR-H/K/R all pass: the SFT-generalization claim has a clear hook, Qwen3-14B-Base plus an 8-epoch finding, and direct relevance to fine-tuning teams. It lacks deployment impact or full benchmark detail, so it stays in the mid-featured band.

r/LocalLLaMA

Prompt caching for RL training: 7.5x speedup on long-prompt, short-response workloads

The author proposes prompt caching for RL training. On Qwen3.5-4B, it reports a 7.5x speedup with 16k-token prompts and 64-token outputs, and the G=8 example with 1000-token prompts and 100-token responses reduces 8800 processed tokens to 1800 unique tokens.

Why it matters: HKR-H/K/R all pass: the angle is novel, and the post gives 16k/64 plus G=8 token-dedup numbers. Kept at 78 because this is a single Reddit post without independent replication or a paper/code artifact disclosed.

May 10Sunday

Synced · WeChat

Turing Award Winner Sutton Uses a 1967 Formula to Improve Streaming Reinforcement Learning

Richard Sutton and coauthors proposed Intentional Updates, which derive the step size from the desired output change; Intentional AC approached SAC on MuJoCo under batch=1 streaming training without replay, while each update used about 1/140 of SAC’s FLOPs.

Why it matters: HKR-H/K/R all pass: Sutton's name, Intentional Updates, MuJoCo conditions, and 1/140 SAC FLOPs give it substance. Strong research signal, but less market-moving than a major LLM product release, so it stays in the 78–84 band.

Xinzhiyuan · WeChat

Next-ToBE Targets Short-Sighted Next-Token Prediction in LLMs at ICLR 2026

East China Normal University and Fudan University researchers proposed Next-ToBE, a training objective that keeps standard autoregressive inference while adding a soft target over future-token windows, and the article reports the method ranked best in 35 of 36 experiments across Qwen2.5-Math-1.5B, Qwen2.5-Math-7B, and Llama3.1-8B-Instruct.

Why it matters: HKR-H and HKR-K pass: the mechanism and 35/36 result are specific, and next-token training is a real debate. The item stays near the featured floor because no artifact, reproduction detail, or production claim is disclosed.

May 6Wednesday

QbitAI · WeChat

Claude Team Tests New Training Method on Qwen

Anthropic proposed MSM training between pretraining and alignment fine-tuning. Tests on Qwen2.5-32B and Qwen3-32B cut misalignment from 68% and 54% to 5% and 7%. The key point is MSM complements AFT rather than replacing it.

Why it matters: HKR-H/K/R all pass: Anthropic offers a concrete MSM alignment method with Qwen2.5-32B and Qwen3-32B rate drops. It is strong safety research, not a model launch or major product update, so 82 fits.

May 1Friday

Synced · WeChat

The Evolution of RL: From PPO to MaxRL in LLM Reasoning Training

Jiqizhixin translated Alexander Weers' article on RL algorithms for LLM reasoning from 2024 to 2026. It covers REINFORCE, PPO, GRPO, RLOO, Dr. GRPO, DAPO, CISPO, MaxRL, DPPO, and ScaleRL, comparing critic removal, clipping, normalization, and pass@k goals. The key signal is mechanism choice, not algorithm names.

Why it matters: A strong technical explainer, not a model or paper release. HKR-H comes from the PPO→MaxRL arc, HKR-K from concrete mechanism comparisons, and HKR-R from live RL-recipe choices; the higher technical bar keeps it in low featured.

Apr 30Thursday

Synced · WeChat

After Generalist, Jianlan Luo’s Team Releases LWD for Embodied AI Training

Jianlan Luo’s team and Agibot released LWD, tested on 16 Agibot G1 robots in real settings. LWD Online scored 0.95 across 8 tasks and 0.91 on long-horizon tasks. Its offline-to-online RL uses failures as data; failed trajectories were 34.8% of a 652.5-hour pool.

Why it matters: HKR-H/K/R all pass: LWD has real-robot scale, task counts, success rates, and failure-trajectory share. Robotics is narrower than a foundation-model launch, so it lands at 78, not P1.