Skip to content

#数据/训练

4 today

May 19Tuesday

Google DeepMind

Fast-tracking genetic leads to reverse cellular aging

Google DeepMind 的 Co-Scientist 正被用于加速细胞衰老研究,它扫描数万篇论文后提出 20 多个可测试的新遗传因子,其中数个经实验室验证能驱动细胞进入更年轻状态并改善整体功能。它还能将原本需长达六个月的筛选数据分析缩短至几天。

AI HOT (Curated Pool)

Cursor releases Composer 2.5 coding model

Cursor released Composer 2.5, claiming up to 10x higher efficiency on long coding tasks; the model is further trained on Moonshot’s Kimi K2.5 and uses text feedback for 100k-token-scale trajectories.

Why it matters: HKR-H/K/R all pass: Cursor is a core AI coding tool, and Composer 2.5 adds concrete claims around 10x long-task gains and Kimi K2.5 tuning. Limited sourcing and no independent eval keep it in the 78–84 band.

AI HOT (Curated Pool)

NVIDIA fine-tunes Cosmos Predict 2.5 with LoRA/DoRA for robot video generation

NVIDIA published a Hugging Face post on fine-tuning Cosmos Predict 2.5 with LoRA and DoRA to generate robot first-person videos from text prompts; the post does not disclose dataset size, training cost, or evaluation results.

Why it matters: HKR-H/K/R pass: the robot POV video angle is clickable, and LoRA/DoRA on Cosmos Predict 2.5 is a concrete mechanism. Missing dataset scale and metrics keep it in the low featured band.

May 18Monday

Synced · WeChat

ICML 2026: Huawei GTS proposes EDCO for dynamic curriculum fine-tuning

Huawei GTS proposed EDCO, a dynamic curriculum method that selects fine-tuning samples by inference entropy; prefix entropy estimation cuts per-sample scoring time from 2.24 seconds to 0.37 seconds.

Why it matters: HKR-H/K/R pass: the story has a lab-race hook, a concrete entropy-based mechanism, and a 2.24s→0.37s efficiency claim. It stays below 78 because it is still a training-method paper, not a major model or product release.

r/LocalLLaMA

I trained TIME: short context-triggered thinking on Qwen instead of overthinking

An independent author trained TIME with QLoRA on Qwen3 4B/8B/14B/32B to trigger short mid-response reasoning when context changes; the post says datasets, notebooks, scripts, curriculum, and TIMEBench are public, with 24GB VRAM enough for training up to 14B.

Why it matters: HKR-H/K/R all pass: the post has a clear tuning hook, concrete reproducible details, and strong local-LLM resonance. Reddit single-post sourcing keeps it in the 72-77 featured band, below lab-level releases.

AI HOT (Curated Pool)

Composer 2.5 release and technical analysis

Cursor released Composer 2.5, built on a Moonshot open-source checkpoint, trained with synthetic data from real codebases at 25 times the previous scale, and updated with text-feedback reinforcement learning and a sharded Muon optimizer.

Why it matters: HKR-H/K/R all pass: Cursor is a core coding-agent surface, and the post gives concrete training details around Moonshot, 25x data, RL, and Muon. It lacks benchmarks, pricing, or user-facing capability limits, so it stays in the 78–84 band.

May 17Sunday

Google DeepMind

Google DeepMind launches Gemini for Science toolset

Google DeepMind released Gemini for Science, which includes three experimental tools on Google Labs: Hypothesis Generation, built on Co-Scientist.

Why it matters: Google is packaging research prototypes like Co-Scientist and AlphaEvolve into apply-to-use science tools, showing what agentic research looks like in practice.

QbitAI · WeChat

TGO Aligns Visual Generative Models with Scalar Feedback Without Preference Pairs | ICML 2026

NUS proposed Threshold-Guided Optimization, which converts scalar feedback into positive or negative updates through a score-distribution threshold and was accepted by ICML 2026; experiments cover Stable Diffusion v1.5, FLUX, Wan 1.3B, and Meissonic across image and video generation settings.

Why it matters: HKR-H/K/R pass: the paper has a concrete mechanism and tests across SD v1.5, FLUX, Wan 1.3B, and Meissonic. Impact is research-heavy, so it lands in featured, not must-write.

Dwarkesh Patel podcast

Notes on Pretraining Parallelisms and Failed Training Runs

Dwarkesh documents pretraining failure modes and parallelism tradeoffs: expert choice and token dropping can break causality in MoE routing, FP16 collectives can bias repeated additions after values exceed 1024, pretraining FLOPs are given as 6ND, B300 HBM is listed as 288GB, and FSDP communication can reach params × 3 with reduce-scatter.

Why it matters: HKR-H/K/R all pass: Dwarkesh’s notes expose concrete pretraining failure modes and numbers. The systems-training focus is specialized, so it sits in the high-quality band rather than same-day must-write.

May 16Saturday

Google DeepMind

Finding the molecular switches behind new infectious diseases

剑桥大学 Clare Bryant 教授利用 Google Co-Scientist 研究流感等病原体跨物种传播时引发脓毒症等重症的分子开关。Co-Scientist 生成并排序假设,优先锁定一个她此前未关注的蛋白,并逐步将假设细化到具体氨基酸。Bryant 团队正构建含氨基酸突变的细胞系验证,原本需两到三年的工作预计六个月完成。

Google DeepMind

Opening new paths in aging research

Calico Life Sciences 的 Matt Onsum 与 Katherine Labbé 正使用 Google DeepMind 的 Co-Scientist 整合衰老生物学中零散的研究发现,将其转化为可验证的假设。在整合应激反应(ISR)研究中,该工具帮助团队生成了一条关于代谢如何调控 ISR 的新假设,并协助优化实验设计。相关实验已产生新发现,团队计划发表这些结果。

Google DeepMind

Accelerating discovery of liver disease mechanisms

爱丁堡大学团队用 Google Co-Scientist 研究 MASH 肝病,系统整合肝生物学与药理学证据,锁定值得关注的机制并筛选出候选组合疗法。针对 resmetirom 仅对少数合格患者有效的问题,Co-Scientist 提出 NLRP3 炎症小体是连接炎症与代谢的分子桥梁,该假说随后经实验验证,有望推动靶向双重疗法。

Google DeepMind

Uncovering repurposed medicines to fight liver fibrosis

斯坦福大学医学院遗传学家 Gary Peltz 团队在《Advanced Science》发表研究,用 Google DeepMind 的 Co-Scientist 从现有药物文献中筛选可重定位治疗肝纤维化的候选药。

May 15Friday

r/LocalLLaMA

I trained Qwen3.5 to jailbreak itself with RL, then used the failures to improve its defenses

The author built an RL-based automated red-teaming loop for Qwen3.5, raising defense rate from 64% to 92% while benign accuracy fell from 92% to 88%, and the attacker found 7 tactic families.

Why it matters: HKR-H/K/R all pass: a named first-person RL red-team loop with concrete rates and failure modes. Source is a single Reddit post without paper/code validation, so it stays below P1.

r/LocalLLaMA

I Let a Small Model Train on Its Own Mistakes; It Reached 80% on HumanEval and Beat GPT-3.5 on Math

The author fine-tuned Qwen 2.5 7B base on self-mined mistake-correction pairs, raising HumanEval from 25/164 to 112/164; Qwen 2.5 14B used 100 pairs and a 95-minute H100 run costing $3.50.

Why it matters: HKR-H/K/R pass: the hook is strong and the post gives samples, H100 time, cost, and HumanEval deltas. Kept at 78 because it is a single Reddit post and the 80% claim differs from 112/164.

r/LocalLLaMA

MOOSE-Star (ICML 2026): 7B Model and 108K-Paper Dataset for Scientific Hypothesis Discovery

MiroMind researchers released the MOOSE-Star collection with three 7B models and TOMATO-Star, a dataset of 108,717 NCBI papers. MS-IR-7B reaches 54.37% inspiration-retrieval accuracy, uses DeepSeek-R1-Distill-Qwen-7B as its base, runs at about 14GB fp16, and supports llama.cpp, vLLM, and SGLang.

Why it matters: HKR-H/K/R all pass via the local 7B research-agent hook and concrete dataset metrics. Single Reddit source and limited lab gravity keep it below the must-write band.

May 14Thursday

r/LocalLLaMA

Automated AI researcher running locally with llama.cpp

Hugging Face’s ml-intern added local-model support through llama.cpp and ollama; the post says Qwen3.6-35B-A3B can orchestrate CPU/GPU sandboxes and Hub jobs to run an end-to-end SFT workflow.

Why it matters: HKR-H/K/R all pass, but this is a Reddit-sourced open-source tool update, not a major model release. Local sandbox and Hub-job orchestration for SFT put it just above the featured threshold.

Xinzhiyuan · WeChat

Yuandong Tian and Seven Co-Founders Launch Recursive Superintelligence at $4.65B Valuation

Recursive Superintelligence, founded by Yuandong Tian and seven other AI researchers, has a 25-person team, $650 million in funding, and a $4.65 billion valuation, with a stated goal to automate evaluation, data filtering, training, post-training, and research-direction selection.

Why it matters: All three HKR axes pass: a $650M raise at a $4.65B valuation for a 25-person recursive-improvement startup is not routine funding. The stated target spans evals, data selection, training, post-training, and research selection.

Synced · WeChat

China in Focus: PsiBot Uses 100,000 Hours of Human Data for Embodied AI

PsiBot says it uses 100,000 hours of human operation data to train robot policies, with the W0 world model acting only as a training-time transfer module while deployment runs R2 alone.

Why it matters: HKR-H/K/R all pass, but the facts come mainly from company framing and lack an artifact link, benchmark, or third-party replication. This fits a solid robotics research/product story, not the 78+ band.

Synced · WeChat

ACL 2026: Alibaba DAMO I²B-LPO Improves RLVR Exploration

Alibaba DAMO Academy introduced I²B-LPO, an RLVR post-training framework that branches rollouts at high-entropy nodes and filters them with an information-bottleneck self-reward, reporting up to 5.3% accuracy gains and 7.4% semantic-diversity gains on math benchmarks using Qwen2.5-7B and Qwen3-14B.

Why it matters: HKR-H/K/R all pass: the ACL 2026 DAMO paper has a clear RLVR exploration hook, concrete I²B-LPO mechanics, and benchmark gains. It is still a training-method paper, not a major model or product release, so 78 fits the lower good-quality band.

May 13Wednesday

AI HOT (Curated Pool)

SenseNova-U1 Technical Report Released: Guide to Native Multimodal Model Building

SenseTime released the SenseNova-U1 technical report, covering six-stage training, RL post-training, and distillation; the open-source SenseNova-U1-A3B-MoT uses an MoE architecture and activates only 3 billion parameters.

Why it matters: HKR-H/K/R all pass: A3B-MoT’s 3B active parameters and six-stage training recipe give concrete signal. The score stays near the featured floor because this is a vendor post with no benchmarks, license terms, or reproduction details disclosed.

Xinzhiyuan · WeChat

Tsinghua-affiliated team open-sources MiniCPM-V 4.6, a 1.3B model tunable on one RTX 4090

ModelBest, Tsinghua University, and OpenBMB open-sourced MiniCPM-V 4.6, a 1.3B multimodal model that supports full fine-tuning on one RTX 4090 and offers 4x/16x visual token compression for accuracy or speed trade-offs.

Why it matters: HKR-H/K/R all pass: the story gives a concrete open-source multimodal release with size, hardware condition, and token-compression details. It lowers local fine-tuning cost, but it is not a frontier-lab flagship release, so 78–84 fits.

May 12Tuesday

AI HOT (Curated Pool)

How Open Model Ecosystems Compound

China’s open AI model community forms a self-reinforcing loop, with domestic open model downloads rising by more than 200% quarter over quarter.

Why it matters: HKR-H/K/R all pass: the flywheel framing is clickable, the article gives a >200% QoQ download claim, and the topic hits China open-model competition. It is strong commentary, not a model launch, so 78 featured.

Google DeepMind

Google DeepMind publishes Co-Scientist multi-agent research system

Google DeepMind published Co-Scientist research in Nature, introducing a Gemini-based multi-agent AI system that iteratively generates, debates and evolves new hypotheses for complex scientific problems.

Why it matters: The post discloses the system's three-stage collaboration mechanism and deployment cases at several labs, showing how AI takes part in scientific hypothesis generation.

QbitAI · WeChat

Shanghai AI Lab Study: SFT Generalizes Under Three Conditions

Shanghai AI Lab, Shanghai Jiao Tong University, and USTC tested Long-CoT SFT on Qwen3-14B-Base and found that cross-domain performance recovered and improved after 8 epochs, with generalization conditioned on optimization depth, data quality and structure, and base-model capability.

Why it matters: HKR-H/K/R all pass: the SFT-generalization claim has a clear hook, Qwen3-14B-Base plus an 8-epoch finding, and direct relevance to fine-tuning teams. It lacks deployment impact or full benchmark detail, so it stays in the mid-featured band.

r/LocalLLaMA

Prompt caching for RL training: 7.5x speedup on long-prompt, short-response workloads

The author proposes prompt caching for RL training. On Qwen3.5-4B, it reports a 7.5x speedup with 16k-token prompts and 64-token outputs, and the G=8 example with 1000-token prompts and 100-token responses reduces 8800 processed tokens to 1800 unique tokens.

Why it matters: HKR-H/K/R all pass: the angle is novel, and the post gives 16k/64 plus G=8 token-dedup numbers. Kept at 78 because this is a single Reddit post without independent replication or a paper/code artifact disclosed.

May 10Sunday

Synced · WeChat

Turing Award Winner Sutton Uses a 1967 Formula to Improve Streaming Reinforcement Learning

Richard Sutton and coauthors proposed Intentional Updates, which derive the step size from the desired output change; Intentional AC approached SAC on MuJoCo under batch=1 streaming training without replay, while each update used about 1/140 of SAC’s FLOPs.

Why it matters: HKR-H/K/R all pass: Sutton's name, Intentional Updates, MuJoCo conditions, and 1/140 SAC FLOPs give it substance. Strong research signal, but less market-moving than a major LLM product release, so it stays in the 78–84 band.

Xinzhiyuan · WeChat

Next-ToBE Targets Short-Sighted Next-Token Prediction in LLMs at ICLR 2026

East China Normal University and Fudan University researchers proposed Next-ToBE, a training objective that keeps standard autoregressive inference while adding a soft target over future-token windows, and the article reports the method ranked best in 35 of 36 experiments across Qwen2.5-Math-1.5B, Qwen2.5-Math-7B, and Llama3.1-8B-Instruct.

Why it matters: HKR-H and HKR-K pass: the mechanism and 35/36 result are specific, and next-token training is a real debate. The item stays near the featured floor because no artifact, reproduction detail, or production claim is disclosed.

May 9Saturday

Xinzhiyuan · WeChat

NVIDIA, AMD, and Intel Back RadixArk’s $100M Seed Round

RadixArk announced a $100 million seed round on May 5 at a $400 million post-money valuation, led by Accel and co-led by Spark Capital, with participation from NVentures, AMD, MediaTek, Databricks, and other investors tied to AI infrastructure.

Why it matters: HKR-H/K/R all pass: a $100M seed round, $400M post-money valuation, and chip/data investors create an AI-infra rivalry angle. It remains a single-company funding item with no product benchmarks or customer data, so it sits near the featured floor.

May 8Friday

Synced · WeChat

SGLang Team Launches RadixArk With $100M Seed Round

RadixArk announced a $100 million seed round on May 5 at a $400 million post-money valuation, while its SGLang inference project has 27K+ GitHub stars and deployments across 400K+ GPUs.

Why it matters: HKR-H/K/R all pass: the round size, valuation, and deployment numbers are concrete, and SGLang is a known inference stack. It is still a startup funding and infra-roadmap story, not a major model release, so it stays in the 78–84 featured band.

QbitAI · WeChat

All Labs Watch ByteDance, Everyone Praises DeepSeek: A U.S. Researcher’s 36-Hour China AI Trip

Ai2 researcher Nathan Lambert visited Zhipu, Moonshot AI, Tsinghua, Meituan, Xiaomi, and 01.AI within 36 hours, and said Chinese labs closely watch ByteDance and respect DeepSeek, while student participation in core work, open source habits, and in-house control of the technical stack mark key differences.

Why it matters: HKR-H/K/R all pass: the piece has a named US researcher’s dense China-lab tour plus concrete claims on ByteDance, DeepSeek, open source, and in-house stacks. It is strong industry field reporting, not a model launch or major deal, so it sits at featured rather than p1.

May 6Wednesday

QbitAI · WeChat

Claude Team Tests New Training Method on Qwen

Anthropic proposed MSM training between pretraining and alignment fine-tuning. Tests on Qwen2.5-32B and Qwen3-32B cut misalignment from 68% and 54% to 5% and 7%. The key point is MSM complements AFT rather than replacing it.

Why it matters: HKR-H/K/R all pass: Anthropic offers a concrete MSM alignment method with Qwen2.5-32B and Qwen3-32B rate drops. It is strong safety research, not a model launch or major product update, so 82 fits.

Financial Times · Technology

Meta and Zuckerberg Sued by Publishers Over ‘Massive’ Copyright Infringement

Five major publishing groups sued Meta and Zuckerberg over copyrighted works allegedly used to train Llama AI models. The RSS snippet does not disclose work counts, damages, court venue, or training-data mechanism.

Why it matters: HKR-H/K/R all pass: FT covers a Meta/Llama copyright suit with Zuckerberg named. Missing court, damages, work counts, and data mechanics keep it at the featured threshold.

May 3Sunday

r/LocalLLaMA

Built a C++17 transformer from scratch with 0.83M params and CPU training

Reddit user Suspicious_Gap1121 released Quadtrix.cpp, a C++17 GPT-style model with 0.83M parameters. It uses 4 layers, 4 heads, 200d width, and a 128-character context; one CPU core trained on 31.4M characters for 76.2 minutes to 1.6371 nats val loss. The key detail is handwritten backprop for LayerNorm, attention, Q/K/V, dropout, and AdamW without PyTorch, BLAS, or autograd.

Why it matters: HKR-H/K/R all pass: the no-framework C++17 build is clickable, the training setup is specific, and local-LLM builders care about dependency-free control. It stays in the 72–77 band because it is a small personal project.

May 2Saturday

MIT Technology Review · AI

Musk v. Altman week 1: Musk says xAI distills OpenAI models

Elon Musk testified in week 1 of Musk v. Altman, saying he gave OpenAI $38 million in funding. He asks the court to remove Sam Altman and Greg Brockman and unwind OpenAI’s for-profit restructuring. The sharp detail: Musk said xAI partly distills OpenAI models, while OpenAI previously accused DeepSeek of similar conduct.

Why it matters: HKR-H/K/R all pass: the trial has conflict, concrete facts include $38M and xAI’s distillation admission, and OpenAI governance is a live nerve. No ruling or product-level change, so it stays below P1.

May 1Friday

Synced · WeChat

The Evolution of RL: From PPO to MaxRL in LLM Reasoning Training

Jiqizhixin translated Alexander Weers' article on RL algorithms for LLM reasoning from 2024 to 2026. It covers REINFORCE, PPO, GRPO, RLOO, Dr. GRPO, DAPO, CISPO, MaxRL, DPPO, and ScaleRL, comparing critic removal, clipping, normalization, and pass@k goals. The key signal is mechanism choice, not algorithm names.

Why it matters: A strong technical explainer, not a model or paper release. HKR-H comes from the PPO→MaxRL arc, HKR-K from concrete mechanism comparisons, and HKR-R from live RL-recipe choices; the higher technical bar keeps it in low featured.

The Verge · AI

Elon Musk confirms xAI used OpenAI’s models to train Grok

Elon Musk testified Thursday in a California federal court that xAI used OpenAI models to improve Grok. The mechanism described is model distillation: a larger teacher model transfers knowledge to a smaller student model. The post does not disclose which OpenAI models, data volume, or training runs.

Why it matters: HKR-H/K/R all pass: Musk confirmed in court that xAI used OpenAI models to train Grok, with distillation as the mechanism. Missing model names, scale, and runs keeps it in 78–84, not P1.

TechCrunch · AI

Elon Musk testifies that xAI trained Grok on OpenAI models

Elon Musk testified that xAI trained Grok on OpenAI models. The post only says distillation concerns frontier labs; it does not disclose scale, model versions, or case context.

Why it matters: All HKR axes pass: Musk’s testimony puts xAI, Grok, OpenAI, and distillation evidence in one story. Missing model versions, scale, and full litigation context keep it at the low end of the 85+ band.

Apr 30Thursday

MIT Technology Review · AI

Goodfire releases Silico, a mechanistic interpretability tool for debugging LLMs

Goodfire released Silico, letting engineers inspect and adjust LLM parameters during training. It maps neurons and pathways; one Qwen 3 neuron triggered trolley-problem-style outputs. Pricing is case-by-case, and the post does not disclose rates.

Why it matters: HKR-H/K/R all pass: Silico offers a concrete interpretability-debugging mechanism. It stays at 76 because this is a startup product preview with no pricing or adoption scale disclosed.