Skip to content

Data & training

The training side: datasets, synthetic data, pre- and post-training methods, compute and training cost.

Latest picks

21–40 of 124

Jun 8Monday

AI HOT (Curated Pool)

Hivemind launches continuous learning for AI coding agents

Hivemind released continuous learning for AI coding agents, collecting trajectories from Claude Code, Codex, Cursor, Hermes, and Pi, converting them into reusable skills stored in users’ cloud storage, with SkillOpt matching or leading all 52 test settings.

Why it matters: HKR-H/K/R all pass, but this is a mid-weight Hivemind feature launch without major-lab weight or cross-source lift. The 52-setting result gives it enough substance for low featured.

AI HOT (Curated Pool)

VoxCPM2 technical report released

OpenBMB released the VoxCPM2 technical report, covering a 2B-parameter speech generation model trained on more than 2 million hours of multilingual speech data, with support for 30 languages and 9 Chinese dialects.

Why it matters: HKR-H/K/R pass via the 2B size, 2M+ training hours, and dialect coverage; the score stays at the low end of 78–84 because the post lacks benchmarks, license terms, and adoption data.

Google DeepMind

Google DeepMind publishes Sierra Leone AI tutoring trial results

Google DeepMind published results from a pre-registered randomized controlled trial in Sierra Leone. Students using Guided Learning gained 0.258 standard deviations in math over the control group, equal to roughly 1.2 to 1.7 years of normal learning progress in eight weeks.

Why it matters: It gives quantified RCT results and interaction data from a real classroom, showing where AI tutoring helps and where it does not.

Jun 7Sunday

QbitAI · WeChat

Kuaishou Kling Proposes VLM-as-Teacher for Test-Time Video Reasoning Optimization

City University and Kuaishou Kling proposed VLM-as-Teacher, which uses VLM feedback to optimize VGM LoRA at test time, reporting a 16.7-point average gain and raising VBVR-Bench from 0.666 to 0.781.

Why it matters: HKR-H/K/R all pass: the story has a novel test-time VLM-teacher hook, concrete VBVR-Bench gains from 0.666 to 0.781, and clear resonance around controllable video generation. It is a strong research release, not a must-write model launch.

AI HOT (Curated Pool)

AI Substitution Wave: Three Forces Reshape Cost Structures

Coinbase, Lindy, Harvey, and Cursor shifted workloads to cheaper models; Harvey reported Kimi 2.6 reached a 15% all-pass rate on Legal Agent Benchmark, versus Opus at 14%, with 100 tasks costing $84 versus $954.

Why it matters: HKR-H/K/R all pass: the $84 vs $954 cost delta and named cases from Coinbase, Lindy, Harvey, and Cursor give it concrete signal. It is a strong cost-structure commentary, not a major model or product release, so it fits the 72-77 band.

AI HOT (Curated Pool)

Five Labs, Five Minds: Building a Multi-Model Financial Drama Game with Small Models

Thousand Token Wood v2 uses four small models from different labs to drive agents in a financial simulation game, with vLLM 0.22.1’s CUDA toolkit dependency identified as the main serving friction, while a fine-tuned 0.5B Qwen reached 0% self-trading and 100% valid quotes.

Why it matters: HKR-H/K/R all pass: the small-model finance game is a real hook, with vLLM and 0.5B Qwen metrics, plus agent-engineering resonance. Scope remains an experiment, so it sits in low featured.

Jun 6Saturday

AI HOT (Curated Pool)

Google Colab CLI Released

Google released the Colab CLI, which lets developers and AI agents connect local terminals to remote Colab runtimes, request high-performance GPUs, run local Python scripts remotely, and retrieve artifacts such as logs or fine-tuned Gemma 3 adapters.

Why it matters: HKR-H/K/R pass: official Google Colab tooling adds terminal-to-remote-runtime GPU workflows for developers and agents. This is a solid developer product update, not a major model or platform release.

Hacker News front page

Launch HN: General Instinct (YC P26) – Frontier Models on Edge Devices

General Instinct open-sourced InstinctRazor, compressing Qwen3.5-122B-A10B from a roughly 245GB BF16 MoE model into a 48GiB GGUF, with a small-GPU mode that streams experts from system RAM and uses about 7.6–8GB peak VRAM at an 8k context window.

Why it matters: HKR-H/K/R all pass: the 122B-to-8GB edge claim is clickable and backed by memory figures. Source authority is still a YC Launch HN, so it fits featured, not must-write.

Jun 5Friday

Xinzhiyuan · WeChat

Anthropic warns of AI self-acceleration as OpenAI is said to cross a reliability threshold

Xinzhiyuan cites a Yann Dubois interview saying OpenAI crossed a reliability threshold around last December, while Anthropic’s internal data says per-person quarterly code contribution reached 8× the Q1 2024 level by Q2 2026.

Why it matters: HKR-H/K/R all pass: the cliff-edge framing is clickable, and the summary includes a timing claim plus Anthropic’s 8x coding metric. Capped at 82 because this is second-hand interview analysis, not an official release or reproducible test.

Jun 3Wednesday

AI HOT (Curated Pool)

Build 2026: Microsoft tops Google in image generation while catching up on reasoning

Microsoft announced seven in-house AI models at Build 2026, including its first reasoning model, one new tuning method, and one autonomous background AI agent; the RSS snippet does not disclose model names, benchmarks, or release dates.

Why it matters: HKR-H/K/R all pass: Microsoft shipped seven in-house AI models across reasoning, tuning, and a background agent. Model names, benchmark details, and availability are not disclosed, so this stays at the top of 78–84, not P1.

The Verge · AI

Google Must Let Publishers Opt Out of AI Search Features, UK Rules

The UK CMA requires Google to let website owners exclude content from AI Search features, including AI Overviews, and prevent that content from being used for fine-tuning Google’s AI models.

Why it matters: HKR-H/K/R all pass: a UK regulator is forcing Google AI Search opt-outs and fine-tuning restrictions. The article lacks timeline and penalty detail, so it stays in the 78–84 band, not p1.

Synced · WeChat

Understanding SFT Mechanisms in LLMs: Resolving Practice Disputes and Avoiding Wasted Compute

Junpeng Zhang and coauthors argue that SFT on highly homogeneous data has an effective window of only hundreds to about 1,000 training steps, and their interaction-based warning signal detects overfitting before loss gaps appear, saving roughly 30%–50% of training compute.

Why it matters: HKR-H/K/R all pass: the paper gives testable SFT windows, earlier overfitting warnings, and 30%-50% compute savings. It is strong research, not a major model or product release, so it stays below 85.

Jun 2Tuesday

r/LocalLLaMA

I spent months inside verl, forked it, then stopped: internals, fork costs, and an NCCL bug

ReinforcedKnowledge analyzes ByteDance’s verl RLHF loop, covering DataProto plus rollout, reward, advantage, and update paths. The author stopped a private fork because near-daily upstream changes made sync cost exceed refactoring work, and describes an NCCL hang fixed on one node by setting NCCL_SOCKET_IFNAME=lo.

Why it matters: Niche but useful RL post-training field report, not an industry release. HKR-H comes from the fork-then-quit twist; HKR-K has verl’s five paths and NCCL_SOCKET_IFNAME=lo; HKR-R hits the cost of maintaining open-source training forks.

AI HOT (Curated Pool)

NVIDIA Cosmos 3 Tops Open-Weight Image and Video Generation Rankings

NVIDIA Cosmos 3 ranked first in Artificial Analysis’s open-weight text-to-image and image-to-video categories, with 16B Nano and 64B Super variants, and the release includes weights, code, curated datasets, and fine-tuning recipes under the OpenMDW 1.1 license.

Why it matters: HKR-H/K/R all pass: Cosmos 3 leads both Artificial Analysis open-weight image and video charts, with 16B/64B variants and OpenMDW 1.1 artifacts disclosed. Single-source benchmark news keeps it in the 78–84 featured band.

Jun 1Monday

AI HOT (Curated Pool)

OpenBMB Releases Two UltraData Open Datasets, Tops HuggingFace Trending

OpenBMB, Tsinghua NLP, and Modelbest released two UltraData open datasets: Ultra-FineWeb-L3 contains 600B+ tokens, including 400B+ English and 200B+ Chinese tokens, while UltraData-SFT-2605 contains 15M+ SFT samples with thinking and non-thinking labels.

Why it matters: HKR-H/K/R pass: two open datasets, 600B+ tokens, and 15M+ SFT samples are concrete practitioner signal. Single-source release with no evals or license detail keeps it at the lower featured band.

r/LocalLLaMA

I bolted an 8-arm reasoning MoE onto a frozen 1.4B Mamba backbone on a single RTX 3060

The author trained Mamba-Titan-1.4B-Reasoning on a 12GB RTX 3060: a frozen 1.4B Mamba-1 backbone with 8 trainable MoE arms, 2.54B total parameters, Top-2 routing at layers 24/25, and about 50% math accuracy.

Why it matters: HKR-H/K/R all pass via a numbered first-person experiment, but it is a single Reddit post with no independent replication and a fairly technical setup, so it stays in the low featured band.

May 31Sunday

QbitAI · WeChat

Fudan and Tongyi introduce ToolCUA for GUI-Tool path selection in agents

Fudan University and Tongyi Lab introduced ToolCUA-8B, which reaches 46.85% accuracy on OSWorld-MCP after training with about 4k synthetic tools and 180k interleaved GUI-Tool trajectory steps.

Why it matters: HKR-H/K/R all pass: the tool-selection failure hook is concrete, with OSWorld-MCP 46.85% and 180k steps. It stays in the 78–84 band because this is a research release, not a major model or product launch.

May 30Saturday

Synced · WeChat

CUHK Pion optimizer updates LLMs on iso-spectral manifolds to address AdamW and Muon instability

CUHK and collaborators introduced Pion, an optimizer that preserves weight singular values through orthogonal equivalence transformations, and reported that it kept a 60M normalization-free LLaMA-like model stable for 9.6B training tokens while AdamW and Muon collapsed with NaNs.

Why it matters: HKR-H/K/R pass: the hook is AdamW/Muon NaN instability, with a concrete isospectral update and 9.6B-token run. Niche optimizer math keeps it in 78–84, not same-day product news.

May 29Friday

AI HOT (Curated Pool)

Adam's Law: Prompts Written with High-Frequency Words Work Better

FaceMind tested 100 languages and four core tasks, finding that, with semantics unchanged, prompts or fine-tuning text using higher-frequency expressions from pretraining data improves large language model performance.

Why it matters: HKR-H/K/R all pass: the claim is counterintuitive and backed by 100 languages and four task types. Missing models, datasets, and effect sizes keep it in the low featured band.

Synced · WeChat

The Ma Jiaqi Failure Exposed an LLM Issue He Spotted in the Shower a Year Earlier

FaceMind links low-frequency token degradation to two papers: SLoW appeared at EMNLP 2025, Adam's Law was accepted as an ACL 2026 Oral, and high-frequency rewriting raised DeepSeek-V3 math accuracy from 63.55% to 71.54%.

Why it matters: HKR-H/K/R all pass: the odd celebrity-token hook is clickable, and the post gives a mechanism plus a 63.55%→71.54% DeepSeek-V3 result. Practical research signal, but not a major model launch.