Skip to content

Data & training

The training side: datasets, synthetic data, pre- and post-training methods, compute and training cost.

Latest picks

41–60 of 124

May 28Thursday

QbitAI · WeChat

Behind DeepSeek V4's Chip-Model Co-Design, China's Compute Ecosystem Gains Speed

QbitAI says DeepSeek V4 validated Ascend chip-model co-design, with CANN open-sourcing 65 repositories and supporting day-zero adaptation for more than 70 mainstream models, while AIGCode reported 65% MFU in MoE pretraining on Ascend.

Why it matters: HKR-H/K/R all pass, but this is mainly a compute-ecosystem progress story, not a DeepSeek V4 capability release. Concrete repo, adaptation, and MFU numbers lift it into featured, below must-write.

Synced · WeChat

Chinese pretrained embodied model Wall-OSS-0.5 is open sourced

X Square Robot open sourced Wall-OSS-0.5, a VLA model whose 400k pretraining checkpoint scored above 80 on 4 of 17 real-robot zero-shot tasks, with weights, code, training recipe, ablations, and a DMuon optimizer implementation released.

Why it matters: Clear HKR-H/K/R: a 400k checkpoint and 17 real-robot zero-shot tasks add substance, while “post-training not required” is a sharp hook. X Square Robot is not a top foundation-model lab, so this stays at 79.

AI HOT (Curated Pool)

NVIDIA Releases AI Framework Polar, Raising Codex Benchmark Score by 594.74%

NVIDIA’s research team open-sourced Polar, an agent reinforcement learning framework that connects GRPO training at the model API boundary without rewriting Codex CLI, Claude Code, Qwen Code, or Pi; on Qwen3.5-4B, Polar raised Codex pass@1 on SWE-Bench Verified from 3.8% to 26.4%, while prefix_merging cut training steps from 1,185 to 218.

Why it matters: HKR-H/K/R all pass: NVIDIA open-sourced Polar with a concrete GRPO mechanism and SWE-Bench Verified numbers. This is a strong research/open-source item, not a major model or product release, so it stays in the 78–84 band.

r/LocalLLaMA

I built a 103B-token Usenet corpus from 1980–2013

OwnerByDane released a 103.1B-token Usenet corpus covering 1980–2013, 408M posts, and 18,347 newsgroups, with free 5K-post-per-hierarchy samples and full-corpus licensing available.

Why it matters: HKR-H/K/R all pass: the zero-contamination corpus has a clear hook, concrete scale, and relevance to training-data scarcity. Score is capped by Reddit-only sourcing, licensed full access, and no third-party validation or benchmark results.

May 27Wednesday

Synced · WeChat

AMD paper: FP4 training instability is not caused by insufficient randomness

AMD and Penn State pretrained Llama 3.1-8B with MXFP4 on MI355X native FP4 hardware, achieving 9-10% end-to-end speedup over an FP8 baseline, while the paper identifies Wgrad quantization as the bottleneck that raises token overhead to 26-27% without deterministic Hadamard stabilization.

Why it matters: HKR-H/K/R all pass: a counterintuitive FP4 claim, concrete Llama 3.1-8B numbers, and a cost/hardware nerve. The topic is narrower training-infra research, so it stays in the 78-84 band.

AI HOT (Curated Pool)

AI Builds AI: ModelBest Open-Sources ForgeTrain, a Training Framework Written by AI

ModelBest, Tsinghua University, and OpenBMB open-sourced ForgeTrain, described as the first production-grade LLM training framework written entirely by AI with zero human code, and ModelBest used it to pretrain MiniCPM5-1B on Huawei Ascend chips.

Why it matters: HKR-H/K/R all pass: an open-source training framework, AI-written code, and MiniCPM5-1B pretraining on Ascend give concrete hooks. This is a strong tooling story, not a top-model launch, so 80 fits featured rather than P1.

AI HOT (Curated Pool)

Shipping a Trillion Parameters With a Hub Bucket: Delta Weight Sync in TRL

Hugging Face merged TRL PR 5417 for delta weight sync, sending only changed weights as sparse safetensors via a Hugging Face Bucket; on Qwen3-0.6B, the per-step payload falls from 1.2GB to 20–35MB.

Why it matters: HKR-H/K/R all pass: TRL gets delta weight sync with a concrete sparse-safetensors mechanism and a 1.2GB to 20–35MB example. Scope is training infra, so it stays below must-write.

May 26Tuesday

AI HOT (Curated Pool)

SenseNova-U1 full training code open-sourced for multimodal multitask training

OpenSenseNova released the full SenseNova-U1 training code on GitHub under Apache-2.0, supporting an 8B dense model, an A3B MoE architecture, and multimodal tasks such as text-to-image generation, image editing, interleaved generation, and text-visual understanding.

Why it matters: HKR-H/K/R all pass, but the source is a short official post with no dataset, training budget, or eval results disclosed. The practical value of full training code puts it in the featured band.

Synced · WeChat

Grok keeps updating after xAI disbandment as Musk announces a new model

Elon Musk said the 1.5T-parameter Grok V9-Medium has finished training, will enter reinforcement learning in a few days, and is planned for release in two to three weeks. Grok Build supports up to 8 parallel sub-agents, a 256K-token context window, Plan Mode, Arena Mode, MCP, and ACP.

Why it matters: HKR-H/K/R all pass, but this is a Grok V9-Medium preview before RL and release, with no benchmarked capability yet. That fits a strong model-race/product update at 82, featured but not p1.

r/LocalLLaMA

Update on a 12×32GB SXM V100 Cluster for Local Legal Drafting

A lawyer runs a local legal-drafting pipeline across 16 GPUs, with Qwen3.5-122B-A10B reaching about 50 tok/s on four V100s, while a verifier blocks ungrounded citations, dates, and Bates numbers before any final document is used.

Why it matters: HKR-H/K/R all pass: this is a first-person local-LLM experiment with concrete numbers, not a vendor post. Reddit source limits authority, so it stays at the low featured band rather than p1.

May 25Monday

r/LocalLLaMA

The Financial Times published an article about Heretic

The Financial Times used Heretic to remove guardrails from Meta Llama 3.3 in under 10 minutes; creator Philipp Emanuel Weidmann said the tool has created over 3,500 decensored models and those modified systems have reached 13 million downloads.

Why it matters: HKR-H/K/R all pass: FT reportedly used Heretic to strip Llama 3.3 guardrails in 10 minutes, with 3,500+ uncensored models and 13M downloads. Capped at 82 because the item is a Reddit summary, not the full FT report or reproducible test log.

May 24Sunday

r/LocalLLaMA

BitCPM-CANN: Native 1.58-Bit Large Language Model Training on Ascend NPU

OpenBMB released BitCPM-CANN, a 1.58-bit QAT training stack on Ascend NPU with 0.5B, 1B, 3B, and 8B models trained from scratch, where the 1B to 8B variants retain 95.7%–97.2% of full-precision MiniCPM4 performance across 11 benchmarks.

Why it matters: HKR-H/K/R pass: low-bit native training on Ascend is novel, and the summary gives sizes plus retention rates. Reddit-only sourcing and no throughput or reproduction details keep it at the featured floor.

Synced · WeChat

Meta layoff survivors face a difficult choice

Meta is pushing some post-layoff employees into new roles: some engineering managers are returning to IC work, while some Infra and AI engineers are being reassigned to data labeling; the article cites a manager-to-report ratio shift from 1:8 to 1:50 and says Meta holds a 49% stake in Scale AI.

Why it matters: HKR-H/K/R all pass: the piece has a concrete oddity, numbers, and a job-security nerve. It is still workforce reporting rather than a model launch or executive departure, so it sits in the lower featured band.

May 23Saturday

Mistral AI

Mistral to acquire physics AI company Emmi AI

Mistral AI said it has reached a definitive agreement to acquire Emmi AI, a physics AI pioneer, to strengthen its position as an AI transformation partner for industrial companies. Austria-based Emmi AI works on physics AI and large engineering models that speed up engineering workflows, replace multi-day computations with real-time simulation and build digital twins. Emmi's co-founders and more than 30 researchers and engineers will join Mistral's Science and Applied AI teams in May.

Why it matters: Mistral is buying physics AI company Emmi to add industrial simulation, showing how it extends into engineering and manufacturing.

May 21Thursday

Xinzhiyuan · WeChat

USTC Papers Study Lifelong Learning for LMMs via Multimodal Knowledge Injection

USTC researchers released MMEVOKE and KORE: MMEVOKE contains 9,422 samples across 159 subcategories, while KORE uses knowledge-tree augmentation and null-space constrained fine-tuning to reduce catastrophic forgetting during multimodal knowledge injection.

Why it matters: HKR-K and HKR-R are solid: the post gives dataset size plus a concrete fine-tuning mechanism. It stays in the low featured band because this is paper-level knowledge injection without production evidence or full reproducibility details.

r/LocalLLaMA

HRM 1B

Sapientinc released HRM-Text 1B Base and its training code, and the paper claims competitive performance against 2–7B open models while using 100–900x fewer training tokens and 96–432x less estimated compute, with training on 16 H100 GPUs taking about 46 hours and costing about $1,472.

Why it matters: HKR-H/K/R all pass: HRM-Text 1B has concrete low-cost training numbers and released code. Capped at 80 because this is a Reddit item and the efficiency claim still lacks independent evaluation.

May 19Tuesday

QbitAI · WeChat

JD and CAS IIE Publish Three Papers Defining Self-Taught RLVR

JD and CAS IIE released three Self-Taught RLVR papers covering RLSD, NPO, and CoPD; RLSD reports that 200 training steps on Qwen3-VL-8B-Instruct exceed GRPO at 400 steps across 8 benchmarks.

Why it matters: HKR-H/K/R pass: self-taught RLVR is a clear hook; RLSD reports 8 benchmarks and a 200-vs-400-step GRPO comparison; it hits reasoning fine-tuning cost. Not a top-lab model launch and replication heat is undisclosed, so it stays low featured.

Synced · WeChat

From Selling Tokens to Selling Outcomes: AI Companies Start Taking KPI Risk

Sierra raised $950 million in May at a valuation above $15 billion, while Lingxi says it reached scaled profitability and positive cash flow in 2025; the article uses both companies to frame RaaS as charging for measurable business outcomes rather than tokens or subscriptions.

Why it matters: HKR-H/K/R all pass: the KPI hook is clickable, Sierra’s $950M raise and RaaS pricing add concrete facts, and the angle hits agent monetization. This is strong business-model signal, not a model-release-level event.

AI HOT (Curated Pool)

Cursor releases Composer 2.5, calling it its strongest model yet

Cursor released Composer 2.5, claiming a 10x efficiency gain at comparable capability, with larger training scale, more complex reinforcement-learning environments, and a text-feedback mechanism.

Why it matters: Cursor Composer 2.5 is a substantive model update for a front-line AI coding tool, with HKR-H/K/R from the 10x efficiency and RL-training details. The single social-source summary lacks benchmarks, pricing, and reproducible tests, keeping it in the 78–84 band.

AI HOT (Curated Pool)

Cursor releases Composer 2.5 coding model

Cursor released Composer 2.5, claiming up to 10x higher efficiency on long coding tasks; the model is further trained on Moonshot’s Kimi K2.5 and uses text feedback for 100k-token-scale trajectories.

Why it matters: HKR-H/K/R all pass: Cursor is a core AI coding tool, and Composer 2.5 adds concrete claims around 10x long-task gains and Kimi K2.5 tuning. Limited sourcing and no independent eval keep it in the 78–84 band.