Skip to content

Data & training

The training side: datasets, synthetic data, pre- and post-training methods, compute and training cost.

Latest picks

101–120 of 124

Apr 20Monday

r/LocalLLaMA

Training LoRA adapters for Apple's on-device 3B model on a free Colab T4 and a Mac

The author built a QLoRA pipeline for Apple’s on-device 3B model, cutting training needs from about 24GB to about 1GB RAM and 5GB GPU, enough for a free Colab T4 or a 24GB Mac. The post says A100 LoRA, T4 QLoRA, and Mac QLoRA adapters perform about the same, raising accuracy from about 40% to 75%, or 86% with retrieval; it also reports a confirmed Apple bug that writes a hidden ~160MB cache copy per CLI call, reaching 269GB over ~300 runs.

Why it matters: A named first-person experiment with reproducible memory and accuracy numbers clears HKR-H/K/R and beats routine tutorial posts. The score stays below the 85 band because this is a single Reddit post with limited source authority and a narrow benchmark scope.

r/LocalLLaMA

Actually put Gemma 4 26B to work on something real: extract trading signals from 2,400 earnings calls

A Reddit user fine-tuned Gemma 4 26B on 800 labeled earnings-call transcripts and ran inference on 2,400 transcripts over 3 years on one RTX 4090 in about 14 hours. On 600 out-of-sample transcripts, one signal linked vaguer CFO guidance to about 1.8% sector-relative underperformance over 5 days with IC 0.04. A stronger signal showed 0.85 correlation with sector returns after checks and was discarded as a ghost factor; the key point is factor sanity checks, not the profit claim.

Why it matters: Strong HKR-H/K/R: this is a named first-person experiment with concrete setup, metrics, and a useful negative result. It stays at featured, not P1, because it is one Reddit test rather than a product release or industry-wide event.

Apr 12Sunday

最佳拍档 (BestPartners)

Breaking RLHF scaling bottlenecks: DeepMind raises data efficiency 10x with information-directed exploration

A Google DeepMind team reports that online RLHF plus information-directed exploration on Gemma 9B reaches about 55% win rate with under 20k preference labels, versus about 200k for offline RLHF. The post describes four algorithms—offline, periodic, online, and information-directed exploration; online training uses batches of 64 prompts and 16 sampled responses per prompt, while the ENN head adds under 5% parameters. The key point is methodological, not that RLHF failed; the post also says results use Gemini 1.5 Pro simulated feedback, and the 1000x gain is an extrapolation toward 1M labels.

Why it matters: HKR-H/K/R all pass: the 10x label-efficiency claim is a strong hook, and the post includes concrete setup details. I kept it at 77 because this is a secondary video summary, feedback is simulated with Gemini 1.5 Pro, and the 1000x figure is an extrapolation.

Apr 8Wednesday

QbitAI · WeChat

Free open-source 2B Chinese speech model reproduces Mangzhuang Ren with high-speed tonguetwisters

ModelBest, OpenBMB, and Tsinghua University released VoxCPM 2, a 2B open speech model that supports 9 Chinese dialects, 30 foreign languages, and 48kHz audio. The post says generation often finishes within 1 second, recommends reference audio of at least 5 seconds, and supports denoising, LoRA, and full fine-tuning; the key detail is its tokenizer-free diffusion autoregressive continuous representation design.

Why it matters: This is a substantive open-source speech release, not a thin demo: the post gives 2B, 48kHz, 9 Chinese dialects, 30 languages, ref audio ≥5s, and a tokenizer-free route. HKR-H/K/R all pass, but the event is not large enough for a must-write P1.

Mar 18Wednesday

MIT Technology Review · AI

The Download: The Pentagon's new AI plans, and next-gen nuclear reactors

The Pentagon plans to create secure environments so generative AI companies can train military-specific models on classified data. The post says Anthropic Claude is already used in classified settings, including analyzing targets in Iran; training on surveillance and battlefield reports would embed sensitive intelligence in the models. It also flags waste challenges from next-gen nuclear reactors, but the post does not disclose reactor designs or disposal parameters.

Why it matters: HKR-H/K/R all pass: the defense-classified training angle is strong, and the post gives one concrete mechanism plus a named Claude use case. I keep it at featured-edge because this is a roundup item, not a primary Pentagon or Anthropic disclosure.

MIT Technology Review · AI

The Pentagon plans to let AI companies train models on classified data, defense official says

The Pentagon is discussing secure facilities where AI firms can train military-specific models on classified data. The post says training would follow tests on nonclassified data; the DoD keeps data ownership, and company staff would access it only rarely with clearance. The key issue is leakage: one shared model may resurface classified information across groups with different access levels.

Why it matters: HKR-H lands on the unusual classified-data-training angle; HKR-K lands on concrete guardrails and ownership terms; HKR-R lands on defense procurement and leakage risk. Score stays below 85 because this is a planning-stage report, not a signed program, budget, or deployment.

Mistral AI

Mistral AI launches Forge, an enterprise model-training system

Mistral AI launched Forge, a system for enterprises to build frontier-class AI models on their own proprietary knowledge. It covers pre-training, post-training and reinforcement learning, supports dense and MoE architectures, and handles multimodal input. Models can be trained and governed on in-house infrastructure. Mistral AI has worked with ASML, Ericsson, the European Space Agency and Singapore's DSO National Laboratories to train models on their proprietary data.

Why it matters: It details the staged capabilities and named partners behind enterprise frontier-model training on private data.

Mar 17Tuesday

NVIDIA Blog

GTC spotlights NVIDIA RTX PCs and DGX Spark running latest open models and AI agents locally

NVIDIA used GTC to showcase RTX PCs and DGX Spark for running local AI agents, and announced Nemotron 3 Nano 4B, Nemotron 3 Super 120B, and the open source NemoClaw stack. The post says DGX Spark has 128GB unified memory for models above 120B parameters; Nemotron 3 Super scored 85.6% on PinchBench, and Qwen 3.5 supports a 262,000-token context window. The key signal is local inference for privacy and zero token cost, while the full “latest open models” lineup and pricing are not disclosed in the post.

Why it matters: HKR-H/K/R all pass: the local-agent hook is strong, and the post includes concrete specs and benchmark numbers. I keep it in featured, not higher, because the full model list and pricing are not disclosed and the source is still a vendor launch post.

Mar 12Thursday

NVIDIA Blog

NVIDIA Nemotron 3 Super delivers 5x higher throughput for agentic AI

NVIDIA launched Nemotron 3 Super, a 120B open model with 12B active parameters, and says it delivers up to 5x higher throughput for agentic AI. It has a 1M-token context window and uses hybrid MoE, latent MoE, and multi-token prediction; the post says Blackwell NVFP4 gives up to 4x faster inference than Hopper FP8, with over 10T training tokens disclosed. What matters is that NVIDIA is releasing open weights, training recipes, and RL environments for reproduction and fine-tuning.

Why it matters: This is a solid model-release story with all three HKR signals, led by strong HKR-K: parameter counts, active params, context length, training scale, and Blackwell/Hopper comparison are all concrete. It stays below 85 because the key performance claims come from NVIDIA's own blog

Feb 12Thursday

MIT Technology Review · AI

What’s next for Chinese open-source AI

MIT Technology Review says that after DeepSeek released R1 in January 2025, Chinese firms kept shipping open-weight models near top Western systems; Moonshot AI’s Kimi K2.5 was close to Anthropic Claude Opus on early benchmarks at about one-seventh the price. The post also says Qwen took over 30% of Hugging Face downloads in 2024 and surpassed Meta Llama in cumulative downloads by 2025–2026; the key shift is from a few general models to many fine-tunable, distillable variants.

Why it matters: All three HKR axes pass. This is not a launch, but it offers concrete market signals—~1/7 pricing, Hugging Face download share, and a clear thesis that Chinese open source is moving toward specialized, distillable variants—so it merits featured, not p1.

Jan 31Saturday

MIT Technology Review · AI

Inside the marketplace powering bespoke AI deepfakes of real women

Researchers from Stanford and Indiana University found that on Civitai, 90% of deepfake bounty requests targeted women and 86% asked for custom LoRAs between mid-2023 and late 2024. Bounties paid $0.50 to $5 and nearly 92% were fulfilled; MIT Technology Review confirmed that even after Civitai's May 2025 deepfake ban, many older requests and purchasable outputs remained live. The key point is that the platform hosts tutorials, payment rails, and matching infrastructure, not just user uploads.

Why it matters: HKR-H lands because the story turns abuse into a visible market. HKR-K lands on four concrete stats and a post-ban moderation gap; HKR-R lands on safety and governance anxiety around open image platforms. Strong featured, not p1.

Jan 6Tuesday

NVIDIA Blog

NVIDIA DGX Spark and DGX Station power the latest open-source and frontier models from the desktop

NVIDIA showed at CES that DGX Spark and DGX Station can run 100B to 1T-parameter models locally on deskside systems. The post cites a 35% average llama.cpp speedup, up to 70% NVFP4 compression, 775GB coherent memory on DGX Station, and a 250,000 token/sec pretraining demo. The real signal is the local dev loop: fine-tuning, inference, RAG, coding assistants, and robotics demos all target replacing some cloud iteration with deskside compute.

Why it matters: HKR-H/K/R all pass: the story pairs a strong desktop-scale hook with concrete specs and demo numbers, and it speaks directly to the local-vs-cloud workflow debate. Still, this is an NVIDIA product post and most performance evidence comes from vendor-run demos, so it stays at 75,.

Jan 3Saturday

TechCrunch · AI

How AI is reshaping work and who gets to do it, according to Mercor's CEO

Mercor reached a $10 billion valuation in 3 years and acts as a talent middleman in AI's data boom. The RSS snippet says it connects labs such as OpenAI and Anthropic with former Goldman Sachs, McKinsey, and elite law firm employees, paying up to $200 an hour to provide domain expertise and train models. The real signal is the labor pipeline: experts from automatable fields are helping build these systems; the post does not disclose scale, contract terms, or task allocation.

Why it matters: Featured on HKR-H/K/R: the angle is displaced experts getting paid up to $200/hour to train models, plus a concrete $10B-in-3-years data point. The post does not disclose scale, contract structure, or task allocation, so it stays in the low-featured band.

Sep 9, 2025Tuesday

Mistral AI

Mistral AI closes €1.7B Series C led by ASML

Mistral AI announced a €1.7B Series C at a €11.7B post-money valuation, led by semiconductor equipment maker ASML, with existing investors DST Global, Andreessen Horowitz, Bpifrance, General Catalyst, Index Ventures, Lightspeed and NVIDIA participating.

Why it matters: Mistral AI closed a €1.7B Series C led by ASML, giving readers a read on its funding plans and semiconductor supply chain ties.

Aug 5, 2025Tuesday

OpenAI News

Estimating Worst-Case Frontier Risks of Open-Weight LLMs

OpenAI says malicious fine-tuning tests on gpt-oss informed its decision to release the model. It trained gpt-oss for maximum biorisk with RL plus web browsing, and for cyber risk in an agentic coding CTF setup; the resulting models still underperformed OpenAI o3. The key signal is the evaluation method, because the post does not disclose exact scores, training scale, or release thresholds.

Why it matters: HKR-H/K/R all pass: the malicious-fine-tuning setup is novel, the paper gives two concrete eval environments, and the open-weight release debate is a live nerve. It stays at 80 because the post omits scores, training scale, and release thresholds.

Jun 16, 2025Monday

OpenAI News

Introducing OpenAI for Government

OpenAI launched OpenAI for Government on June 16, 2025, consolidating its existing US public-sector work under one program for federal, state, and local agencies. Its first partnership is a pilot with the US Department of Defense CDAO under a contract capped at $200 million, offering ChatGPT Enterprise, ChatGPT Gov, secure environments, and limited custom national-security models. The practical signal is deployment: a Pennsylvania pilot reported about 105 minutes saved per employee per day, while the post does not disclose model versions, pricing, or rollout scale.

Why it matters: This is not a model launch, but it is a meaningful OpenAI government push with a DoD pilot capped at $200M and a named 105-min/day productivity claim. HKR-H/K/R all pass, so it clears featured; missing model/version, pricing, and deployment detail keeps it below p1.

Apr 9, 2025Wednesday

OpenAI News

OpenAI Pioneers Program

OpenAI announced the Pioneers Program on April 9, 2025, selecting a handful of startups to build domain-specific evals and custom models for each company’s top three use cases. The program includes public industry evals and reinforcement fine-tuning with OpenAI researchers; the post does not disclose pricing, cohort size, base models, or rollout dates. The key signal is public eval creation, not model specs.

Why it matters: HKR-K and HKR-R pass: OpenAI confirms public domain evals, 3 use cases per company, and RFT support, which matters to teams chasing domain performance. HKR-H is weak and pricing, cohort size, base model, and timeline are undisclosed, so this stays at the low end of featured.

Mar 4, 2025Tuesday

OpenAI News

Introducing NextGenAI: A consortium to advance research and education with AI

OpenAI launched NextGenAI and committed $50M in grants, compute funding, and API access to support 15 research institutions using AI in research and education. The post lists 16 founding members including OpenAI; MIT can train and fine-tune models, and Oxford’s Bodleian Library uses the API to transcribe rare texts. The real signal is not a single product, but OpenAI tying universities, hospitals, and libraries into its tooling stack.

Why it matters: HKR-K is clear: OpenAI says NextGenAI brings $50M plus compute and API access to 15 institutions. HKR-R lands because this is a distribution and talent-pipeline move into academia; HKR-H is weaker since the headline is a generic consortium launch, so this sits at the low end of `

Dec 17, 2024Tuesday

OpenAI News

OpenAI o1 and new tools for developers

OpenAI released o1 in the API, updated the Realtime API, added Preference Fine-Tuning, and shipped beta Go/Java SDKs; o1 is rolling out first to usage tier 5 developers. Disclosed details include 60% fewer reasoning tokens than o1-preview on average, and a 60% GPT-4o audio price cut in Realtime API to $40/1M input and $80/1M output tokens. The key shift is production support for function calling, Structured Outputs, developer messages, vision, and a reasoning_effort parameter; the post is truncated, so some GPT-4o mini realtime pricing details are not disclosed here.

Why it matters: This is a substantive OpenAI developer release: o1 reaches the API with function calling, Structured Outputs, vision, and developer messages, which materially improves production readiness. HKR-H/K/R all pass; the excerpt includes concrete token and pricing data, but later GPT-4o

Oct 3, 2024Thursday

OpenAI News

Introducing canvas, a new way to write and code with ChatGPT

OpenAI launched the canvas beta on October 3, 2024 for ChatGPT Plus and Team users, adding a GPT-4o-based workspace for writing and coding beyond chat. The post says canvas can auto-trigger or open via “use canvas,” supports targeted edits, version restore, and shortcuts like code review and bug fixing. The key signal is model training: across 20+ internal evals, trigger accuracy reached 83% for writing and 94% for coding, targeted edits beat baseline by 18%, and comment accuracy and quality improved by 30% and 16%.