Skip to content

Data & training

The training side: datasets, synthetic data, pre- and post-training methods, compute and training cost.

Latest picks

1–20 of 124

Today · Sep 30Wednesday

The Decoder

AMD to acquire world-model startup World Labs for $8.2B

AMD announced an all-stock deal to acquire spatial-intelligence startup World Labs for about $8.2 billion. The deal is expected to close by the end of 2026 and still needs regulatory approval. World Labs founder Fei-Fei Li will join AMD's leadership as executive vice president and chief scientist, reporting directly to CEO Lisa Su and leading frontier research.

Why it matters: Beyond the price and where the team goes, the original lays out the world-model route and the hardware logic behind AMD chasing Nvidia.

Sep 22Tuesday

Anthropic News

Anthropic, WHO and partners use Claude in DRC Ebola outbreak response

Anthropic's Beneficial Deployments and Applied AI teams worked with CEPI, the WHO African Regional Office and INRB to use Claude in the response to the Bundibugyo ebolavirus (BDBV) outbreak in the Democratic Republic of the Congo.

Why it matters: The post discloses how Claude was used in the DRC Ebola outbreak and how timelines changed, a view of AI's limits in public-health emergencies.

Sep 17Thursday

Hacker News front page

I had Gemini train its own replacement for $9

The author paid Gemini 3.1 Pro $9 to label 4,290 Reddit comments with knife brands, models, and steels, then fine-tuned GLiNER large v2.5 on those labels. The resulting model runs locally and hits 0.83 F1 against Gemini's labels after 24 minutes on a Tesla T4. Zero-shot GLiNER scored roughly 0.65 F1. The hardest bug was words_mask: the docs suggest a binary mask, but it's actually a word index; filling it with ones kept loss flat at 70. Five of ten runs produced no usable model—three config failures, two from words_mask. The post doesn't report human accuracy on Gemini's labels, so 0.83 is measured against Gemini, not ground truth.

Why it matters: A hands-on fine-tuning case with real numbers: $9 to distill Gemini labels into a local GLiNER model, lifting F1 from 0.65 to 0.83. Score capped because the domain (knife NER) is narrow, but the method transfers.

Sep 14Monday

Hacker News front page

The AI job market in 2026: who gets hired, what they earn, and which roles are fading

Maksim Ilin pulls together LinkedIn, WEF, Stanford, PwC, and Bain data for a September 2026 snapshot of the AI labor market. AI Engineer is the most-hired role; Research Scientist at a frontier lab is the most prestigious, with median pay around $746K at Anthropic and $1.15M at OpenAI L5. The fastest-growing niches are agentic systems and Forward Deployed Engineering—postings for the latter jumped over 1,000% YoY. Prompt engineer has faded as a job title; the skill remains but dissolved into other roles. AI skills now command a 62% wage premium in the US. Bain projects more than 1.3M AI jobs in the US by 2027 against roughly 645K available workers. In Europe, over half of AI postings sit outside tech departments, and Germany shows seven AI-user roles for every AI-developer role. Entry-level hiring got harder: employment for 22–25-year-olds in AI-exposed occupations now trails the rest by 19%.

Why it matters: A multi-source synthesis of the 2026 AI job market with concrete salary figures and role trends — high reference value for practitioners. Capped at 72 because it's a personal blog aggregating secondary data, not an original institutional report with primary research.

Sep 8Tuesday

Google DeepMind

Google DeepMind releases AlphaGenome Atlas, predicting every single-base variant in the human genome

Google DeepMind released AlphaGenome Atlas, a platform holding effect predictions for 9 billion single-nucleotide variants across the human genome. It spans 1PB, more than 30 times the size of the AlphaFold Database.

Why it matters: The post gives the 9 billion-variant prediction dataset and its AVI scoring, showing what a new tool for interpreting genomic variants looks like.

Sep 4Friday

Computing Life · Share · Yage

Three ledgers to check before self-hosting open models

Lambda engineer Zach Mueller admits his home GPU rack doesn't save money—the return is skill investment. The article uses H1 2026 data to show open models are viable, but self-hosting math is counterintuitive. Three ledgers: cost (cloud API wins for most, two H100s need ~2B tokens/month to break even), data (commercial agreements often suffice), and capability (fine-tuning and hands-on skills are the real payoff). Three tiers from renting tokens to owning hardware, with a two-to-three-week rental test recommended before buying.

Why it matters: Zach Mueller, a Lambda engineer, debunks the self-hosting cost-saving assumption with a concrete framework — HKR all hit. Deduction because this is a commentary roundup, not a primary release, and the body stops at summary level without full cost breakdown details.

Sep 3Thursday

Hugging Face Blog

A 350M model fine-tuned with GRPO in 100 steps lifts structured-output compliance from 22.6% to 29.7%

A hands-on guide from Hugging Face and Liquid AI that fine-tunes LFM2.5-350M with GRPO via the TRL library. Using only 500 samples and 100 training steps on a free Colab GPU, structured-output compliance on the IFStruct benchmark jumps from 22.6% to 29.7%. The post includes the full notebook, reward-function design, and a local evaluation setup with llama.cpp on a MacBook.

Why it matters: A hands-on guide with concrete numbers and a reproducible recipe — hits H and K. But the audience is narrow and R is absent; tutorial content at the featured threshold gets 72.

Aug 28Friday

Hacker News front page

Free, framework-free Colab notebooks for RAG, agents, and evals on the Groq API

calmrocks published a set of Colab notebooks on GitHub for AI engineers and forward-deployed engineers. They cover model APIs, structured output, tool calling, RAG, evals-as-the-spine, agent loops from scratch, tool design, guardrails, MCP, Skills, fine-tuning vs LoRA, prompt injection, LLMOps, and customer craft. Everything runs on the free Groq API with no frameworks. The post doesn't specify the number of notebooks or an update schedule.

Why it matters: A free Colab notebook suite for frontline engineers covering RAG, agents, fine-tuning, and security — framework-free and evals-first, with high practical value. Score held at 72 because it's a solo open-source project without community validation or cross-source discussion yet.

Aug 27Thursday

Anthropic News

Anthropic opens 10,000 Claude seats to researchers, expands AI for Science

Anthropic announced a new Claude team plan for scientists, opening 10,000 seats to researchers worldwide. Standard seats are free; a higher-tier seat with 5x usage limits costs $15 per month for one year. Anthropic says it plans to grow the program beyond 10,000 seats in the coming months.

Why it matters: Anthropic disclosed the free and discounted seat count, application bar and usage caps, so research teams can judge their actual path in.

Aug 6Thursday

Computing Life · Share · Yage

Fine-tuning is back in 2026, but now it's a cost-engineering play

Engineering teams in 2026 are fine-tuning again—not to make models smarter, but to slash inference costs on high-volume narrow tasks. FermiSense fine-tuned Qwen3.5-9B for e-commerce review, cutting cost from tens of dollars to $0.50 per 1k calls. Intercom's customer-support small model hit 73.1% resolution rate at one-fifth the cost of GPT-5.4. On the vision side, a DINOv3 classifier workflow trained a zero-API-cost local classifier with only 839 reviewed samples, reaching AP 0.9731. The article provides a decision matrix: fine-tuning pays off above ~50k daily requests with automatically verifiable outputs; below that, Prompt Caching plus RAG is the better bet. Most vendor-reported high scores lack third-party reproducible test sets, so hybrid routing remains the pragmatic middle ground.

Why it matters: A well-argued engineering trend piece with concrete numbers from FermiSense, Intercom, and a DINOv3 classifier workflow. It earns featured by making a clear, counterintuitive case backed by data. Held at 78 rather than higher because it's a synthesis/observation piece, not a f...

Aug 4Tuesday

Hacker News front page

Soup: fine-tune an 8B model on a 4 GB laptop GPU with one config and one command

Soup wraps LLM fine-tuning into a single-command workflow that runs on as little as 4 GB of VRAM. It uses QLoRA with Unsloth acceleration and supports mainstream 8B models like Llama, Mistral, and Qwen. You prepare a JSONL dataset, write a YAML config, and run `soup run`. A built-in Streamlit UI lets you test the model in the browser. The README doesn't disclose exact training speed or memory numbers, but it emphasizes that everything works on a regular laptop GPU.

Why it matters: A practical tool that lowers QLoRA fine-tuning to 4 GB VRAM and one command. Hits H and K, but it's a solo dev's fresh project with no community validation, so R is absent — lands right at the featured threshold of 72. If real-world benchmarks or community feedback appear late...

Jul 31Friday

AI HOT (Curated Pool)

Inkling-Small released: 276B total params, 12B active, matches original Inkling performance

Thinking Machines open-sourced Inkling-Small: 276B total parameters, 12B active, one-quarter the size of the original Inkling yet matching its performance. Full weights are available. You can fine-tune it on Tinker or chat with it via text, image, and audio in Tinker Playground. The post doesn't disclose specific benchmark scores or comparison details, so I'd hold off on the 'matching performance' claim until independent evals appear.

Why it matters: Thinking Machines released Inkling-Small: 276B total params, 12B active at inference, full weights open for fine-tuning. The compression ratio and open release are strong, but the post doesn't disclose specific benchmark numbers, so the score stays below 80.

Jul 20Monday

Product Hunt · AI

Thinking Machines releases Inkling, a 975B open-weights multimodal model built for fine-tuning

Thinking Machines launched Inkling on Product Hunt, a 975B MoE open-weights model with 41B active parameters and 1M context window. It handles text, images, and audio natively, with controllable reasoning effort, under Apache 2.0. The companion Tinker API handles LoRA fine-tuning without infrastructure overhead, aimed at researchers who want full control over data and algorithms. The post does not disclose benchmark scores or pricing.

Why it matters: Thinking Machines dropped a 975B MoE open model with only 41B active params — inference cost should be low. 1M context + native multimodal is a strong spec sheet. No benchmarks or real-world latency numbers yet, so holding below 85.

Jul 16Thursday

Hacker News front page

Thinking Machines releases Inkling: a 975B-param open-weights MoE model with 41B active params

Thinking Machines open-sourced Inkling, a 975B-total-param, 41B-active MoE model with a 1M-token context window, pretrained on 45T multimodal tokens. It handles text, images, and audio natively, with controllable thinking effort, and is positioned as a customizable base rather than a benchmark leader. A 12B-active Inkling-Small preview was also announced. The team demoed self-fine-tuning on their Tinker platform: the model wrote a training job, ran it, and switched to the new weights in about 27 minutes. Weights are on Hugging Face; fine-tuning is available on Tinker.

Why it matters: Thinking Machines' first open-weights model: 975B MoE, native multimodal, 1M context, pitched as a hackable base rather than a benchmark leader—clear differentiation. Not p1 because it's a debut from a new team with no ecosystem or real-world validation yet; featured is the ri...

Jul 15Wednesday

Hugging Face Blog

Thinking Machines releases Inkling: a 1T-param, natively multimodal open model

Inkling is an open ~1T-param model that natively accepts image, audio, and text inputs with a 1M context window. Trained on 45T multimodal tokens, it uses a MoE architecture with 975B total and 41B active parameters. It includes MTP speculative decoding layers for faster inference and ships in BF16 and NVFP4 variants. Hugging Face provides day-0 support in transformers, SGLang, vLLM, and llama.cpp, covering agentic coding, multimodal vision, and audio tasks.

Why it matters: A new player, Thinking Machines, open-sources a trillion-parameter multimodal MoE model with solid specs (975B/41B activated, 1M context, 45T tokens trained) and MTP speculative decoding. H and K both hit, but R is weak — the team has no name recognition, no emotional anchor. ...

Jul 14Tuesday

Hacker News front page

Thinking Machines argues the future worth building is human-shaped, not AI-automated

Thinking Machines lays out its technical philosophy: AI should extend human will and judgment, not replace them. The post draws on Hayek's distributed knowledge argument and Toyota's return to craftsmanship, arguing models must be as diverse and distributed as the people they serve. No model specs or release timelines are disclosed.

Why it matters: Thinking Machines published a long-form manifesto arguing AI should extend human will, not replace it, backing 'decentralized alignment' with Hayek and a Toyota case study. The argument is substantive and the stance is sharp, but it's ultimately a vision statement with no prod...

Hacker News front page

DoorDash uses LLM juries and multimodal AI to tag food items, beating human accuracy by 20%

DoorDash's catalog has millions of items with wildly inconsistent names, making manual tagging slow and expensive. They built an AI metadata platform that uses multimodal models—text, images, and web search—to infer attributes like spiciness or cuisine type. Instead of human review, an 'LLM jury' of multiple strong models votes independently and aggregates a consensus, lifting annotation accuracy roughly 20% above typical human reviewers. Context-optimization agents iterate prompts in minutes, adding another 20%+ precision gain and speeding up prompt development 10x. They auto-generate training data to fine-tune small models that match frontier LLM quality at 10% of the inference cost. Distributed inference cut backfill time for millions of items from over a month to a few days. The post doesn't disclose which models, latency, or per-item cost.

Why it matters: DoorDash engineering blog shares a practical LLM-jury + multimodal labeling pipeline with concrete accuracy numbers and auto-prompt-iteration. Useful for AI data pipeline builders, but the food-delivery domain limits audience breadth — R axis missed, so it lands at the feature...

Jun 10Wednesday

r/LocalLLaMA

Fine-tuned Qwen2.5-7B to 96% of Claude Haiku using about $3 of API calls

A Reddit user fine-tuned Qwen2.5-7B with 1,040 DV-DPO preference pairs, costing about $3 at Claude Haiku rates, and reported 96% composite performance versus Claude Haiku on a domain task, with 11-second latency on a 4-bit T4 setup.

Why it matters: HKR-H/K/R all pass: the hook is cheap fine-tuning, with concrete DV-DPO and latency numbers. Kept at 74 because it is a single Reddit post; task scope and eval details are thin.

Jun 9Tuesday

AI HOT (Curated Pool)

Tencent Hunyuan Releases UniRL, a Unified Multimodal RL Infrastructure

Tencent Hunyuan released UniRL, using one post-training loop to cover diffusion and flow-matching models, LLM/VLM systems, and unified multimodal models, while open-sourcing two algorithms, DRPO and Flow-DPPO.

Why it matters: HKR-H/K/R all pass: Tencent Hunyuan names a unified multimodal RL loop and two open-source algorithms. This fits a strong research/open-source infrastructure release, not a flagship model launch, so it stays in the 78–84 band.

Bloomberg Technology

Musk’s xAI Taps Starlink Staffer to Run Grok Training Team

xAI brought in an executive from SpaceX’s Starlink service to run the Grok training team, replacing college-aged engineer Diego Pasini; the RSS snippet does not disclose the executive’s name, tenure, or training process.

Why it matters: HKR-H/K/R pass, but this is a training-team leadership change, not a model release or executive-level departure. Bloomberg sourcing and the Diego Pasini detail put it at the featured floor.