Skip to content

#数据/训练

4 today

Apr 30Thursday

Synced · WeChat

After Generalist, Jianlan Luo’s Team Releases LWD for Embodied AI Training

Jianlan Luo’s team and Agibot released LWD, tested on 16 Agibot G1 robots in real settings. LWD Online scored 0.95 across 8 tasks and 0.91 on long-horizon tasks. Its offline-to-online RL uses failures as data; failed trajectories were 34.8% of a 652.5-hour pool.

Why it matters: HKR-H/K/R all pass: LWD has real-robot scale, task counts, success rates, and failure-trajectory share. Robotics is narrower than a foundation-model launch, so it lands at 78, not P1.

QbitAI · WeChat

OpenAI Explains Why GPT-5.5 Keeps Saying “Goblin”

OpenAI says GPT-5.5’s “goblin” habit came from Nerd-persona rewards and training transfer. After GPT-5.1, ChatGPT’s “goblin” use rose 175%; Nerd replies were 2.5% of all replies but 66.7% of goblin mentions. The key issue is reward bias spreading through RL, rollouts, and SFT.

Why it matters: Strong HKR-H/K/R: an odd model-behavior hook, concrete usage stats, and a clear alignment lesson about reward leakage. It is not a major capability release, so it stays in the 78–84 band.

Apr 28Tuesday

Synced · WeChat

ACL 2026: Huawei Taylor Lab Proposes SHAPE, Adding a Reasoning Tax to LLM Inference

Huawei Taylor Lab, Peking University, and Shanghai University of Finance and Economics proposed SHAPE, accepted by ACL 2026, with about 3% average accuracy gain. It uses entropy segmentation, short rollouts for potential estimation, dynamic length discounts, and token-level credit assignment, cutting token use by about 30%. The key mechanism is a reasoning tax: long high-potential late-stage segments are penalized to reduce verbose confirmation loops.

Why it matters: HKR-H/K/R all pass: the paper gives testable gains of about +3% math accuracy and -30% tokens, with concrete mechanisms. It is a strong research item, not a same-day model-launch story.

Synced · WeChat

Open-source medical video understanding system uAI-NEXUS-MedVLM released

United Imaging Intelligence released uAI-NEXUS-MedVLM for medical video understanding, with a CVPR 2026 paper. MedVidBench has 532k video-instruction pairs across 8 medical sources and 8 tasks. Qwen2.5-VL-7B SFT reached 89.4% CVS accuracy; GPT-5.4 scored 16.4%.

Why it matters: HKR-H/K/R all pass: the story has a real-medical-video open-source hook, concrete 530K+ data scale, 8 tasks, and a 89.4% vs 16.4% result. The medical focus keeps it in the 78–84 band.

X · @op7418

Xiaomi open-sources the MiMo-V2.5 model series

Xiaomi open-sourced the MiMo-V2.5 model series under the MIT license for commercial use, retraining, and fine-tuning. It also launched Orbit 100T Token, offering approved AI builders up to 1.6B credits worth 659 yuan. Agent framework teams can apply for free MiMo token access; the post does not disclose model size or benchmark results.

Why it matters: HKR-H/K/R all pass: Xiaomi MiMo-V2.5 open source, MIT terms, and Orbit 100T credits matter to builders. Missing params and benchmarks keep it in the 78–84 band, below P1.

Apr 27Monday

Google DeepMind

Announcing our partnership with the Republic of Korea

Google DeepMind 与韩国科学技术信息通信部(MSIT)宣布建立合作伙伴关系,将在韩国设立 AI Campus,向韩国学术界开放 AlphaEvolve、AlphaGenome、AlphaFold、AI co-scientist 和 WeatherNext 等前沿模型。

Apr 26Sunday

Hacker News front page

DeepSeek-V4 on Day 0: From Fast Inference to Verified RL with SGLang and Miles

SGLang and Miles added day-0 inference and RL support for DeepSeek-V4, covering 1.6T Pro and 284B Flash. The post cites a 1M-token context, FP4 MoE expert weights, 128-token SWA, and 4:1 or 128:1 KV compression. The key systems detail is ShadowRadix coherence across three KV pools and two compression-state pools.

Why it matters: HKR-H/K/R all pass: a DeepSeek-V4 day-0 systems stack, concrete context/compression mechanisms, and clear deployment-cost stakes. The systems depth narrows reach, but no hard-exclusion rule is triggered.

Apr 22Wednesday

r/LocalLLaMA

ServiceNow-AI/SuperApriel-15B-Instruct · Hugging Face

ServiceNow released SuperApriel-15B-Instruct, a single-checkpoint 15B model with 8 deployment presets spanning 1.0× to 10.7× decode throughput at 32K sequence length. It has 48 decoder layers with 4 mixer variants per layer and up to 262K context positions depending on runtime; the key point is that speed-quality tradeoffs and speculative decoding are exposed from the same weights.

Why it matters: A single checkpoint spanning 8 deployment presets with 1.0x-10.7x decode throughput gives strong HKR-H and HKR-K, and the serving tradeoff gives HKR-R. The blast radius is narrower: this is a 15B inference-focused release, not a frontier-lab flagship update, so 76 and featured.

Apr 20Monday

r/LocalLLaMA

Training LoRA adapters for Apple's on-device 3B model on a free Colab T4 and a Mac

The author built a QLoRA pipeline for Apple’s on-device 3B model, cutting training needs from about 24GB to about 1GB RAM and 5GB GPU, enough for a free Colab T4 or a 24GB Mac. The post says A100 LoRA, T4 QLoRA, and Mac QLoRA adapters perform about the same, raising accuracy from about 40% to 75%, or 86% with retrieval; it also reports a confirmed Apple bug that writes a hidden ~160MB cache copy per CLI call, reaching 269GB over ~300 runs.

Why it matters: A named first-person experiment with reproducible memory and accuracy numbers clears HKR-H/K/R and beats routine tutorial posts. The score stays below the 85 band because this is a single Reddit post with limited source authority and a narrow benchmark scope.

r/LocalLLaMA

Actually put Gemma 4 26B to work on something real: extract trading signals from 2,400 earnings calls

A Reddit user fine-tuned Gemma 4 26B on 800 labeled earnings-call transcripts and ran inference on 2,400 transcripts over 3 years on one RTX 4090 in about 14 hours. On 600 out-of-sample transcripts, one signal linked vaguer CFO guidance to about 1.8% sector-relative underperformance over 5 days with IC 0.04. A stronger signal showed 0.85 correlation with sector returns after checks and was discarded as a ghost factor; the key point is factor sanity checks, not the profit claim.

Why it matters: Strong HKR-H/K/R: this is a named first-person experiment with concrete setup, metrics, and a useful negative result. It stays at featured, not P1, because it is one Reddit test rather than a product release or industry-wide event.

Apr 12Sunday

最佳拍档 (BestPartners)

Breaking RLHF scaling bottlenecks: DeepMind raises data efficiency 10x with information-directed exploration

A Google DeepMind team reports that online RLHF plus information-directed exploration on Gemma 9B reaches about 55% win rate with under 20k preference labels, versus about 200k for offline RLHF. The post describes four algorithms—offline, periodic, online, and information-directed exploration; online training uses batches of 64 prompts and 16 sampled responses per prompt, while the ENN head adds under 5% parameters. The key point is methodological, not that RLHF failed; the post also says results use Gemini 1.5 Pro simulated feedback, and the 1000x gain is an extrapolation toward 1M labels.

Why it matters: HKR-H/K/R all pass: the 10x label-efficiency claim is a strong hook, and the post includes concrete setup details. I kept it at 77 because this is a secondary video summary, feedback is simulated with Gemini 1.5 Pro, and the 1000x figure is an extrapolation.

Apr 8Wednesday

QbitAI · WeChat

Free open-source 2B Chinese speech model reproduces Mangzhuang Ren with high-speed tonguetwisters

ModelBest, OpenBMB, and Tsinghua University released VoxCPM 2, a 2B open speech model that supports 9 Chinese dialects, 30 foreign languages, and 48kHz audio. The post says generation often finishes within 1 second, recommends reference audio of at least 5 seconds, and supports denoising, LoRA, and full fine-tuning; the key detail is its tokenizer-free diffusion autoregressive continuous representation design.

Why it matters: This is a substantive open-source speech release, not a thin demo: the post gives 2B, 48kHz, 9 Chinese dialects, 30 languages, ref audio ≥5s, and a tokenizer-free route. HKR-H/K/R all pass, but the event is not large enough for a must-write P1.

Mar 18Wednesday

MIT Technology Review · AI

The Download: The Pentagon's new AI plans, and next-gen nuclear reactors

The Pentagon plans to create secure environments so generative AI companies can train military-specific models on classified data. The post says Anthropic Claude is already used in classified settings, including analyzing targets in Iran; training on surveillance and battlefield reports would embed sensitive intelligence in the models. It also flags waste challenges from next-gen nuclear reactors, but the post does not disclose reactor designs or disposal parameters.

Why it matters: HKR-H/K/R all pass: the defense-classified training angle is strong, and the post gives one concrete mechanism plus a named Claude use case. I keep it at featured-edge because this is a roundup item, not a primary Pentagon or Anthropic disclosure.

MIT Technology Review · AI

The Pentagon plans to let AI companies train models on classified data, defense official says

The Pentagon is discussing secure facilities where AI firms can train military-specific models on classified data. The post says training would follow tests on nonclassified data; the DoD keeps data ownership, and company staff would access it only rarely with clearance. The key issue is leakage: one shared model may resurface classified information across groups with different access levels.

Why it matters: HKR-H lands on the unusual classified-data-training angle; HKR-K lands on concrete guardrails and ownership terms; HKR-R lands on defense procurement and leakage risk. Score stays below 85 because this is a planning-stage report, not a signed program, budget, or deployment.

Mistral AI

Mistral AI launches Forge, an enterprise model-training system

Mistral AI launched Forge, a system for enterprises to build frontier-class AI models on their own proprietary knowledge. It covers pre-training, post-training and reinforcement learning, supports dense and MoE architectures, and handles multimodal input. Models can be trained and governed on in-house infrastructure. Mistral AI has worked with ASML, Ericsson, the European Space Agency and Singapore's DSO National Laboratories to train models on their proprietary data.

Why it matters: It details the staged capabilities and named partners behind enterprise frontier-model training on private data.

Mar 17Tuesday

NVIDIA Blog

GTC spotlights NVIDIA RTX PCs and DGX Spark running latest open models and AI agents locally

NVIDIA used GTC to showcase RTX PCs and DGX Spark for running local AI agents, and announced Nemotron 3 Nano 4B, Nemotron 3 Super 120B, and the open source NemoClaw stack. The post says DGX Spark has 128GB unified memory for models above 120B parameters; Nemotron 3 Super scored 85.6% on PinchBench, and Qwen 3.5 supports a 262,000-token context window. The key signal is local inference for privacy and zero token cost, while the full “latest open models” lineup and pricing are not disclosed in the post.

Why it matters: HKR-H/K/R all pass: the local-agent hook is strong, and the post includes concrete specs and benchmark numbers. I keep it in featured, not higher, because the full model list and pricing are not disclosed and the source is still a vendor launch post.

Mistral AI

Mistral AI partners with NVIDIA to accelerate open frontier models

Mistral AI 以创始成员身份加入 NVIDIA Nemotron Coalition,双方计划联合开发前沿开源 AI 模型,Mistral AI 提供模型架构、多模态能力与微调工具,NVIDIA 提供算力、模型开发工具和合成数据管线。

Mar 12Thursday

NVIDIA Blog

NVIDIA Nemotron 3 Super delivers 5x higher throughput for agentic AI

NVIDIA launched Nemotron 3 Super, a 120B open model with 12B active parameters, and says it delivers up to 5x higher throughput for agentic AI. It has a 1M-token context window and uses hybrid MoE, latent MoE, and multi-token prediction; the post says Blackwell NVFP4 gives up to 4x faster inference than Hopper FP8, with over 10T training tokens disclosed. What matters is that NVIDIA is releasing open weights, training recipes, and RL environments for reproduction and fine-tuning.

Why it matters: This is a solid model-release story with all three HKR signals, led by strong HKR-K: parameter counts, active params, context length, training scale, and Blackwell/Hopper comparison are all concrete. It stays below 85 because the key performance claims come from NVIDIA's own blog

Feb 12Thursday

MIT Technology Review · AI

What’s next for Chinese open-source AI

MIT Technology Review says that after DeepSeek released R1 in January 2025, Chinese firms kept shipping open-weight models near top Western systems; Moonshot AI’s Kimi K2.5 was close to Anthropic Claude Opus on early benchmarks at about one-seventh the price. The post also says Qwen took over 30% of Hugging Face downloads in 2024 and surpassed Meta Llama in cumulative downloads by 2025–2026; the key shift is from a few general models to many fine-tunable, distillable variants.

Why it matters: All three HKR axes pass. This is not a launch, but it offers concrete market signals—~1/7 pricing, Hugging Face download share, and a clear thesis that Chinese open source is moving toward specialized, distillable variants—so it merits featured, not p1.

Jan 31Saturday

MIT Technology Review · AI

Inside the marketplace powering bespoke AI deepfakes of real women

Researchers from Stanford and Indiana University found that on Civitai, 90% of deepfake bounty requests targeted women and 86% asked for custom LoRAs between mid-2023 and late 2024. Bounties paid $0.50 to $5 and nearly 92% were fulfilled; MIT Technology Review confirmed that even after Civitai's May 2025 deepfake ban, many older requests and purchasable outputs remained live. The key point is that the platform hosts tutorials, payment rails, and matching infrastructure, not just user uploads.

Why it matters: HKR-H lands because the story turns abuse into a visible market. HKR-K lands on four concrete stats and a post-ban moderation gap; HKR-R lands on safety and governance anxiety around open image platforms. Strong featured, not p1.

Jan 6Tuesday

NVIDIA Blog

NVIDIA DGX Spark and DGX Station power the latest open-source and frontier models from the desktop

NVIDIA showed at CES that DGX Spark and DGX Station can run 100B to 1T-parameter models locally on deskside systems. The post cites a 35% average llama.cpp speedup, up to 70% NVFP4 compression, 775GB coherent memory on DGX Station, and a 250,000 token/sec pretraining demo. The real signal is the local dev loop: fine-tuning, inference, RAG, coding assistants, and robotics demos all target replacing some cloud iteration with deskside compute.

Why it matters: HKR-H/K/R all pass: the story pairs a strong desktop-scale hook with concrete specs and demo numbers, and it speaks directly to the local-vs-cloud workflow debate. Still, this is an NVIDIA product post and most performance evidence comes from vendor-run demos, so it stays at 75,.

Jan 3Saturday

TechCrunch · AI

How AI is reshaping work and who gets to do it, according to Mercor's CEO

Mercor reached a $10 billion valuation in 3 years and acts as a talent middleman in AI's data boom. The RSS snippet says it connects labs such as OpenAI and Anthropic with former Goldman Sachs, McKinsey, and elite law firm employees, paying up to $200 an hour to provide domain expertise and train models. The real signal is the labor pipeline: experts from automatable fields are helping build these systems; the post does not disclose scale, contract terms, or task allocation.

Why it matters: Featured on HKR-H/K/R: the angle is displaced experts getting paid up to $200/hour to train models, plus a concrete $10B-in-3-years data point. The post does not disclose scale, contract structure, or task allocation, so it stays in the low-featured band.

Sep 9, 2025Tuesday

Mistral AI

Mistral AI closes €1.7B Series C led by ASML

Mistral AI announced a €1.7B Series C at a €11.7B post-money valuation, led by semiconductor equipment maker ASML, with existing investors DST Global, Andreessen Horowitz, Bpifrance, General Catalyst, Index Ventures, Lightspeed and NVIDIA participating.

Why it matters: Mistral AI closed a €1.7B Series C led by ASML, giving readers a read on its funding plans and semiconductor supply chain ties.

Aug 5, 2025Tuesday

OpenAI News

Estimating Worst-Case Frontier Risks of Open-Weight LLMs

OpenAI says malicious fine-tuning tests on gpt-oss informed its decision to release the model. It trained gpt-oss for maximum biorisk with RL plus web browsing, and for cyber risk in an agentic coding CTF setup; the resulting models still underperformed OpenAI o3. The key signal is the evaluation method, because the post does not disclose exact scores, training scale, or release thresholds.

Why it matters: HKR-H/K/R all pass: the malicious-fine-tuning setup is novel, the paper gives two concrete eval environments, and the open-weight release debate is a live nerve. It stays at 80 because the post omits scores, training scale, and release thresholds.

Aug 1, 2025Friday

Jul 23, 2025Wednesday

Jun 16, 2025Monday

OpenAI News

Introducing OpenAI for Government

OpenAI launched OpenAI for Government on June 16, 2025, consolidating its existing US public-sector work under one program for federal, state, and local agencies. Its first partnership is a pilot with the US Department of Defense CDAO under a contract capped at $200 million, offering ChatGPT Enterprise, ChatGPT Gov, secure environments, and limited custom national-security models. The practical signal is deployment: a Pennsylvania pilot reported about 105 minutes saved per employee per day, while the post does not disclose model versions, pricing, or rollout scale.

Why it matters: This is not a model launch, but it is a meaningful OpenAI government push with a DoD pilot capped at $200M and a named 105-min/day productivity claim. HKR-H/K/R all pass, so it clears featured; missing model/version, pricing, and deployment detail keeps it below p1.

Jun 11, 2025Wednesday

Mistral AI

Mistral Compute

Mistral AI 发布 Mistral Compute,提供从裸金属服务器到全托管 PaaS 的私有集成堆栈,涵盖 GPU、编排、API 与产品服务。该服务由 Mistral 自研训练套件支撑,可训练和部署任意 AI 负载,作为 NVIDIA 合作伙伴将提供最新参考架构与数万块 GPU。

Apr 9, 2025Wednesday

OpenAI News

OpenAI Pioneers Program

OpenAI announced the Pioneers Program on April 9, 2025, selecting a handful of startups to build domain-specific evals and custom models for each company’s top three use cases. The program includes public industry evals and reinforcement fine-tuning with OpenAI researchers; the post does not disclose pricing, cohort size, base models, or rollout dates. The key signal is public eval creation, not model specs.

Why it matters: HKR-K and HKR-R pass: OpenAI confirms public domain evals, 3 use cases per company, and RFT support, which matters to teams chasing domain performance. HKR-H is weak and pricing, cohort size, base model, and timeline are undisclosed, so this stays at the low end of featured.

Mar 4, 2025Tuesday

OpenAI News

Introducing NextGenAI: A consortium to advance research and education with AI

OpenAI launched NextGenAI and committed $50M in grants, compute funding, and API access to support 15 research institutions using AI in research and education. The post lists 16 founding members including OpenAI; MIT can train and fine-tune models, and Oxford’s Bodleian Library uses the API to transcribe rare texts. The real signal is not a single product, but OpenAI tying universities, hospitals, and libraries into its tooling stack.

Why it matters: HKR-K is clear: OpenAI says NextGenAI brings $50M plus compute and API access to 15 institutions. HKR-R lands because this is a distribution and talent-pipeline move into academia; HKR-H is weaker since the headline is a generic consortium launch, so this sits at the low end of `

Dec 17, 2024Tuesday

OpenAI News

OpenAI o1 and new tools for developers

OpenAI released o1 in the API, updated the Realtime API, added Preference Fine-Tuning, and shipped beta Go/Java SDKs; o1 is rolling out first to usage tier 5 developers. Disclosed details include 60% fewer reasoning tokens than o1-preview on average, and a 60% GPT-4o audio price cut in Realtime API to $40/1M input and $80/1M output tokens. The key shift is production support for function calling, Structured Outputs, developer messages, vision, and a reasoning_effort parameter; the post is truncated, so some GPT-4o mini realtime pricing details are not disclosed here.

Why it matters: This is a substantive OpenAI developer release: o1 reaches the API with function calling, Structured Outputs, vision, and developer messages, which materially improves production readiness. HKR-H/K/R all pass; the excerpt includes concrete token and pricing data, but later GPT-4o

Oct 3, 2024Thursday

OpenAI News

Introducing canvas, a new way to write and code with ChatGPT

OpenAI launched the canvas beta on October 3, 2024 for ChatGPT Plus and Team users, adding a GPT-4o-based workspace for writing and coding beyond chat. The post says canvas can auto-trigger or open via “use canvas,” supports targeted edits, version restore, and shortcuts like code review and bug fixing. The key signal is model training: across 20+ internal evals, trigger accuracy reached 83% for writing and 94% for coding, targeted edits beat baseline by 18%, and comment accuracy and quality improved by 30% and 16%.

Oct 1, 2024Tuesday

OpenAI News

Introducing vision to the fine-tuning API

OpenAI launched GPT-4o vision fine-tuning on Oct 1, 2024, letting paid-tier developers train with images plus text, starting from as few as 100 images. The post cites Grab improving lane-count accuracy by 20% and speed-limit sign localization by 13%, while Automat raised RPA success from 16.60% to 61.67%. The notable shift is multimodal customization in the main API; the pricing section is truncated, so full price details are not disclosed.

Why it matters: OpenAI shipped a substantive API update: GPT-4o vision fine-tuning with a 100-image floor and named gains from Grab and Automat, so HKR-H/K/R all pass. Scope is strong for builders, but the blast radius is narrower than a flagship model launch, and pricing is incomplete in the ex

OpenAI News

Model Distillation in the API

OpenAI launched an API distillation workflow on October 1, 2024, letting developers use outputs from GPT-4o and o1-preview to fine-tune cheaper models such as GPT-4o mini. The suite includes Stored Completions, Evals in beta, and fine-tuning; setting store:true auto-saves input-output pairs with no added latency, per the post. Pricing includes 2M free GPT-4o mini training tokens per day and 1M for GPT-4o through October 31; Evals are free up to 7 runs per week through year-end if shared with OpenAI.

Aug 20, 2024Tuesday

OpenAI News

Fine-tuning now available for GPT-4o

OpenAI has opened GPT-4o fine-tuning to developers on all paid tiers, with 1M free training tokens per org per day through September 23. Training costs $25 per 1M tokens, and inference costs $3.75 per 1M input tokens and $15 per 1M output tokens on gpt-4o-2024-08-06. The signal for practitioners: partners reported 43.8% on SWE-bench Verified and 71.83% on BIRD-SQL with fine-tuned GPT-4o.

Why it matters: This is a substantive OpenAI developer release with concrete details: temporary free training quota, train/inference prices, base model version, and two benchmark datapoints. HKR-H/K/R all pass, but this is an API capability expansion, not a new frontier-model launch or platform-

Jul 24, 2024Wednesday

OpenAI News

Improving Model Safety Behavior with Rule-Based Rewards

OpenAI said on July 24, 2024 it uses Rule-Based Rewards in the RLHF pipeline to reduce repeated human feedback for safety alignment. The post defines three response types—hard refusal, soft refusal, and comply—and says the method has been part of OpenAI’s safety stack since GPT-4, including GPT-4o mini. The key point is maintainability when policies change; the post excerpt does not disclose quantitative gains.

Why it matters: HKR-H/K/R all pass: explicit rules inside RLHF is a strong hook, and the post adds three response modes plus paper/code. I keep it in the 78–84 band because the excerpt does not disclose effect sizes, baselines, or failure-case detail.