Skip to content

Research & technical reports

Official research posts and technical reports from labs: architectures, training methods, measurement and safety research. Purely academic papers are not collected.

Latest picks

121–140 of 262

May 12Tuesday

Google DeepMind

Google DeepMind publishes Co-Scientist multi-agent research system

Google DeepMind published Co-Scientist research in Nature, introducing a Gemini-based multi-agent AI system that iteratively generates, debates and evolves new hypotheses for complex scientific problems.

Why it matters: The post discloses the system's three-stage collaboration mechanism and deployment cases at several labs, showing how AI takes part in scientific hypothesis generation.

Latent Space

Thinking Machines' Native Interaction Models: TML-Interaction-Small 276B-A12B Advances Realtime Voice

Thinking Machines released TML-Interaction-Small, a 276B-parameter MoE model with 12B active parameters, and the post says it advances realtime voice through 200ms time-aligned microturns, encoder-free early fusion for audio and images under 200ms, and benchmark wins over GPT-Realtime-2 and Gemini 3.1-Flash.

Why it matters: HKR-H/K/R all pass: TML-Interaction-Small gives architecture, active parameters, 200ms interaction, and named rivals. Benchmarks still need replication, but a real-time voice SOTA claim is same-day material.

QbitAI · WeChat

Shanghai AI Lab Study: SFT Generalizes Under Three Conditions

Shanghai AI Lab, Shanghai Jiao Tong University, and USTC tested Long-CoT SFT on Qwen3-14B-Base and found that cross-domain performance recovered and improved after 8 epochs, with generalization conditioned on optimization depth, data quality and structure, and base-model capability.

Why it matters: HKR-H/K/R all pass: the SFT-generalization claim has a clear hook, Qwen3-14B-Base plus an 8-epoch finding, and direct relevance to fine-tuning teams. It lacks deployment impact or full benchmark detail, so it stays in the mid-featured band.

AI HOT (Curated Pool)

What Parameter Golf Taught Us About AI-Assisted Research

OpenAI’s Parameter Golf brought together over 1,000 participants and more than 2,000 submissions to test AI-assisted machine learning research, coding agents, model quantization, and model design under strict parameter constraints.

Why it matters: OpenAI’s Parameter Golf recap clears HKR-H/K/R with a concrete contest, 1,000+ participants, and 2,000+ submissions. It is research/benchmark signal, not a model or product launch, so 78 fits the lower featured band.

r/LocalLLaMA

Prompt caching for RL training: 7.5x speedup on long-prompt, short-response workloads

The author proposes prompt caching for RL training. On Qwen3.5-4B, it reports a 7.5x speedup with 16k-token prompts and 64-token outputs, and the G=8 example with 1000-token prompts and 100-token responses reduces 8800 processed tokens to 1800 unique tokens.

Why it matters: HKR-H/K/R all pass: the angle is novel, and the post gives 16k/64 plus G=8 token-dedup numbers. Kept at 78 because this is a single Reddit post without independent replication or a paper/code artifact disclosed.

May 11Monday

Synced · WeChat

ICML 2026: PRISM Brings Efficient Test-Time Scaling to dLLMs

PRISM raises LLaDA-8B-Instruct on GSM8K from 67.58% to 85.30% by combining hierarchical trajectory search, partial remasking, and self-verified feedback, reducing dLLM test-time scaling cost from O(NT) toward O(N+KT) under a final candidate width K.

Why it matters: HKR-H/K/R all pass: the hook rejects brute-force scaling, the post gives GSM8K and complexity numbers, and it speaks to inference cost. Still an ICML framework paper, not a mainstream product release, so it sits in 78–84.

AI HOT (Curated Pool)

Qwen-Image-2.0 Technical Report

Qwen-Image-2.0 uses a Qwen3-VL condition encoder and multimodal diffusion transformer for image generation and precise editing, with instruction inputs up to 1K tokens and reported gains in multilingual text rendering, layout quality, and human-rated generation and editing tasks.

Why it matters: HKR-H/K/R all pass: Qwen’s flagship image model report gives concrete architecture, 1K-token instruction input, and editing claims. The domestic flagship-model signal lifts it into the must-write band.

AI HOT (Curated Pool)

Older AI Model Outperforms Human Doctors in Emergency Diagnosis

A Science study reports that OpenAI o1 reached a 67% correct or near-correct diagnosis rate on real emergency department data, exceeding doctors at 50-55%, but the study did not cover long-term inpatient data or imaging diagnosis.

Why it matters: HKR-H/K/R all pass: a Science-linked ER benchmark reports o1 at 67% versus doctors at 50-55%. It stays below P1 because it is one diagnostic study and excludes inpatient and imaging settings.

May 10Sunday

Synced · WeChat

A Framework for Mechanic-Aware Iteration in AI Game Generation

CreativeGame makes an agent write a mechanic contract before four code-generation stages, then evaluates iterations with CreativeProxyReward, two hard gates for runtime and static errors, and lineage-aware memory shared within each game evolution tree.

Why it matters: HKR-H/K/R pass, but this is a game-generation research framework without disclosed open-source status, metrics, or production adoption. It fits the 72–77 band rather than a must-write item.

Synced · WeChat

Turing Award Winner Sutton Uses a 1967 Formula to Improve Streaming Reinforcement Learning

Richard Sutton and coauthors proposed Intentional Updates, which derive the step size from the desired output change; Intentional AC approached SAC on MuJoCo under batch=1 streaming training without replay, while each update used about 1/140 of SAC’s FLOPs.

Why it matters: HKR-H/K/R all pass: Sutton's name, Intentional Updates, MuJoCo conditions, and 1/140 SAC FLOPs give it substance. Strong research signal, but less market-moving than a major LLM product release, so it stays in the 78–84 band.

Xinzhiyuan · WeChat

Next-ToBE Targets Short-Sighted Next-Token Prediction in LLMs at ICLR 2026

East China Normal University and Fudan University researchers proposed Next-ToBE, a training objective that keeps standard autoregressive inference while adding a soft target over future-token windows, and the article reports the method ranked best in 35 of 36 experiments across Qwen2.5-Math-1.5B, Qwen2.5-Math-7B, and Llama3.1-8B-Instruct.

Why it matters: HKR-H and HKR-K pass: the mechanism and 35/36 result are specific, and next-token training is a real debate. The item stays near the featured floor because no artifact, reproduction detail, or production claim is disclosed.

QbitAI · WeChat

Zhejiang University introduces AdaMARP, an AI role-playing framework with scene direction

Zhejiang University and Tencent Youtu proposed AdaMARP for immersive role-playing, using a four-channel message format and a scene manager; its data pipeline includes 81 literary works, 20 synthetic themes, and AdaptiveBench with 100 evaluation seeds.

Why it matters: ACL 2026 role-play agent work brings four-channel messaging, a scene manager, and an 81-book dataset, clearing HKR-H/K. Narrow use cases and missing open-source or production evidence keep it at threshold featured.

r/LocalLLaMA

NVIDIA AI Releases Star Elastic: One Checkpoint Contains 30B, 23B, and 12B Reasoning Models

NVIDIA AI released Star Elastic, a single checkpoint that can zero-shot slice 30B, 23B, and 12B reasoning models in BF16, FP8, and NVFP4; when the 23B submodel handles thinking and the 30B model handles final answers, reported accuracy rises 16% and latency drops 1.9× on AIME-2025, GPQA, LiveCodeBench v5, and MMLU-Pro.

Why it matters: HKR-H/K/R all pass: Star Elastic has a concrete mechanism and testable numbers for inference deployment. Its reach is still narrower than a frontier-model release, so it sits in the high-quality featured band.

Computing Life · Share · Yage

How Anthropic Trained Computer Use: Reading Its Data Pipeline Through a Patent

Anthropic’s patent describes the Computer Use training pipeline: it captures user actions, uses a transformer to infer action intent, and applies a stronger model for synthetic expansion, turning raw UI operations into reasoning data.

Why it matters: HKR-H/K/R all pass: the patent angle is clickable, the three-step data pipeline is concrete, and agent builders care. It is analysis, not an official release or reproducible artifact, so 76 fits the featured threshold.

May 9Saturday

QbitAI · WeChat

Why Perfect AI Agents Do Not Exist: Five Design Philosophies and Trade-offs Behind Claude Code

MBZUAI VILA Lab and UCL analyze Claude Code v2.1.88 source code and identify 5 design philosophies, 13 design principles, 7 permission layers, and 5 context-compaction layers behind its production-agent architecture.

Why it matters: All HKR axes pass: the contrarian Claude Code angle is clickable, the v2.1.88 permission/context mechanisms add substance, and agent tradeoffs resonate with builders. It is third-party analysis, not an Anthropic release, so it stays below must-write.

QbitAI · WeChat

Google AI Co-Mathematician Sets FrontierMath Tier 4 SOTA

Google DeepMind released AI Co-Mathematician, an asynchronous agent workspace for math research, and answered 23 of 48 private FrontierMath Tier 4 problems, scoring 48% under 48-hour, no-token-limit conditions versus GPT-5.5 Pro at 39.6%.

Why it matters: HKR-H/K/R all pass: the story has a hard benchmark number and a concrete research hook. No disclosed product access or cross-source cluster, so it stays at the top of 78–84 rather than p1.

Synced · WeChat

OpenAI's Jiayi Weng: Is the Next AI Training Paradigm Beyond Gradients?

OpenAI researcher Jiayi Weng proposes Heuristic Learning: codex gpt-5.4 reached a perfect 864 score on Breakout and generated 342 search trajectories across Atari 57, with updates applied to code, tests, replays, and memory rather than neural-network weights.

Why it matters: HKR-H/K/R all pass: an OpenAI researcher proposes Heuristic Learning with concrete hooks like Breakout 864 and 342 Atari 57 trajectories. This is strong research/commentary signal, not an official model or product release, so it stays in the 78–84 band.

AI HOT (Curated Pool)

EMO: Expert Mixture Models for Emergent Modular Pretraining

AllenAI introduced EMO, a mixture-of-experts model with 14B total parameters and 1B active parameters, trained on 1 trillion tokens and able to use only 12.5% of its experts for specific tasks while retaining near-full-model performance.

Why it matters: HKR-H/K/R all pass, but this is an AllenAI/Hugging Face research release rather than a frontier model launch. The 14B/1B and 12.5% expert-activation claims justify the low featured band.

May 8Friday

Synced · WeChat

ICLR 2026: NVIDIA and Purdue Use an Agentic Loop for Text-to-3D Scene Generation

NVIDIA Cosmos Lab and Purdue University proposed Scenethesis, a language-and-vision agentic framework for text-to-3D scene generation that uses visual grounding, SDF-based physical constraints, and a judge module; experiments report about 72% first-pass success, 91% after self-checking, and collision rate reduction from 6.1% to 0.8%.

Why it matters: HKR-H/K/R all pass: NVIDIA/Purdue plus an agent loop is clickable, and the post gives SDF constraints, a judge module, and 72%→91% results. Strong research signal, but not a product release, so it stays in 78–84.

AI HOT (Curated Pool)

Adaptive Parallel Reasoning: A New Paradigm for Efficient Reasoning Scaling

BAIR’s post describes adaptive parallel reasoning, where ThreadWeaver and Multiverse dynamically control parallel threads for math and code reasoning; the RSS snippet does not disclose benchmark scores, latency reductions, or reproducible settings.

Why it matters: BAIR authority supports the 72+ band, and HKR-H/K/R all pass. The post names mechanisms and dynamic thread control, but lacks scores, latency gains, and reproducible conditions, so it stays below 78.