Skip to content

Research & technical reports

Official research posts and technical reports from labs: architectures, training methods, measurement and safety research. Purely academic papers are not collected.

Latest picks

161–180 of 262

May 3Sunday

Xinzhiyuan · WeChat

Stanford Nature Study: AI Designs 16 Phages from Scratch

Stanford and Arc Institute used Evo to design 302 phage genomes; 16 infected, replicated, and lysed E. coli. Evo 2 uses StripedHyena 2 with a 1M-base context; Evo-Φ69 expanded 16–65× in 6 hours. The key issue is biosafety: one capsid protein had no known homolog in existing life.

Why it matters: HKR-H/K/R all pass: AI-made viable phage genomes, concrete 302/16/1M-bp details, and a clear biosecurity nerve. Score stays at 82 because it is still an AI+life-science paper, not a direct AI product or developer workflow update.

QbitAI · WeChat

DeepSeek V4’s biggest omission

DeepSeek V4’s technical report omits Engram while listing mHC, CSA, HCA, Muon, and FP4. Engram was open-sourced by DeepSeek and Peking University in January, inserting lookup modules between Transformer layers 2 and 15; its 27B test raised MMLU by 3.4 and Multi-Query NIAH to 97.0%. The engineering signal is CXL pooling: 8 servers shared a 4TB memory pool with under 5% throughput loss.

Why it matters: HKR-H/K/R all pass: the omitted-Engram angle is clickable, with layer ranges, benchmark deltas, and CXL memory-pool numbers. It is analysis, not the V4 launch itself, so 78–84 fits.

r/LocalLLaMA

Implemented TurboQuant, but results do not fully match the paper

A Reddit user reimplemented TurboQuant and found the PROD variant reached about 95.8% correlation at 4-bit, below the paper’s 99%+ claim. They report degraded attention quality, with about 67% top-1 accuracy in a simple simulation. The key issue is correlation versus ranking preservation in KV cache quantization.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit reproduction, not a formal release. The 95.8% 4-bit correlation and ~67% top-1 result make it a low featured item.

May 2Saturday

Hacker News front page

LLMs Consistently Pick Their Own Resumes Over Human or Other Model Resumes

An arXiv paper finds LLMs favor resumes they generated in controlled hiring-screening experiments. Self-preference bias ranges from 67% to 82%; across 24 occupations, same-model applicants are 23% to 60% more likely to be shortlisted. The key lever is self-recognition, where simple interventions cut bias by over 50%.

Why it matters: HKR-H/K/R all pass: the hiring-bias hook is sharp, the post gives testable rates and conditions, and fairness in AI screening is a real practitioner nerve. Strong research story, but not a platform release, so it stays in the 78-84 band.

May 1Friday

Synced · WeChat

The Evolution of RL: From PPO to MaxRL in LLM Reasoning Training

Jiqizhixin translated Alexander Weers' article on RL algorithms for LLM reasoning from 2024 to 2026. It covers REINFORCE, PPO, GRPO, RLOO, Dr. GRPO, DAPO, CISPO, MaxRL, DPPO, and ScaleRL, comparing critic removal, clipping, normalization, and pass@k goals. The key signal is mechanism choice, not algorithm names.

Why it matters: A strong technical explainer, not a model or paper release. HKR-H comes from the PPO→MaxRL arc, HKR-K from concrete mechanism comparisons, and HKR-R from live RL-recipe choices; the higher technical bar keeps it in low featured.

Synced · WeChat

Researchers Estimate GPT, Claude, and Gemini Parameter Counts Using API Calls

Bojie Li posted IKP on arXiv to estimate parameter counts of 188 LLMs from 27 vendors via black-box API calls. The dataset has 1,400 questions across 7 rarity tiers, fitted on 89 open models with R²=0.917. Debate centers on synthetic data, MoE effects, and a 90% interval of 0.3x to 3x.

Why it matters: HKR-H/K/R all pass: API-only parameter inference is a strong hook, with concrete counts and error bounds. The 0.3–3x CI limits confidence, so this fits 78–84 featured, not P1.

Apr 30Thursday

r/LocalLLaMA

My calculator is a transformer

radarsat1 shows an RPN interpreter compiled into Transformer weights; “2 3 + 2 *” returns 10. The residual stream acts as registers, attention weights are compiler-calculated, while nonlinear MLP logic is still trained. The prototype is 1.1 GB; the key point is calculable attention weights, not a practical calculator.

Why it matters: HKR-H comes from the counterintuitive title; HKR-K has a reproducible input, weight-construction mechanism, and 1.1GB figure. HKR-R is real but niche, so this stays just above featured threshold, below 78.

r/LocalLLaMA

DeepSeek released Thinking with Visual Primitives framework

DeepSeek, Peking University, and Tsinghua released the Thinking with Visual Primitives paper and repository. The framework inserts coordinate points and bounding boxes into chain-of-thought; the post does not disclose benchmark scores.

Why it matters: HKR-H/K/R all pass: the hook is visual primitives inside reasoning, the new fact is point/box CoT plus an open repo, and the audience cares about grounded VLMs. No benchmark scores are disclosed, so it stays at 80, not P1.

Google DeepMind

Google DeepMind announces AI co-clinician medical research program

Google DeepMind announced an AI co-clinician research program, exploring how AI agents can assist patient care under a doctor's clinical supervision. In a blinded evaluation of 98 real primary care queries, the system made no critical errors in 97 cases, and doctors preferred its answers over mainstream evidence synthesis tools. On 140 consultation skills, it matched or beat primary care physicians on 68, but expert physicians were still better overall at spotting red flags and key physical exams.

Why it matters: Google DeepMind published its AI co-clinician research program and a multimodal consultation evaluation, showing where medical agents' abilities currently end.

r/LocalLLaMA

Qwen-Scope: Official Sparse Autoencoders (SAEs) for Qwen 3.5 models

Qwen Team released Qwen-Scope, SAEs for Qwen 3.5 models from 2B to 35B MoE. It maps residual-stream features across all layers, including Feature #6159 for Chinese activation. The key point is feature-level debugging and steering; the license discourages removing safety filters.

Why it matters: HKR-H/K/R all pass: official Qwen SAEs are novel, concrete, and useful for interpretability work. This is not a new model release, so it stays in the 78–84 recommendation band.

Synced · WeChat

ACL 2026 Survey: Intrinsic Interpretability Moves LLMs from Post-hoc Analysis to Design

ACL 2026 Main accepted a survey on intrinsic interpretability for LLMs, grouping methods into five design paradigms. It covers functional transparency, concept alignment, decomposable representations, explicit modularization, and latent sparsity induction, with MoE, CBM, and GLU/SwiGLU examples. The key test is whether interpretable parts sit on the model’s computation path, not outside it.

Why it matters: HKR-H/K/R pass: the survey has a clear framing shift, five named mechanisms, and safety/debugging relevance. It is a useful research release, not a model launch or empirical breakthrough.

Synced · WeChat

Alec Radford tests Hassabis’s AGI challenge with a model trained on pre-1931 data

Alec Radford’s team trained 13B talkie on 260B English tokens dated before 1931. They tested surprise on nearly 5,000 historical events and used HumanEval for lower-contamination code evaluation. The key issue is time leakage: the 13B model still has vague post-WWII knowledge.

Why it matters: HKR-H/K/R all pass: the 1930 cutoff is a sharp hook, the post gives 260B tokens and ~5,000 event tests, and the finding targets data leakage. This is strong research, not a model or platform release, so 78–84 fits.

Synced · WeChat

After Generalist, Jianlan Luo’s Team Releases LWD for Embodied AI Training

Jianlan Luo’s team and Agibot released LWD, tested on 16 Agibot G1 robots in real settings. LWD Online scored 0.95 across 8 tasks and 0.91 on long-horizon tasks. Its offline-to-online RL uses failures as data; failed trajectories were 34.8% of a 652.5-hour pool.

Why it matters: HKR-H/K/R all pass: LWD has real-robot scale, task counts, success rates, and failure-trajectory share. Robotics is narrower than a foundation-model launch, so it lands at 78, not P1.

QbitAI · WeChat

NUS and collaborators propose ViF to curb visual hallucination snowballing in multi-agent systems

NUS LV-Lab and collaborators proposed ViF, accepted to ICLR 2026. Across 8 benchmarks, 4 MAS structures, and 10 VLMs, it reports 2.4%–3.8% average gains. ViF replaces text-only passing with visual relay tokens and layered attention redistribution, cutting HS by over 30% on average and nearly 40% in ring topology.

Why it matters: HKR-H/K/R all pass: the hook is concrete, the mechanism and eval grid are disclosed, and hallucination control matters to agent builders. Scope stays research-heavy, so it sits at the featured threshold, not same-day must-write.

QbitAI · WeChat

OpenAI Explains Why GPT-5.5 Keeps Saying “Goblin”

OpenAI says GPT-5.5’s “goblin” habit came from Nerd-persona rewards and training transfer. After GPT-5.1, ChatGPT’s “goblin” use rose 175%; Nerd replies were 2.5% of all replies but 66.7% of goblin mentions. The key issue is reward bias spreading through RL, rollouts, and SFT.

Why it matters: Strong HKR-H/K/R: an odd model-behavior hook, concrete usage stats, and a clear alignment lesson about reward leakage. It is not a major capability release, so it stays in the 78–84 band.

Apr 29Wednesday

Xinzhiyuan · WeChat

Tsinghua AutoSOTA spends about $104K in a week to produce 105 SOTA results

Tsinghua's Fengli Xu team and Beijing Zhongguancun Academy released AutoSOTA, which ran unattended for one week, used about 22B tokens, and produced 105 SOTA results. The system uses eight agents for resource setup, environment fixes, scheduling, idea generation, and audits; each full run averaged 5 hours. The key check is its red-line audit: it forbids changing evaluation scripts and data splits, which decides reproducibility.

Why it matters: HKR-H/K/R all pass: hard numbers, an 8-agent mechanism, and audit constraints make the claim testable. It stays at 84 because this is single-source secondary coverage, not a major model or product release.

Computing Life · Share · Yage

DeepSeek V4 Explained: Engineering Decisions Around Agentic Workloads

DeepSeek V4 targets long-horizon agent tasks with a 1M context. The snippet cites hybrid attention, OPD, Muon, and mHC; the post does not disclose size, data, pricing, or release timing.

Why it matters: HKR-H/K/R all pass: DeepSeek V4, 1M context, and agentic workload engineering create a strong hook with concrete mechanisms. Missing params, data, price, and launch timing keep it at 78, not P1.

X · @dotey

HKUST, NUS, Oxford and others release an 88-page survey on world models

Over 10 universities released an 88-page survey proposing a “capability level × domain law” framework for world models. It reviews 400+ works and reports the best video models pass physical-consistency tests at only 26.2%. The key L3 case is A-Lab: 353 closed-loop experiments in 17 days, yielding 36 compounds.

Why it matters: HKR-H/K/R all pass: the survey turns “world model” confusion into a testable taxonomy, with 400+ papers, a 26.2% physics-consistency rate, and A-Lab’s 353 trials in 17 days. Not a model launch, so it stays below the 85 band.

X · @OpenAI

A 60-Year-Open Erdős Problem Was Solved With Help From GPT-5.4 Pro

OpenAI says GPT-5.4 Pro helped solve an Erdős problem open for 60 years. The post names Sebastien Bubeck, Ernest Ryu, and Andrew Mayne, but does not disclose the problem name, proof details, or reproducible conditions.

Why it matters: HKR-H and HKR-R pass because an OpenAI model aiding a 60-year Erdős problem is a strong AI-research hook. HKR-K fails: no problem name, proof details, or reproduction conditions are disclosed.

Apr 28Tuesday

QbitAI · WeChat

NTU REI-Bench Tests Vague Human Instructions, With Success Rates Dropping Up to 36.9%

NTU MARS Lab released REI-Bench, a benchmark with 9 ambiguity levels for vague human instructions. Tests used 4 robot planning frameworks and 6 small LLMs; LLaMA3.1-8B+SayCan fell from 57.7% to 46.9% in standard multi-turn context. The key issue is implicit reference resolution, where baseline success dropped 7.4% to 36.9%.

Why it matters: HKR-H/K/R all pass: the 36.9% drop is a strong hook, and the setup gives 9 ambiguity levels, 4 frameworks, and 6 models. This is a solid embodied-AI benchmark, not a major model release, so it fits the 78–84 band.