Anthropic used Claude to optimize 36 biomolecular modeling packages, achieving up to 4.1× speedup in four weeks
What happened
Anthropic 的两位研究员带着 Claude,在不到四周内重构了 30 多个开源生物分子模型的代码。他们没让 AI 去替科学家做判断,而是让它干最实在的活:写了一套叫 FlashPairformer 的 GPU 内核,把蛋白质结构预测里最吃资源的三角几何运算合并成高吞吐的流式操作,再给每个模型加上缓存和 CUDA 图重放,清理掉 Python 的...
Coverage
Follow the reports to see the story from different sides.
- Computing Life · Share · YagePickAnthropic used Claude to optimize 36 biomolecular modeling packages, achieving up to 4.1× speedup in four weeks
Two Anthropic researchers with biomodeling expertise but no GPU kernel background spent under four weeks with Claude refactoring 36 open-source biomolecular packages. They built FlashPairformer, a custom GPU kernel that fuses scattered triangle-attention ops into high-throughput streaming, then applied per-model caching and CUDA graph replay. Benchmarked on H100 against a hand-tuned expert baseline, the bitwise-identical exact mode averages 1.6× speedup; the fast mode, which allows noise within the model's own stochastic range, averages 4.1×; the memory-saving big mode averages 3.4×. Exact and fast modes can push memory up to 3×. DockQ acceptable rates stayed at 54–55% across modes, with no systematic accuracy loss. The report draws clear lines: big mode ran a 10,761-token complex at TM-score 0.92–0.997, but on 31k–70k-residue viral capsids the outputs collapsed into dense balls (TM-score 0.08–0.14). The authors attribute this to the model's 768-token training-crop limit, not the optimizations. In protein design, a single Claude instance driving optimized models on one H200 for 24 hours hit a median ipSAE of 0.785, up from 0.749 in the earlier multi-agent campaign, but none of the designs have been wet-lab tested. Code is open-sourced under Apache-2.0 with no ongoing maintenance.