Anthropic shows how Claude Code and the claude-api skill automate eval design and iterative tuning
Anthropic 介绍用 Claude Code 与 claude-api skill 自动化评测设计与 hillclimbing
Anthropic describes building evals and tuning an app round by round with Claude Code's claude-api skill, while checking for eval noise and overfitting. The /claude-api build-eval command walks the user through confirming samples and graders, runs a baseline and reports confidence intervals. /claude-api hillclimb proposes one change per round, checks it against a training set and a held-out test set, and reverts changes that regress or look overfit.
Why it matters: Anthropic turns eval design and round-by-round tuning into a repeatable process, using human confirmation of samples and grading rules plus a held-out test set to cut misjudgment and overfitting risk.