Claude Code: build-eval and hillclimb tune agents without overfitting
Anthropic’s claude-api skill adds /claude-api build-eval (guided eval design in your repo) and /claude-api hillclimb (one-change-per-round tuning with a held-out set to catch overfitting). On an internal support bench, hillclimb lifted search accuracy from 74.4% to 98.9% while cutting cost to about one-fifth.
