Measure whether a skill actually pays off — design and run a controlled with/without benchmark of a target skill across models, then write the result up under `skill-analysis/<skill>/` as a study of one frozen skill revision. Use when asked to "benchmark the X skill", "is this skill worth it", "meas
Measure whether a skill actually pays off — design and run a controlled with/without benchmark of a target skill across models, then write the result up under `skill-analysis/<skill>/` as a study of one frozen skill revision. Use when asked to "benchmark the X skill", "is this skill worth it", "measure / evaluate a skill". On a repeat invocation it re-runs the *same* benchmark already on disk rath