27 — Evaluation, Benchmarks & Reproducibility¶
Driving question: What evidence is sufficient to claim that an automated design technique is useful?
Learning objectives¶
- Define research questions and appropriate baselines.
- Separate internal validity, construct validity, external validity, and conclusion validity.
- Design reproducible experimental packages.
- Report negative and failure cases.
Evaluation dimensions¶
An automated technique may need evaluation along several axes:
| Dimension | Example measure |
|---|---|
| detection quality | precision, recall, F1, MAP, top-k recall |
| transformation validity | parse/compile/test success |
| quality impact | metric delta, smell removal, human judgment |
| developer utility | acceptance rate, time saved, task success |
| efficiency | runtime, memory, API/model cost |
| generalization | cross-project / cross-domain performance |
| robustness | sensitivity to seeds, thresholds, versions |
Reproducibility package¶
A strong artifact includes:
- exact commit/version of source systems;
- environment specification;
- dataset acquisition/cleaning scripts;
- configuration and seeds;
- raw results;
- analysis notebook/script;
- instructions for one-command or minimal-step rerun;
- known limitations.
Threats to validity¶
Never hide them in the final paragraph. Let threats shape the experiment from the beginning.
Design / research exercise¶
Take a research claim such as “Technique X improves maintainability.” Operationalize “maintainability” in three different ways and show how each changes the experiment and its threats.
Suggested reading¶
- Empirical software engineering guidelines and replication literature.