Skip to content

27 — Evaluation, Benchmarks & Reproducibility

Driving question: What evidence is sufficient to claim that an automated design technique is useful?

Learning objectives

  • Define research questions and appropriate baselines.
  • Separate internal validity, construct validity, external validity, and conclusion validity.
  • Design reproducible experimental packages.
  • Report negative and failure cases.

Evaluation dimensions

An automated technique may need evaluation along several axes:

Dimension Example measure
detection quality precision, recall, F1, MAP, top-k recall
transformation validity parse/compile/test success
quality impact metric delta, smell removal, human judgment
developer utility acceptance rate, time saved, task success
efficiency runtime, memory, API/model cost
generalization cross-project / cross-domain performance
robustness sensitivity to seeds, thresholds, versions

Reproducibility package

A strong artifact includes:

  • exact commit/version of source systems;
  • environment specification;
  • dataset acquisition/cleaning scripts;
  • configuration and seeds;
  • raw results;
  • analysis notebook/script;
  • instructions for one-command or minimal-step rerun;
  • known limitations.

Threats to validity

Never hide them in the final paragraph. Let threats shape the experiment from the beginning.

Design / research exercise

Take a research claim such as “Technique X improves maintainability.” Operationalize “maintainability” in three different ways and show how each changes the experiment and its threats.

Suggested reading

  • Empirical software engineering guidelines and replication literature.