Skip to content

Evaluation Protocol

Baselines

At least one meaningful baseline is required. Examples:

  • original threshold/rule method;
  • random search;
  • no-refactoring system;
  • simple lexical model;
  • previous published method;
  • LLM without static-analysis constraints.

Repetitions

Stochastic methods should be run with multiple seeds. Report distribution or uncertainty, not only the best run.

Unit of analysis

State clearly whether metrics are calculated per class, pattern instance, commit, project, refactoring sequence, or developer task.

Error analysis

Inspect concrete false positives, false negatives, invalid transformations, or rejected recommendations. A method that fails systematically on a recognizable case is scientifically more informative than a single aggregate score.

Threats to validity

Address construct, internal, external, and conclusion validity. For LLM-based work, include model/version volatility and contamination where relevant.