Evaluation Protocol¶
Baselines¶
At least one meaningful baseline is required. Examples:
- original threshold/rule method;
- random search;
- no-refactoring system;
- simple lexical model;
- previous published method;
- LLM without static-analysis constraints.
Repetitions¶
Stochastic methods should be run with multiple seeds. Report distribution or uncertainty, not only the best run.
Unit of analysis¶
State clearly whether metrics are calculated per class, pattern instance, commit, project, refactoring sequence, or developer task.
Error analysis¶
Inspect concrete false positives, false negatives, invalid transformations, or rejected recommendations. A method that fails systematically on a recognizable case is scientifically more informative than a single aggregate score.
Threats to validity¶
Address construct, internal, external, and conclusion validity. For LLM-based work, include model/version volatility and contamination where relevant.