Skip to content

22 — Automated Pattern Detection

Driving question: Can we infer pattern intent from implementation evidence?

Learning objectives

  • Explain structural, behavioral, metric, graph-matching, ML, and LLM-based detection approaches.
  • Distinguish instance detection from intent recognition.
  • Define precision/recall and project-level evaluation correctly.
  • Analyze oracle ambiguity and pattern variants.

Detection paradigms

Automated pattern detection has used several families of techniques:

  1. Rule/template matching over UML-like structures.
  2. Graph matching over class/dependency graphs.
  3. Static/dynamic behavioral signatures.
  4. Metric/feature-based classification.
  5. Machine/deep learning over code representations.
  6. LLM-assisted semantic classification/reasoning.

The oracle problem

A detector needs ground truth, but pattern instances may be undocumented, partial, variant, or disputed. Evaluation should therefore document:

  • who labeled instances;
  • whether intent evidence was available;
  • inter-rater agreement;
  • treatment of partial/variant instances;
  • project leakage between train/test sets.

For a binary detector:

\[ Precision = \frac{TP}{TP+FP}, \quad Recall = \frac{TP}{TP+FN} \]

But high instance-level scores do not automatically imply useful developer assistance.

Design / research exercise

Select one GoF pattern. Design a detector using only structural evidence, then list false-positive structures that satisfy the shape but not the intent. Propose one additional semantic/behavioral feature to reduce them.

Suggested reading

  • Research surveys and primary studies on design pattern detection.