Skip to content

Tools & Datasets

Tool choice depends on language and research question. The list below is intentionally modular.

Parsing and structural analysis

  • Eclipse JDT (Java AST and bindings)
  • JavaParser / Spoon (Java source transformation)
  • srcML (multi-language structural representation)
  • tree-sitter (incremental parsing across many languages)
  • ANTLR (custom grammars and language tooling)

Static analysis and smells

  • PMD
  • SonarQube / Sonar analyzers
  • Designite (language-specific editions)
  • Understand or other commercial code-analysis platforms where licensed

Refactoring / transformation

  • IDE refactoring engines (IntelliJ, Eclipse, Roslyn)
  • Spoon / OpenRewrite for Java transformations
  • Roslyn analyzers/code fixes for .NET
  • LibCST/Bowler-style tools for Python, where appropriate

Search and experimentation

  • DEAP / pymoo / jMetal-family frameworks
  • NetworkX for dependency graphs
  • pandas / SciPy / scikit-learn for analysis and baselines
  • Docker/Podman for environment capture

Subject systems

Prefer mature open-source projects with:

  • reproducible tags/commits;
  • buildable tests;
  • clear license;
  • non-trivial history;
  • enough size/variation for the research question.

Avoid constructing the entire evaluation from toy pattern examples unless the project specifically studies synthetic benchmarks.