sigir-experiments
brycewang-stanford/Awesome-Journal-Skills
This comprehensive guide details the rigorous methodology for designing, executing, and reporting experimental results in Information Retrieval (IR). It covers best practices for selecting appropriate test collections, choosing defensible metrics, performing proper statistical significance testing (e.g., paired t-tests), ensuring baseline fairness, and proactively addressing modern pitfalls like LLM contamination. Essential for publishing high-quality, auditable research in IR/NLP.