:PROPERTIES: :ID: 9ad3f3f1-55f7-4114-bc8c-17250b6dd25d :END: #+title: ise_week_7 #+filetags: :uni:notes: * DONE Statistical Test Selection and Use (pages 21-67) * 7.1 Comparing Algorithms in Intelligent Software Engineering ** 1. Motivation - Algorithms and configurations vary in performance. - No universally best algorithm: performance depends on the specific problem ("No Free Lunch" theorem). - To determine which algorithm/configuration is suitable, **comparison is essential**. - However, comparison is challenging due to the **stochastic nature** of computational intelligence algorithms. ** 2. Stochastic Behaviour in Algorithms - Sources of randomness: - In the algorithm (e.g., random initial population, stochastic gradient descent, mutation/crossover probabilities). - In data sampling. - Result: Running the same algorithm multiple times on the same problem yields **different results**. - Therefore, comparisons must account for this randomness. ** 3. Handling Stochastic Behaviour - To compare algorithms meaningfully: - Run each algorithm **multiple times** (e.g., 30+ runs) using **different random seeds**. - This helps capture typical performance and reduce reliance on single-run outliers. ** 4. Methods for Comparison *** a. Mean (Average) - Simple and common. - Problems: - Sensitive to **outliers**. - Does not represent **variability** in results. *** b. Mean + Standard Deviation - Adds information about variation. - Still affected by outliers. - Hard to tell whether differences are statistically significant. *** c. Median - More robust to outliers. - Example: - Sorted list: 0.000001, 0.6, 0.62, 0.65, 0.7, 0.75, 0.8, 0.8, 0.81 - Median = 0.7 - Problem: Ignores **variation** in data. *** d. Median + Quartiles - 1st and 3rd quartiles provide information about data spread. - Still doesn't guarantee ability to distinguish between groups. *** e. Statistical Hypothesis Testing - Scientific method to determine if observed differences are **statistically significant**. - Necessary for robust and credible comparison of algorithms. ** 5. Statistical Hypothesis Testing: Process 1. Define what to compare (e.g., accuracy or fitness). 2. Ensure fair comparison: - Equal number of evaluations or explain why not. - Example: Adjust generations to equate computational budget across algorithms. 3. Formulate hypotheses: - **Null hypothesis (H₀)**: No difference between the two groups. - **Alternative hypothesis (H₁)**: A statistically significant difference exists. 4. Select an appropriate test based on data distribution. ** 6. Choosing the Test *** a. Normality Assumption - Many statistical tests assume a **normal distribution** of values. - Visual inspection or tests (e.g., Shapiro-Wilk) can check this. *** b. Parametric vs Non-parametric Tests - Parametric tests (e.g., t-test): - More powerful. - Require assumptions (e.g., normality, homogeneity of variance). - Non-parametric tests (e.g., Wilcoxon, Mann-Whitney): - Safer for non-normal data. - Widely used in stochastic algorithm comparisons. *** c. Paired vs Unpaired Tests - **Paired**: Use when comparing results from same initial conditions. - **Unpaired**: Use when runs are completely independent. ** 7. Test Outputs - Test produces a **statistic** and a **p-value**. - If **p ≤ 0.05**, reject H₀: significant difference exists. - If **p > 0.05**, do not reject H₀: no significant difference found. - Significance level is usually set to **0.05**, corresponding to 95% confidence. - Lower significance (e.g., 0.01) may be used in critical applications. ** 8. Interpreting P-Values - High p-value → Observed difference likely due to chance → **Do not reject H₀**. - Low p-value → Observed difference unlikely due to chance → **Reject H₀**. ** 9. Test Examples (in R) - Two-tailed Wilcoxon Rank-Sum Test (unpaired). - Two-tailed Wilcoxon Signed-Rank Test (paired). ** 10. Multiple Comparisons Problem - Comparing many algorithms or configurations increases the risk of **Type I errors** (false positives). - Correction methods: - Adjust the significance threshold (e.g., Bonferroni correction). - Downside: Conservative → **reduced power** (risk of missing real differences). ** 11. Tests for N Groups - Stronger than multiple pairwise tests with correction. - Common tests: - **Kruskal-Wallis Test**: for unpaired comparisons across groups. - **Friedman Test**: for paired comparisons across groups. *** Post-hoc Analysis - Needed when the global test finds significant differences but doesn't specify **which pairs** differ. - Kruskal-Wallis → Dunn post-hoc test. - Friedman → Nemenyi post-hoc test.