8.7 KiB
Executable File
ise_week_2
- Different Configuration Sampling Methods (pages 8-19)
- Configuration Encodings (pages 21-28)
- Single Environment Learning (DaL) (pages 39-46)
- 2.1
- 2.2
- 2.3
DONE Different Configuration Sampling Methods (pages 8-19)
DONE Configuration Encodings (pages 21-28)
DONE Single Environment Learning (DaL) (pages 39-46)
2.1
Configuration Sampling
In general machine learning problem, we dont care where the data comes from , but here we do.
Configuration sampling is used to select representative samples for learning performance models.
-
Types of options:
- Binary (e.g., on/off)
- Numeric (e.g., value ranges)
- Goal: Balance model accuracy with sampling effort.
Binary Sampling Strategies
Option-wise Strategy
- Each binary option is selected at least once in some configuration.
- Minimize other options to reduce unknown interaction effects.
- Size: Linear in the number of binary options.
T-wise Strategy
- Covers all T-wise combinations of options (T ≥ 2).
- Example (2-wise): {001}, {010}, {100}, {111}
- Size: Exponential in T.
Negative Option-wise Strategy
- For each option: one configuration where it is disabled, all others enabled.
- Adds one all-yes configuration.
- Size: Linear.
- Example (3 options): {110}, {101}, {011}, {111}
- you can see the 4th one is an all-yes configuration
Random (Binary)
- Select n configurations randomly.
- Simple but may be less representative.
Difference Between Option-wise and Negative Option-wise Strategies
Both strategies are used for sampling configurations in systems with binary options, but they focus on different aspects of option selection.
Option-wise Strategy
- Goal: Ensure each option is enabled (selected) at least once across configurations.
- For every binary option, create a configuration where it is on.
- Other options are minimized to avoid unknown interactions.
- Focus: Testing the presence of each option.
-
Example (3 options):
- {100} → Option 1 enabled, others off
- {010} → Option 2 enabled, others off
- {001} → Option 3 enabled, others off
Negative Option-wise Strategy
- Goal: Ensure each option is disabled (deselected) at least once.
- For each option, create a configuration where it is off, and all others are on.
- Also includes a configuration where all options are on.
- Focus: Testing the absence of each option.
-
Example (3 options):
- {110} → Option 3 disabled
- {101} → Option 2 disabled
- {011} → Option 1 disabled
- {111} → All options enabled
Comparison Summary
| Feature | Option-wise | Negative Option-wise |
|---|---|---|
| Focus | Presence of each option | Absence of each option |
| What is varied | Each option enabled once | Each option disabled once |
| Other options in config | Typically disabled | Typically enabled |
| Additional config? | Not required | Yes, includes all-on config |
| Use case | Minimal presence testing | Influence of removing options |
Non-Binary (Numeric) Sampling Strategies
One-Factor-At-A-Time (OFAT)
- Assumes no interactions among options.
- Varies one option at a time, others fixed at center values.
- Size: Linear in number of options.
- Example (values = 1,3,5): {333}, {533}, {133}, {353}, {313}, {331}, {335}
Box-Behnken Design (BBD)
- Captures quadratic effects and 2-wise interactions.
- Uses subset of 3^k full factorial (min, center, max).
- Size: Exponential in number of options.
- Example: {111}, {113}, {115}, {131}, {151}, etc.
Central Composite Design (CCD)
-
Combines:
- 2^k factorial points
- 2k axial points at α-distance
- 1 center point
- Captures curvature and interactions.
- Example: 8 full factorial + 6 axial + {333}
Plackett-Burman Design (PBD)
- Focus on main effects, assumes negligible interactions.
- Uses predefined seeds, e.g., PBD(9,3)
- First config from seed, rest by right-shifting seed.
- Uses indices only for values.
- Example: If O = {1,100,1000,10000,100000}, index 3 could mean 1, 1000, 100000
Random (Non-Binary)
- Random selection of numeric configurations.
- Risk of non-uniformity and clustering.
- Can negatively impact learning performance.
Mixed Variable Sampling
Some systems include both binary and non-binary (numeric) configuration options. These are referred to as mixed systems.
- Requires hybrid or combined strategies to ensure representative coverage.
- One approach: Permute over the mixed space by combining possible binary and numeric value combinations.
- This can grow combinatorially, so sampling techniques may be needed to reduce the total number of permutations.
Example
- Non-binary configs: {0.1, 0.4, 5}, {0.2, 0.4, 7}, {0.2, 0.7, 5}
- Binary configs: {1,0}, {1,1}
-
Full mixed permutations:
- {0.1, 0.4, 5, 1, 0}
- {0.1, 0.4, 5, 1, 1}
- {0.2, 0.4, 7, 1, 0}
- {0.2, 0.4, 7, 1, 1}
- {0.2, 0.7, 5, 1, 0}
- {0.2, 0.7, 5, 1, 1}
2.2
Single Environment Learning: DeepPerf
Source: Ha & Zhang, ICSE 2019
DeepPerf is an early approach using deep neural networks (>3 layers) to predict software performance in configurable systems.
-
Designed to address:
- Small data size: Limited measurements available.
- Feature sparsity: Only a few configuration options significantly impact performance.
- Network instability: Tackled with tailored hyperparameter tuning.
Limitation of DeepPerf
- Does not handle sample sparsity, a major issue in configuration performance prediction.
Improved Approach: Divide-and-Learn (DaL)
Source: Gong & Chen, ESEC/FSE 2023
Key Problem: Sample Sparsity
-
Caused by:
- Inherited feature sparsity.
- Small configuration changes leading to drastic performance shifts.
- Not all configurations being valid.
- Training data is sparse due to expensive measurements.
Key Properties of Configuration Landscape
- Intra-division smoothness: Configurations in the same division show smooth performance variations.
- Inter-division sharpness: Cross-division configurations differ significantly, possibly on key options.
Risk: Limited data might lead to overfitting within divisions.
Architecture of DaL
Three Goals:
- Divide the configuration data into meaningful divisions → function ϕ
- Learn a local model for each division → function μ
- Assign new configurations to the correct local model → using ϕ and μ
-
Implementation:
- CART (Decision Tree) is used for dividing.
- DeepPerf models are trained within each division.
- Random Forest is used for classifying unseen configurations into divisions.
Trade-off: Number of Divisions
- More divisions → better at tackling sparsity, but less data per model → risks underfitting.
-
Need to balance:
- Generalizability vs.
- Specialization
Results
- DaL outperforms or matches state-of-the-art in 33 out of 40 cases.
- Achieves up to 1.94× improvement.
- Needs fewer training samples for same accuracy.
- Especially beneficial in complex systems or with more training data.
2.3
Single Environment Learning: Encoding
Source: Gong & Chen, MSR 2022
A study conducted by the lab investigates how different encoding schemes impact the software performance learning pipeline.
Three Common Encoding Schemes
- Label encoding
- Scaled label encoding (e.g., max-min normalization)
- One-hot encoding
Encoding Schemes Explained
Label Encoding
- Converts configuration options into numeric values.
-
Example:
- Configuration: (cache_size, interval, ssl, data_strategy)
- Values: cache_size = (1, 10, 10000), interval = (1–4), ssl = (0, 1), data_strategy = (strategy_1, strategy_2, strategy_3)
- Encoded: (10000, 2, 1, 1) → (2, 1, 1, 1) → data_strategy: (0, 1, 2)
Scaled Label Encoding
- Similar to label encoding but normalizes all values to the range [0, 1].
- Example (10000, 2, 1, 1) becomes (1, 1/3, 1, 0.5)
One-Hot Encoding
- Transforms each categorical value into a binary vector.
- Example: (10000, 2, 1, 1) becomes (0, 0, 1, 0, 1, 0, 0, 0, 1, 0, 1, 0)
Community Debate and Justifications
-
Categorical features (e.g., cache_mode = memory, disk, mixed):
- Label encoding implies false ordering (1, 2, 3)
- One-hot encoding avoids this but may introduce multicollinearity.
-
Numeric options (e.g., cache_size = 1, 10, 10000):
- Label encoding maintains order but struggles with large scale differences.
- Scaled label encoding improves numeric stability but weakens interaction with binary features.
Study Protocol
- Evaluated using 7 learning algorithms across 5 software systems.