UNIT 2: SIMULATION METHODOLOGY & TECHNIQUES
2.1 Input Data Analysis & Modeling
Purpose & Importance: Input data defines the stochastic elements of a model. Poor data leads to unreliable results (GIGO principle). The goal is to characterize real-world uncertainty with appropriate probability distributions.
Data Collection & Screening:
-
Sources: Historical records, field observations, expert elicitation, synthetic data.
-
Screening: Identify and handle outliers, data entry errors, and missing values. Use descriptive statistics (mean, std dev, min/max) and visual plots (box-and-whisker).
Distribution Identification:
-
Visual Methods:
-
Histogram: Shape (unimodal, skewed, symmetric).
-
Q-Q Plot (Quantile-Quantile): Compare data quantiles to theoretical distribution quantiles. Points following a straight line indicate a good fit.
[!TIP] A Q-Q plot against a Normal distribution is a primary tool for checking normality.
-
-
Statistical Goodness-of-Fit (GOF) Tests:
-
Chi-Square ($$\displaystyle \chi^2 $$) Test: Bins data, compares observed vs. expected frequencies. Requires sufficient data per bin (≥5).
-
Kolmogorov-Smirnov (K-S) Test: Compares empirical cumulative distribution function (ECDF) to theoretical CDF. More powerful for continuous distributions, sensitive to median.
-
Anderson-Darling Test: A weighted K-S test, more sensitive to tails. Often preferred for Normality.
-
Hypothesis: $$\displaystyle H_0 $$: Data follows specified distribution. Reject $$\displaystyle H_0 $$ if p-value < $\alpha$ (e.g., 0.05).
-
Parameter Estimation:
-
Method of Moments: Set sample moments equal to distribution moments.
-
Example (Normal): $$\displaystyle \hat{\mu} = \bar{x} $$, $$\displaystyle \hat{\sigma}^2 = s^2 $$.
-
Example (Exponential): $$\displaystyle \hat{\lambda} = 1 / \bar{x} $$.
-
-
Maximum Likelihood Estimation (MLE): Finds parameters maximizing the probability of observed data. Often more efficient.
Multivariate & Time-Series Inputs:
-
Correlation: Must be modeled if inputs are correlated (e.g., inter-arrival time and service time). Use copulas or multivariate distributions.
-
Time-Series: For sequential dependent data (e.g., daily demand). Models include ARIMA (AutoRegressive Integrated Moving Average).
2.2 Random Number & Variate Generation
Foundations of Randomness: A good Pseudo-Random Number (PRN) sequence must be:
-
Uniformly distributed over (0,1).
-
Independent (no serial correlation).
-
Have a long period (before sequence repeats).
-
Be repeatable (reproducible with same seed).
Pseudo-Random Number Generators (PRNGs):
- Linear Congruential Generator (LCG):
$$X_{n+1} = (aX_n + c) \mod m$$
* $$\displaystyle X_0 $$ = seed, $a$ = multiplier, $c$ = increment, $m$ = modulus.
* $$\displaystyle U_n = X_n / m $$ gives uniform (0,1).
* **Period** depends on choice of $a, c, m$. Full-period possible if: $c$ and $m$ relatively prime, $a-1$ divisible by all prime factors of $m$, $a-1$ divisible by 4 if $m$ divisible by 4.
* **Issue:** Can exhibit serial correlation on a lattice.
- Combined Generators (e.g., L'Ecuyer): Combine two or more LCGs to achieve a longer period and better statistical properties.
Testing PRNGs (Empirical Tests):
-
Frequency Test: Check counts in intervals for uniformity ($$\displaystyle \chi^2 $$ test).
-
Runs Test: Check sequences of above/below mean for randomness.
-
Autocorrelation Test: Check correlation between $$\displaystyle U_n $$ and $$\displaystyle U_{n+k} $$.
-
Gap Test: Check distances between occurrences in a specified range.
-
Poker Test: Check frequencies of digit patterns (e.g., "pairs" in 5-digit numbers).
Generating Random Variates:
-
Inverse Transform Technique:
-
Method: If $U \sim \text{Uniform}(0,1)$ and $F$ is CDF of target distribution, then $$\displaystyle X = F^{-1}(U) $$ has the target distribution.
-
Applications:
-
Exponential($\lambda$): $$\displaystyle X = -\frac{1}{\lambda} \ln(1-U) $$
-
Uniform($a,b$): $$\displaystyle X = a + (b-a)U $$
-
Weibull($\alpha,\beta$): $$\displaystyle X = \frac{1}{\beta} (-\ln(1-U))^{1/\alpha} $$
-
-
Limitation: Requires invertible CDF.
-
-
Acceptance-Rejection (A-R) Technique:
-
Method: Generate $Y$ from proposal $g(x)$ and $U$. Accept $Y$ if $U \leq f(Y)/(c g(Y))$, where $f$ is target PDF, $$\displaystyle c \geq \sup_x f(x)/g(x) $$.
-
Efficiency: Probability of acceptance = $1/c$. Goal is small $c$.
-
Applications: Normal (using Exponential proposal), Gamma, Beta.
-
-
Convolution: Sum of independent random variables. Example: Erlang($k,\lambda$) = sum of $k$ i.i.d. Exponential($\lambda$) variables.
-
Special-Purpose Generators:
- Box-Muller Transform (Normal):
$$Z_1 = \sqrt{-2\ln U_1} \cos(2\pi U_2), \quad Z_2 = \sqrt{-2\ln U_1} \sin(2\pi U_2)$$
where $$\displaystyle U_1, U_2 \sim \text{Uniform}(0,1) $$. $$\displaystyle Z_1, Z_2 \sim N(0,1) $$.
2.3 Model Verification & Validation (V&V)
| Verification | Validation |
|---|---|
| "Building the model right" | "Building the right model" |
| Is the model implemented correctly? | Is the model an accurate representation of the real system? |
| Code-oriented, "white-box" | Output-oriented, "black-box" |
| Techniques: | Techniques: |
| • Debugging, trace debugging | • Face Validation: Expert review of logic/outputs |
| • Modular testing | • Historical Validation: Compare output to past system data |
| • Animation walk-through | • Sensitivity Analysis: Test output response to input changes |
| • Calibration: Tune parameters to match real behavior | |
| • Statistical Validation: Compare output distributions (e.g., CI overlap) |
[!TIP] Common Pitfall: Confusing verification (correct code) with validation (correct model). Always verify first.
2.4 Experimentation & Output Analysis
Statistical Nature of Output: Simulation output is a random variable. Must use statistical methods for analysis.
Types of Simulations:
-
Terminating: Natural end point (e.g., one day, one project). Output is a finite set of observations.
-
Non-Terminating (Steady-State): No natural end; interest is in long-run average. Requires handling initial transient (warm-up period).
Output Analysis for Terminating Simulations:
- Single Replication: Use all observations from one run. Point estimator $\bar{Y}$, confidence interval:
$$\bar{Y} \pm t_{\alpha/2, n-1} \frac{S}{\sqrt{n}}$$
where $S$ = sample std dev, $n$ = sample size.
- Multiple Replications: Run $k$ independent replications. Use batch means (treat each replication's mean as one observation) to form CI.
Output Analysis for Steady-State Simulations:
-
Warm-up Period Determination: Must delete initial transient to avoid bias.
-
Welch's Method: Plot moving averages of means from multiple replications; look for stabilization.
-
Graphical: Plot time series of output variable; identify stabilization point.
-
-
Analysis Methods:
-
Replication-Deletion: Run $k$ replications, delete warm-up from each, use replication means as independent observations.
-
Batch Means (Single Long Run): Divide post-warm-up data into $m$ large batches. Treat batch means as approximately independent. CI:
-
$$\bar{Y} \pm t_{\alpha/2, m-1} \frac{S_b}{\sqrt{m}}$$
where $$\displaystyle S_b $$ = std dev of batch means.
* **Key Requirement:** Batches must be large enough for batch means to be nearly independent (often $m \geq 10$).
Comparing System Alternatives (Scenarios):
-
Paired-t Method (with Common Random Numbers - CRN):
-
Use same random number streams (seeds) for each alternative to induce positive correlation.
-
For $k$ replications, compute differences $$\displaystyle D_i = Y_{i,A} - Y_{i,B} $$.
-
CI for difference $$\displaystyle \mu_D $$: $$\displaystyle \bar{D} \pm t_{\alpha/2, k-1} \frac{S_D}{\sqrt{k}} $$.
-
Advantage: Reduces variance, fewer replications needed.
-
-
Independent Replications Method: Run separate replications for each alternative. Use two-sample t-test (assuming equal/unequal variances).
Design of Simulation Experiments:
-
OFAT (One-Factor-At-a-Time): Change one input factor at a time. Inefficient; misses interactions.
-
Factorial Designs: Systematically vary multiple factors. Full factorial tests all combinations. Fractional factorial tests a subset.
-
Response Surface Methodology (RSM): Fit a polynomial (e.g., quadratic) model to response vs. factors to find optimal settings.
2.5 Introduction to Simulation Software & Experiment Management
Role of Simulation Software:
-
General-Purpose Simulators: AnyLogic, Simio, Arena. Provide library of modeling constructs (queues, resources, flows), input processors, experiment managers, animators, output analyzers. Faster development, built-in statistics.
-
General-Purpose Programming: Python (SimPy), R, Java, C++. More flexibility, control, and scalability for complex logic or custom algorithms.
Designing Experiments in Software:
-
Key Parameters to Set:
-
Replication Length: How long each run lasts.
-
Number of Replications: For statistical reliability.
-
Warm-up Period: For steady-state models.
-
Output Collection: Specify what variables/statistics to track (e.g., queue length, resource utilization, time-in-system).
-
Documentation & Reporting:
-
Model Documentation: Assumptions, input distributions & parameters, verification/validation steps, experiment design.
-
Results Reporting: Clear presentation of scenarios compared, statistical methods used (CI, test type), confidence intervals, and practical significance of differences. Include run specifications (seed, length, replications) for reproducibility.
[!TIP] Always report both statistical and practical significance. A difference may be statistically significant but not operationally meaningful.