UNIT 4: SIMULATION LAB - COMPREHENSIVE SHORT NOTES
4.1 Foundations & Problem Formulation
4.1.1 Definition & Purpose
-
Simulation: The process of designing a computer-based model of a real-world system and conducting experiments with this model to understand system behavior and evaluate alternative policies.
-
Primary Purpose: Conduct "what-if" analysis and system experimentation without disrupting the actual system. Used for systems analysis, performance evaluation, and decision support.
4.1.2 Types of Simulation
| Type | Description | Key Example |
|---|---|---|
| Discrete-Event (DES) | System state changes at discrete points in time (events). Most common for logistics, manufacturing, services. | Bank queue, manufacturing line. |
| Continuous | System state changes continuously over time. | Chemical process, population dynamics. |
| Monte Carlo | Uses random sampling from probability distributions to estimate deterministic quantities. | Risk analysis, financial modeling. |
| Stochastic | Contains probabilistic (random) components. | DES with random arrivals/service times. |
| Deterministic | No random inputs; outputs are fully predictable from inputs. | Simple kinematic model. |
| Static | Model does not depend on time; single snapshot. | Monte Carlo integration. |
| Dynamic | Model evolves over time. | All DES and continuous models. |
4.1.3 When to Use Simulation: Advantages & Limitations
-
Advantages:
-
Can model complex, stochastic, dynamic systems where analytical solutions are impossible.
-
Allows safe and low-cost experimentation.
-
Provides detailed performance metrics (utilization, queues, throughput).
-
Excellent for "what-if" and sensitivity analysis.
-
-
Limitations:
-
Model building is an art; requires significant expertise.
-
Can be computationally expensive for large, long-run models.
-
Results are statistical estimates, not exact truths (requires careful output analysis).
-
Verification & Validation (V&V) is challenging and time-consuming.
-
4.1.4 The Simulation Process (Steps in a Simulation Study)
A structured, iterative methodology:
-
Problem Formulation & Objectives: Define the problem, goals, and questions to be answered.
-
Data Collection & Input Modeling: Gather data on input processes (arrivals, service times); fit distributions.
-
Model Building & Conceptualization: Translate the real system into a logical model using software constructs.
-
Verification & Validation (V&V): Ensure the model is built correctly (verification) and accurately represents reality (validation).
-
Experimentation & Output Analysis: Design experiments (scenarios, replications), run the model, and statistically analyze outputs.
-
Documentation & Implementation: Document the model, assumptions, and results; present recommendations for implementation.
[!TIP] Exam Focus: Be prepared to define each step and give a one-sentence example. The V&V step (4.5) is a highly frequent exam topic.
4.2 Input Data Analysis & Modeling
4.2.1 Identifying Input Processes
Key stochastic inputs that drive the model:
-
Interarrival Times (for arrival streams).
-
Service Times (for processing/repair).
-
Time Between Failures (for machine breakdowns).
-
Repair Times.
-
Routing Probabilities (for path decisions).
4.2.2 Data Collection
-
Sources: Historical system logs, direct observation, time studies, expert opinion.
-
Considerations: Data quality, quantity, and whether it is stationary (statistical properties constant over time).
4.2.3 Fitting Probability Distributions
-
Identify Distribution Type: Use histograms and Q-Q plots to visually compare data to theoretical distributions (Exponential, Normal, Gamma, Weibull, Poisson, etc.).
-
Goodness-of-Fit Tests: Statistically test if sample data could come from a hypothesized distribution.
- Chi-Square ($$\displaystyle \chi^2 $$) Test: Bins data; compares observed vs. expected frequencies. Requires sufficient sample size per bin.
$$ \chi^2 = \sum_{i=1}^{k} \frac{(O_i - E_i)^2}{E_i} $$
where $$\displaystyle O_i $$ = observed freq., $$\displaystyle E_i $$ = expected freq. in bin $i$. Reject $$\displaystyle H_0 $$ if $$\displaystyle \chi^2 > \chi^2_{\alpha, k-1-p} $$.
* **Kolmogorov-Smirnov (K-S) Test:** Compares **empirical cumulative distribution function (ECDF)** to hypothesized CDF. More powerful for continuous distributions.
$$ D = \max_x |F_n(x) - F(x)| $$
where $$\displaystyle F_n(x) $$ is ECDF, $F(x)$ is hypothesized CDF. Reject $$\displaystyle H_0 $$ if $$\displaystyle D > D_{\alpha} $$.
- Parameter Estimation: For chosen distribution, estimate parameters (e.g., $\lambda$ for Exponential, $\mu, \sigma$ for Normal) using Method of Moments or Maximum Likelihood Estimation (MLE).
4.2.4 Using Empirical Distributions
-
If a theoretical fit is poor or sample size is small, use the raw data directly.
-
Method: Store observed values in a table. Generate a random number $U \sim U(0,1)$, find its position in the cumulative distribution of the data, and interpolate/interview to get a variate.
4.2.5 Input Modeling for Non-Stationary Processes
-
Problem: Arrival or service rates vary systematically with time (e.g., more customers during lunch hour).
-
Solution: Use time-varying (non-homogeneous) Poisson processes or piecewise-constant arrival rates. Define different distributions or rates for different time periods (e.g., by hour of day, day of week).
[!TIP] Common Pitfall: Forcing data to fit a "nice" distribution (like Normal) when empirical or another distribution is better. Always perform and report a GOF test.
4.3 Random Number & Variate Generation
4.3.1 Role of Randomness
Simulation models stochastic systems. Randomness is introduced via random variates from input distributions to mimic real-world uncertainty.
4.3.2 Properties of Good RNGs
-
Uniformity: Numbers should be uniformly distributed on (0,1).
-
Independence: Successive numbers should be statistically independent.
-
Long Period: Sequence should not repeat for a very large number of values.
-
Reproducibility: Same seed should produce same sequence (for debugging).
-
Speed & Portability: Efficient generation across platforms.
4.3.3 Pseudo-Random Number Generators (PRNGs)
- Linear Congruential Generator (LCG): Most common, simple algorithm.
$$ X_{n+1} = (aX_n + c) \mod m $$
where $$\displaystyle X_n $$ is the current integer, $a$ = multiplier, $c$ = increment, $m$ = modulus. The output is $$\displaystyle U_n = X_n / m $$.
* **Period** depends on choice of $a, c, m$. For full period, $c$ and $m$ must be relatively prime, etc.
4.3.4 Testing RNGs for Randomness
-
Runs Test: Checks for too many or too few runs (sequences) of above/below median.
-
Autocorrelation Test: Checks correlation between $$\displaystyle U_n $$ and $$\displaystyle U_{n+k} $$ for lags $k$. Should be near zero.
-
Software: Use built-in tests (e.g., in Arena's Input Analyzer) or dedicated packages (e.g., TestU01).
4.3.5 Transforming U(0,1) to Required Variates
- Inverse Transform Technique: Most intuitive. If $F(x)$ is CDF of desired distribution and $$\displaystyle F^{-1}(u) $$ is its inverse,
$$ X = F^{-1}(U), \quad U \sim U(0,1) $$
* **Works for:** Exponential, Uniform, Weibull, Triangular, Empirical.
* **Example (Exponential, rate $\lambda$):** $$\displaystyle F(x) = 1 - e^{-\lambda x} \Rightarrow X = -\frac{1}{\lambda} \ln(1-U) \approx -\frac{1}{\lambda} \ln(U) $$.
\boxed{X = -\frac{1}{\lambda} \ln(U)}
-
Acceptance-Rejection Technique: Used when inverse CDF is difficult.
-
Find a function $c \cdot g(x)$ that envelops the target PDF $f(x)$.
-
Generate $Y \sim g(y)$ and $U \sim U(0,1)$.
-
If $$\displaystyle U \leq \frac{f(Y)}{c \cdot g(Y)} $$, accept $$\displaystyle X=Y $$; else reject and repeat.
- Efficiency depends on how tight the envelope is (small $c$).
-
-
Convolution Method: Sum of independent random variables.
- Erlang(k, $\lambda$): Sum of $k$ independent Exponential($\lambda$) variates.
$$ X = \sum_{i=1}^{k} \left(-\frac{1}{\lambda} \ln U_i\right) $$
4.3.6 Using Built-in Functions
-
All major simulation software (Arena, Simio, AnyLogic) have built-in RNGs (usually Mersenne Twister) and functions to generate variates (e.g.,
EXPO(mean),NORM(mu, sigma)). -
Control: You can specify and seed random number streams for different model components to ensure reproducibility and enable Variance Reduction Techniques like Common Random Numbers (CRN).
[!TIP] Key Formula: The Inverse Transform for Exponential and Uniform distributions must be memorized. Understand the logic of Acceptance-Rejection (envelope, acceptance condition).
4.4 Model Building: Logic, Entities, Resources & Flow
4.4.1 Core Components of a DES Model
| Component | Description | Example in Software (Arena) |
|---|---|---|
| Entities | Items that flow through the system (customers, parts, jobs). Can be transient (flow through) or permanent (resource). | Create module defines arrival. |
| Attributes | Characteristics of entities (static: type, priority; dynamic: age, time in system). | Assign module sets attribute values. |
| Resources | Facilities that provide service to entities (machines, servers, operators). Have capacity and schedule. | Seize, Release, Resource module. |
| Queues | Places where entities wait for resources. Have discipline (FIFO, LIFO, priority) and capacity. | Implicit in Seize when resource busy; Queue module for custom logic. |
4.4.2 Modeling System Logic
-
Process Flows: Standard sequence of operations.
-
SEIZE → DELAY → RELEASE (basic service).
-
PROCESS module (in Simio/AnyLogic) combines these.
-
-
Routing Decisions: Branching logic based on conditions.
-
DECIDE module (Arena):
2-way(if/else),n-way(based on probability/attribute). -
ROUTE module: Sends entity to a specific station or based on a rule.
-
-
Synchronization: Combining or matching entities.
-
BATCH (Arena): Groups entities (e.g., for a batch oven).
-
MATCH (Arena): Pairs entities from different streams (e.g., mating parts).
-
4.4.3 Modeling Complex Systems
-
Breakdowns & Preventive Maintenance: Model as a resource failure. Use
FAILURE/REPAIRmodules (Arena) or downtime schedules triggered by resource usage (time-based) or number of completions (usage-based). -
Changeovers & Setups: Model as a separate resource (setup crew) or as a delay on a resource after a batch or job change (sequence-dependent setup times).
-
Labor Shifts & Scheduling: Define resource schedules (shift patterns, breaks) using a time table or calendar.
-
Material Handling & Transporters: Use transporter resources (forklifts, AGVs) with routing logic to move entities between locations.
TRANSPORTERmodule (Arena). -
Sub-models & Hierarchical Modeling: Build modular, reusable components (e.g., a "Machine" sub-model with its own logic, breakdowns, setups) and call them from the main model.
[!TIP] Practical Skill: Be able to translate a simple narrative (e.g., "Parts arrive every 5±2 min, are processed on Machine A for 8±3 min, then go to QC with 10% scrap rate") into a flowchart of modules.
4.5 Verification & Validation (V&V)
4.5.1 Verification: "Are we building the model right?"
-
Goal: Ensure the computer program is free of bugs and correctly implements the conceptual model.
-
Techniques:
-
Trace Debugging: Step through the model logic for a few entities manually.
-
Output Verification: Check if outputs make sense for known, simple input cases.
-
Modular Testing: Test each sub-model in isolation before integration.
-
Code Walkthroughs: Have another modeler review the logic.
-
4.5.2 Validation: "Are we building the right model?"
-
Goal: Ensure the conceptual model is an accurate representation of the real-world system for the intended purpose.
-
Techniques:
-
Face Validation: Show model and its logic to domain experts (managers, operators) for review. Do they agree it's realistic?
-
Historical Validation: Compare model output (e.g., average throughput) to actual historical data from the real system. Use statistical tests (t-test) to see if difference is significant.
-
Sensitivity Analysis: Test if model output responds reasonably to changes in inputs (e.g., doubling arrival rate should increase average queue length). Unreasonable response indicates model error.
-
Validation of Input Models: Ensure the fitted input distributions themselves are valid (via GOF tests and face validation).
-
[!TIP] Critical Distinction: Verification is about code correctness (debugging). Validation is about model realism (accuracy). Both are mandatory. Skipping V&V is the #1 reason for simulation project failure.
4.6 Experimentation & Output Analysis
4.6.1 Types of Simulation Experiments
-
Terminating Simulation: Has a natural ending time (e.g., end of a shift, completion of a specific number of jobs). Analysis is straightforward.
-
Steady-State (Non-Terminating) Simulation: Runs indefinitely; interest is in long-run, equilibrium performance (e.g., average server utilization over infinite time). Requires warm-up period elimination.
4.6.2 Warm-Up Period (Initial Transient)
-
Problem: Initial conditions (e.g., empty system) are not representative of steady-state, biasing output averages.
-
Solution: Run the model for an initial warm-up period and discard all data collected during this period. Only analyze data from the steady-state portion.
-
How to Determine Warm-Up Length: Use statistical methods like MSER-5 (Mean Squared Error Reduction) or visual inspection of time-series plots of key outputs (e.g., cumulative average of queue length).
4.6.3 Design of Simulation Experiments
-
One-Factor-at-a-Time (OFAT): Change one input factor while holding others constant. Inefficient and cannot detect interactions.
-
Factorial Designs: Systematically vary multiple factors at different levels to study main effects and interactions. (e.g., 2^k design).
-
Replication Strategy:
-
Number of Replications: More replications → narrower confidence intervals. Use sequential analysis or precision-based stopping (e.g., until CI half-width < 5% of mean).
-
Run Length: For terminating, length = problem horizon. For steady-state, must be long enough after warm-up to get sufficient data.
-
-
Controlling Randomness (Variance Reduction):
- Common Random Numbers (CRN): Use the same random number streams for corresponding random inputs across different system configurations (scenarios). This induces positive correlation between outputs, reducing variance of the difference. Crucial for paired comparisons.
4.6.4 Analyzing Output Data
-
Point Estimates: Sample mean $\bar{Y}$, median, percentiles (e.g., 95th percentile for waiting time).
-
Confidence Interval for Mean (Terminating or Steady-State with Replications):
For $n$ independent replications, each yielding a mean $$\displaystyle \bar{Y}_i $$:
-
Compute overall mean: $$\displaystyle \bar{Y} = \frac{1}{n}\sum_{i=1}^{n} \bar{Y}_i $$
-
Compute sample standard deviation of replication means: $$\displaystyle S = \sqrt{\frac{\sum_{i=1}^{n} (\bar{Y}_i - \bar{Y})^2}{n-1}} $$
-
t-based Confidence Interval:
\boxed{\bar{Y} \pm t_{\alpha/2, n-1} \frac{S}{\sqrt{n}}}
where $$\displaystyle t_{\alpha/2, n-1} $$ is the critical value from t-distribution.
-
-
Comparing System Configurations (Paired-t Test): When using CRN, outputs from the two systems are paired (same RNG streams).
-
Let $$\displaystyle D_i = Y_{1i} - Y_{2i} $$ for replication $i$.
-
Compute $\bar{D}$ and $$\displaystyle S_D $$.
-
Test $$\displaystyle H_0: \mu_D = 0 $$ vs $$\displaystyle H_1: \mu_D \neq 0 $$.
-
CI for difference: $$\displaystyle \bar{D} \pm t_{\alpha/2, n-1} \frac{S_D}{\sqrt{n}} $$. If CI does not contain 0, difference is significant.
-
-
Analyzing Non-Terminating Data: Batch Means Method:
-
After warm-up, run simulation for a long time.
-
Divide the output data (e.g., time-series of queue length) into $k$ non-overlapping batches (e.g., each batch = 1000 time units).
-
Compute the average for each batch, $$\displaystyle \bar{Y}_j $$.
-
Treat these $k$ batch means as approximately independent observations and apply the t-interval formula above ($$\displaystyle n=k $$).
-
-
Key Performance Measures (Outputs):
-
Resource Utilization: % time resource is busy. (Time-persistent average).
-
Queue Lengths & Waiting Times: Average, maximum, percentiles.
-
Throughput: Average number of entities processed per unit time.
-
Cycle Time: Total time an entity spends in the system (from arrival to departure).
-
[!TIP] Exam Core: You must know the t-confidence interval formula and the logic of the paired-t test for comparing two systems. Understand the purpose of warm-up and the batch means method.
4.7 Simulation Software & Practical Implementation (Generic Lab Framework)
4.7.1 Common Packages
-
Arena: Industry-standard, flowchart-based, strong in manufacturing/logistics.
-
AnyLogic: Multi-method (DES, ABS, SD), Java-based, very flexible.
-
Simio: Object-oriented, 3D animation, strong in logistics and healthcare.
-
FlexSim: 3D, strong in material handling and manufacturing.
-
ExtendSim: Icon-based, good for continuous and hybrid models.
4.7.2 Building a Model - Step-by-Step
-
Define Libraries: Place modules/blocks from the template (e.g.,
Basic Process,Advanced Transfer). -
Create Entity Types: Define different kinds of entities (e.g.,
PartA,PartB). -
Define Resources: Create resource sets (e.g.,
Machine1,Operator) with capacities. -
Build Flow: Connect modules in logical order (Create → Process → Decide → Dispose).
-
Set Properties: Double-click modules to define:
-
Process Times: Select distribution (e.g.,
TRIA(5,8,12)for triangular). -
Resource Seizing: Which resource(s) to seize and capacity.
-
Routing: Percentages or rules in
Decidemodules. -
Schedules: Attach time tables to resources for shifts/breaks.
-
-
Define Input Distributions: Use built-in functions or connect to Input Analyzer tool to fit and select distributions.
4.7.3 Running Simulations
-
Run Controls: Set Replication Length (for terminating) or Replication Terminator (e.g., after 10,000 entities). Set Number of Replications (e.g., 30-50 for steady-state).
-
Random Number Streams: Assign different seeds/streams to different stochastic elements (arrivals, service times) to ensure independent randomness. Use CRN streams for scenarios you wish to compare.
4.7.4 Analyzing Results
-
Standard Reports: Most software generates automatic category reports:
-
Resource Report: Utilization, busy time, idle time.
-
Queue Report: Average/maximum length, average wait time.
-
Entity Report: Cycle time (time in system), transfer time.
-
-
Custom Reports & Plots:
-
Use Output Analyzer or Time Series plots to view trends.
-
Create expressions for custom metrics (e.g.,
NQ(Queue1)for current queue length). -
Export tallied data (e.g., all entity cycle times) to Excel/SPSS/R for advanced statistical analysis (histograms, more CIs).
-
-
Animation & 3D Visualization:
-
Purpose: Debugging (see if logic matches reality), communication/presentation to stakeholders.
-
Build: Assign 2D/3D shapes to entities, resources, and locations. Use conveyors, transporter animations.
-
[!TIP] Lab Skill: Be able to build a basic single-server queue model (Create → Seize → Delay → Release → Dispose) in your lab's software within 10 minutes. Know where to find resource utilization and average waiting time in the default output reports.
4.8 Advanced Topics & Applications
4.8.1 Input Modeling for Correlated Data
-
Problem: Successive observations are not independent (e.g., service times today depend on yesterday's workload).
-
Solution: Model using time series (AR(1) process):
$$ X_t = \phi X_{t-1} + \epsilon_t, \quad |\phi| < 1, \quad \epsilon_t \sim \text{IID } N(0, \sigma^2) $$
Generate by: $$\displaystyle X_t = \phi X_{t-1} + \sigma \sqrt{1-\phi^2} \cdot Z_t $$, where $$\displaystyle Z_t \sim N(0,1) $$.
4.8.2 Optimization with Simulation
-
Problem: Find input parameter values (e.g., number of servers, buffer sizes) that optimize an objective (minimize cost, maximize throughput).
-
Solution: Use simulation-based optimization tools:
-
Built-in: Arena's OptQuest, Simio's SimRunner.
-
External: Link simulation model to optimization algorithms in R, Python (SciPy), MATLAB.
-
Method: Optimization engine iteratively changes inputs, runs simulation, evaluates objective, and searches for optimum.
-
4.8.3 Simulation of Specific Systems
-
Manufacturing: Job shops, flow lines, assembly lines. Key: routing logic, batch sizes, setups, breakdowns, WIP limits.
-
Service Systems: Call centers, hospitals (ER), banks, airports. Key: staff scheduling, priority rules, appointment systems, balking/reneging.
-
Logistics & Supply Chains: Warehousing, distribution networks, vehicle routing. Key: transporter logic, inventory policies (s,S), demand forecasting.
-
Computer Systems/Networks: CPU scheduling, network traffic, data centers. Key: job priorities, network protocols, buffer sizes.
4.8.4 Integration with Other Tools
-
Excel: Use Excel Add-in (Arena) or VBA/COM to read input parameters from Excel sheets and write results back. Good for DOE (Design of Experiments) setup.
-
Databases (SQL): Connect model to live or historical databases for real-time input data or to store large output datasets.
-
GIS: For spatial models (e.g., disaster response, vehicle routing on maps).
4.8.5 Current Trends
-
Agent-Based Modeling (ABM): Models autonomous agents (people, vehicles) with behaviors and interactions. Good for social systems, crowd dynamics.
-
Hybrid Simulation: Combines DES (for process flow) with System Dynamics (SD) (for aggregate feedback loops) or ABS.
-
Cloud Simulation: Running large-scale simulations on cloud platforms (AWS, Azure) for scalability and collaboration.
4.9 Lab Report Writing & Presentation
4.9.1 Structure of a Professional Simulation Report
-
Executive Summary: Brief overview of problem, method, key findings, and recommendations (1 page max).
-
Problem Statement & Objectives: Clear statement of the real-world problem and specific questions the simulation must answer.
-
Model Conceptualization & Assumptions: Describe the logical model. List all assumptions (justify why reasonable). Include a high-level process flow diagram.
-
Input Data Analysis & Justification: For each key stochastic input:
-
Show raw data (histogram).
-
State distribution fitted (e.g., Weibull).
-
Report GOF test results (Chi-square value, p-value).
-
Give parameter estimates.
-
Justify choice.
-
-
Verification & Validation Procedures: Detail how you verified the code (trace tests) and validated the model (face validation with manager X, historical comparison with Y data, sensitivity tests). This section proves your model is trustworthy.
-
Experimental Design: Describe scenarios compared (Base case, Alternative 1, Alternative 2). State run length, number of replications, warm-up period determination method, use of CRN.
-
Results & Analysis: Present tables and graphs of key outputs (utilization, wait time, throughput). For each comparison, provide point estimates and 95% confidence intervals. State statistical significance of differences (e.g., "System B's average wait time was 4.2 min (95% CI: 3.8, 4.6), significantly lower than System A's 6.1 min (95% CI: 5.7, 6.5)").
-
Conclusions & Recommendations: Directly answer the original objectives. State which system/config is better based on statistical evidence. Provide actionable recommendations for management. Discuss limitations of the study.
4.9.2 Creating Effective Visualizations
-
Use clear, labeled charts (bar charts for comparing means with error bars for CIs, time-series plots for warm-up analysis).
-
Include screenshots of the model's animation to show the setup.
-
Never include raw, unprocessed output dumps from the software.
4.9.3 Presenting Results to Stakeholders
-
Tailor the message: Focus on business impact (cost savings, reduced waiting time, increased throughput), not statistical jargon.
-
Use visuals: Animation is powerful. Show before/after scenarios.
-
Be prepared to defend: Know your assumptions, validation steps, and statistical certainty (confidence intervals).