UNIT 1: Introduction to Probability and Statistics
I. Basic Probability Concepts
Sample Space (S): Set of all possible outcomes of an experiment. Event (A): Subset of sample space.
Set Operations:
| Operation | Symbol | Definition |
|---|---|---|
| Union | $A \cup B$ | Outcomes in A or B or both |
| Intersection | $A \cap B$ | Outcomes in both A and B |
| Complement | $$\displaystyle A^c $$ or $A'$ | Outcomes not in A |
| Difference | $A - B$ | Outcomes in A but not B |
Axioms of Probability (Kolmogorov):
-
$P(A) \ge 0$ for any event $A$.
-
$$\displaystyle P(S) = 1 $$.
-
For mutually exclusive events $$\displaystyle A_1, A_2, \dots $$, $$\displaystyle P\left(\bigcup_{i=1}^\infty A_i\right) = \sum_{i=1}^\infty P(A_i) $$.
Conditional Probability:
$$P(A|B) = \frac{P(A \cap B)}{P(B)}, \quad P(B) > 0$$
[!TIP] Remember: $$\displaystyle P(A \cap B) = P(A|B)P(B) = P(B|A)P(A) $$.
Independence:
Events A and B are independent iff $$\displaystyle P(A \cap B) = P(A)P(B) $$.
[!CAUTION] Independence $\neq$ disjointness. Disjoint events with positive probability are dependent.
Bayes' Theorem:
For partition $$\displaystyle A_1, \dots, A_n $$ of $S$:
$$P(A_i|B) = \frac{P(B|A_i)P(A_i)}{\sum_{j=1}^n P(B|A_j)P(A_j)}$$
Applications: Medical diagnosis, defective item detection, spam filtering.
II. Random Variables
Definition: Function $X: S \to \mathbb{R}$ assigning real numbers to outcomes.
Types:
-
Discrete: Countable set of possible values (e.g., number of heads).
-
Continuous: Uncountable values, described by intervals (e.g., height, time).
Probability Mass Function (pmf): For discrete $X$,
$$p(x) = P(X = x)$$
Properties:
-
$p(x) \ge 0$ for all $x$.
-
$$\displaystyle \sum_{x} p(x) = 1 $$.
Probability Density Function (pdf): For continuous $X$,
$$f(x) \text{ such that } P(a \le X \le b) = \int_a^b f(x)dx$$
Properties:
-
$f(x) \ge 0$.
-
$$\displaystyle \int_{-\infty}^{\infty} f(x)dx = 1 $$.
[!NOTE] For continuous $X$, $$\displaystyle P(X = x) = 0 $$ for any specific $x$.
Cumulative Distribution Function (cdf):
$$F(x) = P(X \le x)$$
Properties:
-
$$\displaystyle F(-\infty) = 0 $$, $$\displaystyle F(\infty) = 1 $$.
-
Non-decreasing.
-
Right-continuous.
-
For discrete: jumps at points; for continuous: smooth.
III. Expectation and Moments
Expectation (Mean):
-
Discrete: $$\displaystyle E(X) = \sum_{x} x \cdot p(x) $$
-
Continuous: $$\displaystyle E(X) = \int_{-\infty}^{\infty} x f(x) dx $$
Variance:
$$\text{Var}(X) = E[(X - \mu)^2] = E(X^2) - [E(X)]^2$$
Standard deviation: $$\displaystyle \sigma = \sqrt{\text{Var}(X)} $$.
Properties of Expectation:
-
Linearity: $$\displaystyle E(aX + b) = aE(X) + b $$.
-
Additivity: $$\displaystyle E(X + Y) = E(X) + E(Y) $$ (no independence required).
-
$$\displaystyle E(g(X)) = \sum g(x)p(x) $$ or $$\displaystyle \int g(x)f(x)dx $$.
Properties of Variance:
-
$$\displaystyle \text{Var}(aX + b) = a^2 \text{Var}(X) $$.
-
$$\displaystyle \text{Var}(X + Y) = \text{Var}(X) + \text{Var}(Y) $$ if $X, Y$ independent.
[!TIP] $$\displaystyle \text{Var}(X + Y) = \text{Var}(X) + \text{Var}(Y) + 2\text{Cov}(X,Y) $$ in general.
Central Moments:
$$\mu_k = E[(X - \mu)^k]$$
-
$$\displaystyle \mu_1 = 0 $$, $$\displaystyle \mu_2 = \text{Var}(X) $$.
-
$$\displaystyle \mu_3 $$: related to skewness $$\displaystyle \gamma_1 = \mu_3 / \sigma^3 $$.
-
$$\displaystyle \mu_4 $$: related to kurtosis $$\displaystyle \gamma_2 = \mu_4 / \sigma^4 $$.
Moment Generating Function (MGF):
$$M_X(t) = E(e^{tX})$$
Properties:
-
$$\displaystyle M_X(0) = 1 $$.
-
$$\displaystyle M_X^{(k)}(0) = E(X^k) $$ (k-th raw moment).
-
If $X, Y$ independent, $$\displaystyle M_{X+Y}(t) = M_X(t) M_Y(t) $$.
Characteristic Function:
$$\varphi_X(t) = E(e^{itX})$$
Always exists. Used for normal, gamma distributions.
IV. Discrete Probability Distributions
A. Binomial Distribution $\text{Bin}(n, p)$
-
pmf: $$\displaystyle P(X = x) = \binom{n}{x} p^x (1-p)^{n-x} $$, $$\displaystyle x = 0,1,\dots,n $$.
-
Mean: $$\displaystyle \mu = np $$.
-
Variance: $$\displaystyle \sigma^2 = np(1-p) $$.
-
MGF: $$\displaystyle M(t) = (1-p + p e^t)^n $$.
-
Properties: Sum of independent $$\displaystyle \text{Bin}(n_i, p) $$ is $$\displaystyle \text{Bin}(\sum n_i, p) $$.
-
Application: Number of successes in $n$ independent Bernoulli trials.
B. Poisson Distribution $\text{Pois}(\lambda)$
-
pmf: $$\displaystyle P(X = x) = \frac{e^{-\lambda} \lambda^x}{x!} $$, $$\displaystyle x = 0,1,2,\dots $$.
-
Mean & Variance: $\lambda$.
-
MGF: $$\displaystyle M(t) = e^{\lambda(e^t - 1)} $$.
-
Limiting case: $\text{Bin}(n, p) \to \text{Pois}(\lambda)$ as $$\displaystyle n \to \infty, p \to 0, np = \lambda $$.
-
Application: Number of events in fixed interval (time, space).
C. Hypergeometric Distribution $\text{Hyp}(N, K, n)$
-
pmf: $$\displaystyle P(X = x) = \frac{\binom{K}{x} \binom{N-K}{n-x}}{\binom{N}{n}} $$, $$\displaystyle x = \max(0, n+K-N), \dots, \min(n, K) $$.
-
Mean: $$\displaystyle n \frac{K}{N} $$.
-
Variance: $$\displaystyle n \frac{K}{N}\left(1-\frac{K}{N}\right)\frac{N-n}{N-1} $$.
-
Application: Sampling without replacement from finite population.
V. Continuous Probability Distributions
A. Exponential Distribution $\text{Exp}(\theta)$ or $\text{Exp}(\lambda)$
-
pdf: $$\displaystyle f(x) = \frac{1}{\theta} e^{-x/\theta} $$, $x \ge 0$ (scale $\theta$), or $$\displaystyle f(x) = \lambda e^{-\lambda x} $$ (rate $$\displaystyle \lambda = 1/\theta $$).
-
Mean: $\theta$ (or $1/\lambda$).
-
Variance: $$\displaystyle \theta^2 $$ (or $$\displaystyle 1/\lambda^2 $$).
-
MGF: $$\displaystyle M(t) = (1 - \theta t)^{-1} $$, $$\displaystyle t < 1/\theta $$.
-
Key Property (Memoryless): $$\displaystyle P(X > s+t | X > s) = P(X > t) $$.
-
Application: Inter-arrival times in Poisson process.
B. Gamma Distribution $\text{Gamma}(\alpha, \theta)$
-
pdf: $$\displaystyle f(x) = \frac{1}{\Gamma(\alpha)\theta^\alpha} x^{\alpha-1} e^{-x/\theta} $$, $$\displaystyle x > 0 $$.
$$\displaystyle \Gamma(\alpha) = \int_0^\infty x^{\alpha-1} e^{-x} dx $$ (gamma function).
-
Mean: $\alpha\theta$.
-
Variance: $$\displaystyle \alpha\theta^2 $$.
-
MGF: $$\displaystyle M(t) = (1 - \theta t)^{-\alpha} $$, $$\displaystyle t < 1/\theta $$.
-
Special Cases:
-
$$\displaystyle \alpha = 1 $$: Exponential.
-
$$\displaystyle \alpha = k/2, \theta = 2 $$: $$\displaystyle \chi^2 $$ distribution with $k$ degrees of freedom.
-
C. Normal Distribution $$\displaystyle N(\mu, \sigma^2) $$
-
pdf: $$\displaystyle f(x) = \frac{1}{\sigma\sqrt{2\pi}} e^{-\frac{(x-\mu)^2}{2\sigma^2}} $$.
-
Mean: $\mu$, Variance: $$\displaystyle \sigma^2 $$.
-
MGF: $$\displaystyle M(t) = e^{\mu t + \frac{1}{2}\sigma^2 t^2} $$.
-
Properties:
-
Symmetric, bell-shaped.
-
Mean = Median = Mode.
-
Empirical Rule: 68% within $\mu \pm \sigma$, 95% within $\mu \pm 2\sigma$, 99.7% within $\mu \pm 3\sigma$.
-
If $$\displaystyle X \sim N(\mu, \sigma^2) $$, then $$\displaystyle Z = \frac{X-\mu}{\sigma} \sim N(0,1) $$ (standard normal).
-
-
Sum of Independent Normals: If $$\displaystyle X_i \sim N(\mu_i, \sigma_i^2) $$ independent, then $$\displaystyle \sum X_i \sim N(\sum \mu_i, \sum \sigma_i^2) $$.
VI. Bivariate Distributions and Correlation
A. Joint, Marginal, and Conditional
Discrete:
-
Joint pmf: $$\displaystyle p(x,y) = P(X=x, Y=y) $$.
-
Marginal: $$\displaystyle p_X(x) = \sum_y p(x,y) $$.
-
Conditional: $$\displaystyle p_{X|Y}(x|y) = \frac{p(x,y)}{p_Y(y)} $$, $$\displaystyle p_Y(y) > 0 $$.
Continuous:
-
Joint pdf: $f(x,y)$ with $$\displaystyle P((X,Y) \in A) = \iint_A f(x,y) dxdy $$.
-
Marginal: $$\displaystyle f_X(x) = \int_{-\infty}^\infty f(x,y) dy $$.
-
Conditional: $$\displaystyle f_{X|Y}(x|y) = \frac{f(x,y)}{f_Y(y)} $$, $$\displaystyle f_Y(y) > 0 $$.
Bivariate Normal Distribution:
pdf with parameters $$\displaystyle \mu_X, \mu_Y, \sigma_X^2, \sigma_Y^2, \rho $$:
$$f(x,y) = \frac{1}{2\pi\sigma_X\sigma_Y\sqrt{1-\rho^2}} \exp\left(-\frac{1}{2(1-\rho^2)}\left[\frac{(x-\mu_X)^2}{\sigma_X^2} - 2\rho\frac{(x-\mu_X)(y-\mu_Y)}{\sigma_X\sigma_Y} + \frac{(y-\mu_Y)^2}{\sigma_Y^2}\right]\right)$$
B. Independence of Random Variables
$X, Y$ independent iff:
-
Discrete: $$\displaystyle p(x,y) = p_X(x)p_Y(y) $$ for all $x,y$.
-
Continuous: $$\displaystyle f(x,y) = f_X(x)f_Y(y) $$ for all $x,y$. Implication: If independent, $$\displaystyle E(XY) = E(X)E(Y) $$.
C. Covariance and Pearson Correlation
Covariance:
$$\text{Cov}(X,Y) = E[(X-\mu_X)(Y-\mu_Y)] = E(XY) - E(X)E(Y)$$
Properties:
-
$$\displaystyle \text{Cov}(aX+b, cY+d) = ac \cdot \text{Cov}(X,Y) $$.
-
$$\displaystyle \text{Var}(X+Y) = \text{Var}(X) + \text{Var}(Y) + 2\text{Cov}(X,Y) $$.
Pearson Correlation Coefficient:
$$\rho = \frac{\text{Cov}(X,Y)}{\sigma_X \sigma_Y}$$
Properties:
-
$-1 \le \rho \le 1$.
-
Invariant under linear transformations: $$\displaystyle X' = aX+b $$, $$\displaystyle Y' = cY+d $$ ($$\displaystyle a,c>0 $$) does not change $\rho$.
-
$$\displaystyle \rho = 0 $$: uncorrelated (not necessarily independent).
-
$$\displaystyle \rho = \pm 1 $$: perfect linear relationship.
Sample Correlation (r):
$$r = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum (x_i - \bar{x})^2 \sum (y_i - \bar{y})^2}}$$
D. Spearman's Rank Correlation ($$\displaystyle \rho_s $$)
-
No ties: $$\displaystyle \rho_s = 1 - \frac{6\sum d_i^2}{n(n^2-1)} $$, where $$\displaystyle d_i = \text{rank}(x_i) - \text{rank}(y_i) $$.
-
With ties: Use formula with adjusted ranks.
-
Measures monotonic relationship, robust to outliers.
E. Simple Linear Regression
Model: $$\displaystyle y = a + bx $$. Least Squares Estimation:
$$b = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sum (x_i - \bar{x})^2}, \quad a = \bar{y} - b\bar{x}$$
Regression Coefficients:
-
$$\displaystyle b_{yx} = r \frac{\sigma_y}{\sigma_x} $$ (regression of $y$ on $x$).
-
$$\displaystyle b_{xy} = r \frac{\sigma_x}{\sigma_y} $$ (regression of $x$ on $y$). Properties:
-
$$\displaystyle b_{yx} \cdot b_{xy} = r^2 $$.
-
Same sign as $r$.
-
Regression predicts; correlation measures association.
F. Partial and Multiple Correlation
-
Partial Correlation: Correlation between two variables after removing linear effect of others.
-
Multiple Correlation: Correlation between one variable and a set of others (e.g., $$\displaystyle R_{y.x_1,x_2} $$).
VII. Sampling and Statistical Inference
A. Sampling Distribution of Sample Mean
-
For large samples ($n \ge 30$): approximately normal by Central Limit Theorem (CLT), regardless of population distribution.
-
For normal populations: $$\displaystyle \bar{X} \sim N(\mu, \sigma^2/n) $$ exactly.
B. Estimation
Point Estimation:
-
$\mu$: $$\displaystyle \hat{\mu} = \bar{x} $$.
-
$p$: $$\displaystyle \hat{p} = x/n $$.
Confidence Intervals:
-
Mean (known $\sigma$): $$\displaystyle \bar{x} \pm z_{\alpha/2} \frac{\sigma}{\sqrt{n}} $$.
-
Mean (unknown $\sigma$, large $n$): $$\displaystyle \bar{x} \pm z_{\alpha/2} \frac{s}{\sqrt{n}} $$.
-
Proportion: $$\displaystyle \hat{p} \pm z_{\alpha/2} \sqrt{\frac{\hat{p}(1-\hat{p})}{n}} $$.
C. Hypothesis Testing
Key Concepts:
-
$$\displaystyle H_0 $$: null hypothesis (status quo).
-
$$\displaystyle H_1 $$: alternative hypothesis (research claim).
-
Test Statistic: Standardized measure.
-
p-value: Probability of observing data as extreme as sample, assuming $$\displaystyle H_0 $$ true.
-
Significance Level ($\alpha$): Threshold for rejecting $$\displaystyle H_0 $$ (commonly 0.05).
-
Type I Error: Rejecting $$\displaystyle H_0 $$ when true (probability $\alpha$).
-
Type II Error: Failing to reject $$\displaystyle H_0 $$ when false (probability $\beta$).
Common Tests:
- One-sample z-test (large $n$, known $\sigma$):
$$z = \frac{\bar{x} - \mu_0}{\sigma/\sqrt{n}}$$
- Two-sample z-test for difference (large samples):
$$z = \frac{(\bar{x}_1 - \bar{x}_2) - (\mu_1 - \mu_2)}{\sqrt{\frac{\sigma_1^2}{n_1} + \frac{\sigma_2^2}{n_2}}}$$
- One-sample t-test (small $n$, unknown $\sigma$):
$$t = \frac{\bar{x} - \mu_0}{s/\sqrt{n}}, \quad df = n-1$$
- Proportion tests: Normal approximation to binomial.
Steps:
-
State $$\displaystyle H_0 $$ and $$\displaystyle H_1 $$.
-
Choose appropriate test and significance level $\alpha$.
-
Compute test statistic.
-
Find critical value or p-value.
-
Decide: reject $$\displaystyle H_0 $$ if $$\displaystyle |stat| > critical $$ or $$\displaystyle p\text{-value} < \alpha $$.
-
Conclude in context.
D. Chi-Square ($$\displaystyle \chi^2 $$) Tests
Chi-Square Distribution:
-
$$\displaystyle \chi^2_k $$: sum of $k$ independent $$\displaystyle N(0,1)^2 $$.
-
Mean = $k$, Variance = $2k$.
-
Right-skewed, becomes normal for large $k$.
Goodness of Fit Test:
-
Purpose: Test if observed frequencies fit a theoretical distribution.
-
Conditions:
-
Random sample.
-
Categorical data.
-
Expected frequencies $$\displaystyle E_i \ge 5 $$ (generally; some allow up to 20% with $$\displaystyle E_i \ge 1 $$).
-
-
Test Statistic:
$$\chi^2 = \sum_{i=1}^k \frac{(O_i - E_i)^2}{E_i}$$
- Degrees of Freedom: $$\displaystyle df = k - 1 - m $$, where $k$ = number of categories, $m$ = number of parameters estimated from data.
Test of Independence:
-
For contingency table (r rows, c columns).
-
$$\displaystyle H_0 $$: Row and column variables independent.
-
Test Statistic: Same $$\displaystyle \chi^2 = \sum \frac{(O_{ij} - E_{ij})^2}{E_{ij}} $$.
-
df: $(r-1)(c-1)$.
-
Expected: $$\displaystyle E_{ij} = \frac{(\text{row } i \text{ total}) \times (\text{column } j \text{ total})}{\text{grand total}} $$.
E. F-test for Equality of Variances
-
Test Statistic: $$\displaystyle F = \frac{s_1^2}{s_2^2} $$ (larger variance in numerator).
-
Distribution: $$\displaystyle F(df_1 = n_1-1, df_2 = n_2-1) $$.
-
Application: Compare two population variances (assumes normality).
VIII. Measures of Central Tendency and Dispersion
A. Central Tendency
| Measure | Raw Data | Grouped Data | Merits | Demerits |
|---|---|---|---|---|
| Mean $$\displaystyle \bar{x} = \frac{\sum x_i}{n} $$ | $$\displaystyle \bar{x} = \frac{\sum f_i x_i}{\sum f_i} $$ | Uses all data, good for further math | Affected by outliers | |
| Median | Middle value (ordered) | $$\displaystyle L + \frac{\frac{n}{2} - cf}{f} \times c $$ | Robust to outliers | Not for further math, less precise |
| Mode | Most frequent value | $$\displaystyle L + \frac{f_1 - f_0}{2f_1 - f_0 - f_2} \times c $$ | Easy, for categorical data | May not be unique, not stable |
B. Dispersion
| Measure | Formula (Sample) | Formula (Population) |
|---|---|---|
| Range | $\max - \min$ | Same |
| Variance | $$\displaystyle s^2 = \frac{\sum (x_i - \bar{x})^2}{n-1} $$ | $$\displaystyle \sigma^2 = \frac{\sum (x_i - \mu)^2}{N} $$ |
| Std. Dev. | $$\displaystyle s = \sqrt{s^2} $$ | $$\displaystyle \sigma = \sqrt{\sigma^2} $$ |
| Mean Deviation | $$\displaystyle M.D. = \frac{\sum |x_i - \text{median}|}{n} $$ (often about median) | Similar |
| Coefficient of Variation | $$\displaystyle CV = \frac{s}{\bar{x}} \times 100\% $$ | $$\displaystyle CV = \frac{\sigma}{\mu} \times 100\% $$ |
[!TIP] Use $n-1$ for sample variance (unbiased estimator of $$\displaystyle \sigma^2 $$).
IX. Skewness and Kurtosis
A. Skewness (Asymmetry)
-
Symmetric: Mean = Median = Mode.
-
Positive Skew (Right): Tail to right; Mean > Median > Mode.
-
Negative Skew (Left): Tail to left; Mean < Median < Mode.
Measures:
-
Pearson's First Coefficient: $$\displaystyle \frac{\text{Mean} - \text{Mode}}{\sigma} $$.
-
Pearson's Second Coefficient: $$\displaystyle \frac{3(\text{Mean} - \text{Median})}{\sigma} $$.
-
Moment Coefficient:
$$\beta_1 = \frac{\mu_3^2}{\mu_2^3}, \quad \gamma_1 = \sqrt{\beta_1} \text{ (sign from } \mu_3)$$
$$\displaystyle \gamma_1 = 0 $$ implies symmetry.
B. Kurtosis (Peakedness)
-
Mesokurtic: $$\displaystyle \beta_2 = 3 $$ (normal-like).
-
Leptokurtic: $$\displaystyle \beta_2 > 3 $$ (peaked, heavy tails).
-
Platykurtic: $$\displaystyle \beta_2 < 3 $$ (flat, light tails).
Measures:
-
$$\displaystyle \beta_2 = \frac{\mu_4}{\mu_2^2} $$.
-
Excess Kurtosis: $$\displaystyle \beta_2 - 3 $$.
C. Calculation from Moments
Given central moments $$\displaystyle \mu_1=0, \mu_2, \mu_3, \mu_4 $$:
-
Skewness: $$\displaystyle \gamma_1 = \frac{\mu_3}{\mu_2^{3/2}} $$.
-
Kurtosis: $$\displaystyle \beta_2 = \frac{\mu_4}{\mu_2^2} $$.
X. Curve Fitting
A. Method of Least Squares
Minimize sum of squared residuals:
$$S = \sum_{i=1}^n (y_i - \hat{y}_i)^2$$
where $$\displaystyle \hat{y}_i = f(x_i; \text{parameters}) $$.
B. Fitting a Straight Line $$\displaystyle y = a + bx $$
Normal Equations:
$$\begin{cases} \sum y = na + b\sum x \\ \sum xy = a\sum x + b\sum x^2 \end{cases}$$
Solve for $a, b$.
C. Fitting a Second Degree Parabola $$\displaystyle y = a + bx + cx^2 $$
Normal Equations:
$$\begin{cases} \sum y = na + b\sum x + c\sum x^2 \\ \sum xy = a\sum x + b\sum x^2 + c\sum x^3 \\ \sum x^2 y = a\sum x^2 + b\sum x^3 + c\sum x^4 \end{cases}$$
Solve the 3ร3 system.
D. Polynomial Regression (General)
Model: $$\displaystyle y = a_0 + a_1 x + a_2 x^2 + \cdots + a_k x^k $$.
Normal equations: $$\displaystyle \sum x^{i+j} a_j = \sum x^i y $$ for $$\displaystyle i = 0, \dots, k $$.
XI. Additional Important Topics
A. Chebyshev's Inequality
For any random variable $X$ with finite mean $\mu$ and variance $$\displaystyle \sigma^2 $$, and any $$\displaystyle k > 0 $$:
$$P(|X - \mu| \ge k\sigma) \le \frac{1}{k^2}$$
Proof Sketch (Markov): Apply Markov to $$\displaystyle Y = (X-\mu)^2 $$:
$$P(|X-\mu| \ge k\sigma) = P((X-\mu)^2 \ge k^2\sigma^2) \le \frac{E[(X-\mu)^2]}{k^2\sigma^2} = \frac{\sigma^2}{k^2\sigma^2} = \frac{1}{k^2}.$$
Application: Provides probability bounds without distributional assumptions (e.g., at least $$\displaystyle 1 - 1/k^2 $$ of data within $k$ standard deviations).
B. Central Limit Theorem (CLT)
Statement: For independent, identically distributed random variables $$\displaystyle X_1, \dots, X_n $$ with finite mean $\mu$ and variance $$\displaystyle \sigma^2 $$, the distribution of the sample mean $\bar{X}$ approaches normal as $n \to \infty$:
$$\frac{\bar{X} - \mu}{\sigma/\sqrt{n}} \xrightarrow{d} N(0,1).$$
Implication: Justifies normal approximation for sample means in inference, even for non-normal populations, when $n$ is large ($n \ge 30$ typically).
๐ Exam Focus Summary
Based on past RGPV papers, prioritize:
-
Proofs: $$\displaystyle E(X+Y)=E(X)+E(Y) $$, $$\displaystyle \text{Var}(aX+b)=a^2\text{Var}(X) $$, binomial $\to$ Poisson limit, normal MGF.
-
Distributions: Binomial, Poisson, Exponential, Gamma, Normal โ know pmf/pdf, mean, variance, MGF.
-
Correlation vs. Regression: Differences, formulas ($$\displaystyle b_{yx}, b_{xy}, r $$).
-
Chi-Square: Goodness of fit conditions, test statistic, df calculation.
-
Hypothesis Testing: Steps, z-test, t-test, p-value interpretation.
-
Curve Fitting: Normal equations for straight line & parabola.
-
Skewness/Kurtosis: Calculation from moments ($$\displaystyle \beta_1, \beta_2 $$).
-
Bayes' Theorem: Applications (medical diagnosis, defective items).
-
Chebyshev & CLT: Statements and applications.
Final Tip: Always state assumptions (e.g., independence, normality) when applying tests. For discrete distributions, verify $$\displaystyle \sum p(x)=1 $$ when finding $K$. For continuous, ensure pdf integrates to 1.