1. FOUNDATIONS OF PROBABILITY
-
Basic Concepts:
-
Experiment: A process with uncertain outcome.
-
Sample Space (S): Set of all possible outcomes.
-
Event: Subset of sample space.
-
Simple: Single outcome.
-
Compound: Multiple outcomes.
-
Mutually Exclusive: $$\displaystyle A \cap B = \emptyset $$.
-
Exhaustive: $$\displaystyle A_1 \cup A_2 \cup \cdots = S $$.
-
Equally Likely: Each outcome has same probability.
-
-
-
Axioms of Probability:
-
$$\displaystyle P(S) = 1 $$
-
$P(A) \ge 0$ for any event $A$
-
For mutually exclusive $$\displaystyle A_1, A_2, \dots $$, $$\displaystyle P\left(\bigcup_{i=1}^{\infty} A_i\right) = \sum_{i=1}^{\infty} P(A_i) $$
-
-
Conditional Probability & Independence:
-
$$\displaystyle P(A|B) = \frac{P(A \cap B)}{P(B)} $$, provided $$\displaystyle P(B) > 0 $$.
-
$$\displaystyle P(A|B') = \frac{P(A \cap B')}{P(B')} = \frac{P(A) - P(A \cap B)}{1 - P(B)} $$.
-
Independence: $A$ and $B$ are independent iff $$\displaystyle P(A \cap B) = P(A)P(B) $$.
-
-
Theorem of Total Probability: If $$\displaystyle B_1, B_2, \dots, B_n $$ partition $S$, then
$$P(A) = \sum_{i=1}^{n} P(A|B_i) P(B_i).$$
- Bayes' Theorem: For partition $$\displaystyle B_1, \dots, B_n $$,
$$P(B_i|A) = \frac{P(A|B_i) P(B_i)}{\sum_{j=1}^{n} P(A|B_j) P(B_j)}.$$
> [!TIP] Common in medical diagnosis and false positive problems. Always identify the partition and prior probabilities $$\displaystyle P(B_i) $$.
- Combinatorial Probability: Use permutations ($$\displaystyle ^nP_r $$) and combinations ($$\displaystyle ^nC_r $$) to count favorable and total outcomes when equally likely: $$\displaystyle P(\text{event}) = \frac{\text{favorable}}{\text{total}} $$.
2. RANDOM VARIABLES & THEIR DISTRIBUTIONS
2.1 Random Variables (RVs)
-
Definition: Function assigning real numbers to outcomes.
-
Types:
-
Discrete: Countable set of values (e.g., Binomial, Poisson).
-
Continuous: Uncountable values (e.g., Normal, Exponential).
-
-
Probability Mass Function (PMF): $$\displaystyle p(x) = P(X = x) $$ for discrete $X$. Satisfies $p(x) \ge 0$, $$\displaystyle \sum p(x) = 1 $$.
-
Probability Density Function (PDF): $f(x)$ for continuous $X$, with $$\displaystyle P(a \le X \le b) = \int_a^b f(x)dx $$. Satisfies $f(x) \ge 0$, $$\displaystyle \int_{-\infty}^{\infty} f(x)dx = 1 $$.
-
Cumulative Distribution Function (CDF): $$\displaystyle F(x) = P(X \le x) $$.
-
For discrete: $$\displaystyle F(x) = \sum_{t \le x} p(t) $$.
-
For continuous: $$\displaystyle F(x) = \int_{-\infty}^{x} f(t)dt $$.
-
Properties: Non-decreasing, $$\displaystyle \lim_{x \to -\infty} F(x)=0 $$, $$\displaystyle \lim_{x \to \infty} F(x)=1 $$, $$\displaystyle P(a < X \le b) = F(b) - F(a) $$.
-
-
Expectation (Mean):
-
Discrete: $$\displaystyle E(X) = \sum x \cdot p(x) $$.
-
Continuous: $$\displaystyle E(X) = \int_{-\infty}^{\infty} x f(x)dx $$.
-
-
Variance & Standard Deviation:
$$\boxed{V(X) = E(X^2) - [E(X)]^2}$$
* SD: $$\displaystyle \sigma_X = \sqrt{V(X)} $$.
-
Properties of Expectation:
-
Linearity: $$\displaystyle E(aX + b) = aE(X) + b $$.
-
Additivity: $$\displaystyle E(X_1 + X_2 + \cdots + X_n) = E(X_1) + E(X_2) + \cdots + E(X_n) $$ (No independence required).
-
Independence: If $$\displaystyle X_1, X_2, \dots, X_n $$ are independent, $$\displaystyle E(X_1 X_2 \cdots X_n) = E(X_1) E(X_2) \cdots E(X_n) $$.
-
-
Properties of Variance:
-
$$\displaystyle V(aX + b) = a^2 V(X) $$.
-
If independent: $$\displaystyle V(X_1 + X_2) = V(X_1) + V(X_2) $$.
-
If independent: $$\displaystyle V(X_1 - X_2) = V(X_1) + V(X_2) $$.
-
2.2 Discrete Probability Distributions
-
Binomial Distribution $\mathbf{Bin(n, p)}$:
-
Conditions: Fixed $n$ trials, constant $p$, independent trials, two outcomes (success/failure).
-
PMF: $$\displaystyle P(X = x) = \binom{n}{x} p^x q^{n-x} $$, $$\displaystyle x = 0,1,\dots,n $$, $$\displaystyle q=1-p $$.
-
Mean & Variance: $$\displaystyle \boxed{E(X) = np} $$, $$\displaystyle \boxed{V(X) = npq} $$.
-
MGF: $$\displaystyle M_X(t) = (q + p e^t)^n $$.
-
Poisson Approximation: When $n \to \infty$, $p \to 0$, $$\displaystyle np = \lambda $$ (constant), Binomial $\to$ Poisson($\lambda$).
-
-
Poisson Distribution $\mathbf{Pois(\lambda)}$:
-
Conditions: Events in fixed interval/space, independent, constant rate $\lambda$, rare events.
-
PMF: $$\displaystyle P(X = x) = \frac{e^{-\lambda} \lambda^x}{x!} $$, $$\displaystyle x = 0,1,2,\dots $$.
-
Mean & Variance: $$\displaystyle \boxed{E(X) = \lambda} $$, $$\displaystyle \boxed{V(X) = \lambda} $$.
-
MGF: $$\displaystyle M_X(t) = e^{\lambda(e^t - 1)} $$.
-
Additivity: Sum of independent Poissons is Poisson: $$\displaystyle X \sim \text{Pois}(\lambda_1) $$, $$\displaystyle Y \sim \text{Pois}(\lambda_2) $$ $\Rightarrow$ $$\displaystyle X+Y \sim \text{Pois}(\lambda_1+\lambda_2) $$.
-
-
Hypergeometric Distribution (Brief):
-
Conditions: Sampling without replacement from finite population of size $N$ with $K$ successes.
-
PMF: $$\displaystyle P(X = x) = \frac{\binom{K}{x} \binom{N-K}{n-x}}{\binom{N}{n}} $$.
-
Mean: $$\displaystyle E(X) = n \frac{K}{N} $$.
-
Variance: $$\displaystyle V(X) = n \frac{K}{N} \left(1 - \frac{K}{N}\right) \frac{N-n}{N-1} $$.
-
2.3 Continuous Probability Distributions
-
Uniform Distribution $\mathbf{U(a, b)}$:
-
PDF: $$\displaystyle f(x) = \frac{1}{b-a} $$ for $a \le x \le b$, 0 elsewhere.
-
Mean: $$\displaystyle \frac{a+b}{2} $$, Variance: $$\displaystyle \frac{(b-a)^2}{12} $$.
-
-
Exponential Distribution $\mathbf{Exp(\theta)}$ (scale $\theta$ or rate $$\displaystyle \lambda=1/\theta $$):
-
PDF: $$\displaystyle f(x) = \frac{1}{\theta} e^{-x/\theta} $$ for $x \ge 0$.
-
Mean & Variance: $$\displaystyle \boxed{E(X) = \theta} $$, $$\displaystyle \boxed{V(X) = \theta^2} $$.
-
MGF: $$\displaystyle M_X(t) = \frac{1}{1 - \theta t} $$ for $$\displaystyle t < 1/\theta $$.
-
Memoryless Property: $$\displaystyle P(X > s+t | X > s) = P(X > t) $$ for $s,t \ge 0$.
-
-
Gamma Distribution $\mathbf{Gamma(\alpha, \beta)}$ (shape $\alpha$, scale $\beta$):
-
PDF: $$\displaystyle f(x) = \frac{1}{\Gamma(\alpha) \beta^{\alpha}} x^{\alpha-1} e^{-x/\beta} $$ for $$\displaystyle x > 0 $$.
-
Mean & Variance: $$\displaystyle \boxed{E(X) = \alpha\beta} $$, $$\displaystyle \boxed{V(X) = \alpha\beta^2} $$.
-
Relationship: Exponential is Gamma with $$\displaystyle \alpha = 1 $$.
-
-
Normal Distribution $$\displaystyle \mathbf{N(\mu, \sigma^2)} $$:
-
PDF: $$\displaystyle f(x) = \frac{1}{\sigma\sqrt{2\pi}} e^{-\frac{(x-\mu)^2}{2\sigma^2}} $$.
-
Properties: Symmetric, bell-shaped, mean=median=mode, inflection at $\mu \pm \sigma$, $$\displaystyle P(\mu - \sigma < X < \mu + \sigma) \approx 0.68 $$.
-
Standard Normal: $$\displaystyle Z = \frac{X - \mu}{\sigma} \sim N(0,1) $$. Use Z-tables.
-
MGF: $$\displaystyle M_X(t) = e^{\mu t + \frac{1}{2} \sigma^2 t^2} $$.
-
Additivity: Sum of independent normals is normal.
-
-
Chi-Square Distribution $$\displaystyle \mathbf{\chi^2_k} $$:
-
Definition: Distribution of $$\displaystyle \sum_{i=1}^{k} Z_i^2 $$ where $$\displaystyle Z_i \sim N(0,1) $$ independent.
-
PDF: $$\displaystyle f(x) = \frac{1}{2^{k/2} \Gamma(k/2)} x^{k/2 - 1} e^{-x/2} $$ for $$\displaystyle x > 0 $$.
-
Properties: Mean $$\displaystyle = k $$, Variance $$\displaystyle = 2k $$, additive for independent $$\displaystyle \chi^2 $$ variables.
-
3. MOMENTS & MEASURES OF SHAPE
-
Moments:
-
Raw Moments: $$\displaystyle \mu'_r = E(X^r) $$.
-
Central Moments: $$\displaystyle \mu_r = E[(X - \mu)^r] $$.
-
Relationship: $$\displaystyle \mu_2 = \mu'_2 - \mu^2 = V(X) $$.
-
-
Measures of Central Tendency:
-
Mean: $$\displaystyle \bar{x} = \frac{\sum x_i}{n} $$ (or $$\displaystyle \frac{\sum f_i x_i}{N} $$). Uses all data, sensitive to outliers.
-
Median: Middle value when ordered. Robust to outliers.
-
Mode: Most frequent value. May be multiple or absent.
[!TIP] For grouped data, median formula: $$\displaystyle \text{Median} = L + \frac{\frac{N}{2} - C_f}{f} \times h $$.
-
-
Measures of Dispersion:
-
Range: Max $-$ Min. Simple but uses only extremes.
-
Variance: $$\displaystyle s^2 = \frac{\sum (x_i - \bar{x})^2}{n-1} $$ (sample).
-
Standard Deviation: $$\displaystyle s = \sqrt{s^2} $$.
-
Mean Deviation: $$\displaystyle \frac{\sum |x_i - \text{median}|}{n} $$ (or mean).
-
-
Measures of Skewness:
-
Concept: Asymmetry of distribution.
-
Pearson's Coefficient: $$\displaystyle \text{Sk}_p = \frac{\text{Mean} - \text{Mode}}{\text{SD}} \approx \frac{3(\text{Mean} - \text{Median})}{\text{SD}} $$.
-
Quartile Coefficient: $$\displaystyle \frac{Q_3 + Q_1 - 2Q_2}{Q_3 - Q_1} $$.
-
Tests (Moment-based): $$\displaystyle \beta_1 = \frac{\mu_3^2}{\mu_2^3} $$. $$\displaystyle \beta_1 > 0 $$: positive skew; $$\displaystyle \beta_1 < 0 $$: negative skew.
-
-
Measures of Kurtosis:
-
Concept: Peakedness/tailedness relative to normal.
-
Pearson's Kurtosis: $$\displaystyle \beta_2 = \frac{\mu_4}{\mu_2^2} $$.
-
Excess Kurtosis: $$\displaystyle \gamma_2 = \beta_2 - 3 $$.
-
$$\displaystyle \beta_2 > 3 $$: leptokurtic (heavy tails); $$\displaystyle \beta_2 < 3 $$: platykurtic (light tails); $$\displaystyle \beta_2 = 3 $$: mesokurtic (normal).
-
4. BIVARIATE ANALYSIS: CORRELATION & REGRESSION
4.1 Bivariate Distributions
-
Definition: Joint distribution of two RVs $(X, Y)$.
-
Joint PMF/PDF: $$\displaystyle f(x,y) = P(X=x, Y=y) $$ (discrete) or joint density (continuous).
-
Marginal Distributions: $$\displaystyle f_X(x) = \sum_y f(x,y) $$ or $$\displaystyle \int f(x,y)dy $$; similarly for $$\displaystyle f_Y(y) $$.
-
Conditional Distributions: $$\displaystyle f_{Y|X}(y|x) = \frac{f(x,y)}{f_X(x)} $$.
-
Independence: $X$ and $Y$ independent iff $$\displaystyle f(x,y) = f_X(x) f_Y(y) $$ for all $x,y$.
-
Bivariate Normal Distribution:
$$f(x,y) = \frac{1}{2\pi \sigma_x \sigma_y \sqrt{1-\rho^2}} \exp\left(-\frac{1}{2(1-\rho^2)}\left[\frac{(x-\mu_x)^2}{\sigma_x^2} + \frac{(y-\mu_y)^2}{\sigma_y^2} - \frac{2\rho(x-\mu_x)(y-\mu_y)}{\sigma_x \sigma_y}\right]\right)$$
* Marginals: $$\displaystyle X \sim N(\mu_x, \sigma_x^2) $$, $$\displaystyle Y \sim N(\mu_y, \sigma_y^2) $$.
* Conditional distributions are normal.
4.2 Correlation Analysis
-
Karl Pearson's Correlation Coefficient ($r$):
- Definition: Measure of linear association.
$$r = \frac{\text{Cov}(X,Y)}{\sigma_x \sigma_y} = \frac{E[(X-\mu_x)(Y-\mu_y)]}{\sigma_x \sigma_y}$$
* **Raw data formula:**
$$r = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum (x_i - \bar{x})^2 \sum (y_i - \bar{y})^2}}$$
* **Properties:** $-1 \le r \le 1$; sign indicates direction; independent of origin/scale; $$\displaystyle r=0 $$ does not imply independence (only no linear correlation).
-
Spearman's Rank Correlation Coefficient ($\rho$):
-
For ranked data, no ties: $$\displaystyle \rho = 1 - \frac{6 \sum d_i^2}{n(n^2-1)} $$, where $$\displaystyle d_i = \text{rank}(x_i) - \text{rank}(y_i) $$.
-
With ties: Use Pearson's formula on ranks.
-
Maximum Value Proof: When ranks are perfectly concordant, $$\displaystyle d_i = 0 $$ for all $i$, so $$\displaystyle \rho = 1 $$. By Cauchy-Schwarz, $|\rho| \le 1$.
[!TIP] Rank data first, handle ties by assigning average ranks.
-
4.3 Regression Analysis
-
Concept: Predict dependent variable $Y$ from independent $X$.
-
Regression Lines:
-
$Y$ on $X$: $$\displaystyle y = a + b x $$.
-
$X$ on $Y$: $$\displaystyle x = c + d y $$.
-
-
Regression Coefficients:
$$b = r \frac{\sigma_y}{\sigma_x}, \quad d = r \frac{\sigma_x}{\sigma_y}$$
* Properties: $b$ and $d$ have same sign as $r$; $$\displaystyle bd = r^2 $$; geometric mean of $b$ and $d$ is $|r|$.
-
Method of Least Squares (for $$\displaystyle y = a + bx $$):
-
Minimize $$\displaystyle S = \sum (y_i - a - b x_i)^2 $$.
-
Normal Equations:
-
$$\frac{\partial S}{\partial a} = 0 \Rightarrow \sum y_i = n a + b \sum x_i$$
$$\frac{\partial S}{\partial b} = 0 \Rightarrow \sum x_i y_i = a \sum x_i + b \sum x_i^2$$
* Solve for $a, b$: $$\displaystyle b = \frac{n \sum x_i y_i - \sum x_i \sum y_i}{n \sum x_i^2 - (\sum x_i)^2} $$, $$\displaystyle a = \bar{y} - b \bar{x} $$.
-
Fitting Curves:
-
Second-Degree Parabola: $$\displaystyle y = a + b x + c x^2 $$.
Normal equations:
-
$$ \begin{aligned} \sum y &= n a + b \sum x + c \sum x^2 \\ \sum x y &= a \sum x + b \sum x^2 + c \sum x^3 \\ \sum x^2 y &= a \sum x^2 + b \sum x^3 + c \sum x^4 \end{aligned} $$
* **Other Forms (e.g., $$\displaystyle y = \alpha x + \beta x^2 $$):** Set up equations by multiplying by $x$ and $$\displaystyle x^2 $$ and summing.
5. CURVE FITTING & METHOD OF LEAST SQUARES
-
Principle: Choose curve parameters to minimize sum of squared residuals $$\displaystyle S = \sum (y_i - \hat{y}_i)^2 $$.
-
Fitting a Straight Line: As above, $$\displaystyle y = a + bx $$.
-
Fitting a Second-Degree Parabola: As above, $$\displaystyle y = a + bx + cx^2 $$.
-
Fitting Other Curves: Transform to linear form if possible (e.g., $$\displaystyle y = a e^{bx} $$ $\Rightarrow$ $$\displaystyle \ln y = \ln a + b x $$; $$\displaystyle y = a x^b $$ $\Rightarrow$ $$\displaystyle \log y = \log a + b \log x $$). Apply least squares to transformed variables.
6. STATISTICAL INFERENCE: HYPOTHESIS TESTING
6.1 Hypothesis Testing Fundamentals
-
Null Hypothesis ($$\displaystyle H_0 $$): Statement of no effect/difference.
-
Alternative Hypothesis ($$\displaystyle H_1 $$): Statement we want to establish.
-
Type I Error: Rejecting $$\displaystyle H_0 $$ when true. Probability = $\alpha$ (level of significance).
-
Type II Error: Accepting $$\displaystyle H_0 $$ when false. Probability = $\beta$.
-
One-tailed Test: $$\displaystyle H_1 $$: $$\displaystyle \mu > \mu_0 $$ or $$\displaystyle \mu < \mu_0 $$.
-
Two-tailed Test: $$\displaystyle H_1 $$: $$\displaystyle \mu \neq \mu_0 $$.
-
Test Statistic: Standardized measure under $$\displaystyle H_0 $$.
-
Critical Region: Values of test statistic leading to rejection of $$\displaystyle H_0 $$.
-
p-value: Probability of observing test statistic as extreme as sample, assuming $$\displaystyle H_0 $$ true. Reject $$\displaystyle H_0 $$ if p-value $$\displaystyle < \alpha $$.
6.2 Large Sample Tests (Z-test)
-
Test for Single Mean:
-
$\sigma$ known: $$\displaystyle Z = \frac{\bar{x} - \mu_0}{\sigma / \sqrt{n}} \sim N(0,1) $$ under $$\displaystyle H_0 $$.
-
$\sigma$ unknown, $n$ large: $$\displaystyle Z = \frac{\bar{x} - \mu_0}{s / \sqrt{n}} \approx N(0,1) $$ by CLT.
-
-
Test for Difference of Means (Two Independent Samples):
-
$$\displaystyle \sigma_1, \sigma_2 $$ known: $$\displaystyle Z = \frac{(\bar{x}_1 - \bar{x}_2) - (\mu_1 - \mu_2)}{\sqrt{\frac{\sigma_1^2}{n_1} + \frac{\sigma_2^2}{n_2}}} $$.
-
$$\displaystyle \sigma_1 = \sigma_2 = \sigma $$ unknown, large $n$: Use pooled variance $$\displaystyle s_p^2 = \frac{(n_1-1)s_1^2 + (n_2-1)s_2^2}{n_1+n_2-2} $$, then $$\displaystyle Z = \frac{(\bar{x}_1 - \bar{x}_2) - (\mu_1 - \mu_2)}{s_p \sqrt{\frac{1}{n_1} + \frac{1}{n_2}}} $$.
-
-
Confidence Interval (Large Sample): $$\displaystyle \bar{x} \pm z_{\alpha/2} \frac{\sigma}{\sqrt{n}} $$ (or using $s$).
6.3 Tests for Proportions
- Single Proportion: $$\displaystyle H_0: p = p_0 $$. Under $$\displaystyle H_0 $$, $$\displaystyle \hat{p} \sim N\left(p_0, \frac{p_0 q_0}{n}\right) $$ approx. for large $n$.
$$Z = \frac{\hat{p} - p_0}{\sqrt{\frac{p_0 q_0}{n}}}$$
- Difference of Proportions (Two Samples): $$\displaystyle H_0: p_1 = p_2 $$. Pooled $$\displaystyle \hat{p} = \frac{x_1 + x_2}{n_1 + n_2} $$.
$$Z = \frac{\hat{p}_1 - \hat{p}_2}{\sqrt{\hat{p} \hat{q} \left(\frac{1}{n_1} + \frac{1}{n_2}\right)}}$$
6.4 Chi-Square ($$\displaystyle \chi^2 $$) Test
-
Chi-Square Distribution: $$\displaystyle \chi^2_k $$ has $k$ degrees of freedom. Mean $$\displaystyle = k $$, Variance $$\displaystyle = 2k $$, additive for independent $$\displaystyle \chi^2 $$s.
-
Goodness of Fit Test:
-
Purpose: Test if observed frequencies follow a specified theoretical distribution.
-
Conditions:
-
Observations are random and independent.
-
Expected frequencies $$\displaystyle E_i \ge 5 $$ (commonly, at least 80% of $$\displaystyle E_i \ge 5 $$, none $$\displaystyle < 1 $$).
-
Large total sample size.
-
Data in frequency form (grouped).
-
-
Test Statistic:
-
$$\boxed{\chi^2 = \sum_{i=1}^{k} \frac{(O_i - E_i)^2}{E_i}}$$
* **Degrees of Freedom:** $$\displaystyle df = (\text{number of classes}) - 1 - (\text{number of parameters estimated from data}) $$.
* **Decision:** Reject $$\displaystyle H_0 $$ if $$\displaystyle \chi^2_{calc} > \chi^2_{\alpha, df} $$.
- Test for Independence (Contingency Table): $$\displaystyle \chi^2 = \sum \frac{(O - E)^2}{E} $$, $$\displaystyle df = (r-1)(c-1) $$, where $$\displaystyle E = \frac{(\text{row total})(\text{column total})}{\text{grand total}} $$.
6.5 Small Sample Tests (t-test)
-
t-Distribution: Symmetric, heavier tails than normal. Depends on $\nu$ degrees of freedom. Use t-table.
-
Test for Single Mean ($\sigma$ unknown, small $n$, normality assumed):
$$t = \frac{\bar{x} - \mu_0}{s / \sqrt{n}} \sim t_{n-1}$$
-
Test for Difference of Means (Independent Samples):
- Equal variances (pooled t-test):
$$s_p^2 = \frac{(n_1-1)s_1^2 + (n_2-1)s_2^2}{n_1+n_2-2}, \quad t = \frac{(\bar{x}_1 - \bar{x}_2) - (\mu_1 - \mu_2)}{s_p \sqrt{\frac{1}{n_1} + \frac{1}{n_2}}} \sim t_{n_1+n_2-2}$$
* **Unequal variances (Welch's t-test):** Approximate $\nu$ by Satterthwaite formula.
- Paired t-test (Dependent Samples): For matched pairs, compute differences $$\displaystyle d_i = x_i - y_i $$.
$$t = \frac{\bar{d} - \mu_d}{s_d / \sqrt{n}} \sim t_{n-1}$$
7. APPLICATIONS & PROBLEM-SOLVING FRAMEWORK
-
Probability Problems: Identify events, use axioms, conditional probability, Bayes' theorem, combinatorial counting.
-
Distribution Problems: Identify distribution (Binomial, Poisson, Normal, etc.), use PMF/PDF to find probabilities, compute mean/variance/MGF.
-
Moments Problems: Compute raw/central moments, find $$\displaystyle \beta_1 $$ (skewness) and $$\displaystyle \beta_2 $$ (kurtosis).
-
Correlation/Regression Problems:
-
Compute $r$ from raw data or summary statistics.
-
Compute Spearman's $\rho$ from ranks.
-
Fit regression lines: find $b, a$ for $y$ on $x$; interpret coefficients.
-
Fit parabola: set up normal equations, solve for $a, b, c$.
-
-
Hypothesis Testing Problems:
-
State $$\displaystyle H_0 $$ and $$\displaystyle H_1 $$ clearly.
-
Choose appropriate test (Z, t, $$\displaystyle \chi^2 $$) based on sample size, known/unknown $\sigma$, data type.
-
Compute test statistic.
-
Find critical value or p-value.
-
Compare and conclude in context.
-
-
Confidence Intervals: Construct using appropriate distribution (Z, t) and formula.
-
Goodness-of-Fit Problems:
-
State $$\displaystyle H_0 $$ (data follows specified distribution).
-
Compute expected frequencies $$\displaystyle E_i $$ (estimate parameters if needed).
-
Check conditions ($$\displaystyle E_i \ge 5 $$).
-
Compute $$\displaystyle \chi^2 = \sum \frac{(O_i - E_i)^2}{E_i} $$.
-
Determine $df$.
-
Compare with critical $$\displaystyle \chi^2 $$ value.
-
-
Curve Fitting: Use least squares normal equations for specified model (straight line, parabola, etc.). Solve simultaneous equations.