Skip to content
AD-302 · Probability and Statistics for Data Science/Quick Revision Short Notes

Probability and Statistics for Data Science (AD-302) - Unit 1 Short Notes

How unit 1 is examined

The unit covers data science and its life cycle, descriptive and inferential statistics, averages, and measures of dispersion; marks come from averages, standard deviation, and grouped-data numericals.

Data Science: Introduction

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Data science is an interdisciplinary field that uses statistics, programming and domain knowledge to extract useful knowledge and decisions from data.</mark>

Key points.

  1. It combines statistics, computer science and domain expertise.
  2. It handles structured data (tables) and unstructured data (text, images).
  3. Typical outputs are reports, predictions and recommendations, for example fraud detection or product suggestions.

Data Science Life Cycle

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>The data science life cycle is the iterative sequence of stages that turns a business problem into a deployed data-driven solution.</mark>

Key points.

  1. Business understanding fixes the problem and success measure.
  2. Data collection gathers data from databases, sensors and the web.
  3. Data cleaning removes missing values, duplicates and errors.
  4. Exploration uses summaries and plots to find patterns.
  5. Modelling builds a statistical or machine-learning model.
  6. Evaluation checks accuracy on unseen data with metrics such as accuracy, precision, recall or RMSE.
  7. Communication and visualization presents findings to the client through charts and dashboards, so decisions can be taken.
  8. Deployment puts the model into use, for example behind an API, and monitors it; the cycle is iterative and repeats when performance drops.
  9. Tools by stage: SQL for collection, pandas for cleaning, matplotlib plots for exploration, scikit-learn for modelling, dashboards for communication, a REST API for deployment.

Example. Predicting customer churn: business goal is to cut churn by 10%; SQL pulls customer records; pandas fills missing ages; plots show churn is high for short-tenure users; scikit-learn fits a logistic regression; accuracy on test data is 85%; a chart is shown to managers; the model is served by an API and retrained monthly.

Step 1: Business understanding
Step 2: Data collection
Step 3: Data cleaning
Step 4: Exploration
Step 5: Modelling
Step 6: Evaluation
Step 7: Communication and visualization
Step 8: Deployment (then repeat)

Asked: [7 marks] (Dec 2025) Define Data Science. Explain the Data Science life cycle.

Statistics: Descriptive and Inferential Statistics

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Statistics is the science of collecting, presenting, analysing and interpreting numerical data to support decisions.</mark>

Key points.

  1. The four components are collection, presentation, analysis and interpretation of data.
  2. Descriptive statistics summarise the data in hand using averages, dispersion, tables and graphs.
  3. Inferential statistics draw conclusions about a population from a sample, using estimation and hypothesis testing.
  4. Everyday uses, each with an example:
    • Weather: rainfall records of past years give a 70% chance of rain tomorrow.
    • Medical: a trial gives a drug to 100 patients and a placebo to 100; 80 versus 50 cured shows the drug works.
    • Market survey: 400 of 1000 people sampled prefer brand A, so about 40% of the market prefers A.
    • Finance: average monthly return and its SD of a share measure its risk.
    • Quality control: a factory samples 5 bulbs per hour and stops the line if the mean life falls below the limit.
  5. A sampling distribution is the probability distribution of a statistic (mean, proportion, variance) over all possible samples of a fixed size drawn from one population.
  6. Its standard deviation is the standard error: for the mean it is $\sigma/\sqrt{n}$, for a proportion $\sqrt{PQ/n}$ with $Q=1-P$.
  7. Central limit theorem: for large $n$ (about 30 or more) the sampling distribution of $\bar{x}$ is approximately normal with mean $\mu$ and SD $\sigma/\sqrt{n}$, whatever the shape of the population.
  8. Types: normal, Student's $t$, chi-square and $F$ distributions.
  9. Uses: estimation, confidence intervals and hypothesis testing.

Diagram.

Step 1: Population (size N, mean and SD known or unknown)
Step 2: Draw many random samples of size n
Step 3: Compute the statistic (mean) of each sample
Step 4: Plot these values: the sampling distribution

Example (population 2, 4, 6, 8; samples of size 2 with replacement). $\mu=5$, $\sigma=2.236$. The 16 sample means are 2, 3, 3, 4, 4, 4, 5, 5, 5, 5, 6, 6, 6, 7, 7, 8, so 2 to 8 occur 1, 2, 3, 4, 3, 2, 1 times. Their mean is 5 $=\mu$ and their SD is $1.58=2.236/\sqrt{2}$, the standard error.

Example (PCO/STD, n = 16). Sorted: 5, 12, 18, 30, 30, 32, 34, 37, 38, 39, 41, 42, 43, 46, 46, 52; $\sum x = 545$, $\sum x^2 = 21037$.

Measure Working Value
Mean 545/16 34.06
Median (37+38)/2 37.5
Mode most frequent 30 and 46
Range 52 - 5 47
Std. deviation $\sqrt{\sum x^2/n - \bar{x}^2}$ 12.43

Frequency table: 5, 12, 18, 32, 34, 37, 38, 39, 41, 42, 43, 52 once each; 30 twice; 46 twice. $\sum x^2=21037$, so $\sigma=\sqrt{21037/16-(34.0625)^2}=\sqrt{1314.81-1160.25}=\sqrt{154.56}=12.43$.

Interpretation. About 34 customers use the booth per observation, with wide spread (SD 12.4), so demand is uneven.

Answer frame. Open with the definition; list the four components, then the two branches; give one everyday use each; for the numerical, tabulate mean, median, mode, range and SD, then interpret; for sampling distribution, define, draw the population-to-samples flow, give the small example, standard error, CLT, types and uses.

Asked: [7 marks] (Nov 2022) What are the different components of statistics? How is statistics used in everyday life? Explain with suitable examples. Asked: [7 marks] (Dec 2024) A survey of customers using a PCO/STD booth at a college gate in the last week gave: 52 43 30 38 30 42 12 46 39 37 34 46 32 18 41 5. Summarise the observations. Asked: [7 marks] (Dec 2025) What is a sampling distribution and what are the uses of sampling distributions?

Measures of central tendency: Arithmetic Mean, Median and Mode

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>An average is a single value that represents the whole data and lies near its centre.</mark>

Key points.

  1. Arithmetic mean $\bar{x}=\sum x/n$ (grouped: $\sum fx/N$) uses every value but is pulled by extremes.
  2. Median is the middle value of ordered data, the $(n+1)/2$-th; it is unaffected by extremes, so it suits skewed data such as income.
  3. Mode is the most frequent value; it suits popular sizes (shoe size) but may be absent or multiple, is not rigidly defined and cannot be used in algebra.
  4. Geometric mean suits ratios and growth rates; harmonic mean suits rates such as speed.
  5. Median demerits: it needs ordering of data, ignores the size of extreme values, and cannot be combined algebraically across groups.
  6. No average is best: the choice depends on data type and purpose, so each has its own characteristics.
  7. Grouped median $=L+\frac{N/2-cf}{f}\,h$; mode $=L+\frac{f_1-f_0}{2f_1-f_0-f_2}\,h$.

Requirements of an ideal average: rigidly defined, easy to understand and calculate, based on all values, not unduly affected by extremes, and amenable to further algebraic treatment.

Requirement AM Median Mode GM HM
Rigidly defined Yes Yes No Yes Yes
Easy to calculate Yes Yes Yes No No
Uses all values Yes No No Yes Yes
Not affected by extremes No Yes Yes Partly Partly
Algebraic treatment Yes No No Yes Yes

Example (skewed income, thousand rupees). 20, 22, 25, 25, 28, 30, 32, 35, 40, 300: mean $=557/10=55.7$ but median $=(28+30)/2=29$; one rich person drags the mean above nine of ten incomes, so the median is the fair average.

Example (GM, growth rates). Sales grow by 10%, 20% and 50% in three years, factors 1.10, 1.20, 1.50: $GM=(1.10\times1.20\times1.50)^{1/3}=1.98^{1/3}=1.2557$, an average growth of 25.6% per year; AM 26.7% overstates it.

Example (10, 12, 15, 18, 20, 25, 30). $\sum x=130$, so mean $=130/7=18.57$; median is the 4th value $=18$; every value occurs once, so no mode.

Example (500 bulbs). Cumulative failures 12, 40, 108, 242, 346, 428, 500 give weekly $f$ = 12, 28, 68, 134, 104, 82, 72. Failure in week $k$ is taken at midpoint $x=k-0.5$.

Week f x fx
1 12 0.5 6
2 28 1.5 42
3 68 2.5 170
4 134 3.5 469
5 104 4.5 468
6 82 5.5 451
7 72 6.5 468

$\sum fx=2074$; mean life $=2074/500=$ 4.148 weeks. (Using week numbers 1 to 7 gives 2324/500 = 4.648.)

Answer frame. Open by defining an average; take AM, median, mode, GM, HM one by one with merit, demerit and best use; conclude that the choice depends on the data. For numericals, write the formula, tabulate, box the answer.

Asked: [7 marks] (Nov 2022) "Every average has its own peculiar characteristics. It is difficult to say which average is the best." Explain with examples. Asked: [7 marks] (Nov 2022) 500 bulbs installed together; cumulative failures by end of week 1-7 are 12, 40, 108, 242, 346, 428, 500. Calculate the mean life. Asked: [7 marks] (Dec 2025) Calculate Arithmetic Mean, Median and Mode for 10, 12, 15, 18, 20, 25, 30.

Geometric mean

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ==The geometric mean is the $n$-th root of the product of $n$ positive values: $GM=(x_1x_2\cdots x_n)^{1/n}$.==

Key points.

  1. Computed as $\log GM=\frac{1}{n}\sum\log x$ (grouped: $\frac{1}{N}\sum f\log x$).
  2. It suits ratios, index numbers and growth rates.
  3. It is undefined if any value is zero or negative.
  4. For positive unequal values, $AM > GM > HM$.

Harmonic Mean

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ==The harmonic mean is the reciprocal of the mean of reciprocals: $HM=\dfrac{n}{\sum 1/x}$.==

Key points.

  1. Grouped form: $HM=\dfrac{N}{\sum f/x}$.
  2. It suits rates such as average speed over equal distances.
  3. It is undefined if any value is zero.
  4. Example: 40 and 60 km/h over equal distances give $HM=2/(1/40+1/60)=48$ km/h.

Partition values

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Partition values divide ordered data into equal parts: quartiles into 4, deciles into 10, percentiles into 100.</mark>

Key points.

  1. The median is $Q_2$, $D_5$ and $P_{50}$.
  2. Ungrouped: $Q_i$ is the $i(n+1)/4$-th ordered value.
  3. Grouped: $Q_i=L+\frac{iN/4-cf}{f}\,h$; for $D_i$ use $N/10$, for $P_i$ use $N/100$.
  4. Twenty-five percent of the data lies below $Q_1$ and 75% below $Q_3$.

Measures of dispersion: Dispersion, Range

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Dispersion is the extent to which observations scatter around a central value; range is the largest value minus the smallest.</mark>

Key points.

  1. Range $=L-S$, and coefficient of range $=\frac{L-S}{L+S}$.
  2. Range is simple but uses only two values and is hit by outliers.
  3. Other measures are quartile deviation, mean deviation and standard deviation.
  4. Two data sets can share a mean yet differ in dispersion, so averages alone are not enough.

Quartile Deviation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ==Quartile deviation (semi-interquartile range) is half the distance between the third and first quartiles: $QD=\frac{Q_3-Q_1}{2}$.==

Key points.

  1. It uses the middle 50% of the data.
  2. It is unaffected by extreme values.
  3. Relative form: coefficient of QD $=\frac{Q_3-Q_1}{Q_3+Q_1}$.
  4. It ignores the values in the lowest and highest quarters.

Mean deviation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. ==Mean deviation is the mean of the absolute deviations from an average: $MD=\frac{1}{N}\sum f|x-\bar{x}|$.==

Key points.

  1. It is taken from the mean, median or mode; it is least about the median.
  2. Coefficient of MD $=MD/\text{average}$.
  3. For $a,a+d,\dots,a+2nd$: $N=2n+1$, $\bar{x}=a+nd$, deviations are $rd$ for $r=-n,\dots,n$.

$$MD=\frac{2d}{2n+1}\cdot\frac{n(n+1)}{2}=\frac{n(n+1)}{2n+1}\,d$$

$$\sigma^2=\frac{2d^2}{2n+1}\cdot\frac{n(n+1)(2n+1)}{6}=\frac{n(n+1)}{3}d^2,\quad \sigma=d\sqrt{\tfrac{n(n+1)}{3}}$$

  1. Verify: $\sigma^2>MD^2$ because $(2n+1)^2=4n^2+4n+1>3n(n+1)$, so $\sigma>MD$.

Asked: [7 marks] (Jun 2023) Find the mean deviation from the mean and standard deviation of A.P. $a, a+d, \dots, a+2nd$ and verify that the latter is greater than the former.

Standard Deviation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. ==Standard deviation is the positive square root of the mean of squared deviations from the mean: $\sigma=\sqrt{\frac{1}{N}\sum f(x-\bar{x})^2}$.==

Key points.

  1. Shortcut: $\sigma=\sqrt{\frac{\sum fx^2}{N}-\bar{x}^2}$.
  2. Step deviation: $u=\frac{x-A}{h}$, $\bar{x}=A+h\frac{\sum fu}{N}$, $\sigma=h\sqrt{\frac{\sum fu^2}{N}-\left(\frac{\sum fu}{N}\right)^2}$.
  3. It uses every value and is the most widely used measure.
  4. It does not change if a constant is added to all values, but scales with multiplication.
  5. Coefficient of variation $CV=\frac{\sigma}{\bar{x}}\times100$.

Example (542 members). Take $A=55$, $h=10$.

Class x f u fu fu²
20-30 25 3 -3 -9 27
30-40 35 61 -2 -122 244
40-50 45 132 -1 -132 132
50-60 55 153 0 0 0
60-70 65 140 1 140 140
70-80 75 51 2 102 204
80-90 85 2 3 6 18

$\sum fu=-15$, $\sum fu^2=765$. Mean $=55+10(-15/542)=$ 54.72 years. $\sigma=10\sqrt{765/542-(15/542)^2}=10\sqrt{1.4108}=$ 11.88 years.

Turn-round times, air charter (mid-values 1, 3, ..., 13 h; $f$ = 25, 36, 66, 47, 26, 18, 2; $N=220$): $\sum fx=1250$, $\sum fx^2=8924$. Mean $=1250/220=5.68$ h; $\sigma=\sqrt{8924/220-5.68^2}=\sqrt{8.28}=2.88$ h. Mean + 1 SD $=8.56$ h, mean + 2 SD $=11.44$ h. Advice: quoting 8.56 h means the turn-round is met about 84% of the time; quoting 11.44 h is met about 97.5% of the time, so it is safer for the customer but less competitive. Quote about 8.56 h if bids matter, 11.44 h if punctuality matters.

Proof that $\sigma\ge MD$. Let $p_i=f_i/N$ and $d_i=|x_i-\bar{x}|$. By Cauchy-Schwarz, $\left(\sum p_i d_i\cdot 1\right)^2\le\left(\sum p_id_i^2\right)\left(\sum p_i\right)=\sum p_id_i^2$. The left side is $MD^2$ and the right side is $\sigma^2$. Taking square roots, $\sigma\ge MD$. (Equivalently $\text{Var}(|X-\bar{x}|)\ge0$.)

Answer frame. For the numerical: write the mid-values, choose $A$, tabulate $u,fu,fu^2$, substitute, box mean and SD, then advise. For the proof: state both formulas, apply Cauchy-Schwarz, conclude.

Asked: [7 marks] (Dec 2023, Dec 2024) Calculate the mean and standard deviation for the age distribution of 542 members (3, 61, 132, 153, 140, 51, 2 in classes 20-30 to 80-90); also, for the air-charter turn-round times, calculate mean and SD and advise the quote using mean + 1 SD and mean + 2 SD. Asked: [7 marks] (Jun 2023) Prove that for any discrete distribution standard deviation is not less than mean deviation from mean.

Variance

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. ==Variance is the square of the standard deviation, the mean of squared deviations from the mean: $\sigma^2=\frac{1}{N}\sum f(x-\bar{x})^2$.==

Key points.

  1. Combined mean $\bar{x}=\frac{n_1\bar{x}_1+n_2\bar{x}_2}{n_1+n_2}$.
  2. Combined variance $\sigma^2=\frac{n_1(\sigma_1^2+d_1^2)+n_2(\sigma_2^2+d_2^2)}{n_1+n_2}$, with $d_i=\bar{x}_i-\bar{x}$.
  3. Given $n_1=100,\bar{x}_1=15,\sigma_1=3$; $n=250,\bar{x}=15.6,\sigma^2=13.44$: $n_2=150$, and $\bar{x}_2=\frac{250(15.6)-100(15)}{150}=16$.
  4. $d_1=-0.6$, $d_2=0.4$, so $250(13.44)=100(9+0.36)+150(\sigma_2^2+0.16)$, giving $3360=936+150\sigma_2^2+24$, so $\sigma_2^2=16$ and $\sigma_2=4$.

Asked: [7 marks] (Dec 2023) Sample one has 100 items, mean 15, SD 3; the whole group has 250 items, mean 15.6, SD $\sqrt{13.44}$. Find the SD of the second group.

Coefficient of Dispersion

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A coefficient of dispersion is a unit-free relative measure: absolute dispersion divided by an average.</mark>

Key points.

  1. Range: $\frac{L-S}{L+S}$; QD: $\frac{Q_3-Q_1}{Q_3+Q_1}$; MD: $\frac{MD}{\text{average}}$; SD: $\frac{\sigma}{\bar{x}}$.
  2. Coefficient of variation $=\frac{\sigma}{\bar{x}}\times100$.
  3. It compares series with different units or means; the smaller value is more consistent.

Last-minute revision

  • Data science extracts knowledge from data; life cycle has seven iterative stages, business understanding to deployment.
  • Statistics components: collection, presentation, analysis, interpretation.
  • Mean $\sum fx/N$; median at $(n+1)/2$; mode is the most frequent value.
  • 10, 12, 15, 18, 20, 25, 30: mean 18.57, median 18, no mode.
  • $AM>GM>HM$; $HM=n/\sum(1/x)$.
  • $QD=(Q_3-Q_1)/2$; range $=L-S$.
  • $\sigma=\sqrt{\sum fx^2/N-\bar{x}^2}$; $\sigma\ge MD$.
  • Age table: mean 54.72, SD 11.88; bulbs: 4.148 weeks.
  • AP with $2n+1$ terms: $MD=\frac{n(n+1)}{2n+1}d$, $\sigma=d\sqrt{n(n+1)/3}$.
  • Combined variance uses $n_i(\sigma_i^2+d_i^2)$; answer $\sigma_2=4$.

Memory hooks

  • Life cycle: "B-C-C-E-M-E-D" for Business, Collect, Clean, Explore, Model, Evaluate, Deploy.
  • Averages: mean for symmetry, median for skew, mode for popularity, GM for growth, HM for speed.
  • "SD squares first, then roots": square deviations, average, root.
  • Combined variance: add each group's own variance plus its shift squared.
  • Order: $AM\ge GM\ge HM$ (alphabetical A, G, H).

Coverage checklist

  • Data Science: Introduction: definition, outputs.
  • Data Science Life Cycle: Dec 2025 explain.
  • Statistics: Descriptive and Inferential Statistics: Nov 2022 components, Dec 2024 PCO numerical, Dec 2025 sampling distribution.
  • Measures of central tendency: Arithmetic Mean, Median and Mode: Nov 2022 best average, Nov 2022 bulbs, Dec 2025 mean-median-mode.
  • Geometric mean: formula, uses.
  • Harmonic Mean: formula, uses.
  • Partition values: quartile, decile, percentile formulas.
  • Measures of dispersion: Dispersion, Range: range and coefficient.
  • Quartile Deviation: formula and properties.
  • Mean deviation: Jun 2023 AP derivation.
  • Standard Deviation: Dec 2023 and Dec 2024 age table with turn-round times, Jun 2023 proof.
  • Variance: Dec 2023 combined SD.
  • Coefficient of Dispersion: relative measures.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in