How unit 1 is examined
The unit covers data science and its life cycle, descriptive and inferential statistics, averages, and measures of dispersion; marks come from averages, standard deviation, and grouped-data numericals.
Data Science: Introduction
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Data science is an interdisciplinary field that uses statistics, programming and domain knowledge to extract useful knowledge and decisions from data.</mark>
Key points.
- It combines statistics, computer science and domain expertise.
- It handles structured data (tables) and unstructured data (text, images).
- Typical outputs are reports, predictions and recommendations, for example fraud detection or product suggestions.
Data Science Life Cycle
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>The data science life cycle is the iterative sequence of stages that turns a business problem into a deployed data-driven solution.</mark>
Key points.
- Business understanding fixes the problem and success measure.
- Data collection gathers data from databases, sensors and the web.
- Data cleaning removes missing values, duplicates and errors.
- Exploration uses summaries and plots to find patterns.
- Modelling builds a statistical or machine-learning model.
- Evaluation checks accuracy on unseen data with metrics such as accuracy, precision, recall or RMSE.
- Communication and visualization presents findings to the client through charts and dashboards, so decisions can be taken.
- Deployment puts the model into use, for example behind an API, and monitors it; the cycle is iterative and repeats when performance drops.
- Tools by stage: SQL for collection, pandas for cleaning, matplotlib plots for exploration, scikit-learn for modelling, dashboards for communication, a REST API for deployment.
Example. Predicting customer churn: business goal is to cut churn by 10%; SQL pulls customer records; pandas fills missing ages; plots show churn is high for short-tenure users; scikit-learn fits a logistic regression; accuracy on test data is 85%; a chart is shown to managers; the model is served by an API and retrained monthly.
Step 1: Business understanding
Step 2: Data collection
Step 3: Data cleaning
Step 4: Exploration
Step 5: Modelling
Step 6: Evaluation
Step 7: Communication and visualization
Step 8: Deployment (then repeat)
Asked: [7 marks] (Dec 2025) Define Data Science. Explain the Data Science life cycle.
Statistics: Descriptive and Inferential Statistics
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Statistics is the science of collecting, presenting, analysing and interpreting numerical data to support decisions.</mark>
Key points.
- The four components are collection, presentation, analysis and interpretation of data.
- Descriptive statistics summarise the data in hand using averages, dispersion, tables and graphs.
- Inferential statistics draw conclusions about a population from a sample, using estimation and hypothesis testing.
- Everyday uses, each with an example:
- Weather: rainfall records of past years give a 70% chance of rain tomorrow.
- Medical: a trial gives a drug to 100 patients and a placebo to 100; 80 versus 50 cured shows the drug works.
- Market survey: 400 of 1000 people sampled prefer brand A, so about 40% of the market prefers A.
- Finance: average monthly return and its SD of a share measure its risk.
- Quality control: a factory samples 5 bulbs per hour and stops the line if the mean life falls below the limit.
- A sampling distribution is the probability distribution of a statistic (mean, proportion, variance) over all possible samples of a fixed size drawn from one population.
- Its standard deviation is the standard error: for the mean it is $\sigma/\sqrt{n}$, for a proportion $\sqrt{PQ/n}$ with $Q=1-P$.
- Central limit theorem: for large $n$ (about 30 or more) the sampling distribution of $\bar{x}$ is approximately normal with mean $\mu$ and SD $\sigma/\sqrt{n}$, whatever the shape of the population.
- Types: normal, Student's $t$, chi-square and $F$ distributions.
- Uses: estimation, confidence intervals and hypothesis testing.
Diagram.
Step 1: Population (size N, mean and SD known or unknown)
Step 2: Draw many random samples of size n
Step 3: Compute the statistic (mean) of each sample
Step 4: Plot these values: the sampling distribution
Example (population 2, 4, 6, 8; samples of size 2 with replacement). $\mu=5$, $\sigma=2.236$. The 16 sample means are 2, 3, 3, 4, 4, 4, 5, 5, 5, 5, 6, 6, 6, 7, 7, 8, so 2 to 8 occur 1, 2, 3, 4, 3, 2, 1 times. Their mean is 5 $=\mu$ and their SD is $1.58=2.236/\sqrt{2}$, the standard error.
Example (PCO/STD, n = 16). Sorted: 5, 12, 18, 30, 30, 32, 34, 37, 38, 39, 41, 42, 43, 46, 46, 52; $\sum x = 545$, $\sum x^2 = 21037$.
| Measure | Working | Value |
|---|---|---|
| Mean | 545/16 | 34.06 |
| Median | (37+38)/2 | 37.5 |
| Mode | most frequent | 30 and 46 |
| Range | 52 - 5 | 47 |
| Std. deviation | $\sqrt{\sum x^2/n - \bar{x}^2}$ | 12.43 |
Frequency table: 5, 12, 18, 32, 34, 37, 38, 39, 41, 42, 43, 52 once each; 30 twice; 46 twice. $\sum x^2=21037$, so $\sigma=\sqrt{21037/16-(34.0625)^2}=\sqrt{1314.81-1160.25}=\sqrt{154.56}=12.43$.
Interpretation. About 34 customers use the booth per observation, with wide spread (SD 12.4), so demand is uneven.
Answer frame. Open with the definition; list the four components, then the two branches; give one everyday use each; for the numerical, tabulate mean, median, mode, range and SD, then interpret; for sampling distribution, define, draw the population-to-samples flow, give the small example, standard error, CLT, types and uses.
Asked: [7 marks] (Nov 2022) What are the different components of statistics? How is statistics used in everyday life? Explain with suitable examples. Asked: [7 marks] (Dec 2024) A survey of customers using a PCO/STD booth at a college gate in the last week gave: 52 43 30 38 30 42 12 46 39 37 34 46 32 18 41 5. Summarise the observations. Asked: [7 marks] (Dec 2025) What is a sampling distribution and what are the uses of sampling distributions?
Measures of central tendency: Arithmetic Mean, Median and Mode
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>An average is a single value that represents the whole data and lies near its centre.</mark>
Key points.
- Arithmetic mean $\bar{x}=\sum x/n$ (grouped: $\sum fx/N$) uses every value but is pulled by extremes.
- Median is the middle value of ordered data, the $(n+1)/2$-th; it is unaffected by extremes, so it suits skewed data such as income.
- Mode is the most frequent value; it suits popular sizes (shoe size) but may be absent or multiple, is not rigidly defined and cannot be used in algebra.
- Geometric mean suits ratios and growth rates; harmonic mean suits rates such as speed.
- Median demerits: it needs ordering of data, ignores the size of extreme values, and cannot be combined algebraically across groups.
- No average is best: the choice depends on data type and purpose, so each has its own characteristics.
- Grouped median $=L+\frac{N/2-cf}{f}\,h$; mode $=L+\frac{f_1-f_0}{2f_1-f_0-f_2}\,h$.
Requirements of an ideal average: rigidly defined, easy to understand and calculate, based on all values, not unduly affected by extremes, and amenable to further algebraic treatment.
| Requirement | AM | Median | Mode | GM | HM |
|---|---|---|---|---|---|
| Rigidly defined | Yes | Yes | No | Yes | Yes |
| Easy to calculate | Yes | Yes | Yes | No | No |
| Uses all values | Yes | No | No | Yes | Yes |
| Not affected by extremes | No | Yes | Yes | Partly | Partly |
| Algebraic treatment | Yes | No | No | Yes | Yes |
Example (skewed income, thousand rupees). 20, 22, 25, 25, 28, 30, 32, 35, 40, 300: mean $=557/10=55.7$ but median $=(28+30)/2=29$; one rich person drags the mean above nine of ten incomes, so the median is the fair average.
Example (GM, growth rates). Sales grow by 10%, 20% and 50% in three years, factors 1.10, 1.20, 1.50: $GM=(1.10\times1.20\times1.50)^{1/3}=1.98^{1/3}=1.2557$, an average growth of 25.6% per year; AM 26.7% overstates it.
Example (10, 12, 15, 18, 20, 25, 30). $\sum x=130$, so mean $=130/7=18.57$; median is the 4th value $=18$; every value occurs once, so no mode.
Example (500 bulbs). Cumulative failures 12, 40, 108, 242, 346, 428, 500 give weekly $f$ = 12, 28, 68, 134, 104, 82, 72. Failure in week $k$ is taken at midpoint $x=k-0.5$.
| Week | f | x | fx |
|---|---|---|---|
| 1 | 12 | 0.5 | 6 |
| 2 | 28 | 1.5 | 42 |
| 3 | 68 | 2.5 | 170 |
| 4 | 134 | 3.5 | 469 |
| 5 | 104 | 4.5 | 468 |
| 6 | 82 | 5.5 | 451 |
| 7 | 72 | 6.5 | 468 |
$\sum fx=2074$; mean life $=2074/500=$ 4.148 weeks. (Using week numbers 1 to 7 gives 2324/500 = 4.648.)
Answer frame. Open by defining an average; take AM, median, mode, GM, HM one by one with merit, demerit and best use; conclude that the choice depends on the data. For numericals, write the formula, tabulate, box the answer.
Asked: [7 marks] (Nov 2022) "Every average has its own peculiar characteristics. It is difficult to say which average is the best." Explain with examples. Asked: [7 marks] (Nov 2022) 500 bulbs installed together; cumulative failures by end of week 1-7 are 12, 40, 108, 242, 346, 428, 500. Calculate the mean life. Asked: [7 marks] (Dec 2025) Calculate Arithmetic Mean, Median and Mode for 10, 12, 15, 18, 20, 25, 30.
Geometric mean
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ==The geometric mean is the $n$-th root of the product of $n$ positive values: $GM=(x_1x_2\cdots x_n)^{1/n}$.==
Key points.
- Computed as $\log GM=\frac{1}{n}\sum\log x$ (grouped: $\frac{1}{N}\sum f\log x$).
- It suits ratios, index numbers and growth rates.
- It is undefined if any value is zero or negative.
- For positive unequal values, $AM > GM > HM$.
Harmonic Mean
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ==The harmonic mean is the reciprocal of the mean of reciprocals: $HM=\dfrac{n}{\sum 1/x}$.==
Key points.
- Grouped form: $HM=\dfrac{N}{\sum f/x}$.
- It suits rates such as average speed over equal distances.
- It is undefined if any value is zero.
- Example: 40 and 60 km/h over equal distances give $HM=2/(1/40+1/60)=48$ km/h.
Partition values
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Partition values divide ordered data into equal parts: quartiles into 4, deciles into 10, percentiles into 100.</mark>
Key points.
- The median is $Q_2$, $D_5$ and $P_{50}$.
- Ungrouped: $Q_i$ is the $i(n+1)/4$-th ordered value.
- Grouped: $Q_i=L+\frac{iN/4-cf}{f}\,h$; for $D_i$ use $N/10$, for $P_i$ use $N/100$.
- Twenty-five percent of the data lies below $Q_1$ and 75% below $Q_3$.
Measures of dispersion: Dispersion, Range
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Dispersion is the extent to which observations scatter around a central value; range is the largest value minus the smallest.</mark>
Key points.
- Range $=L-S$, and coefficient of range $=\frac{L-S}{L+S}$.
- Range is simple but uses only two values and is hit by outliers.
- Other measures are quartile deviation, mean deviation and standard deviation.
- Two data sets can share a mean yet differ in dispersion, so averages alone are not enough.
Quartile Deviation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ==Quartile deviation (semi-interquartile range) is half the distance between the third and first quartiles: $QD=\frac{Q_3-Q_1}{2}$.==
Key points.
- It uses the middle 50% of the data.
- It is unaffected by extreme values.
- Relative form: coefficient of QD $=\frac{Q_3-Q_1}{Q_3+Q_1}$.
- It ignores the values in the lowest and highest quarters.
Mean deviation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. ==Mean deviation is the mean of the absolute deviations from an average: $MD=\frac{1}{N}\sum f|x-\bar{x}|$.==
Key points.
- It is taken from the mean, median or mode; it is least about the median.
- Coefficient of MD $=MD/\text{average}$.
- For $a,a+d,\dots,a+2nd$: $N=2n+1$, $\bar{x}=a+nd$, deviations are $rd$ for $r=-n,\dots,n$.
$$MD=\frac{2d}{2n+1}\cdot\frac{n(n+1)}{2}=\frac{n(n+1)}{2n+1}\,d$$
$$\sigma^2=\frac{2d^2}{2n+1}\cdot\frac{n(n+1)(2n+1)}{6}=\frac{n(n+1)}{3}d^2,\quad \sigma=d\sqrt{\tfrac{n(n+1)}{3}}$$
- Verify: $\sigma^2>MD^2$ because $(2n+1)^2=4n^2+4n+1>3n(n+1)$, so $\sigma>MD$.
Asked: [7 marks] (Jun 2023) Find the mean deviation from the mean and standard deviation of A.P. $a, a+d, \dots, a+2nd$ and verify that the latter is greater than the former.
Standard Deviation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. ==Standard deviation is the positive square root of the mean of squared deviations from the mean: $\sigma=\sqrt{\frac{1}{N}\sum f(x-\bar{x})^2}$.==
Key points.
- Shortcut: $\sigma=\sqrt{\frac{\sum fx^2}{N}-\bar{x}^2}$.
- Step deviation: $u=\frac{x-A}{h}$, $\bar{x}=A+h\frac{\sum fu}{N}$, $\sigma=h\sqrt{\frac{\sum fu^2}{N}-\left(\frac{\sum fu}{N}\right)^2}$.
- It uses every value and is the most widely used measure.
- It does not change if a constant is added to all values, but scales with multiplication.
- Coefficient of variation $CV=\frac{\sigma}{\bar{x}}\times100$.
Example (542 members). Take $A=55$, $h=10$.
| Class | x | f | u | fu | fu² |
|---|---|---|---|---|---|
| 20-30 | 25 | 3 | -3 | -9 | 27 |
| 30-40 | 35 | 61 | -2 | -122 | 244 |
| 40-50 | 45 | 132 | -1 | -132 | 132 |
| 50-60 | 55 | 153 | 0 | 0 | 0 |
| 60-70 | 65 | 140 | 1 | 140 | 140 |
| 70-80 | 75 | 51 | 2 | 102 | 204 |
| 80-90 | 85 | 2 | 3 | 6 | 18 |
$\sum fu=-15$, $\sum fu^2=765$. Mean $=55+10(-15/542)=$ 54.72 years. $\sigma=10\sqrt{765/542-(15/542)^2}=10\sqrt{1.4108}=$ 11.88 years.
Turn-round times, air charter (mid-values 1, 3, ..., 13 h; $f$ = 25, 36, 66, 47, 26, 18, 2; $N=220$): $\sum fx=1250$, $\sum fx^2=8924$. Mean $=1250/220=5.68$ h; $\sigma=\sqrt{8924/220-5.68^2}=\sqrt{8.28}=2.88$ h. Mean + 1 SD $=8.56$ h, mean + 2 SD $=11.44$ h. Advice: quoting 8.56 h means the turn-round is met about 84% of the time; quoting 11.44 h is met about 97.5% of the time, so it is safer for the customer but less competitive. Quote about 8.56 h if bids matter, 11.44 h if punctuality matters.
Proof that $\sigma\ge MD$. Let $p_i=f_i/N$ and $d_i=|x_i-\bar{x}|$. By Cauchy-Schwarz, $\left(\sum p_i d_i\cdot 1\right)^2\le\left(\sum p_id_i^2\right)\left(\sum p_i\right)=\sum p_id_i^2$. The left side is $MD^2$ and the right side is $\sigma^2$. Taking square roots, $\sigma\ge MD$. (Equivalently $\text{Var}(|X-\bar{x}|)\ge0$.)
Answer frame. For the numerical: write the mid-values, choose $A$, tabulate $u,fu,fu^2$, substitute, box mean and SD, then advise. For the proof: state both formulas, apply Cauchy-Schwarz, conclude.
Asked: [7 marks] (Dec 2023, Dec 2024) Calculate the mean and standard deviation for the age distribution of 542 members (3, 61, 132, 153, 140, 51, 2 in classes 20-30 to 80-90); also, for the air-charter turn-round times, calculate mean and SD and advise the quote using mean + 1 SD and mean + 2 SD. Asked: [7 marks] (Jun 2023) Prove that for any discrete distribution standard deviation is not less than mean deviation from mean.
Variance
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. ==Variance is the square of the standard deviation, the mean of squared deviations from the mean: $\sigma^2=\frac{1}{N}\sum f(x-\bar{x})^2$.==
Key points.
- Combined mean $\bar{x}=\frac{n_1\bar{x}_1+n_2\bar{x}_2}{n_1+n_2}$.
- Combined variance $\sigma^2=\frac{n_1(\sigma_1^2+d_1^2)+n_2(\sigma_2^2+d_2^2)}{n_1+n_2}$, with $d_i=\bar{x}_i-\bar{x}$.
- Given $n_1=100,\bar{x}_1=15,\sigma_1=3$; $n=250,\bar{x}=15.6,\sigma^2=13.44$: $n_2=150$, and $\bar{x}_2=\frac{250(15.6)-100(15)}{150}=16$.
- $d_1=-0.6$, $d_2=0.4$, so $250(13.44)=100(9+0.36)+150(\sigma_2^2+0.16)$, giving $3360=936+150\sigma_2^2+24$, so $\sigma_2^2=16$ and $\sigma_2=4$.
Asked: [7 marks] (Dec 2023) Sample one has 100 items, mean 15, SD 3; the whole group has 250 items, mean 15.6, SD $\sqrt{13.44}$. Find the SD of the second group.
Coefficient of Dispersion
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A coefficient of dispersion is a unit-free relative measure: absolute dispersion divided by an average.</mark>
Key points.
- Range: $\frac{L-S}{L+S}$; QD: $\frac{Q_3-Q_1}{Q_3+Q_1}$; MD: $\frac{MD}{\text{average}}$; SD: $\frac{\sigma}{\bar{x}}$.
- Coefficient of variation $=\frac{\sigma}{\bar{x}}\times100$.
- It compares series with different units or means; the smaller value is more consistent.
Last-minute revision
- Data science extracts knowledge from data; life cycle has seven iterative stages, business understanding to deployment.
- Statistics components: collection, presentation, analysis, interpretation.
- Mean $\sum fx/N$; median at $(n+1)/2$; mode is the most frequent value.
- 10, 12, 15, 18, 20, 25, 30: mean 18.57, median 18, no mode.
- $AM>GM>HM$; $HM=n/\sum(1/x)$.
- $QD=(Q_3-Q_1)/2$; range $=L-S$.
- $\sigma=\sqrt{\sum fx^2/N-\bar{x}^2}$; $\sigma\ge MD$.
- Age table: mean 54.72, SD 11.88; bulbs: 4.148 weeks.
- AP with $2n+1$ terms: $MD=\frac{n(n+1)}{2n+1}d$, $\sigma=d\sqrt{n(n+1)/3}$.
- Combined variance uses $n_i(\sigma_i^2+d_i^2)$; answer $\sigma_2=4$.
Memory hooks
- Life cycle: "B-C-C-E-M-E-D" for Business, Collect, Clean, Explore, Model, Evaluate, Deploy.
- Averages: mean for symmetry, median for skew, mode for popularity, GM for growth, HM for speed.
- "SD squares first, then roots": square deviations, average, root.
- Combined variance: add each group's own variance plus its shift squared.
- Order: $AM\ge GM\ge HM$ (alphabetical A, G, H).
Coverage checklist
- Data Science: Introduction: definition, outputs.
- Data Science Life Cycle: Dec 2025 explain.
- Statistics: Descriptive and Inferential Statistics: Nov 2022 components, Dec 2024 PCO numerical, Dec 2025 sampling distribution.
- Measures of central tendency: Arithmetic Mean, Median and Mode: Nov 2022 best average, Nov 2022 bulbs, Dec 2025 mean-median-mode.
- Geometric mean: formula, uses.
- Harmonic Mean: formula, uses.
- Partition values: quartile, decile, percentile formulas.
- Measures of dispersion: Dispersion, Range: range and coefficient.
- Quartile Deviation: formula and properties.
- Mean deviation: Jun 2023 AP derivation.
- Standard Deviation: Dec 2023 and Dec 2024 age table with turn-round times, Jun 2023 proof.
- Variance: Dec 2023 combined SD.
- Coefficient of Dispersion: relative measures.