How unit 1 is examined
This unit covers the statistical and data-handling foundations of data analytics: framework, pre-processing, statistics, probability, Bayes and the central limit theorem, correlation, regression, outliers and visualization. No topic has been asked recently, so learn each definition and its formula.
Basics of data analytic framework
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. A data analytic framework is the ordered set of stages that turns raw data into decisions. <mark>Data analytics is the process of collecting, cleaning, modelling and interpreting data to discover useful patterns and support decisions.</mark>
Key points.
- The stages are problem definition, data collection, data cleaning, exploration, modelling, evaluation and deployment.
- The process is iterative, so a poor result sends the analyst back to an earlier stage.
- Analytics is descriptive (what happened), diagnostic (why), predictive (what will happen) and prescriptive (what to do).
- Data may be structured, semi-structured or unstructured.
Data pre-processing
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Data pre-processing is the preparation of raw data, by cleaning and transforming it, so that it is fit for analysis and modelling.</mark>
Key points.
- Data cleaning fills or removes missing values, removes duplicates and corrects errors and noise.
- Integration merges data from several sources into one consistent set.
- Transformation includes normalization, e.g. min-max scaling $x' = \frac{x-x_{min}}{x_{max}-x_{min}}$, and z-score standardization $z=\frac{x-\mu}{\sigma}$.
- Reduction lowers the number of attributes or records, e.g. by sampling or dimensionality reduction.
- Categorical values are encoded as numbers.
Statistics
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Statistics is the science of collecting, organising, analysing and interpreting numerical data.</mark>
Key points.
- Descriptive statistics summarise data, while inferential statistics draw conclusions about a population from a sample.
- Central tendency: mean $\bar{x}=\frac{\sum x_i}{n}$, median (middle value) and mode (most frequent value).
- Dispersion: variance $\sigma^2=\frac{\sum (x_i-\mu)^2}{n}$, standard deviation $\sigma=\sqrt{\sigma^2}$ and range.
- Quartiles split ordered data into four parts, and $IQR=Q_3-Q_1$.
Probability
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ==Probability is a number between 0 and 1 that measures how likely an event is: $P(A)=\frac{\text{favourable outcomes}}{\text{total outcomes}}$.==
Key points.
- $P(A)=0$ for an impossible event and $P(A)=1$ for a certain event.
- The complement rule is $P(A')=1-P(A)$.
- Addition rule: $P(A\cup B)=P(A)+P(B)-P(A\cap B)$.
- Conditional probability is $P(A|B)=\frac{P(A\cap B)}{P(B)}$, and for independent events $P(A\cap B)=P(A)P(B)$.
Probability Distribution
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A probability distribution gives every possible value of a random variable together with the probability of each value.</mark>
Key points.
- A discrete variable has a probability mass function, and the probabilities sum to 1.
- A continuous variable has a density function, and the total area under the curve is 1.
- Binomial distribution: $P(X=k)=\binom{n}{k}p^k(1-p)^{n-k}$, with mean $np$.
- Poisson distribution: $P(X=k)=\frac{e^{-\lambda}\lambda^k}{k!}$, with mean $\lambda$.
- Normal distribution is bell-shaped and symmetric about the mean $\mu$, with spread $\sigma$.
Bayes’ Theorem
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ==Bayes' theorem finds the probability of a cause given an observed effect: $P(A|B)=\frac{P(B|A)\,P(A)}{P(B)}$.==
Key points.
- $P(A)$ is the prior, $P(A|B)$ the posterior, $P(B|A)$ the likelihood and $P(B)$ the evidence.
- The evidence is $P(B)=\sum_i P(B|A_i)P(A_i)$ over mutually exclusive, exhaustive causes.
- It updates a belief when new data arrives.
- It is the basis of the Naive Bayes classifier, used in spam filtering.
Central Limit theorem
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>The central limit theorem states that the distribution of sample means approaches a normal distribution as the sample size grows, whatever the shape of the population.</mark>
Key points.
- The mean of the sample means equals the population mean $\mu$.
- The standard error is $\frac{\sigma}{\sqrt{n}}$.
- A sample size of $n\ge 30$ is usually enough.
- It justifies confidence intervals and hypothesis tests on non-normal data.
Data Exploration & preparation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Data exploration is the first look at a dataset, using summaries and plots to understand its structure, quality and patterns before modelling.</mark>
Key points.
- Check the size, data types, missing values and duplicates.
- Univariate analysis uses histograms and box plots, and bivariate analysis uses scatter plots and correlation.
- Exploration reveals outliers, skew and relationships between variables.
- Preparation then imputes missing values, encodes categories, scales features and selects features.
Concepts of Correlation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Correlation measures the strength and direction of the linear relationship between two variables.</mark>
Key points.
- Pearson's coefficient is $r=\frac{\sum (x-\bar{x})(y-\bar{y})}{\sqrt{\sum (x-\bar{x})^2\sum (y-\bar{y})^2}}$.
- $-1\le r\le 1$: $+1$ is perfect positive, $-1$ perfect negative and $0$ no linear relation.
- Spearman's coefficient measures rank correlation.
- Correlation does not imply causation.
Regression
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Regression fits an equation that predicts a dependent variable $y$ from one or more independent variables $x$.</mark>
Key points.
- Simple linear regression is $y=a+bx$.
- The least-squares slope is $b=\frac{\sum (x-\bar{x})(y-\bar{y})}{\sum (x-\bar{x})^2}$ and the intercept is $a=\bar{y}-b\bar{x}$.
- Multiple regression uses several predictors, and logistic regression predicts a class.
- $R^2$ gives the fraction of variance in $y$ explained by the model.
Covariance
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ==Covariance measures how two variables vary together: $Cov(X,Y)=\frac{\sum (x_i-\bar{x})(y_i-\bar{y})}{n}$.==
Key points.
- Positive covariance means the variables rise together, and negative means one rises as the other falls.
- Its size depends on the units, so it is hard to compare.
- Correlation is covariance scaled: $r=\frac{Cov(X,Y)}{\sigma_X\sigma_Y}$.
- A covariance matrix holds the pairwise covariances of many variables.
Outliers
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>An outlier is a data point that lies far from the rest of the data and does not follow the general pattern.</mark>
Key points.
- Causes include measurement error, entry error and genuine rare events.
- The IQR rule flags values below $Q_1-1.5\,IQR$ or above $Q_3+1.5\,IQR$.
- The z-score rule flags $|z|>3$.
- Outliers are removed, capped or transformed, but a genuine one should be kept.
Data visualization
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Data visualization presents data as graphs and charts so that patterns, trends and outliers are seen quickly.</mark>
Key points.
- A histogram shows a distribution, and a bar chart compares categories.
- A scatter plot shows the relation between two variables, and a line chart shows a trend over time.
- A box plot shows the median, quartiles and outliers, and a heat map shows a correlation matrix.
- Choose the chart to suit the data and label the axes.
Last-minute revision
- Analytics stages: define, collect, clean, explore, model, evaluate, deploy.
- Min-max: $x'=\frac{x-x_{min}}{x_{max}-x_{min}}$; z-score: $z=\frac{x-\mu}{\sigma}$.
- Variance $\sigma^2=\frac{\sum (x_i-\mu)^2}{n}$; $IQR=Q_3-Q_1$.
- $P(A\cup B)=P(A)+P(B)-P(A\cap B)$.
- Binomial mean $np$; Poisson mean $\lambda$.
- Bayes: $P(A|B)=\frac{P(B|A)P(A)}{P(B)}$.
- CLT: sample means are normal with standard error $\frac{\sigma}{\sqrt{n}}$.
- $r$ lies between $-1$ and $1$.
- Regression: $y=a+bx$, $a=\bar{y}-b\bar{x}$.
- Outlier fences: $Q_1-1.5\,IQR$ and $Q_3+1.5\,IQR$.
Memory hooks
- Pre-processing order: Clean, Integrate, Transform, Reduce (CITR).
- Bayes: posterior = likelihood times prior, over evidence.
- Correlation is covariance divided by both standard deviations.
- Box plot dots beyond the whiskers are outliers.
Coverage checklist
- Basics of data analytic framework: no past questions.
- Data pre-processing: no past questions.
- Statistics: no past questions.
- Probability: no past questions.
- Probability Distribution: no past questions.
- Bayes’ Theorem: no past questions.
- Central Limit theorem: no past questions.
- Data Exploration & preparation: no past questions.
- Concepts of Correlation: no past questions.
- Regression: no past questions.
- Covariance: no past questions.
- Outliers: no past questions.
- Data visualization: no past questions.