Skip to content
CS-605 · Data Analytics Lab/Quick Revision Short Notes

Data Analytics Lab (CS-605) - Unit 1 Short Notes

How unit 1 is examined

This unit covers the statistical and data-handling foundations of data analytics: framework, pre-processing, statistics, probability, Bayes and the central limit theorem, correlation, regression, outliers and visualization. No topic has been asked recently, so learn each definition and its formula.

Basics of data analytic framework

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. A data analytic framework is the ordered set of stages that turns raw data into decisions. <mark>Data analytics is the process of collecting, cleaning, modelling and interpreting data to discover useful patterns and support decisions.</mark>

Key points.

  1. The stages are problem definition, data collection, data cleaning, exploration, modelling, evaluation and deployment.
  2. The process is iterative, so a poor result sends the analyst back to an earlier stage.
  3. Analytics is descriptive (what happened), diagnostic (why), predictive (what will happen) and prescriptive (what to do).
  4. Data may be structured, semi-structured or unstructured.

Data pre-processing

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Data pre-processing is the preparation of raw data, by cleaning and transforming it, so that it is fit for analysis and modelling.</mark>

Key points.

  1. Data cleaning fills or removes missing values, removes duplicates and corrects errors and noise.
  2. Integration merges data from several sources into one consistent set.
  3. Transformation includes normalization, e.g. min-max scaling $x' = \frac{x-x_{min}}{x_{max}-x_{min}}$, and z-score standardization $z=\frac{x-\mu}{\sigma}$.
  4. Reduction lowers the number of attributes or records, e.g. by sampling or dimensionality reduction.
  5. Categorical values are encoded as numbers.

Statistics

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Statistics is the science of collecting, organising, analysing and interpreting numerical data.</mark>

Key points.

  1. Descriptive statistics summarise data, while inferential statistics draw conclusions about a population from a sample.
  2. Central tendency: mean $\bar{x}=\frac{\sum x_i}{n}$, median (middle value) and mode (most frequent value).
  3. Dispersion: variance $\sigma^2=\frac{\sum (x_i-\mu)^2}{n}$, standard deviation $\sigma=\sqrt{\sigma^2}$ and range.
  4. Quartiles split ordered data into four parts, and $IQR=Q_3-Q_1$.

Probability

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ==Probability is a number between 0 and 1 that measures how likely an event is: $P(A)=\frac{\text{favourable outcomes}}{\text{total outcomes}}$.==

Key points.

  1. $P(A)=0$ for an impossible event and $P(A)=1$ for a certain event.
  2. The complement rule is $P(A')=1-P(A)$.
  3. Addition rule: $P(A\cup B)=P(A)+P(B)-P(A\cap B)$.
  4. Conditional probability is $P(A|B)=\frac{P(A\cap B)}{P(B)}$, and for independent events $P(A\cap B)=P(A)P(B)$.

Probability Distribution

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A probability distribution gives every possible value of a random variable together with the probability of each value.</mark>

Key points.

  1. A discrete variable has a probability mass function, and the probabilities sum to 1.
  2. A continuous variable has a density function, and the total area under the curve is 1.
  3. Binomial distribution: $P(X=k)=\binom{n}{k}p^k(1-p)^{n-k}$, with mean $np$.
  4. Poisson distribution: $P(X=k)=\frac{e^{-\lambda}\lambda^k}{k!}$, with mean $\lambda$.
  5. Normal distribution is bell-shaped and symmetric about the mean $\mu$, with spread $\sigma$.

Bayes’ Theorem

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ==Bayes' theorem finds the probability of a cause given an observed effect: $P(A|B)=\frac{P(B|A)\,P(A)}{P(B)}$.==

Key points.

  1. $P(A)$ is the prior, $P(A|B)$ the posterior, $P(B|A)$ the likelihood and $P(B)$ the evidence.
  2. The evidence is $P(B)=\sum_i P(B|A_i)P(A_i)$ over mutually exclusive, exhaustive causes.
  3. It updates a belief when new data arrives.
  4. It is the basis of the Naive Bayes classifier, used in spam filtering.

Central Limit theorem

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>The central limit theorem states that the distribution of sample means approaches a normal distribution as the sample size grows, whatever the shape of the population.</mark>

Key points.

  1. The mean of the sample means equals the population mean $\mu$.
  2. The standard error is $\frac{\sigma}{\sqrt{n}}$.
  3. A sample size of $n\ge 30$ is usually enough.
  4. It justifies confidence intervals and hypothesis tests on non-normal data.

Data Exploration & preparation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Data exploration is the first look at a dataset, using summaries and plots to understand its structure, quality and patterns before modelling.</mark>

Key points.

  1. Check the size, data types, missing values and duplicates.
  2. Univariate analysis uses histograms and box plots, and bivariate analysis uses scatter plots and correlation.
  3. Exploration reveals outliers, skew and relationships between variables.
  4. Preparation then imputes missing values, encodes categories, scales features and selects features.

Concepts of Correlation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Correlation measures the strength and direction of the linear relationship between two variables.</mark>

Key points.

  1. Pearson's coefficient is $r=\frac{\sum (x-\bar{x})(y-\bar{y})}{\sqrt{\sum (x-\bar{x})^2\sum (y-\bar{y})^2}}$.
  2. $-1\le r\le 1$: $+1$ is perfect positive, $-1$ perfect negative and $0$ no linear relation.
  3. Spearman's coefficient measures rank correlation.
  4. Correlation does not imply causation.

Regression

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Regression fits an equation that predicts a dependent variable $y$ from one or more independent variables $x$.</mark>

Key points.

  1. Simple linear regression is $y=a+bx$.
  2. The least-squares slope is $b=\frac{\sum (x-\bar{x})(y-\bar{y})}{\sum (x-\bar{x})^2}$ and the intercept is $a=\bar{y}-b\bar{x}$.
  3. Multiple regression uses several predictors, and logistic regression predicts a class.
  4. $R^2$ gives the fraction of variance in $y$ explained by the model.

Covariance

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ==Covariance measures how two variables vary together: $Cov(X,Y)=\frac{\sum (x_i-\bar{x})(y_i-\bar{y})}{n}$.==

Key points.

  1. Positive covariance means the variables rise together, and negative means one rises as the other falls.
  2. Its size depends on the units, so it is hard to compare.
  3. Correlation is covariance scaled: $r=\frac{Cov(X,Y)}{\sigma_X\sigma_Y}$.
  4. A covariance matrix holds the pairwise covariances of many variables.

Outliers

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>An outlier is a data point that lies far from the rest of the data and does not follow the general pattern.</mark>

Key points.

  1. Causes include measurement error, entry error and genuine rare events.
  2. The IQR rule flags values below $Q_1-1.5\,IQR$ or above $Q_3+1.5\,IQR$.
  3. The z-score rule flags $|z|>3$.
  4. Outliers are removed, capped or transformed, but a genuine one should be kept.

Data visualization

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Data visualization presents data as graphs and charts so that patterns, trends and outliers are seen quickly.</mark>

Key points.

  1. A histogram shows a distribution, and a bar chart compares categories.
  2. A scatter plot shows the relation between two variables, and a line chart shows a trend over time.
  3. A box plot shows the median, quartiles and outliers, and a heat map shows a correlation matrix.
  4. Choose the chart to suit the data and label the axes.

Last-minute revision

  • Analytics stages: define, collect, clean, explore, model, evaluate, deploy.
  • Min-max: $x'=\frac{x-x_{min}}{x_{max}-x_{min}}$; z-score: $z=\frac{x-\mu}{\sigma}$.
  • Variance $\sigma^2=\frac{\sum (x_i-\mu)^2}{n}$; $IQR=Q_3-Q_1$.
  • $P(A\cup B)=P(A)+P(B)-P(A\cap B)$.
  • Binomial mean $np$; Poisson mean $\lambda$.
  • Bayes: $P(A|B)=\frac{P(B|A)P(A)}{P(B)}$.
  • CLT: sample means are normal with standard error $\frac{\sigma}{\sqrt{n}}$.
  • $r$ lies between $-1$ and $1$.
  • Regression: $y=a+bx$, $a=\bar{y}-b\bar{x}$.
  • Outlier fences: $Q_1-1.5\,IQR$ and $Q_3+1.5\,IQR$.

Memory hooks

  • Pre-processing order: Clean, Integrate, Transform, Reduce (CITR).
  • Bayes: posterior = likelihood times prior, over evidence.
  • Correlation is covariance divided by both standard deviations.
  • Box plot dots beyond the whiskers are outliers.

Coverage checklist

  • Basics of data analytic framework: no past questions.
  • Data pre-processing: no past questions.
  • Statistics: no past questions.
  • Probability: no past questions.
  • Probability Distribution: no past questions.
  • Bayes’ Theorem: no past questions.
  • Central Limit theorem: no past questions.
  • Data Exploration & preparation: no past questions.
  • Concepts of Correlation: no past questions.
  • Regression: no past questions.
  • Covariance: no past questions.
  • Outliers: no past questions.
  • Data visualization: no past questions.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in