Skip to content
AD-404 · DATA SCIENCE/Quick Revision Short Notes

DATA SCIENCE (AD-404) - Unit 4 Short Notes

How unit 4 is examined

Pandas data handling, Python plotting, model evaluation and ensemble methods; plotting (pie, bar, histogram, scatter) carries the most marks, then gradient boosting and random forests.

Introduction to Pandas

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Pandas is an open-source Python library that provides the Series (1-D) and DataFrame (2-D) labelled data structures for fast loading, cleaning, transforming and analysing tabular data.</mark>

Key points.

  1. Install it with pip install pandas and import it with import pandas as pd.
  2. A Series is a one-dimensional labelled array; a DataFrame is a two-dimensional table of rows and columns.
  3. It reads and writes CSV, Excel, JSON and SQL through read_csv(), read_excel() and to_csv().
  4. Common functions are head(), info(), describe(), sort_values(), groupby() and merge().
import pandas as pd
df = pd.DataFrame({"Name": ["Amit", "Riya"], "Marks": [72, 85]})
print(df["Marks"].mean())   # 78.5

Asked: [7 marks] (Jun 2026) Explain Pandas in Python in detail with example.

Understanding DataFrame

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>A DataFrame is a two-dimensional, size-mutable, labelled table whose columns can hold different data types.</mark>

Key points.

  1. It has rows labelled by an index and columns labelled by names, so data is accessed by label with df["col"], df.loc[] and by position with df.iloc[].
  2. Each column is a Series and may have its own type (int, float, string), so it is heterogeneous.
  3. It handles missing data as NaN and offers isnull(), dropna() and fillna().
  4. It supports filtering, grouping, merging, sorting and reshaping directly.
Basis DataFrame NumPy array
Dimensions 2-D table N-dimensional
Data type Different type per column One dtype for all elements
Labels Row index and column names Integer positions only
Missing data Built-in NaN handling Poor support
Use Data analysis and cleaning Fast numerical maths

Asked: [7 marks] (Jun 2025) Explain the structure and features of a Pandas DataFrame. How is it different from a NumPy array?

Missing Values

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Missing values are absent entries, stored as NaN, that must be removed or filled before analysis.

Key points.

  1. df.isnull().sum() counts missing values per column.
  2. df.dropna() deletes rows (or columns with axis=1) that contain NaN.
  3. df.fillna(value) fills gaps with a constant, the mean, median or mode.
  4. Imputation keeps the data size but dropping is safer when few rows are missing.

Data operation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Data operations combine and summarise DataFrames.

Key points.

  1. groupby("col").mean() splits data into groups and applies an aggregate.
  2. merge(a, b, on="key") joins tables like an SQL join (inner, left, right, outer).
  3. pd.concat([a, b]) stacks frames along rows or columns.
  4. sort_values() and boolean filtering such as df[df.Marks > 60] select and order data.

String Manipulation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>The .str accessor applies vectorised string methods to every element of a Series, skipping NaN safely.</mark>

Key points.

  1. s.str.lower() and s.str.upper() change case, which standardises text before comparing.
  2. s.str.split(" ") breaks each string into a list; expand=True makes columns.
  3. s.str.replace("a", "b") substitutes text and s.str.strip() removes extra spaces.
  4. s.str.contains("AD") returns True/False for filtering and s.str.len() gives lengths.
  5. These operations clean messy data, for example fixing case and spaces in names.
s = pd.Series([" Amit Kumar", "riya SHARMA "])
print(s.str.strip().str.title())   # Amit Kumar, Riya Sharma

Asked: [8 marks] (Jun 2024) Write about String manipulation operations in Pandas.

Regular Expressions and Data learning

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. ==A regular expression is a pattern of characters that matches text; Pandas applies it through str.contains(), str.match(), str.extract() and str.replace(regex=True).==

Key points.

  1. contains(pat) tests whether the pattern occurs anywhere; match(pat) tests only from the start.
  2. extract(r"(pattern)") returns the captured group as a new column.
  3. replace(pat, new, regex=True) cleans matched text, such as removing digits.
  4. Common tokens are \d digit, \w word character, + one or more, [A-Z] capital letter.
s = pd.Series(["AD404-2023", "CS301-2022"])
print(s.str.extract(r"([A-Z]{2}\d{3})"))   # AD404, CS301

Asked: [7 marks] (Jun 2025) Write the logic to perform string operations in Pandas and apply regular expressions to extract specific patterns.

Outlier and Error

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. An outlier is a value far from the rest of the data, caused by error or genuine rarity.

Key points.

  1. The z-score rule flags values with $|z|=|x-\mu|/\sigma>3$.
  2. The IQR rule flags values below $Q1-1.5\,IQR$ or above $Q3+1.5\,IQR$.
  3. Box plots show outliers visually.
  4. Handle them by correcting, capping or removing them.

Visualization tool in Python: Pie Chart, Bar Chart, Histogram, Scatterplots

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>Data visualization in Python means drawing charts with libraries such as Matplotlib so that patterns in data are seen quickly.</mark>

Key points.

  1. Matplotlib (matplotlib.pyplot) is the base library and draws all static charts.
  2. Seaborn is built on Matplotlib and gives attractive statistical plots such as heatmaps and box plots.
  3. Plotly makes interactive, zoomable web charts.
  4. Bokeh makes interactive browser dashboards for large data.
  5. A pie chart, plt.pie(sizes, labels=, autopct="%1.1f%%"), shows parts of a whole.
  6. A bar chart, plt.bar(x, height), compares categories.
  7. A histogram, plt.hist(data, bins=5), shows the distribution of one numeric variable.
  8. A scatter plot, plt.scatter(x, y), shows the relationship between two variables.

Steps (scatter).

Step 1: Install and import matplotlib.pyplot as plt.
Step 2: Prepare the x and y data lists.
Step 3: Call plt.scatter(x, y, color, marker).
Step 4: Add xlabel, ylabel and title.
Step 5: Call plt.show() to display, or plt.savefig("a.png") to save.

Example.

import matplotlib.pyplot as plt
lang = ["Python", "Java", "C"]; use = [50, 30, 20]
plt.pie(use, labels=lang, autopct="%1.1f%%"); plt.title("Usage"); plt.show()
plt.bar(lang, use); plt.xlabel("Language"); plt.ylabel("Students"); plt.show()
plt.hist([1,2,2,3,3,3,4,4,5], bins=5); plt.show()
plt.scatter([1,2,3,4], [2,4,5,8]); plt.show()

The pie shows slices of 50%, 30% and 20%; the bar chart shows three bars of height 50, 30 and 20.

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u4-01" viewBox="0 0 302 338" width="302" height="338" role="img" aria-label="Data drawn as Pie chart, Bar chart, Histogram, Scatter plot"><style>#dsfig-u4-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u4-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u4-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u4-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u4-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u4-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u4-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u4-01 .t{fill:#16181D;font-weight:500}#dsfig-u4-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u4-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u4-01 .dot{fill:#16181D}#dsfig-u4-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u4-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u4-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u4-01 .ah{fill:#454C5A}#dsfig-u4-01 .ah.hi{fill:#2340B8}#dsfig-u4-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u4-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u4-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u4-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u4-01 .e{stroke:#B1B7C3}html.dark #dsfig-u4-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u4-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u4-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u4-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u4-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u4-01 .t{fill:#E6E8ED}html.dark #dsfig-u4-01 .t.inv{fill:#0F1115}html.dark #dsfig-u4-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u4-01 .dot{fill:#E6E8ED}html.dark #dsfig-u4-01 .ann{fill:#8FA3FF}html.dark #dsfig-u4-01 .lbl{fill:#858D9C}html.dark #dsfig-u4-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u4-01 .ah{fill:#B1B7C3}html.dark #dsfig-u4-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u4-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u4-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u4-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M64.1,116.3 L235.5,47.8" marker-end="url(#ah1)"/><path class="e" d="M66,126 L234,126" marker-end="url(#ah1)"/><path class="e" d="M64.1,135.7 L235.5,204.2" marker-end="url(#ah1)"/><path class="e" d="M60.3,142.2 L238.6,284.9" marker-end="url(#ah1)"/><rect class="n" x="15" y="111" width="50" height="30" rx="15"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">Data</text><circle class="n" cx="255" cy="40" r="18"/><text class="t" x="255" y="40" dy=".35em" text-anchor="middle">Pie</text><circle class="n" cx="255" cy="126" r="18"/><text class="t" x="255" y="126" dy=".35em" text-anchor="middle">Bar</text><circle class="n" cx="255" cy="212" r="18"/><text class="t" x="255" y="212" dy=".35em" text-anchor="middle">His</text><circle class="n" cx="255" cy="298" r="18"/><text class="t" x="255" y="298" dy=".35em" text-anchor="middle">Sca</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Data drawn as Pie chart, Bar chart, Histogram, Scatter plot</figcaption></figure>

Chart Use when
Pie Showing percentage share of a whole (few categories)
Bar Comparing values across categories

Answer frame. Open with the definition and name Matplotlib; list the libraries in one line; give the code with the imports, data, plotting call, labels and plt.show(); for pie and bar add the comparison table; for scatter write the five steps; close with one line on interpreting the output.

Asked: [8 marks] (Jun 2023) Discuss visualization tools Pie chart and Bar chart in python. Asked: [7 marks] (Dec 2024) Explain the procedure to create scatter plots using Python. Asked: [6 marks] (Jun 2024) List and explain visualization tools in Python. Asked: [7 marks] (Jun 2025) Illustrate how to create pie charts, bar charts, histograms, and scatter plots in Python using matplotlib or seaborn. Asked: [7 marks] (Jun 2026) Write Python programs to generate Pie Chart and Bar Chart.

Data Analysis

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Data analysis is the process of loading, cleaning, exploring, summarising and interpreting data to reach conclusions.

Key points.

  1. Load data with read_csv() and inspect it with head(), info() and describe().
  2. Clean it by handling missing values, duplicates and outliers.
  3. Summarise with groupby, value_counts() and corr().
  4. Visualise and report the findings.

Performance metrics

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Performance metrics are numbers that measure how well a model's predictions match the actual values.</mark>

Formula. With TP, TN, FP, FN from the confusion matrix:

$$Accuracy=\frac{TP+TN}{TP+TN+FP+FN},\quad Precision=\frac{TP}{TP+FP},\quad Recall=\frac{TP}{TP+FN}$$

$$F1=\frac{2\,P\,R}{P+R},\quad MAE=\frac1n\sum|y-\hat y|,\quad RMSE=\sqrt{\frac1n\sum(y-\hat y)^2}$$

Key points.

  1. Accuracy is the share of correct predictions but misleads on imbalanced data.
  2. Precision tells how many predicted positives are truly positive; use it when false alarms are costly.
  3. Recall tells how many actual positives were found; use it when misses are costly, as in disease detection.
  4. F1 is the harmonic mean of precision and recall and balances both.
  5. MAE and RMSE measure regression error; RMSE punishes large errors more.

Example. TP=40, FP=10, FN=20, TN=30 gives accuracy 70/100=0.70, precision 40/50=0.80, recall 40/60=0.667, F1=0.727. For actual 3,5,7 and predicted 2,5,9, errors are 1,0,2, so MAE=1 and RMSE=$\sqrt{5/3}$=1.29.

Asked: [7 marks] (Jun 2026) Explain different Performance Evaluation Metrics used in Machine Learning.

ROC curve

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. The ROC curve plots the true positive rate against the false positive rate at every classification threshold.

Key points.

  1. $TPR=TP/(TP+FN)$ and $FPR=FP/(FP+TN)$.
  2. A curve nearer the top-left corner is better; the diagonal is random guessing.
  3. AUC is the area under the curve: 1 is perfect and 0.5 is random.

Types of errors

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>A Type I error rejects a true null hypothesis (false positive); a Type II error accepts a false null hypothesis (false negative).</mark>

Basis Type I error Type II error
Meaning False positive False negative
Probability $\alpha$ (significance level) $\beta$
Cause Too lenient a threshold Too little data or power
Consequence Claiming an effect that is absent Missing a real effect
Example Healthy patient told he is ill Ill patient told he is healthy

Key points.

  1. In multiple hypothesis testing, many tests are run together, so the chance of at least one Type I error rises above $\alpha$ (about $1-(1-\alpha)^m$ for $m$ tests).
  2. Corrections such as Bonferroni ($\alpha/m$) control this but raise Type II errors.

Asked: [7 marks] (Jun 2025) Explain the difference between Type I and Type II errors in multiple hypothesis testing with an example.

Overfitting & Under fitting

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Overfitting is when a model learns training data including noise, so training error is low but test error is high; underfitting is when a model is too simple to capture the pattern, so both errors are high.</mark>

Basis Overfitting Underfitting
Problem High variance High bias
Training error Very low High
Test error High High
Cause Too complex model, little data Too simple model
Fix Regularisation, cross-validation, more data More features, complex model

Key points.

  1. It is detected by learning curves: a widening gap between training and validation error means overfitting.
  2. Both high and low errors together on the curves mean underfitting.
  3. Regularisation, pruning and early stopping prevent overfitting.

Asked: [6 marks] (Dec 2024) What is over-fitting and under-fitting concept? Discuss.

Evaluating performance of learning model: Holdout

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Holdout splits the data once into a training set and a test set, commonly 70:30 or 80:20.

Key points.

  1. The model learns on the training part and is scored on the unseen test part.
  2. It is fast but the result depends on the particular split.
  3. train_test_split(X, y, test_size=0.2) does it in scikit-learn.

Random sampling

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Random sampling picks items so that each has an equal chance of selection.

Key points.

  1. It avoids selection bias and makes the sample representative.
  2. It can be with or without replacement; df.sample(n=5) does this in Pandas.
  3. Stratified sampling keeps class proportions.

Cross validation and Bootstrap method

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Cross-validation repeatedly splits the data to test a model; bootstrap resamples with replacement.

Key points.

  1. In k-fold cross-validation the data is split into $k$ folds; each fold is the test set once and the $k$ scores are averaged.
  2. It uses all data for both training and testing, so it is more reliable than holdout.
  3. Bootstrap draws $n$ samples with replacement, leaving about 36.8% of the data out of each sample.

Bagging & boosting

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Bagging trains models in parallel on bootstrap samples and averages them; boosting trains models one after another, each correcting the previous errors.</mark>

Basis Bagging Boosting
Training Parallel, independent Sequential, dependent
Data Bootstrap samples Reweighted, hard cases stressed
Reduces Variance Bias
Combining Vote or average Weighted sum
Example Random Forest AdaBoost, Gradient Boosting

Key points.

  1. Bagging suits high-variance models such as deep trees.
  2. Boosting can overfit noisy data.

Asked: [4 marks] (Dec 2024) Explain the following in detail: Bagging and Boosting

Gradient Boosting

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Gradient boosting is an ensemble method that builds weak learners (small trees) sequentially, each new tree fitting the residual errors of the current model by gradient descent on a loss function.</mark>

Steps.

Step 1: Start with a constant prediction F0 (the mean of y).
Step 2: Compute residuals r = y - F(x), the negative gradient of squared loss.
Step 3: Fit a small tree h to the residuals.
Step 4: Update F = F + learning_rate * h.
Step 5: Repeat for M trees; F is the final model.

Key points.

  1. The learners are weak, shallow trees that are added one at a time.
  2. Residuals are the negative gradient of the loss, so each step moves the model downhill.
  3. A small learning rate reduces overfitting but needs more trees.
  4. Examples are XGBoost, LightGBM and scikit-learn GradientBoostingClassifier.
  5. It is accurate on tabular data, used in ranking, fraud detection and risk scoring, but is slow to train and sensitive to noise.

Example. y = 10, 20, 30; F0 = 20; residuals = -10, 0, 10; with learning rate 0.5 and a tree that fits them exactly, F1 = 15, 20, 25, closer to y.

Answer frame. Open with the definition; write the five steps; show the residual example; then applications and advantages; close with the note that it reduces bias. For the "Data Frame" part of the 2023 question, use the DataFrame section.

Asked: [7 marks] (Jun 2023) Explain the following: i) Data Frame ii) Gradient Boosting Asked: [7 marks] (Dec 2024) Give an overview of Gradient Bootstrap.

Random Forests

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>A random forest is an ensemble of many decision trees, each trained on a bootstrap sample with random feature selection, whose outputs are combined by majority vote (classification) or average (regression).</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u4-02" viewBox="0 0 186 134" width="186" height="134" role="img" aria-label="Trees T1, T2, T3 trained on samples, then combined by voting"><style>#dsfig-u4-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u4-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u4-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u4-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u4-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u4-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u4-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u4-02 .t{fill:#16181D;font-weight:500}#dsfig-u4-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u4-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u4-02 .dot{fill:#16181D}#dsfig-u4-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u4-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u4-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u4-02 .ah{fill:#454C5A}#dsfig-u4-02 .ah.hi{fill:#2340B8}#dsfig-u4-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u4-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u4-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u4-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u4-02 .e{stroke:#B1B7C3}html.dark #dsfig-u4-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u4-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u4-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u4-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u4-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u4-02 .t{fill:#E6E8ED}html.dark #dsfig-u4-02 .t.inv{fill:#0F1115}html.dark #dsfig-u4-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u4-02 .dot{fill:#E6E8ED}html.dark #dsfig-u4-02 .ann{fill:#8FA3FF}html.dark #dsfig-u4-02 .lbl{fill:#858D9C}html.dark #dsfig-u4-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u4-02 .ah{fill:#B1B7C3}html.dark #dsfig-u4-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u4-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u4-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u4-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><line class="e" x1="81" y1="39" x2="31" y2="103"/><line class="e" x1="81" y1="39" x2="81" y2="103"/><line class="e" x1="81" y1="39" x2="131" y2="103"/><rect class="n" x="55" y="24" width="52" height="30" rx="8"/><text class="t" x="81" y="39" dy=".35em" text-anchor="middle">Vote</text><circle class="n" cx="31" cy="103" r="17"/><text class="t" x="31" y="103" dy=".35em" text-anchor="middle">T1</text><circle class="n" cx="81" cy="103" r="17"/><text class="t" x="81" y="103" dy=".35em" text-anchor="middle">T2</text><circle class="n" cx="131" cy="103" r="17"/><text class="t" x="131" y="103" dy=".35em" text-anchor="middle">T3</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Trees T1, T2, T3 trained on samples, then combined by voting</figcaption></figure>

Steps.

Step 1: Draw n bootstrap samples from the training data.
Step 2: Grow a tree on each; at every split consider only a random subset of features.
Step 3: For a new input, get each tree's prediction.
Step 4: Take the majority vote (or average).

Key points.

  1. It uses bagging, so trees are trained independently on different samples.
  2. Random feature selection makes the trees different, which lowers correlation.
  3. Voting cancels individual errors and so reduces variance and overfitting.
  4. Out-of-bag samples give a free error estimate.
  5. It gives feature importance, handles missing and mixed data, but is less interpretable than one tree.

Example. Predict loan approval: three trees vote Yes, Yes, No, so the forest predicts Yes.

Answer frame. Open with the definition; draw the tree diagram; write the four steps; then the vote example; close with advantages (accuracy, less overfitting) and uses (banking, medicine).

Asked: [6 marks] (Jun 2023, Jun 2024) What is random forest? Explain with suitable example.

Committee Machines

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. A committee machine combines the outputs of several models (experts) to get a better decision than any single one.

Key points.

  1. Static committees (bagging, boosting) combine outputs without looking at the input.
  2. Dynamic committees (mixture of experts) use a gating network to weight experts by input.
  3. Combining reduces variance and error.

Last-minute revision

  • Pandas: Series is 1-D, DataFrame is 2-D and heterogeneous; import as pd.
  • .str accessor: lower, upper, split, replace, contains, len, extract.
  • Missing values: isnull, dropna, fillna.
  • Plots: pie, bar, hist, scatter, then plt.show().
  • Libraries: Matplotlib, Seaborn, Plotly, Bokeh.
  • Precision=TP/(TP+FP), Recall=TP/(TP+FN), F1=2PR/(P+R).
  • RMSE=$\sqrt{\text{mean squared error}}$; MAE=mean absolute error.
  • Type I = false positive ($\alpha$); Type II = false negative ($\beta$).
  • Overfitting = high variance; underfitting = high bias.
  • Bagging reduces variance; boosting reduces bias.
  • Gradient boosting fits residuals sequentially; random forest votes over bootstrapped trees.

Memory hooks

  • "PRF": Precision punishes false alarms, Recall punishes misses, F1 balances.
  • Type I = "I see what is not there"; Type II = "I miss what is there".
  • Bagging = Bags in parallel; Boosting = Boots one after another.
  • Overfit = memorises; underfit = never learns.
  • Pie = parts, Bar = compare, Hist = spread, Scatter = relation.

Coverage checklist

  • Introduction to Pandas: Jun 2026 Pandas in detail.
  • understanding DataFrame: Jun 2025 DataFrame vs NumPy.
  • Missing Values: isnull, dropna, fillna.
  • Data operation: groupby, merge, concat.
  • String Manipulation: Jun 2024 string operations.
  • Regular Expressions and Data learning: Jun 2025 regex extraction.
  • Outlier and Error: z-score, IQR.
  • Visualization tool in Python: Representation of Pie Chart, Bar Chart, Histogram, Scatterplots using Python: Jun 2023, Dec 2024, Jun 2024, Jun 2025, Jun 2026.
  • Data Analysis: pipeline steps.
  • performance metrics: Jun 2026 metrics.
  • ROC curve: TPR, FPR, AUC.
  • types of errors: Jun 2025 Type I vs II.
  • Overfitting & Under fitting: Dec 2024.
  • evaluating performance of learning model: Holdout: train-test split.
  • Random sampling: with and without replacement.
  • cross validation and Bootstrap method: k-fold, bootstrap.
  • Bagging & boosting: Dec 2024.
  • Gradient Boosting: Jun 2023, Dec 2024.
  • Random Forests: Jun 2023, Jun 2024.
  • Committee Machines: static and dynamic.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in