How unit 4 is examined
Pandas data handling, Python plotting, model evaluation and ensemble methods; plotting (pie, bar, histogram, scatter) carries the most marks, then gradient boosting and random forests.
Introduction to Pandas
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Pandas is an open-source Python library that provides the Series (1-D) and DataFrame (2-D) labelled data structures for fast loading, cleaning, transforming and analysing tabular data.</mark>
Key points.
- Install it with
pip install pandasand import it withimport pandas as pd. - A Series is a one-dimensional labelled array; a DataFrame is a two-dimensional table of rows and columns.
- It reads and writes CSV, Excel, JSON and SQL through
read_csv(),read_excel()andto_csv(). - Common functions are
head(),info(),describe(),sort_values(),groupby()andmerge().
import pandas as pd
df = pd.DataFrame({"Name": ["Amit", "Riya"], "Marks": [72, 85]})
print(df["Marks"].mean()) # 78.5
Asked: [7 marks] (Jun 2026) Explain Pandas in Python in detail with example.
Understanding DataFrame
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>A DataFrame is a two-dimensional, size-mutable, labelled table whose columns can hold different data types.</mark>
Key points.
- It has rows labelled by an index and columns labelled by names, so data is accessed by label with
df["col"],df.loc[]and by position withdf.iloc[]. - Each column is a Series and may have its own type (int, float, string), so it is heterogeneous.
- It handles missing data as NaN and offers
isnull(),dropna()andfillna(). - It supports filtering, grouping, merging, sorting and reshaping directly.
| Basis | DataFrame | NumPy array |
|---|---|---|
| Dimensions | 2-D table | N-dimensional |
| Data type | Different type per column | One dtype for all elements |
| Labels | Row index and column names | Integer positions only |
| Missing data | Built-in NaN handling | Poor support |
| Use | Data analysis and cleaning | Fast numerical maths |
Asked: [7 marks] (Jun 2025) Explain the structure and features of a Pandas DataFrame. How is it different from a NumPy array?
Missing Values
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Missing values are absent entries, stored as NaN, that must be removed or filled before analysis.
Key points.
df.isnull().sum()counts missing values per column.df.dropna()deletes rows (or columns withaxis=1) that contain NaN.df.fillna(value)fills gaps with a constant, the mean, median or mode.- Imputation keeps the data size but dropping is safer when few rows are missing.
Data operation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Data operations combine and summarise DataFrames.
Key points.
groupby("col").mean()splits data into groups and applies an aggregate.merge(a, b, on="key")joins tables like an SQL join (inner, left, right, outer).pd.concat([a, b])stacks frames along rows or columns.sort_values()and boolean filtering such asdf[df.Marks > 60]select and order data.
String Manipulation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>The .str accessor applies vectorised string methods to every element of a Series, skipping NaN safely.</mark>
Key points.
s.str.lower()ands.str.upper()change case, which standardises text before comparing.s.str.split(" ")breaks each string into a list;expand=Truemakes columns.s.str.replace("a", "b")substitutes text ands.str.strip()removes extra spaces.s.str.contains("AD")returns True/False for filtering ands.str.len()gives lengths.- These operations clean messy data, for example fixing case and spaces in names.
s = pd.Series([" Amit Kumar", "riya SHARMA "])
print(s.str.strip().str.title()) # Amit Kumar, Riya Sharma
Asked: [8 marks] (Jun 2024) Write about String manipulation operations in Pandas.
Regular Expressions and Data learning
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. ==A regular expression is a pattern of characters that matches text; Pandas applies it through str.contains(), str.match(), str.extract() and str.replace(regex=True).==
Key points.
contains(pat)tests whether the pattern occurs anywhere;match(pat)tests only from the start.extract(r"(pattern)")returns the captured group as a new column.replace(pat, new, regex=True)cleans matched text, such as removing digits.- Common tokens are
\ddigit,\wword character,+one or more,[A-Z]capital letter.
s = pd.Series(["AD404-2023", "CS301-2022"])
print(s.str.extract(r"([A-Z]{2}\d{3})")) # AD404, CS301
Asked: [7 marks] (Jun 2025) Write the logic to perform string operations in Pandas and apply regular expressions to extract specific patterns.
Outlier and Error
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. An outlier is a value far from the rest of the data, caused by error or genuine rarity.
Key points.
- The z-score rule flags values with $|z|=|x-\mu|/\sigma>3$.
- The IQR rule flags values below $Q1-1.5\,IQR$ or above $Q3+1.5\,IQR$.
- Box plots show outliers visually.
- Handle them by correcting, capping or removing them.
Visualization tool in Python: Pie Chart, Bar Chart, Histogram, Scatterplots
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>Data visualization in Python means drawing charts with libraries such as Matplotlib so that patterns in data are seen quickly.</mark>
Key points.
- Matplotlib (
matplotlib.pyplot) is the base library and draws all static charts. - Seaborn is built on Matplotlib and gives attractive statistical plots such as heatmaps and box plots.
- Plotly makes interactive, zoomable web charts.
- Bokeh makes interactive browser dashboards for large data.
- A pie chart,
plt.pie(sizes, labels=, autopct="%1.1f%%"), shows parts of a whole. - A bar chart,
plt.bar(x, height), compares categories. - A histogram,
plt.hist(data, bins=5), shows the distribution of one numeric variable. - A scatter plot,
plt.scatter(x, y), shows the relationship between two variables.
Steps (scatter).
Step 1: Install and import matplotlib.pyplot as plt.
Step 2: Prepare the x and y data lists.
Step 3: Call plt.scatter(x, y, color, marker).
Step 4: Add xlabel, ylabel and title.
Step 5: Call plt.show() to display, or plt.savefig("a.png") to save.
Example.
import matplotlib.pyplot as plt
lang = ["Python", "Java", "C"]; use = [50, 30, 20]
plt.pie(use, labels=lang, autopct="%1.1f%%"); plt.title("Usage"); plt.show()
plt.bar(lang, use); plt.xlabel("Language"); plt.ylabel("Students"); plt.show()
plt.hist([1,2,2,3,3,3,4,4,5], bins=5); plt.show()
plt.scatter([1,2,3,4], [2,4,5,8]); plt.show()
The pie shows slices of 50%, 30% and 20%; the bar chart shows three bars of height 50, 30 and 20.
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u4-01" viewBox="0 0 302 338" width="302" height="338" role="img" aria-label="Data drawn as Pie chart, Bar chart, Histogram, Scatter plot"><style>#dsfig-u4-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u4-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u4-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u4-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u4-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u4-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u4-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u4-01 .t{fill:#16181D;font-weight:500}#dsfig-u4-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u4-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u4-01 .dot{fill:#16181D}#dsfig-u4-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u4-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u4-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u4-01 .ah{fill:#454C5A}#dsfig-u4-01 .ah.hi{fill:#2340B8}#dsfig-u4-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u4-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u4-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u4-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u4-01 .e{stroke:#B1B7C3}html.dark #dsfig-u4-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u4-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u4-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u4-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u4-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u4-01 .t{fill:#E6E8ED}html.dark #dsfig-u4-01 .t.inv{fill:#0F1115}html.dark #dsfig-u4-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u4-01 .dot{fill:#E6E8ED}html.dark #dsfig-u4-01 .ann{fill:#8FA3FF}html.dark #dsfig-u4-01 .lbl{fill:#858D9C}html.dark #dsfig-u4-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u4-01 .ah{fill:#B1B7C3}html.dark #dsfig-u4-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u4-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u4-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u4-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M64.1,116.3 L235.5,47.8" marker-end="url(#ah1)"/><path class="e" d="M66,126 L234,126" marker-end="url(#ah1)"/><path class="e" d="M64.1,135.7 L235.5,204.2" marker-end="url(#ah1)"/><path class="e" d="M60.3,142.2 L238.6,284.9" marker-end="url(#ah1)"/><rect class="n" x="15" y="111" width="50" height="30" rx="15"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">Data</text><circle class="n" cx="255" cy="40" r="18"/><text class="t" x="255" y="40" dy=".35em" text-anchor="middle">Pie</text><circle class="n" cx="255" cy="126" r="18"/><text class="t" x="255" y="126" dy=".35em" text-anchor="middle">Bar</text><circle class="n" cx="255" cy="212" r="18"/><text class="t" x="255" y="212" dy=".35em" text-anchor="middle">His</text><circle class="n" cx="255" cy="298" r="18"/><text class="t" x="255" y="298" dy=".35em" text-anchor="middle">Sca</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Data drawn as Pie chart, Bar chart, Histogram, Scatter plot</figcaption></figure>
| Chart | Use when |
|---|---|
| Pie | Showing percentage share of a whole (few categories) |
| Bar | Comparing values across categories |
Answer frame. Open with the definition and name Matplotlib; list the libraries in one line; give the code with the imports, data, plotting call, labels and plt.show(); for pie and bar add the comparison table; for scatter write the five steps; close with one line on interpreting the output.
Asked: [8 marks] (Jun 2023) Discuss visualization tools Pie chart and Bar chart in python. Asked: [7 marks] (Dec 2024) Explain the procedure to create scatter plots using Python. Asked: [6 marks] (Jun 2024) List and explain visualization tools in Python. Asked: [7 marks] (Jun 2025) Illustrate how to create pie charts, bar charts, histograms, and scatter plots in Python using matplotlib or seaborn. Asked: [7 marks] (Jun 2026) Write Python programs to generate Pie Chart and Bar Chart.
Data Analysis
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Data analysis is the process of loading, cleaning, exploring, summarising and interpreting data to reach conclusions.
Key points.
- Load data with
read_csv()and inspect it withhead(),info()anddescribe(). - Clean it by handling missing values, duplicates and outliers.
- Summarise with
groupby,value_counts()andcorr(). - Visualise and report the findings.
Performance metrics
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Performance metrics are numbers that measure how well a model's predictions match the actual values.</mark>
Formula. With TP, TN, FP, FN from the confusion matrix:
$$Accuracy=\frac{TP+TN}{TP+TN+FP+FN},\quad Precision=\frac{TP}{TP+FP},\quad Recall=\frac{TP}{TP+FN}$$
$$F1=\frac{2\,P\,R}{P+R},\quad MAE=\frac1n\sum|y-\hat y|,\quad RMSE=\sqrt{\frac1n\sum(y-\hat y)^2}$$
Key points.
- Accuracy is the share of correct predictions but misleads on imbalanced data.
- Precision tells how many predicted positives are truly positive; use it when false alarms are costly.
- Recall tells how many actual positives were found; use it when misses are costly, as in disease detection.
- F1 is the harmonic mean of precision and recall and balances both.
- MAE and RMSE measure regression error; RMSE punishes large errors more.
Example. TP=40, FP=10, FN=20, TN=30 gives accuracy 70/100=0.70, precision 40/50=0.80, recall 40/60=0.667, F1=0.727. For actual 3,5,7 and predicted 2,5,9, errors are 1,0,2, so MAE=1 and RMSE=$\sqrt{5/3}$=1.29.
Asked: [7 marks] (Jun 2026) Explain different Performance Evaluation Metrics used in Machine Learning.
ROC curve
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. The ROC curve plots the true positive rate against the false positive rate at every classification threshold.
Key points.
- $TPR=TP/(TP+FN)$ and $FPR=FP/(FP+TN)$.
- A curve nearer the top-left corner is better; the diagonal is random guessing.
- AUC is the area under the curve: 1 is perfect and 0.5 is random.
Types of errors
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>A Type I error rejects a true null hypothesis (false positive); a Type II error accepts a false null hypothesis (false negative).</mark>
| Basis | Type I error | Type II error |
|---|---|---|
| Meaning | False positive | False negative |
| Probability | $\alpha$ (significance level) | $\beta$ |
| Cause | Too lenient a threshold | Too little data or power |
| Consequence | Claiming an effect that is absent | Missing a real effect |
| Example | Healthy patient told he is ill | Ill patient told he is healthy |
Key points.
- In multiple hypothesis testing, many tests are run together, so the chance of at least one Type I error rises above $\alpha$ (about $1-(1-\alpha)^m$ for $m$ tests).
- Corrections such as Bonferroni ($\alpha/m$) control this but raise Type II errors.
Asked: [7 marks] (Jun 2025) Explain the difference between Type I and Type II errors in multiple hypothesis testing with an example.
Overfitting & Under fitting
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Overfitting is when a model learns training data including noise, so training error is low but test error is high; underfitting is when a model is too simple to capture the pattern, so both errors are high.</mark>
| Basis | Overfitting | Underfitting |
|---|---|---|
| Problem | High variance | High bias |
| Training error | Very low | High |
| Test error | High | High |
| Cause | Too complex model, little data | Too simple model |
| Fix | Regularisation, cross-validation, more data | More features, complex model |
Key points.
- It is detected by learning curves: a widening gap between training and validation error means overfitting.
- Both high and low errors together on the curves mean underfitting.
- Regularisation, pruning and early stopping prevent overfitting.
Asked: [6 marks] (Dec 2024) What is over-fitting and under-fitting concept? Discuss.
Evaluating performance of learning model: Holdout
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Holdout splits the data once into a training set and a test set, commonly 70:30 or 80:20.
Key points.
- The model learns on the training part and is scored on the unseen test part.
- It is fast but the result depends on the particular split.
train_test_split(X, y, test_size=0.2)does it in scikit-learn.
Random sampling
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Random sampling picks items so that each has an equal chance of selection.
Key points.
- It avoids selection bias and makes the sample representative.
- It can be with or without replacement;
df.sample(n=5)does this in Pandas. - Stratified sampling keeps class proportions.
Cross validation and Bootstrap method
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Cross-validation repeatedly splits the data to test a model; bootstrap resamples with replacement.
Key points.
- In k-fold cross-validation the data is split into $k$ folds; each fold is the test set once and the $k$ scores are averaged.
- It uses all data for both training and testing, so it is more reliable than holdout.
- Bootstrap draws $n$ samples with replacement, leaving about 36.8% of the data out of each sample.
Bagging & boosting
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Bagging trains models in parallel on bootstrap samples and averages them; boosting trains models one after another, each correcting the previous errors.</mark>
| Basis | Bagging | Boosting |
|---|---|---|
| Training | Parallel, independent | Sequential, dependent |
| Data | Bootstrap samples | Reweighted, hard cases stressed |
| Reduces | Variance | Bias |
| Combining | Vote or average | Weighted sum |
| Example | Random Forest | AdaBoost, Gradient Boosting |
Key points.
- Bagging suits high-variance models such as deep trees.
- Boosting can overfit noisy data.
Asked: [4 marks] (Dec 2024) Explain the following in detail: Bagging and Boosting
Gradient Boosting
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Gradient boosting is an ensemble method that builds weak learners (small trees) sequentially, each new tree fitting the residual errors of the current model by gradient descent on a loss function.</mark>
Steps.
Step 1: Start with a constant prediction F0 (the mean of y).
Step 2: Compute residuals r = y - F(x), the negative gradient of squared loss.
Step 3: Fit a small tree h to the residuals.
Step 4: Update F = F + learning_rate * h.
Step 5: Repeat for M trees; F is the final model.
Key points.
- The learners are weak, shallow trees that are added one at a time.
- Residuals are the negative gradient of the loss, so each step moves the model downhill.
- A small learning rate reduces overfitting but needs more trees.
- Examples are XGBoost, LightGBM and scikit-learn
GradientBoostingClassifier. - It is accurate on tabular data, used in ranking, fraud detection and risk scoring, but is slow to train and sensitive to noise.
Example. y = 10, 20, 30; F0 = 20; residuals = -10, 0, 10; with learning rate 0.5 and a tree that fits them exactly, F1 = 15, 20, 25, closer to y.
Answer frame. Open with the definition; write the five steps; show the residual example; then applications and advantages; close with the note that it reduces bias. For the "Data Frame" part of the 2023 question, use the DataFrame section.
Asked: [7 marks] (Jun 2023) Explain the following: i) Data Frame ii) Gradient Boosting Asked: [7 marks] (Dec 2024) Give an overview of Gradient Bootstrap.
Random Forests
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>A random forest is an ensemble of many decision trees, each trained on a bootstrap sample with random feature selection, whose outputs are combined by majority vote (classification) or average (regression).</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u4-02" viewBox="0 0 186 134" width="186" height="134" role="img" aria-label="Trees T1, T2, T3 trained on samples, then combined by voting"><style>#dsfig-u4-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u4-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u4-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u4-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u4-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u4-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u4-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u4-02 .t{fill:#16181D;font-weight:500}#dsfig-u4-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u4-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u4-02 .dot{fill:#16181D}#dsfig-u4-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u4-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u4-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u4-02 .ah{fill:#454C5A}#dsfig-u4-02 .ah.hi{fill:#2340B8}#dsfig-u4-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u4-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u4-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u4-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u4-02 .e{stroke:#B1B7C3}html.dark #dsfig-u4-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u4-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u4-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u4-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u4-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u4-02 .t{fill:#E6E8ED}html.dark #dsfig-u4-02 .t.inv{fill:#0F1115}html.dark #dsfig-u4-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u4-02 .dot{fill:#E6E8ED}html.dark #dsfig-u4-02 .ann{fill:#8FA3FF}html.dark #dsfig-u4-02 .lbl{fill:#858D9C}html.dark #dsfig-u4-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u4-02 .ah{fill:#B1B7C3}html.dark #dsfig-u4-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u4-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u4-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u4-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><line class="e" x1="81" y1="39" x2="31" y2="103"/><line class="e" x1="81" y1="39" x2="81" y2="103"/><line class="e" x1="81" y1="39" x2="131" y2="103"/><rect class="n" x="55" y="24" width="52" height="30" rx="8"/><text class="t" x="81" y="39" dy=".35em" text-anchor="middle">Vote</text><circle class="n" cx="31" cy="103" r="17"/><text class="t" x="31" y="103" dy=".35em" text-anchor="middle">T1</text><circle class="n" cx="81" cy="103" r="17"/><text class="t" x="81" y="103" dy=".35em" text-anchor="middle">T2</text><circle class="n" cx="131" cy="103" r="17"/><text class="t" x="131" y="103" dy=".35em" text-anchor="middle">T3</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Trees T1, T2, T3 trained on samples, then combined by voting</figcaption></figure>
Steps.
Step 1: Draw n bootstrap samples from the training data.
Step 2: Grow a tree on each; at every split consider only a random subset of features.
Step 3: For a new input, get each tree's prediction.
Step 4: Take the majority vote (or average).
Key points.
- It uses bagging, so trees are trained independently on different samples.
- Random feature selection makes the trees different, which lowers correlation.
- Voting cancels individual errors and so reduces variance and overfitting.
- Out-of-bag samples give a free error estimate.
- It gives feature importance, handles missing and mixed data, but is less interpretable than one tree.
Example. Predict loan approval: three trees vote Yes, Yes, No, so the forest predicts Yes.
Answer frame. Open with the definition; draw the tree diagram; write the four steps; then the vote example; close with advantages (accuracy, less overfitting) and uses (banking, medicine).
Asked: [6 marks] (Jun 2023, Jun 2024) What is random forest? Explain with suitable example.
Committee Machines
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. A committee machine combines the outputs of several models (experts) to get a better decision than any single one.
Key points.
- Static committees (bagging, boosting) combine outputs without looking at the input.
- Dynamic committees (mixture of experts) use a gating network to weight experts by input.
- Combining reduces variance and error.
Last-minute revision
- Pandas: Series is 1-D, DataFrame is 2-D and heterogeneous; import as
pd. .straccessor: lower, upper, split, replace, contains, len, extract.- Missing values:
isnull,dropna,fillna. - Plots:
pie,bar,hist,scatter, thenplt.show(). - Libraries: Matplotlib, Seaborn, Plotly, Bokeh.
- Precision=TP/(TP+FP), Recall=TP/(TP+FN), F1=2PR/(P+R).
- RMSE=$\sqrt{\text{mean squared error}}$; MAE=mean absolute error.
- Type I = false positive ($\alpha$); Type II = false negative ($\beta$).
- Overfitting = high variance; underfitting = high bias.
- Bagging reduces variance; boosting reduces bias.
- Gradient boosting fits residuals sequentially; random forest votes over bootstrapped trees.
Memory hooks
- "PRF": Precision punishes false alarms, Recall punishes misses, F1 balances.
- Type I = "I see what is not there"; Type II = "I miss what is there".
- Bagging = Bags in parallel; Boosting = Boots one after another.
- Overfit = memorises; underfit = never learns.
- Pie = parts, Bar = compare, Hist = spread, Scatter = relation.
Coverage checklist
- Introduction to Pandas: Jun 2026 Pandas in detail.
- understanding DataFrame: Jun 2025 DataFrame vs NumPy.
- Missing Values: isnull, dropna, fillna.
- Data operation: groupby, merge, concat.
- String Manipulation: Jun 2024 string operations.
- Regular Expressions and Data learning: Jun 2025 regex extraction.
- Outlier and Error: z-score, IQR.
- Visualization tool in Python: Representation of Pie Chart, Bar Chart, Histogram, Scatterplots using Python: Jun 2023, Dec 2024, Jun 2024, Jun 2025, Jun 2026.
- Data Analysis: pipeline steps.
- performance metrics: Jun 2026 metrics.
- ROC curve: TPR, FPR, AUC.
- types of errors: Jun 2025 Type I vs II.
- Overfitting & Under fitting: Dec 2024.
- evaluating performance of learning model: Holdout: train-test split.
- Random sampling: with and without replacement.
- cross validation and Bootstrap method: k-fold, bootstrap.
- Bagging & boosting: Dec 2024.
- Gradient Boosting: Jun 2023, Dec 2024.
- Random Forests: Jun 2023, Jun 2024.
- Committee Machines: static and dynamic.