How unit 4 is examined
This unit covers combining many models into one stronger model; the marks sit in random forests, bagging versus pasting or boosting, and the max voting and averaging techniques.
Introduction to Ensemble Learning
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Ensemble learning combines the predictions of several models, called base learners, so that the combined model is more accurate and more stable than any single one of them.</mark>
Key points.
- The underlying idea is the wisdom of the crowd: many learners that make different errors cancel each other's mistakes when their outputs are combined.
- Base learners are often weak learners, meaning models only slightly better than random guessing, and an ensemble of them can become a strong learner.
- Ensembling reduces variance (bagging), reduces bias (boosting) and reduces overfitting, so generalization improves on unseen data.
- It works best when the learners are diverse, which is obtained by different training samples, different features or different algorithms.
- The main families are voting, bagging, boosting and stacking.
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u4-01" viewBox="0 0 467 252" width="467" height="252" role="img" aria-label="Data goes to base models M1-M3; the combiner C (vote, average) gives the final prediction P"><style>#dsfig-u4-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u4-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u4-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u4-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u4-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u4-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u4-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u4-01 .t{fill:#16181D;font-weight:500}#dsfig-u4-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u4-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u4-01 .dot{fill:#16181D}#dsfig-u4-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u4-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u4-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u4-01 .ah{fill:#454C5A}#dsfig-u4-01 .ah.hi{fill:#2340B8}#dsfig-u4-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u4-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u4-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u4-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u4-01 .e{stroke:#B1B7C3}html.dark #dsfig-u4-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u4-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u4-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u4-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u4-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u4-01 .t{fill:#E6E8ED}html.dark #dsfig-u4-01 .t.inv{fill:#0F1115}html.dark #dsfig-u4-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u4-01 .dot{fill:#E6E8ED}html.dark #dsfig-u4-01 .ann{fill:#8FA3FF}html.dark #dsfig-u4-01 .lbl{fill:#858D9C}html.dark #dsfig-u4-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u4-01 .ah{fill:#B1B7C3}html.dark #dsfig-u4-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u4-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u4-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u4-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah6" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh6" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M55.8,115.5 L151.5,51.6" marker-end="url(#ah6)"/><path class="e" d="M59,126 L148,126" marker-end="url(#ah6)"/><path class="e" d="M55.8,136.5 L151.5,200.4" marker-end="url(#ah6)"/><path class="e" d="M184.8,50.5 L280.5,114.4" marker-end="url(#ah6)"/><path class="e" d="M188,126 L277,126" marker-end="url(#ah6)"/><path class="e" d="M184.8,201.5 L280.5,137.6" marker-end="url(#ah6)"/><path class="e" d="M317,126 L406,126" marker-end="url(#ah6)"/><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">D</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">M1</text><circle class="n" cx="169" cy="126" r="18"/><text class="t" x="169" y="126" dy=".35em" text-anchor="middle">M2</text><circle class="n" cx="169" cy="212" r="18"/><text class="t" x="169" y="212" dy=".35em" text-anchor="middle">M3</text><circle class="n" cx="298" cy="126" r="18"/><text class="t" x="298" y="126" dy=".35em" text-anchor="middle">C</text><circle class="n" cx="427" cy="126" r="18"/><text class="t" x="427" y="126" dy=".35em" text-anchor="middle">P</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Data goes to base models M1-M3; the combiner C (vote, average) gives the final prediction P</figcaption></figure>
Asked: [7 marks] (Nov 2023) Explain the concept of ensemble learning in machine learning. What is the underlying idea behind ensemble methods?
Basic Ensemble Techniques (Max Voting, Averaging, Weighted Average)
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Basic ensemble techniques combine the outputs of several trained models by a simple rule: max voting takes the most frequent class, averaging takes the mean prediction, and weighted average takes the mean with more weight on better models.</mark>
Key points.
- Max voting (majority voting) is used for classification: every model predicts a class, one vote each, and the class with the most votes is the final output.
- Max voting works best with diverse classifiers, since independent errors are outvoted; an odd number of models avoids ties, and a tie is otherwise broken by the class with higher average confidence or by a fixed order.
- Averaging is used for regression (or for class probabilities in soft voting): the final output is the mean of all model predictions, $\hat{y}=\frac{1}{M}\sum_{i=1}^{M}\hat{y}_i$.
- Averaging reduces variance and smooths out the individual model's noise, so the ensemble is more stable than a single model.
- Weighted average gives each model a weight $w_i$ with $\sum w_i=1$, so a more accurate model influences the result more: $\hat{y}=\sum_{i=1}^{M} w_i\hat{y}_i$.
- Weights are chosen from validation accuracy; equal weights reduce weighted average to simple averaging.
Example. Three models predict a house price (lakh) as 40, 50, 60. Simple average $=(40+50+60)/3=50$. With weights 0.5, 0.3, 0.2: $0.5(40)+0.3(50)+0.2(60)=20+15+12=47$. For voting, votes A, B, A, A, B give A three votes and B two, so max voting output is A.
Answer frame. Open with the definition of the technique; write the formula; develop points 1-2 for voting or 3-6 for averaging; give the numeric example; close with "voting suits classification, averaging suits regression, and both cut variance". Draw the Introduction diagram.
Pitfall: Do not use max voting for regression; it needs class labels, while averaging needs numbers or probabilities.
Asked: [7 marks] (Nov 2023) Explain the averaging technique in ensemble learning. How does it combine predictions from multiple models? Asked: [7 marks] (Nov 2023) Describe the max voting technique in ensemble learning. How does it work in the context of classification?
Voting Classifiers
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. A voting classifier trains several different classifiers on the same data and predicts by combining their votes.
Key points.
- In hard voting the predicted class is the one chosen by the majority of classifiers: $\hat{y}=\text{mode}\{h_1(x),\dots,h_M(x)\}$.
- In soft voting the predicted class probabilities are averaged and the class with the highest average probability wins: $\hat{y}=\arg\max_c \frac{1}{M}\sum_i p_i(c\mid x)$.
- Soft voting usually beats hard voting because it gives more weight to highly confident votes, but every classifier must output probabilities.
- The classifiers should be as different as possible (for example logistic regression, SVM, decision tree), so their errors are independent.
Bagging and Pasting
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Bagging (bootstrap aggregating) trains the same algorithm on several random samples drawn with replacement from the training set and combines the models by voting or averaging; pasting does the same but samples without replacement.</mark>
Key points.
- Each base model gets its own random subset of the data, so the models differ and can be trained in parallel on separate cores.
- In bagging a sample is drawn with replacement, so some rows repeat and about 37% are left out of each sample; in pasting no row repeats within a sample.
- Predictions are combined by majority vote for classification and by average for regression.
- Aggregation lowers variance without raising bias much, so bagging suits high-variance, unstable models such as deep decision trees.
- Bagging gives more diversity per model than pasting (slightly higher bias, lower variance overall); pasting needs a large dataset because each subset has no repeats.
- Both scale well, since the training of each model is independent.
Steps.
Step 1: Draw B random samples from the training set (with replacement for bagging, without for pasting).
Step 2: Train one base model on each sample, independently and in parallel.
Step 3: For a new input, collect all B predictions.
Step 4: Output the majority vote (classification) or the mean (regression).
Comparison.
| Basis | Bagging | Pasting |
|---|---|---|
| Sampling | With replacement (bootstrap) | Without replacement |
| Duplicates in a sample | Yes | No |
| Diversity between models | Higher | Lower |
| Bias / variance | Slightly more bias, less variance | Less bias, a little more variance |
| Data needed | Works on small data | Needs large data |
Bagging versus boosting.
| Basis | Bagging | Boosting |
|---|---|---|
| Training | Parallel, independent models | Sequential, each corrects the previous |
| Data | Bootstrap samples, equal weight | Reweighted data, misclassified points get more weight |
| Reduces | Variance | Bias (and variance) |
| Combination | Equal vote or average | Weighted vote by model accuracy |
| Base learner | Strong, high variance (deep trees) | Weak (stumps) |
| Overfitting | Resists it | Can overfit noisy data |
| Example | Random Forest | AdaBoost, Gradient Boosting |
Answer frame. Open with "Bagging is bootstrap aggregating"; draw the Introduction diagram with bootstrap samples; develop points 1-4 in order; add the pasting comparison table; close with "bagging cuts variance by averaging". For bagging versus boosting, define both in one line each and lead with the second table.
Asked: [7 marks] (Nov 2022) What is Bagging and Boosting? Write few differences between them in detail. Asked: [7 marks] (Nov 2023) Explain bagging and pasting as ensemble techniques. What are the key differences between them?
Out-of-Bag Evaluation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Out-of-bag (OOB) evaluation tests each bagged model on the training rows it never saw, giving a free validation score without a separate validation set.
Key points.
- In a bootstrap sample of $n$ rows, each row is missed with probability $(1-\frac1n)^n\approx e^{-1}\approx 0.368$, so about 37% of rows are OOB for that model.
- Each row is predicted only by the models that did not train on it, and those predictions are compared with the true label.
- The averaged OOB accuracy is a good estimate of test accuracy, so cross-validation is not needed.
- In scikit-learn it is switched on with
oob_score=True.
Random Patches and Random Subspaces
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Both methods sample features as well as, or instead of, rows to make base models more diverse.
Key points.
- Random patches samples both training instances and features for each model (
bootstrap=True,max_features<1.0). - Random subspaces keeps all instances but samples only features (
bootstrap=False,max_features<1.0). - Feature sampling adds diversity, trading a little more bias for lower variance.
- They are useful for high-dimensional data such as images.
Random Forests (Extra-Trees, Feature Importance)
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>A random forest is an ensemble of decision trees trained by bagging, where each split considers only a random subset of features, and whose outputs are combined by majority vote (classification) or average (regression).</mark>
Key points.
- A decision tree is the base learner: it splits the data on features into a tree of rules, but a single deep tree overfits and has high variance.
- A random forest builds many such trees, each on a bootstrap sample of the rows, which is bagging.
- At every split only a random subset of features is examined (about $\sqrt{p}$ for classification), which decorrelates the trees.
- For classification the forest outputs the class chosen by most trees; for regression each tree predicts a continuous value and the forest returns their average, $\hat{y}=\frac1T\sum_{t=1}^{T}\hat{y}_t$.
- For example, if five trees predict house prices 40, 44, 46, 50, 60 lakh, the forest predicts $240/5=48$ lakh; a classifier would instead take the majority class.
- Compared with one tree, the forest has lower variance, less overfitting and better generalization, and it handles missing and mixed data well, but it is less interpretable.
- Extra-Trees (extremely randomized trees) also choose split thresholds at random instead of searching for the best one, which trains faster and lowers variance further at slightly higher bias.
- Feature importance: the importance of a feature is the average impurity (Gini) reduction it produces over all splits and trees, normalized to sum to 1, which helps in feature selection.
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u4-02" viewBox="0 0 285 130" width="285" height="130" role="img" aria-label="tree diagram"><style>#dsfig-u4-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u4-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u4-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u4-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u4-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u4-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u4-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u4-02 .t{fill:#16181D;font-weight:500}#dsfig-u4-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u4-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u4-02 .dot{fill:#16181D}#dsfig-u4-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u4-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u4-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u4-02 .ah{fill:#454C5A}#dsfig-u4-02 .ah.hi{fill:#2340B8}#dsfig-u4-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u4-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u4-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u4-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u4-02 .e{stroke:#B1B7C3}html.dark #dsfig-u4-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u4-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u4-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u4-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u4-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u4-02 .t{fill:#E6E8ED}html.dark #dsfig-u4-02 .t.inv{fill:#0F1115}html.dark #dsfig-u4-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u4-02 .dot{fill:#E6E8ED}html.dark #dsfig-u4-02 .ann{fill:#8FA3FF}html.dark #dsfig-u4-02 .lbl{fill:#858D9C}html.dark #dsfig-u4-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u4-02 .ah{fill:#B1B7C3}html.dark #dsfig-u4-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u4-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u4-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u4-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah7" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh7" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><line class="e" x1="130.5" y1="37" x2="47.5" y2="101"/><line class="e" x1="130.5" y1="37" x2="130.5" y2="101"/><line class="e" x1="130.5" y1="37" x2="213.5" y2="101"/><rect class="n" x="97" y="22" width="67" height="30" rx="8"/><text class="t" x="130.5" y="37" dy=".35em" text-anchor="middle">Forest</text><rect class="n" x="14" y="86" width="67" height="30" rx="8"/><text class="t" x="47.5" y="101" dy=".35em" text-anchor="middle">Tree 1</text><rect class="n" x="97" y="86" width="67" height="30" rx="8"/><text class="t" x="130.5" y="101" dy=".35em" text-anchor="middle">Tree 2</text><rect class="n" x="180" y="86" width="67" height="30" rx="8"/><text class="t" x="213.5" y="101" dy=".35em" text-anchor="middle">Tree 3</text></svg></figure>
Answer frame. Open with "A random forest is bagging applied to decision trees with random feature selection"; draw trees feeding a voting or averaging box; develop points 1-6; for regression stress points 4-5 and contrast with majority voting; close with "many uncorrelated trees give low variance and better generalization".
Asked: [8 marks] (Nov 2022) Discuss how Random Forest algorithm give output for Regression problems? Asked: [7 marks] (Nov 2022) How is a Random Forest related to Decision Trees? Discuss.
Boosting (AdaBoost, Gradient Boosting)
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Boosting trains weak learners one after another, each focusing on the mistakes of the previous ones, and combines them into a strong learner.
Key points.
- AdaBoost raises the weights of misclassified samples so the next learner concentrates on them; the learner's vote weight is $\alpha=\frac12\ln\frac{1-\epsilon}{\epsilon}$, so error $\epsilon=0.2$ gives $\alpha=0.693$.
- The final AdaBoost prediction is the sign of the weighted vote $\sum_t \alpha_t h_t(x)$.
- Gradient boosting fits each new tree to the residual errors (negative gradient of the loss) of the current ensemble and adds it with a learning rate $\eta$.
- Boosting is sequential, so it cannot be parallelised, and it can overfit noisy data.
Stacking
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Stacking (stacked generalization) trains a meta-learner to combine the predictions of several base models instead of using a fixed rule like voting.
Key points.
- Level-0 base models of different types are trained on the training data.
- Their predictions, made on held-out data, become the input features of a level-1 meta-model such as logistic regression.
- The meta-model learns how much to trust each base model.
- Held-out (cross-validated) predictions must be used, otherwise the meta-model sees overfitted outputs.
Last-minute revision
- Ensemble learning combines many base learners to lower variance, bias and overfitting.
- Max voting = most frequent class; averaging = mean of predictions (regression); weighted average uses weights that sum to 1.
- Hard voting counts labels; soft voting averages probabilities and usually wins.
- Bagging = bootstrap (with replacement) plus aggregation; pasting = without replacement.
- A bootstrap sample leaves out about 37% of rows; these OOB rows give a free accuracy estimate.
- Random patches sample rows and features; random subspaces sample only features.
- Random forest = bagged trees plus random feature subset at each split; regression output is the average of tree outputs.
- Extra-Trees choose random split thresholds; feature importance is the mean Gini reduction.
- Bagging is parallel and cuts variance; boosting is sequential and cuts bias.
- AdaBoost vote weight $\alpha=\frac12\ln\frac{1-\epsilon}{\epsilon}$; gradient boosting fits residuals.
- Stacking trains a meta-model on base-model predictions.
Memory hooks
- Bagging = Bootstrap AGGregatING: parallel bags of data.
- Boosting = a relay race: each runner fixes the last one's mistakes.
- Forest = many trees + shaken features: bagging plus random splits.
- Vote for labels, average for numbers.
- Stacking = a manager (meta-model) judging the team.
Coverage checklist
- Introduction to Ensemble Learning: Nov 2023 concept of ensemble learning.
- Basic Ensemble Techniques (Max Voting, Averaging, Weighted Average): Nov 2023 averaging; Nov 2023 max voting.
- Voting Classifiers: hard and soft voting, no past questions.
- Bagging and Pasting: Nov 2022 bagging versus boosting; Nov 2023 bagging and pasting.
- Out-of-Bag Evaluation: OOB score, no past questions.
- Random Patches and Random Subspaces: no past questions.
- Random Forests (Extra-Trees, Feature Importance): Nov 2022 regression output; Nov 2022 relation to decision trees.
- Boosting (AdaBoost, Gradient Boosting): no past questions.
- Stacking: no past questions.