How unit 4 is examined
Rough sets handle vague data using only the data itself; the marks sit in Hidden Markov Models, rough membership with the decision tree, and the fuzzy versus rough fuzzy comparison.
Introduction
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Rough set theory, proposed by Pawlak, describes a vague concept by a pair of precise sets, its lower and upper approximation, built from the indiscernibility of objects in a data table.</mark>
Key points.
- Data is kept in an information table whose rows are objects and whose columns are attributes.
- Two objects are indiscernible if they have identical values on the chosen attributes, and this relation splits the universe into equivalence classes called granules.
- A concept that is a union of granules is crisp; any other concept is rough and is bracketed by two approximations.
- Unlike fuzzy sets, it needs no membership function or prior probability supplied by an expert.
Fundamental Concepts
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. An information system is $S=(U,A)$ with $U$ a finite set of objects and $A$ a set of attributes; for $B\subseteq A$ the indiscernibility relation is $IND(B)=\{(x,y): a(x)=a(y)\ \forall a\in B\}$ and $[x]_B$ is the class of $x$.
Key points.
- A fuzzy set $A$ on $U$ is described by a membership function $\mu_A:U\to[0,1]$, so vagueness is a degree of belonging.
- A rough set $X$ is described by $\underline{B}X$ and $\overline{B}X$, so vagueness is the boundary region of objects that cannot be classified with certainty.
- A rough fuzzy set applies the rough approximations to a fuzzy set: the lower and upper approximations of a fuzzy set $\mu$ are $\mu_{\underline{B}}(x)=\min_{y\in[x]}\mu(y)$ and $\mu_{\overline{B}}(x)=\max_{y\in[x]}\mu(y)$, so each class gets a lower and upper membership.
| Basis | Fuzzy set | Rough fuzzy set |
|---|---|---|
| Vagueness modelled by | Graded membership | Granules plus graded membership |
| Needs | Membership function from an expert | Equivalence relation from the data, plus a fuzzy set |
| Description | One value $\mu(x)\in[0,1]$ | Pair $(\mu_{\underline{B}},\mu_{\overline{B}})$ per class |
| Uncertainty type | Vagueness of the concept | Granularity and vagueness together |
| Example | "Tall": 170 cm = 0.5, 180 cm = 0.9 | Persons grouped by age band; "tall" approximated per band by min and max of $\mu$ |
Answer frame. Open with the fuzzy set definition; give the table above; close with one example each: tall with $\mu$ values, and the age band whose lower value is the minimum and upper value the maximum of $\mu$.
Asked: [7 marks] (Nov 2023) Explain the difference between fuzzy set and rough fuzzy sets with the help of examples.
Set approximation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. For a target set $X\subseteq U$ and attributes $B$, ==the lower approximation $\underline{B}X=\{x:[x]_B\subseteq X\}$ holds objects certainly in $X$, and the upper approximation $\overline{B}X=\{x:[x]_B\cap X\neq\emptyset\}$ holds objects possibly in $X$.==
Key points.
- The boundary region is $BN=\overline{B}X-\underline{B}X$, and $X$ is rough exactly when it is non-empty.
- The accuracy of approximation is $\alpha=|\underline{B}X|/|\overline{B}X|$, with $0\le\alpha\le1$ and $\alpha=1$ for a crisp set.
- The outside region $U-\overline{B}X$ holds objects certainly not in $X$.
- Example: classes $\{1,2\},\{3,4,5\},\{6\},\{7,8\}$ and $X=\{1,2,3,6\}$ give lower $\{1,2,6\}$, upper $\{1,2,3,4,5,6\}$, boundary $\{3,4,5\}$, $\alpha=3/6=0.5$.
Rough membership
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. ==The rough membership function $\mu_X^B(x)=\dfrac{|[x]_B\cap X|}{|[x]_B|}$ gives the fraction of the indiscernibility class of $x$ that lies in $X$, so $\mu_X^B(x)\in[0,1]$.==
Key points.
- It expresses the degree to which object $x$ belongs to $X$ given the knowledge in attributes $B$, and it is computed from data counts, not supplied by an expert.
- $\mu=1$ exactly when $x\in\underline{B}X$, and $\mu=0$ exactly when $x$ is outside $\overline{B}X$.
- $0<\mu<1$ exactly for objects in the boundary region, which are the only ones that make the set rough.
- Approximations follow from it: $\underline{B}X=\{x:\mu=1\}$ and $\overline{B}X=\{x:\mu>0\}$.
- Complement rule: $\mu_{U-X}(x)=1-\mu_X(x)$; union and intersection do not obey the fuzzy max and min rules, which separates it from fuzzy membership.
- It is a conditional-probability estimate $P(x\in X\mid [x]_B)$, so it suits classification: an object is assigned to $X$ when $\mu$ passes a threshold.
Example. With classes $\{1,2\},\{3,4,5\},\{6\},\{7,8\}$ and $X=\{1,2,3,6\}$:
| Object | Class | $\lvert[x]\cap X\rvert/\lvert[x]\rvert$ | $\mu$ |
|---|---|---|---|
| 1, 2 | $\{1,2\}$ | 2/2 | 1 |
| 3, 4, 5 | $\{3,4,5\}$ | 1/3 | 0.33 |
| 6 | $\{6\}$ | 1/1 | 1 |
| 7, 8 | $\{7,8\}$ | 0/2 | 0 |
Rough membership of object 3 is 1/3, so it lies in the boundary region.
Answer frame. For Q1 write both parts as 7 marks each. Part (i): open with the formula, list points 1-6, end with the table example. Part (ii): see Decision tree model.
Asked: [14 marks] (Dec 2020) Explain following term: i) Rough Membership ii) Decision tree model
Attributes
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A reduct is a minimal subset of condition attributes that preserves the same classification power, that is the same indiscernibility, as the full attribute set.</mark>
Key points.
- Attributes are split into condition attributes $C$ and a decision attribute $D$.
- The core is the intersection of all reducts, the attributes that can never be removed.
- The dependency degree $\gamma(C,D)=|POS_C(D)|/|U|$, where the positive region is the union of lower approximations of the decision classes; a reduct keeps $\gamma$ unchanged.
- A table can have several reducts, and finding the smallest one is NP-hard, which motivates heuristic search (see Optimization).
Optimization
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. Feature selection uses optimization to find a small attribute subset, a near-minimal reduct, that keeps classifier accuracy.
Key points.
- With $n$ attributes there are $2^n-1$ subsets, so exhaustive search is impossible for large $n$; dimensionality reduction is needed to remove irrelevant and redundant features.
- Ant colony optimization builds subsets step by step, with ants choosing attributes by pheromone and heuristic value (for example dependency degree), and good subsets are reinforced.
- Particle swarm optimization treats a subset as a binary particle that moves toward its personal best and the global best, so the search space is explored without enumeration.
- These techniques and others (genetic algorithms, simulated annealing) escape local optima and cope with the huge search space better than greedy hill climbing.
- The result is fewer features, faster training, less overfitting and usually equal or better classifier performance.
Asked: [7 marks] (Nov 2023) Why the Ant colony optimization, Particle Swarm optimization and other techniques are used for feature selection?
Hidden Markov Models
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>A Hidden Markov Model is a statistical model of a Markov process whose states are hidden and can only be inferred from a sequence of observations emitted by those states.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u4-01" viewBox="0 0 295 252" width="295" height="252" role="img" aria-label="HMM with hidden states R and S (rainy, sunny), each emitting observations W and H (walk, shop)"><style>#dsfig-u4-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u4-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u4-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u4-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u4-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u4-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u4-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u4-01 .t{fill:#16181D;font-weight:500}#dsfig-u4-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u4-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u4-01 .dot{fill:#16181D}#dsfig-u4-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u4-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u4-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u4-01 .ah{fill:#454C5A}#dsfig-u4-01 .ah.hi{fill:#2340B8}#dsfig-u4-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u4-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u4-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u4-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u4-01 .e{stroke:#B1B7C3}html.dark #dsfig-u4-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u4-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u4-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u4-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u4-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u4-01 .t{fill:#E6E8ED}html.dark #dsfig-u4-01 .t.inv{fill:#0F1115}html.dark #dsfig-u4-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u4-01 .dot{fill:#E6E8ED}html.dark #dsfig-u4-01 .ann{fill:#8FA3FF}html.dark #dsfig-u4-01 .lbl{fill:#858D9C}html.dark #dsfig-u4-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u4-01 .ah{fill:#B1B7C3}html.dark #dsfig-u4-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u4-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u4-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u4-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah5" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh5" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M61,212 L234,212" marker-end="url(#ah5)" marker-start="url(#ah5)"/><path class="e" d="M40,193 L40,61" marker-end="url(#ah5)"/><path class="e" d="M255,193 L255,61" marker-end="url(#ah5)"/><g class="wl"><rect x="124" y="203" width="47.1" height="18" rx="9"/><text class="t" x="147.5" y="212" dy=".35em" text-anchor="middle">trans</text></g><g class="wl"><rect x="19.6" y="117" width="40.8" height="18" rx="9"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">emit</text></g><g class="wl"><rect x="234.6" y="117" width="40.8" height="18" rx="9"/><text class="t" x="255" y="126" dy=".35em" text-anchor="middle">emit</text></g><circle class="n" cx="40" cy="212" r="18"/><text class="t" x="40" y="212" dy=".35em" text-anchor="middle">R</text><circle class="n" cx="255" cy="212" r="18"/><text class="t" x="255" y="212" dy=".35em" text-anchor="middle">S</text><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">W</text><circle class="n" cx="255" cy="40" r="18"/><text class="t" x="255" y="40" dy=".35em" text-anchor="middle">H</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">HMM with hidden states R and S (rainy, sunny), each emitting observations W and H (walk, shop)</figcaption></figure>
Key points.
- The components are $\lambda=(A,B,\pi)$: a set of hidden states, a set of observation symbols, transition matrix $A$, emission matrix $B$ and initial distribution $\pi$.
- Transition probability $a_{ij}=P(q_{t+1}=j\mid q_t=i)$ and emission probability $b_j(o)=P(o\mid q_t=j)$; each row sums to 1.
- The Markov assumption says the next state depends only on the current state.
- Problem 1, evaluation: find $P(O\mid\lambda)$ using the forward algorithm.
- Problem 2, decoding: find the most likely state sequence using the Viterbi algorithm.
- Problem 3, learning: adjust $A,B,\pi$ to maximise $P(O\mid\lambda)$ using Baum-Welch (EM).
- Applications: speech recognition, bioinformatics (gene finding, protein sequences), gesture and handwriting recognition, part-of-speech tagging and finance.
Example. States R, S; $\pi=(0.6,0.4)$; $a_{RR}=0.7,a_{RS}=0.3,a_{SR}=0.4,a_{SS}=0.6$; $b_R(w)=0.1,b_R(h)=0.9,b_S(w)=0.8,b_S(h)=0.2$. Observations: walk, shop.
| Step | R | S |
|---|---|---|
| $\alpha_1$ | $0.6\times0.1=0.06$ | $0.4\times0.8=0.32$ |
| $\alpha_2$ | $(0.06\times0.7+0.32\times0.4)\times0.9=0.153$ | $(0.06\times0.3+0.32\times0.6)\times0.2=0.042$ |
$P(O\mid\lambda)=0.153+0.042=0.195$. Viterbi keeps the maximum instead of the sum: the best path is S then R with probability $0.32\times0.4\times0.9=0.1152$.
Answer frame. Open with the definition; draw the state diagram; define the components, then the three problems, then the example calculation; close with the applications list.
Asked: [7 marks] (Dec 2020, Nov 2023) Describe Hidden Markov Model and its application. Explain Hidden Markov Models with its application and with help of suitable example.
Decision tree model
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>A decision tree is a tree classifier whose internal nodes test an attribute, branches carry the test outcomes, and leaves give the class; a rough set-based tree chooses each split attribute by the dependency degree $\gamma$ instead of information gain.</mark>
Key points.
- Splitting divides the objects at a node by the values of one attribute, and pruning removes weak branches to avoid overfitting.
- In the rough set version the split attribute is the one with the largest $\gamma(B,D)=|POS_B(D)|/|U|$, so the branch with the most certainly classified objects comes first.
- A branch becomes a leaf when its objects all have one decision, that is when the boundary region is empty.
- Advantages: it works with inconsistent data, needs no thresholds, gives simple if-then rules and pairs with reducts to drop useless attributes.
Example. Objects (Outlook, Windy, Play): 1 (Sunny, Yes, No), 2 (Sunny, No, Yes), 3 (Rain, Yes, No), 4 (Rain, No, Yes), 5 (Overcast, Yes, Yes), 6 (Overcast, No, Yes). Outlook gives $POS=\{5,6\}$, so $\gamma=2/6=0.33$. Windy gives $POS=\{2,4,6\}$ plus $\{1,3,5\}$ is mixed, so $\gamma=3/6=0.5$. Split on Windy first; the mixed branch $\{1,3,5\}$ is split on Outlook.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u4-02" viewBox="0 0 518 194" width="518" height="194" role="img" aria-label="Rough set-based decision tree for the Play data"><style>#dsfig-u4-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u4-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u4-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u4-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u4-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u4-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u4-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u4-02 .t{fill:#16181D;font-weight:500}#dsfig-u4-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u4-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u4-02 .dot{fill:#16181D}#dsfig-u4-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u4-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u4-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u4-02 .ah{fill:#454C5A}#dsfig-u4-02 .ah.hi{fill:#2340B8}#dsfig-u4-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u4-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u4-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u4-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u4-02 .e{stroke:#B1B7C3}html.dark #dsfig-u4-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u4-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u4-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u4-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u4-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u4-02 .t{fill:#E6E8ED}html.dark #dsfig-u4-02 .t.inv{fill:#0F1115}html.dark #dsfig-u4-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u4-02 .dot{fill:#E6E8ED}html.dark #dsfig-u4-02 .ann{fill:#8FA3FF}html.dark #dsfig-u4-02 .lbl{fill:#858D9C}html.dark #dsfig-u4-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u4-02 .ah{fill:#B1B7C3}html.dark #dsfig-u4-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u4-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u4-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u4-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah6" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh6" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><line class="e" x1="191.6" y1="37" x2="75" y2="101"/><line class="e" x1="191.6" y1="37" x2="308.3" y2="101"/><line class="e" x1="308.3" y1="101" x2="197.5" y2="165"/><line class="e" x1="308.3" y1="101" x2="300.5" y2="165"/><line class="e" x1="308.3" y1="101" x2="419" y2="165"/><rect class="n" x="158.1" y="22" width="67" height="30" rx="8"/><text class="t" x="191.6" y="37" dy=".35em" text-anchor="middle">Windy?</text><rect class="n" x="14" y="86" width="122" height="30" rx="8"/><text class="t" x="75" y="101" dy=".35em" text-anchor="middle">Windy=No: Yes</text><rect class="n" x="223.8" y="86" width="169" height="30" rx="8"/><text class="t" x="308.3" y="101" dy=".35em" text-anchor="middle">Windy=Yes: Outlook?</text><rect class="n" x="152" y="150" width="91" height="30" rx="8"/><text class="t" x="197.5" y="165" dy=".35em" text-anchor="middle">Sunny: No</text><rect class="n" x="259" y="150" width="83" height="30" rx="8"/><text class="t" x="300.5" y="165" dy=".35em" text-anchor="middle">Rain: No</text><rect class="n" x="358" y="150" width="122" height="30" rx="8"/><text class="t" x="419" y="165" dy=".35em" text-anchor="middle">Overcast: Yes</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Rough set-based decision tree for the Play data</figcaption></figure>
Answer frame. Open with the rough set terms (indiscernibility, approximations); explain the $\gamma$ split rule; show the dataset, the two $\gamma$ values and the tree; close with the advantages. For part (ii) of Q1 write the definition, nodes, splitting, pruning and one classification use.
Asked: [7 marks] (Nov 2023) Explain the rough set-based decision tree with help of example.
Last-minute revision
- Rough set = lower approximation (certain) plus upper approximation (possible) of a set.
- $\underline{B}X=\{x:[x]\subseteq X\}$, $\overline{B}X=\{x:[x]\cap X\neq\emptyset\}$, boundary = upper minus lower.
- Accuracy $\alpha=|\text{lower}|/|\text{upper}|$; crisp when $\alpha=1$.
- Rough membership $\mu=|[x]\cap X|/|[x]|$; 1 in lower, 0 outside upper.
- Reduct = minimal attribute subset with same classification; core = intersection of reducts.
- Dependency $\gamma=|POS|/|U|$.
- Fuzzy set uses membership function; rough set uses boundary region; rough fuzzy uses min and max of $\mu$ per class.
- ACO and PSO search the $2^n$ subsets for feature selection.
- HMM = $(A,B,\pi)$; problems: evaluation (forward), decoding (Viterbi), learning (Baum-Welch).
- HMM example: $P(\text{walk, shop})=0.195$; best path S,R = 0.1152.
- Rough decision tree splits on largest $\gamma$; Windy 0.5 beats Outlook 0.33.
Memory hooks
- Lower = Like certain, Upper = Uncertain too: lower certain, upper possible.
- Rough membership is a fraction: inside the class over the whole class.
- HMM problems in order: "E-D-L", Evaluate (forward), Decode (Viterbi), Learn (Baum-Welch).
- Reduct = Reduce without Regret: same power, fewer attributes.
- Fuzzy blurs the edge with degrees; rough draws a boundary band.
Coverage checklist
- Introduction: covered, no past questions.
- Fundamental Concepts: Q3 fuzzy set versus rough fuzzy set (Nov 2023).
- Set approximation: covered, no past questions.
- Rough membership: Q1 part (i) (Dec 2020).
- Attributes: covered, no past questions.
- Optimization: Q5 ACO, PSO for feature selection (Nov 2023).
- Hidden Markov Models: Q4 (Dec 2020, Nov 2023).
- Decision tree model: Q2 (Nov 2023) and Q1 part (ii) (Dec 2020).