How unit 5 is examined
Memory hierarchy, cache, virtual memory, multiprocessors, pipelining, vector and array processing, RISC/CISC and multicore; cache, virtual memory, pipelining, RISC/CISC and vector/array processing carry the marks.
Main memory-RAM, ROM
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. Main memory is the semiconductor memory the CPU addresses directly; RAM is volatile read-write memory and ROM is non-volatile, mostly read-only memory. <mark>Volatile memory loses its contents when power is removed, non-volatile memory keeps them.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-01" viewBox="0 0 672 258" width="672" height="258" role="img" aria-label="Memory hierarchy, top to bottom: speed and cost per bit fall, capacity rises"><style>#dsfig-u5-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-01 .t{fill:#16181D;font-weight:500}#dsfig-u5-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-01 .dot{fill:#16181D}#dsfig-u5-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-01 .ah{fill:#454C5A}#dsfig-u5-01 .ah.hi{fill:#2340B8}#dsfig-u5-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-01 .e{stroke:#B1B7C3}html.dark #dsfig-u5-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-01 .t{fill:#E6E8ED}html.dark #dsfig-u5-01 .t.inv{fill:#0F1115}html.dark #dsfig-u5-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-01 .dot{fill:#E6E8ED}html.dark #dsfig-u5-01 .ann{fill:#8FA3FF}html.dark #dsfig-u5-01 .lbl{fill:#858D9C}html.dark #dsfig-u5-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-01 .ah{fill:#B1B7C3}html.dark #dsfig-u5-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah19" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh19" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><line class="e" x1="556.5" y1="37" x2="401.5" y2="101"/><line class="e" x1="401.5" y1="101" x2="246.5" y2="165"/><line class="e" x1="246.5" y1="165" x2="91.5" y2="229"/><rect class="n" x="511" y="22" width="91" height="30" rx="8"/><text class="t" x="556.5" y="37" dy=".35em" text-anchor="middle">Registers</text><rect class="n" x="372" y="86" width="59" height="30" rx="8"/><text class="t" x="401.5" y="101" dy=".35em" text-anchor="middle">Cache</text><rect class="n" x="193.5" y="150" width="106" height="30" rx="8"/><text class="t" x="246.5" y="165" dy=".35em" text-anchor="middle">Main memory</text><rect class="n" x="19" y="214" width="145" height="30" rx="8"/><text class="t" x="91.5" y="229" dy=".35em" text-anchor="middle">Secondary memory</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Memory hierarchy, top to bottom: speed and cost per bit fall, capacity rises</figcaption></figure>
Key points.
- The hierarchy gives the CPU a memory nearly as fast as the top level and as large and cheap as the bottom level, because of locality of reference.
- SRAM stores a bit in a flip-flop, needs no refresh and is fast but costly, so it is used for cache; DRAM stores a bit as charge on a capacitor, needs periodic refresh, and is slower but dense and cheap, so it is used for main memory.
- ROM types are masked ROM, PROM (programmed once), EPROM (UV erasable) and EEPROM/flash (electrically erasable); ROM holds the bootstrap program.
- Main memory is fast, volatile and small; secondary memory is slow, non-volatile and large.
Volatile memory (RAM) is written freely and used for working storage; non-volatile memory (ROM, flash) is read-only or slow to write and holds boot code and firmware.
Answer frame. Open with volatile versus non-volatile; draw the pyramid for organisation questions; then points 1-4; close with each level's role.
Asked: [8 marks] (Jun 2022, Nov 2023) Memory organization and its types. Asked: [7 marks] (Nov 2023, Dec 2024) Discuss memory hierarchy in computer system. Asked: [7 marks] (Nov 2023) List the differences between volatile and non-volatile memories. Asked: [? marks] (Jun 2023) Semiconductor memories (with other topics). Asked: [? marks] (Jun 2023) Main vs secondary memory; cache vs virtual memory. Asked: [7 marks] (Jun 2026) Any two: auxiliary memory, virtual memory, ROM.
Secondary memory: magnetic tape, disk, optical storage
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. Secondary (auxiliary) memory is non-volatile, high-capacity, low-cost-per-bit storage outside main memory. <mark>Tape gives sequential access, disk gives direct access, and optical discs are portable read-mostly media.</mark>
Key points.
- A hard disk stores data on rotating platters divided into tracks and sectors, and access time is seek time plus rotational latency plus transfer time.
- Magnetic tape is a ribbon coated with magnetic material read sequentially, so it is very slow but has the lowest cost per bit, which suits backup and archives.
- Optical storage (CD, DVD, Blu-ray) is read by a laser sensing pits and lands; it is portable and cheap to copy but slower than disk and mostly read-only or write-once.
- Auxiliary memory is needed because main memory is volatile and too small for all programs.
Tape is sequential and slowest with the lowest cost per bit; disk is direct-access, fast and reliable; optical discs are slower than disk, scratch easily and are cheap to distribute.
Answer frame. Define secondary memory; construction of each device; compare access, cost, reliability; close with uses.
Asked: [7 marks] (Jun 2024, Jun 2025) HDD, tape, optical discs: compare access time, reliability, cost.
Cache memory: structure, mapping, replacement, performance
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. Cache is a small, very fast memory between the CPU and main memory that holds recently used blocks to bridge their speed gap. <mark>Cache works because of locality: temporal locality means a recently used item is used again soon, spatial locality means nearby items are used soon.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-02" viewBox="0 0 434.5 80" width="434.5" height="80" role="img" aria-label="CPU, cache and main memory; a miss fetches a whole block from main memory"><style>#dsfig-u5-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-02 .t{fill:#16181D;font-weight:500}#dsfig-u5-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-02 .dot{fill:#16181D}#dsfig-u5-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-02 .ah{fill:#454C5A}#dsfig-u5-02 .ah.hi{fill:#2340B8}#dsfig-u5-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-02 .e{stroke:#B1B7C3}html.dark #dsfig-u5-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-02 .t{fill:#E6E8ED}html.dark #dsfig-u5-02 .t.inv{fill:#0F1115}html.dark #dsfig-u5-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-02 .dot{fill:#E6E8ED}html.dark #dsfig-u5-02 .ann{fill:#8FA3FF}html.dark #dsfig-u5-02 .lbl{fill:#858D9C}html.dark #dsfig-u5-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-02 .ah{fill:#B1B7C3}html.dark #dsfig-u5-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah20" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh20" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M61,40 L180.5,40" marker-end="url(#ah20)" marker-start="url(#ah20)"/><path class="e" d="M243.5,40 L363,40" marker-end="url(#ah20)" marker-start="url(#ah20)"/><g class="wl"><rect x="109.2" y="31" width="33.6" height="18" rx="9"/><text class="t" x="126" y="40" dy=".35em" text-anchor="middle">hit</text></g><g class="wl"><rect x="277.6" y="31" width="40.8" height="18" rx="9"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">miss</text></g><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">CPU</text><rect class="n" x="183.5" y="25" width="57" height="30" rx="15"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">Cache</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">MM</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">CPU, cache and main memory; a miss fetches a whole block from main memory</figcaption></figure>
Key points.
- A cache hit means the word is found in cache, a miss means the block is fetched from main memory; hit ratio $H$ = hits / total accesses.
- Average access time is $T_{avg}=H\,T_c+(1-H)\,T_m$, with $T_c$ the cache time and $T_m$ the main-memory time.
- Direct mapping puts block $j$ only in line $j \bmod L$; the address is tag, line index, word offset; it needs one comparator but blocks sharing a line evict each other.
- Fully associative mapping puts a block in any line and compares all tags in parallel, giving the best hit ratio but the costliest hardware.
- Set-associative mapping puts block $j$ in set $j \bmod S$ and in any of the $k$ lines of that set; it balances hit ratio and cost.
- Replacement: FIFO removes the oldest block, LRU the least recently used, LFU the least frequently used; direct mapping needs none.
- Write-through updates cache and memory together (slow); write-back updates memory only when a dirty block is replaced (fast).
- Cache coherency is the problem of several caches holding different copies of one block; snooping or directory protocols solve it.
- Performance improves with larger cache, better block size, higher associativity, multilevel L1/L2/L3 and prefetching.
Example. 64K words memory, 1K words cache, 16-word blocks, direct mapped: 16-bit address = tag 6 + index 6 + offset 4. With $H=0.9$, $T_c=20$ ns, $T_m=200$ ns: $T_{avg}=0.9(20)+0.1(200)=$ 38 ns.
| Basis | Direct | Associative | Set-associative |
|---|---|---|---|
| Hit ratio | Lowest | Highest | Between |
| Hardware | Cheapest | Costliest | Moderate |
Answer frame. Mapping: define it, name the three types, draw the address split, develop one with the example, add the table. Role: define cache, draw the figure, locality, $T_{avg}$. Performance: points 6, 7, 9.
Asked: [7 marks] (May 2019, Dec 2020, Jun 2020, Nov 2023, Jun 2026) Name three cache mapping techniques and explain any one in detail. Asked: [7 marks] (Nov 2023, Dec 2024, Jun 2026) Compare direct, set-associative and fully associative mapping. Asked: [7 marks] (Jun 2024, Jun 2025) Cache role, temporal and spatial locality, design. Asked: [7 marks] (May 2019) Cache memory, hit ratio, average access time. Asked: [7 marks] (Nov 2023) Cache hit and miss; cache coherency. Asked: [7 marks] (Jun 2022) Short note on LRU algorithm. Asked: [6 marks] (Jun 2022) Improving cache performance. Asked: [14 marks] (Dec 2020) Define Flynn's taxonomy and replacement algorithm. Asked: [? marks] (Jun 2023) Replacement algorithm; improving cache performance. Asked: [? marks] (Jun 2023) LRU algorithm in brief (with other topics).
Virtual memory
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. Virtual memory lets a program larger than main memory run by keeping only needed pages in RAM and the rest on disk, while the CPU uses virtual addresses. <mark>The MMU translates each virtual address into a physical address using the page table.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-03" viewBox="0 0 424 252" width="424" height="252" role="img" aria-label="VA = page number + offset; a TLB hit gives the frame, a miss reads the page table; PA = frame number + offset"><style>#dsfig-u5-03 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-03 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-03 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-03 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-03 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-03 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-03 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-03 .t{fill:#16181D;font-weight:500}#dsfig-u5-03 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-03 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-03 .dot{fill:#16181D}#dsfig-u5-03 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-03 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-03 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-03 .ah{fill:#454C5A}#dsfig-u5-03 .ah.hi{fill:#2340B8}#dsfig-u5-03 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-03 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-03 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-03 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-03 .e{stroke:#B1B7C3}html.dark #dsfig-u5-03 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-03 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-03 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-03 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-03 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-03 .t{fill:#E6E8ED}html.dark #dsfig-u5-03 .t.inv{fill:#0F1115}html.dark #dsfig-u5-03 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-03 .dot{fill:#E6E8ED}html.dark #dsfig-u5-03 .ann{fill:#8FA3FF}html.dark #dsfig-u5-03 .lbl{fill:#858D9C}html.dark #dsfig-u5-03 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-03 .ah{fill:#B1B7C3}html.dark #dsfig-u5-03 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-03 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-03 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-03 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah21" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh21" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M57,117.5 L193.2,49.4" marker-end="url(#ah21)"/><path class="e" d="M57,134.5 L193.2,202.6" marker-end="url(#ah21)"/><path class="e" d="M229,48.5 L365.2,116.6" marker-end="url(#ah21)"/><path class="e" d="M229,203.5 L365.2,135.4" marker-end="url(#ah21)"/><g class="wl"><rect x="105.6" y="74" width="40.8" height="18" rx="9"/><text class="t" x="126" y="83" dy=".35em" text-anchor="middle">page</text></g><g class="wl"><rect x="105.6" y="160" width="40.8" height="18" rx="9"/><text class="t" x="126" y="169" dy=".35em" text-anchor="middle">miss</text></g><g class="wl"><rect x="274.5" y="74" width="47.1" height="18" rx="9"/><text class="t" x="298" y="83" dy=".35em" text-anchor="middle">frame</text></g><g class="wl"><rect x="274.5" y="160" width="47.1" height="18" rx="9"/><text class="t" x="298" y="169" dy=".35em" text-anchor="middle">frame</text></g><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">VA</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">TLB</text><circle class="n" cx="212" cy="212" r="18"/><text class="t" x="212" y="212" dy=".35em" text-anchor="middle">PT</text><circle class="n" cx="384" cy="126" r="18"/><text class="t" x="384" y="126" dy=".35em" text-anchor="middle">PA</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">VA = page number + offset; a TLB hit gives the frame, a miss reads the page table; PA = frame number + offset</figcaption></figure>
Key points.
- Virtual space is split into equal pages and physical memory into frames of the same size; a virtual address is page number plus offset.
- The page table maps page number to frame number; the physical address is the frame number joined to the offset.
- The TLB is a small cache of recent page-table entries, so most translations skip the page table.
- A page fault occurs when the page is not in memory; the OS loads it from disk, and if no frame is free a replacement algorithm picks a victim.
- FIFO replaces the oldest page and can show Belady's anomaly (more frames, more faults); LRU replaces the page unused longest; Optimal replaces the page needed farthest in the future, gives the fewest faults but cannot be implemented.
- Segmentation divides a program into variable-size logical segments with a segment table of base and limit; it gives protection, sharing and a logical view but causes external fragmentation, so segmented paging combines both.
Example. String 12342156212376321236, 4 frames.
| Algorithm | Faults at string positions | Total |
|---|---|---|
| FIFO | 1,2,3,4,7,8,9,10,12,13,14,16,17,19 | 14 |
| LRU | 1,2,3,4,7,8,12,13,14,17 | 10 |
| Optimal | 1,2,3,4,7,8,13,17 | 8 |
Step 1: For each page, if in a frame, it is a hit: set its last-used time to now.
Step 2: Else count a fault; use a free frame, or replace the frame with the smallest last-used time.
Step 3: Set the new page's last-used time to now; print the fault count.
Cache speeds access to main memory, is hardware-managed SRAM and misses cost nanoseconds; virtual memory gives the illusion of large memory, is managed by the OS and MMU with disk plus RAM, and misses cost milliseconds.
Answer frame. Translation: define both addresses, draw the figure, walk page number, TLB, page table, frame. Replacement: define page fault, explain FIFO, LRU, Optimal on one string, mention Belady's anomaly. Numerical: string, frame states, three totals.
Asked: [7 marks] (May 2019, Dec 2020) Explain any three page replacement methods with an example. Asked: [7 marks] (Jun 2020) With a diagram, explain address translation in virtual memory. Asked: [7 marks] (Dec 2024) Segmentation, advantages and challenges. Asked: [7 marks] (Jun 2025) Pseudocode to simulate page replacement using LRU. Asked: [7 marks] (Jun 2026) Page faults for 12342156212376321236 with LRU, FIFO, Optimal, 4 frames.
Memory management hardware
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. The Memory Management Unit (MMU) is hardware between the CPU and memory that converts virtual addresses to physical addresses. <mark>The MMU translates addresses, protects memory and supports relocation.</mark>
Key points.
- The MMU translates every virtual address using the page table and holds the TLB, a fast cache of recent translations.
- Protection bits (read, write, execute, valid) let it stop illegal access and raise a fault.
- It allows relocation, so a program can be loaded in any frames, and makes virtual memory possible.
Asked: [7 marks] (Jun 2024) Role of the MMU and address translation.
Characteristics of multiprocessor
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. A multiprocessor is a computer with two or more CPUs that share memory and I/O under one operating system. <mark>Multiprocessors raise throughput and reliability by running tasks in parallel on several CPUs.</mark>
Key points.
- Tightly coupled (shared-memory) systems communicate through a common memory; loosely coupled (distributed-memory) systems have local memories and use message passing.
- In UMA every processor has the same memory access time; in NUMA local memory is faster than remote memory.
- The interconnect (bus, crossbar, multistage) decides cost and scalability, and cache coherence must be kept because each processor caches shared data.
- Benefits are higher throughput, speed-up and fault tolerance through graceful degradation.
Answer frame. Define it with coupling types; draw the shared-bus figure below; points 2-4; close with benefits.
Asked: [? marks] (Jun 2023, Jun 2026) Explain characteristics and structure of multiprocessor.
Multiprocessor structure, arbitration, communication, synchronization
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. Inter-processor arbitration decides which processor gets a shared bus or memory, communication exchanges data between processors, and synchronization orders their accesses to shared data. <mark>Interconnection networks link processors to shared memory and decide the scalability of the system.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-04" viewBox="0 0 424 338" width="424" height="338" role="img" aria-label="Shared-bus multiprocessor; P = processor with cache, Bus = common system bus, M = shared memory and I/O"><style>#dsfig-u5-04 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-04 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-04 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-04 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-04 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-04 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-04 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-04 .t{fill:#16181D;font-weight:500}#dsfig-u5-04 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-04 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-04 .dot{fill:#16181D}#dsfig-u5-04 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-04 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-04 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-04 .ah{fill:#454C5A}#dsfig-u5-04 .ah.hi{fill:#2340B8}#dsfig-u5-04 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-04 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-04 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-04 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-04 .e{stroke:#B1B7C3}html.dark #dsfig-u5-04 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-04 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-04 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-04 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-04 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-04 .t{fill:#E6E8ED}html.dark #dsfig-u5-04 .t.inv{fill:#0F1115}html.dark #dsfig-u5-04 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-04 .dot{fill:#E6E8ED}html.dark #dsfig-u5-04 .ann{fill:#8FA3FF}html.dark #dsfig-u5-04 .lbl{fill:#858D9C}html.dark #dsfig-u5-04 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-04 .ah{fill:#B1B7C3}html.dark #dsfig-u5-04 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-04 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-04 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-04 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah22" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh22" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M55.2,51.4 L196.8,157.6"/><path class="e" d="M59,169 L193,169"/><path class="e" d="M55.2,286.6 L196.8,180.4"/><path class="e" d="M231,169 L365,169"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">P1</text><circle class="n" cx="40" cy="169" r="18"/><text class="t" x="40" y="169" dy=".35em" text-anchor="middle">P2</text><circle class="n" cx="40" cy="298" r="18"/><text class="t" x="40" y="298" dy=".35em" text-anchor="middle">P3</text><circle class="n" cx="212" cy="169" r="18"/><text class="t" x="212" y="169" dy=".35em" text-anchor="middle">Bus</text><circle class="n" cx="384" cy="169" r="18"/><text class="t" x="384" y="169" dy=".35em" text-anchor="middle">M</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Shared-bus multiprocessor; P = processor with cache, Bus = common system bus, M = shared memory and I/O</figcaption></figure>
Key points.
- The common bus is cheapest but allows one transfer at a time; the crossbar has a switch at every processor-memory crossing, so no conflicts but cost grows as $n\times m$; the multistage network uses about $\log n$ switch stages at medium cost; the hypercube links $2^n$ nodes, each to $n$ neighbours.
- Bus arbitration uses static or dynamic priority: daisy chain, parallel priority encoder, polling, LRU or rotating priority.
- Shared-memory communication passes data through common locations, which is fast but needs locking; message passing uses explicit send and receive and suits loosely coupled systems.
- Synchronization uses semaphores, locks, test-and-set and barriers so that only one processor is in a critical section at a time.
Answer frame. Define multiprocessor; draw the shared-bus figure; develop networks, arbitration, communication, synchronization; close with coupling type deciding the method.
Asked: [7 marks] (Jun 2022, Jun 2023, Jun 2025) Multiprocessor; inter-processor communication, synchronization, arbitration. Asked: [7 marks] (Jun 2020) Structure of general purpose multiprocessors. Asked: [7 marks] (Dec 2024) Types of interconnection networks in multiprocessors.
Memory in multiprocessor system
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. In a shared-memory multiprocessor all CPUs address one global memory; in a distributed-memory system each CPU has private memory.
Key points.
- Shared memory gives one address space and easy communication but needs cache coherence and synchronization.
- Distributed memory scales better but data moves by message passing.
Concept of pipelining
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. Pipelining divides a task into $k$ sequential segments, each with its own hardware, so that several tasks are in different segments at once. <mark>Pipelining overlaps successive tasks, so throughput reaches one result per clock once the pipeline is full.</mark>
Diagram. Space-time diagram, 6 segments, 8 tasks (number = task):
Clock 1 2 3 4 5 6 7 8 9 10 11 12 13
S1 1 2 3 4 5 6 7 8
S2 1 2 3 4 5 6 7 8
S3 1 2 3 4 5 6 7 8
S4 1 2 3 4 5 6 7 8
S5 1 2 3 4 5 6 7 8
S6 1 2 3 4 5 6 7 8
Formula. Time for $n$ tasks, $k$ segments, clock $t_p$: $T_{pipe}=(k+n-1)\,t_p$; non-pipelined $T_{non}=n\,k\,t_p$; speedup $S=\dfrac{nk}{k+n-1}\to k$ for large $n$. Here $6+8-1=$ 13 cycles against 48, so $S=48/13=$ 3.69.
Key points.
- Instruction pipeline stages are IF (fetch instruction), ID (decode), OF (fetch operands), EX (execute) and WB (write result back); while one instruction executes, the next is decoded and the one after is fetched.
- An arithmetic pipeline splits a floating-point operation into stages; floating-point addition uses compare exponents, align mantissas, add mantissas, normalise the result.
- Each instruction still takes $k$ cycles, but throughput is one instruction per clock once the pipe is full.
- Structural hazards (one resource needed twice), data hazards (result not yet written) and control hazards (branches) cause stalls; remedies are forwarding, stalls and branch prediction.
- In a multiprocessor each CPU has its own pipelines, with registers between segments holding intermediate results under a common clock.
Answer frame. Define pipelining; draw the space-time diagram; state $(k+n-1)t_p$ and speedup; then stages, hazards; close with the limit $k$. For arithmetic pipeline give the floating-point stages.
Asked: [7 marks] (May 2019) Pipelining; space-time diagram for six segments, eight tasks. Asked: [7 marks] (Dec 2020) Explain arithmetic pipeline. Asked: [7 marks] (Jun 2020, Jun 2022, Nov 2023) Concept of pipelining in multiprocessors. Asked: [7 marks] (Jun 2022, Nov 2023, Jun 2024) Concept of pipelining in detail. Asked: [7 marks] (Jun 2022, Jun 2026) Instruction pipelining, its stages and functions. Asked: [6 marks] (Jun 2022) Layout of a pipelined instruction. Asked: [? marks] (Jun 2023, Jun 2026) Concept of pipelining; differentiate vector and array processing.
Vector processing
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. Vector processing performs one operation on whole arrays (vectors) through a deeply pipelined functional unit, so one vector instruction replaces a loop. <mark>A vector instruction operates on a whole array, removing loop overhead and keeping the pipeline full.</mark>
Key points.
- A vector instruction gives the operation, base address and vector length, so it does the work of many scalar instructions.
- Scalar processing handles one item per instruction with fetch and decode each time; vector processing fetches and decodes once and streams elements through the pipeline.
- Memory-to-memory architecture streams operands between memory and the pipeline; register-to-register architecture uses vector registers, which is faster (Cray).
- Chaining feeds the output of one vector pipeline straight into the next without waiting for the whole vector.
- Advantages are speed, throughput and fewer instruction fetches; applications are weather forecasting, seismic analysis, image processing and supercomputing.
Answer frame. Define versus scalar; draw a vector pipeline (vector registers, pipelined adder and multiplier); develop points 1-4; close with applications.
Asked: [7 marks] (Jun 2022, Jun 2026) Vector processing, advantages over scalar. Asked: [7 marks] (Dec 2020) Pipeline vector processing methods. Asked: [7 marks] (Jun 2022) Discuss vector processing in detail. Asked: [14 marks] (Jun 2020) Vector processing; array processing; RISC and CISC.
Array processing
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. An array processor has many identical processing elements (PEs) under one control unit that execute the same instruction on different data at once. <mark>An array processor is a SIMD machine: single instruction, multiple data.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-05" viewBox="0 0 424 338" width="424" height="338" role="img" aria-label="SIMD array processor; CU = control unit, PE = processing element, M = local memory"><style>#dsfig-u5-05 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-05 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-05 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-05 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-05 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-05 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-05 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-05 .t{fill:#16181D;font-weight:500}#dsfig-u5-05 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-05 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-05 .dot{fill:#16181D}#dsfig-u5-05 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-05 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-05 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-05 .ah{fill:#454C5A}#dsfig-u5-05 .ah.hi{fill:#2340B8}#dsfig-u5-05 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-05 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-05 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-05 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-05 .e{stroke:#B1B7C3}html.dark #dsfig-u5-05 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-05 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-05 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-05 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-05 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-05 .t{fill:#E6E8ED}html.dark #dsfig-u5-05 .t.inv{fill:#0F1115}html.dark #dsfig-u5-05 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-05 .dot{fill:#E6E8ED}html.dark #dsfig-u5-05 .ann{fill:#8FA3FF}html.dark #dsfig-u5-05 .lbl{fill:#858D9C}html.dark #dsfig-u5-05 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-05 .ah{fill:#B1B7C3}html.dark #dsfig-u5-05 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-05 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-05 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-05 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah23" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh23" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M55.2,157.6 L195.2,52.6" marker-end="url(#ah23)"/><path class="e" d="M59,169 L191,169" marker-end="url(#ah23)"/><path class="e" d="M55.2,180.4 L195.2,285.4" marker-end="url(#ah23)"/><path class="e" d="M231,40 L365,40"/><path class="e" d="M231,169 L365,169"/><path class="e" d="M231,298 L365,298"/><circle class="n" cx="40" cy="169" r="18"/><text class="t" x="40" y="169" dy=".35em" text-anchor="middle">CU</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">PE1</text><circle class="n" cx="212" cy="169" r="18"/><text class="t" x="212" y="169" dy=".35em" text-anchor="middle">PE2</text><circle class="n" cx="212" cy="298" r="18"/><text class="t" x="212" y="298" dy=".35em" text-anchor="middle">PE3</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">M1</text><circle class="n" cx="384" cy="169" r="18"/><text class="t" x="384" y="169" dy=".35em" text-anchor="middle">M2</text><circle class="n" cx="384" cy="298" r="18"/><text class="t" x="384" y="298" dy=".35em" text-anchor="middle">M3</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">SIMD array processor; CU = control unit, PE = processing element, M = local memory</figcaption></figure>
Key points.
- The control unit broadcasts one instruction and every PE executes it on its own data, so parallelism is in the data, not the pipeline.
- Matrix multiplication $C=A\times B$: each PE computes one element $c_{ij}=\sum_k a_{ik}b_{kj}$ in parallel, or a systolic array passes data between neighbouring cells.
- Instruction format has opcode, mode and address fields, for example
ADD R1, R2; applications are matrix work and image processing.
Flynn's taxonomy by instruction and data streams: SISD (uniprocessor), SIMD (array or vector processor), MISD (rare), MIMD (multiprocessor).
Vector processing gets parallelism from one deep pipeline (lower cost); array processing gets it from many replicated PEs (higher cost).
Answer frame. Define the array processor as SIMD; draw the block diagram; develop points 1-3; close with applications.
Asked: [7 marks] (Jun 2022, Dec 2024) Array processors. Asked: [14 marks] (Dec 2020) Short notes: SIMD, matrix multiplication, instruction format.
RISC and CISC
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. RISC (Reduced Instruction Set Computer) uses a small set of simple, fixed-length instructions that run in one cycle; CISC (Complex Instruction Set Computer) uses many complex, variable-length instructions. <mark>RISC does less per instruction but runs each fast and pipelines well, whereas CISC does more per instruction using microcode.</mark>
Key points.
- RISC is load-store: only LOAD and STORE access memory and all other operations use registers.
- RISC has many registers, few addressing modes, hardwired control and an easily pipelined design, so the compiler does more work.
- CISC has complex instructions, many addressing modes, microprogrammed control and fewer registers, so programs are shorter.
- RISC examples are ARM, MIPS, SPARC; CISC examples are Intel x86 and VAX.
| Basis | RISC | CISC |
|---|---|---|
| Instruction set | Small, simple | Large, complex |
| Addressing modes | Few | Many |
| Registers | Many | Few |
| Control unit | Hardwired | Microprogrammed |
| Memory access | Load and store only | Any instruction |
Horizontal control has a wide control word, no decoding, is faster and needs more storage; vertical control has a narrow word, needs decoding, is slower and needs less storage.
A maskable interrupt can be ignored by a mask and serves I/O requests; a non-maskable interrupt cannot be ignored and serves power failure or memory error.
Answer frame. Define both; table; examples; close with the trade-off. For RISC architecture add registers and pipeline.
Asked: [7 marks] (Nov 2023, Jun 2024, Jun 2025, Jun 2026) RISC vs CISC with examples. Asked: [14 marks] (Dec 2020) Differentiate maskable and non-maskable interrupt; RISC and CISC. Asked: [7 marks] (Jun 2022) Discuss RISC architecture in detail. Asked: [7 marks] (Jun 2026) RISC vs CISC; horizontal vs vertical control unit.
Multicore processor: Intel, AMD
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. A multicore processor places two or more independent CPU cores on a single chip. <mark>Multicore chips gain performance through parallel threads rather than higher clock speed.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-06" viewBox="0 0 467 338" width="467" height="338" role="img" aria-label="Multicore chip; C = core with private L1 and L2 cache, L3 = shared cache, MC = memory controller"><style>#dsfig-u5-06 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-06 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-06 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-06 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-06 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-06 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-06 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-06 .t{fill:#16181D;font-weight:500}#dsfig-u5-06 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-06 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-06 .dot{fill:#16181D}#dsfig-u5-06 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-06 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-06 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-06 .ah{fill:#454C5A}#dsfig-u5-06 .ah.hi{fill:#2340B8}#dsfig-u5-06 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-06 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-06 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-06 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-06 .e{stroke:#B1B7C3}html.dark #dsfig-u5-06 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-06 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-06 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-06 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-06 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-06 .t{fill:#E6E8ED}html.dark #dsfig-u5-06 .t.inv{fill:#0F1115}html.dark #dsfig-u5-06 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-06 .dot{fill:#E6E8ED}html.dark #dsfig-u5-06 .ann{fill:#8FA3FF}html.dark #dsfig-u5-06 .lbl{fill:#858D9C}html.dark #dsfig-u5-06 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-06 .ah{fill:#B1B7C3}html.dark #dsfig-u5-06 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-06 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-06 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-06 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah24" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh24" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M55.7,50.7 L213.5,158.3"/><path class="e" d="M177,57.2 L221.2,151.8"/><path class="e" d="M289.1,56.8 L238.1,152.2"/><path class="e" d="M411.1,50.4 L245.1,158.6"/><path class="e" d="M229.2,188 L229.2,279"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">C1</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">C2</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">C3</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">C4</text><circle class="n" cx="229.2" cy="169" r="18"/><text class="t" x="229.2" y="169" dy=".35em" text-anchor="middle">L3</text><circle class="n" cx="229.2" cy="298" r="18"/><text class="t" x="229.2" y="298" dy=".35em" text-anchor="middle">MC</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Multicore chip; C = core with private L1 and L2 cache, L3 = shared cache, MC = memory controller</figcaption></figure>
Key points.
- Each core has its own L1 instruction and data caches and usually a private L2, while the L3 cache, memory controller and interconnect are shared.
- Intel Core processors use a ring or mesh interconnect and hyper-threading (two threads per core); AMD Ryzen uses chiplets joined by Infinity Fabric.
- Multicore gives parallelism and better performance per watt, but challenges are heat, thread synchronization and load balancing.
Answer frame. Define multicore; draw the chip diagram; develop points 1-3; close with benefits.
Asked: [7 marks] (Jun 2023, Jun 2026) Architecture of Intel multicore processors. Asked: [? marks] (Jun 2023) Discuss multicore processor in detail.
Last-minute revision
- Average access time $T_{avg}=H\,T_c+(1-H)\,T_m$; hit ratio = hits / accesses.
- Direct mapping: line = block mod lines; address = tag, index, offset.
- LRU replaces the least recently used block; FIFO the oldest; Optimal the one used farthest ahead.
- Page string 12342156212376321236 with 4 frames: FIFO 14, LRU 10, Optimal 8 faults.
- Pipeline time $(k+n-1)t_p$; 6 segments, 8 tasks take 13 cycles; stages IF, ID, OF, EX, WB.
- Array processor is SIMD; Flynn: SISD, SIMD, MISD, MIMD.
- RISC: fixed length, load-store, hardwired; CISC: variable length, microprogrammed.
Memory hooks
- Mapping: Direct is Dumb, Associative is Anywhere, Set is the Sensible middle.
- Pipeline stages: IF ID OF EX WB.
- Temporal is Time (same item again), Spatial is Space (neighbours).
- RISC has Registers, CISC has Complexity.
Coverage checklist
- Main memory-RAM, ROM: hierarchy, volatile, semiconductor, main vs secondary.
- Secondary Memory –Magnetic Tape, Disk, Optical Storage: comparison.
- Cache Memory: Cache Structure and Design, Mapping Scheme, Replacement Algorithm, Improving Cache Performance: mapping, hit/miss, coherency, LRU, Flynn.
- Virtual Memory: replacement, translation, segmentation, LRU, fault numerical.
- memory management hardware: MMU.
- Characteristics of Multiprocessor: characteristics, structure.
- Structure of Multiprocessor-Inter-processor Arbitration, Inter-Processor Communication and Synchronization: structure, communication, networks.
- Memory in Multiprocessor System: shared, distributed.
- Concept of Pipelining: space-time, arithmetic, instruction.
- Vector Processing: concept, methods.
- Array Processing: SIMD, matrix multiplication.
- RISC And CISC: comparison, RISC architecture.
- Study of Multicore Processor –Intel, AMD: Intel architecture.