How unit 5 is examined
This unit covers multiprocessor systems, pipelining, vector and array processing, RISC versus CISC, and multicore chips. No topic has been asked in the supplied papers, so each is kept short but complete.
Characteristics of Multiprocessor
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A multiprocessor is a system with two or more CPUs that share memory and I/O and are controlled by one operating system.</mark>
Key points.
- It is an MIMD machine: each processor executes its own instruction stream on its own data.
- Its purpose is higher throughput and better reliability through parallel work.
- If one processor fails, the others keep running, so the system degrades gracefully.
- Coupling is tight when memory is shared and loose when each CPU has local memory and communicates by messages.
- A multiprocessor differs from a multicomputer: the multiprocessor shares one address space, whereas a multicomputer has separate memories joined by a network.
- Multiprocessors improve speed for many independent tasks (multiprogramming) or for one program split into parallel tasks.
- Improved availability comes from redundancy, because a failed CPU can be removed while the rest continue.
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-01" viewBox="0 0 338 338" width="338" height="338" role="img" aria-label="Shared-memory multiprocessor: P1-P3 are CPUs, Bus is the common bus, Mem is shared memory, IO is I/O devices"><style>#dsfig-u5-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-01 .t{fill:#16181D;font-weight:500}#dsfig-u5-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-01 .dot{fill:#16181D}#dsfig-u5-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-01 .ah{fill:#454C5A}#dsfig-u5-01 .ah.hi{fill:#2340B8}#dsfig-u5-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-01 .e{stroke:#B1B7C3}html.dark #dsfig-u5-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-01 .t{fill:#E6E8ED}html.dark #dsfig-u5-01 .t.inv{fill:#0F1115}html.dark #dsfig-u5-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-01 .dot{fill:#E6E8ED}html.dark #dsfig-u5-01 .ann{fill:#8FA3FF}html.dark #dsfig-u5-01 .lbl{fill:#858D9C}html.dark #dsfig-u5-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-01 .ah{fill:#B1B7C3}html.dark #dsfig-u5-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah11" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh11" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M53.4,53.4 L155.6,155.6"/><path class="e" d="M169,59 L169,150"/><path class="e" d="M284.6,53.4 L182.4,155.6"/><path class="e" d="M158.5,184.8 L93.5,282.2"/><path class="e" d="M179.5,184.8 L244.5,282.2"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">P1</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">P2</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">P3</text><circle class="n" cx="169" cy="169" r="18"/><text class="t" x="169" y="169" dy=".35em" text-anchor="middle">Bus</text><circle class="n" cx="83" cy="298" r="18"/><text class="t" x="83" y="298" dy=".35em" text-anchor="middle">Mem</text><circle class="n" cx="255" cy="298" r="18"/><text class="t" x="255" y="298" dy=".35em" text-anchor="middle">IO</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Shared-memory multiprocessor: P1-P3 are CPUs, Bus is the common bus, Mem is shared memory, IO is I/O devices</figcaption></figure>
Structure of Multiprocessor-Interprocessor Arbitration, Inter-Processor Communication and Synchronization
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Arbitration decides which processor gets a shared bus or resource, communication passes data between processors, and synchronization orders their access to shared data.</mark>
Key points.
- Arbitration is done by a static priority (daisy chain, fixed priority) or a dynamic scheme (rotating priority, LRU).
- Processors communicate through shared memory or by sending messages.
- Synchronization uses a semaphore or a lock set by an atomic test-and-set instruction, so only one processor enters the critical section.
- Common structures are time-shared bus, crossbar switch, multiport memory and multistage network.
- In a time-shared bus only one transfer happens at a time, so it is cheap but the bus is the bottleneck.
- A crossbar switch gives every processor a path to every memory module, so it is fastest but needs the most switches, $n\times m$ crosspoints for $n$ CPUs and $m$ modules.
- A multistage (Omega) network uses small 2x2 switches, so it costs less than a crossbar but can block.
| Structure | Cost | Speed |
|---|---|---|
| Time-shared bus | Lowest | Lowest |
| Multiport memory | High | High |
| Crossbar | Highest | Highest |
| Multistage network | Medium | Medium |
Memory in Multiprocessor System
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>In shared memory all processors access one global address space, while in distributed memory each processor owns local memory and exchanges data by messages.</mark>
Key points.
- Shared memory is simple to program, but the bus and memory become a bottleneck as processors grow.
- Distributed memory scales well, but the programmer must send messages explicitly.
- Per-processor caches cut bus traffic but create the cache coherence problem, solved by snooping or directory protocols.
- UMA has equal access time for all processors; NUMA has a faster local and slower remote access.
- In message passing, data moves with send and receive operations, so no shared variable needs locking.
- Write-through and write-back caches must be kept coherent, or two processors may see different values of the same variable.
- Snooping suits bus-based systems because every cache watches the bus; a directory suits large systems because it records which caches hold each block.
Concept of Pipelining
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Pipelining divides a task into sequential stages so that several instructions are in different stages at the same time.</mark>
Formula. For $k$ stages and $n$ tasks, speedup $S=\dfrac{nk}{k+n-1}$, which tends to $k$ for large $n$. For $k=4$, $n=100$: $S=400/103\approx 3.88$.
Key points.
- A typical pipeline has stages fetch, decode, operand fetch, execute and write back.
- Throughput rises to about one instruction per clock, though each instruction's own latency does not fall.
- Hazards stall the pipeline: structural (resource clash), data (dependence) and control (branch).
- Remedies are forwarding, stalls, branch prediction and delayed branches.
- Non-pipelined time for $n$ tasks is $nk$ cycles, while the pipeline needs only $k+n-1$ cycles.
- The clock period is set by the slowest stage, so stages should be balanced.
Diagram.
Cycle: 1 2 3 4 5
I1: IF ID EX WB
I2: IF ID EX WB
I3: IF ID EX
(IF fetch, ID decode, EX execute, WB write back)
Vector Processing
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A vector processor executes one instruction on a whole array of operands, giving SIMD-style parallelism.</mark>
Key points.
- It uses deeply pipelined functional units and vector registers, so one instruction replaces a whole scalar loop.
- It cuts instruction fetch and decode overhead, because one fetch does the work of many.
- Memory is interleaved in banks so that one element can be supplied every clock.
- It suits scientific work such as matrix operations and weather modelling, e.g. the Cray-1.
- A vector instruction names the operation, the base address, the vector length and the stride, so the hardware knows how to walk through memory.
- Example: adding two 100-element arrays needs one vector add rather than a 100-iteration loop of scalar adds.
- Chaining feeds the result of one vector operation straight into the next, so no wait for the whole vector is needed.
Array Processing
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>An array processor has many identical processing elements that execute the same instruction on different data at once, under one control unit.</mark>
Key points.
- It is a SIMD machine: one control unit broadcasts each instruction to all elements.
- Each processing element has local memory, and elements are linked by an interconnection network.
- It differs from a vector processor because it gets its speed from many parallel ALUs, not from a deep pipeline.
- It suits regular data, as in image processing and matrix arithmetic.
- A mask can switch off some processing elements, so conditional operations still work on the array.
- Example: adding two matrices takes one broadcast add per element, all done in one step by the array of PEs.
- Its weakness is irregular or branchy code, where many PEs sit idle.
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-02" viewBox="0 0 338 381" width="338" height="381" role="img" aria-label="Array processor: CU is the control unit, PE1-PE3 are processing elements, Net is the interconnection network"><style>#dsfig-u5-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-02 .t{fill:#16181D;font-weight:500}#dsfig-u5-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-02 .dot{fill:#16181D}#dsfig-u5-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-02 .ah{fill:#454C5A}#dsfig-u5-02 .ah.hi{fill:#2340B8}#dsfig-u5-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-02 .e{stroke:#B1B7C3}html.dark #dsfig-u5-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-02 .t{fill:#E6E8ED}html.dark #dsfig-u5-02 .t.inv{fill:#0F1115}html.dark #dsfig-u5-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-02 .dot{fill:#E6E8ED}html.dark #dsfig-u5-02 .ann{fill:#8FA3FF}html.dark #dsfig-u5-02 .lbl{fill:#858D9C}html.dark #dsfig-u5-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-02 .ah{fill:#B1B7C3}html.dark #dsfig-u5-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah12" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh12" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M40,59 L40,191" marker-end="url(#ah12)"/><path class="e" d="M51.4,55.2 L156.4,195.2" marker-end="url(#ah12)"/><path class="e" d="M55.8,50.5 L280.5,200.4" marker-end="url(#ah12)"/><path class="e" d="M53.4,225.4 L155.6,327.6"/><path class="e" d="M169,231 L169,322"/><path class="e" d="M284.6,225.4 L182.4,327.6"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">CU</text><circle class="n" cx="40" cy="212" r="18"/><text class="t" x="40" y="212" dy=".35em" text-anchor="middle">PE1</text><circle class="n" cx="169" cy="212" r="18"/><text class="t" x="169" y="212" dy=".35em" text-anchor="middle">PE2</text><circle class="n" cx="298" cy="212" r="18"/><text class="t" x="298" y="212" dy=".35em" text-anchor="middle">PE3</text><circle class="n" cx="169" cy="341" r="18"/><text class="t" x="169" y="341" dy=".35em" text-anchor="middle">Net</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Array processor: CU is the control unit, PE1-PE3 are processing elements, Net is the interconnection network</figcaption></figure>
RISC And CISC
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>RISC uses a small set of simple, fixed-length instructions that run in one cycle, while CISC uses many complex instructions of varying length.</mark>
| Point | RISC | CISC |
|---|---|---|
| Instructions | Few, simple | Many, complex |
| Format | Fixed length | Variable length |
| Addressing modes | Few | Many |
| Memory access | Only load and store | Any instruction |
| Control unit | Hardwired | Microprogrammed |
| Registers | Large set | Fewer |
| Pipelining | Easy | Harder |
Key points.
- RISC shifts complexity to the compiler, so it needs longer programs.
- CISC gives shorter programs, but its decoding is slower.
- Examples: ARM and MIPS are RISC; x86 is CISC.
- RISC needs more registers to reduce memory access, and uses register windows in some designs such as SPARC.
- RISC executes about one instruction per clock, which is why it pipelines so well.
- CISC saves memory when memory is costly, since one instruction does more work.
Study of Multicore Processor –Intel, AMD
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A multicore processor places two or more independent cores on one chip, sharing some cache and the memory interface.</mark>
Key points.
- More cores raise throughput without raising the clock, which avoids the heat and power limit of faster single cores.
- Each core has private L1 (and often L2) cache, and cores share an L3 cache.
- Intel offers Core i3, i5, i7 and i9 with hyper-threading, which runs two threads per core.
- AMD offers Ryzen with the Zen design and multi-threading, and both use cache coherence between cores.
- Multicore gains depend on the software being parallel; a serial program uses only one core (Amdahl's law).
- Cores communicate through the shared cache or an on-chip interconnect, which is far faster than an off-chip bus.
- Power falls because several slower cores use less energy than one very fast core.
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-03" viewBox="0 0 467 338" width="467" height="338" role="img" aria-label="Quad-core chip: C0-C3 are cores with private L1/L2, L3 is the shared cache, MC is the memory controller"><style>#dsfig-u5-03 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-03 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-03 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-03 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-03 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-03 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-03 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-03 .t{fill:#16181D;font-weight:500}#dsfig-u5-03 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-03 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-03 .dot{fill:#16181D}#dsfig-u5-03 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-03 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-03 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-03 .ah{fill:#454C5A}#dsfig-u5-03 .ah.hi{fill:#2340B8}#dsfig-u5-03 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-03 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-03 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-03 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-03 .e{stroke:#B1B7C3}html.dark #dsfig-u5-03 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-03 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-03 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-03 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-03 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-03 .t{fill:#E6E8ED}html.dark #dsfig-u5-03 .t.inv{fill:#0F1115}html.dark #dsfig-u5-03 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-03 .dot{fill:#E6E8ED}html.dark #dsfig-u5-03 .ann{fill:#8FA3FF}html.dark #dsfig-u5-03 .lbl{fill:#858D9C}html.dark #dsfig-u5-03 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-03 .ah{fill:#B1B7C3}html.dark #dsfig-u5-03 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-03 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-03 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-03 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah13" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh13" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M55.8,50.5 L217.7,158.5"/><path class="e" d="M177.5,57 L225,152"/><path class="e" d="M289.5,57 L242,152"/><path class="e" d="M411.2,50.5 L249.3,158.5"/><path class="e" d="M233.5,188 L233.5,279"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">C0</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">C1</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">C2</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">C3</text><circle class="n" cx="233.5" cy="169" r="18"/><text class="t" x="233.5" y="169" dy=".35em" text-anchor="middle">L3</text><circle class="n" cx="233.5" cy="298" r="18"/><text class="t" x="233.5" y="298" dy=".35em" text-anchor="middle">MC</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Quad-core chip: C0-C3 are cores with private L1/L2, L3 is the shared cache, MC is the memory controller</figcaption></figure>
Last-minute revision
- Multiprocessor: two or more CPUs, shared memory, one OS; it is MIMD.
- Arbitration: static (daisy chain) or dynamic (rotating priority, LRU).
- Synchronization uses a semaphore or an atomic test-and-set lock.
- UMA has equal access time; NUMA has a faster local memory.
- Pipeline speedup is $nk/(k+n-1)$, which tends to $k$.
- Hazards are structural, data and control.
- A vector processor uses pipelined units and vector registers; an array processor uses many PEs.
- RISC has simple fixed-length instructions and load/store access; CISC has complex variable-length ones.
- Multicore means many cores on one chip with a shared L3 cache.
Memory hooks
- MIMD = Many Instructions, Many Data; SIMD = Single Instruction, Many Data.
- Pipeline hazards: S-D-C (Structural, Data, Control).
- RISC = Reduced; the compiler does more, the hardware less.
- Vector = one deep pipe; array = many parallel ALUs.
Coverage checklist
- Characteristics of Multiprocessor: no past questions.
- Structure of Multiprocessor-Interprocessor Arbitration, Inter-Processor Communication and Synchronization: no past questions.
- Memory in Multiprocessor System: no past questions.
- Concept of Pipelining: no past questions.
- Vector Processing: no past questions.
- Array Processing: no past questions.
- RISC And CISC: no past questions.
- Study of Multicore Processor –Intel, AMD: no past questions.