UNIT 4: Computer Organization & Architecture - Short Notes
Based on RGPV past papers (Dec 2024 & Nov 2023), focusing on high-frequency exam topics.
I. Basic Computer Structure & Organization
Basic Functional Units
A computer consists of three primary functional units:
-
CPU (Central Processing Unit): The "brain." Fetches, decodes, and executes instructions. Contains ALU, control unit, and registers.
-
Memory Unit: Stores data and instructions. Hierarchical (registers, cache, main memory, secondary storage).
-
I/O Unit: Communicates with the external world (keyboard, display, disk). Handles data transfer between I/O devices and memory/CPU.
Exam Tip: Be prepared to draw the basic block diagram showing CPU, Memory, and I/O subsystems connected via a system bus.
System Bus Structure
A bus is a group of electrical lines for communication. Three main types:
| Bus Type | Purpose | Direction |
|---|---|---|
| Data Bus | Carries actual data (instructions, operands, results). | Bi-directional |
| Address Bus | Carries memory/I/O addresses from CPU. | Uni-directional (CPU → Memory/I/O) |
| Control Bus | Carries control signals (Read, Write, Interrupt, Clock). | Bi-directional |
-
Width of data bus determines amount of data transferred per cycle.
-
Width of address bus determines maximum addressable memory space ($$\displaystyle 2^{\text{width}} $$ locations).
General Register Organization
Registers are small, fast storage locations inside the CPU.
-
Common Registers:
-
Program Counter (PC): Holds address of next instruction to fetch.
-
Instruction Register (IR): Holds the current instruction being decoded/executed.
-
Memory Address Register (MAR): Holds address of memory location to be accessed.
-
Memory Data Register (MDR): Holds data read from or to be written to memory.
-
Accumulator (AC): Primary register for ALU operations.
-
General-Purpose Registers (R0, R1...): Used for operands and intermediate results.
-
II. Central Processing Unit (CPU) & Execution
Instruction Cycle (Fetch-Decode-Execute)
The fundamental cycle of CPU operation.
flowchart TD
A[Fetch] --> B[Decode]
B --> C[Execute]
C --> D{Interrupt?}
D -- No --> A
D -- Yes --> E[Service Interrupt] --> A
-
Fetch: PC → MAR → Memory → MDR → IR. PC incremented.
-
Decode: Control unit interprets opcode in IR.
-
Execute: ALU performs operation, or memory/I/O access occurs.
-
Interrupt Check: If an interrupt is pending, the cycle is interrupted to service it.
Stack Operations
A stack is a LIFO (Last-In, First-Out) memory structure, typically implemented in main memory with a Stack Pointer (SP) register pointing to the top.
-
PUSH (Push onto Stack):
-
Decrement SP.
-
Store operand at memory location pointed by SP.
-
-
POP (Pop from Stack):
-
Read operand from memory location pointed by SP.
-
Increment SP.
-
Example:
PUSH AX(Push contents of register AX onto stack).POP BX(Pop top of stack into register BX).
Floating-Point Arithmetic
Handles numbers in scientific notation: $$\displaystyle (-1)^s \times M \times 2^E $$.
-
Addition/Subtraction Flowchart:
-
Align Exponents: Shift mantissa of smaller exponent number until exponents match.
-
Add/Subtract Mantissas: Perform operation on aligned mantissas.
-
Normalize Result: Shift result mantissa left/right to ensure it's in standard form (1.M), adjusting exponent accordingly.
-
Round: Apply rounding rule to fit mantissa into available bits.
-
Check for Overflow/Underflow.
-
Instruction Set Architecture (ISA)
The interface between hardware and software. Defines:
-
Instruction formats and opcodes.
-
Data types and sizes.
-
Registers and their usage.
-
Memory addressing modes.
-
How I/O and interrupts are handled.
Key Point: ISA is abstract. Multiple microarchitectures (e.g., Intel vs AMD) can implement the same ISA (x86).
Instruction Formats
Common fields: Opcode (operation), Address/Register specifiers (operand locations).
-
Zero-Address (Stack): e.g.,
ADD(operands implicit from stack). -
One-Address (Accumulator): e.g.,
ADD M(AC ← AC + M). -
Two-Address: e.g.,
ADD R1, R2(R1 ← R1 + R2). -
Three-Address: e.g.,
ADD R1, R2, R3(R1 ← R2 + R3). More powerful, longer instructions.
III. Instruction Execution & Addressing
Addressing Modes
Specifies how to calculate the effective address (EA) of an operand.
| Mode | How EA is Calculated | Example (ADD) | Use Case |
|---|---|---|---|
| Immediate | Operand is in instruction itself. | ADD #5 |
Load constant. |
| Direct | EA = address field in instruction. | ADD 1000 |
Fast, fixed address. |
| Indirect | EA = contents of memory location whose address is in instruction. | ADD @1000 |
Pointer manipulation. |
| Register | Operand is in specified register. | ADD R1 |
Fastest. |
| Register Indirect | EA = contents of specified register. | ADD (R1) |
Array/string traversal. |
| Displacement | EA = register + constant offset. | ADD 1000(R1) |
Array element access. |
| Relative | EA = PC + constant offset. | ADD +1000 |
Position-independent code (PC-relative). |
| Indexed | EA = index register + constant. | ADD X(Ri) |
Array access. |
| Base-Register | EA = base register + displacement. | Similar to displacement. | Memory protection, relocation. |
Exam Tip: Be able to identify the mode from an example and calculate the EA given register/memory contents.
IV. Control Unit Design
Hardwired Control Unit
-
Principle: Control signals are generated by a fixed combinational logic circuit (decoders, gates). The instruction opcode directly drives the control lines.
-
Diagram: Shows instruction register opcode bits feeding into a control signal generator (logic gates) which outputs all control signals (e.g.,
ALUSrc,RegWrite,MemRead). -
Advantages: Very fast (no microinstruction fetch), efficient for simple ISAs.
-
Disadvantages: Complex to design and modify (adding instructions requires rewiring), inflexible.
Micro-programmed Control Unit
-
Principle: Control signals are stored as microinstructions in a control memory (CM). A micro-program sequencer fetches and executes these microinstructions.
-
Diagram: Shows a Micro-program Counter (μPC) that addresses a Control Memory. The output (a Control Word) drives the control lines. The μPC is updated based on the current microinstruction and instruction opcode.
-
Advantages: Easy to design, modify, and debug. Flexible (new instructions via new microcode).
-
Disadvantages: Slower (extra memory fetch cycle), requires control memory.
Differentiation: Hardwired vs. Micro-programmed
| Feature | Hardwired Control | Micro-programmed Control |
|---|---|---|
| Implementation | Fixed logic (gates, decoders) | Stored microinstructions in control memory |
| Speed | Faster (direct generation) | Slower (microinstruction fetch) |
| Flexibility | Inflexible (hardwired) | Flexible (microcode can be changed) |
| Complexity | Complex design for complex ISAs | Simpler design, easier to modify |
| Cost | Lower (no control memory) | Higher (requires control memory) |
| Usage | Simple, high-speed processors (e.g., RISC) | Complex CISC processors (e.g., x86) |
Micro-programming Concepts
-
Control Word (CW): A binary word where each bit (or group) controls a specific signal in the computer (e.g.,
1meansRegWrite=ON). One microinstruction = one control word. -
Micro-program Sequencer: Generates the address for the next microinstruction.
-
Functions: Increment μPC, jump to a new routine based on opcode or condition, handle subroutine calls.
-
Types: Simple counter-based, branching-capable (using multiplexers).
-
V. Memory Hierarchy & Organization
Memory Hierarchy Concept
Organizes memory into levels (Registers → L1/L2/L3 Cache → Main Memory → Disk) based on speed, size, and cost per bit.
-
Goal: Achieve average access time close to fastest level (cache) with the capacity of the slowest level (disk).
-
Principle of Locality:
-
Temporal Locality: Recently accessed items likely to be accessed again soon.
-
Spatial Locality: Access to an item likely leads to access to nearby addresses.
-
Cache Memory (Crucial for Performance)
A small, fast SRAM memory between CPU and main memory.
-
Role: Holds copies of frequently used main memory blocks. Reduces effective access time dramatically.
-
Key Parameters:
-
Hit: Data found in cache.
-
Miss: Data not in cache, must fetch from main memory.
-
Hit Ratio (h): Fraction of accesses that are hits.
-
Miss Ratio (m): $$\displaystyle m = 1 - h $$.
-
Access Times: $$\displaystyle t_c $$ (cache), $$\displaystyle t_m $$ (main memory).
-
Effective Access Time (EAT) Formula:
$$\boxed{\text{EAT} = h \times t_c + (1 - h) \times (t_c + t_m)}$$
For a simple write-through cache. If $$\displaystyle t_c \ll t_m $$, a high hit ratio yields near-cache speed.
Cache Mapping Techniques
Determines where a main memory block can be placed in cache.
| Technique | How it Works | Pros | Cons |
|---|---|---|---|
| Direct Mapped | Each memory block maps to exactly one cache line (index = block number mod #lines). | Simple, fast hardware. | High conflict misses (multiple blocks map to same line). |
| Fully Associative | A block can be placed in any cache line. | Lowest conflict misses. | Complex, slow (must search all lines in parallel). |
| Set-Associative | Compromise. Cache divided into sets (each with k lines). Block maps to a specific set, can go into any line within it. k-way set-associative. | Balance of cost & performance. | More complex than direct, less flexible than fully associative. |
Common Pitfall: Confusing associative memory (content-addressable, searches by data) with set-associative cache (a cache organization technique).
L1, L2, L3 Cache
-
L1 Cache: Smallest, fastest, integrated on CPU die. Split into L1i (instruction) and L1d (data).
-
L2 Cache: Larger, slightly slower than L1. May be on-die or on-package.
-
L3 Cache: Largest, slowest of on-chip caches. Shared among multiple CPU cores. Acts as a "victim cache" for L2.
VI. Input/Output (I/O) Organization & Data Transfer
Data Transfer Methods Comparison
| Method | CPU Involvement | Data Path | Speed | Use Case |
|---|---|---|---|---|
| Programmed I/O (PIO) | High. CPU executes I/O instructions for each byte/word. | CPU ↔ I/O | Slowest | Simple devices, early systems. |
| Interrupt-Driven I/O | Medium. CPU starts transfer, device interrupts when ready/next byte. | CPU ↔ I/O | Slow-Medium | Reduces CPU polling, but CPU still moves data. |
| Direct Memory Access (DMA) | Low. CPU programs DMA controller, then continues. DMA controller moves data directly between I/O and memory. | I/O ↔ Memory | Fastest | High-speed devices (disk, network, graphics). |
Direct Memory Access (DMA)
-
Working:
-
CPU initializes DMA controller (source, destination address, count).
-
CPU continues other tasks.
-
DMA controller takes control of the system bus (bus mastering).
-
DMA controller transfers data directly between I/O device and memory.
-
DMA controller interrupts CPU upon completion.
-
-
Requirements: DMA controller, ability to temporarily suspend CPU (bus arbitration).
-
Key Difference from PIO/Interrupt: CPU does not move the data bytes. It only sets up and is notified.
Serial vs. Parallel Data Transfer
| Feature | Parallel Transfer | Serial Transfer |
|---|---|---|
| Lines Used | Multiple lines (e.g., 8, 16, 32 bits at once). | Single line (bit-by-bit). |
| Speed (Distance) | Fast for short distances (on a circuit board). | Slower per bit, but scales better over distance. |
| Cost/Complexity | High (more wires, connectors, crosstalk). | Low (1 wire, less interference). |
| Why Use Serial? | Long-distance communication (USB, SATA, PCIe), reduces cost/crosstalk, high clock speeds compensate (e.g., PCIe 5.0). |
Strobe Method of Data Transfer
A simple asynchronous handshaking method for parallel transfer.
-
Source places data on data lines.
-
Source activates strobe signal (e.g.,
DataValidpulse) to indicate data is ready. -
Destination reads data when it sees strobe.
-
Destination may send an acknowledgment pulse back.
Diagram: Shows two devices with data lines and a single strobe line from source to destination.
VII. Interrupts
Definition & Concept
An interrupt is a signal that suspends the CPU's current instruction stream to service an event (I/O completion, timer, error), then resumes.
- Interrupt-Driven I/O: Device interrupts CPU when it's ready to send/receive data, avoiding wasteful polling.
Types of Interrupts
| Type | Description | Example |
|---|---|---|
| Hardware vs. Software | Hardware: External pin signal. Software: INT instruction. |
Hardware: Keyboard press. Software: System call. |
| Maskable vs. Non-Maskable (NMI) | Maskable: Can be ignored/disabled by CPU. NMI: Cannot be ignored, highest priority. | Maskable: Disk I/O. NMI: Power failure, hardware error. |
| Vectored vs. Non-Vectored | Vectored: Interrupt device provides interrupt service routine (ISR) address. Non-Vectored: CPU must poll devices to find source. | x86 uses vectored interrupts (interrupt number → IVT entry). |
Interrupt Handling with Priorities
When multiple interrupts occur:
-
CPU prioritizes them (hardwired or programmable).
-
CPU saves state (PC, PSW) of current program.
-
CPU disables lower-priority interrupts (optional).
-
CPU jumps to ISR of highest-priority pending interrupt.
-
ISR executes, acknowledges the device.
-
ISR restores state and returns (via
IRET).
Interrupt Handling in 8086 Microprocessor
-
Interrupt Vector Table (IVT): Located at physical address 0x00000 in memory. Contains 256 entries (4 bytes each: CS:IP of ISR).
-
Interrupt Types:
-
Hardware Interrupts:
INTR(maskable, vectored via interrupt number on data bus),NMI(non-maskable, fixed vector 2). -
Software Interrupts:
INT n(software-generated, vectored). -
Exceptions: Internal errors (Divide Error, Single Step).
-
-
Sequence:
-
CPU completes current instruction.
-
For
INTR, CPU sendsINTApulse. Interrupt controller (8259A) places interrupt type number on data bus. -
CPU uses type number
nto fetch ISR address from IVT entryn × 4. -
CPU pushes FLAGS, CS, IP, and jumps to ISR.
-
ISR must save/restore all registers it uses and end with
IRET.
-
VIII. Advanced Processor Architectures & Parallelism
RISC vs. CISC
| Feature | CISC (Complex ISA) | RISC (Reduced ISA) |
|---|---|---|
| Goal | Do more per instruction (complex ops). | Simpler, faster instructions. |
| Instruction Size | Variable (1-15 bytes). | Fixed (typically 4 bytes). |
| Addressing Modes | Many (complex, memory-to-memory). | Few (register-register, load/store). |
| Microcode | Typically used (micro-programmed CU). | Not used (hardwired CU). |
| Registers | Few (8-16). | Many (16-32+). |
| Pipeline | Difficult (variable cycles/instruction). | Easy (regular, single-cycle ops). |
| Examples | x86, VAX | ARM, MIPS, RISC-V |
| Why RISC Preferred? | Enables deep pipelining, higher clock speeds, simpler compiler design. Dominates mobile, embedded, and high-performance computing (via ARM). |
Instruction-Level Parallelism (ILP) vs. Thread-Level Parallelism (TLP)
| Aspect | ILP | TLP |
|---|---|---|
| Granularity | Single instruction stream. Exploits parallelism within a single thread/process. | Multiple threads/processes. Exploits parallelism across threads. |
| Goal | Execute multiple instructions from same thread simultaneously. | Execute multiple threads simultaneously. |
| Hardware | Pipelining, Superscalar (multiple ALUs), Out-of-Order Execution (OoO), Speculation. | Multicore, Multithreading (SMT/Hyper-Threading). |
| Example | A superscalar processor issues 4 instructions per cycle from one thread. | A 4-core CPU with SMT runs 8 threads concurrently (2 per core). |
| Key Challenge | Dependencies (data, control) limit parallelism. | Resource contention (cache, memory bandwidth) between threads. |
Pipelining
-
Concept: Divide instruction execution into k stages (Fetch, Decode, Execute, Memory, Writeback). Multiple instructions overlap in different stages.
-
4-Segment Pipeline Space-Time Diagram:
Time | Stage 1 | Stage 2 | Stage 3 | Stage 4 ---- | ------- | ------- | ------- | ------- t1 | I1 | | | t2 | I2 | I1 | | t3 | I3 | I2 | I1 | t4 | I4 | I3 | I2 | I1 t5 | I5 | I4 | I3 | I2-
Speedup (Ideal): ≈ k (for large number of instructions).
-
Hazards Limit Speedup:
-
Structural: Resource conflict (e.g., two instructions need memory at same stage).
-
Data: Dependency (e.g.,
ADD R1, R2followed bySUB R3, R1). -
Control: Branch/jump uncertainty (pipeline fetches wrong instructions).
-
-
Multicore Processor Architecture
-
Basic Structure: A single chip containing multiple independent CPU cores (each with its own ALU, registers, L1 cache).
-
Shared Resources: Typically share L2/L3 cache, memory controller, and I/O.
-
Challenge: Cache coherence (ensuring all cores see a consistent view of memory). Solved by coherence protocols (e.g., MESI).
-
Example: Intel Core i7 (4-8 cores), ARM big.LITTLE (performance + efficiency cores).
Vector Processing
-
Concept: A single instruction operates on multiple data elements simultaneously (SIMD - Single Instruction, Multiple Data).
-
Example:
VADD V1, V2, V3adds 8 pairs of 32-bit integers from vectors V2 and V3, stores result in V1. -
Use Cases: Multimedia, scientific computing, AI/ML (matrix operations).
-
Modern Implementations: SIMD extensions (x86: SSE, AVX; ARM: NEON).
IX. Specific Processor & I/O Device Examples
Architecture of ARM Processor (Basic Overview)
-
RISC Philosophy: Load/Store architecture, fixed 32-bit instructions (ARM state), large register file (16 x 32-bit).
-
Key Features:
-
Registers: R0-R12 (general), R13 (SP), R14 (LR), R15 (PC). Current Program Status Register (CPSR) holds flags/status bits.
-
Pipelining: Classic 3-stage (Fetch, Decode, Execute) or 5-stage.
-
Modes: User, FIQ, IRQ, Supervisor, etc. (for OS/exception handling).
-
Load/Store: Only
LDR/STRaccess memory. ALU ops are register-register. -
Conditional Execution: Most instructions can be conditionally executed (e.g.,
ADDEQ). -
Thumb/Thumb-2: 16-bit compressed instruction set for code density.
-
-
Modern ARM: Supports multi-core (big.LITTLE), SIMD (NEON), virtualization.
8086 Microprocessor Interrupts (Recap from Section VII)
-
IVT at 0x0000:0000, 256 entries (4 bytes each).
-
INTR: Maskable, external. Interrupt number provided by 8259A PIC. -
NMI: Non-maskable, pin on CPU. Fixed vector 2. -
INT n: Software interrupt. -
Exceptions: Internal (Divide Error, INT3, INTO, etc.).
-
Response: CPU pushes FLAGS, clears TF/IF, pushes CS:IP, jumps to ISR from IVT. ISR ends with
IRET.
Optical Disks (CD/DVD/Blu-ray)
-
Principle: Store data as pits and lands on a spiral track. Read by a laser beam reflected from the surface.
-
Types:
-
CD-ROM: ~700 MB, 780 nm laser.
-
DVD: 4.7-17 GB, 650 nm laser (smaller pits).
-
Blu-ray: 25-50 GB, 405 nm blue-violet laser (even smaller pits).
-
-
Key Feature: Rotational speed varies (Constant Linear Velocity - CLV) to maintain constant data density.
-
Use: Primarily for read-only media (software, movies). Rewritable versions exist (CD-RW, DVD±RW).
Key Formulas & Concepts for Calculation
- Effective Access Time (Cache):
$$\text{EAT} = h \cdot t_c + (1 - h) \cdot (t_c + t_m)$$
Where $h$ = hit ratio, $$\displaystyle t_c $$ = cache access time, $$\displaystyle t_m $$ = main memory access time.
- Speedup from Pipelining (Ideal):
$$S \approx k \quad \text{(for large N, where k = number of stages, N = instructions)}$$
Actual speedup reduced by hazards.
- Memory Address Space:
$$\text{Max Addressable Memory} = 2^{\text{Address Bus Width}} \times \text{Addressable Unit (usually 1 byte)}$$
Final Exam Strategy: For 7-mark questions, always start with a clear definition, followed by a diagram/flowchart/table where applicable, then a concise explanation with examples. For calculation questions (like EAT), write the formula first, substitute values, and box the final answer.