UNIT 3: Computer Organization & Architecture
I. FUNDAMENTAL COMPUTER STRUCTURE & CPU ORGANIZATION
Basic Functional Units of a Computer
A computer consists of five core functional units that work together:
-
Input Unit: Accepts data/instructions from external world (e.g., keyboard, mouse). Converts to machine-readable form.
-
Memory Unit (Main/Primary): Stores programs and data. Classified as RAM (volatile, read/write) and ROM (non-volatile, read-only).
-
Arithmetic and Logic Unit (ALU): Performs all arithmetic (+, -, ×, ÷) and logical (AND, OR, NOT, comparisons) operations.
-
Control Unit (CU): Directs and coordinates all operations. Fetches, decodes, and executes instructions. Generates control signals.
-
Output Unit: Converts machine-coded results to human-readable form (e.g., monitor, printer).
[!TIP] Exam Focus: Be prepared to draw the block diagram showing interaction of these five units.
General Register Organization
Registers are small, fast storage locations inside the CPU for temporary data/instruction holding.
-
Special-Purpose Registers:
-
MAR (Memory Address Register): Holds the address of memory location to be accessed.
-
MDR (Memory Data Register): Holds data to be written to or read from memory.
-
IR (Instruction Register): Holds the current instruction being executed.
-
PC (Program Counter): Holds the address of the next instruction to be fetched.
-
-
General-Purpose Registers (GPRs): Used for holding operands and intermediate results during execution (e.g., AX, BX in x86). Their number and size (16/32/64-bit) define the CPU's register architecture.
[!TIP] Common Pitfall: Do not confuse MAR/MDR with cache. They interface with main memory, not cache.
Basic Structure of a Computer System
-
CPU: The brain. Contains CU and ALU.
-
Memory System: Hierarchy (Registers → Cache → Main Memory → Secondary Storage).
-
I/O Subsystem: Manages communication with external devices via I/O controllers and interfaces.
System Bus Structure
A bus is a group of parallel lines carrying data, address, and control signals.
-
Data Bus: Bidirectional. Carries data and instructions between CPU, memory, and I/O. Width (number of lines) determines data transfer volume per cycle.
-
Address Bus: Unidirectional (CPU → Memory/I/O). Carries memory/I/O addresses. Width determines maximum addressable memory space ($$\displaystyle 2^{\text{width}} $$ locations).
-
Control Bus: Carries control signals (Read, Write, Interrupt, Clock, etc.) to coordinate operations.
[!TIP] Key Formula: Maximum Memory = $$\displaystyle 2^{\text{Address Bus Width}} $$ bytes (if each address points to 1 byte).
Instruction Cycle (Fetch-Decode-Execute)
The process of executing a single instruction.
-
Fetch: PC → MAR → Memory → MDR → IR. PC incremented.
-
Decode: CU decodes opcode in IR, determines operation and operands.
-
Execute: CU sends control signals to ALU/memory/I/O to perform the operation.
-
Interrupt: If an interrupt occurs, current PC/PSW saved, ISR executed, then return.
Flowchart:
Start → Fetch → Decode → Execute → Interrupt? → Yes → Service Interrupt → Return → Next Instruction
↓ No
Next Instruction
[!TIP] Exam Question: Always include the Interrupt check as a separate step in the cycle flowchart.
II. INSTRUCTION SET ARCHITECTURE (ISA) & FORMATS
Instruction Set Architecture (ISA)
-
Definition: The interface between hardware and software. It defines all programmer-visible components (registers, memory, data types) and the set of instructions the processor can execute.
-
Role: Acts as a contract. Software (compiler, OS) is written for an ISA; different CPUs can implement the same ISA (e.g., x86: Intel, AMD).
Instruction Formats
The binary layout of an instruction. Common types based on number of address fields:
| Format | Example (A = opcode) | Address Fields | Typical Use |
|---|---|---|---|
| Zero-Address | ADD |
0 | Stack Architecture (operands implicitly from stack) |
| One-Address | ADD A |
1 | Accumulator Architecture (one operand is implicit ACC) |
| Two-Address | ADD R1, A |
2 | One operand is also destination (e.g., R1 ← R1 + A) |
| Three-Address | ADD R1, R2, R3 |
3 | Clear design (e.g., R1 ← R2 + R3), needs more bits |
Design Considerations: Instruction length, number of addresses, addressing modes, and opcode size trade-off between code density and hardware complexity.
Addressing Modes
Specifies how the operand for an instruction is located.
| Mode | How Operand is Found | Example (x86-like) | Key Feature |
|---|---|---|---|
| Implied | Operand is implicit in instruction | CLA (Clear ACC) |
No address field |
| Immediate | Operand is part of instruction | MOV R1, #5 |
Fast, constant value |
| Direct | Address field gives memory address | ADD R1, 1000 |
Single memory access |
| Indirect | Address field points to location holding address | ADD R1, @1000 |
Allows dynamic addressing |
| Register | Operand is in a specified register | ADD R1, R2 |
Very fast |
| Register Indirect | Register holds address of operand | ADD R1, (R2) |
Pointer-like behavior |
| Displacement | Address = Base Reg + Offset | ADD R1, 1000(R2) |
Array/struct access |
| Relative | Address = PC + Offset | JMP +10 |
Position-independent code |
| Stack | Operand is at top of stack | PUSH AX |
Implicit, for stack machines |
[!TIP] Exam Focus: For each mode, be ready to give an example and state how many memory accesses are needed to get the operand.
III. CONTROL UNIT DESIGN
Hardwired Control Unit
-
Principle: Control signals are generated by fixed logic circuits (gates, decoders). The instruction decoder and a sequence counter generate timing signals based on a state diagram.
-
Block Diagram:
DiagramCANVAS: A diagram showing Instruction Register feeding into a Decoder. The Decoder outputs, along with timing signals from a Sequence Counter/Clock, go into a "Hardwired Logic" block (combinational logic). This block outputs the set of Control Signals (e.g., Read, Write, ALUop, etc.). -
Advantages: Fast (no memory access for microcode). Efficient for simple/complex ISAs with good optimization.
-
Disadvantages: Inflexible (difficult to modify/debug). Design complexity increases exponentially with instruction set. Not suitable for complex ISAs.
Micro-programmed Control Unit
-
Principle: Control signals are stored as microinstructions in a Control Memory (CM). A micro-program (sequence of microinstructions) generates control signals for each machine instruction.
-
Key Terms:
-
Control Memory (CM): Stores microprogram. Usually ROM.
-
Micro-instruction: A word in CM. Contains bits for control signals (vertical) or micro-commands (horizontal).
-
Micro-program: A sequence of microinstructions that implements one machine instruction.
-
Micro-program Sequencer: Generates address of next microinstruction (next-address logic).
-
-
Advantages: Flexible, easy to design/debug/modify. Simplifies handling of complex instructions.
-
Disadvantages: Slower (extra memory access per microinstruction). Requires CM storage.
Comparison: Hardwired vs. Micro-programmed
| Feature | Hardwired Control | Micro-programmed Control |
|---|---|---|
| Speed | Faster (no CM access) | Slower (CM access per step) |
| Flexibility | Rigid, hard to change | Flexible, easy to modify |
| Design Complexity | High for complex ISAs | Lower, systematic design |
| Cost | Lower (less hardware) | Higher (CM required) |
| Suitability | Simple, RISC-style ISAs | Complex, CISC-style ISAs |
| Debugging | Difficult | Easier (modify microcode) |
[!TIP] Exam Question: This comparison is very high frequency. Draw a table in your answer.
Microprogram Sequencer
Generates the address for the next microinstruction.
-
Functions: Incrementing, Conditional Branching, Subroutine Call/Return.
-
Types:
-
Incrementing: Simple counter (PC-like).
-
Conditional: Uses condition bits (from ALU flags) for branch decisions.
-
Subroutine: Supports micro-subroutines (push/pop micro-PC).
-
Control Word
-
Definition: The binary word that represents all control signals active at a given time.
-
Hardwired: The output of the combinational logic block at a cycle.
-
Microprogrammed: One field/format of a microinstruction. Can be vertical (encoded, few bits) or horizontal (one bit per signal, wide, fast).
IV. ARITHMETIC & LOGIC UNIT (ALU) OPERATIONS
Floating-Point Arithmetic (Addition/Subtraction)
Representation: $$\displaystyle N = (-1)^s \times M \times 2^E $$ (Sign bit s, Mantissa M, Exponent E).
Algorithm for Addition/Subtraction (for $A \pm B$):
-
Align Exponents: Shift the mantissa of the number with smaller exponent right until exponents are equal. (Loss of precision possible).
-
Add/Subtract Mantissas: Perform operation on aligned mantissas.
-
Normalize Result: Shift result mantissa left/right to have one non-zero digit before the binary point. Adjust exponent accordingly.
-
Round: Apply rounding (e.g., guard, round, sticky bits) to fit mantissa in available bits.
-
Check for Underflow/Overflow: Adjust exponent if it goes out of range.
[!TIP] Diagram: Be ready to draw the flowchart for these 5 steps.
Stack Operations in CPU
-
Implementation: Stack is a LIFO structure in memory. CPU uses a Stack Pointer (SP) register holding the address of the top.
-
PUSH:
SP ← SP - size(for downward stack),M[SP] ← Operand. -
POP:
Operand ← M[SP],SP ← SP + size. -
Example (x86):
PUSH AXdecrements SP by 2, stores AX at[SS:SP].POP BXloads BX from[SS:SP], increments SP by 2. -
Uses: Function call/return (saving return address, parameters), local variables, interrupt handling.
Decimal/BCD Arithmetic
-
BCD (Binary-Coded Decimal): Each decimal digit (0-9) is represented by 4 binary bits.
-
Addition: Add BCD numbers as binary. If result > 9 or carry-out from 4th bit, add
0110(6) to correct and generate carry. -
Subtraction: Use 10's complement or 9's complement methods.
V. MEMORY HIERARCHY & CACHE MEMORY
Memory Hierarchy Concept
A pyramid of storage levels based on speed, cost, and volatility.
Registers (Fastest, Smallest, Costliest)
↓
Cache (SRAM)
↓
Main Memory (DRAM)
↓
Secondary Storage (HDD/SSD) (Slowest, Largest, Cheapest)
- Need: Bridge the speed gap between fast CPU and slow main memory. Exploit Principle of Locality (Temporal: reuse recent data; Spatial: use nearby data).
Cache Memory
-
Importance: Reduces average memory access time. Holds frequently used blocks of main memory.
-
Principle of Locality: Programs tend to access a small set of data/instructions repeatedly over a short time.
Cache Mapping Techniques
Determines where a block from main memory can be placed in cache.
| Technique | How it Works | Pros | Cons |
|---|---|---|---|
| Direct Mapped | Each memory block maps to exactly one cache line (index = block number mod #lines). | Simple, fast hardware. | High conflict misses (multiple hot blocks compete). |
| Fully Associative | Block can be placed in any empty cache line. | Lowest conflict misses. | Complex, slow (search all lines). |
| Set-Associative | Compromise. Cache divided into V sets. Block maps to a set (index), can go in any line within that set (N-way). |
Good balance of cost/performance. | More complex than direct. |
[!TIP] Formula: In
N-way set-associative,#Sets = Cache Size / (Block Size × N).
Cache Levels (L1, L2, L3)
| Level | Location | Size | Speed | Purpose |
|---|---|---|---|---|
| L1 Cache | Inside CPU core | Small (32-64 KB) | Fastest (1-3 cycles) | Critical for single-thread performance. Split into I-Cache & D-Cache. |
| L2 Cache | Inside CPU core (or shared) | Larger (256 KB - 1 MB) | Slower than L1 | Catches misses from L1. Often exclusive or inclusive. |
| L3 Cache | Shared among cores | Largest (8-64 MB) | Slowest cache | Reduces memory traffic between cores. Usually inclusive. |
Associative Memory
-
Concept: Memory where content (not address) determines location. Search is done in parallel.
-
Organization: Each cell has a key (address) and value (data). A comparator matches input key with all stored keys simultaneously.
-
Hardware: Uses match lines and search lines. More expensive than RAM.
-
Comparison with Cache: Cache uses associative lookup within a set/line, but is still indexed by address bits. True associative memory is content-addressable from any input.
Effective Access Time (EAT) Calculation
$$ \boxed{\text{EAT} = (\text{Hit Ratio} \times \text{Cache Access Time}) + (\text{Miss Ratio} \times \text{Miss Penalty})} $$
Where:
-
Miss Penalty = Time to access main memory + transfer block + possibly update cache.
-
Hit Ratio (h) = Fraction of accesses found in cache.
-
Miss Ratio = 1 - h.
Example: If $$\displaystyle h = 0.95 $$, $$\displaystyle t_c = 10\,ns $$, $$\displaystyle t_m = 100\,ns $$, then:
$$ \text{EAT} = (0.95 \times 10) + (0.05 \times 100) = 9.5 + 5 = \boxed{14.5\,ns} $$
Internal Organization of RAM & ROM Chips
-
Address Lines: Select a specific row/column/cell.
-
Data Lines: Bi-directional for read/write.
-
Control Pins:
-
Chip Select (CS): Enables the chip.
-
Read (RD) / Write (WR): Determines operation direction.
-
Output Enable (OE): Tri-states output drivers (for shared bus).
-
VI. INPUT/OUTPUT (I/O) ORGANIZATION
Interrupts
-
Definition: An asynchronous event that causes the CPU to temporarily suspend its current program and execute a special routine (Interrupt Service Routine - ISR).
-
Need: Allows CPU to respond to urgent I/O events (e.g., keypress, data ready) without constant polling (wasting CPU cycles).
-
Types:
-
Maskable vs. Non-maskable (NMI): Can CPU ignore it? NMI (e.g., hardware failure) cannot be masked.
-
Software vs. Hardware: Software interrupt (
INT ninstruction) vs. Hardware signal (from device). -
Vectored vs. Non-Vectored: Does interrupt provide its ISR address? Vectored does (e.g., 8086's
INTR).
-
Interrupt Handling (Process & Priority)
-
Interrupt Occurs: Device sends signal to CPU.
-
Finish Current Instruction: CPU completes current instruction.
-
Save Context: PC and PSW (Program Status Word) pushed onto stack.
-
Branch to ISR: Load PC with ISR address (from Interrupt Vector Table for vectored interrupts).
-
Execute ISR: Service the device.
-
Return:
IRETinstruction pops saved PC/PSW, resumes original program.
Priority Handling: Multiple simultaneous interrupts resolved by:
-
Software Polling: CPU asks each device in priority order.
-
Daisy Chain (Hardware): Devices in series,
INTAsignal cascades. -
Parallel Priority (Interrupt Controller): e.g., 8259 PIC, uses priority encoder.
8086 Specifics:
-
INTR: Maskable, general-purpose interrupt. Type determined by external interrupt controller (like 8259).
-
NMI: Non-maskable. Type 2 interrupt.
-
Interrupt Vector Table (IVT): Located at
0000:0000in memory. Contains 4-byte pointers (CS:IP) for each of 256 interrupt types.
Direct Memory Access (DMA)
-
Working: A DMA Controller takes over the system bus to transfer data directly between I/O device and memory, bypassing the CPU.
-
DMA Controller: Has registers: DR (Data Register), MAR, WC (Word Count), Control/Status.
-
Requirements: CPU must be disabled (or cycles "stolen") during transfer. DMA controller must be initialized by CPU (set source, destination, count).
-
Transfer Modes:
-
Burst/Block: Entire block transferred in one DMA cycle. Fast, CPU idle long.
-
Cycle Stealing: DMA takes one bus cycle at a time, interleaving with CPU. Slower, but CPU not blocked.
-
Transparent: DMA only when CPU not using bus.
-
[!TIP] Key Difference from Programmed I/O: In PIO, CPU executes
IN/OUTfor every byte/word. In DMA, CPU sets up DMA, then DMA handles bulk transfer.
Data Transfer Methods
Serial vs. Parallel Transfer
| Feature | Parallel Transfer | Serial Transfer |
|---|---|---|
| Lines | Multiple data lines (e.g., 8, 16, 32) | Single data line |
| Speed | Faster (one byte/cycle) | Slower (bit-by-bit) |
| Distance | Short (on-board, chip-to-chip) | Long (between systems, USB, SATA) |
| Cost/Complexity | Higher (more wires, skew issues) | Lower (fewer wires, less crosstalk) |
| Why Serial? | N/A | Used for long distances, lower cost, easier clock synchronization (embedded clock). |
Strobe Method (Handshaking)
A simple asynchronous synchronization method for data transfer.
-
Sender: Places data on line, sends a strobe pulse.
-
Receiver: Reads data when it detects the strobe pulse.
-
Limitation: No acknowledgment from receiver (unlike full handshaking with
Request/Acknowledge).
I/O Channels
A specialized processor (channel) that manages I/O operations independently.
-
Concept: Offloads I/O tasks from CPU. CPU sends channel program (list of I/O commands).
-
Types:
-
Selector Channel: Connects one high-speed device (e.g., tape drive) at a time to main memory. Dedicated path.
-
Multiplexor Channel: Connects multiple slower devices (e.g., terminals) to main memory, interleaving their operations. Time-shared path.
-
VII. ADVANCED CPU ARCHITECTURES & PARALLELISM
RISC vs. CISC Architectures
| Feature | CISC (Complex ISA) | RISC (Reduced ISA) |
|---|---|---|
| Goal | Do more per instruction (complex ops) | Simpler, faster instructions |
| Instruction Set | Large, variable-length, many addressing modes | Small, fixed-length, few addressing modes |
| Registers | Few (e.g., 8 in x86) | Many (e.g., 16-32 in ARM, MIPS) |
| Memory Access | Memory-to-memory ops allowed | Load/Store only (memory access only via dedicated load/store instructions) |
| Control Unit | Often micro-programmed | Hardwired for speed |
| Pipelining | Difficult (variable cycles) | Easy (regular, fixed cycles) |
| Examples | x86, VAX | ARM, MIPS, RISC-V |
| Why RISC Preferred? | N/A | Simpler design, easier to pipeline, higher clock speeds, lower power. Dominant in mobile/embedded (ARM) and many servers (RISC-V, MIPS). |
Instruction-Level Parallelism (ILP) vs. Thread-Level Parallelism (TLP)
| Aspect | ILP | TLP |
|---|---|---|
| Granularity | Within a single thread/instruction stream | Across multiple threads |
| Goal | Execute multiple instructions from same thread simultaneously. | Execute multiple threads (or processes) simultaneously. |
| Techniques | Pipelining, Superscalar (multiple ALUs), Out-of-Order Execution, Branch Prediction. | Multi-threading (hardware threads per core), Multi-core (multiple cores). |
| Hardware Focus | CPU core internal resources (multiple functional units). | CPU core count, thread scheduling hardware. |
| Example | A superscalar core issuing 4 instructions/cycle. | A 4-core CPU running 4 threads concurrently. |
Pipelining
-
Concept: Break instruction execution into stages (e.g., Fetch, Decode, Execute, Memory, Writeback). Multiple instructions overlap in different stages.
-
Speedup (Ideal): If
kstages, speedup ≈k(ignoring hazards). -
Hazards:
-
Structural: Resource conflict (two instructions need same unit).
-
Data: Dependency (e.g.,
ADD R1, R2followed bySUB R3, R1). -
Control: Branch/jump changes PC, instructions after branch are wrong.
-
-
Space-Time Diagram (4-stage):
Cycle: 1 2 3 4 5 6 7 I1: F D E W I2: F D E W I3: F D E W I4: F D E WThroughput = 1 instruction/cycle after pipeline fill.
Multicore Processor Architecture
-
Basic Structure: Multiple independent CPU cores (each with own ALU, CU, L1 cache) integrated on a single chip. Share L2/L3 cache and memory controller.
-
Advantages:
-
Increased throughput (true parallelism for multi-threaded apps).
-
Better utilization of silicon area.
-
Lower power per core vs. single huge core.
-
Fault tolerance (one core failure may not crash system).
-
Vector Processing & ARM Architecture
-
Vector Processing: Single instruction operates on entire vectors/arrays (SIMD - Single Instruction, Multiple Data). E.g., add two arrays of 64 floats with one
VADDinstruction. Faster than scalar pipelining for data-parallel tasks (multimedia, scientific). -
ARM Processor Architecture (Key Features):
-
Load-Store Architecture: Only load/store instructions access memory. ALU ops only on registers.
-
Register Set: 16-32 32-bit general-purpose registers (R0-R15). R13=SP, R14=LR, R15=PC.
-
Fixed 32-bit Instruction Length (in ARM mode).
-
Conditional Execution: Most instructions can be conditionally executed (reduces branches).
-
Power Efficient: Dominant in mobile/embedded (phones, IoT).
-
VIII. SECONDARY STORAGE & SPECIALIZED MEMORIES
Optical Disks
-
Principle: Data stored as pits and lands on a spiral track. Read by laser beam reflecting differently from pits/lands.
-
CD (Compact Disc): ~700 MB, 120mm diameter. Single layer, single-sided.
-
DVD (Digital Versatile Disc): Uses smaller pits & dual-layer. 4.7 GB (single-layer), 8.5 GB (dual-layer).
-
Blu-ray: Uses blue-violet laser (shorter wavelength), allows smaller pits. 25 GB (single-layer), 50 GB (dual-layer). Used for HD video.
[!TIP] Low Frequency: This topic appears as a short 3-mark question. Know the trend: Capacity increases with shorter wavelength and multi-layer technology.
END OF UNIT 3 NOTES