UNIT 3: Computer Organization & Architecture (Instrumentation Computing Aspects)
1. Fundamental Computer Structure
Von Neumann Architecture
-
Components: Memory (stores data/instructions), Control Unit (CU), Arithmetic Logic Unit (ALU), Input/Output (I/O) systems.
-
Stored-Program Concept: Instructions and data share the same memory and bus; program is stored as binary in memory.
-
Key Limitation: Von Neumann bottleneck—sequential access to memory limits speed.
Instruction Cycle
-
Fetch: Read instruction from memory address in Program Counter (PC) into Instruction Register (IR). PC incremented.
-
Decode: Control Unit interprets opcode, determines operation and operands.
-
Execute: ALU performs operation; result stored in register/memory.
[!TIP] Common exam question: Draw flowchart showing fetch, decode, execute with register transfers (e.g.,
MAR ← PC,MDR ← Memory[MAR]).
Register Transfer Language (RTL) & Micro-operations
-
Micro-operations: Elementary operations on registers/memory (e.g.,
R1 ← R2,PC ← PC + 1). -
RTL: Symbolic notation to describe micro-operations (e.g.,
MAR ← (PC)). -
Types: Register transfer, arithmetic, logic, shift.
Common Bus System
-
Purpose: Shared pathway for data/address/control signals among registers, memory, ALU.
-
Design: Multiplexers select source; tri-state buffers enable one device at a time.
-
Example: 8-register system with bus; control signals
SELA,SELBselect inputs to ALU.
[!DIAGRAM] Search: "common bus system computer organization diagram"
2. Control Unit Design
Hardwired Control
-
Design: Combinational logic (gates) generates control signals directly from instruction decoder and timing signals.
-
Timing: Fixed by clock pulses; each instruction has predefined sequence.
-
Advantages: Fast, no memory access.
-
Disadvantages: Inflexible; difficult to modify instruction set.
Microprogrammed Control
-
Concept: Control signals stored as microinstructions in control memory.
-
Microinstruction: Low-level control words; each activates specific micro-operations.
-
Advantages: Flexible, easier to design/debug; supports complex instructions.
-
Disadvantages: Slower (extra memory access).
Micro-instruction Format
-
Fields:
-
Operation field: Control signals for parallel micro-operations.
-
Address field: Next microinstruction address (for branching).
-
Next address field: Sequencing logic (e.g.,
CMAR ← address field).
-
-
Vertical vs Horizontal:
-
Vertical: Compact, encoded fields (fewer bits, slower decoding).
-
Horizontal: Wide, one bit per control signal (parallelism, faster).
-
Microprogram Sequencer
-
Block Diagram: Control Memory → Microinstruction Register → Sequencer Logic → Control Memory Address Register (CMAR).
-
Operation: Fetches microinstruction, updates CMAR based on next address logic (conditional branches).
[!TIP] Exam often asks: "Explain microprogram sequencer with block diagram."
Control Signal Logic
-
Derive Boolean expressions from instruction-time step table.
-
Example: For instruction
I1active inT2, signalS5 = I1'·T2 + I3·T4(whereI1'means instruction decode forI1).
[!TIP] Past paper: Given table of control signals for instructions over time steps, find expressions for specific signals (e.g., S5, S6, S10).
3. Arithmetic & Logic Unit (ALU)
Multiplication Circuit
-
Algorithm: Shift-and-add (e.g., Booth’s algorithm reduces steps).
-
Design Challenges:
-
Speed: Sequential addition of partial products is slow.
-
Hardware: Array multipliers use many adders (fast but costly).
-
Overflow: Need wider result (2n bits for n-bit operands).
-
Floating-Point Representation (IEEE 754)
-
Single Precision (32-bit):
-
Sign bit (1 bit)
-
Exponent (8 bits, biased by 127)
-
Mantissa (23 bits, implicit leading 1)
-
-
Double Precision (64-bit): 1 sign, 11 exponent, 52 mantissa.
\boxed{\text{Value} = (-1)^S \times (1.M) \times 2^{(E - \text{bias})}}
Floating-Point Addition/Subtraction
-
Align exponents: Shift mantissa of smaller exponent right.
-
Add/subtract mantissas: Include sign handling.
-
Normalize result: Shift left/right to maintain 1.xxxx form; adjust exponent.
-
Round: Guard bits for precision.
-
Check overflow/underflow.
[!DIAGRAM] Search: "floating point addition alignment normalization flowchart"
ALU Design Considerations
-
Integer operations: Faster, simpler circuits (carry-lookahead adders).
-
Floating-point: More complex (alignment, rounding, special values like NaN/Infinity).
-
Implementation: Separate integer and FP units, or unified with microcode.
4. Instruction Set Architecture
Instruction Format
-
Bit allocation: Opcode field (identifies operation), address fields (source/destination), immediate/constant bits.
-
Example: 16-bit instruction with 6-bit opcode, two 5-bit addresses → 64 instructions, 32 addressable locations.
\boxed{\text{Max instructions} = 2^{\text{opcode bits}}, \quad \text{Address space} = 2^{\text{address bits}}}
Addressing Modes
| Mode | Description | Example (Assume x at address 500) |
|---|---|---|
| Immediate | Operand in instruction | ADD R1, #5 → R1 = R1 + 5 |
| Direct | Address field gives operand address | ADD R1, 500 → R1 = R1 + Memory[500] |
| Indirect | Address field points to address of operand | ADD R1, @500 → R1 = R1 + Memory[Memory[500]] |
| Register | Operand in register | ADD R1, R2 |
| Register-Indirect | Register contains operand address | ADD R1, (R2) |
| Relative | PC-relative addressing for branches | JUMP +10 → PC = PC + 10 |
| Indexed | Address + index register | LOAD R1, 500(R2) → Address = 500 + R2 |
Instruction Encoding
-
Opcode decoding: Hardwired decoder or microprogrammed.
-
Address calculation: Depends on mode (e.g., direct uses address field; relative uses PC + offset).
5. Input/Output Organization
Data Transfer Modes
| Mode | Mechanism | CPU Involvement | Speed | Use Case |
|---|---|---|---|---|
| Programmed I/O (Polling) | CPU checks status register repeatedly | High (busy-wait) | Slow | Simple devices |
| Interrupt-Driven I/O | Device signals CPU when ready | Medium (interrupt handler) | Moderate | Interactive I/O |
| Direct Memory Access (DMA) | DMA controller manages transfer | Low (initiate/interrupt only) | Fast | High-speed disks, video |
I/O Interface
-
Role: Buffering, signal conversion, handshaking (control signals like
READY,BUSY). -
Handshaking: Synchronizes CPU and I/O (e.g.,
STROBEandACKsignals).
I/O Processor (IOP)
-
Dedicated processor handles I/O tasks asynchronously.
-
Operation: CPU programs IOP with memory addresses and transfer count; IOP manages data transfer independently, interrupts CPU on completion.
Synchronous vs Asynchronous Transfer
-
Synchronous: Clock-based; all devices synchronized to common clock. Simple but wastes cycles if devices vary in speed.
-
Asynchronous: Handshaking signals (
REQUEST,ACKNOWLEDGE); flexible for different-speed devices.
Duplex Modes
-
Half-duplex: One direction at a time (e.g., walkie-talkie).
-
Full-duplex: Both directions simultaneously (e.g., telephone).
[!TIP] Applications: Half-duplex for shared channels (radio); full-duplex for high-speed networks.
DMA Controller
-
Block Diagram:
CPU ↔ DMA Controller ↔ Memory ↔ I/O Device
Registers: Command/Status, Address, Count.
-
Operation:
-
CPU programs DMA (address, count, control).
-
DMA requests bus control (holds
BUS REQUEST). -
CPU releases bus (acknowledges
BUS GRANT). -
DMA transfers data directly (read/write cycles).
-
DMA interrupts CPU on completion.
-
DMA Performance (CPU Overhead)
-
CPU overhead = (Cycles to initiate + cycles to respond to interrupt) / (Total cycles for transfer)
-
Example: Transfer
Nwords, each takes 2 cycles (read+write). Total DMA cycles =2N.CPU cycles used =
C_init + C_int.Overhead fraction =
(C_init + C_int) / (2N).
[!TIP] Past paper: Given transfer size, cycles, calculate fraction of CPU time spent.
6. Memory System
Memory Hierarchy
-
Levels: Registers → Cache → Main Memory (RAM) → Secondary (Disk).
-
Principle of Locality:
-
Temporal: Recently accessed data likely reused.
-
Spatial: Nearby addresses likely accessed.
-
-
Significance: Faster, smaller memories closer to CPU reduce average access time.
Cache Memory
-
Organization:
-
Cache size =
Number of blocks × Block size. -
Each block contains tag (memory address), data, valid/dirty bits.
-
-
Address Breakdown (for
n-way set associative):Tag | Set Index | Block OffsetExample: 16 KB cache, 64-byte block, 4-way →
Blocks = 16K/64 = 256; Sets = 256/4 = 64 → Index = log₂(64) = 6 bits.
Mapping Techniques
| Technique | How Address Maps | Pros | Cons |
|---|---|---|---|
| Direct | Index bits select set; tag must match | Simple, fast | High conflict misses |
| Fully Associative | Block can go anywhere; tag compared to all | Low misses | Complex comparator |
| Set-Associative | n blocks per set (e.g., 4-way); tag matches within set |
Balance | Moderate hardware |
[!DIAGRAM] Search: "cache mapping direct set associative diagram"
Cache Performance
-
Hit Ratio (h): Fraction of accesses found in cache.
-
Miss Penalty (p): Time to fetch block from lower level (e.g., main memory).
-
Average Memory Access Time (AMAT):
$$ \text{AMAT} = \text{Hit Time} + (1 - h) \times \text{Miss Penalty} $$
\boxed{\text{AMAT} = t_c + (1 - h) \times t_m}
where \( t_c \) = cache access time, \( t_m \) = main memory access time.
Memory Mapping
-
Significance: Maps program’s logical addresses to physical memory; enables multiprogramming and protection.
-
Effect on Execution: Determines where code/data reside; affects cache behavior and TLB hits.
Virtual Memory (Paging)
-
Paging: Divide memory into fixed-size pages (e.g., 4 KB).
-
Page Table: OS-maintained table mapping virtual pages to physical frames.
-
Translation Lookaside Buffer (TLB): Cache for page table entries (speed up translation).
-
Benefits: Larger address space than physical memory; process isolation; efficient swapping.
Fragmentation
-
Internal: Wasted space within allocated block (fixed partitions).
-
External: Free memory scattered in small blocks (variable partitions).
-
Paging eliminates external fragmentation; internal fragmentation possible (partial page use).
7. Advanced Processor Architectures
Pipelining
-
Principle: Divide instruction execution into
kstages (e.g., IF, ID, EX, MEM, WB); multiple instructions in parallel. -
4-Segment Example:
-
Fetch (F)
-
Decode (D)
-
Execute (E)
-
Write-back (W)
-
-
Advantages: Throughput ≈
1 / (stage time)(ideal speedup ≈k). -
Limitations:
-
Hazards:
-
Structural: Resource conflict.
-
Data: Dependency (e.g.,
ADD R1, R2followed bySUB R3, R1). -
Control: Branches cause pipeline flush.
-
-
Overhead: Pipeline registers, clock skew.
-
[!TIP] Past paper: "Formulate four-segment pipeline" → list stages with operations.
Vector Processing
-
Vector vs Scalar: Vector operates on arrays (e.g.,
C = A + Bfor all elements); scalar one element at a time. -
Architecture: Vector registers, functional units (pipelined), strided memory access.
-
Applications: Scientific computing, simulations, DSP.
Multiprocessor Systems (MIMD)
-
MIMD: Multiple processors execute different instructions on different data.
-
Inter-Processor Communication:
-
Shared memory (with cache coherence protocols).
-
Message passing (networks, buses).
-
-
Challenges: Synchronization, contention, scalability.
RISC Architecture
-
Key Characteristics:
-
Fixed-length instructions (e.g., 32 bits).
-
Load-store design: only load/store access memory; ALU uses registers.
-
Few addressing modes.
-
Hardwired control (fast).
-
Large register set.
-
-
Goal: Simplify instruction set for pipelining efficiency.
8. Specialized Memory & Interconnection
Associative Memory
-
Content-Addressable: Access by data content, not address.
-
Operation: Compare search key with all entries in parallel; return matching address(es).
-
Difference from Cache:
-
Cache uses index+tag mapping; associative memory searches entire memory.
-
Associative memory used for page tables (TLB), caches use set-associative mapping.
-
Shared Bus Architecture
-
Advantages: Simple, low cost, easy to add devices.
-
Contention: Multiple masters compete → need arbitration (centralized: bus arbiter; distributed: daisy-chaining).
-
Performance: Bandwidth shared; arbitration overhead.
Interconnection Structure
-
For multi-processor/I/O systems:
-
Bus: Simple but limited bandwidth.
-
Crossbar: Dedicated paths, high cost.
-
Multistage (e.g., Omega network): Balance cost and performance.
-
Microinstruction Encoding
-
Goal: Minimize control bits while preserving parallelism (mutually exclusive operations can be encoded).
-
Methods:
-
Direct encoding: One bit per signal (horizontal, wide).
-
Field encoding: Group mutually exclusive signals; decode to produce control (vertical, compact).
-
-
Example from Past Paper:
Given microinstructions and activated signals, encode to minimize bits:
-
Group signals that never appear together (e.g.,
aandbinI1andI3? Check table). -
Use field encoding: e.g., Field 1:
a,b,c→ 2 bits; Field 2:d,e,f→ 2 bits, etc. -
Preserve parallelism: Signals in same microinstruction must be in different fields.
-
[!TIP] Past paper: "Find method of encoding microinstructions from table to minimize bits."
Solution Approach:
- List all control signals across microinstructions.
- Identify mutually exclusive groups (signals never active together).
- Assign fields; each field encodes one exclusive group.
- Calculate bits:
ceil(log₂(group size))per field.
- Total bits = sum of field bits.