Skip to content
CY-504 (C) ยท Computer Organization & Architecture/Quick Revision Short Notes

Computer Organization & Architecture (CY-504 (C)) - Unit 5 Short Notes

UNIT 5: Computer Organization & Architecture


I. Fundamental Computer Structure

Basic Functional Units of a Computer System

A digital computer system consists of the following key functional units:

  1. Arithmetic Logic Unit (ALU): Performs all arithmetic (+, -, ร—, รท) and logical (AND, OR, NOT, comparisons) operations.

  2. Control Unit (CU): Directs the operation of the entire system. It fetches, decodes, and executes instructions, generating necessary control signals.

  3. Memory Unit (Main Memory/RAM): Stores data and instructions. It is directly addressable by the CPU.

  4. Input/Output (I/O) Subsystem: Manages communication with the external world (e.g., keyboard, monitor, disk drives).

  5. System Bus: A set of shared electrical paths (wires) that interconnects all major components, allowing them to exchange data, addresses, and control signals.

[!TIP] Exam Focus: Be prepared to draw and label the basic functional units diagram, showing the interconnection via the system bus.

System Bus Structure

A bus is a group of parallel lines carrying information. A typical system bus is divided into three categories:

Bus Type Purpose Lines
Address Bus Carries memory/I/O addresses from CPU to memory/I/O. Unidirectional (CPU โ†’ Memory/I/O)
Data Bus Carries actual data between CPU, memory, and I/O. Bidirectional
Control Bus Carries control signals (Read, Write, Interrupt, Clock, etc.). Mixed direction

Key Point: The width (number of lines) of the address bus determines the maximum addressable memory space ($$\displaystyle 2^n $$ addresses, where $n$ = address bus lines). The width of the data bus determines the amount of data transferred per cycle.

General Register Organization

Registers are small, fast storage locations inside the CPU. Key registers include:

  • Memory Register (MR): Holds data read from or to be written to main memory. Acts as a buffer between CPU and memory.

  • Instruction Register (IR): Holds the current instruction fetched from memory. The CU decodes the instruction in the IR.

  • Program Counter (PC): Also called Instruction Pointer (IP). Holds the memory address of the next instruction to be fetched. Automatically increments after each fetch.

[!TIP] Common Pitfall: Do not confuse the IR (holds the instruction itself) with the PC (holds the address of the instruction).


II. CPU Organization and Instruction Processing

Instruction Set Architecture (ISA)

The ISA is the interface between hardware and software. It defines:

  • The set of instructions the processor can execute.

  • The data types and registers available.

  • The memory addressing modes.

  • The mechanism for memory and I/O access.

  • Examples: x86 (CISC), ARM (RISC), MIPS, RISC-V.

Addressing Modes (with examples)

Addressing mode specifies how the operand of an instruction is located.

Mode How Operand is Specified Example (x86-like) Purpose
Immediate Operand is part of the instruction. ADD R1, #5 Fast, for constants.
Direct Instruction contains the memory address of the operand. ADD R1, 2000 Simple, but address field limits range.
Indirect Instruction contains a register/memory address that points to the operand's address. ADD R1, @R2 Allows dynamic address calculation.
Register Operand is in a CPU register. ADD R1, R2 Very fast.
Register Indirect Register holds the address of the operand in memory. ADD R1, (R2) Efficient for array/string traversal.
Displacement / Indexed Effective Address = Base Address + Displacement. ADD R1, 1000(R2) Good for arrays/structures.
Relative Effective Address = PC + Offset. JUMP +10 Used for position-independent code (PC-relative).
Stack Operand is at the top of the stack. Implicit use of SP. PUSH AX Efficient for procedure calls, local vars.

Instruction Formats

The binary layout of an instruction, typically divided into fields:

  • Opcode (Operation Code): Specifies the operation (e.g., ADD, LOAD).

  • Address Field(s): Specifies register or memory address(es).

  • Mode Field: Specifies the addressing mode.

Common Formats:

  • 3-Address: ADD R1, R2, R3 (R1 = R2 + R3). Program length small, but instruction size large.

  • 2-Address: ADD R1, R2 (R1 = R1 + R2). Common in CISC.

  • 1-Address: ADD R1 (AC = AC + R1). Uses an implicit accumulator (AC).

  • 0-Address (Stack): ADD (Operands from stack top, result pushed back). Used in stack machines.

Instruction Cycle (Fetch-Decode-Execute)

The cycle the CPU repeats for each instruction. A typical flow chart:


START

  |

  V

[FETCH] --> IR <- M[PC]; PC <- PC + 1

  |

  V

[DECODE] --> Decode opcode in IR, determine addressing modes

  |

  V

[FETCH OPERAND(S)] --> If needed, read from memory/register based on addressing mode

  |

  V

[EXECUTE] --> Perform operation in ALU (or memory access/I/O)

  |

  V

[STORE RESULT] --> Write result to destination register/memory

  |

  V

[INTERRUPT CHECK] --> Check for pending interrupts

  |

  V

(Repeat for next instruction)

Stack Operations in CPU

The stack is a LIFO (Last-In-First-Out) data structure, typically implemented in main memory with a Stack Pointer (SP) register holding the address of the top.

  • PUSH (Operand):

    1. Decrement SP.

    2. Store operand at memory location pointed by SP.

  • POP (Operand):

    1. Read operand from memory location pointed by SP.

    2. Increment SP.

Example: PUSH AX (push 16-bit AX register onto stack).

  1. SP = SP - 2 (assuming 16-bit word).
  1. M[SP] โ† AX.

POP AX reverses this.


III. Arithmetic Operations

Floating-Point Arithmetic (Addition/Subtraction)

Flow Chart for Addition/Subtraction (A ยฑ B):


START

  |

  V

1. [ALIGN EXPONENTS] --> Compare exponents. Shift mantissa of smaller exponent number right until exponents are equal. (Loss of precision possible).

  |

  V

2. [ADD/SUBTRACT MANTISSAS] --> Perform addition or subtraction on aligned mantissas.

  |

  V

3. [NORMALIZE RESULT] --> Shift result mantissa left/right to make it a normalized number (1.xxxx for base 2). Adjust exponent accordingly.

  |

  V

4. [ROUND] --> Round the mantissa to fit the format's precision (e.g., guard, round, sticky bits).

  |

  V

5. [CHECK FOR OVERFLOW/UNDERFLOW] --> If exponent too large/small, handle as special case (Infinity, Zero, NaN).

  |

  V

END

Decimal Arithmetic (BCD)

Computers often use Binary-Coded Decimal (BCD) where each decimal digit (0-9) is represented by a 4-bit binary number.

  • Addition/Subtraction: Performed digit-by-digit (like paper-and-pencil). After binary addition of two BCD digits, if the result is >9 or has a carry, +6 (0110) is added to correct it to valid BCD.

  • Example: 9 (1001) + 5 (0101) = 14 (11100 binary). 14 > 9, so add 6 (0110) โ†’ 11100 + 0110 = 101010. The lower 4 bits (1010) = 10 (invalid), so add 6 again โ†’ 1010 + 0110 = 10000. Result: 0001 0000 (BCD for 14).


IV. Control Unit Design

Hardwired Control Unit

The CU is implemented using fixed logic gates and flip-flops. The control signals are generated directly by the decoder based on the opcode and timing signals from a clock.


DiagramCANVAS: A block diagram showing Instruction Register feeding into a Control Signal Generator (combinational logic of gates). The generator outputs multiple control lines (e.g., MemRead, ALUSrc, RegWrite) to datapath components. A Clock signal synchronizes the entire process.
  • Advantages: Very fast (no microinstruction fetch). Efficient for simple, fixed ISA.

  • Disadvantages: Inflexible. Adding new instructions requires rewiring the hardware. Complex to design and test.

Micro-programmed Control Unit

The CU is implemented using a control memory (CM) that stores microinstructions. Each microinstruction is a control word (a set of control signals).


DiagramCANVAS: A block diagram showing a Microprogram Sequencer. It reads a microinstruction address from a micro-program counter (ยตPC). This address points to a location in the Control Memory (CM). The microinstruction (control word) read from CM is loaded into a Microinstruction Register (MIR). The MIR's bits directly become the control signals for the datapath. The sequencer's next address logic (based on IR, condition flags) determines the next ยตPC value.
  • Advantages: Flexible and easy to modify/debug. Changing the control sequence is a matter of updating microcode. Simplifies design of complex ISAs.

  • Disadvantages: Slower than hardwired (extra memory access per microinstruction). Requires additional hardware for CM and sequencer.

Control Word: The binary word stored in CM. Each bit (or group of bits) corresponds to a specific control signal (e.g., 1 = enable, 0 = disable).

Microprogram Sequencer

The hardware that generates the sequence of addresses for reading microinstructions from CM. It uses:

  • Micro-program Counter (ยตPC): Holds address of next microinstruction.

  • Next Address Logic: Determines next ยตPC value based on:

    • Current microinstruction (branching).

    • Opcode from IR (entry point).

    • Condition codes (status flags).

    • External inputs (interrupts).


V. Memory System

Memory Hierarchy

A pyramid of storage technologies with trade-offs between speed, size, and cost per bit.


DiagramCANVAS: A pyramid diagram. Top (smallest, fastest, most expensive): CPU Registers. Next: L1 Cache (on-die). Next: L2 Cache (on-die/off-die). Next: L3 Cache (shared). Next: Main Memory (DRAM). Base (largest, slowest, cheapest): Secondary Storage (SSD/HDD/optical).

Goal: Achieve an average access time close to the fastest level while having the capacity of the slowest level.

Cache Memory

A small, fast SRAM memory placed between CPU and main memory. It stores copies of frequently accessed main memory blocks.

Role in Performance Improvement:

  • Reduces effective access time dramatically due to locality of reference (temporal & spatial).

  • Hides the high latency of main memory (DRAM).

Mapping Techniques:

  1. Direct Mapping:

    • Each memory block maps to exactly one cache line (determined by index bits).

    • Simple, inexpensive hardware.

    • Disadvantage: High conflict misses if two frequently used blocks map to same line.

    
    
    DiagramCANVAS: Main memory blocks numbered 0,1,2,... Cache lines numbered 0,1,2,... Show mapping: Block i maps to line (i mod C), where C = number of cache lines.
  2. Associative Mapping (Fully Associative):

    • A memory block can be placed in any empty cache line.

    • Lowest conflict misses.

    • Disadvantage: Requires searching all cache lines in parallel (expensive comparator hardware).

  3. Set-Associative Mapping (Compromise):

    • Cache is divided into S sets, each with N lines (N-way set-associative).

    • A block maps to a specific set (direct-mapped), but can go into any line within that set (associative).

    • Example: 4-way set-associative: 4 lines per set.

    
    
    DiagramCANVAS: Show cache divided into sets (Set 0, Set 1,...). Main memory block i maps to Set (i mod S). Within that set, it can occupy any of the N lines.

Levels of Cache (L1, L2, L3):

  • L1 Cache: Smallest, fastest, located on the CPU die. Split into L1i (instruction) and L1d (data). Critical for single-thread performance.

  • L2 Cache: Larger, slightly slower than L1. May be on-die or off-die. Often unified (instruction+data).

  • L3 Cache: Largest, slowest among caches. Shared among all CPU cores in a multicore processor. Reduces main memory traffic between cores.

Associative Memory (vs Cache)

  • Associative Memory (Content-Addressable Memory - CAM): You search by content (data), not by address. All locations are searched in parallel. Used for fast lookups (e.g., TLB in MMU, network router tables).

  • Cache Memory: You search by address (tag comparison). It's a speed bridge between CPU and main memory, exploiting locality.

  • Key Difference: Access method. CAM searches by data value; cache searches by memory address (tag). CAM is much more hardware-intensive.

Effective Access Time (EAT) Calculation

$$ \text{EAT} = (\text{Hit Ratio} \times \text{Cache Access Time}) + (\text{Miss Ratio} \times \text{Main Memory Access Time}) $$

Where, $$\displaystyle \text{Miss Ratio} = 1 - \text{Hit Ratio} $$.

If write-back policy is used, write misses may require additional memory access.

Numerical Example (Dec 2024):

Given: Hit Ratio = 95% = 0.95, Cache Time = 10 ns, Main Memory Time = 100 ns.

$$ \text{EAT} = (0.95 \times 10) + (0.05 \times 100) = 9.5 + 5 = \boxed{14.5 \text{ ns}} $$

Internal Organization of RAM and ROM Chips

  • RAM (Random Access Memory - Volatile): Made of memory cells arranged in a 2D array (rows/columns). A decoder selects a row (word line), and sense amplifiers read/write the column (bit line) data. DRAM uses capacitors (needs refresh), SRAM uses flip-flops (faster, more transistors).

  • ROM (Read-Only Memory - Non-Volatile): Data is permanently stored during manufacturing (mask ROM) or programmable (PROM, EPROM, EEPROM, Flash). Internal structure similar to RAM but without write circuitry (or with limited write).

Optical Disks (Secondary Storage)

  • Principle: Data is stored as pits (bumps) on a reflective surface. A laser beam reads the pattern by detecting reflections.

  • Types:

    • CD-ROM: ~700 MB, single layer, single-sided.

    • DVD: 4.7-17 GB (single/double layer, single/double-sided). Uses shorter wavelength laser.

    • Blu-ray: Up to 50 GB (dual-layer). Uses blue-violet laser (shorter wavelength โ†’ higher density).

  • Advantages: Low cost per MB, removable, durable (no head crash).

  • Disadvantages: Slow random access (sequential read optimized), write-once or limited rewrite.


VI. Input/Output Systems

Interrupts

A mechanism for an I/O device or exception to temporarily suspend the current CPU program and execute a special routine (Interrupt Service Routine - ISR).

Types:

  1. Hardware vs Software: Generated by external device (Hardware) or by an instruction (Software - INT).

  2. Maskable (INTR) vs Non-Maskable (NMI): Can be ignored by CPU (maskable) or must be serviced immediately (NMI).

  3. Vectored vs Non-Vectored: Vectored interrupt provides the address of its ISR (interrupt vector) automatically. Non-vectored requires CPU to poll devices to find the source.

Interrupt Handling with Priorities:

When multiple interrupts occur, a priority scheme determines service order.

  • Hardware Priority (Daisy Chain): Devices connected in series. INTA signal cascades until the highest-priority requesting device acknowledges.

  • Software Polling: CPU checks each device's status register in a fixed order.

  • Parallel Priority (Interrupt Controller - e.g., 8259A, APIC): Each device has a dedicated request line to a controller. Controller encodes the highest-priority request and sends it to CPU.

Interrupt Servicing in 8086 Microprocessor:

  1. Interrupt Request (IRQ): Device asserts INTR (maskable) or NMI (non-maskable).

  2. Interrupt Acknowledge: CPU completes current instruction, sends INTA pulse.

  3. Vector Fetch: For vectored interrupts (like NMI or INT n), the interrupting device places an interrupt type number (0-255) on the data bus. For INTR, the 8086 performs two INTA cycles; the external interrupt controller (8259A) provides the type number.

  4. Lookup Interrupt Vector Table (IVT): IVT is in memory at 0000:0000. Each entry is 4 bytes (CS:IP). CPU calculates address = type_number ร— 4 and fetches the ISR's starting address (CS:IP).

  5. Branch to ISR: CPU pushes FLAGS and CS:IP of interrupted program, then jumps to ISR.

  6. ISR Execution: ISR saves registers, services device, restores registers.

  7. Return from Interrupt (IRET): Pops saved CS:IP and FLAGS, resuming interrupted program.

Programmed I/O

CPU executes a dedicated program to control I/O transfer. CPU polls the device's status register in a loop until it's ready, then reads/writes data.

  • Advantage: Simple.

  • Disadvantage: CPU is tied up waiting (wastes cycles). Inefficient for slow devices.

Interrupt-Driven I/O

CPU initiates I/O operation, then continues executing other programs. The I/O device interrupts the CPU when it's ready to transfer data (or on error).

  • Advantage: CPU is not busy-waiting; higher utilization.

  • Disadvantage: Still requires CPU intervention for each byte/word transferred (overhead).

Direct Memory Access (DMA)

A method where an I/O device transfers data directly to/from main memory without CPU intervention, after initial setup.

Requirements:

  1. DMA Controller (DMAC): A dedicated hardware chip with registers for memory address, transfer count, and control.

  2. Bus Arbitration: Mechanism for DMAC to take control of the system bus from the CPU (e.g., HOLD/HLDA signals in 8086).

Working (Burst Mode):

  1. CPU programs DMAC with: source address, destination address, transfer count.

  2. CPU commands DMAC to start.

  3. DMAC requests bus control (HOLD).

  4. CPU finishes current bus cycle, releases bus (HLDA).

  5. DMAC takes over the bus and performs the entire block transfer (read from device/write to memory or vice versa) autonomously.

  6. After transfer, DMAC releases bus back to CPU and may raise an interrupt to signal completion.

Comparison:

Feature Programmed I/O Interrupt-Driven I/O DMA
CPU Involvement For entire transfer For each word/byte Only at start & end
Speed Slowest Medium Fastest
CPU Utilization Very Low Medium High

Data Transfer Methods

Serial vs Parallel Transfer:

Feature Parallel Transfer Serial Transfer
Lines Used Multiple lines (one per bit). Single line (bit-by-bit).
Speed Faster (multiple bits/cycle). Slower (one bit/cycle).
Distance Short (on a PCB, between chips). Long (between systems, cables).
Cost/Complexity Higher (more wires, connectors). Lower (simpler, cheaper).
Why Serial Used? N/A Cost-effective for long distances, less crosstalk, easier synchronization, supports higher clock rates over distance (e.g., USB, SATA, Ethernet).

Strobe Method (Handshaking):

A simple asynchronous serial data transfer method using two control lines: Data Valid (source strobe) and Data Accepted (destination strobe).

  1. Source places data on line, asserts Data Valid.

  2. Destination reads data when it sees Data Valid, then asserts Data Accepted.

  3. Source sees Data Accepted, removes data and de-asserts Data Valid.

  4. Destination sees Data Valid go low, de-asserts Data Accepted.

Ensures source and destination are synchronized without a shared clock.

I/O Channels

A specialized, processor-like I/O subsystem with its own instruction set. It can perform I/O operations and simple data processing (e.g., formatting, error checking) independently of the CPU.

  • CPU loads the channel program (list of I/O commands) into channel memory and starts the channel.

  • Channel fetches and executes its own instructions, managing the I/O device and data transfer to/from main memory.

  • Types: Selector Channel (dedicated to one high-speed device), Multiplexor Channel (time-shares among multiple slow devices).


VII. Advanced Processor Architectures

RISC vs CISC Architectures

Feature CISC (Complex Instruction Set Computer) RISC (Reduced Instruction Set Computer)
Goal Do more with each instruction (complex ops). Simplify instructions for speed.
Instruction Size Variable. Fixed (usually 4 bytes).
Instructions Many, complex, multi-cycle. Few, simple, single-cycle (mostly).
Addressing Modes Many (8-20+). Few (3-5).
Registers Few (e.g., 8-16 general purpose). Many (e.g., 16-32 general purpose).
Control Unit Often microprogrammed. Typically hardwired.
Pipeline Difficult due to variable cycles. Easy, efficient pipelining.
Examples x86, VAX, System/360. ARM, MIPS, RISC-V, SPARC.

Why RISC Preferred in Modern Computing?

  1. Efficient Pipelining: Fixed-length, simple instructions enable deep pipelines and high clock speeds.

  2. Simpler Hardware: Less complex control logic โ†’ smaller die size, lower power, more transistors for cache/registers.

  3. Compiler-Driven Optimization: Complexity moved to compiler, allowing simpler, faster hardware.

  4. Load-Store Architecture: Only load/store instructions access memory. All other ops are register-register. Simplifies design and improves speed.

Pipelining

Basic Concept: Overlap the execution of multiple instructions by dividing the CPU into stages. Each stage works on a different instruction simultaneously.


DiagramCANVAS: A Space-Time Diagram for a 4-Stage Pipeline (Fetch, Decode, Execute, Write-back). Time flows down. Show 4 instructions (I1, I2, I3, I4) progressing across stages. In cycle 1: I1 in Fetch. Cycle 2: I1 in Decode, I2 in Fetch. Cycle 3: I1 in Execute, I2 in Decode, I3 in Fetch. Cycle 4: I1 in Write-back, I2 in Execute, I3 in Decode, I4 in Fetch. Cycle 5+: I2 Write-back, I3 Execute, I4 Decode, I5 Fetch...

Speedup: Ideal speedup โ‰ˆ number of stages (S). After pipeline fill, one instruction completes per cycle. Hazards (Obstacles):

  • Structural: Resource conflict (two instructions need same hardware).

  • Data: Data dependency (e.g., ADD R1, R2 followed by SUB R3, R1).

  • Control: Branch instructions change PC, causing instructions after branch to be wrong (stall until branch resolved).

Instruction-Level Parallelism (ILP) vs Thread-Level Parallelism (TLP)

Feature ILP TLP
Parallelism Granularity Within a single thread/instruction stream. Across multiple threads/processes.
Goal Execute multiple instructions from the same program simultaneously. Execute multiple threads (from same or different programs) simultaneously.
Hardware Support Pipelining, Superscalar (multiple ALUs), Out-of-Order Execution, Speculation. Multicore Processors, Simultaneous Multithreading (SMT - e.g., Intel Hyper-Threading).
Example A superscalar CPU issuing 4 instructions per cycle from a single thread. A dual-core CPU running two independent programs, or SMT core running two threads sharing resources.
Challenges Data dependencies, control dependencies (branches). Resource contention between threads, cache thrashing.

Multicore Processor Architecture

A single physical processor chip containing multiple independent CPU cores (each with its own ALU, CU, L1 cache) that can execute multiple instructions in parallel.

  • Shared Resources: Typically share L2/L3 cache, memory controller, system bus.

  • Advantage: True parallelism at thread level (TLP). Improves throughput for multi-threaded applications.

  • Challenge: Amdahl's Law limits speedup if program has sequential parts. Requires parallel programming (threads, processes, synchronization).

Vector Processing

Processes single instruction on multiple data elements simultaneously (SIMD - Single Instruction, Multiple Data).

  • Vector Processor: Has vector registers (hold arrays of data, e.g., 64x 64-bit elements) and vector functional units (e.g., vector add, multiply).

  • Operation: One vector instruction (VADD V1, V2, V3) performs element-wise addition on entire arrays.

  • Applications: Scientific computing, multimedia (image/audio processing), machine learning.

  • Modern Implementations: SIMD extensions in general CPUs (e.g., Intel SSE/AVX, ARM NEON).

ARM Processor Architecture

A family of RISC architectures, dominant in mobile/embedded systems.

  • Key Features:

    • Load-Store Architecture: Only load/store instructions access memory.

    • Fixed 32-bit (ARM/A32) or 16-bit (Thumb/T16) instruction length (Thumb improves code density).

    • Large register file (16-32 registers, R13=SP, R14=LR, R15=PC).

    • Conditional Execution: Most instructions can be predicated (executed conditionally), reducing branches.

    • 3-Operand Format: ADD Rd, Rn, Rm (Rd = Rn + Rm).

    • Big-Endian/Little-Endian support (configurable).

    • Advanced Power Management: Designed for low power.

    • Modern: ARMv8-A introduces 64-bit (AArch64) state and big.LITTLE heterogeneous multicore (high-performance + high-efficiency cores).

Exam Tip: Be ready to draw the ARM programmer's model (register set) and contrast its simplicity with x86 (CISC).

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in