Skip to content
CY-504 (C) · Computer Organization & Architecture/Quick Revision Short Notes

Computer Organization & Architecture (CY-504 (C)) - Unit 4 Short Notes

UNIT 4: Computer Organization & Architecture - Short Notes

Based on RGPV past papers (Dec 2024 & Nov 2023), focusing on high-frequency exam topics.


I. Basic Computer Structure & Organization

Basic Functional Units

A computer consists of three primary functional units:

  1. CPU (Central Processing Unit): The "brain." Fetches, decodes, and executes instructions. Contains ALU, control unit, and registers.

  2. Memory Unit: Stores data and instructions. Hierarchical (registers, cache, main memory, secondary storage).

  3. I/O Unit: Communicates with the external world (keyboard, display, disk). Handles data transfer between I/O devices and memory/CPU.

Exam Tip: Be prepared to draw the basic block diagram showing CPU, Memory, and I/O subsystems connected via a system bus.

System Bus Structure

A bus is a group of electrical lines for communication. Three main types:

Bus Type Purpose Direction
Data Bus Carries actual data (instructions, operands, results). Bi-directional
Address Bus Carries memory/I/O addresses from CPU. Uni-directional (CPU → Memory/I/O)
Control Bus Carries control signals (Read, Write, Interrupt, Clock). Bi-directional
  • Width of data bus determines amount of data transferred per cycle.

  • Width of address bus determines maximum addressable memory space ($$\displaystyle 2^{\text{width}} $$ locations).

General Register Organization

Registers are small, fast storage locations inside the CPU.

  • Common Registers:

    • Program Counter (PC): Holds address of next instruction to fetch.

    • Instruction Register (IR): Holds the current instruction being decoded/executed.

    • Memory Address Register (MAR): Holds address of memory location to be accessed.

    • Memory Data Register (MDR): Holds data read from or to be written to memory.

    • Accumulator (AC): Primary register for ALU operations.

    • General-Purpose Registers (R0, R1...): Used for operands and intermediate results.


II. Central Processing Unit (CPU) & Execution

Instruction Cycle (Fetch-Decode-Execute)

The fundamental cycle of CPU operation.


flowchart TD

    A[Fetch] --> B[Decode]

    B --> C[Execute]

    C --> D{Interrupt?}

    D -- No --> A

    D -- Yes --> E[Service Interrupt] --> A

  1. Fetch: PC → MAR → Memory → MDR → IR. PC incremented.

  2. Decode: Control unit interprets opcode in IR.

  3. Execute: ALU performs operation, or memory/I/O access occurs.

  4. Interrupt Check: If an interrupt is pending, the cycle is interrupted to service it.

Stack Operations

A stack is a LIFO (Last-In, First-Out) memory structure, typically implemented in main memory with a Stack Pointer (SP) register pointing to the top.

  • PUSH (Push onto Stack):

    1. Decrement SP.

    2. Store operand at memory location pointed by SP.

  • POP (Pop from Stack):

    1. Read operand from memory location pointed by SP.

    2. Increment SP.

Example: PUSH AX (Push contents of register AX onto stack). POP BX (Pop top of stack into register BX).

Floating-Point Arithmetic

Handles numbers in scientific notation: $$\displaystyle (-1)^s \times M \times 2^E $$.

  • Addition/Subtraction Flowchart:

    1. Align Exponents: Shift mantissa of smaller exponent number until exponents match.

    2. Add/Subtract Mantissas: Perform operation on aligned mantissas.

    3. Normalize Result: Shift result mantissa left/right to ensure it's in standard form (1.M), adjusting exponent accordingly.

    4. Round: Apply rounding rule to fit mantissa into available bits.

    5. Check for Overflow/Underflow.

Instruction Set Architecture (ISA)

The interface between hardware and software. Defines:

  • Instruction formats and opcodes.

  • Data types and sizes.

  • Registers and their usage.

  • Memory addressing modes.

  • How I/O and interrupts are handled.

Key Point: ISA is abstract. Multiple microarchitectures (e.g., Intel vs AMD) can implement the same ISA (x86).

Instruction Formats

Common fields: Opcode (operation), Address/Register specifiers (operand locations).

  • Zero-Address (Stack): e.g., ADD (operands implicit from stack).

  • One-Address (Accumulator): e.g., ADD M (AC ← AC + M).

  • Two-Address: e.g., ADD R1, R2 (R1 ← R1 + R2).

  • Three-Address: e.g., ADD R1, R2, R3 (R1 ← R2 + R3). More powerful, longer instructions.


III. Instruction Execution & Addressing

Addressing Modes

Specifies how to calculate the effective address (EA) of an operand.

Mode How EA is Calculated Example (ADD) Use Case
Immediate Operand is in instruction itself. ADD #5 Load constant.
Direct EA = address field in instruction. ADD 1000 Fast, fixed address.
Indirect EA = contents of memory location whose address is in instruction. ADD @1000 Pointer manipulation.
Register Operand is in specified register. ADD R1 Fastest.
Register Indirect EA = contents of specified register. ADD (R1) Array/string traversal.
Displacement EA = register + constant offset. ADD 1000(R1) Array element access.
Relative EA = PC + constant offset. ADD +1000 Position-independent code (PC-relative).
Indexed EA = index register + constant. ADD X(Ri) Array access.
Base-Register EA = base register + displacement. Similar to displacement. Memory protection, relocation.

Exam Tip: Be able to identify the mode from an example and calculate the EA given register/memory contents.


IV. Control Unit Design

Hardwired Control Unit

  • Principle: Control signals are generated by a fixed combinational logic circuit (decoders, gates). The instruction opcode directly drives the control lines.

  • Diagram: Shows instruction register opcode bits feeding into a control signal generator (logic gates) which outputs all control signals (e.g., ALUSrc, RegWrite, MemRead).

  • Advantages: Very fast (no microinstruction fetch), efficient for simple ISAs.

  • Disadvantages: Complex to design and modify (adding instructions requires rewiring), inflexible.

Micro-programmed Control Unit

  • Principle: Control signals are stored as microinstructions in a control memory (CM). A micro-program sequencer fetches and executes these microinstructions.

  • Diagram: Shows a Micro-program Counter (μPC) that addresses a Control Memory. The output (a Control Word) drives the control lines. The μPC is updated based on the current microinstruction and instruction opcode.

  • Advantages: Easy to design, modify, and debug. Flexible (new instructions via new microcode).

  • Disadvantages: Slower (extra memory fetch cycle), requires control memory.

Differentiation: Hardwired vs. Micro-programmed

Feature Hardwired Control Micro-programmed Control
Implementation Fixed logic (gates, decoders) Stored microinstructions in control memory
Speed Faster (direct generation) Slower (microinstruction fetch)
Flexibility Inflexible (hardwired) Flexible (microcode can be changed)
Complexity Complex design for complex ISAs Simpler design, easier to modify
Cost Lower (no control memory) Higher (requires control memory)
Usage Simple, high-speed processors (e.g., RISC) Complex CISC processors (e.g., x86)

Micro-programming Concepts

  • Control Word (CW): A binary word where each bit (or group) controls a specific signal in the computer (e.g., 1 means RegWrite=ON). One microinstruction = one control word.

  • Micro-program Sequencer: Generates the address for the next microinstruction.

    • Functions: Increment μPC, jump to a new routine based on opcode or condition, handle subroutine calls.

    • Types: Simple counter-based, branching-capable (using multiplexers).


V. Memory Hierarchy & Organization

Memory Hierarchy Concept

Organizes memory into levels (Registers → L1/L2/L3 Cache → Main Memory → Disk) based on speed, size, and cost per bit.

  • Goal: Achieve average access time close to fastest level (cache) with the capacity of the slowest level (disk).

  • Principle of Locality:

    • Temporal Locality: Recently accessed items likely to be accessed again soon.

    • Spatial Locality: Access to an item likely leads to access to nearby addresses.

Cache Memory (Crucial for Performance)

A small, fast SRAM memory between CPU and main memory.

  • Role: Holds copies of frequently used main memory blocks. Reduces effective access time dramatically.

  • Key Parameters:

    • Hit: Data found in cache.

    • Miss: Data not in cache, must fetch from main memory.

    • Hit Ratio (h): Fraction of accesses that are hits.

    • Miss Ratio (m): $$\displaystyle m = 1 - h $$.

    • Access Times: $$\displaystyle t_c $$ (cache), $$\displaystyle t_m $$ (main memory).

Effective Access Time (EAT) Formula:

$$\boxed{\text{EAT} = h \times t_c + (1 - h) \times (t_c + t_m)}$$

For a simple write-through cache. If $$\displaystyle t_c \ll t_m $$, a high hit ratio yields near-cache speed.

Cache Mapping Techniques

Determines where a main memory block can be placed in cache.

Technique How it Works Pros Cons
Direct Mapped Each memory block maps to exactly one cache line (index = block number mod #lines). Simple, fast hardware. High conflict misses (multiple blocks map to same line).
Fully Associative A block can be placed in any cache line. Lowest conflict misses. Complex, slow (must search all lines in parallel).
Set-Associative Compromise. Cache divided into sets (each with k lines). Block maps to a specific set, can go into any line within it. k-way set-associative. Balance of cost & performance. More complex than direct, less flexible than fully associative.

Common Pitfall: Confusing associative memory (content-addressable, searches by data) with set-associative cache (a cache organization technique).

L1, L2, L3 Cache

  • L1 Cache: Smallest, fastest, integrated on CPU die. Split into L1i (instruction) and L1d (data).

  • L2 Cache: Larger, slightly slower than L1. May be on-die or on-package.

  • L3 Cache: Largest, slowest of on-chip caches. Shared among multiple CPU cores. Acts as a "victim cache" for L2.


VI. Input/Output (I/O) Organization & Data Transfer

Data Transfer Methods Comparison

Method CPU Involvement Data Path Speed Use Case
Programmed I/O (PIO) High. CPU executes I/O instructions for each byte/word. CPU ↔ I/O Slowest Simple devices, early systems.
Interrupt-Driven I/O Medium. CPU starts transfer, device interrupts when ready/next byte. CPU ↔ I/O Slow-Medium Reduces CPU polling, but CPU still moves data.
Direct Memory Access (DMA) Low. CPU programs DMA controller, then continues. DMA controller moves data directly between I/O and memory. I/O ↔ Memory Fastest High-speed devices (disk, network, graphics).

Direct Memory Access (DMA)

  • Working:

    1. CPU initializes DMA controller (source, destination address, count).

    2. CPU continues other tasks.

    3. DMA controller takes control of the system bus (bus mastering).

    4. DMA controller transfers data directly between I/O device and memory.

    5. DMA controller interrupts CPU upon completion.

  • Requirements: DMA controller, ability to temporarily suspend CPU (bus arbitration).

  • Key Difference from PIO/Interrupt: CPU does not move the data bytes. It only sets up and is notified.

Serial vs. Parallel Data Transfer

Feature Parallel Transfer Serial Transfer
Lines Used Multiple lines (e.g., 8, 16, 32 bits at once). Single line (bit-by-bit).
Speed (Distance) Fast for short distances (on a circuit board). Slower per bit, but scales better over distance.
Cost/Complexity High (more wires, connectors, crosstalk). Low (1 wire, less interference).
Why Use Serial? Long-distance communication (USB, SATA, PCIe), reduces cost/crosstalk, high clock speeds compensate (e.g., PCIe 5.0).

Strobe Method of Data Transfer

A simple asynchronous handshaking method for parallel transfer.

  1. Source places data on data lines.

  2. Source activates strobe signal (e.g., DataValid pulse) to indicate data is ready.

  3. Destination reads data when it sees strobe.

  4. Destination may send an acknowledgment pulse back.

Diagram: Shows two devices with data lines and a single strobe line from source to destination.


VII. Interrupts

Definition & Concept

An interrupt is a signal that suspends the CPU's current instruction stream to service an event (I/O completion, timer, error), then resumes.

  • Interrupt-Driven I/O: Device interrupts CPU when it's ready to send/receive data, avoiding wasteful polling.

Types of Interrupts

Type Description Example
Hardware vs. Software Hardware: External pin signal. Software: INT instruction. Hardware: Keyboard press. Software: System call.
Maskable vs. Non-Maskable (NMI) Maskable: Can be ignored/disabled by CPU. NMI: Cannot be ignored, highest priority. Maskable: Disk I/O. NMI: Power failure, hardware error.
Vectored vs. Non-Vectored Vectored: Interrupt device provides interrupt service routine (ISR) address. Non-Vectored: CPU must poll devices to find source. x86 uses vectored interrupts (interrupt number → IVT entry).

Interrupt Handling with Priorities

When multiple interrupts occur:

  1. CPU prioritizes them (hardwired or programmable).

  2. CPU saves state (PC, PSW) of current program.

  3. CPU disables lower-priority interrupts (optional).

  4. CPU jumps to ISR of highest-priority pending interrupt.

  5. ISR executes, acknowledges the device.

  6. ISR restores state and returns (via IRET).

Interrupt Handling in 8086 Microprocessor

  • Interrupt Vector Table (IVT): Located at physical address 0x00000 in memory. Contains 256 entries (4 bytes each: CS:IP of ISR).

  • Interrupt Types:

    • Hardware Interrupts: INTR (maskable, vectored via interrupt number on data bus), NMI (non-maskable, fixed vector 2).

    • Software Interrupts: INT n (software-generated, vectored).

    • Exceptions: Internal errors (Divide Error, Single Step).

  • Sequence:

    1. CPU completes current instruction.

    2. For INTR, CPU sends INTA pulse. Interrupt controller (8259A) places interrupt type number on data bus.

    3. CPU uses type number n to fetch ISR address from IVT entry n × 4.

    4. CPU pushes FLAGS, CS, IP, and jumps to ISR.

    5. ISR must save/restore all registers it uses and end with IRET.


VIII. Advanced Processor Architectures & Parallelism

RISC vs. CISC

Feature CISC (Complex ISA) RISC (Reduced ISA)
Goal Do more per instruction (complex ops). Simpler, faster instructions.
Instruction Size Variable (1-15 bytes). Fixed (typically 4 bytes).
Addressing Modes Many (complex, memory-to-memory). Few (register-register, load/store).
Microcode Typically used (micro-programmed CU). Not used (hardwired CU).
Registers Few (8-16). Many (16-32+).
Pipeline Difficult (variable cycles/instruction). Easy (regular, single-cycle ops).
Examples x86, VAX ARM, MIPS, RISC-V
Why RISC Preferred? Enables deep pipelining, higher clock speeds, simpler compiler design. Dominates mobile, embedded, and high-performance computing (via ARM).

Instruction-Level Parallelism (ILP) vs. Thread-Level Parallelism (TLP)

Aspect ILP TLP
Granularity Single instruction stream. Exploits parallelism within a single thread/process. Multiple threads/processes. Exploits parallelism across threads.
Goal Execute multiple instructions from same thread simultaneously. Execute multiple threads simultaneously.
Hardware Pipelining, Superscalar (multiple ALUs), Out-of-Order Execution (OoO), Speculation. Multicore, Multithreading (SMT/Hyper-Threading).
Example A superscalar processor issues 4 instructions per cycle from one thread. A 4-core CPU with SMT runs 8 threads concurrently (2 per core).
Key Challenge Dependencies (data, control) limit parallelism. Resource contention (cache, memory bandwidth) between threads.

Pipelining

  • Concept: Divide instruction execution into k stages (Fetch, Decode, Execute, Memory, Writeback). Multiple instructions overlap in different stages.

  • 4-Segment Pipeline Space-Time Diagram:

    
    Time | Stage 1 | Stage 2 | Stage 3 | Stage 4
    
    ---- | ------- | ------- | ------- | -------
    
    t1   | I1      |         |         |
    
    t2   | I2      | I1      |         |
    
    t3   | I3      | I2      | I1      |
    
    t4   | I4      | I3      | I2      | I1
    
    t5   | I5      | I4      | I3      | I2
    
    
    • Speedup (Ideal): ≈ k (for large number of instructions).

    • Hazards Limit Speedup:

      • Structural: Resource conflict (e.g., two instructions need memory at same stage).

      • Data: Dependency (e.g., ADD R1, R2 followed by SUB R3, R1).

      • Control: Branch/jump uncertainty (pipeline fetches wrong instructions).

Multicore Processor Architecture

  • Basic Structure: A single chip containing multiple independent CPU cores (each with its own ALU, registers, L1 cache).

  • Shared Resources: Typically share L2/L3 cache, memory controller, and I/O.

  • Challenge: Cache coherence (ensuring all cores see a consistent view of memory). Solved by coherence protocols (e.g., MESI).

  • Example: Intel Core i7 (4-8 cores), ARM big.LITTLE (performance + efficiency cores).

Vector Processing

  • Concept: A single instruction operates on multiple data elements simultaneously (SIMD - Single Instruction, Multiple Data).

  • Example: VADD V1, V2, V3 adds 8 pairs of 32-bit integers from vectors V2 and V3, stores result in V1.

  • Use Cases: Multimedia, scientific computing, AI/ML (matrix operations).

  • Modern Implementations: SIMD extensions (x86: SSE, AVX; ARM: NEON).


IX. Specific Processor & I/O Device Examples

Architecture of ARM Processor (Basic Overview)

  • RISC Philosophy: Load/Store architecture, fixed 32-bit instructions (ARM state), large register file (16 x 32-bit).

  • Key Features:

    • Registers: R0-R12 (general), R13 (SP), R14 (LR), R15 (PC). Current Program Status Register (CPSR) holds flags/status bits.

    • Pipelining: Classic 3-stage (Fetch, Decode, Execute) or 5-stage.

    • Modes: User, FIQ, IRQ, Supervisor, etc. (for OS/exception handling).

    • Load/Store: Only LDR/STR access memory. ALU ops are register-register.

    • Conditional Execution: Most instructions can be conditionally executed (e.g., ADDEQ).

    • Thumb/Thumb-2: 16-bit compressed instruction set for code density.

  • Modern ARM: Supports multi-core (big.LITTLE), SIMD (NEON), virtualization.

8086 Microprocessor Interrupts (Recap from Section VII)

  • IVT at 0x0000:0000, 256 entries (4 bytes each).

  • INTR: Maskable, external. Interrupt number provided by 8259A PIC.

  • NMI: Non-maskable, pin on CPU. Fixed vector 2.

  • INT n: Software interrupt.

  • Exceptions: Internal (Divide Error, INT3, INTO, etc.).

  • Response: CPU pushes FLAGS, clears TF/IF, pushes CS:IP, jumps to ISR from IVT. ISR ends with IRET.

Optical Disks (CD/DVD/Blu-ray)

  • Principle: Store data as pits and lands on a spiral track. Read by a laser beam reflected from the surface.

  • Types:

    • CD-ROM: ~700 MB, 780 nm laser.

    • DVD: 4.7-17 GB, 650 nm laser (smaller pits).

    • Blu-ray: 25-50 GB, 405 nm blue-violet laser (even smaller pits).

  • Key Feature: Rotational speed varies (Constant Linear Velocity - CLV) to maintain constant data density.

  • Use: Primarily for read-only media (software, movies). Rewritable versions exist (CD-RW, DVD±RW).


Key Formulas & Concepts for Calculation

  1. Effective Access Time (Cache):

$$\text{EAT} = h \cdot t_c + (1 - h) \cdot (t_c + t_m)$$

Where $h$ = hit ratio, $$\displaystyle t_c $$ = cache access time, $$\displaystyle t_m $$ = main memory access time.
  1. Speedup from Pipelining (Ideal):

$$S \approx k \quad \text{(for large N, where k = number of stages, N = instructions)}$$

Actual speedup reduced by hazards.
  1. Memory Address Space:

$$\text{Max Addressable Memory} = 2^{\text{Address Bus Width}} \times \text{Addressable Unit (usually 1 byte)}$$

Final Exam Strategy: For 7-mark questions, always start with a clear definition, followed by a diagram/flowchart/table where applicable, then a concise explanation with examples. For calculation questions (like EAT), write the formula first, substitute values, and box the final answer.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in