UNIT 5: VLSI CIRCUITS AND SYSTEMS
1.0 MOS TRANSISTOR FUNDAMENTALS & SCALING
1.1 Electrical Properties of MOS Transistor
-
Threshold Voltage ($$\displaystyle V_{th} $$): Minimum gate-to-source voltage to form conductive channel.
-
Transconductance ($$\displaystyle g_m $$): Measure of gate voltage control over drain current.
$$g_m = \frac{\partial I_D}{\partial V_{GS}}$$
-
Output Conductance ($$\displaystyle g_{ds} $$): Due to channel length modulation, represents drain current sensitivity to $$\displaystyle V_{DS} $$.
-
Mobility ($\mu$): Carrier velocity per unit electric field. Affected by vertical and lateral fields.
-
Subthreshold Conduction: Weak inversion current when $$\displaystyle V_{GS} < V_{th} $$. Exponential dependence on $$\displaystyle V_{GS} $$.
[!TIP] Exam Focus: Be prepared to derive/explain $$\displaystyle I_D $$ in saturation and linear regions, including $$\displaystyle g_m $$ and $$\displaystyle g_{ds} $$ expressions.
1.2 Need for Scaling & Scaling Principles
-
Need: Higher density, speed, lower cost/power.
-
Constant Field Scaling (Dennard Scaling): Scale all dimensions ($$\displaystyle L, W, t_{ox} $$) and voltages ($$\displaystyle V_{DD} $$) by factor $$\displaystyle S > 1 $$. Electric fields remain constant.
-
Pros: $$\displaystyle I_D $$ constant, power density constant, delay scales as $1/S$.
-
Cons: $$\displaystyle V_{th} $$ doesn't scale ideally, subthreshold leakage increases.
-
-
Constant Voltage Scaling: Scale only dimensions. $$\displaystyle V_{DD} $$ fixed.
-
Pros: Compatible with existing systems.
-
Cons: Electric field increases → reliability issues, power density increases.
-
1.3 Fundamental Units of CMOS Inverter
-
Pull-down Network (PDN): NMOS network, connects output to GND for logic '1' at input.
-
Pull-up Network (PUN): PMOS network, connects output to $$\displaystyle V_{DD} $$ for logic '0' at input.
-
Ratioless Design: Output driver strength independent of input logic state (complementary networks).
-
Static Power: Ideally zero (except leakage).
2.0 CMOS LOGIC GATES & LAYOUT DESIGN
2.1 Layout Design Rules
-
λ-based Rules: Minimum feature sizes expressed in multiples of λ (half of minimum poly width).
-
Well and Implant Rules: n-well/p-well spacing, enclosure of active by well, implant overlap.
-
Key Rules: Minimum width/spacing for diffusion, poly, metal; contact/via sizes and enclosures.
2.2 Layout Diagrams for Standard Cells
-
NAND Gate (2-input):
-
PDN: Series NMOS.
-
PUN: Parallel PMOS.
-
Layout: Poly runs parallel; diffusion for series NMOS shared; metal1 output contacts over diffusion.
-
-
NOR Gate (2-input):
-
PDN: Parallel NMOS.
-
PUN: Series PMOS.
-
Layout: Symmetric to NAND but with PMOS in series.
-
[!TIP] Common Pitfall: Forgetting that PMOS are in n-well. All PMOS in a standard cell share a common n-well.
2.3 Combinational Circuit Design using CMOS
-
Methodology:
-
Derive PDN from logic expression (transistors in series for AND, parallel for OR).
-
Create dual PUN (swap series/parallel, NMOS→PMOS).
-
Verify with truth table: For each input combination, exactly one network (PDN or PUN) is ON.
-
-
Example (AOI21 = $\overline{(A \cdot B) + C}$):
-
PDN: Parallel of (A&B series) and C.
-
PUN: Series of (A parallel B) and C (dual).
-
Verification: Truth table shows no short-circuit path.
-
3.0 CIRCUIT TECHNIQUES & LOGIC FAMILIES
3.1 Transmission Gate (TG)
-
Structure: Parallel NMOS & PMOS, gates controlled by complementary signals (C and $\overline{C}$).
-
Operation: Passes both logic '0' (via NMOS) and '1' (via PMOS) without threshold loss.
-
Applications: Multiplexers, bus switches, low-leakage pass gates, dynamic logic precharge.
3.2 Pass Transistor Logic (PTL)
-
Basic Principle: Use NMOS/PMOS as switches. No complementary pull-up network.
-
Comparison with CMOS:
| Feature | CMOS | PTL | | :--- | :--- | :--- | | Area | Larger (complementary) | Smaller | | Speed | Slower (2 transistors in series) | Faster (single pass) | | Noise Margin | Good (full swing) | Poor (threshold loss) | | Static Power | Very low | Higher (leakage paths) |
-
Variants:
-
Complementary PTL (CPTL): Uses both NMOS and PMOS pass networks.
-
Differential Cascode Voltage Switch Pass Gate (DCVSPG): High-speed, differential.
-
[!TIP] Exam Alert: PTL is faster but suffers from threshold voltage loss ($$\displaystyle V_{th} $$ drop) when passing a '1' through NMOS. TG solves this.
4.0 TIMING ANALYSIS & DELAY MODELS
4.1 Elmore's Constant
- Definition: First-order RC-tree delay approximation. Sum of resistance × downstream capacitance for each capacitor.
$$\tau_{Elmore} = \sum_{i} R_{i} \cdot C_{i}$$
where $$\displaystyle R_i $$ is resistance from root to capacitor $$\displaystyle C_i $$, and $$\displaystyle C_i $$ is capacitance at node $i$.
- Interpretation: Time constant representing effective RC delay.
4.2 Elmore Delay Expression for CMOS Inverter
-
Simple RC Model:
-
Input capacitance $$\displaystyle C_{in} $$ (gate capacitance of driven gate).
-
Output load capacitance $$\displaystyle C_L $$ (gate capacitance of next stage + wire cap).
-
On-resistance of NMOS ($$\displaystyle R_{n} $$) and PMOS ($$\displaystyle R_{p} $$).
-
-
Derivation for Rising Output (PUN ON):
$$\tau_{pLH} = R_p \cdot (C_{int} + C_L) + R_n \cdot C_{int}$$
where $$\displaystyle C_{int} $$ is internal capacitance (diffusion, overlap).
- Simplified (dominant $$\displaystyle C_L $$):
$$\boxed{t_p \approx 0.69 \cdot R_{eq} \cdot C_L}$$
where $$\displaystyle R_{eq} = R_p || R_n $$ for symmetric inverter.
- Extension to Logical Effort: Normalizes delay by $$\displaystyle C_{in} $$ and inverter delay.
5.0 SEQUENTIAL CIRCUIT DESIGN
5.1 Methodology for Latches & Flip-Flops
-
Latch: Level-sensitive. Transparent when clock is active (HIGH/LOW).
-
Flip-Flop: Edge-triggered. Captures input only at clock edge (rising/falling).
-
Basic Latch (Gated D-Latch):
-
Two cross-coupled inverters for storage.
-
Two transmission gates controlled by $\text{CLK}$ and $\overline{\text{CLK}}$ for input gating.
-
When $$\displaystyle \text{CLK}=1 $$, input $D$ propagates to output $Q$ (transparent).
-
5.2 Master-Slave Based Edge-Triggered Register
-
Structure: Two latches in series.
-
Master Latch: Enabled by $\text{CLK}$ (e.g., positive level).
-
Slave Latch: Enabled by $\overline{\text{CLK}}$ (opposite level).
-
-
Operation & Timing:
-
Positive Edge-Triggered D FF:
-
$$\displaystyle \text{CLK}=0 $$: Master transparent, slave opaque. Master follows $D$.
-
$\text{CLK}$ rising edge: Master becomes opaque, slave becomes transparent. Slave captures master's output.
-
$$\displaystyle \text{CLK}=1 $$: Master opaque, slave transparent. $Q$ = master's old value.
-
-
Setup Time ($$\displaystyle t_{su} $$): $D$ must be stable before clock edge.
-
Hold Time ($$\displaystyle t_h $$): $D$ must be stable after clock edge.
-
-
Common Type: Master-Slave D Flip-Flop (most common in pipelines).
[!TIP] Key Point: Master-slave design uses two transparent periods but overall behaves as edge-triggered because slave captures only at the opposite clock phase.
6.0 CLOCK DISTRIBUTION & SYNCHRONOUS DESIGN
6.1 Clock Distribution Techniques
-
Objectives: Minimize clock skew (difference in clock arrival times), reduce jitter, low power.
-
Topologies:
| Topology | Description | Skew | Power | Complexity | | :--- | :--- | :--- | :--- | :--- | | H-tree | Symmetric binary tree | Low | Moderate | Moderate | | Grid | Mesh network over chip | Very Low | High | High | | Fishbone | Main trunk with branches | Moderate | Low | Simple | | Clock Mesh | Dense grid with buffers | Very Low | Very High | Very High |
-
Clock Buffering: Insert buffers to drive loads, balance rise/fall times.
-
Repeater Insertion: For long global wires to reduce RC delay.
-
Clock Gating: Insert enable-controlled AND/OR gates to stop clock to idle modules → saves dynamic power.
7.0 ARITHMETIC CIRCUITS (COMBINATIONAL)
7.1 Ripple Carry Adder (RCA)
-
Structure: Chain of full adders (FA). Carry ripples from LSB to MSB.
-
Limitations: Critical path delay = $$\displaystyle O(n) \cdot t_{FA} $$ (linear with bit-width). Slow for large $n$.
7.2 Carry Look-Ahead Adder (CLA)
- Concept: Generate carry signals in parallel using Generate ($$\displaystyle G_i $$) and Propagate ($$\displaystyle P_i $$) signals.
$$G_i = A_i \cdot B_i, \quad P_i = A_i \oplus B_i$$
$$C_{i+1} = G_i + P_i \cdot C_i$$
- Group Generate/Propagate (for 4-bit block):
$$G_{[3:0]} = G_3 + P_3 G_2 + P_3 P_2 G_1 + P_3 P_2 P_1 G_0$$
$$P_{[3:0]} = P_3 \cdot P_2 \cdot P_1 \cdot P_0$$
-
Block Diagram: Hierarchical (group of 4-bit CLAs → higher-level CLA).
-
Performance: Delay $O(\log n)$ (logarithmic). Area $$\displaystyle O(n^2) $$ (more hardware). Speed-area trade-off.
7.3 Carry Bypass Adder (CBA)
-
Design of 16-bit CBA:
-
Divide into 4-bit blocks.
-
Each block computes its own carry ($$\displaystyle C_{i+3} $$) assuming $$\displaystyle C_i=0 $$ and $$\displaystyle C_i=1 $$.
-
Bypass Logic: If all $$\displaystyle P_i=1 $$ in a block, carry in = carry out. Else, use ripple within block.
-
Global carry chain selects correct block carries.
-
-
Features: Faster than RCA for random inputs (often $$\displaystyle P_i=1 $$). Slower than CLA for worst-case. Area between RCA and CLA.
7.4 Multiplier Design
-
Array Multiplier: Systematic array of FAs. Regular, but slow (ripple-like carry).
-
Booth Multiplier (Radix-2, Double Precision):
-
Booth Encoding: Examine 3 bits (current + previous LSB). Encodes to 0, ±1, ±2.
| $$\displaystyle x_{2i+1} $$ | $$\displaystyle x_{2i} $$ | $$\displaystyle x_{2i-1} $$ | Operation | Partial Product | | :--- | :--- | :--- | :--- | :--- | | 0 | 0 | 0 | 0 | 0 | | 0 | 0 | 1 | +Y | Y | | 0 | 1 | 0 | +Y | Y | | 0 | 1 | 1 | +2Y | Y<<1 | | 1 | 0 | 0 | -2Y | -Y<<1 | | 1 | 0 | 1 | -Y | -Y | | 1 | 1 | 0 | -Y | -Y | | 1 | 1 | 1 | 0 | 0 |
-
Structure: Partial product generation (shifted Y or -Y) → addition tree (e.g., Wallace tree).
-
Sign Extension: For 2's complement, sign bit extended to left.
-
-
Worked Example (4-bit, A=1011 (-5), B=0101 (5)):
-
Pad with 0: A' = 01011, B' = 00101.
-
Booth recoding (3-bit groups, overlap by 1):
-
Group 0 (bits 0,1,2): 110 → -Y
-
Group 1 (bits 1,2,3): 101 → -Y
-
Group 2 (bits 2,3,4): 010 → +Y
-
-
Partial products (shifted, negated using 2's complement):
-
PP0 = -B = -0101 = 1011 (4-bit) → 11011 (5-bit)
-
PP1 = -B<<1 = -1010 = 0110 (4-bit) → 00110 (5-bit)
-
PP2 = +B<<2 = 010100 (6-bit)
-
-
Sum using carry-save/ripple.
-
8.0 PIPELINING & PERFORMANCE ENHANCEMENT
8.1 Concept of Pipelining
-
Break long combinational path into $k$ stages separated by pipeline registers (flip-flops).
-
Throughput: # operations per unit time → increases (ideally by $k$).
-
Latency: Time for one operation to complete → increases (by $$\displaystyle (k-1) \cdot T_{clk} $$).
-
Clock Frequency: $$\displaystyle f_{clk} \leq 1 / (\text{max stage delay} + t_{setup} + t_{skew}) $$.
8.2 Pipeline Design Considerations
-
Pipeline Overhead: Area/power of registers, clock load.
-
Placement: Registers at stage boundaries; balance stage delays.
-
Hazards:
-
Structural: Resource conflict (e.g., two stages need same memory). Mitigation: Stalling, duplication.
-
Data: RAW (Read-After-Write), WAR, WAW. Mitigation: Forwarding/bypassing, stall cycles.
-
Control: Branch decisions not ready. Mitigation: Branch prediction, delayed slots.
-
8.3 Impact
-
Clock Frequency: Increases (shorter critical path).
-
Design Complexity: Increases (hazard handling, verification, clock distribution).
9.0 DESIGN AUTOMATION & PHYSICAL DESIGN TOOLS
9.1 Stick Diagram
-
Purpose: Abstract, symbolic layout for early floorplanning and routing strategy.
-
Symbols:
-
Diffusion: colored rectangles (n-diff, p-diff).
-
Polysilicon: horizontal/vertical lines.
-
Metal1/Metal2: thicker lines.
-
Contact: '+' symbol.
-
-
Use: Quickly estimate area, routing congestion, transistor placement before detailed layout.
9.2 Standard Cell Libraries
-
Components (Cell Views):
-
Symbol: Logical representation.
-
Schematic: Circuit netlist.
-
Layout: Physical geometry (GDSII).
-
Timing: NLDM (Non-Linear Delay Model), CCS (Composite Current Source) – delay vs load/slew.
-
Power: Leakage, dynamic power models.
-
Functional: Verilog/VHDL model.
-
-
Characterization: Extract timing/power for different input slew, output load corners (PVT).
-
Role in ASIC Flow:
-
Synthesis: Maps RTL to library cells.
-
Place & Route: Places cells, connects with routing.
-
Timing/Power Analysis: Uses library models.
-
10.0 FIELD-PROGRAMMABLE GATE ARRAYS (FPGAs)
10.1 FPGA Building Block Architecture
-
Configurable Logic Block (CLB) / Logic Element (LE):
-
Typically: 1 or more LUTs (Look-Up Tables, e.g., 6-input), flip-flops, multiplexers.
-
Implements combinational/sequential logic.
-
-
Programmable Interconnects:
-
Switch Matrix: Connects CLB I/Os to routing wires.
-
Routing Resources: Horizontal/vertical wire segments of varying lengths (single, double, long lines).
-
-
I/O Blocks (IOBs): Programmable I/O standards, drive strength, slew rate.
-
Clock Management Tiles: PLLs/MMCMs for clock generation, deskew, frequency synthesis.
10.2 Programming Technologies
| Technology | Principle | Volatile? | Reconfigurability | Speed | Power | Cost | Example |
|---|---|---|---|---|---|---|---|
| SRAM-based | Configuration SRAM cells control pass transistors | Yes | Unlimited | Fast | Higher | Moderate | Xilinx, Intel |
| Antifuse-based | One-time programmable (fuse blown) | No | No | Very Fast | Low | Low | Actel (Microsemi) |
| Flash-based | Floating-gate transistors (EEPROM-like) | No | Moderate (sector erase) | Moderate | Low | Moderate | Microsemi (SmartFusion) |
[!TIP] Comparison Key: SRAM dominates market due to reconfigurability. Antifuse fastest/low power but one-time. Flash non-volatile, moderate speed.