UNIT 2: VLSI CIRCUITS AND SYSTEMS
Based on the May 2023 VLSI-specific exam paper analysis.
I. MOS TRANSISTOR FUNDAMENTALS & CMOS INVERTER
Electrical Properties of MOS Transistor
-
Threshold Voltage ($$\displaystyle V_{th} $$): Minimum gate-to-source voltage required to create a conductive channel.
-
Mobility ($$\displaystyle \mu_n, \mu_p $$): Carrier mobility in the channel; $$\displaystyle \mu_n > \mu_p $$ (typically 2-3x).
-
Oxide Capacitance ($$\displaystyle C_{ox} $$): Gate capacitance per unit area, $$\displaystyle C_{ox} = \frac{\varepsilon_{ox}}{t_{ox}} $$.
-
Channel Length Modulation (λ): Effect where effective channel length shortens with increased $$\displaystyle V_{DS} $$, modeled by a finite output resistance ($$\displaystyle r_o = \frac{1}{\lambda I_D} $$).
-
Subthreshold Conduction: Weak inversion current flow when $$\displaystyle V_{GS} < V_{th} $$, critical for low-power design.
Scaling in VLSI
-
Need: Increase density, improve speed, reduce power.
-
Scaling Principles & Effects:
| Scaling Type | Voltage ($$\displaystyle V_{DD} $$) | Dimensions (W, L, $$\displaystyle t_{ox} $$) | Doping | Key Effect | | :--- | :--- | :--- | :--- | :--- | | Constant-Field (Dennard) | Scaled by $1/S$ | Scaled by $1/S$ | Scaled by $S$ | Maintains $E$-field, power density constant. | | Constant-Voltage | Constant | Scaled by $1/S$ | Scaled by $S$ | $E$-field increases, hot-carrier effects worsen. | | General Scaling | Scaled by $$\displaystyle 1/S^\alpha $$ | Scaled by $1/S$ | Scaled by $S$ | Trade-off between speed, power, and reliability. |
CMOS Inverter
-
Fundamental Unit: Single-stage, low static power, high noise margin.
-
Static Characteristics (VTC):
-
Noise Margins: $$\displaystyle NM_L = V_{IL} - V_{OL} $$, $$\displaystyle NM_H = V_{OH} - V_{IH} $$.
-
Switching Threshold ($$\displaystyle V_M $$): $$\displaystyle V_{in} = V_{out} $$; ideally $$\displaystyle V_{DD}/2 $$ for symmetric $$\displaystyle \beta_n = \beta_p $$.
-
-
Dynamic Characteristics:
-
Propagation Delays: $$\displaystyle t_{pHL} $$ (HIGH→LOW), $$\displaystyle t_{pLH} $$ (LOW→HIGH).
-
Power Dissipation: $$\displaystyle P_{dynamic} = C_L V_{DD}^2 f $$, $$\displaystyle P_{static} \approx 0 $$ (ideal).
-
-
Elmore Delay for CMOS Inverter:
The RC delay model for a single inverter driving load capacitance $$\displaystyle C_L $$ is:
$$ t_p \approx 0.69 \cdot (R_{eqn} + R_{eqp}) \cdot C_L $$
where $$\displaystyle R_{eqn}, R_{eqp} $$ are equivalent resistances of NMOS/PMOS in their linear region.
[!TIP] Exam Focus: Be prepared to derive $$\displaystyle V_M $$ and explain the effect of $\beta$ ratio on VTC symmetry and noise margins.
II. LOGIC GATE DESIGN & PHYSICAL IMPLEMENTATION
Layout Design Rules (λ-based)
-
Purpose: Ensure manufacturability, avoid shorts/opens, guarantee yield.
-
Key Rules (Minimum Dimensions):
-
Width: Minimum transistor width ($$\displaystyle W_{min} $$), wire width.
-
Spacing: Minimum spacing between diffusion, poly, metal layers.
-
Overlap: Poly must overlap diffusion by at least 2λ to ensure gate formation.
-
Contact: Contact cut size and surrounding spacing.
-
Layout of NAND and NOR Gates
-
2-Input NAND:
-
Pull-Down Network (PDN): Two NMOS in series.
-
Pull-Up Network (PUN): Two PMOS in parallel.
-
Layout: Series NMOS share diffusion; parallel PMOS have separate diffusion connected to $$\displaystyle V_{DD} $$.
-
-
2-Input NOR:
-
PDN: Two NMOS in parallel.
-
PUN: Two PMOS in series.
-
Layout: Parallel NMOS have separate diffusion; series PMOS share diffusion.
-
DiagramCANVAS: Show stick diagrams and full layout for 2-input NAND and NOR, labeling diffusion (active), poly, metal1, and contact layers using λ-rules.
Stick Diagrams
-
Purpose: Abstract, quick planning of layout connectivity and area.
-
Representation:
-
Diffusion: Green (n-well for PMOS, p-sub for NMOS).
-
Poly: Red (crosses diffusion to form gates).
-
Metal1: Blue (connects pins, power rails).
-
Contact: Black dots (via between diffusion/poly and metal1).
-
-
Abstraction: One color/line = one layer; no width/spacing details.
Alternative Logic Styles
-
Pass Transistor Logic (PTL):
-
Concept: Use transistors as switches to pass signals directly to output, reducing transistor count.
-
vs Static CMOS: Faster, smaller area, but suffers from threshold voltage drop ($$\displaystyle V_{th} $$ loss) and charge sharing.
-
Example: Transmission gate XOR.
-
-
Transmission Gate:
-
Structure: Parallel NMOS and PMOS controlled by complementary signals.
-
Use: Bidirectional, low-resistance switch; no $$\displaystyle V_{th} $$ drop; used in multiplexers, registers, and dynamic circuits.
-
[!TIP] Common Pitfall: PTL outputs cannot drive large loads well due to high resistance and degraded swing. Always use a static CMOS inverter as a buffer after PTL blocks.
III. COMBINATIONAL CIRCUIT DESIGN
Design Methodology
-
Pull-Down Network (PDN): Implement logic function $F$ with NMOS (series=AND, parallel=OR).
-
Pull-Up Network (PUN): Implement complement $\overline{F}$ with PMOS (series=OR, parallel=AND) – Dual Network.
-
Verify: Ensure no static current path (PDN and PUN never ON simultaneously).
Arithmetic Circuits: Adders
-
Ripple Carry Adder (RCA): Simple, cascaded full adders. Delay $$\displaystyle \propto n \cdot t_{FA} $$.
-
Carry Look-Ahead Adder (CLA):
- Concept: Generate ($$\displaystyle G_i $$) and Propagate ($$\displaystyle P_i $$) signals.
$$ G_i = A_i \cdot B_i, \quad P_i = A_i \oplus B_i $$
* **Carry Equations:**
$$ C_{i+1} = G_i + P_i C_i $$
$$ C_1 = G_0 + P_0 C_{in} $$
$$ C_2 = G_1 + P_1 G_0 + P_1 P_0 C_{in} \quad \text{(expands)} $$
* **Block Diagram:** DiagramSEARCH: "CLA block diagram generate propagate"
* **Delay:** $O(\log n)$ for carry chain, faster than RCA.
-
Carry Bypass Adder (CBA):
-
Design (16-bit): Group bits into blocks (e.g., 4-bit). Generate block propagate ($$\displaystyle P_{block} $$).
-
Operation: If $$\displaystyle P_{block}=1 $$, carry bypasses internal RCA, $$\displaystyle C_{out} = C_{in} $$. Else, ripple within block.
-
Features: Speed-area trade-off; faster than RCA, simpler than CLA.
-
Arithmetic Circuits: Multipliers
-
Booth Multiplier (Radix-2):
-
Algorithm: Encode multiplier bits to reduce partial products (PP) by examining overlapping triplets (current, previous, next bit).
-
Encoding Table:
| $$\displaystyle m_{i+1} $$ | $$\displaystyle m_i $$ | $$\displaystyle m_{i-1} $$ | Operation | PP | | :--- | :--- | :--- | :--- | :--- | | 0 | 0 | 0 | 0 | 0 | | 0 | 0 | 1 | +Y | +Multiplicand | | 0 | 1 | 0 | +Y | | | 0 | 1 | 1 | +2Y | Shifted | | 1 | 0 | 0 | -2Y | | | 1 | 0 | 1 | -Y | -Multiplicand | | 1 | 1 | 0 | -Y | | | 1 | 1 | 1 | 0 | 0 |
-
Structure: Encoder → Shifter/Adder Tree → Final Adder.
-
Example: Multiply 7 (0111) by 3 (0011) using Booth recoding.
-
[!TIP] Exam Alert: You may be asked to draw the Booth recoding for a given 4-bit multiplier and show partial product generation.
IV. SEQUENTIAL CIRCUIT DESIGN
Latches and Flip-Flops
-
Design: Built from logic gates (e.g., cross-coupled NAND/NOR for SR latch; D latch from 2:1 MUX or gated SR).
-
Types:
-
SR Latch/FF: Asynchronous set/reset; invalid state $$\displaystyle S=R=1 $$.
-
D Latch/FF: Single data input; transparent when clock=1 (latch), edge-triggered (FF).
-
JK FF: Toggles when $$\displaystyle J=K=1 $$; resolves SR invalid state.
-
T FF: Toggles on clock if $$\displaystyle T=1 $$.
-
-
Timing Parameters:
-
Setup Time ($$\displaystyle t_{su} $$): Min $D$ stable before clock edge.
-
Hold Time ($$\displaystyle t_h $$): Min $D$ stable after clock edge.
-
Clock-to-Q Delay ($$\displaystyle t_{cQ} $$): Clock edge to output change.
-
Master-Slave Edge-Triggered Register
-
Operation Principle:
-
Master Latch: Enabled by $$\displaystyle \phi_1 $$ (e.g., clock LOW). Captures input $D$ during $$\displaystyle \phi_1 $$ HIGH.
-
Slave Latch: Enabled by $$\displaystyle \phi_2 $$ (complement of $$\displaystyle \phi_1 $$). Outputs master's state at falling edge of $$\displaystyle \phi_1 $$ (rising edge of $$\displaystyle \phi_2 $$).
-
-
Achieves Edge-Triggering: Output changes only at the negative clock edge (for $$\displaystyle \phi_1 $$ master).
-
Avoids Race-Around: Prevents $Q$ feedback to input within same clock cycle due to two-phase non-overlapping clocks.
DiagramCANVAS: Master-slave D flip-flop using two D latches, showing clock phases $$\displaystyle \phi_1 $$ and $$\displaystyle \phi_2 $$, and signal flow during a clock cycle.
V. TIMING, CLOCKING & PERFORMANCE
Propagation Delay Analysis (Elmore Delay)
- Application to Networks: For an RC tree, Elmore delay is the sum of resistance multiplied by capacitance to ground downstream.
$$ t_p \approx \sum_{i} R_i \cdot C_i \quad \text{(sum over all capacitors $$\displaystyle C_i $$ and resistance $$\displaystyle R_i $$ on path to $$\displaystyle C_i $$)} $$
- Used to estimate delay of complex gate/interconnect networks.
Clock Distribution in Synchronous Design
-
Goals: Low skew (difference in clock arrival times), low power, minimal load.
-
Techniques:
-
H-tree: Symmetric binary tree; equal path lengths from source to all leaves → low skew.
-
Grid (Mesh): Clock wires form a grid; low local skew, high power.
-
Clock Buffering/Repeers: Insert buffers/inverters to drive long wires, reduce RC delay.
-
Clock Gating: Insert AND gates with enable to stop clock to idle blocks → dynamic power saving.
-
[!TIP] Key Trade-off: H-tree minimizes skew but uses more area; grid is robust to local variations but power-hungry.
VI. ADVANCED DESIGN CONCEPTS & FPGA
Pipelining
-
Concept: Insert registers between combinational stages to break critical path.
-
Impact:
-
Throughput: Increases (one output per clock cycle after fill).
-
Latency: Increases (total delay = sum of stage delays + register overheads).
-
Clock Frequency: Increases (determined by slowest stage delay + $$\displaystyle t_{cQ} + t_{su} $$).
-
-
Hazards: Must handle data dependencies (forwarding/stalling) and control hazards.
Standard Cell Libraries
-
Contents:
-
Cell Views: Layout, symbol, schematic.
-
Timing Models: Look-up tables (LUTs) or nonlinear delay models (NLDM) for $$\displaystyle t_{pLH}, t_{pHL}, t_{cQ}, t_{su}, t_h $$ vs load/slew.
-
Power Models: Leakage and dynamic power per cell.
-
Physical Information: Height, pin locations.
-
-
Role: Foundation for automated synthesis (mapping RTL to cells) and place-and-route.
Field-Programmable Gate Arrays (FPGA)
-
Building Block Architecture:
-
Configurable Logic Blocks (CLBs): Contains LUTs (typically 4-6 input), flip-flops, multiplexers.
-
Input/Output Blocks (IOBs): Programmable I/O standards, drive strength.
-
Interconnect: Programmable switching matrix and routing tracks (global, long, short lines).
-
-
Programming Technologies:
| Technology | Volatility | Reprogrammability | Speed | Density | Cost | | :--- | :--- | :--- | :--- | :--- | :--- | | SRAM-based | Volatile | Unlimited | Moderate | Lower | Low | | Antifuse | Non-volatile | One-time | High | High | Moderate | | Flash | Non-volatile | Many (~10k) | Moderate-High | High | Higher | | Fuse | Non-volatile | One-time | High | Moderate | Low |
[!TIP] Exam Comparison: SRAM-FPGAs are most common (Xilinx, Intel) due to flexibility; Antifuse/Flash used for space/rad-hard apps where reprogrammability not needed.
END OF UNIT 2 NOTES