UNIT 3: VLSI CIRCUITS AND SYSTEMS - EX-803(C)
1.0 MOS TRANSISTOR FUNDAMENTALS & SCALING
1.1 Electrical Properties of MOS Transistor
-
Threshold Voltage ($$\displaystyle V_T $$): Minimum gate-to-source voltage to form inversion channel. Affected by body bias, oxide thickness, substrate doping.
-
Transconductance ($$\displaystyle g_m $$): Measure of current control by $$\displaystyle V_{GS} $$. In saturation:
$$g_m = \frac{\partial I_D}{\partial V_{GS}} = \mu_n C_{ox} \frac{W}{L} (V_{GS} - V_T) = \frac{2I_D}{V_{GS} - V_T}$$
- Output Conductance ($$\displaystyle g_{ds} $$): Due to channel length modulation (CLM):
$$g_{ds} = \lambda I_D \quad \text{(in saturation)}$$
- Mobility ($\mu$): Decreases with vertical electric field (mobility degradation):
$$\mu = \frac{\mu_0}{1 + \theta (V_{GS} - V_T)}$$
-
Channel Length Modulation: Effective channel length reduces as $$\displaystyle V_{DS} $$ increases, modeled by $\lambda$.
-
Subthreshold Conduction: Weak inversion current:
$$I_D \approx I_0 e^{(V_{GS} - V_T)/(nV_T)}$$
where $n$ is subthreshold slope factor.
-
Parasitic Capacitances:
-
$$\displaystyle C_{gs} $$: Gate-to-source (overlap + depletion).
-
$$\displaystyle C_{gd} $$: Gate-to-drain (Miller capacitance, critical in high-frequency).
-
$$\displaystyle C_{db} $$, $$\displaystyle C_{sb} $$: Drain/source-to-bulk junction capacitances.
-
[!TIP] In deep submicron, velocity saturation and short-channel effects dominate; use velocity-saturated Id model.
1.2 Scaling Principles and Models
-
Need for Scaling: Higher integration, performance, cost reduction.
-
Dennard Scaling (Constant Field Scaling):
-
Scale all dimensions ($$\displaystyle L, W, T_{ox} $$) by $1/s$, $$\displaystyle V_{DD} $$ by $1/s$, doping by $s$.
-
Electric field constant, power density constant.
-
-
Constant Voltage Scaling: Only dimensions scale, $$\displaystyle V_{DD} $$ constant → higher fields, increased power density.
-
Scaling Effects:
| Parameter | Effect of Scaling (Dennard) | |---|---| | $W, L$ | ↓ (by $1/s$) | | $$\displaystyle T_{ox} $$ | ↓ (by $1/s$) | | $$\displaystyle V_{DD} $$ | ↓ (by $1/s$) | | Doping | ↑ (by $s$) | | Gate capacitance $$\displaystyle C_{ox} $$ | ↓ (by $$\displaystyle 1/s^2 $$) | | Drive current $$\displaystyle I_D $$ | ↓ (by $1/s$) | | Power per transistor | ↓ (by $$\displaystyle 1/s^3 $$) | | Power density | Constant |
-
Short-channel Effects:
-
DIBL: Drain-induced barrier lowering → $$\displaystyle V_T $$ decreases with $$\displaystyle V_{DS} $$.
-
$$\displaystyle V_T $$ roll-off: $$\displaystyle V_T $$ decreases as $L$ reduces.
-
Velocity saturation: Current saturates at lower $$\displaystyle V_{DS} $$.
-
-
Fundamental Limits: Quantum tunneling, dopant fluctuation, variability.
[!TIP] Dennard scaling broke ~0.18µm due to short-channel effects and power density limits.
1.3 Fundamental Units of CMOS Inverter
-
Static CMOS Inverter: Pull-up PMOS ($P$), pull-down NMOS ($N$). No static power when $$\displaystyle V_{in} $$ is logic 0 or 1.
-
Transfer Characteristics: S-shaped curve. Switching point $$\displaystyle V_{M} $$ where $$\displaystyle I_{Dn} = -I_{Dp} $$. For $$\displaystyle \beta_n = \beta_p $$, $$\displaystyle V_M \approx V_{DD}/2 $$.
-
Noise Margins:
$$\boxed{NM_H = V_{OH} - V_{IH}, \quad NM_L = V_{IL} - V_{OL}}$$
For symmetric inverter: $$\displaystyle V_{OH} \approx V_{DD} $$, $$\displaystyle V_{OL} \approx 0 $$, $$\displaystyle V_{IH} \approx V_{IL} \approx V_{DD}/2 $$, so $$\displaystyle NM \approx V_{DD}/2 $$.
-
Power Consumption:
-
Static: Leakage (subthreshold, junction).
-
Dynamic: $$\displaystyle P_{dyn} = \alpha C_L V_{DD}^2 f $$, where $\alpha$ = switching activity.
-
Short-circuit: During transition, both $N$ and $P$ on → current spike.
-
-
Rise/Fall Delays:
$$t_{pdr} \approx 0.69 R_p C_L, \quad t_{pdf} \approx 0.69 R_n C_L$$
$$\displaystyle C_L $$ includes gate capacitances of driven gates and interconnect.
[!TIP] Equal rise/fall delays require $$\displaystyle \beta_n = \beta_p $$ and symmetric loads.
2.0 VLSI CIRCUIT DESIGN & LAYOUT
2.1 Layout Design Rules
-
Purpose: Ensure manufacturability, avoid defects, correct connectivity.
-
Types:
-
Micron Rules: Absolute dimensions (µm).
-
Lambda ($\lambda$) Rules: Scalable, based on minimum feature size. Common: width/spacing = $2\lambda$.
-
-
Well and Substrate Contacts: $n$-well to $$\displaystyle V_{DD} $$, $p$-substrate to $GND$. Minimum contact size/spacing.
-
Layer Rules:
-
Active Area (Diffusion): Minimum width/spacing, must enclose poly gate.
-
Poly: Minimum width/spacing, must overlap active for gate.
-
Metal: Minimum width/spacing, contacts via vias to poly/diffusion.
-
-
Minimum Width and Spacing: Prevent bridging, ensure etch/printability.
2.2 Layout of Basic Gates
-
CMOS NAND (2-input):
-
PDN: NMOS in series.
-
PUN: PMOS in parallel.
-
DiagramCANVAS: Two NMOS in series sharing diffusion; two PMOS in parallel with separate diffusions tied to $$\displaystyle V_{DD} $$; poly gates crossing; metal1 output; contacts to $$\displaystyle V_{DD} $$/$GND$.
-
-
CMOS NOR (2-input):
-
PDN: NMOS in parallel.
-
PUN: PMOS in series.
-
DiagramCANVAS: Two NMOS in parallel with separate diffusions; two PMOS in series sharing diffusion; poly gates; metal1 output.
-
-
Stick Diagrams:
-
Simplified: lines for each layer (poly horizontal, diffusion vertical, metal1/2 alternating).
-
NAND: poly crosses two series NMOS diffusion, two parallel PMOS diffusions.
-
NOR: poly crosses two parallel NMOS diffusions, two series PMOS diffusions.
-
Used for early area/routing estimation.
-
[!TIP] In stick diagrams, use consistent layer ordering: poly (horizontal), diffusion (vertical), metal1 (horizontal), metal2 (vertical).
2.3 Combinational Circuit Design using CMOS
-
Design Methodology:
-
Derive PDN from logic function (sum-of-products → parallel-series).
-
PUN is dual of PDN (series-parallel).
-
Ensure no static power: PDN and PUN never both on.
-
-
Complex Gate Design:
-
AOI (AND-OR-Invert): e.g., $$\displaystyle F = (A·B + C·D)' $$.
-
PDN: parallel of two series pairs ($A·B$ and $C·D$).
-
PUN: series of two parallel pairs ($A+B$ and $C+D$).
-
-
OAI (OR-AND-Invert): e.g., $$\displaystyle F = (A+B·C+D)' $$.
-
PDN: series of two parallel pairs ($A+B$ and $C+D$).
-
PUN: parallel of two series pairs ($A·C$ and $B·D$? Adjust based on function).
-
-
XOR: Can use transmission gates or complex PDN/PUN (e.g., $A'B + AB'$).
-
-
Truth Table Verification: Simulate schematic/layout for all input combinations; check for correct output and no static paths.
[!TIP] Use duality: complement function → swap series/parallel and NMOS/PMOS.
2.4 Alternative Logic Styles
-
Pass Transistor Logic (PTL):
-
Use NMOS/PMOS as switches to pass signals.
-
Example: XOR using NMOS pass network.
-
Advantages: Fewer transistors, smaller area, lower capacitance.
-
Disadvantages: Threshold voltage drop (NMOS passes weak 1), degraded swing, slower for large loads.
-
-
Transmission Gate Logic:
-
Parallel NMOS/PMOS controlled by $C$ and $\overline{C}$.
-
No threshold loss, full swing, bidirectional.
-
-
Comparison:
| Feature | Static CMOS | PTL | Transmission Gate | |---|---|---|---| | Transistor Count | High | Low | Moderate | | Speed | Fast | Slower (threshold) | Fast | | Area | Larger | Smaller | Similar to CMOS | | Power | Moderate | Lower dynamic | Moderate | | Robustness | High | Low (swing loss) | High |
[!TIP] Use transmission gates for multiplexers and level restoration; PTL for area-critical paths with voltage scaling.
2.5 Transmission Gate
-
Structure: NMOS and PMOS in parallel, gates controlled by $C$ and $\overline{C}$.
-
Operation: When $$\displaystyle C=1 $$, conducts both 0 and 1 with low resistance; bidirectional switch.
-
Applications:
-
Multiplexers: Select between inputs.
-
Bus Switches: Connect/disconnect bus lines.
-
Level Restorers: In PTL to restore full swing.
-
-
Advantages over Single Pass Transistor:
-
No $$\displaystyle V_T $$ drop: passes strong 0 and 1.
-
Symmetric on-resistance for high/low.
-
Better for AC and high-frequency signals.
-
[!TIP] Always use complementary control; single NMOS only for passing 0 or in low-power with level shifting.
3.0 TIMING ANALYSIS & INTERCONNECT
3.1 RC Delay Model
- Elmore's Constant: First-order time constant for RC tree. For input node $i$:
$$t_{pd} = \sum_{j} R_i \cdot C_j$$
where $$\displaystyle C_j $$ are all capacitors downstream of $$\displaystyle R_i $$.
- Propagation Delay for Inverter (lumped $$\displaystyle C_L $$):
$$\boxed{t_{pd} \approx 0.69 R_{eq} C_L}$$
$$\displaystyle R_{eq} = R_n $$ for falling, $$\displaystyle R_p $$ for rising.
-
Lumped vs. Distributed:
-
Lumped: All $$\displaystyle C_L $$ at output node; simple but inaccurate for long wires.
-
Distributed: Wire as RC ladder; accurate for interconnect.
-
3.2 Interconnect Delays
-
Impact of Scaling: Wire $W, T$ scale slower than $L$ → $R \uparrow$ (since $R \propto L/(W \cdot T)$), $C$ ↓ slightly → RC product ↑, interconnect delay dominates.
-
Wire Resistance and Capacitance:
-
Resistance: $$\displaystyle R = \rho \cdot L / (W \cdot T) $$.
-
Capacitance: To ground and adjacent wires; use $\pi$-model or distributed RC.
-
-
Effect on Performance: Increased delay, crosstalk, power. Requires repeaters/buffers for long wires.
[!TIP] In deep submicron, interconnect delay often > gate delay; use wire sizing and shielding.
4.0 SEQUENTIAL CIRCUIT DESIGN
4.1 Latches and Flip-Flops
-
Methodology: Feedback loop stores state. Latch = level-sensitive; Flip-flop = edge-triggered.
-
SR Latch:
-
NOR-based (active high): $$\displaystyle Q = S + \overline{R} \cdot Q $$, avoid $$\displaystyle S=R=1 $$.
-
NAND-based (active low): $$\displaystyle Q = \overline{S + \overline{R} \cdot Q} $$.
-
-
Clocked Latch (Transparent Latch):
-
SR latch with enable $E$.
-
When $$\displaystyle E=1 $$, transparent ($Q$ follows $D$); when $$\displaystyle E=0 $$, holds state.
-
Implemented with transmission gates or gated inverters.
-
4.2 Edge-Triggered Registers
-
Master-Slave Flip-Flop:
-
Two latches: master (positive level) and slave (negative level).
-
On rising clock: master captures $D$, slave holds old $Q$.
-
On falling clock: slave updates with master's value.
-
Overall: rising-edge triggered.
-
DiagramCANVAS: Master latch (clk) → slave latch (clk') with feedback.
-
-
Pulse-Triggered Flip-Flops: Use clock pulse to sample and update (e.g., dynamic flip-flops).
-
Timing Parameters:
-
Setup Time ($$\displaystyle t_{su} $$): $D$ stable before clock edge.
-
Hold Time ($$\displaystyle t_h $$): $D$ stable after clock edge.
-
Clock-to-Q Delay ($$\displaystyle t_{cq} $$): Clock edge to $Q$ change.
-
Minimum Clock Period: $$\displaystyle T_{clk} \ge t_{cq} + t_{su} + t_{comb} $$.
-
[!TIP] Master-slave avoids race-through but adds $$\displaystyle t_{cq} $$; use for reliable edge-triggering.
4.3 Clock Distribution
-
Need: Synchronize all registers in synchronous design.
-
Clock Skew: Difference in clock arrival times. Positive skew helps $$\displaystyle t_{su} $$ but hurts $$\displaystyle t_h $$; negative skew opposite.
-
Techniques:
-
Clock Tree Synthesis (CTS): Build balanced H-tree or buffered tree.
-
H-tree: Symmetric structure for equal delay.
-
Buffered Clock Trees: Insert buffers to drive loads and balance delays.
-
Clock Gating: Insert enable gates to stop clock to idle blocks → power saving.
-
-
Goals: Minimize skew, latency, load; avoid glitches.
[!TIP] Use CTS with buffer insertion and clock shielding to reduce skew and noise.
5.0 ARITHMETIC CIRCUITS
5.1 Adders
-
Ripple-Carry Adder (RCA):
-
Chain of full adders; carry ripples.
-
Delay: $O(n)$ gate delays (critical path through all carries).
-
-
Carry Look-Ahead Adder (CLA):
-
Generate: $$\displaystyle G_i = A_i \cdot B_i $$.
-
Propagate: $$\displaystyle P_i = A_i \oplus B_i $$.
-
Carry: $$\displaystyle C_{i+1} = G_i + P_i C_i $$.
-
Block CLA: Group $k$ bits → group $$\displaystyle G_{i:j} $$, $$\displaystyle P_{i:j} $$:
-
$$G_{i:j} = G_j + P_j G_{j-1} + \cdots + P_j \cdots P_{i+1} G_i$$
$$P_{i:j} = P_j \cdot P_{j-1} \cdots P_i$$
-
Delay: $O(\log n)$ with large hardware.
-
Carry Bypass Adder:
-
Skip carry chain if all $$\displaystyle P_i=1 $$ in a block.
-
Speed between RCA and CLA; area moderate.
-
-
Comparison:
| Type | Speed | Area | Power | |---|---|---|---| | RCA | Slow | Small | Low | | Carry Bypass | Medium | Medium | Medium | | CLA | Fast | Large | High |
5.2 Multipliers
-
Array Multiplier:
-
Shift-and-add: generate partial products, sum with adder array.
-
Baugh-Wooley: For signed numbers, adjust signs of partial products to avoid sign extension; uses same array.
-
-
Booth Multiplier:
-
Radix-2 Booth: Encode 2 bits with overlap → recode to reduce partial products by ~2.
-
Radix-4 Booth: 3-bit overlap → further reduction. Encoding:
| $$\displaystyle B_{2i+1}B_{2i}B_{2i-1} $$ | Operation | |---|---| | 000 | 0 | | 001, 010 | $+A$ | | 011 | $+2A$ | | 100 | $-2A$ | | 101, 110 | $-A$ | | 111 | 0 |
-
Example: Multiply $A \times B$ (signed). Show recoding steps and partial product generation.
-
Structure: Booth encoder/decoder, partial product generation, adder tree.
-
-
Wallace Tree / Dadda Tree:
-
Reduce partial products in parallel using carry-save adders (CSAs).
-
Wallace: irregular, faster; Dadda: more regular, slightly slower.
-
Final carry-propagate adder (e.g., CLA) for sum.
-
[!TIP] Booth reduces number of adders; Wallace reduces tree height; both improve multiplier speed.
6.0 PROGRAMMABLE LOGIC & DESIGN METHODOLOGIES
6.1 Field Programmable Gate Arrays (FPGAs)
6.1.1 Programming Technologies
-
SRAM-based: Volatile, unlimited reprogramming, fast, high power, larger area (SRAM cells). Common (Xilinx, Intel).
-
Antifuse-based: One-time programmable, non-volatile, low power, high density, cannot reconfigure. Used in some CPLDs.
-
Flash-based: Non-volatile, reprogrammable (limited cycles), moderate power/area. Used in some FPGAs (Microchip).
-
Comparison:
| Feature | SRAM | Antifuse | Flash | |---|---|---|---| | Volatility | Volatile | Non-volatile | Non-volatile | | Reprogrammability | Unlimited | One-time | Limited | | Speed | Fast | Slower | Moderate | | Power | High | Low | Moderate | | Density | Lower | Higher | Moderate | | Cost | Higher | Lower | Moderate |
6.1.2 Building Block Architecture
-
Configurable Logic Blocks (CLBs): Main logic units. Typically contain:
-
LUTs (Look-Up Tables): e.g., 4-input LUT for any 4-input function.
-
Flip-flops: For sequential logic.
-
-
Input/Output Blocks (IOBs): Interface to external pins. Support I/O standards, tri-state, pull-ups.
-
Programmable Interconnect Resources: Routing channels with switch boxes (crosspoints) and connection boxes (to CLBs/IOBs).
-
Clock Management Resources: PLLs/DCMs for clock synthesis, deskew, frequency multiplication.
[!TIP] FPGA architecture trades flexibility for density; more LUTs/routing increase capacity but also delay.
6.2 Design Methodologies & Supporting Concepts
6.2.1 Pipelining
-
Concept: Insert registers between combinational stages to break critical path → increase throughput.
-
Pipeline Stages: Each stage = combinational logic + register. Clock period = max(stage delay).
-
Hazards:
-
Structural: Resource conflicts (solve with duplication).
-
Data: Read-after-write (solve with forwarding/stalling).
-
Control: Branches (solve with prediction).
-
-
Balancing Pipeline Stages: Adjust logic in each stage to equalize delays → maximize clock frequency.
6.2.2 Standard Cell Libraries
-
Components:
-
Cell Views:
-
Symbol: Schematic representation.
-
Layout: Physical geometry.
-
Abstract: Abstract geometry for P&R (bounding box, pins).
-
Timing: Delay models (e.g., NLDM, CCS).
-
Power: Leakage, internal, switching power.
-
-
-
Characterization:
-
Timing: Delay vs. load capacitance, input transition.
-
Power: Leakage (process/voltage/temp), dynamic (switching activity).
-
Noise: Crosstalk, ground bounce.
-
-
Role in ASIC Flow: Pre-designed, characterized cells used by place-and-route tools to build design.
6.2.3 Stick Diagrams
-
Purpose: Quick, approximate layout planning. Show relative placement and routing without exact geometry.
-
Representation: Colored strips for layers:
-
Poly (horizontal), Diffusion (vertical), Metal1 (horizontal), Metal2 (vertical), etc.
-
Contacts as dots.
-
-
Use in Early Planning: Estimate area, routing congestion, check design rules approximately.
-
Relation to Layout Design Rules: Each strip width = minimum width, spacing = minimum spacing in $\lambda$ rules.
[!TIP] Convert stick diagram to layout by expanding strips to full geometry and adding contacts/vias.