UNIT 4: VLSI Design - Advanced Device Modeling, Fabrication, Digital Structures & Systems
1.0 VLSI Fabrication Process & Technology
1.1 CMOS Fabrication Processes
-
n-well CMOS Process:
-
Start with p-type substrate.
-
Grow field oxide (LOCOS) for isolation.
-
Implant n-well (for pMOS).
-
Grow gate oxide.
-
Deposit and pattern polysilicon gate.
-
Lightly doped drain (LDD) implants.
-
Source/Drain implants (n+ for nMOS, p+ for pMOS).
-
Annealing.
-
Metallization and passivation.
-
-
p-well CMOS Process: Reverse doping; n-well on p-substrate becomes p-well on n-substrate. Generally inferior due to lower nMOS mobility.
-
Twin-Tub (Twin-Well) CMOS Process:
-
Steps: Start with p-substrate. Create isolated n-well and p-well regions using separate masks and implants. Both transistors are formed in their respective wells.
-
Advantages over n-well/p-well:
-
Better matching of nMOS and pMOS threshold voltages.
-
Reduced body effect for both transistors.
-
Higher packing density (no large n-well covering entire pMOS area).
-
Lower substrate resistance, reducing latch-up susceptibility.
-
-
Disadvantages: More complex and costly due to extra well formation steps.
-
-
Enhancement Techniques: LDD, pocket implants (halo), salicide (self-aligned silicide), chemical-mechanical polishing (CMP) for planarization.
[!TIP] Exam often asks for step-by-step cross-sections. Memorize the sequence for n-well and twin-tub. Key difference: twin-tub has both wells defined on a common substrate.
1.2 Key Fabrication Steps & Definitions
| Step | Purpose | Key Concept |
|---|---|---|
| Oxidation | Grow SiO2 layer for gate oxide, field oxide, or mask. | Thermal oxidation; dry (slow, dense) vs. wet (fast, porous). |
| Photolithography | Transfer pattern from mask to wafer. | Photoresist coating, UV exposure through mask, development. Critical for feature size. |
| Metallization | Form interconnects. | Deposition (sputtering, evaporation) & patterning (etch). |
| Ion Implantation & Diffusion | Doping to create source/drain, wells. | Implantation: precise dose/energy. Diffusion: high-temp drive-in. |
| Etching | Remove material selectively. | Wet (chemical) vs. Dry (plasma/RIE - anisotropic). |
1.3 Design Rules & Layout
-
λ-based Design Rules: Express all layout dimensions (width, spacing, overlap) as multiples of a single parameter λ (half the minimum feature size). Ensures design is independent of exact process scaling.
-
Minimum Width (W_min): e.g.,
W_min = 2λfor metal. -
Minimum Spacing (S_min): e.g.,
S_min = 2λbetween polysilicon lines. -
Overlap (O): e.g.,
O_poly_to_contact = λ.
-
-
Importance: Guarantees manufacturability, prevents shorts/opens, and allows automatic design rule checking (DRC).
-
Layout Area Comparison: CMOS requires both nMOS and pMOS transistors, so area ~2x NMOS for same logic function. But CMOS has near-zero static power, justifying area cost.
1.4 Latch-Up in CMOS Circuits
-
Physical Origin: Parasitic p-n-p-n thyristor structure formed by the n-well/p-substrate (for nMOS) and p-substrate/n-well (for pMOS) junctions, with base resistance.
p+ (pMOS source) -> n-well -> p-substrate -> n+ (nMOS source) -
Triggering Mechanisms:
-
Forward-biased junctions: e.g., input voltage > Vdd+0.7V (nMOS) or < Vss-0.7V (pMOS).
-
Substrate/well resistance (R): High R causes voltage drop, forward-biasing parasitic base-emitter junctions.
-
Power supply transients: dI/dt causes voltage spikes on supply lines.
-
-
Internal Prevention Techniques:
-
Guard Rings: Highly doped p+ (around nMOS) and n+ (around pMOS) rings tied to Vss/Vdd to shunt minority carrier current.
-
Epitaxial Substrate: Thin, high-resistivity epi-layer on heavily doped substrate reduces lateral resistance.
-
Substrate Contacts: Frequent contacts to well/substrate to minimize resistance.
-
Triple-Well Isolation: Deep n-well isolates p-well (and its parasitic n-p-n) from p-substrate.
-
[!TIP] Latch-up is a low-impedance, high-current state. Prevention focuses on breaking the thyristor feedback loop (reducing β or R).
1.5 Packaging and Testing
-
Importance: Protects die from environment, provides electrical connection, dissipates heat. Testing ensures only functional ICs are shipped.
-
Packaging Types: DIP, PLCC, QFP, BGA. Steps: die attach, wire bonding (or flip-chip), encapsulation.
-
Testing:
-
Wafer Probing (Parametric/Functional): Test dies on wafer before sawing.
-
Burn-in Test: Operate ICs at elevated temperature (125°C) and voltage for 48-168 hours to accelerate failure mechanisms (infant mortality). Purpose: Identify weak devices before shipment.
-
Final Test: Test packaged IC for functionality, speed, power.
-
2.0 Semiconductor Device Modeling (DC & Small-Signal)
2.1 MOSFET Models
-
DC Models - I-V Characteristics:
-
Cutoff: $$\displaystyle V_{GS} < V_{th} $$, $$\displaystyle I_D = 0 $$.
-
Triode (Linear): $$\displaystyle V_{GS} > V_{th} $$, $$\displaystyle V_{DS} < V_{GS} - V_{th} $$.
-
$$I_D = \mu_n C_{ox} \frac{W}{L} \left[ (V_{GS}-V_{th})V_{DS} - \frac{V_{DS}^2}{2} \right] \left(1 + \lambda V_{DS}\right)$$
* **Saturation:** $$\displaystyle V_{GS} > V_{th} $$, $$\displaystyle V_{DS} \ge V_{GS} - V_{th} $$.
$$I_D = \frac{1}{2} \mu_n C_{ox} \frac{W}{L} (V_{GS}-V_{th})^2 (1 + \lambda V_{DS})$$
-
Level 1 (MOS1) Large Signal Model: Uses above square-law equations. Assumes constant mobility, no short-channel effects. Limitations: Inaccurate for modern short-channel devices.
-
Level 2 (MOS2) Model: Improves upon Level 1 by including:
-
Mobility degradation: $$\displaystyle \mu = \mu_0 / (1 + \theta (V_{GS}-V_{th})) $$.
-
Velocity saturation: $$\displaystyle I_{Dsat} $$ limited by $$\displaystyle v_{sat} $$.
-
Threshold voltage roll-off: $$\displaystyle V_{th} $$ decreases with decreasing L.
-
Drain-induced barrier lowering (DIBL): $$\displaystyle V_{th} $$ decreases with increasing $$\displaystyle V_{DS} $$.
-
-
Body Effect: $$\displaystyle V_{th} $$ increases with substrate bias ($$\displaystyle V_{SB} > 0 $$ for nMOS).
$$V_{th} = V_{th0} + \gamma \left( \sqrt{|2\phi_F + V_{SB}|} - \sqrt{|2\phi_F|} \right)$$
where $$\displaystyle \gamma = \sqrt{2q \varepsilon_{si} N_{sub}} / C_{ox} $$ (body effect coefficient).
-
Short-Channel Devices:
-
Effects: DIBL, velocity saturation, channel length modulation (λ), increased leakage.
-
Advantages: Faster switching, lower capacitance, higher density.
-
Limitations: Higher leakage, lower $$\displaystyle V_{th} $$ control, increased sensitivity to variations.
-
-
Subthreshold Operation: $$\displaystyle V_{GS} < V_{th} $$, weak inversion. Current exponential in $$\displaystyle V_{GS} $$.
$$I_D \approx I_0 e^{(V_{GS}-V_{th})/(nV_T)} \left(1 - e^{-V_{DS}/V_T}\right)$$
where $n$ is subthreshold slope factor (~1.3), $$\displaystyle V_T = kT/q $$. In short-channel devices, DIBL reduces effective $$\displaystyle V_{th} $$, making subthreshold leakage worse.
2.2 BJT Models
- Ebers-Moll Model: Two-diode representation.
$$I_E = I_{ES} \left( e^{V_{BE}/V_T} - 1 \right) - \alpha_R I_{CS} \left( e^{V_{BC}/V_T} - 1 \right)$$
$$I_C = \alpha_F I_{ES} \left( e^{V_{BE}/V_T} - 1 \right) - I_{CS} \left( e^{V_{BC}/V_T} - 1 \right)$$
where $$\displaystyle \alpha_F, \alpha_R $$ are forward/reverse common-base current gains (~0.99-0.999), $$\displaystyle I_{ES}, I_{CS} $$ are saturation currents.
-
Temperature Dependence:
-
$$\displaystyle I_S \propto T^3 e^{-E_g/(kT)} $$ → doubles per ~10°C.
-
$$\displaystyle V_{BE} $$ decreases by ~2mV/°C.
-
Mitigation: Biasing with negative feedback, bandgap reference circuits.
-
-
High-Frequency Behavior: $$\displaystyle f_T = g_m / (2\pi (C_{be} + C_{bc})) $$. Miller effect multiplies $$\displaystyle C_{bc} $$ by $$\displaystyle (1 + g_m R_L) $$, reducing bandwidth.
2.3 Diode Models
- Shockley Diode Equation:
$$I = I_S \left( e^{V/(nV_T)} - 1 \right)$$
* $$\displaystyle I_S $$: Reverse saturation current (material/area dependent).
* $n$: Ideality factor (1 for diffusion, 2 for recombination).
- Significance: $$\displaystyle I_S $$ sets leakage; $n$ indicates dominant current transport mechanism.
2.4 Passive Component Models in ICs
| Component | Implementation | Model/Challenges |
|---|---|---|
| Resistor | Diffusion (high resistivity), Poly (low), Well. | Parasitic capacitances to substrate, temperature coefficient. |
| Capacitor | Parallel-plate (metal-insulator-metal, poly-insulator-poly). | Fringing capacitance becomes significant at small sizes. |
| Inductor | Spiral (Al/Cu) on multiple metal layers. | Very low Q-factor (<10), large area, substrate losses. |
3.0 Circuit Simulation (SPICE)
3.1 Need and Significance
-
Role: Virtual prototyping. Verify functionality, performance (speed, power), and robustness (noise, temperature) before costly fabrication.
-
Significance: Reduces design iterations, catches errors early, enables optimization.
3.2 Types of SPICE Analyses
-
DC Analysis: Operating point, transfer curves (e.g., $$\displaystyle V_{out} $$ vs. $$\displaystyle V_{in} $$). Sweeps sources/temp.
-
AC Analysis: Small-signal frequency response. Linearizes around DC operating point. Gives gain/phase vs. frequency.
-
Transient Analysis: Time-domain response to arbitrary input (pulse, sine). Non-linear, computationally intensive.
-
Noise Analysis: Calculates noise contribution of devices (thermal, flicker) at a given frequency.
3.3 Device Model Implementation in SPICE
-
MOSFET:
.MODEL M1 NMOS (LEVEL=1/2/3 ...)with parameters (VTO, KP, LAMBDA, GAMMA, etc.). -
BJT:
.MODEL Q1 NPN (LEVEL=1 ...)using Ebers-Moll or Gummel-Poon. -
Diode:
.MODEL D1 D (IS=... N=...). -
Noise Modeling: SPICE includes thermal (white) and flicker ($1/f$) noise models for MOSFETs based on process parameters.
4.0 Digital System Structures & Architectures
4.1 Logic Design Styles
| Aspect | Random Logic | Structured Logic |
|---|---|---|
| Definition | Custom, ad-hoc layout for each logic block. | Regular, repeatable structures (arrays, matrices). |
| Design Effort | High (manual). | Low (automated/regular). |
| Area Efficiency | Can be optimal for small blocks. | May have overhead (routing, control). |
| Scalability | Poor. | Excellent. |
| Testability | Difficult. | Easier (regular patterns). |
| Example | Custom adder, small FSM. | PLA, ROM, systolic array, gate array. |
[!TIP] Advantage of Structured Logic: Design reuse, shorter time-to-market, easier verification and testing.
4.2 Register Storage Circuits
-
Static Register Cell: Cross-coupled inverters (bistable). Holds state as long as power is on. Timing Parameters:
-
Setup Time ($$\displaystyle t_{su} $$): Min. time data must be stable before clock edge.
-
Hold Time ($$\displaystyle t_h $$): Min. time data must be stable after clock edge.
-
Propagation Delay ($$\displaystyle t_{pd} $$): Time from clock edge to output change.
-
Derivation: Based on internal gate delays and clock-to-Q delay ($$\displaystyle t_{cQ} $$). $$\displaystyle t_{su} \ge t_{cQ} + t_{logic} - t_{clock\_period} $$.
-
-
Dynamic Register Cell: Stores charge on a capacitance (gate of MOSFET). Requires periodic refresh (2-phase clock). Higher density, but charge leakage limits speed.
-
Quasi-Static Register Cell: Combines static and dynamic. Uses a static master stage and dynamic slave (or vice versa). Avoids some static cell problems (e.g., write margin) while being less sensitive to clock timing than fully dynamic.
-
Astatic Register Cell: Multiplication Process & Linear System Solver:
-
Principle: Uses current-mode logic (CML) or differential pairs. Data represented as currents. No static power consumption (hence "a-static").
-
Multiplication: Implemented using transconductance amplifiers ($$\displaystyle g_m $$). Output current $$\displaystyle I_{out} = g_m \cdot V_{in} $$. Cascaded stages perform polynomial multiplication.
-
Linear System Solver: Systolic array of processing elements (PEs) with local communication. Each PE performs a simple operation (multiply-accumulate). Data flows through array in a wavefront manner, solving $$\displaystyle Ax=b $$ iteratively.
-
4.3 Microcoded Controllers
-
Architecture:
-
Control Store (ROM/PLA): Stores microinstructions.
-
Microinstruction Register (μIR): Holds current microinstruction.
-
Sequencer: Generates address for next microinstruction (based on condition codes, jump fields).
-
-
Operation:
-
Fetch: Sequencer outputs address → ROM outputs microinstruction → loaded into μIR.
-
Decode: μIR fields control datapath (register enables, ALU op, memory R/W).
-
Execute: Datapath performs operation.
-
Next Address: Sequencer determines next address (sequential, branch, call, return).
-
4.4 Systolic Arrays
-
Concept: 2D grid of identical Processing Elements (PEs). Data flows synchronously in a wavefront pattern from one PE to next (e.g., left to right, top to bottom). Each PE performs a simple, fixed operation on incoming data and passes result to neighbor.
-
Design for Parallel Processing:
-
Mesh Topology: Most common. PEs connected to N, S, E, W neighbors.
-
Data Flow: Inputs fed into array edges, results emerge from opposite edges. Regular timing (clocked registers between PEs).
-
-
Advantages:
-
High Throughput: Multiple computations in parallel.
-
Regular Structure: Easy to layout, test, and scale.
-
Local Communication: Short wires → low RC delay, high clock speed.
-
Fault Tolerance: Can bypass faulty PEs.
-
-
Implementation Challenges:
-
Synchronization: All PEs must clock synchronously; clock distribution critical.
-
Data Timing: Precise timing of data arrival at each PE (pipelining).
-
I/O Bottleneck: Limited number of input/output ports for the entire array.
-
Algorithm Mapping: Must map problem to systolic data flow (e.g., matrix multiply).
-
4.5 Specialized Architectures
-
Algotronix Architecture: Refers to a class of systolic array processors designed for high-performance numerical computations (e.g., matrix operations, convolutions). Structure is a linear or 2D array of PEs with nearest-neighbor connections. Significance: Demonstrated the power of regular, parallel architectures for signal processing and scientific computing, influencing designs like the Intel iWarp and many AI accelerators today.
-
Hybrid Technology: Integration of CMOS (logic) with Bipolar (high-speed I/O, analog) on the same chip.
-
Advantages: Best of both worlds: high-density/low-power CMOS logic + high-speed/bipolar drive strength for I/O pads and analog circuits.
-
Challenges: Complex process (different thermal budgets, isolation), higher cost.
-
5.0 Interconnects & Technology Comparison
5.1 Interconnects in CMOS Processing Technology
-
Materials: Aluminum (Al) → Copper (Cu, lower resistivity, better EM), Tungsten (W, for contacts/vias).
-
Layers: Multiple metal layers (M1, M2, ...) separated by dielectric (SiO₂, low-k materials).
-
Impact:
-
RC Delay: $$\displaystyle t_{pd} \propto R \cdot C $$. As dimensions scale, R increases (thinner wires), C decreases but coupling C increases. Interconnect delay dominates gate delay in deep sub-micron.
-
Crosstalk: Capacitive/inductive coupling between adjacent wires → noise, signal integrity issues.
-
Power: Dynamic power $$\displaystyle P = \alpha C V^2 f $$ includes load capacitance from interconnects.
-
5.2 Technology Comparison
| Technology | Area | Speed | Power | Integration Density | Key Use |
|---|---|---|---|---|---|
| NMOS | Small | Moderate | High static power | High (historical) | Obsolete. |
| CMOS | Larger (2x NMOS) | Moderate-High | Very low static | Very High | Dominant digital. |
| Bipolar | Large | Very High | High static & dynamic | Low | High-speed analog, I/O. |
| Hybrid | Largest | High (I/O) | Moderate | Moderate | Mixed-signal systems. |
5.3 Scaling in VLSI Design
-
Definition: Shrinking all dimensions by factor S (>1) while maintaining functionality.
-
Constant Field Scaling (Dennard): Scale voltage $V$ and dimensions by $1/S$. Electric fields constant.
-
Effects:
-
Delay: $t \propto 1/S$ → faster.
-
Power density: $$\displaystyle P/A \propto S^0 $$ → constant (ideal).
-
Switching energy: $$\displaystyle E \propto 1/S^2 $$ → lower.
-
-
Breaks down when $$\displaystyle V_{th} $$ doesn't scale, short-channel effects dominate, leakage increases.
-
-
Generalized Scaling: Allows independent scaling of dimensions, voltage, doping. Used in modern nodes.
-
Effects:
-
Short-channel effects: DIBL, velocity saturation become severe.
-
Leakage: Subthreshold, gate oxide tunneling increase exponentially.
-
Power density: Often increases due to leakage and higher clock rates.
-
Interconnect: RC delay becomes dominant bottleneck.
-
-
[!TIP] Scaling question often asks for trade-offs: Speed ↑, Density ↑, Power/area ↑ (leakage), Design complexity ↑.