UNIT 4: Cloud Computing - Comprehensive Short Notes
1.0 Foundational Concepts & Paradigms
1.1 Evolution of Distributed Computing
Grid Computing: A distributed architecture that pools resources from multiple locations to achieve a common goal. It focuses on coordinated resource sharing across organizational boundaries for large-scale, compute-intensive tasks.
-
Architecture: Consists of Resource Nodes (computers, storage), Middleware (for discovery, scheduling, security), and User Interface.
-
Characteristics: Heterogeneous resources, non-centralized control, collaborative goal-oriented.
Cloud Computing: A model for enabling ubiquitous, convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, services).
-
Essential Characteristics (NIST):
-
On-demand self-service: Provision resources automatically.
-
Broad network access: Available over the network via standard mechanisms.
-
Resource pooling: Multi-tenant model with physical/virtual resources dynamically assigned.
-
Rapid elasticity: Resources scale rapidly outward/inward.
-
Measured service: Resource usage monitored, controlled, and billed.
-
[!TIP] Exam Focus: Grid is for collaborative, large-scale projects (e.g., scientific research), while Cloud is for on-demand, scalable services (e.g., web apps). Grid resources are often dedicated to a single task; Cloud resources are shared (multi-tenant).
Comparative Analysis: Grid vs. Cloud Computing
| Feature | Grid Computing | Cloud Computing |
|---|---|---|
| Primary Goal | Solve large, complex problems | Deliver on-demand IT services |
| Resource Ownership | Resources contributed by various organizations | Owned & managed by a single provider |
| Management | Decentralized, federated | Centralized by provider |
| Scalability | Scale by adding more grid nodes | Elastic scaling (instant up/down) |
| Business Model | Often non-commercial, research-focused | Commercial, utility-based (pay-per-use) |
| Standardization | Less standardized interfaces | Standardized APIs & services (SPI) |
| Similarity | Both use distributed resources for parallel processing. |
Utility Computing & Computing on Demand
-
Utility Computing: A business model where computing resources (processing, storage) are provided as a metered service, similar to traditional utilities (electricity, water). It is the commercial precursor to cloud computing.
-
Computing on Demand: Enables dynamic provisioning of resources based on real-time demand. Users request resources via a portal/API, and the system automatically allocates them from a pooled resource set. Billing is based on consumption (e.g., per hour, per GB).
1.2 Cloud Service Models (SPI Model)
| Service Model | Definition | Control Level (User) | Management Responsibility | Examples |
|---|---|---|---|---|
| IaaS<br>(Infrastructure) | Provides fundamental computing resources: VMs, storage, networks. User controls OS & apps. | High (OS, Apps, Data) | CSP: Physical infra, virtualization. User: OS, middleware, apps. | AWS EC2, Azure VMs, Google Compute Engine |
| PaaS<br>(Platform) | Provides platform/environment to develop, test, deploy applications. User controls app & data. | Medium (Apps, Data) | CSP: OS, runtime, middleware. User: App code & data. | Heroku, Google App Engine, Azure App Services |
| SaaS<br>(Software) | Provides complete, ready-to-use applications over the internet. | Low (Data & Config) | CSP: Everything (app, data, OS, infra). User: Configuration & data input. | Gmail, Salesforce, Microsoft 365 |
[!TIP] Mnemonic: I control A lot in IaaS, Provide App code in PaaS, Simply Use in SaaS.
1.3 Cloud Deployment Models
| Model | Definition | Key Characteristics | Advantages | Disadvantages | Examples |
|---|---|---|---|---|---|
| Public Cloud | Cloud infrastructure made available to the general public over the internet. Owned by CSP. | Multi-tenant, pay-as-you-go, no CapEx. | Cost-effective, no maintenance, high scalability. | Less control, security/concerns, potential downtime. | AWS, Azure Public, Google Cloud |
| Private Cloud | Cloud infrastructure operated solely for a single organization. Can be on-premise or hosted. | Single-tenant, high control, customizable. | Enhanced security & control, compliance-friendly. | High CapEx/OpEx, limited scalability, management overhead. | On-prem VMware, OpenStack, hosted private cloud |
| Hybrid Cloud | Composition of two or more clouds (private/public) with orchestration between them. | Cloud bursting (scale to public during peaks), workload mobility. | Flexibility, cost optimization, avoids vendor lock-in. | Complex integration, security across boundaries, management complexity. | Private cloud + AWS/Azure for burst workloads |
| Community Cloud | Shared infrastructure supporting a specific community with common concerns (security, compliance, mission). | Shared by several organizations with shared goals. | Cost-sharing, tailored to community needs. | Limited scalability, still shared responsibility. | Government clouds, research consortium clouds |
Selecting a Deployment Model โ Decision Factors:
-
Cost: CapEx vs. OpEx, total cost of ownership.
-
Security & Compliance: Regulatory requirements (GDPR, HIPAA), data sensitivity.
-
Control & Customization: Need for deep infrastructure control.
-
Scalability Needs: Predictable vs. spiky workloads.
-
Performance & Latency: Data residency, proximity requirements.
-
Existing IT Investment: Legacy system integration.
2.0 Virtualization: The Core Enabling Technology
2.1 Virtualization Fundamentals
-
Definition: Creation of a virtual (rather than actual) version of something, including virtual computer hardware platforms, storage devices, and network resources.
-
Need in Cloud: Enables resource pooling, multi-tenancy, isolation, and efficient utilization of physical hardware.
-
How it Works: A Hypervisor (VMM) sits between physical hardware and guest OS, abstracting physical resources and presenting virtual versions (vCPUs, vRAM, vDisk) to each VM.
-
Benefits:
-
Resource Utilization: Multiple VMs on one physical server (consolidation).
-
Isolation: VMs are isolated from each other (crash, security).
-
Flexibility: Easy provisioning, cloning, migration.
-
Manageability: Centralized management, snapshots, templates.
-
2.2 Virtualization Architecture & Types
Hypervisors (Virtual Machine Monitors)
| Type | Description | Examples | Key Points |
|---|---|---|---|
| Type-1<br>(Native/Bare-metal) | Runs directly on the host's hardware. No underlying OS. | VMware ESXi, Microsoft Hyper-V, Xen, KVM | High performance & security. Used in production clouds & data centers. |
| Type-2<br>(Hosted) | Runs on top of a conventional OS (like an application). | Oracle VirtualBox, VMware Workstation, Parallels | Easy setup, used for development/testing. Lower performance. |
Hardware-Assisted Virtualization (HVM): Uses CPU extensions (Intel VT-x, AMD-V) to improve virtualization performance and support unmodified guest OSes. Essential for modern Type-1 hypervisors.
Full Virtualization vs. Para-Virtualization
-
Full Virtualization: Guest OS runs unmodified. Hypervisor traps and emulates privileged instructions. Example: VMware ESXi, Hyper-V (with HVM).
-
Para-Virtualization: Guest OS is modified (has special drivers) to make hypercalls for sensitive operations, avoiding emulation. Higher performance. Example: Original Xen, Microsoft Hyper-V (integration services).
2.3 Server Virtualization & Partitioning
Logical Partitioning (LPAR)
-
Definition: Division of a physical server's resources (CPU, memory, I/O) into multiple isolated logical units called partitions. Each partition runs its own OS instance.
-
Implementation: Primarily on IBM POWER systems using PowerVM.
-
Advantages:
-
Precise, static resource allocation (guaranteed minimums).
-
Strong isolation and security between partitions.
-
Enables running different OS types on same physical server.
-
-
Disadvantages:
-
Less dynamic than full virtualization (resources often fixed).
-
Complex management.
-
Vendor-specific (IBM).
-
Virtual Machines (VMs)
-
Lifecycle: Create โ Provision โ Run โ Suspend/Resume โ Migrate โ Delete.
-
VM Sprawl: Uncontrolled proliferation of VMs due to easy creation, leading to management, security, and resource inefficiency issues.
-
VM Migration:
- Live Migration: Move a running VM from one physical host to another with minimal downtime. Requires shared storage and compatible hypervisors.
2.4 Virtualized Data Center Architecture
A virtualized data center abstracts and pools compute, storage, and network resources.
-
Components:
-
Virtualized Servers: VMs on clustered physical hosts.
-
Virtualized Storage: SAN/NAS presented as virtual disks (vDisks) to VMs.
-
Virtualized Networking: Virtual switches (vSwitch), virtual NICs (vNIC), software-defined networking (SDN) for network isolation and policy.
-
-
Management/Orchestration Layer: Software (e.g., vCenter, OpenStack Nova) that automates provisioning, monitoring, and lifecycle management of all virtual resources.
-
DiagramCANVAS: Show physical servers (compute cluster), shared storage (SAN), physical network switches. Overlay: virtual machines connected via virtual switches to virtual networks, all managed by a central management console/orchestrator.
2.5 Storage Virtualization
-
Definition: Abstraction of physical storage resources from their physical location, presenting them as a single, logical pool.
-
Types:
-
Block-level: Presents raw storage blocks (LUNs). Used by VMs and databases. Example: SAN.
-
File-level: Presents files and directories. Used for shared file access. Example: NAS.
-
SAN (Storage Area Network) vs. NAS (Network Attached Storage)
| Feature | SAN | NAS |
|---|---|---|
| Access Method | Block-level (SCSI, Fibre Channel, iSCSI) | File-level (NFS, CIFS/SMB) |
| Protocol | Fibre Channel (FC), iSCSI, FCoE | NFS (Unix/Linux), CIFS/SMB (Windows) |
| Architecture | Dedicated high-speed network (FC). Hosts see storage as local disk. | Standard Ethernet/IP network. Hosts mount remote file systems. |
| Performance | Very high, low latency. | Good, but higher latency than SAN (file protocol overhead). |
| Use Case | Databases, VM storage (vSphere), high-transaction apps. | File sharing, home directories, web content, backups. |
| Example | EMC VMAX, NetApp FAS (in SAN mode) | NetApp FAS (in NAS mode), Isilon, FreeNAS |
Storage Cloud Concepts: Cloud storage services (e.g., AWS S3, Azure Blob) provide object-based storage over HTTP/HTTPS, offering massive scalability, durability, and pay-per-use pricing. They abstract the underlying storage infrastructure completely.
2.6 Requirements for a Virtualization Platform in Cloud Implementation
-
Performance: Low overhead, efficient resource scheduling (CPU, memory, I/O).
-
Scalability: Support thousands of VMs across clustered hosts.
-
Security: Strong isolation between VMs, secure management interfaces, support for encryption (vTPM).
-
Manageability: Centralized console, APIs for automation, monitoring tools.
-
Compatibility: Support for various guest OSes and hardware.
-
API Support: Rich, well-documented APIs for integration with cloud management platforms (CMPs) and orchestration tools (e.g., OpenStack, vRealize).
3.0 Cloud Architecture, Design & Integration
3.1 Cloud Computing Reference Model
NIST Cloud Computing Reference Architecture:
-
Cloud Consumer: Person/organization that uses cloud services.
-
Cloud Provider: Entity that provides cloud services (SPI).
-
Cloud Auditor: Independent assessor of cloud controls.
-
Cloud Broker: Agent that manages use/performance of cloud services.
-
Cloud Carrier: Provides connectivity between consumer/provider.
-
Key Components: Cloud Service Management (orchestration, provisioning, SLA), Cloud Security Management, Cloud Resource Management (physical/virtual resources).
Cloud Computing Stack (Cloud Stack):
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ SaaS (Software) โ โ End-user applications
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ PaaS (Platform) โ โ Dev tools, DBs, middleware
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ IaaS (Infrastructure) โ โ VMs, storage, networks
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ Physical Hardware (Servers, etc.) โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
- Management/Orchestration Layer: Sits above the stack to manage provisioning, metering, and monitoring across all service models.
3.2 Service-Oriented Architecture (SOA) in Cloud
-
SOA Principles: Services are loosely coupled, self-contained, reusable, and composable.
-
Components: Services (business functionality), Orchestration (choreographing services), Registry/Repository (service discovery).
-
Role in Cloud Design:
-
Enables building modular, scalable applications from discrete cloud services.
-
Facilitates interoperability between services from different CSPs or on-premise systems.
-
Supports agile development and integration of heterogeneous systems.
-
-
Web Services & APIs:
-
REST (Representational State Transfer): Lightweight, uses HTTP verbs (GET, POST), stateless. Dominant in modern cloud APIs (AWS, Azure).
-
SOAP: Protocol-heavy, XML-based, WS-* standards for security/transactions. Used in enterprise legacy integrations.
-
[!TIP] SOA vs. Microservices: SOA is an architectural style; Microservices is a finer-grained implementation of SOA principles, often used in cloud-native apps.
3.3 Ecosystem & Stakeholders
-
Independent Software Vendors (ISVs): Develop applications (e.g., ERP, CRM) that run on cloud platforms (PaaS/IaaS). They leverage cloud scalability for deployment and monetize via SaaS or marketplace models. Example: SAP on AWS, Salesforce AppExchange.
-
Cloud Service Provider (CSP): Owns and operates the cloud infrastructure, provides SPI services, manages SLAs, security, and billing. Example: AWS, Microsoft, Google.
-
Cloud Consumer: The end-user organization or individual that subscribes to and uses cloud services.
4.0 Security in Cloud Computing
4.1 Importance & Unique Challenges
-
Shared Responsibility Model:
-
CSP Responsibility: Security OF the cloud (physical infra, hypervisor, network fabric, managed services).
-
Customer Responsibility: Security IN the cloud (data, IAM, OS/network config, applications).
-
Division varies by service model: In SaaS, CSP manages more; in IaaS, customer manages more.
-
-
Challenges:
-
Data Breaches: Unauthorized access to sensitive data.
-
Account Hijacking: Stolen credentials leading to data theft/service abuse.
-
Insider Threats: Malicious or negligent employees (CSP or customer).
-
Compliance: Meeting regulatory standards (GDPR, HIPAA) in a shared, dynamic environment.
-
Multi-tenancy Risks: "Noisy neighbor" attacks, side-channel attacks.
-
Loss of Control: Limited visibility into physical infrastructure.
-
4.2 Multi-Faceted Nature of Cloud Security
-
Data Security:
-
Encryption: At rest (AES-256 on storage), in transit (TLS 1.2/1.3).
-
Data Integrity: Hashes (SHA-256), digital signatures.
-
Data Residency/Locality: Legal requirements about where data can be stored.
-
-
Network Security:
-
Virtual Networks (VPCs, Subnets), Security Groups/Network ACLs (stateful/stateless firewalls).
-
IDS/IPS: Intrusion detection/prevention systems.
-
DDoS Protection: Mitigation services (AWS Shield, Azure DDoS Protection).
-
-
Identity & Access Management (IAM):
-
Authentication: Verifying identity (MFA, SSO).
-
Authorization: Defining permissions (policies, roles).
-
Federation: Trust relationships between identity providers (SAML, OIDC).
-
Role-Based Access Control (RBAC):
-
Definition: Access control based on roles (job function) rather than individual identities.
-
Model Components:
-
Users (subjects)
-
Roles (collections of permissions)
-
Permissions (privileges on resources)
-
Sessions (user activation of a set of roles)
-
-
Implementation: AWS IAM (Users, Groups, Policies, Roles), Azure AD (Roles, Role Assignments). Principle of least privilege is critical.
-
-
-
Service & Application Security: Secure API design (rate limiting, input validation), vulnerability scanning, patching (customer responsibility for IaaS/PaaS).
-
Compliance & Legal Issues:
-
Standards: ISO 27001 (security), SOC 2 (trust services), GDPR (data privacy), HIPAA (healthcare).
-
CSPs provide compliance certifications and shared responsibility matrices.
-
4.3 Securing Virtualized Environments & VMs
VM-Specific Security Risks:
-
VM Escape: Malware in a VM breaks out to compromise the hypervisor/host.
-
VM Hopping: Attacker moves from one VM to another on same host.
-
VM Sprawl: Unmanaged VMs become attack vectors.
-
Image/Snapshot Security: Stolen or tampered VM images containing sensitive data/configs.
-
Hypervisor Attacks: Targeting the privileged VMM layer.
Benefits of Virtualization Security:
-
Isolation: VMs are isolated at the hardware level (stronger than process isolation).
-
Encapsulation: Entire VM state (OS, apps, data) is a single file, enabling easy backup, scanning, and quarantine.
Securing VMs - Recommendations:
-
Hardened VM Images/Templates: Start from minimal, patched, secured base images. Use image scanning.
-
Network Segmentation (Micro-segmentation): Use virtual firewalls (security groups) to control East-West traffic between VMs. Apply zero-trust principles.
-
VM Activity Monitoring & Logging: Monitor VM metrics, process creation, network connections. Centralize logs (SIEM).
-
Secure VM Lifecycle Management: Enforce policies for provisioning, de-provisioning, and patching. Use immutable infrastructure where possible.
4.4 Secure Execution & Communication
-
Secure Execution Environments:
-
Trusted Platform Module (TPM): Hardware chip for cryptographic operations (key storage, attestation).
-
vTPM: Virtual TPM emulated by the hypervisor for VMs, enabling measured boot and disk encryption (BitLocker).
-
Secure Boot: Ensures only signed, trusted bootloaders and OS kernels are loaded. Uses UEFI and certificates.
-
-
Secure Communications:
-
Encryption Protocols: TLS/SSL for data in transit (HTTPS, database connections). Use strong ciphers and certificate management.
-
VPNs over Cloud: Site-to-site VPNs (IPsec) or client VPNs for secure access to cloud VPCs.
-
-
Secure Bootstrapping: Mechanisms to securely initialize cloud instances (e.g., using cloud-init with encrypted user data, signed images).
5.0 Management, Performance & Advanced Topics
5.1 Quality of Service (QoS) in Cloud
-
Definition: The overall performance and reliability of a cloud service, often defined in a Service Level Agreement (SLA).
-
Key Metrics:
-
Availability: Uptime percentage (e.g., 99.9%).
-
Reliability: Mean Time Between Failures (MTBF).
-
Performance: Throughput, response time, latency.
-
Capacity: Resource allocation guarantees.
-
-
Issues & Challenges:
-
Resource Allocation: Fair sharing among multi-tenant users.
-
Performance Isolation: Preventing "noisy neighbor" from degrading others' performance.
-
SLA Negotiation & Monitoring: Defining measurable metrics and continuous monitoring against them.
-
Elasticity Overhead: Rapid scaling can cause transient performance issues.
-
5.2 Cloud Infrastructure Benchmarks
-
Purpose: Objectively evaluate and compare performance, scalability, and cost-effectiveness of cloud services/infrastructures.
-
Key Metrics & Tools:
-
SPEC Cloud: Industry-standard benchmarks for IaaS (CPU, memory, storage, network).
-
YCSB (Yahoo! Cloud Serving Benchmark): For evaluating cloud data services (NoSQL, key-value stores).
-
Custom Workloads: Simulating real application behavior.
-
-
Performance Evaluation: Involves testing under various load conditions, measuring throughput/latency, and analyzing cost-performance ratios.
5.3 Cloud Management Platforms & Tools
-
Open Source Tools:
-
OpenNebula:
-
Architecture: Modular. Core components: Driver (infrastructure interface), Scheduler (resource allocation), Cloud View (API/CLI for users).
-
Use: Provides IaaS cloud management. Supports multiple hypervisors (KVM, LXC, VMware) and data center technologies. Good for private/hybrid clouds.
-
-
Nimbus:
-
Architecture: Lightweight, focused on cloud computing for science. Key components: Nimbus Context Broker (EC2-compatible interface), Workspace Service (VM lifecycle).
-
Features: Strong security (X.509, delegation), support for multiple backends (KVM, Xen).
-
Use Cases: Scientific research clouds, academic institutions.
-
-
-
Proprietary Consoles: AWS Management Console, Azure Portal, Google Cloud Console. Provide integrated, user-friendly interfaces for managing specific CSP services.
5.4 Data Management & Analytics in Cloud
-
OLAP (Online Analytical Processing) in Cloud:
-
Functionality: Enables complex analytical queries against large, historical, aggregated datasets (data warehouses/marts).
-
Core Operations:
-
Roll-up: Aggregating data (sum, avg) by climbing a hierarchy (e.g., city โ country โ region).
-
Drill-down: Opposite of roll-up; viewing finer-grained details.
-
Slice: Selecting a single dimension value (e.g., "sales in 2023").
-
Dice: Selecting on multiple dimensions (e.g., "sales in 2023 for product X in Europe").
-
Pivot (Rotate): Reorienting the multidimensional view (swapping rows/columns).
-
-
Cloud-based Data Warehouses: Fully managed, petabyte-scale analytics services.
-
Examples: Amazon Redshift, Google BigQuery, Snowflake, Azure Synapse.
-
Benefits: Separation of compute/storage, automatic scaling, pay-per-query, built-in high availability.
-
-
5.5 Challenges & Risks of Adoption
Business Perspective:
-
Vendor Lock-in: Difficulty migrating due to proprietary APIs, services, and data egress costs.
-
Cost Management: Unpredictable bills (especially with auto-scaling), need for FinOps practices.
-
Business Continuity: Dependency on CSP's availability; requires multi-region/cloud strategies.
IT Perspective:
-
Integration with Legacy Systems: Connecting old on-premise apps to cloud services can be complex.
-
Skills Gap: Need for new skills (cloud architecture, DevOps, security) different from traditional sysadmin.
-
Performance Variability: "Noisy neighbor" effect, shared resource performance can be inconsistent compared to dedicated hardware.
-
Security & Compliance: Shared responsibility model confusion, meeting regulatory requirements.