UNIT 1: FOUNDATIONS OF CLOUD COMPUTING & VIRTUALIZATION
1.0 Core Concepts & Evolution
1.1 Defining Cloud Computing
Cloud Computing is a model for enabling ubiquitous, convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, services) that can be rapidly provisioned and released with minimal management effort or service provider interaction.
Five Essential Characteristics (NIST):
-
On-demand self-service: Provision computing capabilities automatically without human interaction.
-
Broad network access: Available over the network via standard mechanisms (e.g., web browsers, mobile apps).
-
Resource pooling: Provider's resources pooled to serve multiple consumers using a multi-tenant model.
-
Rapid elasticity: Capabilities can be elastically provisioned and released to scale rapidly.
-
Measured service: Resource usage monitored, controlled, and reported, enabling pay-per-use billing.
Computing on Demand is the paradigm where resources are dynamically allocated as needed, enabling dynamic provisioning—the automatic scaling up or down of resources based on real-time demand.
Comparison with Traditional Utility Models:
| Feature | Electricity/Water Utility | Cloud Computing |
|---|---|---|
| Resource | Physical commodity (electricity, water) | Digital resource (compute, storage, apps) |
| Delivery | Grid/pipeline network | Internet/network |
| Metering | Usage-based (kWh, gallons) | Usage-based (compute hours, GB stored) |
| Provisioning | Always-on, constant flow | On-demand, elastic, can be turned off |
1.2 Grid Computing vs. Cloud Computing
| Aspect | Grid Computing | Cloud Computing |
|---|---|---|
| Primary Goal | Solve large-scale, complex scientific problems (e.g., climate modeling) by aggregating distributed resources. | Deliver on-demand, scalable IT resources and services as a utility. |
| Architecture | Decentralized, heterogeneous resources from multiple administrative domains. | Centralized, large-scale data centers with homogeneous, standardized hardware. |
| Resource Management | Complex, often using middleware (e.g., Globus Toolkit) for job scheduling across domains. | Centralized management by provider; simpler APIs for self-service provisioning. |
| Ownership | Resources owned by different organizations/institutions (federated). | Resources owned and managed by a single cloud provider (or private entity). |
| Target Apps | High-performance computing (HPC), batch processing. | Web applications, enterprise software, development platforms, big data analytics. |
| Scalability Model | Scale by adding more distributed nodes to the grid. | Scale vertically (bigger VMs) and horizontally (more VMs) within a pooled infrastructure. |
| Billing Model | Often free for academic/research use; complex cost allocation. | Simple pay-per-use subscription or utility model. |
| User Experience | Requires expertise to submit jobs to a queue; less focus on UI. | Focus on user-friendly self-service portals and APIs. |
Similarities: Both use distributed computing, aim for resource sharing, and can involve virtualization. Evolution: Grid computing's focus on resource sharing and federation laid conceptual groundwork, but cloud computing evolved to provide a simpler, more commercial, service-oriented model with centralized control and economies of scale.
1.3 Cloud Computing Reference Model
The NIST Cloud Computing Reference Model defines five key functional components:
-
Cloud Consumer: Individual/organization that uses cloud services.
-
Cloud Provider: Entity responsible for making cloud services available.
-
Cloud Auditor: Independent entity that conducts audits.
-
Cloud Broker: Manages use, performance, and delivery of cloud services.
-
Cloud Carrier: Provides connectivity between consumers and providers.
Architectural Layers (Service Stack):
+---------------------+
| Cloud Software | (SaaS - Applications)
+---------------------+
| Cloud Platform | (PaaS - Runtime, Middleware)
+---------------------+
| Cloud Infrastructure | (IaaS - VMs, Storage, Network)
+---------------------+
| Physical Hardware | (Servers, Storage, Networking)
+---------------------+
Interaction: A consumer accesses a service (e.g., SaaS) which may be built on a PaaS platform, which in turn runs on IaaS virtualized infrastructure, all hosted on physical hardware managed by the provider.
2.0 Cloud Service Models (SPI Model)
2.1 Infrastructure as a Service (IaaS)
-
Definition: Provides fundamental computing resources—virtual machines (VMs), storage, and networks—over the internet. Users can provision and manage these resources.
-
Core Offerings: Virtual servers (VMs/containers), block storage (EBS), virtual networks (VPC), load balancers.
-
Examples: Amazon EC2, Microsoft Azure VMs, Google Compute Engine.
-
Management Responsibility:
| Provider Manages | Consumer Manages | | :--- | :--- | | Physical hardware, data center, network backbone | OS, middleware, runtime, applications, data | | Hypervisor, virtualization layer | VM configuration, security groups, patching OS |
2.2 Platform as a Service (PaaS)
-
Definition: Provides a platform allowing customers to develop, run, and manage applications without dealing with the underlying infrastructure.
-
Core Offerings: Development tools, databases (SQL/NoSQL), middleware, operating systems, development environments.
-
Examples: Google App Engine, Microsoft Azure App Service, Heroku.
-
Management Responsibility:
| Provider Manages | Consumer Manages | | :--- | :--- | | Everything in IaaS + OS, runtime, middleware, development tools | Application code, data, application configuration |
2.3 Software as a Service (SaaS)
-
Definition: Delivers complete, ready-to-use software applications over the internet, typically via a web browser.
-
Core Offerings: Complete applications (email, CRM, collaboration tools).
-
Examples: Gmail, Salesforce, Microsoft 365, Dropbox.
-
Management Responsibility:
| Provider Manages | Consumer Manages | | :--- | :--- | | Everything (application, data, runtime, middleware, OS, infrastructure) | User accounts, application-specific data & configuration |
2.4 Differentiating Service Models
Control vs. Management Overhead Trade-off:
| Model | Control (Low -> High) | Flexibility (Low -> High) | Management Overhead (Low -> High) | Best For |
|---|---|---|---|---|
| SaaS | Low | Low | Very Low | End-users needing ready-made apps (email, CRM). |
| PaaS | Medium | Medium | Medium | Developers building apps without infrastructure ops. |
| IaaS | High | High | High | IT admins needing full control over VMs & networks. |
Choosing a Model: Decision based on need for control vs. desire for reduced operational burden. Start with SaaS for standard apps, use PaaS for custom development, choose IaaS for legacy apps or maximum customization.
3.0 Cloud Deployment Models
3.1 Public Cloud
-
Definition: Cloud infrastructure made available to the general public or a large industry group, owned by a cloud service provider.
-
Architecture: Multi-tenant, massive scale data centers.
-
Ownership: Third-party provider (e.g., AWS, Azure, GCP).
-
Benefits: No CapEx, high scalability, broad reach, no maintenance.
-
Risks: Data privacy concerns, limited customization, potential vendor lock-in, shared tenancy risks.
-
Use Cases: Web applications, development/test environments, non-sensitive data processing.
3.2 Private Cloud
-
Definition: Cloud infrastructure operated solely for a single organization. Can be on-premise or hosted.
-
Architecture: Can be on-premise (local data center) or hosted (dedicated servers at provider's site).
-
Ownership: The organization itself or a dedicated provider.
-
Benefits: Maximum control, enhanced security & compliance, customization.
-
Risks: High CapEx/OpEx, requires in-house expertise, limited scalability vs. public.
-
Use Cases: Highly regulated data (finance, government), legacy applications, strict compliance needs.
3.3 Hybrid Cloud
-
Definition: Composition of two or more distinct cloud infrastructures (private & public) that remain unique entities but are bound together by standardized or proprietary technology enabling data and application portability.
-
Architecture: Orchestration between private and public clouds with secure connectivity (VPN, dedicated links).
-
Key Use Cases:
-
Cloud Bursting: Run apps in private cloud, burst to public cloud during demand spikes.
-
Workload Migration: Move non-sensitive workloads to public cloud for cost savings.
-
Disaster Recovery: Use public cloud as a backup site for private cloud.
-
-
Benefits: Flexibility, cost optimization, avoids vendor lock-in.
-
Risks: Complexity in management & security, network latency between clouds, integration challenges.
3.4 Community Cloud
-
Definition: Cloud infrastructure shared by several organizations with shared concerns (e.g., security requirements, policy, compliance considerations). It may be managed by the organizations or a third party.
-
Example: A cloud shared by multiple government agencies for classified projects.
-
Benefits: Cost-sharing, tailored to specific community needs.
-
Risks: Limited scalability, potential for conflicting requirements among members.
3.5 Selecting a Deployment Model
Decision Framework:
-
Security & Compliance: High sensitivity? → Private/Community. Lower sensitivity? → Public/Hybrid.
-
Cost Structure: Avoid CapEx? → Public. Have budget for investment? → Private.
-
Control & Customization: Need deep control? → Private/IaaS. Accept standard services? → Public/SaaS.
-
Scalability Needs: Predictable, steady load? → Private. Variable, spiky load? → Public/Hybrid.
-
Workload Type: Legacy, sensitive apps? → Private. New, scalable web apps? → Public.
Common Pitfall: Assuming "private cloud" is just virtualized on-premise data center. True private cloud includes self-service, automation, and measured service—not just virtualization.
4.0 Virtualization: The Foundational Technology
4.1 Introduction to Virtualization
-
Definition: Creation of a virtual (rather than actual) version of something, including virtual computer hardware platforms, storage devices, and network resources.
-
Purpose:
-
Abstraction: Decouples software from physical hardware.
-
Consolidation: Run multiple VMs on a single physical server, increasing utilization from ~15% to 60-80%.
-
Isolation: VMs are isolated from each other; a crash in one doesn't affect others.
-
Encapsulation: Entire VM state (OS, apps, data) is stored as files (images), enabling easy backup, cloning, and migration.
-
4.2 Hypervisors (Virtual Machine Monitors)
Software that creates and runs VMs.
| Type | Type-1 (Native/Bare-Metal) | Type-2 (Hosted) |
|---|---|---|
| Architecture | Runs directly on host hardware. No underlying OS. | Runs on top of a conventional OS (host OS). |
| Examples | VMware ESXi, Microsoft Hyper-V, Xen, KVM | VMware Workstation, Oracle VirtualBox, Parallels |
| Performance | High (direct hardware access) | Lower (host OS adds overhead) |
| Use Case | Server virtualization, data centers, cloud infrastructure. | Desktop virtualization, software testing, development. |
| Security | Smaller attack surface (no host OS). | Larger attack surface (host OS vulnerabilities). |
Key Hypervisor Functions:
-
Resource Arbitration: Allocates physical CPU, memory, I/O to VMs.
-
Isolation: Ensures VMs cannot access each other's resources.
-
Emulation: Presents virtual hardware ( NIC, disk controller) to guest OS.
4.3 Virtual Machine (VM) Concepts
-
VM Lifecycle:
Create→Start→Suspend(save state to disk) →Resume→Pause→Shutdown→Delete. -
VM Image/Template: A file (or set of files) containing the complete state of a VM (OS, apps, config). Used for rapid provisioning.
-
VM Security Risks in Cloud:
-
VM Escape: Attacker breaks out of VM to access host hypervisor.
-
VM Sprawl: Uncontrolled proliferation of VMs leading to management gaps and attack surfaces.
-
Inter-VM Traffic Monitoring: Sniffing traffic between VMs on same physical host (requires virtual switch security).
-
Insecure VM Images: Templates with default passwords, unpatched OS.
-
Resource Exhaustion: "Noisy neighbor" VM consuming excessive host resources.
-
4.4 Server Virtualization & Virtualized Data Centers
-
Architecture: Physical servers (compute) run a Type-1 hypervisor. Virtual switches (vSwitch) connect VMs. Virtual storage is presented from a centralized pool (SAN/NAS). Management servers orchestrate the entire environment.
-
Benefits:
-
Resource Pooling: CPU, memory, storage pooled across many physical servers.
-
Dynamic Allocation: Live migration (VMotion, Live Migration) moves running VMs between hosts for load balancing.
-
High Availability (HA): If a physical host fails, VMs automatically restart on another host.
-
Disaster Recovery: VM images can be replicated to another site easily.
-
4.5 Storage Virtualization
-
Definition: Abstraction of physical storage from logical storage, pooling resources from multiple storage devices into a single, unified management console.
-
Architecture: A virtualization engine (software or hardware) sits between hosts and physical storage devices.
-
SAN vs. NAS:
| Feature | SAN (Storage Area Network) | NAS (Network-Attached Storage) | | :--- | :--- | :--- | | Access Level | Block-level (raw disks/LUNs) | File-level (NFS, SMB/CIFS) | | Protocol | Fibre Channel (FC), iSCSI | NFS, SMB/CIFS | | Performance | Very high, low latency | Good, but higher latency than SAN | | Use Case | Databases, high-transaction apps, VMs (for boot volumes) | File sharing, home directories, content repositories | | Example | Dell EMC PowerStore, NetApp AFF | NetApp FAS, QNAP, Synology |
4.6 Advanced Partitioning Technologies: Logical Partitioning (LPAR)
-
Definition: Firmware-level partitioning of a single physical server into multiple independent logical partitions, each with its own dedicated resources (CPU, memory, I/O). Common in IBM Power Systems, Oracle SPARC, and mainframes.
-
Implementation: Handled by the system's hypervisor (e.g., PowerVM, Hypervisor for SPARC) at a lower level than typical software hypervisors.
-
Advantages:
-
Strong Isolation: Hardware-enforced separation; partitions cannot affect each other.
-
Fine-grained Resource Allocation: Dedicated or shared resources with strict guarantees.
-
High Security & Reliability: Suitable for consolidating mixed workloads (e.g., production, test, development) on one machine.
-
Dynamic Resource Adjustment: Resources can be moved between partitions without reboot.
-
-
Disadvantages/Limitations:
-
Complexity: Requires specialized knowledge to configure and manage.
-
Vendor Lock-in: Tied to specific hardware architectures (Power, SPARC).
-
Less Flexible than Software Virtualization: Cannot run arbitrary guest OSes (only supported OSes like AIX, Linux, IBM i).
-
4.7 Virtualization Platforms & Tools
-
OpenNebula:
-
Architecture: Open-source cloud management platform. Uses a modular design with a central Sunstone (web UI), OneFlow (orchestration), and OneGate (gateway). Interfaces with multiple hypervisors (KVM, LXC, VMware) and storage/network backends.
-
Features: VM lifecycle management, networking, storage, multi-tenancy, hybrid cloud support.
-
Role: Provides the management layer to transform a virtualized data center (with hypervisors) into an IaaS cloud with self-service portals, APIs, and automated provisioning.
-
-
Nimbus:
-
Overview: Open-source toolkit for providing Infrastructure as a Service (IaaS). Primarily focused on scientific and research clouds.
-
Components: Nimbus Cloud Controller (provides EC2-compatible API), Workspace Service (manages VM lifecycle), Context Broker (injects configuration into VMs).
-
Use Case: Building private research clouds where EC2 API compatibility is desired.
-
-
Requirements for a Virtualization Platform to Implement Cloud:
-
Hypervisor Support: Must manage one or more hypervisors (KVM, ESXi, Hyper-V).
-
API & Self-Service: Provide RESTful APIs and web UI for users to provision resources.
-
Multi-Tenancy & Isolation: Support for separate projects/tenants with network and storage isolation.
-
Orchestration & Automation: Ability to define complex VM deployments (groups of VMs) via templates.
-
Accounting & Metering: Track resource usage per user/tenant for billing or chargeback.
-
High Availability & Fault Tolerance: Ensure VMs restart on failure.
-
Networking: Virtual network creation (VLANs, overlays), IP management, security groups.
-
Storage Management: Integration with block/object storage, attaching volumes to VMs.
-
5.0 Cloud Security Fundamentals
5.1 The Shared Responsibility Model
Core Principle: Security is a shared responsibility between the Cloud Provider and the Cloud Customer. The division of responsibility varies by service model (SPI).
| Responsibility | SaaS | PaaS | IaaS |
|---|---|---|---|
| Provider | Security OF the Cloud: Physical security, network, hypervisor, host OS, application, data (by default). | Security OF the Cloud: Physical, network, hypervisor, host OS, runtime, middleware, development tools. | Security OF the Cloud: Physical, network, hypervisor, host OS. |
| Customer | Security IN the Cloud: User access, application config, data, client device security. | Security IN the Cloud: Application code, data, application config, user access. | Security IN the Cloud: OS configuration & patching, network config (security groups), applications, data, user access. |
Key Takeaway: The more control you have (IaaS), the more responsibility you have. In SaaS, provider handles almost everything; in IaaS, customer is responsible for securing the OS and above.
5.2 Multi-Faceted Nature of Cloud Security
-
Data Security:
-
Encryption: At Rest (storage encryption, database TDE), In Transit (TLS/SSL for all communications).
-
Data Loss Prevention (DLP): Tools to monitor and prevent unauthorized data exfiltration.
-
Key Management: Secure generation, storage, rotation, and destruction of encryption keys (use provider's KMS or bring your own).
-
-
Network Security:
-
Virtual Networks (VPC/VNet): Isolate customer resources in a private network segment.
-
Firewalls: Security Groups (stateful), Network ACLs (stateless) at subnet/VM level.
-
IDS/IPS: Monitor for malicious traffic within the virtual network.
-
DDoS Protection: Provider-level mitigation (e.g., AWS Shield, Azure DDoS Protection).
-
Segmentation: Use multiple subnets/VPCs to separate tiers (web, app, db).
-
-
Identity & Access Management (IAM):
-
Authentication: Verifying identity (MFA, SSO, federated identity).
-
Authorization: Defining permissions (see RBAC below).
-
Single Sign-On (SSO): Allows one set of credentials to access multiple cloud services/apps.
-
Role-Based Access Control (RBAC):
-
Concept: Permissions are assigned to roles (e.g.,
Admin,Developer,ReadOnly), and users/entities are assigned roles. -
Implementation: Cloud IAM services (AWS IAM, Azure AD) use RBAC. Define policies (JSON in AWS) that attach to roles/users/groups, specifying allowed actions (
s3:GetObject) on resources (arn:aws:s3:::bucket/*). -
Principle of Least Privilege: Grant only the minimum permissions necessary.
-
-
-
Service & Application Security: Secure APIs (authentication, rate limiting), secure coding practices, application vulnerability scanning.
-
Compliance & Governance: Understanding provider's compliance certifications (SOC 2, ISO 27001, GDPR, HIPAA). Use audit logs (CloudTrail, Azure Activity Log) for monitoring. Implement governance policies (e.g., "no public S3 buckets").
5.3 Securing the Virtualized Environment
-
Secure VM Execution:
-
Use hardened VM images (minimal OS, no default accounts, patched).
-
Install and update guest OS security (antivirus, host-based firewalls).
-
Apply VM-specific security controls (VMware Tools/guest agents for security).
-
-
Secure Communications:
-
Enforce TLS 1.2+ for all data in transit (client-to-cloud, inter-service, VM-to-VM).
-
Use VPNs (IPsec, SSL VPN) for secure remote access to VPCs.
-
-
Secure Bootstrapping:
-
Trusted Platform Module (TPM): Hardware-based security for key storage and measured boot.
-
Measured Boot/Integrity Monitoring: Hypervisor or host measures boot components (firmware, bootloader, OS kernel) and stores hashes in TPM. Can attest to VM integrity at launch.
-
-
Monitoring & Logging:
-
VM Activity Monitoring: Track process creation, file access, network connections within VMs (via agents).
-
Centralized Logging: Aggregate VM syslogs, application logs, hypervisor logs to a central SIEM (e.g., CloudWatch Logs, Azure Monitor).
-
5.4 Virtualization Security Benefits & Challenges
-
Benefits:
-
Isolation: VMs are isolated from each other and the host (stronger with Type-1 hypervisors).
-
Encapsulation: Entire VM state is a file, enabling easy inspection, backup, and rollback.
-
Fine-Grained Policy Enforcement: Security policies can be applied per VM or per security group.
-
Rapid Recovery: VM images can be quickly redeployed on another host.
-
-
Challenges:
-
VM Sprawl: Uncontrolled creation of VMs increases attack surface and management complexity.
-
Hypervisor Attacks: A vulnerability in the hypervisor (the " crown jewels") could compromise all VMs on the host.
-
Inter-VM Traffic Monitoring: Traffic between VMs on the same host may not pass through physical network taps, requiring virtual network monitoring (port mirroring on vSwitch).
-
Snapshot Risks: VM snapshots may contain sensitive data in memory and should be encrypted.
-
Resource Sharing Risks: "Noisy neighbor" attacks or side-channel attacks (e.g., Spectre/Meltdown) in multi-tenant clouds.
-
6.0 Enabling Architectures & Business Aspects
6.1 Service-Oriented Architecture (SOA) & Cloud
-
SOA Principles:
-
Services: Self-contained, modular units of functionality (e.g., "Payment Service").
-
Loose Coupling: Services interact via well-defined interfaces (APIs), minimizing dependency.
-
Interoperability: Services use standards-based protocols (HTTP, SOAP, REST) to work across platforms.
-
Reusability: Services designed for reuse in different applications.
-
Discoverability: Services can be found and understood via registries (UDDI).
-
-
How SOA Facilitates Cloud Integration:
-
Cloud services are naturally exposed as SOA-compliant web services (RESTful APIs).
-
Loose coupling allows applications to consume cloud services (e.g., AWS S3, Salesforce CRM) without deep integration.
-
Enables composition: Building complex cloud applications by orchestrating multiple cloud-based services (e.g., using AWS Step Functions).
-
Cloud Design with SOA: Design applications as a collection of microservices (an evolution of SOA) deployed as independent, scalable units in the cloud (e.g., containers on Kubernetes).
-
6.2 Business & IT Perspectives on Cloud Adoption
-
Benefits:
-
Cost Savings: Shift from Capital Expenditure (CapEx) to Operational Expenditure (OpEx). Pay only for what you use.
-
Scalability & Elasticity: Instantly scale resources up/down with demand.
-
Agility & Speed: Rapid provisioning (minutes vs. weeks/months) accelerates time-to-market.
-
Focus on Core Business: Offload infrastructure management to provider.
-
Global Reach: Deploy applications in multiple regions with low latency.
-
-
Challenges & Risks:
-
Vendor Lock-in: Difficulty migrating due to proprietary APIs, services, and data egress costs.
-
Data Privacy & Sovereignty: Data location laws (GDPR, data residency requirements).
-
Compliance: Ensuring provider meets industry-specific regulations (HIPAA, PCI-DSS).
-
Downtime & Availability: Dependency on provider's SLA; outages affect your business.
-
Hidden Costs: Data transfer (egress) fees, premium support, over-provisioning.
-
Change Management: Organizational resistance, skill gaps (need cloud/DevOps skills).
-
6.3 Ecosystem Players: Independent Software Vendors (ISVs)
-
Role: Companies that develop and sell software applications that run on various platforms.
-
In Cloud Context:
-
"Lift-and-Shift" ISVs: Package existing on-premise applications for cloud deployment (e.g., as VMs in marketplace).
-
Cloud-Native ISVs: Build applications specifically for cloud platforms, leveraging cloud services (e.g., databases, messaging).
-
Marketplace Publishers: List their software in cloud provider marketplaces (AWS Marketplace, Azure Marketplace) for easy deployment.
-
Partner with Providers: ISVs often get technical and go-to-market support from cloud providers to optimize their apps for the cloud.
-
6.4 Supporting Concepts
-
Utility Computing:
-
Definition: A business model where computing resources (processing, storage) are provided as a metered, on-demand service, analogous to traditional utilities (electricity, water).
-
Key Enablers: Virtualization (for multi-tenancy), broadband networking, standardized hardware.
-
Metering & Billing: Precise measurement of resource consumption (CPU-hours, GB-months) for pay-per-use billing.
-
-
Quality of Service (QoS):
-
Definition: The overall performance and reliability of a cloud service, typically defined and guaranteed in a Service Level Agreement (SLA).
-
Key Issues (SLA Metrics):
-
Availability: Uptime percentage (e.g., 99.9% = ~8.76 hours downtime/year). \boxed{\text{Availability} = \frac{\text{Uptime}}{\text{Total Time}} \times 100%}
-
Performance: Throughput (requests/sec), latency (response time), transaction rates.
-
Reliability: Mean Time Between Failures (MTBF), Mean Time To Repair (MTTR).
-
Capacity: Guaranteed minimum resources (vCPUs, RAM, IOPS).
-
Support: Response time for technical issues.
-
-
Management: Cloud providers use resource scheduling, load balancing, and redundancy to meet SLAs. Customers must architect applications for resilience (multi-AZ, multi-region).
-
7.0 Performance, Management & Analytics
7.1 Cloud Performance & Benchmarks
-
Importance: Enables objective comparison of cloud providers/services, capacity planning, and cost-performance optimization.
-
Key Performance Indicators (KPIs):
-
Compute: CPU utilization (%), instructions per cycle (IPC), benchmark scores (e.g., SPECint).
-
Storage: IOPS (Input/Output Operations Per Second), throughput (MB/s), latency (ms), durability (e.g., 99.999999999% - 11 nines).
-
Network: Bandwidth (Gbps), latency (ms), jitter, packet loss.
-
Overall: Application response time, throughput (transactions/sec), cost per transaction.
-
7.2 Data Management & Analytics in Cloud
-
Online Analytical Processing (OLAP):
-
Functionality: Technology for complex, ad-hoc analytical queries against large, historical, aggregated datasets (data warehouses).
-
Core Operations:
-
Roll-up (Drill-up): Summarize data by climbing up a hierarchy (e.g., daily sales → monthly sales).
-
Drill-down: Navigate from summary to detailed data (e.g., yearly sales → quarterly → monthly).
-
Slice-and-Dice: Select a subset of data (a "slice") and view it from different perspectives ("dice" - changing dimensions).
-
Pivot (Rotate): Reorient the multidimensional view (swap rows and columns in a report).
-
-
Cloud Role: Cloud provides scalable, on-demand data warehousing (e.g., Amazon Redshift, Google BigQuery, Snowflake) and OLAP engines (e.g., Apache Kylin, Druid). Eliminates need for large upfront hardware investment.
-
7.3 Storage Cloud Specifics
-
Storage Service Types:
-
Object Storage: (e.g., AWS S3, Azure Blob). Stores data as objects with metadata. Highly scalable, durable, accessed via HTTP/API. Use: Static assets, backups, data lakes.
-
Block Storage: (e.g., AWS EBS, Azure Disks). Raw block devices (like virtual hard drives) attached to VMs. Low latency. Use: VM boot volumes, databases.
-
File Storage: (e.g., AWS EFS, Azure Files). Network-attached file shares (NFS/SMB). Shared access across multiple VMs. Use: Content repositories, home directories, lift-and-shift apps.
-
-
Key Characteristics:
-
Durability: Probability of not losing an object over a year (e.g., 99.999999999% - 11 nines for S3).
-
Availability: Probability of being able to access data (e.g., 99.9% for S3 Standard).
-
Pricing Models: Based on storage capacity (GB/month), requests (GET/PUT), data transfer (egress), and storage class (Standard, Infrequent Access, Glacier).
-