Skip to content
IT-702 (B) · Cloud Computing/Quick Revision Short Notes

Cloud Computing (IT-702 (B)) - Unit 4 Short Notes

UNIT 4: Cloud Computing

1. Cloud Service Models

Platform as a Service (PaaS)

Definition: A cloud computing model that provides a platform allowing customers to develop, run, and manage applications without the complexity of building and maintaining the underlying infrastructure (hardware, OS, middleware).

Essential Characteristics:

  • Integrated Development Environment (IDE): Built-in tools for coding, testing, and debugging.

  • Built-in Scalability & High Availability: Automatic scaling of applications based on demand.

  • Managed Middleware & Runtime: The provider manages OS patches, database updates, and runtime environments (e.g., Java, .NET, Python).

  • Multi-Tenant Architecture: A single application instance serves multiple users (tenants) with configuration isolation.

  • Service Integration: Seamless connectivity to other cloud services (databases, messaging queues, caching).

  • Pay-per-Use: Billing based on resource consumption (CPU, memory, storage, network I/O).

Problem Domains & Use Cases:

  • Rapid Application Development & Deployment: Ideal for agile teams needing to launch web/mobile apps quickly.

  • DevOps & CI/CD Pipelines: Streamlines build, test, and deployment automation.

  • API & Microservices Development: Provides a managed environment for building and deploying stateless services.

  • Collaborative Development: Multiple developers can work on a shared platform with consistent environments.

  • Examples: Google App Engine, Microsoft Azure App Service, Heroku, AWS Elastic Beanstalk.

[!TIP] Exam Focus: Distinguish PaaS from IaaS (you manage OS/middleware) and SaaS (you use a complete application). PaaS is for developers.


2. Cloud Deployment Models

Community Cloud

Definition: A cloud infrastructure shared by several organizations with shared concerns (e.g., mission, security requirements, policy, compliance considerations). It may be managed by the organizations or a third party and may exist on-premises or off-premises.

Key Features:

  • Shared by a Specific Community: Users are a defined group (e.g., government agencies, universities, healthcare providers).

  • Common Requirements: Addresses specific regulatory, security, or policy needs (e.g., HIPAA for healthcare, FISMA for government).

  • Cost Sharing: Infrastructure costs are distributed among community members, making it more economical than a private cloud for each.

  • Controlled Access: Access is restricted to the member organizations.

  • Governance: Often governed by a consortium or steering committee representing the community.

Comparison: Community Cloud vs. Public Cloud

Feature Community Cloud Public Cloud
Ownership & Management Owned/managed by community consortium or a third-party for the community. Owned/managed by a public cloud provider (AWS, Azure, GCP).
Accessibility Restricted to specific, pre-approved organizations/community members. Open for general public use; anyone can subscribe.
Resource Sharing Shared among a limited, known set of tenants with common goals. Shared among a vast, unknown, and dynamic pool of tenants (multi-tenancy).
Cost Model Shared capital and operational expenses among members. Potential for lower cost than private cloud. Pure operational expense (OpEx) pay-as-you-go model. Economies of scale.
Security & Compliance Can be tailored to meet specific community compliance standards (e.g., government-grade). Generic security controls; compliance is the user's responsibility (Shared Responsibility Model).
Control & Customization Higher degree of control and customization for the community's needs. Limited control over underlying infrastructure; standardized services.
Scalability Scalable within the limits of the community's shared infrastructure. Virtually infinite, on-demand scalability.

3. Cloud Platforms and Services

Google App Engine (GAE)

Definition: A fully managed Platform-as-a-Service (PaaS) offering from Google Cloud Platform (GCP) for building and hosting web applications and services in Google's data centers.

Major Features & Capabilities:

  • Fully Managed Runtime Environments: Supports standard runtimes (Java, Python, Node.js, Go, PHP, .NET) and custom runtimes.

  • Automatic Scaling: Instantly scales applications in response to traffic, from zero to millions of requests.

  • Built-in Services: Integrated NoSQL datastore (Firestore in Datastore mode), Memcache, Task Queues, and Users API.

  • Traffic Splitting & Versioning: Deploy multiple versions and split traffic for A/B testing or gradual rollouts.

  • Microservices Friendly: Ideal for deploying containerized applications using Google Kubernetes Engine (GKE) or Cloud Run (serverless containers).

  • Free Quota & Pay-per-Use: Generous free daily usage quotas; billing only for resources consumed beyond free tier.

Types of Applications & Problems Solvable:

  • Web Applications & APIs: Traditional web apps, RESTful APIs, backend for mobile apps (BFF).

  • Event-Driven & Asynchronous Workloads: Using Task Queues and Cron services for background processing.

  • Data-Intensive Applications: Leveraging integrated datastore and BigQuery for analytics.

  • Microservices & Serverless Architectures: Using Cloud Run or GAE flexible environment for containerized microservices.

Cloud-Specific Features & Development Workflow:

  1. Write Code: Develop application using standard libraries and GAE-specific APIs (e.g., for datastore).

  2. Configure: Define application settings, scaling parameters, and resource requirements in app.yaml (standard) or Dockerfile (flexible).

  3. Deploy: Use gcloud CLI or Cloud Console to deploy. GAE provisions instances, load balancers, and networking automatically.

  4. Manage & Monitor: Use Cloud Console for monitoring (Stackdriver), version management, traffic splitting, and scaling adjustments.

[!TIP] Exam Focus: GAE is a classic PaaS example. Know its auto-scaling and managed runtime features. Contrast with IaaS (like Compute Engine).

Eucalyptus

Definition: An open-source software framework for building AWS-compatible private and hybrid clouds. (Eucalyptus = Elastic Utility Computing Architecture for Linking Your Programs To Useful Systems).

Architectural Features & Components:

  • AWS Compatibility: Provides interfaces (EC2, S3, IAM) compatible with AWS, enabling "cloud bursting" to AWS.

  • Components:

    • CLC (Cloud Controller): The front-end; manages overall cloud, user authentication, and API requests.

    • CC (Cluster Controller): Manages a cluster of nodes; schedules VM instances, manages networking.

    • NC (Node Controller): Runs on each physical host; manages VM instances, networking, and storage.

    • SC (Storage Controller): Provides block storage (EBS-like) and object storage (S3-like) services.

    • Walrus: The object storage service (S3-compatible).

  • Deployment Flexibility: Can be deployed in different modes.

Modes of Operation:

  • Standalone Mode: All components (CLC, CC, SC, NC) run on a single machine. Used for testing, development, and small deployments.

  • Cluster Mode: CLC, CC, and SC run on separate machines; NCs run on multiple compute nodes. This is the typical production deployment for a private cloud within a single administrative domain.

  • Hybrid Mode: Eucalyptus cloud is connected to a public cloud (like AWS), allowing workloads to burst from private to public cloud.

Overview of Cloud Computing Platforms [Short Note]

Comparison of Major Platforms (AWS, Azure, GCP):

Feature Amazon Web Services (AWS) Microsoft Azure Google Cloud Platform (GCP)
Market Position Market Leader, Broadest & deepest service portfolio. Strong Enterprise Integration (MS ecosystem), Hybrid Cloud focus. Strong in Data/AI/ML, Kubernetes, Networking, and Cost-effectiveness.
Core Compute EC2 (VMs), Lambda (serverless), ECS/EKS (containers). Virtual Machines, Azure Functions, AKS (Kubernetes). Compute Engine (VMs), Cloud Functions, GKE (Kubernetes - origin).
Key Storage S3 (object), EBS (block), EFS (file). Blob Storage, Disk Storage, Files Storage. Cloud Storage (object), Persistent Disk (block), Filestore (file).
Database RDS (relational), DynamoDB (NoSQL), Redshift (data warehouse). SQL Database, Cosmos DB (NoSQL), Synapse (data warehouse). Cloud SQL, Firestore (NoSQL), BigQuery (serverless data warehouse).
Networking VPC, CloudFront (CDN), Route 53 (DNS). Virtual Network, Azure CDN, Azure DNS. VPC, Cloud CDN, Cloud DNS.
Pricing Model Complex, per-second billing (Linux), Savings Plans/Reserved Instances. Per-minute billing, Reserved VM Instances, Hybrid Benefit. Per-second billing, Sustained Use Discounts, Committed Use Discounts.
Unique Strength Unmatched service variety & maturity. Seamless integration with Windows Server, Active Directory, .NET. Leadership in Kubernetes, Big Data (BigQuery), and Machine Learning (TensorFlow).

4. Cloud Infrastructure

Virtualization in Cloud

Definition: The creation of a virtual (rather than actual) version of something, including virtual hardware platforms, storage devices, and network resources. It is the foundational technology that enables cloud computing.

Implementation in Microsoft Azure:

  • Hypervisor: Azure uses a customized version of Microsoft Hyper-V as its Type-1 (bare-metal) hypervisor.

  • Isolation: Each customer's virtual machines (VMs) run in isolated partitions on the same physical host, managed by the hypervisor.

  • Resource Management: The hypervisor directly controls CPU, memory, and device access for each VM, ensuring fair sharing and preventing "noisy neighbor" issues.

  • Integration: Azure's virtualization stack is deeply integrated with its software-defined networking (SDN) and storage solutions for automated provisioning and management.

Hypervisor Types & Role:

Type Description Examples Role in Cloud
Type 1 (Bare-Metal) Runs directly on the physical hardware of the host machine. VMware ESXi, Microsoft Hyper-V, Xen, KVM. Primary in public clouds. Provides highest performance, security, and efficiency. Manages all VMs on a host.
Type 2 (Hosted) Runs as a software layer on top of a conventional operating system. VMware Workstation, Oracle VirtualBox, Parallels. Used for development, testing, and desktop virtualization. Not suitable for large-scale cloud production due to OS dependency overhead.

Role in Cloud Infrastructure:

  1. Server Consolidation: Multiple VMs on one physical server, improving utilization.

  2. Isolation & Security: Strong isolation between tenants (multi-tenancy).

  3. Rapid Provisioning & Mobility: VMs can be created, cloned, and migrated (live migration) quickly.

  4. Hardware Abstraction: Applications see virtual hardware, making them portable across underlying physical servers.

  5. Enabler for IaaS: The core technology behind IaaS offerings like AWS EC2, Azure VMs.

Storage Solutions

Storage Cloud: Concepts, Types, and Service Models

  • Concept: Delivery of storage capacity and data management services over the internet on a pay-per-use basis.

  • Primary Service Models:

    • Object Storage: Stores data as objects (blobs) with metadata and a unique ID. Highly scalable, durable, and accessible via HTTP/HTTPS APIs. Use Case: Unstructured data (images, videos, backups, logs). Examples: AWS S3, Azure Blob Storage, GCP Cloud Storage.

    • Block Storage: Provides raw block devices (virtual disks) that can be formatted with a filesystem. High performance, low latency. Use Case: VM boot volumes, databases, enterprise applications. Examples: AWS EBS, Azure Managed Disks, GCP Persistent Disk.

    • File Storage: Provides a hierarchical file system accessible via standard protocols (NFS, SMB/CIFS). Use Case: Shared file systems for legacy apps, content management, home directories. Examples: AWS EFS, Azure Files, GCP Filestore.

Storage Area Network (SAN) in Cloud Environments [Short Note]

  • Definition: A dedicated high-speed network that provides access to consolidated, block-level storage. In the cloud, this is typically offered as a cloud-based SAN service or virtual SAN (vSAN).

  • Cloud Implementation: Instead of a physical Fibre Channel fabric, cloud providers use software-defined storage over high-speed Ethernet (often RDMA over Converged Ethernet - RoCE) to create a virtual SAN.

  • Characteristics: Provides high performance, low latency, and block storage access with features like snapshots, cloning, and replication. It is the underlying technology for cloud block storage services (EBS, Persistent Disk).

  • Key Difference from On-Prem SAN: No need to purchase/manage dedicated hardware switches and disk arrays. It's a fully managed, elastic service.


5. Big Data and Analytics in Cloud

MapReduce Programming Model

Concept: A programming model and associated implementation for processing and generating large data sets with a parallel, distributed algorithm on a cluster.

Workflow:

  1. Input Splitting: Input data is divided into independent splits.

  2. Map Phase: User-defined map() function processes each split, emitting intermediate key-value pairs.

  3. Shuffle & Sort: Framework automatically groups all values associated with the same intermediate key and sorts them.

  4. Reduce Phase: User-defined reduce() function processes each grouped key and its list of values to generate the final output.

Example: Word Count

  • Map Function: For each word in a line, emit (word, 1).

    • Input: "the quick brown fox"

    • Output: ("the",1), ("quick",1), ("brown",1), ("fox",1)

  • Shuffle/Sort: Groups all values for the same word.

    • ("the", [1,1,1]), ("quick", [1]), ("fox", [1,1])
  • Reduce Function: Sum the list of values for each key.

    • Output: ("the", 3), ("quick", 1), ("fox", 2)

Types and Formats of MapReduce Implementations:

  • Hadoop MapReduce (Classic): Java-based, disk-based (HDFS), batch processing. High latency.

  • Apache Spark: In-memory data processing engine. Can run MapReduce jobs but uses Resilient Distributed Datasets (RDDs) for much faster iterative and interactive processing. Supports Java, Scala, Python, R.

  • Other: Amazon EMR (managed Hadoop/Spark), Google Cloud Dataproc, Apache Flink (streaming-focused).

Cloud Analytics

Definition: The practice of using cloud-based platforms and tools to analyze structured and unstructured data to extract valuable insights, support decision-making, and predict trends.

Significance:

  • Scalability on Demand: Handle petabytes of data without upfront hardware investment.

  • Cost-Effectiveness: Pay only for the computing and storage resources used during analysis.

  • Managed Services: No need to install/maintain complex analytics software (e.g., managed Hadoop/Spark, data warehouses).

  • Integration: Seamless integration with various data sources (cloud storage, IoT streams, databases).

  • Advanced Capabilities: Built-in Machine Learning and AI services (e.g., Amazon SageMaker, Azure ML, GCP AI Platform).

Real-World Applications & Tools:

  • Business Intelligence (BI): Cloud-based dashboards and reporting (Tableau Online, Power BI, Looker).

  • Data Warehousing: Scalable, serverless warehouses (Amazon Redshift, Snowflake, Google BigQuery).

  • Stream Processing: Real-time analytics on data streams (Amazon Kinesis, Apache Kafka on Confluent Cloud, Google Pub/Sub + Dataflow).

  • Predictive Analytics & ML: Building and deploying predictive models (Azure Machine Learning Studio, Google Vertex AI).

  • Applications: Customer sentiment analysis, fraud detection, predictive maintenance, recommendation systems, log analysis.


6. Security in Cloud Computing

Security Challenges

Overview of Key Issues:

  1. Data Privacy & Sovereignty: Compliance with regulations (GDPR, HIPAA) regarding data location and processing.

  2. Multi-Tenancy Risks: Isolation failures could lead to data leakage between tenants.

  3. Insecure APIs & Interfaces: Vulnerabilities in cloud service APIs can be exploited.

  4. Account Hijacking: Stolen credentials can lead to data theft, service abuse, or crypto-mining.

  5. Denial-of-Service (DoS) Attacks: Overwhelming cloud resources to cause service disruption.

  6. Insider Threats: Malicious or negligent actions by cloud provider or customer employees.

  7. Data Loss/Leakage: Accidental deletion, insecure storage, or transmission.

  8. Shared Technology Vulnerabilities: Hypervisor bugs, side-channel attacks (e.g., Spectre, Meltdown).

  9. Lack of Visibility & Control: Customers have limited insight into physical security and underlying infrastructure.

Cloud Security Architecture

Components & Layered Security Model (Defense-in-Depth):

A typical cloud security architecture implements controls across multiple layers:


[[DIAGRAM: CANVAS: A layered diagram from bottom to top:

1. Physical Security: Fences, guards, biometrics at data centers.

2. Network Security: Firewalls, IDS/IPS, DDoS protection, VPC segmentation, VPNs.

3. Host/Compute Security: Hypervisor hardening, VM/container security, patching, anti-malware.

4. Application Security: Secure coding, WAF (Web Application Firewall), API security, secrets management.

5. Data Security: Encryption (at rest, in transit), key management (KMS), data loss prevention (DLP), access controls (IAM).

6. Identity & Access Management (IAM): Central pillar. Authentication (MFA), Authorization (least privilege, roles), Federation.

7. Monitoring & Logging: Centralized logging (CloudTrail, Azure Monitor), SIEM integration, audit trails, anomaly detection.

]]

Reference Control Mechanisms:

  • Encryption: TLS/SSL for transit, AES-256 for data at rest (with customer-managed or provider-managed keys).

  • Identity and Access Management (IAM): Granular policies defining who can access what resources and under which conditions.

  • Network Security Groups (NSGs) & Firewalls: Control inbound/outbound traffic at VM/subnet level.

  • Security Groups & NACLs: (AWS/Azure specific) Stateless/stateful packet filtering.

  • Logging & Auditing: Immutable logs of all API activities and access attempts.

  • Shared Responsibility Model: Critical Concept. Security is a shared duty:

    • Cloud Provider: Security of the cloud (physical infrastructure, hypervisor, global network).

    • Customer: Security in the cloud (OS patching, application security, data encryption, IAM configuration, network security).

Trusted Cloud Computing [Short Note]

Definition: A set of technologies and processes designed to ensure that cloud services and infrastructure are trustworthy, reliable, and secure, providing assurance to customers about the integrity and confidentiality of their data and operations.

Principles:

  • Trustworthiness: Evidence-based assurance of security, reliability, and compliance.

  • Transparency: Clear disclosure of security practices, audit reports (SOC 2, ISO 27001), and data handling policies.

  • Accountability: Clear delineation of responsibilities (Shared Responsibility Model).

  • Auditability: Comprehensive logging and monitoring to enable forensic analysis.

Technologies & Implementation Approaches:

  • Hardware Root of Trust: Using TPM (Trusted Platform Module) or Intel SGX/AMD SEV to ensure platform integrity and secure key storage.

  • Attestation: Process where a cloud component (VM, host) provides cryptographic proof of its software/hardware state to a verifier (customer or another service).

  • Secure Multi-Tenancy: Strong isolation via virtualization (hypervisor), containers, and network segmentation (VLANs, VXLANs).

  • Confidential Computing: Encryption of data in use (during processing) using hardware-enforced trusted execution environments (TEEs) like Intel SGX or AMD SEV.

  • Homomorphic Encryption: Allows computation on encrypted data without decryption (still emerging).

  • Implementation: Achieved through a combination of provider-level controls (hardware, hypervisor), customer-level controls (encryption, IAM), and third-party audits/certifications.


7. Advanced Cloud Concepts

Quality of Service (QoS) in Cloud

Definition: The overall performance and reliability characteristics of a cloud service, often defined by measurable parameters and guaranteed through Service Level Agreements (SLAs).

QoS Parameters:

  • Availability/Uptime: % of time service is operational (e.g., 99.9%).

  • Latency: Time taken for a request to get a response.

  • Throughput: Amount of work processed per unit time (requests/sec, GB/sec).

  • Reliability: Probability of failure-free operation over time.

  • Security: Level of protection against threats.

  • Performance Consistency: Predictability of performance (low jitter).

QoS Issues & Challenges:

  • Noisy Neighbor: One tenant's resource-intensive VM degrades performance for others on the same host.

  • Resource Contention: High demand for shared resources (network bandwidth, storage I/O).

  • Multi-Tier Dependencies: Latency in one service (e.g., database) affects the entire application chain.

  • SLA Violations: Failure to meet promised metrics, leading to penalties.

  • Dynamic Workloads: Difficulty in guaranteeing QoS for unpredictable, spiky workloads.

Example & Solution:

  • Challenge: A critical database VM experiences high I/O wait times because a co-located VM is running a batch analytics job.

  • Solutions:

    1. Dedicated Instances/Hosts: Use dedicated VMs or physical hosts for the critical workload.

    2. QoS Policies: Configure storage QoS (IOPS/throughput limits) on the noisy VM's disk.

    3. Placement Groups/Affinity Rules: Place critical VMs on separate physical hardware.

    4. Monitoring & Auto-Scaling: Monitor metrics and scale out the database tier or move workloads.

Elastic Computing [Short Note]

Definition: The ability to automatically and dynamically scale computing resources (up or down) in response to real-time demand, while maintaining optimal performance and cost.

Concept & Scalability:

  • Vertical Scaling (Scale Up/Down): Increase or decrease the power (CPU, RAM) of a single resource (e.g., VM size). Has limits and may require downtime.

  • Horizontal Scaling (Scale Out/In): Increase or decrease the number of resource instances (e.g., VMs, containers). The preferred cloud-native model for stateless applications.

Auto-Scaling Mechanisms:

  1. Metric-Based Scaling: Trigger scaling actions based on cloud monitoring metrics (CPU utilization > 70%, network in, queue length).

  2. Schedule-Based Scaling: Scale based on known traffic patterns (e.g., scale out at 9 AM for business hours).

  3. Predictive Scaling: Use ML to forecast demand and scale proactively (e.g., AWS Predictive Scaling).

  4. Event-Driven Scaling: Scale in response to specific events (e.g., new message in a queue triggers a new worker instance).

Implementation: Configured via cloud provider tools (AWS Auto Scaling Groups, Azure Scale Sets, GCP Instance Groups). Defines min/max/desired capacity and scaling policies.

Cloud for Social Networking Applications

Advantages:

  • Elastic Scalability: Handle massive, unpredictable traffic spikes (viral events) without service degradation.

  • Global Reach: Deploy application instances in multiple geographic regions using CDNs for low latency.

  • Cost-Effective: Pay for infrastructure only during active use; no over-provisioning for peak loads.

  • Managed Services: Offload infrastructure management to focus on core features (user engagement, content).

  • Rapid Feature Deployment: CI/CD pipelines enable frequent updates and A/B testing.

  • Built-in Data Services: Leverage cloud databases (NoSQL for feeds), analytics (user behavior), and media processing (video/image transcoding).

Specific Use Cases:

  • User-Generated Content (UGC): Storing and serving billions of photos/videos via object storage (S3) and CDNs (CloudFront).

  • Real-Time Feeds & Notifications: Using pub/sub systems (Kafka, Pub/Sub) and in-memory data stores (Redis) for real-time updates.

  • Graph Data & Relationships: Using graph databases (Neo4j, Amazon Neptune) for friend recommendations and social graph analysis.

  • Big Data Analytics: Processing petabytes of clickstream and interaction data to derive insights (using Hadoop/Spark on EMR/Dataproc).

  • Mobile Backend: Providing scalable authentication, push notifications, and API gateways for mobile clients.

Cloud Economics

Reduction in Time-to-Market for Applications:

  • On-Demand Resources: Instantly provision development, test, and production environments.

  • Managed Services: Eliminate setup time for databases, messaging queues, and caches.

  • DevOps & Automation: Infrastructure as Code (IaC) enables repeatable, fast environment creation.

  • Global Deployment: Deploy to multiple regions with a few clicks, accelerating international launches.

  • Focus on Innovation: Developers spend less time on infrastructure maintenance and more on coding features.

Capital Expense (CapEx) vs. Operational Expense (OpEx) Savings:

Aspect Traditional On-Premises (CapEx) Cloud Computing (OpEx)
Cost Nature Large upfront investment in hardware, data center, networking. Ongoing, variable operational costs based on usage.
Financial Model Capital Expenditure (purchase of assets). Operational Expenditure (rental/service fee).
Budgeting Requires long-term planning, procurement cycles, and budget approval. Flexible, aligns with actual consumption; easier to scale budgets up/down.
Asset Management Company owns depreciating assets. Responsibility for maintenance, upgrades, disposal. No asset ownership. Provider handles maintenance and upgrades.
Risk High risk of over-provisioning (wasted capacity) or under-provisioning (poor performance). Low risk; pay only for what you use. Easy to experiment and fail fast.
Example Buying a $$\displaystyle 500,000 server rack that will be obsolete in 3 years. | Paying $$500/month for a cluster of VMs that can be resized or shut down anytime.

Core Economic Benefit: Shifts IT from a cost center with fixed, sunk costs to an agile enabler with variable, usage-based costs, directly linking IT spend to business activity and innovation.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in