Skip to content
IT-702 (B) · Cloud Computing/Quick Revision Short Notes

Cloud Computing (IT-702 (B)) - Unit 2 Short Notes

UNIT 2: Cloud Computing and Big Data


1. Fundamentals of Cloud Computing

Definition: A model for enabling ubiquitous, convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, services) that can be rapidly provisioned and released with minimal management effort.

Essential Characteristics (NIST Definition):

  1. On-demand self-service: Provision resources automatically without human interaction.

  2. Broad network access: Available over the network via standard mechanisms.

  3. Resource pooling: Multi-tenant model with physical/virtual resources dynamically assigned.

  4. Rapid elasticity: Capabilities can be elastically provisioned and released.

  5. Measured service: Resource usage monitored, controlled, and reported.

Economic Benefits:

  • Reducing Time-to-Market: Rapid provisioning accelerates development and deployment cycles.

  • Cutting Capital Expenses (CapEx): Shifts to operational expenses (OpEx); no upfront hardware costs. Pay-per-use model.

Cloud Computing Models Overview:

Model Type Examples What it Provides
Service Models IaaS, PaaS, SaaS Level of abstraction/management responsibility.
Deployment Models Public, Private, Hybrid, Community Ownership, location, and access model of the cloud infrastructure.

2. Cloud Service Models

Infrastructure as a Service (IaaS)

Provides fundamental computing resources: processing, storage, networks. User has control over OS, storage, deployed applications. Example: AWS EC2, Azure Virtual Machines.

Platform as a Service (PaaS)

Essential Characteristics:

  • Provides a platform (runtime, middleware, OS) for application development, testing, deployment.

  • User controls: Application and its configuration.

  • Provider controls: Underlying infrastructure (OS, servers, storage, networking).

  • Problem-Solving: Eliminates infrastructure management, enables scalability, supports automated deployment pipelines. Example: Google App Engine, Heroku, Azure App Service.

Software as a Service (SaaS)

Delivers complete, ready-to-use applications over the internet, typically via a web browser. Example: Gmail, Salesforce, Office 365.

[!TIP] Exam Focus: PaaS is frequently asked. Be ready to contrast IaaS/PaaS/SaaS using the shared responsibility model (who manages what).


3. Cloud Deployment Models

Model Ownership/Operation Access Key Use Case
Public Cloud Third-party provider (AWS, Azure, GCP) General public Cost-effective, scalable, no maintenance.
Private Cloud Single organization (on/off-premise) Exclusive High control, security, compliance for sensitive data.
Community Cloud Shared by several organizations with common concerns (security, compliance) Specific community Joint ventures, government agencies, healthcare consortiums.
Hybrid Cloud Composition of two or more clouds (private+public) Bound by proprietary standards Orchestration, bursting, data/application portability.

Community Cloud vs. Public Cloud:

  • Community: Shared infrastructure for a specific group with shared requirements (e.g., law firms sharing legal data). More control and tailored compliance than public.

  • Public: Open to all; standardized, multi-tenant, largest scale and cost benefits.


4. Virtualization in Cloud Environments

Concept: Creation of virtual (rather than actual) versions of resources—OS, servers, storage, networks—to maximize resource utilization and enable isolation.

Types:

  • Server Virtualization: Multiple OS on single physical server (e.g., VMware ESXi, Microsoft Hyper-V).

  • Storage Virtualization: Pooling physical storage from multiple devices.

  • Network Virtualization: Splitting bandwidth into independent channels.

Implementation in Microsoft Azure:

  • Hyper-V: Primary hypervisor for Azure VMs.

  • Azure Virtual Machines (IaaS): Each VM is an isolated instance running on Hyper-V.

  • Azure Virtual Networks: Software-defined networking (SDN) for isolation and connectivity.

  • Role: Foundational for multi-tenancy, resource isolation, live migration, and dynamic resource allocation in Azure.

[!TIP] Common Pitfall: Virtualization is an enabling technology for cloud, not cloud itself. Cloud adds on-demand self-service, elasticity, and measured service.


5. Cloud Storage Solutions

Storage Cloud: A cloud service model where data is stored on remote storage systems accessed via the internet (APIs or web interfaces). Features include durability, scalability, pay-per-use, and global accessibility.

Storage Area Network (SAN) in Cloud:

  • A dedicated high-speed network that connects servers to consolidated block-level storage.

  • Cloud Context: Often virtualized and offered as a service (e.g., Azure Managed Disks, AWS EBS). Provides high performance, low latency storage for VMs and critical applications.

Cloud Storage Architectures:

  • Object Storage: (e.g., AWS S3, Azure Blob) - Unstructured data, massive scale, HTTP/HTTPS access.

  • Block Storage: (e.g., AWS EBS) - Raw storage volumes for VMs, file systems.

  • File Storage: (e.g., Azure Files) - Shared file system accessible via SMB/NFS.


6. Big Data Fundamentals

Definition: Data sets whose size, velocity, or structure challenge traditional data processing systems, requiring new paradigms for capture, storage, analysis, and visualization.

The 3 Vs:

  1. Volume: Scale of data (Terabytes to Zettabytes). Example: Social media feeds, sensor data.

  2. Variety: Different forms (structured, semi-structured, unstructured). Example: Text, images, video, log files.

  3. Velocity: Speed of data generation and processing needs. Example: Real-time fraud detection, stock tickers.

Challenges:

  • Storage and processing at scale.

  • Data heterogeneity and integration.

  • Ensuring data quality and veracity.

  • Real-time analysis requirements.

  • Security and privacy.

Big Data Analytics: The process of examining large, varied data sets to uncover hidden patterns, correlations, and insights.

  • Real-world Applications: Predictive maintenance (manufacturing), personalized recommendations (e-commerce), genomic analysis (healthcare), traffic optimization (smart cities).

Data Preprocessing:

  • Cleaning: Handling missing values, noise removal, outlier detection, consistency checks.

  • Sampling: Selecting a representative subset of data for analysis when full dataset is impractical.

    • Techniques: Simple random sampling, stratified sampling, reservoir sampling (for streaming data).

7. Data Processing Frameworks

Hadoop Architecture (Building Blocks)

DiagramCANVAS: Draw a diagram showing a master node (NameNode, ResourceManager) connected to multiple slave/worker nodes. Each worker node runs a DataNode and NodeManager. HDFS blocks are distributed across DataNodes. YARN manages resource allocation. MapReduce tasks run on worker nodes. Label all components.
  • HDFS (Hadoop Distributed File System): Master-Slave architecture. NameNode (master, metadata) + DataNodes (slaves, store blocks).

  • YARN (Yet Another Resource Negotiator): Cluster resource management. ResourceManager (scheduler) + NodeManager (per-node agent).

  • MapReduce: Programming model for processing.

MapReduce

Concept: Programming model for processing large datasets in parallel across a cluster. Two main functions:

  1. Map: Processes input key-value pairs, emits intermediate key-value pairs.

  2. Reduce: Aggregates intermediate values for each key.

Types & Formats:

  • Input Formats: TextInputFormat (default, line by line), KeyValueTextInputFormat, SequenceFileInputFormat.

  • Output Formats: TextOutputFormat (default), SequenceFileOutputFormat.

  • Example (Word Count):

    
    // Map: (line_offset, line_text) -> (word, 1)
    
    // Reduce: (word, [1,1,1...]) -> (word, sum)
    
    

Apache Hive

Architecture:

DiagramCANVAS: Show a user submitting a HiveQL query via CLI/UI. Query goes to Driver (compiles, optimizes). Driver sends plan to Metastore (gets metadata) and to Execution Engine (compiler). Execution Engine uses Hadoop (MapReduce/Tez/Spark) to process data from HDFS. Result returned to user.
  • Components: Metastore (schema/relational data), Driver (compiler, optimizer), Execution Engine (runs tasks), HiveQL (SQL-like query language).

User-Defined Functions (UDFs):

  • Custom functions written in Java to process Hive data.

  • Procedure:

    1. Extend org.apache.hadoop.hive.ql.exec.UDF class.

    2. Implement evaluate() method with custom logic.

    3. Package into JAR.

    4. Add JAR to Hive session (ADD JAR).

    5. Create temporary/permanent function (CREATE TEMPORARY FUNCTION).

Apache Pig

Architecture:

DiagramCANVAS: Pig Latin script -> Parser -> Logical Plan -> Logical Optimizer -> Physical Plan -> Physical Optimizer -> MapReduce Plan -> Execution on Hadoop cluster. Show interaction with HDFS.
  • Pig Latin: High-level scripting language for data transformation.

  • Application Flow: Script → Parser → Logical Plan → Logical Optimizer → Physical Plan → Physical Optimizer → MapReduce Plan → Execution on Hadoop.


8. Data Analytics and Mining

R Programming Language

Features:

  • Open-source, statistical computing and graphics.

  • Extensive packages (ggplot2, dplyr, caret).

  • Powerful data handling and storage.

  • Array-oriented, vectorized operations.

  • Cross-platform.

Major Components:

  • Base R: Core language and basic stats.

  • R Packages: Extend functionality (CRAN repository).

  • R Studio: Popular IDE.

  • R Graphics: Advanced plotting systems (base, lattice, ggplot2).

Operations on Vectors:

  • Creation: v <- c(1,2,3)

  • Arithmetic: Element-wise (v1 + v2).

  • Indexing: v[2] (second element), v[c(1,3)].

  • Filtering: v[v > 2].

  • Sorting: sort(v).

  • Aggregation: sum(v), mean(v), sd(v).

  • Vectorized Functions: log(v), sqrt(v).

Classification Algorithms

Decision Trees:

  • Tree-like model: internal nodes = feature tests, branches = outcomes, leaf nodes = class labels.

  • Algorithms: ID3 (entropy), C4.5 (gain ratio), CART (Gini index).

  • Process: Recursive partitioning to maximize information gain/minimize impurity.

  • Pruning: To avoid overfitting (pre/post-pruning).

Naive Bayes Classification:

  • Based on Bayes' Theorem with "naive" assumption of feature independence.

  • Formula:

$$P(y|x_1,...,x_n) = \frac{P(y) \prod_{i=1}^{n} P(x_i|y)}{P(x_1,...,x_n)}$$

  • Prediction: Assign class y with highest posterior probability.

  • Types: Gaussian (continuous), Multinomial (discrete counts), Bernoulli (binary).

  • Advantages: Simple, fast, works well with high-dimensional data.

Association Rules

  • Discover interesting relationships (rules) between variables in large databases.

  • Rule: X => Y (if X, then Y). X = antecedent, Y = consequent.

  • Metrics:

    • Support: P(X ∪ Y) - frequency of rule.

    • Confidence: P(Y|X) = Support(X∪Y) / Support(X) - conditional probability.

    • Lift: Confidence / P(Y) - strength of rule over random chance. Lift > 1 indicates positive correlation.

  • Algorithm: Apriori (candidate generation & pruning).

  • Applications: Market basket analysis, cross-selling, recommendation systems, medical diagnosis.


9. Cloud Platforms and Technologies

Google App Engine (GAE)

Major Features:

  • PaaS: Supports multiple languages (Python, Java, Go, PHP, Node.js).

  • Automatic Scaling: Scales apps based on traffic.

  • Managed Services: Integrated with Google Cloud services (Datastore, BigQuery).

  • Sandboxed Environment: Applications run in secure, isolated containers.

Cloud Features:

  • High availability, built-in load balancing, versioning, traffic splitting.

  • No server management.

Problems Solvable:

  • Web applications, mobile backends, APIs.

  • Scalable data-driven apps without infrastructure overhead.

Eucalyptus

Features: Open-source software for building AWS-compatible private/hybrid clouds.

  • AWS Compatibility: Uses same APIs as EC2, S3, EBS, IAM.

  • Components: Cluster Controller (CC), Cloud Controller (CLC), Storage Controller (SC), Node Controller (NC).

Modes:

  • Managed Mode: Eucalyptus manages networking (requires VLANs).

  • System Mode: Uses existing network infrastructure (no VLANs), simpler setup.

  • Managed Mode (Eucalyptus 4+): Advanced networking with security groups, elastic IPs.


10. Cloud Security

Security Challenges:

  • Data breaches and loss.

  • Insecure APIs and interfaces.

  • Account hijacking.

  • Malware/vulnerabilities in shared technology.

  • Abuse of cloud resources (e.g., crypto-mining).

  • Insider threats.

  • Compliance and legal issues (data sovereignty).

Cloud Computing Security Architecture:

[[DIAGRAM: CANVAS: Draw layered architecture from bottom to top:

  1. Physical Security (data center)

  2. Hypervisor/Virtualization Security

  3. Network Security (firewalls, IDS/IPS, VLANs, VPNs)

  4. Host Security (OS hardening, patching)

  5. Application Security (secure coding, WAF)

  6. Data Security (encryption at rest/in transit, key management)

  7. Identity & Access Management (IAM, MFA, SSO)

  8. Governance, Risk & Compliance (GRC) layer on top.]]

  • Key: Defense-in-depth, shared responsibility model (provider secures infrastructure, customer secures data & access).

Trusted Cloud Computing: Concepts and technologies to ensure trustworthiness in cloud environments:

  • Trusted Computing Base (TCB): Minimal set of hardware/software critical to security.

  • Attestation: Verifying platform integrity (e.g., TPM, Intel TXT).

  • Secure/Trusted Execution Environments (TEE): Isolated processing areas (e.g., Intel SGX, ARM TrustZone).

  • Homomorphic Encryption: Compute on encrypted data without decryption.


11. Quality of Service (QoS) in Cloud

QoS Issues:

  • Performance Variability: "Noisy neighbor" problem in multi-tenant clouds.

  • Service Level Agreement (SLA) Violations: Not meeting promised metrics (uptime, response time).

  • Scalability Bottlenecks: Inadequate auto-scaling.

  • Network Latency/Jitter: Affects real-time apps.

Performance Metrics & Guarantees:

  • Availability: (Uptime / Total Time) * 100%. Common SLA: 99.9% ("three nines").

  • Response Time: Time from request to first byte/response.

  • Throughput: Requests processed per unit time.

  • Latency: Delay in data transmission.

  • Scalability: Ability to handle increased load.

  • Guarantees: Specified in SLAs with penalties (service credits) for violation. Requires monitoring and reporting.


12. Cloud Applications and Services

Cloud Technologies for Social Networking:

  • Advantages:

    • Elastic Scalability: Handle viral growth spikes (e.g., trending topics).

    • Global Reach: CDNs and distributed data centers reduce latency.

    • Cost-Effective: Pay for peak usage; no over-provisioning.

    • Rapid Feature Deployment: Continuous integration/ delivery.

    • Data Analytics: Process vast user interaction data for personalization.

Cloud Analytics as a Service (AaaS):

  • Delivery of analytics capabilities (data warehousing, BI, machine learning) as a cloud service.

  • Examples: AWS Redshift, Google BigQuery, Azure Synapse Analytics.

  • Benefits: No infrastructure setup, scalable compute/storage, managed services, pay-per-query.

Application Domains:

  • Enterprise Applications: CRM, ERP, collaboration (Office 365, Salesforce).

  • Data-Intensive Apps: Big data processing, scientific computing.

  • IoT: Device management, data ingestion, analytics.

  • Gaming: Scalable backends for multiplayer games.

  • Media & Entertainment: Content delivery, rendering farms.

  • Healthcare: HIPAA-compliant storage, telemedicine, genomics.

[!TIP] Exam Strategy: For 14m questions (like Community Cloud), use comparative tables. For 7m questions, structure as Definition → Key Features → Examples → Advantages/Challenges. Always link theory to real-world examples.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in