Skip to content
CY-703 (D) · Cloud Computing/Quick Revision Short Notes

Cloud Computing (CY-703 (D)) - Unit 5 Short Notes

UNIT 5: Data Engineering & Cloud Data Systems

I. Foundations of Data Engineering

Data Engineering is the discipline of designing, building, and maintaining the infrastructure and systems that enable the collection, storage, processing, and serving of data for analytical and operational use cases. It focuses on creating reliable, scalable, and efficient data pipelines.

Data Engineering Lifecycle (High Priority - Direct Question)

A cyclical process with six core stages:

  1. Data Generation: Creation of data from sources like applications, IoT sensors, logs, or user interactions.

  2. Data Ingestion: The process of acquiring and importing data for immediate use or storage in a data lake/warehouse. Example: Using Kafka to stream clickstream data.

  3. Data Storage: Persisting data in appropriate systems (Data Lakes, Warehouses, Marts) based on structure, volume, and access needs.

  4. Data Processing: Transforming raw data into a usable format. Includes batch processing (e.g., Spark) and stream processing (e.g., Flink).

  5. Data Serving: Making processed data available to end-users, applications, or BI tools via APIs, query engines, or dashboards.

  6. Data Consumption: The final stage where data scientists, analysts, or applications use the served data for insights, reporting, or ML models.

[!TIP] Exam Focus: Be prepared to draw and label this lifecycle diagram. For each stage, know a real-world cloud service example (e.g., Ingestion: AWS Kinesis, Storage: S3, Processing: Glue).


II. Enterprise Data Frameworks & Maturity Models

Gartner Data Maturity Model

A five-level model assessing an organization's data management sophistication:

Level Description Example
1. Unaware No formal data management; data is an IT by-product. Data siloed in department spreadsheets.
2. Emerging Awareness of need; ad-hoc, project-based efforts. A single team builds a standalone dashboard.
3. Developing Defined standards and some centralized governance. Enterprise data warehouse exists, but usage is inconsistent.
4. Established Data is a managed asset; well-defined processes and metrics. Company-wide data catalog, standardized SLAs for data delivery.
5. Optimized Data drives continuous optimization and innovation. Predictive analytics embedded in all core business processes.

Zachman Framework

A six-by-six schema for enterprise architecture, providing a holistic view from different perspectives (rows) and aspects (columns).

Perspective (Rows) What (Data) How (Function) Where (Network) Who (People) When (Time) Why (Motivation)
Planner (Scope) Inventory Process Inventory Distribution Inventory Role Inventory Event Inventory Goal Inventory
Owner (Business) Conceptual Data Model Business Process Model Business Logistics Organization Chart Master Schedule Business Plan
Designer (System) Logical Data Model Technical Process Model Distribution Logistics Human Interface Process Schedule Design Rationale
Builder (Technology) Physical Data Model Technology Process Technology Logistics Presentation Process Timing Design Constraints
Subcontractor (Details) Data Specification Process Specification Network Specification Organization Spec. Event/Timing Spec. Goal to Implementation
Enterprise (Functioning) Actual Database Running System Running Network Actual Organization Actual Schedule Actual Business Plan

[!TIP] Common Pitfall: Don't confuse the rows (perspectives) with the columns (fundamental questions). Remember the mnemonic "What, How, Where, Who, When, Why" across the six stakeholder perspectives.


III. Data Architecture Patterns (Very High Priority)

Traditional Data Architectures

  • Monolithic: All components (ingestion, storage, processing, serving) are tightly coupled in a single, often on-premise, system (e.g., legacy EDW).

    • Advantages: Simpler initial setup, ACID compliance.

    • Disadvantages: Scalability issues, single point of failure, technology lock-in.

  • Microservices-oriented: Decomposed into independent, loosely-coupled services (e.g., separate services for ingestion, cleansing, analytics).

    • Advantages: Scalable, resilient, technology agnostic.

    • Disadvantages: Increased complexity in orchestration and data consistency.

Lambda Architecture (High Priority - Direct Question)

A unified architecture for batch and real-time processing. It has three layers:

  1. Batch Layer (Speed Layer): Processes the entire historical dataset at rest (e.g., in HDFS/S3) to create accurate, comprehensive batch views (e.g., using Hadoop/Spark). Slow but accurate.

  2. Speed Layer (Hot Path): Processes only the new, real-time data stream to create a real-time view that is low-latency but may be less accurate (e.g., using Kafka/Flink).

  3. Serving Layer: Combines the batch view and real-time view to serve the final, merged result to queries.

Step-by-Step Real-Time Implementation:

  1. Ingest raw data into a master dataset (immutable data lake).
  1. Batch Layer: Periodically (e.g., hourly) recomputes a full aggregate view from the master dataset and overwrites the serving layer's batch table.
  1. Speed Layer: Continuously processes the incoming stream, updating a separate real-time view table in the serving layer with incremental updates.
  1. Serving Layer: A query merges results from the batch table and real-time table on-the-fly.

Kappa Architecture (High Priority - Direct Question)

A simplification of Lambda that treats all data as a stream. It eliminates the separate batch layer.

  • Core Idea: Ingest all data (historical and real-time) into a single, distributed, immutable log (e.g., Apache Kafka). All processing is done via stream processing.

  • To recompute historical data: Simply replay the retained log through the same stream processing job.

  • Comparison with Lambda:

    | Feature | Lambda Architecture | Kappa Architecture | | :--- | :--- | :--- | | Complexity | High (maintain two codebases) | Lower (single stream codebase) | | Historical Recompute | Easy (re-run batch job) | Possible but requires log retention & replay | | Latency | Batch: high, Stream: low | Uniformly low (stream-only) | | Use Case | Need both deep historical accuracy & low-latency streams. | Primarily low-latency stream processing with occasional full recompute. |

Data Lake Patterns (High Priority - Direct Question)

A centralized repository storing raw, unstructured, or semi-structured data in its native format until needed.

  • Characteristics:

    • Schema-on-Read: Schema is applied when data is read/queried, not when written. Allows flexibility.

    • Raw Storage: Stores data in its most granular form (e.g., JSON, CSV, Parquet in S3/ADLS).

    • Cost-Effective: Uses cheap object storage (e.g., S3).

    • Scalable: Can store petabytes of diverse data.

  • Merits:

    • Flexibility: Accommodates any future data type or schema change without upfront design.

    • Scalability & Cost: Cheap, virtually unlimited storage.

    • Foundation for ML/Analytics: Raw data is ideal for exploratory analysis and training ML models.

  • Applications: Machine Learning training datasets, archival storage, exploratory data analysis, data science sandboxes.

[!TIP] Key Distinction: Data Lake = Raw, unprocessed, schema-on-read. Data Warehouse = Processed, structured, schema-on-write.

Data Warehouse Design - Step-by-Step Approach (e.g., University Setup)

  1. Requirement Gathering: Identify key stakeholders (Admissions, Academics, Finance) and their KPIs (e.g., enrollment rates, course pass rates, fee collection).

  2. Dimensional Modeling: Choose a schema (Star Schema is common). Identify:

    • Fact Tables: Central tables containing measurable metrics (e.g., Fact_Enrollment with student_id, course_id, semester_id, grade).

    • Dimension Tables: Descriptive attributes (e.g., Dim_Student with name, program, address; Dim_Course with title, credits, department).

  3. Source System Analysis: Map source systems (Student Information System, Learning Management System) to target facts/dimensions.

  4. ETL/ELT Pipeline Design: Define how to extract, transform (clean, conform, aggregate), and load data from sources into the warehouse.

  5. Tool & Technology Selection: Choose cloud DWH (Snowflake, BigQuery, Redshift), orchestration (Airflow), and BI tools (Tableau, Power BI).

  6. Implementation & Testing: Build pipelines, load initial data, validate accuracy and performance.

  7. Deployment & Maintenance: Deploy to users, set up monitoring, and establish a process for schema changes and pipeline updates.


IV. Data Storage Solutions

Feature Data Lake Data Warehouse Data Mart
Data Type Raw, unstructured, semi-structured Structured, processed Structured, subject-specific
Schema Schema-on-Read Schema-on-Write Schema-on-Write
Users Data scientists, engineers Business analysts, executives Department-specific teams
Cost Low (object storage) High (compute + storage) Medium
Agility Very High Low (rigid schema) Medium
Primary Goal Store everything Fast, reliable analytics Solve specific business problem

Modern Storage Patterns:

  • Lakehouse Architecture: Combines the flexibility and cost of a data lake with the management features and performance of a data warehouse. Uses open table formats (Delta Lake, Apache Iceberg, Apache Hudi) on top of cloud storage (S3) to provide ACID transactions, versioning, and indexing.

  • Cloud-Native Storage: Object storage as the foundational layer (AWS S3, Azure Data Lake Storage, Google Cloud Storage). All other services (DWH, processing) read/write from this central "data landing zone."


V. Data Integration & Acquisition Techniques (High Priority)

ETL vs. ELT

ETL (Extract, Transform, Load) ELT (Extract, Load, Transform)
Transform data before loading into target. Load raw data first into target (e.g., cloud DWH), then transform.
Requires upfront schema design. Leverages powerful target DWH compute for transformation.
Good for structured sources, complex cleansing. Ideal for cloud DWHs, handling large volumes, and agile schema changes.
Example: Extract from Oracle, transform in a staging server, load to Snowflake. Example: Extract logs, load as JSON to BigQuery, use SQL to transform.

Schema Migration Strategies

  • Schema Evolution: The system automatically adapts to changes in incoming data schema (e.g., adding a new column). Example: A streaming job adds a new field user_device to JSON records; the Parquet sink automatically includes it.

  • Schema Enforcement: The system rejects data that does not conform to a predefined schema. Ensures data quality but requires careful change management. Example: A Delta table has a defined schema; a record with an extra field is rejected or triggers an alert.

Data Ingestion Mechanisms

  • Webhooks (High Priority - 11m Question):

    • Definition: An event-driven HTTP callback where a source application (e.g., GitHub) sends an automated HTTP POST message to a pre-configured URL (your server) when a specific event occurs.

    • How it works: 1. You register a URL with the source app. 2. Event happens (e.g., new commit). 3. Source app sends HTTP request with event payload to your URL. 4. Your endpoint processes the data.

    • Examples: GitHub (push events), Stripe (payment succeeded), Slack (message posted).

    • Use Case: Real-time notifications, triggering downstream ETL jobs upon new data arrival.

  • Web Scraping (High Priority - Direct Question):

    • Definition: Programmatically extracting data from websites by parsing HTML/XML content.

    • Tools: BeautifulSoup (Python, parsing), Scrapy (Python, full framework), Selenium (for JS-heavy sites).

    • Ethical/Legal: Check robots.txt, respect rate limits, review Terms of Service. Do not scrape copyrighted/personal data without permission.

    • Scenario Implementation (RGPV Example):

      1. Use requests to fetch https://www.rgpvonline.com.

      2. Use BeautifulSoup to parse HTML.

      3. Search for elements containing "representatives" and "press releases" and "data".

      4. Extract representative names and links to their press releases.

      5. Store results in a CSV/database.

  • Secure Copy Protocol (SCP) (Direct Question):

    • Definition: A secure method for transferring files between a local and a remote host (or between two remote hosts) over an SSH connection.

    • How it works: Uses SSH for authentication and encryption. Command: scp file.txt user@remotehost:/path/.

    • Cloud Use Case: Securely moving log files or backup archives from an EC2 instance to an S3 bucket (often via sftp or scp to a bastion host) or between VPCs.


VI. DataOps & Lineage Tracking (Very High Priority)

DataOps applies DevOps principles (CI/CD, automation, collaboration) to data pipelines to improve quality, speed, and reliability.

Data Lineage Tracking

The ability to trace the origin, movement, and transformation of data throughout its lifecycle. Critical for debugging, compliance (GDPR), and impact analysis.

Pattern-Based Lineage Lineage by Data Tagging
Automated extraction from transformation logic (SQL SELECT statements, Spark code, dbt models). Manual or semi-automated assignment of metadata tags (e.g., PII, source_system) to datasets/columns.
Tools parse code to build a graph of data flow: which table/column feeds into which transformation and output. Relies on a data catalog (e.g., AWS Glue, Alation) where users or systems tag assets. Lineage is inferred from tag propagation.
Example: Parsing a SQL query SELECT user_id, SUM(amount) FROM sales GROUP BY user_id shows sales.amount -> fact_sales.total_amount. Example: Tagging a column ssn as PII. The catalog shows all downstream datasets/columns that inherit this tag via ETL jobs.
Best for: Technical debugging, understanding complex code dependencies. Best for: Compliance, data discovery, business glossaries.

[!TIP] Exam Answer Structure: For lineage questions, define it, then contrast the two methods with a clear example for each. Emphasize that Pattern-Based is code-driven and technical, while Tagging is metadata-driven and business-oriented.


VII. Advanced Data Processing & Analytics

  • Real-Time & Streaming Architectures: Core pattern: Ingest -> Process -> Serve.

    • Ingest: Apache Kafka (decouples producers/consumers), AWS Kinesis.

    • Process: Stream processors like Apache Flink (stateful, exactly-once), Spark Streaming, Kafka Streams.

    • Serve: Processed streams written to low-latency stores (Redis, Cassandra) or stream-table joins in a lakehouse.

    • Integration with Lambda/Kappa: Kafka is often the central log in a Kappa architecture. In Lambda, Kafka feeds the Speed layer, while the Batch layer reads from the data lake.

  • Time-Series Data Pattern Detection:

    • Pattern: Data points indexed in time order (e.g., stock prices, IoT sensor readings).

    • Detection Techniques in Streaming:

      1. Windowing: Tumble/Sliding/Session windows to define computation boundaries (e.g., "average temperature every 5 minutes").

      2. Aggregation: Compute metrics (sum, avg, max) over windows.

      3. Anomaly Detection: Use statistical methods (Z-score) or ML models (Isolation Forest) on the stream to flag outliers.

    • Streaming Architecture for IoT:

      IoT Devices -> MQTT Broker (e.g., HiveMQ) -> Kafka -> Flink/Spark Streaming (windowing, anomaly detection) -> Alerting (SNS) / Storage (S3/TSDB like InfluxDB) -> Dashboard (Grafana).


VIII. Operational Monitoring & Management

Logging, Monitoring, and Alerting (LMA) is the triad for pipeline health.

  1. Logging: Capture granular, timestamped events from pipeline components (e.g., "Job X started," "Row Y failed validation"). Use structured logging (JSON).

  2. Monitoring: Aggregate and visualize key metrics:

    • Throughput: Records/sec processed.

    • Latency: End-to-end delay (ingestion to serving).

    • Error Rates: % of failed records/jobs.

    • Resource Utilization: CPU/Memory of Spark executors, DB connections.

    • Tools: Prometheus (metrics collection), Grafana (dashboards), CloudWatch (AWS), Stackdriver (GCP).

  3. Alerting: Define thresholds on monitored metrics (e.g., "Alert if error rate > 5% for 5 min"). Use PagerDuty, Slack, or Email notifications.

[!TIP] Implementation Example: An Airflow DAG failure triggers a CloudWatch alarm -> sends a Slack message to the #data-ops channel. A Grafana dashboard shows real-time Kafka consumer lag.


IX. Data Governance & Security (High Priority)

Data Governance Framework

A set of policies, standards, roles, and processes to ensure data is secure, private, accurate, and available.

  • Key Components:

    • Policies & Standards: Data quality rules, naming conventions, retention policies.

    • Roles:

      • Data Owner: Business executive accountable for data (defines policies).

      • Data Steward: Manages data day-to-day, ensures quality, implements policies.

      • Data Custodian: IT role managing technical infrastructure (security, backups).

    • Master Data Management (MDM): Creates a single, authoritative source for critical business entities (Customer, Product).

    • Data Quality: Dimensions: Accuracy, Completeness, Consistency, Timeliness, Validity.

Cloud Security Considerations

  • Encryption:

    • At Rest: Storage-level (S3 SSE, Azure Storage Service Encryption), DB-level (TDE).

    • In Transit: TLS/SSL for all data movement (HTTPS, SSL for DB connections).

  • Identity and Access Management (IAM): Principle of Least Privilege. Use roles, policies, and MFA. Federate with corporate IdP.

  • Network Security:

    • VPCs/VNets: Isolate resources in private subnets.

    • Security Groups/NACLs: Control traffic at instance/network level.

    • Private Endpoints: Access cloud services (e.g., S3) without traversing the public internet.

  • Compliance: Understand shared responsibility model. Cloud provider secures infrastructure, you secure data & access.

    • GDPR: Data residency, right to erasure. Use region-locked storage, implement purge pipelines.

    • HIPAA: Protect PHI. Use encryption, strict access controls, audit logging.

[!TIP] Governance vs. Security: Governance is about policies and people (who can do what with data). Security is about technical controls (encryption, IAM) to enforce those policies. They are deeply intertwined.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in