Skip to content
CY-703 (C) · Data Engineering/Quick Revision Short Notes

Data Engineering (CY-703 (C)) - Unit 4 Short Notes

A. FOUNDATIONAL CONCEPTS & THE DATA ENGINEER'S ROLE

The Data Engineering Discipline & Lifecycle

  • Definition: Design, build, and maintain systems for collecting, storing, processing, and serving data for analytical and operational use.

  • Lifecycle Stages:

    1. Ingestion: Acquiring data from source systems.

    2. Storage: Persisting data in systems (lakes, warehouses).

    3. Processing: Transforming and preparing data (ETL/ELT, batch/stream).

    4. Serving: Making data available to consumers via APIs, dashboards, or databases.

    5. Consumption: End-users (analysts, data scientists, applications) using the data.

Data Characteristics & Sources

  • The Five V's of Big Data:

    | V | Definition | Example | |---|---|---| | Volume | Scale of data generated/collected. | Terabytes/Petabytes of IoT sensor logs. | | Velocity | Speed of data generation and processing needs. | Real-time stock market tick data. | | Variety | Different formats and structures. | Structured (SQL DB), semi-structured (JSON logs), unstructured (images, videos). | | Veracity | Quality, accuracy, and trustworthiness of data. | Incomplete customer records, noisy social media data. | | Value | Extracting meaningful insights or business benefit. | Using purchase history to predict customer churn. |

  • Data Types: Structured (relational tables), Semi-structured (XML, JSON), Unstructured (text, audio, video).

  • Data Sources: Internal (CRM, ERP, transactional DBs), External (APIs, social media, public datasets), IoT sensors, Application logs.

Data-Driven Organization & Decision Making

  • Role of Data Engineer: Build reliable, scalable data pipelines and platforms that provide clean, accessible data to analysts and scientists.

  • Impact: Enables evidence-based decisions (e.g., optimizing supply chain using real-time logistics data, personalizing marketing via customer behavior analysis).

[!TIP] Exam Focus: The Five V's and the Data Engineer's role in enabling decisions are repeatedly asked (Dec 2024, Nov 2023). Always link the V's to concrete examples.


B. MODERN DATA ARCHITECTURES & STRATEGIES

Core Architectural Paradigms

  • Lambda Architecture: Processes both batch and stream data in parallel.

    • Batch Layer: Computes accurate results from historical data (e.g., Hadoop/Spark).

    • Speed Layer: Provides low-latency insights from real-time streams (e.g., Spark Streaming, Flink).

    • Serving Layer: Merges outputs from both layers for querying.

    • Challenge: Maintaining two separate codebases.

  • Kappa Architecture: Simplified alternative; handles all data as streams.

    • Uses a single stream processing engine for both real-time and historical data (replays streams).

    • Simpler to maintain but requires robust stream processing for all workloads.

  • Comparison:

    | Feature | Lambda Architecture | Kappa Architecture | |---|---|---| | Complexity | High (two separate layers) | Lower (single stream layer) | | Latency | Batch: high, Stream: low | Low (all stream) | | Data Recalculation | Easy on batch layer | Requires stream replay | | Use Case | Need for absolute batch accuracy + real-time views | Primarily real-time needs with historical context |

Enterprise Architecture Frameworks

  • Zachman Framework: A schema for organizing enterprise architecture artifacts. Uses a 6x6 matrix (Who, What, When, Where, Why, How) across different stakeholder perspectives (Planner, Owner, Designer, Builder, Sub-contractor, Functioning Enterprise).

  • Gartner Data Maturity Model: Stages of organizational data capability evolution: Awareness → Emerging → Consolidating → Integrated → Optimized.

Strategic Approaches

  • Modern Data Strategies: Shift from monolithic, on-premise warehouses to cloud-native, modular stacks (separation of storage & compute), ELT, and self-service analytics.

  • Purpose-Built Ingestion Tools: SaaS tools (Fivetran, Stitch, Airbyte) that automate connector management, schema drift handling, and data loading to cloud destinations. Reduce engineering overhead.

[!TIP] Exam Focus: Lambda vs. Kappa and Zachman Framework appeared in Nov 2023. Be ready to draw/explain the Zachman matrix and contrast the architectures.


C. DATA STORAGE SYSTEMS

Data Lake

  • Definition: Centralized repository storing raw, unprocessed data in its native format (object storage like S3/ADLS).

  • Core Principle: Schema-on-Read (apply schema when data is read, not at write time).

  • Patterns: Basic (single lake), Advanced (zones: Raw, Cleansed, Curated), Multi-Cloud.

  • Merits: Cost-effective storage, schema flexibility, stores all data types.

  • Use Cases: Data science exploration, storing raw logs, archival.

Data Warehouse

  • Definition: Optimized repository for structured, processed data for SQL analytics and BI.

  • Core Principle: Schema-on-Write (data must conform to a predefined schema before loading).

  • Traditional vs. Cloud: Traditional (Teradata, Oracle) vs. Cloud (Snowflake, BigQuery, Redshift) which separate storage/compute and are more elastic.

Data Lake vs. Data Warehouse

Feature Data Lake Data Warehouse
Schema Schema-on-Read Schema-on-Write
Data Type All (raw, structured, unstructured) Structured, processed
Users Data scientists, engineers Business analysts, BI
Cost Low (cheap object storage) Higher (optimized compute/storage)
Agility High (flexible schema) Lower (rigid schema)
Primary Use Exploration, ML, raw storage Reporting, dashboards, SQL analytics
  • Lakehouse Architecture: Convergence combining data lake's low-cost storage with data warehouse's ACID transactions and management features (e.g., Delta Lake, Apache Iceberg on top of Parquet files in a data lake).

[!TIP] Exam Focus: Data Lake merits/applications (Nov 2023) and Lakehouse concept are key. Be prepared to contrast Lake vs. Warehouse in a table.


D. DATA PROCESSING & COMPUTATION FRAMEWORKS

Big Data Processing Fundamentals

  • Need for Distributed Processing: Single machines cannot handle Volume/Velocity. Requires horizontal scaling (adding more machines).

  • Core Concepts:

    • Horizontal Scaling: Scale out by adding commodity servers.

    • Fault Tolerance: System continues operating despite node failures (via data replication).

    • Data Locality: Move computation to where data resides (in HDFS) to minimize network traffic.

Apache Hadoop Ecosystem

  • Hadoop Architecture:

    
    graph LR
    
    A[Client] --> B[HDFS<br/>NameNode/DataNode]
    
    A --> C[YARN<br/>ResourceManager]
    
    C --> D[MapReduce<br/>ApplicationMaster]
    
    D --> B
    
    
    • HDFS: Distributed file system. NameNode (metadata master), DataNodes (block storage).

    • YARN: Resource manager. ResourceManager (cluster resources), NodeManager (per-node agent).

    • MapReduce: Programming model for parallel processing (map -> shuffle -> reduce).

  • Advantages: Scalable, cost-effective, fault-tolerant.

  • Limitations: High disk I/O (MapReduce), not ideal for iterative/real-time processing, complex to manage.

Apache Spark

  • Key Features & Architecture:

    • RDDs (Resilient Distributed Datasets): Immutable, partitioned collection of records. Fault-tolerant via lineage.

    • DAG Scheduler: Creates a Directed Acyclic Graph of operations for optimized execution (vs. MapReduce's fixed map-shuffle-reduce).

    • In-Memory Processing: Caches intermediate data in RAM, 10-100x faster than MapReduce for iterative workloads.

  • Advantages over MapReduce:

    | Feature | MapReduce | Spark | |---|---|---| | Speed | Disk I/O bound | In-memory, DAG optimization | | Ease of Use | Java API (verbose) | Rich APIs (Scala, Python, Java, R) | | Unified Engine | Batch only | Batch, Streaming, SQL, ML, Graph | | Latency | High (minutes/hours) | Low (sub-second to minutes) |

  • Components: Spark SQL (structured data), Spark Streaming (micro-batch), MLlib (machine learning), GraphX (graph processing).

Cloud-Managed Processing Services: Amazon EMR

  • Definition: Managed service that simplifies running big data frameworks (Hadoop, Spark, HBase, Presto) on AWS.

  • Role: Handles cluster provisioning, configuration, tuning, and scaling. Integrates with S3, Glue, Redshift.

  • Key Features: Auto-scaling, pay-as-you-go, easy cluster creation/termination, supports spot instances for cost savings.

[!TIP] Exam Focus: Spark advantages over MapReduce and Amazon EMR are high-frequency (Dec 2024, Nov 2023). Know the DAG and in-memory concepts for Spark.


E. DATA INTEGRATION & TRANSFORMATION (ETL/ELT)

ETL (Extract, Transform, Load)

  • Flow: Extract → Transform (in staging area) → Load.

  • Characteristics: Transformation happens before loading into target. Requires powerful staging servers. Data is cleaned/structured before warehouse entry.

  • Use Cases: Legacy systems, complex transformations, strict data governance needs.

ELT (Extract, Load, Transform)

  • Flow: Extract → Load (raw) → Transform (in target warehouse/lake).

  • Characteristics: Raw data loaded first; transformation leverages cloud data warehouse's scalable compute (e.g., Snowflake, BigQuery).

  • Use Cases: Cloud data platforms, fast ingestion, flexible schema evolution, data lake environments.

ETL vs. ELT

Aspect ETL ELT
Transformation Location Staging area Target system (warehouse/lake)
Performance Can be slower (bottlenecked by staging) Faster (leverages target's scalable compute)
Flexibility Less flexible (schema fixed early) High (raw data kept, transform as needed)
Data Storage Cost Lower (only clean data stored) Higher (raw + transformed data stored)
Best For Small/structured data, strict compliance Big data, cloud platforms, exploratory analysis

Batch vs. Stream Ingestion & Processing

  • Batch Ingestion/Processing:

    • Process data in large, discrete chunks at scheduled intervals (e.g., hourly, daily).

    • Tools: Sqoop, batch Spark jobs, AWS Glue.

    • Pros: Simple, efficient for large volumes, cost-effective.

    • Cons: High latency (data not fresh).

  • Stream Ingestion/Processing (Real-time):

    • Process data continuously as it arrives.

    • Tools: Apache Kafka, AWS Kinesis, Apache Flink.

    • Scaling Considerations:

      1. Throughput: Messages/sec the system can handle.

      2. Latency: Time from event to processed output.

      3. State Management: Handling stateful operations (windowed aggregations) across failures.

      4. Exactly-Once Semantics: Guarantee each record is processed exactly once (critical for financial data).

  • Streaming Analytics Pipeline: Source (Kafka) → Stream Processor (Flink/Spark Streaming) → Sink (Data Lake/DB/Dashboard).

[!TIP] Exam Focus: ETL vs. ELT comparison and stream processing scaling are high-frequency (Nov 2023). Know the trade-offs and state management challenges.


F. CLOUD PLATFORMS & MANAGED SERVICES (AWS FOCUS)

Modern Data Architecture on Cloud Platforms

  • Cloud platforms provide integrated, managed services for each lifecycle stage (S3 for storage, Glue for ETL, Redshift for warehousing, Kinesis for streaming, SageMaker for ML).

  • Benefits: Elastic scalability, pay-per-use, global availability, reduced operational overhead, security/compliance built-in.

  • Comparison: AWS (broadest services), Azure (strong enterprise/MSFT integration), GCP (big data/ML strengths).

AWS Ecosystem for Data & ML

  • AWS SageMaker:

    • Overview: Fully managed service for building, training, and deploying ML models at scale.

    • Supports Scalability: Provides managed Jupyter notebooks, distributed training jobs, one-click deployment to auto-scaling endpoints, and model monitoring.

    • ML Infrastructure on AWS:

      DiagramCANVAS: Show flow: Data in S3 → SageMaker Processing/Feature Store → Training Job (distributed) → Model Registry → Endpoint (auto-scaling) → Application/BI.

  • Key AWS Services:

    • S3: Scalable object storage (Data Lake foundation).

    • Glue: Serverless ETL service (Spark-based).

    • Redshift: Cloud data warehouse.

    • Kinesis: Real-time data streaming (Kinesis Data Streams, Firehose).

    • MSK: Managed Apache Kafka service.

Creating Scalable/Feasible Infrastructure

  • Process:

    1. Requirement Analysis: Data volume, velocity, user concurrency, latency SLAs.

    2. Architecture Design: Choose services (S3 + Glue + Redshift vs. Lakehouse), define data zones (raw, refined).

    3. Cost Estimation: Use AWS Pricing Calculator; consider storage, compute, data transfer.

    4. Security & Governance: IAM roles, encryption, VPC endpoints, data cataloging.

    5. Implementation & Automation: Use IaC (CloudFormation, Terraform) for reproducibility.

    6. Monitoring & Optimization: CloudWatch metrics, cost/performance tuning.

  • Example: For an e-commerce company: S3 (raw clickstream logs) → Kinesis (real-time inventory) → Glue (batch customer data transformation) → Redshift (sales analytics dashboard) → QuickSight (BI).

[!TIP] Exam Focus: AWS SageMaker and scalable infrastructure design are high-frequency (Dec 2024). Be ready to describe the ML infrastructure diagram and give a concrete example.


G. MACHINE LEARNING INTEGRATION

Machine Learning Lifecycle

  • Key Stages:

    1. Problem Definition: Frame business problem as ML task.

    2. Data Collection: Gather relevant data from sources.

    3. Data Pre-processing & Feature Engineering: Clean, transform, create features (most time-consuming).

    4. Modeling: Select algorithm, train, tune hyperparameters.

    5. Evaluation: Assess model performance (accuracy, F1-score, etc.).

    6. Deployment: Serve model as API/batch process.

    7. Monitoring & Maintenance: Track drift, retrain.

Data Preparation for ML

  • Importance: "Garbage in, garbage out." Quality data is foundational. ~80% of ML time is spent here.

  • Common Techniques:

    • Handling missing values (imputation, removal).

    • Scaling/Normalization (StandardScaler, MinMaxScaler).

    • Encoding categorical variables (One-Hot, Label Encoding).

    • Feature creation (deriving new features from existing ones).

    • Feature selection (removing irrelevant features).

MLOps & Deployment

  • Role of Data Engineer: Build and maintain ML pipelines (data versioning, feature stores, training data generation, model deployment infrastructure).

  • CI/CD for ML: Automate testing, building, and deployment of both code and models (e.g., using Jenkins/GitLab CI with SageMaker Pipelines or Kubeflow).

[!TIP] Exam Focus: ML lifecycle stages and pre-processing importance appeared in Dec 2024. Emphasize that data prep is the most critical and time-intensive phase.


H. DATA OPERATIONS (DataOps), GOVERNANCE & QUALITY

DataOps Principles

  • Applying DevOps practices (CI/CD, automation, monitoring) to data pipelines to improve speed, quality, and collaboration.

  • Data Lineage Tracking:

    • Pattern-Based: Traces data flow by analyzing pipeline code/dags (e.g., parsing Spark jobs).

    • Lineage by Data Tagging: Uses metadata tags on datasets/tables to infer relationships (simpler but less granular).

    • Importance: Debugging pipeline failures, impact analysis (what breaks if source changes), compliance (GDPR, CCPA).

Data Governance

  • Definition: Framework of policies, standards, roles, and processes to ensure data is secure, private, accurate, and available.

  • Components:

    • Policies: Data ownership, usage rights.

    • Standards: Naming conventions, data models.

    • Roles: Data owners, stewards, custodians.

    • Tools: Data catalogs (AWS Glue Data Catalog, Alation), metadata management, access control systems.

Data Quality & Wrangling

  • Data Wrangling/Munging: Process of cleaning, structuring, and enriching raw data into a desired format for analysis. Steps: Discovery → Structuring → Cleaning → Enriching → Validating.

  • Data Discovery: Using data catalogs and metadata to find, understand, and assess available datasets for a use case.

  • Tools for Data Quality:

    • Great Expectations: Python-based; defines "expectations" (tests) for data (e.g., column values not null).

    • Deequ: AWS open-source library for data quality checks on Spark.

[!TIP] Exam Focus: Data wrangling/discovery (Dec 2024) and data lineage patterns (Nov 2023) are specific topics. Know the steps of wrangling and how lineage is tracked.


I. SECURITY, MONITORING & ADVANCED TOPICS

Cloud Storage Security

  • Process (Layered Defense):

    1. Encryption:

      • At Rest: Server-side (SSE-S3, SSE-KMS), client-side.

      • In Transit: TLS/SSL.

    2. Access Control (IAM): Fine-grained permissions via IAM users, roles, policies.

    3. Bucket Policies: Resource-based policies on S3 buckets (e.g., allow access only from specific VPC).

    4. VPC Endpoints: Access S3 privately within a VPC without traversing the public internet (Gateway endpoint).

  • Diagram:

    DiagramSEARCH: AWS S3 security architecture diagram showing encryption, IAM, bucket policy, VPC endpoint layers.

Logging, Monitoring, and Alerting

  • Importance: Ensure pipeline reliability, detect failures/performance degradation, cost tracking.

  • Tools & Metrics:

    • CloudWatch (AWS): Logs (CloudWatch Logs), metrics (CPU, memory, custom), alarms.

    • Prometheus/Grafana: Open-source monitoring stack (metrics collection + visualization).

    • Pipeline Health Checks: Data freshness, record counts, error rates, SLA compliance.

Specialized Ingestion & Integration Techniques

  • Webhooks: HTTP callbacks triggered by an event (e.g., GitHub sends POST to your server on a code push). How: Event → Source system makes HTTP request to pre-configured URL → Receiver processes payload.

  • Web Scraping: Programmatically extracting data from websites.

    • Tools: BeautifulSoup (Python, HTML parsing), Scrapy (Python, full framework).

    • Ethical: Respect robots.txt, terms of service, rate limiting; avoid copyright infringement.

  • Secure Copy Protocol (SCP): Command-line tool for secure file transfer between hosts using SSH. scp file user@remote:/path.

  • Schema Migration: Managing changes to data schema over time (evolution).

    • Compatibility: Backward (new schema reads old data), Forward (old schema reads new data).

    • Tools: Avro/Protobuf (schema registry, support for evolution), database migration tools (Flyway, Liquibase).

[!TIP] Exam Focus: Cloud storage security (Dec 2024, Nov 2023) and specialized techniques (webhooks, scraping, SCP, schema migration) from Nov 2023 are specific but guaranteed. Know definitions and simple examples.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in