Skip to content
CY-703 (C) · Data Engineering/Quick Revision Short Notes

Data Engineering (CY-703 (C)) - Unit 3 Short Notes

I. Foundational Concepts & The Data Engineer's Role

Role of Data Engineer in Organizations

  • Core Responsibilities:

    • Design, build, and maintain scalable data pipelines (ingestion → storage → processing → serving).

    • Ensure data quality, reliability, and accessibility for analysts and data scientists.

    • Collaborate with stakeholders to translate business needs into technical specifications.

    • Implement monitoring, alerting, and optimization for data systems.

  • Involvement in Data-Driven Decision-Making:

    • Provides the foundational data infrastructure that enables analytics and ML.

    • Example: For a retail company, a data engineer builds a pipeline that consolidates point-of-sale transactions, website logs, and inventory data into a data warehouse. Analysts then use this clean, integrated data to identify top-selling products and optimize stock levels.

  • Data Engineering Lifecycle:

    1. Ingestion: Collect data from source systems.

    2. Storage: Persist data in lakes/warehouses.

    3. Processing: Clean, transform, and aggregate data.

    4. Serving: Make data available via APIs, dashboards, or databases.

    5. Monitoring: Track pipeline health, data quality, and costs.

[!TIP] Exam Focus: Always link the data engineer's role to enabling downstream analytics and business decisions with a concrete example (retail, healthcare, finance).

Core Data Characteristics

  • The Five V's of Big Data:

    | V | Definition | Example | |---|---|---| | Volume | Scale of data (TB/PB) | Social media posts per day | | Velocity | Speed of data generation/processing | IoT sensor readings per second | | Variety | Different data formats | Structured (SQL), semi-structured (JSON), unstructured (images) | | Veracity | Data quality & trustworthiness | Inconsistent customer addresses from multiple sources | | Value | Extracted business insights | Predicting customer churn from usage logs |

  • Types of Data:

    • Structured: Fixed schema (e.g., relational database tables).

    • Semi-structured: Tags/schema embedded (e.g., XML, JSON).

    • Unstructured: No predefined model (e.g., text files, videos, images).

  • Data Sources:

    • Internal: Transactional DBs, application logs, CRM systems.

    • External: Social media APIs, public datasets, IoT devices, third-party feeds.

Modern Data Strategies

  • Shift from monolithic, on-premise systems to cloud-native, modular architectures.

  • Key Principles:

    • Scalability: Elastic resources (scale up/down on demand).

    • Agility: Rapid deployment of new data products.

    • Cost-Effectiveness: Pay-as-you-go model, separation of compute/storage.

II. Modern Data Architectures & Storage

Cloud-Centric Data Architecture

  • Core Components (Layered Approach):

    1. Ingestion Layer: Tools like Kafka, Kinesis, or cloud services (AWS Kinesis, Azure Event Hubs).

    2. Storage Layer: Data Lakes (S3, ADLS) & Warehouses (Redshift, BigQuery).

    3. Processing Layer: Spark, EMR, Databricks for transformation.

    4. Analytics/Serving Layer: BI tools (Tableau, Power BI), ML services, APIs.

    5. Consumption Layer: End-users (analysts, apps, executives).

  • Cloud Platform Support (AWS Example):

    DiagramCANVAS: A flowchart showing: Data Sources → AWS Kinesis (Ingestion) → S3 (Data Lake) → AWS Glue (Processing/ETL) → Redshift (Data Warehouse) → QuickSight (Analytics/Consumption). Arrows indicate data flow.

Storage Paradigms: Data Lake vs. Data Warehouse

Feature Data Lake Data Warehouse
Schema Schema-on-Read (flexible) Schema-on-Write (rigid)
Data Types All (raw, structured, unstructured) Primarily structured, processed
Storage Cost Low (object storage) Higher (optimized compute/storage)
Users Data scientists, ML engineers Business analysts, executives
Use Case Exploratory analytics, ML, raw data archive BI, reporting, SQL queries
Examples AWS S3, Azure Data Lake Storage Snowflake, Google BigQuery, Redshift
  • Modern Lakehouse Architecture: Combines low-cost storage of lakes with management/performance of warehouses (e.g., Delta Lake, Apache Iceberg). Adds ACID transactions, versioning, and indexing on top of parquet files in a data lake.

Architectural Patterns

  • Lambda Architecture:

    • Batch Layer: Processes all historical data for accuracy (e.g., Hadoop/Spark).

    • Speed Layer: Processes real-time data for low-latency views (e.g., Spark Streaming, Flink).

    • Serving Layer: Merges batch & speed outputs for queries.

    • Drawback: Complex to maintain two codebases.

    DiagramCANVAS: Three horizontal layers. Bottom: Batch Layer (HDFS/Spark) → Serving Layer (serving DB). Top: Speed Layer (stream processing) → same Serving Layer. Arrows from both Batch and Speed converge at Serving.
  • Kappa Architecture: Simplifies Lambda by using only a stream processing layer for all data (historical data replayed through the stream). Reduces complexity but requires powerful stream processors.

  • Zachman Framework: A 6x6 matrix (Who, What, When, Where, Why, How) across perspectives (Executive, Business, Architect, Builder, Subcontractor, Functioning Enterprise). It's an enterprise ontology for organizing architectural artifacts, not a technical blueprint.

III. Data Ingestion & Processing Paradigms

Data Ingestion Strategies

  • Batch Ingestion:

    • Periodic collection of large data volumes (e.g., daily/hourly).

    • Tools: Apache Sqoop, custom scripts, cloud services (AWS DataSync).

    • Use Case: Daily sales report generation.

  • Stream/Real-time Ingestion:

    • Continuous flow of small data records.

    • Tools: Apache Kafka (distributed event streaming), Amazon Kinesis (managed service).

    • Use Case: Fraud detection, live dashboards.

  • Specialized Tools & Patterns:

    • Change Data Capture (CDC): Captures row-level changes in DBs (e.g., Debezium, AWS DMS) for streaming.

    • Webhooks: HTTP callbacks triggered by events (e.g., GitHub push → trigger CI/CD pipeline).

ETL vs. ELT

Aspect ETL (Extract-Transform-Load) ELT (Extract-Load-Transform)
Transformation Location Staging area (before loading) Target data warehouse/lake
Schema Requirement Defined before extraction Flexible; schema-on-read
Performance Slower for large data (bottlenecked on staging) Faster load; leverages target compute
Typical Tools Informatica, Talend, custom scripts dbt, Spark SQL, BigQuery ML
Best For Legacy systems, complex cleansing, small data Cloud data warehouses, large raw datasets, agile analytics

[!TIP] Rule of Thumb: Use ETL when source systems are fragile or transformations are extremely complex. Use ELT with modern cloud warehouses (Snowflake, BigQuery) where compute is cheap and scalable.

Big Data Processing Frameworks

  • Apache Hadoop Ecosystem:

    • HDFS: Distributed file system storing data across clusters.

    • MapReduce: Programming model (Map → Shuffle → Reduce). Disk-based, high I/O latency.

    • YARN: Resource manager for cluster scheduling.

    • Limitations: High latency for iterative tasks; not ideal for real-time.

  • Apache Spark:

    • In-Memory Processing: Caches data in RAM, drastically reducing I/O.

    • APIs: RDDs (resilient distributed datasets, low-level), DataFrames/Datasets (high-level, optimized).

    • Components: Spark SQL (structured data), Spark Streaming (micro-batch), MLlib (machine learning).

    • Advantages over MapReduce:

      1. ~100x faster for in-memory workloads.

      2. Easier API (DataFrames vs. MapReduce code).

      3. Unified engine for batch, streaming, ML.

  • Cloud-Managed Services: Amazon EMR

    • Managed Hadoop/Spark cluster on AWS.

    • Benefits: Auto-scaling, integrated with S3, Glue, Redshift; no cluster management overhead.

    • Use Case: Running large-scale Spark jobs on data in S3 without provisioning servers.

IV. Scalable Infrastructure & Cloud Platforms

Building Scalable Infrastructure

  • Process:

    1. Assess Requirements: Data volume, velocity, query patterns, concurrency.

    2. Choose Architecture: Lakehouse? Lambda? (Based on latency needs).

    3. Select Cloud Services: Compute (EC2, EMR), Storage (S3), DBs (Redshift).

    4. Design for Scale: Use auto-scaling groups, decouple components (queues like SQS), partition data.

    5. Implement Monitoring & Cost Controls.

  • Example (E-Commerce):

    • Problem: Black Friday traffic spike (10x normal load).

    • Solution: Use Kinesis for ingestion (auto-scales), S3 for raw storage, EMR Spark for processing (auto-scaling cluster), Redshift for serving (concurrency scaling). Costs scale with usage.

Cloud Platform Deep Dives

  • AWS Modern Data Architecture:

    DiagramCANVAS: Left to Right: 1. Data Sources (DBs, logs) → 2. Ingestion (Kinesis, DMS, Snowball) → 3. Storage (S3 Data Lake) → 4. Processing (Glue, EMR, Lambda) → 5. Serving (Redshift, Athena, OpenSearch) → 6. Analytics (QuickSight, SageMaker).
    • Key Services: S3 (central lake), Glue (serverless ETL), Redshift (warehouse), Athena (serverless SQL on S3).
  • Azure Modern Data Architecture:

    • Key Services: Azure Data Lake Storage (Gen2), Azure Synapse Analytics (unified analytics), Azure Databricks (Spark), Azure Data Factory (orchestration).
  • GCP Modern Data Architecture:

    • Key Services: BigQuery (serverless warehouse), Cloud Storage (lake), Dataflow (stream/batch processing), Dataproc (managed Spark/Hadoop).

V. Machine Learning Integration in Data Pipelines

Machine Learning Lifecycle

1.  **Problem Definition:** Frame business problem as ML task.

2.  **Data Collection:** Gather relevant data from sources.

3.  **Data Pre-processing & Feature Engineering:** **Critical step** (cleaning, transforming, creating features).

4.  **Modeling:** Select/train algorithms.

5.  **Evaluation:** Measure performance (accuracy, F1-score).

6.  **Deployment:** Serve model as API or batch prediction.

7.  **Monitoring & Retraining:** Track drift, re-train with new data.

AWS SageMaker

  • End-to-End ML Service:

    • Notebooks: JupyterLab for exploration.

    • Training: Built-in algorithms or custom containers; auto-scaling.

    • Hyperparameter Tuning: Automated optimization.

    • Deployment: One-click to REST endpoints (real-time) or batch transform.

    • Monitoring: Detect model drift, data quality issues.

  • ML Infrastructure on AWS:

    DiagramCANVAS: Central S3 Data Lake → SageMaker Processing (feature engineering) → SageMaker Training (built-in/custom) → SageMaker Model Registry → SageMaker Endpoints (real-time) or Batch Transform. Integrated with CloudWatch for monitoring.

[!TIP] Exam Key: Feature engineering often determines model success more than algorithm choice. Data engineers build the feature store and pipelines that feed clean, consistent features to training and inference.

VI. Data Operations (DataOps), Quality & Governance

Data Wrangling & Discovery

  • Process: Cleaning (handle nulls, outliers), transforming (normalization, aggregation), exploring (profiling, visualization).

  • Tools: Python (Pandas, Great Expectations), dbt (tests & docs), cloud services (AWS Glue DataBrew).

Data Lineage Tracking

  • Importance: Debug pipeline failures, impact analysis, compliance (GDPR, HIPAA).

  • Pattern-Based Lineage: Infers lineage by analyzing code (SQL scripts, Spark jobs). Example: Parsing a Spark job to see it reads raw_table → transforms → writes aggregated_table.

  • Lineage by Data Tagging: Attaches metadata tags (source, owner, PII) to datasets. Tools track tag propagation. Example: Tag a source table as PII; all downstream tables inheriting that tag are automatically flagged.

Data Governance

  • Frameworks: Policies, standards, ownership (data stewards), security controls.

  • Gartner Data Maturity Model:

    | Stage | Characteristics | Example | |---|---|---| | Awareness | No formal governance | Ad-hoc reports, unknown data sources | | Reactive | Issues drive actions | Fix data errors after complaints | | Defined | Policies & roles exist | Data dictionary, basic quality rules | | Managed | Proactive monitoring | Automated quality checks, lineage | | Optimized | Business value driven | Data as product, self-service with guardrails |

Schema Management

  • Schema Evolution: Handling changes to data structure over time (e.g., adding a new column). Challenge: Backward/forward compatibility in streaming.

  • Schema Migration: Moving from one schema to another (e.g., denormalizing for a warehouse). Requires data transformation scripts and validation.

VII. Security, Monitoring & CI/CD

Cloud Storage Security (e.g., AWS S3)

  • Controls:

    1. IAM Policies: Fine-grained user/role permissions (e.g., s3:GetObject).

    2. Bucket Policies: Resource-based policies attached to bucket.

    3. Encryption:

      • At Rest: SSE-S3 (AWS managed), SSE-KMS (customer keys).

      • In Transit: HTTPS/TLS.

    4. Access Control Lists (ACLs): Legacy, object-level (use sparingly).

  • Secure Diagram:

    DiagramCANVAS: S3 Bucket with icons: 1. IAM Policy arrow from IAM User → Bucket. 2. Bucket Policy attached to bucket. 3. Lock icon for "Encryption at Rest". 4. HTTPS lock for "Encryption in Transit". 5. ACL as small separate box.

Operational Excellence

  • Logging, Monitoring, Alerting:

    • Tools: AWS CloudWatch, GCP Cloud Logging/Monitoring.

    • Metrics: Pipeline latency, error rates, data freshness, cost.

    • Alerting: Set thresholds (e.g., alert if job fails > 3 times).

  • CI/CD for Data Pipelines:

    • Automate: Code testing (dbt tests), deployment (Airflow DAGs, Glue jobs), infrastructure (Terraform).

    • Pipeline: Git → CI (test) → CD (deploy to staging/prod) → Monitor.

Secure Copy Protocol (SCP)

  • Use Case: Securely transfer files between local machine and remote server (or server-to-server) over SSH.

  • Operation: scp -i key.pem file.txt ec2-user@host:/path. Encrypts both authentication and data transfer.

  • Note: For cloud-to-cloud, use native services (e.g., AWS DataSync) or encrypted channels (TLS).

VIII. Specialized Topics & Applied Scenarios

Advanced Streaming Analytics

  • Pipeline Architecture:

    DiagramCANVAS: 1. Sources (IoT, logs) → 2. Streaming Ingestion (Kafka/Kinesis) → 3. Stream Processing (Spark Streaming/Flink for windowing, aggregation) → 4. Serving (Redis, Cassandra for low-latency) → 5. Consumption (Dashboard like Grafana).
  • Detecting Patterns in Time-Series:

    • Use windowing operations (tumbling, sliding windows) in stream processor.

    • Apply statistical models (moving average, anomaly detection) on each window.

    • Example: Detect sudden spike in server CPU usage over 5-minute window → trigger alert.

Applied Case Studies

  • Building a Data Warehouse for a University:

    1. Identify Sources: Student DB (SQL), course evaluations (CSV), website logs.

    2. Ingest: Batch load DBs nightly via Sqoop/Glue; stream logs via Kafka.

    3. Store: Raw data in Data Lake (S3/ADLS).

    4. Model: Design star schema (Fact: enrollments, Dimensions: students, courses, time).

    5. Process: Use dbt or Spark to transform raw → conformed dimensions/facts.

    6. Serve: Load into Warehouse (Redshift/Snowflake).

    7. Consume: Build dashboards (retention rates, course popularity) in Power BI.

  • ETL from Retail Database to Warehouse (Example):

    • Extract: Daily sales and inventory tables from OLTP DB (PostgreSQL).

    • Transform: Join with product dimension, calculate daily_revenue, filter returns.

    • Load: Append to fact_sales table in warehouse; update dim_product slowly changing dimension (SCD Type 2).

  • Web Scraping for Regulatory Analysis (Scenario):

    • Tool: Python requests + BeautifulSoup or Scrapy.

    • Steps:

      1. Crawl https://www.rgpvonline.com for pages containing "data" in press releases.

      2. Extract: Representative name, date, URL, excerpt.

      3. Store in CSV/JSON → load to data lake for analysis.

    • Legal Note: Check robots.txt and terms of service; respect rate limits.

Data Lake Patterns & Applications

  • Medallion Architecture:

    • Bronze: Raw, immutable data (exact source copy).

    • Silver: Cleaned, validated, conformed data (business-level integrity).

    • Gold: Aggregated, business-level datasets (ready for consumption).

  • Merits: Incremental refinement, traceability, separation of concerns.

  • Applications: Enterprise data lakes, ML feature stores, regulatory reporting.


Final Exam Strategy: For 7-mark questions, structure answers as: Definition → Key Components/Steps → Example/Use Case → Advantages/Limitations. Always relate concepts to real-world cloud services (AWS/Azure/GCP) where possible.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in