UNIT 4: Data Engineering & Related Concepts (Based on Past Exam Analysis)
I. Foundational Concepts of Data Engineering
Definition and Scope of Data Engineering
Data Engineering is the discipline of designing, building, and maintaining the infrastructure and systems that enable the collection, storage, processing, and analysis of large-scale data. Its scope includes:
-
Developing data pipelines for efficient data flow.
-
Ensuring data quality, reliability, and scalability.
-
Supporting data science and analytics workloads.
-
Managing cloud and on-premise data platforms.
Data Engineering Lifecycle
A cyclical process with five core components:
| Component | Description |
|---|---|
| 1. Ingestion | Acquiring data from diverse sources (APIs, logs, IoT sensors). |
| 2. Storage | Persisting data in databases, data lakes, or warehouses. |
| 3. Processing | Transforming raw data (cleaning, aggregating) using batch/streaming engines. |
| 4. Serving | Making processed data available to users/applications via APIs or query interfaces. |
| 5. Consumption | End-users (analysts, ML models) accessing data for insights. |
[!TIP] Exam Focus: Be prepared to draw and label this lifecycle diagram. Link each component to real-world tools (e.g., Ingestion: Apache Kafka; Storage: S3; Processing: Spark).
II. Data Architectures and Storage Patterns
Types of Data Architectures
| Architecture | Advantages | Disadvantages |
|---|---|---|
| Centralized (Monolithic) | Simple management, strong consistency. | Single point of failure, scalability limits. |
| Decentralized (Data Mesh) | Domain-oriented ownership, scalability, agility. | Complex governance, data duplication risks. |
| Hybrid (Lakehouse) | Combines cost-effective storage (lake) with structured querying (warehouse). | Integration complexity, potential performance overhead. |
Data Lake Patterns
A Data Lake stores raw, unstructured/semi-structured data at scale.
-
Merits:
-
Schema-on-read: Flexibility to define structure later.
-
Cost-effective: Uses cheap object storage (e.g., AWS S3).
-
Supports all data types: Logs, images, JSON, etc.
-
-
Applications:
-
Machine Learning: Raw data for feature engineering.
-
Exploratory Analytics: Ad-hoc queries on unprocessed data.
-
Regulatory Archiving: Immutable storage for compliance.
-
[!TIP] Common Pitfall: Do not confuse Data Lake (raw storage) with Data Warehouse (processed, structured). Lakehouse patterns aim to bridge this gap.
III. Data Processing Frameworks & Streaming
Lambda Architecture
A fault-tolerant, low-latency system for processing both real-time and batch data.
Step-by-Step Real-Time Processing:
-
Batch Layer: Precomputes batch views from all historical data (e.g., Hadoop MapReduce).
-
Speed Layer: Processes real-time data streams to compensate for batch latency (e.g., Apache Storm, Spark Streaming). Creates real-time views.
-
Serving Layer: Merges batch views and real-time views to serve complete, up-to-date query results.
[!TIP] Exam Key: Lambda is complex due to dual codebases (batch + speed). Mention trade-offs: accuracy vs. latency.
Kappa Architecture
A simplified alternative to Lambda:
-
Single processing pipeline: Uses only a streaming layer for all data (historical + real-time).
-
Historical data is replayed from storage (e.g., Kafka) through the same stream processor.
-
Advantage: Simpler code maintenance; Disadvantage: Requires powerful streaming engines for large historical reprocessing.
Streaming Data Architecture for Time-Series Data
Pattern Detection Methodology:
-
Windowing: Sliding/tumbling windows over time-series streams (e.g., 5-minute windows).
-
Aggregation: Compute metrics (avg, max, count) per window.
-
Anomaly Detection: Apply statistical/machine learning models (e.g., Z-score, LSTM) on aggregated data.
-
Alerting: Trigger alerts when patterns (spikes, dips) exceed thresholds.
Example: Detecting sudden CPU usage spikes in server logs using a 1-minute tumbling window and standard deviation > 3.
IV. Data Integration, Transformation, and Warehousing
Extract, Transform, Load (ETL) Process
Retail Database Example:
-
Extract: Pull
salesandinventorytables from operational PostgreSQL DB nightly. -
Transform:
-
Clean: Remove null
customer_id. -
Enrich: Join with
product_dimensionto addcategory. -
Aggregate: Compute
daily_sales_per_category.
-
-
Load: Write transformed data into fact_sales table in a Data Warehouse (e.g., Snowflake).
Schema Migration
Definition: Changing the structure of a database (tables, columns, relationships) while preserving data.
Example: Migrating from a denormalized single orders table (with repeated customer_name) to a normalized schema:
-
Create new
customerstable (extract uniquecustomer_id,name). -
Alter
orderstable to referencecustomer_id(foreign key). -
Backfill data and update application queries.
Data Warehouse Design (University Setup)
Step-by-Step Approach:
-
Identify Business Processes: Student enrollment, course scheduling, fee collection.
-
Define Facts: Numeric metrics (e.g.,
enrollment_count,fee_amount). -
Identify Dimensions:
Student(attributes: major, year),Course,Time,Instructor. -
Design Star Schema:
-
Fact Table:
fact_enrollment(keys:student_id,course_id,time_id; measures:credits). -
Dimension Tables:
dim_student,dim_course, etc.
-
-
ETL Pipeline: Extract from university's SIS (Student Information System), transform, load into warehouse.
-
BI Layer: Connect tools (Tableau) for dashboards (e.g., "Enrollment by Department").
[!TIP] Short Answer (3m): Focus on facts vs. dimensions and star schema for the university example.
V. Data Lineage, Governance, and Maturity
Data Lineage Tracking in DataOps
1. Pattern-Based Lineage:
-
Mechanism: Automatically infers lineage by analyzing data pipeline code (e.g., ETL scripts, Spark jobs).
-
Example: A Spark job reads
raw_sales→ filters → aggregates → writesdaily_sales. Lineage tool (e.g., OpenLineage) parses the job DAG to map this flow. -
Use Case: Debugging pipeline failures.
2. Lineage by Data Tagging:
-
Mechanism: Manually/automatically attaches metadata tags (e.g.,
PII,source:CRM) to datasets. Lineage is derived from tag propagation. -
Example: Tag
customer_emailasPIIin source CRM. When ETL copies it todata_warehouse.customers, the tag inherits. Shows impact analysis for GDPR compliance.
Gartner Data Maturity Model
Five stages of organizational data capability:
| Stage | Characteristics | Example |
|---|---|---|
| 1. Aware | Data seen as IT cost, siloed. | Reports generated manually per department. |
| 2. Reactive | Ad-hoc analytics, no standards. | Spreadsheets for sales analysis, inconsistent metrics. |
| 3. Defined | Enterprise-wide strategy, basic governance. | Central data team, documented data dictionary. |
| 4. Managed | Measurable quality, trusted sources. | SLAs for data freshness, automated data quality checks. |
| 5. Optimized | Data as strategic asset, AI-driven. | Real-time personalization using unified customer view. |
Data Governance
Core Principles:
-
Accountability: Clear roles (Data Owners, Stewards).
-
Transparency: Documented policies, data lineage.
-
Integrity: Data quality standards (accuracy, completeness).
-
Compliance: Adherence to regulations (GDPR, HIPAA).
-
Security: Access controls, encryption.
Practices: Establish data council, implement catalog tools (Alation), define data quality rules, conduct regular audits.
VI. Data Acquisition, Transfer, and Extraction Techniques
Web Scraping
Hypothetical Scenario: Scrape https://www.rgpvonline.com for press releases containing "data".
Illustration:
-
Identify URLs: Crawl site to find press release pages.
-
Extract Content: Use Python (
requests,BeautifulSoup) to parse HTML. -
Filter: Search for keyword "data" in title/body.
-
Store: Save representative names, release dates, URLs in CSV/DB.
-
Ethical/Legal: Check
robots.txt, respectrate-limiting, avoid copyright infringement.
[!TIP] Exam Answer Structure: Tools → Steps → Challenges (dynamic JS pages, CAPTCHAs) → Ethics.
Webhooks
Mechanism: Event-driven HTTP callbacks. When event X occurs on Server A, it pushes an HTTP POST to a pre-configured URL (Server B).
Example:
-
GitHub Webhook: On
pushevent, GitHub sends JSON payload to your CI/CD server (e.g., Jenkins) to trigger a build. -
Payment Gateway: On payment success, gateway sends webhook to merchant's server to update order status.
Working:
-
Subscriber registers URL with provider.
-
Event occurs → provider sends HTTP request with event data.
-
Subscriber processes payload (idempotently).
Secure Copy Protocol (SCP)
-
Overview: Uses SSH for secure file transfer between hosts.
-
Working:
scp file.txt user@remote:/pathencrypts file during transit via SSH tunnel. -
Use Case: Transferring sensitive data (e.g., logs, configs) in a controlled network.
-
Limitation: No built-in directory synchronization;
rsyncover SSH preferred for advanced needs.
VII. Monitoring, Logging, and Alerting
Logging, Monitoring, and Alerting Systems
Working Principles:
-
Logging: Record events (errors, transactions) with timestamps.
- Example: Application logs JSON to stdout; collected by Fluentd.
-
Monitoring: Collect & visualize metrics (CPU, latency, error rates).
- Example: Prometheus scrapes metrics from app endpoints; Grafana dashboards.
-
Alerting: Define thresholds (e.g., error rate > 5%) → trigger notifications (Slack, PagerDuty).
- Example: Alertmanager sends alert if
http_requests_failedspikes.
- Example: Alertmanager sends alert if
Example Stack (Modern):
App → Logs (JSON) → Fluentd → Elasticsearch → Kibana (Visualization)
Metrics → Prometheus → Grafana
Alerts → Alertmanager → Slack
[!TIP] Key Point: Emphasize the separation of concerns: logs (what happened?), metrics (how is it performing?), alerts (when to act?).
VIII. Enterprise Frameworks
Zachman Framework
A taxonomy for enterprise architecture using a 6x6 matrix.
Rows (Perspectives):
-
Planner (Scope) – What business entities?
-
Owner (Business) – What business processes?
-
Designer (Architect) – What data models?
-
Builder (Implementer) – What technologies?
-
Subcontractor (Code) – What programs?
-
Functioning Enterprise (Operations) – What outputs?
Columns (Abstractions):
-
What (Data) – e.g., Customer entity.
-
How (Function) – e.g., Order processing.
-
Where (Network) – e.g., Data center locations.
-
Who (People) – e.g., Roles.
-
When (Time) – e.g., Business cycles.
-
Why (Motivation) – e.g., Business goals.
Example (University):
-
Row 2 (Owner), Column 1 (What): "Student" as a business entity.
-
Row 4 (Builder), Column 2 (How): "Student enrollment" implemented via
enroll()API.
[!TIP] Remember: Zachman is descriptive, not prescriptive. It classifies artifacts; it doesn't tell you how to build.
[[END OF UNIT 4 NOTES]]