Skip to content
CY-703 (A) · Ethical Hacking/Quick Revision Short Notes

Ethical Hacking (CY-703 (A)) - Unit 4 Short Notes

UNIT 4: Data Engineering & Related Concepts (Based on Past Exam Analysis)


I. Foundational Concepts of Data Engineering

Definition and Scope of Data Engineering

Data Engineering is the discipline of designing, building, and maintaining the infrastructure and systems that enable the collection, storage, processing, and analysis of large-scale data. Its scope includes:

  • Developing data pipelines for efficient data flow.

  • Ensuring data quality, reliability, and scalability.

  • Supporting data science and analytics workloads.

  • Managing cloud and on-premise data platforms.

Data Engineering Lifecycle

A cyclical process with five core components:

Component Description
1. Ingestion Acquiring data from diverse sources (APIs, logs, IoT sensors).
2. Storage Persisting data in databases, data lakes, or warehouses.
3. Processing Transforming raw data (cleaning, aggregating) using batch/streaming engines.
4. Serving Making processed data available to users/applications via APIs or query interfaces.
5. Consumption End-users (analysts, ML models) accessing data for insights.

[!TIP] Exam Focus: Be prepared to draw and label this lifecycle diagram. Link each component to real-world tools (e.g., Ingestion: Apache Kafka; Storage: S3; Processing: Spark).


II. Data Architectures and Storage Patterns

Types of Data Architectures

Architecture Advantages Disadvantages
Centralized (Monolithic) Simple management, strong consistency. Single point of failure, scalability limits.
Decentralized (Data Mesh) Domain-oriented ownership, scalability, agility. Complex governance, data duplication risks.
Hybrid (Lakehouse) Combines cost-effective storage (lake) with structured querying (warehouse). Integration complexity, potential performance overhead.

Data Lake Patterns

A Data Lake stores raw, unstructured/semi-structured data at scale.

  • Merits:

    • Schema-on-read: Flexibility to define structure later.

    • Cost-effective: Uses cheap object storage (e.g., AWS S3).

    • Supports all data types: Logs, images, JSON, etc.

  • Applications:

    • Machine Learning: Raw data for feature engineering.

    • Exploratory Analytics: Ad-hoc queries on unprocessed data.

    • Regulatory Archiving: Immutable storage for compliance.

[!TIP] Common Pitfall: Do not confuse Data Lake (raw storage) with Data Warehouse (processed, structured). Lakehouse patterns aim to bridge this gap.


III. Data Processing Frameworks & Streaming

Lambda Architecture

A fault-tolerant, low-latency system for processing both real-time and batch data.

Step-by-Step Real-Time Processing:

  1. Batch Layer: Precomputes batch views from all historical data (e.g., Hadoop MapReduce).

  2. Speed Layer: Processes real-time data streams to compensate for batch latency (e.g., Apache Storm, Spark Streaming). Creates real-time views.

  3. Serving Layer: Merges batch views and real-time views to serve complete, up-to-date query results.

[!TIP] Exam Key: Lambda is complex due to dual codebases (batch + speed). Mention trade-offs: accuracy vs. latency.

Kappa Architecture

A simplified alternative to Lambda:

  • Single processing pipeline: Uses only a streaming layer for all data (historical + real-time).

  • Historical data is replayed from storage (e.g., Kafka) through the same stream processor.

  • Advantage: Simpler code maintenance; Disadvantage: Requires powerful streaming engines for large historical reprocessing.

Streaming Data Architecture for Time-Series Data

Pattern Detection Methodology:

  1. Windowing: Sliding/tumbling windows over time-series streams (e.g., 5-minute windows).

  2. Aggregation: Compute metrics (avg, max, count) per window.

  3. Anomaly Detection: Apply statistical/machine learning models (e.g., Z-score, LSTM) on aggregated data.

  4. Alerting: Trigger alerts when patterns (spikes, dips) exceed thresholds.

Example: Detecting sudden CPU usage spikes in server logs using a 1-minute tumbling window and standard deviation > 3.


IV. Data Integration, Transformation, and Warehousing

Extract, Transform, Load (ETL) Process

Retail Database Example:

  • Extract: Pull sales and inventory tables from operational PostgreSQL DB nightly.

  • Transform:

    • Clean: Remove null customer_id.

    • Enrich: Join with product_dimension to add category.

    • Aggregate: Compute daily_sales_per_category.

  • Load: Write transformed data into fact_sales table in a Data Warehouse (e.g., Snowflake).

Schema Migration

Definition: Changing the structure of a database (tables, columns, relationships) while preserving data.

Example: Migrating from a denormalized single orders table (with repeated customer_name) to a normalized schema:

  1. Create new customers table (extract unique customer_id, name).

  2. Alter orders table to reference customer_id (foreign key).

  3. Backfill data and update application queries.

Data Warehouse Design (University Setup)

Step-by-Step Approach:

  1. Identify Business Processes: Student enrollment, course scheduling, fee collection.

  2. Define Facts: Numeric metrics (e.g., enrollment_count, fee_amount).

  3. Identify Dimensions: Student (attributes: major, year), Course, Time, Instructor.

  4. Design Star Schema:

    • Fact Table: fact_enrollment (keys: student_id, course_id, time_id; measures: credits).

    • Dimension Tables: dim_student, dim_course, etc.

  5. ETL Pipeline: Extract from university's SIS (Student Information System), transform, load into warehouse.

  6. BI Layer: Connect tools (Tableau) for dashboards (e.g., "Enrollment by Department").

[!TIP] Short Answer (3m): Focus on facts vs. dimensions and star schema for the university example.


V. Data Lineage, Governance, and Maturity

Data Lineage Tracking in DataOps

1. Pattern-Based Lineage:

  • Mechanism: Automatically infers lineage by analyzing data pipeline code (e.g., ETL scripts, Spark jobs).

  • Example: A Spark job reads raw_sales → filters → aggregates → writes daily_sales. Lineage tool (e.g., OpenLineage) parses the job DAG to map this flow.

  • Use Case: Debugging pipeline failures.

2. Lineage by Data Tagging:

  • Mechanism: Manually/automatically attaches metadata tags (e.g., PII, source:CRM) to datasets. Lineage is derived from tag propagation.

  • Example: Tag customer_email as PII in source CRM. When ETL copies it to data_warehouse.customers, the tag inherits. Shows impact analysis for GDPR compliance.

Gartner Data Maturity Model

Five stages of organizational data capability:

Stage Characteristics Example
1. Aware Data seen as IT cost, siloed. Reports generated manually per department.
2. Reactive Ad-hoc analytics, no standards. Spreadsheets for sales analysis, inconsistent metrics.
3. Defined Enterprise-wide strategy, basic governance. Central data team, documented data dictionary.
4. Managed Measurable quality, trusted sources. SLAs for data freshness, automated data quality checks.
5. Optimized Data as strategic asset, AI-driven. Real-time personalization using unified customer view.

Data Governance

Core Principles:

  • Accountability: Clear roles (Data Owners, Stewards).

  • Transparency: Documented policies, data lineage.

  • Integrity: Data quality standards (accuracy, completeness).

  • Compliance: Adherence to regulations (GDPR, HIPAA).

  • Security: Access controls, encryption.

Practices: Establish data council, implement catalog tools (Alation), define data quality rules, conduct regular audits.


VI. Data Acquisition, Transfer, and Extraction Techniques

Web Scraping

Hypothetical Scenario: Scrape https://www.rgpvonline.com for press releases containing "data".

Illustration:

  1. Identify URLs: Crawl site to find press release pages.

  2. Extract Content: Use Python (requests, BeautifulSoup) to parse HTML.

  3. Filter: Search for keyword "data" in title/body.

  4. Store: Save representative names, release dates, URLs in CSV/DB.

  5. Ethical/Legal: Check robots.txt, respect rate-limiting, avoid copyright infringement.

[!TIP] Exam Answer Structure: Tools → Steps → Challenges (dynamic JS pages, CAPTCHAs) → Ethics.

Webhooks

Mechanism: Event-driven HTTP callbacks. When event X occurs on Server A, it pushes an HTTP POST to a pre-configured URL (Server B).

Example:

  • GitHub Webhook: On push event, GitHub sends JSON payload to your CI/CD server (e.g., Jenkins) to trigger a build.

  • Payment Gateway: On payment success, gateway sends webhook to merchant's server to update order status.

Working:

  1. Subscriber registers URL with provider.

  2. Event occurs → provider sends HTTP request with event data.

  3. Subscriber processes payload (idempotently).

Secure Copy Protocol (SCP)

  • Overview: Uses SSH for secure file transfer between hosts.

  • Working: scp file.txt user@remote:/path encrypts file during transit via SSH tunnel.

  • Use Case: Transferring sensitive data (e.g., logs, configs) in a controlled network.

  • Limitation: No built-in directory synchronization; rsync over SSH preferred for advanced needs.


VII. Monitoring, Logging, and Alerting

Logging, Monitoring, and Alerting Systems

Working Principles:

  1. Logging: Record events (errors, transactions) with timestamps.

    • Example: Application logs JSON to stdout; collected by Fluentd.
  2. Monitoring: Collect & visualize metrics (CPU, latency, error rates).

    • Example: Prometheus scrapes metrics from app endpoints; Grafana dashboards.
  3. Alerting: Define thresholds (e.g., error rate > 5%) → trigger notifications (Slack, PagerDuty).

    • Example: Alertmanager sends alert if http_requests_failed spikes.

Example Stack (Modern):


App → Logs (JSON) → Fluentd → Elasticsearch → Kibana (Visualization)

      Metrics → Prometheus → Grafana

      Alerts → Alertmanager → Slack

[!TIP] Key Point: Emphasize the separation of concerns: logs (what happened?), metrics (how is it performing?), alerts (when to act?).


VIII. Enterprise Frameworks

Zachman Framework

A taxonomy for enterprise architecture using a 6x6 matrix.

Rows (Perspectives):

  1. Planner (Scope) – What business entities?

  2. Owner (Business) – What business processes?

  3. Designer (Architect) – What data models?

  4. Builder (Implementer) – What technologies?

  5. Subcontractor (Code) – What programs?

  6. Functioning Enterprise (Operations) – What outputs?

Columns (Abstractions):

  1. What (Data) – e.g., Customer entity.

  2. How (Function) – e.g., Order processing.

  3. Where (Network) – e.g., Data center locations.

  4. Who (People) – e.g., Roles.

  5. When (Time) – e.g., Business cycles.

  6. Why (Motivation) – e.g., Business goals.

Example (University):

  • Row 2 (Owner), Column 1 (What): "Student" as a business entity.

  • Row 4 (Builder), Column 2 (How): "Student enrollment" implemented via enroll() API.

[!TIP] Remember: Zachman is descriptive, not prescriptive. It classifies artifacts; it doesn't tell you how to build.


[[END OF UNIT 4 NOTES]]

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in