Skip to content
CY-703 (B) · Cyber Security Policies & Standards/Quick Revision Short Notes

Cyber Security Policies & Standards (CY-703 (B)) - Unit 4 Short Notes

UNIT 4: CYBER SECURITY POLICIES & STANDARDS IN DATA ENGINEERING

(Based on RGPV Past Paper Analysis: C DATA ENGINEERING - NOV 2023)


I. FOUNDATIONS OF DATA ENGINEERING & SECURITY CONTEXT

Data Engineering Lifecycle

The end-to-end process of designing, building, and managing data pipelines. Security policies must be integrated at each stage:

Stage Primary Activity Security Policy Implications
Collection Acquiring raw data Data provenance verification, source authentication, initial classification (PII/PHI)
Ingestion Moving data into systems Encryption in transit (TLS/SSL), secure API gateways, input validation against injection
Storage Persisting data Encryption at rest, access controls (RBAC/ABAC), data masking for sensitive fields
Processing Transformation & analysis Secure runtime environments, principle of least privilege for compute resources
Serving Preparing data for use API security, query-level access controls, audit logging of data access
Consumption End-user or application use Output filtering, DLP (Data Loss Prevention) enforcement, user consent management

[!TIP] Exam Focus: Be prepared to map a specific security standard (e.g., GDPR, PCI-DSS) to lifecycle stages. For example, GDPR's "right to erasure" impacts Storage and Consumption policies.

DataOps & Operational Excellence

  • Definition: Extension of DevOps principles to data pipelines, emphasizing automation, monitoring, and collaboration.

  • Security Integration:

    • Automated Compliance: Embedding policy-as-code (e.g., using Open Policy Agent) into CI/CD pipelines to check for insecure configurations before deployment.

    • Secure Pipeline Design: Isolating development, test, and production environments; secrets management (e.g., HashiCorp Vault) in pipeline scripts.

    • Continuous Security: Automated vulnerability scanning of data containers and dependencies.


II. DATA LINEAGE, TRACKING & SECURITY AUDITING

Data Lineage Tracking in DataOps

  • Definition: The ability to trace the flow of data from its origin through various transformations to its final destination.

  • Security Relevance:

    • Audit & Compliance: Provides immutable audit trails for regulations (GDPR Art. 30, HIPAA). Shows who accessed what data and when.

    • Breach Forensics: Quickly identifies affected datasets and systems during a security incident.

    • Impact Analysis: Assesses the blast radius of a policy change or data corruption.

  • Example: Tracking a customer's PII (e.g., email) from a web form → raw data lake (encrypted) → transformation (masked for analytics) → data warehouse (aggregated reports). Lineage proves masking was applied.

Pattern-Based Lineage vs. Lineage by Data Tagging

Feature Pattern-Based Lineage Lineage by Data Tagging
Method Automated analysis of code, queries, job configs Manual/metadata-driven; tags applied to datasets
Effort Low (after initial setup) High (requires consistent manual tagging)
Accuracy High for structured flows; may miss ad-hoc queries High if tags are comprehensive; prone to gaps
Security Application Automated anomaly detection (e.g., unexpected PII flow to public bucket) Policy-based enforcement (e.g., block any job moving "Confidential" tagged data to unencrypted storage)
Tool Examples OpenLineage, Marquez Collibra, Alation

[!TIP] Common Pitfall: Relying solely on tagging for lineage. Untagged "shadow IT" data flows create security blind spots. Pattern-based is essential for comprehensive coverage.


III. DATA MATURITY MODELS & GOVERNANCE FRAMEWORKS

Gartner Data Maturity Model

A five-level model assessing an organization's data management capabilities. Security maturity is a critical dimension.

Level Characteristics Security Integration Example
Basic Siloed, reactive, no standards Ad-hoc firewalls; no data classification
Emerging Initial standards, some collaboration Basic access controls; perimeter security only
Developing Defined processes, cross-functional teams Data classification policies enacted; centralized logging
Established Measured, trusted, business-aligned Automated policy enforcement (e.g., DLP, encryption key management); regular audits
Optimized Innovative, predictive, self-service AI-driven anomaly detection; security embedded in all data products ("security by design")

\boxed{\text{Security Maturity} \propto \text{Data Governance Maturity}}

Higher data maturity enables proactive, integrated security controls.

Zachman Framework

An enterprise architecture framework using a 6x6 matrix (What, How, Where, Who, When, Why × Planner, Owner, Designer, Builder, Subcontractor, Functioning Enterprise).
Security Mapping Example:

  • "Who" (Roles): Defines access control roles (Data Owner, Steward, User).

  • "How" (Processes): Maps to secure data processing policies (e.g., transformation rules that enforce masking).

  • "Where" (Location): Corresponds to data residency and storage security policies (e.g., GDPR data location constraints).

  • "Why" (Motivation): Aligns with compliance drivers (e.g., "Why is this data retained?" → Legal hold policies).

Data Governance

  • Definition: The overall management of data availability, usability, integrity, and security.

  • Core Components & Security Focus:

    • Policies & Standards: Formal documents (e.g., "All PII must be encrypted at rest").

    • Roles:

      • Data Owner: Business executive accountable for data security & compliance.

      • Data Steward: Implements policies (e.g., defines classification labels).

      • Data Custodian: IT role managing technical security (e.g., DB admins managing encryption keys).

    • Processes: Data classification, risk assessment, incident response for data breaches.

  • Security Outcome: Clear accountability for data security, reducing ambiguity during incidents.


IV. DATA ARCHITECTURES & SECURITY IMPLICATIONS

Types of Data Architectures

Architecture Description Advantages Disadvantages / Security Risks
Data Warehouse Structured, schema-on-write, optimized for SQL Strong consistency, mature security tools Rigid schema; expensive scaling; siloed security per warehouse
Data Lake Raw, unstructured/semi-structured, schema-on-read Scalable, flexible, cheap storage "Data Swamp" risk: Lack of governance → sensitive data exposure; complex access control
Lakehouse Hybrid: Lake's storage + Warehouse's management Unified analytics, ACID transactions Newer tech stack; security models still evolving
Hybrid Combination (e.g., lake for raw, warehouse for curated) Best of both worlds Increased attack surface; complex security policy orchestration across systems

Data Lake Patterns

  • Merits: Store everything in native format (logs, JSON, images). Enables centralized security monitoring (e.g., all application logs in one lake for SIEM).

  • Applications: Security use cases like storing threat intelligence feeds, forensic snapshots, or privacy-preserving raw data (encrypted).

  • Security Challenges:

    • Unstructured Data Risks: Hard to apply uniform access controls or discover sensitive data.

    • Solution: Implement fine-grained access control (e.g., AWS Lake Formation, Azure Purview) and automated data classification on ingestion.

Lambda Architecture

Processes both batch (historical, accurate) and speed (real-time, approximate) layers.
Real-Time Security Processing:

  1. Batch Layer: Re-trains fraud detection models on full historical data nightly.

  2. Speed Layer: Applies real-time anomaly detection on streaming transactions (e.g., using Apache Flink + ML).

  3. Serving Layer: Merges results. Security policy: All layers must enforce consistent access controls; batch re-processing must not overwrite security-critical audit logs.

Kappa Architecture

Simplified: stream-only processing. All data treated as a stream.
Security Considerations:

  • Continuous Monitoring: Stream processors (Kafka Streams, Flink) can apply real-time security rules (e.g., "Block any stream containing credit card numbers from leaving the network").

  • Checkpoint Security: Stream checkpoints (state) must be encrypted to prevent tampering.

  • Simplified Policy Stack: One set of security rules for all data, reducing configuration drift.


V. DATA MANAGEMENT PROCESSES & COMPLIANCE

Schema Migration

  • Definition: Changing the structure (schema) of a database or data system over time (e.g., adding a column, changing data type).

  • Security Impact:

    • Data Integrity: Migration scripts must preserve encryption and integrity constraints. A naive ALTER TABLE might break column-level encryption.

    • Policy Updates: Access control policies (e.g., row-level security) must be updated to match new schema.

    • Example: Migrating a customers table to add an encrypted_ssn column. Steps:

      1. Backup with encryption.

      2. Add column with encryption key management.

      3. Migrate data using secure transformation (decrypt old, encrypt new).

      4. Update all dependent ETL jobs and access policies.

      5. Test access controls on new schema.

ETL (Extract, Transform, Load) with Security

Secure ETL Pipeline Example (Retail → Data Warehouse for PCI-DSS compliance):

  1. Extract: Pull from POS systems via mutual TLS (mTLS). Authenticate source systems.

  2. Transform:

    • Data Masking: Replace PAN (Primary Account Number) with token for non-payment analytics.

    • Validation: Reject records failing integrity checks to prevent injection.

    • Encryption: Apply application-layer encryption to sensitive fields before loading.

  3. Load: Write to data warehouse using encrypted connections. Warehouse enforces column-level access control (e.g., only fraud team can see raw PAN).

Building a Data Warehouse: University Case Study

Step-by-Step with Security Policies (FERPA Compliance):

  1. Requirements: Identify Protected Information (student grades, SSN). Define access roles (Registrar, Faculty, Student).

  2. Design:

    • Schema: Separate PII tables from analytics tables.

    • Security Model: Role-Based Access Control (RBAC). Implement row-level security (student sees only their records).

  3. ETL Development:

    • Extract from SIS (Student Information System) via secure API.

    • Transform: Hash student IDs for analytics; mask SSNs.

    • Load to warehouse with audit triggers on all PII tables.

  4. Reporting: All reports must go through approved BI tools with integrated authentication (e.g., SAML). Query-level logging mandatory.

  5. Policy: Data retention policy (e.g., transcripts kept forever, application logs 7 years). Incident response plan for unauthorized access.


VI. SECURE DATA TRANSFER & INTEGRATION TECHNOLOGIES

Secure Copy Protocol (SCP)

  • Function: Secure file transfer based on SSH. Uses SSH for authentication and encryption.

  • How it Works: scp file user@remote:/path → SSH connection established → file encrypted via symmetric cipher (e.g., AES) → transferred.

  • Security Standards: Relies on SSH's strong encryption (AES, ChaCha20) and host/public key authentication.

  • vs. Insecure Alternatives: FTP sends credentials/data in plaintext. SCP is mandatory for transferring sensitive data between trusted hosts.

  • Use Case: Moving daily encrypted backups from a production server to a secure, air-gapped archival system.

Webhooks

  • Function: Event-driven HTTP callbacks. System A sends an HTTP POST to System B's URL when an event occurs.

  • Security Risks:

    • Unauthenticated Triggers: Anyone who knows the URL can trigger the webhook.

    • Data Leakage: Sensitive payload sent over unencrypted HTTP.

    • Replay Attacks: Captured valid requests replayed later.

  • Mitigations:

    1. Secret Tokens/Signatures: Payload signed (HMAC-SHA256). Receiver verifies signature using shared secret.

    2. HTTPS Only: Encrypts payload in transit.

    3. IP Whitelisting: Restrict incoming requests to known sender IPs.

    4. Timestamp + Nonce: Prevent replay.

  • Example: Payment gateway (Stripe) sends payment.succeeded webhook to merchant's server. Merchant verifies Stripe-Signature header using their webhook secret before processing.


VII. DATA SCRAPING, MONITORING & LEGAL POLICIES

Web Scraping & Regulatory Compliance

  • Technique: Programmatically extracting data from websites (parsing HTML/APIs).

  • Legal & Ethical Policy Landscape:

    • Terms of Service (ToS): Many sites prohibit scraping in ToS. Violation can lead to CFAA (Computer Fraud and Abuse Act) liability in US.

    • robots.txt: Ethical guideline (not law) indicating allowed/disallowed paths.

    • Data Protection Laws: GDPR/CCPA apply if scraping personal data of EU/California residents. Requires lawful basis (consent, legitimate interest) and honoring data subject rights.

    • Copyright: Scraped factual data may not be copyrighted, but compilation or database might be (EU Database Directive).

  • Scenario Analysis (rgpvonline.com):

    • Goal: Find press releases about "data" by representatives.

    • Policy Check:

      1. Check robots.txt for /press-releases.

      2. Review ToS for scraping prohibitions.

      3. If personal data (e.g., representative names/contact) is scraped, assess GDPR lawful basis. Government data may have open data licenses.

      4. Recommendation: Use official government APIs if available (more compliant). If scraping, limit rate, identify bot, and store data securely.

Logging, Monitoring, and Alerting (LMA)

  • Components:

    • Log Generation: Applications, OS, network devices emit logs (JSON, syslog).

    • Aggregation: Centralized log collection (e.g., SIEM - Splunk, Elastic, QRadar).

    • Monitoring: Real-time analysis of aggregated logs.

    • Alerting: Threshold-based or anomaly-based notifications (e.g., ">5 failed logins from same IP in 1 min").

  • Security Policies:

    • Log Retention: Define period (e.g., 90 days for security logs, 7 years for financial) per compliance requirement (SOX, PCI-DSS).

    • Integrity: Logs must be tamper-evident (write-once storage, hashing).

    • Alert Escalation: Runbooks defining who is paged, response time (e.g., P1 incident: 15-min response).

    • Incident Response Integration: Alerts must feed into IR ticketing system.

  • Example: Monitoring data warehouse query logs. Alert on:

    • SELECT * FROM customer_pii by non-privileged user.

    • Data export to external IP > 1GB.


VIII. ADVANCED TOPICS: REAL-TIME ANALYTICS & SECURITY

Pattern Detection in Time-Series Streaming Data

  • Streaming Architecture: Data sources → Message Broker (Kafka) → Stream Processor (Flink, Spark Streaming) → Alert/Sink.

  • Security Applications:

    • Fraud Detection: Detect unusual transaction patterns (e.g., "5 transactions in 2 countries in 5 mins").

    • Intrusion Detection (IDS): Identify port scan patterns or beaconing from malware.

  • Pattern Detection Techniques:

    • Sliding Window Aggregates: Count events in last 5 minutes.

    • Complex Event Processing (CEP): Define sequences (e.g., "failed login → success → data download").

    • ML Models: Real-time inference on feature vectors (e.g., isolation forest for anomaly score).

  • Policy Alignment: Real-time response playbooks must be pre-defined. Example: "If fraud score > 0.9, auto-block transaction AND alert SOC."

Real-Time Data Processing Security (Lambda/Kappa)

  • Common Security Requirements:

    • Stream Encryption: Use TLS between brokers and processors. Consider field-level encryption for sensitive data in stream.

    • Secure Checkpoints: Flink/Kafka Streams state checkpoints must be stored encrypted (e.g., S3 with SSE-KMS).

    • Authentication & Authorization: All components (producers, processors, consumers) must authenticate (mTLS, SASL) and be authorized to specific topics.

    • Compliance: Ensure data residency—streams do not cross prohibited geographic boundaries.

    • Audit Trail: Log all stream operations (topic creation, ACL changes) to immutable storage.

[!TIP] Exam Integration Question: "How would you secure a Lambda Architecture processing healthcare data for real-time patient monitoring?"

Answer Structure:

  1. Batch Layer: Encrypted historical data in data lake; ETL jobs run with minimal privileges.
  1. Speed Layer: Kafka topics with TLS & ACLs; Flink jobs apply real-time PHI masking; encrypted state checkpoints.
  1. Serving Layer: API gateway with OAuth2; row-level security on real-time view.
  1. Governance: Unified data lineage across both layers; audit logs from all layers aggregated to SIEM.

Final Exam Strategy:

  • For 7-mark questions, structure answer: Definition → 3-4 key points with examples → Security implication.

  • For scenario questions (e.g., web scraping, university DW), first state relevant policies/laws, then apply step-by-step.

  • Always link back to security policies/standards (GDPR, PCI-DSS, FERPA, NIST CSF) where possible.

  • Use schematics (tables, bullet lists) in your answer script for clarity and marks.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in