Skip to content
CY-703 (C) · Data Engineering/Quick Revision Short Notes

Data Engineering (CY-703 (C)) - Unit 1 Short Notes

UNIT 1: FOUNDATIONS & ARCHITECTURES OF DATA ENGINEERING


1.0 INTRODUCTION & CORE CONCEPTS

1.1 The Data Engineering Discipline & Lifecycle

  • Definition: Data Engineering is the discipline of designing, building, and maintaining the infrastructure, systems, and pipelines that enable the collection, storage, processing, and serving of data for analytical and operational use.

  • The Data Engineering Lifecycle (Core Pipeline):

    1. Ingestion: Acquiring data from various source systems.

    2. Storage: Persisting data in appropriate systems (lakes, warehouses, databases).

    3. Processing: Transforming, cleaning, and preparing data (ETL/ELT).

    4. Serving: Making processed data available to consumers via APIs, dashboards, or databases.

    5. Consumption: End-users (analysts, scientists, apps) utilizing the served data.

  • Evolving Role: From monolithic ETL developer to cloud-native, scalable pipeline architect who enables self-service analytics and ML/AI.

1.2 Fundamental Data Characteristics: The Five V's

V Definition Real-World Example
Volume Scale of data generated/collected. Terabytes/Petabytes of IoT sensor data, social media posts.
Velocity Speed of data generation, ingestion, and processing. Real-time stock ticker data, fraud detection alerts.
Variety Different formats and structures of data. Structured: SQL tables. Semi-structured: JSON logs, XML. Unstructured: Images, videos, PDFs.
Veracity Quality, accuracy, trustworthiness, and uncertainty of data. Incomplete customer records, noisy sensor readings, biased survey data.
Value The ultimate goal: extracting meaningful insights and business value. Predictive maintenance from sensor data, personalized recommendations.

[!TIP] Exam Focus: The Five V's are a very high-frequency question. Be ready to define each and give a distinct, domain-specific example (e.g., healthcare, finance, retail).

1.3 Data Types and Sources

  • By Structure:

    • Structured: Fixed schema, rows & columns (e.g., relational databases).

    • Semi-structured: Tags/schemas but not rigid (e.g., JSON, XML, CSV with headers).

    • Unstructured: No predefined model (e.g., text documents, images, audio, video).

  • By Origin:

    • Internal: Generated within the organization (e.g., transactional DBs, ERP/CRM logs).

    • External: Sourced from outside (e.g., public APIs, social media, government datasets, purchased data).

  • Common Source Systems: Relational DBs (MySQL, PostgreSQL), NoSQL DBs (MongoDB), APIs (REST, GraphQL), Application Logs, IoT Sensors, Social Media Feeds, Web Pages (scraping).


2.0 THE DATA ENGINEER'S ROLE & CONTEXT

2.1 Role in Data-Driven Decision Making

  • Collaboration: Works closely with:

    • Data Scientists: Provides clean, feature-rich, reliable datasets for model training.

    • Data Analysts/BI Developers: Builds data marts/warehouses and dashboards for reporting.

    • Business Stakeholders: Understands business problems to translate them into data requirements.

  • Enabling Self-Service: Builds robust, documented, and discoverable data platforms (Data Catalogs, semantic layers) so analysts can find and use data without constant engineering support.

  • Real-World Example: A retail data engineer builds a pipeline that ingests point-of-sale data, website clickstream logs, and inventory levels into a warehouse. An analyst then uses this integrated view to identify that a specific product's online sales drop when it's out of stock in nearby physical stores, leading to an optimized inventory allocation strategy.

2.2 Data Engineering in the Modern Enterprise

  • ML/AI Pipeline Integration: Data Engineers build the feature stores and data versioning systems that feed ML models. They ensure training/serving data skew is minimized.

  • Real-Time Analytics: Supports operational intelligence (e.g., dashboard for live logistics tracking, real-time fraud alerts) by implementing streaming architectures (Kafka, Kinesis).

  • Domain Understanding: Critical for designing relevant data models (e.g., star schema for retail sales, graph models for social networks) and choosing the right tools.


3.0 DATA ARCHITECTURE & STRATEGY

3.1 Modern Data Architectures (Cloud-Native)

  • Core Principles:

    • Decoupled Storage & Compute: Store data cheaply (e.g., S3) and spin up compute (e.g., Spark cluster) only when needed.

    • Scalability & Elasticity: Automatically scale resources up/down based on workload.

    • Pay-as-you-go: No large upfront capital expenditure.

    • Managed Services: Leverage cloud provider's managed services (Redshift, BigQuery) to reduce ops overhead.

  • Typical Components: Ingestion → Cloud Storage (Lake) → Processing/Transformation → Serving Layer (Warehouse/Marts) → Consumption (BI/ML).

  • Cloud Support (AWS Example):

    DiagramCANVAS: A flowchart showing: Source Systems -> (Kinesis/Kafka for stream, DMS/Glue for batch) -> Amazon S3 (Data Lake) -> (AWS Glue/EMR for processing) -> Amazon Redshift (Data Warehouse) -> (QuickSight, SageMaker, Apps for consumption).

3.2 Purpose-Built Storage: Data Lake vs. Data Warehouse

Feature Data Lake Data Warehouse
Data Type All data: raw, structured, semi-structured, unstructured. Primarily structured, processed data.
Schema Schema-on-Read: Apply structure when data is read. Schema-on-Write: Enforce structure before load.
Cost Very low-cost object storage (e.g., S3). Higher cost, optimized for compute.
Users Data scientists, engineers (exploratory analysis). Business analysts, executives (reporting, BI).
Optimization For storage cost and flexibility. For query performance and concurrency.
Examples Amazon S3, Azure Data Lake Storage (ADLS), Google Cloud Storage (GCS). Snowflake, Amazon Redshift, Google BigQuery, Azure Synapse.
Use Case Store raw data lake for future, unknown use cases. Serve aggregated, business-ready data for fast SQL queries.
  • Lakehouse Architecture: Modern hybrid pattern (e.g., Delta Lake, Apache Iceberg) that adds warehouse-like structures (ACID transactions, indexing) directly on top of data lake files.

3.3 Data Architecture Frameworks & Patterns

  • Lambda Architecture:

    • Batch Layer: Processes all historical data at rest (Hadoop/Spark). Provides accurate, comprehensive views.

    • Speed Layer: Processes real-time data streams (Spark Streaming, Flink). Provides low-latency views.

    • Serving Layer: Merges results from Batch & Speed layers for queries.

    • Use: Complex systems needing both accuracy and low-latency (e.g., real-time dashboard with historical context).

  • Kappa Architecture:

    • Simplifies Lambda by using only a single streaming layer for all processing.

    • All data is treated as a stream. Historical data is re-processed by replaying the stream.

    • Use: Simpler systems where real-time is primary, and batch can be simulated via stream replay.

  • Zachman Framework: A 6x6 matrix (Who, What, When, Where, Why, How) for describing an enterprise's architecture from different stakeholder perspectives (Planner, Owner, Designer, Builder, Subcontractor, Enterprise). It's a taxonomy, not a methodology.

  • Gartner Data Maturity Model: Stages an organization progresses through:

    1. Unaware/Reactive → 2. Aware/Consolidating → 3. Defined/Standardized → 4. Managed/Integrated → 5. Optimized/Innovative.

3.4 Modern Data Strategies

  • Data Mesh: Domain-oriented ownership. Data is treated as a product. Decentralized teams (domains) own their data end-to-end (from ingestion to serving), with a central platform providing self-serve infrastructure tools.

  • Data Fabric: An integrated, unified layer of data and metadata that sits on top of disparate sources and tools, providing a consistent, self-service experience for data discovery, governance, and consumption. Focuses on integration and abstraction.

  • Shift to Composable Stacks: Moving from monolithic, vendor-locked suites (e.g., old-school Informatica) to modular, best-of-breed tools that can be integrated (e.g., Fivetran for ingestion, dbt for transformation, Snowflake for storage, Looker for BI).


4.0 DATA INGESTION & ACQUISITION

4.1 Ingestion Patterns

Pattern Description Pros Cons Use Case
Batch Ingestion Periodic, large-volume loads (hourly, daily). Simple, cost-effective for large volumes, handles backfills. High latency (data not fresh). Daily sales reports, nightly data warehouse loads.
Streaming/Real-time Continuous, record-by-record flow. Low latency (seconds/milliseconds). Complex, requires stateful processing, higher cost. Fraud detection, live dashboards, IoT monitoring.
  • Scaling Stream Processing: Key considerations:

    • Throughput: Records/sec the system can handle.

    • Latency: Time from event generation to processing.

    • Fault Tolerance: Guarantees (at-most-once, at-least-once, exactly-once).

    • State Management: Handling windowed aggregations, joins.

    • Backpressure: Handling upstream surges.

4.2 Purpose-Built Data Ingestion Tools & Technologies

  • Change Data Capture (CDC): Captures row-level changes (INSERT, UPDATE, DELETE) in a source database and streams them in real-time. Tools: Debezium, AWS DMS, Qlik Replicate.

  • Webhooks: User-defined HTTP callbacks triggered by an event.

    • Mechanism: Application A (e.g., GitHub) sends an HTTP POST to a pre-registered URL (your service) when an event occurs (e.g., code push).

    • Example: Receive a notification in your Slack channel when a GitHub issue is created.

  • Web Scraping: Programmatically extracting data from websites.

    • Tools: BeautifulSoup (Python, parsing HTML/XML), Scrapy (Python, full framework), Selenium (browser automation).

    • Ethical/Legal: Check robots.txt, respect rate limits, review Terms of Service, avoid copyrighted/scraping-protected data.

    • Example Scenario (RGPV): To find representatives with press releases about "data", you would:

      1. Identify government websites (e.g., data.gov.in, MP websites).

      2. Use a crawler (Scrapy) to find pages with "press release" in the URL/title.

      3. Parse each press release page and search for the keyword "data" in the text.

      4. Extract representative name and release date/URL into a structured CSV/DB.

4.3 Data Transfer & Security Protocols

  • Secure Copy Protocol (SCP): Uses SSH for secure file transfer between hosts. Role: Simple, secure method for moving files (e.g., logs, CSVs) between on-prem servers or to cloud VMs.

  • Other Methods:

    • APIs (REST/GraphQL): Primary method for application-to-application data exchange.

    • FTP/SFTP: Traditional file transfer (SFTP is secure).

    • Kafka Connect: Framework for streaming data between Kafka and other systems (DBs, file systems).

    • AWS DMS: Fully managed service for database migration and CDC.


5.0 BIG DATA PROCESSING FRAMEWORKS

5.1 Apache Hadoop Ecosystem (Core)

  • HDFS (Hadoop Distributed File System): Distributed, fault-tolerant storage. Splits files into blocks (default 128MB/256MB) and replicates across DataNodes.

  • MapReduce: Programming model for processing large datasets.

    • Map: Processes input key-value pairs to generate intermediate key-value pairs.

    • Shuffle & Sort: Groups intermediate values by key.

    • Reduce: Aggregates intermediate values for each key.

  • YARN (Yet Another Resource Negotiator): Cluster resource manager. Schedules and manages resources for applications (MapReduce, Spark) running on the cluster.

  • Architecture:

    DiagramCANVAS: A diagram with a NameNode (master) managing metadata, connected to multiple DataNodes (slaves) storing actual HDFS blocks. A ResourceManager (YARN) connected to NodeManagers on each DataNode. Applications (MapReduce, Spark) submitted to RM, which allocates containers on NMs.

5.2 Apache Spark

  • Key Features:

    • In-Memory Processing: Caches data in RAM for iterative algorithms → 100x faster than MapReduce for many workloads.

    • DAG (Directed Acyclic Graph) Execution: Optimizes the entire workflow (not just map/reduce phases) into a single job.

    • Unified Engine: Same engine for Batch (Spark Core), Streaming (Spark Streaming/Structured Streaming), ML (MLlib), Graph (GraphX).

    • Ease of Use: Rich APIs in Python (PySpark), Scala, Java, R.

  • Core Components:

    • Spark Core: Base engine, RDDs (Resilient Distributed Datasets), task scheduling.

    • Spark SQL: Structured data processing with DataFrames/Datasets, SQL support.

    • Spark Streaming: Micro-batch processing (or Structured Streaming for true streaming).

    • MLlib: Scalable machine learning library.

  • Advantages over MapReduce:

    | Aspect | MapReduce | Spark | |---|---|---| | Speed | Disk I/O bound (writes intermediate to HDFS). | In-memory → 10-100x faster for iterative ML/ETL. | | Ease of Use | Complex Java API (map/reduce functions). | Simple high-level APIs (DataFrames, transform()). | | Iterative Processing | Very slow (multiple HDFS writes/reads). | Fast (keep data in memory across iterations). | | Real-time | Not designed for it. | Supports micro-batch & continuous processing. |

5.3 Cloud-Managed Processing Services

  • Amazon EMR (Elastic MapReduce):

    • Definition: Fully managed service for running Hadoop, Spark, HBase, Presto etc. on AWS EC2 instances.

    • Role in Simplifying:

      1. Provisioning: Automatically launches and configures a cluster in minutes.

      2. Scaling: Add/remove nodes automatically based on workload.

      3. Management: Handles node failures, software patching, monitoring integration (CloudWatch).

      4. Cost: Pay per second for EC2 instances used. Can use Spot Instances for huge savings.

    • Use Case: Run a large Spark job to process a month of web logs stored in S3, without managing any servers.


6.0 CLOUD PLATFORM INTEGRATION (AWS FOCUS)

6.1 AWS Core Data Services Landscape

Layer Service Purpose
Storage Amazon S3 Scalable object storage (Data Lake foundation).
Amazon EBS Block storage for EC2 (not for shared data).
Amazon Glacier Low-cost archival storage.
Processing Amazon EMR Managed Hadoop/Spark cluster.
AWS Glue Serverless ETL. Serverless Spark for data prep, crawlers for schema discovery.
Warehousing Amazon Redshift Fully managed, petabyte-scale data warehouse.
Orchestration AWS Step Functions Serverless workflow orchestration (state machines).
MWAA (Managed Airflow) Managed Apache Airflow for complex DAG orchestration.

6.2 AWS SageMaker for Machine Learning

  • Definition: Fully managed service to build, train, tune, debug, deploy, and monitor machine learning models at scale.

  • How it Supports Scalable ML:

    1. Build: SageMaker Studio (IDE) with built-in notebooks connected to S3.

    2. Train: Launch distributed training jobs on managed Spot/On-Demand clusters. Automatic model tuning (Hyperparameter Optimization).

    3. Deploy: One-click deployment to auto-scaled, secure HTTPS endpoints (real-time inference) or batch transform jobs.

    4. Monitor: Built-in model monitoring for data drift, concept drift.

  • ML Infrastructure on AWS Diagram:

    DiagramCANVAS: A diagram showing S3 as the central data lake. SageMaker Studio notebooks connected to S3. SageMaker Training jobs pulling data from S3, running on EMR/EC2 clusters, saving model artifacts back to S3. SageMaker Endpoints deployed on EC2/Auto Scaling group, pulling model from S3. SageMaker Ground Truth for labeling, Model Monitor checking endpoint data.


7.0 DATA PREPARATION & MACHINE LEARNING LIFECYCLE

7.1 Machine Learning Pipeline Stages

  1. Problem Definition: Frame the business problem as an ML task (classification, regression).

  2. Data Collection: Gather relevant data from sources (Data Engineer's primary role here).

  3. Data Pre-processing: Cleaning, handling missing values.

  4. Feature Engineering: Creating, selecting, transforming features (Data Engineer/Data Scientist collaboration).

  5. Model Selection & Training: Choosing algorithm, training on prepared dataset.

  6. Evaluation: Measuring model performance (accuracy, F1-score, RMSE).

  7. Deployment: Packaging model, creating inference endpoint (often Data Engineer builds the serving pipeline).

  8. Monitoring: Tracking model performance & data drift in production.

7.2 Data Pre-processing & Feature Engineering

  • Importance: "Garbage in, garbage out." Quality features are often more important than complex models. Can improve model accuracy by 10-30%.

  • Key Techniques:

    • Cleaning: Handle missing values (impute, drop), correct errors, remove duplicates.

    • Transformation: Normalization (min-max), Standardization (z-score), log transforms.

    • Encoding: Convert categoricals → numerical (One-Hot Encoding, Label Encoding).

    • Feature Extraction: Create new features from raw data (e.g., extracting "day of week" from timestamp, text vectorization (TF-IDF)).

    • Feature Selection: Remove irrelevant/redundant features (correlation analysis, mutual information) to reduce overfitting and cost.

  • Data Engineer's Role: Build automated, reproducible pipelines (using dbt, Glue, Spark) that implement these transformations at scale, ensuring the same logic is applied in training and inference.


8.0 DATA GOVERNANCE, QUALITY & OPERATIONS

8.1 Data Governance Fundamentals

  • Definition: The overall management of data availability, usability, integrity, and security in an enterprise.

  • Key Pillars:

    • Policies & Standards: Rules for data quality, naming, retention.

    • Ownership & Stewardship: Data Owners (business, accountable), Data Stewards (technical, implement rules).

    • Metadata Management: "Data about data" (table descriptions, column definitions, lineage).

    • Data Catalog: Searchable inventory of all data assets with business context (e.g., AWS Glue Data Catalog, Alation).

8.2 Data Lineage & Discovery

  • Data Lineage: Tracking data movement and transformation from source to consumption.

    • Why? Debug pipeline errors, assess impact of changes, comply with regulations (GDPR), understand data trust.

    • Lineage in DataOps: Automated, continuous lineage captured from orchestration tools (Airflow) and transformation code (dbt).

    • Pattern-Based Lineage: Infers lineage by analyzing code/query patterns. Example: A Spark job reads table_A, joins with table_B, writes to table_C. The tool parses the job code to draw A→C and B→C.

    • Lineage by Data Tagging: Uses metadata tags (e.g., source_system=CRM, pii=true) to establish relationships. If table_C has tag derived_from=table_A, lineage is inferred.

  • Data Discovery: Process of finding and understanding relevant data assets. Enabled by a good Data Catalog with search, tags, and sample data.

  • Data Wrangling: The hands-on process of cleaning, structuring, and enriching raw data. Often done interactively (Jupyter) before automating the pipeline.

8.3 Schema Management & Migration

  • Schema Evolution: Handling changes to data structure over time.

    • Forward Compatibility: New reader can read old data.

    • Backward Compatibility: Old reader can read new data.

    • Tools: Schema Registries (Confluent Schema Registry for Avro/Protobuf) enforce compatibility rules.

  • Schema Migration: The process of changing a database/warehouse schema (e.g., adding a column, changing a data type).

    • Example (Data Warehouse): Adding a new customer_segment column to a dim_customer table.

      1. Plan: Assess impact on downstream reports/ETLs.

      2. Execute: ALTER TABLE dim_customer ADD COLUMN customer_segment VARCHAR(50);

      3. Backfill: Update ETL to populate new column for historical records.

      4. Validate: Ensure reports render correctly, no errors.

      5. Deploy: Roll out change to production.

8.4 Logging, Monitoring, and Alerting (DataOps)

  • Why? Ensure reliability, performance, and correctness of data pipelines.

  • Key Metrics to Monitor:

    • Freshness/Latency: Time since last successful run/data arrival.

    • Throughput: Records/sec processed.

    • Error Rates: Failed tasks, malformed records.

    • Resource Utilization: CPU, memory, disk I/O of cluster/nodes.

    • Data Quality: Row counts, null percentages, uniqueness checks.

  • Alerting: Set thresholds (e.g., "Alert if job fails" or "Alert if latency > 1 hour") and notify via Slack/Email/PagerDuty.

  • Tools: CloudWatch (AWS), Stackdriver (GCP), Prometheus/Grafana, Datadog, Airflow's built-in alerting.


9.0 SECURITY & DEVOPS FOR DATA

9.1 Securing Cloud Storage & Data (AWS S3 Example)

  • Shared Responsibility Model: AWS secures the infrastructure (physical, hypervisor). You secure your data, access, and configurations.

  • Key Security Mechanisms (Layered Defense):

    DiagramCANVAS: A layered diagram for S3 bucket security. At the center: S3 Bucket with Objects. Surrounding layers: 1. IAM Policies (user/role permissions). 2. Bucket Policies (resource-based). 3. Access Control Lists (ACLs - object-level, legacy). 4. Encryption (at rest: SSE-S3/SSE-KMS; in transit: HTTPS/TLS). 5. VPC Endpoints (private network access, no public internet). 6. CloudTrail for API audit logging.

    • IAM (Identity & Access Management): Define who (users/roles) can do what (actions like s3:GetObject) on which resources (bucket ARN). Principle of Least Privilege.

    • Bucket Policies: JSON policies attached directly to the S3 bucket, controlling access from specific IPs, requiring MFA, etc.

    • Encryption:

      • At Rest: Server-Side Encryption (SSE-S3, SSE-KMS, SSE-C). Client-Side Encryption.

      • In Transit: HTTPS/TLS (enforced by bucket policy).

    • VPC Endpoints (Gateway Endpoint for S3): Allows EC2 instances in a private VPC to access S3 without traversing the public internet (more secure, no data egress costs).

9.2 Infrastructure as Code & CI/CD for Data

  • Infrastructure as Code (IaC): Managing cloud infrastructure (clusters, buckets, IAM roles) using declarative configuration files.

    • Tools: Terraform (cloud-agnostic), AWS CloudFormation (AWS-native).

    • Benefit: Version-controlled, reproducible, consistent environments (Dev/Staging/Prod).

  • CI/CD (Continuous Integration/Continuous Deployment) for Data Pipelines:

    • CI: Automatically test code changes (e.g., run dbt test, unit tests for Spark jobs) on every Git commit.

    • CD: Automatically deploy approved changes to production (e.g., update Airflow DAGs, deploy new Glue job script).

    • Typical Pipeline: Git Commit → CI (Lint, Unit Test) → Build Artifact → CD (Deploy to Staging) → Integration Test → Manual Approval → Deploy to Prod.

    • Tools: Jenkins, GitLab CI/CD, GitHub Actions, Spinnaker.

    • Goal: Fast, reliable, auditable deployments of data infrastructure and code.


10.0 COMPARATIVE ANALYSIS & SYNTHESIS

10.1 ETL vs. ELT

Aspect ETL (Extract-Transform-Load) ELT (Extract-Load-Transform)
Flow Extract → Transform (in staging area) → Load to target. Extract → Load (raw) to target → Transform in target.
Transformation Location Separate staging server/engine (e.g., Informatica, custom code). Inside the target system (e.g., Snowflake, BigQuery, Redshift).
Data Loaded Only transformed, structured data. Raw, unprocessed data first.
Pros Good for complex transforms, legacy systems, strict data quality pre-load. Faster loads, leverages powerful cloud warehouse compute, preserves raw data, flexible for new queries.
Cons Slower (two-step), rigid schema, staging cost. Requires powerful/expensive target, raw data security concerns.
Modern Context Less common in new cloud builds. Dominant pattern in cloud data platforms (Lakehouse).

10.2 Batch vs. Stream Processing

Characteristic Batch Processing Stream Processing
Data Scope Finite, bounded dataset (e.g., a file, a day's data). Infinite, unbounded data stream.
Latency High (minutes to days). Low (milliseconds to seconds).
Volume Very high (process all at once). Continuous, moderate rate per time unit.
Processing Model Process entire dataset together. Process one record/window at a time.
Complexity Simpler (stateless or simple state). Complex (state management, windowing, fault tolerance).
Use Case Daily sales reports, end-of-day risk analysis. Fraud detection, live dashboards, sensor alerts.
Hybrid Lambda/Kappa Architectures combine both.

10.3 Technology Selection & Trade-offs

  • Storage Choice (Lake vs. Warehouse):

    • Choose Data Lake if: Cost is critical, data is raw/multi-structured, users are data scientists, schema is evolving.

    • Choose Data Warehouse if: Users need fast SQL, data is structured, concurrency is high, strong governance needed.

    • Modern Choice: Lakehouse (Delta Lake on S3, Iceberg) for best of both.

  • Processing Engine (Hadoop vs. Spark): Spark is almost always preferred for new projects due to speed (in-memory), ease of use (DataFrames), and unified capabilities. Hadoop/MapReduce is legacy for most new builds.

  • Cloud Service vs. Self-Managed:

    • Managed (EMR, Glue, Redshift): Faster setup, less ops overhead, integrated security, but less control, potentially higher cost at scale.

    • Self-Managed (Spark on EC2): Full control, potentially cheaper at huge scale, but requires deep expertise for setup, tuning, and maintenance.

  • Building Feasible/Scalable Infrastructure: Example: A startup wants to analyze user clickstream.

    1. Cost: Start with serverless (Glue, Athena on S3) to avoid cluster costs. As volume grows, evaluate provisioned (EMR) or dedicated (Redshift) for predictable high load.

    2. Performance: Use columnar formats (Parquet) on S3. For sub-second dashboards, use a warehouse (Redshift) with materialized views.

    3. Manageability: Choose managed services (Glue, MWAA) to keep small engineering team focused on logic, not ops.

    4. Skills: Leverage PySpark (Python) if team knows Python, not Scala/Java. Use dbt for transformation if team is SQL-oriented.

    \boxed{\text{Decision Framework: Start Serverless/Managed, Optimize for Query Patterns, Prioritize Team Skills, Re-architect at Scale.}}

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in