UNIT 1: FOUNDATIONS & ARCHITECTURES OF DATA ENGINEERING
1.0 INTRODUCTION & CORE CONCEPTS
1.1 The Data Engineering Discipline & Lifecycle
-
Definition: Data Engineering is the discipline of designing, building, and maintaining the infrastructure, systems, and pipelines that enable the collection, storage, processing, and serving of data for analytical and operational use.
-
The Data Engineering Lifecycle (Core Pipeline):
-
Ingestion: Acquiring data from various source systems.
-
Storage: Persisting data in appropriate systems (lakes, warehouses, databases).
-
Processing: Transforming, cleaning, and preparing data (ETL/ELT).
-
Serving: Making processed data available to consumers via APIs, dashboards, or databases.
-
Consumption: End-users (analysts, scientists, apps) utilizing the served data.
-
-
Evolving Role: From monolithic ETL developer to cloud-native, scalable pipeline architect who enables self-service analytics and ML/AI.
1.2 Fundamental Data Characteristics: The Five V's
| V | Definition | Real-World Example |
|---|---|---|
| Volume | Scale of data generated/collected. | Terabytes/Petabytes of IoT sensor data, social media posts. |
| Velocity | Speed of data generation, ingestion, and processing. | Real-time stock ticker data, fraud detection alerts. |
| Variety | Different formats and structures of data. | Structured: SQL tables. Semi-structured: JSON logs, XML. Unstructured: Images, videos, PDFs. |
| Veracity | Quality, accuracy, trustworthiness, and uncertainty of data. | Incomplete customer records, noisy sensor readings, biased survey data. |
| Value | The ultimate goal: extracting meaningful insights and business value. | Predictive maintenance from sensor data, personalized recommendations. |
[!TIP] Exam Focus: The Five V's are a very high-frequency question. Be ready to define each and give a distinct, domain-specific example (e.g., healthcare, finance, retail).
1.3 Data Types and Sources
-
By Structure:
-
Structured: Fixed schema, rows & columns (e.g., relational databases).
-
Semi-structured: Tags/schemas but not rigid (e.g., JSON, XML, CSV with headers).
-
Unstructured: No predefined model (e.g., text documents, images, audio, video).
-
-
By Origin:
-
Internal: Generated within the organization (e.g., transactional DBs, ERP/CRM logs).
-
External: Sourced from outside (e.g., public APIs, social media, government datasets, purchased data).
-
-
Common Source Systems: Relational DBs (MySQL, PostgreSQL), NoSQL DBs (MongoDB), APIs (REST, GraphQL), Application Logs, IoT Sensors, Social Media Feeds, Web Pages (scraping).
2.0 THE DATA ENGINEER'S ROLE & CONTEXT
2.1 Role in Data-Driven Decision Making
-
Collaboration: Works closely with:
-
Data Scientists: Provides clean, feature-rich, reliable datasets for model training.
-
Data Analysts/BI Developers: Builds data marts/warehouses and dashboards for reporting.
-
Business Stakeholders: Understands business problems to translate them into data requirements.
-
-
Enabling Self-Service: Builds robust, documented, and discoverable data platforms (Data Catalogs, semantic layers) so analysts can find and use data without constant engineering support.
-
Real-World Example: A retail data engineer builds a pipeline that ingests point-of-sale data, website clickstream logs, and inventory levels into a warehouse. An analyst then uses this integrated view to identify that a specific product's online sales drop when it's out of stock in nearby physical stores, leading to an optimized inventory allocation strategy.
2.2 Data Engineering in the Modern Enterprise
-
ML/AI Pipeline Integration: Data Engineers build the feature stores and data versioning systems that feed ML models. They ensure training/serving data skew is minimized.
-
Real-Time Analytics: Supports operational intelligence (e.g., dashboard for live logistics tracking, real-time fraud alerts) by implementing streaming architectures (Kafka, Kinesis).
-
Domain Understanding: Critical for designing relevant data models (e.g., star schema for retail sales, graph models for social networks) and choosing the right tools.
3.0 DATA ARCHITECTURE & STRATEGY
3.1 Modern Data Architectures (Cloud-Native)
-
Core Principles:
-
Decoupled Storage & Compute: Store data cheaply (e.g., S3) and spin up compute (e.g., Spark cluster) only when needed.
-
Scalability & Elasticity: Automatically scale resources up/down based on workload.
-
Pay-as-you-go: No large upfront capital expenditure.
-
Managed Services: Leverage cloud provider's managed services (Redshift, BigQuery) to reduce ops overhead.
-
-
Typical Components: Ingestion → Cloud Storage (Lake) → Processing/Transformation → Serving Layer (Warehouse/Marts) → Consumption (BI/ML).
-
Cloud Support (AWS Example):
DiagramCANVAS: A flowchart showing: Source Systems -> (Kinesis/Kafka for stream, DMS/Glue for batch) -> Amazon S3 (Data Lake) -> (AWS Glue/EMR for processing) -> Amazon Redshift (Data Warehouse) -> (QuickSight, SageMaker, Apps for consumption).
3.2 Purpose-Built Storage: Data Lake vs. Data Warehouse
| Feature | Data Lake | Data Warehouse |
|---|---|---|
| Data Type | All data: raw, structured, semi-structured, unstructured. | Primarily structured, processed data. |
| Schema | Schema-on-Read: Apply structure when data is read. | Schema-on-Write: Enforce structure before load. |
| Cost | Very low-cost object storage (e.g., S3). | Higher cost, optimized for compute. |
| Users | Data scientists, engineers (exploratory analysis). | Business analysts, executives (reporting, BI). |
| Optimization | For storage cost and flexibility. | For query performance and concurrency. |
| Examples | Amazon S3, Azure Data Lake Storage (ADLS), Google Cloud Storage (GCS). | Snowflake, Amazon Redshift, Google BigQuery, Azure Synapse. |
| Use Case | Store raw data lake for future, unknown use cases. | Serve aggregated, business-ready data for fast SQL queries. |
- Lakehouse Architecture: Modern hybrid pattern (e.g., Delta Lake, Apache Iceberg) that adds warehouse-like structures (ACID transactions, indexing) directly on top of data lake files.
3.3 Data Architecture Frameworks & Patterns
-
Lambda Architecture:
-
Batch Layer: Processes all historical data at rest (Hadoop/Spark). Provides accurate, comprehensive views.
-
Speed Layer: Processes real-time data streams (Spark Streaming, Flink). Provides low-latency views.
-
Serving Layer: Merges results from Batch & Speed layers for queries.
-
Use: Complex systems needing both accuracy and low-latency (e.g., real-time dashboard with historical context).
-
-
Kappa Architecture:
-
Simplifies Lambda by using only a single streaming layer for all processing.
-
All data is treated as a stream. Historical data is re-processed by replaying the stream.
-
Use: Simpler systems where real-time is primary, and batch can be simulated via stream replay.
-
-
Zachman Framework: A 6x6 matrix (Who, What, When, Where, Why, How) for describing an enterprise's architecture from different stakeholder perspectives (Planner, Owner, Designer, Builder, Subcontractor, Enterprise). It's a taxonomy, not a methodology.
-
Gartner Data Maturity Model: Stages an organization progresses through:
- Unaware/Reactive → 2. Aware/Consolidating → 3. Defined/Standardized → 4. Managed/Integrated → 5. Optimized/Innovative.
3.4 Modern Data Strategies
-
Data Mesh: Domain-oriented ownership. Data is treated as a product. Decentralized teams (domains) own their data end-to-end (from ingestion to serving), with a central platform providing self-serve infrastructure tools.
-
Data Fabric: An integrated, unified layer of data and metadata that sits on top of disparate sources and tools, providing a consistent, self-service experience for data discovery, governance, and consumption. Focuses on integration and abstraction.
-
Shift to Composable Stacks: Moving from monolithic, vendor-locked suites (e.g., old-school Informatica) to modular, best-of-breed tools that can be integrated (e.g., Fivetran for ingestion, dbt for transformation, Snowflake for storage, Looker for BI).
4.0 DATA INGESTION & ACQUISITION
4.1 Ingestion Patterns
| Pattern | Description | Pros | Cons | Use Case |
|---|---|---|---|---|
| Batch Ingestion | Periodic, large-volume loads (hourly, daily). | Simple, cost-effective for large volumes, handles backfills. | High latency (data not fresh). | Daily sales reports, nightly data warehouse loads. |
| Streaming/Real-time | Continuous, record-by-record flow. | Low latency (seconds/milliseconds). | Complex, requires stateful processing, higher cost. | Fraud detection, live dashboards, IoT monitoring. |
-
Scaling Stream Processing: Key considerations:
-
Throughput: Records/sec the system can handle.
-
Latency: Time from event generation to processing.
-
Fault Tolerance: Guarantees (at-most-once, at-least-once, exactly-once).
-
State Management: Handling windowed aggregations, joins.
-
Backpressure: Handling upstream surges.
-
4.2 Purpose-Built Data Ingestion Tools & Technologies
-
Change Data Capture (CDC): Captures row-level changes (INSERT, UPDATE, DELETE) in a source database and streams them in real-time. Tools: Debezium, AWS DMS, Qlik Replicate.
-
Webhooks: User-defined HTTP callbacks triggered by an event.
-
Mechanism: Application A (e.g., GitHub) sends an HTTP POST to a pre-registered URL (your service) when an event occurs (e.g., code push).
-
Example: Receive a notification in your Slack channel when a GitHub issue is created.
-
-
Web Scraping: Programmatically extracting data from websites.
-
Tools:
BeautifulSoup(Python, parsing HTML/XML),Scrapy(Python, full framework),Selenium(browser automation). -
Ethical/Legal: Check
robots.txt, respectrate limits, review Terms of Service, avoid copyrighted/scraping-protected data. -
Example Scenario (RGPV): To find representatives with press releases about "data", you would:
-
Identify government websites (e.g.,
data.gov.in, MP websites). -
Use a crawler (Scrapy) to find pages with "press release" in the URL/title.
-
Parse each press release page and search for the keyword "data" in the text.
-
Extract representative name and release date/URL into a structured CSV/DB.
-
-
4.3 Data Transfer & Security Protocols
-
Secure Copy Protocol (SCP): Uses SSH for secure file transfer between hosts. Role: Simple, secure method for moving files (e.g., logs, CSVs) between on-prem servers or to cloud VMs.
-
Other Methods:
-
APIs (REST/GraphQL): Primary method for application-to-application data exchange.
-
FTP/SFTP: Traditional file transfer (SFTP is secure).
-
Kafka Connect: Framework for streaming data between Kafka and other systems (DBs, file systems).
-
AWS DMS: Fully managed service for database migration and CDC.
-
5.0 BIG DATA PROCESSING FRAMEWORKS
5.1 Apache Hadoop Ecosystem (Core)
-
HDFS (Hadoop Distributed File System): Distributed, fault-tolerant storage. Splits files into blocks (default 128MB/256MB) and replicates across DataNodes.
-
MapReduce: Programming model for processing large datasets.
-
Map: Processes input key-value pairs to generate intermediate key-value pairs.
-
Shuffle & Sort: Groups intermediate values by key.
-
Reduce: Aggregates intermediate values for each key.
-
-
YARN (Yet Another Resource Negotiator): Cluster resource manager. Schedules and manages resources for applications (MapReduce, Spark) running on the cluster.
-
Architecture:
DiagramCANVAS: A diagram with a NameNode (master) managing metadata, connected to multiple DataNodes (slaves) storing actual HDFS blocks. A ResourceManager (YARN) connected to NodeManagers on each DataNode. Applications (MapReduce, Spark) submitted to RM, which allocates containers on NMs.
5.2 Apache Spark
-
Key Features:
-
In-Memory Processing: Caches data in RAM for iterative algorithms → 100x faster than MapReduce for many workloads.
-
DAG (Directed Acyclic Graph) Execution: Optimizes the entire workflow (not just map/reduce phases) into a single job.
-
Unified Engine: Same engine for Batch (Spark Core), Streaming (Spark Streaming/Structured Streaming), ML (MLlib), Graph (GraphX).
-
Ease of Use: Rich APIs in Python (PySpark), Scala, Java, R.
-
-
Core Components:
-
Spark Core: Base engine, RDDs (Resilient Distributed Datasets), task scheduling.
-
Spark SQL: Structured data processing with DataFrames/Datasets, SQL support.
-
Spark Streaming: Micro-batch processing (or Structured Streaming for true streaming).
-
MLlib: Scalable machine learning library.
-
-
Advantages over MapReduce:
| Aspect | MapReduce | Spark | |---|---|---| | Speed | Disk I/O bound (writes intermediate to HDFS). | In-memory → 10-100x faster for iterative ML/ETL. | | Ease of Use | Complex Java API (map/reduce functions). | Simple high-level APIs (DataFrames,
transform()). | | Iterative Processing | Very slow (multiple HDFS writes/reads). | Fast (keep data in memory across iterations). | | Real-time | Not designed for it. | Supports micro-batch & continuous processing. |
5.3 Cloud-Managed Processing Services
-
Amazon EMR (Elastic MapReduce):
-
Definition: Fully managed service for running Hadoop, Spark, HBase, Presto etc. on AWS EC2 instances.
-
Role in Simplifying:
-
Provisioning: Automatically launches and configures a cluster in minutes.
-
Scaling: Add/remove nodes automatically based on workload.
-
Management: Handles node failures, software patching, monitoring integration (CloudWatch).
-
Cost: Pay per second for EC2 instances used. Can use Spot Instances for huge savings.
-
-
Use Case: Run a large Spark job to process a month of web logs stored in S3, without managing any servers.
-
6.0 CLOUD PLATFORM INTEGRATION (AWS FOCUS)
6.1 AWS Core Data Services Landscape
| Layer | Service | Purpose |
|---|---|---|
| Storage | Amazon S3 | Scalable object storage (Data Lake foundation). |
| Amazon EBS | Block storage for EC2 (not for shared data). | |
| Amazon Glacier | Low-cost archival storage. | |
| Processing | Amazon EMR | Managed Hadoop/Spark cluster. |
| AWS Glue | Serverless ETL. Serverless Spark for data prep, crawlers for schema discovery. | |
| Warehousing | Amazon Redshift | Fully managed, petabyte-scale data warehouse. |
| Orchestration | AWS Step Functions | Serverless workflow orchestration (state machines). |
| MWAA (Managed Airflow) | Managed Apache Airflow for complex DAG orchestration. |
6.2 AWS SageMaker for Machine Learning
-
Definition: Fully managed service to build, train, tune, debug, deploy, and monitor machine learning models at scale.
-
How it Supports Scalable ML:
-
Build: SageMaker Studio (IDE) with built-in notebooks connected to S3.
-
Train: Launch distributed training jobs on managed Spot/On-Demand clusters. Automatic model tuning (Hyperparameter Optimization).
-
Deploy: One-click deployment to auto-scaled, secure HTTPS endpoints (real-time inference) or batch transform jobs.
-
Monitor: Built-in model monitoring for data drift, concept drift.
-
-
ML Infrastructure on AWS Diagram:
DiagramCANVAS: A diagram showing S3 as the central data lake. SageMaker Studio notebooks connected to S3. SageMaker Training jobs pulling data from S3, running on EMR/EC2 clusters, saving model artifacts back to S3. SageMaker Endpoints deployed on EC2/Auto Scaling group, pulling model from S3. SageMaker Ground Truth for labeling, Model Monitor checking endpoint data.
7.0 DATA PREPARATION & MACHINE LEARNING LIFECYCLE
7.1 Machine Learning Pipeline Stages
-
Problem Definition: Frame the business problem as an ML task (classification, regression).
-
Data Collection: Gather relevant data from sources (Data Engineer's primary role here).
-
Data Pre-processing: Cleaning, handling missing values.
-
Feature Engineering: Creating, selecting, transforming features (Data Engineer/Data Scientist collaboration).
-
Model Selection & Training: Choosing algorithm, training on prepared dataset.
-
Evaluation: Measuring model performance (accuracy, F1-score, RMSE).
-
Deployment: Packaging model, creating inference endpoint (often Data Engineer builds the serving pipeline).
-
Monitoring: Tracking model performance & data drift in production.
7.2 Data Pre-processing & Feature Engineering
-
Importance: "Garbage in, garbage out." Quality features are often more important than complex models. Can improve model accuracy by 10-30%.
-
Key Techniques:
-
Cleaning: Handle missing values (impute, drop), correct errors, remove duplicates.
-
Transformation: Normalization (min-max), Standardization (z-score), log transforms.
-
Encoding: Convert categoricals → numerical (One-Hot Encoding, Label Encoding).
-
Feature Extraction: Create new features from raw data (e.g., extracting "day of week" from timestamp, text vectorization (TF-IDF)).
-
Feature Selection: Remove irrelevant/redundant features (correlation analysis, mutual information) to reduce overfitting and cost.
-
-
Data Engineer's Role: Build automated, reproducible pipelines (using dbt, Glue, Spark) that implement these transformations at scale, ensuring the same logic is applied in training and inference.
8.0 DATA GOVERNANCE, QUALITY & OPERATIONS
8.1 Data Governance Fundamentals
-
Definition: The overall management of data availability, usability, integrity, and security in an enterprise.
-
Key Pillars:
-
Policies & Standards: Rules for data quality, naming, retention.
-
Ownership & Stewardship: Data Owners (business, accountable), Data Stewards (technical, implement rules).
-
Metadata Management: "Data about data" (table descriptions, column definitions, lineage).
-
Data Catalog: Searchable inventory of all data assets with business context (e.g., AWS Glue Data Catalog, Alation).
-
8.2 Data Lineage & Discovery
-
Data Lineage: Tracking data movement and transformation from source to consumption.
-
Why? Debug pipeline errors, assess impact of changes, comply with regulations (GDPR), understand data trust.
-
Lineage in DataOps: Automated, continuous lineage captured from orchestration tools (Airflow) and transformation code (dbt).
-
Pattern-Based Lineage: Infers lineage by analyzing code/query patterns. Example: A Spark job reads
table_A, joins withtable_B, writes totable_C. The tool parses the job code to draw A→C and B→C. -
Lineage by Data Tagging: Uses metadata tags (e.g.,
source_system=CRM,pii=true) to establish relationships. Iftable_Chas tagderived_from=table_A, lineage is inferred.
-
-
Data Discovery: Process of finding and understanding relevant data assets. Enabled by a good Data Catalog with search, tags, and sample data.
-
Data Wrangling: The hands-on process of cleaning, structuring, and enriching raw data. Often done interactively (Jupyter) before automating the pipeline.
8.3 Schema Management & Migration
-
Schema Evolution: Handling changes to data structure over time.
-
Forward Compatibility: New reader can read old data.
-
Backward Compatibility: Old reader can read new data.
-
Tools: Schema Registries (Confluent Schema Registry for Avro/Protobuf) enforce compatibility rules.
-
-
Schema Migration: The process of changing a database/warehouse schema (e.g., adding a column, changing a data type).
-
Example (Data Warehouse): Adding a new
customer_segmentcolumn to adim_customertable.-
Plan: Assess impact on downstream reports/ETLs.
-
Execute:
ALTER TABLE dim_customer ADD COLUMN customer_segment VARCHAR(50); -
Backfill: Update ETL to populate new column for historical records.
-
Validate: Ensure reports render correctly, no errors.
-
Deploy: Roll out change to production.
-
-
8.4 Logging, Monitoring, and Alerting (DataOps)
-
Why? Ensure reliability, performance, and correctness of data pipelines.
-
Key Metrics to Monitor:
-
Freshness/Latency: Time since last successful run/data arrival.
-
Throughput: Records/sec processed.
-
Error Rates: Failed tasks, malformed records.
-
Resource Utilization: CPU, memory, disk I/O of cluster/nodes.
-
Data Quality: Row counts, null percentages, uniqueness checks.
-
-
Alerting: Set thresholds (e.g., "Alert if job fails" or "Alert if latency > 1 hour") and notify via Slack/Email/PagerDuty.
-
Tools: CloudWatch (AWS), Stackdriver (GCP), Prometheus/Grafana, Datadog, Airflow's built-in alerting.
9.0 SECURITY & DEVOPS FOR DATA
9.1 Securing Cloud Storage & Data (AWS S3 Example)
-
Shared Responsibility Model: AWS secures the infrastructure (physical, hypervisor). You secure your data, access, and configurations.
-
Key Security Mechanisms (Layered Defense):
DiagramCANVAS: A layered diagram for S3 bucket security. At the center: S3 Bucket with Objects. Surrounding layers: 1. IAM Policies (user/role permissions). 2. Bucket Policies (resource-based). 3. Access Control Lists (ACLs - object-level, legacy). 4. Encryption (at rest: SSE-S3/SSE-KMS; in transit: HTTPS/TLS). 5. VPC Endpoints (private network access, no public internet). 6. CloudTrail for API audit logging.-
IAM (Identity & Access Management): Define who (users/roles) can do what (actions like
s3:GetObject) on which resources (bucket ARN). Principle of Least Privilege. -
Bucket Policies: JSON policies attached directly to the S3 bucket, controlling access from specific IPs, requiring MFA, etc.
-
Encryption:
-
At Rest: Server-Side Encryption (SSE-S3, SSE-KMS, SSE-C). Client-Side Encryption.
-
In Transit: HTTPS/TLS (enforced by bucket policy).
-
-
VPC Endpoints (Gateway Endpoint for S3): Allows EC2 instances in a private VPC to access S3 without traversing the public internet (more secure, no data egress costs).
-
9.2 Infrastructure as Code & CI/CD for Data
-
Infrastructure as Code (IaC): Managing cloud infrastructure (clusters, buckets, IAM roles) using declarative configuration files.
-
Tools: Terraform (cloud-agnostic), AWS CloudFormation (AWS-native).
-
Benefit: Version-controlled, reproducible, consistent environments (Dev/Staging/Prod).
-
-
CI/CD (Continuous Integration/Continuous Deployment) for Data Pipelines:
-
CI: Automatically test code changes (e.g., run
dbt test, unit tests for Spark jobs) on every Git commit. -
CD: Automatically deploy approved changes to production (e.g., update Airflow DAGs, deploy new Glue job script).
-
Typical Pipeline:
Git Commit→CI (Lint, Unit Test)→Build Artifact→CD (Deploy to Staging)→Integration Test→Manual Approval→Deploy to Prod. -
Tools: Jenkins, GitLab CI/CD, GitHub Actions, Spinnaker.
-
Goal: Fast, reliable, auditable deployments of data infrastructure and code.
-
10.0 COMPARATIVE ANALYSIS & SYNTHESIS
10.1 ETL vs. ELT
| Aspect | ETL (Extract-Transform-Load) | ELT (Extract-Load-Transform) |
|---|---|---|
| Flow | Extract → Transform (in staging area) → Load to target. | Extract → Load (raw) to target → Transform in target. |
| Transformation Location | Separate staging server/engine (e.g., Informatica, custom code). | Inside the target system (e.g., Snowflake, BigQuery, Redshift). |
| Data Loaded | Only transformed, structured data. | Raw, unprocessed data first. |
| Pros | Good for complex transforms, legacy systems, strict data quality pre-load. | Faster loads, leverages powerful cloud warehouse compute, preserves raw data, flexible for new queries. |
| Cons | Slower (two-step), rigid schema, staging cost. | Requires powerful/expensive target, raw data security concerns. |
| Modern Context | Less common in new cloud builds. | Dominant pattern in cloud data platforms (Lakehouse). |
10.2 Batch vs. Stream Processing
| Characteristic | Batch Processing | Stream Processing |
|---|---|---|
| Data Scope | Finite, bounded dataset (e.g., a file, a day's data). | Infinite, unbounded data stream. |
| Latency | High (minutes to days). | Low (milliseconds to seconds). |
| Volume | Very high (process all at once). | Continuous, moderate rate per time unit. |
| Processing Model | Process entire dataset together. | Process one record/window at a time. |
| Complexity | Simpler (stateless or simple state). | Complex (state management, windowing, fault tolerance). |
| Use Case | Daily sales reports, end-of-day risk analysis. | Fraud detection, live dashboards, sensor alerts. |
| Hybrid | Lambda/Kappa Architectures combine both. |
10.3 Technology Selection & Trade-offs
-
Storage Choice (Lake vs. Warehouse):
-
Choose Data Lake if: Cost is critical, data is raw/multi-structured, users are data scientists, schema is evolving.
-
Choose Data Warehouse if: Users need fast SQL, data is structured, concurrency is high, strong governance needed.
-
Modern Choice: Lakehouse (Delta Lake on S3, Iceberg) for best of both.
-
-
Processing Engine (Hadoop vs. Spark): Spark is almost always preferred for new projects due to speed (in-memory), ease of use (DataFrames), and unified capabilities. Hadoop/MapReduce is legacy for most new builds.
-
Cloud Service vs. Self-Managed:
-
Managed (EMR, Glue, Redshift): Faster setup, less ops overhead, integrated security, but less control, potentially higher cost at scale.
-
Self-Managed (Spark on EC2): Full control, potentially cheaper at huge scale, but requires deep expertise for setup, tuning, and maintenance.
-
-
Building Feasible/Scalable Infrastructure: Example: A startup wants to analyze user clickstream.
-
Cost: Start with serverless (Glue, Athena on S3) to avoid cluster costs. As volume grows, evaluate provisioned (EMR) or dedicated (Redshift) for predictable high load.
-
Performance: Use columnar formats (Parquet) on S3. For sub-second dashboards, use a warehouse (Redshift) with materialized views.
-
Manageability: Choose managed services (Glue, MWAA) to keep small engineering team focused on logic, not ops.
-
Skills: Leverage PySpark (Python) if team knows Python, not Scala/Java. Use dbt for transformation if team is SQL-oriented.
\boxed{\text{Decision Framework: Start Serverless/Managed, Optimize for Query Patterns, Prioritize Team Skills, Re-architect at Scale.}}
-