IT-704 Cloud Computing Lab - UNIT 5: Cloud-Native Technologies & DevOps Practices
5.1 Containerization Fundamentals
Core Concept: Containerization packages an application and its dependencies into a standardized unit that runs in an isolated user-space on a shared OS kernel. It provides process and filesystem isolation without the overhead of a full virtual machine (VM).
Container vs. Virtual Machine:
| Feature | Container | Virtual Machine (VM) |
|---|---|---|
| Isolation Level | Process-level (OS-level) | Full OS-level (hardware虚拟化) |
| Kernel | Shares host OS kernel | Has its own kernel (Guest OS) |
| Startup Time | Seconds | Minutes |
| Resource Overhead | Very Low | High (RAM, CPU for each OS) |
| Image Size | MBs | GBs |
| Portability | Very High | Moderate |
Key Underlying Technologies:
-
Namespaces: Provide isolation for process tree, network, mount points, users, etc. (e.g.,
pid,net,mntnamespaces). -
Control Groups (cgroups): Limit and account for resource usage (CPU, memory, disk I/O, network) of a collection of processes.
-
Union File Systems (UnionFS/OverlayFS): Enable layered images. Each Dockerfile instruction creates a new read-only layer. Final container adds a thin writable layer on top.
Docker Deep Dive
Docker Engine Architecture:
-
Docker Daemon (
dockerd): Long-running process managing Docker objects (images, containers, networks, volumes). -
Docker Client (
docker): CLI tool to interact with the daemon. -
Docker Registry: Stores Docker images (e.g., Docker Hub, private registries).
docker pull/pushinteracts with it.
Dockerfile Syntax & Best Practices:
-
FROM <image>:<tag>: Base image. Use specific tags (e.g.,python:3.9-slim) notlatest. -
RUN <command>: Execute commands during build. Chain with&& \to reduce layers. -
COPY <src> <dest>/ADD <src> <dest>: Copy files. PreferCOPYfor clarity. -
WORKDIR /path: Set working directory for subsequent instructions. -
ENV KEY=VAL: Set environment variables. -
EXPOSE <port>: Document which port the app listens on. -
CMD ["executable","param1"]orCMD command param1: Default command. Use exec form (JSON array) for proper signal handling. -
ENTRYPOINT ["executable","param1"]: Configure container as executable. Often used withCMDas default parameters. -
.dockerignore: Exclude files from build context (like.git,*.log, local cache).
Essential Commands:
# Build & Tag
docker build -t myapp:v1.0 .
docker tag myapp:v1.0 myrepo/myapp:v1.0
# Lifecycle
docker run -d -p 8080:80 --name mycontainer myapp:v1.0
docker ps -a # List all containers
docker stop mycontainer
docker rm mycontainer
docker rmi myapp:v1.0 # Remove image
# Data Persistence
docker volume create mydata
docker run -v mydata:/app/data ...
docker run -v /host/path:/container/path ... # Bind mount
# Networking
docker network create mynet
docker run --network mynet ...
docker network ls
Multi-Stage Builds (Optimization):
Use multiple FROM statements. Copy only necessary artifacts from a "builder" stage to a minimal "runtime" stage. Reduces final image size significantly.
# Stage 1: Builder
FROM golang:1.19 AS builder
WORKDIR /app
COPY . .
RUN go build -o myapp .
# Stage 2: Runtime
FROM alpine:latest
COPY --from=builder /app/myapp /usr/local/bin/
CMD ["myapp"]
Container Registries
-
Public: Docker Hub (default), GitHub Container Registry (GHCR).
-
Private/Cloud-Native: AWS ECR, Google GCR, Azure ACR. Provide integrated IAM, scanning, replication.
-
Commands:
docker login,docker push <registry>/<image>:<tag>,docker pull <registry>/<image>:<tag>.
[!TIP] Exam Focus: Know the difference between
CMDandENTRYPOINT. Understand how layers work in Docker images. Be able to write a simple multi-stage Dockerfile.
5.2 Container Orchestration with Kubernetes (K8s)
K8s Architecture (Master/Worker):
[[DIAGRAM: CANVAS: Draw a cluster with a Master Node (control plane) and 3 Worker Nodes.
Master components box: kube-apiserver (frontend), etcd (key-value store), kube-scheduler, kube-controller-manager.
Worker components box on each node: kubelet (node agent), kube-proxy (network proxy), Container Runtime (e.g., containerd).
Pods (1 or more containers) scheduled on Worker Nodes.]]
Core Concepts:
-
Pod: Smallest deployable unit. 1+ containers with shared network/IP and storage.
-
Deployment: Declarative management for Pods/ReplicaSets. Handles rolling updates, rollbacks, scaling.
-
Service: Stable network endpoint to access a set of Pods (load balancing). Types:
-
ClusterIP: Internal cluster access (default). -
NodePort: Exposes service on each Node's IP on a static port. -
LoadBalancer: Provisions cloud load balancer (cloud provider dependent).
-
-
Namespace: Virtual cluster partitioning (e.g.,
default,kube-system,production).
Declarative YAML Manifests (Key Fields):
apiVersion: apps/v1
kind: Deployment
metadata:
name: nginx-deploy
spec:
replicas: 3
selector:
matchLabels:
app: nginx
template: # Pod template
metadata:
labels:
app: nginx
spec:
containers:
- name: nginx
image: nginx:1.21
ports:
- containerPort: 80
---
apiVersion: v1
kind: Service
metadata:
name: nginx-service
spec:
selector:
app: nginx
ports:
- protocol: TCP
port: 80 # Service port
targetPort: 80 # Pod containerPort
type: ClusterIP
kubectl Essential Commands:
| Command | Purpose |
|---|---|
kubectl get pods,svc,deploy,ns |
List resources |
kubectl describe pod <pod-name> |
Detailed state/events (debug) |
kubectl apply -f <file.yaml> |
Create/Update declaratively |
kubectl delete -f <file.yaml> |
Delete resource |
kubectl logs <pod-name> [-c <container>] |
View container logs |
kubectl exec -it <pod-name> -- /bin/bash |
Execute command in container |
kubectl port-forward <pod-name> <local-port>:<pod-port> |
Forward local port to pod |
kubectl scale deploy <name> --replicas=N |
Manual scaling |
Service Discovery & Load Balancing: K8s DNS (CoreDNS) creates DNS records for Services (<service-name>.<namespace>.svc.cluster.local). kube-proxy on each node configures iptables/IPVS rules to route Service traffic to healthy Pod endpoints.
Scaling & Self-Healing:
-
Horizontal Pod Autoscaler (HPA): Automatically scales
replicasbased on CPU/memory/custom metrics.kubectl autoscale deployment <name> --cpu-percent=50 --min=2 --max=10 -
Rolling Updates:
Deploymentupdates Pods in a controlled manner (maxUnavailable, maxSurge).kubectl rollout status deploy/<name>. -
Rollback:
kubectl rollout undo deploy/<name>.
Storage Orchestration:
-
PersistentVolume (PV): Cluster-wide storage resource (provisioned by admin/cloud).
-
PersistentVolumeClaim (PVC): User request for storage (size, access mode). Binds to a suitable PV.
-
StorageClass: Defines "class" of storage (e.g.,
gp2on AWS,standardon GCE). Enables dynamic provisioning.
# PVC Example
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: mypvc
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 1Gi
storageClassName: standard
Pod uses PVC by mounting it in volumeMounts.
[!TIP] Exam Focus: Know the role of each K8s control plane component. Differentiate
ClusterIP,NodePort,LoadBalancer. Understand the PV-PVC binding process. Be able to write basic YAML for Deployment and Service.
5.3 Infrastructure as Code (IaC)
Concept: Managing infrastructure (networks, VMs, databases) through machine-readable definition files (code) rather than manual processes. Benefits: Version control, repeatability, collaboration, reduced drift.
Tool: Terraform (HashiCorp)
-
Language: HCL (HashiCorp Configuration Language). Declarative.
-
Core Workflow:
-
terraform init: Initialize working directory, download providers. -
terraform plan: Show execution plan (what will be created/changed/destroyed). -
terraform apply: Execute plan to provision infrastructure. -
terraform destroy: Destroy all managed infrastructure.
-
-
State Management: Terraform stores infrastructure state in
terraform.tfstate. Critical for mapping config to real resources. Remote backends (S3, Azure Storage) recommended for collaboration/state locking. -
Key Constructs:
-
Provider: Plugin for cloud/platform (AWS, Azure, GCP, etc.).
provider "aws" { region = "us-east-1" } -
Resource: Infrastructure component.
resource "aws_instance" "web" { ami = "ami-0c55b159cbfafe1f0" instance_type = "t2.micro" } -
Data Source: Read-only info from provider.
data "aws_ami" "ubuntu" { most_recent = true filter { name = "name" values = ["ubuntu/images/hvm-ssd/ubuntu-focal-20.04-amd64-server-*"] } } -
Variable & Output: Parameterize config, expose values.
-
Module: Reusable, encapsulated configuration unit.
-
Sample Terraform Config (AWS EC2):
variable "region" { default = "us-east-1" }
provider "aws" { region = var.region }
resource "aws_instance" "web_server" {
ami = "ami-0c55b159cbfafe1f0"
instance_type = "t2.micro"
tags = { Name = "WebServer" }
}
output "public_ip" {
value = aws_instance.web_server.public_ip
}
Cloud-Native Alternatives (Brief):
-
AWS CloudFormation: JSON/YAML templates. Native to AWS. Uses
Stacks. -
Azure Resource Manager (ARM) Templates: JSON templates. Native to Azure.
-
Google Cloud Deployment Manager: YAML/Python/Jinja2 templates.
[!TIP] Exam Focus: Understand the Terraform lifecycle commands (
init,plan,apply,destroy). Know the purpose ofstate. Be able to identify Provider, Resource, Data Source in a snippet.
5.4 CI/CD Pipelines in the Cloud
Pipeline Concepts: Automated process from code commit to production. Stages: Source → Build → Test → Deploy. Triggers (e.g., Git push), Agents/Runners (execution environment).
Tool: Jenkins
-
Architecture: Master (orchestrates) + Agents (execute builds). Plugins extend functionality.
-
Pipeline as Code: Define entire pipeline in a
Jenkinsfile(in repo). Declarative Syntax (recommended) vs. Scripted. -
Declarative Jenkinsfile Structure:
pipeline { agent any stages { stage('Build') { steps { sh 'docker build -t myapp .' } } stage('Test') { steps { sh 'docker run myapp npm test' } } stage('Deploy') { steps { script { // e.g., kubectl apply -f k8s-manifest.yaml sh 'kubectl set image deployment/myapp myapp=myrepo/myapp:${BUILD_ID}' } } } } post { always { echo 'Pipeline finished' } failure { mail to: '[email protected]', subject: 'Build Failed' } } } -
Integrations: Git plugin, Docker plugin, Kubernetes plugin (for dynamic agents/deployments).
Cloud-Native CI/CD Services:
-
AWS: CodeCommit (Source) → CodeBuild (Build) → CodeDeploy (Deploy) → CodePipeline (Orchestrate).
-
Azure: Azure Repos/GitHub → Azure Pipelines (Build/Release).
-
Google Cloud: Cloud Source Repositories/GitHub → Cloud Build → Cloud Deploy.
GitOps Paradigm:
-
Principle: Git repository is the single source of truth for desired infrastructure and application state.
-
Tooling (K8s): ArgoCD or Flux continuously monitor Git repos. When manifest changes are detected, they automatically sync the cluster state to match Git.
-
Workflow: Developer makes PR to Git → ArgoCD detects change → ArgoCD applies K8s manifests to cluster. Fully automated, auditable, reversible.
[!TIP] Exam Focus: Know Jenkinsfile structure (
pipeline,agent,stages,steps). Contrast traditional CI/CD with GitOps. Understand the role of agents/runners.
5.5 Serverless Computing & Function-as-a-Service (FaaS)
Concept: Event-driven execution model. Cloud provider dynamically manages server allocation, scaling, and provisioning. Developer deploys functions (code). Pay-per-execution (duration, memory). No infrastructure to manage.
Major Platforms Comparison:
| Feature | AWS Lambda | Azure Functions | Google Cloud Functions |
|---|---|---|---|
| Supported Languages | Node.js, Python, Java, Go, .NET, Ruby, custom (container) | C#, F#, JavaScript, TypeScript, Python, Java, PowerShell | Node.js, Python, Go, Java, .NET, Ruby |
| Max Duration | 15 minutes | 5 minutes (consumption), unlimited (App Service plan) | 9 minutes (2nd gen), 60s (1st gen) |
| Trigger Examples | API Gateway, S3, DynamoDB, SNS, SQS, CloudWatch Events | HTTP, Blob Storage, Queue, Event Hub, Timer | HTTP, Cloud Storage, Pub/Sub, Firestore |
| Packaging | ZIP, container image | ZIP, custom container | ZIP, source (gcloud) |
Key Components:
-
Function: Code + dependencies. Stateless. Initialization code runs on "cold start".
-
Trigger/Event: Source that invokes function (HTTP request, file upload, message in queue).
-
Execution Context: Reused environment for subsequent invocations (warm start). Initialize DB connections, SDK clients here.
-
Configuration: Memory (also determines CPU), timeout, environment variables, IAM role.
Development & Deployment:
# AWS Lambda (Python) example using AWS CLI
zip function.zip lambda_function.py
aws lambda create-function --function-name myfunc \
--runtime python3.9 --handler lambda_function.handler \
--role arn:aws:iam::123456789012:role/lambda-role \
--zip-file fileb://function.zip
# Invoke
aws lambda invoke --function-name myfunc output.json
Considerations & Trade-offs:
-
Cold Start: Latency when function initializes from scratch (first request or after idle period). Mitigate with provisioned concurrency (AWS), premium plan (Azure).
-
Vendor Lock-in: Tight integration with cloud provider's event sources and services. Portable via frameworks (Serverless Framework, SAM, CDK) but runtime is still provider-specific.
-
Debugging/Monitoring: Distributed tracing (X-Ray, Application Insights), structured logging to cloud watch services.
-
State: Functions are stateless. External state must be in databases (DynamoDB, Cosmos DB), object storage (S3), or caches (ElastiCache).
[!TIP] Exam Focus: Know the core concept: event-driven, pay-per-use, no server mgmt. Identify common triggers. Understand cold start problem and its impact. Differentiate between major providers' key limits (duration, memory).
5.6 Cloud-Native Monitoring, Logging & Observability
The Three Pillars of Observability:
-
Metrics: Quantitative measurements over time (CPU usage, request latency, error rate). PromQL for querying.
-
Logs: Timestamped records of discrete events (application logs, system logs).
-
Traces (Distributed Tracing): Track a request's path through multiple microservices. Jaeger, Zipkin.
Toolchain:
-
Prometheus: Pull-based metrics collection. Scrapes targets (applications, K8s nodes) at intervals. Stores data locally. Powerful query language PromQL.
-
Service Discovery: Automatically discovers K8s Pods/ Services via annotations.
-
Alerting:
Alertmanagerhandles alerts from Prometheus rules, deduplicates, sends to Slack/Email/PagerDuty.
-
-
Grafana: Visualization dashboard. Connects to Prometheus (and many other data sources) to create rich graphs and alerts.
-
ELK/EFK Stack (Centralized Logging):
-
Elasticsearch: Distributed search & analytics engine (stores/indexes logs).
-
Fluentd / Filebeat: Log collector/forwarder.
Fluentd(part of EFK) aggregates logs from various sources, processes, and sends to Elasticsearch.Filebeatis lightweight shipper. -
Kibana: Visualization and exploration for logs (like Grafana for logs).
-
-
Cloud Provider Services:
-
AWS: CloudWatch Metrics, CloudWatch Logs, X-Ray (traces).
-
Azure: Azure Monitor Metrics, Log Analytics, Application Insights.
-
Google Cloud: Cloud Monitoring (formerly Stackdriver), Cloud Logging, Cloud Trace.
-
Sample PromQL Queries:
# HTTP requests per second for a service (rate of increase of counter)
sum(rate(http_requests_total{service="api"}[5m]))
# 95th percentile latency of API requests
histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket{service="api"}[5m])) by (le))
# Memory usage of all pods in a namespace
sum(container_memory_usage_bytes{namespace="production"}) by (pod)
Lab Setup (K8s):
-
Deploy Prometheus, Alertmanager, Grafana (using Helm charts or manifests).
-
Configure Prometheus to scrape K8s API server, kube-state-metrics, node-exporter, and application pods (with
/metricsendpoint). -
Create Grafana dashboards using Prometheus data source.
-
Deploy Fluentd as DaemonSet to collect container logs from
/var/log/containers/*.logand ship to Elasticsearch. -
Access Kibana to search and visualize application logs.
[!TIP] Exam Focus: Differentiate Metrics, Logs, Traces. Know Prometheus is pull-based, ELK is push-based. Understand basic PromQL patterns (rate, histogram_quantile, sum by). Know the role of each EFK component.
5.7 Microservices Patterns & Service Mesh (Conceptual)
API Gateway Pattern:
-
Single Entry Point for all clients.
-
Responsibilities: Request routing, API composition, authentication/authorization, rate limiting, caching, monitoring.
-
Tools: AWS API Gateway, Kong, NGINX Ingress Controller (K8s), Spring Cloud Gateway.
-
K8s Ingress: Defines routing rules (host/path) to Services.
Ingress Controller(like NGINX) implements the rules.
Service Mesh:
-
Problem Solved: Handles service-to-service communication (east-west traffic) concerns transparently, decoupling them from business logic.
-
Architecture: Sidecar Proxy (Envoy, Linkerd-proxy) injected alongside each service pod. All inter-service traffic flows through the sidecar.
-
Key Features:
-
Traffic Management: Canary releases, A/B testing, fault injection, circuit breaking.
-
Security: mTLS (mutual TLS) for automatic encryption & identity.
-
Observability: Automatic metrics, logs, traces for all service calls.
-
Policy Enforcement: Rate limits, access controls.
-
-
Tools: Istio (most feature-rich), Linkerd (simpler, lighter).
-
Control Plane: Manages and configures all sidecars (e.g., Istio's
istiod).
Sidecar Proxy Pattern:
[[DIAGRAM: CANVAS: Show two Pods (Pod-A and Pod-B). Each Pod has two containers:
1. App Container (your microservice code)
2. Sidecar Proxy Container (e.g., Envoy)
Arrows show: App-A talks to localhost:15001 (Sidecar-A). Sidecar-A handles outbound traffic to Sidecar-B (over mTLS). Sidecar-B forwards to App-B on localhost:8080. All inbound traffic to Pod-A also goes through Sidecar-A first.]]
[!TIP] Exam Focus: Distinguish API Gateway (north-south, client-to-service) from Service Mesh (east-west, service-to-service). Know the sidecar pattern. List 2-3 key features of a service mesh (mTLS, traffic shifting, telemetry).
5.8 Practical Lab Scenarios & Troubleshooting
End-to-End Project Flow:
-
IaC: Write Terraform to provision cloud VPC, EKS/GKE/AKS cluster, RDS/Cloud SQL.
-
Containerization:
Dockerfilefor each microservice. Build & push images to registry (ECR/ACR). -
Orchestration: Write K8s YAML (Deployments, Services, Ingress, ConfigMaps, Secrets, PVCs). Apply to cluster.
-
CI/CD: Jenkins pipeline on Git push:
-
Checkout code.
-
Build Docker image, tag with commit hash.
-
Push to registry.
-
kubectl set image/kubectl apply -fto deploy new version.
-
-
Monitoring: Deploy Prometheus stack & EFK. Configure alerts.
Troubleshooting Cheatsheet:
| Problem Area | Key Commands / Checks |
|---|---|
| Pod Failing | kubectl describe pod <pod> (Check Events, State, Conditions) <br> kubectl logs <pod> --previous (if crashed) <br> kubectl get events --sort-by='.lastTimestamp' |
| Service Not Reachable | kubectl get svc,endpoints (Verify endpoints exist) <br> kubectl describe svc <svc> (Check selector matches Pod labels) <br> kubectl exec -it <pod> -- curl <svc-ip>:<port> (Test from inside cluster) |
| Image Pull Error | kubectl describe pod (Check ImagePullBackOff reason) <br> Verify image name/tag, registry credentials (imagePullSecrets), network egress. |
| Docker Build Fails | Check Dockerfile step-by-step. docker build --no-cache . to rule out cache issues. Verify base image exists, apt/yum commands succeed. |
| Terraform Apply Fails | terraform plan to see exact diff. Check state (terraform state list), provider credentials, resource dependencies/depends_on. |
| CI/CD Pipeline Fails | Check Jenkins console output. Common: Git credentials, Docker registry auth, kubectl context/credentials, syntax error in Jenkinsfile. |
| High Latency/Errors | Check Grafana/Prometheus for metrics (pod CPU/mem, request latency, error rate). Check application logs in Kibana for stack traces. |
General Strategy:
-
Isolate: Is it a pod, service, network, or external dependency issue?
-
Inspect: Use
describe,logs,getwith wide output (-o wide). -
Reproduce: Can you
execinto a pod and manually call the failing service? -
Rollback: If deployment caused issue,
kubectl rollout undo deploy/<name>. -
Check Dependencies: Database connectivity? External API? Secrets/ConfigMaps mounted correctly?
[!TIP] Exam Focus: Be prepared with a systematic debugging approach. Memorize critical
kubectl describeandkubectl logsusage. Understand common pod phases (Pending,Running,Failed,CrashLoopBackOff). Know how to check if a Service has endpoints.