UNIT 5: CLOUD COMPUTING & BIG DATA TECHNOLOGIES - EXAM-FOCUSED NOTES
A. CLOUD COMPUTING FUNDAMENTALS
Essential Characteristics of Cloud Computing
Defined by NIST, the five core characteristics are:
-
On-demand self-service: Users can provision computing resources (e.g., server time, network storage) automatically without human interaction.
-
Broad network access: Capabilities are available over the network and accessed through standard mechanisms (e.g., mobile phones, laptops).
-
Resource pooling: Provider's computing resources are pooled to serve multiple consumers using a multi-tenant model.
-
Rapid elasticity: Resources can be rapidly and elastically provisioned/released to scale with demand.
-
Measured service: Cloud systems automatically control and optimize resource use by leveraging a metering capability.
[!TIP] Exam Focus: Be prepared to differentiate these characteristics from traditional IT models. "Elasticity" is a key differentiator.
Service Models: Platform as a Service (PaaS)
-
Definition: A cloud computing model that provides a platform allowing customers to develop, run, and manage applications without the complexity of building and maintaining the underlying infrastructure (hardware, OS, middleware).
-
Essential Characteristics:
-
Includes development tools, database management, middleware, and OS.
-
Handles underlying infrastructure (servers, storage, networking).
-
Enables rapid application development and deployment.
-
-
Problems Solved:
-
Eliminates the cost and complexity of buying, installing, and managing hardware/software for application development.
-
Automates backend management (patches, updates, scaling).
-
Facilitates collaboration among development teams across geographies.
-
Examples: Google App Engine, Microsoft Azure App Services, Heroku.
-
Deployment Models: Community Cloud vs. Public Cloud
| Feature | Community Cloud | Public Cloud |
|---|---|---|
| Ownership & Management | Shared by several organizations with shared concerns (security, compliance, jurisdiction). | Owned and operated by a third-party cloud service provider (e.g., AWS, Azure, GCP). |
| Target Users | Specific community (e.g., government agencies, universities, research institutions). | General public / open to all. |
| Cost Model | Shared cost among community members; often more cost-effective than private cloud for the group. | Pay-as-you-go or subscription model; economies of scale for the provider. |
| Control & Customization | Higher control and customization to meet specific community needs/regulations. | Less control; standardized services and configurations. |
| Security | Tailored to the community's common security requirements. | Security is the provider's responsibility for the infrastructure; user responsibility for their data/apps. |
[!TIP] Common Pitfall: Do not confuse Community Cloud with Private Cloud. A private cloud is for a single organization; a community cloud is for a specific group of organizations.
B. CLOUD PLATFORMS & VIRTUALIZATION
Google App Engine (GAE)
-
Major Features:
-
PaaS offering: Supports multiple languages (Python, Java, Go, PHP, Node.js).
-
Automatic scaling: Scales applications automatically based on traffic.
-
Built-in services: Integrated services for datastore (NoSQL), memcache, task queues, and mail.
-
Managed runtime environment: Google manages servers, OS, and runtime.
-
Sandboxed environment: Applications run in a secure, sandboxed environment.
-
-
Types of Problems Solved:
-
Web applications and mobile backends.
-
APIs and microservices.
-
Applications with variable or unpredictable traffic patterns.
-
Prototyping and rapid development where infrastructure management is a overhead.
-
Microsoft Azure & Virtualization
-
How Virtualization is Employed:
-
Hyper-V: Azure's primary hypervisor. It creates and manages virtual machines (VMs) on physical servers.
-
Abstraction Layer: Virtualization abstracts physical hardware (CPU, memory, storage, networking) into logical resources.
-
Multi-tenancy: Enables multiple isolated VMs from different customers to run on the same physical host securely.
-
Resource Allocation & Isolation: Hyper-V manages resource allocation (CPU shares, memory limits) and ensures strong isolation between VMs.
-
Live Migration: Supports moving running VMs between physical hosts with minimal downtime for maintenance.
-
Eucalyptus
-
Features:
-
Open-source: Provides AWS-compatible cloud computing platform.
-
Compatibility: Implements AWS EC2, S3, and IAM APIs, allowing AWS workloads to be migrated or hybridized.
-
Components: Cloud Controller (CLC), Cluster Controller (CC), Storage Controller (SC), Node Controller (NC).
-
Deployment Flexibility: Can be deployed on-premise (private cloud) or in a hybrid model.
-
-
Modes of Operation:
-
Eucalyptus Mode: The cloud behaves exactly like AWS. Users use AWS-compatible tools (e.g., EC2 API tools) to manage instances.
-
Walrus Mode: Uses Eucalyptus's own object storage (Walrus) which is S3-compatible.
-
Hybrid Mode: A Eucalyptus cloud can be connected to an AWS public cloud, allowing workload bursting or migration.
-
C. CLOUD SECURITY & TRUST
Security Challenges in Cloud Computing
-
Data Breaches: Unauthorized access to sensitive data stored in the cloud.
-
Insecure Interfaces & APIs: Vulnerabilities in cloud service APIs can be exploited.
-
System Vulnerabilities & Exploits: Flaws in the hypervisor, OS, or shared technology components.
-
Account/Service Hijacking: Attackers steal credentials to gain control over cloud resources.
-
Denial of Service (DoS): Overwhelming cloud services to make them unavailable.
-
Data Loss: Permanent loss of data due to deletion, corruption, or disaster.
-
Shared Technology Vulnerabilities: Multi-tenancy risks where one tenant's activity affects another.
-
Insider Threats: Malicious or negligent actions by employees of the cloud provider or customer.
-
Lack of Visibility & Control: Customers have limited insight into the provider's physical security and operations.
-
Compliance & Legal Issues: Meeting regulatory requirements (GDPR, HIPAA) across different jurisdictions.
Cloud Computing Security Architecture (Block Diagram)
[[DIAGRAM: CANVAS: A layered block diagram showing:
-
User Layer: End-users accessing via web/mobile apps.
-
Application Layer: Cloud applications (SaaS) with security controls like authentication, authorization, encryption.
-
Platform Layer (PaaS): Development platforms with runtime security, API security.
-
Infrastructure Layer (IaaS): Virtualization layer (Hypervisor security), virtual networks, storage encryption.
-
Physical Layer: Data center security (biometrics, guards, environmental controls).
Arrows showing security mechanisms (Firewalls, IDS/IPS, Encryption, IAM, Logging/Monitoring) spanning across all layers. A "Cloud Security Broker" or "CASB" layer is often shown between User and Cloud Provider for enterprise visibility and control.]]
[!TIP] Exam Focus: Be able to describe the shared responsibility model: Provider secures the cloud infrastructure; Customer secures what they put in the cloud (data, apps, access).
Trusted Cloud Computing
-
Goal: To establish confidence that a cloud service provider will operate as promised and that the cloud environment is secure and reliable.
-
Key Aspects:
-
Trusted Computing Base (TCB): Minimal set of hardware/software components critical to security.
-
Attestation: Mechanism to verify the integrity of a platform (e.g., "This VM is running on a genuine, untampered hypervisor").
-
Secure Boot & Measured Boot: Ensuring only authenticated software loads during startup.
-
Hardware Roots of Trust: Using technologies like Intel SGX (Secure Guard Extensions) or TPM (Trusted Platform Module) to create isolated, secure execution environments.
-
SLAs with Security Guarantees: Contracts that include measurable security commitments.
-
D. CLOUD ECONOMICS, APPLICATIONS & CHALLENGES
Reducing Time-to-Market & Capital Expenses
-
Capital Expense (CapEx) Reduction:
-
Eliminates upfront costs for hardware, data center space, power, cooling.
-
Converts large fixed costs into smaller, variable operational expenses (OpEx).
-
Pay only for resources consumed (e.g., per hour, per GB).
-
-
Time-to-Market Reduction:
-
Instant Provisioning: Resources are available in minutes vs. weeks/months for procurement and setup.
-
Focus on Core Business: IT staff spend less time on infrastructure maintenance, more on application development.
-
Elastic Scalability: Easily handle launch spikes without over-provisioning.
-
Access to Advanced Services: Use managed databases, AI/ML tools, big data platforms without building expertise from scratch.
-
Advantages for Social Networking Applications
-
Massive Scalability: Handle viral growth and unpredictable user loads (e.g., trending topics).
-
Global Reach: Deploy application instances in multiple geographic regions via CDNs and cloud regions for low latency.
-
Cost-Effective Storage: Cheap, scalable object storage (like S3) for user-generated content (photos, videos).
-
Real-time Analytics: Process vast streams of user activity data for recommendations, ads, and insights using cloud analytics services.
-
High Availability: Built-in redundancy and failover mechanisms ensure 24/7 uptime critical for user engagement.
Specialized Cloud Types
-
Storage Cloud: Provides scalable, durable, and highly available object storage (e.g., AWS S3, Google Cloud Storage). Optimized for storing unstructured data like documents, backups, media files.
-
Cloud Analytics: Cloud-based platforms and services for processing and analyzing large datasets. Includes data warehouses (Redshift, BigQuery), data lakes, and managed Hadoop/Spark services. Enables insights without on-premise infrastructure.
Quality of Service (QoS) Issues in Cloud (with Examples)
-
Definition: The measurable service attributes like performance, availability, reliability, and capacity.
-
Key Issues & Examples:
-
Performance Variability/Noisy Neighbor: One tenant's heavy VM load on a shared physical host degrades performance for others. Example: A VM running a batch analytics job slows down a latency-sensitive web server on the same host.
-
Availability & Downtime: Scheduled maintenance or unplanned outages by the provider. Example: A cloud region outage makes all hosted applications inaccessible.
-
Network Latency & Bandwidth: Distance between user and cloud data center, or intra-cloud network congestion. Example: A real-time gaming application suffers lag because its servers are in a different continent.
-
Scalability Limits: Auto-scaling may have thresholds or delays. Example: A flash sale website cannot scale beyond the provider's account limits fast enough, leading to crashed pages.
-
Service Level Agreement (SLA) Penalties: Often capped at service credits, which may be less than actual business loss.
-
Elastic Computing & Storage Area Network (SAN)
-
Elastic Computing:
-
Definition: The ability to dynamically provision and de-provision computational resources (CPU, memory) in response to workload changes, in an automated, rapid, and fine-grained manner.
-
Key Enabler: Virtualization and orchestration tools (like Kubernetes, auto-scaling groups).
-
Formula (Conceptual): Elasticity = (Speed of Scaling Up/Down) + (Granularity of Scaling) + (Automation Level).
-
-
Storage Area Network (SAN):
-
Definition: A dedicated, high-speed network that provides access to consolidated, block-level storage. Makes storage devices (disk arrays) appear as locally attached devices to servers.
-
In Cloud Context: Cloud providers use massive, distributed SANs (often built on custom hardware and software-defined storage) to provide persistent block storage volumes (e.g., AWS EBS, Azure Disks) for VMs.
-
Benefit: Centralized management, high availability, and improved performance for I/O-intensive applications.
-
E. BIG DATA FUNDAMENTALS
What is Big Data?
- Definition: A term for data sets that are so large or complex that traditional data processing applications are inadequate to deal with them. It's characterized by the need for new architectures and techniques to extract value from the data.
Main Features / Characteristics (3 Vs - in Detail)
-
Volume: The sheer size of data.
-
Detail: Measured in terabytes (TB), petabytes (PB), exabytes (EB). Driven by IoT sensors, social media, transaction logs, multimedia.
-
Challenge: Storage cost, processing time, data management.
-
-
Variety: The different types of data from various sources.
-
Detail: Structured (relational DBs), Semi-structured (XML, JSON), Unstructured (text, images, video, audio). Variety complicates integration and analysis.
-
Challenge: Schema-on-read vs. schema-on-write, data integration.
-
-
Velocity: The speed at which data is generated, collected, and needs to be processed.
-
Detail: Real-time or near-real-time streams (e.g., stock ticks, sensor data, social media feeds). Requires stream processing frameworks.
-
Challenge: Need for low-latency processing, handling continuous data flows.
-
Extended Vs (Often mentioned): Veracity (uncertainty, quality), Value (extracting useful insights), Variability (changing meaning/format).
Challenges with Big Data
-
Storage: Storing massive volumes cost-effectively.
-
Processing: Analyzing data within acceptable timeframes (batch vs. real-time).
-
Data Integration & Variety: Combining disparate data sources and formats.
-
Data Quality & Veracity: Handling missing, noisy, or inconsistent data.
-
Scalability: Ensuring systems can grow seamlessly with data.
-
Security & Privacy: Protecting sensitive information at scale.
-
Visualization: Representing complex, large datasets meaningfully.
-
Skills Gap: Lack of professionals with data science and distributed computing expertise.
Real-World Applications of Big Data Analytics
-
Healthcare: Predictive analytics for patient readmissions, genomic research, drug discovery.
-
Retail & E-commerce: Recommendation engines (Amazon, Netflix), dynamic pricing, inventory optimization, customer sentiment analysis.
-
Finance: Fraud detection, algorithmic trading, risk management, customer segmentation.
-
Telecom: Network optimization, predicting churn, customer experience management.
-
Manufacturing (IoT): Predictive maintenance, supply chain optimization, quality control.
-
Transportation & Logistics: Route optimization, demand forecasting, fleet management.
-
Government: Crime prediction, smart city initiatives, tax revenue analysis.
F. BIG DATA ANALYTICS & PRE-PROCESSING
Data Cleaning
-
Definition: The process of identifying and correcting (or removing) corrupt, inaccurate, incomplete, or irrelevant data from a dataset.
-
Common Tasks:
-
Handling missing values (imputation, deletion).
-
Smoothing noisy data (binning, regression).
-
Resolving inconsistencies (e.g., "NY", "New York", "N.Y.").
-
Removing duplicates.
-
Validating data against known constraints.
-
-
Goal: Improve data quality for accurate and reliable analysis.
Data Sampling
-
Definition: The process of selecting a representative subset of data from a larger population to perform analysis, when processing the entire dataset is infeasible.
-
Common Techniques:
-
Simple Random Sampling: Every data point has equal chance.
-
Stratified Sampling: Population divided into strata (sub-groups); samples taken from each stratum proportionally.
-
Systematic Sampling: Select every k-th record from the dataset.
-
Cluster Sampling: Divide population into clusters; randomly select entire clusters.
-
-
Goal: Reduce computational cost while maintaining statistical significance.
Classification Techniques
Decision Trees (in detail)
-
Concept: A supervised learning method that uses a tree-like model of decisions and their possible consequences. It splits the dataset into smaller subsets based on the value of input features.
-
Key Terms:
-
Root Node: Represents the entire dataset.
-
Decision Node: A node that splits data based on a feature.
-
Leaf/Terminal Node: Represents a class label or output.
-
Branch/Sub-tree: A subsection of the tree.
-
Pruning: Removing sections of the tree to reduce overfitting.
-
-
Popular Algorithms: ID3 (uses Information Gain), C4.5 (uses Gain Ratio), CART (uses Gini Index).
-
Example: Classifying an animal as "Mammal" or "Bird" based on features like "Has Fur?", "Lays Eggs?".
-
Advantages: Easy to understand, interpretable, handles both numerical and categorical data.
-
Disadvantages: Prone to overfitting, unstable (small data changes can change tree structure).
Naive Bayes Classification (in detail)
-
Concept: A probabilistic classifier based on Bayes' Theorem with a "naive" assumption of conditional independence between every pair of features given the class.
-
Bayes' Theorem:
$$P(y|X) = \frac{P(X|y) \cdot P(y)}{P(X)}$$
Where:
* $y$ = class variable
* $$\displaystyle X = (x_1, x_2, ..., x_n) $$ = feature vector
-
"Naive" Assumption: $$\displaystyle P(X|y) = P(x_1|y) \cdot P(x_2|y) \cdot ... \cdot P(x_n|y) $$. This simplifies computation immensely.
-
Classifier: Predicts the class $y$ that maximizes the posterior probability:
$$\hat{y} = \arg\max_y P(y) \prod_{i=1}^{n} P(x_i|y)$$
-
Types: Gaussian (for continuous data), Multinomial (for discrete counts, e.g., text), Bernoulli (for binary features).
-
Advantages: Very fast, works well with high-dimensional data (e.g., text classification), requires less training data.
-
Disadvantages: The independence assumption is rarely true, which can hurt accuracy.
Association Rules
-
Explanation: A rule-based machine learning method for discovering interesting relations (association/correlation) between variables in large databases. Aims to find frequent patterns, correlations, or causal structures.
-
Key Metrics:
- Support: Proportion of transactions in the dataset that contain the itemset.
$$\text{Support}(X \rightarrow Y) = P(X \cup Y)$$
* **Confidence:** Conditional probability of $Y$ given $X$.
$$\text{Confidence}(X \rightarrow Y) = P(Y|X) = \frac{\text{Support}(X \cup Y)}{\text{Support}(X)}$$
* **Lift:** Measures how many times more often $X$ and $Y$ occur together than expected if they were statistically independent.
$$\text{Lift}(X \rightarrow Y) = \frac{\text{Confidence}(X \rightarrow Y)}{\text{Support}(Y)}$$
* Lift > 1: Positive correlation.
* Lift = 1: Independence.
* Lift < 1: Negative correlation.
-
Algorithm: Apriori (most famous), FP-Growth.
-
Applications:
-
Market Basket Analysis: "Customers who buy X also buy Y." (e.g., bread & butter).
-
Web Usage Mining: Predicting next page navigation.
-
Bioinformatics: Finding co-occurring genes.
-
Medical Diagnosis: Finding symptom-disease associations.
-
G. HADOOP ECOSYSTEM & MAPREDUCE
Hadoop: Building Blocks / Architecture
[[DIAGRAM: CANVAS: A layered diagram:
-
Applications Layer: Pig, Hive, HBase, Mahout, etc.
-
MapReduce Engine: JobTracker (master) & TaskTrackers (slaves). (Note: In YARN, this is replaced by ResourceManager & ApplicationMaster).
-
HDFS (Storage Layer): NameNode (master) & DataNodes (slaves). Blocks distributed across DataNodes.
-
Underlying OS: Linux/Unix cluster.
Key flows: Applications submit jobs to MapReduce Engine, which reads/writes data from/to HDFS. HDFS stores data in blocks (default 128MB/256MB) replicated across DataNodes.]]
-
Core Components:
-
HDFS (Hadoop Distributed File System): Primary storage system. NameNode (master, manages metadata) & DataNode (slave, stores data blocks).
-
MapReduce: Programming model for processing large datasets. JobTracker (master, schedules jobs) & TaskTracker (slave, executes tasks). (Note: In Hadoop 2.x+, YARN replaces this).
-
YARN (Yet Another Resource Negotiator): Resource management layer. ResourceManager (master) & NodeManager (slave). Decouples resource management from MapReduce programming model.
-
Common/Hadoop Core: Libraries and utilities required by other Hadoop modules.
-
MapReduce: Concept, Types, Formats, Flow
Concept with Example
-
Idea: Split a large problem into independent sub-problems (map), distribute them across a cluster, then combine the results (reduce).
-
Example: Word Count
-
Map Phase: Input: Text lines. Output:
(word, 1)for each word.-
Input:
"the quick brown fox" -
Map Output:
("the", 1),("quick", 1),("brown", 1),("fox", 1)
-
-
Shuffle & Sort: Group all values by key. Intermediate:
("the", [1,1,1,...]),("quick", [1]), etc. -
Reduce Phase: Sum the counts for each word.
- Reduce Output:
("the", 5),("quick", 2), ...
- Reduce Output:
-
Different Types and Formats (with example each)
| Type / Format | Description | Example Use Case |
|---|---|---|
| InputFormat | Defines how to read input data from source and split into logical InputSplits. | TextInputFormat (default): Each line is a record. Key=byte offset, Value=line. |
| OutputFormat | Defines how to write output data (key-value pairs) to storage. | TextOutputFormat (default): Writes key-value pairs as text lines. |
| Mapper | User-defined class that processes input key-value pairs and emits intermediate key-value pairs. | Mapper<LongWritable, Text, Text, IntWritable> for WordCount. |
| Reducer | User-defined class that processes all values associated with the same intermediate key. | Reducer<Text, IntWritable, Text, IntWritable> summing counts. |
| Combiner | Local, mini-reducer that runs after map phase to reduce network traffic. | Summing counts for a single mapper's output before shuffle. |
| Partitioner | Controls which reducer a particular intermediate key is sent to. Default is HashPartitioner. | Custom partitioner to group related keys to same reducer. |
Application Flow
-
Input Splitting: Input data is split into logical InputSplits.
-
Map Phase:
map()function processes each split. Intermediate(key, value)pairs are written to local disk. -
Shuffle & Sort: Framework transfers map outputs to the appropriate reducer based on partitioner. Sorts the intermediate keys.
-
Reduce Phase:
reduce()function processes all values for a given key. Final output is written to HDFS. -
Commit: Output is atomically moved to final location.
Hive: Procedure to Write User-Defined Functions (UDFs)
-
Write Java Class: Extend
org.apache.hadoop.hive.ql.exec.UDF(for scalar functions) orUDAF/UDTF for aggregation/table-generating functions. Override theevaluate()method.public class MyUDF extends UDF { public String evaluate(String input) { return input.toUpperCase(); // Simple example } } -
Package: Compile and package the class into a JAR file.
-
Add JAR to Hive Session:
ADD JAR /path/to/myudf.jar; -
Create Temporary/Permanent Function:
-
Temporary:
CREATE TEMPORARY FUNCTION my_upper AS 'com.example.MyUDF'; -
Permanent:
CREATE FUNCTION my_upper AS 'com.example.MyUDF' USING JAR '/path/to/myudf.jar';(requires hive.aux.jars.path config).
-
-
Use in HiveQL:
SELECT my_upper(name) FROM users;
Apache Pig: Architecture & Pig Latin Flow
Architecture (with sketch)
[[DIAGRAM: CANVAS: A flowchart:
-
Pig Latin Script (User written)
--> Parser (Checks syntax, validates)
--> Logical Plan (DAG of operators)
--> Logical Optimizer (Combines, reorders operations)
--> Physical Plan (Converts to MapReduce/Tez/Spark jobs)
--> Physical Optimizer (Reorders, combines jobs)
--> Execution Engine (Submits jobs to Hadoop cluster)
--> MapReduce/Tez/Spark Jobs (Actual execution on HDFS)
Key: Pig Latin is a high-level language; Pig runtime compiles it into a series of MapReduce (or other) jobs.]]
- Components: Pig Latin (scripting language), Pig Runtime (compiler/optimizer), Execution Engine (MapReduce/Tez/Spark), Grunt/Shell (interactive shell).
Pig Latin Application Flow
-
Write Script: User writes a Pig Latin script (
.pigfile) with a sequence of operations (LOAD, FILTER, GROUP, FOREACH, STORE). -
Parse & Validate: Pig parser checks syntax and semantics.
-
Logical Plan Creation: Creates a logical plan (logical DAG) of the operations.
-
Logical Optimization: Applies rules (e.g., projection pushdown, filter combination) to optimize the logical plan.
-
Physical Plan Creation: Converts the logical plan into a physical plan consisting of one or more MapReduce/Tez/Spark jobs.
-
Physical Optimization: Optimizes the physical plan (e.g., job chaining, reordering).
-
Execution: The execution engine submits the compiled jobs to the underlying cluster (YARN).
-
Result: Output is stored in HDFS or another specified location.
H. R PROGRAMMING FOR DATA ANALYSIS
Features of R Programming Language
-
Open-source & Free: GNU project.
-
Vectorized Operations: Most operations are applied to entire vectors/matrices without explicit loops.
-
Extensive Statistical & Graphical Capabilities: Huge collection of packages for statistics, machine learning, and visualization (ggplot2, lattice).
-
Cross-platform: Runs on Windows, macOS, Linux.
-
Strong Community & Package Ecosystem: CRAN hosts thousands of packages for diverse domains.
-
Interoperability: Can call code from C, C++, Fortran, Python. Interfaces with databases.
-
Active Development: Regular updates and new features.
-
Object-Oriented & Functional Programming Support: Supports S3, S4, R6, and functional paradigms.
Major Components of R Environment
-
R Console/Interpreter: The core engine that executes R commands.
-
Script Files: Files with
.Rextension containing R code. -
Workspace: The current R session's environment containing all defined objects (variables, functions, data frames).
-
History: Log of previously executed commands.
-
Packages: Collections of functions, data, and documentation (e.g.,
stats,ggplot2,dplyr). Installed in the library path. -
Help System:
?functionorhelp(function)provides documentation. -
Graphics Device: System for generating plots (e.g., X11, windows, quartz, pdf, png).
-
RStudio (IDE): Popular integrated development environment providing a script editor, workspace viewer, console, and plot pane.
Operations on Vectors
Vectors are the basic data type in R (atomic vectors: logical, integer, numeric, character, complex, raw).
-
Creation:
-
c(1, 2, 3)- combine elements. -
1:10- sequence. -
seq(1, 10, by=0.5)- sequence with step. -
rep("a", 5)- repeat element.
-
-
Indexing/Accessing:
-
x[1]- first element. -
x[c(1,3)]- elements at positions 1 and 3. -
x[x > 5]- logical indexing (elements > 5). -
x[-2]- exclude element at position 2.
-
-
Arithmetic Operations (Element-wise):
x + y,x * y,x / y,x %% y(modulo),x %/% y(integer division).
-
Vectorized Functions:
sqrt(x),log(x),exp(x),sin(x),sum(x),mean(x),sd(x),sort(x),rev(x).
-
Logical Operations (Element-wise):
x > 5,x == y,x & y(element-wise AND),x && y(single AND),x | y,x || y.
-
Aggregation:
sum(x),prod(x),min(x),max(x),range(x).
-
Missing Values:
NA(Not Available). Useis.na(x)to detect,na.omit(x)to remove,x[!is.na(x)].
[!TIP] Critical: R is vectorized.
x + 1adds 1 to every element ofx. Loops (for,while) are often slower and should be avoided when vectorized alternatives exist.