UNIT 5: ADVANCED R FOR DATA ANALYSIS, VISUALIZATION, AND REPORTING
5.1 Advanced Data Manipulation & Wrangling with dplyr and tidyr
5.1.1 Core dplyr Verbs for Complex Operations
A tidy data pipeline uses these verbs sequentially, typically within a %>% pipe.
| Verb | Primary Purpose | Key Features / Notes |
|---|---|---|
filter() |
Row subsetting | Use & (AND), ` |
select() |
Column subsetting | starts_with(), ends_with(), contains(), matches() for pattern selection. rename(new = old) renames. relocate() moves columns. |
mutate() |
Create new variables | Creates columns sequentially. Use case_when() for complex conditional logic. lag()/lead() for time-series shifts. |
arrange() |
Sort rows | arrange(col) ascending, arrange(desc(col)) descending. Multiple columns: arrange(col1, desc(col2)). |
summarise() |
Reduce to single value | Must be paired with group_by() for group-level stats. Use na.rm = TRUE to ignore NAs in calculations (e.g., mean(x, na.rm=TRUE)). |
[!TIP] Common Pitfall:
summarise()withoutgroup_by()collapses the entire dataset to one row. Always check your grouping.
5.1.2 Data Reshaping with tidyr
Converts between "wide" and "long" formats for analysis/plotting.
| Function | Purpose | Key Idea |
|---|---|---|
pivot_longer() |
Gather columns into key-value pairs | cols = c(col1, col2) to gather. names_to = "new_key", values_to = "new_value". |
pivot_wider() |
Spread key-value pairs into new columns | names_from = "key_col", values_from = "value_col". |
separate() |
Split one column into multiple | col = "messy_col", into = c("new1", "new2"), sep = "_" (or numeric). |
unite() |
Combine multiple columns into one | col = "new_col", sep = "_". |
replace_na() |
Replace NA with a value |
replace_na(list(col = 0)). |
5.1.3 The Pipe Operator (%>%)
From magrittr (loaded with dplyr). Passes the left-hand side as the first argument to the right-hand side function.
# Without pipe (nested, hard to read)
filtered_data <- filter(mutate(arrange(df, x), new_col = y*2), z > 10)
# With pipe (left-to-right, readable)
processed_df <- df %>%
arrange(x) %>%
mutate(new_col = y * 2) %>%
filter(z > 10)
Use placeholder . when the piped data isn't the first argument: lm(y ~ x, data = .).
5.1.4 Joining/Combining Data Frames
-
Join Functions (merge rows based on key):
-
inner_join(x, y, by = "key"): Keep only matching rows (intersection). -
left_join(x, y, by = "key"): Keep all rows fromx(x is "left"). -
right_join(x, y, by = "key"): Keep all rows fromy. -
full_join(x, y, by = "key"): Keep all rows from both (union).
-
-
Bind Functions (combine without matching keys):
-
bind_rows(x, y): Stack data frames vertically (rows). Requires same/similar column names. -
bind_cols(x, y): Combine data frames horizontally (columns). Requires same number of rows.
-
5.2 Data Visualization with ggplot2 (Grammar of Graphics)
5.2.1 Core Components of a ggplot Object
Built in layers: ggplot(data, aes(x, y)) + geom_...() + ...
| Layer | Function | Purpose |
|---|---|---|
| 1. Data & Aesthetics | ggplot(data, aes(x, y, color=group)) |
Initialize plot, map variables to visual properties (position, color, size, shape). |
| 2. Geometric Objects | geom_point(), geom_line(), geom_bar(stat="identity"), geom_histogram(bins=30), geom_boxplot(), geom_smooth(method="lm") |
The actual visual marks (points, bars, lines). |
| 3. Facets | facet_wrap(~ group) or facet_grid(row ~ col) |
Create small multiples (subplots) by a factor variable. |
5.2.2 Customizing Aesthetics
-
Mapping (inside
aes()): Links a variable to an aesthetic. Color changes by data group.ggplot(mpg, aes(x=displ, y=hwy, color=class)) + geom_point() -
Setting (outside
aes()): Sets an aesthetic to a constant value.ggplot(mpg, aes(x=displ, y=hwy)) + geom_point(color="red", size=3) -
Scale Functions: Fine-tune axes and legends.
-
scale_x_continuous(name="Label", breaks=seq(0,10,2)) -
scale_y_log10() -
scale_color_manual(values=c("setosa"="#F8766D", "versicolor"="#00BA38"))
-
5.2.3 Themes and Labels
-
Pre-built Themes:
theme_minimal(),theme_classic(),theme_bw(). -
Labels:
labs(title="Main", subtitle="Sub", x="X Axis", y="Y Axis", caption="Source"). -
Theme Customization:
theme()modifies non-data elements.theme(plot.title = element_text(hjust=0.5, face="bold"), legend.position = "bottom", panel.grid.major = element_line(color="grey80"))
5.2.4 Saving Plots
ggsave("my_plot.png", plot = last_plot(), width=8, height=6, dpi=300, units="in")
-
If
plotis omitted, saves the last displayed plot. -
File extension (
.png,.pdf,.jpg) determines format. PDF is vector (scalable).
5.3 Basic Statistical Modeling & Inference in R
5.3.1 Linear Regression (lm())
Fits model:
$$ y = \beta_0 + \beta_1 x_1 + \beta_2 x_2 + ... + \epsilon $$
model <- lm(mpg ~ wt + hp, data = mtcars) # Formula interface
summary(model) # Key output below
summary(model) Key Components:
-
Coefficients: Estimate (β), Std. Error, t-value, Pr(>|t|) (p-value).
***p<0.001,**p<0.01,*p<0.05. -
Residual Standard Error: $$\displaystyle \sqrt{\frac{RSS}{df}} $$, estimate of σ.
-
Multiple R-squared:
$$ R^2 = \frac{\text{Explained Variance}}{\text{Total Variance}} $$
. Proportion of variance in y explained by model.
-
Adjusted R-squared: R² penalized for number of predictors.
-
F-statistic: Tests if model is significantly better than intercept-only model (p-value at bottom).
-
Diagnostic Plots:
plot(model)generates 4 plots: Residuals vs Fitted, Normal Q-Q, Scale-Location, Residuals vs Leverage.
5.3.2 Hypothesis Testing Functions
-
t-test:
t.test(x, y = NULL, alternative = "two.sided", var.equal = FALSE)-
x: numeric vector.y: second vector (for two-sample) orNULL(for one-sample). -
var.equal = TRUEassumes equal variances (pooled SD).
-
-
ANOVA:
aov(y ~ group, data = df)thensummary().- Tests if means of >2 groups are equal.
-
Chi-square test:
chisq.test(table_matrix)- Tests independence in a contingency table.
5.3.3 Correlation & Covariance
-
Covariance:
cov(x, y). Measures linear association direction (scale-dependent). -
Pearson Correlation:
cor(x, y, method="pearson").
$$ r = \frac{\text{cov}(x,y)}{s_x s_y} $$
, scale-free [-1,1].
- Test Significance:
cor.test(x, y)gives correlation coefficient and p-value for test $$\displaystyle H_0: \rho = 0 $$.
5.4 Reproducible Reporting with R Markdown (.Rmd)
5.4.1 R Markdown Document Structure
---
title: "Analysis Report"
author: "Name"
date: "`r Sys.Date()`"
output:
html_document:
toc: true
theme: united
---
-
Code Chunks:
{r chunk-name, echo=FALSE, warning=FALSE}-
echo=FALSE: Hide code, show output only. -
results='hide': Hide text output (e.g.,print()). -
fig.cap="Caption": Adds caption, enables figure numbering. -
include=FALSE: Run code but hide all output.
-
-
Inline Code:
`r mean(x)`. Evaluates and inserts result directly in text.
5.4.2 Integrating Text, Code, and Output
-
Text uses Markdown (
# Header,**bold**,- list item,[link](url)). -
Plots/table outputs from code chunks are automatically embedded.
-
Formatted Tables: Use
knitr::kable(df, caption="My Table"). For advanced formatting, usekableExtra::kable_styling().
5.4.3 Rendering the Document
-
.Rmdfile is knit byknitr→ generates a.md(Markdown) file. -
.mdfile is rendered bypandoc→ final output (HTML/PDF/Word).
-
Knit Button in RStudio runs
rmarkdown::render("file.Rmd"). -
Dependency: Changing code in
.Rmdrequires re-knitting to update output.
[!TIP] Debugging: If a plot/table isn't appearing, check chunk options (
echo,include,fig.show). Ensure code in chunk actually produces output.
5.5 Practical Lab Integration: End-to-End Analysis Workflow
5.5.1 Project Structure & Data Import
project_folder/
├── data/
│ ├── raw/ # Original, unmodified data
│ └── processed/ # Cleaned data saved for reuse
├── scripts/ # .R scripts (optional, for modular code)
├── output/ # Generated plots, tables, reports
└── report.Rmd # Main analysis document
- Import:
read.csv("data/file.csv"),read.table(). For Excel:readxl::read_excel("file.xlsx", sheet=1).
5.5.2 Exploratory Data Analysis (EDA) Pipeline
-
Initial Inspection:
glimpse(df),str(df),summary(df),head(df, 10). -
Univariate: Histograms (
ggplot(df, aes(x)) + geom_histogram()), bar plots for factors, descriptive stats (mean(),sd(),quantile()). -
Bivariate/Multivariate: Scatter plots (
geom_point()), box plots (geom_boxplot()), correlation matrix (cor(df[,sapply(df, is.numeric)])),pairs().
5.5.3 Communicating Results (R Markdown Report)
Structure your .Rmd file to tell a story:
-
Introduction: State the question/objective.
-
Methods (Data Wrangling): Document all
dplyr/tidyrsteps. Show key code chunks (useecho=TRUE). -
Results:
-
Present key visualizations (from ggplot2) with clear captions.
-
Present key statistical outputs (tables from
summary(lm),aov,cor.test). Usekable(). -
Interpret numbers in plain language (e.g., "For every 1-unit increase in X, Y increases by β units, p < 0.05").
-
-
Conclusion/Interpretation: Summarize findings, state limitations, suggest implications.
[!TIP] Exam Strategy: In a lab exam, follow this exact workflow: Import → Wrangle (pipe chains) → Visualize (ggplot2 layers) → Model (lm/summary) → Report (Rmd with proper chunk options). Practice debugging common errors: missing packages, column name typos, incorrect
aes()mapping, chunk option syntax.