Skip to content
IT-407 · Open Source Software Lab (Linux and R)/Quick Revision Short Notes

Open Source Software Lab (Linux and R) (IT-407) - Unit 5 Short Notes

UNIT 5: ADVANCED R FOR DATA ANALYSIS, VISUALIZATION, AND REPORTING


5.1 Advanced Data Manipulation & Wrangling with dplyr and tidyr

5.1.1 Core dplyr Verbs for Complex Operations

A tidy data pipeline uses these verbs sequentially, typically within a %>% pipe.

Verb Primary Purpose Key Features / Notes
filter() Row subsetting Use & (AND), `
select() Column subsetting starts_with(), ends_with(), contains(), matches() for pattern selection. rename(new = old) renames. relocate() moves columns.
mutate() Create new variables Creates columns sequentially. Use case_when() for complex conditional logic. lag()/lead() for time-series shifts.
arrange() Sort rows arrange(col) ascending, arrange(desc(col)) descending. Multiple columns: arrange(col1, desc(col2)).
summarise() Reduce to single value Must be paired with group_by() for group-level stats. Use na.rm = TRUE to ignore NAs in calculations (e.g., mean(x, na.rm=TRUE)).

[!TIP] Common Pitfall: summarise() without group_by() collapses the entire dataset to one row. Always check your grouping.

5.1.2 Data Reshaping with tidyr

Converts between "wide" and "long" formats for analysis/plotting.

Function Purpose Key Idea
pivot_longer() Gather columns into key-value pairs cols = c(col1, col2) to gather. names_to = "new_key", values_to = "new_value".
pivot_wider() Spread key-value pairs into new columns names_from = "key_col", values_from = "value_col".
separate() Split one column into multiple col = "messy_col", into = c("new1", "new2"), sep = "_" (or numeric).
unite() Combine multiple columns into one col = "new_col", sep = "_".
replace_na() Replace NA with a value replace_na(list(col = 0)).

5.1.3 The Pipe Operator (%>%)

From magrittr (loaded with dplyr). Passes the left-hand side as the first argument to the right-hand side function.


# Without pipe (nested, hard to read)

filtered_data <- filter(mutate(arrange(df, x), new_col = y*2), z > 10)

# With pipe (left-to-right, readable)

processed_df <- df %>%

  arrange(x) %>%

  mutate(new_col = y * 2) %>%

  filter(z > 10)

Use placeholder . when the piped data isn't the first argument: lm(y ~ x, data = .).

5.1.4 Joining/Combining Data Frames

  • Join Functions (merge rows based on key):

    • inner_join(x, y, by = "key"): Keep only matching rows (intersection).

    • left_join(x, y, by = "key"): Keep all rows from x (x is "left").

    • right_join(x, y, by = "key"): Keep all rows from y.

    • full_join(x, y, by = "key"): Keep all rows from both (union).

  • Bind Functions (combine without matching keys):

    • bind_rows(x, y): Stack data frames vertically (rows). Requires same/similar column names.

    • bind_cols(x, y): Combine data frames horizontally (columns). Requires same number of rows.


5.2 Data Visualization with ggplot2 (Grammar of Graphics)

5.2.1 Core Components of a ggplot Object

Built in layers: ggplot(data, aes(x, y)) + geom_...() + ...

Layer Function Purpose
1. Data & Aesthetics ggplot(data, aes(x, y, color=group)) Initialize plot, map variables to visual properties (position, color, size, shape).
2. Geometric Objects geom_point(), geom_line(), geom_bar(stat="identity"), geom_histogram(bins=30), geom_boxplot(), geom_smooth(method="lm") The actual visual marks (points, bars, lines).
3. Facets facet_wrap(~ group) or facet_grid(row ~ col) Create small multiples (subplots) by a factor variable.

5.2.2 Customizing Aesthetics

  • Mapping (inside aes()): Links a variable to an aesthetic. Color changes by data group.

    
    ggplot(mpg, aes(x=displ, y=hwy, color=class)) + geom_point()
    
    
  • Setting (outside aes()): Sets an aesthetic to a constant value.

    
    ggplot(mpg, aes(x=displ, y=hwy)) + geom_point(color="red", size=3)
    
    
  • Scale Functions: Fine-tune axes and legends.

    • scale_x_continuous(name="Label", breaks=seq(0,10,2))

    • scale_y_log10()

    • scale_color_manual(values=c("setosa"="#F8766D", "versicolor"="#00BA38"))

5.2.3 Themes and Labels

  • Pre-built Themes: theme_minimal(), theme_classic(), theme_bw().

  • Labels: labs(title="Main", subtitle="Sub", x="X Axis", y="Y Axis", caption="Source").

  • Theme Customization: theme() modifies non-data elements.

    
    theme(plot.title = element_text(hjust=0.5, face="bold"),
    
          legend.position = "bottom",
    
          panel.grid.major = element_line(color="grey80"))
    
    

5.2.4 Saving Plots


ggsave("my_plot.png", plot = last_plot(), width=8, height=6, dpi=300, units="in")

  • If plot is omitted, saves the last displayed plot.

  • File extension (.png, .pdf, .jpg) determines format. PDF is vector (scalable).


5.3 Basic Statistical Modeling & Inference in R

5.3.1 Linear Regression (lm())

Fits model:

$$ y = \beta_0 + \beta_1 x_1 + \beta_2 x_2 + ... + \epsilon $$


model <- lm(mpg ~ wt + hp, data = mtcars)  # Formula interface

summary(model)  # Key output below

summary(model) Key Components:

  • Coefficients: Estimate (β), Std. Error, t-value, Pr(>|t|) (p-value). *** p<0.001, ** p<0.01, * p<0.05.

  • Residual Standard Error: $$\displaystyle \sqrt{\frac{RSS}{df}} $$, estimate of σ.

  • Multiple R-squared:

$$ R^2 = \frac{\text{Explained Variance}}{\text{Total Variance}} $$

. Proportion of variance in y explained by model.

  • Adjusted R-squared: R² penalized for number of predictors.

  • F-statistic: Tests if model is significantly better than intercept-only model (p-value at bottom).

  • Diagnostic Plots: plot(model) generates 4 plots: Residuals vs Fitted, Normal Q-Q, Scale-Location, Residuals vs Leverage.

5.3.2 Hypothesis Testing Functions

  • t-test: t.test(x, y = NULL, alternative = "two.sided", var.equal = FALSE)

    • x: numeric vector. y: second vector (for two-sample) or NULL (for one-sample).

    • var.equal = TRUE assumes equal variances (pooled SD).

  • ANOVA: aov(y ~ group, data = df) then summary().

    • Tests if means of >2 groups are equal.
  • Chi-square test: chisq.test(table_matrix)

    • Tests independence in a contingency table.

5.3.3 Correlation & Covariance

  • Covariance: cov(x, y). Measures linear association direction (scale-dependent).

  • Pearson Correlation: cor(x, y, method="pearson").

$$ r = \frac{\text{cov}(x,y)}{s_x s_y} $$

, scale-free [-1,1].

  • Test Significance: cor.test(x, y) gives correlation coefficient and p-value for test $$\displaystyle H_0: \rho = 0 $$.

5.4 Reproducible Reporting with R Markdown (.Rmd)

5.4.1 R Markdown Document Structure

---

title: "Analysis Report"

author: "Name"

date: "`r Sys.Date()`"

output: 

  html_document: 

    toc: true

    theme: united
---

  • Code Chunks: {r chunk-name, echo=FALSE, warning=FALSE}

    • echo=FALSE: Hide code, show output only.

    • results='hide': Hide text output (e.g., print()).

    • fig.cap="Caption": Adds caption, enables figure numbering.

    • include=FALSE: Run code but hide all output.

  • Inline Code: `r mean(x)`. Evaluates and inserts result directly in text.

5.4.2 Integrating Text, Code, and Output

  • Text uses Markdown (# Header, **bold**, - list item, [link](url)).

  • Plots/table outputs from code chunks are automatically embedded.

  • Formatted Tables: Use knitr::kable(df, caption="My Table"). For advanced formatting, use kableExtra::kable_styling().

5.4.3 Rendering the Document

  1. .Rmd file is knit by knitr → generates a .md (Markdown) file.

  2. .md file is rendered by pandoc → final output (HTML/PDF/Word).

  • Knit Button in RStudio runs rmarkdown::render("file.Rmd").

  • Dependency: Changing code in .Rmd requires re-knitting to update output.

[!TIP] Debugging: If a plot/table isn't appearing, check chunk options (echo, include, fig.show). Ensure code in chunk actually produces output.


5.5 Practical Lab Integration: End-to-End Analysis Workflow

5.5.1 Project Structure & Data Import


project_folder/

├── data/

│   ├── raw/      # Original, unmodified data

│   └── processed/ # Cleaned data saved for reuse

├── scripts/      # .R scripts (optional, for modular code)

├── output/       # Generated plots, tables, reports

└── report.Rmd    # Main analysis document

  • Import: read.csv("data/file.csv"), read.table(). For Excel: readxl::read_excel("file.xlsx", sheet=1).

5.5.2 Exploratory Data Analysis (EDA) Pipeline

  1. Initial Inspection: glimpse(df), str(df), summary(df), head(df, 10).

  2. Univariate: Histograms (ggplot(df, aes(x)) + geom_histogram()), bar plots for factors, descriptive stats (mean(), sd(), quantile()).

  3. Bivariate/Multivariate: Scatter plots (geom_point()), box plots (geom_boxplot()), correlation matrix (cor(df[,sapply(df, is.numeric)])), pairs().

5.5.3 Communicating Results (R Markdown Report)

Structure your .Rmd file to tell a story:

  1. Introduction: State the question/objective.

  2. Methods (Data Wrangling): Document all dplyr/tidyr steps. Show key code chunks (use echo=TRUE).

  3. Results:

    • Present key visualizations (from ggplot2) with clear captions.

    • Present key statistical outputs (tables from summary(lm), aov, cor.test). Use kable().

    • Interpret numbers in plain language (e.g., "For every 1-unit increase in X, Y increases by β units, p < 0.05").

  4. Conclusion/Interpretation: Summarize findings, state limitations, suggest implications.

[!TIP] Exam Strategy: In a lab exam, follow this exact workflow: Import → Wrangle (pipe chains) → Visualize (ggplot2 layers) → Model (lm/summary) → Report (Rmd with proper chunk options). Practice debugging common errors: missing packages, column name typos, incorrect aes() mapping, chunk option syntax.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in