How unit 2 is examined
This unit has a single topic, R as a data analytics tool. No question on it appears in the supplied papers, so learn the definition, the features, the data structures and the basic analysis workflow, which is what any question on it would ask.
Introduction to R as a data analytics tool
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>R is a free, open-source programming language and software environment for statistical computing, data analysis and graphics.</mark>
Key points.
- R was created by Ross Ihaka and Robert Gentleman and is now maintained by the R Core Team under the GNU licence, so it costs nothing and runs on Windows, Linux and macOS.
- R is an interpreted, vector-based language, so one command such as
mean(x)orx * 2works on a whole set of values at once without writing a loop. - It has built-in statistics such as mean, variance, correlation, regression, hypothesis tests and probability distributions, and strong graphics through
plot(),hist(),boxplot()and theggplot2package. - CRAN, the Comprehensive R Archive Network, holds thousands of free packages, and each is installed once with
install.packages()and loaded withlibrary(). - RStudio is the usual IDE, giving a console, a script editor, a workspace viewer and a plot pane in one window.
- R can read CSV, Excel, text and database data, so it covers the whole analytics cycle from import and cleaning to modelling and reporting.
Diagram.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-01" viewBox="0 0 527.2 80" width="527.2" height="80" role="img" aria-label="Analysis workflow in R. Imp import data, Cln clean data, Exp explore, Mod model, Vis visualize and report"><style>#dsfig-u2-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-01 .t{fill:#16181D;font-weight:500}#dsfig-u2-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-01 .dot{fill:#16181D}#dsfig-u2-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-01 .ah{fill:#454C5A}#dsfig-u2-01 .ah.hi{fill:#2340B8}#dsfig-u2-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-01 .e{stroke:#B1B7C3}html.dark #dsfig-u2-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-01 .t{fill:#E6E8ED}html.dark #dsfig-u2-01 .t.inv{fill:#0F1115}html.dark #dsfig-u2-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-01 .dot{fill:#E6E8ED}html.dark #dsfig-u2-01 .ann{fill:#8FA3FF}html.dark #dsfig-u2-01 .lbl{fill:#858D9C}html.dark #dsfig-u2-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-01 .ah{fill:#B1B7C3}html.dark #dsfig-u2-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L130.8,40" marker-end="url(#ah1)"/><path class="e" d="M170.8,40 L242.6,40" marker-end="url(#ah1)"/><path class="e" d="M282.6,40 L354.4,40" marker-end="url(#ah1)"/><path class="e" d="M394.4,40 L466.2,40" marker-end="url(#ah1)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Imp</text><circle class="n" cx="151.8" cy="40" r="18"/><text class="t" x="151.8" y="40" dy=".35em" text-anchor="middle">Cln</text><circle class="n" cx="263.6" cy="40" r="18"/><text class="t" x="263.6" y="40" dy=".35em" text-anchor="middle">Exp</text><circle class="n" cx="375.4" cy="40" r="18"/><text class="t" x="375.4" y="40" dy=".35em" text-anchor="middle">Mod</text><circle class="n" cx="487.2" cy="40" r="18"/><text class="t" x="487.2" y="40" dy=".35em" text-anchor="middle">Vis</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Analysis workflow in R. Imp import data, Cln clean data, Exp explore, Mod model, Vis visualize and report</figcaption></figure>
Data types and structures.
| Item | Meaning | Example |
|---|---|---|
| Numeric | decimal or whole numbers | x <- 3.5 |
| Integer | whole numbers marked with L | n <- 5L |
| Character | text in quotes | s <- "R" |
| Logical | TRUE or FALSE | b <- TRUE |
| Vector | ordered values of one type | c(12, 15, 11) |
| Matrix | 2-D table of one type | matrix(1:6, nrow = 2) |
| List | mixed types together | list(1, "a", TRUE) |
| Data frame | table whose columns may differ in type | data.frame(id = 1:3, mark = c(60, 72, 55)) |
| Factor | categorical variable with levels | factor(c("M", "F", "M")) |
Basic commands. Assignment uses <-, help uses ?mean, the working directory is set with setwd(), data is read with read.csv("file.csv"), and head(df), str(df) and summary(df) give a quick look at a data frame.
Example. Descriptive statistics for the data 12, 15, 11, 18, 14.
x <- c(12, 15, 11, 18, 14)
mean(x) # 14
median(x) # 14
var(x) # 7.5
sd(x) # 2.738613
summary(x) # Min 11, 1st Qu. 12, Median 14, Mean 14, 3rd Qu. 15, Max 18
Example. Correlation and simple linear regression for x = 1, 2, 3, 4, 5 and y = 2, 4, 5, 4, 5.
x <- 1:5
y <- c(2, 4, 5, 4, 5)
cor(x, y) # 0.7745967
lm(y ~ x) # intercept 2.2, slope 0.6
The working is $b = \frac{S_{xy}}{S_{xx}} = \frac{6}{10} = 0.6$ and $a = \bar{y} - b\bar{x} = 4 - 0.6 \times 3 = 2.2$, so the line is $\hat{y} = 2.2 + 0.6x$.
Example. Data frame, filtering and plotting.
df <- data.frame(id = 1:4, mark = c(60, 72, 55, 90))
df[df$mark > 58, ] # rows with id 1, 2, 4
mean(df$mark) # 69.25
df$grade <- ifelse(df$mark >= 60, "Pass", "Fail") # new column
hist(df$mark) # histogram
boxplot(df$mark) # boxplot marks outliers as points
Example. Vectors, indexing, a function and a loop.
v <- c(10, 20, 30, 40)
v[2] # 20, indexing starts at 1
v * 2 # 20 40 60 80, no loop needed
sq <- function(a) a^2
sq(5) # 25
for (i in 1:3) print(i) # prints 1, 2, 3
Steps to start.
Step 1: Install R from CRAN, then install RStudio.
Step 2: Open RStudio and write commands in the console or in a script file.
Step 3: Run install.packages("dplyr") once, then library(dplyr) in every session.
Step 4: Import data with read.csv(), inspect it with str() and summary(), then analyse and plot.
Step 5: Save the script and export results or an R Markdown report.
Distributions and statistics in R. Each distribution has four functions named by a prefix: d for density, p for cumulative probability, q for quantile and r for random numbers. For example dnorm(0) gives 0.3989423, pnorm(1.96) gives 0.9750021 and rnorm(100) draws 100 standard normal values. Other tests include t.test(), chisq.test() and cor.test().
R compared with other tools.
| Point | R | Python | MATLAB |
|---|---|---|---|
| Cost | free | free | paid licence |
| Strength | statistics and graphics | general purpose plus machine learning | matrices and engineering |
| Data structure | data frame | DataFrame in pandas | matrix |
| Packages | CRAN | PyPI | toolboxes |
| Typical user | statisticians, researchers | developers, data scientists | engineers |
Common packages.
| Package | Use |
|---|---|
dplyr |
filter, select, group and summarise data |
tidyr |
reshape and tidy data |
ggplot2 |
layered, publication-quality graphics |
caret |
machine-learning models and evaluation |
shiny |
interactive web dashboards |
readr |
fast import of CSV and text files |
Uses in data analytics. R is used for data cleaning (handling missing values with is.na() and na.omit()), exploratory analysis, regression and classification models, time-series forecasting, clustering with kmeans(), and reporting with R Markdown. Banks, healthcare, bioinformatics and research groups use it widely.
Advantages.
- It is free and open source, and a very large community keeps adding packages.
- Its statistical depth and graphics are among the best of any analytics tool.
- It works with other tools, calling C, Python and SQL databases, and produces reports through R Markdown.
Limitations.
- It keeps data in memory, so very large data sets can exhaust RAM unless a database or a big-data package is used.
- Its learning curve is steep for beginners, and base R code is slower than compiled languages for heavy loops.
Answer frame. Open with the definition; list the features in the order free and open source, vector operations, built-in statistics, graphics, CRAN packages, RStudio; give the data-structure table and the small mean and regression example; close with advantages against limitations.
Pitfall: R uses
<-for assignment, indexes from 1 (not 0), and is case sensitive, soMean(x)fails wheremean(x)works.
Last-minute revision
- R is a free, open-source language and environment for statistical computing and graphics.
- R was created by Ross Ihaka and Robert Gentleman.
- R is interpreted and vector-based, so operations apply to whole vectors.
- CRAN is the repository of R packages.
- Packages are installed with
install.packages()and loaded withlibrary(). - Assignment uses
<-, and indexing starts at 1. - The data frame is the main table structure, and each column may have a different type.
- The other structures are vector, matrix, list, array and factor.
read.csv()imports data andsummary()gives the five-number summary and the mean.ggplot2is for graphics anddplyris for data manipulation.lm(y ~ x)fits a linear regression andcor(x, y)gives the correlation.- RStudio is the standard IDE for R.
Memory hooks
- R = Really good at Reports, stats and plots.
- CRAN = the Central store of R packages: install once, library each time.
- Data frame = an Excel sheet inside R.
- Workflow = Import, Clean, Explore, Model, Visualize (I Can Explore Many Views).
- Arrow
<-puts the value into the name; counting starts at 1.
Coverage checklist
- Introduction to R as a data analytics tool: no past questions are tagged; the definition, key points, data structures, examples, packages and advantages above cover any question on it.