Skip to content
CS-605 · Data Analytics Lab/Quick Revision Short Notes

Data Analytics Lab (CS-605) - Unit 2 Short Notes

How unit 2 is examined

This unit has a single topic, R as a data analytics tool. No question on it appears in the supplied papers, so learn the definition, the features, the data structures and the basic analysis workflow, which is what any question on it would ask.

Introduction to R as a data analytics tool

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>R is a free, open-source programming language and software environment for statistical computing, data analysis and graphics.</mark>

Key points.

  1. R was created by Ross Ihaka and Robert Gentleman and is now maintained by the R Core Team under the GNU licence, so it costs nothing and runs on Windows, Linux and macOS.
  2. R is an interpreted, vector-based language, so one command such as mean(x) or x * 2 works on a whole set of values at once without writing a loop.
  3. It has built-in statistics such as mean, variance, correlation, regression, hypothesis tests and probability distributions, and strong graphics through plot(), hist(), boxplot() and the ggplot2 package.
  4. CRAN, the Comprehensive R Archive Network, holds thousands of free packages, and each is installed once with install.packages() and loaded with library().
  5. RStudio is the usual IDE, giving a console, a script editor, a workspace viewer and a plot pane in one window.
  6. R can read CSV, Excel, text and database data, so it covers the whole analytics cycle from import and cleaning to modelling and reporting.

Diagram.

<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-01" viewBox="0 0 527.2 80" width="527.2" height="80" role="img" aria-label="Analysis workflow in R. Imp import data, Cln clean data, Exp explore, Mod model, Vis visualize and report"><style>#dsfig-u2-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-01 .t{fill:#16181D;font-weight:500}#dsfig-u2-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-01 .dot{fill:#16181D}#dsfig-u2-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-01 .ah{fill:#454C5A}#dsfig-u2-01 .ah.hi{fill:#2340B8}#dsfig-u2-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-01 .e{stroke:#B1B7C3}html.dark #dsfig-u2-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-01 .t{fill:#E6E8ED}html.dark #dsfig-u2-01 .t.inv{fill:#0F1115}html.dark #dsfig-u2-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-01 .dot{fill:#E6E8ED}html.dark #dsfig-u2-01 .ann{fill:#8FA3FF}html.dark #dsfig-u2-01 .lbl{fill:#858D9C}html.dark #dsfig-u2-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-01 .ah{fill:#B1B7C3}html.dark #dsfig-u2-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L130.8,40" marker-end="url(#ah1)"/><path class="e" d="M170.8,40 L242.6,40" marker-end="url(#ah1)"/><path class="e" d="M282.6,40 L354.4,40" marker-end="url(#ah1)"/><path class="e" d="M394.4,40 L466.2,40" marker-end="url(#ah1)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Imp</text><circle class="n" cx="151.8" cy="40" r="18"/><text class="t" x="151.8" y="40" dy=".35em" text-anchor="middle">Cln</text><circle class="n" cx="263.6" cy="40" r="18"/><text class="t" x="263.6" y="40" dy=".35em" text-anchor="middle">Exp</text><circle class="n" cx="375.4" cy="40" r="18"/><text class="t" x="375.4" y="40" dy=".35em" text-anchor="middle">Mod</text><circle class="n" cx="487.2" cy="40" r="18"/><text class="t" x="487.2" y="40" dy=".35em" text-anchor="middle">Vis</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Analysis workflow in R. Imp import data, Cln clean data, Exp explore, Mod model, Vis visualize and report</figcaption></figure>

Data types and structures.

Item Meaning Example
Numeric decimal or whole numbers x <- 3.5
Integer whole numbers marked with L n <- 5L
Character text in quotes s <- "R"
Logical TRUE or FALSE b <- TRUE
Vector ordered values of one type c(12, 15, 11)
Matrix 2-D table of one type matrix(1:6, nrow = 2)
List mixed types together list(1, "a", TRUE)
Data frame table whose columns may differ in type data.frame(id = 1:3, mark = c(60, 72, 55))
Factor categorical variable with levels factor(c("M", "F", "M"))

Basic commands. Assignment uses <-, help uses ?mean, the working directory is set with setwd(), data is read with read.csv("file.csv"), and head(df), str(df) and summary(df) give a quick look at a data frame.

Example. Descriptive statistics for the data 12, 15, 11, 18, 14.

x <- c(12, 15, 11, 18, 14)
mean(x)     # 14
median(x)   # 14
var(x)      # 7.5
sd(x)       # 2.738613
summary(x)  # Min 11, 1st Qu. 12, Median 14, Mean 14, 3rd Qu. 15, Max 18

Example. Correlation and simple linear regression for x = 1, 2, 3, 4, 5 and y = 2, 4, 5, 4, 5.

x <- 1:5
y <- c(2, 4, 5, 4, 5)
cor(x, y)        # 0.7745967
lm(y ~ x)        # intercept 2.2, slope 0.6

The working is $b = \frac{S_{xy}}{S_{xx}} = \frac{6}{10} = 0.6$ and $a = \bar{y} - b\bar{x} = 4 - 0.6 \times 3 = 2.2$, so the line is $\hat{y} = 2.2 + 0.6x$.

Example. Data frame, filtering and plotting.

df <- data.frame(id = 1:4, mark = c(60, 72, 55, 90))
df[df$mark > 58, ]            # rows with id 1, 2, 4
mean(df$mark)                 # 69.25
df$grade <- ifelse(df$mark >= 60, "Pass", "Fail")   # new column
hist(df$mark)                 # histogram
boxplot(df$mark)              # boxplot marks outliers as points

Example. Vectors, indexing, a function and a loop.

v <- c(10, 20, 30, 40)
v[2]                 # 20, indexing starts at 1
v * 2                # 20 40 60 80, no loop needed
sq <- function(a) a^2
sq(5)                # 25
for (i in 1:3) print(i)   # prints 1, 2, 3

Steps to start.

Step 1: Install R from CRAN, then install RStudio.
Step 2: Open RStudio and write commands in the console or in a script file.
Step 3: Run install.packages("dplyr") once, then library(dplyr) in every session.
Step 4: Import data with read.csv(), inspect it with str() and summary(), then analyse and plot.
Step 5: Save the script and export results or an R Markdown report.

Distributions and statistics in R. Each distribution has four functions named by a prefix: d for density, p for cumulative probability, q for quantile and r for random numbers. For example dnorm(0) gives 0.3989423, pnorm(1.96) gives 0.9750021 and rnorm(100) draws 100 standard normal values. Other tests include t.test(), chisq.test() and cor.test().

R compared with other tools.

Point R Python MATLAB
Cost free free paid licence
Strength statistics and graphics general purpose plus machine learning matrices and engineering
Data structure data frame DataFrame in pandas matrix
Packages CRAN PyPI toolboxes
Typical user statisticians, researchers developers, data scientists engineers

Common packages.

Package Use
dplyr filter, select, group and summarise data
tidyr reshape and tidy data
ggplot2 layered, publication-quality graphics
caret machine-learning models and evaluation
shiny interactive web dashboards
readr fast import of CSV and text files

Uses in data analytics. R is used for data cleaning (handling missing values with is.na() and na.omit()), exploratory analysis, regression and classification models, time-series forecasting, clustering with kmeans(), and reporting with R Markdown. Banks, healthcare, bioinformatics and research groups use it widely.

Advantages.

  1. It is free and open source, and a very large community keeps adding packages.
  2. Its statistical depth and graphics are among the best of any analytics tool.
  3. It works with other tools, calling C, Python and SQL databases, and produces reports through R Markdown.

Limitations.

  1. It keeps data in memory, so very large data sets can exhaust RAM unless a database or a big-data package is used.
  2. Its learning curve is steep for beginners, and base R code is slower than compiled languages for heavy loops.

Answer frame. Open with the definition; list the features in the order free and open source, vector operations, built-in statistics, graphics, CRAN packages, RStudio; give the data-structure table and the small mean and regression example; close with advantages against limitations.

Pitfall: R uses <- for assignment, indexes from 1 (not 0), and is case sensitive, so Mean(x) fails where mean(x) works.

Last-minute revision

  • R is a free, open-source language and environment for statistical computing and graphics.
  • R was created by Ross Ihaka and Robert Gentleman.
  • R is interpreted and vector-based, so operations apply to whole vectors.
  • CRAN is the repository of R packages.
  • Packages are installed with install.packages() and loaded with library().
  • Assignment uses <-, and indexing starts at 1.
  • The data frame is the main table structure, and each column may have a different type.
  • The other structures are vector, matrix, list, array and factor.
  • read.csv() imports data and summary() gives the five-number summary and the mean.
  • ggplot2 is for graphics and dplyr is for data manipulation.
  • lm(y ~ x) fits a linear regression and cor(x, y) gives the correlation.
  • RStudio is the standard IDE for R.

Memory hooks

  • R = Really good at Reports, stats and plots.
  • CRAN = the Central store of R packages: install once, library each time.
  • Data frame = an Excel sheet inside R.
  • Workflow = Import, Clean, Explore, Model, Visualize (I Can Explore Many Views).
  • Arrow <- puts the value into the name; counting starts at 1.

Coverage checklist

  • Introduction to R as a data analytics tool: no past questions are tagged; the definition, key points, data structures, examples, packages and advantages above cover any question on it.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in