A. Advanced Linux Command-Line & System Management
A.1 Package & Software Management
-
APT (Debian/Ubuntu):
-
apt search <pkg>: Search repositories. -
apt show <pkg>: Display detailed package info. -
apt cache <policy/clean>: Manage local package cache. -
Repository config:
/etc/apt/sources.list&/etc/apt/sources.list.d/.
-
-
YUM/DNF (RHEL/CentOS/Fedora):
-
yum/dnf info <pkg>,yum/dnf provides <file>. -
yum/dnf clean allto clear cache.
-
-
Installing from Source:
-
Standard sequence:
./configure(checks dependencies, sets Makefile) →make(compiles) →sudo make install(installs). -
Key flags:
./configure --prefix=/custom/pathfor non-standard install location.
-
-
Dependency Management:
-
Tools:
apt-get install -f(fix broken deps),yum deplist <pkg>. -
Conflict Resolution: Often requires manual removal or pinning package versions.
-
[!TIP] Common Pitfall: Mixing package managers (e.g.,
aptanddpkg -i) can break dependency chains. Always prefer the native manager.
A.2 Process & System Monitoring
-
ps(Process Snapshot):-
ps aux: Full listing (BSD style). -
ps -eo pid,comm,%cpu,%mem --sort=-%cpu: Custom format, sort by CPU. -
ps -u <username>: Filter by user.
-
-
top/htop:-
Interactive:
P(sort by CPU),M(sort by MEM),k(kill),r(renice). -
Load Average: 1/5/15 min avg. Value > number of CPU cores indicates backlog.
-
-
Signals & Job Control:
-
kill -l: List all signals. -
SIGTERM(15): Graceful termination request. -
SIGKILL(9): Forceful, non-catchable termination. -
cmd &: Run in background.jobs: List bg jobs.fg %1: Bring job 1 to foreground.bg %1: Resume job 1 in background. -
nohup cmd &: Run immune to hangups, output tonohup.out.
-
-
System Load:
uptimeshows load averages.wshows users and their processes.
A.3 Disk & Filesystem Management
-
df(Disk Free) &du(Disk Usage):-
df -h: Human-readable (GB/MB). -
df -i: Show inode usage. -
du -sh <dir>: Summary, human-readable for a directory. -
du -a <dir>: Include all files.
-
-
Filesystem Types & Ops:
-
mkfs -t ext4 /dev/sdX: Create filesystem. -
fsck -f /dev/sdX: Force check filesystem (unmounted!).
-
-
Logical Volume Manager (LVM):
-
Concept: Physical Volume (PV) → Volume Group (VG) → Logical Volume (LV).
-
Key Commands:
-
pvcreate /dev/sdX -
vgcreate <vg_name> /dev/sdX -
lvcreate -L 10G -n <lv_name> <vg_name>
-
-
DiagramLVM: LVM abstraction layers showing PVs building VG, which slices into LVs
-
-
Mounting &
/etc/fstab:-
mount /dev/sdX /mntormount -t ext4 /dev/sdX /mnt. -
/etc/fstabfields:<device> <mountpoint> <fstype> <options> <dump> <pass>. -
mount -a: Mount all from fstab.
-
A.4 Networking & Connectivity
-
IP Configuration (Modern):
-
ip addr show,ip link set <iface> up/down. -
ip route add default via <gateway>. -
Legacy:
ifconfig,route(deprecated but still common).
-
-
Socket/Connection Auditing:
-
ss -tuln: Show listening TCP/UDP ports (modernnetstatreplacement). -
netstat -tunap: Show all connections with process info.
-
-
Connectivity Testing:
-
ping <host>: ICMP echo. -
traceroute <host>: Path tracing. -
mtr <host>: Combinesping+traceroute(continuous).
-
-
Secure Transfer:
-
scp file user@host:/path. -
sftp user@host: Interactive SFTP session (put,get,ls).
-
-
Basic Firewall (
ufw- Ubuntu):-
ufw allow 22/tcp,ufw deny from 192.168.1.100. -
ufw status numbered,ufw delete <rule_num>.
-
A.5 User & Permission Administration
-
chmod(Change Mode):-
Numeric (Octal):
chmod 755 file→rwxr-xr-x.4=r, 2=w, 1=x. Sum for owner/group/others.
-
Symbolic:
chmod u+x,g-w,o+r file. -
Recursive:
chmod -R 755 /dir/.
-
-
chown/chgrp:-
chown user:group file. -
chgrp group file.
-
-
Special Permissions:
-
SUID (4):
chmod 4755 /usr/bin/passwd→ Runs as file owner. -
SGID (2): On file: runs as group. On dir: new files inherit dir's group.
-
Sticky Bit (1): On dir (e.g.,
/tmp): Only file owner can delete/rename.
-
-
sudo&/etc/sudoers:-
Edit Safely:
visudo(syntax check). -
Format:
user host=(runas) command. -
Example:
%admin ALL=(ALL) ALL(group admin can run any cmd as any user on any host). -
Best Practice: Use specific commands, avoid
NOPASSWDunless necessary.
-
-
User/Group Mgmt:
-
useradd -m -s /bin/bash <user>: Create user with home dir and bash. -
usermod -aG <group> <user>: Append user to supplementary group. -
groupadd <group>.
-
A.6 Shell Scripting Deep Dive
-
Functions & Args:
-
myfunc() { echo "Args: $@"; } -
$1,$2`: Positional params. `$@: All args as separate words.$*: All args as single word. -
shift: Discard $1, shift others down.
-
-
Conditional Tests (
testor[ ]):-
File:
-f file(regular),-d dir(directory),-e path(exists). -
String:
-z "$str"(zero length),-n "$str"(non-zero). -
Numeric:
-eq,-ne,-lt,-gt.
-
-
Control Structures:
-
if [ condition ]; then ... elif [ ... ]; then ... else ... fi. -
case $var in pattern) cmd;; esac. -
Loops:
for f in *.txt; do ...; done;while [ condition ]; do ...; done;until [ condition ]; do ...; done.
-
-
Error Handling:
-
$?: Exit status of last command (0 = success). -
set -e: Exit script on any error.set -u: Treat unset vars as error. -
trap 'cleanup' EXIT: Run cleanup on script exit.
-
-
Here Documents:
-
cmd <<EOF ... EOF: Feed multi-line string to command. -
cmd <<< "string": Here string (single line).
-
B. Advanced R for Data Analysis & Visualization
B.1 Data Manipulation with dplyr & tidyr
-
dplyrCore Verbs (Grammar of Data Manipulation):| Verb | Purpose | Example | | :--- | :--- | :--- | |
filter()| Row subsetting |filter(mpg > 20, cyl == 4)| |select()| Column subsetting |select(mpg, cyl, hp)| |mutate()| New columns |mutate(disp_l = disp / 61.02)| |arrange()| Reorder rows |arrange(desc(mpg))| |summarise()| Aggregate |summarise(avg_mpg = mean(mpg))| |group_by()| Define groups |group_by(cyl) %>% summarise(...)| -
Piping (
%>%): Frommagrittr. Passes LHS as first arg to RHS.df %>% filter(cyl == 6) %>% select(mpg, hp) %>% head() -
tidyrReshaping:-
Long to Wide:
pivot_wider(names_from = key, values_from = value). -
Wide to Long:
pivot_longer(cols = c(col1, col2), names_to = "key", values_to = "value"). -
separate(col, into = c("a","b"), sep = "/")/unite(new, col1, col2, sep = "-").
-
-
Joins:
inner_join(x, y, by = "key"),left_join(),full_join(),semi_join(),anti_join().
B.2 Advanced Visualization with ggplot2
-
Grammar of Graphics Layers:
-
Data:
data = df -
Aesthetics (aes):
aes(x = var1, y = var2, color = group) -
Geometries (geom_*):
geom_point(),geom_line(),geom_bar(stat = "identity"). -
Facets (facet_*):
facet_wrap(~group)orfacet_grid(row ~ col). -
Statistics (stat_*): Often implicit in geom.
-
Coordinates (coord_*):
coord_flip(),coord_cartesian(). -
Theme (theme_*):
theme_minimal(),theme_bw().
-
-
Customization:
-
Labels:
labs(title = "Title", x = "X Label", y = "Y Label"). -
Scales:
scale_x_log10(),scale_color_manual(values = c("red","blue")). -
Saving:
ggsave("plot.png", width = 8, height = 6, dpi = 300).
-
B.3 Statistical Modeling & Reporting
-
Linear Model (
lm):model <- lm(mpg ~ wt + hp, data = mtcars) summary(model) # Shows coefficients, p-values, R-squared, F-statistic.- Key Output: Residuals, Coefficients (Estimate, Std. Error, t value, Pr(>|t|)), Residual standard error, Multiple R-squared.
-
Generalized Linear Model (
glm):-
family = binomialfor logistic regression. -
family = poissonfor count data.
-
-
ANOVA:
aov(mpg ~ cyl, data = mtcars). -
R Markdown (
.Rmd):-
YAML Header: Controls output format (
html_document,pdf_document). -
Code Chunks:
```{r chunk_name, echo=FALSE, warning=FALSE} ... ```. -
Parameterized Reports: Use
params:in YAML andparams$varin code. -
Knit:
rmarkdown::render("report.Rmd").
-
B.4 Efficient R Programming
-
Vectorization: Use built-in functions (
log(x),x + y) instead of loops. $O(n)$ vs. slower loop overhead. -
applyFamily:-
lapply(X, FUN): Returns list. -
sapply(X, FUN): Simplifies to vector/matrix. -
vapply(X, FUN, FUN.VALUE): Type-safesapply. -
mapply(FUN, X, Y): Multivariate version.
-
-
data.table(for large data):-
Syntax:
DT[i, j, by].ifor row subset,jfor columns/ops,byfor grouping. -
Key feature:
:=for in-place assignment, no copy.DT[, new_col := col1 + col2] # Fast, memory efficient setkey(DT, col1) # Index for fast binary search
-
-
Profiling:
-
Rprof("profile.out")...Rprof(NULL)thensummaryRprof(). -
profvis::profvis({ code })for interactive visualization.
-
C. Integration: Using R within the Linux Environment
C.1 Executing R from the Shell
-
Non-interactive Scripts:
-
Rscript myscript.R: Preferred. Runs script, sends output to stdout. -
R CMD BATCH myscript.R: Legacy. Saves output tomyscript.Rout.
-
-
Command-Line Arguments in R:
args <- commandArgs(trailingOnly = TRUE) if (length(args) > 0) { input_file <- args[1] } -
One-liners & Capturing Output:
-
R -e "print(summary(lm(mpg~wt, mtcars)))". -
Capture in shell var:
result=$(R -e "cat(1+1)"). -
Redirect:
Rscript analysis.R > output.txt 2>&1.
-
C.2 Automation & Scheduling
-
Cron Syntax:
min hour dom mon dow command.- Example:
0 2 * * * /usr/bin/Rscript /home/user/daily_report.R(Run daily at 2 AM).
- Example:
-
Robust Shell Wrapper (
run_report.sh):#!/bin/bash LOG="/var/log/r_report.log" /usr/bin/Rscript /path/to/script.R >> $LOG 2>&1 if [ $? -ne 0 ]; then echo "ERROR: R script failed on $$\displaystyle (date)" >> $$LOG # Send alert email fi-
Make executable:
chmod +x run_report.sh. -
Cron calls the wrapper, not the
.Rfile directly.
-
C.3 Environment & Package Management
-
Library Paths:
-
.libPaths()in R shows current search paths. -
Set via
R_LIBS_USERenvironment variable orRscript -e ".libPaths('~/Rlibs')".
-
-
CLI Package Installation:
-
In script:
if (!requireNamespace("pkg")) install.packages("pkg", repos = "https://cloud.r-project.org"). -
Command line:
R CMD INSTALL pkg_1.0.tar.gz(from source).
-
-
Project Reproducibility with
renv:-
renv::init(): Createsrenv.locksnapshot of current packages. -
renv::restore(): Reinstalls exact package versions from lockfile. -
Integration: Commit
renv.lockto Git. On new Linux machine:renv::restore().
-
C.4 File & Data Pipeline Operations
-
Shell ↔ R Data Exchange:
-
Shell to R: Pipe data.
cat data.csv | head -n 100 | Rscript -e 'df <- read.csv(file("stdin")); print(summary(df))' -
R to Shell: Capture R output as shell variable/array.
-
-
Efficient File Reading:
-
R:
data.table::fread("large.csv")is faster thanread.csv. -
Shell pre-filter:
awk -F, '$3 > 100' large.csv | Rscript process.R.
-
-
Pipeline Pattern for Logs:
zcat /var/log/syslog.1.gz | grep "ERROR" | cut -d' ' -f5- | Rscript analyze_errors.R- R script reads from
stdin(file("stdin")) for real-time analysis.
- R script reads from
[!TIP] Exam Focus: Integration questions often combine cron scheduling,
RscriptwithcommandArgs, andrenvfor reproducibility. Be ready to write a complete pipeline from shell preprocessing to R analysis and output logging.