Bio Informatics (AL-803 (B)) - Important Questions
-
Unit 414 Marks High Priority
Explain gene finding strategies: ab initio methods and evidence-based (extrinsic) approaches. Compare their principles, advantages and limitations, and describe scenarios where a hybrid approach is preferred.
Core topic: frequently asked in Unit 4 covering foundational strategies for gene finding.
-
Unit 410 Marks High Priority
Describe the probabilistic models used in ab initio gene prediction with emphasis on Hidden Markov Models (HMMs). Explain how states, transitions and emission probabilities are used to model gene structure (exons, introns, intergenic regions) and how the Viterbi algorithm is applied for gene prediction.
Core conceptual question on probabilistic modelling frequently emphasised in past papers.
-
Unit 47 Marks High Priority
List and compare commonly used gene prediction tools such as GENSCAN, AUGUSTUS, GeneMark, Glimmer and SNAP. For each tool indicate the primary algorithmic approach (e.g., HMM, heuristic, SVM), typical input data, and typical applications.
Repeated focus area: tools and algorithms for gene prediction are core to the unit.
-
Unit 47 Marks High Priority
How are splice sites and promoter regions predicted computationally? Describe common sequence signals, position weight matrices (PWMs) or motif models used, scoring schemes, and how these predictions are integrated into gene prediction pipelines.
Important subtopic: signal detection (splice sites, promoters) required for accurate gene models.
-
Unit 414 Marks High Priority
Outline a typical computational pipeline for mining gene expression data from raw sequencing counts to differential expression results. Include steps for quality control, filtering, normalization, statistical testing, multiple testing correction, visualization and mention commonly used tools (e.g., FastQC, DESeq2, edgeR, limma).
Core pipeline question for mining expression data; integrates many common steps and tools.
-
Unit 47 Marks High Priority
Explain methods for normalization of gene expression data. Compare RPKM/FPKM and TPM for within-sample normalization and discuss count-based normalization approaches such as DESeq2 size factors and edgeR TMM for between-sample normalization.
Normalization is repeatedly tested; comparisons between RPKM/FPKM/TPM and count-based normalization matter.
-
Unit 47 Marks High Priority
Describe clustering techniques used in gene expression analysis, specifically hierarchical clustering and k-means clustering. Explain choice of distance measures (e.g., Euclidean, Pearson correlation), linkage criteria for hierarchical clustering, and how results are interpreted via heatmaps and dendrograms.
Clustering and visualization are commonly asked practical analysis topics in Unit 4.
-
Unit 47 Marks High Priority
Discuss functional enrichment analysis after obtaining a set of differentially expressed genes. Explain Gene Ontology (GO) enrichment and pathway analysis approaches and describe the necessity of multiple testing correction (for example Benjamini–Hochberg) when reporting adjusted $p$-values and thresholds such as adjusted $p < 0.05$ and fold change criteria (e.g., $\log_2$ fold change).
Functional interpretation and multiple-testing correction are frequently required follow-ups after DE analysis.
Quick Add to Notes
Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.
Create free accountHave an account? Log in
Notes Panel