Data Analytics Lab (CS-605) - Important Questions
-
Unit 414 Marks High Priority
Write a Python program to perform exploratory data analysis (EDA) on a given CSV dataset. Your program should: load the data using pandas, compute summary statistics for numerical and categorical features, plot histograms for numerical features and bar plots for categorical features, create boxplots to detect outliers, and produce a correlation heatmap. Explain the purpose of each step and interpret the key outputs.
Core practical covering exploratory data analysis (EDA) tasks commonly required in lab exams.
-
Unit 414 Marks High Priority
Implement data preprocessing in Python for a dataset with missing values and categorical variables. Perform the following tasks: (a) handle missing values using mean imputation for numerical features and mode imputation for categorical features; (b) encode categorical variables using one-hot encoding and label encoding; (c) scale numerical features using StandardScaler and MinMaxScaler. Demonstrate how these preprocessing choices affect the performance of a simple classifier and explain your observations.
Standard preprocessing pipeline question frequently asked to assess practical data-cleaning skills.
-
Unit 47 Marks High Priority
Using pandas, write code to load a dataset and perform group-wise aggregations. Compute the mean, median and count for a numerical column grouped by a categorical column. Demonstrate method chaining to perform these operations and discuss any performance considerations.
Typical pandas aggregation and method-chaining task checking proficiency with group operations.
-
Unit 410 Marks High Priority
Demonstrate how to implement linear regression from scratch in Python using NumPy. Derive the normal equation and show it in formula form $\beta = \left(X^{T} X\right)^{-1} X^{T} y$. Implement this solution and compare the coefficients and predictions with scikit-learn's LinearRegression on the same dataset.
Core derivation and implementation question linking linear algebra to Python implementation; often appears in practical examinations.
-
Unit 47 Marks High Priority
Explain how to perform data visualization in Python using Matplotlib and Seaborn. Provide code to plot a scatter plot with a regression line, a Seaborn pairplot for exploratory analysis, and a correlation heatmap. Interpret these plots to identify relationships and potential multicollinearity among features.
Visualization and interpretation skills; frequently required to explain plots and detect relationships.
-
Unit 47 Marks High Priority
Write Python code to split a dataset into training and testing sets, perform k-fold cross-validation using scikit-learn's KFold, and evaluate a classifier using accuracy, precision, recall and F1-score. Explain when and why cross-validation is preferred over a single train-test split.
Standard model evaluation and validation task to assess understanding of train/test splitting and cross-validation.
-
Unit 410 Marks Medium Priority
Demonstrate feature selection techniques in Python: (a) a filter method using correlation thresholding, (b) a wrapper method using Recursive Feature Elimination (RFE), and (c) an embedded method using Lasso. Provide code examples and compare the selected features and resulting model performance.
Feature selection techniques question covering filter, wrapper and embedded methods; common in lab assessments.
-
Unit 410 Marks Medium Priority
Explain and implement Principal Component Analysis (PCA) in Python using scikit-learn. Show how to compute principal components, present the explained variance ratio, and choose the number of components required to retain a specified amount of variance. Demonstrate transforming the dataset and discuss when PCA is appropriate.
Dimensionality reduction via PCA is a standard question to test both theoretical understanding and implementation.
-
Unit 47 Marks Medium Priority
Write Python code to detect and handle outliers using the z-score method and the IQR method. Apply both methods on a sample dataset, remove or cap outliers accordingly, and show the effect on downstream model performance.
Outlier detection and handling is a common preprocessing question focusing on practical impact.
-
Unit 47 Marks High Priority
Create a reproducible data analysis pipeline in Python using scikit-learn's Pipeline. The pipeline should include imputation, scaling and a classifier. Demonstrate fitting the pipeline, performing cross-validated evaluation, and explain the benefits of using pipelines in experiments.
Building reproducible pipelines is an essential practical skill often examined in lab-based questions.
Quick Add to Notes
Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.
Create free accountHave an account? Log in
Notes Panel