Skip to content
CS-605 · Data Analytics Lab/Important Questions

Data Analytics Lab (CS-605) - Important Questions

  1. Unit 414 Marks High Priority

    Write a Python program to perform exploratory data analysis (EDA) on a given CSV dataset. Your program should: load the data using pandas, compute summary statistics for numerical and categorical features, plot histograms for numerical features and bar plots for categorical features, create boxplots to detect outliers, and produce a correlation heatmap. Explain the purpose of each step and interpret the key outputs.

    Core practical covering exploratory data analysis (EDA) tasks commonly required in lab exams.

  2. Unit 414 Marks High Priority

    Implement data preprocessing in Python for a dataset with missing values and categorical variables. Perform the following tasks: (a) handle missing values using mean imputation for numerical features and mode imputation for categorical features; (b) encode categorical variables using one-hot encoding and label encoding; (c) scale numerical features using StandardScaler and MinMaxScaler. Demonstrate how these preprocessing choices affect the performance of a simple classifier and explain your observations.

    Standard preprocessing pipeline question frequently asked to assess practical data-cleaning skills.

  3. Unit 47 Marks High Priority

    Using pandas, write code to load a dataset and perform group-wise aggregations. Compute the mean, median and count for a numerical column grouped by a categorical column. Demonstrate method chaining to perform these operations and discuss any performance considerations.

    Typical pandas aggregation and method-chaining task checking proficiency with group operations.

  4. Unit 410 Marks High Priority

    Demonstrate how to implement linear regression from scratch in Python using NumPy. Derive the normal equation and show it in formula form $\beta = \left(X^{T} X\right)^{-1} X^{T} y$. Implement this solution and compare the coefficients and predictions with scikit-learn's LinearRegression on the same dataset.

    Core derivation and implementation question linking linear algebra to Python implementation; often appears in practical examinations.

  5. Unit 47 Marks High Priority

    Explain how to perform data visualization in Python using Matplotlib and Seaborn. Provide code to plot a scatter plot with a regression line, a Seaborn pairplot for exploratory analysis, and a correlation heatmap. Interpret these plots to identify relationships and potential multicollinearity among features.

    Visualization and interpretation skills; frequently required to explain plots and detect relationships.

  6. Unit 47 Marks High Priority

    Write Python code to split a dataset into training and testing sets, perform k-fold cross-validation using scikit-learn's KFold, and evaluate a classifier using accuracy, precision, recall and F1-score. Explain when and why cross-validation is preferred over a single train-test split.

    Standard model evaluation and validation task to assess understanding of train/test splitting and cross-validation.

  7. Unit 410 Marks Medium Priority

    Demonstrate feature selection techniques in Python: (a) a filter method using correlation thresholding, (b) a wrapper method using Recursive Feature Elimination (RFE), and (c) an embedded method using Lasso. Provide code examples and compare the selected features and resulting model performance.

    Feature selection techniques question covering filter, wrapper and embedded methods; common in lab assessments.

  8. Unit 410 Marks Medium Priority

    Explain and implement Principal Component Analysis (PCA) in Python using scikit-learn. Show how to compute principal components, present the explained variance ratio, and choose the number of components required to retain a specified amount of variance. Demonstrate transforming the dataset and discuss when PCA is appropriate.

    Dimensionality reduction via PCA is a standard question to test both theoretical understanding and implementation.

  9. Unit 47 Marks Medium Priority

    Write Python code to detect and handle outliers using the z-score method and the IQR method. Apply both methods on a sample dataset, remove or cap outliers accordingly, and show the effect on downstream model performance.

    Outlier detection and handling is a common preprocessing question focusing on practical impact.

  10. Unit 47 Marks High Priority

    Create a reproducible data analysis pipeline in Python using scikit-learn's Pipeline. The pipeline should include imputation, scaling and a classifier. Demonstrate fitting the pipeline, performing cross-validated evaluation, and explain the benefits of using pipelines in experiments.

    Building reproducible pipelines is an essential practical skill often examined in lab-based questions.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in