How unit 1 is examined
Covers data types, the data science road map, wrangling, exploratory analysis, basic charts, descriptive statistics and probability; structured vs unstructured data, wrangling, EDA, charts and conditional probability carry the marks.
Types of Data: structured and unstructured data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>Data Science is an interdisciplinary field that combines statistics, computing and domain knowledge to extract useful insight from data.</mark> Structured data follows a fixed schema of rows and columns and is stored in relational tables; unstructured data has no predefined schema.
Key points.
- Structured data has a fixed schema, so every record has the same named fields with defined types, for example a student table with roll number, name and marks.
- It is stored in relational databases and queried with SQL, which makes searching, sorting and joining fast and reliable.
- Its advantages are easy storage, easy analysis and mature tools; its disadvantages are rigidity, since a schema change is costly, and the fact that it captures only a small part of real-world data.
- Unstructured data has no schema or fixed model: text, emails, images, audio, video and social media posts are examples.
- It is stored as files, in data lakes or in NoSQL stores, and needs special processing such as NLP, computer vision or speech recognition before analysis.
- Its advantages are richness and volume, since most real-world data is of this kind; its disadvantages are difficult storage, search and analysis, and higher processing cost.
- Semi-structured data (JSON, XML) sits between the two, with tags but no rigid table.
- Applications: structured data runs banking, payroll and inventory; unstructured data runs sentiment analysis, medical imaging and voice assistants.
| Basis | Structured | Unstructured |
|---|---|---|
| Format | Rows and columns, fixed schema | No fixed format: text, image, audio, video |
| Storage | RDBMS, data warehouse | Data lake, NoSQL, file systems |
| Processing | SQL, simple queries | NLP, image and speech processing, ML |
| Tools | MySQL, Oracle, Excel | Hadoop, MongoDB, Spark |
| Analysis | Easy, direct | Hard, needs preprocessing first |
| Share of data | About 20% | About 80% |
| Example | Bank transactions | Customer reviews, photos |
Answer frame. Open with the definition of Data Science (Q4 only) and then of both data types; draw no figure; give the characteristics, advantages and disadvantages of each in points 1-6; end with the comparison table and applications; close with "Structured data is easy to analyse but limited, while unstructured data is rich but needs heavy processing."
Asked: [7 marks] (Jun 2023, Jun 2025) Explain the difference between structured and unstructured data, with examples. Asked: [8 marks] (Dec 2024) What is Data Science? Explain the difference between structured and unstructured data. Asked: [4 marks] (Jun 2024) Structured Data. Asked: [7 marks] (Jun 2026) What is Structured Data and Unstructured Data. Explain their characteristics, advantages, disadvantages and real-world applications with examples.
Data Science Road Map: Frame the Problem
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. The road map is the life cycle of a data science project: frame the problem, understand the data, wrangle, explore, extract features, then model and deploy.
Key points.
- Framing means turning a vague business question into a precise, measurable data question with a success metric, for example "predict which customers will leave next month".
- The phases run in sequence and loop back: results of one phase often send the team back to an earlier one.
- Each phase has an example: framing (churn prediction), data understanding (check customer tables), wrangling (fix missing ages), EDA (plot churn by plan), features (tenure), model and deploy (classifier behind an API).
- Applications of Data Science: healthcare (disease prediction), finance (fraud detection), marketing (recommendations), transport (route optimisation) and e-commerce (demand forecasting).
Asked: [7 marks] (Jun 2026) Explain the Data Science Road Map in detail. Discuss each phase with suitable examples. Asked: [7 marks] (Jun 2023) Define Data Science? Explain the applications of Data Science.
Understand the Data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Understanding the data means profiling what has been collected before using it: its source, size, fields, types and quality.
Key points.
- Check the source, the collection method and how recent the data is.
- List each variable with its type: numeric, categorical, text or date.
- Profile quality by counting missing values, duplicates and obvious errors.
- Read the data dictionary and confirm what each field means with the domain expert.
Data Wrangling
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Data wrangling (munging) is the process of discovering, structuring, cleaning, enriching and validating raw data so that it becomes reliable and ready for analysis.</mark>
Steps.
- Discovering: study the raw data to see what it contains and what problems it has.
- Structuring: reshape the data into a usable form such as a table with proper columns.
- Cleaning: handle missing values, remove duplicates, correct errors and treat outliers.
- Enriching: add useful data, such as joining another table or deriving a new column.
- Validating: check consistency, types and ranges, then publish the clean data.
Key points.
- It is important because analysis on dirty data gives wrong results: garbage in, garbage out.
- Analysts spend most of a project's time, often around 60-80%, on wrangling.
- Clean data improves model accuracy and makes results repeatable.
- Tools: Python Pandas, Excel, OpenRefine, SQL and R.
- Example workflow: load a CSV with Pandas, drop duplicates, fill missing ages with the median, convert date strings to dates, merge with a city table, save the clean file.
Answer frame. Open with the definition; list the five steps in order with one line each; explain importance; give tools and the example workflow; close with "Wrangling turns raw data into quality data and so decides the quality of every later result."
Asked: [9 marks] (Jun 2023, Jun 2024) What is Data Wrangling in data science and why it is important? Explain. / Write about Data Wrangling.
Exploratory Analysis
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>Exploratory Data Analysis (EDA) is the initial investigation of a dataset, using summary statistics and visualisation, to discover patterns, spot anomalies, test assumptions and form hypotheses before formal modelling.</mark>
Key points.
- Its objectives are to understand the structure and distribution of the data, detect outliers and missing values, and find relationships between variables.
- Step 1 is data collection, loading the data into a tool such as Pandas.
- Step 2 is data cleaning, fixing missing values, duplicates and wrong types.
- Step 3 is summary statistics: mean, median, mode, range, standard deviation, and counts for categories.
- Step 4 is visualisation: histograms and box plots for distribution, scatter plots for relationships, bar and pie charts for categories, heat maps for correlation.
- Step 5 is drawing conclusions, that is, patterns and hypotheses to test, and the choice of features and models.
- Tools include Python (Pandas, Matplotlib, Seaborn), R and Excel.
- EDA is done before modelling, so it prevents wrong assumptions and guides feature selection.
Example. For house prices, a histogram shows skewed prices, a scatter plot of area against price shows a rising trend, and a box plot reveals a few very expensive outliers.
Answer frame. Open with the definition and objectives; sketch a small flow of collect, clean, summarise, visualise, conclude; develop points 2-6 in order and then tools; close with "EDA lets the data suggest hypotheses before any model is built." For the 14-mark choose two of the four parts and write each at this depth.
Asked: [14 marks] (Jun 2023) Explain any two of the following: a) Exploratory Analysis b) Data Validation Techniques c) Regular Expression d) Data Analysis vs. Data Scientist. Asked: [6 marks] (Dec 2024) Discuss about exploratory data analysis.
Extract Features
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Feature extraction (feature engineering) creates or selects the input variables that best represent the problem for a model.
Key points.
- New features are derived from raw ones, for example age from date of birth or day of week from a date.
- Categorical data is encoded, for example one-hot encoding, and numbers are scaled.
- Irrelevant or redundant features are dropped, which reduces noise and overfitting.
- Good features often improve a model more than a fancier algorithm.
Model and Deploy Code
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Modelling fits an algorithm to the features to predict or explain the target; deployment puts the trained model into real use.
Key points.
- Split the data into training and test sets, then train a model such as regression or a decision tree.
- Evaluate it with metrics such as accuracy or RMSE, and tune it.
- Deploy it as an API, app or batch job so users can obtain predictions.
- Monitor it after deployment and retrain when the data changes.
Graphical Summaries of Data: Pie Chart
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. A pie chart is a circle divided into slices whose angles are proportional to each category's share of the whole.
Key points.
- Slice angle = (value / total) $\times 360^\circ$, and the slices add to 100%.
- It shows composition and proportion, not exact comparison.
- It suits few categories, about five or six, and one variable.
- For the branch data (CSE 120, ECE 90, MECH 60, CIVIL 30, AI&DS 100; total 400) the shares are 30%, 22.5%, 15%, 7.5%, 25% and the angles $108^\circ, 81^\circ, 54^\circ, 27^\circ, 90^\circ$.
Asked: [9 marks] (Jun 2024, Jun 2026) Explain the following: i) Pie-Chart ii) Bar Graph iii) Histogram. / Explain graphical methods of data representation with suitable examples.
Bar Graph
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. A bar graph uses separated rectangular bars whose lengths are proportional to the values of categories, to compare categorical data.
Key points.
- Bars have equal width and gaps between them, since the categories are separate.
- It is best for comparing counts or amounts across categories, such as students per branch.
- Label both axes and give a title; bars can be vertical or horizontal.
- For the branch data, CSE is the tallest bar (120) and CIVIL the shortest (30); a pie chart plus a bar graph is a good pair for the paper's question.
Asked: [7 marks] (Jun 2025) Draw two graphical representations (Pie Chart, Bar Graph, Histogram) of the branch-wise data: CSE 120, ECE 90, MECH 60, CIVIL 30, AI & DS 100.
Pareto Chart
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. A Pareto chart is a bar graph with bars sorted in descending order and a cumulative-percentage line on top.
Key points.
- It follows the 80-20 rule: about 80% of effects come from 20% of causes.
- It identifies the few "vital" causes, such as the main defect types, to fix first.
- Bars show frequency on the left axis and the line shows cumulative % on the right axis.
- It is used in quality control and problem prioritisation.
Histogram
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. A histogram shows the frequency distribution of continuous data using adjacent bars over class intervals (bins).
Key points.
- Bars touch each other because the data is continuous, unlike a bar graph.
- Bar height (or area) shows the frequency in each bin.
- It reveals shape: symmetric, skewed, or with several peaks, and outliers.
- The bin width affects the picture, so choose it with care.
Measures of central tendency of Quantitative Data: Mean
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. ==The mean is the sum of all observations divided by their number: $\bar{x} = \frac{\sum x_i}{n}$.== Quantitative data is numeric and is evaluated by descriptive measures (mean, median, mode, spread), visual summaries (histogram, box plot) and inferential methods (estimation, hypothesis tests).
Example. Data 5, 8, 12, 15, 20, 20, 25 ($n = 7$).
| Measure | Working | Result |
|---|---|---|
| Mean | $(5+8+12+15+20+20+25)/7 = 105/7$ | 15 |
| Median | Ordered data, middle (4th) value | 15 |
| Mode | Most frequent value | 20 |
Interpretation. Mean and median are both 15 and the mode is 20, which is slightly above them, so the data is nearly symmetric, with a small left (negative) skew. The mean is affected by extreme values.
Asked: [7 marks] (Jun 2025) Calculate and interpret the mean, median and mode for the given quantitative data: 5, 8, 12, 15, 20, 20, 25. Asked: [8 marks] (Dec 2024) What are the methods to evaluate Quantitative data? Explain.
Median
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. The median is the middle value of the ordered data.
Key points.
- For odd $n$ it is the $\frac{n+1}{2}$-th value; for even $n$ it is the average of the two middle values.
- It is not affected by outliers, so it suits skewed data such as income.
- Data must be sorted first.
Mode
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. The mode is the value that occurs most often.
Key points.
- A dataset can have one mode, several modes, or none.
- It is the only average that works for categorical data, such as the most common branch.
- In a histogram it is the tallest bar.
Measures of Variability of Quantitative Data: Range
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Range = maximum value minus minimum value; it measures the total spread.
Key points.
- For 5 to 25, range = $25 - 5 = 20$.
- It is quick to find but depends only on the two extreme values.
- A single outlier can change it a lot.
Standard Deviation and Variance
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Variance is the average squared deviation from the mean; standard deviation is its square root, in the same unit as the data.
Key points.
- Population: $\sigma^2 = \frac{\sum (x_i - \mu)^2}{N}$; sample: $s^2 = \frac{\sum (x_i - \bar{x})^2}{n-1}$, and $s = \sqrt{s^2}$.
- For 5, 8, 12, 15, 20, 20, 25 with mean 15, $\sum (x - \bar{x})^2 = 308$, so $s^2 = 308/6 \approx 51.33$ and $s \approx 7.16$.
- A large value means widely spread data; zero means all values are equal.
Probability: Introduction to Probability
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Probability measures how likely an event is: $P(A) = \frac{\text{favourable outcomes}}{\text{total outcomes}}$.
Key points.
- The value lies between 0 (impossible) and 1 (certain).
- The sample space is the set of all possible outcomes; an event is a subset of it.
- $P(A) + P(A') = 1$, and $P(A \cup B) = P(A) + P(B) - P(A \cap B)$.
- Example: for a fair die, $P(\text{even}) = 3/6 = 1/2$.
Conditional Probability
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Conditional probability $P(A|B)$ is the probability of event $A$ given that event $B$ has already occurred.</mark>
Formula.
$$P(A|B) = \frac{P(A \cap B)}{P(B)}, \quad P(B) > 0$$
Key points.
- The condition $B$ shrinks the sample space to the outcomes of $B$ only.
- If $A$ and $B$ are independent, $P(A|B) = P(A)$; otherwise the two differ.
- It gives the multiplication rule $P(A \cap B) = P(B)\,P(A|B)$ and leads to Bayes' theorem.
- In data science it underlies spam filters, Naive Bayes classifiers and recommendation systems.
| Basis | Basic probability | Conditional probability |
|---|---|---|
| Meaning | Chance of an event in the whole sample space | Chance of an event given another has occurred |
| Formula | $P(A) = n(A)/n(S)$ | $P(A|B) = P(A \cap B)/P(B)$ |
| Sample space | Full space $S$ | Reduced space $B$ |
| Dependence | Ignores other events | Depends on the given event |
| Example | $P(\text{even}) = 3/6 = 1/2$ | $P(\text{even}|>3) = 2/3$ |
Example. A fair die is rolled. Let $A$ = even number, $B$ = number greater than 3, so $B = \{4,5,6\}$ and $A \cap B = \{4,6\}$. Then $P(B) = 3/6$, $P(A \cap B) = 2/6$, and $P(A|B) = (2/6)/(3/6)$ = 2/3, whereas basic $P(A) = 1/2$.
Answer frame. Open with the definition and formula; give the die example with the working; for the comparison question, add the table and one numerical for each type; close with "Conditioning changes the sample space, so $P(A|B)$ generally differs from $P(A)$."
Asked: [5 marks] (Jun 2024) Explain the concept of Conditional Probability. Asked: [7 marks] (Jun 2026) Differentiate Basic Probability and Conditional Probability with suitable numerical examples.
Last-minute revision
- Data Science = statistics + computing + domain knowledge, used to extract insight from data.
- Structured data has a fixed schema and lives in an RDBMS queried with SQL; unstructured data has no schema (text, image, audio, video).
- Wrangling steps: discover, structure, clean, enrich, validate.
- EDA steps: collect, clean, summarise, visualise, conclude.
- Pie chart angle = value/total $\times 360^\circ$; bar graph has gaps, histogram bars touch.
- Pareto chart is sorted bars plus a cumulative line (80-20 rule).
- Data 5, 8, 12, 15, 20, 20, 25: mean 15, median 15, mode 20, range 20.
- Sample variance divides by $n-1$; SD is the square root of variance.
- $P(A|B) = P(A \cap B)/P(B)$.
- Branch data 120, 90, 60, 30, 100 has total 400.
Memory hooks
- Wrangling steps: "DSCEV" - Discover, Structure, Clean, Enrich, Validate.
- Structured = Table with SQL; Unstructured = Text, Tweets, Time-lapse videos (no table).
- Histogram bars touch, bar graph bars stand apart.
- Mean-median-mode: the mode is the "most popular", the median is the "middle child".
Coverage checklist
- Types of Data: structured and unstructured data - Q Jun 2023/Jun 2025, Dec 2024, Jun 2024, Jun 2026.
- Data Science Road Map: Frame the Problem - Jun 2026 road map; Jun 2023 Data Science applications.
- Understand the Data - covered.
- Data Wrangling - Jun 2023/Jun 2024.
- Exploratory Analysis - Jun 2023 (14), Dec 2024.
- Extract Features - covered.
- Model and Deploy Code - covered.
- Graphical Summaries of Data: Pie Chart - Jun 2024/Jun 2026 (9 marks).
- Bar Graph - Jun 2025 diagram question.
- Pareto Chart - covered.
- Histogram - covered.
- Measures of central tendency of Quantitative Data: Mean - Jun 2025 numerical; Dec 2024 quantitative methods.
- Median - covered.
- Mode - covered.
- Measures of Variability of Quantitative Data: Range - covered.
- Standard Deviation and Variance - covered.
- Probability: Introduction to Probability - covered.
- Conditional Probability - Jun 2024, Jun 2026.