How unit 1 is examined
This unit covers what data visualization is, why and where it is used, and how raw data is prepared for it; the marks sit in data pre-processing (aggregation, integration, k-NN missing-value imputation).
Overview of data visualization, Definition
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Data visualization is the graphical representation of data and information using charts, graphs, maps and dashboards, so that patterns, trends and outliers become easy to see and understand.</mark>
Key points.
- The human eye grasps shapes, colours and positions far faster than rows of numbers, so a picture makes large data understandable at a glance.
- Common forms are bar charts, line charts, scatter plots, histograms, pie charts, heat maps and maps.
- It serves two purposes: exploration (the analyst finds patterns) and explanation (the analyst communicates a finding to others).
- A good visualization is accurate, clear and purposeful; a chart that misleads is worse than none.
Significance in AI and Data Science
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Significance means the role visualization plays in the data science workflow: it turns data and model output into insight that people can act on.
Key points.
- In exploratory analysis, plots reveal distributions, outliers, missing values and relationships before any model is built.
- It guides feature selection and data cleaning by showing which variables matter and which are faulty.
- In AI, plots such as loss curves, confusion matrices and feature-importance charts help explain and debug models.
- It communicates results to non-technical decision makers, which supports trust and better decisions.
Principal of Data Visualization, Methodology
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. The principles are the design rules that make a visualization truthful and readable; the methodology is the step-wise process of producing it.
Key points.
- Know the audience and the question first, then choose the chart type that fits the data (comparison, trend, distribution, relationship, composition).
- Keep it simple: remove clutter and decoration so that the data carries the message (high data-ink ratio, after Tufte).
- Be honest: use a proper axis scale starting from a sensible baseline, and do not distort proportions.
- Use colour, size and position purposefully, label axes and units, and add a title and legend.
- Methodology: acquire data, clean and transform it, choose the visual encoding, build the chart, then refine and interpret.
Applications
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Applications are the fields in which visual representation of data supports analysis and decisions.
Key points.
- Business: sales, marketing and finance dashboards track KPIs, trends and forecasts.
- Healthcare: patient trends, disease spread maps and medical imaging visualization.
- Science and engineering: simulation results, weather and climate maps, sensor data.
- Government and social use: census and election maps, and public-policy reporting.
- AI and machine learning: model diagnostics and monitoring of data quality.
Data pre-processing for visualization: Extraction, Cleaning, Transformation, Aggregation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Data pre-processing is the set of steps that converts raw, incomplete and inconsistent data into clean, consistent, summarised data that is ready to be visualized.</mark>
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u1-01" viewBox="0 0 467 80" width="467" height="80" role="img" aria-label="Pre-processing pipeline: Extraction, Cleaning, Transformation, Aggregation"><style>#dsfig-u1-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u1-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u1-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u1-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u1-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u1-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u1-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u1-01 .t{fill:#16181D;font-weight:500}#dsfig-u1-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u1-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u1-01 .dot{fill:#16181D}#dsfig-u1-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u1-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u1-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u1-01 .ah{fill:#454C5A}#dsfig-u1-01 .ah.hi{fill:#2340B8}#dsfig-u1-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u1-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u1-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u1-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u1-01 .e{stroke:#B1B7C3}html.dark #dsfig-u1-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u1-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u1-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u1-01 .t{fill:#E6E8ED}html.dark #dsfig-u1-01 .t.inv{fill:#0F1115}html.dark #dsfig-u1-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u1-01 .dot{fill:#E6E8ED}html.dark #dsfig-u1-01 .ann{fill:#8FA3FF}html.dark #dsfig-u1-01 .lbl{fill:#858D9C}html.dark #dsfig-u1-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u1-01 .ah{fill:#B1B7C3}html.dark #dsfig-u1-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u1-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u1-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u1-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L148,40" marker-end="url(#ah1)"/><path class="e" d="M188,40 L277,40" marker-end="url(#ah1)"/><path class="e" d="M317,40 L406,40" marker-end="url(#ah1)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Ext</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">Cln</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">Trn</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">Agg</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Pre-processing pipeline: Extraction, Cleaning, Transformation, Aggregation</figcaption></figure>
Key points.
- Extraction pulls the needed records and fields from sources such as databases, files, web pages and APIs.
- Cleaning removes noise and duplicates, corrects errors and inconsistencies, and handles missing values by deleting, filling with mean or median, or imputing with k-NN.
- Transformation changes the form of data: normalisation or scaling, encoding categories, type conversion, log transform and derived columns.
- Aggregation summarises many records into fewer ones using sum, count, average, min or max over groups or time, for example daily sales rolled up to monthly sales per region.
- Aggregation levels are hierarchical (day, month, year; city, state, country), so the analyst can drill down or roll up.
- Aggregation and integration together reduce clutter, speed up rendering of large data, and make trends visible.
Data integration (Q1). Integration combines data from heterogeneous sources (a sales database, a CSV of customers, a web API) into one consistent view by matching keys, resolving naming and unit conflicts and removing duplicates. Example: join Sales(cust_id, amount) with Customers(cust_id, region), then aggregate SUM(amount) by region and draw one bar chart of sales per region.
Missing values with k-NN (Q2). Missing values are absent entries in a record. In heterogeneous (mixed numeric and categorical) data, use a mixed distance such as Gower: numeric features contribute $|x-y|/\text{range}$, categorical features contribute 0 if equal and 1 if different, and the distance is their average.
Step 1: Find the record with the missing value.
Step 2: Compute the mixed distance from it to every complete record using the known features.
Step 3: Select the k nearest records.
Step 4: Fill with the mean of their values (numeric) or the majority value (categorical).
Example. Fill Income for R5 (Age 30, Bhopal), with Age range 20 (25 to 45), k = 2:
| Record | Age | City | Income | Distance to R5 |
|---|---|---|---|---|
| R1 | 25 | Bhopal | 20 | (5/20 + 0)/2 = 0.125 |
| R2 | 32 | Bhopal | 30 | (2/20 + 0)/2 = 0.05 |
| R3 | 45 | Indore | 60 | (15/20 + 1)/2 = 0.875 |
| R4 | 28 | Indore | 24 | (2/20 + 1)/2 = 0.55 |
Nearest two are R2 and R1, so Income of R5 = (30 + 20)/2 = 25.
Answer frame. Open with the definition of aggregation and of integration (or of missing values); draw the pipeline diagram; develop the definition, types or steps, then the example, in that order; close with the benefit: cleaner, faster and more insightful charts.
Pitfall: Do not use plain Euclidean distance on categorical columns, and do not forget to scale numeric features before k-NN.
Asked: [7 marks] (Dec 2024) Explain Aggregation and Data Integration in visualization with help of example. Asked: [7 marks] (Dec 2024) Explain how you will find missing values in heterogeneous data using k-nearest neighbors?
Data Integration
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Data integration combines data from multiple sources into one unified, consistent dataset for analysis.
Key points.
- Sources may differ in format, schema, units and naming, so schema matching and key matching are needed.
- Entity resolution removes duplicates that refer to the same real object.
- Conflicts in values, such as different units or date formats, are resolved by standardising them.
- It is commonly done through ETL (extract, transform, load) into a data warehouse.
Data Reduction
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Data reduction produces a smaller representation of the data that gives nearly the same analytical results.
Key points.
- Dimensionality reduction cuts the number of attributes, for example by PCA or feature selection.
- Numerosity reduction cuts the number of records, using sampling, histograms or clustering.
- Data compression encodes the data in less space.
- Reduction makes large data faster to plot and less cluttered.
Last-minute revision
- Data visualization is the graphical representation of data to reveal patterns, trends and outliers.
- Two purposes: exploration and explanation.
- Pre-processing order: Extraction, Cleaning, Transformation, Aggregation.
- Aggregation summarises records with sum, count, average, min, max over groups or time.
- Integration merges heterogeneous sources into one consistent view (joins, key matching, ETL).
- k-NN imputation: find k nearest complete records, then average (numeric) or take the majority (categorical).
- Mixed data uses Gower distance: numeric $|x-y|/\text{range}$, categorical 0 or 1.
- Worked example: k = 2 neighbours with incomes 30 and 20 give a filled value of 25.
- Data reduction: dimensionality, numerosity, compression.
- Tufte principle: maximise the data-ink ratio.
Memory hooks
- ECTA: Extract, Clean, Transform, Aggregate.
- Gower: numbers by distance, words by match (0 or 1).
- Reduction has three routes: fewer columns, fewer rows, fewer bytes.
- Integration is "many sources, one truth".
Coverage checklist
- Overview of data visualization, Definition: covered.
- Significance in AI and Data Science: covered.
- Principal of Data Visualization, Methodology: covered.
- Applications: covered.
- Data pre-processing for visualization: Extraction, Cleaning, Transformation, Aggregation: covers Dec 2024 Q1 (aggregation and integration) and Dec 2024 Q2 (k-NN missing values).
- Data Integration: covered, also answered under pre-processing (Q1).
- Data Reduction: covered.