Skip to content
AD-703 (A) · Data Visualization/Quick Revision Short Notes

Data Visualization (AD-703 (A)) - Unit 1 Short Notes

How unit 1 is examined

This unit covers what data visualization is, why and where it is used, and how raw data is prepared for it; the marks sit in data pre-processing (aggregation, integration, k-NN missing-value imputation).

Overview of data visualization, Definition

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Data visualization is the graphical representation of data and information using charts, graphs, maps and dashboards, so that patterns, trends and outliers become easy to see and understand.</mark>

Key points.

  1. The human eye grasps shapes, colours and positions far faster than rows of numbers, so a picture makes large data understandable at a glance.
  2. Common forms are bar charts, line charts, scatter plots, histograms, pie charts, heat maps and maps.
  3. It serves two purposes: exploration (the analyst finds patterns) and explanation (the analyst communicates a finding to others).
  4. A good visualization is accurate, clear and purposeful; a chart that misleads is worse than none.

Significance in AI and Data Science

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Significance means the role visualization plays in the data science workflow: it turns data and model output into insight that people can act on.

Key points.

  1. In exploratory analysis, plots reveal distributions, outliers, missing values and relationships before any model is built.
  2. It guides feature selection and data cleaning by showing which variables matter and which are faulty.
  3. In AI, plots such as loss curves, confusion matrices and feature-importance charts help explain and debug models.
  4. It communicates results to non-technical decision makers, which supports trust and better decisions.

Principal of Data Visualization, Methodology

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. The principles are the design rules that make a visualization truthful and readable; the methodology is the step-wise process of producing it.

Key points.

  1. Know the audience and the question first, then choose the chart type that fits the data (comparison, trend, distribution, relationship, composition).
  2. Keep it simple: remove clutter and decoration so that the data carries the message (high data-ink ratio, after Tufte).
  3. Be honest: use a proper axis scale starting from a sensible baseline, and do not distort proportions.
  4. Use colour, size and position purposefully, label axes and units, and add a title and legend.
  5. Methodology: acquire data, clean and transform it, choose the visual encoding, build the chart, then refine and interpret.

Applications

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Applications are the fields in which visual representation of data supports analysis and decisions.

Key points.

  1. Business: sales, marketing and finance dashboards track KPIs, trends and forecasts.
  2. Healthcare: patient trends, disease spread maps and medical imaging visualization.
  3. Science and engineering: simulation results, weather and climate maps, sensor data.
  4. Government and social use: census and election maps, and public-policy reporting.
  5. AI and machine learning: model diagnostics and monitoring of data quality.

Data pre-processing for visualization: Extraction, Cleaning, Transformation, Aggregation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Data pre-processing is the set of steps that converts raw, incomplete and inconsistent data into clean, consistent, summarised data that is ready to be visualized.</mark>

<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u1-01" viewBox="0 0 467 80" width="467" height="80" role="img" aria-label="Pre-processing pipeline: Extraction, Cleaning, Transformation, Aggregation"><style>#dsfig-u1-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u1-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u1-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u1-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u1-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u1-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u1-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u1-01 .t{fill:#16181D;font-weight:500}#dsfig-u1-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u1-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u1-01 .dot{fill:#16181D}#dsfig-u1-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u1-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u1-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u1-01 .ah{fill:#454C5A}#dsfig-u1-01 .ah.hi{fill:#2340B8}#dsfig-u1-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u1-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u1-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u1-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u1-01 .e{stroke:#B1B7C3}html.dark #dsfig-u1-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u1-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u1-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u1-01 .t{fill:#E6E8ED}html.dark #dsfig-u1-01 .t.inv{fill:#0F1115}html.dark #dsfig-u1-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u1-01 .dot{fill:#E6E8ED}html.dark #dsfig-u1-01 .ann{fill:#8FA3FF}html.dark #dsfig-u1-01 .lbl{fill:#858D9C}html.dark #dsfig-u1-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u1-01 .ah{fill:#B1B7C3}html.dark #dsfig-u1-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u1-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u1-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u1-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L148,40" marker-end="url(#ah1)"/><path class="e" d="M188,40 L277,40" marker-end="url(#ah1)"/><path class="e" d="M317,40 L406,40" marker-end="url(#ah1)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Ext</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">Cln</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">Trn</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">Agg</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Pre-processing pipeline: Extraction, Cleaning, Transformation, Aggregation</figcaption></figure>

Key points.

  1. Extraction pulls the needed records and fields from sources such as databases, files, web pages and APIs.
  2. Cleaning removes noise and duplicates, corrects errors and inconsistencies, and handles missing values by deleting, filling with mean or median, or imputing with k-NN.
  3. Transformation changes the form of data: normalisation or scaling, encoding categories, type conversion, log transform and derived columns.
  4. Aggregation summarises many records into fewer ones using sum, count, average, min or max over groups or time, for example daily sales rolled up to monthly sales per region.
  5. Aggregation levels are hierarchical (day, month, year; city, state, country), so the analyst can drill down or roll up.
  6. Aggregation and integration together reduce clutter, speed up rendering of large data, and make trends visible.

Data integration (Q1). Integration combines data from heterogeneous sources (a sales database, a CSV of customers, a web API) into one consistent view by matching keys, resolving naming and unit conflicts and removing duplicates. Example: join Sales(cust_id, amount) with Customers(cust_id, region), then aggregate SUM(amount) by region and draw one bar chart of sales per region.

Missing values with k-NN (Q2). Missing values are absent entries in a record. In heterogeneous (mixed numeric and categorical) data, use a mixed distance such as Gower: numeric features contribute $|x-y|/\text{range}$, categorical features contribute 0 if equal and 1 if different, and the distance is their average.

Step 1: Find the record with the missing value.
Step 2: Compute the mixed distance from it to every complete record using the known features.
Step 3: Select the k nearest records.
Step 4: Fill with the mean of their values (numeric) or the majority value (categorical).

Example. Fill Income for R5 (Age 30, Bhopal), with Age range 20 (25 to 45), k = 2:

Record Age City Income Distance to R5
R1 25 Bhopal 20 (5/20 + 0)/2 = 0.125
R2 32 Bhopal 30 (2/20 + 0)/2 = 0.05
R3 45 Indore 60 (15/20 + 1)/2 = 0.875
R4 28 Indore 24 (2/20 + 1)/2 = 0.55

Nearest two are R2 and R1, so Income of R5 = (30 + 20)/2 = 25.

Answer frame. Open with the definition of aggregation and of integration (or of missing values); draw the pipeline diagram; develop the definition, types or steps, then the example, in that order; close with the benefit: cleaner, faster and more insightful charts.

Pitfall: Do not use plain Euclidean distance on categorical columns, and do not forget to scale numeric features before k-NN.

Asked: [7 marks] (Dec 2024) Explain Aggregation and Data Integration in visualization with help of example. Asked: [7 marks] (Dec 2024) Explain how you will find missing values in heterogeneous data using k-nearest neighbors?

Data Integration

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Data integration combines data from multiple sources into one unified, consistent dataset for analysis.

Key points.

  1. Sources may differ in format, schema, units and naming, so schema matching and key matching are needed.
  2. Entity resolution removes duplicates that refer to the same real object.
  3. Conflicts in values, such as different units or date formats, are resolved by standardising them.
  4. It is commonly done through ETL (extract, transform, load) into a data warehouse.

Data Reduction

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Data reduction produces a smaller representation of the data that gives nearly the same analytical results.

Key points.

  1. Dimensionality reduction cuts the number of attributes, for example by PCA or feature selection.
  2. Numerosity reduction cuts the number of records, using sampling, histograms or clustering.
  3. Data compression encodes the data in less space.
  4. Reduction makes large data faster to plot and less cluttered.

Last-minute revision

  • Data visualization is the graphical representation of data to reveal patterns, trends and outliers.
  • Two purposes: exploration and explanation.
  • Pre-processing order: Extraction, Cleaning, Transformation, Aggregation.
  • Aggregation summarises records with sum, count, average, min, max over groups or time.
  • Integration merges heterogeneous sources into one consistent view (joins, key matching, ETL).
  • k-NN imputation: find k nearest complete records, then average (numeric) or take the majority (categorical).
  • Mixed data uses Gower distance: numeric $|x-y|/\text{range}$, categorical 0 or 1.
  • Worked example: k = 2 neighbours with incomes 30 and 20 give a filled value of 25.
  • Data reduction: dimensionality, numerosity, compression.
  • Tufte principle: maximise the data-ink ratio.

Memory hooks

  • ECTA: Extract, Clean, Transform, Aggregate.
  • Gower: numbers by distance, words by match (0 or 1).
  • Reduction has three routes: fewer columns, fewer rows, fewer bytes.
  • Integration is "many sources, one truth".

Coverage checklist

  • Overview of data visualization, Definition: covered.
  • Significance in AI and Data Science: covered.
  • Principal of Data Visualization, Methodology: covered.
  • Applications: covered.
  • Data pre-processing for visualization: Extraction, Cleaning, Transformation, Aggregation: covers Dec 2024 Q1 (aggregation and integration) and Dec 2024 Q2 (k-NN missing values).
  • Data Integration: covered, also answered under pre-processing (Q1).
  • Data Reduction: covered.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in