Skip to content
CS-503 (A) · Data Analytics/Quick Revision Short Notes

Data Analytics (CS-503 (A)) - Unit 3 Short Notes

How unit 3 is examined

This unit covers how big data is gathered, mapped, extracted, transformed and split before Hadoop MapReduce runs; only extracting data from storage has been asked (7 marks, Jun 2020).

Integrating disparate data stores

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Data integration combines data from different, heterogeneous stores (relational databases, files, logs, NoSQL, APIs) into one unified, consistent view for analysis.</mark>

Key points.

  1. Sources differ in format, schema and location, so structured, semi-structured and unstructured data must be brought together.
  2. Integration is done by ETL into a warehouse or data lake, or by a virtual layer that queries the sources in place.
  3. Schema mismatches, duplicate records and conflicting values must be resolved during integration.
  4. The result is a single source of truth that Hadoop and analytics tools can process.

Mapping data to the programming framework

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Mapping data to the programming framework means shaping the stored data into the input form the framework expects, which in Hadoop MapReduce is key-value pairs.</mark>

Key points.

  1. MapReduce reads every record as a (key, value) pair, so each raw record must be expressed that way.
  2. An InputFormat and RecordReader convert raw files into pairs, for example the byte offset as key and the line as value in TextInputFormat.
  3. Each field is mapped to a data type that the framework can serialise, such as Text, IntWritable or LongWritable.
  4. Correct mapping lets the Mapper process each record independently and in parallel.

Connecting and extracting data from storage

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Data extraction is the first ETL step: pulling data from its source systems and loading it into a store or warehouse for analysis.</mark>

Key points.

  1. Sources include databases (RDBMS, NoSQL), log files, web APIs, sensors and social media.
  2. The framework connects through connectors such as JDBC or ODBC, or through APIs and file readers.
  3. ETL steps are extract from the source, clean and transform (remove errors, duplicates and nulls), then load into HDFS or the warehouse.
  4. Sqoop transfers bulk data between relational databases and HDFS, and Flume collects streaming log data into HDFS.

Asked: [7 marks] (Jun 2020) How can we extract data for data storage?

Transforming data for processing

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Data transformation converts extracted data into a clean, consistent and suitable format for processing; it is the T in ETL.</mark>

Key points.

  1. Cleaning removes duplicates, corrects errors and handles missing values.
  2. Conversion changes data types, formats and units, for example text dates to a standard date format.
  3. Aggregation, filtering, joining and normalisation reshape the data for the analysis required.
  4. In Hadoop this is done by MapReduce, Pig or Hive jobs before the data is analysed.

Subdividing data in preparation for Hadoop MapReduce

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Subdividing data means splitting a large input into input splits, each processed by one map task in parallel; this is also called data partitioning.</mark>

Key points.

  1. HDFS stores files as blocks (128 MB by default), and an input split normally equals one block.
  2. The number of splits decides the number of map tasks, so splitting drives parallelism.
  3. The InputFormat computes the splits and the RecordReader turns each split into key-value pairs.
  4. Map tasks are scheduled on nodes holding the block, which gives data locality and less network traffic.

Last-minute revision

  • Data integration merges heterogeneous stores into one unified view.
  • MapReduce input is always key-value pairs, prepared by InputFormat and RecordReader.
  • ETL means Extract, Transform, Load.
  • Extraction sources are databases, logs and APIs.
  • Sqoop moves data between RDBMS and HDFS; Flume collects streaming logs.
  • Transformation cleans, converts, aggregates and normalises data.
  • One input split gives one map task; the default HDFS block is 128 MB.
  • Data locality means the map task runs where the block is stored.

Memory hooks

  • ETL: Extract, Transform, Load, in that order.
  • Sqoop is SQL to Hadoop; Flume is a flow of logs.
  • Splits equal maps: more splits, more parallel mappers.

Coverage checklist

  • Integrating disparate data stores: no past questions.
  • Mapping data to the programming framework: no past questions.
  • Connecting and extracting data from storage: How can we extract data for data storage? (Jun 2020).
  • Transforming data for processing: no past questions.
  • subdividing data in preparation for Hadoop Map Reduce: no past questions.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in