How unit 3 is examined
This unit covers how big data is gathered, mapped, extracted, transformed and split before Hadoop MapReduce runs; only extracting data from storage has been asked (7 marks, Jun 2020).
Integrating disparate data stores
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Data integration combines data from different, heterogeneous stores (relational databases, files, logs, NoSQL, APIs) into one unified, consistent view for analysis.</mark>
Key points.
- Sources differ in format, schema and location, so structured, semi-structured and unstructured data must be brought together.
- Integration is done by ETL into a warehouse or data lake, or by a virtual layer that queries the sources in place.
- Schema mismatches, duplicate records and conflicting values must be resolved during integration.
- The result is a single source of truth that Hadoop and analytics tools can process.
Mapping data to the programming framework
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Mapping data to the programming framework means shaping the stored data into the input form the framework expects, which in Hadoop MapReduce is key-value pairs.</mark>
Key points.
- MapReduce reads every record as a (key, value) pair, so each raw record must be expressed that way.
- An InputFormat and RecordReader convert raw files into pairs, for example the byte offset as key and the line as value in TextInputFormat.
- Each field is mapped to a data type that the framework can serialise, such as Text, IntWritable or LongWritable.
- Correct mapping lets the Mapper process each record independently and in parallel.
Connecting and extracting data from storage
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Data extraction is the first ETL step: pulling data from its source systems and loading it into a store or warehouse for analysis.</mark>
Key points.
- Sources include databases (RDBMS, NoSQL), log files, web APIs, sensors and social media.
- The framework connects through connectors such as JDBC or ODBC, or through APIs and file readers.
- ETL steps are extract from the source, clean and transform (remove errors, duplicates and nulls), then load into HDFS or the warehouse.
- Sqoop transfers bulk data between relational databases and HDFS, and Flume collects streaming log data into HDFS.
Asked: [7 marks] (Jun 2020) How can we extract data for data storage?
Transforming data for processing
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Data transformation converts extracted data into a clean, consistent and suitable format for processing; it is the T in ETL.</mark>
Key points.
- Cleaning removes duplicates, corrects errors and handles missing values.
- Conversion changes data types, formats and units, for example text dates to a standard date format.
- Aggregation, filtering, joining and normalisation reshape the data for the analysis required.
- In Hadoop this is done by MapReduce, Pig or Hive jobs before the data is analysed.
Subdividing data in preparation for Hadoop MapReduce
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Subdividing data means splitting a large input into input splits, each processed by one map task in parallel; this is also called data partitioning.</mark>
Key points.
- HDFS stores files as blocks (128 MB by default), and an input split normally equals one block.
- The number of splits decides the number of map tasks, so splitting drives parallelism.
- The InputFormat computes the splits and the RecordReader turns each split into key-value pairs.
- Map tasks are scheduled on nodes holding the block, which gives data locality and less network traffic.
Last-minute revision
- Data integration merges heterogeneous stores into one unified view.
- MapReduce input is always key-value pairs, prepared by InputFormat and RecordReader.
- ETL means Extract, Transform, Load.
- Extraction sources are databases, logs and APIs.
- Sqoop moves data between RDBMS and HDFS; Flume collects streaming logs.
- Transformation cleans, converts, aggregates and normalises data.
- One input split gives one map task; the default HDFS block is 128 MB.
- Data locality means the map task runs where the block is stored.
Memory hooks
- ETL: Extract, Transform, Load, in that order.
- Sqoop is SQL to Hadoop; Flume is a flow of logs.
- Splits equal maps: more splits, more parallel mappers.
Coverage checklist
- Integrating disparate data stores: no past questions.
- Mapping data to the programming framework: no past questions.
- Connecting and extracting data from storage: How can we extract data for data storage? (Jun 2020).
- Transforming data for processing: no past questions.
- subdividing data in preparation for Hadoop Map Reduce: no past questions.