Skip to content
AD-801 · Big Data/Quick Revision Short Notes

Big Data (AD-801) - Unit 3 Short Notes

How unit 3 is examined

This unit covers Hive (architecture, data types, HiveQL) and Pig (Pig Latin, execution, operators, functions); no topic was asked in the supplied papers, so learn the definitions and the comparisons.

Introduction to Hive

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Hive is a data-warehouse tool built on Hadoop that lets users query and analyse large data stored in HDFS using an SQL-like language called HiveQL.</mark>

Key points.

  1. Hive was developed at Facebook and is now an Apache project.
  2. HiveQL queries are converted internally into MapReduce, Tez or Spark jobs.
  3. Hive uses schema-on-read: table structure is applied when data is read, not when it is loaded.
  4. It suits batch analysis of huge data, not OLTP, because latency is high and row-level updates are limited.

Hive Architecture

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Hive architecture consists of the user interfaces, the Driver (compiler, optimizer, executor), the Metastore and the Hadoop layer (HDFS and MapReduce).</mark>

Key points.

  1. User interfaces are the CLI, Beeline, web UI and JDBC/ODBC clients connecting through HiveServer2 (Thrift).
  2. The Driver receives the query, creates a session and manages its life cycle.
  3. The Compiler parses HiveQL, checks it against the metastore and builds an execution plan; the Optimizer improves that plan.
  4. The Executor submits the plan as jobs to Hadoop and monitors them.
  5. The Metastore stores table schemas, partitions and locations in an RDBMS such as MySQL or Derby.
  6. Actual data stays in HDFS.

Hive Data types

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Hive data types are of two kinds: primitive types for single values and complex types that hold collections.</mark>

Key points.

  1. Numeric types are TINYINT, SMALLINT, INT, BIGINT, FLOAT, DOUBLE and DECIMAL.
  2. String types are STRING, VARCHAR and CHAR; BOOLEAN holds true or false.
  3. Date and time types are DATE and TIMESTAMP; BINARY holds raw bytes.
  4. Complex types are ARRAY (ordered same-type items), MAP (key-value pairs), STRUCT (named fields) and UNIONTYPE.

Hive Query Language

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>HiveQL is Hive's SQL-like language for defining, loading and querying tables stored on Hadoop.</mark>

Key points.

  1. DDL commands are CREATE, ALTER, DROP and SHOW, for example CREATE TABLE emp (id INT, name STRING).
  2. Data is loaded with LOAD DATA INPATH ... INTO TABLE or INSERT.
  3. Queries use SELECT with WHERE, GROUP BY, ORDER BY, HAVING and JOIN.
  4. Partitions and buckets split a table so queries scan less data.
  5. Tables may be managed (Hive owns the data) or external (dropping keeps the data).

Introduction to Pig

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Pig is a high-level platform for analysing large data sets on Hadoop, using a dataflow language called Pig Latin.</mark>

Key points.

  1. Pig was developed at Yahoo and later became an Apache project.
  2. Pig Latin scripts are compiled automatically into MapReduce jobs, so no Java is needed.
  3. A few lines of Pig Latin replace many lines of MapReduce code.
  4. Pig handles structured, semi-structured and unstructured data and does not force a schema.

Anatomy of Pig

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>The anatomy of Pig is its components: the Pig Latin language, the Grunt shell, and the parser, optimizer, compiler and execution engine.</mark>

Key points.

  1. The Grunt shell is the interactive command prompt for entering Pig Latin statements.
  2. The Parser checks syntax and types and produces a logical plan (a DAG).
  3. The Optimizer applies logical optimizations such as pushing filters early.
  4. The Compiler turns the plan into MapReduce jobs, which the execution engine runs on Hadoop.
  5. The data model has atom, tuple, bag and map, held in relations.

Pig on Hadoop

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Pig on Hadoop means Pig Latin scripts run over data in HDFS, with Pig generating MapReduce jobs that Hadoop executes.</mark>

Key points.

  1. Pig reads its input from HDFS with LOAD and writes results back with STORE.
  2. Pig is a client-side tool; nothing extra is installed on the cluster.
  3. Jobs run on the Hadoop cluster under YARN, so Pig gets parallelism and fault tolerance.
  4. Pig can also run in local mode on one machine.

Use Case for Pig

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A use case for Pig is any large-scale data-processing task, such as log analysis, that is easier as a dataflow than as hand-written MapReduce.</mark>

Key points.

  1. Web-log and clickstream analysis, such as counting hits per page.
  2. Data preparation and cleaning of raw data before analysis.
  3. Iterative processing and ad-hoc research on large data by analysts.
  4. Yahoo and Twitter use Pig for processing huge logs.

ETL Processing

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>ETL is Extract, Transform, Load: data is pulled from sources, cleaned and reshaped, then loaded into a warehouse, and Pig is widely used for the transform step.</mark>

Key points.

  1. Extract: LOAD reads raw data from HDFS or other sources.
  2. Transform: FILTER, FOREACH, JOIN and GROUP clean, combine and summarise the data.
  3. Load: STORE writes the result to HDFS or a Hive table.
  4. Pig suits ETL because it handles messy data without a fixed schema.

Data types in Pig

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Simple (scalar) types in Pig hold a single value.</mark>

Key points.

  1. Numeric types are int, long, float and double.
  2. chararray is a Unicode string and bytearray is a blob of bytes (the default when no type is given).
  3. boolean holds true or false, and datetime holds a date and time.
  4. A schema is optional: A = LOAD 'f' AS (id:int, name:chararray);.

running Pig

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Pig runs in local mode or MapReduce mode, and Pig Latin can be given through the Grunt shell, a script file or an embedded program.</mark>

Key points.

  1. Local mode (pig -x local) uses one JVM and the local file system, and is used for testing.
  2. MapReduce mode (pig -x mapreduce, the default) runs on a Hadoop cluster and reads HDFS.
  3. Grunt shell mode is interactive, one statement at a time.
  4. Script mode runs a .pig file, for example pig script.pig; embedded mode calls Pig from Java or Python.

Execution model of Pig

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Pig uses lazy execution: statements build a logical plan, which becomes a physical plan and then MapReduce jobs only when DUMP or STORE is called.</mark>

Key points.

  1. The parser checks each statement and builds the logical plan.
  2. The logical plan is optimized and converted into a physical plan.
  3. The physical plan is compiled into a sequence of MapReduce jobs.
  4. Nothing runs until DUMP (display) or STORE (save) is reached.

Operators

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Pig operators are the Pig Latin statements that load, transform and store relations.</mark>

Key points.

  1. LOAD and STORE read and write data; DUMP displays it.
  2. FILTER selects rows by a condition; FOREACH ... GENERATE picks or computes columns.
  3. GROUP and COGROUP collect records by key; JOIN combines relations on a key.
  4. ORDER sorts, DISTINCT removes duplicates, LIMIT keeps the first n rows, UNION merges relations, SPLIT divides one.

functions

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Pig functions are built-in or user-written routines applied to data inside Pig Latin statements.</mark>

Key points.

  1. Eval functions include AVG, COUNT, SUM, MAX, MIN, CONCAT and SIZE.
  2. Load/store functions such as PigStorage and TextLoader define how data is read and written.
  3. Filter functions like IsEmpty return a boolean; math and string functions include ABS, ROUND and UPPER.
  4. User Defined Functions (UDFs) can be written in Java, Python or JavaScript and added with REGISTER.

Data types of Pig

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Complex types in Pig hold collections: tuple, bag and map.</mark>

Key points.

  1. A tuple is an ordered set of fields, written (1,'Ravi'); it is like a row.
  2. A bag is a collection of tuples, written {(1,'Ravi'),(2,'Asha')}; it is like a table and may contain duplicates.
  3. A map is a set of key-value pairs, written [name#Ravi], with a chararray key.
  4. A relation is an outer bag of tuples.

Last-minute revision

  • Hive is a Hadoop data warehouse with SQL-like HiveQL, created at Facebook.
  • Hive stores schema in the Metastore and data in HDFS; it uses schema-on-read.
  • Hive complex types: ARRAY, MAP, STRUCT, UNIONTYPE.
  • Hive suits batch analytics, not OLTP.
  • Pig is a dataflow platform with Pig Latin, created at Yahoo, compiled to MapReduce.
  • Pig data model: atom, tuple, bag, map; relation is a bag of tuples.
  • Pig modes: local (-x local) and MapReduce (default).
  • Pig execution is lazy: work starts only at DUMP or STORE.
  • ETL = Extract (LOAD), Transform (FILTER, FOREACH, JOIN), Load (STORE).
  • Common Pig functions: AVG, COUNT, SUM, MAX, MIN; UDFs are added with REGISTER.

Memory hooks

  • Hive is SQL for Hadoop; Pig is a script for Hadoop.
  • Hive architecture: CDEM, meaning Compiler, Driver, Executor, Metastore.
  • Pig data model, smallest to largest: field, tuple, bag, relation.
  • Pig is lazy: it waits for DUMP or STORE.
  • ETL follows LOAD, transform, STORE.

Coverage checklist

  • Introduction to Hive: definition and features (no past questions).
  • Hive Architecture: components and roles (no past questions).
  • Hive Data types: primitive and complex (no past questions).
  • Hive Query Language: DDL, load and query (no past questions).
  • Introduction to Pig: definition and features (no past questions).
  • Anatomy of Pig: components (no past questions).
  • Pig on Hadoop: HDFS and MapReduce (no past questions).
  • Use Case for Pig: log analysis and data preparation (no past questions).
  • ETL Processing: extract, transform, load (no past questions).
  • Data types in Pig: scalar types (no past questions).
  • running Pig: modes (no past questions).
  • Execution model of Pig: plans and lazy execution (no past questions).
  • Operators: relational operators (no past questions).
  • functions: eval, load/store and UDFs (no past questions).
  • Data types of Pig: tuple, bag, map (no past questions).
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in