How unit 3 is examined
This unit covers Hive (architecture, data types, HiveQL) and Pig (Pig Latin, execution, operators, functions); no topic was asked in the supplied papers, so learn the definitions and the comparisons.
Introduction to Hive
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Hive is a data-warehouse tool built on Hadoop that lets users query and analyse large data stored in HDFS using an SQL-like language called HiveQL.</mark>
Key points.
- Hive was developed at Facebook and is now an Apache project.
- HiveQL queries are converted internally into MapReduce, Tez or Spark jobs.
- Hive uses schema-on-read: table structure is applied when data is read, not when it is loaded.
- It suits batch analysis of huge data, not OLTP, because latency is high and row-level updates are limited.
Hive Architecture
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Hive architecture consists of the user interfaces, the Driver (compiler, optimizer, executor), the Metastore and the Hadoop layer (HDFS and MapReduce).</mark>
Key points.
- User interfaces are the CLI, Beeline, web UI and JDBC/ODBC clients connecting through HiveServer2 (Thrift).
- The Driver receives the query, creates a session and manages its life cycle.
- The Compiler parses HiveQL, checks it against the metastore and builds an execution plan; the Optimizer improves that plan.
- The Executor submits the plan as jobs to Hadoop and monitors them.
- The Metastore stores table schemas, partitions and locations in an RDBMS such as MySQL or Derby.
- Actual data stays in HDFS.
Hive Data types
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Hive data types are of two kinds: primitive types for single values and complex types that hold collections.</mark>
Key points.
- Numeric types are TINYINT, SMALLINT, INT, BIGINT, FLOAT, DOUBLE and DECIMAL.
- String types are STRING, VARCHAR and CHAR; BOOLEAN holds true or false.
- Date and time types are DATE and TIMESTAMP; BINARY holds raw bytes.
- Complex types are ARRAY (ordered same-type items), MAP (key-value pairs), STRUCT (named fields) and UNIONTYPE.
Hive Query Language
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>HiveQL is Hive's SQL-like language for defining, loading and querying tables stored on Hadoop.</mark>
Key points.
- DDL commands are CREATE, ALTER, DROP and SHOW, for example
CREATE TABLE emp (id INT, name STRING). - Data is loaded with
LOAD DATA INPATH ... INTO TABLEorINSERT. - Queries use SELECT with WHERE, GROUP BY, ORDER BY, HAVING and JOIN.
- Partitions and buckets split a table so queries scan less data.
- Tables may be managed (Hive owns the data) or external (dropping keeps the data).
Introduction to Pig
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Pig is a high-level platform for analysing large data sets on Hadoop, using a dataflow language called Pig Latin.</mark>
Key points.
- Pig was developed at Yahoo and later became an Apache project.
- Pig Latin scripts are compiled automatically into MapReduce jobs, so no Java is needed.
- A few lines of Pig Latin replace many lines of MapReduce code.
- Pig handles structured, semi-structured and unstructured data and does not force a schema.
Anatomy of Pig
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>The anatomy of Pig is its components: the Pig Latin language, the Grunt shell, and the parser, optimizer, compiler and execution engine.</mark>
Key points.
- The Grunt shell is the interactive command prompt for entering Pig Latin statements.
- The Parser checks syntax and types and produces a logical plan (a DAG).
- The Optimizer applies logical optimizations such as pushing filters early.
- The Compiler turns the plan into MapReduce jobs, which the execution engine runs on Hadoop.
- The data model has atom, tuple, bag and map, held in relations.
Pig on Hadoop
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Pig on Hadoop means Pig Latin scripts run over data in HDFS, with Pig generating MapReduce jobs that Hadoop executes.</mark>
Key points.
- Pig reads its input from HDFS with LOAD and writes results back with STORE.
- Pig is a client-side tool; nothing extra is installed on the cluster.
- Jobs run on the Hadoop cluster under YARN, so Pig gets parallelism and fault tolerance.
- Pig can also run in local mode on one machine.
Use Case for Pig
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A use case for Pig is any large-scale data-processing task, such as log analysis, that is easier as a dataflow than as hand-written MapReduce.</mark>
Key points.
- Web-log and clickstream analysis, such as counting hits per page.
- Data preparation and cleaning of raw data before analysis.
- Iterative processing and ad-hoc research on large data by analysts.
- Yahoo and Twitter use Pig for processing huge logs.
ETL Processing
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>ETL is Extract, Transform, Load: data is pulled from sources, cleaned and reshaped, then loaded into a warehouse, and Pig is widely used for the transform step.</mark>
Key points.
- Extract: LOAD reads raw data from HDFS or other sources.
- Transform: FILTER, FOREACH, JOIN and GROUP clean, combine and summarise the data.
- Load: STORE writes the result to HDFS or a Hive table.
- Pig suits ETL because it handles messy data without a fixed schema.
Data types in Pig
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Simple (scalar) types in Pig hold a single value.</mark>
Key points.
- Numeric types are int, long, float and double.
- chararray is a Unicode string and bytearray is a blob of bytes (the default when no type is given).
- boolean holds true or false, and datetime holds a date and time.
- A schema is optional:
A = LOAD 'f' AS (id:int, name:chararray);.
running Pig
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Pig runs in local mode or MapReduce mode, and Pig Latin can be given through the Grunt shell, a script file or an embedded program.</mark>
Key points.
- Local mode (
pig -x local) uses one JVM and the local file system, and is used for testing. - MapReduce mode (
pig -x mapreduce, the default) runs on a Hadoop cluster and reads HDFS. - Grunt shell mode is interactive, one statement at a time.
- Script mode runs a
.pigfile, for examplepig script.pig; embedded mode calls Pig from Java or Python.
Execution model of Pig
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Pig uses lazy execution: statements build a logical plan, which becomes a physical plan and then MapReduce jobs only when DUMP or STORE is called.</mark>
Key points.
- The parser checks each statement and builds the logical plan.
- The logical plan is optimized and converted into a physical plan.
- The physical plan is compiled into a sequence of MapReduce jobs.
- Nothing runs until DUMP (display) or STORE (save) is reached.
Operators
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Pig operators are the Pig Latin statements that load, transform and store relations.</mark>
Key points.
- LOAD and STORE read and write data; DUMP displays it.
- FILTER selects rows by a condition; FOREACH ... GENERATE picks or computes columns.
- GROUP and COGROUP collect records by key; JOIN combines relations on a key.
- ORDER sorts, DISTINCT removes duplicates, LIMIT keeps the first n rows, UNION merges relations, SPLIT divides one.
functions
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Pig functions are built-in or user-written routines applied to data inside Pig Latin statements.</mark>
Key points.
- Eval functions include AVG, COUNT, SUM, MAX, MIN, CONCAT and SIZE.
- Load/store functions such as PigStorage and TextLoader define how data is read and written.
- Filter functions like IsEmpty return a boolean; math and string functions include ABS, ROUND and UPPER.
- User Defined Functions (UDFs) can be written in Java, Python or JavaScript and added with REGISTER.
Data types of Pig
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Complex types in Pig hold collections: tuple, bag and map.</mark>
Key points.
- A tuple is an ordered set of fields, written
(1,'Ravi'); it is like a row. - A bag is a collection of tuples, written
{(1,'Ravi'),(2,'Asha')}; it is like a table and may contain duplicates. - A map is a set of key-value pairs, written
[name#Ravi], with a chararray key. - A relation is an outer bag of tuples.
Last-minute revision
- Hive is a Hadoop data warehouse with SQL-like HiveQL, created at Facebook.
- Hive stores schema in the Metastore and data in HDFS; it uses schema-on-read.
- Hive complex types: ARRAY, MAP, STRUCT, UNIONTYPE.
- Hive suits batch analytics, not OLTP.
- Pig is a dataflow platform with Pig Latin, created at Yahoo, compiled to MapReduce.
- Pig data model: atom, tuple, bag, map; relation is a bag of tuples.
- Pig modes: local (
-x local) and MapReduce (default). - Pig execution is lazy: work starts only at DUMP or STORE.
- ETL = Extract (LOAD), Transform (FILTER, FOREACH, JOIN), Load (STORE).
- Common Pig functions: AVG, COUNT, SUM, MAX, MIN; UDFs are added with REGISTER.
Memory hooks
- Hive is SQL for Hadoop; Pig is a script for Hadoop.
- Hive architecture: CDEM, meaning Compiler, Driver, Executor, Metastore.
- Pig data model, smallest to largest: field, tuple, bag, relation.
- Pig is lazy: it waits for DUMP or STORE.
- ETL follows LOAD, transform, STORE.
Coverage checklist
- Introduction to Hive: definition and features (no past questions).
- Hive Architecture: components and roles (no past questions).
- Hive Data types: primitive and complex (no past questions).
- Hive Query Language: DDL, load and query (no past questions).
- Introduction to Pig: definition and features (no past questions).
- Anatomy of Pig: components (no past questions).
- Pig on Hadoop: HDFS and MapReduce (no past questions).
- Use Case for Pig: log analysis and data preparation (no past questions).
- ETL Processing: extract, transform, load (no past questions).
- Data types in Pig: scalar types (no past questions).
- running Pig: modes (no past questions).
- Execution model of Pig: plans and lazy execution (no past questions).
- Operators: relational operators (no past questions).
- functions: eval, load/store and UDFs (no past questions).
- Data types of Pig: tuple, bag, map (no past questions).