Skip to content
CS-702 (D) · Big Data/Quick Revision Short Notes

Big Data (CS-702 (D)) - Unit 3 Short Notes

How unit 3 is examined

Hive (architecture, data types, HiveQL joins, ORDER BY and UDFs) and Pig (anatomy, data types, operators, word count) plus ETL processing; HiveQL and ETL carry the most marks, then Hive architecture, Hive and Pig data types and Pig operators.

Introduction to Hive Hive Architecture

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Hive is a data-warehouse system built on top of Hadoop that lets users query data stored in HDFS with an SQL-like language called HiveQL, which is compiled into MapReduce jobs.</mark>

Diagram.

<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-01" viewBox="0 0 603 209" width="603" height="209" role="img" aria-label="Hive working. UI = CLI/Web UI/Thrift server; Drv = Driver; Cmp = Compiler and optimizer; MS = Metastore (RDBMS); Exe = Execution engine (MapReduce) running on Hadoop; HDFS = data storage"><style>#dsfig-u3-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-01 .t{fill:#16181D;font-weight:500}#dsfig-u3-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-01 .dot{fill:#16181D}#dsfig-u3-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-01 .ah{fill:#454C5A}#dsfig-u3-01 .ah.hi{fill:#2340B8}#dsfig-u3-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-01 .e{stroke:#B1B7C3}html.dark #dsfig-u3-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-01 .t{fill:#E6E8ED}html.dark #dsfig-u3-01 .t.inv{fill:#0F1115}html.dark #dsfig-u3-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-01 .dot{fill:#E6E8ED}html.dark #dsfig-u3-01 .ann{fill:#8FA3FF}html.dark #dsfig-u3-01 .lbl{fill:#858D9C}html.dark #dsfig-u3-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-01 .ah{fill:#B1B7C3}html.dark #dsfig-u3-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah6" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh6" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L148,40" marker-end="url(#ah6)"/><path class="e" d="M188,40 L277,40" marker-end="url(#ah6)"/><path class="e" d="M298,61 L298,148" marker-end="url(#ah6)" marker-start="url(#ah6)"/><path class="e" d="M317,40 L406,40" marker-end="url(#ah6)"/><path class="e" d="M446,40 L528,40" marker-end="url(#ah6)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">UI</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">Drv</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">Cmp</text><circle class="n" cx="298" cy="169" r="18"/><text class="t" x="298" y="169" dy=".35em" text-anchor="middle">MS</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">Exe</text><rect class="n" x="531" y="25" width="50" height="30" rx="15"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">HDFS</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Hive working. UI = CLI/Web UI/Thrift server; Drv = Driver; Cmp = Compiler and optimizer; MS = Metastore (RDBMS); Exe = Execution engine (MapReduce) running on Hadoop; HDFS = data storage</figcaption></figure>

Key points.

  1. The user interface (CLI, web UI or JDBC/ODBC through the Thrift server) accepts a HiveQL query and hands it to the driver.
  2. The driver creates a session, manages the query life cycle and passes the query to the compiler.
  3. The compiler parses the query, checks it against the metastore, and produces a logical plan and then a DAG of MapReduce stages.
  4. The metastore is a central repository, usually in an RDBMS such as MySQL or Derby, that stores table definitions, column types, partitions and HDFS locations.
  5. The metastore runs in three modes: embedded (Derby in the same JVM, one user), local (separate database, same JVM) and remote (a separate metastore service, used in production).
  6. The execution engine runs the stages as MapReduce jobs on Hadoop and returns the results to the driver and user.
  7. Hive reads and writes HDFS files through a SerDe (serializer/deserializer) with an InputFormat (for example TextInputFormat) to read and an OutputFormat to write.
  8. Hive suits batch analysis of large data with no row-level updates, and it has high latency; data model: databases, tables, partitions and buckets.

Answer frame. Open with the definition; draw the diagram; then explain steps 1-6 as the query flow (UI, driver, compiler, metastore, execution engine, HDFS); add the metastore modes and SerDe/InputFormat/OutputFormat for the metastore question; close with advantages (SQL-like, scalable) and limits (high latency).

Asked: [7 marks] (Nov 2022, Dec 2024) Explain in detail about HIVE / Explain working of Hive with proper steps and diagram. Asked: [7 marks] (Nov 2023) What is Hive meta store? Which classes are used by Hive to read and write HDFS files? Explain.

Hive Data types

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Hive data types are the types given to table columns, split into primitive types (single values) and complex types (collections built from other types).</mark>

Key points.

  1. Numeric primitives are TINYINT (1 byte), SMALLINT (2), INT (4), BIGINT (8), FLOAT and DOUBLE, and DECIMAL for exact values.
  2. Other primitives are STRING, VARCHAR, CHAR, BOOLEAN, TIMESTAMP, DATE and BINARY.
  3. ARRAY holds an ordered list of same-type items, for example ARRAY<STRING>, accessed as skills[0].
  4. MAP holds key-value pairs, for example MAP<STRING,INT>, accessed as marks['maths'].
  5. STRUCT groups named fields of different types, for example STRUCT<city:STRING,pin:INT>, accessed as addr.city.
  6. UNIONTYPE holds a value of one of several listed types, for example UNIONTYPE<INT,STRING>.
CREATE TABLE emp (id INT, name STRING, skills ARRAY<STRING>,
  marks MAP<STRING,INT>, addr STRUCT<city:STRING,pin:INT>);
Point Hive Pig
Simple types TINYINT, INT, BIGINT, FLOAT, DOUBLE, STRING, BOOLEAN int, long, float, double, chararray, bytearray
Complex types ARRAY, MAP, STRUCT, UNIONTYPE Tuple, Bag, Map
Schema Fixed in the metastore, checked at table creation Optional, declared in LOAD ... AS, defaults to bytearray
Language Declarative HiveQL Procedural dataflow Pig Latin
Users Analysts who know SQL Programmers and ETL developers
Declaration age INT age:int

Answer frame. Open with the definition; list primitives then complex types with one example each; show the CREATE TABLE; for the compare question add the table and one declaration each; close with one line on schema handling.

Asked: [7 marks] (Dec 2020) Describe and compare data types of Hive and Pig. Asked: [7 marks] (Dec 2024) What are the different Hive data types? Explain them briefly.

Hive Query Language

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>HiveQL is Hive's SQL-like query language for creating tables and for selecting, joining and aggregating data, which Hive translates into MapReduce jobs.</mark>

Key points.

  1. SELECT syntax is SELECT cols FROM table [WHERE cond] [GROUP BY cols] [ORDER BY cols] [LIMIT n].
  2. ORDER BY gives a total, global ordering of the whole result, so it sends all data through one reducer and is slow on big data.
  3. SORT BY orders only within each reducer, so output is sorted per reducer but not globally; DISTRIBUTE BY controls which reducer gets a row.
  4. NATURAL JOIN joins on all same-named columns (supported from Hive 2.2); on older versions write an equi-join with ON.
  5. LEFT OUTER keeps all rows of the left table, RIGHT OUTER all of the right, and FULL OUTER all of both, with NULL where there is no match.
  6. A UDF is a Java function added to Hive; the kinds are UDF (one row in, one value out), UDAF (many rows in, one value out, like SUM) and UDTF (one row in, many rows out, like explode).

Example. emp(id, name, dept_id): (1,Asha,10), (2,Ravi,20), (3,Meena,30); dept(dept_id, dept_name): (10,Sales), (20,HR), (40,IT).

SELECT * FROM emp NATURAL JOIN dept;                    -- (10,1,Asha,Sales) (20,2,Ravi,HR)
SELECT e.name, d.dept_name FROM emp e LEFT OUTER JOIN dept d
  ON e.dept_id = d.dept_id;                             -- adds (Meena, NULL)
SELECT e.name, d.dept_name FROM emp e RIGHT OUTER JOIN dept d
  ON e.dept_id = d.dept_id;                             -- adds (NULL, IT)
SELECT e.name, d.dept_name FROM emp e FULL OUTER JOIN dept d
  ON e.dept_id = d.dept_id;                             -- adds both
SELECT name, marks FROM student ORDER BY marks DESC;    -- highest first

Steps (UDF).

Step 1: Write a Java class extending org.apache.hadoop.hive.ql.exec.UDF with a public evaluate() method.
Step 2: Compile it and package it as a jar.
Step 3: In Hive run ADD JAR /path/up.jar;
Step 4: CREATE TEMPORARY FUNCTION up AS 'com.x.Upper';
Step 5: Use it: SELECT up(name) FROM emp;

Answer frame. Joins: define, show both tables, give the five queries and the output of each. UDF: define, give the three types, then the five steps and one usage. ORDER BY: give the SELECT syntax, then ORDER BY versus SORT BY, then the example.

Asked: [7 marks] (Nov 2022, Nov 2023) Write with suitable example the HIVE queries for Natural join and Outer join. Asked: [7 marks] (Dec 2024) Explain procedure to write user defined functions in Hive. Asked: [7 marks] (Jun 2025) Explain the HiveQL Select-Order By with suitable example.

Pitfall: Writing ORDER BY when the question compares it with SORT BY loses the point that ORDER BY is global (one reducer) and SORT BY is per reducer.

Introduction to Pig

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Apache Pig is a high-level platform for analysing large data sets on Hadoop, using the dataflow language Pig Latin.</mark>

Key points.

  1. Pig Latin scripts are compiled into MapReduce jobs, so the user writes no Java.
  2. A Pig script is far shorter than the equivalent MapReduce code.
  3. Pig handles structured, semi-structured and unstructured data.
  4. It was developed at Yahoo and is now an Apache project.

Anatomy of Pig

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Pig is a scripting platform for Hadoop, and its anatomy is the Pig Latin language plus the Grunt shell and the engine (parser, optimizer, compiler, execution engine) that turns scripts into MapReduce jobs.</mark>

Key points.

  1. The parser checks syntax and types and builds a logical plan as a DAG of operators.
  2. The optimizer applies logical optimizations such as pushing filters early and dropping unused columns.
  3. The compiler turns the plan into a chain of MapReduce jobs, and the execution engine submits them to Hadoop and returns the results.
  4. Execution modes are local mode (one JVM, local files) and MapReduce mode (a Hadoop cluster and HDFS).
  5. A script can be run in the Grunt shell, as a script file or embedded in Java.

Asked: [7 marks] (Dec 2020) Give an introduction to Pig. Explain anatomy of pig.

Pig on Hadoop

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Pig on Hadoop means running Pig Latin scripts in MapReduce mode, where each script is converted to MapReduce jobs that read and write HDFS.</mark>

Key points.

  1. Start it with pig -x mapreduce (the default), while pig -x local runs on one machine.
  2. Pig reads its input from HDFS with LOAD and writes results back to HDFS with STORE.
  3. Pig needs the HADOOP_HOME and cluster configuration, and it is installed on the client machine.
  4. The Hadoop cluster provides the distributed storage and parallel processing.

Use Case for Pig

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Pig is used where large raw data must be cleaned, transformed and analysed quickly without writing MapReduce code.</mark>

Key points.

  1. ETL data pipelines use Pig to extract, clean and load data into a warehouse.
  2. Ad hoc analysis of raw data, such as web server logs, is quick to write in Pig.
  3. Research on large data sets, such as building user-behaviour or recommendation data, uses Pig.
  4. Yahoo and Twitter run Pig for log processing.

ETL Processing

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>ETL (Extract, Transform, Load) is the process of extracting data from source systems, transforming it into a clean and consistent form, and loading it into a target such as a data warehouse or Hadoop.</mark>

Key points (i, ETL).

  1. Extract reads data from sources such as databases, files, logs and APIs.
  2. Transform cleans, filters, joins, aggregates and converts formats, for example converting dates or removing duplicates.
  3. Load writes the result to the warehouse or HDFS; Hive and Pig are the usual ETL tools on Hadoop.

Example. Sales files from shops are extracted, prices converted to rupees and duplicate bills removed, then loaded into a Hive table.

Key points (ii, properties of Big Data systems).

  1. Volume is the huge size of the data, measured in terabytes and petabytes.
  2. Velocity is the speed at which data arrives and must be processed.
  3. Variety means structured, semi-structured and unstructured formats.
  4. Veracity is the accuracy and trustworthiness of the data.
  5. Value is the useful insight the data gives; a good system is also scalable, fault tolerant and robust.

Key points (iii, data architectural patterns).

  1. Lambda architecture has a batch layer for accurate results, a speed layer for real-time results and a serving layer that merges both.
  2. Kappa architecture uses only a stream-processing layer and replays the log to reprocess data, so it has one code path.
  3. A data lake stores raw data of every type in its original form, and structure is applied when it is read.
  4. Data warehouse and master-slave/shared-nothing patterns are other options; each pattern is chosen for scale, latency and cost.

Answer frame. This is a 14-mark short note, so give about 4 marks to each part: for ETL a one-line definition with the three steps and an example; for properties the five Vs each in a sentence; for patterns Lambda, Kappa and data lake with one line each; close with one summary line.

Asked: [14 marks] (Jun 2025) Explain the following terms: i) ETL processing ii) Properties of Big data systems iii) Data architectural patterns.

Data types in Pig running Pig

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Pig data types are the scalar and complex types of Pig Latin fields, and running Pig means executing scripts in local or MapReduce mode.</mark>

Key points.

  1. Scalar types are int, long, float, double, chararray and bytearray; complex types are tuple, bag and map.
  2. Pig runs interactively in the Grunt shell, as a script (pig script.pig) or embedded in Java.
  3. The modes are pig -x local and pig -x mapreduce.
  4. DUMP shows the result on screen and STORE saves it.

Execution model of Pig

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>The Pig execution model converts a Pig Latin script into a logical plan, then a physical plan, then a MapReduce plan that runs on Hadoop.</mark>

Key points.

  1. Pig is lazy: nothing runs until DUMP, STORE or another output statement is reached.
  2. The logical plan is an optimized DAG of operators; the physical plan chooses how each runs.
  3. The MapReduce plan groups operators into map and reduce stages, and Hadoop executes the jobs.
  4. Execution runs in local mode or MapReduce mode.

Operators

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Relational operators are the Pig Latin operators that load, filter, group, join and sort relations (bags of tuples).</mark>

Key points.

  1. LOAD reads data into a relation: A = LOAD 'in.txt' AS (name:chararray, age:int);.
  2. FILTER keeps rows that satisfy a condition: B = FILTER A BY age > 18;.
  3. FOREACH ... GENERATE applies an expression to every row, choosing or computing columns: C = FOREACH A GENERATE name;.
  4. GROUP collects rows with the same key into a bag: D = GROUP A BY age;.
  5. JOIN combines two relations on a key: E = JOIN A BY id, B BY id;.
  6. ORDER sorts a relation, DISTINCT removes duplicate rows, LIMIT keeps the first n rows, and UNION, SPLIT and CROSS combine or split relations.
  7. DUMP prints and STORE saves the result.

Example (word count).

lines = LOAD 'input.txt' AS (line:chararray);
words = FOREACH lines GENERATE FLATTEN(TOKENIZE(line)) AS word;
grp = GROUP words BY word;
cnt = FOREACH grp GENERATE group, COUNT(words);
STORE cnt INTO 'out';   -- or DUMP cnt; gives (hello,2) (big,1)

Answer frame. Operators question: define, then list each operator with syntax and one line, then a small example. Word count: load, tokenize and flatten, group, count, store, with one line per statement.

Asked: [7 marks] (Nov 2023) List and explain the relational operators in Pig. Asked: [7 marks] (Dec 2024) Write a word count program in Pig to count the occurrence of similar words in a file.

functions

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Pig functions are built-in (eval, load/store, math and string functions) or user-defined functions (UDFs) that extend Pig Latin.</mark>

Key points.

  1. Eval functions include COUNT, SUM, AVG, MIN, MAX and TOKENIZE; string functions include CONCAT and SUBSTRING.
  2. Load/store functions such as PigStorage and TextLoader read and write data.
  3. UDFs are written in Java (or Python), packaged in a jar, and loaded with REGISTER.
  4. Function names are case sensitive.

Data types of Pig

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Pig has a data model of scalar types (int, long, float, double, chararray, bytearray) and complex types (tuple, bag, map).</mark>

Key points.

  1. A field is a single value, such as 10 or 'Ravi'; chararray is a string and bytearray is raw bytes (the default type).
  2. A tuple is an ordered set of fields, written (1,Asha).
  3. A bag is a collection of tuples, written {(1,Asha),(2,Ravi)}, and a relation is an outer bag.
  4. A map is a set of key-value pairs with chararray keys, written [name#Asha,age#20].
  5. Declaration: A = LOAD 'f' AS (id:int, name:chararray, t:tuple(x:int,y:int), b:bag{t:tuple(z:int)}, m:map[]);.

Asked: [7 marks] (Nov 2022) Discuss the various data types in Pig.

Last-minute revision

  • Hive is a data warehouse on Hadoop whose HiveQL is compiled to MapReduce.
  • Hive architecture: UI, driver, compiler, metastore, execution engine, HDFS.
  • Metastore modes: embedded, local, remote; HDFS I/O uses SerDe, InputFormat and OutputFormat.
  • Hive complex types: ARRAY, MAP, STRUCT, UNIONTYPE; Pig complex types: tuple, bag, map.
  • ORDER BY is global (one reducer); SORT BY is per reducer.
  • Outer joins: LEFT, RIGHT, FULL keep unmatched rows with NULL.
  • UDF steps: extend UDF, write evaluate(), jar, ADD JAR, CREATE TEMPORARY FUNCTION.
  • Pig anatomy: parser, optimizer, compiler, execution engine.
  • Pig word count: LOAD, TOKENIZE and FLATTEN, GROUP, COUNT, STORE.
  • ETL is Extract, Transform, Load; Big Data properties are volume, velocity, variety, veracity, value.
  • Patterns: Lambda (batch and speed layers), Kappa (stream only), data lake (raw storage).

Memory hooks

  • "UDCM-EH": UI, Driver, Compiler, Metastore, Engine, HDFS.
  • ORDER BY = one reducer = one global order; SORT BY = sorted pieces.
  • Pig anatomy "POCE": Parser, Optimizer, Compiler, Engine.
  • Pig data "TBM": Tuple row, Bag of tuples, Map key-value.
  • Word count "LTGCS": Load, Tokenize, Group, Count, Store.

Coverage checklist

  • Introduction to Hive Hive Architecture: Q9 (Nov 2022, Dec 2024), Q10 (Nov 2023).
  • Hive Data types: Q4 (Dec 2020), Q5 (Dec 2024).
  • Hive Query Language: Q6 (Nov 2022, Nov 2023), Q7 (Dec 2024), Q8 (Jun 2025).
  • Introduction to Pig: no past question.
  • Anatomy of Pig: Q2 (Dec 2020).
  • Pig on Hadoop: no past question.
  • Use Case for Pig: no past question.
  • ETL Processing: Q1 (Jun 2025).
  • Data types in Pig running Pig: no past question.
  • Execution model of Pig: no past question.
  • Operators: Q11 (Nov 2023), Q12 (Dec 2024).
  • functions: no past question.
  • Data types of Pig: Q3 (Nov 2022).
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in