High Performance computing (AL-802 (B)) - Important Questions
-
Unit 110 Marks High Priority
Explain the memory hierarchy in modern CPU architectures and derive the formula for Average Memory Access Time (AMAT) for a three-level hierarchy (L1, L2, main memory). Given hit rates and latencies for L1, L2 and main memory, compute AMAT for a numerical example.
Core derivation from Unit 1; frequently asked conceptual and numerical question on memory hierarchies and AMAT.
-
Unit 310 Marks High Priority
Describe the OpenMP parallel for construct and explain data scoping clauses: shared, private and reduction. Discuss common pitfalls such as race conditions and false sharing, and illustrate with a short code example showing correct usage.
Core practical question from Unit 3; examines knowledge of OpenMP directives, data scoping and common pitfalls.
-
Unit 57 Marks High Priority
Compare MPI point-to-point and collective communications. When should non-blocking communication be preferred over blocking communication? Give example scenarios and briefly outline the use of MPI_Send / MPI_Isend and MPI_Bcast.
Standard comparative question from Unit 5; tests understanding of MPI primitives and their use-cases.
-
Unit 27 Marks High Priority
Discuss data access optimizations such as loop blocking (tiling), loop interchange and software prefetching. Explain how these transformations improve cache reuse and reduce memory traffic, and provide a simple example showing tiled matrix multiplication.
Unit 2 practical optimization topic; asks for transformations improving cache performance and locality.
-
Unit 314 Marks High Priority
Apply Amdahl's law to derive the theoretical speedup for a parallel program and discuss the limits to scaling. For a program with parallel fraction $p$ running on $s$ processors, derive the speedup formula and explain implications as $s$ increases. Provide a brief numerical example.
Core scalability law from Unit 3; common high-value question asking derivation and limits (Amdahl's law).
-
Unit 110 Marks High Priority
Define operational intensity and perform a roofline / balance model calculation: given total FLOP count and memory traffic for a kernel, estimate whether the kernel is compute-bound or memory-bound. Show the calculation steps and interpretation.
Important performance-analysis question from Unit 1; roofline / balance model is core for identifying bottlenecks.
-
Unit 17 Marks High Priority
Explain vector processor design principles and how vectorization improves performance. Discuss concepts such as strip-mining, vector length, and data alignment, and mention common challenges when porting scalar code to vector processors.
Vector architecture fundamentals from Unit 1; tests understanding of vectorization benefits and implementation considerations.
-
Unit 47 Marks High Priority
What is false sharing in shared-memory parallel programs? Explain how false sharing arises, how it affects performance, and list practical techniques to detect and mitigate it (for example, padding, alignment, and data privatization).
Shared-memory concurrency issue from Unit 4; false sharing is a frequent short-answer topic with practical mitigation techniques.
Quick Add to Notes
Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.
Create free accountHave an account? Log in
Notes Panel