Skip to content
ME-506 · Python/Quick Revision Short Notes

Python (ME-506) - Unit 5 Short Notes

UNIT 5: ADVANCED PYTHON CONCEPTS & APPLICATIONS


I. ADVANCED LANGUAGE FEATURES & PROGRAMMING PARADIGMS

A. Decorators
  • Function Decorators: A function that takes another function as argument and extends its behavior without modifying it. Syntax uses @decorator_name above the function definition.

    • Use Cases: Logging, timing execution, access control, caching.

    • Basic Pattern:

      
      def decorator(func):
      
          def wrapper(*args, **kwargs):
      
              # pre-processing
      
              result = func(*args, **kwargs)
      
              # post-processing
      
              return result
      
          return wrapper
      
      
  • Class Decorators: A function (or callable) that takes a class as argument and returns a modified class (e.g., adding methods/attributes to all instances).

  • Decorators with Arguments: Requires an extra layer of nesting. The outermost function accepts the decorator's arguments and returns the actual decorator.

    
    def repeat(n):  # Outer: accepts decorator args
    
        def decorator(func):  # Middle: accepts function
    
            def wrapper(*args, **kwargs):
    
                for _ in range(n):
    
                    result = func(*args, **kwargs)
    
                return result
    
            return wrapper
    
        return decorator  # Returns the decorator
    
    
  • functools.wraps: A decorator used inside a custom decorator's wrapper to copy metadata (__name__, __doc__, etc.) from the original function to the wrapper. Crucial for debugging and introspection.

    [!TIP] Always use @functools.wraps(func) on the inner wrapper function in your decorators.

B. Generators & Coroutines
  • Generator Functions: Defined with yield. Produces a sequence of values lazily, one at a time, pausing state between yields. Memory efficient for large/streaming data.

    
    def count_up_to(n):
    
        i = 0
    
        while i < n:
    
            yield i
    
            i += 1
    
    
  • Generator Expressions: Similar syntax to list comprehensions but with parentheses (). Returns a generator object, not a list.

    
    gen_exp = (x**2 for x in range(10))  # Memory efficient
    
    
  • Coroutines & yield as Expression: yield can receive data via .send(value). This enables two-way communication.

    
    def coroutine():
    
        while True:
    
            received = yield  # Pauses, receives value
    
            print(f"Received: {received}")
    
    c = coroutine()
    
    next(c)  # Prime the coroutine
    
    c.send("Hello")  # Sends value into paused yield
    
    
  • asyncio Introduction: Framework for writing single-threaded concurrent code using async def (defines a coroutine) and await (pauses coroutine, yields control to event loop). Used for high-concurrency I/O-bound tasks.

C. Context Managers & The with Statement
  • Protocol: An object must implement __enter__(self) and __exit__(self, exc_type, exc_val, exc_tb).

    • __enter__: Executed at start of with block. Value returned becomes variable after as.

    • __exit__: Executed at end of block. Handles cleanup. If exception occurred, exc_type etc. are set; returning True suppresses it.

  • Implementation:

    1. Class-Based: Define a class with __enter__ and __exit__.

    2. contextlib.contextmanager: Decorator for a generator function. yield statement separates setup from cleanup.

      
      from contextlib import contextmanager
      
      @contextmanager
      
      def managed_file(filename):
      
          file = open(filename, 'w')
      
          try:
      
              yield file
      
          finally:
      
              file.close()
      
      
  • Common Use Cases: Guaranteed resource cleanup (files, DB connections, locks), temporary state changes (e.g., decimal context).

D. Metaclasses
  • What is a Metaclass? The "class of a class." Default metaclass is type. It controls how a class is created.

    • MyClass = type('MyClass', (BaseClass,), {'attr': value}) is the low-level creation.
  • Custom Metaclasses: Inherit from type. Override __new__(mcs, name, bases, attrs) (creates the class dict) or __init__(cls, name, bases, attrs) (initializes the class object).

    
    class Meta(type):
    
        def __new__(mcs, name, bases, attrs):
    
            attrs['added_attr'] = 100
    
            return super().__new__(mcs, name, bases, attrs)
    
    class MyClass(metaclass=Meta):
    
        pass
    
    MyClass.added_attr  # 100
    
    
  • Practical Applications: Enforcing API constraints (e.g., all methods must have docstrings), automatic registration of subclasses (plugins), ORM-like field declaration (Django models).


II. DATA SCIENCE & ENGINEERING LIBRARIES (CORE TOOLKIT)

A. NumPy: Numerical Computing Foundation
  • ndarray Object: N-dimensional, homogeneous-typed array.

    • Key Attributes: .shape (tuple of dimensions), .dtype (data type), .ndim (number of axes).

    • Creation: np.array(), np.zeros(), np.ones(), np.arange(), np.linspace().

  • Vectorized Operations: Operations applied element-wise without explicit Python loops. Universal Functions (ufunc) like np.sin, np.exp operate on arrays.

  • Broadcasting Rules: Allows arithmetic on arrays of different shapes. Stretches smaller array to match larger one's shape if compatible.

    • Rule: Two dimensions are compatible if they are equal or one is 1.
  • Array Manipulation:

    • Reshaping: .reshape(), .flatten().

    • Indexing: Basic arr[i,j], slicing arr[1:3, :].

    • Fancy Indexing: Indexing with integer or boolean arrays (arr[[0,2,4]], arr[arr > 5]).

  • Linear Algebra Basics (np.linalg):

    • np.linalg.inv(A) – Matrix inverse.

    • np.linalg.eig(A) – Eigenvalues/vectors.

    • np.linalg.solve(A, b) – Solve linear system Ax = b.

B. Pandas: Data Manipulation & Analysis
  • Core Data Structures:

    • Series: 1D labeled array (index + values).

    • DataFrame: 2D labeled table (columns of potentially different types). Can be thought of as dict of Series.

  • Data I/O: pd.read_csv(), pd.read_excel(), pd.read_sql(). Corresponding .to_*() methods.

  • Data Selection & Filtering:

    • .loc[]: Label-based (includes end point).

    • .iloc[]: Integer position-based (excludes end point).

    • Boolean Indexing: df[df['col'] > 10].

  • Data Cleaning:

    • Missing: df.isna(), df.fillna(value), df.dropna().

    • Duplicates: df.duplicated(), df.drop_duplicates().

  • Grouping & Aggregation (groupby): "Split-Apply-Combine" pattern.

    
    df.groupby('category')['value'].agg(['mean', 'sum', 'count'])
    
    
C. Matplotlib & Seaborn: Data Visualization
  • Matplotlib Architecture:

    • Figure: The top-level container (window, page).

    • Axes: The actual plot area (x/y axis, data, ticks). A Figure contains one or more Axes.

    • Object-Oriented Interface: fig, ax = plt.subplots() then ax.plot().

  • Basic Plot Types (via Axes methods): plot() (line), scatter(), bar(), hist(), imshow().

  • Seaborn: Statistical visualization library built on Matplotlib. Higher-level, prettier defaults.

    • Functions: seaborn.displot() (hist/kde), seaborn.pairplot() (matrix of scatter/hist), seaborn.heatmap(), seaborn.catplot() (categorical).
  • Customization: Set titles (ax.set_title()), labels (ax.set_xlabel()), legends (ax.legend()). Use plt.style.use('seaborn-v0_8-whitegrid') for styles.


III. PERFORMANCE, PARALLELISM & PROFILING

A. Code Optimization & Profiling
  • timeit: Module for timing small code snippets. Avoids many pitfalls of time.time().

    
    python -m timeit "sum(range(1000))"
    
    
  • cProfile: Deterministic profiler. Gives number of calls, total time, per-call time for each function.

    
    python -m cProfile -s cumtime my_script.py
    
    
  • line_profiler (3rd party): Line-by-line timing. Use @profile decorator on functions, run kernprof.

  • memory_profiler (3rd party): Line-by-line memory usage. Use @profile decorator, run python -m memory_profiler script.py.

B. Parallel & Concurrent Execution
  • multiprocessing: Spawns separate processes, each with its own Python interpreter and memory space. Bypasses GIL, ideal for CPU-bound tasks.

    • Process: Low-level control.

    • Pool: High-level pool of worker processes (pool.map(), pool.apply_async()).

  • threading: Spawns threads within same process. Subject to GIL, so not parallel for CPU-bound Python code. Best for I/O-bound tasks (network, disk).

  • concurrent.futures: High-level abstraction. Provides ThreadPoolExecutor and ProcessPoolExecutor with same interface (submit(), map()). Recommended for new code.

C. Efficient Data Storage
  • Serialization Formats:

    • pickle: Python-specific, can serialize almost any object. Security risk (arbitrary code execution). Fast.

    • JSON: Text-based, universal, human-readable. Limited to basic types (dict, list, str, int, float, bool, None).

    • HDF5 (via h5py or pandas.HDFStore): Binary format designed for large, complex numerical datasets. Supports chunking, compression.

  • Memory Views & Buffer Protocol: memoryview allows slicing and manipulation of binary data without copying. Provides zero-copy access to the underlying buffer (e.g., from bytes, bytearray, array.array, NumPy arrays).


IV. SOFTWARE ENGINEERING & BEST PRACTICES (ADVANCED)

A. Testing & Quality
  • pytest (over unittest): More powerful, less boilerplate.

    • Fixtures: Functions that set up test context (@pytest.fixture). Scope (function, class, module, session).

    • Parametrization: @pytest.mark.parametrize runs a test with multiple input sets.

    • Mocking (unittest.mock): Replace parts of system with mock objects. patch() decorator/context manager to temporarily replace objects.

      
      with patch('module.ClassName') as MockClass:
      
          MockClass.return_value = mock_instance
      
          # test code
      
      
  • Test Coverage (coverage.py): Measures how much code is executed by tests. coverage run -m pytest, coverage report, coverage html.

B. Packaging & Distribution
  • Project Structure (Modern):

    
    project/
    
    ├── src/                    # Source code lives here
    
    │   └── mypackage/
    
    │       ├── __init__.py
    
    │       └── module.py
    
    ├── pyproject.toml          # Central config (build system, dependencies)
    
    ├── README.md
    
    └── tests/
    
    
  • pyproject.toml: Standard file (PEP 518/621). Declares build-system ([build-system]) and project metadata ([project]). Replaces setup.py, setup.cfg, MANIFEST.in.

  • Tools: setuptools (traditional), poetry (modern, handles deps & packaging).

  • Virtual Environments: Isolate project dependencies. Use built-in venv (python -m venv .venv) or pipenv/poetry.

C. Documentation
  • Docstring Conventions: Standardized formats for tools like Sphinx.

    • Google Style: Args:, Returns:, Raises: sections.

    • NumPy Style: Similar, with Parameters section.

    • reStructuredText (reST): Sphinx native, uses :param name:, :return:.

  • Sphinx: Generates documentation from docstrings and .rst files.

    • sphinx-quickstart creates structure.

    • autodoc extension pulls docstrings from code.

    • Build with sphinx-build -b html sourcedir builddir.


V. DOMAIN-SPECIFIC APPLICATIONS (ENGINEERING FOCUS)

A. Symbolic Mathematics & Equation Solving (SymPy)
  • Basics: Define symbols with symbols('x y'). Build expressions using Python operators (x**2 + 2*x + 1).

  • Solving Equations:

    • Algebraic: solve(expr, x) solves expr = 0.

    • Differential: dsolve(eq, f(x)) solves ODE eq.

  • Simplification & Expansion:

    • simplify(expr) – General simplification.

    • expand(expr) – Multiply out.

    • factor(expr) – Factor expression.

B. Interfacing with External Tools & Data
  • Network Requests (requests): Simple, elegant HTTP library.

    
    import requests
    
    response = requests.get('https://api.example.com/data')
    
    data = response.json()  # Parse JSON response
    
    
  • Interprocess Communication (IPC) – subprocess: Run external commands, capture output.

    
    import subprocess
    
    result = subprocess.run(['ls', '-l'], capture_output=True, text=True)
    
    print(result.stdout)
    
    

    [!TIP] Prefer subprocess.run() (Python 3.5+) over older call(), check_output().

C. Basic Simulation & Modeling Concepts
  • Numerical Integration (NumPy): Implement simple methods.

    • Trapezoidal Rule:

      
      def trapezoidal(f, a, b, n):
      
          x = np.linspace(a, b, n+1)
      
          y = f(x)
      
          return np.trapz(y, x)  # Or: (b-a)/(2*n) * (y[0] + 2*y[1:-1].sum() + y[-1])
      
      
  • Optimization (scipy.optimize):

    • minimize(fun, x0) – General-purpose minimizer.

    • curve_fit(f, xdata, ydata) – Fit a function f to data.

  • Monte Carlo Methods: Use random sampling for estimation.

    
    # Estimate pi using unit circle sampling
    
    n = 1000000
    
    x, y = np.random.rand(2, n)
    
    inside = (x**2 + y**2) <= 1
    
    pi_est = 4 * inside.mean()
    
    
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in