Skip to content
IT-702 (A) · Data Science/Quick Revision Short Notes

Data Science (IT-702 (A)) - Unit 2 Short Notes

UNIT 2: DATA SCIENCE - EXAM-FOCUSED SHORT NOTES


I. FOUNDATIONS OF DATA

A. Data Types and Characteristics

  • Structured Data: Organized in fixed fields (rows/columns), typically in relational databases. Examples: Excel sheets, SQL tables.

  • Unstructured Data: No predefined model or organization. Examples: Text documents, images, videos, social media posts, audio files.

  • Semi-structured Data: Contains tags or markers (e.g., JSON, XML) but lacks a rigid table structure. Example: Email messages with headers/body.

  • Key Differences:

    | Feature | Structured | Semi-structured | Unstructured | | :--- | :--- | :--- | :--- | | Schema | Fixed, predefined | Flexible, self-describing | None | | Storage | Relational DBs | NoSQL, XML/JSON files | Data lakes, object storage | | Querying | SQL (simple) | Complex (XPath, JSON paths) | Requires NLP/Computer Vision | | Scalability | Vertical scaling | Horizontal scaling | Horizontal scaling |

  • [!TIP] Exam Focus: Be ready to cite specific examples for healthcare (MRI scans - unstructured, patient records - structured), finance (transaction logs - structured, news sentiment - unstructured), and social media (tweets - semi/unstructured).

B. Data Sources and Context

  • Data Science Definition: An interdisciplinary field using scientific methods, algorithms, and systems to extract knowledge and insights from structured and unstructured data.

  • Core Applications: Predictive analytics, recommendation systems, natural language processing, computer vision, fraud detection.

  • Role of a Data Scientist:

    • Responsibilities: Data collection/wrangling, exploratory analysis, model building/validation, deployment, communication.

    • Key Skills: Statistics, programming (Python/R), ML, data visualization, domain expertise, storytelling.


II. DATA PREPARATION AND WRANGLING

A. Data Wrangling (Data Munging)

  • Definition: The process of cleaning, transforming, and mapping raw data into a usable format for analysis.

  • Key Steps:

    1. Discovery: Understanding data structure, quality, and patterns.

    2. Structuring: Organizing data (e.g., pivoting, merging).

    3. Cleaning: Handling missing values, outliers, errors, inconsistencies.

    4. Enriching: Adding relevant data from external sources.

    5. Validating: Ensuring data consistency and quality rules are met.

    6. Publishing: Preparing wrangled data for downstream use (analysis, ML).

  • [!TIP] Common Pitfall: Wrangling often takes 60-80% of a data scientist's time. Skipping it leads to "garbage in, garbage out."

B. Microsoft Excel for Data Analysis

  • Data Validation:

    • Purpose: Restrict input to a cell to a specific type or list, improving accuracy.

    • Techniques:

      • Whole Number/Decimal/Date/Time/Text Length: Set min/max limits.

      • List (Dropdown): Provide predefined choices (e.g., "Yes", "No").

      • Custom Formula: Use =ISNUMBER() or similar for complex rules.

    • Error Alerts: Stop, Warning, or Information messages on invalid input.

  • Lookup & Reference Functions:

    • VLOOKUP(lookup_value, table_array, col_index_num, [range_lookup])

      • Searches first column of table_array for lookup_value.

      • Limitation: Can only look right; slow on large datasets; breaks if column inserted.

    • XLOOKUP(lookup_value, lookup_array, return_array, [if_not_found], [match_mode])

      • Modern replacement. Searches any column/array, returns from any column. Default exact match.

      • Syntax: =XLOOKUP(A2, Products[ID], Products[Price], "Not Found")

    • INDEX + MATCH Combination:

      • MATCH(lookup_value, lookup_array, [match_type]) returns position.

      • INDEX(return_array, MATCH(...)) returns value at that position.

      • Advantage over VLOOKUP: Look left, insert columns safely, more flexible.

      • Example: =INDEX(C2:C100, MATCH(E2, A2:A100, 0))

  • Pivot Tables & Pivoting:

    • Purpose: Summarize, aggregate, and explore large datasets interactively.

    • Creation: Select data > Insert > PivotTable.

    • Areas:

      • Rows/Columns: Categorical grouping.

      • Values: Aggregated metrics (Sum, Count, Average).

      • Filters/Slicers: Dynamic filtering on fields.

    • Types of Pivoting: Row/Column/Value filters, calculated fields/items, grouping (dates, numbers).

  • Scenario Manager (What-If Analysis):

    • Purpose: Save and compare different sets of input values (scenarios) to see their impact on results.

    • Steps: Data > What-If Analysis > Scenario Manager > Add Scenario (name, changing cells) > Show/Summary.

  • Macros:

    • Definition: A recorded sequence of Excel commands/actions to automate repetitive tasks.

    • Steps to Create: View tab > Macros > Record Macro > Perform actions > Stop Recording.

    • Benefit: Saves time, reduces manual error, ensures consistency.

  • Advanced Formulas for Customer Segmentation:

    • SUMIFS(sum_range, criteria_range1, criteria1, [criteria_range2, criteria2], ...)

      • Sums cells meeting multiple criteria.

      • Example (Segment Sales): =SUMIFS(Sales[Amount], Sales[Region], "East", Sales[Product], "Premium")

    • Dynamic Array Formulas (Excel 365/2021):

      • FILTER(array, include, [if_empty]): Filters data based on condition. Spills results automatically.

      • SORT(array, [sort_index], [sort_order], [by_col]): Sorts array.

      • UNIQUE(array): Returns distinct values.

      • SEQUENCE(rows, [cols], [start], [step]): Generates sequence.

      • Customer Segmentation Logic: Use FILTER to extract a segment (e.g., high-value customers), then SORT by spend, UNIQUE to list distinct products bought.

C. Python with Pandas for Data Manipulation

  • Pandas DataFrame:

    • Structure: 2D labeled data structure with columns of potentially different types. Like a spreadsheet or SQL table.

    • Creation: pd.DataFrame(data, index, columns), from dict, list, NumPy array, CSV/Excel read (pd.read_csv()).

    • Key Features: Indexing (.loc, .iloc), handling missing data (isna(), fillna()), merging/joining (merge, concat), groupby operations.

  • Pandas vs. NumPy:

    | Feature | Pandas DataFrame | NumPy Array | | :--- | :--- | :--- | | Data Types | Heterogeneous (columns can have different dtypes) | Homogeneous (single dtype) | | Axes Labels | Row/column labels (index, columns) | Integer axes only | | Primary Use | Tabular data manipulation, analysis | Numerical computing, linear algebra | | Missing Data | Built-in support (NaN) | Requires masking or special handling | | Performance | Slower for pure math, optimized for heterogeneous data | Faster for large numerical operations |

  • String Operations & Regular Expressions (Regex):

    • Access: Use .str accessor on a Series: df['Column'].str.method()

    • Common Methods: .lower(), .upper(), .strip(), .split(), .replace(), .contains(), .extract().

    • Regex with str.contains() & str.extract():

      
      # Extract area code from phone number (format: (XXX) YYY-ZZZZ)
      
      df['Area_Code'] = df['Phone'].str.extract(r'\((\d{3})\)')
      
      # Find emails from a specific domain
      
      df[df['Email'].str.contains(r'@company\.com$', regex=True)]
      
      
  • Data Cleaning & Transformation Practices:

    • Handle Missing Values: df.dropna(), df.fillna(value/method).

    • Type Conversion: df['Column'].astype('int').

    • Duplicates: df.drop_duplicates().

    • Renaming: df.rename(columns={'old':'new'}).

    • Apply Functions: df['New'] = df['Col'].apply(lambda x: x*2) or vectorized operations.


III. DATA EXPLORATION AND VISUALIZATION

A. Descriptive Statistics

  • Measures of Central Tendency:

    • Mean ($\bar{x}$): Sum / Count. Sensitive to outliers.

$$\bar{x} = \frac{\sum_{i=1}^{n} x_i}{n}$$

*   **Median:** Middle value (sorted). Robust to outliers.

*   **Mode:** Most frequent value. Can be multimodal.
  • Measures of Variability (Spread):

    • Range: Max - Min. Sensitive to outliers.

    • Variance ($$\displaystyle s^2 $$): Average squared deviation from mean.

$$s^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1}$$

*   **Standard Deviation ($s$):** Square root of variance. Same units as data.

$$s = \sqrt{s^2}$$

*   **Interquartile Range (IQR):** $Q3 - Q1$. Measures spread of middle 50%, robust.
  • Correlation Analysis:

    • Pearson Correlation Coefficient ($r$): Measures linear relationship strength/direction (-1 to 1).

$$r = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum (x_i - \bar{x})^2 \sum (y_i - \bar{y})^2}}$$

*   **Interpretation:** |r| > 0.7 strong, 0.3-0.7 moderate, <0.3 weak. **Correlation ≠ Causation.**

*   **Visualization:** Scatter plot shows relationship shape.

B. Graphical Representations

  • Pie Chart: Shows parts of a whole. Use for few categories (<6). Avoid for precise comparison.

  • Bar Graph (Chart): Compares categories. Bars are separated. Use for nominal/ordinal data.

  • Histogram: Shows distribution of a single quantitative variable. Bars touch (bins). Reveals shape, skewness, modality.

  • Box Plot (Whisker Plot):

    • Five-Number Summary: Minimum, Q1, Median (Q2), Q3, Maximum.

    • Box: IQR (Q1 to Q3). Line inside = Median.

    • Whiskers: Extend to min/max within 1.5 * IQR. Points beyond are outliers.

    • Use: Compare distributions across categories, identify outliers.

  • Scatter Plot: Shows relationship between two quantitative variables. Each point is an (x,y) pair. Trend line indicates correlation.

  • t-SNE (t-Distributed Stochastic Neighbor Embedding):

    • Principle: Non-linear dimensionality reduction technique for visualizing high-dimensional data (e.g., text vectors, image pixels) in 2D/3D.

    • Process: Measures similarity (probability) between points in high-D and low-D space, minimizes divergence (Kullback-Leibler) between distributions.

    • Use Case: Visualizing clusters of similar documents, image embeddings, gene expression data.

    • Caveat: Results vary with perplexity parameter; not for cluster density interpretation.

C. Python Visualization Libraries

  • Matplotlib: Foundation library. Highly customizable. Basic plots: plt.plot(), plt.bar(), plt.scatter(), plt.hist().

  • Seaborn: Built on Matplotlib. Statistical focus, nicer defaults. Key functions:

    • sns.histplot() / sns.kdeplot() for distributions.

    • sns.boxplot(), sns.violinplot() for categorical comparisons.

    • sns.scatterplot(), sns.lineplot().

    • sns.heatmap() for correlation matrices.

    • sns.pairplot() for pairwise relationships.

  • Code Examples:

    
    import matplotlib.pyplot as plt
    
    import seaborn as sns
    
    # Histogram (Seaborn)
    
    sns.histplot(data=df, x='Age', kde=True)
    
    plt.title('Age Distribution')
    
    plt.show()
    
    # Scatter Plot with Regression Line (Seaborn)
    
    sns.regplot(data=df, x='Income', y='Spending')
    
    plt.show()
    
    # Bar Chart (Matplotlib)
    
    df['Category'].value_counts().plot(kind='bar')
    
    plt.ylabel('Count')
    
    plt.show()
    
    

IV. STATISTICAL ANALYSIS FOR DATA SCIENCE

A. Probability Concepts

  • Conditional Probability: Probability of event A given that event B has occurred.

$$P(A|B) = \frac{P(A \cap B)}{P(B)}, \quad P(B) > 0$$

  • Example: Probability loan defaults (A) given low credit score (B).

B. Hypothesis Testing Fundamentals

  • Type I Error (False Positive): Rejecting a true null hypothesis ($$\displaystyle H_0 $$). Significance level ($\alpha$) controls this (commonly 0.05).

  • Type II Error (False Negative): Failing to reject a false null hypothesis. Power (1-$\beta$) is probability of avoiding it.

  • Trade-off in Multiple Testing: Running many tests increases chance of at least one Type I error (family-wise error rate). Corrections (Bonferroni) reduce $\alpha$ per test, increasing Type II risk.

  • [!TIP] Mnemonic: Type I = Incorrectly found a effect (false alarm). Type II = II (2) = missed a real effect.

C. Regression Analysis

  • Role: Predict a continuous outcome (dependent variable) based on one or more predictors (independent variables). Also for inference (understanding relationships).

  • Linear Regression:

    • Simple: $$\displaystyle y = \beta_0 + \beta_1 x + \epsilon $$

    • Multiple: $$\displaystyle y = \beta_0 + \beta_1 x_1 + ... + \beta_p x_p + \epsilon $$

    • Assumptions: Linearity, Independence, Homoscedasticity, Normality of errors, No multicollinearity.

  • Logistic Regression (Binary Classification):

    • Purpose: Predict probability of a binary outcome (0/1, Yes/No).

    • Model: Uses sigmoid function to model probability.

$$P(Y=1|X) = \frac{1}{1 + e^{-(\beta_0 + \beta_1 X)}}$$

*   **Odds Ratio:** $$\displaystyle e^{\beta_1} $$ represents change in odds of Y=1 for a 1-unit increase in X.

*   **Example (Loan Approval):** Predict `Approved (1/0)` using features: Income, Credit Score, Debt-to-Income. Output is probability; classify using threshold (e.g., 0.5).

V. MACHINE LEARNING ALGORITHMS

A. Decision Trees

  • Structure: Hierarchical model of nodes (splits), branches (decisions), leaves (predictions).

  • Splitting Criteria: Chooses feature/threshold that best separates classes.

    • Gini Impurity: $$\displaystyle G = 1 - \sum_{i=1}^{C} p_i^2 $$. Measures impurity (0 = pure). CART uses this.

    • Entropy / Information Gain: $$\displaystyle H = -\sum p_i \log_2 p_i $$. ID3/C4.5 use this. **Information Gain = Entropy(parent) - Weighted Avg(Entropy(children))$.

  • CART (Classification and Regression Trees): Produces binary trees. Can handle both classification (majority class in leaf) and regression (mean value in leaf).

B. Ensemble Methods

  • Random Forest (Bagging):

    • Principle: Builds many decision trees on bootstrapped samples (with replacement) and averages/ votes results.

    • Key Feature: Feature Randomness: At each split, considers only a random subset of features (max_features). Decorrelates trees.

    • Output: Feature Importance (mean decrease in impurity/Gini).

    • Example: Predict customer churn. Each tree sees different data/subset of features (tenure, contract type, internet service). Majority vote = final prediction.

  • Gradient Boosting (Boosting):

    • Principle: Builds trees sequentially. Each new tree corrects errors of the ensemble so far.

    • Process: Fit tree to residuals (pseudo-residuals) of current model. Add new tree with learning rate ($\alpha$).

    • Example (Gradient Boosting Machines - GBM): Start with simple prediction (mean). Tree 1 fits residuals. Update predictions: pred_new = pred_old + α * tree1_pred. Tree 2 fits new residuals, and so on.

    • Popular Libraries: XGBoost, LightGBM, CatBoost.

C. Parameter Estimation

  • Maximum Likelihood Estimation (MLE):

    • Principle: Find parameter values ($\theta$) that maximize the likelihood function $L(\theta|data)$, i.e., make the observed data most probable.

    • Steps:

      1. Write likelihood function (product of PDFs for i.i.d. data).

      2. Take log (log-likelihood, easier to maximize).

      3. Differentiate w.r.t. $\theta$, set to zero, solve.

    • Application: In logistic regression, MLE finds $\beta$ coefficients that best fit the binary outcome probabilities. In Gaussian distribution, MLE gives $\mu$ = sample mean, $$\displaystyle \sigma^2 $$ = sample variance.


VI. BUSINESS INTELLIGENCE (BI)

A. BI Fundamentals

  • Definition: Technologies, applications, and practices for collecting, integrating, analyzing, and presenting business information to support better decision-making.

  • Core Objectives: Improve operational efficiency, identify new revenue streams, enhance customer experience, gain competitive advantage.

  • Types of BI:

    | Type | Question | Focus | Example | | :--- | :--- | :--- | :--- | | Descriptive | What happened? | Summarize past data | Sales reports, dashboards | | Diagnostic | Why did it happen? | Drill-down, root cause | Why Q3 sales dropped? | | Predictive | What will happen? | Forecast future trends | Sales forecast, churn risk | | Prescriptive | What should we do? | Recommend actions | Optimal pricing, inventory levels |

B. BI Tools and Ecosystem

  • Comparison: Tableau vs. Power BI

    | Feature | Tableau | Microsoft Power BI | | :--- | :--- | :--- | | Usability | Drag-and-drop, intuitive for ad-hoc analysis. Steeper initial learning for advanced features. | Deeply integrated with Microsoft ecosystem (Excel, Azure). Familiar UI for Office users. | | Key Features | Superior visualizations,地理 mapping, strong community. Tableau Prep for data shaping. | Power Query (excellent ETL), DAX language (powerful measures), Power BI Service (cloud sharing). | | Data Handling | Handles large volumes well, connects to diverse sources. | Efficient with Microsoft sources (SQL Server, Dynamics). Performance can dip with very large datasets. | | Cost | Higher (per-user licensing). | More affordable, especially for organizations with Microsoft agreements. |

C. BI in Decision-Making

  • How BI Supports Strategic Decisions:

    1. Dashboards: Real-time visual KPIs (e.g., revenue, customer acquisition cost).

    2. Ad-hoc Reporting: Self-service exploration by business users.

    3. Data-driven Culture: Replaces gut feeling with evidence.

    4. What-if Analysis: Scenario planning (e.g., impact of price change).

    5. Performance Monitoring: Track progress against goals (OKRs, KPIs).

  • [!TIP] Example: A retail BI dashboard showing sales by region, product category, and sales rep allows VP to quickly identify underperforming regions and allocate marketing budget accordingly.

D. Ethical Considerations in BI

  • Ethical Issues:

    • Privacy: Collecting/using personal data without consent.

    • Misrepresentation: Distorting data in visualizations (truncated axes, inappropriate chart types).

    • Unfair Discrimination: Biased algorithms in hiring, lending, or pricing (e.g., charging more in low-income zip codes).

    • Lack of Transparency: "Black box" models that cannot be explained.

  • Mitigation: Data governance, anonymization, ethical review boards, explainable AI (XAI) techniques.


VII. ETHICS, SECURITY, AND SOCIAL IMPACT

A. Ethical Issues in Data Science

  • Unfair Discrimination (Algorithmic Bias):

    • Definition: Systematic and unfair disadvantage to certain groups (race, gender, age) due to biased data or algorithms.

    • Example: A hiring algorithm trained on historical data from a male-dominated company learns to downgrade resumes with "women's" sports or college names.

  • Reinforcing Human Biases:

    • Feedback Loops: Biased predictions influence future data, amplifying bias.

      • Example: Predictive policing targets minority neighborhoods → more arrests there → data shows "high crime" → more policing. Cycle reinforces itself.
    • Representation Bias: Underrepresented groups in training data lead to poor model performance for them (e.g., facial recognition accuracy lower for darker skin tones).

B. Security Issues

  • Data Breaches: Unauthorized access to sensitive data (PII, financial records).

  • Privacy Violations: Misuse of data beyond consent (e.g., selling user data).

  • Model Theft/Inversion: Stealing proprietary ML models or inferring training data from model outputs (membership inference attacks).

  • Adversarial Attacks: Malicious inputs designed to fool ML models (e.g., slight perturbation to stop sign image classified as speed limit).

C. Social Implications

  • Impact of Data-Driven Decisions:

    • Positive: Improved healthcare diagnostics, efficient resource allocation, personalized education.

    • Negative: Filter bubbles (social media algorithms), credit scoring bias, automated hiring discrimination, surveillance capitalism.

    • Job Displacement: Automation of routine analytical tasks.

    • Digital Divide: Access to data/technology not equal across societies.


VIII. ADDITIONAL CONCEPTS AND INTEGRATIVE TOPICS

A. Exploratory Data Analysis (EDA)

  • Purpose: Understand data, discover patterns, spot anomalies, test hypotheses before formal modeling.

  • Process:

    1. Univariate Analysis: Summary stats, histograms, box plots for each variable.

    2. Bivariate/Multivariate Analysis: Scatter plots, correlation matrices, pivot tables, groupby aggregations.

    3. Missing Data Analysis: Visualize patterns (e.g., sns.heatmap(df.isna())).

    4. Outlier Detection: Box plots, IQR method, Z-score.

  • Tools: Pandas profiling, df.describe(), df.info(), visualization libraries.

B. Data Analysis vs. Data Scientist Role

Aspect Data Analyst Data Scientist
Primary Focus Answering business questions from existing data. Reporting, dashboards. Building predictive/ML models, advanced analytics, product innovation.
Skills SQL, Excel, BI tools (Power BI/Tableau), basic stats, visualization. Python/R, ML libraries (scikit-learn, TensorFlow), advanced stats, software engineering, big data tools.
Output Reports, dashboards, visualizations, insights. Predictive models, ML pipelines, algorithms, research prototypes.
Question Type "What happened last quarter?" "Why did sales drop?" "What will happen next month?" "How can we optimize this process?"

C. Regular Expressions in Practice

  • Pattern Matching: Used for extraction, validation, replacement, splitting text data.

  • Common Patterns:

    • \d - digit, \w - word char, \s - whitespace.

    • + (1 or more), * (0 or more), ? (0 or 1).

    • ^ (start), $ (end).

    • [abc] (set), [^abc] (negation).

    • ( ) - capture group.

  • Integrated with Pandas:

    
    # Extract years from a 'Date' string column (format: YYYY-MM-DD)
    
    df['Year'] = df['Date'].str.extract(r'(\d{4})-\d{2}-\d{2}')
    
    # Validate Indian phone numbers (10 digits starting 6-9)
    
    df['Valid_Phone'] = df['Phone'].str.contains(r'^[6-9]\d{9}$')
    
    # Replace multiple spaces with single space
    
    df['Text'] = df['Text'].str.replace(r'\s+', ' ', regex=True)
    
    

BOXED KEY FORMULAS & CONCEPTS

  • Mean: $$\displaystyle \bar{x} = \frac{\sum x_i}{n} $$

  • Sample Standard Deviation: $$\displaystyle s = \sqrt{\frac{\sum (x_i - \bar{x})^2}{n-1}} $$

  • Pearson Correlation: $$\displaystyle r = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum (x_i - \bar{x})^2 \sum (y_i - \bar{y})^2}} $$

  • Logistic Regression (Sigmoid): $$\displaystyle P(Y=1|X) = \frac{1}{1 + e^{-(\beta_0 + \beta_1 X)}} $$

  • Conditional Probability: $$\displaystyle P(A|B) = \frac{P(A \cap B)}{P(B)} $$

  • Gini Impurity: $$\displaystyle G = 1 - \sum p_i^2 $$

  • Entropy: $$\displaystyle H = -\sum p_i \log_2 p_i $$

DiagramCANVAS: A flowchart of the Data Wrangling process: Discovery -> Structuring -> Cleaning -> Enriching -> Validating -> Publishing. Each step has icons (magnifying glass, database, broom, plus sign, checkmark, export arrow).
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in