UNIT 2: DATA SCIENCE - EXAM-FOCUSED SHORT NOTES
I. FOUNDATIONS OF DATA
A. Data Types and Characteristics
-
Structured Data: Organized in fixed fields (rows/columns), typically in relational databases. Examples: Excel sheets, SQL tables.
-
Unstructured Data: No predefined model or organization. Examples: Text documents, images, videos, social media posts, audio files.
-
Semi-structured Data: Contains tags or markers (e.g., JSON, XML) but lacks a rigid table structure. Example: Email messages with headers/body.
-
Key Differences:
| Feature | Structured | Semi-structured | Unstructured | | :--- | :--- | :--- | :--- | | Schema | Fixed, predefined | Flexible, self-describing | None | | Storage | Relational DBs | NoSQL, XML/JSON files | Data lakes, object storage | | Querying | SQL (simple) | Complex (XPath, JSON paths) | Requires NLP/Computer Vision | | Scalability | Vertical scaling | Horizontal scaling | Horizontal scaling |
-
[!TIP] Exam Focus: Be ready to cite specific examples for healthcare (MRI scans - unstructured, patient records - structured), finance (transaction logs - structured, news sentiment - unstructured), and social media (tweets - semi/unstructured).
B. Data Sources and Context
-
Data Science Definition: An interdisciplinary field using scientific methods, algorithms, and systems to extract knowledge and insights from structured and unstructured data.
-
Core Applications: Predictive analytics, recommendation systems, natural language processing, computer vision, fraud detection.
-
Role of a Data Scientist:
-
Responsibilities: Data collection/wrangling, exploratory analysis, model building/validation, deployment, communication.
-
Key Skills: Statistics, programming (Python/R), ML, data visualization, domain expertise, storytelling.
-
II. DATA PREPARATION AND WRANGLING
A. Data Wrangling (Data Munging)
-
Definition: The process of cleaning, transforming, and mapping raw data into a usable format for analysis.
-
Key Steps:
-
Discovery: Understanding data structure, quality, and patterns.
-
Structuring: Organizing data (e.g., pivoting, merging).
-
Cleaning: Handling missing values, outliers, errors, inconsistencies.
-
Enriching: Adding relevant data from external sources.
-
Validating: Ensuring data consistency and quality rules are met.
-
Publishing: Preparing wrangled data for downstream use (analysis, ML).
-
-
[!TIP] Common Pitfall: Wrangling often takes 60-80% of a data scientist's time. Skipping it leads to "garbage in, garbage out."
B. Microsoft Excel for Data Analysis
-
Data Validation:
-
Purpose: Restrict input to a cell to a specific type or list, improving accuracy.
-
Techniques:
-
Whole Number/Decimal/Date/Time/Text Length: Set min/max limits.
-
List (Dropdown): Provide predefined choices (e.g., "Yes", "No").
-
Custom Formula: Use
=ISNUMBER()or similar for complex rules.
-
-
Error Alerts: Stop, Warning, or Information messages on invalid input.
-
-
Lookup & Reference Functions:
-
VLOOKUP(lookup_value, table_array, col_index_num, [range_lookup])-
Searches first column of
table_arrayforlookup_value. -
Limitation: Can only look right; slow on large datasets; breaks if column inserted.
-
-
XLOOKUP(lookup_value, lookup_array, return_array, [if_not_found], [match_mode])-
Modern replacement. Searches any column/array, returns from any column. Default exact match.
-
Syntax:
=XLOOKUP(A2, Products[ID], Products[Price], "Not Found")
-
-
INDEX+MATCHCombination:-
MATCH(lookup_value, lookup_array, [match_type])returns position. -
INDEX(return_array, MATCH(...))returns value at that position. -
Advantage over VLOOKUP: Look left, insert columns safely, more flexible.
-
Example:
=INDEX(C2:C100, MATCH(E2, A2:A100, 0))
-
-
-
Pivot Tables & Pivoting:
-
Purpose: Summarize, aggregate, and explore large datasets interactively.
-
Creation: Select data > Insert > PivotTable.
-
Areas:
-
Rows/Columns: Categorical grouping.
-
Values: Aggregated metrics (Sum, Count, Average).
-
Filters/Slicers: Dynamic filtering on fields.
-
-
Types of Pivoting: Row/Column/Value filters, calculated fields/items, grouping (dates, numbers).
-
-
Scenario Manager (What-If Analysis):
-
Purpose: Save and compare different sets of input values (scenarios) to see their impact on results.
-
Steps: Data > What-If Analysis > Scenario Manager > Add Scenario (name, changing cells) > Show/Summary.
-
-
Macros:
-
Definition: A recorded sequence of Excel commands/actions to automate repetitive tasks.
-
Steps to Create: View tab > Macros > Record Macro > Perform actions > Stop Recording.
-
Benefit: Saves time, reduces manual error, ensures consistency.
-
-
Advanced Formulas for Customer Segmentation:
-
SUMIFS(sum_range, criteria_range1, criteria1, [criteria_range2, criteria2], ...)-
Sums cells meeting multiple criteria.
-
Example (Segment Sales):
=SUMIFS(Sales[Amount], Sales[Region], "East", Sales[Product], "Premium")
-
-
Dynamic Array Formulas (Excel 365/2021):
-
FILTER(array, include, [if_empty]): Filters data based on condition. Spills results automatically. -
SORT(array, [sort_index], [sort_order], [by_col]): Sorts array. -
UNIQUE(array): Returns distinct values. -
SEQUENCE(rows, [cols], [start], [step]): Generates sequence. -
Customer Segmentation Logic: Use
FILTERto extract a segment (e.g., high-value customers), thenSORTby spend,UNIQUEto list distinct products bought.
-
-
C. Python with Pandas for Data Manipulation
-
Pandas DataFrame:
-
Structure: 2D labeled data structure with columns of potentially different types. Like a spreadsheet or SQL table.
-
Creation:
pd.DataFrame(data, index, columns), from dict, list, NumPy array, CSV/Excel read (pd.read_csv()). -
Key Features: Indexing (
.loc,.iloc), handling missing data (isna(),fillna()), merging/joining (merge,concat), groupby operations.
-
-
Pandas vs. NumPy:
| Feature | Pandas DataFrame | NumPy Array | | :--- | :--- | :--- | | Data Types | Heterogeneous (columns can have different dtypes) | Homogeneous (single dtype) | | Axes Labels | Row/column labels (
index,columns) | Integer axes only | | Primary Use | Tabular data manipulation, analysis | Numerical computing, linear algebra | | Missing Data | Built-in support (NaN) | Requires masking or special handling | | Performance | Slower for pure math, optimized for heterogeneous data | Faster for large numerical operations | -
String Operations & Regular Expressions (Regex):
-
Access: Use
.straccessor on a Series:df['Column'].str.method() -
Common Methods:
.lower(),.upper(),.strip(),.split(),.replace(),.contains(),.extract(). -
Regex with
str.contains()&str.extract():# Extract area code from phone number (format: (XXX) YYY-ZZZZ) df['Area_Code'] = df['Phone'].str.extract(r'\((\d{3})\)') # Find emails from a specific domain df[df['Email'].str.contains(r'@company\.com$', regex=True)]
-
-
Data Cleaning & Transformation Practices:
-
Handle Missing Values:
df.dropna(),df.fillna(value/method). -
Type Conversion:
df['Column'].astype('int'). -
Duplicates:
df.drop_duplicates(). -
Renaming:
df.rename(columns={'old':'new'}). -
Apply Functions:
df['New'] = df['Col'].apply(lambda x: x*2)or vectorized operations.
-
III. DATA EXPLORATION AND VISUALIZATION
A. Descriptive Statistics
-
Measures of Central Tendency:
- Mean ($\bar{x}$): Sum / Count. Sensitive to outliers.
$$\bar{x} = \frac{\sum_{i=1}^{n} x_i}{n}$$
* **Median:** Middle value (sorted). Robust to outliers.
* **Mode:** Most frequent value. Can be multimodal.
-
Measures of Variability (Spread):
-
Range: Max - Min. Sensitive to outliers.
-
Variance ($$\displaystyle s^2 $$): Average squared deviation from mean.
-
$$s^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1}$$
* **Standard Deviation ($s$):** Square root of variance. Same units as data.
$$s = \sqrt{s^2}$$
* **Interquartile Range (IQR):** $Q3 - Q1$. Measures spread of middle 50%, robust.
-
Correlation Analysis:
- Pearson Correlation Coefficient ($r$): Measures linear relationship strength/direction (-1 to 1).
$$r = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum (x_i - \bar{x})^2 \sum (y_i - \bar{y})^2}}$$
* **Interpretation:** |r| > 0.7 strong, 0.3-0.7 moderate, <0.3 weak. **Correlation ≠ Causation.**
* **Visualization:** Scatter plot shows relationship shape.
B. Graphical Representations
-
Pie Chart: Shows parts of a whole. Use for few categories (<6). Avoid for precise comparison.
-
Bar Graph (Chart): Compares categories. Bars are separated. Use for nominal/ordinal data.
-
Histogram: Shows distribution of a single quantitative variable. Bars touch (bins). Reveals shape, skewness, modality.
-
Box Plot (Whisker Plot):
-
Five-Number Summary: Minimum, Q1, Median (Q2), Q3, Maximum.
-
Box: IQR (Q1 to Q3). Line inside = Median.
-
Whiskers: Extend to min/max within
1.5 * IQR. Points beyond are outliers. -
Use: Compare distributions across categories, identify outliers.
-
-
Scatter Plot: Shows relationship between two quantitative variables. Each point is an (x,y) pair. Trend line indicates correlation.
-
t-SNE (t-Distributed Stochastic Neighbor Embedding):
-
Principle: Non-linear dimensionality reduction technique for visualizing high-dimensional data (e.g., text vectors, image pixels) in 2D/3D.
-
Process: Measures similarity (probability) between points in high-D and low-D space, minimizes divergence (Kullback-Leibler) between distributions.
-
Use Case: Visualizing clusters of similar documents, image embeddings, gene expression data.
-
Caveat: Results vary with perplexity parameter; not for cluster density interpretation.
-
C. Python Visualization Libraries
-
Matplotlib: Foundation library. Highly customizable. Basic plots:
plt.plot(),plt.bar(),plt.scatter(),plt.hist(). -
Seaborn: Built on Matplotlib. Statistical focus, nicer defaults. Key functions:
-
sns.histplot()/sns.kdeplot()for distributions. -
sns.boxplot(),sns.violinplot()for categorical comparisons. -
sns.scatterplot(),sns.lineplot(). -
sns.heatmap()for correlation matrices. -
sns.pairplot()for pairwise relationships.
-
-
Code Examples:
import matplotlib.pyplot as plt import seaborn as sns # Histogram (Seaborn) sns.histplot(data=df, x='Age', kde=True) plt.title('Age Distribution') plt.show() # Scatter Plot with Regression Line (Seaborn) sns.regplot(data=df, x='Income', y='Spending') plt.show() # Bar Chart (Matplotlib) df['Category'].value_counts().plot(kind='bar') plt.ylabel('Count') plt.show()
IV. STATISTICAL ANALYSIS FOR DATA SCIENCE
A. Probability Concepts
- Conditional Probability: Probability of event A given that event B has occurred.
$$P(A|B) = \frac{P(A \cap B)}{P(B)}, \quad P(B) > 0$$
- Example: Probability loan defaults (
A) given low credit score (B).
B. Hypothesis Testing Fundamentals
-
Type I Error (False Positive): Rejecting a true null hypothesis ($$\displaystyle H_0 $$). Significance level ($\alpha$) controls this (commonly 0.05).
-
Type II Error (False Negative): Failing to reject a false null hypothesis. Power (1-$\beta$) is probability of avoiding it.
-
Trade-off in Multiple Testing: Running many tests increases chance of at least one Type I error (family-wise error rate). Corrections (Bonferroni) reduce $\alpha$ per test, increasing Type II risk.
-
[!TIP] Mnemonic: Type I = Incorrectly found a effect (false alarm). Type II = II (2) = missed a real effect.
C. Regression Analysis
-
Role: Predict a continuous outcome (dependent variable) based on one or more predictors (independent variables). Also for inference (understanding relationships).
-
Linear Regression:
-
Simple: $$\displaystyle y = \beta_0 + \beta_1 x + \epsilon $$
-
Multiple: $$\displaystyle y = \beta_0 + \beta_1 x_1 + ... + \beta_p x_p + \epsilon $$
-
Assumptions: Linearity, Independence, Homoscedasticity, Normality of errors, No multicollinearity.
-
-
Logistic Regression (Binary Classification):
-
Purpose: Predict probability of a binary outcome (0/1, Yes/No).
-
Model: Uses sigmoid function to model probability.
-
$$P(Y=1|X) = \frac{1}{1 + e^{-(\beta_0 + \beta_1 X)}}$$
* **Odds Ratio:** $$\displaystyle e^{\beta_1} $$ represents change in odds of Y=1 for a 1-unit increase in X.
* **Example (Loan Approval):** Predict `Approved (1/0)` using features: Income, Credit Score, Debt-to-Income. Output is probability; classify using threshold (e.g., 0.5).
V. MACHINE LEARNING ALGORITHMS
A. Decision Trees
-
Structure: Hierarchical model of nodes (splits), branches (decisions), leaves (predictions).
-
Splitting Criteria: Chooses feature/threshold that best separates classes.
-
Gini Impurity: $$\displaystyle G = 1 - \sum_{i=1}^{C} p_i^2 $$. Measures impurity (0 = pure). CART uses this.
-
Entropy / Information Gain: $$\displaystyle H = -\sum p_i \log_2 p_i $$. ID3/C4.5 use this. **Information Gain = Entropy(parent) - Weighted Avg(Entropy(children))$.
-
-
CART (Classification and Regression Trees): Produces binary trees. Can handle both classification (majority class in leaf) and regression (mean value in leaf).
B. Ensemble Methods
-
Random Forest (Bagging):
-
Principle: Builds many decision trees on bootstrapped samples (with replacement) and averages/ votes results.
-
Key Feature: Feature Randomness: At each split, considers only a random subset of features (
max_features). Decorrelates trees. -
Output: Feature Importance (mean decrease in impurity/Gini).
-
Example: Predict customer churn. Each tree sees different data/subset of features (tenure, contract type, internet service). Majority vote = final prediction.
-
-
Gradient Boosting (Boosting):
-
Principle: Builds trees sequentially. Each new tree corrects errors of the ensemble so far.
-
Process: Fit tree to residuals (pseudo-residuals) of current model. Add new tree with learning rate ($\alpha$).
-
Example (Gradient Boosting Machines - GBM): Start with simple prediction (mean). Tree 1 fits residuals. Update predictions:
pred_new = pred_old + α * tree1_pred. Tree 2 fits new residuals, and so on. -
Popular Libraries: XGBoost, LightGBM, CatBoost.
-
C. Parameter Estimation
-
Maximum Likelihood Estimation (MLE):
-
Principle: Find parameter values ($\theta$) that maximize the likelihood function $L(\theta|data)$, i.e., make the observed data most probable.
-
Steps:
-
Write likelihood function (product of PDFs for i.i.d. data).
-
Take log (log-likelihood, easier to maximize).
-
Differentiate w.r.t. $\theta$, set to zero, solve.
-
-
Application: In logistic regression, MLE finds $\beta$ coefficients that best fit the binary outcome probabilities. In Gaussian distribution, MLE gives $\mu$ = sample mean, $$\displaystyle \sigma^2 $$ = sample variance.
-
VI. BUSINESS INTELLIGENCE (BI)
A. BI Fundamentals
-
Definition: Technologies, applications, and practices for collecting, integrating, analyzing, and presenting business information to support better decision-making.
-
Core Objectives: Improve operational efficiency, identify new revenue streams, enhance customer experience, gain competitive advantage.
-
Types of BI:
| Type | Question | Focus | Example | | :--- | :--- | :--- | :--- | | Descriptive | What happened? | Summarize past data | Sales reports, dashboards | | Diagnostic | Why did it happen? | Drill-down, root cause | Why Q3 sales dropped? | | Predictive | What will happen? | Forecast future trends | Sales forecast, churn risk | | Prescriptive | What should we do? | Recommend actions | Optimal pricing, inventory levels |
B. BI Tools and Ecosystem
-
Comparison: Tableau vs. Power BI
| Feature | Tableau | Microsoft Power BI | | :--- | :--- | :--- | | Usability | Drag-and-drop, intuitive for ad-hoc analysis. Steeper initial learning for advanced features. | Deeply integrated with Microsoft ecosystem (Excel, Azure). Familiar UI for Office users. | | Key Features | Superior visualizations,地理 mapping, strong community. Tableau Prep for data shaping. | Power Query (excellent ETL), DAX language (powerful measures), Power BI Service (cloud sharing). | | Data Handling | Handles large volumes well, connects to diverse sources. | Efficient with Microsoft sources (SQL Server, Dynamics). Performance can dip with very large datasets. | | Cost | Higher (per-user licensing). | More affordable, especially for organizations with Microsoft agreements. |
C. BI in Decision-Making
-
How BI Supports Strategic Decisions:
-
Dashboards: Real-time visual KPIs (e.g., revenue, customer acquisition cost).
-
Ad-hoc Reporting: Self-service exploration by business users.
-
Data-driven Culture: Replaces gut feeling with evidence.
-
What-if Analysis: Scenario planning (e.g., impact of price change).
-
Performance Monitoring: Track progress against goals (OKRs, KPIs).
-
-
[!TIP] Example: A retail BI dashboard showing sales by region, product category, and sales rep allows VP to quickly identify underperforming regions and allocate marketing budget accordingly.
D. Ethical Considerations in BI
-
Ethical Issues:
-
Privacy: Collecting/using personal data without consent.
-
Misrepresentation: Distorting data in visualizations (truncated axes, inappropriate chart types).
-
Unfair Discrimination: Biased algorithms in hiring, lending, or pricing (e.g., charging more in low-income zip codes).
-
Lack of Transparency: "Black box" models that cannot be explained.
-
-
Mitigation: Data governance, anonymization, ethical review boards, explainable AI (XAI) techniques.
VII. ETHICS, SECURITY, AND SOCIAL IMPACT
A. Ethical Issues in Data Science
-
Unfair Discrimination (Algorithmic Bias):
-
Definition: Systematic and unfair disadvantage to certain groups (race, gender, age) due to biased data or algorithms.
-
Example: A hiring algorithm trained on historical data from a male-dominated company learns to downgrade resumes with "women's" sports or college names.
-
-
Reinforcing Human Biases:
-
Feedback Loops: Biased predictions influence future data, amplifying bias.
- Example: Predictive policing targets minority neighborhoods → more arrests there → data shows "high crime" → more policing. Cycle reinforces itself.
-
Representation Bias: Underrepresented groups in training data lead to poor model performance for them (e.g., facial recognition accuracy lower for darker skin tones).
-
B. Security Issues
-
Data Breaches: Unauthorized access to sensitive data (PII, financial records).
-
Privacy Violations: Misuse of data beyond consent (e.g., selling user data).
-
Model Theft/Inversion: Stealing proprietary ML models or inferring training data from model outputs (membership inference attacks).
-
Adversarial Attacks: Malicious inputs designed to fool ML models (e.g., slight perturbation to stop sign image classified as speed limit).
C. Social Implications
-
Impact of Data-Driven Decisions:
-
Positive: Improved healthcare diagnostics, efficient resource allocation, personalized education.
-
Negative: Filter bubbles (social media algorithms), credit scoring bias, automated hiring discrimination, surveillance capitalism.
-
Job Displacement: Automation of routine analytical tasks.
-
Digital Divide: Access to data/technology not equal across societies.
-
VIII. ADDITIONAL CONCEPTS AND INTEGRATIVE TOPICS
A. Exploratory Data Analysis (EDA)
-
Purpose: Understand data, discover patterns, spot anomalies, test hypotheses before formal modeling.
-
Process:
-
Univariate Analysis: Summary stats, histograms, box plots for each variable.
-
Bivariate/Multivariate Analysis: Scatter plots, correlation matrices, pivot tables, groupby aggregations.
-
Missing Data Analysis: Visualize patterns (e.g.,
sns.heatmap(df.isna())). -
Outlier Detection: Box plots, IQR method, Z-score.
-
-
Tools: Pandas profiling,
df.describe(),df.info(), visualization libraries.
B. Data Analysis vs. Data Scientist Role
| Aspect | Data Analyst | Data Scientist |
|---|---|---|
| Primary Focus | Answering business questions from existing data. Reporting, dashboards. | Building predictive/ML models, advanced analytics, product innovation. |
| Skills | SQL, Excel, BI tools (Power BI/Tableau), basic stats, visualization. | Python/R, ML libraries (scikit-learn, TensorFlow), advanced stats, software engineering, big data tools. |
| Output | Reports, dashboards, visualizations, insights. | Predictive models, ML pipelines, algorithms, research prototypes. |
| Question Type | "What happened last quarter?" "Why did sales drop?" | "What will happen next month?" "How can we optimize this process?" |
C. Regular Expressions in Practice
-
Pattern Matching: Used for extraction, validation, replacement, splitting text data.
-
Common Patterns:
-
\d- digit,\w- word char,\s- whitespace. -
+(1 or more),*(0 or more),?(0 or 1). -
^(start),$(end). -
[abc](set),[^abc](negation). -
( )- capture group.
-
-
Integrated with Pandas:
# Extract years from a 'Date' string column (format: YYYY-MM-DD) df['Year'] = df['Date'].str.extract(r'(\d{4})-\d{2}-\d{2}') # Validate Indian phone numbers (10 digits starting 6-9) df['Valid_Phone'] = df['Phone'].str.contains(r'^[6-9]\d{9}$') # Replace multiple spaces with single space df['Text'] = df['Text'].str.replace(r'\s+', ' ', regex=True)
BOXED KEY FORMULAS & CONCEPTS
-
Mean: $$\displaystyle \bar{x} = \frac{\sum x_i}{n} $$
-
Sample Standard Deviation: $$\displaystyle s = \sqrt{\frac{\sum (x_i - \bar{x})^2}{n-1}} $$
-
Pearson Correlation: $$\displaystyle r = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum (x_i - \bar{x})^2 \sum (y_i - \bar{y})^2}} $$
-
Logistic Regression (Sigmoid): $$\displaystyle P(Y=1|X) = \frac{1}{1 + e^{-(\beta_0 + \beta_1 X)}} $$
-
Conditional Probability: $$\displaystyle P(A|B) = \frac{P(A \cap B)}{P(B)} $$
-
Gini Impurity: $$\displaystyle G = 1 - \sum p_i^2 $$
-
Entropy: $$\displaystyle H = -\sum p_i \log_2 p_i $$