Imbalanced Data

Myths, Mistakes and Modern Solutions

Imbalanced Data - Myths, Mistakes and Modern Solutions

Imbalanced Data

Myths, Mistakes and Modern Solutions

eBook

Class imbalance isn’t a problem—poor methodology is.

SMOTE, once a go-to solution, is frequently misapplied, introducing bias rather than solving it.

This book challenges outdated practices and provides rigorous, data-driven alternatives. We focus on selecting the right tools—threshold tuning, real costs (not class frequencies), and strategic evaluation metrics—to build models that work.

Ready to learn from mistakes, move beyond myths and master modern solutions? Let’s begin.




Book Description


Class imbalance isn't a problem.

Contrary to popular belief, imbalanced data does not inherently harm model performance—poor methodology does. SMOTE, once a go-to solution, is frequently misapplied, introducing bias rather than solving it.

Similarly, the default 0.5 probability threshold persists despite its misalignment with real-world costs and decision requirements.

This book isn’t about "fixing imbalance". It’s about choosing the right strategy—or none at all—based on data, domain knowledge, and suitable performance metrics.

We focus not on balancing datasets for the sake of it, but on selecting the right tools—threshold tuning, real costs, and strategic evaluation metrics—to build models that work in practice, not just in theory.

If you’re done with oversimplified rules and ready to engineer solutions that actually work, let’s begin.

Table of Content


1. Imbalanced Data: Class Frequency is Not The Problem

  1. Imbalanced Datasets: What Are They?
  2. What Factors Influence the Classification of Imbalanced Datasets?
  3. Resampling in The Age of GBMs
  4. Prediction is Not Classification
  5. How to Approach Imbalanced Learning

2. Metrics that Matter (and Pitfalls to Avoid)

  1. Understanding the Output of Machine Learning Models
  2. Understanding What Metrics Measure
  3. Threshold Dependent vs Threshold Independent Metrics
  4. What Happens When We Threshold Predictions
  5. Why Choosing The Right Metric Matters
  6. How To Choose the Right Metric
  7. Making Sense of Classification Metrics
  8. Choosing the Right Threshold For Classification Metrics
  9. Understanding Ranking Metrics

3. Probability Calibration: When 70% Means 70%

  1. Reliable Probability Estimates: What Are They?
  2. Calibrated Probabilities: Why Do They Matter?
  3. Assessing Probability Calibration: Reliability Diagrams
  4. What Makes Calibration Assessment Hard
  5. What Breaks Probability Calibration
  6. Scoring Functions: Training Models to Be Calibrated
  7. Recalibration: Correcting Biased Probabilities
  8. Recalibrating Models in Python

4. Cost-Sensitive Learning: Thresholds, Weights, and Decisions

  1. Cost-sensitive Learning: What is it?
  2. Making Cost-Sensitive Decisions
  3. Thresholding, Weights, and Resampling Are Equivalent
  4. Reweighting, Resampling, Thresholding: Different Paths to the Same Goal
  5. Thresholds and Weights: Theory Meets Evidence
  6. Empirical Thresholding: Finding the Right Decision Rule
  7. When Cost is Not Class Frequency

5. Undersampling and Cleaning Methods: Discarding Hard-Won Data

  1. Undersampling: What Is It?
  2. The Need for Undersampling: A Historical Perspective
  3. The Value of Undersampling: A Mixed and Fragile Picture
  4. A Modern Reassessment of Undersampling
  5. Resampling is Threshold Shifting
  6. Undersampling in Python

6. Oversampling and SMOTE: The Illusion of Better Models

  1. Oversampling: What Is It?
  2. The Birth of Oversampling and SMOTE
  3. The Rise of Oversampling and SMOTE
  4. The Fall of Oversampling and SMOTE
  5. A Modern Reassessment of Oversampling
  6. Is There Any Place for Oversampling?

7. Imbalanced Learning in the Age of AI

  1. The Workflow We Want the Machine to Follow
  2. Ask AI to Analyse the Data
  3. Ask AI to Build the Model
  4. Audit the Pipeline
  5. Ask AI to Rebuild the Pipeline
  6. How I Work With AI
  7. Last Words on Imbalanced Data

👉 epub and pdf copies
👉 Paperback with our partners
👉 200 pages
👉 English

Author

Soledad Galli, PhD


Sole is a lead data scientist, instructor, and developer of open source software. She created and maintains the Python library Feature-engine, which allows us to impute data, encode categorical variables, transform, create, and select features. Sole is also the author of the"Python Feature Engineering Cookbook," published by Packt.

More about Sole on LinkedIn.

Soledad Galli, PhD

What our readers say

eBook Pricing



Can't afford it? Get in touch.

Paperback


Get a Paperback copy with our Partner Lulu Press.