Lecture 6: Model Evaluation and Data Deception
What Did That Number Leave Out?
Overview
A summary or a high score can hide the part we care about most. We start with the data and the choice of visualization, follow a fraud detector from accuracy to financial consequences, and finish by asking whether each input could have been known at prediction time.
Learning Objectives
By the end of this lecture, you will:
- Detect when summary statistics hide critical patterns using visualization techniques
- Explain how group composition can reverse an aggregate comparison
- Choose appropriate metrics beyond accuracy for imbalanced problems
- Identify future information leaking into a customer-prediction task
- Interpret ML visualizations (ROC curves, confusion matrices) correctly
- State what a classroom test does and does not establish about deployment
Materials
TipQuick Access
Datasets & Acknowledgments
Real Datasets Used
- Datasaurus Dozen: 13 datasets with nearly identical summary statistics but very different spatial patterns
- UC Berkeley Admissions (1973): Historical admissions data illustrating Simpson’s Paradox
- Credit Card Fraud (Kaggle): Transactions from European cardholders over two days; 492 frauds out of 284,807 transactions (0.172%), resulting in a highly imbalanced dataset
- Telco Customer Churn (IBM): The 33-column workbook includes
Churn Reason; it describes a fictional company and supports the prediction-time leakage example
Key Takeaways
WarningCritical Evaluation Principles
- Inspect the observations and relevant groups before trusting a summary.
- Connect each metric to the errors and consequences it counts.
- Ask: “Would I have this information at prediction time?”
- Keep the limits visible: group comparisons alone do not establish causes, and a random split alone does not establish deployment performance.
Previous: ← Lecture 5: Probabilistic Classification | Next: Lecture 7: Overfitting and Regularization →