FirstHack Learn
Log in Sign up free
Lessons in this course 0/6 All courses Machine Learning Foundations

AI & DS

Progress0 / 6 lessons
  1. 1. What learning from data means, and when ML is the wrong tool
  2. 2. Supervised vs unsupervised learning
  3. 3. Linear regression from scratch, then with scikit-learn
  4. 4. Classification, logistic regression and the confusion matrix
  5. 5. Overfitting, train/test split and cross-validation
  6. 6. Why accuracy is a bad metric on imbalanced data

Courses › Machine Learning Foundations

Why accuracy is a bad metric on imbalanced data

A 99 percent model that catches nothing, and what to report instead.

12 min read · Lesson 6 of 6 · Pro

The 99 percent model that is useless

Fraud detection. Ten thousand transactions, one hundred of them fraudulent. That is one percent.

Here is a model:

Python 3
def detect_fraud(transaction):
    return 0        # never fraud

Accuracy: 9,900 correct out of 10,000. 99 percent. It catches zero fraud. It would have been just as accurate if written on paper.

This is not a trick example. Most problems worth solving are imbalanced. Disease screening, fraud, equipment failure, students at risk of dropping out: the interesting class is always the rare one.

Measure it properly

The rest of this lesson is Pro

The free lessons of Machine Learning Foundations finish what they start — read those first if you have not. This one goes further, and it is part of the paid half.

A pass opens the paid lessons of every course, the mock test papers and the company-wise series. It ends on its own date; nothing renews by itself.

Get a pass — from ₹29 for 7 days

Already bought one? Log in.