FirstHack Learn
Log in Sign up free
Lessons in this course 0/6 All courses Machine Learning Foundations

AI & DS

Progress0 / 6 lessons
  1. 1. What learning from data means, and when ML is the wrong tool
  2. 2. Supervised vs unsupervised learning
  3. 3. Linear regression from scratch, then with scikit-learn
  4. 4. Classification, logistic regression and the confusion matrix
  5. 5. Overfitting, train/test split and cross-validation
  6. 6. Why accuracy is a bad metric on imbalanced data

Courses › Machine Learning Foundations

Overfitting, train/test split and cross-validation

Why a model that scores 100 percent on your data is usually the worst one.

13 min read · Lesson 5 of 6 · Pro

Memorising is not learning

A student who memorises last year's question paper scores full marks on last year's paper and fails a new one. Models do exactly this.

Overfitting means the model learned the noise in your training data, not the pattern. It looks excellent on the data it saw and poor on anything new.

Underfitting is the opposite: the model is too simple to capture the real pattern, so it does badly on both.

Watch it happen

The rest of this lesson is Pro

The free lessons of Machine Learning Foundations finish what they start — read those first if you have not. This one goes further, and it is part of the paid half.

A pass opens the paid lessons of every course, the mock test papers and the company-wise series. It ends on its own date; nothing renews by itself.

Get a pass — from ₹29 for 7 days

Already bought one? Log in.