FirstHack Learn
Log in Sign up free
Lessons in this course 0/6 All courses Machine Learning Foundations

AI & DS

Progress0 / 6 lessons
  1. 1. What learning from data means, and when ML is the wrong tool
  2. 2. Supervised vs unsupervised learning
  3. 3. Linear regression from scratch, then with scikit-learn
  4. 4. Classification, logistic regression and the confusion matrix
  5. 5. Overfitting, train/test split and cross-validation
  6. 6. Why accuracy is a bad metric on imbalanced data

Courses › Machine Learning Foundations

What learning from data means, and when ML is the wrong tool

The one idea behind every ML algorithm, and the cases where you should not use one.

11 min read · Lesson 1 of 6 · Free

Normal programming, reversed

Normal programming: you write the rules, the computer applies them to data, you get answers.

Python 3
def fee(units):
    if units <= 100:
        return units * 3
    return 100 * 3 + (units - 100) * 5

print(fee(150))   # 550

Machine learning: you supply data and the answers, and the computer works out the rules.

That is the entire idea. A model is a function with unknown numbers in it. Training means adjusting those numbers until the function's outputs match the answers you already have.

A model is a formula with knobs

The simplest useful model:

Code
predicted_marks = m * hours_studied + b

m and b are the knobs, called parameters. Training searches for the pair that makes the predictions closest to the real marks. Nothing more mystical than that is happening.

Bigger models have more knobs. A large language model has hundreds of billions of them. The principle does not change: adjust numbers until the error goes down.

The vocabulary, once

Word Means
Feature an input column: hours studied, attendance
Label / target the answer column: marks, pass or fail
Sample one row
Model the formula with parameters
Training choosing parameter values from data
Loss a number measuring how wrong the model is
Inference using the trained model on new data

The one assumption everything rests on

ML assumes the future looks like the past. A model trained on your college's 2023 results predicts 2024 results only if nothing important changed. Change the syllabus, change the marking scheme, change who gets admitted, and the model's accuracy quietly falls apart.

This is called distribution shift, and it is the most common reason a model that scored well in testing performs badly in use. The code does not break. It just becomes wrong.

When machine learning is the wrong tool

This section matters more than the rest of the lesson.

When a rule already exists. GST is 18 percent. Write tax = amount * 0.18. Training a model to learn it gives you something slower, less accurate, and impossible to audit.

When you have too little data. Forty rows will not support a model. As a rough guide you want at least a few hundred samples per class for a simple model, and far more as the number of features grows. With small data, descriptive statistics and a good chart are more honest and more useful.

When a wrong answer is expensive and you cannot explain the decision. Loan approval, medical triage, disciplinary action. A model that cannot say why it rejected someone is not acceptable, and in several countries it is not legal.

When your data encodes past unfairness. A model trained on past hiring decisions learns who was hired before, including who was passed over unfairly. The model then repeats it, and the output carries an air of mathematical authority that makes it harder to challenge.

When the goal is understanding, not prediction. If you want to know why attendance and marks move together, a controlled study answers that. A model with 94 percent accuracy does not.

⚠️

"Let us apply machine learning to it" is not a project plan. Before writing any model code, answer three questions in writing: what exactly is being predicted, what a wrong prediction costs, and what a simple rule or an average would score. If a three-line rule gets 88 percent and your model gets 89 percent, the rule wins. It is faster, explainable and it cannot break in strange ways.

Always build a baseline first

Python 3
import numpy as np

y_true = np.array([1, 0, 0, 1, 0, 0, 0, 1, 0, 0])

# Baseline 1: always predict the most common class
majority = np.bincount(y_true).argmax()
baseline = np.full_like(y_true, majority)
print("baseline accuracy:", (baseline == y_true).mean())   # 0.7

Seventy percent, from a model that does no thinking at all. Any real model must beat this clearly before it is worth anything. Students report 72 percent accuracy as a success without ever checking that guessing gets 70.

What this course covers

  • Supervised and unsupervised learning, with real examples
  • Linear regression built by hand, then with scikit-learn
  • Classification and the confusion matrix
  • Overfitting, train/test split and cross-validation
  • Why accuracy is the wrong metric more often than it is the right one
💡

Install what you need now: pip install numpy pandas scikit-learn matplotlib. Everything in this course runs on a normal laptop in a few seconds. You do not need a GPU, and you do not need cloud credits.