Normal programming, reversed
Normal programming: you write the rules, the computer applies them to data, you get answers.
def fee(units):
if units <= 100:
return units * 3
return 100 * 3 + (units - 100) * 5
print(fee(150)) # 550
Machine learning: you supply data and the answers, and the computer works out the rules.
That is the entire idea. A model is a function with unknown numbers in it. Training means adjusting those numbers until the function's outputs match the answers you already have.
A model is a formula with knobs
The simplest useful model:
predicted_marks = m * hours_studied + b
m and b are the knobs, called parameters. Training searches for the pair that makes the predictions closest to the real marks. Nothing more mystical than that is happening.
Bigger models have more knobs. A large language model has hundreds of billions of them. The principle does not change: adjust numbers until the error goes down.
The vocabulary, once
| Word | Means |
|---|---|
| Feature | an input column: hours studied, attendance |
| Label / target | the answer column: marks, pass or fail |
| Sample | one row |
| Model | the formula with parameters |
| Training | choosing parameter values from data |
| Loss | a number measuring how wrong the model is |
| Inference | using the trained model on new data |
The one assumption everything rests on
ML assumes the future looks like the past. A model trained on your college's 2023 results predicts 2024 results only if nothing important changed. Change the syllabus, change the marking scheme, change who gets admitted, and the model's accuracy quietly falls apart.
This is called distribution shift, and it is the most common reason a model that scored well in testing performs badly in use. The code does not break. It just becomes wrong.
When machine learning is the wrong tool
This section matters more than the rest of the lesson.
When a rule already exists. GST is 18 percent. Write tax = amount * 0.18. Training a model to learn it gives you something slower, less accurate, and impossible to audit.
When you have too little data. Forty rows will not support a model. As a rough guide you want at least a few hundred samples per class for a simple model, and far more as the number of features grows. With small data, descriptive statistics and a good chart are more honest and more useful.
When a wrong answer is expensive and you cannot explain the decision. Loan approval, medical triage, disciplinary action. A model that cannot say why it rejected someone is not acceptable, and in several countries it is not legal.
When your data encodes past unfairness. A model trained on past hiring decisions learns who was hired before, including who was passed over unfairly. The model then repeats it, and the output carries an air of mathematical authority that makes it harder to challenge.
When the goal is understanding, not prediction. If you want to know why attendance and marks move together, a controlled study answers that. A model with 94 percent accuracy does not.
"Let us apply machine learning to it" is not a project plan. Before writing any model code, answer three questions in writing: what exactly is being predicted, what a wrong prediction costs, and what a simple rule or an average would score. If a three-line rule gets 88 percent and your model gets 89 percent, the rule wins. It is faster, explainable and it cannot break in strange ways.
Always build a baseline first
import numpy as np
y_true = np.array([1, 0, 0, 1, 0, 0, 0, 1, 0, 0])
# Baseline 1: always predict the most common class
majority = np.bincount(y_true).argmax()
baseline = np.full_like(y_true, majority)
print("baseline accuracy:", (baseline == y_true).mean()) # 0.7
Seventy percent, from a model that does no thinking at all. Any real model must beat this clearly before it is worth anything. Students report 72 percent accuracy as a success without ever checking that guessing gets 70.
What this course covers
- Supervised and unsupervised learning, with real examples
- Linear regression built by hand, then with scikit-learn
- Classification and the confusion matrix
- Overfitting, train/test split and cross-validation
- Why accuracy is the wrong metric more often than it is the right one
Install what you need now: pip install numpy pandas scikit-learn matplotlib. Everything in this course runs on a normal laptop in a few seconds. You do not need a GPU, and you do not need cloud credits.