FirstHack Learn
Log in Sign up free
Lessons in this course 0/6 All courses Machine Learning Foundations

AI & DS

Progress0 / 6 lessons
  1. 1. What learning from data means, and when ML is the wrong tool
  2. 2. Supervised vs unsupervised learning
  3. 3. Linear regression from scratch, then with scikit-learn
  4. 4. Classification, logistic regression and the confusion matrix
  5. 5. Overfitting, train/test split and cross-validation
  6. 6. Why accuracy is a bad metric on imbalanced data

Courses › Machine Learning Foundations

Supervised vs unsupervised learning

The split that decides which algorithm you can even consider.

10 min read · Lesson 2 of 6 · Free

The difference is the label

Supervised learning: your data has the right answers. Each row has features and a known label. The model learns the mapping from one to the other.

Unsupervised learning: there are no answers. You only have features, and you are looking for structure.

Everything else follows from this. If you have labels you can measure accuracy. If you do not, there is nothing to be accurate against, so "accuracy" is not defined and success is a judgement call.

Supervised: two kinds

Regression predicts a number.

  • Marks from hours studied and attendance
  • Electricity units for next month from the past twelve months
  • Time to complete a delivery from distance and time of day

Classification predicts a category.

  • Pass or fail from internal marks
  • Spam or not spam from the words in an email
  • Which of ten digits a handwritten image shows

The difference is only in the label type, and it changes the loss function and the metrics you use.

Python 3
import pandas as pd

df = pd.DataFrame({
    "hours":      [2, 5, 1, 8, 4, 7, 3, 6],
    "attendance": [55, 82, 40, 95, 70, 90, 60, 85],
    "marks":      [38, 71, 30, 88, 60, 84, 48, 76],
})

df["passed"] = (df["marks"] >= 50).astype(int)

X = df[["hours", "attendance"]]   # features
y_reg = df["marks"]               # regression target
y_clf = df["passed"]              # classification target

print(X.shape, y_reg.shape, y_clf.shape)

Same features, two different problems. Capital X for features and small y for the target is the convention in every ML library; follow it so your code reads like everyone else's.

Unsupervised: two kinds you will meet

Clustering groups similar rows.

Python 3
import numpy as np
from sklearn.cluster import KMeans

rng = np.random.default_rng(0)
group_a = rng.normal([30, 40], 4, size=(40, 2))
group_b = rng.normal([70, 80], 5, size=(40, 2))
X = np.vstack([group_a, group_b])

km = KMeans(n_clusters=2, n_init=10, random_state=0)
labels = km.fit_predict(X)

print(labels[:10])
print(km.cluster_centers_.round(1))

The centres come out near [30, 40] and [70, 80]. KMeans never saw a label. It found two dense regions and drew a boundary.

But notice what it cannot do: it cannot tell you what the clusters mean. It gives you cluster 0 and cluster 1. Deciding that cluster 0 is "students who need extra help" is your interpretation, and you can be wrong.

Dimensionality reduction compresses many columns into a few, mainly so you can plot the data or speed up a later model.

Python 3
from sklearn.decomposition import PCA

pca = PCA(n_components=2)
X_small = pca.fit_transform(X)
print(X_small.shape)
print(pca.explained_variance_ratio_.round(3))

Choosing k is a judgement, not a calculation

Python 3
from sklearn.cluster import KMeans

for k in range(1, 7):
    km = KMeans(n_clusters=k, n_init=10, random_state=0).fit(X)
    print(k, round(km.inertia_, 1))

inertia_ always falls as k rises. At k equal to the number of points it reaches zero, and that model is useless. People look for an "elbow" where the fall slows down. That elbow is often unclear, and two reasonable people pick different k values from the same plot.

⚠️

Unsupervised results are easy to over-interpret. KMeans will happily split pure random noise into five neat clusters and report cluster centres with one decimal place. The output always looks convincing. Before you believe a cluster is real, check whether it survives a different random seed, a different k, and a different scaling of your features.

Scaling is not optional for distance-based methods

Python 3
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
from sklearn.pipeline import make_pipeline

pipe = make_pipeline(StandardScaler(), KMeans(n_clusters=2, n_init=10, random_state=0))
labels = pipe.fit_predict(X)

KMeans measures distance. If one column is salary in rupees (range 300000) and another is years of experience (range 10), the salary column decides everything and experience is ignored. Scaling puts every feature on comparable footing. Linear regression and decision trees do not need this; KMeans, KNN and SVM do.

💡

Ask one question before choosing an approach: do I have the answers for past data? Yes means supervised, and you can measure how well you did. No means unsupervised, and you must justify your results by argument. Students often start with clustering because it needs no labels, then find they cannot show their model works.