FirstHack Learn
Log in Sign up free
Lessons in this course 0/6 All courses Python for Data Analysis

AI & DS

Progress0 / 6 lessons
  1. 1. NumPy arrays and why they beat lists
  2. 2. pandas Series and DataFrame
  3. 3. Loading, cleaning and handling missing data honestly
  4. 4. Filtering, groupby and merge
  5. 5. Simple statistics and what they hide
  6. 6. Plotting basics, and how not to mislead with a chart

Courses › Python for Data Analysis

Simple statistics and what they hide

Mean, median, standard deviation and correlation, plus the traps in each one.

12 min read · Lesson 5 of 6 · Pro

Everything at once

Python 3
import pandas as pd

marks = pd.Series([78, 92, 55, 88, 41, 67, 73, 81, 60, 95])

print(marks.describe())
print("median:", marks.median())
print("mode  :", marks.mode().tolist())
print("range :", marks.max() - marks.min())

describe() gives count, mean, standard deviation, minimum, the three quartiles and the maximum. That is usually the right first command on any numeric column.

Mean and median disagree, and that is information

Python 3
import pandas as pd

salaries = pd.Series([3, 3.5, 4, 4, 4.5, 5, 5, 6, 6.5, 90])   # lakhs

print("mean  :", salaries.mean())     # 13.15
print("median:", salaries.median())   # 4.75
The rest of this lesson is Pro

The free lessons of Python for Data Analysis finish what they start — read those first if you have not. This one goes further, and it is part of the paid half.

A pass opens the paid lessons of every course, the mock test papers and the company-wise series. It ends on its own date; nothing renews by itself.

Get a pass — from ₹29 for 7 days

Already bought one? Log in.