FirstHack Learn
Log in Sign up free
Lessons in this course 0/6 All courses Python for Data Analysis

AI & DS

Progress0 / 6 lessons
  1. 1. NumPy arrays and why they beat lists
  2. 2. pandas Series and DataFrame
  3. 3. Loading, cleaning and handling missing data honestly
  4. 4. Filtering, groupby and merge
  5. 5. Simple statistics and what they hide
  6. 6. Plotting basics, and how not to mislead with a chart

Courses › Python for Data Analysis

pandas Series and DataFrame

A labelled column and a labelled table, and the handful of methods you will use all day.

11 min read · Lesson 2 of 6 · Free

Why pandas exists

NumPy is great when everything is numbers of the same type. Real data is not. A student record has a name (text), a roll number (text or int), marks (int), and a fee-paid flag (boolean). Columns need names. Rows need labels.

pandas adds exactly that on top of NumPy.

Series: one labelled column

Python 3
import pandas as pd

marks = pd.Series([78, 92, 55, 88], index=["Ravi", "Sneha", "Arjun", "Meera"])
print(marks)
print(marks["Sneha"])     # 92
print(marks.mean())       # 78.25
print(marks[marks > 70])  # Ravi 78, Sneha 92, Meera 88

A Series is a NumPy array plus an index. The index is what makes lookup by name possible, and it is what pandas uses to line up data when you combine two objects.

DataFrame: a labelled table

Python 3
import pandas as pd

df = pd.DataFrame({
    "name":   ["Ravi", "Sneha", "Arjun", "Meera", "Imran"],
    "branch": ["CSE", "CSE", "ECE", "ECE", "MECH"],
    "sem":    [3, 3, 5, 5, 3],
    "marks":  [78, 92, 55, 88, 41],
})

print(df)
print(df.shape)      # (5, 4)
print(df.columns)
print(df.dtypes)

Each column is a Series. All columns share one index. That is the whole idea.

The first five things to run on any dataset

Python 3
print(df.head())        # first 5 rows
print(df.tail(3))       # last 3 rows
print(df.shape)         # (rows, columns)
print(df.info())        # column names, non-null counts, types
print(df.describe())    # count, mean, std, min, quartiles, max

Run all five before you write a single line of analysis. info() tells you where the missing values are. describe() shows a minimum age of -3 or a maximum mark of 1200 immediately, and those two numbers change what you do next.

Selecting

Python 3
print(df["marks"])                 # one column, a Series
print(df[["name", "marks"]])       # two columns, a DataFrame
print(df.loc[2])                   # row with index label 2
print(df.loc[2, "name"])           # Arjun
print(df.iloc[0])                  # first row by position
print(df.iloc[0:3, 0:2])           # first 3 rows, first 2 columns

loc uses labels. iloc uses positions. They look similar and behave differently, and mixing them up is the most common pandas mistake.

One difference that catches people: df.loc[1:3] includes row 3, because label slicing is inclusive at both ends. df.iloc[1:3] stops before position 3, like every other Python slice.

Filtering

Python 3
print(df[df["marks"] >= 60])
print(df[(df["branch"] == "CSE") & (df["sem"] == 3)])
print(df[df["branch"].isin(["CSE", "ECE"])])
⚠️

Use & and |, not and and or. Python's and asks one yes-or-no question, but here you have five rows and five answers, so pandas raises ValueError: The truth value of a Series is ambiguous. Also wrap each condition in brackets: & binds tighter than ==, so df["sem"] == 3 & df["marks"] > 50 is parsed wrongly and fails.

Adding and changing columns

Python 3
df["percent"] = df["marks"] / 100 * 100
df["passed"] = df["marks"] >= 50
df["grade"] = pd.cut(
    df["marks"],
    bins=[0, 50, 60, 75, 90, 100],
    labels=["F", "D", "C", "B", "A"],
)
print(df)

No loop. The operation applies to the whole column at once, exactly like NumPy.

Sorting and counting

Python 3
print(df.sort_values("marks", ascending=False))
print(df.sort_values(["branch", "marks"]))
print(df["branch"].value_counts())
print(df["branch"].nunique())    # 3

value_counts() is the fastest way to understand a text column. Run it on every categorical column when you meet a new dataset; it exposes typos like "CSE" and "cse" sitting side by side.

⚠️

Most pandas methods return a new object and leave the original alone. df.sort_values("marks") does not sort df. Either assign the result, df = df.sort_values("marks"), or pass inplace=True. Students spend whole evenings wondering why their sort did nothing.

Rows and columns are cheap to reshape

Python 3
df = df.rename(columns={"sem": "semester"})
df = df.drop(columns=["percent"])
df = df.reset_index(drop=True)
💡

Keep a small DataFrame of five rows open while you learn. Print after every operation. Five rows you can check by eye teach you more in an hour than a 50,000-row file where you cannot tell right from wrong.

Next: real files, which are never this tidy.