Why pandas exists
NumPy is great when everything is numbers of the same type. Real data is not. A student record has a name (text), a roll number (text or int), marks (int), and a fee-paid flag (boolean). Columns need names. Rows need labels.
pandas adds exactly that on top of NumPy.
Series: one labelled column
import pandas as pd
marks = pd.Series([78, 92, 55, 88], index=["Ravi", "Sneha", "Arjun", "Meera"])
print(marks)
print(marks["Sneha"]) # 92
print(marks.mean()) # 78.25
print(marks[marks > 70]) # Ravi 78, Sneha 92, Meera 88
A Series is a NumPy array plus an index. The index is what makes lookup by name possible, and it is what pandas uses to line up data when you combine two objects.
DataFrame: a labelled table
import pandas as pd
df = pd.DataFrame({
"name": ["Ravi", "Sneha", "Arjun", "Meera", "Imran"],
"branch": ["CSE", "CSE", "ECE", "ECE", "MECH"],
"sem": [3, 3, 5, 5, 3],
"marks": [78, 92, 55, 88, 41],
})
print(df)
print(df.shape) # (5, 4)
print(df.columns)
print(df.dtypes)
Each column is a Series. All columns share one index. That is the whole idea.
The first five things to run on any dataset
print(df.head()) # first 5 rows
print(df.tail(3)) # last 3 rows
print(df.shape) # (rows, columns)
print(df.info()) # column names, non-null counts, types
print(df.describe()) # count, mean, std, min, quartiles, max
Run all five before you write a single line of analysis. info() tells you where the missing values are. describe() shows a minimum age of -3 or a maximum mark of 1200 immediately, and those two numbers change what you do next.
Selecting
print(df["marks"]) # one column, a Series
print(df[["name", "marks"]]) # two columns, a DataFrame
print(df.loc[2]) # row with index label 2
print(df.loc[2, "name"]) # Arjun
print(df.iloc[0]) # first row by position
print(df.iloc[0:3, 0:2]) # first 3 rows, first 2 columns
loc uses labels. iloc uses positions. They look similar and behave differently, and mixing them up is the most common pandas mistake.
One difference that catches people: df.loc[1:3] includes row 3, because label slicing is inclusive at both ends. df.iloc[1:3] stops before position 3, like every other Python slice.
Filtering
print(df[df["marks"] >= 60])
print(df[(df["branch"] == "CSE") & (df["sem"] == 3)])
print(df[df["branch"].isin(["CSE", "ECE"])])
Use & and |, not and and or. Python's and asks one yes-or-no question, but here you have five rows and five answers, so pandas raises ValueError: The truth value of a Series is ambiguous. Also wrap each condition in brackets: & binds tighter than ==, so df["sem"] == 3 & df["marks"] > 50 is parsed wrongly and fails.
Adding and changing columns
df["percent"] = df["marks"] / 100 * 100
df["passed"] = df["marks"] >= 50
df["grade"] = pd.cut(
df["marks"],
bins=[0, 50, 60, 75, 90, 100],
labels=["F", "D", "C", "B", "A"],
)
print(df)
No loop. The operation applies to the whole column at once, exactly like NumPy.
Sorting and counting
print(df.sort_values("marks", ascending=False))
print(df.sort_values(["branch", "marks"]))
print(df["branch"].value_counts())
print(df["branch"].nunique()) # 3
value_counts() is the fastest way to understand a text column. Run it on every categorical column when you meet a new dataset; it exposes typos like "CSE" and "cse" sitting side by side.
Most pandas methods return a new object and leave the original alone. df.sort_values("marks") does not sort df. Either assign the result, df = df.sort_values("marks"), or pass inplace=True. Students spend whole evenings wondering why their sort did nothing.
Rows and columns are cheap to reshape
df = df.rename(columns={"sem": "semester"})
df = df.drop(columns=["percent"])
df = df.reset_index(drop=True)
Keep a small DataFrame of five rows open while you learn. Print after every operation. Five rows you can check by eye teach you more in an hour than a 50,000-row file where you cannot tell right from wrong.
Next: real files, which are never this tidy.