FirstHack Learn
Log in Sign up free
Lessons in this course 0/6 All courses Python for Data Analysis

AI & DS

Progress0 / 6 lessons
  1. 1. NumPy arrays and why they beat lists
  2. 2. pandas Series and DataFrame
  3. 3. Loading, cleaning and handling missing data honestly
  4. 4. Filtering, groupby and merge
  5. 5. Simple statistics and what they hide
  6. 6. Plotting basics, and how not to mislead with a chart

Courses › Python for Data Analysis

NumPy arrays and why they beat lists

The same sum, timed two ways, and the reason for the 50x gap.

11 min read · Lesson 1 of 6 · Free

Start with the measurement

Do not take anyone's word for it. Run this.

Python 3
import time
import numpy as np

n = 10_000_000
py_list = list(range(n))
np_array = np.arange(n)

start = time.perf_counter()
total_list = sum(py_list)
t_list = time.perf_counter() - start

start = time.perf_counter()
total_array = np_array.sum()
t_array = time.perf_counter() - start

print("list sum :", total_list, round(t_list, 4), "s")
print("numpy sum:", total_array, round(t_array, 4), "s")
print("numpy is", round(t_list / t_array, 1), "times faster")

On an ordinary laptop this prints something close to:

Code
list sum : 49999995000000 0.0598 s
numpy sum: 49999995000000 0.0075 s
numpy is 8.0 times faster

That is with sum(), which is itself written in C. Write a manual for loop over the list and the gap grows to 50x or more. Your exact numbers will differ. The order of magnitude will not.

Why the gap exists

A Python list of 10 million integers is 10 million separate objects scattered in memory, plus an array of 10 million pointers to them. Every addition means: follow a pointer, check the object's type, unbox the integer, add, box the result back into a new object.

A NumPy array is one continuous block of memory holding raw 8-byte integers. No pointers, no type checks, no boxing. The loop runs in compiled C, and the CPU can add several numbers per instruction.

Memory tells the same story:

Python 3
import sys
import numpy as np

py_list = list(range(1_000_000))
np_array = np.arange(1_000_000)

objects = sum(sys.getsizeof(v) for v in py_list)
print(sys.getsizeof(py_list) / 1e6, "MB for the list object alone")
print(objects / 1e6, "MB for the integer objects inside it")
print(np_array.nbytes / 1e6, "MB for the numpy array, total")

Both first numbers are about 8.0 MB, but they are not the same 8 MB. For the list that is only the table of pointers; the million integer objects it points at add about 28 MB more. For the array, 8.0 MB is the whole thing. Roughly 36 MB against 8 MB for the same numbers.

Creating arrays

Python 3
import numpy as np

a = np.array([2, 4, 6, 8])
b = np.zeros(5)
c = np.ones((2, 3))
d = np.arange(0, 10, 2)          # [0 2 4 6 8]
e = np.linspace(0, 1, 5)         # [0.   0.25 0.5  0.75 1.  ]

print(a.shape, a.dtype)          # (4,) int64
print(c.shape)                   # (2, 3)

Two attributes matter constantly. shape is the size in each dimension. dtype is the element type, decided once for the whole array.

Vectorised operations

This is the habit change. You stop writing loops.

Python 3
import numpy as np

marks = np.array([78, 92, 55, 88, 41])

print(marks + 5)          # [83 97 60 93 46]
print(marks * 2)          # [156 184 110 176  82]
print(marks / 100)        # [0.78 0.92 0.55 0.88 0.41]
print(marks > 60)         # [ True  True False  True False]
print(marks[marks > 60])  # [78 92 88]
print(marks.mean())       # 70.8

marks > 60 gives an array of True and False. Using that inside [ ] is called boolean masking, and it is how you filter data in NumPy and pandas.

Two arrays combine element by element:

Python 3
internal = np.array([18, 20, 15, 19, 12])
external = np.array([60, 72, 40, 69, 29])
total = internal + external
print(total)                        # [ 78  92  55  88  41]
print(np.round(total / 100 * 100))  # percentage, same numbers here

Two dimensions

Python 3
import numpy as np

# rows = students, columns = subjects
marks = np.array([
    [78, 65, 90],
    [55, 72, 60],
    [88, 91, 84],
])

print(marks.shape)          # (3, 3)
print(marks[1, 2])          # 60, row 1 column 2
print(marks[0])             # first student's marks
print(marks[:, 1])          # everyone's second subject
print(marks.mean(axis=0))   # average per subject
print(marks.mean(axis=1))   # average per student

axis=0 collapses rows, giving one number per column. axis=1 collapses columns, giving one number per row.

⚠️

axis confuses almost everyone at first, including people who have used pandas for a year. Remember it as "the axis that disappears". A (3, 3) array with axis=0 gives a result of shape (3,) with the rows gone. If your averages look wrong, print .shape before and after. That check takes two seconds and settles the argument.

Views, not copies

Python 3
import numpy as np

a = np.array([1, 2, 3, 4, 5])
b = a[1:4]
b[0] = 99
print(a)      # [ 1 99  3  4  5]

Slicing a NumPy array gives a view into the same memory, not a copy. This makes slicing free even on huge arrays, but it means editing the slice edits the original. Python lists behave the opposite way. When you need a separate copy, ask for one with a[1:4].copy().

💡

Whenever you catch yourself writing for i in range(len(arr)) with NumPy, stop. There is almost always a vectorised version that is shorter, faster and less likely to have an off-by-one bug.

Next: pandas, which puts labels and mixed types on top of these arrays.