The one sentence version
Data that came from outside your program is not data yet. It is a request to be data. Your job is to decide, on purpose, what you will accept.
Almost every security bug in this course is the same mistake wearing a different shirt: a value crossed a boundary and nobody checked it.
"Input" is bigger than you think
Students hear "user input" and picture a text box. The real list is longer.
- Form fields, query strings, URL path segments
- HTTP headers, cookies, the
User-Agent, theReferer - Uploaded file names and file contents
- JSON bodies from a mobile app
- Data from a third-party API
- Rows in your own database that a user put there yesterday
- Environment variables and command line arguments
- QR codes and barcodes
The most missed one is your own database. Developers say "it is our data, it is clean". It is not. If a user typed it, it is still user input every time you read it back out. Stored XSS lives entirely in this gap.
Validate on the server. Always.
Browser-side checks with required, maxlength and a JavaScript regex are good for the user. They are worth nothing for security, because the browser is fully under the attacker's control. Anyone can send a request straight to your endpoint without ever loading your page.
Keep the client checks for the friendly error message. Repeat every one of them on the server, where the decision actually counts.
Allow-list, not blocklist
A blocklist tries to name the bad things: strip <script>, remove DROP, reject quotes. It fails because the list of bad things is infinite and you are guessing at it. Every blocklist ever shipped has been incomplete.
An allow-list names the good things and rejects everything else. It works because you actually know what a valid pincode looks like.
Ask three questions of every field: what type is it, what range or length is allowed, and what format must it match?
import re
PINCODE = re.compile(r"^[1-9][0-9]{5}$")
def clean_pincode(raw):
if not isinstance(raw, str):
raise ValueError("pincode must be text")
value = raw.strip()
if not PINCODE.fullmatch(value):
raise ValueError("pincode must be 6 digits and not start with 0")
return value
print(clean_pincode(" 641004 "))
try:
clean_pincode("64100")
except ValueError as e:
print("rejected:", e)
That prints 641004 and then rejected: pincode must be 6 digits and not start with 0. Notice what it does not do: it does not try to "fix" a bad value. Rejecting is safer than repairing, because a repair can turn one invalid value into a different valid one you did not intend.
Convert types at the edge
A page number arrives as the text "3". Convert it once, at the boundary, and let the rest of your code work with a real integer.
def page_number(raw, max_page=500):
try:
n = int(raw)
except (TypeError, ValueError):
return 1
if n < 1:
return 1
if n > max_page:
return max_page
return n
print(page_number("3"), page_number("abc"), page_number("-7"), page_number("9999"))
Output: 3 1 1 500. Clamping is fine for a page number because every value in the range is harmless. Clamping would be wrong for a bank amount, where you must reject instead.
Canonicalise before you check
Check the value in one agreed form, or two spellings of the same thing will disagree — one passes your check, the other reaches the dangerous operation. File paths are the classic case.
import os
UPLOAD_DIR = os.path.realpath("/srv/app/uploads")
def safe_path(user_filename):
candidate = os.path.realpath(os.path.join(UPLOAD_DIR, user_filename))
if os.path.commonpath([UPLOAD_DIR, candidate]) != UPLOAD_DIR:
raise ValueError("path escapes the upload directory")
return candidate
print(safe_path("notes.pdf"))
try:
safe_path("../../etc/passwd")
except ValueError as e:
print("rejected:", e)
realpath resolves .. and symlinks first, and only then do we compare. Checking the raw string for .. before resolving is exactly the blocklist mistake, and it misses encodings and symlinks.
The strongest version of this idea is not to put user data in the path at all. Store the uploaded file under a generated name such as a UUID, and keep the original name in a database column for display only.
Validation is not encoding
Two different jobs, both needed.
Validation happens at the entrance and answers "is this acceptable at all?" Encoding happens at the exit and answers "how do I write this safely into HTML, into SQL, into a shell command?"
A person's name can legitimately contain an apostrophe. You must accept O'Brien as valid, and you must still write it into SQL and into HTML safely. Validation cannot save you at the exit, and encoding cannot save you at the entrance.
The next two lessons are about the exit.