The same bug in a different place
In the last lesson a value became SQL because it was pasted into SQL text. Cross-site scripting is that story again, with HTML as the language and the victim's browser as the interpreter.
A comment, a name, a search term goes into a page. The browser parses the whole page as HTML. If the value contains HTML syntax, it becomes part of the document instead of remaining text.
Why it matters more than it looks: the injected code runs inside your origin, with your user's session. It can read the page, change the page, and act as the user.
Three shapes to recognise
Stored XSS. The value is saved — a profile bio, a comment — and served to everyone who opens the page. The most damaging kind.
Reflected XSS. The value comes in on the URL and is echoed straight back in the response, so the victim has to open a crafted link.
DOM-based XSS. The server sends clean HTML, and JavaScript on the page puts untrusted text into innerHTML or document.write. The server logs look perfectly normal.
Escaping on input is the wrong instinct, and it is the one most students have. If you escape when saving, your database fills with & and ', values look wrong in reports and exports, and the day someone writes a second view that skips the escaping the bug is back. Store the real text. Encode at the moment of output, every time.
The fix: encode for the context you are writing into
import html
def render_comment(author, text):
return "<p><b>{}</b>: {}</p>".format(
html.escape(author), html.escape(text)
)
print(render_comment("Anu", 'Nice notes on "trees" & graphs'))
print(render_comment("Bharat", "<b>hi</b>"))
The first line prints <p><b>Anu</b>: Nice notes on "trees" & graphs</p> and the second prints <p><b>Bharat</b>: <b>hi</b></p>.
The second line is the point. The characters are still there, the user sees <b>hi</b> on screen as text, and the browser's parser never sees a tag. html.escape converts &, <, > and, with quote=True which is the default, " and '.
Context decides the encoding
HTML escaping is correct for text between tags and for quoted attribute values. It is not correct everywhere.
| Where the value lands | What to do |
|---|---|
Between tags, e.g. inside a p |
HTML-escape |
| Inside a quoted attribute | HTML-escape, and always keep the quotes |
| In a URL or query string | Percent-encode with urllib.parse.quote |
| Inside a JavaScript block | Do not build it; pass it as JSON data |
| Inside a CSS block | Do not put user data there at all |
An unquoted attribute is a hole even after HTML escaping, because a space alone can start a new attribute. Quote every attribute.
import html
from urllib.parse import quote
def profile_link(username):
return '<a href="/user/{}" title="{}">{}</a>'.format(
quote(username, safe=""), html.escape(username), html.escape(username)
)
print(profile_link("anu kumar"))
Prints <a href="/user/anu%20kumar" title="anu kumar">anu kumar</a>. Same value, two different encodings, because it lands in two different places.
For JavaScript, do not paste text into a script at all. Put it in a data attribute and read it back.
import html, json
def page_with_config(user_name):
config = json.dumps({"name": user_name})
return (
'<div id="app" data-config="{}"></div>'.format(html.escape(config))
+ '<script>const cfg = JSON.parse('
+ 'document.getElementById("app").dataset.config);</script>'
)
print(page_with_config('Anu "the topper"'))
Let the template engine do it
You should rarely be writing html.escape by hand. Every serious template engine escapes by default. Jinja2 (used by Flask) escapes when autoescape is on; Django templates escape always.
from jinja2 import Environment, select_autoescape
env = Environment(autoescape=select_autoescape(["html"]))
tpl = env.from_string("<p>Hello {{ name }}</p>")
print(tpl.render(name="<b>Anu</b>"))
Prints <p>Hello <b>Anu</b></p>.
The danger in every template engine is the "trust me" switch — |safe in Jinja, {% autoescape off %} in Django, dangerouslySetInnerHTML in React, v-html in Vue. Each one turns the protection off for that value. Search your codebase for them and justify every single use.
If you genuinely must accept rich text from users, do not write your own filter. Run it through a maintained sanitiser such as bleach or nh3 with a small allow-list of tags and attributes.
Two layers behind the encoding
Content-Security-Policy. A response header that tells the browser which sources of script it may execute. A policy without unsafe-inline stops injected inline script even if a bug slips through. It is a seatbelt, not a substitute for encoding.
HttpOnly cookies. Injected JavaScript cannot read a cookie marked HttpOnly, so a session cookie cannot simply be copied out.
Review checklist
- Store raw text; encode at output.
- Match the encoding to the context; quote every attribute.
- Autoescape on, and every
|safejustified in writing. - No user data reaching
innerHTML. - CSP header set, HttpOnly on session cookies.