Why URLs Have All Those %20s
September 29, 2026 · 5 min read
Copy a link to a file named my report (final).pdf and paste it somewhere, and you'll often get back my%20report%20(final).pdf — or worse, my%2520report. Search Google for anything in Chinese and the URL bar fills with %E4%B8%AD%E6%96%87 soup. This is URL encoding (properly: percent-encoding), and it's one of those invisible mechanisms holding the web together. Here's what's going on.
URLs only speak a limited alphabet
The original URL spec was written for ASCII, and even within ASCII, a bunch of characters are "reserved" because they have jobs: ? starts the query string, & separates parameters, = separates keys from values, # starts a fragment, / separates path segments, and spaces… spaces just break things.
So what happens when your data contains one of these characters? If you put fish & chips raw into a query parameter — ?q=fish & chips — the server reads q=fish and then a mysterious second parameter chips. Your search is broken. The fix: replace each problematic character with a % followed by its two-digit hex code. Space is byte 0x20, so it becomes %20. & is 0x26, so %26. Now ?q=fish%20%26%20chips arrives intact, and the server decodes it back.
Non-English text: UTF-8 first, then encode
What about Chinese, emoji, Arabic? The process is two steps: first convert the character to its UTF-8 bytes, then percent-encode each byte. The character 中 is three bytes in UTF-8 — E4 B8 AD — so it becomes %E4%B8%AD. An emoji like 🎉 is four bytes, producing four %XX groups. That's why non-English URLs look so long: every character costs 9 characters of encoding.
Modern browsers hide this from you — the address bar shows 中文 while the actual URL underneath is percent-encoded. Copy the URL and paste it into a text editor, and you'll see the raw %E4%B8%AD form. Both are the same URL.
What gets encoded, and what doesn't
Not everything needs encoding. Letters, digits, and - _ . ~ are always safe. Everything else is context-dependent: / is fine in a path (it's a separator) but must be encoded inside a query value. This is where bugs live.
The classic JavaScript trap: encodeURI() vs encodeURIComponent(). The first encodes a whole URL and deliberately leaves ? & = # / alone — because in a full URL, those are structural. The second encodes a single value and encodes everything except the always-safe characters. Use encodeURIComponent for query parameter values, always. Using encodeURI on a value containing & is how you get the "my search term got split into two parameters" bug.
Double encoding: the %2520 horror
See %2520? That's a space that got encoded twice: space → %20 → the % itself got encoded → %2520. It happens when one layer of your stack encodes and another layer encodes again — a proxy, a redirect, a framework "helpfully" encoding an already-encoded URL. The symptom is unmistakable once you know the pattern: %25 is an encoded percent sign, so %25XX always means "something got encoded twice." Fix it by finding which layer is double-dipping, not by decoding twice (that just masks it until the next redirect).
When do you need to care?
Building URLs by string concatenation in code — encode the values. Accepting filenames or search terms from users and putting them in links — encode. Debugging a broken API call where parameters look right but the server sees garbage — check the encoding. Most HTTP libraries encode query params automatically these days, but the moment you hand-roll a URL, it's on you.