Copy a link to a search results page and you may get something like https://example.com/search?q=caf%C3%A9%20%26%20bakery. Those percent signs and hex digits are URL encoding, also called percent-encoding. It is how URLs carry spaces, symbols and non-English characters without breaking. Knowing how it works helps you build correct links, debug broken API calls, and avoid some subtle security bugs.
Why URLs need encoding
The URL standard, RFC 3986, allows only a limited set of ASCII characters in a URL, and several of those have special jobs:
/separates path segments.?starts the query string.&separates query parameters, and=separates a name from its value.#starts the fragment (the part the browser keeps to itself).:and@appear in the scheme and in user information.
So what if a value itself contains one of these? A search for fish & chips placed directly into ?q=fish & chips would be read as a parameter q=fish followed by a separate, nameless parameter chips. Spaces, quotes, non-ASCII letters and control characters are not allowed at all. Percent-encoding solves both problems.
How percent-encoding works
Each byte that needs encoding is written as a percent sign followed by its value in two hexadecimal digits:
- Convert the text to bytes using UTF-8.
- Leave "unreserved" characters as they are.
- Replace every other byte with
%plus its hex value.
The unreserved characters, which never need encoding, are A–Z, a–z, 0–9, and - . _ ~.
Some common encodings:
| Character | Encoded | Why it matters |
|---|---|---|
| space | %20 (or + in form data) | Spaces are not allowed in URLs |
& | %26 | Would start a new parameter |
= | %3D | Would split name and value |
? | %3F | Would start the query string |
# | %23 | Would cut off the rest of the URL as a fragment |
/ | %2F | Would create a new path segment |
+ | %2B | May be read as a space in query strings |
% | %25 | The escape character itself |
é | %C3%A9 | Non-ASCII: two UTF-8 bytes, each encoded |
The é example shows the UTF-8 step. The character is stored as two bytes, C3 and A9, so it becomes two percent-encoded triplets. Characters in Hindi, Chinese and other scripts typically take three bytes each, so a short word can become a long encoded string.
The plus sign versus %20
This is the most common source of confusion. When an HTML form is submitted with the default application/x-www-form-urlencoded format, spaces are encoded as +. In the path part of a URL, a + is just a literal plus sign, and spaces must be %20.
The practical consequences:
- A literal plus in a query value, such as a phone number
+91…or an email aliasname+tag@example.com, must be sent as%2B. Otherwise the server decodes it as a space. - When unsure,
%20for spaces is understood correctly everywhere.
Encode components, not whole URLs
You should encode each value you insert into a URL, not the finished URL. Encoding a whole URL would turn its own /, ? and & into %2F, %3F and %26 and break it. The right functions in common languages:
// JavaScript
const url = 'https://example.com/search?q=' + encodeURIComponent('fish & chips');
// → https://example.com/search?q=fish%20%26%20chips
// Better: let the URL API handle it
const u = new URL('https://example.com/search');
u.searchParams.set('q', 'fish & chips'); // encodes as q=fish+%26+chips
In JavaScript, encodeURIComponent() is for individual values; encodeURI() leaves reserved characters intact and is only for tidying an already-structured URL.
// PHP
$q = rawurlencode('fish & chips'); // fish%20%26%20chips
$qs = http_build_query(['q' => 'fish & chips']); // q=fish+%26+chips
# Python
from urllib.parse import quote, urlencode
quote('fish & chips', safe='') # fish%20%26%20chips
urlencode({'q': 'fish & chips'}) # q=fish+%26+chips
On the receiving side, web frameworks decode query parameters for you. Decoding them a second time is a bug.
Double encoding and other pitfalls
- Double encoding. Encoding a value that is already encoded turns
%20into%2520(because%becomes%25). If you see%25sequences in URLs or logs, something has encoded twice. - Security filters. Attackers use encoding, and double encoding, to sneak characters like
../or<script>past naive filters that check the raw string. Validate input after decoding, exactly once, and rely on proper output encoding for HTML. - Encoding is not escaping for HTML. A percent-encoded value placed in a web page still needs HTML escaping where it appears in markup; different contexts need different encoding.
- Case and normalization.
%2fand%2Fare equivalent, but some caches and signature schemes compare URLs as raw strings. Use uppercase hex consistently. - Internationalized domain names are not percent-encoded. A domain like
café.exampleis converted to an ASCII form called Punycode (xn--caf-dma.example) instead.
Tools for quick checks
When debugging, compare exactly what is sent with what the server receives. Your browser's developer tools (Network tab) show the encoded request URL. To see redirects and the exact Location URLs a server returns, use our HTTP header checker. For other encodings you will meet alongside percent-encoding, such as Base64 tokens in query strings, our Base64 encoder and decoder helps, and the full set of utilities is on the tools page.
Key takeaways
- URL encoding replaces unsafe bytes with
%plus two hex digits, after converting text to UTF-8. - Only letters, digits and
- . _ ~never need encoding. - Spaces are
+in form data but%20elsewhere; send a literal plus as%2B. - Encode individual values with the right library function, and decode exactly once.