URL encoding (percent-encoding): what it is and how to decode it

- URL encoding (percent-encoding) replaces 'unsafe' characters with a combination of the % sign and two hexadecimal digits, for example, a space becomes %20.
- You need to encode spaces, Cyrillic characters, special characters (?, &, #, =, /) and anything that can disrupt the structure of the address.
- Cyrillic is encoded using UTF-8 encoding: one Russian letter takes 2 bytes and turns into 6 characters like %D0%9F.
- Decoding is the reverse process: %20 becomes a space again, and %D0%9F becomes the letter 'П'.
You have probably seen addresses where instead of familiar letters, mysterious "%D0%BF%D1%80%D0%B8" and percent signs flash. This is not an error or a malfunction — this is how URL encoding works, a mechanism that allows spaces, Cyrillic letters, and special characters to be transmitted in web addresses without the risk of 'breaking' the link. In this article, we will explain in simple terms what percent-encoding is, why it is needed, which characters must be encoded, and how to perform the reverse transformation — decoding.
Encode and decode URLs online — for free
What is URL encoding
URL encoding, also known as percent-encoding, is a way to represent any character in a page address using a safe set of characters. The address standard (RFC 3986) allows only a limited alphabet to be used in URLs: Latin letters A–Z and a–z, digits 0–9, and a few symbols (- _ . ~). Everything else — spaces, Cyrillic, quotes, parentheses — is considered 'unsafe' and must be encoded.
It works like this: the program takes the byte of the character, converts it to a hexadecimal number, and places a percent sign before it. A space has the code 32, which is 20 in hexadecimal, so in the address, a space looks like %20. This is precisely why the method got its name.
Why encode a URL
The main reason is the reliability of data transmission. An address has a strict structure where certain characters act as separators: a question mark separates parameters, an ampersand separates key-value pairs, and a hash indicates an anchor. If a user enters a word with an ampersand or space in the website's search bar, and you insert it into a link 'as is', the browser or server will misinterpret the address.
Here are typical situations where encoding is mandatory:
- Spaces. There should be none in the address — they turn into %20 (or, in form parameters, into a + sign).
- Cyrillic and other national alphabets. Russian letters are not part of the basic ASCII set, so they are encoded byte by byte.
- Special character delimiters. If ? & = # / are part of the data itself and not the address structure, they need to be encoded; otherwise, the meaning of the link will be distorted.
- Punctuation marks and symbols. Quotes, brackets, percentages, plus signs — all of these are also encoded to keep the address unambiguous.
Which characters are encoded
It's easier to remember the opposite: only 'safe' characters are not encoded. These are Latin letters, digits, and the symbols - _ . ~ (they are called unreserved). Everything else is either reserved or unsafe and must be encoded if used as regular data. Below is a table of common examples.
| Character | Code | Explanation |
|---|---|---|
| space | %20 | in form parameters can be + |
| ! | %21 | exclamation mark |
| # | %23 | hash (anchor) |
| % | %25 | the percent sign itself |
| & | %26 | ampersand (parameter separator) |
| + | %2B | plus |
| / | %2F | slash |
| = | %3D | equal |
| ? | %3F | question mark |
| П | %D0%9F | Cyrillic, 2 bytes UTF-8 |
How Cyrillic is encoded
The most common confusion is related to Russian letters. Unlike the Latin alphabet, where one letter equals one byte, Cyrillic in UTF-8 encoding takes two bytes per character. Therefore, one letter turns into six characters: two bytes, each represented as %XX. For example, the word 'привет' is encoded as %D0%BF%D1%80%D0%B8%D0%B2%D0%B5%D1%82.
It is important that both encoding and decoding are performed in the same encoding — today this is almost always UTF-8. If a site uses a different encoding (for example, the outdated Windows-1251), the same letters will produce completely different sequences, and with incorrect decoding, instead of text, you will see 'garbage'. Therefore, the modern web is standardized on UTF-8.
How to decode a URL
Decoding is the reverse operation. The algorithm searches the string for each construct of the form %XX, converts the hexadecimal number back to a byte, and restores the original character. Thus, %20 becomes a space again, and the pair %D0%9F assembles into the letter 'П'. Doing this manually using a table is time-consuming, so ready-made tools are usually used.
Ways to decode (and encode) an address:
- Online tools — you paste the string, get the result in one click, no installation needed.
- JavaScript — the functions encodeURIComponent() and decodeURIComponent() encode and decode individual values; encodeURI() and decodeURI() work with the entire address.
- PHP — the functions urlencode()/urldecode() for form parameters and rawurlencode()/rawurldecode() for strict percent-encoding.
- Python — the urllib.parse module with the methods quote() and unquote().
Practical advice: encode only individual parameter values, not the entire address. Otherwise, control characters like ? and & will also be encoded, and the link will stop working. For an individual value in JavaScript, use encodeURIComponent(), and for assembling the entire address — use encodeURI().
Frequently Asked Questions
What does %20 mean in a link?
This is an encoded space. The code for a space is 32, which is 20 in hexadecimal, so it is represented in the address as %20. In web form parameters, a space is sometimes replaced by a plus sign (+).
Why do Russian letters in the URL look like %D0%BF?
Cyrillic is encoded according to the UTF-8 standard, where each letter takes up two bytes. Each byte is represented as a percent sign followed by two hexadecimal digits, so one Russian letter turns into six characters like %D0%BF.
Do I need to encode the entire address or just part of it?
Only values are encoded — for example, the text of a search query or a file name. Special characters that separate parts of the address (?, &, /, =) should remain unencoded, otherwise the browser will not understand the structure of the link.
What is the difference between encoding and decoding?
Encoding transforms ordinary characters into a safe representation like %XX, while decoding performs the reverse transformation — restoring spaces, letters, and signs from %XX. These are two sides of the same percent-encoding mechanism.
Encode and decode URLs online — for free
Updated: July 9, 2026. Edition of MegaInet.art.



