URL encode and decode
See the same text run through encodeURIComponent, encodeURI, form encoding (space as +) and strict RFC 3986 at once, decode with exact error positions, and check whether a string was encoded twice. It runs in your browser; nothing is uploaded.
Results
Each row is the same input run through a different encoder. Pick by what you are building, not by which looks shortest.
- encodeURIComponent (one value, one path segment)
- encodeURI (a whole URL; keeps ; / ? : @ & = + $ , #)
- Form encoding, application/x-www-form-urlencoded (space becomes +)
- Strict RFC 3986 (only A-Z a-z 0-9 - . _ ~ kept)
Decoded
- As a URL component or path (
+stays+) - As form data (
+becomes a space)
Double-encoding check
Looks like a query string: parameters
| Name | Value (form rules) |
|---|
Runs in your browser. Nothing is uploaded or stored. The share link puts your text after # in the address, which browsers do not send to servers; anyone you give the link to can read it, so do not share URLs that contain tokens.
How to use
- Choose Encode and type the text. Four results appear, one per method. Copy the one that fits:
encodeURIComponentfor a query value or path segment,encodeURIfor a complete URL, form encoding for an HTML form body or a?key=valuequery built the way browsers build it. - Choose Decode and paste percent-encoded text, a path or a whole query string. Both readings are shown: as a URL component (
+stays+) and as form data (+becomes a space). If they differ, use the one that matches where the text came from. - Read the Double-encoding check. It decodes up to five times and lists each step. A result of 2 or more layers with
%25XXin the text usually means something encoded already-encoded text. - If the input looks like a query string (
a=1&b=2), a table splits it into parameters using the WHATWG parser rules, which is how a browser'sURLSearchParamsreads it. - If decoding fails, the message names the position of the bad
%or the invalid UTF-8 byte. Use result as input chains a decode after an encode, which is a quick way to test a round trip.
Worked examples
Every row below comes from a data file that a unit test runs through the tool and, separately, through the browser's own encodeURIComponent, encodeURI and URLSearchParams. The expected strings were first worked out by hand from the byte values.
A space and a plus
Input: Tom & Jerry + friends
| Method | Output |
|---|---|
| encodeURIComponent | Tom%20%26%20Jerry%20%2B%20friends |
| encodeURI | Tom%20&%20Jerry%20+%20friends |
| Form (+) | Tom+%26+Jerry+%2B+friends |
| Strict RFC 3986 | Tom%20%26%20Jerry%20%2B%20friends |
encodeURI keeps & and + because they are reserved delimiters; encoding a value with it would let the & start a new parameter. Form encoding writes the space as + and the real plus as %2B.
Non-ASCII text
Input: café €5
| Method | Output |
|---|---|
| encodeURIComponent | caf%C3%A9%20%E2%82%AC5 |
| encodeURI | caf%C3%A9%20%E2%82%AC5 |
| Form (+) | caf%C3%A9+%E2%82%AC5 |
| Strict RFC 3986 | caf%C3%A9%20%E2%82%AC5 |
é is UTF-8 bytes C3 A9 and € is E2 82 AC, so each byte becomes one %XX.
The characters ! ' ( ) * ~
Input: wow!(it's*great~)
| Method | Output |
|---|---|
| encodeURIComponent | wow!(it's*great~) |
| encodeURI | wow!(it's*great~) |
| Form (+) | wow%21%28it%27s*great%7E%29 |
| Strict RFC 3986 | wow%21%28it%27s%2Agreat~%29 |
The three encoders disagree on these. ECMAScript leaves ! ' ( ) * ~ alone, the WHATWG form encoder escapes all but *, and the strict RFC 3986 column escapes the sub-delimiters ! ' ( ) * but keeps ~.
A whole URL placed inside a query parameter
Input: https://a.example/p?q=1&r=2
| Method | Output |
|---|---|
| encodeURIComponent | https%3A%2F%2Fa.example%2Fp%3Fq%3D1%26r%3D2 |
| encodeURI | https://a.example/p?q=1&r=2 |
| Form (+) | https%3A%2F%2Fa.example%2Fp%3Fq%3D1%26r%3D2 |
| Strict RFC 3986 | https%3A%2F%2Fa.example%2Fp%3Fq%3D1%26r%3D2 |
Use the component form for a value. encodeURI leaves the URL unchanged, so its & and = would be read as part of the outer URL.
RFC 3986 reserved characters and what each encoder does
RFC 3986 section 2.2 splits the characters into gen-delims (: / ? # [ ] @) and sub-delims (! $ & ' ( ) * + , ; =), and section 2.3 defines unreserved characters as letters, digits and - . _ ~. A reserved character used as data must be encoded or it can change what the URL means. The table is generated by the same code as the tool, and a test checks the class counts (7 gen-delims, 11 sub-delims).
| Char | Name | RFC 3986 class | encodeURIComponent | encodeURI | Form | Strict |
|---|---|---|---|---|---|---|
(space) | space | not allowed as data | %20 | %20 | + | %20 |
! | exclamation mark | reserved: sub-delim | kept | kept | %21 | %21 |
" | double quote | not allowed as data | %22 | %22 | %22 | %22 |
# | number sign | reserved: gen-delim | %23 | kept | %23 | %23 |
$ | dollar | reserved: sub-delim | %24 | kept | %24 | %24 |
% | percent | not allowed as data | %25 | %25 | %25 | %25 |
& | ampersand | reserved: sub-delim | %26 | kept | %26 | %26 |
' | apostrophe | reserved: sub-delim | kept | kept | %27 | %27 |
( | left parenthesis | reserved: sub-delim | kept | kept | %28 | %28 |
) | right parenthesis | reserved: sub-delim | kept | kept | %29 | %29 |
* | asterisk | reserved: sub-delim | kept | kept | kept | %2A |
+ | plus | reserved: sub-delim | %2B | kept | %2B | %2B |
, | comma | reserved: sub-delim | %2C | kept | %2C | %2C |
- | hyphen | unreserved | kept | kept | kept | kept |
. | full stop | unreserved | kept | kept | kept | kept |
/ | slash | reserved: gen-delim | %2F | kept | %2F | %2F |
: | colon | reserved: gen-delim | %3A | kept | %3A | %3A |
; | semicolon | reserved: sub-delim | %3B | kept | %3B | %3B |
< | less-than | not allowed as data | %3C | %3C | %3C | %3C |
= | equals | reserved: sub-delim | %3D | kept | %3D | %3D |
> | greater-than | not allowed as data | %3E | %3E | %3E | %3E |
? | question mark | reserved: gen-delim | %3F | kept | %3F | %3F |
@ | at sign | reserved: gen-delim | %40 | kept | %40 | %40 |
[ | left bracket | reserved: gen-delim | %5B | %5B | %5B | %5B |
\ | backslash | not allowed as data | %5C | %5C | %5C | %5C |
] | right bracket | reserved: gen-delim | %5D | %5D | %5D | %5D |
^ | caret | not allowed as data | %5E | %5E | %5E | %5E |
_ | underscore | unreserved | kept | kept | kept | kept |
` | backtick | not allowed as data | %60 | %60 | %60 | %60 |
{ | left brace | not allowed as data | %7B | %7B | %7B | %7B |
| | vertical bar | not allowed as data | %7C | %7C | %7C | %7C |
} | right brace | not allowed as data | %7D | %7D | %7D | %7D |
~ | tilde | unreserved | kept | kept | %7E | kept |
"kept" means the character is written as-is. Characters above U+007F are always written as the percent-escaped bytes of their UTF-8 form by all four encoders.
What goes wrong
Each case is a real input with the exact output of the tool (messages quoted verbatim).
A literal percent sign: "100%"
Input: 100%
Error: The "%" at position 4 is not followed by two hexadecimal digits (the text ends after "%").
decodeURIComponent("100%") throws URIError: URI malformed. The fix is to encode the percent sign as %25 when the text is produced.
A percent sign before letters: "50%off"
Input: 50%off
Error: The "%" at position 3 is followed by "of", which is not two hexadecimal digits.
"%of" is not an escape because "o" is not a hex digit. Native decodeURIComponent throws; the WHATWG form parser would keep "50%off" unchanged.
A Latin-1 escape in a UTF-8 world: "caf%E9"
Input: caf%E9
Error: The escape sequences starting at position 4 do not form valid UTF-8: the data ends inside a 3-byte sequence that starts with 0xE9.
E9 is é in Latin-1 but, in UTF-8, a lead byte that needs two more bytes. Old systems produced such links. No generic decoder can know the intended charset.
Encoded twice: "caf%25C3%25A9"
Input: caf%25C3%25A9
After the first decode: caf%C3%A9
One decode gives "caf%C3%A9" (what the page shows is still escaped); a second gives "café". The detector reports 2 layers.
A plus that means different things
Input: a+b%2Bc
Readings: URL rule: "a+b+c" Form rule: "a b+c"
In a path, + is a literal plus (RFC 3986 gives it no special meaning as data). In a query read as application/x-www-form-urlencoded it is a space. The same text, two answers.
encodeURI on a value
Input: a&b=c
encodeURI output: a&b=c
encodeURI("a&b=c") returns it unchanged, so the receiving server splits it into two parameters. encodeURIComponent gives "a%26b%3Dc".
Limits & gotchas
- Everything is UTF-8. The tool encodes text as UTF-8 and decodes assuming UTF-8, as RFC 3986 section 2.5 recommends for new URI schemes and as the WHATWG URL Standard requires. It does not decode other charsets such as Latin-1 or Shift_JIS. If an old system sends
%E9for é, the decoder reports invalid UTF-8 instead of guessing. - This is text-level encoding, not URL parsing. The tool does not split a URL into scheme, host, path and query and encode each part with its own rules. The WHATWG URL Standard uses different percent-encode sets for the path, query, fragment and userinfo (for example the path set also encodes
?,^,`,{and}), and hosts use IDNA/punycode instead of percent-encoding. For a whole URL, paste it into your own code'sURLclass. - Lone surrogates. A string containing half of a surrogate pair cannot be written as UTF-8. Native
encodeURIComponentthrowsURIError;URLSearchParamssilently substitutes U+FFFD. This tool stops and tells you which character. - Strict decoding versus the browser form parser. The decode result shown above is strict, like
decodeURIComponent: a bad%is an error. The WHATWG form parser instead keeps a bad%as a literal and replaces invalid UTF-8 with U+FFFD. The query table uses the lenient parser, so it can show values that the strict panel rejects. - The double-encoding check is a heuristic. It counts how often the text can be decoded in a row. It cannot know your intent: a tutorial that shows
%2520as an example is "double-encoded" on purpose. - Case of hex digits. Output uses uppercase (
%C3), as RFC 3986 section 2.1 recommends. Lowercase input decodes fine and is equivalent. - The share link carries your text after
#, up to about 3,000 characters. Do not share links that contain tokens, passwords or private URLs.
FAQ
Should I use encodeURI or encodeURIComponent?
Use encodeURIComponent for a single value or path segment that you insert into a URL, because it escapes the delimiters ? & = / # + that would otherwise change the URL's structure. Use encodeURI only on a complete URL that is already assembled, to escape spaces and non-ASCII characters while leaving ; / ? : @ & = + $ , # alone (ECMAScript 19.2.6). Applying encodeURI to a value is the classic bug: "a&b=c" comes out unchanged and the server sees two parameters.
Why is a space sometimes %20 and sometimes +?
They are two different formats. RFC 3986 percent-encoding writes a space as %20 and gives + no special meaning in data. The application/x-www-form-urlencoded format used by HTML forms (WHATWG URL Standard) writes a space as + and a real plus as %2B, and its parser turns + back into a space. Query strings are very often read with the form rules, so + in a query usually means a space, while + in a path means a plus.
What does "URI malformed" mean?
decodeURIComponent throws URIError: URI malformed when a % is not followed by two hex digits (for example "100%" or "50%off"), or when the escaped bytes are not valid UTF-8 (for example "%E9" from a Latin-1 system). The decoder above reports which position failed and why. The fix on the producing side is to encode % as %25 and to use UTF-8.
How do I know my text was encoded twice?
A tell-tale is %25 followed by two hex digits, such as %2520 or %25C3. %25 is an encoded percent sign, so "%2520" decodes to "%20", which decodes to a space. RFC 3986 section 2.4 says implementations must not percent-encode or decode the same string more than once. The detector above decodes repeatedly and reports how many layers it found, but it is a heuristic: a page that documents URL encoding can contain "%2520" on purpose.
Are ! ' ( ) * ~ safe to leave unencoded?
It depends on whose rule you follow. ECMAScript leaves ! ' ( ) * ~ alone in encodeURIComponent. RFC 3986 lists ! ' ( ) * as reserved sub-delimiters, not as unreserved, so a strict producer encodes them if they are data; only ~ is unreserved. The WHATWG form encoder escapes all of them except *. If a system is picky, the strict RFC 3986 row is the safest choice, and it decodes correctly everywhere.
Sources
- IETF: RFC 3986: Uniform Resource Identifier (URI): Generic Syntax Used for: Section 2.1 percent-encoding, 2.2 reserved characters (gen-delims and sub-delims), 2.3 unreserved characters, 2.4 when to encode and decode, 2.5 UTF-8 for new schemes, 3.4 query, 6.2.2.2 normalising percent-encoding.
- WHATWG: URL Standard: percent-encode sets and application/x-www-form-urlencoded Used for: The component, path, query and application/x-www-form-urlencoded percent-encode sets; the form serializer (space to +, UTF-8) and parser (+ to space, then percent-decode).
- Ecma International: ECMAScript Language Specification: URI Handling Functions (encodeURI, encodeURIComponent, decodeURI, decodeURIComponent) Used for: Which characters encodeURI and encodeURIComponent leave alone, and that lone surrogates throw URIError.
- IETF: RFC 3629: UTF-8, a transformation format of ISO 10646 Used for: Valid UTF-8 byte sequences: lead bytes, continuation bytes, overlong forms, surrogates and U+10FFFF are not allowed.
Every document above was opened and read on 2026-10-02. Documentation changes; if a page here disagrees with the current docs, trust the docs and tell us.