All Articles
Base64 & Encoding·5 min read·October 2, 2026

Handling UTF-8 & Unicode in Base64: Avoiding Character Corruption

How to properly encode Arabic, Chinese, Urdu, Hindi, Cyrillic, and emojis in Base64 without Mojibake or decoding errors.

TBy Toolstack Engineering
Recommended Free Tool

Try Toolstack's Base64 Encoder

Free, instantaneous, and processes 100% locally in your browser.

Open Base64 Encoder

A frequent bug in legacy JavaScript applications occurs when converting text with non-ASCII characters (such as Arabic, Urdu, Japanese, or emoji like 🚀) into Base64 using naive `btoa()` calls.

Why Naive btoa() Fails

JavaScript strings are stored internally as UTF-16 code units. The browser `btoa()` function expects each character code to fit within a single 8-bit byte (0x00 to 0xFF). When encountering a character like `🚀` (code point U+1F680), `btoa()` immediately throws an uncaught error.

The Correct Modern Solution: TextEncoder & TextDecoder

The modern W3C standard approach converts the JavaScript string into a strict UTF-8 `Uint8Array` first: ```javascript // Safe Encoding const utf8Bytes = new TextEncoder().encode("Hello 🌍 مرحبا"); const base64 = btoa(String.fromCharCode(...utf8Bytes)); // Safe Decoding const rawBytes = Uint8Array.from(atob(base64), c => c.charCodeAt(0)); const originalString = new TextDecoder().decode(rawBytes); ```

Toolstack's Base64 tools implement this standard natively, guaranteeing 100% accurate translation across all global languages.

Tags:#Unicode#UTF-8#Internationalization#JavaScript

More Guides in Base64 & Encoding