josiete.com

When bytes masquerade as text: Unicode, Python, and encoding errors

Streams of green characters on a black background, with the words decode and encode highlighted in the center

A file contains José, but after passing through a pipeline it becomes José. Another preserves accents, yet searching for Jose also returns José. In a third, emojis turn into ?.

All three look like text problems, but they call for different investigations: how we interpret bytes, how we compare strings, and what information was lost along the way. Switching to UTF-8 without locating the failure can simply store already damaged text correctly.

What is an encoding, and why do we need one?

A file does not store a drawing of ñ: it stores bytes. Turning those numbers into text requires an agreement about which sequences represent each character. That agreement is a character encoding.

Think of it as the rules for writing and reading a message. Sender and receiver must share them: receiving identical bytes does not guarantee identical text when each side uses different rules.

Encodings arose from this need to represent and transmit text using machines. Available space, languages, and compatibility with existing equipment shaped the solutions. The Unicode character encoding model distinguishes the character repertoire, its codes, and the ways of mapping them to storage units.

A reproducible Python example shows why bytes need context:

print("ñ".encode("utf-8").hex())     # c3b1
print("ñ".encode("latin-1").hex())   # f1

Both sequences represent the same letter under their respective encodings. Neither carries a label that, by itself, tells us how to read it.

From ASCII to a world of many alphabets

Although we associate ASCII with computers in the eighties, it dates back to the sixties. RFC 20, published in 1969, already proposed it for network interchange. Its seven bits allow 128 values: unaccented Latin letters, digits, punctuation, and control characters. ñ, á, and Chinese ideographs are outside that repertoire.

Other tables, such as ISO-8859-1 and Windows-1252, appeared to cover more languages, alongside encodings for other writing systems. Single-byte tables had little room, and different environments used it differently. A solution for one language could be insufficient for another; exchanging files between systems required knowing the source table.

ISO-8859-1, also known as Latin-1, and Windows-1252 are not interchangeable: they differ in the 0x80–0x9F range. For example, 0x80 represents € in Windows-1252 and a control character in Latin-1. Python provides separate codecs for both in its list of encodings.

Unicode emerged to build a shared repertoire. Early discussions began in 1987 among Xerox and Apple engineers, and the consortium was incorporated in 1991, according to its official history. The goal was to mix languages without switching tables each time.

Unicode assigns code points to characters. For example, ñ corresponds to U+00F1. UTF-8 and UTF-16 specify how to represent that text using units that can be stored and transmitted. There is no historical sequence of “ASCII, then UTF-8, then UTF-16, then Unicode”: both UTFs represent Unicode.

ConceptPurposeExample for ñ
UnicodeIdentify the character with a code pointU+00F1
UTF-8Use one to four bytes per Unicode scalar valueC3 B1
UTF-16LEUse one or two 16-bit units, with bytes in little-endian orderF1 00

UTF-8 preserves ASCII values. UTF-16 needs a pair of units for characters such as 😀: it does not mean “two bytes for every character.” Byte order also matters when serializing UTF-16, indicated by LE, BE, or, depending on the format, a BOM. The Unicode FAQ on UTFs explains these differences.

Why Unicode matters in Python

In Python 3, str represents Unicode text and bytes represents bytes. Keeping that boundary clear lets us process names, languages, and symbols without constantly converting them to a local table. A str is not a miniature UTF-8 file. The Python Unicode HOWTO develops this distinction.

text = "España 😀"                  # str
data = text.encode("utf-8")         # bytes
recovered = data.decode("utf-8")    # str

assert recovered == text
assert len(text) == 8
assert len(data) == 12

The length of a str counts code points; the length of bytes counts bytes. Neither always equals the number of visible symbols: a letter with a combining accent or a compound emoji can contain multiple code points. The Python type reference describes these sequences.

The practical rule is to decode on input, work with text, and encode on output:

source bytes ── decode(source) ──► Python str
Python str   ── encode(target) ──► target bytes

decode turns bytes into text; encode turns text into bytes. Displaying that text is the job of the output system and renderer. If a library already returns str, it may have handled that boundary: check the type before adding another conversion.

Handling an error: recover the source and interpret it correctly

For uninterpreted bytes, we need decode, using the actual source encoding. Returning to the original binary representation helps the investigation, but there is no “more basic binary encoding” that removes the need to know how those bytes were written.

Suppose a legacy system exports a name in Windows-1252 and we must deliver it as UTF-8:

source = b"Jos\xe9"                 # bytes received from the legacy system
text = source.decode("cp1252")      # "José"
output = text.encode("utf-8")        # b'Jos\xc3\xa9'

First interpret the source; then serialize for the target. For a large file, we can let text input and output handle the conversion:

with open("input.csv", encoding="cp1252", errors="strict") as source:
    with open("output.csv", "w", encoding="utf-8", errors="strict") as target:
        for line in source:
            target.write(line)

Specifying encoding avoids relying on environment defaults. Text and binary modes are documented in open.

When we already have José

This phenomenon is called mojibake: bytes were interpreted using the wrong rules. We can reproduce one specific cause:

original = "José".encode("utf-8")
wrong = original.decode("latin-1")
assert wrong == "José"

# Only because we know the wrong conversion and no information was lost:
recovered_bytes = wrong.encode("latin-1")
correct = recovered_bytes.decode("utf-8")
assert correct == "José"

This repair starts with encode to undo a known incorrect decoding. It does not contradict the input rule: we are reconstructing bytes that had already been turned into text. Applying encode(...).decode(...) indiscriminately is not a general solution. If the error used Windows-1252, we would need to investigate that transformation rather than assume Latin-1.

A decoding that raises no exception is not necessarily correct. Latin-1 assigns a character to every byte value: it can accept data that produces meaningless text. Automatic detectors offer hypotheses; the source contract, its metadata, and known samples provide better evidence.

If characters were replaced with ? or �, or removed using errors="ignore", information may have been lost. Recovering the original value then requires an earlier copy. In a pipeline, I recommend failing or quarantining the record before accepting a silent modification. That data quality decision should be explicit.

In a database: storage and collation

Encoding answers “how do I represent the text?” A collation answers “how do I sort and compare this text?” It defines linguistic rules and sensitivities, such as distinguishing case or accents.

A name can be stored perfectly yet produce unexpected results in WHERE, JOIN, ORDER BY, GROUP BY, DISTINCT, or a UNIQUE constraint. If a comparison treats two names as equivalent, they may be grouped or collide even when their characters differ. The effect depends on the engine and the specific collation.

In MySQL, ci and cs indicate case insensitivity and sensitivity; ai and as indicate accent insensitivity and sensitivity. Its collation naming documentation defines these suffixes. This example targets MySQL 8.0/8.4:

SELECT
    _utf8mb4'José' COLLATE utf8mb4_0900_ai_ci = _utf8mb4'jose'
        AS equal_names; -- 1

SELECT
    _utf8mb4'José' COLLATE utf8mb4_0900_as_cs = _utf8mb4'jose'
        AS equal_names; -- 0

To support the full Unicode repertoire in MySQL, use utf8mb4 rather than the older utf8mb3, restricted to characters representable in up to three bytes. The utf8mb4 documentation explains why an emoji such as 😀 requires the former. Check the connection configuration as well as the column.

In SQL Server, collation also determines the code page of varchar when it does not use UTF-8. Since SQL Server 2019, _UTF8 collations allow UTF-8 in varchar; nvarchar uses UCS-2 or UTF-16 depending on collation and supplementary character support. Unicode literals use N'José'. See collation and Unicode support in SQL Server.

Choosing the right collation

Start with the product rules. A name search may need to find José when someone types jose. An identifier may require ABC and abc to be different. Search and uniqueness do not necessarily need identical rules.

Test real names from the supported languages: José/Jose, ABC/abc, Peña/Pena, and Turkish or German cases when relevant. Do not assume every collation treats ñ as an n with a disposable accent: linguistic rules matter.

Before changing collation, review duplicates under the new comparison, indexes, constraints, and queries. Applying COLLATE in a query can solve an individual comparison, but check its execution plan. Changing collation will not reconstruct a José already stored as José.

Debugging the failure inside a pipeline

The useful question is: at which first boundary does the value change? Follow a known sample through the entire journey:

CSV → Python reader → transformation → database → API → browser

Preserve the source bytes and, for a synthetic sample, record the type and exact representation before and after each boundary. This small inspector avoids depending on how a console draws characters:

def inspect_value(value):
    if isinstance(value, bytes):
        return {"type": "bytes", "length": len(value), "hex": value.hex()}
    if isinstance(value, str):
        return {
            "type": "str",
            "length": len(value),
            "escaped": ascii(value),
            "points": [f"U+{ord(c):04X}" for c in value],
        }
    raise TypeError("Expected str or bytes")

print(inspect_value("José"))
print(inspect_value("José".encode("utf-8")))

In production, limit and protect these samples: personal data need not be dumped into logs. Add the source, codec, and stage so that the transformation can be reproduced.

SymptomHypothesis to investigate
JoséUTF-8 read as another table, or a repeated transformation
UnicodeDecodeErrorWrong codec, damaged bytes, or an incomplete multibyte sequence
UnicodeEncodeErrorThe output encoding cannot represent a character
Unexpected � or ?Lossy substitution somewhere along the way
Empty boxes, but correct code pointsMissing glyphs in the font
Correct data, unexpected search resultsCollation or normalization rules

Check the BOM: utf-8-sig consumes a UTF-8 signature if the source includes one. If the contract specifies UTF-16, check byte order. An HTTP header announcing a charset different from the actual bytes also deserves attention. In JSON, "Jos\u00e9" can be a perfectly valid escaped representation; compare the parsed value rather than just looking at the file. RFC 8259 defines escapes and UTF-8 for JSON interchange between systems.

When reading in chunks, a character can be split across two reads. Use a text reader or an incremental decoder, which retains pending bytes:

import codecs

decoder = codecs.getincrementaldecoder("utf-8")(errors="strict")
assert decoder.decode(b"\xc3") == ""
assert decoder.decode(b"\xb1", final=True) == "ñ"

Do not treat an incomplete chunk as evidence that the file uses a different encoding.

Preventing failures with data tests

A test using Alice or Madrid barely exercises these boundaries. I suggest a small corpus containing accents, ñ, €, typographic quotes, non-Latin alphabets, emojis, and combining characters. These unit tests provide a starting point and can be run with pytest:

import unicodedata
import pytest

@pytest.mark.parametrize("text", [
    "José", "España", "€", "“hola”", "東京", "العربية", "😀", "e\u0301",
])
def test_lossless_utf8(text):
    assert text.encode("utf-8").decode("utf-8") == text

def test_legacy_input_against_expected_value():
    assert b"Jos\xe9".decode("cp1252") == "José"

def test_invalid_utf8_is_rejected():
    with pytest.raises(UnicodeDecodeError):
        b"\xff".decode("utf-8", errors="strict")

def test_normalization_when_contract_requires_nfc():
    assert unicodedata.normalize("NFC", "e\u0301") == "é"

Unicode normalization resolves equivalent representations, such as é and e plus a combining accent; it does not repair mojibake. unicodedata.normalize applies NFC when the contract requires it. Do not remove accents from the original data merely to simplify a search.

The UTF-8 round trip is a basic check: José would pass it too. The decisive test follows the actual pipeline: read a fixture containing known bytes, run the transformation, insert using the same driver and schema as production, retrieve through the API, and compare against independent expected values. Include the streaming path and its chunk boundaries when applicable.

Add cases for searches and uniqueness under the chosen collation, rejection or quarantine of invalid input, and the absence of substitutions introduced by processing. For the source, document encoding and BOM policy; during processing, keep Unicode text; for each output, specify serialization. Once those agreements are written and tested, the next José becomes a specific boundary we can fix.