josiete.com

Data Quality

When bytes masquerade as text: Unicode, Python, and encoding errors

Streams of green characters on a black background, with the words decode and encode highlighted in the center

A file contains José, but after passing through a pipeline it becomes José. Another preserves accents, yet searching for Jose also returns José. In a third, emojis turn into ?.

All three look like text problems, but they call for different investigations: how we interpret bytes, how we compare strings, and what information was lost along the way. Switching to UTF-8 without locating the failure can simply store already damaged text correctly.