<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>UTF-8 on josiete.com</title><link>https://www.josiete.com/en/tags/utf-8/</link><description>Recent content in UTF-8 on josiete.com</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 28 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://www.josiete.com/en/tags/utf-8/index.xml" rel="self" type="application/rss+xml"/><item><title>When bytes masquerade as text: Unicode, Python, and encoding errors</title><link>https://www.josiete.com/en/posts/encoding-errors-unicode/</link><pubDate>Mon, 28 Sep 2026 00:00:00 +0000</pubDate><guid>https://www.josiete.com/en/posts/encoding-errors-unicode/</guid><description>&lt;p&gt;&lt;img src="https://www.josiete.com/images/encoding-errors-unicode.webp" alt="Streams of green characters on a black background, with the words decode and encode highlighted in the center"&gt;&lt;/p&gt;&#10;&lt;p&gt;A file contains &lt;code&gt;José&lt;/code&gt;, but after passing through a pipeline it becomes &lt;code&gt;JosÃ©&lt;/code&gt;. Another preserves accents, yet searching for &lt;code&gt;Jose&lt;/code&gt; also returns &lt;code&gt;José&lt;/code&gt;. In a third, emojis turn into &lt;code&gt;?&lt;/code&gt;.&lt;/p&gt;&#10;&lt;p&gt;All three look like text problems, but they call for different investigations: how we interpret bytes, how we compare strings, and what information was lost along the way. Switching to UTF-8 without locating the failure can simply store already damaged text correctly.&lt;/p&gt;&#10;&lt;h2 id="what-is-an-encoding-and-why-do-we-need-one"&gt;What is an encoding, and why do we need one?&lt;/h2&gt;&#10;&lt;p&gt;A file does not store a drawing of &lt;code&gt;ñ&lt;/code&gt;: it stores bytes. Turning those numbers into text requires an agreement about which sequences represent each character. That agreement is a &lt;strong&gt;character encoding&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;Think of it as the rules for writing and reading a message. Sender and receiver must share them: receiving identical bytes does not guarantee identical text when each side uses different rules.&lt;/p&gt;&#10;&lt;p&gt;Encodings arose from this need to represent and transmit text using machines. Available space, languages, and compatibility with existing equipment shaped the solutions. The &lt;a href="https://www.unicode.org/reports/tr17/"&gt;Unicode character encoding model&lt;/a&gt; distinguishes the character repertoire, its codes, and the ways of mapping them to storage units.&lt;/p&gt;&#10;&lt;p&gt;A reproducible Python example shows why bytes need context:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;print(&lt;span style="color:#e6db74"&gt;&amp;#34;ñ&amp;#34;&lt;/span&gt;&lt;span style="color:#f92672"&gt;.&lt;/span&gt;encode(&lt;span style="color:#e6db74"&gt;&amp;#34;utf-8&amp;#34;&lt;/span&gt;)&lt;span style="color:#f92672"&gt;.&lt;/span&gt;hex()) &lt;span style="color:#75715e"&gt;# c3b1&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;print(&lt;span style="color:#e6db74"&gt;&amp;#34;ñ&amp;#34;&lt;/span&gt;&lt;span style="color:#f92672"&gt;.&lt;/span&gt;encode(&lt;span style="color:#e6db74"&gt;&amp;#34;latin-1&amp;#34;&lt;/span&gt;)&lt;span style="color:#f92672"&gt;.&lt;/span&gt;hex()) &lt;span style="color:#75715e"&gt;# f1&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Both sequences represent the same letter under their respective encodings. Neither carries a label that, by itself, tells us how to read it.&lt;/p&gt;&#10;&lt;h2 id="from-ascii-to-a-world-of-many-alphabets"&gt;From ASCII to a world of many alphabets&lt;/h2&gt;&#10;&lt;p&gt;Although we associate ASCII with computers in the eighties, it dates back to the sixties. &lt;a href="https://www.rfc-editor.org/rfc/rfc20"&gt;RFC 20, published in 1969&lt;/a&gt;, already proposed it for network interchange. Its seven bits allow 128 values: unaccented Latin letters, digits, punctuation, and control characters. &lt;code&gt;ñ&lt;/code&gt;, &lt;code&gt;á&lt;/code&gt;, and Chinese ideographs are outside that repertoire.&lt;/p&gt;&#10;&lt;p&gt;Other tables, such as ISO-8859-1 and Windows-1252, appeared to cover more languages, alongside encodings for other writing systems. Single-byte tables had little room, and different environments used it differently. A solution for one language could be insufficient for another; exchanging files between systems required knowing the source table.&lt;/p&gt;&#10;&lt;p&gt;ISO-8859-1, also known as Latin-1, and Windows-1252 are not interchangeable: they differ in the &lt;code&gt;0x80–0x9F&lt;/code&gt; range. For example, &lt;code&gt;0x80&lt;/code&gt; represents &lt;code&gt;€&lt;/code&gt; in Windows-1252 and a control character in Latin-1. Python provides separate codecs for both in its &lt;a href="https://docs.python.org/3/library/codecs.html#standard-encodings"&gt;list of encodings&lt;/a&gt;.&lt;/p&gt;&#10;&lt;p&gt;Unicode emerged to build a shared repertoire. Early discussions began in 1987 among Xerox and Apple engineers, and the consortium was incorporated in 1991, according to its &lt;a href="https://www.unicode.org/history/"&gt;official history&lt;/a&gt;. The goal was to mix languages without switching tables each time.&lt;/p&gt;&#10;&lt;h2 id="unicode-utf-8-and-utf-16-are-related-concepts"&gt;Unicode, UTF-8, and UTF-16 are related concepts&lt;/h2&gt;&#10;&lt;p&gt;&lt;strong&gt;Unicode assigns code points to characters&lt;/strong&gt;. For example, &lt;code&gt;ñ&lt;/code&gt; corresponds to &lt;code&gt;U+00F1&lt;/code&gt;. &lt;strong&gt;UTF-8 and UTF-16 specify how to represent that text&lt;/strong&gt; using units that can be stored and transmitted. There is no historical sequence of “ASCII, then UTF-8, then UTF-16, then Unicode”: both UTFs represent Unicode.&lt;/p&gt;&#10;&lt;table&gt;&#10;&#9;&lt;thead&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Concept&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Purpose&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Example for &lt;code&gt;ñ&lt;/code&gt;&lt;/th&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/thead&gt;&#10;&#9;&lt;tbody&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Unicode&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Identify the character with a code point&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;&lt;code&gt;U+00F1&lt;/code&gt;&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;UTF-8&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Use one to four bytes per Unicode scalar value&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;&lt;code&gt;C3 B1&lt;/code&gt;&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;UTF-16LE&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Use one or two 16-bit units, with bytes in little-endian order&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;&lt;code&gt;F1 00&lt;/code&gt;&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;p&gt;UTF-8 preserves ASCII values. UTF-16 needs a pair of units for characters such as &lt;code&gt;😀&lt;/code&gt;: it does not mean “two bytes for every character.” Byte order also matters when serializing UTF-16, indicated by LE, BE, or, depending on the format, a BOM. The &lt;a href="https://www.unicode.org/faq/utf_bom.html"&gt;Unicode FAQ on UTFs&lt;/a&gt; explains these differences.&lt;/p&gt;&#10;&lt;h2 id="why-unicode-matters-in-python"&gt;Why Unicode matters in Python&lt;/h2&gt;&#10;&lt;p&gt;In Python 3, &lt;code&gt;str&lt;/code&gt; represents Unicode text and &lt;code&gt;bytes&lt;/code&gt; represents bytes. Keeping that boundary clear lets us process names, languages, and symbols without constantly converting them to a local table. A &lt;code&gt;str&lt;/code&gt; is not a miniature UTF-8 file. The &lt;a href="https://docs.python.org/3/howto/unicode.html"&gt;Python Unicode HOWTO&lt;/a&gt; develops this distinction.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;text &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;España 😀&amp;#34;&lt;/span&gt; &lt;span style="color:#75715e"&gt;# str&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;data &lt;span style="color:#f92672"&gt;=&lt;/span&gt; text&lt;span style="color:#f92672"&gt;.&lt;/span&gt;encode(&lt;span style="color:#e6db74"&gt;&amp;#34;utf-8&amp;#34;&lt;/span&gt;) &lt;span style="color:#75715e"&gt;# bytes&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;recovered &lt;span style="color:#f92672"&gt;=&lt;/span&gt; data&lt;span style="color:#f92672"&gt;.&lt;/span&gt;decode(&lt;span style="color:#e6db74"&gt;&amp;#34;utf-8&amp;#34;&lt;/span&gt;) &lt;span style="color:#75715e"&gt;# str&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;assert&lt;/span&gt; recovered &lt;span style="color:#f92672"&gt;==&lt;/span&gt; text&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;assert&lt;/span&gt; len(text) &lt;span style="color:#f92672"&gt;==&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;8&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;assert&lt;/span&gt; len(data) &lt;span style="color:#f92672"&gt;==&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;12&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The length of a &lt;code&gt;str&lt;/code&gt; counts code points; the length of &lt;code&gt;bytes&lt;/code&gt; counts bytes. Neither always equals the number of visible symbols: a letter with a combining accent or a compound emoji can contain multiple code points. The &lt;a href="https://docs.python.org/3/library/stdtypes.html#text-sequence-type-str"&gt;Python type reference&lt;/a&gt; describes these sequences.&lt;/p&gt;&#10;&lt;p&gt;The practical rule is to &lt;strong&gt;decode on input, work with text, and encode on output&lt;/strong&gt;:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;source bytes ── decode(source) ──► Python str&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Python str ── encode(target) ──► target bytes&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;code&gt;decode&lt;/code&gt; turns bytes into text; &lt;code&gt;encode&lt;/code&gt; turns text into bytes. Displaying that text is the job of the output system and renderer. If a library already returns &lt;code&gt;str&lt;/code&gt;, it may have handled that boundary: check the type before adding another conversion.&lt;/p&gt;&#10;&lt;h2 id="handling-an-error-recover-the-source-and-interpret-it-correctly"&gt;Handling an error: recover the source and interpret it correctly&lt;/h2&gt;&#10;&lt;p&gt;For uninterpreted bytes, we need &lt;strong&gt;&lt;code&gt;decode&lt;/code&gt;, using the actual source encoding&lt;/strong&gt;. Returning to the original binary representation helps the investigation, but there is no “more basic binary encoding” that removes the need to know how those bytes were written.&lt;/p&gt;&#10;&lt;p&gt;Suppose a legacy system exports a name in Windows-1252 and we must deliver it as UTF-8:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;source &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#e6db74"&gt;b&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;Jos&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;\xe9&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt; &lt;span style="color:#75715e"&gt;# bytes received from the legacy system&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;text &lt;span style="color:#f92672"&gt;=&lt;/span&gt; source&lt;span style="color:#f92672"&gt;.&lt;/span&gt;decode(&lt;span style="color:#e6db74"&gt;&amp;#34;cp1252&amp;#34;&lt;/span&gt;) &lt;span style="color:#75715e"&gt;# &amp;#34;José&amp;#34;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;output &lt;span style="color:#f92672"&gt;=&lt;/span&gt; text&lt;span style="color:#f92672"&gt;.&lt;/span&gt;encode(&lt;span style="color:#e6db74"&gt;&amp;#34;utf-8&amp;#34;&lt;/span&gt;) &lt;span style="color:#75715e"&gt;# b&amp;#39;Jos\xc3\xa9&amp;#39;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;First interpret the source; then serialize for the target. For a large file, we can let text input and output handle the conversion:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;with&lt;/span&gt; open(&lt;span style="color:#e6db74"&gt;&amp;#34;input.csv&amp;#34;&lt;/span&gt;, encoding&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;cp1252&amp;#34;&lt;/span&gt;, errors&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;strict&amp;#34;&lt;/span&gt;) &lt;span style="color:#66d9ef"&gt;as&lt;/span&gt; source:&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;with&lt;/span&gt; open(&lt;span style="color:#e6db74"&gt;&amp;#34;output.csv&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;w&amp;#34;&lt;/span&gt;, encoding&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;utf-8&amp;#34;&lt;/span&gt;, errors&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;strict&amp;#34;&lt;/span&gt;) &lt;span style="color:#66d9ef"&gt;as&lt;/span&gt; target:&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; line &lt;span style="color:#f92672"&gt;in&lt;/span&gt; source:&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; target&lt;span style="color:#f92672"&gt;.&lt;/span&gt;write(line)&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Specifying &lt;code&gt;encoding&lt;/code&gt; avoids relying on environment defaults. Text and binary modes are documented in &lt;a href="https://docs.python.org/3/library/functions.html#open"&gt;&lt;code&gt;open&lt;/code&gt;&lt;/a&gt;.&lt;/p&gt;&#10;&lt;h3 id="when-we-already-have-josã"&gt;When we already have &lt;code&gt;JosÃ©&lt;/code&gt;&lt;/h3&gt;&#10;&lt;p&gt;This phenomenon is called &lt;em&gt;mojibake&lt;/em&gt;: bytes were interpreted using the wrong rules. We can reproduce one specific cause:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;original &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;José&amp;#34;&lt;/span&gt;&lt;span style="color:#f92672"&gt;.&lt;/span&gt;encode(&lt;span style="color:#e6db74"&gt;&amp;#34;utf-8&amp;#34;&lt;/span&gt;)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;wrong &lt;span style="color:#f92672"&gt;=&lt;/span&gt; original&lt;span style="color:#f92672"&gt;.&lt;/span&gt;decode(&lt;span style="color:#e6db74"&gt;&amp;#34;latin-1&amp;#34;&lt;/span&gt;)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;assert&lt;/span&gt; wrong &lt;span style="color:#f92672"&gt;==&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;JosÃ©&amp;#34;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Only because we know the wrong conversion and no information was lost:&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;recovered_bytes &lt;span style="color:#f92672"&gt;=&lt;/span&gt; wrong&lt;span style="color:#f92672"&gt;.&lt;/span&gt;encode(&lt;span style="color:#e6db74"&gt;&amp;#34;latin-1&amp;#34;&lt;/span&gt;)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;correct &lt;span style="color:#f92672"&gt;=&lt;/span&gt; recovered_bytes&lt;span style="color:#f92672"&gt;.&lt;/span&gt;decode(&lt;span style="color:#e6db74"&gt;&amp;#34;utf-8&amp;#34;&lt;/span&gt;)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;assert&lt;/span&gt; correct &lt;span style="color:#f92672"&gt;==&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;José&amp;#34;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This repair starts with &lt;code&gt;encode&lt;/code&gt; to undo a known incorrect decoding. It does not contradict the input rule: we are reconstructing bytes that had already been turned into text. &lt;strong&gt;Applying &lt;code&gt;encode(...).decode(...)&lt;/code&gt; indiscriminately is not a general solution.&lt;/strong&gt; If the error used Windows-1252, we would need to investigate that transformation rather than assume Latin-1.&lt;/p&gt;&#10;&lt;p&gt;A decoding that raises no exception is not necessarily correct. Latin-1 assigns a character to every byte value: it can accept data that produces meaningless text. Automatic detectors offer hypotheses; the source contract, its metadata, and known samples provide better evidence.&lt;/p&gt;&#10;&lt;p&gt;If characters were replaced with &lt;code&gt;?&lt;/code&gt; or &lt;code&gt;�&lt;/code&gt;, or removed using &lt;code&gt;errors=&amp;quot;ignore&amp;quot;&lt;/code&gt;, information may have been lost. Recovering the original value then requires an earlier copy. In a pipeline, I recommend failing or quarantining the record before accepting a silent modification. That data quality decision should be explicit.&lt;/p&gt;&#10;&lt;h2 id="in-a-database-storage-and-collation"&gt;In a database: storage and collation&lt;/h2&gt;&#10;&lt;p&gt;Encoding answers “how do I represent the text?” A &lt;strong&gt;collation&lt;/strong&gt; answers “how do I sort and compare this text?” It defines linguistic rules and sensitivities, such as distinguishing case or accents.&lt;/p&gt;&#10;&lt;p&gt;A name can be stored perfectly yet produce unexpected results in &lt;code&gt;WHERE&lt;/code&gt;, &lt;code&gt;JOIN&lt;/code&gt;, &lt;code&gt;ORDER BY&lt;/code&gt;, &lt;code&gt;GROUP BY&lt;/code&gt;, &lt;code&gt;DISTINCT&lt;/code&gt;, or a &lt;code&gt;UNIQUE&lt;/code&gt; constraint. If a comparison treats two names as equivalent, they may be grouped or collide even when their characters differ. The effect depends on the engine and the specific collation.&lt;/p&gt;&#10;&lt;p&gt;In MySQL, &lt;code&gt;ci&lt;/code&gt; and &lt;code&gt;cs&lt;/code&gt; indicate case insensitivity and sensitivity; &lt;code&gt;ai&lt;/code&gt; and &lt;code&gt;as&lt;/code&gt; indicate accent insensitivity and sensitivity. Its &lt;a href="https://dev.mysql.com/doc/refman/8.4/en/charset-collation-names.html"&gt;collation naming documentation&lt;/a&gt; defines these suffixes. This example targets MySQL 8.0/8.4:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-sql" data-lang="sql"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;SELECT&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; _utf8mb4&lt;span style="color:#e6db74"&gt;&amp;#39;José&amp;#39;&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;COLLATE&lt;/span&gt; utf8mb4_0900_ai_ci &lt;span style="color:#f92672"&gt;=&lt;/span&gt; _utf8mb4&lt;span style="color:#e6db74"&gt;&amp;#39;jose&amp;#39;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;AS&lt;/span&gt; equal_names; &lt;span style="color:#75715e"&gt;-- 1&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;SELECT&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; _utf8mb4&lt;span style="color:#e6db74"&gt;&amp;#39;José&amp;#39;&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;COLLATE&lt;/span&gt; utf8mb4_0900_as_cs &lt;span style="color:#f92672"&gt;=&lt;/span&gt; _utf8mb4&lt;span style="color:#e6db74"&gt;&amp;#39;jose&amp;#39;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;AS&lt;/span&gt; equal_names; &lt;span style="color:#75715e"&gt;-- 0&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;To support the full Unicode repertoire in MySQL, use &lt;code&gt;utf8mb4&lt;/code&gt; rather than the older &lt;code&gt;utf8mb3&lt;/code&gt;, restricted to characters representable in up to three bytes. The &lt;a href="https://dev.mysql.com/doc/refman/8.4/en/charset-unicode-utf8mb4.html"&gt;utf8mb4 documentation&lt;/a&gt; explains why an emoji such as &lt;code&gt;😀&lt;/code&gt; requires the former. Check the connection configuration as well as the column.&lt;/p&gt;&#10;&lt;p&gt;In SQL Server, collation also determines the code page of &lt;code&gt;varchar&lt;/code&gt; when it does not use UTF-8. Since SQL Server 2019, &lt;code&gt;_UTF8&lt;/code&gt; collations allow UTF-8 in &lt;code&gt;varchar&lt;/code&gt;; &lt;code&gt;nvarchar&lt;/code&gt; uses UCS-2 or UTF-16 depending on collation and supplementary character support. Unicode literals use &lt;code&gt;N'José'&lt;/code&gt;. See &lt;a href="https://learn.microsoft.com/en-us/sql/relational-databases/collations/collation-and-unicode-support"&gt;collation and Unicode support in SQL Server&lt;/a&gt;.&lt;/p&gt;&#10;&lt;h3 id="choosing-the-right-collation"&gt;Choosing the right collation&lt;/h3&gt;&#10;&lt;p&gt;Start with the product rules. A name search may need to find &lt;code&gt;José&lt;/code&gt; when someone types &lt;code&gt;jose&lt;/code&gt;. An identifier may require &lt;code&gt;ABC&lt;/code&gt; and &lt;code&gt;abc&lt;/code&gt; to be different. Search and uniqueness do not necessarily need identical rules.&lt;/p&gt;&#10;&lt;p&gt;Test real names from the supported languages: &lt;code&gt;José/Jose&lt;/code&gt;, &lt;code&gt;ABC/abc&lt;/code&gt;, &lt;code&gt;Peña/Pena&lt;/code&gt;, and Turkish or German cases when relevant. &lt;strong&gt;Do not assume every collation treats &lt;code&gt;ñ&lt;/code&gt; as an &lt;code&gt;n&lt;/code&gt; with a disposable accent&lt;/strong&gt;: linguistic rules matter.&lt;/p&gt;&#10;&lt;p&gt;Before changing collation, review duplicates under the new comparison, indexes, constraints, and queries. Applying &lt;code&gt;COLLATE&lt;/code&gt; in a query can solve an individual comparison, but check its execution plan. Changing collation will not reconstruct a &lt;code&gt;José&lt;/code&gt; already stored as &lt;code&gt;JosÃ©&lt;/code&gt;.&lt;/p&gt;&#10;&lt;h2 id="debugging-the-failure-inside-a-pipeline"&gt;Debugging the failure inside a pipeline&lt;/h2&gt;&#10;&lt;p&gt;The useful question is: &lt;strong&gt;at which first boundary does the value change?&lt;/strong&gt; Follow a known sample through the entire journey:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;CSV → Python reader → transformation → database → API → browser&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Preserve the source bytes and, for a synthetic sample, record the type and exact representation before and after each boundary. This small inspector avoids depending on how a console draws characters:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;inspect_value&lt;/span&gt;(value):&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; isinstance(value, bytes):&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; {&lt;span style="color:#e6db74"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;bytes&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;length&amp;#34;&lt;/span&gt;: len(value), &lt;span style="color:#e6db74"&gt;&amp;#34;hex&amp;#34;&lt;/span&gt;: value&lt;span style="color:#f92672"&gt;.&lt;/span&gt;hex()}&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; isinstance(value, str):&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; {&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;str&amp;#34;&lt;/span&gt;,&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;length&amp;#34;&lt;/span&gt;: len(value),&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;escaped&amp;#34;&lt;/span&gt;: ascii(value),&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;points&amp;#34;&lt;/span&gt;: [&lt;span style="color:#e6db74"&gt;f&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;U+&lt;/span&gt;&lt;span style="color:#e6db74"&gt;{&lt;/span&gt;ord(c)&lt;span style="color:#e6db74"&gt;:&lt;/span&gt;&lt;span style="color:#e6db74"&gt;04X&lt;/span&gt;&lt;span style="color:#e6db74"&gt;}&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; c &lt;span style="color:#f92672"&gt;in&lt;/span&gt; value],&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;raise&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;TypeError&lt;/span&gt;(&lt;span style="color:#e6db74"&gt;&amp;#34;Expected str or bytes&amp;#34;&lt;/span&gt;)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;print(inspect_value(&lt;span style="color:#e6db74"&gt;&amp;#34;José&amp;#34;&lt;/span&gt;))&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;print(inspect_value(&lt;span style="color:#e6db74"&gt;&amp;#34;José&amp;#34;&lt;/span&gt;&lt;span style="color:#f92672"&gt;.&lt;/span&gt;encode(&lt;span style="color:#e6db74"&gt;&amp;#34;utf-8&amp;#34;&lt;/span&gt;)))&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;In production, limit and protect these samples: personal data need not be dumped into logs. Add the source, codec, and stage so that the transformation can be reproduced.&lt;/p&gt;&#10;&lt;table&gt;&#10;&#9;&lt;thead&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Symptom&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Hypothesis to investigate&lt;/th&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/thead&gt;&#10;&#9;&lt;tbody&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;&lt;code&gt;JosÃ©&lt;/code&gt;&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;UTF-8 read as another table, or a repeated transformation&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;&lt;code&gt;UnicodeDecodeError&lt;/code&gt;&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Wrong codec, damaged bytes, or an incomplete multibyte sequence&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;&lt;code&gt;UnicodeEncodeError&lt;/code&gt;&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;The output encoding cannot represent a character&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Unexpected &lt;code&gt;�&lt;/code&gt; or &lt;code&gt;?&lt;/code&gt;&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Lossy substitution somewhere along the way&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Empty boxes, but correct code points&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Missing glyphs in the font&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Correct data, unexpected search results&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Collation or normalization rules&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;p&gt;Check the BOM: &lt;code&gt;utf-8-sig&lt;/code&gt; consumes a UTF-8 signature if the source includes one. If the contract specifies UTF-16, check byte order. An HTTP header announcing a charset different from the actual bytes also deserves attention. In JSON, &lt;code&gt;&amp;quot;Jos\u00e9&amp;quot;&lt;/code&gt; can be a perfectly valid escaped representation; compare the parsed value rather than just looking at the file. &lt;a href="https://www.rfc-editor.org/rfc/rfc8259"&gt;RFC 8259&lt;/a&gt; defines escapes and UTF-8 for JSON interchange between systems.&lt;/p&gt;&#10;&lt;p&gt;When reading in chunks, a character can be split across two reads. Use a text reader or an &lt;a href="https://docs.python.org/3/library/codecs.html#incrementaldecoder-objects"&gt;incremental decoder&lt;/a&gt;, which retains pending bytes:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;import&lt;/span&gt; codecs&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;decoder &lt;span style="color:#f92672"&gt;=&lt;/span&gt; codecs&lt;span style="color:#f92672"&gt;.&lt;/span&gt;getincrementaldecoder(&lt;span style="color:#e6db74"&gt;&amp;#34;utf-8&amp;#34;&lt;/span&gt;)(errors&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;strict&amp;#34;&lt;/span&gt;)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;assert&lt;/span&gt; decoder&lt;span style="color:#f92672"&gt;.&lt;/span&gt;decode(&lt;span style="color:#e6db74"&gt;b&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;\xc3&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;) &lt;span style="color:#f92672"&gt;==&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;#34;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;assert&lt;/span&gt; decoder&lt;span style="color:#f92672"&gt;.&lt;/span&gt;decode(&lt;span style="color:#e6db74"&gt;b&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;\xb1&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;, final&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#66d9ef"&gt;True&lt;/span&gt;) &lt;span style="color:#f92672"&gt;==&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;ñ&amp;#34;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Do not treat an incomplete chunk as evidence that the file uses a different encoding.&lt;/p&gt;&#10;&lt;h2 id="preventing-failures-with-data-tests"&gt;Preventing failures with data tests&lt;/h2&gt;&#10;&lt;p&gt;A test using &lt;code&gt;Alice&lt;/code&gt; or &lt;code&gt;Madrid&lt;/code&gt; barely exercises these boundaries. I suggest a small corpus containing accents, &lt;code&gt;ñ&lt;/code&gt;, &lt;code&gt;€&lt;/code&gt;, typographic quotes, non-Latin alphabets, emojis, and combining characters. These unit tests provide a starting point and can be run with &lt;code&gt;pytest&lt;/code&gt;:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;import&lt;/span&gt; unicodedata&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;import&lt;/span&gt; pytest&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#a6e22e"&gt;@pytest.mark.parametrize&lt;/span&gt;(&lt;span style="color:#e6db74"&gt;&amp;#34;text&amp;#34;&lt;/span&gt;, [&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;José&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;España&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;€&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;“hola”&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;東京&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;العربية&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;😀&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;e&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;\u0301&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;,&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;])&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;test_lossless_utf8&lt;/span&gt;(text):&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;assert&lt;/span&gt; text&lt;span style="color:#f92672"&gt;.&lt;/span&gt;encode(&lt;span style="color:#e6db74"&gt;&amp;#34;utf-8&amp;#34;&lt;/span&gt;)&lt;span style="color:#f92672"&gt;.&lt;/span&gt;decode(&lt;span style="color:#e6db74"&gt;&amp;#34;utf-8&amp;#34;&lt;/span&gt;) &lt;span style="color:#f92672"&gt;==&lt;/span&gt; text&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;test_legacy_input_against_expected_value&lt;/span&gt;():&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;assert&lt;/span&gt; &lt;span style="color:#e6db74"&gt;b&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;Jos&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;\xe9&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;&lt;span style="color:#f92672"&gt;.&lt;/span&gt;decode(&lt;span style="color:#e6db74"&gt;&amp;#34;cp1252&amp;#34;&lt;/span&gt;) &lt;span style="color:#f92672"&gt;==&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;José&amp;#34;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;test_invalid_utf8_is_rejected&lt;/span&gt;():&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;with&lt;/span&gt; pytest&lt;span style="color:#f92672"&gt;.&lt;/span&gt;raises(&lt;span style="color:#a6e22e"&gt;UnicodeDecodeError&lt;/span&gt;):&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;b&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;\xff&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;&lt;span style="color:#f92672"&gt;.&lt;/span&gt;decode(&lt;span style="color:#e6db74"&gt;&amp;#34;utf-8&amp;#34;&lt;/span&gt;, errors&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;strict&amp;#34;&lt;/span&gt;)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;test_normalization_when_contract_requires_nfc&lt;/span&gt;():&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;assert&lt;/span&gt; unicodedata&lt;span style="color:#f92672"&gt;.&lt;/span&gt;normalize(&lt;span style="color:#e6db74"&gt;&amp;#34;NFC&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;e&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;\u0301&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;) &lt;span style="color:#f92672"&gt;==&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;é&amp;#34;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Unicode normalization resolves equivalent representations, such as &lt;code&gt;é&lt;/code&gt; and &lt;code&gt;e&lt;/code&gt; plus a combining accent; it does not repair mojibake. &lt;a href="https://docs.python.org/3/library/unicodedata.html#unicodedata.normalize"&gt;&lt;code&gt;unicodedata.normalize&lt;/code&gt;&lt;/a&gt; applies NFC when the contract requires it. Do not remove accents from the original data merely to simplify a search.&lt;/p&gt;&#10;&lt;p&gt;The UTF-8 &lt;em&gt;round trip&lt;/em&gt; is a basic check: &lt;code&gt;JosÃ©&lt;/code&gt; would pass it too. The decisive test follows &lt;strong&gt;the actual pipeline&lt;/strong&gt;: read a fixture containing known bytes, run the transformation, insert using the same driver and schema as production, retrieve through the API, and compare against independent expected values. Include the streaming path and its chunk boundaries when applicable.&lt;/p&gt;&#10;&lt;p&gt;Add cases for searches and uniqueness under the chosen collation, rejection or quarantine of invalid input, and the absence of substitutions introduced by processing. For the source, document encoding and BOM policy; during processing, keep Unicode text; for each output, specify serialization. Once those agreements are written and tested, the next &lt;code&gt;JosÃ©&lt;/code&gt; becomes a specific boundary we can fix.&lt;/p&gt;&#10;</description></item></channel></rss>