<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>OpenMetadata on josiete.com</title><link>https://www.josiete.com/en/tags/openmetadata/</link><description>Recent content in OpenMetadata on josiete.com</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sat, 10 Oct 2026 00:00:00 +0200</lastBuildDate><atom:link href="https://www.josiete.com/en/tags/openmetadata/index.xml" rel="self" type="application/rss+xml"/><item><title>Schema Drift in ETL: when data changes and your code does not notice</title><link>https://www.josiete.com/en/posts/schema-drift-schema-evolution-etl/</link><pubDate>Sat, 10 Oct 2026 00:00:00 +0200</pubDate><guid>https://www.josiete.com/en/posts/schema-drift-schema-evolution-etl/</guid><description>&lt;p&gt;&lt;img src="https://www.josiete.com/images/schema-drift-etl-timeline.svg" alt="September timeline: data moves from Schema A to B and C before the ETL changes from v1.0 to v1.1 and v1.2; orange bands mark periods whose compatibility must be checked"&gt;&lt;/p&gt;&#10;&lt;p&gt;Every morning, an ETL receives a CSV, converts ages to integers, and loads customers. Yesterday it worked with this file:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-csv" data-lang="csv"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt;id&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;name&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;age&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt;1&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;John&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;25&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt;2&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;Mary&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;30&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Today, the provider adds a country:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-csv" data-lang="csv"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt;id&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;name&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;age&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;country&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt;1&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;John&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;25&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;USA&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt;2&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;Mary&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;30&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;UK&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A few days later, another surprise arrives:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-csv" data-lang="csv"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt;id&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;name&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;age&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;country&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt;1&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;John&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;25&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;USA&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt;2&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;Mary&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;unknown&lt;/span&gt;,&lt;span style="color:#e6db74"&gt;UK&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;An additional column may be harmless if we select fields by name and allow extra columns. If we unpack each row into three variables, it may break the load. &lt;code&gt;unknown&lt;/code&gt; may cause an age conversion error, silently become null, or lead a subsequent inference to treat the entire column as text.&lt;/p&gt;&#10;&lt;p&gt;Compatibility depends on the contract and the code. Even the last case may be a violation involving individual values without the provider changing its declared schema. The question is: &lt;strong&gt;if we know which version of the code processed a file, can we also establish its data structure and when it stopped being compatible with our ETL?&lt;/strong&gt;&lt;/p&gt;&#10;&lt;h2 id="detecting-a-change-does-not-approve-it"&gt;Detecting a change does not approve it&lt;/h2&gt;&#10;&lt;p&gt;&lt;strong&gt;Schema Drift&lt;/strong&gt; describes divergence between the expected and observed structure: new columns, missing fields, or different types. In a CSV, whose values are text, observed types depend on how we interpret them.&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Schema Evolution&lt;/strong&gt; means managing those changes: agreeing on versions, deciding what changes we accept, and adapting consumers. Adding an optional field may preserve compatibility; removing a key we need will usually break it.&lt;/p&gt;&#10;&lt;p&gt;A &lt;strong&gt;Data Contract&lt;/strong&gt; records the agreement between producer and consumer: fields, types, nullability, formats, units, quality expectations, and delivery conditions. Making it operational requires checks and a policy for violations. Recording a change, accepting it, and transforming the data are separate decisions.&lt;/p&gt;&#10;&lt;p&gt;For example, we might allow new columns, reject a missing &lt;code&gt;id&lt;/code&gt;, and quarantine rows with invalid ages. That policy should be versioned: “the file could be read” tells us less than “it satisfied contract 2.1.”&lt;/p&gt;&#10;&lt;h2 id="git-tells-only-part-of-the-story"&gt;Git tells only part of the story&lt;/h2&gt;&#10;&lt;p&gt;Git can identify deployed code, but it does not automatically record what an external provider sent. Consider this sequence, where A, B, and C represent observed structures rather than already approved contracts:&lt;/p&gt;&#10;&lt;table&gt;&#10;&#9;&lt;thead&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Reception date&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Received schema&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;ETL version&lt;/th&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/thead&gt;&#10;&#9;&lt;tbody&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;01 Sep&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Schema A&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;v1.0&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;12 Sep&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Schema B&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;v1.0&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;15 Sep&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Schema B&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;v1.1&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;20 Sep&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Schema C&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;v1.1&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;25 Sep&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Schema C&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;v1.2&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;p&gt;If v1.0 only supported A and v1.1 only supported A and B, batches received between 12 and 15 September, and between 20 and 25 September, are potentially affected. The diagram&amp;rsquo;s bands mark those review periods; they do not establish incompatibility on their own. We must consult contracts, validations, and each run&amp;rsquo;s results.&lt;/p&gt;&#10;&lt;p&gt;For each batch, we should record:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Executed code:&lt;/strong&gt; a Git commit or immutable artifact, as well as a readable label.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Observed schema and applied contract:&lt;/strong&gt; version identifiers and their descriptors.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;File identity:&lt;/strong&gt; source, batch identifier, and content hash; also the stored object version when available.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Reception and processing:&lt;/strong&gt; separate timestamps with time zones and a run identifier.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Outcome:&lt;/strong&gt; checks, rejections, and produced outputs.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;There are also two clocks: when the source changed and when we detected it. A nightly inspection only establishes what was observed that night. Without provider events or historical files, we may be able to bound the change between two observations, but cannot assign an exact time.&lt;/p&gt;&#10;&lt;p&gt;A metadata history does not automatically preserve original content either. Reproducing the calculation also requires the data, dependencies, and ETL configuration, subject to the relevant retention policy.&lt;/p&gt;&#10;&lt;h2 id="failures-that-leave-the-pipeline-green"&gt;Failures that leave the pipeline green&lt;/h2&gt;&#10;&lt;p&gt;An amount moves from euros to cents: &lt;code&gt;19.95&lt;/code&gt; becomes &lt;code&gt;1995&lt;/code&gt;. Both can be valid numbers. A list of columns and types will not detect the unit change; we need declared units, business rules, or a statistical comparison that flags the jump.&lt;/p&gt;&#10;&lt;p&gt;A date changes from &lt;code&gt;2026-10-09&lt;/code&gt; to &lt;code&gt;09/10/2026&lt;/code&gt;. If the contract requires a format, we can reject it. If the schema merely says “text,” everything looks correct; a permissive parser may also confuse day and month for ambiguous dates.&lt;/p&gt;&#10;&lt;p&gt;&lt;code&gt;La Coruña&lt;/code&gt; appears as &lt;code&gt;La CoruÃ±a&lt;/code&gt; because bytes were interpreted incorrectly. It is still text, and may even subsequently be stored in a perfectly valid UTF-8 file. Specifying the encoding helps, but does not repair earlier corruption. This connects to &lt;a href="https://www.josiete.com/en/posts/encoding-errors-unicode/"&gt;encoding errors and Unicode&lt;/a&gt;.&lt;/p&gt;&#10;&lt;p&gt;These problems require combining structure, semantics, and &lt;strong&gt;Data Quality&lt;/strong&gt;. An age range, an explicit currency, or a list of valid places catches issues outside schema analysis. A different distribution may also reflect a real business change: an alert calls for investigation, not automatic correction.&lt;/p&gt;&#10;&lt;h2 id="what-each-tool-contributes"&gt;What each tool contributes&lt;/h2&gt;&#10;&lt;h3 id="openmetadata-a-shared-memory-for-the-ecosystem"&gt;OpenMetadata: a shared memory for the ecosystem&lt;/h3&gt;&#10;&lt;p&gt;&lt;a href="https://open-metadata.org/"&gt;OpenMetadata&lt;/a&gt; centralizes metadata about assets, owners, and relationships. Its &lt;a href="https://docs.open-metadata.org/v2.0.x/how-to-guides/guide-for-data-users/versions"&gt;version history&lt;/a&gt; includes schema changes and other metadata. It uses major and minor versions; deleting a column is a documented example of an incompatible change. That catalog classification does not replace the compatibility requirements of our particular consumer.&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://docs.open-metadata.org/v2.0.x/how-to-guides/data-quality-observability/alerts-notifications/data-observability-alerts"&gt;Observability alerts&lt;/a&gt; can monitor added, deleted, or updated columns, test results, and pipeline states. We can filter events and route them to email, chat channels, or webhooks. Its &lt;a href="https://docs.open-metadata.org/v2.0.x/how-to-guides/data-quality-observability/quality/tests-ui"&gt;table tests&lt;/a&gt; check expected columns and counts, among other things; profiling supplies additional metrics.&lt;/p&gt;&#10;&lt;p&gt;It fits situations where several teams need a common catalog spanning databases, storage, pipelines, and analytics. It requires a platform deployment and configured ingestion and connectors. A change not yet ingested remains outside its knowledge: checking every CSV on arrival requires instrumenting that entry point.&lt;/p&gt;&#10;&lt;h3 id="openlineage-connecting-data-code-and-runs"&gt;OpenLineage: connecting data, code, and runs&lt;/h3&gt;&#10;&lt;p&gt;&lt;a href="https://openlineage.io/docs/spec/object-model/"&gt;OpenLineage&lt;/a&gt; defines an open standard for lineage. A &lt;strong&gt;job&lt;/strong&gt; identifies logical work; a &lt;strong&gt;run&lt;/strong&gt;, one execution; a &lt;strong&gt;dataset&lt;/strong&gt;, an input or output. &lt;strong&gt;Facets&lt;/strong&gt; add metadata to these entities and to a run&amp;rsquo;s inputs or outputs.&lt;/p&gt;&#10;&lt;p&gt;The &lt;a href="https://openlineage.io/docs/spec/facets/dataset-facets/schema/"&gt;schema facet&lt;/a&gt; describes fields and types. The &lt;a href="https://openlineage.io/docs/spec/facets/job-facets/source-code-location/"&gt;source code location facet&lt;/a&gt; supports repository and version information. When our event producer supplies them, we can connect a run to its datasets and executed code. We must emit the actual reference for the deployed artifact rather than assume it matches the current branch.&lt;/p&gt;&#10;&lt;p&gt;Clients and integrations publish events that a backend receives and retains. OpenLineage does not inspect a CSV on its own or decide whether a change is compatible. It requires a reader or integration that supplies the schema, plus a system that compares observations, stores history, and alerts. Its value grows when we need to trace impact across several engines.&lt;/p&gt;&#10;&lt;h3 id="datahub-catalog-and-history-with-differences-between-editions"&gt;DataHub: catalog and history, with differences between editions&lt;/h3&gt;&#10;&lt;p&gt;&lt;a href="https://docs.datahub.com/docs/schema-history"&gt;DataHub&lt;/a&gt; provides schema history for inspecting added or removed fields and type changes. This capability is available in Core, its open source edition, and Cloud. The catalog and lineage help investigate dependencies, provided we ingest those metadata.&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://docs.datahub.com/docs/managed-datahub/observe/data-contract"&gt;Data Contracts&lt;/a&gt; group assertions about an asset. The &lt;a href="https://docs.datahub.com/docs/managed-datahub/managed-datahub-overview"&gt;official Core and Cloud comparison&lt;/a&gt; includes contracts in both editions; with Core, we can execute checks outside DataHub and publish their results, for example from dbt or Great Expectations.&lt;/p&gt;&#10;&lt;p&gt;Native &lt;a href="https://docs.datahub.com/docs/managed-datahub/observe/schema-assertions"&gt;Schema Assertions&lt;/a&gt; belong to DataHub Cloud Observe. They can require the schema to contain specified columns and types or to match exactly. They are evaluated when the ingested schema changes; compared types are high level. Cloud adds native monitoring and anomaly detection. Having the open assertion model does not mean having the commercial execution service.&lt;/p&gt;&#10;&lt;p&gt;It is an option for organizations with multiple sources and governance and discovery needs. Operating the catalog and its ingestion processes costs more than adding a library to a small ETL.&lt;/p&gt;&#10;&lt;h3 id="frictionless-start-with-the-file-that-just-arrived"&gt;Frictionless: start with the file that just arrived&lt;/h3&gt;&#10;&lt;p&gt;&lt;a href="https://framework.frictionlessdata.io/"&gt;Frictionless&lt;/a&gt; can describe, read, and validate tabular data from Python. It can inspect a CSV, infer names and types, save a schema descriptor, and validate another delivery against it. It is useful for checking files close to reception without first deploying a catalog.&lt;/p&gt;&#10;&lt;p&gt;This example uses &lt;strong&gt;Frictionless 5.20.0&lt;/strong&gt;, checked against the &lt;a href="https://framework.frictionlessdata.io/docs/framework/resource.html"&gt;Resource&lt;/a&gt;, &lt;a href="https://framework.frictionlessdata.io/docs/framework/schema.html"&gt;Schema&lt;/a&gt;, and &lt;a href="https://framework.frictionlessdata.io/docs/guides/validating-data.html"&gt;validation&lt;/a&gt; documentation. Run it in a new working directory: it creates three files. The Spanish filenames are kept so both versions use the same executable example.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;python3 -m venv .venv&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;source .venv/bin/activate&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;python -m pip install frictionless&lt;span style="color:#f92672"&gt;==&lt;/span&gt;5.20.0&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;from&lt;/span&gt; pathlib &lt;span style="color:#f92672"&gt;import&lt;/span&gt; Path&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;from&lt;/span&gt; frictionless &lt;span style="color:#f92672"&gt;import&lt;/span&gt; Resource, Schema&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Path(&lt;span style="color:#e6db74"&gt;&amp;#34;clientes.csv&amp;#34;&lt;/span&gt;)&lt;span style="color:#f92672"&gt;.&lt;/span&gt;write_text(&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;id,name,age,country&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;\n&lt;/span&gt;&lt;span style="color:#e6db74"&gt;1,John,25,USA&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;\n&lt;/span&gt;&lt;span style="color:#e6db74"&gt;2,Mary,30,UK&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;\n&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;,&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; encoding&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;utf-8&amp;#34;&lt;/span&gt;,&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Path(&lt;span style="color:#e6db74"&gt;&amp;#34;clientes-cambio.csv&amp;#34;&lt;/span&gt;)&lt;span style="color:#f92672"&gt;.&lt;/span&gt;write_text(&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;id,name,age,country&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;\n&lt;/span&gt;&lt;span style="color:#e6db74"&gt;1,John,25,USA&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;\n&lt;/span&gt;&lt;span style="color:#e6db74"&gt;2,Mary,unknown,UK&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;\n&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;,&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; encoding&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;utf-8&amp;#34;&lt;/span&gt;,&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;original &lt;span style="color:#f92672"&gt;=&lt;/span&gt; Resource(path&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;clientes.csv&amp;#34;&lt;/span&gt;, encoding&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;utf-8&amp;#34;&lt;/span&gt;)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;original&lt;span style="color:#f92672"&gt;.&lt;/span&gt;infer()&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;print([(field&lt;span style="color:#f92672"&gt;.&lt;/span&gt;name, field&lt;span style="color:#f92672"&gt;.&lt;/span&gt;type) &lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; field &lt;span style="color:#f92672"&gt;in&lt;/span&gt; original&lt;span style="color:#f92672"&gt;.&lt;/span&gt;schema&lt;span style="color:#f92672"&gt;.&lt;/span&gt;fields])&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;original&lt;span style="color:#f92672"&gt;.&lt;/span&gt;schema&lt;span style="color:#f92672"&gt;.&lt;/span&gt;to_json(&lt;span style="color:#e6db74"&gt;&amp;#34;clientes.schema.json&amp;#34;&lt;/span&gt;)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;report &lt;span style="color:#f92672"&gt;=&lt;/span&gt; Resource(&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; path&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;clientes-cambio.csv&amp;#34;&lt;/span&gt;,&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; encoding&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;utf-8&amp;#34;&lt;/span&gt;,&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; schema&lt;span style="color:#f92672"&gt;=&lt;/span&gt;Schema(&lt;span style="color:#e6db74"&gt;&amp;#34;clientes.schema.json&amp;#34;&lt;/span&gt;),&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;)&lt;span style="color:#f92672"&gt;.&lt;/span&gt;validate()&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;print(report&lt;span style="color:#f92672"&gt;.&lt;/span&gt;valid)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;print(report&lt;span style="color:#f92672"&gt;.&lt;/span&gt;flatten([&lt;span style="color:#e6db74"&gt;&amp;#34;rowNumber&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;fieldNumber&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;]))&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;It infers &lt;code&gt;id&lt;/code&gt; and &lt;code&gt;age&lt;/code&gt; as &lt;code&gt;integer&lt;/code&gt;, and &lt;code&gt;name&lt;/code&gt; and &lt;code&gt;country&lt;/code&gt; as &lt;code&gt;string&lt;/code&gt;. The second file returns &lt;code&gt;False&lt;/code&gt; and &lt;code&gt;[[3, 3, 'type-error']]&lt;/code&gt;: its third row, including the header, contains an incompatible age.&lt;/p&gt;&#10;&lt;p&gt;The key detail is &lt;strong&gt;validating against the earlier schema&lt;/strong&gt;. If we infer afresh for every delivery and accept the result, we may normalize the problem as a new type. Inference is a proposal rather than a contract: it depends on the sample and &lt;a href="https://framework.frictionlessdata.io/docs/framework/detector.html"&gt;detector configuration&lt;/a&gt;. Identifiers such as &lt;code&gt;00123&lt;/code&gt; should be declared as text when their meaning requires it.&lt;/p&gt;&#10;&lt;p&gt;Frictionless also validates constraints and supports additional checks. We would need to implement or integrate history, compatibility policy, run records, and alerts. Its readers can handle tabular JSON; nested documents require a defined representation for objects, arrays, and optional fields.&lt;/p&gt;&#10;&lt;h3 id="whylogs-observe-content-through-profiles"&gt;whylogs: observe content through profiles&lt;/h3&gt;&#10;&lt;p&gt;&lt;a href="https://github.com/whylabs/whylogs"&gt;whylogs&lt;/a&gt; generates compact profiles with statistical properties, missing values, and configurable metrics. We can retain profiles from different batches, compare them, and apply constraints to investigate changes in distributions, ranges, or proportions of nulls.&lt;/p&gt;&#10;&lt;p&gt;A profile can help show that amounts jumped in scale, for example. It does not prove they are now cents. It also records column properties that allow us to observe changes in types or presence, but a profile does not replace a contract or decide schema compatibility. Summarized or approximate statistics require interpreting the signal.&lt;/p&gt;&#10;&lt;p&gt;The library works locally. Historical storage, periodic comparisons, and alert delivery require an application or integration; the commercial WhyLabs service is an additional option. Profiles reduce the need to retain raw values, although that does not guarantee all metadata are harmless from a confidentiality perspective.&lt;/p&gt;&#10;&lt;h2 id="compare-responsibilities-not-just-checkboxes"&gt;Compare responsibilities, not just checkboxes&lt;/h2&gt;&#10;&lt;p&gt;In this table, “integration” means we must provide or connect that component. Complexity is an indicative assessment for this use case, not a performance measurement.&lt;/p&gt;&#10;&lt;table&gt;&#10;&#9;&lt;thead&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Tool&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Main purpose&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Schema changes&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Schema history&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Run traceability&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Statistical analysis&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Deployment&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Open source&lt;/th&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/thead&gt;&#10;&#9;&lt;tbody&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;OpenMetadata&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Catalog and observability&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;On ingested metadata; configured tests and alerts&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Asset and metadata versions&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Lineage and pipelines through connectors; batch detail needs instrumentation&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Configurable profiling&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Platform and ingestion: greater complexity&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Yes; additional commercial services&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;OpenLineage&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Lineage event standard&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Carries schema; comparison in integration/backend&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Depends on backend and retained events&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Emitted jobs, runs, datasets, and facets&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Can carry metrics; external calculation&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Instrumentation plus backend&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Yes&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;DataHub&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Catalog and governance&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;History in Core; native Schema Assertions in Cloud Observe&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Available in Core and Cloud&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Lineage and job metadata depending on integration&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Ingested profiles; native anomalies in Cloud&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Core platform or Cloud service&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Open Core; commercial Cloud&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Frictionless&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Describe and validate files&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Inference and validation against schema; integrate historical comparison&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Exportable descriptors; registry to build&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Run association to build&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Basic statistics and checks; no continuous statistical monitor&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Library/CLI: lower complexity&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Yes, MIT&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;whylogs&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Data profiles&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Column signals; comparison and contract to integrate&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Storable profiles; structural versioning to integrate&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Run association to integrate&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Its central function&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Library: lower complexity; additional monitoring&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Yes; WhyLabs is an additional service&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;p&gt;These tools complement one another. Frictionless can validate inputs; whylogs can profile them; OpenLineage can carry their relationship with a run; a catalog can retain shared context. We do not need all five to start.&lt;/p&gt;&#10;&lt;h2 id="do-we-need-an-llm"&gt;Do we need an LLM?&lt;/h2&gt;&#10;&lt;p&gt;An LLM can help interpret a change or propose a transformation. Many basic checks can be handled with type inference, rules, regular expressions, and statistical analysis, producing results that are easy to review.&lt;/p&gt;&#10;&lt;p&gt;Adding one introduces infrastructure cost, latency, and outputs that are not necessarily deterministic. It also requires deciding which data or metadata it may receive and validating proposed transformations. Local models exist: avoiding an external service changes infrastructure requirements but does not remove these issues. A suggested conversion should never automatically become a new business rule.&lt;/p&gt;&#10;&lt;h2 id="a-lightweight-architecture-we-could-build"&gt;A lightweight architecture we could build&lt;/h2&gt;&#10;&lt;p&gt;A Python library independent of the ETL engine could inspect CSV/JSON and maintain historical metadata without executing transformations. This is a conceptual proposal:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;CSV / JSON → Inspection → Observed schema + validation&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; Version registry&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; Batch + hash + run + commit + timestamps&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; Unknown change → Alert → Review&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;First, we would normalize the descriptor and generate a &lt;strong&gt;structural fingerprint&lt;/strong&gt;. We would need to decide whether column order matters, how nested types are represented, and which properties enter the fingerprint. We would also record the algorithm version and inference options so a detector change is not confused with a data change. A schema fingerprint and a file hash serve different purposes.&lt;/p&gt;&#10;&lt;p&gt;Next, we would record immutable observations: the same schema across several batches does not mean the same content. The approved contract would have its own identifier. An unknown structure would trigger a review; a known one would still need row validation and semantic checks.&lt;/p&gt;&#10;&lt;p&gt;Associations between batches, runs, and code would let us query which executions combined Schema C with v1.1. Those batches would be &lt;strong&gt;review candidates&lt;/strong&gt;, rather than automatic evidence of incorrect results. An explicit policy would decide whether to stop loading, isolate rows, or continue with an alert. SQLite could serve as a local starting point; process coordination and retention would require further decisions.&lt;/p&gt;&#10;&lt;h2 id="reconstructing-what-happened"&gt;Reconstructing what happened&lt;/h2&gt;&#10;&lt;p&gt;Versioning code is necessary, but insufficient to explain a pipeline. We also need to know the received data, their structure, the applied contract, and each run&amp;rsquo;s outcome. Catalogs, lineage events, and validation and profiling libraries cover different parts of that history.&lt;/p&gt;&#10;&lt;p&gt;If a file changes tomorrow without warning, could you identify which batches were processed with each schema and ETL version, and justify which ones need review?&lt;/p&gt;&#10;&lt;h2 id="article-credits"&gt;Article credits&lt;/h2&gt;&#10;&lt;p&gt;The idea and editorial supervision of this article belong to José Pérez. AI assisted its preparation, with the contributions detailed in these credits.&lt;/p&gt;&#10;&lt;table&gt;&#10;&#9;&lt;thead&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Metadata&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Value&lt;/th&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/thead&gt;&#10;&#9;&lt;tbody&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Editing&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;José Pérez&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Research&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;AI assisted&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Writing&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;OpenAI Codex&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Model&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;GPT-6&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Generation time&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Not recorded&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Human review&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Pending&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Sources verified&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Yes&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;p&gt;Sources were verified with AI assistance by checking the technical claims against official documentation. Human review of the final text is awaiting the author&amp;rsquo;s confirmation.&lt;/p&gt;&#10;</description></item></channel></rss>