Schema Drift in ETL: when data changes and your code does not notice
Every morning, an ETL receives a CSV, converts ages to integers, and loads customers. Yesterday it worked with this file:
id,name,age
1,John,25
2,Mary,30
Today, the provider adds a country:
id,name,age,country
1,John,25,USA
2,Mary,30,UK
A few days later, another surprise arrives:
id,name,age,country
1,John,25,USA
2,Mary,unknown,UK
An additional column may be harmless if we select fields by name and allow extra columns. If we unpack each row into three variables, it may break the load. unknown may cause an age conversion error, silently become null, or lead a subsequent inference to treat the entire column as text.
