Data Lineage Preservation Under Automated Schema Evolution
Keywords:
Data lineage, schema evolution, graph-based modeling, traceability.Abstract
Data lineage preservation in modern data pipelines has become increasingly complex due to the widespread adoption of automated schema inference and continuous schema evolution. Existing approaches primarily assume stable schemas and direct attribute mappings, which are insufficient for handling dynamic structural changes across multi-stage data transformations. This study addresses this limitation by proposing a graph-based lineage modeling framework that captures node-level dependencies and transformation semantics while incorporating schema-aware mapping strategies. The analysis reveals that lineage degradation is stage-dependent, with transformation and aggregation layers exhibiting higher susceptibility to schema-induced disruptions. A distribution-based evaluation further demonstrates that adaptive reconstruction mechanisms can effectively restore lineage continuity, even when direct mappings are lost. The findings highlight that lineage preservation must be treated as a dynamic reconstruction problem rather than a static tracking task. The proposed approach improves traceability, enhances data integrity, and supports compliance requirements in AI-driven data architectures, offering practical relevance for enterprise-scale data systems operating under evolving schema conditions.