Prompt Lineage and Governance in LLM-Enabled DataEngineering: A Reference Architecture

Authors

  • Shambhu Adhikari Sr. Data Engineer - United Airlines Author

Keywords:

Large Language Models, Data Engineering, Prompt Governance, Prompt Lineage, LLMOps 

Abstract

Background and Purpose: Large Language Models (LLMs) are increasingly embedded within modern data engineering ecosystems to automate data transformation, quality assurance, metadata generation, and analytical reasoning. Although these models enhance productivity and adaptability, their probabilistic behavior and reliance on prompts as executable control artifacts introduce significant governance challenges. In many LLM-enabled systems, prompts are embedded within orchestration layers without formal lifecycle management, lineage tracking, or policy enforcement, limiting reproducibility, auditability, and regulatory compliance. This study proposes a reference architecture for prompt lineage and governance in LLM-enabled data engineering environments.

Methods: A design science research approach was used to develop the reference architecture through structured synthesis of principles from DataOps, MLOps, metadata management, data lineage, and responsible AI. The architecture models prompts as governed assets and incorporates prompt registry and versioning, context-aware execution, lineage and metadata integration, governance and policy enforcement, and observability. Scenario-based validation was conducted across LLM-assisted data quality validation, automated schema and metadata enrichment, and natural-language-driven analytical query generation.

Findings: The proposed architecture provided end-to-end traceability across data assets, prompt versions, execution contexts, model configurations, and downstream outputs. Across the evaluated scenarios, it enabled prompt version traceability, recovery of execution context, governance controls, and operational monitoring. The framework improved reproducibility and auditability, supported policy enforcement and compliance readiness, and reduced the difficulty of root-cause analysis compared with unmanaged prompt use.

Theoretical Contributions: This study extends established data lineage, DataOps, and MLOps concepts by treating prompts as firstclass governed assets within enterprise data engineering. It introduces a prompt-aware lineage model that links data, prompts, models, execution metadata, and outputs, providing a structured foundation for prompt lifecycle governance and the emerging practice of LLMOps.

Conclusions and Policy Implications: Prompt lineage and governance should be treated as foundational capabilities for scalable and trustworthy LLM-enabled data engineering. Organizations adopting LLMs in operational data workflows should implement version control, metadata capture, access control, validation, audit logging, and continuous observability for prompts. Embedding these controls within existing data governance infrastructure can strengthen reproducibility, accountability, compliance, and operational resilience while preserving development flexibility.

Downloads

Published

2022-09-15

Similar Articles

1-10 of 17

You may also start an advanced similarity search for this article.