Large Language Models in Modern Data Engineering: A Systematic Review of Architectures, Use Cases, and Limitations

Authors

  • Shambhu Adhikari Sr. Data Engineer - United Airlines, NW, NJ Author

Keywords:

Large language models, Data engineering, Retrieval-augmented generation, Data pipelines, Model governance

Abstract

Background and Purpose: Large language models (LLMs) are increasingly integrated into modern data engineering workflows, including data ingestion, schema inference, metadata generation, transformation logic synthesis, data quality monitoring, and natural-language interaction with analytical systems. Their probabilistic behavior introduces new reliability, governance, and reproducibility concerns. This review examines the architectural patterns, practical data-engineering use cases, and operational limitations associated with LLM integration.

Methods: A PRISMA-informed systematic review was conducted across IEEE Xplore, ACM Digital Library, Scopus, Web of Science, PubMed, arXiv, and selected industrial and technical reports. For the historical 2023 edition, the evidence window was restricted to studies available from January through November 2022. Sources were screened for relevance to LLM-enabled data engineering, architecture, practical deployment, and governance.

Findings: The review identifies direct LLM integration, Retrieval-Augmented Generation (RAG), agent-based or tool-calling systems, and hybrid governed architectures as the principal integration patterns. Across the source synthesis, RAG was the most frequently reported architecture, while common use cases included transformation code generation, metadata and documentation support, conversational analytics, ingestion assistance, and governance-oriented checks. Reported benefits included productivity, accessibility, flexibility, and knowledge reuse, while recurring limitations involved hallucination, privacy, cost and scalability, explainability, and governance gaps.

Theoretical Contributions: This review organizes LLM-enabled data engineering into a structured architectural taxonomy and links deployment patterns with stages of the data engineering lifecycle. It integrates foundational work on transformers, retrieval, production machine learning, natural-language data interfaces, security, and explainability to frame LLM adoption as an architectural and governance problem rather than a model-only problem.

Conclusions and Policy Implications: LLMs are most appropriate as governed, assistive components embedded within deterministic data platforms rather than as fully autonomous replacements for production pipelines. Organizations should prioritize retrieval grounding, human validation, access control, lineage, observability, cost monitoring, and privacy safeguards. Responsible adoption requires explicit architectural boundaries and governance mechanisms that preserve reliability, auditability, and operational control.

Downloads

Published

2026-08-26

Similar Articles

1-10 of 20

You may also start an advanced similarity search for this article.