Abstract
Cloud data lakes and ETL pipelines process large volumes of raw, structured, and regulated data across storage layers, transformation jobs, workflow schedulers, and analytics services. Vulnerabilities in data-processing components can expose sensitive datasets, disrupt data lineage, or allow unauthorized modification of transformation logic. This study proposes a policy-aware vulnerability risk assessment framework for cloud data lakes and ETL pipeline security. The framework evaluates dataset sensitivity, pipeline privilege scope, transformation-job vulnerabilities, storage access policies, data lineage depth, scheduler exposure, encryption status, and policy violation records. A lineage-aware scoring mechanism is used to propagate vulnerability risk from processing jobs to upstream and downstream datasets. Experiments are conducted on a cloud data-lake environment containing 14,800 datasets, 6,240 ETL jobs, 1,760 workflow DAGs, 480 analytics workspaces, 32,900 storage-policy rules, and 25,600 vulnerability findings from pipeline images, dependency libraries, and scheduler components. The proposed framework identifies 2,310 lineage-linked data-security risk paths, including vulnerable transformation jobs writing to regulated tables, overprivileged pipeline service accounts, exposed workflow schedulers, and unencrypted intermediate storage locations. It converts 25,600 raw findings into 3,140 pipeline-level remediation units based on data lineage and policy impact. The full data-lake assessment completes in 16.2 minutes, with a median scoring time of 49 ms per dataset or pipeline node. After one remediation cycle, exposed intermediate datasets decrease by 860 cases, and vulnerable jobs connected to regulated data outputs decrease from 1,120 to 437. Risk ranking remains stable across five daily pipeline snapshots, with a maximum rank shift of 38 positions among the top 500 risks. The results demonstrate that policy-aware vulnerability risk assessment can connect infrastructure vulnerabilities with data lineage, helping teams prioritize risks that affect sensitive data movement and transformation integrity.
References
Liu, H., Xu, D., Ma, Q., Xu, S., & Qiu, D. (2026). Memory Poisoning Propagation and Repair Mechanism in Multi-Agent Collaborative Environments.
Georgiades, M., & Hussain, F. (2025). An explainable ai approach for interpretable cross-layer intrusion detection in internet of medical things. Electronics, 14(16), 3218.
Jiao, Y., Wang, A., Zhao, B., & Shi, T. (2026). The Impact of Visual Language Strategies in Public Art Creation on Community Spatial Perception and Public Behavior.
Adjewa, F., Esseghir, M., & Merghem-Boulahia, L. (2026). From Edge Transformer to IoT Decisions: Offloaded Embeddings for Lightweight Intrusion Detection. Sensors, 26(2), 356.
Abdulganiyu, O. H., Fadi, O., Moukafih, Y., Ait Tchakoucht, T., Saheed, Y. K., Chukwuere, J. E., & Yau, S. (2026). Explainable attention based few shot LSTM for intrusion detection in imbalanced cyber physical system networks. Scientific Reports.
Qi, C., & Qiao, X. (2026, May). Design Patterns for Multi-Agent Systems in Production. In 2026 International Conference on Intelligent Systems, Automation and Control (ISAC) (pp. 198-202). IEEE.
Kayit, A. D., & Ismail, M. T. (2025). Leveraging hybrid ensemble models in stock market prediction: A data-driven approach. Data Science in Finance and Economics, 5(3), 355.
Su, D., & Dong, Y. (2026). Classroom-Based Assessment with Bayesian Learning Analytics for Instructional Decision-Making in ASD Inclusive Education.
Suwannaphong, T., Jovan, F., Craddock, I., & McConville, R. (2025). Optimising TinyML with quantization and distillation of transformer and mamba models for indoor localisation on edge devices. Scientific Reports, 15(1), 10081.
Gao, G., Gao, R., Lu, C., Gao, R., & Kuang, Y. (2026, March). Security Governance Methods and Quantitative Evaluation for Enterprise SMS and Verification Code Systems. In 2026 International Conference on Generative Artificial Intelligence and Information Security (GAIIS) (pp. 455-458). IEEE.
Alatawi, M. N. (2025). SAFEL-IoT: secure adaptive federated learning with explainability for anomaly detection in 6G-enabled Smart Industry 5.0. Electronics, 14(11), 2153.
Li, Y., & Liu, S. (2026, May). A Study on Dynamic Optimization of Alerting Policies and Multi-Agent Decision-Making Mechanisms in Cloud Environments. In 2026 7th International Seminar on Artificial Intelligence, Networking and Information Technology (AINIT) (pp. 703-706). IEEE.
Mohale, V. Z., & Obagbuwa, I. C. (2025). A systematic review on the integration of explainable artificial intelligence in intrusion detection systems to enhancing transparency and interpretability in cybersecurity. Frontiers in Artificial Intelligence, 8, 1526221.
Zhao, J., Fan, J., & Li, L. (2026). A Study on an Explainable Causal-Enhanced LLM Agent for Predicting the Forming Quality of Automotive Component Materials.
Naif Alatawi, M. (2025). Enhancing intrusion detection systems with advanced machine learning techniques: an ensemble and explainable artificial intelligence (AI) approach. Security and Privacy, 8(1), e496.
Xu, T., Zhu, W., & Zhang, J. (2026). A Study on the Application of Alternative Data in Credit Assessment for the Unbanked Population.
Adiputra, I. N. M., Wanchai, P., & Lin, P. C. (2025). Optimized customer churn prediction using tabular generative adversarial network (GAN)-based hybrid sampling method and cost-sensitive learning. PeerJ Computer Science, 11, e2949.
Yan, B. (2026). Robust Day-Night Image Matching Across Extreme Illumination Variations: A Comparative Study of Deep Learning and Classical Methods. International Academic Journal of Engineering and Technology Science, 2, 65-72.
Liao, T., Fang, X., Feng, Y., & Wang, S. (2026). Reliable and Efficient Decision-Making for Large Language Model Agents through Uncertainty-Guided Adaptive Reasoning. Available at SSRN 6844659.
Eren, K. K., Kucuk, K., Saleh, R. A., Konyar, M. Z., Hardy, O. M., & Khan, S. A. (2026). Distributional drift in IoT intrusion detection systems: implications for cross-dataset generalization. Electronics, 15(11).
Yin, J., Rao, H., & Huang, Y. (2026, March). Dynamic Modeling and Heterogeneity Analysis of Platform User Behavior Time Series. In 2026 International Conference on AI in Education Technology and Applications (AIETA) (pp. 30-33). IEEE.
Panopio, A. J. N., Abushahla, H. A., Sajun, A. R., Alawnah, S., Aloul, F., & Zualkernan, I. (2026). From Sensor to Server: Deployable Lightweight ML for IoT Intrusion Detection Across Network Layers. IEEE Internet of Things Journal.
Yang, J. (2026). Computational Analysis of How Digital Non-Clinical Communication Tools Influence Social Participation Among Older Adults in Community-Based Elderly Care. Available at SSRN 6682838.
Kiliç, C., & Şengül, G. (2026). SHAP-Guided Feature Selection for Cross-Dataset Generalization in Network Intrusion Detection Systems. IEEE Access.
Zhang, Z. (2026). A Study on the Identification of Manipulative Design in Subscription and Payment Interfaces of Digital Consumer Platforms and Its Behavioral Effects. Available at SSRN 6734760.
Sorostinean, R., Neghina, C., & Gellert, A. (2025). Boosting anomaly detection with unsupervised K-Means and SOM for energy-efficient factory machines. Journal of Intelligent Manufacturing, 1-17.
Zhang, Z., Tong, Y., & Gao, Y. (2026). Retrieval-Augmented Generation with Low-Latency Deployment for Vertical Domains Question Answering: A Case Study on Economic Resource Platforms.
Ahmed, A. M., Sadiq, A. T., Farhan, A. K., Ahmad, M., & Shakeel, M. R. (2026). A Comprehensive Survey of Lightweight Large Language Models: Architectures, Techniques, and Challenges. Archives of Computational Methods in Engineering, 1-28.
Qi, C., & Qiao, X. (2026). Efficient Data Sampling and Feature Selection Algorithms for Scalable Machine Learning Pipelines.
Potharaju, S., Tambe, S. N., Tirandasu, R. K., Kumar, D. A., & Kantipudi, M. P. (2026). Enhancing Cybersecurity in Industrial Internet of Things Systems Using Ensemble Learning Against False Data Injection Attacks. Journal of Current Science and Technology, 16(1), 151-151.
