Generative AI and Large Language Models in Banking and Financial Services: A Systematic Review
Abstract
The emergence of Generative Artificial Intelligence (GenAI) and Large Language Models (LLMs) is creating new opportunities for automation, knowledge processing, and intelligent decision support in banking and financial services. Unlike conventional machine learning systems designed for specific predictive tasks, generative models can process and produce natural language, summarize complex financial documents, support customer interactions, and assist employees with knowledge-intensive activities. This review examines the emerging applications of GenAI and LLMs across banking, including financial document analysis, customer service, regulatory reporting, risk analysis, financial research, code generation, knowledge management, and compliance operations. The paper evaluates the potential benefits of LLM-based systems in improving operational efficiency, personalization, accessibility, and employee productivity. At the same time, significant risks involving hallucination, data leakage, privacy, cybersecurity, model bias, explainability, and regulatory compliance are examined. The review also discusses retrieval-augmented generation, domain-specific language models, human-in-the-loop systems, and private enterprise AI architectures as approaches for improving reliability in financial environments. Future research directions are identified around trustworthy GenAI, financial-domain benchmarking, model governance, explainability, and secure deployment. The review emphasizes that responsible implementation and strong governance will be essential for realizing the benefits of generative AI while managing its risks in highly regulated financial institutions.
References
Sandra, K. (2025). Master data management in multi-cloud environments: A survey with operational evidence from banking and insurance deployments. International Journal of Emerging Research in Engineering and Technology, 6(3), 152–157.
Baesens, B., Van Gestel, T., Viaene, S., Stepanova, M., Suykens, J., & Vanthienen, J. (2003). Benchmarking state-of-the-art classification algorithms for credit scoring. Journal of the Operational Research Society, 54(6), 627–635. https://doi.org/10.1057/palgrave.jors.2601545
Paruchuri, J. K. (2022). Real-time fraud detection and feature store design patterns for streaming ML in financial services. International Journal of Computer Science Engineering Techniques (IJCSE), 6(2), 1–31.
Ozbayoglu, A. M., Gudelek, M. U., & Sezer, O. B. (2020). Deep learning for financial applications: A survey. Applied Soft Computing, 93, 106384. https://doi.org/10.1016/j.asoc.2020.106384
Sandra, K. (2024). Large language models for data catalog enrichment: A survey with operational evidence from enterprise deployments. International Journal of Computer Science Engineering Techniques, 12(3), 1–8.
Khandani, A. E., Kim, A. J., & Lo, A. W. (2010). Consumer credit-risk models via machine-learning algorithms. Journal of Banking & Finance, 34(11), 2767–2787. https://doi.org/10.1016/j.jbankfin.2010.06.001
Paruchuri, J. K. (2021). Lakehouse architecture: Unifying data lakes and data warehouses. International Journal of Computer Techniques (IJCT), 8(1), 1–15.
Ngai, E. W. T., Hu, Y., Wong, Y. H., Chen, Y., & Sun, X. (2011). The application of data mining techniques in financial fraud detection: A classification framework and an academic review of literature. Decision Support Systems, 50(3), 559–569. https://doi.org/10.1016/j.dss.2010.08.006
Sandra, K. (2022). Intelligent data workbench design for multi-language code compatibility. International Journal of Computer Science Engineering Techniques, 12(1), 1–17.
Bussmann, N., Giudici, P., Marinelli, D., & Papenbrock, J. (2021). Explainable machine learning in credit risk management. Computational Economics, 57, 203–216. https://doi.org/10.1007/s10614-020-10042-0
Paruchuri, J. K. (2021). Exactly-once semantics in distributed stream processing at scale. International Journal of Computer Science Engineering Techniques (IJCSE), 5(1), 1–15.
Lessmann, S., Baesens, B., Seow, H.-V., & Thomas, L. C. (2015). Benchmarking state-of-the-art classification algorithms for credit scoring: An update of research. European Journal of Operational Research, 247(1), 124–136. https://doi.org/10.1016/j.ejor.2015.05.030
Paruchuri, J. K. (2023). Design and performance evaluation of an elastic multi-tenant Spark SQL platform using Apache Kyuubi on Kubernetes. Journal of Advanced Research in Technology and Management Sciences, 5(3), 33–45.
Abdallah, A., Maarof, M. A., & Zainal, A. (2016). Fraud detection system: A survey. Journal of Network and Computer Applications, 68, 90–113. https://doi.org/10.1016/j.jnca.2016.04.007
Barredo Arrieta, A., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., García, S., Gil-López, S., Molina, D., Benjamins, R., Chatila, R., & Herrera, F. (2020). Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Information Fusion, 58, 82–115. https://doi.org/10.1016/j.inffus.2019.12.012
Altman, E. I., Marco, G., & Varetto, F. (1994). Corporate distress diagnosis: Comparisons using linear discriminant analysis and neural networks. Journal of Banking & Finance, 18(3), 505–529. https://doi.org/10.1016/0378-4266(94)90007-8
Fuster, A., Goldsmith-Pinkham, P., Ramadorai, T., & Walther, A. (2022). Predictably unequal? The effects of machine learning on credit markets. The Journal of Finance, 77(1), 5–47. https://doi.org/10.1111/jofi.13090
West, J., & Bhattacharya, M. (2016). Intelligent financial fraud detection: A comprehensive review. Computers & Security, 57, 47–66. https://doi.org/10.1016/j.cose.2015.09.005
Leo, M., Sharma, S., & Maddulety, K. (2019). Machine learning in banking risk management: A literature review. Risks, 7(1), 29. https://doi.org/10.3390/risks7010029
Arner, D. W., Barberis, J. N., & Buckley, R. P. (2016). The evolution of FinTech: A new post-crisis paradigm? Georgetown Journal of International Law, 47(4), 1271–1319.