Generative AI for Scenario-Based Financial Stress Testing and Risk Management
Keywords:
generative artificial intelligence; financial stress testing; risk management; scenario generation; model governance; systemic risk; regulatory technology 1 Introduction The global financial crisis of 2007 to 2009 transformed stress testing from a largely internal risk management exercise into a prominent supervisory and macroprudential instrument. The Basel Committee on Banking Supervision has formalized stress testing as a core supervisory expectation, requiring banks to assess capital adequacy under severe but plausible adverse conditions and to integrate the results into capital planning and risk appetite frameworks [1]. This regulatory turn has made the scenario generation process a matter of institutional and systemic importance. The scenarios selected by supervisors and firms influence capital buffers, risk limits, and strategic decisions under uncertainty. However, the conceptual foundations of scenario design have evolved more slowly than the regulatory expectations surrounding them. Schuermann observed that stress tests can produce misleading confidence when the underlying scenarios are insufficiently severe or coherent with respect to the firm's actual exposures [2]. The difficulty is not merely technical but structural: a stress scenario must be extreme enough to reveal vulnerabilities, coherent enough to preserve economic relationships, and diverse enough to capture complex tail dynamics. Macroprudential research has shown that systemic risk emerges from interconnected balance sheets, common exposures, and endogenous feedback effects that are difficult to represent in a small number of expert-defined scenarios [3]. The practical challenge of selecting stress paths with tail dependence has motivated empirical likelihood methods that choose scenarios according to their plausibility and severity [4]. Bayesian approaches have further allowed stress testing to incorporate judgment while maintaining probabilistic coherence [5]. Yet these methods remain constrained by the parametric assumptions, factor structures, and historical data windows upon which they are built. They are generally better at refining a known scenario space than at discovering novel but plausible configurations of stress. This is precisely where generative artificial intelligence offers a different kind of analytical contribution. Rather than imposing a low-dimensional structure on the world, generative models can learn high-dimensional joint distributions over market variables, credit factors, liquidity indicators, and macroeconomic conditions and then sample synthetic states from those distributions. The resulting scenarios can expose risk concentrations that fixed-factor analysis may miss. The integration of generative AI into financial stress testing raises a series of system-level questions. These include the design of model architectures, the construction of data pipelines, the governance of synthetic data, the validation of generated scenarios, the fairness implications of automated risk inference, and the energy and operational costs of large-scale generative systems. This paper addresses these questions through an interdisciplinary systems perspective. It argues that generative AI should be understood not as a substitute for established stress testing disciplines but as a socio-technical infrastructure that reshapes how scenarios are produced, contested, and governed. The remaining sections examine the evolution and constraints of conventional scenario design, the capabilities of generative architectures, the supporting data and computational infrastructures, the regulatory and validation challenges, and the broader implications for robustness, fairness, sustainability, and policy. 2 The Evolution and Structural Constraints of Scenario-Based Stress Testing Scenario-based stress testing has passed through several distinct phases. Early implementations focused on simple sensitivity analyses and historical scenarios, in which a past crisis was replayed against current positions. Supervisory programs such as the Comprehensive Capital Analysis and Review in the United States and the European Banking Authority stress tests subsequently standardized the use of common macroeconomic scenarios across firms. These programs improved comparability and strengthened the link between stress results and capital adequacy. However, they also institutionalized a particular mode of scenario production that relies heavily on expert committees, reduced-form macro-financial models, and a small set of baseline, adverse, and severely adverse paths. The structural consequence is that the scenario space is often narrow relative to the true dimensionality of financial risk. A central limitation of conventional scenario design is that it struggles to represent nonlinear interactions and regime shifts. Financial systems exhibit thresholds, fire sales, liquidity spirals, and confidence effects that are not well captured by linear or mildly nonlinear models estimated over normal periods. When scenarios are constructed from a limited number of macroeconomic variables and mapped to firm-level losses through satellite models, the mapping can understate the joint severity of market, credit, and funding shocks. Glasserman and coauthors identified the problem of selecting stress scenarios that are both extreme and plausible, showing that the choice of scenario probabilities and severity measures materially affects the risk assessment [4]. Bayesian coherent stress testing addresses some of these issues by treating scenario construction as an inference problem, but its effectiveness depends on the prior distributions and dependence structures supplied by experts [5]. These approaches remain fundamentally analytical in orientation; they do not autonomously generate new combinations of stress factors from high-dimensional data. Another structural constraint arises from the organizational separation between scenario design and model execution. In many institutions, macroeconomic scenario teams, risk modeling teams, and business lines operate under different governance structures and often with different assumptions about model behavior. This separation can produce scenarios that are internally inconsistent or that fail to connect with the institutional vulnerabilities that matter most. Moreover, supervisory scenarios are necessarily standardized, and firm-specific scenarios may be insufficiently challenged because they are developed within existing risk taxonomies. The result is a scenario generation system that is conservative in form but potentially fragile in content, precisely because it cannot systematically explore the large space of possible stress configurations. Generative AI has the potential to alter this equilibrium by offering a machine-readable representation of high-dimensional scenario spaces that can be sampled, audited, and updated as new data arrive. However, realizing that potential requires more than algorithmic sophistication; it requires a rethinking of how scenario knowledge is represented and governed. 3 Generative AI Architectures for Scenario Generation and Risk Inference Generative AI introduces a family of architectures capable of learning distributions over complex financial and macroeconomic data. Adversarial generative models have been particularly influential because they frame the learning process as a contest between a generator that produces synthetic samples and a discriminator that attempts to distinguish them from real observations [6]. This adversarial framework can be extended to conditional generation, in which scenarios are produced subject to specified macroeconomic or market conditions. The resulting scenarios are not constrained by a fixed factor model and can capture nonlinear dependencies that are difficult to encode analytically. Variational autoencoders provide an alternative generative framework based on probabilistic latent variable representations [7]. In stress testing, such representations can be used to encode high-dimensional market states into lower-dimensional latent factors and then generate new states by interpolating or extrapolating within the latent space. Diffusion models, which progressively denoise random inputs into structured outputs, have recently emerged as another powerful approach for synthesizing realistic time series and panel data [8]. These architectures differ in their training dynamics, sample quality, and ease of conditional control, but they share a common capacity to learn from large datasets without requiring a complete pre-specified structural model. The application of generative models to financial scenario generation is not limited to market data. Transformer-based architectures have shown substantial capacity for modeling sequences and high-dimensional dependencies across modalities [9]. Large language models trained on broad text corpora can support narrative scenario construction, extract risk-relevant signals from unstructured documents, and assist analysts in articulating plausible stress narratives [10]. Foundation models, which are trained on broad data and adapted to downstream tasks, offer a further layer of abstraction that could support scenario generation, risk classification, and causal querying within a unified representational space [11]. However, the use of foundation models in financial risk management raises important questions about domain adaptation, factual reliability, and the stability of generated outputs under distribution shift. The financial domain is characterized by evolving regulatory definitions, changing market structures, and feedback effects from the models themselves, which may differ significantly from the general domains used to pretrain foundation models. From a systems perspective, the key architectural trade-off is between expressiveness and controllability. Highly expressive generative models can produce novel scenarios that challenge existing risk taxonomies, but they may also generate outputs that are economically implausible or operationally unusable. Conversely, models with strong conditioning and regularization may remain too close to historical patterns to provide genuine stress discovery. Effective deployment therefore requires an orchestration layer that combines generative sampling with economic constraints, plausibility filters, and supervisory review. This orchestration is not simply a technical pipeline; it is an institutional mechanism through which generated scenarios are negotiated, interpreted, and integrated into decision processes. The generative model is one component of a broader scenario production system, and its value depends on how well its outputs are embedded within the governance and validation structures of the organization. 4 System Architecture and Data Infrastructure The deployment of generative AI for financial stress testing requires a data and computational infrastructure that differs substantially from traditional risk modeling environments. Financial stress testing depends on longitudinal data across multiple frequencies, including daily market prices, monthly macroeconomic indicators, quarterly firm-level financial statements, and high-frequency liquidity metrics. These data are heterogeneous in quality, coverage, and vintage, and they are subject to revisions that can change the meaning of historical observations. A generative scenario engine must therefore be connected to a data infrastructure that supports lineage tracking, temporal consistency, and point-in-time retrieval. Without such capabilities, the model may learn spurious relationships from data that reflect backfilled or as-of values that would not have been available at the time of the stress event. This data governance challenge is often underestimated in proofs of concept but becomes critical in supervisory and capital planning contexts. Computational infrastructure also shapes the feasible space of generative modeling. Training large generative models requires significant accelerator capacity, distributed storage, and efficient orchestration of data loading and checkpointing. Inference for scenario generation is comparatively less expensive but may still be nontrivial when large numbers of conditional scenarios are required across different business lines and legal entities. Institutions must decide whether to train models in-house, fine-tune open-source foundation models, or access external model services. Each choice carries different implications for control, explainability, data confidentiality, and operational risk. The rise of foundation models has made it possible to benefit from large-scale pretraining without owning the full training pipeline, but it also introduces dependencies on external model providers and reduces transparency into the data and objectives used during pretraining [11]. These dependencies must be managed through contractual safeguards, technical audits, and clear model ownership boundaries. The system architecture should distinguish between exploratory scenario generation and production-grade stress analytics. Exploratory environments can tolerate greater model volatility because their purpose is to surface novel risk configurations for expert review. Production environments, by contrast, require reproducible inputs, versioned models, and deterministic post-processing so that results can be audited and compared across reporting cycles. A layered architecture that separates data ingestion, feature engineering, generative sampling, scenario filtering, loss projection, and reporting can support both modes while maintaining control over the provenance of each scenario. This architecture must also support feedback loops through which model outputs are evaluated against realized outcomes and used to refine the scenario generation process. Machine behavior research suggests that such feedback can produce emergent behaviors that are not obvious from the model specification alone [12]. Monitoring generated scenario distributions over time is therefore essential to detect drift, mode collapse, or unintended narrowing of the scenario space. 5 Governance, Validation, and Regulatory Alignment The governance of generative AI in financial stress testing must address both model risk management and the specific challenges of generative outputs. Traditional model risk guidance emphasizes conceptual soundness, data quality, ongoing monitoring, and independent validation. Generative models complicate each of these dimensions because their outputs are stochastic, their internal representations are difficult to inspect, and their failure modes may not be captured by standard backtesting metrics. Validation must therefore move beyond point accuracy and examine the distributional properties of generated scenarios, their economic plausibility, their coverage of tail events, and their stability under perturbations of the input data or conditioning variables. The potential for generative models to produce confident but misleading outputs has been widely documented in the context of large language models [13]. In stress testing, similar risks arise when a model generates a coherent-looking scenario that embeds hidden inconsistencies or implausible assumptions. Regulatory alignment remains an open issue. Supervisory authorities have issued principles and reports on the use of artificial intelligence in finance, emphasizing the need for accountability, transparency, and sound data governance [15]. The European Commission has proposed a risk-based regulatory framework for artificial intelligence that classifies certain uses as high risk and imposes obligations related to data governance, technical documentation, human oversight, and robustness [16]. Stress testing models that influence capital adequacy and systemic risk assessment are likely to fall within a high-risk category under any reasonable taxonomy. This means that institutions must be able to demonstrate not only that generative models perform well in aggregate but also that they can be audited at the level of individual scenarios and model versions. Explainable artificial intelligence methods offer partial support by identifying the features and training examples that most influence a given output [17]. However, the interpretability of a generative scenario is different from the interpretability of a classifier or regression model, and Lipton has cautioned against treating post hoc explanations as substitutes for intrinsic model transparency [18]. Governance frameworks must therefore combine technical explanation with domain expert review, model documentation, and structured challenge processes. A further governance challenge concerns the distinction between scenario generation and loss projection. In many risk frameworks, the scenario defines the external environment, while the loss projection model translates that environment into financial outcomes. If a generative model is used to produce the scenario but a separate model is used to estimate losses, the interface between the two must be carefully specified. Inconsistencies in variable definitions, timing conventions, or accounting treatments can produce results that are difficult to reproduce and defend. Model risk management should therefore treat the generative scenario engine and the downstream loss models as a coupled system and evaluate their joint behavior under stress. This coupled perspective is more demanding than validating each component in isolation, but it better reflects the actual operating environment of stress testing. 6 Robustness, Fairness, and Operational Resilience Robustness is a first-order concern for generative AI in financial stress testing. Financial time series are nonstationary, and relationships that hold in one regime may reverse in another. A generative model trained primarily on data from a prolonged expansion may underestimate the frequency and severity of crisis dynamics. Adversarial testing can help identify sensitivities to input perturbations and conditioning assumptions, but formal guarantees are difficult to provide in high-dimensional, nonstationary environments. The machine behavior literature emphasizes that algorithmic systems can produce unforeseen behaviors when embedded in social and economic contexts [12]. In stress testing, such behaviors may manifest as a systematic narrowing of generated scenarios toward historically frequent patterns, or as an overrepresentation of extreme but economically irrelevant states. Ongoing monitoring of scenario diversity, tail coverage, and conditional calibration is therefore necessary. Scenario databases should be examined not only for individual plausibility but also for collective coverage of the risk surface. Fairness is an equally important but less discussed dimension of generative stress testing. The outputs of stress scenarios influence capital allocation, credit availability, and risk pricing, all of which have distributional consequences for households, firms, and regions. If generative models are trained on historical data that encode existing inequalities, they may reproduce or amplify those inequalities in synthetic scenarios. Fairness in machine learning has largely focused on classification and prediction, but the concepts translate imperfectly to generative risk modeling [20]. A survey of bias and fairness research shows that bias can enter at many stages, including data collection, representation learning, and evaluation [21]. For stress testing, a key question is whether generated scenarios systematically disadvantage certain portfolios, customer segments, or geographic areas, and whether those disadvantages reflect legitimate risk differences or artifacts of the data and model. Addressing this requires disaggregated evaluation, fairness-aware scenario audits, and the involvement of risk committees that include nontechnical perspectives on the social implications of stress scenarios. Policy-oriented research has similarly emphasized that algorithmic decisions are not merely technical predictions but interventions in social systems [19]. Operational resilience concerns the ability of institutions to maintain scenario generation capabilities under disruption. Generative AI systems may depend on specialized hardware, cloud services, external model providers, and scarce technical personnel. A failure in any of these dependencies could delay stress testing, impair regulatory submissions, or force reliance on less capable fallback models. Institutions must therefore design generative scenario engines with graceful degradation paths. They should maintain the ability to run conventional scenario generation methods when advanced generative systems are unavailable and should document the conditions under which fallback is triggered. This operational requirement interacts with model governance, because the fallback model will generally have different risk characteristics and may require separate validation. Operational resilience also extends to cybersecurity and data integrity, since the value of generative scenario engines depends on the trustworthiness of the training data and conditioning inputs. A compromised or manipulated scenario generation system could produce outputs that appear plausible but are designed to hide or exaggerate specific risks, which would be a new form of model risk. 7 Sustainability and Deployment Trade-Offs The sustainability of generative AI in stress testing has both environmental and institutional dimensions. Training large generative models is computationally intensive and can impose significant energy costs, as documented in studies of natural language processing pipelines [22]. While financial institutions may train smaller domain-specific models or fine-tune existing foundation models, the cumulative energy demand across a global financial system with many institutions and frequent retraining cycles could be substantial. This environmental footprint must be weighed against the risk management benefits of improved scenario exploration. Institutional sustainability is a broader concern. Generative AI systems require continuous investment in data engineering, model maintenance, validation, and expert oversight. If these investments are not sustained, the models may degrade in quality over time, producing scenarios that are less relevant to current market conditions. The financial sector has experienced cycles of enthusiasm for advanced analytics followed by cost rationalization, and generative AI systems are vulnerable to the same dynamics. A sustainable deployment strategy should therefore prioritize modular architectures, reusable data assets, and clear ownership of the model lifecycle. Deployment trade-offs also arise between centralization and decentralization. Centralized scenario generation can ensure consistency, comparability, and efficient use of computational resources, but it may reduce the diversity of perspectives that is valuable for risk identification. Decentralized deployment across business lines and jurisdictions can support local relevance but may lead to fragmented governance, inconsistent assumptions, and duplicated costs. A hybrid approach may be most effective, in which a central platform provides base generative models, data standards, and validation tools, while business units and subsidiaries contribute domain-specific conditioning and review. Such a platform must be designed with strong technical boundaries so that local adaptations do not compromise the integrity of the central model. The architecture should also allow for controlled experimentation, because the exploration of new scenario families is essential for the long-term value of the system. A purely control-oriented deployment that suppresses novelty may fail to realize the central benefit of generative AI, while a purely exploratory deployment may produce an unmanageable volume of low-quality scenarios. The choice of modeling paradigm also influences sustainability. Large autoregressive models trained on broad text may be attractive because of their flexibility, but they can be inefficient for structured financial data and difficult to align with supervisory expectations. Smaller generative architectures trained specifically on risk-relevant data may offer better interpretability, lower energy use, and stronger domain alignment, albeit with less generative diversity. The recent literature on aligning language models with human feedback illustrates the broader tension between capability and control [14]. In stress testing, alignment must extend beyond human preference to include economic plausibility, regulatory consistency, and institutional risk appetite. This suggests that the deployment of generative AI should be treated as an ongoing socio-technical negotiation rather than a one-time implementation. The optimal system at one point in time may become suboptimal as market structures, regulatory expectations, and institutional capabilities evolve. 8 Policy Implications and Future Directions The policy implications of generative AI in stress testing extend beyond model risk guidance to the organization of financial supervision itself. Supervisors may need to develop their own generative scenario capabilities in order to evaluate the outputs of regulated institutions and to generate independent views of systemic risk. If only large financial institutions possess advanced scenario generation tools, the asymmetry in analytical capacity could influence the supervisory dialogue. A shared or open scenario infrastructure could mitigate this asymmetry, but it would raise difficult questions about data confidentiality, intellectual property, and the appropriate boundaries between supervisory and private sector modeling. The machine behavior perspective suggests that the interaction between supervisory models and institutional models can produce feedback loops that affect market behavior and risk taking [12]. Supervisors should therefore consider how the introduction of generative models changes the strategic environment in which stress testing operates. Regulatory frameworks will need to evolve to accommodate generative methods without prematurely codifying specific technical choices. Principle-based approaches that emphasize outcomes such as model explainability, scenario diversity, and operational soundness may prove more durable than rules tied to particular architectures [16]. At the same time, principles alone may not be sufficient to ensure comparability across institutions. Standardized reporting of generative scenario methodology, training data provenance, and validation results could help supervisors compare different implementations. International coordination is also important because financial stress can propagate across borders and because large institutions operate under multiple regulatory regimes. Divergent national rules on artificial intelligence could fragment the global stress testing landscape and create regulatory arbitrage. Policy-oriented research has emphasized the need to align technical capabilities with broader social goals and institutional accountability [19]. The integration of generative AI into stress testing should therefore be accompanied by public consultation, interdisciplinary review, and careful attention to the distributional effects of risk decisions. Future research should explore the use of generative models for multi-agent and network-based stress simulations that go beyond scenario generation in the narrow sense. Financial crises arise from interactions among heterogeneous agents, institutions, and markets, and generative models could support simulation environments in which those interactions are represented more explicitly. Another important direction is the development of uncertainty quantification methods that are compatible with generative outputs. Supervisors and risk managers need to know not only whether a generated scenario is severe but also how much confidence can be placed in its plausibility and relevance. This requires progress in distribution-free inference, robustness evaluation, and calibration under regime shift. The continued evolution of foundation models also suggests that scenario generation may become more integrated with textual analysis, regulatory interpretation, and automated reporting [11]. However, these advances should not displace the human judgment that remains essential for interpreting stress scenarios and translating them into risk decisions. The most promising future is one in which generative AI expands the collective imagination of risk while remaining subject to the disciplines of evidence, accountability, and institutional memory. 9 Conclusion Generative AI offers a substantial opportunity to improve scenario-based financial stress testing by expanding the space of plausible stress configurations and by reducing reliance on rigid factor models and expert-defined scenarios. The technical capacity of adversarial, variational, diffusion, and transformer-based architectures to learn high-dimensional dependencies and synthesize new states represents a genuine advance over conventional scenario design methods. However, the value of these techniques cannot be realized through algorithms alone. The system-level analysis presented in this paper emphasizes that generative scenario engines must be embedded within data architectures, governance frameworks, validation processes, and operational controls that are robust to model failure, distribution shift, and institutional misalignment. The structural trade-offs between expressiveness and controllability, centralization and local adaptation, and exploratory novelty and supervisory transparency are not incidental features but defining tensions of the deployment problem. Fairness, sustainability, and policy coordination are equally constitutive of whether generative AI strengthens or weakens the broader risk management system. A careful integration that treats generative models as components of socio-technical infrastructure, rather than as autonomous sources of risk insight, is therefore essential. With appropriate governance and institutional investment, generative AI can enhance the capacity of financial institutions and supervisors to imagine, evaluate, and prepare for severe stress. Without such safeguards, it may add another layer of opacity and operational complexity to an already intricate risk landscape. The task ahead is to build the institutional, regulatory, and technical conditions under which generative scenario generation can be trusted, challenged, and improved over timeAbstract
Scenario-based financial stress testing has become a central instrument of prudential supervision, capital planning, and enterprise risk management, yet its effectiveness remains constrained by the limited novelty, coherence, and tail coverage of analytically designed scenarios. Generative artificial intelligence introduces a different class of modeling capability, one that can learn high-dimensional dependencies across markets, sectors, and time horizons and synthesize plausible stress configurations that resist easy specification through traditional econometric methods. This paper provides a system-level examination of generative AI for scenario-based stress testing and risk management. It analyzes the structural limitations of conventional scenario design, the architectural features of generative models that are relevant for financial risk inference, and the data and computational infrastructures required for dependable deployment. The discussion emphasizes that the value of generative scenario engines depends less on raw generative capacity than on the surrounding governance apparatus, including model validation, auditability, explainability, adversarial robustness, and regulatory alignment. The paper further addresses fairness and distributional effects, operational resilience, sustainability trade-offs, and policy implications. By treating generative AI as a socio-technical infrastructure rather than an isolated modeling technique, the analysis identifies design tensions between exploratory scenario generation and risk control, between model expressiveness and supervisory transparency, and between computational intensity and institutional accountability. The paper concludes with forward-looking perspectives on the integration of generative methods into the broader financial stability toolkit.
References
1. Bank for International Settlements. (2018). Stress testing principles. Basel Committee on Banking Supervision.
2. Schuermann, T. (2014). Stress testing banks. International Journal of Forecasting, 30(3), 717–728.
3. Acharya, V. V., Pedersen, L. H., Philippon, T., & Richardson, M. (2017). Measuring systemic risk. Review of Financial Studies, 30(1), 2–47.
4. Glasserman, P., Kang, C., & Kang, W. (2015). Stress scenario selection by empirical likelihood. Quantitative Finance, 15(1), 25–41.
5. Rebonato, R. (2010). Coherent stress testing: A Bayesian approach to the analysis of financial stress. John Wiley & Sons.
6. Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial nets. Advances in Neural Information Processing Systems, 27.
7. Kingma, D. P., & Welling, M. (2014). Auto-encoding variational Bayes. International Conference on Learning Representations.
8. Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33, 6840–6851.
9. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.
10. Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shinn, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., & Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.
11. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N., Chen, A., Creel, K., Davis, J. Q., Demszky, D., Donahue, C., Doumbouya, M., Durmus, E., Ermon, S., Etchemendy, J., Ethayarajh, K., Fedus, W., Finn, C., Gale, T., Gillespie, L., Goel, S., Goodman, N., Grossman, S., Guha, N., Hashimoto, T., Henderson, P., Hewitt, J., Ho, D. E., Huang, J., Icard, T., Jain, S., Jurafsky, D., Kalluri, P., Karamcheti, S., Kiela, D., Khashabi, D., Koh, P. W., Koyejo, S., Kraska, T., Krishna, R., Kuditipudi, R., Kumar, A., Ladhak, F., Lee, M., Lee, T., Leskovec, J., Leventhal, I., Li, X. L., Li, X., Ma, T., Malik, A., Manning, C. D., Mirhoseini, S., Mitchell, M., Munyikwa, Z., Narla, A., Narayanan, D., Newman, B., Nie, A., Niu, Y., Nilforoshan, H., Nyarko, J., Ogut, G., Orr, L., Papadimitriou, I., Park, J. S., Piech, C., Portelance, E., Potts, C., Raghunathan, A., Ratner, A., Re, C., Rolnick, D., Rosasco, L., Ruder, S., Sasse, K., Schick, T., Srebro, N., Tamkin, A., Tandon, N., Thomas, A., Tramer, F., Wang, R., Webster, A., Williams, A., Wu, S., Xie, S. M., Yacoby, E., Yang, S., Yogatama, D., Zettlemoyer, L., & Liang, P. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
12. Rahwan, I., Cebrian, M., Obradovich, N., Bongard, J., Bonnefon, J.-F., Breazeal, C., Crandall, J. W., Christakis, N. A., Couzin, I. D., Jackson, M. O., Jennings, N. R., Kamar, E., Kloumann, I. M., Larocque, L., Leibo, J. Z., Loewenstein, G., Nunez, M. A., Nushi, B., Olson, S. A., Park, S. E., Pentland, A., Price, M., Rand, D. G., Risi, S., Russakovsky, O., Sherif, Y., Simon, A., Sloman, S. A., Tennant, J., Walker, S., Wernsing, T., Woolley, A., Yeung, K., & Wellman, M. (2019). Machine behaviour. Nature, 568(7753), 477–486.
13. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610–623.
14. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744.
15. Financial Stability Board. (2017). Artificial intelligence and machine learning in financial services: Market developments and financial stability implications. Financial Stability Board.
16. European Commission. (2021). Proposal for a Regulation laying down harmonised rules on artificial intelligence. COM(2021) 206 final.
17. Arrieta, A. B., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., Garcia, S., Gil-Lopez, S., Molina, D., Benjamins, R., Chatila, R., & Herrera, F. (2020). Explainable artificial intelligence: Concepts, taxonomies, opportunities and challenges toward responsible AI. Information Fusion, 58, 82–115.
18. Lipton, Z. C. (2018). The mythos of model interpretability. Communications of the ACM, 61(10), 36–43.
19. Athey, S. (2017). Beyond prediction: Using big data for policy problems. Science, 355(6324), 483–485.
20. Kasy, M., & Abebe, R. (2021). Fairness, equality, and power in algorithmic decision-making. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 576–586.
21. Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), 1–35.
22. Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 3645–3650.
23. Buckmann, M., Haldane, A. G., & Hinterschweiger, M. (2021). Stress testing with concurrent scenarios. Bank of England Staff Working Paper No. 943.
24. Haldane, A. G., & Madouros, V. (2012). The dog and the frisbee. Federal Reserve Bank of Kansas City Economic Policy Symposium.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Advanced Financial Research

This work is licensed under a Creative Commons Attribution 4.0 International License.