LLM-Based Strategic Risk Q & A over SEC Filings for Top Management Decision Support: A Reproducible Evaluastion on the SecQue Benchmark

Qi Xin

Abstract


This paper studies strategic question answering over SEC filings for top-management decision support through a comprehensive evaluation on the SecQue benchmark, encompassing expert-written comparison, ratio, risk, and analysis tasks. The study operationalizes a board-facing retrieval-control stack with four integrated components: evidence retrieval, answer composition, refusal control, and explanation tracing. The benchmark contains 565 expert-written questions distributed across 220 comparison questions, 188 ratio questions, 85 risk questions, and 72 analysis questions. Empirical results demonstrate that metadata-aware retrieval significantly outperforms conventional lexical baselines, while specialized synthesis models achieve superior textual alignment and numerical fidelity. Furthermore, negative-testing and explanation diagnostics confirm robust cross-company refusal mechanisms and complete metadata provenance tracing. By filtering ungrounded assertions and highlighting precise metadata provenance, these automated control layers significantly reduce cognitive overload for corporate directors evaluating dense filings under tight time constraints. Ultimately, grounding board-level synthesis in auditable, trace-backed financial evidence mitigates informational asymmetry between executive teams and non-executive board members

Keywords


Corporate Governance; fFnancial Question Answering; Retrieval Augmented Generation; SEC Filings; Strategic Risk Analysis

Full Text:

PDF

References


Araci, D. (2019). FinBERT: Financial sentiment analysis with pre-trained language models. arXiv preprint arXiv:1908.10063. https://doi.org/10.48550/arXiv.1908.10063

Asai, A., Min, S., Zhong, Z., & Chen, D. (2024). Self-RAG: Learning to retrieve, generate, and critique through self-reflection. In Proceedings of the International Conference on Learning Representations. https://doi.org/10.20944/preprints202508.1211.v1

Beltagy, I., Peters, M. E., & Cohan, A. (2020). Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150. https://doi.org/10.48550/arXiv.2004.05150

Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N. S., Chen, A., Creel, K., Davis, J. Q., Demszky, D., ... Liang, P. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258. https://doi.org/10.48550/arXiv.2108.07258

Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., ... Amodei, D. (2020). Language models are few-shot learners. In Advances in Neural Information Processing Systems (Vol. 33, pp. 1877–1901). https://doi.org/10.48550/arXiv.2005.14165

Campbell, J. L., Chen, H., Dhaliwal, D. S., Lu, H., & Steele, L. B. (2014). The information content of mandatory risk factor disclosures in corporate filings. Review of Accounting Studies, 19(1), 396–455. https://doi.org/10.1007/s11142-013-9258-3

Chalkidis, I., Jana, A., Hartung, D., Bommarito, M., Androutsopoulos, I., Katz, D. M., & Aletras, N. (2022). LexGLUE: A benchmark dataset for legal language understanding in English. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Vol. 1, pp. 4310–4330). https://doi.org/10.18653/v1/2022.acl-long.297

Chen, Z., Chen, W., Smiley, C., Shah, S., Borova, I., Langdon, D., Moussa, R., Beane, M., Huang, T.-H., Routledge, B. R., & Wang, W. Y. (2021). FinQA: A dataset of numerical reasoning over financial data. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (pp. 3697–3711). https://doi.org/10.18653/v1/2021.emnlp-main.300

DeYoung, J., Jain, S., Rajani, N. F., Lehman, E., Xiong, C., Socher, R., & Wallace, B. C. (2020). ERASER: A benchmark to evaluate rationalized NLP models. Transactions of the Association for Computational Linguistics, 8, 444–461. https://doi.org/10.18653/v1/2020.acl-main.408

Eppler, M. J., & Mengis, J. (2004). The concept of information overload: A review of literature from organization science, accounting, marketing, MIS, and related disciplines. The International Journal of Information Management, 24(4), 325–344. https://doi.org/10.1016/j.ijinfomgt.2004.04.006

Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Wang, M., & Wang, H. (2023). Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997. https://doi.org/10.2139/ssrn.4895062

Guu, K., Lee, K., Tung, Z., Pasupat, P., & Chang, M.-W. (2020). REALM: Retrieval-augmented language model pre-training. In Proceedings of the 37th International Conference on Machine Learning (pp. 3929–3938). https://doi.org/10.1145/3581783.3613848

Islam, P., Kannappan, A., Derrick, H., Fodeh, S. J., & Caragea, C. (2023). FinanceBench: A new benchmark for financial question answering. arXiv preprint arXiv:2311.11944. https://doi.org/10.48550/arXiv.2311.11944

Izacard, G., & Grave, E. (2021). Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics (pp. 874–880). https://doi.org/10.18653/v1/2021.eacl-main.74

Jacovi, A., & Goldberg, Y. (2020). Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 4198–4205). https://doi.org/10.18653/v1/2020.acl-main.386

Jensen, M. C., & Meckling, W. H. (1976). Theory of the firm: Managerial behavior, agency costs and ownership structure. Journal of Financial Economics, 3(4), 305–360. https://doi.org/10.1016/0304-405X(76)90026-X

Kamath, A., Jia, R., & Liang, P. (2020). Selective question answering under domain shift. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 5684–5696). https://doi.org/10.18653/v1/2020.acl-main.503

Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., & Yih, W.-t. (2020). Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (pp. 6769–6781). https://doi.org/10.18653/v1/2020.emnlp-main.550

Khattab, O., & Zaharia, M. (2020). ColBERT: Efficient and effective passage search via contextualized late interaction over BERT. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 39–48). https://doi.org/10.1145/3397271.3401075

Kogan, S., Levin, D., Routledge, B. R., Sagi, J. S., & Smith, N. A. (2009). Predicting risk from financial reports with regression. In Proceedings of Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the Association for Computational Linguistics (pp. 272–280). https://doi.org/10.3115/1620754.1620794

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems (33). 9459–9474). https://doi.org/10.48550/arXiv.2005.11401

Li, F. (2010). The information content of forward-looking statements in corporate filings—A naïve Bayesian machine learning approach. Journal of Accounting Research, 48(5), 1049–1102. https://doi.org/10.1111/j.1475-679x.2010.00382.x

Loughran, T., & McDonald, B. (2011). When is a liability not a liability? Textual analysis, dictionaries, and 10-Ks. Journal of Finance, 66(1), 35–65. https://doi.org/10.1111/j.1540-6261.2010.01625.x

Loughran, T., & McDonald, B. (2015). The use of word lists in textual analysis. Journal of Behavioral Finance, 16(1), 1–11. https://doi.org/10.1080/15427560.2015.1000335

Mialon, G., Dessì, R., Lomeli, M., Nalmpantis, C., Pascal, F., Ranaldi, L., Stojnic, R., Penedo, G., & Scialom, T. (2023). Augmented language models: A survey. Transactions on Machine Learning Research, 2023, 1–34. https://doi.org/10.48550/arXiv.2302.07842

Open AI. (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774. https://doi.org/10.48550/arXiv.2303.08774

Roberts, A., Raffel, C., & Shazeer, N. (2020). How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (pp. 5418–5426). https://doi.org/10.18653/v1/2020.emnlp-main.437

Wu, S., Irsoy, O., Lu, S., Dabravolski, V., Dredze, M., Gehrmann, S., Kambadur, P., Rosenberg, D., & Mann, G. (2023). BloombergGPT: A large language model for finance. arXiv preprint arXiv:2303.17564. https://doi.org/10.48550/arXiv.2303.17564

Zaheer, M., Guruganesh, G., Dubey, A., Ainslie, J., Alberti, C., Ontañón, S., Pham, P., Ravula, A., Wang, Q., Yang, L., & Ahmed, A. (2020). Big Bird: Transformers for longer sequences. In Advances in Neural Information Processing Systems (Vol. 33, pp. 17283–17297). https://doi.org/10.48550/arXiv.2007.14062




DOI: https://doi.org/10.18860/mec-j.v10i2.41794

Refbacks

  • There are currently no refbacks.




Editorial Office:
Faculty of Economics,
State Islamic University of Maulana Malik Ibrahim Malang
Gajayana Street 50, Malang-East Java, Indonesia 65144
Phone (+62) 341 558881, Facsimile (+62) 341 558881
e-mail: mecjournal@uin-malang.ac.id

 

 

P-ISSN 2599-3402
E-ISSN 2598-9537

Lisensi Creative Commons

MEC-J is licensed under Creative Commons Attribution–ShareAlike 4.0 International License (CC BY-SA 4.0)

 

MEC-J INDEXED IN:  

                        

SINTA

Member of:

 

View My Stats