Conference paper
Same Claim, Different Judgment: Benchmarking Scenario-Induced Bias in Multilingual Financial Misinformation Detection
AcceptedFindings of ACL 2026
Overview & figures
In short
MFMD-Scen asks whether large language models reach the same verdict on the same financial claim once a realistic scenario is added to the prompt: an investor’s role and personality, the market they work in, or their ethnicity and faith. Across 22 models and four languages, the answer is often no. The claim stays fixed, yet the scenario shifts the verdict.
Full abstract
Large language models (LLMs) have been widely applied across various domains of finance. Since their training data are largely derived from human-authored corpora, LLMs may inherit a range of human biases. Behavioral biases can lead to instability and uncertainty in decision-making, particularly when processing financial information. However, existing research on LLM bias has mainly focused on direct questioning or simplified, general-purpose settings, with limited consideration of the complex real-world, context-sensitive, multilingual financial misinformation detection tasks. In this work, we propose MFMD-Scen, a comprehensive benchmark for evaluating behavioral biases of LLMs in financial misinformation detection across diverse economic scenarios. In collaboration with financial experts, we construct three types of complex financial scenarios: (i) role- and personality-based, (ii) role- and region-based, and (iii) role-based scenarios incorporating ethnicity and religious beliefs. We further develop a multilingual financial misinformation dataset covering English, Chinese, Greek, and Bengali. By integrating these scenarios with misinformation claims, MFMD-Scen enables a systematic evaluation of 22 mainstream LLMs. Our findings reveal that pronounced behavioral biases persist across both commercial and open-source models.





BibTeX
@inproceedings{liu-etal-2026-claim,
title = "Same Claim, Different Judgment: Benchmarking Scenario-Induced Bias in Multilingual Financial Misinformation Detection",
author = "Liu, Zhiwei and
Cao, Yupeng and
Jiang, Yuechen and
Kabir, Mohsinul and
Giannouris, Polydoros and
Xu, Chen and
Xu, Ziyang and
Zhu, Tianlei and
Tariquzzaman, Md. and
Papadopoulos, Triantafillos and
Wang, Yan and
Qian, Lingfei and
Peng, Xueqing and
Xie, Zhuohan and
Yuan, Ye and
Almheiri, Saeed and
Alnajjar, Abdulrazzaq and
Chen, Ming-Bin and
Stuart, Harry and
Thompson, Paul and
Tiwari, Prayag and
Lopez-Lira, Alejandro and
Liu, Xue and
Huang, Jimin and
Ananiadou, Sophia",
editor = "Liakata, Maria and
Moreira, Viviane P. and
Zhang, Jiajun and
Jurgens, David",
booktitle = "Findings of the {A}ssociation for {C}omputational {L}inguistics: {ACL} 2026",
month = jul,
year = "2026",
address = "San Diego, California, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.findings-acl.479/",
doi = "10.18653/v1/2026.findings-acl.479",
pages = "9838--9864",
ISBN = "979-8-89176-395-1",
abstract = "Large language models (LLMs) have been widely applied across various domains of finance. Since their training data are largely derived from human-authored corpora, LLMs may inherit a range of human biases. Behavioral biases can lead to instability and uncertainty in decision-making, particularly when processing financial information. However, existing research on LLM bias has mainly focused on direct questioning or simplified, general-purpose settings, with limited consideration of the complex real-world financial environments and high-risk, context-sensitive, multilingual financial misinformation detection tasks (MFMD). In this work, we propose MFMDScen, a comprehensive benchmark for evaluating behavioral biases of LLMs in MFMD across diverse economic scenarios. In collaboration with financial experts, we construct three types of complex financial scenarios: (i) role- and personality-based, (ii) role- and region-based, and (iii) role-based scenarios incorporating ethnicity and religious beliefs. We further develop a multilingual financial misinformation dataset covering English, Chinese, Greek, and Bengali. By integrating these scenarios with misinformation claims, MFMDScen enables a systematic evaluation of 22 mainstream LLMs. Our findings reveal that pronounced behavioral biases persist across both commercial and open-source models. This project is available at \url{https://github.com/lzw108/FMD}."
}






