Research output

Publications

A record of my research, in print and in progress.

Publications
4
Citations
18
h-index
3
i10-index
0

Google Scholar metrics refreshed on 2026-09-24.

Conference papers 1

C1

Conference paper

Same Claim, Different Judgment: Benchmarking Scenario-Induced Bias in Multilingual Financial Misinformation Detection

Zhiwei Liu, Yupen Cao, Yuechen Jiang, Mohsinul Kabir, Polydoros Giannouris, Chen Xu, Ziyang Xu, Tianlei Zhu, Md. Tariquzzaman, Triantafillos Papadopoulos, Yan Wang, Lingfei Qian, Xueqing Peng, Zhuohan Xie, Ye Yuan, Saeed Almheiri, Abdulrazzaq Alnajjar, Mingbin Chen, Harry Stuart, Paul Thompson, Prayag Tiwari, Alejandro Lopez-Lira, Xue Liu, Jimin Huang, Sophia Ananiadou

AcceptedFindings of ACL 2026

LLM evaluationMisinformation
Overview & figures

In short

MFMD-Scen asks whether large language models reach the same verdict on the same financial claim once a realistic scenario is added to the prompt: an investor’s role and personality, the market they work in, or their ethnicity and faith. Across 22 models and four languages, the answer is often no. The claim stays fixed, yet the scenario shifts the verdict.

Full abstract

Large language models (LLMs) have been widely applied across various domains of finance. Since their training data are largely derived from human-authored corpora, LLMs may inherit a range of human biases. Behavioral biases can lead to instability and uncertainty in decision-making, particularly when processing financial information. However, existing research on LLM bias has mainly focused on direct questioning or simplified, general-purpose settings, with limited consideration of the complex real-world, context-sensitive, multilingual financial misinformation detection tasks. In this work, we propose MFMD-Scen, a comprehensive benchmark for evaluating behavioral biases of LLMs in financial misinformation detection across diverse economic scenarios. In collaboration with financial experts, we construct three types of complex financial scenarios: (i) role- and personality-based, (ii) role- and region-based, and (iii) role-based scenarios incorporating ethnicity and religious beliefs. We further develop a multilingual financial misinformation dataset covering English, Chinese, Greek, and Bengali. By integrating these scenarios with misinformation claims, MFMD-Scen enables a systematic evaluation of 22 mainstream LLMs. Our findings reveal that pronounced behavioral biases persist across both commercial and open-source models.

BibTeX
@inproceedings{liu-etal-2026-claim,
    title = "Same Claim, Different Judgment: Benchmarking Scenario-Induced Bias in Multilingual Financial Misinformation Detection",
    author = "Liu, Zhiwei  and
      Cao, Yupeng  and
      Jiang, Yuechen  and
      Kabir, Mohsinul  and
      Giannouris, Polydoros  and
      Xu, Chen  and
      Xu, Ziyang  and
      Zhu, Tianlei  and
      Tariquzzaman, Md.  and
      Papadopoulos, Triantafillos  and
      Wang, Yan  and
      Qian, Lingfei  and
      Peng, Xueqing  and
      Xie, Zhuohan  and
      Yuan, Ye  and
      Almheiri, Saeed  and
      Alnajjar, Abdulrazzaq  and
      Chen, Ming-Bin  and
      Stuart, Harry  and
      Thompson, Paul  and
      Tiwari, Prayag  and
      Lopez-Lira, Alejandro  and
      Liu, Xue  and
      Huang, Jimin  and
      Ananiadou, Sophia",
    editor = "Liakata, Maria  and
      Moreira, Viviane P.  and
      Zhang, Jiajun  and
      Jurgens, David",
    booktitle = "Findings of the {A}ssociation for {C}omputational {L}inguistics: {ACL} 2026",
    month = jul,
    year = "2026",
    address = "San Diego, California, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2026.findings-acl.479/",
    doi = "10.18653/v1/2026.findings-acl.479",
    pages = "9838--9864",
    ISBN = "979-8-89176-395-1",
    abstract = "Large language models (LLMs) have been widely applied across various domains of finance. Since their training data are largely derived from human-authored corpora, LLMs may inherit a range of human biases. Behavioral biases can lead to instability and uncertainty in decision-making, particularly when processing financial information. However, existing research on LLM bias has mainly focused on direct questioning or simplified, general-purpose settings, with limited consideration of the complex real-world financial environments and high-risk, context-sensitive, multilingual financial misinformation detection tasks (MFMD). In this work, we propose MFMDScen, a comprehensive benchmark for evaluating behavioral biases of LLMs in MFMD across diverse economic scenarios. In collaboration with financial experts, we construct three types of complex financial scenarios: (i) role- and personality-based, (ii) role- and region-based, and (iii) role-based scenarios incorporating ethnicity and religious beliefs. We further develop a multilingual financial misinformation dataset covering English, Chinese, Greek, and Bengali. By integrating these scenarios with misinformation claims, MFMDScen enables a systematic evaluation of 22 mainstream LLMs. Our findings reveal that pronounced behavioral biases persist across both commercial and open-source models. This project is available at \url{https://github.com/lzw108/FMD}."
}

Journal articles 0

No journal articles listed yet.

Workshop papers 2

W2

Workshop paper

Prompting with Sign Parameters for Low-resource Sign Language Instruction Generation

Md. Tariquzzaman, Md Farhan Ishmam, Saiyma Sittul Muna, Md Kamrul Hasan, Hasan Mahmud

AcceptedVision Foundation Models and Gen AI for Accessibility (CV4A11y) at ICCV 2025

Accessibility
Overview & figures

In short

Sign-language instruction generation writes step-by-step text that lets a hearing learner reproduce a sign. BdSLIG is the first dataset for this task in Bengali Sign Language, and Sign Parameter-Infused (SPI) prompting writes standard sign parameters into the prompt so vision-language models produce more structured, reproducible instructions.

Full abstract

Sign Language Instruction Generation (SLIG) produces step-by-step textual instructions that enable non-SL users to imitate and learn sign language gestures, promoting two-way interaction. We introduce BdSLIG, the first Bengali SLIG dataset, used to evaluate Vision Language Models (VLMs) on under-resourced SLIG tasks and long-tail visual concepts. To enhance zero-shot performance, we introduce Sign Parameter-Infused (SPI) prompting, which integrates standard sign parameters, like hand shape, motion, and orientation, directly into textual prompts. Subsuming standard sign parameters into the prompt makes the instructions more structured and reproducible than free-form natural text from vanilla prompting. We envision that our work would promote inclusivity and advancement in sign language learning systems for under-resourced communities.

BibTeX
@misc{tariquzzaman2026promptingsignparameterslowresource,
      title={Prompting with Sign Parameters for Low-resource Sign Language Instruction Generation},
      author={Md Tariquzzaman and Md Farhan Ishmam and Saiyma Sittul Muna and Md Kamrul Hasan and Hasan Mahmud},
      year={2026},
      eprint={2508.16076},
      archivePrefix={arXiv},
      primaryClass={cs.HC},
      url={https://arxiv.org/abs/2508.16076},
}
W1

Workshop paper

the_linguists at BLP-2023 Task 1: A Novel Informal Bangla Fasttext Embedding for Violence Inciting Text Detection

Md. Tariquzzaman, Md. Wasif Kader, Audwit Nafi Anam, Naimul Haque, Mohsinul Kabir, Hasan Mahmud, Md Kamrul Hasan

Accepted1st Workshop on Bangla Language Processing (BLP) at EMNLP 2023

Best Shared Task Paper Award

Bangla NLP
Overview & figures

In short

An informal Bangla FastText embedding, trained on 3.8 million social-media comments, lets a lightweight BiLSTM detect violence-inciting text within about four macro-F1 points of BanglaBERT while training 17 times faster. The paper received the Best Shared Task Paper Award at BLP 2023.

Full abstract

This paper introduces a novel informal Bangla word embedding for designing a cost-efficient solution for the task “Violence Inciting Text Detection,” which focuses on developing classification systems to categorize violence that can potentially incite further violent actions. We propose a semi-supervised learning approach by training an informal Bangla FastText embedding, which is further fine-tuned on lightweight models on the task-specific dataset and yielded competitive results to our initial method using BanglaBERT, which secured the 7th position with an F1-score of 73.98%. We conduct extensive experiments to assess the efficiency of the proposed embedding and how well it generalizes in terms of violence classification, along with its coverage on the task’s dataset. Our proposed Bangla IFT embedding achieved a competitive macro average F1 score of 70.45%. Additionally, we provide a detailed analysis of our findings, delving into potential causes of misclassification in the detection of violence-inciting text.

BibTeX
@inproceedings{tariquzzaman-etal-2023-linguists,
    title = "the{\_}linguists at {BLP}-2023 Task 1: A Novel Informal {B}angla {F}asttext Embedding for Violence Inciting Text Detection",
    author = "Tariquzzaman, Md.  and
      Kader, Md Wasif  and
      Anam, Audwit  and
      Haque, Naimul  and
      Kabir, Mohsinul  and
      Mahmud, Hasan  and
      Hasan, Md Kamrul",
    editor = "Alam, Firoj  and
      Kar, Sudipta  and
      Chowdhury, Shammur Absar  and
      Sadeque, Farig  and
      Amin, Ruhul",
    booktitle = "Proceedings of the First Workshop on Bangla Language Processing (BLP-2023)",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.banglalp-1.26/",
    doi = "10.18653/v1/2023.banglalp-1.26",
    pages = "214--219"
}

Preprints 1

P1

Preprint

BDA: Bangla Text Data Augmentation Framework

Md. Tariquzzaman, Audwit Nafi Anam, Naimul Haque, Mohsinul Kabir, Hasan Mahmud, Md Kamrul Hasan

arXiv preprint arXiv:2412.08753 (2024)

Text augmentation
Overview & figures

In short

BDA augments Bangla text with two rule-based and two transformer-based methods, then keeps only outputs that preserve the meaning while changing the wording. Across five classification datasets, training on half the data plus BDA matches or beats full-data training on three.

Full abstract

Data augmentation involves generating synthetic samples that resemble those in a given dataset. In resource-limited fields where high-quality data is scarce, augmentation plays a crucial role in increasing the volume of training data. This paper introduces a Bangla Text Data Augmentation (BDA) Framework that uses both pre-trained models and rule-based methods to create new variants of the text. A filtering process is included to ensure that the new text keeps the same meaning as the original while also adding variety in the words used. We conduct a comprehensive evaluation of the framework’s effectiveness in Bangla text classification tasks. Our framework achieved significant improvement in F1 scores across five distinct datasets, delivering performance equivalent to models trained on 100% of the data while utilizing only 50% of the training dataset. Additionally, we explore the impact of data scarcity by progressively reducing the training data and augmenting it through BDA, resulting in notable F1 score enhancements. The study offers a thorough examination of BDA’s performance, identifying key factors for optimal results and addressing its limitations through detailed analysis.

BibTeX
@misc{tariquzzaman2024bdabanglatextdata,
      title={BDA: Bangla Text Data Augmentation Framework},
      author={Md. Tariquzzaman and Audwit Nafi Anam and Naimul Haque and Mohsinul Kabir and Hasan Mahmud and Md Kamrul Hasan},
      year={2024},
      eprint={2412.08753},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2412.08753},
}

Code & data

BdSLIG

Bangla Sign Language instruction generation dataset.

SPIP

Sign Parameter Informed Prompting: reference implementation.

VITD

Informal Bangla embeddings and violence-inciting text detection.