Cross-lingual multimodal mathematical reasoning: benchmarking frontier models in Bengali
View/ Open
Date
2026-08Author
Nondon, Nishorgo
Naurin, Sabah
Muzaffar, Shams Habib
Metadata
Show full item recordAbstract
This thesis introduces the Bangladesh Mathematical Olympiad (BdMO 2022–2024, 2026) benchmark, containing 1,336 mathematical problems, including 201 image-based problems, to evaluate the mathematical and visual reasoning abilities of Large Language Models (LLMs) and Vision-Language Models (VLMs) in Bangla and English. Four AI models were assessed across seven mathematical topics and four educational levels. Results showed that image-based problems reduced accuracy by 18–19%, while switching from English to Bangla caused only a 0.46% decrease, indicating that visual reasoning presents a greater challenge than language processing. The study also examined bilingual consistency, model performance, and common reasoning failures, providing insights into the limitations of current AI models in mathematical and multimodal reasoning.
Collections
- Undergraduate Thesis [63]
Publisher:
Independent University, Bangladesh (IUB)
Department:
Department of Computer Science and Engineering
Type:
Thesis
Keywords:
Mathematical Reasoning, Large Language Models (LLMs), Vision-Language Models (VLMs), Bangla-English Benchmark, Multimodal AI Evaluation