MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding Evaluation

  • Weihua Zheng ,
  • Zhengyuan Liu ,
  • Tanmoy Chakraborty ,
  • Weiwen Xu ,
  • Xiaoxue Gao ,
  • Bryan Chen Zhengyu Tan ,
  • Bowei Zou ,
  • Chang Liu ,
  • Yujia Hu ,
  • ,
  • ,
  • ,
  • Chaojun Wang ,
  • Long Li ,
  • Rui Liu ,
  • Huiya Liu ,
  • K. Inoue ,
  • Ryuichi Sumida ,
  • Tatsuya Kawahara ,
  • Fan Xu ,
  • Lingyu Ye ,
  • Wei Tian ,
  • Dongjun Kim ,
  • Jimin Jung ,
  • J. Seo ,
  • Nadya Yuki Wangsajaya ,
  • P. M. Duc ,
  • Ojasva Saxena ,
  • Palash Nandi ,
  • Xiyan Tao ,
  • Wiwik Karlina ,
  • T. Luong ,
  • Keertana Arun Vasan ,
  • Roy Ka-Wei Lee ,
  • Nancy F. Chen

Annual Meeting of the Association for Computational Linguistics |

PDF

The global deployment of Large Language Models (LLMs) underscores the urgent need to evaluate their cultural alignment. How ever, assessing genuine “cultural awareness” across modalities (text, vision, speech) and languages remains a significant challenge. To comprehensively investigate this domain, we propose a Multilingual, Multimodal Alignment framework for Cultural grounding evaluation (MMAC). This systematic framework encompasses a tri-modally aligned cultural bench mark creation pipeline and a five-dimensional evaluation protocol to assess cross-country awareness disparities, evaluate cross-lingual and cross-modal consistency, and verify cultural knowledge generalization and grounding validity. Given the prevailing Western cultural bias in current models, we focus on 8 Asian countries as our dataset foundation to more acutely reveal potential cultural deficiencies in LLMs. Our dataset, MMAC-bench, features 27,000 human-curated questions across 10 languages. Crucially, it is the first dataset aligned at the input level across text, image, and speech, enabling direct cross-modal transfer tests. Each question consists of multiple-choice options accompanied by open-ended generated explanations, where 79% require multi-step reasoning grounded in cultural context, moving beyond simple memorization. We probe the causes of modal divergence, offering insights into foster ing culturally robust MLLMs.