MMMMM:多语言多模态虚假信息机制研究的统一分类法
MMMMM: A Unified Taxonomy for Investigating the Mechanisms of Multilingual MultiModal Misinformation
August 30, 2026
作者: Nadav Borenstein, Greta Warren, Desmond Elliott, Isabelle Augenstein
cs.AI
摘要
社交媒体上的多模态错误信息极为普遍、影响显著且危害严重,然而相较于纯文本错误信息,其检测与应对更加困难,且目前仍缺乏充分理解。对其特征与欺骗策略的研究,受到缺乏基于真实世界情境的分类体系以及当前多模态机器学习模型局限性的阻碍,这导致无法在大规模层面实现自动化标注与分析。我们通过三个步骤解决上述不足。首先,我们收集了来自Twitter/X的七种语言的大规模、高质量真实世界错误信息数据集。其次,我们基于对数据的深入定性分析及先前理论工作,开发了一个新颖且全面的多模态错误信息分类体系。最后,我们利用视觉-语言模型(VLM)通过自动化多步骤标注流程将该分类体系操作化,并进行人工验证。我们的新方法揭示了关于社交媒体用户如何将图像与文本结合以在现实中传播错误信息的此前未被记录的洞察,例如,AI生成内容在科技与科学领域尤为普遍,而疫苗错误信息则不成比例地利用新闻机构的图像来彰显可信度。我们的方法与发现为多模态错误信息的针对性检测提供了指导,并表明缓解措施应进行战略性、差异化的制定与应用,而非统一采用。
English
Multimodal misinformation on social media is highly prevalent, potent, and harmful, yet difficult to detect and counter, and still poorly understood compared to its text-only counterpart. Research on the properties and deceptive strategies of multimodal misinformation is hindered by a lack of taxonomies grounded in real-world contexts and by the limitations of current multimodal machine learning models, which prevent the automation of annotation and analysis at scale. We address these shortcomings in three steps. First, we collect a large-scale, high-quality dataset of real-world misinformation instances from Twitter/X in seven languages. Second, we develop a novel, comprehensive taxonomy of multimodal misinformation grounded in an in-depth qualitative analysis of the data and prior theoretical work. Finally, we operationalise the taxonomy through an automated multi-step annotation pipeline using a Vision-Language Model (VLM), and perform human-validation. Our novel approach leads to previously undocumented insights about how social media users combine images with text to spread misinformation in the wild, e.g., that AI-generated content is particularly prevalent in technology and science, while vaccination misinformation disproportionately utilises images from news outlets to assert credibility. Our method and findings provide guidance for targeted approaches for detecting multimodal misinformation, and suggest that mitigation efforts should be developed and applied strategically rather than uniformly.