MMMMM:探究多語言多模態虛假資訊機制的統一分類法
MMMMM: A Unified Taxonomy for Investigating the Mechanisms of Multilingual MultiModal Misinformation
August 30, 2026
作者: Nadav Borenstein, Greta Warren, Desmond Elliott, Isabelle Augenstein
cs.AI
摘要
社群媒體上的多模態錯誤資訊極為普遍、影響力強大且具危害性,然而相較於純文字的錯誤資訊,其偵測與反制难度更高,且目前的理解仍相當有限。針對多模態錯誤資訊之特性與欺騙策略的研究,長期受到兩項因素阻礙:缺乏根植於真實世界脈絡的分類法,以及當前多模態機器學習模型的局限性,導致無法大規模自動化標註與分析。我們透過三個步驟來解決這些不足。首先,我們從Twitter/X收集了一個大規模、高品質的真實世界錯誤資訊資料集,涵蓋七種語言。其次,我們基於對資料的深入質性分析及先前的理論研究,建立了一套新穎且全面的多模態錯誤資訊分類法。最後,我們運用視覺語言模型(VLM)透過自動化的多步驟標註流程將該分類法操作化,並進行人工驗證。我們的新穎方法帶來了以往未被記錄的洞見,揭示了社群媒體使用者如何結合圖片與文字在實際環境中傳播錯誤資訊,例如:AI生成內容在科技與科學領域尤為普遍,而疫苗相關錯誤資訊則不成比例地利用新聞媒體的圖片來建立可信度。我們的方法與發現為針對性的多模態錯誤資訊偵測策略提供了指引,並建議緩解措施應採取策略性且差異化的方式來制定與實施,而非齊頭式地統一應用。
English
Multimodal misinformation on social media is highly prevalent, potent, and harmful, yet difficult to detect and counter, and still poorly understood compared to its text-only counterpart. Research on the properties and deceptive strategies of multimodal misinformation is hindered by a lack of taxonomies grounded in real-world contexts and by the limitations of current multimodal machine learning models, which prevent the automation of annotation and analysis at scale. We address these shortcomings in three steps. First, we collect a large-scale, high-quality dataset of real-world misinformation instances from Twitter/X in seven languages. Second, we develop a novel, comprehensive taxonomy of multimodal misinformation grounded in an in-depth qualitative analysis of the data and prior theoretical work. Finally, we operationalise the taxonomy through an automated multi-step annotation pipeline using a Vision-Language Model (VLM), and perform human-validation. Our novel approach leads to previously undocumented insights about how social media users combine images with text to spread misinformation in the wild, e.g., that AI-generated content is particularly prevalent in technology and science, while vaccination misinformation disproportionately utilises images from news outlets to assert credibility. Our method and findings provide guidance for targeted approaches for detecting multimodal misinformation, and suggest that mitigation efforts should be developed and applied strategically rather than uniformly.