MMMMM:多言語マルチモーダル誤情報のメカニズム解明のための統一分類法
MMMMM: A Unified Taxonomy for Investigating the Mechanisms of Multilingual MultiModal Misinformation
August 30, 2026
著者: Nadav Borenstein, Greta Warren, Desmond Elliott, Isabelle Augenstein
cs.AI
要旨
ソーシャルメディア上のマルチモーダル誤情報は、非常に蔓延しており、影響力が強く、有害である一方、テキストのみの誤情報と比較すると、検出や対抗が難しく、未だ十分に理解されていない。マルチモーダル誤情報の特性と欺瞞的戦略に関する研究は、実世界の文脈に基づいた分類体系の欠如と、現在のマルチモーダル機械学習モデルの限界によって妨げられている。これらの限界により、大規模な注釈付けと分析の自動化が不可能になっている。我々は、これらの欠点に三つの段階で対処する。第一に、七言語のTwitter/Xから、実世界の誤情報事例からなる大規模かつ高品質なデータセットを収集する。第二に、データの詳細な質的分析と先行理論研究に基づいた、新規かつ包括的なマルチモーダル誤情報の分類体系を開発する。最後に、視覚言語モデル(VLM)を用いた自動化された多段階注釈パイプラインを通じて分類体系を運用可能にし、人間による検証を行う。我々の新しいアプローチは、ソーシャルメディアユーザーが実際の環境で誤情報を拡散するために画像とテキストをどのように組み合わせるかについて、これまで記録されていなかった洞察をもたらす。例えば、AI生成コンテンツはテクノロジーと科学の分野で特に顕著である一方、ワクチン誤情報は信頼性を主張するためにニュースメディアの画像を不均衡に利用している。我々の手法と発見は、マルチモーダル誤情報を検出するための標的型アプローチへの指針を提供し、緩和策は均一的ではなく戦略的に開発・適用されるべきであることを示唆している。
English
Multimodal misinformation on social media is highly prevalent, potent, and harmful, yet difficult to detect and counter, and still poorly understood compared to its text-only counterpart. Research on the properties and deceptive strategies of multimodal misinformation is hindered by a lack of taxonomies grounded in real-world contexts and by the limitations of current multimodal machine learning models, which prevent the automation of annotation and analysis at scale. We address these shortcomings in three steps. First, we collect a large-scale, high-quality dataset of real-world misinformation instances from Twitter/X in seven languages. Second, we develop a novel, comprehensive taxonomy of multimodal misinformation grounded in an in-depth qualitative analysis of the data and prior theoretical work. Finally, we operationalise the taxonomy through an automated multi-step annotation pipeline using a Vision-Language Model (VLM), and perform human-validation. Our novel approach leads to previously undocumented insights about how social media users combine images with text to spread misinformation in the wild, e.g., that AI-generated content is particularly prevalent in technology and science, while vaccination misinformation disproportionately utilises images from news outlets to assert credibility. Our method and findings provide guidance for targeted approaches for detecting multimodal misinformation, and suggest that mitigation efforts should be developed and applied strategically rather than uniformly.