ChatPaper.aiChatPaper

MMMMM: 다국어 멀티모달 오정보(Misinformation)의 메커니즘을 규명하기 위한 통합 분류체계

MMMMM: A Unified Taxonomy for Investigating the Mechanisms of Multilingual MultiModal Misinformation

August 30, 2026
저자: Nadav Borenstein, Greta Warren, Desmond Elliott, Isabelle Augenstein
cs.AI

초록

소셜 미디어상의 멀티모달 허위정보는 매우 널리 퍼져 있고, 강력하며, 해롭지만, 탐지하고 대응하는 것이 어렵고, 텍스트만으로 구성된 허위정보에 비해 여전히 제대로 이해되지 않고 있다. 멀티모달 허위정보의 특성과 기만적 전략에 대한 연구는 실제 맥락에 기반한 분류체계의 부재와, 대규모 주석 및 분석 자동화를 불가능하게 하는 현재 멀티모달 기계학습 모델의 한계로 인해 어려움을 겪고 있다. 우리는 세 단계로 이러한 문제를 해결한다. 첫째, 7개 언어로 작성된 트위터(X)의 실제 허위정보 사례를 대규모 고품질 데이터셋으로 수집한다. 둘째, 데이터에 대한 심층 질적 분석과 기존 이론 연구에 기반한 새롭고 포괄적인 멀티모달 허위정보 분류체계를 개발한다. 마지막으로, 비전-언어 모델(VLM)을 활용한 자동화된 다단계 주석 파이프라인을 통해 분류체계를 구현하고 인간 검증을 수행한다. 이러한 새로운 접근 방식은 실제 환경에서 소셜 미디어 사용자가 이미지와 텍스트를 결합하여 허위정보를 유포하는 방식에 대해 이전에 보고된 바 없는 통찰을 제공한다. 예컨대 AI 생성 콘텐츠는 기술 및 과학 분야에서 특히 두드러지는 반면, 백신 허위정보는 신뢰성을 내세우기 위해 뉴스 매체의 이미지를 불균형적으로 활용한다. 우리의 방법과 발견은 멀티모달 허위정보 탐지를 위한 맞춤형 접근 방식에 지침을 제공하며, 완화 노력은 일률적으로 적용하기보다 전략적으로 개발·적용되어야 함을 시사한다.
English
Multimodal misinformation on social media is highly prevalent, potent, and harmful, yet difficult to detect and counter, and still poorly understood compared to its text-only counterpart. Research on the properties and deceptive strategies of multimodal misinformation is hindered by a lack of taxonomies grounded in real-world contexts and by the limitations of current multimodal machine learning models, which prevent the automation of annotation and analysis at scale. We address these shortcomings in three steps. First, we collect a large-scale, high-quality dataset of real-world misinformation instances from Twitter/X in seven languages. Second, we develop a novel, comprehensive taxonomy of multimodal misinformation grounded in an in-depth qualitative analysis of the data and prior theoretical work. Finally, we operationalise the taxonomy through an automated multi-step annotation pipeline using a Vision-Language Model (VLM), and perform human-validation. Our novel approach leads to previously undocumented insights about how social media users combine images with text to spread misinformation in the wild, e.g., that AI-generated content is particularly prevalent in technology and science, while vaccination misinformation disproportionately utilises images from news outlets to assert credibility. Our method and findings provide guidance for targeted approaches for detecting multimodal misinformation, and suggest that mitigation efforts should be developed and applied strategically rather than uniformly.