방언도 언어처럼 제어될 수 있을까? 아랍어 LLM에서의 희소 뉴런과 분산 방향
Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs
July 4, 2026
저자: Kareem Elozeiri, Mervat Abassy, Omar Kallas, Fahim Dalvi, Preslav Nakov, Kentaro Inui, Nadir Durrani
cs.AI
초록
아랍어 자연어 처리(NLP)의 주요 과제 중 하나는 현대 표준 아랍어(MSA)에 비해 방언 데이터가 부족하다는 점이며, 이로 인해 LLM은 MSA를 과도하게 생성하고 방언적으로 정확한 생성에 어려움을 겪는다. 해석 가능성 관점에서 이는 근본적인 질문을 제기한다: 방언적 특징이 모델 내부에서 어디에, 어떻게 인코딩되어 있으며, 미세 조정 없이도 이러한 표현을 활용하여 방언 생성을 개선할 수 있는가? 본 연구는 해석 가능성 탐사 도구이자 제어 메커니즘으로 동시에 기능하는 두 가지 보완적 추론 시간 접근법을 조사한다. 첫째, 뉴런 수준 분석을 수행하여 방언별 특징을 인코딩하는 희소 뉴런 집단을 식별하고, 이러한 뉴런을 증폭 또는 억제함으로써 모델 출력을 목표 방언으로 유도할 수 있음을 보여준다. 둘째, 단일 뉴런 수준에서 방언적 특징이 얽혀 있다는 점에 착안하여, 방언별 활성화 방향을 추출하고 추론 중에 이를 주입하는 벡터 유도 접근법을 적용한다. 이 두 방법을 통해 아랍어 LLM에서 방언 지식의 기하학적 구조를 조명하고, 방언별 미세 조정 없이도 방언 제어를 가능하게 하는 원칙적이고 해석 가능성에 기반한 프레임워크를 제공한다.
English
A key challenge in Arabic NLP is the scarcity of dialectal data relative to Modern Standard Arabic (MSA), causing LLMs to overproduce MSA and struggle with dialectally accurate generation. From an interpretability perspective, this raises a fundamental question: where and how are dialectal features encoded within model internals, and can these representations be leveraged to improve dialect generation without fine-tuning? This study investigates two complementary inference-time approaches that serve simultaneously as interpretability probes and control mechanisms. First, we conduct a neuron-level analysis, identifying sparse neuron populations that encode dialect-specific features and showing that amplifying or suppressing these neurons can steer model outputs toward target dialects. Second, motivated by the entanglement of dialectal features at the single-neuron level, we apply a vector-steering approach that extracts dialect-specific activation directions and injects them during inference. Together, these methods illuminate the geometry of dialectal knowledge in Arabic LLMs and offer a principled, interpretability-grounded framework for dialect control without requiring dialect-specific fine-tuning.