ChatPaper.aiChatPaper

方言能否像語言一樣被引導?阿拉伯語大語言模型中的稀疏神經元與分佈方向

Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs

July 4, 2026
作者: Kareem Elozeiri, Mervat Abassy, Omar Kallas, Fahim Dalvi, Preslav Nakov, Kentaro Inui, Nadir Durrani
cs.AI

摘要

阿拉伯语自然语言处理面临的一个关键挑战是,相较于现代标准阿拉伯语(MSA),方言数据的稀缺性,这导致大语言模型(LLMs)过度生成MSA,而难以生成准确符合方言特征的文本。从可解释性的角度来看,这引出一个根本性问题:方言特征在模型内部编码于何处、如何编码,以及这些表征能否在不进行微调的情况下用于改进方言生成?本研究探讨了两种互为补充的推理阶段方法,这些方法同时充当可解释性探针和控制机制。首先,我们进行神经元层级的分析,识别出编码方言特定特征的稀疏神经元群体,并证明放大或抑制这些神经元可以将模型输出导向目标方言。其次,鉴于方言特征在单神经元层面存在纠缠现象,我们应用一种向量引导方法,提取方言特定的激活方向,并在推理过程中注入这些方向。综合来看,这些方法揭示了阿拉伯语大语言模型中方言知识的空间结构特征,并提供了一个基于可解释性的原则性框架,用于在不经方言特定微调的情况下实现方言控制。
English
A key challenge in Arabic NLP is the scarcity of dialectal data relative to Modern Standard Arabic (MSA), causing LLMs to overproduce MSA and struggle with dialectally accurate generation. From an interpretability perspective, this raises a fundamental question: where and how are dialectal features encoded within model internals, and can these representations be leveraged to improve dialect generation without fine-tuning? This study investigates two complementary inference-time approaches that serve simultaneously as interpretability probes and control mechanisms. First, we conduct a neuron-level analysis, identifying sparse neuron populations that encode dialect-specific features and showing that amplifying or suppressing these neurons can steer model outputs toward target dialects. Second, motivated by the entanglement of dialectal features at the single-neuron level, we apply a vector-steering approach that extracts dialect-specific activation directions and injects them during inference. Together, these methods illuminate the geometry of dialectal knowledge in Arabic LLMs and offer a principled, interpretability-grounded framework for dialect control without requiring dialect-specific fine-tuning.