ChatPaper.aiChatPaper

技能问题:大型语言模型中的技能是否具有语言不变性?

Skill Issue: Are Skills Language-Invariant in LLMs?

August 26, 2026
作者: Bobby Cheng, Adam Gaber, Zhengyuan Liu, Catherine Arnett, Omer Goldman, Cheston Tan, Leshem Choshen
cs.AI

摘要

大语言模型在不同语言间的知识获取并不一致,但同一模型在使用不同语言交互时,其技能集究竟存在多大差异?本研究从知识与通用基准性能两个维度之外,正交地量化了跨语言技能不一致性。我们通过多语言自我对弈实现这一目标:同一模型的两个实例在基于文本的游戏中相互竞争,各自通过不同的语言界面进行交互。由于模型、对手、规则、状态空间和可用动作均保持不变,该设置能够隔离语言对模型实际行为的影响。我们构建了TextArena的多语言扩展,并在八种语言、六类游戏中评估了三个开放权重模型,涵盖空间推理、不完美信息、资源分配和重复交互等场景。研究发现,同一模型在不同语言下可能表现出显著差异的博弈水平,胜负差距、无效动作和策略倾向均呈现系统性变化。详细分析揭示了空间推理、基于牌面条件的决策以及最优落子选择中的语言特异性缺陷。在某些设置下,仅改变中间推理语言即可恢复大部分性能损失,这表明语言可能影响决策过程的不同阶段。上述结果表明,技能差异是可测量的重大障碍,阻碍了真正多语言模型的发展。深入理解这些差异有助于我们设计在不同语言间表现更公平的模型。
English
Large language models access knowledge inconsistently across languages, but to what extent do they differ in their skill sets when interacting with different languages? This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchmark performance. We do this via multilingual self-play: two instances of the same model compete in a text-based game, each interacting through a different language interface. Since the model, opponent, rules, state space, and available actions remain fixed, this setting isolates the effect of language on the model's realized behavior. We build a multilingual extension to TextArena and evaluate three open-weight models across eight languages and six games covering spatial reasoning, imperfect information, resource allocation, and repeated interaction. We find that the same model can exhibit markedly different playing strength across languages, with systematic variation in win--loss margins, invalid actions, and strategic tendencies. Detailed analyses reveal language-specific failures in spatial reasoning, card-conditioned decisions, and optimal move selection. In some settings, changing only the intermediate reasoning language recovers much of the lost performance, suggesting that language can affect different stages of the decision process. These results show that skill discrepancies are a measurable major roadblock in the development of truly multilingual models. Better understanding these discrepancies can help us design models that perform more equitably across languages.