技能問題:技能在大型語言模型中是否具有語言不變性?
Skill Issue: Are Skills Language-Invariant in LLMs?
August 26, 2026
作者: Bobby Cheng, Adam Gaber, Zhengyuan Liu, Catherine Arnett, Omer Goldman, Cheston Tan, Leshem Choshen
cs.AI
摘要
大型語言模型在不同語言間的知識獲取並不 Consistent,但它們在使用不同語言互動時,其技能組合的差異程度為何?本研究從知識與一般基準測試表現中正交地量化跨語言的技能不一致性。我們透過多語言自我對弈來達成此目標:同一模型的兩個實例在文字型遊戲中競爭,各自透過不同的語言介面進行互動。由於模型、對手、規則、狀態空間與可用動作皆保持固定,此設定得以隔離語言對模型實際表現行為的影響。我們建構了 TextArena 的多語言擴充版本,並評估三個開放權重模型在八種語言、六種遊戲中的表現,涵蓋空間推理、不完全資訊、資源分配與重複互動。我們發現,同一模型在不同語言下可能展現出顯著不同的遊戲實力,勝負差距、無效動作與策略傾向皆呈現系統性變化。詳細分析揭示了空間推理、基於卡牌的決策以及最佳動作選擇中的語言特定失敗模式。在某些設定中,僅改變中間推理語言即可恢復大部分失去的效能,這顯示語言可能影響決策過程的不同階段。這些結果表明,技能差異是發展真正的多語言模型時一個可量測的重大障礙。更深入理解這些差異,有助於我們設計能在不同語言間表現更公平的模型。
English
Large language models access knowledge inconsistently across languages, but to what extent do they differ in their skill sets when interacting with different languages? This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchmark performance. We do this via multilingual self-play: two instances of the same model compete in a text-based game, each interacting through a different language interface. Since the model, opponent, rules, state space, and available actions remain fixed, this setting isolates the effect of language on the model's realized behavior. We build a multilingual extension to TextArena and evaluate three open-weight models across eight languages and six games covering spatial reasoning, imperfect information, resource allocation, and repeated interaction. We find that the same model can exhibit markedly different playing strength across languages, with systematic variation in win--loss margins, invalid actions, and strategic tendencies. Detailed analyses reveal language-specific failures in spatial reasoning, card-conditioned decisions, and optimal move selection. In some settings, changing only the intermediate reasoning language recovers much of the lost performance, suggesting that language can affect different stages of the decision process. These results show that skill discrepancies are a measurable major roadblock in the development of truly multilingual models. Better understanding these discrepancies can help us design models that perform more equitably across languages.