ChatPaper.aiChatPaper

スキル問題:LLMにおけるスキルは言語非依存なのか?

Skill Issue: Are Skills Language-Invariant in LLMs?

August 26, 2026
著者: Bobby Cheng, Adam Gaber, Zhengyuan Liu, Catherine Arnett, Omer Goldman, Cheston Tan, Leshem Choshen
cs.AI

要旨

大規模言語モデルは言語によって知識へのアクセスが一貫していないが、異なる言語で対話するとき、そのスキルセットはどの程度異なるのだろうか。本研究は、知識や一般的なベンチマーク性能から直交する形で、言語横断的なスキル不一致を定量化する。このために、多言語自己対戦を用いる。すなわち、同じモデルの2つのインスタンスがテキストベースのゲームで対戦し、それぞれが異なる言語インターフェースを通じて対話する。モデル、対戦相手、ルール、状態空間、利用可能なアクションはすべて固定されているため、この設定はモデルが実際に発現する行動に対する言語の影響を分離する。我々はTextArenaに多言語拡張を構築し、空間推論、不完全情報、資源配分、反復的相互作用を対象とする6つのゲームにわたり、8言語で3つのオープンウェイトモデルを評価する。その結果、同じモデルでも言語によってプレイの強さが著しく異なることがあり、勝敗の差、無効なアクション、戦略的傾向に系統的な変動が見られる。詳細な分析により、空間推論、カード条件付きの意思決定、最適手の選択における言語固有の失敗が明らかになった。一部の設定では、中間推論言語を変更するだけで失われたパフォーマンスの大部分が回復し、言語が意思決定プロセスの異なる段階に影響を及ぼし得ることが示唆される。これらの結果は、スキルの不一致が、真の多言語モデルの開発における測定可能な大きな障害であることを示している。このような不一致をより深く理解することは、言語間でより公平に性能を発揮するモデルの設計に役立つだろう。
English
Large language models access knowledge inconsistently across languages, but to what extent do they differ in their skill sets when interacting with different languages? This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchmark performance. We do this via multilingual self-play: two instances of the same model compete in a text-based game, each interacting through a different language interface. Since the model, opponent, rules, state space, and available actions remain fixed, this setting isolates the effect of language on the model's realized behavior. We build a multilingual extension to TextArena and evaluate three open-weight models across eight languages and six games covering spatial reasoning, imperfect information, resource allocation, and repeated interaction. We find that the same model can exhibit markedly different playing strength across languages, with systematic variation in win--loss margins, invalid actions, and strategic tendencies. Detailed analyses reveal language-specific failures in spatial reasoning, card-conditioned decisions, and optimal move selection. In some settings, changing only the intermediate reasoning language recovers much of the lost performance, suggesting that language can affect different stages of the decision process. These results show that skill discrepancies are a measurable major roadblock in the development of truly multilingual models. Better understanding these discrepancies can help us design models that perform more equitably across languages.