ChatPaper.aiChatPaper

ロボットポリシーにおける言語転移の測定:Cosmos3視覚言語行動ポリシーへのギリシャ語の追加

Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy

September 7, 2026
著者: Ayoub Kirouane, Georgios Giaples, Christos Petrocheilos
cs.AI

要旨

ロボット基盤モデルは主に英語で訓練および評価されており、ほとんどの言語にはロボット実演コーパスが存在しない。我々は、機械的に言い換えられた指示のみを用い、アーキテクチャを変更せずに、オープンな視覚–言語–行動スタックへギリシャ語を追加することを検討する。主な課題は翻訳ではなく測定である。いくつかのもっともらしい評価手段は誤った結論を導く。すなわち、カラーヒストグラム指標はノイズを高く評価し、単一目標ベンチマークは正しいギリシャ語で84.6%、意図的に誤った指示で82.6%を記録し、訓練損失はギリシャ語での成功を予測できず、単一実行の比較はシード変動に支配される。識別的な90タスクスイートにおいて、各アームにつき3シードを用いると、ギリシャ語の実演を含まない多言語テキストタワーは誤指示時の下限にとどまる一方、ギリシャ語のみの訓練は対照条件を最大2.7ポイント上回る。バイリンガル訓練は対照条件に対して一貫して6.7~7.1ポイントのマージンをもたらし、英語性能の約5分の2に達する。このポリシーは翻訳者の言い回しにも過適合する。タスクごとに7つの言い回しで訓練すると、このペナルティはおよそ半減する。言語適応済み世界モデルからのウォームスタートとテキストタワーの凍結解除はいずれも性能を低下させる。この結果は、低リソースのロボットポリシーをローカライズするための2つの実用的要件を支持する。すなわち、指標を信頼する前に保証されたヌルを構築すること、および低リソース言語の結果をシード間で再現することである。
English
Robot foundation models are trained and evaluated predominantly in English, and robot demonstration corpora do not exist for most languages. We study the addition of Greek to an open vision-language-action stack using only machine-rephrased instructions and no architecture changes. The main challenge is measurement rather than translation. Several plausible instruments produce false conclusions: a color-histogram metric rewards noise, a single-goal benchmark scores 84.6% under correct Greek and 82.6% under deliberately wrong instructions, training loss fails to predict Greek success, and single-run comparisons are dominated by seed variation. On a discriminative ninety-task suite with three seeds per arm, a multilingual text tower without Greek demonstrations remains at its wrong-instruction floor, while Greek-only training exceeds its control by at most 2.7 points. Bilingual training yields a consistent 6.7-7.1 point margin over its control and reaches about two fifths of English performance. The policy also overfits the translator's phrasing; training on seven phrasings per task approximately halves this penalty. Warm-starting from a language-adapted world model and unfreezing the text tower both degrade performance. The results support two practical requirements for low-resource robot-policy localization: build a guaranteed null before trusting a metric, and replicate low-resource-language results across seeds.