ChatPaper.aiChatPaper

衡量機器人策略中的語言遷移:將希臘語加入 Cosmos3 視覺-語言-動作策略

Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy

September 7, 2026
作者: Ayoub Kirouane, Georgios Giaples, Christos Petrocheilos
cs.AI

摘要

機器人基礎模型主要以英文訓練與評估,且大多數語言並不存在機器人示範語料庫。我們研究在一個開放式視覺-語言-動作堆疊中加入希臘語,並且僅使用機器改寫的指令,不改變架構。主要挑戰在於量測,而非翻譯。數種看似可行的工具會產生錯誤結論:色彩直方圖指標會獎勵雜訊;單一目標基準在正確希臘語下得分 84.6%,在刻意錯誤的指令下得分 82.6%;訓練損失無法預測希臘語的成功表現;單次執行比較則受隨機種子變異主導。在一套具區辨力的九十項任務套件中,每個實驗組別使用三個隨機種子,未使用希臘語示範的多語言文字塔仍停留在其錯誤指令下限,而僅以希臘語訓練最多僅超越其對照組 2.7 個百分點。雙語訓練則相較其對照組呈現一致的 6.7 至 7.1 個百分點領先幅度,並達到約英文表現的五分之二。此策略也會過度擬合翻譯器的措辭;每項任務以七種措辭訓練,約可將此效能懲罰減半。從語言調適過的世界模型進行暖啟動,以及解凍文字塔,兩者都會降低效能。結果支持低資源機器人策略在地化的兩項實務要求:在信任某項指標之前,先建立一個有保證的無效基準;並跨隨機種子重複驗證低資源語言的結果。
English
Robot foundation models are trained and evaluated predominantly in English, and robot demonstration corpora do not exist for most languages. We study the addition of Greek to an open vision-language-action stack using only machine-rephrased instructions and no architecture changes. The main challenge is measurement rather than translation. Several plausible instruments produce false conclusions: a color-histogram metric rewards noise, a single-goal benchmark scores 84.6% under correct Greek and 82.6% under deliberately wrong instructions, training loss fails to predict Greek success, and single-run comparisons are dominated by seed variation. On a discriminative ninety-task suite with three seeds per arm, a multilingual text tower without Greek demonstrations remains at its wrong-instruction floor, while Greek-only training exceeds its control by at most 2.7 points. Bilingual training yields a consistent 6.7-7.1 point margin over its control and reaches about two fifths of English performance. The policy also overfits the translator's phrasing; training on seven phrasings per task approximately halves this penalty. Warm-starting from a language-adapted world model and unfreezing the text tower both degrade performance. The results support two practical requirements for low-resource robot-policy localization: build a guaranteed null before trusting a metric, and replicate low-resource-language results across seeds.