ChatPaper.aiChatPaper

로봇 정책에서의 언어 전이 측정: Cosmos3 비전-언어-행동 정책에 그리스어 추가하기

Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy

September 7, 2026
저자: Ayoub Kirouane, Georgios Giaples, Christos Petrocheilos
cs.AI

초록

로봇 파운데이션 모델은 주로 영어로 학습되고 평가되며, 대부분의 언어에는 로봇 시연 코퍼스가 존재하지 않는다. 우리는 아키텍처 변경 없이 기계가 다시 표현한 지시문만을 사용하여 오픈 비전-언어-행동 스택에 그리스어를 추가하는 것을 연구한다. 주요 과제는 번역이 아니라 측정이다. 몇 가지 그럴듯한 측정 도구는 잘못된 결론을 낳는다: 색상 히스토그램 지표는 잡음에 보상을 주고, 단일 목표 벤치마크는 올바른 그리스어에서 84.6%, 의도적으로 잘못된 지시문에서 82.6%를 기록하며, 훈련 손실은 그리스어 성공을 예측하지 못하고, 단일 실행 비교는 시드 변동에 지배된다. 실험군당 세 개의 시드를 사용한 판별적 90개 과제 스위트에서, 그리스어 시연이 없는 다국어 텍스트 타워는 잘못된 지시문 하한에 머무는 반면, 그리스어만으로 훈련한 경우는 대조군을 최대 2.7점 상회한다. 이중 언어 훈련은 대조군 대비 일관되게 6.7–7.1점의 우위를 보이며 영어 성능의 약 5분의 2에 도달한다. 또한 정책은 번역기의 표현 방식에 과적합하며, 과제당 일곱 가지 표현으로 훈련하면 이 페널티가 대략 절반으로 줄어든다. 언어 적응된 월드 모델에서 웜 스타트하는 것과 텍스트 타워의 동결을 해제하는 것은 모두 성능을 저하시킨다. 결과는 저자원 로봇 정책 현지화를 위한 두 가지 실용적 요구사항을 뒷받침한다: 지표를 신뢰하기 전에 보장된 널을 구축할 것, 그리고 저자원 언어 결과를 여러 시드에 걸쳐 재현할 것.
English
Robot foundation models are trained and evaluated predominantly in English, and robot demonstration corpora do not exist for most languages. We study the addition of Greek to an open vision-language-action stack using only machine-rephrased instructions and no architecture changes. The main challenge is measurement rather than translation. Several plausible instruments produce false conclusions: a color-histogram metric rewards noise, a single-goal benchmark scores 84.6% under correct Greek and 82.6% under deliberately wrong instructions, training loss fails to predict Greek success, and single-run comparisons are dominated by seed variation. On a discriminative ninety-task suite with three seeds per arm, a multilingual text tower without Greek demonstrations remains at its wrong-instruction floor, while Greek-only training exceeds its control by at most 2.7 points. Bilingual training yields a consistent 6.7-7.1 point margin over its control and reaches about two fifths of English performance. The policy also overfits the translator's phrasing; training on seven phrasings per task approximately halves this penalty. Warm-starting from a language-adapted world model and unfreezing the text tower both degrade performance. The results support two practical requirements for low-resource robot-policy localization: build a guaranteed null before trusting a metric, and replicate low-resource-language results across seeds.