PonderPounce: 로봇 제어를 위한 에피소드 컨텍스트 엔진으로서의 사전 훈련된 MLLM
PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control
August 25, 2026
저자: Suhwan Choi, Jaeyoon Jung, Sungkyung Kim, Yunsung Lee, Youngjae Yu
cs.AI
초록
다중모달 대규모 언어 모델(MLLM)은 긴 시각적 이력을 통합하고, 부분 관측 가능성 하에서 추론하며, 소수의 예시만으로 행동을 추론할 수 있다. 그러나 비전-언어-행동(VLA) 모델은 일반적으로 사전 학습된 표현을 상속받을 뿐, 이러한 맥락적 능력을 에피소드 메모리로 활용하지 않는다. 메모리 의존적 정책은 특수 목적의 이력 메커니즘을 통해 이러한 격차를 해소한다. 반면 PonderPounce는 MLLM의 고유한 인과적 맥락을 로봇 메모리로 재사용한다. System2 MLLM인 Ponder는 에피소드 관측치, 시연, 이전 인지를 고유한 인과적 맥락에 축적하고, 내부 사용을 위한 하위 목표 텍스트와 시연 추론을 생성할 수 있다. System1 VLA인 Pounce는 현재 관측치, 지시, 고유수용감각을 직접 수신하며, Ponder–Pounce 인터페이스를 통해 비동기적으로 가장 최신의 연속 인지 토큰과 그 연령(age)만을 수신한다. 두 모델 모두 특수 목적의 메모리 모듈이나 별도의 브리지 사전 학습 없이 종단 간(end-to-end)으로 공동 학습된다. 최적화된 서빙은 인지 갱신에 78ms, 행동 모델 호출에 25ms의 p50 지연 시간을 달성하여 20Hz 행동 재생을 지원한다. 기본 규모 학습 데이터로 RoboMME에서 PonderPounce는 동일한 Pounce 아키텍처와 인터페이스 하에 9B에서 60.83%, 0.8B에서 50.04%를 달성한 반면, FrameSamp+Modul은 44.51%, 현재 관측치 기반 π_{0.5}는 17.93%를 기록했다. 9배 데이터로는 FrameSamp+Modul의 57.88% 대비 75.54%에 도달한다. RoboCasa-DC에서는 동일한 인터페이스가 행동 감독만으로 학습하여 12.5%를 달성했으며, 이는 가장 강력한 공개 시연 조건부 기준선의 11.6%를 상회한다. 인지가 학습된 널 상태로 대체되면 8.6%로 하락한다.
English
Multimodal large language models (MLLMs) can integrate long visual histories, reason under partial observability, and infer behavior from a few examples. Yet vision-language-action (VLA) models generally inherit pretrained representations without using this contextual capacity as episode memory. Memory-dependent policies address this gap through purpose-built history mechanisms. PonderPounce instead reuses an MLLM's native causal context as robot memory. Ponder, a System2 MLLM, accumulates episode observations, demonstrations, and prior cognition in its native causal context and can generate subgoal text and demonstration reasoning for internal use. Pounce, a System1 VLA, receives the current observation, instruction, and proprioception directly; through the Ponder--Pounce interface, it asynchronously receives only the newest continuous cognition token and its age. Both are jointly trained end to end without a purpose-built memory module or separate bridge pretraining. Optimized serving achieves p50 latencies of 78ms for cognition refresh and 25ms for action-model invocation, supporting 20Hz action playback. On RoboMME with base-scale training data, PonderPounce reaches 60.83% with 9B and 50.04% with 0.8B under the same Pounce architecture and interface, versus 44.51% for FrameSamp+Modul and 17.93% for the current-observation π_{0.5}. With 9x data, it reaches 75.54% versus 57.88% for FrameSamp+Modul. On RoboCasa-DC, the same interface learns from action supervision alone and reaches 12.5% versus 11.6% for the strongest published demonstration-conditioned baseline, falling to 8.6% when cognition is replaced by a learned null state.