옴니 인터랙션 에이전트 기술 보고서
Omni Interaction Agent Technical Report
September 8, 2026
저자: Orantqing, Shengpeng Ji, Junlong Tong, Jialong Zuo, Dongjie Fu, Di Cao, Yangzhuo Li, Shangda Wu, Franz, Evan, Theron Veyra, Changhao Pan, Jingyu Lu, Dongchao Yang, Zhifei Xie, Yang Tan, Xiaoyu Shen, Xiaoda Yang, Wenfu Wang, Teddysun, Steveyves, Zhou Zhao, Bryanytian
cs.AI
초록
본 연구에서 우리는 옴니 지각, 실시간 상호작용, 에이전트 능력을 단일 프레임워크로 통합한 엔드투엔드 모델인 Gander를 제안한다. 턴 기반의 전통적 패러다임과 달리, Gander는 비디오, 음성, 텍스트를 포함한 여러 모달리티에 걸쳐 스트리밍 입력을 지속적으로 수신하여 일상 대화와 복잡한 워크플로 지향 에이전트 시나리오 모두에서 자연스러운 전이중 상호작용을 가능하게 한다. 사용자는 언제든지 모델을 중단시킬 수 있으며, 모델은 또한 능동적으로 중간 피드백을 제공하거나 후속 질문을 할 수 있다. 이러한 능력을 기본적으로 지원하기 위해 Gander는 두 가지 핵심 아키텍처 설계를 채택한다. 1) Gander는 Cerebellum-Brain 협업 프레임워크를 사용하며, 여기서 Cerebellum은 실시간 상호작용과 옴니 대화 능력을 담당하고 Brain은 복잡한 추론과 상위 수준의 에이전트 작업을 처리한다. 두 구성 요소는 도구 호출과 에이전트 오케스트레이션 런타임을 통해 지속적으로 상호작용한다. 2) Cerebellum은 스트리밍 Thinker-Talker 아키텍처에 기반하며, 사용자 입력과 모델 출력은 청크 수준에서 순서가 있는 토큰 스트림으로 추가로 평탄화되어 낮은 지연 시간과 연속적 상호작용을 위한 통합 표현을 제공한다. 우리는 대화 능력, 옴니 이해, 상호작용 능력, 에이전트 지능의 네 가지 차원에 걸쳐 Gander를 포괄적으로 평가한다. 내부 인간 평가는 Gander가 SOTA 오픈 소스 모델의 자연스럽고 표현력 있는 음성 대화 능력을 유지하면서 옴니 상호작용에서 경쟁력 있는 성능을 달성함을 보여준다. Gander는 또한 배경 잡음 간섭, 다자간 상호작용, 백채널 커뮤니케이션을 포함한 어려운 실제 시나리오에서 강건성을 입증한다. 우리는 커뮤니티의 추가 연구와 개발을 촉진하기 위해 Gander를 모델, 코드, 데이터와 함께 공개한다.
English
In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and agentic capabilities within a single framework. In contrast to turn-based conventional paradigms, Gander continuously receives streaming inputs across multiple modalities, including video, speech, and text, enabling natural full-duplex interaction in both everyday conversations and complex workflow-oriented agent scenarios. Users can interrupt the model at any time, while the model can also proactively provide intermediate feedback or ask follow up questions. To natively support these capabilities, Gander adopts two key architectural designs: 1) It employs a Cerebellum-Brain collaborative framework, in which the Cerebellum is responsible for realtime interaction and omni conversational capabilities, while the Brain handles complex reasoning and higher-level agentic tasks. The two components interact continuously through tool calling and the agent orchestration runtime. 2) The Cerebellum is built upon a streaming Thinker-Talker architecture, user inputs and model outputs are further flattened into an ordered token stream at the chunk level, providing a unified representation for low latency, continuous interaction. We conduct comprehensive evaluations of Gander across four dimensions: conversational ability, omni understanding, interactive capability, and agentic intelligence. Internal human evaluations demonstrate that Gander maintains the natural and expressive spoken dialogue capabilities of SOTA open source models while achieving competitive performance in omni interaction. Gander also demonstrates robustness in challenging real-world scenarios, including background noise interference, multi-party interactions, and backchannel communication. We release Gander together with its models, code, and data to facilitate further research and development in the community.