선택, 압축, 재투자: 장시간 비디오 MLLM에서의 시각 토큰 할당에 관한 통제 연구
Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs
September 3, 2026
저자: Prakhar Khatri
cs.AI
초록
장시간 비디오 언어 모델은 모든 프레임을 볼 수 없다. 1시간짜리 비디오를 초당 한 번 샘플링하면 3,600개의 이미지가 되며, 시스템은 그 이미지 풀에서 고정된 아주 작은 일부만 유지한다. 어떤 프레임이 그 일부에 남는지는 보통 전처리 세부 사항으로 취급된다. 우리는 그것이 과연 전처리 세부 사항인지 검증한다. 기존에 발표된 선택 방식들은 프레임 점수화 방식, 프롬프트 경계, 해상도 정책, 응답 모델을 한꺼번에 바꾸기 때문에 비교가 어렵다. 따라서 우리는 각 요소를 고정한 채 선택, 공간 압축, 절약분의 재투자라는 세 결정을 하나씩 변경하면서, 학습이 필요 없는 여섯 가지 선택 규칙, 세 가지 장시간 비디오 벤치마크, 두 가지 응답 모델에 걸쳐 실험했다. 선택이 가장 큰 단일 영향 요소였다. LongVideoBench의 1시간 길이 버킷에서 쿼리로 선택된 8개 프레임은 균등 간격으로 뽑은 16개 프레임보다 6.9포인트 높은 성능을 보였다. 또한 수정 없이 사용된 수십 년 된 희소 근사 알고리즘인 Orthogonal Matching Pursuit은 세 벤치마크 전반에서 우리가 비교한 모든 특수 목적 선택 방식과 성능이 동일하거나 1포인트 이내의 차이를 보였다. 공간 압축은 비용이 거의 들지 않는다. 타임스탬프를 고정한 채 각 프레임의 공간 예산을 절반으로 줄이면 정확도 손실은 최대 0.44포인트다. 그 예산이 다시 정확도로 전환되는 지점은 재투자다. 절약된 토큰을 두 배로 많은 압축 프레임에 사용하면, 측정 비용이 원래 8개 프레임을 사용할 때보다 높지 않은데도 정확도가 추가로 2~3포인트 오른다. 압축은 절약분을 이렇게 사용할 때만 이득이 된다. 실험 과정에서 우리 AKS 기준 구현에서 발견된 버그와, 동일한 예산에서 동일한 공개 규칙을 실행하는 두 평가 환경 간의 0.07~3.74포인트 차이는, 이러한 비교가 서로 다른 논문에 걸쳐서가 아니라 하나의 통제된 평가 환경 안에서 수행되어야 하는 이유를 보여준다.
English
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a system keeps only a small fixed slice of that pool. Which frames survive that slice is usually treated as a preprocessing detail; we test whether it should be. Published selectors make the comparison hard because they change the frame scorer, the prompt boundary, the resolution policy, and the answering model all at once. We hold each fixed and vary one decision at a time: selection, spatial compression, and reinvestment of the savings, across six training-free selection rules, three long-video benchmarks, and two answering models. Selection is the largest single lever: on LongVideoBench's hour-long bin, eight query-selected frames beat sixteen uniformly spaced ones by 6.9 points, and Orthogonal Matching Pursuit, an unmodified decades-old sparse-approximation algorithm, matches or comes within a point of every purpose-built selector we compare it against, across all three benchmarks. Compression is close to free: halving each frame's spatial budget at fixed timestamps costs at most 0.44 points. Reinvestment is where that budget turns back into accuracy: spending the freed tokens on twice as many compressed frames, at a measured cost no higher than the original eight, returns a further two to three points; compression only pays off once its savings are spent this way. Along the way, an implementation bug in our own AKS baseline and a 0.07 to 3.74 point gap between two harnesses running the same published rules at the same budget show why these comparisons need to happen inside one controlled harness rather than across papers.