RaysUp: Ultralichte universele feature-omhoogbemonstering via geometriebewuste straalrepresentatie
RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation
June 22, 2026
Auteurs: Yuchuan Ding, Linfei Li, Lin Zhang, Ying Shen
cs.AI
Samenvatting
Voorgetrainde Vision Foundation Modellen (VFM's) zijn centraal komen te staan in moderne computervisie vanwege hun krachtige semantische representaties en sterke generalisatievermogen. Hun patch-gebaseerde of gepoolde uitvoer is echter inherent laag in resolutie, wat hun effectiviteit beperkt in taken die fijnmazige, pixel-niveau redenering vereisen. Bestaande benaderingen voor kenmerkopwaardering verminderen de semantische betrouwbaarheid of zijn afhankelijk van VFM-specifieke hertraining en zware architecturen, wat de efficiëntie en schaalbaarheid belemmert. Om deze uitdagingen aan te pakken, stellen we RaysUp voor, een ultralicht, taakonafhankelijk en VFM-onafhankelijk raamwerk voor kenmerkopwaardering dat kenmerken op hoge resolutie reconstrueert bij willekeurige resoluties. In tegenstelling tot conventionele 2D-interpolatie of aandachtsgebaseerde schema's tilt RaysUp kenmerkreconstructie naar een geometriebewust stralendomein. Specifiek introduceren we een Ruimtelijk Ontkoppelde Geleidingsencoder voor richtingsbewuste geleidingscodering, een Willekeurige-Resolutie Kruisaandachtmechanisme voor resolutieflexibele reconstructie, en een nieuwe Stralenpositiecodering (RayPE) die impliciete 3D-geometrische voorkennis injecteert via 6D Plücker-straalcoördinaten. Ten slotte zorgt een Geometriebewuste Buurtaandachtmodule voor inhoudadaptieve bilaterale aggregatie met behoud van geometrische consistentie. Uitgebreide experimenten over diverse dichte voorspellingstaken tonen aan dat RaysUp de nieuwste prestaties levert terwijl het slechts 16% van de parameters van AnyUp gebruikt en ongeveer 7x snellere inferentie levert. Deze resultaten benadrukken een aanzienlijk verbeterde nauwkeurigheid-efficiëntie afweging en vestigen RaysUp als een praktische en schaalbare oplossing voor universele kenmerkopwaardering. Code is beschikbaar op https://github.com/MAP-RaysUp/RaysUp.
English
Pre-trained Vision Foundation Models (VFMs) have become central to modern computer vision due to their powerful semantic representations and strong generalization ability. However, their patchified or pooled outputs are inherently low-resolution, limiting their effectiveness in tasks requiring fine-grained, pixel-level reasoning. Existing feature upsampling approaches either degrade semantic fidelity or rely on VFM-specific retraining and heavy architectures, hindering efficiency and scalability. To address these challenges, we propose RaysUp, an ultra-lightweight, task-agnostic, and VFM-agnostic feature upsampling framework that reconstructs high-resolution feature maps at arbitrary resolutions. Unlike conventional 2D interpolation or attention-based schemes, RaysUp lifts feature reconstruction into a geometry-aware ray domain. Specifically, we introduce a Spatially Decoupled Guidance Encoder for direction-aware guidance encoding, an Any-Resolution Cross-Attention mechanism for resolution-flexible reconstruction, and a novel Ray Positional Encoding (RayPE) that injects implicit 3D geometric priors via 6D Plucker ray coordinates. Finally, a Geometry-Aware Neighborhood Attention module further ensures content-adaptive bilateral aggregation while preserving geometric consistency. Extensive experiments across diverse dense prediction tasks demonstrate that RaysUp achieves state-of-the-art performance while using only 16% of the parameters of AnyUp and delivering approximately 7x faster inference. These results highlight a substantially improved accuracy-efficiency trade-off and establish RaysUp as a practical and scalable solution for universal feature upsampling. Code is available at https://github.com/MAP-RaysUp/RaysUp.