ChatPaper.aiChatPaper

EgoForce: Onderarm-geleide 3D-handpose in cameraruimte vanuit een monoscopische egocentrische camera

EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera

May 12, 2026
Auteurs: Christen Millerdurai, Shaoxiang Wang, Yaxu Xie, Vladislav Golyanik, Didier Stricker, Alain Pagani
cs.AI

Samenvatting

Het reconstrueren van de absolute 3D-houding en -vorm van de handen vanuit het perspectief van de gebruiker met behulp van een enkele hoofdmontagecamera is cruciaal voor praktische egocentrische interactie in AR/VR, telepresence en handgerichte manipulatie taken, waarbij sensoren compact en onopvallend moeten blijven. Hoewel monoculaire RGB-methoden vooruitgang hebben geboekt, worden ze beperkt door diepte-schaalambiguïteit en hebben ze moeite om te generaliseren over de uiteenlopende optische configuraties van hoofdmontageapparaten. Hierdoor vereisen modellen doorgaans uitgebreide training op apparaatspecifieke datasets, die kostbaar en arbeidsintensief zijn om te verkrijgen. Dit artikel pakt deze uitdagingen aan door EgoForce te introduceren, een moniculair 3D-handreconstructieframework dat robuuste, absolute 3D-handhoudingen en hun positie vanuit het perspectief van de gebruiker (cameraruimte) herstelt. EgoForce werkt met vissenoog-, perspectivische en vervormde groothoek-cameramodellen met behulp van één enkel verenigd netwerk. Onze aanpak combineert een differentieerbare onderarmrepresentatie die de handhouding stabiliseert, een verenigde arm-handtransformator die zowel hand- als onderarmgeometrie voorspelt vanuit één enkel egocentrisch beeld, waardoor diepte-schaalambiguïteit wordt verminderd, en een gesloten-vormoplosser in stralingsruimte die absolute 3D-houdingsherstel mogelijk maakt voor verschillende hoofdmontagecameramodellen. Experimenten op drie egocentrische benchmarks tonen aan dat EgoForce de modernste 3D-nauwkeurigheid bereikt, met een reductie van de MPJPE in cameraruimte tot 28% op de HOT3D-dataset in vergelijking met eerdere methoden, en consistente prestaties levert over cameraconfiguraties heen. Voor meer details, bezoek de projectpagina op https://dfki-av.github.io/EgoForce.
English
Reconstructing the absolute 3D pose and shape of the hands from the user's viewpoint using a single head-mounted camera is crucial for practical egocentric interaction in AR/VR, telepresence, and hand-centric manipulation tasks, where sensing must remain compact and unobtrusive. While monocular RGB methods have made progress, they remain constrained by depth-scale ambiguity and struggle to generalize across the diverse optical configurations of head-mounted devices. As a result, models typically require extensive training on device-specific datasets, which are costly and laborious to acquire. This paper addresses these challenges by introducing EgoForce, a monocular 3D hand reconstruction framework that recovers robust, absolute 3D hand pose and its position from the user's (camera-space) viewpoint. EgoForce operates across fisheye, perspective, and distorted wide-FOV camera models using a single unified network. Our approach combines a differentiable forearm representation that stabilizes hand pose, a unified arm-hand transformer that predicts both hand and forearm geometry from a single egocentric view, mitigating depth-scale ambiguity, and a ray space closed-form solver that enables absolute 3D pose recovery across diverse head-mounted camera models. Experiments on three egocentric benchmarks show that EgoForce achieves state-of-the-art 3D accuracy, reducing camera-space MPJPE by up to 28% on the HOT3D dataset compared to prior methods and maintaining consistent performance across camera configurations. For more details, visit the project page at https://dfki-av.github.io/EgoForce.