ACID: Actieconsistentie via Inverse Dynamica voor Planning met Wereldmodellen
ACID: Action Consistency via Inverse Dynamics for Planning with World Models
July 2, 2026
Auteurs: Gawon Seo, Dongwon Kim, Suha Kwak
cs.AI
Samenvatting
Planning op beslissingsmoment met actie-geconditioneerde wereldmodellen is een populaire paradigma geworden voor belichaamde besturing. Echter, de standaard planningskosten beoordelen een kandidaat uitsluitend op basis van hoe dicht de voorspelde eindtoestand bij het doel ligt, waarbij de realiseerbaarheid van de tussentijdse overgangen onopgemerkt blijft – een voorspeld traject kan overtuigend lijken terwijl de uitrol in de omgeving ervan afdwaalt. In dit artikel stellen wij ACID voor, een raamwerk voor planning op beslissingsmoment dat cyclische actieconsistentie introduceert: de actie die door een invers dynamisch model terugwaarts wordt afgeleid uit een voorspelde overgang, zou de actie moeten herstellen waarop die overgang was geconditioneerd. We vouwen deze stapsgewijze residu in de planningskosten via een schaalinvariante adaptieve weging. Over vier actie-geconditioneerde wereldmodellen en zes taken, variërend van rigide en vervormbare manipulatie, gearticuleerde besturing tot visuele navigatie, verbetert ACID consistent de planning en evenaart het de nauwkeurigheid van de basislijn met aanzienlijk minder rekenkracht voor planning.
English
Decision-time planning with action-conditioned world models has become a popular paradigm for embodied control. However, the standard planning cost judges a candidate solely by how close its predicted terminal state lies to the goal, leaving the realizability of the intermediate transitions unchecked -- a predicted trajectory can look convincing while the environment rollout drifts away from it. In this paper, we propose ACID, a decision-time planning framework that introduces cycle action consistency: the action inferred backward from a predicted transition by an inverse dynamics model should recover the one that was conditioned on. We fold this per-step residual into the planning cost via a scale-invariant adaptive weight. Across four action-conditioned world models and six tasks spanning rigid and deformable manipulation, articulated control, and visual navigation, ACID consistently improves planning and matches the baseline's accuracy with substantially less planning compute.