ChatPaper.aiChatPaper

Leren vouwen: prijswinnende oplossing op de LeHome Challenge 2026 (1e plaats online, 2e offline)

Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)

June 25, 2026
Auteurs: Ilia Larchenko
cs.AI

Samenvatting

Ik beschrijf mijn oplossing voor de LeHome Challenge 2026, een ICRA 2026-competitie over bimanueel kledingstukvouwen. Het systeem eindigde als 1e van 62 teams in de online (simulatie)ronde en als 2e in de finale in de echte wereld. Het verbetert een visie-taal-actiebeleid (VLA) met een versterkingsleer-lus. Het beleid fungeert als zijn eigen waardefunctie: hetzelfde netwerk dat acties voorspelt, voorspelt ook succes, voortgang en een aantal taakrelevante toekomstige grootheden, en deze voorspellingen drijven voordeelschatting, foutdetectie in real-time en kandidaatselectie aan. Het werk combineert grotendeels bestaande RL-ideeën met technische en optimalisatiebijdragen die samen als één recept of afzonderlijk kunnen worden gebruikt: AWR + RECAP gecombineerd voor stroomafstemmende VLA; een asynchrone gedistribueerde trainings-/uitrolpijplijn via HuggingFace Hub; optimalisatie van hyperparameters tijdens inferentie via Thompson-steekproef; een sim-naar-real-recept met gereedschap voor camera-uitlijning, zware augmentatie en DAgger-achtige HIL-gegevensverzameling.
English
I describe my solution to the LeHome Challenge 2026, an ICRA 2026 competition on bimanual garment folding. The system placed 1st of 62 teams in the online (simulation) round and 2nd in the real-world final. It improves a vision-language-action (VLA) policy with a reinforcement-learning loop. The policy is its own value function: the same network that predicts actions also predicts success, progress, and a few task-relevant future quantities, and those predictions drive advantage estimation, live failure detection, and candidate selection. The work mostly recombines existing RL ideas with engineering and optimization contributions that can be used together as one recipe or individually: AWR + RECAP combined for flow-matching VLA; an asynchronous distributed training / rollout pipeline through HuggingFace Hub; inference-time hyperparameters optimization via Thompson sampling; a sim-to-real recipe with camera-alignment tooling, heavy augmentation and DAgger-like HIL data collection.