ChatPaper.aiChatPaper

FoldingAgent:從示範影片中推斷參數化摺紙程序

FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos

August 31, 2026
作者: Maya Moriya, Sigal Raab, Yael Vinker, Tali Dekel
cs.AI

摘要

我們提出 FoldingAgent,一個代理式框架,能直接從摺紙示範影片中推斷出明確的參數化摺疊程式。我們的框架利用預先訓練的視覺語言模型(VLM)的推理能力,該模型配備了一套專門工具,使代理能模擬幾何轉換、驗證物理合理性、檢索並比較視覺內容,以及評估自身預測。為了將視覺內容轉化為摺疊程式,我們定義了一個參數空間,其中包含紙張的幾何結構以及一組參數化摺疊動作。不同於預測靜態摺痕圖案的模型,我們的代理依序運作,並具備重新規劃其動作的能力,從而有效減輕多步驟摺疊中固有的累積誤差。我們的方法朝縮小人類摺紙知識與計算方法之間的差距邁出了一步:人類摺紙知識主要透過非結構化的視覺示範來分享,而計算方法通常依賴結構化的參數化表示,例如摺痕圖或可執行的參數化計畫。我們在 PurelandFold 上評估我們的方法,這是一個新策劃的基準,包含多樣化的 Pureland 摺紙影片,並帶有真實幾何與動作標籤。我們的結果表明,透過結合 VLM 推理、一套專門工具與物理模擬,我們能成功將非結構化的視覺示範轉化為可執行且物理上合理的摺疊程序。
English
We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from origami demonstration videos. Our framework leverages the reasoning power of a pre-trained Vision-Language Model (VLM) equipped with a suite of specialized tools that enable the agent to simulate geometric transitions, verify physical plausibility, retrieve and compare visual content, and evaluate its own predictions. To translate visual content into folding programs, we define a parametric space that consists of the paper's geometry and a set of parametric folding actions. Unlike models that predict static crease patterns, our agent operates sequentially and possesses the ability to re-plan its actions, effectively mitigating the compounding errors inherent in multi-step folding. Our approach takes a step toward closing the gap between human origami knowledge, which is primarily shared through unstructured visual demonstrations, and computational methods, which typically rely on structured, parametric representations such as a crease pattern or an executable parametric plan. We evaluate our approach on PurelandFold, a newly curated benchmark of diverse Pureland origami videos with ground-truth geometry and action labels. Our results demonstrate that by combining VLM reasoning with a set of specialized tools and physical simulation, we can successfully transform unstructured visual demonstrations into executable, physically plausible folding procedures.