ChatPaper.aiChatPaper

MBA:真實世界商業構想的多模態基準與智能體

MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

August 12, 2026
作者: Hojun Choi, Jaeyo Shin, Suin Lee, Hyunjung Shim
cs.AI

摘要

由大型語言模型(LLM)驅動的代理系統為商業構思開闢了新的契機。然而,現有方法仍局限於純文字模式,儘管現實世界情境本質上具有多模態特性。為此,我們提出MBA-Bench,這是首個用於訓練與評估商業構思代理的多模態基準,涵蓋六個領域共30,000個樣本,每個領域皆具備無法僅靠文字完整傳達的獨特視覺線索。具體而言,我們自動為圖像生成標註,並透過檢索查詢生成、市場證據檢索與證據增強合成,使用GPT-4o針對三個商業問題各生成五個參考構思。遵循既有研究,我們以多模態大型語言模型作為評審(MLLM-as-a-Judge),依六項商業導向標準評估代理效能。為考量標準隱藏或公開的不同設定,我們分別提出MBA-b(盲測)與MBA-k(已知標準)。兩者皆以兩個新穎的獎勵目標——創造性與可行性——進行訓練,而MBA-k另針對六項公開標準進行優化,共計八個目標。兩者均採用基於LoRA的有監督微調,再以針對各設定獎勵的群體相對策略優化進行訓練。為在MBA-Bench上進行全面實驗,我們設置了兩個基線,分別容納僅標註輸入與多模態輸入,其中後者在多項指標上接近閉源效能。MBA-b與MBA-k相較標註基線分別提升63.9%與77.1%,相較多模態基線則分別提升25.6%與35.8%。
English
Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We thus introduce MBA-Bench, the first multimodal benchmark for training and evaluating business ideation agents, comprising 30K samples across six domains, each domain characterized by distinct visual cues not fully conveyed by text alone. Concretely, we automatically caption images and employ GPT-4o to generate five reference ideas for each of three business questions through retrieval query generation, market evidence retrieval, and evidence-augmented synthesis. Following prior work, we evaluate agents across six business-oriented criteria using MLLM-as-a-Judge. To consider settings where criteria are hidden or disclosed, we present MBA-b and MBA-k for blind and known, respectively. We train both with two novel reward objectives---creativity and feasibility---while MBA-k further optimizes the six disclosed criteria for eight in total. Both are trained via LoRA-based supervised fine-tuning followed by group relative policy optimization with these setting-specific rewards. For extensive experiments on MBA-Bench, we set up two baselines accommodating either captions only or multimodal inputs, with the latter nearing closed-source performance on several metrics. MBA-b and MBA-k outperform caption baselines by 63.9% and 77.1%, and multimodal baselines by 25.6% and 35.8%, respectively.