ChatPaper.aiChatPaper

MBA: 実世界ビジネスアイデア創出のためのマルチモーダルベンチマークとエージェント

MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

August 12, 2026
著者: Hojun Choi, Jaeyo Shin, Suin Lee, Hyunjung Shim
cs.AI

要旨

大規模言語モデル(LLM)を基盤とするエージェント型システムは、ビジネスアイデア創出に新たな可能性をもたらしている。しかし、実世界の文脈が本質的にマルチモーダルであるにもかかわらず、既存の手法はテキストのみのパラダイムに留まっている。そこで我々は、ビジネスアイデア創出エージェントの訓練と評価のための初のマルチモーダルベンチマークであるMBA-Benchを導入する。MBA-Benchは、6つのドメインにわたる30Kサンプルから構成され、各ドメインはテキストだけでは完全には伝わらない特徴的な視覚的手がかりによって特徴づけられる。具体的には、画像に自動的にキャプションを付与し、GPT-4oを用いて、検索クエリ生成、市場エビデンス検索、エビデンス拡張合成を通じて、3つのビジネス質問のそれぞれに対して5つの参照アイデアを生成する。先行研究に従い、MLLM-as-a-Judgeを用いて、6つのビジネス指向の基準にわたってエージェントを評価する。基準が非表示または開示される設定を考慮するため、それぞれブラインド用のMBA-bと既知用のMBA-kを提示する。両者は、新規の報酬目標である創造性と実現可能性を用いて訓練され、MBA-kはさらに開示された6つの基準を最適化して合計8つとする。両者とも、LoRAベースの教師ありファインチューニングと、それに続く設定固有の報酬を用いたグループ相対方策最適化によって訓練される。MBA-Benchでの広範な実験のために、キャプションのみまたはマルチモーダル入力のいずれかに対応する2つのベースラインを設定し、後者は複数のメトリクスにおいてクローズドソースの性能に迫る。MBA-bとMBA-kは、それぞれキャプションベースラインを63.9%および77.1%、マルチモーダルベースラインを25.6%および35.8%上回る。
English
Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We thus introduce MBA-Bench, the first multimodal benchmark for training and evaluating business ideation agents, comprising 30K samples across six domains, each domain characterized by distinct visual cues not fully conveyed by text alone. Concretely, we automatically caption images and employ GPT-4o to generate five reference ideas for each of three business questions through retrieval query generation, market evidence retrieval, and evidence-augmented synthesis. Following prior work, we evaluate agents across six business-oriented criteria using MLLM-as-a-Judge. To consider settings where criteria are hidden or disclosed, we present MBA-b and MBA-k for blind and known, respectively. We train both with two novel reward objectives---creativity and feasibility---while MBA-k further optimizes the six disclosed criteria for eight in total. Both are trained via LoRA-based supervised fine-tuning followed by group relative policy optimization with these setting-specific rewards. For extensive experiments on MBA-Bench, we set up two baselines accommodating either captions only or multimodal inputs, with the latter nearing closed-source performance on several metrics. MBA-b and MBA-k outperform caption baselines by 63.9% and 77.1%, and multimodal baselines by 25.6% and 35.8%, respectively.