ChatPaper.aiChatPaper

MBA: 실세계 비즈니스 아이디어 창출을 위한 멀티모달 벤치마크 및 에이전트

MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

August 12, 2026
저자: Hojun Choi, Jaeyo Shin, Suin Lee, Hyunjung Shim
cs.AI

초록

대규모 언어 모델(LLM) 기반 에이전트 시스템은 비즈니스 아이디어 창출에 새로운 기회를 열어주었다. 하지만 기존 접근 방식은 실세계 맥락이 본질적으로 다중 모달임에도 불구하고 여전히 텍스트 전용 패러다임에 머물러 있다. 이에 우리는 비즈니스 아이디어 창출 에이전트를 훈련하고 평가하기 위한 최초의 다중 모달 벤치마크인 MBA-Bench를 소개한다. MBA-Bench는 여섯 개 도메인에 걸쳐 30K개의 샘플로 구성되며, 각 도메인은 텍스트만으로는 완전히 전달되지 않는 고유한 시각적 단서를 특징으로 한다. 구체적으로, 우리는 이미지를 자동으로 캡셔닝하고 GPT-4o를 활용하여 검색 쿼리 생성, 시장 증거 검색, 증거 증강 합성 과정을 통해 세 가지 비즈니스 질문 각각에 대해 다섯 개의 참조 아이디어를 생성한다. 이전 연구에 따라, 우리는 MLLM-as-a-Judge를 사용하여 여섯 가지 비즈니스 중심 기준으로 에이전트를 평가한다. 기준이 숨겨지거나 공개되는 설정을 모두 고려하기 위해, 우리는 각각 블라인드(blind) 및 노운(known) 설정을 위한 MBA-b와 MBA-k를 제시한다. 두 모델 모두 창의성과 실현 가능성이라는 두 가지 새로운 보상 목표로 훈련되며, MBA-k는 추가로 공개된 여섯 가지 기준을 최적화하여 총 8개의 목표를 최적화한다. 두 모델은 LoRA 기반 지도 미세 조정으로 훈련한 후, 설정별 보상을 사용하는 그룹 상대 정책 최적화(GRPO)를 적용한다. MBA-Bench에서의 광범위한 실험을 위해 우리는 캡션 전용 또는 다중 모달 입력을 처리하는 두 가지 베이스라인을 구축했으며, 후자는 여러 지표에서 클로즈드소스 모델의 성능에 근접한다. MBA-b와 MBA-k는 캡션 베이스라인을 각각 63.9%와 77.1% 능가하고, 다중 모달 베이스라인을 각각 25.6%와 35.8% 능가한다.
English
Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We thus introduce MBA-Bench, the first multimodal benchmark for training and evaluating business ideation agents, comprising 30K samples across six domains, each domain characterized by distinct visual cues not fully conveyed by text alone. Concretely, we automatically caption images and employ GPT-4o to generate five reference ideas for each of three business questions through retrieval query generation, market evidence retrieval, and evidence-augmented synthesis. Following prior work, we evaluate agents across six business-oriented criteria using MLLM-as-a-Judge. To consider settings where criteria are hidden or disclosed, we present MBA-b and MBA-k for blind and known, respectively. We train both with two novel reward objectives---creativity and feasibility---while MBA-k further optimizes the six disclosed criteria for eight in total. Both are trained via LoRA-based supervised fine-tuning followed by group relative policy optimization with these setting-specific rewards. For extensive experiments on MBA-Bench, we set up two baselines accommodating either captions only or multimodal inputs, with the latter nearing closed-source performance on several metrics. MBA-b and MBA-k outperform caption baselines by 63.9% and 77.1%, and multimodal baselines by 25.6% and 35.8%, respectively.