MatrAIx:83億のペルソナエージェントによる世界シミュレーション
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
August 4, 2026
著者: Xiaomin Li, Yuexing Hao, Jianheng Hou, Jintao Huang, Qianfeng Wen, Shirley Huang, Yifan Liu, Xiaoyi Liu, Yilan Fan, Yijun Wang, Koutian Wu, Ruoqi Gao, Muhammad Ahmed Mohsin, Jing Tang, Brihi Joshi, Heming Liu, Zheyuan Deng, Zonglin Di, Sankalp Jajee, Jiuyao Lu, Zhiwei Zhang, Saksham Kapoor, Ishan Gupta, Yunhan Zhao, Chanwoo Park, Yucheng Lu, Bing Hu, Weihang Xiao, Aravind Mohan, Hanwen Xing, Runyu Zhang, Mihir Kulshreshtha, Yuanda Xu, Qianyu Zhu, Dianzhuo Wang, Yuxin Xiao, Bowen Jiang, Yongye Su, Wenhao Chai, Zuxin Liu, Lawrence Yunliang Chen, Xuandong Zhao, Ethan Ye, Shivam Patel, Jason Xie, Alex Martin Richmond, Weixiang Ding, Emre Okcular, Diya Mathew, Ziheng Wang, Rana M. Shahroz Khan, Zhejian Peng, Fang Wu, Fan Nie, Xinyang Han, Yubin Kim, Jiawei Zhang, Zhenting Qi, Huangyuan Su, Xu Pan, Abinitha Gourabathina, Hyewon Jeong, Hemanth Neelgund Ramesh, Kumail Alhamoud, Kimia Hamidieh, Zidi Xiong, Samuel Schmidgall, Pengrui Han, Yepeng Huang, Yongheng Wang, Bowen Yang, Alex Gu, Yuchu Wang, Akshay Paruchuri, Brenna Li, Hejie Cui, Jiayuan Ding, Chaosheng Dong, Jiahao Wang, Yixuan He, Chi Wang, Pamela Bhattacharya, Tianyi Peng, Paul Pu Liang, Mitchell Gordon, Yilun Du, Marinka Zitnik, James Zou, Prasanna Tambe, Philip Torr, Emily Fox, Asu Ozdaglar, Dawn Song
cs.AI
要旨
AIシステムやデジタル製品の人間による評価は、コストが高く、時間がかかり、スケールが難しい。オフライン評価はよりスケーラブルだが、人間の多様性や対話的な行動を抽象化してしまうことが多い。そこで我々は、異種ユーザーを用いてAIシステムやデジタル製品をテストするための、人口規模のシミュレートユーザー評価インフラであるMatrAIxを紹介する。MatrAIxは3つの中核的要素から構成される。第一に、Persona 8Bは、1,290のカテゴリカル次元で表現される83億件のペルソナレコードを含む。レコードは、相関する属性を保存する依存グラフからサンプリングされるか、または人間が作成したプロフィールに由来する。我々は、599,847件の人間由来レコードと400,000件の合成レコードからなる、品質フィルタリング済みの約100万ペルソナのコアセットを公開する。第二に、MatrAIx Playgroundは、多様なユーザーがデジタル製品を評価し対話する4つの環境を提供する。すなわち、サーベイ、AIチャットボット、Web、アプリである。第三に、MatrAIxは、コマース、ソフトウェア、ファイナンス、ヘルスケアなど25以上のドメインにわたる1,010のアプリケーションタスクを提供する。我々は、8つの代表的なタスクにわたって18,189回の評価トライアルを実施した。ペルソナエージェントは、Claude Opus 4.8、GPT 5.5、Claude Haiku 4.5の3つのLLMによって駆動された。得られたフィードバックは、値上げ後のためらい、AIアシスタントの失敗後の継続意欲、レイテンシ許容度など、ペルソナの背景に応じて決定や嗜好がどのように変化するかを捉えている。我々は2つの主要な検証研究を実施した。第一に、400トライアルの統制研究では、10の行動属性と4つの環境すべてにわたってペルソナの順守を評価した。宣言された行動は、366トライアル(91.5%)で表現されるか、正しく抑制された。第二に、人間とLLMの評価者が、人間由来ペルソナの抽出品質を評価した。全体として、MatrAIxは、多様なシミュレートされた人間ユーザーを用いてAIシステムやデジタル製品を評価するためのエンドツーエンドのインフラを提供する。
English
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.