MatrAIx:用83亿人格智能体模拟世界
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
August 4, 2026
作者: Xiaomin Li, Yuexing Hao, Jianheng Hou, Jintao Huang, Qianfeng Wen, Shirley Huang, Yifan Liu, Xiaoyi Liu, Yilan Fan, Yijun Wang, Koutian Wu, Ruoqi Gao, Muhammad Ahmed Mohsin, Jing Tang, Brihi Joshi, Heming Liu, Zheyuan Deng, Zonglin Di, Sankalp Jajee, Jiuyao Lu, Zhiwei Zhang, Saksham Kapoor, Ishan Gupta, Yunhan Zhao, Chanwoo Park, Yucheng Lu, Bing Hu, Weihang Xiao, Aravind Mohan, Hanwen Xing, Runyu Zhang, Mihir Kulshreshtha, Yuanda Xu, Qianyu Zhu, Dianzhuo Wang, Yuxin Xiao, Bowen Jiang, Yongye Su, Wenhao Chai, Zuxin Liu, Lawrence Yunliang Chen, Xuandong Zhao, Ethan Ye, Shivam Patel, Jason Xie, Alex Martin Richmond, Weixiang Ding, Emre Okcular, Diya Mathew, Ziheng Wang, Rana M. Shahroz Khan, Zhejian Peng, Fang Wu, Fan Nie, Xinyang Han, Yubin Kim, Jiawei Zhang, Zhenting Qi, Huangyuan Su, Xu Pan, Abinitha Gourabathina, Hyewon Jeong, Hemanth Neelgund Ramesh, Kumail Alhamoud, Kimia Hamidieh, Zidi Xiong, Samuel Schmidgall, Pengrui Han, Yepeng Huang, Yongheng Wang, Bowen Yang, Alex Gu, Yuchu Wang, Akshay Paruchuri, Brenna Li, Hejie Cui, Jiayuan Ding, Chaosheng Dong, Jiahao Wang, Yixuan He, Chi Wang, Pamela Bhattacharya, Tianyi Peng, Paul Pu Liang, Mitchell Gordon, Yilun Du, Marinka Zitnik, James Zou, Prasanna Tambe, Philip Torr, Emily Fox, Asu Ozdaglar, Dawn Song
cs.AI
摘要
对AI系统和数字产品进行人工评估成本高昂、速度缓慢且难以规模化。离线评估更具可扩展性,但往往忽略了人类多样性和交互行为。为此,我们引入MatrAIx——一个用于测试AI系统和数字产品、面向异构用户的大规模模拟用户评估基础设施。MatrAIx包含三个核心组件:第一,Persona 8B包含83亿条用户画像记录,由1,290个分类维度表示。这些记录要么从保留属性相关性的依赖图中采样,要么来源于人工撰写的画像。我们发布了一个经过质量过滤、约100万条用户画像的核心子集,其中包括599,847条基于人类数据的记录和400,000条合成记录。第二,MatrAIx Playground提供四种环境,让多样化用户评估数字产品并与之交互:问卷调查、AI聊天机器人、网页和应用。第三,MatrAIx提供1,010个应用任务,覆盖超过25个领域,包括商务、软件、金融和医疗健康。我们在八个代表性任务上进行了18,189次评估试验。用户画像智能体由三个大语言模型驱动:Claude Opus 4.5、GPT-5.5和Claude Haiku 4.5。所产生的反馈刻画了决策和偏好如何随用户画像背景而变化,包括价格上涨后的犹豫、AI助手失败后是否愿意继续,以及延迟容忍度。我们开展了两项主要验证研究:第一,一项包含400次试验的对照研究评估了跨十个行为属性和全部四种环境的用户画像遵从性。在366次试验(91.5%)中,声明的行为得到表达或被正确抑制。第二,人类评审员与LLM评审员评估了基于人类数据的用户画像的提取质量。总体而言,MatrAIx为使用多样化模拟人类用户评估AI系统和数字产品提供了一种端到端基础设施。
English
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.