AISPA:大型語言模型應用中以使用者為中心的系統提示詞審計
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
July 30, 2026
作者: Xiangning Lin, Shenzhe Zhu, Shu Yang, Zhenyu Zhang, Haoqian Zhang, Yipeng Zhao, Chengxuan Qian, Tianwei Wang, Ziheng Zhang, Zhenlong Yuan, Dingcheng Wang, Juncheng Wu, Yuan Si, Jiaxin Liu, Baolong Bi, Robert Mahari, Tobin South, Dazza Greenwood, Zexue He, Rishi Bommasani, Sophia Kazinnik, Andreas Haupt, Samuele Marro, Erik Brynjolfsson, Alex Pentland, Jiaxin Pei
cs.AI
摘要
系統提示詞(system prompts)是開發者設定的指令,用以規範AI應用程式中基礎模型的行為。雖然系統提示詞廣泛應用於商業AI產品,但鮮少對公眾或監管機構公開,這在AI系統大規模部署中造成了嚴重的信任與問責缺口。本文提出人工智慧系統提示詞保證框架(Artificial Intelligence System Prompt Assurance, AISPA),這是一個以使用者為中心的框架,用於系統性地稽核AI系統中的系統提示詞。AISPA檢視系統提示詞的特定組成部分,並沿著八個對使用者至關重要的面向進行評估。我們接著運用此框架審查了88個商業AI產品中系統提示詞內的3,249條指令,將每條指令分類為保護性(保護使用者)或問題性(有礙使用者權益)。我們的稽核揭示了四項核心發現。首先,系統提示詞的設計在不同產品和開發者之間存在顯著差異,部分組織每個產品平均包含超過60條保護性指令,而其他組織平均不到5條。其次,保護性指令雖被廣泛採用但涵蓋範圍淺薄:98.9%的產品至少包含一條保護性指令,但僅有24%的產品涵蓋AISPA分類架構的全部八個面向。第三,系統提示詞已持續變得更長且更具使用者保護性,顯示使用者保護正成為商業提示詞設計中日益受關注的議題。第四,儘管有此進展,問題性指令仍然普遍存在:約40%的產品至少包含一條有違使用者利益的指令,且保護性與問題性指令經常共存於同一提示詞中。我們的研究發現凸顯了商業AI產品中系統提示詞需要更高的透明度、標準化以及獨立監督機制。
English
System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are rarely disclosed to the public or regulators, creating a serious trust and accountability gap in the wide deployment of AI systems. In this paper, we introduce Artificial Intelligence System Prompt Assurance (AISPA), a user-centric framework for systematically auditing system prompts in AI systems. AISPA examines specific parts of a system prompt and evaluates them along eight dimensions that matter to users. We then use this framework to review 3,249 instructions from system prompts in 88 commercial AI products, classifying each instruction as either protective (of users) or problematic. Our audit surfaces four core findings. First, system prompt design varies substantially across products and developers, with some organizations averaging over 60 protective instructions per product while others average fewer than 5. Second, protective instructions are widely adopted but shallow in scope: 98.9% of products contain at least one, yet only 24% cover all eight dimensions of the AISPA taxonomy. Third, system prompts have grown steadily longer and more protective of users, suggesting that user protection is becoming a more visible concern in commercial prompt design. Fourth, despite this progress, problematic instructions remain pervasive: roughly 40% of products contain at least one instruction that works against user interests, and protective and problematic instructions frequently coexist within the same prompt. Our findings highlight the need for greater transparency, standardization, and independent oversight for system prompts in commercial AI products.