AISPA: 대규모 언어 모델 애플리케이션을 위한 사용자 중심 시스템 프롬프트 감사
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
July 30, 2026
저자: Xiangning Lin, Shenzhe Zhu, Shu Yang, Zhenyu Zhang, Haoqian Zhang, Yipeng Zhao, Chengxuan Qian, Tianwei Wang, Ziheng Zhang, Zhenlong Yuan, Dingcheng Wang, Juncheng Wu, Yuan Si, Jiaxin Liu, Baolong Bi, Robert Mahari, Tobin South, Dazza Greenwood, Zexue He, Rishi Bommasani, Sophia Kazinnik, Andreas Haupt, Samuele Marro, Erik Brynjolfsson, Alex Pentland, Jiaxin Pei
cs.AI
초록
시스템 프롬프트(system prompt)는 AI 애플리케이션에서 기반 모델(foundation model)의 동작을 제어하기 위해 개발자가 구성하는 지침이다. 이들은 상용 AI 제품 전반에 걸쳐 사용되고 있으나 대중이나 규제 기관에 거의 공개되지 않아, AI 시스템의 광범위한 배포에 있어 심각한 신뢰 및 책임성 공백을 초래하고 있다. 본 논문에서는 AI 시스템의 시스템 프롬프트를 체계적으로 감사하기 위한 사용자 중심 프레임워크인 AISPA(Artificial Intelligence System Prompt Assurance, 인공지능 시스템 프롬프트 보증)를 제안한다. AISPA는 시스템 프롬프트의 특정 부분을 검토하고 사용자에게 중요한 여덟 가지 차원을 기준으로 평가한다. 이 프레임워크를 활용하여 88개 상용 AI 제품의 시스템 프롬프트에 포함된 3,249개 지침을 검토하고, 각 지침을 사용자 보호적(protective) 또는 문제적(problematic)으로 분류하였다. 본 감사는 네 가지 핵심 결과를 도출한다. 첫째, 시스템 프롬프트 설계는 제품 및 개발자에 따라 상당히 상이하여, 일부 조직은 제품당 평균 60개 이상의 보호적 지침을 포함하는 반면 다른 조직은 평균 5개 미만을 포함한다. 둘째, 보호적 지침은 널리 채택되었으나 범위가 얕다. 제품의 98.9%가 최소 하나 이상의 보호적 지침을 포함하지만, AISPA 분류 체계의 여덟 가지 차원을 모두 충족하는 제품은 24%에 불과하다. 셋째, 시스템 프롬프트는 지속적으로 길어지고 사용자 보호 성격이 강화되고 있어, 사용자 보호가 상용 프롬프트 설계에서 점점 더 중요한 고려 사항이 되고 있음을 시사한다. 넷째, 이러한 진전에도 불구하고 문제적 지침은 여전히 만연하다. 약 40%의 제품이 사용자 이익에 반하는 지침을 최소 하나 이상 포함하고 있으며, 보호적 지침과 문제적 지침이 동일한 프롬프트 내에 빈번히 공존한다. 본 연구 결과는 상용 AI 제품의 시스템 프롬프트에 대한 더 큰 투명성, 표준화, 그리고 독립적 감독의 필요성을 강조한다.
English
System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are rarely disclosed to the public or regulators, creating a serious trust and accountability gap in the wide deployment of AI systems. In this paper, we introduce Artificial Intelligence System Prompt Assurance (AISPA), a user-centric framework for systematically auditing system prompts in AI systems. AISPA examines specific parts of a system prompt and evaluates them along eight dimensions that matter to users. We then use this framework to review 3,249 instructions from system prompts in 88 commercial AI products, classifying each instruction as either protective (of users) or problematic. Our audit surfaces four core findings. First, system prompt design varies substantially across products and developers, with some organizations averaging over 60 protective instructions per product while others average fewer than 5. Second, protective instructions are widely adopted but shallow in scope: 98.9% of products contain at least one, yet only 24% cover all eight dimensions of the AISPA taxonomy. Third, system prompts have grown steadily longer and more protective of users, suggesting that user protection is becoming a more visible concern in commercial prompt design. Fourth, despite this progress, problematic instructions remain pervasive: roughly 40% of products contain at least one instruction that works against user interests, and protective and problematic instructions frequently coexist within the same prompt. Our findings highlight the need for greater transparency, standardization, and independent oversight for system prompts in commercial AI products.