危険な能力に対するフロンティアモデルの評価

要旨

新たなAIシステムがもたらすリスクを理解するためには、そのシステムが何をでき、何ができないかを理解する必要がある。先行研究を基盤として、我々は新たな「危険な能力」評価プログラムを導入し、Gemini 1.0モデルにおいてそのパイロット評価を実施した。我々の評価は以下の4つの領域をカバーしている：(1) 説得と欺瞞、(2) サイバーセキュリティ、(3) 自己増殖、(4) 自己推論。評価したモデルにおいて強い危険な能力の証拠は見つからなかったが、早期警告の兆候を指摘した。我々の目標は、将来のモデルに備えて、危険な能力評価の厳密な科学を進展させることに貢献することである。

English

To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evaluations and pilot them on Gemini 1.0 models. Our evaluations cover four areas: (1) persuasion and deception; (2) cyber-security; (3) self-proliferation; and (4) self-reasoning. We do not find evidence of strong dangerous capabilities in the models we evaluated, but we flag early warning signs. Our goal is to help advance a rigorous science of dangerous capability evaluation, in preparation for future models.

危険な能力に対するフロンティアモデルの評価

Evaluating Frontier Models for Dangerous Capabilities

要旨

Support