ChatPaper.aiChatPaper

LegalPincite: 多層的な法律情報検索データセット

LegalPincite: Multi-level Legal Information Retrieval Dataset

August 4, 2026
著者: Theresia Veronika Rampisela, Henrik Palmer Olsen, Giovanni Colavizza
cs.AI

要旨

法的情報検索(IR)における一般的なタスクは、判例集から関連する法的資料を見つけることである。法的実務では特定の判決段落へのピンポイント引用(ピンサイト)がしばしば求められる一方で、既存の公開法的IRデータセットの大半は段落レベルの引用注釈を備えていない。しかし、そのような情報を含む公開データセットは、クエリ文にデータ漏洩が含まれており、またコーパスから引用する側でも引用される側でもない段落を除外しているため、非現実的で過度に単純化された検索設定となっており、性能を過大評価する可能性がある。これらの限界に対処するため、我々は欧州連合司法裁判所(CJEU)の判決から構築した大規模な法的IRデータセットを提供する。このデータセットには、(i) 引用情報を除去したマスク済みの事例/段落クエリ、(ii) すべての段落を含むコーパス、(iii) 部分的な人間専門家による検証を経た事例レベルおよび段落レベルの正解引用が含まれる。我々のデータセットは、複数のクエリ-文書レベル(事例間、段落から事例、段落間の検索)において、法的IR手法の開発と厳密な評価の両方を支援する。データセットへのリンク: https://huggingface.co/datasets/theresiavr/legalpincite
English
A common task in legal Information Retrieval (IR) is to find relevant legal sources from case-law collections. While legal practice often requires pinpoint citations (pincites) to specific case paragraphs, most existing public legal IR datasets lack paragraph-level citation annotations. Yet, publicly available datasets with such information contain data leakage in the query text and exclude paragraphs that are neither citing nor cited from the corpora, creating an unrealistic and oversimplified retrieval setting, potentially leading to inflated performance. To address these limitations, we contribute a large-scale legal IR dataset constructed from Court of Justice of the European Union (CJEU) judgments. The dataset contains: (i) masked case/paragraph queries, with removed citation information; (ii) a corpus that includes all paragraphs; and (iii) case- and paragraph-level ground-truth citations, with partial human expert validation. Our dataset supports both the development and rigorous evaluation of legal IR methods, at multiple query-document levels (case-to-case, paragraph-to-case, and paragraph-to-paragraph retrieval). Link to dataset: https://huggingface.co/datasets/theresiavr/legalpincite