LegalPincite:多层级法律信息检索数据集
LegalPincite: Multi-level Legal Information Retrieval Dataset
August 4, 2026
作者: Theresia Veronika Rampisela, Henrik Palmer Olsen, Giovanni Colavizza
cs.AI
摘要
法律信息检索(IR)中的一个常见任务是从判例法文集中查找相关的法律来源。虽然法律实践通常需要对特定案件段落的精确定位引用(pincites),但现有的大多数公开法律IR数据集缺乏段落级引注标注。然而,包含此类信息的公开数据集在查询文本中存在数据泄漏,并且将既不引用也不被引用的段落排除在语料库之外,这造成了不现实且过度简化的检索环境,可能导致性能虚高。为解决这些局限,我们构建了一个基于欧盟法院(CJEU)判决的大规模法律IR数据集。该数据集包含:(i)掩码的案件/段落查询,其中移除了引注信息;(ii)包含所有段落的语料库;以及(iii)案例级和段落级的基准真值引注,并经过部分人类专家验证。我们的数据集支持法律IR方法在多个查询-文档层级(案例到案例、段落到案例以及段落到段落的检索)上的开发和严格评估。数据集链接:https://huggingface.co/datasets/theresiavr/legalpincite
English
A common task in legal Information Retrieval (IR) is to find relevant legal sources from case-law collections. While legal practice often requires pinpoint citations (pincites) to specific case paragraphs, most existing public legal IR datasets lack paragraph-level citation annotations. Yet, publicly available datasets with such information contain data leakage in the query text and exclude paragraphs that are neither citing nor cited from the corpora, creating an unrealistic and oversimplified retrieval setting, potentially leading to inflated performance. To address these limitations, we contribute a large-scale legal IR dataset constructed from Court of Justice of the European Union (CJEU) judgments. The dataset contains: (i) masked case/paragraph queries, with removed citation information; (ii) a corpus that includes all paragraphs; and (iii) case- and paragraph-level ground-truth citations, with partial human expert validation. Our dataset supports both the development and rigorous evaluation of legal IR methods, at multiple query-document levels (case-to-case, paragraph-to-case, and paragraph-to-paragraph retrieval). Link to dataset: https://huggingface.co/datasets/theresiavr/legalpincite