네모트론에 그리스어 가르치기: 전문 분야 전반의 현대 그리스어를 위한 코퍼스 마이닝, 검색 적응, 및 생성 접지
Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains
August 5, 2026
저자: Ayoub Kirouane, Christos Petrocheilos
cs.AI
초록
현대 그리스어는 법률, 에너지, 금융, 의료 응용 분야에서 검색 증강 생성(RAG)에 중요함에도 불구하고 NVIDIA의 Nemotron 검색 모델과 주요 다국어 검색 벤치마크에서 누락되어 있다. 본 연구는 코퍼스 마이닝, 합성 지도 학습, 검색 모델 훈련, 리랭커 적응, 리더 미세 조정, 그리고 HERA라는 새로운 벤치마크를 포함하여 현대 그리스어를 위한 Nemotron 검색 스택의 종단 간 적응을 제시한다. 본 연구는 파라미터가 없는 BM25 기준선이 전문 그리스어 코퍼스에서 여러 오프더셸 다국어 밀집 검색 모델을 능가함을 보여준다. 65,773개의 그리스어 검색 쌍으로 미세 조정한 후, Nemotron 1B 임베더는 nDCG@10을 0.362에서 0.835로 개선하며 적응되지 않은 대응 모델을 상당히 능가한다. 학습된 언어 능력은 일반 영역 그리스어로 전이되지만, BM25에 대한 우위는 도메인에 따라 달라진다. 또한 교차 인코더 리랭커를 적응시켜 전문 도메인 전반에 걸쳐 일관된 개선을 입증한다. 마지막으로, 근거 기반 생성을 위해 Nemotron 30B-A3B 혼합 전문가 리더를 LoRA 튜닝하여 판정된 답변 정확도를 29.4%에서 66.9%로 향상시키는 동시에 충실도와 인용 품질을 유의미하게 개선한다. 또한 그리스어 RAG를 위한 최초의 대규모 벤치마크인 HERA를 도입하고, 그리스어 RAG 시스템에 대한 향후 연구를 지원하기 위해 적응된 모델과 벤치마크를 공개한다.
English
Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial, and medical applications. We present an end-to-end adaptation of the Nemotron retrieval stack for Modern Greek, including corpus mining, synthetic supervision, retrieval model training, reranker adaptation, reader fine-tuning, and a new benchmark called HERA. Our study shows that a parameter-free BM25 baseline outperforms several off-the-shelf multilingual dense retrieval models on specialist Greek corpora. After fine-tuning on 65,773 Greek retrieval pairs, a Nemotron 1B embedder improves nDCG@10 from 0.362 to 0.835 and substantially outperforms its unadapted counterpart. The learned language competence transfers to general-domain Greek, although the advantage over BM25 remains domain-dependent. We further adapt a cross-encoder reranker and demonstrate consistent improvements across specialist domains. Finally, we LoRA-tune a Nemotron 30B-A3B mixture-of-experts reader for grounded generation, increasing judged answer correctness from 29.4% to 66.9% while significantly improving faithfulness and citation quality. We also introduce HERA, the first large-scale Greek benchmark for retrieval-augmented generation, and release our adapted models and benchmark to support future research on Greek-language RAG systems.