教導 Nemotron 希臘語:跨專業領域之現代希臘語語料庫挖掘、檢索調適與生成錨定
Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains
August 5, 2026
作者: Ayoub Kirouane, Christos Petrocheilos
cs.AI
摘要
現代希臘語在NVIDIA的Nemotron檢索模型及主流多語言檢索基準中均付之闕如,儘管其在法律、能源、金融及醫療應用中的檢索增強生成(RAG)領域至關重要。我們提出了一套針對現代希臘語的Nemotron檢索框架端到端改適方案,涵蓋語料庫挖掘、合成監督、檢索模型訓練、重排序器改適、閱讀器微調,以及一個名為HERA的新基準。我們的研究顯示,在專業希臘語語料庫上,無參數的BM25基線模型表現優於多個現成的多語言密集檢索模型。經由65,773個希臘語檢索配對微調後,Nemotron 1B嵌入模型將nDCG@10從0.362提升至0.835,並大幅優於未改適的對應模型。所習得的語言能力可遷移至通用領域的希臘語,但相較於BM25的優勢仍取決於領域。我們進一步改適了交叉編碼器重排序器,並在專業領域中展現一致的效能提升。最後,我們以LoRA微調了一個Nemotron 30B-A3B混合專家閱讀器以進行接地生成,將評判答案正確率從29.4%提升至66.9%,同時顯著改善了忠實度與引用品質。我們亦推出HERA——首個大規模希臘語檢索增強生成基準,並公開我們的改適模型及基準,以支援未來希臘語RAG系統的研究。
English
Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial, and medical applications. We present an end-to-end adaptation of the Nemotron retrieval stack for Modern Greek, including corpus mining, synthetic supervision, retrieval model training, reranker adaptation, reader fine-tuning, and a new benchmark called HERA. Our study shows that a parameter-free BM25 baseline outperforms several off-the-shelf multilingual dense retrieval models on specialist Greek corpora. After fine-tuning on 65,773 Greek retrieval pairs, a Nemotron 1B embedder improves nDCG@10 from 0.362 to 0.835 and substantially outperforms its unadapted counterpart. The learned language competence transfers to general-domain Greek, although the advantage over BM25 remains domain-dependent. We further adapt a cross-encoder reranker and demonstrate consistent improvements across specialist domains. Finally, we LoRA-tune a Nemotron 30B-A3B mixture-of-experts reader for grounded generation, increasing judged answer correctness from 29.4% to 66.9% while significantly improving faithfulness and citation quality. We also introduce HERA, the first large-scale Greek benchmark for retrieval-augmented generation, and release our adapted models and benchmark to support future research on Greek-language RAG systems.