Wnuan:企業専有知識に対する質問応答のための段階的ポストトレーニング
Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge
August 3, 2026
著者: Xiaofeng Shi, Xiaosong Qiu, Wenxin Ma, Qian Kou, Yiming Pan, Longbin Yu, Ying Liu, Haiping Wang, Hua Zhou
cs.AI
要旨
エンタープライズ質問応答では、モデルが汎用能力を失うことなく専有知識を獲得することが求められる。本稿では、文書からタスク指向の教師信号を構築し、汎用データのリプレイを伴う教師ありファインチューニングを実施し、残差エラーに対して強化学習を適用する3段階のパイプラインであるWnuanを提示する。707問のWnuanBenchでは、主要な32Bルートにおいて、許容回答率(AAR)が適応前の52.76%から、SFT後には80.06%、RL後には91.51%へと向上する。一致させた100更新プロトコルの下では、残差エラーサンプリングは、全プールサンプリングとサイズを一致させたランダムサンプリングをそれぞれ3.11ポイントおよび2.97ポイント上回る。ソースクラスタのブートストラップ区間は、両対比においてゼロより上にあり、同一ドメインの検証セットでもその順序が維持される。汎用ベンチマークの平均は、このルート全体で5.17ポイント低下し、その低下は指示追従に集中している。自動評価アンサンブルは、層化されたWnuan-Inst応答サンプルの90.5%において、権威あるドメイン専門家と一致する。これらの結果は、段階的エンタープライズ適応による利点と汎用能力コストの両方を特徴づけるものである。
English
Enterprise question answering requires models to acquire proprietary knowledge without discarding general capabilities. We present Wnuan, a three-stage pipeline that constructs task-oriented supervision from documents, performs supervised fine-tuning with general-data replay, and applies reinforcement learning to residual errors. On the 707-question WnuanBench, the primary 32B route raises acceptable-answer rate (AAR) from 52.76% before adaptation to 80.06% after SFT and 91.51% after RL. Under a matched 100-update protocol, residual-error sampling outperforms full-pool and size-matched random sampling by 3.11 and 2.97 points, respectively. Source-cluster bootstrap intervals remain above zero for both contrasts, and a same-domain validation set preserves the ordering. The general-benchmark average decreases by 5.17 points across the route, concentrated in instruction following. The automatic evaluation ensemble agrees with an authoritative domain expert on 90.5% of a stratified Wnuan-Inst response sample. These results characterize both the gains and the general-capability cost of staged enterprise adaptation.