Wnuan:針對專有企業知識問答的分階段後訓練
Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge
August 3, 2026
作者: Xiaofeng Shi, Xiaosong Qiu, Wenxin Ma, Qian Kou, Yiming Pan, Longbin Yu, Ying Liu, Haiping Wang, Hua Zhou
cs.AI
摘要
企業問答要求模型獲取專有知識,同時不喪失通用能力。我們提出 Wnuan,這是一個三階段流程:從文件中建構任務導向監督、以通用資料重播進行監督式微調,並對殘餘錯誤應用強化學習。在包含 707 道問題的 WnuanBench 上,主要 32B 路線將可接受答案率(AAR)從適配前的 52.76% 提升至 SFT 後的 80.06%,並於 RL 後達到 91.51%。在匹配的 100 次更新協議下,殘餘錯誤採樣分別優於全池採樣與規模匹配隨機採樣 3.11 與 2.97 個百分點。兩個對比的來源叢集自助法信賴區間均高於零,且同領域驗證集保持了相同的排序。沿著該路線,通用基準平均分數下降 5.17 分,主要集中在指令遵循方面。自動評估集成與權威領域專家在分層 Wnuan-Inst 回應樣本上有 90.5% 的一致性。這些結果同時刻畫了分階段企業適配的收益與通用能力成本。
English
Enterprise question answering requires models to acquire proprietary knowledge without discarding general capabilities. We present Wnuan, a three-stage pipeline that constructs task-oriented supervision from documents, performs supervised fine-tuning with general-data replay, and applies reinforcement learning to residual errors. On the 707-question WnuanBench, the primary 32B route raises acceptable-answer rate (AAR) from 52.76% before adaptation to 80.06% after SFT and 91.51% after RL. Under a matched 100-update protocol, residual-error sampling outperforms full-pool and size-matched random sampling by 3.11 and 2.97 points, respectively. Source-cluster bootstrap intervals remain above zero for both contrasts, and a same-domain validation set preserves the ordering. The general-benchmark average decreases by 5.17 points across the route, concentrated in instruction following. The automatic evaluation ensemble agrees with an authoritative domain expert on 90.5% of a stratified Wnuan-Inst response sample. These results characterize both the gains and the general-capability cost of staged enterprise adaptation.