ChatPaper.aiChatPaper

Wnuan:面向专有企业知识问答的分阶段后训练

Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge

August 3, 2026
作者: Xiaofeng Shi, Xiaosong Qiu, Wenxin Ma, Qian Kou, Yiming Pan, Longbin Yu, Ying Liu, Haiping Wang, Hua Zhou
cs.AI

摘要

企业问答要求模型在不丧失通用能力的前提下获取专有知识。我们提出Wnuan——一种三阶段流水线:从文档中构建任务导向的监督信号,利用通用数据回放进行监督微调,并对残差错误应用强化学习。在包含707个问题的WnuanBench上,主要的32B路径将可接受答案率(AAR)从适配前的52.76%提升至SFT后的80.06%及RL后的91.51%。在匹配的100次更新协议下,残差错误采样分别比全池采样和规模匹配的随机采样高出3.11和2.97个百分点。两种对比的源聚类自助法置信区间均保持在零以上,同域验证集亦保持了该排序。通用基准平均分在整个路径中下降5.17分,主要集中在指令遵循维度。自动评估集成与权威领域专家在分层Wnuan-Inst响应样本上的一致率为90.5%。这些结果同时刻画了分阶段企业适配带来的收益与通用能力代价。
English
Enterprise question answering requires models to acquire proprietary knowledge without discarding general capabilities. We present Wnuan, a three-stage pipeline that constructs task-oriented supervision from documents, performs supervised fine-tuning with general-data replay, and applies reinforcement learning to residual errors. On the 707-question WnuanBench, the primary 32B route raises acceptable-answer rate (AAR) from 52.76% before adaptation to 80.06% after SFT and 91.51% after RL. Under a matched 100-update protocol, residual-error sampling outperforms full-pool and size-matched random sampling by 3.11 and 2.97 points, respectively. Source-cluster bootstrap intervals remain above zero for both contrasts, and a same-domain validation set preserves the ordering. The general-benchmark average decreases by 5.17 points across the route, concentrated in instruction following. The automatic evaluation ensemble agrees with an authoritative domain expert on 90.5% of a stratified Wnuan-Inst response sample. These results characterize both the gains and the general-capability cost of staged enterprise adaptation.