TY - GEN
T1 - Dialogue Language Model with Large-Scale Persona Data Engineering
AU - Hong, Mengze
AU - Zhang, Chen Jason
AU - Chen, Chaotao
AU - Lian, Rongzhong
AU - Jiang, Di
N1 - Publisher Copyright:
© 2025 Association for Computational Linguistics.
PY - 2025/4
Y1 - 2025/4
N2 - Maintaining persona consistency is paramount in the application of open-domain dialogue systems, as exemplified by models like ChatGPT. Despite significant advancements, the limited scale and diversity of current persona dialogue datasets remain challenges to achieving robust persona-consistent dialogue models. In this study, drawing inspiration from the success of large-scale pre-training, we introduce PPDS, an open-domain persona dialogue system that employs extensive generative pre-training on a persona dialogue dataset to enhance persona consistency. Specifically, we present a persona extraction model designed to autonomously and precisely generate vast persona dialogue datasets. Additionally, we unveil a pioneering persona augmentation technique to address the invalid persona bias inherent in the constructed dataset. Both quantitative and human evaluations consistently highlight the superior response quality and persona consistency of our proposed model, underscoring its effectiveness.
AB - Maintaining persona consistency is paramount in the application of open-domain dialogue systems, as exemplified by models like ChatGPT. Despite significant advancements, the limited scale and diversity of current persona dialogue datasets remain challenges to achieving robust persona-consistent dialogue models. In this study, drawing inspiration from the success of large-scale pre-training, we introduce PPDS, an open-domain persona dialogue system that employs extensive generative pre-training on a persona dialogue dataset to enhance persona consistency. Specifically, we present a persona extraction model designed to autonomously and precisely generate vast persona dialogue datasets. Additionally, we unveil a pioneering persona augmentation technique to address the invalid persona bias inherent in the constructed dataset. Both quantitative and human evaluations consistently highlight the superior response quality and persona consistency of our proposed model, underscoring its effectiveness.
UR - https://www.scopus.com/pages/publications/105027129688
U2 - 10.18653/v1/2025.naacl-industry.71
DO - 10.18653/v1/2025.naacl-industry.71
M3 - Conference article published in proceeding or book
AN - SCOPUS:105027129688
T3 - Proceedings of the 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies: Long Papers, NAACL-HLT 2025
SP - 961
EP - 970
BT - Industry Track
A2 - Chen, Weizhu
A2 - Yang, Yi
A2 - Kachuee, Mohammad
A2 - Fu, Xue-Yong
PB - Association for Computational Linguistics (ACL)
T2 - 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2025
Y2 - 29 April 2025 through 4 May 2025
ER -