TY - GEN
T1 - Agentic Audio Moderator vs Human Moderator in Think-Aloud Usability Testing Results from a Randomized Controlled Trial
AU - Zhu, Wangda
AU - Chen, Guang
AU - Wang, Yao
AU - An, Pengcheng
AU - Du, Jiachun
AU - Li, Chen
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s).
PY - 2026/4/13
Y1 - 2026/4/13
N2 - Agentic AI holds promise for usability testing, yet its role as an audio moderator in think-aloud protocols is not well understood. This study explores: (1) how to design and develop an agentic audio moderator for think-aloud usability testing, and (2) how participants moderated by an agentic moderator differ from those moderated by a human regarding task performance, verbalization behaviors, user experience, and social perceptions of the moderator. Using a design-based research approach, we interviewed nine UX experts, iteratively developed an AI moderator, and evaluated it in a randomized controlled trial (N = 60) with a note-taking application. Results suggest that significant differences were not observed between AI and human moderators in task performance or verbalization behaviors, though AI moderators received lower social perception ratings. This work contributes the first design-oriented evaluation of AI moderators in usability testing, offering implications for developing more acceptable and effective agentic audio moderators.
AB - Agentic AI holds promise for usability testing, yet its role as an audio moderator in think-aloud protocols is not well understood. This study explores: (1) how to design and develop an agentic audio moderator for think-aloud usability testing, and (2) how participants moderated by an agentic moderator differ from those moderated by a human regarding task performance, verbalization behaviors, user experience, and social perceptions of the moderator. Using a design-based research approach, we interviewed nine UX experts, iteratively developed an AI moderator, and evaluated it in a randomized controlled trial (N = 60) with a note-taking application. Results suggest that significant differences were not observed between AI and human moderators in task performance or verbalization behaviors, though AI moderators received lower social perception ratings. This work contributes the first design-oriented evaluation of AI moderators in usability testing, offering implications for developing more acceptable and effective agentic audio moderators.
KW - Agentic audio AI
KW - Design-based research
KW - Randomized controlled trial
KW - Think-aloud usability testing
KW - User study moderator
UR - https://www.scopus.com/pages/publications/105038757349
U2 - 10.1145/3772318.3791653
DO - 10.1145/3772318.3791653
M3 - Conference article published in proceeding or book
AN - SCOPUS:105038757349
T3 - Conference on Human Factors in Computing Systems - Proceedings
SP - 1
EP - 19
BT - CHI 2026 - Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems
A2 - Oliver, Nuria
A2 - Shamma, David A.
A2 - Candello, Heloisa
A2 - Cesar, Pablo
A2 - Lopes, Pedro
A2 - Bozzon, Alessandro
A2 - Kosch, Thomas
A2 - Liao, Vera
A2 - Ma, Xiaojuan
A2 - Artizzu, Valentino
A2 - Draxler, Fiona
A2 - Lopez, Gustavo
A2 - Reinschluessel, Anke V.
A2 - Tong, Xin
A2 - Toups Dugas, Phoebe O.
PB - Association for Computing Machinery
T2 - 2026 CHI Conference on Human Factors in Computing Systems, CHI 2026
Y2 - 13 April 2026 through 17 April 2026
ER -