TY - GEN
T1 - Recognizing Textual Entailment by Hierarchical Crowdsourcing with Diverse Labor Costs
AU - Zhang, Haodi
AU - Yang, Junyu
AU - Huang, Wenxi
AU - Cai, Min
AU - Li, Jiahong
AU - Zhang, Chen
AU - Wu, Kaishun
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024/7
Y1 - 2024/7
N2 - With the rapid advancement of supervised learning and the rise of large language models, the demand for high-quality labeled datasets has surged. However, crowdsourced datasets often suffer from label noise. To address this challenge, we present a hierarchical crowdsourcing framework with diverse labor costs that aims to mitigate label noise. Our framework models crowdsourcing workers with varying labor costs and leverages hierarchical crowdsourcing under limited budget constraints to enhance data quality. We establish a loop for label selection and checking, strategically selecting checkers from different cost levels, including perfect workers (Oracles) and regular experts, to optimize the checking process. Additionally, we tackle an NP-hard problem in label selection. Experimental evaluation on a real-world dataset for the Recognizing Textual Entailment (RTE) task demonstrates a significant improvement in labeled dataset quality, leading to state-of-the-art performance in downstream tasks.
AB - With the rapid advancement of supervised learning and the rise of large language models, the demand for high-quality labeled datasets has surged. However, crowdsourced datasets often suffer from label noise. To address this challenge, we present a hierarchical crowdsourcing framework with diverse labor costs that aims to mitigate label noise. Our framework models crowdsourcing workers with varying labor costs and leverages hierarchical crowdsourcing under limited budget constraints to enhance data quality. We establish a loop for label selection and checking, strategically selecting checkers from different cost levels, including perfect workers (Oracles) and regular experts, to optimize the checking process. Additionally, we tackle an NP-hard problem in label selection. Experimental evaluation on a real-world dataset for the Recognizing Textual Entailment (RTE) task demonstrates a significant improvement in labeled dataset quality, leading to state-of-the-art performance in downstream tasks.
KW - diverse labor costs
KW - hierarchical crowd-sourcing
KW - oracle
UR - https://www.scopus.com/pages/publications/85199027775
U2 - 10.1109/CSCWD61410.2024.10580016
DO - 10.1109/CSCWD61410.2024.10580016
M3 - Conference article published in proceeding or book
AN - SCOPUS:85199027775
T3 - Proceedings of the 2024 27th International Conference on Computer Supported Cooperative Work in Design, CSCWD 2024
SP - 453
EP - 458
BT - Proceedings of the 2024 27th International Conference on Computer Supported Cooperative Work in Design, CSCWD 2024
A2 - Shen, Weiming
A2 - Shen, Weiming
A2 - Barthes, Jean-Paul
A2 - Luo, Junzhou
A2 - Qiu, Tie
A2 - Zhou, Xiaobo
A2 - Zhang, Jinghui
A2 - Zhu, Haibin
A2 - Peng, Kunkun
A2 - Xu, Tianyi
A2 - Chen, Ning
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 27th International Conference on Computer Supported Cooperative Work in Design, CSCWD 2024
Y2 - 8 May 2024 through 10 May 2024
ER -