TY - GEN
T1 - Investigation into the Impact of Speaker Adversarial Perturbation on Speech Recognition
AU - Guo, Chenyang
AU - Chen, Liping
AU - Lee, Kong Aik
AU - Ling, Zhen Hua
AU - Guo, Wu
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025.
PY - 2024/12
Y1 - 2024/12
N2 - Recent development in adversarial perturbation has shown its efficacy for voice privacy protection. This paper further explores the impact of speaker adversarial perturbation on speech in downstream automatic speech recognition (ASR) tasks. To be specific, the perturbation is generated by attacking a speaker embedding extractor in an untargeted manner and added to the original speech, resulting in the adversarial version. Additionally, we examine the efficacy of incorporating the supervision from an ASR model into the perturbation generation process. Experiments were conducted on the LibriSpeech dataset, where two ASR models with different levels of robustness were examined. Firstly, the results showed a decline in the ASR performance caused by the speaker adversarial perturbation, inferring the negative influence of the speaker perturbation on speech recognition. With the supervision of the ASR model during perturbation generation, its impact on speech recognition could be mitigated. Moreover, the ASR model with a lower robustness level provided a better constraint for generating perturbations, compared to the one with a higher robustness level. Audio samples can be found in (https://voiceprivacy.github.io/Speaker-Adversarial-Perturbation-Impact-on-ASR/).
AB - Recent development in adversarial perturbation has shown its efficacy for voice privacy protection. This paper further explores the impact of speaker adversarial perturbation on speech in downstream automatic speech recognition (ASR) tasks. To be specific, the perturbation is generated by attacking a speaker embedding extractor in an untargeted manner and added to the original speech, resulting in the adversarial version. Additionally, we examine the efficacy of incorporating the supervision from an ASR model into the perturbation generation process. Experiments were conducted on the LibriSpeech dataset, where two ASR models with different levels of robustness were examined. Firstly, the results showed a decline in the ASR performance caused by the speaker adversarial perturbation, inferring the negative influence of the speaker perturbation on speech recognition. With the supervision of the ASR model during perturbation generation, its impact on speech recognition could be mitigated. Moreover, the ASR model with a lower robustness level provided a better constraint for generating perturbations, compared to the one with a higher robustness level. Audio samples can be found in (https://voiceprivacy.github.io/Speaker-Adversarial-Perturbation-Impact-on-ASR/).
KW - Speaker adversarial perturbation
KW - Speech recognition
KW - Voice privacy protection
UR - https://www.scopus.com/pages/publications/85219192268
U2 - 10.1007/978-981-96-1045-7_16
DO - 10.1007/978-981-96-1045-7_16
M3 - Conference article published in proceeding or book
AN - SCOPUS:85219192268
SN - 9789819610440
T3 - Communications in Computer and Information Science
SP - 191
EP - 199
BT - Man-Machine Speech Communication - 19th National Conference, NCMMSC 2024, Proceedings
A2 - Ling, Zhenhua
A2 - Chen, Xie
A2 - Hamdulla, Askar
A2 - He, Liang
A2 - Li, Ya
PB - Springer Science and Business Media Deutschland GmbH
T2 - 19th National Conference on Man-Machine Speech Communication, NCMMSC 2024
Y2 - 15 August 2024 through 18 August 2024
ER -