TY - GEN
T1 - Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding
AU - Wang, Rui
AU - Chen, Liping
AU - Lee, Kong Aik
AU - Ling, Zhen Hua
N1 - Publisher Copyright:
© 2024 International Speech Communication Association. All rights reserved.
PY - 2024/9
Y1 - 2024/9
N2 - Voice anonymization has been developed as a technique for preserving privacy by replacing the speaker's voice in a speech signal with that of a pseudo-speaker, thereby obscuring the original voice attributes from machine recognition and human perception. In this paper, we focus on altering the voice attributes against machine recognition while retaining human perception. We referred to this as the asynchronous voice anonymization. To this end, a speech generation framework incorporating a speaker disentanglement mechanism is employed to generate the anonymized speech. The speaker attributes are altered through adversarial perturbation applied on the speaker embedding, while human perception is preserved by controlling the intensity of perturbation. Experiments conducted on the LibriSpeech dataset showed that the speaker attributes were obscured with their human perception preserved for 60.71% of the processed utterances. Audio samples can be found in.
AB - Voice anonymization has been developed as a technique for preserving privacy by replacing the speaker's voice in a speech signal with that of a pseudo-speaker, thereby obscuring the original voice attributes from machine recognition and human perception. In this paper, we focus on altering the voice attributes against machine recognition while retaining human perception. We referred to this as the asynchronous voice anonymization. To this end, a speech generation framework incorporating a speaker disentanglement mechanism is employed to generate the anonymized speech. The speaker attributes are altered through adversarial perturbation applied on the speaker embedding, while human perception is preserved by controlling the intensity of perturbation. Experiments conducted on the LibriSpeech dataset showed that the speaker attributes were obscured with their human perception preserved for 60.71% of the processed utterances. Audio samples can be found in.
KW - adversarial perturbation on speaker embedding
KW - asynchronous anonymization
KW - human perception preservation
KW - voice privacy
UR - https://www.scopus.com/pages/publications/85214832957
U2 - 10.21437/Interspeech.2024-1888
DO - 10.21437/Interspeech.2024-1888
M3 - Conference article published in proceeding or book
AN - SCOPUS:85214832957
T3 - Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
SP - 4443
EP - 4447
BT - English
T2 - 25th Interspeech Conferece 2024
Y2 - 1 September 2024 through 5 September 2024
ER -