Skip to main navigation Skip to search Skip to main content

Uncertainty-Aware Multi-Head Multi-Mode Knowledge Distillation for Self-Supervised Speaker Verification

Research output: Journal article publicationJournal articleAcademic researchpeer-review

Abstract

In recent years, Self-Distillation with NO labels (DINO) has achieved promising results in training speaker embedding networks. However, DINO requires the teacher and student networks to have identical architecture, limiting the flexibility of the student network. To address this limitation, we propose Uncertainty-Aware Multi-Head, Multi-Mode (UA-MeMo) self-supervised learning based on knowledge distillation. UAMeMo extends vanilla DINO by using two distinct architectures for both teacher and student, enabling selfdistillation when they share the same architecture and crossdistillation when their architectures differ. To prevent architectural discrepancies from derailing student learning, the projection head is split into separate self and crossheads, and a weighted crossdistillation scheme reduces the weights for samples with greater discrepancy between the teacher's and student's outputs. To better balance self and crossdistillation, we propose an uncertaintyweighted adaptive loss strategy that enables the model to dynamically assign weights based on distillation uncertainty. To prevent false negatives from interfering with knowledge transfer, UA-MeMo applies distillation with contrastive learning at the embedding level during the early stages of training only; in the later training stage, it dynamically discontinues the contrastive learning to focus on pure knowledge distillation. UA-MeMo achieves an impressive EER of 2.71% on VoxCeleb and 12.85% on CN-Celeb using a small ECAPA-TDNN backbone. This corresponds to a relative EER reduction of about 32.8% on Voxceleb1 compared to the baseline.

Original languageEnglish
Pages (from-to)2101-2114
Number of pages14
JournalIEEE Transactions on Audio, Speech and Language Processing
Volume34
DOIs
Publication statusPublished - Apr 2026

Keywords

  • cross-distillation
  • DINO
  • knowledge distillation
  • self-supervised learning
  • Speaker verification
  • uncertainty-aware distillation SL-SPR Speaker recognition, identification and verification

ASJC Scopus subject areas

  • Acoustics and Ultrasonics
  • Electrical and Electronic Engineering

Fingerprint

Dive into the research topics of 'Uncertainty-Aware Multi-Head Multi-Mode Knowledge Distillation for Self-Supervised Speaker Verification'. Together they form a unique fingerprint.

Cite this