Multi-Scale Speaker Diarization with Neural Affinity Score Fusion

Tae Jin Park,Manoj Kumar,Shrikanth S. Narayanan

Multi-Scale Speaker Diarization with Neural Affinity Score Fusion

2021

Tae Jin Park
Manoj Kumar
Shrikanth S. Narayanan

Identifying the identity of the speaker of short segments in human dialogue has been considered one of the most challenging problems in speech signal processing. Speaker representations of short speech segments tend to be unreliable, resulting in poor fidelity of speaker representations in tasks requiring speaker recognition. In this paper, we propose an unconventional method that tackles the trade-off between temporal resolution and the quality of the speaker representations. To find a set of weights that balance the scores from multiple temporal scales of segments, a neural affinity score fusion model is presented. Using the CALLHOME dataset, we show that our proposed multi-scale segmentation and integration approach can achieve a state-of-the-art diarization performance.

Keywords:

Segmentation
Speech recognition
identity
Signal processing
Temporal resolution
Speaker diarisation
Computer science
Speaker recognition
Fidelity
Set (psychology)

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations