A qualitative systematic review of intra-speaker variation in the human voice
Audio deepfake detection is essential for addressing societal challenges such as differentiating real news from fake content or authenticating voice recordings in legal contexts. However, identifying whether a voice is human or AI-generated requires knowing which characteristics to examine, and the...
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Online |
| Language: | English |
| Published: |
Universidade Estadual de Campinas
2025
|
| Subjects: | |
| Online Access: | https://econtents.sbu.unicamp.br/inpec/index.php/joss/article/view/20738 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| Summary: | Audio deepfake detection is essential for addressing societal challenges such as differentiating real news from fake content or authenticating voice recordings in legal contexts. However, identifying whether a voice is human or AI-generated requires knowing which characteristics to examine, and the choice of voice features for this task is relatively unguided. This justifies the systematic review presented in this paper. Hypothesizing that human voices exhibit more intra-speaker variation than deepfakes, the aim of this review has been to summarize and analyze the published studies on the topic of intra-speaker variation in human voice. A systematic search was conducted in Web of Science, the Cochrane Library, and the electronic database of the International Journal of Speech Language and the Law, initially identifying 305 studies. After removing duplicates and applying inclusion/exclusion criteria, 36 articles were selected for analysis. Findings highlight speaking style as a major factor in intra-speaker variation affecting various acoustic parameters. This review suggests that experts may prioritize features that show higher within-speaker variation, while noting that their utility for deepfake detection must be verified on deepfake datasets. |
|---|---|
| ISSN: | 2236-9740 |