TY - GEN
T1 - Demographic Attributes Prediction from Speech Using WavLM Embeddings
AU - Yang, Yuchen
AU - Thebaud, Thomas
AU - Dehak, Najim
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - This paper introduces a general classifier based on WavLM features, to infer demographic characteristics, such as age, gender, native language, education, and country, from speech. Demographic feature prediction plays a crucial role in applications like language learning, accessibility, and digital forensics, enabling more personalized and inclusive technologies. Leveraging pretrained models for embedding extraction, the proposed framework identifies key acoustic and linguistic features associated with demographic attributes, achieving a Mean Absolute Error (MAE) of 4.94 for age prediction and over 99.81% accuracy for gender classification across various datasets. Our system improves upon existing models by up to relative 30% in MAE and up to relative 10% in accuracy and F1 scores across tasks, leveraging a diverse range of datasets and large pretrained models to ensure robustness and generalizability. This study offers new insights into speaker diversity and provides a strong foundation for future research in speech-based demographic profiling.
AB - This paper introduces a general classifier based on WavLM features, to infer demographic characteristics, such as age, gender, native language, education, and country, from speech. Demographic feature prediction plays a crucial role in applications like language learning, accessibility, and digital forensics, enabling more personalized and inclusive technologies. Leveraging pretrained models for embedding extraction, the proposed framework identifies key acoustic and linguistic features associated with demographic attributes, achieving a Mean Absolute Error (MAE) of 4.94 for age prediction and over 99.81% accuracy for gender classification across various datasets. Our system improves upon existing models by up to relative 30% in MAE and up to relative 10% in accuracy and F1 scores across tasks, leveraging a diverse range of datasets and large pretrained models to ensure robustness and generalizability. This study offers new insights into speaker diversity and provides a strong foundation for future research in speech-based demographic profiling.
UR - https://www.scopus.com/pages/publications/105002711341
UR - https://www.scopus.com/pages/publications/105002711341#tab=citedBy
U2 - 10.1109/CISS64860.2025.10944678
DO - 10.1109/CISS64860.2025.10944678
M3 - Conference contribution
AN - SCOPUS:105002711341
T3 - 2025 59th Annual Conference on Information Sciences and Systems, CISS 2025
BT - 2025 59th Annual Conference on Information Sciences and Systems, CISS 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 59th Annual Conference on Information Sciences and Systems, CISS 2025
Y2 - 19 March 2025 through 21 March 2025
ER -