Skip to main navigation Skip to search Skip to main content

Classification of Accented English Using CNN Model Trained on Amplitude Mel-Spectrograms

  • Mariia Lesnichaia
  • , Veranika Mikhailava
  • , Natalia Bogach
  • , Iurii Lezhenin
  • , John Blake
  • , Evgeny Pyshkin
  • St. Petersburg State Polytechnical University
  • Center for Language Research (CLR)
  • Speech Technology Center (STC)

Research output: Contribution to journalConference articlepeer-review

Abstract

Automatic speech recognition is hindered by the linguistic differences occurring in accented speech. This paper advances a classification method for accented speech using a CNN-based model trained and tested on English with Germanic, Romance and Slavic accents. The input feature set was examined to find the optimal combination of time-frequency and energy characteristics of speech fed into the machine learning model. We also tuned model hyperparameters and the dimensionality of input features. We argue that mel-scale amplitude spectrograms on a liner scale appear more powerful in accent classification tasks compared to conventional feature sets based on MFCCs and raw spectrograms. Our models used only sparse data from the Speech Accent Archive, yet produced state-of-the-art classification results for English with Germanic, Romance and Slavic accents. The accuracy of our models trained on linear scale amplitude mel-spectrograms ranged from 0.964 to 0.987, outperforming existing models classifying accents using the same dataset.

Original languageEnglish
Pages (from-to)3669-3673
Number of pages5
JournalProceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
Volume2022
Issue numberSeptember
DOIs
Publication statusPublished - 18 Sept 2022
Event23rd Annual Conference of the International Speech Communication Association, INTERSPEECH 2022 - Incheon, Korea, Republic of
Duration: 18 Sept 202222 Sept 2022

Keywords

  • Amplitude mel-spectrogram
  • Automatic accent identification
  • Convolutional neural networks (CNN)
  • Mel-frequency cepstral coefficients (MFCC)

Fingerprint

Dive into the research topics of 'Classification of Accented English Using CNN Model Trained on Amplitude Mel-Spectrograms'. Together they form a unique fingerprint.

Cite this