Regularized Urdu Speech Recognition with Semi-Supervised Deep Learning

Clicks: 171
ID: 268289
2019
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal
Steady

Ranked #159 of 187 articles by views in applied sciences

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 187 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Automatic Speech Recognition, (ASR) has achieved the best results for English, with end-to-end neural network based supervised models. These supervised models need huge amounts of labeled speech data for good generalization, which can be quite a challenge to obtain for low-resource languages like Urdu. Most models proposed for Urdu ASR are based on Hidden Markov Models (HMMs). This paper proposes an end-to-end neural network model, for Urdu ASR, regularized with dropout, ensemble averaging and Maxout units. Dropout and ensembles are averaging techniques over multiple neural network models while Maxout are units in a neural network which adapt their activation functions. Due to limited labeled data, Semi Supervised Learning (SSL) techniques are also incorporated to improve model generalization. Speech features are transformed into a lower dimensional manifold using an unsupervised dimensionality-reduction technique called Locally Linear Embedding (LLE). Transformed data along with higher dimensional features is used to train neural networks. The proposed model also utilizes label propagation-based self-training of initially trained models and achieves a Word Error Rate (WER) of 4% less than that reported as the benchmark on the same Urdu corpus using HMM. The decrease in WER after incorporating SSL is more significant with an increased validation data size.
Reference Key
humayun2019appliedregularized Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Mohammad Ali Humayun;Ibrahim A. Hameed;Syed Muslim Shah;Sohaib Hassan Khan;Irfan Zafar;Saad Bin Ahmed;Junaid Shuja;Ali Humayun, Mohammad;Hameed, Ibrahim A.;Muslim Shah, Syed;Hassan Khan, Sohaib;Zafar, Irfan;Bin Ahmed, Saad;Shuja, Junaid;
Journal applied sciences
Year 2019
DOI
10.3390/app9091956
URL
Keywords

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.