Regularized Urdu Speech Recognition with Semi-Supervised Deep Learning
Clicks: 171
ID: 268289
2019
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
Steady Performance
30.0
/100
171 views
27 readers
AI Quality Assessment
Not analyzed
Readership in this journal
SteadyRanked #159 of 187 articles by views in applied sciences
Most read
Least read
Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 187 in total.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
Automatic Speech Recognition, (ASR) has achieved the best results for English, with end-to-end neural network based supervised models. These supervised models need huge amounts of labeled speech data for good generalization, which can be quite a challenge to obtain for low-resource languages like Urdu. Most models proposed for Urdu ASR are based on Hidden Markov Models (HMMs). This paper proposes an end-to-end neural network model, for Urdu ASR, regularized with dropout, ensemble averaging and Maxout units. Dropout and ensembles are averaging techniques over multiple neural network models while Maxout are units in a neural network which adapt their activation functions. Due to limited labeled data, Semi Supervised Learning (SSL) techniques are also incorporated to improve model generalization. Speech features are transformed into a lower dimensional manifold using an unsupervised dimensionality-reduction technique called Locally Linear Embedding (LLE). Transformed data along with higher dimensional features is used to train neural networks. The proposed model also utilizes label propagation-based self-training of initially trained models and achieves a Word Error Rate (WER) of 4% less than that reported as the benchmark on the same Urdu corpus using HMM. The decrease in WER after incorporating SSL is more significant with an increased validation data size.
| Reference Key |
humayun2019appliedregularized
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | Mohammad Ali Humayun;Ibrahim A. Hameed;Syed Muslim Shah;Sohaib Hassan Khan;Irfan Zafar;Saad Bin Ahmed;Junaid Shuja;Ali Humayun, Mohammad;Hameed, Ibrahim A.;Muslim Shah, Syed;Hassan Khan, Sohaib;Zafar, Irfan;Bin Ahmed, Saad;Shuja, Junaid; |
| Journal | applied sciences |
| Year | 2019 |
| DOI |
10.3390/app9091956
|
| URL | |
| Keywords |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.