accuracy increase for automatic visual russian speech recognition: viseme classes optimization

Clicks: 90
ID: 224199
2018
Article Quality & Performance Metrics
Overall Quality Improving Quality
0.0 /100
Combines engagement data with AI-assessed academic quality
AI Quality Assessment
Not analyzed
Abstract
Nowadays there are a lot of continuous studies on the correct viseme classes to be used for the most effective automatic lip-reading. The paper proposes a structured approach for the development of speaker-dependent classes of visemes. This method gives the possibility to create a set of phoneme-viseme correspondence maps, where each class has a different number of visemes from two to forty-eight with a constant number of phonemes. Viseme classes are based on their mapping from phonemes, which are converted into viseme groups during speech recognition process. With the usage of the obtained correspondence maps together with the database of audio-visual Russian speech HAVRUS the paper demonstrates the dependence of the visual speech recognition accuracy on the number of used viseme classes. The application of high-speed video data made it possible to expand the optimal set of viseme classes to twenty that resulted in recognition accuracy improvement by 1.34% compared to the standard set of fourteen classes.
Reference Key
2018nauno-tehnieskijaccuracy Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors ;D. V. Ivanko ;D. V. Fedotov;A. A. Karpov
Journal bmc pharmacology and toxicology
Year 2018
DOI
10.17586/2226-1494-2018-18-2-346-349
URL
Keywords

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.