Bilingual Spectral Emotion Learning Through Patch-Encoded VGG-16 Features and a Full Vision Transformer Pipeline

Clicks: 2
ID: 312660
2025
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal
Emerging

Ranked #633 of 705 articles by views in Journal of Computing & Biomedical Informatics

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 705 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
The research introduces a new bilingual speech emotion recognition system which integrates patch-encoded VGG-16 spectral cues with a whole Vision Transformer pipeline to learn the affective cues on English and Gujarati speech. Mel-spectrograms are initially inputted into a frozen VGG-16 backbone to obtain high-level spatial spectral features, and these features are then split into regular patches and transformed into an embedding space to be used to represent them in a transformer-based global attention model. It is tested on four reference English emotional speech datasets, including RAVDESS, CREMA-D, SAVEE, and TESS, with the results of accuracy 99% for all. To evaluate robustness on non-controlled data, a hand-collected bilingual corpus of student recordings was created, where the model was able to assess English and Gujarati speech with accuracy on 90% and 88% percent respectively. Such findings show that the convolutional spectral extraction with contextual learning by transformers is an effective way of modeling cross-lingual emotional differences and outperforms traditional convolution-only or transformer-only models. The bilingual results also show that the model can be used to achieve stable performance with languages that have different phonetics and prosody and thus is applicable to scalable and inclusive emotion-sensitive speech technologies in practice through interactive assistants, call-center analytics and affect sensitive human-machine interfaces.
Reference Key
imported_1777055855_69ebb86f86d34 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Sheshang Degadwala
Journal Journal of Computing & Biomedical Informatics
Year 2025
DOI
10.56979/1001/2025/1146
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.