Synergistic Fusion of Clinical Interview EEG and Video for Depression Detection: A Cross-Modal Attention Approach
Fusion synergique de l'EEG et de la vidéo d'entretiens cliniques pour la détection de la dépression : une approche par attention intermodale
Clicks: 27
ID: 312652
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
Emerging Content
7.5
/100
27 views
5 readers
AI Quality Assessment
Not analyzed
Readership in this journal
EmergingRanked #20 of 705 articles by views in Journal of Computing & Biomedical Informatics
Most read
Least read
Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 705 in total.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
Objective quantification of Major Depressive Disorder (MDD) remains a substantial clinical challenge due to the inherent subjectivity of traditional diagnostic interviews. This paper presents a novel multimodal deep learning framework that synergistically integrates neurophysiological signals and behavioural cues for automated depression detection. Utilizing the Multi-modal Open Dataset for Mental-disorder Analysis (MODMA), we analyze synchronized 128-channel EEG and video recordings obtained during professional clinical assessments. Our architecture employs a dual-stream approach: a Graph Convolutional Network (GCN) combined with a Long Short-Term Memory (LSTM) network to capture the spatiotemporal dynamics of brain activity, and a 3D Convolutional Neural Network (3D-CNN) with a temporal attention mechanism to extract behavioral markers from facial expressions. A sophisticated cross-modal attention module is implemented to fuse these modalities, allowing the model to learn the complex interdependencies between neural states and overt behavior. To ensure clinical generalizability and prevent data leakage, the framework was evaluated using a strict subject-independent 10-fold cross-validation scheme. Experimental results demonstrate latest performance, achieving an Accuracy of 92.1 % and an F1-Score of 92.5 %. These findings suggest that the proposed multimodal integration offers a powerful and objective tool for mental health screening, enhancing diagnostic precision through the fusion of brain and behavioral biomarkers.
La quantification objective du trouble dépressif majeur (TDM) demeure un défi clinique substantiel en raison de la subjectivité inhérente aux entretiens diagnostiques traditionnels. Cet article présente un nouveau cadre d'apprentissage profond multimodal intégrant de manière synergique des signaux neurophysiologiques et des indices comportementaux pour la détection automatisée de la dépression. En nous appuyant sur le jeu de données *Multi-modal Open Dataset for Mental-disorder Analysis* (MODMA), nous analysons des enregistrements synchronisés d'EEG à 128 canaux et de vidéos acquis lors d'évaluations cliniques professionnelles. Notre architecture repose sur une approche à double flux : un réseau de neurones convolutifs sur graphe (GCN) combiné à un réseau à mémoire à long et court terme (LSTM) afin de capturer la dynamique spatio-temporelle de l'activité cérébrale, ainsi qu'un réseau de neurones convolutifs 3D (3D-CNN) doté d'un mécanisme d'attention temporelle pour extraire les marqueurs comportementaux issus des expressions faciales. Un module sophistiqué d'attention intermodale est implémenté pour fusionner ces modalités, permettant au modèle d'apprendre les interdépendances complexes entre les états neuronaux et les comportements manifestes. Afin de garantir la généralisabilité clinique et de prévenir toute fuite de données, le cadre a été évalué selon un schéma rigoureux de validation croisée à 10 plis indépendante des sujets. Les résultats expérimentaux démontrent des performances de pointe, atteignant une exactitude de 92,1 % et un score F1 de 92,5 %. Ces conclusions suggèrent que l'intégration multimodale proposée constitue un outil puissant et objectif pour le dépistage en santé mentale, améliorant la précision diagnostique grâce à la fusion de biomarqueurs cérébraux et comportementaux.
| Reference Key |
imported_1777055799_69ebb837a45a5
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | Chokka Anuradha |
| Journal | Journal of Computing & Biomedical Informatics |
| Year | 2026 |
| DOI |
10.56979/1002/2026/1222
|
| URL | |
| Keywords | Keywords not found |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.