MMAD: Multi-label Micro-Action Detection in Videos
Clicks: 31
ID: 282432
2024
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
Emerging Content
9.0
/100
31 views
12 readers
AI Quality Assessment
Not analyzed
Readership in this journal
EmergingRanked #693 of 803 articles by views in arXiv
Most read
Least read
Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 803 in total.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
Human body actions are an important form of non-verbal communication in
social interactions. This paper specifically focuses on a subset of body
actions known as micro-actions, which are subtle, low-intensity body movements
with promising applications in human emotion analysis. In real-world scenarios,
human micro-actions often temporally co-occur, with multiple micro-actions
overlapping in time, such as concurrent head and hand movements. However,
current research primarily focuses on recognizing individual micro-actions
while overlooking their co-occurring nature. To address this gap, we propose a
new task named Multi-label Micro-Action Detection (MMAD), which involves
identifying all micro-actions in a given short video, determining their start
and end times, and categorizing them. Accomplishing this requires a model
capable of accurately capturing both long-term and short-term action
relationships to detect multiple overlapping micro-actions. To facilitate the
MMAD task, we introduce a new dataset named Multi-label Micro-Action-52
(MMA-52) and propose a baseline method equipped with a dual-path
spatial-temporal adapter to address the challenges of subtle visual change in
MMAD. We hope that MMA-52 can stimulate research on micro-action analysis in
videos and prompt the development of spatio-temporal modeling in human-centric
video understanding. The proposed MMA-52 dataset is available at:
https://github.com/VUT-HFUT/Micro-Action.
| Reference Key |
wang2024mmad
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | Kun Li; Pengyu Liu; Dan Guo; Fei Wang; Zhiliang Wu; Hehe Fan; Meng Wang |
| Journal | arXiv |
| Year | 2024 |
| DOI |
DOI not found
|
| URL | |
| Keywords |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.