Are Visual-Language Models Effective in Action Recognition? A Comparative Study
Clicks: 54
ID: 282442
2024
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
Emerging Content
15.9
/100
54 views
40 readers
AI Quality Assessment
Not analyzed
Readership in this journal
EmergingRanked #516 of 803 articles by views in arXiv
Most read
Least read
Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 803 in total.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
Current vision-language foundation models, such as CLIP, have recently shown
significant improvement in performance across various downstream tasks.
However, whether such foundation models significantly improve more complex
fine-grained action recognition tasks is still an open question. To answer this
question and better find out the future research direction on human behavior
analysis in-the-wild, this paper provides a large-scale study and insight on
current state-of-the-art vision foundation models by comparing their transfer
ability onto zero-shot and frame-wise action recognition tasks. Extensive
experiments are conducted on recent fine-grained, human-centric action
recognition datasets (e.g., Toyota Smarthome, Penn Action, UAV-Human, TSU,
Charades) including action classification and segmentation.
| Reference Key |
brémond2024are
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | Mahmoud Ali; Di Yang; François Brémond |
| Journal | arXiv |
| Year | 2024 |
| DOI |
DOI not found
|
| URL | |
| Keywords |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.