What matters beyond model choice for wearable sleep staging? How personalization, evaluation choices, and easy-to-classify wake impact performance
Clicks: 1
ID: 314982
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
0.0
/100
1 views
0 readers
AI Quality Assessment
Not analyzed
Readership in this journal
Ranked #6 of 39 articles by views in SLEEP Advances
Most read
Least read
Bar heights use a square-root scale.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
Abstract Study Objectives Methods to improve sleep classification of wearable data often emphasize model choice. However, without widely used benchmark datasets, it is difficult to determine factors driving performance gains (model architecture, dataset selection, or evaluation decisions). This study examines how personalization, validation strategy, and dataset characteristics influence sleep classification performance beyond model choice. Methods We collected SleepAccel Clinical, a dataset of Apple Watch acceleration and polysomnography from 31 individuals with sleep apnea. Neural network models were trained on this dataset alongside SleepAccel (31 healthy individuals) and DREAMT (100 individuals with sleep disorders) to evaluate how training and testing set choices impact performance. All models were compared against an Easy To Classify (ETC) wake model, which estimates wake probability by smoothing and scaling activity in a brief time window. The ETC model’s average area under the recever operating curve (AUROC) on each dataset was used to quantify intrinsic ease of classification. Results The ETC model accounted for substantial amount of the models’ performance across the architectures, datasets, and training conditions, with significant correlations (p < .01) between ETC performance and all other scenarios. Including individuals with obstructive sleep apnea (OSA) in the training data improved performance when testing on datasets in which sleep disorders were suspected. Conclusions Sleep classification performance depends heavily on training and testing dataset characteristics, not solely on model choice. The presence of ETC wake epochs can markedly inflate performance metrics. Researchers should benchmark models on widely available datasets and quantify intrinsic class separability when introducing new datasets. This paper is part of the Consumer Sleep Technology Collection.
| Reference Key |
openalex_W7162425698
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | Eric Cantón, Franco Tavella, Christopher Drake, Olivia Walch, Philip Cheng |
| Journal | SLEEP Advances |
| Year | 2026 |
| DOI |
10.1093/sleepadvances/zpag051
|
| URL | |
| Keywords | Keywords not found |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.