What matters beyond model choice for wearable sleep staging? How personalization, evaluation choices, and easy-to-classify wake impact performance

Clicks: 1
ID: 314982
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal

Ranked #6 of 39 articles by views in SLEEP Advances

Most read Least read

Bar heights use a square-root scale.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Abstract Study Objectives Methods to improve sleep classification of wearable data often emphasize model choice. However, without widely used benchmark datasets, it is difficult to determine factors driving performance gains (model architecture, dataset selection, or evaluation decisions). This study examines how personalization, validation strategy, and dataset characteristics influence sleep classification performance beyond model choice. Methods We collected SleepAccel Clinical, a dataset of Apple Watch acceleration and polysomnography from 31 individuals with sleep apnea. Neural network models were trained on this dataset alongside SleepAccel (31 healthy individuals) and DREAMT (100 individuals with sleep disorders) to evaluate how training and testing set choices impact performance. All models were compared against an Easy To Classify (ETC) wake model, which estimates wake probability by smoothing and scaling activity in a brief time window. The ETC model’s average area under the recever operating curve (AUROC) on each dataset was used to quantify intrinsic ease of classification. Results The ETC model accounted for substantial amount of the models’ performance across the architectures, datasets, and training conditions, with significant correlations (p < .01) between ETC performance and all other scenarios. Including individuals with obstructive sleep apnea (OSA) in the training data improved performance when testing on datasets in which sleep disorders were suspected. Conclusions Sleep classification performance depends heavily on training and testing dataset characteristics, not solely on model choice. The presence of ETC wake epochs can markedly inflate performance metrics. Researchers should benchmark models on widely available datasets and quantify intrinsic class separability when introducing new datasets. This paper is part of the Consumer Sleep Technology Collection.
Reference Key
openalex_W7162425698 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Eric Cantón, Franco Tavella, Christopher Drake, Olivia Walch, Philip Cheng
Journal SLEEP Advances
Year 2026
DOI
10.1093/sleepadvances/zpag051
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.