regression models tolerant to massively missing data: a case study in solar-radiation nowcasting
Clicks: 125
ID: 256021
2014
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
Steady Performance
30.0
/100
125 views
15 readers
AI Quality Assessment
Not analyzed
Readership in this journal
SteadyRanked #133 of 183 articles by views in bioorganic & medicinal chemistry
Most read
Least read
Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 183 in total.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
Statistical models for environmental monitoring strongly rely on automatic
data acquisition systems that use various physical sensors. Often, sensor
readings are missing for extended periods of time, while model outputs need
to be continuously available in real time. With a case study in solar-radiation nowcasting, we investigate how to deal with massively missing data
(around 50% of the time some data are unavailable) in such situations.
Our goal is to analyze characteristics of missing data and recommend
a strategy for deploying regression models which would be robust to missing
data in situations where data are massively missing. We are after one model
that performs well at all times, with and without data gaps. Due to the need
to provide instantaneous outputs with minimum energy consumption for
computing in the data streaming setting, we dismiss computationally demanding
data imputation methods and resort to a mean replacement, accompanied with a
robust regression model. We use an established strategy for assessing
different regression models and for determining how many missing sensor readings
can be tolerated before model outputs become obsolete. We experimentally
analyze the accuracies and robustness to missing data of seven linear regression
models. We recommend using the regularized PCA regression with our established
guideline in training regression models, which themselves are robust to
missing data.
| Reference Key |
liobait2014atmosphericregression
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | ;I. Žliobaitė;J. Hollmén;H. Junninen |
| Journal | bioorganic & medicinal chemistry |
| Year | 2014 |
| DOI |
10.5194/amt-7-4387-2014
|
| URL | |
| Keywords |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.