Identifying Multiple Outliers in Multivariate Data

Clicks: 2
ID: 304804
1992
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal
Emerging

Ranked #128 of 145 articles by views in Journal of the Royal Statistical Society Series B (Statistical Methodology)

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 145 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
SUMMARY We propose a procedure for the detection of multiple outliers in multivariate data. Let X be an n × p data matrix representing n observations on p variates. We first order the n observations, using an appropriately chosen robust measure of outlyingness, then divide the data set into two initial subsets: A ‘basic’ subset which contains p +1 ‘good’ observations and a ‘non-basic’ subset which contains the remaining n - p - 1 observations. Second, we compute the relative distance from each point in the data set to the centre of the basic subset, relative to the (possibly singular) covariance matrix of the basic subset. Third, we rearrange the n observations in ascending order accordingly, then divide the data set into two subsets: A basic subset which contains the first p + 2 observations and a non-basic subset which contains the remaining n - p - 2 observations. This process is repeated until an appropriately chosen stopping criterion is met. The final non-basic subset of observations is declared an outlying subset. The procedure proposed is illustrated and compared with existing methods by using several data sets. The procedure is simple, computationally inexpensive, suitable for automation, computable with widely available software packages, effective in dealing with masking and swamping problems and, most importantly, successful in identifying multivariate outliers.
Reference Key
openalex_W2107554480 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Ali S. Hadi
Journal Journal of the Royal Statistical Society Series B (Statistical Methodology)
Year 1992
DOI
10.1111/j.2517-6161.1992.tb01449.x
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.