Estimating the Number of Clusters in a Data Set Via the Gap Statistic

Clicks: 1
ID: 289431
2001
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal

Ranked #98 of 145 articles by views in Journal of the Royal Statistical Society Series B (Statistical Methodology)

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 145 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Summary We propose a method (the ‘gap statistic’) for estimating the number of clusters (groups) in a set of data. The technique uses the output of any clustering algorithm (e.g. K-means or hierarchical), comparing the change in within-cluster dispersion with that expected under an appropriate reference null distribution. Some theory is developed for the proposal and a simulation study shows that the gap statistic usually outperforms other methods that have been proposed in the literature.
Reference Key
openalex_W2071949631 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Robert Tibshirani, Guenther Walther, Trevor Hastie
Journal Journal of the Royal Statistical Society Series B (Statistical Methodology)
Year 2001
DOI
10.1111/1467-9868.00293
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.