Multi-class protein fold recognition using support vector machines and neural networks

Clicks: 1
ID: 302263
2001
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal

Ranked #748 of 829 articles by views in BMC Bioinformatics

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 829 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Abstract Motivation: Protein fold recognition is an important approach to structure discovery without relying on sequence similarity. We study this approach with new multi-class classification methods and examined many issues important for a practical recognition system. Results: Most current discriminative methods for protein fold prediction use the one-against-others method, which has the well-known ‘False Positives’ problem. We investigated two new methods: the unique one-against-others and the all-against-all methods. Both improve prediction accuracy by 14–110% on a dataset containing 27 SCOP folds. We used the Support Vector Machine (SVM) and the Neural Network (NN) learning methods as base classifiers. SVMs converges fast and leads to high accuracy. When scores of multiple parameter datasets are combined, majority voting reduces noise and increases recognition accuracy. We examined many issues involved with large number of classes, including dependencies of prediction accuracy on the number of folds and on the number of representatives in a fold. Overall, recognition systems achieve 56% fold prediction accuracy on a protein test dataset, where most of the proteins have below 25% sequence identity with the proteins used in training. Supplementary information: The protein parameter datasets used in this paper are available online (http://www.nersc.gov/~cding/protein). Contact: chqding@lbl.gov; ildubchak@lbl.gov * To whom correspondence should be addressed.
Reference Key
openalex_W2109109045 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Chris Ding, Inna Dubchak
Journal BMC Bioinformatics
Year 2001
DOI
10.1093/bioinformatics/17.4.349
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.