Faculty Publications

Predicting Protein Residue-Residue Contacts Using Random Forests and Deep Networks

Joseph Luttrell IV, University of Southern MississippiFollow
Tong Liu, Univerity of MiamiFollow
Chaoyang Zhang, University of Southern MississippiFollow
Zheng Wang, University of MiamiFollow

Document Type

Article

Publication Date

3-14-2019

Department

Computing

School

Computing Sciences and Computer Engineering

Abstract

Background: The ability to predict which pairs of amino acid residues in a protein are in contact with each other offers many advantages for various areas of research that focus on proteins. For example, contact prediction can be used to reduce the computational complexity of predicting the structure of proteins and even to help identify functionally important regions of proteins. These predictions are becoming especially important given the relatively low number of experimentally determined protein structures compared to the amount of available protein sequence data.

Results: Here we have developed and benchmarked a set of machine learning methods for performing residue-residue contact prediction, including random forests, direct-coupling analysis, support vector machines, and deep networks (stacked denoising autoencoders). These methods are able to predict contacting residue pairs given only the amino acid sequence of a protein. According to our own evaluations performed at a resolution of +/− two residues, the predictors we trained with the random forest algorithm were our top performing methods with average top 10 prediction accuracy scores of 85.13% (short range), 74.49% (medium range), and 54.49% (long range). Our ensemble models (stacked denoising autoencoders combined with support vector machines) were our best performing deep network predictors and achieved top 10 prediction accuracy scores of 75.51% (short range), 60.26% (medium range), and 43.85% (long range) using the same evaluation. These tests were blindly performed on targets from the CASP11 dataset; and the results suggested that our models achieved comparable performance to contact predictors developed by groups that participated in CASP11.

Conclusions: Due to the challenging nature of contact prediction, it is beneficial to develop and benchmark a variety of different prediction methods. Our work has produced useful tools with a simple interface that can provide contact predictions to users without requiring a lengthy installation process. In addition to this, we have released our C++ implementation of the direct-coupling analysis method as a standalone software package. Both this tool and our RFcon web server are freely available to the public at http://dna.cs.miami.edu/RFcon/.

Comments

Published by 'BMC Bioinformatics' at 10.1186/s12859-019-2627-6.

Publication Title

BMC Bioinformatics

Volume

Issue

First Page

Last Page

Recommended Citation

Luttrell, J., Liu, T., Zhang, C., Wang, Z. (2019). Predicting Protein Residue-Residue Contacts Using Random Forests and Deep Networks. BMC Bioinformatics, 20(S2), 1-13.
Available at: https://aquila.usm.edu/fac_pubs/15984

Download

Find in your library

Included in

Bioinformatics Commons

COinS

Faculty Publications

Predicting Protein Residue-Residue Contacts Using Random Forests and Deep Networks

Document Type

Publication Date

Department

School

Abstract

Comments

Publication Title

Volume

Issue

First Page

Last Page

Recommended Citation

Included in

Search

Browse

Author Corner

Faculty Publications

Predicting Protein Residue-Residue Contacts Using Random Forests and Deep Networks

Authors

Document Type

Publication Date

Department

School

Abstract

Comments

Publication Title

Volume

Issue

First Page

Last Page

Recommended Citation

Included in

Share

Search

Browse

Author Corner