Assessing the low complexity of protein sequences via the low complexity triangle.

6533b86ffe1ef96bd12cdd0d

RESEARCH PRODUCT

Assessing the low complexity of protein sequences via the low complexity triangle.

subject

Proteome Proteomes Computer science Protein Sequencing Biochemistry Database and Informatics Methods Sequence Analysis Protein Protein methods Peptide sequence chemistry.chemical_classification 0303 health sciences Sequence Multidisciplinary 030302 biochemistry & molecular biology Q R Genomics Amino acid Tandem Repeats Proteome Amino Acid Analysis Medicine Sequence Analysis Research Article Repetitive Sequences Amino Acid Bioinformatics Sequence analysis Science Research and Analysis Methods Genome Complexity 03 medical and health sciences Protein Domains Amino Acid Sequence Analysis Tandem repeat Genetics Humans Fraction (mathematics)Repeated Sequences Amino Acid Sequence Molecular Biology Techniques Sequencing Techniques Representation (mathematics)Molecular Biology 030304 developmental biology Molecular Biology Assays and Analysis Techniques business.industry Biology and Life Sciences Proteins Computational Biology Pattern recognition chemistry Globular Proteins Artificial intelligence business

description

Background Proteins with low complexity regions (LCRs) have atypical sequence and structural features. Their amino acid composition varies from the expected, determined proteome-wise, and they do not follow the rules of structural folding that prevail in globular regions. One way to characterize these regions is by assessing the repeatability of a sequence, that is, calculating the local propensity of a region to be part of a repeat. Results We combine two local measures of low complexity, repeatability (using the RES algorithm) and fraction of the most frequent amino acid, to evaluate different proteomes, datasets of protein regions with specific features, and individual cases of proteins with extreme compositions. We apply a representation called ‘low complexity triangle’ as a proof-of-concept to represent the low complexity measured values. Results show that proteomes have distinct signatures in the low complexity triangle, and that these signatures are associated to complexity features of the sequences. We developed a web tool called LCT (http://cbdm-01.zdv.uni-mainz.de/~munoz/lct/) to allow users to calculate the low complexity triangle of a given protein or region of interest. Conclusions The low complexity triangle proves to be a suitable procedure to represent the general low complexity of a sequence or protein dataset. Homorepeats, direpeats, compositionally biased regions and globular regions occupy characteristic positions in the triangle. The described pipeline can be used to characterize LCRs and may help in quantifying the content of degenerated tandem repeats in proteins and proteomes.

year	journal	country	edition	language
2020-12-01	PLoS ONE

10.1371/journal.pone.0239154 https://doaj.org/article/c7e6900a33e7410b8da56cedd9018b8c