Disentangling the complexity of low complexity proteins

6533b7dcfe1ef96bd1271f8c

RESEARCH PRODUCT

Disentangling the complexity of low complexity proteins

Pablo Mier Pau Bernadó Zoltán Gáspári Christos A. Ouzounis Vasilis J. Promponas Andrey V. Kajava John M. Hancock Silvio C. E. Tosatto Zsuzsanna Dosztanyi Miguel A. Andrade-navarro Pablo Mier Lisanna Paladin Lisanna Paladin Stella Tamana Sophia Petrosian Borbála Hajdu-soltész Annika Urbanek Aleksandra Gruca Dariusz Plewczynski Marcin Grynberg Pau Bernadó Zoltán Gáspári Stella Tamana Christos A. Ouzounis Vasilis J. Promponas Andrey V. Kajava John M. Hancock Silvio C. E. Tosatto Zsuzsanna Dosztanyi Miguel A. Andrade-navarro Sophia Petrosian Borbála Hajdu-soltész Annika Urbanek Aleksandra Gruca Dariusz Plewczynski Marcin Grynberg

subject

Protein Conformation Computer science Review Article Computational biology Measure (mathematics)Evolution Molecular Low complexity 03 medical and health sciences Protein Domains Amino Acid Sequence structure [SDV.BBM.BC]Life Sciences [q-bio]/Biochemistry Molecular Biology/Biochemistry [q-bio.BM]Databases Protein Molecular Biology 030304 developmental biology Structure (mathematical logic)0303 health sciences Sequence [SCCO.NEUR]Cognitive science/Neuroscience composition bias 030302 biochemistry & molecular biology Proteins disorder low complexity regions Structure and function [INFO.INFO-BI]Computer Science [cs]/Bioinformatics [q-bio.QM]Algorithms Information Systems

description

Abstract There are multiple definitions for low complexity regions (LCRs) in protein sequences, with all of them broadly considering LCRs as regions with fewer amino acid types compared to an average composition. Following this view, LCRs can also be defined as regions showing composition bias. In this critical review, we focus on the definition of sequence complexity of LCRs and their connection with structure. We present statistics and methodological approaches that measure low complexity (LC) and related sequence properties. Composition bias is often associated with LC and disorder, but repeats, while compositionally biased, might also induce ordered structures. We illustrate this dichotomy, and more generally the overlaps between different properties related to LCRs, using examples. We argue that statistical measures alone cannot capture all structural aspects of LCRs and recommend the combined usage of a variety of predictive tools and measurements. While the methodologies available to study LCRs are already very advanced, we foresee that a more comprehensive annotation of sequences in the databases will enable the improvement of predictions and a better understanding of the evolution and the connection between structure and function of LCRs. This will require the use of standards for the generation and exchange of data describing all aspects of LCRs. Short abstract There are multiple definitions for low complexity regions (LCRs) in protein sequences. In this critical review, we focus on the definition of sequence complexity of LCRs and their connection with structure. We present statistics and methodological approaches that measure low complexity (LC) and related sequence properties. Composition bias is often associated with LC and disorder, but repeats, while compositionally biased, might also induce ordered structures. We illustrate this dichotomy, plus overlaps between different properties related to LCRs, using examples.

year	journal	country	edition	language
2020-03-23	Briefings in Bioinformatics

https://doi.org/10.1093/bib/bbz007