6533b857fe1ef96bd12b5166
RESEARCH PRODUCT
Normalised compression distance and evolutionary distance of genomic sequences: comparison of clustering results
Massimo La RosaAlfonso UrsoSalvatore GaglioRiccardo Rizzosubject
Settore ING-INF/05 - Sistemi Di Elaborazione Delle Informazionibusiness.industryCompression (functional analysis)Metric (mathematics)Normalized compression distanceuniversal similarity metric USM clustering DNA sequences normalised compression distance evolutionary distance genomic sequences nonlinear mapping bioinformaticsPattern recognitionArtificial intelligenceCluster analysisbusinessDistance matrices in phylogenyMathematicsdescription
Genomic sequences are usually compared using evolutionary distance, a procedure that implies the alignment of the sequences. Alignment of long sequences is a time consuming procedure and the obtained dissimilarity results is not a metric. Recently, the normalised compression distance was introduced as a method to calculate the distance between two generic digital objects and it seems a suitable way to compare genomic strings. In this paper, the clustering and the non-linear mapping obtained using the evolutionary distance and the compression distance are compared, in order to understand if the two distances sets are similar.
year | journal | country | edition | language |
---|---|---|---|---|
2009-01-01 |