Search results for "Sequence analysis"

showing 10 items of 1349 documents

The Power of Word-Frequency Based Alignment-Free Functions: a Comprehensive Large-Scale Experimental Analysis

2021

Abstract Motivation Alignment-free (AF) distance/similarity functions are a key tool for sequence analysis. Experimental studies on real datasets abound and, to some extent, there are also studies regarding their control of false positive rate (Type I error). However, assessment of their power, i.e. their ability to identify true similarity, has been limited to some members of the D2 family. The corresponding experimental studies have concentrated on short sequences, a scenario no longer adequate for current applications, where sequence lengths may vary considerably. Such a State of the Art is methodologically problematic, since information regarding a key feature such as power is either mi…

Statistics and ProbabilitySequenceSimilarity (geometry)Settore INF/01 - Informaticasequence analysisComputer sciencepower statisticsAlignment-Free Genomic Analysis Big Data Software Platforms Bioinformatics AlgorithmsScale (descriptive set theory)Function (mathematics)computer.software_genreBiochemistryComputer Science ApplicationsSet (abstract data type)Computational MathematicsRange (mathematics)Computational Theory and Mathematicssequence analysis; power statistics; alignment-free functionsalignment-free functionsData miningCompleteness (statistics)Molecular BiologycomputerType I and type II errors
researchProduct

Long read alignment based on maximal exact match seeds

2012

Abstract Motivation: The explosive growth of next-generation sequencing datasets poses a challenge to the mapping of reads to reference genomes in terms of alignment quality and execution speed. With the continuing progress of high-throughput sequencing technologies, read length is constantly increasing and many existing aligners are becoming inefficient as generated reads grow larger. Results: We present CUSHAW2, a parallelized, accurate, and memory-efficient long read aligner. Our aligner is based on the seed-and-extend approach and uses maximal exact matches as seeds to find gapped alignments. We have evaluated and compared CUSHAW2 to the three other long read aligners BWA-SW, Bowtie2 an…

Statistics and ProbabilitySequencing and Sequence AnalysisTheoretical computer scienceGenomicsBiologyBiochemistrySoftwareHumansMolecular BiologyAlignment-free sequence analysisExact matchSupplementary dataGenome Humanbusiness.industryChromosome MappingHigh-Throughput Nucleotide SequencingGenomicsSequence Analysis DNAOriginal PapersComputer Science ApplicationsComputational MathematicsComputational Theory and MathematicsComputer engineeringScalabilitybusinessSequence AlignmentAlgorithmsSoftwareBioinformatics
researchProduct

Overlap and diversity in antimicrobial peptide databases: Compiling a non-redundant set of sequences

2015

Abstract Motivation: The large variety of antimicrobial peptide (AMP) databases developed to date are characterized by a substantial overlap of data and similarity of sequences. Our goals are to analyze the levels of redundancy for all available AMP databases and use this information to build a new non-redundant sequence database. For this purpose, a new software tool is introduced. Results: A comparative study of 25 AMP databases reveals the overlap and diversity among them and the internal diversity within each database. The overlap analysis shows that only one database (Peptaibol) contains exclusive data, not present in any other, whereas all sequences in the LAMP_Patent database are inc…

Statistics and ProbabilitySimilarity (geometry)Computer scienceSequence analysisAntimicrobial peptidesPeptaibolPeptidecomputer.software_genreProceduresBiochemistrySet (abstract data type)chemistry.chemical_compoundProtein methodsSequence Analysis ProteinRedundancy (engineering)HumansDatabases ProteinMolecular BiologyAntimicrobial cationic peptideschemistry.chemical_classificationSequenceAntimicrobial cationic peptideDatabaseSequence databaseSequence analysisComputer Science ApplicationsAlgorithmComputational MathematicsChemistryProtein databaseComputational Theory and MathematicschemistryData miningNucleic acid databaseDatabases Nucleic AcidcomputerSoftwareAlgorithmsHuman
researchProduct

SKINK: a web server for string kernel based kink prediction in α-helices

2014

Abstract Motivation: The reasons for distortions from optimal α-helical geometry are widely unknown, but their influences on structural changes of proteins are significant. Hence, their prediction is a crucial problem in structural bioinformatics. Here, we present a new web server, called SKINK, for string kernel based kink prediction. Extending our previous study, we also annotate the most probable kink position in a given α-helix sequence. Availability and implementation: The SKINK web server is freely accessible at http://biows-inf.zdv.uni-mainz.de/skink. Moreover, SKINK is a module of the BALL software, also freely available at www.ballview.org. Contact:  benny.kneissl@roche.com

Statistics and ProbabilitySkinkWeb serverTheoretical computer scienceComputer scienceReal-time computingcomputer.software_genreBiochemistryProtein Structure SecondaryStructural bioinformaticsSoftwareSequence Analysis ProteinString kernelPosition (vector)Ball (mathematics)Molecular BiologyInternetSequencebiologybusiness.industryComputational BiologyProteinsbiology.organism_classificationComputer Science ApplicationsComputational MathematicsComputational Theory and MathematicsbusinesscomputerSoftwareBioinformatics
researchProduct

kmcEx: memory-frugal and retrieval-efficient encoding of counted k-mers.

2018

Abstract Motivation K-mers along with their frequency have served as an elementary building block for error correction, repeat detection, multiple sequence alignment, genome assembly, etc., attracting intensive studies in k-mer counting. However, the output of k-mer counters itself is large; very often, it is too large to fit into main memory, leading to highly narrowed usability. Results We introduce a novel idea of encoding k-mers as well as their frequency, achieving good memory saving and retrieval efficiency. Specifically, we propose a Bloom filter-like data structure to encode counted k-mers by coupled-bit arrays—one for k-mer representation and the other for frequency encoding. Exper…

Statistics and ProbabilitySource codeComputer sciencemedia_common.quotation_subject0206 medical engineeringHash function02 engineering and technologyBiochemistry03 medical and health sciencesEncoding (memory)Molecular BiologyTime complexity030304 developmental biologyBlock (data storage)media_common0303 health sciencesSequence Analysis DNAData structureComputer Science ApplicationsComputational MathematicsComputational Theory and MathematicsError detection and correctionAlgorithmSequence Alignment020602 bioinformaticsAlgorithmsSoftwareBioinformatics (Oxford, England)
researchProduct

ArtiFuse—computational validation of fusion gene detection tools without relying on simulated reads

2019

Abstract Motivation Gene fusions are an important class of transcriptional variants that can influence cancer development and can be predicted from RNA sequencing (RNA-seq) data by multiple existing tools. However, the real-world performance of these tools is unclear due to the lack of known positive and negative events, especially with regard to fusion genes in individual samples. Often simulated reads are used, but these cannot account for all technical biases in RNA-seq data generated from real samples. Results Here, we present ArtiFuse, a novel approach that simulates fusion genes by sequence modification to the genomic reference, and therefore, can be applied to any RNA-seq dataset wit…

Statistics and ProbabilitySource codeSequence analysisComputer sciencemedia_common.quotation_subjectValue (computer science)Genomicscomputer.software_genreBiochemistryFusion gene03 medical and health sciences0302 clinical medicineSoftwareMolecular BiologyGene030304 developmental biologymedia_common0303 health sciencesSequence Analysis RNAbusiness.industryHigh-Throughput Nucleotide SequencingRNAGenomicsComputer Science ApplicationsComputational MathematicsComputational Theory and Mathematics030220 oncology & carcinogenesisBenchmark (computing)RNAData miningGene FusionbusinesscomputerSoftwareBioinformatics
researchProduct

RNA-Seq Atlas—a reference database for gene expression profiling in normal tissue by next-generation sequencing

2012

Abstract Motivation: Next-generation sequencing technology enables an entirely new perspective for clinical research and will speed up personalized medicine. In contrast to microarray-based approaches, RNA-Seq analysis provides a much more comprehensive and unbiased view of gene expression. Although the perspective is clear and the long-term success of this new technology obvious, bioinformatics resources making these data easily available especially to the biomedical research community are still evolving. Results: We have generated RNA-Seq Atlas, a web-based repository of RNA-Seq gene expression profiles and query tools. The website offers open and easy access to RNA-Seq gene expression pr…

Statistics and ProbabilitySystems biologyRNA-SeqComputational biologyBiologycomputer.software_genreBiochemistryNeoplasmsGene expressionHumansMicroarray databasesMolecular BiologyGeneOligonucleotide Array Sequence AnalysisInternetSequence Analysis RNAbusiness.industryGene Expression ProfilingHigh-Throughput Nucleotide SequencingComputer Science ApplicationsGene expression profilingComputational MathematicsComputational Theory and MathematicsGene chip analysisData miningPersonalized medicineDatabases Nucleic AcidbusinesscomputerSoftwareBioinformatics
researchProduct

Structure Learning in Nested Effects Models

2007

Nested Effects Models (NEMs) are a class of graphical models introduced to analyze the results of gene perturbation screens. NEMs explore noisy subset relations between the high-dimensional outputs of phenotyping studies, e.g., the effects showing in gene expression profiles or as morphological features of the perturbed cell. In this paper we expand the statistical basis of NEMs in four directions. First, we derive a new formula for the likelihood function of a NEM, which generalizes previous results for binary data. Second, we prove model identifiability under mild assumptions. Third, we show that the new formulation of the likelihood allows efficiency in traversing model space. Fourth, we…

Statistics and ProbabilityTraverseComputer scienceMolecular Networks (q-bio.MN)Genes MHC Class IIPerturbation (astronomy)Genes InsectFeature selectionQuantitative Biology - Quantitative Methods03 medical and health sciences0302 clinical medicineGeneticsAnimalsheterocyclic compoundsQuantitative Biology - Molecular NetworksGraphical modelMolecular BiologyQuantitative Methods (q-bio.QM)Oligonucleotide Array Sequence Analysis030304 developmental biologyLikelihood Functions0303 health sciencesNanoelectromechanical systemsModels StatisticalModels GeneticGene Expression ProfilingGenomicsComputational MathematicsDrosophila melanogasterPhenotypeFOS: Biological sciencesBinary dataIdentifiabilityRNA InterferenceLikelihood functionAlgorithmAlgorithms030217 neurology & neurosurgery
researchProduct

Efficient change point detection in genomic sequences of continuous measurements

2010

Abstract Motivation: Knowing the exact locations of multiple change points in genomic sequences serves several biological needs, for instance when data represent aCGH profiles and it is of interest to identify possibly damaged genes involved in cancer and other diseases. Only a few of the currently available methods deal explicitly with estimation of the number and location of change points, and moreover these methods may be somewhat vulnerable to deviations of model assumptions usually employed. Results: We present a computationally efficient method to obtain estimates of the number and location of the change points. The method is based on a simple transformation of data and it provides re…

Statistics and Probabilitymodel selectionBreast Neoplasmscomputer.software_genreBiochemistryCell LineSimple (abstract algebra)Cell Line TumorHumansComputer Simulationpiecewise constant modelMolecular BiologyMathematicsOligonucleotide Array Sequence AnalysisSupplementary dataComparative Genomic HybridizationModels StatisticalSeries (mathematics)Model selectionGenomicsComputer Science ApplicationsComputational MathematicsR packageTransformation (function)Computational Theory and MathematicsChange pointsChangepointaCGH analysiFemaleData miningSettore SECS-S/01 - StatisticacomputerChange detection
researchProduct

Characterization of poultry egg-white avidins and their potential as a tool in pretargeting cancer treatment.

2003

Chicken avidin and bacterial streptavidin are proteins used in a wide variety of applications in the life sciences due to their strong affinity for biotin. A new and promising use for them is in medical pretargeting cancer treatments. However, their pharmacokinetics and immunological properties are not always optimal, thereby limiting their use in these applications. To search for potentially beneficial new candidates, we screened egg white from four different poultry species for avidin. Avidin proteins, isolated from the duck, goose, ostrich and turkey, showed a similar tetrameric structure, similar glycosylation and stability against both temperature and proteolytic activity of proteinase…

StreptavidinGlycosylationanimal structuresBiotinBiochemistryAntibodiesBirds03 medical and health scienceschemistry.chemical_compound0302 clinical medicineGooseBiotinstomatognathic systemSequence Analysis Proteinbiology.animalNeoplasmsAnimalsMolecular BiologyPhylogeny030304 developmental biologyPretargeting0303 health sciencesbiologyCell Biologyrespiratory systemProteinase KAvidinMolecular biology3. Good healthchemistryBiochemistry030220 oncology & carcinogenesisbiology.proteinAvidinEgg whiteResearch ArticleProtein Binding
researchProduct