Mar 17, 2016

Computational Performance Assessment of k-mer Counting Algorithms

Journal of Computational Biology : a Journal of Computational Molecular Cell Biology
Nelson PérezNelson Vera

Abstract

This article is about the assessment of several tools for k-mer counting, with the purpose to create a reference framework for bioinformatics researchers to identify computational requirements, parallelizing, advantages, disadvantages, and bottlenecks of each of the algorithms proposed in the tools. The k-mer counters evaluated in this article were BFCounter, DSK, Jellyfish, KAnalyze, KHMer, KMC2, MSPKmerCounter, Tallymer, and Turtle. Measured parameters were the following: RAM occupied space, processing time, parallelization, and read and write disk access. A dataset consisting of 36,504,800 reads was used corresponding to the 14th human chromosome. The assessment was performed for two k-mer lengths: 31 and 55. Obtained results were the following: pure Bloom filter-based tools and disk-partitioning techniques showed a lesser RAM use. The tools that took less execution time were the ones that used disk-partitioning techniques. The techniques that made the major parallelization were the ones that used disk partitioning, hash tables with lock-free approach, or multiple hash tables.

  • References7
  • Citations1

Mentioned in this Paper

Jellyfish
Chromosomes, Human, Pair 14
Research Personnel
Chromosomes, Human
Bio-Informatics
Anatomical Space Structure
Evaluation
Computer Programs and Programming
Cyanea capillata preparation
Filter - Medical Device

Trending Feeds

COVID-19

Coronaviruses encompass a large family of viruses that cause the common cold as well as more serious diseases, such as the ongoing outbreak of coronavirus disease 2019 (COVID-19; formally known as 2019-nCoV). Coronaviruses can spread from animals to humans; symptoms include fever, cough, shortness of breath, and breathing difficulties; in more severe cases, infection can lead to death. This feed covers recent research on COVID-19.

Coronavirus Protein Structures

Deciphering and comparing the proteins of different coronaviruses forms a basis for understanding SARS-CoV-2 evolution and virus-receptor interactions. This feed follows studies analyzing the structures of coronavirus proteins, thereby revealing potential drug target sites.

DDX3X Syndrome

DDX3X syndrome is caused by a spontaneous mutation at conception that primarily affects girls due to its location on the X-chromosome. DDX3X syndrome has been linked to intellectual disabilities, seizures, autism, low muscle tone, brain abnormalities, and slower physical developments. Here is the latest research.

ALS: Stress Granules

Amyotrophic Lateral Sclerosis (ALS) is a neurodegenerative disease characterized by cytoplasmic protein aggregates within motor neurons. TDP-43 is an ALS-linked protein that is known to regulate splicing and storage of specific mRNAs into stress granules, which have been implicated in formation of ALS protein aggregates. Here is the latest research.

Fusion Oncoproteins in Childhood Cancers

This feed explores the function of fusion oncoproteins in specific childhood cancers, including those from racial/ethnic minority and underserved groups, and to provide preclinical assessment of potential therapeutics and how fusion oncoproteins influence gene expression to perturb normal cellular programs to block lineage differentiation and development

Applications of Molecular Barcoding

The concept of molecular barcoding is that each original DNA or RNA molecule is attached to a unique sequence barcode. Sequence reads having different barcodes represent different original molecules, while sequence reads having the same barcode are results of PCR duplication from one original molecule. Discover the latest research on molecular barcoding here.

Regulation of Vocal-Motor Plasticity

Dopaminergic projections to the basal ganglia and nucleus accumbens shape the learning and plasticity of motivated behaviors across species including the regulation of vocal-motor plasticity and performance in songbirds. Discover the latest research on the regulation of vocal-motor plasticity here.

Mitotic-exit networks with cytokinesis

Cytokinesis is the highly regulated process that physically separates daughter and mother cells in late mitosis. The mitotic-exit network (MEN), the signalling pathway that drives mitotic exit, directly regulates cytokinesis. Discover the latest research on mitotic-exit networks with cytokinesis here.

DNA Replication Origin

DNA replication is initiated as specific gene sequences, called origins, that function to start DNA replication. Pre-replication complexes are assembled at these origins during the G1 phase of the cell cycle. These sequences allow for targeted activation or deactivation of replication. Discover the latest research on DNA replication origins here.

Related Papers

Western Journal of Nursing Research
Vicki S Conn, Tamara G Coon Sells
Tijdschrift voor psychiatrie
P N van Harten
© 2020 Meta ULC. All rights reserved