Aug 13, 2011

Efficient counting of k-mers in DNA sequences using a bloom filter

BMC Bioinformatics
Páll Melsted, Jonathan K Pritchard

Abstract

Counting k-mers (substrings of length k in DNA sequence data) is an essential component of many methods in bioinformatics, including for genome and transcriptome assembly, for metagenomic sequencing, and for error correction of sequence reads. Although simple in principle, counting k-mers in large modern sequence data sets can easily overwhelm the memory capacity of standard computers. In current data sets, a large fraction-often more than 50%-of the storage capacity may be spent on storing k-mers that contain sequencing errors and which are typically observed only a single time in the data. These singleton k-mers are uninformative for many algorithms without some kind of error correction. We present a new method that identifies all the k-mers that occur more than once in a DNA sequence data set. Our method does this using a Bloom filter, a probabilistic data structure that stores all the observed k-mers implicitly in memory with greatly reduced memory requirements. We then make a second sweep through the data to provide exact counts of all nonunique k-mers. For example data sets, we report up to 50% savings in memory usage compared to current software, with modest costs in computational speed. This approach may reduce memory r...Continue Reading

  • References16
  • Citations45

Citations

Mentioned in this Paper

Chromosomes, Human, Pair 21
Severe Acute Respiratory Syndrome
Jellyfish
HapMap
Genome
Genome Assembly Sequence
Bio-Informatics
Computer Programs and Programming
Pandas, Giant
Sequence Determinations, DNA

Trending Feeds

COVID-19

Coronaviruses encompass a large family of viruses that cause the common cold as well as more serious diseases, such as the ongoing outbreak of coronavirus disease 2019 (COVID-19; formally known as 2019-nCoV). Coronaviruses can spread from animals to humans; symptoms include fever, cough, shortness of breath, and breathing difficulties; in more severe cases, infection can lead to death. This feed covers recent research on COVID-19.

Bone Marrow Neoplasms

Bone Marrow Neoplasms are cancers that occur in the bone marrow. Discover the latest research on Bone Marrow Neoplasms here.

IGA Glomerulonephritis

IgA glomerulonephritis is a chronic form of glomerulonephritis characterized by deposits of predominantly Iimmunoglobin A in the mesangial area. Discover the latest research on IgA glomerulonephritis here.

Cryogenic Electron Microscopy

Cryogenic electron microscopy (Cryo-EM) allows the determination of biological macromolecules and their assemblies at a near-atomic resolution. Here is the latest research.

STING Receptor Agonists

Stimulator of IFN genes (STING) are a group of transmembrane proteins that are involved in the induction of type I interferon that is important in the innate immune response. The stimulation of STING has been an active area of research in the treatment of cancer and infectious diseases. Here is the latest research on STING receptor agonists.

LRRK2 & Immunity During Infection

Mutations in the LRRK2 gene are a risk-factor for developing Parkinson’s disease. However, LRRK2 has been shown to function as a central regulator of vesicular trafficking, infection, immunity, and inflammation. Here is the latest research on the role of this kinase on immunity during infection.

Antiphospholipid Syndrome

Antiphospholipid syndrome or antiphospholipid antibody syndrome (APS or APLS), is an autoimmune, hypercoagulable state caused by the presence of antibodies directed against phospholipids.

Meningococcal Myelitis

Meningococcal myelitis is characterized by inflammation and myelin damage to the meninges and spinal cord. Discover the latest research on meningococcal myelitis here.

Alzheimer's Disease: MS4A

Variants within membrane-spanning 4-domains subfamily A (MS4A) gene cluster have recently been implicated in Alzheimer's disease by recent genome-wide association studies. Here is the latest research.