Uniform approximation is more appropriate for Wilcoxon Rank-Sum Test in gene set analysis.

PloS One
Zhide FangXiangqin Cui

Abstract

Gene set analysis is widely used to facilitate biological interpretations in the analyses of differential expression from high throughput profiling data. Wilcoxon Rank-Sum (WRS) test is one of the commonly used methods in gene set enrichment analysis. It compares the ranks of genes in a gene set against those of genes outside the gene set. This method is easy to implement and it eliminates the dichotomization of genes into significant and non-significant in a competitive hypothesis testing. Due to the large number of genes being examined, it is impractical to calculate the exact null distribution for the WRS test. Therefore, the normal distribution is commonly used as an approximation. However, as we demonstrate in this paper, the normal approximation is problematic when a gene set with relative small number of genes is tested against the large number of genes in the complementary set. In this situation, a uniform approximation is substantially more powerful, more accurate, and less intensive in computation. We demonstrate the advantage of the uniform approximations in Gene Ontology (GO) term analysis using simulations and real data sets.

References

Feb 14, 2004·Bioinformatics·Tim Beissbarth, Terence P Speed
Feb 17, 2007·Bioinformatics·Jelle J Goeman, Peter Bühlmann
Nov 9, 2007·BMC Bioinformatics·Qi LiuYutaka Yasui
Jun 29, 2010·Nature Medicine·Hail KimMichael S German
Oct 20, 2010·BMC Genomics·Daniel M GattiFred A Wright
Apr 1, 1946·Journal of Economic Entomology·F WILCOXON
Apr 19, 2011·Briefings in Bioinformatics·Zhide Fang, Xiangqin Cui
Sep 9, 2011·Briefings in Bioinformatics·Jui-Hung HungCharles DeLisi

❮ Previous
Next ❯

Citations

Apr 22, 2015·Computational and Mathematical Methods in Medicine·Congwei SunZheming Yuan
Jan 10, 2018·Interdisciplinary Sciences, Computational Life Sciences·José A Castellanos-GarzónJuan M Corchado

❮ Previous
Next ❯

Datasets Mentioned

BETA
GSE11045
GSE7869
GSE21860

Methods Mentioned

BETA
RNA-seq

Software Mentioned

R
SAS
glmLRT
Affymetrix
Bioconductor package affy
STATA
R package
Illumina genome analyzer
edgeR
R geneSetTest

Related Concepts

Trending Feeds

COVID-19

Coronaviruses encompass a large family of viruses that cause the common cold as well as more serious diseases, such as the ongoing outbreak of coronavirus disease 2019 (COVID-19; formally known as 2019-nCoV). Coronaviruses can spread from animals to humans; symptoms include fever, cough, shortness of breath, and breathing difficulties; in more severe cases, infection can lead to death. This feed covers recent research on COVID-19.

Blastomycosis

Blastomycosis fungal infections spread through inhaling Blastomyces dermatitidis spores. Discover the latest research on blastomycosis fungal infections here.

Nuclear Pore Complex in ALS/FTD

Alterations in nucleocytoplasmic transport, controlled by the nuclear pore complex, may be involved in the pathomechanism underlying multiple neurodegenerative diseases including Amyotrophic Lateral Sclerosis and Frontotemporal Dementia. Here is the latest research on the nuclear pore complex in ALS and FTD.

Applications of Molecular Barcoding

The concept of molecular barcoding is that each original DNA or RNA molecule is attached to a unique sequence barcode. Sequence reads having different barcodes represent different original molecules, while sequence reads having the same barcode are results of PCR duplication from one original molecule. Discover the latest research on molecular barcoding here.

Chronic Fatigue Syndrome

Chronic fatigue syndrome is a disease characterized by unexplained disabling fatigue; the pathology of which is incompletely understood. Discover the latest research on chronic fatigue syndrome here.

Evolution of Pluripotency

Pluripotency refers to the ability of a cell to develop into three primary germ cell layers of the embryo. This feed focuses on the mechanisms that underlie the evolution of pluripotency. Here is the latest research.

Position Effect Variegation

Position Effect Variagation occurs when a gene is inactivated due to its positioning near heterochromatic regions within a chromosome. Discover the latest research on Position Effect Variagation here.

STING Receptor Agonists

Stimulator of IFN genes (STING) are a group of transmembrane proteins that are involved in the induction of type I interferon that is important in the innate immune response. The stimulation of STING has been an active area of research in the treatment of cancer and infectious diseases. Here is the latest research on STING receptor agonists.

Microbicide

Microbicides are products that can be applied to vaginal or rectal mucosal surfaces with the goal of preventing, or at least significantly reducing, the transmission of sexually transmitted infections. Here is the latest research on microbicides.