A common approach in analyzing gene expression profiles was identifying differential expressed genes that are deemed interesting. The enrichment analysis we demonstrated in Disease enrichment analysis vignette were based on these differential expressed genes. This approach will find genes where the difference is large, but it will not detect a situation where the difference is small, but evidenced in coordinated way in a set of related genes. Gene Set Enrichment Analysis (GSEA)1 directly addresses this limitation. All genes can be used in GSEA; GSEA aggregates the per gene statistics across genes within a gene set, therefore making it possible to detect situations where all genes in a predefined set change in a small but coordinated way. Since it is likely that many relevant phenotypic differences are manifested by small but consistent changes in a set of genes.
Genes are ranked based on their phenotypes. Given a priori defined set of gens S (e.g., genes shareing the same DO category), the goal of GSEA is to determine whether the members of S are randomly distributed throughout the ranked gene list (L) or primarily found at the top or bottom.
There are three key elements of the GSEA method:
We implemented GSEA algorithm proposed by Subramanian1. Alexey Sergushichev implemented an algorithm for fast GSEA analysis in the fgsea2 package.
In DOSE3, user can use GSEA algorithm implemented in DOSE
or fgsea
by specifying the parameter by="DOSE"
or by="fgsea"
. By default, DOSE use fgsea
since it is much more fast.
Leading edge analysis reports Tags
to indicate the percentage of genes contributing to the enrichment score, List
to indicate where in the list the enrichment score is attained and Signal
for enrichment signal strength.
It would also be very interesting to get the core enriched genes that contribute to the enrichment.
DOSE supports leading edge analysis and report core enriched genes in GSEA analysis.
gseDO
fuctionIn the following example, in order to speedup the compilation of this document, only gene sets with size above 120 were tested and only 100 permutations were performed.
library(DOSE)
data(geneList)
y <- gseDO(geneList,
nPerm = 100,
minGSSize = 120,
pvalueCutoff = 0.2,
pAdjustMethod = "BH",
verbose = FALSE)
head(y, 3)
## ID Description setSize enrichmentScore
## DOID:114 DOID:114 heart disease 462 -0.2978223
## DOID:1492 DOID:1492 eye and adnexa disease 459 -0.3105160
## DOID:5614 DOID:5614 eye disease 450 -0.3125247
## NES pvalue p.adjust qvalues rank
## DOID:114 -1.316185 0.01234568 0.07433712 0.03090112 1904
## DOID:1492 -1.368189 0.01265823 0.07433712 0.03090112 1793
## DOID:5614 -1.379267 0.01265823 0.07433712 0.03090112 1768
## leading_edge
## DOID:114 tags=22%, list=15%, signal=19%
## DOID:1492 tags=22%, list=14%, signal=19%
## DOID:5614 tags=22%, list=14%, signal=19%
## core_enrichment
## DOID:114 6649/10268/3567/4882/3910/3371/6548/3082/4153/29119/3791/182/3554/5813/1129/5624/3240/8743/7450/947/78987/1843/4179/7168/948/4314/10272/4881/2628/5021/4018/4256/187/6403/4322/2308/3752/1907/1511/283/3953/7078/2247/2281/10398/5468/10411/10203/1281/4023/83700/11167/7056/3952/126/6310/4313/5502/2944/6444/3075/2273/2099/3480/1471/7079/775/1909/2690/1363/4306/23414/5167/213/5350/5744/11188/2152/2697/185/2952/367/4982/7349/2200/4056/3572/2053/7122/1489/3479/2006/10266/9370/10699/4629/2167/652/1524/7021
## DOID:1492 3371/3082/5914/2878/4153/3791/23247/1543/80184/6750/1958/2098/7450/596/9187/2034/482/948/1490/1280/3931/5737/4314/4881/2261/3426/187/629/6403/7042/6785/7507/2934/5176/4060/1277/7078/5950/2057/727/10516/4311/2247/1295/358/10203/2192/582/10218/57125/3485/585/1675/6310/2202/4313/2944/4254/3075/1501/2099/3480/4653/1195/6387/3305/1471/857/4016/1909/4053/6678/1296/7033/4915/55812/1191/5654/10631/2152/2697/7043/2952/6935/2200/3572/7177/7031/3479/2006/10451/9370/771/3117/125/652/4693/5346/1524
## DOID:5614 3082/5914/2878/4153/3791/23247/1543/80184/6750/1958/2098/7450/596/9187/2034/482/948/1490/1280/3931/5737/4314/4881/2261/3426/187/629/6403/7042/6785/7507/2934/5176/4060/1277/7078/5950/2057/727/10516/4311/2247/1295/358/10203/2192/582/10218/57125/3485/585/1675/6310/2202/4313/2944/4254/3075/1501/2099/3480/4653/6387/3305/1471/857/4016/1909/4053/6678/1296/7033/4915/55812/1191/5654/10631/2152/2697/7043/2952/6935/2200/3572/7177/7031/3479/2006/10451/9370/771/3117/125/652/4693/5346/1524
gseNCG
fuctionncg <- gseNCG(geneList,
nPerm = 100,
minGSSize = 120,
pvalueCutoff = 0.2,
pAdjustMethod = "BH",
verbose = FALSE)
ncg <- setReadable(ncg, 'org.Hs.eg.db')
head(ncg, 3)
## ID Description setSize enrichmentScore NES pvalue
## breast breast breast 133 -0.4869070 -1.943255 0.01351351
## lung lung lung 173 -0.3880662 -1.573163 0.01351351
## lymphoma lymphoma lymphoma 188 0.2999589 1.306967 0.03448276
## p.adjust qvalues rank leading_edge
## breast 0.04054054 0.02133713 2930 tags=33%, list=23%, signal=26%
## lung 0.04054054 0.02133713 2775 tags=31%, list=22%, signal=25%
## lymphoma 0.06896552 0.03629764 2087 tags=21%, list=17%, signal=18%
## core_enrichment
## breast KMT2A/ERBB3/SETD2/ARID1A/GPS2/NCOR1/RB1/MAP2K4/NF1/TP53/PIK3R1/STK11/CDKN1B/PTGFR/APC/CCND1/TRAF5/MAP3K1/ESR1/TBX3/FOXA1/GATA3
## lung SETD2/ATXN3L/LRP1B/BRD3/ARID1A/INHBA/RB1/ADCY1/LYRM9/NF1/CTNNB1/TP53/SATB2/STK11/CTIF/CTNNA3/KDR/COL11A1/FLT3/APC/ADGRL3/FGFR3/NCAM2/DIP2C/APLNR/SLIT2/EPHA3/RUNX1T1/ZMYND10/ZFHX4/GLI3/TNN/PLSCR4/DACH1/ERBB4
## lymphoma DUSP2/EZH2/PRDM1/MYC/ZWILCH/IKZF3/PLCG2/IDH2/HIST1H1C/MAGEC3/CD79B/ETV6/HIST1H1E/HIST1H1B/IRF8/CD28/SLC29A2/DUSP9/TNFAIP3/DNMT3A/SYK/TNF/BCR/HIST1H1D/DSC3/UBE2A/PABPC1
gseDGN
fuctiondgn <- gseDGN(geneList,
nPerm = 100,
minGSSize = 120,
pvalueCutoff = 0.2,
pAdjustMethod = "BH",
verbose = FALSE)
dgn <- setReadable(dgn, 'org.Hs.eg.db')
head(dgn, 3)
## ID Description setSize enrichmentScore
## umls:C0011581 umls:C0011581 Depressive disorder 464 -0.2963136
## umls:C0151744 umls:C0151744 Myocardial Ischemia 418 -0.3013524
## umls:C0023267 umls:C0023267 Fibroid Tumor 282 -0.3222256
## NES pvalue p.adjust qvalues rank
## umls:C0011581 -1.348234 0.01204819 0.1143293 0.07408002 2587
## umls:C0151744 -1.344959 0.01250000 0.1143293 0.07408002 2309
## umls:C0023267 -1.398665 0.01265823 0.1143293 0.07408002 2105
## leading_edge
## umls:C0011581 tags=25%, list=21%, signal=21%
## umls:C0151744 tags=26%, list=18%, signal=22%
## umls:C0023267 tags=27%, list=17%, signal=23%
## core_enrichment
## umls:C0011581 ETS2/HDAC5/RGN/GRIA1/PTGS1/PDE4A/SNCA/ADAMTS2/EHD3/NR5A1/SORCS3/CRY1/ADRB2/FZD1/MYOM2/ADCY1/POU6F1/MAPK3/BICC1/SLC6A4/AHI1/TP53/RNF103/SLC12A2/BDNF/NR3C1/SRSF5/PCLO/GABRA6/WWC1/IL5/GLUL/ELK3/GAD1/RARA/GRM5/KDR/ASAH1/IMPACT/CHRM2/WFS1/TSPAN31/HP/PVALB/HTR1A/BCL2/GPM6A/CYP2A6/DUSP1/NLGN4Y/F2R/CD36/NGFR/NPY2R/DBH/BECN1/CCND1/OXTR/SGCE/SELP/NGF/LPAR1/NRP1/AVPR1B/IFT88/ARSD/FAAH/NEFL/FGF2/CD1C/ABCB1/SRPX/RAPGEF3/CRHBP/HSPA2/LEP/FTO/PER2/ALPK1/GSTM1/DIXDC1/XBP1/ESR1/IGF1R/NTF3/CACNA1C/NR3C2/SLC18A2/NTRK2/SPDEF/RAPGEF4/ALB/NPY1R/F3/AGTR1/TAC1/AR/UCN/FBN1/MAOA/CARTPT/TAT/ADRA2A/MUC1/TGFBR3/TPH1/IGF1/ABAT/MAOB/ADIPOQ/TBC1D9/ADH1B/CRY2/GATA3/TFAP2B
## umls:C0151744 ADRB2/HSPB1/ABCC6/ADD1/PECAM1/MAPK3/GRK5/VEGFC/AMPD1/AES/F7/HSPB2/ENTPD1/ID3/PRKAA2/ATP2B1/SOD3/PRKAB1/AMH/STAT6/RGCC/RXRG/GDF10/SLC9A1/HGF/SERPINA3/MBL2/KDR/EGR1/HSPB6/HBB/STAT5A/EEF1A2/VWF/BCL2/CD34/DUSP1/PRKG1/CD36/CTGF/MMP3/BECN1/NPR1/CCND1/GATM/LPA/EDIL3/RTN1/APLNR/PYGM/SELP/FGF1/NEDD4/ID1/ALDH6A1/FOXO1/SULT1A1/SNAP23/FGF2/DUSP6/ABCB1/PPARG/PDK4/SHH/HSPA2/BHLHE40/LPL/THBD/COL5A2/UGCG/KL/ADH1C/GSTM2/THBS2/PER2/ATXN1/MMP2/TXNIP/KITLG/CFH/ESR1/CXCL12/CIRBP/EDNRA/GHR/SPARC/GPD1L/ENPP1/ALB/F13A1/MEOX2/F3/AGTR1/ZEB1/TNFRSF11B/UCN/DCN/LTC4S/IL6ST/EPHX2/THBS4/IGF1/FXYD1/SFRP4/ELN/RAMP2/ADIPOQ/ADH1B/HMGCS2
## umls:C0023267 CTNNB1/TP53/FZD2/SMAD3/ADAM12/COL4A6/TEK/GLI1/HSD17B7/CYP1A1/BCL6/CDKN1B/EGR1/SALL1/IGFBP7/VWF/BCL2/CD34/CTGF/HPGDS/MMP3/AHR/CCND1/HOXA5/OXTR/FGFR3/FERMT2/NR4A2/FGF1/LAMB1/ADGRV1/XPA/FOXO1/FOS/COL1A1/MME/FGF2/PPARG/TAGLN/SHH/CCNG1/ALDH1A1/FBLN1/COL3A1/IGFBP2/WNT5B/TIE1/THBS2/MMP2/GSTM1/ESR1/IGF1R/CAV1/VCAN/EDNRA/GHR/LTBP2/SLC7A8/PTHLH/NTS/DPT/MST1/ZKSCAN7/F3/GJA1/ANO1/TGFB3/AR/FBN1/COL4A5/XIST/IGF1/WISP2/PGR
cnetplot(ncg, categorySize="pvalue", foldChange=geneList)
enrichMap(y, n=20)
gseaplot(y, geneSetID = y$ID[1], title=y$Description[1])
1. Subramanian, A. et al. Gene set enrichment analysis: A knowledge-based approach for interpreting genome-wide expression profiles. Proceedings of the National Academy of Sciences of the United States of America 102, 15545–15550 (2005).
2. S., A. An algorithm for fast preranked gene set enrichment analysis using cumulative statistic calculation. biorxiv doi:10.1101/060012
3. Yu, G., Wang, L.-G., Yan, G.-R. & He, Q.-Y. DOSE: An r/bioconductor package for disease ontology semantic and enrichment analysis. Bioinformatics 31, 608–609 (2015).