Showing posts with label human network. Show all posts
Showing posts with label human network. Show all posts

Saturday, September 24, 2016

cancer network analysis, 2014 Leiserson et al, Nature genetics



2014 Nature genetics
 Pan-cancer network analysis identifies combinations of rare somatic mutations across pathways and protein complexes

Mark D M Leiserson1,2,14, Fabio Vandin1,2,13,14, Hsin-Ta Wu1,2, Jason R Dobson1–3, Jonathan V Eldridge1, Jacob L Thomas1, Alexandra Papoutsaki1, Younhun Kim1, Beifang Niu4, Michael McLellan4, Michael S Lawrence5, Abel Gonzalez-Perez6, David Tamborero6, Yuwei Cheng7, Gregory A Ryslik8, Nuria Lopez-Bigas6,9, Gad Getz5,10, Li Ding4,11,12 & Benjamin J Raphael1,2



"URLs. HI2012 interactome, http://interactome.dfci.harvard.edu/; HotNet2 pan-cancer analysis website, http://compbio.cs.brown.edu/pancancer/hotnet2/; RNA expression data used for the TCGA pan-cancer data set, https://www.synapse.org/#!Synapse:syn1734155; pan-cancer mutations with additional germline variant filtering, https://www.synapse.org/#!Synapse:syn1729383; HotNet2 software release, http://compbio.cs.brown.edu/software." 

Tuesday, February 3, 2015

admixture mapping

admixture mapping
two ancestral population

https://www.youtube.com/watch?v=syKWlFjzUss

https://www.youtube.com/watch?v=zGha8i8Xzpk


Sunday, December 28, 2014

toread, interaction based discovery of cancer genes

2014 Feb;42(3):e18. doi: 10.1093/nar/gkt1305. Epub 2013 Dec 19.

Interaction-based discovery of functionally important genes in cancers.


http://www.ncbi.nlm.nih.gov/pubmed/24362839

Monday, December 8, 2014

Li14, BMC Medical Genomics, predict disease genes using weighted tissue-specific networks

Li et al. BMC Medical Genomics 2014, 7(Suppl 2):S4 Prediction of disease-related genes based on weighted tissue-specific networks by using DNA methylation

Min Li1, Jiayi Zhang1, Qing Liu1, Jianxin Wang1*, Fang-Xiang Wu1,2*

From IEEE International Conference on Bioinformatics and Biomedicine (BIBM 2013)

 "Considering the fact that the majority of genetic disorders tend to manifest only in a single or a few
tissues, we constructed tissue-specific networks (TSN) by integrating PIN and tissue-specific data." Qin: The observation is true, but it does not mean disease genes are only found in tissues with clinic phenotypes. These tissues show phenotypes only probably because the disease-causing genes play limiting roles in this tissues. 

 "In this paper, we treated known aberrant methylation genes as seed nodes and set initial quantity value with the use of the seed set, which will enhance the importance of seed nodes in network and solve defects of initial PageRank algorithm. The aberrant methylation data related to specific diseases in PubMeth database [44] were used in this paper." 

Li14 used page-rank method to predict disease genes, but seemed to modify it with centrality related measures.

Li14 calculated precision, how?

[52] Culhane AC, Schröder MS, Sultana R, et al: GeneSigDB: a manually curated database and resource for analysis of gene expression signatures. Nucleic acids research 2012, 40(D1):1060-1066.
Qin: Using a known-database to evaluate prediction. The precision of 60-80%. Did the authors address over-fitting problems? 

Tissue specific networks were constructed by removing unexpressed nodes.

References on removal method to generate tissue-specific networks
Waldman YY, Tuller T, Shlomi T, et al: Translation efficiency in humans: tissue specificity global optimization and differences between developmental stages. Nucleic Acids Research 2010, 38(9):2964-2974.

Bossi A, Lehner B: Tissue specificity and the human protein interaction network. Molecular Systems Biology 2009, 5(1):260.

Lopes TJ, Schaefer M, Shoemaker J, et al: Tissue-specific subnetworks and characteristics of publicly available human protein interaction databases. Bioinformatics 2011, 27(17):2414-2421.

Probabilistic model of human PPI, Rhodes, 2005, Nat Biotech

Probabilistic model of the human protein-protein interaction network
Daniel R Rhodes1,2,7, Scott A Tomlins2,7, Sooryanarayana Varambally2,7, Vasudeva Mahavisno2, Terrence Barrette2, Shanker Kalyana-Sundaram2, Debashis Ghosh3, Akhilesh Pandey6 & Arul M Chinnaiyan1,2,4,5
2005, Nature Biotechnology

Rhodes05 proposed a probabilistic model integrating PPI, protein domain, and expression data, and annotation data for 40K PPI.

Rhodes05 used otholog interaction from Sce, worm, and fruit fly to predict interactions in human proteins.

A semi-naive Bayes model was used to predict human protein interactions. A maximum likelihood ratio was used as the final score. Co-expression data from oncomine were used for the prediction. (My critique: oncomine may contain abnormal co-expressions).

My  summary: The probabilistic network model of Rhodes05 is used for prediction, not to describe the stochastic nature of protein-interactions or conditional protein-interactions depending on cell types or states, such as cell cycles, growth stages. 


Human Interaction Map  www.himap.org



Rhodes05 used both gold positive control and gold negative control. The gold negative control is an good idea which I can add into my future network analysis project.


Tuesday, September 16, 2014

Network resources, human, (in progress)

*** GeneSigDB (used by Li14BMC, should contain disease genes). This site provide download for R analysis.   http://compbio.dfci.harvard.edu/genesigdb/

List of human disease genes in gene Card
http://www.genecards.org/cgi-bin/listdiseasecards.pl?type=full
(not sure whether this includes haploid type association)

Genome Research Genome-wide map of regulatory interactions in the human genome,
http://genome.cshlp.org/content/early/2014/09/15/gr.176586.114.abstract

Farmington study
http://videocast.nih.gov/Summary.asp?File=18760&bhcp=1

Network analysis of GWAS data, 2013 Current Opinion in Genetics and Development
Mark DM Leiserson1,2, Jonathan V Eldridge1,2, Sohini Ramachandran2,3 and Benjamin J Raphael1,2 


MIF Parsers


The human protein interaction network data can be found from http://www.hprd.org/download
Reference: Peri S, Navarro JD, Amanchy R, Kristiansen TZ, Jonnalagadda CK,
Surendranath V, Niranjan V, Muthusamy B, Gandhi TKB, GronborgM,
Ibarrola N, Deshpande N, Shanker K, Shivashankar HN, Rashmi BP,
Ramya MA, Zhao ZX, Chandrika KN, Padma N, Harsha HC et al
(2003) Development of human protein reference database as an initial
platform for approaching systems biology in humans. Genome Res 13:
2363–237

Gerstein lab's human networks
  human multinet 
http://homes.gersteinlab.org/Khurana-PLoSCompBio-2013/
http://www.ploscompbiol.org/article/info%3Adoi%2F10.1371%2Fjournal.pcbi.1002886#s4

http://encodenets.gersteinlab.org/


dbCline
UChicago

peptieAtlas
http://www.nature.com/embor/journal/v9/n5/pdf/embor200856.pdf

ensemble

human twin aging expression
http://genomebiology.com/2013/14/7/R75?utm_campaign=10_12_13_genomebiol_Article_Mailing_Reg&utm_content=7387379393&utm_medium=BMCemail&utm_source=Emailvision


http://hongqinlab.blogspot.com/2013/10/mutation-tolerance-in-human-genes-data.html

http://hongqinlab.blogspot.com/2013/03/candidate-rojects-for-r-based-data.html

expression, aging, human, kidny,
http://www.plosbiology.org/article/info%3Adoi%2F10.1371%2Fjournal.pbio.0020427



9. Keshava Prasad TS, Goel R, Kandasamy K, Keerthikumar S, Kumar S, Mathivanan S, Telikicherla D, Raju R, Shafreen B, Venugopal A et al.: Human Protein Reference Database- 2009 update. [Internet]. Nucleic Acids Res 2009, 37: D767-D772. 
10. Stark C, Breitkreutz B-J, Reguly T, Boucher L, Breitkreutz A, Tyers M: BioGRID: a general repository for interaction datasets. [Internet]. Nucleic Acids Res 2006, 34:D535-D539. 
11. Franceschini A, Szklarczyk D, Frankild S, Kuhn M, Simonovic M, Roth A, Lin J, Minguez P, Bork P, von Mering C et al.: STRING v9.1: proteinprotein interaction networks, with increased coverage and integration. [Internet]. Nucleic Acids Res 2013, 41:D808-D815. 
12. Razick S, Magklaras G, Donaldson IM: iRefIndex: a consolidated protein interaction database with provenance. [Internet]. BMC Bioinformatics 2008, 9:405. 
13. Croft D, O’Kelly G, Wu G, Haw R, Gillespie M, Matthews L, Caudy M, Garapati P, Gopinath G, Jassal B et al.: Reactome: a database of reactions, pathways and biological processes. [Internet]. Nucleic Acids Research 2011, 39:D691-D697. 
14. Ewing RM et al.: Large-scale mapping of human proteinprotein interactions by mass spectrometry. Molecular Systems Biology 2007, 3:89. 
15. Hutchins JRa et al.: Systematic analysis of human protein complexes identifies chromosome segregation proteins. Science 2010, 328:593-599. 
16. Rual J-F et al.: Towards a proteome-scale map of the human proteinprotein interaction network. Nature 2005, 437:1173-1178. 
17. Stelzl U et al.: A human proteinprotein interaction network: a resource for annotating the proteome. Cell 2005, 122:957-968. 
18. Yu H et al.: Next-generation sequencing to generate interactome datasets. Nat Methods 2011, 8:478-480. 608 Genetics of system biology Current Opinion 


From Gilman 2011 Neuron, netbag on autism
Downloaded Information
The data described in previous sections was downloaded from the following public
resources:
GeneOntology annotations from NCBI (01/2009 – ftp://ftp.ncbi.nlm.nih.gov/gene/)
Pathways and enzyme codes from Kyoto Encyclopedia of Genes and Genomes (KEGG)
database (01/2009 – ftp://ftp.genome.jp/pub/kegg/genes/organisms/hsa/)
Domains from InterPro database (01/2009 – ftp://ftp.ebi.ac.uk/pub/databases/interpro/)
Tissue indicators from the TiGER database (09/2009 – http://bioinfo.wilmer.jhu.edu/tiger/)
Protein-protein Interactions:
o BIND with protein complexes (08/2009 – http://bond.unleashedinformatics.com/)
o BioGRID interactions (08/2009 – http://www.thebiogrid.org/downloads.php)
o DIP interactions (10/2009 – http://dip.doe-mbi.ucla.edu/dip/)
o HPRD interactions (10/2009 – http://www.hprd.org/)
o InNetDB interactions (05/2009 – http://hanlab.genetics.ac.cn/sys/intnetdb)
o IntAct (10/2009 – http://www.ebi.ac.uk/intact/)
o BiGG metabolic interactions (04/2009 – http://gcrg.ucsd.edu/Downloads )
o MINT (05/2009 – http://mint.bio.uniroma2.it/mint/)

o MIPS (05/2009 – http://mips.gsf.de/proj/ppi/)


gene expression between human and mouse, PNAS, mike snyder lab, Standford
http://www.pnas.org/content/111/48/17224.full

Tuesday, August 26, 2014

GWAS meta analysis

Gilman et al, Neuron, 2011. p898-907. NetBag on Autism
NetBag is a greedy approaches. The clustering methods started with one or two genes in CNV as ‘seeds’.

Gilman11 generated a weighted background human gene network for their study.

Gilman11 compared the cluster raw pvalue, called local pvalue to the p-values from random networks. The adjusted p-value is called global p-value.



-----------------------------------------------------------------------------------------------------------------------

AIS13 categorize pathway association methods into canonical and de nov pathway methods.

For de novo pathway discovery, integer linear program (ILP) is used in Leiserson , Blokh, Plos Comput Biol. Simultaneous identification of multiple driver pathway in cancer.

Steiner tree problem where one seeks the lowest cost pathway that connect the associated genes. See Liu et al, BMC Sys Biol 2012, Gene, pathway and nework frameworks to identify epistatic interactions of single nucleotide polymorphisms derived from GWAS data.

-----------------------------------------------------------------------------------------------------------------------
LERR13 review the method on protein-protein and protein-DNA networks to identify 'causal' genetic variant. (Their 'causal' definition is a narrowly defined one).

LERR13 argues that GWAS mostly find SNP that are LD with the actual 'causal' gene. This problem is of less concern to CNV analysis. One solution is to use network to 'rank' genes in the same haplotype known to be associate to the phenotype of interest or similar phenotypes. (This method is in the spirit of our recent CNV paper).

The green square represents the 'known' 'causal' gene. So, this is largely a traversal-measures based method.

LERR13 argues that networks contribute to 'missing heritability'.

LERR13 seems to suggest that protein-DNA networks are better suited for expression QTL (eQTL).

LERR13 shows that OMIM is the source of 'causal' gene information for most network based GWAS (table 1). Only one paper use GeneCards as an alternative source.

LERR13 cited several pathway enrichment analysis of GWAS. It argues that interaction are treated equally in these enrichment analysis. (This can be cited in our CNV replies). The authors then show several method use weighted networks to identify network modules using iterative 'seed and extend' method. (For comparison, our CNV paper did not use seed explicitly, and avoid some 'prejudice').

LERR13 also discussed subnetwork modules with mutation hotspots in cancer genomes.

----------------------------------------------------------------------------------------------------------

BGTF12: GWAS use meta-analysis of multiple data sets to reduce false positives and increase statistical power.

A major concern of GWAS meta analysis is the heterogeneity in the data sets, such LD difference among data sets, chip differences.  However, Lin and Zeng 2010 (Gene epidemil) show heterogeneity is not a significant factor using simulation studies.

Combination across data sets is the frequentist approach, cumulative studies is the Bayesian approach.


In R, GWAS meta-analysis package: Metrafor, rmeta, and CATMAP.

BGTF12 argues that GWAS data should be 'cleaned' and imputed before meta-analysis.

Reference:
[AIS13] Atias, Istrail, Sharan 2013, Current Opinion in Genetics and Development. Pathway-based analysis of genomic variation data.

[LERR13] Leiserson, Eldrige, Ramachandran, Raphael, 2013, Current Opinion in Genetics and Development. Network analysis of GWAS data.

[BGTF12], begum, ghosh, tseng, feigold, 2012 NAS, comprehensive literature review and statistical consideration for GWAS meta analysis

Tuesday, February 18, 2014

clustering based prediction of human disease genes


http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3030143/pdf/ijbsv07p0061.pdf

multiNet

Things to do:
 various clustering methods
 compare I=1.4 and 2.0.
 visualization, say by cytoscape.

What is Disease ID, and where to obtain this information?