EDBT 2026 Demo / reviewers in the wild / expert
Jianrong Li
dblp:63/5716 · also Jianrong John Li
· DBLP profile ↗
43ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 36 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-scale Dual-Attention Gating Fusion for Thymoma Segmentation in CT ImagesabstractThymoma CT images exhibit significant scale variations and blurred boundaries, posing a challenge to automatic segmentation. The proposed model is a multi-scale dual attention and gated skip connection network (MSDACG-Net), which suppresses redundant features in skip connections through context gating (CoT Gate) and enhances context modeling capabilities at high resolution using a multi-scale dual attention aggregation (MSDA) module. The effectiveness of the method was validated through experimentation on a thymoma CT dataset that had been self-built. A comparison of the proposed model with a strong baseline reveals that MSDACG-Net improves Dice by 1.01%, reduces HD95 by 24%, and improves Recall by 0.59%, demonstrating superior overall accuracy, boundary approximation ability, and detection robustness. Moreover, cross-modal experiments on the publicly available polysegmentation dataset Kvasir-SEG demonstrate the proposed method’s capacity for effective generalisation. Chenfei Wu, Jianrong Li, Chuanlei Zhang, Wenchao Xia, Yaoyu Zhou |
ICIC | 2 |
| 2025 | EGAP-YOLO: An Efficient Crack Detection Model Based on YOLO Architecture
Jianrong Li, Haifeng Fan, Di Sun 0001, Chuanlei Zhang, Yinglun Dong |
ICIC (2) | 1 |
| 2025 | LSGNSF: A Graph-Based Time Series Anomaly Detection Algorithm
Chuanlei Zhang, Yinglun Dong, Jianrong Li, Haifeng Fan, Di Sun 0001 |
ICIC (20) | 4 |
| 2024 | AES Improvement Algorithm Based on the Chaotic System in IIOT
Jianrong Li, Pengyu Han, Huiying Sun, Ting Ke, Wei Chen 0036, Chuanlei Zhang |
ICIC (8) | 1 |
| 2024 | Anomaly Detection of Transmission Line Large Metal Based on EGFPN-YOLO and UAVs
Gongcheng Shi, Jianrong Li, Yicong Li 0009, Di Sun 0001, Chuanlei Zhang |
ICIC (5) | 2 |
| 2024 | DGAP-YOLO: A Crack Detection Method Based on UAV Images and YOLO
Yunyi Li, Jianrong Li, Di Sun 0001, Chuanlei Zhang |
ICIC (11) | 5 |
| 2024 | A Multi-dimensional Camera Image Stitching Method Under Large Parallax Conditions
Chuanlei Zhang, Tianxiang Cheng, Jianrong Li, Haifeng Fan, Zhanjun Si |
ICIC (7) | 4 |
| 2024 | A general maximal margin hyper-sphere SVM for multi-class classification
Ting Ke, Xuechun Ge, Feifei Yin, Yaozong Zheng, Chuanlei Zhang, Jianrong Li |
Expert Syst. Appl. | 7 |
| 2023 | Time Series Prediction of 5G Network Data Based on Improved EEMD-BiLSTM Prediction Model
Jianrong Li, Gongcheng Shi, Chuanlei Zhang |
ICIC (5) | 1 |
| 2023 | Maximal margin hyper-sphere SVM for binary pattern classification
Ting Ke, Yangyang Liao, Mengyan Wu, Xuechun Ge, Xinyi Huang 0011, Chuanlei Zhang, Jianrong Li |
Eng. Appl. Artif. Intell. | 7 |
| 2022 | A Multi-step Attention and Multi-level Structure Network for Multimodal Sentiment Analysis
Chuanlei Zhang, Ting Ke, Jianrong Li |
NLPCC (1) | 6 |
| 2019 | Multi-modal Recognition of Mental Workload Using Empirical Mode Decomposition and Semi-Supervised LearningabstractReal-time monitoring and analysis of human operator's mental workload (MWL) is crucial for development of adaptive/intelligent human-machine cooperative systems in various safety/mission-critical application fields. Although data-driven machine learning (ML) approach has shown promise in MWL recognition, it is usually difficult to acquire sufficient labeled data to train the ML model. This paper proposes semi-supervised extreme learning machines (SS-ELM) for MWL pattern classification using solely a small number of labeled data. The experimental data analysis results are presented to show the effectiveness of the proposed SS-ELM paradigm for the 3-class MWL classification. Jianhua Zhang 0004, Jianrong Li |
CW | 2 |
| 2019 | CMEP: a database for circulating microRNA expression profilingabstractMOTIVATION: In recent years, several experimental studies have revealed that the microRNAs (miRNAs) in serum, plasma, exosome and whole blood are dysregulated in various types of diseases, indicating that the circulating miRNAs may serve as potential noninvasive biomarkers for disease diagnosis and prognosis. However, no database has been constructed to integrate the large-scale circulating miRNA profiles, explore the functional pathways involved and predict the potential biomarkers using feature selection between the disease conditions. Although there have been several studies attempting to generate a circulating miRNA database, they have not yet integrated the large-scale circulating miRNA profiles or provided the biomarker-selection function using machine learning methods. RESULTS: To fill this gap, we constructed the Circulating MicroRNA Expression Profiling (CMEP) database for integrating, analyzing and visualizing the large-scale expression profiles of phenotype-specific circulating miRNAs. The CMEP database contains massive datasets that were manually curated from NCBI GEO and the exRNA Atlas, including 66 datasets, 228 subsets and 10 419 samples. The CMEP provides the differential expression circulating miRNAs analysis and the KEGG functional pathway enrichment analysis. Furthermore, to provide the function of noninvasive biomarker discovery, we implemented several feature-selection methods, including ridge regression, lasso regression, support vector machine and random forests. Finally, we implemented a user-friendly web interface to improve the user experience and to visualize the data and results of CMEP. AVAILABILITY AND IMPLEMENTATION: CMEP is accessible at http://syslab5.nchu.edu.tw/CMEP. Jianrong Li, Chun-Yip Tong, Tsai-Jung Sung, Ting-Yu Kang, Xianghong Jasmine Zhou, Chun-Chi Liu |
Bioinform. | 1 |
| 2018 | Systems of Phenome-Exposome Associations Unveiled by Mining Practice-Based Evidence with Environmental Exposures
Jung-wei Fan, Samir Rachid Zaim, Walter W. Piegorsch, Jianrong Li, Colleen Kenost, Yves A. Lussier |
AMIA | 4 |
| 2017 | Mental Workload Classification Based on Semi-Supervised Extreme Learning Machine
Jianrong Li, Jianhua Zhang 0004 |
ICANN (2) | 1 |
| 2017 | A genome-by-environment interaction classifier for precision medicine: personal transcriptome response to rhinovirus identifies children prone to asthma exacerbationsabstractOBJECTIVE: To introduce a disease prognosis framework enabled by a robust classification scheme derived from patient-specific transcriptomic response to stimulation. MATERIALS AND METHODS: Within an illustrative case study to predict asthma exacerbation, we designed a stimulation assay that reveals individualized transcriptomic response to human rhinovirus. Gene expression from peripheral blood mononuclear cells was quantified from 23 pediatric asthmatic patients and stimulated in vitro with human rhinovirus. Responses were obtained via the single-subject gene set testing methodology "N-of-1-pathways." The classifier was trained on a related independent training dataset (n = 19). Novel visualizations of personal transcriptomic responses are provided. RESULTS: Of the 23 pediatric asthmatic patients, 12 experienced recurrent exacerbations. Our classifier, using individualized responses and trained on an independent dataset, obtained 74% accuracy (area under the receiver operating curve of 71%; 2-sided P = .039). Conventional classifiers using messenger RNA (mRNA) expression within the viral-exposed samples were unsuccessful (all patients predicted to have recurrent exacerbations; accuracy of 52%). DISCUSSION: Prognosis based on single time point, static mRNA expression alone neglects the importance of dynamic genome-by-environment interplay in phenotypic presentation. Individualized transcriptomic response quantified at the pathway (gene sets) level reveals interpretable signals related to clinical outcomes. CONCLUSION: The proposed framework provides an innovative approach to precision medicine. We show that quantifying personal pathway-level transcriptomic response to a disease-relevant environmental challenge predicts disease progression. This genome-by-environment interaction assay offers a noninvasive opportunity to translate omics data to clinical practice by improving the ability to predict disease exacerbation and increasing the potential to produce more effective treatment decisions. Vincent Gardeux, Joanne Berghout, Ikbel Achour, A. Grant Schissler, Qike Li, Colleen Kenost, Jianrong Li, Yuan Shang, Anthony Bosco, Donald Saner, Marilyn J. Halonen, Daniel Jackson 0003, Haiquan Li, Fernando D. Martinez, Yves A. Lussier |
J. Am. Medical Informatics Assoc. | 7 |
| 2015 | Metrics and tools for consistent cohort discovery and financial analyses post-transition to ICD-10-CMabstractIn the United States, International Classification of Disease Clinical Modification (ICD-9-CM, the ninth revision) diagnosis codes are commonly used to identify patient cohorts and to conduct financial analyses related to disease. In October 2015, the healthcare system of the United States will transition to ICD-10-CM (the tenth revision) diagnosis codes. One challenge posed to clinical researchers and other analysts is conducting diagnosis-related queries across datasets containing both coding schemes. Further, healthcare administrators will manage growth, trends, and strategic planning with these dually-coded datasets. The majority of the ICD-9-CM to ICD-10-CM translations are complex and nonreciprocal, creating convoluted representations and meanings. Similarly, mapping back from ICD-10-CM to ICD-9-CM is equally complex, yet different from mapping forward, as relationships are likewise nonreciprocal. Indeed, 10 of the 21 top clinical categories are complex as 78% of their diagnosis codes are labeled as "convoluted" by our analyses. Analysis and research related to external causes of morbidity, injury, and poisoning will face the greatest challenges due to 41 745 (90%) convolutions and a decrease in the number of codes. We created a web portal tool and translation tables to list all ICD-9-CM diagnosis codes related to the specific input of ICD-10-CM diagnosis codes and their level of complexity: "identity" (reciprocal), "class-to-subclass," "subclass-to-class," "convoluted," or "no mapping." These tools provide guidance on ambiguous and complex translations to reveal where reports or analyses may be challenging to impossible.Web portal: http://www.lussierlab.org/transition-to-ICD9CM/Tables annotated with levels of translation complexity: http://www.lussierlab.org/publications/ICD10to9. Andrew D. Boyd, Jianrong Li, Colleen Kenost, Binoy Joese, Young Min Yang, Olympia A. Kalagidis, Ilir Zenku, Donald Saner, Neil Bahroos, Yves A. Lussier |
J. Am. Medical Informatics Assoc. | 2 |
| 2015 | Challenges and remediation for Patient Safety Indicators in the transition to ICD-10-CMabstractReporting of hospital adverse events relies on Patient Safety Indicators (PSIs) using International Classification of Diseases, Ninth Edition, Clinical Modification (ICD-9-CM) codes. The US transition to ICD-10-CM in 2015 could result in erroneous comparisons of PSIs. Using the General Equivalent Mappings (GEMs), we compared the accuracy of ICD-9-CM coded PSIs against recommended ICD-10-CM codes from the Centers for Medicaid/Medicare Services (CMS). We further predict their impact in a cohort of 38,644 patients (1,446,581 visits and 399 hospitals). We compared the predicted results to the published PSI related ICD-10-CM diagnosis codes. We provide the first report of substantial hospital safety reporting errors with five direct comparisons from the 23 types of PSIs (transfusion and anesthesia related PSIs). One PSI was excluded from the comparison between code sets due to reorganization, while 15 additional PSIs were inaccurate to a lesser degree due to the complexity of the coding translation. The ICD-10-CM translations proposed by CMS pose impending risks for (1) comparing safety incidents, (2) inflating the number of PSIs, and (3) increasing the variability of calculations attributable to the abundance of coding system translations. Ethical organizations addressing 'data-, process-, and system-focused' improvements could be penalized using the new ICD-10-CM Agency for Healthcare Research and Quality PSIs because of apparent increases in PSIs bearing the same PSI identifier and label, yet calculated differently. Here we investigate which PSIs would reliably transition between ICD-9-CM and ICD-10-CM, and those at risk of under-reporting and over-reporting adverse events while the frequency of these adverse events remain unchanged. Andrew D. Boyd, Young Min Yang, Jianrong Li, Colleen Kenost, Mike D. Burton, Bryan Becker, Yves A. Lussier |
J. Am. Medical Informatics Assoc. | 3 |
| 2015 | Towards a PBMC "virogram assay" for precision medicine: Concordance between ex vivo and in vivo viral infection transcriptomesabstractBACKGROUND: Understanding individual patient host-response to viruses is key to designing optimal personalized therapy. Unsurprisingly, in vivo human experimentation to understand individualized dynamic response of the transcriptome to viruses are rarely studied because of the obvious limitations stemming from ethical considerations of the clinical risk. OBJECTIVE: In this rhinovirus study, we first hypothesized that ex vivo human cells response to virus can serve as a proxy for otherwise controversial in vivo human experimentation. We further hypothesized that the N-of-1-pathways framework, previously validated in cancer, can be effective in understanding the more subtle individual transcriptomic response to viral infection. METHOD: N-of-1-pathways computes a significance score for a given list of gene sets at the patient level, using merely the 'omics profiles of two paired samples as input. We extracted the peripheral blood mononuclear cells (PBMC) of four human subjects, aliquoted in two paired samples, one subjected to ex vivo rhinovirus infection. Their dysregulated genes and pathways were then compared to those of 9 human subjects prior and after intranasal inoculation in vivo with rhinovirus. Additionally, we developed the Similarity Venn Diagram, a novel visualization method that goes beyond conventional overlap to show the similarity between two sets of qualitative measures. RESULTS: We evaluated the individual N-of-1-pathways results using two established cohort-based methods: GSEA and enrichment of differentially expressed genes. Similarity Venn Diagrams and individual patient ROC curves illustrate and quantify that the in vivo dysregulation is recapitulated ex vivo both at the gene and pathway level (p-values⩽0.004). CONCLUSION: We established the first evidence that an interpretable dynamic transcriptome metric, conducted as an ex vivo assays for a single subject, has the potential to predict individualized response to infectious disease without the clinical risks otherwise associated to in vivo challenges. These results serve as a foundational work for personalized "virograms". Vincent Gardeux, Anthony Bosco, Jianrong Li, Marilyn J. Halonen, Daniel Jackson 0003, Fernando D. Martinez, Yves A. Lussier |
J. Biomed. Informatics | 3 |
| 2015 | eQTL networks unveil enriched mRNA master integrators downstream of complex disease-associated SNPsabstractThe causal and interplay mechanisms of Single Nucleotide Polymorphisms (SNPs) associated with complex diseases (complex disease SNPs) investigated in genome-wide association studies (GWAS) at the transcriptional level (mRNA) are poorly understood despite recent advancements such as discoveries reported in the Encyclopedia of DNA Elements (ENCODE) and Genotype-Tissue Expression (GTex). Protein interaction network analyses have successfully improved our understanding of both single gene diseases (Mendelian diseases) and complex diseases. Whether the mRNAs downstream of complex disease genes are central or peripheral in the genetic information flow relating DNA to mRNA remains unclear and may be disease-specific. Using expression Quantitative Trait Loci (eQTL) that provide DNA to mRNA associations and network centrality metrics, we hypothesize that we can unveil the systems properties of information flow between SNPs and the transcriptomes of complex diseases. We compare different conditions such as naïve SNP assignments and stringent linkage disequilibrium (LD) free assignments for transcripts to remove confounders from LD. Additionally, we compare the results from eQTL networks between lymphoblastoid cell lines and liver tissue. Empirical permutation resampling (p<0.001) and theoretic Mann-Whitney U test (p<10(-30)) statistics indicate that mRNAs corresponding to complex disease SNPs via eQTL associations are likely to be regulated by a larger number of SNPs than expected. We name this novel property mRNA hubness in eQTL networks, and further term mRNAs with high hubness as master integrators. mRNA master integrators receive and coordinate the perturbation signals from large numbers of polymorphisms and respond to the personal genetic architecture integratively. This genetic signal integration contrasts with the mechanism underlying some Mendelian diseases, where a genetic polymorphism affecting a single protein hub produces a divergent signal that affects a large number of downstream proteins. Indeed, we verify that this property is independent of the hubness in protein networks for which these mRNAs are transcribed. Our findings provide novel insights into the pleiotropy of mRNAs targeted by complex disease polymorphisms and the architecture of the information flow between the genetic polymorphisms and transcriptomes of complex diseases. Haiquan Li, Nima Pouladi, Ikbel Achour, Vincent Gardeux, Jianrong Li, Qike Li, Hao Helen Zhang 0001, Fernando D. Martinez, Joe G. N. 'Skip' Garcia, Yves A. Lussier |
J. Biomed. Informatics | 5 |
| 2014 | COPD Hospitalization Risk Increased with Distinct Patterns of Multiple Systems Comorbidities Unveiled by Network Modeling
Young Ji Lee, Andrew D. Boyd, Jianrong Li, Vincent Gardeux, Colleen Kenost, Donald Saner, Haiquan Li, Ivo L. Abraham, Jerry A. Krishnan, Yves A. Lussier |
AMIA | 3 |
| 2014 | PatientNarr: Towards generating patient-centric summaries of hospital staysabstractBarbara Di Eugenio, Andrew Boyd, Camillo Lugaresi, Abhinaya Balasubramanian, Gail Keenan, Mike Burton, Tamara Goncalves Rezende Macieira, Jianrong Li, Yves Lussier, Yves Lussier. Proceedings of the 8th International Natural Language Generation Conference (INLG). 2014. Barbara Di Eugenio, Andrew D. Boyd, Camillo Lugaresi, Abhinaya Balasubramanian, Gail M. Keenan, Mike D. Burton, Tamara Goncalves Rezende Macieira, Jianrong Li, Yves A. Lussier |
INLG | 8 |
| 2014 | 'N-of-1-pathways' unveils personal deregulated mechanisms from a single pair of RNA-Seq samples: towards precision medicineabstractBACKGROUND: The emergence of precision medicine allowed the incorporation of individual molecular data into patient care. Indeed, DNA sequencing predicts somatic mutations in individual patients. However, these genetic features overlook dynamic epigenetic and phenotypic response to therapy. Meanwhile, accurate personal transcriptome interpretation remains an unmet challenge. Further, N-of-1 (single-subject) efficacy trials are increasingly pursued, but are underpowered for molecular marker discovery. METHOD: 'N-of-1-pathways' is a global framework relying on three principles: (i) the statistical universe is a single patient; (ii) significance is derived from geneset/biomodules powered by paired samples from the same patient; and (iii) similarity between genesets/biomodules assesses commonality and differences, within-study and cross-studies. Thus, patient gene-level profiles are transformed into deregulated pathways. From RNA-Seq of 55 lung adenocarcinoma patients, N-of-1-pathways predicts the deregulated pathways of each patient. RESULTS: Cross-patient N-of-1-pathways obtains comparable results with conventional genesets enrichment analysis (GSEA) and differentially expressed gene (DEG) enrichment, validated in three external evaluations. Moreover, heatmap and star plots highlight both individual and shared mechanisms ranging from molecular to organ-systems levels (eg, DNA repair, signaling, immune response). Patients were ranked based on the similarity of their deregulated mechanisms to those of an independent gold standard, generating unsupervised clusters of diametric extreme survival phenotypes (p=0.03). CONCLUSIONS: The N-of-1-pathways framework provides a robust statistical and relevant biological interpretation of individual disease-free survival that is often overlooked in conventional cross-patient studies. It enables mechanism-level classifiers with smaller cohorts as well as N-of-1 studies. SOFTWARE: http://lussierlab.org/publications/N-of-1-pathways. Vincent Gardeux, Ikbel Achour, Jianrong Li, Mark Maienschein-Cline, Haiquan Li, Lorenzo L. Pesce, Gurunadh Parinandi, Neil Bahroos, Robert Winn, Ian T. Foster, Joe G. N. 'Skip' Garcia, Yves A. Lussier |
J. Am. Medical Informatics Assoc. | 3 |
| 2013 | HospSum: Integrating physician discharge notes with coded nursing care data to generate patient-centric summaries
Barbara Di Eugenio, Camillo Lugaresi, Gail M. Keenan, Yves A. Lussier, Jianrong Li, Mike D. Burton, Carol Friedman, Andrew D. Boyd |
AMIA | 5 |
| 2013 | Research and applications: The discriminatory cost of ICD-10-CM transition between clinical specialties: metrics, case study, and mitigating toolsabstractOBJECTIVE: Applying the science of networks to quantify the discriminatory impact of the ICD-9-CM to ICD-10-CM transition between clinical specialties. MATERIALS AND METHODS: Datasets were the Center for Medicaid and Medicare Services ICD-9-CM to ICD-10-CM mapping files, general equivalence mappings, and statewide Medicaid emergency department billing. Diagnoses were represented as nodes and their mappings as directional relationships. The complex network was synthesized as an aggregate of simpler motifs and tabulation per clinical specialty. RESULTS: We identified five mapping motif categories: identity, class-to-subclass, subclass-to-class, convoluted, and no mapping. Convoluted mappings indicate that multiple ICD-9-CM and ICD-10-CM codes share complex, entangled, and non-reciprocal mappings. The proportions of convoluted diagnoses mappings (36% overall) range from 5% (hematology) to 60% (obstetrics and injuries). In a case study of 24 008 patient visits in 217 emergency departments, 27% of the costs are associated with convoluted diagnoses, with 'abdominal pain' and 'gastroenteritis' accounting for approximately 3.5%. DISCUSSION: Previous qualitative studies report that administrators and clinicians are likely to be challenged in understanding and managing their practice because of the ICD-10-CM transition. We substantiate the complexity of this transition with a thorough quantitative summary per clinical specialty, a case study, and the tools to apply this methodology easily to any clinical practice in the form of a web portal and analytic tables. CONCLUSIONS: Post-transition, successful management of frequent diseases with convoluted mapping network patterns is critical. The http://lussierlab.org/transition-to-ICD10CM web portal provides insight in linking onerous diseases to the ICD-10 transition. Andrew D. Boyd, Jianrong Li, Mike D. Burton, Michael Jonen, Vincent Gardeux, Ikbel Achour, Roger Q. Luo, Ilir Zenku, Neil Bahroos, Stephen B. Brown, Terry L. Vanden Hoek, Yves A. Lussier |
J. Am. Medical Informatics Assoc. | 2 |
| 2013 | Research and applications: Network models of genome-wide association studies uncover the topological centrality of protein interactions in complex diseasesabstractBACKGROUND: While genome-wide association studies (GWAS) of complex traits have revealed thousands of reproducible genetic associations to date, these loci collectively confer very little of the heritability of their respective diseases and, in general, have contributed little to our understanding the underlying disease biology. Physical protein interactions have been utilized to increase our understanding of human Mendelian disease loci but have yet to be fully exploited for complex traits. METHODS: We hypothesized that protein interaction modeling of GWAS findings could highlight important disease-associated loci and unveil the role of their network topology in the genetic architecture of diseases with complex inheritance. RESULTS: Network modeling of proteins associated with the intragenic single nucleotide polymorphisms of the National Human Genome Research Institute catalog of complex trait GWAS revealed that complex trait associated loci are more likely to be hub and bottleneck genes in available, albeit incomplete, networks (OR=1.59, Fisher's exact test p < 2.24 × 10(-12)). Network modeling also prioritized novel type 2 diabetes (T2D) genetic variations from the Finland-USA Investigation of Non-Insulin-Dependent Diabetes Mellitus Genetics and the Wellcome Trust GWAS data, and demonstrated the enrichment of hubs and bottlenecks in prioritized T2D GWAS genes. The potential biological relevance of the T2D hub and bottleneck genes was revealed by their increased number of first degree protein interactions with known T2D genes according to several independent sources (p<0.01, probability of being first interactors of known T2D genes). CONCLUSION: Virtually all common diseases are complex human traits, and thus the topological centrality in protein networks of complex trait genes has implications in genetics, personal genomics, and therapy. Younghee Lee, Haiquan Li, Jianrong Li, Ellen Rebman, Ikbel Achour, Kelly Regan-Fendt, Eric R. Gamazon, James L. Chen, Xinan Yang, Nancy J. Cox, Yves A. Lussier |
J. Am. Medical Informatics Assoc. | 3 |
| 2012 | Towards Mechanism Classifiers: Expression-anchored Gene Ontology Signature Predicts Clinical Outcome in Lung Adenocarcinoma Patients
Xinan Yang, Haiquan Li, Kelly Regan-Fendt, Jianrong Li, H. Rosie Xing, Yves A. Lussier |
AMIA | 4 |
| 2012 | Complex-disease networks of trait-associated single-nucleotide polymorphisms (SNPs) unveiled by information theoryabstractOBJECTIVE: Thousands of complex-disease single-nucleotide polymorphisms (SNPs) have been discovered in genome-wide association studies (GWAS). However, these intragenic SNPs have not been collectively mined to unveil the genetic architecture between complex clinical traits. The authors hypothesize that biological annotations of host genes of trait-associated SNPs may reveal the biomolecular modularity across complex-disease traits and offer insights for drug repositioning. METHODS: Trait-to-polymorphism (SNPs) associations confirmed in GWAS were used. A novel method to quantify trait-trait similarity anchored in Gene Ontology annotations of human proteins and information theory was developed. The results were then validated with the shortest paths of physical protein interactions between biologically similar traits. RESULTS: A network was constructed consisting of 280 significant intertrait similarities among 177 disease traits, which covered 1438 well-validated disease-associated SNPs. Thirty-nine percent of intertrait connections were confirmed by curators, and the following additional studies demonstrated the validity of a proportion of the remainder. On a phenotypic trait level, higher Gene Ontology similarity between proteins correlated with smaller 'shortest distance' in protein interaction networks of complexly inherited diseases (Spearman p<2.2×10(-16)). Further, 'cancer traits' were similar to one another, as were 'metabolic syndrome traits' (Fisher's exact test p=0.001 and 3.5×10(-7), respectively). CONCLUSION: An imputed disease network by information-anchored functional similarity from GWAS trait-associated SNPs is reported. It is also demonstrated that small shortest paths of protein interactions correlate with complex-disease function. Taken together, these findings provide the framework for investigating drug targets with unbiased functional biomolecular networks rather than worn-out single-gene and subjective canonical pathway approaches. Haiquan Li, Younghee Lee, James L. Chen, Ellen Rebman, Jianrong Li, Yves A. Lussier |
J. Am. Medical Informatics Assoc. | 5 |
| 2012 | Translating Mendelian and complex inheritance of Alzheimer's disease genes for predicting unique personal genome variantsabstractOBJECTIVE: Although trait-associated genes identified as complex versus single-gene inheritance differ substantially in odds ratio, the authors nonetheless posit that their mechanistic concordance can reveal fundamental properties of the genetic architecture, allowing the automated interpretation of unique polymorphisms within a personal genome. MATERIALS AND METHODS: An analytical method, SPADE-gen, spanning three biological scales was developed to demonstrate the mechanistic concordance between Mendelian and complex inheritance of Alzheimer's disease (AD) genes: biological functions (BP), protein interaction modeling, and protein domain implicated in the disease-associated polymorphism. RESULTS: Among Gene Ontology (GO) biological processes (BP) enriched at a false detection rate <5% in 15 AD genes of Mendelian inheritance (Online Mendelian Inheritance in Man) and independently in those of complex inheritance (25 host genes of intragenic AD single-nucleotide polymorphisms confirmed in genome-wide association studies), 16 overlapped (empirical p=0.007) and 45 were similar (empirical p<0.009; information theory). SPAN network modeling extended the canonical pathway of AD (KEGG) with 26 new protein interactions (empirical p<0.0001). DISCUSSION: The study prioritized new AD-associated biological mechanisms and focused the analysis on previously unreported interactions associated with the biological processes of polymorphisms that affect specific protein domains within characterized AD genes and their direct interactors using (1) concordant GO-BP and (2) domain interactions within STRING protein-protein interactions corresponding to the genomic location of the AD polymorphism (eg, EPHA1, APOE, and CD2AP). CONCLUSION: These results are in line with unique-event polymorphism theory, indicating how disease-associated polymorphisms of Mendelian or complex inheritance relate genetically to those observed as 'unique personal variants'. They also provide insight for identifying novel targets, for repositioning drugs, and for personal therapeutics. Kelly Regan-Fendt, Kanix Wang, Emily Doughty, Haiquan Li, Jianrong Li, Younghee Lee, Maricel G. Kann, Yves A. Lussier |
J. Am. Medical Informatics Assoc. | 5 |
| 2012 | Single Sample Expression-Anchored Mechanisms Predict Survival in Head and Neck CancerabstractGene expression signatures that are predictive of therapeutic response or prognosis are increasingly useful in clinical care; however, mechanistic (and intuitive) interpretation of expression arrays remains an unmet challenge. Additionally, there is surprisingly little gene overlap among distinct clinically validated expression signatures. These "causality challenges" hinder the adoption of signatures as compared to functionally well-characterized single gene biomarkers. To increase the utility of multi-gene signatures in survival studies, we developed a novel approach to generate "personal mechanism signatures" of molecular pathways and functions from gene expression arrays. FAIME, the Functional Analysis of Individual Microarray Expression, computes mechanism scores using rank-weighted gene expression of an individual sample. By comparing head and neck squamous cell carcinoma (HNSCC) samples with non-tumor control tissues, the precision and recall of deregulated FAIME-derived mechanisms of pathways and molecular functions are comparable to those produced by conventional cohort-wide methods (e.g. GSEA). The overlap of "Oncogenic FAIME Features of HNSCC" (statistically significant and differentially regulated FAIME-derived genesets representing GO functions or KEGG pathways derived from HNSCC tissue) among three distinct HNSCC datasets (pathways:46%, p<0.001) is more significant than the gene overlap (genes:4%). These Oncogenic FAIME Features of HNSCC can accurately discriminate tumors from control tissues in two additional HNSCC datasets (n = 35 and 91, F-accuracy = 100% and 97%, empirical p<0.001, area under the receiver operating characteristic curves = 99% and 92%), and stratify recurrence-free survival in patients from two independent studies (p = 0.0018 and p = 0.032, log-rank). Previous approaches depending on group assignment of individual samples before selecting features or learning a classifier are limited by design to discrete-class prediction. In contrast, FAIME calculates mechanism profiles for individual patients without requiring group assignment in validation sets. FAIME is more amenable for clinical deployment since it translates the gene-level measurements of each given sample into pathways and molecular function profiles that can be applied to analyze continuous phenotypes in clinical outcome studies (e.g. survival time, tumor volume). Xinan Yang, Kelly Regan-Fendt, Qingbei Zhang, Jianrong Li, Tanguy Y. Seiwert, Ezra E. W. Cohen, H. Rosie Xing, Yves A. Lussier |
PLoS Comput. Biol. | 5 |
| 2011 | GO-Module: functional synthesis and improved interpretation of Gene Ontology patternsabstractUNLABELLED: GO-Module is a web-accessible synthesis and visualization tool developed for end-user biologists to greatly simplify the interpretation of prioritized Gene Ontology (GO) terms. GO-Module radically reduces the complexity of raw GO results into compact biomodules in two distinct ways, by (i) constructing biomodules from significant GO terms based on hierarchical knowledge, and (ii) refining the GO terms in each biomodule to contain only true positive results. Altogether, the features (biomodules) of GO-Module outputs are better organized and on average four times smaller than the input GO terms list (P = 0.0005, n = 16). AVAILABILITY: http://lussierlab.org/GO-Module. Xinan Yang, Jianrong Li, Younghee Lee, Yves A. Lussier |
Bioinform. | 2 |
| 2011 | Protein-network modeling of prostate cancer gene signatures reveals essential pathways in disease recurrenceabstractOBJECTIVE: Uncovering the dominant molecular deregulation among the multitude of pathways implicated in aggressive prostate cancer is essential to intelligently developing targeted therapies. Paradoxically, published prostate cancer gene expression signatures of poor prognosis share little overlap and thus do not reveal shared mechanisms. The authors hypothesize that, by analyzing gene signatures with quantitative models of protein-protein interactions, key pathways will be elucidated and shown to be shared. DESIGN: The authors statistically prioritized common interactors between established cancer genes and genes from each prostate cancer signature of poor prognosis independently via a previously validated single protein analysis of network (SPAN) methodology. Additionally, they computationally identified pathways among the aggregated interactors across signatures and validated them using a similarity metric and patient survival. MEASUREMENT: Using an information-theoretic metric, the authors assessed the mechanistic similarity of the interactor signature. Its prognostic ability was assessed in an independent cohort of 198 patients with high-Gleason prostate cancer using Kaplan-Meier analysis. RESULTS: Of the 13 prostate cancer signatures that were evaluated, eight interacted significantly with established cancer genes (false discovery rate <5%) and generated a 42-gene interactor signature that showed the highest mechanistic similarity (p<0.0001). Via parameter-free unsupervised classification, the interactor signature dichotomized the independent prostate cancer cohort with a significant survival difference (p=0.009). Interpretation of the network not only recapitulated phosphatidylinositol-3 kinase/NF-κB signaling, but also highlighted less well established relevant pathways such as the Janus kinase 2 cascade. CONCLUSIONS: SPAN methodology provides a robust means of abstracting disparate prostate cancer gene expression signatures into clinically useful, prioritized pathways as well as useful mechanistic pathways. James L. Chen, Jianrong Li, Walter M. Stadler, Yves A. Lussier |
J. Am. Medical Informatics Assoc. | 2 |
| 2010 | Protein interaction network underpins concordant prognosis among heterogeneous breast cancer signaturesabstractCharacterizing the biomolecular systems' properties underpinning prognosis signatures derived from gene expression profiles remains a key clinical and biological challenge. In breast cancer, while different "poor-prognosis" sets of genes have predicted patient survival outcome equally well in independent cohorts, these prognostic signatures have surprisingly little genetic overlap. We examine 10 such published expression-based signatures that are predictors or distinct breast cancer phenotypes, uncover their mechanistic interconnectivity through a protein-protein interaction network, and introduce a novel cross-"gene expression signature" analysis method using (i) domain knowledge to constrain multiple comparisons in a mechanistically relevant single-gene network interactions and (ii) scale-free permutation re-sampling to statistically control for hubness (SPAN - Single Protein Analysis of Network with constant node degree per protein). At adjusted p-values<5%, 54-genes thus identified have a significantly greater connectivity than those through meticulous permutation re-sampling of the context-constrained network. More importantly, eight of 10 genetically non-overlapping signatures are connected through well-established mechanisms of breast cancer oncogenesis and progression. Gene Ontology enrichment studies demonstrate common markers of cell cycle regulation. Kaplan-Meier analysis of three independent historical gene expression sets confirms this network-signature's inherent ability to identify "poor outcome" in ER(+) patients without the requirement of machine learning. We provide a novel demonstration that genetically distinct prognosis signatures, developed from independent clinical datasets, occupy overlapping prognostic space of breast cancer via shared mechanisms that are mediated by genetically different yet mechanistically comparable interactions among proteins of differentially expressed genes in the signatures. This is the first study employing a networks' approach to aggregate established gene expression signatures in order to develop a phenotype/pathway-based cancer roadmap with the potential for (i) novel drug development applications and for (ii) facilitating the clinical deployment of prognostic gene signatures with improved mechanistic understanding of biological processes and functions associated with gene expression changes. http://www.lussierlab.org/publication/networksignature/. James L. Chen, Lee T. Sam, Younghee Lee, Jianrong Li, Yang Liu 0023, H. Rosie Xing, Yves A. Lussier |
J. Biomed. Informatics | 5 |
| 2010 | Kinase inhibition-related adverse events predicted from in vitro kinome and clinical trial data
Xinan Yang, Matthew G. Crowson, Jianrong Li, Michael L. Maitland, Yves A. Lussier |
J. Biomed. Informatics | 4 |
| 2010 | Network Modeling Identifies Molecular Functions Targeted by miR-204 to Suppress Head and Neck Tumor MetastasisabstractDue to the large number of putative microRNA gene targets predicted by sequence-alignment databases and the relative low accuracy of such predictions which are conducted independently of biological context by design, systematic experimental identification and validation of every functional microRNA target is currently challenging. Consequently, biological studies have yet to identify, on a genome scale, key regulatory networks perturbed by altered microRNA functions in the context of cancer. In this report, we demonstrate for the first time how phenotypic knowledge of inheritable cancer traits and of risk factor loci can be utilized jointly with gene expression analysis to efficiently prioritize deregulated microRNAs for biological characterization. Using this approach we characterize miR-204 as a tumor suppressor microRNA and uncover previously unknown connections between microRNA regulation, network topology, and expression dynamics. Specifically, we validate 18 gene targets of miR-204 that show elevated mRNA expression and are enriched in biological processes associated with tumor progression in squamous cell carcinoma of the head and neck (HNSCC). We further demonstrate the enrichment of bottleneckness, a key molecular network topology, among miR-204 gene targets. Restoration of miR-204 function in HNSCC cell lines inhibits the expression of its functionally related gene targets, leads to the reduced adhesion, migration and invasion in vitro and attenuates experimental lung metastasis in vivo. As importantly, our investigation also provides experimental evidence linking the function of microRNAs that are located in the cancer-associated genomic regions (CAGRs) to the observed predisposition to human cancers. Specifically, we show miR-204 may serve as a tumor suppressor gene at the 9q21.1-22.3 CAGR locus, a well established risk factor locus in head and neck cancers for which tumor suppressor genes have not been identified. This new strategy that integrates expression profiling, genetics and novel computational biology approaches provides for improved efficiency in characterization and modeling of microRNA functions in cancer as compared to the state of art and is applicable to the investigation of microRNA functions in other biological processes and diseases. Younghee Lee, Xinan Yang, Hanli Fan, Qingbei Zhang, Youngfei Wu, Jianrong Li, Rifat Hasina, Mark W. Lingen, Mark Gerstein, Ralph R. Weichselbaum, H. Rosie Xing, Yves A. Lussier |
PLoS Comput. Biol. | 7 |
| 2009 | Robust methods for accurate diagnosis using pan-microbiological oligonucleotide microarraysabstractBACKGROUND: To address the limitations of traditional virus and pathogen detection methodologies in clinical diagnosis, scientists have developed high-throughput oligonucleotide microarrays to rapidly identify infectious agents. However, objectively identifying pathogens from the complex hybridization patterns of these massively multiplexed arrays remains challenging. METHODS: In this study, we conceived an automated method based on the hypergeometric distribution for identifying pathogens in multiplexed arrays and compared it to five other methods. We evaluated these metrics: 1) accurate prediction, whether the top ranked prediction(s) match the real virus(es); 2) four accuracy scores. RESULTS: Though accurate prediction and high specificity and sensitivity can be achieved with several methods, the method based on hypergeometric distribution provides a significant advantage in term of positive predicting value with two to sixty folds the positive predicting values of other methods. CONCLUSION: The proposed multi-specie array analysis based on the hypergeometric distribution addresses shortcomings of previous methods by enhancing signals of positively hybridized probes. Yang Liu 0023, Lee T. Sam, Jianrong Li, Yves A. Lussier |
BMC Bioinform. | 3 |
| 2009 | PhenoGO: an integrated resource for the multiscale mining of clinical and biological dataabstractThe evolving complexity of genome-scale experiments has increasingly centralized the role of a highly computable, accurate, and comprehensive resource spanning multiple biological scales and viewpoints. To provide a resource to meet this need, we have significantly extended the PhenoGO database with gene-disease specific annotations and included an additional ten species. This a computationally-derived resource is primarily intended to provide phenotypic context (cell type, tissue, organ, and disease) for mining existing associations between gene products and GO terms specified in the Gene Ontology Databases Automated natural language processing (BioMedLEE) and computational ontology (PhenOS) methods were used to derive these relationships from the literature, expanding the database with information from ten additional species to include over 600,000 phenotypic contexts spanning eleven species from five GO annotation databases. A comprehensive evaluation evaluating the mappings (n = 300) found precision (positive predictive value) at 85%, and recall (sensitivity) at 76%. Phenotypes are encoded in general purpose ontologies such as Cell Ontology, the Unified Medical Language System, and in specialized ontologies such as the Mouse Anatomy and the Mammalian Phenotype Ontology. A web portal has also been developed, allowing for advanced filtering and querying of the database as well as download of the entire dataset http://www.phenogo.org. Lee T. Sam, Eneida A. Mendonça, Jianrong Li, Judith A. Blake, Carol Friedman, Yves A. Lussier |
BMC Bioinform. | 3 |
| 2007 | Information-Theoretic Classification of SNOMED Improves the Organization of Context-Sensitive Excerpts from Cochrane Reviews
Lee T. Sam, Tara Borlawsky, Ying Tao, Jianrong Li, Carol Friedman, Barry Smith 0001, Yves A. Lussier |
AMIA | 4 |
| 2007 | Evaluation of high-throughput functional categorization of human disease genesabstractBACKGROUND: Biological data that are well-organized by an ontology, such as Gene Ontology, enables high-throughput availability of the semantic web. It can also be used to facilitate high throughput classification of biomedical information. However, to our knowledge, no evaluation has been published on automating classifications of human diseases genes using Gene Ontology. In this study, we evaluate automated classifications of well-defined human disease genes using their Gene Ontology annotations and compared them to a gold standard. This gold standard was independently conceived by Valle's research group, and contains 923 human disease genes organized in 14 categories of protein function. RESULTS: Two automated methods were applied to investigate the classification of human disease genes into independently pre-defined categories of protein function. One method used the structure of Gene Ontology by pre-selecting 74 Gene Ontology terms assigned to 11 protein function categories. The second method was based on the similarity of human disease genes clustered according to the information-theoretic distance of their Gene Ontology annotations. Compared to the categorization of human disease genes found in the gold standard, our automated methods can achieve an overall 56% and 47% precision with 62% and 71% recall respectively. However, approximately 15% of the studied human disease genes remain without GO annotations. CONCLUSION: Automated methods can recapitulate a significant portion of classification of the human disease genes. The method using information-theoretic distance performs slightly better on the precision with some loss in recall. For some protein function categories, such as 'hormone' and 'transcription factor', the automated methods perform particularly well, achieving precision and recall levels above 75%. In summary, this study demonstrates that for semantic webs, methods to automatically classify or analyze a majority of human disease genes require significant progress in both the Gene Ontology annotations and particularly in the utilization of these annotations. James L. Chen, Yang Liu 0023, Lee T. Sam, Jianrong Li, Yves A. Lussier |
BMC Bioinform. | 4 |
| 2007 | Partitioning knowledge bases between advanced notification and clinical decision support systems
Yves A. Lussier, Rose Williams, Jianrong Li, Srikant Jalan, Tara Borlawsky, Edie Stern, Inderpal Kohli |
Decis. Support Syst. | 3 |
| 2006 | An Integrative Genomic Approach to Uncover Molecular Mechanisms of Prokaryotic TraitsabstractWith mounting availability of genomic and phenotypic databases, data integration and mining become increasingly challenging. While efforts have been put forward to analyze prokaryotic phenotypes, current computational technologies either lack high throughput capacity for genomic scale analysis, or are limited in their capability to integrate and mine data across different scales of biology. Consequently, simultaneous analysis of associations among genomes, phenotypes, and gene functions is prohibited. Here, we developed a high throughput computational approach, and demonstrated for the first time the feasibility of integrating large quantities of prokaryotic phenotypes along with genomic datasets for mining across multiple scales of biology (protein domains, pathways, molecular functions, and cellular processes). Applying this method over 59 fully sequenced prokaryotic species, we identified genetic basis and molecular mechanisms underlying the phenotypes in bacteria. We identified 3,711 significant correlations between 1,499 distinct Pfam and 63 phenotypes, with 2,650 correlations and 1,061 anti-correlations. Manual evaluation of a random sample of these significant correlations showed a minimal precision of 30% (95% confidence interval: 20%-42%; n = 50). We stratified the most significant 478 predictions and subjected 100 to manual evaluation, of which 60 were corroborated in the literature. We furthermore unveiled 10 significant correlations between phenotypes and KEGG pathways, eight of which were corroborated in the evaluation, and 309 significant correlations between phenotypes and 166 GO concepts evaluated using a random sample (minimal precision = 72%; 95% confidence interval: 60%-80%; n = 50). Additionally, we conducted a novel large-scale phenomic visualization analysis to provide insight into the modular nature of common molecular mechanisms spanning multiple biological scales and reused by related phenotypes (metaphenotypes). We propose that this method elucidates which classes of molecular mechanisms are associated with phenotypes or metaphenotypes and holds promise in facilitating a computable systems biology approach to genomic and biomedical research. Yang Liu 0023, Jianrong Li, Lee T. Sam, Chern-Sing Goh, Mark Gerstein, Yves A. Lussier |
PLoS Comput. Biol. | 2 |
| 2005 | Partitioning Knowledge Bases between Advanced Notification and Clinical Decision Support Systems
Tara Borlawsky, Jianrong Li, Srikant Jalan, Edie Stern, Rose Williams, Yves A. Lussier |
AMIA | 2 |
| 2003 | Guideline Interaction: a study of interactions among drug-disease contraindication rules
Te-Hui Kuo, Eneida A. Mendonça, Jianrong Li, Yves A. Lussier |
AMIA | 3 |