VLDB 2026 Research / reviewers in the wild / expert
Simon M. Lin
dblp:35/5195
· DBLP profile ↗
30ranked-venue papers
1as first author
2since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 24 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Security and privacy · 1Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
11 papers |
Bioinformatics and computational biology · 60% Medical and health informatics · 21% Computational science and engineering · 19% | |
| Artificial intelligence
2 papers |
Trustworthy machine learning · 52% Knowledge representation and reasoning · 26% Information extraction and text analysis · 22% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Network and information security
1 paper |
Privacy and data protection · 100% |
Topics — the 24 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › question answering
FAQ retrieval |
0.5 | 1 | 2021 | COUGH: A Challenge Dataset and Models for COVID-19 FAQ Retrieval · EMNLP (1) 2021 |
Information retrieval
question answering |
0.5 | 1 | 2021 | COUGH: A Challenge Dataset and Models for COVID-19 FAQ Retrieval · EMNLP (1) 2021 |
Machine learning › Trustworthy machine learning
interpretability |
0.4 | 1 | 2020 | Rationalizing Medical Relation Prediction from Corpus-level Statistics · ACL 2020 |
Machine learning › Trustworthy machine learning › interpretability › rationalization
rationale extraction |
0.4 | 1 | 2020 | Rationalizing Medical Relation Prediction from Corpus-level Statistics · ACL 2020 |
Bioinformatics and computational biology › network bioinformatics
biomedical network analysis |
0.4 | 1 | 2020 | Graph embedding on biomedical networks: methods, applications and evaluations · Bioinform. 2020 |
Computational science and engineering › graph learning
link prediction |
0.4 | 1 | 2020 | Graph embedding on biomedical networks: methods, applications and evaluations · Bioinform. 2020 |
Natural language and speech › Information extraction and text analysis › lexical resources › lexical resource construction
synonym discovery |
0.4 | 1 | 2019 | SurfCon: Synonym Discovery on Privacy-Aware Clinical Data · KDD 2019 |
Medical and health informatics › electronic health records
electronic health record analysis |
0.2 | 1 | 2015 | Application of clinical text data for phenome-wide association studies (PheWASs) · Bioinform. 2015 |
Bioinformatics and computational biology › statistical genetics › post-GWAS analysis
phenome-wide association study |
0.2 | 1 | 2015 | Application of clinical text data for phenome-wide association studies (PheWASs) · Bioinform. 2015 |
Privacy and data protection
privacy-preserving data analysis |
0.2 | 1 | 2014 | Privacy in Pharmacogenetics: An End-to-End Case Study of Personalized Warfarin Dosing · USENIX Security Symposium 2014 |
Information retrieval › retrieval models › neural retrieval
dense retrieval |
0.1 | 1 | 2021 | COUGH: A Challenge Dataset and Models for COVID-19 FAQ Retrieval · EMNLP (1) 2021 |
Information retrieval
retrieval models |
0.1 | 1 | 2021 | COUGH: A Challenge Dataset and Models for COVID-19 FAQ Retrieval · EMNLP (1) 2021 |
Medical and health informatics
clinical decision support |
0.1 | 1 | 2020 | Rationalizing Medical Relation Prediction from Corpus-level Statistics · ACL 2020 |
Medical and health informatics
clinical text processing |
0.1 | 1 | 2019 | SurfCon: Synonym Discovery on Privacy-Aware Clinical Data · KDD 2019 |
Bioinformatics and computational biology › gene expression analysis › microarray data preprocessing
microarray data processing |
0.1 | 1 | 2008 | lumi: a pipeline for processing Illumina microarray · Bioinform. 2008 |
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis |
0.1 | 1 | 2006 | Improved peak detection in mass spectrum by incorporating continuous wavelet transform-based pattern matching · Bioinform. 2006 |
Bioinformatics and computational biology › epigenomics › ChIP-seq analysis
peak detection |
0.1 | 1 | 2006 | Improved peak detection in mass spectrum by incorporating continuous wavelet transform-based pattern matching · Bioinform. 2006 |
Bioinformatics and computational biology › genomics › pharmacogenomics
pharmacogenetics |
0.1 | 1 | 2014 | Privacy in Pharmacogenetics: An End-to-End Case Study of Personalized Warfarin Dosing · USENIX Security Symposium 2014 |
Bioinformatics and computational biology
biomedical text mining |
0.0 | 1 | 2004 | MedlineR: an open source library in R for Medline literature data mining · Bioinform. 2004 |
Bioinformatics and computational biology
biological data visualization |
0.0 | 1 | 2002 | Applications of Tree-Maps to hierarchical biological data · Bioinform. 2002 |
Bioinformatics and computational biology › gene expression analysis
microarray data analysis |
0.0 | 1 | 2001 | Critical Assessment of Microarray Data Analysis: the 2001 challenge · Bioinform. 2001 |
Bioinformatics and computational biology › biomarker discovery
disease-gene association |
0.0 | 1 | 2009 | From disease ontology to disease-ontology lite: statistical methods to adapt a general-purpose ontology for the test of gene-ontology associations · Bioinform. 2009 |
Bioinformatics and computational biology
gene expression analysis |
0.0 | 1 | 2002 | Applications of Tree-Maps to hierarchical biological data · Bioinform. 2002 |
Bioinformatics and computational biology › protein function prediction
gene ontology annotation |
0.0 | 1 | 2002 | Applications of Tree-Maps to hierarchical biological data · Bioinform. 2002 |
Methods — techniques the papers use, named apart from their topics
recall and recognition theory · 0.9cooccurrence graph analysis · 0.9surface form matching · 0.8global context modeling · 0.8BM25 · 0.5BERT · 0.5random walk · 0.4neural network · 0.4matrix factorization · 0.4graph embedding · 0.4association testing · 0.2SNP genotyping · 0.2differential privacy · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | CliniQG4QA: Generating Diverse Questions for Domain Adaptation of Clinical Question AnsweringabstractClinical question answering (QA) aims to automatically answer questions from medical professionals based on clinical texts. Studies show that neural QA models trained on one corpus may not generalize well to new clinical texts from a different institute or a different patient group, where largescale QA pairs are not readily available for model retraining. To address this challenge, we propose a simple yet effective framework, CliniQG4QA, which leverages question generation (QG) to synthesize QA pairs on new clinical contexts and boosts QA models without requiring manual annotations. In order to generate diverse types of questions that are essential for training QA models, we further introduce a seq2seq-based question phrase prediction (QPP) module that can be used together with most existing QG models to diversify the generation. Our comprehensive experiment results show that the QA corpus generated by our framework can improve QA models on the new contexts (up to 8% absolute gain in terms of Exact Match), and that the QPP module plays a crucial role in achieving the gain.11Our dataset and code are available at: https://github.com/sunlabosu/CliniQG4QA/. Xiang Yue, Xinliang Frederick Zhang, Ziyu Yao 0002, Simon M. Lin, Huan Sun 0001 |
BIBM | 4 |
| 2021 | COUGH: A Challenge Dataset and Models for COVID-19 FAQ RetrievalabstractWe present a large, challenging dataset, COUGH, for COVID-19 FAQ retrieval.Similar to a standard FAQ dataset, COUGH consists of three parts: FAQ Bank, Query Bank and Relevance Set.The FAQ Bank contains ∼16K FAQ items scraped from 55 credible websites (e.g., CDC and WHO).For evaluation, we introduce Query Bank and Relevance Set, where the former contains 1,236 human-paraphrased queries while the latter contains ∼32 humanannotated FAQ items for each query.We analyze COUGH by testing different FAQ retrieval models built on top of BM25 and BERT, among which the best model achieves 48.8 under P@5, indicating a great challenge presented by COUGH and encouraging future research for further improvement.Our COUGH dataset is available at https://github. com/sunlab-osu/covid-faq. *Work was done when the first two authors were at OSU. 1 q and a are question and answer fields in an FAQ item.Question1: Should children wear masks?Answer1: In general, children 2 years and older should wear a mask...Appropriate and consistent use of masks...FAQ Bank Question2: Coping with Self-Quarantine Answer2: Remind yourself that difficult emotions are normal during self-quarantine... Query1: Is it possible for human beings to get sick with COVID-19 transmitted to them from animals?Query2: Is it possible to get infected by COVID 19 if I touch food surface packaging?Query Bank Question3: COVID-19是如何在⼈与⼈之间传播的? (How does COVID-19 spread between people?) Answer3: . Xinliang Frederick Zhang, Heming Sun, Xiang Yue, Simon M. Lin, Huan Sun 0001 |
EMNLP (1) | 4 |
| 2020 | Rationalizing Medical Relation Prediction from Corpus-level StatisticsabstractNowadays, the interpretability of machine learning models is becoming increasingly important, especially in the medical domain.Aiming to shed some light on how to rationalize medical relation prediction, we present a new interpretable framework inspired by existing theories on how human memory works, e.g., theories of recall and recognition.Given the corpus-level statistics, i.e., a global cooccurrence graph of a clinical text corpus, to predict the relations between two entities, we first recall rich contexts associated with the target entities, and then recognize relational interactions between these contexts to form model rationales, which will contribute to the final prediction.We conduct experiments on a real-world public clinical dataset and show that our framework can not only achieve competitive predictive performance against a comprehensive list of neural baseline models, but also present rationales to justify its prediction.We further collaborate with medical experts deeply to verify the usefulness of our model rationales for clinical decision making 1 . Zhen Wang 0041, Jennifer Lee, Simon M. Lin, Huan Sun 0001 |
ACL | 3 |
| 2020 | Tasks Associated With Extreme After-hours EHR Use Among Ambulatory Care Pediatricians
Selasi Attipoe, Yungui Huang, Sharon B. Schweikhart, Jennifer L. Hefner, Daniel M. Walker, Steve Rust, Jeffrey Hoffman, Simon M. Lin |
AMIA | 8 |
| 2020 | Clinical Phrase Mining with Language ModelsabstractA vast amount of vital clinical data is available within unstructured texts such as discharge summaries and procedure notes in Electronic Medical Records (EMRs). Automatically transforming such unstructured data into structured units is crucial for effective data analysis in the field of clinical informatics. Recognizing phrases that reveal important medical information in a concise and thorough manner is a fundamental step in this process. Existing systems that are built for opendomain texts are designed to detect mostly non-medical phrases, while tools designed specifically for extracting concepts from clinical texts are not scalable to large corpora and often leave out essential context surrounding those detected clinical concepts. We address these issues by proposing a framework, CliniPhrase, which adapts domain-specific deep neural network based language models (such as ClinicalBERT) to effectively and efficiently extract high-quality phrases from clinical documents with a limited amount of training data. Experimental results on the MIMIC-III dataset show that our method can outperform the current state-of-the-art techniques by up to 18% in terms of F1measure while being very efficient (up to 48 times faster). Kaushik Mani, Xiang Yue, Bernal Jimenez Gutierrez, Yungui Huang, Simon M. Lin, Huan Sun 0001 |
BIBM | 5 |
| 2020 | Graph embedding on biomedical networks: methods, applications and evaluationsabstractMOTIVATION: Graph embedding learning that aims to automatically learn low-dimensional node representations, has drawn increasing attention in recent years. To date, most recent graph embedding methods are evaluated on social and information networks and are not comprehensively studied on biomedical networks under systematic experiments and analyses. On the other hand, for a variety of biomedical network analysis tasks, traditional techniques such as matrix factorization (which can be seen as a type of graph embedding methods) have shown promising results, and hence there is a need to systematically evaluate the more recent graph embedding methods (e.g. random walk-based and neural network-based) in terms of their usability and potential to further the state-of-the-art. RESULTS: We select 11 representative graph embedding methods and conduct a systematic comparison on 3 important biomedical link prediction tasks: drug-disease association (DDA) prediction, drug-drug interaction (DDI) prediction, protein-protein interaction (PPI) prediction; and 2 node classification tasks: medical term semantic type classification, protein function prediction. Our experimental results demonstrate that the recent graph embedding methods achieve promising results and deserve more attention in the future biomedical graph analysis. Compared with three state-of-the-art methods for DDAs, DDIs and protein function predictions, the recent graph embedding methods achieve competitive performance without using any biological features and the learned embeddings can be treated as complementary representations for the biological features. By summarizing the experimental results, we provide general guidelines for properly selecting graph embedding methods and setting their hyper-parameters for different biomedical tasks. AVAILABILITY AND IMPLEMENTATION: As part of our contributions in the paper, we develop an easy-to-use Python package with detailed instructions, BioNEV, available at: https://github.com/xiangyue9607/BioNEV, including all source code and datasets, to facilitate studying various graph embedding methods on biomedical tasks. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiang Yue, Zhen Wang 0041, Jingong Huang, Srinivasan Parthasarathy 0001, Soheil Moosavinasab, Yungui Huang, Simon M. Lin, Wen Zhang 0008, Ping Zhang 0016, Huan Sun 0001 |
Bioinform. | 7 |
| 2019 | SurfCon: Synonym Discovery on Privacy-Aware Clinical DataabstractUnstructured clinical texts contain rich health-related information. To better utilize the knowledge buried in clinical texts, discovering synonyms for a medical query term has become an important task. Recent automatic synonym discovery methods leveraging raw text information have been developed. However, to preserve patient privacy and security, it is usually quite difficult to get access to large-scale raw clinical texts. In this paper, we study a new setting named synonym discovery on privacy-aware clinical data (i.e., medical terms extracted from the clinical texts and their aggregated co-occurrence counts, without raw clinical texts). To solve the problem, we propose a new framework SurfCon that leverages two important types of information in the privacy-aware clinical data, i.e., the surface form information, and the global context information for synonym discovery. In particular, the surface form module enables us to detect synonyms that look similar while the global context module plays a complementary role to discover synonyms that are semantically similar but in different surface forms, and both allow us to deal with the OOV query issue (i.e., when the query is not found in the given data). We conduct extensive experiments and case studies on publicly available privacy-aware clinical data, and show that SurfCon can outperform strong baseline methods by large margins under various settings. Zhen Wang 0041, Xiang Yue, Soheil Moosavinasab, Yungui Huang, Simon M. Lin, Huan Sun 0001 |
KDD | 5 |
| 2018 | Variation in After-hours EHR Usage among Ambulatory Care Physicians
Selasi Attipoe, Yungui Huang, Sharon B. Schweikhart, Simon M. Lin |
AMIA | 4 |
| 2018 | Char2Vec: Learning the Semantic Embedding of Rare and Unseen Words in the Biomedical Literature
Syed-Amad A. Hussain, Soheil Moosavinasab, Emre Sezgin, Yungui Huang, Simon M. Lin |
AMIA | 5 |
| 2018 | A roadmap of transition to health independence: A patient-centered study of self-management and digital-health solutions among teens with chronic conditions
Emre Sezgin, Monica Weiler, Anthony Weiler, Simon M. Lin |
AMIA | 4 |
| 2018 | DeepChild: Hospitalization Prediction via Neural Network
Xianlong Zeng, Soheil Moosavinasab, Enju Lin, Yungui Huang, Chang Liu 0028, Simon M. Lin |
AMIA | 6 |
| 2017 | Vitals Risk Index: A Simplified Pediatric Early Warning Index that Relies Solely on Objective Measures
Tyler Gorham, Steve Rust, Richard Hoyt, Swan Bee Liu, Jeffrey Hoffman, Tensing Maa, Yungui Huang, Simon M. Lin |
AMIA | 10 |
| 2017 | Genome Dashboard: An Interactive Tool for the Exploration of Human Genomic Variants and Phenotypic Associations
Katherine Miller, Rajeswari Swaminathan, Soheil Moosavinasab, Robert Strouse, Matthew Bailey 0002, Yungui Huang, Simon M. Lin |
AMIA | 7 |
| 2017 | DeepSuggest: Query Expansion for Clinical Notes by Deep Learning
Emre Sezgin, Koushik Jinna, Tran Bourgeois, Soheil Moosavinasab, Yungui Huang, Simon M. Lin |
AMIA | 6 |
| 2017 | Sharing exome sequencing data between Clinical Sequencing Labs and Healthcare Providers
Rajeswari Swaminathan, Yungui Huang, Katherine Miller, Matthew Pastore, Sayaka Hashimoto, Theodora Jacobson, Danielle Mouhlas, Simon M. Lin |
AMIA | 8 |
| 2017 | Clinical exome sequencing reports: current informatics practice and future opportunitiesabstractThe increased adoption of clinical whole exome sequencing (WES) has improved the diagnostic yield for patients with complex genetic conditions. However, the informatics practice for handling information contained in whole exome reports is still in its infancy, as evidenced by the lack of a common vocabulary within clinical sequencing reports generated across genetic laboratories. Genetic testing results are mostly transmitted using portable document format, which can make secondary analysis and data extraction challenging. This paper reviews a sample of clinical exome reports generated by Clinical Laboratory Improvement Amendments-certified genetic testing laboratories at tertiary-care facilities to assess and identify common data elements. Like structured radiology reports, which enable faster information retrieval and reuse, structuring genetic information within clinical WES reports would help facilitate integration of genetic information into electronic health records and enable retrospective research on the clinical utility of WES. We identify elements listed as mandatory according to practice guidelines but are currently missing from some of the clinical reports, which might help to organize the data when stored within structured databases. We also highlight elements, such as patient consent, that, although they do not appear within any of the current reports, may help in interpreting some of the information within the reports. Integrating genetic and clinical information would assist the adoption of personalized medicine for improved patient care and outcomes. Rajeswari Swaminathan, Yungui Huang, Caroline Astbury, Sara Fitzgerald-Butt, Katherine Miller, Justin Cole, Christopher W. Bartlett, Simon M. Lin |
J. Am. Medical Informatics Assoc. | 8 |
| 2015 | Application of clinical text data for phenome-wide association studies (PheWASs)abstractMOTIVATION: Genome-wide association studies (GWASs) are effective for describing genetic complexities of common diseases. Phenome-wide association studies (PheWASs) offer an alternative and complementary approach to GWAS using data embedded in the electronic health record (EHR) to define the phenome. International Classification of Disease version 9 (ICD9) codes are used frequently to define the phenome, but using ICD9 codes alone misses other clinically relevant information from the EHR that can be used for PheWAS analyses and discovery. RESULTS: As an alternative to ICD9 coding, a text-based phenome was defined by 23 384 clinically relevant terms extracted from Marshfield Clinic's EHR. Five single nucleotide polymorphisms (SNPs) with known phenotypic associations were genotyped in 4235 individuals and associated across the text-based phenome. All five SNPs genotyped were associated with expected terms (P<0.02), most at or near the top of their respective PheWAS ranking. Raw association results indicate that text data performed equivalently to ICD9 coding and demonstrate the utility of information beyond ICD9 coding for application in PheWAS. Scott J. Hebbring, Majid Rastegar-Mojarad, Zhan Ye, John Mayer, Crystal Jacobson, Simon M. Lin |
Bioinform. | 6 |
| 2014 | Privacy in Pharmacogenetics: An End-to-End Case Study of Personalized Warfarin Dosing
Matt Fredrikson, Eric Lantz, Somesh Jha, Simon M. Lin, David Page, Thomas Ristenpart |
USENIX Security Symposium | 4 |
| 2011 | PLoS Computational Biology Conference Postcards from ISMB/ECCB 2011abstractThis July, PLoS Computational Biology invited attendees of ISMB/ECCB 2011 (http://www.iscb.org/ismbeccb2011) to send us short reports of conference highlights in the guise of PLoS Conference Postcards. Philip E. Bourne, Editor-in-Chief, selected three Postcards, which we received from Poland, Germany, and the United States of America. If the reports below capture your interest, you can find Postcards from past conferences in our recent collection: http://collections.plos.org/ploscompbiol/conferencepostcards. Pedro Madrigal, Noa Sela, Simon M. Lin |
PLoS Comput. Biol. | 3 |
| 2010 | Comparison of Beta-value and M-value methods for quantifying methylation levels by microarray analysisabstractBACKGROUND: High-throughput profiling of DNA methylation status of CpG islands is crucial to understand the epigenetic regulation of genes. The microarray-based Infinium methylation assay by Illumina is one platform for low-cost high-throughput methylation profiling. Both Beta-value and M-value statistics have been used as metrics to measure methylation levels. However, there are no detailed studies of their relations and their strengths and limitations. RESULTS: We demonstrate that the relationship between the Beta-value and M-value methods is a Logit transformation, and show that the Beta-value method has severe heteroscedasticity for highly methylated or unmethylated CpG sites. In order to evaluate the performance of the Beta-value and M-value methods for identifying differentially methylated CpG sites, we designed a methylation titration experiment. The evaluation results show that the M-value method provides much better performance in terms of Detection Rate (DR) and True Positive Rate (TPR) for both highly methylated and unmethylated CpG sites. Imposing a minimum threshold of difference can improve the performance of the M-value method but not the Beta-value method. We also provide guidance for how to select the threshold of methylation differences. CONCLUSIONS: The Beta-value has a more intuitive biological interpretation, but the M-value is more statistically valid for the differential analysis of methylation levels. Therefore, we recommend using the M-value method for conducting differential methylation analysis and including the Beta-value statistics when reporting the results to investigators. Pan Du 0004, Chiang-Ching Huang, Nadereh Jafari, Warren A. Kibbe, Lifang Hou, Simon M. Lin |
BMC Bioinform. | 7 |
| 2010 | ChIPpeakAnno: a Bioconductor package to annotate ChIP-seq and ChIP-chip dataabstractBACKGROUND: Chromatin immunoprecipitation (ChIP) followed by high-throughput sequencing (ChIP-seq) or ChIP followed by genome tiling array analysis (ChIP-chip) have become standard technologies for genome-wide identification of DNA-binding protein target sites. A number of algorithms have been developed in parallel that allow identification of binding sites from ChIP-seq or ChIP-chip datasets and subsequent visualization in the University of California Santa Cruz (UCSC) Genome Browser as custom annotation tracks. However, summarizing these tracks can be a daunting task, particularly if there are a large number of binding sites or the binding sites are distributed widely across the genome. RESULTS: We have developed ChIPpeakAnno as a Bioconductor package within the statistical programming environment R to facilitate batch annotation of enriched peaks identified from ChIP-seq, ChIP-chip, cap analysis of gene expression (CAGE) or any experiments resulting in a large number of enriched genomic regions. The binding sites annotated with ChIPpeakAnno can be viewed easily as a table, a pie chart or plotted in histogram form, i.e., the distribution of distances to the nearest genes for each set of peaks. In addition, we have implemented functionalities for determining the significance of overlap between replicates or binding sites among transcription factors within a complex, and for drawing Venn diagrams to visualize the extent of the overlap between replicates. Furthermore, the package includes functionalities to retrieve sequences flanking putative binding sites for PCR amplification, cloning, or motif discovery, and to identify Gene Ontology (GO) terms associated with adjacent genes. CONCLUSIONS: ChIPpeakAnno enables batch annotation of the binding sites identified from ChIP-seq, ChIP-chip, CAGE or any technology that results in a large number of enriched genomic regions within the statistical programming environment R. Allowing users to pass their own annotation data such as a different Chromatin immunoprecipitation (ChIP) preparation and a dataset from literature, or existing annotation packages, such as GenomicFeatures and BSgenome, provides flexibility. Tight integration to the biomaRt package enables up-to-date annotation retrieval from the BioMart database. Lihua Julie Zhu, Claude Gazin, Nathan D. Lawson, Hervé Pagès, Simon M. Lin, David S. Lapointe, Michael R. Green |
BMC Bioinform. | 5 |
| 2009 | From disease ontology to disease-ontology lite: statistical methods to adapt a general-purpose ontology for the test of gene-ontology associationsabstractSubjective methods have been reported to adapt a general-purpose ontology for a specific application. For example, Gene Ontology (GO) Slim was created from GO to generate a highly aggregated report of the human-genome annotation. We propose statistical methods to adapt the general purpose, OBO Foundry Disease Ontology (DO) for the identification of gene-disease associations. Thus, we need a simplified definition of disease categories derived from implicated genes. On the basis of the assumption that the DO terms having similar associated genes are closely related, we group the DO terms based on the similarity of gene-to-DO mapping profiles. Two types of binary distance metrics are defined to measure the overall and subset similarity between DO terms. A compactness-scalable fuzzy clustering method is then applied to group similar DO terms. To reduce false clustering, the semantic similarities between DO terms are also used to constrain clustering results. As such, the DO terms are aggregated and the redundant DO terms are largely removed. Using these methods, we constructed a simplified vocabulary list from the DO called Disease Ontology Lite (DOLite). We demonstrated that DOLite results in more interpretable results than DO for gene-disease association tests. The resultant DOLite has been used in the Functional Disease Ontology (FunDO) Web application at http://www.projects.bioinformatics.northwestern.edu/fundo. Pan Du 0004, Jared Flatow, Michelle Holko, Warren A. Kibbe, Simon M. Lin |
Bioinform. | 7 |
| 2008 | lumi: a pipeline for processing Illumina microarrayabstractUNLABELLED: Illumina microarray is becoming a popular microarray platform. The BeadArray technology from Illumina makes its preprocessing and quality control different from other microarray technologies. Unfortunately, most other analyses have not taken advantage of the unique properties of the BeadArray system, and have just incorporated preprocessing methods originally designed for Affymetrix microarrays. lumi is a Bioconductor package especially designed to process the Illumina microarray data. It includes data input, quality control, variance stabilization, normalization and gene annotation portions. In specific, the lumi package includes a variance-stabilizing transformation (VST) algorithm that takes advantage of the technical replicates available on every Illumina microarray. Different normalization method options and multiple quality control plots are provided in the package. To better annotate the Illumina data, a vendor independent nucleotide universal identifier (nuID) was devised to identify the probes of Illumina microarray. The nuID annotation packages and output of lumi processed results can be easily integrated with other Bioconductor packages to construct a statistical data analysis pipeline for Illumina data. AVAILABILITY: The lumi Bioconductor package, www.bioconductor.org Pan Du 0004, Warren A. Kibbe, Simon M. Lin |
Bioinform. | 3 |
| 2007 | Application of Wavelet Transform to the MS-based Proteomics Data PreprocessingabstractMass spectrometry (MS) has become one of the major detection technologies for high-throughput proteomics. The preprocessing of mass spectra is crucial for its subsequent analysis like biomarker discovery or protein identification. Wavelet transform is gradually becoming an important methodology in the MS data preprocessing. This paper reviews the application of wavelet transforms in quality control, smoothing and peak detection of MS data preprocessing. It also proposes an improved Discrete Wavelet Transform (DWT) smoothing algorithm, which utilizes the cross-level DWT coefficients information during smoothing. Most of the algorithms described in this paper are included or will be included in the BioconductorMassSpecWaveletpackage. Pan Du 0004, Simon M. Lin, Warren A. Kibbe, Haihui Wang |
BIBE | 2 |
| 2006 | Improved peak detection in mass spectrum by incorporating continuous wavelet transform-based pattern matchingabstractMOTIVATION: A major problem for current peak detection algorithms is that noise in mass spectrometry (MS) spectra gives rise to a high rate of false positives. The false positive rate is especially problematic in detecting peaks with low amplitudes. Usually, various baseline correction algorithms and smoothing methods are applied before attempting peak detection. This approach is very sensitive to the amount of smoothing and aggressiveness of the baseline correction, which contribute to making peak detection results inconsistent between runs, instrumentation and analysis methods. RESULTS: Most peak detection algorithms simply identify peaks based on amplitude, ignoring the additional information present in the shape of the peaks in a spectrum. In our experience, 'true' peaks have characteristic shapes, and providing a shape-matching function that provides a 'goodness of fit' coefficient should provide a more robust peak identification method. Based on these observations, a continuous wavelet transform (CWT)-based peak detection algorithm has been devised that identifies peaks with different scales and amplitudes. By transforming the spectrum into wavelet space, the pattern-matching problem is simplified and in addition provides a powerful technique for identifying and separating the signal from the spike noise and colored noise. This transformation, with the additional information provided by the 2D CWT coefficients can greatly enhance the effective signal-to-noise ratio. Furthermore, with this technique no baseline removal or peak smoothing preprocessing steps are required before peak detection, and this improves the robustness of peak detection under a variety of conditions. The algorithm was evaluated with SELDI-TOF spectra with known polypeptide positions. Comparisons with two other popular algorithms were performed. The results show the CWT-based algorithm can identify both strong and weak peaks while keeping false positive rate low. AVAILABILITY: The algorithm is implemented in R and will be included as an open source module in the Bioconductor project. Pan Du 0004, Warren A. Kibbe, Simon M. Lin |
Bioinform. | 3 |
| 2004 | MedlineR: an open source library in R for Medline literature data miningabstractSUMMARY: We describe an open source library written in the R programming language for Medline literature data mining. This MedlineR library includes programs to query Medline through the NCBI PubMed database; to construct the co-occurrence matrix; and to visualize the network topology of query terms. The open source nature of this library allows users to extend it freely in the statistical programming language of R. To demonstrate its utility, we have built an application to analyze term-association by using only 10 lines of code. We provide MedlineR as a library foundation for bioinformaticians and statisticians to build more sophisticated literature data mining applications. AVAILABILITY: The library is available from http://dbsr.duke.edu/pub/MedlineR. Simon M. Lin, Patrick McConnell, Kimberly F. Johnson, Jennifer Shoemaker |
Bioinform. | 1 |
| 2002 | ICA and PLS modeling for functional analysis and drug sensitivity for DNA microarray signalsabstractThe DNA microarray technique offers an ability to analyze the expression profile of a genome. The complex correlation between the large number of genes present in the genome undermines straightforward understanding of their functionality. In this paper, we have proposed a pair of modeling schemes to recognize the functional identities of the known genes. In Independent Component Analysis (ICA), each of the microarray signals is modeled as a linear combination of some underlying independent components having specific biological interpretation. The second algorithm, Partial Least Squares (PLS) is proposed to identify the latent functional units contributing to drug sensitivity from the microarray data. Applications of this research include prediction of drug responses based on gene expressions, and also to identify the function(s) of a new gene. We consider the Rosetta compendium data set with yeast gene profiles, and the NCI-60 data set of human gene expressions as a function of drug type (cancer drugs are considered). Xuejun Liao, Nilanjan Dasgupta, Simon M. Lin, Lawrence Carin |
ICASSP | 3 |
| 2002 | Applications of Tree-Maps to hierarchical biological dataabstractUNLABELLED: A brief overview of Tree-Maps provides the basis for understanding two new implementations of Tree-Map methods. TreeMapClusterView provides a new way to view microarray gene expression data, and GenePlacer provides a view of gene ontology annotation data. We also discuss the benefits of Tree-Maps to visualize complex hierarchies in functional genomics. AVAILABILITY: Java class files are freely available at http://mendel.mc.duke.edu/bioinformatics/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: For more information on TreeMapClusterView (see http://mendel.mc.duke.edu/bioinformatics/software/boxclusterview/), and http://mendel.mc.duke.edu/bioinformatics/software/geneplacer/). Patrick McConnell, Kimberly F. Johnson, Simon M. Lin |
Bioinform. | 3 |
| 2002 | Sequential modeling for identifying CpG island locations in human genomeabstractWe consider several sequential processing algorithms for identifying genes in human DNA, based on detecting CpG ("C proceeds G") islands. The algorithms are designed to capture the underlying statistical structure in a DNA sequence. Sequential processing using a Markov model and a hidden Markov model are shown to identify most CpG islands in annotated (marked) DNA subsequences available from publicly available DNA datasets. We also consider a wavelet-based hidden Markov tree (HMT). In the context of the HMT, we address design of adaptive wavelets matched to CpG islands, this accomplished via lifting and genetic-algorithm optimization. Nilanjan Dasgupta, Simon M. Lin, Lawrence Carin |
IEEE Signal Process. Lett. | 2 |
| 2001 | Critical Assessment of Microarray Data Analysis: the 2001 challengeabstractUNLABELLED: We initiated the Critical Assessment of Microarray Data Analysis (CAMDA) conference to stimulate and evaluate the development of advanced data analysis techniques for microarrays. A standard data set has been released for this data analysis challenge. The goal of this challenge is to assess the performance of different analytical methods and at the same time to determine how such methods should be evaluated. We hope this effort will catalyze the discussion of microarray data analysis among the research community of biologists, statisticians, mathematicians, and computer scientists. AVAILABILITY: http://camda.duke.edu. Kimberly F. Johnson, Simon M. Lin |
Bioinform. | 2 |