EDBT 2026 Demo / reviewers in the wild / expert
Jin Chen 0004
dblp:03/5287-4
· DBLP profile ↗
37ranked-venue papers
0as first author
10since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 27 · 10 since 2021Artificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 3Systems, architecture and hardware · 2Graphics, computer vision, multimedia, augmented reality and games · 2Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Accurately deciphering spatial domains for spatially resolved transcriptomics with stClusterabstractSpatial transcriptomics provides valuable insights into gene expression within the native tissue context, effectively merging molecular data with spatial information to uncover intricate cellular relationships and tissue organizations. In this context, deciphering cellular spatial domains becomes essential for revealing complex cellular dynamics and tissue structures. However, current methods encounter challenges in seamlessly integrating gene expression data with spatial information, resulting in less informative representations of spots and suboptimal accuracy in spatial domain identification. We introduce stCluster, a novel method that integrates graph contrastive learning with multi-task learning to refine informative representations for spatial transcriptomic data, consequently improving spatial domain identification. stCluster first leverages graph contrastive learning technology to obtain discriminative representations capable of recognizing spatially coherent patterns. Through jointly optimizing multiple tasks, stCluster further fine-tunes the representations to be able to capture complex relationships between gene expression and spatial organization. Benchmarked against six state-of-the-art methods, the experimental results reveal its proficiency in accurately identifying complex spatial domains across various datasets and platforms, spanning tissue, organ, and embryo levels. Moreover, stCluster can effectively denoise the spatial gene expression patterns and enhance the spatial trajectory inference. The source code of stCluster is freely available at https://github.com/hannshu/stCluster. Tao Wang 0082, Han Shu, Jialu Hu, Yongtian Wang, Jin Chen 0004, Jiajie Peng, Xuequn Shang 0001 |
Briefings Bioinform. | 5 |
| 2023 | A Comparative Effectiveness Study on Opioid Use Disorder Prediction Using Artificial Intelligence and Existing Risk ModelsabstractOpioid use disorder (OUD) is a leading cause of death in the United States placing a tremendous burden on patients, their families, and health care systems. Artificial intelligence (AI) can be harnessed with available healthcare data to produce automated OUD prediction tools. In this retrospective study, we developed AI based models for OUD prediction and showed that AI can predict OUD more effectively than existing clinical tools including the unweighted opioid risk tool (ORT). Data include 474,208 patients' data over 10 years; 269,748 were females with an average age of 56.78 years. Cases are prescription opioid users with at least one diagnosis of OUD or at least one prescription for buprenorphine or methadone. Controls are prescription opioid users with no OUD diagnoses or buprenorphine or methadone prescriptions. On 100 randomly selected test sets including 47,396 patients, our proposed transformer-based AI model can predict OUD more efficiently (AUC = 0.742 ± 0.021) compared to logistic regression (AUC = 0.651 ± 0.025), random forest (AUC = 0.679 ± 0.026), xgboost (AUC = 0.690 ± 0.027), long short-term memory model (AUC = 0.706 ± 0.026), transformer (AUC = 0.725 ± 0.024), and unweighted ORT model (AUC = 0.559 ± 0.025). Our results show that embedding AI algorithms into clinical care may assist clinicians in risk stratification and management of patients receiving opioid therapy. Sajjad Fouladvand, Jeffery C. Talbert, Linda P. Dwoskin, Heather Bush, Amy Lynn Meadows, Lars E. Peterson, Yash R. Mishra, Steven K. Roggenkamp, Ramakanth Kavuluru, Jin Chen 0004 |
IEEE J. Biomed. Health Informatics | 11 |
| 2022 | KIT-LSTM: Knowledge-guided Time-aware LSTM for Continuous Clinical Risk PredictionabstractRapid accumulation of temporal Electronic Health Record (EHR) data and recent advances in deep learning have shown high potential in precisely and timely predicting patients' risks using AI. However, most existing risk prediction approaches ignore the complex asynchronous and irregular problems in real-world EHR data. This paper proposes a novel approach called Knowledge-guIded Time-aware LSTM (KIT-LSTM) for continuous mortality predictions using EHR. KIT-LSTM extends LSTM with two time-aware gates and a knowledge-aware gate to better model EHR and interprets results. Experiments on real-world data for patients with acute kidney injury with dialysis (AKI-D) demonstrate that KIT-LSTM performs better than the state-of-the-art methods for predicting patients' risk trajectories and model interpretation. KIT-LSTM can better support timely decision-making for clinicians. Lucas Jing Liu, Victor Ortiz-Soriano, Javier A. Neyra, Jin Chen 0004 |
BIBM | 4 |
| 2022 | UDA-CT: A General Framework for CT Image StandardizationabstractLarge-scale CT image studies often suffer from a lack of homogeneity regarding radiomic characteristics due to the images acquired with scanners from different vendors or with different reconstruction algorithms. We propose a deep learning-based framework called UDA-CT to tackle the homogeneity issue by leveraging both paired and unpaired images. Using UDA-CT, the CT images can be standardized both from different acquisition protocols of the same scanner and CT images acquired using a similar protocol but scanners from different vendors. UDA-CT incorporates recent advances in deep learning including domain adaptation and adversarial augmentation. It includes a unique design for model training batch which integrates nonstandard images and their adversarial variations to enhance model generalizability. The experimental results show that UDA-CT significantly improves the performance of the cross-scanner image standardization by utilizing both paired and unpaired data. Md. Selim, Jie Zhang 0092, Baowei Fei, Matthew A. Lewis, Guo-Qiang Zhang 0001, Jin Chen 0004 |
BIBM | 6 |
| 2022 | Enhancing discoveries of molecular QTL studies with small sample size using summary statistic imputationabstractQuantitative trait locus (QTL) analyses of multiomic molecular traits, such as gene transcription (eQTL), DNA methylation (mQTL) and histone modification (haQTL), have been widely used to infer the functional effects of genome variants. However, the QTL discovery is largely restricted by the limited study sample size, which demands higher threshold of minor allele frequency and then causes heavy missing molecular trait-variant associations. This happens prominently in single-cell level molecular QTL studies because of sample availability and cost. It is urgent to propose a method to solve this problem in order to enhance discoveries of current molecular QTL studies with small sample size. In this study, we presented an efficient computational framework called xQTLImp to impute missing molecular QTL associations. In the local-region imputation, xQTLImp uses multivariate Gaussian model to impute the missing associations by leveraging known association statistics of variants and the linkage disequilibrium (LD) around. In the genome-wide imputation, novel procedures are implemented to improve efficiency, including dynamically constructing a reused LD buffer, adopting multiple heuristic strategies and parallel computing. Experiments on various multiomic bulk and single-cell sequencing-based QTL datasets have demonstrated high imputation accuracy and novel QTL discovery ability of xQTLImp. Finally, a C++ software package is freely available at https://github.com/stormlovetao/QTLIMP. Tao Wang 0082, Yongzhuang Liu, Quanwei Yin, Jiaquan Geng, Jin Chen 0004, Xipeng Yin, Yongtian Wang, Xuequn Shang 0001, Chunwei Tian, Yadong Wang 0001, Jiajie Peng |
Briefings Bioinform. | 5 |
| 2022 | Correction to: Enhancing discoveries of molecular QTL studies with small sample size using summary statistic imputationabstractIn the originally published version of this manuscript, there was an error in the Funding section; ‘National Natural Science Foundation of China (6210071334, 62072376)’ has now been corrected to ‘National Natural Science Foundation of China (62102319, 62072376)’. Tao Wang 0082, Yongzhuang Liu, Quanwei Yin, Jiaquan Geng, Jin Chen 0004, Xipeng Yin, Yongtian Wang, Xuequn Shang 0001, Chunwei Tian, Yadong Wang 0001, Jiajie Peng |
Briefings Bioinform. | 5 |
| 2021 | Identifying Opioid Use Disorder from Longitudinal Healthcare Data using a Multi-stream Transformer
Sajjad Fouladvand, Jeffery C. Talbert, Linda P. Dwoskin, Heather Bush, Amy Lynn Meadows, Lars E. Peterson, Steven K. Roggenkamp, Ramakanth Kavuluru, Jin Chen 0004 |
AMIA | 9 |
| 2021 | Cross-Vendor CT Image Data Harmonization Using CVH-CT
Md. Selim, Jie Zhang 0092, Baowei Fei, Guo-Qiang Zhang 0001, Gary Yeeming Ge, Jin Chen 0004 |
AMIA | 6 |
| 2021 | Alzheimer's Disease Classification Using Genetic DataabstractThere has been a recent surge of interest in using genetic data to build ML-based accurate and interpretable disease classification models. In this line of research, we separately assess the potential of the peripheral blood gene expression data as well as the Single Nucleotide Polymorphism (SNP) data in building ML models for AD classification. We present a systematic approach on feature selection and ML model design using both types of genetic data provided by the Alzheimer’s Disease Neuroimaging Initiatives (ADNI). Our two-step feature selection produced a curated list of important genes. In addition to these selected genetic features, to examine the role of non-genetic covariates, we included age and number of education years (EDU) as extra features. In the Control (CN) vs. AD classification, the best performing classifier, XGBoost, trained with gene expression features only and that with extra features included had Area Under Curve (AUC) of 0.64 and 0.65 respectively. However, AUC for the same task using SNP data only and that with extra features included was 0.56 and 0.64 respectively. The just above chance results of classifier trained with SNP features and the improvement when used along with additional covariates indicate low potential of SNP data in AD classification when used alone while also indicating the importance of non-genetic factors associated with AD. Nevertheless, with well above chance performance, gene expression features show great potential especially between groups of AD progression, i.e., CN vs. AD, CN vs. EMCI, EMCI vs. AD and LMCI vs. AD. The source code and manual are available at https://github.com/mvrl/ADNI_Genetics. Subash Khanal, Jin Chen 0004, Nathan Jacobs, Ai-Ling Lin |
BIBM | 2 |
| 2021 | CT Image Harmonization for Enhancing Radiomics StudiesabstractWhile remarkable advances have been made in Computed Tomography (CT), most of the existing efforts focus on imaging enhancement while reducing radiation dose. How to normalize CT images acquired using non-standard protocols is vital for decision-making in cross-center large-scale radiomics studies but remains the boundary to explore. We develop a novel GAN-based image standardization algorithm called RadiomicGAN to mitigate the discrepancy caused by using non-standard acquisition protocols. In RadiomicGAN, a pre-trained U-Net has been adopted as part of the generator to learn radiomic feature distributions efficiently, and a novel training approach, called Window Training, has been developed to smoothly transform the pre-trained model to the medical imaging domain. In the experiments, we compared RadiomicGAN with four state-of-the-art CT image standardization approaches on both patient and phantom CT images acquired using three different reconstruction kernels. We objectively evaluated model performance based on more than 1,000 radiomic features. The results show that RadiomicGAN clearly outperforms the compared models. The source code, manual, and sample data are available at https://github.con selim-iitdu/radiomicGAN. Md. Selim, Jie Zhang 0092, Baowei Fei, Guo-Qiang Zhang 0001, Jin Chen 0004 |
BIBM | 5 |
| 2020 | STAN-CT: Standardizing CT Image using Generative Adversarial Networks
Md. Selim, Jie Zhang 0092, Baowei Fei, Guo-Qiang Zhang 0001, Jin Chen 0004 |
AMIA | 5 |
| 2020 | Identifying emerging phenomenon in long temporal phenotyping experimentsabstractMOTIVATION: The rapid improvement of phenotyping capability, accuracy and throughput have greatly increased the volume and diversity of phenomics data. A remaining challenge is an efficient way to identify phenotypic patterns to improve our understanding of the quantitative variation of complex phenotypes, and to attribute gene functions. To address this challenge, we developed a new algorithm to identify emerging phenomena from large-scale temporal plant phenotyping experiments. An emerging phenomenon is defined as a group of genotypes who exhibit a coherent phenotype pattern during a relatively short time. Emerging phenomena are highly transient and diverse, and are dependent in complex ways on both environmental conditions and development. Identifying emerging phenomena may help biologists to examine potential relationships among phenotypes and genotypes in a genetically diverse population and to associate such relationships with the change of environments or development. RESULTS: We present an emerging phenomenon identification tool called Temporal Emerging Phenomenon Finder (TEP-Finder). Using large-scale longitudinal phenomics data as input, TEP-Finder first encodes the complicated phenotypic patterns into a dynamic phenotype network. Then, emerging phenomena in different temporal scales are identified from dynamic phenotype network using a maximal clique based approach. Meanwhile, a directed acyclic network of emerging phenomena is composed to model the relationships among the emerging phenomena. The experiment that compares TEP-Finder with two state-of-art algorithms shows that the emerging phenomena identified by TEP-Finder are more functionally specific, robust and biologically significant. AVAILABILITY AND IMPLEMENTATION: The source code, manual and sample data of TEP-Finder are all available at: http://phenomics.uky.edu/TEP-Finder/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jiajie Peng, Junya Lu, Donghee Hoh, Ayesha S. Dina, Xuequn Shang 0001, David M. Kramer 0001, Jin Chen 0004 |
Bioinform. | 7 |
| 2020 | Mining Relationships among Multiple Entities in Biological NetworksabstractIdentifying topological relationships among multiple entities in biological networks is critical towards the understanding of the organizational principles of network functionality. Theoretically, this problem can be solved using minimum Steiner tree (MSTT) algorithms. However, due to large network size, it remains to be computationally challenging, and the predictive value of multi-entity topological relationships is still unclear. We present a novel solution called Cluster-based Steiner Tree Miner (CST-Miner) to instantly identify multi-entity topological relationships in biological networks. Given a list of user-specific entities, CST-Miner decomposes a biological network into nested cluster-based subgraphs, on which multiple minimum Steiner trees are identified. By merging all of them into a minimum cost tree, the optimal topological relationships among all the user-specific entities are revealed. Experimental results showed that CST-Miner can finish in nearly log-linear time and the tree constructed by CST-Miner is close to the global minimum. Jiajie Peng, Linjiao Zhu, Yadong Wang 0001, Jin Chen 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2019 | DoRC: Discovery of rare cells from ultra-large scRNA-seq dataabstractThe advent of droplet-based transcriptomics platforms has enabled parallel screening over thousands or millions of cells. One of the challenging issues is to identify the rare cells from the ultra-large scRNA-seq data. Existing algorithms to find rare cells are time consuming or memory-exhausting. We propose an efficient and accurate method, Discovery of Rare Cells (DoRC). The rareness scores generated by DoRC can help biologists focus the downstream analyses only on a fraction of expression profiles within ultra-large scRNA-seq data. We also demonstrate the efficacy of DoRC in delineating human blood dendritic cell sub-types using ~68k single-cell expression profiles of human blood cells. DoRC can recover artificially planted rare cells and is sensitive to cell type identities as well. Xiang Chen 0029, Fang-Xiang Wu, Jin Chen 0004, Min Li 0007 |
BIBM | 3 |
| 2019 | A Knowledge Graph Enhanced Topic Modeling Approach for Herb Recommendation
Xinyu Wang 0017, Xiaoling Wang 0004, Jin Chen 0004 |
DASFAA (1) | 4 |
| 2019 | MC-eLDA: Towards Pathogenesis Analysis in Traditional Chinese Medicine by Multi-Content Embedding LDA
Wendi Ji, Haofen Wang, Xiaoling Wang 0004, Jin Chen 0004 |
PAKDD (1) | 5 |
| 2018 | Enhancing Radiomic Features of CT Images using Generative Adversarial Network with Alternative Improvement
Gongbo Liang, Jie Zhang 0092, Michael A. Brooks, Jessica Howard, Jin Chen 0004 |
AMIA | 5 |
| 2018 | OLIVER: A Tool for Visual Data Analysis on Longitudinal Plant Phenomics Data
Oliver L. Tessmer, David M. Kramer 0001, Jin Chen 0004 |
BIBM | 3 |
| 2018 | Identifying Representative Network Motifs for Inferring Higher-order Structure of Biological Networks
Tao Wang 0082, Jiajie Peng, Yadong Wang 0001, Jin Chen 0004 |
BIBM | 4 |
| 2018 | Chrysanthemum Abnormal Petal Type Classification using Random Forest and Over-sampling
Peisen Yuan, Shougang Ren, Huanliang Xu, Jin Chen 0004 |
BIBM | 4 |
| 2018 | Joint Multi-Leaf Segmentation, Alignment, and Tracking for Fluorescence Plant VideosabstractThis paper proposes a novel framework for fluorescence plant video processing. The plant research community is interested in the leaf-level photosynthetic analysis within a plant. A prerequisite for such analysis is to segment all leaves, estimate their structures, and track them over time. We identify this as a joint multi-leaf segmentation, alignment, and tracking problem. First, leaf segmentation and alignment are applied on the last frame of a plant video to find a number of well-aligned leaf candidates. Second, leaf tracking is applied on the remaining frames with leaf candidate transformation from the previous frame. We form two optimization problems with shared terms in their objective functions for leaf alignment and tracking respectively. A quantitative evaluation framework is formulated to evaluate the performance of our algorithm with four metrics. Two models are learned to predict the alignment accuracy and detect tracking failure respectively in order to provide guidance for subsequent plant biology analysis. The limitation of our algorithm is also studied. Experimental results show the effectiveness, efficiency, and robustness of the proposed method. Xi Yin 0001, Xiaoming Liu 0002, Jin Chen 0004, David M. Kramer 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | PhenoCurve: capturing dynamic phenotype-environment relationships using phenomics dataabstractMotivation: Phenomics is essential for understanding the mechanisms that regulate or influence growth, fitness, and development. Techniques have been developed to conduct high-throughput large-scale phenotyping on animals, plants and humans, aiming to bridge the gap between genomics, gene functions and traits. Although new developments in phenotyping techniques are exciting, we are limited by the tools to analyze fully the massive phenotype data, especially the dynamic relationships between phenotypes and environments. Results: We present a new algorithm called PhenoCurve, a knowledge-based curve fitting algorithm, aiming to identify the complex relationships between phenotypes and environments, thus studying both values and trends of phenomics data. The results on both real and simulated data showed that PhenoCurve has the best performance among all the six tested methods. Its application to photosynthesis hysteresis pattern identification reveals new functions of core genes that control photosynthetic efficiency in response to varying environmental conditions, which are critical for understanding plant energy storage and improving crop productivity. Availability and Implementation: Software is available at phenomics.uky.edu/PhenoCurve. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Zheyun Feng, Jeffrey A. Cruz, Linda J. Savage, David M. Kramer 0001, Jin Chen 0004 |
Bioinform. | 7 |
| 2016 | Measuring phenotype semantic similarity using Human Phenotype OntologyabstractIt is critical yet remains to be challenging to make right disease diagnosis based on complex clinical characteristic and heterogeneous genetic background. Recently, Human Phenotype Ontology (HPO)-based phenotype similarity has been widely used to aid disease diagnosis. However, the existing measurements are revised based on the Gene Ontology-based term similarity models, which are not optimized for human phenotype ontologies. We propose a new similarity measure called PhenoSim. Our model includes a noise reduction component to model the noisy patient phenotype data, and a path-constrained Information Content-based method for measuring phenotype semantics similarity. Evaluation tests showed that PhenoSim could improve the performance of HPO-based phenotype similarity measurement. Jiajie Peng, Hansheng Xue, Yukai Shao, Xuequn Shang 0001, Yadong Wang 0001, Jin Chen 0004 |
BIBM | 6 |
| 2016 | Inter-functional analysis of high-throughput phenotype data by non-parametric clustering and its application to photosynthesisabstractMOTIVATION: Phenomics is the study of the properties and behaviors of organisms (i.e. their phenotypes) on a high-throughput scale. New computational tools are needed to analyze complex phenomics data, which consists of multiple traits/behaviors that interact with each other and are dependent on external factors, such as genotype and environmental conditions, in a way that has not been well studied. RESULTS: We deployed an efficient framework for partitioning complex and high dimensional phenotype data into distinct functional groups. To achieve this, we represented measured phenotype data from each genotype as a cloud-of-points, and developed a novel non-parametric clustering algorithm to cluster all the genotypes. When compared with conventional clustering approaches, the new method is advantageous in that it makes no assumption about the parametric form of the underlying data distribution and is thus particularly suitable for phenotype data analysis. We demonstrated the utility of the new clustering technique by distinguishing novel phenotypic patterns in both synthetic data and a high-throughput plant photosynthetic phenotype dataset. We biologically verified the clustering results using four Arabidopsis chloroplast mutant lines. AVAILABILITY AND IMPLEMENTATION: Software is available at www.msu.edu/~jinchen/NPM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: [email protected], [email protected] or [email protected]. Qiaozi Gao, Elisabeth Ostendorf, Jeffrey A. Cruz, David M. Kramer 0001, Jin Chen 0004 |
Bioinform. | 6 |
| 2016 | Extending gene ontology with gene association networksabstractMOTIVATION: Gene ontology (GO) is a widely used resource to describe the attributes for gene products. However, automatic GO maintenance remains to be difficult because of the complex logical reasoning and the need of biological knowledge that are not explicitly represented in the GO. The existing studies either construct whole GO based on network data or only infer the relations between existing GO terms. None is purposed to add new terms automatically to the existing GO. RESULTS: We proposed a new algorithm 'GOExtender' to efficiently identify all the connected gene pairs labeled by the same parent GO terms. GOExtender is used to predict new GO terms with biological network data, and connect them to the existing GO. Evaluation tests on biological process and cellular component categories of different GO releases showed that GOExtender can extend new GO terms automatically based on the biological network. Furthermore, we applied GOExtender to the recent release of GO and discovered new GO terms with strong support from literature. AVAILABILITY AND IMPLEMENTATION: Software and supplementary document are available at www.msu.edu/%7Ejinchen/GOExtender CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jiajie Peng, Tao Wang 0082, Jixuan Wang, Yadong Wang 0001, Jin Chen 0004 |
Bioinform. | 5 |
| 2016 | Multi-modality imagery database for plant phenotyping
Jeffrey A. Cruz, Xi Yin 0001, Xiaoming Liu 0002, Saif Muhammad Imran, Daniel D. Morris, David M. Kramer 0001, Jin Chen 0004 |
Mach. Vis. Appl. | 7 |
| 2015 | Incremental Distributed Weighted Class Discriminant Analysis on Interval-Valued Emitter ParametersabstractIn the age of big data, the emitter parameter measurement data is generally characteristic of uncertainty in the form of normally-distributed intervals, enormous size and continuous growth. However, existing interval-valued data analysis methods generally assume a uniform distribution instead and are unable to adapt to the rapid growth of volume. To address the above problems, we have brought forward an incremental distributed weighted class discriminant analysis method on interval-valued emitter parameters. Extensive experiments indicate that our method is able to cope with these new characteristics effectively. Wei Wang 0002, Jiaheng Lu, Jin Chen 0004 |
KSEM | 4 |
| 2015 | Plant photosynthesis phenomics data quality controlabstractMOTIVATION: Plant phenomics, the collection of large-scale plant phenotype data, is growing exponentially. The resources have become essential component of modern plant science. Such complex datasets are critical for understanding the mechanisms governing energy intake and storage in plants, and this is essential for improving crop productivity. However, a major issue facing these efforts is the determination of the quality of phenotypic data. Automated methods are needed to identify and characterize alterations caused by system errors, all of which are difficult to remove in the data collection step and distinguish them from more interesting cases of altered biological responses. RESULTS: As a step towards solving this problem, we have developed a coarse-to-refined model called dynamic filter to identify abnormalities in plant photosynthesis phenotype data by comparing light responses of photosynthesis using a simplified kinetic model of photosynthesis. Dynamic filter employs an expectation-maximization process to adjust the kinetic model in coarse and refined regions to identify both abnormalities and biological outliers. The experimental results show that our algorithm can effectively identify most of the abnormalities in both real and synthetic datasets. AVAILABILITY AND IMPLEMENTATION: Software available at www.msu.edu/%7Ejinchen/DynamicFilter . Jeffrey A. Cruz, Linda J. Savage, David M. Kramer 0001, Jin Chen 0004 |
Bioinform. | 5 |
| 2015 | Measuring semantic similarities by combining gene ontology annotations and gene co-function networksabstractBACKGROUND: Gene Ontology (GO) has been used widely to study functional relationships between genes. The current semantic similarity measures rely only on GO annotations and GO structure. This limits the power of GO-based similarity because of the limited proportion of genes that are annotated to GO in most organisms. RESULTS: We introduce a novel approach called NETSIM (network-based similarity measure) that incorporates information from gene co-function networks in addition to using the GO structure and annotations. Using metabolic reaction maps of yeast, Arabidopsis, and human, we demonstrate that NETSIM can improve the accuracy of GO term similarities. We also demonstrate that NETSIM works well even for genomes with sparser gene annotation data. We applied NETSIM on large Arabidopsis gene families such as cytochrome P450 monooxygenases to group the members functionally and show that this grouping could facilitate functional characterization of genes in these families. CONCLUSIONS: Using NETSIM as an example, we demonstrated that the performance of a semantic similarity measure could be significantly improved after incorporating genome-specific information. NETSIM incorporates both GO annotations and gene co-function network data as a priori knowledge in the model. Therefore, functional similarities of GO terms that are not explicitly encoded in GO but are relevant in a taxon-specific manner become measurable when GO annotations are limited. Supplementary information and software are available at http://www.msu.edu/~jinchen/NETSIM . Jiajie Peng, Sahra Uygun, Taehyong Kim, Yadong Wang 0001, Seung Y. Rhee, Jin Chen 0004 |
BMC Bioinform. | 6 |
| 2014 | Multi-leaf tracking from fluorescence plant videosabstractDriven by the plant phenotyping application, this paper proposes a new leaf tracking framework to jointly segment, align and track multiple leaves from fluorescence plant videos. Our framework consists of two steps. First, leaf alignment is applied to one video frame to generate a collection of leaf candidates. Second, we define a set of transformation parameters operated on the leaf candidates in order to optimize the alignment in the subsequent video frame according to an objective function. Gradient descent is employed to solve this optimization problem. Experimental results show that the proposed multi-leaf tracking algorithm is superior to the image-based leaf alignment method in terms of three quantitative metrics. Xi Yin 0001, Xiaoming Liu 0002, Jin Chen 0004, David M. Kramer 0001 |
ICIP | 3 |
| 2014 | Multi-leaf alignment from fluorescence plant imagesabstractIn this paper, we propose a multi-leaf alignment framework based on Chamfer matching to study the problem of leaf alignment from fluorescence images of plants, which will provide a leaf-level analysis of photosynthetic activities. Different from the naive procedure of aligning leaves iteratively using the Chamfer distance, the new algorithm aims to find the best alignment of multiple leaves simultaneously in an input image. We formulate an optimization problem of an objective function with three terms: the average of chamfer distances of aligned leaves, the number of leaves, and the difference between the synthesized mask by the leaf candidates and the original image mask. Gradient descent is used to minimize our objective function. A quantitative evaluation framework is also formulated to test the performance of our algorithm. Experimental results show that the proposed multi-leaf alignment optimization performs substantially better than the baseline of the Chamfer matching algorithm in terms of both accuracy and efficiency. Xi Yin 0001, Xiaoming Liu 0002, Jin Chen 0004, David M. Kramer 0001 |
WACV | 3 |
| 2014 | Towards integrative gene functional similarity measurementabstractBACKGROUND: In Gene Ontology, the "Molecular Function" (MF) categorization is a widely used knowledge framework for gene function comparison and prediction. Its structure and annotation provide a convenient way to compare gene functional similarities at the molecular level. The existing gene similarity measures, however, solely rely on one or few aspects of MF without utilizing all the rich information available including structure, annotation, common terms, lowest common parents. RESULTS: We introduce a rank-based gene semantic similarity measure called InteGO by synergistically integrating the state-of-the-art gene-to-gene similarity measures. By integrating three GO based seed measures, InteGO significantly improves the performance by about two-fold in all the three species studied (yeast, Arabidopsis and human). CONCLUSIONS: InteGO is a systematic and novel method to study gene functional associations. The software and description are available at http://www.msu.edu/~jinchen/InteGO. Jiajie Peng, Yadong Wang 0001, Jin Chen 0004 |
BMC Bioinform. | 3 |
| 2014 | Wireless Spectrum Occupancy Prediction Based on Partial Periodic Pattern MiningabstractCognitive radio appears as a promising technology to allocate wireless spectrum between licensed and unlicensed users in an efficient way. When unlicensed users opportunistically utilize spectrum holes, prediction models that infer the availability of spectrum holes can help to improve the spectrum extraction rate and reduce the collision rate. In this paper, a spectrum occupancy prediction model based on Partial Periodic Pattern Mining (PPPM) is introduced. The mining aims at identifying frequent spectrum occupancy patterns that are hidden in the spectrum usage of a channel. The mined frequent patterns are then used to predict future channel states (i.e., busy or idle). Based on the prediction, unlicensed users are able to utilize spectrum holes aggressively without introducing significant interference to licensed users. PPPM outperforms traditional Frequent Pattern Mining (FPM) by considering real patterns that do not repeat perfectly due to noise, sensing errors, and irregular behaviors. Using real-world Wi-Fi and personal communication service (PCS) activities, we show a significant reduction on miss rate in channel state prediction. With the proposed prediction mechanism, the performance of Dynamic Spectrum Access (DSA) is substantially improved. Further, we extend the three-state PPPM to an N-state PPPM to predict the duration of high/low utilization in a channel. The frequent patterns of channel utilization duration are critical in optimizing channel switch strategies. The high prediction accuracy is validated with data collected in the paging bands. Pei Huang 0001, Chin-Jung Liu 0001, Xi Yang 0016, Li Xiao 0001, Jin Chen 0004 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2013 | Identifying cross-category relations in gene ontology and constructing genome-specific term association networksabstractBACKGROUND: Gene Ontology (GO) has been widely used in biological databases, annotation projects, and computational analyses. Although the three GO categories are structured as independent ontologies, the biological relationships across the categories are not negligible for biological reasoning and knowledge integration. However, the existing cross-category ontology term similarity measures are either developed by utilizing the GO data only or based on manually curated term name similarities, ignoring the fact that GO is evolving quickly and the gene annotations are far from complete. RESULTS: In this paper we introduce a new cross-category similarity measurement called CroGO by incorporating genome-specific gene co-function network data. The performance study showed that our measurement outperforms the existing algorithms. We also generated genome-specific term association networks for yeast and human. An enrichment based test showed our networks are better than those generated by the other measures. CONCLUSIONS: The genome-specific term association networks constructed using CroGO provided a platform to enable a more consistent use of GO. In the networks, the frequently occurred MF-centered hub indicates that a molecular function may be shared by different genes in multiple biological processes, or a set of genes with the same functions may participate in distinct biological processes. And common subgraphs in multiple organisms also revealed conserved GO term relationships. Software and data are available online at http://www.msu.edu/~jinchen/CroGO. Jiajie Peng, Jin Chen 0004, Yadong Wang 0001 |
BMC Bioinform. | 2 |
| 2012 | Mining frequent partial periodic patterns in spectrum usage dataabstractCognitive radio appears as a promising technology to allocate wireless spectrum between licensed and unlicensed users. Predictive methods for inferring the availability of spectrum holes can help to reduce collision and improve spectrum extraction. This paper introduces a Partial Periodic Pattern Mining (PPPM) algorithm to identify frequent spectrum occupancy patterns that are hidden in the spectrum usage of a channel. The mined frequent patterns are then used to predict future channel states (i.e., busy or idle). PPPM outperforms traditional Frequent Pattern Mining (FPM) by considering real patterns that do not repeat perfectly. Using real life network activities, we show a significant reduction on miss rate in channel state prediction. Pei Huang 0001, Chin-Jung Liu 0001, Li Xiao 0001, Jin Chen 0004 |
IWQoS | 4 |
| 2012 | Wireless Spectrum Occupancy Prediction Based on Partial Periodic Pattern MiningabstractCognitive radio appears as a promising technology to allocate wireless spectrum between licensed and unlicensed users in an efficient way. The availability of spectrum holes vastly affects the throughput and delay of unlicensed users. Predictive methods for inferring the availability of spectrum holes can help to improve spectrum extraction rate and reduce collision rate. In this paper, a spectrum occupancy prediction model based on Partial Periodic Pattern Mining (PPPM) is introduced. The mining aims to identify frequent spectrum occupancy patterns that are hidden in the spectrum usage of a channel. The mined frequent patterns are then used to predict future channel states (i.e., busy or idle). Based on the prediction, unlicensed users will be able to make use of spectrum holes efficiently without introducing significant interference to licensed users. PPPM outperforms traditional Frequent Pattern Mining (FPM) by considering real patterns that do not repeat perfectly due to noise, sensing errors, and irregular behaviors. Using real life network activities we show a significant reduction on miss rate in channel state prediction. With the proposed prediction mechanism, the performance of Dynamic Spectrum Access (DSA) is substantially improved. Pei Huang 0001, Chin-Jung Liu 0001, Li Xiao 0001, Jin Chen 0004 |
MASCOTS | 4 |
| 2012 | Optimal timepoint sampling in high-throughput gene expression experimentsabstractMOTIVATION: Determining the best sampling rates (which maximize information yield and minimize cost) for time-series high-throughput gene expression experiments is a challenging optimization problem. Although existing approaches provide insight into the design of optimal sampling rates, our ability to utilize existing differential gene expression data to discover optimal timepoints is compelling. RESULTS: We present a new data-integrative model, Optimal Timepoint Selection (OTS), to address the sampling rate problem. Three experiments were run on two different datasets in order to test the performance of OTS, including iterative-online and a top-up sampling approaches. In all of the experiments, OTS outperformed the best existing timepoint selection approaches, suggesting that it can optimize the distribution of a limited number of timepoints, potentially leading to better biological insights about the resulting gene expression patterns. AVAILABILITY: OTS is available at www.msu.edu/∼jinchen/OTS. Bruce A. Rosa, Ian T. Major, Wensheng Qin, Jin Chen 0004 |
Bioinform. | 5 |