Jinze Liu

dblp:73/1534 · DBLP profile ↗
← Back
38ranked-venue papers
11as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 20 · 6 since 2021Artificial intelligence and machine learning · 12 · 9 first-author · 4 since 2021Databases, data management, data science and information retrieval · 11 · 7 first-authorSystems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Sample size requirements for machine learning classification of binary outcomes in bulk RNA-Seq data
abstract
BACKGROUND: Bulk RNA sequencing data is often leveraged to build machine learning (ML)-based predictive models for classification of disease groups or subtypes, but the sample size needed to adequately train these models is unknown. METHODS: We collected 27 experimental datasets from the Gene Expression Omnibus and the Cancer Genome Atlas. In 24/27 datasets, pseudo-data were simulated using Bayesian Network Generation. Three ML algorithms were assessed: XGBoost (XGB), Random Forest (RF), and Neural Networks (NN). Learning curves were fit, and sample sizes needed to reach the full-dataset AUC minus 0.02 were determined and compared across the datasets/algorithms. Multivariable negative binomial regression models quantified relationships between dataset-level characteristics and required sample sizes within each algorithm. These models were validated in independent experimental datasets. RESULTS: Across the datasets studied, median required sample sizes were 480 (XGB)/190 (RF)/269 (NN). Higher effect sizes, less class imbalance/dispersion, and less complex data were associated with lower required sample size. Validation demonstrated that predictions were accurate in new data. CONCLUSIONS: Comparison of results to sample sizes obtained from differential analysis power analysis methods showed that ML methods generally required larger sample sizes. In conclusion, incorporating ML-based sample size planning alongside traditional power analysis can provide more robust results.
Scott Silvey, Amy Olex, Shaojun Tang, Jinze Liu
BMC Bioinform.4
2025 Offline Reinforcement Learning with Koopman Operators for Control of Soft Robots
abstract
Soft robots are promising to offer flexibility in environmental interaction tasks through compliant deformations. However, the infinite degrees of freedom and high nonlinearity of dynamics pose significant challenges in dynamic modeling and control in soft robots. While online reinforcement learning (RL) is promising for designing policies directly from data, the black-box policy learning process suffers from data inefficiency and sim-to-real gap, limiting its applications in soft robots. To address these challenges, we propose a novel offline RL with Koopman operators (KORL) framework to generate control policies for soft robots without using physical simulators or real-world interactions. In particular, we first utilize a deep neural network to map dynamics of soft robots to a lifted Koopman observable space, which is inherently linear. Then, an offline RL algorithm with a control-informed actor is designed to learn the robotic policy in the linear observable space. This is significantly different from the black-box policy design in existing offline RL paradigms. The designed Koopman observable enables efficient model-free policy learning with linear control theory, improving control performance while preserving interpretability in policy learning. The effectiveness of our KORL framework is validated in a real-world soft robotic system. Comparative experimental results demonstrate that our method outperforms state-of-the-art methods in target-reaching and trajectory-tracking tasks.
Yihe Yang, Wenyu Cao, Jinze Liu
IROS6
2025 UAV-Assisted Vehicle Edge Computing Resource Allocation Under Blockchain Application
abstract
ABSTRACT The synergy of Unmanned Aerial Vehicles (UAVs) and edge computing provides a dynamic platform for real‐time data processing, enabling various applications such as autonomous driving and intelligent traffic management. In order to provide users with higher and satisfactory service quality, it is necessary to allocate edge computing resources between edge computing stations and UAVs. However, because the vehicle data collected by many onboard sensors contains sensitive and personal information, and there is a lack of financial incentives, vehicles are reluctant to upload data to edge servers. Unlike sharing data for free, encrypted data transactions mitigate security and privacy concerns while providing an incentive for car owners to share data. Edge servers pay a price in data transactions, and reputation management is an effective way to help them trade with reliable and available vehicles. This paper proposes a resource pricing and trading scheme based on Stackelberg dynamic game and adopts a reputation management scheme based on multi‐armed bandit (MAB) so that edge servers can choose vehicles with high reputation for data trading to ensure the credibility and reliability of data. An optimization problem is presented in this paper to maximize vehicle revenue in data transactions under the constraints of delay, energy consumption, and security level. Experiments show that the proposed scheme is effective in vehicle reputation management, data transaction selection, and resource allocation.
Fengna Ji, Jinze Liu
Concurr. Comput. Pract. Exp.5
2025 A novel multi-level hierarchy optimization algorithm for pipeline inner detector speed control
Jinze Liu, Jian Feng 0001, Huaguang Zhang, Shengxiang Yang
Neurocomputing1
2025 Deciphering genomic codes using advanced natural language processing techniques: a scoping review
abstract
Objectives: The vast and complex nature of human genomic sequencing data presents challenges for effective analysis. This review aims to investigate the application of Natural Language Processing (NLP) techniques, particularly Large Language Models (LLMs) and transformer architectures, in deciphering genomic codes, focusing on tokenization, transformer models, and regulatory annotation prediction. This review aims to assess data and model accessibility in the most recent literature, gaining a better understanding of the existing capabilities and constraints of these tools in processing genomic sequencing data. Methods: Following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, our scoping review was conducted across PubMed, Medline, Scopus, Web of Science, Embase, and ACM Digital Library. Studies were included if they focused on NLP methodologies applied to genomic sequencing data analysis, without restrictions on publication date or article type. Results: A total of 26 studies published between 2021 and April 2024 were selected for review. The review highlights that tokenization and transformer models enhance the processing and understanding of genomic data, with applications in predicting regulatory annotations like transcription-factor binding sites and chromatin accessibility. Discussion: The application of NLP and LLMs to genomic sequencing data interpretation is a promising field that can help streamline the processing of large-scale genomic data while providing a better understanding of its complex structures. It can potentially drive advancements in personalized medicine by offering more efficient and scalable solutions for genomic analysis. Further research is needed to discuss and overcome limitations, enhancing model transparency and applicability.
Shuyan Cheng, Yishu Wei, Yiliang Zhou, Drew N. Wright, Jinze Liu, Yifan Peng 0002
J. Am. Medical Informatics Assoc.6
2025 Machine learning applications related to suicide in military and Veterans: A scoping literature review
Yishu Wei, Yanshan Wang, Yunyu Xiao, Ronald K. Poropatich, Gretchen L. Haas, Yiye Zhang, Chunhua Weng, Jinze Liu, Lisa A. Brenner, James M. Bjork, Yifan Peng 0002
J. Biomed. Informatics9
2025 Dynamic Category Preference Learning: Tackling Long-Tailed Semi-Supervised Segmentation in Remote Sensing
abstract
Remote sensing (RS) image segmentation faces persistent challenges due to limited labeled data and the presence of long-tailed distribution. While semi-supervised learning (SSL) can leverage unlabeled data, it often struggles to address the severe class imbalance that manifests as a long-tailed distribution. To tackle this, we propose the Dynamic Category Preference-Aware Semantic Segmentation Framework (CPSeg). Unlike existing methods that estimate model’s category preference on pseudo-labels or rely on prior information, CPSeg introduces a simple yet effective approach to directly model this preference. By feeding a Pattern-Free Image (PFI) into the model, we capture its current tendencies toward certain categories without any prior knowledge. This dynamically learned category preference drives two key components of CPSeg. First, it enables adaptive weight allocation, which adjusts the training focus by assigning higher weights to low-frequency categories, promoting balanced learning across all categories. Second, CPSeg introduces a Category-wise Pairwise Memory Bank (CPMB) and integrates category preference information to perform dynamic preference-aware data augmentation, making the augmentation process more suitable for addressing the long-tailed distribution problem. When compared with other advanced approaches, CPSeg exhibits superior performance, achieving mean intersection over union (mIoU) scores of 74.50% on the GID-15 and 59.26% on the MSl dataset, respectively, under a challenging 1/8 labeled data setting, highlighting its strong potential for addressing long-tailed distribution and class imbalance in RS segmentation.
Yong Liu 0048, Shiheng Zhang, Baoqi Yu, Zijun Zhou, Jinze Liu
IEEE Trans. Geosci. Remote. Sens.6
2024 scFed: federated learning for cell type classification with scRNA-seq
abstract
The advent of single-cell RNA sequencing (scRNA-seq) has revolutionized our understanding of cellular heterogeneity and complexity in biological tissues. However, the nature of large, sparse scRNA-seq datasets and privacy regulations present challenges for efficient cell identification. Federated learning provides a solution, allowing efficient and private data use. Here, we introduce scFed, a unified federated learning framework that allows for benchmarking of four classification algorithms without violating data privacy, including single-cell-specific and general-purpose classifiers. We evaluated scFed using eight publicly available scRNA-seq datasets with diverse sizes, species and technologies, assessing its performance via intra-dataset and inter-dataset experimental setups. We find that scFed performs well on a variety of datasets with competitive accuracy to centralized models. Though Transformer-based model excels in centralized training, its performance slightly lags behind single-cell-specific model within the scFed framework, coupled with a notable time complexity concern. Our study not only helps select suitable cell identification methods but also highlights federated learning's potential for privacy-preserving, collaborative biomedical research.
Shuang Wang 0002, Bochen Shen, Lanting Guo, Mengqi Shang, Jinze Liu, Bairong Shen
Briefings Bioinform.5
2024 BTR: a bioinformatics tool recommendation system
abstract
MOTIVATION: The rapid expansion of Bioinformatics research has led to a proliferation of computational tools for scientific analysis pipelines. However, constructing these pipelines is a demanding task, requiring extensive domain knowledge and careful consideration. As the Bioinformatics landscape evolves, researchers, both novice and expert, may feel overwhelmed in unfamiliar fields, potentially leading to the selection of unsuitable tools during workflow development. RESULTS: In this article, we introduce the Bioinformatics Tool Recommendation system (BTR), a deep learning model designed to recommend suitable tools for a given workflow-in-progress. BTR leverages recent advances in graph neural network technology, representing the workflow as a graph to capture essential context. Natural language processing techniques enhance tool recommendations by analyzing associated tool descriptions. Experiments demonstrate that BTR outperforms the existing Galaxy tool recommendation system, showcasing its potential to streamline scientific workflow construction. AVAILABILITY AND IMPLEMENTATION: The Python source code is available at https://github.com/ryangreenj/bioinformatics_tool_recommendation.
Ryan Green, Xufeng Qu, Jinze Liu
Bioinform.3
2024 N-Level Hierarchy-Based Optimal Control to Develop Therapeutic Strategies for Ecological Evolutionary Dynamics Systems
abstract
This article mainly proposes an evolutionary algorithm and its first application to develop therapeutic strategies for ecological evolutionary dynamics systems (EEDS), obtaining the balance between tumor cells and immune cells by rationally arranging chemotherapeutic drugs and immune drugs. First, an EEDS nonlinear kinetic model is constructed to describe the relationship between tumor cells, immune cells, dose, and drug concentration. Second, the N-level hierarchy optimization (NLHO) algorithm is designed and compared with five algorithms on 20 benchmark functions, which proves the feasibility and effectiveness of NLHO. Finally, we apply NLHO into EEDS to give a dynamic adaptive optimal control policy and develop therapeutic strategies to reduce tumor cells, while minimizing the harm of chemotherapy drugs and immune drugs to the human body. The experimental results prove the validity of the research method.
Jinze Liu, Jiayue Sun, Huaguang Zhang, Shun Xu, Zifang Zou
IEEE Trans. Neural Networks Learn. Syst.1
2023 CLF-CBF Constraints for Real-Time Avoidance of Multiple Obstacles in Bipedal Locomotion and Navigation
abstract
This paper presents a reactive planning system that allows a Cassie-series bipedal robot to avoid multiple non-overlapping obstacles via a single, continuously differentiable control barrier function (CBF). The overall system detects an individual obstacle via a height map derived from a LiDAR point cloud and computes an elliptical outer approximation, which is then turned into a CBF. The QP-CLF-CBF formalism developed by Ames et al. is applied to ensure that safe trajectories are generated. Safe planning in environments with multiple obstacles is demonstrated both in simulation and experimentally on the Cassie biped.
Jinze Liu, Minzhe Li, Jessy W. Grizzle, Jiunn-Kai Huang
IROS1
2020 HRV-Spark: Computing Heart Rate Variability Measures Using Apache Spark
abstract
Heart rate variability (HRV) analysis has been serving as a significant promising marker in clinical research over the last few decades. The rapidly growing heart rate data generated from various devices, particularly the electrocardiograph (ECG), need to be stored properly and processed timely. There is a pressing need to develop efficient approaches for performing HRV analyses based on ECG signals. In this paper, we introduce a cloud computing approach (called HRV-Spark) to compute HRV measures in parallel by leveraging Apache Spark and a QRS detection algorithm in [1]. We ran HRV-Spark on Amazon Web Services (AWS) clusters using large-scale datasets in the National Sleep Research Resource. We evaluated the performance and scalability of HRV-Spark in terms of the number of computing nodes in the AWS cluster, the size of the input datasets, and the hardware configuration of the computing nodes. The results show that HRV-Spark is an efficient and scalable approach for computing HRV measures.
Xufeng Qu, Jinze Liu, Licong Cui
BIBM3
2020 Inferring Gene Regulatory Networks of Metabolic Enzymes Using Gradient Boosted Trees
abstract
Metabolic reprogramming is a hallmark of cancer. In cancer cells, transcription factors (TFs) govern metabolic reprogramming through abnormally increasing or decreasing the transcription rate of metabolic enzymes, which provides cancer cells growth advantages and concurrently leads to the altered metabolic phenotypes observed in many cancers. Consequently, targeting TFs that govern metabolic reprogramming can be highly effective for novel cancer therapeutics. In this paper, we present TFmeta, a machine learning approach to uncover TFs that govern reprogramming of cancer metabolism. Our approach achieves the state-of-the-art performance in reconstructing relations between TFs and their target genes on public benchmark datasets. Leveraging TF binding profiles inferred from genome-wide ChIP-seq experiments and 150 RNA-seq samples from 75 paired cancerous and non-cancerous human lung tissues, our approach predicted 19 key TFs that may be the major regulators of the gene expression changes of metabolic enzymes of the central metabolic pathway glycolysis, which may underlie the dysregulation of glycolysis in non-small-cell lung cancer patients.
Yi Zhang 0078, Andrew N. Lane, Teresa W.-M. Fan, Jinze Liu
IEEE J. Biomed. Health Informatics5
2019 Hadoop-EDF: Large-scale Distributed Processing of Electrophysiological Signal Data in Hadoop MapReduce
abstract
Rapidly growing volume of electrophysiological signals has been generated for clinical research in neurological disorders. European Data Format (EDF) is a standard format for storing electrophysiological signals. However, the bottleneck of existing signal analysis tools for handling large-scale datasets is the sequential way of loading large EDF files before performing signal analyses. To overcome this, we develop Hadoop-EDF, a distributed signal processing tool to load EDF data in a parallel manner using Hadoop MapReduce. Hadoop-EDF uses a robust data partition algorithm making EDF data parallelly processable. We evaluate Hadoop-EDF's scalability and performance by leveraging two datasets from the National Sleep Research Resource and running experiments on Amazon Web Service clusters. The performance of Hadoop-EDF on a 20-node cluster achieved about 26 times and 47 times faster than the sequential processing of 200 small-size files and 200 large-size files, respectively. The results demonstrate that Hadoop-EDF is more suitable and effective in processing large EDF files.
Jinze Liu, Licong Cui
BIBM3
2018 Toward data-driven identification of kingdom-specific protein sequence motifs
Corrine F. Elliott, Kristin Linscott, Satrio Husodo, Joseph Chappell, Jinze Liu
BIBM5
2018 Adaptive software search toward users' customized requirements in GitHub
abstract
Because of a tremendous growth of Open Source Software (OSS) scale and the diversity of users' requirements, users now face the problem of finding OSS that meets their expectations in a huge number of OSS resources.However, current GitHub-provided search service has a shortage in adapting to user needs.When facing diverse users' requirements, it cannot always return satisfactory results.In this paper, we provide a more efficient search service for OSS on GitHub.We first design a multi-dimensional measurement model for OSS, which forms a corresponding metric system and quantitative measurement method.Then we propose a ranking algorithm based on fuzzy synthetic evaluation in order to implement an adaptive metric ranking method that is oriented to user requirements.We verify that our work is useful by setting up experiments.The experiment results show that compared with GitHub-provided search service (searching by "Best Match" & searching by "Most Stars"), the effectiveness of our method improved by 97.6% and 13.8% respectively, which means our method returns search results which meet users' expectations more, and has high self-adaptive ability.
Jinze Liu, Tao Wang 0006, Yue Yu 0001, Gang Yin
SEKE1
2018 A novel data structure to support ultra-fast taxonomic classification of metagenomic sequences with k-mer signatures
abstract
Motivation: Metagenomic read classification is a critical step in the identification and quantification of microbial species sampled by high-throughput sequencing. Although many algorithms have been developed to date, they suffer significant memory and/or computational costs. Due to the growing popularity of metagenomic data in both basic science and clinical applications, as well as the increasing volume of data being generated, efficient and accurate algorithms are in high demand. Results: We introduce MetaOthello, a probabilistic hashing classifier for metagenomic sequencing reads. The algorithm employs a novel data structure, called l-Othello, to support efficient querying of a taxon using its k-mer signatures. MetaOthello is an order-of-magnitude faster than the current state-of-the-art algorithms Kraken and Clark, and requires only one-third of the RAM. In comparison to Kaiju, a metagenomic classification tool using protein sequences instead of genomic sequences, MetaOthello is three times faster and exhibits 20-30% higher classification sensitivity. We report comparative analyses of both scalability and accuracy using a number of simulated and empirical datasets. Availability and implementation: MetaOthello is a stand-alone program implemented in C ++. The current version (1.0) is accessible via https://doi.org/10.5281/zenodo.808941. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Xinan Liu, Ye Yu 0001, Corrine F. Elliott, Chen Qian 0001, Jinze Liu
Bioinform.6
2017 Whole mammogram image classification with convolutional neural networks
abstract
Due to the high variability in tumor morphology and the low signal-to-noise ratio inherent to mammography, manual classification of mammogram yields a significant number of patients being called back, and subsequent large number of biopsies performed to reduce the risk of missing cancer. The convolutional neural network (CNN) is a popular deep-learning construct used in image classification. This technique has achieved significant advancements in large-set image-classification challenges in recent years. In this study, we had obtained over 3000 high-quality original mammograms with approval from an institutional review board at the University of Kentucky. Different classifiers based on CNNs were built, and each classifier was evaluated based on its performance relative to truth values generated by histology results from biopsy and two-year negative mammogram follow-up confirmed by expert radiologists. Our results showed that CNN model we had built and optimized via data augmentation and transfer learning have a great potential for automatic breast cancer detection using mammograms.
Yi Zhang 0078, Erik Y. Han, Nathan Jacobs, Qiong Han, Jinze Liu
BIBM7
2016 DeepSplice: Deep classification of novel splice junctions revealed by RNA-seq
abstract
Alternative splicing (AS) is a regulated process that enables the production of multiple mRNA transcripts from a single multi-exon gene. The availability of large-scale RNA-seq datasets has made it possible to predict splice junctions, as well as splice sites through spliced alignment to the reference genome. This greatly enhances the capability to decipher gene structures and explore the diversity of splicing variants. However, existing ab initio aligners are vulnerable to false positive spliced alignments as a result of sequence errors and random sequence matches. These spurious alignments can lead to a significant set of false positive splice junction predictions, confusing downstream analyses of splice variant detection and abundance estimation. In this work, we illustrate that splice junction sequence characteristics can be ascertained from experimental data with deep learning techniques. We employ deep convolutional neural networks for a novel splice junction classification tool named DeepSplice that (i) outperforms state-of-the-art methods for predicting splice sites, (ii) shows high computational efficiency and (iii) can be applied to self-defined training data by users.
Yi Zhang 0078, Xinan Liu, James N. MacLeod, Jinze Liu
BIBM4
2014 Piecing the puzzle together: a revisit to transcript reconstruction problem in RNA-seq
abstract
The advancement of RNA sequencing (RNA-seq) has provided an unprecedented opportunity to assess both the diversity and quantity of transcript isoforms in an mRNA transcriptome. In this paper, we revisit the computational problem of transcript reconstruction and quantification. Unlike existing methods which focus on how to explain the exons and splice variants detected by the reads with a set of isoforms, we aim at reconstructing transcripts by piecing the reads into individual effective transcript copies. Simultaneously, the quantity of each isoform is explicitly measured by the number of assembled effective copies, instead of estimated solely based on the collective read count. We have developed a novel method named Astroid that solves the problem of effective copy reconstruction on the basis of a flow network. The RNA-seq reads are represented as vertices in the flow network and are connected by weighted edges that evaluate the likelihood of two reads originating from the same effective copy. A maximum likelihood set of transcript copies is then reconstructed by solving a minimum-cost flow problem on the flow network. Simulation studies on the human transcriptome have demonstrated the superior sensitivity and specificity of Astroid in transcript reconstruction as well as improved accuracy in transcript quantification over several existing approaches. The application of Astroid on two real RNA-seq datasets has further demonstrated its accuracy through high correlation between the estimated isoform abundance and the qRT-PCR validations.
Yan Huang 0006, Jinze Liu
BMC Bioinform.3
2012 A Robust Method for Transcript Quantification with RNA-seq Data
Yan Huang 0006, Corbin D. Jones, James N. MacLeod, Derek Y. Chiang, Jan F. Prins, Jinze Liu
RECOMB8
2011 FDM: a graph-based statistical method to detect differential transcription using RNA-seq data
abstract
MOTIVATION: In eukaryotic cells, alternative splicing expands the diversity of RNA transcripts and plays an important role in tissue-specific differentiation, and can be misregulated in disease. To understand these processes, there is a great need for methods to detect differential transcription between samples. Our focus is on samples observed using short-read RNA sequencing (RNA-seq). METHODS: We characterize differential transcription between two samples as the difference in the relative abundance of the transcript isoforms present in the samples. The magnitude of differential transcription of a gene between two samples can be measured by the square root of the Jensen Shannon Divergence (JSD*) between the gene's transcript abundance vectors in each sample. We define a weighted splice-graph representation of RNA-seq data, summarizing in compact form the alignment of RNA-seq reads to a reference genome. The flow difference metric (FDM) identifies regions of differential RNA transcript expression between pairs of splice graphs, without need for an underlying gene model or catalog of transcripts. We present a novel non-parametric statistical test between splice graphs to assess the significance of differential transcription, and extend it to group-wise comparison incorporating sample replicates. RESULTS: Using simulated RNA-seq data consisting of four technical replicates of two samples with varying transcription between genes, we show that (i) the FDM is highly correlated with JSD* (r=0.82) when average RNA-seq coverage of the transcripts is sufficiently deep; and (ii) the FDM is able to identify 90% of genes with differential transcription when JSD* >0.28 and coverage >7. This represents higher sensitivity than Cufflinks (without annotations) and rDiff (MMD), which respectively identified 69 and 49% of the genes in this region as differential transcribed. Using annotations identifying the transcripts, Cufflinks was able to identify 86% of the genes in this region as differentially transcribed. Using experimental data consisting of four replicates each for two cancer cell lines (MCF7 and SUM102), FDM identified 1425 genes as significantly different in transcription. Subsequent study of the samples using quantitative real time polymerase chain reaction (qRT-PCR) of several differential transcription sites identified by FDM, confirmed significant differences at these sites. AVAILABILITY: http://csbio-linux001.cs.unc.edu/nextgen/software/FDM CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Darshan Singh, Christian F. Orellana, Corbin D. Jones, Derek Y. Chiang, Jinze Liu, Jan F. Prins
Bioinform.7
2010 A probabilistic framework for aligning paired-end RNA-seq data
abstract
MOTIVATION: The RNA-seq paired-end read (PER) protocol samples transcript fragments longer than the sequencing capability of today's technology by sequencing just the two ends of each fragment. Deep sampling of the transcriptome using the PER protocol presents the opportunity to reconstruct the unsequenced portion of each transcript fragment using end reads from overlapping PERs, guided by the expected length of the fragment. METHODS: A probabilistic framework is described to predict the alignment to the genome of all PER transcript fragments in a PER dataset. Starting from possible exonic and spliced alignments of all end reads, our method constructs potential splicing paths connecting paired ends. An expectation maximization method assigns likelihood values to all splice junctions and assigns the most probable alignment for each transcript fragment. RESULTS: The method was applied to 2 x 35 bp PER datasets from cancer cell lines MCF-7 and SUM-102. PER fragment alignment increased the coverage 3-fold compared to the alignment of the end reads alone, and increased the accuracy of splice detection. The accuracy of the expectation maximization (EM) algorithm in the presence of alternative paths in the splice graph was validated by qRT-PCR experiments on eight exon skipping alternative splicing events. PER fragment alignment with long-range splicing confirmed 8 out of 10 fusion events identified in the MCF-7 cell line in an earlier study by (Maher et al., 2009). AVAILABILITY: Software available at http://www.netlab.uky.edu/p/bioinfo/MapSplice/PER.
Xiaping He, Derek Y. Chiang, Jan F. Prins, Jinze Liu
Bioinform.6
2010 Analysis of equine protein-coding gene structure and expression by RNA-sequencing
abstract
RNA-sequencing (RNA-seq) data from eight equine tissue samples (34-day whole embryo, full term placental villous, adult testes, adult cerebellum, adult articular cartilage, adult LPS-stimulated articular cartilage, adult synovial membrane, and adult LPS-stimulated synovial membrane) were used to refine the structural annotation of protein-coding genes in the horse and for a preliminary assessment of tissue-specific expression patterns.
Stephen J. Coleman, Jinze Liu, James N. MacLeod
BMC Bioinform.3
2010 Functional neighbors: inferring relationships between nonhomologous protein families using family-specific packing motifs
abstract
We describe a new approach for inferring the functional relationships between nonhomologous protein families by looking at statistical enrichment of alternative function predictions in classification hierarchies such as Gene Ontology (GO) and Structural Classification of Proteins (SCOP). Protein structures are represented by robust graph representations, and the fast frequent subgraph mining algorithm is applied to protein families to generate sets of family-specific packing motifs, i.e., amino acid residue-packing patterns shared by most family members but infrequent in other proteins. The function of a protein is inferred by identifying in it motifs characteristic of a known family. We employ these family-specific motifs to elucidate functional relationships between families in the GO and SCOP hierarchies. Specifically, we postulate that two families are functionally related if one family is statistically enriched by motifs characteristic of another family, i.e., if the number of proteins in a family containing a motif from another family is greater than expected by chance. This function-inference method can help annotate proteins of unknown function, establish functional neighbors of existing families, and help specify alternate functions for known proteins.
Deepak Bandyopadhyay, Jun Huan, Jinze Liu, Jan F. Prins, Jack Snoeyink, Wei Wang 0010, Alexander Tropsha
IEEE Trans. Inf. Technol. Biomed.3
2009 Unsupervised learning of high-order structural semantics from images
abstract
Structural semantics are fundamental to understanding both natural and man-made objects from languages to buildings. They are manifested as repeated structures or patterns and are often captured in images. Finding repeated patterns in images, therefore, has important applications in scene understanding, 3D reconstruction, and image retrieval as well as image compression. Previous approaches in visual-pattern mining limited themselves by looking for frequently co-occurring features within a small neighborhood in an image. However, semantics of a visual pattern are typically defined by specific spatial relationships between features regardless of the spatial proximity. In this paper, semantics are represented as visual elements and geometric relationships between them. A novel unsupervised learning algorithm finds pair-wise associations of visual elements that have consistent geometric relationships sufficiently often. The algorithms are efficient - maximal matchings are determined without combinatorial search. High-order structural semantics are extracted by mining patterns that are composed of pairwise spatially consistent associations of visual elements. We demonstrate the effectiveness of our approach for discovering repeated visual patterns on a variety of image collections.
Jizhou Gao, Jinze Liu, Ruigang Yang
ICCV3
2009 Privacy Preservation in Social Networks with Sensitive Edge Weights
abstract
With the development of emerging social networks, such as Facebook and MySpace, security and privacy threats arising from social network analysis bring a risk of disclosure of confidential knowledge when the social network data is shared or made public. In addition to the current social network anonymity de-identification techniques, we study a situation, such as in a business transaction network, in which weights are attached to network edges that are considered to be confidential (e.g., transactions). We consider perturbing the weights of some edges to preserve data privacy when the network is published, while retaining the shortest path and the approximate cost of the path between some pairs of nodes in the original network. We develop two privacy-preserving strategies for this application. The first strategy is based on a Gaussian randomization multiplication, the second one is a greedy perturbation algorithm based on graph theory. In particular, the second strategy not only yields an approximate length of the shortest path while maintaining the shortest path between selected pairs of nodes, but also maximizes privacy preservation of the original weights. We present experimental results to support our mathematical analysis.
Jie Wang 0008, Jinze Liu, Jun Zhang 0001
SDM3
2008 Functional Neighbors: Inferring Relationships between Non-Homologous Protein Families Using Family-Specific Packing Motifs
abstract
We describe a new approach for inferring the functional relationships between non-homologous protein families by looking at statistical enrichment of alternative function predictions in classification hierarchies such as Gene Ontology (GO) and Structural Classification of Proteins (SCOP). Protein structures are represented by robust graphs, and the Fast Frequent Subgraph Mining algorithm is applied to protein families to generate sets of family-specific packing motifs, i.e. amino acid residue packing patterns shared by most family members but infrequent in other proteins. The function of a protein is inferred by identifying in it motifs characteristic of a known family. We employ these family-specific motifs to elucidate functional relationships between families in the GO and SCOP hierarchies. Specifically, we postulate that two families are functionally related if one family is statistically enriched by motifs characteristic of another family, i.e. if the number of proteins in a family containing a motif from another family is greater than expected by chance. This function inference method can help annotate proteins of unknown function, establish functional neighbors of existing families, and help specify alternate functions for known proteins.
Deepak Bandyopadhyay, Jun Huan, Jinze Liu, Jan F. Prins, Jack Snoeyink, Wei Wang 0010, Alexander Tropsha
BIBM3
2008 Approximate Clustering on Distributed Data Streams
abstract
We investigate the problem of clustering on distributed data streams. In particular, we consider the k-median clustering on stream data arriving at distributed sites which communicate through a routing tree. Distributed clustering on high speed data streams is a challenging task due to limited communication capacity, storage space, and computing power at each site. In this paper, we propose a suite of algorithms for computing (1 + epsiv) -approximate k-median clustering over distributed data streams under three different topology settings: topology-oblivious, height-aware, and path-aware. Our algorithms reduce the maximum per node transmission topolylogN(opposed to Omega(N) for transmitting the raw data). We have simulated our algorithms on a distributed stream system with both real and synthetic datasets composed of millions of data. In practice, our algorithms are able to reduce the data transmission to a small fraction of the original data. Moreover, our results indicate that the algorithms are scalable with respect to the data volume, approximation factor, and the number of sites.
Qi Zhang 0025, Jinze Liu, Wei Wang 0010
ICDE2
2008 Mining Approximate Order Preserving Clusters in the Presence of Noise
abstract
Subspace clustering has attracted great attention due to its capability of finding salient patterns in high dimensional data. Order preserving subspace clusters have been proven to be important in high throughput gene expression analysis, since functionally related genes are often co-expressed under a set of experimental conditions. Such co-expression patterns can be represented by consistent orderings of attributes. Existing order preserving cluster models require all objects in a cluster have identical attribute order without deviation. However, real data are noisy due to measurement technology limitation and experimental variability which prohibits these strict models from revealing true clusters corrupted by noise. In this paper, we study the problem of revealing the order preserving clusters in the presence of noise. We propose a noise-tolerant model called approximate order preserving cluster (AOPC). Instead of requiring all objects in a cluster have identical attribute order, we require that (1) at least a certain fraction of the objects have identical attribute order; (2) other objects in the cluster may deviate from the consensus order by up to a certain fraction of attributes. We also propose an algorithm to mine AOPC. Experiments on gene expression data demonstrate the efficiency and effectiveness of our algorithm.
Mengsheng Zhang, Wei Wang 0010, Jinze Liu
ICDE3
2007 Incremental Subspace Clustering over Multiple Data Streams
abstract
Data streams are often locally correlated, with a subset of streams exhibiting coherent patterns over a subset of time points. Subspace clustering can discover clusters of objects in different subspaces. However, traditional sub- space clustering algorithms for static data sets are not readily used for incremental clustering, and is very expensive for frequent re-clustering over dynamically changing stream data. In this paper, we present an efficient incremental sub- space clustering algorithm for multiple streams over sliding windows. Our algorithm detects all the delta-CC-Clusters, which capture the coherent changing patterns among a set of streams over a set of time points. delta-CC'-Cluster s are incrementally generated by traversing a directed acyclic graph pDAG. We propose efficient insertion and deletion operations to update the pDAG dynamically. In addition, effective pruning techniques are applied to reduce the search space. Experiments on real data sets demonstrate the performance of our algorithm.
Qi Zhang 0025, Jinze Liu, Wei Wang 0010
ICDM2
2007 PoClustering: Lossless Clustering of Dissimilarity Data
abstract
Given a set of objects V with a dissimilarity measure between pairs of objects in V, a PoCluster is a collection of sets P ⊂ powerset(V) partially ordered by the ⊂ relation such that S ⊂ T iff the maximal dissimilarity among objects in S is less than the maximal dissimilarity among objects in T. PoClusters capture categorizations of objects that are not strictly hierarchical, such as those found in ontologies. PoClusters can not, in general, be constructed using hierarchical clustering algorithms. In this paper, we examine the relationship between PoClusters and dissimilarity matrices and prove that PoClusters are in one-to-one correspondence with the set of dissimilarity matrices. The PoClustering problem is NP-Complete, and we present a heuristic algorithm for it in this paper. Experiments on both synthetic and real datasets demonstrate the quality and scalability of the algorithms.
Jinze Liu, Qi Zhang 0025, Wei Wang 0010, Leonard McMillan, Jan F. Prins
SDM1
2006 Clustering pair-wise dissimilarity data into partially ordered sets
abstract
Ontologies represent data relationships as hierarchies of possibly overlapping classes. Ontologies are closely related to clustering hierarchies, and in this article we explore this relationship in depth. In particular, we examine the space of ontologies that can be generated by pairwise dissimilarity matrices. We demonstrate that classical clustering algorithms, which take dissimilarity matrices as inputs, do not incorporate all available information. In fact, only special types of dissimilarity matrices can be exactly preserved by previous clustering methods. We model ontologies as a partially ordered set (poset) over the subset relation. In this paper, we propose a new clustering algorithm, that generates a partially ordered set of clusters from a dissimilarity matrix.
Jinze Liu, Qi Zhang 0025, Wei Wang 0010, Leonard McMillan, Jan F. Prins
KDD1
2006 Mining Approximate Frequent Itemsets In the Presence of Noise: Algorithm and Analysis
abstract
Frequent itemset mining is a popular and important first step in the analysis of data arising in a broad range of applications. The traditional “exact” model for frequent itemsets requires that every item occur in each supporting transaction. However, real data is typically subject to noise and measurement error. To date, the effect of noise on exact frequent pattern mining algorithms have been addressed primarily through simulation studies, and there has been limited attention to the development of noise tolerant algorithms. In this paper we propose a noise tolerant itemset model, which we call approximate frequent itemsets (AFI). Like frequent itemsets, the AFI model requires that an itemset has a minimum number of supporting transactions. However, the AFI model tolerates a controlled fraction of errors in each item and each supporting transaction. Motivating this model are theoretical results (and a supporting simulation study presented here) which state that, in the presence of even low levels of noise, large frequent itemsets are broken into fragments of logarithmic size; thus the itemsets cannot be recovered by a routine application of frequent itemset mining. By contrast, we provide theoretical results showing that the AFI criterion is well suited to recovery of block structures subject to noise. We developed and implemented an algorithm to mine AFIs that generalizes the level-wise enumeration of frequent itemsets by allowing noise. We propose the noise-tolerant support threshold, a relaxed version of support, which varies with the length of the itemset and the noise threshold. We exhibit an Apriori property that permits the pruning of an itemset if any of its sub-itemset is not sufficiently supported. Several experiments presented demonstrate that the AFI algorithm enables better recoverability of frequent patterns under noisy conditions than existing frequent itemset mining approaches. Noise-tolerant support pruning also renders an order of magnitude performance gain over existing methods.
Jinze Liu, Susan Paulsen, Xing Sun 0002, Wei Wang 0010, Andrew B. Nobel, Jan F. Prins
SDM1
2005 Mining Approximate Frequent Itemsets from Noisy Data
abstract
Frequent itemset mining is a popular and important first step in analyzing data sets across a broad range of applications. The traditional, "exact" approach for finding frequent itemsets requires that every item in the itemset occurs in each supporting transaction. However, real data is typically subject to noise, and in the presence of such noise, traditional itemset mining may fail to detect relevant itemsets, particularly those large itemsets that are more vulnerable to noise. In this paper we propose approximate frequent itemsets (AFI), as a noise-tolerant itemset model. In addition to the usual requirement for sufficiently many supporting transactions, the AFI model places constraints on the fraction of errors permitted in each item column and the fraction of errors permitted in a supporting transaction. Taken together, these constraints winnow out the approximate itemsets that exhibit systematic errors. In the context of a simple noise model, we demonstrate that AFI is better at recovering underlying data patterns, while identifying fewer spurious patterns than either the exact frequent itemset approach or the existing error tolerant itemset approach of Yang et al.
Jinze Liu, Susan Paulsen, Wei Wang 0010, Andrew B. Nobel, Jan F. Prins
ICDM1
2004 Revealing True Subspace Clusters in High Dimensions
abstract
Subspace clustering is one of the best approaches for discovering meaningful clusters in high dimensional space. One cluster in high dimensional space may be transcribed into multiple distinct maximal clusters by projecting onto different subspaces. A direct consequence of clustering independently in each subspace is an overwhelmingly large set of overlapping clusters which may be significantly similar. To reveal the true underlying clusters, we propose a similarity measurement of the overlapping clusters. We adopt the model of Gaussian tailed hyper-rectangles to capture the distribution of any subspace cluster. A set of experiments on a synthetic dataset demonstrates the effectiveness of our approach. Application to real gene expression data also reveals impressive meta-clusters expected by biologists.
Jinze Liu, Karl Strohmaier, Wei Wang 0010
ICDM1
2004 A framework for ontology-driven subspace clustering
abstract
Traditional clustering is a descriptive task that seeks to identify homogeneous groups of objects based on the values of their attributes. While domain knowledge is always the best way to justify clustering, few clustering algorithms have ever take domain knowledge into consideration. In this paper, the domain knowledge is represented by hierarchical ontology. We develop a framework by directly incorporating domain knowledge into clustering process, yielding a set of clusters with strong ontology implication. During the clustering process, ontology information is utilized to efficiently prune the exponential search space of the subspace clustering algorithms. Meanwhile, the algorithm generates automatical interpretation of the clustering result by mapping the natural hierarchical organized subspace clusters with significant categorical enrichment onto the ontology hierarchy. Our experiments on a set of gene expression data using gene ontology demonstrate that our pruning technique driven by ontology significantly improve the clustering performance with minimal degradation of the cluster quality. Meanwhile, many hierarchical organizations of gene clusters corresponding to a sub-hierarchies in gene ontology were also successfully captured.
Jinze Liu, Wei Wang 0010, Jiong Yang 0001
KDD1
2003 OP-Cluster: Clustering by Tendency in High Dimensional Space
abstract
Clustering is the process of grouping a set of objects into classes of similar objects. Because of unknownness of the hidden patterns in the data sets, the definition of similarity is very subtle. Until recently, similarity measures are typically based on distances, e.g Euclidean distance and cosine distance. We propose a flexible yet powerful clustering model, namely OP-cluster (Order Preserving Cluster). Under this new model, two objects are similar on a subset of dimensions if the values of these two objects induce the same relative order of those dimensions. Such a cluster might arise when the expression levels of (coregulated) genes can rise or fall synchronously in response to a sequence of environment stimuli. Hence, discovery of OP-Cluster is essential in revealing significant gene regulatory networks. A deterministic algorithm is designed and implemented to discover all the significant OP-Clusters. A set of extensive experiments has been done on several real biological data sets to demonstrate its effectiveness and efficiency in detecting coregulated patterns.
Jinze Liu, Wei Wang 0010
ICDM1