Zhi Wei 0001

dblp:48/5555-1 · DBLP profile ↗
← Back
39ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0001-6059-4267ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 11 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Semi-supervised disentangled representation learning for single-cell RNA sequencing data
abstract
Single-cell RNA sequencing (scRNA-seq) data are inherently high-dimensional, and most analysis tools reduce this complexity by projecting the data into a low-dimensional latent space before performing downstream analyses. However, the resulting representations are often entangled, with biological or technical factors such as batch effects and disease stages mixed together, which complicates interpretation. Recent methods have introduced disentanglement mechanisms to improve interpretability, but they typically require large amounts of well-annotated data to perform well or are limited to factors with only a few categories. To address these challenges, we propose SCDRL (Semi-Supervised Disentangled Representation Learning for Single-Cell RNA Sequencing Data), a method that uses gene expression profiles together with a small proportion of labeled samples to learn disentangled representations that separate batch effects, cell types, and other biological signals, thereby enhancing interpretability. Unlike existing approaches, SCDRL generalizes from factors with only a few categories to complex settings involving more than 10 cell types. Experiments on both simulated and real-world datasets demonstrate that SCDRL consistently outperforms existing disentangled representation learning methods for scRNA-seq data, even when only 5% of labeled samples are available.
Yuanjie Zou, Zhi Wei 0001
Briefings Bioinform.3
2026 Integrating feature selection with unsupervised deep embedding for clustering single-cell RNA-seq data
abstract
Single-cell RNA sequencing (scRNA-seq) enables high-resolution analysis of gene expression at the individual cell level, with clustering serving as a critical step for identifying distinct cell populations. Due to the high dimensionality and sparsity of scRNA-seq data, existing approaches typically perform gene selection prior to clustering. However, treating feature selection as a separate preprocessing step can overlook latent clustering structure and often results in suboptimal outcomes, as it does not guarantee that the selected genes are informative for clustering. To address this limitation, we propose FSSC (Feature Selection for scRNA-seq Clustering), a unified framework for joint feature selection and clustering in scRNA-seq analysis. FSSC integrates a zero-inflated negative binomial (ZINB) autoencoder with a group Lasso penalty and a dedicated clustering loss. This joint optimization enables the model to simultaneously learn low-dimensional representations and select a compact set of cluster-discriminatory genes, preserving both the statistical characteristics of scRNA-seq data and its underlying cluster structure. Extensive experiments on both simulated and real scRNA-seq datasets demonstrate that FSSC consistently outperforms state-of-the-art methods in clustering accuracy and effectively identifies a compact, biologically meaningful set of marker genes.
Siqi Jiang, Zhi Wei 0001
Briefings Bioinform.3
2025 Modeling TCR-pMHC Binding with Dual Encoders and Cross-Attention Fusion
abstract
Accurately modeling the binding between T-cell receptors (TCRs) and peptide-MHC (pMHC) complexes is essential for guiding immunotherapy development and personalized vaccine design. However, the vast diversity of TCR repertoires and the scarcity of experimentally validated interactions make generalization to unseen epitopes challenging. This paper proposes TIDE, a cross-attention-driven dual-encoder framework that leverages large protein and molecular language models to learn discriminative representations of TCRs and peptides. In TIDE, TCR sequences are encoded using Evolutionary Scale Modeling (ESM), while peptides are transformed into SMILES strings and processed by MolFormer to capture chemical and structural properties. Multi-layer cross-attention then refines and integrates these embeddings, highlighting interaction-relevant patterns without requiring explicit structural alignment. Evaluated on the TCHard benchmark under both zero-shot and few-shot settings, TIDE achieves superior predictive accuracy and robustness compared to state-of-the-art baselines such as ChemBERTa, TITAN, and NetTCR. These results demonstrate that combining pretrained language models with cross-attention fusion offers a powerful approach for TCR-pMHC binding prediction and paves the way for more reliable computational immunology applications.
Cong Qi, Zhi Wei 0001
BIBM3
2025 ESPNet: Edge-Aware Graph Representation Learning Over Analyst-Firm Bipartite Networks for Earnings Surprise Prediction
Siqi Jiang, Xinyuan Tao, Ajim Uddin, Zhi Wei 0001, Dantong Yu
IEEE Big Data4
2025 Two-stage Risk Control with Application to Ranked Retrieval
abstract
Practical machine learning systems often operate in multiple sequential stages, as seen in ranking and recommendation systems, which typically include a retrieval phase followed by a ranking phase. Effectively assessing prediction uncertainty and ensuring effective risk control in such systems pose significant challenges due to their inherent complexity. To address these challenges, we developed two-stage risk control methods based on the recently proposed learn-then-test (LTT) and conformal risk control (CRC) frameworks. Unlike the methods in prior work that address multiple risks, our approach leverages the sequential nature of the problem, resulting in reduced computational burden. We provide theoretical guarantees for our proposed methods and design novel loss functions tailored for ranked retrieval tasks. The effectiveness of our approach is validated through experiments on two large-scale, widely-used datasets: MSLR-Web and Yahoo LTRC.
Yunpeng Xu, Mufang Ying, Wenge Guo, Zhi Wei 0001
IJCAI4
2025 scDILT: A Model-Based and Constrained Deep Learning Framework for Single-Cell Data Integration, Label Transferring, and Clustering
abstract
The scRNA-seq technology enables high-resolution profiling and analysis of individual cells. The increasing availability of datasets and advancements in technology have prompted researchers to integrate existing annotated datasets with newly sequenced datasets for a more comprehensive analysis. It is important to ensure that the integration of new datasets does not alter the cell clusters defined in the old/reference datasets. Although several methods have been developed for scRNA-seq data integration, there is currently a lack of tools that can simultaneously achieve the aforementioned objectives. Therefore, in this study, we have introduced a novel tool called scDILT, which leverages a conditional autoencoder and deep embedding clustering to effectively remove batch effects among different datasets. Moreover, scDILT utilizes homogeneous constraints to preserve the cell-type/clustering patterns observed in the reference datasets, while employing heterogeneous constraints to map cells in the new datasets to the annotated cell clusters in the reference datasets. We have conducted extensive experiments to demonstrate that scDILT outperforms other methods in terms of data integration, as confirmed by evaluations on simulated and real datasets. Furthermore, we have shown that scDILT can be successfully applied to integrate multi-omics single-cell datasets. Based on these findings, we conclude that scDILT holds great promise as a tool for integrating single-cell datasets derived from different batches, experiments, times, or interventions.
Jianlan Ren, Junwen Wang, Zhi Wei 0001
IEEE Trans. Comput. Biol. Bioinform.5
2024 MultiSC: a deep learning pipeline for analyzing multiomics single-cell data
abstract
Single-cell technologies enable researchers to investigate cell functions at an individual cell level and study cellular processes with higher resolution. Several multi-omics single-cell sequencing techniques have been developed to explore various aspects of cellular behavior. Using NEAT-seq as an example, this method simultaneously obtains three kinds of omics data for each cell: gene expression, chromatin accessibility, and protein expression of transcription factors (TFs). Consequently, NEAT-seq offers a more comprehensive understanding of cellular activities in multiple modalities. However, there is a lack of tools available for effectively integrating the three types of omics data. To address this gap, we propose a novel pipeline called MultiSC for the analysis of MULTIomic Single-Cell data. Our pipeline leverages a multimodal constraint autoencoder (single-cell hierarchical constraint autoencoder) to integrate the multi-omics data during the clustering process and a matrix factorization-based model (scMF) to predict target genes regulated by a TF. Moreover, we utilize multivariate linear regression models to predict gene regulatory networks from the multi-omics data. Additional functionalities, including differential expression, mediation analysis, and causal inference, are also incorporated into the MultiSC pipeline. Extensive experiments were conducted to evaluate the performance of MultiSC. The results demonstrate that our pipeline enables researchers to gain a comprehensive view of cell activities and gene regulatory networks by fully leveraging the potential of multiomics single-cell data. By employing MultiSC, researchers can effectively integrate and analyze diverse omics data types, enhancing their understanding of cellular processes.
Siqi Jiang, Zhi Wei 0001, Junwen Wang
Briefings Bioinform.4
2023 Conformal Risk Control for Ordinal Classification
abstract
As a natural extension to the standard conformal prediction method, several conformal risk control methods have been recently developed and applied to various learning problems. In this work, we seek to control the conformal risk in expectation for ordinal classification tasks, which have broad applications to many real problems. For this purpose, we firstly formulated the ordinal classification task in the conformal risk control framework, and provided theoretic risk bounds of the risk control method. Then we proposed two types of loss functions specially designed for ordinal classification tasks, and developed corresponding algorithms to determine the prediction set for each case to control their risks at a desired level. We demonstrated the effectiveness of our proposed methods, and analyzed the difference between the two types of risks on three different datasets, including a simulated dataset, the UTKFace dataset and the diabetic retinopathy detection dataset.
Yunpeng Xu, Wenge Guo, Zhi Wei 0001
UAI3
2023 Hidden Markov random field models for cell-type assignment of spatially resolved transcriptomics
abstract
MOTIVATION: The recent development of spatially resolved transcriptomics (SRT) technologies has facilitated research on gene expression in the spatial context. Annotating cell types is one crucial step for downstream analysis. However, many existing algorithms use an unsupervised strategy to assign cell types for SRT data. They first conduct clustering analysis and then aggregate cluster-level expression based on the clustering results. This workflow fails to leverage the marker gene information efficiently. On the other hand, other cell annotation methods designed for single-cell RNA-seq data utilize the cell-type marker genes information but fail to use spatial information in SRT data. RESULTS: We introduce a statistical spatial transcriptomics cell assignment model, SPAN, to annotate clusters of cells or spots into known types in SRT data with prior knowledge of predefined marker genes and spatial information. The SPAN model annotates cells or spots from SRT data using predefined overexpressed marker genes and combines a mixture model with a hidden Markov random field to model the spatial dependency between neighboring spots. We demonstrate the effectiveness of SPAN against spatial and nonspatial clustering algorithms through extensive simulation and real data experiments. AVAILABILITY AND IMPLEMENTATION: https://github.com/ChengZ352/SPAN.
Zhi Wei 0001
Bioinform.3
2023 MGEL: Multigrained Representation Analysis and Ensemble Learning for Text Moderation
abstract
In this work, we describe our efforts in addressing two typical challenges involved in the popular text classification methods when they are applied to text moderation: the representation of multibyte characters and word obfuscations. Specifically, a multihot byte-level scheme is developed to significantly reduce the dimension of one-hot character-level encoding caused by the multiplicity of instance-scarce non-ASCII characters. In addition, we introduce a simple yet effective weighting approach for fusing n-gram features to empower the classical logistic regression. Surprisingly, it outperforms well-tuned representative neural networks greatly. As a continual effort toward text moderation, we endeavor to analyze the current state-of-the-art (SOTA) algorithm bidirectional encoder representations from transformers (BERT), which works well in context understanding but performs poorly on intentional word obfuscations. To resolve this crux, we then develop an enhanced variant and remedy this drawback by integrating byte and character decomposition. It advances the SOTA performance on the largest abusive language datasets as demonstrated by our comprehensive experiments. Our work offers a feasible and effective framework to tackle word obfuscations.
Fei Tan 0002, Changwei Hu, Yifan Hu 0001, Kevin Yen, Zhi Wei 0001, Aasish Pappu, Se Rim Park, Keqian Li
IEEE Trans. Neural Networks Learn. Syst.5
2022 Fa-Mb-ResNet for Grounding Fault Identification and Line Selection in the Distribution Networks
abstract
Accurate and fast identification of the fault types and the fault feeders can improve the distribution networks’ power supply reliability. This article focuses on two issues of classifiers in performing fault identification and line selection of the distribution networks, namely, the low utilization rate of fault information and the insufficient accuracy. We propose to use multilabel and multiclassification and build a fast-multibranch residual network (Fa-Mb-ResNet) to accomplish the identification and line selection of the distribution network grounding fault simultaneously. Our work has the following contributions. First, we propose a method of frequency division and time division for learning the features of the time–frequency matrix based on wavelet transformation. Second, we propose an improved residual unit (IRU) structure, which employs different small branches and convolution kernels to achieve the fusion of abstract fault feature information in different dimensions and enhance learning efficiency. Finally, the IRU structure is connected end to end. The new approach fully exploits the side fault information. Our extensive experiments show that the Fa-Mb-ResNet is faster, more adaptable, and has better anti-interference than the state-of-the-art methods in fault identification and line selection of the distribution network.
Liulin Yang, Zhi Wei 0001
IEEE Internet Things J.3
2021 DeepCNV: a deep learning approach for authenticating copy number variations
abstract
Copy number variations (CNVs) are an important class of variations contributing to the pathogenesis of many disease phenotypes. Detecting CNVs from genomic data remains difficult, and the most currently applied methods suffer from an unacceptably high false positive rate. A common practice is to have human experts manually review original CNV calls for filtering false positives before further downstream analysis or experimental validation. Here, we propose DeepCNV, a deep learning-based tool, intended to replace human experts when validating CNV calls, focusing on the calls made by one of the most accurate CNV callers, PennCNV. The sophistication of the deep neural network algorithm is enriched with over 10 000 expert-scored samples that are split into training and testing sets. Variant confidence, especially for CNVs, is a main roadblock impeding the progress of linking CNVs with the disease. We show that DeepCNV adds to the confidence of the CNV calls with an optimal area under the receiver operating characteristic curve of 0.909, exceeding other machine learning methods. The superiority of DeepCNV was also benchmarked and confirmed using an experimental wet-lab validation dataset. We conclude that the improvement obtained by DeepCNV results in significantly fewer false positive results and failures to replicate the CNV association results.
Joseph Glessner, Xiurui Hou, Jie Zhang 0049, Munir Khan, Fabian Brand, Peter M. Krawitz, Patrick Sleiman, Hakon Hakonarson, Zhi Wei 0001
Briefings Bioinform.10
2020 DeepVar: An End-to-End Deep Learning Approach for Genomic Variant Recognition in Biomedical Literature
abstract
We consider the problem of Named Entity Recognition (NER) on biomedical scientific literature, and more specifically the genomic variants recognition in this work. Significant success has been achieved for NER on canonical tasks in recent years where large data sets are generally available. However, it remains a challenging problem on many domain-specific areas, especially the domains where only small gold annotations can be obtained. In addition, genomic variant entities exhibit diverse linguistic heterogeneity, differing much from those that have been characterized in existing canonical NER tasks. The state-of-the-art machine learning approaches heavily rely on arduous feature engineering to characterize those unique patterns. In this work, we present the first successful end-to-end deep learning approach to bridge the gap between generic NER algorithms and low-resource applications through genomic variants recognition. Our proposed model can result in promising performance without any hand-crafted features or post-processing rules. Our extensive experiments and results may shed light on other similar low-resource NER applications.
Chaoran Cheng, Fei Tan 0002, Zhi Wei 0001
AAAI3
2020 AirNet: A Calibration Model for Low-Cost Air Monitoring Sensors Using Dual Sequence Encoder Networks
abstract
Air pollution monitoring has attracted much attention in recent years. However, accurate and high-resolution monitoring of atmospheric pollution remains challenging. There are two types of devices for air pollution monitoring, i.e., static stations and mobile stations. Static stations can provide accurate pollution measurements but their spatial distribution is sparse because of their high expense. In contrast, mobile stations offer an effective solution for dense placement by utilizing low-cost air monitoring sensors, whereas their measurements are less accurate. In this work, we propose a data-driven model based on deep neural networks, referred to as AirNet, for calibrating low-cost air monitoring sensors. Unlike traditional methods, which treat the calibration task as a point-to-point regression problem, we model it as a sequence-to-point mapping problem by introducing historical data sequences from both a mobile station (to be calibrated) and the referred static station. Specifically, AirNet first extracts an observation trend feature of the mobile station and a reference trend feature of the static station via dual encoder neural networks. Then, a social-based guidance mechanism is designed to select periodic and adjacent features. Finally, the features are fused and fed into a decoder to obtain a calibrated measurement. We evaluate the proposed method on two real-world datasets and compare it with six baselines. The experimental results demonstrate that our method yields the best performance.
Haomin Yu, Qingyong Li, Zhi Wei 0001
AAAI5
2019 User Response Driven Content Understanding with Causal Inference
abstract
Content understanding with many potential industrial applications, is spurring interest by researchers in many areas in artificial intelligence. We propose to revisit the content understanding problem in digital marketing from three novel perspectives. First, our problem is to explore the way how user experience is delivered with divergent key multimedia elements. Second, we treat understanding as to elucidate their causal implications in driving user responses. Third, we propose to understand content based on observational audience visit logs. To approach this problem, we measure and generate heterogeneous content features and model them as binary, multivalued or continuous genres. Multiple key performance indicators (KPIs) are introduced to quantify user responses. We then develop a flexible and adaptive doubly robust estimator to identify the causality between these features and user responses from observational data. The comprehensive experiments are performed on real-world data sets. We show that the further analysis of the experimental results can shed actionable insights on how to improve KPIs. Our work will benefit content distribution and optimization in digital marketing.
Fei Tan 0002, Zhi Wei 0001, Abhishek Pani, Zhenyu Yan 0001
ICDM2
2019 Success Prediction on Crowdfunding with Multimodal Deep Learning
abstract
We consider the problem of project success prediction on crowdfunding platforms. Despite the information in a project profile can be of different modalities such as text, images, and metadata, most existing prediction approaches leverage only the text dominated modality. Nowadays rich visual images have been utilized in more and more project profiles for attracting backers, little work has been conducted to evaluate their effects towards success prediction. Moreover, meta information has been exploited in many existing approaches for improving prediction accuracy. However, such meta information is usually limited to the dynamics after projects are posted, e.g., funding dynamics such as comments and updates. Such a requirement of using after-posting information makes both project creators and platforms not able to predict the outcome in a timely manner. In this work, we designed and evaluated advanced neural network schemes that combine information from different modalities to study the influence of sophisticated interactions among textual, visual, and metadata on project success prediction. To make pre-posting prediction possible, our approach requires only information collected from the pre-posting profile. Our extensive experimental results show that the image features could improve success prediction performance significantly, particularly for project profiles with little text information. Furthermore, we identified contributing elements.
Chaoran Cheng, Fei Tan 0002, Xiurui Hou, Zhi Wei 0001
IJCAI4
2019 Modeling and elucidation of housing price
Fei Tan 0002, Chaoran Cheng, Zhi Wei 0001
Data Min. Knowl. Discov.3
2019 A Feature Sampling Strategy for Analysis of High Dimensional Genomic Data
abstract
With the development of high throughput technology, it has become feasible and common to profile tens of thousands of gene activities simultaneously. These genomic data typically have sample size of hundreds or fewer, which is much less than the feature size (number of genes). In addition, the genes, in particular the ones from the same pathway, are often highly correlated. These issues impose a great challenge for selecting meaningful genes from a large number of (correlated) candidates in many genomic studies. Quite a few methods have been proposed to attack this challenge. Among them, regularization-based techniques, e.g., lasso, become much more appealing, because they can do model fitting and variable selection at the same time. However, the lasso regression has its known limitations. One is that the number of genes selected by the lasso couldn't exceed the number of samples. Another limitation is that, if causal genes are highly correlated, the lasso tends to select only one or few genes from them. Biologists, however, desire to identify them all. To overcome these limitations, we present here a novel, robust, and stable variable selection method. Through simulation studies and a real application to the transcriptome data, we demonstrate the superiority of the proposed method in selecting highly correlated causal genes. We also provide some theoretical justifications for this feature sampling strategy based on the mean and variance analyses.
Jie Zhang 0049, Zhigen Zhao, Kai Zhang 0001, Zhi Wei 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2019 A Deep Learning Approach to Competing Risks Representation in Peer-to-Peer Lending
abstract
Online peer-to-peer (P2P) lending is expected to benefit both investors and borrowers due to their low transaction cost and the elimination of expensive intermediaries. From the lenders' perspective, maximizing their return on investment is an ultimate goal during their decision-making procedure. In this paper, we explore and address a fundamental problem underlying such a goal: how to represent the two competing risks, charge-off and prepayment, in funded loans. We propose to model both potential risks simultaneously, which remains largely unexplored until now. We first develop a hierarchical grading framework to integrate two risks of loans both qualitatively and quantitatively. Afterward, we introduce an end-to-end deep learning approach to solve this problem by breaking it down into multiple binary classification subproblems that are amenable to both feature representation and risks learning. Particularly, we leverage deep neural networks to jointly solve these subtasks, which leads to the in-depth exploration of the interaction involved in these tasks. To the best of our knowledge, this is the first attempt to characterize competing risks for loans in P2P lending via deep neural networks. The comprehensive experiments on real-world loan data show that our methodology is able to achieve an appealing investment performance by modeling the competition within and between risks explicitly and properly. The feature analysis based on saliency maps provides useful insights into payment dynamics of loans for potential investors intuitively.
Fei Tan 0002, Xiurui Hou, Jie Zhang 0049, Zhi Wei 0001, Zhenyu Yan 0001
IEEE Trans. Neural Networks Learn. Syst.4
2019 Online Change-Point Detection in Sparse Time Series With Application to Online Advertising
abstract
Online advertising delivers promotional marketing messages to consumers through online media. Advertisers often have the desire to optimize their advertising spending strategies in order to gain the highest return on investment and maximize their key performance indicator. To build accurate advertisement performance predictive models, it is crucial to detect the change-points in the historical data and apply appropriate strategies to address a data pattern shift problem. However, with sparse data, which is common in online advertising and some other applications, online change-point detection is very challenging. We present a novel collaborated online change-point detection method in this paper. Through efficiently leveraging and coordinating with auxiliary time series, we can quickly and accurately identify the change-points in sparse and noisy time series. Simulation studies as well as real data experiments have justified the proposed method's effectiveness in detecting change-points in sparse time series. Therefore, it can be used to improve the accuracy of predictive models.
Jie Zhang 0049, Zhi Wei 0001, Zhenyu Yan 0001, MengChu Zhou, Abhishek Pani
IEEE Trans. Syst. Man Cybern. Syst.2
2018 A Blended Deep Learning Approach for Predicting User Intended Actions
abstract
User intended actions are widely seen in many areas. Forecasting these actions and taking proactive measures to optimize business outcome is a crucial step towards sustaining the steady business growth. In this work, we focus on predicting attrition, which is one of typical user intended actions. Conventional attrition predictive modeling strategies suffer a few inherent drawbacks. To overcome these limitations, we propose a novel end-to-end learning scheme to keep track of the evolution of attrition patterns for the predictive modeling. It integrates user activity logs, dynamic and static user profiles based on multi-path learning. It exploits historical user records by establishing a decaying multi-snapshot technique. And finally it employs the precedent user intentions via guiding them to the subsequent learning procedure. As a result, it addresses all disadvantages of conventional methods. We evaluate our methodology on two public data repositories and one private user usage dataset provided by Adobe Creative Cloud. The extensive experiments demonstrate that it can offer the appealing performance in comparison with several existing approaches as rated by different popular metrics. Furthermore, we introduce an advanced interpretation and visualization strategy to effectively characterize the periodicity of user activity logs. It can help to pinpoint important factors that are critical to user attrition and retention and thus suggests actionable improvement targets for business practice. Our work will provide useful insights into the prediction and elucidation of other user intended actions as well.
Fei Tan 0002, Zhi Wei 0001, Zhenyu Yan 0001
ICDM2
2018 Network Inference from Contrastive Groups Using Discriminative Structural Regularization
abstract
Gaussian graphical models (GGMs) are a popular tool for exploring conditional dependence among high dimensional data. We consider developing an estimator for GGMs for multiple graph analysis, wherein the graphs are assumed to come from two (or more) contrastive groups, and exhibit not only major global similarity, but also substantial between-group disparity. Under this setting, inferring each group of networks separately ignores the common structure, while simply assuming a global common network structure would mask the critical disparity. We propose a novel approach to pursue simultaneous network inference using discriminative and adaptive structural regularizations. We introduce a heterogeneity ratio parameter to balance the within group similarity and the between group disparity. This formulation for the first time, to our knowledge, generalizes the existing single-group network analysis to multiple-group network analysis. In other words, our proposed multiple-group network analysis reduces to single-group network analysis, when the heterogeneity ratio equal to 1. By iteratively updating a global regularization template with individual network structures, together with a feature screening module specifying relevant dimensions to satisfy the group-level constraints, our generalized approach can recover the underlying conditional independence with greater flexibility and improved accuracy. Theoretically, we show the asymptotic consistency for the proposed method in joint reconstruction of multiple network structures. We demonstrate its superior performance via extensive simulation studies. We also illustrate its practical usage in an application to polychromatic flow cytometry data sets for protein interactions under different conditions.
Ruihua Cheng, Zhi Wei 0001, Kai Zhang 0001
SDM2
2018 Modeling Item-specific Effects for Video Click
abstract
Prediction is widely employed to improve the number of video clicks and views, which are the key important indicators (KPIs) due to their contribution to revenue. The available predictive features, however, are generally limited as compared to the expected prediction capability from the algorithm side. Inspired by the intrinsic dependence among multiple clicks for the same video, we hypothesize that there exist some consistent effects involved in grouped click records. We then propose to recover such effects from the associated hidden features, which are likely to alleviate the insufficiency of features. The simulation studies are performed to elucidate how the derived grouped effects empower a model with additional discriminating capacity compared with the original one. The proposed methodology is further examined on the repository of PPTV (a leading video service provider in China) click records comprehensively. The results confirm the existence of the hypothesized effects and demonstrate their critical role in the performance improvement of video click prediction.
Fei Tan 0002, Kuang Du, Zhi Wei 0001, Chenguang Qin
SDM3
2018 An omnibus test for differential distribution analysis of microbiome sequencing data
abstract
Motivation: One objective of human microbiome studies is to identify differentially abundant microbes across biological conditions. Previous statistical methods focus on detecting the shift in the abundance and/or prevalence of the microbes and treat the dispersion (spread of the data) as a nuisance. These methods also assume that the dispersion is the same across conditions, an assumption which may not hold in presence of sample heterogeneity. Moreover, the widespread outliers in the microbiome sequencing data make existing parametric models not overly robust. Therefore, a robust and powerful method that allows covariate-dependent dispersion and addresses outliers is still needed for differential abundance analysis. Results: We introduce a novel test for differential distribution analysis of microbiome sequencing data by jointly testing the abundance, prevalence and dispersion. The test is built on a zero-inflated negative binomial regression model and winsorized count data to account for zero-inflation and outliers. Using simulated data and real microbiome sequencing datasets, we show that our test is robust across various biological conditions and overall more powerful than previous methods. Availability and implementation: R package is available at https://github.com/jchen1981/MicrobiomeDDA. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Emily King, Rebecca Deek, Zhi Wei 0001, Diane E. Grill, Karla V. Ballman
Bioinform.4
2018 A distance-based approach for testing the mediation effect of the human microbiome
abstract
Motivation: Recent studies have revealed a complex interplay between environment, the human microbiome and health and disease. Mediation analysis of the human microbiome in these complex relationships could potentially provide insights into the role of the microbiome in the etiology of disease and, more importantly, lead to novel clinical interventions by modulating the microbiome. However, due to the high dimensionality, sparsity, non-normality and phylogenetic structure of microbiome data, none of the existing methods are suitable for testing such clinically important mediation effect. Results: We propose a distance-based approach for testing the mediation effect of the human microbiome. In the framework, the nonlinear relationship between the human microbiome and independent/dependent variables is captured implicitly through the use of sample-wise ecological distances, and the phylogenetic tree information is conveniently incorporated by using phylogeny-based distance metrics. Multiple distance metrics are utilized to maximize the power to detect various types of mediation effect. Simulation studies demonstrate that our method has correct Type I error control, and is robust and powerful under various mediation models. Application to a real gut microbiome dataset revealed that the association between the dietary fiber intake and body mass index was mediated by the gut microbiome. Availability and implementation: An R package 'MedTest' is freely available at https://github.com/jchen1981/MedTest. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Jie Zhang 0049, Zhi Wei 0001
Bioinform.2
2018 A Distance-Based Weighted Undersampling Scheme for Support Vector Machines and its Application to Imbalanced Classification
abstract
A support vector machine (SVM) plays a prominent role in classic machine learning, especially classification and regression. Through its structural risk minimization, it has enjoyed a good reputation in effectively reducing overfitting, avoiding dimensional disaster, and not falling into local minima. Nevertheless, existing SVMs do not perform well when facing class imbalance and large-scale samples. Undersampling is a plausible alternative to solve imbalanced problems in some way, but suffers from soaring computational complexity and reduced accuracy because of its enormous iterations and random sampling process. To improve their classification performance in dealing with data imbalance problems, this work proposes a weighted undersampling (WU) scheme for SVM based on space geometry distance, and thus produces an improved algorithm named WU-SVM. In WU-SVM, majority samples are grouped into some subregions (SRs) and assigned different weights according to their Euclidean distance to the hyper plane. The samples in an SR with higher weight have more chance to be sampled and put to use in each learning iteration, so as to retain the data distribution information of original data sets as much as possible. Comprehensive experiments are performed to test WU-SVM via 21 binary-class and six multiclass publically available data sets. The results show that it well outperforms the state-of-the-art methods in terms of three popular metrics for imbalanced classification, i.e., area under the curve, F-Measure, and G-Mean.
Qi Kang 0001, MengChu Zhou, Xuesong Wang 0002, Qidi Wu, Zhi Wei 0001
IEEE Trans. Neural Networks Learn. Syst.6
2017 Time-Aware Latent Hierarchical Model for Predicting House Prices
abstract
It is widely acknowledged that the value of a house is the mixture of a large number of characteristics. House price prediction thus presents a unique set of challenges in practice. While a large body of works are dedicated to this task, their performance and applications have been limited by the shortage of long time span of transaction data, the absence of real-world settings and the insufficiency of housing features. To this end, a time-aware latent hierarchical model is introduced to capture underlying spatiotemporal interactions behind the evolution of house prices. The hierarchical perspective obviates the need for historical transaction data of exactly same houses when temporal effects are considered. The proposed framework is examined on a large-scale dataset of the property transaction in Beijing. The whole experimental procedure strictly complies with the real-world scenario. The empirical evaluation results demonstrate the outperformance of our approach over alternative competitive methods.
Fei Tan 0002, Chaoran Cheng, Zhi Wei 0001
ICDM3
2017 REMOLD: An Efficient Model-Based Clustering Algorithm for Large Datasets with Spark
abstract
Density-based clustering algorithms have the distinctive advantage of discovering arbitrarily shaped clusters, but they usually require a procedure to compute the distance between every pair of data points, and this procedure is prohibitive for large datasets since it has quadratic computation complexity. In this paper, we propose a new distributed clustering algorithm, named REstore MOdel with Local Density estimation (REMOLD). Firstly, REMODL applies a balanced partitioning method to evenly divide an large dataset based on Local Sensitive Hashing (LSH). Then, it locally clusters each partition of the dataset, and uses a Gaussian model to represent each local cluster based on the observation that the density distribution of each local cluster shares similar shape with Gaussian distribution. Finally, these models are aggregated on a server where REMOLD restores global clusters based on these local Gaussian models. More specifically, model connection, which measures the density connectivity between two models, are defined to merge local models with an optimized procedure. In this aggregation, REMOLD requires low cost of network transmission for local Gaussian models, since the number of Gaussian models is often less than that of core objects for each partition. We evaluate REMOLD on three synthetic datasets and three real-world datasets on Spark, and the experiment results demonstrate that REMOLD is efficient and effective to find out clusters with complex shapes and it outperforms the established methods.
Mingfei Liang, Qingyong Li, Jianzhu Wang, Zhi Wei 0001
ICPADS5
2016 Modeling Real Estate for School District Identification
abstract
The affiliated school district of a real estate property is often a crucial concern. How to automate the identification of residential homes located in a favorable educational environment, however, is largely unexplored until now. The availability of heterogeneous estate-related data offers a great opportunity for this task. Nevertheless, it is such heterogeneity that poses significant challenges to their amalgamation in a unified fashion. To this end, we develop G-LRMM model to integrate digital price, textual comments, and geographical location information together. The proposed approach is able to capture the in-depth interaction among multi-type data greatly. The evaluation on the dataset of Beijing property market justifies the benefits of our approach over baselines. The further comparison among different components is also conducted and demonstrates their important roles. Moreover, the proposed model can offer useful insights into modeling heterogeneous data sources.
Fei Tan 0002, Chaoran Cheng, Zhi Wei 0001
ICDM3
2016 Reliable Gender Prediction Based on Users' Video Viewing Behavior
abstract
With the growth of the digital advertising market, it has become more important than ever to target the desired audiences. Among various demographic traits, gender information plays a key role in precisely targeting the potential consumers in online advertising and ecommerce. However, such personal information is generally unavailable to digital media sellers. In this paper, we propose a novel task-specific multi-task learning algorithm to predict users' gender information from their video viewing behaviors. To detect as many desired users as possible, while controlling the Type I error rate at a user-specified level, we further propose Bayes testing and decision procedures to efficiently identify male and female users, respectively. Comprehensive experiments have justified the effectiveness and reliability of our framework.
Jie Zhang 0049, Kuang Du, Ruihua Cheng, Zhi Wei 0001, Chenguang Qin, Huaxin You
ICDM4
2016 Annealed Sparsity via Adaptive and Dynamic Shrinking
abstract
Sparse learning has received tremendous amount of interest in high-dimensional data analysis due to its model interpretability and the low-computational cost. Among the various techniques, adaptive l1-regularization is an effective framework to improve the convergence behaviour of the LASSO, by using varying strength of regularization across different features. In the meantime, the adaptive structure makes it very powerful in modelling grouped sparsity patterns as well, being particularly useful in high-dimensional multi-task problems. However, choosing an appropriate, global regularization weight is still an open problem. In this paper, inspired by the annealing technique in material science, we propose to achieve "annealed sparsity" by designing a dynamic shrinking scheme that simultaneously optimizes the regularization weights and model coefficients in sparse (multi-task) learning. The dynamic structures of our algorithm are twofold. Feature-wise (spatially), the regularization weights are updated interactively with model coefficients, allowing us to improve the global regularization structure. Iteration-wise (temporally), such interaction is coupled with gradually boosted l1-regularization by adjusting an equality norm-constraint, achieving an annealing effect to further improve model selection. This renders interesting shrinking behaviour in the whole solution path. Our method competes favorably with state-of-the-art methods in sparse (multi-task) learning. We also apply it in expression quantitative trait loci analysis (eQTL), which gives useful biological insights in human cancer (melanoma) study.
Kai Zhang 0001, Shandian Zhe, Chaoran Cheng, Zhi Wei 0001, Zhengzhang Chen, Guofei Jiang, Yuan Qi 0001, Jieping Ye
KDD4
2016 An empirical Bayes change-point model for identifying 3′ and 5′ alternative splicing by next-generation RNA sequencing
abstract
MOTIVATION: Next-generation RNA sequencing (RNA-seq) has been widely used to investigate alternative isoform regulations. Among them, alternative 3 ': splice site (SS) and 5 ': SS account for more than 30% of all alternative splicing (AS) events in higher eukaryotes. Recent studies have revealed that they play important roles in building complex organisms and have a critical impact on biological functions which could cause disease. Quite a few analytical methods have been developed to facilitate alternative 3 ': SS and 5 ': SS studies using RNA-seq data. However, these methods have various limitations and their performances may be further improved. RESULTS: We propose an empirical Bayes change-point model to identify alternative 3 ': SS and 5 ': SS. Compared with previous methods, our approach has several unique merits. First of all, our model does not rely on annotation information. Instead, it provides for the first time a systematic framework to integrate various information when available, in particular the useful junction read information, in order to obtain better performance. Second, we utilize an empirical Bayes model to efficiently pool information across genes to improve detection efficiency. Third, we provide a flexible testing framework in which the user can choose to address different levels of questions, namely, whether alternative 3 ': SS or 5 ': SS happens, and/or where it happens. Simulation studies and real data application have demonstrated that our method is powerful and accurate. AVAILABILITY AND IMPLEMENTATION: The software is implemented in Java and can be freely downloaded from http://ebchangepoint.sourceforge.net/ CONTACT: [email protected].
Jie Zhang 0049, Zhi Wei 0001
Bioinform.2
2015 Collaborated Online Change-Point Detection in Sparse Time Series for Online Advertising
abstract
Online advertising delivers promotional marketing messages to consumers through online media. Advertisers often have the desire to optimize their advertising spending and strategies in order to maximize their KPI (Key performance indicator). To build accurate ad performance predictive models, it is crucial to detect the change-points in historical data and therefore apply appropriate strategies to address the data pattern shift. However, with sparse data, which is common in online advertising, online change-point detection often becomes challenging. We propose a novel collaborated online change-point detection method in this paper. Through efficiently leveraging and coordinating with auxiliary time series, it can quickly and accurately identify the change-points in sparse and noisy time series. Simulation studies as well as real data applications have demonstrated its effectiveness in detecting change-point in sparse time series and therefore improving the accuracy of predictive models.
Jie Zhang 0049, Zhi Wei 0001, Zhenyu Yan 0001, Abhishek Pani
ICDM2
2015 Scalable quality assurance for large SNOMED CT hierarchies using subject-based subtaxonomies
abstract
OBJECTIVE: Standards terminologies may be large and complex, making their quality assurance challenging. Some terminology quality assurance (TQA) methodologies are based on abstraction networks (AbNs), compact terminology summaries. We have tested AbNs and the performance of related TQA methodologies on small terminology hierarchies. However, some standards terminologies, for example, SNOMED, are composed of very large hierarchies. Scaling AbN TQA techniques to such hierarchies poses a significant challenge. We present a scalable subject-based approach for AbN TQA. METHODS: An innovative technique is presented for scaling TQA by creating a new kind of subject-based AbN called a subtaxonomy for large hierarchies. New hypotheses about concentrations of erroneous concepts within the AbN are introduced to guide scalable TQA. RESULTS: We test the TQA methodology for a subject-based subtaxonomy for the Bleeding subhierarchy in SNOMED's large Clinical finding hierarchy. To test the error concentration hypotheses, three domain experts reviewed a sample of 300 concepts. A consensus-based evaluation identified 87 erroneous concepts. The subtaxonomy-based TQA methodology was shown to uncover statistically significantly more erroneous concepts when compared to a control sample. DISCUSSION: The scalability of TQA methodologies is a challenge for large standards systems like SNOMED. We demonstrated innovative subject-based TQA techniques by identifying groups of concepts with a higher likelihood of having errors within the subtaxonomy. Scalability is achieved by reviewing a large hierarchy by subject. CONCLUSIONS: An innovative methodology for scaling the derivation of AbNs and a TQA methodology was shown to perform successfully for the largest hierarchy of SNOMED.
Christopher Ochs, James Geller, Yehoshua Perl, Yan Chen 0009, Junchuan Xu, Hua Min, James T. Case, Zhi Wei 0001
J. Am. Medical Informatics Assoc.8
2014 A change-point model for identifying 3′UTR switching by next-generation RNA sequencing
abstract
MOTIVATION: Next-generation RNA sequencing offers an opportunity to investigate transcriptome in an unprecedented scale. Recent studies have revealed widespread alternative polyadenylation (polyA) in eukaryotes, leading to various mRNA isoforms differing in their 3' untranslated regions (3'UTR), through which, the stability, localization and translation of mRNA can be regulated. However, very few, if any, methods and tools are available for directly analyzing this special alternative RNA processing event. Conventional methods rely on annotation of polyA sites; yet, such knowledge remains incomplete, and identification of polyA sites is still challenging. The goal of this article is to develop methods for detecting 3'UTR switching without any prior knowledge of polyA annotations. RESULTS: We propose a change-point model based on a likelihood ratio test for detecting 3'UTR switching. We develop a directional testing procedure for identifying dramatic shortening or lengthening events in 3'UTR, while controlling mixed directional false discovery rate at a nominal level. To our knowledge, this is the first approach to analyze 3'UTR switching directly without relying on any polyA annotations. Simulation studies and applications to two real datasets reveal that our proposed method is powerful, accurate and feasible for the analysis of next-generation RNA sequencing data. CONCLUSIONS: The proposed method will fill a void among alternative RNA processing analysis tools for transcriptome studies. It can help to obtain additional insights from RNA sequencing data by understanding gene regulation mechanisms through the analysis of 3'UTR switching. AVAILABILITY AND IMPLEMENTATION: The software is implemented in Java and can be freely downloaded from http://utr.sourceforge.net/. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wei Wang 0080, Zhi Wei 0001, Hongzhe Li
Bioinform.2
2009 Multiple testing in genome-wide association studies via hidden Markov models
abstract
MOTIVATION: Genome-wide association studies (GWAS) interrogate common genetic variation across the entire human genome in an unbiased manner and hold promise in identifying genetic variants with moderate or weak effect sizes. However, conventional testing procedures, which are mostly P-value based, ignore the dependency and therefore suffer from loss of efficiency. The goal of this article is to exploit the dependency information among adjacent single nucleotide polymorphisms (SNPs) to improve the screening efficiency in GWAS. RESULTS: We propose to model the linear block dependency in the SNP data using hidden Markov models (HMMs). A compound decision-theoretic framework for testing HMM-dependent hypotheses is developed. We propose a powerful data-driven procedure [pooled local index of significance (PLIS)] that controls the false discovery rate (FDR) at the nominal level. PLIS is shown to be optimal in the sense that it has the smallest false negative rate (FNR) among all valid FDR procedures. By re-ranking significance for all SNPs with dependency considered, PLIS gains higher power than conventional P-value based methods. Simulation results demonstrate that PLIS dominates conventional FDR procedures in detecting disease-associated SNPs. Our method is applied to analysis of the SNP data from a GWAS of type 1 diabetes. Compared with the Benjamini-Hochberg (BH) procedure, PLIS yields more accurate results and has better reproducibility of findings. CONCLUSION: The genomic rankings based on our procedure are substantially different from the rankings based on the P-values. By integrating information from adjacent locations, the PLIS rankings benefit from the increased signal-to-noise ratio, hence our procedure often has higher statistical power and better reproducibility. It provides a promising direction in large-scale GWAS. AVAILABILITY: An R package PLIS has been developed to implement the PLIS procedure. Source codes are available upon request and will be available on CRAN (http://cran.r-project.org/). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhi Wei 0001, Wenguang Sun, Kai Wang 0049, Hakon Hakonarson
Bioinform.1
2008 A spatially varying two-sample recombinant coalescent, with applications to HIV escape response
abstract
Statistical evolutionary models provide an important mechanism for describing and understanding the escape response of a viral population under a particular therapy. We present a new hierarchical model that incorporates spatially varying mutation and recombination rates at the nucleotide level. It also maintains sep- arate parameters for treatment and control groups, which allows us to estimate treatment effects explicitly. We use the model to investigate the sequence evolu- tion of HIV populations exposed to a recently developed antisense gene therapy, as well as a more conventional drug therapy. The detection of biologically rele- vant and plausible signals in both therapy studies demonstrates the effectiveness of the method.
Alexander Braunstein, Zhi Wei 0001, Shane T. Jensen, Jon D. McAuliffe
NIPS2
2007 A Markov random field model for network-based analysis of genomic data
abstract
MOTIVATION: A central problem in genomic research is the identification of genes and pathways involved in diseases and other biological processes. The genes identified or the univariate test statistics are often linked to known biological pathways through gene set enrichment analysis in order to identify the pathways involved. However, most of the procedures for identifying differentially expressed (DE) genes do not utilize the known pathway information in the phase of identifying such genes. In this article, we develop a Markov random field (MRF)-based method for identifying genes and subnetworks that are related to diseases. Such a procedure models the dependency of the DE patterns of genes on the networks using a local discrete MRF model. RESULTS: Simulation studies indicated that the method is quite effective in identifying genes and subnetworks that are related to disease and has higher sensitivity and lower false discovery rates than the commonly used procedures that do not use the pathway structure information. Applications to two breast cancer microarray gene expression datasets identified several subnetworks on several of the KEGG transcriptional pathways that are related to breast cancer recurrence or survival due to breast cancer. CONCLUSIONS: The proposed MRF-based model efficiently utilizes the known pathway structures in identifying the DE genes and the subnetworks that might be related to phenotype. As more biological networks are identified and documented in databases, the proposed method should find more applications in identifying the subnetworks that are related to diseases and other biological processes.
Zhi Wei 0001, Hongzhe Li
Bioinform.1
2006 GAME: detecting cis-regulatory elements using a genetic algorithm
abstract
MOTIVATION: Identification of a transcription factor binding sites is an important aspect of the analysis of genetic regulation. Many programs have been developed for the de novo discovery of a binding motif (collection of binding sites). Recently, a scoring function formulation was derived that allows for the comparison of discovered motifs from different programs [S.T. Jensen, X.S. Liu, Q. Zhou and J.S. Liu (2004) Stat. Sci., 19, 188-204.] A simple program, BioOptimizer, was proposed in [S.T. Jensen and J.S. Liu (2004) Bioinformatics, 20, 1557-1564.] that improved discovered motifs by optimizing a scoring function. However, BioOptimizer is a very simple algorithm that can only make local improvements upon an already discovered motif and so BioOptimizer can only be used in conjunction with other motif-finding software. RESULTS: We introduce software, GAME, which utilizes a genetic algorithm to find optimal motifs in DNA sequences. GAME evolves motifs with high fitness from a population of randomly generated starting motifs, which eliminate the reliance on additional motif-finding programs. In addition to using standard genetic operations, GAME also incorporates two additional operators that are specific to the motif discovery problem. We demonstrate the superior performance of GAME compared with MEME, BioProspector and BioOptimizer in simulation studies as well as several real data applications where we use an extended version of the GAME algorithm that allows the motif width to be unknown.
Zhi Wei 0001, Shane T. Jensen
Bioinform.1