VLDB 2026 Research / reviewers in the wild / expert
Chengxin He
dblp:197/5322
· DBLP profile ↗
21ranked-venue papers
4as first author
17since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MVIC: Multi-view Information Collaborative Fusion for Drug-Drug Interaction Prediction
Xianxian Zhao, Chengxin He, Lei Duan |
DASFAA (2) | 2 |
| 2025 | Learning Decision Boundaries for Multidimensional Anomaly DetectionabstractIn this article, we attempt to study a problem of identifying outliers from multiple dimensions, terming it as multidimensional anomaly detection. This task assumes that each sample corresponds to heterogeneous discriminative spaces, where each space characterizes distinct semantic information along one dimension. Consequently, the abnormalities exhibited by a sample will vary under these semantically different dimensions. In contrast to traditional anomaly detection, multidimensional anomaly detection offers a more holistic assessment of a sample's abnormalities. However, the heterogeneity of discriminative spaces leads to incomparability of the outputs from different dimensions which is the major difficulty in designing multidimensional anomaly detection methods. This article introduces a novel model, maximum margin multidimensional anomaly detection (ALOE), specifically tailored for multidimensional anomaly detection. ALOE constructs a convex optimization problem with nonlinear constraints. The primary objective is to simultaneously learn multiple decision boundaries, utilizing the maximum margin principle and covariance regularization, while distinguishing between outliers and normal samples under multiple dimensions by capturing the correlation among multiple dimensions. To obtain the optimal decision boundary under each dimension, we devise an alternating optimization method for this convex optimization problem. To validate the effectiveness of ALOE, we conduct extensive experiments on 12 real-world datasets, comparing its performance against 34 anomaly detection methods. The experimental results demonstrate the superior performance of ALOE. Xinye Wang, Lei Duan, Lili Guan, Jiaxuan Xu 0001, Chengxin He |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | SCOP: A Sequence-Structure Contrast-Aware Framework for Protein Function PredictionabstractImproving the ability to predict protein function can potentially facilitate research in the fields of drug discovery and precision medicine. Technically, the properties of proteins are directly or indirectly reflected in their sequence and structure information, especially as the protein function is largely determined by its spatial properties. Existing approaches mostly focus on protein sequences or topological structures, while rarely exploiting the spatial properties and ignoring the relevance between sequence and structure information. Moreover, obtaining annotated data to improve protein function prediction is often time-consuming and costly. To this end, this work proposes a novel contrast-aware pre-training framework, called SCOP, for protein function prediction. We first design a simple yet effective encoder to integrate the protein topological and spatial features under the structure view. Then a convolutional neural network is utilized to learn the protein features under the sequence view. Finally, we pretrain SCOP by leveraging two types of auxiliary supervision to explore the relevance between these two views and thus extract informative representations to better predict protein function. Experimental results on four benchmark datasets and one self-built dataset demonstrate that SCOP provides more specific results, while using less pre-training data. Chengxin He, Huiru Zheng, Xinye Wang, Yidan Zhang 0001, Lei Duan |
BIBM | 2 |
| 2024 | Towards Better Zero-Shot Anomaly Detection under Distribution Shift with CLIP
Jiyao Gao, Chengxin He, Lei Duan, Jie Zuo |
BMVC | 2 |
| 2024 | Community-Guided Contrastive Learning with Anomaly-Aware Reconstruction for Anomaly Detection on Attributed Networks
Xinye Wang, Chengxin He, Xiaocong Chen, Zhaohang Luo, Lei Duan, Jie Zuo |
DASFAA (7) | 3 |
| 2024 | An Efficient Adaptive Multi-Kernel Learning With Safe Screening Rule for Outlier DetectionabstractRecent advances in multi-kernel-based methods for outlier detection have positioned them as an attractive way to detect instances that are markedly different from the remaining data in a dataset. Currently, most outlier detection approaches based on multi-kernel learning are simply a convex combination of various kernels with handcrafted weights, meaning that these weights may not be suitable. Meanwhile, this combination of weights does not sufficiently consider the intrinsic correlations of instances when fusing different kernels. Thus, a key challenge is how to adaptively learn an appropriate combination of weights for capturing a new feature space in which outliers can be better detected than the original space. Simultaneously, it is still a burning issue to get the optimal combination of weights due to considerable computational cost and memory usage when the feature or instance size is large. In this paper, we propose a novel method forefficientadaptivemulti-kernel foroutlierdetection (EAMOD), which automatically learns the optimal weight for each training instance under different kernels using a non-negative function. In addition, we design a safe screening rule (SSR) for EAMOD to improve its training efficiency without any loss of accuracy. To the best of our knowledge, it is the first attempt to develop SSR for multi-kernel-based outlier detection methods. Extensive experiments show that EAMOD is effective and efficient. Xinye Wang, Lei Duan, Chengxin He, Yuanyuan Chen 0006, Xindong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Robust Multi-Kernel Nearest Neighborhood for Outlier DetectionabstractOutlier detection methods based on distance measure have been used in numerous applications due to their effectiveness and interpretability. However, distances among instances heavily depend on the feature space in which they reside. For an outlier, distances from it to the normal instances may be extremely close in one feature space, failing to separate them from each other, while this situation is reversed in another space. Meanwhile, the distance measure is sensitive to a few “marginal instances” (i.e., normal instances located very close to outliers in the feature space) during the estimation of whether a test instance is an outlier or not. In this paper, we propose a robust multi-kernel nearest neighborhood (RMKN) method for outlier detection. Specifically, in the training phase, we only consider normal instances and transform them into a Polynomial kernel function weighted digraph to capture their geometric relationships in the original feature space. Then, we develop an objective function based on the weighted digraph to find a latent feature space via multi-kernel learning such that distances among normal instances in this latent feature space are as close as possible while preserving their original distributions. In the detecting phase, we design an outlying score based on the two-stage multi-kernel k-nearest nearest neighbors to detect outliers. Extensive experiments with ten datasets show that RMKN is effective and robust Xinye Wang, Lei Duan, Zhenyang Yu, Chengxin He, Zhifeng Bao |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | SIDE: Sequence-Interaction-Aware Dual Encoder for Predicting circRNA Back-Splicing EventsabstractCircular RNAs (circRNAs) play a critical role in gene regulation and association with diseases due to their specialized structure, which is formed as a closed loop structure during a non-canonical splicing process where the donor site back-spliced to an upstream acceptor site. As fundamental work to clarify their functions and mechanisms, a large number of computational methods for predicting circRNA formation have been proposed, among which, in particular, deep learning is utilized to capture relevant patterns from raw RNA sequences and model their interactions to facilitate prediction. However, these methods fail to fully utilize the important characteristics of back-splicing events, i.e., the positional information of the splice sites and the interaction features of its flanking sequences, for prediction. To this end, we hereby propose a novel approach called SIDE for predicting circRNA back-splicing events using only nucleotide sequences. Our model employs a dual encoder to capture global and interactive features of the sequence, and then a decoder designed by the contrastive learning to fuse out discriminative features improving the prediction of circRNAs formation. Empirical results on three real-world datasets have shown the effectiveness of SIDE. Our code is publicly available at https://github.com/scu-kdde/Bioinfo-SIDE-2023. Chengxin He, Lei Duan, Huiru Zheng, Yuening Qu, Zhenyang Yu |
BIBM | 1 |
| 2023 | DARE: Sequence-Structure Dual-Aware Encoder for RNA-Protein Binding PredictionabstractPredicting RNA-protein binding sites helps to explore the mechanisms of the interaction between RNA and proteins. Numerous deep learning methods have been applied to predict RNA-protein binding sites. Some of these methods use only sequence information for prediction which could lose information about the topology. And there may be a loss of important information if the secondary structure features are simply represented as one-hot matrices. Furthermore, existing deep learning methods are usually based on convolutional neural networks for feature extraction, which tend to focus on local features. As for the information of the whole sequence, existing methods usually ignore global features. Therefore, we propose a novel deep learning model called DARE for RNA-protein binding sites prediction using both sequence and secondary structure information of RNA. DARE employs the secondary structure feature extraction module to capture the features of the RNA secondary structure and learn the topological information. Therefore, we design a local feature extraction module and a global feature integration module to capture the whole information of RNA. Thus we can achieve the purpose of complementary information. Extensive experiments demonstrate that DARE outperforms baselines. Our analysis of the case study further confirm the effectiveness of DARE. Luhan Shen, Chengxin He, Haiying Wang 0001, Yuening Qu, Lei Duan |
BIBM | 2 |
| 2023 | MGDTI: Graph Transformer with Meta-Learning for Drug-Target Interaction PredictionabstractDrug-target interaction (DTI) prediction is of great importance for drug discovery and development. With the rapid development of biological and chemical technologies, computational methods for DTI prediction are becoming a promising strategy. However, there are few methods which explore solving the cold-start problem in DTI prediction scenarios due to most of existing methods require modeling under the existing interaction that can’t effectively capture information from new drugs and new targets which have few interactions in existing literature. In this paper, we propose a graph transformer method based on meta-learning named MGDTI to fill the gap. In particular, we employ drug-drug similarity and target-target similarity as additional information for network to mitigate the scarcity of interactions. Besides, we trained our model via meta-learning to be adaptive to cold-start tasks. Moreover, we introduced graph transformer to prevent over-smoothing by capturing long-range dependencies. Comparison results on the benchmark dataset demonstrate that our proposed MGDTI is effective in the DTI prediction. Chengxin He, Yuening Qu, Huiru Zheng, Lei Duan, Jie Zuo |
BIBM | 2 |
| 2023 | Enhancing GNN-based Fraud Detector via Semantic Extraction and Max-Representation-MarginabstractFraud detection aims to identify fraudsters from normal users. In graph environments, both fraudsters and normal users are modeled as nodes, while edges represent the connections between them. However, fraudulent nodes in the real world often camouflage themselves by establishing numerous fake connections with normal nodes, making them challenging to be identified. Existing fraud detection methods struggle to address this issue, they utilize graph neural networks to aggregate normal informations from normal neighbors, which leads to the smoothing of the fraudulent information. Furthermore, these methods exhibit poor generalization performance as they are unable to detect new fraudsters which not present in the training process. To overcome these limitations, this paper proposes GFAN, a novel model based on Graph Feature enhAncement Network. Specifically, GFAN introduces a specific semantic extraction module to screen and delete fake connections by evaluating the confidence level of edge presence. Additionally, GFAN provides a representation enhanced co-training module that highlights camouflaged fraudulent representations by training the small sphere and large margin support vector data description. Experimental results show that GFAN outperforms other competitive graph-based fraud detectors on public datasets. The GFAN code is available at: https://github.com/scu-kdde/OAM-GFAN-2023. Bingzhe Zhang, Xinye Wang, Zhenyang Yu, Yuanhao Zhang, Chengxin He, Song Deng, Zhaohang Luo, Lei Duan |
ICDM | 5 |
| 2023 | Memory-Enhanced Transformer for Representation Learning on Temporal Heterogeneous GraphsabstractAbstract Temporal heterogeneous graphs can model lots of complex systems in the real world, such as social networks and e-commerce applications, which are naturally time-varying and heterogeneous. As most existing graph representation learning methods cannot efficiently handle both of these characteristics, we propose a Transformer-like representation learning model, named THAN, to learn low-dimensional node embeddings preserving the topological structure features, heterogeneous semantics, and dynamic patterns of temporal heterogeneous graphs, simultaneously. Specifically, THAN first samples heterogeneous neighbors with temporal constraints and projects node features into the same vector space, then encodes time information and aggregates the neighborhood influence in different weights via type-aware self-attention. To capture long-term dependencies and evolutionary patterns, we design an optional memory module for storing and evolving dynamic node representations. Experiments on three real-world datasets demonstrate that THAN outperforms the state-of-the-arts in terms of effectiveness with respect to the temporal link prediction task. Longhai Li, Lei Duan, Junchen Wang, Chengxin He, Guicai Xie, Song Deng, Zhaohang Luo |
Data Sci. Eng. | 4 |
| 2022 | Efficient Gene Community Search to Discover Similar Aspects for Similarity ExplanationabstractGene similar aspects provide reliable explanation in understanding the biological roles and gene functions. As the volume of biomedical data expands, most of the current methods for similar explanation among genes are no longer applicable. Limited by information sources and search effciency, these methods cannot be flexible and effcient for the similarity analysis. We hereby propose a flexible method VENUS to analyze gene similar aspect among multiple genes on heterogeneous information networks, which constructed from public biomedicine databases and literature. VENUS infers the semantic and structural similarity of the query genes by gene community search. In this way, VENUS narrows the search space when searching information network within an acceptable time cost. Besides, VENUS is not limited by inherent domain knowledge and is adaptive to large-scale networks. Through experiments on multiple different public data sources, it demonstrates that VENUS is effective and effcient. Lei Duan, Chengxin He, Yuening Qu, Yidan Zhang 0001 |
BIBM | 3 |
| 2022 | MOVE: Integrating Multi-source Information for Predicting DTI via Cross-view Contrastive LearningabstractDrug-target interaction (DTI) prediction serves as the foundation of new drug findings and drug repositioning. For drugs/targets, the sequence data contains the biological structural information, while the heterogeneous network contains the biochemical functional information. These two types of information describe different aspects of drugs and targets. Due to the complexity of DTI machinery, it is necessary to learn the representation from multiple perspectives. We hereby try to design a way to leverage information from multi-source data to the maximum extent and find a strategy to fuse them. To address the above challenges, we propose a model, named MOVE (short for integrating multi-source information for predicting DTI via cross-view contrastive 1earning), for learning comprehensive representations of each drug and target from multi-source data. MOVE extracts information from the sequence view and the network view, then utilizes a fusion module with auxiliary contrastive learning to facilitate the fusion of representations. Experimental results on the benchmark dataset demonstrate that MOVE is effective in DTI prediction. Yuening Qu, Chengxin He, Jin Yin, Lei Duan |
BIBM | 2 |
| 2022 | An explainable framework for drug repositioning from disease information networkabstractExploring efficient and high-accuracy computational drug repositioning methods has become a popular and attractive topic in drug development. This technology can systematically identify potential drug-disease interactions, which could greatly alleviate the pressures from the high cost and long period taken by traditional drug research and discovery. However, plenty of current computational drug repositioning approaches lack interpretability in predicting drug-disease associations, which will not be friendly to their subsequent in-depth research. To this end, we hereby propose a novel computational framework, called EDEN, for exploring explainable drug repositioning from the disease information network (DIN). EDEN is a graph neural network framework that learns the local semantics and global structure of the DIN, and models the drug-disease associations into the DIN by maximizing the mutual information of both and an end-to-end manner. In this way, the learned biomedical entity and link embeddings are enabled to retain the ability to drug repositioning with the semantical structure of external knowledge, thereby making interpretation possible. Meanwhile, we also propose a matching score based on the final embeddings to generate the predictive drug repositioning explanation. Empirical results on the real-world dataset show that EDEN outperforms other state-of-the-art baselines on most of the metrics. Further studies reveal the effectiveness of the explainability of our approach. Chengxin He, Lei Duan, Huiru Zheng, Linlin Song, Menglin Huang |
Neurocomputing | 1 |
| 2022 | Efficient mining of concept-hierarchy aware distinguishing sequential patterns
Chengxin He, Lei Duan, Guozhu Dong, Jyrki Nummenmaa, Tingting Wang 0009, Tinghai Pang |
Knowl. Based Syst. | 1 |
| 2022 | Mining Similar Aspects for Gene Similarity Explanation Based on Gene Information NetworkabstractAnalysis of gene similarity not only can provide information on the understanding of the biological roles and functions of a gene, but may also reveal the relationships among various genes. In this paper, we introduce a novel idea of mining similar aspects from a gene information network, i.e., for a given gene pair, we want to know in which aspects (meta paths) they are most similar from the perspective of the gene information network. We defined a similarity metric based on the set of meta paths connecting the query genes in the gene information network and used the rank of similarity of a gene pair in a meta path set to measure the similarity significance in that aspect. A minimal set of gene meta paths where the query gene pair ranks the highest is a similar aspect, and the similar aspect of a query gene pair is far from trivial. We proposed a novel method, SCENARIO, to investigate minimal similar aspects. Our empirical study on the gene information network, constructed from six public gene-related databases, verified that our proposed method is effective, efficient, and useful. Yidan Zhang 0001, Lei Duan, Huiru Zheng, Jesse Li-Ling, Ruiqi Qin 0001, Chengxin He, Tingting Wang 0009 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2020 | DRAMA: Discovering Disease-related circRNA-miRNA-mRNA Axes from Disease-RNA Information NetworkabstractNon-coding RNAs are gaining prominence in biology and medicine, as they play major roles in cellular homeostasis and disease. A large number of computational methods have been recently developed for the prediction of the relationship between ncRNAs and diseases, which can alleviate the time-consuming and labor-intensive exploration among biological experiments. However, such methods have mainly focused on the association between the disease and certain types of ncRNAs such as miRNA or circRNA, thereby ignoring the impact of the interactions among ncRNAs on the diseases. We hereby propose a novel approach called DRAMA for discovering disease-related circRNA-miRNA-mRNA axes from the disease-RNA information network we constructed. Our method, using graph convolutional network, learns the characteristic representation of each biological entity by propagating and aggregating local neighbor information based on the global structure of the network. And then we design a favorable measurement to infer disease-related circRNA-miRNA-mRNA axes based on the learned embeddings. To evaluate the effectiveness of DRAMA, we conduct experiments on real-world datasets. Further analysis reveals that DRAMA outperforms other state-of-the-art baselines on most of the metrics. Chengxin He, Lei Duan, Huiru Zheng, Jesse Li-Ling, Longhai Li |
BIBM | 1 |
| 2020 | TSGYE: Two-Stage Grape Yield Estimation
Geng Deng, Tianyu Geng, Chengxin He, Xinao Wang, Bangjun He, Lei Duan |
ICONIP (4) | 3 |
| 2019 | ATOM: Construction of Anti-tumor Biomaterial Knowledge Graph by Biomedicine LiteratureabstractWith the rapid development of anti-tumor biomaterials, biomedicine literature with respect to anti-tumor biomaterials has been leveraged for tumor treatment as it provides abundant and useful information. A large number of biomedicine literature contains unstructured data, making it difficult for researchers to obtain desired messages from it. Knowledge Graphs (KGs) provides structured relationships among entities and can be served as a solution. However, no existing tool can be found in constructing an anti-tumor biomaterial knowledge graph from biomedicine literature. To fill this gap, a novel approach, ATOM, was proposed to construct an anti-tumor biomaterial knowledge graph from biomedicine literature through a series of process including the recognition of anti-tumor entities, the simplification of sentences, the extraction of triples, and the predicate mapping. Experiments demonstrated that ATOM is able to effectively express the extracted anti-tumor entities and their relationships. Tingting Wang 0009, Lei Duan, Chengxin He, Geng Deng, Ruiqi Qin 0001, Yidan Zhang 0001 |
BIBM | 3 |
| 2019 | SCENARIO: Discovery of Similar Aspects for Gene Similarity Explanation from Gene Information NetworkabstractGene similarity analysis not only provides information on understanding the biological roles and functions of a gene, but also reveals the relationships among different genes. In this paper, we identify the novel idea of mining similar aspects from gene information network, i.e., given a pair of genes, we want to know, in which aspects (meta paths) the two genes are mostly similar from the perspective of gene information network? We define a similarity metric based on the set of meta paths connecting the query genes in the gene information network, and use the rank of the similarity of a gene pair in a meta path set to measure the similarity significance in the aspect. A minimal set of meta paths where the query gene pair is ranked the best is a similar aspect. Computing the similar aspects of a query gene pair is far from trivial. In this paper, we propose a novel heuristic based-mining method, SCENARIO, to investigate minimal similar aspects. Our empirical study on the gene information network, constructed from seven public gene-related databases, verified that our proposed method is effective, efficient, and useful. Yidan Zhang 0001, Lei Duan, Huiru Zheng, Jesse Li-Ling, Ruiqi Qin 0001, Chengxin He |
BIBM | 7 |