VLDB 2026 Research / reviewers in the wild / expert
Chengkun Wu
dblp:06/7238
· DBLP profile ↗
42ranked-venue papers
3as first author
27since 2021 · last 2026
0000-0002-9688-5311ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 23 · 2 first-author · 18 since 2021Systems, architecture and hardware · 7 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ZeroForge: Zero-Shot 6D Pose Estimation via Generative 3D Reconstruction and Vision-Language Models
Binglin Wang, Boping Ran, Chengkun Wu |
ICIC (8) | 4 |
| 2026 | A Gated Semantic-Graph Network for Accurate Drug-Drug Interaction Prediction
Yilou Zhang, Haoyi Wang, Chengkun Wu |
ICIC (8) | 3 |
| 2025 | DTRAG: Triple-Fusion Retrieval-Augmented Generation Leveraging a Knowledge Graph for Answering Diabetes Questions
Wengen LI, Jianquan Wen, Chengkun Wu |
IEEE Big Data | 8 |
| 2025 | Scaling Deep Learning Molecular Dynamics to 500M Atoms on 4096-Node ARMv8 ClustersabstractMolecular dynamics (MD) simulations are essential tools for investigating large-scale molecular systems, yet achieving high performance and scalability on CPU-based architectures remains challenging. In this study, we present a highly optimized framework based on DeepMD-kit for conducting 500 millionatom MD simulations on an ARMv8 SVE high-performance computing (HPC) system. Key optimizations include leveraging OpenMP for multi-threaded acceleration of DeepMD-kit and utilizing the ARMv8 SVE instruction set to optimize doubleprecision matrix multiplication in PyTorch. These enhancements enable single ARMv8 SVE 64-core processors to achieve 1.3x the training performance of NVIDIA V100 GPU, and two ARMv8 SVE 64-core processors to achieve 1.05x the inference performance of NVIDIA V100 GPU. Leveraging this optimized framework, we achieve large-scale MD simulations across 4,096 computing nodes. Qi Du, Feng Wang 0050, Chengkun Wu, Han Wang 0006, Yongpeng Liu, Zhaoyin Zhou, Kenli Li 0001 |
CLUSTER | 3 |
| 2025 | Long-Tailed Recognition via Multi-sampling Classifier Fusion with Frozen CLIP
Yilou Zhang, Chengkun Wu |
ICIC (12) | 4 |
| 2025 | Pushing the boundaries of few-shot learning for low-data drug discovery with a Bayesian meta-learning hypernetwork frameworkabstractHunting for candidate compounds with favorable pharmacological, toxicological, and pharmacokinetic properties in drug discovery is essentially a low-data problem, as data acquisition is both challenging and costly. This inherent data limitation clashes with the requirements of many powerful deep learning models, which typically require large datasets. Here, we present Meta-Mol, a novel few-shot learning framework based on Bayesian Model-Agnostic Meta-Learning. Meta-Mol introduces a novel atom-bond graph isomorphism encoder that captures molecular structure information at the atomic and bond levels. This representation is further enhanced by a Bayesian meta-learning strategy, allowing for task-specific parameter adaptation and reducing overfitting risks. Additionally, a hypernetwork is employed to dynamically adjust weight updates across tasks, facilitating more complex posterior estimation. Our results demonstrate that Meta-Mol significantly outperforms existing models on several benchmarks, providing a robust solution to address data scarcity in drug discovery. Jiacai Yi, Dejun Jiang 0002, Chengkun Wu, Xiao-Chen Zhang, Weixing He, Dong-Sheng Cao 0001 |
Briefings Bioinform. | 3 |
| 2025 | Parallelization Strategies for DeepMD-Kit Using OpenMP: Enhancing Efficiency in Machine Learning-Based Molecular SimulationsabstractDeepMD-kit enables deep learning-based molecular dynamics (MD) simulations that require efficient parallelization to leverage modern HPC architectures. In this work, we optimize DeepMD-kit using advanced OpenMP strategies to improve scalability and computational efficiency on an ARMv8 processor-based server. Our optimizations include data parallelism for neural network inference, force calculation acceleration, NUMAaware memory management, and synchronization reductions, leading to up to 4.1× speedup and 82% higher memory band-width efficiency compared to the baseline implementation. Strong scaling analysis demonstrates superlinear speedup at mid-range core counts, with improved workload balancing and vectorized computations. However, challenges remain at ultra-large scales due to increasing synchronization overhead. Qi Du, Feng Wang 0050, Chengkun Wu |
IEEE Trans. Computers | 3 |
| 2024 | Exploring Natural Language Processing Model Acceleration in Molecular Dynamics Simulation Using High-Performance Computing and Machine LearningabstractIn molecular dynamics simulations, the integral methods used to calculate molecular trajectories, such as the Verlet integral method, involve significant computational costs. On the premise of ensuring the accuracy of simulation, how to effectively reduce the amount of computation is always a challenging problem. This paper applies Natural Language Processing models to simulate molecular motion trajectories in molecular dynamics, integrating the MPI programming model and machine learning on a high-performance computing platform to achieve MPI parallelization during model training and inference. The results indicate that different types of neural network architectures have a significant impact on inference performance and accuracy. When using the deep spatio-temporal networks model, the mean absolute error is 0.0032. At an atomic scale of 3M, the parallel efficiency with 1024 computational nodes exceeds 90%, demonstrating excellent parallel performance. Qi Du, Feng Wang 0050, Heng Wan, Chengkun Wu |
BIBM | 6 |
| 2024 | ChemMORT: an automatic ADMET optimization platform using deep learning and multi-objective particle swarm optimizationabstractDrug discovery and development constitute a laborious and costly undertaking. The success of a drug hinges not only good efficacy but also acceptable absorption, distribution, metabolism, elimination, and toxicity (ADMET) properties. Overall, up to 50% of drug development failures have been contributed from undesirable ADMET profiles. As a multiple parameter objective, the optimization of the ADMET properties is extremely challenging owing to the vast chemical space and limited human expert knowledge. In this study, a freely available platform called Chemical Molecular Optimization, Representation and Translation (ChemMORT) is developed for the optimization of multiple ADMET endpoints without the loss of potency (https://cadd.nscc-tj.cn/deploy/chemmort/). ChemMORT contains three modules: Simplified Molecular Input Line Entry System (SMILES) Encoder, Descriptor Decoder and Molecular Optimizer. The SMILES Encoder can generate the molecular representation with a 512-dimensional vector, and the Descriptor Decoder is able to translate the above representation to the corresponding molecular structure with high accuracy. Based on reversible molecular representation and particle swarm optimization strategy, the Molecular Optimizer can be used to effectively optimize undesirable ADMET properties without the loss of bioactivity, which essentially accomplishes the design of inverse QSAR. The constrained multi-objective optimization of the poly (ADP-ribose) polymerase-1 inhibitor is provided as the case to explore the utility of ChemMORT. Jiacai Yi, Wen-Tao Zhao, Zhi-Jiang Yang, Xiao-Chen Zhang, Chengkun Wu, Aiping Lu, Dong-Sheng Cao 0001 |
Briefings Bioinform. | 6 |
| 2024 | HFN: Heterogeneous feature network for multivariate time series anomaly detection
Chengkun Wu, Canqun Yang, Qiucheng Miao, Xiandong Ma |
Inf. Sci. | 2 |
| 2023 | DrugProtKGE: Weakly Supervised Knowledge Graph Embedding for Highly-Effective Drug-Protein Interaction RepresentationabstractWith the exponential growth of biomedical knowledge in unstructured text repositories such as PubMed, it is imminent to establish a knowledge graph-style, efficient searchable and targeted database that can support the need of information retrieval from researchers and clinicians. To mine knowledge from graph databases, most previous methods view a triple in a graph (see Fig. 1) as the basic processing unit and embed the triplet element (i.e. drugs/chemicals, proteins/genes and their interaction) as separated embedding matrices, which cannot capture the semantic correlation among triple elements. To remedy the loss of semantic correlation caused by disjoint embeddings, we propose a novel approach to learn triple embeddings by combining entities and interactions into a unified representation. Furthermore, traditional methods usually learn triple embeddings from scratch, which cannot take advantage of the rich domain knowledge embedded in pre-trained models, and is also another significant reason for the fact that they cannot distinguish the differences implied by the same entity in the multi-interaction triples. In this paper, we propose a novel fine-tuning based approach to learn better triple embeddings by creating weakly supervised signals from pre-trained knowledge graph embeddings. The method automatically samples triples from knowledge graphs and estimates their pairwise similarity from pre-trained embedding models. The triples are then fed pairwise into a Siamese-like neural architecture, where the triple representation is fine-tuned in the manner bootstrapped by triple similarity scores. Finally, we demonstrate that triple embeddings learned with our method can be readily applied to several downstream applications (e.g. triple classification and triple clustering). We evaluated the proposed method on two open-source drug-protein knowledge graphs constructed from PubMed abstracts, as provided by BioCreative. Our method achieves consistent improvement in both triple classification and triple clustering tasks when compared to other state-of-the-art triple embedding methods, with an average 35% improvement of F1 score for the multi-interaction triples. Siqi Wang 0001, Xi Yang 0020, Xinyuan Qiu, Chengkun Wu, Yingbo Cui 0001, Canqun Yang |
BIBM | 5 |
| 2022 | Deep Anomaly Discovery from Unlabeled Videos via Normality Advantage and Self-Paced RefinementabstractWhile classic video anomaly detection (VAD) requires labeled normal videos for training, emerging unsupervised VAD (UVAD) aims to discover anomalies directly from fully unlabeled videos. However, existing UVAD methods still rely on shallow models to perform detection or initialization, and they are evidently inferior to classic VAD methods. This paper proposes a full deep neural network (DNN) based solution that can realize highly effective UVAD. First, we, for the first time, point out that deep reconstruction can be surprisingly effective for UVAD, which inspires us to unveil a property named “normality advantage”, i.e., normal events will enjoy lower reconstruction loss when DNN learns to reconstruct unlabeled videos. With this property, we propose Localization based Reconstruction (LBR) as a strong UVAD baseline and a solid foundation of our solution. Second, we propose a novel self-paced refinement (SPR) scheme, which is synthesized into LBR to conduct UVAD. Unlike ordinary self-paced learning that injects more samples in an easy-to-hard manner, the proposed SPR scheme gradually drops samples so that suspicious anomalies can be removed from the learning process. In this way, SPR consolidates normality advantage and enables better UVAD in a more proactive way. Finally, we further design a variant solution that explicitly takes the motion cues into account. The solution evidently enhances the UVAD performance, and it sometimes even surpasses the best classic VAD methods. Experiments show that our solution not only significantly outperforms existing UVAD methods by a wide margin (5% to 9% AUROC), but also enables UVAD to catch up with the mainstream performance of classic VAD. Siqi Wang 0001, Zhiping Cai, Xinwang Liu 0002, Chuanfu Xu, Chengkun Wu |
CVPR | 6 |
| 2022 | Stgat-Mad : Spatial-Temporal Graph Attention Network For Multivariate Time Series Anomaly DetectionabstractAnomaly detection in multivariate time series data is challenging due to complex temporal and feature correlations. This paper proposes a novel unsupervised multi-scale stacked spatial-temporal graph attention network for multivariate time series anomaly detection (STGAT-MAD). The core of our framework is to coherently capture the feature and temporal correlations among multivariate time-series data by stackable STGAT networks. Meanwhile, a multi-scale input network is exploited to capture the temporal correlations in different time-scales. Besides, a new dataset derived from a real-world wind farm is built and released for multivariate time series anomaly detection. Experiments on the proprietary dataset and three public datasets show that our method significantly outperforms existing baseline approaches, and provides interpretability for anomaly location. Siqi Wang 0001, Xiandong Ma, Chengkun Wu, Canqun Yang, Detian Zeng, Shi-Lin Wang |
ICASSP | 4 |
| 2022 | An Unsupervised Short- and Long-Term Mask Representation for Multivariate Time Series Anomaly Detection
Qiucheng Miao, Chuanfu Xu, Chengkun Wu |
ICONIP (6) | 5 |
| 2022 | Effective Video Abnormal Event Detection by Learning A Consistency-Aware High-Level Feature ExtractorabstractWith pure normal training videos, video abnormal event detection (VAD) aims to build a normality model, and then detect abnormal events that deviate from this model. Despite of some progress, existing VAD methods typically train the normality model by a low-level learning objective (e.g. pixel-wise reconstruction/prediction), which often overlooks the high-level semantics in videos. To better exploit high-level semantics for VAD, we propose a novel paradigm that performs VAD by learning a Consistency-Aware high-level Feature Extractor (CAFE). Specifically, with a pre-trained deep neural network (DNN) as teacher network, we first feed raw video events into the teacher network and extract the outputs of multiple hidden layers as their high-level features, which contain rich high-level semantics. Guided by high-level features extracted from normal training videos, we train a student network to be the high-level feature extractor of normal events, so as to explicitly consider high-level semantics in training. For inference, a video event can be viewed as normal if the student extractor produces similar high-level features to the teacher network. Second, based on the fact that consecutive video frames usually enjoy minor differences, we propose a consistency-aware scheme that requires high-level features extracted from neighboring frames to be consistent. Our consistency-aware scheme not only encourages the student extractor to ignore low-level differences and capture more high-level semantics, but also enables better anomaly scoring. Last, we also design a generic framework that can bridge high-level and low-level learning in VAD to further ameliorate VAD performance. By flexibly embedding one or more low-level learning objectives into CAFE, the framework makes it possible to combine the strengths of both high-level and low-level learning. The proposed method attains state-of-the-art results on commonly-used benchmark datasets. Siqi Wang 0001, Zhiping Cai, Xinwang Liu 0002, Chengkun Wu |
ACM Multimedia | 5 |
| 2022 | BioNet: a large-scale and heterogeneous biological network model for interaction prediction with graph convolutionabstractMOTIVATION: Understanding chemical-gene interactions (CGIs) is crucial for screening drugs. Wet experiments are usually costly and laborious, which limits relevant studies to a small scale. On the contrary, computational studies enable efficient in-silico exploration. For the CGI prediction problem, a common method is to perform systematic analyses on a heterogeneous network involving various biomedical entities. Recently, graph neural networks become popular in the field of relation prediction. However, the inherent heterogeneous complexity of biological interaction networks and the massive amount of data pose enormous challenges. This paper aims to develop a data-driven model that is capable of learning latent information from the interaction network and making correct predictions. RESULTS: We developed BioNet, a deep biological networkmodel with a graph encoder-decoder architecture. The graph encoder utilizes graph convolution to learn latent information embedded in complex interactions among chemicals, genes, diseases and biological pathways. The learning process is featured by two consecutive steps. Then, embedded information learnt by the encoder is then employed to make multi-type interaction predictions between chemicals and genes with a tensor decomposition decoder based on the RESCAL algorithm. BioNet includes 79 325 entities as nodes, and 34 005 501 relations as edges. To train such a massive deep graph model, BioNet introduces a parallel training algorithm utilizing multiple Graphics Processing Unit (GPUs). The evaluation experiments indicated that BioNet exhibits outstanding prediction performance with a best area under Receiver Operating Characteristic (ROC) curve of 0.952, which significantly surpasses state-of-theart methods. For further validation, top predicted CGIs of cancer and COVID-19 by BioNet were verified by external curated data and published literature. Xi Yang 0020, Jing-Lun Ma, Kai Lu 0001, Dong-Sheng Cao 0001, Chengkun Wu |
Briefings Bioinform. | 7 |
| 2022 | ABC-Net: a divide-and-conquer based deep learning architecture for SMILES recognition from molecular imagesabstractStructural information for chemical compounds is often described by pictorial images in most scientific documents, which cannot be easily understood and manipulated by computers. This dilemma makes optical chemical structure recognition (OCSR) an essential tool for automatically mining knowledge from an enormous amount of literature. However, existing OCSR methods fall far short of our expectations for realistic requirements due to their poor recovery accuracy. In this paper, we developed a deep neural network model named ABC-Net (Atom and Bond Center Network) to predict graph structures directly. Based on the divide-and-conquer principle, we propose to model an atom or a bond as a single point in the center. In this way, we can leverage a fully convolutional neural network (CNN) to generate a series of heat-maps to identify these points and predict relevant properties, such as atom types, atom charges, bond types and other properties. Thus, the molecular structure can be recovered by assembling the detected atoms and bonds. Our approach integrates all the detection and property prediction tasks into a single fully CNN, which is scalable and capable of processing molecular images quite efficiently. Experimental results demonstrate that our method could achieve a significant improvement in recognition performance compared with publicly available tools. The proposed method could be considered as a promising solution to OCSR problems and a starting point for the acquisition of molecular information in the literature. Xiao-Chen Zhang, Jiacai Yi, Chengkun Wu, Tingjun Hou, Dong-Sheng Cao 0001 |
Briefings Bioinform. | 4 |
| 2022 | DeepKG: an end-to-end deep learning-based workflow for biomedical knowledge graph extraction, optimization and applicationsabstractSUMMARY: DeepKG is an end-to-end deep learning-based workflow that helps researchers automatically mine valuable knowledge in biomedical literature. Users can utilize it to establish customized knowledge graphs in specified domains, thus facilitating in-depth understanding on disease mechanisms and applications on drug repurposing and clinical research. To improve the performance of DeepKG, a cascaded hybrid information extraction framework is developed for training model of 3-tuple extraction, and a novel AutoML-based knowledge representation algorithm (AutoTransX) is proposed for knowledge representation and inference. The system has been deployed in dozens of hospitals and extensive experiments strongly evidence the effectiveness. In the context of 144 900 COVID-19 scholarly full-text literature, DeepKG generates a high-quality knowledge graph with 7980 entities and 43 760 3-tuples, a candidate drug list, and relevant animal experimental studies are being carried out. To accelerate more studies, we make DeepKG publicly available and provide an online tool including the data of 3-tuples, potential drug list, question answering system, visualization platform. AVAILABILITY AND IMPLEMENTATION: All the results are publicly available at the website (http://covidkg.ai/). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zongren Li, Qin Zhong, Yongjie Duan, Chengkun Wu, Kunlun He |
Bioinform. | 6 |
| 2022 | MICER: a pre-trained encoder-decoder architecture for molecular image captioningabstractMOTIVATION: Automatic recognition of chemical structures from molecular images provides an important avenue for the rediscovery of chemicals. Traditional rule-based approaches that rely on expert knowledge and fail to consider all the stylistic variations of molecular images usually suffer from cumbersome recognition processes and low generalization ability. Deep learning-based methods that integrate different image styles and automatically learn valuable features are flexible, but currently under-researched and have limitations, and are therefore not fully exploited. RESULTS: MICER, an encoder-decoder-based, reconstructed architecture for molecular image captioning, combines transfer learning, attention mechanisms and several strategies to strengthen effectiveness and plasticity in different datasets. The effects of stereochemical information, molecular complexity, data volume and pre-trained encoders on MICER performance were evaluated. Experimental results show that the intrinsic features of the molecular images and the sub-model match have a significant impact on the performance of this task. These findings inspire us to design the training dataset and the encoder for the final validation model, and the experimental results suggest that the MICER model consistently outperforms the state-of-the-art methods on four datasets. MICER was more reliable and scalable due to its interpretability and transfer capacity and provides a practical framework for developing comprehensive and accurate automated molecular structure identification tools to explore unknown chemical space. AVAILABILITY AND IMPLEMENTATION: https://github.com/Jiacai-Yi/MICER. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jiacai Yi, Chengkun Wu, Xiao-Chen Zhang, Xinyi Xiao, Tingjun Hou, Dong-Sheng Cao 0001 |
Bioinform. | 2 |
| 2022 | A risk factor attention-based model for cardiovascular disease predictionabstractBACKGROUND: Cardiovascular disease (CVD) is a serious disease that endangers human health and is one of the main causes of death. Therefore, using the patient's electronic medical record (EMR) to predict CVD automatically has important application value in intelligent assisted diagnosis and treatment, and is a hot issue in intelligent medical research. However, existing methods based on natural language processing can only predict CVD according to the whole or part of the context information of EMR. RESULTS: Given the deficiencies of the existing research on CVD prediction based on EMRs, this paper proposes a risk factor attention-based model (RFAB) to predict CVD by utilizing CVD risk factors and general EMRs text, which adopts the attention mechanism of a deep neural network to fuse the character sequence and CVD risk factors contained in EMRs text. The experimental results show that the proposed method can significantly improve the prediction performance of CVD, and the F-score reaches 0.9586, which outperforms the existing related methods. CONCLUSIONS: RFAB focuses on the key information in EMR that leads to CVD, that is, 12 risk factors. In the stage of risk factor identification and extraction, risk factors are labeled with category information and time attribute information by BiLSTM-CRF model. In the stage of CVD prediction, the information contained in risk factors and their labels is fused with the information of character sequence in EMR to predict CVD. RFAB makes well use of the fine-grained information contained in EMR, and also provides a reliable idea for predicting CVD. Chengkun Wu, Zhichang Zhang |
BMC Bioinform. | 3 |
| 2021 | Joint motion context and clip augmentation for spatio-temporal action detectionabstractThis paper endeavors to leverage spatio-temporal visual cues to improve video-based action detection. As a result, a NOn-Local Action detector based on anchor-free called NOLA is proposed, which is built off a recent moving center detector (MOC) and further extends it by efficiently aggregating long-range spatio-temporal information. In detail, a significantly efficient spatio-temporal motion-aware non-local block is explored to provide global motion contexts for the entire predictive branches of MOC. This byproduct can make the large batch samples run on a resource limited device. Besides, a light-weighted data augmentation method termed clip augmentation designed for video-based tasks is proposed, which serves to improve the generalization ability of the detector with economical scale-and-addition operation. NOLA works with two above schemes in real-time as well. Experiments on two benchmark datasets show that NOLA significantly exceeds MOC. Compared to other existing methods,, NOLA reaches the state-of-the-art, in terms of video-level mean of average precision (video mAP). Xurui Ma, Xiang Zhang 0008, Chengkun Wu, Chuanfu Xu, Jie Liu 0002, Zhigang Luo |
ICMV | 3 |
| 2021 | AA-LSTM: An Adversarial Autoencoder Joint Model for Prediction of Equipment Remaining Useful Life
Chengkun Wu, Chuanfu Xu, Zhenghua Wang |
PAKDD (1) | 2 |
| 2021 | Learning to SMILES: BAN-based strategies to improve latent representation learning from moleculesabstractComputational methods have become indispensable tools to accelerate the drug discovery process and alleviate the excessive dependence on time-consuming and labor-intensive experiments. Traditional feature-engineering approaches heavily rely on expert knowledge to devise useful features, which could be costly and sometimes biased. The emerging deep learning (DL) methods deliver a data-driven method to automatically learn expressive representations from complex raw data. Inspired by this, researchers have attempted to apply various deep neural network models to simplified molecular input line entry specification (SMILES) strings, which contain all the composition and structure information of molecules. However, current models usually suffer from the scarcity of labeled data. This results in a low generalization ability of SMILES-based DL models, which prevents them from competing with the state-of-the-art computational methods. In this study, we utilized the BiLSTM (bidirectional long short term merory) attention network (BAN) in which we employed a novel multi-step attention mechanism to facilitate the extracting of key features from the SMILES strings. Meanwhile, SMILES enumeration was utilized as a data augmentation method in the training phase to substantially increase the number of labeled data and enlarge the probability of mining more patterns from complex SMILES. We again took advantage of SMILES enumeration in the prediction phase to rectify model prediction bias and provide a more accurate prediction. Combined with the BAN model, our strategies can greatly improve the performance of latent features learned from SMILES strings. In 11 canonical absorption, distribution, metabolism, excretion and toxicity-related tasks, our method outperformed the state-of-the-art approaches. Chengkun Wu, Xiao-Chen Zhang, Zhi-Jiang Yang, Aiping Lu, Tingjun Hou, Dong-Sheng Cao 0001 |
Briefings Bioinform. | 1 |
| 2021 | MG-BERT: leveraging unsupervised atomic representation learning for molecular property predictionabstractMOTIVATION: Accurate and efficient prediction of molecular properties is one of the fundamental issues in drug design and discovery pipelines. Traditional feature engineering-based approaches require extensive expertise in the feature design and selection process. With the development of artificial intelligence (AI) technologies, data-driven methods exhibit unparalleled advantages over the feature engineering-based methods in various domains. Nevertheless, when applied to molecular property prediction, AI models usually suffer from the scarcity of labeled data and show poor generalization ability. RESULTS: In this study, we proposed molecular graph BERT (MG-BERT), which integrates the local message passing mechanism of graph neural networks (GNNs) into the powerful BERT model to facilitate learning from molecular graphs. Furthermore, an effective self-supervised learning strategy named masked atoms prediction was proposed to pretrain the MG-BERT model on a large amount of unlabeled data to mine context information in molecules. We found the MG-BERT model can generate context-sensitive atomic representations after pretraining and transfer the learned knowledge to the prediction of a variety of molecular properties. The experimental results show that the pretrained MG-BERT model with a little extra fine-tuning can consistently outperform the state-of-the-art methods on all 11 ADMET datasets. Moreover, the MG-BERT model leverages attention mechanisms to focus on atomic features essential to the target property, providing excellent interpretability for the trained model. The MG-BERT model does not require any hand-crafted feature as input and is more reliable due to its excellent interpretability, providing a novel framework to develop state-of-the-art models for a wide range of drug discovery tasks. Xiao-Chen Zhang, Chengkun Wu, Zhi-Jiang Yang, Zhen-Xing Wu, Jiacai Yi, Chang-Yu Hsieh, Tingjun Hou, Dong-Sheng Cao 0001 |
Briefings Bioinform. | 2 |
| 2021 | Mining microbe-disease interactions from literature via a transfer learning modelabstractBACKGROUND: Interactions of microbes and diseases are of great importance for biomedical research. However, large-scale of microbe-disease interactions are hidden in the biomedical literature. The structured databases for microbe-disease interactions are in limited amounts. In this paper, we aim to construct a large-scale database for microbe-disease interactions automatically. We attained this goal via applying text mining methods based on a deep learning model with a moderate curation cost. We also built a user-friendly web interface that allows researchers to navigate and query required information. RESULTS: Firstly, we manually constructed a golden-standard corpus and a sliver-standard corpus (SSC) for microbe-disease interactions for curation. Moreover, we proposed a text mining framework for microbe-disease interaction extraction based on a pretrained model BERE. We applied named entity recognition tools to detect microbe and disease mentions from the free biomedical texts. After that, we fine-tuned the pretrained model BERE to recognize relations between targeted entities, which was originally built for drug-target interactions or drug-drug interactions. The introduction of SSC for model fine-tuning greatly improved detection performance for microbe-disease interactions, with an average reduction in error of approximately 10%. The MDIDB website offers data browsing, custom searching for specific diseases or microbes, and batch downloading. CONCLUSIONS: Evaluation results demonstrate that our method outperform the baseline model (rule-based PKDE4J) with an average [Formula: see text]-score of 73.81%. For further validation, we randomly sampled nearly 1000 predicted interactions by our model, and manually checked the correctness of each interaction, which gives a 73% accuracy. The MDIDB webiste is freely avaliable throuth http://dbmdi.com/index/. Chengkun Wu, Xinyi Xiao, Canqun Yang, Jiacai Yi |
BMC Bioinform. | 1 |
| 2021 | Mining a stroke knowledge graph from literatureabstractBACKGROUND: Stroke has an acute onset and a high mortality rate, making it one of the most fatal diseases worldwide. Its underlying biology and treatments have been widely studied both in the "Western" biomedicine and the Traditional Chinese Medicine (TCM). However, these two approaches are often studied and reported in insolation, both in the literature and associated databases. RESULTS: To aid research in finding effective prevention methods and treatments, we integrated knowledge from the literature and a number of databases (e.g. CID, TCMID, ETCM). We employed a suite of biomedical text mining (i.e. named-entity) approaches to identify mentions of genes, diseases, drugs, chemicals, symptoms, Chinese herbs and patent medicines, etc. in a large set of stroke papers from both biomedical and TCM domains. Then, using a combination of a rule-based approach with a pre-trained BioBERT model, we extracted and classified links and relationships among stroke-related entities as expressed in the literature. We construct StrokeKG, a knowledge graph includes almost 46 k nodes of nine types, and 157 k links of 30 types, connecting diseases, genes, symptoms, drugs, pathways, herbs, chemical, ingredients and patent medicine. CONCLUSIONS: Our Stroke-KG can provide practical and reliable stroke-related knowledge to help with stroke-related research like exploring new directions for stroke research and ideas for drug repurposing and discovery. We make StrokeKG freely available at http://114.115.208.144:7474/browser/ (Please click "Connect" directly) and the source structured data for stroke at https://github.com/yangxi1016/Stroke. Xi Yang 0020, Chengkun Wu, Goran Nenadic, Wei Wang 0169, Kai Lu 0001 |
BMC Bioinform. | 2 |
| 2021 | Correction to: Mining a stroke knowledge graph from literature
Xi Yang 0020, Chengkun Wu, Goran Nenadic, Wei Wang 0169, Kai Lu 0001 |
BMC Bioinform. | 2 |
| 2020 | CGINet: graph convolutional network-based model for identifying chemical-gene interaction in an integrated multi-relational graphabstractBACKGROUND: Elucidation of interactive relation between chemicals and genes is of key relevance not only for discovering new drug leads in drug development but also for repositioning existing drugs to novel therapeutic targets. Recently, biological network-based approaches have been proven to be effective in predicting chemical-gene interactions. RESULTS: We present CGINet, a graph convolutional network-based method for identifying chemical-gene interactions in an integrated multi-relational graph containing three types of nodes: chemicals, genes, and pathways. We investigate two different perspectives on learning node embeddings. One is to view the graph as a whole, and the other is to adopt a subgraph view that initial node embeddings are learned from the binary association subgraphs and then transferred to the multi-interaction subgraph for more focused learning of higher-level target node representations. Besides, we reconstruct the topological structures of target nodes with the latent links captured by the designed substructures. CGINet adopts an end-to-end way that the encoder and the decoder are trained jointly with known chemical-gene interactions. We aim to predict unknown but potential associations between chemicals and genes as well as their interaction types. CONCLUSIONS: We study three model implementations CGINet-1/2/3 with various components and compare them with baseline approaches. As the experimental results suggest, our models exhibit competitive performances on identifying chemical-gene interactions. Besides, the subgraph perspective and the latent link both play positive roles in learning much more informative node embeddings and can lead to improved prediction. Wei Wang 0130, Xi Yang 0020, Chengkun Wu, Canqun Yang |
BMC Bioinform. | 3 |
| 2018 | Collaborative Subspace Graph Hashing for Cross-modal RetrievalabstractCurrent hashing methods for cross-modal retrieval generally attempt to learn the separate modality-specific transformation matrices to embed multi-modality data into a latent common subspace, and usually ignore the fact that respecting the diversity of multi-modality features in the latent subspace could be beneficial for retrieval improvements. To this, we propose a collaborative subspace graph hashing method (CSGH) to perform a two-stage collaborative learning framework for cross-modal retrieval. Particularly, CSGH first embeds multi-modality data into separate latent subspaces through individual modality-specific transformation matrices, and then connects these latent subspaces to a common Hamming space through a shared transformation matrix. In this framework, CSGH considers the modality-specific neighborhood structure and the cross-modal correlation within multi-modality data through the Laplacian regularization and the graph based correlation constraint, respectively. To solve CSGH, we develop an alternative procedure to optimize it, and fortunately, each sub-problem of CSGH has the elegant analytical solution. Experiments of cross-modal retrieval on Wiki, NUS-WIDE, Flickr25K and Flickr1M datasets show the effectiveness of CSGH compared with the state-of-the-art cross-modal hashing methods. Xiang Zhang 0008, Guohua Dong, Yimo Du, Chengkun Wu, Zhigang Luo, Canqun Yang |
ICMR | 4 |
| 2018 | Constructing a database for the relations between CNV and human genetic diseases via systematic text miningabstractBACKGROUND: The detection and interpretation of CNVs are of clinical importance in genetic testing. Several databases and web services are already being used by clinical geneticists to interpret the medical relevance of identified CNVs in patients. However, geneticists or physicians would like to obtain the original literature context for more detailed information, especially for rare CNVs that were not included in databases. RESULTS: The resulting CNVdigest database includes 440,485 sentences for CNV-disease relationship. A total number of 1582 CNVs and 2425 diseases are involved. Sentences describing CNV-disease correlations are indexed in CNVdigest, with CNV mentions and disease mentions annotated. CONCLUSIONS: In this paper, we use a systematic text mining method to construct a database for the relationship between CNVs and diseases. Based on that, we also developed a concise front-end to facilitate the analysis of CNV/disease association, providing a user-friendly web interface for convenient queries. The resulting system is publically available at http://cnv.gtxlab.com /. Xi Yang 0020, Chengkun Wu, Wei Wang 0130, Gen Li 0007, Wei Zhang 0027, Lingqian Wu, Kai Lu 0001 |
BMC Bioinform. | 3 |
| 2018 | Distributed and asynchronous Stochastic Gradient Descent with variance reduction
Yuewei Ming, Chengkun Wu, Kuan Li, Jianping Yin |
Neurocomputing | 3 |
| 2018 | mSNP: A Massively Parallel Algorithm for Large-Scale SNP DetectionabstractSingle Nucleotide Polymorphism (SNP) detection is a fundamental procedure of whole genome analysis. SOAPsnp, a classic tool for detection, would take more than one week to analyze one typical human genome, which limits the efficiency of downstream analyses. In this paper, we present mSNP, an optimized version of SOAPsnp, which leverages Intel Xeon Phi coprocessors for large-scale SNP detection. Firstly, we redesigned the essential data structures of SOAPsnp, which significantly reduces memory footprint and improves computing efficiency. Then we developed a coordinated parallel framework for a higher hardware utilization of both CPU and Xeon Phi. Also, we tailored the data structures and operations to utilize the wide VPU of Xeon Phi to improve data throughput. Last but not the least, we proposed a read-based window division strategy to improve throughput and obtain better load balance. mSNP is the first SNP detection tool empowered by Xeon Phi. We achieved a 38x single thread speedup on CPU, without any loss in precision. Moreover, mSNP successfully scaled to 4,096 nodes on Tianhe-2. Our experiments demonstrate that mSNP is efficient and scalable for large-scale human genome SNP detection. Yingbo Cui 0001, Shaoliang Peng, Yutong Lu, Xiaoqian Zhu, Bingqiang Wang, Chengkun Wu, Xiangke Liao |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2017 | SVRG with adaptive epoch sizeabstractStochastic gradient descent (SGD) is a commonly used technique in large-scale machine learning tasks, but its convergence is slow due to the inherent variance. In recent years, a popular method, Stochastic Variance Reduced Gradient (SVRG), addresses this shortcoming via computing the full gradient of the entire dataset in each epoch. However, conventional SVRG and its variants usually need to identify a hyperparameter - the epoch size, which is essential to the convergence performance. Few previous studies discuss how to systematically find a suitable value for that hyper-parameter, which makes it hard to gain a good convergence performance in practical machine learning applications. In this paper, we propose a new stochastic gradient descent named AESVRG, which introduces variance reduction and computes the full gradient adaptively. Its enhanced implementation, AESVRG+, has a convergence performance that can outplay existing SVRG with fine-tuned epoch sizes. An extensive evaluation illustrates the significant performance improvement of our method. Erxue Min, Chengkun Wu, Kuan Li, Jianping Yin |
IJCNN | 4 |
| 2017 | Dependency-based long short term memory network for drug-drug interaction extractionabstractBACKGROUND: Drug-drug interaction extraction (DDI) needs assistance from automated methods to address the explosively increasing biomedical texts. In recent years, deep neural network based models have been developed to address such needs and they have made significant progress in relation identification. METHODS: We propose a dependency-based deep neural network model for DDI extraction. By introducing the dependency-based technique to a bi-directional long short term memory network (Bi-LSTM), we build three channels, namely, Linear channel, DFS channel and BFS channel. All of these channels are constructed with three network layers, including embedding layer, LSTM layer and max pooling layer from bottom up. In the embedding layer, we extract two types of features, one is distance-based feature and another is dependency-based feature. In the LSTM layer, a Bi-LSTM is instituted in each channel to better capture relation information. Then max pooling is used to get optimal features from the entire encoding sequential data. At last, we concatenate the outputs of all channels and then link it to the softmax layer for relation identification. RESULTS: To the best of our knowledge, our model achieves new state-of-the-art performance with the F-score of 72.0% on the DDIExtraction 2013 corpus. Moreover, our approach obtains much higher Recall value compared to the existing methods. CONCLUSIONS: The dependency-based Bi-LSTM model can learn effective relation information with less feature engineering in the task of DDI extraction. Besides, the experimental results show that our model excels at balancing the Precision and Recall values. Wei Wang 0130, Xi Yang 0020, Canqun Yang, Xiang Zhang 0008, Chengkun Wu |
BMC Bioinform. | 6 |
| 2017 | GTZ: a fast compression and cloud transmission tool optimized for FASTQ filesabstractBACKGROUND: The dramatic development of DNA sequencing technology is generating real big data, craving for more storage and bandwidth. To speed up data sharing and bring data to computing resource faster and cheaper, it is necessary to develop a compression tool than can support efficient compression and transmission of sequencing data onto the cloud storage. RESULTS: This paper presents GTZ, a compression and transmission tool, optimized for FASTQ files. As a reference-free lossless FASTQ compressor, GTZ treats different lines of FASTQ separately, utilizes adaptive context modelling to estimate their characteristic probabilities, and compresses data blocks with arithmetic coding. GTZ can also be used to compress multiple files or directories at once. Furthermore, as a tool to be used in the cloud computing era, it is capable of saving compressed data locally or transmitting data directly into cloud by choice. We evaluated the performance of GTZ on some diverse FASTQ benchmarks. Results show that in most cases, it outperforms many other tools in terms of the compression ratio, speed and stability. CONCLUSIONS: GTZ is a tool that enables efficient lossless FASTQ data compression and simultaneous data transmission onto to cloud. It emerges as a useful tool for NGS data storage and transmission in the cloud environment. GTZ is freely available online at: https://github.com/Genetalks/gtz . Yuting Xing, Gen Li 0007, Bolun Feng, Chengkun Wu |
BMC Bioinform. | 6 |
| 2016 | Predicting Unknown Interactions Between Known Drugs and Targets via Matrix Completion
Qing Liao 0001, Naiyang Guan, Chengkun Wu, Qian Zhang 0001 |
PAKDD (1) | 3 |
| 2015 | mAMBER: Accelerating Explicit Solvent Molecular Dynamic with Intel Xeon Phi Many-Integrated Core CoprocessorsabstractMolecular dynamics (MD) is a computer simulation of physical movements of atoms and molecules, which is a very important research technique for the study of biological and chemical systems at micro-scale. Assisted Model Building with Energy Refinement (AMBER) is one of the most commonly used software for MD. However, the microsecond MD simulation of large-scale atom system requires a lot of computation power. In this paper, we propose mAMBER: an Intel Xeon Phi Many-Integrated Core (MIC) Coprocessors accelerated implementation of explicit solvent all-atom classical molecular dynamics (MD) within the AMBER program package. We mAMBER also includes new parallel algorithm using CPUs and MIC coprocessors on Tianhe-2 supercomputer. With several optimizing techniques including CPU/MIC collaborated parallelization, factorization and asynchronous data transfer framework, we can accelerate the sander program of AMBER (version 12) in 'offload' mode, and achieves a 4.17-fold overall speedup compared with the CPU-only sander program. Shaoliang Peng, Canqun Yang, Chengkun Wu, Haiqiang Wang, Weiliang Zhu, Jinan Wang |
CCGRID | 4 |
| 2015 | The Challenge of Scaling Genome Big Data Analysis Software on TH-2 SupercomputerabstractWhole genome re-sequencing plays a crucial role in biomedical studies. The emergence of genomic big data calls for an enormous amount of computing power. However, current computational methods are inefficient in utilizing available computational resources. In this paper, we address this challenge by optimizing the utilization of the fastest supercomputer in the world - TH-2 supercomputer. TH-2 is featured by its neo-heterogeneous architecture, in which each compute node is equipped with 2 Intel Xeon CPUs and 3 Intel Xeon Phi coprocessors. The heterogeneity and the massive amount of data to be processed pose great challenges for the deployment of the genome analysis software pipeline on TH-2. Runtime profiling shows that SOAP3-dp and SOAPsnp are the most time-consuming components (up to 70% of total runtime) in a typical genome-analyzing pipeline. To optimize the whole pipeline, we first devise a number of parallel and optimization strategies for SOAP3-dp and SOAPsnp, respectively targeting each node to fully utilize all sorts of hardware resources provided both by CPU and MIC. We also employ a few scaling methods to reduce communication between different nodes. We then scaled up our method on TH-2. With 8192 nodes, the whole analyzing procedure took 8.37 hours to finish the analysis of a 300 TB dataset of whole genome sequences from 2,000 human beings, which can take as long as 8 months on a commodity server. The speedup is about 700x. Shaoliang Peng, Xiangke Liao, Canqun Yang, Yutong Lu, Jie Liu 0002, Yingbo Cui 0001, Chengkun Wu, Bingqiang Wang |
CCGRID | 8 |
| 2015 | A Method to Accelerate GROMACS in Offload Mode on Tianhe-2 SupercomputerabstractMolecular Dynamics(MD) is a computer simulation of physical movements of atoms and molecules in the context of N-body simulation, and is an important part of pharmaceutical industry. GROMACS, which is the most popular software for MD, could not perform satisfactorily with large-scale for the limit of computing resources. In this paper, we proposed a method to accelerate GROMACS with offload mode. In this mode, GROMACS could be arranged efficiently with CPU and the Intel® Xeon PhiTM Many Integrated Core (MIC) coprocessors at the same time, making the full use of Tianhe-2 supercomputer resources. To promote the efficiency of GROMACS, we proposed a series of methods, such as synchronization, data reassemble and array reuse. As we known, we are the first to accelerate GROMACS in offload mode on MIC. Haiqiang Wang, Shaoliang Peng, Xiaoqian Zhu, Chengkun Wu, Weiliang Zhu, Jinan Wang, Huaiyu Yang |
CCGRID | 4 |
| 2011 | A Parallel Fusion Method for Heterogeneous Multi-sensor Transportation Data
Yingjie Xia, Chengkun Wu, Qing-Jie Kong, Zhenyu Shan, Li Kuang |
MDAI | 2 |
| 2009 | A Hybrid Parallel Signature Matching Model for Network Security Applications Using SIMD GPU
Chengkun Wu, Jianping Yin, Zhiping Cai, En Zhu, Jieren Cheng |
APPT | 1 |
| 2009 | DDoS Attack Detection Method Based on Linear Prediction Model
Jieren Cheng, Jianping Yin, Chengkun Wu, Boyun Zhang |
ICIC (1) | 3 |