VLDB 2026 Research / reviewers in the wild / expert
Xiangyu Meng 0005
dblp:16/8083-5
· DBLP profile ↗
14ranked-venue papers
2as first author
14since 2021 · last 2025
0000-0001-5696-3090ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 9 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GASNet: Progressive Resolution-Aware Supervision and Gabor Guidance for Accurate Liver Vessel SegmentationabstractHigh-precision segmentation of liver vessels is crucial for surgical planning and clinical diagnosis. Yet, it remains challenging due to intricate vascular structures and the low contrast inherent in CT images. We propose GASNet, an innovative liver vessel segmentation network that integrates progressive resolution-aware supervision with a learnable Gabor multi-filter module. GASNet generates hierarchical segmentation masks and applies adaptive supervision at intermediate layers to enhance multi-scale feature learning. Built upon a ConvNeXt backbone, the network integrates a Multi-Scale Feature Refiner and Dilated Convolutional modules to improve structural modeling capacity. We evaluate GASNet on two public datasets (LiVS and MSD) as well as aselfconstructed dataset, LVTSD. GASNet consistently outperforms state-of-the-art methods and demonstrates superior capability in distinguishing hepatic and portal veins on the revised MSD dataset, achieving a Dice score of 0.836 for hepatic veins and 0.820 for portal veins. The implementation is available at https://github.com/HappyBot516/GASNet. Xiangyu Meng 0005, Huanhuan Dai, Xun Wang 0010 |
BIBM | 4 |
| 2025 | An interpretable DeePMD-kit performance model for emerging supercomputersabstractAbstract Deep potential (DP) scheme has increased the simulation temporal and spatial scales while maintaining the ab initio accuracy of the molecular dynamics. DeePMD-kit is an outstanding application that implements DP scheme efficiently. However, current performance model cannot accurately measure the resource utilization of DeePMD-kit operators and predict the execution time. We introduce DP-perf, an interpretable performance model for DeePMD-kit. DP-perf can accurately measure the resource utilization of the individual DeePMD-kit operators, communication pattern, and the overall application by exploiting physical system properties and machine configurations. It can be easily applied to mainstream supercomputers including Tianhe-3F, the new Sunway, Fugaku, and Summit. With DP-perf, users can select the optimal machine and decide the corresponding configuration for various purposes (e.g., lower cost, less time) without real runs. Evaluation of four top supercomputers shows that DP-perf can fit overall execution time with a low mean absolute percentage error of 5.7 %/8.1%/14.3%/13.1% on Tianhe-3F/new Sunway/Fugaku/Summit. On the prediction scenario, DP-perf can predict the total execution time with a mean absolute percentage error of less than 20%. Xiangyu Meng 0005, Xun Wang 0010, Mingzhen Li 0001, Guangming Tan, Weile Jia |
CCF Trans. High Perform. Comput. | 1 |
| 2025 | 29-Billion Atoms Molecular Dynamics Simulation With Ab Initio Accuracy on 35 Million Cores of New Sunway SupercomputerabstractPhysical phenomena such as bond breaking and phase transitions require molecular dynamics (MD) withab initioaccuracy, involving up to billions of atoms and over nanosecond timescales. Previous state-of-the-art work has demonstrated that neural network molecular dynamics (NNMD) like deep potential molecular dynamics (DeePMD), can successfully extend the temporal and spatial scales of MD withab initioaccuracy on both ARM and GPU platforms. However, the DeePMD-kit package is currently unable to fully exploit the computational potential of the new Sunway supercomputer due to its unique many-core architecture, memory hierarchy, and low precision capability. In this paper, we re-design the DeePMD-kit to harness the massive computing power of the new Sunway, enabling the MD with over ten billion atoms. We first design a large-scale parallelization scheme to exploit the massive parallelism of the new Sunway. Then we devise specialized optimizations for the time-consuming operators. Finally, we design a novel mixed precision method for DeePMD-kit customized operators to leverage the low precision computing power of the new Sunway. The optimized DeePMD-kit achieves 67.6 / 56.5$\boldsymbol{\times}$speedup for water / copper systems on the new Sunway. Meanwhile, it can perform 29 billion atoms simulation for the water system on 35 million cores (i.e., 90,000 computing nodes, around 84% of the whole supercomputer) with a peak performance of 57.1 PFLOPs, which is 7.9$\boldsymbol{\times}$bigger and 1.2$\boldsymbol{\times}$faster than state-of-the-art results. This paves the way for investigating more realistic scenarios, such as studying the mechanical properties of metals, semiconductor devices, batteries, and other materials and physical systems. Xun Wang 0010, Xiangyu Meng 0005, Zhuoqiang Guo, Mingzhen Li 0001, Mingfan Li, Ninghui Sun, Guangming Tan, Weile Jia |
IEEE Trans. Computers | 2 |
| 2025 | Gene-MOE: A Sparsely Gated Cancer Diagnosis and Prognosis Framework Exploiting Pan-Cancer Genomic InformationabstractImproved cancer genomic diagnosis and prognosis are vital to accurate medical therapy. Deep learning methods offered an end-to-end solution to enhance the precision of analysis. With the fast pace of pre-trained Transformer models, it remains uncertain whether some novel approaches such as the sparsely gated mixture of expert (MOE) and self-attention mechanisms can further improve the precision of cancer prognosis and classification. In this paper, we introduce a novel sparsely gated cancer diagnosis and prognosis framework called Gene-MOE exploiting the potential of the MOE layers and the proposed mixture of attention expert (MOAE) layers to enhance the analysis accuracy. Additionally, we address overfitting challenges by integrating pan-cancer information from 33 distinct cancer types through pre-training. For survival analysis, Gene-MOE achieves the best Concordance Index compared with state-of-the-art models on 12 of 14 cancer types. For cancer classification, the total accuracy of the classification model for 33 cancer classifications reached 95.8%, representing the best performance compared to state-of-the-art models. For cancer subtyping, Gene-MOE achieves the best result on at least one metric of the log10 P-values and the number of significant clinical on seven of nine cancers. These results indicate that Gene-MOE holds strong potential for these downstream tasks. Xiangyu Meng 0005, Xue Li 0019, Huanhuan Dai, Lian Qiao, Hongzhen Ding, Long Hao, Xun Wang 0010 |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2025 | scSwinTNet: A Cell Type Annotation Method for Large-Scale Single-Cell RNA-Seq Data Based on Shifted Window AttentionabstractThe annotation of cell types based on single-cell RNA sequencing (scRNA-seq) data is a critical downstream task in single-cell analysis, with significant implications for a deeper understanding of biological processes. Most analytical methods cluster cells by unsupervised clustering, which requires manual annotation for cell type determination. This procedure is time-overwhelming and non-repeatable. To accommodate the exponential growth of sequencing cells, reduce the impact of data bias, and integrate large-scale datasets for further improvement of type annotation accuracy, we proposed scSwinTNet. It is a pre-trained tool for annotating cell types in scRNA-seq data, which uses self-attention based on shifted windows and enables intelligent information extraction from gene data. We demonstrated the effectiveness and robustness of scSwinTNet by using 399 760 cells from human and mouse tissues. To the best of our knowledge, scSwinTNet is the first model to annotate cell types in scRNA-seq data using a pre-trained shifted window attention-based model. It does not require a priori knowledge and accurately annotates cell types without manual annotation. Huanhuan Dai, Xiangyu Meng 0005, Zhiyi Pan 0003, Haonan Song, Yuan Gao 0048, Xun Wang 0010 |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | Predicting Mutation-Disease Associations Through Protein Interactions Via Deep LearningabstractDisease is one of the primary factors affecting life activities, with complex etiologies often influenced by gene expression and mutation. Currently, wet lab experiments have analyzed the mechanisms of mutations, but these are usually limited by the costs of wet experiments and constraints in sample types and scales. Therefore, this paper constructs a real-world mutation-induced disease dataset and proposes Capsule and Graph topology networks with Multi-head attention (CGM) to predict the mutation-disease associations. CGM can accurately predict protein mutation-disease associations, and to further elucidate the pathogenicity of protein mutations, we also verified that protein mutations lead to protein structural alterations by the model, which suggests that mutation-induced conformational changes may be an important pathogenic factor. Limited by the size of the mutated protein dataset, we also performed experiments on benchmark and imbalanced datasets, where CGM mined 22 unknown protein interaction pairs from the benchmark dataset, better illustrating the potential of CGM in predicting mutation-disease associations. In summary, this paper curates a real dataset. It proposes that CGM predicts protein mutations and disease associations, providing a novel tool for further understanding of biomolecular pathways and disease mechanisms. Xue Li 0019, Ben Cao, Jianmin Wang 0016, Xiangyu Meng 0005, Yu Huang 0004, Enrico Petretto, Tao Song 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Inference and training acceleration of deep learning partial differential equation solver
Xun Wang 0010, Xianxi Zhu, Xiangyu Meng 0005, Zeyang Zhu, Tao Song 0001 |
J. Supercomput. | 3 |
| 2024 | AEG-PPIS: A Dual-Branch Protein-protein Interaction Site Predictor Based on Augmented Graph Attention Network and Equivariant Graph Neural NetworkabstractThe identification of protein-protein interaction sites (PPIS) plays a crucial role in understanding the mechanisms of biological processes. Traditional biological experimental methods for PPIS prediction are both expensive and time-consuming, developing computational methods can effectively reduce costs. However, existing approaches often focus on single-scale features and pay little attention to spatial neighborhood features, leading to unsatisfactory prediction performance. To address these challenges, we propose a dual-branch PPIS predictor (AEG-PPIS) based on augmented graph attention network (AGAT) and E(n) equivariant graph neural network (EGNN). AEG-PPIS extracts global features through EGNN, ensuring rotational and translational invariance of the protein graph. For local feature extraction, it employs an enhanced Augmented Graph Attention Network, which integrates initial node features, previous layer outputs, and features extracted by GraphSAGE using residual connections and identity mapping. Our model realizes the modeling of multi-scale information. Comparative experimental results show that the performance of AEG-PPIS is better than that of state-of-the-art methods. Ablation experiments and case studies demonstrate the effectiveness of AEG-PPIS in predicting PPIS. Huanhuan Dai, Haonan Song, Tongyu Han, Xiangyu Meng 0005, Xun Wang 0010 |
BIBM | 5 |
| 2024 | DCUI-MGraphDTA: Enabling Efficient Inference of a Drug-Target Binding Affinity Prediction Model on DCUsabstractEfficient and accurate identification of drug-target affinity (DTA) is crucial in the virtual screening stage of drug discovery and the reuse of existing drugs. Deep neural networks are becoming popular tools for providing fast and accurate binding affinity prediction. MGraphDTA is one of the best performing models for DTA prediction, utilizing a super-deep graph neural network to extract features, demonstrating good generalization and interpretation capabilities. However, the complex deep neural networks come with a higher computational cost of model inference. Meanwhile, due to the expansion of drug databases, rapid completion of large-scale virtual screening still presents challenges. In this work, we propose an efficient binding affinity prediction method called DCUI-MGraphDTA through a series of optimizations on MGraphDTA to achieve large-scale virtual screening for millions of molecules on the Deep Computing Unit (DCU) cluster. The optimized DCUI-MGraphDTA model achieves a speedup of approximately 5.83 times on a single DCU compared to the original MGraphDTA. As the number of DCUs expands from 1 to 32, the total time for the model to process 1.28 million data significantly decreases, achieving a speedup of approximately 11.9 times. Tao Song 0001, Xiangyu Meng 0005, Zeyang Zhu, Xianxi Zhu, Xun Wang 0010 |
BIBM | 3 |
| 2024 | Training one DeePMD Model in Minutes: a Step towards Online LearningabstractNeural Network Molecular Dynamics (NNMD) has become a major approach in material simulations, which can speedup the molecular dynamics (MD) simulation for thousands of times, while maintaining ab initio accuracy, thus has a potential to fundamentally change the paradigm of material simulations. However, there are two time-consuming bottlenecks of the NNMD developments. One is the data access of ab initio calculation results. The other, which is the focus of the current work, is reducing the training time of NNMD model. The training of NNMD model is different from most other neural network training because the atomic force (which is related to the gradient of the network) is an important physical property to be fit. Tests show the traditional stochastic gradient methods, like the Adam algorithms, cannot efficiently deploy the multisample minibatch algorithm. As a result, a typical training (taking the Deep Potential Molecular Dynamics (DeePMD) as an example) can take many hours. In this work, we designed a heuristic minibatch quasi-Newtonian optimizer based on Extended Kalman Filter method. An early reduction of gradient and error is adopted to reduce memory footprint and communication. The memory footprint, communication and settings of hyper-parameters of this new method are analyzed in detail. Computational innovations such as customized kernels of the symmetry-preserving descriptor are applied to exploit the computing power of the heterogeneous architecture. Experiments are performed on 8 different datasets representing different real case situations, and numerical results show that our new method has an average speedup of 32.2 compared to the Reorganized Layer-wised Extended Kalman Filter with 1 GPU, reducing the absolute training time of one DeePMD model from hours to several minutes, making it one step toward online training. Siyu Hu, Qiuchen Sha, Enji Li, Xiangyu Meng 0005, Lin-Wang Wang, Guangming Tan, Weile Jia |
PPoPP | 5 |
| 2024 | Exploring Efficient Partial Differential Equation Solution Using Speed Galerkin TransformerabstractFourier Neural Operator (FNO) has been proven to be a universal and effective deep learning framework capable of achieving remarkable accuracy on Partial Differential Equation (PDE) solution problem. However, certain key components of emerging FNO-based models cannot leverage hardware potential, which makes it difficult to apply in high resolution and high realtime demand scenario. This paper presents a high optimized model called Speed Galerkin Transformer, including multilevel parallel SliceK-SplitK-ReduceK strategy for batched skinny matrix multiplication, memory layout optimization for QKV matrices and positional encodings and multi-head layer normalization fusion, as well as batched transposition optimization with strided scattering and gathering in 2D FNO, and these strategies can achieve $10.29 \mathrm{x}, 4.41 \mathrm{x}$ and 2.38 x speedup respectively under specific configuration. When solving the Darcy Flow equation at 512x512 resolution, the Speed Galerkin Transformer model can achieve about 1.72 x speedup, and achieve more than $\mathbf{9 0 \%}$ parallel efficiency on 8 GPUs. Xun Wang 0010, Zeyang Zhu, Xiangyu Meng 0005, Tao Song 0001 |
SC | 3 |
| 2023 | MulAxialGO: Multi-Modal Feature-Enhanced Deep Learning Model for Protein Function PredictionabstractPredicting protein function from sequences through machine learning can improve the understanding of novel proteins and biological mechanisms. Existing methods mainly rely on one-dimensional convolution or natural language processing (NLP) techniques to extract features from sequences, but they suffer from limited predictive performance. To address this challenge, we propose MulAxialGO, a new method that leverages multi-modal feature fusion to improve prediction accuracy. MulAxialGO integrates the prior features of a large-scale pre-trained protein language model and the posterior features of dynamic embedding coding and sequence homology. In addition, MulAxialGO employs a comprehensive image feature encoder to extract features from sequences, providing a novel perspective for protein function prediction. MulAxialGO is tested on two benchmark datasets and achieves state-of-the-art results. On the 2016 dataset, MulAxialGO significantly outperforms DeepGOPlus, improving molecular function by 4.5 points, biological process by 2.4 points and cellular component by 1.6 points for the AUPR metric. Similarly, on the NetGO dataset, MulAxialGO outperforms the state-of-the-art NetGO2.0, improving Fmax by 1.1 points for biological process and 2.3 points for cellular component. Xun Wang 0010, Peng Qu 0002, Xiangyu Meng 0005, Lian Qiao, Chaogang Zhang, Xianjin Xie |
BIBM | 3 |
| 2023 | TransFusionNet: Semantic and Spatial Features Fusion Framework for Liver Tumor and Vessel Segmentation Under JetsonTX2abstractLiver cancer is one of the most common malignant diseases worldwide. Segmentation and reconstruction of liver tumors and vessels in CT images can provide convenience for physicians in preoperative planning and surgical intervention. In this paper, we introduced a TransFusionNet framework, which consists of a semantic feature extraction module, a local spatial feature extraction module, an edge feature extraction module, and a multi-scale feature fusion module to achieve fine-grained segmentation of liver tumors and vessels. In addition, we applied the transfer learning approach to pre-train using public datasets and then fine-tune the model to further improve the fitting effect. Furthermore, we proposed an intelligent quantization scheme to compress the model weights and achieved high performance inference on JetsonTX2. The TransFusionNet framework achieved mean IoU of 0.854 in vessel segmentation task, and achieved mean IoU of 0.927 in liver tumor segmentation task. When profiling the Computational Performance of the quantized inference, our quantized model achieved 4TFLOPs on Node with NVIDIA RTX3090 and 132GFLOPs on JetsonTX2. This unprecedented segmentation effect solves the accuracy and performance bottleneck of automated segmentation to a certain extent. Xun Wang 0010, Gan Wang, Huanhuan Dai, Zixuan Wang 0012, Xiangyu Meng 0005 |
IEEE J. Biomed. Health Informatics | 9 |
| 2022 | Molormer: a lightweight self-attention-based method focused on spatial structure of molecular graph for drug-drug interactions predictionabstractMulti-drug combinations for the treatment of complex diseases are gradually becoming an important treatment, and this type of treatment can take advantage of the synergistic effects among drugs. However, drug-drug interactions (DDIs) are not just all beneficial. Accurate and rapid identifications of the DDIs are essential to enhance the effectiveness of combination therapy and avoid unintended side effects. Traditional DDIs prediction methods use only drug sequence information or drug graph information, which ignores information about the position of atoms and edges in the spatial structure. In this paper, we propose Molormer, a method based on a lightweight attention mechanism for DDIs prediction. Molormer takes the two-dimension (2D) structures of drugs as input and encodes the molecular graph with spatial information. Besides, Molormer uses lightweight-based attention mechanism and self-attention distilling to process spatially the encoded molecular graph, which not only retains the multi-headed attention mechanism but also reduces the computational and storage costs. Finally, we use the Siamese network architecture to serve as the architecture of Molormer, which can make full use of the limited data to train the model for better performance and also limit the differences to some extent between networks dealing with drug features. Experiments show that our proposed method outperforms state-of-the-art methods in Accuracy, Precision, Recall and F1 on multi-label DDIs dataset. In the case study section, we used Molormer to make predictions of new interactions for the drugs Aliskiren, Selexipag and Vorapaxar and validated parts of the predictions. Code and models are available at https://github.com/IsXudongZhang/Molormer. Gan Wang, Xiangyu Meng 0005, Alfonso Rodríguez-Patón, Jianmin Wang 0016, Xun Wang 0010 |
Briefings Bioinform. | 3 |