VLDB 2026 Research / reviewers in the wild / expert
Shaoliang Peng
dblp:07/4720
· DBLP profile ↗
138ranked-venue papers
20as first author
85since 2021 · last 2026
0000-0002-4647-2615ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 94 · 13 first-author · 66 since 2021Systems, architecture and hardware · 16 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 12 · 8 since 2021Computer networks · 11 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DRAMA: A Dual-Phase Reliability-Aware Mamba Architecture for Multimodal Depression
Sijun Tan, Shaoliang Peng |
ICIC (15) | 2 |
| 2026 | MORM: Multi-omics-Guided Region Mining and Cross-Modal Interaction for Multimodal Survival Prediction
Jiadi Luo, Qingchun Liang, Shaoliang Peng |
ISBRA (1) | 5 |
| 2026 | QSyncFold: quantum neural network for multidimensional sync-discovery in protein foldingabstractQuantum computing provides alternative encoding and sampling paradigms for protein structure prediction (PSP), but existing quantum-PSP methods are often limited by resource-scaling issues and by discrete or inefficient encodings for continuous coordinates. To address these limitations, we propose QSyncFold, a hybrid quantum-classical neural network framework that combines quantum superposition with differentiable learning. QSyncFold employs ProtaQode to simultaneously achieve reversible continuous-space encoding of residue coordinates and parameterized interaction modeling. This is realized by encoding residue-pair interactions in superposition via a decomposable Any-State RY (ASRY) operator that is efficient for a limited qubit budget. Algorithmically, QSyncFold trades register size for iteration count, reducing the qubit requirement for each iteration from $O(N)$ to $3+\lceil \log _{2} N \rceil $, where $N$ is the number of residues. This design ensures the framework is experimentally viable under NISQ constraints. On short peptide structure prediction, QSyncFold achieved a 5.25-fold improvement in the lDDT metric compared with the Variational Quantum Eigensolver baseline and demonstrated a clear trade-off between qubit budget and convergence speed. While using quantum baselines as the primary comparison, the method performance approaches AlphaFold2 in the short peptide domain, with classical methods serving as background reference. This study advances the precision and methodology of quantum computing in PSP, illustrating a viable pathway for quantum algorithms in biomolecular modeling. Jinjing Shi, Wenwu Zeng, Shaoliang Peng |
Briefings Bioinform. | 5 |
| 2026 | Para-FDS: a scalable multilevel parallel scheme for fire dynamic simulator on multicore architectures
Dazheng Liu, Sheng Xiao, Xiaoli Ren, Wenjuan Liu, Dajiang Yi, Ze'an Tian, Yongan Wu, Zuodong Niu, Keqin Li 0001, Shaoliang Peng |
CCF Trans. High Perform. Comput. | 11 |
| 2026 | TransEHR: Alignment-free electronic health records continual learning across feature spaces
Xiongjun Zhao, Peng Xi, Shaoliang Peng |
Expert Syst. Appl. | 5 |
| 2025 | IBAS: Imperceptible Backdoor Attacks in Split Learning with Limited InformationabstractSplit learning, as a distributed learning framework, effectively addresses the issue of limited computing resources. However, despite achieving a separation of data and computation, recent studies have pointed out that this framework still faces two major security challenges: privacy leakage and model security. Most current research focuses on the problem of privacy leakage, emphasizing how to prevent malicious servers from recovering or inferring the client's private data. However, the issue of model security in split learning has not received sufficient attention. This paper reveals the vulnerability of split learning to backdoor attacks. Since split learning cannot access client data directly, it can only guide the client model to incorporate backdoors through gradients. To address this issue, we design an attack framework that modifies intermediate activations to influence the gradients. We designed a parrot model that learns the client’s feature space, enabling the server to obtain the intermediate activations of poisoned data. During the forward pass, some of the intermediate activations and labels transmitted from the client to the server are replaced with poisoned activations and target labels. This replacement method effectively integrates the backdoor task into the model while partially retaining the main task. This approach ensures that the main task is preserved while seamlessly embedding the backdoor task. Our attack framework minimizes reliance on client knowledge and ensures that the attack process remains undetectable by the client. Through extensive experiments, we demonstrated high attack success rates using triggers such as BadNet, SIG, Blended, and WaNet, while minimizing the impact on the main task. Peng Xi, Shaoliang Peng, Wenjuan Tang |
AAAI | 2 |
| 2025 | Accurate Nucleic Acid-Binding Residue Identification Based Domain-Adaptive Protein Language Model and Explainable Geometric Deep LearningabstractProtein-nucleic acid interactions play a fundamental and critical role in a wide range of life activities. Accurate identification of nucleic acid-binding residues helps to understand the intrinsic mechanisms of the interactions. However, the accuracy and interpretability of existing computational methods for recognizing nucleic acid-binding residues need to be further improved. Here, we propose a novel method called GeSite based the domain-adaptive protein language model and E(3)-equivariant graph neural network. Prediction results across multiple benchmark test sets demonstrate that GeSite is superior or comparable to state-of-the-art prediction methods. The MCC values of GeSite are 0.522 and 0.326 for the one DNA-binding residue test set and one RNA-binding resi-due test set, which are 0.57 and 38.14% higher than that of the second-best method, respectively. Detailed experi-mental results suggest that the advanced performance of GeSite lies in the well-designed nucleic acid-binding pro-tein adaptive language model. Additionally, interpretabil-ity analysis exposes the perception of the prediction mod-el on various remote and close functional domains, which is the source of its discernment ability. Wenwu Zeng, Liangrui Pan, Shaoliang Peng |
AAAI | 5 |
| 2025 | MP-MIL: Multi-View Multiple Instance Learning with Positional Embedding to Predict PIK3CA MutationabstractPhosphatidylinositol-4, 5-Bisphosphate 3-Kinase Catalytic Subunit Alpha (PIK3CA) gene mutations are among the most common somatic mutations in cancer, particularly in hormone receptor-positive breast cancer, with a mutation rate as high as 40 %. They are crucial for guiding targeted therapies in precision medicine. However, traditional detection methods, such as tissue-based next-generation sequencing, are challenged by high costs, time consumption, and insufficient detection of low-frequency mutations. This study proposes a novel multi-view multiple instance learning model, named MP-MIL, for non-invasively predicting PIK3CA mutation status from whole slide images (WSIs). MP-MIL effectively captures the complex spatial relationships and heterogeneity of the tumor microenvironment by fusing multi-scale features with patch sizes of$256 \times 256$and$512 \times 512$. It introducing a regional multi-head self-attention mechanism (RMSA) and a position embedding for attention-based (PEAT). Furthermore, the multi-view feature concatenation (MVC) module integrates microscopic details with macroscopic contextual information, improving the model's adaptability to heterogeneous pathological data. Experimental results on three datasets, TCGA-LUAD, TCGA-LUSC, and TCGA-BRCA, demonstrate that MP-MIL outperforms seven baseline models on most performance metrics. Ablation experiments further validated the key role of multi-view feature integration and the PEAT module in improving prediction performance. MP-MIL provides an efficient and non-invasive method for PIK3CA mutation prediction, providing important support for the advancement of precision medicine. Guanting Li, Liangrui Pan, Xiaoyu Li 0008, Jiadi Luo, Qingchun Liang, Shaoliang Peng |
BIBM | 6 |
| 2025 | Multi-Task Prompt-Aware Therapeutic Peptide Generation by Protein Language ModelabstractTherapeutic peptides, such as antimicrobial peptides(AMPs) and anticancer peptides (ACPs), are highly selective, low-toxicity agents that hold strong clinical potential. Designing peptides with specific biological functions remains challenging due to the vast combinatorial sequence space and the difficulty of modeling functional specificity, despite increasing interest. While deep generative models have recently shown promise in peptide design, most of them lack mechanisms for controllable functional output and underutilize protein language models (PLMs), which capture rich sequence-level dependencies. In this work, we propose MPTPep, a multi-task, prompt-aware fine-tuning framework built upon the pretrained PLM ProGen2, enabling controllable and function-specific therapeutic peptide generation. We introduce symbolic hard prompt tokens to encode peptide activity types, allowing the model to learn explicit function-sequence mappings within a unified generative architecture. Moreover, by jointly training on both AMP and ACP datasets, we exploit their biological overlap to boost the model to transfer knowledge from the AMP-rich corpus to boost ACP task performance. Extensive experiments demonstrate that MPTPep consistently outperforms state-of-the-art peptide generation models across both AMP and ACP sequence generation tasks, achieving superior functional scores, better sequence stability, and greater sequence diversity. Haitao Zou 0001, Zhijin Wang, Shaoliang Peng |
BIBM | 4 |
| 2025 | ChaiSite: An Accurate Method for Protein-Nucleic Acid Binding Site Prediction Based on a Complex Structure Prediction NetworkabstractPredicting protein-nucleic acid binding sites is crucial for understanding the intrinsic mechanisms of protein-nucleic acid interactions, thereby aiding drug development and gene engineering. Although many computational methods have attempted to address this issue, accuracy and generalizability remain to be improved due to the lack of exploration of complex structure and sequence evolutionary information. Here, we present ChaiSite, a novel method that integrates the embedded knowledge of the large protein language model and a complex structure prediction model by a customized hybrid model that integrates an Equivariant Graph Neural Network and the Transformer for the first time. The proposed ChaiSite consistently outperforms existing approaches in predicting protein-nucleic acid binding sites across multiple datasets. On independent benchmark datasets DNA-129_Test, DNA-181_Test, RNA-117_Test, and RNA-285_Test, ChaiSite achieves AUC scores of$0.945,0.930$, 0.898, and 0.859, respectively, representing improvements of 0.43, 1.20, 4.30, and 0.94% over the second-best methods, demonstrating superior predictive accuracy. Wenwu Zeng, Haitao Zou 0001, Shaoliang Peng |
BIBM | 6 |
| 2025 | A-Mel: A Resource-Efficient Deep Learning Agent for Precise Early Melanoma DiagnosisabstractEarly and precise diagnosis of melanoma significantly improves patient prognosis. However, current diagnostic methods are challenged by hair occlusion, diverse lesion morphologies, and limited computational resources in clinical settings. In this paper, we propose A-Mel, a resource-efficient deep learning agent designed for precise early melanoma diagnosis. A-Mel integrates a unified agent framework comprising three sequential modules: hair artifact removal preprocessing, pathological lesion segmentation, and melanoma classification. Specifically, we employ a modified$\mathrm{U}^{2} \text{Net}++$architecture enhanced with depthwise separable convolutions to significantly reduce model parameters and computational load. Furthermore, we introduce a novel Dual-Path Spatial Attention (DPSA) mechanism integrated with Efficient Channel Attention (ECA), achieving comprehensive attention enhancement across multiple scales and dimensions. The network is further augmented with an Atrous Spatial Pyramid Pooling (ASPP) module to strengthen multi-scale feature representation. Experimental evaluations conducted on a test set of approximately 700 dermoscopic images with hair artifacts selected from the ISIC 2019 and ISIC 2020 datasets demonstrate that our A-Mel model achieves a state-of-the-art binary classification accuracy of 99.12 %, outperforming existing approaches through extensive comparative analyses. Our segmentation model is “extremely lightweight” in terms of parameters, volume, and speed, with excellent performance for edge deployment scenarios. The codes for A-Mel are made publicly available at https://github.com/Yaoooyu/A-Mel. Yaoyu Liu, Yingjie Cao, Denan Liu, Sike Chen, Shaoliang Peng |
BIBM | 7 |
| 2025 | DLiPath: A Benchmark for the Comprehensive Assessment of Donor Liver Based on Histopathological Image DatasetabstractPathologists' comprehensive evaluation of donor liver biopsies provides crucial information for accepting or discarding potential grafts. However, rapidly and accurately obtaining these assessments intraoperatively poses a significant challenge for pathologists. Features in donor liver biopsies, such as portal tract fibrosis, total steatosis, macrovesicular steatosis, and hepatocellular ballooning are correlated with transplant outcomes, yet quantifying these indicators suffers from substantial inter- and intra-observer variability. To address this, we introduce DLiPath, the first benchmark for comprehensive donor liver assessment based on a histopathology image dataset. We collected and publicly released 636 whole slide images from 304 donor liver patients at the Department of Pathology, the Third Xiangya Hospital, with expert annotations for key pathological features (including cholestasis, portal tract fibrosis, portal inflammation, total steatosis, macrovesicular steatosis, and hepatocellular ballooning). We selected nine state-of-the-art multiple-instance learning (MIL) models based on the DLiPath dataset as baselines for extensive comparative analysis. The experimental results demonstrate that several MIL models achieve high accuracy across donor liver assessment indicators on DLiPath, charting a clear course for future automated and intelligent donor liver assessment research. Data and code are available at https://github.com/panliangrui/liver. Liangrui Pan, Zhongyi Chen, Chenchen Nie, Ling Chu, Shaoliang Peng |
BIBM | 6 |
| 2025 | SpaceSeg: Spatially Feature-Aware Segmentation and Classification Model for Cell NucleiabstractDetecting and segmenting cell nuclei in Hematoxylin and Eosin (H&E) stained tissue images is a critical clinical task with broad applications. However, it is a challenging problem due to variations in staining and size, overlapping boundaries, cell clustering, and the high morphological and size variability of lesion regions in medical images. Accurate segmentation in medical imaging requires precise global contour localization and careful handling of local boundaries. Existing CNN-based and Transformer-based models are often limited by high parameter counts and computational complexity, making it difficult to effectively integrate these features. To address this challenge, we propose a spatially feature-aware segmentation and classification model for cell nuclei, named SpaceSeg. The Partial Gated Feed-forward Network module in SpaceSeg enhances feature representations, enabling the model to focus on key features while reducing attention to redundant information. The Spatial Attention Block enhances the spatial information of features passed to the decoder, allowing the network to automatically focus on the most critical regions in the image, especially those related to nuclear boundaries. The Partial Gated CNN module more accurately fuses low-level and high-level features, helping the model learn finer semantic information. Experimental results on the PanNuke and MoNuSeg datasets demonstrate that SpaceSeg outperforms all existing state-of-the-art models. Yijun Peng, Liangrui Pan, Jiadi Luo, Christopher Wang, Qingchun Liang, Shaoliang Peng |
BIBM | 7 |
| 2025 | MSI-GNN: A Graph Neural Network for Pathogenicity Prediction of Microsatellite InsertionsabstractMicrosatellite insertions (MSIs) are a common type of genetic variation implicated in various hereditary disorders and cancers. However, due to their repetitive structure, high sequence variability, and limited annotation resources, the pathogenicity of MSIs remains difficult to determine in clinical settings, hindering their utility in genetic diagnosis and disease mechanism studies. To address this challenge, we propose MSIGNN, a graph neural network model that integrates multidimensional annotations and graph attention mechanisms to predict the pathogenicity of MSI events. MSI-GNN constructs a comprehensive feature profile for each MSI by incorporating heterogeneous annotations, including genomic functions, epigenetic signals, deleteriousness scores, functional constraints, evolutionary conservation, splicing effects, and molecular consequences. A multi-layer graph attention network is then employed to model biological correlations among variants and extract informative pathogenicity representations. Experimental results demonstrate that MSI-GNN significantly outperforms existing general-purpose models across multiple evaluation metrics, offering superior predictive performance and interpretability. Our model provides a promising tool for elucidating the pathogenic mechanisms of MSIs and advancing precision medicine. Yaning Yang, Yadong Fan, Liangrui Pan, Shaoliang Peng |
BIBM | 5 |
| 2025 | GEMIL: A GELU-Enhanced Multiple-Instance Learning Model for Predicting Gene Mutations in Lung CancerabstractLung cancer remains a leading cause of cancer mortality globally. Recently, targeted therapies have significantly improved clinical outcomes in lung cancer patients, making accurate identification of driver gene mutations crucial for precision treatment. Artificial intelligence approaches leveraging routinely acquired whole-slide histopathology images (WSIs) offer a promising and cost-effective means for molecular biomarker prediction, potentially enhancing clinical decision-making. However, mutation prediction from WSIs faces substantial technical challenges, including high data dimensionality, weakly supervised labels, and severe class imbalance inherent in clinical datasets. To address these issues, we propose GEMIL, a novel multipleinstance learning (MIL) model. GEMIL features a hierarchical attention-based encoder that employs deep non-linear projections and structured regularization to learn discriminative patch-level representations. These features are then aggregated by a querydriven decoder, which efficiently consolidates instance information into a robust slide-level prediction. We validated GEMIL through extensive experiments on the large-scale PathGene dataset, with further evaluation on an independent TCGA cohort. The results demonstrate that GEMIL consistently outperforms state-of-the-art MIL methods across multiple gene mutation prediction tasks (TP53, EGFR, etc.), improving the average accuracy by 2.8% and the F1-score by 3.1%. Consequently, GEMIL provides a robust and generalizable computational tool for WSI-based biomarker prediction, holding significant potential to advance precision oncology. Haihua Zhu 0003, Liangrui Pan, Christopher Wang, Jiadi Luo, Qingchun Liang, Shaoliang Peng |
BIBM | 6 |
| 2025 | UniOMP: Unified Optimization Framework for OpenMP Offload Under Machine Learning GuidanceabstractHeterogeneous programming is a critical approach to unleashing the parallel computing potential of GPUs, playing a vital role in fields requiring massive computational resources, such as scientific computing, artificial intelligence, and computer graphics. While OpenMP simplifies parallel programming and offers greater flexibility and portability compared to CUDA and Triton, it still faces challenges in fully leveraging hardware resources for efficient task execution. We propose UniOMP, an OpenMP-based optimization framework designed to enhance kernel launch efficiency and deeply optimize parallel tasks across diverse computational patterns. By integrating machine learning algorithms to select optimal parameters and orchestrate optimization passes, UniOMP seamlessly combines traditional compilation with ML-driven strategies, achieving holistic coordination of fine-grained and global optimizations. Experimental results demonstrate that UniOMP achieves speedups of 1.74x on AMDGPU and 1.34x on NVIDIA platforms for PolyBench, and 1.33x (AMDGPU) and 1.24x (NVIDIA) for SPEC ACCEL. Additionally, it reduces average register usage by 25 %, maximizing the efficient utilization of GPU hardware resources. Shaoliang Peng |
HiPC | 3 |
| 2025 | SMILE: A Scale-aware Multiple Instance Learning Method for Multicenter STAS Lung Cancer Histopathology DiagnosisabstractSpread through air spaces (STAS) represents a newly identified aggressive pattern in lung cancer, which is known to be associated with adverse prognostic factors and complex pathological features. Pathologists currently rely on time-consuming manual assessments, which are highly subjective and prone to variation. This highlights the urgent need for automated and precise diagnostic solutions. 2,970 lung cancer tissue slides are comprised from multiple centers, re-diagnosed them, and constructed and publicly released three lung cancer STAS datasets: STAS-CSU (hospital), STAS-TCGA, and STAS-CPTAC. All STAS datasets provide corresponding pathological feature diagnoses and related clinical data. To address the bias, sparse and heterogeneous nature of STAS, we propose an scale-aware multiple instance learning(SMILE) method for STAS diagnosis of lung cancer. By introducing a scale-adaptive attention mechanism, the SMILE can adaptively adjust high-attention instances, reducing over-reliance on local regions and promoting consistent detection of STAS lesions. Extensive experiments show that SMILE achieved competitive diagnostic results on STAS-CSU, diagnosing 251 and 319 STAS samples in CPTAC and TCGA, respectively, surpassing clinical average AUC. The 11 open baseline results are the first to be established for STAS research, laying the foundation for the future expansion, interpretability, and clinical integration of computational pathology technologies. The datasets and code are available at https://github.com/panliangrui/IJCAI25. Liangrui Pan, Xiaoyu Li 0008, Yutao Dou, Qiya Song, Jiadi Luo, Qingchun Liang, Shaoliang Peng |
IJCAI | 7 |
| 2025 | Pre-trained Latent Diffusion Model-based Image Enhancement with Cross-Fusion Transformer Network for Expression RecognitionabstractFacial Expression Recognition (FER) has broad application potential in fields such as education, human-computer interaction, healthcare, and Online monitoring. However, the FER task faces significant challenges, including missing, incomplete, and tilted facial samples in the data, as well as inter-class similarity and intra-class disparity in facial expressions, which pose challenges to recognition. To address these issues, we propose a two-stage method called LDM-POSTER, which integrates Latent Diffusion Models (LDM) and Prompt Engineering for data augmentation. The enhanced images are then fed into a dual-stream Pyramid Cross-Fusion Transformer Network (POSTER) for recognition. Specifically, the LDM and prompt engineering stages effectively repair missing or incomplete facial regions and correct tilted poses, resulting in more complete, clear, and information-rich input data. In the recognition stage, POSTER utilizes a transformer-based cross-fusion structure to integrate facial landmark features with global image features, effectively distinguishing inter-class similarities. By focusing on salient facial regions and employing a pyramid structure to achieve scale invariance, it alleviates the impact of intra-class disparity on recognition accuracy.Within extensive experiments demonstrate that LDM-POSTER achieves significant performance improvements in FER scenarios with missing facial key parts and complex structures, providing a robust and efficient solution for FER. Yutao Dou, Changfeng He, Xianliang Chen, Jiansong Zhou, Shaoliang Peng |
IJCNN | 6 |
| 2025 | DepMambaformer: Integrating Bidirectional State Space Duality Model with Multimodal Attention for Depression Detection
Changfeng He, Yutao Dou, Shaoliang Peng |
ISBRA (1) | 5 |
| 2025 | MegSite: an accurate nucleic acid-binding residue prediction method based on multimodal protein language modelabstractAccurate identification of nucleic acid-binding residues is crucial for understanding protein-nucleic acid interactions, which play a key role in gene expression research and the discovery of regulatory mechanisms. Despite numerous computational efforts to address this challenge, achieving high accuracy remains difficult due to the complexity of extracting meaningful insights from proteins. Here, we introduce MegSite, a novel multimodal protein language model-informed method that integrates discriminative knowledge from protein sequence, structure, and function. This work presents the first integration of ESM3 multimodal features for nucleic acid-binding site prediction. MegSite significantly outperforms existing prediction methods, as evidenced by its performance on multiple independent test sets. The Matthews correlation coefficient values achieved by MegSite on DNA-129_Test, DNA-181_Test, RNA-117_Test, and RNA-285_Test are 0.567, 0.444, 0.411, and 0.421, representing the improvements of 2.72%, 7.66%, 1.22% and 6.58% over the second-best method separately. Notably, MegSite demonstrates robust performance even on proteins with low structural similarity, surpassing the previous structure-based methods. Furthermore, this method is seamlessly extendable to the predicted protein structure and a newly released RNA-binding residue test set with high accuracy, highlighting its broad applicability. Comprehensive experimental results reveal that the superior performance of MegSite is attributed to its effective integration of multimodal protein knowledge. Wenwu Zeng, Shaoliang Peng |
Briefings Bioinform. | 3 |
| 2025 | scDCA: deciphering the dominant cell communication assembly of downstream functional events from single-cell RNA-seq dataabstractCell-cell communications (CCCs) involve signaling from multiple sender cells that collectively impact downstream functional processes in receiver cells. Currently, computational methods are lacking for quantifying the contribution of pairwise combinations of cell types to specific functional processes in receiver cells (e.g. target gene expression or cell states). This limitation has impeded understanding the underlying mechanisms of cancer progression and identifying potential therapeutic targets. Here, we proposed a deep learning-based method, scDCA, to decipher the dominant cell communication assembly (DCA) that have a higher impact on a particular functional event in receiver cells from single-cell RNA-seq data. Specifically, scDCA employed a multi-view graph convolution network to reconstruct the CCCs landscape at single-cell resolution, and then identified DCA by interpreting the model with the attention mechanism. Taking the samples from advanced renal cell carcinoma as a case study, the scDCA was successfully applied and validated in revealing the DCA affecting the crucial gene expression in immune cells. The scDCA was also applied and validated in revealing the DCA responsible for the variation of 14 typical functional states of malignant cells. Furthermore, the scDCA was applied and validated to explore the alteration of CCCs under clinical intervention by comparing the DCA for certain cytotoxic factors between patients with and without immunotherapy. In summary, scDCA provides a valuable and practical tool for deciphering the cell type combinations with the most dominant impact on a specific functional process of receiver cells, which is of great significance for precise cancer treatment. Our data and code are free available at a public GitHub repository: https://github.com/pengsl-lab/scDCA.git. Shaoliang Peng |
Briefings Bioinform. | 5 |
| 2025 | CellMsg: graph convolutional networks for ligand-receptor-mediated cell-cell communication analysisabstractThe role of cell-cell communications (CCCs) is increasingly recognized as being important to differentiation, invasion, metastasis, and drug resistance in tumoral tissues. Developing CCC inference methods using traditional experimental methods are time-consuming, labor-intensive, cannot handle large amounts of data. To facilitate inference of CCCs, we proposed a computational framework, called CellMsg, which involves two primary steps: identifying ligand-receptor interactions (LRIs) and measuring the strength of LRIs-mediated CCCs. Specifically, CellMsg first identifies high-confident LRIs based on multimodal features of ligands and receptors and graph convolutional networks. Then, CellMsg measures the strength of intercellular communication by combining the identified LRIs and single-cell RNA-seq data using a three-point estimation method. Performance evaluation on four benchmark LRI datasets by five-fold cross validation demonstrated that CellMsg accurately captured the relationships between ligands and receptors, resulting in the identification of high-confident LRIs. Compared with other methods of identifying LRIs, CellMsg has better prediction performance and robustness. Furthermore, the LRIs identified by CellMsg were successfully validated through molecular docking. Finally, we examined the overlap of LRIs between CellMsg and five other classical CCC databases, as well as the intercellular crosstalk among seven cell types within a human melanoma tissue. In summary, CellMsg establishes a complete, reliable, and well-organized LRI database and an effective CCC strength evaluation method for each single-cell RNA-seq data. It provides a computational tool allowing researchers to decipher intercellular communications. CellMsg is freely available at https://github.com/pengsl-lab/CellMsg. Hong Xia, Debin Qiao, Shaoliang Peng |
Briefings Bioinform. | 4 |
| 2025 | Application of deep learning-based multimodal fusion technology in cancer diagnosis: A survey
Liangrui Pan, Yijun Peng, Xiaoyu Li 0008, Limeng Qu, Qiya Song, Qingchun Liang, Shaoliang Peng |
Eng. Appl. Artif. Intell. | 9 |
| 2025 | ECG-I2S: a method for extracting heartbeat cycle and numerical signals from ECG captured images
Xiongjun Zhao, Linzhuang Zou, Shaoliang Peng |
Frontiers Comput. Sci. | 4 |
| 2025 | Parallel Acceleration of Genome Variation Detection on Multi-Zone Heterogeneous SystemabstractGenomic variation is critical for understanding the genetic basis of disease. Pindel, a widely used structural variant caller, leverages short-read sequencing data to detect variation at single-base resolution; however, its hotspot module imposes substantial computational demands, limiting efficiency in large-scale whole-genome analyses. Heterogeneous architectures offer a promising solution, yet disparities in hardware design and programming models preclude direct porting of the original algorithm. To address this, we introduce MTPindel, a novel heterogeneous parallel optimization framework tailored to the MT-3000 processor. Focusing on Pindel's most compute-intensive modules, we design multi-core and task-level parallel algorithms that exploit the MT-3000's accelerator domains to balance and accelerate workload distribution. On 128 MT-3000–equipped nodes of the Tianhe next-generation supercomputer, MTPindel achieves an impressive 122.549 times of speedup and 95.74% parallel efficiency, with only a 0.74% error margin relative to the original implementation. This work represents a pioneering effort in heterogeneous parallelization for variant detection, paving the way for rapid, large-scale genomic analyses in research and clinical settings. Yaning Yang, Chengqing Li, Shaoliang Peng |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2024 | MRF-XGBLC: Large-scale gene regulatory network inference based on multi-model fusionabstractGene regulatory networks (GRNs) are crucial for revealing gene interactions and understanding cellular biological mechanisms. However, the high dimensionality and nonlinearity of gene expression data make accurate inference and reconstruction of large-scale GRNs a core computational challenge in systems biology. This study introduces a novel approach, termed MRF-XGBLC, for the reconstruction of large-scale Gene Regulatory Networks (GRNs) utilizing steady-state and time-series gene expression data through nonlinear ordinary differential equations. Firstly, MRF-XGBLC uses the maximum information coefficient (MIC) for dimensionality reduction and eliminates redundant regulatory relationships by calculating the MIC between factors as a prior step in model processing. Furthermore, recognizing the superior performance of the Lasso-Cox model in survival analysis, the feature fusion algorithm of this paper incorporates a hybrid model of XGBoost (eXtreme Gradient Boosting), RF (Random Forest), and Lasso-Cox (Least Absolute Shrinkage and Selection Operator-Cox proportional hazards regression model) integration to effectively train the nonlinear ordinary differential equations, thus improving the accuracy and stability of the inference algorithm. Extensive experiments on datasets of varying sizes demonstrate significant improvements over state-of-the-art methods. Cross-validation experiments on real gene datasets confirm the robustness and effectiveness of MRF-XGBLC. Shulin Wang, Shaoliang Peng |
BIBM | 5 |
| 2024 | Autonomous Pharmaceutical Care with Large Language ModelsabstractIn modern healthcare systems, physicians often diagnose and prescribe within their specific fields of expertise. However, this compartmentalized approach may not fully assess the overall health condition of patients, leading to potential risks associated with polypharmacy, such as drug-drug interactions. Consequently, pharmaceutical care becomes a critical systemic task, typically managed by clinical pharmacists. Despite their crucial role, clinical pharmacists face challenges in providing personalized medication guidance due to the vast and rapidly evolving body of medical knowledge. With the advancement of artificial intelligence and large language models (LLMs), these technologies have demonstrated significant potential in pharmaceutical care by effectively reducing medication errors and adverse drug events, and alleviating the workload of healthcare professionals. However, existing LLMs still struggle with the accurate interpretation and automated execution of tasks in complex clinical scenarios like pharmaceutical care. To address these challenges, we developed the Shennong-Agent system, a multi-agent framework for LLMs with integrated multimodal inputs. Shennong-Agent is enabled to analyze and segment tasks using chained reasoning and autonomously perform complex pharmacy care tasks using tools such as knowledge retrieval and web search. Rigorously evaluated by medical experts, this system not only surpasses existing LLMs in performance but also enhances its capabilities through reinforcement learning with human feedback. Yutao Dou, Zike Deng, Tao Xing, Shaoliang Peng |
BIBM | 5 |
| 2024 | Interpretable deep learning quantifies the impact of genomic variations on CD8+ T cell infiltrationabstractImmune checkpoint inhibitors (ICIs) have revolutionized cancer treatment, but challenges remain in identifying molecular determinants of clinical response. To address this, we developed a Transformer-based model, ImmRegVar, to predict CD8+ T cell infiltration using somatic genomic variations from 7972 solid tumor samples. ImmRegVar outperforms traditional regression methods in predictive accuracy and model performance. We introduce an innovative explanatory factor that combines Transformer attention weights with feature ablation techniques, allowing precise quantification of each gene mutation’s contribution. This factor offers an interpretable metric for understanding how specific mutations affect CD8+ T cell infiltration, shedding light on immune regulatory mechanisms. Validation through random permutation tests confirms its robustness and stability in high-dimensional sparse data. Our analysis also aligns with known immune regulatory genes, supporting the biological credibility of ImmRegVar. This research advances the understanding of key factors in immunotherapy and opens new pathways for exploring immune regulation and clinical applications. ImmRegVar and the explanatory factor are available at //github.com/Liwen-Liberty/ImmRegVar. Shaoliang Peng |
BIBM | 4 |
| 2024 | SeqCCC: A Sequence-Informed Ligand-Receptor Interaction Prediction Method for Cell-Cell Communication InferenceabstractCell-cell communications (CCCs) mediated by ligand-receptor interactions (LRIs) play a pivotal role in coordination and function of biological systems. The two primary steps of CCC inference methods are usually filtering significant LRIs and measuring the intercellular communication strength. Biological language models have demonstrated notable progress in bioinformatics across several biological problems, inspiring by the success of language models in the field of natural language processing. Here, we proposed a computational approach for CCC inference called SeqCCC. First, SeqCCC employed two sophisticated protein language models and the protein sequences of ligands and receptors to predict potential LRIs with the help of existing LRIs. Then, SeqCCC employed permutation test to filter significantly expressed LRIs from the single-cell expression matrix based on these known and predicted LRIs. Finally, SeqCCC calculated communication strength by applying the molecular diffusion and law of mass action in chemistry, and then it improves this by eliminating non-specific CCCs using a permutation test. In the results, SeqCCC yielded an average AUC of 0.93 (mean of the five folds, with a standard deviation of 0.02). The validity and reliability of SeqCCC are further supported by a comparative analysis with common CCC inference methods, which reveals high concordance in inference results. Additionally, SeqCCC provided different visualization methods for CCC results, including Circos Plot, Heatmap Plot, Dot Plot and Heatmap for LRIs. In conclusion, SeqCCC provides a new option for CCC inference by fusing biological language models with sequence information to improve our comprehension of intercellular communications. Hong Xia, Dezun Dong, Shaoliang Peng |
BIBM | 4 |
| 2024 | ParaSAT: a scalable parallel framework for sequence alignmentabstractSequence alignment is a fundamental step in genomic data analysis. Third-generation sequencing technology facilitates the acquisition of high-quality genomic data but the explosive growth of sequencing data poses huge challenges to current sequence alignment. To reduce sequence alignment time and enhance alignment performance, a parallelization framework based on Minimap2 is proposed, called ParaSAT, aiming to expedite sequence alignment and offer insights to researchers in the field. In order to achieve load balancing among multi-nodes, We design a task pool scheduling strategy which can dynamically distribute tasks according to the status of compute nodes. To evaluate the performance of the framework, We choose 6 dataset to conduct experiments on TH-1 supercomputer. The outcomes confirm that the parallel framework ensures sequence alignment accuracy, and demonstrates a notable speedup on large datasets, approaching linearity, with parallel efficiency consistently above 80%. The framework also exhibits strong and weak scalability, effectively enhancing the efficiency of sequence alignment and offering guidance for high-performance genomic data processing. Wenjuan Liu, Jianbang Xu, Liangrui Pan, Tie Cai, Shaoliang Peng |
BIBM | 5 |
| 2024 | FedDP: Privacy-preserving method based on federated learning for histopathology image segmentationabstractHematoxylin and Eosin (H&E) staining of whole slide images (WSIs) is considered the gold standard for pathologists and medical practitioners for tumor diagnosis, surgical planning, and post-operative assessment. With the rapid advancement of deep learning technologies, the development of numerous models based on convolutional neural networks and transformer-based models has been applied to the precise segmentation of WSIs. However, due to privacy regulations and the need to protect patient confidentiality, centralized storage and processing of image data are impractical. Training a centralized model directly is challenging to implement in medical settings due to these privacy concerns.This paper addresses the dispersed nature and privacy sensitivity of medical image data by employing a federated learning framework, allowing medical institutions to collaboratively learn while protecting patient privacy. Additionally, to address the issue of original data reconstruction through gradient inversion during the federated learning training process, differential privacy introduces noise into the model updates, preventing attackers from inferring the contributions of individual samples, thereby protecting the privacy of the training data.Experimental results show that the proposed method, FedDP, minimally impacts model accuracy while effectively safeguarding the privacy of cancer pathology image data, with only a slight decrease in Dice, Jaccard, and Acc indices by 0.55%, 0.63%, and 0.42%, respectively. This approach facilitates cross-institutional collaboration and knowledge sharing while protecting sensitive data privacy, providing a viable solution for further research and application in the medical field. Liangrui Pan, Mao Huang, Pinle Qin, Shaoliang Peng |
BIBM | 5 |
| 2024 | SpaChat: Integrating Single-Cell Foundation Model with Cell Graph Network for Spatially Resolved Cell-Cell Communication InferenceabstractCurrently, single-cell foundation models, which are based on extensive single-cell sequencing data, are highly effective in extracting crucial biological information about genes and cells. They consistently demonstrate exceptional performance across a wide range of downstream applications. However, when it comes to spatially resolved transcriptomic data (ST), the availability of methods for inferring cell-cell communications (CCCs), in conjunction with foundation models, remains limited. In this work, we present SpaChat, which integrating a fine-tuned single-cell foundation model with a cell graph network for spatially resolved CCCs inference. SpaChat defines two scoring strategies to infer significant ligand-receptor (LR) pairs, including intracellular score based on gene-level attention of fine-tuned single-cell foundation model and inter-cellular score based on the KNN algorithm in a cell-cell graph network. On this basis, SpaChat employs the law of molecular diffusion and mass action in chemistry and the strategy of permutation test to calculate communication strengths and filter out communications with low specificity. The benchmarked performance of SpaChat on public spatial transcriptomic dataset is superior to that of existing inference methods. Furthermore, SpaChat subsequently identifies the communication patterns of specific cell types and provides a variety of options for visualizing the results of CCC analyses. In summary, SpaChat enables the inference of spatially resolved CCCs from spatial transcriptomic data, providing valuable insights into understanding CCCs in tissues. Debin Qiao, Hong Xia, Shaoliang Peng |
BIBM | 4 |
| 2024 | MinimapPool: an improved flexible and efficient parallel algorithm based on minimap2abstractThird-generation sequencing techniques have achieved major breakthroughs in sequencing long reads and speed. Continuous improvements in sequencing techniques have reduced sequencing costs, and the number of sequencing data files has shown explosive growth. In terms of sequence alignment, to deal with these high numbers and large-scale data, the conventional serial alignment method can no longer effectively meet the research requirements, therefore, it is of great importance to develop a faster, low-load, and compatible parallel alignment program. In this paper, we propose a parallel task pool algorithm based on the minimap2, a sequence alignment tool, and develop the task pool parallel alignment program based on this algorithm. We compare the program’s work with the average segmentation parallel alignment program. The results show that the task pool parallel alignment program has significant improvement in speedup, memory load, segmentation flexibility, and computational efficiency, it also has good scalability and computational stability. MinimapPool is available at https://github.com/krkrcc/MinimapPool. Zhenang Wang, Yingbo Cui 0001, Jiandong Shang, Shaoliang Peng |
BIBM | 6 |
| 2024 | EMO-Mamba: Multimodal Selective Structured State Space Model for Depression DetectionabstractDepression is a severe mental disorder that significantly affects health and quality of life, making early detection critical for effective treatment and management. In traditional clinical diagnostics, doctors and mental health professionals typically assess emotional states by observing patients’ facial expressions and listening to their speech. Changes in facial expressions can reflect emotional responses, while vocal features such as tone and pace may indicate underlying emotional distress or psychological conditions. Artificial intelligence technologies are adept at efficiently extracting and analyzing these multimedia signals. However, existing methods often overlook critical information when processing long-time series data and struggle to handle multiple modalities simultaneously. To address these issues, we propose the EMO-Mamba, which utilizes selective mechanisms and state space models (SSM) to filter important features and effectively maintain memory over long time-series. Additionally, we introduced a multimodal data fusion framework to integrate key features from various modals, thereby enhancing model performance. Evaluations on the multimodal public datasets D-Vlog showed that our method achieved accuracy of 75.54%, demonstrating its effectiveness across diverse data environments and optimizing detection outcomes. Tao Xing, Yutao Dou, Jiansong Zhou, Xianliang Chen, Shaoliang Peng |
BIBM | 6 |
| 2024 | TranSVPath: A TabTransformer-Based Model for Predicting the Pathogenicity of Structural VariantsabstractGenomic structural variants are recognized as critical molecular factors contributing to various major diseases, including cancer and genetic disorders. However, accurately determining whether a variant can cause a disease in clinical settings is extremely challenging due to limited sample sizes, the diversity of variant types, and the complexity of the mechanisms linking variants to diseases. Existing computational tools attempt to predict the pathogenic effects of these variants, but they often consider only single-layer biological data, limiting their ability to comprehensively explain the functional impacts of the variants. To address this, we propose TranSVPath, a pathogenicity scoring tool for structural variants based on the Transformer framework. TranSVPath provides a more comprehensive biological annotation of structural variations by integrating multi-dimensional data, including overlap with specific genomic regions, single nucleotide variant deleteriousness scores, phylogenetic conservation scores, Mendelian clinical application pathogenicity scores, and gene function loss. Additionally, TranSVPath employs a variant of the Transformer architecture tailored to this integrated data. Its attention mechanism captures key biological features, facilitating molecular-level interpretation of the pathogenic mechanisms of variants. Consequently, TranSVPath can accurately predict the pathogenicity of deletions, insertions, tandem duplications, inversions, and microsatellite insertions. Evaluations on large-scale human structural variant datasets demonstrate that TranSVPath outperforms state-of-the-art algorithms across relevant metrics. Yaning Yang, Liangrui Pan, Shaoliang Peng |
BIBM | 6 |
| 2024 | Multivariate Time-Series Representation Learning for Continuous Medical DiagnosisabstractMultivariate time series (MTS) data in electronic health records (EHR) pose unique challenges due to their sparsity and irregular time intervals. Existing methods often tend towards imputation or isolated encoding, resulting in suboptimal representation learning. Moreover, most existing research solely focuses on single-shot diagnosis, neglecting the importance of continuous diagnosis, particularly for critically ill patients. Continuous diagnosis provides significant opportunities for timely intervention and rational resource allocation. To address these challenges, we propose an innovative multivariate time-series representation learning for continuous medical diagnosis. Specifically, we first address sparsity issues by combining feature names and record values encoded by time. Then, we utilize a transformer variant with gated units to extract contextual features. Additionally, we introduce Time Update Block, a component that combines the strengths of long short-term memory and attention mechanisms, aimed at improving the model's ability for continuous diagnosis. Based on extensive experimental evaluations on real-world medical datasets, we demonstrate the superior performance of the proposed method. Xiongjun Zhao, Linzhuang Zou, Long Ye, Shaoliang Peng |
BIBM | 5 |
| 2024 | Pipe-AGCM: A Fine-Grain Pipelining Scheme for Optimizing the Parallel Atmospheric General Circulation Model
Dazheng Liu, Xiaoli Ren, Wenjuan Liu, Juan Zhao 0006, Shaoliang Peng |
Euro-Par (3) | 6 |
| 2024 | Fighting Fire with Fire: Medical AI Models Defend Against Backdoor Attacks via Self-learning
Peng Xi, Wenjuan Tang, Shaoliang Peng |
ISBRA (1) | 3 |
| 2024 | Report-Concept Textual-Prompt Learning for Enhancing X-ray DiagnosisabstractDespite significant advances in image-text medical visual language modeling, the high cost of fine-grained annotation of images to align radiology reports has led current approaches to focus primarily on semantic alignment between the image and the full report, neglecting the critical diagnostic information contained in the text. This is insufficient in medical scenarios demanding high explainability. To address this problem, in this paper, we introduce radiology reports as images in prompt learning. Specifically, we extract key clinical concepts, lesion locations, and positive labels from easily accessible radiology reports and combine them with an external medical knowledge base to form fine-grained self-supervised signals. Moreover, we propose a novel Report-Concept Textual-Prompt Learning ( RC-TPL ), which aligns radiology reports at multiple levels. In the inference phase, the report-level and concept-level prompts provide rich global and local semantic understanding for X-ray images. Extensive experiments on X-ray image datasets demonstrate the superior performance of our approach with respect to various baselines, especially in the presence of scarce imaging data. Our study not only significantly improves the accuracy of data-constrained medical X-ray diagnosis, but also demonstrates how the integration of domain-specific conceptual knowledge can enhance the explainability of medical image analysis. Xiongjun Zhao, Guanting Li, Yutao Dou, Shaoliang Peng |
ACM Multimedia | 6 |
| 2024 | MUSCLE: multi-view and multi-scale attentional feature fusion for microRNA-disease associations predictionabstractMicroRNAs (miRNAs) synergize with various biomolecules in human cells resulting in diverse functions in regulating a wide range of biological processes. Predicting potential disease-associated miRNAs as valuable biomarkers contributes to the treatment of human diseases. However, few previous methods take a holistic perspective and only concentrate on isolated miRNA and disease objects, thereby ignoring that human cells are responsible for multiple relationships. In this work, we first constructed a multi-view graph based on the relationships between miRNAs and various biomolecules, and then utilized graph attention neural network to learn the graph topology features of miRNAs and diseases for each view. Next, we added an attention mechanism again, and developed a multi-scale feature fusion module, aiming to determine the optimal fusion results for the multi-view topology features of miRNAs and diseases. In addition, the prior attribute knowledge of miRNAs and diseases was simultaneously added to achieve better prediction results and solve the cold start problem. Finally, the learned miRNA and disease representations were then concatenated and fed into a multi-layer perceptron for end-to-end training and predicting potential miRNA-disease associations. To assess the efficacy of our model (called MUSCLE), we performed 5- and 10-fold cross-validation (CV), which got average the Area under ROC curves of 0.966${\pm }$0.0102 and 0.973${\pm }$0.0135, respectively, outperforming most current state-of-the-art models. We then examined the impact of crucial parameters on prediction performance and performed ablation experiments on the feature combination and model architecture. Furthermore, the case studies about colon cancer, lung cancer and breast cancer also fully demonstrate the good inductive capability of MUSCLE. Our data and code are free available at a public GitHub repository: https://github.com/zht-code/MUSCLE.git. Haitao Zou 0001, Shaoliang Peng |
Briefings Bioinform. | 5 |
| 2024 | LBi-DBP, an accurate DNA-binding protein prediction method based lightweight interpretable BiLSTM network
Wenwu Zeng, Jiandong Shang, Wenjuan Liu, Shaoliang Peng |
Expert Syst. Appl. | 7 |
| 2024 | 3DSGIMD: An accurate and interpretable molecular property prediction method using 3D spatial graph focusing network and structure-based feature fusion
Chenbin Wang, Ruiqiang Lu, Henry H. Y. Tong, Xiaoqing Gong, Jiayue Qiu, Shaoliang Peng, Huanxiang Liu |
Future Gener. Comput. Syst. | 7 |
| 2023 | An Attention-based Label Mapping and Multi-factor Domain Adaptation Approach for ACS PredictionabstractAcute Coronary Syndrome (ACS), an emergent medical condition, is intricately linked to environmental factors like air pollution and meteorological conditions. Harnessing regional environmental data, such as weather metrics, can promptly forecast ACS incidence rates, enabling optimised medical resource allocation and increased patient recovery rates. However, the prediction task is rendered complex due to disparities in data collection capabilities across institutions, yielding datasets with analogous features but significant label variations, impeding the application of universal models. Challenges abound due to the heterogeneity of multi-factor data, temporal alignment disparities, and the intricacies of sparse data. To address these challenges, this paper introduces the Domain Adaptation with Multi-factor Associative Structures (DAMAS), a time-series domain adaptation approach based on multi-factor sparse associative frameworks. Augmented by an isomorphic attention-driven variable label mapping scheme and combined with multi-layer perceptrons, our approach skilfully negotiates label imbalances. This results in refined prediction precision connecting environmental factors to regional ACS incidences. Yutao Dou, Xiongjun Zhao, Kun Xie 0001, Guo Chen 0001, Shaoliang Peng |
BIBM | 6 |
| 2023 | ParaMET: A Parallel Framework for Efficient Medical Data Extraction on Tianhe-NG SupercomputerabstractIn the burgeoning realm of data-driven medical research, the escalating scale and intricacy of contemporary medical datasets frequently surpass the processing capabilities of traditional computational environments. Specifically, I/O bottlenecks have emerged as pivotal constraints in several research areas. In this paper, a data service framework is introduced that harnesses supercomputers to parallelly access multi-modal datasets and supports multi-node processing, called ParaMET. To enhance user accessibility, a web-based user interface has been integrated, allowing permitted researchers to effortlessly interact via their laptops, complemented by an API suite available through SDK for those adept with supercomputing for deeper data manipulation. Through extensive empirical validation, our framework manifests a remarkable performance elevation in multi-node supercomputing settings, achieving acceleration of up to approximately 1000x compared to existing methods. Yutao Dou, Yangtao Zheng, Dazheng Liu, Keqin Li 0001, Sheng Xiao, Shaoliang Peng |
BIBM | 6 |
| 2023 | LDCSF: Local depth convolution-based Swim framework for classifying multi-label histopathology imagesabstractHistopathological images are the gold standard for diagnosing liver cancer. However, the accuracy of fully digital diagnosis in computational pathology needs to be improved. In this paper, in order to solve the problem of multi-label and low classification accuracy of histopathology images, we propose a locally deep convolutional Swim framework (LDCSF) to classify multi-label histopathology images. In order to be able to provide local field of view diagnostic results, we propose the LDCSF model, which consists of a Swin transformer module, a local depth convolution (LDC) module, a feature reconstruction (FR) module, and a ResNet module. The Swin transformer module reduces the amount of computation generated by the attention mechanism by limiting the attention to each window. The LDC then reconstructs the attention map and performs convolution operations in multiple channels, passing the resulting feature map to the next layer. The FR module uses the corresponding weight coefficient vectors obtained from the channels to dot product with the original feature map vector matrix to generate representative feature maps. Finally, the residual network undertakes the final classification task. As a result, the classification accuracy of LDCSF for interstitial area, necrosis, non-tumor and tumor reached 0.9460, 0.9960, 0.9808, 0.9847, respectively. Liangrui Pan, Guo Chen 0001, Wenjuan Liu, Xuan Liu 0001, Shaoliang Peng |
BIBM | 6 |
| 2023 | CVFC: Attention-Based Cross-View Feature Consistency for Weakly Supervised Semantic Segmentation of Pathology ImagesabstractHistopathology image segmentation is the gold standard for diagnosing cancer, and can indicate cancer prognosis. However, histopathology image segmentation requires high-quality masks, so many studies now use image-level labels to achieve pixel-level segmentation to reduce the need for fine-grained annotation. To solve this problem, we propose an attention-based cross-view feature consistency end-to-end pseudo-mask generation framework named CVFC based on the attention mechanism. Specifically, CVFC is a three-branch joint framework composed of two Resnet38 and one Resnet50, and the independent branch multi-scale integrated feature map to generate a class activation map (CAM); in each branch, through down-sampling and The expansion method adjusts the size of the CAM; the middle branch projects the feature matrix to the query and key feature spaces, and generates a feature space perception matrix through the connection layer and inner product to adjust and refine the CAM of each branch; finally, through the feature consistency loss and feature cross loss to optimize the parameters of CVFC in co-training mode. After a large number of experiments, An IoU of 0.7122 and a fwIoU of 0.7018 are obtained on the WSSS4LUAD dataset, which outperforms HistoSegNet, SEAM, C-CAM, WSSS-Tissue, and OEEM, respectively. Liangrui Pan, Keqin Li 0001, Wenjuan Liu, Zhichao Feng, Shaoliang Peng |
BIBM | 6 |
| 2023 | PACS: Prediction and analysis of cancer subtypes from multi-omics data based on a multi-head attention mechanism modelabstractDue to the high heterogeneity and clinical characteristics of cancer, there are significant differences in multi-omic data and clinical characteristics among different cancer subtypes. Therefore, accurate classification of cancer subtypes can help doctors choose the most appropriate treatment options, improve treatment outcomes, and provide more accurate patient survival predictions. In this study, we propose a supervised multi-head attention mechanism model (SMA) to classify cancer subtypes successfully. The attention mechanism and feature sharing module of the SMA model can successfully learn the global and local feature information of multi-omics data. Second, it enriches the parameters of the model by deeply fusing multi-head attention encoders from Siamese through the fusion module. Validated by extensive experiments, the SMA model achieves the highest accuracy, F1 macroscopic, F1 weighted, and accurate classification of cancer subtypes in simulated, single-cell, and cancer multi-omics datasets compared to AE, CNN, and GNN-based models. Therefore, we contribute to future research on multi-omics data using our attention-based approach. Liangrui Pan, Pinle Qin, Pengfei Rong, Xiangxiang Zeng, Dazheng Liu, Shaoliang Peng |
BIBM | 6 |
| 2023 | RobustHealthFL: Robust Strategy Against Malicious Clients in Non-iid Healthcare Federated LearningabstractDue to the sensitive and confidential nature of healthcare data, it cannot be freely transmitted, resulting in the challenge of data silos. While federated learning offers a solution to this data island dilemma, it also faces potential threats from malicious client attacks. The system becomes more vulnerable when dealing with non-independent and identically distributed (non-IID) data. Hence, a robust federated learning framework is indispensable to defend against malicious attacks. In this study, we employ a dynamic feature extractor based on the sparse representation, and the parameters of each iteration construct the feature extractor for the subsequent parameters. Then, we use optimized k-means clustering to get benign clients and aggregate them. The experimental results show that the system robustness is lower for non-IID datasets compared to IID healthcare datasets. And our RobustHealthFL can significantly enhance the robustness of the system when facing non-IID data. Our code is available at https://github.com/xipengp/RobustHealthFL. Peng Xi, Wenjuan Tang, Kun Xie 0001, Xuan Liu 0001, Shaoliang Peng |
BIBM | 6 |
| 2023 | ImmRegInformer: an interpretable Transformer-based method for prioritizing immune-regulatory cancer driversabstractThe identification and prioritization of immune-regulatory cancer driver mutations present a promising study for precision immunotherapy of cancer but remain considerable challenges. Here we introduced a novel method ImmRegInformer to systematically explore the regulatory relationship between cancer driver mutations and immune response by leveraging the powerful Transformer model and the lasso-regularised ordinal regression. In particular, our method integrated the mutation co-occurrence information with the self-attention weight to discern the underlying relationships between different driver mutations when regulating the immune cytolytic activity (CYT). Using ImmRegInformer, we identified 250 immune-regulating driver mutations in 8223 pan-cancer samples. They were verified in terms of the mutation frequency and interactions with the cytolytic signature genes. Further, we found the complementary roles of self-attention weight and mutation co-occurrence in prioritizing the driver mutations exhibited dominant associations with CYT. In conclusion, this study underscored the importance of employing deep learning methods like Transformer to unlock hidden insights into the biological complexities of cancer immunity, offering a new avenue in dissecting the immune regulatory mechanism and potential clinical applications. ImmRegInformer is freely available at https://github.com/Liwen-Liberty/ImmregInformer. Yijun Peng, Wending Pi, Xiangxiang Zeng, Shaoliang Peng |
BIBM | 7 |
| 2023 | ESM-NBR: fast and accurate nucleic acid-binding residue prediction via protein language model feature representation and multi-task learningabstractProtein-nucleic acid interactions play a very important role in a variety of biological activities. Accurate identification of nucleic acid-binding residues is a critical step in understanding the interaction mechanisms. Although many computationally based methods have been developed to predict nucleic acid-binding residues, challenges remain. In this study, a fast and accurate sequence-based method, called ESM-NBR, is proposed. In ESM-NBR, we first use the large protein language model ESM2 to extract discriminative biological properties feature representation from protein primary sequences; then, a multi-task deep learning model composed of stacked bidirectional long short-term memory (BiLSTM) and multi-layer perceptron (MLP) networks is employed to explore common and private information of DNA- and RNA-binding residues with ESM2 feature as input. Experimental results on benchmark data sets demonstrate that the prediction performance of ESM2 feature representation comprehensively outperforms evolutionary information-based hidden Markov model (HMM) features. Meanwhile, the ESM-NBR obtains the MCC values for DNA-binding residues prediction of 0.427 and 0.391 on two independent test sets, which are 18.61 and 10.45% higher than those of the second-best methods, respectively. Moreover, by completely discarding the time-cost multiple sequence alignment process, the prediction speed of ESM-NBR far exceeds that of existing methods (5.52s for a protein sequence of length 500, which is about 16 times faster than the second-fastest method). A user-friendly standalone package and the data of ESM-NBR are freely available for academic use at: https://github.com/wwzll123/ESM-NBR. Wenwu Zeng, Dafeng Lv, Xuan Liu 0001, Guo Chen 0001, Wenjuan Liu, Shaoliang Peng |
BIBM | 6 |
| 2023 | Performance Evaluation of Spark, Ray and MPI: A Case Study on Long Read Alignment Algorithm
Kun Ran, Yingbo Cui 0001, Shaoliang Peng |
ICA3PP (3) | 4 |
| 2023 | One Adapter for All Programming Languages? Adapter Tuning for Code Search and SummarizationabstractAs pre-trained models automate many code intel-ligence tasks, a widely used paradigm is to fine-tune a model on the task dataset for each programming language. A recent study reported that multilingual fine-tuning benefits a range of tasks and models. However, we find that multilingual fine-tuning leads to performance degradation on recent models UniXcoder and CodeT5. To alleviate the potentially catastrophic forgetting issue in multilingual models, we fix all pre-trained model parameters, insert the parameter-efficient structure adapter, and fine-tune it. Updating only 0.6% of the overall parameters compared to full-model fine-tuning for each programming language, adapter tuning yields consistent improvements on code search and sum-marization tasks, achieving state-of-the-art results. In addition, we experimentally show its effectiveness in cross-lingual and low-resource scenarios. Multilingual fine-tuning with 200 samples per programming language approaches the results fine-tuned with the entire dataset on code summarization. Our experiments on three probing tasks show that adapter tuning significantly outperforms full-model fine-tuning and effectively overcomes catastrophic forgetting. Deze Wang, Boxing Chen, Shanshan Li 0001, Shaoliang Peng, Wei Dong 0006, Xiangke Liao |
ICSE | 5 |
| 2023 | Understanding and Detecting On-The-Fly Configuration BugsabstractSoftware systems introduce an increasing number of configuration options to provide flexibility, and support updating the options on the fly to provide persistent services. This mechanism, however, may affect the system reliability, leading to unexpected results like software crashes or functional errors. In this paper, we refer to the bugs caused by on-the-fly configuration updates as on-the-fly configuration bugs, or OCBugs for short. In this paper, we conducted the first in-depth study on 75 real-world OCBugs from 5 widely used systems to understand the symptoms, root causes, and triggering conditions of OCBugs. Based on our study, we designed and implemented Parachute, an automated testing framework to detect OCBugs. Our key insight is that the value of one configuration option, either loaded at the startup phase or updated on the fly, should have the same effects on the target program. Parachute generates tests for on-the-fly configuration updates by mutating the existing tests and conducts differential analysis to identify OCBugs. We evaluated Parachute on 7 real-world software systems. The results show that Parachute detected 75% (42/56) of the known OCBugs, and reported 13 unknown bugs, 11 of which have been confirmed or fixed by developers until the time of writing. Teng Wang 0004, Zhouyang Jia, Shanshan Li 0001, Si Zheng 0003, Yue Yu 0001, Erci Xu, Shaoliang Peng, Xiangke Liao |
ICSE | 7 |
| 2023 | GADRP: graph convolutional networks and autoencoders for cancer drug response predictionabstractDrug response prediction in cancer cell lines is of great significance in personalized medicine. In this study, we propose GADRP, a cancer drug response prediction model based on graph convolutional networks (GCNs) and autoencoders (AEs). We first use a stacked deep AE to extract low-dimensional representations from cell line features, and then construct a sparse drug cell line pair (DCP) network incorporating drug, cell line, and DCP similarity information. Later, initial residual and layer attention-based GCN (ILGCN) that can alleviate over-smoothing problem is utilized to learn DCP features. And finally, fully connected network is employed to make prediction. Benchmarking results demonstrate that GADRP can significantly improve prediction performance on all metrics compared with baselines on five datasets. Particularly, experiments of predictions of unknown DCP responses, drug-cancer tissue associations, and drug-pathway associations illustrate the predictive power of GADRP. All results highlight the effectiveness of GADRP in predicting drug responses, and its potential value in guiding anti-cancer drug selection. Chong Dai, Yuqi Wen, Wenjuan Liu, Xiaochen Bo, Shaoliang Peng |
Briefings Bioinform. | 8 |
| 2023 | A peer-to-peer file storage and sharing system based on consortium blockchainabstractIn the era of big data, data is playing an increasingly important role in scientific study, and reliable storage and secure sharing of data have become a research hotspot. At present, centralized solutions based on data centers and cloud storage have problems with data-right confirmation and center trust. A large number of decentralized storage solutions are public systems, in which blockchain technology, as a tool for value exchange, does not solve the problems of data verification and system supervision. We propose a peer-to-peer storage system with identity access, which achieves data validation, cross-organizational data retrieval, trusted authorization, and sharing based on the consortium blockchain. Our solution proposes a peer-to-peer data storage scheme based on the consortium blockchain and a set of identity authentication mechanisms compatible with the consortium blockchain. Based on this, we propose a blockchain-based permission control scheme and a set of retrieval, authorization, and sharing processes. Finally, we implemented and tested the system to prove the feasibility of the scheme. Shaoliang Peng, Jiandong Shang |
Future Gener. Comput. Syst. | 1 |
| 2022 | MGTUNet: An new UNet for colon nuclei instance segmentation and quantificationabstractColorectal cancer (CRC) is among the top three malignant tumor types in terms of morbidity and mortality. Histopathological images are the gold standard for diagnosing colon cancer. Cellular nuclei instance segmentation and classification, and nuclear component regression tasks can aid in the analysis of the tumor microenvironment in colon tissue. Traditional methods are still unable to handle both types of tasks end-to-end at the same time, and have poor prediction accuracy and high application costs. This paper proposes a new UNet model for handling nuclei based on the UNet framework, called MGTUNet, which uses Mish, Group normalization and transposed convolution layer to improve the segmentation model, and a ranger optimizer to adjust the SmoothL1Loss values. Secondly, it uses different channels to segment and classify different types of nucleus, ultimately completing the nuclei instance segmentation and classification task, and the nuclei component regression task simultaneously. Finally, we did extensive comparison experiments using eight segmentation models. By comparing the three evaluation metrics and the parameter sizes of the models, MGTUNet obtained 0.6254 on PQ, 0.6359 on mPQ, and 0.8695 on R2. Thus, the experiments demonstrated that MGTUNet is now a state-of-the-art method for quantifying histopathological images of colon cancer. Liangrui Pan, Zhichao Feng, Zhujun Xu, Shaoliang Peng |
BIBM | 6 |
| 2022 | deepDGA: Biomedical Heterogeneous Network-based Deep Learning Framework for Disease-Gene Association PredictionsabstractAccurate prediction of disease-gene associations is a crucial tissue in the treatment of diseases. Currently, deep learning-based methods have been proposed to determine the associations between diseases and genes. However, previous network-based models do not consider the semantic characteristics of various biomedical entities and suffer from the problems of cold-start. To this end, this study proposes a heterogeneous network-based deep learning framework (termed deepDGA) to predict disease-gene associations. First, a heterogeneous network with four kinds of biological nodes and eight kinds of edges is constructed. Second, we develop a meta path-driven deep Transformer encoder to learn node representations which contains semantic characteristics of nodes in the heterogeneous network. Finally, the inductive matrix completion algorithm that can solve problem of cold-start, is used for disease-gene association prediction. The results of 5-flod cross-validation and top-ranked predictions suggest that deepDGA is superior to other methods. In addition, we further observe that deepDGA performs the highest predictive ability for specific diseases via the literature verification, KEGG human pathway analyses, and GO enrichment analyses. In summary, deepDGA is an effective framework for predicting the diseases-gene associations. Wenjuan Liu, Shaoliang Peng |
BIBM | 5 |
| 2022 | MSVF: Multi-task Structure Variation Filter with Transfer Learning in High-throughput SequencingabstractThe single molecule real-time sequencing technologies, such as PacBio and Nanopore, have higher throughput and produce longer reads, which promote the discovery of more structure variations that cannot be discovered by the second-generation sequencing data. However, compared with the second-generation sequencing data, the PacBio data lacks paired-end sequencing information, making traditional structure variations filter fail to process the new data. To solve this problem, this paper proposes a universal multi-tasking structure variation filtering model MSVF. MSVF adopts the CIGAR string defined in SAM format. CIGAR is not limited by sequencing technology or alignment algorithms, so MSVF is suitable for not only the second-generation but also the third-generation sequencing data. Moreover, CIGAR string preserves the complete sequence alignment information, which makes MSVF a highly precise model. Besides, MSVF uses deep learning methods, making it supports more structure variation types, including deletion and insertion. We trained and tested the models on the open-access NCBI datasets. The experiments proved that ShuffleNet, MobileNet, ResNet transfer learning models achieve better classification results on SVs task. The average AUC reaches more than 90% and the AUC of each category reach more than 87%. The accuracy and AUC of deletion and insertion structure variations were above 90% and above 92%, respectively. The code and data can be obtained at https://github.con weimingxiang/MSVF. Weiming Xiang 0003, Yingbo Cui 0001, Yaning Yang, Shaoliang Peng |
BIBM | 6 |
| 2022 | ECGNN: Enhancing Abnormal Recognition in 12-Lead ECG with Graph Neural NetworkabstractThe 12-lead Electrocardiography (ECG) is one of the most commonly used diagnostic tools for cardiovascular disease. Widely available ECG databases and deep learning algorithms present an opportunity to substantially improve the accuracy and scalability of automated ECG abnormal identification. However, existing methods mainly model leads individually and then aggregate them for prediction, ignoring the relationship between leads, which is an important diagnostic reference for clinicians. In this paper, we propose a novel model, called ECGNN, which main consists a feature extractor backbone and a graph neural network module. The feature extractor backbone is a neural network used to extract features of ECG signal for subsequent prediction and initialization of the graph nodes. Specifically, the proposed graph neural network module combines graph convolution and graph pooling into a unified module to generate hierarchical representations of graphs and can be integrated into various feature extractor backbones. Experimental results on two largescale 12-lead ECG databases demonstrate the effectiveness of our proposed model. Xiongjun Zhao, Shaoliang Peng |
BIBM | 4 |
| 2022 | UniMed: Multimodal Multitask Learning for Medical PredictionsabstractRecently, deep learning techniques based on electronic health record (EHR) data have achieved success in medical prediction. However, due to the complexity, heterogeneity nature of EHR data, most previous studies build models based on single-modal data (e.g. the structured data or the unstructured free-text data). Although some studies have trained the models based on multimodal EHR data and achieved more advanced performance, they still suffer from the clinical practicability problems, as they require separate modeling for each medical prediction task. Moreover, they ignore the potential correlation between clinical prediction tasks. In this work, we propose UniMed, a Unified model handles multiple Medical prediction tasks simultaneously by learning from multimodal EHR data. Our UniMed model encodes each input modality separately and uses a transformer decoder followed by task-specific prediction heads to predict each medical task. Experimental results conducted on publicly available EHR dataset demonstrate that there is a time-progressive correlation between medical prediction tasks and show the effectiveness of our method. Xiongjun Zhao, Fenglei Yu, Jiandong Shang, Shaoliang Peng |
BIBM | 5 |
| 2022 | A Multimodal Data Fusion-Based Deep Learning Approach for Drug-Drug Interaction Prediction
An Huang 0004, Shaoliang Peng |
ISBRA | 4 |
| 2022 | MPCDDI: A Secure Multiparty Computation-Based Deep Learning Framework for Drug-Drug Interaction Predictions
Shengyun Liu, Shaoliang Peng |
ISBRA | 4 |
| 2022 | EMRShareChain: A Privacy-Preserving EMR Sharing System Model Based on the Consortium Blockchain
Peng Xi, Wenjuan Liu, Shaoliang Peng |
ISBRA | 4 |
| 2022 | Channel pruning guided by global channel relation
Yingjie Cheng, Shaoliang Peng |
Appl. Intell. | 5 |
| 2022 | SVPath: an accurate pipeline for predicting the pathogenicity of human exon structural variantsabstractAlthough there are a large number of structural variations in the chromosomes of each individual, there is a lack of more accurate methods for identifying clinical pathogenic variants. Here, we proposed SVPath, a machine learning-based method to predict the pathogenicity of deletions, insertions and duplications structural variations that occur in exons. We constructed three types of annotation features for each structural variation event in the ClinVar database. First, we treated complex structural variations as multiple consecutive single nucleotide polymorphisms events, and annotated them with correlation scores based on single nucleic acid substitutions, such as the impact on protein function. Second, we determined which genes the variation occurred in, and constructed gene-based annotation features for each structural variation. Third, we also calculated related features based on the transcriptome, such as histone signal, the overlap ratio of variation and genomic element definitions, etc. Finally, we employed a gradient boosting decision tree machine learning method, and used the deletions, insertions and duplications in the ClinVar database to train a structural variation pathogenicity prediction model SVPath. These structural variations are clearly indicated as pathogenic or benign. Experimental results show that our SVPath has achieved excellent predictive performance and outperforms existing state-of-the-art tools. SVPath is very promising in evaluating the clinical pathogenicity of structural variants. SVPath can be used in clinical research to predict the clinical significance of unknown pathogenicity and new structural variation, so as to explore the relationship between diseases and structural variations in a computational way. Yaning Yang, Deshan Zhou, Shaoliang Peng |
Briefings Bioinform. | 5 |
| 2022 | D3AI-CoV: a deep learning platform for predicting drug targets and for virtual screening against COVID-19abstractTarget prediction and virtual screening are two powerful tools of computer-aided drug design. Target identification is of great significance for hit discovery, lead optimization, drug repurposing and elucidation of the mechanism. Virtual screening can improve the hit rate of drug screening to shorten the cycle of drug discovery and development. Therefore, target prediction and virtual screening are of great importance for developing highly effective drugs against COVID-19. Here we present D3AI-CoV, a platform for target prediction and virtual screening for the discovery of anti-COVID-19 drugs. The platform is composed of three newly developed deep learning-based models i.e., MultiDTI, MPNNs-CNN and MPNNs-CNN-R models. To compare the predictive performance of D3AI-CoV with other methods, an external test set, named Test-78, was prepared, which consists of 39 newly published independent active compounds and 39 inactive compounds from DrugBank. For target prediction, the areas under the receiver operating characteristic curves (AUCs) of MultiDTI and MPNNs-CNN models are 0.93 and 0.91, respectively, whereas the AUCs of the other reported approaches range from 0.51 to 0.74. For virtual screening, the hit rate of D3AI-CoV is also better than other methods. D3AI-CoV is available for free as a web application at http://www.d3pharma.com/D3Targets-2019-nCoV/D3AI-CoV/index.php, which can serve as a rapid online tool for predicting potential targets for active compounds and for identifying active molecules against a specific target protein for COVID-19 treatment. Yanqing Yang, Deshan Zhou, Xinben Zhang, Yulong Shi, Jiaxin Han, Leyun Wu, Minfei Ma, Jintian Li, Shaoliang Peng, Weiliang Zhu |
Briefings Bioinform. | 10 |
| 2022 | Clustering-Evolutionary Random Support Vector Machine Ensemble for fMRI-Based Asperger Syndrome DiagnosisabstractAbstract It is a hot spot in the field of computer application to diagnose complex brain diseases such as Asperger syndrome (AS) using machine learning technology. To identify AS patients and detect lesions, this paper proposes a novel clustering-evolutionary random support vector machine (SVM) ensemble (CERSVME) based on graph theory. Firstly, we extract graph theory indexes from the resting-state functional magnetic resonance imaging (fMRI) data as sample features and construct an ensemble learner by integrating multiple SVM classifiers. Secondly, the base learners with high redundancy and poor classification ability are deleted through clustering evolutions to improve the performance of the model. Then the CERSVME model is used to classify fMRI image of AS patients and healthy controls. According to the classification results, a multi-stage analysis scheme is designed to find the AS-related brain areas. We validate the proposed approach on 135 participants from autism brain imaging data exchange cohort. The highest accuracy reported by the CERSVME reaches 95.24%. More importantly, the diseased brain areas such as middle frontal gyrus, hippocampus and precuneus are found based on their contributions to classification performances of the CERSVME. Our study provides useful assistances for the clinical detection of patients with AS. Xia-an Bi, Hao Wu 0143, Shaoliang Peng |
Comput. J. | 5 |
| 2021 | DFL-PiDA: Prediction of Piwi-interacting RNA-Disease Associations based on Deep Feature LearningabstractPiwi-interacting RNAs (piRNAs) fulfill the necessary requirements of epigenetic mechanisms, working to regulate gene expression in diseases and homeostasis in a coordinated manner. Hence, predicting new piRNAs that are associated with diseases conduces to understanding the pathogenicity mechanisms. In this study, we presented a deep feature learning model (DFLPiDA) to predict potential piRNA-disease associations based on the multi-model similarity features of piRNAs and diseases and the convolutional denoising auto-encoder. In particular, we firstly calculated four types of similarity features of piRNAs and diseases. Then, the convolutional denoising auto-encoder was utilized to perform deep learning on the fused similarity features. Finally, the extreme learning machine was employed as the training model as well as to predict unknown associations. The empirical results of five-fold cross-validation experiments show that the DFL-PiDA is efficient for predicting potential piRNA-disease associations. Furthermore, we proved the effectiveness of convolutional denoising auto-encoder neural network in piRNA and disease association prediction. Case studies also demonstrate the practical application of DFL-PiDA to discover potential associations. Jiawei Luo 0001, Liangrui Pan, Shaoliang Peng |
BIBM | 5 |
| 2021 | LADstackING: Stacking Ensemble Learning-based Computational Model for Predicting Potential LncRNA-disease AssociationsabstractIn recent years, accumulation of researches have proved many diseases that seriously endanger human health originate from mutations or dysfunctions in LncRNA (Long non-encoding RNA). Therefore, it is important to discover the intrinsic associations between the LncRNAs and the diseases. Meantime, accurately identifying potential associations between diseases and LncRNAs remains a highly challenging task. In this paper, we proposed a model based on the stacking ensemble learning framework called LADstackING to predict the potential LncRNA associated disease. LADstackING effectively integrates different types of strong predictive performance models rather than the same type models with weak predictive performance.LADstackING is able to exploit the respective advantages of different base models in its framework and significantly improve the overall predictive performance. The multi-perspective features bring by different base models allow LADstackING remain stable in facing of sparse data sources. Moreover, the overall predictive performance of LADstackING is greatly improved compare to the stat-of-art models. Experimental results and case study result demonstrate that LADstackING performs promising in predicting the potential LncRNA-disease associations. Jiechen Li, Xiangxiang Zeng, Yong Dou, Fei Xia 0003, Shaoliang Peng |
BIBM | 5 |
| 2021 | FEDI: Few-shot learning based on Earth Mover's Distance algorithm combined with deep residual network to identify diabetic retinopathyabstractDiabetic retinopathy(DR) is the main cause of blindness in diabetic patients. However, DR can easily delay the occurrence of blindness through the diagnosis of the fundus. In view of the reality, it is difficult to collect a large amount of diabetic retina data in clinical practice. This paper proposes a few-shot learning model of a deep residual network based on Earth Mover's Distance algorithm to assist in diagnosing DR. We build training and validation classification tasks for few-shot learning based on 39 categories of 1000 sample data, train deep residual networks, and obtain experience maximization pre-training models. Based on the weights of the pre-trained model, the Earth Mover's Distance algorithm calculates the distance between the images, obtains the similarity between the images, and changes the model's parameters to improve the accuracy of the training model. Finally, the experimental construction of the small sample classification task of the test set to optimize the model further, and finally, an accuracy of 93.5667% on the 3wayl0shot task of the diabetic retina test set. For the experimental code and results, please refer to: https://github.com/panliangrui/few-shot-learning-funds. Liangrui Pan, Peng Zhang 0035, Fei Xia 0003, Wenjuan Liu, Hetian Wang, Mitchai Chongcheawchamnan, Shaoliang Peng |
BIBM | 9 |
| 2021 | H-VAE: A Hybrid Variational AutoEncoder with Data Augmentation in Predicting CRISPR/Cas9 Off-targetabstractCRISPR/Cas9-based gene editing technology has been widely used in various cells and organisms. However, the off-target effects will bring unpredictable consequences to the organism edited. One of the main obstacles to predict CRISPR/Cas9 off-target is the imbalance of the number of positive and negative samples, which puts forward a challenge for the training of traditional deep learning algorithms. In this paper, we proposed H-VAE, a hybrid variational autoencoder model with data augmentation. This model can extract more abundant sgRNA-DNA base pair matching information, and reduce the risk of overfitting. Moreover, the sample imbalance is resolved. H-VAE can make use of underlying information of training sample, extracted by VAE, to alleviate data-imbalance problem. In view of the weak ability to extract base pair matching information of existing models, a different encoding scheme based on pair encoding is proposed, which enables the model to make full use of sgRNA-DNA base pair matching information. On the Mismatch data set, compared with DeepCRISPR, the ROC-AUC and PR-AUC increased by 0.6% and 41.9%, respectively. In the new Indels data set test scenario, compared with CRISPR-Net, the ROC-AUC and PR-AUC were increased by 1.5% and 133.4% respectively. This proves that H-VAE can improve off-target prediction in various scenarios. The improvement of PR-AUC shows that H-VAE can significantly improve the effect of unbalanced classification. The experimental results demonstrate that H-VAE could achieve a better effect compared with state-of-the-art CRISPR/Cas9 off-target methods on various types of data sets. The code and data can be obtained at https://github.com/weimingxiang/H-VAE. Weiming Xiang 0003, Dong Chen 0013, Yingbo Cui 0001, Shaoliang Peng |
BIBM | 4 |
| 2021 | ParaPindel: a scalable coordinated parallel detection framework for human genome-wide structural variationabstractDetecting the existence of variation from massive human genome data, and determining the breakpoints and types of variations is essential for analyzing structural variation. Pindel, an accurate detection tool based on pattern growth approach, is commonly used for discovering indels and other types of structural variations from next-generation sequencing data. The explosive growth of sequencing data poses new challenges to current implementation of Pindel. Here, we proposed ParaPindel, an optimized version of Pindel that utilizes distributed multiprocess, for efficient large-scale detection of structural variation for human whole-genome sequencing data. ParaPindel divides the chromosome into multiple small windows with a fixed-length window size, so as to realize the parallel detection between different windows and different chromosomes. A crosswindow with a smaller length is introduced to cope with possible structural variations at the edge of the window. The experimental results show that ParaPindel shortens the time to detect an individual’s genome-wide structural variation from 186 hours to 33 minutes under the premise that the detection results are basically consistent. Employing 256 processes on 128 nodes on the TH-IHN supercomputer, the speedup ratio has reached 163 times, and the parallel efficiency has reached 69.74%. Yaning Yang, Chao Yang 0015, Bin Jiang 0006, Shaoliang Peng |
BIBM | 6 |
| 2021 | A Knowledge-aware Machine Reading Comprehension Framework for Dialogue Symptom DiagnosisabstractSymptom diagnosis in dialogue remains a challenging task because the symptom entities and their status need to be extracted correctly at the same time. Most previous studies treat symptom diagnosis as a classification or sequence labeling task and focus on using single-sentence dialogue as input. Unique from past studies, in this paper, we propose a new framework for dialogue symptom diagnosis, which formulate it as a machine reading comprehension (MRC) task. We first use window-level multi-turn of dialogue as input and extract the symptom entities. Then, we generate a question for each entity to infer the symptom status in the form of question answering (QA). Benefit from the MRC formalization, our proposed framework can encode more informative prior knowledge, which can effectively improve the performance of symptom status inference. Experiments on the Chinese medical dialogue dataset show that the proposed framework outperforms the previous best model and several competitive baselines, which indicates that our framework provides a useful direction for dialogue symptom diagnosis. The code and data are publicly available at https://github.com/zhaoxiongjun/DSD. Xiongjun Zhao, Yingjie Cheng, Weiming Xiang 0003, Jiandong Shang, Shaoliang Peng |
BIBM | 7 |
| 2021 | IDOS: Improved D3DOCK on Spark
Yonghui Cui, Shaoliang Peng |
ISBRA | 3 |
| 2021 | MIFS: A Peer-to-Peer Medical Images Storage and Sharing System Based on Consortium Blockchain
Kenli Li 0001, Shaoliang Peng |
ISBRA | 5 |
| 2021 | CFCN: A Multi-scale Fully Convolutional Network with Dilated Convolution for Nuclei Classification and Localization
Yaning Yang, Shaoliang Peng |
ISBRA | 4 |
| 2021 | MDA-GCNFTG: identifying miRNA-disease associations based on graph convolutional networks via graph sampling through the feature and topology graphabstractAccurate identification of the miRNA-disease associations (MDAs) helps to understand the etiology and mechanisms of various diseases. However, the experimental methods are costly and time-consuming. Thus, it is urgent to develop computational methods towards the prediction of MDAs. Based on the graph theory, the MDA prediction is regarded as a node classification task in the present study. To solve this task, we propose a novel method MDA-GCNFTG, which predicts MDAs based on Graph Convolutional Networks (GCNs) via graph sampling through the Feature and Topology Graph to improve the training efficiency and accuracy. This method models both the potential connections of feature space and the structural relationships of MDA data. The nodes of the graphs are represented by the disease semantic similarity, miRNA functional similarity and Gaussian interaction profile kernel similarity. Moreover, we considered six tasks simultaneously on the MDA prediction problem at the first time, which ensure that under both balanced and unbalanced sample distribution, MDA-GCNFTG can predict not only new MDAs but also new diseases without known related miRNAs and new miRNAs without known related diseases. The results of 5-fold cross-validation show that the MDA-GCNFTG method has achieved satisfactory performance on all six tasks and is significantly superior to the classic machine learning methods and the state-of-the-art MDA prediction methods. Moreover, the effectiveness of GCNs via the graph sampling strategy and the feature and topology graph in MDA-GCNFTG has also been demonstrated. More importantly, case studies for two diseases and three miRNAs are conducted and achieved satisfactory performance. Yanyi Chu, Xuhong Wang, Qiuying Dai, Yanjing Wang 0003, Shaoliang Peng, Xiaoyong Wei, Jingfei Qiu, Dennis R. Salahub, Yi Xiong 0002 |
Briefings Bioinform. | 6 |
| 2021 | DeepR2cov: deep representation learning on heterogeneous drug networks to discover anti-inflammatory agents for COVID-19abstractRecent studies have demonstrated that the excessive inflammatory response is an important factor of death in coronavirus disease 2019 (COVID-19) patients. In this study, we propose a deep representation on heterogeneous drug networks, termed DeepR2cov, to discover potential agents for treating the excessive inflammatory response in COVID-19 patients. This work explores the multi-hub characteristic of a heterogeneous drug network integrating eight unique networks. Inspired by the multi-hub characteristic, we design 3 billion special meta paths to train a deep representation model for learning low-dimensional vectors that integrate long-range structure dependency and complex semantic relation among network nodes. Based on the representation vectors and transcriptomics data, we predict 22 drugs that bind to tumor necrosis factor-α or interleukin-6, whose therapeutic associations with the inflammation storm in COVID-19 patients, and molecular binding model are further validated via data from PubMed publications, ongoing clinical trials and a docking program. In addition, the results on five biomedical applications suggest that DeepR2cov significantly outperforms five existing representation approaches. In summary, DeepR2cov is a powerful network representation approach and holds the potential to accelerate treatment of the inflammatory responses in COVID-19 patients. The source code and data can be downloaded from https://github.com/pengsl-lab/DeepR2cov.git. Weihong Tan, Kenli Li 0001, Fei Li 0040, Wu Zhong, Shaoliang Peng |
Briefings Bioinform. | 8 |
| 2021 | BioERP: biomedical heterogeneous network-based self-supervised representation learning approach for entity relationship predictionsabstractMOTIVATION: Predicting entity relationship can greatly benefit important biomedical problems. Recently, a large amount of biomedical heterogeneous networks (BioHNs) are generated and offer opportunities for developing network-based learning approaches to predict relationships among entities. However, current researches slightly explored BioHNs-based self-supervised representation learning methods, and are hard to simultaneously capturing local- and global-level association information among entities. RESULTS: In this study, we propose a BioHN-based self-supervised representation learning approach for entity relationship predictions, termed BioERP. A self-supervised meta path detection mechanism is proposed to train a deep Transformer encoder model that can capture the global structure and semantic feature in BioHNs. Meanwhile, a biomedical entity mask learning strategy is designed to reflect local associations of vertices. Finally, the representations from different task models are concatenated to generate two-level representation vectors for predicting relationships among entities. The results on eight datasets show BioERP outperforms 30 state-of-the-art methods. In particular, BioERP reveals great performance with results close to 1 in terms of AUC and AUPR on the drug-target interaction predictions. In summary, BioERP is a promising bio-entity relationship prediction approach. AVAILABILITY AND IMPLEMENTATION: Source code and data can be downloaded from https://github.com/pengsl-lab/BioERP.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yaning Yang, Kenli Li 0001, Fei Li 0040, Shaoliang Peng |
Bioinform. | 6 |
| 2021 | MultiDTI: drug-target interaction prediction based on multi-modal representation learning to bridge the gap between new chemical entities and known heterogeneous networkabstractMOTIVATION: Predicting new drug-target interactions is an important step in new drug development, understanding of its side effects and drug repositioning. Heterogeneous data sources can provide comprehensive information and different perspectives for drug-target interaction prediction. Thus, there have been many calculation methods relying on heterogeneous networks. Most of them use graph-related algorithms to characterize nodes in heterogeneous networks for predicting new drug-target interactions (DTI). However, these methods can only make predictions in known heterogeneous network datasets, and cannot support the prediction of new chemical entities outside the heterogeneous network, which hinder further drug discovery and development. RESULTS: To solve this problem, we proposed a multi-modal DTI prediction model named 'MultiDTI' which uses our proposed joint learning framework based on heterogeneous networks. It combines the interaction or association information of the heterogeneous network and the drug/target sequence information, and maps the drugs, targets, side effects and disease nodes in the heterogeneous network into a common space. In this way, 'MultiDTI' can map the new chemical entity to this learned common space based on the chemical structure of the new entity. That is, bridging the gap between new chemical entities and known heterogeneous network. Our model has strong predictive performance, and the area under the receiver operating characteristic curve of the model is 0.961 and the area under the precision recall curve is 0.947 with 10-fold cross validation. In addition, some predicted new DTIs have been confirmed by ChEMBL database. Our results indicate that 'MultiDTI' is a powerful and practical tool for predicting new DTI, which can promote the development of drug discovery or drug repositioning. AVAILABILITY AND IMPLEMENTATION: Python codes and dataset are available at https://github.com/Deshan-Zhou/MultiDTI/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Deshan Zhou, Shaoliang Peng |
Bioinform. | 5 |
| 2021 | VISPR-online: a web-based interactive tool to visualize CRISPR screening experimentsabstractBACKGROUND: VISPR is an interactive visualization and analysis framework for CRISPR screening experiments. However, it only supports the output of MAGeCK, and requires installation and manual configuration. Furthermore, VISPR is designed to run on a single computer, and data sharing between collaborators is challenging. RESULTS: To make the tool easily accessible to the community, we present VISPR-online, a web-based general application allowing users to visualize, explore, and share CRISPR screening data online with a few simple steps. VISPR-online provides an exploration of screening results and visualization of read count changes. Apart from MAGeCK, VISPR-online supports two more popular CRISPR screening analysis tools: BAGEL and JACKS. It provides an interactive environment for exploring gene essentiality, viewing guide RNA (gRNA) locations, and allowing users to resume and share screening results. CONCLUSIONS: VISPR-online allows users to visualize, explore and share CRISPR screening data online. It is freely available at http://vispr-online.weililab.org , while the source code is available at https://github.com/lemoncyb/VISPR-online . Yingbo Cui 0001, Johannes Köster, Xiangke Liao, Shaoliang Peng, Tao Tang 0001, Chun Huang 0006, Canqun Yang |
BMC Bioinform. | 5 |
| 2021 | D2D-Enabled Mobile-Edge Computation Offloading for Multiuser IoT NetworkabstractThe emerging mobile-edge computing paradigm provides opportunities for the resource-hungry mobile devices (MDs) to migrate computation. In order to satisfy the requirements of MDs in terms of latency and energy consumption, recent researches proposed diverse computation offloading schemes. However, they either fail to consider the potential computing resources at the edge, or ignore the selfish behavior of users and the dynamic resource adaptability. To this end, we study the computation offloading problem and take into consideration the dynamic available resource of idle devices and the selfish behavior of users. Furthermore, we propose a game theoretic offloading method by regarding the computation offloading process as a resource contention game, which minimizes the individual task execution cost and the system overhead. Utilizing the potential game, we prove the existence of Nash equilibrium (NE), and give a lightweight algorithm to help the game reach a NE, wherein each user can find an optimal offloading strategy based on three contention principles. Additionally, we conduct analysis of computational complexity and the Price of Anarchy (PoA), and deploy three baseline methods to compare with our proposed scheme. Numerical results illustrate that our scheme can provide high-quality services to users, and also demonstrate the effectiveness, scalability and dynamic resource adaptability of our proposed algorithm in a multiuser network. Chengnian Long, Jing Wu 0006, Shaoliang Peng, Bo Li 0001 |
IEEE Internet Things J. | 4 |
| 2021 | Discriminant Projection Shared Dictionary Learning for Classification of Tumors Using Gene Expression DataabstractWith a variety of tumor subtypes, personalized treatments need to identify the subtype of a tumor as accurately as possible. The development of DNA microarrays provides an opportunity to predict tumor classification. One strategy is to use gene expression profiling to extend current biological insights into the disease. However, overfitting problems exist in most machine learning methods when classifying tumor gene expression profile data characterized by high dimensional, small samples and nonlinearities. As a new machine learning methods, dictionary learning has become a more effective algorithm for gene expression profile classification. Here, a new method called discriminant projection shared dictionary learning (DPSDL) is proposed for classifying tumor subtypes using LINCS gene expression profile data. The method trains a shared dictionary, embeds Fisher discriminant criteria to obtain a class-specific sub-dictionary and coding coefficients. At the same time, a projection matrix is trained to widen the distance between different classes of samples. Experimental results show that our method performs better classification based on gene expression profile than the other dictionary learning methods and machine learning methods. Shaoliang Peng, Yaning Yang, Fei Li 0040, Xiangke Liao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | Guest Editorial for the 17th Asia Pacific Bioinformatics ConferenceabstractThe eight papers in this special section were presented at the 17th Asia Pacific Bioinformatics Conference (APBC), which was held in Wuhan, China, 14-16 January 2019. Louxin Zhang, Shaoliang Peng, Yi-Ping Phoebe Chen, David Sankoff, Guoliang Li 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | LUNAR : Drug Screening for Novel Coronavirus Based on Representation Learning Graph Convolutional NetworkabstractAn outbreak of COVID-19 that began in late 2019 was caused by a novel coronavirus(SARS-CoV-2). It has become a global pandemic. As of June 9, 2020, it has infected nearly 7 million people and killed more than 400,000, but there is no specific drug. Therefore, there is an urgent need to find or develop more drugs to suppress the virus. Here, we propose a new nonlinear end-to-end model called LUNAR. It uses graph convolutional neural networks to automatically learn the neighborhood information of complex heterogeneous relational networks and combines the attention mechanism to reflect the importance of the sum of different types of neighborhood information to obtain the representation characteristics of each node. Finally, through the topology reconstruction process, the feature representations of drugs and targets are forcibly extracted to match the observed network as much as possible. Through this reconstruction process, we obtain the strength of the relationship between different nodes and predict drug candidates that may affect the treatment of COVID-19 based on the known targets of COVID-19. These selected candidate drugs can be used as a reference for experimental scientists and accelerate the speed of drug development. LUNAR can well integrate various topological structure information in heterogeneous networks, and skillfully combine attention mechanisms to reflect the importance of neighborhood information of different types of nodes, improving the interpretability of the model. The area under the curve(AUC) of the model is 0.949 and the accurate recall curve (AUPR) is 0.866 using 10-fold cross-validation. These two performance indexes show that the model has superior predictive performance. Besides, some of the drugs screened out by our model have appeared in some clinical studies to further illustrate the effectiveness of the model. Deshan Zhou, Shaoliang Peng, Wu Zhong, Yutao Dou |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2020 | Predicting CRISPR-Cas9 Off-target with Self-supervised Neural NetworksabstractCRISPR-Cas9 is causing a new revolution in many fields s uch a s b asic b iological r esearch, m edicine, a nd biotechnology as the third-generation gene-editing tool. However, the phenomenon of off-target is a stumbling block to the vigorous development of gene-editing technology. In this paper, we proposed DNA-BERT by adding more meaningful tasks that learn regulatory sequence code from genomic sequence and remove useless tasks based on original Bidirectional Encoder Representations from Transformers (BERT) model to make it more suitable for DNA sequence tasks. Due to the lack of training samples, we use it to pre-training from massive genome data and use LightGBM(Light Gradient Boosting Model) to build a classification and regression model using DNA-BERT embeddings combine with hand-crafted features including mismatches, the secondary structure and so on. The empirical results from the public benchmark demonstrate that our method achieves better performance compared with state-of-art off-target methods (i.e. Elevation, DeepCRISPR, CNN-based method, CFD, MIT, CROPIT, CCTop) on benchmark studies. Dong Chen 0013, Wenjie Shu 0001, Shaoliang Peng |
BIBM | 3 |
| 2020 | Predicting Drugs for COVID-19/SARS-CoV-2 via Heterogeneous Graph Attention NetworksabstractCoronavirus Disease-19 (COVID-19) has led to global epidemics with high morbidity and mortality. However, there are currently no proven effective drugs targeting COVID19. Identifying drug-virus associations can not only provide insights into the understanding of drug-virus interaction mechanism, but also guide and facilitate the screening of compound candidates for antiviral drug discovery. In this work, we propose a novel framework of Heterogeneous Graph Attention Networks for Drug-Virus Association predictions, named HGATDVA. First, we fully incorporate multiple sources of biomedical data to construct abundant features for drugs and viruses. Second, we construct two drug-virus heterogeneous graphs. For each graph, we design a self-enhanced graph attention network (SGAT) to explicitly model the dependency between a node and its local neighbors and derive the graph-specific representations for nodes. Third, we further develop a neural network architecture with tri-aggregator to aggregate the graph-specific representations to generate the final node representations. Experiments on two datasets were conducted to demonstrate the effectiveness of our proposed method in identifying candidate drugs for viruses. Yahui Long, Yu Zhang 0084, Min Wu 0008, Shaoliang Peng, Chee Keong Kwoh 0001, Jiawei Luo 0001, Xiaoli Li 0001 |
BIBM | 4 |
| 2020 | A drug information embedding method based on graph convolution neural networkabstractNew drug development is an extremely time-consuming and high-risk process. [1]It has been widely valued by the biomedical industry to fully explore the new uses of existing drugs and reorientate them. [2]How to find drug disease with potential therapeutic relationship from a large number of unproven relationship pairs is the research focus of drug reorientation. With the help of machine learning model, we can improve the enrichment degree of potential drug disease relationship pairs, and reduce the false positive rate of prediction. In the past few years, a series of graph based convolutional network models have been developed to calculate the information latent feature representation of nodes and links. Researchers at home and abroad have done a lot of research on network embedding technology based on biomedical data, and have achieved a series of important research results. Among them, the research methods used can be divided into two categories: one is the traditional machine learning algorithm based on artificial feature extraction, the other is the method based on deep learning. For example, kipf and welling [3]proposed a new graph convolution network (GCN) with parts of existing models, DeepDR [4] and DTINet [5] based on node characteristics and their connections, which can be used for node classification. Aiming at the problem of imbalance of drug information data samples, the invention provides a drug relocation method based on deep learning multi-source heterogeneous network. In order to avoid the limitations of traditional feature extraction methods, such as highly dependent on the experience and knowledge of medical staff, strong subjectivity, consuming a lot of time and energy to complete, and extracting high-quality features with distinguishing features often exists In this paper, with the help of graph convolution encoder model and variational auto encoder neural network, we can automatically learn the characteristics of multi-source and heterogeneous drug low-dimensional network, and complete the drug relocation of drug disease association prediction. Xiaoyi Feng, Shaoliang Peng, Fei Li 0040, Xiangxiang Zeng, Yunhao Liu 0001 |
HealthCom | 2 |
| 2020 | Segmented Encryption: A Quality and Safety Supervisory Model for Herbal Medicine Based on Blockchain TechnologyabstractThe quality of herbal medicine has an important impact on human health. In this paper, we proposed a blockchain-based herbal quality and safety supervisory model for the current frequent herbal counterfeiting phenomenon. We manage the production, processing, and trading processes of herbal medicines by exploiting the blockchain's immutable and traceable properties. We proposed a segmented encryption method for information to encrypt the private information of enterprises. We use shared cloud storage to reduce waste of local storage space, and we proposed a verifiable random chain cutting mechanism based on the verifiable random function to handle the redundant blocks of the chain. Our article addressed the problem of herbal source falsification. Automated recording of key factors such as soil and temperature that affect the quality of herbal medicines is done from the seedling stage. Our herbal quality and safety supervisory model used blockchain technology to increase control over the production and distribution of herbal products, reduce herbal counterfeiting, and improve the efficiency of the system. Jiameng Liu, Shaoliang Peng, Jiawei Luo 0001, Zhuo Tang |
HealthCom | 2 |
| 2020 | Predicting functional elements and variants effects in non-coding regions based on deep learningabstractAccurate recognition and annotation of the important functional elements in the genome is an important prerequisite to understand the coding mode of complex regulatory networks in the one-dimensional genome. Despite rapid advances in sequencing and recognition technologies, accurately calling non-coding variant effects from large-scale sequence reads remains challenging. Here we present a deep neural network-based algorithmic framework, DeepMSA, which directly learns a regulatory sequence code from large-scale chromatin-profiling data,enabling to evaluate chromatin effects caused by SNP(single nucleotide polymorphism). Yunhao Liu 0001, Shaoliang Peng, Wenjie Shu 0001, Bin Jiang 0006, Chao Yang 0015, Kun Xie 0001 |
HealthCom | 2 |
| 2020 | GeoAI-based Epidemic Control with Geo-Social Data Sharing on BlockchainabstractEpidemics especially those caused by major contagious diseases have entailed huge losses in human history. The fights have thus never stopped to prevent pandemics. Due to its acute outbreak, is generally susceptible to the population regardless of ages, so strict quarantine of the infections becomes the most effective means for the epidemic control, which has been proved in the prevention of other contagious diseases such as SARS and H1N1. The key strategy widely used to find infected and suspected patients is still the epidemiological tracking of confirmed cases. However, this may fail to identify infections especially when patients do not show any symptoms. Therefore, the approach to rapid, effective, and simple infection identification is essential to prevent the spread of a contagious disease. This paper proposes to leverage a social apps and Geospatial artificial intelligence (GeoAI) with Blockchain to effectively identify infections with privacy concern. Since people widely use social apps, a large scale of social data with geospatial information could be easily collected and kept on Blockchain with privacy preservation, which thus provides a framework of decentralized, tamper-proof, and privacy-preserved information sharing. With the support of GeoAI, which analyzes the spatial distribution of diseases from the shared data, we could study the influence factors based on spatial propagation of contagious diseases for infection identification. Since WeChat is widely used in China, we take COVID-19 as an example to use the experiments on real-life datasets demonstrate the effectiveness of our method, and provide insight into epidemic control in terms of geo-social data sharing. Shaoliang Peng, Li Xiong 0015, Qiang Qu 0001, Shulin Wang |
HealthCom | 1 |
| 2020 | Efficiently recognition of vaginal micro-ecological environment based on Convolutional Neural NetworkabstractVaginal diseases caused by vaginal micro-ecological abnormalities mainly include Vulvovaginal Candidiasis (VVC), Aerobic Vaginitis (AV), and Bacterial Vaginosis (BV). Severe cases can lead to poor pregnancy outcomes and infertility. AI-based technologies are being deployed with an expectation to relieve doctors of routine, tedious work when implemented correctly in daily microscopy of vaginal micro-ecological abnormalities. In this paper, we built a clinical image dataset of the Gram stain of the vaginal discharge. By comparing the performance of state of art convolutional neural network models, we found the fine-tuning Inception ResNet V2 shows the best classification performance for vaginal diseases. It achieves 96%, 94%, 86% AUC in VVC, AV, BV classification respectively. The result shows that compared with human visual inspection, the method based on deep learning greatly improves the screening sensitivity. Besides, we found that transfer learning can reduce the required manual labeling by roughly 73% (about more than one thousand samples). But for BV, which is difficult to diagnose for both humans and AI. Unlike AV and VVC, it requires more labeled data and is insensitive to the transfer fine-tuning. Shaoliang Peng, Minxia Cheng, Yaning Yang, Fei Li 0040 |
HealthCom | 1 |
| 2020 | Multi-View Weighted Feature Fusion Using CNN for Pneumonia Detection on Chest X-RaysabstractChest X-ray is still the most common and important method for diagnosing pneumonia. However, the analysis of chest radiographs requires professional radiologists, and overreliance on radiologists may lead to erroneous diagnosis or missed diagnosis. Using convolutional neural networks(CNNs) for diagnosis chest diseases on chest X-ray has achieved better results, but most of the previous models are only trained by frontal-view X-ray images. Unique from past studies, in this paper, we proposed a model that can learn multi-view semantic information from chest X-rays to detect pneumonia. Our model includes two stages of feature extraction and feature fusion, and is trained on MIMIC-CXR-JPG dataset, currently the largest publicly available chest x-ray dataset, containing 377,110 JPG format images. We demonstrate that such multi-view weighted feature fusion model outperforms the models that use features only from one view. Our results are better than previous models for pneumonia detection. Shaoliang Peng, Xiongjun Zhao, Xiaoyong Wei, Donqing Wei, Yuehua Peng |
HealthCom | 1 |
| 2020 | A deep metric learning algorithm for similarity measure of the gene expression profileabstractClustering gene expression profiles is a fundamental task in the genome and biomedical research. With the development of RNA-seq and gene chip technology, mass gene expression profile data has been generated, which puts forward two requirements for related research of gene expression profile: i) accurate analysis of drug R&D requires high accuracy of similarity analysis, ii) large-scale analysis of data requires as little running time as possible. We propose a faster, more accurate method called DeepCDNet, which is based on the framework of the Siamese network. DeepCDNet uses the DenseNet structure and optimized loss function to achieve rapid convergence, and the similarity between expression spectra is calculated by a cosine function. The experiment results show that: i) our method breaks through the limitation of high dimensions of gene expression profile and can quickly and accurately learn the required gene characteristics, ii) The accuracy of our method in similarity analysis is greatly improved, iii) as the dimension of data increases, the advantage of our method on time cost gradually becomes more prominent, and time consumption is less. Shaoliang Peng, Yaning Yang, Fei Li 0040, Hao Hong, Kenli Li 0001, Shulin Wang |
HealthCom | 1 |
| 2020 | A Dynamic Protection Mechanism for GPU Memory Overflow
Yaning Yang, Shaoliang Peng |
NPC | 3 |
| 2020 | Re-ranking Answer Selection with Similarity AggregationabstractAnswer selection plays a crucial role in natural language processing. and thus has received much attention. Many recent works treat it as an ad-hoc retrieval problem where ranking optimization accounts for a large proportion. Previous works mainly consider the similarity between answer and question, but rarely utilize similarity and dissimilarity relationship in the answers candidate set. In this paper, we propose a similarity aggregation method to rerank the results produced by different baseline neural networks. The key idea of similarity aggregation is that true matches should not only similar to other true matches, but also dissimilar with false matches, and inspired by multi-view verification, the true answers should have the same ranking to the question in different baseline methods and false answers are the same. The empirical results, from the public benchmark task of answer selection, demonstrate that our method has significant improvement over the baseline methods. Dong Chen 0013, Shaoliang Peng, Kenli Li 0001 |
SIGIR | 2 |
| 2020 | A hybrid two-stage financial stock forecasting algorithm based on clustering and ensemble learning
Cuijuan Yang, Shaoliang Peng, Yusuke Nojima |
Appl. Intell. | 3 |
| 2020 | An efficient framework for generating robust adversarial examplesabstractRecent studies show that deep neural networks (DNNs) suffer adversarial examples. That is, attackers can mislead the output of a DNN by adding subtle perturbation to a benign input image. In addition, researchers propose new generation of technologies to produce robust adversarial examples. Robust adversarial examples can consistently fool DNN models under predefined hyperparameter space, which can break through some defenses against adversarial examples or even generate physical adversarial examples against real-world applications. Behind these achievements, expectation over transformation (EOT) algorithm plays as the backbone framework for generating robust adversarial examples. Though EOT framework is powerful, we know little about why such a framework can generate robust adversarial examples. To address this issue, we do the first work to explain the principle behind robust adversarial examples. Then, based on the findings, we point out that traditional EOT framework has a performance problem and propose an adaptive sampling algorithm to overcome such a problem. By modeling the sampling process as classic Coupon Collector Problem, we prove that our new framework reduces the cost from O ( n ∗ log ( n ) ) to O ( n ), where n denotes the number of sampling points. Under the view of computational complexity, the algorithm is optimal for this problem. The experimental results show that our algorithm can save up to 23% overhead in average. This is significant for black-box attack, where the cost is charged by the amount of queries. Kai Lu 0001, Shaoliang Peng |
Int. J. Intell. Syst. | 4 |
| 2020 | Joint consensus and diversity for multi-view semi-supervised classification
Wenzhang Zhuge, Chenping Hou, Shaoliang Peng, Dongyun Yi |
Mach. Learn. | 3 |
| 2020 | High-Scalable Collaborated Parallel Framework for Large-Scale Molecular Dynamic Simulation on Tianhe-2 SupercomputerabstractMolecular dynamics (MD) is a computer simulation method of studying physical movements of atoms and molecules that provide detailed microscopic sampling on molecular scale. With the continuous efforts and improvements, MD simulation gained popularity in materials science, biochemistry and biophysics with various application areas and expanding data scale. Assisted Model Building with Energy Refinement (AMBER) is one of the most widely used software packages for conducting MD simulations. However, the speed of AMBER MD simulations for system with millions of atoms in microsecond scale still need to be improved. In this paper, we propose a parallel acceleration strategy for AMBER on the Tianhe-2 supercomputer. The parallel optimization of AMBER is carried out on three different levels: fine grained OpenMP parallel on a single CPU, single node CPU/MIC parallel optimization and multi-node multi-MIC collaborated parallel acceleration. By the three levels of parallel acceleration strategy above, we achieved the highest speedup of 25-33 times compared with the original program. Shaoliang Peng, Xiaoyu Zhang 0008, Wenhe Su, Yutong Lu, Xiangke Liao, Kai Lu 0001, Canqun Yang, Jie Liu 0002, Weiliang Zhu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2019 | vGuard: A Spatiotemporal Efficiency Supervision Method For Vaccine Production Based On Double-level BlockchainabstractA vaccine is a biological production that is related to people's lives. Currently, vaccine production supervision is very rough. The vaccine production records are completely controlled by the enterprise. Enterprises only submit production records to review agency for review when they need to sell vaccines. Production records are easy to forge and modify. In order to solve the shortcomings of traditional centralized management. We propose a supervision method for vaccine production based on double-level blockchain. We have designed a double-level blockchain structure. The first level is private data of vaccine production enterprise, including production records and corresponding hash. The next level is public data, including production records hash and vaccine information. In this way, we make vaccine enterprise to submit production records in a timely manner without fear of privacy leaks. We avoid enterprise tampering or falsification of production records through the non-tampering features and time stamps of the blockchain. Through these methods, we have realized efficiency supervision of vaccine production. Shaoliang Peng, Chengnian Long, Hongbo Jiang 0001, Lijun Wei |
BIBM | 2 |
| 2019 | LysoPhD: predicting functional prophages in bacterial genomes from high-throughput sequencingabstractMotivation: When a temperate bacteriophage (phage) integrates into the chromosome of its host or to remain as latent episomal DNA, the prophage plays an important role in bacterial virulence acquisition and increase. With the recent release of thousands of bacterial high throughput sequencing data, there has been a growing interest in analyzing the impact of interdependency between a prophage and its host. One of the essential tasks is to acquire and identify the complete sequence of the prophage, i.e., the functional prophages. However, no existing tools can identify functional prophages from bacterial genomes. A few tools can predict putative prophage in bacterial genomes, but they are unable to distinguish functional prophages from defective prophages. To reduce the cost and to relieve the tedious biological experiments for functional prophage analysis, computational methods for predicting functional prophages based on bacterial high-throughput sequencing data would be an effective alternation. Results: This paper presents a method, LysoPhD, the first tool to predict functional prophage sequences using bacterial HTS data automatically and accurately. The main idea is to obtain whole genome of functional prophages from the bacterial genome based on phage gene annotation and reads-mediated contig connection. LysoPhD obtains whole genome sequences of functional prophages based on improved models. The obtained the sequences are verified by wet lab experiments. Specifically, LysoPhD is applied to a set of staphylococcus HTS data in the case studies. Ten functional prophages are predicted from 72 staphylococcus isolates. The 72 bacterial strains are then induced with Mitomycin C and subsequent deep sequencing verified of the 10 functional prophages. This shows that LysoPhD is a sensitive and accurate program for functional prophage prediction. Qi Niu, Shaoliang Peng, Xianglilan Zhang, Shuaicheng Li 0001, Xiang-Cheng Xie, Yi-Gang Tong |
BIBM | 2 |
| 2019 | A General Fine-tuned Transfer Learning Model for Predicting Clinical Task Acrossing Diverse EHRs DatasetsabstractData analysis of electronic health record (EHRs) system using machine learning, statistical methods can predict relevant clinical tasks. However, there is no uniform standard for current electronic health record systems, and the clinical outcome prediction models trained on one EHR dataset cannot be applied well on other EHR datasets from different medical institutions. Data differences between different medical institutions pose a huge challenge to the study of electronic health records. In this study, we proposed a general transfer learning strategy which can enable models to make clinical prediction acrossing diverse EHRs datasets and validated its strong versatility on three deep learning models. Two different intensive care units (ICU) databases (MIMIC-III and eICU) and one clinical task (in-hospital mortality) are used to evaluate our method. At first, we trained the deep learning models on the source dataset and saved the model states after each epoch. Then, we selected the best performing model as the pre-training model, transferred it to the target dataset and fine-tuned the whole network on target dataset. Finally, we use the fine-tuned models to make predictions on the target dataset. Experiment results show that AUROC score increased by 3%-20% with transfer strategy, which indicated that the general strategy can provide more reliable predictions acrossing EHRs databases to predict clinical tasks. Shaoliang Peng, Yaning Yang, Fei Li 0040 |
BIBM | 2 |
| 2019 | DMCM: a Data-adaptive Mutation Clustering Method to identify cancer-related mutation clustersabstractMotivation: Functional somatic mutations within coding amino acid sequences confer growth advantage in pathogenic process. Most existing methods for identifying cancer-related mutations focus on the single amino acid or the entire gene level. However, gain-of-function mutations often cluster in specific protein regions instead of existing independently in the amino acid sequences. Some approaches for identifying mutation clusters with mutation density on amino acid chain have been proposed recently. But their performance in identification of mutation clusters remains to be improved. Results: Here we present a Data-adaptive Mutation Clustering Method (DMCM), in which kernel density estimate (KDE) with a data-adaptive bandwidth is applied to estimate the mutation density, to find variable clusters with different lengths on amino acid sequences. We apply this approach in the mutation data of 571 genes in over twenty cancer types from The Cancer Genome Atlas (TCGA). We compare the DMCM with M2C, OncodriveCLUST and Pfam Domain and find that DMCM tends to identify more significant clusters. The cross-validation analysis shows DMCM is robust and cluster cancer type enrichment analysis shows that specific cancer types are enriched for specific mutation clusters. Availability and implementation: DMCM is written in Python and analysis methods of DMCM are written in R. They are all released online, available through https://github.com/XinguoLu/DMCM. Supplementary information: Supplementary data are available at Bioinformatics online. Xinguo Lu, Qiumai Miao, Shaoliang Peng |
Bioinform. | 5 |
| 2019 | A CPU/MIC Collaborated Parallel Framework for GROMACS on Tianhe-2 SupercomputerabstractMolecular Dynamics (MD) is the simulation of the dynamic behavior of atoms and molecules. As the most popular software for molecular dynamics, GROMACS cannot work on large-scale data because of limit computing resources. In this paper, we propose a CPU and Intel® Xeon Phi Many Integrated Core (MIC) collaborated parallel framework to accelerate GROMACS using the offload mode on a MIC coprocessor, with which the performance of GROMACS is improved significantly, especially with the utility of Tianhe-2 supercomputer. Furthermore, we optimize GROMACS so that it can run on both the CPU and MIC at the same time. In addition, we accelerate multi-node GROMACS so that it can be used in practice. Benchmarking on real data, our accelerated GROMACS performs very well and reduces computation time significantly. Source code: https://github.com/tianhe2/gromacs-mic. Shaoliang Peng, Yingbo Cui 0001, Shunyun Yang, Wenhe Su, Xiaoyu Zhang 0008, Tenglilang Zhang, Xingming Zhao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2018 | Efficient computation of motif discovery on Intel Many Integrated Core (MIC) ArchitectureabstractBACKGROUND: Novel sequence motifs detection is becoming increasingly essential in computational biology. However, the high computational cost greatly constrains the efficiency of most motif discovery algorithms. RESULTS: In this paper, we accelerate MEME algorithm targeted on Intel Many Integrated Core (MIC) Architecture and present a parallel implementation of MEME called MIC-MEME base on hybrid CPU/MIC computing framework. Our method focuses on parallelizing the starting point searching method and improving iteration updating strategy of the algorithm. MIC-MEME has achieved significant speedups of 26.6 for ZOOPS model and 30.2 for OOPS model on average for the overall runtime when benchmarked on the experimental platform with two Xeon Phi 3120 coprocessors. CONCLUSIONS: Furthermore, MIC-MEME has been compared with state-of-arts methods and it shows good scalability with respect to dataset size and the number of MICs. Source code: https://github.com/hkwkevin28/MIC-MEME . Shaoliang Peng, Minxia Cheng, Yingbo Cui 0001, Runxin Guo, Xiaoyu Zhang 0008, Shunyun Yang, Xiangke Liao, Yutong Lu, Quan Zou 0001, Benyun Shi |
BMC Bioinform. | 1 |
| 2018 | cmFSM: a scalable CPU-MIC coordinated drug-finding tool by frequent subgraph miningabstractBACKGROUND: Frequent subgraphs mining is a significant problem in many practical domains. The solution of this kind of problem can particularly used in some large-scale drug molecular or biological libraries to help us find drugs or core biological structures rapidly and predict toxicity of some unknown compounds. The main challenge is its efficiency, as (i) it is computationally intensive to test for graph isomorphisms, and (ii) the graph collection to be mined and mining results can be very large. Existing solutions often require days to derive mining results from biological networks even with relative low support threshold. Also, the whole mining results always cannot be stored in single node memory. RESULTS: In this paper, we implement a parallel acceleration tool for classical frequent subgraph mining algorithm called cmFSM. The core idea is to employ parallel techniques to parallelize extension tasks, so as to reduce computation time. On the other hand, we employ multi-node strategy to solve the problem of memory constraints. The parallel optimization of cmFSM is carried out on three different levels, including the fine-grained OpenMP parallelization on single node, multi-node multi-process parallel acceleration and CPU-MIC collaborated parallel optimization. CONCLUSIONS: Evaluation results show that cmFSM clearly outperforms the existing state-of-the-art miners even if we only hold a few parallel computing resources. It means that cmFSM provides a practical solution to frequent subgraph mining problem with huge number of mining results. Specifically, our solution is up to one order of magnitude faster than the best CPU-based approach on single node and presents a promising scalability of massive mining tasks in multi-node scenario. More source code are available at:Source Code: https://github.com/ysycloud/cmFSM . Shunyun Yang, Runxin Guo, Xiangke Liao, Quan Zou 0001, Benyun Shi, Shaoliang Peng |
BMC Bioinform. | 7 |
| 2018 | mSNP: A Massively Parallel Algorithm for Large-Scale SNP DetectionabstractSingle Nucleotide Polymorphism (SNP) detection is a fundamental procedure of whole genome analysis. SOAPsnp, a classic tool for detection, would take more than one week to analyze one typical human genome, which limits the efficiency of downstream analyses. In this paper, we present mSNP, an optimized version of SOAPsnp, which leverages Intel Xeon Phi coprocessors for large-scale SNP detection. Firstly, we redesigned the essential data structures of SOAPsnp, which significantly reduces memory footprint and improves computing efficiency. Then we developed a coordinated parallel framework for a higher hardware utilization of both CPU and Xeon Phi. Also, we tailored the data structures and operations to utilize the wide VPU of Xeon Phi to improve data throughput. Last but not the least, we proposed a read-based window division strategy to improve throughput and obtain better load balance. mSNP is the first SNP detection tool empowered by Xeon Phi. We achieved a 38x single thread speedup on CPU, without any loss in precision. Moreover, mSNP successfully scaled to 4,096 nodes on Tianhe-2. Our experiments demonstrate that mSNP is efficient and scalable for large-scale human genome SNP detection. Yingbo Cui 0001, Shaoliang Peng, Yutong Lu, Xiaoqian Zhu, Bingqiang Wang, Chengkun Wu, Xiangke Liao |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2018 | ConfVD: System Reactions Analysis and Evaluation Through Misconfiguration InjectionabstractIn recent years, misconfigurations have become one of the major causes of software system failures, resulting in numerous service outages. What is worse, misconfigurations are also costly to diagnose and troubleshoot. This remains a great challenge for sysadmins (system administrators) to detect, diagnose, or troubleshoot these misconfigurations. Unlike software bugs, misconfigurations are more vulnerable to sysadmins' mistakes. Developers and researchers are attempting to improve system reactions to misconfigurations to ease the burden of sysadmins' diagnoses. Such efforts would greatly benefit from the techniques that can comprehensively detect bad system reactions through injected misconfigurations. Unfortunately, few such studies have achieved the above goal in the past, primarily because they only relied on generic alterations and failed to find a way to systematically generate misconfigurations. In this paper, we study eight mature open-source and commercial software packages and summarize a fine-grained classification of option types. Based on this classification, we use Augmented Backus-Naur Form to summarize and extract syntactic and semantic constraints of each type. In order to generate comprehensive misconfigurations in the test systems, we propose misconfiguration generation methods for our constraints. We implement a tool named Configuration Vulnerability Detector (ConfVD) to conduct misconfiguration injection and further analyze the systems' reaction abilities to various misconfigurations. We carried out comprehensive analyses upon Apache Httpd, MySQL, PostgreSQL, and Yum. The results of our analysis show that our option classification covers 96% of 1582 options from the above-mentioned systems. Our constraints are more fine grained than previous works and their accuracy was found to be 91% (ascertained by manual verification). Our technique could improve generic alteration approaches without constraints, and we found that ConfVD could find nearly three times the bad reactions that were found by ConfErr. In total, we found 65 bad reactions from the systems being tested and our fine-grained constraints contributed 27.7% more bad reactions than techniques only using coarse-grained constraints. Shanshan Li 0001, Wang Li 0003, Xiangke Liao, Shaoliang Peng, Shulin Zhou, Zhouyang Jia, Teng Wang 0004 |
IEEE Trans. Reliab. | 4 |
| 2017 | A novel algorithm for detecting co-evolutionary domains in protein and nucleotide sequencesabstractCo-evolution exists ubiquitously in biological systems. At the molecular level, interacting proteins, such as ligands and their receptors and components in protein complexes, co-evolve to maintain their structural and functional interactions. Many proteins contain multiple functional domains interacting with different partners, making co-evolution of interacting domains occur more prominently. Multiple methods have been developed to predict interacting proteins or domains within proteins by detecting their co-variation. This strategy neglects the fact that interacting domains can be highly co-conserved due to their functional interactions. Here we report a novel algorithm to detect signals of both co-positive selection (co-variation) and co-purifying selection (co-conservation). Preliminary results show that our algorithm performs well and outperforms the popular co-variation analysis program CAPS. Our algorithm can be widely used to predict interacting domains in protein and nucleotide sequences and to analyze protein-ncRNA complexes. Xiaoyu Zhang 0008, Xiangke Liao, Kenli Li 0001, Benyun Shi, Shaoliang Peng |
BIBM | 6 |
| 2017 | mD3DOCKxb: An Ultra-Scalable CPU-MIC Coordinated Virtual Screening FrameworkabstractMolecular docking is an important method in computational drug discovery. In large-scale virtual screening, millions of small drug-like molecules (chemical compounds) are compared against a designated target protein (receptor). Depending on the utilized docking algorithm for screening, this can take several weeks on conventional HPC systems. However, for certain applications including large-scale screening tasks for newly emerging infectious diseases such high runtimes can be highly prohibitive. In this paper, we investigate how the massively parallel neo-heterogeneous architecture of Tianhe-2 Supercomputer consisting of thousands of nodes comprising CPUs and MIC coprocessors that can efficiently be used for virtual screening tasks. Our proposed approach is based on a coordinated parallel framework called mD3DOCKxb in which CPUs collaborate with MICs to achieve high hardware utilization. mD3DOCKxb comprises a novel efficient communication engine for dynamic task scheduling and load balancing between nodes in order to reduce communication and I/O latency. This results in a highly scalable implementation with parallel efficiency of over 84% (strong scaling) when executing on 8,000 Tianhe-2 nodes comprising 192,000 CPU cores and 1,368,000 MIC cores. Shaoliang Peng, Xiaoyu Zhang 0008, Shunyun Yang, Wenhe Su, Kai Lu 0001, Yutong Lu, Xiangke Liao, Bertil Schmidt, Weiliang Zhu, Kuanching Li |
CCGrid | 1 |
| 2016 | High performance computational biology and drug design on TianHe SupercomputersabstractSummary form only given. Extremely powerful computers are needed to help scientists to handle high performance computational biology and drug design problems. The world's largest genomics institute BGI currently generates 6 TB data each day. The European Bioinformatics Institute (EBI) in Hinxton currently stores 20 petabytes (1 petabyte is 1015 bytes) of data and back-ups about genes, proteins and small molecules. TianHe supercomputers can speed up computational biology and drug design processing. In 2013, 2014, and 2015, Tianhe-2 topped the TOP500 list of fastest supercomputers in the world. Many well-known bioinformatics and drug design softwares (BWA, DOCK, SOAP3-dp, SOAPdenovo, SOAPsnp etc.) are developed and running on TH-2. The talk focuses on two main areas: 1. Drug Design: mD3DOCKxb is a largest high throughput molecular docking platform and finishes the docking of all the purchasable molecules (about 42 million) on earth within 24 hours.It has a parallel efficiency of over 70% using 192,000 CPU cores and 1,368,000 MIC cores. It gains the Gold Award of PAC 2015 (Parallel Application Challenge Competition) and is reported by CCTV 1, ScienceNet, China Science and Technology News, and 2015 Top 10 News of Hunan Province of China. 2. Genetic Engineering: The “Human Whole Genome Re-sequencing Analysis Software Pipeline” is firstly designed by applicant. The whole analyzing procedure takes 4 hours to finish the analysis of a 300 TB dataset of whole genome sequences from 2,000 human beings. The speedup is about 1200X. TianHe Supercomputers can handle 3 kinds of computational biology and drug design problems: computation intensive, memory intensive, and communication intensive. In future, TH-2 will be open online to all the scientists not only in China but also all over the world. Shaoliang Peng |
BIBM | 1 |
| 2016 | mAMBER: A CPU/MIC collaborated parallel framework for AMBER on Tianhe-2 supercomputerabstractMolecular dynamics (MD) is a computer simulation method of studying physical movements of atoms and molecules that provide detailed microscopic sampling on molecular scale. With the continuous efforts and improvements, MD simulation gained popularity in materials science, biochemistry and biophysics with various application areas and expanding data scale. Assisted Model Building with Energy Refinement (AMBER) is one of the most widely used software packages for conducting MD simulations. However, the speed of AMBER MD simulations for system with millions of atoms in microsecond scale still need to be improved. In this paper, we propose a parallel acceleration strategy for AMBER on Tianhe-2 supercomputer. The parallel optimization of AMBER is carried out on three different levels: fine grained OpenMP parallel on a single MIC, single-node CPU/MIC collaborated parallel optimization and multi-node multi-MIC collaborated parallel acceleration. By the three levels of parallel acceleration strategy above, we achieved the highest speedup of 25-33 times compared with the original program. Source Code: https://github.com/tianhe2/mAMBER. Shaoliang Peng, Xiaoyu Zhang 0008, Yutong Lu, Xiangke Liao, Kai Lu 0001, Canqun Yang, Jie Liu 0002, Weiliang Zhu |
BIBM | 1 |
| 2016 | Parallel algorithms for large-scale biological sequence alignment on Xeon-Phi based clustersabstractBACKGROUND: Computing alignments between two or more sequences are common operations frequently performed in computational molecular biology. The continuing growth of biological sequence databases establishes the need for their efficient parallel implementation on modern accelerators. RESULTS: This paper presents new approaches to high performance biological sequence database scanning with the Smith-Waterman algorithm and the first stage of progressive multiple sequence alignment based on the ClustalW heuristic on a Xeon Phi-based compute cluster. Our approach uses a three-level parallelization scheme to take full advantage of the compute power available on this type of architecture; i.e. cluster-level data parallelism, thread-level coarse-grained parallelism, and vector-level fine-grained parallelism. Furthermore, we re-organize the sequence datasets and use Xeon Phi shuffle operations to improve I/O efficiency. CONCLUSIONS: Evaluations show that our method achieves a peak overall performance up to 220 GCUPS for scanning real protein sequence databanks on a single node consisting of two Intel E5-2620 CPUs and two Intel Xeon Phi 7110P cards. It also exhibits good scalability in terms of sequence length and size, and number of compute nodes for both database scanning and multiple sequence alignment. Furthermore, the achieved performance is highly competitive in comparison to optimized Xeon Phi and GPU implementations. Our implementation is available at https://github.com/turbo0628/LSDBS-mpi . Haidong Lan, Yuandong Chan, Bertil Schmidt, Shaoliang Peng |
BMC Bioinform. | 5 |
| 2015 | mD3DOCKxb: A Deep Parallel Optimized Software for Molecular Docking with Intel Xeon Phi CoprocessorsabstractMolecular docking is a time consuming process, and it requires a substantial amount of computing power. D3DOCkxb was developed for investigating the effects of halogen bond in drug discovery by adding two precise score functions to Auto Dock. The docking accuracy of D3DOCkxb is better than Auto Dock, which can be attributed to a more complicated processing logic of D3DOCkxb. Consequently, it is an even more challenging task to do parallel optimization on D3DOCkxb. In this paper, we developed mD3DOCkxb, a MIC enabled version of D3DOCkxb, which utilizes Intel Xeon Phi, a Many-Integrated Core (MIC) accelerator, to boost the docking performance. We parallelized the Lamarckian Genetic Algorithm (LGA) in D3DOCKxb with OpenMP and port it to MIC with a number of optimization. And 12x to 18x speedup can be achieved, depending on the number of LGA iterations. Shaoliang Peng, Yutong Lu, Weiliang Zhu, Xinben Zhang |
CCGRID | 2 |
| 2015 | mAMBER: Accelerating Explicit Solvent Molecular Dynamic with Intel Xeon Phi Many-Integrated Core CoprocessorsabstractMolecular dynamics (MD) is a computer simulation of physical movements of atoms and molecules, which is a very important research technique for the study of biological and chemical systems at micro-scale. Assisted Model Building with Energy Refinement (AMBER) is one of the most commonly used software for MD. However, the microsecond MD simulation of large-scale atom system requires a lot of computation power. In this paper, we propose mAMBER: an Intel Xeon Phi Many-Integrated Core (MIC) Coprocessors accelerated implementation of explicit solvent all-atom classical molecular dynamics (MD) within the AMBER program package. We mAMBER also includes new parallel algorithm using CPUs and MIC coprocessors on Tianhe-2 supercomputer. With several optimizing techniques including CPU/MIC collaborated parallelization, factorization and asynchronous data transfer framework, we can accelerate the sander program of AMBER (version 12) in 'offload' mode, and achieves a 4.17-fold overall speedup compared with the CPU-only sander program. Shaoliang Peng, Canqun Yang, Chengkun Wu, Haiqiang Wang, Weiliang Zhu, Jinan Wang |
CCGRID | 2 |
| 2015 | The Challenge of Scaling Genome Big Data Analysis Software on TH-2 SupercomputerabstractWhole genome re-sequencing plays a crucial role in biomedical studies. The emergence of genomic big data calls for an enormous amount of computing power. However, current computational methods are inefficient in utilizing available computational resources. In this paper, we address this challenge by optimizing the utilization of the fastest supercomputer in the world - TH-2 supercomputer. TH-2 is featured by its neo-heterogeneous architecture, in which each compute node is equipped with 2 Intel Xeon CPUs and 3 Intel Xeon Phi coprocessors. The heterogeneity and the massive amount of data to be processed pose great challenges for the deployment of the genome analysis software pipeline on TH-2. Runtime profiling shows that SOAP3-dp and SOAPsnp are the most time-consuming components (up to 70% of total runtime) in a typical genome-analyzing pipeline. To optimize the whole pipeline, we first devise a number of parallel and optimization strategies for SOAP3-dp and SOAPsnp, respectively targeting each node to fully utilize all sorts of hardware resources provided both by CPU and MIC. We also employ a few scaling methods to reduce communication between different nodes. We then scaled up our method on TH-2. With 8192 nodes, the whole analyzing procedure took 8.37 hours to finish the analysis of a 300 TB dataset of whole genome sequences from 2,000 human beings, which can take as long as 8 months on a commodity server. The speedup is about 700x. Shaoliang Peng, Xiangke Liao, Canqun Yang, Yutong Lu, Jie Liu 0002, Yingbo Cui 0001, Chengkun Wu, Bingqiang Wang |
CCGRID | 1 |
| 2015 | A Method to Accelerate GROMACS in Offload Mode on Tianhe-2 SupercomputerabstractMolecular Dynamics(MD) is a computer simulation of physical movements of atoms and molecules in the context of N-body simulation, and is an important part of pharmaceutical industry. GROMACS, which is the most popular software for MD, could not perform satisfactorily with large-scale for the limit of computing resources. In this paper, we proposed a method to accelerate GROMACS with offload mode. In this mode, GROMACS could be arranged efficiently with CPU and the Intel® Xeon PhiTM Many Integrated Core (MIC) coprocessors at the same time, making the full use of Tianhe-2 supercomputer resources. To promote the efficiency of GROMACS, we proposed a series of methods, such as synchronization, data reassemble and array reuse. As we known, we are the first to accelerate GROMACS in offload mode on MIC. Haiqiang Wang, Shaoliang Peng, Xiaoqian Zhu, Chengkun Wu, Weiliang Zhu, Jinan Wang, Huaiyu Yang |
CCGRID | 2 |
| 2015 | MICA: A fast short-read aligner that takes full advantage of Many Integrated Core Architecture (MIC)abstractBACKGROUND: Short-read aligners have recently gained a lot of speed by exploiting the massive parallelism of GPU. An uprising alterative to GPU is Intel MIC; supercomputers like Tianhe-2, currently top of TOP500, is built with 48,000 MIC boards to offer ~55 PFLOPS. The CPU-like architecture of MIC allows CPU-based software to be parallelized easily; however, the performance is often inferior to GPU counterparts as an MIC card contains only ~60 cores (while a GPU card typically has over a thousand cores). RESULTS: To better utilize MIC-enabled computers for NGS data analysis, we developed a new short-read aligner MICA that is optimized in view of MIC's limitation and the extra parallelism inside each MIC core. By utilizing the 512-bit vector units in the MIC and implementing a new seeding strategy, experiments on aligning 150 bp paired-end reads show that MICA using one MIC card is 4.9 times faster than BWA-MEM (using 6 cores of a top-end CPU), and slightly faster than SOAP3-dp (using a GPU). Furthermore, MICA's simplicity allows very efficient scale-up when multiple MIC cards are used in a node (3 cards give a 14.1-fold speedup over BWA-MEM). SUMMARY: MICA can be readily used by MIC-enabled supercomputers for production purpose. We have tested MICA on Tianhe-2 with 90 WGS samples (17.47 Tera-bases), which can be aligned in an hour using 400 nodes. MICA has impressive performance even though MIC is only in its initial stage of development. AVAILABILITY AND IMPLEMENTATION: MICA's source code is freely available at http://sourceforge.net/projects/mica-aligner under GPL v3. SUPPLEMENTARY INFORMATION: Supplementary information is available as "Additional File 1". Datasets are available at www.bio8.cs.hku.hk/dataset/mica. Ruibang Luo, Jeanno Cheung, Edward Wu, Sze-Hang Chan, Wai-Chun Law, Guangzhu He, Chi-Man Liu, Dazong Zhou, Yingrui Li, Ruiqiang Li, Jun Wang 0004, Xiaoqian Zhu, Shaoliang Peng, Tak Wah Lam |
BMC Bioinform. | 15 |
| 2015 | HeMatch: A redundancy layout placement scheme for erasure-coded storages in practical heterogeneous failure patterns
Shanshan Li 0001, Xiangke Liao, Shaoliang Peng, Xiaodong Liu 0004, Zhouyang Jia |
Sci. China Inf. Sci. | 4 |
| 2014 | IMGPU: GPU-Accelerated Influence Maximization in Large-Scale Social NetworksabstractInfluence Maximization aims to find the top-$(K)$ influential individuals to maximize the influence spread within a social network, which remains an important yet challenging problem. Proven to be NP-hard, the influence maximization problem attracts tremendous studies. Though there exist basic greedy algorithms which may provide good approximation to optimal result, they mainly suffer from low computational efficiency and excessively long execution time, limiting the application to large-scale social networks. In this paper, we present IMGPU, a novel framework to accelerate the influence maximization by leveraging the parallel processing capability of graphics processing unit (GPU). We first improve the existing greedy algorithms and design a bottom-up traversal algorithm with GPU implementation, which contains inherent parallelism. To best fit the proposed influence maximization algorithm with the GPU architecture, we further develop an adaptive K-level combination method to maximize the parallelism and reorganize the influence graph to minimize the potential divergence. We carry out comprehensive experiments with both real-world and sythetic social network traces and demonstrate that with IMGPU framework, we are able to outperform the state-of-the-art influence maximization algorithm up to a factor of 60, and show potential to scale up to extraordinarily large-scale networks. Xiaodong Liu 0004, Mo Li 0001, Shanshan Li 0001, Shaoliang Peng, Xiangke Liao, Xiaopei Lu |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2014 | Know by a handful the whole sack: efficient sampling for top-k influential user identification in large graphs
Xiaodong Liu 0004, Shanshan Li 0001, Xiangke Liao, Shaoliang Peng, Zhiyin Kong |
World Wide Web | 4 |
| 2013 | The architecture and traffic management of wireless collaborated hybrid data center networkabstractThis paper introduces a novel wireless collaborated hybrid data center architecture called RF-HYBRID that could optimize the effect of wireless transmission while reduce the complexity of wired network. RF-HYBRID improves throughput and packet delivery latency through flexible wireless detours and shortcuts, with a comprehensive routing and congestion control method. Xiangke Liao, Shanshan Li 0001, Shaoliang Peng, Xiaodong Liu 0004, Bin Lin 0011 |
SIGCOMM | 4 |
| 2013 | INCOME: Practical land monitoring in precision agriculture with sensor networks
Shanshan Li 0001, Shaoliang Peng, Xiaopei Lu |
Comput. Commun. | 2 |
| 2012 | A scalable code dissemination protocol in heterogeneous wireless sensor networks
Shaoliang Peng, Shanshan Li 0001, Xiangke Liao, Yuxing Peng 0001, Nong Xiao 0001 |
Sci. China Inf. Sci. | 1 |
| 2012 | Fast Release/Capture Sampling in Large-Scale Sensor NetworksabstractEfficient estimation of global information is a common requirement for many wireless sensor network applications. Examples include counting the number of nodes alive in the network and measuring the scale of physically correlated events. These tasks must be accomplished at extremely low overhead due to the severe resource limitation of sensor nodes, which poses a challenge for large-scale sensor networks. In this paper, we develop a novel protocol FLAKE to efficiently and accurately estimate the global information of large-scale sensor networks based on the sparse sampling theory. Specially, FLAKE disseminates a small number of messages called seeds to the network and issues a query about which nodes receive a seed. The number of nodes that have the information of interest can be estimated by counting the seeds disseminated, the nodes queried, and the nodes that receive a seed. FLAKE can be easily implemented in a distributed manner due to its simplicity. Moreover, desirable tradeoffs can be achieved between the accuracy of estimation and the system overhead. Our simulations show that FLAKE significantly outperforms several existing schemes on accuracy, delay, and message overhead. Shaoliang Peng, Guoliang Xing, Shanshan Li 0001, Weijia Jia 0001, Yuxing Peng 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2010 | Exploring the practicability of mobile sensors in complex environment surveillanceabstractMobile sensors are often employed for enhancing the sensing coverage and assisting the data gathering in wireless sensor networks. Equipped with unlimited mobility, they are able to move anywhere within the monitored field. Despite the promising simulation (or testbed) result and theoretical conclusion in paper, we have to realize the assumption on unlimited mobility has limitations and is practically unrealistic in many applications. Shanshan Li 0001, Si Zheng 0003, Xiangke Liao, Shaoliang Peng |
IWQoS | 4 |
| 2010 | MPIActor: A thread-based MPI program acceleratorabstractTowards gaining the performance improvement benefited from threaded MPI while supporting MPI standard well, in this paper, we propose a thread-based MPI program accelerator (MPIActor). MPIActor is a transparent middleware to assist general MPI libraries. People can choose to adopt or abandon MPIActor freely in compiling time for any MPI program (Currently only support C code). With the join of MPIActor, in each node, the MPI processes will be mapped as several threads of one process, and the intra-node point-to-point communication and collective communication will have been enhanced by take advantage of thread based mechanism. We have implemented the point-to-point communication module of our design and evaluated it on a real platform. Comparing with MVAPICH2, the experimental results of OSU PINGPONG benchmark show a significant performance improvement from 114% to 321% for transferring messages which size is between 4KB and 2MB. Junqiang Song, Shaoliang Peng |
IWQoS | 3 |
| 2010 | Fish a lake: Fast release/capture sampling in large-scale sensor networksabstractEfficient estimation of global information is a common requirement for many wireless sensor network applications. Examples include counting the number of nodes alive in the network and measuring the scale of physically correlated events. These tasks must be accomplished at extremely low overhead due to the severe resource limitation of sensor nodes, which poses a challenge for large-scale sensor networks. In this paper, we develop a novel protocol called FLAKE that can efficiently and accurately estimate the global information of large-scale sensor networks based on the sparse sampling theory. Specially, FLAKE disseminates a small number of messages called seeds to the network and issues a query about which nodes receive a seed. The number of nodes that have the information of interest can be estimated by counting the seeds disseminated, the nodes queried, and the nodes that receive a seed. FLAKE can be easily implemented in a distributed manner due to its simplicity. Moreover, desirable trade-offs can be achieved between the accuracy of estimation and the system overhead. Our simulations show that FLAKE significantly outperforms several existing schemes on accuracy, delay and message overhead. Shaoliang Peng, Guoliang Xing, Shanshan Li 0001, Weijia Jia 0001, Yuxing Peng 0001 |
IWQoS | 1 |
| 2009 | EDEVS : A Scalable DEVS Formalism for Event-Scheduling Based Parallel and Distributed SimulationsabstractScalability is very important for parallel and distributed simulations. Several techniques have been proposed to develop scalable synchronization strategies, communication services or fundamental algorithms, while little has been seen to deal with the modeling stage of the application. Learning from the HPC (High Performance Computing) lesson, it is clear that the time spent in developing a simulation application must be considered in evaluating the scalability of the application. There are many discrete event simulation platforms built for large parallel and distributed simulations, such as SPEEDES (Synchronous Parallel Environment for Emulation and Discrete Event Simulation), GTW (Georgia tech Time Warp), and YHSUPE, etc. They take Event-Scheduling as their modeling paradigm and have achieved great runtime performance, but lack in providing efficient modeling methods. To deal with this issue, a component-based specification, which can support hierarchical decomposition of large models and facilitate model reuse, is presented. This paper extends the DEVS (Discrete Event simulation specification) and proposes a component-based formalism, called EDEVS (Event-Scheduling Discrete Event simulation Specification) for the existing Event-Scheduling parallel and distributed simulation platforms. Yiping Yao, Shaoliang Peng |
DS-RT | 3 |
| 2009 | Estimation of a Population Size in Large-Scale Wireless Sensor Networks
Shaoliang Peng, Shanshan Li 0001, Xiangke Liao, Yuxing Peng 0001, Nong Xiao 0001 |
J. Comput. Sci. Technol. | 1 |
| 2008 | SenCast: Scalable multicast in wireless sensor networksabstractMulticast is essential for wireless sensor network (WSN) applications. Existing multicast protocols in WSNs are often designed in a P2P pattern, assuming small number of destination nodes and frequent changes on network topologies. In order to truly adopt multicast in WSNs, we propose a base-station model- based multicast, SenCast, to meet the general requirements of applications. SenCast is scalable and energy-efficient for large group communications in WSNs. Theoretical analysis shows that SenCast is able to approximate the Minimum Nonleaf Nodes (MNN) problem to a ratio of ln\R\ (R is the set of all destinations), best known lowest bound. We evaluate our design through comprehensive simulations. Experimental results demonstrate that SenCast outperforms previous multicast protocols including the most recent work uCast. Shaoliang Peng, Shanshan Li 0001, Lei Chen 0002, Nong Xiao 0001, Yuxing Peng 0001 |
IPDPS | 1 |
| 2008 | Using cable-based mobile sensors to assist environment surveillanceabstractThere have been works done by utilizing mobile sensors as supplementary to assist the sensing coverage for the static sensor nodes in possible event happenings. Most of them assume that the mobile sensors are equipped with unlimited mobility and thus can move anywhere within the monitored field. However, the assumption of unlimited mobility has its own limitations and is unrealistic in many practical applications. Alternatively, we consider pre-deploying cables within the monitored field so that mobile sensors move along cables to destinations. In this case, we can far more relax the requirements on the mobile sensors and achieve more realistic usage despite of the complex field landforms. We find that existing cables deployed in the tunnel are perfect carriers for deploying mobile nodes which help get rid of the complex circumstance in the underground tunnel. Shanshan Li 0001, Shaoliang Peng, Mo Li 0001, Xiangke Liao |
MASS | 2 |
| 2008 | Scalable Base-Station Model-Based Multicast in Wireless Sensor Networks
Shaoliang Peng, Shanshan Li 0001, Lei Chen 0002, Yuxing Peng 0001, Nong Xiao 0001 |
J. Comput. Sci. Technol. | 1 |
| 2007 | A Framework for Congestion Control for Reliable Data Delivery in Wireless Sensor NetworksabstractWSN congestion occurs when offered traffic load exceeds available capacity. It causes overall channel quality to degrade and drop rates to rise. Furthermore, redundant transmissions are always adopted to guarantee reliable data delivery, which may deteriorate congestion since they bring on more contention and in reverse hampers the reliability. In this paper, we propose a framework to avoid, detect and mitigate congestion effectively. In this framework, a congestion aware traffic allocation (COTA) is used in multipath routing to balance traffic around the whole network COTA uses some heuristic information to analyze the potential congestion region and avoid traversing these regions. A runtime traffic adjustment CODEM is presented to use accurate metrics to detect and mitigate congestion. Compared with previous works, our work can control congestion while achieving the desired reliability at the same time. Comprehensive simulations have validated the distinguished performance in several aspects of our framework. Shanshan Li 0001, Shaoliang Peng, Xiangke Liao, Peidong Zhu, Yuxing Peng 0001 |
Integrated Network Management | 2 |
| 2007 | Real-Time Data Delivery in Wireless Sensor Networks: A Data-Aggregated, Cluster-Based Adaptive Approach
Shaoliang Peng, Shanshan Li 0001, Yuxing Peng 0001, Wen-sheng Tang, Nong Xiao 0001 |
UIC | 1 |
| 2006 | A Trust-Based Routing Framework in Energy-Constrained Wireless Sensor Networks
Wei-Fang Cheng, Xiangke Liao, Changxiang Shen, Shanshan Li 0001, Shaoliang Peng |
WASA | 5 |
| 2006 | Path Selection of Reliable Data Delivery in Wireless Sensor Networks
Xiangke Liao, Shanshan Li 0001, Peidong Zhu, Shaoliang Peng, Wei-Fang Cheng, Dezun Dong |
WASA | 4 |