Weiming Xiang 0003

dblp:72/5686-3 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-9307-6965ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SVScope: Structural Variation Detection for Short Reads via Multi-Source Fusion and Visual Filtering
abstract
Structural variations (SVs) are one of the major sources of genomic diversity and are closely associated with human disease. Existing short-read-based SV detection tools often rely on limited alignment features, which restricts their ability to fully capture variation signals. While many multi-source fusion methods have improved recall rates, they also introduce a large number of false positives. To address these issues, we present SVScope, an integrated SV detection tool that fuses multi-source signals, structure-sensitive image encoding, and deep visual filtering. It first consolidates candidate variations derived from complementary alignment signals across tools into a standardized candidate set via a harmonized detection and merging pipeline. To further improve specificity, SVScope utilizes a dynamic window to extract alignment features from candidate regions, then encodes them into a seven-channel image with CIGAR operations and read pair orientations to capture local structural details. It also combines a convolutional neural network with an embedded attention mechanism to enhance effective signals and suppress redundant noise, thereby achieving precise filtering of false positives. Benchmarking on real datasets demonstrates that SVScope consistently improves recall compared to individual detection tools while substantially enhancing overall precision through deep learning-based filtering. These results highlight SVScope's capacity to balance sensitivity and specificity, offering a precise and robust solution for SV analysis in large-scale short-read sequencing studies. The code and documentation of SVScope are publicly available at https://github.com/nudt-bioinfo/SVScope
Weiming Xiang 0003, Tao Tang 0001, Yingbo Cui 0001
BIBM4
2025 MMF-SV: A Multi-Modal Feature Fusion-Based Structural Variant Caller
abstract
Structural variant (SV) calling plays a critical role in understanding genome diversity and disease mechanisms. Although deep learning techniques have been increasingly applied to SV identification, existing general-purpose models still face significant challenges, including incomplete extraction of alignment signals, limited accuracy and efficiency, and poor performance in highly polymorphic or structurally complex genomic regions. These limitations lead to suboptimal detection accuracy in current SV callers. In this work, we present MMF-SV, a multi-modal feature fusion-based model (MMF) for SV calling. MMF-SV integrates matching patterns and statistical information from CIGAR signals with textual features extracted from alignment information, enabling comprehensive representation of diverse SV signals. We trained MMF-SV using CLIP, and the trained model achieved over 96% F1 score for classifying various types of variations. We validated the stability and robustness of the MMF-SV model through 5-fold cross-validation. Compared to existing long-read SV callers, MMF-SV achieves higher accuracy and can be effectively integrated with them to significantly reduce the number of false positives in the calling results.
Canqun Yang, Haoang Chi, Tao Tang 0001, Weiming Xiang 0003, Yingbo Cui 0001
ACM Multimedia5
2024 CSV-Filter: a deep learning-based comprehensive structural variant filtering method for both short and long reads
abstract
MOTIVATION: Structural variants (SVs) play an important role in genetic research and precision medicine. As existing SV detection methods usually contain a substantial number of false positive calls, approaches to filter the detection results are needed. RESULTS: We developed a novel deep learning-based SV filtering tool, CSV-Filter, for both short and long reads. CSV-Filter uses a novel multi-level grayscale image encoding method based on CIGAR strings of the alignment results and employs image augmentation techniques to improve SV feature extraction. CSV-Filter also utilizes self-supervised learning networks for transfer as classification models, and employs mixed-precision operations to accelerate training. The experiments showed that the integration of CSV-Filter with popular SV detection tools could considerably reduce false positive SVs for short and long reads, while maintaining true positive SVs almost unchanged. Compared with DeepSVFilter, a SV filtering tool for short reads, CSV-Filter could recognize more false positive calls and support long reads as an additional feature. AVAILABILITY AND IMPLEMENTATION: https://github.com/xzyschumacher/CSV-Filter.
Weiming Xiang 0003, Qingzhe Wang, Xingze Li, Junyu Gao 0005, Tao Tang 0001, Canqun Yang, Yingbo Cui 0001
Bioinform.2
2022 MSVF: Multi-task Structure Variation Filter with Transfer Learning in High-throughput Sequencing
abstract
The single molecule real-time sequencing technologies, such as PacBio and Nanopore, have higher throughput and produce longer reads, which promote the discovery of more structure variations that cannot be discovered by the second-generation sequencing data. However, compared with the second-generation sequencing data, the PacBio data lacks paired-end sequencing information, making traditional structure variations filter fail to process the new data. To solve this problem, this paper proposes a universal multi-tasking structure variation filtering model MSVF. MSVF adopts the CIGAR string defined in SAM format. CIGAR is not limited by sequencing technology or alignment algorithms, so MSVF is suitable for not only the second-generation but also the third-generation sequencing data. Moreover, CIGAR string preserves the complete sequence alignment information, which makes MSVF a highly precise model. Besides, MSVF uses deep learning methods, making it supports more structure variation types, including deletion and insertion. We trained and tested the models on the open-access NCBI datasets. The experiments proved that ShuffleNet, MobileNet, ResNet transfer learning models achieve better classification results on SVs task. The average AUC reaches more than 90% and the AUC of each category reach more than 87%. The accuracy and AUC of deletion and insertion structure variations were above 90% and above 92%, respectively. The code and data can be obtained at https://github.con weimingxiang/MSVF.
Weiming Xiang 0003, Yingbo Cui 0001, Yaning Yang, Shaoliang Peng
BIBM1
2021 H-VAE: A Hybrid Variational AutoEncoder with Data Augmentation in Predicting CRISPR/Cas9 Off-target
abstract
CRISPR/Cas9-based gene editing technology has been widely used in various cells and organisms. However, the off-target effects will bring unpredictable consequences to the organism edited. One of the main obstacles to predict CRISPR/Cas9 off-target is the imbalance of the number of positive and negative samples, which puts forward a challenge for the training of traditional deep learning algorithms. In this paper, we proposed H-VAE, a hybrid variational autoencoder model with data augmentation. This model can extract more abundant sgRNA-DNA base pair matching information, and reduce the risk of overfitting. Moreover, the sample imbalance is resolved. H-VAE can make use of underlying information of training sample, extracted by VAE, to alleviate data-imbalance problem. In view of the weak ability to extract base pair matching information of existing models, a different encoding scheme based on pair encoding is proposed, which enables the model to make full use of sgRNA-DNA base pair matching information. On the Mismatch data set, compared with DeepCRISPR, the ROC-AUC and PR-AUC increased by 0.6% and 41.9%, respectively. In the new Indels data set test scenario, compared with CRISPR-Net, the ROC-AUC and PR-AUC were increased by 1.5% and 133.4% respectively. This proves that H-VAE can improve off-target prediction in various scenarios. The improvement of PR-AUC shows that H-VAE can significantly improve the effect of unbalanced classification. The experimental results demonstrate that H-VAE could achieve a better effect compared with state-of-the-art CRISPR/Cas9 off-target methods on various types of data sets. The code and data can be obtained at https://github.com/weimingxiang/H-VAE.
Weiming Xiang 0003, Dong Chen 0013, Yingbo Cui 0001, Shaoliang Peng
BIBM1
2021 A Knowledge-aware Machine Reading Comprehension Framework for Dialogue Symptom Diagnosis
abstract
Symptom diagnosis in dialogue remains a challenging task because the symptom entities and their status need to be extracted correctly at the same time. Most previous studies treat symptom diagnosis as a classification or sequence labeling task and focus on using single-sentence dialogue as input. Unique from past studies, in this paper, we propose a new framework for dialogue symptom diagnosis, which formulate it as a machine reading comprehension (MRC) task. We first use window-level multi-turn of dialogue as input and extract the symptom entities. Then, we generate a question for each entity to infer the symptom status in the form of question answering (QA). Benefit from the MRC formalization, our proposed framework can encode more informative prior knowledge, which can effectively improve the performance of symptom status inference. Experiments on the Chinese medical dialogue dataset show that the proposed framework outperforms the previous best model and several competitive baselines, which indicates that our framework provides a useful direction for dialogue symptom diagnosis. The code and data are publicly available at https://github.com/zhaoxiongjun/DSD.
Xiongjun Zhao, Yingjie Cheng, Weiming Xiang 0003, Jiandong Shang, Shaoliang Peng
BIBM3