EDBT 2026 Demo / reviewers in the wild / expert
Jingyang Gao
dblp:164/0467 · also Jing-Yang Gao
· DBLP profile ↗
19ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0003-1270-6257ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 7 since 2021Systems, architecture and hardware · 1Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VirNucPro: an identifier for the identification of viral short sequences using six-frame translation and large language modelsabstractViruses are ubiquitous in nature, yet our understanding of them remains limited. High-throughput sequencing technology facilitates the unbiased revelation of genetic composition in samples; however, viral sequences typically make up a small proportion of the entire sequencing data, making it challenging to accurately identify the few or fragmented viral sequences present in a sample. The limited features and information provided by short sequences result in insufficient resolution of viral sequences by existing models. Therefore, we propose a new model, VirNucPro, for short viral sequence identification. Based on a six-frame translation strategy and large language models, we combine nucleotide and amino acid sequence information to enhance feature extraction for short sequences, achieving high accuracy in identifying short viral sequences. Ablation experiments compared the contributions of nucleotide and amino acid sequence features to the model, confirming that the introduced amino acid features significantly contribute to the classification results. Our model outperforms others, such as GCNFrame, DeepVirFinder, DETIRE, and Virtifier, which have demonstrated good performance in identifying short viral sequences of 300 and 500 bp. Our model demonstrates excellent performance on carefully created real-world datasets. Additionally, it can scan for prophage regions within long bacterial fragments, offering a wide range of applications. The codes are available at: https://github.com/Li-Jing-1997/VirNucPro. Fengjuan Tian, Jingyang Gao, Yigang Tong |
Briefings Bioinform. | 6 |
| 2025 | DeepPFP: a multi-task-aware architecture for protein function predictionabstractDeriving protein function from protein sequences poses a significant challenge due to the intricate relationship between sequence and function. Deep learning has made remarkable strides in predicting sequence-function relationships. However, models tailored for specific tasks or protein types encounter difficulties when using transfer learning across domains. This is attributed to the fact that protein function relies heavily on structural characteristics rather than mere sequence information. Consequently, there is a pressing need for a model capable of capturing shared features among diverse sequence-function mapping tasks to address the generalization issue. In this study, we explore the potential of Model-Agnostic Meta-Learning combined with a protein language model called Evolutionary Scale Modeling to tackle this challenge. Our approach involves training the architecture on five out-domain deep mutational scanning (DMS) datasets and evaluating its performance across four key dimensions. Our findings demonstrate that the proposed architecture exhibits satisfactory performance in terms of generalization and employs an effective few-shot learning strategy. To explain further, Compared to the best results, the Pearson's correlation coefficient (PCC) in the final stage increased by ~0.31%. Furthermore, we leverage the trained architecture to predict binding affinity scores of the DMS dataset of SARS-CoV-2 using transfer learning. Notably, training on a subset of the Ube4b dataset with 500 samples resulted in a notable improvement of 0.11 in the PCC. These results underscore the potential of our conceptual architecture as a promising methodology for multi-task protein function prediction. Zilin Ren, Jinghong Sun, Yongbing Chen, Xiaochen Bo, Jiguo Xue, Jingyang Gao |
Briefings Bioinform. | 7 |
| 2024 | GGN-GO: geometric graph networks for predicting protein function by multi-scale structure featuresabstractRecent advances in high-throughput sequencing have led to an explosion of genomic and transcriptomic data, offering a wealth of protein sequence information. However, the functions of most proteins remain unannotated. Traditional experimental methods for annotation of protein functions are costly and time-consuming. Current deep learning methods typically rely on Graph Convolutional Networks to propagate features between protein residues. However, these methods fail to capture fine atomic-level geometric structural features and cannot directly compute or propagate structural features (such as distances, directions, and angles) when transmitting features, often simplifying them to scalars. Additionally, difficulties in capturing long-range dependencies limit the model's ability to identify key nodes (residues). To address these challenges, we propose a geometric graph network (GGN-GO) for predicting protein function that enriches feature extraction by capturing multi-scale geometric structural features at the atomic and residue levels. We use a geometric vector perceptron to convert these features into vector representations and aggregate them with node features for better understanding and propagation in the network. Moreover, we introduce a graph attention pooling layer captures key node information by adaptively aggregating local functional motifs, while contrastive learning enhances graph representation discriminability through random noise and different views. The experimental results show that GGN-GO outperforms six comparative methods in tasks with the most labels for both experimentally validated and predicted protein structures. Furthermore, GGN-GO identifies functional residues corresponding to those experimentally confirmed, showcasing its interpretability and the ability to pinpoint key protein regions. The code and data are available at: https://github.com/MiJia-ID/GGN-GO. Jinghong Sun, Jingyang Gao |
Briefings Bioinform. | 8 |
| 2024 | Deletion variants calling in third-generation sequencing data based on a dual-attention mechanismabstractDeletion is a crucial type of genomic structural variation and is associated with numerous genetic diseases. The advent of third-generation sequencing technology has facilitated the analysis of complex genomic structures and the elucidation of the mechanisms underlying phenotypic changes and disease onset due to genomic variants. Importantly, it has introduced innovative perspectives for deletion variants calling. Here we propose a method named Dual Attention Structural Variation (DASV) to analyze deletion structural variations in sequencing data. DASV converts gene alignment information into images and integrates them with genomic sequencing data through a dual attention mechanism. Subsequently, it employs a multi-scale network to precisely identify deletion regions. Compared with four widely used genome structural variation calling tools: cuteSV, SVIM, Sniffles and PBSV, the results demonstrate that DASV consistently achieves a balance between precision and recall, enhancing the F1 score across various datasets. The source code is available at https://github.com/deconvolution-w/DASV. Jingyang Gao |
Briefings Bioinform. | 4 |
| 2024 | ViroISDC: a method for calling integration sites of hepatitis B virus based on feature encodingabstractBACKGROUND: Hepatitis B virus (HBV) integrates into human chromosomes and can lead to genomic instability and hepatocarcinogenesis. Current tools for HBV integration site detection lack accuracy and stability. RESULTS: This study proposes a deep learning-based method, named ViroISDC, for detecting integration sites. ViroISDC generates corresponding grammar rules and encodes the characteristics of the language data to predict integration sites accurately. Compared with Lumpy, Pindel, Seeksv, and SurVirus, ViroISDC exhibits better overall performance and is less sensitive to sequencing depth and integration sequence length, displaying good reliability, stability, and generality. Further downstream analysis of integrated sites detected by ViroISDC reveals the integration patterns and features of HBV. It is observed that HBV integration exhibits specific chromosomal preferences and tends to integrate into cancerous tissue. Moreover, HBV integration frequency was higher in males than females, and high-frequency integration sites were more likely to be present on hepatocarcinogenesis- and anti-cancer-related genes, validating the reliability of the ViroISDC. CONCLUSIONS: ViroISDC pipeline exhibits superior precision, stability, and reliability across various datasets when compared to similar software. It is invaluable in exploring HBV infection in the human body, holding significant implications for the diagnosis, treatment, and prognosis assessment of HCC. Xiaoqi He, Yigang Tong, Jingyang Gao |
BMC Bioinform. | 7 |
| 2024 | MTAF-DTA: multi-type attention fusion network for drug-target affinity predictionabstractBACKGROUND: The development of drug-target binding affinity (DTA) prediction tasks significantly drives the drug discovery process forward. Leveraging the rapid advancement of artificial intelligence, DTA prediction tasks have undergone a transformative shift from wet lab experimentation to machine learning-based prediction. This transition enables a more expedient exploration of potential interactions between drugs and targets, leading to substantial savings in time and funding resources. However, existing methods still face several challenges, such as drug information loss, lack of calculation of the contribution of each modality, and lack of simulation regarding the drug-target binding mechanisms. RESULTS: We propose MTAF-DTA, a method for drug-target binding affinity prediction to solve the above problems. The drug representation module extracts three modalities of features from drugs and uses an attention mechanism to update their respective contribution weights. Additionally, we design a Spiral-Attention Block (SAB) as drug-target feature fusion module based on multi-type attention mechanisms, facilitating a triple fusion process between them. The SAB, to some extent, simulates the interactions between drugs and targets, thereby enabling outstanding performance in the DTA task. Our regression task on the Davis and KIBA datasets demonstrates the predictive capability of MTAF-DTA, with CI and MSE metrics showing respective improvements of 1.1% and 9.2% over the state-of-the-art (SOTA) method in the novel target settings. Furthermore, downstream tasks further validate MTAF-DTA's superiority in DTA prediction. CONCLUSIONS: Experimental results and case study demonstrate the superior performance of our approach in DTA prediction tasks, showing its potential in practical applications such as drug discovery and disease treatment. Jinghong Sun, Jingyang Gao |
BMC Bioinform. | 5 |
| 2023 | Lossless segmentation of cardiac medical images by a resolution consistent network with nondamage data preprocessing
Chenglizhao Chen, Jingyang Gao |
Multim. Tools Appl. | 3 |
| 2021 | An efficient scRNA-seq dropout imputation method using graph attention networkabstractBACKGROUND: Single-cell sequencing technology can address the amount of single-cell library data at the same time and display the heterogeneity of different cells. However, analyzing single-cell data is a computationally challenging problem. Because there are low counts in the gene expression region, it has a high chance of recognizing the non-zero entity as zero, which are called dropout events. At present, the mainstream dropout imputation methods cannot effectively recover the true expression of cells from dropout noise such as DCA, MAGIC, scVI, scImpute and SAVER. RESULTS: In this paper, we propose an autoencoder structure network, named GNNImpute. GNNImpute uses graph attention convolution to aggregate multi-level similar cell information and implements convolution operations on non-Euclidean space on scRNA-seq data. Distinct from current imputation tools, GNNImpute can accurately and effectively impute the dropout and reduce dropout noise. We use mean square error (MSE), mean absolute error (MAE), Pearson correlation coefficient (PCC) and Cosine similarity (CS) to measure the performance of different methods with GNNImpute. We analyze four real datasets, and our results show that the GNNImpute achieves 3.0130 MSE, 0.6781 MAE, 0.9073 PCC and 0.9134 CS. Furthermore, we use Adjusted rand index (ARI) and Normalized mutual information (NMI) to measure the clustering effect. The GNNImpute achieves 0.8199 (ARI) and 0.8368 (NMI), respectively. CONCLUSIONS: In this investigation, we propose a single-cell dropout imputation method (GNNImpute), which effectively utilizes shared information for imputing the dropout of scRNA-seq data. We test it with different real datasets and evaluate its effectiveness in MSE, MAE, PCC and CS. The results show that graph attention convolution and autoencoder structure have great potential in single-cell dropout imputation. Jingyang Gao |
BMC Bioinform. | 3 |
| 2020 | scSNVIndel. accurate and efficient calling of SNVs and indels from single cell sequencing using integrated Bi-LSTMabstractSingle-cell data are sparse and have coverage fluctuations, making it difficult, in comparison with data obtained from next-generation sequencing (NGS), to call single nucleotide variants (SNVs) and indels. Furthermore, most existing sequencing methods are unable to effectively call whole-genome SNVs and indels from single cell sequencing (SCS) data. In this study, we propose a new method for the efficient identification of SNVs and indels from SCS data, called scSNVIndel. scSNVIndel uses bidirectional long short-term memory (Bi-LSTM) as its base and integrates new natural language processing (NLP) technology. It automatically extracts features and accurately calls SNVs and indels when using SCS data, which is characterized by uneven and discontinuous coverage. Moreover, scSNVIndel can call variants from the sequence directly, retaining valuable information from the SCS data, as it does not convert the sequence into an image like the DeepVariant method. The results show that scSNVIndel performs better in terms of accuracy and recall for calling variants, when compared with other existing methods. scSNVIndel is currently an open-source method, available at https://github.com/CSuperlei/scSNVIndel, and its usage methods are published on the following website: https://www.aiguqu.com/2020/06/18/scSNVIndel/. Yufeng Wu 0001, Jingyang Gao |
BIBM | 3 |
| 2020 | A novel synonymous processing method based on amino acid substitution matrics for the classification of G-protein-coupled receptorsabstractExtracting valuable features and filtering out redundancy are the key challenges to determine the overall classification performance for G-protein-coupled receptors (GPCRs). In this study, we consider improving the feature synonym problem, and put forward a novel feature knowledge mining strategy based on functional word clustering and integration. The essence behind the method is the novel feature knowledge mining strategy. Through evaluating the independence of each candidate feature using the evolutionary hypothesis based on residue substitution matrices, clustering candidate features, and fusing them by retaining the main functional words, the proposed strategy adds a layer between the feature extraction layer and the prediction layer. Based on the proposed method, four classic machine learning algorithms in conjunction with the feature extraction method were applied to classify GPCRs at all family levels. Surprisingly, these classifiers achieve considerable performance in almost all evaluation criteria which indicated the validity and superiority of the proposed molecular evolution based feature extraction method. Cheng Ling, Yitian Shen, Jingyang Gao |
BIBM | 3 |
| 2019 | DeepSV: accurate calling of genomic deletions from high-throughput sequencing data using deep convolutional neural networkabstractBACKGROUND: Calling genetic variations from sequence reads is an important problem in genomics. There are many existing methods for calling various types of variations. Recently, Google developed a method for calling single nucleotide polymorphisms (SNPs) based on deep learning. Their method visualizes sequence reads in the forms of images. These images are then used to train a deep neural network model, which is used to call SNPs. This raises a research question: can deep learning be used to call more complex genetic variations such as structural variations (SVs) from sequence data? RESULTS: In this paper, we extend this high-level approach to the problem of calling structural variations. We present DeepSV, an approach based on deep learning for calling long deletions from sequence reads. DeepSV is based on a novel method of visualizing sequence reads. The visualization is designed to capture multiple sources of information in the sequence data that are relevant to long deletions. DeepSV also implements techniques for working with noisy training data. DeepSV trains a model from the visualized sequence reads and calls deletions based on this model. We demonstrate that DeepSV outperforms existing methods in terms of accuracy and efficiency of deletion calling on the data from the 1000 Genomes Project. CONCLUSIONS: Our work shows that deep learning can potentially lead to effective calling of different types of genetic variations that are complex than SNPs. Yufeng Wu 0001, Jingyang Gao |
BMC Bioinform. | 3 |
| 2018 | cnnCNV: A Sensitive and Efficient Method for Detecting Copy Number Variation based on Convolutional Neural Networks
Miaosen Ding, Jingyang Gao, Cheng Ling, Liwei Gao |
BIBM | 2 |
| 2018 | Classification of G-protein Coupled Receptors Based on Semi-navïe Bayesian Inference
Cheng Ling, Lin Yue, Jingyang Gao |
BIBM | 3 |
| 2018 | An Integrated Method of Detecting Copy Number Variation Based on Sequence Assembly
Jingyang Gao |
ICIC (1) | 2 |
| 2017 | An efficient CNN-based classification on G-protein Coupled Receptors using TF-IDF and N-gramabstractProtein sequence classification is increasingly crucial in the current “biological information sciences” epoch, where researchers hammer at functional genomics and proteomics technologies for predicting the function of large-scale new proteins. This has sparked interest in the methods which do not rely on traditional sequence alignment, but prefer machine learning approaches. In this paper, we present a Convolutional Neural Network (CNN) based method to perform the classification on the different levels of G-protein Coupled Receptors (GPCRs). The method is implemented in conjunction with an improved feature extraction method and TF-IDF feature weighting strategy. Experimental results indicate that the proposed method makes significant improvements over previous methods, which attains an accuracy of up to 98.34%, 98.13% and 96.47% in the classification of family level, subfamily level I and II, respectively. In comparison to the other well-known classification methods for GPCRs, the classification error rate of the proposed method is reduced by of at least 55.14% (family level), 72.86% (level I) and 52.63% (Level II). Cheng Ling, Jingyang Gao |
ISCC | 3 |
| 2016 | A high-precision shallow Convolutional Neural Network based strategy for the detection of Genomic DeletionsabstractGenomic Deletion holds the largest proportion of the structural variation (SV). There are many methods for detection of SVs using next-generation data, such as Pindel, SVseq2, BreakDancer, DELLY and so on. However, each method has advantages on only some kind of SVs. For deletions, existing tools usually produce variations with accurate results under 0.5. In this paper, we present CNNdel, a tool based on shallow convolutional neural network to detect genomic deletions with real data from the 1000 Genomes Project. The experimental results show that the accuracy and sensitivity is both improved compared with other existing methods. Jing Wang 0043, Cheng Ling, Jingyang Gao |
BIBM | 3 |
| 2016 | Concod: Accurate consensus-based approach of calling deletions from high-throughput sequencing dataabstractAccurate calling of structural variations such as deletions with short sequence reads from high-throughput sequencing is an important but challenging problem in the field of genome analysis. There are many existing methods for calling deletions. At present, not a single method clearly outperforms all other methods in precision and sensitivity. A popular strategy used by several authors is combining different signatures left by deletions in order to achieve more accurate deletion calling. However, most existing methods using the combining approach are heuristic and the called deletions by these tools still contain many wrongly called deletions. In this paper, we present Concod, a machine learning based framework for calling deletions with consensus, which is able to more accurately detect and distinguish true deletions from falsely called ones. First, Concod collects candidate deletions by merging the output of multiple existing deletion calling tools. Then, features of each candidate are extracted from aligned reads based on multiple detection theories. Finally, a machine learning model is trained with these features and used to classify the true and false candidates. We test our approach on different coverage of real data and compare with existing tools, including Pindel, SVseq2, BreakDancer, and DELLY. Results show that Concod improves both precision and sensitivity of deletion calling significantly. Chong Chu, Yufeng Wu 0001, Jingyang Gao |
BIBM | 5 |
| 2016 | MrBayes 3.2.6 on Tianhe-1A: A High Performance and Distributed Implementation of Phylogenetic AnalysisabstractPhylogenetic analysis has achieved extraordinary results in domains like species delimitation and evolutionary biology. An essential element behind this success has been the introduction of high performance computing techniques in the step of estimating the phylogenetic likelihoods. This paper describes the design and implementation of a distributed and CPU-GPU based heterogeneous computing system on parallelizing the analysis. The parallelization has been implemented in the state-of-the-art version of MrBayes, a widespread phylogeny reconstruction program. We benchmarked the method and another two GPU-based methods by using 8 distributed computing nodes on Tianhe-1A. The experimental results indicate that the proposed method outstrips BEAGLE and the nMC3 method by speedup factors of up to 1.98× and 1.68×, respectively. In comparison to the serially implemented MrBayes, a peak speedup of 188× is finally achieved by using 8 Tesla M 2050 GPUs. The proposed method is publicly available to facilitate further research on phylogenetic analysis. Cheng Ling, Arong Luo, Jingyang Gao |
ICPADS | 3 |
| 2016 | MrBayes tgMC3++: A High Performance and Resource-Efficient GPU-Oriented Phylogenetic Analysis MethodabstractMrBayes is a widespread phylogenetic inference tool harnessing empirical evolutionary models and Bayesian statistics. However, the computational cost on the likelihood estimation is very expensive, resulting in undesirably long execution time. Although a number of multi-threaded optimizations have been proposed to speed up MrBayes, there are bottlenecks that severely limit the GPU thread-level parallelism of likelihood estimations. This study proposes a high performance and resource-efficient method for GPU-oriented parallelization of likelihood estimations. Instead of having to rely on empirical programming, the proposed novel decomposition storage model implements high performance data transfers implicitly. In terms of performance improvement, a speedup factor of up to 178 can be achieved on the analysis of simulated datasets by four Tesla K40 cards. In comparison to the other publicly available GPU-oriented MrBayes, the tgMC3++ method (proposed herein) outperforms the tgMC3(v1.0), nMC3(v2.1.1) and oMC3(v1.00) methods by speedup factors of up to 1.6, 1.9 and 2.9, respectively. Moreover, tgMC3++ supports more evolutionary models and gamma categories, which previous GPU-oriented methods fail to take into analysis. Cheng Ling, Tsuyoshi Hamada, Jingyang Gao, Guoguang Zhao, Donghong Sun, Weifeng Shi |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |