EDBT 2026 Demo / reviewers in the wild / expert
Changyong Yu
dblp:94/6965
· DBLP profile ↗
19ranked-venue papers
13as first author
12since 2021 · last 2026
0000-0002-2803-4291ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 9 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Different Graph-Level Attention Based on Multi-Scale for Predicting lncRNA-Disease AssociationsabstractThe relationship between long non-coding RNAs (lncRNAs) and diseases is crucial for understanding biological processes, as well as the onset, progression, prevention, and treatment of diseases. Accurate prediction of associations between lncRNAs and diseases holds significant potential, offering new insights for biological research and identifying novel therapeutic targets for clinical applications. These predictions enhance research efficiency, reduce unnecessary experimental costs, and improve diagnostic and treatment precision. However, existing methods often fail to simultaneously consider global structural information, local subgraph details, and multi-scale graph information. In this study, we design a graph neural network prediction framework based on multi-scale graph-level attention, designed to predict disease-related candidate lncRNAs, named GLALDA. It employs a dual-level attention mechanism to integrate both global structural and local subgraph information. To enhance the ability of the model to capture the graph structure, edge feature information is incorporated into the attention calculations. Furthermore, we utilize a cross-attention mechanism to deeply fuse feature representations from graphs of different scales, effectively combining local node details with global context. The resulting integrated features are then fed into a scoring network for evaluation. Experimental results on public datasets demonstrate that GLALDA achieves an AUC of 0.949 and an AUPR of 0.947, outperforming six other state-of-the-art methods. Furthermore, through in-depth case studies of three cancers, we further validate GLALDA's capability to identify potential disease-related lncRNA candidates. These findings underscore the framework's potential to advance both biological research and clinical applications. Zhenkang Hu, Jiale Tu, Xuxian Zhou, Yuhai Zhao, Changyong Yu |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2026 | Prediction of Cancer Drug Response Based on Hypergraph Convolutional Network and Contrastive LearningabstractOBJECTIVE: Accurate prediction of cancer drug responses is essential for advancing precision oncology. This work aims to improve generalization and robustness in drug response prediction by modeling complex drug-cell line interactions. METHODS: We propose HypergraphCDR, a Hypergraph Convolutional Network model with Hypergraph Contrastive Learning for cancer drug response prediction. Multiomics features of cancer cell lines are first compressed using an autoencoder. A Hypergraph is then constructed to capture high-order relationships between drugs and cancer cell lines, with Hypergraph Convolutional Networks generating drug embeddings. In parallel, cell line embeddings are generated via a neural network. Drug and cell line embeddings are jointly optimized and integrated into a regression framework using a combination of supervised regression loss and contrastive loss. RESULTS: Extensive experiments demonstrate that HypergraphCDR consistently outperforms state-of-the-art methods in terms of Pearson Correlation Coefficient (${PCC}$), Spearman Correlation Coefficient (${SCC}$), and Coefficient of Determination (${R}^{2}$). Moreover, in independent experiments involving unseen drugs and unseen cell lines, as well as tissue-specific evaluations, HypergraphCDR exhibits superior generalization performance compared with recent baselines. CONCLUSION: By explicitly modeling higher-order drug-cell line relationships, HypergraphCDR substantially enhances the accuracy and robustness of cancer drug response prediction. SIGNIFICANCE: This study provides an effective and generalizable computational framework for drug response prediction, supporting reliable drug screening and treatment strategy development in precision medicine. Yukai Jia, Yuhai Zhao, Changyong Yu |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2025 | GPU-Powered Evolutionary Auxiliary Multitasking for Fast SNP Interaction DetectionabstractIdentifying complex interactions among millions of single nucleotide polymorphisms (SNPs) is a key challenge in Genome-Wide Association Studies (GWAS), offering crucial insights into the genetic architecture of complex diseases. Evolutionary algorithm (EA)-based methods have gained significant attention for their global search capabilities, controllable runtime, and multi-objective optimization potential. However, when applied to high-dimensional GWAS datasets, many existing EA-based methods encounter challenges such as getting trapped in local optima and facing high computational demands. To address these issues, the evolutionary multitasking (EMT) paradigm presents a promising solution, enhancing population diversity and convergence speed through collaborative, cross-task knowledge sharing. Furthermore, the multi-tasking framework and EA can be seamlessly deployed across multiple Graphics Processing Units (GPUs), leveraging their high parallelism and aggregated memory bandwidth. Therefore, we introduce a GPU-powered evolutionary auxiliary multitasking algorithm (GEAMT) for fast SNP interaction detection. GEAMT first constructs a main task along with several low-dimensional auxiliary tasks to redefine the original task. The main task explores the entire search space, while the auxiliary tasks search distinct subspaces to enhance local optimization capabilities. In each iteration, the auxiliary tasks transfer high-quality information to the main task via an information transfer mechanism. Subsequently, an auxiliary task update strategy based on feature regrouping is employed to switch the search subspaces of the auxiliary tasks. The final results are derived from the Pareto-optimal solutions of the main task. Implemented across multiple GPUs, GEAMT achieves notable scalability and efficiency. Comprehensive experiments on both synthetic and real-world datasets demonstrate that GEAMT can significantly enhance search accuracy and speed up the search process. Ying Yin 0001, Xin Wang 0124, Changyong Yu, Yuhai Zhao |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | gShapeLnoc: A Graph Network Incorporating Shapelet Embedding Model for LncRNA Subcellular LocalizationabstractThe subcellular localization of Long non-coding RNAs (LncRNAs) is a pivotal research area with profound implications for understanding underlying molecular mechanisms, involvement in pathological processes, and regulation of gene expression. Traditional machine learning based methods often rely on k-mer frequencies for classification, ignoring the global features of LncRNAs. More recent methods based on deep learning have utilized sequence and graph models for LncRNA classification. However, while these methods could improve their combination of LncRNA features, they still possess limitations, for example, ignoring the fact that mutations could occur in LncRNAs. Simultaneously, it employs the Shapelet model to extract local features of the most representative k-mer among different LncRNA classes. Furthermore, gShapeLnoc combines global and local feature representations for predicting the subcellular localization of LncRNAs. We have evaluated the performance of the gShapeLnoc algorithm on a real dataset, and the results demonstrate that it outperforms existing state-of-the-art methods in terms of accuracy. Changyong Yu, Xuxian Zhou, Yuhai Zhao |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2024 | BL: An Efficient Index for Reachability Queries on Large GraphsabstractReachability query has important applications in many fields such as social networks, semantic web, and biological information networks. How to improve the query efficiency in directed acyclic graph (DAG) has always been the main problem of reachability query research. Existing methods either can't prune unreachable pairs enough or can't perform well on both index size and query time. In this paper, we propose BL (Bridging Label), a general index framework for reachability queries in large graphs. First, we summarize the relationships between BL and existing label indices. Second, we propose a kind of specific index, named minBL, which can avoid redundant labels. Moreover, we propose TFD-minBL and CTFD-minBL, which generate minBL under the TFD-based permutation single-pass and in incremental, respectively. Finally, we conduct a large number of extensive experiments on real datasets. The experimental results show that our methods are much faster and use less storage overhead than state-of-the-art methods. The source codes of BL can be downloaded from web sitehttps://github.com/BioLab310/BL Changyong Yu, Tianmei Ren, Yuhai Zhao |
IEEE Trans. Big Data | 1 |
| 2024 | dwMLCS: An Efficient MLCS Algorithm Based on Dynamic and Weighted Directed Acyclic GraphabstractThe problem of finding the longest common subsequence (MLCS) for multiple sequences is a computationally intensive and challenging problem that has significant applications in various fields such as text comparison, pattern recognition, and gene diagnosis. Currently, the dominant point-based MLCS algorithms have become popular and extensively studied. Generally, they construct the directed acyclic graph (DAG) of matching points and convert the MLCS problem into a search for the longest paths in the DAG. Several improvements have been made, focusing on decreasing model size and reducing redundant computations. These include 1) hash methods for eliminating duplicated nodes, 2) dynamic structures for supporting smaller DAG and 3) path pruning strategy and so on. However, the algorithms are still too limited when facing large-scale MLCS problem due to 1) the dynamic structures are too time-consuming to maintain and 2) the path pruning relies heavily on the tightness of the lower and upper bound of the MLCS. These factors contribute to the large-scale MLCS problem remaining a challenge. We propose a novel algorithm for the large-scale MLCS problem, named dwMLCS. It is based on two models: one is a dynamic DAG model which is both space and time efficient. It can decrease the size of the DAG significantly. The other is a weighted DAG model with new successor strategies. With this model, we design the algorithm for finding a tighter lower bound of the MLCS. Then, the path pruning is conducted to further reduce the size of the DAG and eliminate redundant computation. Additionally, we propose an upper bound method for improving the efficiency of the path pruning strategy. The experimental results demonstrate that the effectiveness and efficiency of the models and algorithms proposed are better than state-of-the-art algorithms. Changyong Yu, Dekuan Gao, Yuhai Zhao, Guoren Wang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2024 | MiniDBG: A Novel and Minimal De Bruijn Graph for Read MappingabstractThe De Bruijn graph (DBG) has been widely used in the algorithms for indexing or organizing read and reference sequences in bioinformatics. However, a DBG model that can locate each node, edge and path on sequence has not been proposed so far. Recently, DBG has been used for representing reference sequences in read mapping tasks. In this process, it is not a one-to-one correspondence between the paths of DBG and the substrings of reference sequence. This results in the false path on DBG, which means no substrings of reference producing the path. Moreover, if a candidate path of a read is true, we need to locate it and verify the candidate on sequence. To solve these problems, we proposed a DBG model, called MiniDBG, which stores the position lists of a minimal set of edges. With the position lists, MiniDBG can locate any node, edge and path efficiently. We also proposed algorithms for generating MiniDBG based on an original DBG and algorithms for locating edges or paths on sequence. We designed and ran experiments on real datasets for comparing them with BWT-based and position list-based methods. The experimental results show that MiniDBG can locate the edges and paths efficiently with lower memory costs. Changyong Yu, Yuhai Zhao, Chu Zhao, Jianyu Jin, Keming Mao, Guoren Wang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2023 | AKGIN: An API Knowledge Graph and Intent Network based Mashup-Oriented API Recommendation MethodabstractWith the continuous development of the Web API ecosystem, mashup-oriented API recommendation gets a lot of attention. Collaborative filtering, deep learning and their combination based methods are recently proposed for API recommendation. However, the recommendation results often perform poorly when data is sparse. Therefore, we propose an API knowledge graph (AKG) and intent network based mashup-oriented API recommendation method, whose main target is to recruit the semantic representation ability to improve the precision of API recommendation. Firstly, Latent Dirichlet Allocation (LDA) is used to extract API topics, which are important entities in the AKG. Then, the mashup intention is modeled based on the mashup-API interactions. The high-order relationship paths are recursively aggregated into the representation vector of Mashup and API. Finally, the inner product of their representations present the probability of the API recommendation for the mashup. The experimental results on real-world datasets show that our proposed method outperforms several baseline methods in mashup-oriented API recommendations. Changyong Yu, Bangjun Wang |
CSCWD | 1 |
| 2023 | Global triangle estimation based on first edge sampling in large graph streams
Changyong Yu, Fazal Wahab, Zihan Ling, Tianmei Ren, Yuhai Zhao |
J. Supercomput. | 1 |
| 2022 | A fast and efficient path elimination algorithm for large-scale multiple common longest sequence problemsabstractBACKGROUND: In various fields, searching for the Longest Common Subsequences (LCS) of Multiple (i.e., three or more) sequences (MLCS) is a classic but difficult problem to solve. The primary bottleneck in this problem is that present state-of-the-art algorithms require the construction of a huge graph (called a direct acyclic graph, or DAG), which the computer usually has not enough space to handle. Because of their massive time and space consumption, present algorithms are inapplicable to issues with lengthy and large-scale sequences. RESULTS: A mini Directed Acyclic Graph (mini-DAG) model and a novel Path Elimination Algorithm are proposed to address large-scale MLCS issues efficiently. In mini-DAG, we employ the branch and bound approach to reduce paths during DAG construction, resulting in a very mini DAG (mini-DAG), which saves memory space and search time. CONCLUSION: Empirical experiments have been performed on a standard benchmark set of DNA sequences. The experimental results show that our model outperforms the leading algorithms, especially for large-scale MLCS problems. Changyong Yu, Pengxi Lin, Yuhai Zhao, Tianmei Ren, Guoren Wang |
BMC Bioinform. | 1 |
| 2022 | StLiter: A Novel Algorithm to Iteratively Build the Compacted de Bruijn Graph From Many Complete GenomesabstractRecently, the compacted de Bruijn graph (cDBG) of complete genome sequences was successfully used in read mapping due to its ability to deal with the repetitions in genomes. However, current approaches are not flexible enough to fit frequently building the graphs with different k-mer lengths. Instead of building the graph directly, how can we build the compacted de Bruijin graph of longer k-mer based on the one of short k-mer? In this article, we present StLiter, a novel algorithm to build the compacted de Bruijn graph either directly from genome sequences or indirectly based on the graph of a short k-mer. For 100 simulated human genomes, StLiter can construct the graph of k-mer length 15-18 in 2.5-3.2 hours with maximal ∼70GB memory in the case of without considering the reverese complements of the reference genomes. And it costs 4.5-5.9 hours when considering the reverse complements. In experiments, we compared StLiter with TwoPaCo, the state-of-art method for building the graph, on 4 datasets. For k-mer length 15-18, StLiter can build the graph 5-9 times faster than TwoPaCo using less maximal memory cost. For k-mer length larger than 18, given the graph of a short (k- x)-mer, such as x= 1-2, compared with TwoPaCo building the graph directly, StLiter can also build the graph more efficiently. The source codes of StLiter can be downloaded from web site https://github.com/BioLab-cz/StLiter. Changyong Yu, Keming Mao, Yuhai Zhao, Guoren Wang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | DiagAF: A More Accurate and Efficient Pre-Alignment Filter for Sequence AlignmentabstractSequence alignment is an essential step in computational genomics. More accurate and efficient sequence pre-alignment methods that run before conducting expensive computation for final verification are still urgently needed. In this article, we propose a more accurate and efficient pre-alignment algorithm for sequence alignment, called DiagAF. Firstly, DiagAF uses a new lower bound of edit distance based on shift hamming masks. The new lower bound makes use of fewer shift hamming masks comparing with state-of-the-art algorithms such as SHD and MAGNET. Moreover, it takes account the information of edit distance path exchanging on shift hamming masks. Secondly, DiagAF can deal with alignments of sequence pairs with not equal length, rather than state-of-the-art methods just for equal length. Thirdly, DiagAF can align sequences with early termination for true alignments. In the experiment, we compared DiagAF with state-of-the-art methods. DiagAF can achieve a much smaller error rate than them, meanwhile use less time than them. We believe that DiagAF algorithm can further improve the performance of state-of-the-art sequence alignment softwares. The source codes of DiagAF can be downloaded from web site https://github.com/BioLab-cz/DiagAF. Changyong Yu, Yuhai Zhao, Chu Zhao, Guoren Wang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | LLR: Learning learning rates by LSTM for training neural networks
Changyong Yu, Xin Qi 0004, Xin He 0045, Cuirong Wang, Yuhai Zhao |
Neurocomputing | 1 |
| 2019 | Diffusion-based kernel matrix model for face liveness detection
Changyong Yu, Chengtang Yao, Mingtao Pei, Yunde Jia |
Image Vis. Comput. | 1 |
| 2015 | Postfix automata
Maohua Jing, Yixian Yang, Ning Lu 0005, Changyong Yu |
Theor. Comput. Sci. | 5 |
| 2010 | A Multi-stage Spectral Alignment Strategy for Unrestrictive PTM Peptide IdentificationabstractSpectral alignment, which studies the matching of ion peaks between the investigated spectrum and theoretical spectrum of peptide in the peptide database, is a very useful topic in computational proteomics. So far, the efficient, accurate and practical spectral alignment algorithm is still urgently needed due to its important application in the PTM unrestrictive peptide identification. In this paper, a multi-stage spectral alignment algorithm called MS-SA is proposed with the following two features: (a) it provided four different levels of alignment aims according to the alignment quality which can be specified by users, (b) it provided the capability of analyzing the detail modification types and locations for spectrum with multiple PTM sites. Therefore, MS-SA is of high practicality and can be applied to different specific applications such as being a filter in the large-scale database searching, a tool for detail modification types and locations analysis in small-scale spectral alignment and so on. A large number of experiments on real MS/MS data have been done for testing the performance of MS-SA. Also, the results of MS-SA are compared with those of same type of algorithms such as SA and SPC. The results show that MS-SA possesses strong practicality and outperforms the SA and SPC algorithms on several aspects. Changyong Yu, Guoren Wang, Yuhai Zhao, Keming Mao |
BIBE | 1 |
| 2009 | Generating Peptide Sequence Tags for Peptide Identification via Tandem Mass SpectrometryabstractLarge-scale, rapid and accurate protein identification is the crucial basis for further protein analysis in computational proteomics. Searching protein database by use of the protein tandem mass spectra has been a standard solution for solving this problem. Though several algorithms have been proposed, more sensitive and accurate approaches are still needed. In this paper, an effective database search approach is proposed. Prior to searching sequence database, an approach based on a graph-theoretic model is proposed to infer the peptide sequence tag (PST) from the tandem mass spectra data which is the partial sequence of the peptide. Also, an index approach for the protein sequence database is proposed for speeding up the database search and filtering out the incorrect protein sequences. Then, a novel scoring method for evaluating the match between the peptide sequence tag and the protein sequence is proposed for improving the accuracy of the database search result. Finally, we develop an algorithm for solving the problem and implement it as a computer program PepCheck. All the results fore-Check are compared with those of the famous algorithms. Experimental results demonstrate that PepCheck is as accurate as or more accurate than them with the test datasets. Changyong Yu, Guoren Wang, Yuhai Zhao, Keming Mao, Wendan Zhai |
BIBE | 1 |
| 2009 | A Novel Multi-reference Points Fingerprint Matching Method
Keming Mao, Guoren Wang, Changyong Yu |
MMM | 3 |
| 2006 | Partition Frequency Distance based Filter Method for Finding Approximate Repetitions in DNA SequencesabstractSearching for approximate repetitions in a DNA sequence has been an important topic in gene analysis. One of the problems in the study is that because of the varying lengths of patterns, the similarity between patterns cannot be judged accurately if we use only the concept of ED (edit distance). In this paper we shall use the function similar to compute similarity, which considers both the difference and sameness between patterns at the same time. Seeing the computational complexity, we shall also propose a new distance PFD (partition frequency distance) and design a new filter based on PFD, with which we can sort out candidate set of approximate repetitions efficiently. We use SUA instead of sliding window to get the fragments in a DNA sequence, so that the patterns of an approximate repetition have no limitation on length. The results show that with this technique we are able to find a bigger number of approximate repetitions than that of those found with tandem repeat finder Guoren Wang, Qingquan Wu, Baichen Chen, Changyong Yu, Ge Yu 0001 |
BIBE | 5 |