VLDB 2026 Research / reviewers in the wild / expert
Chaokun Yan
dblp:15/9378
· DBLP profile ↗
14ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0002-6246-7242ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DeepHapNet: a haplotype assembly method based on RetNet and deep spectral clusteringabstractGene polymorphism originates from single-nucleotide polymorphisms (SNPs), and the analysis and study of SNPs are of great significance in the field of biogenetics. The haplotype, which consists of the sequence of SNP loci, carries more genetic information than a single SNP. Haplotype assembly plays a significant role in understanding gene function, diagnosing complex diseases, and pinpointing species genes. We propose a novel method, DeepHapNet, for haplotype assembly through the clustering of reads and learning correlations between read pairs. We employ a sequence model called Retentive Network (RetNet), which utilizes a multiscale retention mechanism to extract read features and learn the global relationships among them. Based on the feature representation of reads learned from the RetNet model, the clustering process of reads is implemented using the SpectralNet model, and, finally, haplotypes are constructed based on the read clusters. Experiments with simulated and real datasets show that the method performs well in the haplotype assembly problem of diploid and polyploid based on either long or short reads. The code implementation of DeepHapNet and the processing scripts for experimental data are publicly available at https://github.com/wjj6666/DeepHapNet. Jing-Jing Wei, Chaokun Yan, Huimin Luo |
Briefings Bioinform. | 4 |
| 2025 | DconnLoop: a deep learning model for predicting chromatin loops based on multi-source data integrationabstractBACKGROUND: Chromatin loops are critical for the three-dimensional organization of the genome and gene regulation. Accurate identification of chromatin loops is essential for understanding the regulatory mechanisms in disease. However, current mainstream detection methods rely primarily on single-source data, such as Hi-C, which limits these methods' ability to capture the diverse features of chromatin loop structures. In contrast, multi-source data integration and deep learning approaches, though not yet widely applied, hold significant potential. RESULTS: In this study, we developed a method called DconnLoop to integrate Hi-C, ChIP-seq, and ATAC-seq data to predict chromatin loops. This method achieves feature extraction and fusion of multi-source data by integrating residual mechanisms, directional connectivity excitation modules, and interactive feature space decoders. Finally, we apply density estimation and density clustering to the genome-wide prediction results to identify more representative loops. The code is available from https://github.com/kuikui-C/DconnLoop . CONCLUSIONS: The results demonstrate that DconnLoop outperforms existing methods in both precision and recall. In various experiments, including Aggregate Peak Analysis and peak enrichment comparisons, DconnLoop consistently shows advantages. Extensive ablation studies and validation across different sequencing depths further confirm DconnLoop's robustness and generalizability. Kuikui Cheng, Chaokun Yan, Huimin Luo |
BMC Bioinform. | 3 |
| 2024 | ResDeepGS: A Deep Learning-Based Method for Crop Phenotype Prediction
Chaokun Yan, Huimin Luo |
ISBRA (2) | 1 |
| 2024 | Enhancing Privacy and Preserving Accuracy in Medical Image Classification with Limited Labeled Samples
Chaokun Yan, Menghan Yin, Wenjuan Liang, Haicao Yan, Huimin Luo |
ISBRA (1) | 1 |
| 2023 | SLHSD: hybrid scaffolding method based on short and long readsabstractIn genome assembly, scaffolding can obtain more complete and continuous scaffolds. Current scaffolding methods usually adopt one type of read to construct a scaffold graph and then orient and order contigs. However, scaffolding with the strengths of two or more types of reads seems to be a better solution to some tricky problems. Combining the advantages of different types of data is significant for scaffolding. Here, a hybrid scaffolding method (SLHSD) is present that simultaneously leverages the precision of short reads and the length advantage of long reads. Building an optimal scaffold graph is an important foundation for getting scaffolds. SLHSD uses a new algorithm that combines long and short read alignment information to determine whether to add an edge and how to calculate the edge weight in a scaffold graph. In addition, SLHSD develops a strategy to ensure that edges with high confidence can be added to the graph with priority. Then, a linear programming model is used to detect and remove remaining false edges in the graph. We compared SLHSD with other scaffolding methods on five datasets. Experimental results show that SLHSD outperforms other methods. The open-source code of SLHSD is available at https://github.com/luojunwei/SLHSD. Ting Guan, Guolin Chen, Zhonghua Yu, Haixia Zhai, Chaokun Yan, Huimin Luo |
Briefings Bioinform. | 6 |
| 2023 | KGANSynergy: knowledge graph attention network for drug synergy predictionabstractCombination therapy is widely used to treat complex diseases, particularly in patients who respond poorly to monotherapy. For example, compared with the use of a single drug, drug combinations can reduce drug resistance and improve the efficacy of cancer treatment. Thus, it is vital for researchers and society to help develop effective combination therapies through clinical trials. However, high-throughput synergistic drug combination screening remains challenging and expensive in the large combinational space, where an array of compounds are used. To solve this problem, various computational approaches have been proposed to effectively identify drug combinations by utilizing drug-related biomedical information. In this study, considering the implications of various types of neighbor information of drug entities, we propose a novel end-to-end Knowledge Graph Attention Network to predict drug synergy (KGANSynergy), which utilizes neighbor information of known drugs/cell lines effectively. KGANSynergy uses knowledge graph (KG) hierarchical propagation to find multi-source neighbor nodes for drugs and cell lines. The knowledge graph attention network is designed to distinguish the importance of neighbors in a KG through a multi-attention mechanism and then aggregate the entity's neighbor node information to enrich the entity. Finally, the learned drug and cell line embeddings can be utilized to predict the synergy of drug combinations. Experiments demonstrated that our method outperformed several other competing methods, indicating that our method is effective in identifying drug combinations. Zhijie Gao, Chaokun Yan, Wenjuan Liang, Huimin Luo |
Briefings Bioinform. | 3 |
| 2022 | Deep learning approach for cancer subtype classification using high-dimensional gene expression dataabstractMOTIVATION: Studies have shown that classifying cancer subtypes can provide valuable information for a range of cancer research, from aetiology and tumour biology to prognosis and personalized treatment. Current methods usually adopt gene expression data to perform cancer subtype classification. However, cancer samples are scarce, and the high-dimensional features of their gene expression data are too sparse to allow most methods to achieve desirable classification results. RESULTS: In this paper, we propose a deep learning approach by combining a convolutional neural network (CNN) and bidirectional gated recurrent unit (BiGRU): our approach, DCGN, aims to achieve nonlinear dimensionality reduction and learn features to eliminate irrelevant factors in gene expression data. Specifically, DCGN first uses the synthetic minority oversampling technique algorithm to equalize data. The CNN can handle high-dimensional data without stress and extract important local features, and the BiGRU can analyse deep features and retain their important information; the DCGN captures key features by combining both neural networks to overcome the challenges of small sample sizes and sparse, high-dimensional features. In the experiments, we compared the DCGN to seven other cancer subtype classification methods using breast and bladder cancer gene expression datasets. The experimental results show that the DCGN performs better than the other seven methods and can provide more satisfactory classification results. Jiquan Shen, Haixia Zhai, Zhengjiang Wu, Chaokun Yan, Huimin Luo |
BMC Bioinform. | 7 |
| 2021 | A comprehensive review of scaffolding methods in genome assemblyabstractIn the field of genome assembly, scaffolding methods make it possible to obtain a more complete and contiguous reference genome, which is the cornerstone of genomic research. Scaffolding methods typically utilize the alignments between contigs and sequencing data (reads) to determine the orientation and order among contigs and to produce longer scaffolds, which are helpful for genomic downstream analysis. With the rapid development of high-throughput sequencing technologies, diverse types of reads have emerged over the past decade, especially in long-range sequencing, which have greatly enhanced the assembly quality of scaffolding methods. As the number of scaffolding methods increases, biology and bioinformatics researchers need to perform in-depth analyses of state-of-the-art scaffolding methods. In this article, we focus on the difficulties in scaffolding, the differences in characteristics among various kinds of reads, the methods by which current scaffolding methods address these difficulties, and future research opportunities. We hope this work will benefit the design of new scaffolding methods and the selection of appropriate scaffolding methods for specific biological studies. Yawei Wei, Mengna Lyu, Zhengjiang Wu, Huimin Luo, Chaokun Yan |
Briefings Bioinform. | 7 |
| 2021 | BreakNet: detecting deletions using long reads and a deep learning approachabstractBACKGROUND: Structural variations (SVs) occupy a prominent position in human genetic diversity, and deletions form an important type of SV that has been suggested to be associated with genetic diseases. Although various deletion calling methods based on long reads have been proposed, a new approach is still needed to mine features in long-read alignment information. Recently, deep learning has attracted much attention in genome analysis, and it is a promising technique for calling SVs. RESULTS: In this paper, we propose BreakNet, a deep learning method that detects deletions by using long reads. BreakNet first extracts feature matrices from long-read alignments. Second, it uses a time-distributed convolutional neural network (CNN) to integrate and map the feature matrices to feature vectors. Third, BreakNet employs a bidirectional long short-term memory (BLSTM) model to analyse the produced set of continuous feature vectors in both the forward and backward directions. Finally, a classification module determines whether a region refers to a deletion. On real long-read sequencing datasets, we demonstrate that BreakNet outperforms Sniffles, SVIM and cuteSV in terms of their F1 scores. The source code for the proposed method is available from GitHub at https://github.com/luojunwei/BreakNet . CONCLUSIONS: Our work shows that deep learning can be combined with long reads to call deletions more effectively than existing methods. Hongyu Ding, Jiquan Shen, Haixia Zhai, Zhengjiang Wu, Chaokun Yan, Huimin Luo |
BMC Bioinform. | 6 |
| 2020 | Correction to: SLR: a scaffolding algorithm based on long reads and contig classificationabstractFollowing publication of the original article [1], the author reported that there is an error in the original article. Mengna Lyu, Ranran Chen, Huimin Luo, Chaokun Yan |
BMC Bioinform. | 6 |
| 2020 | Quality-guided lane detection by deeply modeling sophisticated traffic context
Chaokun Yan |
Signal Process. Image Commun. | 2 |
| 2019 | SLR: a scaffolding algorithm based on long reads and contig classificationabstractBACKGROUND: Scaffolding is an important step in genome assembly that orders and orients the contigs produced by assemblers. However, repetitive regions in contigs usually prevent scaffolding from producing accurate results. How to solve the problem of repetitive regions has received a great deal of attention. In the past few years, long reads sequenced by third-generation sequencing technologies (Pacific Biosciences and Oxford Nanopore) have been demonstrated to be useful for sequencing repetitive regions in genomes. Although some stand-alone scaffolding algorithms based on long reads have been presented, scaffolding still requires a new strategy to take full advantage of the characteristics of long reads. RESULTS: Here, we present a new scaffolding algorithm based on long reads and contig classification (SLR). Through the alignment information of long reads and contigs, SLR classifies the contigs into unique contigs and ambiguous contigs for addressing the problem of repetitive regions. Next, SLR uses only unique contigs to produce draft scaffolds. Then, SLR inserts the ambiguous contigs into the draft scaffolds and produces the final scaffolds. We compare SLR to three popular scaffolding tools by using long read datasets sequenced with Pacific Biosciences and Oxford Nanopore technologies. The experimental results show that SLR can produce better results in terms of accuracy and completeness. The open-source code of SLR is available at https://github.com/luojunwei/SLR. CONCLUSION: In this paper, we describes SLR, which is designed to scaffold contigs using long reads. We conclude that SLR can improve the completeness of genome assembly. Mengna Lyu, Ranran Chen, Huimin Luo, Chaokun Yan |
BMC Bioinform. | 6 |
| 2019 | Application research of image compression and wireless network traffic video streaming
Chaokun Yan |
J. Vis. Commun. Image Represent. | 3 |
| 2011 | A Deadline Satisfaction Enhanced Workflow Scheduling AlgorithmabstractMeeting users' deadline constraint is usually the most important goal of workflow scheduling in Grid environment. In order to consider the dynamism of Grid resource, we adopted a stochastic model to describe dynamic workloads of Grid resources. A concept called Deadline Satisfaction Degree of Workflow (DSDW) was defined to represent the probability that a workflow could be completed before its deadline. We calculated task execution priorities based on their precedence relations in the workflow, then determined the candidate resource for each task so as to maximize DSDW, finally converted distribution problem of overall workflow deadline into a nonlinear programming problem with constraints and resolved it with known solutions. A Deadline Satisfaction Enhanced Scheduling Algorithm for Workflow (DSESAW) involving deadline distribution and resource selection was presented. The extensive simulation experiments using a practical medical image analysis application was conducted to verify our algorithm. Experimental results indicated that our algorithm could adapt to dynamic Grid environment and provide a good guarantee for user's deadline requirements. Zhigang Hu 0001, Chaokun Yan |
PDP | 3 |