Fang Du

dblp:14/2676 · DBLP profile ↗
← Back
29ranked-venue papers
7as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Metric-aware multi-objective deep reinforcement learning for database knob tuning
Chuitian Rong, Jitai Li, Chunbin Lin, Fang Du, Wei Lu 0015
Future Gener. Comput. Syst.4
2025 Enhancing Few Shot Named Entity Recognition via Label Semantic Description and Diversity Text
abstract
Large Language Models (LLMs) have demonstrated remarkable few-shot learning capabilities on Named Entity Recognition (NER) tasks, particularly through prompt-based approaches that avoid additional fine-tuning. However, despite the powerful generative and reasoning abilities of LLMs, two critical challenges remain: (1) semantic discrepancy of the same label across different datasets, which leads to recognition errors when using general labels to guide model outputs, and (2) contextual homogeneity in in-context examples, which limits the model’s ability to distinguish fine-grained entity types during inference. To address these challenges, we propose LSDNER, a dual-faceted prompt construction strategy that integrates structured label semantic descriptions and promotes context diversity. Specifically, to tackle label semantic inconsistency, we introduce a structured framework that organizes label semantic descriptions into definitions, attributes, relational features, and behavioral characteristics. This representation enables LLMs to better understand the dataset-specific semantic meanings of entity labels. To address contextual monotony, we devise a diversity-driven sampling strategy for selecting in-context demonstrations, thereby expanding semantic coverage and promoting reasoning capabilities. Experiments on three general and four domain-specific NER datasets demonstrate that our approach surpasses prompt-based methods and achieves competitive results with supervised baselines. Our code is available at: https://github.com/hui68633/LSDNER.
Hui Wang 0170, Fang Du
ECAI2
2025 GHS-VDG:Graph and Hybrid Spatio-Temporal Attention for Video Diffusion Generation
abstract
Video generation is a core task in computer vision, its key challenge is the effective modeling of the temporal and spatial dependencies between video frames. This paper proposes a diffusion model-based video generation framework, GHS-VDG, which introduces the Video Frame Temporal Graph Aggregator (VFTGA) and Hybrid Spatio-Temporal Attention (HSTA) modules to capture complex spatiotemporal dependencies between video frames, thereby enhancing the quality and accuracy of predicted frames. Specifically, in the VFTGA module, we innovatively represent video frames as graph nodes and construct edge connections according to their temporal order, which effectively enhances the capability of capturing temporal information. The HSTA module combines temporal and spatial attention mechanisms to optimize the learning of spatiotemporal features. Experimental results demonstrate that GHS-VDG performs excellently on datasets such as SMMNIST, BAIR, KTH, and UCF101, validating its effectiveness in video generation tasks.
Gao Xinyu, Fang Du, Song Lijuan, Zhang Xu
ICIP2
2025 scDCT: a conditional diffusion-based deep learning model for high-fidelity single-cell cross-modality translation
abstract
Single-cell multi-omics technologies enable comprehensive molecular profiling, offering insights into cellular heterogeneity and biological mechanisms. However, current cross-modality translation methods struggle with high-dimensional, noisy, and sparse single-cell data. We propose single-cell Diffusion models for Cross-modality Translation (scDCT), a probabilistic framework for bidirectional cross-modality translation in single-cell data, including single-cell RNA sequencing, single-cell assay for transposase-accessible chromatin sequencing, and protein expression. scDCT integrates modality-specific autoencoders with conditional denoising diffusion probabilistic models to map inputs to latent spaces and perform probabilistic translation across modalities. This design captures cell-type heterogeneity, accounts for data sparsity, and models uncertainty during translation. Extensive experiments on eight benchmark datasets demonstrate that scDCT outperforms state-of-the-art methods across paired, unpaired, cross-type, and cross-tissue settings, offering a robust and interpretable solution for single-cell multi-omics integration.
Junlei Zhou, Jialiang Xue, Furui Liu, Fang Du, Zhenhua Yu 0002
Briefings Bioinform.5
2025 scCMP: A Deep Learning Method for Identifying Clonal Mutational Profiles From Single-Cell Genomic Data
abstract
Accurately inferring clonal mutational profiles is essential for understanding intra-tumor heterogeneity and clonal selection during tumor evolution. Single-cell multi-modal genomic data, such as copy numbers and point mutations, can be integrated to deliver multiple views of the clonal mutational patterns. Despite of the fact that integration of single-cell multi-modal data has been extensively explored in existing studies, computational methods specifically developed to integrate copy number and point mutation data of single cells are still highly needed. We introduce a deep joint representation learning framework called scCMP, to accurately identify clonal mutational profiles. scCMP employs hybrid Transformer-CNN architectures and graph convolutional networks to integrate single-cell copy number and point mutation data. By fusing individual and commonality information among the two modalities, it generates meaningful cell embeddings for identifying clonal clusters. We comprehensively evaluate the effectiveness of scCMP on five real single-cell DNA sequencing datasets, and further showcase its good scalability on datasets generated from other omics technologies. The results show scCMP accurately aggregates the cells with similar mutational profiles into a same cluster, and surpasses the state-of-the-art methods, indicating its advantage in integrating single-cell genomic data.
Junlei Zhou, Fangyuan Shi, Xianhao Huo, Fang Du, Zhenhua Yu 0002
IEEE Trans. Comput. Biol. Bioinform.5
2024 CoT: a transformer-based method for inferring tumor clonal copy number substructure from scDNA-seq data
abstract
Single-cell DNA sequencing (scDNA-seq) has been an effective means to unscramble intra-tumor heterogeneity, while joint inference of tumor clones and their respective copy number profiles remains a challenging task due to the noisy nature of scDNA-seq data. We introduce a new bioinformatics method called CoT for deciphering clonal copy number substructure. The backbone of CoT is a Copy number Transformer autoencoder that leverages multi-head attention mechanism to explore correlations between different genomic regions, and thus capture global features to create latent embeddings for the cells. CoT makes it convenient to first infer cell subpopulations based on the learned embeddings, and then estimate single-cell copy numbers through joint analysis of read counts data for the cells belonging to the same cluster. This exploitation of clonal substructure information in copy number analysis helps to alleviate the effect of read counts non-uniformity, and yield robust estimations of the tumor copy numbers. Performance evaluation on synthetic and real datasets showcases that CoT outperforms the state of the arts, and is highly useful for deciphering clonal copy number substructure.
Furui Liu, Fangyuan Shi, Fang Du, Xiangmei Cao, Zhenhua Yu 0002
Briefings Bioinform.3
2024 Convolutional neural network incorporating misclassification information for image recognition
Junying Hu, Rongrong Fei, Fang Du, Peiju Chang, Jiangshe Zhang 0001
Soft Comput.3
2023 rcCAE: a convolutional autoencoder method for detecting intra-tumor heterogeneity and single-cell copy number alterations
abstract
Intra-tumor heterogeneity (ITH) is one of the major confounding factors that result in cancer relapse, and deciphering ITH is essential for personalized therapy. Single-cell DNA sequencing (scDNA-seq) now enables profiling of single-cell copy number alterations (CNAs) and thus aids in high-resolution inference of ITH. Here, we introduce an integrated framework called rcCAE to accurately infer cell subpopulations and single-cell CNAs from scDNA-seq data. A convolutional autoencoder (CAE) is employed in rcCAE to learn latent representation of the cells as well as distill copy number information from noisy read counts data. This unsupervised representation learning via the CAE model makes it convenient to accurately cluster cells over the low-dimensional latent space, and detect single-cell CNAs from enhanced read counts data. Extensive performance evaluations on simulated datasets show that rcCAE outperforms the existing CNA calling methods, and is highly effective in inferring clonal architecture. Furthermore, evaluations of rcCAE on two real datasets demonstrate that it is able to provide a more refined clonal structure, of which some details are lost in clonal inference based on integer copy numbers.
Zhenhua Yu 0002, Furui Liu, Fangyuan Shi, Fang Du
Briefings Bioinform.4
2023 Analytical Study of the Changes in Brightness Temperature Based on the Tectonic Field Associated With Three Earthquakes in the Eastern Tibetan Plateau
abstract
The thermal infrared brightness temperature (BT) of the eastern Tibetan Plateau (TP) was retrieved from the Moderate Resolution Imaging Spectroradiometer (MODIS) level-1B data. The multiyear averaged BT background field was subtracted from the punctual BT data to yield monthly BT spatial anomaly, and calculated time series of BT for the secondary blocks. Then, the spatial and temporal changes in the BT of the study area before the Menyuan M6.4, Zaduo M6.2, and Jiuzhaigou M7.0 earthquakes were investigated and analyzed based on the tectonic setting. The results show the following. The spatial BT radiation enhancement frequency rose remarkably before strong earthquakes; each of the three earthquakes was preceded by marked spatiotemporal continuous BT anomalies. The tectonic setting significantly influences the BT anomaly feature. The spatial BT anomaly was not notable in the Qaidam and Qilian block before the Menyuan earthquake; the spatial BT anomaly mainly appeared in the Qiangtang and Bayan Har blocks before the Zaduo and Jiuzhaigou earthquakes. The Qiangtang and Bayan Har block’s BT time series curves have similar features. The Qaidam and Qilian block’s BT time series curves have analogous shapes. The three earthquakes may be regarded as one seismic event induced by a stage of tectonic stress enhancement rather than three independent occasions. The spatial BT anomalous behavior before earthquakes is, to a great extent, like the result of the rock stress loading experiment; the rock compression and the lithosphere-atmosphere-ionosphere coupling (LAIC) may be the main reasons for the intensification of the BT radiation.
Tiebao Zhang, Fang Du, Feng Long
IEEE Trans. Geosci. Remote. Sens.4
2022 AMC: accurate mutation clustering from single-cell DNA sequencing data
abstract
SUMMARY: Single-cell DNA sequencing (scDNA-seq) now enables high-resolution profiles of intra-tumor heterogeneity. Existing methods for phylogenetic inference from scDNA-seq data perform acceptably well on small datasets but suffer from low computational efficiency and/or degraded accuracy on large datasets. Motivated by the fact that mutations sharing common states over single cells can be grouped together, we introduce a new software called AMC (accurate mutation clustering) to accurately cluster mutations, thus improve the efficiency of phylogenetic inference. AMC first employs principal component analysis followed by K-means clustering to find mutation clusters, then infers the maximum likelihood estimates of the genotypes of each cluster. The inferred genotypes can subsequently be used to reconstruct the phylogenetic tree with high efficiency. Comprehensive evaluations on various simulated datasets demonstrate AMC is particularly useful to efficiently reason the mutation clusters on large scDNA-seq datasets. AVAILABILITY AND IMPLEMENTATION: AMC is freely available at https://github.com/qasimyu/amc. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhenhua Yu 0002, Fang Du
Bioinform.2
2022 SCViT: A Spatial-Channel Feature Preserving Vision Transformer for Remote Sensing Image Scene Classification
abstract
Convolutional neural network (CNN)-based methods are widely used in remote sensing image scene classification and can obtain excellent performances. However, the stacked receptive fields in the CNN-based methods have limitations in modeling the long-range dependencies of local features. The vision transformer (ViT) model provides a good solution as it directly considers the global interactions of local patches by the self-attention mechanism. However, the vanilla ViT model, which simply splits images into fixed-size patches treated as tokens, mainly considers the global information in the spatial domain. In this article, a spatial-channel feature preserving ViT (SCViT) model is proposed, which considers both the detailed geometric information of the high-spatial-resolution (HSR) imagery and the contribution of the different channels contained in the classification token. First, in the proposed method, tokens are generated by progressively aggregating the neighboring overlapping patches to extract the local structural features of the imagery. Second, a multihead self-attention (MSA) mechanism is used to model the global interactions of the tokens in the encoder. A lightweight channel attention (LCA) module is then introduced to consider the importance of the different channels in the classification token. Finally, a multilayer perceptron (MLP) is used to acquire the final results. Compared with the state-of-the-art scene classification methods, the experimental results confirm the potential of using ViT models in remote sensing image scene classification.
Pengyuan Lv, Yanfei Zhong, Fang Du, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4
2021 Enhanced Bayesian detection for copy number alterations from next-generation sequencing data
abstract
Next-generation sequencing (NGS) promises highresolution landscapes of cancer genomes, especially the genomewide copy number alterations (CNAs). Detecting CNAs from tumor NGS data often encounters several critical issues that heavily affect the accuracy of the results. For instance, the over-dispersed distribution of read depth signals gives rise to a challenge of designing appropriate statistical model to effectively explain the data; aneuploidy of tumor genome often causes ambiguity in interpreting copy number status for a genomic fragment. We introduce a new bioinformatics tool called EBHMM to accurately infer CNAs from single tumor sample. EBHMM employs a novel hidden Markov model (HMM) to jointly analyze read depth and read counts signals derived from tumor NGS data. To improve CNA detection robustness against severe signal fluctuation and distribution shift of read counts induced by aneuploidy, EBHMM exploits more appropriate emission models in the HMM and prior knowledge to enable accurate inference. Experimental results on real data suggest the proposed method is highly effective in detecting CNAs as well as genome ploidy, and achieves comparable performance to the state-of-the-art methods. The EBHMM software is freely available at https://github.com/qasimyu/ebhmm.
Zhenhua Yu 0002, Fang Du
BIBM2
2020 Network Architecture Reasoning Via Deep Deterministic Policy Gradient
abstract
In this paper, we introduce global compression learning (GCL) for finding reduced network architecture from a pre-trained network by removing both intra-layer and inter-layer structural redundancy. To accomplish this, we first derive architecture features from a binary representation of the network structure that effectively characterize the relationships between different layers. We then leverage reinforcement learning to iteratively compress the network via deep deterministic policy gradient based on the learned architecture features. To void extensive exploration of the huge space of network architectures, we bound feasible solutions within a small subspace by following a strict accuracy loss tolerance. Benchmarking tests show GCL outperforms the state-of-the-art models. On CIFAR-10 dataset, our model reduces 60.5% FLOPs and 93.3% parameters on VGG-16 without hurting the network accuracy, and yields a significantly compressed architecture for ResNet-110 by reductions of 71.92% FLOPs and 79.62% parameters with the cost of only 0.11% accuracy loss.
Huidong Liu, Fang Du, Xiaofen Tang, Hao Liu 0019, Zhenhua Yu 0002
ICME2
2020 SCSsim: an integrated tool for simulating single-cell genome sequencing data
abstract
MOTIVATION: Allele dropout (ADO) and unbalanced amplification of alleles are main technical issues of single-cell sequencing (SCS), and effectively emulating these issues is necessary for reliably benchmarking SCS-based bioinformatics tools. Unfortunately, currently available sequencing simulators are free of whole-genome amplification involved in SCS technique and therefore not suited for generating SCS datasets. We develop a new software package (SCSsim) that can efficiently simulate SCS datasets in a parallel fashion with minimal user intervention. SCSsim first constructs the genome sequence of single cell by mimicking a complement of genomic variations under user-controlled manner, and then amplifies the genome according to MALBAC technique and finally yields sequencing reads from the amplified products based on inferred sequencing profiles. Comprehensive evaluation in simulating different ADO rates, variation detection efficiency and genome coverage demonstrates that SCSsim is a very useful tool in mimicking single-cell sequencing data with high efficiency. AVAILABILITY AND IMPLEMENTATION: SCSsim is freely available at https://github.com/qasimyu/scssim. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhenhua Yu 0002, Fang Du, Xuehong Sun, Ao Li 0001
Bioinform.2
2020 SimuSCoP: reliably simulate Illumina sequencing data based on position and context dependent profiles
abstract
BACKGROUND: A number of simulators have been developed for emulating next-generation sequencing data by incorporating known errors such as base substitutions and indels. However, their practicality may be degraded by functional and runtime limitations. Particularly, the positional and genomic contextual information is not effectively utilized for reliably characterizing base substitution patterns, as well as the positional and contextual difference of Phred quality scores is not fully investigated. Thus, a more effective and efficient bioinformatics tool is sorely required. RESULTS: Here, we introduce a novel tool, SimuSCoP, to reliably emulate complex DNA sequencing data. The base substitution patterns and the statistical behavior of quality scores in Illumina sequencing data are fully explored and integrated into the simulation model for reliably emulating datasets for different applications. In addition, an integrated and easy-to-use pipeline is employed in SimuSCoP to facilitate end-to-end simulation of complex samples, and high runtime efficiency is achieved by implementing the tool to run in multithreading with low memory consumption. These features enable SimuSCoP to gets substantial improvements in reliability, functionality, practicality and runtime efficiency. The tool is comprehensively evaluated in multiple aspects including consistency of profiles, simulation of genomic variations and complex tumor samples, and the results demonstrate the advantages of SimuSCoP over existing tools. CONCLUSIONS: SimuSCoP, a new bioinformatics tool is developed to learn informative profiles from real sequencing data and reliably mimic complex data by introducing various genomic variations. We believe that the presented work will catalyse new development of downstream bioinformatics methods for analyzing sequencing data.
Zhenhua Yu 0002, Fang Du, Rongjun Ban, Yuanwei Zhang
BMC Bioinform.2
2020 Weighted-capsule routing via a fuzzy gaussian model
Ouafa Amira, Fang Du, Jiangshe Zhang 0001, Chunxia Zhang 0002, Rafik Hamza
Pattern Recognit. Lett.3
2019 Discriminative multi-modal deep generative models
Fang Du, Jiangshe Zhang 0001, Junying Hu, Rongrong Fei
Knowl. Based Syst.1
2019 Convolutional Sparse Representation of Injected Details for Pansharpening
abstract
In this letter, we address the pansharpening problem, which focuses on constructing a high-resolution (HR) multispectral (MS) image from a low-resolution (LR) MS and an HR panchromatic (Pan) image. The accuracy of pansharpening method based on sparse representation (SR) mainly depends on the construction of dictionary and the learning of sparse coefficients, while the details injection (DI)-based pansharpening method sharpens the MS bands by adding the proper spatial details from Pan. The combination of SR and DI has been put forward as the pansharpening method based on SR of injected details (SR-D). However, limited to the patch-based manner, pansharpening with traditional SR model faces two disadvantages, i.e., limited ability in detail preservation and high sensitivity to misregistration. In this letter, we replace the traditional SR model with convolutional SR (CSR) as a global SR model in the SR-D method and propose a new pansharpening method called CSR of injected details (CSR-D) to overcome the above-mentioned two drawbacks. Experimental results on the IKONOS and WorldView2 data sets show that the proposed method can achieve remarkable spectral and spatial quality on both reduced scale and full scale.
Rongrong Fei, Jiangshe Zhang 0001, Junmin Liu, Fang Du, Peiju Chang, Junying Hu
IEEE Geosci. Remote. Sens. Lett.4
2019 Discriminative Representation Learning with Supervised Auto-encoder
Fang Du, Jiangshe Zhang 0001, Nannan Ji, Junying Hu, Chunxia Zhang 0002
Neural Process. Lett.1
2018 An effective hierarchical extreme learning machine based multimodal fusion framework
Fang Du, Jiangshe Zhang 0001, Nannan Ji, Chunxia Zhang 0002
Neurocomputing1
2018 Symmetric low-rank representation with adaptive distance penalty for semi-supervised learning
Changpeng Wang, Jiangshe Zhang 0001, Fang Du
Neurocomputing3
2016 Drug target path discovery on semantic biomedical big data
abstract
Systems chemical biology integrate chemistry, biology and computation tools as a whole system, which can help researchers to deeply study the interaction and relationship among small molecules, such as genes, proteins, targets, compounds and so on. With systems chemical biology, researchers can concentrate on new way of drug discovery, including drug target path discovery, which can not only help biomedical researchers to find evidences for existing disease associate genes, but also to design new effect medicine based on targets. Network based approaches are the state-of-art solutions for drug target path discovery, however, there are still some challenges: 1) The quality of the network dominate the efficiency and accuracy of the results, therefore a well designed network is quite important on drug target path discovery mission; 2) the existing network based approaches only work on small graph, it can not handle massive data well. In the paper, we designed a novel framework of systems chemical biology based on semantic big data. In the paper, we proposed a novel drug target path discovery approach. It can identify targets associated with specific medicines (disease) and the path of relationship based on a RDF semantic D-T network. The ranking of candidate targets is performed through an improved parallel random walk with restart algorithm. The experimental studies show that the proposed approaches can efficiently discover drug target relationship path, meanwhile, the approaches have good scalability which are suitable for big data analysis.
Fang Du, Yingjie Shi, Lijuan Song, Xiaojun Gu
IEEE BigData1
2013 Linking Entities in Unstructured Texts with RDF Knowledge Bases
Fang Du, Yueguo Chen, Xiaoyong Du 0001
APWeb1
2012 Efficient SPARQL Query Processing in MapReduce through Data Partitioning and Indexing
Zhi Nie, Fang Du, Yueguo Chen, Xiaoyong Du 0001, Linhao Xu
APWeb2
2012 Partitioned Indexes for Entity Search over RDF Knowledge Bases
Fang Du, Yueguo Chen, Xiaoyong Du 0001
DASFAA (1)1
2012 Functional mapping of ontogeny in flowering plants
abstract
All organisms face the problem of how to perform a sequence of developmental changes and transitions during ontogeny. We revise functional mapping, a statistical model originally derived to map genes that determine developmental dynamics, to take into account the entire process of ontogenetic growth from embryo to adult and from the vegetative to reproductive phase. The revised model provides a framework that reconciles the genetic architecture of development at different stages and elucidates a comprehensive picture of the genetic control mechanisms of growth that change gradually from a simple to a more complex level. We use an annual flowering plant, as an example, to demonstrate our model by which to map genes and their interactions involved in embryo and postembryonic growth. The model provides a useful tool to study the genetic control of ontogenetic growth in flowering plants and any other organisms through proper modifications based on their biological characteristics.
Xiyang Zhao, Chunfa Tong, Xiaoming Pang, Zhong Wang 0001, Yunqian Guo, Fang Du, Rongling Wu
Briefings Bioinform.6
2011 A Framework to Model the Topological Structure of Supply Networks
abstract
Topological structure is considered more and more important in managing a supply network or predicting its development. In this paper, a new framework is proposed to model the topological structure of supply networks, where different types of supply networks can be created just by introducing different supplier-customer connecting rules. Generally, the networks created in the framework are much different from the random networks with the same degree sequences. The revealed phenomenon suggests that real-world supply networks may benefit from its intrinsic mechanism on flexibility, efficiency, and robustness to target attacks. Note to Practitioners-The topological structure of supply networks is considered more and more important in managing a supply network or predicting its development. In this paper, we introduce a framework to model and analyze the topological structure of supply networks. This work aims to characterize supply networks by statistical methods and can help researchers better understand the material dynamics on supply networks and further conveniently create their own supply networks by summarizing practical supplier-customer connecting rules or analyzing real-world supply network data. The work should be further expanded in other aspects, such as simulating material dynamics on supply networks, designing optimal structure by introducing proper supplier-customer connecting rules, rearranging local connections to enhance the competi tiveness and further ensure the long-term benefit of a target firm, and so on, all of which are of much interest for governors, investors, and managers and can be studied in the present framework in the future.
Qi Xuan 0001, Fang Du, Tie-Jun Wu
IEEE Trans Autom. Sci. Eng.2
2004 ShreX: Managing XML Documents in Relational Databases
Fang Du, Sihem Amer-Yahia, Juliana Freire
VLDB1
2002 Achieving a More Robust Neural Network Model for Control of a MR Damper by Signal Sensitivity Analysis
Chih-Chen Chang, Fang Du
Neural Comput. Appl.3