EDBT 2026 Demo / reviewers in the wild / expert
Jiaquan Gao
dblp:78/2677
· DBLP profile ↗
31ranked-venue papers
11as first author
21since 2021 · last 2026
0000-0002-2983-9921ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 9 since 2021Systems, architecture and hardware · 13 · 8 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GFSAI: An Adaptive Factorized Sparse Approximate Inverse Preconditioning Algorithm on GPUabstractABSTRACT The factorized sparse approximate inverse (FSAI) preconditioner has been proven to be effective in accelerating the convergence of iterative methods. Due to the high cost of constructing the FSAI preconditioner, accelerating it on graphics processing unit (GPU) has attracted considerable attention. However, despite the development of some existing FSAI preconditioning algorithms on GPU, their performance will significantly decrease when they encounter matrix types that are not suitable for them. This motivates us to investigate how to design an effective FSAI preconditioning algorithm on GPU. In this paper, we propose an adaptive FSAI preconditioning algorithm on GPU, called GFSAI‐Adaptive, to address the above problem. In GFSAI‐Adaptive, first, two adaptive thread allocation strategies are proposed for two special types of SPD matrices to ensure that the allocated threads can be fully utilized. Second, based on the proposed two thread allocation strategies, two FSAI kernels, called GFSAII and GFSAIII, are presented. Third, we construct a new graph convolutional network, and thus propose a search engine to select the optimal kernel from GFSAII and GFSAIII for matrices that do not belong to two special types based on it. Experimental results show that our proposed GFSAI‐Adaptive is effective and outperforms a popular preconditioning algorithm in the public CUSPARSE library and a recent parallel static FSAI preconditioning algorithm on GPU. Yige Zhang, Jiaquan Gao |
Concurr. Comput. Pract. Exp. | 3 |
| 2026 | Unsupervised pattern image retrieval via dual-encoder architecture with multi-head attention
Ru Han, Chunming Guan, Jiaquan Gao, Ying Li 0016 |
Neurocomputing | 4 |
| 2026 | RA-GCN: Residual attention based graph convolutional network for multi-label pattern image retrieval
Ying Li 0016, Longye Du, Chunming Guan, Jiaquan Gao |
Pattern Recognit. | 5 |
| 2025 | An improved approximation algorithm for Hypergraph Max p-Section
Guangfeng Li, Jian Sun 0022, Jiaquan Gao, Zhiren Sun, Xiaoyan Zhang 0001 |
Discret. Appl. Math. | 3 |
| 2025 | Federated learning-based private medical knowledge graph for epidemic surveillance in internet of thingsabstractAbstract With the explosive development of the Internet of Things (IoT), it is convenient and important to collect health data from medical sensors and smart devices and construct medical knowledge graph. The knowledge graph contributes to investigating the connection between patient and disease, especially for epidemic surveillance. However, it is possible to cause the leakage of sensitive health information due to the untrusted data collector or various malicious attackers. In this paper, we attempt to utilise federated learning to construct a special knowledge graph, that is, individual‐symptom relationship diagram with local differential privacy (LDP‐ISRD), for epidemic risk surveillance, which presents the underlying infectious relationship among individuals. At first, we propose a federated learning‐based framework of LDP‐ISRD by utilising individuals' smart devices in IoT. Then, we leverage locations to determine the connection among individuals in terms of physical contact. Next, we propose a randomised algorithm PrivISRD to implement federated learning‐based LDP‐ISRD, which consists of symptom perturbation and aggregation. Finally, extensive experiments evaluate the impact of various parameters and results demonstrate that LDP‐ISRD has good performance. Xiaotong Wu, Jiaquan Gao, Muhammad Bilal 0003, Fei Dai 0002, Xiaolong Xu 0001, Lianyong Qi, Wan-Chun Dou |
Expert Syst. J. Knowl. Eng. | 2 |
| 2025 | HTS-LB: Hypergraph tree search for learning branch
Yige Zhang, Ying Li 0016, Jiaquan Gao |
Neural Networks | 5 |
| 2025 | Complementary two-branch Transformer for multi-label image retrieval
Ying Li 0016, Shuaiyu Deng, Chunming Guan, Jiaquan Gao |
Pattern Recognit. | 4 |
| 2025 | Acceleration of Timing-Aware Gate-Level Logic Simulation Through One-Pass GPU ParallelismabstractWitnessing the advancements in the scale and complexity of chip design, along with the benefits from high-performance computing technologies, the simulation of Very Large Scale Integration (VLSI) circuits increasingly demands acceleration through parallel computing with GPU devices. However, conventional parallel strategies fail to fully leverage modern GPU capabilities, introducing new challenges in GPU-based parallelism for VLSI simulations despite previous demonstrations of significant acceleration. In this paper, we propose a novel approach for accelerating the simulation of 4-value logic timing-aware gate-level circuits through waveform-based GPU parallelism. Our approach introduces an innovative strategy that effectively manages task dependencies during the parallelism of combinational circuits, significantly reducing the synchronization requirement between CPU and GPU. The proposed approach achieves one-pass parallelism by requiring only a single round of data transfer. Moreover, to address the implementation challenges associated with our strategy on GPU devices, we have developed and optimized a series of data structures that dynamically allocate and store newly generated outputs of uncertain scale. Finally, we conduct experiments on industrial-scale open-source benchmarks to demonstrate our approach’s performance gains over several state-of-the-art baselines. Weijie Fang, Yanggeng Fu, Jiaquan Gao, Longkun Guo, Gregory Z. Gutin, Xiaoyan Zhang 0001 |
IEEE Trans. Computers | 3 |
| 2025 | Energy-Efficiency Oriented Distributed Heterogeneous Hybrid Flow Shop Scheduling With Multilevelled Mixed-Model AssemblyabstractThis article studies an energy-efficient scheduling problem in a two-stage manufacturing system with distributed heterogeneous hybrid flow shops and mixed-model assembly lines (EDHHFSP-MMAL). A mixed-integer linear programming model is proposed that simultaneously optimizes total tardiness and energy consumption (including operational, idle, and common energy components). To solve this multiobjective problem, a learning competitive swarm optimizer (LCSO) is proposed that integrates two novel mechanisms: 1) environmental-competitive learning through probability models capturing product-task relationships and 2) comprehensive learning utilizing reinforcement learning to guide local search based on nondominated solution states. The hybrid approach balances convergence speed and solution diversity by combining solution-space and policy-space learning perspectives. Experimental results demonstrate LCSO’s superior performance over compared methods, achieving 25% improvement in energy-time tradeoff compared to other state-of-the-art multiobjective optimizers in solving related problems. The proposed method particularly excels in optimizing complex energy-time tradeoffs while maintaining better solution diversity and convergence across different problem scales. Weishi Shao, Zhongshi Shao, Dechang Pi, Jiaquan Gao |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2023 | TsP-Tran: Two-Stage Pure Transformer for Multi-Label Image RetrievalabstractImage retrieval aims to find similar images given the query. Most of existing retrieval works are based on the pre-trained model of single-label image classification. In practice, the query usually contains more than one instance, and the single label is far from enough for fully depicting the attributes of an open-world image. Due to the complicated similarity relationships between multiple semantics, the multi-label image retrieval task is not so well solved as the single-label task. In this work, we propose a two-stage pure Transformer model for multi-label image retrieval, which leverages a Transformer encoder to exploit the complex dependencies among visual features and labels. Except for the Transformer encoder, the image feature embedding module is also based on Transformer, so that the optimal model weights could be learned in an end-to-end manner. To be specific, inputs of the Transformer encoder mainly consist of an Vision Transformer branch and a label embedding branch, which generates suitable image features and label descriptions, respectively. Given an input set of visual features and text labels, the developed Transformer encoder could be accordingly optimized in the training stage with compressed multi-label output layer. In order to obtain sufficient outputs to accurately find images containing similar semantics with the query from the database, we adjust the network by removing the last fully connected layer in the retrieval stage. Specially, images and labels are used for the training stage in a randomly masked manner to enhance the model performance, and no labels are visible in the content-based image retrieval stage. Comprehensive experiments are performed on three multi-label datasets including MS-COCO, NUS-WIDE and VOC2007, demonstrating promising results of our proposed method against the state-of-the-arts for multi-label image retrieval. Ying Li 0016, Chunming Guan, Jiaquan Gao |
ICMR | 3 |
| 2023 | Tran-GCN: Multi-label Pattern Image Retrieval via Transformer Driven Graph Convolutional NetworkabstractPattern images are artificially designed images that possess distinctiveness in their elements, styles, and arrangements. With the ever-growing number of pattern images, pattern image retrieval emerges as a promising technique with significant potential for commercial and industrial applications, such as fashion and home decoration, facilitating rapid identification of preferred print patterns by users. The main purpose of multi-label pattern image retrieval is to effectively represent and match images with their corresponding labels. Compared to conventional image retrieval, multi-label pattern image retrieval faces greater challenges due to the richer semantic information contained within the abstract print patterns and the complex relationships between multiple labels. To tackle these challenges, we propose a model specifically designed for multi-label pattern image retrieval, called Tran-GCN. Our proposed model is built upon a Transformer-based autoregressive architecture, which leverages image information to guide the exploration of correlations between different labels through the textual modality. By utilizing this correlation information, we construct a graph convolutional network (GCN) model to further enhance the correlations between image and label representations. To be more specific, our Tran-GCN model utilizes a cross-modal attention mechanisms at each layer to effectively aggregate visual features from the input image and update label semantics through residual connections. The GCN module is updated based on the correlation between textual features, as represented in a relationship matrix. Extensive experiments on two widely used public visual benchmarks, MS-COCO and NUS-WIDE, as well as a multi-label pattern image dataset, Pattern 2, consistently demonstrate the ability of our proposed Tran-GCN model for general use and its superior performance in multi-label pattern image retrieval tasks as well. Ying Li 0016, Chunming Guan, Erwan Ye, Ding Yuxiang, Jiaquan Gao |
ACM Multimedia | 6 |
| 2023 | HeuriSPAI: a heuristic sparse approximate inverse preconditioning algorithm on GPU
Jiaquan Gao, Xinyue Chu |
CCF Trans. High Perform. Comput. | 1 |
| 2022 | An ensemble of random decision trees with local differential privacy in edge computing
Xiaotong Wu, Lianyong Qi, Jiaquan Gao, Genlin Ji, Xiaolong Xu 0001 |
Neurocomputing | 3 |
| 2022 | Parallel Dynamic Sparse Approximate Inverse Preconditioning Algorithm on GPUabstractThe dynamic sparse approximate inverse (SPAI) preconditioner has proven to be effective in accelerating the convergence of iterative methods for large linear systems. Recently, accelerating it on graphics processing unit (GPU) has attracted considerable attention due to the fact that the cost of constructing the preconditioner is high. However, the existing parallel dynamic SPAI preconditioning algorithms on GPU are usually ineffective because of the out-of-memory error for large matrices. This motivates us to investigate how to accelerate the construction of dynamic SPAI preconditioners on GPU. In this article, we propose an efficient dynamic SPAI preconditioning algorithm on GPU, called GDSPAI. For our proposed GDSPAI, there are the following novelties: (1) a well-known dynamic SPAI preconditioning algorithm is substantially modified to address the main challenges of parallelization on GPU, (2) a parallel framework of constructing the dynamic SPAI preconditioner on GPU is presented on the basis of the modified dynamic SPAI preconditioning algorithm; and (3) each component of the preconditioner is computed in parallel inside a group of threads. Experimental results show that the proposed GDSPAI is effective for large matrices, and outperforms the popular preconditioning algorithms in three public libraries, as well as a recent parallel static SPAI preconditioning algorithm. Jiaquan Gao, Xinyue Chu, Xiaotong Wu, Jun Wang 0077, Guixia He |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | Adversarial Example Defense Based on the SupervisionabstractIn recent years, deep learning has developed rapidly and has shown great performance on many challenging machine learning tasks, such as image classification, natural language processing, and speech recognition. However, researchers have recently discovered that deep learning models have security risks and are easily affected by adversarial examples. The adversarial example is a sample formed by deliberately adding subtle perturbation that is invisible to the human in the dataset. It can make the classification classify incorrectly with a high degree of confidence, which poses more challenges for deep learning research. In this paper, we propose a defense model based on the supervision mechanism. The model adds supervision layers to the original convolutional neural network and improves the robustness and defense ability of the model by improving the loss function. The LeNet-5 and VGG networks are used as the original network models. The experimental results on MNIST and CIFAR-10 confirm that the method proposed in this paper will effectively increase the difficulty of the attackers. Ziyu Yao 0003, Jiaquan Gao |
IJCNN | 2 |
| 2021 | Multi-label Pattern Image Retrieval via Attention Mechanism Driven Graph Convolutional NetworkabstractPattern images are artificially designed images which are discriminative in aspects of elements, styles, arrangements and so on. Pattern images are widely used in fields like textile, clothing, art, fashion and graphic design. With the growth of image numbers, pattern image retrieval has great potential in commercial applications and industrial production. However, most of existing content-based image retrieval works mainly focus on describing simple attributes with clear conceptual boundaries, which are not suitable for pattern image retrieval. It is difficult to accurately represent and retrieve pattern images which include complex details and multiple elements. Therefore, in this paper, we collect a new pattern image dataset with multiple labels per image for the pattern image retrieval task. To extract discriminative semantic features of multi-label pattern images and construct high-level topology relationships between features, we further propose an Attention Mechanism Driven Graph Convolutional Network (AMD-GCN). Different layers of the multi-semantic attention module activate regions of interest corresponding to multiple labels, respectively. By embedding the learned labels from attention module into the graph convolutional network, which can capture the dependency of labels on the graph manifold, the AMD-GCN builds an end-to-end framework to extract high-level semantic features with label semantics and inner relationships for retrieval. Experiments on the pattern image dataset show that the proposed method highlights the relevant semantic regions of multiple labels, and achieves higher accuracy than state-of-the-art image retrieval methods. Ying Li 0016, Yeyu Yin, Jiaquan Gao |
ACM Multimedia | 4 |
| 2021 | Adversarial Examples Defense via Combining Data Transformations and RBF Layers
Jiaquan Gao, Xiaoxin Li 0001 |
PRICAI (2) | 2 |
| 2021 | A new diagonal storage for efficient implementation of sparse matrix-vector multiplication on graphics processing unitabstractSummary The sparse matrix–vector multiplication (SpMV) is of great importance in computational science. For multidiagonal sparse matrices that have many long zero sections or scatter points, a great number of zeros are filled to maintain the diagonal structure when using the popular DIA format to store them. This leads to the performance degradation of the DIA kernel. To alleviate the drawback of DIA, we present a novel diagonal storage format, called RBDCS (diagonal compressed storage based on row‐blocks), for multidiagonal sparse matrices, and thus propose an efficient SpMV kernel that corresponds to RBDCS. Given that the RBDCS kernel codes must be manually rewritten for different multidiagonal sparse matrices, a code generator is presented to automatically generate RBDCS kernel codes. Experimental results show that the proposed RBDCS kernel is effective, and outperforms HYBMV in the CUSPARSE library, and three popular diagonal SpMV kernels: DIA, HDI, and CRSD. Guixia He, Jiaquan Gao |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | A feature-based intelligent deduplication compression system with extreme resemblance detectionabstractWith the fast development of various computing paradigms, the amount of data is rapidly increasing that brings the huge storage overhead. However, the existing data deduplication techniques do not make full use of similarity detection to improve the storage efficiency and data transmission rate. In this paper, we study the problem of utilising the duplicate and resemblance detection techniques to further compress data. We first present a framework of FIDCS-ERD, a feature-based intelligent deduplication compression system with extreme resemblance detection. We also introduce the main components and the detailed workflow of our compression system. We propose a content-defined chunking algorithm for duplicate detection and a Bloom filter-based resemblance detection algorithm. FIDCS-ERD implements the intelligent file chunking and the fast duplicate and resemblance detection. By extensive experiments over the real datasets, we demonstrate that FIDCS-ERD has better compression effect and more accurate resemblance detection compared to the existing approaches. Xiaotong Wu, Jiaquan Gao, Genlin Ji, Taotao Wu, Yuan Tian 0003, Najla Al-Nabhan |
Connect. Sci. | 2 |
| 2021 | Adaptive diagonal sparse matrix-vector multiplication on GPU
Jiaquan Gao, Renjie Yin, Guixia He |
J. Parallel Distributed Comput. | 1 |
| 2021 | A thread-adaptive sparse approximate inverse preconditioning algorithm on multi-GPUs
Jiaquan Gao, Guixia He |
Parallel Comput. | 1 |
| 2020 | An efficient sparse approximate inverse preconditioning algorithm on GPUabstractSummary The sparse approximate inverse (SPAI) preconditioner has proven to be effective in accelerating the convergence of iterative methods. Recently, accelerating it on the graphics processing unit (GPU) has attracted considerable attention due to the fact that the cost of constructing it is high. This motivates us to investigate how to accelerate the construction of SPAI preconditioners on GPU in this paper. We propose an efficient sparse approximate inverse algorithm on GPU, called SPAI‐Adaptive. For our proposed SPAI‐Adaptive, there are the following novelties: (1) an adaptive thread allocation strategy for SPAI‐Adaptive is proposed to assign the optimal thread number for each column of the preconditioner, and (2) Each component of the preconditioner, which includes finding indices I and J , constructing local submatrix, decomposing the local matrix into QR, and solving the upper triangular linear system, is computed in parallel inside a thread group of GPU. Experimental results show that the proposed SPAI‐Adaptive is effective, and has good performance and high parallelism. Guixia He, Renjie Yin, Jiaquan Gao |
Concurr. Comput. Pract. Exp. | 3 |
| 2018 | Efficient dense matrix-vector multiplication on GPUabstractSummary Given that the dense matrix‐vector multiplication (Ax or ATx) is of great importance in scientific computations, how to accelerate it is investigated on the graphics processing unit (GPU) in this paper. We present a warp‐based implementation of Ax on the GPU, called GEMV‐Adaptive, and a thread‐based implementation of ATx on the GPU, called GEMV‐T‐Adaptive. For our proposed GEMV‐Adaptive and GEMV‐T‐Adaptive, there are the following novelties: (1) an adaptive warp allocation strategy for GEMV‐Adaptive is proposed to assign the optimal warp number for each matrix row, (2) an adaptive thread allocation strategy for GEMV‐T‐Adaptive is designed to assign the optimal thread number to each matrix row, and (3) several optimization schemes are formulated. Experimental results show that the proposed GEMV‐Adaptive and GEMV‐T‐Adaptive mitigate the performance fluctuations of the implementations in the CUBLAS library, always have high performance, and outperform the most recently proposed GEMV and GEMV‐T kernels by Gao et al, respectively, for all test matrices. Guixia He, Jiaquan Gao, Jun Wang 0077 |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | A novel multi-graphics processing unit parallel optimization framework for the sparse matrix-vector multiplicationabstractSummary The sparse matrix‐vector multiplication (SpMV) is of great importance in scientific computations. Graphics processing unit (GPU)‐accelerated SpMVs for large‐sized problems have attracted considerable attention recently. We observe that on a specific multi‐GPU platform, the SpMV performance can usually be greatly improved when a matrix is partitioned into several blocks according to a predetermined rule and each block is assigned to a GPU with an appropriate storage format. This motivates us to propose a novel multi‐GPU parallel SpMV optimization framework, which involves the following parts: (1) a simple rule is defined to divide any given matrix among multiple GPUs; (2) a performance model, which is independent of the problems and dependent on the resources of devices, is proposed to accurately predict the execution time of SpMV kernels; and (3) a selection algorithm is suggested to automatically select the most appropriate one from the storage formats that are involved in the framework for the matrix block that is assigned to each GPU on the basis of the performance model. The objective of our framework does not construct a new storage format or algorithm but automatically and rapidly generates an optimally parallel SpMV for any sparse matrix on a specific multi‐GPU platform by integrating the existing storage formats and their corresponding kernels. We take 5 popular storage formats, for example, to present the idea of constructing the framework. Theoretically, we validate the correctness of our proposed SpMV performance model. This model is constructed only once for each type of GPU. Moreover, this framework is general and easy to be extensible. For a storage format that is not included in our framework, once the performance model of its corresponding SpMV kernel is successfully constructed, it can be incorporated into our framework. The experiments validate the efficiency of our proposed framework. Jiaquan Gao, Jun Wang 0077 |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | A multi-GPU parallel optimization model for the preconditioned conjugate gradient algorithm
Jiaquan Gao, Yuanshen Zhou, Guixia He |
Parallel Comput. | 1 |
| 2014 | Research on the conjugate gradient algorithm with a modified incomplete Cholesky preconditioner on GPU
Jiaquan Gao, Ronghua Liang, Jun Wang 0077 |
J. Parallel Distributed Comput. | 1 |
| 2013 | Modified Incomplete Cholesky Preconditioned Conjugate Gradient Algorithm on GPU for the 3D Parabolic Equation
Jiaquan Gao, Guixia He |
NPC | 1 |
| 2010 | A Quantum-Inspired Artificial Immune System for Multiobjective 0-1 Knapsack Problems
Jiaquan Gao, Guixia He |
ISNN (1) | 1 |
| 2009 | A Novel Artificial Immune System for Multiobjective Optimization Problems
Jiaquan Gao |
ISNN (3) | 1 |
| 2009 | A Novel Weight-Based Immune Genetic Algorithm for Multiobjective Optimization Problems
Guixia He, Jiaquan Gao |
ISNN (2) | 2 |
| 2008 | Multi-objective scheduling problems subjected to special process constraintabstractThe problem of parallel machine multi-objective scheduling subjected to special process constraint in the textile industries, as one of the most important combinational optimization problems, is different from other parallel machine scheduling problems in the following characteristics. On one hand, processing machines are non-identical; on the other hand, the sort of job processed on every machine can be restricted Considering one of the multi-objective problems, either minimizing the maximum completion time among all the machines(makespan) or minimizing the total earliness/tardiness penalty of all the jobs has been cornerstone of most studies done so far. However, under special process constraint, taking them into account as a multi-objective problem has not been well studied Therefore, in this paper, a multi-objective model based on them is presented and a new parallel genetic algorithm based on a vector group coding method is also proposed in order to effectively solve this model. The algorithm shows the following advantages: the coding method is simple and can effectively reflect the virtual scheduling policy, which can vividly reflect the numbers and sequences of these processed jobs on every machine, and then enables the individuals generated by crossover and mutation to satisfy process constraint. Numerical experiments show that it is efficient, and is better than the common genetic algorithm, and has the better parallel efficiency. A much better prospect of application can be optimistically expected. Jiaquan Gao, Guixia He, Yushun Wang |
IEEE Congress on Evolutionary Computation | 1 |