VLDB 2026 Research / reviewers in the wild / expert
Jianqiang Huang 0002
dblp:207/1901-2
· DBLP profile ↗
32ranked-venue papers
3as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SVSIG: Incremental Streaming Graph Processing with Source Vertex SuppressionabstractIn practice, graphs are often massive and continuously changing, with updates ranging from a single edge to hundreds of thousands of edges at once. Reusing previously computed intermediate values significantly reduces processing time. However, current state-of-the-art streaming graph processing systems that guarantee Bulk Synchronous Parallel (BSP) semantics often exhibit lower performance, especially when dealing with small-batch mutations. To address this challenge, we propose source vertex suppression and develop SVSIG, a streaming graph processing system that incorporates this technique. Unlike existing approaches that check value changes independently for each batch, source vertex suppression tracks accumulated deviations across multiple batches, propagating updates only when the accumulated deviation exceeds a predefined threshold. Implemented with efficient lock-free atomic operations, this cross-batch accumulation paradigm fundamentally reduces redundant vertex and edge activations. We provide theoretical analysis with provable error bounds, showing that SVSIG’s approximation error is bounded and converges. SVSIG supports iterative and aggregation-based graph algorithms while ensuring results closely approximate BSP semantics. Experimental results demonstrate that SVSIG outperforms state-of-the-art streaming graph processing systems by an average of 11.68 × (up to 66.3 ×) for small-batch mutations, with average error far below the theoretical bound. Jianqiang Huang 0002, Dapeng Fu, Haodong Bian |
ICS | 1 |
| 2026 | PANA: A Fine-Grained Runtime-Adaptive Load Balancing for Parallel SpMV on Multicore CPUsabstractSpMV has been widely utilized and is regarded as a significant kernel in various scientific and engineering computing applications, where its parallel performance is heavily influenced by matrix sparsity and hardware architecture. Despite extensive prior research, static partitioning strategies that narrowly target computation or memory access remain a key performance bottleneck, severely stifling performance advancement of SpMV on modern multicore CPUs. Haodong Bian, Youhui Zhang, Jianqiang Huang 0002, Xiaoying Wang 0002 |
PPoPP | 4 |
| 2026 | Qklu: a two-dimensional block-cyclic sparse direct solver
Renqian Wan, Jianqiang Huang 0002, Haodong Bian |
CCF Trans. High Perform. Comput. | 2 |
| 2026 | Robust low-rank representation on anchor graph with self-attention
Wenyi Feng, Jianqiang Huang 0002, Jinfang Jia, Ruirui Pu |
Frontiers Comput. Sci. | 2 |
| 2026 | Hierarchical fusion of local and global visual features with mixture-of-experts for remote sensing image scene classification
Yuanhao Tang, Xuechao Zou, Zhengpei Hu, Junliang Xing, Jianqiang Huang 0002 |
Neurocomputing | 6 |
| 2026 | Dynamic pivotal node attention for spatio-temporal wind speed forecasting
Guojing Zhang, Jianqiang Huang 0002, Xiaoying Wang 0002 |
Neurocomputing | 4 |
| 2026 | EDTC: Exact Triangle Counting for Dynamic Graphs on GPUabstractIn the process of updating a dynamic graph, an update to one edge may result in the addition or deletion of multiple triangles, while an update to multiple edges may only result in the addition or deletion of a single triangle. Consequently, accurately counting triangles on a dynamic graph is a challenging undertaking. As dynamic graphs are continuously updated, the GPU's memory may be insufficient to accommodate the storage of larger graphs. This presents a challenge when the graph, which is constantly growing, cannot be stored. The hash-based and binary search-based triangle counting algorithm is regarded as the most efficient for static graphs. However, when vertices with high degrees are encountered, the hash-based triangle counting method results in significant memory wastage due to the traditional construction of a hash table, leading to a shortage of memory. This issue remains unresolved. In this paper a triangle counting system EDTC is developed for dynamic graphs while ensuring the accuracy of counting. The system addresses three main problems: (1) An efficient EHTC algorithm is introduced to rapidly and accurately count the number of triangles in a graph. (2) The concept of an Update Activation CSR(UA-CSR) is introduced, along with a data structure to facilitate its implementation. This structure loads only the subgraph portion affected by the updated edge into the GPU, allowing calculations to be performed on this specific subgraph. (3) A compressed hash table is designed to reduce memory consumption, along with a dynamic shared memory assignment(DSA) strategy to fully utilize the shared memory of the GPU. Jiahao Tang, Jinxing Tu, Wei Xue 0003, Jianqiang Huang 0002 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2025 | A Fast and Lightweight Model for Causal Audio-Visual Speech SeparationabstractAudio-visual speech separation (AVSS) aims to extract a target speech signal from a mixed signal by leveraging both auditory and visual (lip movement) cues. However, most existing AVSS methods exhibit complex architectures and rely on future context, operating offline, which renders them unsuitable for real-time applications. Inspired by the pipeline of RTFSNet, we propose a novel streaming AVSS model, named Swift-Net, which enhances the causal processing capabilities required for real-time applications. Swift-Net adopts a lightweight visual feature extraction module and an efficient fusion module for audio-visual integration. Additionally, Swift-Net employs Grouped SRUs to integrate historical information across different feature spaces, thereby improving the utilization efficiency of historical information. We further propose a causal transformation template to facilitate the conversion of non-causal AVSS models into causal counterparts. Experiments on three standard benchmark datasets (LRS2, LRS3, and VoxCeleb2) demonstrated that under causal conditions, our proposed Swift-Net exhibited outstanding performance, highlighting the potential of this method for processing speech in complex environments. Wendi Sang, Kai Li 0047, Runxuan Yang, Jianqiang Huang 0002, Xiaolin Hu 0001 |
ECAI | 4 |
| 2025 | Research on Parallel Weighted Back-Projection Algorithm on Multi-CPU
Kaige Zheng, Kuangzheng Wu, Haodong Bian, Jianqiang Huang 0002, Xiaoying Wang 0002 |
ICIC (14) | 5 |
| 2025 | Exploring Linear Attention Alternative for Single Image Super-ResolutionabstractDeep learning-based single-image super-resolution (SISR) technology focuses on enhancing low-resolution (LR) images into high-resolution (HR) ones. Although significant progress has been made, challenges remain in computational complexity and quality, particularly in remote sensing image processing. To address these issues, we propose our Omni-Scale RWKV Super-Resolution (OmniRWKVSR) model which presents a novel approach that combines the Receptance Weighted Key Value (RWKV) architecture with feature extraction techniques such as Visual RWKV Spatial Mixing (VRSM) and Visual RWKV Channel Mixing (VRCM), aiming to overcome the limitations of existing methods and achieve superior SISR performance. Our work demonstrates the ability to provide effective solutions for high-quality image reconstruction. Under the 4x Super-Resolution tasks, compared to the MambaIR model, we achieved an average improvement of 0.26% in PSNR and 0.16% in SSIM. Rongchang Lu, Changyu Li, Donghang Li, Guojing Zhang, Jianqiang Huang 0002, Xilai Li |
IJCNN | 5 |
| 2025 | Octree-STCM: Octree-Based Spatio-Temporal Context Model for Lossless Geometry Compression of Dynamic Point CloudabstractDeep learning approaches have demonstrated remarkable effectiveness in point cloud geometry compression. However, existing octree-based methods face limitations due to insufficient contextual utilization within temporal sequences of dynamic point clouds. This paper proposes a spatio-temporal context model under an octree structure to enhance lossless compression of dynamic point cloud geometry. Firstly, a context extraction module is employed to capture the intra-contexts based on spatial correlations and the inter-contexts based on temporal dependencies. Subsequently, a context network employing 3D convolutional layers and fully connected layers is designed to extract spatio-temporal features from various contexts. After the context features integration, a multilayer perceptron is used to approximate the probability distribution of the occupancy symbol. The derived probability distributions finally optimize the arithmetic coding efficiency. Experimental results demonstrate that the proposed method outperforms the state-of-the-art octree-based approaches across multiple benchmark datasets. Zhecheng Wang 0002, Shuai Wan, Jianqiang Huang 0002 |
ICMR | 3 |
| 2025 | EAGLE: An Efficient Global Attention Lesion Segmentation Model for Hepatic Echinococcosis
Jiayan Chen, Kai Li 0047, Yulu Zhao, Jianqiang Huang 0002 |
PRCV (14) | 4 |
| 2025 | Imputing Missing Temperature Data of Meteorological Stations Based on Global Spatiotemporal Attention Neural NetworkabstractImputing missing meteorological site temperature data is necessary and valuable for researchers to analyze climate change and predict related natural disasters. Prior research often used interpolation-based methods, which basically ignored the temporal correlation existing in the site itself. Recently, researchers have attempted to leverage deep learning techniques. However, these models cannot fully utilize the spatiotemporal correlation in meteorological stations data. Therefore, this paper proposes a global spatiotemporal attention neural network (GSTA-Net), which consists of two sub networks, including the global spatial attention network and the global temporal attention network, respectively. The global spatial attention network primarily addresses the global spatial correlations among meteorological stations. The global temporal attention network predominantly captures the global temporal correlations inherent in meteorological stations. To further fully exploit and utilize spatiotemporal information from meteorological station data, adaptive weighting is applied to the outputs of the two sub-networks, thereby enhancing the imputation performance. Additionally, a progressive gated loss function has been designed to guide and accelerate GSTA-Net’s convergence. Finally, GSTA-Net has been validated through a large number of experiments on public dataset TND and QND with missing rates of 25%, 50%, and 75%, respectively. The experimental results indicate that GSTA-Net outperforms the latest models, including Linear, NLinear, DLinear, PatchTST, and STA-Net, across both the mean absolute error (MAE) and the root mean square error (RMSE) metrics. Tianrui Hou, Xinshuai Guo, Xiaoying Wang 0002, Guojing Zhang, Jianqiang Huang 0002 |
SMC | 6 |
| 2025 | Collaborative pseudo-label transfer for few-shot unsupervised domain adaptation
Jinfang Jia, Wandong Xue, Jianqiang Huang 0002 |
CCF Trans. High Perform. Comput. | 4 |
| 2025 | Research on GPU transplantation optimization of PRM scalar advection scheme in GRAPES global forecast system
Zhangjie Tan, Jinfang Jia, Zhengsheng Ning, Jianqiang Huang 0002, Xiaoying Wang 0002 |
CCF Trans. High Perform. Comput. | 4 |
| 2025 | An Efficient Parallel Ordered Depth-First Search Strategy for Directed Acyclic GraphsabstractABSTRACT With the advent of the big data era, accelerating the parallelization of Depth‐First Search (DFS) has become pivotal for addressing the challenges posed by large‐scale datasets and complex problems in contemporary applications. To improve the parallel processing performance of DFS on Directed Acyclic Graphs (DAGs) while maintaining the orderliness of traversal outcomes, this paper introduces an efficient Parallel Ordered Depth‐First Search (PODFS). By leveraging the novel concepts of Clue Path, ParallelList, and the node attributes Level and Dis, PODFS achieves precise subgraph partitioning while preserving the ordered nature of parallel search results. After performing a one‐time preprocessing on a specific graph, the proposed algorithm enables more efficient global traversals, achieving a speedup ranging from 6× to 12× on various real‐world graph datasets while maintaining the orderedness of traversal results. These performance improvements are crucial for applications that require frequent, in‐depth graph searches with a strict need to preserve traversal order. Chuqi Yan, Jianqiang Huang 0002 |
Concurr. Comput. Pract. Exp. | 2 |
| 2025 | Vector knowledge transfer-driven representation learning for heterogeneous hypernetworks
Yijian Chen, Xiaoying Wang 0002, Jianqiang Huang 0002 |
Knowl. Inf. Syst. | 4 |
| 2025 | CPAT: cross-patch aggregated transformer for time series forecasting
Xiaoying Wang 0002, Jianqiang Huang 0002, Guojing Zhang |
Mach. Learn. | 4 |
| 2025 | GDSR: Global-Detail Integration Through Dual-Branch Network With Wavelet Losses for Remote Sensing Image Super-ResolutionabstractIn recent years, deep neural networks, including Convolutional Neural Networks, Transformers, and State Space Models, have achieved significant progress in Remote Sensing Image (RSI) Super-Resolution (SR). However, existing SR methods typically overlook the complementary relationship between global and local dependencies. These methods either focus on capturing local information or prioritize global information, which results in models that are unable to effectively capture both global and local features simultaneously. Moreover, their computational cost becomes prohibitive when applied to large-scale RSIs. To address these challenges, we introduce the novel application of Receptance Weighted Key Value (RWKV) to RSI-SR, which captures long-range dependencies with linear complexity. To simultaneously model global and local features, we propose the Global-Detail dual-branch structure, GDSR, which performs SR by paralleling RWKV and convolutional operations to handle large-scale RSIs. Furthermore, we introduce the Global-Detail Reconstruction Module (GDRM) as an intermediary between the two branches to bridge their complementary roles. In addition, we propose the Dual-Group Multi-Scale Wavelet Loss, a wavelet-domain constraint mechanism via dual-group subband strategy and cross-resolution frequency alignment for enhanced reconstruction fidelity in RSI-SR. Extensive experiments under two degradation methods on several benchmarks, including AID, UCMerced, and RSSRD-QH, demonstrate that GSDR outperforms the state-of-the-art Transformer-based method HAT by an average of 0.09 dB in PSNR, while using only 63% of its parameters and 51% of its FLOPs, achieving an inference speed 3.2 times faster. Qiwei Zhu, Kai Li 0047, Guojing Zhang, Xiaoying Wang 0002, Jianqiang Huang 0002, Xilai Li |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | ARFA: An Asymmetric Receptive Field Autoencoder Model for Spatiotemporal PredictionabstractSpatiotemporal prediction aims to generate future sequences by paradigms learned from historical contexts. It is essential in numerous domains, such as traffic flow prediction and weather forecasting. Recently, research in this field has been predominantly driven by deep neural networks based on autoencoder architectures. However, existing methods commonly adopt autoencoder architectures with identical receptive field sizes. To address this issue, we propose an Asymmetric Receptive Field Autoencoder (ARFA) model, which introduces corresponding sizes of receptive field modules tailored to the distinct functionalities of the encoder and decoder. In the encoder, we present a large kernel module for global spatiotemporal feature extraction. In the decoder, we develop a small kernel module for local spatiotemporal information reconstruction. Experimental results demonstrate that ARFA consistently achieves state-of-the-art performance on popular datasets. Additionally, we construct the RainBench, a large-scale radar echo dataset for precipitation prediction, to address the scarcity of meteorological data in the domain. Xuechao Zou, Xiaoying Wang 0002, Jianqiang Huang 0002, Junliang Xing |
ICASSP | 5 |
| 2024 | DisRot: boosting the generalization capability of few-shot learning via knowledge distillation and self-supervised learning
Chenyu Ma, Jinfang Jia, Jianqiang Huang 0002, Xiaoying Wang 0002 |
Mach. Vis. Appl. | 3 |
| 2024 | Heterogeneous Hypernetwork Representation Learning With Hyperedge FusionabstractMost of the existing hypernetwork representation learning methods fail to fully consider the hyperedges, leading to the untapped potential of information contained within the hyperedges. To address this issue, this article proposes a heterogeneous hypernetwork representation learning method with hyperedge fusion abbreviated as HRHF. First, this method incorporates the hyperedges into random walk node sequences by means of incidence graph to enhance tuple relationships, i.e., the hyperedges among the nodes. Second, under the condition of the above random walk node sequences, the cognitive structure model, cognitive set model, and cognitive hyperedge model are jointly optimized to comprehensively consider pairwise relationships and tuple relationships among the nodes to learn high-quality node representation vectors. The experimental results demonstrate that, for the link prediction task, this method outperforms other optimal baseline methods, i.e., hyper-path-based random walks + hyper-gram (HPHG) by 0.99% points on the drug dataset, is comparable to the other optimal method, i.e., Event2vec on the global positioning system (GPS) dataset, is close to the performance of other optimal methods, i.e., hyper-path-based random walks + skip-gram (HPSG) on the MovieLens, and is close to the performance of other optimal methods, i.e., Hyper2vec on the WordNet dataset. For the hypernetwork reconstruction task, this method achieves superior average performance on the drug and GPS datasets compared with other baseline methods. Xiaoying Wang 0002, Jianqiang Huang 0002 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2023 | STA-Net: Reconstruct Missing Temperature Data of Meteorological Stations Using a Spatiotemporal Attention Neural Network
Tianrui Hou, Xinzhong Zhang, Xiaoying Wang 0002, Jianqiang Huang 0002 |
ICONIP (7) | 5 |
| 2023 | PAS: A new powerful and simple quantum computing simulatorabstractAbstract In recent years, many researchers have been using CPU for quantum computing simulation. However, in reality, the simulation efficiency of the large‐scale simulator is low on a single node. Therefore, striving to improve the simulator efficiency on a single node has become a serious challenge that many researchers need to solve. After many experiments, we found that much computational redundancy and frequent memory access are important factors that hinder the efficient operation of the CPU. This paper proposes a new powerful and simple quantum computing simulator: PAS (power and simple). Compared with existing simulators, PAS introduces four novel optimization methods: efficient hybrid vectorization, fast bitwise operation, memory access filtering, and quantum tracking. In the experiment, we tested the QFT (quantum Fourier transform) and RQC (random quantum circuits) of 21 to 30 qubits and selected the state‐of‐the‐art simulator QuEST (quantum exact simulation toolkit) as the benchmark. After experiments, we have concluded that PAS compared with QuEST can achieve a mean speedup of (QFT), (RQC) (up to , ) on the Intel Xeon E5‐2670 v3 CPU. Haodong Bian, Jianqiang Huang 0002, Jiahao Tang, Runting Dong, Xiaoying Wang 0002 |
Softw. Pract. Exp. | 2 |
| 2022 | A Two-stage Algorithm Based on Prediction and Search for Maxk-Truss DecompositionabstractThe cohesive subgraph K-truss is often used to detect communities in large-scale social networks. The truss decomposition algorithm calculates the number of triangles formed by each edge, then peels off the edges with the number of triangles less than k-2 iteratively. The incremental maxk-truss algorithm must detect k-truss one by one making it inefficient. We propose a two-stage maxk-ktruss decomposition algorithm PSKT. The innovation points of PSKT are as follows: 1) The k-core structure is used to predict the maxk value, and the structure helps us avoid unnecessary mass calculations. 2) The characteristic that the k-truss of graph G must be included in the (k-l)-core of graph G is utilized to ensure that the proposed algorithm can safely prune in the graph. 3) A suitable triangle counting algorithm and the dynamic stream compression method are used to improve the algorithm efficiency. Comprehensive experiments demonstrate that PSKT is 15.26 times faster than the incremental algorithm FMT-dec and 3 times faster than the state-of-the-art maxk-truss algorithm FMT-max. Jiahao Tang, Lingbin Liu, Jinfang Jia, Xiaoying Wang 0002, Jianqiang Huang 0002 |
ICPADS | 6 |
| 2022 | Simulation of three-dimensional phase field model with LBM method using OpenCL
Chenyu Ma, Jinfang Jia, Jianqiang Huang 0002, Xiaoying Wang 0002 |
J. Supercomput. | 5 |
| 2022 | $TC-Stream$TC-Stream: Large-Scale Graph Triangle Counting on a Single Machine Using GPUsabstractIn this paper, we build a TC-Stream, a high-performance graph processing system specific for a triangle counting algorithm on graph data with up to tens of billions of edges, which significantly exceeds the device memory capacity of Graphics Processing Units (GPUs). The triangle counting problem is a broad research topic in data mining and social network analysis in the graph processing field. As the scale of the graph data grows, a portion of the graph data must be loaded iteratively. To solve the above problem, we propose TC-Stream. It focuses on three issues: 1) For power-law graphs, because the amount of tasks of each vertex or edge is inconsistent, it is bound to cause different demands of computing and memory resources for different task types. We propose a parallel vertex approach and the reordering of vertices for graph data that can be placed in the GPU device memory to ensure the maximum workload balancing; 2) A binary-search-based set intersection method is designed to achieve the maximum parallelism in GPU; 3) For the graph data that exceeds the GPU device memory capacity, we develop a novel vertical partition algorithm to guarantee the independent computing on each partition so that the three computation processes, i.e., the computation on GPU, the data transmission between main memory of CPU and SSD, and the communication between the CPU and the GPU can be perfectly overlapped. Extensive experiments conducted on large-scale datasets showed that the TC-stream running on a single Tesla V100 GPU performs 2.4 6 and 1.8 4.4 faster than the state-of-the-art single-machine in-memory triangle counting system and GPU-based triangle counting system, respectively, and achieves 2.4faster than the state-of-the-art out-of-core distributed system PDTL running on an 8-node cluster when processing the graph data with 42.5 billion edges. Jianqiang Huang 0002, Haojie Wang 0004, Xiaoying Wang 0002 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | ALBUS: A method for efficiently processing SpMV using SIMD and Load balancing
Haodong Bian, Jianqiang Huang 0002, Lingbin Liu, Dongqiang Huang, Xiaoying Wang 0002 |
Future Gener. Comput. Syst. | 2 |
| 2020 | CSR2: A New Format for SIMD-accelerated SpMVabstractSpMV (Sparse matrix-vector multiplication) has attracted the attention of researchers in related fields at home and abroad. Of course, improving SpMV performance has also been a research hot spot for researchers in related fields. In this paper, we propose a new sparse matrix storage format CSR2 (Compressed Sparse Row 2) suitable for SIMD (Single Instruction Multiple Data)-accelerated SpMV. First, the format operation of CSR2 is easy to implement and has a low overhead of conversion. Second, CSR2 is a new single format and suitable for use on processor platforms with SIMD vectorization. We compare the SpMV algorithm based on CSR21with the one based on the current most advanced single format CSR5 (Compressed Sparse Row 5) on two mainstream high-performance processors: Intel Core i7-7700HQ CPU and Intel Xeon CPU E5-2670 v3. We choose 10 sets of regular matrices and 3 sets of irregular matrices to be used as benchmark suit. Experiments show that for the 13 sets of regular and irregular matrices in the benchmark suite, CSR2 has an average performance improvement of more than 50% compared to CSR5 (up to 125% on Intel Core i77700HQ CPU and 303% on Intel Xeon CPU E5-2670 v3). For applications with multiple iterations, in reality, using our CSR2 can bring low-overhead format conversion and high-throughput computing performance. Haodong Bian, Jianqiang Huang 0002, Runting Dong, Lingbin Liu, Xiaoying Wang 0002 |
CCGRID | 2 |
| 2020 | HpQC: A New Efficient Quantum Computing Simulator
Haodong Bian, Jianqiang Huang 0002, Runting Dong, Yuluo Guo, Xiaoying Wang 0002 |
ICA3PP (2) | 2 |
| 2020 | Survey of single image super-resolution reconstructionabstractImage super‐resolution reconstruction refers to a technique of recovering a high‐resolution (HR) image (or multiple images) from a low‐resolution (LR) degraded image (or multiple images). Due to the breakthrough progress in deep learning in other computer vision tasks, people try to introduce deep neural network and solve the problem of image super‐resolution reconstruction by constructing a deep‐level network for end‐to‐end training. The currently used deep learning models can divide the SISR model into four types: interpolation‐based preprocessing‐based model, original image processing based model, hierarchical feature‐based model, and high‐frequency detail‐based model, or shared the network model. The current challenges for super‐resolution reconstruction are mainly reflected in the actual application process, such as encountering an unknown scaling factor, losing paired LR–HR images, and so on. Kai Li 0047, Shenghao Yang 0003, Runting Dong, Xiaoying Wang 0002, Jianqiang Huang 0002 |
IET Image Process. | 5 |
| 2019 | Pimiento: A Vertex-Centric Graph-Processing Framework on a Single Machine
Jianqiang Huang 0002, Xiaoying Wang 0002 |
ICA3PP (2) | 1 |