VLDB 2026 Research / reviewers in the wild / expert
Qicong Wang
dblp:79/6000
· DBLP profile ↗
24ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0001-7324-0433ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 2 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ProPL: Universal Semi-Supervised Ultrasound Image Segmentation via Prompt-Guided Pseudo-LabelingabstractExisting approaches for the problem of ultrasound image segmentation, whether supervised or semi-supervised, are typically specialized for specific anatomical structures or tasks, limiting their practical utility in clinical settings. In this paper, we pioneer the task of universal semi-supervised ultrasound image segmentation and propose ProPL, a framework that can handle multiple organs and segmentation tasks while leveraging both labeled and unlabeled data. At its core, ProPL employs a shared vision encoder coupled with prompt-guided dual decoders, enabling flexible task adaptation through a prompting-upon-decoding mechanism and reliable self-training via an uncertainty-driven pseudo-label calibration (UPLC) module. To facilitate research in this direction, we introduce a comprehensive ultrasound dataset spanning 5 organs and 8 segmentation tasks. Extensive experiments demonstrate that ProPL outperforms state-of-the-art methods across various metrics, establishing a new benchmark for universal ultrasound image segmentation. Yaxiong Chen, Qicong Wang, Jingliang Hu, Yilei Shi, Shengwu Xiong 0001, Xiao Xiang Zhu 0001, Lichao Mou |
AAAI | 2 |
| 2026 | HGFormer: Hyperbolic Graph Transformer for Unsupervised Smart Contract Vulnerability Detection
Haiming Zhu, Yaoxin Chen, Qicong Wang |
ICIC (11) | 5 |
| 2026 | RGCNet: Riemannian graph convolutional networks for end-to-end smart contract vulnerability detectionabstractFrequent security issues with smart contract vulnerabilities have become a pressing challenge in the industry. Conventional program analysis methods lack flexibility and extensibility, leading to high false positive rates. Deep learning approaches are emerging as a new trend to address this issue. Compared to other neural networks, graph convolutional networks can better capture the structural and logical information of smart contracts. However, existing methods do not fully consider the scale-free characteristics of smart contracts and fail to leverage their complex hierarchical structures and semantic information. Therefore, we develop an end-to-end vulnerability detection framework using Riemannian Graph Convolutional Networks (RGCNet). We first construct smart contract graphs that are rich in semantic and structural information. Next, we learn features of the smart contract graph in the Riemannian manifold, thereby better reflecting its actual topology. Simultaneously, the word embedding network extracts semantic features, forming an end-to-end network where modules promote one another. Extensive experiments are conducted on three vulnerabilities using real-world smart contracts. The results show that the proposed approach exhibits superior performance over state-of-the-art methodologies in terms of accuracy, precision, and recall. Yaoxin Chen, Haiming Zhu, Qicong Wang, Maozhen Li 0001 |
Neurocomputing | 5 |
| 2025 | ITW-DehazeFormer: Imaging through Turbid Water Using Improved DehazeFormerabstractLight scattering and absorption degrade the quality of underwater images, and various image enhancement methods have been explored. However, the existing underwater image datasets lack corresponding high-quality references, and the degree of scattering and absorption is not strictly controlled. In this study, we constructed an image dataset with different degrees of light scattering and controlled water turbidity via a water tank. The Swin Transformer based dehazing network DehazeFormer has been improved, termed ITW-DehazeFormer, to enhance images acquired through turbid water. First, a histogram equalization pre-enhancement block is added. Second, the SKfusion block is replaced by a content-guided attention based fusion block to combine channel and spatial attention so that information interactions between different channels are guaranteed. Finally, a hybrid loss function combining space and frequency domain information is introduced. Experimental results show that ITW-DehazeFormer outperforms seven existing image enhancement methods in terms of several image quality metrics, including Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measure (SSIM) and multi-scale SSIM. Qicong Wang, Xiaopin Zhong, Dajiang Lu, Yibin Tian |
ICASSP | 1 |
| 2025 | Multi-level domain adaptation for improved generalization in electroencephalogram-based driver fatigue detection
Fuzhong Huang, Qicong Wang, Wang Mei, Zhenchang Zhang, Zelong Chen |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | PaperEval: A universal, quantitative, and explainable paper evaluation method powered by a multi-agent system
Shengzhi Huang, Qicong Wang, Wei Lu 0019, Lingyu Liu, Zhenzhen Xu, Yong Huang 0008 |
Inf. Process. Manag. | 2 |
| 2025 | Bilinear Parallel Fourier Transformer for Multimodal Remote Sensing ClassificationabstractVision Transformers (ViTs) have shown promise in multimodal fusion image classification, yet face performance challenges in complex remote sensing scenarios. Single fusion frameworks often fail to fully utilize multimodal diversity, and the uneven distribution of image categories complicates the accurate construction of spatial structures by Transformers. Additionally, traditional cross-entropy tends to favor majority classes, neglecting minority classes, resulting in suboptimal predictions and reduced overall accuracy (OA). To solve these challenges, we propose a novel deep neural network, a bilinear parallel Fourier Transformer (BPFT). We propose a novel dual-fusion feature interaction (DFFI) module that utilizes two distinct types of fused features for learning, namely the spatial-spectral fusion feature and the global fusion feature. Besides, we introduce a dual-feature interaction (DFI) module to improve the utilization of fused feature information. To enable the Transformer to better establish spatial structural relationships, we employ the Fourier transform in place of the self-attention mechanism. To address the focus on minority class labels, we propose an exponential label smoothing cross-entropy loss function. This loss function comprises two components: exponential cross-entropy and label smoothing. The exponential cross-entropy component applies a strong penalty to misclassified samples, thereby increasing attention on minority class labels. To validate the efficacy of our approach, extensive experiments are conducted across two multimodal remote sensing datasets: Augsburg and Berlin, encompassing hyperspectral imaging (HSI) data and synthetic aperture radar (SAR) data. The results of these experiments affirm the superior performance of our proposed BPFT model compared to existing state-of-the-art models in multimodal remote sensing image classification tasks. Yaxiong Chen, Qicong Wang, Yichen Zhao, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | SSRL: Self-Supervised Spatial-Temporal Representation Learning for 3D Action RecognitionabstractFor 3D action recognition, the main challenge is to extract long-range semantic information in both temporal and spatial dimensions. In this paper, in order to better excavate long-range semantic information from large number of unlabelled skeleton sequences, we propose Self-supervised Spatial-temporal Representation Learning (SSRL), a contrastive learning framework to learn skeleton representation. SSRL consists of two novel inference tasks that enable the network to learn global semantic information in the temporal and spatial dimensions, respectively. The temporal inference task learns the temporal persistence of human actions through temporally incomplete skeleton sequences. And the spatial inference task learns the spatially coordinated nature of human action through spatially partially skeleton sequence. We design two transformation modules to efficiently realize these two tasks while fitting the encoder network. To avoid the difficulty of constructing and maintaining high-quality negative samples, our proposed framework learns by maintaining consistency among positive samples without the need of any negative sample. Experiments demonstrate that our proposed method can achieve better results in comparison with state-of-the-art methods under a variety of evaluation protocols on NTU RGB+D 60, PKU-MMD and NTU RGB+D 120 datasets. Zhihao Jin, Qicong Wang, Yehu Shen, Hongying Meng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Self-Supervised 3D Behavior Representation Learning Based on Homotopic Hyperbolic EmbeddingabstractBehavior sequences are generated by a series of spatio-temporal interactions and have a high-dimensional nonlinear manifold structure. Therefore, it is difficult to learn 3D behavior representations without relying on supervised signals. To this end, self-supervised learning methods can be used to explore the rich information contained in the data itself. Context-context contrastive self-supervised methods construct the manifold embedded in Euclidean space by learning the distance relationship between data, and find the geometric distribution of data. However, traditional Euclidean space is difficult to express context joint features. In order to obtain an effective global representation from the relationship between data under unlabeled conditions, this paper adopts contrastive learning to compare global feature, and proposes a self-supervised learning method based on hyperbolic embedding to mine the nonlinear relationship of behavior trajectories. This method adopts the framework of discarding negative samples, which overcomes the shortcomings of the paradigm based on positive and negative samples that pull similar data away in the feature space. Meanwhile, the output of the network is embedded in a hyperbolic space, and a multi-layer perceptron is added to convert the entire module into a homotopic mapping by using the geometric properties of operations in the hyperbolic space, so as to obtain homotopy invariant knowledge. The proposed method combines the geometric properties of hyperbolic manifolds and the equivariance of homotopy groups to promote better supervised signals for the network, which improves the performance of unsupervised learning. Jinghong Chen, Zhihao Jin, Qicong Wang, Hongying Meng |
IEEE Trans. Image Process. | 3 |
| 2022 | Unsupervised visual feature learning based on similarity guidance
Zhihao Jin, Qicong Wang, Wenming Yang, Qingmin Liao, Hongying Meng |
Neurocomputing | 3 |
| 2022 | An end-to-end heterogeneous network for graph similarity learning
Yan Huang 0030, Qicong Wang, Hongying Meng |
Neurocomputing | 4 |
| 2022 | Self-Supervised Representation Learning for Videos by Segmenting via Sampling Rate Order PredictionabstractSelf-supervised representation learning for videos has been very attractive recently because these methods exploit the information inherently obtained from the video itself instead of annotated labels that is quite time-consuming. However, existing methods ignore the importance of global observation while performing spatio-temporal transformation perception, which highly limits the expression capabilities of the video representation. This paper proposes a novel pretext task that combines the temporal information perception of the video with the motion amplitude perception of moving objects to learn the spatio-temporal representation of the video. Specifically, given a video clip containing several video segments, each video segment is sampled by different sampling rates and the order of video segments is disrupted. Then, the network is used to regress the sampling rate of each video segment and classify the order of input video segments. In the pre-training stage, the network can learn rich spatio-temporal semantic information where content-related contrastive learning is introduced to make the learned video representation more discriminative. To alleviate the appearance dependency caused by contrastive learning, we design a novel and robust vector similarity measurement approach, which can take feature alignment into consideration. Moreover, a view synthesis framework is proposed to further improve the performance of contrastive learning by automatically generating reasonable transformed views. We conduct benchmark experiments with several 3D backbone networks on two datasets. The results show that our proposed method outperforms the existing state-of-the-art methods across the three backbones on two downstream tasks of human action recognition and video retrieval. Yan Huang 0030, Qicong Wang, Wenming Yang, Hongying Meng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | A voxelized point clouds representation for object classification and segmentation on 3D data
Abubakar Sulaiman Gezawa, Zikirillahi A. Bello, Qicong Wang |
J. Supercomput. | 3 |
| 2020 | Multi-GPU Parallel Implementation of Spatial-Spectral Kernel Sparse Representation for Hyperspectral Image ClassificationabstractClassification is one of the major research fields in hyperspectral imagery. Due to the fact that neighboring pixels are more likely to share the same label, it is practical to use spatial information in hyperspectral image to achieve higher accuracy. On the other hand, however, spatial information also leads to higher computational complexity. This paper proposes an efficient implementation of a spatial-spectral kernel sparse representation for hyperspectral image classification base on the multi-GPU platform. The proposed implementation takes advantage of the capability of compute-unified device architecture (CUDA), such as shared memory, streams and peer-to-peer (P2P) transfer of data. In addition, an improvement of performance can be achieved by calculation reorganization and bandwidth usage optimization. Experimental results demonstrate that the proposed method achieves an up to 56.81X speedup in computation time while guaranteeing the classification accuracy. Weishi Deng, Zebin Wu 0001, Qicong Wang, Jin Sun 0001, Yang Xu 0006, Jiandong Yang, Zhihui Wei, Hongyi Liu 0001 |
IGARSS | 4 |
| 2019 | Reweighted sparse representation with residual compensation for 3D human pose estimation from a single RGB image
Mengxi Jiang, Zhu Liang Yu, Yan Zhang 0059, Qicong Wang, Cuihua Li |
Neurocomputing | 4 |
| 2018 | Temporal sparse feature auto-combination deep network for video action recognitionabstractSummary In order to deal with action recognition for large‐scale video data, we present a spatio‐temporal auto‐combination deep network, which is able to extract deep features from short video segments by making full use of temporal contextual correlation of corresponding pixels among successive video frames. Based on conventional sparse encoding, we further consider the representative features in adjacent nodes of the hidden layers according to activation states similarities. A sparse auto‐combination strategy is applied to multiple input maps in each convolution stage. An information constraint of the representative features of hidden layer nodes is imposed to handle the adaptive sparse encoding of the topology. As a result, the learned features can represent the spatio‐temporal transition relationships better and the number of hidden nodes can be restricted to a certain range. We conduct a series of experiments on two public data sets. The experimental results show that our approach is more effective and robust in video action recognition compared with traditional methods. Qicong Wang, Dingxi Gong, Man Qi, Yehu Shen |
Concurr. Comput. Pract. Exp. | 1 |
| 2018 | Transfer learning-based online multiperson tracking with Gaussian process regressionabstractSummary Most existing tracking‐by‐detection approaches are affected by abrupt pedestrian pose changes, lighting conditions, scale changes, and real‐time processing, which leads to issues such as detection errors and drifts. To deal with these issues, we present a novel multi‐person tracking framework by introducing a new Gaussian Process Regression based observation model, which learns in a semi‐supervised manner. The background information is taken into consideration to build the discriminative tracker, training samples are re‐weighted appropriately to ease the impact of the potential sample misalignment and noisy during model updating. Unlabeled samples from the current frame provide rich information, which is used for enhancing the tracking inference. Experimental results show that the proposed approach outperforms a number of state‐of‐the‐art methods on some benchmark datasets. Baobing Zhang, Siguang Li, Zhengwen Huang, Babak H. Rahi, Qicong Wang, Maozhen Li 0001 |
Concurr. Comput. Pract. Exp. | 5 |
| 2016 | Face recognition by decision fusion of two-dimensional linear discriminant analysis and local binary pattern
Qicong Wang, Xinjie Hao, Lisheng Chen, Jingmin Cui, Rongrong Ji |
Frontiers Comput. Sci. | 1 |
| 2015 | Real-Time Implementation of the Sparse Multinomial Logistic Regression for Hyperspectral Image Classification on GPUsabstractIn this letter, a real-time implementation of the logistic regression via variable splitting and augmented Lagrangian (LORSAL) algorithm for sparse multinomial logistic regression is presented on commodity graphics processing units (GPUs) using Nvidia's compute unified device architecture. The proposed parallel method properly exploits the GPU architecture at the low level, including its shared memory, and takes full advantage of the computational power of GPUs to achieve real-time classification performance of hyperspectral images for the first time in the hyperspectral imaging literature. Our experimental results reveal remarkable acceleration factors and real-time performance, while retaining exactly the same classification accuracy with regard to the serial and multicore versions of the classifier. Zebin Wu 0001, Qicong Wang, Antonio Plaza, Jun Li 0009, Le Sun 0002, Zhihui Wei |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2015 | Supervised sparse manifold regression for head pose estimation in 3D space
Qicong Wang, Yuxiang Wu, Yehu Shen |
Signal Process. | 1 |
| 2014 | Supervised locality discriminant manifold learning for head pose estimation
Qicong Wang |
Knowl. Based Syst. | 2 |
| 2009 | Symmetry Detection for Multi-object Using Local Polar Coordinate
Yuanhao Gong, Qicong Wang, Chenhui Yang, Yahui Gao, Cuihua Li |
CAIP | 2 |
| 2006 | Enhancing Particle Swarm Optimization Based Particle Filter Tracker
Qicong Wang, Jilin Liu, Zhiyu Xiang |
ICIC (2) | 1 |
| 2006 | Object Tracking Using Genetic Evolution Based Kernel Particle Filter
Qicong Wang, Jilin Liu |
IWCIA | 1 |