VLDB 2026 Research / reviewers in the wild / expert
Quan Tang 0001
dblp:150/5249-1
· DBLP profile ↗
22ranked-venue papers
8as first author
22since 2021 · last 2026
0000-0003-4011-6166ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Computer networks · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multitask Cooperative Genetic Programming for Co-Scheduling Online-Offline Workflows in the Cloud
Zaixing Sun, Quan Tang 0001, Jun Jiang 0003, Chonglin Gu, Bin Wang 0048 |
INFOCOM | 3 |
| 2026 | CaSS: Category-Aware Semantic Segmentation With Vision-Language PriorsabstractSemantic image segmentation typically relies on static pixel- wise classifiers that operate over a fixed category space, making them insensitive to the actual semantic composition of a real-world image. In this work, we propose CaSS, a category-aware semantic segmentation framework that introduces image-level semantic priors to dynamically adapt pixel-level classification. Specifically, a pre-trained vision–language model is employed to infer the set of semantic categories present in an image, which is encoded as a structured category prior. This prior is then used to drive a lightweight dynamic parsing network that generates image-conditioned classifier parameters for pixel- wise segmentation. By explicitly constraining the classifier with category-aware priors, CaSS reduces interference from absent classes and enhances both intra-class consistency and inter-class discriminability. The proposed approach follows a large–small model collaboration paradigm, leveraging the strong semantic understanding of vision–language models while preserving efficient pixel-level inference. Extensive experiments on standard semantic segmentation benchmarks demonstrate that CaSS consistently improves segmentation accuracy over state-of-the-art methods with minimal parameter overhead. Quan Tang 0001, Dengke Zhang, Xuhao Tang 0001, Bin Wang 0048, Cuifeng Du, Jun Jiang 0003 |
IEEE Signal Process. Lett. | 1 |
| 2025 | CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic Segmentation
Dengke Zhang, Fagui Liu, Quan Tang 0001 |
ICCV | 3 |
| 2025 | EK-Net++: Real-time scene text detection with expand kernel distance and Epoch Adaptive Weight
Boyuan Zhu, Quan Tang 0001, C. L. Philip Chen, Fagui Liu |
Expert Syst. Appl. | 3 |
| 2025 | Increase the sensitivity of moderate examples for semantic image segmentation
Quan Tang 0001, Fagui Liu, Dengke Zhang, Jun Jiang 0003, Xuhao Tang 0001, C. L. Philip Chen |
Image Vis. Comput. | 1 |
| 2025 | Exploring Token-Level Augmentation in Vision Transformer for Semi-Supervised Semantic SegmentationabstractSemi-supervised semantic segmentation has witnessed remarkable advancements in recent years. However, existing algorithms are based on convolutional neural networks, and directly applying them to Vision Transformers poses certain limitations due to conceptual disparities. To this end, we propose TokenSwap, a data augmentation technique designed explicitly for semi-supervised semantic segmentation with Vision Transformers. TokenSwap aligns well with the global attention mechanism by mixing images at the token level, enhancing the learning capability for contextual information among image patches and the utilization of unlabeled data. We further incorporate image augmentation and feature augmentation to promote the diversity of augmentation. Moreover, to enhance consistency regularization, we propose a dual-branch framework where each branch applies image and feature augmentation to the input image. We conduct extensive experiments across multiple benchmark datasets, including Pascal VOC 2012, Cityscapes, and COCO. Results suggest that the proposed method outperforms state-of-the-art algorithms with notably observed accuracy improvement, especially under limited fine annotations. Dengke Zhang, Quan Tang 0001, Fagui Liu, Haiqing Mei, C. L. Philip Chen |
IEEE Signal Process. Lett. | 2 |
| 2025 | Cost and Makespan-Aware Task Scheduling With Deep Reinforcement Learning in Multicloud EnvironmentsabstractThe multicloud environments (MCE) represent a novel paradigm encompassing multiple infrastructure as a service (IaaS) providers, enabling users to tailor and optimize cloud services according to their specific requirements. This approach effectively addresses the limitations of a single cloud environment (SCE) regarding technical constraints, geographical coverage deficiencies, and cost-effectiveness concerns while catering to the increasingly diverse and expanding user demands. In MCE, users must employ appropriate strategies to efficiently allocate diverse tasks across multiple cloud service providers (CSPs) by leveraging the best available resources. Traditional scheduling algorithms are inadequate for addressing the complexities of such MCE. This study introduces a framework for the task scheduling procedure in MCE, treating independent task scheduling as a Markov decision process (MDP). We propose a novel agent environment framework that is designed based on the distinctive characteristics of MCE and enables independent task scheduling. Furthermore, we propose a task scheduling algorithm for MCE based on deep reinforcement learning (DRL) to optimize cost and makespan according to diverse user requirements. The simulation experiments are conducted using both simulated datasets and real-world datasets, demonstrating that our proposed algorithm surpasses the other five algorithms in terms of cost minimization and makespan optimization. Xuhao Tang 0001, Fagui Liu, Bin Wang 0048, Jun Jiang 0003, Quan Tang 0001, Qingbo Wu 0003, C. L. Philip Chen |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2025 | Rethinking Feature Reconstruction via Category Prototype in Semantic SegmentationabstractThe encoder-decoder architecture is a prevailing paradigm for semantic segmentation. It has been discovered that aggregation of multi-stage encoder features plays a significant role in capturing discriminative pixel representation. In this work, we rethink feature reconstruction for scale alignment of multi-stage pyramidal features and treat it as a Query Update (Q-UP) task. Pixel-wise affinity scores are calculated between the high-resolution query map and low-resolution feature map to dynamically broadcast low-resolution pixel features to match a higher resolution. Unlike prior works (e.g. bilinear interpolation) that only exploit sub-pixel neighborhoods, Q-UP samples contextual information within a global receptive field via a data-dependent manner. To alleviate intra-category feature variance, we substitute source pixel features for feature reconstruction with their corresponding category prototype that is assessed by averaging all pixel features belonging to that category. Besides, a memory module is proposed to explore the capacity of category prototypes at the dataset level. We refer to the method as Category Prototype Transformer (CPT). We conduct extensive experiments on popular benchmarks. Integrating CPT into a feature pyramid structure exhibits superior performance for semantic segmentation even with low-resolution feature maps, e.g. 1/32 of the input size, significantly reducing computational complexity. Specifically, the proposed method obtains a compelling 55.5% mIoU with greatly reduced model parameters and computations on the challenging ADE20K dataset. Quan Tang 0001, Chuanjian Liu, Fagui Liu, Jun Jiang 0003, Bowen Zhang 0009, C. L. Philip Chen, Kai Han 0002, Yunhe Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Tightly Coupled RTK-Visual-Inertial Integration With a Novel Sliding Ambiguity Window Optimization FrameworkabstractAccurate and reliable navigation is fundamental for intelligent transportation applications such as autonomous driving. Current research community has gained advances in global navigation satellite system (GNSS) and its fusion with visual-inertial navigation system (VINS). However, existing optimization-based integrations fail to fully utilize the constant characteristic of carrier phase ambiguity when facing frequent cycle slips and degraded GNSS signals, leading to erratic navigation output in complex environments. To address these limitations, this work proposes a tightly coupled GNSS real-time kinematic (RTK)-visual-inertial integration with a novel sliding ambiguity window optimization framework to achieve high-precision and robust positioning. Specifically, a sliding ambiguity window framework is built to associate float single-differenced ambiguities estimated by factor graph optimization and integer double-differenced ambiguities obtained by LAMBDA ambiguity resolution (AR). The framework can maintain continuous AR constraints across epochs and extend ambiguity-fixed solutions even during AR failures. Additionally, a VINS-aided single&dual-frequency hybrid cycle slip detection method is proposed, combining available dual-frequency GNSS observations and VINS information to perform reliable cycle slip detection for all single- and dual-frequency carrier phases. The superiority of the proposed method is verified by real-world vehicle traveling experiments in campus and urban scenarios. Results show that our proposed method can achieve robust centimeter-level positioning accuracy in both trajectories, outperforming the state-of-the-art optimization-based integration method by 72.2% and 66.0%, respectively. Chufeng Duan, Shengquan Li 0002, Quan Tang 0001, Zhiqiang Dai, Xiangwei Zhu |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Incremental Semi-Supervised Learning for Data Streams Classification in Internet of ThingsabstractData stream classification is widely used in Internet of Things (IoT) scenarios such as health monitoring, anomaly detection and online diagnosis. Due to the continuous data stream changing dynamically over time, it is impossible to classify all the data simultaneously. Moreover, labeling each sample in practical data stream applications is time-and resource-consuming. The realistic situation is that only a few instances in a data stream are labeled. Therefore, classifying data streams with limited labels has become challenging in IoT scenarios. In this paper, we propose an incremental dynamic weighted semi-supervised method for classifying IoT data streams. Considering the dynamics and continuity in data streams, we use a chunk-based approach to learn the features in the data stream and assign weights to the classifier dynamically. Moreover, we deploy incremental learning methods to continuously learn from the sampled labeled data stream to update the classifier model, which can take advantage of newly incoming labeled data to improve learning performance. Experimental evaluations on seven IoT datasets show that the proposed method outperforms semi-supervised methods in accuracy, precision, and geometric mean (Gmean) by 10% and 5% over supervised methods, respectively. Jun Jiang 0003, Bin Wang 0048, Quan Tang 0001, Guoxiang Zhong, Xuhao Tang 0001, Joel J. P. C. Rodrigues |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2024 | EK-Net: Real-Time Scene Text Detection with Expand Kernel DistanceabstractRecently, scene text detection has received significant attention due to its wide application. However, accurate detection in complex scenes of multiple scales, orientations, and curvature remains a challenge. Numerous detection methods adopt the Vatti clipping (VC) algorithm for multiple-instance training to address the issue of arbitrary-shaped text. Yet we identify several bias results from these approaches called the "shrinked kernel". Specifically, it refers to a decrease in accuracy resulting from an output that overly favors the text kernel. In this paper, we propose a new approach named Expand Kernel Network (EK-Net) with expand kernel distance to compensate for the previous deficiency, which includes three-stages regression to complete instance detection. Moreover, EK-Net not only realize the precise positioning of arbitrary-shaped text, but also achieve a trade-off between performance and speed. Evaluation results demonstrate that EK-Net achieves state-of-the-art or competitive performance compared to other advanced methods, e.g., F-measure of 85.72% at 35.42 FPS on ICDAR 2015, F-measure of 85.75% at 40.13 FPS on CTW1500. Boyuan Zhu, Fagui Liu, Quan Tang 0001 |
ICASSP | 4 |
| 2024 | ACP-Net: Asymmetric Center Positioning Network for Real-Time Text Detection
Boyuan Zhu, Fagui Liu, Quan Tang 0001, C. L. Philip Chen |
Knowl. Based Syst. | 4 |
| 2023 | Dynamic Token Pruning in Plain Vision Transformers for Semantic SegmentationabstractVision transformers have achieved leading performance on various visual tasks yet still suffer from high computational complexity. The situation deteriorates in dense prediction tasks like semantic segmentation, as high-resolution inputs and outputs usually imply more tokens involved in computations. Directly removing the less attentive tokens has been discussed for the image classification task but can not be extended to semantic segmentation since a dense prediction is required for every patch. To this end, this work introduces a Dynamic Token Pruning (DToP) method based on the early exit of tokens for semantic segmentation. Motivated by the coarse-to-fine segmentation process by humans, we naturally split the widely adopted auxiliary-loss-based network architecture into several stages, where each auxiliary block grades every token’s difficulty level. We can finalize the prediction of easy tokens in advance without completing the entire forward pass. Moreover, we keep k highest confidence tokens for each semantic category to uphold the representative context information. Thus, computational complexity will change with the difficulty of the input, akin to the way humans do segmentation. Experiments suggest that the proposed DToP architecture reduces on average 20% ∼ 35% of computational cost for current semantic segmentation methods based on plain vision transformers without accuracy degradation. The code is available through the following link: https://github.com/zbwxp/Dynamic-Token-Pruning. Quan Tang 0001, Bowen Zhang 0009, Jiajun Liu 0004, Fagui Liu, Yifan Liu 0001 |
ICCV | 1 |
| 2023 | Boosting Semantic Segmentation from the Perspective of Explicit Class EmbeddingsabstractSemantic segmentation is a computer vision task that associates a label with each pixel in an image. Modern approaches tend to introduce class embeddings into semantic segmentation for deeply utilizing category semantics, and regard supervised class masks as final predictions. In this paper, we explore the mechanism of class embeddings and have an insight that more explicit and meaningful class embeddings can be generated based on class masks purposely. Following this observation, we propose ECENet, a new segmentation paradigm, in which class embeddings are obtained and enhanced explicitly during interacting with multi-stage image features. Based on this, we revisit the traditional decoding process and explore inverted information flow between segmentation masks and class embeddings. Furthermore, to ensure the discriminability and informativity of features from backbone, we propose a Feature Reconstruction module, which combines intrinsic and diverse branches together to ensure the concurrence of diversity and redundancy in features. Experiments show that our ECENet outperforms its counterparts on the ADE20K dataset with much less computational cost and achieves new state-of-the-art results on PASCALContext dataset. The code will be released at https://gitee.com/mindspore/models and https://github.com/Carol-lyh/ECENet. Yuhe Liu, Chuanjian Liu, Kai Han 0002, Quan Tang 0001, Zengchang Qin |
ICCV | 4 |
| 2023 | AERF: Adaptive ensemble random fuzzy algorithm for anomaly detection in cloud computing
Jun Jiang 0003, Fagui Liu, Wing W. Y. Ng, Quan Tang 0001, Guoxiang Zhong, Xuhao Tang 0001, Bin Wang 0048 |
Comput. Commun. | 4 |
| 2022 | Alleviating Overconfident Failure Predictions via Masking Predictive Logits in Semantic Segmentation
Quan Tang 0001, Fagui Liu, Jun Jiang 0003, Yu Zhang 0144, Xuhao Tang 0001 |
ICANN (2) | 1 |
| 2022 | Utilize Spatial Prior in Ground Truth: Spatial-Enhanced Loss for Semantic Segmentation
Yu Zhang 0144, Fagui Liu, Quan Tang 0001 |
ICANN (3) | 3 |
| 2022 | SegViT: Semantic Segmentation with Plain Vision TransformersabstractWe explore the capability of plain Vision Transformers (ViTs) for semantic segmentation and propose the SegViT. Previous ViT-based segmentation networks usually learn a pixel-level representation from the output of the ViT. Differently, we make use of the fundamental component—attention mechanism, to generate masks for semantic segmentation. Specifically, we propose the Attention-to-Mask (ATM) module, in which the similarity maps between a set of learnable class tokens and the spatial feature maps are transferred to the segmentation masks. Experiments show that our proposed SegViT using the ATM module outperforms its counterparts using the plain ViT backbone on the ADE20K dataset and achieves new state-of-the-art performance on COCO-Stuff-10K and PASCAL-Context datasets. Furthermore, to reduce the computational cost of the ViT backbone, we propose query-based down-sampling (QD) and query-based up-sampling (QU) to build a Shrunk structure. With our Shrunk structure, the model can save up to 40% computations while maintaining competitive performance. Bowen Zhang 0009, Zhi Tian, Quan Tang 0001, Xiangxiang Chu, Xiaolin Wei, Chunhua Shen, Yifan Liu 0001 |
NeurIPS | 3 |
| 2022 | A dynamic ensemble algorithm for anomaly detection in IoT imbalanced data streams
Jun Jiang 0003, Fagui Liu, Yongheng Liu, Quan Tang 0001, Bin Wang 0048, Guoxiang Zhong, Weizheng Wang 0001 |
Comput. Commun. | 4 |
| 2022 | EPRNet: Efficient Pyramid Representation Network for Real-Time Street Scene SegmentationabstractCurrent scene segmentation methods suffer from cumbersome model structures and high computational complexity, impeding their applications to real-world scenarios that require real-time processing. This paper proposes a novel Efficient Pyramid Representation Network (EPRNet), which strikes an innovative record on segmentation accuracy, model lightness and inference efficiency. Unlike existing methods delivering transfer learning based on pixel features of limited receptive fields encoded by shallow image classification backbones, EPRNet distributes multi-scale representations throughout the feature encoding flow to quickly enlarge and enrich receptive fields. Specifically, we introduce an extremely lightweight and efficient Multi-scale Processing Unit (MPU) that encodes multi-scale features through parallel convolutions of different kernels. By combining MPU and residual learning, we propose a core Pyramid Representation Module (PRM) to correctly acquire and aggregate region-based contexts in both shallow and deep layers. In this way, EPRNet can encode discriminative and comprehensive representations of multi-scale objects with a compact structure. We conduct extensive experiments on Cityscapes and CamVid datasets, demonstrating the superiority. Without any extra and coarse labeled data, EPRNet obtains mIoU 73.9% on the Cityscapes test set with only 0.9 million parameters at a speed of 42 FPS. Quan Tang 0001, Fagui Liu, Jun Jiang 0003, Yu Zhang 0144 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Compensating for Local Ambiguity With Encoder-Decoder in Urban Scene SegmentationabstractSemantic segmentation plays a critical role in scene understanding for self-driving vehicles. A line of efforts has proven that global context matters in urban scene segmentation due to massive scale changes. However, we find that existing methods suffer from local ambiguities when dissipating continuous local context, i.e. scrambling to a huge receptive field of global cues by coarse pooling. To this end, this paper proposes a new Context Aggregation Module (CAM) that consists of two primary components: context encoding using no coarse pooling but encoder-decoders with appropriate sampling scales and gated fusion that extends gate attention mechanism to balance different-scale context during feature fusion. Weeding out coarse pooling and applying the encoder-decoder inherits the merits of exploring global context while avoiding the drawback of losing local contextual continuity. We then construct a Context Aggregation Network (CANet) and conduct extensive evaluations on challenging autonomous driving benchmarks of Cityscapes, CamVid and BDD100K. Consistently improved results evidence the effectiveness. Notably, we attain competitive mIoU 82.7% on Cityscapes and optimal mIoU 80.5% on CamVid. Quan Tang 0001, Fagui Liu, Tong Zhang 0015, Jun Jiang 0003, Yu Zhang 0144, Boyuan Zhu, Xuhao Tang 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Attention-guided chained context aggregation for semantic segmentation
Quan Tang 0001, Fagui Liu, Tong Zhang 0015, Jun Jiang 0003, Yu Zhang 0144 |
Image Vis. Comput. | 1 |