EDBT 2026 Demo / reviewers in the wild / expert
Qiong Wang 0003
dblp:65/3144-3
· DBLP profile ↗
25ranked-venue papers
3as first author
20since 2021 · last 2026
0000-0003-4193-0960ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 2 first-author · 17 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Monocular 3D lane detection with geometry-guided transformation and contextual enhancement
Chunying Song, Qiong Wang 0003, Zeren Sun, Huafeng Liu 0004 |
Pattern Recognit. Lett. | 2 |
| 2025 | DynaPlane-Lane: Dynamic Multi-Plane Geometry Learning for Robust Monocular 3D Lane DetectionabstractAccurate monocular 3D lane detection remains a fundamental challenge in autonomous driving, due to the need to infer spatial geometry from single-view images under diverse road and environmental conditions. Existing approaches often struggle with limited geometric adaptability and inconsistent structural predictions, particularly in the presence of sloped or uneven road surfaces. In this paper, we propose DynaPlane-Lane, a geometry-aware framework that explicitly models road topology through a dynamic multi-plane representation. By partitioning the scene into near, transition, and far regions with learnable geometric transformations, our method effectively captures height changes and surface irregularities. To enhance spatial reasoning, we design a Multi-Scale Adaptive Feature Aggregation module with channel and spatial attention to fuse encoder features across multiple resolutions. Additionally, a Topology-Aware Lane Optimization objective is formulated to enforce continuity and smoothness of predicted lane curves under partial visibility. Comprehensive experiments on the OpenLane and Apollo 3D Synthetic benchmarks demonstrate that our approach achieves robust performance across diverse scenarios, validating the effectiveness of incorporating adaptive geometry and structural constraints into monocular 3D lane detection. Chunying Song, Huafeng Liu 0004, Tao Chen 0012, Qiong Wang 0003 |
MMAsia | 4 |
| 2025 | PASS-SAM: Integration of Segment Anything Model for Large-Scale Unsupervised Semantic SegmentationabstractLarge-scale unsupervised semantic segmentation (LUSS) is a sophisticated process that aims to segment similar areas within an image without relying on labeled training data. While existing methodologies have made substantial progress in this area, there is ample scope for enhancement. We thus introduce the PASS-SAM model, a comprehensive solution that amalgamates the benefits of various models to improve segmentation performance. Specifically, we enhance a baseline model utilizing self-attention and external attention modules. In the fine-tuning phase, we make use of conditional random fields (CRF) and the segment anything model (SAM) to refine and retrain the baseline model. During inferencing, we employ a model ensemble to blend predictions from different models, thereby enhancing segmentation accuracy. This approach secured first place in the LUSS track of the Third Jittor Artificial Intelligence Challenge. Our model, which makes use of the Jittor framework, is publicly available at https://github.com/PGSmall/jittor-PGSmall-LUSS. Gensheng Pei, Qiong Wang 0003 |
Comput. Vis. Media | 4 |
| 2025 | UncertainBEV: Uncertainty-aware BEV fusion for roadside 3D object detection
Chunying Song, Huafeng Liu 0004, Qiong Wang 0003 |
Image Vis. Comput. | 5 |
| 2024 | SMP-Track: SAM in Multi-Pedestrian TrackingabstractMultiple Object Tracking (MOT) plays a crucial role in security data analysis as a fundamental problem for video surveillance. Our goal is to design a robust tracker for data damaged by attacks, while also emphasizing privacy protection. The mainstream paradigm for MOT is tracking-by-detection (TBD), which involves object detection followed by target association. In the association stage, most models rely on Intersection over Union (IoU) similarity of bounding boxes for short-range matching and cosine similarity of appearance features for long-range matching. However, both of these similarities contain a lot of redundant background regions except the target. To this end, we propose a new tracker named SMP-Track that integrates Segment Anything Model (SAM) into Multi-Pedestrian tracking method. Firstly we extract the pedestrian masks based on box prompt, focusing solely on the foreground information. Then we introduce a new similarity metric that combines the advantages of motion and foreground information (i.e., box-mask similarity). Extensive experiments demonstrate that SMP-Track increases main metrics on the MOT17 validation set, and achieves comparable performance to other state-of-the-art methods on the MOT17 and MOT20 test sets. Furthermore, by incorporating pedestrian masks, we reduce reliance on raw pedestrian images or features, making the model robust to corrupted data and mitigating the risk of privacy leakage. Shiyin Wang, Huafeng Liu 0004, Qiong Wang 0003, Yazhou Yao |
DSAA | 3 |
| 2024 | Universal Organizer of Segment Anything Model for Unsupervised Semantic SegmentationabstractUnsupervised semantic segmentation (USS) aims to achieve high-quality segmentation without manual pixel-level annotations. Existing USS models provide coarse category classifi-cation for regions, but the results often have blurry and imprecise edges. Recently, a robust framework called the segment anything model (SAM) has been proven to deliver precise boundary object masks. Therefore, this paper proposes a universal organizer based on SAM, termed as UO-SAM, to enhance the mask quality of USS models. Specifically, using only the original image and the masks generated by the USS model, we extract visual features to obtain positional prompts for target objects. Then, we activate a local region optimizer that performs segmentation using SAM on a per-object basis. Finally, we employ a global region optimizer to incorporate global image information and refine the masks to obtain the final fine-grained masks. Compared to existing methods, our UO-SAM achieves state-of-the-art performance. Our codes are available at https://github.com/NUST-Machine-Intelligence-Laboratory/UO-SAM. Gensheng Pei, Xinhao Cai, Qiong Wang 0003, Huafeng Liu 0004, Yazhou Yao |
ICME | 4 |
| 2023 | Semi-Supervised Semantic Segmentation With Region RelevanceabstractSemi-supervised semantic segmentation aims to learn from a small amount of labeled data and plenty of unlabeled ones for the segmentation task. The most common approach is to generate pseudo-labels for unlabeled images to augment the training data. However, the noisy pseudo-labels will lead to cumulative classification errors and aggravate the local inconsistency in prediction. This paper proposes a Region Relevance Network (RRN) to alleviate the problem mentioned above. Specifically, we first introduce a local pseudo-label filtering module that leverages discriminator networks to assess the accuracy of the pseudo-label at the region level. A local selection loss is proposed to mitigate the negative impact of wrong pseudo-labels in consistency regularization training. In addition, we propose a dynamic region-loss correction module, which takes the merit of network diversity to further rate the reliability of pseudo-labels and correct the convergence direction of the segmentation network with a dynamic region loss. Extensive experiments are conducted on PASCAL VOC 2012 and Cityscapes datasets with varying amounts of labeled data, demonstrating that our proposed approach achieves state-of-the-art performance compared to current counterparts. Our code is available at: https://github.com/NUST-Machine-Intelligence-Laboratory/TorchSemiSeg2. Tao Chen 0012, Qiong Wang 0003, Yazhou Yao |
ICME | 3 |
| 2023 | Saliency Guided Inter- and Intra-Class Relation Constraints for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation with only image-level labels aims to reduce annotation costs for the segmentation task. Existing approaches generally leverage class activation maps (CAMs) to locate the object regions for pseudo label generation. However, CAMs can only discover the most discriminative parts of objects, thus leading to inferior pixel-level pseudo labels. To address this issue, we propose a saliency guidedInter- andIntra-ClassRelationConstrained (I$^{2}$CRC) framework to assist the expansion of the activated object regions in CAMs. Specifically, we propose a saliency guided class-agnostic distance module to pull the intra-category features closer by aligning features to their class prototypes. Further, we propose a class-specific distance module to push the inter-class features apart and encourage the object region to have a higher activation than the background. Besides strengthening the capability of the classification network to activate more integral object regions in CAMs, we also introduce an object guided label refinement module to take a full use of both the segmentation prediction and the initial labels for obtaining superior pseudo-labels. Extensive experiments on PASCAL VOC 2012 and COCO datasets demonstrate well the effectiveness of I$^{2}$CRC over other state-of-the-art counterparts. Tao Chen 0012, Yazhou Yao, Lei Zhang 0054, Qiong Wang 0003, Guosen Xie, Fumin Shen |
IEEE Trans. Multim. | 4 |
| 2023 | FECANet: Boosting Few-Shot Semantic Segmentation With Feature-Enhanced Context-Aware NetworkabstractFew-shot semantic segmentation is the task of learning to locate each pixel of the novel class in the query image with only a few annotated support images. The current correlation-based methods construct pair-wise feature correlations to establish the many-to-many matching because the typical prototype-based approaches cannot learn fine-grained correspondence relations. However, the existing methods still suffer from the noise contained in naive correlations and the lack of context semantic information in correlations. To alleviate these problems mentioned above, we propose a Feature-Enhanced Context-Aware Network (FECANet). Specifically, a feature enhancement module is proposed to suppress the matching noise caused by inter-class local similarity and enhance the intra-class relevance in the naive correlation. In addition, we propose a novel correlation reconstruction module that encodes extra correspondence relations between foreground and background and multi-scale context semantic features, significantly boosting the encoder to capture a reliable matching pattern. Experiments on PASCAL-$5^{i}$and COCO-$20^{i}$datasets demonstrate that our proposed FECANet leads to remarkable improvement compared to previous state-of-the-arts, demonstrating its effectiveness. The source codes and models have been made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/FECANET. Huafeng Liu 0004, Tao Chen 0012, Qiong Wang 0003, Yazhou Yao, Xian-Sheng Hua 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Guided by Meta-Set: A Data-Driven Method for Fine-Grained Visual RecognitionabstractThe lack of sufficient training data has been one obstacle to fine-grained visual classification research because labeling subcategories generally requires specialist knowledge. As one optional approach to alleviating the data-hunger problem, leveraging web images as training data is drawing increasing attention. Nevertheless, web images potentially have false labels, which can misguide the training process. Although several works have been proposed to deal with label noise, it still can be difficult for the network to tackle complex real-world noisy labels without any prior knowledge. In the literature, we propose to leverage a small and clean meta-set to provide reliable prior knowledge for tackling noisy web images. Specifically, our method trains a network with two peer predicting heads, which learn from noisy web images (web head) and meta ones (meta head), respectively. The meta head produces pseudo soft labels for web images to revise their training loss, which can overcome the high noise ratio problem. Furthermore, a selection net is trained in a meta-learning strategy to identify in- and out-of-distribution noisy images. Then in-distribution ones are reused for training with pseudo soft labels produced by the meta head as supervision, while out-of-distribution ones are discarded. In this manner, the misguidance caused by label noise is remarkably alleviated and in-distribution noisy samples are properly exploited to boost model performance. The superiority of our proposed approach is demonstrated by mathematical theory with great interpretability as well as extensive experimental results on the real-world dataset WebFG-496. Chuanyi Zhang, Guosheng Lin, Qiong Wang 0003, Fumin Shen, Yazhou Yao, Zhenmin Tang |
IEEE Trans. Multim. | 3 |
| 2022 | PNP: Robust Learning from Noisy Labels by Probabilistic Noise PredictionabstractLabel noise has been a practical challenge in deep learning due to the strong capability of deep neural networks in fitting all training data. Prior literature primarily resorts to sample selection methods for combating noisy labels. However, these approaches focus on dividing samples by order sorting or threshold selection, inevitably introducing hyperparameters (e.g., selection ratio / threshold) that are hard-to-tune and dataset-dependent. To this end, we propose a simple yet effective approach named PNP (Probabilistic Noise Prediction) to explicitly model label noise. Specifically, we simultaneously train two networks, in which one predicts the category label and the other predicts the noise type. By predicting label noise probabilistically, we identify noisy samples and adopt dedicated optimization objectives accordingly. Finally, we establish a joint loss for network update by unifying the classification loss, the auxiliary constraint loss, and the in-distribution consistency loss. Comprehensive experimental results on synthetic and realworld datasets demonstrate the superiority of our proposed method. The source code and models have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/PNP. Zeren Sun, Fumin Shen, Qiong Wang 0003, Xiangbo Shu, Yazhou Yao, Jinhui Tang 0001 |
CVPR | 4 |
| 2022 | Multi-scale Self-attention-based Few-shot Object Detection for Remote Sensing ImagesabstractFor object detection on Remote Sensing Images (RSI), numerous methods based on deep convolutional neural networks have been developed by researchers(CNN) and and re-markable achievements have been made in detection performance and efficiency. Current CNN-based methods usually require a large number of annotated samples for training. However, labeling RSI is time-consuming, making it difficult to obtain large-scale annotated training samples. In this paper, we introduce a transfer learning-based method for few-shot object detection on RSI. In our method, only a few annotated samples are required for unseen classes. More specifically, our model adopts a two-stage fine-tuning scheme and contains two modules: a multi-scale self-attention module and a copy-paste with diminishing edge transparency module. Our design enables the model to learn transferable knowledge from seen classes and generalizes well to unseen classes. Experiments on two benchmark datasets demonstrate the effectiveness of our proposed method in few-shot object detection for RSI. Qiong Wang 0003, Jiaxing Tong |
MMSP | 2 |
| 2022 | Unsupervised Pre-training for 3D Object Detection with Transformer
Maosheng Sun, Xiaoshui Huang, Zeren Sun, Qiong Wang 0003, Yazhou Yao |
PRCV (3) | 4 |
| 2022 | Few-Shot Object Detection via Understanding Convolution and Attention
Jiaxing Tong, Tao Chen 0012, Qiong Wang 0003, Yazhou Yao |
PRCV (1) | 3 |
| 2022 | DBFC-Net: a uniform framework for fine-grained cross-media retrieval
Qiong Wang 0003, Youdong Guo, Yazhou Yao |
Multim. Syst. | 1 |
| 2022 | Enhanced Feature Alignment for Unsupervised Domain Adaptation of Semantic SegmentationabstractUnsupervised domain adaptation for semantic segmentation aims to transfer knowledge from a labeled source domain to another unlabeled target domain. However, due to the label noise and domain mismatch, learning directly from source domain data tends to have poor performance. Though adversarial learning methods strive to reduce domain discrepancies by aligning feature distributions, traditional methods suffer from the training imbalance and feature distortion problems. Besides, due to the absence of target domain labels, the classifier is blind to features from the target domain during training. Consequently, the final classifier overfits the source domain features and usually fails to predict the structured outputs of the target domain. To alleviate these problems, we focus on enhancing the adversarial learning based feature alignment from three perspectives. First, a classification constrained discriminator is proposed to balance the adversarial training and alleviate the feature distortion problem. Next, to alleviate the classifier overfitting problem, self-training is collaboratively used to learn a domain robust classifier with target domain pseudo labels. Moreover, an efficient class centroid calculation module is proposed and the domain discrepancy is further reduced by aligning the feature centroids of the same class from different domains. Experimental evaluations on GTA5$\rightarrow$Cityscapes and SYNTHIA$\rightarrow$Cityscapes demonstrate state-of-the-art results compared to other counterpart methods. The source code and models have been made available at.11[Online]. Available:https://github.com/NUST-Machine-Intelligence-Laboratory/EFA. Tao Chen 0012, Shuihua Wang, Qiong Wang 0003, Zheng Zhang 0006, Guosen Xie, Zhenmin Tang |
IEEE Trans. Multim. | 3 |
| 2022 | Semantically Meaningful Class Prototype Learning for One-Shot Image SegmentationabstractOne-shot semantic image segmentation aims to segment the object regions for the novel class with only one annotated image. Recent works adopt the episodic training strategy to mimic the expected situation at testing time. However, these existing approaches simulate the test conditions too strictly during the training process, and thus cannot make full use of the given label information. Besides, these approaches mainly focus on the foreground-background target class segmentation setting. They only utilize binary mask labels for training. In this paper, we propose to leverage the multi-class label information during the episodic training. It will encourage the network to generate more semantically meaningful features for each category. After integrating the target class cues into the query features, we then propose a pyramid feature fusion module to mine the fused features for the final classifier. Furthermore, to take more advantage of the support image-mask pair, we propose a self-prototype guidance branch to support image segmentation. It can constrain the network for generating more compact features and a robust prototype for each semantic class. For inference, we propose a fused prototype guidance branch for the segmentation of the query image. Specifically, we leverage the prediction of the query image to extract the pseudo-prototype and combine it with the initial prototype. Then we utilize the fused prototype to guide the final segmentation of the query image. Extensive experiments demonstrate the superiority of our proposed approach. The source codes and models have been made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/SMCP. Tao Chen 0012, Guosen Xie, Yazhou Yao, Qiong Wang 0003, Fumin Shen, Zhenmin Tang, Jian Zhang 0002 |
IEEE Trans. Multim. | 4 |
| 2022 | Co-LDL: A Co-Training-Based Label Distribution Learning Method for Tackling Label NoiseabstractPerformances of deep neural networks are prone to be degraded by label noise due to their powerful capability in fitting training data. Deeming low-loss instances as clean data is one of the most promising strategies in tackling label noise and has been widely adopted by state-of-the-art methods. However, prior works tend to drop high-loss instances directly, neglecting their valuable information. To address this issue, we propose an end-to-end framework named Co-LDL, which incorporates the low-loss sample selection strategy with label distribution learning. Specifically, we simultaneously train two deep neural networks and let them communicate useful knowledge by selecting low-loss and high-loss samples for each other. Low-loss samples are leveraged conventionally for updating network parameters. On the contrary, high-loss samples are trained in a label distribution learning manner to update network parameters and label distributions concurrently. Moreover, we propose a self-supervised module to further boost the model performance by enhancing the learned representations. Comprehensive experiments on both synthetic and real-world noisy datasets are provided to demonstrate the superiority of our Co-LDL method over state-of-the-art approaches in learning with noisy labels. The source code and models have been made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/CoLDL. Zeren Sun, Huafeng Liu 0004, Qiong Wang 0003, Tianfei Zhou, Qi Wu 0001, Zhenmin Tang |
IEEE Trans. Multim. | 3 |
| 2022 | Robust Learning From Noisy Web Images Via Data Purification for Fine-Grained RecognitionabstractManually labeling fine-grained datasetsis laborious and typically requires domain-specific expert knowledge. Conversely, a vast amount of web data is relatively easy to obtain with nearly no human effort. Therefore, learning from noisy web data for fine-grained tasks is attracting increasing attention in recent years. However, the presence of noise in web images is a huge obstacle for training robust fine-grained recognition models. To this end, we propose a novel approach to identify noisy images as well as specifically distinguish in- and out-of-distribution samples. It can purify the noisy web training set by discarding out-of-distribution noise and relabeling in-distribution noisy samples. Then we can train the model on the purified dataset to alleviate the harmful effects of noise and make the most of web images to achieve better performance. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is far superior to current state-of-the-art web-supervised methods. The data and source code of this work have been made publicly available at:https://github.com/NUST-Machine-Intelligence-Laboratory/Dataset-Purification. Chuanyi Zhang, Qiong Wang 0003, Guosen Xie, Qi Wu 0001, Fumin Shen, Zhenmin Tang |
IEEE Trans. Multim. | 2 |
| 2021 | Local Self-Attention on Fine-grained Cross-media RetrievalabstractDue to the heterogeneity gap, the data representation of different media is inconsistent and belongs to different feature spaces. Therefore, it is challenging to measure the fine-grained gap between them. To this end, we propose an attention space training method to learn common representations of different media data. Specifically, we utilize local self-attention layers to learn the common attention space between different media data. We propose a similarity concatenation method to understand the content relationship between features. To further improve the robustness of the model, we also train a local position encoding to capture the spatial relationships between features. In this way, our proposed method can effectively reduce the gap between different feature distributions on cross-media retrieval tasks. It also improves the fine-grained recognition performance by attaching attention to high-level semantic information. Extensive experiments and ablation studies demonstrate that our proposed method achieves state-of-the-art performance. At the same time, our approach provides a new pipeline for fine-grained cross-media retrieval. The source code and models are publicly available at: https://github.com/NUST-Machine-Intelligence-Laboratory/SAFGCMHN. Yazhou Yao, Qiong Wang 0003, Zhenmin Tang |
MMAsia | 3 |
| 2020 | Multi-model Network for Fine-Grained Cross-Media Retrieval
Jiemi Bai, Yazhou Yao, Qiong Wang 0003, Wankou Yang, Fumin Shen |
PRCV (2) | 3 |
| 2019 | Exploiting textual and visual features for image categorization
Yazhou Yao, Wankou Yang, Qiong Wang 0003, Yunfei Cai, Zhenmin Tang |
Pattern Recognit. Lett. | 4 |
| 2016 | Automated classification of brain images using wavelet-energy and biogeography-based optimization
Gelan Yang, Yudong Zhang 0001, Jiquan Yang, Genlin Ji, Zhengchao Dong, Shuihua Wang, Chunmei Feng, Qiong Wang 0003 |
Multim. Tools Appl. | 8 |
| 2009 | Robust Facial Feature Location on Gray Intensity Face
Qiong Wang 0003, Chunxia Zhao, Jing-Yu Yang 0001 |
PSIVT | 1 |
| 2006 | Face Detection Using Binary Template Matching and SVM
Qiong Wang 0003, Wankou Yang, Huan Wang 0013, Jing-Yu Yang 0001, Yu-Jie Zheng |
PRICAI | 1 |