Rong Quan

dblp:191/4667 · DBLP profile ↗
← Back
27ranked-venue papers
4as first author
24since 2021 · last 2026
0000-0002-1494-6193ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 4 first-author · 19 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Dual-view Driver Gaze Estimation via Mutual Enhancement
Chengzheng Fu, Rong Quan, Siyu Chen 0004, Yiming Ni, Jie Qin 0004
ICMR2
2026 Object Referring-Guided Scanpath Prediction with Perception-Enhanced Vision-Language Models
Rong Quan, Yantao Lai, Dong Liang 0008, Jie Qin 0004
ICMR1
2026 Frequency-aware optimization for floating artifact removal in 3D Gaussian splatting
Cen Li, Rong Quan
Vis. Comput.4
2025 Doubly Contrastive Learning for Source-Free Domain Adaptive Person Search
abstract
Domain Adaptive Person Search (DAPS) aims to improve the generalization capability of person search models by training on both labeled source data and unlabeled target data, which is not that practical in real-world applications considering the storage/transmission costs and the privacy of source data. In this paper, we investigate a more practical and efficient person search setting, Source-Free Domain Adaptive Person Search (SFDA-PS), which seeks to generalize an existing source person search model to any unseen domain without requiring source data. Considering the absence of effective annotations in SFDA-PS, we propose a Doubly Contrastive Learning (DCL) method to adapt the target domain knowledge to the source model in a mutual learning and contrastive learning way. Specifically, we employ a mutual learning-based mean-teacher model as our baseline to incorporate target domain knowledge by pursuing prediction consistency between the teacher and student. Then, a Relation-embedded Contrastive (ReC) learning strategy is introduced to the detection head to ensure semantic consistency among proposals related to the same person while maintaining semantic distinction among proposals from different categories or persons. Furthermore, a Memory-aided Constrative (MaC) learning strategy is integrated into the re-identification (Re-ID) head to enhance its discriminative capability on target person embeddings. Extensive experiments on existing state-of-the-art person search models and two widely used benchmarks demonstrate the superiority of the proposed SFDA-PS task, as well as our proposed DCL.
Rong Quan, Haiyan Chen 0001, Jie Qin 0004
AAAI2
2025 MarkEditor: Precise and Controllable 3D Gaussian Editing with Text-Driven Optimization
abstract
3D Gaussian Splatting (3DGS) has gained significant attention in the 3D field; however, achieving precise and controllable editing remains a challenging problem. This paper introduces MarkEditor, a framework for precise and localized editing of 3D Gaussians. It uses a text-driven approach, beginning with a marked region as a coarse editing guide. The process consists of three stages. First, we use text-to-3D model to generate an initial 3DGS from a textual description. Second, we align it with the marked area through geometric optimization. Third, we employ a freezing mechanism and utilize Score Distillation Sampling (SDS) for global optimization of the appearance, while updating only the content within the marked region, resulting in the final edited 3D Gaussian scene. MarkEditor provides a precise and controllable approach for 3D Gaussians editing, achieving results that outperform state-of-the-art methods.
Cen Li, Hanfei Hu, Rong Quan
CW5
2025 Low-Frequency First: Eliminating Floating Artifacts in 3D Gaussian Splatting
abstract
3D Gaussian Splatting (3DGS) is a powerful and computationally efficient representation for 3D reconstruction. Despite its strengths, 3DGS often produces floating artifacts, which are erroneous structures detached from the actual geometry and significantly degrade visual fidelity. The underlying mechanisms causing these artifacts, particularly in low-quality initialization scenarios, have not been fully explored. In this paper, we investigate the origins of floating artifacts from a frequency-domain perspective and identify under-optimized Gaussians as the primary source. Based on our analysis, we propose Eliminating-Floating-Artifacts Gaussian Splatting (EFA-GS), which selectively expands under-optimized Gaussians to prioritize accurate low-frequency learning. Additionally, we introduce complementary depth-based and scale-based strategies to dynamically refine Gaussian expansion, effectively mitigating detail erosion. Extensive experiments on both synthetic and real-world datasets demonstrate that EFA-GS substantially reduces floating artifacts while preserving high-frequency details, achieving an improvement of$\text{1. 6 8 ~ d B}$in PSNR over baseline method on our RWLQ dataset. Furthermore, we validate the effectiveness of our approach in downstream 3D editing tasks. We provide our implementation in https://jcwang-gh.github.io/EFA-GS.
Cen Li, Rong Quan
CW4
2025 BEVFog: Enhancing Vision-Based Roadside 3D Object Detection Robustness Under Foggy Conditions
abstract
Roadside cameras play a crucial role in extending the perception range of intelligent transportation infrastructures, offering a cost-effective means for long-range 3D object detection. However, under foggy conditions, the visibility degradation severely diminishes the performance of vision-based roadside perception systems, as reduced contrast and distorted monocular cues impair bird’s-eye-view (BEV) depth estimation. To address this critical challenge, we propose BEVFog, a novel end-to-end framework that simultaneously restores fog-degraded semantic features and leverages fog density as an auxiliary depth prior to enhance detection robustness. BEVFog integrates two key components: a lightweight, region-aware defogging module with CNN-predicted, differentiable filters, and a depth adjustment module that uses Fog-Cue Conditional Kernels (FCK) to modulate BEV features based on local scattering, correcting depth distortions and enhancing geometric cues. To facilitate rigorous evaluation, we construct Foggy DAIR-V2X-I, a synthetic benchmark that introduces ten fog levels into the large-scale DAIR-V2X-I roadside dataset using an atmospheric scattering model. Extensive experiments demonstrate that BEVFog narrows the performance gap between heavy-fog and clear-weather conditions from over 10% AP to less than 2% AP for vehicles, while achieving up to 8% AP improvements for pedestrians and cyclists compared to the best dehazing-plus-detection baselines. This work paves the way for safer, all-weather autonomous infrastructure perception, and highlights the importance of integrating low-level restoration with high-level 3D reasoning in adverse conditions.
Gaoyuan Miao, Wentong Li 0001, Rong Quan, Jie Qin 0004
ECAI3
2025 CLIPGaze: Zero-Shot Goal-Directed Scanpath Prediction Using CLIP
abstract
Goal-directed scanpath prediction aims to predict people’s gaze shift path when searching for objects in a visual scene. Most existing goal-directed scanpath prediction methods cannot generalize to target classes not present during training. Besides, they usually exploit different pre-trained models to extract features for the target prompt and image, resulting in big feature gap and making the subsequent feature matching and fusion very difficult. To solve the above problems, we propose a novel zero-shot goal-directed scanpath prediction model named CLIPGaze. We use CLIP to extract pre-matched features for the target prompt and input image, making the feature fusion easier to receive. Using large model like CLIP can also enhance the whole model’s generalization ability on target classes not present during training. We propose a hierarchical visual-semantic feature fusion module to fuse the target and image features more comprehensively. Furthermore, due to the limited number of classes in goal-directed scanpath dataset, we employ image segmentation as a proxy task to help train the feature fusion module, significantly enhancing our model’s performance in zeroshot setting. Extensive experiments demonstrate the effectiveness of our method on both seen and unseen target classes.
Yantao Lai, Rong Quan, Dong Liang 0008, Jie Qin 0004
ICASSP2
2025 DCN: Decoupled-Coupled Network for Text-based Person Search
abstract
Text-based person search aims to identify a person based on textual descriptions, by simultaneously addressing person detection and cross-modal alignment between text queries and person images. Existing approaches often struggle with conflicts in exploiting proposals across these two sub-tasks. Specifically, cross-modal alignment requires highly precise proposals, while person detection can tolerate a certain degree of proposal inaccuracy but always needs a large number of proposals. In this paper, we propose the Decoupled-Coupled Network (DCN) to tackle the above conflicts. We first attempt to resolve the above conflicts by proposing a Decoupled Proposal Selection (DPS) strategy, inspired by the divide-and-conquer principle. DPS adaptively selects the most suitable proposals for each sub-task, ensuring their distinct requirements are adequately met. We further present a Coupled Cascade Refinement (CCR) module to jointly optimize both sub-tasks in a multi-stage manner, progressively improving detection accuracy and fostering cross-modal alignment between text and person image. In addition, we introduce two types of objective functions to optimize the inherently multi-positive contrastive learning challenge. Extensive experiments conducted on two benchmarks demonstrate the effectiveness and superiority of our DCN over existing competitors.
Rong Quan, Liangxu Su, Wentong Li 0001, Yichao Yan, Jie Qin 0004
MMAsia2
2025 Effective Text-Directed Scanpath Prediction via Comprehensive Multi-modal Information Fusion
Rong Quan
PRCV (9)2
2025 Structurally and Semantically Guided Human-Object Interaction Detection
Lele Sun, Rong Quan
PRCV (17)2
2025 Large Models are Good Annotators for Zero-Shot Learning
abstract
Human-annotated attributes serve as effective semantic label embeddings for zero-shot learning (ZSL); however, their annotation is labor-intensive and difficult to scale. Recent studies have explored weakly supervised semantic label embeddings to reduce human effort, but these methods often fail to capture visual similarity and underperform compared to human-annotated semantics. In this work, we propose a minimally supervised yet effective approach: GPT- and CLIP-powered attributes (GCAtt). Specifically, we introduce a three-step interaction process with ChatGPT-comprising preliminary design, hierarchical refinement, and specific value determination-to generate attributes that are both category-shared and discriminative for classification. Additionally, we develop a method that encodes attributes and their values as potential text pairings, leveraging CLIP's retrieval capabilities for annotation. Experimental results on four widely used benchmarks demonstrate that GCAtt consistently outperforms human-annotated semantics. Code and data are available at https://github.com/RowenaHe/GCAtt.
Qingzhi He, Wentong Li 0001, Shengcai Liao, Rong Quan, Tong Cui, Jie Qin 0004
SIGIR5
2025 E²GO : Free Your Hands for Smartphone Interaction
abstract
Current eye-gaze interaction technologies for smartphones are considered inflexible, inaccurate, and power-hungry. These methods typically rely on hand involvement and accomplish partial interactions. In this paper, we propose a novel eye-gaze smartphone interaction method named Event-driven Eye-Gaze Operation (E2GO), which can realize comprehensive interaction using only eyes and gazes to cover various interaction types. Before the interaction, an anti-jitter gaze estimation method was exploited to stabilize human eye fixation and predict accurate and stable human gaze positions on smartphone screens to further explore refined time-dependent eye-gaze interactions. We also integrated an event-triggering mechanism in E2GO to significantly decrease its power consumption to deploy on smartphones. We have implemented the prototype of E2GO on different brands of smartphones and conducted a comprehensive user study to validate its efficacy, demonstrating E2GO‘s superior smartphone control capabilities across various scenarios. Demo videos
Shaoming Yan, Yuanliang Ju, Rong Quan, Huawei Tu, Dong Liang 0008
Int. J. Hum. Comput. Interact.3
2025 Disaggregation Distillation for Person Search
abstract
Person search is a challenging task in computer vision and multimedia understanding, which aims at localizing and identifying target individuals in realistic scenes. State-of-the-art models achieve remarkable success but suffer from overloaded computation and inefficient inference, making them impractical in most real-world applications. A promising approach to tackle this dilemma is to compress person search models with knowledge distillation (KD). Previous KD-based person search methods typically distill the knowledge from the re-identification (re-id) branch, completely overlooking the useful knowledge from the detection branch. In addition, we elucidate that the imbalance between person and background regions in feature maps has a negative impact on the distillation process. To this end, we propose a novel KD-based approach, namely Disaggregation Distillation for Person Search (DDPS), which disaggregates the distillation process and feature maps, respectively. Firstly, the distillation process is disaggregated into two task-oriented sub-processes,i.e., detection distillation and re-id distillation, to help the student learn both accurate localization capability and discriminative person embeddings. Secondly, we disaggregate each feature map into person and background regions, and distill these two regions independently to alleviate the imbalance problem. More concretely, three types of distillation modules,i.e., logit distillation (LD), correlation distillation (CD), and disaggregation feature distillation (DFD), are particularly designed to transfer comprehensive information from the teacher to the student. Note that such a simple yet effective distillation scheme can be readily applied to both homogeneous and heterogeneous teacher-student combinations. We conduct extensive experiments on two person search benchmarks, where the results demonstrate that, surprisingly, our DDPS enables the student model to surpass the performance of the corresponding teacher model, even achieving comparable results with general person search models.
Rong Quan, Haiyan Chen 0001, Jiamei Liu, Yichao Yan, Song Bai 0001, Jie Qin 0004
IEEE Trans. Multim.2
2024 Pathformer3D: A 3D Scanpath Transformer for $360^{\circ }$ Images
Rong Quan, Yantao Lai, Mengyu Qiu, Dong Liang 0008
ECCV (35)1
2024 Camouflaged Object Detection Via Style Transfer-Based Data Augmentation
abstract
Infrared (IR) images can be seen as complementary to visible light (RGB) images, as they can capture accurate targets in low-visibility conditions. However, camouflaged object detection (COD) based on RGB and IR images is expensive. To this end, we propose to exploit a style transfer-based data augmentation method to generate pseudo-IR images by absorbing the style information of IR images into RGB images, and to perform COD based on RGB and the pseudo-IR images. For RGB and IR-based COD, we propose a novel Edge-guided Uncertainty-aware Fusion Network (EUFNet), to make better use of the complementarity between the two kinds of images. Specifically, an uncertainty-aware fusion module is first proposed to aggregate RGB and IR features by estimating their uncertainties. Then, an edge enhancement module is proposed to extract and enhance the edge information in multiple stages. Lastly, a hierarchical integration module is designed to integrate RGB and IR features with edge cues. Extensive experiments demonstrate the effectiveness of the generated pseudo-IR images as well as the proposed EUFNet. The code is available at https://github.com/csdahunzi/COD.
Dongni Lu, Jiaxuan Chen 0006, Haiyan Chen 0001, Ziyi Peng, Rong Quan, Jie Qin 0004
ICIP5
2024 MTA-PS: Towards Practical Person Search in Videos
abstract
Person search (PS) aims to simultaneously localize and identify a target person from natural, uncropped images. Existing PS datasets and research works are mostly based on individual images, exhibiting limited practicability in real-world surveillance scenarios. We contend that videos, compared to static images, offer additional temporal information, making searching for the trajectory of the target person from videos more realistic and accurate. In this paper, we propose a new practical and realistic task, namely person search in videos, and a new evaluation metric specifically tailored for it. To fulfill this, we introduce a new PS dataset, namely MTAPS, based on an existing large-scale simulated video dataset. MTA-PS is the first cross-camera PS dataset in virtual videos, consisting of 6 cameras, 60 videos, 1.8K identities, 295.2K frames, 7.3M bounding boxes, and more than 20 minutes per camera, which is challenging and comprehensive, and meanwhile avoids privacy issues. To validate the effectiveness of PS in videos and make full use of the temporal information on our dataset, we also propose a novel framework by seamlessly integrating the three sub-tasks of person detection, tracking, and re-identification. Extensive experiments demonstrate that our method performs favorably over existing counterparts on the newly-introduced MTA-PS dataset. Codes and datasets are available at https://github.com/mtmyyy/MTA-PS.
Tiancheng Ying, Rong Quan, Peng Zheng 0004, Yichao Yan, Jie Qin 0004
ICIP2
2024 Visual Feature Disentanglement for Zero-Shot Learning
abstract
Generative model-based zero-shot learning (ZSL) approaches usually transform ZSL to supervised learning by generating full visual features for unseen classes, which may contain redundant or noisy information harmful for ZSL classification. In this work, we propose a novel generative framework that fully exploits the visual features by disentangling them into three distinct components, i.e., semantically correlative (SC), visually discriminative (VD), and residual features, respectively. In particular, we propose an encoder-decoder disentanglement framework to learn SC features that align closely with their corresponding semantic label embeddings, as well as VD features that contribute significantly to visual classification accuracy. Additionally, we introduce a mutual information-based loss to disperse the distributions between the discriminative (SC and VD) features and the residual ones, thereby eliminating the redundant information that may hinder the final ZSL classification. Our proposed method is extensively evaluated on four popular ZSL benchmarks, where the experimental results demonstrate its superiority over existing counterparts. Code is available at https://github.com/RowenaHe/VFD-ZSL.
Qingzhi He, Rong Quan, Weifeng Yang, Jie Qin 0004
ICME2
2024 Cross-Domain Few-Shot Semantic Segmentation via Doubly Matching Transformation
Rong Quan
IJCAI2
2024 BEVTemp: Enhancing Vision-Based Roadside 3D Object Detection with Temporal Information
Gaoyuan Miao, Rong Quan, Cong Pan 0001, Zhiheng Hu, Jie Qin 0004
PRICAI (3)2
2024 MACA: Memory-aided Coarse-to-fine Alignment for Text-based Person Search
abstract
Text-based person search (TBPS) aims to search for the target person in the full image through textual descriptions. The key to addressing this task is to effectively perform cross-modality alignment between text and images. In this paper, we propose a novel TBPS framework, named Memory-Aided Coarse-to-fine Alignment (MACA), to learn an accurate and reliable alignment between the two modalities. Firstly, we introduce a proposal-based alignment module, which performs contrastive learning to accurately align the textual modality with different pedestrian proposals at a coarse-grained level. Secondly, for the fine-grained alignment, we propose an attribute-based alignment module to mitigate unreliable features by aligning text-wise details with image-wise global features. Moreover, we introduce an intuitive memory bank strategy to supplement useful negative samples for more effective contrastive learning, improving the convergence and generalization ability of the model based on the learned discriminative features. Extensive experiments on CUHK-SYSU-TBPS and PRW-TBPS demonstrate the superiority of MACA over state-of-the-art approaches. The code is available at https://github.com/suliangxu/MACA.
Liangxu Su, Rong Quan, Jie Qin 0004
SIGIR2
2023 Movienet-PS: A Large-Scale Person Search Dataset in the Wild
abstract
Person search (PS) aims to jointly localize and identify a query person from natural, uncropped images. Existing works unintentionally adopt pedestrians (with similar poses and unchanging clothing) as the query and restrict the application scenarios in surveillance. This is due to that most PS datasets are collected from surveillance cameras with a limited diversity of views, scenes, appearances, etc. In this paper, we study a more general and realistic task in the wild, where we aim to search target persons with a much higher degree of diversity. To this end, we introduce a new PS dataset, namely MovieNet-PS, based on an existing large-scale movie dataset. MovieNet-PS is currently the largest and most diverse PS dataset, consisting of 160K images (100K for training), 274K bounding boxes, and 3K identities. It stands out from existing counterparts from two levels of diversities, i.e., scene-level and identity-level, with 92,043 scenes and significant variations in poses, clothing, scales, etc. for the same identity. To validate the rich context information on our dataset and make full use of it, we propose a novel global-local context network, which exploits scene and group context to boost the search performance. Extensive experiments demonstrate that MovieNet-PS is more challenging and comprehensive than existing datasets, and our approach further pushes the state of the art by a large margin (relatively 34% in mAP) on this dataset. Codes, models, and the dataset are available at: https://github.com/ZhengPeng7/GLCNet.
Jie Qin 0004, Peng Zheng 0004, Yichao Yan, Rong Quan, Xiaogang Cheng, Bingbing Ni
ICASSP4
2023 RefineTAD: Learning Proposal-free Refinement for Temporal Action Detection
abstract
Temporal action detection (TAD) aims to localize the start and end frames of actions in untrimmed videos, which is a challenging task due to the similarity of adjacent frames and the ambiguity of action boundaries. Previous methods often generate coarse proposals first and then perform proposal-based refinement, which is coupled with prior action detectors and leads to proposal-oriented offsets. However, this paradigm increases the training difficulty of the TAD model and is heavily influenced by the quantity and quality of the proposals. To address the above issues, we decouple the refinement process from conventional TAD methods and propose a learnable, proposal-free refinement method for fine boundary localization, named RefineTAD. We first propose a multi-level refinement module to generate multi-scale boundary offsets, score offsets and boundary-aware probability at each time point based on the feature pyramid. Then, we propose an offset focusing strategy to progressively refine the predicted results of TAD models in a coarse-to-fine manner with our multi-scale offsets. We perform extensive experiments on three challenging datasets and demonstrate that our RefineTAD significantly improves the state-of-the-art TAD methods with minimal computational overhead.
Zhengye Zhang, Rong Quan, Limin Wang 0002, Jie Qin 0004
ACM Multimedia3
2022 More Than Accuracy: An Empirical Study of Consistency Between Performance and Interpretability
Dong Liang 0008, Rong Quan, Songlin Du, Yaping Yan
PRICAI (3)3
2018 Robust Object Co-Segmentation Using Background Prior
abstract
Given a set of images that contain objects from a common category, object co-segmentation aims at automatically discovering and segmenting such common objects from each image. During the past few years, object co-segmentation has received great attention in the computer vision community. However, the existing approaches are usually designed with misleading assumptions, unscalable priors, or subjective computational models, which do not have sufficient robustness for dealing with complex and unconstrained real-world image contents. This paper proposes a novel two-stage co-segmentation framework, mainly for addressing the robustness issue. In the proposed framework, we first introduce the concept of union background and use it to improve the robustness for suppressing the image backgrounds contained by the given image groups. Then, we also weaken the requirement for the strong prior knowledge by using the background prior instead. This can improve the robustness when scaling up for the unconstrained image contents. Based on the weak background prior, we propose a novel MR-SGS model, i.e., manifold ranking with the self-learned graph structure, which can infer suitable graph structures in a data-driven manner rather than building the fixed graph structure relying on the subjective design. Such capacity is critical for further improving the robustness in inferring the foreground/background probability of each image pixel. Comprehensive experiments and comparisons with other state-of-the-art approaches can demonstrate the effectiveness of the proposed work.
Junwei Han 0001, Rong Quan, Dingwen Zhang, Feiping Nie 0001
IEEE Trans. Image Process.2
2018 Unsupervised Salient Object Detection via Inferring From Imperfect Saliency Models
abstract
Visual saliency detection has become an active research direction in recent years. A large number of saliency models, which can automatically locate objects of interest in images, have been developed. As these models take advantage of different kinds of prior assumptions, image features, and computational methodologies, they have their own strengths and weaknesses and may cope with only one or a few types of images well. Inspired by these facts, this paper proposes a novel salient object detection approach with the idea of inferring a superior model from a variety of previous imperfect saliency models via optimally leveraging the complementary information among them. The proposed approach mainly consists of three steps. First, a number of existing unsupervised saliency models are adopted to provide weak/imperfect saliency predictions for each region in the image. Then, a fusion strategy is used to fuse each image region's weak saliency predictions into a strong one by simultaneously considering the performance differences among various weak predictions and various characteristics of different image regions. Finally, a local spatial consistency constraint that ensures high similarity of the saliency labels for neighboring image regions with similar features is proposed to refine the results. Comprehensive experiments on five public benchmark datasets and comparisons with a number of state-of-the-art approaches can demonstrate the effectiveness of the proposed work.
Rong Quan, Junwei Han 0001, Dingwen Zhang, Feiping Nie 0001, Xueming Qian, Xuelong Li 0001
IEEE Trans. Multim.1
2016 Object Co-segmentation via Graph Optimized-Flexible Manifold Ranking
abstract
Aiming at automatically discovering the common objects contained in a set of relevant images and segmenting them as foreground simultaneously, object co-segmentation has become an active research topic in recent years. Although a number of approaches have been proposed to address this problem, many of them are designed with the misleading assumption, unscalable prior, or low flexibility and thus still suffer from certain limitations, which reduces their capability in the real-world scenarios. To alleviate these limitations, we propose a novel two-stage co-segmentation framework, which introduces the weak background prior to establish a globally close-loop graph to represent the common object and union background separately. Then a novel graph optimized-flexible manifold ranking algorithm is proposed to flexibly optimize the graph connection and node labels to co-segment the common objects. Experiments on three image datasets demonstrate that our method outperforms other state-of-the-art methods.
Rong Quan, Junwei Han 0001, Dingwen Zhang, Feiping Nie 0001
CVPR1