VLDB 2026 Research / reviewers in the wild / expert
Rongrong Gao
dblp:168/8117
· DBLP profile ↗
10ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scribble-Based Weakly Supervised Camouflaged Object Detection via SAM-Guided Feature Correlation TransformerabstractWeakly Supervised Camouflaged Object Detection (WS-COD) aims to locate camouflaged objects with only sparse supervision, thereby substantially reducing the reliance on costly pixel-level annotations. This task poses two major challenges: limited supervision arising from sparse annotations (e.g., scribbles), and weak discriminability due to the inherent high visual similarity between camouflaged objects and their surroundings. To tackle these challenges, this paper proposes a novel multi-scale feature correlation transformer guided by the Segment Anything Model (SAM) for scribble-based WSCOD. Specifically, we introduce a cross-scale correlation module built upon Transformers, which exploits enriched cross-attention mechanisms to capture long-range global correlations and multi-scale discriminative cues, enabling accurate segmentation of camouflaged objects. In addition, we develop a SAM-based pseudo-label generation module that leverages sparse annotations as prompts to produce high-quality object masks, thereby enhancing supervision. Extensive experiments on three challenging datasets demonstrate that our proposed method consistently and significantly surpasses existing state-of-the-art approaches for scribble-based WSCOD. The code will be available at: http://github.com/ farewellIamLoser/FCT-SAM-WSCOD. Zi-Jie Wu, Rongrong Gao, Tian-Zhu Xiang |
ECAI | 2 |
| 2025 | Uncertainty-Aware Transformer for Referring Camouflaged Object DetectionabstractReferring camouflaged object detection (Ref-COD) is a recently proposed task, aiming to segment specified camouflaged objects by leveraging visual reference, i.e., a small set of referring images with salient target objects. Ref-COD poses a considerable challenge due to the difficulty of discerning camouflaged objects from their highly similar backgrounds, as well as the significant feature differences between the camouflaged objects and the provided visual reference. To tackle the above dilemma, we propose a novel uncertainty-aware transformer for the Ref-COD task, termed UAT. UAT first utilizes a cross-attention mechanism to align and integrate visual reference to guide camouflaged feature learning, and then models dependencies between patches in a probabilistic manner to learn predictive uncertainty and excavate discriminative camouflaged features. Specifically, we first design a referring feature aggregation (RFA) module to align and incorporate referring features with camouflaged features, guiding targeted specific feature learning within the feature space of camouflaged images. Then, to enhance multi-level feature extraction, we develop a cross-attention encoder (CAE) to integrate global information and multi-scale semantics between adjacent layers to excavate critical camouflage cues. More importantly, we propose a transformer probabilistic decoder (TPD) to model the dependencies between patches as Gaussian random variables to capture uncertainty-aware camouflaged features. Extensive experiments on the golden Ref-COD benchmark demonstrate the superiority of UAT over existing state-of-the-art competitors. The proposed UAT also achieves competitive performance on several conventional COD datasets, further demonstrating its scalability. The source code is available at https://github.com/CVL-hub/UAT. Ranwan Wu, Tian-Zhu Xiang, Guosen Xie, Rongrong Gao, Xiangbo Shu, Fang Zhao 0006, Ling Shao 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | Diffusion Model for Camouflaged Object DetectionabstractCamouflaged object detection is a challenging task that aims to identify objects that are highly similar to their background. Due to the powerful noise-to-image denoising capability of denoising diffusion models, in this paper, we propose a diffusion-based framework for camouflaged object detection, termed diffCOD, a new framework that considers the camouflaged object segmentation task as a denoising diffusion process from noisy masks to object masks. Specifically, the object mask diffuses from the ground-truth masks to a random distribution, and the designed model learns to reverse this noising process. To strengthen the denoising learning, the input image prior is encoded and integrated into the denoising diffusion model to guide the diffusion process. Furthermore, we design an injection attention module (IAM) to interact conditional semantic features extracted from the image with the diffusion noise embedding via the cross-attention mechanism to enhance denoising learning. Extensive experiments on four widely used COD benchmark datasets demonstrate that the proposed method achieves favorable performance compared to the existing 11 state-of-the-art methods, especially in the detailed texture segmentation of camouflaged objects. Our code will be made publicly available at: https://github.com/ZNan-Chen/diffCOD. Zhennan Chen, Rongrong Gao, Tian-Zhu Xiang, Fan Lin |
ECAI | 2 |
| 2023 | Scene-level Point Cloud Colorization with Semantics-and-geometry-aware NetworksabstractIn robotic applications, we often obtain tons of 3D point cloud data without color information, and it is difficult to visualize point clouds in a meaningful and colorful way. Can we colorize 3D point clouds for better visualization? Existing deep learning-based colorization methods usually only take simple 3D objects as input, and their performance for complex scenes with multiple objects is limited. To this end, this paper proposes a novel semantics-and-geometry-aware colorization network, termed SGNet, for vivid scene-level point cloud colorization. Specifically, we propose a novel pipeline that explores geometric and semantic cues from point clouds containing only coordinates for color prediction. We also design two novel losses, including a colorfulness metric loss and a pairwise consistency loss, to constrain model training for genuine colorization. To the best of our knowledge, our work is the first to generate realistic colors for point clouds of large-scale indoor scenes. Extensive experiments on the widely used ScanNet benchmarks demonstrate that the proposed method achieves state-of-the-art performance on point cloud colorization. Rongrong Gao, Tian-Zhu Xiang, Chenyang Lei, Jaesik Park, Qifeng Chen 0001 |
ICRA | 1 |
| 2023 | A Unified Query-based Paradigm for Camouflaged Instance SegmentationabstractDue to the high similarity between camouflaged instances and the background, the recently proposed camouflaged instance segmentation (CIS) faces challenges in accurate localization and instance segmentation. To this end, inspired by query-based transformers, we propose a unified query-based multi-task learning framework for camouflaged instance segmentation, termed UQFormer, which builds a set of mask queries and a set of boundary queries to learn a shared composed query representation and efficiently integrates global camouflaged object region and boundary cues, for simultaneous instance segmentation and instance boundary detection in camouflaged scenarios. Specifically, we design a composed query learning paradigm that learns a shared representation to capture object region and boundary features by the cross-attention interaction of mask queries and boundary queries in the designed multi-scale unified learning transformer decoder. Then, we present a transformer-based multi-task learning framework for simultaneous camouflaged instance segmentation and camouflaged instance boundary detection based on the learned composed query representation, which also forces the model to learn a strong instance-level query representation. Notably, our model views the instance segmentation as a query-based direct set prediction problem, without other post-processing such as non-maximal suppression. Compared with 14 state-of-the-art approaches, our UQFormer significantly improves the performance of camouflaged instance segmentation. Our code will be available at: https://github.com/dongbo811/UQFormer. Jialun Pei, Rongrong Gao, Tian-Zhu Xiang, Shuo Wang 0010, Huan Xiong |
ACM Multimedia | 3 |
| 2021 | Joint Depth and Normal Estimation from Real-world Time-of-flight Raw DataabstractWe present a novel approach to joint depth and normal estimation for time-of-flight (ToF) sensors. Our model learns to predict the high-quality depth and normal maps jointly from ToF raw sensor data. To achieve this, we meticulously constructed the first large-scale dataset (named ToF-100) with paired raw ToF data and ground-truth high-resolution depth maps provided by an industrial depth camera. In addition, we also design a simple but effective framework for joint depth and normal estimation, applying a robust Chamfer loss via jittering to improve the performance of our model. Our experiments demonstrate that our proposed method can efficiently reconstruct high-resolution depth and normal maps and significantly outperforms state-of-the-art approaches. Rongrong Gao, Na Fan 0002, Wentao Liu 0002, Qifeng Chen 0001 |
IROS | 1 |
| 2021 | A Novel Method of Cropped Images Forensics in Social Networks
Rongrong Gao, Xiaolong Li 0001, Yao Zhao 0001 |
PRCV (2) | 1 |
| 2020 | Future Video Synthesis With Object Motion PredictionabstractWe present an approach to predict future video frames given a sequence of continuous video frames in the past. Instead of synthesizing images directly, our approach is designed to understand the complex scene dynamics by decoupling the background scene and moving objects. The appearance of the scene components in the future is predicted by non-rigid deformation of the background and affine transformation of moving objects. The anticipated appearances are combined to create a reasonable video in the future. With this procedure, our method exhibits much less tearing or distortion artifact compared to other approaches. Experimental results on the Cityscapes and KITTI datasets show that our model outperforms the state-of-the-art in terms of visual quality and accuracy. Yue Wu 0012, Rongrong Gao, Jaesik Park, Qifeng Chen 0001 |
CVPR | 2 |
| 2016 | Feature extraction framework in class space for hyperspectral image classificationabstractIn this paper, a novel feature extraction framework is proposed for hyperspectral image classification. Inspired by the role of discriminant function in classifier, which intends to learn a mapping from the input features to label information in class space, we develop a feature extraction framework to learn the new feature representation of original input features in class space, by establishing the relevance between feature extraction and discriminative classifier. The new learning features integrate the input features and the discrimination information of used classifier with available training samples, which reveal the cues of class in class space. Therefore, the new features are called as the features of class-in-class. Several experiments were conducted to illustrate the availability of the proposed features. Ji Zhao 0006, Yanfei Zhong, Rongrong Gao, Liangpei Zhang 0001, Hong Shu |
IGARSS | 3 |
| 2016 | Multiscale and Multifeature Normalized Cut Segmentation for High Spatial Resolution Remote Sensing ImageryabstractIn this paper, a framework for multiscale and multifeature normalized cut (MMNCut) segmentation is proposed for high spatial resolution (HSR) remote sensing images. Normalized cuts (NCuts), as a widely used segmentation method for natural images, can obtain a globally optimized segmentation result corresponding to the optimized partitions of a graph. However, it is difficult to apply the traditional NCuts directly to HSR images because of the huge computational complexity and the diversity of the characteristics of the land covers. In order to solve these problems, the proposed MMNCuts builds a multiscale graph based on superpixels, which can provide powerful grouping cues to guide the segmentation. Generated by different algorithms with varying parameters, superpixels can capture diverse and multiscale visual patterns of HSR images. In addition, the newly constructed graph integrates the multiscale information by considering various connection relationships. Meanwhile, the successful integration of the multifeature cues, including the spectral information, texture information, and structure information, from a large number of superpixels, helps to enhance the expression ability of the graph. Computationally, this leads to a much more efficient algorithm than the traditional NCuts, and in effect, the proposed method achieves a significantly better performance than the traditional approaches. The experimental results with three HSR image data sets demonstrate that the proposed MMNCut algorithm shows a competitive performance in both qualitative and quantitative evaluations when compared with the other state-of-the-art segmentation algorithms for HSR images. Yanfei Zhong, Rongrong Gao, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |