VLDB 2026 Research / reviewers in the wild / expert
Wenxuan Cheng
dblp:370/2615
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Vision and language · 44% Segmentation and scene understanding · 38% Image recognition and object detection · 19% |
Topics — the 4 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
referring image segmentation |
2.1 | 3 | 2026 | Improving Generalized Visual Grounding With Instance-Aware Joint Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2026 DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation Through Loopback Synergy · ICCV 2025 PropVG: End-To-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination · ICCV 2025 |
Computer vision › Vision and language
visual grounding |
1.9 | 2 | 2026 | Improving Generalized Visual Grounding With Instance-Aware Joint Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2026 PropVG: End-To-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination · ICCV 2025 |
Computer vision › Vision and language
multimodal understanding |
0.9 | 1 | 2025 | DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation Through Loopback Synergy · ICCV 2025 |
Computer vision › Image recognition and object detection › object detection
object proposal generation |
0.9 | 1 | 2025 | PropVG: End-To-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination · ICCV 2025 |
Methods — techniques the papers use, named apart from their topics
multi-task learning · 1.0joint learning · 1.0instance query · 1.0proposal-based framework · 0.9non-referent sample conversion · 0.9multi-granularity discrimination · 0.9loopback synergy · 0.9data augmentation · 0.9contrastive learning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Generalized Visual Grounding With Instance-Aware Joint LearningabstractGeneralized visual grounding tasks, including Generalized Referring Expression Comprehension (GREC) and Segmentation (GRES), extend the classical visual grounding paradigm by accommodating multi-target and non-target scenarios. Specifically, GREC focuses on accurately identifying all referential objects at the coarse bounding box level, while GRES aims for achieve fine-grained pixel-level perception. However, existing approaches typically treat these tasks independently, overlooking the benefits of jointly training GREC and GRES to ensure consistent multi-granularity predictions and streamline the overall process. Moreover, current methods often treat GRES as a semantic segmentation task, neglecting the crucial role of instance-aware capabilities and the necessity of ensuring consistent predictions between instance-level boxes and masks. To address these limitations, we propose InstanceVG, a multi-task generalized visual grounding framework equipped with instance-aware capabilities, which leverages instance queries to unify the joint and consistency predictions of instance-level boxes and masks. To the best of our knowledge, InstanceVG is the first framework to simultaneously tackle both GREC and GRES while incorporating instance-aware capabilities into generalized visual grounding. To instantiate the framework, we assign each instance query a prior reference point, which also serves as an additional basis for target matching. This design facilitates consistent predictions of points, boxes, and masks for the same instance. Extensive experiments obtained on ten datasets across four tasks demonstrate that InstanceVG achieves state-of-the-art performance, significantly surpassing the existing methods in various evaluation metrics. The code and model will be publicly available at https://github.com/Dmmm1997/InstanceVG. Wenxuan Cheng, Jiang-Jiang Liu 0001, Lingfeng Yang, Zhenhua Feng 0001, Wankou Yang, Jingdong Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | PLRVG: Progressive layer-wise refinement for visual grounding via deep-to-shallow decoding
Wenxuan Cheng, Wankou Yang |
Pattern Recognit. | 1 |
| 2026 | DRL: An efficient heterogeneous spatial feature interaction framework for UAV self-localization
Enhui Zheng, Wenxuan Cheng, Zhenhua Feng 0001, Wankou Yang |
Pattern Recognit. | 3 |
| 2026 | GC3VG: Generalized Multi-Task Visual Grounding With Coarse-to-Fine Consistency ConstraintsabstractIn this work, we propose an efficient and streamlined paradigm to address the challenge of consistency prediction in generalized multi-task visual grounding. While most existing approaches primarily focus on integrating multi-modal information and employing multi-task learning to enhance both visual and linguistic understanding, they often rely on joint supervision at the region and pixel levels to exploit task complementarities. In contrast, C3VG explores the relatively under-addressed problem ofconsistency across multi-task predictions. To this end, a multi-task visual grounding framework based on a coarse-to-fine architecture is introduced. Empirical studies demonstrate that the incorporation of both implicit and explicit consistency constraints substantially enhances the coherence between detection and segmentation outputs. However, C3VG is restricted to single-referent visual grounding scenarios and exhibits limited generalizability to real-world applications, which often involve multi-referents or even absent referent. To overcome these limitations, we proposeGC3VG, which incorporates three key advancements: (1) extension to generalized scenarios, including both multi-referent and non-referent cases; (2) aUnified Coherent Refinement Modulethat implicitly encodes region- and instance-level features while explicitly modeling their relational alignment through an IoUbased constraint; and (3) aGranularity-aware Hard-mining Alignmentstrategy that enforces prediction consistency in the feature space and simultaneously enhances the discriminative power of visual and linguistic representations. Extensive experiments on RefCOCO/+/g and gRefCOCO demonstrate the effectiveness and generalizability of the proposed framework. Kai Chen 0037, Wenxuan Cheng, Jiedong Zhuang, Zhenhua Feng 0001, Pengfei Zhu 0001, Wankou Yang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation Through Loopback SynergyabstractReferring Image Segmentation (RIS) is a challenging task that aims to segment objects in an image based on natural language expressions. While prior studies have predominantly concentrated on improving vision-language interactions and achieving fine-grained localization, a systematic analysis of the fundamental bottlenecks in existing RIS frameworks remains underexplored. To bridge this gap, we propose DeRIS, a novel framework that decomposes RIS into two key components: perception and cognition. This modular decomposition facilitates a systematic analysis of the primary bottlenecks impeding RIS performance. Our findings reveal that the predominant limitation lies not in perceptual deficiencies, but in the insufficient multi-modal cognitive capacity of current models. To mitigate this, we propose a Loopback Synergy mechanism, which enhances the synergy between the perception and cognition modules, thereby enabling precise segmentation while simultaneously improving robust image-text comprehension. Additionally, we analyze and introduce a simple non-referent sample conversion data augmentation to address the long-tail distribution issue related to target existence judgement in general scenarios. Notably, DeRIS demonstrates inherent adaptability to both non- and multi-referents scenarios without requiring specialized architectural modifications, enhancing its general applicability. The codes and models are available at https://github.com/Dmmm1997/DeRIS. Wenxuan Cheng, Jiang-Jiang Liu 0001, Wenxiao Cai, Yanpeng Sun, Wankou Yang |
ICCV | 2 |
| 2025 | PropVG: End-To-End Proposal-Driven Visual Grounding with Multi-Granularity DiscriminationabstractRecent advances in visual grounding have largely shifted away from traditional proposal-based two-stage frameworks due to their inefficiency and high computational complexity, favoring end-to-end direct reference paradigms. However, these methods rely exclusively on the referred target for supervision, overlooking the potential benefits of prominent prospective targets. Moreover, existing approaches often fail to incorporate multi-granularity discrimination, which is crucial for robust object identification in complex scenarios. To address these limitations, we propose PropVG, an end-to-end proposal-based framework that, to the best of our knowledge, is the first to seamlessly integrate foreground object proposal generation with referential object comprehension without requiring additional detectors. Furthermore, we introduce a Contrastive-based Refer Scoring (CRS) module, which employs contrastive learning at both sentence and word levels to enhance the capability in understanding and distinguishing referred objects. Additionally, we design a Multi-granularity Target Discrimination (MTD) module that fuses object- and semantic-level information to improve the recognition of absent targets. Extensive experiments on gRefCOCO (GREC/GRES), Ref-ZOM, R-RefCOCO, and RefCOCO (REC/RES) benchmarks demonstrate the effectiveness of PropVG. The codes and models are available at https://github.com/Dmmm1997/PropVG. Wenxuan Cheng, Jiedong Zhuang, Jiang-jiang Liu, Hongshen Zhao, Zhenhua Feng 0001, Wankou Yang |
ICCV | 2 |
| 2023 | Robust Medical Data Sharing System Based on Blockchain and Threshold Rroxy Re-encryption
Wenxuan Cheng, Bo Zhang 0020, Zhongtao Li |
ICA3PP (5) | 1 |