Wenxiao Cai

dblp:348/4892 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
9since 2021 · last 2026
0009-0000-0522-9654ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 HippNet: A Hippocampus-Inspired Architecture for Spatial Memory and Inference Using CMOS Oscillator Networks
Zongru Li, Wenxiao Cai, Thomas H. Lee
ISCAS2
2026 Probabilistic modeling of disparity uncertainty for robust and efficient stereo matching
Wenxiao Cai, Dongting Hu, Ruoyan Yin, Jiankang Deng, Huan Fu, Wankou Yang, Mingming Gong
Pattern Recognit.1
2025 Object-level Geometric Structure Preserving for Natural Image Stitching
abstract
The topic of stitching images with globally natural structures holds paramount significance, with two main goals: pixel-level alignment and distortion prevention. The existing approaches exhibit the ability to align well, yet fall short in maintaining object structures. In this paper, we endeavour to safeguard the overall OBJect-level structures within images based on Global Similarity Prior (OBJ-GSP), on the basis of good alignment performance. Our approach leverages semantic segmentation models like the family of Segment Anything Model to extract the contours of any objects in a scene. Triangular meshes are employed in image transformation to protect the overall shapes of objects within images. The balance between alignment and distortion prevention is achieved by allowing the object meshes to strike a balance between similarity and projective transformation. We also demonstrate that object-level semantic information is necessary in low-altitude aerial image stitching. Additionally, we propose StitchBench, the largest image stitching benchmark with most diverse scenarios. Extensive experimental results demonstrate that OBJ-GSP outperforms existing methods in both pixel alignment and shape preservation.
Wenxiao Cai, Wankou Yang
AAAI1
2025 DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation Through Loopback Synergy
abstract
Referring Image Segmentation (RIS) is a challenging task that aims to segment objects in an image based on natural language expressions. While prior studies have predominantly concentrated on improving vision-language interactions and achieving fine-grained localization, a systematic analysis of the fundamental bottlenecks in existing RIS frameworks remains underexplored. To bridge this gap, we propose DeRIS, a novel framework that decomposes RIS into two key components: perception and cognition. This modular decomposition facilitates a systematic analysis of the primary bottlenecks impeding RIS performance. Our findings reveal that the predominant limitation lies not in perceptual deficiencies, but in the insufficient multi-modal cognitive capacity of current models. To mitigate this, we propose a Loopback Synergy mechanism, which enhances the synergy between the perception and cognition modules, thereby enabling precise segmentation while simultaneously improving robust image-text comprehension. Additionally, we analyze and introduce a simple non-referent sample conversion data augmentation to address the long-tail distribution issue related to target existence judgement in general scenarios. Notably, DeRIS demonstrates inherent adaptability to both non- and multi-referents scenarios without requiring specialized architectural modifications, enhancing its general applicability. The codes and models are available at https://github.com/Dmmm1997/DeRIS.
Wenxuan Cheng, Jiang-Jiang Liu 0001, Wenxiao Cai, Yanpeng Sun, Wankou Yang
ICCV5
2025 STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
abstract
The use of Multimodal Large Language Models (MLLMs) as an end-to-end solution for Embodied AI and Autonomous Driving has become a prevailing trend. While MLLMs have been extensively studied for visual semantic understanding tasks, their ability to perform precise and quantitative spatial-temporal understanding in real-world applications remains largely unexamined, leading to uncertain prospects. To evaluate models' Spatial-Temporal Intelligence, we introduce STI-Bench, a benchmark designed to evaluate MLLMs' spatial-temporal understanding through challenging tasks such as estimating and predicting the appearance, pose, displacement, and motion of objects. Our benchmark encompasses a wide range of robot and vehicle operations across desktop, indoor, and outdoor scenarios. The extensive experiments reveals that the state-of-the-art MLLMs still struggle in real-world spatial-temporal understanding, especially in tasks requiring precise distance estimation and motion analysis.
XiangRui Liu, Wenxiao Cai
ICCV5
2025 SpatialBot: Precise Spatial Understanding with Vision Language Models
abstract
Vision Language Models (VLMs) have achieved impressive performance in 2D image understanding; however, they still struggle with spatial understanding, which is fundamental to embodied AI. In this paper, we propose SpatialBot, a model designed to enhance spatial understanding by utilizing both RGB and depth images. To train VLMs for depth perception, we introduce the SpatialQA and SpatialQA$\boldsymbol{E}$datasets, which include multi-level depth-related questions spanning various scenarios and embodiment tasks. SpatialBench is also developed to comprehensively evaluate VLMs' spatial understanding capabilities across different levels. Extensive experiments on our spatial-understanding benchmark, general VLM benchmarks, and embodied AI tasks demonstrate the remarkable improvements offered by SpatialBot. The model, code, and datasets are available at https://github.com/BAAI-DCAI/SpatialBot.
Wenxiao Cai, Iaroslav Ponomarenko, Jianhao Yuan, Xiaoqi Li 0020, Wankou Yang, Hao Dong 0003, Bo Zhao 0037
ICRA1
2025 VDD: Varied Drone Dataset for semantic segmentation
Wenxiao Cai, Jinyan Hou, Letian Wu, Wankou Yang
J. Vis. Commun. Image Represent.1
2025 Knowledge NeRF: Few-shot novel view synthesis for dynamic articulated objects
Wenxiao Cai, Xinyue Lei, Junming Leo Chen, Yuzhi Hao
J. Vis. Commun. Image Represent.1
2023 UAV image stitching by estimating orthograph with RGB cameras
abstract
In the field of image stitching, cases with large camera optical center movement and large parallax have been the virgin territory of research. The goal of image stitching is to overcome the parallax and stitch a natural image. We look into this problem in the context of ultra-low altitude flight of a UAV . We model the 3D world in this scenario and quickly estimate orthographic projection by pairs of homography matrices. Our stitching method can achieve precise alignment since it takes parallax well into consideration. The stitching results are natural and the extra time consumed is short.
Wenxiao Cai, Songlin Du, Wankou Yang
J. Vis. Commun. Image Represent.1