Xingrui Wang

dblp:280/8952 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2025 TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video Generation
abstract
Text-driven Image to Video Generation (TI2V) aims to generate controllable video given the first frame and corresponding textual description. The primary challenges of this task lie in two parts: (i) how to identify the target objects and ensure the consistency between the movement trajectory and the textual description. (ii) how to improve the subjective quality of generated videos. To tackle the above challenges, we propose a new diffusion-based TI2V framework, termed TIV-Diffusion, via object-centric textual-visual alignment, intending to achieve precise control and high-quality video generation based on textual-described motion for different objects. Concretely, we enable our TIV-Diffuion model to perceive the textual-described objects and their motion trajectory by incorporating the fused textual and visual knowledge through scale-offset modulation. Moreover, to mitigate the problems of object disappearance and misaligned objects and motion, we introduce an object-centric textual-visual alignment module, which reduces the risk of misaligned objects/motion by decoupling the objects in the reference image and aligning textual features with each object individually. Based on the above innovations, our TIV-Diffusion achieves state-of-the-art high-quality video generation compared with existing TI2V methods.
Xingrui Wang, Xin Li 0082, Yaosi Hu, Hanxin Zhu, Chen Hou, Cuiling Lan, Zhibo Chen 0001
AAAI1
2025 Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Mutimodal Models
abstract
Although large multimodal models (LMMs) have demonstrated remarkable capabilities in visual scene interpretation and reasoning, their capacity for complex and precise 3-dimensional spatial reasoning remains uncertain. Existing benchmarks focus predominantly on 2D spatial understanding and lack a framework to comprehensively evaluate 6D spatial reasoning across varying complexities. To address this limitation, we present Spatial457, a scalable and unbiased synthetic dataset designed with 4 key capability for spatial reasoning: multi-object recognition, 2D location, 3D location, and 3D orientation. We develop a cascading evaluation structure, constructing 7 question types across 5 difficulty levels that range from basic single object recognition to our new proposed complex 6D spatial reasoning tasks. We evaluated various large multimodal models (LMMs) on Spatial457, observing a general decline in performance as task complexity increases, particularly in 3D reasoning and 6D spatial tasks. To quantify these challenges, we introduce the Relative Performance Dropping Rate (RPDR), highlighting key weaknesses in 3D reasoning capabilities. Leveraging the unbiased attribute design of our dataset, we also uncover prediction biases across different attributes, with similar patterns observed in real-world image settings.1The code is released in https://github.com/XingruiWang/Spatial457.
Xingrui Wang, Wufei Ma, Tiezheng Zhang, Celso de Melo, Jieneng Chen, Alan L. Yuille
CVPR1
2025 Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
abstract
For vision-language models (VLMs), understanding the dynamic properties of objects and their interactions in 3D scenes from videos is crucial for effective reasoning about high-level temporal and action semantics. Although humans are adept at understanding these properties by constructing 3D and temporal (4D) representations of the world, current video understanding models struggle to extract these dynamic semantics, arguably because these models use cross-frame reasoning without underlying knowledge of the 3D/4D scenes. In this work, we introduce **DynSuperCLEVR**, the first video question answering dataset that focuses on language understanding of the dynamic properties of 3D objects. We concentrate on three physical concepts—*velocity*, *acceleration*, and *collisions*—within 4D scenes. We further generate three types of questions, including factual queries, future predictions, and counterfactual reasoning that involve different aspects of reasoning on these 4D dynamic properties. To further demonstrate the importance of explicit scene representations in answering these 4D dynamics questions, we propose **NS-4DPhysics**, a **N**eural-**S**ymbolic VideoQA model integrating **Physics** prior for **4D** dynamic properties with explicit scene representation of videos. Instead of answering the questions directly from the video text input, our method first estimates the 4D world states with a 3D generative model powered by a physical prior, and then uses neural symbolic reasoning to answer the questions based on the 4D world states. Our evaluation on all three types of questions in DynSuperCLEVR shows that previous video question answering models and large multimodal models struggle with questions about 4D dynamics, while our NS-4DPhysics significantly outperforms previous state-of-the-art models.
Xingrui Wang, Wufei Ma, Angtian Wang, Adam Kortylewski, Alan L. Yuille
ICLR1
2025 SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
abstract
Despite recent advances on multi-modal models, 3D spatial reasoning remains a challenging task for state-of-the-art open-source and proprietary models. Recent studies explore data-driven approaches and achieve enhanced spatial reasoning performance by fine-tuning models on 3D-related visual question-answering data. However, these methods typically perform spatial reasoning in an implicit manner and often fail on questions that are trivial to humans, even with long chain-of-thought reasoning. In this work, we introduce SpatialReasoner, a novel large vision-language model (LVLM) that addresses 3D spatial reasoning with explicit 3D representations shared between multiple stages--3D perception, computation, and reasoning. Explicit 3D representations provide a coherent interface that supports advanced 3D spatial reasoning and improves the generalization ability to novel question types. Furthermore, by analyzing the explicit 3D representations in multi-step reasoning traces of SpatialReasoner, we study the factual errors and identify key shortcomings of current LVLMs. Results show that our SpatialReasoner achieves improved performance on a variety of spatial reasoning benchmarks, outperforming Gemini 2.0 by 9.2% on 3DSRBench, and generalizes better when evaluating on novel 3D spatial reasoning questions. Our study bridges the 3D parsing capabilities of prior visual foundation models with the powerful reasoning abilities of large language models, opening new directions for 3D spatial reasoning.
Wufei Ma, Yu-Cheng Chou, Qihao Liu, Xingrui Wang, Celso de Melo, Jianwen Xie, Alan L. Yuille
NeurIPS4
2025 Diffusion Models for Image Restoration and Enhancement: A Comprehensive Survey
Xin Li 0082, Yulin Ren, Xin Jin 0014, Cuiling Lan, Xingrui Wang, Wenjun Zeng 0001, Xinchao Wang, Zhibo Chen 0001
Int. J. Comput. Vis.5
2025 Deep Learning Research on Quantitative Evaluation Model of Tea Taste Based on NAR Neural Network
abstract
The demand for tea from consumers exhibits a diversified and personalized trend. Initiating from the scientific analysis of tea taste, this paper addresses issues such as the incomplete cognition standard of tea market taste, varying evaluation criteria for tea, and the development of a young tea consumption group, utilizing the Nonlinear Auto-Regressive (NAR) model, a quantitative evaluation of tea taste is conducted. In this study, raw and ripe tea samples from Pu-erh tea were used as training data, and the NAR was employed for deep learning, error analysis, and comparison. The aim was to predict the taste resulting from different content ratios of chemical components in tea, further verifying the feasibility and scientific accuracy of the quantitative evaluation of tea taste based on the NAR. This study not only broadens the research field of NAR neural network, but also further enables tea enterprises to better provide consumers with diversified tea taste, and provides an important reference for tea taste evaluation.
Mingxin Ji, Xingrui Wang
Int. J. Pattern Recognit. Artif. Intell.2
2024 MoE-DiffIR: Task-Customized Diffusion Priors for Universal Compressed Image Restoration
Yulin Ren, Xin Li 0082, Bingchen Li 0001, Xingrui Wang, Mengxi Guo, Shijie Zhao 0001, Li Zhang 0006, Zhibo Chen 0001
ECCV (9)4
2023 Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning
abstract
Visual Question Answering (VQA) models often perform poorly on out-of-distribution data and struggle on domain generalization. Due to the multi-modal nature of this task, multiple factors of variation are intertwined, making generalization difficult to analyze. This motivates us to introduce a virtual benchmark, Super-CLEVR, where different factors in VQA domain shifts can be isolated in order that their effects can be studied independently. Four factors are considered: visual complexity, question redundancy, concept distribution and concept compositionality. With controllably generated data, Super-CLEVR enables us to test VQA methods in situations where the test data differs from the training data along each of these axes. We study four existing methods, including two neural symbolic methods NSCL [45] and NSVQA [59], and two non-symbolic methods FiLM [50] and mDETR [29]; and our proposed method, probabilistic NSVQA (P-NSVQA), which extends NSVQA with uncertainty reasoning. P-NSVQA outperforms other methods on three of the four domain shift factors. Our results suggest that disentangling reasoning and perception, combined with probabilistic uncertainty, form a strong VQA model that is more robust to domain shifts. The dataset and code are released at https://github.com/Lizw14/Super-CLEVR.
Zhuowan Li, Xingrui Wang, Elias Stengel-Eskin, Adam Kortylewski, Wufei Ma, Benjamin Van Durme, Alan L. Yuille
CVPR2
2023 3D-Aware Visual Question Answering about Parts, Poses and Occlusions
abstract
Despite rapid progress in Visual question answering (\textit{VQA}), existing datasets and models mainly focus on testing reasoning in 2D. However, it is important that VQA models also understand the 3D structure of visual scenes, for example to support tasks like navigation or manipulation. This includes an understanding of the 3D object pose, their parts and occlusions. In this work, we introduce the task of 3D-aware VQA, which focuses on challenging questions that require a compositional reasoning over the 3D structure of visual scenes. We address 3D-aware VQA from both the dataset and the model perspective. First, we introduce Super-CLEVR-3D, a compositional reasoning dataset that contains questions about object parts, their 3D poses, and occlusions. Second, we propose PO3D-VQA, a 3D-aware VQA model that marries two powerful ideas: probabilistic neural symbolic program execution for reasoning and deep neural networks with 3D generative representations of objects for robust visual recognition. Our experimental results show our model PO3D-VQA outperforms existing methods significantly, but we still observe a significant performance gap compared to 2D VQA benchmarks, indicating that 3D-aware VQA remains an important open research area.
Xingrui Wang, Wufei Ma, Zhuowan Li, Adam Kortylewski, Alan L. Yuille
NeurIPS1
2022 Contributions of Shape, Texture, and Color in Visual Recognition
Yunhao Ge, Zhi Xu 0013, Xingrui Wang, Laurent Itti
ECCV (12)4
2022 Exploring Coarse-grained Pre-guided Attention to Assist Fine-grained Attention Reinforcement Learning Agents
abstract
Recently, people have applied the attention mechanism to deep reinforcement learning (DRL), which commits to helping agents focus on crucial factors to learn the task more effectively. However, there is still some margin between the current attention methods and natural human attention since evidence suggests that human attention can be pre-guided before they perform a task, allowing humans to quickly catch areas of important factors at the beginning of the task and then gradually refine fine-grained attention to learn the details during training. This allows humans to use their attention more efficiently. In this paper, we propose an attention method that mimics human attention for DRL in the Atari Games. The proposed method contains a fusion attention module, for which we build a simulated human coarse-grained pre-guided (SHCP) attention module to assist the original fine-grained attention of RL agents. The proposed SHCP attention module contains information about key objects for game tasks and is implemented as a coarse-grained attention region. The experimental results demonstrate that our method can quickly boost performance in the early stages and then outperform the current state-of-the-art fine-grained attention methods significantly in sample efficiency, just like human attention. Further analysis shows that, with fusion attention, agents can not only capture rich features of pre-guided attention but also extend to more improved features after training, which suggests the pre-guided attention signal acts as a good initializer. Therefore, we consider our work reveals a potential and promising direction that combines human attention signals to affect agents' behavior via attention mechanisms.
Yang Liu 0482, Xingrui Wang, Hanfang Yang
IJCNN3
2022 A dark image enhancement method based on multiscale features and dilated residual networks
Xingrui Wang, Yan Piao, Yumo Wang
Neural Process. Lett.1