VLDB 2026 Research / reviewers in the wild / expert
Xin Wang 0118
dblp:10/5630-118
· DBLP profile ↗
11ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0002-7977-6586ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NoiseController: Towards Consistent Multi-View Video Generation via Noise Decomposition and Collaboration
Haotian Dong, Xin Wang 0118, Di Lin 0002, Yipeng Wu, Kairui Yang, Ping Li 0016, Qing Guo 0005 |
ICCV | 2 |
| 2025 | Beyond Single Images: Retrieval Self-Augmented Unsupervised Camouflaged Object DetectionabstractAt the core of Camouflaged Object Detection (COD) lies segmenting objects from their highly similar surroundings. Previous efforts navigate this challenge primarily through image-level modeling or annotation-based optimization. Despite advancing considerably, this commonplace practice hardly taps valuable dataset-level contextual information or relies on laborious annotations. In this paper, we propose RISE, a RetrIeval SElf-augmented paradigm that exploits the entire training dataset to generate pseudo-labels for single images, which could be used to train COD models. RISE begins by constructing prototype libraries for environments and camouflaged objects using training images (without ground truth), followed by K-Nearest Neighbor (KNN) retrieval to generate pseudo-masks for each image based on these libraries. It is important to recognize that using only training images without annotations exerts a pronounced challenge in crafting high-quality prototype libraries. In this light, we introduce a Clustering-then-Retrieval (CR) strategy, where coarse masks are first generated through clustering, facilitating subsequent histogram-based image filtering and cross-category retrieval to produce high-confidence prototypes. In the KNN retrieval stage, to alleviate the effect of artifacts in feature maps, we propose Multi-View KNN Retrieval (MVKR), which integrates retrieval results from diverse views to produce more robust and precise pseudo-masks. Extensive experiments demonstrate that RISE outperforms state-of-the-art unsupervised and prompt-based methods. Code is available at https://github.com/xiaohainku/RISE. Ji Du, Xin Wang 0118, Fangwei Hao, Mingyang Yu 0001, Chunyuan Chen, Jiesheng Wu, Jing Xu 0008, Ping Li 0016 |
ICCV | 2 |
| 2025 | Trajectory-LLM: A Language-based Data Generator for Trajectory Prediction in Autonomous DrivingabstractVehicle trajectory prediction is a crucial aspect of autonomous driving, which requires extensive trajectory data to train prediction models to understand the complex, varied, and unpredictable patterns of vehicular interactions. However, acquiring real-world data is expensive, so we advocate using Large Language Models (LLMs) to generate abundant and realistic trajectories of interacting vehicles efficiently. These models rely on textual descriptions of vehicle-to-vehicle interactions on a map to produce the trajectories. We introduce Trajectory-LLM (Traj-LLM), a new approach that takes brief descriptions of vehicular interactions as input and generates corresponding trajectories. Unlike language-based approaches that translate text directly to trajectories, Traj-LLM uses reasonable driving behaviors to align the vehicle trajectories with the text. This results in an "interaction-behavior-trajectory" translation process. We have also created a new dataset, Language-to-Trajectory (L2T), which includes 240K textual descriptions of vehicle interactions and behaviors, each paired with corresponding map topologies and vehicle trajectory segments. By leveraging the L2T dataset, Traj-LLM can adapt interactive trajectories to diverse map topologies. Furthermore, Traj-LLM generates additional data that enhances downstream prediction models, leading to consistent performance improvements across public benchmarks. The source code is released at https://github.com/TJU-IDVLab/Traj-LLM. Kairui Yang, Gengjie Lin, Haotian Dong, Yipeng Wu, Die Zuo, Jibin Peng, Ziyuan Zhong, Xin Wang 0118, Qing Guo 0005, Xiaosong Jia, Junchi Yan, Di Lin 0002 |
ICLR | 10 |
| 2025 | Accurate-PGNet: Learning to Assemble Perceptual Body Parts for Accurate Human Skeleton EstablishmentabstractThe human skeleton establishment aims to provide accurate localization information of the human body from RGB images and establish a complete human skeleton for many applications, such as action recognition, video surveillance, and human-computer interaction. Considering the inherent human body structure, many recent methods group the relevant body parts and utilize the deep convolutional network to learn the visual context from the part groups. However, the grouping approaches used in these methods heavily rely on prior knowledge of the human body shape but lose important relationships between parts. In this paper, we introduce the Accurate Part Grouping Network (Accurate-PGNet), a novel network for hierarchically grouping body parts in a data-driven manner. In contrast to the previous methods, we use neural architecture search (NAS) to optimize the architecture of Accurate-PGNet and properly group the body parts. The part grouping respects the diverse visual patterns of parts, producing groups containing different body parts. From each group, we learn the visual feature map. It helps to capture the correlation between parts and predict their locations. The feature maps of the part groups are merged hierarchically to capture the higher-order context of parts in larger groups. We extensively evaluated our method on the challenging benchmarks, demonstrating that Accurate-PGNet effectively helps to achieve state-of-the-art results. Di Lin 0002, Xin Wang 0118, George Baciu, C. L. Philip Chen, Ping Li 0016 |
IEEE Trans. Multim. | 3 |
| 2025 | Temporal-Interim Pose Synthesis and Distillation for Dynamic Human Pose EstimationabstractIn the task of dynamic human pose estimation (dynamic HPE), the temporal relationships between human body parts should be captured comprehensively to understand the dynamic human motions, where the correlated motion information eventually helps to recognize body parts. The popular methods are successful in terms of utilizing long-term motion information captured by low-speed cameras. Yet they neglect the underlying intermediate motions between captured frames, which comprise the temporal-interim poses lost in the video. In this article, we introduce a novel framework, temporal-interim pose synthesis and distillation, to produce and leverage the intermediate motion information for dynamic motion establishment. The pose synthesis yields the visual feature maps of the intermediate poses, which appear between the existing video frames. It allows the synthesized and current poses to form richer motion patterns. Next, the pose distillation divides the body parts into several groups, where it learns the specific part-wise relationship within each group. It degrades the complexity of learning useful part-wise relationships from rich motion patterns and extracts more detailed motion information for fine-grained part groups. We extensively evaluate our method on challenging datasets for dynamic pose estimation, achieving state-of-the-artresults. Di Lin 0002, Xin Wang 0118, Bin Sheng 0001, George Baciu, C. L. Philip Chen, Ping Li 0016 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | HRC-Net: Learning Visual Hypothesis, Representative, and Collaboration for Multi-Domain Image InpaintingabstractMulti-domain image inpainting utilizes complementary contextual information from auxiliary domain images to restore corrupted regions. While existing methods reconstruct auxiliary images to provide additional guidance, they face fundamental limitations: recovered pixels with complex patterns often lack representative details, while oversimplified patterns offer insufficient contextual information. To address these challenges, we propose HRC-Net, a novel framework incorporating three generative sub-networks for the comprehensive image inpainting task. Our architecture consists of: (1) A Hypothesis Sub-network that enables robust samplings of pixel-wise hypotheses from multi-domain inputs; (2) A Representative Sub-network that learns to score hypothesis quality based on contextual relevance; and (3) a Collaboration Sub-network that optimizes adaptive fusion kernels to integrate the most pertinent details. Together, these components model the joint distribution of representative scores and convolutional kernels, fostering a precise interaction between auxiliary hypotheses and target image corruption to meticulously repair the target image. Extensive evaluations across multiple benchmark datasets demonstrate HRC-Net's superior performance, significantly outperforming state-of-the-art methods in both quantitative metrics and visual quality. Xin Wang 0118, Di Lin 0002, Wanchao Su, Ji Du, Jie Zhang 0090, Haotian Dong, Ke Xu 0010, Qing Guo 0005, Ping Li 0016 |
ACM Trans. Graph. | 1 |
| 2025 | Distilling complementary information from temporal context for enhancing human appearance in human-specific NeRFabstractAbstract Reconstructing and animating digital avatars with free views from monocular videos have been an interesting research task in the computer vision field for a long time. Recently, some methods have introduced a novel category method of leveraging the neural radiance field to represent the human body in a canonical space with the help of the SMPL model. With the deformation of the points from an observation space into a canonical space, the human appearance can be learned in various poses and viewpoints. However, previous methods highly rely on pose-dependent representation learned from frame-independent optimization and ignore the temporal contexts across the continuous motion video, causing a bad influence on the dynamic appearance texture generation. To overcome these problems, we propose a novel free-viewpoint rendering framework, TMIHuman. It aims at introducing temporal information into NeRF-based rendering and distilling task-relevant information from complex pixel-wise representations. To be specific, we build a temporal fusion encoder that imports timestamps into the learning of non-rigid deformation and fuses the visual features of other frames into human representation. Then, we propose to disentangle the fused features and extract useful visual cues via mutual information objectives. We have extensively evaluated our method and achieved state-of-the-art performance on different public datasets. Xin Wang 0118, George Baciu, Ping Li 0016 |
Vis. Comput. | 2 |
| 2024 | Two-Stage Video Shadow Detection via Temporal-Spatial Adaption
Xin Duan, Yu Cao 0019, Lei Zhu 0003, Gang Fu 0003, Xin Wang 0118, Ping Li 0016 |
ECCV (48) | 5 |
| 2024 | Shadow-aware image colorizationabstractAbstract Significant advancements have been made in colorization in recent years, especially with the introduction of deep learning technology. However, challenges remain in accurately colorizing images under certain lighting conditions, such as shadow. Shadows often cause distortions and inaccuracies in object recognition and visual data interpretation, impacting the reliability and effectiveness of colorization techniques. These problems often lead to unsaturated colors in shadowed images and incorrect colorization of shadows as objects. Our research proposes the first shadow-aware image colorization method, addressing two key challenges that previous studies have overlooked: integrating shadow information with general semantic understanding and preserving saturated colors while accurately colorizing shadow areas. To tackle these challenges, we develop a dual-branch shadow-aware colorization network. Additionally, we introduce our shadow-aware block, an innovative mechanism that seamlessly integrates shadow-specific information into the colorization process, distinguishing between shadow and non-shadow areas. This research significantly improves the accuracy and realism of image colorization, particularly in shadow scenarios, thereby enhancing the practical application of colorization in real-world scenarios. Xin Duan, Yu Cao 0019, Xin Wang 0118, Ping Li 0016 |
Vis. Comput. | 4 |
| 2022 | Generative Status Estimation and Information Decoupling for Image Rain RemovalabstractImage rain removal requires the accurate separation between the pixels of the rain streaks and object textures. But the confusing appearances of rains and objects lead to the misunderstanding of pixels, thus remaining the rain streaks or missing the object details in the result. In this paper, we propose SEIDNet equipped with the generative Status Estimation and Information Decoupling for rain removal. In the status estimation, we embed the pixel-wise statuses into the status space, where each status indicates a pixel of the rain or object. The status space allows sampling multiple statuses for a pixel, thus capturing the confusing rain or object. In the information decoupling, we respect the pixel-wise statuses, decoupling the appearance information of rain and object from the pixel. Based on the decoupled information, we construct the kernel space, where multiple kernels are sampled for the pixel to remove the rain and recover the object appearance. We evaluate SEIDNet on the public datasets, achieving state-of-the-art performances of image rain removal. The experimental results also demonstrate the generalization of SEIDNet, which can be easily extended to achieve state-of-the-art performances on other image restoration tasks (e.g., snow, haze, and shadow removal). Di Lin 0002, Xin Wang 0118, Miaohui Wang, Wuyuan Xie, Qing Guo 0005, Ping Li 0016 |
NeurIPS | 2 |
| 2018 | Efficient image super-resolution integration
Ke Xu 0010, Xin Wang 0118, Xin Yang 0011, Shengfeng He, Qiang Zhang 0008, Xiaopeng Wei, Rynson W. H. Lau |
Vis. Comput. | 2 |