Yuxi Wang 0002

dblp:29/8340-2 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0001-9358-5893ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 DiffPC: Diffusion-Based Projector Photometric Compensation
abstract
Projector photometric compensation corrects color distortions introduced by surface texture, reflection, and ambient lighting. Existing deep learning-based methods usually require professional scene-specific data collection and lack consideration for perceptual quality. To address this limitation, we present a diffusion-based photometric compensation method that reconstructs compensation images under photometric and content-aware guidance. Specifically, we fi rst mo del th e ph otometric distortions introduced during projection as environment-dependent additive noise, thereby reformulating the photometric compensation problem as a denoising task with physical constraints. Next, we introduce a diffusion model, which generates compensation images by following an additive trajectory to iteratively remove the noise. Finally, to accurately estimate the noise at each timestep, by analyzing the factors that contribute to distortions in the physical process of projection and capturing, we design a noise estimation network that incorporates features of both photometry-aware and content conditions. Experiments show that our method achieves superior visual performance in unknown scenarios, thereby exhibiting significant practical advantages over prior art. Our source code is available at https://github.com/cyxwang/DiffPC.
Yuxi Wang 0002, Haibin Ling, Bingyao Huang
IEEE Trans. Vis. Comput. Graph.1
2025 Multi-Scale Video Demoiréing Based on Selective Time-Domain Fusion
abstract
Video demoiréing aims to remove Moiré patterns from video footage captured by cameras when recording electronic displays. Considering the limitation that existing methods often fail to effectively exploit temporal information across video frames, thereby hindering their ability to recover details and maintain temporal consistency, this letter proposed a novel Multi-scale Video Demoiréing network (MVDS) based on Selective Temporal Fusion. By employing deformable convolutions for frame alignment and a selective temporal fusion mechanism, MVDS efficiently exploits temporal information to improve demoiréing performance. An embedded multi-scale U-net with attention fusion further refines image details at different scales. Extensive experiments on the VDMoiré dataset demonstrate that MVDS outperforms state-of-the-art methods in terms of both qualitative and quantitative metrics, effectively removing Moiré patterns while preserving image details and temporal consistency.
Yinfeng Fang, Xiaohao Pan, Yuxi Wang 0002, Dalin Zhou, Zhaojie Ju
IEEE Signal Process. Lett.3
2024 ViComp: Video Compensation for Projector-Camera Systems
abstract
Projector video compensation aims to cancel the geometric and photometric distortions caused by non-ideal projection surfaces and environments when projecting videos. Most existing projector compensation methods start by projecting and capturing a set of sampling images, followed by an offline compensation model training step. Thus, abundant user effort is required before the users can watch the video. Moreover, the sampling images have little prior knowledge of the video content and may lead to suboptimal results. To address these issues, this paper builds a video compensation system that can online adapt the compensation parameters. Our approach consists of five threads and can perform compensation, projection, capturing, and short-term and long-term model updates in parallel. Due to the parallel mechanism, rather than projecting and capturing hundreds of sampling images and training the model offline, we can directly use the projected and captured video frames for model updates on the fly. To quickly apply to the new environment, we introduce a deep learning-based compensation model that integrates a fixed transformer-based method and a novel CNN-based network. Moreover, for fast convergence and to reduce error accumulation during fine-tuning, we present a strategy that cooperates with short-term and long-term memory model updates. Experiments show that it significantly outperforms state-of-the-art baselines.
Yuxi Wang 0002, Haibin Ling, Bingyao Huang
IEEE Trans. Vis. Comput. Graph.1
2023 CompenHR: Efficient Full Compensation for High-resolution Projector
abstract
Full projector compensation is a practical task of projector-camera systems. It aims to find a projector input image, named compensation image, such that when projected it cancels the geometric and photometric distortions due to the physical environment and hardware. State-of-the-art methods use deep learning to address this problem and show promising performance for low-resolution setups. However, directly applying deep learning to high-resolution setups is impractical due to the long training time and high memory cost. To address this issue, this paper proposes a practical full compensation solution. Firstly, we design an attention-based grid refinement network to improve geometric correction quality. Secondly, we integrate a novel sampling scheme into an end-to-end compensation network to alleviate computation and introduce attention blocks to preserve key features. Finally, we construct a benchmark dataset for high-resolution projector full compensation. In experiments, our method demonstrates clear advantages in both efficiency and quality.
Yuxi Wang 0002, Haibin Ling, Bingyao Huang
VR1
2023 Enhanced YOLOv5 algorithm for helmet wearing detection via combining bi-directional feature pyramid, attention mechanism and transfer learning
Yinfeng Fang, Yuxi Wang 0002
Multim. Tools Appl.4
2023 Lightweight Asymmetric Convolutional Distillation Network for Single Image Super-Resolution
abstract
Recent single image super-resolution methods based on various complex deep neural networks have achieved remarkable success. However, these methods require a large amount of computational overhead while improving performance, and thus are difficult to apply to mobile devices in real-world scenarios. In this letter, we design an efficient asymmetric convolutional distillation block (ACDB). Especially in this block, introducing an asymmetric convolution block (ACB) and reusing shallow distillation features can effectively improve the performance of the model and reduce the model complexity. Our model achieves efficient performance while maintaining low complexity.>
Yuxi Wang 0002
IEEE Signal Process. Lett.2
2018 Simultaneous Trajectory Association and Clustering for Motion Segmentation
abstract
Trajectory association and clustering are two key problems in motion analysis. While association links the points of interest to form trajectories, clustering discovers motion patterns of these trajectories and group them into clusters. Despite mutually related, the two problems have been typically studied separately in the literature. In this letter, we formulate them as a unified optimization problem and take the advantage of high-order information to capture the interrelations for Motion Segmentation by Trajectory Association and Clustering (MSTAC). To solve this unified problem, we propose an alternating optimization strategy to improve the association and clustering in each iteration. Specifically, a tensor-based multidimensional assignment method with high-order motion context information is proposed for trajectory association; and a minimum cost multicut-based trajectory clustering method is introduced for trajectory clustering. While the association process provides incomplete trajectories to clustering, the clustering method presents high-order context information to improve the performance of association; and thus they benefit from each other. Experiments on the Hopkins 155 dataset and a realistic airport sequence demonstrate that the proposed MSTAC framework obtains high accuracy on both trajectory association and clustering.
Yuxi Wang 0002, Yue Liu 0005, Erik Blasch, Haibin Ling
IEEE Signal Process. Lett.1
2016 Visual tracking via sparsity pattern learning
abstract
Recently sparse representation has been applied to visual tracking by modeling the target appearance using a sparse approximation over the template set. However, this approach is limited by the high computational cost of the ℓ1-norm minimization involved, which also impacts on the amount of particle samples that we can have. This paper introduces a basic constraint on the self-representation of the target set. The sparsity pattern in the self-representation allows us to recover the “sparse coefficients” of the candidate samples by some small-scale ℓ2-norm minimization; this results in a fast tracking algorithm. It also leads to a principled dictionary update mechanism which is crucial for good performance. Experiments on a recently released benchmark with 50 challenging video sequences show significant runtime efficiency and tracking accuracy achieved by the proposed algorithm.
Yuxi Wang 0002, Yue Liu 0005, Zhuwen Li, Loong Fah Cheong, Haibin Ling
ICPR1
2013 Diminished reality using appearance and 3D geometry of internet photo collections
abstract
This paper presents a new system level framework for Diminished Reality, leveraging for the first time both the appearance and 3D information provided by large photo collections on the Internet. Recent computer vision techniques have made it possible to automatically reconstruct 3-D structure-from-motion points from large and unordered photo collections. Using these point clouds and a prior provided by GPS, reasonably accurate 6 degree of freedom camera poses can be obtained, thus allowing localization. Once the camera (and hence the user) is correctly localized, photos depicting scenes visible from the user's viewpoint can be used to remove unwanted objects indicated by the user in the video sequences. Existing methods based on texture synthesis bring undesirable artifacts and video inconsistency when the background is heterogeneous; the task is rendered even harder for these methods when the background contains complex structures. On the other hand, methods based on plane warping fail when the background has arbitrary shape. Unlike these methods, our algorithm copes with these problems by making use of internet photos, registering them in 3D space and obtaining the 3D scene structure in an offline process. We carefully design the various components during the online phase so as to meet both speed and quality requirements of the task. Experiments on real data collected demonstrate the superiority of our system.
Zhuwen Li, Yuxi Wang 0002, Jiaming Guo, Loong Fah Cheong, Steven Zhiying Zhou
ISMAR2
2013 Robust structure from motion with affine camera via low-rank matrix recovery
abstract
Abstract We present a novel approach to structure from motion that can deal with missing data and outliers with an affine camera. We model the corruptions as sparse error. Therefore the structure from motion problem is reduced to the problem of recovering a low-rank matrix from corrupted observations. We first decompose the matrix of trajectories of features into low-rank and sparse components by nuclear-norm and ℓ 1-norm minimization, and then obtain the motion and structure from the low-rank components by the classical factorization method. Unlike pervious methods, which have some drawbacks such as depending on the initial value selection and being sensitive to the large magnitude errors, our method uses a convex optimization technique that is guaranteed to recover the low-rank matrix from highly corrupted and incomplete observations. Experimental results demonstrate that the proposed approach is more efficient and robust to large-scale outliers.
Lun Wu, Yongtian Wang, Yue Liu 0005, Yuxi Wang 0002
Sci. China Inf. Sci.4