EDBT 2026 Demo / reviewers in the wild / expert
Yuesong Wang 0001
dblp:67/6192-1
· DBLP profile ↗
17ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-8832-4530ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PATexGS: Perceptual-Adaptive Texture Scheduling for Visual Coherence in Textured Gaussian Splattingabstract3D Gaussian Splatting (3DGS) has emerged as a mainstream solution for real-time rendering and high-fidelity novel view synthesis. Building on this foundation, methods based on Textured Gaussians further improve the expression ability by incorporating explicit texture mapping into Gaussians. However, their reliance on fixed texture resolution often results in noticeable visual incoherence, triggering artifacts such as aliasing or inconsistent sharpness under different viewpoints. To address these issues, we propose PATexGS, a perceptual-adaptive texture scheduling framework designed to improve visual coherence for Textured Gaussians. Specifically, we introduce an entropy-guided texture allocation strategy that dynamically adjusts texture resolution based on each Gaussian’s spatial gradient and rendering contribution, constantly preserving details while being memory efficiency. Furthermore, we incorporate a mipmap-inspired hierarchical scheduling mechanism that adaptively schedule texture levels according to view-dependent projection scale, effectively suppressing aliasing and further enhancing perceptual consistency. Extensive experiments on diverse real-world scenes demonstrate that PATexGS significantly improves visual coherence while maintaining high rendering quality, outperforming existing TexturedGS variants in both perceptual fidelity and storage efficiency. Yuesong Wang 0001, Dounian Ma |
AAAI | 1 |
| 2026 | Matching ambiguity-resilient multi-view stereo via adaptive patch deformation
Zhaojie Zeng, Yuesong Wang 0001 |
Pattern Recognit. | 2 |
| 2025 | Frequency-Aware Density Control via Reparameterization for High-Quality Rendering of 3D Gaussian SplattingabstractBy adaptively controlling the density and generating more Gaussians in regions with high-frequency information, 3D Gaussian Splatting (3DGS) can better represent scene details. From the signal processing perspective, representing details usually needs more Gaussians with relatively smaller scales. However, 3DGS currently lacks an explicit constraint linking the density and scale of 3D Gaussians across the domain, leading to 3DGS using improper-scale Gaussians to express frequency information, resulting in the loss of accuracy. In this paper, we propose to establish a direct relation between density and scale through the reparameterization of the scaling parameters and ensure the consistency between them via explicit constraints (i.e., density responds well to changes in frequency). Furthermore, we develop a frequency-aware density control strategy, consisting of densification and deletion, to improve representation quality with fewer Gaussians. A dynamic threshold encourages densification in high-frequency regions, while a scale-based filter deletes Gaussians with improper scale. Experimental results on various datasets demonstrate that our method outperforms existing state-of-the-art methods quantitatively and qualitatively. Zhaojie Zeng, Yuesong Wang 0001, Lili Ju |
AAAI | 2 |
| 2025 | IndoorGS: Geometric Cues Guided Gaussian Splatting for Indoor Scene Reconstructionabstract3D Gaussian Splatting (3DGS) has shown impressive performance in scene reconstruction, offering high rendering quality and rapid rendering speed with short training time. However, it often yields unsatisfactory results when applied to indoor scenes due to its poor ability to learn geometries without enough textural information. In this paper, we propose a new 3DGS-based method "IndoorGS", that leverages the commonly found yet important geometric cues in indoor scenes to improve the reconstruction quality. Specifically, we first extract 2D lines from input images and fuse them into 3D line cues via feature-based matching, which can provide a structural understanding of the target scene. We then apply the statistical outlier removal to refine Structure-from-Motion (SfM) points, ensuring robust cues in texture-rich areas. Based on these two types of cues, we further extract reliable 3D plane-like cues for textureless regions. Such geometric information will be utilized not only for initialization but also in the realization of a geometric-cue-guided adaptive density control (ADC) strategy. The proposed ADC approach is grounded in the principle of divide-and-conquer and optimizes the use of each type of geometric cues to enhance overall reconstruction performance. Extensive experiments on multiple indoor datasets show that our method can deliver much more accurate geometry and higher rendering quality for indoor scenes than existing 3DGS approaches. Cong Ruan, Yuesong Wang 0001, Lili Ju |
CVPR | 2 |
| 2025 | Instant Gaussianimage: A Generalizable and Self-Adaptive Image Representation via 2D Gaussian Splatting
Zhaojie Zeng, Yuesong Wang 0001, Lili Ju |
ICCV | 2 |
| 2024 | Entangled View-Epipolar Information Aggregation for Generalizable Neural Radiance FieldsabstractGeneralizable NeRF can directly synthesize novel views across new scenes, eliminating the need for scene-specific re-training in vanilla NeRF. A critical enabling factor in these approaches is the extraction of a generalizable 3D representation by aggregating source-view features. In this paper, we propose an Entangled View-Epipolar Information Aggregation method dubbed EVE-NeRF. Differentfrom existing methods that consider cross-view and along-epipolar information independently, EVE-NeRF conducts the view-epipolar feature aggregation in an entangled manner by injecting the scene-invariant appearance continuity and geometry consistency priors to the aggregation process. Our approach effectively mitigates the potential lack of inherent geometric and appearance constraints resulting from one-dimensional interactions, thus further boosting the 3D representation generalizability. EVE-NeRF attains state-of-the-art performance across various evaluation scenarios. Extensive experiments demonstrate that, compared to pre-vailing single-dimensional aggregation, the entangled network excels in the accuracy of 3D scene geometry and appearance reconstruction. Our code is publicly available at https://github.com/tatakai1/EVENeRF. Zhiyuan Min, Yawei Luo, Wei Yang 0011, Yuesong Wang 0001, Yi Yang 0001 |
CVPR | 4 |
| 2024 | ReinforceNS: Reinforcement Learning-based Multi-start Neighborhood Search for Solving the Traveling Thief Problem
Huachao Cui, Yuesong Wang 0001 |
IJCAI | 4 |
| 2023 | Adaptive Patch Deformation for Textureless-Resilient Multi-View StereoabstractIn recent years, deep learning-based approaches have shown great strength in multi-view stereo because of their outstanding ability to extract robust visual features. However, most learning-based methods need to build the cost volume and increase the receptive field enormously to get a satisfactory result when dealing with large-scale textureless regions, consequently leading to prohibitive memory consumption. To ensure both memory-friendly and textureless-resilient, we innovatively transplant the spirit of deformable convolution from deep learning into the traditional PatchMatch-based method. Specifically, for each pixel with matching ambiguity (termed unreliable pixel), we adaptively deform the patch centered on it to extend the receptive field until covering enough correlative reliable pixels (without matching ambiguity) that serve as anchors. When performing PatchMatch, constrained by the anchor pixels, the matching cost of an unreliable pixel is guaranteed to reach the global minimum at the correct depth and therefore increases the robustness of multi-view stereo significantly. To detect more anchor pixels to ensure better adaptive patch deformation, we propose to evaluate the matching ambiguity of a certain pixel by checking the convergence of the estimated depth as optimization proceeds. As a result, our method achieves state-of-the-art performance on ETH3D and Tanks and Temples while preserving low memory consumption. Yuesong Wang 0001, Zhaojie Zeng, Wei Yang 0011, Zhuo Chen 0054, Luoyuan Xu, Yawei Luo |
CVPR | 1 |
| 2023 | C2F2NeUS: Cascade Cost Frustum Fusion for High Fidelity and Generalizable Neural Surface ReconstructionabstractThere is an emerging effort to combine the two popular 3D frameworks using Multi-View Stereo (MVS) and Neural Implicit Surfaces (NIS) with a specific focus on the few-shot / sparse view setting. In this paper, we introduce a novel integration scheme that combines the multi-view stereo with neural signed distance function representations, which potentially overcomes the limitations of both methods. MVS uses per-view depth estimation and cross-view fusion to generate accurate surfaces, while NIS relies on a common coordinate volume. Based on this strategy, we propose to construct per-view cost frustum for finer geometry estimation, and then fuse cross-view frustums and estimate the implicit signed distance functions to tackle artifacts that are due to noise and holes in the produced surface reconstruction. We further apply a cascade frustum fusion strategy to effectively captures global-local information and structural consistency. Finally, we apply cascade sampling and a pseudo-geometric loss to foster stronger integration between the two architectures. Extensive experiments demonstrate that our method reconstructs robust surfaces and outperforms existing state-of-the-art methods. Luoyuan Xu, Yuesong Wang 0001, Zhaojie Zeng, Junle Wang, Wei Yang 0011 |
ICCV | 3 |
| 2022 | Self-Supervised Multi-view Stereo via Adjacent Geometry Guided Volume CompletionabstractExisting self-supervised multi-view stereo (MVS) approaches largely rely on photometric consistency for geometry inference, and hence suffer from low-texture or non-Lambertian appearances. In this paper, we observe that adjacent geometry shares certain commonality that can help to infer the correct geometry of the challenging or low-confident regions. Yet exploiting such property in a non-supervised MVS approach remains challenging for the lacking of training data and necessity of ensuring consistency between views. To address the issues, we propose a novel geometry inference training scheme by selectively masking regions with rich textures, where geometry can be well recovered and used for supervisory signal, and then lead a deliberately designed cost volume completion network to learn how to recover geometry of the masked regions. During inference, we then mask the low-confident regions instead and use the cost volume completion network for geometry correction. To deal with the different depth hypotheses of the cost volume pyramid, we design a three-branch volume inference structure for the completion network. Further, by considering plane as a special geometry, we first identify planar regions from pseudo labels and then correct the low-confident pixels by high-confident labels through plane normal consistency. Extensive experiments on DTU and Tanks & Temples demonstrate the effectiveness of the proposed framework and the state-of-the-art performance. Luoyuan Xu, Yuesong Wang 0001, Yawei Luo, Zhuo Chen 0054, Wei Yang 0011 |
ACM Multimedia | 3 |
| 2022 | Transformer-Based 3D Face Reconstruction With End-to-End Shape-Preserved Domain TransferabstractLearning-based face reconstruction methods have recently shown promising performance in recovering face geometry from a single image. However, the lack of training data with 3D annotations severely limits the performance. To tackle this problem, we proposed a novel end-to-end 3D face reconstruction network consisting of a conditional GAN (cGAN) for cross-domain face synthesis and a novel mesh transformer for face reconstruction. Our method first uses cGAN to translate the realistic face images to the specific rendered style, with a 2D facial edge consistency loss function. The domain-transferred images are then fed into face reconstruction network which uses a novel mesh transformer to output 3D mesh vertices. To exploit the domain-transferred in-the-wild images, we further propose a reprojection consistency loss to restrict face reconstruction network in a self-supervised way. Our approach can be trained with annotated dataset, synthetic dataset and in-the-wild images to learn a unified face model. Extensive experiments have demonstrated the effectiveness of our method. Zhuo Chen 0054, Yuesong Wang 0001, Luoyuan Xu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | DeepFusion: A simple way to improve traditional multi-view stereo methods using deep learning
Yuesong Wang 0001, Keyang Luo, Zhuo Chen 0054, Lili Ju |
Knowl. Based Syst. | 1 |
| 2020 | Attention-Aware Multi-View StereoabstractMulti-view stereo is a crucial task in computer vision, that requires accurate and robust photo-consistency among input images for depth estimation. Recent studies have shown that learning-based feature matching and confidence regularization can play a vital role in this task. Nevertheless, how to design good matching confidence volumes as well as effective regularizers for them are still under in-depth study. In this paper, we propose an attention-aware deep neural network “AttMVS” for learning multi-view stereo. In particular, we propose a novel attention-enhanced matching confidence volume, that combines the raw pixel-wise matching confidence from the extracted perceptual features with the contextual information of local scenes, to improve the matching robustness. Furthermore, we develop an attention-guided regularization module, which consists of multilevel ray fusion modules, to hierarchically aggregate and regularize the matching confidence volume into a latent depth probability volume.Experimental results show that our approach achieves the best overall performance on the DTU dataset and the intermediate sequences of Tanks & Temples benchmark over many state-of-the-art MVS algorithms. Keyang Luo, Lili Ju, Yuesong Wang 0001, Zhuo Chen 0054, Yawei Luo |
CVPR | 4 |
| 2020 | Mesh-Guided Multi-View Stereo With Pyramid ArchitectureabstractMulti-view stereo (MVS) aims to reconstruct 3D geometry of the target scene by using only information from 2D images. Although much progress has been made, it still suffers from textureless regions. To overcome this difficulty, we propose a mesh-guided MVS method with pyramid architecture, which makes use of the surface mesh obtained from coarse-scale images to guide the reconstruction process. Specifically, a PatchMatch-based MVS algorithm is first used to generate depth maps for coarse-scale images and the corresponding surface mesh is obtained by a surface reconstruction algorithm. Next we project the mesh onto each of depth maps to replace unreliable depth values and the corrected depth maps are fed to fine-scale reconstruction for initialization. To alleviate the influence of possible erroneous faces on the mesh, we further design and train a convolutional neural network to remove incorrect depths. In addition, it is often hard for the correct depth values for low-textured regions to survive at the fine-scale, thus we also develop an efficient method to seek out these regions and further enforce the geometric consistency in these regions. Experimental results on the ETH3D high-resolution dataset demonstrate that our method achieves state-of-the-art performance, especially in completeness. Yuesong Wang 0001, Zhuo Chen 0054, Yawei Luo, Keyang Luo, Lili Ju |
CVPR | 1 |
| 2020 | PC-Net: A Deep Network for 3D Point Clouds AnalysisabstractDue to the irregularity and sparsity of 3D point clouds, applying convolutional neural networks directly on them can be nontrivial. In this work, we propose a simple but effective approach for 3D Point Clouds analysis, named PC-Net. PC-Net directly learns on point sets and is equipped with three new operations: first, we apply a novel scale-aware neighbor search for adaptive neighborhood extracting; second, for each neighboring point, we learn a local spatial feature as a complement to their associated features; finally, at the end we use a distance re-weighted pooling to aggregate all the features from local structure. With this module, we design hierarchical neural network for point cloud understanding. For both classification and segmentation tasks, our architecture proves effective in the experiments and our models demonstrate state-of-the-art performance over existing deep learning methods on popular point cloud benchmarks. Zhuo Chen 0054, Yawei Luo, Yuesong Wang 0001, Keyang Luo, Luoyuan Xu |
ICPR | 4 |
| 2016 | Accurate localization for mobile device using a multi-planar city modelabstractThis paper presents a novel method for estimating the unknown 6DOF pose of a mobile device. The method is based on matching between the mobile image and the virtual city model which is merely composed of 3D points on planar building facade. The main contributions of this paper are as follows: firstly, we design a new plane generation strategy which fuses the 3D model points, photo homography and the orientation of the buildings together within RANSAC framework. Secondly, we propose a novel energy-based method which can parallel solve the mobile poses as well as the best 2D-3D matches. Thirdly, a client/server mode is established to support a speedy localization experience on the mobile devices. To the best of our knowledge, this is the first implementation that uses such a multi-planar model to accurately locate the mobile device in a scene of city scale. Experiment shows that the localization performance becomes faster and more robust comparing with other methods when adding the multi-planar information of the buildings to localization algorithm. Yawei Luo, Hailong Pan, Yuesong Wang 0001, Junqing Yu |
ICPR | 4 |
| 2015 | On-Device Mobile Landmark Recognition Using Binarized Descriptor with Multifeature FusionabstractAlong with the exponential growth of high-performance mobile devices, on-device Mobile Landmark Recognition (MLR) has recently attracted increasing research attention. However, the latency and accuracy of automatic recognition remain as bottlenecks against its real-world usage. In this article, we introduce a novel framework that combines interactive image segmentation with multifeature fusion to achieve improved MLR with high accuracy. First, we propose an effective vector binarization method to reduce the memory usage of image descriptors extracted on-device, which maintains comparable recognition accuracy to the original descriptors. Second, we design a location-aware fusion algorithm that can fuse multiple visual features into a compact yet discriminative image descriptor to improve on-device efficiency. Third, a user-friendly interaction scheme is developed that enables interactive foreground/background segmentation to largely improve recognition accuracy. Experimental results demonstrate the effectiveness of the proposed algorithms for on-device MLR applications. Yuesong Wang 0001, Liya Duan, Rongrong Ji |
ACM Trans. Intell. Syst. Technol. | 2 |