Yuanbo Xiangli

dblp:186/4450 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Computer networks · 2Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 DCSHARP: 3D Gaussian Splatting with Direction Cosine Spherical Harmonics and Shape-Aware Pruning
abstract
3D Gaussian Splatting (3DGS) shows outstanding rendering quality for novel view synthesis. Despite its performance, the massive amount of Gaussian blobs leads to expensive run-time sorting and irregular memory access during rendering. Although 3DGS-based pruning algorithm has been widely explored, most of the current research has mainly focused on designing a proper pruning metric and the root cause behind the inevitable quality degradation remains underexplored for highly-sparse 3DGS. In particular, our investigation shows that the Spherical Harmonics (SH) of 3DGS is insufficient to capture high-frequency anisotropic reflections and specular highlights during rendering, especially with sparsified Gaussians. Motivated by that, this work proposes Direction Cosine Spherical Harmonics with Shape-Aware Pruning (DCSHARP). Specifically, the proposed Direction Cosine Spherical Harmonics (DCSH) replaces the vanilla spherical harmonics by facilitating the expressiveness of 3DGS on high-frequency and highly reflective scenes. Unlike recent works that rely on trainable masks or pseudo-rendering scores, the proposed Shape-aware Pruning method enables "pruning on-the-fly" while achieving high quality rendering. As a combined scheme, the proposed DCSHARP reduces the number of active Gaussians by up to 3.9× and improves rendering throughput by 1.9× with ZERO quality degradation compared to the vanilla 3DGS. Furthermore, the proposed DCSH scheme outperforms the vanilla 3DGS on all the mainstream benchmarks by simply replacing the vanilla SH with the DCSH. The source code of the proposed method will be open-sourced.
Ahmed Hasssan, Jian Meng, Yuanbo Xiangli, Jae-sun Seo
WACV3
2025 Doppelgangers++: Improved Visual Disambiguation with Geometric 3D Features
abstract
Accurate 3D reconstruction is frequently hindered by visual aliasing, where visually similar but distinct surfaces (aka, doppelgangers), are incorrectly matched. These spurious matches distort the structure-from-motion (SfM) process, leading to misplaced model elements and reduced accuracy. Prior efforts addressed this with CNN classifiers trained on curated datasets, but these approaches struggle to generalize across diverse real-world scenes and can require extensive parameter tuning. In this work, we present Doppelgangers++, a method to enhance doppelganger detection and improve 3D reconstruction accuracy. Our contributions include a diversified training dataset that incorporates geo-tagged images from everyday scenes to expand robustness beyond landmark-based datasets. We further propose a Transformer-based classifier that leverages 3D-aware features from the MASt3R model, achieving superior precision and recall across both in-domain and out-of-domain tests. Doppelgangers++ integrates seamlessly into standard SfM and MASt3R-SfM pipelines, offering efficiency and adaptability across varied scenes. To evaluate SfM accuracy, we introduce an automated, geotag-based method for validating reconstructed models, eliminating the need for manual inspection. Through extensive experiments, we demonstrate that Doppelgangers++ significantly enhances pairwise vi sual disambiguation and improves 3D reconstruction quality in complex and diverse scenarios.
Yuanbo Xiangli, Ruojin Cai, Hanyu Chen 0002, Jeffrey Byrne, Noah Snavely
CVPR1
2024 Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering
abstract
Neural rendering methods have significantly advanced photo-realistic 3D scene rendering in various academic and industrial applications. The recent 3D Gaussian Splatting method has achieved the state-of-the-art rendering quality and speed combining the benefits of both primitive-based representations and volumetric representations. However, it often leads to heavily redundant Gaussians that try to fit every training view, neglecting the underlying scene ge-ometry. Consequently, the resulting model becomes less robust to significant view changes, texture-less area and lighting effects. We introduce Scaffold-GS, which uses an-chor points to distribute local 3D Gaussians, and predicts their attributes on-the-fly based on viewing direction and distance within the view frustum. Anchor growing and pruning strategies are developed based on the importance of neural Gaussians to reliably improve the scene cover-age. We show that our method effectively reduces redun-dant Gaussians while delivering high-quality rendering. We also demonstrates an enhanced capability to accommodate scenes with varying levels-of-detail and view-dependent ob-servations, without sacrificing the rendering speed. Project page: https://city-super.github.iolscaffold-gsl.
Tao Lu 0005, Mulin Yu, Linning Xu, Yuanbo Xiangli, Limin Wang 0002, Dahua Lin, Bo Dai 0002
CVPR4
2024 GS-LRM: Large Reconstruction Model for 3D Gaussian Splatting
Kai Zhang 0045, Sai Bi, Hao Tan 0002, Yuanbo Xiangli, Nanxuan Zhao, Kalyan Sunkavalli, Zexiang Xu
ECCV (22)4
2024 Neural Gaffer: Relighting Any Object via Diffusion
abstract
Single-image relighting is a challenging task that involves reasoning about the complex interplay between geometry, materials, and lighting. Many prior methods either support only specific categories of images, such as portraits, or require special capture conditions, like using a flashlight. Alternatively, some methods explicitly decompose a scene into intrinsic components, such as normals and BRDFs, which can be inaccurate or under-expressive. In this work, we propose a novel end-to-end 2D relighting diffusion model, called Neural Gaffer, that takes a single image of any object and can synthesize an accurate, high-quality relit image under any novel environmental lighting condition, simply by conditioning an image generator on a target environment map, without an explicit scene decomposition. Our method builds on a pre-trained diffusion model, and fine-tunes it on a synthetic relighting dataset, revealing and harnessing the inherent understanding of lighting present in the diffusion model. We evaluate our model on both synthetic and in-the-wild Internet imagery and demonstrate its advantages in terms of generalization and accuracy. Moreover, by combining with other generative methods, our model enables many downstream 2D tasks, such as text-based relighting and object insertion. Our model can also operate as a strong relighting prior for 3D tasks, such as relighting a radiance field.
Haian Jin, Fujun Luan, Yuanbo Xiangli, Sai Bi, Kai Zhang 0045, Zexiang Xu, Jin Sun 0009, Noah Snavely
NeurIPS4
2024 GSDF: 3DGS Meets SDF for Improved Neural Rendering and Reconstruction
abstract
Representing 3D scenes from multiview images remains a core challenge in computer vision and graphics, requiring both reliable rendering and reconstruction, which often conflicts due to the mismatched prioritization of image quality over precise underlying scene geometry. Although both neural implicit surfaces and explicit Gaussian primitives have advanced with neural rendering techniques, current methods impose strict constraints on density fields or primitive shapes, which enhances the affinity for geometric reconstruction at the sacrifice of rendering quality. To address this dilemma, we introduce GSDF, a dual-branch architecture combining 3D Gaussian Splatting (3DGS) and neural Signed Distance Fields (SDF). Our approach leverages mutual guidance and joint supervision during the training process to mutually enhance reconstruction and rendering. Specifically, our method guides the Gaussian primitives to locate near potential surfaces and accelerates the SDF convergence. This implicit mutual guidance ensures robustness and accuracy in both synthetic and real-world scenarios. Experimental results demonstrate that our method boosts the SDF optimization process to reconstruct more detailed geometry, while reducing floaters and blurry edge artifacts in rendering by aligning Gaussian primitives with the underlying geometry.
Mulin Yu, Tao Lu 0005, Linning Xu, Lihan Jiang, Yuanbo Xiangli, Bo Dai 0002
NeurIPS5
2023 OmniCity: Omnipotent City Understanding with Multi-Level and Multi-View Images
abstract
This paper presents OmniCity, a new dataset for omnipotent city understanding from multi-level and multi-view images. More precisely, OmniCity contains multi-view satellite images as well as street-level panorama and mono-view images, constituting over 100K pixel-wise annotated images that are well-aligned and collected from 25K geo-locations in New York City. To alleviate the substantial pixel-wise annotation efforts, we propose an efficient street-viewimage annotation pipeline that leverages the existing label maps of satellite view and the transformation relations between different views (satellite, panorama, and mono-view). With the new OmniCity dataset, we provide benchmarks for a variety of tasks including building footprint extraction, height estimation, and building plane/instance/fine-grained segmentation. Compared with existing multi-level and multi-view benchmarks, OmniCity contains a larger number of images with richer annotation types and more views, provides more benchmark results of state-of-the-art models, and introduces a new task for fine-grained building instance segmentation on street-level panorama images. Moreover, OmniCity provides new problem settings for existing tasks, such as cross-view image matching, synthesis, segmentation, detection, etc., and facilitates the developing of new methods for large-scale city understanding, reconstruction, and simulation. The OmniCity dataset as well as the benchmarks will be released at https://city-super.github.io/mnicity/.
Yawen Lai, Linning Xu, Yuanbo Xiangli, Conghui He, Gui-Song Xia, Dahua Lin
CVPR4
2023 Grid-guided Neural Radiance Fields for Large Urban Scenes
abstract
Purely MLP-based neural radiance fields (NeRF-based methods) often suffer from underfitting with blurred renderings on large-scale scenes due to limited model capacity. Recent approaches propose to geographically divide the scene and adopt multiple sub-NeRFs to model each region individually, leading to linear scale-up in training costs and the number of sub-NeRFs as the scene expands. An alternative solution is to use a feature grid representation, which is computationally efficient and can naturally scale to a large scene with increased grid resolutions. However, the feature grid tends to be less constrained and often reaches suboptimal solutions, producing noisy artifacts in renderings, especially in regions with complex geometry and texture. In this work, we present a new framework that realizes high-fidelity rendering on large urban scenes while being computationally efficient. We propose to use a compact multi-resolution ground feature plane representation to coarsely capture the scene, and complement it with positional encoding inputs through another NeRF branch for rendering in a joint learning fashion. We show that such an integration can utilize the advantages of two alternative solutions: a light-weighted NeRF is sufficient, under the guidance of the feature grid representation, to render photorealistic novel views with fine details; and the jointly optimized ground feature planes, can meanwhile gain further refinements, forming a more accurate and compact feature space and output much more natural rendering results.
Linning Xu, Yuanbo Xiangli, Sida Peng, Xingang Pan, Nanxuan Zhao, Christian Theobalt, Bo Dai 0002, Dahua Lin
CVPR2
2023 MatrixCity: A Large-scale City Dataset for City-scale Neural Rendering and Beyond
abstract
Neural radiance fields (NeRF) and its subsequent variants have led to remarkable progress in neural rendering. While most of recent neural rendering works focus on objects and small-scale scenes, developing neural rendering methods for city-scale scenes is of great potential in many real-world applications. However, this line of research is impeded by the absence of a comprehensive and high-quality dataset, yet collecting such a dataset over real city-scale scenes is costly, sensitive, and technically infeasible. To this end, we build a large-scale, comprehensive, and high-quality synthetic dataset for city-scale neural rendering researches. Leveraging the Unreal Engine 5 City Sample project, we developed a pipeline to easily collect aerial and street city views, accompanied by ground-truth camera poses and a range of additional data modalities. Flexible controls on environmental factors like light, weather, human and car crowd are also available in our pipeline, supporting the need of various tasks covering city-scale neural rendering and beyond. The resulting pilot dataset, MatrixCity, contains 67k aerial images and 452k street images from two city maps of total size 28km2. On top of MatrixCity, a thorough benchmark is also conducted, which not only reveals unique challenges of the task of city-scale neural rendering, but also highlights potential improvements for future works. The dataset and code will be publicly available at the project page: https://city-super.github.io/matrixcity/.
Yixuan Li 0002, Lihan Jiang, Linning Xu, Yuanbo Xiangli, Zhenzhi Wang 0001, Dahua Lin, Bo Dai 0002
ICCV4
2023 AssetField: Assets Mining and Reconfiguration in Ground Feature Plane Representation
abstract
Both indoor and outdoor environments are inherently structured and repetitive. Traditional modeling pipelines keep an asset library storing unique object templates, which is both versatile and memory efficient in practice. Inspired by this observation, we propose AssetField, a novel neural scene representation that learns a set of object-aware ground feature planes to represent the scene, where an asset library storing template feature patches can be constructed in an unsupervised manner. Unlike existing methods which require object masks to query spatial points for object editing, our ground feature plane representation offers a natural visualization of the scene in the bird-eye view, allowing a variety of operations (e.g. translation, duplication, deformation) on objects to configure a new scene. With the template feature patches, group editing is enabled for scenes with many recurring items to avoid repetitive work on object individuals. We show that AssetField not only achieves competitive performance for novel-view synthesis but also generates realistic renderings for new scene configurations.
Yuanbo Xiangli, Linning Xu, Xingang Pan, Nanxuan Zhao, Bo Dai 0002, Dahua Lin
ICCV1
2022 BungeeNeRF: Progressive Neural Radiance Field for Extreme Multi-scale Scene Rendering
Yuanbo Xiangli, Linning Xu, Xingang Pan, Nanxuan Zhao, Anyi Rao, Christian Theobalt, Bo Dai 0002, Dahua Lin
ECCV (32)1
2021 BlockPlanner: City Block Generation with Vectorized Graph Representation
abstract
City modeling is the foundation for computational urban planning, navigation, and entertainment. In this work, we present the first generative model of city blocks named BlockPlanner, and showcase its ability to synthesize valid city blocks with varying land lots configurations. We propose a novel vectorized city block representation utilizing a ring topology and a two-tier graph to capture the global and local structures of a city block. Each land lot is abstracted into a vector representation covering both its 3D geometry and land use semantics. Such vectorized representation enables us to deploy a lightweight network to capture the underlying distribution of land lots configurations in a city block. To enforce intrinsic spatial constraints of a valid city block, a set of effective loss functions are imposed to shape rational results. We contribute a pilot city block dataset to demonstrate the effectiveness and efficiency of our representation and framework over the state-of-the-art. Notably, our BlockPlanner is also able to edit and manipulate city blocks, enabling several useful applications, e.g., topology refinement and footprint generation.
Linning Xu, Yuanbo Xiangli, Anyi Rao, Nanxuan Zhao, Bo Dai 0002, Ziwei Liu 0002, Dahua Lin
ICCV2
2020 Real or Not Real, that is the Question
Yuanbo Xiangli, Yubin Deng, Bo Dai 0002, Chen Change Loy, Dahua Lin
ICLR1
2020 Nowhere to Hide: Cross-modal Identity Leakage between Biometrics and Devices
abstract
Along with the benefits of Internet of Things (IoT) come potential privacy risks, since billions of the connected devices are granted permission to track information about their users and communicate it to other parties over the Internet. Of particular interest to the adversary is the user identity which constantly plays an important role in launching attacks. While the exposure of a certain type of physical biometrics or device identity is extensively studied, the compound effect of leakage from both sides remains unknown in multi-modal sensing environments. In this work, we explore the feasibility of the compound identity leakage across cyber-physical spaces and unveil that co-located smart device IDs (e.g., smartphone MAC addresses) and physical biometrics (e.g., facial/vocal samples) are side channels to each other. It is demonstrated that our method is robust to various observation noise in the wild and an attacker can comprehensively profile victims in multi-dimension with nearly zero analysis effort. Two real-world experiments on different biometrics and device IDs show that the presented approach can compromise more than 70% of device IDs and harvests multiple biometric clusters with purity at the same time.
Xiaoxuan Lu 0001, Yang Li 0073, Yuanbo Xiangli, Zhengxiong Li
WWW3
2019 Autonomous Learning of Speaker Identity and WiFi Geofence From Noisy Sensor Data
abstract
A fundamental building block toward intelligent environments is the ability to understand who is present in a certain area. A ubiquitous way of detecting this is to exploit unique vocal characteristics as people interact with one another in common spaces. However, manually enrolling users into a biometric database is time-consuming and not robust to vocal deviations over time. Instead, consider audio features sampled during a meeting, yielding a noisy set of possible voiceprints. With a number of meetings and knowledge of participation, e.g., sniffed wireless media access control (MAC) addresses, can we learn to associate a specific identity with a particular voiceprint? To address this problem, this paper advocates an Internet of Things (IoT) solution and proposes to use co-located WiFi as supervisory weak labels to automatically bootstrap the labeling process. In particular, a novel cross-modality labeling algorithm is proposed that jointly optimizes the clustering and association process, which solves the inherent mismatching issues arising from heterogeneous sensor data. At the same time, we further propose to reuse the labeled data to iteratively update wireless geofence models and curate device specific thresholds. The extensive experimental results from two different scenarios demonstrate that our proposed method is able to achieve twofold improvement in labeling compared with conventional methods and can achieve reliable speaker recognition in the wild.
Xiaoxuan Lu 0001, Yuanbo Xiangli, Peijun Zhao, Changhao Chen, Agathoniki Trigoni, Andrew Markham
IEEE Internet Things J.2
2016 In-device, runtime cellular network information extraction and analysis: demo
abstract
We present the demonstration of MobileInsight, a software tool that collects, analyzes and exploits runtime information from operational cellular network. MobileInsight runs on commercial off-the-shelf phones without extra hardware or additional support from cellular network operators. It exposes cellular protocol messages from the 3G/4G chipset, and performs in-device protocol analysis. We demonstrate the in-device runtime cellular message collection, analysis of the protocol states, visualization of runtime wireless channel and mobility dynamics, and how mobile applications benefit from MobileInsight.
Yuanjie Li, Haotian Deng 0001, Yuanbo Xiangli, Zengwen Yuan, Chunyi Peng 0001, Songwu Lu
MobiCom3