Yunsong Wang

dblp:182/0203 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Theory of computation · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FreeSplat++: Generalizable 3D Gaussian Splatting for Efficient Indoor Scene Reconstruction
abstract
Recently, the integration of the efficient feed-forward scheme into 3D Gaussian Splatting (3DGS) has been actively explored. However, most existing methods focus on sparse view reconstruction of small regions and cannot produce eligible whole-scene reconstruction results in terms of either quality or efficiency. In this paper, we propose FreeSplat++, which focuses on extending the generalizable 3DGS to become an alternative approach to large-scale indoor whole-scene reconstruction, which has the potential of significantly accelerating the reconstruction speed and improving the geometric accuracy. To facilitate whole-scene reconstruction, we initially propose the Low-cost Cross-View Aggregation framework to efficiently process extremely long input sequences. Subsequently, we introduce a carefully designed pixel-wise triplet fusion method to incrementally aggregate the overlapping 3D Gaussian primitives from multiple views, adaptively reducing their redundancy. Furthermore, given the fused 3DGS primitives with accumulated weights after the fusion step, we propose a weighted floater removal strategy that can effectively reduce floaters, which serves as an explicit depth fusion approach that is tailored for generalizable 3DGS methods and becomes crucial in whole-scene reconstruction. After the feed-forward reconstruction of 3DGS primitives, we investigate a depth-regularized per-scene fine-tuning process. Leveraging the dense, multi-view consistent depth maps obtained during the feed-forward prediction phase for an extra constraint, we refine the entire scene's 3DGS primitive to enhance rendering quality while preserving geometric accuracy. Extensive experiments confirm that our FreeSplat++ significantly outperforms existing generalizable 3DGS methods, especially in whole scene reconstructions. Compared to conventional per-scene optimized 3DGS approaches, our method with depth-regularized per-scene fine-tuning demonstrates substantial improvements in reconstruction accuracy and a notable reduction in training time.
Yunsong Wang, Tianxin Huang, Gim Hee Lee
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Enhancing Generalizability of Representation Learning for Data-Efficient 3D Scene Understanding
abstract
The field of self-supervised $3 D$ representation learning has emerged as a promising solution to alleviate the challenge presented by the scarcity of extensive, well-annotated datasets. However, it continues to be hindered by the lack of diverse, large-scale, real-world 3D scene datasets for source data. To address this shortfall, we propose Generalizable Representation Learning (GRL), where we devise a generative Bayesian network to produce diverse synthetic scenes with real-world patterns, and conduct pre-training with a joint objective. By jointly learning a coarse-to-fine contrastive learning task and an occlusion-aware reconstruction task, the model is primed with transferable, geometry-informed representations. Post pre-training on synthetic data, the acquired knowledge of the model can be seamlessly transferred to two principal downstream tasks associated with 3D scene understanding, namely 3D object detection and 3D semantic segmentation, using real-world benchmark datasets. A thorough series of experiments robustly display our method’s consistent superiority over existing state-of-the-art pre-training approaches.
Yunsong Wang, Na Zhao 0004, Gim Hee Lee
3DV1
2024 Syn-to-Real Unsupervised Domain Adaptation for Indoor 3D Object Detection
Yunsong Wang, Na Zhao 0004, Gim Hee Lee
BMVC1
2024 GOV-NeSF: Generalizable Open-Vocabulary Neural Semantic Fields
abstract
Recent advancements in vision-language foundation models have significantly enhanced open-vocabulary 3D scene understanding. However, the generalizability of existing methods is constrained due to their framework designs and their reliance on 3D data. We address this limitation by introducing Generalizable Open-Vocabulary Neural Semantic Fields (GOV-NeSF), a novel approach offering a generalizable implicit representation of 3D scenes with open-vocabulary semantics. We aggregate the geometry-aware features using a cost volume, and propose a Multi-view Joint Fusion module to aggregate multi-view features through a cross-view attention mechanism, which effectively predicts view-specific blending weights for both colors and open-vocabulary features. Remarkably, our GOV-NeSF exhibits state-of-the-art performance in both 2D and 3D open-vocabulary semantic segmentation, eliminating the need for ground truth semantic labels or depth priors, and effectively generalize across scenes and datasets without fine-tuning.
Yunsong Wang, Gim Hee Lee
CVPR1
2024 VCR-GauS: View Consistent Depth-Normal Regularizer for Gaussian Surface Reconstruction
abstract
Although 3D Gaussian Splatting has been widely studied because of its realistic and efficient novel-view synthesis, it is still challenging to extract a high-quality surface from the point-based representation. Previous works improve the surface by incorporating geometric priors from the off-the-shelf normal estimator. However, there are two main limitations: 1) Supervising normal rendered from 3D Gaussians updates only the rotation parameter while neglecting other geometric parameters; 2) The inconsistency of predicted normal maps across multiple views may lead to severe reconstruction artifacts. In this paper, we propose a Depth-Normal regularizer that directly couples normal with other geometric parameters, leading to full updates of the geometric parameters from normal regularization. We further propose a confidence term to mitigate inconsistencies of normal predictions across multiple views. Moreover, we also introduce a densification and splitting strategy to regularize the size and distribution of 3D Gaussians for more accurate surface modeling. Compared with Gaussian-based baselines, experiments show that our approach obtains better reconstruction quality and maintains competitive appearance quality at faster training speed and 100+ FPS rendering. Our code will be made open-source upon paper acceptance.
Fangyin Wei, Chen Li 0038, Tianxin Huang, Yunsong Wang, Gim Hee Lee
NeurIPS5
2024 FreeSplat: Generalizable 3D Gaussian Splatting Towards Free View Synthesis of Indoor Scenes
abstract
Empowering 3D Gaussian Splatting with generalization ability is appealing. However, existing generalizable 3D Gaussian Splatting methods are largely confined to narrow-range interpolation between stereo images due to their heavy backbones, thus lacking the ability to accurately localize 3D Gaussian and support free-view synthesis across wide view range. In this paper, we present a novel framework FreeSplat that is capable of reconstructing geometrically consistent 3D scenes from long sequence input towards free-view synthesis.Specifically, we firstly introduce Low-cost Cross-View Aggregation achieved by constructing adaptive cost volumes among nearby views and aggregating features using a multi-scale structure. Subsequently, we present the Pixel-wise Triplet Fusion to eliminate redundancy of 3D Gaussians in overlapping view regions and to aggregate features observed across multiple views. Additionally, we propose a simple but effective free-view training strategy that ensures robust view synthesis across broader view range regardless of the number of views. Our empirical results demonstrate state-of-the-art novel view synthesis peformances in both novel view rendered color maps quality and depth maps accuracy across different numbers of input views. We also show that FreeSplat performs inference more efficiently and can effectively reduce redundant Gaussians, offering the possibility of feed-forward large scene reconstruction without depth priors. Our code will be made open-source upon paper acceptance.
Yunsong Wang, Tianxin Huang, Gim Hee Lee
NeurIPS1
2024 Rethinking Visibility in Human Pose Estimation: Occluded Pose Reasoning via Transformers
abstract
Occlusion is a common challenge in human pose estimation. Curiously, learning from occluded keypoints hinders a model to detect visible keypoints. We speculate that the impairment is likely due to a forced correlation between keypoints and visual features of the occluders. As such, we propose a novel visibility-aware attention mechanism to eliminate unreliable occluding features. The explicit occlusion handling encourages the model to reason about occluded keypoints using evidence and contextual information from the visible keypoints. It also mitigates the damage of unreliable correlations of the occluded keypoints. Our method, when added to the strong baseline SimCC, improves by 1.3 AP and 0.7 AP with ResNet and HRNet respectively. It also surpasses the state-of-the-art I2R-Net on CrowdPose by 0.3 AP and 0.6 APhard. The improvements highlight that rethinking visibility information is critical for developing effective human pose estimation systems.
Pengzhan Sun 0001, Kerui Gu, Yunsong Wang, Linlin Yang 0001, Angela Yao
WACV3
2023 Accelerating User-Defined Aggregate Functions (UDAF) with Block-wide Execution and JIT Compilation on GPUs
abstract
The GPU-accelerated DataFrame library cuDF has become increasingly popular for data analytics applications due to its superior performance against CPU-based DataFrame libraries such as Pandas. One of the frequently-used operations in dataframe manipulation is user-defined aggregate functions (UDAFs). UDAFs allow users to define custom aggregate routines outside of the pre-defined aggregate operations (Sum(), Max(), Avg(), etc.)
Bobbi W. Yogatama, Brandon Miller, Yunsong Wang, Graham R. Markall, Jacob Hemstad, Gregory Kimball, Xiangyao Yu
DaMoN3
2022 An Epistemic Interpretation of Tensor Disjunction
Yanjing Wang 0001, Yunsong Wang
AiML2
2022 Inquisitive logic as an epistemic logic of knowing how
Yanjing Wang 0001, Yunsong Wang
Ann. Pure Appl. Log.3
2021 Instance Similarity Learning for Unsupervised Feature Representation
abstract
In this paper, we propose an instance similarity learning (ISL) method for unsupervised feature representation. Conventional methods assign close instance pairs in the feature space with high similarity, which usually leads to wrong pairwise relationship for large neighborhoods because the Euclidean distance fails to depict the true semantic similarity on the feature manifold. On the contrary, our method mines the feature manifold in an unsupervised manner, through which the semantic similarity among instances is learned in order to obtain discriminative representations. Specifically, we employ the Generative Adversarial Networks (GAN) to mine the underlying feature manifold, where the generated features are applied as the proxies to progressively explore the feature manifold so that the semantic similarity among instances is acquired as reliable pseudo supervision. Extensive experiments on image classification demonstrate the superiority of our method compared with the state-of-the-art methods. The code is available at https://github.com/ZiweiWangTHU/ISL.git.
Ziwei Wang 0010, Yunsong Wang, Ziyi Wu 0002, Jiwen Lu, Jie Zhou 0001
ICCV2
2021 A 320×240 I-ToF CMOS Image Sensor with 2-Tap 5.6µm Pixel and Mismatch-Nonlinearity Suppression
abstract
This paper presents a 320×240 indirect time of flight (I-ToF) image sensor with 5.6μm×5.6μm 2-Tap pixel in 110nm process. The readout channel offset cancellation and nonlinearity suppression techniques are proposed to achieve high-precision detection. The measured relative precision is 1% at a 5m target distance and non-linearity is below 1.02%. The chip also integrates LVDS and I2C interface for data transmission and Laser control. This work effectively improved the ranging accuracy with a simple method.
Youze Xin, Bing Zhang 0019, Congzhen Hu, Li Dong 0007, Dan Li 0011, Yunsong Wang, Shuyu Lei, Li Geng
ISCAS6