EDBT 2026 Demo / reviewers in the wild / expert
Yuliang Guo
dblp:117/8269
· DBLP profile ↗
19ranked-venue papers
5as first author
15since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 10 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Depth Any Camera: Zero-Shot Metric Depth Estimation from Any CameraabstractWhile recent depth foundation models exhibit strong zero-shot generalization, achieving accurate metric depth across diverse camera types—particularly those with large fields of view (FoV) such as fisheye and 360-degree cameras—remains a significant challenge. This paper presents Depth Any Camera (DAC), a powerful zero-shot metric depth estimation framework that extends a perspective-trained model to effectively handle cameras with varying FoVs. The framework is designed to ensure that all existing 3D data can be leveraged, regardless of the specific camera types used in new applications. Remarkably, DAC is trained exclusively on perspective images but generalizes seamlessly to fisheye and 360-degree cameras without the need for specialized training data. DAC employs Equi-Rectangular Projection (ERP) as a unified image representation, enabling consistent processing of images with diverse FoVs. Its core components include pitch-aware Image-to-ERP conversion with efficient online augmentation to simulate distorted ERP patches from undistorted inputs, FoV alignment operations to enable effective training across a wide range of FoVs, and multi-resolution data augmentation to further address resolution disparities between training and testing. DAC achieves state-of-the-art zero-shot metric depth estimation, improving δ1accuracy by up to 50% on multiple fisheye and 360-degree datasets compared to prior metric depth foundation models, demonstrating robust generalization across camera types. Yuliang Guo, Sparsh Garg, S. Mahdi H. Miangoleh, Xinyu Huang 0001, Liu Ren 0001 |
CVPR | 1 |
| 2025 | Online Language SplattingabstractTo enable AI agents to interact seamlessly with both humans and 3D environments, they must not only perceive the 3D world accurately but also align human language with 3D spatial representations. While prior work has made significant progress by integrating language features into geometrically detailed 3D scene representations using 3D Gaussian Splatting (GS), these approaches rely on computationally intensive offline preprocessing of language features for each input image, limiting adaptability to new environments. In this work, we introduce Online Language Splatting, the first framework to achieve online, near real-time, open-vocabulary language mapping within a 3DGS-SLAM system without requiring pre-generated language features. The key challenge lies in efficiently fusing high-dimensional language features into 3D representations while balancing the computation speed, memory usage, rendering quality and open-vocabulary capability. To this end, we innovatively design: (1) a high-resolution CLIP embedding module capable of generating detailed language feature maps in 18ms per frame, (2) a two-stage online auto-encoder that compresses 768-dimensional CLIP features to 15 dimensions while preserving open-vocabulary capabilities, and (3) a color-language disentangled optimization approach to improve rendering quality. Experimental results show that our online method not only surpasses the state-of-the-art offline methods in accuracy but also achieves more than 40x efficiency boost, demonstrating the potential for dynamic and interactive AI applications. Saimouli Katragadda, Cho-Ying Wu, Yuliang Guo, Xinyu Huang 0001, Guoquan Huang 0001, Liu Ren 0001 |
ICCV | 3 |
| 2025 | CHARM3R: Towards Unseen Camera Height Robust Monocular 3D Detector
Abhinav Kumar 0004, Yuliang Guo, Xinyu Huang 0001, Liu Ren 0001, Xiaoming Liu 0002 |
ICCV | 2 |
| 2025 | SMART: Advancing Scalable Map Priors for Driving Topology ReasoningabstractTopology reasoning is crucial for autonomous driving as it enables comprehensive understanding of connec-tivity and relationships between lanes and traffic elements. While recent approaches have shown success in perceiving driving topology using vehicle-mounted sensors, their scalability is hindered by the reliance on training data captured by consistent sensor configurations. We identify that the key factor in scalable lane perception and topology reasoning is the elimination of this sensor-dependent feature. To address this, we propose SMART, a scalable solution that leverages easily available standard-definition (SD) and satellite maps to learn a map prior model, supervised by large-scale geo-referenced high-definition (HD) maps independent of sensor settings. Attributed to scaled training, SMART alone achieves superior offline lane topology understanding using only SD and satellite inputs. Extensive experiments further demonstrate that SMART can be seamlessly integrated into any online topology reasoning methods, yielding significant improvements of up to 28% on the OpenLane-V2 benchmark. Project page: https://jay-ye.github.io/smart. Junjie Ye 0007, David Paz, Hengyuan Zhang 0001, Yuliang Guo, Xinyu Huang 0001, Henrik I. Christensen, Yue Wang 0041, Liu Ren 0001 |
ICRA | 4 |
| 2025 | MapGS: Generalizable Pretraining and Data Augmentation for Online Mapping via Novel View SynthesisabstractOnline mapping reduces the reliance of au-tonomous vehicles on high-definition (HD) maps, significantly enhancing scalability. However, recent advancements often overlook cross-sensor configuration generalization, leading to performance degradation when models are deployed on vehicles with different camera intrinsics and extrinsics. With the rapid evolution of novel view synthesis methods, we investigate the extent to which these techniques can be leveraged to address the sensor configuration generalization challenge. We propose a novel framework leveraging Gaussian splatting to reconstruct scenes and render camera images in target sensor configurations. The target config sensor data, along with labels mapped to the target config, are used to train online mapping models. Our proposed framework on the nuScenes and Ar-goverse 2 datasets demonstrates a performance improvement of 18 % through effective dataset augmentation, achieves faster convergence and efficient training, and exceeds state-of-the-art performance when using only 25 % of the original training data. This enables data reuse and reduces the need for laborious data labeling. Project page at https://henryzhangzhy.github.io/mapgs. Hengyuan Zhang 0001, David Paz, Yuliang Guo, Xinyu Huang 0001, Henrik I. Christensen, Liu Ren 0001 |
IV | 3 |
| 2024 | Multiscale Representation Enhanced Temporal Flow Fusion Model for Long-Term Workload ForecastingabstractAccurate workload forecasting is critical for efficient resource management in cloud computing systems, enabling effective scheduling and autoscaling. Despite recent advances with transformer-based forecasting models, challenges remain due to the non-stationary, nonlinear characteristics of workload time series and the long-term dependencies. In particular, inconsistent performance between long-term history and near-term forecasts hinders long-range predictions. This paper proposes a novel framework leveraging self-supervised multiscale representation learning to capture both long-term and near-term workload patterns. The long-term history is encoded through multiscale representations while the near-term observations are modeled via temporal flow fusion. These representations of different scales are fused using an attention mechanism and characterized with normalizing flows to handle non-Gaussian/non-linear distributions of time series. Extensive experiments on 9 benchmarks demonstrate superiority over existing methods. Shiyu Wang 0001, Zhixuan Chu, Yinbo Sun, Yu Liu 0071, Yuliang Guo, Huiyang Jian, Lintao Ma, Xingyu Lu 0004, Jun Zhou 0011 |
CIKM | 5 |
| 2024 | SeaBird: Segmentation in Bird's View with Dice Loss Improves Monocular 3D Detection of Large ObjectsabstractMonocular 3D detectors achieve remarkable performance on cars and smaller objects. However, their performance drops on larger objects, leading to fatal accidents. Some attribute the failures to training data scarcity or the receptive field requirements of large objects. In this paper, we highlight this understudied problem of generalization to large objects. We find that modern frontal detectors struggle to generalize to large objects even on nearly balanced datasets. We argue that the cause of failure is the sensitivity of depth regression losses to noise of larger objects. To bridge this gap, we comprehensively investigate regression and dice losses, examining their robustness under varying error levels and object sizes. We mathematically prove that the dice loss leads to superior noise-robustness and model convergence for large objects compared to regression losses for a simplified case. Leveraging our theoretical insights, we propose SeaBird (Segmentation in Bird's View) as the first step towards generalizing to large objects. SeaBird effectively integrates BEV segmentation on foreground objects for 3D detection, with the segmentation head trained with the dice loss. SeaBird achieves SoTA results on the KITTI-360 leaderboard and improves existing detectors on the nuScenes leaderboard, particularly for large objects. Abhinav Kumar 0004, Yuliang Guo, Xinyu Huang 0001, Liu Ren 0001, Xiaoming Liu 0002 |
CVPR | 2 |
| 2024 | Behind the Veil: Enhanced Indoor 3D Scene Reconstruction with Occluded Surfaces CompletionabstractIn this paper, we present a novel indoor 3D reconstruction method with occluded surface completion, given a sequence of depth readings. Prior state-of-the-art (SOTA) methods only focus on the reconstruction of the visible areas in a scene, neglecting the invisible areas due to the occlusions, e.g., the contact surface between furniture, occluded wall and floor. Our method tackles the task of completing the occluded scene surfaces, resulting in a complete 3D scene mesh. The core idea of our method is learning 3D geometry prior from various complete scenes to infer the occluded geometry of an unseen scene from solely depth measurements. We design a coarse-fine hierarchical octree representation coupled with a dual-decoder architecture, i.e., Geo-decoder and 3D Inpainter, which jointly reconstructs the complete 3D scene geometry. The Geo-decoder with detailed representation at fine levels is optimized online for each scene to reconstruct visible surfaces. The 3D Inpainter with abstract representation at coarse levels is trained offline using various scenes to complete occluded surfaces. As a result, while the Geo-decoder is specialized for an individual scene, the 3D Inpainter can be generally applied across different scenes. We evaluate the proposed method on the 3D Completed Room Scene (3D-CRS) and iTHOR datasets, significantly outperforming the SOTA methods by a gain of 16.8% and 24.2% in terms of the completeness of 3D reconstruction. 3D-CRS dataset including a complete 3D mesh of each scene is provided on project webpage11https://github.com/BoschRHI3NA/3D-CRS-dataset. Su Sun, Cheng Zhao 0002, Yuliang Guo, Ruoyu Wang 0012, Xinyu Huang 0001, Victor Y. Chen, Liu Ren 0001 |
CVPR | 3 |
| 2024 | SUP-NeRF: A Streamlined Unification of Pose Estimation and NeRF for Monocular 3D Object Reconstruction
Yuliang Guo, Abhinav Kumar 0004, Cheng Zhao 0002, Ruoyu Wang 0012, Xinyu Huang 0001, Liu Ren 0001 |
ECCV (69) | 1 |
| 2024 | TCLC-GS: Tightly Coupled LiDAR-Camera Gaussian Splatting for Autonomous Driving: Supplementary Materials
Cheng Zhao 0002, Su Sun, Ruoyu Wang 0012, Yuliang Guo, Jun-Jun Wan, Xinyu Huang 0001, Victor Y. Chen, Liu Ren 0001 |
ECCV (63) | 4 |
| 2024 | Enhancing Online Road Network Perception and Reasoning with Standard Definition MapsabstractAutonomous driving for urban and highway driving applications often requires High Definition (HD) maps to generate a navigation plan. Nevertheless, various challenges arise when generating and maintaining HD maps at scale. While recent online mapping methods have started to emerge, their performance especially for longer ranges is limited by heavy occlusion in dynamic environments. With these considerations in mind, our work focuses on leveraging lightweight and scalable priors–Standard Definition (SD) maps–in the development of online vectorized HD map representations. We first examine the integration of prototypical rasterized SD map representations into various online mapping architectures. Furthermore, to identify lightweight strategies, we extend the OpenLane-V2 dataset with OpenStreetMaps and evaluate the benefits of graphical SD map representations. A key finding from designing SD map integration components is that SD map encoders are model agnostic and can be quickly adapted to new architectures that utilize bird’s eye view (BEV) encoders. Our results show that making use of SD maps as priors for the online mapping task can significantly speed up convergence and boost the performance of the online centerline perception task by 30% (mAP). Furthermore, we show that the introduction of the SD maps leads to a reduction of the number of parameters in the perception and reasoning task by leveraging SD map graphs while improving the overall performance. Project Page: https://henryzhangzhy.github.io/sdhdmap/. Hengyuan Zhang 0001, David Paz, Yuliang Guo, Arun Das 0007, Xinyu Huang 0001, Karsten Haug, Henrik I. Christensen, Liu Ren 0001 |
IROS | 3 |
| 2023 | 3D Copy-Paste: Physically Plausible Object Insertion for Monocular 3D DetectionabstractA major challenge in monocular 3D object detection is the limited diversity and quantity of objects in real datasets. While augmenting real scenes with virtual objects holds promise to improve both the diversity and quantity of the objects, it remains elusive due to the lack of an effective 3D object insertion method in complex real captured scenes. In this work, we study augmenting complex real indoor scenes with virtual objects for monocular 3D object detection. The main challenge is to automatically identify plausible physical properties for virtual assets (e.g., locations, appearances, sizes, etc.) in cluttered real scenes. To address this challenge, we propose a physically plausible indoor 3D object insertion approach to automatically copy virtual objects and paste them into real scenes. The resulting objects in scenes have 3D bounding boxes with plausible physical locations and appearances. In particular, our method first identifies physically feasible locations and poses for the inserted objects to prevent collisions with the existing room layout. Subsequently, it estimates spatially-varying illumination for the insertion location, enabling the immersive blending of the virtual objects into the original scene with plausible appearances and cast shadows. We show that our augmentation method significantly improves existing monocular 3D object models and achieves state-of-the-art performance. For the first time, we demonstrate that a physically plausible 3D object insertion, serving as a generative data augmentation technique, can lead to significant improvements for discriminative downstream tasks such as monocular 3D object detection. Project website: https://gyhandy.github.io/3D-Copy-Paste/. Yunhao Ge, Hong-Xing Yu, Cheng Zhao 0002, Yuliang Guo, Xinyu Huang 0001, Liu Ren 0001, Laurent Itti, Jiajun Wu 0001 |
NeurIPS | 4 |
| 2022 | OmniFusion: 360 Monocular Depth Estimation via Geometry-Aware FusionabstractA well-known challenge in applying deep-learning methods to omnidirectional images is spherical distortion. In dense regression tasks such as depth estimation, where structural details are required, using a vanilla CNN layer on the distorted 360 image results in undesired information loss. In this paper, we propose a 360 monocular depth estimation pipeline, OmniFusion, to tackle the spherical distortion issue. Our pipeline transforms a 360 image into less-distorted perspective patches (i.e. tangent images) to obtain patch-wise predictions via CNN, and then merge the patch-wise results for final output. To handle the discrepancy between patch-wise predictions which is a major issue affecting the merging quality, we propose a new framework with the following key components. First, we propose a geometry-aware feature fusion mechanism that combines 3D geometric features with 2D image features to compensate for the patch-wise discrepancy. Second, we employ the self-attention-based transformer architecture to conduct a global aggregation of patch-wise information, which further improves the consistency. Last, we introduce an iterative depth refinement mechanism, to further refine the estimated depth based on the more accurate geometric features. Experiments show that our method greatly mitigates the distortion issue, and achieves state-of-the-art performances on several 360 monocular depth estimation benchmark datasets. Our code is available at https://github.com/yuyanli0831/OmniFusion. Yuliang Guo, Zhixin Yan, Xinyu Huang 0001, Ye Duan, Liu Ren 0001 |
CVPR | 2 |
| 2022 | Symmetry and Uncertainty-Aware Object SLAM for 6DoF Object Pose EstimationabstractWe propose a keypoint-based object-level SLAM framework that can provide globally consistent 6DoF pose estimates for symmetric and asymmetric objects alike. To the best of our knowledge, our system is among the first to utilize the camera pose information from SLAM to provide prior knowledge for tracking keypoints on symmetric objects - ensuring that new measurements are consistent with the current 3D scene. Moreover, our semantic key-point network is trained to predict the Gaussian covariance for the keypoints that captures the true error of the prediction, and thus is not only useful as a weight for the residuals in the system's optimization problems, but also as a means to detect harmful statistical outliers without choosing a manual threshold. Experiments show that our method provides competitive performance to the state of the art in 6DoF object pose estimation, and at a real-time speed. Our code, pre-trained models, and keypoint labels are available https://github.com/rpng/suo_slam. Nathaniel W. Merrill, Yuliang Guo, Xingxing Zuo 0001, Xinyu Huang 0001, Stefan Leutenegger, Liu Ren 0001, Guoquan Huang 0001 |
CVPR | 2 |
| 2022 | PoP-Net: Pose over Parts Network for Multi-Person 3D Pose Estimation from a Depth ImageabstractIn this paper, a real-time method called PoP-Net is proposed to predict multi-person 3D poses from a depth image. PoP-Net learns to predict bottom-up part representations and top-down global poses in a single shot. Specifically, a new part-level representation, called Truncated Part Displacement Field (TPDF), is introduced which enables an explicit fusion process to unify the advantages of bottom-up part detection and global pose detection. Meanwhile, an effective mode selection scheme is introduced to automatically resolve the conflicting cases between global pose and part detections. Finally, due to the lack of high-quality depth datasets for developing multi-person 3D pose estimation, we introduce Multi-Person 3D Human Pose Dataset (MP-3DHP) as a new benchmark. MP-3DHP is designed to enable effective multi-person and background data augmentation in model training, and to evaluate 3D human pose estimators under uncontrolled multi-person scenarios. We show that PoP-Net achieves the state-of-the-art results both on MP-3DHP and on the widely used ITOP dataset, and has significant advantages in efficiency for multi-person processing. MP-3DHP Dataset and the evaluation code have been made available at: https://github.com/oppo-us-research/PoP-Net. Yuliang Guo, Zhong Li 0007, Zekun Li 0011, Xiangyu Du, Shuxue Quan, Yi Xu 0002 |
WACV | 1 |
| 2020 | Gen-LaneNet: A Generalized and Scalable Approach for 3D Lane Detection
Yuliang Guo, Peitao Zhao, Weide Zhang, Jinghao Miao, Jingao Wang, Tae Eun Choe |
ECCV (21) | 1 |
| 2019 | Differential Geometry in Edge Detection: Accurate Estimation of Position, Orientation and CurvatureabstractThe vast majority of edge detection literature has aimed at improving edge recall and precision, with relatively few addressing the accuracy of edge orientation estimates which are often based on gradient. We show that first-order estimates of orientation can have significant error and this can be remedied by employing Third-Order estimates. This paper aims at estimating differential geometry attributes of an edge, namely, localization, orientation, and curvature, as well as edge topology, and develop robust numerical techniques in gray-scale and color images, applicable to a variety of popular edge detectors, such as gradient-based, gPb and SE. Second, a combinatorial model of edge grouping in a small neighborhood is developed to capture all geometrically consistent grouping called curvels, which establish: (i) edge topology in the form of potential links between an edge and other edges; (ii) an accurate curvature estimate for each possible grouping, whose performance is comparable to methods which use global and multi-scale methods; (iii) a more accurate localization of an edge. These have been evaluated using four distinct methodologies (i) traditional human annotated datasets; (ii) using coherence measure; (iii) stability analysis under visual perturbation, and (iv) utilitarian evaluation, and show meaningful improvements. Benjamin B. Kimia, Yuliang Guo, Amir Tamrakar |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | BoMW: Bag of Manifold Words for One-Shot Learning Gesture Recognition From KinectabstractIn this paper, we study one-shot learning gesture recognition on RGB-D data recorded from Microsoft's Kinect. To this end, we propose a novel bag of manifold words (BoMW)-based feature representation on symmetric positive definite (SPD) manifolds. In particular, we use covariance matrices to extract local features from RGB-D data due to its compact representation ability as well as the convenience of fusing both RGB and depth information. Since covariance matrices are SPD matrices and the space spanned by them is the SPD manifold, traditional learning methods in the Euclidean space, such as sparse coding, cannot be directly applied to them. To overcome this problem, we propose a unified framework to transfer the sparse coding on SPD manifolds to the one on the Euclidean space, which enables any existing learning method to be used. After building BoMW representation on a video from each gesture class, a nearest neighbor classifier is adopted to perform the one-shot learning gesture recognition. Experimental results on the ChaLearn gesture data set demonstrate the outstanding performance of the proposed one-shot learning gesture recognition method compared against the state-of-the-art methods. The effectiveness of the proposed feature extraction method is also validated on a new RGB-D action recognition data set. Lei Zhang 0036, Shengping Zhang, Feng Jiang 0001, Yuankai Qi, Jun Zhang 0017, Yuliang Guo, Huiyu Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2014 | A Multi-stage Approach to Curve Extraction
Yuliang Guo, Naman Kumar 0001, Maruthi Narayanan, Benjamin B. Kimia |
ECCV (1) | 1 |