EDBT 2026 Demo / reviewers in the wild / expert
Yi Xu 0002
dblp:14/5580-2
· DBLP profile ↗
46ranked-venue papers
5as first author
37since 2021 · last 2025
0000-0003-2126-6054ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 5 first-author · 31 since 2021Artificial intelligence and machine learning · 24 · 1 first-author · 20 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dynamic Voxel Grid Optimization for High-Fidelity RGB-D Supervised Surface Reconstruction
Xiangyu Xu 0004, Qingan Yan, Changjiang Cai, Huangying Zhan, Pan Ji, Junsong Yuan 0001, Yi Xu 0002 |
CGI (2) | 8 |
| 2025 | ActiveGAMER: Active GAussian Mapping through Efficient RenderingabstractWe introduce ActiveGAMER, an active mapping system that utilizes 3D Gaussian Splatting (3DGS) to achieve high-quality scene mapping and efficient exploration. Unlike recent NeRF-based methods, which are computationally demanding and limit mapping performance, our approach leverages the efficient rendering capabilities of 3DGS to enable effective and efficient exploration in complex environments. The core of our system is a rendering-based information gain module that identifies the most informative viewpoints for next-best-view planning, enhancing both geometric and photometric reconstruction accuracy. ActiveGAMER also integrates a carefully balanced framework, combining coarse-to-fine exploration, post-refinement, and a global-local keyframe selection strategy to maximize reconstruction completeness and fidelity. Our system autonomously explores and reconstructs environments with state-of-the-art geometric and photometric accuracy and completeness, significantly surpassing existing approaches in both aspects. Extensive evaluations on benchmark datasets such as Replica and MP3D highlight ActiveGAMER’s effectiveness in active mapping tasks. Huangying Zhan, Xiangyu Xu 0004, Qingan Yan, Changjiang Cai, Yi Xu 0002 |
CVPR | 7 |
| 2025 | PlanarNeRF: Online Learning of Planar Primitives with Neural Radiance FieldsabstractIdentifying spatially complete planar primitives from visual data is a crucial task in computer vision. Prior methods are largely restricted to either 2D segment recovery or simplifying 3D structures, even with extensive plane annotations. We present PlanarNeRF, a novel framework capable of detecting dense 3D planes through online learning. Drawing upon the neural field representation, PlanarNeRF brings three major contributions. First, it enhances 3D plane detection with concurrent appearance and geometry knowledge. Second, a lightweight plane fitting module is used to estimate plane parameters. Third, a novel global memory bank structure with an update mechanism is introduced, ensuring consistent cross-frame correspondence. The flexible architecture of PlanarNeRF allows it to function in both 2D-supervised and self-supervised solutions, in each of which it can effectively learn from sparse training signals, significantly improving training efficiency. Through extensive experiments, we demonstrate the effectiveness of PlanarNeRF in various real-world scenarios and remarkable improvement in 3D plane detection over existing works. Zheng Chen 0016, Qingan Yan, Huangying Zhan, Changjiang Cai, Xiangyu Xu 0004, Yuzhong Huang, Ziyue Feng, Yi Xu 0002, Lantao Liu |
ICRA | 9 |
| 2025 | DualMat: PBR Material Estimation via Coherent Dual-Path DiffusionabstractWe present DualMat, a novel dual-path diffusion framework for estimating Physically Based Rendering (PBR) materials from single images under complex lighting conditions. Our approach operates in two distinct latent spaces: an albedo-optimized path leveraging pretrained visual knowledge through RGB latent space, and a material-specialized path operating in a compact latent space designed for precise metallic and roughness estimation. To ensure coherent predictions between the albedo-optimized and material-specialized paths, we introduce feature distillation during training. We employ rectified flow to enhance efficiency by reducing inference steps while maintaining quality. Our framework extends to high-resolution and multi-view inputs through patch-based estimation and cross-view attention, enabling seamless integration into image-to-3D pipelines. DualMat achieves state-of-the-art performance on both Objaverse and real-world data, significantly outperforming existing methods with up to 28% improvement in albedo estimation and 39% reduction in metallic-roughness prediction errors. Our project can be found at yifehuang97.github.io/DualMatProjPage/. Yi Xu 0002, Minh Hoai, Zhong Li 0007 |
ACM Multimedia | 3 |
| 2025 | Understanding while Exploring: Semantics-driven Active MappingabstractEffective robotic autonomy in unknown environments demands proactive exploration and precise understanding of both geometry and semantics. In this paper, we propose ActiveSGM, an active semantic mapping framework designed to predict the informativeness of potential observations before execution. Built upon a 3D Gaussian Splatting (3DGS) mapping backbone, our approach employs semantic and geometric uncertainty quantification, coupled with a sparse semantic representation, to guide exploration. By enabling robots to strategically select the most beneficial viewpoints, ActiveSGM efficiently enhances mapping completeness, accuracy, and robustness to noisy semantic data, ultimately supporting more adaptive scene exploration. Our experiments on the Replica and Matterport3D datasets highlight the effectiveness of ActiveSGM in active semantic mapping tasks. Huangying Zhan, Hairong Yin, Yi Xu 0002, Philippos Mordohai |
NeurIPS | 4 |
| 2025 | Scalable High-Fidelity 3D Hand Shape Reconstruction via Graph-Image Frequency Mapping and Graph Frequency DecompositionabstractDespite the impressive performance obtained by recent single-image hand modeling techniques, they lack the capability to capture sufficient details of the 3D hand mesh. This deficiency greatly limits their applications when high-fidelity hand modeling is required, e.g., personalized hand modeling. To address this problem, we design a frequency split network to generate 3D hand meshes using different frequency bands in a coarse-to-fine manner. To capture high-frequency personalized details, we transform the 3D mesh into the frequency domain, and proposed a novel frequency decomposition loss to supervise each frequency component. By leveraging such a coarse-to-fine scheme, hand details that correspond to the higher frequency domain can be preserved. In addition, the proposed network is scalable, and can stop the inference at any resolution level to accommodate different hardware with varying computational powers. To feed the scalable frequency network with frequency split image features, we proposed an image-graph ring feature mapping strategy. To train our network with per-vertex supervision, we use a bidirectional registration strategy to generate a topology-fixed ground-truth. To quantitatively evaluate the performance of our method in terms of recovering personalized shape details, we introduce a new evaluation metric named Mean-frequency Signal-to-Noise Ratio (MSNR) to measure the mean signal-to-noise ratio of mesh signal on each frequency component. Extensive experiments demonstrate that our approach generates fine-grained details for high-fidelity 3D hand reconstruction, and our evaluation metric is more effective than traditional metrics for measuring mesh details. Tianyu Luan, Yuanhao Zhai 0001, Jingjing Meng, Zhong Li 0007, Yi Xu 0002, Junsong Yuan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Weakly Supervised Segmentation on Outdoor 4D Point Clouds With Progressive 4D GroupingabstractRecently, some weakly supervised 3D point cloud segmentation methods have been proposed to develop effective models with minimum annotation efforts. Our previous work, W4DTS, proposes a challenging task that utilizes only 0.001% points in outdoor point cloud datasets to achieve an effective segmentation model. However, under an extremely limited annotation budget, the quality of pseudo labels generated by W4DTS is unsatisfactory, which limits the segmentation performance in such scenarios. To solve this issue, we propose a progressive 4D grouping approach to group the annotated and unannotated points across space and time, which can generate high-quality pseudo labels with very sparse annotated points. Moreover, to further improve our progressive 4D grouping approach, we design a cross-frame contrastive learning and a local consistency learning to improve the quality of our 4D grouping. Experimental results reveal that with only 0.001% annotations, our solution significantly outperforms the previous best approach on SemanticKITTI. We also evaluate our framework on the SemanticPOSS dataset and ScribbleKITTI dataset, and achieve performances close to our fully supervised backbone models. Hanyu Shi 0002, Fayao Liu, Yi Xu 0002, Guosheng Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Diverse and Stable 2D Diffusion Guided Text to 3D Generation with Noise RecalibrationabstractIn recent years, following the success of text guided image generation, text guided 3D generation has gained increasing attention among researchers. Dreamfusion is a notable approach that enhances generation quality by utilizing 2D text guided diffusion models and introducing SDS loss, a technique for distilling 2D diffusion model information to train 3D models. However, the SDS loss has two major limitations that hinder its effectiveness. Firstly, when given a text prompt, the SDS loss struggles to produce diverse content. Secondly, during training, SDS loss may cause the generated content to overfit and collapse, limiting the model's ability to learn intricate texture details. To overcome these challenges, we propose a novel approach called Noise Recalibration algorithm. By incorporating this technique, we can generate 3D content with significantly greater diversity and stunning details. Our approach offers a promising solution to the limitations of SDS loss. Fayao Liu, Yi Xu 0002, Hanjing Su, Qingyao Wu, Guosheng Lin |
AAAI | 3 |
| 2024 | NARUTO: Neural Active Reconstruction from Uncertain Target ObservationsabstractWe present NARUTO, a neural active reconstruction system that combines a hybrid neural representation with uncertainty learning, enabling high-fidelity surface reconstruction. Our approach leverages a multi-resolution hashgrid as the mapping backbone, chosen for its exceptional convergence speed and capacity to capture high-frequency local features. The centerpiece of our work is the incorporation of an uncertainty learning module that dynamically quantifies reconstruction uncertainty while actively reconstructing the environment. By harnessing learned uncertainty, we propose a novel uncertainty aggregation strategy for goal searching and efficient path planning. Our system autonomously explores by targeting uncertain observations and reconstructs environments with remarkable completeness and fidelity. We also demonstrate the utility of this uncertainty-aware approach by enhancing SOTA neural SLAM systems through an active ray sampling strategy. Extensive evaluations of NARUTO in various environments, using an indoor scene simulator, confirm its superior performance and state-of-the-art status in active reconstruction, as evidenced by its impressive results on benchmark datasets like Replica and MP3D. Project page: oppo-usresearch.github.io/NARUTO-website/ Ziyue Feng, Huangying Zhan, Zheng Chen 0016, Qingan Yan, Xiangyu Xu 0004, Changjiang Cai, Qilun Zhu, Yi Xu 0002 |
CVPR | 9 |
| 2024 | Spectrum AUC Difference (SAUCD): Human-Aligned 3D Shape EvaluationabstractExisting 3D mesh shape evaluation metrics mainly focus on the overall shape but are usually less sensitive to local details. This makes them inconsistent with human evaluation, as human perception cares about both overall and detailed shape. In this paper, we propose an analytic metric named Spectrum Area Under the Curve Difference (SAUCD) that demonstrates better consistency with human evaluation. To compare the difference between two shapes, we first transform the 3D mesh to the spectrum domain using the discrete Laplace-Beltrami operator and Fourier transform. Then, we calculate the Area Under the Curve (AUC) difference between the two spectrums, so that each frequency band that captures either the overall or detailed shape is equitably considered. Taking human sensitivity across frequency bands into account, we further extend our metric by learning suitable weights for each frequency band which better aligns with human perception. To measure the performance of SAUCD, we build a 3D mesh evaluation dataset called Shape Grading, along with manual annotations from more than 800 subjects. By measuring the correlation between our metric and human evaluation, we demonstrate that SAUCD is well aligned with human evaluation, and outperforms previous 3D mesh metrics. Our project page: https://bit.ly/saucd. Tianyu Luan, Zhong Li 0007, Lichang Chen, Yi Xu 0002, Junsong Yuan 0001 |
CVPR | 6 |
| 2024 | PanoFree: Tuning-Free Holistic Multi-view Image Generation with Cross-View Self-guidance
Aoming Liu, Zhong Li 0007, Nannan Li 0004, Yi Xu 0002, Bryan A. Plummer |
ECCV (27) | 5 |
| 2024 | Learn to Optimize Denoising Scores: A Unified and Improved Diffusion Prior for 3D Generation
Chi Zhang 0007, Yi Xu 0002, Xulei Yang, Fayao Liu, Guosheng Lin |
ECCV (44) | 5 |
| 2024 | Stereo-NEC: Enhancing Stereo Visual-Inertial SLAM Initialization with Normal Epipolar ConstraintsabstractWe propose an accurate and robust initialization approach for stereo visual-inertial SLAM systems. Unlike the current state-of-the-art method, which heavily relies on the accuracy of a pure visual SLAM system to estimate inertial variables without updating camera poses, potentially compromising accuracy and robustness, our approach offers a different solution. We realize the crucial impact of precise gyroscope bias estimation on rotation accuracy. This, in turn, affects trajectory accuracy due to the accumulation of translation errors. To address this, we first independently estimate the gyroscope bias and use it to formulate a maximum a posteriori problem for further refinement. After this refinement, we proceed to update the rotation estimation by performing IMU integration with gyroscope bias removed from gyroscope measurements. We then leverage robust and accurate rotation estimates to enhance translation estimation via 3-DoF bundle adjustment. Moreover, we introduce a novel approach for determining the success of the initialization by evaluating the residual of the normal epipolar constraint. Extensive evaluations on the EuRoC dataset illustrate that our method excels in accuracy and robustness. It outperforms ORB-SLAM3, the current leading stereo visual-inertial initialization method, in terms of absolute trajectory error and relative rotation error, while maintaining competitive computational speed. Notably, even with 5 keyframes for initialization, our method consistently surpasses the state-of-the-art approach using 10 keyframes in rotation accuracy. The open source code is available at https://github.com/ApdowJN/Stereo-NEC.git. Chieh Chou, Ganesh Sevagamoorthy, Zheng Chen 0016, Ziyue Feng, Youjie Xia, Feiyang Cai, Yi Xu 0002, Philippos Mordohai |
ICRA | 9 |
| 2024 | Show Your Face: Restoring Complete Facial Images from Partial Observations for VR MeetingabstractVirtual Reality (VR) headsets allow users to interact with the virtual world. However, the device physically blocks visual connections among users, causing huge inconveniences for VR meetings. To address this issue, studies have been conducted to restore human faces from images captured by Headset Mounted Cameras (HMC). Unfortunately, existing approaches heavily rely on high-resolution person-specific 3D models which are prohibitively expensive to apply to large-scale scenarios. Our goal is to design an efficient framework for restoring users’ facial data in VR meetings. Specifically, we first build a new dataset, named Facial Image Composition (FIC) data which approximates the real HMC images from a VR headset. By leveraging the heterogeneity of the HMC images, we decompose the restoration problem into a local geometry transformation and global color/style fusion. Then we propose a 2D light-weight facial image composition network (FIC-Net), where three independent local models are responsible for transforming raw HMC patches and the global model performs a fusion of the transformed HMC patches with a pre-recorded reference image. Finally, we also propose a stage-wise training strategy to optimize the generalization of our FIC-Net. We have validated the effectiveness of our proposed FIC-Net through extensive experiments. Zheng Chen 0016, Junsong Yuan 0001, Yi Xu 0002, Lantao Liu |
WACV | 4 |
| 2023 | RIAV-MVS: Recurrent-Indexing an Asymmetric Volume for Multi-View StereoabstractThis paper presents a learning-based method for multi-view depth estimation from posed images. Our core idea is a “learning-to-optimize” paradigm that iteratively indexes a plane-sweeping cost volume and regresses the depth map via a convolutional Gated Recurrent Unit (GRU). Since the cost volume plays a paramount role in encoding the multi-view geometry, we aim to improve its construction both at pixel- and frame- levels. At the pixel level, we propose to break the symmetry of the Siamese network (which is typically used in MVS to extract image features) by introducing a transformer block to the reference image (but not to the source images). Such an asymmetric volume allows the network to extract global features from the reference image to predict its depth map. Given potential inaccuracies in the poses between reference and source images, we propose to incorporate a residual pose network to correct the relative poses. This essentially rectifies the cost volume at the frame level. We conduct extensive experiments on real-world MVS datasets and show that our method achieves state-of-the-art performance in terms of both within-dataset evaluation and cross-dataset generalization. Changjiang Cai, Pan Ji, Qingan Yan, Yi Xu 0002 |
CVPR | 4 |
| 2023 | High Fidelity 3D Hand Shape Reconstruction via Scalable Graph Frequency DecompositionabstractDespite the impressive performance obtained by recent single-image hand modeling techniques, they lack the capability to capture sufficient details of the 3D hand mesh. This deficiency greatly limits their applications when high-fidelity hand modeling is required, e.g., personalized hand modeling. To address this problem, we design a frequency split network to generate 3D hand mesh using different frequency bands in a coarse-to-fine manner. To capture high-frequency personalized details, we transform the 3D mesh into the frequency domain, and propose a novel frequency decomposition loss to supervise each frequency component. By leveraging such a coarse-to-fine scheme, hand details that correspond to the higher frequency domain can be preserved. In addition, the proposed network is scalable, and can stop the inference at any resolution level to accommodate different hardware with varying computational powers. To quantitatively evaluate the performance of our method in terms of recovering personalized shape details, we introduce a new evaluation metric named Mean Signal-to-Noise Ratio (MSNR) to measure the signal-to-noise ratio of each mesh frequency component. Extensive experiments demonstrate that our approach generates fine-grained details for high-fidelity 3D hand reconstruction, and our evaluation metric is more effective for measuring mesh details compared with traditional metrics. The code is available at https://github.com/tyluann/FreqHand. Tianyu Luan, Yuanhao Zhai 0001, Jingjing Meng, Zhong Li 0007, Yi Xu 0002, Junsong Yuan 0001 |
CVPR | 6 |
| 2023 | 3D-aware Facial Landmark Detection via Multi-view Consistent Training on Synthetic DataabstractAccurate facial landmark detection on wild images plays an essential role in human-computer interaction, entertainment, and medical applications. Existing approaches have limitations in enforcing 3D consistency while detecting 3D/2D facial landmarks due to the lack of multi-view in-the-wild training data. Fortunately, with the recent advances in generative visual models and neural rendering, we have witnessed rapid progress towards high quality 3D image synthesis. In this work, we leverage such approaches to construct a synthetic dataset and propose a novel multi-view consistent learning strategy to improve 3D facial landmark detection accuracy on in-the-wild images. The proposed 3D-aware module can be plugged into any learning-based landmark detection algorithm to enhance its accuracy. We demonstrate the superiority of the proposed plug-in module with extensive comparison against state-of-the-art methods on several real and synthetic datasets. Libing Zeng, Wentao Bao, Zhong Li 0007, Yi Xu 0002, Junsong Yuan 0001, Nima Khademi Kalantari |
CVPR | 5 |
| 2023 | Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory ForecastingabstractHand trajectory forecasting from egocentric views is crucial for enabling a prompt understanding of human intentions when interacting with AR/VR systems. However, existing methods handle this problem in a 2D image space which is inadequate for 3D real-world applications. In this paper, we set up an egocentric 3D hand trajectory forecasting task that aims to predict hand trajectories in a 3D space from early observed RGB videos in a first-person view. To fulfill this goal, we propose an uncertainty-aware state space Transformer (USST) that takes the merits of the attention mechanism and aleatoric uncertainty within the framework of the classical state-space model. The model can be further enhanced by the velocity constraint and visual prompt tuning (VPT) on large vision transformers. Moreover, we develop an annotation workflow to collect 3D hand trajectories with high quality. Experimental results on H2O and EgoPAT3D datasets demonstrate the superiority of USST for both 2D and 3D trajectory forecasting. The code and datasets are publicly released: https://actionlab-cv.github.io/EgoHandTrajPred. Wentao Bao, Libing Zeng, Zhong Li 0007, Yi Xu 0002, Junsong Yuan 0001, Yu Kong 0001 |
ICCV | 5 |
| 2023 | NeuRBF: A Neural Fields Representation with Adaptive Radial Basis FunctionsabstractWe present a novel type of neural fields that uses general radial bases for signal representation. State-of-the-art neural fields typically rely on grid-based representations for storing local neural features and N-dimensional linear kernels for interpolating features at continuous query points. The spatial positions of their neural features are fixed on grid nodes and cannot well adapt to target signals. Our method instead builds upon general radial bases with flexible kernel position and shape, which have higher spatial adaptivity and can more closely fit target signals. To further improve the channel-wise capacity of radial basis functions, we propose to compose them with multi-frequency sinusoid functions. This technique extends a radial basis to multiple Fourier radial bases of different frequency bands without requiring extra parameters, facilitating the representation of details. Moreover, by marrying adaptive radial bases with grid-based ones, our hybrid combination inherits both adaptivity and interpolation smoothness. We carefully designed weighting schemes to let radial bases adapt to different types of signals effectively. Our experiments on 2D image and 3D signed distance field representation demonstrate the higher accuracy and compactness of our method than prior arts. When applied to neural radiance field reconstruction, our method achieves state-of-the-art rendering quality, with small model size and comparable training speed. Zhong Li 0007, Liangchen Song, Jingyi Yu 0001, Junsong Yuan 0001, Yi Xu 0002 |
ICCV | 7 |
| 2023 | Relit-NeuLF: Efficient Relighting and Novel View Synthesis via Neural 4D Light FieldabstractIn this paper, we address the problem of simultaneous relighting and novel view synthesis of a complex scene from multi-view images with a limited number of light sources. We propose an analysis-synthesis approach called Relit-NeuLF. Following the recent neural 4D light field network (NeuLF)[22], Relit-NeuLF first leverages a two-plane light field representation to parameterize each ray in a 4D coordinate system, enabling efficient learning and inference. Then, we recover the spatially-varying bidirectional reflectance distribution function (SVBRDF) of a 3D scene in a self-supervised manner. A DecomposeNet learns to map each ray to its SVBRDF components: albedo, normal, and roughness. Based on the decomposed BRDF components and conditioning light directions, a RenderNet learns to synthesize the color of the ray. To self-supervise the SVBRDF decomposition, we encourage the predicted ray color to be close to the physically-based rendering result using the microfacet model. Comprehensive experiments demonstrate that the proposed method is efficient and effective on both synthetic data and real-world human face data, and outperforms the state-of-the-art results. Zhong Li 0007, Liangchen Song, Xiangyu Du, Junsong Yuan 0001, Yi Xu 0002 |
ACM Multimedia | 7 |
| 2023 | OpenIllumination: A Multi-Illumination Dataset for Inverse Rendering Evaluation on Real ObjectsabstractWe introduce OpenIllumination, a real-world dataset containing over 108K images of 64 objects with diverse materials, captured under 72 camera views and a large number of different illuminations. For each image in the dataset, we provide accurate camera parameters, illumination ground truth, and foreground segmentation masks. Our dataset enables the quantitative evaluation of most inverse rendering and material decomposition methods for real objects. We examine several state-of-the-art inverse rendering methods on our dataset and compare their performances. The dataset and code can be found on the project page: https://oppo-us-research.github.io/OpenIllumination. Isabella Liu, Ziyang Fu, Liwen Wu, Haian Jin, Zhong Li 0007, Chin Ming Ryan Wong, Yi Xu 0002, Ravi Ramamoorthi, Zexiang Xu, Hao Su 0001 |
NeurIPS | 8 |
| 2023 | MyStyle++: A Controllable Personalized Generative PriorabstractIn this paper, we propose an approach to obtain a personalized generative prior with explicit control over a set of attributes. We build upon MyStyle, a recently introduced method, that tunes the weights of a pre-trained StyleGAN face generator on a few images of an individual. This system allows synthesizing, editing, and enhancing images of the target individual with high fidelity to their facial features. However, MyStyle does not demonstrate precise control over the attributes of the generated images. We propose to address this problem through a novel optimization system that organizes the latent space in addition to tuning the generator. Our key contribution is to formulate a loss that arranges the latent codes, corresponding to the input images, along a set of specific directions according to their attributes. We demonstrate that our approach, dubbed MyStyle++, is able to synthesize, edit, and enhance images of an individual with great control over the attributes, while preserving the unique facial characteristics of that individual. Libing Zeng, Yi Xu 0002, Nima Khademi Kalantari |
SIGGRAPH Asia | 3 |
| 2023 | Semantics-Depth-Symbiosis: Deeply Coupled Semi-Supervised Learning of Semantics and DepthabstractMulti-task learning (MTL) paradigm focuses on jointly learning two or more tasks, aiming for an improvement w.r.t model’s generalizability, performance, and training/inference memory footprint. The aforementioned benefits become ever so indispensable in the case of training for vision-related dense prediction tasks. In this work, we tackle the MTL problem of two dense tasks, i.e., semantic segmentation and depth estimation, and present a novel attention module called Cross-Channel Attention Module (CCAM), which facilitates effective feature sharing along each channel between the two tasks, leading to mutual performance gain with a negligible increase in trainable parameters. In a symbiotic spirit, we also formulate novel data augmentations for the semantic segmentation task using predicted depth called AffineMix, and one using predicted semantics called ColorAug, for depth estimation task. Finally, we validate the performance gain of the proposed method on the Cityscapes and ScanNet dataset. which helps us achieve state-of-the-art results for a semi-supervised joint model based on depth estimation and semantic segmentation. Nitin Bansal, Pan Ji, Junsong Yuan 0001, Yi Xu 0002 |
WACV | 4 |
| 2023 | MonoIndoor++: Towards Better Practice of Self-Supervised Monocular Depth Estimation for Indoor EnvironmentsabstractSelf-supervised monocular depth estimation has seen significant progress in recent years, especially in outdoor environments, i.e., autonomous driving scenes. However, depth prediction results are not satisfying in indoor scenes where most of the existing data are captured with hand-held devices. As compared to outdoor environments, estimating depth of monocular videos for indoor environments, using self-supervised methods, results in two additional challenges: (i) the depth range of indoor video sequences varies a lot across different frames, making it difficult for the depth network to induce consistent depth cues for training, whereas the maximum distance in outdoor scenes mostly stays the same as the camera usually sees the sky; (ii) the indoor sequences recorded with handheld devices often contain much more rotational motions, which cause difficulties for the pose network to predict accurate relative camera poses, while the motions of outdoor sequences are pre-dominantly translational, especially for street-scene driving datasets such as KITTI. In this work, we propose a novel framework-MonoIndoor++ by giving special considerations to those challenges and consolidating a set of good practices for improving the performance of self-supervised monocular depth estimation for indoor environments. First, a depth factorization module with transformer-based scale regression network is proposed to estimate a global depth scale factor explicitly, and the predicted scale factor can indicate the maximum depth values. Second, rather than using a single-stage pose estimation strategy as in previous methods, we propose to utilize a residual pose estimation module to estimate relative camera poses across consecutive frames iteratively. Third, to incorporate extensive coordinates guidance for our residual pose estimation module, we propose to perform coordinate convolutional encoding directly over the inputs to pose networks. The proposed method is validated on a variety of benchmark indoor datasets, i.e., EuRoC MAV, NYUv2, ScanNet and 7-Scenes, demonstrating the state-of-the-art performance. In addition, the effectiveness of each module is shown through a carefully conducted ablation study and the good generalization and universality of our trained model is also demonstrated, specifically on ScanNet and 7-Scenes datasets. Runze Li 0003, Pan Ji, Yi Xu 0002, Bir Bhanu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Efficient Few-Shot Object Detection via Knowledge InheritanceabstractFew-shot object detection (FSOD), which aims at learning a generic detector that can adapt to unseen tasks with scarce training samples, has witnessed consistent improvement recently. However, most existing methods ignore the efficiency issues, e.g., high computational complexity and slow adaptation speed. Notably, efficiency has become an increasingly important evaluation metric for few-shot techniques due to an emerging trend toward embedded AI. To this end, we present an efficient pretrain-transfer framework (PTF) baseline with no computational increment, which achieves comparable results with previous state-of-the-art (SOTA) methods. Upon this baseline, we devise an initializer named knowledge inheritance (KI) to reliably initialize the novel weights for the box classifier, which effectively facilitates the knowledge transfer process and boosts the adaptation speed. Within the KI initializer, we propose an adaptive length re-scaling (ALR) strategy to alleviate the vector length inconsistency between the predicted novel weights and the pretrained base weights. Finally, our approach not only achieves the SOTA results across three public benchmarks, i.e., PASCAL VOC, COCO and LVIS, but also exhibits high efficiency with $1.8-100\times $ faster adaptation speed against the other methods on COCO/LVIS benchmark during few-shot transfer. To our best knowledge, this is the first work to consider the efficiency problem in FSOD. We hope to motivate a trend toward powerful yet efficient few-shot technique development. The codes are publicly available at https://github.com/Ze-Yang/Efficient-FSOD. Ze Yang 0002, Chi Zhang 0007, Ruibo Li, Yi Xu 0002, Guosheng Lin |
IEEE Trans. Image Process. | 4 |
| 2023 | Real-Time Lighting Estimation for Augmented Reality via Differentiable Screen-Space RenderingabstractAugmented Reality (AR) applications aim to provide realistic blending between the real-world and virtual objects. One of the important factors for realistic AR is the correct lighting estimation. In this article, we present a method that estimates the real-world lighting condition from a single image in real time, using information from an optional support plane provided by advanced AR frameworks (e.g., ARCore, ARKit, etc.). By analyzing the visual appearance of the real scene, our algorithm can predict the lighting condition from the input RGB photo. In the first stage, we use a deep neural network to decompose the scene into several components: lighting, normal, and Bidirectional Reflectance Distribution Function (BRDF). Then we introduce differentiable screen-space rendering, a novel approach to providing the supervisory signal for regressing lighting, normal, and BRDF jointly. We recover the most plausible real-world lighting condition using Spherical Harmonics and the main directional lighting. Through a variety of experimental results, we demonstrate that our method can provide improved results than prior works quantitatively and qualitatively, and it can enhance the real-time AR experiences. Celong Liu, Zhong Li 0007, Shuxue Quan, Yi Xu 0002 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | NeRFPlayer: A Streamable Dynamic Scene Representation with Decomposed Neural Radiance FieldsabstractVisually exploring in a real-world 4D spatiotemporal space freely in VR has been a long-term quest. The task is especially appealing when only a few or even single RGB cameras are used for capturing the dynamic scene. To this end, we present an efficient framework capable of fast reconstruction, compact modeling, and streamable rendering. First, we propose to decompose the 4D spatiotemporal space according to temporal characteristics. Points in the 4D space are associated with probabilities of belonging to three categories: static, deforming, and new areas. Each area is represented and regularized by a separate neural field. Second, we propose a hybrid representations based feature streaming scheme for efficiently modeling the neural fields. Our approach, coined NeRFPlayer, is evaluated on dynamic scenes captured by single hand-held cameras and multi-camera arrays, achieving comparable or superior rendering performance in terms of quality and speed comparable to recent state-of-the-art methods, achieving reconstruction in 10 seconds per frame and interactive rendering. Project website: https://bit.ly/nerfplayer. Liangchen Song, Anpei Chen, Zhong Li 0007, Junsong Yuan 0001, Yi Xu 0002, Andreas Geiger 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2022 | PlaneMVS: 3D Plane Reconstruction from Multi-View StereoabstractWe present a novel framework named PlaneMVS for 3D plane reconstruction from multiple input views with known camera poses. Most previous learning-based plane reconstruction methods reconstruct 3D planes from single images, which highly rely on single-view regression and suffer from depth scale ambiguity. In contrast, we reconstruct 3D planes with a multi-view-stereo (MVS) pipeline that takes advantage of multi-view geometry. We decouple plane reconstruction into a semantic plane detection branch and a plane MVS branch. The semantic plane detection branch is based on a single-view plane detection framework but with differences. The plane MVS branch adopts a set of slanted plane hypotheses to replace conventional depth hypotheses to perform plane sweeping strategy and finally learns pixel-level plane parameters and its planar depth map. We present how the two branches are learned in a balanced way, and propose a soft-pooling loss to associate the outputs of the two branches and make them benefit from each other. Extensive experiments on various indoor datasets show that PlaneMVS significantly outperforms state-of-the-art (SOTA) single-view plane reconstruction methods on both plane detection and 3D geometry metrics. Our method even outperforms a set of SOTA learning-based MVS methods thanks to the learned plane priors. To the best of our knowledge, this is the first work on 3D plane reconstruction within an end-to-end MVS framework. Pan Ji, Nitin Bansal, Changjiang Cai, Qingan Yan, Sharon X. Huang, Yi Xu 0002 |
CVPR | 7 |
| 2022 | GeoRefine: Self-supervised Online Depth Refinement for Accurate Dense Mapping
Pan Ji, Qingan Yan, Yi Xu 0002 |
ECCV (1) | 4 |
| 2022 | SightX: A 3D Selection Technique for XR
Yi Xu 0002 |
EuroXR | 3 |
| 2022 | Deformable VisTR: Spatio Temporal Deformable Attention for Video Instance SegmentationabstractVideo instance segmentation (VIS) task requires classifying, segmenting, and tracking object instances over all frames in a video clip. Recently, VisTR [1] has been proposed as end-to-end transformer-based VIS framework, while demonstrating state-of-the-art performance. However, VisTR is slow to converge during training, requiring around 1000 GPU hours due to the high computational cost of its transformer attention module. To improve the training efficiency, we propose Deformable VisTR, leveraging spatio-temporal deformable attention module that only attends to a small fixed set of key spatio-temporal sampling points around a reference point. This enables Deformable VisTR to achieve linear computation in the size of spatio-temporal feature maps. Moreover, it can achieve on par performance as the original VisTR with 10× less GPU training hours. We validate the effectiveness of our method on the Youtube-VIS benchmark. Code is available at https://github.com/skrya/DefVIS. Sudhir Yarram, Jialian Wu, Pan Ji, Yi Xu 0002, Junsong Yuan 0001 |
ICASSP | 4 |
| 2022 | NeuLF: Efficient Novel View Synthesis with Neural 4D Light FieldabstractIn this paper, we present an efficient and robust deep learning solution for novel view synthesis of complex scenes. In our approach, a 3D scene is represented as a light field, i.e., a set of rays, each of which has a corresponding color when reaching the image plane. For efficient novel view rendering, we adopt a two-plane parameterization of the light field, where each ray is characterized by a 4D parameter. We then formulate the light field as a function that indexes rays to corresponding color values. We train a deep fully connected network to optimize this implicit function and memorize the 3D scene. Then, the scene-specific model is used to synthesize novel views. Different from previous light field approaches which require dense view sampling to reliably render novel views, our method can render novel views by sampling rays and querying the color for each ray from the network directly, thus enabling high-quality light field rendering with a sparser set of training images. Per-ray depth can be optionally predicted by the network, thus enabling applications such as auto refocus. Our novel view synthesis results are comparable to the state-of-the-arts, and even superior in some challenging scenes with refraction and reflection. We achieve this while maintaining an interactive frame rate and a small memory footprint. Zhong Li 0007, Liangchen Song, Celong Liu, Junsong Yuan 0001, Yi Xu 0002 |
EGSR (ST) | 5 |
| 2022 | PoP-Net: Pose over Parts Network for Multi-Person 3D Pose Estimation from a Depth ImageabstractIn this paper, a real-time method called PoP-Net is proposed to predict multi-person 3D poses from a depth image. PoP-Net learns to predict bottom-up part representations and top-down global poses in a single shot. Specifically, a new part-level representation, called Truncated Part Displacement Field (TPDF), is introduced which enables an explicit fusion process to unify the advantages of bottom-up part detection and global pose detection. Meanwhile, an effective mode selection scheme is introduced to automatically resolve the conflicting cases between global pose and part detections. Finally, due to the lack of high-quality depth datasets for developing multi-person 3D pose estimation, we introduce Multi-Person 3D Human Pose Dataset (MP-3DHP) as a new benchmark. MP-3DHP is designed to enable effective multi-person and background data augmentation in model training, and to evaluate 3D human pose estimators under uncontrolled multi-person scenarios. We show that PoP-Net achieves the state-of-the-art results both on MP-3DHP and on the widely used ITOP dataset, and has significant advantages in efficiency for multi-person processing. MP-3DHP Dataset and the evaluation code have been made available at: https://github.com/oppo-us-research/PoP-Net. Yuliang Guo, Zhong Li 0007, Zekun Li 0011, Xiangyu Du, Shuxue Quan, Yi Xu 0002 |
WACV | 6 |
| 2021 | MonoIndoor: Towards Good Practice of Self-Supervised Monocular Depth Estimation for Indoor EnvironmentsabstractSelf-supervised depth estimation for indoor environments is more challenging than its outdoor counterpart in at least the following two aspects: (i) the depth range of indoor sequences varies a lot across different frames, making it difficult for the depth network to induce consistent depth cues, whereas the maximum distance in outdoor scenes mostly stays the same as the camera usually sees the sky; (ii) the indoor sequences contain much more rotational motions, which cause difficulties for the pose network, while the motions of outdoor sequences are pre-dominantly translational, especially for driving datasets such as KITTI. In this paper, special considerations are given to those challenges and a set of good practices are consolidated for improving the performance of self-supervised monocular depth estimation in indoor environments. The proposed method mainly consists of two novel modules, i.e., a depth factorization module and a residual pose estimation module, each of which is designed to respectively tackle the aforementioned challenges. The effectiveness of each module is shown through a carefully conducted ablation study and the demonstration of the state-of-the-art performance on three indoor datasets, i.e., EuRoC, NYUv2 and 7-Scenes. Pan Ji, Runze Li 0003, Bir Bhanu, Yi Xu 0002 |
ICCV | 4 |
| 2021 | Learning Kinematic Formulas from Multiple View VideosabstractGiven a set of multiple view videos, which records the motion trajectory of an object, we propose to find out the objects' kinematic formulas with neural rendering techniques. For example, if the input multiple view videos record the free fall motion of an object with different initial speed v, the network aims to learn its kinematics: Δ=vt-1over 2 gt2, where Δ, g and t are displacement, gravitational acceleration and time. To achieve this goal, we design a novel framework consisting of a motion network and a differentiable renderer. For the differentiable renderer, we employ Neural Radiance Field (NeRF) since the geometry is implicitly modeled by querying coordinates in the space. The motion network is composed of a series of blending functions and linear weights, enabling us to analytically derive the kinematic formulas after training. The proposed framework is trained end to end and only requires knowledge of cameras' intrinsic and extrinsic parameters. To validate the proposed framework, we design three experiments to demonstrate its effectiveness and extensibility. The first experiment is the video of free fall and the framework can be easily combined with the principle of parsimony, resulting in the correct free fall kinematics. The second experiment is on the large angle pendulum which does not have analytical kinematics. We use the differential equation controlling pendulum dynamics as a physical prior in the framework and demonstrate that the convergence speed becomes much faster. Finally, we study the explosion animation and demonstrate that our framework can well handle such black-box-generated motions. Liangchen Song, Sheng Liu 0017, Celong Liu, Zhong Li 0007, Yuqi Ding, Yi Xu 0002, Junsong Yuan 0001 |
ACM Multimedia | 6 |
| 2021 | NeCH: Neural Clothed Human ModelabstractExisting human models, e.g., SMPL and STAR, represent 3D geometry of a human body in the form of a polygon mesh obtained by deforming a template mesh according to a set of shape and pose parameters. The appearance, however, is not directly modeled by most existing human models. We present a novel 3D human model that faithfully models both the 3D geometry and the appearance of a clothed human body with a continuous volumetric representation, i.e., volume densities and emitted colors of continuous 3D locations in the volume encompassing the human body. In contrast to the mesh-based representation whose resolution is limited by a mesh's fixed number of polygons, our volumetric representation does not limit the resolution of our model. Moreover, our volumetric represen-tation can be rendered via differentiable volume rendering, thus enabling us to train the model only using 2D images (without using ground truth 3D geometries of human bodies) by minimizing a loss function which measures the differences between rendered images and ground truth images. On the contrary, existing human models are trained using ground truth 3D geometries of human bodies. Thanks to the ability of our model to jointly model both the geometries and the appearances of clothed people, our model can benefit applications including human image synthesis, gaming and 3D television and telepresence. Sheng Liu 0017, Liangchen Song, Yi Xu 0002, Junsong Yuan 0001 |
VCIP | 3 |
| 2021 | Animated 3D human avatars from a single image with GAN-based texture inference
Zhong Li 0007, Celong Liu, Fuyao Zhang, Zekun Li 0011, Yuanzhou Ha, Chenliang Xu, Shuxue Quan, Yi Xu 0002 |
Comput. Graph. | 10 |
| 2020 | Talking-Head Generation with Rhythmic Head Motion
Guofeng Cui, Celong Liu, Zhong Li 0007, Ziyi Kou, Yi Xu 0002, Chenliang Xu |
ECCV (9) | 6 |
| 2020 | Object Detection in the Context of Mobile Augmented RealityabstractIn the past few years, numerous Deep Neural Network (DNN) models and frameworks have been developed to tackle the problem of real-time object detection from RGB images. Ordinary object detection approaches process information from the images only, and they are oblivious to the camera pose with regard to the environment and the scale of the environment. On the other hand, mobile Augmented Reality (AR) frameworks can continuously track a camera's pose within the scene and can estimate the correct scale of the environment by using Visual-Inertial Odometry (VIO). In this paper, we propose a novel approach that combines the geometric information from VIO with semantic information from object detectors to improve the performance of object detection on mobile devices. Our approach includes three components: (1) an image orientation correction method, (2) a scale-based filtering approach, and (3) an online semantic map. Each component takes advantage of the different characteristics of the VIO-based AR framework. We implemented the AR-enhanced features using ARCore and the SSD Mobilenet model on Android phones. To validate our approach, we manually labeled objects in image sequences taken from 12 room-scale AR sessions. The results show that our approach can improve on the accuracy of generic object detectors by 12% on our dataset. Xiang Li 0102, Fuyao Zhang, Shuxue Quan, Yi Xu 0002 |
ISMAR | 5 |
| 2010 | High-resolution modeling of moving and deforming objects using sparse geometric and dense photometric measurementsabstractModeling moving and deforming objects requires capturing as much information as possible during a very short time. When using off-the-shelf hardware, this often hinders the resolution and accuracy of the acquired model. Our key observation is that in as little as four frames both sparse surface-positional measurements and dense surface-orientation measurements can be acquired using a combination of structured light and photometric stereo, resulting in high-resolution models of moving and deforming objects. Our system projects alternating geometric and photometric patterns onto the object using a set of three projectors and captures the object using a synchronized camera. Small motion among temporally close frames is compensated by estimating the optical flow of images captured under the uniform illumination of the photometric light. Then spatial-temporal photogeometric reconstructions are performed to obtain dense and accurate point samples with a sampling resolution equal to that of the camera. Temporal coherence is also enforced. We demonstrate our system by successfully modeling several moving and deforming real-world objects. Yi Xu 0002, Daniel G. Aliaga |
CVPR | 1 |
| 2010 | A Self-Calibrating Method for Photogeometric Acquisition of 3D ObjectsabstractWe present a self-calibrating photogeometric method using only off--the-shelf hardware that enables quickly and robustly obtaining multimillion point-sampled and colored models of real-world objects. Some previous efforts use a priori calibrated systems to separately acquire geometric and photometric information. Our key enabling observation is that a digital projector can be simultaneously used as either an active light source or as a virtual camera (as opposed to a digital camera, which cannot be used for both). We present our self--calibrating and multiviewpoint 3D acquisition method, based on structured light, which simultaneously obtains mutually registered surface position and surface normal information and produces a single high-quality model. Acquisition processing freely alternates between using a geometric setup and using a photometric setup with the same hardware configuration. Further, our approach generates reconstructions at the resolution of the camera and not only the projector. We show the results of capturing several high-quality models of real--world objects. Daniel G. Aliaga, Yi Xu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Modeling Repetitive Motions Using Structured LightabstractObtaining models of dynamic 3D objects is an important part of content generation for computer graphics. Numerous methods have been extended from static scenarios to model dynamic scenes. If the states or poses of the dynamic object repeat often during a sequence (but not necessarily periodically), we call such a repetitive motion. There are many objects, such as toys, machines, and humans, undergoing repetitive motions. Our key observation is that when a motion-state repeats, we can sample the scene under the same motion state again but using a different set of parameters; thus, providing more information of each motion state. This enables robustly acquiring dense 3D information difficult for objects with repetitive motions using only simple hardware. After the motion sequence, we group temporally disjoint observations of the same motion state together and produce a smooth space-time reconstruction of the scene. Effectively, the dynamic scene modeling problem is converted to a series of static scene reconstructions, which are easier to tackle. The varying sampling parameters can be, for example, structured-light patterns, illumination directions, and viewpoints resulting in different modeling techniques. Based on this observation, we present an image-based motion-state framework and demonstrate our paradigm using either a synchronized or an unsynchronized structured-light acquisition method. Yi Xu 0002, Daniel G. Aliaga |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2009 | An Adaptive Correspondence Algorithm for Modeling Scenes with Strong InterreflectionsabstractModeling real-world scenes, beyond diffuse objects, plays an important role in computer graphics, virtual reality, and other commercial applications. One active approach is projecting binary patterns in order to obtain correspondence and reconstruct a densely sampled 3D model. In such structured-light systems, determining whether a pixel is directly illuminated by the projector is essential to decoding the patterns. When a scene has abundant indirect light, this process is especially difficult. In this paper, we present a robust pixel classification algorithm for this purpose. Our method correctly establishes the lower and upper bounds of the possible intensity values of an illuminated pixel and of a non-illuminated pixel. Based on the two intervals, our method classifies a pixel by determining whether its intensity is within one interval but not in the other. Our method performs better than standard method due to the fact that it avoids gross errors during decoding process caused by strong inter-reflections. For the remaining uncertain pixels, we apply an iterative algorithm to reduce the inter-reflection within the scene. Thus, more points can be decoded and reconstructed after each iteration. Moreover, the iterative algorithm is carried out in an adaptive fashion for fast convergence. Yi Xu 0002, Daniel G. Aliaga |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2008 | Photogeometric structured light: A self-calibrating and multi-viewpoint framework for accurate 3D modelingabstractStructured-light methods actively generate geometric correspondence data between projectors and cameras in order to facilitate robust 3D reconstruction. In this paper, we present photogeometric structured light whereby a standard structured light method is extended to include photometric methods. Photometric processing serves the double purpose of increasing the amount of recovered surface detail and of enabling the structured-light setup to be robustly self-calibrated. Further, our framework uses a photogeometric optimization that supports the simultaneous use of multiple cameras and projectors and yields a single and accurate multi-view 3D model which best complies with photometric and geometric data. Daniel G. Aliaga, Yi Xu 0002 |
CVPR | 2 |
| 2007 | Robust pixel classification for 3D modeling with structured lightabstractModeling 3D objects and scenes is an important part of computer graphics. One approach to modeling is projecting binary patterns onto the scene in order to obtain correspondences and reconstruct a densely sampled 3D model. In such structured light systems, determining whether a pixel is directly illuminated by the projector is essential to decoding the patterns. In this paper, we introduce a robust, efficient, and easy to implement pixel classification algorithm for this purpose. Our method correctly establishes the lower and upper bounds of the possible intensity values of an illuminated pixel and of a non-illuminated pixel. Based on the two intervals, our method classifies a pixel by determining whether its intensity is within one interval and not in the other. Experiments show that our method improves both the quantity of decoded pixels and the quality of the final reconstruction producing a dense set of 3D points, inclusively for complex scenes with indirect lighting effects. Furthermore, our method does not require newly designed patterns; therefore, it can be easily applied to previously captured data. Yi Xu 0002, Daniel G. Aliaga |
Graphics Interface | 1 |
| 2007 | Efficient multi-viewpoint acquisition of 3D objects undergoing repetitive motionsabstractComputer graphics applications such as movie effects, video gaming, and product demonstration demand 3D models of dynamic objects. For this purpose, numerous methods, such as light fields, stereo reconstruction and visual hulls have been extended to model dynamic objects. These methods use multiple cameras to acquire images simultaneously and use the synchronized samples to reconstruct the model for each time instance. However, a large number of cameras are required to obtain compelling results. We introduce an efficient acquisition and modeling schema for dynamic objects with repetitive motions. Our method requires as few as two cameras. The key idea is that repetitive motions can be described by a finite number of states. Images capturing the same state can be grouped together and fed to the later modeling phase as if they are captured from multiple cameras simultaneously. Our work includes an acquisition system with interactive feedback, a graph traversal algorithm to help obtain a near minimum subset of images to sample the object and its motion, and a space-time image optimization method. We demonstrate this system using several datasets with different complexity of motion, and different number of desired viewpoints. Yi Xu 0002, Daniel G. Aliaga |
SI3D | 1 |