VLDB 2026 Research / reviewers in the wild / expert
Xinghui Li
dblp:248/2144
· DBLP profile ↗
27ranked-venue papers
6as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 18 since 2021Artificial intelligence and machine learning · 17 · 4 first-author · 16 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep learning-enabled phase error correction in structured light 3D measurement systems for overexposure, non-linearity, and textured scenes
Haozhen Huang, Zongyang Zhang, Xiaojun Liang, Xinghui Li, Limei Song |
Expert Syst. Appl. | 5 |
| 2026 | Semantic Correspondence: Unified Benchmarking and a Strong BaselineabstractEstablishing semantic correspondence is a challenging task in computer vision, aiming to match keypoints with the same semantic information across different images. Benefiting from the rapid development of deep learning, remarkable progress has been made over the past decade. However, a comprehensive review and analysis of this task remains absent. In this paper, we present the first extensive survey of semantic correspondence methods. We first propose a taxonomy to classify existing methods based on the type of their method designs. These methods are then categorized accordingly, and we provide a detailed analysis of each approach. Furthermore, we aggregate and summarize the results of methods in the literature across various benchmarks into a unified comparative table, with detailed configurations to highlight performance variations. Additionally, to provide a detailed understanding of existing methods for semantic matching, we thoroughly conduct controlled experiments to analyze the effectiveness of the components of different methods. Finally, we propose a simple yet effective baseline that achieves state-of-the-art performance on multiple benchmarks, providing a solid foundation for future research in this field. We hope this survey serves as a comprehensive reference and consolidated baseline for future development. Xinghui Li, Kai Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | FESS-3D: Target foreground enhancement single-shot fringe projection 3D reconstruction
Zinan Li, Xiaojun Liang, Weikang Chen, Cong Liu 0036, Xiaohao Wang, Weihua Gui 0001, Wen Gao 0001, Xinghui Li |
Pattern Recognit. | 10 |
| 2025 | AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion ModelsabstractRecent advances in garment-centric image generation from text and image prompts based on diffusion models are impressive. However, existing methods lack support for various combinations of attire, and struggle to preserve the garment details while maintaining faithfulness to the text prompts, limiting their performance across diverse scenarios. In this paper, we focus on a new task, i.e., Multi-Garment Virtual Dressing, and we propose a novel Any-Dressing method for customizing characters conditioned on any combination of garments and any personalized text prompts. AnyDressing comprises two primary networks named GarmentsNet and DressingNet, which are respectively dedicated to extracting detailed clothing features and generating customized images. Specifically, we propose an efficient and scalable module called Garment-Specific Feature Extractor in GarmentsNet to individually encode garment textures in parallel. This design prevents garment confusion while ensuring network efficiency. Meanwhile, we design an adaptive Dressing-Attention mechanism and a novel Instance-Level Garment Localization Learning strategy in DressingNet to accurately inject multi-garment features into their corresponding regions. This approach efficiently integrates multi-garment texture cues into generated images and further enhances text-image consistency. Additionally, we introduce a Garment-Enhanced Texture Learning strategy to improve the fine-grained texture details of garments. Thanks to our well-craft design, Any-Dressing can serve as a plug-in module to easily integrate with any community control extensions for diffusion models, improving the diversity and controllability of synthesized images. Extensive experiments show that AnyDressing achieves state-of-the-art results. Xinghui Li, Qichao Sun, Pengze Zhang, Fulong Ye, Zhichao Liao, Wanquan Feng, Songtao Zhao |
CVPR | 1 |
| 2025 | Training-Free Point Cloud Recognition Based on Geometric and Semantic Information FusionabstractThe trend of employing training-free methods for point cloud recognition is becoming increasingly popular due to its significant reduction in computational resources and time costs. However, existing approaches are limited as they typically extract either geometric or semantic features. To address this limitation, we are the first to propose a novel training-free method that integrates both geometric and semantic features. For the geometric branch, we adopt a non-parametric strategy to extract geometric features. In the semantic branch, we leverage a model aligned with text features to obtain semantic features. Additionally, we introduce the GFE module to complement the geometric information of point clouds and the MFF module to improve performance in few-shot settings. Experimental results demonstrate that our method outperforms existing state-of-the-art training-free approaches on mainstream benchmark datasets, including ModelNet and ScanObiectNN. Zhichao Liao, Xinghui Li, Long Zeng 0001 |
ICASSP | 5 |
| 2025 | GaussianRoom: Improving 3D Gaussian Splatting with SDF Guidance and Monocular Cues for Indoor Scene ReconstructionabstractEmbodied intelligence requires precise reconstruction and rendering to simulate large-scale real-world data. Although 3D Gaussian Splatting (3DGS) has recently demonstrated high-quality results with real-time performance, it still faces challenges in indoor scenes with large, textureless regions, resulting in incomplete and noisy reconstructions due to poor point cloud initialization and underconstrained optimization. Inspired by the continuity of signed distance field (SDF), which naturally has advantages in modeling surfaces, we propose a unified optimization framework that integrates neural signed distance fields (SDFs) with 3DGS for accurate geometry reconstruction and real-time rendering. This framework incorporates a neural SDF field to guide the densification and pruning of Gaussians, enabling Gaussians to model scenes accurately even with poor initialized point clouds. Simultaneously, the geometry represented by Gaussians improves the efficiency of the SDF field by piloting its point sampling. Additionally, we introduce two regularization terms based on normal and edge priors to resolve geometric ambiguities in textureless areas and enhance detail accuracy. Extensive experiments in ScanNet and ScanNet++ show that our method achieves state-of-the-art performance in both surface reconstruction and novel view synthesis. Project page: https://xhd0612.github.io/GaussianRoom.github.io/ Haodong Xiang, Xinghui Li, Xiansong Lai, Wanting Zhang, Zhichao Liao, Long Zeng 0001, Xueping Liu 0003 |
ICRA | 2 |
| 2025 | Seeing in the Dark: Benchmarking Egocentric 3D Vision with the Oxford Day-and-Night DatasetabstractWe introduce Oxford Day-and-Night, a large-scale, egocentric dataset for novel view synthesis (NVS) and visual relocalisation under challenging lighting conditions. Existing datasets often lack crucial combinations of features such as ground-truth 3D geometry, wide-ranging lighting variation, and full 6DoF motion. Oxford Day-and-Night addresses these gaps by leveraging Meta ARIA glasses to capture egocentric video and applying multi-session SLAM to estimate camera poses, reconstruct 3D point clouds, and align sequences captured under varying lighting conditions, including both day and night. The dataset spans over 30 km of recorded trajectories and covers an area of $40{,}000\mathrm{m}^2$, offering a rich foundation for egocentric 3D vision research. It supports two core benchmarks, NVS and relocalisation, providing a unique platform for evaluating models in realistic and diverse environments. Project page: https://oxdan.active.vision/ Wenjing Bian, Xinghui Li, Yifu Tao, Jianeng Wang, Maurice Fallon, Victor Adrian Prisacariu |
NeurIPS | 3 |
| 2025 | DreamO: A Unified Framework for Image CustomizationabstractRecently, extensive research on image customization (e.g., identity, subject, style, background, etc.) demonstrates strong customization capabilities in large-scale generative models. However, most approaches are designed for specific tasks, restricting their generalizability to combine different types of condition. Developing a unified framework for image customization remains an open challenge. In this paper, we present DreamO, an image customization framework designed to support a wide range of tasks while facilitating seamless integration of multiple conditions. Specifically, DreamO utilizes a diffusion transformer (DiT) framework to uniformly process input of different types. During training, we introduce a feature routing constraint to facilitate the precise querying of relevant information from reference images. Additionally, we design a placeholder strategy that associates specific placeholders with conditions at particular positions, enabling control over the placement of conditions in the generated results. Moreover, we employ a progressive training strategy to ensure smooth model convergence and correct the generation quality of the final output. Extensive experiments demonstrate that the proposed DreamO can effectively perform various image customization tasks with high quality and flexibly integrate different types of control conditions. Project page: https://mc-e.github.io/project/DreamO Chong Mou, Yanze Wu, Wenxu Wu, Pengze Zhang, Yufeng Cheng, Xinghui Li, Mengtian Li 0003, Mingcong Liu, Yunsheng Jiang, Shaojin Wu, Songtao Zhao, Jian Zhang 0018 |
SIGGRAPH Asia | 10 |
| 2025 | DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group LearningabstractIn this paper, we introduce DreamID, a diffusion-based face swapping model that achieves high levels of ID similarity, attribute preservation, image fidelity, and fast inference speed. Unlike the typical face swapping training process, which often relies on implicit supervision and struggles to achieve satisfactory results. DreamID establishes explicit supervision for face swapping by constructing Triplet ID Group data, significantly enhancing identity similarity and attribute preservation. The iterative nature of diffusion models poses challenges for utilizing efficient image-space loss functions, as performing time-consuming multi-step sampling to obtain the generated image during training is impractical. To address this issue, we leverage the accelerated diffusion model SD Turbo, reducing the inference steps to a single iteration, enabling efficient pixel-level end-to-end training with explicit Triplet ID Group supervision. Additionally, we propose an improved diffusion-based model architecture comprising SwapNet, FaceNet, and ID Adapter. This robust architecture fully unlocks the power of the Triplet ID Group explicit supervision. Finally, to further extend our method, we explicitly modify the Triplet ID Group data during training to fine-tune and preserve specific attributes, such as glasses and face shape. Extensive experiments demonstrate that DreamID outperforms state-of-the-art methods in terms of identity similarity, pose and expression preservation, and image fidelity. Overall, DreamID achieves high-quality face swapping results at 512×512 resolution in just 0.6 seconds and performs exceptionally well in challenging scenarios such as complex lighting, large angles, and occlusions. Our project: https://superhero-7.github.io/DreamID/. Fulong Ye, Miao Hua, Pengze Zhang, Xinghui Li, Qichao Sun, Songtao Zhao |
SIGGRAPH Asia | 4 |
| 2025 | SL3D-BF: A Real-World Structured Light 3D Dataset With Background-to-Foreground EnhancementabstractDeep-learning-based structured light 3D reconstruction technology (SL3D) provides excellent solutions for intelligent manufacturing. However, the scarcity of real-world datasets covering full-process data and diverse objects hampers the validation of new ideas. Moreover, limited research on dataset construction strategies, such as scene backgrounds and sample distribution, reduces network performance. We investigate the impact of background stability on foreground accuracy (BS-FA) and find a whiteboard background improved foreground prediction accuracy by up to 82% over a black background. Guided by BS-FA, we develop the SL3D-BF, a background-effective SL3D dataset for industrial use, featuring approximately 2,100 scenes with diverse objects like metal/plastic workpieces, plaster sculptures, and standard parts for precise evaluation. It uniquely includes shadow and foreground masks, absent in prior datasets, and offers full-process data from gratings to 3D point clouds, totaling 100,800 gratings. We also establish an initial benchmark for future research by conducting evaluation experiments with advanced methods. Furthermore, we investigate the relationship between the spatial frequency of sample occurrence and the model predictive ability to minimize the time and resource demands of dataset construction. Most importantly, SL3D-BF is also a valuable resource for tasks like depth estimation, defect detection, and semantic segmentation. The dataset is available at: https://github.com/LiYiMingM/Dataset_SL3D_BF. Weikang Chen, Zinan Li, Xiaohao Wang, Weihua Gui 0001, Wen Gao 0001, Xiaojun Liang, Xinghui Li |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2025 | Attention Mono-Depth: Attention-Enhanced Transformer for Monocular Depth Estimation of Volatile Kiln Burden SurfaceabstractAccurate estimation of burden surface depth plays a crucial role in constructing the temperature field and optimizing reaction control in volatile kilns. However, most image-based depth estimation techniques require high-quality input images and achieve limited accuracy, which restrict their applications in actual harsh working conditions such as high temperature, heavy dust and dense smoke. In this study, a deep learning-based monocular depth estimation model is proposed to measure the burden surface depth in the volatile kiln head zone. The proposed model integrates an encoder-decoder network with an attention module. The encoder-decoder network outputs a set of deep semantic features, while the attention module intelligently fuses multi-level features to predict a probability distribution over depth intervals for each pixel. A volatile kiln prototype is designed and constructed to generate image datasets of the kiln head zone which approximate real data collected from industrial production sites. Results demonstrate that the proposed model has a depth prediction error of RMSE = 11.008 mm for the burden surface region, outperforming state-of-the-art neural networks and the traditional depth-from-defocus method. Code and datasets are available athttps://github.com/LLLcong/Attention-MonoDepth. Cong Liu 0036, Xiaojun Liang, Zhiming Han, Chunhua Yang 0001, Weihua Gui 0001, Wen Gao 0001, Xiaohao Wang, Xinghui Li |
IEEE Trans. Circuits Syst. Video Technol. | 10 |
| 2025 | HDRSL Net for Accurate High Dynamic Range Imaging-Based Structured Light 3D ReconstructionabstractIn fringe projection profilometry systems, accurately reconstructing 3D objects with varying surface reflectivity requires high dynamic range (HDR) imaging. However, the limited dynamic range of single-exposure cameras poses challenges for capturing HDR fringe patterns efficiently. This paper introduces a deep learning-based HDR structured light 3D reconstruction pipeline, comprising an HDR Fringe Generation Module and a Phase Calculation Module. The HDR Fringe Generation Module employs an end-to-end network with attention guidance and feature distillation to reconstruct HDR fringe images from short- and long-exposure low dynamic range (LDR) inputs. The Phase Calculation Module processes the phase information from HDR fringes to enable 3D reconstruction. On a metallic HDR dataset, the method achieved a phase error of 0.105, comparable to the 4-exposure 6-step Phase Shifting Profilometry (PSP) method (0.069), with only 8.3% of the projection time. Experimental results demonstrate the robustness of our approach under diverse object geometries, exposure levels, and challenging global illumination environments. In quantitative measurements, our method achieved accuracies of sub-50 $\mu $ m on ceramic spheres, flat plates and metal step object. Ablation experiments confirmed that feature distillation and attention module effectively enhance the HDR Fringe Generation Module, producing high-quality HDR fringe patterns critical for reconstructing objects with HDR surface reflectivity. Furthermore, we constructed an HDR imaging metal dataset comprising 1,700 samples of machined metal parts with diverse shapes, sizes, and materials, making it a benchmark in the field of HDR structured light measurement. Our method offers a general HDR imaging-based structured light 3D reconstruction approach, integrating the two modules into an efficient, end-to-end solution for objects with HDR reflective surfaces. Hao Wang 0234, Xiang Qian, Xiaohao Wang, Weihua Gui 0001, Wen Gao 0001, Xiaojun Liang, Xinghui Li |
IEEE Trans. Image Process. | 8 |
| 2024 | Neural Refinement for Absolute Pose Regression with Feature SynthesisabstractAbsolute Pose Regression (APR) methods use deep neural networks to directly regress camera poses from RGB images. However, the predominant APR architectures only rely on 2D operations during inference, resulting in limited accuracy of pose estimation due to the lack of 3D geometry constraints or priors. In this work, we propose a test-time refinement pipeline that leverages implicit geometric constraints using a robust feature field to enhance the ability of APR methods to use 3D information during inference. We also introduce a novel Neural Feature Synthesizer (NeFeS) model, which encodes 3D geometric features during training and directly renders dense novel view features at test time to refine APR methods. To enhance the robustness of our model, we introduce a feature fusion module and a progressive training strategy. Our proposed method achieves state-of-the-art single-image APR accuracy on indoor and outdoor datasets. Code will be released at https://github.com/ActiveVisionLab/NeFeS. Yash Bhalgat, Xinghui Li, Jiawang Bian, Kejie Li, Victor Adrian Prisacariu |
CVPR | 3 |
| 2024 | GenesisTex: Adapting Image Denoising Diffusion to Texture SpaceabstractWe present GenesisTex, a novel method for synthesizing textures for 3D geometries from text descriptions. GenesisTex adapts the pretrained image diffusion model to texture space by texture space sampling. Specifically, we maintain a latent texture map for each viewpoint, which is updated with predicted noise on the rendering of the corresponding viewpoint. The sampled latent texture maps are then decoded into a final texture map. During the sampling process, we focus on both global and local consistency across multiple viewpoints: global consistency is achieved through the integration of style consistency mechanisms within the noise prediction network, and low-level consistency is achieved by dynamically aligning latent textures. Finally, we apply reference-based inpainting and img2img on denser views for texture refinement. Our approach overcomes the limitations of slow optimization in distillation-based methods and instability in inpainting-based methods. Experiments on meshes from various sources demonstrate that our method surpasses the baseline methods quantitatively and qualitatively. Chenjian Gao, Boyan Jiang, Xinghui Li, Yingpeng Zhang |
CVPR | 3 |
| 2024 | SD4Match: Learning to Prompt Stable Diffusion Model for Semantic MatchingabstractIn this paper, we address the challenge of matching semantically similar keypoints across image pairs. Existing research indicates that the intermediate output of the UNet within the Stable Diffusion (SD) can serve as robust image feature maps for such a matching task. We demonstrate that by employing a basic prompt tuning technique, the inherent potential of Stable Diffusion can be harnessed, resulting in a significant enhancement in accuracy over previous approaches. We further introduce a novel conditional prompting module that conditions the prompt on the local details of the input image pairs, leading to a further improvement in performance. We designate our approach as SD4Match, short for Stable Diffusion for Semantic Matching. Comprehensive evaluations of SD4Match on the PF-Pascal, PF-Willow, and SPair-71k datasets show that it sets new benchmarks in accuracy across all these datasets. Particularly, SD4Match outperforms the previous state-of-the-art by a margin of 12 percentage points on the challenging SPair-71k dataset. Code is available at the project website: https://sd4match.active.vision/. Xinghui Li, Kai Han 0001, Victor Adrian Prisacariu |
CVPR | 1 |
| 2024 | RegionDrag: Fast Region-Based Image Editing with Diffusion Models
Xinghui Li, Kai Han 0001 |
ECCV (17) | 2 |
| 2024 | GaussCtrl: Multi-view Consistent Text-Driven 3D Gaussian Splatting Editing
Jiawang Bian, Xinghui Li, Guangrun Wang, Ian D. Reid 0001, Philip Torr 0001, Victor Adrian Prisacariu |
ECCV (14) | 3 |
| 2024 | Fine-Detailed Neural Indoor Scene Reconstruction Using Multi-Level Importance Sampling And Multi-View ConsistencyabstractRecently, neural implicit 3D reconstruction in indoor scenarios has become popular due to its simplicity and impressive performance. Previous works could produce complete results leveraging monocular priors of normal or depth. However, they may suffer from over-smoothed reconstructions and long-time optimization due to unbiased sampling and inaccurate monocular priors. In this paper, we propose a novel neural implicit surface reconstruction method, named FD-NeuS, to learn fine-detailed 3D models using multi-level importance sampling strategy and multi-view consistency methodology. Specifically, we leverage segmentation priors to guide region-based ray sampling, and use piecewise exponential functions as weights to pilot 3 D points sampling along the rays, ensuring more attention on important regions. In addition, we introduce multi-view feature consistency and multi-view normal consistency as supervision and uncertainty respectively, which further improve the reconstruction of details. Extensive quantitative and qualitative results show that FD-NeuS outperforms existing methods in various scenes. Xinghui Li, Yuchen Ji, Xiansong Lai, Wanting Zhang, Long Zeng 0001 |
ICIP | 1 |
| 2024 | UW-SDF: Exploiting Hybrid Geometric Priors for Neural SDF Reconstruction from Underwater Multi-view Monocular ImagesabstractDue to the unique characteristics of underwater environments, accurate 3D reconstruction of underwater objects poses a challenging problem in tasks such as underwater exploration and mapping. Traditional methods that rely on multiple sensor data for 3D reconstruction are time-consuming and face challenges in data acquisition in underwater scenarios. We propose UW-SDF, a framework for reconstructing target objects from multi-view underwater images based on neural SDF. We introduce hybrid geometric priors to optimize the reconstruction process, markedly enhancing the quality and efficiency of neural SDF reconstruction. Additionally, to address the challenge of segmentation consistency in multi-view images, we propose a novel few-shot multi-view target segmentation strategy using the general-purpose segmentation model (SAM), enabling rapid automatic segmentation of unseen objects. Through extensive qualitative and quantitative experiments on diverse datasets, we demonstrate that our proposed method outperforms the traditional underwater 3D reconstruction method and other neural rendering approaches in the field of underwater 3D reconstruction. Jingyi Tang, Gu Wang 0001, Shengquan Li 0001, Xinghui Li, Xiangyang Ji, Xiu Li 0001 |
IROS | 5 |
| 2024 | Freehand Sketch Generation from Mechanical ComponentsabstractDrawing freehand sketches of mechanical components on multimedia devices for AI-based engineering modeling has become a new trend. However, its development is being impeded because existing works cannot produce suitable sketches for data-driven research. These works either generate sketches lacking a freehand style or utilize generative models not originally designed for this task resulting in poor effectiveness. To address this issue, we design a two-stage generative framework mimicking the human sketching behavior pattern, called MSFormer, which is the first time to produce humanoid freehand sketches tailored for mechanical components. The first stage employs Open CASCADE technology to obtain multi-view contour sketches from mechanical components, filtering perturbing signals for the ensuing generation process. Meanwhile, we design a view selector to simulate viewpoint selection tasks during human sketching for picking out information-rich sketches. The second stage translates contour sketches into freehand sketches by a transformer-based generator. To retain essential modeling features as much as possible and rationalize stroke distribution, we introduce a novel edge-constraint stroke initialization. Furthermore, we utilize a CLIP vision encoder and a new loss function incorporating the Hausdorff distance to enhance the generalizability and robustness of the model. Extensive experiments demonstrate that our approach achieves state-of-the-art performance for generating freehand sketches in the mechanical domain. Project page: https://mcfreeskegen.github.io/. Zhichao Liao, Fengyuan Piao, Xinghui Li, Yue Ma 0033, Pingfa Feng, Heming Fang, Long Zeng 0001 |
ACM Multimedia | 4 |
| 2024 | Towards the generalization of time series classification: A feature-level style transfer and multi-source transfer learning perspective
Baihan Chen, Qiaolin Li, Rui Ma 0037, Xiang Qian, Xiaohao Wang, Xinghui Li |
Knowl. Based Syst. | 6 |
| 2024 | DualRC: A Dual-Resolution Learning Framework With Neighbourhood Consensus for Visual CorrespondencesabstractWe address the problem of establishing accurate correspondences between two images. We present a flexible framework that can easily adapt to both geometric and semantic matching. Our contribution consists of three parts. Firstly, we propose an end-to-end trainable framework that uses the coarse-to-fine matching strategy to accurately find the correspondences. We generate feature maps in two levels of resolution, enforce the neighbourhood consensus constraint on the coarse feature maps by 4D convolutions and use the resulting correlation map to regulate the matches from the fine feature maps. Secondly, we present three variants of the model with different focuses. Namely, a universal correspondence model named DualRC that is suitable for both geometric and semantic matching, an efficient model named DualRC-L tailored for geometric matching with a lightweight neighbourhood consensus module that significantly accelerates the pipeline for high-resolution input images, and the DualRC-D model in which we propose a novel dynamically adaptive neighbourhood consensus module (DyANC) that dynamically selects the most suitable non-isotropic 4D convolutional kernels with the proper neighbourhood size to account for the scale variation. Last, we thoroughly experiment on public benchmarks for both geometric and semantic matching, showing superior performance in both cases. Xinghui Li, Kai Han 0001, Shuda Li, Victor Adrian Prisacariu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Volumetric 3D Reconstruction with Window-Wise Global Feature AggregationabstractVolumetric 3D reconstruction methods have shown great performance in reconstructing indoor scenarios from monocular videos. However, as such approaches utilize discrete feature voxels to encode the observed scenes, the global feature interaction within and across different voxels is ignored, leading to imperfect reconstructions. To solve this problem, we propose a novel volumetric 3D reconstruction method named VolGARecon. The core portion of VolGARecon includes two parts: first, we use an MLP-based weighted fusion module (WFM) to unproject the extracted features to each voxel, which considers the visibility and is capable to reduce the noise caused by occlusion; second, a 3D transformer module (3DTR) is used to perform window-wise global feature interaction in a local sliding window, which strengthens the feature expression in 3D space and benefits estimating more complete and spatially coherent 3D models. In addition, we propose a multi-dimensional hybrid loss (MHL) that incorporates the 3D supervision in classical volumetric methods and the 2D supervision in novel view synthesis works. Extensive experiments show our method achieves superior performance on multiple datasets. Shihao Ren, Yikang Ding, Jinli Liao, Xinghui Li, Wensen Feng, Xueqian Wang 0001 |
ICASSP | 4 |
| 2023 | Edge-aware Neural Implicit Surface ReconstructionabstractRecently, neural implicit 3D reconstruction in indoor scenarios has achieved impressive performance. Utilizing the volume rendering method and neural implicit representation to learn 3D scenes, such per-scene optimization methods could reconstruct pretty complete models but also suffer from missing details and overly-smoothed reconstructions. In this paper, we propose a novel edge-aware neural implicit surface reconstruction method, named Ea-NeuS, to learn high-quality 3D models with fine details. Specifically, we use the edge of objects to locate the important areas, and propose a simple yet effective edge-guided ray-sampling strategy to learn the 3D models. The aforementioned edge information further guides the normal prior supervision, which helps reduce inaccurate optimization in detailed regions. We additionally use the visibility-aware sparse points to pilot the 3D points sampling along the rays and perform explicit supervision. As a result, our method achieves superior performance compared with existing methods on various scenes. Xinghui Li, Yikang Ding, Xiansong Lai, Shihao Ren, Wensen Feng, Long Zeng 0001 |
ICME | 1 |
| 2022 | Disentangling 3D Attributes from a Single 2D Image: Human Pose, Shape and Garment
Xinghui Li, Benjamin Busam, Yiren Zhou, Ales Leonardis, Shanxin Yuan |
BMVC | 2 |
| 2022 | DFNet: Enhance Absolute Pose Regression with Direct Feature Matching
Xinghui Li, Victor Adrian Prisacariu |
ECCV (10) | 2 |
| 2020 | Dual-Resolution Correspondence NetworksabstractWe tackle the problem of establishing dense pixel-wise correspondences between a pair of images. In this work, we introduce Dual-Resolution Correspondence Networks (DualRC-Net), to obtain pixel-wise correspondences in a coarse-to-fine manner. DualRC-Net extracts both coarse- and fine- resolution feature maps. The coarse maps are used to produce a full but coarse 4D correlation tensor, which is then refined by a learnable neighbourhood consensus module. The fine-resolution feature maps are used to obtain the final dense correspondences guided by the refined coarse 4D correlation tensor. The selected coarse-resolution matching scores allow the fine-resolution features to focus only on a limited number of possible matches with high confidence. In this way, DualRC-Net dramatically increases matching reliability and localisation accuracy, while avoiding to apply the expensive 4D convolution kernels on fine-resolution feature maps. We comprehensively evaluate our method on large-scale public benchmarks including HPatches, InLoc, and Aachen Day-Night. It achieves state-of-the-art results on all of them. Xinghui Li, Kai Han 0001, Shuda Li, Victor Adrian Prisacariu |
NeurIPS | 1 |