EDBT 2026 Demo / reviewers in the wild / expert
Shuaifeng Zhi
dblp:209/3436
· DBLP profile ↗
22ranked-venue papers
3as first author
20since 2021 · last 2026
0000-0002-5927-5426ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 2 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PlaneRecTR++: Unified Query Learning for Joint 3D Planar Reconstruction and Pose EstimationabstractThe challenging task of 3D planar reconstruction from images involves several sub-tasks including frame-wise plane detection, segmentation, parameter regression and possibly depth prediction, along with cross-frame plane correspondence and relative camera pose estimation. Previous works adopt a divide and conquer strategy, addressing above sub-tasks with distinct network modules in a two-stage paradigm. Specifically, given an initial camera pose and per-frame plane predictions from the first stage, further exclusively designed modules relying on external plane correspondence labeling are applied to merge multi-view plane entities and produce refined camera pose. Notably, existing work fails to integrate these closely related sub-tasks into a unified framework, and instead addresses them separately and sequentially, which we identify as a primary source of performance limitations. Motivated by this finding and the success of query-based learning in enriching reasoning among semantic entities, in this paper, we propose PlaneRecTR++, a Transformer-based architecture, which for the first time unifies all tasks of multi-view planar reconstruction and pose estimation within a compact single-stage framework, eliminating the need for the initial pose estimation and supervision of plane correspondence. Extensive quantitative and qualitative experiments demonstrate that our proposed unified learning achieves mutual benefits across sub-tasks, achieving a new state-of-the-art performance on the public ScanNetv1, ScanNetv2, NYUv2-Plane, and MatterPort3D datasets. Jingjia Shi, Shuaifeng Zhi, Kai Xu 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Step-Wise Distribution-Aligned Style Prompt Tuning for Source-Free Cross-Domain Few-Shot LearningabstractExisting cross-domain few-shot learning (CDFSL) methods, which develop training strategies in the source domain to enhance model transferability, face challenges when applied to large-scale pre-trained models (LMs), as their source domains and training strategies are not accessible. Besides, fine-tuning LMs specifically for CDFSL requires substantial computational resources, which limits their practicality. Therefore, this paper investigates the source-free CDFSL (SF-CDFSL) problem to solve the few-shot learning (FSL) task in target domain using only a pre-trained model and a few target samples, without requiring source data or training strategies. However, the inaccessibility of source data prevents explicitly reducing the domain gaps between the source and target. To tackle this challenge, this paper proposes a novel approach, Step-wise Distribution-aligned Style Prompt Tuning (StepSPT), to implicitly narrow the domain gaps from the perspective of prediction distribution optimization. StepSPT initially proposes a style prompt that adjusts the target samples to mirror the expected distribution. Furthermore, StepSPT tunes the style prompt and classifier by exploring a dual-phase optimization process (external and internal processes). In the external process, a step-wise distribution alignment strategy is introduced to tune the proposed style prompt by factorizing the prediction distribution optimization problem into the multi-step distribution alignment problem. In the internal process, the classifier is updated via standard cross-entropy loss. Evaluation on 5 datasets illustrates the superiority of StepSPT over existing prompt tuning-based methods and state-of-the-art methods (SOTAs). Furthermore, ablation studies and performance analyzes highlight the efficacy of StepSPT. Huali Xu, Li Liu 0002, Tianpeng Liu, Shuaifeng Zhi, Shuzhou Sun, Ming-Ming Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | RemixFusion: Residual-based Mixed Representation for Large-scale Online RGB-D ReconstructionabstractThe introduction of the neural implicit representation has notably propelled the advancement of online dense reconstruction techniques. Compared to traditional explicit representations, such as TSDF, it substantially improves the mapping completeness and memory efficiency. However, the lack of reconstruction details and the time-consuming learning of neural representations hinder the widespread application of neural-based methods to large-scale online reconstruction. We introduce RemixFusion, a novel residual-based mixed representation for scene reconstruction and camera pose estimation dedicated to high-quality and large-scale online RGB-D reconstruction. In particular, we propose a residual-based map representation comprised of an explicit coarse TSDF grid and an implicit neural module that produces residuals representing fine-grained details to be added to the coarse grid. Such mixed representation allows for detail-rich reconstruction with bounded time and memory budget, contrasting with the overly-smoothed results by the purely implicit representations, thus paving the way for high-quality camera tracking. Furthermore, we extend the residual-based representation to handle multi-frame joint pose optimization via bundle adjustment (BA). In contrast to the existing methods, which optimize poses directly, we opt to optimize pose changes. Combined with a novel technique for adaptive gradient amplification, our method attains better optimization convergence and global optimality. Furthermore, we adopt a local moving volume to factorize the whole mixed scene representation with a divide-and-conquer design to facilitate efficient online learning in our residual-based framework. Extensive experiments demonstrate that our method surpasses all state-of-the-art ones, including those based either on explicit or implicit representations, in terms of the accuracy of both mapping and tracking on large-scale scenes. Project page can be found at https://lanlan96.github.io/RemixFusion/ . Yuqing Lan, Chenyang Zhu 0002, Shuaifeng Zhi, Jiazhao Zhang, Zhoufeng Wang, Renjiao Yi, Yijie Wang 0001, Kai Xu 0004 |
ACM Trans. Graph. | 3 |
| 2025 | Luminance-Aware Statistical Quantization: Unsupervised Hierarchical Learning for Illumination EnhancementabstractLow-light image enhancement (LLIE) faces persistent challenges in balancing reconstruction fidelity with cross-scenario generalization. While existing methods predominantly focus on deterministic pixel-level mappings between paired low/normal-light images, they often neglect the continuous physical process of luminance transitions in real-world environments, leading to performance drop when normal-light references are unavailable. Inspired by empirical analysis of natural luminance dynamics revealing power-law distributed intensity transitions, this paper introduces Luminance-Aware Statistical Quantification (LASQ), a novel framework that reformulates LLIE as a statistical sampling process over hierarchical luminance distributions. Our LASQ re-conceptualizes luminance transition as a power-law distribution in intensity coordinate space that can be approximated by stratified power functions, therefore, replacing deterministic mappings with probabilistic sampling over continuous luminance layers. A diffusion forward process is designed to autonomously discover optimal transition paths between luminance layers, achieving unsupervised distribution emulation without normal-light references.
In this way, it considerably improves the performance in practical situations, enabling more adaptable and versatile light restoration. This framework is also readily applicable to cases with normal-light references, where it achieves superior performance on domain-specific datasets alongside better generalization-ability across non-reference datasets. The code is available at: https://github.com/XYLGroup/LASQ. Derong Kong, Zhixiong Yang 0001, Shengxi Li, Shuaifeng Zhi, Li Liu 0002, Zhen Liu 0004, Jingyuan Xia |
NeurIPS | 4 |
| 2025 | A Causal Adjustment Module for Debiasing Scene Graph GenerationabstractWhile recent debiasing methods for Scene Graph Generation (SGG) have shown impressive performance, these efforts often attribute model bias solely to the long-tail distribution of relationships, overlooking the more profound causes stemming from skewed object and object pair distributions. In this paper, we employ causal inference techniques to model the causality among these observed skewed distributions. Our insight lies in the ability of causal inference to capture the unobservable causal effects between complex distributions, which is crucial for tracing the roots of model bias. Specifically, we introduce the Mediator-based Causal Chain Model (MCCM), which, in addition to modeling causality among objects, object pairs, and relationships, incorporates mediator variables, i.e., cooccurrence distribution, for complementing the causality. Following this, we propose the Causal Adjustment Module (CAModule) to estimate the modeled causal structure, using variables from MCCM as inputs to produce a set of adjustment factors aimed at correcting biased model predictions. Moreover, our method enables the composition of zero-shot relationships, thereby enhancing the model's ability to recognize such relationships. Experiments conducted across various SGG backbones and popular benchmarks demonstrate that CAModule achieves state-of-the-art mean recall rates, with significant improvements also observed on the challenging zero-shot recall rate metric. Li Liu 0002, Shuzhou Sun, Shuaifeng Zhi, Fan Shi 0003, Zhen Liu 0004, Janne Heikkilä, Yongxiang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | A Reverse Causal Framework to Mitigate Spurious Correlations for Debiasing Scene Graph GenerationabstractExisting two-stage Scene Graph Generation (SGG) frameworks typically incorporate a detector to extract relationship features and a classifier to categorize these relationships; therefore, the training paradigm follows a causal chain structure, where the detector's inputs determine the classifier's inputs, which in turn influence the final predictions. However, such a causal chain structure can yield spurious correlations between the detector's inputs and the final predictions, i.e., the prediction of a certain relationship may be influenced by other relationships. This influence can induce at least two observable biases: tail relationships are predicted as head ones, and foreground relationships are predicted as background ones; notably, the latter bias is seldom discussed in the literature. To address this issue, we propose reconstructing the causal chain structure into a reverse causal structure, wherein the classifier's inputs are treated as the confounder, and both the detector's inputs and the final predictions are viewed as causal variables. Specifically, we term the reconstructed causal paradigm as the Reverse causal Framework for SGG (RcSGG). RcSGG initially employs the proposed Active Reverse Estimation (ARE) to intervene on the confounder to estimate the reverse causality, i.e., the causality from final predictions to the classifier's inputs. Then, the Maximum Information Sampling (MIS) is suggested to enhance the reverse causality estimation further by considering the relationship information. Theoretically, RcSGG can mitigate the spurious correlations inherent in the SGG framework, subsequently eliminating the induced biases. Comprehensive experiments on popular benchmarks and diverse SGG frameworks show the state-of-the-art mean recall rate. Shuzhou Sun, Li Liu 0002, Tianpeng Liu, Shuaifeng Zhi, Ming-Ming Cheng, Janne Heikkilä, Yongxiang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | GDROS: A Geometry-Guided Dense Registration Framework for Optical-SAR Images Under Large Geometric TransformationsabstractRegistration of optical and synthetic aperture radar (SAR) remote sensing images serves as a critical foundation for image fusion and visual navigation tasks. This task is particularly challenging because of their modal discrepancy, primarily manifested as severe nonlinear radiometric differences (NRD), geometric distortions, and noise variations. Under large geometric transformations, existing classical template-based and sparse keypoint-based strategies struggle to achieve reliable registration results for optical-SAR image pairs. To address these limitations, we propose GDROS, a geometry-guided dense registration framework leveraging global cross-modal image interactions. First, we extract cross-modal deep features from optical and SAR images through a CNN-Transformer hybrid feature extraction module, upon which a multi-scale 4D correlation volume is constructed and iteratively refined to establish pixel-wise dense correspondences. Subsequently, we implement a least squares regression (LSR) module to geometrically constrain the predicted dense optical flow field. Such geometry guidance mitigates prediction divergence by directly imposing an estimated affine transformation on the final flow predictions. Extensive experiments have been conducted on three representative datasets WHU-Opt-SAR dataset, OS dataset, and UBCv2 dataset with different spatial resolutions, demonstrating robust performance of our proposed method across different imaging resolutions. Qualitative and quantitative results show that GDROS significantly outperforms current state-of-the-art methods in all metrics. Our source code will be released at: https://github.com/Zi-Xuan-Sun/GDROS. Zixuan Sun, Shuaifeng Zhi, Ruize Li, Jingyuan Xia, Yongxiang Liu, Weidong Jiang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Looking Beneath More: A Sequence-based Localizing Ground Penetrating Radar FrameworkabstractLocalizing ground penetrating radar (LGPR) has been proven to be a promising technology for robot localization in various dynamic environments. However, the extreme scarcity of underground features introduces false candidate matches and brings unique challenges to this task. In this paper, we propose a sequence-based framework for LGPR to address the aforementioned issues. Specifically, we first introduce a trainable strategy to extract robust underground features in multi-weather conditions. By further using sequential information, our LGPR system can observe richer underground scene contexts, and the associated multi-frame scans could also improve the performance of underground place recognition. We demonstrate the superiority of our proposed method by comparing it against several recent state-of-the-art baseline methods applied to GPR image tasks. Experimental results on large public and self-collected datasets show that our proposed framework significantly improves the performance of various baselines in different scenarios. Shuaifeng Zhi, Yuelin Yuan, Beizhen Bi, Qin Xin 0004, Xiaotao Huang 0001, Liang Shen 0003 |
ICRA | 2 |
| 2024 | OS3Flow: Optical and SAR Image Registration Using Symmetry-Guided Semi-Dense Optical FlowabstractRegistration of optical and synthetic aperture radar (SAR) image pairs is a fundamental task in various remote sensing applications, including image fusion, target localization, and object detection. Unlike homogeneous image pairs, optical and SAR image pairs exhibit a significant modality gap, making it exceptionally challenging to extract consistent and reliable features. Particularly for optical and SAR image pairs with substantial geometric differences, few methods can achieve high-precision registration. To address this challenging task, we introduce a novel registration framework, called OS3Flow, leveraging on the implicit symmetry between heterogeneous image pairs to extract high-quality semi-dense flow estimations. We start by training the network in a multi-task manner using a standard flow regression loss as well as a symmetry loss with reverse input order. A confidence mask thus can be generated to measure the similarity between predictions at inference time. We then perform a linear regression upon selected flows with high confidence to estimate the parameters of underlying affine transformation. Under large transformations, our proposed method achieves an average registration error of less than 3 pixels on the public OS dataset and WHU-OPT-SAR dataset, demonstrating superior accuracy and robustness compared to state-of-the-art methods. Zixuan Sun, Shuaifeng Zhi, Kai Huo, Xuecong Liu, Weidong Jiang, Yongxiang Liu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Meta-learning based blind image super-resolution approach to different degradations
Zhixiong Yang 0001, Jingyuan Xia, Shengxi Li, Wende Liu, Shuaifeng Zhi, Shuanghui Zhang, Li Liu 0002, Yaowen Fu, Deniz Gündüz |
Neural Networks | 5 |
| 2024 | SSR-2D: Semantic 3D Scene Reconstruction From 2D ImagesabstractMost deep learning approaches to comprehensive semantic modeling of 3D indoor spaces require costly dense annotations in the 3D domain. In this work, we explore a central 3D scene modeling task, namely, semantic scene reconstruction without using any 3D annotations. The key idea of our approach is to design a trainable model that employs both incomplete 3D reconstructions and their corresponding source RGB-D images, fusing cross-domain features into volumetric embeddings to predict complete 3D geometry, color, and semantics with only 2D labeling which can be either manual or machine-generated. Our key technical innovation is to leverage differentiable rendering of color and semantics to bridge 2D observations and unknown 3D space, using the observed RGB images and 2D semantics as supervision, respectively. We additionally develop a learning pipeline and corresponding method to enable learning from imperfect predicted 2D labels, which could be additionally acquired by synthesizing in an augmented set of virtual training views complementing the original real captures, enabling more efficient self-supervision loop for semantics. As a result, our end-to-end trainable solution jointly addresses geometry completion, colorization, and semantic mapping from limited RGB-D images, without relying on any 3D ground-truth information. Our method achieves state-of-the-art performance of semantic scene completion on two large-scale benchmark datasets MatterPort3D and ScanNet, surpasses baselines even with costly 3D annotations in predicting both geometry and semantics. To our knowledge, our method is also the first 2D-driven method addressing completion and semantic segmentation of real-world 3D scans simultaneously. Junwen Huang 0001, Alexey Artemov, Yujin Chen, Shuaifeng Zhi, Kai Xu 0004, Matthias Nießner |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Enhancing Information Maximization With Distance-Aware Contrastive Learning for Source-Free Cross-Domain Few-Shot LearningabstractExisting Cross-Domain Few-Shot Learning (CDFSL) methods require access to source domain data to train a model in the pre-training phase. However, due to increasing concerns about data privacy and the desire to reduce data transmission and training costs, it is necessary to develop a CDFSL solution without accessing source data. For this reason, this paper explores a Source-Free CDFSL (SF-CDFSL) problem, in which CDFSL is addressed through the use of existing pretrained models instead of training a model with source data, avoiding accessing source data. However, due to the lack of source data, we face two key challenges: effectively tackling CDFSL with limited labeled target samples, and the impossibility of addressing domain disparities by aligning source and target domain distributions. This paper proposes an Enhanced Information Maximization with Distance-Aware Contrastive Learning (IM-DCL) method to address these challenges. Firstly, we introduce the transductive mechanism for learning the query set. Secondly, information maximization (IM) is explored to map target samples into both individual certainty and global diversity predictions, helping the source model better fit the target data distribution. However, IM fails to learn the decision boundary of the target task. This motivates us to introduce a novel approach called Distance-Aware Contrastive Learning (DCL), in which we consider the entire feature set as both positive and negative sets, akin to Schrödinger's concept of a dual state. Instead of a rigid separation between positive and negative sets, we employ a weighted distance calculation among features to establish a soft classification of the positive and negative sets for the entire feature set. We explore three types of negative weights to enhance the performance of CDFSL. Furthermore, we address issues related to IM by incorporating contrastive constraints between object features and their corresponding positive and negative sets. Evaluations of the 4 datasets in the BSCD-FSL benchmark indicate that the proposed IM-DCL, without accessing the source domain, demonstrates superiority over existing methods, especially in the distant domain task. Additionally, the ablation study and performance analysis confirmed the ability of IM-DCL to handle SF-CDFSL. The code will be made public at https://github.com/xuhuali-mxj/IM-DCL. Huali Xu, Li Liu 0002, Shuaifeng Zhi, Shaojing Fu, Zhuo Su 0002, Ming-Ming Cheng, Yongxiang Liu |
IEEE Trans. Image Process. | 3 |
| 2023 | ROFusion: Efficient Object Detection Using Hybrid Point-Wise Radar-Optical Fusion
Shuaifeng Zhi, Zhenhua Du, Li Liu 0002, Xinyu Zhang 0010, Kai Huo, Weidong Jiang |
ICANN (7) | 2 |
| 2023 | PlaneRecTR: Unified Query Learning for 3D Plane Recovery from a Single Viewabstract3D plane recovery from a single image can usually be divided into several subtasks of plane detection, segmentation, parameter estimation and possibly depth estimation. Previous works tend to solve it by either extending the RCNN-based segmentation network or the dense pixel embedding-based clustering framework. However, none of them tried to integrate above related subtasks into a unified framework but treated them separately and sequentially, which we suspect is potentially a main source of performance limitation for existing approaches. Motivated by this finding and the success of query-based learning in enriching reasoning among semantic entities, in this paper, we propose PlaneRecTR, a Transformer-based architecture, which for the first time unifies all subtasks related to single-view plane recovery with a single compact model. Extensive quantitative and qualitative experiments demonstrate that our proposed unified learning achieves mutual benefits across subtasks, obtaining a new state-of-the-art performance on public ScanNet and NYUv2-Plane datasets. Codes are available at https://github.com/SJingjia/PlaneRecTR. Jingjia Shi, Shuaifeng Zhi, Kai Xu 0004 |
ICCV | 2 |
| 2023 | Cross-Domain Few-Shot Classification Via Inter-Source StylizationabstractThe goal of Cross-Domain Few-Shot Classification (CDFSC) is to accurately classify a target dataset with limited labelled data by exploiting the knowledge of a richly labelled auxiliary dataset, despite the differences between the domains of the two datasets. Some existing approaches require labelled samples from multiple domains for model training. However, these methods fail when the sample labels are scarce. To overcome this challenge, this paper proposes a solution that makes use of multiple source domains without the need for additional labeling costs. Specifically, one of the source domains is completely tagged, while the others are untagged. An Inter-Source Stylization Network (ISSNet) is then introduced to enhance stylisation across multiple source domains, enriching data distribution and model’s generalization capabilities. Experiments on 8 target datasets show that ISSNet leverages unlabelled data from multiple source data and significantly reduces the negative impact of domain gaps on classification performance compared to several baseline methods. Huali Xu, Shuaifeng Zhi, Li Liu 0002 |
ICIP | 2 |
| 2023 | Evidential Uncertainty and Diversity Guided Active Learning for Scene Graph Generation
Shuzhou Sun, Shuaifeng Zhi, Janne Heikkilä, Li Liu 0002 |
ICLR | 2 |
| 2023 | Unbiased Scene Graph Generation via Two-Stage Causal ModelingabstractDespite the impressive performance of recent unbiased Scene Graph Generation (SGG) methods, the current debiasing literature mainly focuses on the long-tailed distribution problem, whereas it overlooks another source of bias, i.e., semantic confusion, which makes the SGG model prone to yield false predictions for similar relationships. In this paper, we explore a debiasing procedure for the SGG task leveraging causal inference. Our central insight is that the Sparse Mechanism Shift (SMS) in causality allows independent intervention on multiple biases, thereby potentially preserving head category performance while pursuing the prediction of high-informative tail relationships. However, the noisy datasets lead to unobserved confounders for the SGG task, and thus the constructed causal models are always causal-insufficient to benefit from SMS. To remedy this, we propose Two-stage Causal Modeling (TsCM) for the SGG task, which takes the long-tailed distribution and semantic confusion as confounders to the Structural Causal Model (SCM) and then decouples the causal intervention into two stages. The first stage is causal representation learning, where we use a novel Population Loss (P-Loss) to intervene in the semantic confusion confounder. The second stage introduces the Adaptive Logit Adjustment (AL-Adjustment) to eliminate the long-tailed distribution confounder to complete causal calibration learning. These two stages are model agnostic and thus can be used in any SGG model that seeks unbiased predictions. Comprehensive experiments conducted on the popular SGG backbones and benchmarks show that our TsCM can achieve state-of-the-art performance in terms of mean recall rate. Furthermore, TsCM can maintain a higher recall rate than other debiasing methods, which indicates that our method can achieve a better tradeoff between head and tail relationships. Shuzhou Sun, Shuaifeng Zhi, Qing Liao 0001, Janne Heikkilä, Li Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | RaNeRF: Neural 3-D Reconstruction of Space Targets From ISAR Image SequencesabstractCompared to 2D inverse synthetic aperture radar (ISAR) images of a space target, its 3D model can provide adequate details and accurate measurement parameters. However, it is challenging to tackle the problem of feature extraction and correlation during 3D reconstruction of space targets purely based on radar image sequences, due to their lack of clear evidence in imaging similarity compared to optical images. To address this problem, this paper proposes radar neural radiance fields (i.e. RaNeRF), which is a novel 3D reconstruction method using only observed ISAR image sequences. Firstly, the 3D structure of a target is represented as a continuous 6D function of space positions and viewing directions using a fully-connected deep network. Secondly, the relationship between the 3D structure and 2D ISAR images of the target is constructed to enable differential rendering of ISAR images. Our overall pipeline can thus be trained using the discrepancy between the modulus of rendered and observed ISAR images in a purely self-supervised manner without 3D supervision. Finally, the 3D mesh model of the target can be retrieved from the learned density field via marching cube. As a result, the proposed RaNeRF can directly reconstruct the 3D structure of targets without explicit feature extraction and correlation of ISAR image sequences. Both quantitative and qualitative results verify the effectiveness of the proposed method. Compared to conventional baseline methods using point clouds, our reconstructed structure is more complete and accurate. In addition, the optimized model can synthesize ISAR images at novel observation direction, which can be used for downstream tasks including data augmentation and target recognition. Afei Liu, Shuanghui Zhang, Chi Zhang 0045, Shuaifeng Zhi, Xiang Li 0014 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Bootstrapping Semantic Segmentation with Regional Contrast
Shikun Liu, Shuaifeng Zhi, Edward Johns, Andrew J. Davison |
ICLR | 2 |
| 2021 | In-Place Scene Labelling and Understanding with Implicit Scene RepresentationabstractSemantic labelling is highly correlated with geometry and radiance reconstruction, as scene entities with similar shape and appearance are more likely to come from similar classes. Recent implicit neural reconstruction techniques are appealing as they do not require prior training data, but the same fully self-supervised approach is not possible for semantics because labels are human-defined properties.We extend neural radiance fields (NeRF) to jointly encode semantics with appearance and geometry, so that complete and accurate 2D semantic labels can be achieved using a small amount of in-place annotations specific to the scene. The intrinsic multi-view consistency and smoothness of NeRF benefit semantics by enabling sparse labels to efficiently propagate. We show the benefit of this approach when labels are either sparse or very noisy in room-scale scenes. We demonstrate its advantageous properties in various interesting applications such as an efficient scene labelling tool, novel semantic view synthesis, label denoising, super-resolution, label interpolation and multi-view semantic label fusion in visual semantic mapping systems. Shuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger, Andrew J. Davison |
ICCV | 1 |
| 2019 | SceneCode: Monocular Dense Semantic Reconstruction Using Learned Encoded Scene RepresentationsabstractSystems which incrementally create 3D semantic maps from image sequences must store and update representations of both geometry and semantic entities. However, while there has been much work on the correct formulation for geometrical estimation, state-of-the-art systems usually rely on simple semantic representations which store and update independent label estimates for each surface element (depth pixels, surfels, or voxels). Spatial correlation is discarded, and fused label maps are incoherent and noisy. We introduce a new compact and optimisable semantic representation by training a variational auto-encoder that is conditioned on a colour image. Using this learned latent space, we can tackle semantic label fusion by jointly optimising the low-dimenional codes associated with each of a set of overlapping images, producing consistent fused label maps which preserve spatial correlation. We also show how this approach can be used within a monocular keyframe based semantic mapping system where a similar code approach is used for geometry. The probabilistic formulation allows a flexible formulation where we can jointly estimate motion, geometry and semantics in a unified optimisation. Shuaifeng Zhi, Michael Bloesch, Stefan Leutenegger, Andrew J. Davison |
CVPR | 1 |
| 2018 | Toward real-time 3D object recognition: A lightweight volumetric CNN framework using multitask learning
Shuaifeng Zhi, Yongxiang Liu, Xiang Li 0014, Yulan Guo |
Comput. Graph. | 1 |