Yanqi Bao

dblp:294/6944 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0001-5298-7087ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 FISN: FInding Spatial Neighborhoods for Generalizable Novel View Synthesis
abstract
We present FISN, a generalizable novel view synthesis algorithm that enables feedforward inference of Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS) from reference images. Unlike existing work that either separately model the 3D feature space on each view or process multiview reference features by 3D-point-based view aggregation, FISN integrates multi-reference 3D cost volumes into a unified high-dimensional entity. Specifically, we reconceptualize the generalizable novel view synthesis task as a feedforward process of FInding Spatial Neighborhoods across this unified 4D feature space, comprising both view and spatial dimensions, and introduce View-Spatial Convolutions for direct 4D feature aggregation. This enhances the correlation among multiview neighboring points in a window-to-window manner and incorporates 3D spatial awareness. However, this approach poses two intertwined challenges: high computational expense for high-dimensional features and degraded rendering performance with low-resolution features. To address these challenges, FISN constructs a new efficient convolution paradigm, Decomposable View-Spatial Convolution, which includes a Spatial Cross Decomposition strategy as well as a Feature Compression and Upscaling module. This paradigm maintains multiview geometric consistency better than existing decomposition methods and achieves a balance between efficiency and fine-grained spatial features. Furthermore, by integrating Depth Refinement modules based on this paradigm, FISN further improves global depth understanding. Comprehensive evaluations on mainstream datasets and benchmarks demonstrate that FISN achieves state-of-the-art performance for both NeRF and 3DGS, and remains robust in challenging scenarios where existing 3DGS-based methods struggle, such as those with noisy poses or dense references. The code will be released soon.
Yanqi Bao, Tianyu Ding, Jing Huo, Wenbin Li 0006, Yang Gao 0001
IEEE Trans. Vis. Comput. Graph.1
2025 3D Gaussian Splatting: Survey, Technologies, Challenges, and Opportunities
abstract
3D Gaussian Splatting (3DGS) has emerged as a prominent technique with the potential to become a mainstream method for 3D representations. It can effectively transform multi-view images into explicit 3D Gaussian through efficient training, and achieve real-time rendering of novel views. This survey aims to analyze existing 3DGS-related works from multiple intersecting perspectives, including related tasks, technologies, challenges, and opportunities. The primary objective is to provide newcomers with a rapid understanding of the field and to assist researchers in methodically organizing existing technologies and challenges. Specifically, we delve into the optimization, application, and extension of 3DGS, categorizing them based on their focuses or motivations. Additionally, we summarize and classify nine types of technical modules and corresponding improvements identified in existing works. Based on these analyses, we further examine the common challenges and technologies across various tasks, proposing potential research opportunities.
Yanqi Bao, Tianyu Ding, Jing Huo, Yaoli Liu, Wenbin Li 0006, Yang Gao 0001, Jiebo Luo 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 InsertNeRF: Instilling Generalizability into NeRF with HyperNet Modules
abstract
Generalizing Neural Radiance Fields (NeRF) to new scenes is a significant challenge that existing approaches struggle to address without extensive modifications to vanilla NeRF framework. We introduce **InsertNeRF**, a method for **INS**tilling g**E**ne**R**alizabili**T**y into **NeRF**. By utilizing multiple plug-and-play HyperNet modules, InsertNeRF dynamically tailors NeRF's weights to specific reference scenes, transforming multi-scale sampling-aware features into scene-specific representations. This novel design allows for more accurate and efficient representations of complex appearances and geometries. Experiments show that this method not only achieves superior generalization performance but also provides a flexible pathway for integration with other NeRF-like systems, even in sparse input settings. Code will be available at: https://github.com/bbbbby-99/InsertNeRF.
Yanqi Bao, Tianyu Ding, Jing Huo, Wenbin Li 0006, Yang Gao 0001
ICLR1
2023 Region-Awared Transformer with Asymmetric Loss in Multi-Label Classification
abstract
Multi-label image classification (MLIC) deals with assigning multiple labels to each image, a easy task for human being while still a open problem in machine learning. The greatest challenge in MLIC lies in that different target objects in one image keep distinct viewpoints and scales. One effective way is to borrow the label-related information to guide the selection of interesting region, which will act an important role in classification. By leveraging the attention mechanism in transformer, we propose a region-awared transformer to focus on top related regions and neglect background interference. Furthermore, our approach can cope with the positive-negative imbalance by assigning them different exponential decay factors of positive and negative samples separately. Experiments on MS-COCO show a competitive performance against other state-of-the-art methods.
Yanqi Bao
ICASSP3
2023 Where and How: Mitigating Confusion in Neural Radiance Fields from Sparse Inputs
abstract
Neural Radiance Fields from Sparse inputs (NeRF-S) have shown great potential in synthesizing novel views with a limited number of observed viewpoints. However, due to the inherent limitations of sparse inputs and the gap between non-adjacent views, rendering results often suffer from over-fitting and foggy surfaces, a phenomenon we refer to as "CONFUSION" during volume rendering. In this paper, we analyze the root cause of this confusion and attribute it to two fundamental questions: "WHERE" and "HOW". To this end, we present a novel learning framework, WaH-NeRF, which effectively mitigates confusion by tackling the following challenges: (i) "WHERE" to Sample? in NeRF-S-we introduce a Deformable Sampling strategy and a Weight-based Mutual Information Loss to address sample-position confusion arising from the limited number of viewpoints; and (ii) "HOW" to Predict? in NeRF-S-we propose a Semi-Supervised NeRF learning Paradigm based on pose perturbation and a Pixel-Patch Correspondence Loss to alleviate prediction confusion caused by the disparity between training and testing viewpoints. By integrating our proposed modules and loss functions, WaH-NeRF outperforms previous methods under the NeRF-S setting. Code is available https://github.com/bbbbby-99/WaH-NeRF.
Yanqi Bao, Jing Huo, Tianyu Ding, Wenbin Li 0006, Yang Gao 0001
ACM Multimedia1
2022 Few-shot Semantic Segmentation with Support-induced Graph Convolutional Network
Jie Liu 0043, Yanqi Bao, Wenzhe Yin, Yang Gao 0001, Jan-Jakob Sonke, Efstratios Gavves
BMVC2
2022 Dynamic Prototype Convolution Network for Few-Shot Semantic Segmentation
abstract
The key challenge for few-shot semantic segmentation (FSS) is how to tailor a desirable interaction among sup-port and query features and/or their prototypes, under the episodic training scenario. Most existing FSS methods im-plement such support/query interactions by solely leveraging plain operations - e.g., cosine similarity and feature concatenation - for segmenting the query objects. How-ever, these interaction approaches usually cannot well capture the intrinsic object details in the query images that are widely encountered in FSS, e.g., if the query object to be segmented has holes and slots, inaccurate segmentation al-most always happens. To this end, we propose a dynamic prototype convolution network (DPCN) to fully capture the aforementioned intrinsic details for accurate FSS. Specifi-cally, in DPCN, a dynamic convolution module (DCM) is firstly proposed to generate dynamic kernels from support foreground, then information interaction is achieved by con-volution operations over query features using these kernels. Moreover, we equip DPCN with a support activation mod-ule (SAM) and a feature filtering module (FFM) to generate pseudo mask and filter out background information for the query images, respectively. SAM and FFM together can mine enriched context information from the query features. Our DPCN is also flexible and efficient under the k-shot FSS setting. Extensive experiments on PASCAL-5iand COCO 20ishow that DPCN yields superior performances under both 1-shot and 5-shot settings.
Jie Liu 0043, Yanqi Bao, Guosen Xie, Huan Xiong, Jan-Jakob Sonke, Efstratios Gavves
CVPR2
2022 Unidirectional RGB-T salient object detection with intertwined driving of encoding and fusion
Jie Wang 0095, Kechen Song, Yanqi Bao, Yunhui Yan, Yahong Han
Eng. Appl. Artif. Intell.3
2022 CGFNet: Cross-Guided Fusion Network for RGB-T Salient Object Detection
abstract
RGB salient object detection (SOD) has made great progress. However, the performance of this single-modal salient object detection will be significantly decreased when encountering some challenging scenes, such as low light or darkness. To deal with the above challenges, thermal infrared (T) image is introduced into the salient object detection. This fused modal is called RGB-T salient object detection. To achieve deep mining of the unique characteristics of single modal and the full integration of cross-modality information, a novel Cross-Guided Fusion Network (CGFNet) for RGB-T salient object detection is proposed. Specifically, a Cross-Scale Alternate Guiding Fusion (CSAGF) module is proposed to mine the high-level semantic information and provide global context support. Subsequently, we design a Guidance Fusion Module (GFM) to achieve sufficient cross-modality fusion by using single modal as the main guidance and the other modal as auxiliary. Finally, the Cross-Guided Fusion Module (CGFM) is presented and serves as the main decoding block. And each decoding block is consists of two parts with two modalities information of each being the main guidance, i.e., cross-shared Cross-Level Enhancement (CLE) and Global Auxiliary Enhancement (GAE). The main difference between the two parts is that the GFM using different modalities as the main guide. The comprehensive experimental results prove that our method achieves better performance than the state-of-the-art salient detection methods. The source code has released at:https://github.com/wangjie0825/CGFNet.git.
Jie Wang 0095, Kechen Song, Yanqi Bao, Liming Huang, Yunhui Yan
IEEE Trans. Circuits Syst. Video Technol.3
2021 Visible and thermal images fusion architecture for few-shot semantic segmentation
Yanqi Bao, Kechen Song, Jie Wang 0095, Liming Huang, Hongwen Dong, Yunhui Yan
J. Vis. Commun. Image Represent.1