Yuantao Chen

dblp:140/1870 · DBLP profile ↗
← Back
25ranked-venue papers
15as first author
23since 2021 · last 2025
0000-0003-2277-1765ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 11 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Computer networks · 2 · 2 first-author
YearPublicationVenuePosition
2025 PUGS: Zero-Shot Physical Understanding with Gaussian Splatting
abstract
Current robotic systems can understand the categories and poses of objects well. But understanding physical properties like mass, friction, and hardness, in the wild, remains challenging. We propose a new method that reconstructs 3D objects using the Gaussian splatting representation and predicts various physical properties in a zero-shot manner. We propose two techniques during the reconstruction phase: a geometryaware regularization loss function to improve the shape quality and a region-aware feature contrastive loss function to promote region affinity. Two other new techniques are designed during inference: a feature-based property propagation module and a volume integration module tailored for the Gaussian representation. Our framework is named as zero-shot physical understanding with Gaussian splatting, or PUGS. PUGS achieves new state-of-the-art results on the standard benchmark of ABO-500 mass prediction. We provide extensive quantitative ablations and qualitative visualization to demonstrate the mechanism of our designs. We show the proposed methodology can help address challenging real-world grasping tasks. Our codes, data, and models are available at https://github.com/EverNorif/PUGS
Yinghao Shuai, Yuantao Chen, Zijian Jiang, Nan Wang 0041, Jv Zheng, Jianzhu Ma, Meng Yang 0035, Zhicheng Wang 0022, Wenbo Ding 0001, Hao Zhao 0002
ICRA3
2025 Unifying Appearance Codes and Bilateral Grids for Driving Scene Gaussian Splatting
abstract
Neural rendering techniques, including NeRF and Gaussian Splatting (GS), rely on photometric consistency to produce high-quality reconstructions. However, in real-world driving scenarios, it is challenging to guarantee perfect photometric consistency in acquired images. Appearance codes have been widely used to address this issue, but their modeling capability is limited, as a single code is applied to the entire image. Recently, the bilateral grid was introduced to perform pixel-wise color mapping, but it is difficult to optimize and constrain effectively. In this paper, we propose a novel multi-scale bilateral grid that unifies appearance codes and bilateral grids. We demonstrate that this approach significantly improves geometric accuracy in dynamic, decoupled autonomous driving scene reconstruction, outperforming both appearance codes and bilateral grids. This is crucial for autonomous driving, where accurate geometry is important for obstacle avoidance and control. Our method shows strong results across four datasets: Waymo, NuScenes, Argoverse, and PandaSet. We further demonstrate that the improvement in geometry is driven by the multi-scale bilateral grid, which effectively reduces floaters caused by photometric inconsistency.
Nan Wang 0041, Lixing Xiao, Yuantao Chen, Weiqing Xiao, Pierre Merriaux, Ziyang Yan, Saining Zhang, Shaocong Xu, Chongjie Ye, Bohan Li 0015, Zhaoxi Chen 0009, Tianfan Xue, Hao Zhao 0002
NeurIPS3
2025 Self-Aligning Depth-Regularized Radiance Fields for Asynchronous RGB-D Sequences
abstract
It has been shown that learning radiance fields with depth rendering and depth supervision can effectively promote the quality and convergence of view synthesis. However, this paradigm requires input RGB-D sequences to be synchronized. In the UAV city modeling scenario, there exists asynchrony between RGB images and depth images due to the different frequencies of the solid-state LiDAR and RGB sensors. To synthesize high-quality views in such a scenario, we propose a novel time-pose function, which is an implicit network that maps timestamps to SE(3) elements. To train this function, we also design a joint optimization scheme to jointly learn the large-scale depth-regularized radiance fields and the time-pose function. Furthermore, we propose a large synthetic dataset with diverse controlled mismatches and ground truth to evaluate this new problem setting systematically. The proposed approach has been evaluated on both datasets and in a real drone. To evaluate the impact of view density, each algorithm was test on three different trajectories with different view densities. Compared to state-of-the-art baseline methods, the proposed approach reduces reconstruction error by 35.26% in city modeling scenarios. Our code is available at github.com/saythe17/AsyncNeRF.
Andong Yang, Yuantao Chen, Runyi Yang, Zhenxin Zhu, Hao Zhao 0002, Guyue Zhou
WACV3
2025 ATM-DEN: Image Inpainting via attention transfer module and Decoder-Encoder network
Yuantao Chen
Signal Process. Image Commun.2
2024 Drone-assisted Road Gaussian Splatting with Cross-view Uncertainty
Saining Zhang, Baijun Ye, Xiaoxue Chen, Yuantao Chen, Zongzheng Zhang, Yongliang Shi, Hao Zhao 0002
BMVC4
2024 GaussReg: Fast 3D Registration with Gaussian Splatting
Yinglin Xu, Yuantao Chen, Wensen Feng, Xiaoguang Han 0001
ECCV (15)4
2024 Camera Relocalization in Shadow-free Neural Radiance Fields
abstract
Camera relocalization is a crucial problem in computer vision and robotics. Recent advancements in neural radiance fields (NeRFs) have shown promise in synthesizing photo-realistic images. Several works have utilized NeRFs for refining camera poses, but they do not account for lighting changes that can affect scene appearance and shadow regions, causing a degraded pose optimization process. In this paper, we propose a two-staged pipeline that normalizes images with varying lighting and shadow conditions to improve camera relocalization. We implement our scene representation upon a hash-encoded NeRF which significantly boosts up the pose optimization process. To account for the noisy image gradient computing problem in grid-based NeRFs, we further propose a re-devised truncated dynamic low-pass filter (TDLF) and a numerical gradient averaging technique to smoothen the process. Experimental results on several datasets with varying lighting conditions demonstrate that our method achieves state-of-the-art results in camera relocalization under varying lighting conditions. Code and data will be made publicly available.
Shiyao Xu, Caiyun Liu 0004, Yuantao Chen, Zhenxin Zhu, Zike Yan, Yongliang Shi, Hao Zhao 0002, Guyue Zhou
ICRA3
2024 Blending Distributed NeRFs with Tri-stage Robust Pose Optimization
abstract
Due to the limited model capacity, leveraging distributed Neural Radiance Fields (NeRFs) for modeling extensive urban environments has become a necessity. However, current distributed NeRF registration approaches encounter aliasing artifacts, arising from discrepancies in rendering resolutions and suboptimal pose precision. These factors collectively deteriorate the fidelity of pose estimation within NeRF frameworks, resulting in occlusion artifacts during the NeRF blending stage. In this paper, we present a distributed NeRF system with tri-stage pose optimization. In the first stage, precise poses of images are achieved by bundle adjusting Mip-NeRF 360 with a coarse-to-fine strategy. In the second stage, we incorporate the inverting Mip-NeRF 360, coupled with the truncated dynamic low-pass filter, to enable the achievement of robust and precise poses, termed Frame2Model optimization. On top of this, we obtain a coarse transformation between NeRFs in different coordinate systems. In the third stage, we fine-tune the transformation between NeRFs by Model2Model pose optimization. After obtaining precise transformation parameters, we proceed to implement NeRF blending, showcasing superior performance metrics in both real-world and simulation scenarios. Codes and data will be publicly available at https://github.com/boilcy/Distributed-NeRF.
Baijun Ye, Caiyun Liu 0004, Xiaoyu Ye, Yuantao Chen, Yuhai Wang, Zike Yan, Yongliang Shi, Hao Zhao 0002, Guyue Zhou
IROS4
2024 MFMAM: Image inpainting via multi-scale feature module with attention module
Yuantao Chen, Runlong Xia, Kai Yang 0010, Ke Zou
Comput. Vis. Image Underst.1
2024 Image inpainting algorithm based on inference attention module and two-stage network
Yuantao Chen, Runlong Xia, Kai Yang 0010, Ke Zou
Eng. Appl. Artif. Intell.1
2024 MICU: Image super-resolution via multi-level information compensation and U-net
Yuantao Chen, Runlong Xia, Kai Yang 0010, Ke Zou
Expert Syst. Appl.1
2024 MFFN: image super-resolution via multi-level features fusion network
Yuantao Chen, Runlong Xia, Kai Yang 0010, Ke Zou
Vis. Comput.1
2023 LATITUDE: Robotic Global Localization with Truncated Dynamic Low-pass Filter in City-scale NeRF
abstract
Neural Radiance Fields (NeRFs) have made great success in representing complex 3D scenes with high-resolution details and efficient memory. Nevertheless, current NeRF - based pose estimators have no initial pose prediction and are prone to local optima during optimization. In this paper, we present LATITUDE: Global Localization with Truncated Dynamic Low-pass Filter, which introduces a two-stage localization mechanism in city-scale NeRF. In place recognition stage, we train a regressor through images generated from trained NeRFs, which provides an initial value for global localization. In pose optimization stage, we minimize the residual between the observed image and rendered image by directly optimizing the pose on the tangent plane. To avoid falling into local optimum, we introduce a Truncated Dynamic Low-pass Filter (TDLF) for coarse-to-fine pose registration. We evaluate our method on both synthetic and real-world data and show its potential applications for high-precision navigation in large-scale city scenes. Codes and dataset will be publicly available at https://github.com/jike5/LATITUDE.
Zhenxin Zhu, Yuantao Chen, Zirui Wu, Yongliang Shi, Chuxuan Li, Pengfei Li 0007, Hao Zhao 0002, Guyue Zhou
ICRA2
2023 FFTI: Image inpainting algorithm via features fusion and two-steps inpainting
Yuantao Chen, Runlong Xia, Ke Zou, Kai Yang 0010
J. Vis. Commun. Image Represent.1
2023 Corrigendum to "FFTI: Image inpainting algorithm via features fusion and two-steps inpainting" [J. Visual Commun. Image Represent. 91 (2023) 103776]
Yuantao Chen, Runlong Xia, Ke Zou, Kai Yang 0010
J. Vis. Commun. Image Represent.1
2023 DGCA: high resolution image inpainting via DR-GAN and contextual attention
Yuantao Chen, Runlong Xia, Kai Yang 0010, Ke Zou
Multim. Tools Appl.1
2022 ETBRec: a novel recommendation algorithm combining the double influence of trust relationship and expert users
Zhenchun Duan, Yuantao Chen
Appl. Intell.3
2022 Retracted: Multiscale fast correlation filtering tracking algorithm based on a feature fusion model
abstract
Retraction: Multiscale fast correlation filtering tracking algorithm based on a feature fusion model Yuantao Chen, Jin Wang, Songjie Liu, Xi Chen, Jie Xiong, Jingbo Xie, Kai Yang, 2021, 33 (15), (https://doi.org/10.1002/cpe.5533) The above article, published online on 23 October 2019 in Wiley Online Library (wileyonlinelibrary.com), has been retracted by agreement between the authors, the journal Editors, David W. Walker, Jinjun Chen, Nitin Auluck and Martin Berzins, and John Wiley and Sons Ltd. The retraction has been agreed due to scientific errors arising from the incorrect use of materials and data, which have led to conclusions that are unreliable.
Yuantao Chen, Jin Wang 0001, Songjie Liu, Jingbo Xie, Kai Yang 0010
Concurr. Comput. Pract. Exp.1
2021 Image super-resolution reconstruction based on feature map attention mechanism
Yuantao Chen, Linwu Liu, Volachith Phonevilay, Ke Gu 0002, Runlong Xia, Jingbo Xie, Qian Zhang 0079, Kai Yang 0010
Appl. Intell.1
2021 Research on image Inpainting algorithm of improved GAN based on two-discriminations networks
Yuantao Chen, Haopeng Zhang 0010, Linwu Liu, Qian Zhang 0079, Kai Yang 0010, Runlong Xia, Jingbo Xie
Appl. Intell.1
2021 The image annotation algorithm using convolutional features from intermediate layer of deep learning
Yuantao Chen, Linwu Liu, Jiajun Tao, Runlong Xia, Qian Zhang 0079, Kai Yang 0010, Jingbo Xie
Multim. Tools Appl.1
2021 The face image super-resolution algorithm based on combined representation learning
Yuantao Chen, Volachith Phonevilay, Jiajun Tao, Runlong Xia, Qian Zhang 0079, Kai Yang 0010, Jingbo Xie
Multim. Tools Appl.1
2021 The improved image inpainting algorithm via encoder and similarity constraint
Yuantao Chen, Linwu Liu, Jiajun Tao, Runlong Xia, Qian Zhang 0079, Kai Yang 0010
Vis. Comput.1
2020 Saliency Detection via the Improved Hierarchical Principal Component Analysis Method
abstract
Aiming at the problems of intensive background noise, low accuracy, and high computational complexity of the current significant object detection methods, the visual saliency detection algorithm based on Hierarchical Principal Component Analysis (HPCA) has been proposed in the paper. Firstly, the original RGB image has been converted to a grayscale image, and the original grayscale image has been divided into eight layers by the bit surface stratification technique. Each image layer contains significant object information matching the layer image features. Secondly, taking the color structure of the original image as the reference image, the grayscale image is reassigned by the grayscale color conversion method, so that the layered image not only reflects the original structural features but also effectively preserves the color feature of the original image. Thirdly, the Principal Component Analysis (PCA) has been performed on the layered image to obtain the structural difference characteristics and color difference characteristics of each layer of the image in the principal component direction. Fourthly, two features are integrated to get the saliency map with high robustness and to further refine our results; the known priors have been incorporated on image organization, which can place the subject of the photograph near the center of the image. Finally, the entropy calculation has been used to determine the optimal image from the layered saliency map; the optimal map has the least background information and most prominently saliency objects than others. The object detection results of the proposed model are closer to the ground truth and take advantages of performance parameters including precision rate (PRE), recall rate (REC), and F -measure (FME). The HPCA model’s conclusion can obviously reduce the interference of redundant information and effectively separate the saliency object from the background. At the same time, it had more improved detection accuracy than others.
Yuantao Chen, Jiajun Tao, Qian Zhang 0079, Kai Yang 0010, Runlong Xia, Jingbo Xie
Wirel. Commun. Mob. Comput.1
2019 The Image Annotation Method by Convolutional Features from Intermediate Layer of Deep Learning Based on Internet of Things
abstract
Existing image annotation methods that employ convolutional features of deep learning methods from the Internet of Things (IoT) have a number of limitations, including complex training and high space/time expenses associated with the image annotation procedure. Accordingly, this paper proposes an innovative method in which the visual features of the image are presented by the intermediate layer features of deep learning, while semantic concepts are represented by mean vectors of positive samples. Firstly, the convolutional result is directly output in the form of low-level visual features through the mid-level of the pre-trained deep learning model, with the image being represented by sparse coding in the IoT. Secondly, the positive mean vector method is used to construct visual feature vectors for each text vocabulary item, so that a visual feature vector database is created. Finally, the visual feature vector similarity between the testing image and all text vocabulary is calculated, and the vocabulary with the largest similarity is taken from the IoT as the words used for annotation. Experiments on multiple datasets demonstrate the effectiveness of the proposed method; in terms of F1 score, the proposed method's performance on the Corel5k and IAPR TC-12 datasets is superior to that of MBRM, JEC-AF, JEC-DF, and 2PKNN with end-to-end deep features.
Yuantao Chen, Jiajun Tao, Jin Wang 0001, Zhuofan Liao, Lei Wang 0143
MSN1