EDBT 2026 Demo / reviewers in the wild / expert
Yuantao Chen
dblp:140/1870
· DBLP profile ↗
25ranked-venue papers
15as first author
23since 2021 · last 2025
0000-0003-2277-1765ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 11 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Computer networks · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PUGS: Zero-Shot Physical Understanding with Gaussian SplattingabstractCurrent robotic systems can understand the categories and poses of objects well. But understanding physical properties like mass, friction, and hardness, in the wild, remains challenging. We propose a new method that reconstructs 3D objects using the Gaussian splatting representation and predicts various physical properties in a zero-shot manner. We propose two techniques during the reconstruction phase: a geometryaware regularization loss function to improve the shape quality and a region-aware feature contrastive loss function to promote region affinity. Two other new techniques are designed during inference: a feature-based property propagation module and a volume integration module tailored for the Gaussian representation. Our framework is named as zero-shot physical understanding with Gaussian splatting, or PUGS. PUGS achieves new state-of-the-art results on the standard benchmark of ABO-500 mass prediction. We provide extensive quantitative ablations and qualitative visualization to demonstrate the mechanism of our designs. We show the proposed methodology can help address challenging real-world grasping tasks. Our codes, data, and models are available at https://github.com/EverNorif/PUGS Yinghao Shuai, Yuantao Chen, Zijian Jiang, Nan Wang 0041, Jv Zheng, Jianzhu Ma, Meng Yang 0035, Zhicheng Wang 0022, Wenbo Ding 0001, Hao Zhao 0002 |
ICRA | 3 |
| 2025 | Unifying Appearance Codes and Bilateral Grids for Driving Scene Gaussian SplattingabstractNeural rendering techniques, including NeRF and Gaussian Splatting (GS), rely on photometric consistency to produce high-quality reconstructions. However, in real-world driving scenarios, it is challenging to guarantee perfect photometric consistency in acquired images. Appearance codes have been widely used to address this issue, but their modeling capability is limited, as a single code is applied to the entire image. Recently, the bilateral grid was introduced to perform pixel-wise color mapping, but it is difficult to optimize and constrain effectively. In this paper, we propose a novel multi-scale bilateral grid that unifies appearance codes and bilateral grids. We demonstrate that this approach significantly improves geometric accuracy in dynamic, decoupled autonomous driving scene reconstruction, outperforming both appearance codes and bilateral grids. This is crucial for autonomous driving, where accurate geometry is important for obstacle avoidance and control. Our method shows strong results across four datasets: Waymo, NuScenes, Argoverse, and PandaSet. We further demonstrate that the improvement in geometry is driven by the multi-scale bilateral grid, which effectively reduces floaters caused by photometric inconsistency. Nan Wang 0041, Lixing Xiao, Yuantao Chen, Weiqing Xiao, Pierre Merriaux, Ziyang Yan, Saining Zhang, Shaocong Xu, Chongjie Ye, Bohan Li 0015, Zhaoxi Chen 0009, Tianfan Xue, Hao Zhao 0002 |
NeurIPS | 3 |
| 2025 | Self-Aligning Depth-Regularized Radiance Fields for Asynchronous RGB-D SequencesabstractIt has been shown that learning radiance fields with depth rendering and depth supervision can effectively promote the quality and convergence of view synthesis. However, this paradigm requires input RGB-D sequences to be synchronized. In the UAV city modeling scenario, there exists asynchrony between RGB images and depth images due to the different frequencies of the solid-state LiDAR and RGB sensors. To synthesize high-quality views in such a scenario, we propose a novel time-pose function, which is an implicit network that maps timestamps to SE(3) elements. To train this function, we also design a joint optimization scheme to jointly learn the large-scale depth-regularized radiance fields and the time-pose function. Furthermore, we propose a large synthetic dataset with diverse controlled mismatches and ground truth to evaluate this new problem setting systematically. The proposed approach has been evaluated on both datasets and in a real drone. To evaluate the impact of view density, each algorithm was test on three different trajectories with different view densities. Compared to state-of-the-art baseline methods, the proposed approach reduces reconstruction error by 35.26% in city modeling scenarios. Our code is available at github.com/saythe17/AsyncNeRF. Andong Yang, Yuantao Chen, Runyi Yang, Zhenxin Zhu, Hao Zhao 0002, Guyue Zhou |
WACV | 3 |
| 2025 | ATM-DEN: Image Inpainting via attention transfer module and Decoder-Encoder network
Yuantao Chen |
Signal Process. Image Commun. | 2 |
| 2024 | Drone-assisted Road Gaussian Splatting with Cross-view Uncertainty
Saining Zhang, Baijun Ye, Xiaoxue Chen, Yuantao Chen, Zongzheng Zhang, Yongliang Shi, Hao Zhao 0002 |
BMVC | 4 |
| 2024 | GaussReg: Fast 3D Registration with Gaussian Splatting
Yinglin Xu, Yuantao Chen, Wensen Feng, Xiaoguang Han 0001 |
ECCV (15) | 4 |
| 2024 | Camera Relocalization in Shadow-free Neural Radiance FieldsabstractCamera relocalization is a crucial problem in computer vision and robotics. Recent advancements in neural radiance fields (NeRFs) have shown promise in synthesizing photo-realistic images. Several works have utilized NeRFs for refining camera poses, but they do not account for lighting changes that can affect scene appearance and shadow regions, causing a degraded pose optimization process. In this paper, we propose a two-staged pipeline that normalizes images with varying lighting and shadow conditions to improve camera relocalization. We implement our scene representation upon a hash-encoded NeRF which significantly boosts up the pose optimization process. To account for the noisy image gradient computing problem in grid-based NeRFs, we further propose a re-devised truncated dynamic low-pass filter (TDLF) and a numerical gradient averaging technique to smoothen the process. Experimental results on several datasets with varying lighting conditions demonstrate that our method achieves state-of-the-art results in camera relocalization under varying lighting conditions. Code and data will be made publicly available. Shiyao Xu, Caiyun Liu 0004, Yuantao Chen, Zhenxin Zhu, Zike Yan, Yongliang Shi, Hao Zhao 0002, Guyue Zhou |
ICRA | 3 |
| 2024 | Blending Distributed NeRFs with Tri-stage Robust Pose OptimizationabstractDue to the limited model capacity, leveraging distributed Neural Radiance Fields (NeRFs) for modeling extensive urban environments has become a necessity. However, current distributed NeRF registration approaches encounter aliasing artifacts, arising from discrepancies in rendering resolutions and suboptimal pose precision. These factors collectively deteriorate the fidelity of pose estimation within NeRF frameworks, resulting in occlusion artifacts during the NeRF blending stage. In this paper, we present a distributed NeRF system with tri-stage pose optimization. In the first stage, precise poses of images are achieved by bundle adjusting Mip-NeRF 360 with a coarse-to-fine strategy. In the second stage, we incorporate the inverting Mip-NeRF 360, coupled with the truncated dynamic low-pass filter, to enable the achievement of robust and precise poses, termed Frame2Model optimization. On top of this, we obtain a coarse transformation between NeRFs in different coordinate systems. In the third stage, we fine-tune the transformation between NeRFs by Model2Model pose optimization. After obtaining precise transformation parameters, we proceed to implement NeRF blending, showcasing superior performance metrics in both real-world and simulation scenarios. Codes and data will be publicly available at https://github.com/boilcy/Distributed-NeRF. Baijun Ye, Caiyun Liu 0004, Xiaoyu Ye, Yuantao Chen, Yuhai Wang, Zike Yan, Yongliang Shi, Hao Zhao 0002, Guyue Zhou |
IROS | 4 |
| 2024 | MFMAM: Image inpainting via multi-scale feature module with attention module
Yuantao Chen, Runlong Xia, Kai Yang 0010, Ke Zou |
Comput. Vis. Image Underst. | 1 |
| 2024 | Image inpainting algorithm based on inference attention module and two-stage network
Yuantao Chen, Runlong Xia, Kai Yang 0010, Ke Zou |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | MICU: Image super-resolution via multi-level information compensation and U-net
Yuantao Chen, Runlong Xia, Kai Yang 0010, Ke Zou |
Expert Syst. Appl. | 1 |
| 2024 | MFFN: image super-resolution via multi-level features fusion network
Yuantao Chen, Runlong Xia, Kai Yang 0010, Ke Zou |
Vis. Comput. | 1 |
| 2023 | LATITUDE: Robotic Global Localization with Truncated Dynamic Low-pass Filter in City-scale NeRFabstractNeural Radiance Fields (NeRFs) have made great success in representing complex 3D scenes with high-resolution details and efficient memory. Nevertheless, current NeRF - based pose estimators have no initial pose prediction and are prone to local optima during optimization. In this paper, we present LATITUDE: Global Localization with Truncated Dynamic Low-pass Filter, which introduces a two-stage localization mechanism in city-scale NeRF. In place recognition stage, we train a regressor through images generated from trained NeRFs, which provides an initial value for global localization. In pose optimization stage, we minimize the residual between the observed image and rendered image by directly optimizing the pose on the tangent plane. To avoid falling into local optimum, we introduce a Truncated Dynamic Low-pass Filter (TDLF) for coarse-to-fine pose registration. We evaluate our method on both synthetic and real-world data and show its potential applications for high-precision navigation in large-scale city scenes. Codes and dataset will be publicly available at https://github.com/jike5/LATITUDE. Zhenxin Zhu, Yuantao Chen, Zirui Wu, Yongliang Shi, Chuxuan Li, Pengfei Li 0007, Hao Zhao 0002, Guyue Zhou |
ICRA | 2 |
| 2023 | FFTI: Image inpainting algorithm via features fusion and two-steps inpainting
Yuantao Chen, Runlong Xia, Ke Zou, Kai Yang 0010 |
J. Vis. Commun. Image Represent. | 1 |
| 2023 | Corrigendum to "FFTI: Image inpainting algorithm via features fusion and two-steps inpainting" [J. Visual Commun. Image Represent. 91 (2023) 103776]
Yuantao Chen, Runlong Xia, Ke Zou, Kai Yang 0010 |
J. Vis. Commun. Image Represent. | 1 |
| 2023 | DGCA: high resolution image inpainting via DR-GAN and contextual attention
Yuantao Chen, Runlong Xia, Kai Yang 0010, Ke Zou |
Multim. Tools Appl. | 1 |
| 2022 | ETBRec: a novel recommendation algorithm combining the double influence of trust relationship and expert users
Zhenchun Duan, Yuantao Chen |
Appl. Intell. | 3 |
| 2022 | Retracted: Multiscale fast correlation filtering tracking algorithm based on a feature fusion modelabstractRetraction: Multiscale fast correlation filtering tracking algorithm based on a feature fusion model Yuantao Chen, Jin Wang, Songjie Liu, Xi Chen, Jie Xiong, Jingbo Xie, Kai Yang, 2021, 33 (15), (https://doi.org/10.1002/cpe.5533) The above article, published online on 23 October 2019 in Wiley Online Library (wileyonlinelibrary.com), has been retracted by agreement between the authors, the journal Editors, David W. Walker, Jinjun Chen, Nitin Auluck and Martin Berzins, and John Wiley and Sons Ltd. The retraction has been agreed due to scientific errors arising from the incorrect use of materials and data, which have led to conclusions that are unreliable. Yuantao Chen, Jin Wang 0001, Songjie Liu, Jingbo Xie, Kai Yang 0010 |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | Image super-resolution reconstruction based on feature map attention mechanism
Yuantao Chen, Linwu Liu, Volachith Phonevilay, Ke Gu 0002, Runlong Xia, Jingbo Xie, Qian Zhang 0079, Kai Yang 0010 |
Appl. Intell. | 1 |
| 2021 | Research on image Inpainting algorithm of improved GAN based on two-discriminations networks
Yuantao Chen, Haopeng Zhang 0010, Linwu Liu, Qian Zhang 0079, Kai Yang 0010, Runlong Xia, Jingbo Xie |
Appl. Intell. | 1 |
| 2021 | The image annotation algorithm using convolutional features from intermediate layer of deep learning
Yuantao Chen, Linwu Liu, Jiajun Tao, Runlong Xia, Qian Zhang 0079, Kai Yang 0010, Jingbo Xie |
Multim. Tools Appl. | 1 |
| 2021 | The face image super-resolution algorithm based on combined representation learning
Yuantao Chen, Volachith Phonevilay, Jiajun Tao, Runlong Xia, Qian Zhang 0079, Kai Yang 0010, Jingbo Xie |
Multim. Tools Appl. | 1 |
| 2021 | The improved image inpainting algorithm via encoder and similarity constraint
Yuantao Chen, Linwu Liu, Jiajun Tao, Runlong Xia, Qian Zhang 0079, Kai Yang 0010 |
Vis. Comput. | 1 |
| 2020 | Saliency Detection via the Improved Hierarchical Principal Component Analysis MethodabstractAiming at the problems of intensive background noise, low accuracy, and high computational complexity of the current significant object detection methods, the visual saliency detection algorithm based on Hierarchical Principal Component Analysis (HPCA) has been proposed in the paper. Firstly, the original RGB image has been converted to a grayscale image, and the original grayscale image has been divided into eight layers by the bit surface stratification technique. Each image layer contains significant object information matching the layer image features. Secondly, taking the color structure of the original image as the reference image, the grayscale image is reassigned by the grayscale color conversion method, so that the layered image not only reflects the original structural features but also effectively preserves the color feature of the original image. Thirdly, the Principal Component Analysis (PCA) has been performed on the layered image to obtain the structural difference characteristics and color difference characteristics of each layer of the image in the principal component direction. Fourthly, two features are integrated to get the saliency map with high robustness and to further refine our results; the known priors have been incorporated on image organization, which can place the subject of the photograph near the center of the image. Finally, the entropy calculation has been used to determine the optimal image from the layered saliency map; the optimal map has the least background information and most prominently saliency objects than others. The object detection results of the proposed model are closer to the ground truth and take advantages of performance parameters including precision rate (PRE), recall rate (REC), and F -measure (FME). The HPCA model’s conclusion can obviously reduce the interference of redundant information and effectively separate the saliency object from the background. At the same time, it had more improved detection accuracy than others. Yuantao Chen, Jiajun Tao, Qian Zhang 0079, Kai Yang 0010, Runlong Xia, Jingbo Xie |
Wirel. Commun. Mob. Comput. | 1 |
| 2019 | The Image Annotation Method by Convolutional Features from Intermediate Layer of Deep Learning Based on Internet of ThingsabstractExisting image annotation methods that employ convolutional features of deep learning methods from the Internet of Things (IoT) have a number of limitations, including complex training and high space/time expenses associated with the image annotation procedure. Accordingly, this paper proposes an innovative method in which the visual features of the image are presented by the intermediate layer features of deep learning, while semantic concepts are represented by mean vectors of positive samples. Firstly, the convolutional result is directly output in the form of low-level visual features through the mid-level of the pre-trained deep learning model, with the image being represented by sparse coding in the IoT. Secondly, the positive mean vector method is used to construct visual feature vectors for each text vocabulary item, so that a visual feature vector database is created. Finally, the visual feature vector similarity between the testing image and all text vocabulary is calculated, and the vocabulary with the largest similarity is taken from the IoT as the words used for annotation. Experiments on multiple datasets demonstrate the effectiveness of the proposed method; in terms of F1 score, the proposed method's performance on the Corel5k and IAPR TC-12 datasets is superior to that of MBRM, JEC-AF, JEC-DF, and 2PKNN with end-to-end deep features. Yuantao Chen, Jiajun Tao, Jin Wang 0001, Zhuofan Liao, Lei Wang 0143 |
MSN | 1 |