Shen Yan 0002

dblp:51/8939-2 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2025
0000-0002-1415-5113ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 NTR-Gaussian: Nighttime Dynamic Thermal Reconstruction with 4D Gaussian Splatting Based on Thermodynamics
abstract
Thermal infrared imaging enables a non-invasive measurement of the surface temperature of objects with all-weather applicability. Leveraging such techniques for 3D reconstruction can accurately reflect the temperature distribution of a scene, thereby supporting applications such as building monitoring and energy management. However, existing approaches predominantly focus on static 3D reconstruction for a single time period, overlooking the dynamic nature of thermal radiation phenomena, and failing to predict or analyze temperature variations over time. In this paper, we introduce a novel method, termed NTR-Gaussian, grounded in thermodynamics to address the challenge of nighttime dynamic thermal reconstruction using 4D Gaussian Splatting. Specifically, We utilize neural networks to predict thermodynamic parameters, such as emissivity, convective heat transfer coefficient, and heat capacity, etc. By means of integration, we numerically solve the infrared temperature of the scene at each moment during the night, so as to predict the temperature of the nighttime scene more accurately. To further advance research in this domain, we release a comprehensive dataset of dynamic thermal reconstruction spanning four distinct regions. Extensive experiments demonstrate that NTR-Gaussian significantly outperforms comparison methods in thermal reconstruction, achieving a predicted temperature error within 1 degree Celsius. The code is available at https://github.com/NPUCVPG/NTR-Gaussian.
Zeyu Cui, Yu Liu 0008, Maojun Zhang, Shen Yan 0002
CVPR6
2025 LoD-Loc v2: Aerial Visual Localization Over Low Level-of-Detail City Models using Explicit Silhouette Alignment
abstract
We propose a novel method for aerial visual localization over low Level-of-Detail (LoD) city models. Previous wireframe-alignment-based method LoD-Loc has shown promising localization results leveraging LoD models. However, LoD-Loc mainly relies on high-LoD (LoD3 or LoD2) city models, but the majority of available models and those many countries plan to construct nationwide are low-LoD (LoD1). Consequently, enabling localization on low-LoD city models could unlock drones' potential for global urban localization. To address these issues, we introduce LoD-Loc v2, which employs a coarse-to-fine strategy using explicit silhouette alignment to achieve accurate localization over low-LoD city models in the air. Specifically, given a query image, LoD-Loc v2 first applies a building segmentation network to shape building silhouettes. Then, in the coarse pose selection stage, we construct a pose cost volume by uniformly sampling pose hypotheses around a prior pose to represent the pose probability distribution. Each cost of the volume measures the degree of alignment between the projected and predicted silhouettes. We select the pose with maximum value as the coarse pose. In the fine pose estimation stage, a particle filtering method incorporating a multi-beam tracking approach is used to efficiently explore the hypothesis space and obtain the final pose estimation. To further facilitate research in this field, we release two datasets with LoD1 city models covering 10.7 km , along with real RGB queries and ground-truth pose annotations. Experimental results show that LoD-Loc v2 improves estimation accuracy with high-LoD models and enables localization with low-LoD models for the first time. Moreover, it outperforms state-of-the-art baselines by large margins, even surpassing texture-model-based methods, and broadens the convergence basin to accommodate larger prior errors.
Juelin Zhu, Shuaibang Peng, Hanlin Tan, Yu Liu 0008, Maojun Zhang, Shen Yan 0002
ICCV7
2024 UAVD4L: A Large-Scale Dataset for UAV 6-DoF Localization
abstract
Despite significant progress in global localization of Unmanned Aerial Vehicles (UAVs) in GPS-denied environments, existing methods remain constrained by the availability of datasets. Current datasets often focus on small-scale scenes and lack viewpoint variability, accurate ground truth (GT) pose, and UAV build-in sensor data. To address these limitations, we introduce a large-scale 6-DoF UAV dataset for localization (UAVD4L) and develop a two-stage 6-DoF localization pipeline (UAVLoc), which consists of offline synthetic data generation and online visual localization. Additionally, based on the 6DoF estimator, we design a hierarchical system for tracking ground target in $3 D$ space. Experimental results on the new dataset demonstrate the effectiveness of the proposed approach. Code and dataset are available at https://github.com/RingoWRW/UAVD4L.
Rouwan Wu, Xiaoya Cheng, Juelin Zhu, Maojun Zhang, Shen Yan 0002
3DV6
2024 LoD-Loc: Aerial Visual Localization using LoD 3D Map with Neural Wireframe Alignment
abstract
We propose a new method named LoD-Loc for visual localization in the air. Unlike existing localization algorithms, LoD-Loc does not rely on complex 3D representations and can estimate the pose of an Unmanned Aerial Vehicle (UAV) using a Level-of-Detail (LoD) 3D map. LoD-Loc mainly achieves this goal by aligning the wireframe derived from the LoD projected model with that predicted by the neural network. Specifically, given a coarse pose provided by the UAV sensor, LoD-Loc hierarchically builds a cost volume for uniformly sampled pose hypotheses to describe pose probability distribution and select a pose with maximum probability. Each cost within this volume measures the degree of line alignment between projected and predicted wireframes. LoD-Loc also devises a 6-DoF pose optimization algorithm to refine the previous result with a differentiable Gaussian-Newton method. As no public dataset exists for the studied problem, we collect two datasets with map levels of LoD3.0 and LoD2.0, along with real RGB queries and ground-truth pose annotations. We benchmark our method and demonstrate that LoD-Loc achieves excellent performance, even surpassing current state-of-the-art methods that use textured 3D models for localization. The code and dataset will be made available upon publication.
Juelin Zhu, Shen Yan 0002, Shengyue Zhang, Yu Liu 0008, Maojun Zhang
NeurIPS2
2024 CLaSP: Cross-view 6-DoF localisation assisted by synthetic panorama
abstract
Abstract Despite the impressive progress in visual localisation, 6‐DoF cross‐view localisation is still a challenging task in the computer vision community due to the huge appearance changes. To address this issue, the authors propose the CLaSP, a coarse‐to‐fine framework, which leverages a synthetic panorama to facilitate cross‐view 6‐DoF localisation in a large‐scale scene. The authors first leverage a segmentation map to correct the prior pose, followed by a synthetic panorama on the ground to enable coarse pose estimation combined with a template matching method. The authors finally formulate the refine localisation process as feature matching and pose refinement to obtain the final result. The authors evaluate the performance of the CLaSP and several state‐of‐the‐art baselines on the Airloc dataset, which demonstrates the effectiveness of our proposed framework.
Juelin Zhu, Shen Yan 0002, Xiaoya Cheng, Rouwan Wu, Maojun Zhang
IET Comput. Vis.2
2024 Spectral-Spatial Adversarial Multidomain Synthesis Network for Cross-Scene Hyperspectral Image Classification
abstract
Cross-scene hyperspectral image (HSI) classification has received widespread attention due to its practicality. However, domain adaptation-based cross-scene HSI classification methods are typically tailored for a specific target scene involved in model training and require retraining for new scenes. We instead propose an novel spectral-spatial adversarial multi-domain synthetic network (S2AMSnet) that can be trained on a single source domain (SD) and generalized to unseen domains. S2AMSnet improves the robustness of the model to the unseen domain by expanding the diverse distribution of the SD. Specifically, to spatially and spectrally generate diversified generative domain (GD), the spectral-spatial domain generation network (S2DGN) is designed, and two S2DGNs with the same structure but not shared parameters are enabled to generate diversified GD through two-step min-max strategy. A Multi-domain mixing module is employed to expand the diversity of the GD further and enhance their class-domain semantic consistency information. Additionally, a multi-scale mutual information regularization network is used to constrain the S2DGN so that the intrinsic class semantic information of its generated GD does not deviate from the SD. A Semantic consistency discriminator with spectral-spatial feature extraction capability is utilized to capture class-domain semantic consistency information from diverse GD to obtain cross-domain invariant knowledge. Comparative analysis with eight state-of-the-art transfer learning methods on three real HSI datasets, along with an ablation study, validates the effectiveness of the proposed S2AMSnet in the cross-scene HSI classification task. The codes of this work will be available at https://github.com/daxichen/S2AMSnet.
Xi Chen 0077, Maojun Zhang, Chen Chen 0127, Shen Yan 0002
IEEE Trans. Geosci. Remote. Sens.5
2024 ATLoc: Aerial Thermal Images Localization via View Synthesis
abstract
While visual localization has made significant advances in recent years, it still lacks robustness in low-light situations. Thermal camera images, which capture temperature data, provide a potential solution for these environments. However, the scarcity of well-annotated, publicly available datasets for thermal localization, particularly those focused on absolute pose estimation, impedes further advancement in this field. In this research, we introduce a novel dataset that includes six-degree-of-freedom (6-DoF) absolute poses of query images for large-scale, realistic aerial localization of thermal images. Besides, we introduce a render-to-localization pipeline tailored for thermal image localization. This pipeline predicts the 6-DoF pose of a query using a synthetic technique based on geometric refined thermal model. Experimental results demonstrate the effectiveness of our method on this newly proposed dataset. Notably, our method achieves a median position error of less than 1.5 m and a median angle error of less than 1.5° under diverse test conditions. A comprehensive analysis of factors influencing localization accuracy is also provided. Our code and dataset will be available athttps://github.com/RingoWRW/ATLoc.
Rouwan Wu, Shen Yan 0002, Xiaoya Cheng, Juelin Zhu, Yu Liu 0008, Maojun Zhang
IEEE Trans. Geosci. Remote. Sens.3
2023 Long-Term Visual Localization with Mobile Sensors
abstract
Despite the remarkable advances in image matching and pose estimation, image-based localization of a camera in a temporally-varying outdoor environment is still a challenging problem due to huge appearance disparity between query and reference images caused by illumination, seasonal and structural changes. In this work, we propose to leverage additional sensors on a mobile phone, mainly GPS, compass, and gravity sensor, to solve this challenging problem. We show that these mobile sensors provide decent initial poses and effective constraints to reduce the searching space in image matching and final pose estimation. With the initial pose, we are also able to devise a direct 2D-3D matching network to efficiently establish 2D-3D correspondences instead of tedious 2D-2D matching in existing systems. As no public dataset exists for the studied problem, we collect a new dataset that provides a variety of mobile sensor data and significant scene appearance variations, and develop a system to acquire ground-truth poses for query images. We benchmark our method as well as several state-of-the-art baselines and demonstrate the effectiveness of the proposed approach. Our code and dataset are available on the project page: https://zju3dv.github.io/sensloc/
Shen Yan 0002, Yu Liu 0008, Zehong Shen, Haomin Liu, Maojun Zhang, Guofeng Zhang 0001, Xiaowei Zhou 0001
CVPR1
2023 Deep Active Contours for Real-time 6-DoF Object Tracking
abstract
This paper solves the problem of real-time 6-DoF object tracking from an RGB video. Prior optimization-based methods optimize the object pose by aligning the projected model to the image based on handcrafted features, which are prone to suboptimal solutions. Recent learning-based methods use neural networks to predict the pose, which suffer from limited generalizability or computational efficiency. We propose a learning-based active contour model to make the best use of both worlds. Specifically, given an initial pose, we project the object model to the image plane to obtain the initial contour and use a lightweight network to predict how the contour should move to match the true object boundary, which provides the gradients to optimize the object pose. We also devise an efficient optimization algorithm to train our model end-to-end with pose supervision. Experimental results on semi-synthetic and real-world 6-DoF object tracking datasets demonstrate that our model outperforms state-of-the-art methods by a substantial margin in pose accuracy, while achieving real-time performance on mobile devices. Code is available on our project page: https://zju3dv.github.io/deep_ac/.
Shen Yan 0002, Jianan Zhen, Yu Liu 0008, Maojun Zhang, Guofeng Zhang 0001, Xiaowei Zhou 0001
ICCV2
2023 Render-and-Compare: Cross-view 6-DoF Localization from Noisy Prior
abstract
Despite the significant progress in 6-DoF visual localization, researchers are mostly driven by ground-level benchmarks. Compared with aerial oblique photography capture, ground-level map collection lacks scalability and complete coverage. In this work, we propose to go beyond the traditional ground-level setting and exploit cross-view 6-DoF localization from aerial to ground. We address this problem by formulating camera pose estimation as an iterative render-and-compare pipeline and enhancing the algorithm robustness through augmenting seeds from noisy initial priors. As no public dataset exists for the studied problem, we have collected a new dataset that provides a variety of cross-view images from smartphones and low-altitude drones and developed a semi-automatic system to acquire ground-truth poses for query images. We benchmark our method as well as several state-of-the-art baselines and demonstrate that our method outperforms other approaches by a large margin. Code is available at https://github.com/Choyaa/Render2Loc.
Shen Yan 0002, Xiaoya Cheng, Juelin Zhu, Rouwan Wu, Yu Liu 0008, Maojun Zhang
ICME1
2022 View graph construction for scenes with duplicate structures via graph convolutional network
abstract
Abstract View graph construction aims to effectively organise disordered image dataset through image retrieval technique before structure from motion (SfM). Existing view graph construction methods usually fail to handle scenes with duplicate structure, because these methods solely treat the construction of view graph as a process of image‐pair‐wise matching and lack in exploiting images' topological details in dataset. In this paper, we handle this problem from a novel perspective to construct view graph in a global paradigm by introducing an end‐to‐end graph convolutional network (GCN). First, a location‐aware embedding module is introduced to encode images into a feature space that takes into account the feature's location by using Vision Transformer architecture, improving the distinction between features of duplicate structure. Second, graph convolutional network that consists of topological relationship preserving module and feature metric learning module is proposed. Topological relationship preserving network is proposed to help nodes maintain their connected neighbourhood features. By merging the topological connected information into images' embedding, our method can process image matching in a global mode, thus improving the disambiguation ability for images with duplicate scenes. Then a feature metric learning network is embedded into GCN to dynamically compute the linkage prediction among nodes based on their features. Finally, our method combines these three parts to jointly optimise nodes' features and linkage prediction in an end‐to‐end paradigm. We make qualitative and quantitative comparisons based on three public benchmark datasets and demonstrate that our proposed method performs favourably against other state‐of‐the‐art methods.
Shen Yan 0002, Yu Liu 0008, Maojun Zhang
IET Comput. Vis.2
2021 Image retrieval for Structure-from-Motion via Graph Convolutional Network
abstract
Conventional image retrieval techniques for Structure-from-Motion (SfM) are limited in their ability to effectively distinguish symmetric or repetitive textured patterns and cannot guarantee an accurate generation of pairwise matches without costly redundancy. In this paper, we formulate the image retrieval task as a node binary classification problem with graph data: if a candidate node is marked as positive, it is believed to share the same scene with the query image. The key idea of our approach is that the local context in the feature space around a query image contains abundant information about the matchable relation between the image and its neighbours. By constructing a subgraph surrounding the query image as input data, we adopt a learnable Graph Convolutional Network (GCN) to determine whether nodes in the subgraph have overlapping regions with the query photograph. Experiments demonstrate that our method performs remarkably well on a challenging dataset of highly ambiguous and duplicated scenes. Furthermore, compared with state-of-the-art matchable retrieval methods , the proposed approach significantly reduces unnecessary attempted matches without sacrificing the accuracy and completeness of reconstruction.
Shen Yan 0002, Maojun Zhang, Shiming Lai, Yu Liu 0008
Inf. Sci.1