Wentao Sun

dblp:146/9056 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 BCWildfire: A Long-term Multi-factor Dataset and Deep Learning Benchmark for Boreal Wildfire Risk Prediction
abstract
Wildfire risk prediction remains a critical yet challenging task due to the complex interactions among fuel conditions, meteorology, topography, and human activity. Despite growing interest in data-driven approaches, publicly available benchmark datasets that support long-term temporal modeling, large-scale spatial coverage, and multimodal drivers remain scarce. To address this gap, we present a 25-year, daily-resolution wildfire dataset covering 240 million hectares across British Columbia and surrounding regions. The dataset includes 38 covariates, encompassing active fire detections, weather variables, fuel conditions, terrain features, and anthropogenic factors. Using this benchmark, we evaluate a diverse set of time-series forecasting models, including CNN-based, linear-based, Transformer-based, and Mamba-based architectures. We also investigate effectiveness of position embedding and the relative importance of different fire-driving factors.
Zhengsen Xu, Sibo Cheng, Hongjie He 0003, Wentao Sun, Jonathan Li 0001, Lincoln Linlin Xu
AAAI5
2025 CNSv2: Probabilistic Correspondence Encoded Neural Image Servo
abstract
Visual servo based on traditional image matching methods often requires accurate keypoint correspondence for high precision control. However, keypoint detection or matching tends to fail in challenging scenarios with inconsistent illuminations or textureless objects, resulting significant performance degradation. Previous approaches, including our proposed Correspondence encoded Neural image Servo policy (CNS), attempted to alleviate these issues by integrating neural control strategies. While CNS shows certain improvement against error correspondence over conventional image-based controllers, it could not fully resolve the limitations arising from poor keypoint detection and matching. In this paper, we continue to address this problem and propose a new solution: Probabilistic Correspondence Encoded Neural Image Servo (CNSv2). CNSv2 leverages probabilistic feature matching to improve robustness in challenging scenarios. By redesigning the architecture to condition on multimodal feature matching, CNSv2 achieves high precision, improved robustness across diverse scenes and runs in real-time. We validate CNSv2 with simulations and real-world experiments, demonstrating its effectiveness in overcoming the limitations of detector-based methods in visual servo tasks.
Anzhe Chen, Hongxiang Yu, Zhongxiang Zhou, Wentao Sun, Rong Xiong, Yue Wang 0020
ICRA6
2025 Disambiguate Gripper State in Grasp-Based Tasks: Pseudo-Tactile as Feedback Enables Pure Simulation Learning
abstract
Grasp-based manipulation tasks are fundamental to robots interacting with their environments, yet gripper state ambiguity significantly reduces the robustness of imitation learning policies for these tasks. Data-driven solutions face the challenge of high real-world data costs, while simulation data, despite its low costs, is limited by the sim-to-real gap. We identify the root cause of gripper state ambiguity as the lack of tactile feedback. To address this, we propose a novel approach employing pseudo-tactile as feedback, inspired by the idea of using a force-controlled gripper as a tactile sensor. This method enhances policy robustness without additional data collection and hardware involvement, while providing a noise-free binary gripper state observation for the policy and thus facilitating pure simulation learning to unleash the power of simulation. Experimental results across three real-world grasp-based tasks demonstrate the necessity, effectiveness, and efficiency of our approach. Videos are available on Project Page.
Zherui Song, Yenan Chen, Wentao Sun, Zhongxiang Zhou, Rong Xiong, Yue Wang 0020
IROS5
2025 PCBNet: positional crossing and broad features network for indoor scene semantic segmentation
Huifang Hou, Wentao Sun, Yale Yang, Haipeng Han, Xiaofeng Liu 0006
J. Supercomput.4
2025 Image registration of pointer gauges based on improved reliable and repeatable detector and descriptor algorithm
Wentao Sun, Tiancheng Zhang 0009
J. Supercomput.5
2023 A Click-Based Interactive Segmentation Network for Point Clouds
abstract
Interactive segmentation plays an essential role in several tasks involving point clouds. However, existing methods suffer from low segmentation accuracy and cannot adjust the segmentation results according to the user’s personal demands. This paper presents a novel deep learning-based interactive segmentation method, named Click Rough Segmentation Network (CRSNet), designed to handle point clouds. The method allows users to iteratively click to segment interesting objects. CRSNet consists of two key parts: a CRS module and a feature extraction module. First, the CRS module transforms click operations into an appropriate representation to input into the feature extraction module. The CRS module takes raw point clouds and click operations as input and outputs 3D Gaussian vectors and roughly segmented blocks, which adapt to different-sized and densely-distributed objects in complex environments. Second, the feature extraction module, which uses a novel mix loss-based analysis algorithm, extracts deep features and obtains instance segmentation results. The module is highly compatible because its backbones can be replaced by different deep learning architectures. Experimental results on the KITTI, Apolloscape, Roadmarking, Scannet, and SemanticKITTI datasets show that our method outperforms state-of-the-art semantic segmentation methods with one click. Moreover, our method can generalize well to unseen objects and datasets.
Wentao Sun, Yiping Chen 0002, Huxiong Li, José Marcato Junior, Wesley Nunes Gonçalves, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.1