Guotao Xie

dblp:216/3428 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0003-4617-7349ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
3D vision · 93% Autonomous driving · 7%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing › super-resolution
image super-resolution
1.012026
Domain-Complementary Prior With Fine-Grained Feedback for Scene Text Image Super-Resolution · IEEE Trans. Image Process. 2026
Image and video processing › super-resolution › image super-resolution
scene text image super-resolution
1.012026
Domain-Complementary Prior With Fine-Grained Feedback for Scene Text Image Super-Resolution · IEEE Trans. Image Process. 2026
Computer vision › 3D vision
3d object detection
0.912025
LiDARDustX: A LiDAR Dataset for Dusty Unstructured Road Environments · ICRA 2025
Computer vision › 3D vision
3d scene understanding
0.912025
LiDARDustX: A LiDAR Dataset for Dusty Unstructured Road Environments · ICRA 2025
Computer vision › 3D vision › 3d object detection › point cloud object detection
LiDAR-based 3D object detection
0.912025
LiDARDustX: A LiDAR Dataset for Dusty Unstructured Road Environments · ICRA 2025
Computer vision › 3D vision › point cloud segmentation
point cloud semantic segmentation
0.912025
LiDARDustX: A LiDAR Dataset for Dusty Unstructured Road Environments · ICRA 2025
Robotics › Autonomous driving › perception › perception robustness
perception in adverse weather
0.312025
LiDARDustX: A LiDAR Dataset for Dusty Unstructured Road Environments · ICRA 2025

Methods — techniques the papers use, named apart from their topics

fine-grained feedback mechanism · 1.0domain-complementary prior · 1.0convolutional neural network · 1.0
YearPublicationVenuePosition
2026 Domain-Complementary Prior With Fine-Grained Feedback for Scene Text Image Super-Resolution
abstract
Enhancing the resolution of scene text images is a critical preprocessing step that can substantially improve the accuracy of downstream text recognition in low-quality images. Existing methods primarily rely on auxiliary text features to guide the super-resolution process. However, these features often lack rich low-level information, making them insufficient for faithfully reconstructing both the global structure and fine-grained details of text. Moreover, previous methods often learn suboptimal feature representations from the original low-quality landmark images, which cannot provide precise guidance for super-resolution. In this study, we propose a Fine-Grained Feedback Domain-Complementary Network (FDNet) for scene text image super-resolution. Specifically, we first employ a fine-grained feedback mechanism to selectively refine landmark images, thereby enhancing feature representations. Then, we introduce a novel domain-trace prior interaction generator, which integrates domain-specific traces with a text prior to comprehensively complement the clear edges and structural coverage of the text. Finally, motivated by the limitations of existing datasets, which often exhibit limited scene scales and insufficient challenging scenarios, we introduce a new dataset, MDRText. The proposed dataset, MDRText, features multi-scale and diverse characteristics and is designed to support challenging text image recognition and super-resolution tasks. Extensive experiments on the MDRText and TextZoom datasets demonstrate that our method achieves superior performance in scene text image super-resolution and further improves the accuracy of subsequent recognition tasks.
Yang Li 0093, Pengwen Dai, Guotao Xie
IEEE Trans. Image Process.5
2026 MOAT: Multi-Scale Group Interaction Transformer for Trajectory Prediction in Crowded Scenes
abstract
Predicting pedestrian trajectories in complex environments presents a significant challenge for autonomous driving systems, primarily due to the intricate social interactions among pedestrians and their surrounding groups. Existing methods often struggle to fully capture the impact of group behavior on individual movement. To address these limitations, we propose the Multi-scale Group Interaction Transformer (MOAT) for pedestrian trajectory prediction in high-density, complex scenarios. Our approach introduces a dynamic group interaction module (DGIM) that clusters pedestrians based on proximity, determined by a distance matrix of neighboring pedestrians. In constructing interaction representations, we go beyond traditional features like speed, distance, and direction by incorporating crowd density, thus providing a more comprehensive understanding of group dynamics. To effectively process these features, we employ a multi-branch attention fusion (MBAF) module, which independently analyzes each feature set to capture the unique dynamics and density characteristics of each group and their varying effects on the target pedestrian. These spatial features are then combined with temporal information, allowing our model to account for both spatial and temporal dependencies. Additionally, we leverage a multi-scale Transformer to adaptively partition input trajectories, enhancing the model’s ability to capture dynamic patterns across various scales. Extensive evaluations on benchmark datasets show that our approach consistently outperforms state-of-the-art methods in terms of both prediction accuracy and robustness.
Ming Gao 0012, Yinlin Wang, Guotao Xie, Chi Ding, Yougang Bian
IEEE Trans. Intell. Transp. Syst.4
2025 LiDARDustX: A LiDAR Dataset for Dusty Unstructured Road Environments
abstract
Autonomous driving datasets are essential for validating the progress of intelligent vehicle algorithms, which include localization, perception, and prediction. However, existing datasets are predominantly focused on structured urban environments, which limits the exploration of unstructured and specialized scenarios, particularly those characterized by significant dust levels. This paper introduces the LiDARDustX dataset, which is specifically designed for perception tasks under high-dust conditions, such as those encountered in mining areas. The LiDARDustX dataset consists of 30,000 LiDAR frames captured by six different LiDAR sensors, each accompanied by 3D bounding box annotations and point cloud semantic segmentation. Notably, over 80% of the dataset comprises dust-affected scenes. By utilizing this dataset, we have established a benchmark for evaluating the performance of state-of-the-art 3D detection and segmentation algorithms. Additionally, we have analyzed the impact of dust on perception accuracy and delved into the causes of these effects. The data and further information can be accessed at: https://github.com/vincentweikey/LiDARDustX.
Chenfeng Wei, Si Zuo, Guotao Xie, Shenhong Wang
ICRA7
2024 Dynamic multi-scale spatial-temporal graph convolutional network for traffic flow prediction
abstract
This paper proposes a dynamic multi-scale spatial-temporal graph convolutional network (DS-STGCN) for traffic flow prediction . The network aims to comprehensively extract global and local dependencies in dynamic spatial-temporal data by inputting traffic network flow data to construct node feature graphs, topology graphs , and time slot feature graphs, capturing the complexity and dynamics of traffic flow. DS-STGCN interprets feature information of the traffic network from both spatial and temporal dimensions through dynamic multi-scale graph convolutional blocks. In the spatial dimension, these blocks use constraints at different levels to balance fine-grained local features and extensive global features, revealing the intrinsic structure of traffic flow data. In the temporal dimension, these blocks jointly learn with temporal convolutional blocks to capture multi-frequency time patterns and handle long sequence data, effectively extracting potential dependencies of time series . Furthermore, DS-STGCN effectively models the changing spatial-temporal relationships in road network flow by constructing dynamically adaptive updated adjacency tensors, generating dynamic graph structures to address the challenge of changing spatial-temporal relationships in the transportation system. Experimental results show that our method significantly outperforms other competing methods on five real traffic datasets (PEMS03, PEMS04, PEMS07, PEMS08 and METR-LA).
Ming Gao 0012, Zhuoran Du, Hongmao Qin, Guangyin Jin, Guotao Xie
Knowl. Based Syst.6
2024 PPF-Det: Point-Pixel Fusion for Multi-Modal 3D Object Detection
abstract
Multi-modal fusion can take advantage of the LiDAR and camera to boost the robustness and performance of 3D object detection. However, there are still of great challenges to comprehensively exploit image information and perform accurate diverse feature interaction fusion. In this paper, we proposed a novel multi-modal framework, namely Point-Pixel Fusion for Multi-Modal 3D Object Detection (PPF-Det). The PPF-Det consists of three submodules, Multi Pixel Perception (MPP), Shared Combined Point Feature Encoder (SCPFE), and Point-Voxel-Wise Triple Attention Fusion (PVW-TAF) to address the above problems. Firstly, MPP can make full use of image semantic information to mitigate the problem of resolution mismatch between point cloud and image. In addition, we proposed SCPFE to preliminary extract point cloud features and point-pixel features simultaneously reducing time-consuming on 3D space. Lastly, we proposed a fine alignment fusion strategy PVW-TAF to generate multi-level voxel-fused features based on attention mechanism. Extensive experiments on KITTI benchmarks, conducted on September 24, 2023, demonstrate that our method shows excellent performance.
Guotao Xie, Ming Gao 0012, Manjiang Hu, Xiaohui Qin 0001
IEEE Trans. Intell. Transp. Syst.1
2022 Situational Assessment for Intelligent Vehicles Based on Stochastic Model and Gaussian Distributions in Typical Traffic Scenarios
abstract
In intelligent driving, situational assessment (SA) is an important technology, which helps to improve the cognitive ability of intelligent vehicles in the environment. Uncertainty analysis is very significant in situation assessment. This article proposes an SA method based on uncertainty risk analysis. Under uncertain conditions, according to the random environment model and Gaussian distribution model, the collision probability between multiple vehicles is estimated by comprehensive trajectory prediction. The proposed method considers collision probabilities of different prediction points within and outside the prediction range and obtains long-term accurate prediction results. The method is suitable for the situation risk assessment of sensor systems in the presence of unexpected dynamic obstacles, sensor failures or communication losses in traffic, and different environmental sensing accuracy. The experimental results show that in the dynamic traffic environment, the proposed scenario assessment method can not only accurately predict and assess the situation risks within the prediction range, but also provide accurate scenario risk assessment outside the prediction range.
Hongbo Gao 0001, Juping Zhu, Tong Zhang 0015, Guotao Xie, Zhen Kan, Zhengyuan Hao, Kang Liu 0023
IEEE Trans. Syst. Man Cybern. Syst.4