EDBT 2026 Demo / reviewers in the wild / expert
Ting Han 0001
dblp:22/5110-1
· DBLP profile ↗
18ranked-venue papers
4as first author
18since 2021 · last 2026
0009-0002-9474-8337ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SGS-3D: High-Fidelity 3D Instance Segmentation via Reliable Semantic Mask Splitting and GrowingabstractAccurate 3D instance segmentation is crucial for high-quality scene understanding in the 3D vision domain. However, 3D instance segmentation based on 2D-to-3D lifting approaches struggle to produce precise instance-level segmentation, due to accumulated errors introduced during the lifting process from ambiguous semantic guidance and insufficient depth constraints. To tackle these challenges, we propose Splitting and Growing reliable Semantic mask for high-fidelity 3D instance segmentation (SGS-3D), a novel "split-then-grow" framework that first purifies and splits ambiguous lifted masks using geometric primitives, and then grows them into complete instances within the scene. Unlike existing approaches that directly rely on raw lifted masks and sacrifice segmentation accuracy, SGS-3D serves as a training-free refinement method that jointly fuses semantic and geometric information, enabling effective cooperation between the two levels of representation. Specifically, for semantic guidance, we introduce a mask filtering strategy that leverages the co-occurrence of 3D geometry primitives to identify and remove ambiguous masks, thereby ensuring more reliable semantic consistency with the 3D object instances. For the geometric refinement, we construct fine-grained object instances by exploiting both spatial continuity and high-level features, particularly in the case of semantic ambiguity between distinct objects. Experimental results on ScanNet200, ScanNet++, and KITTI-360 demonstrate that SGS-3D substantially improves segmentation accuracy and robustness against inaccurate masks from pre-trained models, yielding high-fidelity object instances while maintaining strong generalization across diverse indoor and outdoor environments. Chaolei Wang, Yang Luo 0002, Siyu Chen 0004, Yiping Chen 0002, Ting Han 0001 |
AAAI | 6 |
| 2026 | LiDAR-DHMT: LiDAR-Adaptive Dual Hierarchical Mask Transformer for Robust Freespace Detection and Semantic SegmentationabstractInaccurate freespace detection remains a significant challenge to the safety of autonomous driving. However, we observe that current multisource fusion approaches rely on converting LiDAR point clouds into depth maps, often lose crucial 3D geometric cues. This compromises the spatial consistency of predictions, especially in complex urban scenes. To address this limitation, we propose LiDAR-DHMT (LiDAR-Adaptive Dual-branch Hierarchical Mask Transformer), a novel framework designed for spatial-consistent freespace detection and semantic segmentation. Our key innovation lies in the introduction of a 3D Relative Position Bias module, which effectively captures LiDAR’s inherent spatial priors. This is coupled with a Dynamic Bias Attention mechanism that adaptively incorporates the 3D positional cues into the Transformer’s attention computation, enhancing spatial coherence. Additionally, we employ a Mask Interaction module and a global-local fusion strategy to jointly model contextual semantics and fine-grained structural details. Extensive experiments conducted on the KITTI Road, KITTI-360, Cityscapes datasets demonstrate that LiDAR-DHMT consistently outperforms existing state-of-the-art methods, achieving a competitive 97.59% F1 score in freespace detection and 69.45% and 84.4% mIoU in semantic segmentation. Our findings suggest that LiDAR-DHMT offers a practical solution for deploying robust freespace perception in complex urban driving environments. Siyu Chen 0004, Ting Han 0001, Changshe Zhang, Huan Chen 0025, Meiliu Wu, Guo-Rong Cai, Jinhe Su |
WACV | 2 |
| 2026 | Structural Safety Condition Prediction of rural houses based on Bidirectional Encoder Representations from Transformers and Multimodal Feature Fusion model
Yifei Jiang, Xiaofei Wei, Ting Han 0001, Guiwen Liu |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | An efficient single-branch network for semantic scene completion of building spacesabstractSemantic scene completion (SSC) aims to reconstruct a complete 3D scene from a sparse point cloud while simultaneously predicting semantic labels. However, existing SSC methods still struggle with two major types of completion errors. First, when the geometric completion branch lacks an immediate understanding of semantic transitions at object boundaries, it tends to prioritize surface continuity or smoothness, leading to erroneous overfilling. Second, in severely occluded regions, insufficient understanding of scene context and object semantics often leads to implausible or incomplete reconstructions. To address these challenges, we propose Efficient-SSC , an efficient framework composed of three complementary modules. First, we introduce the Structural Semantic Joint ( SSJ ) module, which implicitly fuses structural cues and semantic features through a Structural Semantic Attention mechanism, thereby preserving geometric consistency and reducing superfluous filling in regions with clear boundaries. Second, the Hierarchical Region Structure Extraction ( HRSE ) module captures scene layout and structural relationships through multi-scale graph convolution to provide semantic priors for scene completion, thereby improving completion integrity. Finally, we develop the Awareness Offset Estimation ( AOE ) module, which jointly refines geometric structure and semantic information through a heterogeneous densification strategy, enabling predictive refinement and improving the fidelity of fine-grained scene details. Our network adopts a single-branch architecture for semantic scene completion. The proposed approach is computationally efficient and achieves competitive performance on widely used benchmark datasets, including SSC-PC and NYUCAD-PC. For example, on the SSC-PC dataset, compared with previous single-modality methods, it improves mIoU and mAcc by 2.54% and 1.82% , respectively, while reducing the parameter count by 28% . Duxin Zhu, Jinhe Su, Zhaohong Huang, Ting Han 0001, Jiasheng Su, Guo-Rong Cai |
Neurocomputing | 4 |
| 2026 | A Policy-Driven Black-Box Adversarial Example With Location Optimization Against 3D Object DetectionabstractAdversarial attack strategies for 3D object detection have highlighted the critical importance of addressing security concerns in this domain. However, white-box methods require full access to the victim model in large-scale point cloud applications. To this end, we propose a novel Policy-Driven Black-box Attack (BAT) that is designed to optimize attack locations without necessitating detailed knowledge of the victim models. First, we introduce a density-aware pattern generator that creates scene-adaptive attack clusters. Second, we leverage the deep deterministic policy gradient in deep reinforcement learning to train an attack agent capable of targeting the victim model. Ultimately, the attack agent is iteratively directed towards optimal attack locations through the joint application of critic loss and actor loss. To the best of our knowledge, this represents the first reinforcement learning-based black-box attack applied to practical 3D object detection. Experimental results on the KITTI, nuScenes, and Waymo datasets demonstrate that BAT effectively diminishes the accuracy of notable models. Importantly, BAT significantly enhances the attack success rate (surpassing state-of-the-art both white-box and black-box methods) and increases transferability (by 20 times) through simple deep deterministic policy gradient, thus establishing a new baseline for adversarial attacks in 3D object detection. Ting Han 0001, Xiaobin Wu, Chaolei Wang, Huan Luo 0001, Xiaochun Cao, Li Liu 0002, Yiping Chen 0002 |
IEEE Trans. Image Process. | 1 |
| 2025 | Scene4U: Hierarchical Layered 3D Scene Reconstruction from Single Panoramic Image for Your Immerse ExplorationabstractThe reconstruction of immersive and realistic 3D scenes holds significant practical importance in various fields of computer vision and computer graphics. Typically, immersive and realistic scenes should be free from obstructions by dynamic objects, maintain global texture consistency, and allow for unrestricted exploration. The current mainstream methods for image-driven scene construction involves iteratively refining the initial image using a moving virtual camera to generate the scene. However, previous methods struggle with visual discontinuities due to global texture inconsistencies under varying camera poses, and they frequently exhibit scene voids caused by foreground-background occlusions. To this end, we propose a novel layered 3D scene reconstruction framework from panoramic image, named Scene4U. Specifically, Scene4U integrates an open-vocabulary segmentation model with a large language model to decompose a real panorama into multiple layers. Then, we employs a layered repair module based on diffusion model to restore occluded regions using visual cues and depth information, generating a hierarchical representation of the scene. The multi-layer panorama is then initialized as a 3D Gaussian Splatting representation, followed by layered optimization, which ultimately produces an immersive 3D scene with semantic and structural consistency that supports free exploration. Scene4U outperforms state-of-the-art method, improving by 24.24% in LPIPS and 24.40% in BRISQUE, while also achieving the fastest training speed. Additionally, to demonstrate the robustness of Scene4U and allow users to experience immersive scenes from various landmarks, we build WorldVista3D dataset for 3D scene reconstruction, which contains panoramic images of globally renowned sites. The implementation code and dataset will be made publicly available. Junyan Ye, Lihan Jiang, Yiping Chen 0002, Ting Han 0001 |
CVPR | 7 |
| 2025 | VoxelFlow: 2D Semantic Mask-Guided Voxel Flow for Open-Vocabulary 3D Instance SegmentationabstractOpen-vocabulary 3D instance segmentation (OV3IS) has emerged as a promising field, successfully bridging point cloud data and text descriptions through intermediate image modalities. However, existing approaches often rely on strong 3D clustering priors or a single, isolated 2D mask branch, which typically results in insufficient and inefficient feature interaction between the point cloud and image domains. To address these limitations, we propose VoxelFlow, a novel framework that effectively combines 2D mask semantics with 3D geometric features. Our core insight is to leverage the implicit semantic and spatial information embedded in 2D masks and leverage them to replace rigid 3D priors. Specifically, voxel masks, initially derived from 2D projections, are meticulously refined through a unique two-stage process: growth and cross-frame merging, ensuring precise alignment with true instance masks under both semantic and geometric constraints. The methodology unfolds in three key steps. Firstly, we project 2D masks onto a voxelized 3D scene to establish initial seed voxel masks in 3D. Secondly, these voxel masks are grown based on local 3D geometric features and progressively merged across frames using a dynamic threshold, yielding robust super voxel masks. Finally, these semantically consistent and geometrically aligned super voxel masks are integrated with a visual language model and text embeddings to achieve comprehensive open-vocabulary instance segmentation. Extensive experiments on ScanNet200 and ScanNet++ demonstrate that VoxelFlow not only significantly outperforms robust 3D prior baselines but also achieves superior performance in both class-agnostic and open-vocabulary 3D instance segmentation. Chaolei Wang, Huan Chen 0025, Ting Han 0001, Yiping Chen 0002 |
CW | 4 |
| 2025 | Towards A New Era of Geo-Foundation Models: Expert-Guided Multimodal Alignment and Geospatial Context AwarenessabstractFoundation Models (FMs) have demonstrated significant potential for geospatial analysis. As geospatial big data (e.g., geo-tagged text and images) continues to grow, FMs have evolved into geospatial foundation models (GeoFMs), the FMs capable of geospatial reasoning. However, two critical tasks remain challenging: (1) how to align various geospatial modalities for model training, and (2) how to enhance FMs' geospatial context awareness. To fill these gaps, this work develops CLIP4Geo, a unified GeoFM that integrates satellite imagery, LiDAR point clouds, geo-tagged points of interest (POIs), and textual descriptions into a shared representation space. Specifically, to empower the model with geospatial intelligence, we design a Geospatial Expert Knowledge System that tailors both textual and visual information and encodes them into shared geospatial embeddings. Furthermore, we propose a novel Geospatial Context Awareness Transformer to guide the cross-attention mechanism, which leverages geospatial context position embedding to effectively capture spatial dependencies across modalities. We also introduce CityVerse, a large-scale, spatially aligned, and multimodal benchmark dataset covering diverse urban regions in the U.S. and China, supporting a range of geospatial downstream tasks (e.g., zero-shot classification, semantic segmentation / localization, and cross-modality retrieval and generation). Extensive experiments demonstrate that CLIP4Geo consistently outperforms state-of-the-art FM baselines, highlighting the importance of expert-guided multimodal alignment and geospatial context embeddings. Our work provides a scalable framework and comprehensive benchmark to facilitate the future development of GeoFMs for various geospatial applications. Code and dataset are available at https://github.com/Ting-Devin-Han/CLIP4Geo.git. Ting Han 0001, Huan Chen 0025, Chaolei Wang, Yilan Ren, Meiliu Wu |
SIGSPATIAL/GIS | 1 |
| 2025 | Edge First: Edge-Guided Geometry for Superior 3D Roof Wireframe ReconstructionabstractRoof wireframe reconstruction has shown great success in 3D building reconstruction due to its lightweight nature and straightforward representation. However, previous methods consider all roof points, which result in edge redundancy and omissions. In this paper, we propose a novel and streamlined Edge-guided Geometric wireframe reconstruction framework, named EDGE. We find that points distributed along the roof edges make a significant contribution to the precise geometric structure of wireframe. Therefore, we design an edge point extractor (EPE) to capture the spatial relationship between points and edges, filtering out internal plane points. Moreover, we discover that the previous edge detectors rely solely on corner points, leading to error accumulation. To address this, we present the Hybrid Edge Detector (HED) feeding corner points with edge contextual features, which not only enhances edge completeness but also mitigates edge redundancy. Comprehensive experiments demonstrate EDGE outperforms existing wireframe reconstruction methods with 0.86 Corner F1-score and 0.71 Edge F1-score on Building3D dataset, striking the significant improvement of accuracy between corner and edge. Notably, our EDGE achieves a significant improvement of over 11% in Edge Recall, demonstrating the effectiveness and robustness of the proposed method. Qiaoqiao Hao, Ting Han 0001, Yujun Liu 0005, Shangfeng Huang, Duxin Zhu, Jinhe Su, Yun-Dong Wu, Guo-Rong Cai |
ICASSP | 2 |
| 2025 | Stronger, Steadier & Superior: Geometric Consistency in Depth VFM Forges Domain Generalized Semantic Segmentation
Siyu Chen 0004, Ting Han 0001, Changshe Zhang, Meiliu Wu, Guo-Rong Cai, Jinhe Su |
ICCV | 2 |
| 2025 | Depth Matters: Exploring Deep Interactions of RGB-D for Semantic Segmentation in Traffic ScenesabstractRGB-D has gradually become a crucial data source for understanding complex scenes in assisted driving. However, existing studies have paid insufficient attention to the intrinsic spatial properties of depth maps. This oversight significantly impacts the attention representation, leading to prediction errors caused by attention shift issues. To this end, we propose a novel learnable Depth interaction Pyramid Transformer (DiPFormer) to explore the effectiveness of depth. Firstly, we introduce Depth Spatial-Aware Optimization (Depth SAO) as offset to represent real-world spatial relationships. Secondly, the similarity in the feature space of RGB-D is learned by Depth Linear Cross-Attention (Depth LCA) to clarify spatial differences at the pixel level. Finally, an MLP Decoder is utilized to effectively fuse multi-scale features for meeting real-time requirements. Comprehensive experiments demonstrate that the proposed DiPFormer significantly addresses the issue of attention misalignment in both road detection (+7.5%) and semantic segmentation (+4.9% / +1.5%) tasks. DiPFormer achieves state-of-the-art performance on the KITTI (97.57% F-score on KITTI road and 68.74% mIoU on KITTI-360) and Cityscapes (83.4% mIoU) datasets. Siyu Chen 0004, Ting Han 0001, Changshe Zhang, Weiquan Liu, Jinhe Su, Zongyue Wang, Guo-Rong Cai |
IROS | 2 |
| 2025 | Leveraging Depth and Language for Open-Vocabulary Domain-Generalized Semantic SegmentationabstractOpen-Vocabulary semantic segmentation (OVSS) and domain generalization in semantic segmentation (DGSS) highlight a subtle complementarity that motivates Open-Vocabulary Domain-Generalized Semantic Segmentation (OV-DGSS). OV-DGSS aims to generate pixel-level masks for unseen categories while maintaining robustness across unseen domains, a critical capability for real-world scenarios such as autonomous driving in adverse conditions. We introduce Vireo, a novel single-stage framework for OV-DGSS that unifies the strengths of OVSS and DGSS for the first time. Vireo builds upon the frozen Visual Foundation Models (VFMs) and incorporates scene geometry via Depth VFMs to extract domain-invariant structural features. To bridge the gap between visual and textual modalities under domain shift, we propose three key components: (1) GeoText Query, which align geometric features with language cues and progressively refine VFM encoder representations; (2) Coarse Mask Prior Embedding (CMPE) for enhancing gradient flow for faster convergence and stronger textual influence; and (3) the Domain-Open-Vocabulary Vector Embedding Head (DOV-VEH), which fuses refined structural and semantic features for robust prediction. Comprehensive evaluation on these components demonstrates the effectiveness of our designs. Our proposed Vireo achieves the state-of-the-art performance and surpasses existing methods by a large margin in both domain generalization and open-vocabulary recognition, offering a unified and scalable solution for robust visual understanding in diverse and dynamic environments. Code is available at https://github.com/SY-Ch/Vireo. Siyu Chen 0004, Ting Han 0001, Chengzheng Fu, Changshe Zhang, Chaolei Wang, Jinhe Su, Guo-Rong Cai, Meiliu Wu |
NeurIPS | 2 |
| 2025 | Hyperbolic Multi-Criteria Rating RecommendationabstractMulti-criteria (MC) ratings as auxiliary supervisory signals can improve the prediction accuracy of recommender systems. The existing MC methods learn the representations of users and items in Euclidean space to estimate the interaction probabilities. However, this modeling paradigm ignores two important aspects. Firstly, when embedding power-law distribution data and personalized MC preferences in Euclidean space, the model may produce suboptimal solutions due to the distortion of the hierarchical structure. Secondly, the inevitable noise in MC ratings may hinder the recommendation quality of the model. To address the above issues, we propose a novel framework called Hyperbolic Multi-Criteria Recommendation (HMCR), which aims to mine users' MC behavioral features on hyperbolic manifolds and mitigate the noise interference through knowledge transfer among the criteria. Specifically, we map the representations on each criterion view to a hyperbolic space with adjustable curvature based on the Lorentz model, which is used to capture the hierarchical structure of collective user behavior. The MC preferences of individual users are fused by calculating the hyperbolic attention among each criterion and the overall rating. Moreover, we design a self-supervised contrastive loss to suppress the negative impact of noise interactions on the model. The experimental results on four real-world datasets show that the HMCR significantly outperforms the existing baselines. Ting Han 0001, Peng Song 0004, Chenjiao Feng, Kaixuan Yao, Jiye Liang |
SIGIR | 2 |
| 2025 | CityInsight: Incorporating Dual-Condition-Based Diffusion Model Into Building Footprint Segmentation From Remote Sensing ImageryabstractAccurately identifying urban building layouts plays a crucial role in understanding the complexity of urban construction and the level of economic development. Previous footprint segmentation methods have struggled to adapt to the diverse morphology of buildings, limiting the accurate extraction of building footprints from remote sensing imagery and impeding insights and understanding of urban areas. To this end, we propose a framework named CityInsight for analyzing urban building morphology from remote sensing imagery. First, we establish a semantic segmentation network, dual-condition diffusion network (DC-Net), based on a diffusion model to accurately identify building footprints from remote sensing images. Second, we use uncertainty attention and condition attention to generate spatial and semantic priors. Finally, we design a condition injection module to incorporate spatial and semantic information into the diffusion learning. Comprehensive experiments demonstrate the accuracy, robustness, and generalization of the proposed method. The$F_{1}$-scores of the DC-Net on the large-scale remote sensing datasets SpaceNet, WHU Building, Inria, and Massachusetts are 92.05%, 96.59%, 92.17%, and 92.86%, respectively. Furthermore, the footprint segmentation is utilized for subsequent urban function identification and urban analysis of the Zona Oeste of Rio de Janeiro, to emphasize the application value of CityInsight. Our code is available athttps://github.com/Ting-Devin-Han/CityInsight Ting Han 0001, Chaolei Wang, Yang Luo 0002, Hongchao Fan, José Marcato Junior, Xinchang Zhang 0002, Yiping Chen 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | CSFNet: Cross-Modal Semantic Focus Network for Semantic Segmentation of Large-Scale Point CloudsabstractSemantic segmentation of large-scale point clouds is an indispensable component of outdoor scene perception, providing essential 3-D semantic insights for applications in scene reconstruction, urban planning, autonomous driving, and more. However, the discriminative capability of point clouds features declines with increasing distance from the sensor, causing current methods to usually perform poorly in segmenting distant objects. To overcome this challenge and improve the differentiation between classes with similar geometric features, we propose the cross-modal semantic focus network (CSFNet). Firstly, we design a multiscale feature dynamic fusion (MDF) module to leverage multiscale image features, thereby enriching the feature representation of point clouds with additional images color and texture information. Then, in order to extract the distinguishing features of distant and different categories of objects more efficiently, we propose a semantic focus module (SFM) that employs a multiclass contrastive learning strategy to enhance feature discrimination. Finally, we introduce cross-modal knowledge distillation (KD) to augment the model’s comprehension of point clouds. Extensive experiments conducted on the SemanticKITTI and nuScenes datasets demonstrate the effectiveness of our method. Notably, our method achieves superior segmentation accuracy across multiple classes at various distances compared to current methods. Yang Luo 0002, Ting Han 0001, Yujun Liu 0005, Jinhe Su, Yiping Chen 0002, Yun-Dong Wu, Guo-Rong Cai |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | HSPFormer: Hierarchical Spatial Perception Transformer for Semantic SegmentationabstractSemantic perception in driving scenarios plays a crucial role in intelligent transportation systems. However, existing Transformer-based semantic segmentation methods often do not fully exploit their potential in understanding driving scene dynamically. These methods typically lack spatial reasoning, failing to effectively correlate image pixels with their spatial positions, leading to attention drift. To address this issue, we propose a novel architecture, the Hierarchical Spatial Perception Transformer (HSPFormer), which integrates monocular depth estimation and semantic segmentation into a unified framework for the first time. We introduce the Spatial Depth Perception Auxiliary Network (SDPNet), a framework for multiscale feature extraction and multilayer depth map prediction to establish hierarchical spatial coherence. Additionally, we design the Hierarchical Pyramid Transformer Network (HPTNet), which uses depth estimation as learnable position embeddings to form spatially correlated semantic representations and generate global contextual information. Experiments on benchmark datasets such as KITTI-360, Cityscapes, and NYU Depth V2, demonstrate that HSPFormer outperforms several state-of-the-art networks, and achieves promising performance with 66.82% top-1 mIoU on KITTI-360, 83.8% mIoU on Cityscapes, and 57.7% mIoU on NYU Depth V2, respectively. The code will be made publicly available athttps://github.com/SY-Ch/HSPFormer. Siyu Chen 0004, Ting Han 0001, Changshe Zhang, Jinhe Su, Ruisheng Wang 0001, Yiping Chen 0002, Zongyue Wang, Guo-Rong Cai |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Incorporating Building Information into Road Extraction from Remote Sensing ImagesabstractRoad extraction is a critical task in the field of remote sensing, and is widely applied in various tasks. However, significant challenges arise in road extraction due to occlusions from surrounding objects and chaotic backgrounds. Signed Distance Map (SDM) of buildings contains distance information from pixels to building edges, which can serve as auxiliary information for road extraction tasks. To address these challenges, a Road-SDM Fusion Network (RSDMNet) is proposed to explore the latent relationship between roads and buildings to enhance contextual reasoning capabilities. By incorporating the distribution information of buildings, the proposed method provides additional prior features and more global information for road extraction. Experimental results demonstrate that the proposed method achieves a higher IoU and outperforms state-of-the-art approaches. Xiangyi Xie, Ting Han 0001, Yumeng Du, Yiping Chen 0002 |
IGARSS | 2 |
| 2024 | Epurate-Net: Efficient Progressive Uncertainty Refinement Analysis for Traffic Environment Urban Road DetectionabstractHigh-performance and real-time road detection plays an essential role in Advanced Driver Assistance Systems (ADAS) of intelligent transportation. However, existing approaches still suffer from ambiguous road contour in traffic environment because deep learning methods lack explicit constraints on road boundaries with similar textures and structures. To address the unsatisfactory boundaries, we propose an efficient architecture for urban road detection to refine road edges adaptively. First, we design a lightweight symmetrical data-fusion network to merge spatial responses into visual features. Second, we construct cross-layer attention transformation to aggregate non-local contextual information. Moreover, a progressive uncertainty analysis module eliminates indistinct road and obstacle edges. Finally, we introduce upgrade uncertainty loss and improved deep supervision to constrain margin error for multi-scale predictions. Results of experiments using three famous datasets confirm the superiority of our method (F1-measure of 96.91% in KITTI, 98.86% in Cityscapes, and 95.18% in R2D, processing speed of 0.02s) over previous approaches. We demonstrate that, to ensure the safety of autonomous driving, the Epurate-Net adaptively refines road contour to reach exquisite road margins. The source code will be available soon. Ting Han 0001, Siyu Chen 0004, Chuanmu Li, Zongyue Wang, Jinhe Su, Min Huang 0004, Guo-Rong Cai |
IEEE Trans. Intell. Transp. Syst. | 1 |