VLDB 2026 Research / reviewers in the wild / expert
Wanlong Li
dblp:118/4591
· DBLP profile ↗
18ranked-venue papers
0as first author
15since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 11 since 2021Systems, architecture and hardware · 12 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Semi-Supervised Learning for Visual Bird's Eye View Semantic SegmentationabstractVisual bird’s eye view (BEV) semantic segmentation helps autonomous vehicles understand the surrounding environment only from front-view (FV) images, including static elements (e.g., roads) and dynamic elements (e.g., vehicles, pedestrians). However, the high cost of annotation procedures of full-supervised methods limits the capability of the visual BEV semantic segmentation, which usually needs HD maps, 3D object bounding boxes, and camera extrinsic matrixes. In this paper, we present a novel semi-supervised framework for visual BEV semantic segmentation to boost performance by exploiting unlabeled images during the training. A consistency loss that makes full use of unlabeled data is then proposed to constrain the model on not only semantic prediction but also the BEV feature. Furthermore, we propose a novel and effective data augmentation method named conjoint rotation which reasonably augments the dataset while maintaining the geometric relationship between the FV images and the BEV semantic segmentation. Extensive experiments on the nuScenes dataset show that our semi-supervised framework can effectively improve prediction accuracy. To the best of our knowledge, this is the first work that explores improving visual BEV semantic segmentation performance using unlabeled data. The code is available at https://github.com/Junyu-Z/Semi-BEVseg. Junyu Zhu, Lina Liu 0010, Wanlong Li, Yong Liu 0007 |
ICRA | 5 |
| 2023 | FG-Depth: Flow-Guided Unsupervised Monocular Depth EstimationabstractThe great potential of unsupervised monocular depth estimation has been demonstrated by many works due to low annotation cost and impressive accuracy comparable to supervised methods. To further improve the performance, recent works mainly focus on designing more complex network structures and exploiting extra supervised information, e.g., semantic segmentation. These methods optimize the models by exploiting the reconstructed relationship between the target and reference images in varying degrees. However, previous methods prove that this image reconstruction optimization is prone to get trapped in local minima. In this paper, our core idea is to guide the optimization with prior knowledge from pretrained Flow-Net. And we show that the bottleneck of unsupervised monocular depth estimation can be broken with our simple but effective framework named FG-Depth. In particular, we propose (i) a flow distillation loss to replace the typical photometric loss that limits the capacity of the model and (ii) a prior flow based mask to remove invalid pixels that bring the noise in training loss. Extensive experiments demonstrate the effectiveness of each component, and our approach achieves state-of-the-art results on both KITTI and NYU-Depth-v2 datasets. Junyu Zhu, Lina Liu 0010, Yong Liu 0007, Wanlong Li, Hongbo Zhang 0004 |
ICRA | 4 |
| 2023 | Self-Supervised Event-Based Monocular Depth Estimation Using Cross-Modal ConsistencyabstractAn event camera is a novel vision sensor that can capture per-pixel brightness changes and output a stream of asynchronous “events”. It has advantages over conventional cameras in those scenes with high-speed motions and challenging lighting conditions because of the high temporal resolution, high dynamic range, low bandwidth, low power consumption, and no motion blur. Therefore, several supervised monocular depth estimation from events is proposed to address scenes difficult for conventional cameras. However, depth annotation is costly and time-consuming. In this paper, to lower the annotation cost, we propose a self-supervised event-based monocular depth estimation framework named EMoDepth. EMoDepth constrains the training process using the cross-modal consistency from intensity frames that are aligned with events in the pixel coordinate. Moreover, in inference, only events are used for monocular depth prediction. Additionally, we design a multi-scale skip-connection architecture to effectively fuse features for depth estimation while maintaining high inference speed. Experiments on MVSEC and DSEC datasets demonstrate that our contributions are effective and that the accuracy can outperform existing supervised event-based and unsupervised frame-based methods. Junyu Zhu, Lina Liu 0010, Bofeng Jiang, Hongbo Zhang 0004, Wanlong Li, Yong Liu 0007 |
IROS | 6 |
| 2023 | Relative order constraint for monocular depth estimation
Chunpu Liu, Wangmeng Zuo, Guanglei Yang, Wanlong Li, Hongbo Zhang 0004, Tianyi Zang |
Appl. Intell. | 4 |
| 2023 | A Mapping Approach for Eucalyptus Plantations Canopy and Single Tree Using High-Resolution Satellite Images in Liuzhou, ChinaabstractAccurate canopy and single-tree mapping is important to obtain information on the ecological structure and biogeophysical parameters for forests. Although some airborne radars can retrieve canopy and single-tree information within a smaller area, the optical satellite imagery-based approaches for rapidly and accurately mapping them over a large region are still limited. In this study, based onEucalyptuscanopy and single-tree texture and spectral features, we proposed a mapping approach using the combinations of image morphology, the Otsu method, and an adaptive iterative erosion algorithm (EUMAP). Then, we applied the commonly used red/green/blue bands from the high-resolution satellite images, which are freely available, to map the canopy and single-tree inEucalyptusplantations in southern China. EUMAP consists of two steps: (i)Eucalyptuscanopy identification for various canopy density regions; (ii) adaptive iterative erosion to separate single-tree. Our study was conducted in the Chengzhong and Liubei districts of Liuzhou city, China. The accuracy evaluation was carried out in the state-owned Sanmenjiang Forest Farm. The results showed that the average F1 score for mapping canopy and single-tree reached 88.34% and 86.40%, respectively. For the whole study area, there were 7033021Eucalyptustrees and the average density was 819 trees per hectare. The approach adopted in this study, combining the prior knowledges about image morphology and single-tree texture features ofEucalyptusplantations, was highly efficient for satellite image processing and had excellent applicability to large-scaleEucalyptusplantations mapping. Our study highlights the necessary of prior knowledges for forest mapping using satellite images without requiring a training sample and provides a universal approach of accurately large-scale mapping for specific forest species with common red/green/blue images. Yaoping Cui, Junwu Dong, Wanlong Li, Bailu Liu, Jinwei Dong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | LTSR: Long-term Semantic Relocalization based on HD Map for Autonomous VehiclesabstractHighly accurate and robust relocalization or localization initialization ability is of great importance for autonomous vehicles (AVs). Traditional GNSS-based methods are not reliable enough in occlusion and multipath conditions. In this paper we propose a novel long-term semantic relocalization algorithm based on HD map and semantic features which are compact in representation. Semantic features appear widely on urban roads, and are robust to illumination, weather, view-point and appearance changes. Repeated structures, missed and false detections make data association (DA) highly ambiguous. To this end, a robust semantic feature matching method based on a new local semantic descriptor which encodes the spatial and normal relationship between semantic features is performed. Further, we introduce an accurate, efficient, yet simple outlier removal method which works by assessing the local and global geometric consistencies and temporal consistency of semantic matching pairs. The experimental results on our urban dataset demonstrate that our approach performs better in accuracy and robustness compared with the current state-of-the-art methods. Huayou Wang, Changliang Xue, Wanlong Li, Hongbo Zhang 0004 |
ICRA | 4 |
| 2022 | Dynamic Multi-projection Mapping Based on Parallel Intensity ControlabstractProjection mapping using multiple projectors is promising for spatial augmented reality; however, it is difficult to apply it to dynamic scenes. This is because the conventional method decides all pixel intensities of multiple images simultaneously based on the global optimization method, and it is hard to reduce the latency from motion to projection. To mitigate this, we propose a novel method of controlling the intensity based on a pixel-parallel calculation for each projector in real-time with low latency. This parallel calculation leverages the insight that the projected pixels from different projectors in overlapping areas can be approximated independently if the pixel is sufficiently small relative to the surface structure. Additionally, our pixel-parallel calculation method allows a distributed system configuration, such that the number of projectors can be increased to form a network for high scalability. We demonstrate a seamless mapping into dynamic scenes at 360 fps with a 9.5-ms latency using ten cameras and four projectors. Takashi Nomoto, Wanlong Li, Hao-Lun Peng, Yoshihiro Watanabe |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | PocoNet: SLAM-oriented 3D LiDAR Point Cloud Online Compression NetworkabstractIn this paper, we present PocoNet: Point cloud Online COmpression NETwork to address the task of SLAM-oriented compression. The aim of this task is to select a compact subset of points with high priority to maintain localization accuracy. The key insight is that points with high priority have similar geometric features in SLAM scenarios. Hence, we tackle this task as point cloud segmentation to capture complex geometric information. We calculate observation counts by matching between maps and point clouds and divide them into different priority levels. Trained by labels annotated with such observation counts, the proposed network could evaluate the point-wise priority. Experiments are conducted by integrating our compression module into an existing SLAM system to evaluate compression ratios and localization performances. Experimental results on two different datasets verify the feasibility and generalization of our approach. Jinhao Cui, Xin Kong, Xuemeng Yang, Xiangrui Zhao, Yong Liu 0007, Wanlong Li, Hongbo Zhang 0004 |
ICRA | 7 |
| 2021 | SA-LOAM: Semantic-aided LiDAR SLAM with Loop ClosureabstractLiDAR-based SLAM system is admittedly more accurate and stable than others, while its loop closure detection is still an open issue. With the development of 3D semantic segmentation for point cloud, semantic information can be obtained conveniently and steadily, essential for high-level intelligence and conductive to SLAM. In this paper, we present a novel semantic-aided LiDAR SLAM with loop closure based on LOAM, named SA-LOAM, which leverages semantics in odometry as well as loop closure detection. Specifically, we propose a semantic-assisted ICP, including semantically matching, downsampling and plane constraint, and integrates a semantic graph-based place recognition method in our loop closure detection module. Benefitting from semantics, we can improve the localization accuracy, detect loop closures effectively, and construct a global consistent semantic map even in large-scale scenes. Extensive experiments on KITTI and Ford Campus dataset show that our system significantly improves baseline performance, has generalization ability to unseen data and achieves competitive results compared with state-of-the-art methods. Lin Li 0091, Xin Kong, Xiangrui Zhao, Wanlong Li, Hongbo Zhang 0004, Yong Liu 0007 |
ICRA | 4 |
| 2021 | SSC: Semantic Scan Context for Large-Scale Place RecognitionabstractPlace recognition gives a SLAM system the ability to correct cumulative errors. Unlike images that contain rich texture features, point clouds are almost pure geometric information which makes place recognition based on point clouds challenging. Existing works usually encode low-level features such as coordinate, normal, reflection intensity, etc., as local or global descriptors to represent scenes. Besides, they often ignore the translation between point clouds when matching descriptors. Different from most existing methods, we explore the use of high-level features, namely semantics, to improve the descriptor’s representation ability. Also, when matching descriptors, we try to correct the translation between point clouds to improve accuracy. Concretely, we propose a novel global descriptor, Semantic Scan Context, which explores semantic information to represent scenes more effectively. We also present a two-step global semantic ICP to obtain the 3D pose (x, y, yaw) used to align the point cloud to improve matching performance. Our experiments on the KITTI dataset show that our approach outperforms the state-of-the- art methods with a large margin. Our code is available at: https://github.com/lilin-hitcrt/SSC. Lin Li 0091, Xin Kong, Xiangrui Zhao, Tianxin Huang, Wanlong Li, Hongbo Zhang 0004, Yong Liu 0007 |
IROS | 5 |
| 2021 | Semantic Segmentation-assisted Scene Completion for LiDAR Point CloudsabstractOutdoor scene completion is a challenging issue in 3D scene understanding, which plays an important role in intelligent robotics and autonomous driving. Due to the sparsity of LiDAR acquisition, it is far more complex for 3D scene completion and semantic segmentation. Since semantic features can provide constraints and semantic priors for completion tasks, the relationship between them is worth exploring. Therefore, we propose an end-to-end semantic segmentation-assisted scene completion network, including a 2D completion branch and a 3D semantic segmentation branch. Specifically, the network takes a raw point cloud as input, and merges the features from the segmentation branch into the completion branch hierarchically to provide semantic information. By adopting BEV representation and 3D sparse convolution, we can benefit from the lower operand while maintaining effective expression. Besides, the decoder of the segmentation branch is used as an auxiliary, which can be discarded in the inference stage to save computational consumption. Extensive experiments demonstrate that our method achieves competitive performance on SemanticKITTI dataset with low latency. Code and models will be released at https://github.com/jokester-zzz/SSA-SC. Xuemeng Yang, Xin Kong, Tianxin Huang, Yong Liu 0007, Wanlong Li, Hongbo Zhang 0004 |
IROS | 6 |
| 2021 | Up-to-Down Network: Fusing Multi-Scale Context for 3D Semantic Scene CompletionabstractAn efficient 3D scene perception algorithm is a vital component for autonomous driving and robotics systems. In this paper, we focus on semantic scene completion, which is a task of jointly estimating the volumetric occupancy and semantic labels of objects. Since the real-world data is sparse and occluded, this is an extremely challenging task. We propose a novel framework, named Up-to-Down network (UDNet), to achieve the large-scale semantic scene completion with an encoder-decoder architecture for voxel grids. The novel up-to-down block can effectively aggregate multi-scale context information to improve labeling coherence, and the atrous spatial pyramid pooling module is leveraged to expand the receptive field while preserving detailed geometric information. Besides, the proposed multi-scale fusion mechanism efficiently aggregates global background information and improves the semantic completion accuracy. Moreover, to further satisfy the needs of different tasks, our UDNet can accomplish the multi-resolution semantic completion, achieving faster but coarser completion. Detailed experiments in the semantic scene completion benchmark of SemanticKITTI illustrate that our proposed framework surpasses the state-of-the-art methods with remarkable margins and a real-time inference speed by using only voxel grids as input. Xuemeng Yang, Tianxin Huang, Chujuan Zhang, Yong Liu 0007, Wanlong Li, Hongbo Zhang 0004 |
IROS | 6 |
| 2021 | PointSiamRCNN: Target-aware Voxel-based Siamese Tracker for Point CloudsabstractCurrently, there have been many kinds of pointbased 3D trackers, while voxel-based methods are still underexplored. In this paper, we first propose a voxel-based tracker, named PointSiamRCNN, improving tracking performance by embedding target information into the search region. Our framework is composed of two parts for achieving proposal generation and proposal refinement, which fully releases the potential of the two-stage object tracking. Specifically, it takes advantage of efficient feature learning of the voxel-based Siamese network and high-quality proposal generation of the Siamese region proposal network head. In the search region, the groundtruth annotations are utilized to realize semantic segmentation, which leads to more discriminative feature learning with pointwise supervisions. Furthermore, we propose the Self and Cross Attention Module for embedding target information into the search region. Finally, the multi-scale RoI pooling module is proposed to obtain compact representations from target-aware features for proposal refinement. Exhaustive experiments on the KITTI tracking dataset demonstrate that our framework reaches the competitive performance with the state-of-the-art 3D tracking methods and achieves the state-of-the-art in terms of BEV tracking. Chujuan Zhang, Yong Liu 0007, Wanlong Li, Hongbo Zhang 0004 |
IROS | 4 |
| 2021 | Co-graph Attention Reasoning Based Imaging and Clinical Features Integration for Lymph Node Metastasis Prediction
Hui Cui 0002, Ping Xuan, Qiangguo Jin, Mingjun Ding, Butuo Li, Bing Zou, Yiyue Xu, Bingjie Fan, Wanlong Li, Jinming Yu, Henry Been-Lirn Duh |
MICCAI (5) | 9 |
| 2021 | Predicting Esophageal Fistula Risks Using a Multimodal Self-attention Network
Yulu Guan, Hui Cui 0002, Yiyue Xu, Qiangguo Jin, Tian Feng 0001, Huawei Tu, Ping Xuan, Wanlong Li, Henry Been-Lirn Duh |
MICCAI (5) | 8 |
| 2020 | Semantic Graph Based Place Recognition for 3D Point CloudsabstractDue to the difficulty in generating the effective descriptors which are robust to occlusion and viewpoint changes, place recognition for 3D point cloud remains an open issue. Unlike most of the existing methods that focus on extracting local, global, and statistical features of raw point clouds, our method aims at the semantic level that can be superior in terms of robustness to environmental changes. Inspired by the perspective of humans, who recognize scenes through identifying semantic objects and capturing their relations, this paper presents a novel semantic graph based approach for place recognition. First, we propose a novel semantic graph representation for the point cloud scenes by reserving the semantic and topological information of the raw point cloud. Thus, place recognition is modeled as a graph matching problem. Then we design a fast and effective graph similarity network to compute the similarity. Exhaustive evaluations on the KITTI dataset show that our approach is robust to the occlusion as well as viewpoint changes and outperforms the state-of-the-art methods with a large margin. Our code is available at: https://github.com/kxhit/SG_PR. Xin Kong, Xuemeng Yang, Guangyao Zhai, Xiangrui Zhao, Xianfang Zeng, Mengmeng Wang 0005, Yong Liu 0007, Wanlong Li |
IROS | 8 |
| 2020 | F-Siamese Tracker: A Frustum-based Double Siamese Network for 3D Single Object TrackingabstractThis paper presents F-Siamese Tracker, a novel approach for single object tracking prominently characterized by more robustly integrating 2D and 3D information to reduce redundant search space. A main challenge in 3D single object tracking is how to reduce search space for generating appropriate 3D candidates. Instead of solely relying on 3D proposals, firstly, our method leverages the Siamese network applied on RGB images to produce 2D region proposals which are then extruded into 3D viewing frustums. Besides, we perform an on-line accuracy validation on the 3D frustum to generate refined point cloud searching space, which can be embedded directly into the existing 3D tracking backbone. For efficiency, our approach gains better performance with fewer candidates by reducing search space. In addition, benefited from introducing the online accuracy validation, for occasional cases with strong occlusions or very sparse points, our approach can still achieve high precision, even when the 2D Siamese tracker loses the target. This approach allows us to set a new state-of-the-art in 3D single object tracking by a significant margin on a sparse outdoor dataset (KITTI tracking). Moreover, experiments on 2D single object tracking show that our framework boosts 2D tracking performance as well. Jinhao Cui, Xin Kong, Chujuan Zhang, Yong Liu 0007, Wanlong Li |
IROS | 7 |
| 2020 | Collaborative Learning of Cross-channel Clinical Attention for Radiotherapy-Related Esophageal Fistula Prediction from CT
Hui Cui 0002, Yiyue Xu, Wanlong Li, Henry Been-Lirn Duh |
MICCAI (1) | 3 |