EDBT 2026 Demo / reviewers in the wild / expert
Weiquan Liu
dblp:03/1188
· DBLP profile ↗
44ranked-venue papers
6as first author
37since 2021 · last 2026
0000-0002-5934-1139ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 2 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OmniEvent: Unified Event Representation LearningabstractEvent cameras have gained increasing popularity in computer vision due to their ultra-high dynamic range and temporal resolution. However, event networks heavily rely on task-specific designs due to the unstructured data distribution and spatial-temporal (S-T) inhomogeneity, making it hard to reuse existing architectures for new tasks. We propose OmniEvent, an innovative unified event representation learning framework that achieves SOTA performance across diverse tasks, fully removing the need for task-specific designs. Unlike previous methods that treat event data as 3D point clouds with manually tuned S-T scaling weights, OmniEvent proposes a decouple-enhance-fuse paradigm, where the local feature aggregation and enhancement are done independently on the spatial and temporal domains to avoid inhomogeneity issues. Space-filling curves are applied to enable large receptive fields while improving memory and compute efficiency. The features from individual domains are then fused by attention to learn S-T interactions. The output of OmniEvent is a grid-shaped tensor, which enables standard vision models to process event data without architectural changes. With a unified framework and similar hyperparameters, OmniEvent outperforms (task-specific) SOTA by up to 68.2% across 3 representative tasks and 10 datasets (Fig. 1). Weiqi Yan 0005, Chenlu Lin, Youbiao Wang, Zhipeng Cai 0003, Xiuhong Lin, Yangyang Shi, Weiquan Liu |
AAAI | 7 |
| 2026 | Physically-Based LiDAR Smoke Simulation for Robust 3D Object Detectionabstract3D object detection in adverse weather is crucial for autonomous driving, especially in smoke where LiDAR data becomes sparse and noisy. Due to the lack of real smoke data, this paper introduces a physics-based simulation framework to generate realistic LiDAR point clouds of smoke and augment large-scale driving datasets. First, we present a 3D fluid dynamics-based smoke simulation framework in Unity, which models the realistic spatial diffusion and temporal evolution of smoke particles. Coupled with a physically accurate LiDAR perception module, our system captures complex light interactions—such as beam attenuation, scattering, and multi-path effects—to generate high-fidelity, physically consistent smoke point clouds. Second, we propose a range image-based data fusion strategy that seamlessly integrates the simulated smoke point clouds into large-scale real-world LiDAR datasets (e.g., Waymo). This approach accurately emulates LiDAR scanning characteristics and naturally incorporates occlusion effects, enabling realistic smoke integration without compromising spatial consistency. To validate our approach, we collect a real-world LiDAR smoke dataset (LiSmoke) and conduct extensive experiments using state-of-the-art 3D detectors. Results demonstrate that models trained with our augmented synthetic data achieve significant improvements in smoke-affected scenarios, while maintaining competitive performance in clear-weather conditions. Our work provides a cost-effective solution for enhancing perception robustness in safety-critical environments. Shijun Zheng, Weiquan Liu, Ming Cheng 0002, Cheng Wang 0003 |
AAAI | 3 |
| 2026 | IO-LIO: Information-Oriented Voxel Mapping for Efficient and Precise LiDAR-Inertial OdometryabstractIn dynamic urban scenarios involving autonomous vehicles and vehicle-infrastructure interactions, LiDAR-inertial odometry (LIO) methods are widely adopted to provide real-time vehicle pose estimation with robust and accurate positioning, particularly in urban environments where GPS signals are frequently disrupted. However, existing methods typically struggle with balancing real-time computational efficiency and localization accuracy, limiting their practical applicability. To address this challenge, we propose IO-LIO, an information-oriented voxel-based LIO system specifically designed for diverse real-world transportation scenarios. IO-LIO enhances real-time performance through concurrent processing and targeted computational optimizations, while efficiently extracting and utilizing high-value voxel information. Specifically, a lookup table-based method is introduced for rapid point cloud undistortion, coupled with an incremental distribution computation approach for information-preserving voxel downsampling. The mapping module adopts an adaptive voxel merging strategy with a multilevel Least Recently Used (M-LRU) mechanism, effectively reducing memory usage. Moreover, we propose a weighted GICP residual construction method, where residuals are weighted by voxel information values, quantitatively improving trajectory accuracy compared to state-of-the-art methods. Additionally, IO-LIO systematically addresses the critical but often overlooked challenge of memory management in LIO systems through a multi-threaded, template-based object pool. Extensive experiments were conducted in realistic ITS-relevant scenarios, including autonomous vehicles operating in dense urban traffic, vehicles navigating campus and park environments, and pedestrian-based backpack mapping. These experiments were complemented by a dedicated ablation study that quantified the incremental benefit of each module. Collectively, the results demonstrate that IO-LIO significantly surpasses state-of-the-art methods in localization accuracy, real-time performance, and reliability, highlighting its strong potential for practical deployment in intelligent transportation applications. Junyuan Lu, Shichun Yi, Weiquan Liu, Yue Wang 0020, Rong Xiong, Yu Zhang 0018 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | A New Adversarial Perspective for LiDAR-based 3D Object DetectionabstractAutonomous vehicles (AVs) rely on LiDAR sensors for environmental perception and decision-making in driving scenarios. However, ensuring the safety and reliability of AVs in complex environments remains a pressing challenge. To address this issue, we introduce a real-world dataset (ROLiD) comprising LiDAR-scanned point clouds of two random objects: water mist and smoke. In this paper, we introduce a novel adversarial perspective by proposing an attack framework that utilizes water mist and smoke to simulate environmental interference. Specifically, we propose a point cloud sequence generation method using a motion and content decomposition generative adversarial network named PCS-GAN to simulate the distribution of random objects. Furthermore, leveraging the simulated LiDAR scanning characteristics implemented with Range Image, we examine the effects of introducing random object perturbations at various positions on the target vehicle. Extensive experiments demonstrate that adversarial perturbations based on random objects effectively deceive vehicle detection and reduce the recognition rate of 3D object detection models. Shijun Zheng, Weiquan Liu, Cheng Wang 0003 |
AAAI | 2 |
| 2025 | Boosting Adversarial Transferability through Augmentation in Hypothesis SpaceabstractAdversarial examples can mislead deep neural networks with subtle perturbations, causing them to make incorrect predictions. Notably, adversarial examples crafted for one model can also deceive other models, a phenomenon known as the transferability of adversarial examples. To improve transferability, existing studies have designed increasingly complex mechanisms, but the improvements achieved remain relatively limited and are often difficult to adapt to other modalities, further restricting the scalability of these methods. In this work, we observe a mirroring relationship between model generalization and adversarial example transferability. Motivated by this observation, we propose an augmentation-based attack, called OPS (OperatorPerturbation-based Stochastic optimization), which constructs a stochastic optimization problem by input transformation operators and random perturbations, and solves this problem to generate adversarial examples with better transferability. Extensive experiments on both images and 3D point clouds demonstrate that OPS significantly outperforms existing state-of-the-art methods in terms of both performance and cost, showcasing the universality and superiority of our approach. The code is available at https://github.com/the-full/OPS. Weiquan Liu, Qingshan Xu 0001, Shijun Zheng, Shujun Huang, Chenglu Wen, Cheng Wang 0003 |
CVPR | 2 |
| 2025 | DoF: A Diffusion Factorization Framework for Offline Multi-Agent Reinforcement LearningabstractDiffusion models have been widely adopted in image and language generation and are now being applied to reinforcement learning. However, the application of diffusion models in offline cooperative Multi-Agent Reinforcement Learning (MARL) remains limited. Although existing studies explore this direction, they suffer from scalability or poor cooperation issues due to the lack of design principles for diffusion-based MARL. The Individual-Global-Max (IGM) principle is a popular design principle for cooperative MARL. By satisfying this principle, MARL algorithms achieve remarkable performance with good scalability. In this work, we extend the IGM principle to the Individual-Global-identically-Distributed (IGD) principle. This principle stipulates that the generated outcome of a multi-agent diffusion model should be identically distributed as the collective outcomes from multiple individual-agent diffusion models. We propose DoF, a diffusion factorization framework for Offline MARL. It uses noise factorization function to factorize a centralized diffusion model into multiple diffusion models. We theoretically show that the noise factorization functions satisfy the IGD principle. Furthermore, DoF uses data factorization function to model the complex relationship among data generated by multiple diffusion models. Through extensive experiments, we demonstrate the effectiveness of DoF. The source code is available at [https://github.com/xmu-rl-3dv/DoF](https://github.com/xmu-rl-3dv/DoF). Ziwei Deng, Chenxing Lin, Yongquan Fu, Weiquan Liu, Chenglu Wen, Cheng Wang 0003 |
ICLR | 6 |
| 2025 | CRS3D: Consistency Regularization for Sparsely-supervised 3D object detectionabstract3D object detection is an indispensable component of autonomous driving. Due to the expensive and labor-intensive annotation required for full supervision, sparsely supervised 3D object detection is emerging as a promising alternative. Although some existing sparse supervision methods have achieved encouraging detection results, they do not perform well with distant or occluded objects. To address this issue, we propose a Consistency Regularization method for Sparsely-supervised 3D object detection(CRS3D). CRS3D consists of three modules: the Point Cloud Adjust module and the Consistency Loss module, which enhance the model’s ability to perceive distant objects by aligning predictions between the original and perturbed point clouds; and the Priori Ratio module, which optimizes the perception of occluded objects by imposing constraints based on priori information. In the KITTI benchmark, CRS3D achieves 87.4% 3D-mAP for the car class at easy difficulty using only 2% of the annotations, surpassing the accuracy of fully supervised methods. Binghui Zeng, Zongyue Wang, Zhaoliang Liu, Yidong Chen 0001, Weiquan Liu |
IJCNN | 5 |
| 2025 | Depth Matters: Exploring Deep Interactions of RGB-D for Semantic Segmentation in Traffic ScenesabstractRGB-D has gradually become a crucial data source for understanding complex scenes in assisted driving. However, existing studies have paid insufficient attention to the intrinsic spatial properties of depth maps. This oversight significantly impacts the attention representation, leading to prediction errors caused by attention shift issues. To this end, we propose a novel learnable Depth interaction Pyramid Transformer (DiPFormer) to explore the effectiveness of depth. Firstly, we introduce Depth Spatial-Aware Optimization (Depth SAO) as offset to represent real-world spatial relationships. Secondly, the similarity in the feature space of RGB-D is learned by Depth Linear Cross-Attention (Depth LCA) to clarify spatial differences at the pixel level. Finally, an MLP Decoder is utilized to effectively fuse multi-scale features for meeting real-time requirements. Comprehensive experiments demonstrate that the proposed DiPFormer significantly addresses the issue of attention misalignment in both road detection (+7.5%) and semantic segmentation (+4.9% / +1.5%) tasks. DiPFormer achieves state-of-the-art performance on the KITTI (97.57% F-score on KITTI road and 68.74% mIoU on KITTI-360) and Cityscapes (83.4% mIoU) datasets. Siyu Chen 0004, Ting Han 0001, Changshe Zhang, Weiquan Liu, Jinhe Su, Zongyue Wang, Guo-Rong Cai |
IROS | 4 |
| 2025 | MSSA-Net: A Multi-Scale Structure-Aware Network for Edge Detection in Point CloudsabstractEdge detection in point clouds is a fundamental problem in 3D vision. Previous approaches based on point-based neural networks treat edge detection as a binary segmentation task, directly applying semantic segmentation frameworks. However, due to the intrinsic differences between edge detection and semantic segmentation, these methods have limitations compared to traditional methods that rely on hand-crafted features, as well as recent volumetric edge representation methods. To address these challenges, we propose three key constraints: Local Structure Disruption, Encoding Region Offset, and Information Loss in Feature Propagation. Based on these constraints, we introduce a novel Multi-Scale Structure-Aware Network (MSSA-Net). MSSA-Net is designed around local multi-scale neighborhoods, learning intact local structural features at different scales within highly relevant neighborhoods for each point. This strategy effectively avoids Local Structure Disruption and Encoding Region Offset. During feature learning, the MSSA Core module generates adaptive encoding branches for different scales of neighborhoods. The learned multi-scale local structural features are then concatenated to encode the local structural information of each point in the point cloud. MSSA-Net eliminates the need for propagating high-level features to low-level features, thereby preventing Information Loss in Feature Propagation. Extensive experiments demonstrate that our MSSA-Net surpasses existing works by a large margin and achieves state-of-the-art performance on various datasets. Yunzhou Xia, Weiqi Yan 0002, Weiquan Liu, Cheng Wang 0003 |
ICMR | 4 |
| 2025 | Interpreting Hidden Semantics in the Intermediate Layers of 3D Point Cloud Classification Neural NetworkabstractAlthough 3D point cloud classification neural network models have been widely used, the in-depth interpretation of the activation of the neurons and layers is still a challenge. We propose a novel approach, named Relevance Flow, to interpret the hidden semantics of 3D point cloud classification neural networks. It delivers the class Relevance to the activated neurons in the intermediate layers in a back-propagation manner, and associates the activation of neurons with the input points to visualize the hidden semantics of each layer. Specially, we reveal that the 3D point cloud classification neural network has learned the plane-level and part-level hidden semantics in the intermediate layers, and utilize the normal and IoU to evaluate the consistency of both levels' hidden semantics. Besides, by using the hidden semantics, we generate the adversarial attack samples to attack 3D point cloud classifiers. Experiments show that our proposed method reveals the hidden semantics of the 3D point cloud classification neural network on ModelNet40 and ShapeNet, which can be used for the unsupervised point cloud part segmentation without labels and attacking the 3D point cloud classifiers. Weiquan Liu, Minghao Liu 0007, Shijun Zheng, Xuesheng Bian, Ping Zhong 0001, Cheng Wang 0003 |
IEEE Trans. Multim. | 1 |
| 2024 | Density-guided Translator Boosts Synthetic-to-Real Unsupervised Domain Adaptive Segmentation of 3D Point Cloudsabstract3D synthetic-to-real unsupervised domain adaptive seg-mentation is crucial to annotating new domains. Self-training is a competitive approach for this task, but its performance is limited by different sensor sampling patterns (i.e., variations in point density) and incomplete training strate-gies. In this work, we propose a density-guided translator (DGT), which translates point density between domains, and integrates it into a two-stage self-training pipeline named DGT-ST. First, in contrast to existing works that simulta-neously conduct data generation and feature/output align-ment within unstable adversarial training, we employ the non-learnable DGT to bridge the domain gap at the in-put level. Second, to provide a well-initialized model for self-training, we propose a category-level adversarial net-work in stage one that utilizes the prototype to prevent neg-ative transfer. Finally, by leveraging the designs above, a domain-mixed self-training method with source-aware consistency loss is proposed in stage two to narrow the domain gap further. Experiments on two synthetic-to-real segmentation tasks (SynLiDAR → semanticKITTI and SynL- iDAR → semanticPOSS) demonstrate that DGT-ST outper-forms state-of-the-art methods, achieving 9.4% and 4.3% mIoU improvements, respectively. Code is available at https://github.com/yuan-zm/DGT-ST. Zhimin Yuan, Wankang Zeng, Yanfei Su, Weiquan Liu, Ming Cheng 0002, Yulan Guo, Cheng Wang 0003 |
CVPR | 4 |
| 2024 | Domain adaptive remote sensing image semantic segmentation with prototype guidance
Wankang Zeng, Ming Cheng 0002, Zhimin Yuan, Youming Wu, Weiquan Liu, Cheng Wang 0003 |
Neurocomputing | 6 |
| 2024 | SR-Adv: Salient Region Adversarial Attacks on 3D Point Clouds for Autonomous DrivingabstractAutonomous driving safety based on LiDAR perception is increasingly becoming a hot spot. Specifically, 3D adversarial examples always make the prediction results of deep neural network models unpredictable, which poses a major security risk to autonomous driving systems. However, the vulnerability of 3D neural network models to adversarial examples is less explored. At present, existing adversarial attack methods often obtain 3D adversarial examples by perturbing the entire point cloud, which requires a large number of perturbed points, i.e. requires a large perturbation budget. In this paper, we propose a Salient Region Adversarial attack method (SR-Adv) to generate adversarial point clouds, by perturbing fewer regions and fewer points. To our knowledge, we are the first to propose region-based attacks for 3D point clouds. First, the proposed SR-Adv employs game theory to extract salient regions of point clouds. This mechanism assigns a value to each region to measure its importance to the 3D neural network model prediction results and realizes the vulnerability analysis of the 3D model. Second, we propose a novel optimization-based gradient attack algorithm to achieve adversarial attacks on salient regions. We evaluate the proposed SR-Adv attack method on the synthetic datasets ModelNet40 and ShapeNetPart as well as the real-world dataset KITTI and NuScenes. Experimental results show that the proposed SR-Adv achieves a state-of-the-art attack success rate and better imperceptibility by perturbing fewer points on 3D point clouds. Shijun Zheng, Weiquan Liu, Chenglu Wen, Ming Cheng 0002, Ping Zhong 0001, Cheng Wang 0003 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | E2PNet: Event to Point Cloud Registration with Spatio-Temporal Representation LearningabstractEvent cameras have emerged as a promising vision sensor in recent years due to their unparalleled temporal resolution and dynamic range. While registration of 2D RGB images to 3D point clouds is a long-standing problem in computer vision, no prior work studies 2D-3D registration for event cameras. To this end, we propose E2PNet, the first learning-based method for event-to-point cloud registration.
The core of E2PNet is a novel feature representation network called Event-Points-to-Tensor (EP2T), which encodes event data into a 2D grid-shaped feature tensor. This grid-shaped feature enables matured RGB-based frameworks to be easily used for event-to-point cloud registration, without changing hyper-parameters and the training procedure. EP2T treats the event input as spatio-temporal point clouds. Unlike standard 3D learning architectures that treat all dimensions of point clouds equally, the novel sampling and information aggregation modules in EP2T are designed to handle the inhomogeneity of the spatial and temporal dimensions. Experiments on the MVSEC and VECtor datasets demonstrate the superiority of E2PNet over hand-crafted and other learning-based methods. Compared to RGB-based registration, E2PNet is more robust to extreme illumination or fast motion due to the use of event data. Beyond 2D-3D registration, we also show the potential of EP2T for other vision tasks such as flow estimation, event-to-image reconstruction and object recognition. The source code can be found at: https://github.com/Xmu-qcj/E2PNet. Xiuhong Lin, Changjie Qiu, Zhipeng Cai 0003, Weiquan Liu, Xuesheng Bian, Matthias Müller 0011, Cheng Wang 0003 |
NeurIPS | 6 |
| 2023 | RiskQ: Risk-sensitive Multi-Agent Reinforcement Learning Value FactorizationabstractMulti-agent systems are characterized by environmental uncertainty, varying policies of agents, and partial observability, which result in significant risks. In the context of Multi-Agent Reinforcement Learning (MARL), learning coordinated and decentralized policies that are sensitive to risk is challenging. To formulate the coordination requirements in risk-sensitive MARL, we introduce the Risk-sensitive Individual-Global-Max (RIGM) principle as a generalization of the Individual-Global-Max (IGM) and Distributional IGM (DIGM) principles. This principle requires that the collection of risk-sensitive action selections of each agent should be equivalent to the risk-sensitive action selection of the central policy. Current MARL value factorization methods do not satisfy the RIGM principle for common risk metrics such as the Value at Risk (VaR) metric or distorted risk measurements. Therefore, we propose RiskQ to address this limitation, which models the joint return distribution by modeling quantiles of it as weighted quantile mixtures of per-agent return distribution utilities. RiskQ satisfies the RIGM principle for the VaR and distorted risk metrics. We show that RiskQ can obtain promising performance through extensive experiments. The source code of RiskQ is available in https://github.com/xmu-rl-3dv/RiskQ. Chennan Ma, Weiquan Liu, Yongquan Fu, Songzhu Mei, Cheng Wang 0003 |
NeurIPS | 4 |
| 2023 | Feature Matching in the Changed Environments for Visual Localization
Xuelun Shen, Zijun Li 0002, Weiquan Liu, Cheng Wang 0003 |
PRCV (11) | 4 |
| 2023 | DSMNet: Deep High-Precision 3-D Surface Modeling From Sparse Point Cloud FramesabstractExisting point cloud modeling datasets primarily express the modeling precision by pose or trajectory precision rather than the point cloud modeling effect itself. Under this demand, we first independently construct a set of LiDAR system with an optical stage, and then we build a HPMB dataset based on the constructed LiDAR system, a High-Precision, Multi-Beam, real-world dataset. Second, we propose an modeling evaluation method based on HPMB for object-level modeling to overcome this limitation. In addition, the existing point cloud modeling methods tend to generate continuous skeletons of the global environment, hence lacking attention to the shape of complex objects. To tackle this challenge, we propose a novel learning-based joint framework, DSMNet, for high-precision 3D surface modeling from sparse point cloud frames. DSMNet comprises density-aware Point Cloud Registration (PCR) and geometry-aware Point Cloud Sampling (PCS) to effectively learn the implicit structure feature of sparse point clouds. Extensive experiments demonstrate that DSMNet outperforms the state-of-the-art methods in PCS and PCR on Multi-View Partial Point Cloud (MVP) database. Furthermore, the experiments on the open source KITTI and our proposed HPMB datasets show that DSMNet can be generalized as a post-processing of Simultaneous Localization And Mapping (SLAM), thereby improving modeling precision in environments with sparse point clouds. Changjie Qiu, Xiuhong Lin, Cheng Wang 0003, Weiquan Liu |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2023 | Adaptive local adversarial attacks on 3D point clouds
Shijun Zheng, Weiquan Liu, Chenglu Wen, Ming Cheng 0002, Cheng Wang 0003 |
Pattern Recognit. | 2 |
| 2023 | Multistage Scene-Level Constraints for Large-Scale Point Cloud Weakly Supervised Semantic SegmentationabstractCompared to fully supervised 3D large-scale point cloud segmentation methods, which necessitate extensive manual point-wise annotations, weakly supervised segmentation has emerged as a popular approach for significantly reducing labeling costs while maintaining effectiveness. However, the existing methods have exhibited inferior segmentation performance and unsatisfactory generalization capabilities in some scenarios with unique structures (e.g., building facades). In this paper, we propose an effective and generalized weakly supervised semantic segmentation framework, called multi-stage scene-level constraints (MSC), to solve the above problem. To address the issue regarding inadequate labeled data, we use pseudo-labels for unlabeled data and propose an uncertainty-guided adaptive reweighting strategy to reduce the negative impact of erroneous pseudo-labeled data on the model learning process. To address the class imbalance issue, we employ multi-stage scene-level constraints (i.e., encoder, decoder, and classifier stages) to treat each class equally and improve perception ability of the model for each class. Evaluations conducted on multiple large-scale point cloud datasets collected in different scenarios, including building facades, indoor scenes, outdoor scenes, and UAV scenes, show that our MSC achieves a large gain over the existing weakly supervised methods and even surpasses some fully supervised methods. Yanfei Su, Ming Cheng 0002, Zhimin Yuan, Weiquan Liu, Wankang Zeng, Cheng Wang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Spatial Adaptive Fusion Consistency Contrastive Constraint: Weakly Supervised Building Facade Point Cloud Semantic SegmentationabstractSemantic segmentation of building facade point clouds has diverse applications. The development of semantic segmentation methods is inextricably linked to datasets. The available building facade datasets suffer from a lack of abundant semantic categories and data completeness. To compensate for these shortcomings, we propose a new building facade dataset characterized by various categories and relatively complete 3-D building facades. In addition, most existing methods focus on fully supervised learning, which relies on manually labeling large-scale point cloud data and results in high time and labor costs. In this article, we propose an effective weakly supervised building facade segmentation approach, called spatial adaptive fusion consistency contrastive constraint (SAF-C3), to solve the above problem. We first design a multirandom point cloud augmentor as an auxiliary supervision branch to enhance the learning ability of the original network branch. Then, we present a spatial adaptive fusion (SAF) module to extract discriminative features for building facade point clouds. Finally, we propose a spatial consistency contrastive constraint to explore the contrastive property in feature space and to ensure the predictive consistency among the augmentation and original branches. The proposed method achieves a significant performance improvement against the state-of-the-art methods on two building facade point cloud datasets through extensive experiments. In particular, the performance of SAF-C3 with 1% labels significantly surpasses the baseline network with 100% labels. Yanfei Su, Ming Cheng 0002, Zhimin Yuan, Weiquan Liu, Wankang Zeng, Zhihong Zhang 0001, Cheng Wang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Prototype-Guided Multitask Adversarial Network for Cross-Domain LiDAR Point Clouds Semantic SegmentationabstractUnsupervised domain adaptation (UDA) segmentation aims to leverage labeled source data to make accurate predictions on unlabeled target data. The key is to make the segmentation network learn domain-invariant representations. In this work, we propose a prototype-guided multitask adversarial network (PMAN) to achieve this. First, we propose an intensity-aware segmentation network (IAS-Net) that leverages the private intensity information of target data to substantially facilitate feature learning of the target domain. Second, the category-level cross-domain feature alignment strategy is introduced to flee the side effects of global feature alignment. It employs the prototype (class centroid) and includes two essential operations: 1) build an auxiliary nonparametric classifier to evaluate the semantic alignment degree of each point based on the prediction consistency between the main and auxiliary classifiers and 2) introduce two class-conditional point-to-prototype learning objectives for better alignment. One is to explicitly perform category-level feature alignment in a progressive manner, and the other aims to shape the source feature representation to be discriminative. Extensive experiments reveal that our PMAN outperforms state-of-the-art results on two benchmark datasets. Zhimin Yuan, Ming Cheng 0002, Wankang Zeng, Yanfei Su, Weiquan Liu, Shangshu Yu, Cheng Wang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | LCE-NET: Contour Extraction for Large-Scale 3-D Point CloudsabstractThe contours, one of the most significant human perceptual features, have a significant impact on point cloud processing. In urban scenes, contour extraction is quite challenging due to the enormous number of unstructured and irregular points (typically greater than 107points). In this paper, we propose a Large-scale 3D point cloud Contour Extraction Network (LCE-NET) to generate contours consistent with human perception of outdoor scenes. To our knowledge, it is the first time that an end-to-end learning-based framework has been proposed for contour extraction on point cloud at this scale. The proposed LCE-Net is essentially a two-phase system. In the first phase, potential vertexes are detected from the input point cloud by the vertex detection module, then in the second phase, a designed overcomplete line proposal set is generated, and invalid line segments are further suppressed by the line proposal discrimination module. The two phases are jointly trained by a uniform loss function to promote the information interchange, thus leading to extraction results with satisfied accurate and false alarm ratings. Since there is hardly any available dataset with labeled contours for the large-scale outdoor scene, we open sourced SemanticLine, the first dataset for large-scale point clouds with labeled contour information, based on re-annotation of previous mapping level point cloud dataset semantic3D. Experimental results demonstrate that LCE-NET can effectively extract parametric contour lines from large-scale point clouds of urban scenes. Additionally, it outperforms the state-of-the-art approaches. The code will be open source on GitHub soon. Binjie Chen, Yunzhou Xia, Hanyun Guo, Yunuo Yang, Weiquan Liu, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Category-Level Adversaries for Outdoor LiDAR Point Clouds Cross-Domain Semantic SegmentationabstractUnsupervised domain adaptation (UDA) is a low-cost way to deal with the lack of annotations in a new domain. For outdoor point clouds in urban transportation scenes, the mismatch of sampling patterns and the transferability difference between classes make cross-domain segmentation extremely difficult. To overcome these challenges, we propose a category-level adversarial framework. Firstly, we propose a multi-scale domain conditioned block that facilitates to extract the critical low-level domain-dependent knowledge and reduce the domain gap caused by distinct LiDAR sampling patterns. Secondly, we make full use of multiple representation forms (i.e., point-based sets and voxel-based cells) and utilize the prediction consistency between the two forms to measure how well each point is semantically aligned. The model then focuses on the poorly-aligned points without affecting the well-aligned points. Experimental results on three autonomous driving point cloud datasets show that the proposed method outperforms existing methods by a large margin, especially on the low-beam to high-beam cross-domain segmentation task. Zhimin Yuan, Chenglu Wen, Ming Cheng 0002, Yanfei Su, Weiquan Liu, Shangshu Yu, Cheng Wang 0003 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Qrelation: an Agent Relation-Based Approach for Multi-Agent Reinforcement Learning Value Function FactorizationabstractThe Centralized Training with Decentralized Execution paradigm (CTDE), which trains policies centrally with additional information, is important for Multi-Agent Reinforcement Learning (MARL). For CTDE, value function factorization methods make use of state during training and factorize the value function into multiple local value functions for decentralized execution. These approaches do not fully consider the relational information among agents, resulting in sub-optimal models for complex tasks. To remedy this issue, we propose QRelation which is a graph neural network approach for value function factorization. It considers both the static relations (e.g., agent types) and dynamic relations (e.g., close-by). We show that QRelation can obtain better results than state-of-the-art methods on challenging StarCraft II benchmarks. Mengwei Qiu, Weiquan Liu, Cheng Wang 0003, Yongquan Fu, Peng Qiao |
ICASSP | 4 |
| 2022 | A LiDAR-Based 3D Indoor Mapping Framework with Mismatch Detection and OptimizationabstractIn this paper, we propose a novel LiDAR-based mapping framework for geometry-featureless scenarios, combining learning-based mismatch detection and intensity-assisted registration optimization. The mismatch detection method hybridizes point cloud features and temporal features of pose to detect the mismatch position, and the optimization method uses the intensity of captured road markers by LiDAR to optimize mismatch pose. Experiments on underground garages data demonstrate that the proposed method performs better localization and mapping accuracy than the LOAM. Weiquan Liu, Chenglu Wen, Yongfei Shi, Xiaocheng Yan, Jinbin Tan, Cheng Wang 0003, Jonathan Li 0001 |
IGARSS | 2 |
| 2022 | TopoSeg: Topology-aware Segmentation for Point CloudsabstractPoint cloud segmentation plays an important role in AI applications such as autonomous driving, AR, and VR. However, previous point cloud segmentation neural networks rarely pay attention to the topological correctness of the segmentation results. In this paper, focusing on the perspective of topology awareness. First, to optimize the distribution of segmented predictions from the perspective of topology, we introduce the persistent homology theory in topology into a 3D point cloud deep learning framework. Second, we propose a topology-aware 3D point cloud segmentation module, TopoSeg. Specifically, we design a topological loss function embedded in TopoSeg module, which imposes topological constraints on the segmentation of 3D point clouds. Experiments show that our proposed TopoSeg module can be easily embedded into the point cloud segmentation network and improve the segmentation performance. In addition, based on the constructed topology loss function, we propose a topology-aware point cloud edge extraction algorithm, which is demonstrated that has strong robustness. Weiquan Liu, Hanyun Guo, Weini Zhang, Cheng Wang 0003, Jonathan Li 0001 |
IJCAI | 1 |
| 2022 | ResQ: A Residual Q Function-based Approach for Multi-Agent Reinforcement Learning Value FactorizationabstractThe factorization of state-action value functions for Multi-Agent Reinforcement Learning (MARL) is important. Existing studies are limited by their representation capability, sample efficiency, and approximation error. To address these challenges, we propose, ResQ, a MARL value function factorization method, which can find the optimal joint policy for any state-action value function through residual functions. ResQ masks some state-action value pairs from a joint state-action value function, which is transformed as the sum of a main function and a residual function. ResQ can be used with mean-value and stochastic-value RL. We theoretically show that ResQ can satisfy both the individual global max (IGM) and the distributional IGM principle without representation limitations. Through experiments on matrix games, the predator-prey, and StarCraft benchmarks, we show that ResQ can obtain better results than multiple expected/stochastic value factorization methods. Mengwei Qiu, Weiquan Liu, Yongquan Fu, Cheng Wang 0003 |
NeurIPS | 4 |
| 2022 | 2D3D-MVPNet: Learning cross-domain feature descriptors for 2D-3D matching based on multi-view projections of point clouds
Baiqi Lai, Weiquan Liu, Cheng Wang 0003, Xiaoliang Fan, Yangbin Lin, Xuesheng Bian, Shangbin Wu, Ming Cheng 0002, Jonathan Li 0001 |
Appl. Intell. | 2 |
| 2022 | EAGAN: Event-based attention generative adversarial networks for optical flow and depth estimationabstractAbstract Event camera is a new vision sensor that produces independent asynchronous responses to each pixel's change of illumination intensity. The unique principle of event camera has many advantages over traditional cameras, such as low latency, high temporal resolution, and high dynamic range (HDR). These advantages make event camera ideal for dealing with high speed, HDR visual tasks, especially in automatic driving scenes. In this study, we propose an image generation network named Event‐based attention generative adversarial networks (EAGAN), which simultaneously deals with optical flow and depth estimation of monocular event camera data. In addition to the innovative network architecture and loss function suitable for depth estimation, we are also the first to process incomplete training data to obtain more dense and uniform prediction results. Experiments on the multi‐vehicle stereo event camera dataset show that our EAGAN is competitive on the depth estimation task and achieves the state‐of‐the‐art effect in the optical flow estimation task. Xiuhong Lin, Chenhui Yang, Xuesheng Bian, Weiquan Liu, Cheng Wang 0003 |
IET Comput. Vis. | 4 |
| 2022 | Semantic Segmentation of Coastal Zone on Airborne Lidar Bathymetry Point CloudsabstractLarge-scale semantic segmentation point cloud is an ongoing research topic for on-land environments. However, there is a rare deep learning research study for the sub-surface environment. Although, PointNet and its successor PointNet++ have become the cornerstone of point cloud segmentation. However, these techniques handle a relatively small number of points. This poses a natural difficulty in a large spatial scene with millions of possible points. In particular, for shallow water of coastal zone, the small number of points where the seabed and water surface meet, close points may belong to different classes. In our work, we present the semantic segmentation on a large-scale airborne Lidar bathymetry (ALB) point cloud containing millions of sample points into two classes of water surface and seabed with the voxel sampling pre-processing (VSP) approach. The proposed approach will allow us to capture the complicated outdoor natural scene components of water surface and seabed more accurately and more realistic through nonuniform voxelization in the mixture of dense and sparse points of the ALB point cloud. The performance of validation results show improvement in a per-point accuracy of 72.45% compared with other state-of-the-art deep learning-based methods. Sajjad Roshandel, Weiquan Liu, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Local Fusion Attention Network for Semantic Segmentation of Building Facade Point CloudsabstractAutomatic building facade point cloud semantic segmentation is an important step in 3-D urban building reconstruction. How to correctly segment the components (e.g., windows, walls, and columns) from the building facade is still a challenging task. According to the characteristics of building facade point clouds, we introduce local fusion attention network (LFA-Net), an efficient neural network that learns LFA features from building facade point clouds, for better capturing the local neighborhood structure information of each point. The core of LFA-Net is the LFA module, which consists of three neural units: local graph attention (LGA), local aggregation attention (LAA), and fusion attention (FA). The LFA-Net is the standard encoder-decoder architecture. Experiments demonstrate that our LFA-Net outperforms the state-of-the-art methods on the large-scale building facade point cloud dataset. Yanfei Su, Weiquan Liu, Ming Cheng 0002, Zhimin Yuan, Cheng Wang 0003 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | DLA-Net: Learning dual local attention features for semantic segmentation of large-scale building facade point clouds
Yanfei Su, Weiquan Liu, Zhimin Yuan, Ming Cheng 0002, Zhihong Zhang 0001, Xuelun Shen, Cheng Wang 0003 |
Pattern Recognit. | 2 |
| 2022 | Intracity Temperature Estimation by Physics Informed Neural Network Using Modeled Forcing Meteorology and Multispectral Satellite ImageryabstractEstimating urban surface temperature at high resolution is crucial for effective urban planning for climate-driven risks. This high-resolution surface temperature over broader scales can usually be obtained via satellite remote sensing for historical period. However, it can be hard for future predictions. This paper presents a Physics Informed Hierarchical Perception (PIHP) network, a novel approach for accurate, high-resolution and generalizable urban surface temperature estimation. The key to our approach is leveraging the implied temperature-related physics information of the land surface structure from high-resolution multi-spectral satellite images, thus achieving precise estimation or prediction for high spatial resolution urban surface temperature. Specifically, a semantic category histogram is first designed to describe the land surface structures. Based on this, a hierarchical urban surface perception network is proposed to capture the complex relationship between the underlying land surface features, upper atmosphere conditions and the intracity temperature. The proposed PIHP-Net makes it possible to generate models that can generalize across different cities, thus to estimating or predicting high-resolution urban surface temperature when the satellite land surface temperature (LST) observation is not available. Experiments over various cities in different climate regions in China show, for the first time, errors less than 2 Kelvin (for most of the cases) at the high resolution (60-by-60 meters grids), thus making it possible to predict futureintracity temperaturefrom forcing meteorology and multi-spectral satellite imagery. Donghang Wu, Weiquan Liu, Lei Zhao 0023, Shenlong Wang, Cheng Wang 0003, José Marcato Junior, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Learning Cross-Domain Descriptors for 2D-3D Matching with Hard Triplet Loss and Spatial Transformer Network
Baiqi Lai, Weiquan Liu, Cheng Wang 0003, Xuesheng Bian, Yanfei Su, Xiuhong Lin, Zhimin Yuan, Ming Cheng 0002 |
ICIG (3) | 2 |
| 2021 | Metric Learning for 2D Image Patch and 3D Point Cloud Volume MatchingabstractSimilarity measure of cross-domain descriptors (2D descriptors and 3D descriptors) between 2D image patches and 3D point cloud volumes provides stable retrieval performance and establishes the spatial relationship between 2D and 3D space, which plays the potential applications in geospatial space, such as 2D and 3D interaction of remote sensing, Augmented Reality (AR) and robot navigation. However, the mature handcrafted descriptors of 2D image patches and 3D point cloud volumes are extremely different, resulting in the huge challenge for 2D image patch and 3D point cloud volume matching. In this paper, we propose a novel network which combines both unified descriptor training and descriptor comparison function training for 2D image patch and 3D point cloud volume matching. First, two feature extraction networks are applied for jointly learning the local descriptors for 2D image patches and 3D point cloud volumes, respectively. Second, a fully connected network is introduced to compute the similarity between 2D descriptors and 3D descriptors. Motivated by the successful indicator system on evaluating 2D patch feature representation, we use the false positive rate at 95% recall (FPR95) and precision based on cross-domain descriptors as the measured metric. The experimental results show that our proposed network achieve state-of-the-art performance in the matching of 2D image patches and 3D point cloud volumes. Baiqi Lai, Weiquan Liu, Cheng Wang 0003, Xuesheng Bian, Xiuhong Lin, Chenglu Wen, Jonathan Li 0001 |
IGARSS | 2 |
| 2021 | OrgaNet: A Deep Learning Approach for Automated Evaluation of Organoids Viability in Drug Screening
Xuesheng Bian, Cheng Wang 0003, Weiquan Liu, Xiuhong Lin, Zexin Chen, Mancheung Cheung, Xióngbiao Luó |
ISBRA | 5 |
| 2021 | Y-Net: Learning Domain Robust Feature Representation for ground camera image and large-scale image-based point cloud registration
Weiquan Liu, Cheng Wang 0003, Xuesheng Bian, Baiqi Lai, Xuelun Shen, Ming Cheng 0002, Shang-Hong Lai, Dongdong Weng, Jonathan Li 0001 |
Inf. Sci. | 1 |
| 2020 | Learning to Match Ground Camera Image and UAV 3-D Model-Rendered Image Based on Siamese Network With Attention MechanismabstractDifferent domain image sensors or imaging mechanisms provide cross-domain images when sensing the same scene. There is a domain shift between cross-domain images so that the image gap between different domains is the major challenge for measuring the similarity of the feature descriptors extracted from different domain images. Specifically, matching ground camera images and unmanned aerial vehicle (UAV) 3-D model-rendered images, which are two kinds of extremely challenging cross-domain images, is a way to establish indirectly the spatial relationship between 2-D and 3-D spaces. This provides a solution for the virtual-real registration of augmented reality (AR) in outdoor environments. However, during matching, handcrafted descriptors and existing learning-based feature descriptors limit the rendered images. In this letter, first, to learn robust and invariant 128-D local feature descriptors for ground camera and rendered images, we present a novel network structure, SiamAM-Net, which embeds the autoencoders with an attention mechanism into the Siamese network. Then, to narrow the gap between the cross-domain images during the optimizing of SiamAM-Net, we design an adaptive margin for the loss function. Finally, we match the ground camera-rendered images by using the learned local feature descriptors and explore the outdoor AR virtual-real registration. Experiments show that the local feature descriptors, learned by SiamAM-Net, are robust and achieve state-of-the-art retrieval performance on the cross-domain image data set of ground camera and rendered images. In addition, several outdoor AR applications also demonstrate the usefulness of the proposed outdoor AR virtual-real registration. Weiquan Liu, Cheng Wang 0003, Xuesheng Bian, Shangshu Yu, Xiuhong Lin, Shang-Hong Lai, Dongdong Weng, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2019 | Unsupervised Optic Disc Segmentation for Cross Domain Fundus Image Based on Structure Consistency Constraint
Xuesheng Bian, Cheng Wang 0003, Weiquan Liu, Xiuhong Lin |
ICIG (1) | 3 |
| 2019 | Ground Camera Images and UAV 3D Model Registration for Outdoor Augmented RealityabstractThis paper presents a novel virtual-real registration approach for augmented reality (AR) in large-scale outdoor environments. Essentially, it is a pose estimation for the mobile camera images (ground camera images) in 3D model recovered by Unmanned Aerial Vehicle (UAV) image sequence via Structure-From-Motion (SFM) technology. The approach considers to indirectly establish the spatial relationship between 2D and 3D space by inferring the transformation relationship between the ground camera images and the UAV 3D model rendered images. Specifically, the proposed approach can overcome the positioning errors, which are deterioration and drift in the GPS, and deviation of orientation. The experimental results demonstrate the possibility of the proposed virtual-real registration approach, and show that the approach is robust, efficient and intuitive for AR in large-scale outdoor environments. Weiquan Liu, Cheng Wang 0003, Shang-Hong Lai, Dongdong Weng, Xuesheng Bian, Xiuhong Lin, Xuelun Shen, Jonathan Li 0001 |
VR | 1 |
| 2018 | Discriminative Learning of Point Cloud Feature Descriptors Based on Siamese NetworkabstractIt is challenging to direct extract the feature descriptors of the object in the point cloud, although deep learning has been widely used with the classification and detection in the point cloud, those methods hidden feature presentation in the network. Since the point cloud scanned by the Laser Scanner usually have different point density, unordered and even the different occlusion, which go beyond the reach of handcrafted descriptors, e.g. FPH, FPFH, VFH, ROPS. In this paper, we aim to direct extract the feature descriptors of the point cloud object through the raw point cloud. Inspired by the recent success of the Siamese networks[6], PointNet[7] and PointNet++[8], we propose a novel network to direct extract the feature descriptors of the whole point cloud object. We train our network with the Euclidean distance as the loss function which reflects feature descriptors similarity. The experiment object datasets were acquired by Mobile Laser Scanning (MLS) system which contains 6 categories. Experiment result shows that our network has a robust generalization, which can well direct extract the feature descriptors of the whole point cloud object. Xuelun Shen, Cheng Wang 0003, Chenglu Wen, Weiquan Liu, Xiaotian Sun 0005, Jonathan Li 0001 |
IGARSS | 4 |
| 2018 | H-Net: Neural Network for Cross-domain Image Patch MatchingabstractDescribing the same scene with different imaging style or rendering image from its 3D model gives us different domain images. Different domain images tend to have a gap and different local appearances, which raise the main challenge on the cross-domain image patch matching. In this paper, we propose to incorporate AutoEncoder into the Siamese network, named as H-Net, of which the structural shape resembles the letter H. The H-Net achieves state-of-the-art performance on the cross-domain image patch matching. Furthermore, we improved H-Net to H-Net++. The H-Net++ extracts invariant feature descriptors in cross-domain image patches and achieves state-of-the-art performance by feature retrieval in Euclidean space. As there is no benchmark dataset including cross-domain images, we made a cross-domain image dataset which consists of camera images, rendering images from UAV 3D model, and images generated by CycleGAN algorithm. Experiments show that the proposed H-Net and H-Net++ outperform the existing algorithms. Our code and cross-domain image dataset are available at https://github.com/Xylon-Sean/H-Net. Weiquan Liu, Xuelun Shen, Cheng Wang 0003, Zhihong Zhang 0001, Chenglu Wen, Jonathan Li 0001 |
IJCAI | 1 |
| 2017 | An efficiently volumetric fusing method for structure-frome-motion and terrestrial point cloudabstractAirborne acquisition and ground-view 3D point cloud provide complementary 3D information at city scale. A complete but lacks ground-view details, while the latter is incomplete for higher floors and severe occlusion. In this paper, First, a volumetric fusion method based on graph cuts were applied for fusing of airborne and terrestrial 3D LiDAR data. Second, we propose a method of constraints based on the local centroid of point cloud to eliminate the gap of fusion boundary. Finally, the experiments show that the improved fusion algorithm implement blending effectively. Wei Li 0151, Cheng Wang 0003, Dawei Zai, Pengdi Huang, Weiquan Liu, Jonathan Li 0001 |
IGARSS | 5 |
| 2000 | A Real-time Integration Of Concept-based Search and Summarization of Chinese WebsitesabstractThis paper introduces an intuitive search environment for casual and novice Chinese users over Internet. The system consists of four components, a concept network, a query reformulation model, a standard search engine, and an automatic summarizer. When the user enters one or more fairly general and vague terms, the search engine returns an initial answer set, and at the same time pipes the query to the concept network that connects thousands of conceptual nodes, each referring to a specific concept for a given domain and pointing to a number of associated conceptual terms. If the concept is located in the network, the related conceptual terms are displayed. The user has the option of using one or more of these specific terms to reformulate the next round of searches. Such search iterations continue until the user' s ultimate information seeking goal is reached. For each search iteration, auto summarizer presents the main theme of the document retrieved and an optional text-to-speech engine can read out the output summary if the user prefers. Joe F. Zhou, Weiquan Liu |
EMNLP | 2 |