Haisong Liu

dblp:24/146 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-0687-3713ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 5 since 2021Computer networks · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 SparseBEV: A Fully Sparse Framework for Multi-View 3D Object Detection
Haisong Liu, Limin Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Fully Sparse 3D Occupancy Prediction
Haisong Liu, Zetong Yang, Tianyu Li 0004, Li Chen 0008, Hongyang Li 0001, Limin Wang 0002
ECCV (25)1
2024 Learning Optical Flow and Scene Flow With Bidirectional Camera-LiDAR Fusion
abstract
In this paper, we study the problem of jointly estimating the optical flow and scene flow from synchronized 2D and 3D data. Previous methods either employ a complex pipeline that splits the joint task into independent stages, or fuse 2D and 3D information in an "early-fusion" or "late-fusion" manner. Such one-size-fits-all approaches suffer from a dilemma of failing to fully utilize the characteristic of each modality or to maximize the inter-modality complementarity. To address the problem, we propose a novel end-to-end framework, which consists of 2D and 3D branches with multiple bidirectional fusion connections between them in specific layers. Different from previous work, we apply a point-based 3D branch to extract the LiDAR features, as it preserves the geometric structure of point clouds. To fuse dense image features and sparse point features, we propose a learnable operator named bidirectional camera-LiDAR fusion module (Bi-CLFM). We instantiate two types of the bidirectional fusion pipeline, one based on the pyramidal coarse-to-fine architecture (dubbed CamLiPWC), and the other one based on the recurrent all-pairs field transforms (dubbed CamLiRAFT). On FlyingThings3D, both CamLiPWC and CamLiRAFT surpass all existing methods and achieve up to a 47.9% reduction in 3D end-point-error from the best published result. Our best-performing model, CamLiRAFT, achieves an error of 4.26% on the KITTI Scene Flow benchmark, ranking 1st among all submissions with much fewer parameters. Besides, our methods have strong generalization performance and the ability to handle non-rigid motion.
Haisong Liu, Tao Lu 0005, Jia Liu 0008, Limin Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 LinK: Linear Kernel for LiDAR-based 3D Perception
abstract
Extending the success of 2D Large Kernel to 3D perception is challenging due to: 1. the cubically-increasing overhead in processing 3D data; 2. the optimization difficulties from data scarcity and sparsity. Previous work has taken the first step to scale up the kernel size from 3 × 3 × 3 to 7 × 7 × 7 by introducing block-shared weights. However, to reduce the feature variations within a block, it only employs modest block size and fails to achieve larger kernels like the 21 × 21 × 21. To address this issue, we propose a new method, called LinK, to achieve a wider-range perception receptive field in a convolution-like manner with two core designs. The first is to replace the static kernel matrix with a linear kernel generator, which adaptively provides weights only for non-empty voxels. The second is to reuse the pre-computed aggregation results in the overlapped blocks to reduce computation complexity. The proposed method successfully enables each voxel to perceive context within a range of 21 × 21 × 21. Extensive experiments on two basic perception tasks, 3D object detection and 3D semantic segmentation, demonstrate the effectiveness of our method. Notably, we rank 1st on the public leaderboard of the 3D detection benchmark of nuScenes (LiDAR track), by simply incorporating a LinK-based backbone into the basic detector, CenterPoint. We also boost the strong segmentation baseline's mIoU with 2.7% in the SemanticKITTI test set. Code is available at https://github.com/MCG-NJU/LinK.
Tao Lu 0005, Haisong Liu, Gangshan Wu, Limin Wang 0002
CVPR3
2023 SparseBEV: High-Performance Sparse 3D Object Detection from Multi-Camera Videos
abstract
Camera-based 3D object detection in BEV (Bird’s Eye View) space has drawn great attention over the past few years. Dense detectors typically follow a two-stage pipeline by first constructing a dense BEV feature and then performing object detection in BEV space, which suffers from complex view transformations and high computation cost. On the other side, sparse detectors follow a query-based paradigm without explicit dense BEV feature construction, but achieve worse performance than the dense counterparts. In this paper, we find that the key to mitigate this performance gap is the adaptability of the detector in both BEV and image space. To achieve this goal, we propose SparseBEV, a fully sparse 3D object detector that outperforms the dense counterparts. SparseBEV contains three key designs, which are (1) scale-adaptive self attention to aggregate features with adaptive receptive field in BEV space, (2) adaptive spatio-temporal sampling to generate sampling locations under the guidance of queries, and (3) adaptive mixing to decode the sampled features with dynamic weights from the queries. On the test split of nuScenes, SparseBEV achieves the state-of-the-art performance of 67.5 NDS. On the val split, SparseBEV achieves 55.8 NDS while maintaining a real-time inference speed of 23.5 FPS. Code is available at https://github.com/MCG-NJU/SparseBEV.
Haisong Liu, Yao Teng, Tao Lu 0005, Limin Wang 0002
ICCV1
2023 StageInteractor: Query-based Object Detector with Cross-stage Interaction
abstract
Previous object detectors make predictions based on dense grid points or numerous preset anchors. Most of these detectors are trained with one-to-many label assignment strategies. On the contrary, recent query-based object detectors are based a sparse set of learnable queries refined by a series of decoder layers. The one-to-one label assignment is independently applied on each layer for deep supervision during training. Despite the great success of query-based object detection, however, this vanilla one-to-one label assignment strategy requires the detectors to have strong fine-grained discrimination and modeling capacity. In this paper, we propose a new query-based object detector with cross-stage interaction, coined as StageInter-actor. During the forward pass, we come up with an efficient way to improve this modeling ability by reusing dynamic operators with lightweight adapters. As for the label assignment, a cross-stage label assigner is designed to improve the one-to-one label assignment. With this assigner, the training target class labels are gathered across stages and then reallocated to proper predictions at each decoder layer. On MS COCO benchmark, our model improves the baseline counterpart by 2.2 AP, and achieves a 44.8 AP with ResNet-50 as backbone, 100 queries and 12 training epochs. With longer training time and 300 queries, StageIn-teractor achieves 51.3 AP and 52.7 AP with ResNeXt-101-DCN and Swin-S, respectively. The code and models are made available at https://github.com/MCG-NJU/StageInteractor.
Yao Teng, Haisong Liu, Sheng Guo 0005, Limin Wang 0002
ICCV2
2023 Online deep Bingham network for probabilistic orientation estimation
abstract
Abstract Orientation estimation is one of the core problems in several computer vision tasks. Recently deep learning techniques combined with the Bingham distribution have attracted considerable interest towards this problem when considering ambiguities and rotational symmetries of objects. However, existing works suffer from two issues. First, the computational overhead for calculating the normalisation constant of the Bingham distribution is relatively high. Second, the choice of loss functions is uncertain. In light of these problems, we present an online deep Bingham network to estimate the orientation of objects. We sharply reduce the computational overhead of the normalisation constant by directly applying a numerical integration formula. Additionally, we are the first to give theorems on the convexity and Lipschitz continuity of the Bingham distribution's negative log‐likelihood, which formally indicates that it is a proper choice of the loss function. We test our method on three public datasets, namely the UPNA, the T‐LESS and Pascal3D+, showing that our method outperforms the state‐of‐the‐art in terms of orientation accuracy and time efficiency, which can reduce the runtime by more than 6 h compared to the offline methods. The ablation experiments further demonstrate the effectiveness and robustness of our model.
Wenjie Li 0002, Jia Liu 0008, Haisong Liu, Dayong Ren, Yanyan Wang 0001, Lijun Chen 0006
IET Comput. Vis.4
2022 CamLiFlow: Bidirectional Camera-LiDAR Fusion for Joint Optical Flow and Scene Flow Estimation
abstract
In this paper, we study the problem of jointly estimating the optical flow and scene flow from synchronized 2D and 3D data. Previous methods either employ a complex pipeline that splits the joint task into independent stages, or fuse 2D and 3D information in an “early-fusion“ or “late-fusion“ manner. Such one-size-fits-all approaches suffer from a dilemma of failing to fully utilize the characteristic of each modality or to maximize the inter-modality complementarity. To address the problem, we propose a novel end-to-end framework, called CamLiFlow. It consists of 2D and 3D branches with multiple bidirectional connections between them in specific layers. Different from previous work, we apply a point-based 3D branch to better extract the geometric features and design a symmetric learnable operator to fuse dense image features and sparse point features. Experiments show that CamLiFlow achieves better performance with fewer parameters. Our method ranks 1st on the KITTI Scene Flow benchmark, outperforming the previous art with 1/7 parameters. Code is available at https://github.com/MCG-NJU/CamLiFlow.
Haisong Liu, Tao Lu 0005, Jia Liu 0008, Wenjie Li 0002, Lijun Chen 0006
CVPR1
2022 Time-Efficient Range Detection in Commodity RFID Systems
abstract
RFID is becoming ubiquitously available in our daily life. After RFID tags are deployed to make attached objects identifiable, a natural next step is to communicate with the tags and collect their information for the purpose of tracking tagged objects or monitoring their surroundings in real-time. In this paper, we study an under-investigated problem range detection in a commodity RFID system, which aims to check if there are any tags with the data between an upper and lower boundary in a time-efficient way. This is important especially in a large RFID system, which can help users quickly pinpoint the target tags (if any) and give an early warning to users for taking urgent actions and reducing the potential risk in the nascent stage. We propose two tailored protocols, selective query and range query, to achieve range detection within the scope of the C1G2 standard. The novelty is that, instead of querying each tag, we exploit the capability of C1G2-compatible selection and quickly separate target tags from others by silencing most of tags. The final result is that range query is able to achieve a range detection with only one query command. We implement the proposed protocols in commodity RFID systems, with no need for any hardware modifications. Extensive experiments show that range query is able to improve the time efficiency by an order of magnitude, compared with the baseline.
Jia Liu 0008, Xuan Liu 0001, Haisong Liu, Yanyan Wang 0001, Lijun Chen 0006
IEEE/ACM Trans. Netw.5
2021 MTPGait: Multi-person Gait Recognition with Spatio-temporal Information via Millimeter Wave Radar
abstract
As one of the important methods of identity recognition, gait recognition has a wide range of applications in the fields of new human-computer interaction, smart home, smart office and health monitoring. In this paper, we propose a system for multi-person gait recognition (MTPGait) with spatio-temporal information via millimeter wave radar. We specially design a neural network that can extract multi-scale spatio-temporal features along space and time dimensions of 3D point cloud concisely and efficiently. In addition, we construct and release a millimeter wave radar 3D point cloud data set, which consists of 960-minute gait data of 25 volunteers. The experimental results show that MTPGait is able to achieve 96.7% recognition accuracy in a single-person scene on random routes, and 90.2 % recognition accuracy when two people coexist, while the accuracy of the existing methods can not reach 90 % in either scenario.
Haisong Liu, Chenqi Shi
ICPADS3
2020 Time-efficient Range Detection in Commodity RFID Systems
abstract
RFID is becoming ubiquitously available in our daily life. After RFID tags are deployed to make attached objects identifiable, a natural next step is to communicate with the tags and collect their information for the purpose of tracking tagged objects or monitoring their surroundings in real time. In this paper, we study an under-investigated problem range detection in a commodity RFID system, which aims to check if there are any tags with the data between an upper and lower boundary in a time-efficient way. This is important especially in a large RFID system, which can help users quickly pinpoint the target tags (if any) and give an early warning to users for taking urgent actions and reducing the potential risk in the nascent stage. We propose two tailored protocols, selective query and range query (RQ), to achieve range detection within the scope of the C1G2 standard. The novelty is that, instead of querying each tag, we exploit the capability of C1G2-compatible selection and quickly separate target tags from others by silencing most of tags. The final result is that our best protocol RQ is able to achieve a range detection with only one query command. We implement the proposed protocols in commodity RFID systems, without any modifications of hardware. Extensive experiments show that RQ is able to improve the time efficiency by near 30×, compared with the baseline.
Jia Liu 0008, Haisong Liu, Hualin Gong, Yanyan Wang 0001, Lijun Chen 0006
ICNP3
2020 Eingerprint: Robust Energy-related Fingerprinting for Passive RFID Tags
Jia Liu 0008, Xia Wang 0004, Haisong Liu, Lijun Chen 0006
NSDI4
2020 Localization and Completion for 3D Object Interactions
abstract
Finding where and what objects to put into an existing scene is a common task for scene synthesis and robot/character motion planning. Existing frameworks require development of hand-crafted features suitable for the task, or full volumetric analysis that could be memory intensive and imprecise. In this paper, we propose a data-driven framework to discover a suitable location and then place the appropriate objects in a scene. Our approach is inspired by computer vision techniques for localizing objects in images: using an all directional depth image (ADD-image) that encodes the 360-degree field of view from samples in the scene, our system regresses the images to the positions where the new object can be located. Given several candidate areas around the host object in the scene, our system predicts the partner object whose geometry fits well to the host object. Our approach is highly parallel and memory efficient, and is especially suitable for handling interactions between large and small objects. We show examples where the system can hang bags on hooks, fit chairs in front of desks, put objects into shelves, insert flowers into vases, and put hangers onto laundry rack.
Xi Zhao 0002, Ruizhen Hu, Haisong Liu, Taku Komura, Xinyu Yang 0001
IEEE Trans. Vis. Comput. Graph.3
2006 Seed Localization Using Trus and Gre Based Gaussian Fil Tering For Brachytherapy Applications
abstract
In this paper, we propose a novel algorithm for detecting needles and their corresponding implanted radioactive seed locations in the prostate during brachytherapy from transrectal ultrasound images. This is carried out efficiently using separable Gaussian filters in a probabilistic Gibbs random field framework. An approximation of the needle path through the prostate volume is obtained using a polynomial fit. The seeds are then detected and assigned to their corresponding needles by calculating local maxima. In our experiments, we were able to successfully localize over 85% of the implanted seeds
Varsha Sampath, Vladimir Misic, Eli Saber, Haisong Liu
ICASSP (2)4
2001 Dynamic Brachytherapy of the Prostate Under Active Image Guidance
Haisong Liu, Lydia Liao
MICCAI2