EDBT 2026 Demo / reviewers in the wild / expert
Haisheng Su
dblp:222/2080
· DBLP profile ↗
15ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0002-4228-7439ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 5 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | UniMamba: Unified Spatial-Channel Representation Learning with Group-Efficient Mamba for LiDAR-based 3D Object DetectionabstractRecent advances in LiDAR 3D detection have demonstrated the effectiveness of Transformer-based frameworks in capturing the global dependencies from point cloud spaces, which serialize the 3D voxels into the flattened 1D sequence for iterative self-attention. However, the spatial structure of 3D voxels will be inevitably destroyed during the serialization process. Besides, due to the considerable number of 3D voxels and quadratic complexity of Transformers, multiple sequences are grouped before feeding to Transformers, leading to a limited receptive field. Inspired by the impressive performance of State Space Models (SSM), in this paper, we propose a novel Unified Mamba (UniMamba), which seamlessly integrates the merits of 3D convolution and SSM in a concise multi-head manner, aiming to perform "local and global" spatial context aggregation efficiently and simultaneously. Specifically, a Uni-Mamba block is designed which mainly consists of spatial locality modeling, complementary Z-order serialization and local-global sequential aggregator. The spatial locality modeling module integrates 3D submanifold convolution to capture the dynamic spatial position embedding before serialization. Then the efficient Z-order curve is adopted for serialization both horizontally and vertically. Furthermore, the local-global sequential aggregator adopts the channel grouping strategy to efficiently encode both "local and global" spatial inter-dependencies using multi-head SSM. Additionally, an encoder-decoder architecture with stacked UniMamba blocks is formed to facilitate multi-scale spatial learning hierarchically. Extensive experiments are conducted on three popular datasets: nuScenes, Waymo and Argoverse 2. Particularly, our UniMamba achieves 70.2 mAP on the nuScenes dataset. Xin Jin 0014, Haisheng Su, Wei Wu 0021, Fei Hui, Junchi Yan |
CVPR | 2 |
| 2025 | RoboSense: Large-scale Dataset and Benchmark for Egocentric Robot Perception and Navigation in Crowded and Unstructured EnvironmentsabstractReliable embodied perception from an egocentric perspective is challenging yet essential for autonomous navigation technology of intelligent mobile agents. With the growing demand of social robotics, near-field scene understanding becomes an important research topic in the areas of egocentric perceptual tasks related to navigation in both crowded and unstructured environments. Due to the complexity of environmental conditions and difficulty of surrounding obstacles owing to truncation and occlusion, the perception capability under this circumstance is still inferior. To further enhance the intelligence of mobile robots, in this paper, we setup an egocentric multi-sensor data collection platform based on 3 main types of sensors (Camera, LiDAR and Fisheye), which supports flexible sensor configurations to enable dynamic sight of view from ego-perspective, capturing either near or farther areas. Meanwhile, a large-scale multimodal dataset is constructed, named RoboSense, to facilitate egocentric robot perception. Specifically, RoboSense contains more than 133K synchronized data with 1.4M 3D bounding box and IDs annotated in the full 360° view, forming 216K trajectories across 7.6K temporal sequences. It has 270× and 18× as many annotations of surrounding obstacles within near ranges as the previous datasets collected for autonomous driving scenarios such as KITTI and nuScenes. Moreover, we define a novel matching criterion for near-field 3D perception and prediction metrics. Based on RoboSense, we formulate 6 popular tasks to facilitate the future research development, where the detailed analysis as well as benchmarks are also provided accordingly. Data desensitization measures have been conducted for privacy protection. Haisheng Su, Feixiang Song, Wei Wu 0021, Junchi Yan |
CVPR | 1 |
| 2025 | GeoFormer: Geometry Point Encoder for 3D Object Detection with Graph-Based Transformer
Xin Jin 0014, Haisheng Su, Wei Wu 0021, Fei Hui, Junchi Yan |
ICCV | 2 |
| 2025 | FreqPDE: Rethinking Positional Depth Embedding for Multi-View 3D Object Detection TransformersabstractDetecting 3D objects accurately from multi-view 2D images is a challenging yet essential task in the field of autonomous driving. Current methods resort to integrating depth prediction to recover the spatial information for object query decoding, which necessitates explicit supervision from LiDAR points during the training phase. However, the predicted depth quality is still unsatisfactory such as depth discontinuity of object boundaries and indistinction of small objects, which are mainly caused by the sparse supervision of projected points and the use of high-level image features for depth prediction. Besides, cross-view consistency and scale invariance are also overlooked in previous methods. In this paper, we introduce Frequency-aware Positional Depth Embedding (FreqPDE) to equip 2D image features with spatial information for 3D detection transformer decoder, which can be obtained through three main modules. Specifically, the Frequency-aware Spatial Pyramid Encoder (FSPE) constructs a feature pyramid by combining high-frequency edge clues and low-frequency semantics from different levels respectively. Then the Cross-view Scale-invariant Depth Predictor (CSDP) estimates the pixel-level depth distribution with cross-view and efficient channel attention mechanism. Finally, the Positional Depth Encoder (PDE) combines the 2D image features and 3D position embeddings to generate the 3D depth-aware features for query decoding. Additionally, hybrid depth supervision is adopted for complementary depth learning from both metric and distribution aspects. Extensive experiments conducted on the nuScenes dataset demonstrate the effectiveness and superiority of our proposed method. Haisheng Su, Feixiang Song, Sanping Zhou, Wei Wu 0021, Junchi Yan, Nanning Zheng 0001 |
ICCV | 1 |
| 2024 | Regularity Learning via Explicit Distribution Modeling for Skeletal Video Anomaly DetectionabstractAnomaly detection in surveillance videos is challenging but important for ensuring public security. Different from pixel-based anomaly detection methods, pose-based methods utilize highly-structured skeleton data, which decreases the computational burden and also avoids the negative impact of background noise. However, pose-based methods lack an alternative dynamic representation akin to the explicit motion features, such as optical flow, employed by pixel-based methods. In this paper, a novel Motion Embedder (ME), a label-efficient scheme without extra annotation efforts, is proposed to provide a pose motion representation for the structured posed data from a probability perspective. Furthermore, a novel task-specific Spatial-Temporal Transformer (STT) is deployed for self-supervised pose sequence reconstruction. These two modules are then integrated into a unified framework for pose regularity learning, which is referred to as Motion Prior Regularity Learner (MoPRL). MoPRL achieves competitive results on multiple challenging datasets while minimizing computational costs. Extensive experiments validate the versatility of the proposed modules and provide insights for future research. Shoubin Yu, Zhongyin Zhao, Haoshu Fang, Andong Deng, Haisheng Su, Weihao Gan, Cewu Lu, Wei Wu 0021 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Learning Video Representations of Human Motion from Synthetic DataabstractIn this paper, we take an early step towards video representation learning of human actions with the help of large-scale synthetic videos, particularly for human motion representation enhancement. Specifically, we first introduce an automatic action-related video synthesis pipeline based on a photorealistic video game. A large-scale human action dataset named GATA (GTA Animation Transformed Actions) is then built by the proposed pipeline, which includes 8.1 million action clips spanning over 28K action classes. Based on the presented dataset, we design a contrastive learning framework for human motion representation learning, which shows significant performance improvements on several typical video datasets for action recognition, e.g., Charades, HAA 500 and NTU-RGB. Besides, we further explore a domain adaptation method based on cross-domain positive pairs mining to alleviate the domain gap between synthetic and realistic data. Extensive properties analyses of learned representation are conducted to demonstrate the effectiveness of the proposed dataset for enhancing human motion representation learning. Wei Wu 0021, Jing Su 0005, Haisheng Su, Weihao Gan |
CVPR | 5 |
| 2021 | BSN++: Complementary Boundary Regressor with Scale-Balanced Relation Modeling for Temporal Action Proposal GenerationabstractGenerating human action proposals in untrimmed videos is an important yet challenging task with wide applications. Current methods often suffer from the noisy boundary locations and the inferior quality of confidence scores used for proposal retrieving. In this paper, we present BSN++, a new framework which exploits complementary boundary regressor and relation modeling for temporal proposal generation. First, we propose a novel boundary regressor based on the complementary characteristics of both starting and ending boundary classifiers. Specifically, we utilize the U-shaped architecture with nested skip connections to capture rich contexts and introduce bi-directional boundary matching mechanism to improve boundary precision. Second, to account for the proposal-proposal relations ignored in previous methods, we devise a proposal relation block to which includes two self-attention modules from the aspects of position and channel. Furthermore, we find that there inevitably exists data imbalanced problems in the positive/negative proposals and temporal durations, which harm the model performance on tail distributions. To relieve this issue, we introduce the scale-balanced re-sampling strategy. Extensive experiments are conducted on two popular benchmarks: ActivityNet-1.3 and THUMOS14, which demonstrate that BSN++ achieves the state-of-the-art performance. Not surprisingly, the proposed BSN++ ranked 1st place in the CVPR19 - ActivityNet challenge leaderboard on temporal action localization task. Haisheng Su, Weihao Gan, Wei Wu 0021, Yu Qiao 0001 |
AAAI | 1 |
| 2021 | Temporal Context Aggregation Network for Temporal Action Proposal RefinementabstractTemporal action proposal generation aims to estimate temporal intervals of actions in untrimmed videos, which is a challenging yet important task in the video understanding field. The proposals generated by current methods still suffer from inaccurate temporal boundaries and inferior confidence used for retrieval owing to the lack of efficient temporal modeling and effective boundary context utilization. In this paper, we propose Temporal Context Aggregation Network (TCANet) to generate high-quality action proposals through "local and global" temporal context aggregation and complementary as well as progressive boundary refinement. Specifically, we first design a Local-Global Temporal Encoder (LGTE), which adopts the channel grouping strategy to efficiently encode both "local and global" temporal inter-dependencies. Furthermore, both the boundary and internal context of proposals are adopted for frame-level and segment-level boundary regressions, respectively. Temporal Boundary Regressor (TBR) is designed to combine these two regression granularities in an end-to-end fashion, which achieves the precise boundaries and reliable confidence of proposals through progressive refinement. Extensive experiments are conducted on three challenging datasets: HACS, ActivityNet-v1.3, and THUMOS-14, where TCANet can generate proposals with high precision and recall. By combining with the existing action classifier, TCANet can obtain remarkable temporal action detection performance compared with other methods. Not surprisingly, the proposed TCANet won the 1stplace in the CVPR 2020 - HACS challenge leaderboard on temporal action localization task. Zhiwu Qing, Haisheng Su, Weihao Gan, Wei Wu 0021, Xiang Wang 0012, Yu Qiao 0001, Changxin Gao, Nong Sang |
CVPR | 2 |
| 2021 | Transferable Knowledge-Based Multi-Granularity Fusion Network for Weakly Supervised Temporal Action DetectionabstractDespite remarkable progress, temporal action detection is still limited for real application due to the great amount of manual annotations. This issue motivates interest in addressing this task under weak supervision, namely, locating the action instances using only video-level class labels. Many current works on this task are mainly based on the Class Activation Sequence (CAS), which is generated by the video classification network to describe the probability of each snippet being in a specific action class of the video. However, the CAS generated by a simple classification network can only focus on local discriminative parts instead of locating the entire interval of target actions. In this paper, we present a novel framework to handle this issue. Specifically, we propose to utilize convolutional kernels with varied dilation rates to enlarge the receptive fields, which can transfer the discriminative information to the surrounding non-discriminative regions. Then, we design a cascaded module with the proposed Online Adversarial Erasing (OAE) mechanism to further mine more relevant regions of target actions by feeding the erased-feature maps of discovered regions back into the system. In addition, inspired by the transfer learning method, we adopt an additional module to transfer the knowledge from trimmed videos to untrimmed videos to promote the classification performance on untrimmed videos. Finally, we employ a boundary regression module embedded with Outer-Inner-Contrastive (OIC) loss to automatically predict the boundaries based on the enhanced CAS. Extensive experiments are conducted on two challenging datasets, THUMOS14 and ActivityNet-1.3, and the experimental results clearly demonstrate the superiority of our unified framework. Haisheng Su, Xu Zhao 0001, Shuming Liu 0001, Zhilan Hu |
IEEE Trans. Multim. | 1 |
| 2020 | TSI: Temporal Scale Invariant Network for Action Proposal Generation
Shuming Liu 0001, Xu Zhao 0001, Haisheng Su, Zhilan Hu |
ACCV (5) | 3 |
| 2020 | Joint Learning of Local and Global Context for Temporal Action Proposal GenerationabstractTemporal action proposal generation is an important yet challenging problem, since temporal proposals with rich action content are indispensable for analysing real-world videos with long duration and high proportion irrelevant content. This problem requires methods not only generating proposals with precise temporal boundaries, but also retrieving proposals to cover ground truth action instances with high recall and high overlap using relatively fewer proposals. To address these difficulties, we introduce an effective and efficient proposal generation method, named Local-Global Network (LGN), by which local and global contexts are jointly learned to generate high quality proposals. Locally, LGN first locates temporal boundaries with high starting and ending probabilities separately, then directly combines these boundaries as proposals. Globally, LGN evaluates the actionness probability of multiple-durations temporal regions simultaneously using temporal convolutional layers and anchor mechanism. Finally, we combine the boundary probabilities of each proposal with actionness probability of matched temporal regions as the confidence score, which is used for retrieving proposals. We conduct experiments on two datasets: ActivityNet-1.3 and THUMOS-14, where LGN outperforms other state-of-the-art methods with both high recall and high temporal precision. Finally, further experiments demonstrate that by combining existing action classifiers, our method significantly improves the state-of-the-art temporal action detection performance. Xu Zhao 0001, Haisheng Su |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Attention-Based Multiview Re-Observation Fusion Network for Skeletal Action RecognitionabstractAction recognition is an important and popular area in computer vision. Because of the helpfulness of action recognition of the skeleton and the development of related pose estimation techniques, action recognition based on skeleton data has drawn considerable attention and has been widely studied in recent years. In this paper, we propose an attention-based multiview re-observation fusion model for skeletal action recognition. The proposed model focuses on the factor of observation view of actions, which greatly influences action recognition. The model utilizes action information from multiple observation views to improve the recognition performance. In this method, we re-observe input skeleton data from several possible viewpoints, process these augmented observation data with a long short-term memory (LSTM) network separately, and, finally, fuse the outputs to generate the final recognition result. In the multiview fusion process, an attention mechanism is applied to regulate the fusion operation according to the helpfulness for the recognition of all views. In this way, the model can fuse information from multiple viewpoints to recognize actions and can learn to evaluate observation views to improve fusion performance. We also propose a multilayer feature attention method to improve the performance of the LSTM in our model. We utilize an attention mechanism to enhance the feature expression by finding and focusing on informative feature dimensions according to contextual action information. Moreover, we propose stacking multiple layers of attention operation in a multilayer LSTM network to further improve network performance. The final model is integrated into an end-to-end trainable network. Experiments conducted on two popular datasets, NTU RGB+D and SBU Kinect interaction, show that our model achieves state-of-the-art performance. Zhaoxuan Fan, Xu Zhao 0001, Haisheng Su |
IEEE Trans. Multim. | 4 |
| 2018 | Cascaded Pyramid Mining Network for Weakly Supervised Temporal Action Localization
Haisheng Su, Xu Zhao 0001 |
ACCV (2) | 1 |
| 2018 | BSN: Boundary Sensitive Network for Temporal Action Proposal Generation
Xu Zhao 0001, Haisheng Su, Chongjing Wang, Ming Yang 0002 |
ECCV (4) | 3 |
| 2018 | Weakly Supervised Temporal Action Detection with Shot-Based Temporal Pooling Network
Haisheng Su, Xu Zhao 0001, Haiping Fei |
ICONIP (4) | 1 |