Shuangjie Xu

dblp:205/3160 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
9since 2021 · last 2024
0000-0003-0150-7068ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021
YearPublicationVenuePosition
2024 Learning High-Resolution Vector Representation from Multi-camera Images for 3D Object Detection
Shuangjie Xu, Maosheng Ye, Zian Qian, Xiaoyi Zou, Dit-Yan Yeung, Qifeng Chen 0001
ECCV (35)2
2024 PPAD: Iterative Interactions of Prediction and Planning for End-to-End Autonomous Driving
Maosheng Ye, Shuangjie Xu, Tongyi Cao, Qifeng Chen 0001
ECCV (35)3
2023 SVQNet: Sparse Voxel-Adjacent Query Network for 4D Spatio-Temporal LiDAR Semantic Segmentation
abstract
LiDAR-based semantic perception tasks are critical yet challenging for autonomous driving. Due to the motion of objects and static/dynamic occlusion, temporal information plays an essential role in reinforcing perception by enhancing and completing single-frame knowledge. Previous approaches either directly stack historical frames to the current frame or build a 4D spatio-temporal neighborhood using KNN, which duplicates computation and hinders real-time performance. Based on our observation that stacking all the historical points would damage performance due to a large amount of redundant and misleading information, we propose the Sparse Voxel-Adjacent Query Network (SVQNet) for 4D LiDAR semantic segmentation. To take full advantage of the historical frames high-efficiently, we shunt the historical points into two groups with reference to the current points. One is the Voxel-Adjacent Neighborhood carrying local enhancing knowledge. The other is the Historical Context completing the global knowledge. Then we propose new modules to select and extract the instructive features from the two groups. Our SVQNet achieves state-of-the-art performance in LiDAR semantic segmentation of the SemanticKITTI benchmark and the nuScenes dataset.
Xuechao Chen, Shuangjie Xu, Xiaoyi Zou, Tongyi Cao, Dit-Yan Yeung
ICCV2
2022 Sparse Cross-Scale Attention Network for Efficient LiDAR Panoptic Segmentation
abstract
Two major challenges of 3D LiDAR Panoptic Segmentation (PS) are that point clouds of an object are surface-aggregated and thus hard to model the long-range dependency especially for large instances, and that objects are too close to separate each other. Recent literature addresses these problems by time-consuming grouping processes such as dual-clustering, mean-shift offsets and etc., or by bird-eye-view (BEV) dense centroid representation that downplays geometry. However, the long-range geometry relationship has not been sufficiently modeled by local feature learning from the above methods. To this end, we present SCAN, a novel sparse cross-scale attention network to first align multi-scale sparse features with global voxel-encoded attention to capture the long-range relationship of instance context, which is able to boost the regression accuracy of the over-segmented large objects. For the surface-aggregated points, SCAN adopts a novel sparse class-agnostic representation of instance centroids, which can not only maintain the sparsity of aligned features to solve the under-segmentation on small objects, but also reduce the computation amount of the network through sparse convolution. Our method outperforms previous methods by a large margin in the SemanticKITTI dataset for the challenging 3D PS task, achieving 1st place with a real-time inference speed.
Shuangjie Xu, Rui Wan, Maosheng Ye, Xiaoyi Zou, Tongyi Cao
AAAI1
2022 Efficient Point Cloud Segmentation with Geometry-Aware Sparse Networks
Maosheng Ye, Rui Wan, Shuangjie Xu, Tongyi Cao, Qifeng Chen 0001
ECCV (39)3
2021 Spatiotemporal Graph Neural Network based Mask Reconstruction for Video Object Segmentation
abstract
This paper addresses the task of segmenting class-agnostic objects in semi-supervised setting. Although previous detection based methods achieve relatively good performance, these approaches extract the best proposal by a greedy strategy, which may lose the local patch details outside the chosen candidate. In this paper, we propose a novel spatiotemporal graph neural network (STG-Net) to reconstruct more accurate masks for video object segmentation, which captures the local contexts by utilizing all proposals. In the spatial graph, we treat object proposals of a frame as nodes and represent their correlations with an edge weight strategy for mask context aggregation. To capture temporal information from previous frames, we use a memory network to refine the mask of current frame by retrieving historic masks in a temporal graph. The joint use of both local patch details and temporal relationships allow us to better address the challenges such as object occlusions and missing. Without online learning and fine-tuning, our STG-Net achieves state-of-the-art performance on four large benchmarks, demonstrating the effectiveness of the proposed approach.
Daizong Liu, Shuangjie Xu, Xiao-Yang Liu, Zichuan Xu, Wei Wei 0002, Pan Zhou 0001
AAAI2
2021 DRINet: A Dual-Representation Iterative Learning Network for Point Cloud Segmentation
abstract
We present a novel and flexible architecture for point cloud segmentation with dual-representation iterative learning. In point cloud processing, different representations have their own pros and cons. Thus, finding suitable ways to represent point cloud data structure while keeping its own internal physical property such as permutation and scale-invariant is a fundamental problem. Therefore, we propose our work, DRINet, which serves as the basic network structure for dual-representation learning with great flexibility at feature transferring and less computation cost, especially for large-scale point clouds. DRINet mainly consists of two modules called Sparse Point-Voxel Feature Extraction and Sparse Voxel-Point Feature Extraction. By utilizing these two modules iteratively, features can be propagated between two different representations. We further propose a novel multi-scale pooling layer for pointwise locality learning to improve context information propagation. Our network achieves state-of-the-art results for point cloud classification and segmentation tasks on several datasets while maintaining high runtime efficiency. For large-scale outdoor scenarios, our method outperforms state-of-the-art methods with a real-time inference time of 62ms per frame.
Maosheng Ye, Shuangjie Xu, Tongyi Cao, Qifeng Chen 0001
ICCV2
2021 Coarse to Fine: Domain Adaptive Crowd Counting via Adversarial Scoring Network
abstract
Recent deep networks have convincingly demonstrated high capability in crowd counting, which is a critical task attracting widespread attention due to its various industrial applications. Despite such progress, trained data-dependent models usually can not generalize well to unseen scenarios because of the inherent domain shift. To facilitate this issue, this paper proposes a novel adversarial scoring network (ASNet) to gradually bridge the gap across domains from coarse to fine granularity. In specific, at the coarse-grained stage, we design a dual-discriminator strategy to adapt source domain to be close to the targets from the perspectives of both global and local feature space via adversarial learning. The distributions between two domains can thus be aligned roughly. At the fine-grained stage, we explore the transferability of source characteristics by scoring how similar the source samples are to target ones from multiple levels based on generative probability derived from coarse stage. Guided by these hierarchical scores, the transferable source features are properly selected to enhance the knowledge transfer during the adaptation process. With the coarse-to-fine design, the generalization bottleneck induced from the domain discrepancy can be effectively alleviated. Three sets of migration experiments show that the proposed methods achieve state-of-the-art counting performance compared with major unsupervised methods.
Zhikang Zou, Xiaoye Qu, Pan Zhou 0001, Shuangjie Xu, Xiaoqing Ye, Jin Ye 0006
ACM Multimedia4
2021 Computer vision and long short-term memory: Learning to predict unsafe behaviour in construction
Ting Kong, Weili Fang, Peter E. D. Love, Hanbin Luo, Shuangjie Xu, Heng Li 0001
Adv. Eng. Informatics5
2020 HVNet: Hybrid Voxel Network for LiDAR Based 3D Object Detection
abstract
We present Hybrid Voxel Network (HVNet), a novel one-stage unified network for point cloud based 3D object detection for autonomous driving. Recent studies show that 2D voxelization with per voxel PointNet style feature extractor leads to accurate and efficient detector for large 3D scenes. Since the size of the feature map determines the computation and memory cost, the size of the voxel becomes a parameter that is hard to balance. A smaller voxel size gives a better performance, especially for small objects, but a longer inference time. A larger voxel can cover the same area with a smaller feature map, but fails to capture intricate features and accurate location for smaller objects. We present a Hybrid Voxel network that solves this problem by fusing voxel feature encoder (VFE) of different scales at point-wise level and project into multiple pseudo-image feature maps. We further propose an attentive voxel feature encoding that outperforms plain VFE and a feature fusion pyramid network to aggregate multi-scale information at feature map level. Experiments on the KITTI benchmark show that a single HVNet achieves the best mAP among all existing methods with a real time inference speed of 31Hz.
Maosheng Ye, Shuangjie Xu, Tongyi Cao
CVPR2
2020 Crowd Counting via Hierarchical Scale Recalibration Network
abstract
The task of crowd counting is extremely challenging due to complicated difficulties, especially the huge variation in vision scale.Previous works tend to adopt a naive concatenation of multiscale information to tackle it, while the scale shifts between the feature maps are ignored.In this paper, we propose a novel Hierarchical Scale Recalibration Network (HSRNet), which addresses the above issues by modeling rich contextual dependencies and recalibrating multiple scale-associated information.Specifically, a Scale Focus Module (SFM) first integrates global context into local features by modeling the semantic inter-dependencies along channel and spatial dimensions sequentially.In order to reallocate channel-wise feature responses, a Scale Recalibration Module (SRM) adopts a step-bystep fusion to generate final density maps.Furthermore, we propose a novel Scale Consistency loss to constrain that the scale-associated outputs are coherent with groundtruth of different scales.With the proposed modules, our approach can ignore various noises selectively and focus on appropriate crowd scales automatically.Extensive experiments on crowd counting datasets (ShanghaiTech, MALL, WorldEXPO'10, and UCSD) show that our HSRNet can deliver superior results over all state-of-the-art approaches.More remarkably, we extend experiments on an extra vehicle dataset , whose results indicate that the proposed model is generalized to other applications.
Zhikang Zou, Yifan Liu 0004, Shuangjie Xu, Wei Wei 0002, Shiping Wen 0001, Pan Zhou 0001
ECAI3
2020 Automated text classification of near-misses from safety reports: An improved deep learning approach
Weili Fang, Hanbin Luo, Shuangjie Xu, Peter E. D. Love, Zhenchuan Lu
Adv. Eng. Informatics3
2020 SAANet: Siamese action-units attention network for improving dynamic facial expression recognition
Daizong Liu, Xi Ouyang, Shuangjie Xu, Pan Zhou 0001, Kun He 0001, Shiping Wen 0001
Neurocomputing3
2019 MHP-VOS: Multiple Hypotheses Propagation for Video Object Segmentation
abstract
We address the problem of semi-supervised video object segmentation (VOS), where the masks of objects of interests are given in the first frame of an input video. To deal with challenging cases where objects are occluded or missing, previous work relies on greedy data association strategies that make decisions for each frame individually. In this paper, we propose a novel approach to defer the decision making for a target object in each frame, until a global view can be established with the entire video being taken into consideration. Our approach is in the same spirit as Multiple Hypotheses Tracking (MHT) methods, making several critical adaptations for the VOS problem. We employ the bounding box (bbox) hypothesis for tracking tree formation, and the multiple hypotheses are spawned by propagating the preceding bbox into the detected bbox proposals within a gated region starting from the initial object mask in the first frame. The gated region is determined by a gating scheme which takes into account a more comprehensive motion model rather than the simple Kalman filtering model in traditional MHT. To further design more customized algorithms tailored for VOS, we develop a novel mask propagation score instead of the appearance similarity score that could be brittle due to large deformations. The mask propagation score, together with the motion score, determines the affinity between the hypotheses during tree pruning. Finally, a novel mask merging strategy is employed to handle mask conflicts between objects. Extensive experiments on challenging datasets demonstrate the effectiveness of the proposed method, especially in the case of object missing.
Shuangjie Xu, Daizong Liu, Linchao Bao, Wei Liu 0005, Pan Zhou 0001
CVPR1
2019 A deep learning-based approach for mitigating falls from height with computer vision: Convolutional neural network
Weili Fang, Botao Zhong, Neng Zhao, Peter E. D. Love, Hanbin Luo, Jiayue Xue, Shuangjie Xu
Adv. Eng. Informatics7
2019 Recognizing people's identity in construction sites with computer vision: A spatial and temporal attention pooling network
Peter E. D. Love, Weili Fang, Hanbin Luo, Shuangjie Xu
Adv. Eng. Informatics5
2017 Jointly Attentive Spatial-Temporal Pooling Networks for Video-Based Person Re-identification
abstract
Person Re-Identification (person re-id) is a crucial task as its applications in visual surveillance and human-computer interaction. In this work, we present a novel joint Spatial and Temporal Attention Pooling Network (ASTPN) for video-based person re-identification, which enables the feature extractor to be aware of the current input video sequences, in a way that interdependency from the matching items can directly influence the computation of each other's representation. Specifically, the spatial pooling layer is able to select regions from each frame, while the attention temporal pooling performed can select informative frames over the sequence, both pooling guided by the information from distance matching. Experiments are conduced on the iLIDS-VID, PRID-2011 and MARS datasets and the results demonstrate that this approach outperforms existing state-of-art methods. We also analyze how the joint pooling in both dimensions can boost the person re-id performance more effectively than using either of them separately 1.
Shuangjie Xu, Yu Cheng 0001, Kang Gu, Yang Yang 0002, Shiyu Chang, Pan Zhou 0001
ICCV1