Hui Shuai

dblp:214/6233 · DBLP profile ↗
← Back
23ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0001-8840-5069ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 VLM-driven fine-grained semantic regularization for low-light image enhancement
Zixuan Sun, Chuanwei Zhou, Hui Shuai, Qingshan Liu 0001
Multim. Syst.3
2026 Adaptive frequency collaboration for remote sensing change detection
Feng Zhou 0006, Hui Shuai, Qingshan Liu 0001, Renlong Hang
Neural Networks3
2026 Enhanced Spatiotemporal Consistency for Image-to-LiDAR Data Pretraining
abstract
LiDAR representation learning has emerged as a promising approach to reducing reliance on costly and labor-intensive human annotations. While existing methods primarily focus on spatial alignment between LiDAR and camera sensors, they often overlook the temporal dynamics critical for capturing motion and scene continuity in driving scenarios. To address this limitation, we propose SuperFlow++, a novel framework that integrates spatiotemporal cues in both pretraining and downstream tasks using consecutive LiDAR-camera pairs. SuperFlow++ introduces four key components: (1) a view consistency alignment module to unify semantic information across camera views, (2) a dense-to-sparse consistency regularization mechanism to enhance feature robustness across varying point cloud densities, (3) a flow-based contrastive learning approach that models temporal relationships for improved scene understanding, and (4) a temporal voting strategy that propagates semantic information across LiDAR scans to improve prediction consistency. Extensive evaluations on 11 heterogeneous LiDAR datasets demonstrate that SuperFlow++ outperforms state-of-the-art methods across diverse tasks and driving conditions. Furthermore, by scaling both 2D and 3D backbones during pretraining, we uncover emergent properties that provide deeper insights into developing scalable 3D foundation models. With strong generalizability and computational efficiency, SuperFlow++ establishes a new benchmark for data-efficient LiDAR-based perception in autonomous driving.
Xiang Xu 0009, Lingdong Kong, Hui Shuai, Liang Pan, Kai Chen 0026, Ziwei Liu 0002, Qingshan Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Adversarial Pruning Networks for Compact 3D Gaussian Splatting
abstract
3D Gaussian Splatting holds significant potential for high-quality visual scene rendering. However, the large number of Gaussian primitives it requires poses challenges in memory consumption and practical deploy. Existing methods often rely on empirical criteria to prune Gaussians, which inevitably compromises visual quality. To address this, we propose Adversarial Pruning Networks (APNet), a framework that employs adversarial learning to balances the reduction of redundant Gaussians with the preservation of visual fidelity. APNet comprises a Gaussian Learning and Pruning Network (GLPN) and a Discriminative Network. GLPN incorporates the geometric information into the learning of Gaussians and prunes these Gaussians through a data-driven mask. Meanwhile, the Discriminative Network is trained to distinguish between synthesized and real images, acting as an adversary. Through adversarial pruning, APNet significantly reduces the number of Gaussians while rendering high-quality images. Extensive experiments on the Mip-NeRF360, Tanks & Temples, and Deep Blending datasets demonstrate that APNet achieves up to a 90% reduction in the original 3DGS while maintaining high rendering quality.
Hui Shuai, Yubao Sun, Qingshan Liu 0001
IEEE Trans. Multim.1
2025 LiMoE: Mixture of LiDAR Representation Learners from Automotive Scenes
abstract
LiDAR data pretraining offers a promising approach to leveraging large-scale, readily available datasets for enhanced data utilization. However, existing methods predominantly focus on sparse voxel representation, overlooking the complementary attributes provided by other LiDAR representations. In this work, we propose ${\color{Red}\text{Li}}{\color{Green}\text{MoE}}$, a framework that integrates the Mixture of Experts (MoE) paradigm into LiDAR data representation learning to synergistically combine multiple representations, such as range images, sparse voxels, and raw points. Our approach consists of three stages: i) Image-to-LiDAR Pretraining, which transfers prior knowledge from images to point clouds across different representations; ii) Contrastive Mixture Learning (CML), which uses MoE to adaptively activate relevant attributes from each representation and distills these mixed features into a unified 3D network; iii) Semantic Mixture Supervision (SMS), which combines semantic logits from multiple representations to boost downstream segmentation performance. Extensive experiments across eleven large-scale LiDAR datasets demonstrate our effectiveness and superiority. The code has been made publicly accessible.
Xiang Xu 0009, Lingdong Kong, Hui Shuai, Liang Pan, Ziwei Liu 0002, Qingshan Liu 0001
CVPR3
2025 FRNet: Frustum-Range Networks for Scalable LiDAR Segmentation
abstract
LiDAR segmentation has become a crucial component of advanced autonomous driving systems. Recent range-view LiDAR segmentation approaches show promise for real-time processing. However, they inevitably suffer from corrupted contextual information and rely heavily on post-processing techniques for prediction refinement. In this work, we propose FRNet, a simple yet powerful method aimed at restoring the contextual information of range image pixels using corresponding frustum LiDAR points. First, a frustum feature encoder module is used to extract per-point features within the frustum region, which preserves scene consistency and is critical for point-level predictions. Next, a frustum-point fusion module is introduced to update per-point features hierarchically, enabling each point to extract more surrounding information through the frustum features. Finally, a head fusion module is used to fuse features at different levels for final semantic predictions. Extensive experiments conducted on four popular LiDAR segmentation benchmarks under various task setups demonstrate the superiority of FRNet. Notably, FRNet achieves 73.3% and 82.5% mIoU scores on the testing sets of SemanticKITTI and nuScenes. While achieving competitive performance, FRNet operates 5 times faster than state-of-the-art approaches. Such high efficiency opens up new possibilities for more scalable LiDAR segmentation. The code has been made publicly available at https://github.com/Xiangxu-0103/FRNet.
Xiang Xu 0009, Lingdong Kong, Hui Shuai, Qingshan Liu 0001
IEEE Trans. Image Process.3
2024 4D Contrastive Superflows are Dense 3D Representation Learners
Xiang Xu 0009, Lingdong Kong, Hui Shuai, Liang Pan, Kai Chen 0026, Ziwei Liu 0002, Qingshan Liu 0001
ECCV (1)3
2023 ACLM: Adaptive Compensatory Label Mining for Facial Expression Recognition
Chengguang Liu, Shanmin Wang, Hui Shuai, Qingshan Liu 0001
ICIG (4)3
2023 Adaptive Multi-View and Temporal Fusing Transformer for 3D Human Pose Estimation
abstract
This article proposes a unified framework dubbed Multi-view and Temporal Fusing Transformer (MTF-Transformer) to adaptively handle varying view numbers and video length without camera calibration in 3D Human Pose Estimation (HPE). It consists of Feature Extractor, Multi-view Fusing Transformer (MFT), and Temporal Fusing Transformer (TFT). Feature Extractor estimates 2D pose from each image and fuses the prediction according to the confidence. It provides pose-focused feature embedding and makes subsequent modules computationally lightweight. MFT fuses the features of a varying number of views with a novel Relative-Attention block. It adaptively measures the implicit relative relationship between each pair of views and reconstructs more informative features. TFT aggregates the features of the whole sequence and predicts 3D pose via a transformer. It adaptively deals with the video of arbitrary length and fully unitizes the temporal information. The migration of transformers enables our model to learn spatial geometry better and preserve robustness for varying application scenarios. We report quantitative and qualitative results on the Human3.6M, TotalCapture, and KTH Multiview Football II. Compared with state-of-the-art methods with camera parameters, MTF-Transformer obtains competitive results and generalizes well to dynamic capture with an arbitrary number of unseen views.
Hui Shuai, Lele Wu, Qingshan Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Bias-Based Soft Label Learning for Facial Expression Recognition
abstract
Facial Expression Recognition (FER) suffers from misrecognition due to the similarities between expressions. To address this issue, popular works replace original annotations with soft labels to reflect expression similarities. However, existing soft label learning (SLL) modules are independent of FER modules. In this article, inspired by automatic control theory, we propose a bias-based soft label learning network for FER named EC-Net. For optimizing FER and SLL modules jointly, EC-Net constitutes the closed-loop feedback between the two modules by designing a module measuring and transmitting the bias between FER module predictions and target labels. Specifically, EC-Net contains three modules: E-subNet, C-subNet, and L-Transmitter. First, E-subNet, i.e., the FER module, attempts to converge to target labels under the supervision of soft labels, acting as the executor. Then, L-Transmitter measures the bias between E-subNet predictions and target labels. It converts multiple discrete biases to the bias-based label through spectral clustering and transmits it to C-subNet. Finally, C-SubNet, i.e., the SLL module, generates soft labels from the bias-based label with a cascaded learner and progressively distinguishes similar expressions. It updates the learned soft labels for E-subNet, performing like the controller. Supervised by the bias-based soft label, E-subNet effectively reduces the dominant bias caused by similar expressions. We conduct extensive experiments on four popular benchmarks, demonstrating the effectiveness of applying closed-loop feedback in the FER task.
Shanmin Wang, Hui Shuai, Chengguang Liu, Qingshan Liu 0001
IEEE Trans. Affect. Comput.2
2023 Geometry-Injected Image-Based Point Cloud Semantic Segmentation
abstract
Image-based methods have replicated the success from 2D domain to 3D point cloud semantic segmentation. However, when we directly apply 2D techniques to the projected pseudo-image, inherent differences between the point cloud and the image cause geometric distortion. This paper analyzes the geometric distortion between the point cloud and the pseudo-image, including truncation, dislocation, and hole. To ensure geometric fidelity, we propose the Geometry-injected Image-based point cloud semantic segmentation Network (GINet). We design a Cyclic Convolution to optimize the convolution operation, dealing with truncation. For dislocation and hole, we propose Dual Geometric Constraints, including Local Spatial Attention and Local Affinity Regularization, to incorporate the geometric information into semantic feature learning. Local Spatial Attention generates an attention map from the point coordinates to modulate the feature map before convolution. Local Affinity Regularization supervises the semantic similarity of pixels in the convolution kernel range. GINet rectifies the geometric distortion with these mechanisms while taking advantage of the successful 2D semantic segmentation methods. Quantitative and qualitative experiments on SemanticKITTI and SemanticPOSS demonstrate the effectiveness of GINet.
Hui Shuai, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Multilevel Spatial-Temporal Excited Graph Network for Skeleton-Based Action Recognition
abstract
The ability to capture joint connections in complicated motion is essential for skeleton-based action recognition. However, earlier approaches may not be able to fully explore this connection in either the spatial or temporal dimension due to fixed or single-level topological structures and insufficient temporal modeling. In this paper, we propose a novel multilevel spatial-temporal excited graph network (ML-STGNet) to address the above problems. In the spatial configuration, we decouple the learning of the human skeleton into general and individual graphs by designing a multilevel graph convolution (ML-GCN) network and a spatial data-driven excitation (SDE) module, respectively. ML-GCN leverages joint-level, part-level, and body-level graphs to comprehensively model the hierarchical relations of a human body. Based on this, SDE is further introduced to handle the diverse joint relations of different samples in a data-dependent way. This decoupling approach not only increases the flexibility of the model for graph construction but also enables the generality to adapt to various data samples. In the temporal configuration, we apply the concept of temporal difference to the human skeleton and design an efficient temporal motion excitation (TME) module to highlight the motion-sensitive features. Furthermore, a simplified multiscale temporal convolution (MS-TCN) network is introduced to enrich the expression ability of temporal features. Extensive experiments on the four popular datasets NTU-RGB+D, NTU-RGB+D 120, Kinetics Skeleton 400, and Toyota Smarthome demonstrate that ML-STGNet gains considerable improvements over the existing state of the art.
Yisheng Zhu, Hui Shuai, Guangcan Liu, Qingshan Liu 0001
IEEE Trans. Image Process.2
2023 Multi-level feature fusion pyramid network for object detection
Zebin Guo, Hui Shuai, Guangcan Liu, Yisheng Zhu
Vis. Comput.2
2022 Waterfall-Net: Waterfall Feature Aggregation for Point Cloud Semantic Segmentation
Hui Shuai, Xiang Xu 0009, Qingshan Liu 0001
PRCV (3)1
2022 A globally convergent approximate Newton method for non-convex sparse learning
Fanfan Ji, Hui Shuai, Xiao-Tong Yuan
Pattern Recognit.2
2022 Phase Space Reconstruction Driven Spatio-Temporal Feature Learning for Dynamic Facial Expression Recognition
abstract
Automatic Dynamic Facial Expression Recognition (DFER) is a challenging task, since how to effectively capture facial temporal dynamics is still an open problem. In this article, we regard variations of facial expressions as a dynamic system in accord with certain rules, and try to explore the fundamental temporal properties for recognizing dynamic expressions. Inspired by the phase space reconstruction method for time series analysis, we propose a novel network named Phase Space Reconstruction Network (PSRNet) for learning spatio-temporal features of facial expressions. First, 3D convolutional neural networks are used to extract spatial and short-term temporal features, which indicate the state of each frame and are termed as observations in the phase space. All the observations compose the trajectory of the dynamical system. Then, a data-driven across-correlation matrix is inferred to reveal the relationship of the observations. With this matrix, the phase space reconstruction module reconstructs the trajectory by aggregating the observations adaptively in the phase space. Reconstructed observations represent the gradual process of dynamic facial expressions, which is beneficial to recognize these expressions. The experiment results on three databases (Oulu, MMI, and CK+) demonstrate that the proposed PSRNet can extract more informative and representative spatio-temporal features for DFER. Moreover, the visualization of intermediate features reveals that the reconstructed features have global consistency in facial regions and the underlying evolutionary pattern of dynamic facial expression.
Shanmin Wang, Hui Shuai, Qingshan Liu 0001
IEEE Trans. Affect. Comput.2
2022 Self-Supervised Video Representation Learning Using Improved Instance-Wise Contrastive Learning and Deep Clustering
abstract
Instance-wise contrastive learning (Instance-CL), which learns to map similar instances closer and different instances farther apart in the embedding space, has achieved considerable progress in self-supervised video representation learning. However, canonical Instance-CL does not handle properly the temporal similarities between different videos, limiting the representation capabilities of learned models. This paper presents a novel two-stage framework that combines Instance-CL and unsupervised clustering to progressively learn desirable temporal representations with high intra-class compactness. Specifically, (a) we first introduce a new consistency-preserving sampling strategy to generate positive/negative pairs. Compared to the traditional sampling methods, our sampling strategy focuses more on motion dynamics, resulting in more temporal-related feature representations. (b) To further explore the temporal similarities between videos so as to encourage intra-class compactness, we set temporal representations extracted from Instance-CL as an initializer, and iteratively use k-means clustering to generate pseudo-labels for training the encoder. We term our method as Improved Instance-CL with Deep Clustering (ICDC) and apply it to two downstream tasks, including action recognition and video retrieval. Extensive experimental results show that ICDC gains considerable improvements compared to the existing self-supervised methods.
Yisheng Zhu, Hui Shuai, Guangcan Liu, Qingshan Liu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Hierarchical Context Network for Airborne Image Segmentation
abstract
Most of the recent methods focus on capturing contextual information by measuring relations (e.g., feature similarity) between each pixel and all the others for airborne image segmentation. Nevertheless, these methods have difficulty in handling confusing objects with a partially similar appearance. In this article, we attempt to simultaneously explore pixel-to-pixel (P2P) and pixel-to-object (P2O) relations to learn contextual information. For this purpose, a hierarchical context network (HCNet) is proposed. It consists of a P2P subnetwork and a P2O subnetwork. The P2P subnetwork learns the P2P relation (detail-grained context) for better preservation of the details (e.g., boundary) of the objects. Meanwhile, the P2O subnetwork models the P2O relation (semantic-grained context), aiming at improving the intraobject semantic consistency. When inferring the segmentation results, outputs of these two subnetworks are aggregated to obtain the hierarchical contextual information. Experimental results demonstrate that the proposed model achieves competitive performance on three challenging benchmarks.
Feng Zhou 0006, Renlong Hang, Hui Shuai, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.3
2021 Backward Attentive Fusing Network With Local Aggregation Classifier for 3D Point Cloud Semantic Segmentation
abstract
In this paper, a Backward Attentive Fusing Network with Local Aggregation Classifier (BAF-LAC) is proposed to improve the performance of 3D point cloud semantic segmentation. It consists of a Backward Attentive Fusing Encoder-Decoder (BAF-ED) to learn semantic features and a Local Aggregation Classifier (LAC) to maintain the context-awareness of points. BAF-ED narrows the semantic gap between the encoder and the decoder via fusing multi-layer encoder features with the decoder features. High-level encoder features are transformed into an attention map to modulate low-level encoder features backward. LAC adaptively enhances the intermediate features in point-wise MLPs via aggregating the features of neighboring points into the center point. It takes the place of commonly used post-processing techniques and retains context consistency into the classifier. Equipped with these modules, BAF-LAC can extract discriminative semantic features and predict smoother results. Extensive experiments on Semantic3D, SemanticKITTI, and S3DIS demonstrate that the proposed method can achieve competitive results against the state-of-the-art methods.
Hui Shuai, Xiang Xu 0009, Qingshan Liu 0001
IEEE Trans. Image Process.1
2020 Flow driven attention network for video salient object detection
abstract
Salient object detection has been revolutionised by convolutional neural network (CNN) recently. However, it is hard to transfer the state‐of‐the‐art still‐image based saliency detectors to videos directly, owing to the neglect of temporal contexts between frames. In this study, the authors propose a flow‐driven attention network (FDAN) to exploit motion information for video salient object detection. FDAN consists of an appearance feature extractor, a motion‐guided attention module and a saliency map regression module. It extracts the appearance feature per frame, refines appearance feature with optical flow and infers the ultimate saliency map, respectively. Motion‐guided attention module is the core of FDAN, which extracts motion information in the form of attention. This attention mechanism is a two‐branch CNN, fusing optical flow and appearance features. In addition, a shortcut connection is applied to the attention multiplied feature map for noise suppression intensively. Experimental results show that the proposed method can achieve performance on par with the state‐of‐the‐art method flow‐guided recurrent neural encoder on challenging benchmarks of Densely Annotated Video Segmentation and Freiburg–Berkeley Motion Segmentation while being two times faster in detection.
Feng Zhou 0006, Hui Shuai, Qingshan Liu 0001, Guodong Guo
IET Image Process.2
2019 PointNet-Based Channel Attention VLAD Network
Rongrong Fan, Hui Shuai, Qingshan Liu 0001
PRCV (3)2
2019 Quadratic Approximation Greedy Pursuit for Cardinality-Constrained Sparse Learning
Fanfan Ji, Hui Shuai, Xiao-Tong Yuan
PRCV (1)2
2019 Pruning Convolutional Neural Networks via Stochastic Gradient Hard Thresholding
Haiwei Lu, Hui Shuai, Xiao-Tong Yuan
PRCV (1)3