Shucheng Huang

dblp:75/2069 · DBLP profile ↗
← Back
38ranked-venue papers
5as first author
31since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 20 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 2 · 2 since 2021
YearPublicationVenuePosition
2026 DriveLegal: Toward legally compliant driving via trustworthy hybrid retrieval-augmented LLMs
abstract
• Modular legal-interpretation layer with hybrid vector–graph RAG for AV guidance. • Two datasets: SFT and RAG for multilingual, cross-jurisdiction evaluation. • Hybrid retrieval improves faithfulness and reduces hallucination vs single modes. • Trust module scores context, groundedness, and answer relevance online. • Validated in smart-cabin, V2X intersection monitoring, and offline auditing. Autonomous vehicles (AVs) face persistent challenges in complying with complex and evolving traffic laws. Existing approaches, including rule-based, learning-based, and large language model (LLM) methods, each face limits in adaptability, generalizability, or trustworthiness. We present DriveLegal , a modular legal-interpretation framework for downstream autonomous driving applications. DriveLegal pairs fine-tuned multilingual large language models (LLMs) with an intelligent hybrid retrieval module that routes between vector search and knowledge graph, then returns concise, cited answers. A trust layer scores context relevance, groundedness, and answer relevance and supports continuous improvement through periodic automatic signals and targeted human review. We introduce the DriveLegal datasets for supervised fine-tuning and for retrieval and graph reasoning. Across benchmarks and case studies in smart cabin and vehicle-to-everything (V2X) settings, the hybrid retrieval strategy improves contextual accuracy and reduces hallucination while producing jurisdiction-aware outputs suitable for compliance checks, incident analysis, and reporting.
Shucheng Huang, Chen Sun 0008, Minghao Ning, Changye Ma, Jiaming Zhong, Keqi Shu, Freda Shi, Amir Khajepour
Expert Syst. Appl.1
2026 Mul-FES: A multimodal facial expression spotting method that integrates textual information
Henian Yang, Shucheng Huang
Expert Syst. Appl.2
2026 Fine-grained feature enhancement and occlusion-aware dense pedestrian detection in complex scenes
Shucheng Huang
Multim. Syst.2
2026 A variational causal inference-based method for recognizing object state changes in videos
Wenliang Ge, Shucheng Huang
Multim. Syst.3
2026 LEPD-Net: A Lightweight and Efficient Network for Pedestrian Detection
abstract
The pedestrian detection is crucial in practical applications, such as autonomous driving and video surveillance. However, the existing research mainly focuses on improving detection accuracy, with relatively little attention paid to model complexity and operational efficiency. In scenarios with high real-time requirements, the practical deployment of pedestrian detectors still faces many difficulties. To this end, we propose a lightweight and efficient pedestrian detection network (LEPD-Net). First, we design a PoolFormer-based detection head (PDH) to reduce the model computation and inference time. Second, to compensate for the deficiency of PDH in global context modeling, we design a triple-branch joint attention module (TJAM). TJAM uses only a small number of parameters and strengthens the model's contextual representation by capturing spatial location dependencies and global semantic information between channels. Finally, after incorporating PDH and TJAM into the backbone network, a lightweight and efficient pedestrian detector is constructed. We benchmarked the model on mainstream pedestrian datasets Caltech and CityPersons. The results show that our model achieves the current state-of-the-art performance level. In addition, our model reduces inference time by 25% while maintaining accuracy.
Wenliang Ge, Shucheng Huang, Mingxing Li 0001, Yifan Jiao
IEEE Trans. Neural Networks Learn. Syst.2
2026 Facial core anchoring triangle: enhancing micro-expression spotting through geometric alignment
Henian Yang, Shucheng Huang, Hualong Yu
Vis. Comput.2
2025 Ms-VLPD: A multi-scale VLPD based method for pedestrian detection
Shucheng Huang, Senbao Zhang, Yifan Jiao
Expert Syst. Appl.1
2025 MSOF: A main and secondary bi-directional optical flow feature method for spotting micro-expression
Henian Yang, Shucheng Huang, Mingxing Li 0001
Neurocomputing2
2025 UA-FER: Uncertainty-aware representation learning for facial expression recognition
Haoliang Zhou, Shucheng Huang, Yuqiao Xu
Neurocomputing2
2025 Vehicle Trajectory Prediction Based on Driver's Cognitive Mechanism
abstract
The vehicle trajectory prediction (VTP) is important for autonomous vehicles to make decisions. However, it is challenging because of the inherent uncertainty, dynamic nature, and interactions within driver-vehicle-traffic systems. Moreover, existing methods inadequately balance interpretability and accuracy. Furthermore, few works have been dedicated to analyzing VTP on the drivers cognitive level, while drivers have powerful reasoning abilities due to extensive driving experience and knowledge. To alleviate these, a spatial-temporal graph convolutional network based on domain-knowledge guided learning is proposed by analyzing the drivers cognitive mechanism for VTP on highway. Specifically, a topological graph is constructed to represent the interactions of driving scenarios. A spatial-temporal graph convolutional network, incorporating spatial-temporal attention and domain knowledge, is developed to model the interactions and improve the interpretability. The experimental results suggest that our proposed method outperforms previous methods in both accuracy and interpretability of VTP for the next 5 seconds.
Lisheng Jin, Zejian Deng, Shucheng Huang
IEEE Internet Things J.4
2025 A multi-scale network with multi-view correlation for vehicle re-identification
Shucheng Huang, Yifan Jiao, Mingxing Li 0001
Multim. Syst.2
2025 A Multimodal framework for 3D few-shot class-incremental learning
Senbao Zhang, Shucheng Huang, Li Pengyi, Mingxing Li 0001
Multim. Syst.2
2025 Toward Human-Vehicle Collaboration for Automated Vehicles: A Review and Perspective
abstract
The human-vehicle collaboration in automated vehicles is an effective transitional means to overcome the difficulty of rapidly transitioning to a highly automated level of intelligence. Furthermore, it can fully leverage the strengths of both drivers and autonomous driving systems, embodying a design philosophy of human-centered. Therefore, this paper provides a review and perspectives of the human-vehicle collaboration for automated vehicles. First, the concept, forms and methods of human-vehicle collaboration are reviewed. Then, a human-vehicle mutual trust collaboration framework based on complementary advantages of humans and vehicles and brain-like intelligence is proposed. Specifically, the framework focuses on driver behavior understanding and brain-like cognitive decision planning. After that, the methods of driver behavior understanding and brain-like cognitive decision planning are summarized. Finally, challenges and future works are analyzed to contribute the develop of understandable, trustable, and acceptable human-vehicle collaboration systems.
Qinyu Sun, Zejian Deng, Shucheng Huang, Lisheng Jin
IEEE Trans. Intell. Transp. Syst.5
2025 Multi-Objective Agent-Based Model Predictive Controller for Plug-and-Play Vehicle Control
abstract
Functional integration is a growing trend in vehicle control, often involving the coordination of multiple controllers to achieve various objectives simultaneously. The need for flexibility and reliability has led to a “plug-and-play” approach in control system design, which presents challenges for traditional integrated model predictive control (MPC). Agent-based model predictive control (AMPC) has recently emerged as a distributed solution that treats controllers as agents, creating a collaborative framework among them to reach a common goal. However, this approach struggles to manage distributed conflicting objectives when agents are coupled or interdependent. To address this, we propose a novel, practical distributed control scheme called multi-objective AMPC, which adapts the alternating direction method of multipliers (ADMM) into a general control strategy that approximates global optimization while decoupling objectives. We systematically develop three formulations that maintain convergence while addressing control regularization and inequality constraints, applying them to complex vehicle control systems for the first time. The proposed method has been tested on two vehicle control scenarios with a multi-objective topology. Different formulations are compared through simulations, and the most computationally efficient one was implemented on an electric vehicle for real-world evaluations. The results demonstrate that the proposed multi-objective AMPC can converge approximately to the same global optimum as integrated MPC with greater flexibility and the potential to reduce computational costs.
Jiaming Zhong, Ladan Khoshnevisan, Shucheng Huang, Mohammad Pirani, Yash Pant, Amir Khajepour
IEEE Trans. Intell. Transp. Syst.3
2025 Ellipsoid-SLAM: enhancing dynamic scene understanding through ellipsoidal object representation and trajectory tracking
Haowei Zhu, Suqin Bai, Shucheng Huang
Vis. Comput.6
2025 IOFusion: instance segmentation and optical-flow guided 3D reconstruction in dynamic scenes
Haowei Zhu, Suqin Bai, Chenggen Wang, Yunhan Sun, Shucheng Huang
Vis. Comput.8
2024 MPFC-Net: A multi-perspective feature compensation network for medical image segmentation
Xianghu Wu, Shucheng Huang, Xin Shu 0001, Chunlong Hu, Xiaojun Wu 0001
Expert Syst. Appl.2
2024 Learning shared features from specific and ambiguous descriptions for text-based person search
Qikai Geng, Shucheng Huang, Juanjuan Tu, Hu Lu
Multim. Syst.3
2024 Dual-stream network with cross-layer attention and similarity constraint for micro-expression recognition
Shucheng Huang
Multim. Syst.2
2024 CA-CLIP: category-aware adaptation of CLIP model for few-shot class-incremental learning
Yuqiao Xu, Shucheng Huang, Haoliang Zhou
Multim. Syst.2
2024 Class Incremental Learning for Light-Weighted Networks
abstract
Despite deep neural networks (DNNs) show impressive performance across diverse tasks, they suffer from catastrophic forgetting when dealing with continuous data streams. Incremental learning aims to alleviate this phenomenon and enable DNNs to accumulate new knowledge to cope with the ever-changing world. Recently numerous advanced methods have been developed to enhance the incremental learning capabilities of neural networks. However, these methods mainly focus on the large networks, neglecting the unique needs of edged-device applications, which is surprisingly under-investigated in previous literature. In this paper, we propose two strategies for transferring knowledge from large teacher networks to light-weighted networks in class incremental learning. Specifically, in cases where the initial task contains a large number of categories, our static teacher strategy involves transferring knowledge from the teacher to the student network on the initial task to enhance the plasticity of the student network, and applying regularization constraints on the subsequent task to improve its stability. In a more challenging scenario where each task includes an equal number of categories, the dynamic teacher strategy continuously guides the student network on each task. We evaluate the proposed methods on CIFAR100, Tiny-ImageNet and ImageNet-subset datasets with different types of light-weighted networks (MobileNet, ShuffleNet). We observed that effective knowledge transfer resulting in the student network achieving performance comparable or even outperform the teacher network. Extensive and detailed experiments conducted on three datasets demonstrated the simplicity and effectiveness of our proposed method. Comprehensive analysis are also conducted including different factors and visualization.
Zhe Tao, Lu Yu 0004, Hantao Yao, Shucheng Huang, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.4
2024 CEPrompt: Cross-Modal Emotion-Aware Prompting for Facial Expression Recognition
abstract
Facial expression recognition (FER) remains a challenging task due to the ambiguity and subtlety of expressions. To address this challenge, current FER methods predominantly prioritize visual cues while inadvertently neglecting the potential insights that can be gleaned from other modalities. Recently, vision-language pre-training (VLP) models integrated textual cues as guidance, culminating in a powerful multi-modal solution that has proven effective for a range of computer vision tasks. In this paper, we propose a Cross-Modal Emotion-Aware Prompting (CEPrompt) framework for FER based on VLP models. To make VLP models sensitive to expression-relevant visual discrepancies, we devise an Emotion Conception-guided Visual Adapter (EVA) to capture the category-specific appearance representations with emotion conception guidance. Moreover, knowledge distillation is employed to prevent the model from forgetting the pre-trained category-invariant knowledge. In addition, we design a Conception-Appearance Tuner (CAT) to facilitate the interaction of multi-modal information via cooperatively tuning between emotion conception and appearance prompts. In this way, semantic information about emotion text conception is infused directly into facial appearance images, thereby enhancing a comprehensive and precise understanding of expression-related facial details. Quantitative and qualitative experiments show that our CEPrompt outperforms state-of-the-art approaches on three real-world FER datasets. The code is available athttps://github.com/HaoliangZhou/CEPrompt.
Haoliang Zhou, Shucheng Huang, Feifei Zhang 0001, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.2
2024 Learning Commonsense-aware Moment-Text Alignment for Fast Video Temporal Grounding
abstract
Grounding temporal video segments described in natural language queries effectively and efficiently is a crucial capability needed in vision-and-language fields. In this article, we deal with the fast video temporal grounding (FVTG) task, aiming at localizing the target segment with high speed and favorable accuracy. Most existing approaches adopt elaborately designed cross-modal interaction modules to improve the grounding performance, which suffer from the test-time bottleneck. Although several common space-based methods enjoy the high-speed merit during inference, they can hardly capture the comprehensive and explicit relations between visual and textual modalities. In this article, to tackle the dilemma of the speed–accuracy tradeoff, we propose a commonsense-aware cross-modal alignment network (C 2 AN) that incorporates commonsense-guided visual and text representations into a complementary common space for fast video temporal grounding. Specifically, the commonsense concepts are explored and exploited by extracting the structural semantic information from a language corpus. Then, a commonsense-aware interaction module is designed to obtain bridged visual and text features by utilizing the learned commonsense concepts. Finally, to maintain the original semantic information of textual queries, a cross-modal complementary common space is optimized to obtain matching scores for performing FVTG. Extensive results on two challenging benchmarks show that our C 2 AN method performs favorably against states of the art while running at high speed. Our code is available at https://github.com/ZiyueWu59/CCA
Ziyue Wu, Junyu Gao 0002, Shucheng Huang, Changsheng Xu
ACM Trans. Multim. Comput. Commun. Appl.3
2023 RES-CapsNet: an improved capsule network for micro-expression recognition
Xin Shu 0001, Shucheng Huang
Multim. Syst.4
2023 Shallow multi-branch attention convolutional neural network for micro-expression recognition
Shucheng Huang, Zhe Tao
Multim. Syst.2
2023 A LiDAR point cloud registration method combining linear feature extraction and TrICP algorithm
Chuanwang Wen, Shucheng Huang
Multim. Syst.2
2023 Inceptr: micro-expression recognition integrating inception-CBAM and vision transformer
Haoliang Zhou, Shucheng Huang, Yuqiao Xu
Multim. Syst.2
2022 Intentional-Deception Detection Based on Facial Muscle Movements in an Interactive Social Context
abstract
Micro-expression, which is generated by facial muscle movements, could be a crucial cue for deception detection. In the existed research investigating the relationship between facial muscles and deception detection, researchers have focused almost exclusively on two muscles, i.e., zygomaticus and corrugator supercilii, based on the theoretical basis that they are highly associated with positive and negative expressions. However, the aim of this study is to demonstrate the direct relationship between facial muscle movements and deception detection. Addressing this issue, this paper proposes an experimental paradigm with high ecological validity that uses electromyography (EMG) signals to precisely examine the role of facial muscle movements in deception detection. Moreover, we propose a vector-based sequential forward selection (VSFS) algorithm to identify the muscle (or muscle combination) most closely associated with lying. Based on our proposed approach, the importance of seven selected facial muscles was explored by comparing the corresponding facial EMG (fEMG) between truth and lying conditions. First, the present study found that the zygomaticus and corrugator supercilii could play important roles in deception detection, and our findings are consistent with existed research. Second, the experiment result verified that the muscles related to deception detection were consistent with those with higher frequency occurring in micro-expression. Moreover, the present study provides a theoretical basis that intelligent micro-expressions analysis could improve the lie detection performance by focusing on the area of the forehead, eyebrows, and cheeks.
Zizhao Dong, Shaoyuan Lu, Luyao Dai, Shucheng Huang, Ye Liu 0010
Pattern Recognit. Lett.5
2021 Diving Into The Relations: Leveraging Semantic and Visual Structures For Video Moment Retrieval
abstract
Existing dominant approaches for video moment retrieval task are to learn semantic correlation between a given query and the video. However, these methods rarely explore the fine-grained semantic structure and comprehensive visual structure, leading to insufficient utilization of textual and visual relations. In this paper, we propose a unified framework for video moment retrieval, which considers to simultaneously encode semantic and visual structures. Specifically, a semantic role tree is built to reveal the fine-grained semantic information by generating hierarchical textual embeddings. Then the semantic structure is adopted to facilitate the visual structure learning with a contextual attention-based proposal interaction module. Finally, we adaptively aggregate and obtain the visual-semantic matching information through a multi-level fusion strategy to select the best matching moment proposal. Extensive experiments on two popular benchmarks (Charades-STA and ActivityNet Captions) show that our proposed method achieves state-of-the-art performance. Codes are available in the Supplementary Material.
Ziyue Wu, Junyu Gao 0002, Shucheng Huang, Changsheng Xu
ICME3
2021 Face spoofing detection based on chromatic ED-LBP texture feature
Xin Shu 0001, Shucheng Huang
Multim. Syst.3
2021 Multiple channels local binary pattern for color texture representation and classification
Xin Shu 0001, Zhigang Song, Shucheng Huang, Xiaojun Wu 0001
Signal Process. Image Commun.4
2020 Deep Sentiment Classification and Topic Discovery on Novel Coronavirus or COVID-19 Online Discussions: NLP Using LSTM Recurrent Neural Network Approach
abstract
Internet forums and public social media, such as online healthcare forums, provide a convenient channel for users (people/patients) concerned about health issues to discuss and share information with each other. In late December 2019, an outbreak of a novel coronavirus (infection from which results in the disease named COVID-19) was reported, and, due to the rapid spread of the virus in other parts of the world, the World Health Organization declared a state of emergency. In this paper, we used automated extraction of COVID-19-related discussions from social media and a natural language process (NLP) method based on topic modeling to uncover various issues related to COVID-19 from public opinions. Moreover, we also investigate how to use LSTM recurrent neural network for sentiment classification of COVID-19 comments. Our findings shed light on the importance of using public opinions and suitable computational techniques to understand issues surrounding COVID-19 and to guide related decision-making. In addition, experiments demonstrated that the research model achieved an accuracy of 81.15% - a higher accuracy than that of several other well-known machine-learning algorithms for COVID-19-Sentiment Classification.
Hamed Jelodar, Yongli Wang 0002, Rita Orji, Shucheng Huang
IEEE J. Biomed. Health Informatics4
2019 Video Highlight Detection via Region-Based Deep Ranking Model
abstract
The video highlight detection task is to localize key elements (moments of user’s major or special interest) in a video. Most of the existing highlight detection approaches extract features from the video segment as a whole without considering the difference of local features spatially. In spatial extent, not all regions are worth watching because some of them only contain the background of the environment without human or other moving objects, especially when there is lots of clutter in the background. To deal with this issue, we propose a novel region-based model which can automatically localize the key elements in a video without any extra supervised annotations. Specifically, the proposed model produces position-sensitive score maps for local regions in the spatial dimension of the video segment, and then aggregates all position-wise scores with position-pooling operation. The regions with higher response values will be extracted as key elements. Thus more effective features of the video segment are obtained to predict the highlight score. The proposed position-sensitive scheme can be easily integrated into an end-to-end fully convolutional network which aims to update parameters via stochastic gradient descent method in the backward propagation to improve the robustness of the model. Extensive experimental results on the YouTube and SumMe datasets demonstrate that the proposed approach achieves significant improvement over state-of-the-art methods.
Yifan Jiao, Tianzhu Zhang 0001, Shucheng Huang, Bin Liu 0014, Changsheng Xu
Int. J. Pattern Recognit. Artif. Intell.3
2018 Three-Dimensional Attention-Based Deep Ranking Model for Video Highlight Detection
abstract
The video highlight detection task is to localize key elements (moments of user's major or special interest) in a video. Most of existing highlight detection approaches extract features from the video segment as a whole without considering the difference of local features both temporally and spatially. Due to the complexity of video content, this kind of mixed features will impact the final highlight prediction. In temporal extent, not all frames are worth watching because some of them only contain the background of the environment without human or other moving objects. In spatial extent, it is similar that not all regions in each frame are highlights especially when there are lots of clutters in the background. To solve the above problem, we propose a novel three-dimensional (3-D) (spatial+temporal) attention model that can automatically localize the key elements in a video without any extra supervised annotations. Specifically, the proposed attention model produces attention weights of local regions along both the spatial and temporal dimensions of the video segment. The regions of key elements in the video will be strengthened with large weights. Thus, the more effective feature of the video segment is obtained to predict the highlight score. The proposed 3-D attention scheme can be easily integrated into a conventional end-to-end deep ranking model that aims to learn a deep neural network to compute the highlight score of each video segment. Extensive experimental results on the YouTube and SumMe datasets demonstrate that the proposed approach achieves significant improvement over state-of-the-art methods. With the proposed 3-D attention model, video highlights can be accurately retrieved in spatial and temporal dimensions without human supervision in several domains, such as gymnastics, parkour, skating, skiing, surfing, and dog activities, on the public datasets.
Yifan Jiao, Zhetao Li, Shucheng Huang, Xiaoshan Yang, Bin Liu 0014, Tianzhu Zhang 0001
IEEE Trans. Multim.3
2017 Video Highlight Detection via Deep Ranking Modeling
Yifan Jiao, Xiaoshan Yang, Tianzhu Zhang 0001, Shucheng Huang, Changsheng Xu
PSIVT4
2016 Multi-object tracking via discriminative appearance modeling
Shucheng Huang
Comput. Vis. Image Underst.1
2016 Exponential Discriminant Locality Preserving Projection for face recognition
Shucheng Huang, Zhuang Lu
Neurocomputing1
2007 An active learning system for mining time-changing data streams
Shucheng Huang, Yisheng Dong
Intell. Data Anal.1