VLDB 2026 Research / reviewers in the wild / expert
Shucheng Huang
dblp:75/2069
· DBLP profile ↗
38ranked-venue papers
5as first author
31since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 20 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DriveLegal: Toward legally compliant driving via trustworthy hybrid retrieval-augmented LLMsabstract• Modular legal-interpretation layer with hybrid vector–graph RAG for AV guidance. • Two datasets: SFT and RAG for multilingual, cross-jurisdiction evaluation. • Hybrid retrieval improves faithfulness and reduces hallucination vs single modes. • Trust module scores context, groundedness, and answer relevance online. • Validated in smart-cabin, V2X intersection monitoring, and offline auditing. Autonomous vehicles (AVs) face persistent challenges in complying with complex and evolving traffic laws. Existing approaches, including rule-based, learning-based, and large language model (LLM) methods, each face limits in adaptability, generalizability, or trustworthiness. We present DriveLegal , a modular legal-interpretation framework for downstream autonomous driving applications. DriveLegal pairs fine-tuned multilingual large language models (LLMs) with an intelligent hybrid retrieval module that routes between vector search and knowledge graph, then returns concise, cited answers. A trust layer scores context relevance, groundedness, and answer relevance and supports continuous improvement through periodic automatic signals and targeted human review. We introduce the DriveLegal datasets for supervised fine-tuning and for retrieval and graph reasoning. Across benchmarks and case studies in smart cabin and vehicle-to-everything (V2X) settings, the hybrid retrieval strategy improves contextual accuracy and reduces hallucination while producing jurisdiction-aware outputs suitable for compliance checks, incident analysis, and reporting. Shucheng Huang, Chen Sun 0008, Minghao Ning, Changye Ma, Jiaming Zhong, Keqi Shu, Freda Shi, Amir Khajepour |
Expert Syst. Appl. | 1 |
| 2026 | Mul-FES: A multimodal facial expression spotting method that integrates textual information
Henian Yang, Shucheng Huang |
Expert Syst. Appl. | 2 |
| 2026 | Fine-grained feature enhancement and occlusion-aware dense pedestrian detection in complex scenes
Shucheng Huang |
Multim. Syst. | 2 |
| 2026 | A variational causal inference-based method for recognizing object state changes in videos
Wenliang Ge, Shucheng Huang |
Multim. Syst. | 3 |
| 2026 | LEPD-Net: A Lightweight and Efficient Network for Pedestrian DetectionabstractThe pedestrian detection is crucial in practical applications, such as autonomous driving and video surveillance. However, the existing research mainly focuses on improving detection accuracy, with relatively little attention paid to model complexity and operational efficiency. In scenarios with high real-time requirements, the practical deployment of pedestrian detectors still faces many difficulties. To this end, we propose a lightweight and efficient pedestrian detection network (LEPD-Net). First, we design a PoolFormer-based detection head (PDH) to reduce the model computation and inference time. Second, to compensate for the deficiency of PDH in global context modeling, we design a triple-branch joint attention module (TJAM). TJAM uses only a small number of parameters and strengthens the model's contextual representation by capturing spatial location dependencies and global semantic information between channels. Finally, after incorporating PDH and TJAM into the backbone network, a lightweight and efficient pedestrian detector is constructed. We benchmarked the model on mainstream pedestrian datasets Caltech and CityPersons. The results show that our model achieves the current state-of-the-art performance level. In addition, our model reduces inference time by 25% while maintaining accuracy. Wenliang Ge, Shucheng Huang, Mingxing Li 0001, Yifan Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2026 | Facial core anchoring triangle: enhancing micro-expression spotting through geometric alignment
Henian Yang, Shucheng Huang, Hualong Yu |
Vis. Comput. | 2 |
| 2025 | Ms-VLPD: A multi-scale VLPD based method for pedestrian detection
Shucheng Huang, Senbao Zhang, Yifan Jiao |
Expert Syst. Appl. | 1 |
| 2025 | MSOF: A main and secondary bi-directional optical flow feature method for spotting micro-expression
Henian Yang, Shucheng Huang, Mingxing Li 0001 |
Neurocomputing | 2 |
| 2025 | UA-FER: Uncertainty-aware representation learning for facial expression recognition
Haoliang Zhou, Shucheng Huang, Yuqiao Xu |
Neurocomputing | 2 |
| 2025 | Vehicle Trajectory Prediction Based on Driver's Cognitive MechanismabstractThe vehicle trajectory prediction (VTP) is important for autonomous vehicles to make decisions. However, it is challenging because of the inherent uncertainty, dynamic nature, and interactions within driver-vehicle-traffic systems. Moreover, existing methods inadequately balance interpretability and accuracy. Furthermore, few works have been dedicated to analyzing VTP on the drivers cognitive level, while drivers have powerful reasoning abilities due to extensive driving experience and knowledge. To alleviate these, a spatial-temporal graph convolutional network based on domain-knowledge guided learning is proposed by analyzing the drivers cognitive mechanism for VTP on highway. Specifically, a topological graph is constructed to represent the interactions of driving scenarios. A spatial-temporal graph convolutional network, incorporating spatial-temporal attention and domain knowledge, is developed to model the interactions and improve the interpretability. The experimental results suggest that our proposed method outperforms previous methods in both accuracy and interpretability of VTP for the next 5 seconds. Lisheng Jin, Zejian Deng, Shucheng Huang |
IEEE Internet Things J. | 4 |
| 2025 | A multi-scale network with multi-view correlation for vehicle re-identification
Shucheng Huang, Yifan Jiao, Mingxing Li 0001 |
Multim. Syst. | 2 |
| 2025 | A Multimodal framework for 3D few-shot class-incremental learning
Senbao Zhang, Shucheng Huang, Li Pengyi, Mingxing Li 0001 |
Multim. Syst. | 2 |
| 2025 | Toward Human-Vehicle Collaboration for Automated Vehicles: A Review and PerspectiveabstractThe human-vehicle collaboration in automated vehicles is an effective transitional means to overcome the difficulty of rapidly transitioning to a highly automated level of intelligence. Furthermore, it can fully leverage the strengths of both drivers and autonomous driving systems, embodying a design philosophy of human-centered. Therefore, this paper provides a review and perspectives of the human-vehicle collaboration for automated vehicles. First, the concept, forms and methods of human-vehicle collaboration are reviewed. Then, a human-vehicle mutual trust collaboration framework based on complementary advantages of humans and vehicles and brain-like intelligence is proposed. Specifically, the framework focuses on driver behavior understanding and brain-like cognitive decision planning. After that, the methods of driver behavior understanding and brain-like cognitive decision planning are summarized. Finally, challenges and future works are analyzed to contribute the develop of understandable, trustable, and acceptable human-vehicle collaboration systems. Qinyu Sun, Zejian Deng, Shucheng Huang, Lisheng Jin |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Multi-Objective Agent-Based Model Predictive Controller for Plug-and-Play Vehicle ControlabstractFunctional integration is a growing trend in vehicle control, often involving the coordination of multiple controllers to achieve various objectives simultaneously. The need for flexibility and reliability has led to a “plug-and-play” approach in control system design, which presents challenges for traditional integrated model predictive control (MPC). Agent-based model predictive control (AMPC) has recently emerged as a distributed solution that treats controllers as agents, creating a collaborative framework among them to reach a common goal. However, this approach struggles to manage distributed conflicting objectives when agents are coupled or interdependent. To address this, we propose a novel, practical distributed control scheme called multi-objective AMPC, which adapts the alternating direction method of multipliers (ADMM) into a general control strategy that approximates global optimization while decoupling objectives. We systematically develop three formulations that maintain convergence while addressing control regularization and inequality constraints, applying them to complex vehicle control systems for the first time. The proposed method has been tested on two vehicle control scenarios with a multi-objective topology. Different formulations are compared through simulations, and the most computationally efficient one was implemented on an electric vehicle for real-world evaluations. The results demonstrate that the proposed multi-objective AMPC can converge approximately to the same global optimum as integrated MPC with greater flexibility and the potential to reduce computational costs. Jiaming Zhong, Ladan Khoshnevisan, Shucheng Huang, Mohammad Pirani, Yash Pant, Amir Khajepour |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Ellipsoid-SLAM: enhancing dynamic scene understanding through ellipsoidal object representation and trajectory tracking
Haowei Zhu, Suqin Bai, Shucheng Huang |
Vis. Comput. | 6 |
| 2025 | IOFusion: instance segmentation and optical-flow guided 3D reconstruction in dynamic scenes
Haowei Zhu, Suqin Bai, Chenggen Wang, Yunhan Sun, Shucheng Huang |
Vis. Comput. | 8 |
| 2024 | MPFC-Net: A multi-perspective feature compensation network for medical image segmentation
Xianghu Wu, Shucheng Huang, Xin Shu 0001, Chunlong Hu, Xiaojun Wu 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Learning shared features from specific and ambiguous descriptions for text-based person search
Qikai Geng, Shucheng Huang, Juanjuan Tu, Hu Lu |
Multim. Syst. | 3 |
| 2024 | Dual-stream network with cross-layer attention and similarity constraint for micro-expression recognition
Shucheng Huang |
Multim. Syst. | 2 |
| 2024 | CA-CLIP: category-aware adaptation of CLIP model for few-shot class-incremental learning
Yuqiao Xu, Shucheng Huang, Haoliang Zhou |
Multim. Syst. | 2 |
| 2024 | Class Incremental Learning for Light-Weighted NetworksabstractDespite deep neural networks (DNNs) show impressive performance across diverse tasks, they suffer from catastrophic forgetting when dealing with continuous data streams. Incremental learning aims to alleviate this phenomenon and enable DNNs to accumulate new knowledge to cope with the ever-changing world. Recently numerous advanced methods have been developed to enhance the incremental learning capabilities of neural networks. However, these methods mainly focus on the large networks, neglecting the unique needs of edged-device applications, which is surprisingly under-investigated in previous literature. In this paper, we propose two strategies for transferring knowledge from large teacher networks to light-weighted networks in class incremental learning. Specifically, in cases where the initial task contains a large number of categories, our static teacher strategy involves transferring knowledge from the teacher to the student network on the initial task to enhance the plasticity of the student network, and applying regularization constraints on the subsequent task to improve its stability. In a more challenging scenario where each task includes an equal number of categories, the dynamic teacher strategy continuously guides the student network on each task. We evaluate the proposed methods on CIFAR100, Tiny-ImageNet and ImageNet-subset datasets with different types of light-weighted networks (MobileNet, ShuffleNet). We observed that effective knowledge transfer resulting in the student network achieving performance comparable or even outperform the teacher network. Extensive and detailed experiments conducted on three datasets demonstrated the simplicity and effectiveness of our proposed method. Comprehensive analysis are also conducted including different factors and visualization. Zhe Tao, Lu Yu 0004, Hantao Yao, Shucheng Huang, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | CEPrompt: Cross-Modal Emotion-Aware Prompting for Facial Expression RecognitionabstractFacial expression recognition (FER) remains a challenging task due to the ambiguity and subtlety of expressions. To address this challenge, current FER methods predominantly prioritize visual cues while inadvertently neglecting the potential insights that can be gleaned from other modalities. Recently, vision-language pre-training (VLP) models integrated textual cues as guidance, culminating in a powerful multi-modal solution that has proven effective for a range of computer vision tasks. In this paper, we propose a Cross-Modal Emotion-Aware Prompting (CEPrompt) framework for FER based on VLP models. To make VLP models sensitive to expression-relevant visual discrepancies, we devise an Emotion Conception-guided Visual Adapter (EVA) to capture the category-specific appearance representations with emotion conception guidance. Moreover, knowledge distillation is employed to prevent the model from forgetting the pre-trained category-invariant knowledge. In addition, we design a Conception-Appearance Tuner (CAT) to facilitate the interaction of multi-modal information via cooperatively tuning between emotion conception and appearance prompts. In this way, semantic information about emotion text conception is infused directly into facial appearance images, thereby enhancing a comprehensive and precise understanding of expression-related facial details. Quantitative and qualitative experiments show that our CEPrompt outperforms state-of-the-art approaches on three real-world FER datasets. The code is available athttps://github.com/HaoliangZhou/CEPrompt. Haoliang Zhou, Shucheng Huang, Feifei Zhang 0001, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Learning Commonsense-aware Moment-Text Alignment for Fast Video Temporal GroundingabstractGrounding temporal video segments described in natural language queries effectively and efficiently is a crucial capability needed in vision-and-language fields. In this article, we deal with the fast video temporal grounding (FVTG) task, aiming at localizing the target segment with high speed and favorable accuracy. Most existing approaches adopt elaborately designed cross-modal interaction modules to improve the grounding performance, which suffer from the test-time bottleneck. Although several common space-based methods enjoy the high-speed merit during inference, they can hardly capture the comprehensive and explicit relations between visual and textual modalities. In this article, to tackle the dilemma of the speed–accuracy tradeoff, we propose a commonsense-aware cross-modal alignment network (C 2 AN) that incorporates commonsense-guided visual and text representations into a complementary common space for fast video temporal grounding. Specifically, the commonsense concepts are explored and exploited by extracting the structural semantic information from a language corpus. Then, a commonsense-aware interaction module is designed to obtain bridged visual and text features by utilizing the learned commonsense concepts. Finally, to maintain the original semantic information of textual queries, a cross-modal complementary common space is optimized to obtain matching scores for performing FVTG. Extensive results on two challenging benchmarks show that our C 2 AN method performs favorably against states of the art while running at high speed. Our code is available at https://github.com/ZiyueWu59/CCA Ziyue Wu, Junyu Gao 0002, Shucheng Huang, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | RES-CapsNet: an improved capsule network for micro-expression recognition
Xin Shu 0001, Shucheng Huang |
Multim. Syst. | 4 |
| 2023 | Shallow multi-branch attention convolutional neural network for micro-expression recognition
Shucheng Huang, Zhe Tao |
Multim. Syst. | 2 |
| 2023 | A LiDAR point cloud registration method combining linear feature extraction and TrICP algorithm
Chuanwang Wen, Shucheng Huang |
Multim. Syst. | 2 |
| 2023 | Inceptr: micro-expression recognition integrating inception-CBAM and vision transformer
Haoliang Zhou, Shucheng Huang, Yuqiao Xu |
Multim. Syst. | 2 |
| 2022 | Intentional-Deception Detection Based on Facial Muscle Movements in an Interactive Social ContextabstractMicro-expression, which is generated by facial muscle movements, could be a crucial cue for deception detection. In the existed research investigating the relationship between facial muscles and deception detection, researchers have focused almost exclusively on two muscles, i.e., zygomaticus and corrugator supercilii, based on the theoretical basis that they are highly associated with positive and negative expressions. However, the aim of this study is to demonstrate the direct relationship between facial muscle movements and deception detection. Addressing this issue, this paper proposes an experimental paradigm with high ecological validity that uses electromyography (EMG) signals to precisely examine the role of facial muscle movements in deception detection. Moreover, we propose a vector-based sequential forward selection (VSFS) algorithm to identify the muscle (or muscle combination) most closely associated with lying. Based on our proposed approach, the importance of seven selected facial muscles was explored by comparing the corresponding facial EMG (fEMG) between truth and lying conditions. First, the present study found that the zygomaticus and corrugator supercilii could play important roles in deception detection, and our findings are consistent with existed research. Second, the experiment result verified that the muscles related to deception detection were consistent with those with higher frequency occurring in micro-expression. Moreover, the present study provides a theoretical basis that intelligent micro-expressions analysis could improve the lie detection performance by focusing on the area of the forehead, eyebrows, and cheeks. Zizhao Dong, Shaoyuan Lu, Luyao Dai, Shucheng Huang, Ye Liu 0010 |
Pattern Recognit. Lett. | 5 |
| 2021 | Diving Into The Relations: Leveraging Semantic and Visual Structures For Video Moment RetrievalabstractExisting dominant approaches for video moment retrieval task are to learn semantic correlation between a given query and the video. However, these methods rarely explore the fine-grained semantic structure and comprehensive visual structure, leading to insufficient utilization of textual and visual relations. In this paper, we propose a unified framework for video moment retrieval, which considers to simultaneously encode semantic and visual structures. Specifically, a semantic role tree is built to reveal the fine-grained semantic information by generating hierarchical textual embeddings. Then the semantic structure is adopted to facilitate the visual structure learning with a contextual attention-based proposal interaction module. Finally, we adaptively aggregate and obtain the visual-semantic matching information through a multi-level fusion strategy to select the best matching moment proposal. Extensive experiments on two popular benchmarks (Charades-STA and ActivityNet Captions) show that our proposed method achieves state-of-the-art performance. Codes are available in the Supplementary Material. Ziyue Wu, Junyu Gao 0002, Shucheng Huang, Changsheng Xu |
ICME | 3 |
| 2021 | Face spoofing detection based on chromatic ED-LBP texture feature
Xin Shu 0001, Shucheng Huang |
Multim. Syst. | 3 |
| 2021 | Multiple channels local binary pattern for color texture representation and classification
Xin Shu 0001, Zhigang Song, Shucheng Huang, Xiaojun Wu 0001 |
Signal Process. Image Commun. | 4 |
| 2020 | Deep Sentiment Classification and Topic Discovery on Novel Coronavirus or COVID-19 Online Discussions: NLP Using LSTM Recurrent Neural Network ApproachabstractInternet forums and public social media, such as online healthcare forums, provide a convenient channel for users (people/patients) concerned about health issues to discuss and share information with each other. In late December 2019, an outbreak of a novel coronavirus (infection from which results in the disease named COVID-19) was reported, and, due to the rapid spread of the virus in other parts of the world, the World Health Organization declared a state of emergency. In this paper, we used automated extraction of COVID-19-related discussions from social media and a natural language process (NLP) method based on topic modeling to uncover various issues related to COVID-19 from public opinions. Moreover, we also investigate how to use LSTM recurrent neural network for sentiment classification of COVID-19 comments. Our findings shed light on the importance of using public opinions and suitable computational techniques to understand issues surrounding COVID-19 and to guide related decision-making. In addition, experiments demonstrated that the research model achieved an accuracy of 81.15% - a higher accuracy than that of several other well-known machine-learning algorithms for COVID-19-Sentiment Classification. Hamed Jelodar, Yongli Wang 0002, Rita Orji, Shucheng Huang |
IEEE J. Biomed. Health Informatics | 4 |
| 2019 | Video Highlight Detection via Region-Based Deep Ranking ModelabstractThe video highlight detection task is to localize key elements (moments of user’s major or special interest) in a video. Most of the existing highlight detection approaches extract features from the video segment as a whole without considering the difference of local features spatially. In spatial extent, not all regions are worth watching because some of them only contain the background of the environment without human or other moving objects, especially when there is lots of clutter in the background. To deal with this issue, we propose a novel region-based model which can automatically localize the key elements in a video without any extra supervised annotations. Specifically, the proposed model produces position-sensitive score maps for local regions in the spatial dimension of the video segment, and then aggregates all position-wise scores with position-pooling operation. The regions with higher response values will be extracted as key elements. Thus more effective features of the video segment are obtained to predict the highlight score. The proposed position-sensitive scheme can be easily integrated into an end-to-end fully convolutional network which aims to update parameters via stochastic gradient descent method in the backward propagation to improve the robustness of the model. Extensive experimental results on the YouTube and SumMe datasets demonstrate that the proposed approach achieves significant improvement over state-of-the-art methods. Yifan Jiao, Tianzhu Zhang 0001, Shucheng Huang, Bin Liu 0014, Changsheng Xu |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2018 | Three-Dimensional Attention-Based Deep Ranking Model for Video Highlight DetectionabstractThe video highlight detection task is to localize key elements (moments of user's major or special interest) in a video. Most of existing highlight detection approaches extract features from the video segment as a whole without considering the difference of local features both temporally and spatially. Due to the complexity of video content, this kind of mixed features will impact the final highlight prediction. In temporal extent, not all frames are worth watching because some of them only contain the background of the environment without human or other moving objects. In spatial extent, it is similar that not all regions in each frame are highlights especially when there are lots of clutters in the background. To solve the above problem, we propose a novel three-dimensional (3-D) (spatial+temporal) attention model that can automatically localize the key elements in a video without any extra supervised annotations. Specifically, the proposed attention model produces attention weights of local regions along both the spatial and temporal dimensions of the video segment. The regions of key elements in the video will be strengthened with large weights. Thus, the more effective feature of the video segment is obtained to predict the highlight score. The proposed 3-D attention scheme can be easily integrated into a conventional end-to-end deep ranking model that aims to learn a deep neural network to compute the highlight score of each video segment. Extensive experimental results on the YouTube and SumMe datasets demonstrate that the proposed approach achieves significant improvement over state-of-the-art methods. With the proposed 3-D attention model, video highlights can be accurately retrieved in spatial and temporal dimensions without human supervision in several domains, such as gymnastics, parkour, skating, skiing, surfing, and dog activities, on the public datasets. Yifan Jiao, Zhetao Li, Shucheng Huang, Xiaoshan Yang, Bin Liu 0014, Tianzhu Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2017 | Video Highlight Detection via Deep Ranking Modeling
Yifan Jiao, Xiaoshan Yang, Tianzhu Zhang 0001, Shucheng Huang, Changsheng Xu |
PSIVT | 4 |
| 2016 | Multi-object tracking via discriminative appearance modeling
Shucheng Huang |
Comput. Vis. Image Underst. | 1 |
| 2016 | Exponential Discriminant Locality Preserving Projection for face recognition
Shucheng Huang, Zhuang Lu |
Neurocomputing | 1 |
| 2007 | An active learning system for mining time-changing data streams
Shucheng Huang, Yisheng Dong |
Intell. Data Anal. | 1 |