Yaocong Hu

dblp:181/3373 · DBLP profile ↗
← Back
19ranked-venue papers
13as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 10 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Unifying Modality and Scale: Visual Mamba for Feature Fusion in RGB-X Crowd Counting
abstract
Crowd counting has long been a crucial topic in the domains of computer vision and video surveillance. In particular, with the widespread use of thermal or depth cameras, RGB-X crowd counting has emerged as a prominent research focus. Although depth or infrared images provide complementary information, the core challenge remains in effectively unifying heterogeneous cross-modality and cross-scale information to form a comprehensive representation of crowd distributions in complex scenes. To address this problem, we propose a novel Mamba-based framework, termed UMS-VMamba-CC for multi-modal (RGB-X) crowd counting. Specifically, we design the cross-modality disentanglement fusion visual Mamba (CMDF-VMamba) that uses self-supervised learning to decompose modality-invariant and modality-specific features in spatial-frequency domains, followed by multi-modal feature aggregation via the gating mechanism. For cross-scale fusion, we design the cross-scale pyramid fusion visual Mamba (CSPF-VMamba), which adopts the bi-directional pyramid structure to incorporate low-level features into high-level representations during downsampling, while upsampling high-level features and aggregating them with low-level features through state space contextual modeling. Comprehensive experiments on multiple mainstream datasets demonstrate that the UMS-VMamba-CC framework achieves competitive performance for RGB-X crowd counting.
Yaocong Hu, Mengbo Jia, Pindeng Wang, Wenbo Zhu 0002, Huanjie Tao, Tianming Ni, Teng Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 EMFFTrans: Efficient Multi-Scale Feature Fusion Transformer for Road Scene Semantic Segmentation
Yaocong Hu, Pindeng Wang, Mengbo Jia, Jinwen Hong, Guoyang Wan, Huicheng Yang, Bingyou Liu, Tianming Ni, Teng Li 0001
IEEE Trans. Intell. Transp. Syst.1
2025 Residual Quotient Learning for Zero-Reference Low-Light Image Enhancement
abstract
Recently, neural networks have become the dominant approach to low-light image enhancement (LLIE), with at least one-third of them adopting a Retinex-related architecture. However, through in-depth analysis, we contend that this most widely accepted LLIE structure is suboptimal, particularly when addressing the non-uniform illumination commonly observed in natural images. In this paper, we present a novel variant learning framework, termed residual quotient learning, to substantially alleviate this issue. Instead of following the existing Retinex-related decomposition-enhancement-reconstruction process, our basic idea is to explicitly reformulate the light enhancement task as adaptively predicting the latent quotient with reference to the original low-light input using a residual learning fashion. By leveraging the proposed residual quotient learning, we develop a lightweight yet effective network called ResQ-Net. This network features enhanced non-uniform illumination modeling capabilities, making it more suitable for real-world LLIE tasks. Moreover, due to its well-designed structure and reference-free loss function, ResQ-Net is flexible in training as it allows for zero-reference optimization, which further enhances the generalization and adaptability of our entire framework. Extensive experiments on various benchmark datasets demonstrate the merits and effectiveness of the proposed residual quotient learning, and our trained ResQ-Net outperforms state-of-the-art methods both qualitatively and quantitatively. Furthermore, a practical application in dark face detection is explored, and the preliminary results confirm the potential and feasibility of our method in real-world scenarios.
Linfeng Fei, Huanjie Tao, Yaocong Hu, Wei Zhou 0042, Jiun Tian Hoe, Weipeng Hu, Yap-Peng Tan
IEEE Trans. Image Process.4
2024 CLDE-Net: crowd localization and density estimation based on CNN and transformer network
Yaocong Hu, Huicheng Yang, Bingyou Liu, Guoyang Wan, Jinwen Hong, Xiaobo Lu
Multim. Syst.1
2024 ILSR-Diff: joint face illumination normalization and super-resolution via diffusion models
Minghao Mu, Yaocong Hu, Xiaobo Lu
Multim. Syst.4
2024 ESDAR-net: towards high-accuracy and real-time driver action recognition for embedded systems
Yaocong Hu, Zhen Shuai, Huicheng Yang, Guoyang Wan, MingQi Lu, Xiaobo Lu
Multim. Tools Appl.1
2022 A pose-aware dynamic weighting model using feature integration for driver action recognition
MingQi Lu, Yaocong Hu, Xiaobo Lu
Eng. Appl. Artif. Intell.2
2022 Pose-guided model for driving behavior recognition using keypoint action learning
MingQi Lu, Yaocong Hu, Xiaobo Lu
Signal Process. Image Commun.2
2021 FIN-GAN: Face illumination normalization via retinex-based self-supervised learning and conditional generative adversarial network
Yaocong Hu, MingQi Lu, Xiaobo Lu
Neurocomputing1
2021 Video-based driver action recognition via hybrid spatial-temporal deep learning framework
Yaocong Hu, MingQi Lu, Xiaobo Lu
Multim. Syst.1
2020 Driver action recognition using deformable and dilated faster R-CNN with optimized region proposals
MingQi Lu, Yaocong Hu, Xiaobo Lu
Appl. Intell.2
2020 Feature refinement for image-based driver action recognition via multi-scale attention convolutional neural network
Yaocong Hu, MingQi Lu, Xiaobo Lu
Signal Process. Image Commun.1
2020 Driver Drowsiness Recognition via 3D Conditional GAN and Two-Level Attention Bi-LSTM
abstract
Driver drowsiness has currently been a severe issue threatening road safety, hence it is vital to develop an effective drowsiness recognition algorithm to avoid traffic accidents. However, recognizing drowsiness is still very challenging, due to the large intra-class variations in facial expression, head pose and illumination condition. In this paper, a new deep learning framework based on the hybrid of 3D conditional generative adversarial network and two-level attention bidirectional long short-term memory network (3DcGAN-TLABiLSTM) has been proposed for robust driver drowsiness recognition. Aiming at extracting short-term spatial-temporal features with abundant drowsiness-related information, we design a 3D encoder-decoder generator with the condition of auxiliary information to generate high-quality fake image sequences and devise a 3D discriminator to learn drowsiness-related representation from spatial-temporal domain. In addition, for long-term spatial-temporal fusion, we investigate the use of two-level attention mechanism to guide the bidirectional long short-term memory learn the saliency of short-term memory information and long-term temporal information. For experiment, we evaluate our 3DcGAN-TLABiLSTM framework on a public NTHU-DDD dataset. Experimental results show that the proposed approach achieves higher precision of drowsiness recognition compared to the state-of-the-art.
Yaocong Hu, MingQi Lu, Xiaobo Lu
IEEE Trans. Circuits Syst. Video Technol.1
2019 Dilated Light-Head R-CNN using tri-center loss for driving behavior recognition
MingQi Lu, Yaocong Hu, Xiaobo Lu
Image Vis. Comput.2
2019 Driving behaviour recognition from still images by using multi-stream fusion CNN
Yaocong Hu, MingQi Lu, Xiaobo Lu
Mach. Vis. Appl.1
2018 Spatial-Temporal Fusion Convolutional Neural Network for Simulated Driving Behavior Recognition
abstract
Abnormal driving behaviour is one of the leading cause of terrible traffic accidents endangering human life. Therefore, study on driving behaviour surveillance has become essential to traffic security and public management. In this paper, we conduct this promising research and employ a two stream CNN framework for video-based driving behaviour recognition, in which spatial stream CNN captures appearance information from still frames, whilst temporal stream CNN captures motion information with pre-computed optical flow displacement between a few adjacent video frames. We investigate different spatial-temporal fusion strategies to combine the intra frame static clues and inter frame dynamic clues for final behaviour recognition. So as to validate the effectiveness of the designed spatial-temporal deep learning based model, we create a simulated driving behaviour dataset, containing 1237 videos with 6 different driving behavior for recognition. Experiment result shows that our proposed method obtains noticeable performance improvements compared to the existing methods.
Yaocong Hu, MingQi Lu, Xiaobo Lu
ICARCV1
2018 Learning spatial-temporal features for video copy detection by the combination of CNN and RNN
Yaocong Hu, Xiaobo Lu
J. Vis. Commun. Image Represent.1
2018 Real-time video fire smoke detection by utilizing spatial-temporal ConvNet features
Yaocong Hu, Xiaobo Lu
Multim. Tools Appl.1
2016 Dense crowd counting from still images with convolutional neural networks
Yaocong Hu, Fudong Nian, Yan Wang 0059, Teng Li 0001
J. Vis. Commun. Image Represent.1