Huicheng Lai

dblp:224/2516 · DBLP profile ↗
← Back
16ranked-venue papers
0as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MTC-VAD: cross-modal temporal coherence modeling for video anomaly detection
Huicheng Lai, Guxue Gao
Multim. Syst.2
2024 Efficient Guided Query Network for Human-Object Interaction Detection
abstract
Recently, Transformer-based one-stage methods have demonstrated excellent efficiency in Human-Object Interaction (HOI) tasks. However, these methods often utilize semantically ambiguous initial queries, thus constraining the model’s ability for set prediction. In addition, currently widely used HOI datasets suffer from long-tail distribution issues, so accurately identifying rare interaction categories remains challenging. To address these challenges, we propose an Efficient Guided Query Network (EGQ-Net). The network introduces a forward-guided relational queries approach, which accurately captures the triplets of interaction relationships by effectively integrating the initial queries predicted by the encoder and the output features of each decoder layer. Furthermore, we used the visual language pre-training models CLIP and BLIP2 to design interaction position query guidance and interaction content query guidance to achieve accurate recognition and localization of interactive areas by queries. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on widely used HOI benchmarks (V-COCO and HICO-DET).
Junkai Li, Huicheng Lai, Tongguan Wang, Hutuo Quan, Dongji Chen
ICME2
2024 UFSRNet: U-shaped face super-resolution reconstruction network based on wavelet transform
Tongguan Wang, Yang Xiao 0018, Yuxi Cai, Guxue Gao, Xiaocong Jin, Huicheng Lai
Multim. Tools Appl.7
2023 Scoreformer: Score Fusion-Based Transformers for Weakly-Supervised Violence Detection
abstract
Violence detection is an application of anomaly detection, which is used to detect violence content in video clips. Using multimodal as input can improve the performance of violence detection. However, the existing MML Transformers-based fusion methods do not take into account the differences between non-homologous modals. The fusion of non-homologous modals makes features become noise between each other. This paper proposes a score fusion-based transformer framework, named Scoreformer. First of all, the optical flow, RGB and audio features pass through the independent self-attention transformer blocks. Second, the optical flow and RGB features pass through the cross-modal transformer blocks, after that they are fused with the audio features through the score fusion block. This method avoids the noise interference caused by the direct fusion of audio features and visual features. Experiments on the XD-Violence dataset show that the proposed method achieves 84.54% of the AP value, which exceeds at least 2.85% compared with the most advanced method (e. g. MSL, CRFD).
Yang Xiao 0018, Tongguan Wang, Huicheng Lai
ICASSP4
2023 HATFNet: Hierarchical adaptive trident fusion network for RGBT tracking
Huicheng Lai, Guxue Gao
Appl. Intell.2
2023 Video anomaly detection based on cross-frame prediction mechanism and spatio-temporal memory-enhanced pseudo-3D encoder
Xiaopeng Wen, Huicheng Lai, Guxue Gao, Yang Xiao 0018, Tongguan Wang, Zhenhong Jia
Eng. Appl. Artif. Intell.2
2023 Video abnormal behaviour detection based on pseudo-3D encoder and multi-cascade memory mechanism
abstract
Abstract Frame prediction methods based on Auto‐Encoder (AE) composed of convolutional neural networks (CNN) are very popular in detecting abnormal behaviour. The methods predict normal behaviour accurately and abnormal behaviour incorrectly, which is considered a criterion for abnormality discrimination. However, the emergence of problems such as too strong AE representation leading to detection failure, the insufficient ability of the network to extract spatio‐temporal information, a large number of model parameters and slow running speed leads to the need for the method to be further improved. In this work, the authors propose a network framework for abnormal behaviour detection in video based on a pseudo‐3D encoder and a multi‐cascade memory mechanism (MMP3D). First of all, the encoder consisting of pseudo‐3D convolution is used to extract spatio‐temporal information from the video. Then, the multi‐cascade memory mechanism (MM) and the multi‐headed prototype attention mechanism are used to store and aggregate features of normal behaviour, which solves to some extent the problem of detection failure caused by strong AE representation power. Finally, the decoder designed by the 2D deconvolution layers is used to recover the prediction information. The efficiency and superiority of our method is validated on the Ped2 dataset, Avenue dataset, and ShanghaiTech dataset.
Xiaopeng Wen, Huicheng Lai, Guxue Gao
IET Image Process.2
2023 Two-Step Unsupervised Approach for Sand-Dust Image Enhancement
abstract
In sand‐dust environments, light is scattered and absorbed, and sand‐dust images thus suffer from severe image degradation problems, such as color shifts, low contrast, and blurred details. To address these problems, we propose a two‐step unsupervised sand‐dust image enhancement algorithm. In the first step, a convenient and competent color correction method is put forward to solve the color shift problem. Considering the wavelength attenuation features of sand‐dust images, a linear stretching and blue channel compensation method is designed, and an adaptive color shift correction factor is developed to remove the color shift. In the second step, to enhance the clarity and details of the images, an unsupervised generative adversarial network is proposed, which does not require pairs of data for training. To reduce detail loss, the detail enhancement branch is designed, and the generator considers to more details through the constructed coarse‐grained and fine‐grained discriminators. The introduced multiscale perceptual loss promotes the image fidelity well. Experiments show that the proposed method achieves better color correction, enhances image details and clarity, has a better subjective effect, and outperforms existing sand‐dust image enhancement methods both quantitatively and qualitatively. Similarly, our method promotes the application capability of the target detection algorithm and also has a good enhancement effect on underwater images and haze images.
Guxue Gao, Huicheng Lai, Zhenhong Jia
Int. J. Intell. Syst.2
2023 Sand-dust image enhancement based on light attenuation and transmission compensation
Zhenhong Jia, Huicheng Lai, Nikola K. Kasabov, Sensen Song
Multim. Tools Appl.3
2023 A real time target face tracking algorithm based on saliency detection and Camshift
Zhenhong Jia, Huicheng Lai
Multim. Tools Appl.3
2023 Echo-Enhanced Embodied Visual Navigation
abstract
Visual navigation involves a movable robotic agent striving to reach a point goal (target location) using vision sensory input. While navigation with ideal visibility has seen plenty of success, it becomes challenging in suboptimal visual conditions like poor illumination, where traditional approaches suffer from severe performance degradation. We propose E3VN (echo-enhanced embodied visual navigation) to effectively perceive the surroundings even under poor visibility to mitigate this problem. This is made possible by adopting an echoer that actively perceives the environment via auditory signals. E3VN models the robot agent as playing a cooperative Markov game with that echoer. The action policies of robot and echoer are jointly optimized to maximize the reward in a two-stream actor-critic architecture. During optimization, the reward is also adaptively decomposed into the robot and echoer parts. Our experiments and ablation studies show that E3VN is consistently effective and robust in point goal navigation tasks, especially under nonideal visibility.
Yinfeng Yu, Le-le Cao, Fuchun Sun 0001, Chao Yang 0026, Huicheng Lai, Wenbing Huang 0001
Neural Comput.5
2022 Lightweight spatial-channel adaptive coordination of multilevel refinement enhancement network for image reconstruction
Yuxi Cai, Huicheng Lai, Zhenhong Jia
Knowl. Based Syst.2
2022 Color balance and sand-dust image enhancement in lab space
Guxue Gao, Huicheng Lai, Zhenhong Jia
Multim. Tools Appl.2
2022 Target re-location kernel correlation filtered visual tracking with fused deep feature
Qingzhong Shu, Huicheng Lai, Zhenhong Jia
Multim. Tools Appl.2
2019 Cells Planning of VLC Networks using Non-Circular Symmetric Optical Beam
abstract
In the typical indoor office environments, the existing site locations resource on ceiling for visible light communications (VLC) access points (AP) is relatively limited due to the originally illumination & fire-fighting oriented ceiling planning. For reducing the capital expenditure of VLC networking in most situations, the straight forward solution is making full use of the limited linear location of luminaries. In this work, for the first time, the non-circular symmetric optical beam is introduced to the cells planning of VLC networks. The coverage performance is compared between the typical noncircular symmetric optical beam pattern i.e. BBE LED pattern and the well investigated circular symmetric optical beam pattern i.e. Lambertian pattern. It is sufficiently demonstrated that the satisfying matching capability of this non-circular symmetric optical beam to indoor scenario with just 3 linear site locations. By appropriately placing the BBE LED pattern, more than 12 dB the signal-to-interference-plus-noise ratio (SINR) enhancement is provided for more than 70% receiver positions while the total optical transmitted power of all VLC AP remains unchanged. Moreover, the superior performance uniformity of this optical beam pattern is separately identified in cellular cells configuration case and in electrical frequency reuse case.
Jupeng Ding, Chih-Lin I, Xifeng Chen, Baoshan Yu, Huicheng Lai
ICC6
2019 A novel deep hashing method for fast image retrieval
Shuli Cheng, Huicheng Lai, Jiwei Qin
Vis. Comput.2