Cheng Xu 0005

dblp:74/5952-5 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0003-4913-5371ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 12 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Computer networks · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 CBHA-DETR: multi-kernel attention and deformable fusion network for behavior recognition in classroom monitoring
Cheng Xu 0005, Bingxin Xu
Multim. Syst.3
2026 Sgd-det:structure-guided oracle character detection in degraded rubbing images
Xuxing Qi, Cheng Xu 0005, Jiazheng Yuan, Bofeng Mo, Hongzhe Liu 0001
Multim. Syst.3
2026 TSPPO: transformer-based sequential proximal policy optimization for multi-agent systems
Yuxiao Gao, Cheng Xu 0005, Hongzhe Liu 0001
Multim. Syst.3
2025 YOLODF: YOLO-Based Spatial-Frequency Interaction Mining for General Deepfake Detection
Xin Li 0184, Bingxin Xu, Hongzhe Liu 0001, Weiguo Pan, Cheng Xu 0005
PRCV (7)5
2025 Self-attention enhanced dynamic semantic multi-scale graph convolutional network for skeleton-based action recognition
Cheng Xu 0005, Songyin Dai, Nuoya Li, Weiguo Pan, Bingxin Xu, Hongzhe Liu 0001
Image Vis. Comput.2
2025 DAN: Distortion-aware Network for fisheye image rectification using graph reasoning
Yongjia Yan, Hongzhe Liu 0001, Cheng Xu 0005, Bingxin Xu, Weiguo Pan, Songyin Dai, Yiqing Song
Image Vis. Comput.4
2025 Optimization methods in fully cooperative scenarios: a review of multiagent reinforcement learning
abstract
Multiagent reinforcement learning (MARL) has become a dazzling new star in the field of reinforcement learning in recent years, demonstrating its immense potential across many application scenarios. The reward function directs agents to explore their environments and make optimal decisions within them by establishing evaluation criteria and feedback mechanisms. Concurrently, cooperative objectives at the macro level provide a trajectory for agents’ learning, ensuring alignment between individual behavioral strategies and the overarching system goals. The interplay between reward structures and cooperative objectives not only bolsters the effectiveness of individual agents but also fosters interagent collaboration, offering both momentum and direction for the development of swarm intelligence and the harmonious operation of multiagent systems. This review delves deeply into the methods for designing reward structures and optimizing cooperative objectives in MARL, along with the most recent scientific advancements in this field. The article meticulously reviews the application of simulation environments in cooperative scenarios and discusses future trends and potential research directions in the field, providing a forward-looking perspective and inspiration for subsequent research efforts.
Xinhao Shi, Qinghan Zeng, Cheng Xu 0005, Hongzhe Liu 0001
Frontiers Inf. Technol. Electron. Eng.5
2025 Ihenet: an illumination invariant hierarchical feature enhancement network for low-light object detection
Nuoya Li, Weiguo Pan, Bingxin Xu, Hongzhe Liu 0001, Songyin Dai, Cheng Xu 0005
Multim. Syst.6
2025 Generalization-oriented face forgery detection via discriminative feature analysis and normalization
Xin Li 0184, Bingxin Xu, Hongzhe Liu 0001, Weiguo Pan, Cheng Xu 0005
Multim. Syst.5
2025 Flood scenarios vehicle detection algorithm based on improved YOLOv9
Jiwu Sun, Cheng Xu 0005, Hongzhe Liu 0001
Multim. Syst.2
2024 ISE-UFDS: A Dataset for Detecting the Degree of Danger to Vehicles in Urban Flooding and Performance Assessment
Jiwu Sun, Cheng Xu 0005, Hongzhe Liu 0001
ICIC (7)3
2024 TSD-YOLO: Small traffic sign detection based on improved YOLO v8
abstract
Abstract Traffic sign detection is critical for autonomous driving technology. However, accurately detecting traffic signs in complex traffic environments remains challenge despite the widespread use of one‐stage detection algorithms known for their real‐time processing capabilities. In this paper, the authors propose a traffic sign detection method based on YOLO v8. Specifically, this study introduces the Space‐to‐Depth (SPD) module to address missed detections caused by multi‐scale variations of traffic signs in traffic scenes. The SPD module compresses spatial information into depth channels, expanding the receptive field and enhancing the detection capabilities for objects of varying sizes. Furthermore, to address missed detections caused by complex backgrounds such as trees, this paper employs the Select Kernel attention mechanism. This mechanism enables the model to dynamically adjust its focus and more effectively concentrate on key features. Additionally, considering the uneven distribution of training data, the authors adopted the WIoUv3 loss function, which optimizes loss calculation through a weighted approach, thereby improving the model's detection performance across various sizes and frequencies of instances. The proposed methods were validated on the CCTSDB and TT100K datasets. Experimental results demonstrate that the authors’ method achieves substantial improvements of 3.2% and 5.1% on the mAP50 metric compared to YOLOv8s, while maintaining high detection speed, significantly enhancing the overall performance of the detection system. The code for this paper is located at https://github.com/dusongjie/TSD‐YOLO‐Small‐Traffic‐Sign‐Detection‐Based‐on‐Improved‐YOLO‐v8
Songjie Du, Weiguo Pan, Nuoya Li, Songyin Dai, Bingxin Xu, Hongzhe Liu 0001, Cheng Xu 0005, Xuewei Li 0006
IET Image Process.7
2024 FSKT-GE: Feature maps similarity knowledge transfer for low-resolution gaze estimation
abstract
Abstract The limited of texture details information in low‐resolution facial or eye images presents a challenge for gaze estimation. To address this, FSKT‐GE (feature maps similarity knowledge transfer for low‐resolution gaze estimation) is proposed, a gaze estimation framework consisting of both a high resolution (HR) network and low resolution (LR) network with the identical structure. Rather than mere feature imitation, this issue is addressed by assessing the cosine similarity of feature layers, emphasizing the distribution similarity between the HR and LR networks. This enables the LR network to acquire richer knowledge. This framework utilizes a combination loss function, incorporating cosine similarity measurement, soft loss based on probability distribution difference and gaze direction output, along with a hard loss from the LR network output layer. This approach on low‐resolution datasets derived from Gaze360 and RT‐Gene datasets is validated, demonstrating excellent performance in low‐resolution gaze estimation. Evaluations on low‐resolution images obtained through 2×, 4×, and 8× down‐sampling are conducted on two datasets. On the Gaze360 dataset, the lowest mean angular errors of 10.97°, 11.22°, and 13.61° were achieved, while on the RT‐Gene dataset, the lowest mean angular errors of 6.73°, 6.83°, and 7.75° were obtained.
Weiguo Pan, Songyin Dai, Bingxin Xu, Cheng Xu 0005, Hongzhe Liu 0001, Xuewei Li 0006
IET Image Process.5
2024 Spatial-temporal hypergraph based on dual-stage attention network for multi-view data lightweight action recognition
Zhixuan Wu, Nan Ma 0008, Cheng Xu 0005, Genbao Xu
Pattern Recognit.4
2022 W-ART: Action Relation Transformer for Weakly-Supervised Temporal Action Localization
abstract
Weakly-supervised temporal action localization (WTAL) is a long-standing and challenging research problem in video signal analysis. It is to localize the action segments in the video given only video-level labels. The key to this task is understanding how the diverse actions interact. In this paper, we propose W-ART, a relation Transformer to explicitly capture the relationships between action segments. We devise a new effective Transformer architecture and construct new training loss functions for WTAL. Further, we propose a dedicated query mechanism to satisfy the different feature preferences between classification and localization. Thanks to these designs, our W-ART can accurately localize the diverse actions even in weakly-supervised setting. Extensive evaluation and empirical analysis show that our method outperforms the state of the arts on two challenging benchmarks, Charades and THUMOS14.
Mengzhu Li, Hongjun Wu 0002, Yongcheng Liu, Hongzhe Liu 0001, Cheng Xu 0005, Xuewei Li 0006
ICASSP5
2022 Transtl: Spatial-Temporal Localization Transformer for Multi-Label Video Classification
abstract
Multi-label video classification (MLVC) is a long-standing and challenging research problem in video signal analysis. Generally, there exist many complex action labels in real-world videos and these actions are with inherent dependencies at both spatial and temporal domains. Motivated by this observation, we propose TranSTL, a spatial-temporal localization Transformer framework for MLVC task. In addition to leverage global action label co-occurrence, we also propose a novel plug-and-play Spatial Temporal Label Dependency (STLD) layer in TranSTL. STLD not only dynamically models the label co-occurrence in a video by self-attention mechanism, but also fully captures spatial-temporal label dependencies using cross-attention strategy. As a result, our TranSTL is able to explicitly and accurately grasp the diverse action labels at both spatial and temporal domains. Extensive evaluation and empirical analysis show that TranSTL achieves superior performance over the state of the arts on two challenging benchmarks, Charades and Multi-Thumos.
Hongjun Wu 0002, Mengzhu Li, Yongcheng Liu, Hongzhe Liu 0001, Cheng Xu 0005, Xuewei Li 0006
ICASSP5
2022 SMART: Semantic-Aware Masked Attention Relational Transformer for Multi-label Image Recognition
abstract
As objects usually co-exist in an image, learning the label co-occurrence is a compelling approach to improving the performance of multi-label image recognition. However, the dependencies among the non-exist categories in an image cannot be effectively evaluated. These redundant label dependencies may bring noise and further decrease the performance of classification. Therefore, we proposed SMART, a Semantic-aware Masked Attention Relational Transformer for Multi-label Image Recognition for multi-label image recognition tasks. In addition to leveraging Transformer to model the inter-class dependencies, the proposed masked attention filters out the redundant dependencies among the non-exist categories. SMART is able to explicitly and accurately capture the label dependencies without extra word embeddings. Moreover, our method achieves new state-of-the-art results on two benchmarks for multi-label image recognition: MS-COCO 2014 and NUS-WIDE. In addition, extensive ablation studies and empirical analysis are provided to demonstrate the effectiveness of the essential components of our method under different factors.
Hongjun Wu 0002, Cheng Xu 0005, Hongzhe Liu 0001
IEEE Signal Process. Lett.2
2022 A Real-Time Complex Road AI Perception Based on 5G-V2X for Smart City Security
abstract
The Internet of Vehicles and information security are key components of a smart city. Real‐time road perception is one of the most difficult tasks. Traditional detection methods require manual adjustment of parameters, which is difficult, and is susceptible to interference from object occlusion, light changes, and road wear. Designing a robust road perception algorithm is still challenging. On this basis, we combine artificial intelligence algorithms and the 5G‐V2X framework to propose a real‐time road perception method. First, an improved model based on Mask R‐CNN is implemented to improve the accuracy of detecting lane line features. Then, the linear and polynomial fitting methods of feature points in different fields of view are combined. Finally, the optimal parameter equation of the lane line can be obtained. We tested our method in complex road scenes. Experimental results show that, combined with 5G‐V2X, this method ultimately has a faster processing speed and can sense road conditions robustly under various complex actual conditions.
Cheng Xu 0005, Hongjun Wu 0002, Yinong Zhang, Songyin Dai, Hongzhe Liu 0001
Wirel. Commun. Mob. Comput.1
2020 Mutual authentication for vehicular network in complex and uncertain driving
Cheng Xu 0005, Hongzhe Liu 0001
Neural Comput. Appl.1
2018 An Anonymous Handover Authentication Scheme Based on LTE-A for Vehicular Networks
abstract
Vehicular networks play an important role in the intelligent transportation systems which have gained technical supports from car industry. Due to the mobility and the broadcast nature of wireless communication, security of the vehicular networks is a critical issue for the academia and industry. Many solutions have been proposed to target the security provisioning. However, most of them have various shortcomings. Based on the elliptic curve public key cryptography algorithm, in this paper, we propose a new anonymous roaming authentication protocol for the Long Term Evolution‐Advanced (LTE‐A) supported vehicular networks. For a vehicular LTE‐A network, an authentication protocol should be able to fulfill a variety of security requirements, which can be met by our proposal and proved by using Burrows–Abadi–Needham (BAN) logic. Compared with some existing solutions, our scheme has lower communication costs with stronger security functionality. The analyses on the security functions and the performance of the proposed solution show that our scheme is secure and efficient with ability against various types of malicious attacks.
Cheng Xu 0005, Xiaohong Huang 0003, Maode Ma, Hong Bao
Wirel. Commun. Mob. Comput.1