Chengtao Cai

dblp:39/7697 · DBLP profile ↗
← Back
48ranked-venue papers
7as first author
35since 2021 · last 2026
0000-0002-3475-6098ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 16 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 9 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Point-Voxel Transformer for point cloud object detection with spatial and channel attention
Guangyu Ji, Chengtao Cai, Kaibin Qin
Eng. Appl. Artif. Intell.3
2026 Advanced deep features fusion network for partial overlapping registration
Zhuoran Tian, Chengtao Cai
Neurocomputing3
2026 VLMAR: Maritime scene anomaly detection via retrieval-augmented vision-language models
Chunsheng Yang, Chengtao Cai
J. Vis. Commun. Image Represent.3
2026 A Proposal-Based Transformer With Point-Voxel Cross Attention for Accurate 3D Object Detection
abstract
Existing point-voxel hybrid two-stage 3D detectors mainly rely on point-based feature extraction in the second stage, while insufficiently exploiting the voxel features generated in the first stage. To address this limitation, we propose a novel two-stage 3D object detection framework that combines a voxel-based Region Proposal Network with a Proposal-Based Transformer using Point-Voxel Cross Attention (PVCAT). PVCAT integrates point and voxel features within the detection head, enhancing the detection performance of small objects while introducing only marginal computational overhead. The module consists of a Proposal Embedding Module, three Spatial-Aware Encoder layers, and three Channel-Aware Decoder layers in a three-layer encoder-decoder architecture. For each proposal, point-wise embeddings are first generated and then enhanced with voxel features from the RPN stage. In the encoder, Voxel features are leveraged as positional cues to model the local spatial context among points. In the decoder, a weighted fusion strategy strengthens query-key interactions by integrating global and local channel information from both point and voxel features. Experiments indicate that PVCAT is compatible with different voxel-based backbones and remains effective in improving detection performance, demonstrating its potential to enhance the accuracy of 3D object detectors.
Guangyu Ji, Chengtao Cai, Kaibin Qin
IEEE Signal Process. Lett.3
2026 Wavelet-Guided Token-Focus Transformer for Fine-Grained Ship Recognition
abstract
Fine-grained ship recognition in complex backgrounds is challenged by significant texture variations and the frequent occlusion of the ship's core structure by noise. Existing studies have not effectively addressed complex background interference and lack the ability to dynamically focus on the most discriminative features, particularly in small target recognition tasks. To overcome these issues, this paper proposes a novel end-to-end network architecture that integrates Multi-Scale Wavelet Block (MSWB) and Token Focus Attention Block (TFAB), which effectively focuses on key ship features while mitigating the impact of background noise. The proposed method is evaluated on the constructed Complex Background Ships (GCS) dataset, as well as the publicly available MAR-ships and Game-of-Ships datasets. Experimental results demonstrate that our method achieves leading accuracy of 85.7%, 95.7%, and 98.1% on these three datasets, respectively. The method significantly enhances the robustness and precision of ship recognition in real-world scenarios and provides valuable insights for dynamic monitoring applications.
Runtian Wang, Renjie Qiao, Chengtao Cai
IEEE Signal Process. Lett.3
2026 HALT: Hierarchical Attention Learning for Visual Tracking
abstract
Siamese tracking algorithms have gained widespread recognition due to their exceptional efficiency and scalability. However, they exhibit suboptimal performance when localizing arbitrary targets under various disturbances, particularly in complex environments involving challenges such as illumination variation, deformation, and background clutter. Therefore, this paper proposes a Hierarchical Attention Learning network (HAL) to enhance tracking performance. Drawing inspiration from the hybrid attention mechanism, HAL designs a Feature Mining Attention module (FMA), a Global Feature Attention module (GFA), and a hierarchical network structure. Concretely, FMA employs parallel branches to fully extract channel and spatial features, enabling preliminary enhancement of target features and establishing global associations. Since lower-level and higher-level features emphasize positional and semantic information, respectively, GFA comprehensively integrates multi-level features to obtain more accurate tracking predictions. In particular, to improve the model's representation capacity, a hierarchical network structure is developed to deepen the HAL network and strengthen the feature dependencies between the template and the search region. Finally, based on HAL, we propose a Hierarchical Attention Learning Tracker (HALT) for visual tracking in complex environments, which is capable of learning rich hierarchical features. Extensive experiments demonstrate that, compared to state-of-the-art trackers, our HALT achieves outstanding tracking performance across multiple benchmarks while maintaining a real-time speed of 48.5 fps.
Fengwei Gu, Ao Li 0002, Chen Chen 0086, Chengtao Cai, Renjie Qiao, Guangyao Zhai, Zhaojie Ju
IEEE Trans Autom. Sci. Eng.5
2025 Multi-UAV intelligent decision-making method with layer delay dual-center MAPPO for air combat
Zhengkun Ding, Xingmei Wang 0002, Chengtao Cai, Luyu Jia
Appl. Intell.3
2025 A Transformer based on Voxel Spatial-Channel Attention for 3D object detection
Guangyu Ji, Chengtao Cai, Kaibin Qin
Pattern Recognit. Lett.3
2025 Practical Prescribed-Time Control for Underactuated Marine Surface Vessels: Theory and Experiment
abstract
This paper proposes a practical prescribed-time control strategy for underactuated marine surface vessels (UMSVs) with fore-aft asymmetry. A sufficient condition for practical prescribed-time stability (PPTS) is established through a novel time-varying gain function that maps the settling time into a single user-defined parameter. First, the nonholonomic constraint of the UMSV is transformed into an integral cascade system via the hand position approach, which addresses the design challenge from underactuation. Subsequently, input saturation nonlinearity is approximated by a sigmoid function, while model uncertainties are compensated through an adaptive parametric design paradigm. A barrier function-integrated adaptive law is further developed to bound the adaptive parameter without overestimation. Rigorous stability analysis confirms that the tracking error converges to an adjustable bounded domain within the prescribed time. Experimental results validate the efficacy of the proposed controller.
Daohui Zeng, Chengtao Cai, Yongchao Liu 0002, Jie Zhao 0044
IEEE Trans Autom. Sci. Eng.2
2025 Adaptive Output Feedback Control of Underactuated Marine Surface Vehicles Under Input Saturation
abstract
This study addresses the tracking control issue of underactuated marine surface vehicles (UMSVs) with parameter and external uncertainties, input saturation, and unmeasurable velocity. An adaptive output feedback control scheme is developed without assuming the fore-aft symmetry of the hull. First, a state observer is developed to estimate the unmeasurable velocity. Next, the UMSV model is transformed into an integral cascade form using the hand position approach to overcome the design difficulties caused by the underactuated feature and asymmetric hull characteristics. Then, an adaptive auxiliary dynamic system is designed to solve the problem of input saturation caused by actuator constraints. In addition, the Lyapunov theory is applied to demonstrate the capability of the proposed control scheme to ensure the boundedness of the observation and tracking errors in the control system. Finally, the effectiveness of the developed control scheme is verified through simulation.
Daohui Zeng, Chengtao Cai, Yongchao Liu 0002, Jie Zhao 0044
IEEE Trans. Intell. Transp. Syst.2
2025 Dbanet: a dual branch aggregation network for real-time semantic segmentation of omnidirectional images in maritime environments
abstract
Abstract We introduce DBANet, a dual-branch aggregation network designed for efficient and real-time semantic segmentation of omnidirectional images in maritime environments. To support research and evaluation in this area, we also present the maritime omnidirectional semantic segmentation dataset, which fills the gap in maritime omnidirectional image segmentation. While omnidirectional vision systems are increasingly popular for their 360-degree perception capabilities, their large field of view imposes significant computational demands, and comprehensive evaluation methods for semantic segmentation in such scenarios remain limited. Our approach addresses these challenges by providing a robust and computationally efficient solution applicable to intelligent perception for maritime surface vehicles. Experimental results highlight the performance of DBANet, achieving 92.36 mIoU at 4.94 FPS on the MODSS dataset and 85.08 mIoU at 30.25 FPS on the MaSTr1325 dataset, outperforming state-of-the-art models in both accuracy and efficiency.
Chengtao Cai, Jinwhan Kim, Renjie Qiao
J. Supercomput.2
2025 ADH-YOLO: a small object detection based on improved YOLOv8 for airport scene images in hazy weather
Chengtao Cai, Sutthiphong Srigrarom, Zijian Cui
J. Supercomput.2
2025 Multi-Scale Dynamic Fusion for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) aims to match persons across visible and infrared modalities; however, its performance is prone to complex dynamic scenes, such as occlusions, background shifts, and pose changes. In this paper, we propose a Multi-scale Dynamic Fusion Network (MDFN) to address these challenges in the VI-ReID task. Specifically, the proposed MDFN consists of the Dynamic Feature Fusion (DFF), Dynamic Perception Enhancement (DPE), and Feature Reweighting with Similarity (FRS) modules. The DFF module dynamically extracts local and long-range dependencies among features to obtain finer-grained discriminative features. The DPE module extracts multi-scale features from both visible and infrared modalities to generate diverse embeddings. The FRS module mitigates the impact of information imbalance between modalities, thereby further improving performance. Extensive experiments on the SYSU-MM01 and RegDB datasets show that our MDFN outperforms other state-of-the-art methods, especially in complex dynamic scenes with occlusions, background shifts, and pose changes.
Yu Wang 0292, Renjie Qiao, Kejun Wu, Chia-Wen Lin, Chengtao Cai
ACM Trans. Multim. Comput. Commun. Appl.6
2024 CFENet: Cost-effective underwater image enhancement network via cascaded feature extraction
Xun Ji, Chengtao Cai
Eng. Appl. Artif. Intell.4
2024 Language conditioned multi-scale visual attention networks for visual grounding
Haibo Yao, Lipeng Wang 0002, Chengtao Cai, Zhi Zhang 0008, Xiaobing Shang
Image Vis. Comput.3
2024 Propagating prior information with transformer for robust visual object tracking
Chengtao Cai, Chai Kiat Yeo
Multim. Syst.2
2024 An improved personal protective equipment detection method based on YOLOv4
Rengjie Qiao, Chengtao Cai, Haiyang Meng, Kejun Wu
Multim. Tools Appl.2
2024 OARPD: occlusion-aware rotated people detection in overhead fisheye images
Rengjie Qiao, Chengtao Cai, Haiyang Meng
Multim. Tools Appl.2
2024 Hierarchical Attention Networks for Fact-based Visual Question Answering
Haibo Yao, Zhi Zhang 0008, Jianhang Yang, Chengtao Cai
Multim. Tools Appl.5
2024 ASSD-YOLO: a small object detection method based on improved YOLOv7 for airport surface surveillance
Chengtao Cai, Liying Zheng, Daohui Zeng
Multim. Tools Appl.2
2024 Reinforced Res-Unet transformer for underwater image enhancement
Peitong Li, Chengtao Cai
Signal Process. Image Commun.3
2024 EANTrack: An Efficient Attention Network for Visual Tracking
abstract
Recently, Siamese trackers have gained widespread attention in visual tracking due to their exceptional performance. However, many trackers still suffer from limitations in challenging scenarios, such as fast motion and scale variation, which hinder the full exploitation of target features. Consequently, the accuracy and efficiency of the trackers are limited. Therefore, this paper proposes an efficient attention network, called EAN, to improve tracking performance. The EAN comprises three primary components, namely a Transformer-s subnetwork, a Transformer-t subnetwork, and a Feature-Fused Attention Module (FFAM). The designed Transformer-s and Transformer-t subnetworks adopt complementary structures and functions to fully integrate and emphasize the relevant feature information, including channel and spatial features. The FFAM is responsible for fusing the multi-level features from both subnetworks, which establishes the global dependencies between the templates and search regions and enhances the discriminative power of the model. To further improve the tracking accuracy, a novel Feature-Aware Attention Module (FAAM) is introduced into the tracking prediction head to enhance the feature representation capability of the model. Finally, we propose an efficient EANTrack tracker based on EAN for robust tracking in complex scenarios, which exhibits significant advantages in challenging attributes. Experimental results on multiple benchmarks indicate that our approach achieves remarkable tracking performance with a real-time running speed of 55.6fps.Note to Practitioners—Siamese trackers have garnered considerable attention in the field of visual tracking due to their impressive performance. However, these trackers often face limitations in challenging scenarios, which impede the complete exploitation of target features. As a result, the accuracy and efficiency of many trackers are compromised. To address these issues, we propose an efficient tracker called EANTrack to enable robust tracking in complex scenarios. Our EANTrack exhibits significant advantages in handling challenging attributes. Please refer to our complete paper for detailed information on the EANTrack tracker and experimental results. Practitioners in the field can benefit from our research by leveraging our findings and methodologies in their work. We encourage further exploration and experimentation to enhance the performance and applicability of visual tracking systems.
Fengwei Gu, Chengtao Cai, Qidan Zhu, Zhaojie Ju
IEEE Trans Autom. Sci. Eng.3
2024 RTSformer: A Robust Toroidal Transformer With Spatiotemporal Features for Visual Tracking
abstract
In complex environments, trackers are extremely susceptible to some interference factors, such as fast motions, occlusion, and scale changes, which result in poor tracking performance. The reason is that trackers cannot sufficiently utilize the target feature information in these cases. Therefore, it has become a particularly critical issue in the field of visual tracking to utilize the target feature information efficiently. In this article, a composite transformer involving spatiotemporal features is proposed to achieve robust visual tracking. Our method develops a novel toroidal transformer to fully integrate features while designing a template refresh mechanism to provide temporal features efficiently. Combined with the hybrid attention mechanism, the composite of temporal and spatial feature information is more conducive to mining feature associations between the template and search region than a single feature. To further correlate the global information, the proposed method adopts a closed-loop structure of the toroidal transformer formed by the cross-feature fusion head to integrate features. Moreover, the designed score head is used as a basis for judging whether the template is refreshed. Ultimately, the proposed tracker can achieve the tracking task only through a simple network framework, which especially simplifies the existing tracking architectures. Experiments show that the proposed tracker outperforms extensive state-of-the-art methods on seven benchmarks at a real-time speed of 56.5 fps.
Fengwei Gu, Chengtao Cai, Qidan Zhu, Zhaojie Ju
IEEE Trans. Hum. Mach. Syst.3
2024 Las-yolo: a lightweight detection method based on YOLOv7 for small objects in airport surveillance
Chengtao Cai, Kejun Wu, Biqin Gao
J. Supercomput.2
2024 Convex hull regression strategy for people detection on top-view fisheye images
Rengjie Qiao, Chengtao Cai, Haiyang Meng, Kejun Wu
Vis. Comput.2
2024 Target-aware pooling combining global contexts for aerial tracking
Chengtao Cai, Chai Kiat Yeo, Kejun Wu
Vis. Comput.2
2023 Multi-intent autonomous decision-making for air combat with deep reinforcement learning
Luyu Jia, Chengtao Cai, Xingmei Wang 0002, Zhengkun Ding, Junzheng Xu, Kejun Wu
Appl. Intell.2
2023 Multi-modal spatial relational attention networks for visual question answering
Haibo Yao, Lipeng Wang 0002, Chengtao Cai, Zhi Zhang 0008
Image Vis. Comput.3
2023 A robust attention-enhanced network with transformer for visual tracking
Fengwei Gu, Chengtao Cai
Multim. Tools Appl.3
2023 Repformer: a robust shared-encoder dual-pipeline transformer for visual tracking
Fengwei Gu, Chengtao Cai, Qidan Zhu, Zhaojie Ju
Neural Comput. Appl.3
2023 Siamese Centerness Prediction Network for Real-Time Visual Object Tracking
Chengtao Cai, Chai Kiat Yeo
Neural Process. Lett.2
2022 A hybrid algorithm for underwater image restoration based on color correction and image sharpening
Haiyang Meng, Yongjie Yan, Chengtao Cai, Renjie Qiao
Multim. Syst.3
2022 Keypoint matching using salient regions and GMM in images with weak textures and repetitive patterns
Qidan Zhu, Chengtao Cai, Haiyang Meng, Renjie Qiao
Multim. Tools Appl.3
2021 Water-air imaging: distorted image reconstruction based on a twice registration algorithm
Chengtao Cai, Haiyang Meng, Renjie Qiao
Mach. Vis. Appl.1
2021 Adaptive cropping and deskewing of scanned documents based on high accuracy estimation of skew angle and cropping value
Chengtao Cai, Haiyang Meng, Renjie Qiao
Vis. Comput.1
2018 Partially supervised anchored neighborhood regression for image super-resolution through FoE features
Fangfang Han, Chengtao Cai
Neurocomputing3
2017 Optimal Control of Carrier-Based Aircraft Steam Launching Valve
Chengtao Cai, Yujia Cui, Yanhua Liang
CISIS1
2017 Simulation of Upward Underwater Image Distortion Correction
Chengtao Cai, Yanhua Liang
CISIS1
2016 Motion deblurring from a single image
abstract
In recent years, the image processing technology has been used in various fields, such as the detection and monitoring and so on. Because the exposure time of camera is not infinitesimal, there is relative motion between the camera and the object being captured during the short exposing time, it causes the image blurry .The blurred image that we capture is worthless and useless. For dealing with this challenging but imperative issue, in this paper, we restore the blurry image from a single image. We analysize the spectrum of the Fourier transform to estimate the point spread function of the blurred image and restore the image in winner filter. In Wiener filter algorithm, different parameter has different influence on the image, we analysize the influence and how to adjust parameter according to the restored results. Some experiments have also been conducted for validating that winner filter is simple, efficient and has a good influence on restoring.
Chengtao Cai, Baolu Zhang
CSCWD1
2016 Fast image stitching based on improved SURF
abstract
This paper presents an improved method based on Speed up Robust Features (SURF) algorithm to achieve fast image stitching. As the variability of scenes lead to instability of features, expecting to obtain accurate number of features is pretty difficult and time-consuming. support vector machine (SVM) applied in this paper to predict primary threshold of determinant of Hessian matrix can conspicuously reduce detected feature points and simplify the process of features matching. This paper also combines an optimized method of image preprocessing-cylindrical projection and image interpolation to weigh the final quality of stitching image and stitching time. Several experiments are conducted to verify the performance of improved SURF.
Chengtao Cai, Yanhua Liang
CSCWD1
2016 Ship diesel engine fault diagnosis based on the SVM and association rule mining
abstract
Ship diesel engine's structure is complex, and its fault has high coupling. For this reason, we study ship diesel engine fault diagnosis from two aspects. First of all, we divided ship diesel engine system into four parts according its basic structure and fault features, the fuel system, the lubrication system, the intake and exhaust system and the cooling system, and then analyzed the fault features for the each subsystem respectively. We used the support vector machine (SVM) algorithm to classify fault data for each subsystem of ship diesel engine. So that, we could implement the fault diagnosis for the each subsystem. It reduced the complexity of the whole system fault diagnosis. Secondly, Ship diesel engine fault often occurs between different the subsystems, and the occurrence of a fault is often accompanied by other fault. We could solve this high coupling by using association rule mining and then found out the implicit association rules of the fault in whole system.
Chengtao Cai, Hongri Zong, Baolu Zhang
CSCWD1
2016 Local environments modelling and path planning for patrol robot in the substation
abstract
Substation is the hub of the power grid. Regular inspection is very crucial in order to confirm the electrical equipment in substation to operate normally. Recently, manual inspection is gradually replaced by inspection robot. Local environment modelling and path planning are mainly key problems when patrol robot carry out the inspection work. For dealing with this challenging but imperative issue, there are numerous researchers have strove for this scientific field and have proposed some valuable approaches. The LIDAR is one of excellent sensors for environment perception and collision avoidance for robot. For enhancing the suitability of path planning, a novel local environment modelling method is proposed in which the safety, accessibility, stability and reachability are taken into account when the patrol robot moves in unknown environment, one robot control algorithm which meet the kinematics and dynamics motion principle is investigate as well. Some simulation experiments have also been conducted for validating modelling and planning performance of the proposed approach.
Yanhua Liang, Chengtao Cai, Guo-xiang Chang, Xiao-long Lv
CSCWD2
2016 A non-linear combination filtering algorithm in MINS/GNSS navigation system
abstract
Considering the nonlinear and uncertainty in the MINS/GNSS navigation system, a nonlinear Sage-Husa noise maximum posterior estimator was designed. Since the estimator cannot solve problems both of system noise and observation, a Bi-parallel BP neural network controller is designed to approximate the estimator. Then an adaptive UKF algorithm based on Bi-parallel neural network is proposed. In the case of uncertain noise, the simulation and analysis were shown that data saturation was emerged in the preceding filtering algorithm. A strong tracking UKF algorithm based on variance inflation factor was combined with the preceding algorithm in some conversion condition. The simulation in the paper was shown that the combined adaptive nonlinear filtering algorithm could suppress divergence and ensure precision.
Chengtao Cai
CSCWD3
2016 An adaptive factor-based method for improving dark channel prior dehazing
abstract
The scattering effects of the atmospheric particles in the air affects significantly contrast reduction and color fading. To address this challenging, many attention have been paid to this issue. The foggy image generally contains the sky and non-sky regions while the pixel values in this two distinguished regions is different. The dark channel prior algorithm has been considered as one effective dehazing method which only uses one constant factor for the overall image regardless of the scene pattern. This imprudent procedure results in more darkness image color and fails to accomplish excellent results. In this paper we propose one adaptive factor-based approach to improving dark channel prior dehazing. In our methods, the foggy image is segmented into sky region and non-sky region by Otsu, the critical parameters i.e. light intensity and transmission ratio are obtained based on different factors. Some experiments have been conducted for validating dehazing performance of the proposed approach.
Chengtao Cai
CSCWD2
2016 A measurement methods for angular motion of body rotation based on only-accelerometer coplanar configuration
abstract
In the paper, an angular velocity measurement based on only-acceleration was proposed to deal with the problem that the angular motion of autobiography carrier could not be measured through gyro. Comparing the merit and demerit of traditional configurations, a new coplanar configuration of only-acceleration was designed based on the relativity between Vertical accelerometer and parallel accelerometer. Given the observing accelerometers selection principle, the angular motion measurement based on only-acceleration coplanar configuration was designed. Finally, according to the aerodynamic coefficients, simulation experiment was performed. And the result of the simulation was shown that, in the 180s time of flight, the angular velocity measurement error of this method was equivalent to 10(°)/h gyro drift.
Chengtao Cai
CSCWD3
2016 A novel mismatching elimination algorithm based on distribution of features
abstract
Catadioptric panoramic image's application to computer visual field gains its popularity in recent years. However, due to its complicated imaging relationship, most existing mismatching elimination algorithms cannot directly operate on the unprocessed panoramic images. Those above algorithms usually need to unwarp the panoramic images before further processing. In order to solve the above problems, based on the distribution characteristics of features in the panoramic image, a novel mismatching elimination algorithm is proposed in this paper. Under different scene conditions, the novel algorithm can eliminate the mismatching features and improve the matching accuracy effectively. Experiments on the image databases confirm its effectiveness.
Qidan Zhu, Chuanjia Liu, Chengtao Cai
CSCWD3
2015 Identifying Connectome Module Patterns via New Balanced Multi-graph Normalized Cut
Hongchang Gao, Chengtao Cai, Lin Yan 0003, Joaquín Goñi, Feiping Nie 0001, John D. West, Andrew J. Saykin, Li Shen 0001, Heng Huang 0001
MICCAI (2)2
2014 Non-local neighbor embedding for image super-resolution through FoE features
Qidan Zhu, Chengtao Cai
Neurocomputing3