Aibin Chen

dblp:96/7820 · DBLP profile ↗
← Back
34ranked-venue papers
0as first author
34since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 16 since 2021Artificial intelligence and machine learning · 12 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Decoding dog barking emotion from non-periodicity: A heterogeneous dual-driven mixture model with selective state space model-enhanced frequency representation
Choujun Yang, Jizheng Yi, Guoxiong Zhou, Aibin Chen
Eng. Appl. Artif. Intell.6
2026 Bridging modal gaps in multimodal sentiment analysis: a self-supervised text-guided fusion approach
Hepeng Zhong, Jizheng Yi, Ronglong Hu, Aibin Chen, Guangjie Han
Expert Syst. Appl.4
2026 DACNN: A dual-attention convolutional neural network for aspect term extraction
Zhongxiang Lin, Jingming Dai, Aibin Chen, Yurong Sun
Knowl. Based Syst.5
2026 Efficient industrial anomaly detection via cross-scale distillation with enhanced feature compression
Ronglong Hu, Jizheng Yi, Aibin Chen, Guangjie Han
Pattern Recognit.4
2025 LRM-MVSR: A lightweight birdsong recognition model based on multi-view feature extraction enhancement and spatial relationship capture
Zhongxiang Lin, Zhiqi Zhu, Wanhong Yang, Aibin Chen, Yurong Sun
Expert Syst. Appl.5
2025 MoVis: When 3D Object Detection Is Like Human Monocular Vision
abstract
Monocular 3D object detection has garnered significant attention for its outstanding cost effectiveness compared with multi-sensor systems. However, previous work mainly acquires object 3D properties in a heuristic way, with less emphasis on the cues between objects. Inspired by the mechanisms of monocular vision, we propose MoVis, an innovative 3D object detection framework that skillfully combines object hierarchy and color sequence cues. Specifically, a decoupled Spatial Relationship Encoder (SRE) is designed to effectively feed back the high-level encoding results with object hierarchical relationships to low-level features. This method not only effectively reduces the computational overhead of multi-scale coding, but also significantly improves the detection accuracy of occluded objects by incorporating the hierarchical relationship between objects into multi-scale features. Moreover, to obtain more precise object depth information, an Object-level Depth Modulator (ODM) based on the concept of conditional random fields is designed, which employs color sequences. Ultimately, the results of the SRE and ODM are efficiently fused by our Spatial Context Processor (SCP) to accurately perceive the 3D attributes of the objects. Extensive experiments on the KITTI and Rope3D benchmarks show that MoVis achieves state-of-the-art performance. Our MoVis represents a progressive approach that emulates how human monocular vision utilizes monocular cues to perceive 3D scenes.
Jizheng Yi, Aibin Chen, Guangjie Han
IEEE Trans. Image Process.3
2025 Hybrid DQN-Based Low-Computational Reinforcement Learning Object Detection With Adaptive Dynamic Reward Function and ROI Align-Based Bounding Box Regression
abstract
Deep reinforcement learning-based object detection approaches center around a pivotal concept: hierarchically scaling image segments that harbor more intricate details. Compared with the traditional object detection approaches, this approach significantly curbs the quantity of region proposals. This reduction holds paramount significance in curtailing the computational overhead. However, common deep reinforcement learning-based approaches suffer from a significant defect in terms of precision. This issue arises from inadequacies in representing image states appropriately and the unstable learning ability exhibited by the agent. To address these issues, we present the LHAR-RLD. First, we design the Low-dimensional RepVGG(LDR) feature extractor to reduce memory consumption and to reduce the difficulty of fitting downstream networks. Second, we propose the Hybrid DQN(HDQN) to enhance the agent's ability to determine the state-action of images in complex environments. Then, the Adaptive Dynamic Reward Function(ADR) is crafted to dynamically adjust the reward based on shifts within the agent's exploration environment. Finally, the ROI Align-based bounding box regression network (RABRNet) is proposed, which aims at further regressing the localization results of reinforcement learning to improve the detection precision. Our method accomplishes 74.4% mAP on the VOC2007, 76.2% mAP on the COCO2017, 75.2% Precision on the SF dataset, with 1.43G FLOPs. The precision outperforms the advanced deep reinforcement learning approaches and the computational cost is far lower than theirs and mainstream object detection methods. This method facilitates highly accurate object localization with minimal computational demands, which means it has notable applications on resource-constrained devices.
Guangjie Han, Guoxiong Zhou, Yongfei Xue, Mingjie Lv, Aibin Chen
IEEE Trans. Image Process.6
2024 Buffer-text: Detecting arbitrary shaped text in natural scene image
Jizheng Yi, Aibin Chen, Ze Jin
Eng. Appl. Artif. Intell.3
2024 A high-quality self-supervised image denoising method based on SDDW-GAN and CHRNet
Guoxiong Zhou, Aibin Chen, Liujun Li
Expert Syst. Appl.4
2024 JL-TFMSFNet: A domestic cat sound emotion recognition method based on jointly learning the time-frequency domain and multi-scale features
Shipeng Hu, Choujun Yang, Aibin Chen, Guoxiong Zhou
Expert Syst. Appl.5
2024 A barking emotion recognition method based on Mamba and Synchrosqueezing Short-Time Fourier Transform
Choujun Yang, Shipeng Hu, Guoxiong Zhou, Jizheng Yi, Aibin Chen
Expert Syst. Appl.7
2024 SMFE-Net: a saliency multi-feature extraction framework for VHR remote sensing image classification
Junsong Chen, Jizheng Yi, Aibin Chen, Ze Jin
Multim. Tools Appl.3
2024 Combined CNN LSTM with attention for speech emotion recognition based on feature-level fusion
Yanlin Liu, Aibin Chen, Guoxiong Zhou, Jizheng Yi, Jin Xiang
Multim. Tools Appl.2
2024 Parallel attention of representation global time-frequency correlation for music genre classification
Zhifang Wen, Aibin Chen, Guoxiong Zhou, Jizheng Yi, Weixiong Peng
Multim. Tools Appl.2
2024 Multichannel Cross-Modal Fusion Network for Multimodal Sentiment Analysis Considering Language Information Enhancement
abstract
With the popularity of short videos, analyzing human emotions is crucial for understanding individual attitudes and guiding social public opinions. Consequently, multimodal sentiment analysis (MSA) has garnered significant attention in the field of human–computer interaction. The main challenge of MSA is to explore a high-quality multimodal fusion framework, as multiple modalities contribute inconsistently to sentiment prediction. However, most of the existing methods assume equal importance among different modalities, resulting in inadequate expression of the main modality. In addition, auxiliary modalities often contain redundant information, which hinders the multimodal fusion process. Therefore, we propose the multichannel cross-modal fusion network (MCFNet) to promote the multimodal fusion procedure by constructing a multichannel various modality fusion framework comprising three channels: obtaining multimodal representation through the first channel; eliminating information redundancy from auxiliary modalities via the second channel; and enhancing significance attributed to the main modality adopting the third channel. Subsequently, we design a multichannel information fusion gate to integrate feature representations from these three channels for downstream sentiment classification tasks. Numerous experiments on three benchmark datasets, CMU-multimodal opinion sentiment intensity (MOSI), CMU-multimodal opinion sentiment and emotion intensity (MOSEI), and Twitter2019, show that the MCFNet has made a significant progress compared to recent state-of-the-art methods.
Ronglong Hu, Jizheng Yi, Aibin Chen, Lijiang Chen
IEEE Trans. Ind. Informatics3
2023 Identification of tomato leaf diseases based on LMBRNet
Guoxiong Zhou, Aibin Chen, Liujun Li, Yahui Hu
Eng. Appl. Artif. Intell.3
2023 MMFNet: Forest Fire Smoke Detection Using Multiscale Convergence Coordinated Pyramid Network With Mixed Attention and Fast-Robust NMS
abstract
There is a problem in the field of early automatic detection of forest fire smoke that due to low concentration or tiny size, some smoke is difficult to capture. This article proposes a multiscale convergence coordinated pyramid network (MCCPN) with mixed attention and Fast-robust NMS (MMFNet) for the fast detection of forest fire smoke. First, an MCCPN is designed, which combines a dual-attention feature pyramid network and a coordinated convergence module. It improves the detection rate of targets of different sizes. Second, a mixed attention module is designed to focus more on the smoke in the image and enhance the extraction of horizontal and vertical features of smoke. Then, a Fast-robust nonmaximum suppression is proposed to accelerate the convergence of bounding boxes and increase the accuracy of the prediction box. Finally, a forest fire detection system of the Internet of Things using MMFNet is built. The experimental results show that our method achieves 80.72% AP, 88.52% AP50, 83.45% AP75, 46.88% AR, and 154 FPS, which is superior to the state-of-art forest fire smoke detection methods.
Liangji Zhang, Haiwen Xu, Aibin Chen, Liujun Li, Guoxiong Zhou
IEEE Internet Things J.4
2023 Fabric defect detection based on separate convolutional UNet
Jizheng Yi, Aibin Chen
Multim. Tools Appl.3
2023 Tomato Leaf Disease Detection System Based on FC-SNDPN
Xibei Huang, Aibin Chen, Guoxiong Zhou, Xin Zhang 0056, Jianwu Wang 0002, Ning Peng, Canhui Jiang
Multim. Tools Appl.2
2023 Classification methods of butterfly images based on U-net and STL-MSDNet
Jin Xiang, Aibin Chen, Guoxiong Zhou
Multim. Tools Appl.3
2023 Rapid computer vision detection of apple diseases based on AMCFNet
Liangji Zhang, Guoxiong Zhou, Aibin Chen, Ning Peng
Multim. Tools Appl.3
2023 MDMASNet: A dual-task interactive semi-supervised remote sensing image segmentation method
abstract
Remote sensing image (RSIs) segmentation is widely used in urban planning, natural disaster detection and many other fields. Compared with natural scene images, RSIs have higher resolution, complex imaging, and diverse object shapes and sizes, while semantic segmentation methods based on deep learning often require many data labels. In this paper, we propose a semi-supervised RSIs segmentation network with multi-scale deformable threshold feature extraction module and mixed attention (MDMANet). First, a pyramid ensemble structure is used, which incorporates deformable convolution and bole convolution, to extract features of objects with different shapes and sizes and reduce the influence of redundant features. Meanwhile, a mixed attention (MA) is proposed to aggregate long-range contextual relationships and fuse low-level features with high-level features. Second, an FCN-based full convolution discriminator task network is designed to help evaluate the feasibility of unlabeled image prediction results. We performed experimental validation on three datasets, and the results show that MDMANet segmentation provides more significant improvement in accuracy and better generalization than existing segmentation networks.
Liangji Zhang, Zaichun Yang, Guoxiong Zhou, Aibin Chen, Yao Ding 0010, Liujun Li, Weiwei Cai 0001
Signal Process.5
2023 EFCOMFF-Net: A Multiscale Feature Fusion Architecture With Enhanced Feature Correlation for Remote Sensing Image Scene Classification
abstract
Remote sensing images have the essential attribute of large-scale spatial variation and complex scene information, as well as the high similarity between various classes and the significant differences within same class, which are easy to cause misclassification. To solve this problem, an efficient systematic architecture named EFCOMFF-Net (Multi-scale Feature Fusion Network with Enhanced Feature Correlation) is proposed to reduce the gap among multi-scale features and fuse them to improve the representation ability of remote sensing images. Firstly, to strengthen the correlation of multi-scale features, a Feature Correlation Enhancement Module (FCEM) is specifically developed, which takes the features of different stages of the backbone network as input data to obtain multi-scale features with enhanced correlation. Considering the differences between the shallow features and the deep features, the EFCOMFF-Net-v1 related to shallow features and EFCOMFF-Net-v2 related to deep features with different structures are proposed. Secondly, the designed two versions of the deep learning network focus on the global contour information and need to encode more accurate spatial information. A Feature Aggregation Attention Module (FAAM) is designed and embedded into the network to encode the deep features by applying the spatial information aggregation features. Finally, considering that the simple integration strategy cannot reduce the gap between the shallow multi-scale features and the deep features, a Feature Refinement Module (FRM) is presented to optimize the network. ResNet50, DenseNet121, and ResNet152 are selected to conduct a considerable number of experiments on four datasets, which show the superiority of our method compared to recent methods.
Junsong Chen, Jizheng Yi, Aibin Chen, Ze Jin
IEEE Trans. Geosci. Remote. Sens.3
2023 SRCBTFusion-Net: An Efficient Fusion Architecture via Stacked Residual Convolution Blocks and Transformer for Remote Sensing Image Semantic Segmentation
abstract
Convolutional neural network (CNN) and Transformer-based self-attention models have their advantages in extracting local information and global semantic information, and it is a trend to design a model combining stacked residual convolution blocks (SRCB) and Transformer. How to efficiently integrate the two mechanisms to improve the segmentation effect of remote sensing (RS) images is an urgent problem to be solved. An efficient fusion via SRCB and Transformer (SRCBTFusion-Net) is proposed as a new semantic segmentation architecture for RS images. The SRCBTFusion-Net adopts an encoder-decoder structure, and the Transformer is embedded into SRCB to form a double coding structure, then the coding features are up-sampled and fused with multi-scale features of SRCB to form a decoding structure. Firstly, a semantic information enhancement module (SIEM) is proposed to get global clues for enhancing deep semantic information. Subsequently, the relationship guidance module (RGM) is incorporated to re-encode the decoder’s upsampled feature maps, enhancing the edge segmentation performance. Secondly, a multipath atrous self-attention module (MASM) is developed to enhance the effective selection and weighting of low-level features, effectively reducing the potential confusion introduced by the skip connections between low-level and high-level features. Finally, a multi-scale feature aggregation module (MFAM) is developed to enhance the extraction of semantic and contextual information, thus alleviating the loss of image feature information and improving the ability to identify similar categories. The proposed SRCBTFusion-Net’s performance on the Vaihingen and Potsdam datasets is superior to the state-of-the-art methods. The code will be freely available at https://github.com/js257/SRCBTFusion-Net.
Junsong Chen, Jizheng Yi, Aibin Chen
IEEE Trans. Geosci. Remote. Sens.3
2023 DGPF-RENet: A Low Data Dependence Network With Low Training Iterations for Hyperspectral Image Classification
abstract
The classification of ground objects from hyperspectral images (HSIs) is of great importance for human perception of information about the terrain and landscape. HSIs have numerous dimensions, and obtaining the data is difficult. The issue of slow convergence of neural network training is brought on by high dimensional data, and the neural network’s performance is impacted by the challenging data acquisition process. In order to achieve the effects of low data dependence and rapid convergence, we propose a redundancy elimination network architecture with decoupled-gaze attention mechanism and phantom fractal modules (DGPF-RENet) for HSIs classification. First, we propose the decoupled-gaze attention mechanism (DGA) to make full use of correlation between adjacent bands and the continuity of neighboring pixels in HSIs. Then, a redundancy elimination module (REM) is proposed to reduce the number of feature points and eliminate redundant information while preserving the contextual information and relationships between pixels. Finally, the phantom fractal module (PFM) is proposed, which improves the scale of feature learning by fractalising convolutions at multiple scales. Four publicly available HSIs datasets, including Indian Pines, Salinas, DFC2018, and WHUHi-HongHu, were used in our experiments. According to experimental findings, when compared to other state-of-the-art methods, our method performs best with a small number of training samples and few iterations. We have released our code and models at https://github.com/yuhua666/DGPF-RENet.
Jialei Zhan, Yaowen Hu, Guoxiong Zhou, Weiwei Cai 0001, Aibin Chen, Liu Xie, Maopeng Li, Liujun Li
IEEE Trans. Geosci. Remote. Sens.8
2022 ConvPatchTrans: A script identification network with global and local semantics deeply integrated
Jizheng Yi, Aibin Chen, Ze Jin
Eng. Appl. Artif. Intell.3
2022 Fast forest fire smoke detection using MVMNet
abstract
Forest fires are a huge ecological hazard, and smoke is an early characteristic of forest fires. Smoke is present only in a tiny region in images that are captured in the early stages of smoke occurrence or when the smoke is far from the camera. Furthermore, smoke dispersal is uneven, and the background environment is complicated and changing, thereby leading to inconspicuous pixel-based features that complicate smoke detection. In this paper, we propose a detection method called multioriented detection based on a value conversion-attention mechanism module and Mixed-NMS (MVMNet). First, a multioriented detection method is proposed. In contrast to traditional detection techniques, this method includes an angle parameter in the data loading process and calculates the target’s rotation angle using the classification prediction method, which has reference significance for determining the direction of the fire source. Then, to address the issue of inconsistent image input size while preserving more feature information, Softpool-spatial pyramid pooling (Soft-SPP) is proposed. Next, we construct a value conversion-attention mechanism module (VAM) based on the joint weighting strategy in the horizontal and vertical directions, which can specifically extract the colour and texture of the smoke. Ultimately, the DIoU-NMS and Skew-NMS hybrid nonmaximum suppression methods are employed to address the issues of smoke false detection and missed detection. Experiments are conducted using the homemade forest fire multioriented detection dataset, and the results demonstrate that compared to the traditional detection method, our model’s mAP reaches 78.92%, mAP 50 reaches 88.05%, and FPS reaches 122.
Yaowen Hu, Jialei Zhan, Guoxiong Zhou, Aibin Chen, Yahui Hu, Liujun Li
Knowl. Based Syst.4
2022 ConDinet++: Full-Scale Fusion Network Based on Conditional Dilated Convolution to Extract Roads From Remote Sensing Images
abstract
Extracting roads from aerial images is an issue that has attracted much attention. Using semantic segmentation methods to extract roads often faces the problem of narrow and occluded roads. In this letter, we propose a network called ConDinet++, which improves the general codec architecture. In the encoder part, the VGG16 with pretraining parameters is utilized for the feature extraction. In the decoder part, we perform a feature fusion mechanism on the full-scale feature map. In order to improve the ability of the network to extract and integrate semantic information and further increase the receptive field, we recommend adopting the conditional dilated convolution blocks (CDBs) in the encoder, and each CDB consists of a group of cascaded conditional dilated convolutions. More importantly, the designed codec architecture can adjust the number of convolutions and the parameters of the convolution kernel according to the input data. For a slender area like a road, which occupies a small area in the picture, we use the joint loss function and introduce the joint loss of Lovasz loss and cross-entropy loss to avoid the segmentation model having a serious bias caused by highly unbalanced object sizes between roads and background. The proposed method was tested on two public datasets Massachusetts Roads Dataset and Mini DeepGlobe Road Extraction Challenge. Compared with some previous semantic segmentation networks, the proposed ConDinet++ achieved the best values of recall, F-score, and mIoU.
Jizheng Yi, Aibin Chen
IEEE Geosci. Remote. Sens. Lett.3
2022 Remote sensing scene classification with multi-spatial scale frequency covariance pooling
Yuan Gao 0065, Aibin Chen, Guoxiong Zhou, Jianwu Wang 0002
Multim. Tools Appl.3
2022 Pulmonary nodules recognition based on parallel cross-convolution
Yaowen Hu, Jialei Zhan, Guoxiong Zhou, Aibin Chen, Jiayong Li
Multim. Tools Appl.4
2022 Birdsong classification based on multi feature channel fusion
Aibin Chen, Guoxiong Zhou, Jizheng Yi
Multim. Tools Appl.3
2021 The fruit classification algorithm based on the multi-optimization convolutional neural network
Guoxiong Zhou, Aibin Chen, Ling Pu
Multim. Tools Appl.3
2021 Dermatoscopic image melanoma recognition based on CFLDnet fusion network
Aibin Chen, Guoxiong Zhou, Ning Peng
Multim. Tools Appl.2
2021 Birdsong classification based on multi-feature fusion
Aibin Chen, Guoxiong Zhou, Jianwu Wang 0002
Multim. Tools Appl.2