EDBT 2026 Demo / reviewers in the wild / expert
Jizheng Yi
dblp:155/6728
· DBLP profile ↗
22ranked-venue papers
2as first author
20since 2021 · last 2026
0000-0002-2360-905XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decoding dog barking emotion from non-periodicity: A heterogeneous dual-driven mixture model with selective state space model-enhanced frequency representation
Choujun Yang, Jizheng Yi, Guoxiong Zhou, Aibin Chen |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | UKANCNet: Multi-scale feature fusion with UKAN enhancement for micro wind turbine blade defect segmentation
Jizheng Yi, Xiangyu Shen, Lijiang Chen, Ze Jin |
Expert Syst. Appl. | 2 |
| 2026 | Bridging modal gaps in multimodal sentiment analysis: a self-supervised text-guided fusion approach
Hepeng Zhong, Jizheng Yi, Ronglong Hu, Aibin Chen, Guangjie Han |
Expert Syst. Appl. | 2 |
| 2026 | Efficient industrial anomaly detection via cross-scale distillation with enhanced feature compression
Ronglong Hu, Jizheng Yi, Aibin Chen, Guangjie Han |
Pattern Recognit. | 2 |
| 2025 | Agricultural surface water extraction in environmental remote sensing: A novel semantic segmentation model emphasizing contextual information enhancement and foreground detail attention
Pengyu Lei, Jizheng Yi, Yasheng Li |
Neurocomputing | 2 |
| 2025 | A Lightweight Semantic Segmentation Network Based on Self-Attention Mechanism and State Space Model for Efficient Urban Scene SegmentationabstractIn the semantic segmentation of remote sensing images, methods based on Convolutional Neural Networks (CNN) and Transformers have been extensively studied. Nevertheless, CNN struggles to capture global context due to its local feature extraction, while Transformer is constrained by the complexity of quadratic calculations. Recently, there has been a great deal of interest in Mamba-based state space models. However, the existing Mamba-based methods do not adequately consider the significance of local information in remote sensing image segmentation tasks. In this paper, a codec style network UMFormer is constructed for the semantic segmentation of remote sensing image. Specifically, UMFormer employs the ResNet18 as the encoder, with the objective of performing a preliminary image feature extraction. Subsequently, a self-attention mechanism is optimized to extract the global information pertaining to the objects of disparate sizes within the context of a multi-scale condition. For fusing the codec feature map information, another attention structure is built to reconstruct the space information and to capture the relative position relationship. Finally, a decoder based on Mamba is designed to effectively model both global and local information. Concurrently, a feature fusion mechanism utilizing feature similarity is devised with the objective of embedding local information into global ones. Numerous experiments on UAVid, Vaihingen and Potsdam datasets have demonstrated that the proposed UMFormer exhibits enhanced accuracy while maintaining an efficient running speed. The code will be freely available at https://github.com/takeyoutime/UMFormer. Langping Li, Jizheng Yi |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Graph Reconstruction Attention Fusion Network for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) has become increasingly popular due to the exponential surge of user comments on social media. The MSA aims to efficiently integrate various modalities through a superior fusion framework. However, previous studies have primarily focused on the integration of sequence data while neglecting its structural information. In addition, effectively modeling the continuous expression of human sentiment polarity remains a significant challenge. Therefore, we propose the graph reconstruction attention fusion network, which availably promotes the multimodal fusion process by combining sequence learning with graph learning. First, we design a graph reconstruction learning module to obtain multimodal graph embeddings. Second, a text-guided cross-modal enhancement architecture is adopted to acquire multimodal representations, where a sentiment attenuation factor is introduced to promote emotional continuity modeling. Finally, we propose a feature-wised attention structure adapted for the classifier, it dynamically adjusts weights of multimodal features that are beneficial for downstream tasks. Extensive experiments on three challenging datasets, CMU-MOSI, CMU-MOSEI, and CH-SIMS demonstrate that our model significantly outperforms existing state-of-the-art methods. Ronglong Hu, Jizheng Yi, Lijiang Chen, Ze Jin |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | MoVis: When 3D Object Detection Is Like Human Monocular VisionabstractMonocular 3D object detection has garnered significant attention for its outstanding cost effectiveness compared with multi-sensor systems. However, previous work mainly acquires object 3D properties in a heuristic way, with less emphasis on the cues between objects. Inspired by the mechanisms of monocular vision, we propose MoVis, an innovative 3D object detection framework that skillfully combines object hierarchy and color sequence cues. Specifically, a decoupled Spatial Relationship Encoder (SRE) is designed to effectively feed back the high-level encoding results with object hierarchical relationships to low-level features. This method not only effectively reduces the computational overhead of multi-scale coding, but also significantly improves the detection accuracy of occluded objects by incorporating the hierarchical relationship between objects into multi-scale features. Moreover, to obtain more precise object depth information, an Object-level Depth Modulator (ODM) based on the concept of conditional random fields is designed, which employs color sequences. Ultimately, the results of the SRE and ODM are efficiently fused by our Spatial Context Processor (SCP) to accurately perceive the 3D attributes of the objects. Extensive experiments on the KITTI and Rope3D benchmarks show that MoVis achieves state-of-the-art performance. Our MoVis represents a progressive approach that emulates how human monocular vision utilizes monocular cues to perceive 3D scenes. Jizheng Yi, Aibin Chen, Guangjie Han |
IEEE Trans. Image Process. | 2 |
| 2024 | Buffer-text: Detecting arbitrary shaped text in natural scene image
Jizheng Yi, Aibin Chen, Ze Jin |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | A barking emotion recognition method based on Mamba and Synchrosqueezing Short-Time Fourier Transform
Choujun Yang, Shipeng Hu, Guoxiong Zhou, Jizheng Yi, Aibin Chen |
Expert Syst. Appl. | 6 |
| 2024 | SMFE-Net: a saliency multi-feature extraction framework for VHR remote sensing image classification
Junsong Chen, Jizheng Yi, Aibin Chen, Ze Jin |
Multim. Tools Appl. | 2 |
| 2024 | Combined CNN LSTM with attention for speech emotion recognition based on feature-level fusion
Yanlin Liu, Aibin Chen, Guoxiong Zhou, Jizheng Yi, Jin Xiang |
Multim. Tools Appl. | 4 |
| 2024 | Parallel attention of representation global time-frequency correlation for music genre classification
Zhifang Wen, Aibin Chen, Guoxiong Zhou, Jizheng Yi, Weixiong Peng |
Multim. Tools Appl. | 4 |
| 2024 | Multichannel Cross-Modal Fusion Network for Multimodal Sentiment Analysis Considering Language Information EnhancementabstractWith the popularity of short videos, analyzing human emotions is crucial for understanding individual attitudes and guiding social public opinions. Consequently, multimodal sentiment analysis (MSA) has garnered significant attention in the field of human–computer interaction. The main challenge of MSA is to explore a high-quality multimodal fusion framework, as multiple modalities contribute inconsistently to sentiment prediction. However, most of the existing methods assume equal importance among different modalities, resulting in inadequate expression of the main modality. In addition, auxiliary modalities often contain redundant information, which hinders the multimodal fusion process. Therefore, we propose the multichannel cross-modal fusion network (MCFNet) to promote the multimodal fusion procedure by constructing a multichannel various modality fusion framework comprising three channels: obtaining multimodal representation through the first channel; eliminating information redundancy from auxiliary modalities via the second channel; and enhancing significance attributed to the main modality adopting the third channel. Subsequently, we design a multichannel information fusion gate to integrate feature representations from these three channels for downstream sentiment classification tasks. Numerous experiments on three benchmark datasets, CMU-multimodal opinion sentiment intensity (MOSI), CMU-multimodal opinion sentiment and emotion intensity (MOSEI), and Twitter2019, show that the MCFNet has made a significant progress compared to recent state-of-the-art methods. Ronglong Hu, Jizheng Yi, Aibin Chen, Lijiang Chen |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | Fabric defect detection based on separate convolutional UNet
Jizheng Yi, Aibin Chen |
Multim. Tools Appl. | 2 |
| 2023 | EFCOMFF-Net: A Multiscale Feature Fusion Architecture With Enhanced Feature Correlation for Remote Sensing Image Scene ClassificationabstractRemote sensing images have the essential attribute of large-scale spatial variation and complex scene information, as well as the high similarity between various classes and the significant differences within same class, which are easy to cause misclassification. To solve this problem, an efficient systematic architecture named EFCOMFF-Net (Multi-scale Feature Fusion Network with Enhanced Feature Correlation) is proposed to reduce the gap among multi-scale features and fuse them to improve the representation ability of remote sensing images. Firstly, to strengthen the correlation of multi-scale features, a Feature Correlation Enhancement Module (FCEM) is specifically developed, which takes the features of different stages of the backbone network as input data to obtain multi-scale features with enhanced correlation. Considering the differences between the shallow features and the deep features, the EFCOMFF-Net-v1 related to shallow features and EFCOMFF-Net-v2 related to deep features with different structures are proposed. Secondly, the designed two versions of the deep learning network focus on the global contour information and need to encode more accurate spatial information. A Feature Aggregation Attention Module (FAAM) is designed and embedded into the network to encode the deep features by applying the spatial information aggregation features. Finally, considering that the simple integration strategy cannot reduce the gap between the shallow multi-scale features and the deep features, a Feature Refinement Module (FRM) is presented to optimize the network. ResNet50, DenseNet121, and ResNet152 are selected to conduct a considerable number of experiments on four datasets, which show the superiority of our method compared to recent methods. Junsong Chen, Jizheng Yi, Aibin Chen, Ze Jin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | SRCBTFusion-Net: An Efficient Fusion Architecture via Stacked Residual Convolution Blocks and Transformer for Remote Sensing Image Semantic SegmentationabstractConvolutional neural network (CNN) and Transformer-based self-attention models have their advantages in extracting local information and global semantic information, and it is a trend to design a model combining stacked residual convolution blocks (SRCB) and Transformer. How to efficiently integrate the two mechanisms to improve the segmentation effect of remote sensing (RS) images is an urgent problem to be solved. An efficient fusion via SRCB and Transformer (SRCBTFusion-Net) is proposed as a new semantic segmentation architecture for RS images. The SRCBTFusion-Net adopts an encoder-decoder structure, and the Transformer is embedded into SRCB to form a double coding structure, then the coding features are up-sampled and fused with multi-scale features of SRCB to form a decoding structure. Firstly, a semantic information enhancement module (SIEM) is proposed to get global clues for enhancing deep semantic information. Subsequently, the relationship guidance module (RGM) is incorporated to re-encode the decoder’s upsampled feature maps, enhancing the edge segmentation performance. Secondly, a multipath atrous self-attention module (MASM) is developed to enhance the effective selection and weighting of low-level features, effectively reducing the potential confusion introduced by the skip connections between low-level and high-level features. Finally, a multi-scale feature aggregation module (MFAM) is developed to enhance the extraction of semantic and contextual information, thus alleviating the loss of image feature information and improving the ability to identify similar categories. The proposed SRCBTFusion-Net’s performance on the Vaihingen and Potsdam datasets is superior to the state-of-the-art methods. The code will be freely available at https://github.com/js257/SRCBTFusion-Net. Junsong Chen, Jizheng Yi, Aibin Chen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | ConvPatchTrans: A script identification network with global and local semantics deeply integrated
Jizheng Yi, Aibin Chen, Ze Jin |
Eng. Appl. Artif. Intell. | 2 |
| 2022 | ConDinet++: Full-Scale Fusion Network Based on Conditional Dilated Convolution to Extract Roads From Remote Sensing ImagesabstractExtracting roads from aerial images is an issue that has attracted much attention. Using semantic segmentation methods to extract roads often faces the problem of narrow and occluded roads. In this letter, we propose a network called ConDinet++, which improves the general codec architecture. In the encoder part, the VGG16 with pretraining parameters is utilized for the feature extraction. In the decoder part, we perform a feature fusion mechanism on the full-scale feature map. In order to improve the ability of the network to extract and integrate semantic information and further increase the receptive field, we recommend adopting the conditional dilated convolution blocks (CDBs) in the encoder, and each CDB consists of a group of cascaded conditional dilated convolutions. More importantly, the designed codec architecture can adjust the number of convolutions and the parameters of the convolution kernel according to the input data. For a slender area like a road, which occupies a small area in the picture, we use the joint loss function and introduce the joint loss of Lovasz loss and cross-entropy loss to avoid the segmentation model having a serious bias caused by highly unbalanced object sizes between roads and background. The proposed method was tested on two public datasets Massachusetts Roads Dataset and Mini DeepGlobe Road Extraction Challenge. Compared with some previous semantic segmentation networks, the proposed ConDinet++ achieved the best values of recall, F-score, and mIoU. Jizheng Yi, Aibin Chen |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Birdsong classification based on multi feature channel fusion
Aibin Chen, Guoxiong Zhou, Jizheng Yi |
Multim. Tools Appl. | 5 |
| 2016 | Illumination compensation for facial feature point localization in a single 2D face image
Jizheng Yi, Xia Mao, Lijiang Chen, Alberto Rovetta |
Neurocomputing | 1 |
| 2014 | Facial expression recognition considering individual differences in facial structure and textureabstractFacial expression recognition (FER) plays an important role in human–computer interaction. The recent years have witnessed an increasing trend of various approaches for the FER, but these approaches usually do not consider the effect of individual differences to the recognition result. When the face images change from neutral to a certain expression, the changing information constituted of the structural characteristics and the texture information can provide rich important clues not seen in either face image. Therefore it is believed to be of great importance for machine vision. This study proposes a novel FER algorithm by exploiting the structural characteristics and the texture information hiding in the image space. Firstly, the feature points are marked by an active appearance model. Secondly, three facial features, which are feature point distance ratio coefficient, connection angle ratio coefficient and skin deformation energy parameter, are proposed to eliminate the differences among the individuals. Finally, a radial basis function neural network is utilised as the classifier for the FER. Extensive experimental results on the Cohn–Kanade database and the Beihang University (BHU) facial expression database show the significant advantages of the proposed method over the existing ones. Jizheng Yi, Xia Mao, Lijiang Chen, Yu-Li Xue, Angelo Compare |
IET Comput. Vis. | 1 |