VLDB 2026 Research / reviewers in the wild / expert
Zhenhong Jia
dblp:18/3664
· DBLP profile ↗
78ranked-venue papers
0as first author
65since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 32 since 2021Artificial intelligence and machine learning · 27 · 21 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 10 since 2021Databases, data management, data science and information retrieval · 7 · 7 since 2021Systems, architecture and hardware · 4Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | TM-Net: A temporal-aware Mamba-attention hybrid network for video desnowing
Shenhan Feng, Yanyun Zhou, Jiaohao Li, Zhenhong Jia |
Expert Syst. Appl. | 4 |
| 2026 | PGCNet: A Prototype-Guided Network for Few-Shot Pest Counting
Zhenhong Jia |
ICIC (18) | 4 |
| 2026 | The intelligent detection method for similar cotton diseases based on high- and low-frequency feature enhancement and attention-guided fusion
Zhenhong Jia, Jiajia Wang 0002 |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | An enhanced segmentation network built upon the you only look once framework for precise weed recognition in early-stage cotton
Jiajia Wang 0006, Zhenhong Jia |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | TSRA-Net: A pest detection method based on token statistics and reparameterized adaptive networks
Guohong Chen, Wu Le, Zhenhong Jia, Jiajia Wang 0006 |
Expert Syst. Appl. | 3 |
| 2026 | LW-CD: A dual-domain unsupervised framework and dataset for low-illumination wide-field change detection
Dayong Ren, Zhenhong Jia, Sensen Song, Yani Guo |
Expert Syst. Appl. | 3 |
| 2025 | Difference Bonds Consistency and Complementarity to Enhance Multimodal Representation LearningabstractIn the field of multimodal representation learning, existing research has primarily focused on exploring modal consistency and modal complementarity, while overlooking the positive role of modal difference. Moreover, modal difference establishes a bonding relationship between modal consistency and modal complementarity. However, existing algorithms lack the study of this relationship, resulting in an accuracy that still needs to be improved on multimodal classification tasks. To tackle the above issues, We propose a novel multimodal representation learning framework. It enhances multimodal representation learning by modal difference constructing connecting bonds of modal consistency and modal complementarity. We conducted experiments on two widely used multimodal emotion recognition datasets, IEMOCAP and MELD. The results demonstrate that our method outperforms existing multimodal representation learning approaches in terms of accuracy on the multimodal emotion recognition task. Congbing He, Sensen Song, Zhenhong Jia |
ICASSP | 3 |
| 2025 | Small Target Insect Detection Based on Improved YOLOv8nabstractInsect pests greatly affect the growth and harvest of crops, and accurate identification of insect species is very important to the agricultural industry. The insect dataset has the problem that the target is too small to locate, detect, and identify. To solve this problem, We designed a modified YOLOv8n for detecting small target insects. First, we add a small object detection head to the network structure of the YOLOv8n (You Only Look Once) model. Then, we improve C2f and Conv in YOLOv8n, construct C2f-RFAConv by adding Receptive-Field Attention Convolutional Operation (RFAConv) into C2f for the first time, and replace the original convolution with RFAConv for feature extraction. Finally, we improve Wise-IoU (WIoU) loss function to solve the problem of small object detection difficulty. When the number of parameters is reduced, the performance of our model YOLOv8n-Improved on the Yellow Sticky Traps dataset and self-built dataset is greatly improved compared with YOLOv8n, and it also has great advantages compared with other models in the YOLO series. This work plays an important role in our efforts to achieve intelligent pest management in agriculture. Jianyu Shi, Yuan Jia, Zhenhong Jia |
ICASSP | 5 |
| 2025 | Change Detection for Wide-Field Video Images in Foggy Weather Based on Enhanced K-Means Clustering
Yankai Cao, Xiaoqian Qu, Sensen Song, Zhenhong Jia |
ICIC (1) | 5 |
| 2025 | TFRNet: A Text-Focused Snow Removal Network for Scene Text Recognition
Zhenhong Jia, Mengnan Zhang |
ICIC (2) | 4 |
| 2025 | RGPest-YOLO: A YOLOv8 Pest Detection Method Based on Image Preprocessing
Xiaoqian Qu, Yankai Cao, Zhenhong Jia |
ICIC (1) | 3 |
| 2025 | A Image Restoration Network for Nighttime Variegated Haze Conditions
Yuwei Feng, Linghui Ma, Zhenhong Jia |
ICIC (3) | 6 |
| 2025 | A dual branch graphic text detection network based on progressive Domain adaptationabstractIn recent years, scene text detection methods have been widely studied and have achieved good results on various common text datasets. However, graphic text commonly found in movie posters and clothing prints tends to have irregular appearances, leading to a decline in the performance of existing text detection methods. To address this issue, we propose a dual-branch graphic text detection network based on progressive domain adaptation (GTDNet++). We achieve feature alignment between the source and target domains through progressive domain adaptation in one branch. Additionally, we use an attention fusion module to fully combine the features from both branches, further reducing domain shift and addressing the issue of source domain forgetting. Extensive experiments on this dataset demonstrate that GTDNet++ achieves state-of-the-art performance on graphic text and mixed text scenarios. Yuwei Feng, Zhenhong Jia |
ICIP | 6 |
| 2025 | A Depth Semantic Perception Network for Camouflage Object DetectionabstractCamouflage object detection (COD) aims to identify objects that blend seamlessly into complex backgrounds, making it inherently more challenging than conventional object detection. However, most existing COD methods fail to strike a balance between local details and global understanding, leading to false positives or missed detections by the model. To address this issue, we propose a Depth Semantic Perception Network (DSP-Net), which can capture depth semantic information from images and accurately model the spatial relationship between objects and their backgrounds, thereby improving detection accuracy. Specifically, we first propose a Semantic Localization Module (SLM) that integrates multi-scale features from the backbone network and obtains rich semantic information, which is used to modulate the fused features obtained in the Cross-Level Fusion Module (CLFM) to generate higher-level semantic representations. Next, we propose an Adaptive Channel Fusion Module (ACFM), which aggregates multi-scale features through weighted fusion of input features and dynamically learns the weights for each channel. Finally, we propose a Cross-Level Fusion Module (CLFM), which captures depth semantic information through prior guidance and cross-level feature fusion, and balances the information against high spatial resolution information, thereby enhancing the accuracy of COD. Extensive experiments on four benchmark COD datasets show that our DSP-Net outperforms other state-of-the-art models. Zijun Wei, Zhenhong Jia, Haochu Ku |
ICME | 5 |
| 2025 | Real-World Video Dehazing based on Optical Flow Deformable Attention Fusion and Contrastive LearningabstractVideo dehazing aims to restore clear, haze-free frames from hazy videos while maintaining temporal continuity. However, existing deep learning methods for video dehazing do not take into account the gap between real and synthesized data. To address this issue, we propose a new method for achieving superior dehazing performance in real scenes. Specially, we construct a synthesizing hazy video datasets, which simulates the atmospheric light of realistic hazy videos. Additionally, we propose a optical flow deformable fusion module that utilizes optical flow to guide frame alignment and employs neighboring frame similarity for frame fusion. Furthermore, a contrastive learning loss function helps recovered frames closer to clear frames in feature space. Experimental results demonstrate the effectiveness of our proposed method, outperforming state-of-the-art methods in real hazy videos. Mengnan Zhang, Linghui Ma, Zhaoxi Liu, Zhenhong Jia |
ICME | 6 |
| 2025 | DedustNet:A Large-Scale Benchmark and Baseline for Sand Dust Video Enhancement
Linghui Ma, Mengzhen Xue, Gulisidan Abulizi, Zhenhong Jia |
ICONIP (2) | 6 |
| 2025 | Change Detection Method for Nighttime Large-Field Video Images Combining Anisotropic Diffusion and Non-Negative Matrix FactorizationabstractIn nighttime large-field environments, surveillance video images often suffer from poor illumination and complex random noise, which complicates the timely detection of small changes in targets. Traditional change detection algorithms exhibit poor performance in such scenarios due to their limited robustness. To address these challenges, this paper proposes a novel change detection method for nighttime large-field video images that integrates anisotropic diffusion and non-negative matrix factorization (NMF). The method begins by applying anisotropic diffusion to edge-difference images, effectively reducing noise while preserving critical structural information. Next, a fusion strategy based on non-negative matrix factorization is introduced to combine two difference images, enhancing quality and generating a more robust difference image. Finally, clustering techniques are employed to extract the final change targets. To validate the effectiveness of the proposed method, extensive experiments were conducted on a custom-built nighttime large-field video image dataset. The experimental results demonstrate that the proposed method surpasses existing algorithms in terms of accuracy, robustness, and noise resistance, particularly in low-light nighttime conditions, significantly enhancing the reliability of change detection. Zhenhong Jia, Xiaohui Huang 0005, Sensen Song, Jiajia Wang 0002, Ming Lv |
IJCNN | 2 |
| 2025 | STPYOLO: A Lightweight and Efficient Model for Small Target Pest DetectionabstractVision-based pest detection is an important task in smart agriculture. However, existing visual detection algorithms for small target pests often suffer from low detection accuracy, high complexity, and insufficient generalisation. In order to solve these existing algorithms’ deficiencies for small target pest detection. A pest detection model Small Target Pest YOLO (STPYOLO) based on the improved YOLOv10 model is proposed. To enhance the performance of small-target pest detection, this paper introduces a novel feature fusion method called MicroTargetScaleFusion(MTSFusion). This method finely adjusts feature weights, allowing the model to better focus on key areas where small-target pests are located. It also reduces the complexity of the model, making it more lightweight. To compensate for the lack of small-target pest datasets, a self-constructed small-target pest dataset is built, including six pests. Compared to YOLOv10, STPYOLO shows strong performance on two public datasets and the self-built dataset. The model size is reduced by 36.2%, and the number of parameters by 40%. Meanwhile on the self-built dataset mAP50 increased to 72.6%, an improvement of 5.8%. The excellent detection performance is still maintained while lightweighting. Finally, we integrated the STPYOLO model into the designed pest detection and early warning system to realize the real-time pest detection function for farm cotton fields. Jianyi Wang, Zhenhong Jia |
IJCNN | 2 |
| 2025 | Assisted Refinement Network Based on Channel Information Interaction for Camouflaged Object DetectionabstractCurrent camouflaged object detection (COD) methods predominantly focus on cross-layer feature fusion while neglecting cross-channel information interaction within the same layer. To address this limitation, we propose ARNet, a novel channel-aware refinement network featuring three key innovations. First, our Channel Information Interaction Module (CIIM) enables bidirectional horizontal-vertical integration of split-channel features, effectively capturing complementary cross-channel information through channel dimension fusion. Second, the Assisted Guidance Module (AGM) generates semantic-aware guidance maps from backbone features to dynamically regulate feature decoding and enhance spatial localization. Third, a Multi-scale Enhancement (MSE) module employs multi-scale convolution and residual-connected strategies to refine feature representations and expand contextual perception. Extensive experiments on COD10K, CAMO, and NC4K datasets demonstrate ARNet's superiority over 12 state-of-the-art methods. Code and results are available at: https://github.com/akuan1234/ARNet Kuan Wang 0007, Mengge Lu, Zhenhong Jia |
ICMR | 6 |
| 2025 | Efficient Camouflaged Object Detection Network Based on Channel Reconstruction and Hybrid AttentionabstractCamouflaged object detection is designed to segment objects that are highly integrated into the background and is highly challenging. At present, the main problem is how to effectively maintain the consistency of object semantics and avoid the loss of object information in the process of multi-scale feature fusion. To address this issue, we propose an efficient camouflaged object detection network (CHNet) based on channel reconstruction and hybrid attention. In our method, the feature channel is reconstructed using the strategy of Fusion-Separation-Transform-Fusion, and the adjacent features with strong semantic correlation are fused. Then, the global and local hybrid attention mechanism is used for layer-by-layer refinement, to maintain the consistency of object semantics to the largest extent. Experimental results on three large benchmark datasets show that our CHNet outperforms 11 state-of-the-art (SOTA) models, making it the most advanced COD model with 11.19M parameters. Our code is available at https://github.com/akuan1234/CHNet Kuan Wang 0007, Mengge Lu, Zhenhong Jia |
ICMR | 7 |
| 2025 | Infrared Small Target Detection with Feature Refinement and Context Enhancement
Xinyue Zhu, Zhenhong Jia |
MMM (2) | 6 |
| 2025 | A Fog Text Detection Network with Visual-Language Models and Fog Feature Decoupling
Zhaoxi Liu, Jiakun Tian, Zhenhong Jia |
PRCV (18) | 6 |
| 2025 | PATVTN: Period-Aware Time-Varying Topological Graph Neural Network for Traffic Flow ForecastingabstractTraffic flow forecasting has become a key technology to alleviate urban congestion. The core challenge is to accurately model the spatial-temporal coupling correlation in data. At present, many methods have made great progress, but most methods still face two important challenges: (i) The internal periodic pattern of traffic flow has not been accurately modeled. (ii) Failure to adapt effectively to the dynamic nature of traffic flow data. Accurate modeling of periodic patterns helps uncover the evolution of traffic flow, while improved adaptability to dynamic topologies enhances the capture of changing spatial dependencies. Therefore, this paper proposes a Periodic-Aware Time-Varying Topological Network (PATVTN) to enhance the precision of traffic flow prediction. In particular, PATVTN employs a Period-Aware Time Encoder to capture periodic traffic patterns, and introduces a Time-Varying Topological Space Encoder to adapt to the time-varying topological structure. Experiments conducted on four real-world traffic datasets show that our model consistently outperforms existing state-of-the-art methods. Xizhong Qin, Zhenhong Jia |
SMC | 4 |
| 2025 | RLCFE-Net: A reparameterization large convolutional kernel feature extraction network for weed detection in multiple scenarios
Zhenhong Jia, Baoquan Ge, Sensen Song, Congbing He, Jiajia Wang 0006, Xiaoyi Lv |
Expert Syst. Appl. | 2 |
| 2025 | CDTDNet: A neural network for capturing deep temporal dependencies in time series
Congbing He, Zhenhong Jia, Xiaohui Huang 0005 |
Inf. Sci. | 2 |
| 2025 | Scene Text Detection in Foggy Weather Utilizing Knowledge Distillation of Diffusion ModelsabstractAdverse weather conditions can significantly hinder the performance of deep learning-based object detection models. Traditional approaches often rely on image restoration techniques to enhance the quality of degraded images prior to detection. However, these methods frequently struggle to balance image enhancement and detection tasks effectively, often overlooking latent information that could be beneficial for detection. To address these challenges, we propose a novel framework: Knowledge Distillation based on Diffusion Models (KDDM). This framework incorporates a Dehaze Network (DN), which employs large kernel convolution to remove weather-specific artifacts, thereby revealing more latent information. The DN, together with a text detector, forms an end-to-end scene text detection network, acting as the student network. Additionally, the nuanced internal representations of text-to-image diffusion models adeptly capture and integrate higher-order visual semantic concepts. Given the rich textual and visual content inherent in scene text, there is a fundamental connection to text-to-image diffusion models. As such, we utilize diffusion models as a teacher network to distill high-level visual semantic knowledge into the student network. Notably, we introduce an innovative distillation technique using a “Threshold_Mask”, which ensures that the student network focuses on text regions while minimizing interference from irrelevant background elements. Comprehensive experimental evaluations demonstrate that our KDDM framework significantly outperforms baseline models under foggy weather conditions, marking a substantial advancement in the field. Zhaoxi Liu, Zhenhong Jia |
IEEE Signal Process. Lett. | 3 |
| 2025 | Efficient Mamba-Attention Network for Remote Sensing Image Super-ResolutionabstractLightweight remote sensing image super-resolution (RSISR) methods aim to reconstruct remote sensing images (RSIs) while reducing computational complexity. Previous lightweight model development has primarily focused on the design of convolutional neural networks (CNNs). While CNNs excel at capturing local features, they are limited in establishing long-range dependencies. Mamba, as a model for long-range modeling, has linear computational complexity, making it a viable option for lightweight models. Based on these considerations, this paper proposes an efficient mamba-attention network (EMAN) that can efficiently capture the intricate details and broader semantic information in RSIs. Specifically, we designed a multi-scale detail extraction unit (MDEU) and a multi-dimensional mamba-attention (MDMA). In MDEU, we introduced a multi-scale mechanism and local variance to focus on structural information in RSIs. In MDMA, we integrated spatial expansion and an atrous-based selective scan mechanism to design an efficient scanning method. This method ensures the lightweight nature of the model while establishing global correlations. Additionally, MDMA establishes inter-channel correlations to enhance information exchange. We conducted a comprehensive evaluation of the proposed method on two remote sensing datasets and five benchmark super-resolution (SR) datasets. Extensive experiments demonstrate that our method can achieve superior performance while maintaining a model complexity similar to other lightweight models. Tianren Wu, Rundong Zhao, Ming Lv, Zhenhong Jia, Liangliang Li 0001, Minqin Liu, Xiaobin Zhao, Hongbing Ma, Gemine Vivone |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Adaptive Gaussian Regularization Constrained Sparse Subspace Clustering for Image SegmentationabstractSparse Subspace Clustering (SSC) is integral to image processing, drawing from spectral clustering foundations. However, prevalent methods, relying on an l1-norm constraint, fail to capture nuanced inter-region correlations, affecting segmentation efficacy. To remedy this, we introduce an Adaptive Gaussian Regularization Constrained SSC for enhanced image segmentation. This method begins with superpixel preprocessing to enrich local information. Given the Gaussian nature of the SSC’s sparse coefficient matrix, a Gaussian probability density function is infused as a regularization term, reinforcing regional image ties and facilitating similarity matrix creation. Using spectral clustering, we then define superpixel clusters leading to the final segmentation. When tested against the BSDS500 and SBD datasets and other leading algorithms, our model showcases marked improvements in natural image segmentation. Sensen Song, Dayong Ren, Zhenhong Jia |
ICASSP | 3 |
| 2024 | Beyond the Snowfall: Enhancing Snowy Day Object Detection Through Progressive Restoration and Multi-Feature FusionabstractIn the field of computer vision, object detection is a prominent and challenging task. Despite the favorable performance of deep learning-based object detection techniques on clear images, it fails in inclement weather conditions like snow because of image degradation. Recent efforts have explored using image restoration methods to enhance degraded images before object detection. However, direct restoration can sometimes cause new disturbances, impeding detection performance improvements. To address this issue, we propose a joint framework that connects the iterative desnow module and detection module in an end-to-end manner. Specially, we design an Advantage Union structure for multi-feature fusion, which effectively combines original, intermediate, and restored features, reducing potential information loss from restoration. Experimental results show that our method achieves higher accuracy compared to the recent state-of-the-art methods in both synthetic dataset and real-to-world snowy images. Tianhao Xue, Zhenhong Jia |
ICASSP | 5 |
| 2024 | RVDNet: A Two-Stage Network for Real-World Video Desnowing with Domain AdaptationabstractVideo snow removal is an important task in computer vision, as the snowflakes in videos reduce visibility and negatively affect the performance of outdoor visual systems. However, due to the complexity of real snowy scenarios, it is difficult to apply existing supervised learning-based methods to process real-world snowy videos. In this paper, we propose a novel two-stage video desnow network for the real world, called RVDNet. The first stage of RVDNet utilizes Spatial Feature Extraction Modules (SFEM) to extract the spatial features of the input frames. In the second stage, we design Spatial-Temporal Desnowing Modules (STDM) to remove snowflakes via spatio-temporal learning. Furthermore, we introduce the unsupervised domain adaptation module, which is embedded for aligning the feature space of real and synthetic data in the spatial and spatio-temporal domains, respectively. Experiments on the proposed SnowScape dataset prove that our method has superior desnow performance not only on synthetic data, but also in the real world. Tianhao Xue, Runlin He, Zhenhong Jia |
ICASSP | 6 |
| 2024 | Joint Image Restoration For Domain Adaptive Object Detection In Foggy Weather ConditionabstractDriven by deep learning, object detection methods have made significant progress in recent years. However, there is still a domain shift between synthetic foggy data and real foggy data, this leads to a undesirable decrease in detection results when applying the algorithm model trained on synthetic foggy datasets. In this article, we design a domain-adaptive YOLOX object detection algorithm by joint image restoration, in order to improve object detection performance in foggy scenes. Specially, we design an end-to-end domain adaptive framework that combines dehazing module and YOLOX together. To achieve feature alignment, we introduced a domain classifier in the feature spaceand discuss its optimal placement in the framework. Experimental results show that the proposed domain adaptation method achieved 62.60 percent mAP on real foggy image datasets RTTs, outperforming other state-of-the-art methods. Zhenhong Jia |
ICIP | 4 |
| 2024 | U-Convnext Network for Infrared Small Target DetectionabstractDue to the inherent weakness and difficulty in extracting features of infrared small targets, there is a risk of information loss in the deep layers of the network.We propose a new network model called U-Convnext. Specifically, we design a novel multi-scale Convnext module (Mcnt) based on the Convnext network, aiding in better feature extraction of infrared small targets. To mitigate deep-layer information loss, we introduce parallel dilated convolution module (Pdconv) and serial dilated convolution module (Sdconv). Pdconv captures surrounding information from multiple scales during downsampling, while Sdconv enables finer processing in the deep layers of the network. Experimental results demonstrate the superiority of the U-Convnext network over other methods. Yuye Zhang, Dangxuan Wu, Zhenhong Jia |
ICIP | 6 |
| 2024 | Intermediate Domain Meets Natural Hazy TrackingabstractBenefit from the development of deep learning, recent advances in object tracking have achieved compelling results on the normal sequences. However, in adverse weather, such as haze, it is difficult to extract effective features due to the videos being severely degraded, making existing object tracking methods extremely ineffective. To address this issue, this work introduces an unsupervised domain adaptation framework for natural hazy weather tracking (NHT). Specifically, we employ a Fastvit domain discriminator with a Gradient Reverse Layer (GRL) in the feature extractor to align image features from normal to natural hazy weather. To tackle the significant domain shift between normal and natural hazy weather, we incorporate a synthesized haze dataset as an intermediate domain, achieving progressive domain adaptation. Moreover, we establish a hazy weather dataset namely Haze2023 for training and testing. Comprehensive experimental results demonstrate the generalization ability of our proposed NHT in natural hazy weather. Yuwei Feng, Zhenhong Jia |
ICME | 6 |
| 2024 | Deformable Multi-Scale Network for Snow Removal in Video
Runlin He, Tianhao Xue, Zhaoxi Liu, Zhenhong Jia |
ICPR (32) | 5 |
| 2024 | TBIA-DBNet: A Two-Branch Image-Adaptive DBNet for Scene Text Detection in Real-World Foggy Scenes
Zhaoxi Liu, Runlin He, Mengnan Zhang, Zhenhong Jia |
ICPR (31) | 5 |
| 2024 | Dual-MambaNet: A Lightweight Dual-Branch Brain Image Segmentation Network Based on Local Attention and Mamba
Dayong Ren, Zhenhong Jia, Jianyi Wang |
ICPR (28) | 4 |
| 2024 | Detecting Video Image Changes Based on Improved Difference Map and Image FusionabstractTo resolve the problems of complex random noise and poor image quality in hawk-eye surveillance video images under low illumination, a surveillance video image change detection method based on K-means clustering and differential image adaptive fusion was developed. First, the log ratio operator and extreme average ratio operator were used to obtain two difference maps for two video surveillance images at different times. Then, the Laplacian pyramid was used for adaptive fusion. Next, the TVL1 dual method was used to denoise the fused image, PCA was used to fuse the denoised image and the enhanced image of the changed region, and then the adaptive median filter was used for filtering. Finally, the two-stage clustering method was used to obtain the final change detection results for the difference map. Based on our experimental findings, our method outperforms comparative algorithms in many metrics, including accuracy, robustness, time consumption, and false alarm rate. Ang Tian, Zhenhong Jia, Xiaohui Huang 0005, Sensen Song, Jiajia Wang 0002 |
IJCNN | 2 |
| 2024 | BD-YOLO: Optimising Insect Imbalance Data Detection ModelsabstractVision-based insect detection is an important task in smart agriculture. Field-collected insect image datasets usually show severe imbalances in categories. Previous models have insufficient accuracy as well as low real-time performance for detecting unbalanced data of insects. To address the issues described above, an improved framework Balance Detect YOLO(BD-YOLO) based on YOLOv8 is proposed in this paper. We innovatively designed Focaler-CPIoU loss to increase the weight of a few classes of samples to achieve the optimisation of pest imbalance data detection. We introduced lightweight convolutional MBConv to improve the recognition accuracy while reducing the number of model parameters. In order to verify the effectiveness of the method in this paper, we established an insect dataset containing six types of insects, in which the most numerous thrips is 431.5 times more than the least numerous chrysopaperla. The experimental results show that compared with YOLOv8, the mAP50 reaches 61.7% with a 73.1% reduction in the number of model parameters, an improvement of 5.4%, and an improvement of 769 in FPS. A feasible approach to optimise the detection of unbalanced number of insects is provided in this paper along with the ability to be deployed on lightweight hardware devices. Jian-Yi Wang, Zhenhong Jia |
SMC | 2 |
| 2024 | Joint Optimization of Maximum Achievable Rate in SWIPT Systems Assisted by Active STAR-RIS
Junlong Yang, Xizhong Qin, Zhenhong Jia, Lamu Mao |
WASA (2) | 3 |
| 2024 | A lightweight weed detection model with global contextual joint features
Zhenhong Jia, Baoquan Ge |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Adaptive fuzzy weighted C-mean image segmentation algorithm combining a new distance metric and prior entropy
Sensen Song, Zhenhong Jia, Dongdong Ni |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Lightweight Remote Sensing Image Super-Resolution via Background-Based Multiscale Feature Enhancement NetworkabstractIn the field of remote sensing image super-resolution (RSISR), most methods based on convolutional neural networks (CNNs) tend to focus on high-weight features in the convolutional kernels, thus overlooking low-weight background features. This bias may result in the neglect of some important information in the background. To address this challenge, we propose a background-based multiscale feature enhancement network (BMFENet), which can extract and supplement missing features from different scale backgrounds to improve the reconstruction of remote sensing images (RSIs). Specifically, we constructed a large kernel feature supplement block (LFSB). The LFSB uses large kernel attention mechanism and multiscale mechanism to expand the receptive field, aggregating global information. Meanwhile, it generates background feature weights to increase the attention to neglected information, thereby reducing the distortion of detailed features. Furthermore, to enhance the nonlinear expression capability of the model, we designed a lattice gated unit (LGU). The LGU removes redundant information through a gating mechanism, efficiently aggregates useful channel information through interchannel interactions and attention mechanisms, and introduces directional convolution to make the model more adaptable to super-resolution (SR) tasks in complex scenes. We validated our method on two remote sensing and four SR benchmark datasets, and the results show that our approach achieves a good balance between performance and complexity. Tianren Wu, Rundong Zhao, Ming Lv, Zhenhong Jia, Liangliang Li 0001, Zheyuan Wang, Hongbing Ma |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Saliency optimization fused background feature with frequency domain features
Sensen Song, Zhenhong Jia, Jie Yang 0002, Nikola K. Kasabov |
Multim. Tools Appl. | 2 |
| 2024 | Constructing New Backbone Networks via Space-Frequency Interactive Convolution for Deepfake DetectionabstractThe serious concerns over the negative impacts of Deepfakes have attracted wide attentions in the community of multimedia forensics. The existing detection works achieve deepfake detection by improving the traditional backbone networks to capture subtle manipulation traces. However, there is no attempt to construct new backbone networks with different structures for Deepfake detection by improving the internal feature representation of convolution. In this work, we propose a novel Space-Frequency Interactive Convolution (SFIConv) to efficiently model the manipulation clues left by Deepfake. To obtain high-frequency features from tampering traces, a Multichannel Constrained Separable Convolution (MCSConv) is designed as the component of the proposed SFIConv, which learns space-frequency features via three stages, namely generation, interaction and fusion. In addition, SFIConv can replace the vanilla convolution in any backbone networks without changing the network structure. Extensive experimental results show that seamlessly equipping SFIConv into the backbone network greatly improves the accuracy for Deepfake detection. In addition, the space-frequency interaction mechanism does benefit to capturing common artifact features, thus achieving better results in cross-dataset evaluation. Our code will be available athttps://github.com/EricGzq/SFIConv. Zhiqing Guo, Zhenhong Jia, Dewang Wang, Gaobo Yang, Nikola K. Kasabov |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Online Low-Light Sand-Dust Video Enhancement Using Adaptive Dynamic Brightness Correction and a Rolling Guidance FilterabstractSand-dust videos obtained in a low-light environment are characterized by low contrast, nonuniform illumination, color cast, and considerable noise. To realize sand-dust removal and brightness enhancement simultaneously, this article proposes an online low-light sand-dust video enhancement method using adaptive dynamic brightness correction and a rolling guidance filter. The proposed dual-threshold interframe detection strategy involves two methods to treat low-light sand-dust video frames. The first method involves two components: an adaptive dynamic brightness correction algorithm to correct the color deviation of the low-light video frame and improve its brightness and a rolling guidance filter combined with guided image filtering to enhance the frame details. The second method enhances the quality of the incoming frame by reducing the amount of calculation. The first frame of the video is processed using the first method. The processing method of each subsequent frame is determined according to its interframe detection value with the buffer frame. Through qualitative and quantitative comprehensive experiments on low-light sand-dust images and videos, the performance of the proposed method is compared with those of state-of-the-art methods. The proposed method for frame quality improvement achieves the best visual effect in enhancing the quality of low-light sand-dust images, as indicated by the best objective evaluation indicators. Moreover, compared with the framewise enhancement method, the video processing efficiency associated with the dual-threshold interframe detection strategy is 2.77 times higher. Dongdong Ni, Zhenhong Jia, Jie Yang 0002, Nikola K. Kasabov |
IEEE Trans. Multim. | 2 |
| 2023 | Text Enhancement: Scene Text Recognition in Hazy Weather
En Deng, Jiakun Tian, Yangxin Liu, Zhenhong Jia |
ICDAR (6) | 5 |
| 2023 | FTDNet: Joint Semantic Learning for Scene Text Detection in Adverse Weather Conditions
Jiakun Tian, Yangxin Liu, En Deng, Zhenhong Jia |
ICDAR (5) | 5 |
| 2023 | Unsupervised Domain Adaptive Learning for Image Desnowing with Real-World DataabstractSnow images usually contain snow grains, snow streaks, and mist, which greatly affect the visibility of images. Currently, supervised learning with synthetic data often faces limitations when it comes to handling real-world snow images. To address this crucial issue, this work proposes an unsupervised domain adaptation image snow removal framework. The framework improves the performance on real-world images by learning a domain classifier in adversarial training manner. Additionally, considering the diversity of snowflake shapes and sizes in real-world snow images, we design a multiple-kernel dilated convolution module. Extensive experiments on three representative datasets have validated that our model can achieve better results than existing desnowing methods. More importantly, experiments on real datasets show that the proposed method obtains state-of-the-art performance in real-world desnowing. Jingxu Ren, Yusen Zhu, Yangxin Liu, Zhenhong Jia |
ICIP | 6 |
| 2023 | Multi-Task Model Based on Vision Task Level for Saliency Object Detection in Foggy ConditionsabstractIn recent years, saliency object detection methods based on convolutional neural networks have been widely studied, and have achieved excellent performance in clear images. However, due to the low visibility of images in foggy conditions, the existing saliency object detection methods will be seriously affected or even ineffective. To address this problem, we introduce an end-to-end multi-task learning network. We design two subetworks for depth estimation and image restoration as auxiliary tasks to improve saliency object detection in foggy conditions. According to different characteristics of vision tasks, different shared layers are assigned to improve the performance of saliency object detection. Experiments show that our method has been greatly improved on both synthetic foggy datasets and real-to-world foggy datasets, outperforming many state-to-the-art saliency object detection methods. Yusen Zhu, Jingxu Ren, Jiakun Tian, Zhenhong Jia |
ICIP | 5 |
| 2023 | Video anomaly detection based on cross-frame prediction mechanism and spatio-temporal memory-enhanced pseudo-3D encoder
Xiaopeng Wen, Huicheng Lai, Guxue Gao, Yang Xiao 0018, Tongguan Wang, Zhenhong Jia |
Eng. Appl. Artif. Intell. | 6 |
| 2023 | Two-Step Unsupervised Approach for Sand-Dust Image EnhancementabstractIn sand‐dust environments, light is scattered and absorbed, and sand‐dust images thus suffer from severe image degradation problems, such as color shifts, low contrast, and blurred details. To address these problems, we propose a two‐step unsupervised sand‐dust image enhancement algorithm. In the first step, a convenient and competent color correction method is put forward to solve the color shift problem. Considering the wavelength attenuation features of sand‐dust images, a linear stretching and blue channel compensation method is designed, and an adaptive color shift correction factor is developed to remove the color shift. In the second step, to enhance the clarity and details of the images, an unsupervised generative adversarial network is proposed, which does not require pairs of data for training. To reduce detail loss, the detail enhancement branch is designed, and the generator considers to more details through the constructed coarse‐grained and fine‐grained discriminators. The introduced multiscale perceptual loss promotes the image fidelity well. Experiments show that the proposed method achieves better color correction, enhances image details and clarity, has a better subjective effect, and outperforms existing sand‐dust image enhancement methods both quantitatively and qualitatively. Similarly, our method promotes the application capability of the target detection algorithm and also has a good enhancement effect on underwater images and haze images. Guxue Gao, Huicheng Lai, Zhenhong Jia |
Int. J. Intell. Syst. | 3 |
| 2023 | A fast sand-dust video quality improvement method using simple color balance and dynamic guided filtering
Dongdong Ni, Zhenhong Jia, Jie Yang 0002, Nikola K. Kasabov |
Multim. Tools Appl. | 2 |
| 2023 | Sand-dust image enhancement based on light attenuation and transmission compensation
Zhenhong Jia, Huicheng Lai, Nikola K. Kasabov, Sensen Song |
Multim. Tools Appl. | 2 |
| 2023 | A real time target face tracking algorithm based on saliency detection and Camshift
Zhenhong Jia, Huicheng Lai |
Multim. Tools Appl. | 2 |
| 2023 | A fast and effective algorithm for specular reflection image enhancement
Ye Xin, Yifei Wei, Zhuang Huang, Zhenhong Jia, Jie Yang 0002, Nikola K. Kasabov |
Multim. Tools Appl. | 4 |
| 2023 | A dual channel decomposition and remapping fusion model for low illumination images with a wide field of view
Wei Zhang 0362, Zhenhong Jia, Jie Yang 0002, Nikola K. Kasabov |
Signal Process. Image Commun. | 2 |
| 2023 | Unsupervised Change Detection in Wide-Field Video Images Under Low IlluminationabstractIn low-illumination environments such as at night, due to factors such as the large monitoring field of an eagle eye, short sensor exposure time, and high-density random noise, the video images collected by image sensors generally have poor visual quality and low signal-to-noise ratio, which makes it difficult for surveillance systems to detect weak changes. To solve this problem, we propose a method for image change detection (CD) in surveillance video based on optimized k-medoids clustering and adaptive fusion of difference images (DIs). First, for the input multitemporal video surveillance images, two DIs are obtained by log-ratio and extremum pixel ratio operators. Then, the two DIs are adaptively fused by combining the local energy of DIs and the Laplacian pyramid. Simultaneously, the fused DI is compressed by the normalization function, and the final DI is obtained via the improved adaptive median filter. Finally, the changed image is obtained by using the optimized k-medoids clustering algorithm. The experimental results show that the proposed method can accurately and effectively detect weak changes in the eagle eye surveillance picture in a low-illumination environment. Compared with those of other methods, the accuracy and robustness of the proposed method are higher, and the running time of the algorithm is shorter. Moreover, it will not generate a false alarm due to the influence of noise in unchanged scenes. Baoqiang Shi, Zhenhong Jia, Jie Yang 0002, Nikola K. Kasabov |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Image Segmentation Based on Fuzzy Low-Rank Structural ClusteringabstractFuzzy clustering is an essential algorithm in image segmentation, and most of them are based on fuzzy c-mean algorithms. However, it is sensitive to noise, center point selection, cluster number, and distance metric. To address this problem, we propose a new fuzzy clustering method based on low-rank representation (LRR) for image segmentation, which integrates low-rank structure with fuzzy theory. First, we improve the morphological reconstruction superpixel method based on edge detection by introducing anisotropy to enhance the image edge. Thus, on the one hand, the improved morphological reconstruction superpixel method can improve its noise-resistance performance; on the other hand, the complexity of the subsequent low-rank computation can be reduced by enhancing the superpixels constructed by the edges. Second, inspired by the fact that rank can represent correlation, we propose the concept of fuzzy low-rank structure, which is not dealing with data directly but with the relationship between data. Specifically, we perform rank minimization on the constructed membership matrix to obtain the optimal matrix. To obtain better clustering results, we added the Frobenius norm of the fuzzy matrix as a fuzzy regularization term in the LRR model to achieve global convergence and obtain a membership matrix with a strong element correlation. Finally, we obtain the final clustering results by clustering the processed membership matrix using a subspace clustering with a low-rank structure constraint. Experiments performed on artificial and real-world images show that the proposed method is more effective and efficient than the current state-of-the-art methods. Sensen Song, Zhenhong Jia, Jie Yang 0002, Nikola K. Kasabov |
IEEE Trans. Fuzzy Syst. | 2 |
| 2022 | Salient detection via the fusion of background-based and multiscale frequency-domain features
Sensen Song, Zhenhong Jia, Jie Yang 0002, Nikola K. Kasabov |
Inf. Sci. | 2 |
| 2022 | Lightweight spatial-channel adaptive coordination of multilevel refinement enhancement network for image reconstruction
Yuxi Cai, Huicheng Lai, Zhenhong Jia |
Knowl. Based Syst. | 3 |
| 2022 | Multispectral Image Enhancement Based on Weighted Principal Component Analysis and Improved Fractional Differential MaskabstractCompared with single-band images, multispectral images with multiple wavelengths contain more spectral information and can fully reflect the detailed features of ground objects in different bands. To synthesize the feature information, a method of multispectral image enhancement based on improved weighted principal component analysis (WPCA) and improved fractional differential (IFD) filtering is proposed. First, the optimum index factor (OIF) model is modified to select the bands for easy follow-up processing. Then an improved WPCA transform that uses the average gradient and texture roughness is applied to compensate for each band. This approach preserves the main information, compresses the amount of image data, and obtains uncorrelated principal components. According to the correlation of neighboring pixels, a new mask is introduced to enhance the first principal component. Finally, the brightness values of the image are adjusted after inverse WPCA transform to obtain the final enhanced image. The experimental results demonstrate the superiority of the proposed method over related methods. Its future practical applications are discussed in the conclusion. Zhenhong Jia, Jie Yang 0002, Nikola K. Kasabov |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Color balance and sand-dust image enhancement in lab space
Guxue Gao, Huicheng Lai, Zhenhong Jia |
Multim. Tools Appl. | 4 |
| 2022 | Target re-location kernel correlation filtered visual tracking with fused deep feature
Qingzhong Shu, Huicheng Lai, Zhenhong Jia |
Multim. Tools Appl. | 3 |
| 2021 | A novel multiscale transform decomposition based multi-focus image fusion framework
Liangliang Li 0001, Hongbing Ma, Zhenhong Jia, Yujuan Si |
Multim. Tools Appl. | 3 |
| 2021 | Energy efficiency resource allocation for D2D communication network based on relay selection
Xizhong Qin, Zhenhong Jia |
Wirel. Networks | 3 |
| 2020 | A novel approach for multi-focus image fusion based on SF-PAPCNN and ISML in NSST domain
Liangliang Li 0001, Yujuan Si, Linli Wang, Zhenhong Jia, Hongbing Ma |
Multim. Tools Appl. | 4 |
| 2020 | Remote sensing image enhancement based on the combination of adaptive nonlinear gain and the PLIP model in the NSST domain
Lanhua Zhang, Zhenhong Jia, Lucien Koefoed, Jie Yang 0002, Nikola K. Kasabov |
Multim. Tools Appl. | 2 |
| 2020 | Moving object detection in video sequence images based on an improved visual background extraction algorithm
Junhui Zuo, Zhenhong Jia, Jie Yang 0002, Nikola K. Kasabov |
Multim. Tools Appl. | 2 |
| 2015 | Combining CNN and MIL to Assist Hotspot Segmentation in Bone Scintigraphy
Shijie Geng, Shaoyong Jia, Yu Qiao 0003, Jie Yang 0002, Zhenhong Jia |
ICONIP (4) | 5 |
| 2015 | Scene text detection method based on the hierarchical modelabstractAs an important step in text‐based information extraction systems, scene text detection has become a popular subject of research in recent years. In this study, the authors present a novel approach to robustly detect texts which are variable in scales, colours, fonts, languages and orientations in scene images. To segment candidate text connected components (CCs) from images, both local contrast and colour consistency are considered in superpixel level. To filter out the non‐text CCs, a hierarchical model is designed. This hierarchical model groups the CCs into three cascaded stages, and is equipped with a well‐designed classifier in each stage. Experimental results on the public ICDAR 2005 dataset and the MSRA‐TD500 dataset show that their approach obtains better performance than other state‐of‐the‐art methods. Yuehu Liu, Zhenhong Jia |
IET Comput. Vis. | 4 |
| 2014 | Automatic Image Annotation Exploiting Textual and Visual Saliency
Yun Gu, Haoyang Xue, Jie Yang 0002, Zhenhong Jia |
ICONIP (3) | 4 |
| 2014 | Shape Preserving RGB-D Depth Map Restoration
Wei Liu 0044, Haoyang Xue, Yun Gu, Jie Yang 0002, Qiang Wu 0001, Zhenhong Jia |
ICONIP (3) | 6 |
| 2014 | Unsupervised Segmentation Using Cluster Ensembles
Wei Zhang 0362, Jie Yang 0002, Wenjing Jia, Nikola K. Kasabov, Zhenhong Jia, Lei Zhou 0003 |
ICONIP (3) | 5 |
| 2014 | The remote sensing image enhancement based on nonsubsampled contourlet transform and unsharp maskingabstractSUMMARY To restrain pseudo‐Gibbs phenomenon, low contrast and blurred phenomenon in the process of image enhancement, a new method based on the nonsubsampled contourlet transform and the unsharp masking is proposed in this paper. The proposed method utilizes the shift‐invariance of nonsubsampled contourlet transform to restrain the pseudo‐Gibbs phenomenon, and then enhance details of the image by unsharp masking. We achieved an increase in image definition by 54.5%, the mean increased by 15.6%, whereas the standard deviation increased by 54.5% compared with the unsharp masking method. To the noisy image, we achieved an increase in image definition by 35.4%, the mean increased by 2.2%, whereas the standard deviation increased by 34.9% compared with the unsharp masking method. Copyright © 2013 John Wiley & Sons, Ltd. Xiaoting Pu, Zhenhong Jia, Jie Yang 0002 |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | An enhanced multiphase Chan-Vese model for the remote sensing image segmentationabstractSUMMARY The level set method has been widely used in image segmentation; however, the complexity of the computation has restricted its application field. Also, it is a big challenge to segment remote sensing image mainly because of the complex terrain. In this paper, an enhanced multiphase phase level set method based on the Chan–Vese (C‐V) model is proposed for segmenting remote sensing images. Compared with the C‐V model, two main contributions of the proposed model mainly include the following: First, we introduce a new strategy of initialization in which the contours of the first k biggest connected regions are extracted as the initial curves (k is the number of level set functions); Second, to increase the accuracy, a morphological gradient component is added to the original intensity image. To investigate the effectiveness and efficiency of the proposed model, we have applied it to analyze different kinds of images, including synthetic, real, and remote sensing images. The experimental results have shown that our method is able to achieve better segmentation with less computational consumption compared with the traditional multiphase C‐V model and local and global intensity fitting model. Copyright © 2013 John Wiley & Sons, Ltd. Zhenhong Jia, Jie Yang 0002, Nikola K. Kasabov |
Concurr. Comput. Pract. Exp. | 3 |
| 2013 | Face detection algorithm based on hybrid Monte Carlo method and Bayesian support vector machineabstractSUMMARY With distinct advantages in resolving the problems of small sample, nonlinear, high dimension learning, the support vector machine (SVM) has been widely applied in face detection and face recognition. In fact, a large number of facial images were needed to train the SVM algorithms. With the rising of training image numbers, the training complexity of SVM was increased by way of geometric series. In this paper, the hybrid Monte Carlo method of the Bayesian support vector machine is proposed. This method solves the problems of high‐dimension and long training time effectively. Experimental results show that the method greatly reduces the training time of face detection algorithm and obtains more accurate face detection effect. Copyright © 2012 John Wiley & Sons, Ltd. Taiyi Zhang, Zhenhong Jia |
Concurr. Comput. Pract. Exp. | 3 |
| 2012 | Multi-scale image segmentation algorithm based on support vector machine approximation criteriaabstractSUMMARY A new multi‐scale image segmentation algorithm based on support vector machine (SVM) approximation criteria has been discussed in this paper. Most current multi‐scale image segmentation algorithms are based on the restricted empirical risk minimization, and the approximation of multi‐scale image segmentation was poor. As the SVM theory was one based on the structural risk minimization, the best approximation results could be reached. So, it was combined with multi‐scale image segmentation algorithms, and one‐image multi‐resolution analysis approximation algorithms based on the SVM theory were presented in this paper, which could obtain more accurate multi‐scale image segmentation. By numerical results, the algorithm was further verified that more accurate image segmentation results were unfolded. Copyright © 2011 John Wiley & Sons, Ltd. Zhenhong Jia |
Concurr. Comput. Pract. Exp. | 2 |
| 2007 | A Novel Chaotic Neural Network for Function Optimization
Zhenhong Jia |
ICONIP (2) | 2 |