Zhongjie Zhu

dblp:53/9171 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 1 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DIR-RSCLIP: A Degradation-Invariant Cross-Modal Retrieval Network for Remote Sensing Images
abstract
Remote sensing image-text retrieval locates target scenes in large-scale repositories through natural language descriptions, serving as an indispensable technique for disaster monitoring and related applications. Existing methods have achieved notable progress under idealized clear-image settings, yet image degradation is pervasive in real-world remote sensing scenarios. Degradations such as low illumination, haze, sensor noise, motion blur, and cloud occlusion severely disrupt semantic information, causing sharp performance drops for mainstream retrieval models. To address this challenge, we propose DIR-RSCLIP, a cross-modal retrieval network designed for degraded conditions. DIR-RSCLIP introduces a degradation-aware dual-branch architecture: the semantic branch preserves high-level scene semantics while the degradation branch explicitly models degradation attributes; the two branches are adaptively fused via a type-conditioned gating mechanism. Additionally, we adopt a multi-positive contrastive learning framework that treats both clean and degraded versions of the same scene as positives, thereby establishing degradation-invariant cross-modal alignment at the loss-function level. Experiments on RSICD and RSITMD validate the effectiveness of the proposed method.
Changbai Chen, Zhongjie Zhu
ICMR4
2026 BigCounter: A Bidirectional-Guided Network With Scene-Semantics-Driven Fusion for RGB-Thermal Crowd Counting
abstract
Accurate crowd counting has become increasingly essential for public safety management and Internet of Video Things (IoVT) applications, driven by rapid population growth and urbanization. However, RGB-Thermal (RGB-T) crowd counting remains challenging due to poor recognition of small targets and degraded performance under extreme conditions such as low-light environments. To address these issues, we propose a bidirectional-guided network with scene-semantics driven fusion for RGB-T crowd counting (BigCounter) that enhances robustness and generalization in complex scenes. BigCounter comprises of three parallel branches: a primary branch, a dynamic illumination auxiliary enhancement branch (DIAEB), and a high-resolution auxiliary enhancement branch (HAEB), which respectively improve robustness under illumination variations and accuracy for small target detection. Moreover, a cross-layer scene-driven fusion module (CLSFM) and a cross-modal semantic-driven fusion module (CMSFM) are designed to strengthen structural consistency and explore semantic complementarity between modalities. Through multi-branch collaboration and semantic-aware fusion, BigCounter significantly enhances feature representation. Extensive experiments on two benchmark RGB-T datasets demonstrate that BigCounter achieves superior accuracy and generalization compared with state-of-the-art methods.
Xiaomin Fan, Feng Shao 0001, Baoyang Mu, Xiongli Chai, Zhongjie Zhu, Zhiyi Mo
IEEE Internet Things J.5
2026 PU-TransMamba: A Hybrid Point Cloud Upsampling Framework With Detail-Aware Transformer and Spatially Coherent Mamba
abstract
High-quality 3D point clouds are essential for high-fidelity perception in Internet of Things (IoT)-enabled intelligent systems. While Point Cloud Upsampling (PCU) is widely used to mitigate data sparsity, existing methods often struggle to balance the preservation of fine-grained local details with the maintenance of global topological consistency. Transformer-based approaches frequently suffer from excessive computational overhead and high-frequency detail loss, whereas emerging state space models like Mamba, despite their efficiency, inevitably sacrifice spatial coherence due to the 1D serialization of irregular 3D points. To address these critical bottlenecks, we introduce PU-TransMamba, a hybrid framework that synergistically leverages a Detail-Aware Transformer and a Spatially-Coherent Mamba. Each component is designed to resolve specific PCU limitations: a Complexity-Aware Bilateral Decoder is developed to adaptively recover sharp geometric edges by processing features across dual domains, while a Sequence-Aligned Mamba Encoder utilizes multiple spatial curvature descriptors to compensate for the spatial information loss inherent in serialization. Additionally, a Global Geometry Injector and a Local Neighbor Injector are designed to ensure structural integrity by infusing holistic skeletal priors and neighborhood context, respectively. To minimize feature discrepancies between the hybrid branches, we also propose a Self-Distillation Loss. Extensive experiments on five benchmark datasets demonstrate that PU-TransMamba outperforms state-of-the-art methods in both reconstruction accuracy and computational scalability. The results confirm its ability to recover intricate geometries, indicating significant potential for IoT-driven 3D perception and communication systems.
Feng Shao 0001, Xiongli Chai, Hangwei Chen, Zhongjie Zhu, Zhiyi Mo
IEEE Internet Things J.5
2026 SGNet: A Structure-Guided Lightweight Network for VDT Salient Object Detection
abstract
Visual-Depth-Thermal (VDT) salient object detection (SOD) aims to jointly exploit RGB, depth and thermal cues to segment the most visually significant regions. However, most existing VDT SOD models are heavy in parameters and computational cost, limiting their deployment on real-world and edge devices. To tackle this, we propose SGNet, a structure-guided lightweight network for efficient VDT SOD. Specifically, we design a lightweight Tri-modal Fusion Module (TFM) to integrate three modalities at the semantic level, and a Shared Structure Extraction Module (SSEM) to extract common structural information from depth and thermal modalities. A Structure Refine Module (SRM) further injects the extracted structure into the deepest semantic features, while a Multiscale Feature Refinement Module (MFRM) progressively decodes multi-level features under deep supervision to produce saliency maps with clear boundaries. Benefiting from these modules, SGNet achieves competitive performance on the VDT2048 benchmark with only 5.51 M parameters and a real-time speed of 120 FPS at 320 × 320 resolution, surpassing state-of-the-art methods while remaining deployment-friendly.
Huizhi Wang, Feng Shao 0001, Xuebin Wei, Xiongli Chai, Hangwei Chen, Zhongjie Zhu
IEEE Internet Things J.6
2025 Rethinking decomposition for time series forecasting: Learning from temporal patterns
Di Ge, Bo Liu 0103, Yuhang Cheng, Zhongjie Zhu
Expert Syst. Appl.4
2025 MISF-Net: Modality-Invariant and -Specific Fusion Network for RGB-T Crowd Counting
abstract
To accurately perform crowd counting, utilizing the complementary relationship between RGB and thermal images to analyze the crowd has become the focus of current research. Due to different imaging principles, multi-modal images often contain different contents, which are their modality-specific information. For example, RGB images contain more texture and color details, while thermal images contain thermal radiation information. Meanwhile, they also describe the same target content, e.g., crowds, which are modality-invariant. However, existing methods only design different modules to directly fuse RGB and thermal image features, which did not fully consider the above facts. In this paper, by analyzing the similarities and differences between multi-modal images, we propose a Modality-Invariant and -Specific Fusion Network (MISF-Net) for RGB-T Crowd Counting. Specifically, we design a modality decomposition and fusion module (MDFM), which decomposes RGB and thermal image features into modality-invariant and -specific features by using the similarity and difference supervision between multi-modal features. Besides, reconstruction supervision is also used to prevent network learning from generating bias. After that, different fusion strategies are applied to the invariant and specific features, respectively. In addition, to adapt to the variations in size of different pedestrians, we design a modality-invariant fusion module (MIFM). Finally, after the fusion decoder, MISF-Net can obtain a more accurate crowd density map. Comprehensive experiments on the RGB-T crowd counting dataset show that our MISF-Net can achieve competitive performance.
Baoyang Mu, Feng Shao 0001, Zhengxuan Xie, Hangwei Chen, Zhongjie Zhu, Qiuping Jiang
IEEE Trans. Multim.5
2024 HDR Video Coding Based on Perceptual Optimization
Jiamin Sun, Zhongjie Zhu, Weifeng Cu, Yongqiang Bai, Zhijing Yu
ICIC (6)2
2024 Octree-Retention Fusion: A High-Performance Context Model for Point Cloud Geometry Compression
abstract
Point cloud compression is a pivotal technology for efficient storage and transmission of 3D point cloud data, which has significant implications for practical applications in virtual reality, autonomous driving, and cultural heritage preservation. In this paper, we propose a new learning-based model using the Retentive Network (RetNet) for point cloud compression, which achieves a lower bitrate while maintaining a high peak signal-to-noise ratio (PSNR). We first use an octree structure to segment the point cloud objects. Then, we use octree-based contextual windows to extract pivotal features from relevant sibling and ancestor nodes. Finally, we employ our proposed Octree-Retention model to effectively exploit the prior information between the spatially adjacent nodes for compression. The experimental results show that our method outperforms the state-of-the-art methods on both the LIDAR dataset(SemanticKITTI) and the object dataset(MPEG 8i), demonstrating its effectiveness.
Zhongjie Zhu, Yongqiang Bai, Zhijing Yu
ICMR2
2024 Oriented Object Detection Based on Adaptive Feature Learning and Enrichment
abstract
Oriented object detection has broad utilization in many fields, including urban traffic monitoring, land utilization assessment, and environmental monitoring. However, current oriented object detecting methods are limited in leveraging multiscale information, failing to fully exploit the rich scale variation within images and resulting in suboptimal performance when detecting multiscale targets. Herein, an innovative method SH-Net is proposed based on adaptive feature learning and enrichment. First, an adaptive feature learning module (AFLM) is constructed to enhance the feature learning capability for multiscale objects. Second, a high-resolution feature pyramidal network (HRFPN) is constructed to enhance deep feature fusion for dense and small targets. Finally, a rotated proposal generation (RPG) module and rotated box refinement (RBR) module are proposed to generate and refine the bounding box for extracted oriented objects. The experimental results obtained on the DOTA dataset show that SH-Net can achieve a mAP of 82.67% and surpasses most state-of-the-art methods.
Zhongjie Zhu, Yongqiang Bai, Yuer Wang
IEEE Signal Process. Lett.2
2023 Fast Prediction of Ternary Tree Partition for Efficient VVC Intra Coding
Jiamin Sun, Zhongjie Zhu, Yongqiang Bai, Yuer Wang
CGI (1)2
2022 Distracted driving detection based on the improved CenterNet with attention mechanism
Zhongjie Zhu, Yongqiang Bai, Guanglong Liao, Tingna Liu
Multim. Tools Appl.2
2022 Tensor Product and Tensor-Singular Value Decomposition Based Multi-Exposure Fusion of Images
abstract
Considering multidimensional structure of the multi-exposure images, a new Tensor product and Tensor-singular value decomposition based Multi-Exposure image Fusion (TT-MEF) method is proposed. The main innovation of this work is to explore a new feature representation of multi-exposure images in the new tensor domain and design the fusion strategy on this basis. Specifically, the luminance and the chrominance channels are fused separately to maintain color consistency. For the luminance fusion, the luminance channel of multi-exposure images is divided into two parts, that is, de-mean term and mean term. The de-mean term is represented as a tensor to extract the feature. Then, the tensor product and tensor-singular value decomposition (T-SVD) are used to design a tensor feature extractor. Furthermore, a fusion strategy of the de-mean term is presented according to the visual saliency model, and a fusion strategy of the mean term is defined by the local and the global visual weights to control counterpoise between the local and global luminance. For the chrominance fusion, a new fusion strategy is also designed by the tensor product and T-SVD, similar to the luminance fusion. Finally, the fused image is obtained by combining the luminance and chrominance fusion. Experimental results show that the proposed TT-MEF method generally outperforms the existing state-of-the-art in terms of subjective visual quality and objective evaluation.
Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Zhongjie Zhu, Yongqiang Bai, Yang Song 0015, Huifang Sun
IEEE Trans. Multim.4
2021 No-reference light field image quality assessment based on depth, structural and angular information
Jianjun Xiang, Gangyi Jiang, Mei Yu 0001, Yongqiang Bai, Zhongjie Zhu
Signal Process.5
2021 Reversible data hiding scheme for high dynamic range images based on multiple prediction error expansion
Yongqiang Bai, Gangyi Jiang, Zhongjie Zhu, Haiyong Xu, Yang Song 0015
Signal Process. Image Commun.3
2021 Blind Quality Assessment of Screen Content Images Via Macro-Micro Modeling of Tensor Domain Dictionary
abstract
Screen content images (SCIs) have been rapidly and widely applied in interactive multimedia applications. The problem of quality assessment for SCIs is an interesting research topic. Most of the existing methods use subjective and independent features in gray domain to predict the image quality, which cannot comprehensively characterize the image properties or lack unified mathematical explanation for SCIs. To address these problems, we propose a novel blind quality assessment method based on macro-micro modeling of tensor domain dictionary for SCIs in this article. In the proposed method, the tensor decomposition is explored first to avoid the loss of color information, and then a target dictionary is learned more effectively with the principal components. Furthermore, a macro-micro model is established to characterize the micro and macro features in the target dictionary space, which can provide a systematic mathematical interpretation for feature extraction. For the micro features, a log-normal pooling scheme is designed to enhance the effectiveness of feature aggregation by analyzing the particularity of the statistical distribution of sparse codes. Additionally, the statistical properties are mainly discussed and studied based on the Bernoulli law of large numbers, and then a reliable macro feature is generated to describe the relationship between the statistical distribution and quality degradation of SCIs. Experimental results determined by using three public SCI databases show that the proposed method can perform better than relevant existing methods in the prediction of the visual quality of SCIs, especially in terms of the generalization for distortion type and interpretability for feature generation.
Yongqiang Bai, Zhongjie Zhu, Gangyi Jiang, Huifang Sun
IEEE Trans. Multim.2
2020 Deep Learning-Based Segmentation of Key Objects of Transmission Lines
Yongteng Li, Renwei Tu, Zhongjie Zhu
ICEC5
2019 Learning content-specific codebooks for blind quality assessment of screen content images
Yongqiang Bai, Mei Yu 0001, Qiuping Jiang, Gangyi Jiang, Zhongjie Zhu
Signal Process.5
2019 Efficient Shape Coding for Object-Based 3D Video Applications
abstract
Shape is a popular way to define objects and shape coding is a key technique for object-based 3D video applications. In this paper, the issue of efficient shape coding for object-based 3D video applications is addressed, and a novel contour-based and chain-represented scheme is proposed. For a given 3D shape video, contour extraction and preprocessing are first implemented followed by chain-based representation. Then, to achieve high coding efficiency, a chain-based prediction and compensation technique is developed based on joint motion-compensated prediction and disparity-compensated prediction to effectively exploit the intra-view temporal correlation and the inter-view spatial correlation. Experiments are conducted, and the results demonstrate that the proposed scheme is more efficient than the existing methods, including state-of-the-art methods.
Zhongjie Zhu, Yuer Wang, Gangyi Jiang, Yueping Yang
IEEE Trans. Circuits Syst. Video Technol.1
2017 Hybrid scheme for accurate stereo matching
Zhongjie Zhu, Qin-yan Dai
Neurocomputing1
2017 Unsupervised segmentation of natural images based on statistical modeling
Zhongjie Zhu, Yu-er Wang, Gangyi Jiang
Neurocomputing1
2015 Highly efficient contour-based predictive shape coding
Zhongjie Zhu, Yu-er Wang, Gangyi Jiang
Pattern Recognit. Lett.1
2013 On Perceptually Consistent Image Binarization
abstract
Compared with gray images, bi-level images have only two values for each pixel and are more convenient for transmission and storing. Hence, they are very useful in many practical applications such as in image printing and display. Of various image binarization methods, error diffusion is a popular one that can produce perceptually consistent halftone images. However, most of the existing error diffusion techniques have not given rigid theoretical foundation or explicit theoretical derivation, and most of their diffusion domains and diffusion weights are constant, which make them inconvenient to be used and inefficient in some practical applications. In this paper, a new error diffusion scheme for image binarization is derived and established based on the analysis of features of the human visual system (HVS) and the heat transfer theory, where the local image information and the local pixels' value distribution are kept stable or little changed during the binarization process. As a result, the overall visual quality of binarized images can be kept perceptually similar to the original ones. Experiments are conducted and convincing results are acquired.
Zhongjie Zhu, Yuer Wang, Gangyi Jiang
IEEE Signal Process. Lett.1