Jinpu Zhang

dblp:263/1392 · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
21since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Retinex-Based Self-Conditioned Diffusion Model for Low-Light Image Enhancement
abstract
The conditional diffusion models have made significant progress in image synthesis, leveraging human annotations such as class labels or text descriptions to guide the generative process. However, different from image synthesis, low-light image enhancement(LLIE) lacks strictly calibrated conditional priors to guide the enhancement process, often resulting in unsatisfactory results. To address the issue, we propose Retinex-Based Self-Conditioned Diffusion Models, dubbed RSCDM, which utilizes self-conditioned illumination representation learning and representation guidance enhancement to generate high-quality image. To be specific, in the first stage, we pretrain a retinex decomposed model (RDM) to capture illumination representation and devise a illumination-representation restoration model (IRM) to accurately reconstruct the representation from noisy images. Moreover, we further design dynamic resblock (DRB) and dynamic simplified attention gated block (DSAGB) as basic units of IRM for better fine-grained restoration. In the second stage, we employ a self-conditioned diffusion model (SDM) to generate realistic results conditioned on the illumination representation. Extensive experiments demonstrates our method outperforms the existing SOTA methods both quantitatively and qualitatively. The codes will be publicly available.
Ziwen Li 0005, Jinpu Zhang, Yuehuan Wang
ICASSP3
2025 HGSLoc: 3DGS-Based Heuristic Camera Pose Refinement
abstract
Visual localization refers to the process of determining camera poses and orientation within a known scene representation. This task is often complicated by factors such as changes in illumination and variations in viewing angles. In this paper, we propose HGSLoc, a novel lightweight plug-and-play pose optimization framework, which integrates 3D reconstruction with a heuristic refinement strategy to achieve higher pose estimation accuracy. Specifically, we introduce an explicit geometric map for 3D representation and high-fidelity rendering, allowing the generation of high-quality synthesized views to support accurate visual localization. Our method demonstrates higher localization accuracy compared to NeRFbased neural rendering localization approaches. We introduce a heuristic refinement strategy, its efficient optimization capability can quickly locate the target node, while we set the steplevel optimization step to enhance the pose accuracy in the scenarios with small errors. With carefully designed heuristic functions, it offers efficient optimization capabilities, enabling rapid error reduction in rough localization estimations. Our method mitigates the dependence on complex neural network models while demonstrating improved robustness against noise and higher localization accuracy in challenging environments, as compared to neural network joint optimization strategies. The optimization framework proposed in this paper introduces novel approaches to visual localization by integrating the advantages of 3D reconstruction and the heuristic refinement strategy, which demonstrates strong performance across multiple benchmark datasets, including 7Scenes and Deep Blending dataset. The implementation of our method has been released at https://github.com/anchang699/HGSLoc.
Zhongyan Niu, Zhen Tan 0002, Jinpu Zhang, Xueliang Yang, Dewen Hu
ICRA3
2025 Adaptive Language-Aware Image Reflection Removal Network
abstract
Existing image reflection removal methods struggle to handle complex reflections. Accurate language descriptions can help the model understand the image content to remove complex reflections. However, due to blurred and distorted interferences in reflected images, machine-generated language descriptions of the image content are often inaccurate, which harms the performance of language-guided reflection removal. To address this, we propose the Adaptive Language-Aware Network (ALANet) to remove reflections even with inaccurate language inputs. Specifically, ALANet integrates both filtering and optimization strategies. The filtering strategy reduces the negative effects of language while preserving its benefits, whereas the optimization strategy enhances the alignment between language and visual features. ALANet also utilizes language cues to decouple specific layer content from feature maps, improving its ability to handle complex reflections. To evaluate the model's performance under complex reflections and varying levels of language accuracy, we introduce the Complex Reflection and Language Accuracy Variance (CRLAV) dataset. Experimental results demonstrate that ALANet surpasses state-of-the-art methods for image reflection removal. The code and dataset are available at https://github.com/fashyon/ALANet.
Siyan Fang, Jinpu Zhang, Ziwen Li 0005, Yuehuan Wang
IJCAI3
2025 Tracking Any Point with Frame-Event Fusion Network at High Frame Rate
abstract
Tracking any point based on image frames is constrained by frame rates, leading to instability in high-speed scenarios and limited generalization in real-world applications. To overcome these limitations, we propose an image-event fusion point tracker, FE-TAP, which combines the contextual information from image frames with the high temporal resolution of events, achieving high frame rate and robust point tracking under various challenging conditions. Specifically, we designed an Evolution Fusion module (EvoFusion) to model the image generation process guided by events. This module can effectively integrate valuable information from both modalities operating at different frequencies. To achieve smoother point trajectories, we employed a transformer-based refinement strategy that updates the point’s trajectories and features iteratively. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches, particularly improving expected feature age by 24% on EDS datasets. Finally, we qualitatively validated the robustness of our algorithm in real driving scenarios using our custom-designed image-event synchronization device.
Jiaxiong Liu, Bo Wang 0144, Zhen Tan 0002, Jinpu Zhang, Hui Shen 0004, Dewen Hu
IROS4
2025 Fully Spiking Neural Networks for Unified Frame-Event Object Tracking
abstract
The integration of image and event streams offers a promising approach for achieving robust visual object tracking in complex environments. However, current fusion methods achieve high performance at the cost of significant computational overhead and struggle to efficiently extract the sparse, asynchronous information from event streams, failing to leverage the energy-efficient advantages of event-driven spiking paradigms. To address this challenge, we propose the first fully Spiking Frame-Event Tracking framework called SpikeFET. This network achieves synergistic integration of convolutional local feature extraction and Transformer-based global modeling within the spiking paradigm, effectively fusing frame and event data. To overcome the degradation of translation invariance caused by convolutional padding, we introduce a Random Patchwork Module (RPM) that eliminates positional bias through randomized spatial reorganization and learnable type encoding while preserving residual structures. Furthermore, we propose a Spatial-Temporal Regularization (STR) strategy that overcomes similarity metric degradation from asymmetric features by enforcing spatio-temporal consistency among temporal template features in latent space. Extensive experiments across multiple benchmarks demonstrate that the proposed framework achieves superior tracking accuracy over existing methods while significantly reducing power consumption, attaining an optimal balance between performance and efficiency.
Jingjun Yang, Liangwei Fan, Jinpu Zhang, Xiangkai Lian, Hui Shen 0004, Dewen Hu
NeurIPS3
2025 Augment One With Others: Generalizing to Unforeseen Variations for Visual Tracking
abstract
Unforeseen appearance variation is a challenging factor for visual tracking. This paper provides a novel solution from semantic data augmentation, which facilitates offline training of trackers for better generalization. We utilize existing samples to obtain knowledge to augment another in terms of diversity and hardness. First, we propose that the similarity matching space in Siamese-like models has class-agnostic transferability. Based on this, we design the Latent Augmentation (LaAug) to transfer relevant variations and suppress irrelevant ones between training similarity embeddings of different classes. Thus the model can generalize across a more diverse semantic distribution. Then, we propose the Semantic Interaction Mix (SIMix), which interacts moments between different feature samples to contaminate structure and texture attributes and retain other semantic attributes. SIMix simulates the occlusion and complements the training distribution with hard cases. The mixed features with adversarial perturbations can empirically enable the model against external environmental disturbances. Experiments on six challenging benchmarks demonstrate that three representative tracking models, i.e., SiamBAN, TransT and OSTrack, can be consistently improved by incorporating the proposed methods without extra parameters and inference cost.
Jinpu Zhang, Ziwen Li 0005, Ruonan Wei, Yuehuan Wang
IEEE Trans. Multim.1
2024 Real-Time Exposure Correction via Collaborative Transformations and Adaptive Sampling
abstract
Most of the previous exposure correction methods learn dense pixel-wise transformations to achieve promising results, but consume huge computational resources. Recently, Learnable 3D lookup tables (3D LUTs) have demon-strated impressive performance and efficiency for image enhancement. However, these methods can only perform global transformations and fail to finely manipulate local regions. Moreover, they uniformly downsample the input image, which loses the rich color information and limits the learning of color transformation capabilities. In this paper, we present a collaborative transformation framework (CoTF) for real-time exposure correction, which integrates global transformation with pixel-wise transformations in an efficient manner. Specifically, the global transformation adjusts the overall appearance using image-adaptive 3D LUTs to provide decent global contrast and sharp details, while the pixel transformation compensates for local context. Then, a relation-aware modulation module is designed to combine these two components effectively. In addition, we propose an adaptive sampling strategy to preserve more color information by predicting the sampling intervals, thus providing higher quality input data for the learning of 3D LUTs. Extensive experiments demonstrate that our method can process high-resolution images in real-time on GPUs while achieving comparable performance against current state-of-the-art methods. The code is avail-able at https://github.com/HUST-IAL/CoTF.
Ziwen Li 0005, Feng Zhang 0039, Jinpu Zhang, Yuanjie Shao, Yuehuan Wang, Nong Sang
CVPR4
2024 Joint Language Prompt and Object Tracking
abstract
Recently, utilizing natural language descriptions to assist in object tracking is becoming a trend. However, The tracking performance is compromised with the absence of the Vision-Language (VL) datasets. In this paper, we develop Joint Language Prompt and Object Tracking (JLPT), the first application of prompt learning in VL tracking. JLPT preserves knowledge from large-scale pretrained vision-based models and adopts language descriptions as prompts to aid tracking with a limited amount of datasets. Specifically, the Multimodal Fusion Prompter (MFP) constructs precise language prompts by dimension compression and correlation operation on heterologous modal. It can adaptively reinforce target features and attenuate interference features. Additionally, the Language Rectification Moudle (LRM) is introduced to dynamically adjust language prompts based on target appearance variations, providing temporal adaptability of the language prompts. JLPT achieves notable performance with minimal trainable language prompt parameters and limited VL datasets. Extensive experiments on TNL2K, LaSOT, LaSOText, and OTB99-L confirm the superiority of JLPT in VL tracking over existing trackers.
Zhimin Weng, Jinpu Zhang, Yuehuan Wang
ICME2
2024 MFRGN: Multi-scale Feature Representation Generalization Network for Ground-to-Aerial Geo-localization
abstract
Cross-area evaluation poses a significant challenge for ground-to-aerial geo-localization, in which the training and testing data are captured from entirely distinct areas. However, current methods struggle in cross-area evaluation due to their emphasis solely on learning global information from single-scale features. Some efforts alleviate this problem but rely on complex and specific technologies like pre-processing and hard sample mining. To this end, we propose a pure end-to-end solution, free from task-specific techniques, termed the Multi-scale Feature Representation Generalization Network (MFRGN) to improve generalization. Specifically, we introduce multi-scale features and explicitly utilize them by an novel global-local information representation structure with two flows, to bolster feature representations. In the global flow, we present a lightweight Self and Cross Attention Module (SCAM) to efficiently learn global embeddings. In the local flow, we develop a Global-Prompt Attention Block (GPAB) to capture discriminative features under the global embeddings as prompts. As a result, our approach generates robust descriptors representing multi-scale global and local information, thereby enhancing the model's invariance to scene variations. Extensive experiments on benchmarks show our MFRGN achieves competitive performance in same-area evaluation and improves cross-area generalization by a significant margin compared to SOTA methods. Our code is available at https://github.com/ytao-wang/MFRGN.
Jinpu Zhang, Ruonan Wei, Yuehuan Wang
ACM Multimedia2
2024 Difficulty-Aware Dynamic Network for Lightweight Exposure Correction
abstract
Recently, deep learning-based methods have been successfully applied to the field of exposure correction. However, most of the existing methods treat different locations of an image in the same way, ignoring the inhomogeneous recovery difficulty and spatially-varying visual patterns in the image, which is sub-optimal and not perfectly efficient. In this paper, we propose a difficulty-aware dynamic network (DDNet) for lightweight exposure correction. Specifically, we propose a difficulty-aware strategy that determines the difficulty of feature patches according to a difficulty mask. Then, only the difficult patches are further refined instead of the whole features, which greatly reduces the overall computational complexity. Moreover, in order to achieve spatially-varying processing with a minimal computational burden, we design a spatial-aware dynamic convolution (SDConv), which is generated by predicting a set of basic kernels and a spatial-aware weight map. Benefiting from these designs, our method can strike a good trade-off between performance and complexity. Extensive experiments on several datasets demonstrate that our approach outperforms the state-of-the-art methods both qualitatively and quantitatively while requiring cheaper computational costs.
Ziwen Li 0005, Yuanjie Shao, Feng Zhang 0039, Jinpu Zhang, Yuehuan Wang, Nong Sang
IEEE Trans. Circuits Syst. Video Technol.4
2023 Learning Mutually in Crowd Scenes for Pedestrian Detection
abstract
Pedestrian detection in crowded scenes is a challenging problem due to the diverse occlusion patterns and highly overlap. To tackle this critical problem, we propose a mutual learning detection network. First, a self-attention mechanism is proposed to achieve mutual learning between individuals by capturing similar semantics among pedestrians. Feature representation of occluded individuals is enhanced by locally fusing similar semantics. Second, mutual loss is designed to improve the consistency of regression and classification. Specifically, regression results are leveraged to make classification score aware of the quality of predicted boxes, and the classification scores help the regression head to accelerate convergence of redundant boxes. Finally, we evaluate our proposed method on MOT20 and CityPersons datasets and achieve comparable state-of-the-art performance using less data. Compared to baseline, our detector obtains 14.4% AP and 11.7% AR gains on challenging MOT20 dataset.
Ruonan Wei, Yuehuan Wang, Jinpu Zhang
ICIP3
2023 Progressive Domain-style Translation for Nighttime Tracking
abstract
Nighttime tracking is challenging due to the lack of sufficient training data and scene diversity. Unsupervised domain adaptation is a solution by transferring knowledge from day (source domain) to night (target domain). It typically involves adversarial training with a domain discriminator on the source and target data to learn domain-invariant features. However, the imbalanced source/target distribution can cause overfitting of the domain discriminator, hindering the domain adaptability. To address this issue, we propose a Progressive Domain-Style Translation (PDST) for domain adaptive nighttime tracking. PDST decomposes and recombines domain-invariant content encodings and domain-specific style encodings of different domains. Thus the rich source domain content is translated to the target domain, expanding the inter-class diversity of the target domain to alleviate overfitting. Moreover, a momentum update manner is introduced to progressively estimate the domain-style encoding from multiple features, which more accurately reflects the statistical domain attribute than an individual image-style. Finally, we incorporate two regularization terms to constrain the content and domain-style consistency in the translation process, ensuring the generated source-like target features are valid to facilitate the training of domain adaptation. Exhaustive experiments demonstrate the domain adaptability and SOTA performance of the proposed method in nighttime tracking.
Jinpu Zhang, Ziwen Li 0005, Ruonan Wei, Yuehuan Wang
ACM Multimedia1
2023 Half Aggregation Transformer for Exposure Correction
Ziwen Li 0005, Jinpu Zhang, Yuehuan Wang
PRCV (10)2
2023 A Dynamic Tracking Framework Based on Scene Perception
Jinpu Zhang, Ziwen Li 0005, Yuehuan Wang
PRCV (12)1
2023 Low-light image enhancement with knowledge distillation
Ziwen Li 0005, Yuehuan Wang, Jinpu Zhang
Neurocomputing3
2023 Corrigendum to "Low-light image enhancement with knowledge distillation" [Neurocomputing 518 (2023) 332-343]
Ziwen Li 0005, Yuehuan Wang, Jinpu Zhang
Neurocomputing3
2023 Spatio-temporal matching for siamese visual tracking
Jinpu Zhang, Kaiheng Dai, Ziwen Li 0005, Ruonan Wei, Yuehuan Wang
Neurocomputing1
2023 SMTN: Multidimensional Fusion and Time-Domain Coding for Object Tracking in Satellite Videos
abstract
Object tracking in satellite videos contains various small targets, such as cars and ships. However, since small targets always lack salient texture features, and have low contrast to the background, it is difficult to detect targets, and distinguish different target instances, which results in tracking failure. In this letter, we propose a Siamese multidimensional fusion and time domain coding network (SMTN) with an efficient attention-based multidimensional information fusion (MDF) module and a time domain information fusion (TDF) module. The MDF module fuses multi-scale template map and search map information to make the network focus on the target, thus improving the target discriminability. In the TDF module, the aggregated temporal information of previous frames is used to adjust the current frame response map through a local-sensing Metaformer module, which suppresses the response of similar interferences. Different from other temporal fusion methods in object tracking in satellite videos, the TDF enables video satellite trackers to be optimized end-to-end, and improves the tracking success rate. Experiments demonstrate that the proposed method outperforms state-of-the-art trackers.
Jinpu Zhang, Yuehuan Wang
IEEE Geosci. Remote. Sens. Lett.2
2022 A systematic knowledge-based method for design of transformable product
Jinpu Zhang, Guozhong Cao, Qingjin Peng, Runhua Tan, Wei Liu 0093, Huangao Zhang
Adv. Eng. Informatics1
2021 A function-oriented biologically analogical approach for constructing the design concept of smart product in Industry 4.0
Guozhong Cao, Yindi Sun, Runhua Tan, Jinpu Zhang, Wei Liu 0093
Adv. Eng. Informatics4
2021 YOLSO: You Only Look Small Object
Jinpu Zhang, Lei Zhang 0145, Yuehuan Wang
J. Vis. Commun. Image Represent.1