Jin Zhan

dblp:66/8990 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 Bridging CLIP and CLAP for open-vocabulary audio-visual segmentation with semantic coherence
Yunzhi Zhuge, Mengyuan Zhu, Yizhuang Peng, Lu Zhang 0053, Jin Zhan, Huchuan Lu
Pattern Recognit.5
2026 A cross self-attention feature fusion module for 2D multiple human pose estimation
Jin Zhan, Zhenmeng Yue, Weili Tian, Huimin Zhao 0001, Guiyuan Xie, Bo Hu 0023, Fangyuan Lei, Guozhu Liang
Signal Process. Image Commun.1
2026 Context-Infused Trajectories: Enhancing Context and Frame Consistency in Reasoning Video Object Segmentation
abstract
Reasoning video object segmentation (ReaVOS) aims to segment referred objects in video sequences based on implicit and complex linguistic queries. Existing methods typically compress limited video frames into pooled representations and prompt multimodal large language models (MLLMs) to generate a single global segmentation token. However, this strategy lacks explicit contextual guidance and causes substantial loss of spatial details, limiting capability and segmentation consistency. To overcome these limitations, we introduce Context-infused Consistent Video Segmentor (CiCVS), a novel framework leveraging contextual information to guide generation of temporally coherent and accurate mask trajectories. CiCVS incorporates a Hierarchical Frame Sampling (HFS) module, which globally samples support frames across the entire video to ensure broad temporal coverage, and then uniformly selects target frames within the support set. It also employs a Contextual Token Prompting (CTP) module, which utilizes contextual cues from support frames to guide the MLLM in generating specialized tokens for various target frames, enabling the model to capture intricate temporal patterns and ensure consistency across long-range sequences. At the core of CTP is the Multimodal Injection Compressor (MIC) block, which efficiently integrates support frame features and textual semantic information into a compact set of latent queries, enhancing temporal-level object perception. To further advance the ReaVOS field, we introduce the CoCoRVOS benchmark, which features more temporally intricate reasoning instructions and a diverse set of video scenarios. Extensive experiments demonstrate that CiCVS establishes a new state-of-the-art on multiple benchmarks, achieving significant improvements in $\mathcal {J}\& \mathcal {F}$ scores, including +2.7 on CoCoRVOS, +1.4 on ReVOS, and +7.0 on ReasonVOS, underscoring its superior contextual reasoning and segmentation capabilities.
Yunzhi Zhuge, Sitong Gong, Lu Zhang 0053, Qi Xu 0008, Wenda Zhao 0003, Jin Zhan, Huchuan Lu
IEEE Trans. Image Process.6
2026 Exploiting Cross-Task Synergy via Frequency-Driven Hierarchical Learning for Multi-Task Dense Prediction
abstract
Multi-task dense prediction improves pixel-level performance by leveraging shared representations and inter-task collaboration. However, existing approaches either rely on implicit task relationships or neglect frequency-domain cues that are essential for preserving fine-grained details and enhancing cross-task feature learning at multiple scales. As a result, they face persistent challenges in multi-scale feature fusion, effective task interaction, and accurate decoding. To address these issues, we propose a hierarchical frequency-driven framework, termed Hierarchical Frequency-Adaptive Network (HiFAN), that facilitates cross-task collaborative optimization via frequency-domain analysis. Specifically, we first design a task-adaptive fusion module that exploits multi-scale frequency-domain information to enhance spatial details. This module generates dynamic convolutional kernels with task-specific parameters and positional biases to adaptively accommodate diverse task requirements. Next, we introduce an efficient cross-task interaction module that leverages compact low-frequency representations to enable global context exchange across tasks. Finally, we present a high-frequency-aware decoder that mitigates feature smoothing and detail loss commonly introduced by Transformer-based decoders. We demonstrate the effectiveness of HiFAN on two standard multi-task learning benchmarks, PASCAL-Context and NYUD-v2, achieving strong and competitive performance across multiple tasks. The code and model weights are available in HiFAN.
Yunzhi Zhuge, Xinzhuo Yu, Lu Zhang 0053, Xu Jia 0012, Jin Zhan, Huchuan Lu
IEEE Trans. Image Process.5
2026 Local feature enhancement for robust 2D multi-person pose estimation via pose refinement network
abstract
Accurate 2D multi-person pose estimation remains challenging due to issues such as occlusion, missing body parts, and low resolution, particularly in complex backgrounds. This paper proposes an refinement network for multi-person pose estimation through the complementary fusion of extended local receptive fields and contextual information. The proposed cascaded dilated convolution module (DCM) expands the local receptive field through geometric perception, addressing the issue of feature ambiguity in low-resolution and small-scale human bodies. Simultaneously, a hybrid self-attention module (HSM) is introduced to integrate the semantic relevance of joints and precise spatial location information by parallelly combining convolutional self-attention (CSA) and coordinate attention (CA). This optimizes localization through semantic association, not only reducing background interference but also resolving the problem of overlapping joints in multi-person scenarios. Consequently, the network framework achieves an effective balance between the accuracy of human feature extraction at different scales and computational speed. Extensive experiments conducted on the MS COCO and CrowdPose datasets demonstrate that the proposed network architecture outperforms comparable methods, exhibiting superior robustness and computational performance in high-density crowd scenes, uneven lighting conditions, and complex texture scenarios. The related code and models are available at https://github.com/Twl-GZ/Human-pose .
Weili Tian, Jin Zhan, ZhaoKang Guan, Chensheng Yi, Fangyuan Lei, Xiaoyong Liu 0001, Yufeng Zeng
Vis. Comput.2
2025 GaitBranch: A multi-branch refinement model combined with frame-channel attention mechanism for gait recognition
Huakang Li, Yidan Qiu, Huimin Zhao 0001, Jin Zhan, Rongjun Chen 0001, Jinchang Ren, Ying Gao 0004, Wing W. Y. Ng
Comput. Vis. Image Underst.4
2025 Dual-channel hypergraph networks in the time-frequency domain for learning advanced spatiotemporal dependencies in multivariate time series
Jianjian Jiang, Xiangmin Luo, Fangyuan Lei, Xiaochen Yuan, Jin Zhan
Neurocomputing6
2022 GaitSlice: A gait recognition model based on spatio-temporal slice features
Huakang Li, Yidan Qiu, Huimin Zhao 0001, Jin Zhan, Rongjun Chen 0001, Tuanjie Wei
Pattern Recognit.4
2022 A multivariate intersection over union of SiamRPN network for visual tracking
abstract
Abstract SiamPRN algorithm performs well in visual tracking, but it is easy to drift under occlusion and fast motion scenes because it uses $$\ell _1$$ ℓ 1 -smooth loss function to measure the regression location of bounding box. In this paper, we propose a multivariate intersection over union (MIOU) loss in SiamRPN tracking framework. Firstly, MIOU loss includes three geometric factors in regression: the overlap area ratio, the center distance ratio, and the aspect ratio, which can better reflect the coincidence degree of target box and prediction box. Secondly, we improve the definition of aspect ratio loss to avoid gradient explosion, improve the optimization performance of prediction box. Finally, based on SiamPRN tracker, we compared the tracking performance of $$\ell _1$$ ℓ 1 -smooth loss, IOU loss, GIOU loss, DIOU loss, and MIOU loss. Experimental results show that the MIOU loss has better target location regression than other loss functions on the OTB2015 and VOT2016 benchmark, especially for the challenges of occlusion, illumination change and fast motion.
Huimin Zhao 0001, Jin Zhan, Huakang Li
Vis. Comput.3
2019 Compressive sensing based secret signals recovery for effective image Steganalysis in secure communications
Huimin Zhao 0001, Jinchang Ren, Jin Zhan, Yinyin Xiao, Sophia Zhao, Fangyuan Lei, Maher Assaad
Multim. Tools Appl.3
2018 Discriminative Visual Tracking Using Multi-feature and Adaptive Dictionary Learning
Penggen Zheng, Jin Zhan, Huimin Zhao 0001, Jujian Lv
PRCV (4)2
2018 Unsupervised image saliency detection with Gestalt-laws guided optimization and visual attention based refinement
Yijun Yan, Jinchang Ren, Genyun Sun, Huimin Zhao 0001, Junwei Han 0001, Xuelong Li 0001, Stephen Marshall, Jin Zhan
Pattern Recognit.8
2018 Weak-structure-aware visual object tracking with bottom-up and top-down context exploration
Hefeng Wu, Hengzheng Zhu, Jin Zhan
Signal Process. Image Commun.5
2015 Cascaded probabilistic tracking with supervised dictionary learning
Jin Zhan, Hefeng Wu, Huifang Zhang
Signal Process. Image Commun.1
2015 Robust tracking via discriminative sparse feature selection
Jin Zhan, Zhuo Su 0001, Hefeng Wu
Vis. Comput.1
2010 Random noise SAR based on compressed sensing
abstract
Recent theory of compressed sensing (CS) suggested that exact recovery of an unknown sparse signal can be achieved from few measurements with overwhelming probability. In this paper, we combine CS technology with a random noise SAR and proposed the concept of random noise SAR based on CS. The block diagram of the radar system and the collected data processing procedure was presented. Theoretic analysis show that the sensing matrix of the random noise SAR exhibits good restricted isometry property (RIP).When the target scene is sparse or sparse in any basis, the random noise radar based on CS can get high accuracy image by collecting far less amount of echo data than traditional noise radar does. The conclusions are all demonstrated by simulation experiments.
Bingchen Zhang, Yueguan Lin, Wen Hong, Yirong Wu, Jin Zhan
IGARSS6
2010 Collaborative processing of acoustic and magnetic data via heterogeneous sensor networks
abstract
For wireless sensor networks, collaboration of multi-sensors (acoustic, magnetic, radiation, mechanical) enhances the ability of detection. In this paper, an effective approach is proposed in order to process collaboratively acoustic and magnetic data from heterogeneous nodes in underwater sensor networks. And theoretic analyze and simulation of this method are shown in this paper.
Yongqiang Chen 0001, Jin Zhan, Minhui Zhu
IGARSS3