VLDB 2026 Research / reviewers in the wild / expert
Ke Zhang 0046
dblp:20/4152-46
· DBLP profile ↗
9ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0003-2415-1519ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Rethinking Pseudo-Label Guided Learning for Weakly Supervised Temporal Action Localization from the Perspective of Noise CorrectionabstractPseudo-label learning methods have been widely applied in weakly-supervised temporal action localization. Existing works directly utilize weakly-supervised base model to generate instance-level pseudo-labels for training the fully-supervised detection head. We argue that the noise in pseudo-labels would interfere with the learning of fully-supervised detection head, leading to significant performance leakage. Issues with noisy labels include:(1) inaccurate boundary localization; (2) undetected short action clips; (3) multiple adjacent segments incorrectly detected as one segment. To target these issues, we introduce a two-stage noisy label learning strategy to harness every potential useful signal in noisy labels. First, we propose a frame-level pseudo-label generation model with a context-aware denoising algorithm to refine the boundaries. Second, we introduce an online-revised teacher-student framework with a missing instance compensation module and an ambiguous instance correction module to solve the short-action-missing and many-to-one problems. Besides, we apply a high-quality pseudo-label mining loss in our online-revised teacher-student framework to add different weights to the noisy labels to train more effectively. Our model outperforms the previous state-of-the-art method in detection accuracy and inference speed greatly upon the THUMOS14 and ActivityNet v1.2 benchmarks. Yuxin Qi 0001, Xi Lin 0003, Ke Zhang 0046, Chun Yuan 0003 |
AAAI | 6 |
| 2025 | Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language ModelsabstractRecent breakthroughs in Multimodal Large Language Models (MLLMs) have gained significant recognition within the deep learning community, where the fusion of the Video Foundation Models (VFMs) and Large Language Models(LLMs) has proven instrumental in constructing robust video understanding systems, effectively surmounting constraints associated with predefined visual tasks. These sophisticated MLLMs exhibit remarkable proficiency in comprehending videos, swiftly attaining unprecedented performance levels across diverse benchmarks. However, their operation demands substantial memory and computational resources, underscoring the continued importance of traditional models in video comprehension tasks. In this paper, we introduce a novel learning paradigm termed MLLM4WTAL. This paradigm harnesses the potential of MLLM to offer temporal action key semantics and complete semantic priors for conventional Weakly-supervised Temporal Action Localization (WTAL) methods. MLLM4WTAL facilitates the enhancement of WTAL by leveraging MLLM guidance. It achieves this by integrating two distinct modules: Key Semantic Matching (KSM) and Complete Semantic Reconstruction (CSR). These modules work in tandem to effectively address prevalent issues like incomplete and over-complete outcomes common in WTAL methods. Rigorous experiments are conducted to validate the efficacy of our proposed approach in augmenting the performance of various heterogeneous WTAL models. Jinwei Fang, Yuxin Qi 0001, Ke Zhang 0046, Chun Yuan 0003 |
CVPR | 6 |
| 2025 | Clip-Ae: Clip-Assisted Cross-View Audio-Visual Enhancement for Unsupervised Temporal Action LocalizationabstractTemporal Action Localization (TAL) has garnered significant attention in information retrieval. Existing supervised or weakly supervised methods heavily rely on labeled temporal boundaries and action categories, which are labor-intensive and time-consuming. Consequently, unsupervised temporal action localization (UTAL) has gained popularity. However, current methods face two main challenges: 1) Classification pre-trained features overly focus on highly discriminative regions; 2) Solely relying on visual modality information makes it difficult to determine contextual boundaries. To address these issues, we propose a CLIP-assisted cross-view audio-visual enhanced UTAL method. Specifically, we introduce visual language pre-training (VLP) and classification pre-training-based collaborative enhancement to avoid excessive focus on highly discriminative regions; we also incorporate audio perception to provide richer contextual boundary information. Finally, we introduce a self-supervised cross-view learning paradigm to achieve multi-view perceptual enhancement without additional annotations. Extensive experiments on two public datasets demonstrate our model’s superiority over several state-of-the-art competitors. Ke Zhang 0046, Chun Yuan 0003 |
ICIP | 4 |
| 2025 | IMDPrompter: Adapting SAM to Image Manipulation Detection by Cross-View Automated Prompt LearningabstractUsing extensive training data from SA-1B, the Segment Anything Model (SAM) has demonstrated exceptional generalization and zero-shot capabilities, attracting widespread attention in areas such as medical image segmentation and remote sensing image segmentation. However, its performance in the field of image manipulation detection remains largely unexplored and unconfirmed. There are two main challenges in applying SAM to image manipulation detection: a) reliance on manual prompts, and b) the difficulty of single-view information in supporting cross-dataset generalization. To address these challenges, we develops a cross-view prompt learning paradigm called IMDPrompter based on SAM. Benefiting from the design of automated prompts, IMDPrompter no longer relies on manual guidance, enabling automated detection and localization. Additionally, we propose components such as Cross-view Feature Perception, Optimal Prompt Selection, and Cross-View Prompt Consistency, which facilitate cross-view perceptual learning and guide SAM to generate accurate masks. Extensive experimental results from five datasets (CASIA, Columbia, Coverage, IMD2020, and NIST16) validate the effectiveness of our proposed method. Yuxin Qi 0001, Jinwei Fang, Xi Lin 0003, Ke Zhang 0046, Chun Yuan 0003 |
ICLR | 6 |
| 2025 | EAV-Mamba: Efficient Audio-Visual Representation Learning for Weakly-Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization aims to learn to locate actions in videos from video-level or point-level labels, avoiding the need for costly frame-level annotations. Unlike previous work that relies solely on visual modality information, we propose incorporating audio information into the weakly supervised temporal action localization task. While audio-visual localization tasks combine audio and visual information for video localization, temporal action localization often deals with action categories that have weak audio cues. To address this, we propose EAV-Mamba, the first audio-visual perception modeling method based on Mamba. Leveraging Mamba’s powerful audio-visual perception capabilities, we developed modules such as Audio-Perceptive Flow Enhancement, Audio-Perceptive RGB Enhancement, and Audio Self-Perceptive Enhancement. Extensive experiments on two publicly available temporal action localization datasets demonstrate that EAV-Mamba achieves efficient audio-visual perception modeling and state-of-the-art performance in weakly supervised temporal action localization tasks. Jinwei Fang, Yuxin Qi 0001, Mingyang Wan, Guojun Ma, Ke Zhang 0046, Chun Yuan 0003 |
ICME | 6 |
| 2024 | PatchNet: Maximize the Exploration of Congeneric Semantics for Weakly Supervised Semantic SegmentationabstractWith the increase in the number of image data and the lack of corresponding labels, weakly supervised learning has drawn a lot of attention recently in computer vision tasks, especially in the fine-grained semantic segmentation problem. To alleviate human efforts from expensive pixel-by-pixel annotations, our method focuses on weakly supervised semantic segmentation (WSSS) with image-level labels, which are much easier to obtain. As a considerable gap exists between pixel-level segmentation and image-level labels, how to reflect the image-level semantic information on each pixel is an important question. To explore the congeneric semantic regions from the same class to the maximum, we construct the patch-level semantic augmentation network (PatchNet) based on the self-detected patches from different images that contain the same class labels. Patches can frame the objects as much as possible and include as little background as possible. The patch-level semantic augmentation network that is established with patches as the nodes can maximize the mutual learning of similar objects. We regard the embedding vectors of patches as nodes and use a transformer-based complementary learning module to construct weighted edges according to the embedding similarity between different nodes. Moreover, to better supplement semantic information, we propose softcomplementary loss functions matched with the whole network structure. We conduct experiments on the popular PASCAL VOC 2012 and MS COCO 2014 benchmarks, and our model yields the state-of-the-art performance. Ke Zhang 0046, Chen Chen 0015, Chun Yuan 0003, Xinfeng Wang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | High-Frequency Normalizing Flow for Image RescalingabstractIt is desirable to develop efficient image rescaling methods to transmit digital images with different resolutions between devices and assure visual quality. In image downscaling, the inevitable loss of high-frequency information makes the reverse upscaling highly ill-posed. Recent approaches focus on joint learning of image downscaling and upscaling (e.g., rescaling). However, existing methods still fail to recover satisfactory high-frequency signals when upscaling. To solve it, we propose high-frequency flow (HfFlow), which learns the distribution of high-frequency signals during rescaling. HfFlow is an overall invertible framework with a conditional flow on the high-frequency space to compensate for the information lost during downscaling. To facilitate finding the optimal upscaling solution, we introduce a reference low-resolution (LR) manifold and propose a cross-entropy Gaussian loss (CGloss) to force the downscaled manifold closer to the reference LR manifold and simultaneously fulfill recovering missing details. HfFlow can be generalized to other scale transformation tasks such as image colorization with its excellent rescaling capacity. Qualitative and quantitative experimental evaluations demonstrate that HfFlow restores rich high-frequency details and outperforms state-of-the-art rescaling methods in PSNR, SSIM, and perceptual quality metrics. Cairong Wang, Chenyu Dong, Ke Zhang 0046, Hongyang Gao, Chun Yuan 0003 |
IEEE Trans. Image Process. | 4 |
| 2023 | Weakly Supervised Instance Segmentation by Exploring Entire Object RegionsabstractWeakly supervised instance segmentation with image-level class supervision is a challenging task as it associates the highest-level instances to the lowest-level appearance. Previous approaches for the task utilize classification networks to obtain rough discriminative parts as seed regions and use distance as a metric to cluster pixels of the same instances. Unlike previous approaches, we provide a novel self-supervised joint learning framework as the basic network and consider the clustering problem as calculating the probability that pixels belong to each instance. To this end, we propose our self-supervised joint learning two-stream network (SJLT Net) to finish this task. In the first stream, we leverage a joint learning framework to implement image-level supervised semantic segmentation with self-supervised saliency detection. In the second stream, we propose a Center Detection Network to detect different instances’ centers with the gaussian loss function to cluster instances pixels. Besides, an integration module is utilized to combine information of both streams and get precise pseudo instances labels. Our approach generates pseudo instance segmentation labels of training images, which are used to train a fully supervised model. Our model achieves excellent performance on the PASCAL VOC 2012 dataset, surpassing the best baseline trained with the same labels by 4.6$\%$$AP^r_{50}$on the train set and 2.6$\%$$AP^r_{50}$on the validation set. Ke Zhang 0046, Chun Yuan 0003, Yong Jiang 0001, Lishu Luo |
IEEE Trans. Multim. | 1 |
| 2020 | Double Shot: Preserve and Erase Based Class Attention Networks for Weakly Supervised Localization (Peca-Net)abstractWeakly supervised localization has attracted increasing attention since only image-wise labels are needed. One mainstream approach, CAM based top-down localization method, suffers from poor resolution and localizing only the most discriminative regions. Another kind, model agnostic perturbation based method, suffers from multiple iterations for each sample. In this paper, we introduce PECA-Net: Preserve and Erase Based Class Attention Networks, which adopts preserve and erase perturbed U-net as the basis, with class activation mechanism as attention to enhance localization capability. Class attention module strengthens informative features and achieves a basic localization. Preserve and erase perturbed U-net replaces the random and iterative extrinsic perturbation with meaningful erasing. In addition, this structure refines the preliminary localization. Since the target object is hit twice, therefore, entitled as double shot. Experiments validate that localization error of both CUB-200 and ILSVRC ImageNet dataset is the new state-of-the-art. Lishu Luo, Chun Yuan 0003, Ke Zhang 0046, Yong Jiang 0001 |
ICME | 3 |