Zheng Wang 0059

dblp:181/2834-59 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
20since 2021 · last 2026
0000-0002-6753-6569ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 16 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 PMPGuard: Catching Pseudo-Matched Pairs in Remote Sensing Image-Text Retrieval
abstract
Remote sensing (RS) image–text retrieval faces significant challenges in real-world datasets due to the presence of Pseudo-Matched Pairs (PMPs), semantically mismatched or weakly aligned image–text pairs, which hinder the learning of reliable cross-modal alignments. To address this issue, we propose a novel retrieval framework that leverages Cross-Modal Gated Attention and a Positive–Negative Awareness Attention mechanism to mitigate the impact of such noisy associations. The gated module dynamically regulates cross-modal information flow, while the awareness mechanism explicitly distinguishes informative (positive) cues from misleading (negative) ones during alignment learning. Extensive experiments on three benchmark RS datasets, i.e., RSICD, RSITMD, and RS5M, demonstrate that our method consistently achieves state-of-the-art performance, highlighting its robustness and effectiveness in handling real-world mismatches and PMPs in RS image–text retrieval tasks.
Pengxiang Ouyang, Zheng Wang 0059, Cong Bai
AAAI3
2026 Open-Vocabulary Camouflaged Object Segmentation with Cascaded Vision Language Models
abstract
Open-vocabulary camouflaged object segmentation (OVCOS) seeks to segment and classify camouflaged objects in arbitrary categories, presenting unique challenges due to visual ambiguity and unseen categories. Recent approaches typically adopt a two-stage paradigm: they first segment objects, and then classify the segmented regions using vision language models (VLMs). However, such methods (i) suffer from a domain gap caused by the mismatch between VLMs' full-image training and cropped-region inferencing, and (ii) depend on generic segmentation models optimized for well-delineated objects which are less effective for camouflaged objects. Without explicit guidance, generic segmentation models often overlook subtle boundaries, leading to imprecise segmentation. In this paper, we introduce a novel VLM-guided cascaded framework to address these issues in OVCOS. For segmentation, we leverage the segment anything model (SAM), guided by the VLM. Our framework uses VLM-derived features as explicit prompts to SAM, effectively directing attention to camouflaged regions and significantly improving localization accuracy. For classification, we avoid the domain gap introduced by hard cropping. Instead, we treat the segmentation output as a soft spatial prior using the alpha channel. This retains the full image context while providing precise spatial guidance, leading to more accurate and context-aware classification of camouflaged objects. The same VLM is shared between segmentation and classification to ensure efficiency and semantic consistency. Extensive experiments on both OVCOS and conventional camouflaged object segmentation benchmarks demonstrate the clear superiority of our method, highlighting the effectiveness of leveraging rich VLM semantics for both segmentation and classification of camouflaged objects. Our code and models are open-sourced at https://github.com/intcomp/camouflaged-vlm.
Kai Zhao 0012, Wubang Yuan, Zheng Wang 0059, Guanyi Li, Xiaoqiang Zhu, Deng-Ping Fan, Dan Zeng 0001
Comput. Vis. Media3
2026 Prototype-based latent space distance optimization on vehicle re-identification
Sixian Chan 0001, Jiaao Cui, Zheng Wang 0059, Xiaolong Zhou 0001, Xiaoqin Zhang 0002
Expert Syst. Appl.5
2026 SeGDP: Source-free Cross-domain Few-shot Learning via Semantic Guided Diversity Prompting
abstract
Source-free cross-domain few-shot learning (SF-CDFSL) aims to transfer pre-trained models to target domains with minimal samples, eliminating the need for source domain data. However, limited samples constrain visual diversity and cross-domain images lack inherent semantic context or prior knowledge, impairing the feature discriminability and generalization of large-scale pre-trained models, affecting transfer performance. To tackle these problems, this article introduces Semantic Guided Diversity Prompting (SeGDP), a method that utilizes semantic guided visual prompts to enhance input diversity. Specifically, SeGDP obtains additional diversity features by concatenating different visual prompts to each support sample, guided by randomly combined and sampled text descriptions during training. Additionally, deep prompt tuning and adapter are introduced to learn static knowledge and further enhance the model’s cross-domain adaptation capability. Extensive experimental results across multiple benchmarks demonstrate that the proposed SeGDP achieves state-of-the-art (SOTA) performance under SF-CDFSL task, and rivals the performance of leading source-utilized models. Our code is available on https://github.com/qwzlh/TOMM_submission .
Linhai Zhuo, Zheng Wang 0059, Tianwen Qian, Yuqian Fu
ACM Trans. Multim. Comput. Commun. Appl.2
2025 NeighborRetr: Balancing Hub Centrality in Cross-Modal Retrieval
abstract
Cross-modal retrieval aims to bridge the semantic gap between different modalities, such as visual and textual data, enabling accurate retrieval across them. Despite significant advancements with models like CLIP that align cross-modal representations, a persistent challenge remains: the hubness problem, where a small subset of samples (hubs) dominate as nearest neighbors, leading to biased representations and degraded retrieval accuracy. Existing methods often mitigate hubness through post-hoc normalization techniques, relying on prior data distributions that may not be practical in real-world scenarios. In this paper, we directly mitigate hubness during training and introduce NeighborRetr, a novel method that effectively balances the learning of hubs and adaptively adjusts the relations of various kinds of neighbors. Our approach not only mitigates the hubness problem but also enhances retrieval performance, achieving state-of-the-art results on multiple cross-modal retrieval benchmarks. Furthermore, Neighbor-Retr demonstrates robust generalization to new domains with substantial distribution shifts, highlighting its effectiveness in real-world applications. We make our code publicly available at: https://github.com/NeighborRetr.
Zengrong Lin, Zheng Wang 0059, Tianwen Qian, Pan Mu, Sixian Chan 0001, Cong Bai
CVPR2
2025 PiCNet: Physics-infused Convolution Network for Radar-Based Precipitation Nowcasting
abstract
Meteorological disasters, especially extreme precipitation, cause significant socioeconomic damage, highlighting the need for effective quantitative precipitation nowcasting. Existing methods, often data-driven and resource-intensive, struggle to capture the underlying physical laws of meteorology. This paper introduces a simple yet effective model using an advection simulator to learn precipitation’s physical dynamics, making the predictions more interpretable. Our model also incorporates a physics-guided module to enhance sensitivity to high-intensity rainfall, improving rainfall prediction accuracy. Experiments on the KNMI radar echo dataset demonstrate that our model outperforms state-of-the-art methods, offering better insights into physics-infused precipitation nowcasting.
Zheng Wang 0059, Hanyi Zhang, Cong Bai
ICASSP1
2025 From Swath to Full-Disc: Advancing Precipitation Retrieval with Multimodal Knowledge Expansion
abstract
Accurate near-real-time precipitation retrieval has been enhanced by satellite-based technologies.However, infrared-based algorithms have low accuracy due to weak relations with surface precipitation, whereas passive microwave and radar-based methods are more accurate but limited in range.This challenge motivates the Precipitation Retrieval Expansion (PRE) task, which aims to enable accurate, infrared-based full-disc precipitation retrievals beyond the scanning swath.We introduce Multimodal Knowledge Expansion, a two-stage pipeline with the proposed PRE-Net model.In the Swath-Distilling stage, PRE-Net transfers knowledge from a multimodal data integration model to an infrared-based model within the scanning swath via Coordinated Masking and Wavelet Enhancement (CoMWE).In the Full-Disc Adaptation stage, Self-MaskTune refines predictions across the full disc by balancing multimodal and full-disc infrared knowledge.Experiments on the introduced PRE benchmark demonstrate that PRE-Net significantly advanced precipitation retrieval performance, outperforming leading products like PERSIANN-CCS, PDIR, and IMERG.The code will be available at https://github.com/Zjut-MultimediaPlus/PRE-Net.
Zheng Wang 0059, Kai Ying, Bin Xu 0017, Chunjiao Wang, Cong Bai
KDD (2)1
2025 Event-Driven Hybrid and Cross-Stage Guide for Video Corpus Moment Retrieval
Zheng Wang 0059, Zengrong Lin, Cong Bai
ICMR1
2025 DiffusionAD: Norm-Guided One-Step Denoising Diffusion for Anomaly Detection
abstract
Anomaly detection has garnered extensive applications in real industrial manufacturing due to its remarkable effectiveness and efficiency. However, previous generative-based models have been limited by suboptimal reconstruction quality, hampering their overall performance. We introduce DiffusionAD, a novel anomaly detection pipeline comprising a reconstruction sub-network and a segmentation sub-network. A fundamental enhancement lies in our reformulation of the reconstruction process using a diffusion model into a noise-to-norm paradigm. Here, the anomalous region loses its distinctive features after being disturbed by Gaussian noise and is subsequently reconstructed into an anomaly-free one. Afterward, the segmentation sub-network predicts pixel-level anomaly scores based on the similarities and discrepancies between the input image and its anomaly-free reconstruction. Additionally, given the substantial decrease in inference speed due to the iterative denoising nature of diffusion models, we revisit the denoising process and introduce a rapid one-step denoising paradigm. This paradigm achieves hundreds of times acceleration while preserving comparable reconstruction quality. Furthermore, considering the diversity in the manifestation of anomalies, we propose a norm-guided paradigm to integrate the benefits of multiple noise scales, enhancing the fidelity of reconstructions. Comprehensive evaluations on four standard and challenging benchmarks reveal that DiffusionAD outperforms current state-of-the-art approaches and achieves comparable inference speed, demonstrating the effectiveness and broad applicability of the proposed pipeline.
Hui Zhang 0090, Zheng Wang 0059, Dan Zeng 0001, Zuxuan Wu, Yu-Gang Jiang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Precipitation Retrieval Integrating Multiple Satellite Observations: A Dataset and a Framework
abstract
Multimodal satellite observations have been widely used for precipitation retrieval. Numerous retrieval algorithms and precipitation products have been developed based on these data. However, the integrated retrieval of multimodal data remains challenging due to the modality heterogeneity caused by different data characteristics reflecting precipitation patterns. To address these issues, effectively integrating multimodal data is crucial. We propose a framework named Precipitation Retrieval Integrating Multiple Satellite Observations And Geographical Information (PRMG), consisting of two networks for precipitation identification and estimation respectively. Specifically, PRMG includes a Multi-Branch Fusion (MBF) module for integrating three types of satellite observations: infrared (IR), passive microwave (PMW), and spaceborne precipitation radar (PR), and a geographical information correction (GIC) module to incorporate the geographical information to calibrate precipitation features. We also collect a new precipitation retrieval dataset for multimodal precipitation retrieval, called Precipitation-MG, which includes satellite observations, corresponding geographical information, and precipitation products. Extensive experiments on Precipitation-MG demonstrate the effectiveness of the multimodal fusion method and the geographic correction method. The retrieval performance of PRMG achieves significant improvements compared to the Global Precipitation Measurement currently in operation, i.e., (GPM) Level-2 DPR and GMI Combined (2B-CMB) product. The source code and dataset are publicly available at https://github.com/Zjut-MultimediaPlus/PRMG.
Zheng Wang 0059, Boxian He, Chunjiao Wang, Bin Xu 0017, Cong Bai
IEEE Trans. Geosci. Remote. Sens.1
2025 Diversity-Representativeness Replay and Knowledge Alignment for Lifelong Vehicle Re-identification
abstract
Lifelong Vehicle Re-Identification (LVReID) aims to match a target vehicle across multiple cameras, considering non-stationary and continuous data streams, which fits the needs of the practical application better than traditional vehicle re-identification. Nonetheless, this area has received relatively little attention. Recently, methods for Lifelong Person Re-Identification (LPReID) have been emerging, with replay-based methods achieving the best results by storing a small number of instances from previous tasks for retraining, thus effectively reducing catastrophic forgetting. However, these methods cannot be directly applied to LVReID because they fail to simultaneously consider the diversity and representativeness of replayed data, resulting in biases between the subset stored in the memory buffer and the original data. They randomly sample classes, which may not adequately represent the distribution of the original data. Additionally, these methods fail to consider the rich variation in instances of the same vehicle class due to factors such as vehicle orientation and lighting conditions. Therefore, preserving more informative classes and instances for replay helps maintain information from previous tasks and may mitigate the model's forgetting of old knowledge. In view of this, we propose a novel Diversity-Representativeness Dual-Stage Sampling Replay (DDSR) strategy for LVReID that constructs an effective memory buffer through two stages, i.e. , Cluster-Centric Class Selection and Diverse Instance Mining. Specifically, we first perform class-level sampling based on density in the clustered class-centered feature space and then further mine the diverse, high-quality instances within the selected classes. In addition, we introduce Maximum Mean Discrepancy loss to align the feature distribution between replay data and the new arrivals and apply L2 regularization in the parameter space to facilitate knowledge transfer, thus enhancing the model's generalization ability to new tasks. Extensive experiments demonstrate effective improvements of our method compared to current state-of-the-art lifelong ReID methods on the VeRi-776, VehicleID, and VERI-Wild datasets.
Zhijing Wan, Xiao Wang 0029, Wei Liu 0183, Wei Wang 0170, Zheng Wang 0059, Xin Xu 0007
ACM Trans. Multim. Comput. Commun. Appl.6
2024 Open-Vocabulary Video Relation Extraction
Wentao Tian, Zheng Wang 0059, Yuqian Fu, Jingjing Chen 0001, Lechao Cheng
AAAI2
2024 Distinguishing Visually Similar Images: Triplet Contrastive Learning Framework for Image-text Retrieval
abstract
In recent years, contrastive learning techniques, particularly InfoNCE loss, have propelled advancements in image-text alignment. However, aligning indistinguishable image-text pairs remains a challenge for conventional methods, often overlooking semantically similar content. To overcome these limitations, we introduce the Triplet Contrast Learning Framework (TCLF). Accompanying TCLF is the Stricter Noise Contrast Estimation (SNCE) loss, designed to minimize mutual information between positive and negative image-text pairs. Additionally, we propose the intra-modal mutual information (IMI) loss to encourage a uniform distribution of image features and enhance discriminative capacity for similar images. SNCE aligns two enhanced image versions, and collectively, SNCE and IMI empower TCLF for effective learning with challenging image-text pairs. Experimental evaluations on MSCOCO and Flickr30K datasets demonstrate superior performance compared to state-of-the-art methods. Ablation studies confirm the efficacy of SNCE and IMI in overcoming identified challenges.
Pengxiang Ouyang, Jianan Chen 0002, Zheng Wang 0059, Cong Bai
ICME4
2024 Cross-Modal Quantization for Co-Speech Gesture Generation
abstract
Learning proper representations for speech and gesture is essential for co-speech gesture generation. Existing approaches either utilize direct representations or independently encode the speech and gesture, which neglect the joint representation to highlight the interplay between these two modalities. In this work, we propose a novel Cross-modal Quantization (CMQ) to jointly learn the quantized codes for speech and gesture together. Such representation highlights the speech-gesture interaction before actually learning the complex mapping, and thus better suits the intricate mapping between speech and gesture. Specifically, the Cross-modal Quantizer jointly encodes speech and gesture as discrete codebooks, enabling better cross-modal interaction. Cross-modal Predictor subsequently utilizes the learned codebooks to autoregressively predict the next-step gesture. With cross-modal quantization, our approach yields much higher codebook usage and generates more realistic and diverse gestures in practice. Extensive experiments are conducted on both 3D and 2D datasets as well as the subjective user study, demonstrating a clear performance gain compared to several baseline models in terms of audio-visual alignment and gesture diversity. In particular, our method demonstrates a three-fold improvement in diversity compared to baseline models, while simultaneously maintaining high motion fidelity.
Zheng Wang 0059, Wei Zhang 0031, Long Ye, Dan Zeng 0001, Tao Mei 0001
IEEE Trans. Multim.1
2023 Prototypical Residual Networks for Anomaly Detection and Localization
abstract
Anomaly detection and localization are widely used in industrial manufacturing for its efficiency and effectiveness. Anomalies are rare and hard to collect and supervised models easily over-fit to these seen anomalies with a handful of abnormal samples, producing unsatisfactory performance. On the other hand, anomalies are typically subtle, hard to discern, and of various appearance, making it difficult to detect anomalies and let alone locate anomalous regions. To address these issues, we propose a framework called Prototypical Residual Network (PRN), which learns feature residuals of varying scales and sizes between anomalous and normal patterns to accurately reconstruct the segmentation maps of anomalous regions. PRN mainly consists of two parts: multi-scale prototypes that explicitly represent the residual features of anomalies to normal patterns; a multisize self-attention mechanism that enables variable-sized anomalous feature learning. Besides, we present a variety of anomaly generation strategies that consider both seen and unseen appearance variance to enlarge and diversify anomalies. Extensive experiments on the challenging and widely used MVTec AD benchmark show that PRN outperforms current state-of-the-art unsupervised and supervised methods. We further report SOTA results on three additional datasets to demonstrate the effectiveness and generalizability of PRN.
Hui Zhang 0090, Zuxuan Wu, Zheng Wang 0059, Zhineng Chen, Yu-Gang Jiang 0001
CVPR3
2023 A Generalized Physical-knowledge-guided Dynamic Model for Underwater Image Enhancement
abstract
Underwater images often suffer from color distortion and low contrast resulting in various image types, due to the scattering and absorption of light by water. While it is difficult to obtain high-quality paired training samples with a generalized model. To tackle these challenges, we design a Generalized Underwater image enhancement method via a Physical-knowledge-guided Dynamic Model (short for GUPDM). In particular, to cover complex underwater scenes, this study changes the global atmosphere light and the transmission to simulate various underwater image types through the formation model. We then design an Atmosphere-based Dynamic Structure (ADS) and Transmission-guided Dynamic Structure (TDS) that use dynamic convolutions to adaptively extract prior information from underwater images and generate parameters for Prior-based Multi-scale Structure (PMS). These two modules enable the network to select appropriate parameters for various water types adaptively. Besides, the multi-scale feature extraction module in PMS uses convolution blocks with different kernel sizes and obtains weights for each feature map via channel attention block. The source code will be available at https://github.com/shiningZZ/GUPDM
Pan Mu, Hanning Xu, Zheyuan Liu 0009, Zheng Wang 0059, Sixian Chan 0001, Cong Bai
ACM Multimedia4
2022 Balanced Contrastive Learning for Long-Tailed Visual Recognition
abstract
Real-world data typically follow a long-tailed distribution, where a few majority categories occupy most of the data while most minority categories contain a limited number of samples. Classification models minimizing crossentropy struggle to represent and classify the tail classes. Although the problem of learning unbiased classifiers has been well studied, methods for representing imbalanced data are under-explored. In this paper, we focus on representation learning for imbalanced data. Recently, supervised contrastive learning has shown promising performance on balanced data recently. However, through our theoretical analysis, we find that for long-tailed data, it fails to form a regular simplex which is an ideal geometric configuration for representation learning. To correct the optimization behavior of SCL and further improve the performance of long-tailed visual recognition, we propose a novel loss for balanced contrastive learning (BCL). Compared with SCL, we have two improvements in BCL: classaveraging, which balances the gradient contribution of negative classes; class-complement, which allows all classes to appear in every mini-batch. The proposed balanced contrastive learning (BCL) method satisfies the condition of forming a regular simplex and assists the optimization of cross-entropy. Equipped with BCL, the proposed two-branch framework can obtain a stronger feature representation and achieve competitive performance on long-tailed benchmark datasets such as CIFAR-10-LT, CIFAR-100-LT, ImageNet-LT, and iNaturalist2018.
Jianggang Zhu, Zheng Wang 0059, Jingjing Chen 0001, Yi-Ping Phoebe Chen, Yu-Gang Jiang 0001
CVPR2
2022 Data-Free Network Debiasing for Long-Tailed Visual Recognition
abstract
Real-world data is often unbalanced and exhibits long-tailed distribution over classes. Vanilla classification models trained on imbalanced datasets inherently exhibit bias towards dominant classes. Existing debiasing methods mostly balance the data or the loss during training. Nevertheless, these data-acquiring methods are not suitable for situations where training data are unavailable. In this paper, we appeal to solutions without access to training data and propose a datafree debiasing (Free-D) method that serves as a plug-and-play module for any standard classification model. Specifically, our method adjusts both the feature representation via feature representation shifting and the classifier weight via class prior compensation in a data-free manner. We evaluate and compare our methods on four long-tailed visual recognition datasets, i.e., long-tailed CIFAR-10/-100, ImageNet-LT, and Places-LT. Extensive experiments demonstrate that the proposed data-free method achieves comparable results of other data-acquired methods.
Jinmian Cai, Zheng Wang 0059, Huazhu Fu, Jingjing Chen 0001, Yu-Gang Jiang 0001
ICME2
2021 Visual Co-Occurrence Alignment Learning for Weakly-Supervised Video Moment Retrieval
abstract
Video moment retrieval aims to localize the most relevant video moment given the text query. Weakly supervised approaches leverage video-text pairs only for training, without temporal annotations. Most current methods align the proposed video moment and the text in a joint embedding space. However, in lack of temporal annotations, the semantic gap between these two modalities makes it predominant to learn joint feature representation for most methods, with less emphasis on learning visual feature representation. This paper aims to improve the visual feature representation with supervisions in the visual domain, obtaining discriminative visual features for cross-modal learning. Based on the observation that relevant video moments (i.e., share similar activities) from different videos are commonly described by similar sentences; hence the visual features of these relevant video moments should also be similar despite that they come from different videos. Therefore, to obtain more discriminative and robust visual features for video moment retrieval, we propose to align the visual features of relevant video moments from different videos that co-occurred in the same training batch. Besides, a contrastive learning approach is introduced for learning the moment-level alignment of these videos. Through extensive experiments, we demonstrate that the proposed visual co-occurrence alignment learning method outperforms the cross-modal alignment learning counterpart and achieves promising results for video moment retrieval.
Zheng Wang 0059, Jingjing Chen 0001, Yu-Gang Jiang 0001
ACM Multimedia1
2021 Story-driven Video Editing
abstract
This paper proposes a novel multimedia task: story-driven video editing. Given a story paragraph, this task aims to retrieve related video segments from a gallery of collected video segments and compose them into a video sequence by the storyline order. Our proposed baseline solution consists of three modules: a retrieval module, which returns lists of candidate segments for all query sentences in the story paragraph using an object-aware sentence-segment matching method; a sequence candidate proposal module, which aggregates the retrieved segment sets into a sequence proposal by the submodular optimization method; a sorting module, which arranges the candidates according to the storyline of the paragraph using the Sinkhorn network. We build a benchmark for this task including a reorganized version of the ActivityNet Captions dataset, a well-defined quantitative metric called Evaluation of Segment-to-Sequence Matching (ESSM) for measuring the difference between the generated video segment sequence and the ground truth. Quantitative results of the proposed baseline solution are reported. We hope this new task and benchmark will bring broad research attention and push forward a lot of novel online short video editing applications.
Zheng Wang 0059, Yu-Gang Jiang 0001
IEEE Trans. Multim.1
2019 Composite Binary Decomposition Networks
You Qiaoben, Zheng Wang 0059, Yinpeng Dong, Yu-Gang Jiang 0001, Jun Zhu 0001
AAAI2