Zheng Wang 0044

dblp:181/2834-44 · DBLP profile ↗
← Back
59ranked-venue papers
18as first author
46since 2021 · last 2026
0000-0002-9318-0084ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 8 first-author · 32 since 2021Artificial intelligence and machine learning · 21 · 8 first-author · 16 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies
abstract
Developing intelligent agents capable of operating a wide range of Graphical User Interfaces (GUIs) with human-level proficiency is a key milestone on the path toward Artificial General Intelligence. While most existing datasets and benchmarks for training and evaluating GUI agents are static and idealized, failing to reflect the complexity and unpredictability of real-world environments, particularly the presence of anomalies. To bridge this research gap, we propose D-GARA, a dynamic benchmarking framework, to evaluate Android GUI agent robustness in real-world anomalies. D-GARA introduces a diverse set of real-world anomalies that GUI agents commonly face in practice, including interruptions such as permission dialogs, battery warnings, and update prompts. Based on D-GARA framework, we construct and annotate a benchmark featuring commonly used Android applications with embedded anomalies to support broader community research. Comprehensive experiments and results demonstrate substantial performance degradation in state-of-the-art GUI agents when exposed to anomaly-rich environments, highlighting the need for robustness-aware learning. D-GARA is modular and extensible, supporting the seamless integration of new tasks, anomaly types, and interaction scenarios to meet specific evaluation goals.
Yi Bin, Fei Ma 0006, Wenqi Shao, Zheng Wang 0044
AAAI6
2026 Hyper-Opinion Vagueness Quantification for Robust Multimodal Learning
abstract
Robust Multimodal Learning (RML) aims to address the issues of unreliable predictions of multimodal models. Nevertheless, previous RML works often struggle to distinguish between different categories that rely on identical intra-modal cues, making ambiguous predictions. We defined this degree of ``uncertain'' in extracting discriminative features of a multimodal model as vagueness. Neglecting such vagueness, as previous RML works commonly do, will undermine the ability to extract unique semantics of each category in multimodal models, further resulting in worse robustness under disturbances that affect semantic representations. Additionally, this vagueness will lead the parameter updating processes towards unreliable fusion, thus diverting the learning processes of the multimodal model from learning unique features of each category. Based on the above insight, we propose a novel robust multimodal learning approach, termed Hyper-Opinion Quantifying Vagueness (HOQV). Specifically, we first introduce hyper-opinion to capture and quantify the vagueness of multimodal learning in discriminating representations of different categories. Moreover, to mitigate the interference in parameter updating of unreliable representations with high vagueness, we also design the Hyper-Opinion Gradient Modulation to guide the optimization processes. We evaluate our HOQV on six datasets with different disturbances, including noise and adversarial attack, and demonstrate that our proposed method achieves state-of-the-art performance consistently.
Disen Hu, Xun Jiang 0001, Xiaofeng Cao 0002, Zheng Wang 0044, Jingkuan Song, Heng Tao Shen, Xing Xu 0001
AAAI4
2026 Text answer guided RGB-D saliency detection
Zheng Wang 0044, Jingkuan Song
Expert Syst. Appl.2
2026 Generalizable Egocentric Task Verification via Cross-Modal Hybrid Hypergraph Matching
abstract
Egocentric Task Verification (ETV) aims to determine if the operation flows of procedural tasks in egocentric videos align with the logic of given rules. Early works adopt the video-based verification paradigm that compares a reference video to the testing video, which limits the flexibility of model deployment. Recent researches incorporate reference textual rules instead of videos, describing the operational logic with natural language, but also raises the challenges of cross-modal heterogeneity and hierarchical misalignment between the two modalities. While previous works mainly address the cross-modal heterogeneity between vision and text modalities, they inevitably suffer from two additional key challenges: (1) Existing methods are mostly developed in synthetic domains, yet have not considered the issues of synthetic-to-realistic generalization challenges in real-world applications. (2) The intricate relations between visual content and textual rule involve multiple matching correlations, indicating high-order matching interactions. To address these issues, we proposed the Generalizable Egocentric Task Verification (GETV), and construct a cross-domain ETV benchmark dataset, EgoCross. It features synthetic-to-real cross-domain evaluation, covering both synthetic datasets for training and realistic datasets for testing, across three different types of tasks. Furthermore, we also propose a novel method for this challenge, termed Cross-modal Hybrid Hypergraph Matching (CHHM), which models the logical cross-modal matching in the GETV challenge as a heterogeneous hybrid hypergraph learning process, thus addressing intrinsic multiple matching correlations. Additionally, to tackle the problems of synthetic-to-realistic generalization, we enhance the cross-modal matching process with prototype-based graph representation alignment, which effectively mitigates the cross-domain gap. Extensive experiments on the existing two ETV benchmark datasets, i.e., EgoTV and CSV-NL, and our proposed GETV dataset EgoCross, demonstrate our approach establishes new state-of-the-art performance on both intra-domain and cross-domain challenges.
Xun Jiang 0001, Xing Xu 0001, Zheng Wang 0044, Jingkuan Song, Fumin Shen, Heng Tao Shen
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Distribution-to-Points Matching for Image Text Retrieval
abstract
Eliminating semantic discrepancy between different modalities is the ultimate goal of image text retrieval. However, most of the existing methods only focus on retrieval of the ground-truth instance while ignoring those semantically similar instances yet unlabeled as positives, which causes the phenomenon of one-to-many correspondence. The mainstream solutions of this research are mainly based on uncertainty learning and the exploration of one-to-many correspondence is still insufficient albeit their significant progress. Therefore, this work develops a novel Distribution-to-Points (termed D2P) matching mechanism for image-text retrieval to capture the one-to-many correspondence between multiple samples and a given query via hypergraph modeling. Specifically, a given query is first mapped as a probabilistic embedding to learn its true semantic distribution based on Mahalanobis distance. Then each candidate instance in a mini-batch is regarded as a hypergraph node with its mean semantics while a Gaussian query is modeled as a hyperedge to capture the semantic correlations beyond the pair between candidate points and the query. Moreover, an energy-based semantic modeling framework is developed to pull all similar candidates (not only the ground truth) close to their query while pushing those dissimilar ones far away. In the end, distribution-to-points matching is learned based on the similarity measurement over the Mahalanobis distance, which considers semantic variance to perform many-to-one correspondence well. Experimental results on several widely used datasets and under various evaluation metrics confirm our superiority and effectiveness in improving the retrieval ability of the baseline including ground-truth matching and semantic multiplicity for image text retrieval.
Zheng Wang 0044, Xing Xu 0001, Lei Zhu 0002, Jingkuan Song, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 A2VAD: Attribute-augmented prompt learning for weakly supervised video anomaly detection
Zheng Wang 0044, Xing Xu 0001, Jingkuan Song, Zhe Sun 0009, Andrzej Cichocki
Pattern Recognit.2
2026 Causal-Inspired Fourier Representation Learning for Wearable IMUs and Egocentric Action Recognition
abstract
Inertial Measurement Units (IMUs) can capture intricate kinematic behaviors, thereby enhancing the performance of human action recognition methods. Consequently, this technology has recently garnered considerable attention within this domain. However, existing methods either encounter limitations in instance-level visual representation due to self-occlusion or fail to fully utilize the potential of complex kinematic information, making it challenging to adequately capture the intricate relationships between the two data sources. In this paper, we tackle this issue by addressing the problem through causal learning and Fourier learning. Specifically, we introduce a novel framework calledCausal-Inspired Fourier Representation Learning (CIFRL)for Wearable IMUs and Egocentric Action Recognition, which aims to enhance cross-modal feature alignment. The framework consists of two key components: (1) Temporal Causal Modeling (TCM), designed for video interpretation from a causal perspective; (2) Spectral-Temporal Learning (STL), which aims to decompose the inertial data using Fourier representation and align cross-modal features. We evaluate our proposed framework on the WEAR and CMU-MMAC benchmarks. Empirical results demonstrate the superior performance of our CIFRL approach compared to state-of-the-art methods. Our code is available at https://github.com/Adrianos1219/CIFRL.
Xing Xu 0001, Zheng Wang 0044, Jingkuan Song, Fumin Shen, Heng Tao Shen
IEEE Trans. Circuits Syst. Video Technol.3
2026 Egocentric Online Action Segmentation via Parametric Context Memory Learning
abstract
To facilitate smart wearable devices or human-like robotics with real-time first-person perspective perception ability, recent researchers proposed the Egocentric Online Action Segmentation (EOAS) task. It requires models to recognize what is happening in egocentric streaming videos and discriminate the starting and ending times of an activity in a real-time manner. However, compared with offline-recorded exocentric videos, egocentric streaming videos cannot provide equivalent sufficient temporal-spatial cues due to the limited perspective and unknown coming frames. Hence, it raises a high demand for the long-term episodic memory ability of models. To this end, most previous approaches work on compressing long-term memory into feature representations. In this paper, we propose a novel EOAS paradigm, termed Parametric Context Memory Learning (PCML), which integrates episodic memory into learnable parameters and keeps dynamic updates according to real-time frames. Concretely, we design the Parametric Context Perception layer and construct a novel Episodic Semantic Memorization Network (ESMN) based on it, which integrates episodic memory into learnable parameters and keeps dynamic updates with real-time frames. We evaluate our proposed method on three public egocentric streaming video benchmarks including EgoPER, EgoProceL, and GTEA. Extensive experiments demonstrate the ESMN model significantly outperforms recent state-of-the-art methods. Our code is available at https://github.com/XunCHN/PCML.
Xun Jiang 0001, Xing Xu 0001, Zheng Wang 0044, Jingkuan Song, Zhe Sun 0009, Andrzej Cichocki, Heng Tao Shen
IEEE Trans. Image Process.4
2026 Progressively Alleviating Noise for Unsupervised Cross-Domain Image Retrieval
abstract
Unsupervised cross-domain image retrieval (UCDIR) aims to obtain images with the same category from other domains yet without labels for all data. The prevailing strategy first enhances the intra-domain representation by clustering, and then realizes cross-domain alignment based on various similarity comparisons. Despite their significant advancements, current methods are hindered by substantial ambiguity resulting from the absence of explicit semantic information, exacerbated by various sources of noise such as background interference and attribute composition. This difficulty can compromise the performance of existing clustering-based methods. Therefore, we propose a straightforward yet effective strategy to Progressively Alleviate Noise (termedPAN) for UCDIR. We initially utilize information entropy to quantify the certainty of prediction, which also serves as an indicator of uncertainty level resulting from noise. We further refine a progressive learning approach that specifically targets data with low confidence to enhance the semantic comprehension. The model can gradually focus on data with appropriate uncertainty for semantic learning and cross-domain alignment via dynamic reweighing, thereby excluding the adverse effects of noisy data in an incremental manner. Extensive experiments on three widely used datasets for UCDIR demonstrate our effectiveness and superiority.
Zheng Wang 0044, Xing Xu 0001, Lei Zhu 0002, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Multim.1
2026 SCL: Semantic Coherence Learning for Video Question Answering
abstract
Video Question Answering (VideoQA) requires clarifying the true question intent and understanding dynamic video content to predict the answer. However, most existing methods overlook the temporal semantic exploration of fine-grained clues across frames, and neglect the interrogative properties of question text, leading to suboptimal grounding of video frames. To address these limitations, in this paper, we propose a novel Semantic Coherence Learning (SCL) framework, which explicitly models the fine-grained spatio-temporal coherence of critical visual regions and progressively mines multi-modal semantic clues aligned with the reasoning intent. Specifically, we design a spatio-temporal clue reasoning module to adaptively model the fine-grained coherent transition semantics of critical regions across temporal frames while suppressing irrelevant noisy regions within frames. Additionally, we propose a multi-modal clue reasoning module to progressively mine multi-modal coherent semantics associated with reasoning intent, which iteratively grounds critical frames in videos based on refined question semantics and mines critical phrases in questions based on refined video semantics. Finally, the answer is derived through an answer decoder with intra and inter-sample contrastive strategies. Extensive experimental results on six widely used benchmarks verify our effectiveness and superiority. The experimental source codes are available at https://github.com/XizeWu/SCL.
Xize Wu, Zheng Wang 0044, Jiasong Wu, Lei Zhu 0002, Heng Tao Shen
IEEE Trans. Multim.2
2025 Egocentric Online Action Segmentation with Behavior-Centred Feature Augmentation
abstract
The Egocentric Online Action Segmentation (EOAS) task aims to sequentially segment untrimmed egocentric videos into distinct action segments in a streaming manner. Previous methods primarily focused on improving contextual information utilization, which highly relied on leveraging the prior context. However, under online constraints, the absence of post context limits the effectiveness of the prior context in learning the action semantics. Excessive reliance on prior context may lead to insufficient feature representations of current presented behavior. To tackle this problem, we propose a novel EOAS method, termed Behavior-Centred Feature Augmentation (BCFA), which consists of two key modules: (1) Behavior Prototype Learning models the common sense of each action across different surroundings, enhancing the model’s ability to capture the shared characteristics of behaviors. (2) Presented Behavior Enhancement leverages both the intrinsic characteristics of the current presented behavior itself and the common sense captured by BPL for feature enhancement, mitigating the absence of post contextualization. We evaluate our proposed BCFA method on three public EOAS benchmark datasets, GTEA, EgoProceL, and EgoPER, and demonstrate that our proposed BCFA approach outperforms recent state-of-the-art methods.
Zhangye Han, Xun Jiang 0001, Zheng Wang 0044, Xin Liu 0011, Fumin Shen, Xing Xu 0001
ICME3
2025 Probabilistic Embeddings with Causal Constraint for Error Detection in Egocentric Procedural Videos
abstract
Error detection in egocentric procedural task videos aims to identify deviations to support intelligent monitoring and task automation. Despite making significant progress, existing methods that leverage prototypes for egocentric error detection have two drawbacks: (1) The neglect of inherent data traits, i.e., large intra-class variance and minimal inter-class distinction. (2) The absence of causal consistency in temporal modeling. To address these challenges, we introduce a novel framework termed Probabilistic Embeddings with Causal Constraint (PECC) for error detection in egocentric procedural videos. Specifically, we first integrated a causal dilated convolution module in temporal action segmentation model to capture temporal causal consistency. We then train Gaussian Mixture Models (GMMs) for each action class to get frame-level probabilistic embeddings. Finally, We evaluate test frames using log-likelihood values to detect erroneous actions. Extensive experiments conducted on EgoPER and HoloAssist demonstrate that our method achieves state-of-the-art performance, significantly surpassing existing methods in error detection. Our code is available at https://github.com/HouTong-s/PECC-for-Error-Detection-in-Egocentric-Videos.
Tong Hou, Shenshen Li, Xun Jiang 0001, Zheng Wang 0044, Fumin Shen, Xing Xu 0001
ICME4
2025 Cross-Modal Task Verification via Hypergraph-based Sequential Matching
abstract
Cross-Modal Task Verification (CMTV) assesses whether a procedural task is executed accurately according to language-based rules, presenting challenges due to its multi-modal and chronological nature. Existing methods using graph or neuro-symbolic approaches face two issues: (1) Conventional methods only model linear sequential relationships among intra-modal nodes, ignoring implicit relationships between non-neighboring nodes. (2) Directed graphs model pairwise cross-modal relationships but overlook cases where a step corresponds to multiple video segments. To address these issues, we propose Hypergraph-based Sequential Matching (HSM) with two components: (1) Temporal Complementary Hypergraph Module (TCHM), a hierarchical sequential hyperedge construction method that focuses on both sequential connections and implicit relationships across nodes. (2) Step-wise Hypergraph Modeling (SHM), a novel hypergraph-based alignment mechanism that better aligns an action description with multiple video segments, improving task verification accuracy. We evaluate HSM on EgoTV and CTV datasets, demonstrating its superiority over state-of-the-art methods.
Xun Jiang 0001, Zheng Wang 0044, Fumin Shen, Jingkuan Song, Xing Xu 0001
ICME3
2025 Noise Mitigation for Unsupervised Cross-Domain Image Retrieval
abstract
Cross-domain image retrieval task is derived from traditional image retrieval task, wherein the model aims to find images in another domain that share the same semantic meaning. Existing methods first achieve the pseudo labels through clustering algorithm, and then perform the cross-domain alignment. However, in unsupervised scenarios, data lacking explicit semantic information can easily induce the model to produce erroneous predictions, which can significantly deteriorate existing clustering-based methods. To mitigate the influence of these noisy instances, we propose the Noise Mitigation (NM) method, utilizing information entropy to separate noisy instances from clear data. Moreover, our label adaption strategy can enhance the prediction accuracy of noisy data by leveraging clear data. We subsequently apply the explicit semantic maximization strategy, selectively construct an intermediate domain through fusing the explicit semantic in clear data, further reducing the influence of noisy instances. Our approach is evaluated on three datasets, and the experimental results demonstrate the overwhelming performance superiority of our noise mitigation strategy.
Zheng Wang 0044, Xin Liu 0011, Fumin Shen, Xing Xu 0001
ICME3
2025 Fast Policy: Accelerating Visuomotor Policies without Re-training
abstract
Diffusion models are increasingly employed in visuomotor policies to achieve promising performance of behavior cloning. However, the slow inference caused by iterative denoising is a notorious disadvantage, which greatly limits its application in resource-limited and real-time interactive robot systems. The prevailing strategy to this problem is distillation, but it still requires considerable resources to retrain a student model. To this end, we take another training-free view to develop a novel Fast Policy (termed FP), which can be regarded as a powerful and accelerated alternative to Diffusion Policy for learning visuomotor robot control. Specifically, our comprehensive study of UNet encoder shows that its features change little during inference, prompting us to reuse encoder features in non-critical denoising steps. In addition, we design strategies based on Fourier energy to screen critical and non-critical steps dynamically according to different tasks. Importantly, to mitigate performance degradation caused by the repeated use of non-critical steps, we further introduce a noise correction strategy. Our FP is evaluated on multiple simulation benchmarks and the comparison results with existing speed-up methods demonstrate our effectiveness and superiority with state-of-the-art success rates in visuomotor inference speed. The code is available at https://github.com/xwccchong/Fast-Policy
Tongshu Wu, Zheng Wang 0044
IROS2
2025 Composed Query-Based Event Retrieval in Video Corpus with Multimodal Episodic Perceptron
abstract
Event retrieval involves searching for specific events from untrimmed video galleries and has garnered significant attention in recent years. However, most existing works follow a text-based video retrieval paradigm only, limited by two main drawbacks: (1) The episodic information presented in described events is not fully perceived, leading to declines in retrieval performance facing variable query intentions. (2) Current models are prone to returning false positive results with similar semantics, as simple text queries can hardly accurately describe the target video content users seek. In this paper, we propose a novel event retrieval framework termed Composed Query-Based Event Retrieval (CQBER). Specifically, we first construct two CQBER benchmark datasets, namely ActivityNet-CQ and TVR-CQ, which cover TV shows and open-world scenarios, respectively. Additionally, we propose an initial CQBER method, termed Multimodal Episodic Perceptron (MEP), which excavates complete query semantics from both observed static visual cues and various descriptions. Extensive experiments demonstrate that our proposed framework significantly boosts event retrieval accuracy across different existing methods. Our code and datasets are available at https://github.com/VincentVanNF/CQBER.
Fan Ni, Xun Jiang 0001, Hao Yang 0015, Zheng Wang 0044, Fumin Shen, Xing Xu 0001
ICMR6
2025 PSCon: Product Search Through Conversations
abstract
Conversational Product Search ( CPS ) systems interact with users via natural language to offer personalized and context-aware product lists. However, most existing research on CPS is limited to simulated conversations, due to the lack of a real CPS dataset driven by human-like language. Moreover, existing conversational datasets for e-commerce are constructed for a particular market or a particular language and thus can not support cross-market and multi-lingual usage. In this paper, we propose a CPS data collection protocol and create a new CPS dataset, called PSCon, which assists product search through conversations with human-like language. The dataset is collected by a coached human-human data collection protocol and is available for dual markets and two languages. By formulating the task of CPS, the dataset allows for comprehensive and in-depth research on six subtasks: user intent detection, keyword extraction, system action prediction, question selection, item ranking, and response generation. Moreover, we present a concise analysis of the dataset and propose a benchmark model on the proposed CPS dataset. Our proposed dataset and model will be helpful for facilitating future research on CPS.
Jie Zou 0001, Mohammad Aliannejadi, Evangelos Kanoulas, Shuxi Han, Heli Ma, Zheng Wang 0044, Yang Yang 0002, Heng Tao Shen
SIGIR6
2025 Chatting with interactive memory for text-based person retrieval
Shenshen Li, Zheng Wang 0044, Fumin Shen, Xing Xu 0001
Multim. Syst.3
2025 Evidence-Based Multi-Feature Fusion for Adversarial Robustness
abstract
The accumulation of adversarial perturbations in the feature space makes it impossible for Deep Neural Networks (DNNs) to know what features are robust and reliable, and thus DNNs can be fooled by relying on a single contaminated feature. Numerous defense strategies attempt to improve their robustness by denoising, deactivating, or recalibrating non-robust features. Despite their effectiveness, we still argue that these methods are under-explored in terms of determining how trustworthy the features are. To address this issue, we propose a novel Evidence-based Multi-Feature Fusion (termed EMFF) for adversarial robustness. Specifically, our EMFF approach introduces evidential deep learning to help DNNs quantify the belief mass and uncertainty of the contaminated features. Subsequently, a novel multi-feature evidential fusion mechanism based on Dempster's rule is proposed to fuse the trusted features of multiple blocks within an architecture, which further helps DNNs avoid the induction of a single manipulated feature and thus improve their robustness. Comprehensive experiments confirm that compared with existing defense techniques, our novel EMFF method has obvious advantages and effectiveness in both scenarios of white-box and black-box attacks, and also prove that by integrating into several adversarial training strategies, we can improve the robustness of across distinct architectures, including traditional CNNs and recent vision Transformers with a few extra parameters and almost the same cost.
Zheng Wang 0044, Xing Xu 0001, Lei Zhu 0002, Yi Bin, Guoqing Wang 0001, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 CMNet: Cross-Modal Coarse-to-Fine Network for Point Cloud Completion Based on Patches
abstract
Point clouds serve as the foundational representation of 3D objects, playing a pivotal role in both computer vision and computer graphics. Recently, the acquisition of point clouds has been effortless because of the development of hardware devices. However, the collected point clouds may be incomplete due to environmental conditions, such as occlusion. Therefore, completing partial point clouds becomes an essential task. The majority of current methods address point cloud completion via the utilization of shape priors. While these methods have demonstrated commendable performance, they often encounter challenges in preserving the global structural and geometric details of the 3D shape. In contrast to those mentioned earlier, we propose a novel cross-modal coarse-to-fine network (CMNet) for point cloud completion. Our method utilizes additional image information to provide global information, thus avoiding the loss of structure. To ensure that the generated results contain sufficient geometric details, we propose a coarse-to-fine learning approach based on multiple patches. Specifically, we encode the image and use multiple generators to generate multiple coarse patches, which are combined into a complete shape. Subsequently, based on the coarse patches generated in advance, we generate fine patches by combining partial point cloud information. Experimental results show that our method achieves state-of-the-art performance on point cloud completion.
Zhenjiang Du, Zhitao Liu, Jiwei Wei, Sophyani Banaamwini Yussif, Zheng Wang 0044, Ning Xie 0003, Yang Yang 0002
IEEE Trans. Circuits Syst. Video Technol.6
2025 Topology Learning for Two-View Correspondence Filtering
abstract
In this paper, we propose a novel neural network called Topology Learning Network (TL-Net), that exploits local and global geometric relation by topology graphs to handle the problem of correspondence filtering in complex scenes. Specifically, we first design a Multi-level Topology Encoder (MLTE), which fuses local and global topology graphs by a channel attention, to sufficiently extract the geometric relation among correspondences. MLTE not only includes local topology graphs by gathering the information of relative motion and multi-resolution group convolution, but also includes a global topology graph by aggregating the information of the similarity and the Graph Laplacian. In addition, inspired by Transformer, we design the backbone of TL-Net to generate enriched fdeature maps for correspondence filtering. Meanwhile, by simplifying the global context aggregation, we maintain the lightweight of the backbone, introducing the superiority of Transformer while avoiding extra parameters and calculations. Empirical experiments on several computer vision tasks show that the performance and generalization ability of TL-Net are significantly superior to the state of the art methods. Notably, on relative pose estimation, we achieve 5.63% and 5.03% mAP improvements under an error threshold of$5^{\circ }$outdoors and indoors, respectively.
Ziwei Shi, Xiangyang Miao, Guobao Xiao, Songlin Du, Zheng Wang 0044, Heng Tao Shen
IEEE Trans. Multim.5
2025 Geometric Matching for Cross-Modal Retrieval
abstract
Despite its significant progress, cross-modal retrieval still suffers from one-to-many matching cases, where the multiplicity of semantic instances in another modality could be acquired by a given query. However, existing approaches usually map heterogeneous data into the learned space as deterministic point vectors. In spite of their remarkable performance in matching the most similar instance, such deterministic point embedding suffers from the insufficient representation of rich semantics in one-to-many correspondence. To address the limitations, we intuitively extend a deterministic point into a closed geometry and develop geometric representation learning methods for cross-modal retrieval. Thus, a set of points inside such a geometry could be semantically related to many candidates, and we could effectively capture the semantic uncertainty. We then introduce two types of geometric matching for one-to-many correspondence, i.e., point-to-rectangle matching (dubbed P2RM) and rectangle-to-rectangle matching (termed R2RM). The former treats all retrieved candidates as rectangles with zero volume (equivalent to points) and the query as a box, while the latter encodes all heterogeneous data into rectangles. Therefore, we could evaluate semantic similarity among heterogeneous data by the Euclidean distance from a point to a rectangle or the volume of intersection between two rectangles. Additionally, both strategies could be easily employed for off-the-self approaches and further improve the retrieval performance of baselines. Under various evaluation metrics, extensive experiments and ablation studies on several commonly used datasets, two for image-text matching and two for video-text retrieval, demonstrate our effectiveness and superiority.
Zheng Wang 0044, Zhenwei Gao, Yang Yang 0002, Guoqing Wang 0001, Chengbo Jiao, Heng Tao Shen
IEEE Trans. Neural Networks Learn. Syst.1
2024 Ensemble Diversity Facilitates Adversarial Transferability
abstract
With the advent of ensemble-based attacks, the transfer-ability of generated adversarial examples is elevated by a noticeable margin despite many methods only employing superficial integration yet ignoring the diversity between ensemble models. However, most of them compromise the latent value of the diversity between generated perturbation from distinct models which we argue is also able to increase the adversarial transferability, especially heterogeneous at-tacks. To address the issues, we propose a novel method of Stochastic Mini-batch black-box attack with Ensemble Reweighing using reinforcement learning (SMER) to produce highly transferable adversarial examples. We emphasize the diversity between surrogate models achieving indi-vidual perturbation iteratively. In order to customize the individual effect between surrogates, ensemble reweighing is introduced to refine ensemble weights by maximizing attack loss based on reinforcement learning which functions on the ultimate transferability elevation. Extensive exper-iments demonstrate our superiority to recent ensemble at-tacks with a significant margin across different black-box attack scenarios, especially on heterogeneous conditions. https://github.com/tangbwb/SMER
Zheng Wang 0044, Yi Bin, Qi Dou 0001, Yang Yang 0002, Heng Tao Shen
CVPR2
2024 Diverse Embedding Modeling with Adaptive Noise Filter for Text-based Person Retrieval
abstract
Text-based person retrieval (TBPR) involves retrieving pedestrian images from a gallery using textual queries. Existing methods assume that all the training pairs are correct and the textual query only corresponds to one image. However, in practical scenarios of TBPR, there indeed exists data noise and many-to-many matching relationships between semantically similar images and textual queries. To address these problems, we propose a novel approach termed Diverse Embedding Modeling (DEM) with Adaptive Noise Filter for the TBPR task. Firstly, we propose a dynamic margin to measure the degree of noise, which can adaptively reduce the weights of image-text pairs with severe noise during training, thereby effectively mitigating the impact of noisy pairs. Moreover, we model diverse visual and textual embeddings from learnable parameterized distributions, which aim to simulate the many-to-many matching scenarios. Extensive experiments conducted on three TBPR datasets demonstrate the superior performance of our DEM method compared to recent state-of-the-art methods.
Shenshen Li, Zheng Wang 0044, Fumin Shen, Yang Yang 0002, Xing Xu 0001
ICME3
2024 Shapley Ensemble Adversarial Attack
abstract
An intuitive strategy for generating more transferable adversarial examples is to absorb the advantages of different surrogates in an ensemble. However, existing ensemble adversarial attacks simply average the outputs of various models while ignoring their different contributions. To quantify the importance of various surrogates, we propose a novel Shapley Ensemble Adversarial Attack (dubbed SEAA) – an effective algorithm that allocates weights based on Shapley values of the surrogates. Specifically, the synthesis of adversarial examples in an ensemble adversarial attack is firstly regarded as a cooperative game process of multiple surrogates from the perspective of contribution. As the Shapley value is known as an effective decision metric in cooperative game theory, it is thus intuitively introduced into this work to accurately evaluate the contribution of a surrogate by reweighing its importance at each iteration, thus avoiding local optimality and steering the generation of adversarial examples. Comprehensive experimental results demonstrate our effectiveness.
Zheng Wang 0044, Yi Bin, Lei Zhu 0002, Guoqing Wang 0001, Yang Yang 0002
ICME1
2024 PTAN: Principal Token-aware Adjacent Network for Compositional Temporal Grounding
abstract
Compositional temporal grounding (CTG) aims to localize the most relevant segment from an untrimmed video based on a given natural language sentence, and the test samples for this task contain novel components not seen in training. However, existing CTG methods suffer from two shortcomings: (1) Most methods adopt transformers to model global video information only, thus failing to balance the long-range perception and regional representation of video sequences; (2) Due to the lack of aligning videos and sentences at a fine-grained level, the model's capacity for compositional generalization is limited, particularly when query sentences contain novel components. To address these problems, we propose a novel method called Principal Token-aware Adjacent Network (PTAN), which consists of three parts: (1) Principal Temporal Token Recomposition combining video clip-level features obtained from the transformer backbone to capture more significant local features while retaining enough contextual information. (2) Regional Semantic-Aware Learning, which exploits regional representations of videos for cross-modal semantic alignment on the feature space. (3) Principal Semantic-Aware Learning that facilitates fine-grained alignment between visual and textual by sensing principal visual and textual tokens in a self-supervised manner. Extensive experiments on two widely used benchmarks (i.e., Charades-CG and ActivityNet-CG) show that our PTAN method outperforms recent CTG state-of-the-art methods, achieving remarkable improvements in compositional generalization. Our code is available at https://github.com/rushzy/PTAN.
Zhuoyuan Wei, Xun Jiang 0001, Zheng Wang 0044, Fumin Shen, Xing Xu 0001
ICMR3
2024 GalleryGPT: Analyzing Paintings with Large Multimodal Models
abstract
Artwork analysis is important and fundamental skill for art appreciation, which could enrich personal aesthetic sensibility and facilitate the critical thinking ability. Understanding artworks is challenging due to its subjective nature, diverse interpretations, and complex visual elements, requiring expertise in art history, cultural background, and aesthetic theory. However, limited by the data collection and model ability, previous works for automatically analyzing artworks mainly focus on classification, retrieval, and other simple tasks, which is far from the goal of AI. To facilitate the research progress, in this paper, we step further to compose comprehensive analysis inspired by the remarkable perception and generation ability of large multimodal models. Specifically, we first propose a task of composing paragraph analysis for artworks, i.e., painting in this paper, only focusing on visual characteristics to formulate more comprehensive understanding of artworks. To support the research on formal analysis, we collect a large dataset PaintingForm, with about 19k painting images and 50k analysis paragraphs. We further introduce a superior large multimodal model for painting analysis composing, dubbed GalleryGPT, which is slightly modified and fine-tuned based on LLaVA architecture leveraging our collected data. We conduct formal analysis generation and zero-shot experiments across several datasets to assess the capacity of our model. The results show remarkable performance improvements comparing with powerful baseline LMMs, demonstrating its superb ability of art analysis and generalization. \textcolor{blue}{The codes and model are available at: https://github.com/steven640pixel/GalleryGPT.
Yi Bin, Yujuan Ding, Zheng Wang 0044, Yang Yang 0002, See-Kiong Ng, Heng Tao Shen
ACM Multimedia5
2024 Cascaded Adversarial Attack: Simultaneously Fooling Rain Removal and Semantic Segmentation Networks
abstract
When applying high-level visual algorithms to rainy scenes, it is customary to preprocess the rainy images using low-level rain removal networks, followed by visual networks to achieve the desired objectives. Such a setting has never been explored by adversarial attack methods, which are only limited to attacking one kind of them. Considering the deficiency of multi-functional attacking strategies and the significance for open-world perception scenarios, we are the first to propose a Cascaded Adversarial Attack (CAA) setting, where the adversarial example can simultaneously attack different-level tasks, such as rain removal and semantic segmentation in an integrated system. Specifically, our attack on the rain removal network aims to preserve rain streaks in the output image, while for the semantic segmentation network, we employ powerful existing adversarial attack methods to induce misclassification of the image content. Importantly, CAA innovatively utilizes binary masks to effectively concentrate the aforementioned two significantly disparate perturbation distributions on the input image, enabling attacks on both networks. Additionally, we propose two variants of CAA, which minimize the differences between the two generated perturbations by introducing a carefully designed perturbation interaction mechanism, resulting in enhanced attack performance. Extensive experiments validate the effectiveness of our methods, demonstrating their superior ability to significantly degrade the performance of the downstream task compared to methods that solely attack a single network.
Zhiwen Wang 0004, Yuhui Wu 0001, Zheng Wang 0044, Jiwei Wei, Tianyu Li 0003, Guoqing Wang 0001, Yang Yang 0002, Heng Tao Shen
ACM Multimedia3
2024 Semantics Disentangling for Cross-Modal Retrieval
abstract
Cross-modal retrieval (e.g., query a given image to obtain a semantically similar sentence, and vice versa) is an important but challenging task, as the heterogeneous gap and inconsistent distributions exist between different modalities. The dominant approaches struggle to bridge the heterogeneity by capturing the common representations among heterogeneous data in a constructed subspace which can reflect the semantic closeness. However, insufficient consideration is taken into the fact that learned latent representations are actually heavily entangled with those semantic-unrelated features, which obviously further compounds the challenges of cross-modal retrieval. To alleviate the difficulty, this work makes an assumption that the data are jointly characterized by two independent features: semantic-shared and semantic-unrelated representations. The former presents characteristics of consistent semantics shared by different modalities, while the latter reflects the characteristics with respect to the modality yet unrelated to semantics, such as background, illumination, and other low-level information. Therefore, this paper aims to disentangle the shared semantics from the entangled features, andthus the purer semantic representation can promote the closeness of paired data. Specifically, this paper designs a novel Semantics Disentangling approach for Cross-Modal Retrieval (termed as SDCMR) to explicitly decouple the two different features based on variational auto-encoder. Next, the reconstruction is performed by exchanging shared semantics to ensure the learning of semantic consistency. Moreover, a dual adversarial mechanism is designed to disentangle the two independent features via a pushing-and-pulling strategy. Comprehensive experiments on four widely used datasets demonstrate the effectiveness and superiority of the proposed SDCMR method by achieving a new bar on performance when compared against 15 state-of-the-art methods.
Zheng Wang 0044, Xing Xu 0001, Jiwei Wei, Ning Xie 0003, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Image Process.1
2024 Estimating the Semantics via Sector Embedding for Image-Text Retrieval
abstract
Based on deterministic single-point embedding, most extant image-text retrieval methods only focus on the match of ground truth while suffering from one-to-many correspondence, where besides annotated positives, many similar instances of another modality should be retrieved by a given query. Recent solutions of probabilistic embedding and rectangle mapping still encounter some drawbacks, albeit their promising effectiveness at multiple matches. Meanwhile, the exploration of one-to-many correspondence is still insufficient. Therefore, this paper proposes a novel geometric representation toEstimate theSemantics of heterogeneous data viaSectorEmbedding (dubbed ESSE). Specifically, a given image/text can be projected as a sector, where its symmetric axis represents mean semantics and the aperture estimates uncertainty. Further, a sector matching loss is introduced to better handle the multiplicity by considering the sine of included angles as distance calculation, which encourages candidates to be contained by the apertures of a query sector. The experimental results on three widely used benchmarks CUB, Flickr30K and MS-COCO reveal that sector embedding can achieve competitive performance on multiple matches and also improve the traditional ground-truth matching of the baselines. Additionally, we also verify the generalization to video-text retrieval on two extensively used datasets of MSRVTT and MSVD, and to text-based person retrieval on CUHK-PEDES. This superiority and effectiveness can also demonstrate that the bounded property of the aperture can better estimate semantic uncertainty when compared to prior remedies.
Zheng Wang 0044, Zhenwei Gao, Mengqun Han, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Multim.1
2024 Complex Relation Embedding for Scene Graph Generation
abstract
Given an input image, scene graph generation (SGG) aims to generate comprehensive visual relationships between objects in the form of graphs. Recently, more attention to the design of complex networks and complicated strategies has been paid to the long tail issue caused by the imbalanced class distribution. However, most existing methods adopt the concatenated features of two objects in real space as the final relation representation for a given triplet. We mainly argue that such a simple concatenation may neglect the importance of complex interactions between objects, which results in the diversity of visual relations. In addition, the representation learning in real space is also inadequate to express this property. To alleviate these issues, we seamlessly incorporate Hermitian inner product into existing models to facilitate the generation of scene graphs by learning Relation Embedding in Complex space (CoRE). More specifically, we first introduce the concept of complex-valued representations for entities and then formulate the relation triplets with Hermitian inner product in complex space. Finally, we investigate the effect of utilizing only real component or both of Hermitian inner product on inferring more reasonable interaction between objects for scene graphs. Comprehensive experiments on two widely used benchmark datasets, Visual Genome (VG) and Open Image, demonstrate our effectiveness, superiority, and generalization on various metrics for biased or unbiased inference.
Zheng Wang 0044, Xing Xu 0001, Yin Zhang 0002, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Neural Networks Learn. Syst.1
2023 Multilateral Semantic Relations Modeling for Image Text Retrieval
abstract
Image-text retrieval is a fundamental task to bridge vision and language by exploiting various strategies to finegrained alignment between regions and words. This is still tough mainly because of one-to-many correspondence, where a set of matches from another modality can be accessed by a random query. While existing solutions to this problem including multi-point mapping, probabilistic distribution, and geometric embedding have made promising progress, one-to-many correspondence is still under-explored. In this work, we develop a Multilateral Semantic Relations Modeling (termed MSRM) for image-text retrieval to capture the one-to-many correspondence between multiple samples and a given query via hypergraph modeling. Specifically, a given query is first mapped as a probabilistic embedding to learn its true semantic distribution based on Mahalanobis distance. Then each candidate instance in a mini-batch is regarded as a hypergraph node with its mean semantics while a Gaussian query is modeled as a hyperedge to capture the semantic correlations beyond the pair between candidate points and the query. Comprehensive experimental results on two widely used datasets demonstrate that our MSRM method can outper-form state-of-the-art methods in the settlement of multiple matches while still maintaining the comparable performance of instance-level matching.
Zheng Wang 0044, Zhenwei Gao, Kangshuai Guo, Yang Yang 0002, Heng Tao Shen
CVPR1
2023 Revisiting Domain-Adaptive 3D Object Detection by Reliable, Diverse and Class-balanced Pseudo-Labeling
abstract
Unsupervised domain adaptation (DA) with the aid of pseudo labeling techniques has emerged as a crucial approach for domain-adaptive 3D object detection. While effective, existing DA methods suffer from a substantial drop in performance when applied to a multi-class training setting, due to the co-existence of low-quality pseudo labels and class imbalance issues. In this paper, we address this challenge by proposing a novel ReDB framework tailored for learning to detect all classes at once. Our approach produces Reliable, Diverse, and class-Balanced pseudo 3D boxes to iteratively guide the self-training on a distributionally different target domain. To alleviate disruptions caused by the environmental discrepancy (e.g., beam numbers), the proposed cross-domain examination (CDE) assesses the correctness of pseudo labels by copy-pasting target instances into a source environment and measuring the prediction consistency. To reduce computational overhead and mitigate the object shift (e.g., scales and point densities), we design an overlapped boxes counting (OBC) metric that allows to uniformly downsample pseudo-labeled objects across different geometric characteristics. To confront the issue of inter-class imbalance, we progressively augment the target point clouds with a class-balanced set of pseudo-labeled target instances and source objects, which boosts recognition accuracies on both frequently appearing and rare classes. Experimental results on three benchmark datasets using both voxel-based (i.e., SECOND) and point-based 3D detectors (i.e., PointRCNN) demonstrate that our proposed ReDB approach outperforms existing 3D domain adaptation methods by a large margin, improving 23.15% mAP on the nuScenes → KITTI task. The code is available at https://github.com/zhuoxiao-chen/ReDB-DA-3Ddet.
Zhuoxiao Chen, Yadan Luo, Zheng Wang 0044, Mahsa Baktash, Zi Huang
ICCV3
2023 Quaternion Representation Learning for cross-modal matching
Zheng Wang 0044, Xing Xu 0001, Jiwei Wei, Ning Xie 0003, Jie Shao 0001, Yang Yang 0002
Knowl. Based Syst.1
2023 Hypercomplex context guided interaction modeling for scene graph generation
Zheng Wang 0044, Xing Xu 0001, Yadan Luo, Guoqing Wang 0001, Yang Yang 0002
Pattern Recognit.1
2023 Visual Embedding Augmentation in Fourier Domain for Deep Metric Learning
abstract
Deep Metric Learning (DML) is very effective for many computer vision applications such as image retrieval or cross-modal matching. The common paradigm for DML is to seek metric spaces that can encode semantically similar objects close while locating the dissimilar ones far away from each other. To make features more discriminative, the mainstream methods usually design various specific loss functions to seek the help of hard negatives through complex hard mining strategies or hard synthesizing with additional networks. In spite of their fruitfulness, these approaches ignore the impact of low-level information in images on the performance, which may degrade the discerning ability of learned embedding. To alleviate these problems, we introduce a simple yet effective augmentation method to generate more hard negatives by swapping the low-frequency spectra of negative instances with anchors in the Fourier domain. Specifically, unlike previous methods, our proposed approach does not involve any complex design strategies but enriches hard negatives by manipulating the low-level variability of images only with simple Fourier transforms. In addition, our method is treated as a universal plug-in, which can be incorporated into different models for performance improvement. In the end, we conduct extensive experiments to evaluate our method on the widely-used datasets including CUB-200–2011, CARS-196, and Stanford Online Products. Our quantitative results demonstrate that the proposed plug-in outperforms previous approaches consistently and significantly across different datasets and evaluation metrics.
Zheng Wang 0044, Zhenwei Gao, Guoqing Wang 0001, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Circuits Syst. Video Technol.1
2023 Category Alignment Adversarial Learning for Cross-Modal Retrieval
abstract
Cross-modal retrieval aims to retrieve one semantically similar media from multiple media types based on queries entered by another type of media. An intuitive idea is to map different media data into a common space and then directly measure content similarity between different types of data. In this paper, we present a novel method, called Category Alignment Adversarial Learning (CAAL) for cross-modal retrieval. It aims to find a common representation space supervised by category information, in which the samples from different modalities can be compared directly. Specifically, CAAL firstly employs two parallel encoders to generate common representations for image and text features respectively. Furthermore, we employ two parallel GANs with category information to generate fake image and text features which next will be utilized with already generated embedding to reconstruct the common representation. At last, two joint discriminators are utilized to reduce the gap between the mapping of the first stage and the embedding of the second stage. Comprehensive experimental results on four widely-used benchmark datasets demonstrate the superior performance of our proposed method compared with the state-of-the-art approaches.
Shiyuan He, Weiyang Wang, Zheng Wang 0044, Xing Xu 0001, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Knowl. Data Eng.3
2023 Quaternion Relation Embedding for Scene Graph Generation
abstract
As an important visual understanding task, scene graph generation has been drawing widespread attention and could boost a broad range of downstream vision applications. Traditional scene graph generation methods based on different context refinements are trained with probabilistic chain rule, which treats objects and relationships as independent entities. Despite their surprisingly great progress, such a plain formulation unconsciously ignores the latent geometric structure of entities and relationships. To address this issue, we move beyond the traditional real-valued representations and useQuaternionRelationEmbedding (QuatRE) to generate scene graphs with more expressive hypercomplex representations. More specifically, we introduce the concept of quaternion representations, hyper-complex valued with three imaginary components for objects entities, then formulate the relation triplets with Hamilton product. Benefiting from explicitly modeling the latent inter-dependencies among all imaginary components and strong expressive capacity, our proposed QuatRE method could better capture the interactions between entities. More importantly, our novel QuatRE method can be treated as a plug-in and well generalized into other methods for performance improvement as it involves no additional layers. Finally, extensive comparisons of our proposed method against the state-of-the-art methods on two large-scale and widely-used datasets, i.e. Visual Genome and Open Images, demonstrated our superiority and generalization capability on various metrics for biased or unbiased inference.
Zheng Wang 0044, Xing Xu 0001, Guoqing Wang 0001, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Multim.1
2022 Point to Rectangle Matching for Image Text Retrieval
abstract
The difficulty of image-text retrieval is further exacerbated by the phenomenon of one-to-many correspondence, where multiple semantic manifestations of the other modality could be obtained by a given query. However, the prevailing methods adopt the deterministic embedding strategy to retrieve the most similar candidate, which encodes the representations of different modalities as single points in vector space. We argue that such a deterministic point mapping is obviously insufficient to represent a potential set of retrieval results for one-to-many correspondence, despite its noticeable progress. As a remedy to this issue, we propose a Point to Rectangle Matching (abbreviated as P2RM) mechanism, which actually is a geometric representation learning method for image-text retrieval. Specifically, our intuitive insight is that the representations of different modalities could be extended to rectangles, then a set of points inside such a rectangle embedding could be semantically related to many candidate correspondences. Thus our P2RM method could essentially address the one-to-many correspondence. Besides, we design a novel semantic similarity measurement method from the perspective of distance for our rectangle embedding. Under the evaluation metric for multiple matches, extensive experiments and ablation studies on two commonly used benchmarks demonstrate our effectiveness and superiority in tackling the multiplicity of image-text retrieval.
Zheng Wang 0044, Zhenwei Gao, Xing Xu 0001, Yadan Luo, Yang Yang 0002, Heng Tao Shen
ACM Multimedia1
2022 MRA-Net: Improving VQA Via Multi-Modal Relation Attention Network
abstract
Visual Question Answering (VQA) is a task to answer natural language questions tied to the content of visual images. Most recent VQA approaches usually apply attention mechanism to focus on the relevant visual objects and/or consider the relations between objects via off-the-shelf methods in visual relation reasoning. However, they still suffer from several drawbacks. First, they mostly model the simple relations between objects, which results in many complicated questions cannot be answered correctly, because of failing to provide sufficient knowledge. Second, they seldom leverage the harmony cooperation of visual appearance feature and relation feature. To solve these problems, we propose a novel end-to-end VQA model, termed Multi-modal Relation Attention Network (MRA-Net). The proposed model explores both textual and visual relations to improve performance and interpretability. In specific, we devise 1) a self-guided word relation attention scheme, which explore the latent semantic relations between words; 2) two question-adaptive visual relation attention modules that can extract not only the fine-grained and precise binary relations between objects but also the more sophisticated trinary relations. Both kinds of question-related visual relations provide more and deeper visual semantics, thereby improving the visual reasoning ability of question answering. Furthermore, the proposed model also combines appearance feature with relation feature to reconcile the two types of features effectively. Extensive experiments on five large benchmark datasets, VQA-1.0, VQA-2.0, COCO-QA, VQA-CP v2, and TDIUC, demonstrate that our proposed model outperforms state-of-the-art approaches.
Yang Yang 0002, Zheng Wang 0044, Zi Huang, Heng Tao Shen
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Universal adversarial perturbations generative network
Zheng Wang 0044, Yang Yang 0002, Jingjing Li 0001, Xiaofeng Zhu 0001
World Wide Web1
2021 Attention-Based Relation Reasoning Network for Video-Text Retrieval
abstract
In the field of video-text matching, there are several potential and effective internal relations within a single modal data, which the existing approaches always ignore. In this paper, we propose a novel model named Attention-based Relation Reasoning Network (ARRN), that can robustly learn and reason the word relations of a sentence and temporal relations between video frames. It can jointly capture the local and global characteristics of video and text, thus significantly improves the performance on video-text retrieval. In ARRN, with global-to-local attention strategy, we could attend to important relations of multi-scales, then learn more reasonable local relation features. These features, generated at distinct levels, are powerful and complementary to each other, allowing us to obtain effective video and text representations by very simple fusion. The extensive experiments on two widely-used video-text datasets MSVD and TGIF show that our proposed ARRN approach establishes a substantial improvement.
Zheng Wang 0044, Xing Xu 0001, Fumin Shen, Yang Yang 0002, Heng Tao Shen
ICME2
2021 Multi-scale Dynamic Network for Temporal Action Detection
abstract
In recent years, as the fundamental task in video understanding, Temporal Action Detection is attracting extensive attention. Most existing approaches use the same model parameters to process all input videos, which are not adaptive to the input video during the inference stage. In this paper, we propose a novel model termed Multi-scale Dynamic Network (MDN) to tackle this problem. The proposed MDN model incorporates multiple Multi-scale Dynamic Modules (MDMs). Each MDM can generate video-specific and segment-specific convolution kernels based on video content from different scales and adaptively capture rich semantic information for the prediction. Besides, we also design a new Edge Suppression Loss (ESL) function for MDN to pay more attention to hard examples. Extensive experiments conducted on two popular benchmarks ActivityNet-1.3 and THUMOS-14 show that the proposed MDN model achieves the state-of-the-art performance.
Yifan Ren, Xing Xu 0001, Fumin Shen, Zheng Wang 0044, Yang Yang 0002, Heng Tao Shen
ICMR4
2021 Relationship-Preserving Knowledge Distillation for Zero-Shot Sketch Based Image Retrieval
abstract
Zero-shot sketch-based image retrieval is challenging for the modal gap between distributions of sketches and images and the inconsistency of label spaces during training and testing. Previous methods mitigate the modal gap by projecting sketches and images into a joint embedding space. Most of them also bridge seen and unseen classes by leveraging semantic embeddings, i.e., word vectors and hierarchical similarities. In this paper, we propose Relationship-Preserving Knowledge Distillation (RPKD) to study generalizable embeddings from the perspective of knowledge distillation bypassing the usage of semantic embeddings. In particular, we firstly distill the instance-level knowledge to preserve inter-class relationships without semantic similarities that require extra effort to collect. We also reconcile the contrastive relationships among instances between different embedding spaces, which is complementary to instance-level relationships. Furthermore, embedding-induced supervision, which measures the similarities of an instance to partial class embedding centers from the teacher, is developed to align the student's classification confidences. Extensive experiments conducted on three benchmark ZS-SBIR datasets, i.e., Sketchy, TU-Berlin, and QuickDraw, demonstrate the superiority of our proposed RPKD approach comparing to the state-of-the-art methods.
Xing Xu 0001, Zheng Wang 0044, Fumin Shen, Xin Liu 0011
ACM Multimedia3
2021 Disentangled Representation Learning and Enhancement Network for Single Image De-Raining
abstract
In this paper, we present a disentangled representation learning and enhancement network (DRLE-Net) to address the challenging single image de-raining problems, i.e., raindrop and rain streak removal. Specifically, the DRLE-Net is formulated as a multi-task learning framework, and an elegant knowledge transfer strategy is designed to train the encoder of DRLE-Net to embed a rainy image into two separated latent spaces representing the task (clean image reconstruction in this paper) relevant and irrelevant variations respectively, such that only the essential task-relevant factors will be used by the decoder of DRLE-Net to generate high-quality de-raining results. Furthermore, visual attention information is modeled and fed into the disentangled representation learning network to enhance the task-relevant factor learning. To facilitate the optimization of the hierarchical network, a new adversarial loss formulation is proposed and used together with the reconstruction loss to train the proposed DRLE-Net. Extensive experiments are carried out for removing raindrops or rainstreaks from both synthetic and real rainy images, and DRLE-Net is demonstrated to produce significantly better results than state-of-the-art models.
Guoqing Wang 0001, Changming Sun, Xing Xu 0001, Jingjing Li 0001, Zheng Wang 0044, Zeyu Ma 0002
ACM Multimedia5
2021 Meta Self-Paced Learning for Cross-Modal Matching
abstract
Cross-modal matching has attracted growing attention due to the rapid emergence of the multimedia data on the web and social applications. Recently, many re-weighting methods have been proposed for accelerating model training by designing a mapping function from similarity scores to weights. However, these re-weighting methods are difficult to be universally applied in practice since manually pre-set weighting functions inevitably involve hyper-parameters. In this paper, we propose a Meta Self-Paced Network (Meta-SPN) that automatically learns a weighting scheme from data for cross-modal matching. Specifically, a meta self-paced network composed of a fully connected neural network is designed to fit the weight function, which takes the similarity score of the sample pairs as input and outputs the corresponding weight value. Our meta self-paced network considers not only the self-similarity scores, but also their potential interactions (e.g., relative-similarity) when learning the weights. Motivated by the success of meta-learning, we use the validation set to update the meta self-paced network during the training of the matching network. Experiments on two image-text matching benchmarks and two video-text matching benchmarks demonstrate the generalization and effectiveness of our method.
Jiwei Wei, Xing Xu 0001, Zheng Wang 0044, Guoqing Wang 0001
ACM Multimedia3
2020 Learning Cross-Aligned Latent Embeddings for Zero-Shot Cross-Modal Retrieval
abstract
Zero-Shot Cross-Modal Retrieval (ZS-CMR) is an emerging research hotspot that aims to retrieve data of new classes across different modality data. It is challenging for not only the heterogeneous distributions across different modalities, but also the inconsistent semantics across seen and unseen classes. A handful of recently proposed methods typically borrow the idea from zero-shot learning, i.e., exploiting word embeddings of class labels (i.e., class-embeddings) as common semantic space, and using generative adversarial network (GAN) to capture the underlying multimodal data structures, as well as strengthen relations between input data and semantic space to generalize across seen and unseen classes. In this paper, we propose a novel method termed Learning Cross-Aligned Latent Embeddings (LCALE) as an alternative to these GAN based methods for ZS-CMR. Unlike using the class-embeddings as the semantic space, our method seeks for a shared low-dimensional latent space of input multimodal features and class-embeddings by modality-specific variational autoencoders. Notably, we align the distributions learned from multimodal input features and from class-embeddings to construct latent embeddings that contain the essential cross-modal correlation associated with unseen classes. Effective cross-reconstruction and cross-alignment criterions are further developed to preserve class-discriminative information in latent space, which benefits the efficiency for retrieval and enable the knowledge transfer to unseen classes. We evaluate our model using four benchmark datasets on image-text retrieval tasks and one large-scale dataset on image-sketch retrieval tasks. The experimental results show that our method establishes the new state-of-the-art performance for both tasks on all datasets.
Kaiyi Lin, Xing Xu 0001, Lianli Gao, Zheng Wang 0044, Heng Tao Shen
AAAI4
2020 Universal Weighting Metric Learning for Cross-Modal Matching
abstract
Cross-modal matching has been a highlighted research topic in both vision and language areas. Learning appropriate mining strategy to sample and weight informative pairs is crucial for the cross-modal matching performance. However, most existing metric learning methods are developed for unimodal matching, which is unsuitable for cross-modal matching on multimodal data with heterogeneous features. To address this problem, we propose a simple and interpretable universal weighting framework for cross-modal matching, which provides a tool to analyze the interpretability of various loss functions. Furthermore, we introduce a new polynomial loss under the universal weighting framework, which defines a weight function for the positive and negative informative pairs respectively. Experimental results on two image-text matching benchmarks and two video-text matching benchmarks validate the efficacy of the proposed method.
Jiwei Wei, Xing Xu 0001, Yang Yang 0002, Yanli Ji, Zheng Wang 0044, Heng Tao Shen
CVPR5
2020 Fooled by Imagination: Adversarial Attack to Image Captioning Via Perturbation in Complex Domain
abstract
Adversarial attacks are very successful on image classification, but there are few researches on vision-language systems, such as image captioning. In this paper, we study the robustness of a CNN+RNN based image captioning system being subjected to adversarial noises in complex domain. In particular, we propose Fooled-by-Imagination, a novel algorithm for crafting adversarial examples with semantic embedding of targeted caption as perturbation in complex domain. The proposed algorithm explores the great merit of complex values in introducing imaginary part for modeling adversarial perturbation, and maintains the similarity of the image in real part. Our approach provides two evaluation approaches, which check whether neural image captioning systems can be fooled to output some randomly chosen captions or keywords. Besides, our method has good transferability under black-box setting. At last, our extensive experiments show that our algorithm can successfully craft visually-similar adversarial examples with randomly targeted captions or keywords at a higher success rate.
Shaofeng Zhang, Zheng Wang 0044, Xing Xu 0001, Xiang Guan, Yang Yang 0002
ICME2
2020 Ocean: A Dual Learning Approach For Generalized Zero-Shot Sketch-Based Image Retrieval
abstract
Sketch-Based Image Retrieval (SBIR) is an emerging research area with many real-world applications. Recent studies have approached this research task under the more challenging zero-shot learning setting (ZS-SBIR), which assume classes in the target domain are unseen during the training stage. Many of the existing ZS-SBIR studies transferred the learned cross-modal (i.e., sketch and image) representations from the source domain to the target domain by leveraging side information in semantic embeddings. However, these ZS-SBIR methods are not able to generalize well to a more realistic setting to retrieve images from seen and unseen classes. To address the limitation of existing methods, we propose the cOmmon Conditional Encoder Adversarial Network (OCEAN) to perform generalized zero-shot sketch-based image retrieval (GZS-SBIR). The OCEAN model utilizes a dual learning framework to cyclically map the sketch and image features to a common semantic space, and project semantic features back to the relevant visual space by adversarial training. We conduct experiments on two publicly available datasets and demonstrate that our proposed model outperformed the state-of-the-arts baselines in both ZS-SBIR and GZS-SBIR tasks.
Xing Xu 0001, Fumin Shen, Roy Ka-Wei Lee, Zheng Wang 0044, Heng Tao Shen
ICME5
2020 Learning Optimization-based Adversarial Perturbations for Attacking Sequential Recognition Models
abstract
A large number of recent studies on adversarial attack have verified that a Deep Neural Network (DNN) model designed for non-sequential recognition (NSR) tasks (e.g., classification, detection and segmentation) can be easily fooled by adversarial examples. However, only a few researches pay attention to the adversarial attack on sequential recognition (SR). They either apply the attack methods proposed for NSR to SR by neglecting the sequential dependencies, or focus on attacking specific SR models without considering the generality. In this paper, we study the adversarial attack on the general and popular DNN structure of CNN+RNN, i.e., the combination of convolutional neural network (CNN) and recurrent neural network (RNN), which has been widely used in various SR tasks. We take the scene text recognition (STR) and image captioning (IC) as case study, and derive the objective function for attacking the CNN+RNN based models with targeted and untargeted attack modes, and then developed an optimization-based algorithm to learn adversarial perturbations from the derived gradients of each character (or word) in sequence by incorporating the sequential dependencies. Extensive experiments show that our proposed method can effective fool several state-of-the-arts including four STR models and two IC models with higher successful rate and less time consumption, comparing to three latest attack methods.
Xing Xu 0001, Jiefu Chen, Jinhui Xiao, Zheng Wang 0044, Yang Yang 0002, Heng Tao Shen
ACM Multimedia4
2020 Semantic feature augmentation for fine-grained visual categorization with few-sample training
abstract
Small data challenges have emerged in many learning problems, since the success of deep neural networks often relies on the availability of a huge number of labeled data that is expensive to collect. We explore a highly challenging task, few-sample training, which uses a small number of labeled images of each category and corresponding textual descriptions to train a model for fine-grained visual categorization. In order to tackle overfitting caused by small data, in this paper, we propose two novel feature augmentation approaches, Semantic Gate Feature Augmentation (SGFA) and Semantic Boundary Feature Augmentation (SBFA). Instead of generating a new image instance, we propose to directly synthesize instance features by leveraging semantic information, and its main novelties are: (1) The SGFA method is proposed to reduce the overfitting of small data by adding random noise to different regions of the image's feature maps through a gating mechanism. (2) The SBFA approach is proposed to optimize the decision boundary of the classifier. Technically, the decision boundary of the image feature is estimated through the assistance of semantic information, and then feature augmentation is performed by sampling in this region. Experiments in fine-grained visual categorization benchmark demonstrate that our proposed approach can significantly improve the categorization performance.
Xiang Guan, Yang Yang 0002, Zheng Wang 0044, Jingjing Li 0001
MMAsia3
2020 Scene graph generation via multi-relation classification and cross-modal attention coordinator
abstract
Scene graph generation intends to build graph-based representation from images, where nodes and edges respectively represent objects and relationships between them. However, scene graph generation today is heavily limited by imbalanced class prediction. Specifically, most of existing work achieves satisfying performance on simple and frequent relation classes (e.g. on), yet leaving poor performance with fine-grained and infrequent ones (e.g. walk on, stand on). To tackle this problem, in this paper, we redesign the framework as two branches, representation learning branch and classifier learning branch, for a more balanced scene graph generator. Furthermore, for representation learning branch, we propose Cross-modal Attention Coordinator (CAC) to gather consistent features from multi-modal using dynamic attention. For classifier learning branch, we first transfer relation classes' knowledge from large scale corpus, then we leverage Multi-Relationship classifier via Graph Attention neTworks (MR-GAT) to bridge the gap between frequent relations and infrequent ones. The comprehensive experimental results on VG200, a challenge dataset, indicate the competitiveness and the significant superiority of our proposed approach.
Zheng Wang 0044, Xing Xu 0001, Jiwei Wei, Yang Yang 0002
MMAsia2
2020 Discovering attractive segments in the user-generated video streams
Zheng Wang 0044, Jie Zhou 0001, Jing Ma 0004, Jingjing Li 0001, Jiangbo Ai, Yang Yang 0002
Inf. Process. Manag.1
2020 Leveraging unpaired out-of-domain data for image captioning
Xinghan Chen, Zheng Wang 0044, Lin Zuo, Bo Li 0072, Yang Yang 0002
Pattern Recognit. Lett.3
2020 Bidirectional Discrete Matrix Factorization Hashing for Image Search
abstract
Unsupervised image hashing has recently gained significant momentum due to the scarcity of reliable supervision knowledge, such as class labels and pairwise relationship. Previous unsupervised methods heavily rely on constructing sufficiently large affinity matrix for exploring the geometric structure of data. Nevertheless, due to lack of adequately preserving the intrinsic information of original visual data, satisfactory performance can hardly be achieved. In this article, we propose a novel approach, called bidirectional discrete matrix factorization hashing (BDMFH), which alternates two mutually promoted processes of 1) learning binary codes from data and 2) recovering data from the binary codes. In particular, we design the inverse factorization model, which enforces the learned binary codes inheriting intrinsic structure from the original visual data. Moreover, we develop an efficient discrete optimization algorithm for the proposed BDMFH. Comprehensive experimental results on three large-scale benchmark datasets show that the proposed BDMFH not only significantly outperforms the state-of-the-arts but also provides the satisfactory computational efficiency.
Shiyuan He, Bokun Wang, Zheng Wang 0044, Yang Yang 0002, Fumin Shen, Zi Huang, Heng Tao Shen
IEEE Trans. Cybern.3
2020 Exploring nonnegative and low-rank correlation for noise-resistant spectral clustering
Zheng Wang 0044, Lin Zuo, Jing Ma 0004, Jingjing Li 0001, Zhao Kang 0001, Lei Zhang 0038
World Wide Web1
2019 CRA-Net: Composed Relation Attention Network for Visual Question Answering
abstract
The task of Visual Question Answering (VQA) is to answer a natural language question tied to the content of a visual image. Most existing VQA models either apply attention mechanism to locate the relevant object regions and/or utilize the off-the-shelf methods of the relation reasoning to detect object relations. However, they 1) mostly encode the simple relations which cannot sufficiently provide sophisticated knowledge for answering complicated visual questions; 2) seldom leverage the harmony cooperation of the object appearance feature and relation feature. To address these problems, we propose a novel end-to-end VQA model, termed Composed Relation Attention Network (CRA-Net ). In specific, we devise two question-adaptive relation attention modules that can extract not only the fine-grained and precise binary relations but also the more sophisticated trinary relations. Both kinds of question-related relations can reveal deeper semantics, thereby enhancing the reasoning ability in question answering. Furthermore, our CRA-Net also combines the object appearance feature with the relation feature under the guidance of the corresponding question, which can reconcile the two types of features effectively. Extensive experiments on two large benchmark datasets, VQA-1.0 and VQA-2.0, demonstrate that our proposed model outperforms state-of-the-art approaches.
Yang Yang 0002, Zheng Wang 0044, Xiao Wu 0001, Zi Huang
ACM Multimedia3
2019 Multi-scale aggregation network for temporal action proposals
Zheng Wang 0044, Yang Yang 0002
Pattern Recognit. Lett.1