Zhengwei Yang 0001

dblp:04/3687-1 · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
21since 2021 · last 2025
0000-0002-8190-1438ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 VEGAS: Towards Visually Explainable and Grounded Artificial Social Intelligence
abstract
Social Intelligence Queries (Social-IQ) serve as the primary multimodal benchmark for evaluating a model’s social intelligence level. While impressive multiple-choice question (MCQ) accuracy is achieved by current solutions, increasing evidence shows that they are largely, and in some cases entirely, dependent on language modality, overlooking visual context. Additionally, the closed-set nature further prevents the exploration of whether and to what extent the reasoning path behind selection is correct. To address these limitations, we propose the Visually Explainable and Grounded Artificial Social Intelligence (VEGAS) model. As a generative multimodal model, VEGAS leverages open-ended answering to provide explainable responses, which enhances the clarity and evaluation of reasoning paths. To enable visually grounded answering, we propose a novel sampling strategy to provide the model with more relevant visual frames. We then enhance the model’s interpretation of these frames through Generalist Instruction Fine-Tuning (GIFT), which aims to: i) learn multimodal language transformations for fundamental emotional social traits, and ii) establish multimodal joint reasoning capabilities. Extensive experiments, comprising modality ablation, open-ended assessments, and supervised MCQ evaluations, consistently show that VEGAS effectively utilizes visual information in reasoning to produce correct and also credible answers. We expect this work to offer a new perspective on Social-IQ and advance the development of human-like social AI.
Hao Li 0093, Hao Fei 0001, Zechao Hu 0003, Zhengwei Yang 0001, Zheng Wang 0007
AAAI4
2025 CCIN: Compositional Conflict Identification and Neutralization for Composed Image Retrieval
abstract
Composed Image Retrieval (CIR) is a multi-modal task that seeks to retrieve target images by harmonizing a reference image with a modified instruction. A key challenge in CIR lies in compositional conflicts between the reference image (e.g., blue, long sleeve) and the modified instruction (e.g., grey, short sleeve). Previous works attempt to mitigate such conflicts through feature-level manipulation, commonly employing learnable masks to obscure conflicting features within the reference image. However, the inherent complexity of feature spaces poses significant challenges in precise conflict neutralization, thereby leading to uncontrollable results. To this end, this paper proposes the Compositional Conflict Identification and Neutralization (CCIN) framework, which sequentially identifies and neutralizes compositional conflicts for effective CIR. Specifically, CCIN comprises two core modules: 1) Compositional Conflict Identification module, which utilizes LLM-based analysis to identify specific conflicting attributes, and 2) Compositional Conflict Neutralization module, which first generates a kept instruction to preserve non-conflicting attributes, then neutralizes conflicts under collaborative guidance of both the kept and modified instructions. Extensive experiments demonstrate the superiority of CCIN over the state-of-the-arts. Code repository: https://github.com/LikaiTian/CCIN.
Likai Tian, Jian Zhao 0006, Zechao Hu 0003, Zhengwei Yang 0001, Hao Li 0093, Lei Jin 0003, Zheng Wang 0007, Xuelong Li 0001
CVPR4
2025 Cross-Category Subjectivity Generalization for Style-Adaptive Sketch Re-ID
Zechao Hu 0003, Zhengwei Yang 0001, Hao Li 0093, Zheng Wang 0007, Yixiong Zou
ICCV2
2025 Beyond Preferences: Enriching User Profiles for Effective E-commerce Recommendations
abstract
Recommender systems are essential in E-Commerce platforms, with recent advancements leveraging users’ historical records to extract multi-interests. However, beyond these records, user profiles contain semantic information that inherently shapes their interests. Existing works mainly overlook that a user’s interests have: group influence, multi-level preference, and time relevance. To this end, a novel Enhanced User Profile-based Multi-interest Model (E-UPMiM) for recommendation is proposed to integrate enhanced user profiles with social relationships to model users’ multi-interests effectively. We propose to extract user preferences with three components: 1) integrated input containing enhanced profiles with social relationships to meet users’ grouping needs; 2) a multi-interest extraction module to obtain complex interest representations; and 3) a time-aware ranking module to adjust the recommendations dynamically. Extensive experiments on three public datasets show that E-UPMiM significantly outperforms state-of-the-art recommendation models. Codes are publicly available at: https://github.com/KevinXu-01/E-UPMiM.
Jingyu Xu 0002, Zhengwei Yang 0001, Zheng Wang 0007
ISCAS2
2025 From Language to Instance: Generative Visual Prompting for Zero-shot Camouflaged Object Detection
abstract
Traditional Camouflaged Object Detection (COD) methods heavily depend on labor-intensive annotated datasets which require extensive manual effort, resulting in limited generalization. While recent studies have combined Multimodal Large Language Models (MLLMs) and Vision Foundation Models (VFMs) to achieve zero-shot COD, their performance is hindered by modality gap between linguistic semantics and fine-grained visual cues, especially in complex camouflage scenarios. In this paper, we propose Language-to-instance generative visual Prompting (LiP), a novel framework that addresses this limitation by transforming text prompts generated by MLLMs into instance-level visual prompts through a text-to-image generative process. Specifically, we introduce a Diffusion-driven Visual Prompt Generation (DVPG) module that leverages Stable Diffusion model to synthesize visual references, enabling robust homogeneous modality matching for COD. Additionally, we introduce Instruction Contrastive Reasoning (ICR) module to enhance the semantic reliability of prompts by suppressing hallucinated concepts during MLLM inference. To the best of our knowledge, LiP is the first framework that utilize text-to-image generative model to construct instance-level visual prompts in COD task. Extensive experiments on four benchmark datasets demonstrate the effectiveness and strong generalization ability of our approach.
Zihou Zhang, Hao Li 0093, Zhengwei Yang 0001, Zechao Hu 0003, Liang Li 0003, Zheng Wang 0007
ACM Multimedia3
2025 Unified Category and Style Generalization for Instance-Level Sketch Retrieval
abstract
Zero-shot instance-level sketch retrieval addresses a practical retrieval scenario in which sketches from unseen categories during training serve as queries to retrieve matching RGB images. The core challenges of this task lie in two aspects: unknown category generalization and subjective style adaptation. Existing methods either focus solely on category generalization or apply simplistic style elimination techniques within a specific category, leading to suboptimal performance when both challenges are present. To this end, we propose the Dual-Attentive Prompt (DAP) method, which unifies category generalization and style adaptation into a single, interpretable framework. Central to DAP is a dual-attentive prompt composer, consisting of two self-attention-based modules. This composer dynamically integrates pre-learned category-specific knowledge with instance-specific prompts that adapt to sketch-specific styles. By cooperating with additional style alignment loss, the proposed method ensures robust generalization of unseen categories while mitigating the impact of subjective style variations. Extensive experimental results demonstrate the state-of-the-art performance of the proposed method. Additionally, some insights are provided into the challenges of traditional training processes when handling multi-style sketches, along with quantitative and qualitative evidence showing how the proposed approach effectively mitigates the negative impact of subjective style variations.
Zechao Hu 0003, Zhengwei Yang 0001, Hao Li 0093, Yixiong Zou, Fengbin Zhu, Zheng Wang 0007
SIGIR2
2025 Clothing Purification with Causality Meets Vision-Language Pretraining Models
Zhengwei Yang 0001, Huilin Zhu, Nan Lei, Basura Fernando, Zheng Wang 0007
Int. J. Comput. Vis.1
2025 Contrastive-Generative-Contrastive: Neutralize Subjectivity in Sketch Re-Identification
abstract
Sketch-based person re-identification (Sketch re-ID) aims to match pedestrian figures in hand-drawn sketches with their corresponding RGB photos. This technique allows for person retrieval or tracking in surveillance systems when the target person’s RGB photo is not available. While previous research predominantly focused on bridging the modality gap between sketches and RGB photos, the influence of the inherent subjectivity in hand-drawn sketches on re-ID performance remains under-explored. This subjectivity, originating from the artist’s unique style, perceptions, and interpretations, introduces inaccuracies in depicting pedestrian appearances, thereby posing additional challenges such as feature distortion and stylistic variation. This paper introduces a Contrastive-Generative-Contrastive (CGC) framework for subjective style-insensitive re-ID. The framework employs a generative model optimized through self-supervision by contrasting positive and negative pairs of pedestrian sketches and RGB photos. In this manner, it simulates an additional artist specializing in transforming original sketches from various subjective styles into uniform ones. Besides, a simple yet effective weighted contrastive learning loss is proposed to further enhance the model’s focus on pedestrian ID-relevant features. Experimental results demonstrate that the proposed method significantly reduces the influence of subjectivity in feature extraction, achieving new state-of-the-art results on benchmark datasets.
Zechao Hu 0003, Zhengwei Yang 0001, Hao Li 0093, Zheng Wang 0007
IEEE Trans. Inf. Forensics Secur.2
2024 Zero-Shot Object Counting with Good Exemplars
Huilin Zhu, Jingling Yuan, Zhengwei Yang 0001, Yu Guo 0008, Zheng Wang 0007, Xian Zhong, Shengfeng He
ECCV (5)3
2024 Expressiveness is Effectiveness: Self-supervised Fashion-aware CLIP for Video-to-Shop Retrieval
Likai Tian, Zhengwei Yang 0001, Zechao Hu 0003, Hao Li 0093, Yifang Yin, Zheng Wang 0007
IJCAI2
2023 Good is Bad: Causality Inspired Cloth-debiasing for Cloth-changing Person Re-identification
abstract
Entangled representation of clothing and identity (ID)-intrinsic clues are potentially concomitant in conventional person Re- IDentification (ReID). Nevertheless, eliminating the negative impact of clothing on ID remains challenging due to the lack of theory and the difficulty of isolating the exact implications. In this paper, a causality-based Auto-Intervention Model, referred to as AIM11Codes will publicly available at https://github.com/BoomShakaY/AIM-CCReID., is first proposed to mitigate clothing bias for robust cloth-changing person ReID (CC-ReID). Specifically, we analyze the effect of clothing on the model inference and adopt a dual-branch model to simulate causal intervention. Progressively, clothing bias is eliminated automatically with model training. AIM is encouraged to learn more discriminative ID clues that are free from clothing bias. Extensive experiments on two standard CC-ReID datasets demonstrate the superiority of the proposed AIM over other state-of-the-art methods.
Zhengwei Yang 0001, Xian Zhong, Zheng Wang 0007
CVPR1
2023 Implicit Attention-Based Cross-Modal Collaborative Learning for Action Recognition
abstract
Human action recognition is an active research topic in recent years. Multiple modalities often convey heterogeneous but potentially complementary action information that single modality does not hold. Some efforts have been resoted to explore cross-modal representation to promote the modeling capability, but with limited improvement due to the simple fusion of different modalities. To this end, we propose an impliCit attention-based Cross-modal Collaborative Learning (C3L) for action recognition. Specifically, we apply a Modality Generalization network with Grayscale enhancement (MGG) to learn specific modality representation and interaction (infrared and RGB). Then, we construct a unified representation space through the Uniform Modality Representation module (UMR), which preserves the modality information while enhancing the overall representation ability. Finally, feature extractors adaptively leverage modality-specific knowledge to realize cross-modal collaborative learning. Extensive experiments conducted on three widely-used public benchmarks InfAR, HMDB51, and UCF101, demonstrate the effectiveness and strength of our proposed method.
Jianghao Zhang, Xian Zhong, Wenxuan Liu 0008, Kui Jiang, Zhengwei Yang 0001, Zheng Wang 0007
ICIP5
2023 Striking a Balance: Unsupervised Cross-Domain Crowd Counting via Knowledge Diffusion
abstract
Supervised crowd counting relies on manual labeling, which is costly and time-consuming. This led to an increased interest in unsupervised methods. However, there is a significant domain gap issue in unsupervised methods, which is manifested by a model trained on one dataset serving dramatic performance drops when being transferred to another. This phenomenon can be attributed to the diverse domain knowledge making it difficult for the unsupervised models to transfer between general (e.g., similar distribution) and domain-specific (e.g., unique density, perspective, illumination, etc.) knowledge, leading to knowledge bias. Existing methods focus on exploring distinguishable relationships and establishing connections between the source and target domains. However, the similar knowledge transfer cannot perfectly simulate the contents of the target domain, leading to the model's inability to generalize to domain-specific knowledge. In this paper, we propose a Self-awareness Knowledge Diffusion method (SaKnD) that leverages the self-knowledge without establishing cross-domain knowledge relationships, which aims to balance the knowledge bias between general and domain-specific knowledge. Specifically, we propose a strategy to evaluate the uncertainty and consistency to define the clueless and informed areas, which determine the location and orientation of knowledge diffusion. These clueless areas serve as domain-specific knowledge that needs to be optimized, and these informed areas serve as general knowledge across domains. Extensive experiments on three standard crowd-counting benchmarks, ShanghaiTech PartA, ShanghaiTech PartB, and UCF_QNRF, show that the proposed SaKnD achieves state-of-the-art performance.
Haiyang Xie, Zhengwei Yang 0001, Huilin Zhu, Zheng Wang 0007
ACM Multimedia2
2023 DAOT: Domain-Agnostically Aligned Optimal Transport for Domain-Adaptive Crowd Counting
abstract
Domain adaptation is commonly employed in crowd counting to bridge the domain gaps between different datasets. However, existing domain adaptation methods tend to focus on inter-dataset differences while overlooking the intra-differences within the same dataset, leading to additional learning ambiguities. These domain-agnostic factors,e.g., density, surveillance perspective, and scale, can cause significant in-domain variations, and the misalignment of these factors across domains can lead to a drop in performance in cross-domain crowd counting. To address this issue, we propose a Domain-agnostically Aligned Optimal Transport (DAOT) strategy that aligns domain-agnostic factors between domains. The DAOT consists of three steps. First, individual-level differences in domain-agnostic factors are measured using structural similarity (SSIM). Second, the optimal transfer (OT) strategy is employed to smooth out these differences and find the optimal domain-to-domain misalignment, with outlier individuals removed via a virtual "dustbin'' column. Third, knowledge is transferred based on the aligned domain-agnostic factors, and the model is retrained for domain adaptation to bridge the gap across domains. We conduct extensive experiments on five standard crowd-counting benchmarks and demonstrate that the proposed method has strong generalizability across diverse datasets. Our code will be available at: https://github.com/HopooLinZ/DAOT/.
Huilin Zhu, Jingling Yuan, Xian Zhong, Zhengwei Yang 0001, Zheng Wang 0007, Shengfeng He
ACM Multimedia4
2023 Win-Win by Competition: Auxiliary-Free Cloth-Changing Person Re-Identification
abstract
Recent person Re-IDentification (ReID) systems have been challenged by changes in personnel clothing, leading to the study of Cloth-Changing person ReID (CC-ReID). Commonly used techniques involve incorporating auxiliary information (e.g., body masks, gait, skeleton, and keypoints) to accurately identify the target pedestrian. However, the effectiveness of these methods heavily relies on the quality of auxiliary information and comes at the cost of additional computational resources, ultimately increasing system complexity. This paper focuses on achieving CC-ReID by effectively leveraging the information concealed within the image. To this end, we introduce an Auxiliary-free Competitive IDentification (ACID) model. It achieves a win-win situation by enriching the identity (ID)-preserving information conveyed by the appearance and structure features while maintaining holistic efficiency. In detail, we build a hierarchical competitive strategy that progressively accumulates meticulous ID cues with discriminating feature extraction at the global, channel, and pixel levels during model inference. After mining the hierarchical discriminative clues for appearance and structure features, these enhanced ID-relevant features are crosswise integrated to reconstruct images for reducing intra-class variations. Finally, by combing with self- and cross-ID penalties, the ACID is trained under a generative adversarial learning framework to effectively minimize the distribution discrepancy between the generated data and real-world data. Experimental results on four public cloth-changing datasets (i.e., PRCC-ReID, VC-Cloth, LTCC-ReID, and Celeb-ReID) demonstrate the proposed ACID can achieve superior performance over state-of-the-art methods. The code is available soon at: https://github.com/BoomShakaY/Win-CCReID.
Zhengwei Yang 0001, Xian Zhong, Zhun Zhong, Hong Liu 0009, Zheng Wang 0007, Shin'ichi Satoh 0001
IEEE Trans. Image Process.1
2023 Visual Exposes You: Pedestrian Trajectory Prediction Meets Visual Intention
abstract
Pedestrian trajectory prediction in multiple scenarios is of immense importance in autonomous driving and disentanglement of human behavior but is limited in catching human intention and initiative. Most previous works tend to predict the trajectory using only 2D coordinates, which generally cause two common problems: a) Overlooking the subjective initiative, including sudden swerve and erratic movement; b) A potential challenge called abnormal collision caused by unlabeled pedestrians on dataset is not being identified and resolved, which would ruin the model prediction. To break those limitations, we introduce visual localization and orientation as Visual Intention Knowledge to help the trajectory prediction, which is learned directly from visual scenarios. It benefits to comprehend human intention and formulates decision-making processes. Moreover, by learning from the visual information and decision-making policy, we construct the Visual Intention Knowledge associated spatio-temporal Transformer (VIKT) to predict human trajectory by combining the intention knowledge with the novel Transformer. Extensive experimental results demonstrate that our VIKT model could achieve competitive performance by the Visual Intention Knowledge through optimizing the model prediction compared with state-of-the-art methods in terms of prediction accuracy on ETH/UCY and SDD benchmarks.
Xian Zhong, Zhengwei Yang 0001, Wenxin Huang, Kui Jiang, Ryan Wen Liu, Zheng Wang 0007
IEEE Trans. Intell. Transp. Syst.3
2022 Attentive Decoupling Network for Cloth-Changing Re-Identification
abstract
Recently, Cloth-Changing person Re-IDentification (CC-ReID) plays a vital role in the public security system and social livelihood, and suffers the problem of considerable intra-class variation. This paper demonstrates that coarse-grained appearance and body shape features are helpful for CC-ReID. We propose an Attentive DeCoupling (ADC) Network for CC-ReID without auxiliary information. The proposed network is built on two core designs. First, a joint identification structure is proposed to retain ID-relevant information at appearance and shape levels. Second, Competitive Attention (CA) is adopted, where the model progressively updates attention to accumulate sound cues for discriminating identity (ID). The proposed decoupling process is continuously improved through constant self-defeating competition of the network. Experimental results on the public cloth-changing dataset show the proposed method's effectiveness and generalizability.
Zhengwei Yang 0001, Xian Zhong, Hong Liu 0009, Zhun Zhong, Zheng Wang 0007
ICME1
2022 Fine-Grained Fragment Diffusion for Cross Domain Crowd Counting
abstract
Deep learning improves the performance of crowd counting, but model migration remains a tricky challenge. Due to the reliance on training data and inherent domain shift, model application to unseen scenarios is tough. To facilitate the problem, this paper proposes a cross-domain Fine-Grained Fragment Diffusion model (FGFD) that explores feature-level fine-grained similarities of crowd distributions between different fragments to bridge the cross-domain gap (content-level coarse-grained dissimilarities). Specifically, we obtain features of fragments in both source and target domains, and then perform the alignment of the crowd distribution across different domains. With the assistance of the diffusion of crowd distribution, it is able to label unseen domain fragments and make source domain close to target domain, which is fed back to the model to reduce the domain discrepancy. By monitoring the distribution alignment, the distribution perception model is updated, then the performance of distribution alignment is improved. During the model inference, the gap between different domains is gradually alleviated. Multiple sets of migration experiments show that the proposed method achieves competitive results with other state-of-the-art domain-transfer methods.
Huilin Zhu, Jingling Yuan, Zhengwei Yang 0001, Xian Zhong, Zheng Wang 0007
ACM Multimedia3
2022 Global Temporal Attention Optimization for Human Trajectory Prediction
abstract
Predicting human trajectory is one of the key knowledge required for autonomous driving and social robots in real scenarios. Recent studies based on Transformer networks have shown a great ability to model social behaviors. As far as we know, global trajectory information has an essential influence on prediction at a certain step. However, these methods only rely on the previous trajectory states/attention but ignore the important following states/attention of the trajectory for each pedestrian, which will generally collapse on some irregular movements (e.g. acceleration, deceleration, and motionless). To solve this issue, we propose a Global Temporal Attention optimization model (GTAO), which activates the utilization of the following states/attention of the trajectory, and jointly and iteratively optimizes the preliminary trajectory prediction through a global temporal attention (GTA) module. To effectively address the decline in the generalizability and abnormal processing of the model, we further introduce global temporal guidance (GTG) module to instruct the GTA to learn the features closer to realistic trajectories. Experimental results on commonly used real-world human trajectory prediction datasets (ETH and UCY) indicate that our GTAO can achieve better performance in terms of prediction accuracy.
Xian Zhong, Zhengwei Yang 0001, Wenxin Huang, Zheng Wang 0007
SMC3
2021 Auxiliary Bi-Level Graph Representation for Cross-Modal Image-Text Retrieval
abstract
Image-text retrieval is one of the most common tasks in multimodal retrieval. It suffers from the problem of information imbalance between modalities, which is so-called modality gap. It remains challenging because prior methods cannot bridge the gap reasonably. With the help of scene graph, we start by designing an auxiliary bi-level graph representation (ABGR) pipeline that can fully mine the potential information and reduce the information redundancy. By doing so, each modality will be represented by lexical word graph that carries the main content of the information. Specifically, we design a graph feature enhancement (GFE) module to embed the graph-structured information in a common subspace while exploring the relationship between lexical words. As a result, a better representation for both image and text can be obtained, which helps us to evaluate the similarity between images and texts more reasonably. Experimental results conducted on two benchmark datasets Flickr30K and MS-COCO demonstrate the effectiveness of our proposed model for cross-modal retrieval task.
Xian Zhong, Zhengwei Yang 0001, Mang Ye, Wenxin Huang, Jingling Yuan, Chia-Wen Lin
ICME2
2021 Subspace Enhancement and Colorization Network for Infrared Video Action Recognition
Xian Zhong, Wenxuan Liu 0008, Zhengwei Yang 0001, Luo Zhong
PRICAI (3)5