VLDB 2026 Research / reviewers in the wild / expert
Wenxin Ding
dblp:254/8202
· DBLP profile ↗
10ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Security and privacy · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RFG-VTON: Reward function-guided high-fidelity virtual try-on method based on diffusion model
Shufang Zhang, Wenxin Ding |
Neurocomputing | 3 |
| 2026 | DRFusionRec: Enhancing Rationality and Diversity in Garment RecommendationsabstractFashion recommendation is crucial for consumers to express their self-image and personal style. To boost recommendation accuracy and rationality, researchers have explored state-of-the-art (SOTA) methods that incorporate items’ visual and textual information, along with their pairing records. However, these methods fail to fully explore how latent semantic-stylistic correlations between items influence recommendations, and inadequately tackle exposure bias, which results in skewed, restricted recommendations that prioritize popular or frequently observed outfits. To address these limitations, we propose DRFusionRec, a fashion recommender that enhances recommendation rationality and diversity by fully leveraging item semantic-stylistic correlations while mitigating exposure bias. Specifically, to harness rich item correlations, we introduce a novel Multi-Factor Relationship Measurement (MRM) matrix. It integrates semantic and stylistic features by mining the synergistic interaction probabilities across semantically and stylistically adjacent items, capturing the latent compatibility patterns. This matrix is then used to refine item features for richer details. To address homogeneity from exposure bias, we propose an Adaptive Propensity Score (APS) strategy. By dynamically weighting item popularity (direct influence) and popularity of style-similar neighbors (indirect contextual influence), we model exposure confounders to derive item propensity scores. Integrating these scores into item features effectively mitigates the bias. Lastly, MRM and APS synergistically optimize item representations for rationale-based and diverse recommendations. Experimental results validate that the DRFusionRec outperforms SOTA methods in capturing item compatibility, ensuring diverse recommendations, and maintaining reasonable complexity. Wenxin Ding, Xiangdong Huang 0002, Shufang Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Diffusion model-based size variable virtual try-on technology and evaluation method
Shufang Zhang, Hang Qian, Minxue Ni, Wenxin Ding |
Comput. Graph. | 5 |
| 2024 | Understanding Implosion in Text-to-Image Generative ModelsabstractRecent works show that text-to-image generative models are surprisingly vulnerable to a variety of poisoning attacks. Empirical results find that these models can be corrupted by altering associations between individual text prompts and associated visual features. Furthermore, a number of concurrent poisoning attacks can induce "model implosion," where the model becomes unable to produce meaningful images for unpoisoned prompts. These intriguing findings highlight the absence of an intuitive framework to understand poisoning attacks on these models. In this work, we establish the first analytical framework on robustness of image generative models to poisoning attacks, by modeling and analyzing the behavior of the cross-attention mechanism in latent diffusion models. We model cross-attention training as an abstract problem of "supervised graph alignment" and formally quantify the impact of training data by the hardness of alignment, measured by an Alignment Difficulty (AD) metric. The higher the AD, the harder the alignment. We prove that AD increases with the number of individual prompts (or concepts) poisoned. As AD grows, the alignment task becomes increasingly difficult, yielding highly distorted outcomes that frequently map meaningful text prompts to undefined or meaningless visual representations. As a result, the generative model implodes and outputs random, incoherent images at large. We validate our analytical framework through extensive experiments, and we confirm and explain the unexpected (and unexplained) effect of model implosion while producing new, unforeseen insights. Our work provides a useful tool for studying poisoning attacks against diffusion models and their defenses. Wenxin Ding, Cathy Yuanchen Li, Shawn Shan, Ben Y. Zhao, Hai-Tao Zheng 0002 |
CCS | 1 |
| 2024 | Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative ModelsabstractTrained on billions of images, diffusion-based text-to-image models seem impervious to traditional data poisoning attacks, which typically require poison samples approaching 20% of the training set. In this paper, we show that state-of-the-art text-to-image generative models are in fact highly vulnerable to poisoning attacks. Our work is driven by two key insights. First, while diffusion models are trained on billions of samples, the number of training samples associated with a specific concept or prompt is generally on the order of thousands. This suggests that these models will be vulnerable to prompt-specific poisoning attacks that corrupt a model’s ability to respond to specific targeted prompts. Second, poison samples can be carefully crafted to maximize poison potency to ensure success with very few samples.We introduce Nightshade, a prompt-specific poisoning attack optimized for potency that can completely control the output of a prompt in Stable Diffusion’s newest model (SDXL) with less than 100 poisoned training samples. Nightshade also generates stealthy poison images that look visually identical to their benign counterparts, and produces poison effects that "bleed through" to related concepts. More importantly, a moderate number of Nightshade attacks on independent prompts can destabilize a model and disable its ability to generate images for any and all prompts. Finally, we propose the use of Nightshade and similar tools as a defense for content owners against web scrapers that ignore opt-out/do-not-crawl directives, and discuss potential implications for both model trainers and content owners. Shawn Shan, Wenxin Ding, Josephine Passananti, Stanley Wu, Haitao Zheng 0001, Ben Y. Zhao |
SP | 2 |
| 2024 | A Two-Stage Personalized Virtual Try-On Framework With Shape Control and Texture GuidanceabstractThe Diffusion model has a strong ability to generate wild images. However, the model can just generate inaccurate images with the guidance of text, which makes it very challenging to directly apply the text-guided generative model for virtual try-on scenarios. Taking images as guiding conditions of the diffusion model, this paper proposes a brand new personalized virtual try-on model (PE-VITON), which uses the two stages (shape control and texture guidance) to decouple the clothing attributes. Specifically, the proposed model adaptively matches the clothing to human body parts through the Shape Control Module (SCM) to mitigate the misalignment of the clothing and the human body parts. The semantic information of the input clothing is parsed by the Texture Guided Module (TGM), and the corresponding texture is generated by directional guidance. Therefore, this model can effectively solve the problems of weak reduction of clothing folds, poor generation effect under complex human posture, blurred edges of clothing, and unclear texture styles in traditional try-on methods. Meanwhile, the model can automatically enhance the generated clothing folds and textures according to the human posture, and improve the authenticity of the virtual try-on. In this paper, qualitative and quantitative experiments are carried out on high-resolution paired and unpaired datasets, the results show that the proposed model outperforms the state-of-the-art model. Shufang Zhang, Minxue Ni, Lei Wang 0293, Wenxin Ding, Yuhong Liu 0003 |
IEEE Trans. Multim. | 5 |
| 2023 | Characterizing the Optimal 0-1 Loss for Multi-class Classification with a Test-time AttackerabstractFinding classifiers robust to adversarial examples is critical for their safe
deployment. Determining the robustness of the best possible classifier under a
given threat model for a fixed data distribution and comparing it to that
achieved by state-of-the-art training methods is thus an important diagnostic
tool. In this paper, we find achievable information-theoretic lower bounds on
robust loss in the presence of a test-time attacker for *multi-class
classifiers on any discrete dataset*. We provide a general framework for finding
the optimal $0-1$ loss that revolves around the construction of a conflict
hypergraph from the data and adversarial constraints. The prohibitive cost of
this formulation in practice leads us to formulate other variants of the attacker-classifier
game that more efficiently determine the range of the optimal loss. Our
valuation shows, for the first time, an analysis of the gap to optimal
robustness for classifiers in the multi-class setting on benchmark datasets. Sihui Dai, Wenxin Ding, Arjun Nitin Bhagoji, Daniel Cullina, Haitao Zheng 0001, Ben Zhao, Prateek Mittal |
NeurIPS | 2 |
| 2022 | Post-breach Recovery: Protection against White-box Adversarial Examples for Leaked DNN ModelsabstractServer breaches are an unfortunate reality on today's Internet. In the context of deep neural network (DNN) models, they are particularly harmful, because a leaked model gives an attacker "white-box'' access to generate adversarial examples, a threat model that has no practical robust defenses. For practitioners who have invested years and millions into proprietary DNNs, e.g. medical imaging, this seems like an inevitable disaster looming on the horizon. Shawn Shan, Wenxin Ding, Emily Wenger, Haitao Zheng 0001, Ben Y. Zhao |
CCS | 2 |
| 2022 | Calibration with Privacy in Peer ReviewabstractThis paper is eligible for the Jack Keil Wolf ISIT Student Paper Award. Reviewers in peer review are often miscalibrated: they may be strict, lenient, extreme, moderate, etc. A number of algorithms have previously been proposed to calibrate reviews. Such attempts of calibration can however leak sensitive information about which reviewer reviewed which paper. In this paper, we identify this problem of calibration with privacy, and provide a foundational building block to address it. Specifically, we present a theoretical study of this problem under a simplified-yet-challenging model involving two reviewers, two papers, and an MAP-computing adversary. Our main results establish the Pareto frontier of the tradeoff between privacy (preventing the adversary from inferring reviewer identity) and utility (accepting better papers), and design explicit computationally-efficient algorithms that we prove are Pareto optimal. Wenxin Ding, Gautam Kamath 0001, Weina Wang 0001, Nihar B. Shah |
ISIT | 1 |
| 2019 | Virtual Reality Video Quality Assessment Based on 3d Convolutional Neural NetworksabstractAs a new medium, Virtual Reality (VR) has attracted widespread attentions and research interests. More and more researchers have built their VR image/video database and devise related algorithms. However, the existing methods of VR video quality assessment are not very effective, and one of the most important reasons is that the database is not suitable. To this end, this paper proposes an efficient VR quality assessment method on self-built database. Firstly, we establish a VR video quality assessment database with subjective scores, and add the projection format to the production of the database. Secondly, the database is proved to be valid by using some traditional image quality assessment metrics. Lastly, we design a 3D convolutional neural network to predict the VR video quality without reference VR video. Meanwhile, taking the pre-processed VR video patches as input, different quality score strategy is applied to get the final score. The experimental results surface that the network we designed has good results, and the performance is improved after the weight calculation combined with the projection format. Wenxin Ding, Zhixiang You |
ICIP | 2 |