Shibin Mei

dblp:324/0613 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
8since 2021 · last 2025
0009-0005-6315-4642ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021
YearPublicationVenuePosition
2025 GeoMM: On Geodesic Perspective for Multi-modal Learning
abstract
Geodesic distance serves as a reliable means of measuring distance in nonlinear spaces, and such nonlinear manifolds are prevalent in the current multimodal learning. In these scenarios, some samples may exhibit high similarity, yet they convey different semantics, making traditional distance metrics inadequate for distinguishing between positive and negative samples. This paper introduces geodesic distance as a novel distance metric in multi-modal learning for the first time, to mine correlations between samples, aiming to address the limitations of common distance metric. Our approach incorporates a comprehensive series of strategies to adapt geodesic distance for the current multi-modal learning. Specifically, we construct a graph structure to represent the adjacency relationships among samples by thresholding distances between them and then apply the shortest-path algorithm to obtain geodesic distance within this graph. To facilitate efficient computation, we further propose a hierarchical graph structure through clustering and combined with incremental update strategies for dynamic status updates. Extensive experiments across various downstream tasks validate the effectiveness of our proposed method, demonstrating its capability to capture complex relationships between samples and improve the performance of multimodal learning models.
Shibin Mei, Bingbing Ni
CVPR1
2025 PhysDiff-VTON: Cross-Domain Physics Modeling and Trajectory Optimization for Virtual Try-On
abstract
We present PhysDiff-VTON, a diffusion-based framework for image-based virtual try-on that systematically addresses the dual challenges of garment deformation modeling and high-frequency detail preservation. The core innovation lies in integrating physics-inspired mechanisms into the diffusion process: a pose-guided deformable warping module simulates fabric dynamics by predicting spatial offsets conditioned on human pose semantics, while wavelet-enhanced feature decomposition explicitly preserves texture fidelity through frequency-aware attention. Further enhancing generation quality, a novel sampling strategy optimizes the denoising trajectory via least action principles, enforcing temporal coherence, spatial smoothness, and multi-scale structural consistency. Comprehensive evaluations across multiple datasets demonstrate significant improvements in both geometric plausibility and perceptual quality compared to existing approaches. The framework establishes a new paradigm for synthesizing photorealistic try-on images that adhere to physical constraints while maintaining intricate garment details, advancing the practical applicability of diffusion models in fashion technology.
Shibin Mei, Bingbing Ni
NeurIPS1
2024 Object-Oriented Anchoring and Modal Alignment in Multimodal Learning
Shibin Mei, Bingbing Ni, Chenglong Zhao, Fengfa Hu, Zhiming Pi, Bilian Ke
ECCV (50)1
2024 Variational Adversarial Defense: A Bayes Perspective for Adversarial Training
abstract
Various methods have been proposed to defend against adversarial attacks. However, there is a lack of enough theoretical guarantee of the performance, thus leading to two problems: First, deficiency of necessary adversarial training samples might attenuate the normal gradient's back-propagation, which leads to overfitting and gradient masking potentially. Second, point-wise adversarial sampling offers an insufficient support region for adversarial data and thus cannot form a robust decision-boundary. To solve these issues, we provide a theoretical analysis to reveal the relationship between robust accuracy and the complexity of the training set in adversarial training. As a result, we propose a novel training scheme called Variational Adversarial Defense. Based on the distribution of adversarial samples, this novel construction upgrades the defend scheme from local point-wise to distribution-wise, yielding an enlarged support region for safeguarding robust training, thus possessing a higher promising to defense attacks. The proposed method features the following advantages: 1) Instead of seeking adversarial examples point-by-point (in a sequential way), we draw diverse adversarial examples from the inferred distribution; and 2) Augmenting the training set by a larger support region consolidates the smoothness of the decision boundary. Finally, the proposed method is analyzed via the Taylor expansion technique, which casts our solution with natural interpretability.
Chenglong Zhao, Shibin Mei, Bingbing Ni, Shengchao Yuan, Zhenbo Yu, Jun Wang 0159
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Towards Interpreting and Utilizing Symmetry Property in Adversarial Examples
abstract
In this paper, we identify symmetry property in adversarial scenario by viewing adversarial attack in a fine-grained manner. A newly designed metric called attack proportion, is thus proposed to count the proportion of the adversarial examples misclassified between classes. We observe that the distribution of attack proportion is unbalanced as each class shows vulnerability to particular classes. Further, some class pairs correlate strongly and have the same degree of attack proportion for each other. We call this intriguing phenomenon symmetry property. We empirically prove this phenomenon is widespread and then analyze the reason behind the existence of symmetry property. This explanation, to some extent, could be utilized to understand robust models, which also inspires us to strengthen adversarial defenses.
Shibin Mei, Chenglong Zhao, Bingbing Ni, Shengchao Yuan
AAAI1
2023 Exploring and Utilizing Pattern Imbalance
abstract
In this paper, we identify pattern imbalance from several aspects, and further develop a new training scheme to avert pattern preference as well as spurious correlation. In contrast to prior methods which are mostly concerned with category or domain granularity, ignoring the potential finer structure that existed in datasets, we give a new definition of seed category as an appropriate optimization unit to distinguish different patterns in the same category or domain. Extensive experiments on domain generalization datasets of diverse scales demonstrate the effectiveness of the proposed method.
Shibin Mei, Chenglong Zhao, Shengchao Yuan, Bingbing Ni
CVPR1
2022 Towards Bridging Sample Complexity and Model Capacity
abstract
In this paper, we give a new definition for sample complexity, and further develop a theoretical analysis to bridge the gap between sample complexity and model capacity. In contrast to previous works which study on some toy samples, we conduct our analysis on more general data space, and build a qualitative relationship from sample complexity to model capacity required to achieve comparable performance. Besides, we introduce a simple indicator to evaluate the sample complexity based on continuous mapping. Moreover, we further analysis the relationship between sample complexity and data distribution, which paves the way to understand the present representation learning. Extensive experiments on several datasets well demonstrate the effectiveness of our evaluation method.
Shibin Mei, Chenglong Zhao, Shengchao Yuan, Bingbing Ni
AAAI1
2022 Explore Adversarial Attack via Black Box Variational Inference
abstract
From the perspective of probability, we propose a new method for black-box adversarial attack via black-box variational inference (BBVI), where the knowledge of victim model is unavailable. Instead of obtaining a single point, the proposed method focuses on approximating the probability distribution of adversarial examples. Thus, infinite adversarial examples can be drawn from the inferred distribution. Although the Monte Carlo estimator in BBVI is unbiased, its variance brings unstable gradient estimation, which leads to poor attack performance and low query efficiency. To reduce variance, we improve the BBVI with importance sampling which guided by a surrogate model to obtain a better estimator of gradient, which enhances both success rate and query efficiency. Extensive experiments on ImageNet dataset well demonstrate the outperformance of the proposed method compared with prior arts.
Chenglong Zhao, Bingbing Ni, Shibin Mei
IEEE Signal Process. Lett.3