VLDB 2026 Research / reviewers in the wild / expert
Yuhe Ding
dblp:278/3256
· DBLP profile ↗
10ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-7572-3049ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-level alignment network for unsupervised domain adaptive multi-modality object re-identification
Yusong Sheng, Yuhe Ding, Aihua Zheng, Zi Wang 0013, Jin Tang 0001 |
Knowl. Based Syst. | 2 |
| 2026 | Harmonizing class uniformity and separability for transferability estimation
Yuhe Ding, Bo Jiang 0002, Lijun Sheng, Aihua Zheng, Jian Liang 0001 |
Pattern Recognit. | 1 |
| 2026 | Ranking Vision-Language Models in Fully Unlabeled TasksabstractVision language models (VLMs) like CLIP show stellar zero-shot capability on classification benchmarks. However, selecting the VLM with the highest performance on the unlabeled downstream task is non-trivial. Existing VLM selection methods focus on the class-name-only setting, relying on supervised auxiliary datasets and large language models, which may not be accessible or feasible during deployment. This paper introduces the problem ofunsupervised vision-language model selection, where only unsupervised downstream datasets are available, with no additional information provided. To solve this problem, we propose a method termed Visual-tExtual Graph Alignment (VEGA), to select VLMs without any annotations by measuring the alignment of the VLM between the two modalities on the downstream task. VEGA is motivated by the pretraining paradigm of VLMs, which aligns features with the same semantics from the visual and textual modalities, thereby mapping both modalities into a shared representation space. Specifically, we first construct two graphs on the vision and textual features, respectively. VEGA is then defined as the overall similarity between the visual and textual graphs at both node and edge levels. Extensive experiments across three different benchmarks, covering a variety of application scenarios and downstream datasets, demonstrate that VEGA consistently provides reliable and accurate estimates of VLMs' performance on unlabeled downstream tasks. Yuhe Ding, Bo Jiang 0002, Aihua Zheng, Jian Liang 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Exploring Vacant Classes in Label-Skewed Federated LearningabstractLabel skews, characterized by disparities in local label distribution across clients, pose a significant challenge in federated learning. As minority classes suffer from worse accuracy due to overfitting on local imbalanced data, prior methods often incorporate class-balanced learning techniques during local training. Although these methods improve the mean accuracy across all classes, we observe that vacant classes—referring to categories absent from a client's data distribution—remain poorly recognized. Besides, there is still a gap in the accuracy of local models on minority classes compared to the global model. This paper introduces FedVLS, a novel approach to label-skewed federated learning that integrates both vacant-class distillation and logit suppression simultaneously. Specifically, vacant-class distillation leverages knowledge distillation during local training on each client to retain essential information related to vacant classes from the global model. Moreover, logit suppression directly penalizes network logits for non-label classes, effectively addressing misclassifications in minority classes that may be biased toward majority classes. Extensive experiments validate the efficacy of FedVLS, demonstrating superior performance compared to previous state-of-the-art (SOTA) methods across diverse datasets with varying degrees of label skews. Kuangpu Guo, Yuhe Ding, Jian Liang 0001, Zilei Wang, Ran He 0001, Tieniu Tan |
AAAI | 2 |
| 2025 | Visual and text prompt learning for multi-modal brain disease diagnosis
Yumiao Zhao, Bo Jiang 0002, Yuhe Ding, Xixi Wan, Jin Tang 0001 |
Sci. China Inf. Sci. | 3 |
| 2024 | MAPS: A Noise-Robust Progressive Learning Approach for Source-Free Domain Adaptive Keypoint DetectionabstractExisting cross-domain keypoint detection methods always require accessing the source data during adaptation, which may violate the data privacy law and pose serious security concerns. Instead, this paper considers a realistic problem setting called source-free domain adaptive keypoint detection, where only the well-trained source model is provided to the target domain. For the challenging problem, we first construct a teacher-student learning baseline by stabilizing the predictions under data augmentation and network ensembles. Built on this, we further propose a unified approach, Mixup Augmentation and Progressive Selection (MAPS), to fully exploit the noisy pseudo labels of unlabeled target data during training. On the one hand, MAPS regularizes the model to favor simple linear behavior in-between the target samples via self-mixup augmentation, preventing the model from over-fitting to noisy predictions. On the other hand, MAPS employs the self-paced learning paradigm and progressively selects pseudo-labeled samples from ‘easy’ to ‘hard’ into the training process to reduce noise accumulation. Results on four keypoint detection datasets show that MAPS outperforms the baseline and achieves comparable or even better results in comparison to previous non-source-free counterparts. The code is available athttps://github.com/YuheD/MAPS. Yuhe Ding, Jian Liang 0001, Bo Jiang 0002, Aihua Zheng, Ran He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Modify: Model-Driven Face Stylization Without Style ImagesabstractExisting face stylization methods always acquire the presence of the target (style) domain during the translation process, which violates privacy regulations and limits their applicability in real-world systems. To address this issue, we propose a new method called MODel-drIven Face stYlization (MODIFY), which relies on the generative model to bypass the dependence of the target images. Briefly, MODIFY first trains a generative model in the target domain and then translates a source input to the target domain via the provided style model. To preserve the multimodal style information, MODIFY further introduces an additional remapping network, mapping a known continuous distribution into the encoder’s embedding space. During translation in the source domain, MODIFY fine-tunes the encoder module within the target style-persevering model to capture the content of the source input as precisely as possible. Our method is extremely simple and satisfies versatile training modes for face stylization. Experimental results on several different datasets validate the effectiveness of MODIFY for unsupervised face stylization. Code will be released at https://github.com/YuheD/MODIFY. Yuhe Ding, Jian Liang 0001, Jie Cao 0002, Aihua Zheng, Ran He 0001 |
ICASSP | 1 |
| 2023 | Where to Focus: Central Attention-Based Face Forgery Detection
Jinghui Sun, Yuhe Ding, Jie Cao 0002, Junxian Duan, Aihua Zheng |
PRCV (5) | 2 |
| 2023 | ProxyMix: Proxy-based Mixup training with label refinery for source-free domain adaptation
Yuhe Ding, Lijun Sheng, Jian Liang 0001, Aihua Zheng, Ran He 0001 |
Neural Networks | 1 |
| 2020 | Unsupervised Contrastive Photo-to-Caricature Translation based on Auto-distortionabstractPhoto-to-caricature translation aims to synthesize the caricature as a rendered image exaggerating the features through sketching, pencil strokes, or other artistic drawings. Style rendering and geometry deformation are the most important aspects in photo-to-caricature translation task. To take both into consideration, we propose an unsupervised contrastive photo-to-caricature translation architecture. Considering the intuitive artifacts in the existing methods, we propose a contrastive style loss for style rendering to enforce the similarity between the style of rendered photo and the caricature, and simultaneously enhance its discrepancy to the photos. To obtain an exaggerating deformation in an unpaired/unsupervised fashion, we propose a Distortion Prediction Module (DPM) to predict a set of displacements vectors for each input image while fixing some controlling points, followed by the thin plate spline interpolation for warping. The model is trained on unpaired photo and caricature while can offer bidirectional synthesizing via inputting either a photo or a caricature. Extensive experiments demonstrate that the proposed model is effective to generate hand-drawn like caricatures compared with existing competitors. Yuhe Ding, Xin Ma 0031, Mandi Luo, Aihua Zheng, Ran He 0001 |
ICPR | 1 |