VLDB 2026 Research / reviewers in the wild / expert
Lijun Sheng
dblp:321/3477
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-8240-9736ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Harmonizing class uniformity and separability for transferability estimation
Yuhe Ding, Bo Jiang 0002, Lijun Sheng, Aihua Zheng, Jian Liang 0001 |
Pattern Recognit. | 3 |
| 2025 | Protecting Model Adaptation from Trojans in the Unlabeled DataabstractModel adaptation tackles the distribution shift problem with a pre-trained model instead of raw data, which has become a popular paradigm due to its great privacy protection. Existing methods always assume adapting to a clean target domain, overlooking the security risks of unlabeled samples. This paper for the first time explores the potential trojan attacks on model adaptation launched by well-designed poisoning target data. Concretely, we provide two trigger patterns with two poisoning strategies for different prior knowledge owned by attackers. These attacks achieve a high success rate while maintaining the normal performance on clean samples in the test stage. To defend against such backdoor injection, we propose a plug-and-play method named DiffAdapt, which can be seamlessly integrated with existing adaptation algorithms. Experiments across commonly used benchmarks and adaptation methods demonstrate the effectiveness of DiffAdapt. We hope this work will shed light on the safety of transfer learning with unlabeled data. Lijun Sheng, Jian Liang 0001, Ran He 0001, Zilei Wang, Tieniu Tan |
AAAI | 1 |
| 2025 | R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt TuningabstractVision-language models (VLMs), such as CLIP, have gained significant popularity as foundation models, with numerous fine-tuning methods developed to enhance performance on downstream tasks. However, due to their inherent vulnerability and the common practice of selecting from a limited set of open-source models, VLMs suffer from a higher risk of adversarial attacks than traditional vision models. Existing defense techniques typically rely on adversarial fine-tuning during training, which requires labeled data and lacks of flexibility for downstream tasks. To address these limitations, we propose robust test-time prompt tuning (R-TPT), which mitigates the impact of adversarial attacks during the inference stage. We first reformulate the classic marginal entropy objective by eliminating the term that introduces conflicts under adversarial conditions, retaining only the pointwise entropy minimization. Furthermore, we introduce a plug-and-play reliability-based weighted ensembling strategy, which aggregates useful information from reliable augmented views to strengthen the defense. R-TPT enhances defense against adversarial attacks without requiring labeled training data while offering high flexibility for inference tasks. Extensive experiments on widely used benchmarks with various attacks demonstrate the effectiveness of R-TPT. The code is available in https://github.com/TomSheng21/R-TPT. Lijun Sheng, Jian Liang 0001, Zilei Wang, Ran He 0001 |
CVPR | 1 |
| 2025 | Cooperative Pseudo Labeling for Unsupervised Federated Classification
Kuangpu Guo, Lijun Sheng, Yongcan Yu, Jian Liang 0001, Zilei Wang, Ran He 0001 |
ICCV | 2 |
| 2025 | The Illusion of Progress? A Critical Look at Test-Time Adaptation for Vision-Language ModelsabstractTest-time adaptation (TTA) methods have gained significant attention for enhancing the performance of vision-language models (VLMs) such as CLIP during inference, without requiring additional labeled data. However, current TTA researches generally suffer from major limitations such as duplication of baseline results, limited evaluation metrics, inconsistent experimental settings, and insufficient analysis. These problems hinder fair comparisons between TTA methods and make it difficult to assess their practical strengths and weaknesses. To address these challenges, we introduce TTA-VLM, a comprehensive benchmark for evaluating TTA methods on VLMs. Our benchmark implements 8 episodic TTA and 7 online TTA methods within a unified and reproducible framework, and evaluates them across 15 widely used datasets. Unlike prior studies focused solely on CLIP, we extend the evaluation to SigLIP—a model trained with a Sigmoid loss—and include training-time tuning methods such as CoOp, MaPLe, and TeCoA to assess generality. Beyond classification accuracy, TTA-VLM incorporates various evaluation metrics, including robustness, calibration, out-of-distribution detection, and stability, enabling a more holistic assessment of TTA methods. Through extensive experiments, we find that 1) existing TTA methods produce limited gains compared to the previous pioneering work; 2) current TTA methods exhibit poor collaboration with training-time fine-tuning methods; 3) accuracy gains frequently come at the cost of reduced model trustworthiness. We release TTA-VLM to provide fair comparison and comprehensive evaluation of TTA methods for VLMs, and we hope it encourages the community to develop more reliable and generalizable TTA strategies. The code is available in https://github.com/TomSheng21/tta-vlm. Lijun Sheng, Jian Liang 0001, Ran He 0001, Zilei Wang, Tieniu Tan |
NeurIPS | 1 |
| 2024 | STAMP: Outlier-Aware Test-Time Adaptation with Stable Memory Replay
Yongcan Yu, Lijun Sheng, Ran He 0001, Jian Liang 0001 |
ECCV (81) | 2 |
| 2024 | A Hard-to-Beat Baseline for Training-free CLIP-based AdaptationabstractContrastive Language-Image Pretraining (CLIP) has gained popularity for its remarkable zero-shot capacity.
Recent research has focused on developing efficient fine-tuning methods, such as prompt learning and adapter, to enhance CLIP's performance in downstream tasks.
However, these methods still require additional training time and computational resources, which is undesirable for devices with limited resources.
In this paper, we revisit a classical algorithm, Gaussian Discriminant Analysis (GDA), and apply it to the downstream classification of CLIP.
Typically, GDA assumes that features of each class follow Gaussian distributions with identical covariance.
By leveraging Bayes' formula, the classifier can be expressed in terms of the class means and covariance, which can be estimated from the data without the need for training.
To integrate knowledge from both visual and textual modalities, we ensemble it with the original zero-shot classifier within CLIP.
Extensive results on 17 datasets validate that our method surpasses or achieves comparable results with state-of-the-art methods on few-shot classification, imbalanced learning, and out-of-distribution generalization.
In addition, we extend our method to base-to-new generalization and unsupervised learning, once again demonstrating its superiority over competing approaches.
Our code is publicly available at https://github.com/mrflogs/ICLR24. Zhengbo Wang, Jian Liang 0001, Lijun Sheng, Ran He 0001, Zilei Wang, Tieniu Tan |
ICLR | 3 |
| 2024 | Realistic Unsupervised CLIP Fine-tuning with Universal Entropy OptimizationabstractThe emergence of vision-language models, such as CLIP, has spurred a significant research effort towards their application for downstream supervised learning tasks. Although some previous studies have explored the unsupervised fine-tuning of CLIP, they often rely on prior knowledge in the form of class names associated with ground truth labels. This paper explores a realistic unsupervised fine-tuning scenario, considering the presence of out-of-distribution samples from unknown classes within the unlabeled data. In particular, we focus on simultaneously enhancing out-of-distribution detection and the recognition of instances associated with known classes. To tackle this problem, we present a simple, efficient, and effective approach called Universal Entropy Optimization (UEO). UEO leverages sample-level confidence to approximately minimize the conditional entropy of confident instances and maximize the marginal entropy of less confident instances. Apart from optimizing the textual prompt, UEO incorporates optimization of channel-wise affine transformations within the visual branch of CLIP. Extensive experiments across 15 domains and 4 different types of prior knowledge validate the effectiveness of UEO compared to baseline methods. The code is at https://github.com/tim-learn/UEO. Jian Liang 0001, Lijun Sheng, Zhengbo Wang, Ran He 0001, Tieniu Tan |
ICML | 2 |
| 2024 | Learning Spatiotemporal Inconsistency via Thumbnail Layout for Face Deepfake Detection
Jian Liang 0001, Lijun Sheng, Xiaoyu Zhang 0002 |
Int. J. Comput. Vis. | 3 |
| 2023 | ProxyMix: Proxy-based Mixup training with label refinery for source-free domain adaptation
Yuhe Ding, Lijun Sheng, Jian Liang 0001, Aihua Zheng, Ran He 0001 |
Neural Networks | 2 |