Yinqi Li 0001

dblp:244/8144-1 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
4since 2021 · last 2026
0000-0002-4481-0895ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Generative modeling · 37% Representation and self-supervised learning · 33% Trustworthy machine learning · 30%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model
conditional diffusion model
1.012026
DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks · IEEE Trans. Multim. 2026
Machine learning › Generative modeling
diffusion model
1.012026
DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks · IEEE Trans. Multim. 2026
Machine learning › Trustworthy machine learning › robustness
corruption robustness
0.912025
PIT: A Plug-and-Play Image Translator for Making Off-the-Shelf Models Adapt to Corruptions · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
PIT: A Plug-and-Play Image Translator for Making Off-the-Shelf Models Adapt to Corruptions · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Visual content generation and editing
image-to-image translation
0.912025
PIT: A Plug-and-Play Image Translator for Making Off-the-Shelf Models Adapt to Corruptions · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.612022
Optimal Positive Generation via Latent Transformation for Contrastive Learning · NeurIPS 2022
Machine learning › Representation and self-supervised learning › latent space › latent space manipulation
latent space translation
0.612022
Optimal Positive Generation via Latent Transformation for Contrastive Learning · NeurIPS 2022
Machine learning › Representation and self-supervised learning › contrastive learning
positive pair construction
0.612022
Optimal Positive Generation via Latent Transformation for Contrastive Learning · NeurIPS 2022
Machine learning › Generative modeling › trustworthy generative modeling
semantic consistency
0.212022
Optimal Positive Generation via Latent Transformation for Contrastive Learning · NeurIPS 2022
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning
0.212022
Optimal Positive Generation via Latent Transformation for Contrastive Learning · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

unsupervised image translation · 1.7data augmentation · 1.7diffusion model inversion · 1.0mutual information minimization · 0.6latent transformation · 0.6generative model · 0.6
YearPublicationVenuePosition
2026 DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks
Yinqi Li 0001, Hong Chang 0001, Ruibing Hou, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Multim.1
2025 un2CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP
abstract
Contrastive Language-Image Pre-training (CLIP) has become a foundation model and has been applied to various vision and multimodal tasks. However, recent works indicate that CLIP falls short in distinguishing detailed differences in images and shows suboptimal performance on dense-prediction and vision-centric multimodal tasks. Therefore, this work focuses on improving existing CLIP models, aiming to capture as many visual details in images as possible. We find that a specific type of generative models, unCLIP, provides a suitable framework for achieving our goal. Specifically, unCLIP trains an image generator conditioned on the CLIP image embedding. In other words, it inverts the CLIP image encoder. Compared to discriminative models like CLIP, generative models are better at capturing image details because they are trained to learn the data distribution of images. Additionally, the conditional input space of unCLIP aligns with CLIP's original image-text embedding space. Therefore, we propose to invert unCLIP (dubbed un$^2$CLIP) to improve the CLIP model. In this way, the improved image encoder can gain unCLIP's visual detail capturing ability while preserving its alignment with the original text encoder simultaneously. We evaluate our improved CLIP across various tasks to which CLIP has been applied, including the challenging MMVP-VLM benchmark, the dense-prediction open-vocabulary segmentation task, and multimodal large language model tasks. Experiments show that un$^2$CLIP significantly improves the original CLIP and previous CLIP improvement methods. Code and models are available at https://github.com/LiYinqi/un2CLIP.
Yinqi Li 0001, Jiahe Zhao, Hong Chang 0001, Ruibing Hou, Shiguang Shan, Xilin Chen 0001
NeurIPS1
2025 PIT: A Plug-and-Play Image Translator for Making Off-the-Shelf Models Adapt to Corruptions
abstract
Visual recognition models pretrained on clean images usually do not perform well in the presence of image corruptions, such as blurring or noise, which limits their applicability in real-world scenarios. To solve this problem, existing approaches usually design complex data augmentations to train a robust model from scratch or adapt a pretrained model to corrupted scenarios. These approaches ignore the existence of the large number of deployed models in our community, causing extensive computation and storage costs for making deployed models adapted. Based on this consideration, this paper focuses on solving a practical problem of making many clean-image-pretrained models adapt to unlabeled corrupted images through one training procedure. To this end, we aim to learn a Plug-and-play Image Translator (PIT) that can be directly combined with recognition models after training. Existing approaches, such as vanilla image translation and restoration, are not proper for solving this problem, as they are mostly based on supervised training and are not recognition-oriented. To address this issue, we propose a recognition-oriented unsupervised image translation framework to make PIT produce images with indistinguishable recognition predictions from the clean ones. We verify the effectiveness of PIT on several recognition tasks and show that PIT boosts the performance of clean-image-pretrained models significantly in the presence of image corruptions.
Yinqi Li 0001, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Optimal Positive Generation via Latent Transformation for Contrastive Learning
abstract
Contrastive learning, which learns to contrast positive with negative pairs of samples, has been popular for self-supervised visual representation learning. Although great effort has been made to design proper positive pairs through data augmentation, few works attempt to generate optimal positives for each instance. Inspired by semantic consistency and computational advantage in latent space of pretrained generative models, this paper proposes to learn instance-specific latent transformations to generate Contrastive Optimal Positives (COP-Gen) for self-supervised contrastive learning. Specifically, we formulate COP-Gen as an instance-specific latent space navigator which minimizes the mutual information between the generated positive pair subject to the semantic consistency constraint. Theoretically, the learned latent transformation creates optimal positives for contrastive learning, which removes as much nuisance information as possible while preserving the semantics. Empirically, using generated positives by COP-Gen consistently outperforms other latent transformation methods and even real-image-based methods in self-supervised contrastive learning.
Yinqi Li 0001, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001
NeurIPS1
2019 Deep Conditional Variational Estimation for Depth-Based Hand Poses
abstract
We propose a novel and effective approach for 3D hand pose estimation on single depth image. Instead of doing deterministic regression from depth images, our model focuses on learning a latent distribution to model the high dimensional space of pose joints, which can also be interpreted as a kinematics model for human hands. Specifically, the proposed network combines the framework of conditional variational autoencoder which learns an encoder and a decoder with standard convolutional network. The encoder models the latent variable as a prior or a regularization for the pose joints. Then probabilistic inference is performed by the decoder to generate the output prediction conditioned on input depth images. In addition, we introduce a pool-convolution module to improve the localization regression of the network. The architecture can be trained end-to-end. In experiments, we demonstrate the effectiveness of our proposed approach in comparison to various state-of-art holistic regression approaches.
Yinqi Li 0001, Ji'an Tao, Jianru Xue
FG3
2019 A Novel Compression Algorithm for Hardware-Oriented Gradient Boosting Decision Tree Classification Model
Xiao Wang 0019, Yinqi Li 0001
ICIC (3)3