Yuqi Lin

dblp:117/7752 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Vision and language · 30% Generative modeling · 22% Segmentation and scene understanding · 16%
Network and information security
2 papers
Security and privacy of machine learning · 75% Network security · 25%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%

Topics — the 29 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.012026
Forecast Then Calibrate: Feature Caching as ODE for Efficient Diffusion Transformers · AAAI 2026
Machine learning › Generative modeling › diffusion model
diffusion transformer
1.012026
Forecast Then Calibrate: Feature Caching as ODE for Efficient Diffusion Transformers · AAAI 2026
Machine learning › Efficient and distributed learning › inference acceleration
feature caching
1.012026
Forecast Then Calibrate: Feature Caching as ODE for Efficient Diffusion Transformers · AAAI 2026
Machine learning › Efficient and distributed learning
inference acceleration
1.012026
Forecast Then Calibrate: Feature Caching as ODE for Efficient Diffusion Transformers · AAAI 2026
Machine learning › Efficient and distributed learning
model acceleration
1.012026
Forecast Then Calibrate: Feature Caching as ODE for Efficient Diffusion Transformers · AAAI 2026
Computer vision › Segmentation and scene understanding › semantic segmentation
weakly supervised semantic segmentation
0.922024
CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation · CVPR 2023
TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP without Training · AAAI 2024
Computer vision › Vision and language › vision-language generation
interleaved image-text generation
0.912025
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation · CVPR 2025
Computer vision › Segmentation and scene understanding › semantic segmentation › segmentation refinement
mask refinement
0.912025
SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement · ICLR 2025
Computer vision › Vision and language
multimodal evaluation
0.912025
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation · CVPR 2025
Machine learning › Generative modeling
multimodal generation
0.912025
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation · CVPR 2025
Computer vision › Segmentation and scene understanding › prompt-based segmentation
segment anything model adaptation
0.912025
SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement · ICLR 2025
Computer vision › Vision and language › vision-language model
CLIP
0.812024
TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP without Training · AAAI 2024
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.812024
Few-shot Hybrid Domain Adaptation of Image Generator · ICLR 2024
Machine learning › Generative modeling
generative adversarial network
0.812024
Few-shot Hybrid Domain Adaptation of Image Generator · ICLR 2024
Machine learning › Transfer learning and domain adaptation › domain adaptation
hybrid domain adaptation
0.812024
Few-shot Hybrid Domain Adaptation of Image Generator · ICLR 2024
Machine learning › Learning paradigms
multi-label classification
0.812024
TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP without Training · AAAI 2024
Computer vision › Vision and language
multimodal benchmark
0.812024
MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI · ICML 2024
Computer vision › Vision and language › vision-language model
multimodal large language model
0.812024
MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI · ICML 2024
Machine learning › Learning paradigms › multi-label classification
open-vocabulary multi-label classification
0.812024
TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP without Training · AAAI 2024
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.812024
Position: Towards Implicit Prompt For Text-To-Image Models · ICML 2024
Computer vision › Vision and language
vision-language model
0.812024
TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP without Training · AAAI 2024
Computer vision › Vision and language › vision-language model
vision-language model evaluation
0.812024
ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Ablation Capability for Large Vision-Language Models · NeurIPS 2024
Security and privacy of machine learning › adversarial attack › large language model attack
adversarial prompts
0.812024
Position: Towards Implicit Prompt For Text-To-Image Models · ICML 2024
Performance modeling and evaluation
benchmarking
0.812024
MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI · ICML 2024
Computer vision › Segmentation and scene understanding
semantic segmentation
0.712023
CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation · CVPR 2023
Computer vision › Vision and language
vision-language pretraining
0.712023
CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation · CVPR 2023
Network security
covert channel
0.212016
Designing and Modeling of Covert Channels in Operating Systems · IEEE Trans. Computers 2016
Natural language and speech › Language models and text generation › large language model evaluation
automatic evaluation
0.212024
ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Ablation Capability for Large Vision-Language Models · NeurIPS 2024
Machine learning › Trustworthy machine learning
safety evaluation
0.212024
Position: Towards Implicit Prompt For Text-To-Image Models · ICML 2024

Methods — techniques the papers use, named apart from their topics

prompt engineering · 1.5ordinary differential equation modeling · 1.0feature forecasting · 1.0calibration · 1.0split-then-merge · 0.9judge model · 0.9iou adaption · 0.9human annotation · 0.9task map · 0.8patch-level classification · 0.8benchmarking · 0.8attention refinement · 0.8high-level petri nets · 0.5finite state machine · 0.5
YearPublicationVenuePosition
2026 Forecast Then Calibrate: Feature Caching as ODE for Efficient Diffusion Transformers
abstract
Diffusion Transformers (DiTs) have demonstrated exceptional performance in high-fidelity image and video generation. To reduce their substantial computational costs, feature caching techniques have been proposed to accelerate inference by reusing hidden representations from previous timesteps. However, current methods often struggle to maintain generation quality at high acceleration ratios, where prediction errors increase sharply due to the inherent instability of long-step forecasting. In this work, we adopt an ordinary differential equation (ODE) perspective on the hidden-feature sequence, modeling layer representations along the trajectory as a feature-ODE. We attribute the degradation of existing caching strategies to their inability to robustly integrate historical features under large skipping intervals. To address this, we propose FoCa (Forecast-then-Calibrate), which treats feature caching as a feature-ODE solving problem. Extensive experiments on image, video generation, and super-resolution tasks demonstrate the effectiveness of FoCa, especially under aggressive acceleration. Without additional training, FoCa achieves near-lossless speedups of 5.50× on FLUX, 6.45× on HunyuanVideo, 3.17× on Inf-DiT, and maintains high quality with a 4.53× speedup on DiT.
Shikang Zheng, Qinming Zhou, Peiliang Cai, Chang Zou, Yuqi Lin, Linfeng Zhang 0001
AAAI8
2025 OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
abstract
Multimodal Large Language Models (MLLMs) have made significant strides in visual understanding and generation tasks. However, generating interleaved image-text content remains a challenge, which requires integrated multimodal understanding and generation abilities. While the progress in unified models offers new solutions, existing benchmarks are insufficient for evaluating these methods due to data size and diversity limitations. To bridge this gap, we introduce OpenING, a comprehensive benchmark comprising 5,400 high-quality human-annotated instances across 56 real-world tasks. OpenING covers diverse daily scenarios such as travel guide, design, and brainstorming, offering a robust platform for challenging interleaved generation methods. In addition, we present IntJudge, a judge model for evaluating open-ended multimodal generation methods. Trained with a novel data pipeline, our IntJudge achieves an agreement rate of 82.42% with human judgments, outperforming GPT-based evaluators by 11.34%. Extensive experiments on OpenING reveal that current interleaved generation methods still have substantial room for improvement. Key findings on interleaved image-text generation are further presented to guide the development of next-generation models.
Xiaopeng Peng 0001, Jiajun Song, Chuanhao Li 0001, Zhaopan Xu, Ziyao Guo, Hao Zhang 0117, Yuqi Lin, Yefei He, Lirui Zhao, Xiaojun Chang, Yu Qiao 0001, Wenqi Shao, Kaipeng Zhang
CVPR9
2025 SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement
abstract
In this paper, we explore a principal way to enhance the quality of widely pre-existing coarse masks, enabling them to serve as reliable training data for segmentation models to reduce the annotation cost. In contrast to prior refinement techniques that are tailored to specific models or tasks in a close-world manner, we propose SAMRefiner, a universal and efficient approach by adapting SAM to the mask refinement task. The core technique of our model is the noise-tolerant prompting scheme. Specifically, we introduce a multi-prompt excavation strategy to mine diverse input prompts for SAM (\ie, distance-guided points, context-aware elastic bounding boxes, and Gaussian-style masks) from initial coarse masks. These prompts can collaborate with each other to mitigate the effect of defects in coarse masks. In particular, considering the difficulty of SAM to handle the multi-object case in semantic segmentation, we introduce a split-then-merge (STM) pipeline. Additionally, we extend our method to SAMRefiner++ by introducing an additional IoU adaption step to further boost the performance of the generic SAMRefiner on the target dataset. This step is self-boosted and requires no additional annotation. The proposed framework is versatile and can flexibly cooperate with existing segmentation methods. We evaluate our mask framework on a wide range of benchmarks under different settings, demonstrating better accuracy and efficiency. SAMRefiner holds significant potential to expedite the evolution of refinement tools. Our code is available at https://github.com/linyq2117/SAMRefiner.
Yuqi Lin, Hengjia Li, Wenqi Shao, Zheng Yang 0008, Jun Zhao 0009, Xiaofei He 0001, Ping Luo 0002, Kaipeng Zhang
ICLR1
2024 TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP without Training
abstract
Contrastive Language-Image Pre-training (CLIP) has demonstrated impressive capabilities in open-vocabulary classification. The class token in the image encoder is trained to capture the global features to distinguish different text descriptions supervised by contrastive loss, making it highly effective for single-label classification. However, it shows poor performance on multi-label datasets because the global feature tends to be dominated by the most prominent class and the contrastive nature of softmax operation aggravates it. In this study, we observe that the multi-label classification results heavily rely on discriminative local features but are overlooked by CLIP. As a result, we dissect the preservation of patch-wise spatial information in CLIP and proposed a local-to-global framework to obtain image tags. It comprises three steps: (1) patch-level classification to obtain coarse scores; (2) dual-masking attention refinement (DMAR) module to refine the coarse scores; (3) class-wise reidentification (CWR) module to remedy predictions from a global perspective. This framework is solely based on frozen CLIP and significantly enhances its multi-label classification performance on various benchmarks without dataset-specific training. Besides, to comprehensively assess the quality and practicality of generated tags, we extend their application to the downstream task, i.e., weakly supervised semantic segmentation (WSSS) with generated tags as image-level pseudo labels. Experiments demonstrate that this classify-then-segment paradigm dramatically outperforms other annotation-free segmentation methods and validates the effectiveness of generated tags. Our code is available at https://github.com/linyq2117/TagCLIP.
Yuqi Lin, Minghao Chen 0001, Kaipeng Zhang, Hengjia Li, Zheng Yang 0008, Dongqin Lv, Binbin Lin 0001, Haifeng Liu 0001, Deng Cai 0001
AAAI1
2024 Few-shot Hybrid Domain Adaptation of Image Generator
abstract
Can a pre-trained generator be adapted to the hybrid of multiple target domains and generate images with integrated attributes of them? In this work, we introduce a new task -- Few-shot $\textit{Hybrid Domain Adaptation}$ (HDA). Given a source generator and several target domains, HDA aims to acquire an adapted generator that preserves the integrated attributes of all target domains, without overriding the source domain's characteristics. Compared with $\textit{Domain Adaptation}$ (DA), HDA offers greater flexibility and versatility to adapt generators to more composite and expansive domains. Simultaneously, HDA also presents more challenges than DA as we have access only to images from individual target domains and lack authentic images from the hybrid domain. To address this issue, we introduce a discriminator-free framework that directly encodes different domains' images into well-separable subspaces. To achieve HDA, we propose a novel directional subspace loss comprised of a distance loss and a direction loss. Concretely, the distance loss blends the attributes of all target domains by reducing the distances from generated images to all target subspaces. The direction loss preserves the characteristics from the source domain by guiding the adaptation along the perpendicular to subspaces. Experiments show that our method can obtain numerous domain-specific attributes in a single adapted generator, which surpasses the baseline methods in semantic similarity, image fidelity, and cross-domain consistency.
Hengjia Li, Yang Liu 0212, Linxuan Xia, Yuqi Lin, Wenxiao Wang 0001, Tu Zheng, Zheng Yang 0008, Xiaohui Zhong, Xiaobo Ren, Xiaofei He 0001
ICLR4
2024 Position: Towards Implicit Prompt For Text-To-Image Models
abstract
Recent text-to-image (T2I) models have had great success, and many benchmarks have been proposed to evaluate their performance and safety. However, they only consider explicit prompts while neglecting implicit prompts (hint at a target without explicitly mentioning it). These prompts may get rid of safety constraints and pose potential threats to the applications of these models. This position paper highlights the current state of T2I models toward implicit prompts. We present a benchmark named ImplicitBench and conduct an investigation on the performance and impacts of implicit prompts with popular T2I models. Specifically, we design and collect more than 2,000 implicit prompts of three aspects: General Symbols, Celebrity Privacy, and Not-Safe-For-Work (NSFW) Issues, and evaluate six well-known T2I models’ capabilities under these implicit prompts. Experiment results show that (1) T2I models are able to accurately create various target symbols indicated by implicit prompts; (2) Implicit prompts bring potential risks of privacy leakage for T2I models. (3) Constraints of NSFW in most of the evaluated T2I models can be bypassed with implicit prompts. We call for increased attention to the potential and risks of implicit prompts in the T2I community and further investigation into the capabilities and impacts of implicit prompts, advocating for a balanced approach that harnesses their benefits while mitigating their risks.
Yuqi Lin, Wenqi Shao, Runjian Chen, Hailong Shang, Yu Wang 0002, Yu Qiao 0001, Kaipeng Zhang, Ping Luo 0002
ICML2
2024 MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
abstract
Large Vision-Language Models (LVLMs) show significant strides in general-propose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of multimodal tasks testing rudimentary capabilities, falling short in tracking LVLM development. In this study, we present MMT-Bench, a comprehensive benchmark designed to assess LVLMs across massive multimodal tasks requiring expert knowledge and deliberate visual recognition, localization, and reasoning. MMT-Bench comprises $31,325$ meticulously curated multi-choice visual questions from various multimodal scenarios such as vehicle driving and embodied navigation, covering $32$ core meta-tasks and $162$ subtasks in multimodal understanding. Due to its extensive task coverage, MMT-Bench enables the evaluation of LVLMs using a task map, facilitating the discovery of in- and out-of-domain tasks. Evaluation results involving $20$ publicly available LVLMs such as the proprietary GeminiProVision model, underscore the significant challenges posed by MMT-Bench. We anticipate that MMT-Bench will inspire the community to develop next-generation multimodal foundation models aimed at achieving general-purpose multimodal intelligence.
Kaining Ying, Fanqing Meng, Zhiqian Li, Hao Zhang 0117, Wenbo Zhang 0009, Yuqi Lin, Jiayi Lei, Quanfeng Lu, Runjian Chen, Peng Xu 0035, Renrui Zhang, Haozhe Zhang 0002, Peng Gao 0007, Yali Wang 0001, Yu Qiao 0001, Ping Luo 0002, Kaipeng Zhang, Wenqi Shao
ICML9
2024 ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Ablation Capability for Large Vision-Language Models
abstract
Multi-turn visual conversation is an important ability of real-world AI assistants. However, the related evaluation benchmark is missed. This paper presents ConvBench, a multi-turn conversation benchmark with hierarchical capabilities ablation evaluation for Large Vision-Language Models (LVLMs). ConvBench comprises 577 curated multi-turn conversations, encompassing 215 tasks. These tasks are broad and open-ended, which resemble real-world user behaviors. ConvBench progressively examines the LVLMs' perception, reasoning, and creativity capabilities in each conversation and can decouple these capabilities in evaluations and thus perform reliable error attribution. Besides, considering the diversity of open-ended questions, we introduce an efficient and reliable automatic evaluation framework. Experimental results reveal that ConvBench is a significant challenge for current LVLMs, even for GPT4V, which achieves only a 39.51% score. Besides, we have some insightful findings, such as the weak perception of LVLMs inhibits authentic strengths in reasoning and creation. We believe our design of hierarchical capabilities, decoupling capabilities evaluation, and multi-turn conversation can blaze a new trail in LVLMs evaluation. Code and benchmark are released at https://github.com/shirlyliu64/ConvBench.
Kaining Ying, Hao Zhang 0117, Yuqi Lin, Chuanhao Li 0001, Yu Qiao 0001, Ping Luo 0002, Wenqi Shao, Kaipeng Zhang
NeurIPS5
2023 CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation (WSSS) with image-level labels is a challenging task. Mainstream approaches follow a multi-stage framework and suffer from high training costs. In this paper, we explore the potential of Contrastive Language-Image Pre-training models (CLIP) to localize different categories with only image-level labels and without further training. To efficiently generate high-quality segmentation masks from CLIP, we propose a novel WSSS framework called CLIP-ES. Our framework improves all three stages of WSSS with special designs for CLIP: 1) We introduce the softmax function into GradCAM and exploit the zero-shot ability of CLIP to suppress the confusion caused by non-target classes and backgrounds. Mean-while, to take full advantage of CLIP, we re-explore text inputs under the WSSS setting and customize two text-driven strategies: sharpness-based prompt selection and synonym fusion. 2) To simplify the stage of CAM refinement, we propose a real-time class-aware attention-based affinity (CAA) module based on the inherent multi-head self-attention (MHSA) in CLIP- ViTs. 3) When training the final segmentation model with the masks generated by CLIP, we introduced a confidence-guided loss (CGL) focus on confident regions. Our CLIP-ES achieves SOTA performance on Pascal VOC 2012 and MS COCO 2014 while only taking 10% time of previous methods for the pseudo mask generation. Code is available at https://github.com/linyq2117/CLIP-ES.
Yuqi Lin, Minghao Chen 0001, Wenxiao Wang 0001, Boxi Wu 0001, Binbin Lin 0001, Haifeng Liu 0001, Xiaofei He 0001
CVPR1
2019 An Efficient Approach for Mitigating Covert Storage Channel Attacks in Virtual Machines by the Anti-Detection Criterion
Nasro Min-Allah, Bei Guan, Yuqi Lin, JingZheng Wu, Yongji Wang 0002
J. Comput. Sci. Technol.4
2016 Designing and Modeling of Covert Channels in Operating Systems
abstract
Covert channels are widely considered as a major risk of information leakage in various operating systems, such as desktop, cloud, and mobile systems. The existing works of modeling covert channels have mainly focused on using finite state machines (FSMs) and their transforms to describe the process of covert channel transmission. However, a FSM is rather an abstract model, where information about the shared resource, synchronization, and encoding/decoding cannot be presented in the model, making it difficult for researchers to realize and analyze the covert channels. In this paper, we use the high-level Petri Nets (HLPN) to model the structural and behavioral properties of covert channels. We use the HLPN to model the classic covert channel protocol. Moreover, the results from the analysis of the HLPN model are used to highlight the major shortcomings and interferences in the protocol. Furthermore, we propose two new covert channel models, namely: (a) two channel transmission protocol (TCTP) model and (b) self-adaptive protocol (SAP) model. The TCTP model circumvents the mutual inferences in encoding and synchronization operations; whereas the SAP model uses sleeping time and redundancy check to ensure correct transmission in an environment with strong noise. To demonstrate the correctness and usability of our proposed models in heterogeneous environments, we implement the TCTP and SAP in three different systems: (a) Linux, (b) Xen, and (c) Fiasco.OC. Our implementation also indicates the practicability of the models in heterogeneous, scalable and flexible environments.
Yuqi Lin, Saif Ur Rehman Malik, Kashif Bilal, Qiusong Yang, Yongji Wang 0002, Samee Ullah Khan
IEEE Trans. Computers1
2012 XenPump: A New Method to Mitigate Timing Channel in Cloud Computing
abstract
Cloud computing security has become the focus in information security, where much attention has been drawn to the user privacy leakage. Although isolation and some other security policies have been provided to protect the security of cloud computing, confidential information can be still stolen by timing channels without being detected. In this paper, a new method named XenPump is presented aiming to mitigate the threat of the timing channels by adding latency. XenPump is designed as a module located in hypervisor, monitoring the hypercalls used by the timing channels and adding latencies to lower the threat into an acceptable level. The prototype of XenPump has been implemented in Xen virtualization platform, and the performance is evaluated by the shared memory based timing channel. The experiment results show that XenPump can mitigate the threat of the timing channel by interrupting both the capacity and transmission accuracy. It is believed that after small extension, XenPump can mitigate the incoming timing channels.
JingZheng Wu, Liping Ding, Yuqi Lin, Nasro Min-Allah, Yongji Wang 0002
IEEE CLOUD3