Yanshu Li

dblp:244/6489 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning
abstract
Modern large vision-language models (LVLMs) convert each input image into a large set of tokens that far outnumber the text tokens. Although this improves visual perception, it also introduces severe image token redundancy. Because image tokens contain sparse information, many contribute little to reasoning but greatly increase inference cost. Recent image token pruning methods address this issue by identifying important tokens and removing the rest. These methods improve efficiency with only small performance drops. However, most of them focus on single-image tasks and overlook multimodal in-context learning (ICL), where redundancy is higher and efficiency is more important. Redundant tokens weaken the advantage of multimodal ICL for rapid domain adaptation and lead to unstable performance. When existing pruning methods are applied in this setting, they cause large accuracy drops, which exposes a clear gap and the need for new approaches. To address this, we propose Contextually Adaptive Token Pruning (CATP), a training-free pruning method designed for multimodal ICL. CATP uses two stages of progressive pruning that fully reflect the complex cross-modal interactions in the input sequence. After removing 77.8% of the image tokens, CATP achieves an average performance gain of 0.6% over the vanilla model on four LVLMs and eight benchmarks, clearly outperforming all baselines. At the same time, it improves efficiency by reducing inference latency by an average of 10.78%. CATP strengthens the practical value of multimodal ICL and lays the foundation for future progress in interleaved image-text settings.
Yanshu Li, Jianjiang Yang, Zhennan Shen, Ligong Han, Haoyan Xu, Ruixiang Tang
AAAI1
2026 Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
abstract
Multimodal in-context learning (ICL) is becoming a key capability that allows large vision-language models (LVLMs) to adapt to novel tasks without parameter updates, which expands their usefulness in many real-world applications. However, ICL performance remains unstable even when the in-context demonstrations (ICDs) are well matched, showing that LVLMs still struggle to make full use of the provided context. While existing work mainly focuses on prompt engineering or post-hoc logit calibration, we study the attention mechanisms inside LVLMs to address their inherent limitations. We identify two important weaknesses in their self-attention that hinder effective ICL. To address these weaknesses, we propose Context-Aware Modulated Attention (CAMA), a training-free and plug-and-play method that dynamically adjusts attention logits based on the input in-context sequence. CAMA uses a two-stage modulation process that strengthens attention to semantically important tokens, especially visual ones. Across four LVLMs and seven benchmarks, CAMA consistently outperforms vanilla models and baselines, showing clear effectiveness and generalization. It can also activate the intended benefits of prompt engineering methods and remains robust across different sequence configurations. Therefore, CAMA opens up new directions for improving multimodal reasoning through a deeper understanding of attention dynamics.
Yanshu Li, Jianjiang Yang, Ziteng Yang, Bozheng Li, Ligong Han, Hongyang He, Zhengtao Yao, Victor Y. Chen, Songlin Fei, Dongfang Liu, Ruixiang Tang
AAAI1
2026 MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance Detection
abstract
Multimodal Stance Detection (MSD) is crucial for understanding public discourse, yet effectively fusing text and image, especially with conflicting signals, remains challenging.Existing methods often face difficulties with contextual grounding, cross-modal interpretation ambiguity, and single-pass reasoning fragility.To address these, we propose Retrieval-Augmented Multi-modal Multiagent Stance Detection (MM-StanceDet), a novel multi-agent framework integrating Retrieval Augmentation for contextual grounding, specialized Multimodal Analysis agents for nuanced interpretation, a Reasoning-Enhanced Debate stage for exploring perspectives, and Self-Reflection for robust adjudication.Extensive experiments on five datasets demonstrate MM-StanceDet significantly outperforms state-of-the-art baselines, validating the efficacy of its multi-agent architecture and structured reasoning stages in addressing complex multimodal stance challenges.
Weihai Lu, Zhejun Zhao, Yanshu Li
ACL (1)3
2026 Not All Directions Matter: Towards Structured and Task-Aware Low-Rank Model Adaptation
abstract
Xi Xiao, Chenrui Ma, Yunbei Zhang, Chen Liu, Zhuxuanzi Wang, Yanshu Li, Lin Zhao, Guosheng Hu, Tianyang Wang, Hao Xu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xi Xiao 0003, Chenrui Ma, Yunbei Zhang, Chen Liu 0020, Zhuxuanzi Wang, Yanshu Li, Guosheng Hu, Tianyang Wang 0004
ACL (1)6
2026 Self-Supervised Visual Prompting for Cross-Domain Road Damage Detection
abstract
The deployment of automated pavement defect detection is often hindered by poor cross-domain generalization. Supervised detectors achieve strong in-domain accuracy but require costly re-annotation for new environments, while standard self-supervised methods capture generic features and remain vulnerable to domain shift. We propose PROBE, a self-supervised framework that visually probes target domains without labels. PROBE introduces a Self-supervised Prompt Enhancement Module (SPEM), which derives defect-aware prompts from unlabeled target data to guide a frozen ViT backbone, and a Domain-Aware Prompt Alignment (DAPA) objective, which aligns prompt-conditioned source and target representations. Experiments on four challenging benchmarks show that PROBE consistently outperforms strong supervised, self-supervised, and adaptation baselines, achieving robust zero-shot transfer, improved resilience to domain variations, and high data efficiency in few-shot adaptation. These results highlight self-supervised prompting as a practical direction for building scalable and adaptive visual inspection systems. Source code is publicly available: https://github.com/xixiaouab/PROBE/tree/main
Xi Xiao 0003, Zhuxuanzi Wang, Mingqiao Mo, Chen Liu 0020, Chenrui Ma, Yanshu Li, Smita Krishnaswamy, Xiao Wang 0004, Tianyang Wang 0004
WACV6
2026 RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage Understanding
abstract
Accurate road damage detection is crucial for timely infrastructure maintenance and public safety, but existing vision-only datasets and models lack the rich contextual understanding that textual information can provide. To address this limitation, we introduce RoadBench, the first multimodal benchmark for comprehensive road damage understanding. This dataset pairs high-resolution images of road damages with detailed textual descriptions, providing a richer context for model training. We also present RoadCLIP, a novel vision-language model that builds upon CLIP by integrating domain-specific enhancements. It includes a disease-aware positional encoding that captures spatial patterns of road defects and a mechanism for injecting road-condition priors to refine the model’s understanding of road damages. We further employ a GPT-driven data generation pipeline to expand the image–text pairs in Road-Bench, greatly increasing data diversity without exhaustive manual annotation. Experiments demonstrate that Road-CLIP achieves state-of-the-art performance on road damage recognition tasks, significantly outperforming existing vision-only models by 19.2%. These results highlight the advantages of integrating visual and textual information for enhanced road condition analysis, setting new benchmarks for the field and paving the way for more effective infrastructure monitoring through multimodal learning.
Xi Xiao 0003, Yunbei Zhang, Janet Wang, Yuxiang Wei 0004, Hengjia Li, Yanshu Li, Xiao Wang 0004, Swalpa Kumar Roy, Tianyang Wang 0004
WACV7
2025 A Jailbreak Prompt Detector Based on Selective Perturbation and Contrastive Learning
abstract
Jailbreak attacks pose a significant threat to the reliable deployment of large language models (LLMs) in critical applications. Although existing LLMs are supervised fine-tuning and aligned through reinforcement learning from human feedback, automated jailbreak attack algorithms can still identify potential jailbreak prompts that lead to harmful outputs. In this paper, we propose JPS, a jailbreak attack detector based on selective perturbation and contrastive learning. JPS leverages the robustness of jailbreak attacks, which is achieved through complex multi-step optimization, by using perturbation methods to enhance the training data for jailbreak prompts. In order to mitigate noise from perturbation, we introduce a selective strategy based on token importance. Additionally, we employ supervised contrastive learning to effectively differentiate between jailbreak and benign samples. Extensive experiments on the popular jailbreak attacks and benign datasets show that JPS outperforms all the baseline approaches according to F1-score. Furthermore, detailed ablation experiments were conducted to analyze each module of the model, demonstrating the effectiveness of our approach.
Yanshu Li, Yan Wang 0081, Haitian Yang
CSCWD1
2025 TACO: Enhancing Multimodal In-context Learning via Task Mapping-Guided Sequence Configuration
abstract
Multimodal in-context learning (ICL) has emerged as a key mechanism for harnessing the capabilities of large vision-language models (LVLMs).However, its effectiveness remains highly sensitive to the quality of input ICL sequences, particularly for tasks involving complex reasoning or open-ended generation.A major limitation is our limited understanding of how LVLMs actually exploit these sequences during inference.To bridge this gap, we systematically interpret multimodal ICL through the lens of task mapping, which reveals how local and global relationships within and among demonstrations guide model reasoning.Building on this insight, we present TACO, a lightweight transformer-based model equipped with task-aware attention that dynamically configures ICL sequences.By injecting task-mapping signals into the autoregressive decoding process, TACO creates a bidirectional synergy between sequence construction and task reasoning.Experiments on five LVLMs and nine datasets demonstrate that TACO consistently surpasses baselines across diverse ICL tasks.These results position task mapping as a novel and valuable perspective for interpreting and improving multimodal ICL.
Yanshu Li, Jianjiang Yang, Tian Yun 0001, Pinyuan Feng, Jinfa Huang, Ruixiang Tang
EMNLP1
2025 M-ABSA: A Multilingual Dataset for Aspect-Based Sentiment Analysis
abstract
ChengYan Wu, Bolei Ma, Yihong Liu, Zheyu Zhang, Ningyuan Deng, Yanshu Li, Baolan Chen, Yi Zhang, Yun Xue, Barbara Plank. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
ChengYan Wu, Bolei Ma, Yihong Liu 0001, Zheyu Zhang 0007, Ningyuan Deng, Yanshu Li, Baolan Chen, Barbara Plank
EMNLP6
2025 Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram Generation
abstract
Scientific diagrams are vital tools for communicating structured knowledge across disciplines. However, they are often published as static raster images, losing symbolic semantics and limiting reuse. While Multimodal Large Language Models (MLLMs) offer a pathway to bridging vision and structure, existing methods lack semantic control and structural interpretability, especially on complex diagrams. We propose Draw with Thought (DwT), a training-free framework that guides MLLMs to reconstruct diagrams into editable mxGraph XML code through cognitively inspired Chain-of-Thought reasoning. DwT enables interpretable and controllable outputs without model fine-tuning by dividing the task into two stages: Coarse-to-Fine Planning, which handles perceptual structuring and semantic specification, and Structure-Aware Code Generation, enhanced by format-guided refinement. To support evaluation, we release Plot2XML, a benchmark of 247 real-world scientific diagrams with gold-standard XML annotations. Extensive experiments across eight MLLMs show that our approach yields high-fidelity, semantically aligned, and structurally valid reconstructions, with human evaluations confirming strong alignment in both accuracy and visual aesthetics, offering a scalable solution for converting static visuals into structurally valid and renderable representations and advancing machine understanding of scientific graphics.
Zhiqing Cui, Yanshu Li, Chenxu Du, Zhenglong Ding
ACM Multimedia4
2025 TRiCo: Triadic Game-Theoretic Co-Training for Robust Semi-Supervised Learning
abstract
We introduce TRiCo, a novel triadic game-theoretic co-training framework that rethinks the structure of semi-supervised learning by incorporating a teacher, two students, and an adversarial generator into a unified training paradigm. Unlike existing co-training or teacher-student approaches, TRiCo formulates SSL as a structured interaction among three roles: (i) two student classifiers trained on frozen, complementary representations, (ii) a meta-learned teacher that adaptively regulates pseudo-label selection and loss balancing via validation-based feedback, and (iii) a non-parametric generator that perturbs embeddings to uncover decision boundary weaknesses. Pseudo-labels are selected based on mutual information rather than confidence, providing a more robust measure of epistemic uncertainty. This triadic interaction is formalized as a Stackelberg game, where the teacher leads strategy optimization and students follow under adversarial perturbations. By addressing key limitations in existing SSL frameworks—such as static view interactions, unreliable pseudo-labels, and lack of hard sample modeling—TRiCo provides a principled and generalizable solution. Extensive experiments on CIFAR-10, SVHN, STL-10, and ImageNet demonstrate that TRiCo consistently achieves state-of-the-art performance in low-label regimes, while remaining architecture-agnostic and compatible with frozen vision backbones.
Hongyang He, Xinyuan Song 0002, Yangfan He, Yanshu Li, Haochen You, Lifan Sun, Wenqiao Zhang
NeurIPS5
2024 ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLMs
Zhaochen Su, Jun Zhang 0069, Xiaoye Qu, Tong Zhu 0002, Yanshu Li, Jiashuo Sun, Juntao Li 0005, Min Zhang 0005, Yu Cheng 0001
NeurIPS5
2023 ASGNet: Adaptive Semantic Gate Networks for Log-Based Anomaly Diagnosis
Haitian Yang, Degang Sun, Yanshu Li, Yan Wang 0081, Weiqing Huang
ICONIP (4)4
2022 Exploring real-time fault detection of high-speed train traction motor based on machine learning and wavelet analysis
Yanshu Li
Neural Comput. Appl.1
2019 Analysis of disease organ as a novel phenotype towards disease genetics understanding
Lingyun Luo, Chunlei Zheng, Jiaolong Wang, Minsheng Tan, Yanshu Li
J. Biomed. Informatics5