Yingyue Li

dblp:219/0673 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0003-3285-1507ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Segmentation and scene understanding · 44% Vision and language · 31% Trustworthy machine learning · 12%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
semantic segmentation
1.922026
WeakTr: Exploring Plain Vision Transformer for Weakly-Supervised Semantic Segmentation · IEEE Trans. Image Process. 2026
WeakCLIP: Adapting CLIP for Weakly-Supervised Semantic Segmentation · Int. J. Comput. Vis. 2025
Computer vision › Segmentation and scene understanding › semantic segmentation
weakly supervised semantic segmentation
1.922026
WeakTr: Exploring Plain Vision Transformer for Weakly-Supervised Semantic Segmentation · IEEE Trans. Image Process. 2026
WeakCLIP: Adapting CLIP for Weakly-Supervised Semantic Segmentation · Int. J. Comput. Vis. 2025
Computer vision › Vision and language
vision-language model
1.722025
WeakCLIP: Adapting CLIP for Weakly-Supervised Semantic Segmentation · Int. J. Comput. Vis. 2025
MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling · ICCV 2025
Machine learning › Trustworthy machine learning › interpretability › visual explanation
class activation map
1.012026
WeakTr: Exploring Plain Vision Transformer for Weakly-Supervised Semantic Segmentation · IEEE Trans. Image Process. 2026
Computer vision › Vision and language › vision-language model › vision-language model adaptation
CLIP adaptation
0.912025
WeakCLIP: Adapting CLIP for Weakly-Supervised Semantic Segmentation · Int. J. Comput. Vis. 2025
Machine learning › Deep learning architectures and training
hybrid architecture
0.912025
MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling · ICCV 2025

Methods — techniques the papers use, named apart from their topics

vision transformer · 1.0multi-head self-attention · 1.0gradient clipping · 1.0weakly supervised learning · 0.9vision-language pretraining · 0.9transformer · 0.9state space model · 0.9CLIP · 0.9
YearPublicationVenuePosition
2026 WeakTr: Exploring Plain Vision Transformer for Weakly-Supervised Semantic Segmentation
abstract
Transformer has been very successful in various computer vision tasks and understanding the working mechanism of transformer is important. As touchstones, weakly-supervised semantic segmentation (WSSS) and class activation map (CAM) are useful tasks for analyzing vision transformers (ViT). Based on the plain ViT pre-trained with ImageNet classification, we find that multi-layer, multi-head self-attention maps can provide rich and diverse information for weakly-supervised semantic segmentation and CAM generation, e.g., different attention heads of ViT focus on different image areas and object categories. Thus we propose a novel method to end-to-end estimate the importance of attention heads, where the self-attention maps are adaptively fused for high-quality CAM results that tend to have more complete objects. Besides, we propose a ViT-based gradient clipping decoder for online retraining with the CAM results efficiently and effectively. Furthermore, the gradient clipping decoder can make good use of the knowledge in large-scale pre-trained ViT and has a scalable ability. The proposed plain Transformer-based Weakly-supervised learning method (WeakTr) obtains the superior WSSS performance on standard benchmarks, i.e., 78.5% mIoU on the $val$ set of PASCAL VOC 2012 and 51.1% mIoU on the $val$ set of COCO 2014. Source code and checkpoints are available at https://github.com/hustvl/WeakTr.
Lianghui Zhu, Yingyue Li, Jiemin Fang, Yan Liu 0069, Xin Hao, Wenyu Liu 0001, Xinggang Wang
IEEE Trans. Image Process.2
2025 MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling
Yingyue Li, Bencheng Liao, Wenyu Liu 0001, Xinggang Wang
ICCV1
2025 WeakCLIP: Adapting CLIP for Weakly-Supervised Semantic Segmentation
Lianghui Zhu, Xinggang Wang, Jiapei Feng, Tianheng Cheng, Yingyue Li, Bo Jiang 0011, Dingwen Zhang, Junwei Han 0001
Int. J. Comput. Vis.5
2021 Dynamic Design of the Turbomachinery Blade with the Joint Application of the Newly Developed Constrained Zone and the PSO Algorithm
abstract
New design approach to improve the turbomachinery performance is the research hot in various areas. In order to make up for the shortcomings about the uncertain design variables ranges and the simple 2D blade profile design method in previous studies, a modified inverse design approach established on the newly derived constrained zone and the PSO algorithm has been proposed in this study, and a developed design program with the Matlab platform was developed. With the adoption of the scaled 1:2.5 hydraulic model of the CAP1400 impeller as the target, the proposed joint design method was applied and verified.
Yeming Lu, Yingyue Li, Yongqi Yan
CSCWD4