Xiangyan Qu

dblp:295/4241 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0003-3658-8099ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Image recognition and object detection · 29% Vision and language · 27% Language models and text generation · 23%
Computer graphics and multimedia
2 papers
Multimedia analysis and retrieval · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
0.912025
ProAPO: Progressively Automatic Prompt Optimization for Visual Classification · CVPR 2025
Computer vision › Image recognition and object detection
image classification
0.912025
ProAPO: Progressively Automatic Prompt Optimization for Visual Classification · CVPR 2025
Natural language and speech › Language models and text generation › prompting › prompt engineering
prompt optimization
0.912025
ProAPO: Progressively Automatic Prompt Optimization for Visual Classification · CVPR 2025
Computer vision › Vision and language
vision-language model
0.912025
ProAPO: Progressively Automatic Prompt Optimization for Visual Classification · CVPR 2025
Multimedia analysis and retrieval › video retrieval
text-to-video retrieval
0.912025
T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video Retrieval · ACM Multimedia 2025
Computer vision › Vision and language
cross-modal alignment
0.812024
Visual-Semantic Decomposition and Partial Alignment for Document-based Zero-Shot Learning · ACM Multimedia 2024
Machine learning › Transfer learning and domain adaptation
zero-shot learning
0.812024
Visual-Semantic Decomposition and Partial Alignment for Document-based Zero-Shot Learning · ACM Multimedia 2024
Multimedia analysis and retrieval › cross-modal retrieval
image-text retrieval
0.812024
Towards Fast and Accurate Image-Text Retrieval With Self-Supervised Fine-Grained Alignment · IEEE Trans. Multim. 2024
Natural language and speech › Language models and text generation
large language model
0.312025
ProAPO: Progressively Automatic Prompt Optimization for Visual Classification · CVPR 2025
Natural language and speech › Language models and text generation › prompting › prompt engineering
prompt generation
0.312025
ProAPO: Progressively Automatic Prompt Optimization for Visual Classification · CVPR 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.212024
Towards Fast and Accurate Image-Text Retrieval With Self-Supervised Fine-Grained Alignment · IEEE Trans. Multim. 2024
Computer vision › Segmentation and scene understanding
semantic decomposition
0.212024
Visual-Semantic Decomposition and Partial Alignment for Document-based Zero-Shot Learning · ACM Multimedia 2024

Methods — techniques the papers use, named apart from their topics

cross-modal alignment · 2.4self-supervised contrastive learning · 1.5independent-embedding framework · 1.5evolution-based algorithm · 0.9entropy-constrained fitness score · 0.9edit-based operations · 0.9adaptive decomposition tokens · 0.9CLIP · 0.9variance loss · 0.8orthogonality regularization · 0.8contrastive learning · 0.8
YearPublicationVenuePosition
2025 ProAPO: Progressively Automatic Prompt Optimization for Visual Classification
abstract
Vision-language models (VLMs) have made significant progress in image classification by training with large-scale paired image-text data. Their performances largely depend on the prompt quality. While recent methods show that visual descriptions generated by large language models (LLMs) enhance the generalization of VLMs, class-specific prompts may be inaccurate or lack discrimination due to the hallucination in LLMs. In this paper, we aim to find visually discriminative prompts for fine-grained categories with minimal supervision and no human-in-the-loop. An evolution-based algorithm is proposed to progressively optimize language prompts from task-specific templates to class-specific descriptions. Unlike optimizing templates, the search space shows an explosion in class-specific candidate prompts. This increases prompt generation costs, iterative times, and the overfitting problem. To this end, we first introduce several simple yet effective edit-based and evolution-based operations to generate diverse candidate prompts by one-time query of LLMs. Then, two sampling strategies are proposed to find a better initial search point and reduce traversed categories, saving iteration costs. Moreover, we apply a novel fitness score with entropy constraints to mitigate overfitting. In a challenging one-shot image classification setting, our method outperforms existing textual prompt-based methods and improves LLM-generated description methods across 13 datasets. Meanwhile, we demonstrate that our optimal prompts improve adapter-based methods and transfer effectively across different backbones. Our code is available at here.
Xiangyan Qu, Gaopeng Gou, Jiamin Zhuang, Jing Yu 0007, Qihao Wang, Yili Li, Gang Xiong 0001
CVPR1
2025 T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video Retrieval
abstract
Text-to-video retrieval essentially aims to train models to align visual content with textual descriptions accurately. Due to the impressive general multimodal knowledge demonstrated by image-text pretrained models such as CLIP, existing work has primarily focused on extending CLIP knowledge for video-text tasks. However, videos typically contain richer information than images. In current video-text datasets, textual descriptions can only reflect a portion of the video content, leading to partial misalignment in video-text matching. Therefore, directly aligning text representations with video representations can result in incorrect supervision, ignoring the inequivalence of information. In this work, we propose T2VParser to extract multiview semantic representations from text and video, achieving adaptive semantic alignment rather than aligning the entire representation. To extract corresponding representations from different modalities, we introduce Adaptive Decomposition Tokens, which consist of a set of learnable tokens shared across modalities. The goal of T2VParser is to emphasize precise alignment between text and video while retaining the knowledge of pretrained models. Experimental results demonstrate that T2VParser achieves accurate partial alignment through effective cross-modal content decomposition. The code is available at https://github.com/Lilidamowang/T2VParser.
Yili Li, Gang Xiong 0001, Gaopeng Gou, Xiangyan Qu, Jiamin Zhuang, Zhen Li 0011, Junzheng Shi
ACM Multimedia4
2025 Soft Multi-view Representation Learning for Disambiguating Text-Based Person Retrieval
Jiamin Zhuang, Jing Yu 0007, Xiangyan Qu, Yuanmin Tang, Gaopeng Gou, Gang Xiong 0001, Qi Wu 0001
WASA (1)3
2024 Visual-Semantic Decomposition and Partial Alignment for Document-based Zero-Shot Learning
abstract
Recent work shows that documents from encyclopedias serve as helpful auxiliary information for zero-shot learning. Existing methods align the entire semantics of a document with corresponding images to transfer knowledge. However, they disregard that semantic information is not equivalent between them, resulting in a suboptimal alignment. In this work, we propose a novel network to extract multi-view semantic concepts from documents and images and align the matching rather than entire concepts. Specifically, we propose a semantic decomposition module to generate multi-view semantic embeddings from visual and textual sides, providing the basic concepts for partial alignment. To alleviate the issue of information redundancy among embeddings, we propose the local-to-semantic variance loss to capture distinct local details and multiple semantic diversity loss to enforce orthogonality among embeddings. Subsequently, two losses are introduced to partially align visual-semantic embedding pairs according to their semantic relevance at the view and word-to-patch levels. Consequently, we consistently outperform state-of-the-art methods under two document sources in three standard benchmarks for document-based zero-shot learning. Qualitatively, we show that our model learns the interpretable partial association. Code is available at https://github.com/MorningStarOvO/EmDepart.
Xiangyan Qu, Jing Yu 0007, Keke Gai, Jiamin Zhuang, Yuanmin Tang, Gang Xiong 0001, Gaopeng Gou, Qi Wu 0001
ACM Multimedia1
2024 Towards Fast and Accurate Image-Text Retrieval With Self-Supervised Fine-Grained Alignment
abstract
Image-text retrieval requires the system to bridge the heterogenous gap between vision and language for accurate retrieval while keeping the network lightweight-enough for efficient retrieval. Existing trade-off solutions mainly study from the view of incorporating cross-modal interactions with the independent-embedding framework or leveraging stronger pre-trained encoders, which still demand time-consuming similarity measurement or heavyweight model structure in the retrieval stage. In this work, we propose an image-text alignment module SelfAlign on top of the independent-embedding framework, which improves the retrieval accuracy while maintains the retrieval efficiency without extra supervision. SelfAlign contains two collaborative sub-modules that force image-text alignment at both the concept level and context level by self-supervised contrastive learning. It doesn't require cross-modal embedding interactions during training while maintaining independent image and text encoders during retrieval. With comparable time cost, SelfAlign consistently boosts the accuracy of state-of-the-art non-pre-training independent-embedding models respectively by 9.1%, 4.2%, and 6.6% in terms of R@sum score on Flickr30 K, MS-COCO 1 K and MS-COCO 5 K datasets. The retrieval accuracy also outperforms most existing interactive-embedding models with orders of magnitude decrease in retrieval time. The source code is available at:https://github.com/Zjamie813/SelfAlign.
Jiamin Zhuang, Jing Yu 0007, Xiangyan Qu, Yue Hu 0002
IEEE Trans. Multim.4
2021 An Effective Deep Neural Network for Lung Lesions Segmentation From COVID-19 CT Images
abstract
Automatic segmentation of lung lesions from COVID-19 computed tomography (CT) images can help to establish a quantitative model for diagnosis and treatment. For this reason, this article provides a new segmentation method to meet the needs of CT images processing under COVID-19 epidemic. The main steps are as follows: First, the proposed region of interest extraction implements patch mechanism strategy to satisfy the applicability of 3-D network and remove irrelevant background. Second, 3-D network is established to extract spatial features, where 3-D attention model promotes network to enhance target area. Then, to improve the convergence of network, a combination loss function is introduced to lead gradient optimization and training direction. Finally, data augmentation and conditional random field are applied to realize data resampling and binary segmentation. This method was assessed with some comparative experiment. By comparison, the proposed method reached the highest performance. Therefore, it has potential clinical applications.
Cheng Chen 0024, Kangneng Zhou, Muxi Zha, Xiangyan Qu, Xiaoyu Guo 0004, Ruoxiu Xiao
IEEE Trans. Ind. Informatics4