Yili Li

dblp:28/7476 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing fake news video detection with self-driven question-answer from LMMs
Yili Li, Jian Lang, Rongpei Hong, Fan Zhou 0002
Inf. Process. Manag.2
2026 CapSRA: LMM-guided caption semantic retrieval aggregation for hateful meme detection
Yili Li, Yong Wang 0046, Jin Wu 0002, Kunpeng Zhang 0001, Fan Zhou 0002
Pattern Recognit.3
2025 ProAPO: Progressively Automatic Prompt Optimization for Visual Classification
abstract
Vision-language models (VLMs) have made significant progress in image classification by training with large-scale paired image-text data. Their performances largely depend on the prompt quality. While recent methods show that visual descriptions generated by large language models (LLMs) enhance the generalization of VLMs, class-specific prompts may be inaccurate or lack discrimination due to the hallucination in LLMs. In this paper, we aim to find visually discriminative prompts for fine-grained categories with minimal supervision and no human-in-the-loop. An evolution-based algorithm is proposed to progressively optimize language prompts from task-specific templates to class-specific descriptions. Unlike optimizing templates, the search space shows an explosion in class-specific candidate prompts. This increases prompt generation costs, iterative times, and the overfitting problem. To this end, we first introduce several simple yet effective edit-based and evolution-based operations to generate diverse candidate prompts by one-time query of LLMs. Then, two sampling strategies are proposed to find a better initial search point and reduce traversed categories, saving iteration costs. Moreover, we apply a novel fitness score with entropy constraints to mitigate overfitting. In a challenging one-shot image classification setting, our method outperforms existing textual prompt-based methods and improves LLM-generated description methods across 13 datasets. Meanwhile, we demonstrate that our optimal prompts improve adapter-based methods and transfer effectively across different backbones. Our code is available at here.
Xiangyan Qu, Gaopeng Gou, Jiamin Zhuang, Jing Yu 0007, Qihao Wang, Yili Li, Gang Xiong 0001
CVPR7
2025 Commonality Augmented Disentanglement for Multimodal Crowdfunding Success Prediction
abstract
Online crowdfunding platforms have been gaining increasing popularity due to their convenience in soliciting social capital from the public. These platforms offer valuable opportunities for fundraisers to bring their creative products to life and support pro-social projects. However, the relatively low success rate of crowdfunding campaigns highlights the need for better strategies. While existing studies have explored various factors that contribute to crowdfunding success, they often overlook the intricate relationships between different aspects of crowdfunding. In this paper, we propose a novel model called Commonality Augmented Multimodal Disentanglement (CAMD) for predicting crowdfunding success. It can disentangle the roles of different data modalities, separating common factors from specific ones. Besides, we enhance the disentangled commonality using an augmentation network to achieve balanced representation of the skewed interrelations between different modalities. At last, we introduce a cross-attention-based multimodal fusion mechanism that further improves model performance by highlighting the crucial role of crowdfunding attributes. Experiments conducted on two large-scale crowdfunding datasets demonstrate the effectiveness and generalizability of our model.
Jiayang Li 0006, Xovee Xu, Yili Li, Ting Zhong, Kunpeng Zhang 0001, Fan Zhou 0002
ICASSP3
2025 REAL: Retrieval-Augmented Prototype Alignment for Improved Fake News Video Detection
abstract
Detecting fake news videos has emerged as a critical task due to their profound implications in politics, finance, and public health. However, existing methods often fail to distinguish real videos from their subtly manipulated counterparts, resulting in suboptimal performance. To address this limitation, we propose REAL, a novel model-agnostic REtrieval-Augmented prototype-aLignment framework. REAL first introduces an LLM-driven video retriever to identify contextually relevant samples for a given target video. Subsequently, a dual-prototype aligner is carefully developed to model two distinct prototypes: one representing authentic patterns from retrieved real news videos and the other encapsulating manipulation-specific patterns from fake samples. By aligning the target video’s representations with its ground-truth prototype while distancing them from the opposing prototype, the aligner captures manipulation-aware representations capable of detecting even subtle video manipulations. Finally, these enriched representations are seamlessly integrated into existing detection models in a plug-and-play manner. Extensive experiments on three benchmarks demonstrate that REAL largely enhances the detection ability of existing methods. The code and data for reproducing the results are available at https://github.com/Jian-Lang/REAL.
Yili Li, Jian Lang, Rongpei Hong, Zhangtao Cheng, Ting Zhong, Fan Zhou 0002
ICME1
2025 T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video Retrieval
abstract
Text-to-video retrieval essentially aims to train models to align visual content with textual descriptions accurately. Due to the impressive general multimodal knowledge demonstrated by image-text pretrained models such as CLIP, existing work has primarily focused on extending CLIP knowledge for video-text tasks. However, videos typically contain richer information than images. In current video-text datasets, textual descriptions can only reflect a portion of the video content, leading to partial misalignment in video-text matching. Therefore, directly aligning text representations with video representations can result in incorrect supervision, ignoring the inequivalence of information. In this work, we propose T2VParser to extract multiview semantic representations from text and video, achieving adaptive semantic alignment rather than aligning the entire representation. To extract corresponding representations from different modalities, we introduce Adaptive Decomposition Tokens, which consist of a set of learnable tokens shared across modalities. The goal of T2VParser is to emphasize precise alignment between text and video while retaining the knowledge of pretrained models. Experimental results demonstrate that T2VParser achieves accurate partial alignment through effective cross-modal content decomposition. The code is available at https://github.com/Lilidamowang/T2VParser.
Yili Li, Gang Xiong 0001, Gaopeng Gou, Xiangyan Qu, Jiamin Zhuang, Zhen Li 0011, Junzheng Shi
ACM Multimedia1
2025 Biting Off More Than You Can Detect: Retrieval-Augmented Multimodal Experts for Short Video Hate Detection
abstract
Short Video Hate Detection (SVHD) is increasingly vital as hateful content -such as racial and gender-based discriminationspreads rapidly across platforms like TikTok, YouTube Shorts, and Instagram Reels.Existing approaches face significant challenges: hate expressions continuously evolve, hateful signals are dispersed across multiple modalities (audio, text, and vision), and the contribution of each modality varies across different hate content.To address these issues, we introduce MoRE (Mixture of Retrievalaugmented multimodal Experts), a novel framework designed to enhance SVHD.MoRE employs specialized multimodal experts for each modality, leveraging their unique strengths to identify hateful content effectively.To ensure model's adaptability to rapidly evolving hate content, MoRE leverages contextual knowledge extracted from relevant instances retrieved by a powerful joint multimodal video retriever for each target short video.Moreover, a dynamic sample-sensitive integration network adaptively adjusts the importance of each modality on a per-sample basis, optimizing the detection process by prioritizing the most informative modalities for each instance.Our MoRE adopts an end-to-end training strategy that jointly optimizes both expert networks and the overall framework, resulting in nearly a twofold improvement in training efficiency, which in turn enhances its applicability to real-world scenarios.Extensive experiments on three benchmarks demonstrate that MoRE surpasses state-of-the-art baselines, achieving an average improvement of 6.91% in macro-F1 score across all datasets.
Jian Lang, Rongpei Hong, Yili Li, Xovee Xu, Fan Zhou 0002
WWW4
2024 IIU: Independent Inference Units for Knowledge-Based Visual Question Answering
Yili Li, Jing Yu 0007, Keke Gai, Gang Xiong 0001
KSEM (4)1
2024 T2VIndexer: A Generative Video Indexer for Efficient Text-Video Retrieval
abstract
Current text-video retrieval methods mainly rely on cross-modal matching between queries and videos to calculate their similarity scores, which are then sorted to obtain retrieval results. This method considers the matching between each candidate video and the query, but it incurs a significant time cost and will increase notably with the increase of candidates. Generative models are common in natural language processing and computer vision, and have been successfully applied in document retrieval, but their application in multimodal retrieval remains unexplored. To enhance retrieval efficiency, in this paper, we introduce a model-based video indexer named T2VIndexer, which is a sequence-to-sequence generative model directly generating video identifiers and retrieving candidate videos with constant time complexity. T2VIndexer aims to reduce retrieval time while maintaining high accuracy. To achieve this goal, we propose video identifier encoding and query-identifier augmentation approaches to represent videos as short sequences while preserving their semantic information. Our method consistently enhances the retrieval efficiency of current state-of-the-art models on four standard datasets. It enables baselines with only 30%-50% of the original retrieval time to achieve better retrieval performance on MSR-VTT (+1.0%), MSVD (+1.8%), ActivityNet (+1.5%), and DiDeMo (+0.2%). The code is available at https://github.com/Lilidamowang/T2VIndexer-generativeSearch.
Yili Li, Jing Yu 0007, Keke Gai, Bang Liu 0003, Gang Xiong 0001, Qi Wu 0001
ACM Multimedia1
2012 Electroencephalogram signals classification for sleep-state decision - a Riemannian geometry approach
abstract
In this work, the authors study the classification of electroencephalogram (EEG) signals for the determination of the state of sleep of a patient. They employ the power spectral density (PSD) matrices as the feature for the distinction between different classes of EEG signals. This not only allows us to examine the power spectrum contents of each signal as well as the correlation between the multi-channel signals, but also complies with what clinical experts use in their visual judgement of EEG signals. To establish a metric facilitating the classification, the authors exploit the specific geometric properties, and develop, with the aid of fibre bundle theory, an appropriate metric in the Riemannian manifold described by the PSD matrices. To use this new metric effectively for the EEG signal classification, the authors further need to find a weighting for the PSD matrices so that the distances of similar features are minimised whereas those for dissimilar features are maximised. A closed form of this weighting matrix is obtained by solving an equivalent convex optimisation problem. The effectiveness of using these new metrics is examined by applying them to a collection of recorded EEG signals for sleep pattern classification based on the k-nearest neighbour decision algorithm with excellent outcome.
Yili Li, Kon Max Wong, Hubert de Bruin
IET Signal Process.1