Qisheng Xu

dblp:335/8976 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0003-4141-5950ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2025 SSAST-Adapter: A Parameter-efficient Incremental Learning Algorithm for Underwater Acoustic Target Recognition
abstract
Underwater acoustic target recognition involves identifying and classifying targets in underwater environments using acoustic signals. In recent years, deep learning has made significant progress in this field. However, the models require the entire dataset to be available upfront, and classification categories must be predefined. In practical scenarios, new objects or species may appear in underwater environments over time, making it impractical to retrain a model from scratch each time new data is introduced. At the same time, training on new data inevitably leads to catastrophic forgetting of past data. To address these challenges, we propose a large-scale pre-training strategy combined with adapters to enable incremental learning without the need for complete retraining. To minimize the number of parameters during model fine-tuning, we employ a adapter structure at various network layers, reducing the number of trainable parameters to less than 2%. Experimental results demonstrate that our proposed method effectively reduces the need for full retraining by allowing the model to update in a resource-efficient manner using only new data, and the model has achieved high recognition accuracy in underwater acoustic target recognition tasks.
Qisheng Xu, Boqing Zhu, Zijian Gao, Lingbin Zeng, Kele Xu
ICASSP2
2025 Multi-view Fusion and Parameter Perturbation for Few-Shot Class-Incremental Audio Classification
Yulu Fang, Mingyue He, Qisheng Xu, Jianqiao Zhao, Cheng Yang 0004, Kele Xu, Yong Dou
INTERSPEECH3
2025 AudioSet-R: A Refined AudioSet with Multi-Stage LLM Label Reannotation
abstract
AudioSet is a widely used benchmark in the audio research community and has significantly advanced various audio-related tasks. However, persistent issues with label accuracy and completeness remain critical bottlenecks that limit performance in downstream applications. To address the aforementioned challenges, we propose a three-stage reannotation framework that harnesses general-purpose audio-language foundation models to systematically improve the label quality of AudioSet. The framework employs a cross-modal prompting strategy, inspired by the concept of prompt chaining, wherein prompts are sequentially composed to execute subtasks (audio comprehension, label synthesis, and semantic alignment). Leveraging this framework, we construct a high-quality, structured relabeled version of AudioSet-R. Extensive experiments conducted on representative audio classification models-including AST, PANNs, SSAST, and AudioMAE-consistently demonstrate substantial performance improvements, thereby validating the generalizability and effectiveness of the proposed approach in enhancing label reliability.The code is publicly available at: https://github.com/colaudiolab/AudioSet-R.
Qisheng Xu, Yi Su 0010, Yong Dou, Xinwang Liu 0002, Kele Xu
ACM Multimedia2
2025 Higher-Order Vision-Language Fusion for Video Popularity Prediction
abstract
Predicting the popularity of social media videos involves estimating user engagement based on rich multimodal information embedded within the posts. Unlike static images, videos incorporate temporally evolving visual signals that, alongside associated metadata such as descriptions, hashtags, timestamps, and user attributes, offer valuable insights into their potential audience reach. Prior approaches typically extract features from different modalities independently and merge them via naïve concatenation, which overlooks the semantic discrepancy and interaction dynamics across modalities. To address these limitations, we propose a feature fusion framework that encodes and aligns video content and associated textual cues into a shared semantic space. By jointly modeling temporally structured visual features with context-aware textual embeddings, our method effectively captures cross-modal correlations that are crucial for discerning content virality patterns. In addition, we incorporate user-centric behavioral profiles and content creation dynamics, enriching the representation with personalized signals that reflect audience-specific preferences. Notably, our method achieves top-tier performance in the 2025 SMP challenge, ranking among the highest-performing entries. This strong empirical result underscores the value of deep semantic alignment across video, text, and user domains in accurately forecasting social media video popularity.
Kele Xu, Qisheng Xu, Binli Luo, Han Zhou 0003, Zengming Lin, Hui Geng, Xianhan Tan
ACM Multimedia2
2024 Adapter-Based Incremental Learning for Face Forgery Detection
abstract
Many existing face forgery detection methods primarily revolve around learning general representations on predefined datasets and subsequently crossing these static representations to other datasets. However, these approaches could lead to catastrophic forgetting in real-world scenarios, especially when new forgery methods continually emerge. In this paper, we proposed a novel incremental learning framework for face forgery detection, where we design an adapter-based incremental learning scheme combined with a confidence-based ensemble prediction mechanism. When confronted with new forgery methods, we incorporate small trainable adapter modules, which are retrained along with their corresponding classification layers, yielding a series of task-specific modules. Then we incorporate a confidence-based ensemble prediction mechanism to aggregate all predictions. Through comprehensive evaluations on multiple benchmark datasets (FF++, DFD, and Celeb-DF), our method successfully mitigates the catastrophic forgetting problem in a cost-effective manner and attains state-of-the-art performance in cross-dataset scenario.
Caili Gao, Qisheng Xu, Peng Qiao, Kele Xu, Xifu Qian, Yong Dou
ICASSP2
2024 Higher-Order Vision-Language Alignment for Social Media Prediction
abstract
The prediction task of social media popularity aims to automatically forecast the future popularity of the posts by leveraging vast amounts of social media data. This data encompasses diverse visual and textual content, including photos, categories, custom tags, temporal information, and geographical data. Existing methods have explored multiple feature types to enhance popularity prediction. Despite their success, visual and textual features-both crucial pieces of information-are often simply concatenated after extraction, ignoring the divergence between these two feature spaces. In this paper, we propose a method to project visual and language information into an aligned semantic representation, thereby uncovering intricate associations between these two modalities. Specifically, we leverage the BLIP-2 model to understand and generate visual description text that encapsulates the content of photos. Semantic embeddings are then extracted from all available visual and textual information. Additionally, we deeply exploit user-related behavior and characteristic information to extract features, uncovering hidden clues for post popularity prediction. Leveraging these improvements, we conduct extensive experiments to demonstrate the effectiveness of our proposed method.
Mingsheng Tu, Tianjiao Wan, Qisheng Xu, Xinhao Jiang, Kele Xu, Cheng Yang 0004
ACM Multimedia3
2023 Raw Ultrasound-Based Phonetic Segments Classification Via Mask Modeling
abstract
Ultrasound tongue imaging is widely used in clinical linguistics and phonetics. Recently, deep neural networks, especially convolutional neural networks, have been widely used in the interpretation and analysis of ultrasound tongue images (UTI). Despite achieving satisfactory performance, deep models rely on a large amount of manually labeled data, which is often difficult to obtain in practical settings. To address this issue, this paper focuses on how to utilize a large amount of unlabeled UTI data to improve the performance of UTI classification task. Specifically, we explore self-supervised learning with masking modeling strategy. By predicting the masked part, our pre-trained model enables the neural network to infer contextual information. Then, we fine-tune the pre-trained model with a small amount of labeled data. Compared with the previous competing algorithms, our method can improve the classification accuracy by an average of 13.33% in four different scenarios.
Kang You, Bo Liu 0014, Kele Xu, Yunsheng Xiong, Qisheng Xu, Ming Feng, Tamás Gábor Csapó, Boqing Zhu
ICASSP5
2023 Spatial and Frequency Domains Inconsistency Learning for Face Forgery Detection
Caili Gao, Peng Qiao, Yong Dou, Qisheng Xu, Xifu Qian
ICONIP (12)4