Lijian Gao

dblp:132/0272 · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0002-6458-0660ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Semantic-Guided Visual Byte-Pair Encoding for Unified Autoregressive Multimodal Modeling
Wenlong Dong, Lijian Gao, Qirong Mao
ICIC (12)3
2026 Hierarchical Temporal Sequence Segmentation for weakly supervised video anomaly detection
Nuku Atta Kordzo Abiew, Lijian Gao, Qirong Mao
Expert Syst. Appl.2
2026 Attribute-calibrated local embeddings for prompt learning in vision-language models
Zhongchen Ma, Xingchen Wu, Yintong Wang, Lijian Gao
Inf. Sci.5
2026 HCATRE-AVAD: Hierarchical cross-alignment and temporal relational encoding for weakly supervised audio-visual anomaly detection
Nuku Atta Kordzo Abiew, Lijian Gao, Godbless Mensah, Wenlong Dong, Qirong Mao
Image Vis. Comput.2
2025 Global Enhanced Frame Prompt Tuning for Sound Event Detection
abstract
Sound Event Detection (SED) often employs pre-trained models to address data scarcity issues. However, existing systems usually treat the pretrained models as frozen feature extractors, resulting in suboptimal efficiency, or fully fine-tune the pretrained models, which requires substantial computational resources. To fully leverage the knowledge from pretrained models, we propose a novel Global Enhanced Frame Prompt Tuning (GE-FPT) framework, providing global and local insights tailored for SED tasks. Additionally, Frame Prompt Tuning (FPT) is proposed in our GE-FPT to effectively explore local temporal information, i.e., temporal details and context, which is essential for SED tasks, and in particular, for precise event boundary detection. Extensive experiments claim that our approach significantly outperforms full fine-tuning methods while substantially reducing computational costs. Our system achieves new state-of-the-art results, with PSDS1/PSDS2 scores of 0.628/0.845 on the DCASE2023 Challenge Task4 dataset. The source code is publicly available1.
Shiyu Yu, Lijian Gao, Qirong Mao
ICASSP2
2025 Scattering-Conditioned Diffusion Models for Multiple Appropriate Facial Reaction Generation
abstract
As embodied intelligence has become a new hot topic in current artificial intelligence research, facial reaction generation has increasingly become a key technology for achieving natural human-computer interaction. Existing methods typically rely on bimodal inputs of audio and visual signals, but they still suffer from poor cross-modal consistency and insufficient feature fusion, making it difficult to effectively capture complex facial features. To enhance the representational capacity of fused features, this paper proposes a Scattering-Conditioned Diffusion Model (SC-Diff), which extracts stable multi-scale structural features in the frequency domain via wavelet scattering transform and injects them into the diffusion generation process as conditional information, thereby enhancing the modeling ability of representing local facial variations. Furthermore, considering that different prediction tasks exhibit varying sensitivity to target changes during training, we introduce an uncertainty-based adaptive loss weighting strategy to dynamically balance three types of supervision targets: facial action units, facial affect, and facial expressions. Experimental results on the REACT 2025 dataset demonstrate that the proposed method outperforms existing state-of-the-art approaches across multiple evaluation metrics.
Qirong Mao, Qiwei Wu 0002, Yakui Ding, Lijian Gao
ACM Multimedia5
2025 StyU-STD: Style-Diverse Sample Generation from Unlabeled Data for Query-by-Example Spoken Term Detection
abstract
In recent years, query-by-example spoken term detection (QbE-STD) techniques have made significant progress in detection accuracy and speed. However, this task also encounters situations where labeled data is scarce or even nonexistent, with only unlabeled data available. Although some solutions exist, they still struggle to effectively handle highly variable speech, especially when it comes to differing styles. To address this issue, we propose a self-supervised learning method named Style-diverse sample generation from Unlabeled data for query-by-example Spoken Term Detection (StyU-STD). The core idea is to generate samples with the same content but different styles for learning. Specifically, we randomly extract segments from the speech to be tested as positive samples, while segments randomly extracted from other speech data are labeled as negative samples of the speech to be tested. In addition, various transformations are applied to alter the style of both positive and negative samples while preserving their original content. Then, the generated sample pairs are used to train the Style Suppressed Convolutional Network, which focuses more on content-related information in speech and effectively reduces the interference caused by style differences. The experimental results show that, across multiple datasets, our method outperforms existing methods, achieving higher accuracy and robustness.
Hanyu Ding, Lijian Gao, Wenlong Dong, Xiangrui Li, Qirong Mao
SMC2
2025 Dynamic prompting class distribution optimization for semi-supervised sound event detection
abstract
Semi-supervised sound event detection (SSED) tasks typically leverage a large amount of unlabeled and synthetic data to facilitate model generalization during training, reducing overfitting on a limited set of labeled data. However, the generalization training process often encounters challenges from noisy interference introduced by pseudo-labels or domain knowledge gaps. To alleviate noisy interference in class distribution learning, we propose an efficient semi-supervised class distribution learning method through dynamic prompt tuning, named prompting class distribution optimization (PADO). Specifically, when modeling real labeled data, PADO dynamically incorporates independent learnable prompt tokens to explore prior knowledge about the true distribution. Then, the prior knowledge serves as prompt information, dynamically interacting with the posterior noisy-class distribution information. In this case, PADO achieves class distribution optimization while maintaining model generalization, leading to a significant improvement in the efficiency of class distribution learning. Compared with state-of-the-art methods on the SSED datasets from DCASE 2019, 2020, and 2021 challenges, PADO achieves significant performance improvements. Furthermore, it is readily extendable to other benchmark models.
Lijian Gao, Qing Zhu 0002, Yaxin Shen, Qirong Mao, Yongzhao Zhan 0001
Frontiers Inf. Technol. Electron. Eng.1
2024 On Learning Frequency-Instance Correlations by Model-Agnostic Training for Synthetic Speech Detection
Lijian Gao, Qirong Mao
ACML2
2024 Leveraging Contrastive Language-Image Pre-Training and Bidirectional Cross-attention for Multimodal Keyword Spotting
Dong Liu 0037, Qirong Mao, Lijian Gao, Gang Wang 0023
Eng. Appl. Artif. Intell.3
2024 A novel conversational hierarchical attention network for speech emotion recognition in dyadic conversation
Mohammed Tellai, Lijian Gao, Qirong Mao, Mounir Abdelaziz
Multim. Tools Appl.2
2024 On Local Temporal Embedding for Semi-Supervised Sound Event Detection
abstract
Semi-supervised sound event detection (SSED) task requires recognizing the categories of events and marking each event's onset and offset times in a mixed audio recording using a small amount of weakly labeled and a large scale of unlabeled data. So, exploring local temporal information, i.e., local discrimination and local correlations in the time domain, is essential for SSED, and in particular, for precise event boundary detection. Besides, as manual-labeled datasets are scarce, SSED tasks require effectively exploiting unlabelled data to reduce overfitting, typically through regularization techniques. Recently, self-supervised learning provided a viable solution to leverage unlabeled data for effective feature learning in various downstream tasks. In this paper, we propose LTE-Net, a novel multitask framework, to learn the Local Temporal Embedding for SSED. Specifically, LTE-Net first locally down-samples the input spectrogram and learns the token embeddings with a high temporal resolution (i.e., local discrimination). Then, LTE-Net effectively models the local correlations among the token embeddings through self-supervised masked spectrogram modeling. Finally, a novel joint (self- and semi-supervision) regularization framework is employed for the training of LTE-Net to effectively leverage unlabeled data in SSED. Extensive experiments on DCASE 2019, 2020 and 2021 SSED datasets show that LTE-Net significantly outperformed existing methods and achieved 2.1% to 8.7%, 2.1% to 3.9% and 1.2% to 6.1% performance gains on the evaluation set in 2019, 2020 and 2021 datasets, respectively.
Lijian Gao, Qirong Mao, Ming Dong 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Joint-Former: Jointly Regularized and Locally Down-sampled Conformer for Semi-supervised Sound Event Detection
Lijian Gao, Qirong Mao, Ming Dong 0001
INTERSPEECH1
2023 TE-KWS: Text-Informed Speech Enhancement for Noise-Robust Keyword Spotting
abstract
Keyword spotting (KWS) presents a formidable challenge, particularly in high-noise environments. Traditional denoising algorithms that rely solely on speech have difficulty recovering speech that has been severely corrupted by noise. In this investigation, we develop an adaptive text-informed denoising model to bolster reliable keyword identification in the presence of considerable noise degradation. The whole proposed TE-KWS incorporates a tripartite branch structure, where the speech branch (SB) takes noisy speech as input which provides the raw speech information, the alignment branch (AB) accommodates aligned text input which facilitates accurate restoration of the corresponding speech when text with alignment is preserved, and the text branch (TB) handles unaligned text which prompts the model to autonomously learn the alignment between speech and text. To make the proposed denoising model more beneficial for KWS, following the training of the whole model,the alignment branch (AB) is frozen, and the model is fine-tuned by leveraging its speech restoration and forced alignment capabilities. Subsequently, the input for the text branch (TB) is supplanted with designated keywords, and a heavier denoising penalty is applied on the keywords period, thereby explicitly intensifying the speech restoration ability of the model for keywords. Finally, the Combined Adversarial Domain Adaptation (CADA) is implemented to enhance the robustness of KWS with regard to data pre-and post-speech enhancement (SE). Experimental results indicate that our approach not only markedly ameliorates highly corrupted speech, achieving SOTA performance for marginally corrupted speech, but also bolsters the efficacy and generalizability of prevailing mainstream KWS models.
Dong Liu 0037, Qirong Mao, Lijian Gao, Qinghua Ren, Zhenghan Chen, Ming Dong 0001
ACM Multimedia3
2022 Efficient Monaural Speech Separation with Multiscale Time-Delay Sampling
abstract
Recently, the segmented sample-level modeling approach based on Dual-Path Recurrent Neural Network (DPRNN) has been proved to be effective in Monaural Speech Separation (MSS). Many dual-path networks such as Dual-Path Transformer Network (DPTNet), with a series of improvements to DPRNN, have also improved the separation performance since these methods are effective to process long sequences. However, the receptive fields of these methods are fixed during local and global features learning, which makes it difficult to capture different scale local and global information in long sequences. In this paper, we propose a novel Multiscale Time-Delay Sampling method (MTDS) for the dual-path networks in MSS to learn sequence features from fine to coarse by multiscale time-delay sampling, which effectively integrates different scale local and global information for long sequences. Our experiments on the notable benchmark WSJ0-2mix data corpus result in 21.7dB SDRi and 21.5dB SI-SNRi, which obviously outperforms the state-of-the-arts without data augmentation.
Shuang-qing Qian, Lijian Gao, Hongjie Jia, Qirong Mao
ICASSP2
2022 Adaptive Hierarchical Pooling for Weakly-supervised Sound Event Detection
abstract
In Weakly-supervised Sound Event Detection (WSED), the ground truth of training data contains the presence or absence of each sound event only at the clip-level (i.e., no frame-level annotations). Recently, WSED has been formulated under the multi-instance learning framework, and a critical component within this formulation is the design of the temporal pooling function. In this paper, we propose an adaptive hierarchical pooling (HiPool) for WSED, which combines the advantages of max pooling in audio tagging and weighted average pooling in audio localization through a novel hierarchical structure and learns event-wise optimal pooling functions through continuous relaxation-based joint optimization. Extensive experiments on benchmark datasets show that HiPool outperforms the current pooling methods and greatly improves the performance of WSED. HiPool also has great generality - ready to be plugged into any WSED models.
Lijian Gao, Qirong Mao, Ming Dong 0001
ACM Multimedia1
2022 Weakly Supervised Sentiment-Specific Region Discovery for VSA
abstract
Abstract Local information has significant contributions to visual sentiment analysis (VSA). Recent studies about local region discovery need manually annotate region location. Affective local information learning and automatic discovery of sentiment-specific region are still the challenges in VSA. In this paper, we propose an end-to-end VSA method for weakly supervised sentiment-specific region discovery. Our method contains two branches: an automatic sentiment-specific region discovery branch and a sentiment analysis branch. In the sentiment-specific region discovery branch, a region proposal network with multiple convolution kernels is proposed to generate candidate affective regions. Then, we design the multiple instance learning (MIL) loss to remove redundant and noisy candidate regions. Finally, the sentiment analysis branch integrates both holistic and localized information obtained in the first branch by feature map coupling for final sentiment classification. Our method automatically discovers sentiment-specific regions by the constraint of MIL loss function without object-level labels. Quantitative and qualitative evaluations on four benchmark affective datasets demonstrate that our proposed method outperforms the state-of-the-art methods.
Luoyang Xue, Ang Xu, Qirong Mao, Lijian Gao, Jie Chen 0069
Comput. J.4
2021 Reproducibility Companion Paper: On Learning Disentangled Representation for Acoustic Event Detection
abstract
This companion paper is provided to describe the major experiments reported in our paper "On Learning Disentangled Representation for Acoustic Event Detection" published in ACM Multimedia 2019. To make the replication of our work easier, we first give an introduction of the computing environment where all of our experiments are conducted. Furthermore, we provide an environmental configuration file to setup the compiling environment and other artifacts including the source code, datasets and the files generated during our experiments. Finally, we summarize the structure and usage of the source code. For more details, please consult the README file in the archive of artifacts on GitHub: https://github.com/mastergofujs/SED_PyTorch.
Lijian Gao, Qirong Mao, Ming Dong 0001, Ratna Babu Chinnam, Lucile Sassatelli, Miguel Fabián Romero Rondón, Ujjwal Sharma 0001
ACM Multimedia1
2021 Learning to disentangle emotion factors for facial expression recognition in the wild
abstract
Facial expression recognition (FER) in the wild is a very challenging problem due to different expressions under complex scenario (e.g., large head pose, illumination variation, occlusions, etc.), leading to suboptimal FER performance. Accuracy in FER heavily relies on discovering superior discriminative, emotion-related features. In this paper, we propose an end-to-end module to disentangle latent emotion discriminative factors from the complex factors variables for FER to obtain salient emotion features. The training of proposed method contains two stages. First of all, emotion samples are used to obtain the latent representation using a variational auto-encoder with reconstruction penalization. Furthermore, the latent representation as the input is thrown into a disentangling layer to learn a set of discriminative emotion factors through the attention mechanism (e.g., a Squeeze-and-Excitation block) that encourages to separate emotion-related factors and nonaffective factors. Experimental results on public benchmark databases (RAF-DB and FER2013) show that our approach has remarkable performance in complex scenes than current state-of-the-art methods.
Qing Zhu 0002, Lijian Gao, Heping Song, Qirong Mao
Int. J. Intell. Syst.2
2019 On Learning Disentangled Representation for Acoustic Event Detection
abstract
Polyphonic Acoustic Event Detection (AED) is a challenging task as the sounds are mixed with the signals from different events, and the features extracted from the mixture do not match well with features calculated from sounds in isolation, leading to suboptimal AED performance. In this paper, we propose a supervised β-VAE model for AED, which adds a novel event-specific disentangling loss in the objective function of disentangled learning. By incorporating either latent factor blocks or latent attention in disentangling, supervised β-VAE learns a set of discriminative features for each event. Extensive experiments on benchmark datasets show that our approach outperforms the current state-of-the-arts (top-1 performers in the Detection and Classification of Acoustic Scenes and Events (DCASE) 2017 AED challenge). Supervised β-VAE has great success in challenging AED tasks with a large variety of events and imbalanced data.
Lijian Gao, Qirong Mao, Ming Dong 0001, Yu Jing, Ratna Babu Chinnam
ACM Multimedia1
2009 A New Decomposition Algorithm of DCT-IV/DST-IV for Realizing Fast IMDCT Computation
abstract
In this letter, a new decomposition algorithm of type-IV discrete cosine transform/type-IV discrete sine transform (DCT-IV/DST-IV) and architecture of a hardware accelerator for realizing fast inverse modified discrete cosine transform (IMDCT) computation are presented. The comparisons of computational complexity with some well-known algorithms show that the proposed algorithm possesses the advantages of higher computational efficiency and simpler hardware implementation. A real-time audio decoding experiment was performed to verify the efficiency of the proposed algorithm. Experimental results show more than one third of computational cycles are saved compared with one of reported fast algorithms for IMDCT computation.
Ping Li 0014, Lijian Gao
IEEE Signal Process. Lett.5