VLDB 2026 Research / reviewers in the wild / expert
Shengchen Li
dblp:49/10857
· DBLP profile ↗
17ranked-venue papers
2as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LBOR: Laplace-Beltrami Operator Regularization for Robust Skeleton-based Isolated Sign Language Recognition
Peihong Zhang, Rui Sang, Shengchen Li |
FG | 5 |
| 2026 | KCAM-SENet: Speech enhancement network with KAN-based channel attention module
Linhui Sun, Zhaowei Ding, Yuhang Qin, Shengchen Li, Xi Shao, Chng Eng Siong |
Speech Commun. | 5 |
| 2024 | Language-based Audio Retrieval with GPT-Augmented Captions and Self-Attended Audio ClipsabstractWith the explosion of user-generated content in recent years, efficient methods for organizing multimedia databases based on content and retrieving relevant items have become essential. Language-based audio retrieval seeks to find relevant audio clips based on natural language queries. However, there exists a scarcity of datasets specifically developed for this task. Moreover, the language annotations often carry biases, leading to unsatisfactory retrieval accuracy. In this work, we propose a novel framework for language-based audio retrieval that aims to: 1) utilize GPT-generated text to augment audio captions, thereby improving language diversity; 2) employ audio self-attention mechanisms to capture intricate acoustic features and temporal dependencies. Experiments conducted on two public datasets, containing both short- and long-term audios, demonstrate that our framework can achieve significant performance improvements compared with other methods. Specifically, the proposed framework can achieve a 27% increase in mean average precision (mAP) on the Clotho dataset, and a 31% improvement in mAP on the AudioCaps dataset compared with the baseline. Fuyu Gu, Yiyan Xu, Yushan Pan, Shengchen Li, Haiyang Zhang 0004 |
CSCWD | 6 |
| 2024 | TF-SepNet: An Efficient 1D Kernel Design in Cnns for Low-Complexity Acoustic Scene ClassificationabstractRecent studies focus on developing efficient systems for acoustic scene classification (ASC) using convolutional neural networks (CNNs), which typically consist of consecutive kernels. This paper highlights the benefits of using separate kernels as a more powerful and efficient design approach in ASC tasks. Inspired by the time-frequency nature of audio signals, we propose TF-SepNet, a CNN architecture that separates the feature processing along the time and frequency dimensions. Features resulted from the separate paths are then merged by channels and directly forwarded to the classifier. Instead of the conventional two dimensional (2D) kernel, TF-SepNet incorporates one dimensional (1D) kernels to reduce the computational costs. Experiments have been conducted using the TAU Urban Acoustic Scene 2022 Mobile development dataset. The results show that TF-SepNet outperforms similar state-of-the-arts that use consecutive kernels. A further investigation reveals that the separate kernels lead to a larger effective receptive field (ERF), which enables TF-SepNet to capture more time-frequency features. Yiqiang Cai, Peihong Zhang, Shengchen Li |
ICASSP | 3 |
| 2024 | Speech Formants Integration for Generalized Detection of Synthetic Speech Spoofing Attacks
Kexu Liu, Shengchen Li, Xi Shao |
INTERSPEECH | 3 |
| 2024 | Intelligent Fish Detection System with Similarity-Aware TransformerabstractFish detection in water-land transfer has significantly contributed to the fishery. However, manual fish detection in crowd-collaboration performs inefficiently and expensively, involving insufficient accuracy. To further enhance the water-land transfer efficiency, improve detection accuracy, and reduce labor costs, this work designs a new type of lightweight and plug-and-play edge intelligent vision system to automatically conduct fast fish detection with high-speed camera. Moreover, a novel similarity-aware vision Transformer for fast fish detection (FishViT) is proposed to onboard identify every single fish in a dense and similar group. Specifically, a novel similarity-aware multi-level encoder is developed to enhance multi-scale features in parallel, thereby yielding discriminative representations for varying-size fish. Additionally, a new soft-threshold attention mechanism is introduced, which not only effectively eliminates background noise from images but also accurately recognizes both the edge details and overall features of different similar fish. 85 challenging video sequences with high framerate and high-resolution are collected to establish a benchmark from real fish water-land transfer scenarios. Exhaustive evaluation conducted with this challenging benchmark has proved the robustness and effectiveness of FishViT with over 80 FPS. Real work scenario tests validate the practicality of the proposed method. The code and demo video are available at https://github.com/vision4robotics/FishViT. Shengchen Li, Haobo Zuo, Changhong Fu 0001 |
IROS | 1 |
| 2024 | Acoustic Scene Classification Across Cities and Devices via Feature DisentanglementabstractAcoustic Scene Classification (ASC) is a task that classifies a scene according to environmental acoustic signals. Audios collected from different cities and devices often exhibit biases in feature distributions, which may negatively impact ASC performance. Taking the city and device of the audio collection as two types of data domain, this paper attempts to disentangle the audio features of each domain to remove the related feature biases. A dual-alignment framework is proposed to generalize the ASC system on new devices or cities, by aligning boundaries across domains and decision boundaries within each domain. During the alignment, the maximum classifier discrepancy and gradient reversed layer are used for the feature disentanglement of scene, city and device, while four candidate domain classifiers are proposed to explore the optimal solution of feature disentanglement. To evaluate the dual-alignment framework, three experiments of biased ASC tasks are designed: 1) cross-city ASC in new cities; 2) cross-device ASC in new devices; 3) cross-city-device ASC in new cities and new devices. Results demonstrate the superiority of the proposed framework, showcasing performance improvements of 0.9%, 19.8%, and 10.7% on classification accuracy, respectively. The effectiveness of the proposed feature disentanglement approach is further evaluated in both biased and unbiased ASC problems, and the results demonstrate that better-disentangled audio features can lead to a more robust ASC system across different devices and cities. This paper advocates for the integration of feature disentanglement in ASC systems to achieve more reliable performance. Yizhou Tan, Haojun Ai, Shengchen Li, Mark D. Plumbley |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Visually-Aware Audio Captioning With Adaptive Audio-Visual AttentionabstractAudio captioning aims to generate text descriptions of audio clips.In the real world, many objects produce similar sounds.How to accurately recognize ambiguous sounds is a major challenge for audio captioning.In this work, inspired by inherent human multimodal perception, we propose visuallyaware audio captioning, which makes use of visual information to help the description of ambiguous sounding objects.Specifically, we introduce an off-the-shelf visual encoder to extract video features and incorporate the visual features into an audio captioning system.Furthermore, to better exploit complementary audio-visual contexts, we propose an audio-visual attention mechanism that adaptively integrates audio and visual context and removes the redundant information in the latent space.Experimental results on AudioCaps, the largest audio captioning dataset, show that our proposed method achieves state-of-theart results on machine translation metrics. Xubo Liu 0001, Qiushi Huang, Xinhao Mei, Haohe Liu, Qiuqiang Kong, Jianyuan Sun, Shengchen Li, Tom Ko, Yu Zhang 0006, Lilian Tang, Mark D. Plumbley, Volkan Kilic, Wenwu Wang 0001 |
INTERSPEECH | 7 |
| 2023 | Transductive Feature Space Regularization for Few-shot Bioacoustic Event Detection
Yizhou Tan, Haojun Ai, Shengchen Li |
INTERSPEECH | 3 |
| 2022 | DRVAT: Exploring RSSI series representation and attention model for indoor positioningabstractAlthough Bluetooth Low Energy (BLE) fingerprinting localization has become a hot research topic with encouraging results, it is difficult to predict the location depending on a short duration received signal strength indication (RSSI) sequence in realistic scenarios due to the severe fluctuation of RSSI. We introduce a new perspective to view the indoor positioning problem by radio map fingerprint. We argue that even though beacons may be independently deployed, the RSSI series bear certain spatial relation because of their copresence in the same physical space. The latent relation implicitly conveyed by the coexistence of their signals at various indoor locations. Unlike existing approaches that try to find a direct mapping between sensed signals and the corresponding location, we explore the spatial relation of beacons from the input data to estimate location. We propose a deep learning localization system, termed DRVAT, which is based on the distributed representation vector (DRV) and self-attention (AT) among the pairs of MAC-RSSI. First, we obtain DRVs which represent dense features in low dimensionality through pre-training on all MAC-RSSIs. Then we exploit self-attention mechanism to learn the latent spatial relation of beacons. Finally, MAC-RSSIs labeled with locations are used to fine-tune the model for estimating location. Localization accuracy results demonstrated the superior performance as compared with other positioning methods, and the visualization of DRV and attention mechanism are consistent with the spatial deployment of BLE. Haojun Ai, Xu Sun 0010, Jingjie Tao, Shengchen Li |
Int. J. Intell. Syst. | 5 |
| 2022 | Error model and simulation for multisource fusion indoor positioningabstractSeamless positioning services are of a critical concern in building smart cities. In a multisource fusion indoor positioning system, providing the guidance information for the deployment of positioning sources is a key technology, which can optimize the infrastructure resources to provide higher positioning accuracy. The error models of single-source positioning such as the received signal strength (RSS) fingerprint and the pedestrian dead reckoning (PDR) should be extended to meet the requirement of multisource indoor positioning for positioning error estimation. This paper proposes a model that combines the RSS fingerprint and PDR positioning error models for fusion positioning error simulation, which weights the PDR and RSS fingerprint positioning results and calculates the mean square error for the fusion positioning according to their positioning variances. This model is also used to establish an indoor positioning simulation system. To validate the proposed model, an experiment is performed which compared the actual positioning errors using the fusion positioning with the errors of the simulate model. The results show that the actual positioning error curves and the error curve predicted by the model are consistent. As a result, the proposed error model provides a solution for optimizing the deployment of positioning sources. Haojun Ai, Jingjie Tao, Shan Ai, Tianshui Xu, Ning Li 0050, Kaifeng Tang, Yuhong Yang 0001, Shengchen Li |
Int. J. Intell. Syst. | 9 |
| 2022 | A Deeply Supervised Attention Metric-Based Network and an Open Aerial Image Dataset for Remote Sensing Change DetectionabstractChange detection (CD) aims to identify surface changes from bitemporal images. In recent years, deep learning (DL)-based methods have made substantial breakthroughs in the field of CD. However, CD results can be easily affected by external factors, including illumination, noise, and scale, which leads to pseudo-changes and noise in the detection map. To deal with these problems and achieve more accurate results, a deeply supervised (DS) attention metric-based network (DSAMNet) is proposed in this article. A metric module is employed in DSAMNet to learn change maps by means of deep metric learning, in which convolutional block attention modules (CBAM) are integrated to provide more discriminative features. As an auxiliary, a DS module is introduced to enhance the feature extractor’s learning ability and generate more useful features. Moreover, another challenge encountered by data-driven DL algorithms is posed by the limitations in change detection datasets (CDDs). Therefore, we create a CD dataset, Sun Yat-Sen University (SYSU)-CD, for bitemporal image CD, which contains a total of 20 000 aerial image pairs of size$256\times256$. Experiments are conducted on both the CDD and the SYSU-CD dataset. Compared to other state-of-the-art methods, our network achieves the highest accuracy on both datasets, with an F1 of 93.69% on the CDD dataset and 78.18% on the SYSU-CD dataset. Qian Shi 0001, Mengxi Liu 0001, Shengchen Li, Xiaoping Liu 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Unsupervised Heart Abnormality Detection Based on Phonocardiogram Analysis with Beta Variational Auto-EncodersabstractHeart Sound (also known as phonocardiogram (PCG)) analysis, is a popular way that detects cardiovascular diseases (CVDs). Most PCG analysis uses supervised way, which demands both normal and abnormal samples. This paper proposes a method of unsupervised PCG analysis that uses beta variational auto-encoder (β – VAE) to model the normal PCG signals. The best performed model reaches an AUC (Area Under Curve) value of 0.91 in ROC (Receiver Operating Characteristic) test for PCG signals collected from the same source. Unlike majority of β – VAEs that are used as generative models, the best-performed β – VAE has a β value smaller than 1. This fact demonstrates that the resampling process helps the improvements on anomaly PCG detection through reconstruction loss worth a heavier weight. Further investigations suggest that anomaly score based on reconstruction loss may be better than anomaly scores based on latent vectors of samples in PCG analysis based on VAE systems. Shengchen Li, Ke Tian |
ICASSP | 1 |
| 2020 | Transfer Learning for Improving Singing-Voice Detection in Polyphonic Instrumental MusicabstractDetecting singing-voice in polyphonic instrumental music is critical to music information retrieval.To train a robust vocal detector, a large dataset marked with vocal or non-vocal label at frame-level is essential.However, frame-level labeling is time-consuming and labor expensive, resulting there is little well-labeled dataset available for singing-voice detection (S-VD).Hence, we propose a data augmentation method for S-VD by transfer learning.In this study, clean speech clips with voice activity endpoints and separate instrumental music clips are artificially added together to simulate polyphonic vocals to train a vocal /non-vocal detector.Due to the different articulation and phonation between speaking and singing, the vocal detector trained with the artificial dataset does not match well with the polyphonic music which is singing vocals together with the instrumental accompaniments.To reduce this mismatch, transfer learning is used to transfer the knowledge learned from the artificial speech-plus-music training set to a small but matched polyphonic dataset, i.e., singing vocals with accompaniments.By transferring the related knowledge to make up for the lack of well-labeled training data in S-VD, the proposed data augmentation method by transfer learning can improve S-VD performance with an F-score improvement from 89.5% to 93.2%. Yuanbo Hou, Frank K. Soong, Jian Luan 0001, Shengchen Li |
INTERSPEECH | 4 |
| 2020 | Peking Opera Synthesis via Duration Informed Attention NetworkabstractPeking Opera has been the most dominant form of Chinese performing art since around 200 years ago.A Peking Opera singer usually exhibits a very strong personal style via introducing improvisation and expressiveness on stage which leads the actual rhythm and pitch contour to deviate significantly from the original music score.This inconsistency poses a great challenge in Peking Opera singing voice synthesis from a music score.In this work, we propose to deal with this issue and synthesize expressive Peking Opera singing from the music score based on the Duration Informed Attention Network (DurIAN) framework.To tackle the rhythm mismatch, Lagrange multiplier is used to find the optimal output phoneme duration sequence with the constraint of the given note duration from music score.As for the pitch contour mismatch, instead of directly inferring from music score, we adopt a pseudo music score generated from the real singing and feed it as input during training.The experiments demonstrate that with the proposed system we can synthesize Peking Opera singing voice with high-quality timbre, pitch and expressiveness. Yusong Wu, Shengchen Li, Chengzhu Yu, Heng Lu 0004, Chao Weng, Dong Yu 0001 |
INTERSPEECH | 2 |
| 2019 | Sound Event Detection with Sequentially Labelled Data Based on Connectionist Temporal Classification and Unsupervised ClusteringabstractSound event detection (SED) methods typically rely on either strongly labelled data or weakly labelled data. As an alternative, sequentially labelled data (SLD) was proposed. In SLD, the events and the order of events in audio clips are known, without knowing the occurrence time of events. This paper proposes a connectionist temporal classification (CTC) based SED system that uses SLD instead of strongly labelled data, with a novel unsupervised clustering stage. Experiments on 41 classes of sound events show that the proposed two-stage method trained on SLD achieves performance comparable to the previous state-of-the-art SED system trained on strongly labelled data, and is far better than another state-of-the-art SED system trained on weakly labelled data, which indicates the effectiveness of the proposed two-stage method trained on SLD without any onset/offset time of sound events. Yuanbo Hou, Qiuqiang Kong, Shengchen Li, Mark D. Plumbley |
ICASSP | 3 |
| 2018 | Comparing the Influence of Depth and Width of Deep Neural Network Based on Fixed Number of Parameters for Audio Event DetectionabstractDeep Neural Network (DNN) is a basic method used for the rare Acoustic Event Detection (AED) in synthesised audio. The structure of DNNs including Multi-Layer Perceptron (MLP) and Recurrent Neural Network (RNN) for AED tasks has rather fewer hidden layers compared with computer vision systems. This paper tries to demonstrate that a DNN with more hidden layers does not necessarily guarantee a better performance in AED tasks. Taking the rare AED in synthesised audio with MLPs as an example and simulating a fixed budget of memory in an embedded system, various structures of MLPs are tested with fixed number of parameters engaged. Comparing the importance of neuron numbers in a hidden layer (i.e. the width of DNNs) and the importance of layer numbers in DNNs (i.e. the depth of DNNs) for AED tasks, the performance of the candidate DNN systems are evaluated by the event-based error rate. The results illustrate that a shallower network may outperform a deeper network when enough parameters are engaged and a larger number of parameters introduces a better performance in general. Shengchen Li |
ICASSP | 2 |