EDBT 2026 Demo / reviewers in the wild / expert
Baoxiang Li
dblp:118/4815
· DBLP profile ↗
13ranked-venue papers
2as first author
10since 2021 · last 2024
0009-0009-4490-2157ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | CIF-T: A Novel CIF-Based Transducer Architecture for Automatic Speech RecognitionabstractRNN-T models are widely used in ASR, which rely on the RNN-T loss to achieve length alignment between input audio and target sequence. However, the implementation complexity and the alignment-based optimization target of RNN-T loss lead to computational redundancy and a reduced role for predictor network, respectively. In this paper, we propose a novel model named CIF-Transducer (CIF-T) which incorporates the Continuous Integrate-and-Fire (CIF) mechanism with the RNN-T model to achieve efficient alignment. In this way, the RNN-T loss is abandoned, thus bringing a computational reduction and allowing the predictor network a more significant role. We also introduce Funnel-CIF, Context Blocks, Unified Gating and Bilinear Pooling joint network, and auxiliary training strategy to further improve performance. Experiments on the 178-hour AISHELL-1 and 10000-hour WenetSpeech datasets show that CIF-T achieves state-of-the-art results with lower computational overhead compared to RNN-T models. Tian-Hao Zhang, Dinghao Zhou, Guiping Zhong, Jiaming Zhou 0001, Baoxiang Li |
ICASSP | 5 |
| 2024 | Balancing Multimodal Learning via Online Logit Modulation
Daoming Zong, Chaoyue Ding, Baoxiang Li, Jiakui Li, Ken Zheng |
IJCAI | 3 |
| 2023 | Stable Speech Emotion Recognition with Head-k-Pooling Loss
Chaoyue Ding, Jiakui Li, Daoming Zong, Baoxiang Li, Tian-Hao Zhang, Qunyan Zhou 0002 |
INTERSPEECH | 4 |
| 2023 | Unsupervised Adaptation with Quality-Aware Masking to Improve Target-Speaker Voice Activity Detection for Speaker Diarization
Shutong Niu, Jun Du 0002, Maokui He, Chin-Hui Lee 0001, Baoxiang Li, Jiakui Li |
INTERSPEECH | 5 |
| 2023 | AD-TUNING: An Adaptive CHILD-TUNING Approach to Efficient Hyperparameter Optimization of Child Networks for Speech Processing Tasks in the SUPERB Benchmark
Gaobin Yang, Jun Du 0002, Maokui He, Shutong Niu, Baoxiang Li, Jiakui Li, Chin-Hui Lee 0001 |
INTERSPEECH | 5 |
| 2023 | AcFormer: An Aligned and Compact Transformer for Multimodal Sentiment AnalysisabstractMultimodal Sentiment Analysis (MSA) is a popular research topic aimed at utilizing multimodal signals for understanding human emotions. The primary approach to solving this task is to develop complex fusion techniques. However, the heterogeneity and unaligned nature between modalities pose significant challenges to fusion. Additionally, existing methods lack consideration for the efficiency of modal fusion. To tackle these issues, we propose AcFormer, which contains two core ingredients: i) contrastive learning within and across modalities to explicitly align different modality streams before fusion; and ii) pivot attention for multimodal interaction/fusion. The former encourages positive triplets of image-audio-text to have similar representations in contrast to negative ones. The latter introduces attention pivots that can serve as cross-modal information bridges and limit cross-modal attention to a certain number of fusion pivot tokens. We evaluate AcFormer on multiple MSA tasks, including multimodal emotion recognition, humor detection, and sarcasm detection. Empirical evidence shows that AcFormer achieves the optimal performance with minimal computation cost compared to previous state-of-the-art methods. Our code is publicly available at https://github.com/dingchaoyue/AcFormer. Daoming Zong, Chaoyue Ding, Baoxiang Li, Jiakui Li, Ken Zheng, Qunyan Zhou 0002 |
ACM Multimedia | 3 |
| 2023 | Building Robust Multimodal Sentiment Recognition via a Simple yet Effective Multimodal TransformerabstractIn this paper, we present the solutions to the MER-MULTI and MER-NOISE sub-challenges of the Multimodal Emotion Recognition Challenge (MER 2023). For the tasks MER-MULTI and MER-NOISE, participants are required to recognize both discrete and dimensional emotions. Particularly, in MER-NOISE, the test videos are corrupted with noise, necessitating the consideration of modality robustness. Our empirical findings indicate that different modalities contribute differently to the tasks, with a significant impact from the audio and visual modalities, while the text modality plays a weaker role in emotion prediction. To facilitate subsequent multimodal fusion, and considering that language information is implicitly embedded in large pre-trained speech models, we have made the deliberate choice to abandon the text modality and solely utilize visual and acoustic modalities for these sub-challenges. To address the potential underfitting of individual modalities during multimodal training, we propose to jointly train all modalities via a weighted blending of supervision signals. Furthermore, to enhance the robustness of our model, we employ a range of data augmentation techniques at the image level, waveform level, and spectrogram level. Experimental results show that our model ranks 1st in both MER-MULTI (0.7005) and MER-NOISE (0.6846) sub-challenges, validating the effectiveness of our method. Our code is publicly available at https://github.com/dingchaoyue/Multimodal-Emotion-Recognition-MER-and-MuSe-2023-Challenges. Daoming Zong, Chaoyue Ding, Baoxiang Li, Dinghao Zhou, Jiakui Li, Ken Zheng, Qunyan Zhou 0002 |
ACM Multimedia | 3 |
| 2022 | LETR: A Lightweight and Efficient Transformer for Keyword SpottingabstractTransformer recently has achieved impressive success in a number of domains, including machine translation, image recognition, and speech recognition. Most of the previous work on Keyword Spotting (KWS) is built upon convolutional or recurrent neural networks. In this paper, we explore a family of Transformer architectures for keyword spotting, optimizing the trade-off between accuracy and efficiency in a high-speed regime. We also studied the effectiveness and summarized the principles of applying key components in vision Transformers to KWS, including patch embedding, position encoding, attention mechanism, and class token. On top of the findings, we propose the LeTR: a lightweight and highly efficient Transformer for KWS. We consider different efficiency measures on different edge devices so as to reflect a wide range of application scenarios best. Experimental results on two common benchmarks demonstrate that LeTR has achieved state-of-the-art results over competing methods with respect to the speed/accuracy trade-off. Kevin Ding, Martin Zong, Jiakui Li, Baoxiang Li |
ICASSP | 4 |
| 2022 | A polyphone BERT for Polyphone Disambiguation in Mandarin ChineseabstractGrapheme-to-phoneme (G2P) conversion is an indispensable part of the Chinese Mandarin text-to-speech (TTS) system, and the core of G2P conversion is to solve the problem of polyphone disambiguation, which is to pick up the correct pronunciation for several candidates for a Chinese polyphonic character.In this paper, we propose a Chinese polyphone BERT model to predict the pronunciations of Chinese polyphonic characters.Firstly, we create 741 new Chinese monophonic characters from 354 source Chinese polyphonic characters by pronunciation.Then we get a Chinese polyphone BERT by extending a pre-trained Chinese BERT with 741 new Chinese monophonic characters and adding a corresponding embedding layer for new tokens, which is initialized by the embeddings of source Chinese polyphonic characters.In this way, we can turn the polyphone disambiguation task into a pre-training task of the Chinese polyphone BERT.Experimental results demonstrate the effectiveness of the proposed model, and the polyphone BERT model obtain 2% (from 92.1% to 94.1%) improvement of average accuracy compared with the BERT-based classifier model, which is the prior state-of-the-art in polyphone disambiguation. Ken Zheng, Xiaoxu Zhu, Baoxiang Li |
INTERSPEECH | 4 |
| 2022 | Speed-Robust Keyword Spotting Via Soft Self-Attention on Multi-Scale FeaturesabstractIn this work, we focus on the robustness of keyword spotting (KWS) at various speech speeds. First, to enable small-footprint KWS, we graft a depthwise separable convolution and a dilated temporal convolution to build our basic model block. Second, to make KWS desensitized to speech rate, a simple yet effective soft self-attention is proposed to operate between different hierarchical features, offering our model the ability to be aware of the speaker's speech rate and to dynamically integrate multi-scale features from varying sizes of receptive fields. Besides, we construct two frame-level annotated Chinese intelligent in-car speech commands datasets, termed Car-C1 and Car-C2, for model evaluation. Experimental results show that our model achieves the highest accuracy on both datasets with a comparable number of model parameters and computational cost. Meanwhile, the sensitivity analysis suggests that our model outperforms baselines when tested on a fast corpus and a slow corpus. Chaoyue Ding, Jiakui Li, Martin Zong, Baoxiang Li |
SLT | 4 |
| 2019 | Route Planning for a Fleet of Electric Vehicles with Waiting Times at Charging Stations
Baoxiang Li, Shashi Shekhar Jha, Hoong Chuin Lau |
EvoCOP | 1 |
| 2012 | Robust lyric search based on weighted syllable confusion matrix
Baoxiang Li, Fengxiang Chang, Qiang Wang 0048, Gang Liu 0008, Jun Guo 0002 |
ICPR | 1 |
| 2012 | Tempo variation based multilayer filters for query by humming
Qiang Wang 0048, Baoxiang Li, Gang Liu 0008, Jun Guo 0002 |
ICPR | 3 |