EDBT 2026 Demo / reviewers in the wild / expert
Jiakui Li
dblp:189/4442
· DBLP profile ↗
9ranked-venue papers
0as first author
8since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Balancing Multimodal Learning via Online Logit Modulation
Daoming Zong, Chaoyue Ding, Baoxiang Li, Jiakui Li, Ken Zheng |
IJCAI | 4 |
| 2023 | Stable Speech Emotion Recognition with Head-k-Pooling Loss
Chaoyue Ding, Jiakui Li, Daoming Zong, Baoxiang Li, Tian-Hao Zhang, Qunyan Zhou 0002 |
INTERSPEECH | 2 |
| 2023 | Unsupervised Adaptation with Quality-Aware Masking to Improve Target-Speaker Voice Activity Detection for Speaker Diarization
Shutong Niu, Jun Du 0002, Maokui He, Chin-Hui Lee 0001, Baoxiang Li, Jiakui Li |
INTERSPEECH | 6 |
| 2023 | AD-TUNING: An Adaptive CHILD-TUNING Approach to Efficient Hyperparameter Optimization of Child Networks for Speech Processing Tasks in the SUPERB Benchmark
Gaobin Yang, Jun Du 0002, Maokui He, Shutong Niu, Baoxiang Li, Jiakui Li, Chin-Hui Lee 0001 |
INTERSPEECH | 6 |
| 2023 | AcFormer: An Aligned and Compact Transformer for Multimodal Sentiment AnalysisabstractMultimodal Sentiment Analysis (MSA) is a popular research topic aimed at utilizing multimodal signals for understanding human emotions. The primary approach to solving this task is to develop complex fusion techniques. However, the heterogeneity and unaligned nature between modalities pose significant challenges to fusion. Additionally, existing methods lack consideration for the efficiency of modal fusion. To tackle these issues, we propose AcFormer, which contains two core ingredients: i) contrastive learning within and across modalities to explicitly align different modality streams before fusion; and ii) pivot attention for multimodal interaction/fusion. The former encourages positive triplets of image-audio-text to have similar representations in contrast to negative ones. The latter introduces attention pivots that can serve as cross-modal information bridges and limit cross-modal attention to a certain number of fusion pivot tokens. We evaluate AcFormer on multiple MSA tasks, including multimodal emotion recognition, humor detection, and sarcasm detection. Empirical evidence shows that AcFormer achieves the optimal performance with minimal computation cost compared to previous state-of-the-art methods. Our code is publicly available at https://github.com/dingchaoyue/AcFormer. Daoming Zong, Chaoyue Ding, Baoxiang Li, Jiakui Li, Ken Zheng, Qunyan Zhou 0002 |
ACM Multimedia | 4 |
| 2023 | Building Robust Multimodal Sentiment Recognition via a Simple yet Effective Multimodal TransformerabstractIn this paper, we present the solutions to the MER-MULTI and MER-NOISE sub-challenges of the Multimodal Emotion Recognition Challenge (MER 2023). For the tasks MER-MULTI and MER-NOISE, participants are required to recognize both discrete and dimensional emotions. Particularly, in MER-NOISE, the test videos are corrupted with noise, necessitating the consideration of modality robustness. Our empirical findings indicate that different modalities contribute differently to the tasks, with a significant impact from the audio and visual modalities, while the text modality plays a weaker role in emotion prediction. To facilitate subsequent multimodal fusion, and considering that language information is implicitly embedded in large pre-trained speech models, we have made the deliberate choice to abandon the text modality and solely utilize visual and acoustic modalities for these sub-challenges. To address the potential underfitting of individual modalities during multimodal training, we propose to jointly train all modalities via a weighted blending of supervision signals. Furthermore, to enhance the robustness of our model, we employ a range of data augmentation techniques at the image level, waveform level, and spectrogram level. Experimental results show that our model ranks 1st in both MER-MULTI (0.7005) and MER-NOISE (0.6846) sub-challenges, validating the effectiveness of our method. Our code is publicly available at https://github.com/dingchaoyue/Multimodal-Emotion-Recognition-MER-and-MuSe-2023-Challenges. Daoming Zong, Chaoyue Ding, Baoxiang Li, Dinghao Zhou, Jiakui Li, Ken Zheng, Qunyan Zhou 0002 |
ACM Multimedia | 5 |
| 2022 | LETR: A Lightweight and Efficient Transformer for Keyword SpottingabstractTransformer recently has achieved impressive success in a number of domains, including machine translation, image recognition, and speech recognition. Most of the previous work on Keyword Spotting (KWS) is built upon convolutional or recurrent neural networks. In this paper, we explore a family of Transformer architectures for keyword spotting, optimizing the trade-off between accuracy and efficiency in a high-speed regime. We also studied the effectiveness and summarized the principles of applying key components in vision Transformers to KWS, including patch embedding, position encoding, attention mechanism, and class token. On top of the findings, we propose the LeTR: a lightweight and highly efficient Transformer for KWS. We consider different efficiency measures on different edge devices so as to reflect a wide range of application scenarios best. Experimental results on two common benchmarks demonstrate that LeTR has achieved state-of-the-art results over competing methods with respect to the speed/accuracy trade-off. Kevin Ding, Martin Zong, Jiakui Li, Baoxiang Li |
ICASSP | 3 |
| 2022 | Speed-Robust Keyword Spotting Via Soft Self-Attention on Multi-Scale FeaturesabstractIn this work, we focus on the robustness of keyword spotting (KWS) at various speech speeds. First, to enable small-footprint KWS, we graft a depthwise separable convolution and a dilated temporal convolution to build our basic model block. Second, to make KWS desensitized to speech rate, a simple yet effective soft self-attention is proposed to operate between different hierarchical features, offering our model the ability to be aware of the speaker's speech rate and to dynamically integrate multi-scale features from varying sizes of receptive fields. Besides, we construct two frame-level annotated Chinese intelligent in-car speech commands datasets, termed Car-C1 and Car-C2, for model evaluation. Experimental results show that our model achieves the highest accuracy on both datasets with a comparable number of model parameters and computational cost. Meanwhile, the sensitivity analysis suggests that our model outperforms baselines when tested on a fast corpus and a slow corpus. Chaoyue Ding, Jiakui Li, Martin Zong, Baoxiang Li |
SLT | 2 |
| 2016 | Sonar-based place recognition using joint sparse coding methodabstractThe problem of place recognition is central to robot navigation. The robot needs to be able to recognize or at least to be able to estimate the likelihood that it has been at a place before when it has returned to a previously visited place. We cast the place recognition problem as one of classifying among multiple linear regression models, and argue that new theory from sparse signal representation offers the key to addressing the problem. In this paper, a joint kernel sparse coding model is developed to tackle the multivariate sonar samples place recognition problem. The experimental results show that the joint sparse coding achieves better performance than 1-Nearest Neighborhood (1-NN) method. Xiangmei Zheng, Huaping Liu 0001, Fuchun Sun 0001, Meng Gao 0002, Jiakui Li |
IJCNN | 5 |