VLDB 2026 Research / reviewers in the wild / expert
Ruizhe Li 0001
dblp:14/10102-1
· DBLP profile ↗
21ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0003-2512-845XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DFWe: Efficient knowledge distillation of fine-tuned Whisper encoder for speech emotion recognition
Yujian Ma, Xianquan Jiang, Jinqiu Sang, Ruizhe Li 0001 |
Pattern Recognit. | 4 |
| 2025 | Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large LanguageabstractBo Zeng, Chenyang Lyu, Sinuo Liu, Mingyan Zeng, Minghao Wu, Xuanfan Ni, Tianqi Shi, Yu Zhao, Yefeng Liu, Chenyu Zhu, Ruizhe Li, Jiahui Geng, Qing Li, Yu Tong, Longyue Wang, Weihua Luo, Kaifu Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Chenyang Lyu, Sinuo Liu, Mingyan Zeng, Minghao Wu, Xuanfan Ni, Tianqi Shi, Yefeng Liu, Chenyu Zhu, Ruizhe Li 0001, Jiahui Geng, Longyue Wang, Weihua Luo, Kaifu Zhang |
ACL (1) | 11 |
| 2025 | FADERec: Fine-Grained Attribute Distillation Enhanced by Collaborative Fusion for LLM-Based Recommendation
Mingzhi Xu, Ruizhe Li 0001, Huimin Deng |
NLPCC (1) | 2 |
| 2025 | Aspect-aware semantic feature enhanced networks for multimodal aspect-based sentiment analysis
Liangqi Xie, Ruizhe Li 0001, Yongtao Yao, Huimin Deng |
J. Supercomput. | 3 |
| 2024 | GenTranslate: Large Language Models are Generative Multilingual Speech and Machine TranslatorsabstractYuchen Hu, Chen Chen, Chao-Han Huck Yang, Ruizhe Li, Dong Zhang, Zhehuai Chen, Eng Siong Chng. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Chen Chen 0075, Chao-Han Huck Yang, Ruizhe Li 0001, Zhehuai Chen, Chng Eng Siong |
ACL (1) | 4 |
| 2024 | Simulated Task Oriented Dialogues for Developing Versatile Conversational Agents
Xi Wang 0012, Procheta Sen, Ruizhe Li 0001, Emine Yilmaz |
ECIR (1) | 3 |
| 2024 | LLM-Generated Personalized Analogies to Foster AI Literacy in Adult NovicesabstractBroad Al literacy is essential in today's rapidly advancing technological landscape, extending beyond Al specialists to encompass the general public. However, the complexity of Al concepts poses significant barriers to learning for individuals without prior Al knowledge. While teaching through analogies is a well-recognized method to simplify complex information by connecting it to familiar concepts, adapting these analogies to match individual learner profiles remains a substantial challenge. This paper addresses this gap by proposing a novel method for personalizing educational analogies, enhancing the accessibility and engagement of AI concepts for a diverse audience. Our approach uses Large language models (LLMs) to dynamically tailor content to each learner's cognitive and cultural contexts, grounded in educational theories and practices. Utilizing a crowdsourced AIB testing framework through Prolific (N-60), this research contrasts conventional instructional methods with content incorporating LLM-enhanced personalized analogies. Data collection comprised pre- and post-tests, activity logs, and surveys featuring Likert-scale and open-ended questions. Quantitative analysis of key learning outcomes revealed significant improvements in comprehension and retention, evidenced by enhanced pre-and post- test scores (p < 0.01 and p < 0,05, respectively) and motivation, as indicated by increased engagement in survey responses (p < 0.05). Qualitative analysis revealed a need for more examples and visual aids to complement analogies and a preference for balancing analogies with detailed technical content. This study demonstrates the potential of Al-generated analogies to make complex Al concepts more accessible and engaging. Future research should refine analogy generation. incorporate multimedia elements, and explore long-term and cross-cultural impacts to further enhance Al education. Chen Cao 0005, Eason Chen, Zoe Fang, Lydia Y. Cao, Jionghao Lin, Ruizhe Li 0001 |
ICCE | 6 |
| 2024 | It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech RecognitionabstractRecent studies have successfully shown that large language models (LLMs) can be successfully used for generative error correction (GER) on top of the automatic speech recognition (ASR) output. Specifically, an LLM is utilized to carry out a direct mapping from the N-best hypotheses list generated by an ASR system to the predicted output transcription. However, despite its effectiveness, GER introduces extra data uncertainty since the LLM is trained without taking into account acoustic information available in the speech signal. In this work, we aim to overcome such a limitation by infusing acoustic information before generating the predicted transcription through a novel late fusion solution termed Uncertainty-Aware Dynamic Fusion (UADF). UADF is a multimodal fusion approach implemented into an auto-regressive decoding process and works in two stages: (i) It first analyzes and calibrates the token-level LLM decision, and (ii) it then dynamically assimilates the information from the acoustic modality. Experimental evidence collected from various ASR tasks shows that UADF surpasses existing fusion mechanisms in several ways. It yields significant improvements in word error rate (WER) while mitigating data uncertainty issues in LLM and addressing the poor generalization relied with sole modality during fusion. We also demonstrate that UADF seamlessly adapts to audio-visual speech recognition. Chen Chen 0075, Ruizhe Li 0001, Sabato Marco Siniscalchi, Chng Eng Siong, Chao-Han Huck Yang |
ICLR | 2 |
| 2024 | Large Language Models are Efficient Learners of Noise-Robust Speech RecognitionabstractRecent advances in large language models (LLMs) have promoted generative error correction (GER) for automatic speech recognition (ASR), which leverages the rich linguistic knowledge and powerful reasoning ability of LLMs to improve recognition results. The latest work proposes a GER benchmark with "HyPoradise" dataset to learn the mapping from ASR N-best hypotheses to ground-truth transcription by efficient LLM finetuning, which shows great effectiveness but lacks specificity on noise-robust ASR. In this work, we extend the benchmark to noisy conditions and investigate if we can teach LLMs to perform denoising for GER just like what robust ASR do, where one solution is introducing noise information as a conditioner into LLM. However, directly incorporating noise embeddings from audio encoder could harm the LLM tuning due to cross-modality gap. To this end, we propose to extract a language-space noise embedding from the N-best list to represent the noise conditions of source speech, which can promote the denoising process in GER. Furthermore, in order to enhance its representation ability of audio noise, we design a knowledge distillation (KD) approach via mutual information estimation to distill the real noise information in audio embeddings to our language embedding. Experiments on various latest LLMs demonstrate our approach achieves a new breakthrough with up to 53.9% correction improvement in terms of word error rate while with limited training data. Analysis shows that our language-space noise embedding can well represent the noise conditions of source speech, under which off-the-shelf LLMs show strong ability of language-space denoising. Chen Chen 0075, Chao-Han Huck Yang, Ruizhe Li 0001, Chao Zhang 0031, Chng Eng Siong |
ICLR | 4 |
| 2024 | Noise-aware Speech Enhancement using Diffusion Probabilistic ModelabstractWith recent advances of diffusion model, generative speech enhancement (SE) has attracted a surge of research interest due to its great potential for unseen testing noises. However, existing efforts mainly focus on inherent properties of clean speech, underexploiting the varying noise information in real world. In this paper, we propose a noise-aware speech enhancement (NASE) approach that extracts noise-specific information to guide the reverse process in diffusion model. Specifically, we design a noise classification (NC) model to produce acoustic embedding as a noise conditioner to guide the reverse denoising process. Meanwhile, a multi-task learning scheme is devised to jointly optimize SE and NC tasks to enhance the noise specificity of conditioner. NASE is shown to be a plug-and-play module that can be generalized to any diffusion SE models. Experiments on VB-DEMAND dataset show that NASE effectively improves multiple mainstream diffusion SE models, especially on unseen noises. Chen Chen 0075, Ruizhe Li 0001, Qiushi Zhu, Chng Eng Siong |
INTERSPEECH | 3 |
| 2024 | Leveraging Large Language Models Knowledge Enhancement Dual-Stage Fine-Tuning Framework for Recommendation
Yangyu Li, Ruizhe Li 0001, Huimin Deng |
NLPCC (2) | 4 |
| 2023 | MIR-GAN: Refining Frame-Level Modality-Invariant Representations with Adversarial Network for Audio-Visual Speech RecognitionabstractAudio-visual speech recognition (AVSR) attracts a surge of research interest recently by leveraging multimodal signals to understand human speech.Mainstream approaches addressing this task have developed sophisticated architectures and techniques for multi-modality fusion and representation learning.However, the natural heterogeneity of different modalities causes distribution gap between their representations, making it challenging to fuse them.In this paper, we aim to learn the shared representations across modalities to bridge their gap.Different from existing similar methods on other multimodal tasks like sentiment analysis, we focus on the temporal contextual dependencies considering the sequence-to-sequence task setting of AVSR.In particular, we propose an adversarial network to refine framelevel modality-invariant representations (MIR-GAN), which captures the commonality across modalities to ease the subsequent multimodal fusion process.Extensive experiments on public benchmarks LRS3 and LRS2 show that our approach outperforms the state-of-the-arts 1 . Chen Chen 0075, Ruizhe Li 0001, Heqing Zou, Chng Eng Siong |
ACL (1) | 3 |
| 2023 | Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech RecognitionabstractAudio-visual speech recognition (AVSR) provides a promising solution to ameliorate the noise-robustness of audio-only speech recognition with visual information.However, most existing efforts still focus on audio modality to improve robustness considering its dominance in AVSR task, with noise adaptation techniques such as front-end denoise processing.Though effective, these methods are usually faced with two practical challenges: 1) lack of sufficient labeled noisy audio-visual training data in some real-world scenarios and 2) less optimal model generality to unseen testing noises.In this work, we investigate the noiseinvariant visual modality to strengthen robustness of AVSR, which can adapt to any testing noises while without dependence on noisy training data, a.k.a., unsupervised noise adaptation.Inspired by human perception mechanism, we propose a universal viseme-phoneme mapping (UniVPM) approach to implement modality transfer, which can restore clean audio from visual signals to enable speech recognition under any noisy conditions.Extensive experiments on public benchmarks LRS3 and LRS2 show that our approach achieves the state-of-the-art under various noisy as well as clean conditions.In addition, we also outperform previous stateof-the-arts on visual speech recognition task 1 . Ruizhe Li 0001, Chen Chen 0075, Chengwei Qin, Qiushi Zhu, Chng Eng Siong |
ACL (1) | 2 |
| 2023 | Gradient Remedy for Multi-Task Learning in End-to-End Noise-Robust Speech RecognitionabstractSpeech enhancement (SE) is proved effective in reducing noise from noisy speech signals for downstream automatic speech recognition (ASR), where multi-task learning strategy is employed to jointly optimize these two tasks. However, the enhanced speech learned by SE objective may not always yield good ASR results. From the optimization view, there sometimes exists interference between the gradients of SE and ASR tasks, which could hinder the multi-task learning and finally lead to sub-optimal ASR performance. In this paper, we propose a simple yet effective approach called gradient remedy (GR) to solve interference between task gradients in noise-robust speech recognition, from perspectives of both angle and magnitude. Specifically, we first project the SE task's gradient onto a dynamic surface that is at acute angle to ASR gradient, in order to remove the conflict between them and assist in ASR optimization. Furthermore, we adaptively rescale the magnitude of two gradients to prevent the dominant ASR task from being misled by SE gradient. Experimental results show that the proposed approach well resolves the gradient interference and achieves relative word error rate (WER) reductions of 9.3% and 11.1% over multi-task learning baseline, on RATS and CHiME-4 datasets, respectively. Our code is available at GitHub1. Chen Chen 0075, Ruizhe Li 0001, Qiushi Zhu, Chng Eng Siong |
ICASSP | 3 |
| 2023 | Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech RecognitionabstractAudio-visual speech recognition (AVSR) research has gained a great success recently by improving the noise-robustness of audio-only automatic speech recognition (ASR) with noise-invariant visual information. However, most existing AVSR approaches simply fuse the audio and visual features by concatenation, without explicit interactions to capture the deep correlations between them, which results in sub-optimal multimodal representations for downstream speech recognition task. In this paper, we propose a cross-modal global interaction and local alignment (GILA) approach for AVSR, which captures the deep audio-visual (A-V) correlations from both global and local perspectives. Specifically, we design a global interaction model to capture the A-V complementary relationship on modality level, as well as a local alignment approach to model the A-V temporal consistency on frame level. Such a holistic view of cross-modal correlations enable better multimodal representations for AVSR. Experiments on public benchmarks LRS3 and LRS2 show that our GILA outperforms the supervised learning state-of-the-art. Code is at https://github.com/YUCHEN005/GILA. Ruizhe Li 0001, Chen Chen 0075, Heqing Zou, Qiushi Zhu, Chng Eng Siong |
IJCAI | 2 |
| 2021 | Affective Decoding for Empathetic Response GenerationabstractUnderstanding speaker's feelings and producing appropriate responses with emotion connection is a key communicative skill for empathetic dialogue systems.In this paper, we propose a simple technique called Affective Decoding for empathetic response generation.Our method can effectively incorporate emotion signals during each decoding step, and can additionally be augmented with an auxiliary dual emotion encoder, which learns separate embeddings for the speaker and listener given the emotion base of the dialogue.Extensive empirical studies show that our models are perceived to be more empathetic by human evaluations, in comparison to several strong mainstream methods for empathetic responding. Chengkun Zheng, Guanyi Chen, Chenghua Lin 0002, Ruizhe Li 0001 |
INLG | 4 |
| 2020 | Improving Variational Autoencoder for Text Modelling with Timestep-Wise RegularisationabstractThe Variational Autoencoder (VAE) is a popular and powerful model applied to text modelling to generate diverse sentences.However, an issue known as posterior collapse (or KL loss vanishing) happens when the VAE is used in text modelling, where the approximate posterior collapses to the prior, and the model will totally ignore the latent variables and be degraded to a plain language model during text generation.Such an issue is particularly prevalent when RNN-based VAE models are employed for text modelling.In this paper, we propose a simple, generic architecture called Timestep-Wise Regularisation VAE (TWR-VAE), which can effectively avoid posterior collapse and can be applied to any RNN-based VAE models.The effectiveness and versatility of our model are demonstrated in different tasks, including language modelling and dialogue response generation. Ruizhe Li 0001, Xiao Li 0041, Guanyi Chen, Chenghua Lin 0002 |
COLING | 1 |
| 2020 | DGST: a Dual-Generator Network for Text Style TransferabstractWe propose DGST, a novel and simple Dual-Generator network architecture for text Style Transfer.Our model employs two generators only, and does not rely on any discriminators or parallel corpus for training.Both quantitative and qualitative experiments on the Yelp and IMDb datasets show that our model gives competitive performance compared to several strong baselines with more complicated architecture designs. Xiao Li 0041, Guanyi Chen, Chenghua Lin 0002, Ruizhe Li 0001 |
EMNLP (1) | 4 |
| 2020 | Latent Space Factorisation and Manipulation via Matrix Subspace ProjectionabstractWe tackle the problem disentangling the latent space of an autoencoder in order to separate labelled attribute information from other characteristic information. This then allows us to change selected attributes while preserving other information. Our method, matrix subspace projection, is much simpler than previous approaches to latent space factorisation, for example not requiring multiple discriminators or a careful weighting among their loss functions. Furthermore our new model can be applied to autoencoders as a plugin, and works across diverse domains such as images or text. We demonstrate the utility of our method for attribute manipulation in autoencoders trained across varied domains, using both human evaluation and automated methods. The quality of generation of our new model (e.g. reconstruction, conditional generation) is highly competitive to a number of strong baselines. Xiao Li 0041, Chenghua Lin 0002, Ruizhe Li 0001, Chaozheng Wang, Frank Guerin |
ICML | 3 |
| 2019 | A Dual-Attention Hierarchical Recurrent Neural Network for Dialogue Act ClassificationabstractThis is a repository copy of A dual-attention hierarchical recurrent neural network for dialogue act classification. Ruizhe Li 0001, Chenghua Lin 0002, Matthew Collinson, Xiao Li 0041, Guanyi Chen |
CoNLL | 1 |
| 2019 | A Stable Variational Autoencoder for Text ModellingabstractVariational Autoencoder (VAE) is a powerful method for learning representations of highdimensional data.However, VAEs can suffer from an issue known as latent variable collapse (or KL loss vanishing), where the posterior collapses to the prior and the model will ignore the latent codes in generative tasks.Such an issue is particularly prevalent when employing VAE-RNN architectures for text modelling (Bowman et al., 2016).In this paper, we present a simple architecture called holistic regularisation VAE (HR-VAE), which can effectively avoid latent variable collapse.Compared to existing VAE-RNN architectures, we show that our model can achieve much more stable training process and can generate text with significantly better quality. Ruizhe Li 0001, Xiao Li 0041, Chenghua Lin 0002, Matthew Collinson, Rui Mao 0010 |
INLG | 1 |