EDBT 2026 Demo / reviewers in the wild / expert
Ruiyu Liang
dblp:42/6262
· DBLP profile ↗
18ranked-venue papers
2as first author
13since 2021 · last 2025
0000-0002-6813-4203ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Lightweight Attentive ConvNeXt-TCN for Causal Target Sound ExtractionabstractTarget sound extraction (TSE) aims to isolate specific sounds from complex acoustic mixtures. While various causal TSE models have been developed for real-time processing, most existing causal models operate in the time domain and do not effectively leverage frequency-domain information. This paper proposes a novel time-frequency domain model, ACN-TCN, which integrates the ConvNeXt design paradigm into the temporal convolutional network (TCN) to jointly model temporal and spectral information. Additionally, a convolutional self-attention mechanism is introduced to improve feature selection. Since shallow information is easily lost in a deep network, we incorporate a feature enhancement module to effectively integrate shallow features. Experimental results demonstrate that the proposed model improves signal-to-noise ratio (SNR) by 2.5 dB and scale-invariant signal-to-noise ratio (SI-SNR) by 3.4 dB compared to the state-of-the-art causal TSE method in single-target extraction task. Furthermore, ACN-TCN reduces the number of parameters by approximately 40% compared to previous models. We provide code and audio samples:https://github.com/Xiang-M-J/ACN-TCN MinJie Xiang, Ruiyu Liang, Ye Ni, Li Zhao 0003, Björn W. Schuller |
IEEE Signal Process. Lett. | 2 |
| 2024 | An improved TF-GSC for dual-microphone interference suppression in the specific direction
Cong Pang, Jingjie Fan, Ruiyu Liang, Li Zhao 0003, Jiaming Cheng 0005 |
Multim. Tools Appl. | 3 |
| 2024 | A strategy scheme of self-fitting based on gain adjustment for digital hearing aids
Yang Yang 0112, Ruxue Guo, Cairong Zou, Ruiyu Liang |
Multim. Tools Appl. | 4 |
| 2024 | Residual Fusion Probabilistic Knowledge Distillation for Speech EnhancementabstractIn recent years, a great deal of research has focused on in developing neural network (NN)-based speech enhancement (SE) models, which have achieved promising results. However, NN-based models typically require expensive computations to achieve remarkable performance, constraining their deployment in real-world scenarios, especially when hardware resources are limited or when latency requirements are strict. To reduce this computational burden, we propose a unified residual fusion probabilistic knowledge distillation (KD) method for the SE task, in which knowledge is transferred from a deep teacher to a shallower student model. Previous KD approaches commonly focused on narrowing the output distances between teachers and students, but research on the intermediate representation of these models is lacking. In this paper, we first study the cross-layer residual feature fusion strategy, which enables the student model to distill knowledge contained in multiple teacher layers from shallow to deep. Second, a frame weighting probabilistic distillation loss is proposed to assign more emphasis to frames containing essential information and preserve pairwise probabilistic similarities in the representation space. The proposed distillation framework is applied to the dual-path dilated convolutional recurrent network (DPDCRN), which won the championship of the SE track in the L3DAS23 challenge. Extensive experiments are conducted on single-channel and multichannel SE datasets. Objective evaluations show that the proposed KD strategy outperforms other distillation methods and considerably improves the enhancement effect of the low-complexity student model (with only 17% of the teacher's parameters). Jiaming Cheng 0005, Ruiyu Liang, Lin Zhou 0001, Li Zhao 0003, Chengwei Huang, Björn W. Schuller |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | A Non-Invasive Speech Quality Evaluation Algorithm for Hearing Aids With Multi-Head Self-Attention and Audiogram-Based FeaturesabstractThe speech quality delivered by hearing aids plays a crucial role in determining the acceptance and satisfaction of users. Compared with invasive speech quality evaluation methods that require pure signals as a reference, this paper proposes a non-invasive speech quality evaluation algorithm for hearing aids with multi-head self-attention and audiogram-based features. Initially, the audiogram of hearing-impaired individuals is extended along the frequency axis, enabling the speech quality evaluation model to learn the gain requirements specific to frequency bands for hearing-impaired individuals. Subsequently, the spectrogram is extracted from the speech signals to be evaluated. These features are combined with the transformed audiogram to create input features. To extract deep frame-level feature, a network employing multiple two-dimensional convolutional modules is utilized. Then, the temporal features are modeled using bidirectional long short-term memory networks (BiLSTM), while a multi-head self-attention mechanism is employed to integrate contextual information. This mechanism enables the model to focus on key frame information. Experimental results demonstrate that, compared to currently available advanced algorithms, the proposed network exhibits a higher correlation with the Hearing Aid Speech Quality Index (HASQI) and demonstrates robustness under various noise conditions. Ruiyu Liang, Jiaming Cheng 0005, Cong Pang, Björn W. Schuller |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Dual-Path Dilated Convolutional Recurrent Network with Group Attention for Multi-Channel Speech EnhancementabstractThis paper proposes a dual-path convolutional recurrent network with group attention for ICASSP Signal Processing Grand Challenge: L3DAS23 Challenge. We design a structure based on convolutional encoder-decoder, and frequency-time blocks based on group attention are introduced in the middle. The encoder is used to extract the local representation from the complex spectrum, the correlation along the frequency axis and the time axis are captured through groups of time-frequency processing modules and the key information in the feature flow is extracted by the group attention. As a result, our system ranks the 1st place of the 3D speech enhancement task in L3DAS23 Challenge, and significantly outperforms the baseline, while achieving 0.101 WER and 0.902 STOI on the blind test-set. Jiaming Cheng 0005, Cong Pang, Ruiyu Liang, Jingjie Fan, Li Zhao 0003 |
ICASSP | 3 |
| 2023 | Hearing loss classification algorithm based on the insertion gain of hearing aidabstractAbstract Hearing loss is one of the most prevalent chronic health problems worldwide and a common intervention is the wearing of hearing aids. However, the tedious fitting procedures and limited hearing experts pose restrictions for the popularity of hearing aids. This paper introduced a hearing loss classification method based on the insertion gain of hearing aids, which aims to simplify the fitting procedure and achieve a fitting-free effect of the hearing aid, in line with current research trends in key algorithms for fitting-free hearing aids. The proposed method innovatively combines the insertion gain of hearing aids with the covariates of patient’s gender, age, wearing history to form a new set of hearing loss vectors, and then classifies the hearing loss into six categories by unsupervised cluster analysis method. Each category of representative parameters characterizes a typical type of hearing loss, which can be used as the initial parameter to improve the efficiency of hearing aid fitting. Compared with the traditional audiogram classification method AMCLASS (Automated Audiogram Classification System), the proposed classification method reflect the actual hearing loss of hearing impaired patients better. Moreover, the effectiveness of the new classification method was verified by the comparison between the obtained six sets of representative insertion gains and the inferred hearing personalization information. Ruxue Guo, Ruiyu Liang, Qingyun Wang 0004, Cairong Zou |
Multim. Tools Appl. | 2 |
| 2023 | Multimodal emotion recognition from facial expression and speech based on feature fusion
Guichen Tang, Ke Li 0049, Ruiyu Liang, Li Zhao 0003 |
Multim. Tools Appl. | 4 |
| 2023 | Speech Denoising and Compensation for Hearing Aids Using an FTCRN-Based Metric GANabstractHearing aids aims to improve speech intelligibility for hearing impaired patients to levels comparable to those for normal hearing listeners. However, the interference of environmental noises greatly increase the difficulty of hearing loss compensation. Most related research only focuses on one aspect of noise reduction and hearing loss compensation. In this letter, we propose a metric generative adversarial framework based on a frequency-time convolution recurrent network for joint noise reduction and hearing loss compensation. The audiogram is extended along the frequency axis to form embedded features. A metric discriminator is introduced and the optimization of the generator is guided by an evaluation score related to hearing loss compensation. Additional perceptual-based losses are set to stabilize optimization. Experimental results show that the proposed method can better reduce noise and compensate for hearing loss compared with other algorithms. Jiaming Cheng 0005, Ruiyu Liang, Li Zhao 0003, Chengwei Huang, Björn W. Schuller |
IEEE Signal Process. Lett. | 2 |
| 2022 | Cross-Layer Similarity Knowledge Distillation for Speech EnhancementabstractSpeech enhancement (SE) algorithms based on deep neural networks (DNNs) often encounter challenges of limited hardware resources or strict latency requirements when deployed in real-world scenarios. However, a strong enhancement effect typically requires a large DNN. In this paper, a knowledge distillation framework for SE is proposed to compress the DNN model. We study the strategy of cross-layer connection paths, which fuses multi-level information from the teacher and transfers it to the student. To adapt to the SE task, we propose a frame-level similarity distillation loss. We apply this method to the deep complex convolution recurrent network (DCCRN) and make targeted adjustments. Experimental results show that the proposed method considerably improves the enhancement effect of the compressed DNN and outperforms other distillation methods. Jiaming Cheng 0005, Ruiyu Liang, Li Zhao 0003, Björn W. Schuller, Yiyuan Peng |
INTERSPEECH | 2 |
| 2021 | Real-time speech enhancement algorithm for transient noise suppression
Ruiyu Liang, Guichen Tang, Shinuo Sun |
Multim. Tools Appl. | 1 |
| 2021 | A frequency-domain nonlinear echo processing algorithm for high quality hands-free voice communication devices
Ruiyu Liang, Haicheng Liu |
Multim. Tools Appl. | 3 |
| 2021 | A Deep Adaptation Network for Speech Enhancement: Combining a Relativistic Discriminator With Multi-Kernel Maximum Mean DiscrepancyabstractIn deep-learning-based speech enhancement (SE) systems, trained models are often used to handle unseen noise types and language environments in real-life scenarios. However, since production environments differ from training conditions, mismatch problems arise that may cause a serious decrease in the performance of an SE system. In this study, a domain adaptive method combining two adaptation strategies is proposed to improve the generalization of unlabeled noisy speech. In the proposed encoder-decoder-based SE framework, a domain discriminator and a domain confusion adaptation layer are introduced to conduct adversarial training. The model has two main innovations. First, the algorithm optimizes adversarial training by introducing a relativistic discriminator that relies on relative values by applying the difference, thus avoiding possible bias and better reflecting domain differences. Second, the multi-kernel maximum mean discrepancy (MK-MMD) between domains is taken as the regularization term of the domain adversarial loss, thereby further decreasing the edge distribution distance between domains. The proposed model improves the adaptability to unseen noises by encouraging the feature encoder to generate domain-invariant features. The model was evaluated using cross-noise and cross-language-and-noise experiments, and the results show that the proposed method provides considerable improvements over the baseline without an adaptation in the perceptual evaluation of speech quality (PESQ), the short time objective intelligibility (STOI) and the frequency-weighted signal-to-noise ratio (FWSNR). Jiaming Cheng 0005, Ruiyu Liang, Zhenlin Liang, Li Zhao 0003, Chengwei Huang, Björn W. Schuller |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | DNN-based speech enhancement with self-attention on feature dimension
Jiaming Cheng 0005, Ruiyu Liang, Li Zhao 0003 |
Multim. Tools Appl. | 2 |
| 2019 | Improved Convolutional Neural Networks for Acoustic Event Classification
Guichen Tang, Ruiyu Liang, Yongqiang Bao |
Multim. Tools Appl. | 2 |
| 2019 | Speech Emotion Classification Using Attention-Based LSTMabstractAutomatic speech emotion recognition has been a research hotspot in the field of human-computer interaction over the past decade. However, due to the lack of research on the inherent temporal relationship of the speech waveform, the current recognition accuracy needs improvement. To make full use of the difference of emotional saturation between time frames, a novel method is proposed for speech recognition using frame-level speech features combined with attention-based long short-term memory (LSTM) recurrent neural networks. Frame-level speech features were extracted from waveform to replace traditional statistical features, which could preserve the timing relations in the original speech through the sequence of frames. To distinguish emotional saturation in different frames, two improvement strategies are proposed for LSTM based on the attention mechanism: first, the algorithm reduces the computational complexity by modifying the forgetting gate of traditional LSTM without sacrificing performance and second, in the final output of the LSTM, an attention mechanism is applied to both the time and the feature dimension to obtain the information related to the task, rather than using the output from the last iteration of the traditional algorithm. Extensive experiments on the CASIA, eNTERFACE, and GEMEP emotion corpora demonstrate that the performance of the proposed approach is able to outperform the state-of-the-art algorithms reported to date. Ruiyu Liang, Zhenlin Liang, Chengwei Huang, Cairong Zou, Björn W. Schuller |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Unsupervised learning of phonemes of whispered speech in a noisy environment based on convolutive non-negative matrix factorization
Jian Zhou 0006, Ruiyu Liang, Li Zhao 0003, Cairong Zou |
Inf. Sci. | 2 |
| 2011 | A vision inspection system for the surface defects of strongly reflected metal based on multi-class SVM
Xuewu Zhang 0001, Yan-Qiong Ding, Yan-yun Lv, Aiye Shi, Ruiyu Liang |
Expert Syst. Appl. | 5 |