Yong Ren 0006

dblp:00/440-6 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
16since 2021 · last 2026
0009-0000-9015-000XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021
YearPublicationVenuePosition
2026 MERBench: A Unified Evaluation Benchmark for Multimodal Emotion Recognition
abstract
Multimodal emotion recognition plays a vital role in enhancing user experience in human-computer interaction. Over the past few decades, researchers have developed a range of algorithms and made remarkable progress. While each approach demonstrates certain advantages, inconsistent choices in feature extraction methods, evaluation protocols, and experimental settings have hindered fair comparisons among them. These inconsistencies significantly impede the advancement of the field. To address this issue, we introduce MERBench, a unified evaluation benchmark for multimodal emotion recognition. Our goal is to assess the contributions of several key techniques commonly used in prior studies, such as feature selection, multimodal fusion, robustness analysis, fine-tuning, and pre-training. We believe this work offers clear and comprehensive guidance for future research. Based on the evaluation results of MERBench, we further point out some promising research directions. In addition, we present a new emotion dataset, MER2023, specifically designed for the Chinese language environment. This dataset serves as a benchmark for research in multi-label learning, noise robustness, and semi-supervised learning.
Zheng Lian 0004, Licai Sun, Yong Ren 0006, Haiyang Sun 0004, Lan Chen 0005, Bin Liu 0041, Jianhua Tao 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Region-Based Optimization in Continual Learning for Audio Deepfake Detection
abstract
Rapid advancements in speech synthesis and voice conversion bring convenience but also new security risks, creating an urgent need for effective audio deepfake detection. Although current models perform well, their effectiveness diminishes when confronted with the diverse and evolving nature of real-world deepfakes. To address this issue, we propose a continual learning method named Region-Based Optimization (RegO) for audio deepfake detection. Specifically, we use the Fisher information matrix to measure important neuron regions for real and fake audio detection, dividing them into four regions. First, we directly fine-tune the less important regions to quickly adapt to new tasks. Next, we apply gradient optimization in parallel for regions important only to real audio detection, and in orthogonal directions for regions important only to fake audio detection. For regions that are important to both, we use sample proportion-based adaptive gradient optimization. This region-adaptive optimization ensures an appropriate trade-off between memory stability and learning plasticity. Additionally, to address the increase of redundant neurons from old tasks, we further introduce the Ebbinghaus forgetting mechanism to release them, thereby promoting the model’s ability to learn more generalized discriminative features. Experimental results show our method achieves a 21.3 percent improvement in EER over the state-of-the-art continual learning approach RWM for audio deepfake detection. Moreover, the effectiveness of RegO extends beyond the audio deepfake detection domain, showing potential significance in other tasks, such as image recognition.
Yujie Chen 0006, Jiangyan Yi, Cunhang Fan, Jianhua Tao 0001, Yong Ren 0006, Siding Zeng, Chu Yuan Zhang, Xinrui Yan, Jun Xue 0001, Chenglong Wang 0001, Zhao Lv, Xiaohui Zhang 0006
AAAI5
2025 Enhancing Multimodal Continual Instruction Tuning with BranchLoRA
abstract
Multimodal Continual Instruction Tuning (MCIT) aims to finetune Multimodal Large Language Models (MLLMs) to continually align with human intent across sequential tasks. Existing approaches often rely on the Mixture-of-Experts (MoE) LoRA framework to preserve previous instruction alignments. However, these methods are prone to Catastrophic Forgetting (CF), as they aggregate all LoRA blocks via simple summation, which compromises performance over time. In this paper, we identify a critical parameter inefficiency in the MoELoRA framework within the MCIT context. Based on this insight, we propose BranchLoRA, an asymmetric framework to enhance both efficiency and performance. To mitigate CF, we introduce a flexible tuning-freezing mechanism within BranchLoRA, enabling branches to specialize in intra-task knowledge while fostering inter-task collaboration. Moreover, we incrementally incorporate task-specific routers to ensure an optimal branch distribution over time, rather than favoring the most recent task. To streamline inference, we introduce a task selector that automatically routes test inputs to the appropriate router without requiring task identity. Extensive experiments on the latest MCIT benchmark demonstrate that BranchLoRA significantly outperforms MoELoRA and maintains its superiority across various MLLM sizes.
Duzhen Zhang, Yong Ren 0006, Zhongzhi Li, Yahan Yu, Jiahua Dong 0001, Chenxing Li, Zhilong Ji, Jinfeng Bai
ACL (1)2
2025 STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
abstract
Visual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmonious audio in generated video limits its broader applications. In this paper, we propose Semantic and Temporal Aligned Video-to-Audio (STA-V2A), an approach that enhances audio generation from videos by extracting both local temporal and global semantic video features and combining these refined video features with text as cross-modal guidance. To address the issue of information redundancy in videos, we propose an onset prediction pretext task for local temporal feature extraction and an attentive pooling module for global semantic feature extraction. To supplement the insufficient semantic information in videos, we propose a Latent Diffusion Model with Text-to-Audio priors initialization and cross-modal guidance. We also introduce Audio-Audio Align, a new metric to assess audio-temporal alignment. Subjective and objective metrics demonstrate that our method surpasses existing Video-to-Audio models in generating audio with better quality, semantic consistency, and temporal alignment. The ablation experiment validated the effectiveness of each module. Audio samples are available at https://y-ren16.github.io/STAV2A.
Yong Ren 0006, Chenxing Li, Manjie Xu, Rilin Chen, Dong Yu 0001
ICASSP1
2025 WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
abstract
Recent advances in speech spoofing necessitate stronger verification mechanisms in neural speech codecs to ensure authenticity. Current methods embed numerical watermarks before compression and extract them from reconstructed speech for verification, but face limitations such as separate training processes for the watermark and codec, and insufficient cross-modal information integration, leading to reduced watermark imperceptibility, extraction accuracy, and capacity. To address these issues, we propose WMCodec, the first neural speech codec to jointly train compression-reconstruction and watermark embedding-extraction in an end-to-end manner, optimizing both imperceptibility and extractability of the watermark. Furthermore, We design an iterative Attention Imprint Unit (AIU) for deeper feature integration of watermark and speech, reducing the impact of quantization noise on the watermark. Experimental results show WMCodec outperforms AudioSeal with Encodec in most quality metrics for watermark imperceptibility and consistently exceeds both AudioSeal with Encodec and reinforced TraceableSpeech in extraction accuracy of watermark. At bandwidth of 6 kbps with a watermark capacity of 16 bps, WMCodec maintains over 99% extraction accuracy under common attacks, demonstrating strong robustness.
Junzuo Zhou, Jiangyan Yi, Yong Ren 0006, Jianhua Tao 0001, Tao Wang 0074, Chuyuan Zhang
ICASSP3
2025 AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models
abstract
The emergence of multimodal large language models (MLLMs) advances multimodal emotion recognition (MER) to the next level—from naive discriminative tasks to complex emotion understanding with advanced video understanding abilities and natural language description. However, the current community suffers from a lack of large-scale datasets with intensive, descriptive emotion annotations, as well as a multimodal-centric framework to maximize the potential of MLLMs for emotion understanding. To address this, we establish a new benchmark for MLLM-based emotion understanding with a novel dataset (MER-Caption) and a new model (AffectGPT). Utilizing our model-based crowd-sourcing data collection strategy, we construct the largest descriptive emotion dataset to date (by far), featuring over 2K fine-grained emotion categories across 115K samples. We also introduce the AffectGPT model, designed with pre-fusion operations to enhance multimodal integration. Finally, we present MER-UniBench, a unified benchmark with evaluation metrics tailored for typical MER tasks and the free-form, natural language output style of MLLMs. Extensive experimental results show AffectGPT's robust performance across various MER tasks. We have released both the code and the dataset to advance research and development in emotion understanding: https://github.com/zeroQiaoba/AffectGPT.
Zheng Lian 0004, Haoyu Chen 0001, Lan Chen 0005, Haiyang Sun 0004, Licai Sun, Yong Ren 0006, Zebang Cheng, Bin Liu 0041, Rui Liu 0008, Xiaojiang Peng, Jiangyan Yi, Jianhua Tao 0001
ICML6
2025 Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
Yong Ren 0006, Chenxing Li, Duzhen Zhang, Yujie Chen 0006, Manjie Xu, Ruibo Fu, Shan Yang 0001, Dong Yu 0001
INTERSPEECH1
2025 Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
Chenxing Li, Yong Ren 0006, Yujie Chen 0006, Ruibo Fu, Shan Yang 0001, Dong Yu 0001
INTERSPEECH3
2025 Towards Diverse and Efficient Audio Captioning via Diffusion Models
Manjie Xu, Chenxing Li, Yong Ren 0006, Ruibo Fu, Dong Yu 0001
INTERSPEECH3
2025 ALLM4ADD: Unlocking the Capabilities of Audio Large Language Models for Audio Deepfake Detection
abstract
Audio deepfake detection (ADD) has grown increasingly important due to the rise of high-fidelity audio generative models and their potential for misuse. Given that audio large language models (ALLMs) have made significant progress in various audio processing tasks, a heuristic question arises: Can ALLMs be leveraged to solve ADD?. In this paper, we first conduct a comprehensive zero-shot evaluation of ALLMs on ADD, revealing their ineffectiveness. To this end, we propose ALLM4ADD, an ALLM-driven framework for ADD. Specifically, we reformulate ADD task as an audio question answering problem, prompting the model with the question: ''Is this audio fake or real?''. We then perform supervised fine-tuning to enable the ALLM to assess the authenticity of query audio. Extensive experiments are conducted to demonstrate that our ALLM-based method can achieve superior performance in fake audio detection, particularly in data-scarce scenarios. As a pioneering study, we anticipate that this work will inspire the research community to leverage ALLMs to develop more effective ADD systems. Code is available at https://github.com/ucas-hao/qwen_audio_for_add.git.
Jiangyan Yi, Chenglong Wang 0001, Jianhua Tao 0001, Zheng Lian 0004, Yong Ren 0006, Yujie Chen 0006, Zhengqi Wen
ACM Multimedia7
2025 Information-theoretic complementary prompts for improved continual text classification
Duzhen Zhang, Yong Ren 0006, Chenxing Li, Dong Yu 0001, Tielin Zhang
Neural Networks2
2024 SpikeVoice: High-Quality Text-to-Speech Via Efficient Spiking Neural Network
abstract
Brain-inspired Spiking Neural Network (SNN) has demonstrated its effectiveness and efficiency in vision, natural language, and speech understanding tasks, indicating their capacity to “see”, “listen”, and “read”. In this paper, we design SpikeVoice, which performs high-quality Text-To-Speech (TTS) via SNN, to explore the potential of SNN to “speak”. A major obstacle to using SNN for such generative tasks lies in the demand for models to grasp long-term dependencies. The serial nature of spiking neurons, however, leads to the invisibility of information at future spiking time steps, limiting SNN models to capture sequence dependencies solely within the same time step. We term this phenomenon “partial-time dependency”. To address this issue, we introduce Spiking Temporal-Sequential Attention (STSA) in the SpikeVoice. To the best of our knowledge, SpikeVoice is the first TTS work in the SNN field. We perform experiments using four well-established datasets that cover both Chinese and English languages, encompassing scenarios with both single-speaker and multi-speaker configurations. The results demonstrate that SpikeVoice can achieve results comparable to Artificial Neural Networks (ANN) with only 10.5% energy consumption of ANN. Both our demo and code are available as supplementary material.
Jiahong Zhang, Yong Ren 0006, Man Yao, Richard D. Shang, Bo Xu 0002, Guoqi Li 0002
ACL (1)3
2024 Fewer-Token Neural Speech Codec with Time-Invariant Codes
abstract
Language model based text-to-speech (TTS) models, like VALL-E, have gained attention for their outstanding in-context learning capability in zero-shot scenarios. Neural speech codec is a critical component of these models, which can convert speech into discrete token representations. However, excessive token sequences from the codec may negatively affect prediction accuracy and restrict the progression of Language model based TTS models. To address this issue, this paper proposes a novel neural speech codec with time-invariant codes named TiCodec. By encoding and quantizing time-invariant information into a separate code, TiCodec can reduce the amount of frame-level information that needs encoding, effectively decreasing the number of tokens as codes of speech. Furthermore, this paper introduces a time-invariant encoding consistency loss to enhance the consistency of time-invariant code within an utterance, which can benefit the zero-shot TTS task. Experimental results demonstrate that TiCodec can not only enhance the quality of reconstruction speech with fewer tokens but also increase the similarity and naturalness, as well as reduce the word error rate of the synthesized speech by the TTS model. The code is publicly available at https://github.com/y-ren16/TiCodec.
Yong Ren 0006, Tao Wang 0074, Jiangyan Yi, Jianhua Tao 0001, Chu Yuan Zhang, Junzuo Zhou
ICASSP1
2024 Residual Speaker Representation for One-Shot Voice Conversion
abstract
International audience
Jiangyan Yi, Tao Wang 0074, Yong Ren 0006, Rongxiu Zhong, Zhengqi Wen, Jianhua Tao 0001
INTERSPEECH4
2024 TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
Junzuo Zhou, Jiangyan Yi, Tao Wang 0074, Jianhua Tao 0001, Ye Bai 0001, Chu Yuan Zhang, Yong Ren 0006, Zhengqi Wen
INTERSPEECH7
2024 Utilizing Speaker Profiles for Impersonation Audio Detection
abstract
Fake audio detection is an emerging active topic. A growing number of literatures have aimed to detect fake utterance, which are mostly generated by Text-to-speech (TTS) or voice conversion (VC). However, countermeasures against impersonation remain an underexplored area. Impersonation is a fake type that involves an imitator replicating specific traits and speech style of a target speaker. Unlike TTS and VC, which often leave digital traces or signal artifacts, impersonation involves live human beings producing entirely natural speech, rendering the detection of impersonation audio a challenging task. Thus, we propose a novel method that integrates speaker profiles into the process of impersonation audio detection. Speaker profiles are inherent characteristics that are challenging for impersonators to mimic accurately, such as speaker's age, job. We aim to leverage these features to extract discriminative information for detecting impersonation audio. Moreover, there is no large impersonated speech corpora available for quantitative study of impersonation impacts. To address this gap, we further design the first large-scale, diverse-speaker Chinese impersonation dataset, named ImPersonation Audio Detection (IPAD), to advance the community's research on impersonation audio detection. We evaluate several existing fake audio detection methods on our proposed dataset IPAD, demonstrating its necessity and the challenges. Additionally, our findings reveal that incorporating speaker profiles can significantly enhance the model's performance in detecting impersonation audio.
Jiangyan Yi, Chenglong Wang 0001, Yong Ren 0006, Jianhua Tao 0001, Xinrui Yan, Yujie Chen 0006, Xiaohui Zhang 0006
ACM Multimedia4