EDBT 2026 Demo / reviewers in the wild / expert
Chen Zhang 0020
dblp:94/4084-20
· DBLP profile ↗
51ranked-venue papers
16as first author
45since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 10 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VoiceBench: Benchmarking LLM-Based Voice AssistantsabstractAbstract Recent advancements in large language models (LLMs) like GPT-4o have enabled real-time speech interactions through LLM-based voice assistants, offering an improved user experience over text-based interactions. However, a suitable benchmark to rigorously evaluate such speech interactions systems is currently lacking. To bridge this gap, we introduce VoiceBench, the first benchmark specifically designed to assess LLM-based voice assistants. VoiceBench comprises 6,783 synthetic and real spoken instructions recorded from diverse speakers across eight distinct tasks. These instructions are meticulously crafted to assess three crucial capability areas: general knowledge, instruction-following, and safety compliance. Furthermore, VoiceBench systematically incorporates realistic variations common in spoken interactions, including differences in speaker characteristics (e.g., accents), heterogeneous environmental conditions (e.g., reverberation), and content complexities such as mispronunciations. Extensive experiments reveal the limitations of current LLM-based voice assistant models and offer valuable insights for future research and development in this field.1 Yiming Chen 0010, Xianghu Yue, Chen Zhang 0020, Xiaoxue Gao, Robby T. Tan, Haizhou Li 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2025 | Towards the Law of Capacity Gap in Distilling Language ModelsabstractLanguage model (LM) distillation aims at distilling the knowledge in a large teacher LM to a small student one. As a critical issue facing LM distillation, a superior student often arises from a teacher of a relatively small scale instead of a larger one, especially in the presence of substantial capacity gap between the teacher and student. This issue, often referred to as the curse of capacity gap, suggests that there is likely an optimal teacher yielding the best-performing student along the scaling course of the teacher. Consequently, distillation trials on teachers of a wide range of scales are called for to determine the optimal teacher, which becomes computationally intensive in the context of large LMs (LLMs). This paper addresses this critical bottleneck by providing the law of capacity gap inducted from a preliminary study on distilling a broad range of small-scale (<3B) LMs, where the optimal teacher consistently scales linearly with the student scale across different model and data scales. By extending the law to LLM distillation on a larger scale (7B), we succeed in obtaining versatile LLMs that outperform a wide array of competitors. Chen Zhang 0020, Qiuchi Li, Dawei Song 0001, Zheyu Ye, Yan Gao 0017, Yao Hu 0002 |
ACL (1) | 1 |
| 2025 | A Compressive Memory-based Retrieval Approach for Event Argument ExtractionabstractRecent works have demonstrated the effectiveness of retrieval augmentation in the Event Argument Extraction (EAE) task. However, existing retrieval-based EAE methods have two main limitations: (1) input length constraints and (2) the gap between the retriever and the inference model. These issues limit the diversity and quality of the retrieved information. In this paper, we propose a Compressive Memory-based Retrieval (CMR) mechanism for EAE, which addresses the two limitations mentioned above. Our compressive memory, designed as a dynamic matrix that effectively caches retrieved information and supports continuous updates, overcomes the limitations of input length. Additionally, after pre-loading all candidate demonstrations into the compressive memory, the model further retrieves and filters relevant information from the memory based on the input query, bridging the gap between the retriever and the inference model. Extensive experiments show that our method achieves new state-of-the-art performance on three public datasets (RAMS, WikiEvents, ACE05), significantly outperforming existing retrieval-based EAE methods. Wanlong Liu, Enqi Zhang, Shaohuan Cheng, Dingyi Zeng, Li Zhou 0010, Chen Zhang 0020, Malu Zhang, Wenyu Chen 0001 |
COLING | 6 |
| 2025 | ZigZagKV: Dynamic KV Cache Compression for Long-context Modeling based on Layer UncertaintyabstractLarge Language models (LLMs) have become a research hotspot. To accelerate the inference of LLMs, storing computed caches in memory has become the standard technique. However, as the inference length increases, growing KV caches might lead to out-of-memory issues. Many existing methods address this issue through KV cache compression, primarily by preserving key tokens throughout all layers to reduce information loss. Most of them allocate a uniform budget size for each layer to retain. However, we observe that the minimum budget sizes needed to retain essential information vary across layers and models based on the perspectives of attention and hidden state output. Building on this observation, this paper proposes a simple yet effective KV cache compression method that leverages layer uncertainty to allocate budget size for each layer. Experimental results show that the proposed method can reduce memory usage of the KV caches to only ~20% when compared to full KV inference while achieving nearly lossless performance. Meizhi Zhong, Xikai Liu, Chen Zhang 0020, Yikun Lei, Yan Gao 0017, Yao Hu 0002, Kehai Chen, Min Zhang 0005 |
COLING | 3 |
| 2025 | Understanding the RoPE Extensions of Long-Context LLMs: An Attention PerspectiveabstractEnabling LLMs to handle lengthy context is currently a research hotspot. Most LLMs are built upon rotary position embedding (RoPE), a popular position encoding method. Therefore, a prominent path is to extrapolate the RoPE trained on comparably short texts to far longer texts. A heavy bunch of efforts have been dedicated to boosting the extrapolation via extending the formulations of the RoPE, however, few of them have attempted to showcase their inner workings comprehensively. In this paper, we are driven to offer a straightforward yet in-depth understanding of RoPE extensions from an attention perspective and on two benchmarking tasks. A broad array of experiments reveals several valuable findings: 1) Maintaining attention patterns to those at the pretrained length improves extrapolation; 2) Large attention uncertainty leads to retrieval errors; 3) Using longer continual pretraining lengths for RoPE extensions could reduce attention uncertainty and significantly enhance extrapolation. Meizhi Zhong, Chen Zhang 0020, Yikun Lei, Xikai Liu, Yan Gao 0017, Yao Hu 0002, Kehai Chen, Min Zhang 0005 |
COLING | 2 |
| 2025 | Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference OptimizationabstractCurrent emotional text-to-speech (TTS) models pre-dominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a single emotion per text-speech pair. These models only learn the correct emotional outputs without fully comprehending other emotion characteristics, which limits their capabilities of capturing the nuances between different emotions. We propose a controllable Emo-DPO approach, which employs direct preference optimization to differentiate subtle emotional nuances between emotions through optimizing towards preferred emotions over less preferred emotional ones. Instead of relying on traditional neural architectures used in existing emotional TTS models, we propose utilizing the emotion-aware LLM-TTS neural architecture to leverage LLMs’ in-context learning and instruction-following capabilities. Comprehensive experiments confirm that our proposed method outperforms the existing baselines. Xiaoxue Gao, Chen Zhang 0020, Yiming Chen 0010, Huayun Zhang, Nancy F. Chen |
ICASSP | 2 |
| 2025 | MoDification: Mixture of Depths Made EasyabstractChen Zhang, Meizhi Zhong, Qimeng Wang, Xuantao Lu, Zheyu Ye, Chengqiang Lu, Yan Gao, Yao Hu, Kehai Chen, Min Zhang, Dawei Song. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Chen Zhang 0020, Meizhi Zhong, Qimeng Wang, Xuantao Lu, Zheyu Ye, Chengqiang Lu, Yan Gao 0017, Yao Hu 0002, Kehai Chen, Min Zhang 0005, Dawei Song 0001 |
NAACL (Long Papers) | 1 |
| 2025 | REMAST: Real-Time Emotion-Based Music Arrangement With Soft TransitionabstractMusic as an emotional intervention media has important applications in scenarios such as music therapy, games, and movies. However, music needs real-time arrangement according to changing emotions, bringing challenges to balance emotion real-time fit and soft emotion transition due to the fine-grained and mutable nature of the target emotion. Existing studies mainly focus on achieving emotion real-time fit, while the issue of smooth transition remains understudied, affecting the overall emotional coherence of the music. In this paper, we propose REMAST to address this trade-off. Specifically, we recognize the last timestep's music emotion and fuse it with the current timestep's input emotion. The fused emotion then guides REMAST to generate the music based on the input melody. To adjust music similarity and emotion real-time fit flexibly, we downsample the original melody and feed it into the generation model. Furthermore, we design four music theory features by domain knowledge to enhance emotion information and employ semi-supervised learning to mitigate the subjective bias introduced by manual dataset annotation. According to the evaluation results, REMAST surpasses the state-of-the-art methods in objective and subjective metrics. These results demonstrate that REMAST achieves real-time fit and smooth transition simultaneously, enhancing the coherence of the generated music. Le Ma 0002, Chen Zhang 0020, Bo Han 0015, HaoRong Hong, Xinda Wu |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue EvaluatorsabstractAutomatic evaluation is an integral aspect of dialogue system research. The traditional reference-based NLG metrics are generally found to be unsuitable for dialogue assessment. Consequently, recent studies have suggested various unique, reference-free neural metrics that better align with human evaluations. Notably among them, large language models (LLMs), particularly the instruction-tuned variants like ChatGPT, are shown to be promising substitutes for human judges. Yet, existing works on utilizing LLMs for automatic dialogue evaluation are limited in their scope in terms of the number of meta-evaluation datasets, mode of evaluation, coverage of LLMs, etc. Hence, it remains inconclusive how effective these LLMs are. To this end, we conduct a comprehensive study on the application of LLMs for automatic dialogue evaluation. Specifically, we analyze the multi-dimensional evaluation capability of 30 recently emerged LLMs at both turn and dialogue levels, using a comprehensive set of 12 meta-evaluation datasets. Additionally, we probe the robustness of the LLMs in handling various adversarial perturbations at both turn and dialogue levels. Finally, we explore how model-level and dimension-level ensembles impact the evaluation performance. All resources are available at https://github.com/e0397123/comp-analysis. Chen Zhang 0020, Luis Fernando D'Haro, Yiming Chen 0010, Malu Zhang, Haizhou Li 0001 |
AAAI | 1 |
| 2024 | How Speculative Can Speculative Decoding Be?abstractLarge language models (LLMs) have drawn great attention from the field of natural language processing and beyond, due to their impressive capability of autoregressive modeling, yet bringing an obvious problem, i.e., the largely increased latency. An emerging idea to alleviate this problem is speculative decoding, which first uses a draft model to draft tokens autoregressively and then makes the target model verify these tokens in parallel. The draft model is typically smaller than the target model, and it essentially trades generation quality for speed. Thereby, speculative decoding can be viewed as a speculative game for the target model in term of verification failures. That is, the lengthy draft tokens proposed by the small draft models could fail in the verification stage. Naturally, a critical question arises: how speculative can speculative decoding be, or in other words, how small can an adequate draft model be and how large can an appropriate number of draft tokens be? This work aims to investigate these questions and demonstrate how the scale of the draft model and the number of draft tokens would have an impact on the overall latency of the speculative decoding. We theoretically show that neither of above two factors will be infinitely speculative. Namely, there is a certain turning point for each of them. We then empirically show that the scale of the draft model could be 10-20\times smaller than the target model and the optimal number of draft tokens should lie in 3-5. Zhuorui Liu, Chen Zhang 0020, Dawei Song 0001 |
LREC/COLING | 2 |
| 2024 | CrossTune: Black-Box Few-Shot Classification with Label EnhancementabstractTraining or finetuning large-scale language models (LLMs) requires substantial computation resources, motivating recent efforts to explore parameter-efficient adaptation to downstream tasks. One approach is to treat these models as black boxes and use forward passes (Inference APIs) to interact with them. Current research focuses on adapting these black-box models to downstream tasks using gradient-free prompt optimization, but this often involves an expensive process of searching task-specific prompts. Therefore, we are motivated to study black-box language model adaptation without prompt search. Specifically, we introduce a label-enhanced cross-attention network called CrossTune, which models the semantic relatedness between the input text sequence and task-specific label descriptions. Its effectiveness is examined in the context of few-shot text classification. To improve the generalization of CrossTune, we utilize ChatGPT to generate additional training data through in-context learning. A switch mechanism is implemented to exclude low-quality ChatGPT-generated data. Through extensive experiments on seven benchmark text classification datasets, we demonstrate that our proposed approach outperforms the previous state-of-the-art gradient-free black-box tuning method by 5.7% on average. Even without using ChatGPT-augmented data, CrossTune performs better or comparably than previous black-box tuning methods, suggesting the effectiveness of our approach. Danqing Luo, Chen Zhang 0020, Yan Zhang 0004, Haizhou Li 0001 |
LREC/COLING | 2 |
| 2024 | Task-agnostic Distillation of Encoder-Decoder Language ModelsabstractFinetuning pretrained language models (LMs) have enabled appealing performance on a diverse array of tasks. The intriguing task-agnostic property has driven a shifted focus from task-specific to task-agnostic distillation of LMs. While task-agnostic, compute-efficient, performance-preserved LMs can be yielded by task-agnostic distillation, previous studies mainly sit in distillation of either encoder-only LMs (e.g., BERT) or decoder-only ones (e.g., GPT) yet largely neglect that distillation of encoder-decoder LMs (e.g., T5) can posit very distinguished behaviors. Frustratingly, we discover that existing task-agnostic distillation methods can fail to handle the distillation of encoder-decoder LMs. To the demand, we explore a few paths and uncover a path named as MiniEnD that successfully tackles the distillation of encoder-decoder LMs in a task-agnostic fashion. We examine MiniEnD on language understanding and abstractive summarization. The results showcase that MiniEnD is generally effective and is competitive compared to other alternatives. We further scale MiniEnD up to distillation of 3B encoder-decoder language models with interpolated distillation. The results imply the opportunities and challenges in distilling large language models (e.g., LLaMA). Chen Zhang 0020, Yang Yang 0129, Qiuchi Li, Jingang Wang, Dawei Song 0001 |
LREC/COLING | 1 |
| 2024 | Eliminating Contextual Bias in Aspect-Based Sentiment Analysis
Ruize An, Chen Zhang 0020, Dawei Song 0001 |
ECIR (1) | 2 |
| 2024 | DynaThink: Fast or Slow? A Dynamic Decision-Making Framework for Large Language ModelsabstractLarge language models (LLMs) have demonstrated emergent capabilities across diverse reasoning tasks via popular Chains-of-Thought (COT) prompting.However, such a simple and fast COT approach often encounters limitations in dealing with complicated problems, while a thorough method, which considers multiple reasoning pathways and verifies each step carefully, results in slower inference.This paper addresses the challenge of enabling LLMs to autonomously select between fast and slow inference methods, thereby optimizing both efficiency and effectiveness.We introduce a dynamic decision-making framework that categorizes tasks into two distinct pathways: 'Fast', designated for tasks where the LLM quickly identifies a high-confidence solution, and 'Slow', allocated for tasks that the LLM perceives as complex and for which it has low confidence in immediate solutions as well as requiring more reasoning paths to verify.Experiments on five popular reasoning benchmarks demonstrated the superiority of the Dyna-Think over baselines. Yan Zhang 0004, Chen Zhang 0020, Zuozhu Liu, Hongwei Wang 0001, Haizhou Li 0001 |
EMNLP | 3 |
| 2024 | Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech SynthesisabstractZero-shot text-to-speech (TTS) aims to synthesize voices with unseen speech prompts, which significantly reduces the data and computation requirements for voice cloning by skipping the fine-tuning process. However, the prompting mechanisms of zero-shot TTS still face challenges in the following aspects: 1) previous works of zero-shot TTS are typically trained with single-sentence prompts, which significantly restricts their performance when the data is relatively sufficient during the inference stage. 2) The prosodic information in prompts is highly coupled with timbre, making it untransferable to each other.
This paper introduces Mega-TTS 2, a generic prompting mechanism for zero-shot TTS, to tackle the aforementioned challenges. Specifically, we design a powerful acoustic autoencoder that separately encodes the prosody and timbre information into the compressed latent space while providing high-quality reconstructions. Then, we propose a multi-reference timbre encoder and a prosody latent language model (P-LLM) to extract useful information from multi-sentence prompts. We further leverage the probabilities derived from multiple P-LLM outputs to produce transferable and controllable prosody.
Experimental results demonstrate that Mega-TTS 2 could not only synthesize identity-preserving speech with a short prompt of an unseen speaker from arbitrary sources but consistently outperform the fine-tuning method when the volume of data ranges from 10 seconds to 5 minutes. Furthermore, our method enables to transfer various speaking styles to the target timbre in a fine-grained and controlled manner. Audio samples can be found in https://boostprompt.github.io/boostprompt/. Ziyue Jiang 0001, Jinglin Liu, Yi Ren 0006, Jinzheng He, Zhenhui Ye, Shengpeng Ji, Qian Yang 0006, Chen Zhang 0020, Pengfei Wei 0001, Xiang Yin 0006, Zejun Ma 0001, Zhou Zhao 0001 |
ICLR | 8 |
| 2024 | Real3D-Portrait: One-shot Realistic 3D Talking Portrait SynthesisabstractOne-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simultaneously achieve the goals of accurate 3D avatar reconstruction and stable talking face animation. Besides, while the existing works mainly focus on synthesizing the head part, it is also vital to generate natural torso and background segments to obtain a realistic talking portrait video. To address these limitations, we present Real3D-Potrait, a framework that (1) improves the one-shot 3D reconstruction power with a large image-to-plane model that distills 3D prior knowledge from a 3D face generative model; (2) facilitates accurate motion-conditioned animation with an efficient motion adapter; (3) synthesizes realistic video with natural torso movement and switchable background using a head-torso-background super-resolution model; and (4) supports one-shot audio-driven talking face generation with a generalizable audio-to-motion model. Extensive experiments show that Real3D-Portrait generalizes well to unseen identities and generates more realistic talking portrait videos compared to previous methods. Video samples are available at https://real3dportrait.github.io. Zhenhui Ye, Tianyun Zhong, Yi Ren 0006, Jiaqi Yang 0008, Weichuang Li, Jiawei Huang 0008, Ziyue Jiang 0001, Jinzheng He, Rongjie Huang 0001, Jinglin Liu, Chen Zhang 0020, Xiang Yin 0006, Zejun Ma 0001, Zhou Zhao 0001 |
ICLR | 11 |
| 2024 | MimicTalk: Mimicking a personalized and expressive 3D talking face in minutesabstractTalking face generation (TFG) aims to animate a target identity's face to create realistic talking videos. Personalized TFG is a variant that emphasizes the perceptual identity similarity of the synthesized result (from the perspective of appearance and talking style). While previous works typically solve this problem by learning an individual neural radiance field (NeRF) for each identity to implicitly store its static and dynamic information, we find it inefficient and non-generalized due to the per-identity-per-training framework and the limited training data. To this end, we propose MimicTalk, the first attempt that exploits the rich knowledge from a NeRF-based person-agnostic generic model for improving the efficiency and robustness of personalized TFG. To be specific, (1) we first come up with a person-agnostic 3D TFG model as the base model and propose to adapt it into a specific identity; (2) we propose a static-dynamic-hybrid adaptation pipeline to help the model learn the personalized static appearance and facial dynamic features; (3) To generate the facial motion of the personalized talking style, we propose an in-context stylized audio-to-motion model that mimics the implicit talking style provided in the reference video without information loss by an explicit style representation. The adaptation process to an unseen identity can be performed in 15 minutes, which is 47 times faster than previous person-dependent methods. Experiments show that our MimicTalk surpasses previous baselines regarding video quality, efficiency, and expressiveness. Video samples are available at https://mimictalk.github.io . Zhenhui Ye, Tianyun Zhong, Yi Ren 0006, Ziyue Jiang 0001, Jiawei Huang 0008, Rongjie Huang 0001, Jinglin Liu, Jinzheng He, Chen Zhang 0020, Zehan Wang 0001, Xize Cheng, Xiang Yin 0006, Zhou Zhao 0001 |
NeurIPS | 9 |
| 2024 | Intrusion detection for Industrial Internet of Things based on deep learning
Yaoyao Lu, Senchun Chai, Yuhan Suo, Fenxi Yao, Chen Zhang 0020 |
Neurocomputing | 5 |
| 2024 | NaturalSpeech: End-to-End Text-to-Speech Synthesis With Human-Level QualityabstractText-to-speech (TTS) has made rapid progress in both academia and industry in recent years. Some questions naturally arise that whether a TTS system can achieve human-level quality, how to define/judge that quality, and how to achieve it. In this paper, we answer these questions by first defining the human-level quality based on the statistical significance of subjective measure and introducing appropriate guidelines to judge it, and then developing a TTS system called NaturalSpeech that achieves human-level quality on benchmark datasets. Specifically, we leverage a variational auto-encoder (VAE) for end-to-end text-to-waveform generation, with several key modules to enhance the capacity of the prior from text and reduce the complexity of the posterior from speech, including phoneme pre-training, differentiable duration modeling, bidirectional prior/posterior modeling, and a memory mechanism in VAE. Experimental evaluations on the popular LJSpeech dataset show that our proposed NaturalSpeech achieves -0.01 CMOS (comparative mean opinion score) to human recordings at the sentence level, with Wilcoxon signed rank test at p-level p >> 0.05, which demonstrates no statistically significant difference from human recordings for the first time. Xu Tan 0003, Jiawei Chen 0008, Haohe Liu, Jian Cong, Chen Zhang 0020, Xi Wang 0016, Yichong Leng, Yuanhao Yi, Lei He 0005, Sheng Zhao 0002, Tao Qin 0001, Frank K. Soong, Tie-Yan Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Overview of the Tenth Dialog System Technology Challenge: DSTC10abstractThis article introduces the Tenth Dialog System Technology Challenge (DSTC-10). This edition of the DSTC focuses on applying end-to-end dialog technologies for five distinct tasks in dialog systems, namely 1. Incorporation of Meme images into open domain dialogs, 2. Knowledge-grounded Task-oriented Dialogue Modeling on Spoken Conversations, 3. Situated Interactive Multimodal dialogs, 4. Reasoning for Audio Visual Scene-Aware Dialog, and 5. Automatic Evaluation and Moderation of Open-domainDialogue Systems. This article describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks. Koichiro Yoshino, Yun-Nung Chen, Paul A. Crook, Satwik Kottur, Jinchao Li, Behnam Hedayatnia, Seungwhan Moon, Zhengcong Fei, Zekang Li, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016, Seokhwan Kim, Yang Liu 0004, Di Jin 0005, Alexandros Papangelis, Karthik Gopalakrishnan 0001, Dilek Hakkani-Tür, Babak Damavandi, Alborz Geramifard, Chiori Hori, Chen Zhang 0020, Haizhou Li 0001, João Sedoc, Luis Fernando D'Haro, Rafael E. Banchs, Alexander I. Rudnicky |
IEEE ACM Trans. Audio Speech Lang. Process. | 23 |
| 2024 | RefXVC: Cross-Lingual Voice Conversion With Enhanced Reference LeveragingabstractThis paper proposes RefXVC, a method for cross-lingual voice conversion (XVC) that leverages reference information to improve conversion performance. Previous XVC works generally take an average speaker embedding to condition the speaker identity, which does not account for the changing timbre of speech that occurs with different pronunciations. To address this, our method uses both global and local speaker embeddings to capture the timbre changes during speech conversion. Additionally, we observed a connection between timbre and pronunciation in different languages and utilized this by incorporating a timbre encoder and a pronunciation matching network into our model. Furthermore, we found that the variation in tones is not adequately reflected in a sentence, and therefore, we used multiple references to better capture the range of a speaker's voice. The proposed method outperformed existing systems in terms of both speech quality and speaker similarity, highlighting the effectiveness of leveraging reference information in cross-lingual voice conversion. Mingyang Zhang 0003, Yi Zhou 0020, Yi Ren 0006, Chen Zhang 0020, Xiang Yin 0006, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | SDMuse: Stochastic Differential Music Editing and Generation via Hybrid RepresentationabstractWhile deep generative models have empowered music generation, it remains a challenging and under-explored problem to edit an existing musical piece at fine granularity. In this paper, we propose SDMuse, a unified Stochastic Differential Music editing and generation framework, which can not only compose a whole musical piece from scratch, but also modify existing musical pieces in many ways, such as combination, continuation, inpainting, and style transferring. The proposed SDMuse follows a two-stage pipeline to achieve music generation and editing on top of a hybrid representation including pianoroll and MIDI-event. In particular, SDMuse first generates/edits pianoroll by iteratively denoising through a stochastic differential equation (SDE) based on a diffusion model generative prior, and then refines the generated pianoroll and predicts MIDI-event tokens auto-regressively. We evaluate the generated music of our method onailabs1k7pop music dataset in terms of quality and controllability on various music editing and generation tasks. Experimental results demonstrate the effectiveness of our proposed stochastic differential music editing and generation process, as well as the hybrid representations. Chen Zhang 0020, Yi Ren 0006, Shuicheng Yan |
IEEE Trans. Multim. | 1 |
| 2024 | On Elastic Language ModelsabstractLarge-scale pretrained language models have achieved compelling performance in a wide range of language understanding and information retrieval tasks. While their large scales ensure capacity, they also hinder deployment. Knowledge distillation offers an opportunity to compress a large language model to a small one, in order to reach a reasonable latency-performance tradeoff. However, for scenarios where the number of requests (e.g., queries submitted to a search engine) is highly variant, the static tradeoff attained by the compressed language model might not always fit. Once a model is assigned with a static tradeoff, it could be inadequate in that the latency is too high when the number of requests is large, or the performance is too low when the number of requests is small. To this end, we propose an elastic language model ( ElasticLM ) that elastically adjusts the tradeoff according to the request stream. The basic idea is to introduce a compute elasticity to the compressed language model, so that the tradeoff could vary on-the-fly along a scalable and controllable compute. Specifically, we impose an elastic structure to equip ElasticLM with compute elasticity and design an elastic optimization method to learn ElasticLM under compute elasticity. To serve ElasticLM , we apply an elastic schedule. Considering the specificity of information retrieval, we adapt ElasticLM to dense retrieval and reranking, and present an ElasticDenser and an ElasticRanker, respectively. Offline evaluation is conducted on a language understanding benchmark GLUE, and several information retrieval tasks including Natural Question, Trivia QA and MS MARCO. The results show that ElasticLM along with ElasticDenser and ElasticRanker can perform correctly and competitively compared with an array of static baselines. Furthermore, an online simulation with concurrency is also carried out. The results demonstrate that ElasticLM can provide elastic tradeoffs with respect to varying request stream. Chen Zhang 0020, Benyou Wang, Dawei Song 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2023 | VideoDubber: Machine Translation with Speech-Aware Length Control for Video DubbingabstractVideo dubbing aims to translate the original speech in a film or television program into the speech in a target language, which can be achieved with a cascaded system consisting of speech recognition, machine translation and speech synthesis. To ensure the translated speech to be well aligned with the corresponding video, the length/duration of the translated speech should be as close as possible to that of the original speech, which requires strict length control. Previous works usually control the number of words or characters generated by the machine translation model to be similar to the source sentence, without considering the isochronicity of speech as the speech duration of words/characters in different languages varies. In this paper, we propose VideoDubber, a machine translation system tailored for the task of video dubbing, which directly considers the speech duration of each token in translation, to match the length of source and target speech. Specifically, we control the speech length of generated sentence by guiding the prediction of each word with the duration information, including the speech duration of itself as well as how much duration is left for the remaining words. We design experiments on four language directions (German -> English, Spanish -> English, Chinese English), and the results show that VideoDubber achieves better length control ability on the generated speech than baseline methods. To make up the lack of real-world datasets, we also construct a real-world test set collected from films to provide comprehensive evaluations on the video dubbing task. Yihan Wu 0008, Junliang Guo, Xu Tan 0003, Chen Zhang 0020, Bohan Li 0003, Ruihua Song, Lei He 0005, Sheng Zhao 0002, Arul Menezes, Jiang Bian 0002 |
AAAI | 4 |
| 2023 | Lifting the Curse of Capacity Gap in Distilling Language ModelsabstractChen Zhang, Yang Yang, Jiahao Liu, Jingang Wang, Yunsen Xian, Benyou Wang, Dawei Song. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Chen Zhang 0020, Yang Yang 0129, Jingang Wang, Yunsen Xian, Benyou Wang, Dawei Song 0001 |
ACL (1) | 1 |
| 2023 | LeanSpeech: The Microsoft Lightweight Speech Synthesis System for Limmits Challenge 2023abstractThis paper describes the Microsoft Text-to-Speech (TTS) system: LeanSpeech for LIMMITS (Lightweight, Multi-speaker, Multi-lingual Indic TTS) Challenge 20231, which is part of ICASSP2023 to encourage the advance of TTS in Indian Languages. We propose a lightweight encoder-decoder acoustic model composed of 1-D convolution and LSTM blocks, which is trained with knowledge distillation from a multi-speaker multi-lingual teacher model, DelightfulTTS [1]. The speech corpus is reprocessed and used in both AM training and vocoder fine-tuning. In Track-2 of the challenge, our system achieves MOS 4.56 and SMOS 3.98, which indicates the efficiency of the proposed model and training strategy. Chen Zhang 0020, Shubham Bansal, Aakash Lakhera, Jinzhu Li, Gang Wang 0001, Sandeepkumar Satpal, Sheng Zhao 0002, Lei He 0005 |
ICASSP | 1 |
| 2023 | Bag of Tricks for Unsupervised Text-to-Speech
Yi Ren 0006, Chen Zhang 0020, Shuicheng Yan |
ICLR | 2 |
| 2023 | Stance-level Sarcasm Detection with BERT and Stance-centered Graph Attention NetworksabstractComputational Linguistics (CL) associated with the Internet of Multimedia Things (IoMT)-enabled multimedia computing applications brings several research challenges, such as real-time speech understanding, deep fake video detection, emotion recognition, home automation, and so on. Due to the emergence of machine translation, CL solutions have increased tremendously for different natural language processing (NLP) applications. Nowadays, NLP-enabled IoMT is essential for its success. Sarcasm detection, a recently emerging artificial intelligence (AI) and NLP task, aims at discovering sarcastic, ironic, and metaphoric information implied in texts that are generated in the IoMT. It has drawn much attention from the AI and IoMT research community. The advance of sarcasm detection and NLP techniques will provide a cost-effective, intelligent way to work together with machine devices and high-level human-to-device interactions. However, existing sarcasm detection approaches neglect the hidden stance behind texts, thus insufficient to exploit the full potential of the task. Indeed, the stance, i.e., whether the author of a text is in favor of, against, or neutral toward the proposition or target talked in the text, largely determines the text’s actual sarcasm orientation. To fill the gap, in this research, we propose a new task: stance-level sarcasm detection (SLSD), where the goal is to uncover the author’s latent stance and based on it to identify the sarcasm polarity expressed in the text. We then propose an integral framework, which consists of Bidirectional Encoder Representations from Transformers (BERT) and a novel stance-centered graph attention networks (SCGAT). Specifically, BERT is used to capture the sentence representation, and SCGAT is designed to capture the stance information on specific target. Extensive experiments are conducted on a Chinese sarcasm sentiment dataset we created and the SemEval-2018 Task 3 English sarcasm dataset. The experimental results prove the effectiveness of the SCGAT framework over state-of-the-art baselines by a large margin. Yazhou Zhang 0001, Dan Ma 0010, Prayag Tiwari, Chen Zhang 0020, Mehedi Masud, Mohammad Shorfuzzaman, Dawei Song 0001 |
ACM Trans. Internet Techn. | 4 |
| 2022 | Structural Bias for Aspect Sentiment Triplet ExtractionabstractStructural bias has recently been exploited for aspect sentiment triplet extraction (ASTE) and led to improved performance. On the other hand, it is recognized that explicitly incorporating structural bias would have a negative impact on efficiency, whereas pretrained language models (PLMs) can already capture implicit structures. Thus, a natural question arises: Is structural bias still a necessity in the context of PLMs? To answer the question, we propose to address the efficiency issues by using an adapter to integrate structural bias in the PLM and using a cheap-to-compute relative position structure in place of the syntactic dependency structure. Benchmarking evaluation is conducted on the SemEval datasets. The results show that our proposed structural adapter is beneficial to PLMs and achieves state-of-the-art performance over a range of strong baselines, yet with a light parameter demand and low latency. Meanwhile, we give rise to the concern that the current evaluation default with data of small scale is under-confident. Consequently, we release a large-scale dataset for ASTE. The results on the new dataset hint that the structural adapter is confidently effective and efficient to a large scale. Overall, we draw the conclusion that structural bias shall still be a necessity even with PLMs. Chen Zhang 0020, Fang Ma, Jingang Wang, Wei Wu 0014, Dawei Song 0001 |
COLING | 1 |
| 2022 | TeleMelody: Lyric-to-Melody Generation with a Template-Based Two-Stage MethodabstractZeqian Ju, Peiling Lu, Xu Tan, Rui Wang, Chen Zhang, Songruoyao Wu, Kejun Zhang, Xiang-Yang Li, Tao Qin, Tie-Yan Liu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Zeqian Ju, Peiling Lu, Xu Tan 0003, Rui Wang 0031, Chen Zhang 0020, Songruoyao Wu, Xiang-Yang Li 0001, Tao Qin 0001, Tie-Yan Liu |
EMNLP | 5 |
| 2022 | XPrompt: Exploring the Extreme of Prompt TuningabstractPrompt tuning learns soft prompts to condition the frozen Pre-trained Language Models (PLMs) for performing downstream tasks in a parameter-efficient manner.While prompt tuning has gradually reached the performance level of fine-tuning as the model scale increases, there is still a large performance gap between prompt tuning and fine-tuning for models of moderate and small scales (typically less than 11B parameters).In this paper, we empirically show that the trained prompt tokens can have a negative impact on a downstream task and thus degrade its performance.To bridge the gap, we propose a novel PROMPT tuning model with an eXtremely small scale (XPROMPT) under the regime of lottery tickets hypothesis.Specifically, XPROMPT eliminates the negative prompt tokens at different granularity levels through a hierarchical structured pruning, yielding a more parameter-efficient prompt yet with a competitive performance.Comprehensive experiments are carried out on the SuperGLUE tasks, and the results indicate that XPROMPT is able to close the performance gap at smaller model scales. 1 Recently, Prompt-Tuning (Lester et al., 2021; Liu et al., 2021b) has been proposed to address this issue by prepending a soft prompt to the input and only updating the parameters of prompt tokens during tuning.Prompt-Tuning provides a parameter-efficient alternative to fine-tuning, since the scale of the soft prompt is tens of thousand smaller.It is also conceptually simpler and more flexible than other parameter-efficient tuning methods (such as Adapters), that require intrusive modifications to transformer layers (Houlsby et al., 2019;Guo et al., 2021).Using fewer tunable parameters, prompt tuning achieves competitive performance to fine-tuning with the increase of the model scale.However, there is still a large performance gap between prompt tuning and fine-tuning for models of smaller scales (as shown in Figure 1).This paper aims to fill the gap, from the perspective of the lottery tickets hypothesis (LTH) (Frankle and Carbin, 2019).We are motivated by an observation that, on a specific task, not all prompt tokens contribute equally to the task performance, while certain prompt tokens may even bring a negative Fang Ma, Chen Zhang 0020, Jingang Wang, Qifan Wang 0001, Wei Wu 0014, Xiaojun Quan, Dawei Song 0001 |
EMNLP | 2 |
| 2022 | Analyzing and Evaluating Faithfulness in Dialogue SummarizationabstractDialogue summarization is abstractive in nature, making it suffer from factual errors.The factual correctness of summaries has the highest priority before practical applications.Many efforts have been made to improve faithfulness in text summarization.However, there is a lack of systematic study on dialogue summarization systems.In this work, we first perform the fine-grained human analysis on the faithfulness of dialogue summaries and observe that over 35% of generated summaries are faithfully inconsistent respective the source dialogues.Furthermore, we present a new model-level faithfulness evaluation method.It examines generation models with multi-choice questions created by rule-based transformations.Experimental results show that our evaluation schema is a strong proxy for the factual correctness of summarization models.The humanannotated faithfulness samples and the evaluation toolkit are released to facilitate future research toward faithful dialogue summarization.Code Bin Wang 0040, Chen Zhang 0020, Yan Zhang 0004, Yiming Chen 0010, Haizhou Li 0001 |
EMNLP | 2 |
| 2022 | Sparse Teachers Can Be Dense with KnowledgeabstractRecent advances in distilling pretrained language models have discovered that, besides the expressiveness of knowledge, the studentfriendliness should be taken into consideration to realize a truly knowledgeable teacher.Based on a pilot study, we find that overparameterized teachers can produce expressive yet student-unfriendly knowledge and are thus limited in overall knowledgeableness.To remove the parameters that result in studentunfriendliness, we propose a sparse teacher trick under the guidance of an overall knowledgeable score for each teacher parameter.The knowledgeable score is essentially an interpolation of the expressiveness and studentfriendliness scores.The aim is to ensure that the expressive parameters are retained while the student-unfriendly ones are removed.Extensive experiments on the GLUE benchmark show that the proposed sparse teachers can be dense with knowledge and lead to students with compelling performance in comparison with a series of competitive baselines.1 Chen Zhang 0020, Dawei Song 0001 |
EMNLP | 2 |
| 2022 | Making Pretrained Language Models Good Long-tailed LearnersabstractPrompt-tuning has shown appealing performance in few-shot classification by virtue of its capability in effectively exploiting pre-trained knowledge.This motivates us to check the hypothesis that prompt-tuning is also a promising choice for long-tailed classification, since the tail classes are intuitively few-shot ones.To achieve this aim, we conduct empirical studies to examine the hypothesis.The results demonstrate that prompt-tuning makes pretrained language models at least good long-tailed learners.For intuitions on why prompt-tuning can achieve good performance in long-tailed classification, we carry out in-depth analyses by progressively bridging the gap between prompttuning and commonly used finetuning.The summary is that the classifier structure and parameterization form the key to making good long-tailed learners, in comparison with the less important input structure.Finally, we verify the applicability of our finding to few-shot classification. 1 Chen Zhang 0020, Jingang Wang, Wei Wu 0014, Dawei Song 0001 |
EMNLP | 1 |
| 2022 | S3T: Self-Supervised Pre-Training with Swin Transformer For Music ClassificationabstractIn this paper, we propose S3T, a self-supervised pre-training method with Swin Transformer for music classification, aiming to learn meaningful music representations from massive easily accessible unlabeled music data. S3T introduces a momentum-based paradigm, MoCo, with Swin Transformer as its feature extractor to music time-frequency domain. For better music representations learning, S3T contributes a music data augmentation pipeline and two specially designed pre-processors. To our knowledge, S3T is the first method combining the Swin Transformer with a self-supervised learning method for music classification. We evaluate S3T on music genre classification and music tagging tasks with linear classifiers trained on learned representations. Experimental results show that S3T outperforms the previous self-supervised method (CLMR) by 12.5 percents top-1 accuracy and 4.8 percents PR-AUC on two tasks respectively, and also surpasses the task-specific state-of-the-art supervised methods. Besides, S3T shows advances in label efficiency using only 10% labeled data exceeding CLMR on both tasks with 100% labeled data. Chen Zhang 0020, Bilei Zhu, Zejun Ma 0001 |
ICASSP | 2 |
| 2022 | SongDriver: Real-time Music Accompaniment Generation without Logical Latency nor Exposure BiasabstractReal-time music accompaniment generation has a wide range of applications in the music industry, such as music education and live performances. However, automatic real-time music accompaniment generation is still understudied and often faces a trade-off between logical latency and exposure bias. In this paper, we propose SongDriver, a real-time music accompaniment generation system without logical latency nor exposure bias. Specifically, SongDriver divides one accompaniment generation task into two phases: 1) The arrangement phase, where a Transformer model first arranges chords for input melodies in real-time, and caches the chords for the next phase instead of playing them out. 2) The prediction phase, where a CRF model generates playable multi-track accompaniments for the coming melodies based on previously cached chords. With this two-phase strategy, SongDriver directly generates the accompaniment for the upcoming melody, achieving zero logical latency. Furthermore, when predicting chords for a timestep, SongDriver refers to the cached chords from the first phase rather than its previous predictions, which avoids the exposure bias problem. Since the input length is often constrained under real-time conditions, another potential problem is the loss of long-term sequential information. To make up for this disadvantage, we extract four musical features from a long-term music piece before the current time step as global information. In the experiment, we train SongDriver on some open-source datasets and an original àiMusic Dataset built from Chinese-style modern pop music sheets. The results show that SongDriver outperforms existing SOTA (state-of-the-art) models on both objective and subjective metrics, meanwhile significantly reducing the physical latency. Chen Zhang 0020, Qihao Liang, Yongsheng Feng, Yuntai Bao |
ACM Multimedia | 4 |
| 2022 | ReLyMe: Improving Lyric-to-Melody Generation by Incorporating Lyric-Melody RelationshipsabstractLyric-to-melody generation, which generates melody according to given lyrics, is one of the most important automatic music composition tasks. With the rapid development of deep learning, previous works address this task with end-to-end neural network models. However, deep learning models cannot well capture the strict but subtle relationships between lyrics and melodies, which compromises the harmony between lyrics and generated melodies. In this paper, we propose ReLyMe, a method that incorporates Relationships between Lyrics and Melodies from music theory to ensure the harmony between lyrics and melodies. Specifically, we first introduce several principles that lyrics and melodies should follow in terms of tone, rhythm, and structure relationships. These principles are then integrated into neural network lyric-to-melody models by adding corresponding constraints during the decoding process to improve the harmony between lyrics and melodies. We use a series of objective and subjective metrics to evaluate the generated melodies. Experiments on both English and Chinese song datasets show the effectiveness of ReLyMe, demonstrating the superiority of incorporating lyric-melody relationships from the music domain into neural lyric-to-melody generation. Chen Zhang 0020, LuChin Chang, Songruoyao Wu, Xu Tan 0003, Tao Qin 0001, Tie-Yan Liu |
ACM Multimedia | 1 |
| 2022 | Towards Effective Multi-Modal Interchanges in Zero-Resource Sounding Object LocalizationabstractAiming to locate the object that emits a specified sound in complex scenes, the task of sounding object localization bridges two perception-oriented modalities of vision and acoustics, and brings enormous research value to the comprehensive perceptual understanding of machine intelligence. Although there are massive training data collected in this field, few of them contain accurate bounding box annotations, hindering the learning process and further application of proposed models. In order to address this problem, we try to explore an effective multi-modal knowledge transfer strategy to obtain precise knowledge from other similar tasks and transfer it through well-aligned multi-modal data to deal with this task in a zero-resource manner. Concretely, we design and propose a novel \textit{Two-stream Universal Referring localization Network} (TURN), which is composed of a localization stream and an alignment stream to carry out different functions. The former is utilized to extract the knowledge related to referring object localization from the image grounding task, while the latter is devised to learn a universal semantic space shared between texts and audios. Moreover, we further develop an adaptive sampling strategy to automatically identify the overlap between different data domains, thus boosting the performance and stability of our model. The extensive experiments on various publicly-available benchmarks demonstrate that TURN can achieve competitive performance compared with the state-of-the-art approaches without using any data in this field, which verifies the feasibility of our proposed mechanisms and strategies. Yang Zhao 0022, Chen Zhang 0020, Haifeng Huang 0001, Haoyuan Li 0002, Zhou Zhao 0001 |
NeurIPS | 2 |
| 2022 | Aspect-Specific Context Modeling for Aspect-Based Sentiment Analysis
Fang Ma, Chen Zhang 0020, Dawei Song 0001 |
NLPCC (1) | 2 |
| 2022 | Doge Tickets: Uncovering Domain-General Language Models by Playing Lottery Tickets
Chen Zhang 0020, Benyou Wang, Dawei Song 0001 |
NLPCC (1) | 2 |
| 2022 | Adaptable Text Matching via Meta-Weight RegulatorabstractNeural text matching models have been used in a range of applications such as question answering and natural language inference, and have yielded a good performance. However, these neural models are of a limited adaptability, resulting in a decline in performance when encountering test examples from a different dataset or even a different task. The adaptability is particularly important in the few-shot setting: in many cases, there is only a limited amount of labeled data available for a target dataset or task, while we may have access to a richly labeled source dataset or task. However, adapting a model trained on the abundant source data to a few-shot target dataset or task is challenging. To tackle this challenge, we propose a Meta-Weight Regulator (MWR), which is a meta-learning approach that learns to assign weights to the source examples based on their relevance to the target loss. Specifically, MWR first trains the model on the uniformly weighted source examples, and measures the efficacy of the model on the target examples via a loss function. By iteratively performing a (meta) gradient descent, high-order gradients are propagated to the source examples. These gradients are then used to update the weights of source examples, in a way that is relevant to the target performance. As MWR is model-agnostic, it can be applied to any backbone neural model. Extensive experiments are conducted with various backbone text matching models, on four widely used datasets and two tasks. The results demonstrate that our proposed approach significantly outperforms a number of existing adaptation methods and effectively improves the cross-dataset and cross-task adaptability of the neural text matching models in the few-shot setting. Chen Zhang 0020, Fang Ma, Dawei Song 0001 |
SIGIR | 2 |
| 2021 | UWSpeech: Speech to Speech Translation for Unwritten LanguagesabstractExisting speech to speech translation systems heavily rely on the text of target language: they usually translate source language either to target text and then synthesize target speech from text, or directly to target speech with target text for auxiliary training. However, those methods cannot be applied to unwritten target languages, which have no written text or phoneme available. In this paper, we develop a translation system for unwritten languages, named as UWSpeech, which converts target unwritten speech into discrete tokens with a converter, and then translates source-language speech into target discrete tokens with a translator, and finally synthesizes target speech from target discrete tokens with an inverter. We propose a method called XL-VAE, which enhances vector quantized variational autoencoder (VQ-VAE) with cross-lingual (XL) speech recognition, to train the converter and inverter of UWSpeech jointly. Experiments on Fisher Spanish-English conversation translation dataset show that UWSpeech outperforms direct translation and VQ-VAE baseline by about 16 and 10 BLEU points respectively, which demonstrate the advantages and potentials of UWSpeech. Chen Zhang 0020, Xu Tan 0003, Yi Ren 0006, Tao Qin 0001, Tie-Yan Liu |
AAAI | 1 |
| 2021 | Revisiting Self-training for Few-shot Learning of Language ModelabstractAs unlabeled data carry rich task-relevant information, they are proven useful for fewshot learning of language model.The question is how to effectively make use of such data.In this work, we revisit the self-training technique for language model fine-tuning and present a state-of-the-art prompt-based fewshot learner, SFLM.Given two views of a text sample via weak and strong augmentation techniques, SFLM generates a pseudo label on the weakly augmented version.Then, the model predicts the same pseudo label when fine-tuned with the strongly augmented version.This simple approach is shown to outperform other state-of-the-art supervised and semi-supervised counterparts on six sentence classification and six sentence-pair classification benchmarking tasks.In addition, SFLM only relies on a few in-domain unlabeled data.We conduct a comprehensive analysis to demonstrate the robustness of our proposed approach under various settings, including augmentation techniques, model scale, and fewshot knowledge transfer across tasks. Yiming Chen 0010, Yan Zhang 0004, Chen Zhang 0020, Grandee Lee, Haizhou Li 0001 |
EMNLP (1) | 3 |
| 2021 | Denoispeech: Denoising Text to Speech with Frame-Level Noise ModelingabstractWhile neural-based text to speech (TTS) models can synthesize natural and intelligible voice, they usually require high-quality speech data, which is costly to collect. In many scenarios, only noisy speech of a target speaker is available, which presents challenges for TTS model training for this speaker. Previous works usually address the challenge using two methods: 1) training the TTS model using the speech denoised with an enhancement model; 2) taking a single noise embedding as input when training with noisy speech. However, they usually cannot handle speech with real-world complicated noise such as those with high variations along time. In this paper, we develop DenoiSpeech, a TTS system that can synthesize clean speech for a speaker with noisy speech data. In DenoiSpeech, we handle real-world noisy speech by modeling the fine-grained frame-level noise with a noise condition module, which is jointly trained with the TTS model. Experimental results on real-world data show that DenoiSpeech outperforms the previous two methods by 0.31 and 0.66 MOS respectively.1 Chen Zhang 0020, Yi Ren 0006, Xu Tan 0003, Jinglin Liu, Tao Qin 0001, Sheng Zhao 0002, Tie-Yan Liu |
ICASSP | 1 |
| 2021 | A Simple Baseline for Cross-Domain Few-Shot Text Classification
Chen Zhang 0020, Dawei Song 0001 |
NLPCC (1) | 1 |
| 2020 | SimulSpeech: End-to-End Simultaneous Speech to Text TranslationabstractIn this work, we develop SimulSpeech, an endto-end simultaneous speech to text translation system which translates speech in source language to text in target language concurrently.SimulSpeech consists of a speech encoder, a speech segmenter and a text decoder, where 1) the segmenter builds upon the encoder and leverages a connectionist temporal classification (CTC) loss to split the input streaming speech in real time, 2) the encoder-decoder attention adopts a wait-k strategy for simultaneous translation.SimulSpeech is more challenging than previous cascaded systems (with simultaneous automatic speech recognition (ASR) and simultaneous neural machine translation (NMT)).We introduce two novel knowledge distillation methods to ensure the performance: 1) Attention-level knowledge distillation transfers the knowledge from the multiplication of the attention matrices of simultaneous NMT and ASR models to help the training of the attention mechanism in SimulSpeech; 2) Data-level knowledge distillation transfers the knowledge from the full-sentence NMT model and also reduces the complexity of data distribution to help on the optimization of Simul-Speech.Experiments on MuST-C English-Spanish and English-German spoken language translation datasets show that SimulSpeech achieves reasonable BLEU scores and lower delay compared to full-sentence end-to-end speech to text translation (without simultaneous translation), and better performance than the two-stage cascaded simultaneous translation model in terms of BLEU scores and translation delay. Yi Ren 0006, Jinglin Liu, Xu Tan 0003, Chen Zhang 0020, Tao Qin 0001, Zhou Zhao 0001, Tie-Yan Liu |
ACL | 4 |
| 2020 | Task-Level Curriculum Learning for Non-Autoregressive Neural Machine TranslationabstractNon-autoregressive translation (NAT) achieves faster inference speed but at the cost of worse accuracy compared with autoregressive translation (AT). Since AT and NAT can share model structure and AT is an easier task than NAT due to the explicit dependency on previous target-side tokens, a natural idea is to gradually shift the model training from the easier AT task to the harder NAT task. To smooth the shift from AT training to NAT training, in this paper, we introduce semi-autoregressive translation (SAT) as intermediate tasks. SAT contains a hyperparameter k, and each k value defines a SAT task with different degrees of parallelism. Specially, SAT covers AT and NAT as its special cases: it reduces to AT when k=1 and to NAT when k=N (N is the length of target sentence). We design curriculum schedules to gradually shift k from 1 to N, with different pacing functions and number of tasks trained at the same time. We called our method as task-level curriculum learning for NAT (TCL-NAT). Experiments on IWSLT14 De-En, IWSLT16 En-De, WMT14 En-De and De-En datasets show that TCL-NAT achieves significant accuracy improvements over previous NAT baselines and reduces the performance gap between NAT and AT models to 1-2 BLEU points, demonstrating the effectiveness of our proposed method. Jinglin Liu, Yi Ren 0006, Xu Tan 0003, Chen Zhang 0020, Tao Qin 0001, Zhou Zhao 0001, Tie-Yan Liu |
IJCAI | 4 |
| 2020 | FastLR: Non-Autoregressive Lipreading Model with Integrate-and-FireabstractLipreading is an impressive technique and there has been a definite improvement of accuracy in recent years. However, existing methods for lipreading mainly build on autoregressive (AR) model, which generate target tokens one by one and suffer from high inference latency. To breakthrough this constraint, we propose FastLR, a non-autoregressive (NAR) lipreading model which generates all target tokens simultaneously. NAR lipreading is a challenging task that has many difficulties: 1) the discrepancy of sequence lengths between source and target makes it difficult to estimate the length of the output sequence; 2) the conditionally independent behavior of NAR generation lacks the correlation across time which leads to a poor approximation of target distribution; 3) the feature representation ability of encoder can be weak due to lack of effective alignment mechanism; and 4) the removal of AR language model exacerbates the inherent ambiguity problem of lipreading. Thus, in this paper, we introduce three methods to reduce the gap between FastLR and AR model: 1) to address challenges 1 and 2, we leverage integrate-and-fire (I&F) module to model the correspondence between source video frames and output text sequence. 2) To tackle challenge 3, we add an auxiliary connectionist temporal classification (CTC) decoder to the top of the encoder and optimize it with extra CTC loss. We also add an auxiliary autoregressive decoder to help the feature extraction of encoder. 3) To overcome challenge 4, we propose a novel Noisy Parallel Decoding (NPD) for I&F and bring Byte-Pair Encoding (BPE) into lipreading. Our experiments exhibit that FastLR achieves the speedup up to 10.97× comparing with state-of-the-art lipreading model with slight WER absolute increase of 1.5% and 5.5% on GRID and LRS2 lipreading datasets respectively, which demonstrates the effectiveness of our proposed method. Jinglin Liu, Yi Ren 0006, Zhou Zhao 0001, Chen Zhang 0020, Baoxing Huai, Nicholas Jing Yuan |
ACM Multimedia | 4 |
| 2019 | Aspect-based Sentiment Classification with Aspect-specific Graph Convolutional NetworksabstractChen Zhang, Qiuchi Li, Dawei Song. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Chen Zhang 0020, Qiuchi Li, Dawei Song 0001 |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Discriminative and Correlative Partial Multi-Label LearningabstractIn partial label learning (PML), each instance is associated with a candidate label set that contains multiple relevant labels and other false positive labels. The most challenging issue for the PML is that the training procedure is prone to be affected by the labeling noise. We observe that state-of-the-art PML methods are either powerless to disambiguate the correct labels from the candidate labels or incapable of extracting the label correlations sufficiently. To fill this gap, a two-stage DiscRiminative and correlAtive partial Multi-label leArning (DRAMA) algorithm is presented in this work. In the first stage, a confidence value is learned for each label by utilizing the feature manifold, which indicates how likely a label is correct. In the second stage, a gradient boosting model is induced to fit the label confidences. Specifically, to explore the label correlations, we augment the feature space by the previously elicited labels on each boosting round. Extensive experiments on various real-world datasets clearly validate the superiority of our proposed method. Haobo Wang 0001, Weiwei Liu 0003, Yang Zhao 0022, Chen Zhang 0020, Tianlei Hu, Gang Chen 0001 |
IJCAI | 4 |
| 2019 | Syntax-Aware Aspect-Level Sentiment Classification with Proximity-Weighted Convolution NetworkabstractIt has been widely accepted that Long Short-Term Memory (LSTM) network, coupled with attention mechanism and memory module, is useful for aspect-level sentiment classification. However, existing approaches largely rely on the modelling of semantic relatedness of an aspect with its context words, while to some extent ignore their syntactic dependencies within sentences. Consequently, this may lead to an undesirable result that the aspect attends on contextual words that are descriptive of other aspects. In this paper, we propose a proximity-weighted convolution network to offer an aspect-specific syntax-aware representation of contexts. In particular, two ways of determining proximity weight are explored, namely position proximity and dependency proximity. The representation is primarily abstracted by a bidirectional LSTM architecture and further enhanced by a proximity-weighted convolution. Experiments conducted on the SemEval 2014 benchmark demonstrate the effectiveness of our proposed approach compared with a range of state-of-the-art models is available at https://github.com/GeneZC/PWCN. Chen Zhang 0020, Qiuchi Li, Dawei Song 0001 |
SIGIR | 1 |