EDBT 2026 Demo / reviewers in the wild / expert
Yu Wang 0027
dblp:02/5889-27
· DBLP profile ↗
64ranked-venue papers
12as first author
46since 2021 · last 2026
0000-0001-9500-081XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 47 · 6 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 10 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MedS³: Towards Medical Slow Thinking with Self-Evolved Soft Dual-sided Process SupervisionabstractMedical language models face critical barriers to real-world clinical reasoning applications. However, mainstream efforts, which fall short in task coverage, lack fine-grained supervision for intermediate reasoning steps, and rely on proprietary systems, are still far from a versatile, credible and efficient language model for clinical reasoning usage. To this end, we propose MedS3, a self-evolving framework that imparts robust reasoning capabilities to small, deployable models. Starting with 8,000 curated instances sampled via a curriculum strategy across five medical domains and 16 datasets, we use a small base policy model to conduct Monte Carlo Tree Search (MCTS) for constructing rule-verifiable reasoning trajectories. Self-explored reasoning trajectories ranked by node values are used to bootstrap the policy model via reinforcement fine-tuning and preference learning. Moreover, we introduce a soft dual process reward model that incorporates value dynamics: steps that degrade node value are penalized, enabling fine-grained identification of reasoning errors even when the final answer is correct. Experiments on eleven benchmarks show that MedS3 outperforms the previous state-of-the-art medical model by +6.45 accuracy points and surpasses 32B-scale general-purpose reasoning models by +8.57 points. Additional empirical analysis further demonstrates that MedS3 achieves robust and faithful reasoning behavior. Shuyang Jiang, Yusheng Liao, Zhe Chen 0024, Ya Zhang 0002, Yanfeng Wang 0001, Yu Wang 0027 |
AAAI | 6 |
| 2026 | Miner: Mining Intrinsic Mastery for Data-Efficient RL in Large Reasoning ModelsabstractCurrent critic-free RL methods for large reasoning models suffer from severe inefficiency when training on positive homogeneous prompts (where all rollouts are correct), resulting in waste of rollouts due to zero advantage estimates.We introduce a radically simple yet powerful solution to Mine intrinsic mastery (MINER), that repurposes the policy's intrinsic uncertainty as a self-supervised reward signal, with no external supervision, auxiliary models, or additional inference cost.Our method pioneers two key innovations: (1) a token-level focal credit assignment mechanism that dynamically amplifies gradients on critical uncertain tokens while suppressing overconfident ones, and (2) adaptive advantage calibration to seamlessly integrate intrinsic and verifiable rewards.Evaluated across six reasoning benchmarks on Qwen3-4B and Qwen3-8B base models, MINER achieves state-of-theart performance among the other four algorithms, yielding up to 4.58 absolute gains in Pass@1 and 6.66 gains in Pass@K compared to GRPO.Comparison with other methods targeted at exploration enhancement further discloses the superiority of the two newly proposed innovations.This demonstrates that latent uncertainty exploitation is both necessary and sufficient for efficient and scalable RL training of reasoning models.Code is available at https://github.com/pixas/Miner. Shuyang Jiang, Ya Zhang 0002, Yanfeng Wang 0001, Yu Wang 0027 |
ACL (1) | 5 |
| 2026 | Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMsabstractHongcheng Liu, Yuhao Wang, Zhe Chen, Pingjie Wang, Zhiyuan Zhu, Yixuan Hou, Yanfeng Wang, Yu Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhe Chen 0024, Pingjie Wang, Yixuan Hou, Yanfeng Wang 0001, Yu Wang 0027 |
ACL (1) | 8 |
| 2026 | When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMsabstractMultimodal large language models (MLLMs) have shown strong capabilities across a broad range of benchmarks. However, most existing evaluations focus on passive inference, where models perform step-by-step reasoning under complete information. This setup is misaligned with real-world use, where seeing is not enough. This raises a fundamental question: Can MLLMs actively acquire missing evidence under incomplete information? To bridge this gap, we require the MLLMs to actively acquire missing evidence and iteratively refine decisions under incomplete information, by selecting a target image from a candidate pool without task-specific priors. To support systematic study, we propose GuessBench, a benchmark with both perception-oriented and knowledge-oriented images for evaluating active reasoning in MLLMs. We evaluate 20 superior MLLMs and find that performance on active reasoning lags far behind it on passive settings, indicating substantial room for improvement. Further analysis identifies fine-grained perception and timely decision-making as key challenges. Ablation studies show that perceptual enhancements benefit smaller models, whereas thinking-oriented methods provide consistent gains across model sizes. These results suggest promising directions for future research on multimodal active reasoning. Pingjie Wang, Siqu Ou, Yanfeng Wang 0001, Yu Wang 0027 |
ACL (1) | 6 |
| 2026 | SLoRA: Balancing Plasticity and Forgetting in Large Language Models for Continual LearningabstractLarge language models (LLMs) have achieved remarkable success across diverse tasks through large-scale pretraining.However, they remain prone to catastrophic forgetting in continual learning.To the best of our knowledge, this is the first work to identify noise accumulation in LoRA updates as a key cause of forgetting in continual learning.A preliminary two-task experiment demonstrates that removing less important components of the second task's LoRA parameters improves performance on the first task, suggesting that later updates introduce noisy interference.Building on this insight, we propose Subspace-Denoised Low-Rank Adaptation (SLoRA), a simple and effective framework that filters noisy components from LoRA updates via subspace similarity with the base model.SLoRA is a regularizationfree method without accessing data or gradients from previous tasks or modifying the training process.It offers two variants, SLoRA-Pre and SLoRA-Post, for online and offline continual learning, respectively.Extensive experiments across tasks and models validate the effectiveness of SLoRA.It improves final accuracy by up to 12%, reduces forgetting by 29%, and filters out over 30% of LoRA parameters identified as noisy.Our code is available at https://github.com/alina1031/SLoRA. Yusheng Liao, Yanfeng Wang 0001, Yu Wang 0027 |
ACL (1) | 4 |
| 2026 | Disentangled diffusion model for 3D molecular generation with protein-ligand interaction priorsabstractMOTIVATION: Structure-based drug design (SBDD) aims to generate ligand molecules that tightly bind to specific protein targets, a critical step in drug discovery. Diffusion models have shown promise for this task, yet existing methods struggle to effectively incorporate protein-ligand interaction priors during generation. Most approaches rely on protein-specific structural priors that remain fixed throughout generation, limiting molecular diversity and failing to capture the dynamic interplay between protein pockets and ligand atoms, which is essential for achieving high binding affinity. RESULTS: We propose DPDiff, a disentangled prior-conditioned diffusion model for protein-specific 3D molecular generation. DPDiff introduces two complementary interaction prior networks that capture geometry-based spatial interactions and sequence-based interactions robust to structural noise. During generation, the model dynamically extracts interaction priors using intermediate diffusion predictions and adaptively fuses them via a time-dependent adapter. A disentangled denoising network balances prior guidance with generative flexibility. Experiments on the CrossDocked2020 dataset demonstrate that DPDiff generates molecules with more realistic 3D structures and state-of-the-art binding affinities, achieving an average Vina Dock score of -8.58 and a high affinity ratio of 69.4%, outperforming existing methods while maintaining favorable drug-likeness and synthetic accessibility. AVAILABILITY AND IMPLEMENTATION: The source code of DPDiff is available at https://github.com/ZerinHwang03/DPDiff. Zhilin Huang, Ling Yang 0006, Chujun Qin, Yifei Xing 0001, Xiangxin Zhou, Yu Wang 0027, Xin Gao 0001, Wenming Yang |
Bioinform. | 8 |
| 2025 | AnyTalk: Multi-modal Driven Multi-domain Talking Head GenerationabstractCross-domain talking head generation, such as animating a static cartoon animal photo with real human video, is crucial for personalized content creation. However, prior works typically rely on domain-specific frameworks and paired videos, limiting its utility and complicating its architecture with additional motion alignment modules. Addressing these shortcomings, we propose Anytalk, a unified framework that eliminates the need for paired data and learns a shared motion representation across different domains. The motion is represented by canonical 3D keypoints extracted using an unsupervised 3D keypoint detector. Further, we propose an expression consistency loss to improve the accuracy of facial dynamics in video generation. Additionally, we present AniTalk, a comprehensive dataset designed for advanced multi-modal cross-domain generation. Our experiments demonstrate that Anytalk excels at generating high-quality, multi-modal talking head videos, showcasing remarkable generalization capabilities across diverse domains. Yu Wang 0027, Yunfei Liu 0001, Fa-Ting Hong, Lijian Lin, Yu Li 0003 |
AAAI | 1 |
| 2025 | Towards Omni-RAG: Comprehensive Retrieval-Augmented Generation for Large Language Models in Medical ApplicationsabstractLarge language models hold promise for addressing medical challenges, such as medical diagnosis reasoning, research knowledge acquisition, clinical decision-making, and consumer health inquiry support. However, they often generate hallucinations due to limited medical knowledge. Incorporating external knowledge is therefore critical, which necessitates multi-source knowledge acquisition. We address this challenge by framing it as a source planning problem, which is to formulate context-appropriate queries tailored to the attributes of diverse sources. Existing approaches either overlook source planning or fail to achieve it effectively due to misalignment between the model’s expectation of the sources and their actual content. To bridge this gap, we present MedOmniKB, a repository comprising multigenre and multi-structured medical knowledge sources. Leveraging these sources, we propose the Source Planning Optimisation method, which enhances multi-source utilisation. Our approach involves enabling an expert model to explore and evaluate potential plans while training a smaller model to learn source alignment. Experimental results demonstrate that our method substantially improves multi-source planning performance, enabling the optimised small model to achieve state-of-the-art results in leveraging diverse medical knowledge sources. Zhe Chen 0024, Yusheng Liao, Shuyang Jiang, Pingjie Wang, Yiqiu Guo, Yanfeng Wang 0001, Yu Wang 0027 |
ACL (1) | 7 |
| 2025 | ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical AgentsabstractLarge Language Models (LLMs) have shown promising potential in the medical domain, assisting with tasks like clinical note generation and patient communication.However, current LLMs are limited to text-based communication, hindering their ability to interact with diverse forms of information in clinical environments.Despite clinical agents succeeding in diverse signal interaction, they are oriented to a single clinical scenario and hence fail for broader applications.To evaluate clinical agents holistically, we propose ClinicalAgent Bench (CAB), a comprehensive medical agent benchmark consisting of 18 tasks across five key realistic clinical dimensions.Building on this, we introduce REFLECTOOL, a novel framework that excels at utilizing domain-specific tools within two stages.The first optimization stage progressively enlarges a long-term memory by saving successful solving processes and toolwise experience of agents in a tiny pre-defined training set.In the following inference stage, REFLECTOOL can search for supportive successful demonstrations from already built longterm memory to guide the tool selection strategy, and a verifier improves the tool usage according to the tool-wise experience with two verification methods-iterative refinement and candidate selection.Extensive experiments on CAB demonstrate that REFLECTOOL surpasses the pure LLMs with more than 10 points and the well-established agent-based methods with 3 points, highlighting its adaptability and effectiveness in solving complex clinical tasks.Our code and datasets are available at https: //github.com/BlueZeros/ReflecTool. Yusheng Liao, Shuyang Jiang, Yanfeng Wang 0001, Yu Wang 0027 |
ACL (1) | 4 |
| 2025 | EvolveBench: A Comprehensive Benchmark for Assessing Temporal Awareness in LLMs on Evolving KnowledgeabstractLarge language models (LLMs) are trained on extensive historical corpora, but their ability to understand time and maintain temporal awareness of time-evolving factual knowledge remains limited. Previous studies often neglect the critical aspect of utilizing knowledge from various sources. To address this gap, we introduce EvolveBench, a comprehensive benchmark that evaluates temporal competence along five key dimensions: Cognition, which examines the ability to recall and contextualize historical facts. Awareness, which tests LLMs’ awareness of temporal misalignment between external inputs and the temporal context of a query. Trustworthiness, which assesses whether models can identify and appropriately refuse queries based on invalid timestamps. Understanding, which focuses on interpreting both explicit dates and implicit historical markers. Finally, reasoning evaluates the capacity to analyze temporal relationships and draw accurate inferences. Evaluating 15 widely used LLMs on EvolveBench shows that GPT-4o achieves the highest average EM score of 79.36, while the open-source Llama3.1-70B demonstrates notable strength in handling temporally misaligned contexts with an average score of 72.47. Despite these advances, all models still struggle with handling temporal misaligned context. Our code and dataset are available at https://github.com/zzysjtuiwct/EvolveBench. Yusheng Liao, Zhe Chen 0024, Yunfeng Guan 0001, Yanfeng Wang 0001, Yu Wang 0027 |
ACL (1) | 7 |
| 2025 | DSVD: Dynamic Self-Verify Decoding for Faithful Generation in Large Language ModelsabstractThe reliability of large language models remains a critical challenge, particularly due to their susceptibility to hallucinations and factual inaccuracies during text generation.Existing solutions either underutilize models' selfcorrection with preemptive strategies or use costly post-hoc verification.To further explore the potential of real-time self-verification and correction, we present Dynamic Self-Verify Decoding (DSVD), a novel decoding framework that enhances generation reliability through real-time hallucination detection and efficient error correction.DSVD integrates two key components: (1) parallel self-verification architecture for continuous quality assessment, (2) dynamic rollback mechanism for targeted error recovery.Extensive experiments across five benchmarks demonstrate DSVD's effectiveness, achieving significant improvement in truthfulness (Quesetion-Answering) and factual accuracy (FActScore).Results show the DSVD can be further incorporated with existing faithful decoding methods to achieve stronger performance.Our work establishes that real-time self-verification during generation offers a viable path toward more faithful language models without sacrificing practical deployability. YiQiu Guo, Zhe Chen 0024, Pingjie Wang, Yusheng Liao, Ya Zhang 0002, Yanfeng Wang 0001, Yu Wang 0027 |
EMNLP | 8 |
| 2025 | DICE: Structured Reasoning in LLMs through SLM-Guided Chain-of-Thought CorrectionabstractWhen performing reasoning tasks with userspecific requirements, such as strict output formats, large language models (LLMs) often prioritize reasoning over adherence to detailed instructions.Fine-tuning LLMs on supervised datasets to address this is impractical due to high computational costs and limited parameter access.To tackle this, we propose DICE, a lightweight framework that guides small language models (SLMs) to refine LLMs' outputs through chain-of-thought (CoT) correction.DICE decouples the process by first prompting LLMs to generate natural language responses, then using trained SLMs to analyze and refine these outputs to meet structured output specifications.This framework preserves LLMs' broad knowledge and reasoning capabilities while ensuring the outputs conform to user demands.Specifically, DICE first constructs structured CoT adaptation datasets via a two-stage method and subsequently applies a dual-tuning strategy to fine-tune SLMs for generating structured outputs in an analyze-thenanswer pattern. 1 Experiments demonstrate that DICE improves the average format accuracy and content correctness of LLM outputs by 35.4% and 29.4%, respectively, achieving stateof-the-art (SOTA) performance over other competitive baselines. Yusheng Liao, Zhe Chen 0024, Yanfeng Wang 0001, Yu Wang 0027 |
EMNLP | 5 |
| 2025 | AuscMLLM: Bridging Classification and Reasoning in Heart Sound Analysis with a Multimodal Large Language ModelabstractThis study introduces a multimodal large language model capable of not only accomplishing various heart sound tasks but also providing reasoning, marking an advancement in the field of medical diagnostics. The model’s innovation stems from a collaboration with experts to collect a novel dataset designed specifically for reasoning tasks, addressing the limitations of existing datasets that lacked this capability. Our model integrates multiple novel methodologies to enhance diagnostic accuracy, including the incorporation of knowledge from relevant textbooks through pre-training, the employment of an audio feature extractor optimized for heart sound-text alignment, and a logit adjustment loss tailored for large language model to mitigate the challenge of imbalanced data categories. This approach not only sets a new standard for heart sound analysis but also paves the way for more interpretable and comprehensive diagnostic models in healthcare. Pingjie Wang, Liudan Zhao, Ya Zhang 0002, Xin Sun 0020, Yanfeng Wang 0001, Yu Wang 0027 |
ICASSP | 9 |
| 2025 | Fine-tuning with Reserved Majority for Noise ReductionabstractParameter-efficient fine-tuning (PEFT) has revolutionized supervised fine-tuning, where LoRA and its variants gain the most popularity due to their low training costs and zero inference latency.
However, LoRA tuning not only injects knowledgeable features but also noisy hallucination during fine-tuning, which hinders the utilization of tunable parameters with the increasing LoRA rank.
In this work, we first investigate in-depth the redundancies among LoRA parameters with substantial empirical studies.
Aiming to resemble the learning capacity of high ranks from the findings, we set up a new fine-tuning framework, \textbf{P}arameter-\textbf{Re}dundant \textbf{F}ine-\textbf{T}uning (\preft), which follows the vanilla LoRA tuning process but is required to reduce redundancies before merging LoRA parameters back to pre-trained models.
Based on this framework, we propose \textbf{No}ise reduction with \textbf{R}eserved \textbf{M}ajority~(\norm), which decomposes the LoRA parameters into majority parts and redundant parts with random singular value decomposition.
The major components are determined by the proposed \search method, specifically employing subspace similarity to confirm the parameter groups that share the highest similarity with the base weight.
By employing \norm, we enhance both the learning capacity and benefits from larger ranks, which consistently outperforms both LoRA and other \preft-based methods on various downstream tasks, such as general instruction tuning, math reasoning and code generation.
Code is available at \url{https://github.com/pixas/NoRM}. Shuyang Jiang, Yusheng Liao, Ya Zhang 0002, Yanfeng Wang 0001, Yu Wang 0027 |
ICLR | 5 |
| 2025 | SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
Yixuan Hou, Heyang Liu, Ziyang Cheng 0002, Ronghua Wu, Qunshan Gu, Yanfeng Wang 0001, Yu Wang 0027 |
INTERSPEECH | 8 |
| 2025 | TaxDiff: taxonomic-guided diffusion model for protein sequence generation
Zongying Lin, Hao Li 0073, Liuzhenghao Lv, Yu Wang 0027, Bin Lin 0014, Junwu Zhang, Calvin Yu-Chian Chen, Li Yuan 0007, Yonghong Tian 0001 |
Sci. China Inf. Sci. | 4 |
| 2025 | Adaptive Fuzzy Positive Learning for Annotation-Scarce Semantic Segmentation
Pengchong Qiao, Yu Wang 0027, Chang Liu 0030, Baigui Sun, Zhennan Wang 0001, Xiawu Zheng, Rongrong Ji, Jie Chen 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | Redundancy-Adaptive Multimodal Learning for imperfect data
Mengxi Chen, Jiangchao Yao, Linyu Xing, Yu Wang 0027, Ya Zhang 0002, Yanfeng Wang 0001 |
Neural Networks | 4 |
| 2025 | Dual-Level Masked Semantic Inference for Semi-Supervised Semantic SegmentationabstractSemi-supervised semantic segmentation pursues a holistic pixel-wise understanding of unseen images with limited annotation. To this end, existing methods focus on regularizing per-pixel prediction consistency within unlabeled data, while rarely modeling contextual relationships. But in fact, contextual semantics can provide valuable clues for scene understanding like inner-object continuity and spatial relationships' causality. Thus, in this paper, we propose a Dual-level Masked Semantics Inference (DMSI) that takes the initiative to explicitly learn contextual relationships via enforcing our model to infer the semantics of a pixel according to its surrounding contexts. This allows our model to exhaust accurate semantics by incorporating inter-pixel context clues, further leading to comprehensive segmentation. Specifically, DMSI comprises two main components. 1) Dual-level mask consistency regularization (DMCR) that learns the ability of semantics inference by aligning the predictions of masked views with the prediction of the complete view. The masked views here come from both the image level and feature level, where our model captures low-level attributes and high-level representations respectively. 2) AdaMask that provides a proper mask position and ratio for each image, guiding our model to focus on semantic-rich regions while providing balanced training between hard and easy samples. Through learning the ability of semantic inferring, DMSI remarkably enhances the interaction between pixels, further progressively intensifying the understanding of semantics. Extensive experiments under various settings on Cityscapes and Pascal VOC 2012 show that DMSI achieves new state-of-the-art performances. Furthermore, analysis indicates that our method has superiority in mining inter-pixel semantic relationships and improving robustness facing noise corruption. Qiankun Ma, Ziyao Zhang 0003, Pengchong Qiao, Yu Wang 0027, Rongrong Ji, Chang Liu 0047, Jie Chen 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | M³AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture DatasetabstractZhe Chen, Heyang Liu, Wenyi Yu, Guangzhi Sun, Hongcheng Liu, Ji Wu, Chao Zhang, Yu Wang, Yanfeng Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhe Chen 0024, Heyang Liu, Wenyi Yu, Guangzhi Sun, Ji Wu 0002, Chao Zhang 0031, Yu Wang 0027, Yanfeng Wang 0001 |
ACL (1) | 8 |
| 2024 | MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in PerceptionabstractRecent advancements in Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in visual perception and understanding.However, these models also suffer from hallucinations, which limit their reliability as AI systems.We believe that these hallucinations are partially due to the models' struggle with understanding what they can and cannot perceive from images, a capability we refer to as self-awareness in perception.Despite its importance, this aspect of MLLMs has been overlooked in prior studies.In this paper, we aim to define and evaluate the selfawareness of MLLMs in perception.To do this, we first introduce the knowledge quadrant in perception, which helps define what MLLMs know and do not know about images.Using this framework, we propose a novel benchmark, the Self-Awareness in Perception for MLLMs (MM-SAP), specifically designed to assess this capability.We apply MM-SAP to a variety of popular MLLMs, offering a comprehensive analysis of their self-awareness and providing detailed insights.The experiment results reveal that current MLLMs possess limited selfawareness capabilities, pointing to a crucial area for future advancement in the development of trustworthy MLLMs. Yusheng Liao, Heyang Liu, Yanfeng Wang 0001, Yu Wang 0027 |
ACL (1) | 6 |
| 2024 | CE-VDG: Counterfactual Entropy-based Bias Reduction for Video-grounded Dialogue GenerationabstractThe Video-Grounded Dialogue generation (VDG) is a challenging task requiring a comprehensive understanding of the multi-modal information to produce a pertinent response. However, VDG models may rely on dataset bias as a shortcut and fail to learn the multi-modal knowledge from both video and audio. Counterfactual reasoning is an effective method that can estimate and eliminate bias on some special aspects of classification tasks. However, conventional counterfactual reasoning cannot be applied to VDG tasks directly due to the BPE algorithm. In this paper, we reformulate the counterfactual reasoning from the information entropy perspective and extend it from the classification task to the generative task, which can effectively reduce the question-related bias in the auto-regressive generation task. We design CE-VDG to demonstrate the effectiveness in bias elimination of the reformulated counterfactual reasoning by using the proposed counterfactual entropy as an external loss. Extensive experiment results on two popular VDG datasets show the superiority of CE-VDG over the existing baseline method, demonstrating the effective debiasing capability in our model considering counterfactual entropy. Pingjie Wang, Yanfeng Wang 0001, Yu Wang 0027 |
LREC/COLING | 5 |
| 2024 | Pruning before Fine-tuning: A Retraining-free Compression Framework for Pre-trained Language ModelsabstractStructured pruning is an effective technique for compressing pre-trained language models (PLMs), reducing model size and improving inference speed for efficient deployment. However, most of existing pruning algorithms require retraining, leading to additional computational overhead. While some retraining-free approaches have been proposed for classification tasks, they still require a fully fine-tuned model for the task, and may cause catastrophic performance degradation on generative tasks. To address these challenges, we propose P-pruning (pre-pruning), an innovative task-specific compression framework. P-pruning prunes redundant modules of PLMs before fine-tuning, reducing the costs associated with fine-tuning. We also introduce a pruning algorithm for this framework, which includes two techniques: (1) module clustering, which clusters the outputs of all heads and neurons based on the task input; and (2) centroid selection, which identifies the most salient element in each cluster and prunes the others. We apply our method to BERT and GPT-2 and evaluate its effectiveness on GLUE, SQuAD, WikiText-2, WikiText-103, and PTB datasets. Experimental results demonstrate that our approach achieves higher performance in both classification and generative tasks, while also reducing the time required for fine-tuning. Pingjie Wang, Yanfeng Wang 0001, Yu Wang 0027 |
LREC/COLING | 4 |
| 2024 | AdaShield : Safeguarding Multimodal Large Language Models from Structure-Based Attack via Adaptive Shield Prompting
Yu Wang 0027, Xiaogeng Liu, Yu Li 0003, Muhao Chen 0001, Chaowei Xiao |
ECCV (20) | 1 |
| 2024 | ParCo: Part-Coordinating Text-to-Motion Synthesis
Qiran Zou, Shangyuan Yuan, Shian Du, Yu Wang 0027, Chang Liu 0030, Yi Xu 0008, Jie Chen 0001, Xiangyang Ji |
ECCV (56) | 4 |
| 2024 | RA2FD: Distilling Faithfulness into Efficient Dialogue SystemsabstractGenerating faithful and fast responses is crucial in the knowledge-grounded dialogue.Retrieval Augmented Generation (RAG) strategies are effective but are inference inefficient, while previous Retrieval Free Generations (RFG) are more efficient but sacrifice faithfulness.To solve this faithfulness-efficiency trade-off dilemma, we propose a novel retrieval-free model training scheme named Retrieval Augmented to Retrieval Free Distillation (RA2FD) to build a retrieval-free model that achieves higher faithfulness than the previous RFG method while maintaining inference efficiency.The core idea of RA2FD is to use a teacher-student framework to distill the faithfulness capacity of a teacher, which is an oracle RAG model that generates multiple knowledge-infused responses.The student retrieval-free model learns how to generate faithful responses from these teacher labels through sequence-level distillation and contrastive learning.Experiment results show that RA2FD let the faithfulness performance of an RFG model surpass the previous SOTA RFG baseline on three knowledge-grounded dialogue datasets by an average of 33% and even matching an RAG model's performance while significantly improving inference efficiency.Our code is available at https:// github.com/zzysjtuiwct/RA2FD. Yusheng Liao, Chenxin Xu, Yunfeng Guan 0001, Yanfeng Wang 0001, Yu Wang 0027 |
EMNLP | 6 |
| 2024 | MSG-BART: Multi-Granularity Scene Graph-Enhanced Encoder-Decoder Language Model for Video-Grounded Dialogue GenerationabstractGenerating dialogue grounded in videos requires a high level of understanding and reasoning about the visual scenes in the videos. However, existing large visual-language models are not effective due to their latent features and decoder-only structure, especially with respect to spatio-temporal relationship reasoning. In this paper, we propose a novel approach named MSG-BART, which enhances the integration of video information by incorporating a multi-granularity spatio-temporal scene graph into an encoder-decoder pre-trained language model. Specifically, we integrate the global and local scene graph into the encoder and decoder, respectively, to improve both overall perception and target reasoning capability. To further improve the information selection capability, we propose a multi-pointer network to facilitate selection between text and video. Extensive experiments are conducted on three video-grounded dialogue benchmarks, which show the significant superiority of the proposed MSG-BART compared to a range of state-of-the-art approaches. Zhe Chen 0024, Pingjie Wang, Yanfeng Wang 0001, Yu Wang 0027 |
ICASSP | 6 |
| 2024 | TAIA: Large Language Models are Out-of-Distribution Data LearnersabstractFine-tuning on task-specific question-answer pairs is a predominant method for enhancing the performance of instruction-tuned large language models (LLMs) on downstream tasks. However, in certain specialized domains, such as healthcare or harmless content generation, it is nearly impossible to obtain a large volume of high-quality data that matches the downstream distribution. To improve the performance of LLMs in data-scarce domains with domain-mismatched data, we re-evaluated the Transformer architecture and discovered that not all parameter updates during fine-tuning contribute positively to downstream performance. Our analysis reveals that within the self-attention and feed-forward networks, only the fine-tuned attention parameters are particularly beneficial when the training set's distribution does not fully align with the test set. Based on this insight, we propose an effective inference-time intervention method: \uline{T}raining \uline{A}ll parameters but \uline{I}nferring with only \uline{A}ttention (TAIA). We empirically validate TAIA using two general instruction-tuning datasets and evaluate it on seven downstream tasks involving math, reasoning, and knowledge understanding across LLMs of different parameter sizes and fine-tuning techniques. Our comprehensive experiments demonstrate that TAIA achieves superior improvements compared to both the fully fine-tuned model and the base model in most scenarios, with significant performance gains. The high tolerance of TAIA to data mismatches makes it resistant to jailbreaking tuning and enhances specialized tasks using general data. Code is available in \url{https://github.com/pixas/TAIA_LLM}. Shuyang Jiang, Yusheng Liao, Ya Zhang 0002, Yanfeng Wang 0001, Yu Wang 0027 |
NeurIPS | 5 |
| 2024 | SubgDiff: A Subgraph Diffusion Model to Improve Molecular Representation LearningabstractMolecular representation learning has shown great success in advancing AI-based drug discovery. A key insight of many recent works is that the 3D geometric structure of molecules provides essential information about their physicochemical properties. Recently, denoising diffusion probabilistic models have achieved impressive performance in molecular 3D conformation generation. However, most existing molecular diffusion models treat each atom as an independent entity, overlooking the dependency among atoms within the substructures. This paper introduces a novel approach that enhances molecular representation learning by incorporating substructural information in the diffusion model framework. We propose a novel diffusion model termed SubgDiff for involving the molecular subgraph information in diffusion. Specifically, SubgDiff adopts three vital techniques: i) subgraph prediction, ii) expectation state, and iii) k-step same subgraph diffusion, to enhance the perception of molecular substructure in the denoising network. Experiments on extensive downstream tasks, especially the molecular force predictions, demonstrate the superior performance of our approach. Jiying Zhang, Zijing Liu, Yu Wang 0027, Bin Feng 0001, Yu Li 0003 |
NeurIPS | 3 |
| 2024 | Annotation-free Audio-Visual SegmentationabstractThe objective of Audio-Visual Segmentation (AVS) is to localise the sounding objects within visual scenes by accurately predicting pixel-wise segmentation masks. To tackle the task, it involves a comprehensive consideration of both the data and model aspects. In this paper, first, we initiate a novel pipeline for generating artificial data for the AVS task without extra manual annotations. We leverage existing image segmentation and audio datasets and match the image-mask pairs with its corresponding audio samples using category labels in segmentation datasets, that allows us to effortlessly compose (image, audio, mask) triplets for training AVS models. The pipeline is annotation-free and scalable to cover a large number of categories. Additionally, we introduce a lightweight model SAMA-AVS which adapts the pre-trained segment anything model (SAM) to the AVS task. By introducing only a small number of trainable parameters with adapters, the proposed model can effectively achieve adequate audio-visual fusion and interaction in the encoding stage with vast majority of parameters fixed. We conduct extensive experiments, and the results show our proposed model remarkably surpasses other competing methods. Moreover, by using the proposed model pretrained with our synthetic data, the performance on real AVSBench data is further improved, achieving 83.17 mIoU on S4 subset and 66.95 mIoU on MS3 set. The project page is https://jinxiang-liu.github.io/anno-free-AVS/. Jinxiang Liu, Yu Wang 0027, Chen Ju, Chaofan Ma, Ya Zhang 0002, Weidi Xie |
WACV | 2 |
| 2024 | DialogMCF: Multimodal Context Flow for Audio Visual Scene-Aware DialogabstractIn recent years, Audio Visual Scene-Aware Dialog (AVSD) has been an active research task in the multimodal dialogue community and has also been a core part of the Dialog System Technology Challenge (DSTC). This task is an extension of conventional visual question answering, where video-relevant answers must be generated taking into account multimodal contextual information from previous dialogue rounds. Despite recent advances in the AVSD task, there are still two major challenges in developing such a system: how to model the multimodal contextual information of multiple rounds of dialogues and how to integrate audio-visual information into the generation of textual responses. To tackle these two challenges, in this paper we propose a novel model, named DialogMCF, which constructs a multimodal context flow model to generate responses that are relevant to video scenes. This proposed context flow modeling can track the dynamics of the topic information across multiple rounds of dialogue history. To achieve an effective fusion of multimodal information, we propose an audio-visual memory network with cross-modality aligned features to model long multimodal dialogue context, and thus enhance the flow modeling. Furthermore, we make attempts to improve the performance of the proposed DialogMCF model with manual descriptions and explore the incorporation of temporal reasoning information. Extensive experiments on the DSTC AVSD datasets show that, compared to a range of baseline methods, the proposed method can yield state-of-art dialogue generation performance on most metrics when integrating the video descriptions. Zhe Chen 0024, Yu Wang 0027 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | Leveraging Diverse Modeling Contexts With Collaborating Learning for Neural Machine TranslationabstractAutoregressive (AR) and Non-autoregressive (NAR) models are two types of generative models for Neural Machine Translation (NMT). AR models predict tokens in a word-byword manner and can effectively capture the distribution of real translations. NAR models predict tokens by extracting bidirectional contextual information which can improve the inference speed but they suffer from performance degradation. Previous works utilized AR models to enhance NAR models by reducing the training data's complexity or incorporating the global information into AR models by virtue of NAR models. However, those investigated methods only take advantage of the contextual information of a single type of model while neglecting the diversity in the contextual information that can be provided by different types of models. In this paper, we propose a novel generic collaborative learning method, DCMCL, where AR and NAR models are treated as collaborators instead of teachers and students. To hierarchically leverage the bilateral contextual information, token-level mutual learning and sequence-level contrastive learning are adopted between AR and NAR models. Extensive experiments on four widely used benchmarks show that the proposed DCMCL method can simultaneously improve both AR and NAR models with up to 1.38 and 2.98 BLEU scores, respectively, and can also outperform the current best-unified model with up to 0.97 BLEU scores for both AR and NAR decoding. Yusheng Liao, Yanfeng Wang 0001, Yu Wang 0027 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Enhanced Multimodal Representation Learning with Cross-modal KDabstractThis paper explores the tasks of leveraging auxiliary modalities which are only available at training to enhance multimodal representation learning through cross-modal Knowledge Distillation (KD). The widely adopted mutual information maximization-based objective leads to a short-cut solution of the weak teacher, i.e., achieving the maximum mutual information by simply making the teacher model as weak as the student model. To prevent such a weak solution, we introduce an additional objective term, i.e., the mutual information between the teacher and the auxiliary modality model. Besides, to narrow down the information gap between the student and teacher, we further propose to minimize the conditional entropy of the teacher given the student. Novel training schemes based on contrastive learning and adversarial learning are designed to optimize the mutual information and the conditional entropy, respectively. Experimental results on three popular multimodal benchmark datasets have shown that the proposed method outperforms a range of state-of-the-art approaches for video recognition, video retrieval and emotion classification. Mengxi Chen, Linyu Xing, Yu Wang 0027, Ya Zhang 0002 |
CVPR | 3 |
| 2023 | Fuzzy Positive Learning for Semi-Supervised Semantic SegmentationabstractSemi-supervised learning (SSL) essentially pursues class boundary exploration with less dependence on human annotations. Although typical attempts focus on ameliorating the inevitable error-prone pseudo-labeling, we think differently and resort to exhausting informative semantics from multiple probably correct candidate labels. In this paper, we introduce Fuzzy Positive Learning (FPL) for accurate SSL semantic segmentation in a plug-and-play fashion, targeting adaptively encouraging fuzzy positive predictions and suppressing highly-probable negatives. Being conceptually simple yet practically effective, FPL can remarkably alleviate interference from wrong pseudo labels and progressively achieve clear pixel-level semantic discrimination. Concretely, our FPL approach consists of two main components, including fuzzy positive assignment (FPA) to provide an adaptive number of labels for each pixel and fuzzy positive regularization (FPR) to restrict the predictions of fuzzy positive categories to be larger than the rest under different perturbations. Theoretical analysis and extensive experiments on Cityscapes and VOC 2012 with consistent performance gain justify the superiority of our approach. Codes are provided in https://github.com/qpc1611094/FPL. Pengchong Qiao, Zhidan Wei, Yu Wang 0027, Zhennan Wang 0001, Guoli Song, Xiangyang Ji, Chang Liu 0030, Jie Chen 0001 |
CVPR | 3 |
| 2023 | Out-of-Distributed Semantic Pruning for Robust Semi-Supervised LearningabstractRecent advances in robust semi-supervised learning (SSL) typically filter out-of-distribution (OOD) information at the sample level. We argue that an overlooked problem of robust SSL is its corrupted information on semantic level, practically limiting the development of the field. In this paper, we take an initial step to explore and propose a unified framework termed OOD Semantic Pruning (OSP), which aims at pruning OOD semantics out from in-distribution (ID) features. Specifically, (i) we propose an aliasing OOD matching module to pair each ID sample with an OOD sample with semantic overlap. (ii) We design a soft orthogonality regularization, which first transforms each ID feature by suppressing its semantic component that is collinear with paired OOD sample. It then forces the predictions before and after soft orthogonality decomposition to be consistent. Being practically simple, our method shows a strong performance in OOD detection and ID classification on challenging benchmarks. In particular, OSP surpasses the previous state-of-the-art by 13.7% on accuracy for ID classification and 5.9% on AUROC for OOD detection on TinyImageNet dataset. The source codes are publicly available at https://github.com/rain305f/OSP. Yu Wang 0027, Pengchong Qiao, Chang Liu 0030, Guoli Song, Xiawu Zheng, Jie Chen 0001 |
CVPR | 1 |
| 2023 | Self-Improvement of Non-autoregressive Model via Sequence-Level DistillationabstractAlthough Non-autoregressive Transformer (NAT) models have achieved great success in terms of fast inference speed, this speedup comes with a performance drop due to the inherent multi-modality problem of the NAT model.Previous works commonly alleviate this problem by replacing the target side of the raw data with distilled data generated by Autoregressive Transformer (AT) models.However, the multimodality problem in the distilled data is still significant and thus limits further improvement of the NAT models.In this paper, we propose a method called Sequence-Level Self-Distillation (SLSD), which aims to generate distilled data by the NAT model itself, eliminating the need for additional teacher networks.Furthermore, SLSD can adapt to different NAT models without precise adjustments since the self-distilled data is generated from the same types of NAT models.We conduct extensive experiments on WMT14 EN↔DE and WMT16 EN↔RO and choose five classic NAT models as the backbones to validate the generality and effectiveness of SLSD.The results show that our approach can consistently improve all models on both raw data and distilled data without sacrificing the inference speed. Yusheng Liao, Shuyang Jiang, Yu Wang 0027, Yanfeng Wang 0001 |
EMNLP | 4 |
| 2023 | Knowledge-Aware Bayesian Co-Attention for Multimodal Emotion RecognitionabstractMultimodal emotion recognition is a challenging research area that aims to fuse different modalities to predict human emotion. However, most existing models that are based on attention mechanisms have difficulty in learning emotionally relevant parts on their own. To solve this problem, we propose to incorporate external emotion-related knowledge in the co-attention based fusion of pre-trained models. To effectively incorporate this knowledge, we enhance the co-attention model with a Bayesian attention module (BAM) where a prior distribution is estimated using the emotion-related knowledge. Experimental results on the IEMOCAP dataset show that the proposed approach can outperform several state-of-the-art approaches by at least 0.7% unweighted accuracy (UA). Yu Wang 0027, Yanfeng Wang 0001 |
ICASSP | 2 |
| 2023 | Pushing the Limits of Unsupervised Unit Discovery for SSL Speech Representation
Ziyang Ma 0001, Zhisheng Zheng, Guanrou Yang, Yu Wang 0027, Chao Zhang 0031, Xie Chen 0001 |
INTERSPEECH | 4 |
| 2023 | Unsupervised Active Learning: Optimizing Labeling Cost-Effectiveness for Automatic Speech Recognition
Zhisheng Zheng, Ziyang Ma 0001, Yu Wang 0027, Xie Chen 0001 |
INTERSPEECH | 3 |
| 2023 | Contrastive Learning Based ASR Robust Knowledge Selection For Spoken Dialogue System
Yusheng Liao, Yu Wang 0027, Yunfeng Guan 0001 |
INTERSPEECH | 3 |
| 2023 | Uncertainty-Guided End-to-End Audio-Visual Speaker Diarization for Far-Field RecordingsabstractAudio-visual speaker diarization refers to the task of identifying "who spoke when" by using both audio and video data. Although previous fusion-based approaches have shown exceptional performance over audio-only methods, they have mainly focused on high-quality data and have not accounted for the impacts of acoustic noise or missing faces. To address these limitations, we propose a novel uncertainty-aware end-to-end audio-visual speaker diarization (UAV-SD) approach in this paper. Our approach leverages both framewise inter- and intra-modal confidence to achieve more effective and robust speaker diarization. By taking into account the uncertainty of the data, UAV-SD can achieve better diarization performance even in noisy or low-quality recordings. Additionally, our approach is compatible with multi-channel audio signals without the need to retrain the model, making it a more versatile solution. To evaluate the effectiveness of our approach, we conduct extensive experiments on the Multi-modal Information Based Speech Processing (MISP) 2022 Challenge datasets which consist of far-field audio and video data. The results show that UAV-SD is able to yield significant performance gains compared to baseline methods for both single and multi-channel data, demonstrating its effectiveness in real-world scenarios. Mengxi Chen, Yanfeng Wang 0001, Yu Wang 0027 |
ACM Multimedia | 4 |
| 2023 | Discover and Align Taxonomic Context Priors for Open-world Semi-Supervised LearningabstractOpen-world Semi-Supervised Learning (OSSL) is a realistic and challenging task, aiming to classify unlabeled samples from both seen and novel classes using partially labeled samples from the seen classes.
Previous works typically explore the relationship of samples as priors on the pre-defined single-granularity labels to help novel class recognition. In fact, classes follow a taxonomy and samples can be classified at multiple levels of granularity, which contains more underlying relationships for supervision. We thus argue that learning with single-granularity labels results in sub-optimal representation learning and inaccurate pseudo labels, especially with unknown classes. In this paper, we take the initiative to explore and propose a uniformed framework, called Taxonomic context prIors Discovering and Aligning (TIDA), which exploits the relationship of samples under various granularity. It allows us to discover multi-granularity semantic concepts as taxonomic context priors (i.e., sub-class, target-class, and super-class), and then collaboratively leverage them to enhance representation learning and improve the quality of pseudo labels.
Specifically, TIDA comprises two components: i) A taxonomic context discovery module that constructs a set of hierarchical prototypes in the latent space to discover the underlying taxonomic context priors; ii) A taxonomic context-based prediction alignment module that enforces consistency across hierarchical predictions to build the reliable relationship between classes among various granularity and provide additions supervision. We demonstrate that these two components are mutually beneficial for an effective OSSL framework, which is theoretically explained from the perspective of the EM algorithm. Extensive experiments on seven commonly used datasets show that TIDA can significantly improve the performance and achieve a new state of the art. The source codes are publicly available at https://github.com/rain305f/TIDA. Yu Wang 0027, Zhun Zhong, Pengchong Qiao, Xuxin Cheng, Xiawu Zheng, Chang Liu 0030, Nicu Sebe, Rongrong Ji, Jie Chen 0001 |
NeurIPS | 1 |
| 2023 | Self-Supervised Masking for Unsupervised Anomaly Detection and LocalizationabstractRecently, anomaly detection and localization in multimedia data have received significant attention among the machine learning community. In real-world applications such as medical diagnosis and industrial defect detection, anomalies only present in a fraction of the images. To extend the reconstruction-based anomaly detection architecture to the localized anomalies, we propose a self-supervised learning approach throughrandom maskingand thenrestoring, namedSelf-SupervisedMasking(SSM) for unsupervised anomaly detection and localization. SSM not only enhances the training of the inpainting network but also leads to great improvement in the efficiency of mask prediction at inference. Through random masking, each image is augmented into a diverse set of training triplets, thus enabling the autoencoder to learn to reconstruct with masks of various sizes and shapes during training. To improve the efficiency and effectiveness of anomaly detection and localization at inference, we propose a novel progressive mask refinement approach that progressively uncovers the normal regions and finally locates the anomalous regions. The proposed SSM method outperforms several state-of-the-arts for both anomaly detection and anomaly localization, achieving 98.3% AUC on Retinal-OCT and 93.9% AUC on MVTec AD, respectively. Chaoqin Huang, Qinwei Xu, Yanfeng Wang 0001, Yu Wang 0027, Ya Zhang 0002 |
IEEE Trans. Multim. | 4 |
| 2022 | LAR-SR: A Local Autoregressive Model for Image Super-ResolutionabstractPrevious super-resolution (SR) approaches often formulate SR as a regression problem and pixel wise restoration, which leads to a blurry and unreal SR output. Recent works combine adversarial loss with pixel-wise loss to train a GAN-based model or introduce normalizing flows into SR problems to generate more realistic images. As another powerful generative approach, autoregressive (AR) model has not been noticed in low level tasks due to its limitation. Based on the fact that given the structural in-formation, the textural details in the natural images are locally related without long term dependency, in this paper we propose a novel autoregressive model-based SR approach, namely LAR-SR, which can efficiently generate realistic SR images using a novel local autoregressive (LAR) module. The proposed LAR module can sample all the patches of textural components in parallel, which greatly reduces the time consumption. In addition to high time efficiency, it is also able to leverage contextual information of pixels and can be optimized with a consistent loss. Experimental results on the widely-used datasets show that the proposed LAR-SR approach achieves superior performance on the vi-sual quality and quantitative metrics compared with other generative models such as GAN, Flow, and is competitive with the mixture generative model. Baisong Guo, Xiaoyun Zhang 0001, Haoning Wu 0002, Yu Wang 0027, Ya Zhang 0002, Yanfeng Wang 0001 |
CVPR | 4 |
| 2022 | Multi-level Fusion of Wav2vec 2.0 and BERT for Multimodal Emotion RecognitionabstractThe research and applications of multimodal emotion recognition have become increasingly popular recently.However, multimodal emotion recognition faces the challenge of lack of data.To solve this problem, we propose to use transfer learning which leverages state-of-the-art pre-trained models including wav2vec 2.0 and BERT for this task.Multi-level fusion approaches including coattention-based early fusion and late fusion with the models trained on both embeddings are explored.Also, a multi-granularity framework which extracts not only frame-level speech embeddings but also segment-level embeddings including phone, syllable and word-level speech embeddings is proposed to further boost the performance.By combining our coattention-based early fusion model and late fusion model with the multi-granularity feature extraction framework, we obtain result that outperforms best baseline approaches by 1.3% unweighted accuracy (UA) on the IEMOCAP dataset. Yanfeng Wang 0001, Yu Wang 0027 |
INTERSPEECH | 3 |
| 2021 | Efficient Use of End-to-End Data in Spoken Language ProcessingabstractFor many challenging tasks there is often limited data to train the systems in an end-to-end fashion, which has become increasingly popular for deep-learning. However, these tasks can normally be split into multiple separate modules, with significant quantities of data associated with each module. Spoken language processing applications fit into this scenario, as they usually start with a speech recognition module, followed by multiple task specific modules to achieve the end goal. This work examines how the best use can be made of limited end-to-end training for sequence-to-sequence tasks. The key to improving the use of the data is to more tightly integrate the modules via embeddings, rather than simply propagating words between modules. In this work speech translation is considered as the spoken language application. When significant quantities of in-domain, end-to-end data is available, cascade approaches operate well. When the in-domain data is limited, how-ever, tighter integration between modules enables better use of the data to be made. One of the challenges with tighter integration is how to ensure embedding consistency between the modules. A novel form of embedding-passing between modules is proposed that shows improved performance over both cascade and standard embedding-passing approaches for limited in-domain data. Yiting Lu, Yu Wang 0027, Mark J. F. Gales |
ICASSP | 2 |
| 2020 | Non-Native Children's Automatic Speech Recognition: The INTERSPEECH 2020 Shared Task ALTA SystemsabstractAutomatic spoken language assessment (SLA) is a challenging problem due to the large variations in learner speech combined with limited resources. These issues are even more problematic when considering children learning a language, with higher levels of acoustic and lexical variability, and of code-switching compared to adult data. This paper describes the ALTA system for the INTERSPEECH 2020 Shared Task on Automatic Speech Recognition for Non-Native Children’s Speech. The data for this task consists of examination recordings of Italian school children aged 9-16, ranging in ability from minimal, to basic, to limited but effective command of spoken English. A variety of systems were developed using the limited training data available, 49 hours. State-of-the-art acoustic models and language models were evaluated, including a diversity of lexical representations, handling code-switching and learner pronunciation errors, and grade specific models. The best single system achieved a word error rate (WER) of 16.9% on the evaluation data. By combining multiple diverse systems, including both grade independent and grade specific models, the error rate was reduced to 15.7%. This combined system was the best performing submission for both the closed and open tasks. Kate M. Knill, Yu Wang 0027, Xixin Wu, Mark J. F. Gales |
INTERSPEECH | 3 |
| 2020 | Spoken Language 'Grammatical Error Correction'abstractSpoken language ‘grammatical error correction’ (GEC) is an important mechanism to help learners of a foreign language, here English, improve their spoken grammar. GEC is challeng- ing for non-native spoken language due to interruptions from disfluent speech events such as repetitions and false starts and issues in strictly defining what is acceptable in spoken language. Furthermore there is little labelled data to train models. One way to mitigate the impact of speech events is to use a disflu- ency detection (DD) model. Removing the detected disfluencies converts the speech transcript to be closer to written language, which has significantly more labelled training data. This paper considers two types of approaches to leveraging DD models to boost spoken GEC performance. One is sequential, a separately trained DD model acts as a pre-processing module providing a more structured input to the GEC model. The second approach is to train DD and GEC models in an end-to-end fashion, simul- taneously optimising both modules. Embeddings enable end- to-end models to have a richer information flow. Experimen- tal results show that DD effectively regulates GEC input; end- to-end training works well when fine-tuned on limited labelled in-domain data; and improving DD by incorporating acoustic information helps improve spoken GEC. Yiting Lu, Mark J. F. Gales, Yu Wang 0027 |
INTERSPEECH | 3 |
| 2019 | Learning Between Different Teacher and Student Models in ASRabstractTeacher-student learning can be applied in automatic speech recognition for model compression and domain adaptation. This trains a student model to emulate the behaviour of a teacher model, and only the student is used to perform recognition. Depending on the application, the teacher and student may differ in their model types, complexities, input contexts, and input features. In previous works, it is often shown that learning from a strong teacher allows the student to perform better than an equivalent model trained with only the reference transcriptions. However, there has not been much investigation into whether a particular form of teacher is appropriate for the student to learn from. This paper aims to study how effectively the student is able to learn from the teacher, when differences exist between their designs. The Augmented Multi-party Interaction (AMI) meeting transcription and Multi-Genre Broadcast (MGB-3) television broadcast audio tasks are used in this analysis. Experimental results suggest that a student can effectively learn from a more complex teacher, but may struggle when it lacks input information. It is therefore important to carefully consider the design of the student for each application. Jeremy H. M. Wong, Mark J. F. Gales, Yu Wang 0027 |
ASRU | 3 |
| 2019 | Impact of ASR Performance on Spoken Grammatical Error DetectionabstractComputer assisted language learning (CALL) systems aidlearners to monitor their progress by providing scoring andfeedback on language assessment tasks. Free speaking tests al-low assessment of what a learner has said, as well as how theysaid it. For these tasks, Automatic Speech Recognition (ASR)is required to generate transcriptions of a candidate’s responses,the quality of these transcriptions is crucial to provide reliablefeedback in downstream processes. This paper considers theimpact of ASR performance on Grammatical Error Detection(GED) for free speaking tasks, as an example of providing feed-back on a learner’s use of English. The performance of an ad-vanced deep-learning based GED system, initially trained onwritten corpora, is used to evaluate the influence of ASR errors.One consequence of these errors is that grammatical errors canresult from incorrect transcriptions as well as learner errors, thismay yield confusing feedback. To mitigate the effect of theseerrors, and reduce erroneous feedback, ASR confidence scoresare incorporated into the GED system. By additionally adaptingthe written text GED system to the speech domain, using ASRtranscriptions, significant gains in performance can be achieved.Analysis of the GED performance for different grammatical er-ror types and across grade is also presented. Yiting Lu, Mark J. F. Gales, Kate M. Knill, P. P. Manakul, Yu Wang 0027 |
INTERSPEECH | 6 |
| 2019 | Exploiting Future Word Contexts in Neural Network Language Models for Speech RecognitionabstractLanguage modeling is a crucial component in a wide range of applications including speech recognition. Language models (LMs) are usually constructed by splitting a sentence into words and computing the probability of a word based on its word history. This sentence probability calculation, making use of conditional probability distributions, assumes that there is little impact from approximations used in the LMs, including the word history representations and finite training data. This motivates examining models that make use of additional information from the sentence. In this paper, future word information, in addition to the history, is used to predict the probability of the current word. For recurrent neural network LMs (RNNLMs), this information can be encapsulated in a bi-directional model. However, if used directly, this form of model is computationally expensive when trained on large quantities of data, and can be problematic when used with word lattices. This paper proposes a novel neural network language model structure, the succeeding-word RNNLM, su-RNNLM, to address these issues. Instead of using a recurrent unit to capture the complete future word contexts, a feedforward unit is used to model a fixed finite number of succeeding words. This is more efficient in training than bi-directional models and can be applied to lattice rescoring. The generated lattices can be used for downstream applications, such as confusion network decoding and keyword search. Experimental results on speech recognition and keyword spotting tasks illustrate the empirical usefulness of future word information, and the flexibility of the proposed model to represent this information. Xie Chen 0001, Xunying Liu, Yu Wang 0027, Anton Ragni, Jeremy H. M. Wong, Mark J. F. Gales |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | General Sequence Teacher-Student LearningabstractIn automatic speech recognition, performance gains can often be obtained by combining an ensemble of multiple models. However, this can be computationally expensive when performing recognition. Teacher-student learning alleviates this cost by training a single student model to emulate the combined ensemble behaviour. Only this student needs to be used for recognition. Previously investigated teacher-student criteria often limit the forms of diversity allowed in the ensemble, and only propagate information from the teachers to the student at the frame level. This paper addresses both of these issues by examining teacher-student learning within a sequence-level framework, and assessing the flexibility that these approaches offer. Various sequence-level teacher-student criteria are examined in this work, to propagate sequence posterior information. A training criterion based on the Kullback-Leibler (KL)-divergence between context-dependent state sequence posteriors is proposed that allows for a diversity of state cluster sets to be present in the ensemble. This criterion is shown to be an upper bound to a more general KL-divergence between word sequence posteriors, which places even fewer restrictions on the ensemble diversity, but whose gradient can be expensive to compute. These methods are evaluated on the augmented multi-party interaction (AMI) meeting transcription and MGB-3 television broadcast audio tasks. Jeremy H. M. Wong, Mark J. F. Gales, Yu Wang 0027 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Phonetic and Graphemic Systems for Multi-Genre Broadcast TranscriptionabstractState-of-the-art English automatic speech recognition systems typically use phonetic rather than graphemic lexicons. Graphemic systems are known to perform less well for English as the mapping from the written form to the spoken form is complicated. However, in recent years the representational power of deep-learning based acoustic models has improved, raising interest in graphemic acoustic models for English, due to the simplicity of generating the lexicon. In this paper, phonetic and graphemic models are compared for an English Multi-Genre Broadcast transcription task. A range of acoustic models based on lattice-free MMI training are constructed using phonetic and graphemic lexicons. For this task, it is found that having a long-span temporal history reduces the difference in performance between the two forms of models. In addition, system combination is examined, using parameter smoothing and hypothesis combination. As the combination approaches become more complicated the difference between the phonetic and graphemic systems further decreases. Finally, for all configurations examined the combination of phonetic and graphemic systems yields consistent gains. Yu Wang 0027, Xie Chen 0001, Mark J. F. Gales, Anton Ragni, Jeremy H. M. Wong |
ICASSP | 1 |
| 2018 | Impact of ASR Performance on Free Speaking Language AssessmentabstractIn free speaking tests candidates respond in spontaneous speech to prompts. This form of test allows the spoken language proficiency of a non-native speaker of English to be assessed more fully than read aloud tests. As the candidate's responses are unscripted, transcription by automatic speech recognition (ASR) is essential for automated assessment. ASR will never be 100% accurate so any assessment system must seek to minimise and mitigate ASR errors. This paper considers the impact of ASR errors on the performance of free speaking test auto-marking systems. Firstly rich linguistically related features, based on part-of-speech tags from statistical parse trees, are investigated for assessment. Then, the impact of ASR errors on how well the system can detect whether a learner's answer is relevant to the question asked is evaluated. Finally, the impact that these errors may have on the ability of the system to provide detailed feedback to the learner is analysed. In particular, pronunciation and grammatical errors are considered as these are important in helping a learner to make progress. As feedback resulting from an ASR error would be highly confusing, an approach to mitigate this problem using confidence scores is also analysed. Kate M. Knill, Mark J. F. Gales, Konstantinos Kyriakopoulos, Andrey Malinin, Anton Ragni, Yu Wang 0027, Andrew Caines |
INTERSPEECH | 6 |
| 2018 | Speaker Adaptation and Adaptive Training for Jointly Optimised Tandem SystemsabstractSpeaker independent (SI) Tandem systems trained by joint optimisation of bottleneck (BN) deep neural networks (DNNs) and Gaussian mixture models (GMMs) have been found to produce similar word error rates (WERs) to Hybrid DNN systems. A key advantage of using GMMs is that existing speaker adaptation methods, such as maximum likelihood linear regression (MLLR), can be used which to account for diverse speaker variations and improve system robustness. This paper investigates speaker adaptation and adaptive training (SAT) schemes for jointly optimised Tandem systems. Adaptation techniques investigated include constrained MLLR (CMLLR) transforms based on BN features for SAT as well as MLLR and parameterised sigmoid functions for unsupervised test-time adaptation. Experiments using English multi-genre broadcast (MGB3) data show that CMLLR SAT yields a 4% relative WER reduction over jointly trained Tandem and Hybrid SI systems, and further reductions in WER are obtained by system combination. Yu Wang 0027, Chao Zhang 0031, Mark J. F. Gales, Philip C. Woodland |
INTERSPEECH | 1 |
| 2018 | Sequence Teacher-Student Training of Acoustic Models for Automatic Free Speaking Language AssessmentabstractA high performance automatic speech recognition (ASR) system is an important constituent component of an automatic language assessment system for free speaking language tests. The ASR system is required to be capable of recognising non-native spontaneous English speech and to be deployable under real-time conditions. The performance of ASR systems can often be significantly improved by leveraging upon multiple systems that are complementary, such as an ensemble. Ensemble methods, however, can be computationally expensive, often requiring multiple decoding runs, which makes them impractical for deployment. In this paper, a lattice-free implementation of sequence-level teacher-student training is used to reduce this computational cost, thereby allowing for real-time applications. This method allows a single student model to emulate the performance of an ensemble of teachers, but without the need for multiple decoding runs. Adaptations of the student model to speakers from different first languages (L1s) and grades are also explored. Yu Wang 0027, Jeremy H. M. Wong, Mark J. F. Gales, Kate M. Knill, Anton Ragni |
SLT | 1 |
| 2018 | Towards automatic assessment of spontaneous spoken English
Yu Wang 0027, Mark J. F. Gales, Kate M. Knill, Konstantinos Kyriakopoulos, Andrey Malinin, Rogier C. van Dalen, M. Rashid |
Speech Commun. | 1 |
| 2018 | Model-Based Speech Enhancement in the Modulation DomainabstractThis paper presents an algorithm for modulation-domain speech enhancement using a Kalman filter. The proposed estimator jointly models the estimated dynamics of the spectral amplitudes of speech and noise to obtain an MMSE estimation of the speech amplitude spectrum with the assumption that the speech and noise are additive in the complex domain. In order to include the dynamics of noise amplitudes with those of speech amplitudes, we propose a statistical “Gaussring” model that comprises a mixture of Gaussians whose centers lie in a circle on the complex plane. The performance of the proposed algorithm is evaluated using the perceptual evaluation of speech quality measure, segmental SNR measure, and short-time objective intelligibility measure. For speech quality measures, the proposed algorithm is shown to give a consistent improvement over a wide range of SNRs when compared to competitive algorithms. Speech recognition experiments also show that the Gaussring-model-based algorithm performs well for two types of noise. Yu Wang 0027, Mike Brookes |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Use of Graphemic Lexicons for Spoken Language AssessmentabstractCopyright © 2017 ISCA. Automatic systems for practice and exams are essential to support the growing worldwide demand for learning English as an additional language. Assessment of spontaneous spoken English is, however, currently limited in scope due to the difficulty of achieving sufficient automatic speech recognition (ASR) accuracy. "Off-the-shelf" English ASR systems cannot model the exceptionally wide variety of accents, pronunications and recording conditions found in non-native learner data. Limited training data for different first languages (L1s), across all proficiency levels, often with (at most) crowd-sourced transcriptions, limits the performance of ASR systems trained on non-native English learner speech. This paper investigates whether the effect of one source of error in the system, lexical modelling, can be mitigated by using graphemic lexicons in place of phonetic lexicons based on native speaker pronunications. Graphemicbased English ASR is typically worse than phonetic-based due to the irregularity of English spelling-to-pronunciation but here lower word error rates are consistently observed with the graphemic ASR. The effect of using graphemes on automatic assessment is assessed on different grader feature sets: audio and fluency derived features, including some phonetic level features; and phone/grapheme distance features which capture a measure of pronunciation ability. Kate M. Knill, Mark J. F. Gales, Konstantinos Kyriakopoulos, Anton Ragni, Yu Wang 0027 |
INTERSPEECH | 5 |
| 2016 | Off-topic Response Detection for Spontaneous Spoken English AssessmentabstractAutomatic spoken language assessment systems are becoming increasingly important to meet the demand for English second language learning.This is a challenging task due to the high error rates of, even state-of-the-art, non-native speech recognition.Consequently current systems primarily assess fluency and pronunciation.However, content assessment is essential for full automation.As a first stage it is important to judge whether the speaker responds on topic to test questions designed to elicit spontaneous speech.Standard approaches to off-topic response detection assess similarity between the response and question based on bag-of-words representations.An alternative framework based on Recurrent Neural Network Language Models (RNNLM) is proposed in this paper.The RNNLM is adapted to the topic of each test question.It learns to associate example responses to questions with points in a topic space constructed using these example responses.Classification is done by ranking the topic-conditional posterior probabilities of a response.The RNNLMs associate a broad range of responses with each topic, incorporate sequence information and scale better with additional training data, unlike standard methods.On experiments conducted on data from the Business Language Testing Service (BULATS) this approach outperforms standard approaches. Andrey Malinin, Rogier C. van Dalen, Kate M. Knill, Yu Wang 0027, Mark J. F. Gales |
ACL (1) | 4 |
| 2016 | Speech enhancement using an MMSE spectral amplitude estimator based on a modulation domain Kalman filter with a Gamma priorabstractIn this paper, we propose a minimum mean square error spectral estimator for clean speech spectral amplitudes that uses a Kalman filter to model the temporal dynamics of the spectral amplitudes in the modulation domain. Using a two-parameter Gamma distribution to model the prior distribution of the speech spectral amplitudes, we derive closed form expressions for the posterior mean and variance of the spectral amplitudes as well as for the associated update step of the Kalman filter. The performance of the proposed algorithm is evaluated on the TIMIT core test set using the perceptual evaluation of speech quality (PESQ) measure and segmental SNR measure and is shown to give a consistent improvement over a wide range of SNRs when compared to competitive algorithms. Yu Wang 0027, Mike Brookes |
ICASSP | 1 |
| 2016 | A data-driven non-intrusive measure of speech quality and intelligibility
Dushyant Sharma, Yu Wang 0027, Patrick A. Naylor, Mike Brookes |
Speech Commun. | 2 |
| 2014 | Speech enhancement usinga modulation domain Kalman filter post-processor with a Gaussian Mixture noise modelabstractWe propose a speech enhancement algorithm that applies a Kalman filter in the modulation domain to the output of a conventional enhancer operating in the time-frequency domain. We show that the prediction residual signal of the spectral amplitude errors at the output of the baseline MMSE enhancer do not follow a Gaussian distribution. Accordingly, the Kalman filter used in our enhancement algorithm combines a colored noise model with a Gaussian mixture model of the residual noise. We evaluate the performance of the speech enhancement algorithm on the core TIMIT test set and demonstrate that it gives consistent performance improvements over the baseline enhancer and over a previously proposed Kalman filter post-processor. Yu Wang 0027, Mike Brookes |
ICASSP | 1 |
| 2013 | Speech enhancement using a robust Kalman filter post-processor in the modulation domainabstractWe propose a speech enhancement algorithm that applies a Kalman filter in the modulation domain to the output of a conventional enhancer operating in the time-frequency domain. The speech model required by the Kalman filter is obtained by performing linear predictive analysis in each frequency bin of the modulation domain signal. We show, however, that the corresponding speech synthesis filter can have a very high gain at low frequencies and may approach instability. To improve the stability of the synthesis filter, we propose two alternative methods of limiting its low frequency gain. We evaluate the performance of the speech enhancement algorithm on the core TIMIT test set and demonstrate that it gives consistent performance improvements over the baseline enhancer. Yu Wang 0027, Mike Brookes |
ICASSP | 1 |