Ji Wu 0002

dblp:91/4957-2 · DBLP profile ↗
← Back
100ranked-venue papers
7as first author
48since 2021 · last 2025
0000-0001-6170-726XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 61 · 5 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 55 · 6 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 13 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Enhancing Elusive Clues in Knowledge Learning by Contrasting Attention of Language Models
abstract
Causal language models acquire vast amount of knowledge from general text corpus during pretraining, but the efficiency of knowledge learning is known to be unsatisfactory, especially when learning from knowledge-dense and small-sized corpora. The deficiency can come from long-distance dependencies which are hard to capture by language models, and overfitting to co-occurrence patterns and distracting clues in the training text. To address these issues, the paper proposes a method to enhance knowledge learning during language model pretraining, by enhancing elusive but important clues in text discovered by the language model themselves. We found that larger language models pay more attention to non-obvious but important clues, which are often overlooked by smaller language models. Therefore, we can identify these clues by contrasting the attention weights of large and small language models. We use the identified clues as a guide to perform token-dropout data augmentation on the training text, and observed a significant boost in both small and large models' performance in fact memorization. This shows that the behavior contrast between more and less-performant language models contains important clues for knowledge learning, and it can be "amplified" for a straight-forward improvement in knowledge learning efficiency.
Xiao Zhang 0001, Miao Li 0003, Ji Wu 0002
AAAI4
2025 Med-2E3: A 2D-Enhanced 3D Medical Multimodal Large Language Model
abstract
3D medical image analysis is essential for modern healthcare, yet traditional task-specific models are inadequate due to limited generalizability across diverse clinical scenarios. Multimodal large language models (MLLMs) offer a promising solution to these challenges. However, existing MLLMs have limitations in fully leveraging the rich, hierarchical information embedded in 3D medical images. Inspired by clinical practice, where radiologists focus on both 3D spatial structure and 2D planar content, we propose Med-2E3, a 3D medical MLLM that integrates a dual 3D-2D encoder architecture. To aggregate 2D features effectively, we design a Text-Guided Inter-Slice (TG-IS) scoring module, which scores the attention of each 2D slice based on slice contents and task instructions. To the best of our knowledge, Med-2E3 is the first MLLM to integrate both 3D and 2D features for 3D medical image analysis. Experiments on large-scale, open-source 3D medical multimodal datasets demonstrate that TG- IS exhibits task-specific attention distribution and sig-nificantly outperforms current state-of-the-art models. The code is available at: https://github.com/MSIIPlMed-2E3
Yiming Shi, Chenyi Guo, Miao Li 0003, Ji Wu 0002
BIBM7
2025 3D-HSPA: Integrating 3D Spatial Information with Hierarchical Slice-Patch Attention for Knee MRI Analysis
abstract
Magnetic Resonance Imaging (MRI) is a crucial modality for diagnosing knee joint diseases. However, accurately extracting disease-relevant features from complex multi-slice, multi-sequence MRI scans remains a considerable challenge. To address this, we propose 3D-HSPA, a novel diagnostic framework for multi-slice, multi-sequence knee MRI, which integrates disease-specific information at both slice and patch levels and establishes intrinsic spatial connections among different sequences. Specifically, we introduce a patch-level and slice-level label attention mechanism, guiding the model to automatically learn a precise alignment between image regions and disease labels. Furthermore, by mapping 2D images from various sequences into a unified 3D spatial coordinate system, we enhance the spatial consistency and robustness of the attention distributions. We validated 3D-HSPA on a large-scale MRI dataset comprising 50 fine-grained types of knee joint diseases. The experimental results demonstrate that 3D-HSPA not only achieves superior diagnostic performance but also exhibits strong model interpretability.
Jingzhi Yang, Yiming Shi, Ji Wu 0002, Huishu Yuan, Miao Li 0003, Xiangling Fu
BIBM5
2025 LLM Sensitivity Evaluation Framework for Clinical Diagnosis
abstract
Large language models (LLMs) have demonstrated impressive performance across various domains. However, for clinical diagnosis, higher expectations are required for LLM’s reliability and sensitivity: thinking like physicians and remaining sensitive to key medical information that affects diagnostic reasoning, as subtle variations can lead to different diagnosis results. Yet, existing works focus mainly on investigating the sensitivity of LLMs to irrelevant context and overlook the importance of key information. In this paper, we investigate the sensitivity of LLMs, i.e. GPT-3.5, GPT-4, Gemini, Claude3 and LLaMA2-7b, to key medical information by introducing different perturbation strategies. The evaluation results highlight the limitations of current LLMs in remaining sensitive to key medical information for diagnostic decision-making. The evolution of LLMs must focus on improving their reliability, enhancing their ability to be sensitive to key information, and effectively utilizing this information. These improvements will enhance human trust in LLMs and facilitate their practical application in real-world scenarios. Our code and dataset are available at https://github.com/chenwei23333/DiagnosisQA.
Chenwei Yan, Xiangling Fu, Yuxuan Xiong, Siu Cheung Hui, Ji Wu 0002, Xien Liu
COLING6
2025 SSSL-HAR: Synthetic-Data-Driven Self-Supervised Learning for flexible IMU-Based Human Activity Recognition
abstract
The scarcity of labeled training data have significantly hindered the deployment of Inertial Measurement Unit-based Human Activity Recognition (IMU-HAR) in real-world scenarios. To address this limitation, we pro-pose SSSL-HAR (Synthetic-data-driven Self-Supervised-Learning HAR), a novel framework that leverages large amount synthetic IMU data for self-supervised pre-training, followed by fine-tuning with minimal real-world data. This approach mitigates the reliance on large-scale real IMU data collection compared to the traditional self-supervised-learning frameworks and also bypasses the need for labor-intensive annotation or cross-modality alignment compared to the traditional cross-modality methods. Furthermore, by adopting multi-view contrastive learning (MVCL) architectures, our method effectively captures the intrinsic relationships between synthetic sensor views, enabling robust generalization to diverse sensor placements and configurations. Experiments on the PAMAP2 dataset and a more complex custom fitness monitoring dataset demonstrate that SSSL-HAR achieves performance comparable to models pre-trained on real data, highlighting its potential for scalable and adaptive HAR deployment.
Timin Li, Zhuangzhuang Li, Ji Wu 0002, Yuepeng Chen, Xuefeng Feng, Chenyi Guo
IJCB4
2025 Reliable and Diverse Evaluation of LLM Medical Knowledge Mastery
abstract
Mastering medical knowledge is crucial for medical-specific LLMs. However, despite the existence of medical benchmarks like MedQA, a unified framework that fully leverages existing knowledge bases to evaluate LLMs' mastery of medical knowledge is still lacking. We propose PretexEval, a novel framework that dynamically generates reliable and diverse test samples to evaluate LLMs for any given medical knowledge base. We notice that test samples produced directly from knowledge bases by templates or LLMs may introduce factual errors and also lack diversity. To address these issues, our framework employs predicate equivalence transformations to produce a series of variants for any given medical knowledge point. Finally, these produced predicate variants are converted into textual language, resulting in a series of reliable and diverse test samples. Here, we use our proposed framework to systematically investigate the mastery of medical factual knowledge of 12 well-known LLMs, based on two knowledge bases that are crucial for clinical diagnosis and treatment. The evaluation results illustrate that current LLMs still exhibit significant deficiencies in fully mastering medical knowledge, despite achieving considerable success on some famous public benchmarks. These new findings provide valuable insights for developing medical-specific LLMs, highlighting that current LLMs urgently need to strengthen their comprehensive and in-depth mastery of medical knowledge before being applied to real-world medical scenarios.
Yuxuan Zhou 0002, Xien Liu, Chen Ning, Xiao Zhang 0001, Ji Wu 0002
ICLR5
2025 Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
abstract
Large language models (LLMs) have demonstrated remarkable performance on various medical benchmarks, but their capabilities across different cognitive levels remain underexplored. Inspired by Bloom's Taxonomy, we propose a multi-cognitive-level evaluation framework for assessing LLMs in the medical domain in this study. The framework integrates existing medical datasets and introduces tasks targeting three cognitive levels: preliminary knowledge grasp, comprehensive knowledge application, and scenario-based problem solving. Using this framework, we systematically evaluate state-of-the-art general and medical LLMs from six prominent families: Llama, Qwen, Gemma, Phi, GPT, and DeepSeek. Our findings reveal a significant performance decline as cognitive complexity increases across evaluated models, with model size playing a more critical role in performance at higher cognitive levels. Our study highlights the need to enhance LLMs' medical capabilities at higher cognitive levels and provides insights for developing LLMs suited to real-world medical applications.
Yuxuan Zhou 0002, Xien Liu, Chenwei Yan, Chen Ning, Xiao Zhang 0001, Boxun Li, Xiangling Fu, Shijin Wang 0001, Yu Wang 0002, Ji Wu 0002
ICML11
2025 Connector-S: A Survey of Connectors in Multi-modal Large Language Models
abstract
With the rapid advancements in multi-modal large language models (MLLMs), connectors play a pivotal role in bridging diverse modalities and enhancing model performance. However, the design and evolution of connectors have not been comprehensively analyzed, leaving gaps in understanding how these components function and hindering the development of more powerful connectors. In this survey, we systematically review the current progress of connectors in MLLMs and present a structured taxonomy that categorizes connectors into atomic operations (mapping, compression, mixture of experts) and holistic designs (multi-layer, multi-encoder, multi-modal scenarios), highlighting their technical contributions and advancements. Furthermore, we discuss several promising research frontiers and challenges, including high-resolution input, dynamic compression, guide information selection, combination strategy, and interpretability. This survey is intended to serve as a foundational reference and a clear roadmap for researchers, providing valuable insights into the design and optimization of next-generation connectors to enhance the performance and adaptability of MLLMs.
Xi Chen 0009, Yiming Shi, Miao Li 0003, Ji Wu 0002
IJCAI6
2025 SEmgFormer: Muscle Synergy Channel Attention Enhanced Vision Transformer Network For sEMG Motion Recognition Using STFT Spectrogram
abstract
The recognition of motions based on surface electromyography (sEMG) has been extensively studied, yielding promising results from initial machine learning approaches to contemporary deep learning methods. However, most previous research has concentrated on the classification of movements from individual body parts, such as the widely used gesture dataset, NinaPro. Furthermore, much of work has been restricted to convolutional neural networks (CNNs) and their variants, without a thorough exploration of the synergistic effects of muscles from different body parts, often assigning equal weights to all muscles. This study presents the collection of electromyographic data from sixteen major muscles across the entire body, acquiring the MultiMotion-sEMG Dataset, which includes forty-three full-body movements from thirteen participants. According to the current knowledge, this is the first dataset designed to synchronize the capture of full-body surface electromyography (sEMG) signals. Based on this dataset, a novel sEMG recognition network, SEmgFormer, is proposed, which is augmented by a vision transformer (ViT). The short-time Fourier transform (STFT) is utilized to transform conventional time-domain sEMG signal recognition tasks into visual understanding tasks of time-frequency spectrograms, utilizing the Cutmix method for data augmentation. In addition, a novel Muscle Synergy Channel Attention (MS-CA) mechanism is introduced, improving the channel attention mechanism (CA). The results indicate that the proposed model surpasses other methods in performance, including CNN-based networks, achieving optimal accuracy. This validates the efficacy of the proposed ViT classifier using time-frequency spectrograms as input, enhancing the accuracy of sEMG recognition based on full-body signals, and paving new avenues for research in this field.
Zhuangzhuang Li, Chenyi Guo, Ji Wu 0002
IJCNN4
2025 The 1st SpeechWellness Challenge: Detecting Suicide Risk Among Adolescents
abstract
The 1st SpeechWellness Challenge (SW1) aims to advance methods for detecting current suicide risk in adolescents using speech analysis techniques. Suicide among adolescents is a critical public health issue globally. Early detection of suicidal tendencies can lead to timely intervention and potentially save lives. Traditional methods of assessment often rely on self-reporting or clinical interviews, which may not always be accessible. The SW1 challenge addresses this gap by exploring speech as a non-invasive and readily available indicator of mental health. We release the SW1 dataset which contains speech recordings from 600 adolescents aged 10-18 years. By focusing on speech generated from natural tasks, the challenge seeks to uncover patterns and markers that correlate with current suicide risk.
Wen Wu 0007, Ziyun Cui, Chang Lei, Yinan Duan, Diyang Qu, Ji Wu 0002, Bowen Zhou 0001, Runsen Chen, Chao Zhang 0031
INTERSPEECH6
2025 Medical Contrastive Learning of Positive and Negative Mentions
WeiLong Wu, Jingzhi Yang, Xiao Zhang 0001, ZiYu Liu, Miao Li 0003, Ji Wu 0002
MICCAI (11)7
2025 DML-FitAR: A Deep Metric Learning Approach for IMU-Based Fitness Activity Recognition
abstract
This paper proposes DML-FitAR, a novel deep metric learning framework for IMU-based fitness action recognition, addressing critical challenges in real-world deployment. Unlike traditional transfer learning methods requiring fine-tuning for new action types, DML-FitAR achieves competitive accuracy on unseen actions through a retraining-free paradigm. Evaluated on a custom dataset (560+ fitness actions) and the MyoGYM dataset, DML-FitAR demonstrates superior performance over contrastive learning and visual backbone-based approaches, achieving cross-action-type recognition accuracy ranging from 80% to 90%. Besides that, the framework also exhibits robustness to sensor placement variations and noteworthy cross-dataset generalization.
Timin Li, Yuepeng Chen, Zhuangzhuang Li, Xuefeng Feng, Ji Wu 0002, Chenyi Guo
ICMR8
2025 MIPS: A Multimodal Infinite Polymer Sequence Pre-training Framework for Polymer Property Prediction
abstract
Polymers, composed of repeating structural units called monomers, are fundamental materials with a wide range of applications in daily life and industry. Accurate property prediction for polymers is essential for their design, development, and application. However, existing modeling approaches, which typically represent polymers by the constituent monomers, struggle to capture the whole properties of polymer, since the properties change during the polymerization process. In this study, we propose a Multimodal Infinite Polymer Sequence (MIPS) pre-training framework, which represents polymers as infinite sequences of monomers and integrates both topological and spatial information for comprehensive modeling. From the topological perspective, we generalize message passing mechanism (MPM) and graph attention mechanism (GAM) to infinite polymer sequences. For MPM, we demonstrate that applying MPM to infinite polymer sequences is equivalent to applying MPM on the induced star-linking graph of monomers. For GAM, we propose to further replace global graph attention with localized graph attention (LGA). Moreover, we show the robustness of the ''star linking'' strategy through an adversarial evaluation method named Repeat and Shift Invariance Test (RSIT). Despite its robustness, ''star linking'' strategy exhibits limitations when monomer side chains contain ring structures, a common characteristic of polymers, as it fails the Weisfeiler-Lehman (WL) test. To overcome this issue, we propose backbone embedding to enhance the capability of MPM and LGA on infinite polymer sequences. From the spatial perspective, we extract 3D descriptors of repeating monomers to capture spatial information. Finally, we design a cross-modal fusion mechanism to unify the topological and spatial information. Experimental validation across eight diverse polymer property prediction tasks reveals that MIPS achieves state-of-the-art performance. Ablation studies further comfirm the efficacy of our infinite polymer sequence modeling approach and multimodal pre-training framework.
Yaosen Min, Miao Li 0003, Ji Wu 0002
ACM Multimedia5
2025 Enhancing Multi-task Learning Capability of Medical Generalist Foundation Model via Image-centric Multi-annotation Data
abstract
The emergence of medical generalist foundation models has revolutionized conventional task-specific model development paradigms, aiming to better handle multiple tasks through joint training on large-scale medical datasets. However, recent advances prioritize simple data scaling or architectural component enhancement, while neglecting to re-examine multi-task learning from a data-centric perspective. Critically, simply aggregating existing data resources leads to decentralized image-task alignment, which fails to cultivate comprehensive image understanding or align with clinical needs for multi-dimensional image interpretation. In this paper, we introduce the image-centric multi-annotation X-ray dataset (IMAX), the first attempt to enhance the multi-task learning capabilities of medical multi-modal large language models (MLLMs) from the data construction level. To be specific, IMAX is featured from the following attributes: 1) High-quality data curation. A comprehensive collection of more than 354K entries applicable to seven different medical tasks. 2) Image-centric dense annotation. Each X-ray image is associated with an average of 4.10 tasks and 7.46 training entries, ensuring multi-task representation richness per image. Compared to the general decentralized multi-annotation X-ray dataset (DMAX), IMAX consistently demonstrates significant multi-task average performance gains ranging from 3.20% to 21.05% across seven open-source state-of-the-art medical MLLMs. Moreover, we investigate differences in statistical patterns exhibited by IMAX and DMAX training processes, exploring potential correlations between optimization dynamics and multi-task performance. Finally, leveraging the core concept of IMAX data construction, we propose an optimized DMAX-based training strategy to alleviate the dilemma of obtaining high-quality IMAX data in practical scenarios. Related resources are available at https://github.com/MSIIP/IMAX.
Fanbin Mo, Yiming Shi, Ming Wu 0001, Miao Li 0003, Ji Wu 0002
ACM Multimedia9
2025 FACT: Mitigating Inconsistent Hallucinations in LLMs via Fact-Driven Alternating Code-Text Training
abstract
Inconsistent hallucinations remain a major challenge for large language models (LLMs), undermining the accuracy and reliability of fact-based reasoning in real-world applications. Existing approaches often rely on task-specific training or adaptation, such as hand-crafted synthetic datasets for domain tasks or solutions mainly focused on numerical reasoning, thereby limiting generalizability to broader, unseen NLP tasks. Inspired by the structural rigor and logical consistency of programming languages, we observe that fact-based texts can be mapped to programming structures due to their inherent patterns. We further propose FACT, a novel Fact-driven Alternating Code-text Training framework that alternates between text-to-code and code-to-text prediction. FACT is the first task-agnostic paradigm that embeds code and natural language in a shared semantic space, thereby transferring the logical consistency of code to LLM outputs in NLP tasks. Experiments show that with only a small subset of Wiki-40B-en for training, FACT reduces inconsistent hallucinations by 2.7%–8.0% and improves overall performance by 2.5%–6.1% in three leading LLMs and four diverse datasets covering QA and summarization tasks. This framework offers a new perspective on addressing challenging hallucinations in LLMs, contributing to more reliable AI.
Xinxin You, Qixin Sun, Chenwei Yan, Xiao Zhang 0001, Chen Ning, Xiangling Fu, Shijin Wang 0001, Ji Wu 0002, Xien Liu
NeurIPS10
2025 Investigating and Mitigating Catastrophic Forgetting in Medical Knowledge Injection through Internal Knowledge Augmentation Learning
abstract
Large Language Models (LLMs) are expected to possess comprehensive medical knowledge to support real-world clinical applications. While domain-specific fine-tuning effectively injects medical knowledge into LLMs, it often causes catastrophic forgetting of previously acquired knowledge and instruction-following capabilities. In this paper, we investigate this issue and reveal a pattern of proximity-dependent forgetting: knowledge that is semantically or topically close to the injected content is more likely to be forgotten, while unrelated knowledge shows minimal degradation. Moreover, we observe that existing mitigation techniques fail to address this type of forgetting effectively. Motivated by this observation and inspired by human learning mechanisms, we proposeInternAL (\Internal Knowledge Augmentation Learning), a novel approach that leverages LLMs' own internal knowledge to mitigate forgetting. InternAL first probes internal knowledge closely related to the injection by prompting the model with questions derived from the injected knowledge. This knowledge is then used to augment the original injection dataset, guiding the model to retain related prior knowledge during training. Experimental results on multiple LLMs (LLaMA, Qwen) demonstrate that InternAL significantly mitigates proximity-related forgetting while maintaining strong knowledge injection performance. Our findings provide new insights into the nature of catastrophic forgetting in medical knowledge injection and highlight a promising direction for robust domain adaptation in LLMs. Code and datasets are available at https://github.com/THUMLP/InternAL.
Yuxuan Zhou 0002, Xien Liu, Xiao Zhang 0001, Chen Ning, Shijin Wang 0001, Ji Wu 0002
NeurIPS7
2024 M³AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset
abstract
Zhe Chen, Heyang Liu, Wenyi Yu, Guangzhi Sun, Hongcheng Liu, Ji Wu, Chao Zhang, Yu Wang, Yanfeng Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zhe Chen 0024, Heyang Liu, Wenyi Yu, Guangzhi Sun, Ji Wu 0002, Chao Zhang 0031, Yu Wang 0027, Yanfeng Wang 0001
ACL (1)6
2024 Data Augmentation Techniques for Chinese Disease Name Normalization
abstract
Disease name normalization is an important task in the medical domain. It classifies disease names written in various formats into standardized names, serving as a fundamental component in smart healthcare systems for various disease-related functions. Nevertheless, the most significant obstacle to existing disease name normalization systems is the severe shortage of training data. Consequently, we present a novel data augmentation approach that includes a series of data augmentation techniques and some supporting modules to help mitigate the problem. Through extensive experimentation, we illustrate that our proposed approach exhibits significant performance improvements across various baseline models and training objectives, particularly in scenarios with limited training data1.
Wenqian Cui, Xiangling Fu, Shaohui Liu, Mingjun Gu, Xien Liu, Ji Wu 0002, Irwin King
BIBM6
2024 Slice-Level Label Attention with Global-Guided Attention Regularization for Multi-Label Classification in Knee MRI Sequences
abstract
Magnetic Resonance Imaging (MRI) is crucial for diagnosing various knee-related diseases, and developing automatic diagnostic models based on knee MRI data is highly valuable. However, this task presents significant challenges due to the need to manage MRI data with multiple sequences and numerous images, where different diseases are often associated with specific images within certain sequences. To address these challenges, we propose a multi-label classification framework designed to effectively process MRI data and handle a large-scale label space encompassing hundreds of disease categories. Our approach introduces a Slice-Level Label Attention mechanism, which enables the model to learn the alignment between labels and images within sequences, thereby enhancing both performance and interpretability. Additionally, we present a Global-Guided Attention Regularization mechanism that further improves the consistency and robustness of the Slice-Level Label Attention results. We validate our framework on a large-scale MRI dataset involving multi-label classification across hundreds of fine-grained disease categories. Experimental results demonstrate that our method not only achieves superior performance but also provides more robust and consistent interpretability.
Jingzhi Yang, Weilong Wu, Ji Wu 0002, Huishu Yuan, Xiangling Fu, Miao Li 0003
IEEE Big Data5
2024 UniFS: Universal Few-Shot Instance Perception with Point Representations
Sheng Jin 0007, Ruijie Yao, Lumin Xu, Wentao Liu 0002, Chen Qian 0006, Ji Wu 0002, Ping Luo 0002
ECCV (29)6
2024 GKGNet: Group K-Nearest Neighbor Based Graph Convolutional Network for Multi-label Image Recognition
Ruijie Yao, Sheng Jin 0007, Lumin Xu, Wentao Liu 0002, Chen Qian 0006, Ping Luo 0002, Ji Wu 0002
ECCV (18)8
2024 3D Human Pose Estimation via Non-causal Retentive Networks
Kaili Zheng, Feixiang Lu, Yihao Lv, Liangjun Zhang, Chenyi Guo, Ji Wu 0002
ECCV (33)6
2024 Bayesian Example Selection Improves In-Context Learning for Speech, Text and Visual Modalities
abstract
Large language models (LLMs) can adapt to new tasks through in-context learning (ICL) based on a few examples presented in dialogue history without any model parameter update.Despite such convenience, the performance of ICL heavily depends on the quality of the incontext examples presented, which makes the in-context example selection approach a critical choice.This paper proposes a novel Bayesian in-Context example Selection method (ByCS) for ICL.Extending the inference probability conditioned on in-context examples based on Bayes' theorem, ByCS focuses on the inverse inference conditioned on test input.Following the assumption that accurate inverse inference probability (likelihood) will result in accurate inference probability (posterior), incontext examples are selected based on their inverse inference results.Diverse and extensive cross-tasking and cross-modality experiments are performed with speech, text, and image examples.Experimental results show the efficacy and robustness of our ByCS method on various models, tasks and modalities.
Siyin Wang, Chao-Han Huck Yang, Ji Wu 0002, Chao Zhang 0031
EMNLP3
2024 Can Whisper Perform Speech-Based In-Context Learning?
abstract
This paper investigates the in-context learning abilities of the Whisper automatic speech recognition (ASR) models released by OpenAI. A novel speech-based in-context learning (SICL) approach is proposed for test-time adaptation, which can reduce the word error rates (WERs) with only a small number of labelled speech samples without gradient descent. Language-level adaptation experiments using Chinese dialects showed that when applying SICL to isolated word ASR, consistent and considerable relative WER reductions can be achieved using Whisper models of any size on two dialects, which is on average 32.3%. A k-nearest-neighbours-based in-context example selection technique can be applied to further improve the efficiency of SICL, which can increase the average relative WER reduction to 36.4%. The findings are verified using speaker adaptation or continuous speech recognition tasks, and both achieved considerable relative WER reductions. Detailed quantitative analyses are also provided to shed light on SICL’s adaptability to phonological variances and dialect-specific lexical nuances.
Siyin Wang, Chao-Han Huck Yang, Ji Wu 0002, Chao Zhang 0031
ICASSP3
2024 Conditional Language Learning with Context
abstract
Language models can learn sophisticated language understanding skills from fitting raw text. They also unselectively learn useless corpus statistics and biases, especially during finetuning on domain-specific corpora. In this paper, we propose a simple modification to causal language modeling called conditional finetuning, which performs language modeling conditioned on a context. We show that a context can "explain away" certain corpus statistics and make the model avoid learning them. In this fashion, conditional finetuning achieves selective learning from a corpus, learning knowledge useful for downstream tasks while avoiding learning useless corpus statistics like topic biases. This selective learning effect leads to less forgetting and better stability-plasticity tradeoff in domain finetuning, potentially benefitting lifelong learning with language models.
Xiao Zhang 0001, Miao Li 0003, Ji Wu 0002
ICML3
2024 MultifacetEval: Multifaceted Evaluation to Probe LLMs in Mastering Medical Knowledge
Yuxuan Zhou 0002, Xien Liu, Chen Ning, Ji Wu 0002
IJCAI4
2024 Dual Dynamic Attention Network for Flexible Job Scheduling with Reinforcement Learning
abstract
The flexible job shop problem (FJSP) is a classic combinatorial optimization problem that is strongly NP-hard. Recent studies have utilized deep reinforcement learning (DRL) methods for scheduling operations in FJSP problems, achieving results comparable to accurate methods such as OR tools. However, there are still limitations in extracting global representations of machines and operations. This paper proposes the Dual Dynamic Attention Network (DDAN), which addresses these limitations. The proposed method utilizes an in-channel dynamic attention mechanism to capture the global representation of machines and operations. This allows for accurate and efficient representation of complex dependencies between operations and machines, providing effective support for subsequent dispatching model. Assessments using synthetic datasets as well as public benchmarks corroborate the proposed approach’s superiority over traditional priority dispatching rules (PDRs) and state-of-the-art DRL algorithms. In certain cases, it even surpasses deterministic algorithms. Additionally, this method demonstrates superior performance and stronger generalization capabilities compared to current state-of-the-art DRL methods on large-scale FJSP problems that have not been previously encountered.
Yuepeng Chen, Ji Wu 0002, Chenyi Guo
IJCNN3
2024 Spontaneous Speech-Based Suicide Risk Detection Using Whisper and Large Language Models
Ziyun Cui, Chang Lei, Wen Wu 0007, Yinan Duan, Diyang Qu, Ji Wu 0002, Runsen Chen, Chao Zhang 0031
INTERSPEECH6
2024 Co-occurrence is not Factual Association in Language Models
abstract
Pretrained language models can encode a large amount of knowledge and utilize it for various reasoning tasks, yet they can still struggle to learn novel factual knowledge effectively from finetuning on limited textual demonstrations. In this work, we show that the reason for this deficiency is that language models are biased to learn word co-occurrence statistics instead of true factual associations. We identify the differences between two forms of knowledge representation in language models: knowledge in the form of co-occurrence statistics is encoded in the middle layers of the transformer model and does not generalize well to reasoning scenarios beyond simple question answering, while true factual associations are encoded in the lower layers and can be freely utilized in various reasoning tasks. Based on these observations, we propose two strategies to improve the learning of factual associations in language models. We show that training on text with implicit rather than explicit factual associations can force the model to learn factual associations instead of co-occurrence statistics, significantly improving the generalization of newly learned knowledge. We also propose a simple training method to actively forget the learned co-occurrence statistics, which unblocks and enhances the learning of factual associations when training on plain narrative text. On both synthetic and real-world corpora, the two proposed strategies improve the generalization of the knowledge learned during finetuning to reasoning scenarios such as indirect and multi-hop question answering.
Xiao Zhang 0001, Miao Li 0003, Ji Wu 0002
NeurIPS3
2024 Uni-Med: A Unified Medical Generalist Foundation Model For Multi-Task Learning Via Connector-MoE
abstract
Multi-modal large language models (MLLMs) have shown impressive capabilities as a general-purpose interface for various visual and linguistic tasks. However, building a unified MLLM for multi-task learning in the medical field remains a thorny challenge. To mitigate the tug-of-war problem of multi-modal multi-task optimization in MLLMs, recent advances primarily focus on improving the LLM components, while neglecting the connector that bridges the gap between modalities. In this paper, we introduce Uni-Med, a novel medical generalist foundation model which consists of a universal visual feature extraction module, a connector mixture-of-experts (CMoE) module, and an LLM. Benefiting from the proposed CMoE that leverages a well-designed router with a mixture of projection experts at the connector, Uni-Med achieves efficient solution to the tug-of-war problem and can perform six different medical tasks including question answering, visual question answering, report generation, referring expression comprehension, referring expression generation and image classification. To the best of our knowledge, Uni-Med is the first effort to tackle multi-task interference at the connector in MLLMs. Extensive ablation experiments validate the effectiveness of introducing CMoE under any configuration, with up to an average 8% performance gains. We further provide interpretation analysis of the tug-of-war problem from the perspective of gradient optimization and parameter statistics. Compared to previous state-of-the-art medical MLLMs, Uni-Med achieves competitive or superior evaluation metrics on diverse tasks. Code and resources are available at https://github.com/MSIIP/Uni-Med.
Fanbin Mo, Miao Li 0003, Ji Wu 0002
NeurIPS5
2024 A survey of label-noise deep learning for medical image analysis
Jialin Shi, Kailai Zhang, Chenyi Guo, Youquan Yang, Yali Xu, Ji Wu 0002
Medical Image Anal.6
2024 Graph-Based Cross-Granularity Message Passing on Knowledge-Intensive Text
abstract
In knowledge-intensive fields such as medicine, the text often contains numerous professional terms, specific text fragments, and multidimensional information. However, most existing text representation methods ignore this specialized knowledge and instead adopt methods similar to those used in the general domain. In this paper, we focus on developing a learning module to enhance the representation ability of knowledge-intensive text by leveraging a graph-based cross-granularity message passing mechanism. To this end, we propose a novel learning framework, theMulti-GranularityGraphNeuralNetwork (MG-GNN), to integrate fine-grained and coarse-grained knowledge at the character, word, and phase levels. The MG-GNN performs learning in two stages: 1) inter-granularity learning and 2) intra-granularity learning. During inter-granularity learning, semantic knowledge is extracted from character, word, and phrase granularity graphs, whereas intra-granularity learning focuses on fusing knowledge across different granularity graphs to achieve comprehensive message integration. To enhance the fusion performance, we propose a context-based gating mechanism to guide cross-graph propagation learning. Furthermore, we apply MG-GNN to address two important medical applications. Experimental results demonstrate that our proposed MG-GNN model significantly enhances the performance in both diagnosis prediction and medical named entity recognition tasks.
Chenwei Yan, Xiangling Fu, Xinxin You, Ji Wu 0002, Xien Liu
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Transferring Speech-Generic and Depression-Specific Knowledge for Alzheimer's Disease Detection
abstract
The detection of Alzheimer’s disease (AD) from spontaneous speech has attracted increasing attention while the sparsity of training data remains an important issue. This paper handles the issue by knowledge transfer, specifically from both speech-generic and depression-specific knowledge. The paper first studies sequential knowledge transfer from generic foundation models pretrained on large amounts of speech and text data. A block-wise analysis is performed for AD diagnosis based on the representations extracted from different intermediate blocks of different foundation models. Apart from the knowledge from speech-generic representations, this paper also proposes to simultaneously transfer the knowledge from a speech depression detection task based on the high comorbidity rates of depression and AD. A parallel knowledge transfer framework is studied that jointly learns the information shared between these two tasks. Experimental results show that the proposed method improves AD and depression detection, and produces a state-of-the-art F1 score of 0.928 for AD diagnosis on the commonly used ADReSSo dataset.
Ziyun Cui, Wen Wu 0007, Weiqiang Zhang 0001, Ji Wu 0002, Chao Zhang 0031
ASRU4
2023 Learning to Generate Radiology Findings from Impressions Based on Large Language Model
abstract
Medical imaging plays a pivotal role in clinical diagnosis, and the textual reports associated with these images are of paramount importance in aiding image comprehension and supporting treatment decisions. Automated report generation serves to alleviate the burden on radiologists and has garnered significant attention in the field of medical artificial intelligence. Previous research in text-based report generation primarily focused on generating impressions statements from radiology findings. However, the benefits in terms of reducing the workload on radiologists were not particularly evident. In this article, we propose a novel task of generating findings from radiology impressions. Leveraging advanced large language models, we trained a set of report generation models using a real dataset of knee MRI reports. Additionally, we incorporated various strategies, including data augmentation and efficient parameter fine-tuning. Objective experiments affirm the effectiveness of the methods we introduced. Furthermore, we conducted subjective assessments by radiologists, and the results demonstrate that our trained large language models significantly outperform professional radiologists in terms of overall report quality and content consistency.
Weilong Wu, Miao Li 0003, Ji Wu 0002, Huishu Yuan
IEEE Big Data3
2023 Knowledge Distillation Approach for Efficient Internal Language Model Estimation
Haihua Xu 0001, Yerbolat Khassanov, Lu Lu 0015, Zejun Ma 0001, Ji Wu 0002
INTERSPEECH7
2023 Obstructive Sleep Apnea Detection using Pre-trained Speech Representations
Kaibo Zhang, Lili Cao, Yanru Li, Chao Zhang 0031, Ji Wu 0002, Demin Han
INTERSPEECH6
2022 Understanding the Failure of Batch Normalization for Transformers in NLP
abstract
Batch Normalization (BN) is a core and prevalent technique in accelerating the training of deep neural networks and improving the generalization on Computer Vision (CV) tasks. However, it fails to defend its position in Natural Language Processing (NLP), which is dominated by Layer Normalization (LN). In this paper, we are trying to answer why BN usually performs worse than LN in NLP tasks with Transformer models. We find that the inconsistency between training and inference of BN is the leading cause that results in the failure of BN in NLP. We define Training Inference Discrepancy (TID) to quantitatively measure this inconsistency and reveal that TID can indicate BN's performance, supported by extensive experiments, including image classification, neural machine translation, language modeling, sequence labeling, and text classification tasks. We find that BN can obtain much better test performance than LN when TID keeps small through training. To suppress the explosion of TID, we propose Regularized BN (RBN) that adds a simple regularization term to narrow the gap between batch statistics and population statistics of BN. RBN improves the performance of BN consistently and outperforms or is on par with LN on 17 out of 20 settings, including ten datasets and two common variants of Transformer.
Ji Wu 0002, Lei Huang 0015
NeurIPS2
2022 Meta joint optimization: a holistic framework for noisy-labeled visual recognition
Jialin Shi, Ji Wu 0002
Appl. Intell.3
2022 MPF-net: An effective framework for automated cobb angle estimation
Kailai Zhang, Nanfang Xu, Chenyi Guo, Ji Wu 0002
Medical Image Anal.4
2021 SMP-Graph: Structure-Enhanced Unsupervised Semantic Graph Representation for Precise Medical Procedure Coding on EMRs
abstract
Automatic ICD coding, as a fundamental task in the field of healthcare management, has been paid much attention by researchers. However, the current deep learning-based ICD coding research mostly focus on the introduction of external diagnostic description text or the imposition of rules, while ignoring the structured features of the coding text itself. Especially for short texts of medical procedure codes, it is much more important to mine the information value in the texts. In this paper, we propose a structure-enhanced unsupervised semantic graph representation for precise medical procedure coding (SMP-Graph). The SMP-Graph representation method constructs each medical procedure text with an inductive heterogeneous graph and particularly enhances the kernel knowledge by extracting the axis words and chapter title and allocating them with distinct node representations. Both the nodes and edges are generated by the unsupervised pretrained model and then interact information in a bidirectionally weighted graph structure. Therefore, the SMP-Graph really realizes the intra-integration of unsupervised contextualized information and graph-based global information from the medical procedure code. Experiments conducted on the Chinese ICD-9-CM-3 procedure text dataset we collected from EMRs demonstrate that the SMP-Graph is a better representation method that outperforms other representative methods for medical procedure coding. Characteristic analysis is also conducted to prove the interpretability and adaptability of the SMP-Graph on the medical procedure coding task.
Yue Gao 0012, Xiangling Fu, Xien Liu, Kaiyin Zhou, Ji Wu 0002
BIBM5
2021 MolCloze: A Unified Cloze-style Self-supervised Molecular Structure Learning Model for Chemical Property Prediction
abstract
Machine Learning approaches are required to predict accurately on test samples that are distributionally different from training ones in the fields of drug discovery, computational biology, and cheminformatics. However, (i) labeled task-specific molecule data are often scarce, and (ii) poor generalization due to test molecules that are structurally different from those seen during training. To alleviate the problems, we propose a cloze-style self-supervised learning model (MolCloze) to obtain universal informative representations for molecular property prediction tasks. With carefully designed self-supervised tasks unifying generative- and discriminative-paradigm, MolCloze can learn rich structural and semantic information of molecules from enormous unlabelled molecular data. To capture such complex information, we design two novel strategies - Structural Fingerprint Tokenization (SFT) for better tokenizing molecule graphs, and Normalized Graph Raw Shortcut-connection (NGRS) for better latent representations by training a deeper model. We pretrain the MolCloze model via three tasks, which are Unordered Masked Language Modeling (UMLM), Replaced Masked Token Detection (RMTD), and Contrastive Energy-based Unmasked Token Clozing (CE-UTC). Then, we transfer the pre-trained model to a broad range of downstream molecular property prediction tasks via minor architecture modification. Extensive experiments demonstrate the generalizability of MolCloze by predicting a broad range of chemical properties which are related to drug discovery. We also observe significant performance boost on different downstream molecular property prediction datasets, achieving higher performance than the state-of-the-art baseline approaches and previous pre-training techniques developed for molecule data.
Yingheng Wang, Yaosen Min, Ji Wu 0002
BIBM4
2021 Molecular Graph Contrastive Learning with Parameterized Explainable Augmentations
abstract
Learning generalizable, transferable, and robust representations for molecule data has always been a challenge. The recent success of contrastive learning (CL) for self-supervised graph representation learning provides a novel perspective to learn molecule representations. However, existing graph CL frameworks usually adopt stochastic augmentations or schemes according to pre-defined rules ont he input graph to obtain different graph views in various scales, which may destroy topological semantemes and domain prior in molecule data, leading to suboptimal performance. Therefore, a well-designed parameterized augmentation scheme that preserves chemically meaningful structural information and intrinsically essential attributes is crucial for molecular graph contrastive learning, helping to learn representations that are insensitive to perturbation on unimportant atoms and bonds. In this paper, we propose a novel method, Molecular Graph Contrastive Learning with Parameterized Explainable Augmentations, that adaptively incorporates chemically significative information from both topological and semantic aspects of molecular graphs. Specifically, we apply deep neural networks to parameterize the augmentation process for both the molecular graph topology and atom attributes, to highlight contributive molecular substructures and recognize underlying chemical semantemes. Comprehensive experiments demonstrate that our method consistently outperforms compared baselines, verifying the effectiveness of the proposed framework. Our self-supervised model only uses one percent of the parameters to achieve comparative results against the state-of-the-art baseline, which has hundreds of millions of parameters. We also provide detailed case studies to validate the explainability of augmented views.
Yingheng Wang, Yaosen Min, Erzhuo Shao, Ji Wu 0002
BIBM4
2021 Multi-features-Based Automatic Clinical Coding for Chinese ICD-9-CM-3
Yue Gao 0012, Xiangling Fu, Xien Liu, Ji Wu 0002
ICANN (5)4
2021 Image Periodization for Convolutional Neural Networks
Kailai Zhang, Ji Wu 0002
ICONIP (1)3
2021 Training Graph Convolutional Neural Network Against Label Noise
Yuxin Zhuo, Xuesi Zhou, Ji Wu 0002
ICONIP (3)3
2021 Distilling Effective Supervision for Robust Medical Image Segmentation with Noisy Labels
Jialin Shi, Ji Wu 0002
MICCAI (1)2
2021 Multi-view Graph Contrastive Representation Learning for Drug-Drug Interaction Prediction
abstract
Potential Drug-Drug Interactions (DDI) occur while treating complex or co-existing diseases with drug combinations, which may cause changes in drugs’ pharmacological activity. Therefore, DDI prediction has been an important task in the medical health machine learning community. Graph-based learning methods have recently aroused widespread interest and are proved to be a priority for this task. However, these methods are often limited to exploiting the inter-view drug molecular structure and ignoring the drug’s intra-view interaction relationship, vital to capturing the complex DDI patterns. This study presents a new method, multi-view graph contrastive representation learning for drug-drug interaction prediction, MIRACLE for brevity, to capture inter-view molecule structure and intra-view interactions between molecules simultaneously. MIRACLE treats a DDI network as a multi-view graph where each node in the interaction graph itself is a drug molecular graph instance. We use GCN to encode DDI relationships and a bond-aware attentive message propagating method to capture drug molecular structure information in the MIRACLE learning stage. Also, we propose a novel unsupervised contrastive learning component to balance and integrate the multi-view information. Comprehensive experiments on multiple real datasets show that MIRACLE outperforms the state-of-the-art DDI prediction models consistently.
Yingheng Wang, Yaosen Min, Ji Wu 0002
WWW4
2021 An automated estimator for Cobb angle measurement using multi-task networks
Xiangling Fu, Guosheng Yang, Kailai Zhang, Nanfang Xu, Ji Wu 0002
Neural Comput. Appl.5
2020 Tensor Graph Convolutional Networks for Text Classification
abstract
Compared to sequential learning models, graph-based neural networks exhibit some excellent properties, such as ability capturing global information. In this paper, we investigate graph-based neural networks for text classification problem. A new framework TensorGCN (tensor graph convolutional networks), is presented for this task. A text graph tensor is firstly constructed to describe semantic, syntactic, and sequential contextual information. Then, two kinds of propagation learning perform on the text graph tensor. The first is intra-graph propagation used for aggregating information from neighborhood nodes in a single graph. The second is inter-graph propagation used for harmonizing heterogeneous information between graphs. Extensive experiments are conducted on benchmark datasets, and the results illustrate the effectiveness of our proposed framework. Our proposed TensorGCN presents an effective way to harmonize and integrate heterogeneous information from different kinds of graphs.
Xien Liu, Xinxin You, Xiao Zhang 0001, Ji Wu 0002, Ping Lv
AAAI4
2020 Learning Conceptual-Contextual Embeddings for Medical Text
abstract
External knowledge is often useful for natural language understanding tasks. We introduce a contextual text representation model called Conceptual-Contextual (CC) embeddings, which incorporates structured knowledge into text representations. Unlike entity embedding methods, our approach encodes a knowledge graph into a context model. CC embeddings can be easily reused for a wide range of tasks in a similar fashion to pre-trained language models. Our model effectively encodes the huge UMLS database by leveraging semantic generalizability. Experiments on electronic health records (EHRs) and medical text processing benchmarks showed our model gives a major boost to the performance of supervised medical NLP tasks.
Xiao Zhang 0001, Dejing Dou, Ji Wu 0002
AAAI3
2020 Reversal No Longer Matters: Attention-Based Arrhythmia Detection with Lead-Reversal ECG Data
abstract
In this paper, we propose an attention-based multi-scale neural network for arrhythmia detection with lead-reversal electrocardiogram data. Electrocardiogram with a set of 12 waveforms(known as 12-lead ECG) measures myocardial electro-physiological activity, which is important in clinical diagnosis of arrhythmia. However, lead reversals caused by electrode interchange may cause great interference to the interpretation, leading to significant accuracy decline of automated algorithms and possible faulty diagnosis by cardiologists. To address this problem, we design a multi-scale neural network using attention mechanism to reduce the influence of lead reversals. The proposed model is evaluated on a dataset which consists of ECG data from 3,658 patients. In experiments, the proposed method shows high performance on both lead-reversal data and normal data, which proves great robustness of the method.
Jialin Shi, Ji Wu 0002
ICASSP3
2020 HKA: A Hierarchical Knowledge Attention Mechanism for Multi-Turn Dialogue System
abstract
Generating informative responses by incorporating external knowledge into dialogue system attracts more and more attention. Most previous works facilitate single-turn dialogue system on generating such responses. However, few works focus on incorporating knowledge for multi-turn system, since the hierarchy of knowledge, from the words and utterances in context, is ignored. Motivated by this, we propose a novel hierarchical knowledge attention (HKA) mechanism for open-domain multi-turn dialogue system in this paper, which utilizes both word and utterance level attention jointly. Experiments demonstrate that the proposed HKA can incorporate more appropriate knowledge and make the state-of-the-art models generate more informative responses. Further analysis shows that our HKA can improve the model's ability of dialogue state management, especially when the number of dialogue turns is large.
Kailai Zhang, Xuesi Zhou, Ji Wu 0002
ICASSP4
2020 FPB: Improving Multi-Scale Feature Representation Inside Convolutional Layer Via Feature Pyramid Block
abstract
Multi-scale features exist widely in biomedical images. For example, the scale of lesions may vary greatly according to different diseases. Effective representation of multi-scale features is essential for fully perceiving and understanding objects, which guarantees the performance of models. However, in biomedical image tasks, the insufficiency of data may prevent models from effectively capturing multi-scale features. In this paper, we propose Feature Pyramid Block (FPB), a novel structure to improve multi-scale feature representation within a single convolutional layer, which can be easily plugged into existing convolutional networks. Experiments on public biomedical image datasets prove consistent performance improvement with FPB. Furthermore, the convergence speed is faster and the computational costs are lower when using FPB, which proves high efficiency of our method.
Kailai Zhang, Ji Wu 0002
ICIP3
2020 Circular Shift: An Effective Data Augmentation Method For Convolutional Neural Network On Image Classification
abstract
In this paper, we present a novel and effective data augmentation method for convolutional neural network(CNN) on image classification tasks. CNN-based models such as VGG, Resnet and Densenet have achieved great success on image classification tasks. The common data augmentation methods such as rotation, crop and flip are always used for CNN, especially under the lack of data. However, in some cases such as small images and dispersed feature of objects, these methods have limitations and even can decrease the classification performance. In this case, an operation that has lower risk is important for the performance improvement. Addressing this problem, we design a data augmentation method named circular shift, which provides variations for the CNN-based models but does not lose too much information. Three commonly used image datasets are chosen for the evaluation of our proposed operation, and the experiment results show consistent improvement on different CNN-based models. What is more, our operation can be added to the current set of augmentation operation and achieves further performance improvement.
Kailai Zhang, Ji Wu 0002
ICIP3
2020 Multi-modal Feature Attention for Cervical Lymph Node Segmentation in Ultrasound and Doppler Images
Xiangling Fu, Mengke Zhang, Chenyi Guo, Ji Wu 0002
ICONIP (4)6
2020 A Landmark Estimation and Correction Network for Automated Measurement of Sagittal Spinal Parameters
Guosheng Yang, Xiangling Fu, Nanfang Xu, Kailai Zhang, Ji Wu 0002
ICONIP (4)5
2020 An automatic multi-camera-based event extraction system for real soccer videos
Kailai Zhang, Ji Wu 0002, Xiaofeng Tong
Pattern Anal. Appl.2
2019 Exploiting Sentence Embedding for Medical Question Answering
abstract
Despite the great success of word embedding, sentence embedding remains a not-well-solved problem. In this paper, we present a supervised learning framework to exploit sentence embedding for the medical question answering task. The learning framework consists of two main parts: 1) a sentence embedding producing module, and 2) a scoring module. The former is developed with contextual self-attention and multi-scale techniques to encode a sentence into an embedding tensor. This module is shortly called Contextual self-Attention Multi-scale Sentence Embedding (CAMSE). The latter employs two scoring strategies: Semantic Matching Scoring (SMS) and Semantic Association Scoring (SAS). SMS measures similarity while SAS captures association between sentence pairs: a medical question concatenated with a candidate choice, and a piece of corresponding supportive evidence. The proposed framework is examined by two Medical Question Answering(MedicalQA) datasets which are collected from real-world applications: medical exam and clinical diagnosis based on electronic medical records (EMR). The comparison results show that our proposed framework achieved significant improvements compared to competitive baseline approaches. Additionally, a series of controlled experiments are also conducted to illustrate that the multi-scale strategy and the contextual self-attention layer play important roles for producing effective sentence embedding, and the two kinds of scoring strategies are highly complementary to each other for question answering problems.
Xien Liu, Ji Wu 0002, Ping Lv
AAAI3
2019 Delta Embedding Learning
abstract
Unsupervised word embeddings have become a popular approach of word representation in NLP tasks.However there are limitations to the semantics represented by unsupervised embeddings, and inadequate fine-tuning of embeddings can lead to suboptimal performance.We propose a novel learning technique called Delta Embedding Learning, which can be applied to general NLP tasks to improve performance by optimized tuning of the word embeddings.A structured regularization is applied to the embeddings to ensure they are tuned in an incremental way.As a result, the tuned word embeddings become better word representations by absorbing semantic information from supervision without "forgetting."We apply the method to various NLP tasks and see a consistent improvement in performance.Evaluation also confirms the tuned word embeddings have better semantic properties.
Xiao Zhang 0001, Ji Wu 0002, Dejing Dou
ACL (1)2
2019 Drug-drug Interaction Prediction with Graph Representation Learning
abstract
Pharmacological activity of one drug may be altered due to the concomitant administration of another drug, leading to unanticipated drug-drug interactions(DDIs). However, existing DDI prediction approaches are lacking in the following aspects: (1)scalability: they rely heavily on diverse drug-related features, leading to the unavailability of important features for most of the drugs when it comes to large-scale datasets. (2)robustness: they aim to approximate the interaction probability with the integration of diverse features. The model may be sensitive to pairwise similarity information of the test set. In this paper, we explore the promising application of graph representation learning for more accurate DDI prediction, establishing a brand new model to solve the two problems, achieving greater performance and keeping certain interpretability. Our experiments on the small-scale DDI dataset as well as the large-scale one illustrate that our model can achieve higher performance compared to various existing state-of-the-art approaches, which can indicate the scalability of our model. Moreover, our model can find the most important local atoms with the attention mechanism, which conform to domain knowledge with certain interpretability. Furthermore, the robust analysis show that the proposed method is insensitive to the pairwise similarity information of test datasets, and can retrieve interacting drug pairs even though their pairwise similarities are extremely low with a high recall rate and a considerable precision rate.
Xien Liu, Ji Wu 0002
BIBM3
2019 U-Module: Better Parameters Initialization of Convolutional Neural Network for Medical Image Classification
abstract
In this paper, we present a novel U-module for better parameters initialization of convolutional neural network(CNN) on medical image classification tasks. CNN-based models such as VGG and Resnet have achieved great success on natural image classification tasks. However, these methods need a large number of training data to get good performance. In medical image field, the performance of the CNN-based models is always limited by the lack of annotated medical images. In this case, the parameters initialization plays an important role for the model performance. Addressing this problem, we design an up-sample structure with unsupervised loss function named U-module, which can be easily inserted into different CNN-based models. In addition, we propose a special training method for the CNN with our U-modules. Two challenging datasets of medical images are chosen for the evaluation, and the experiment results show consistent improvement on different models. Particularly, our method is even better than using the pre-train parameters on large natural image dataset.
Kailai Zhang, Xuesi Zhou, Ji Wu 0002
ICIP3
2019 An Automated Cobb Angle Estimation Method Using Convolutional Neural Network with Area Limitation
Kailai Zhang, Nanfang Xu, Guosheng Yang, Ji Wu 0002, Xiangling Fu
MICCAI (6)4
2019 Pick-and-Learn: Automatic Quality Evaluation for Noisy-Labeled Image Segmentation
Haidong Zhu, Jialin Shi, Ji Wu 0002
MICCAI (6)3
2019 Constrained Learned Feature Extraction for Acoustic Scene Classification
abstract
Deep neural networks (DNNs) have been proven to be powerful models for acoustic scene classification tasks. State-of-the-art DNNs have millions of connections and are computationally intensive, making them difficult to deploy on systems with limited resources. With a focus on acoustic scene classification, we describe a new learnable module, the simulated Fourier transform module, which allows deep neural networks to implement the discrete Fourier transform operation 8x faster on a graphics processing unit (GPU). We frame the signal processing procedure as an adaptive machine learning problem and introduce learnable parameters in the module to facilitate fast adaptation for the complex and variable acoustic signal. This module gives neural networks the ability to model audio signals from raw waveforms, without extra fast Fourier transform and filter bank patches. Then, we use the temporal transformer module, which has been previously published, to alleviate the information loss caused by the simulated Fourier transform module. These techniques can be integrated into an existing fully connected neural network (FCNN), convolutional neural network (CNN), or recurrent neural network (RNN) models. We evaluate the proposed strategy using four acoustic scene datasets (LITIS Rouen, DCASE2016, DCASE2017, and DCASE2018) as target tasks. We show that the proposed approach significantly outperforms the vanilla FCNN, CNN, and RNN approach on both efficiency and performance. For instance, the proposed approach can reduce inference time by 8× while reducing the classification error on LITIS Rouen dataset from 3.21% to 1.81%.
Ji Wu 0002
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Medical Exam Question Answering with Large-scale Reading Comprehension
abstract
Reading and understanding text is one important component in computer aided diagnosis in clinical medicine, also being a major research problem in the field of NLP. In this work, we introduce a question-answering task called MedQA to study answering questions in clinical medicine using knowledge in a large-scale document collection. The aim of MedQA is to answer real-world questions with large-scale reading comprehension. We propose our solution SeaReader---a modular end-to-end reading comprehension model based on LSTM networks and dual-path attention architecture. The novel dual-path attention models information flow from two perspectives and has the ability to simultaneously read individual documents and integrate information across multiple documents. In experiments our SeaReader achieved a large increase in accuracy on MedQA over competing models. Additionally, we develop a series of novel techniques to demonstrate the interpretation of the question answering process in SeaReader.
Xiao Zhang 0001, Ji Wu 0002, Zhiyang He, Xien Liu
AAAI2
2018 Data Independent Sequence Augmentation Method for Acoustic Scene Classification
Kailai Zhang, Ji Wu 0002
INTERSPEECH3
2018 Temporal Transformer Networks for Acoustic Scene Classification
Kailai Zhang, Ji Wu 0002
INTERSPEECH3
2018 Multi-modal Attention Mechanisms in LSTM and Its Application to Acoustic Scene Classification
Kailai Zhang, Ji Wu 0002
INTERSPEECH3
2017 A Rescoring Approach for Keyword Search Using Lattice Context Information
Ji Wu 0002
INTERSPEECH2
2017 Multi-label text classification based on the label correlation mixture model
abstract
In the current paper, we propose a probabilistic generative model, the label correlation mixture model (LCMM), to depict multi-labeled document data, which can be utilized for multi-label text classification. LCMM assumes two stochastic generative processes, which correspond to two submodels: 1) a label correlation model; and 2) a label mixture model. The former model formulates labels’ generative process, in which a label correlation network is created to depict the dependency between labels. Moreover, we present an efficient inference algorithm for calculating the generative probability of a multi-label class. Furthermore, in order to optimize the label correlation network, we propose a parameter-learning algorithm based on gradient descent. The second submodel in the LCMM depicts the generative process of words in a document with the given labels. Different traditional mixture models can be adopted in this generative process, such as the mixture of language models, or topic models. In the multi-label classification stage, we propose a two-step strategy to most efficiently utilize the LCMM based on the framework of Bayes decision theory. We conduct extensive multi-label classification experiments on three standard text data sets. The experimental results show significant performance improvements comparing to existing approaches. For example, the improvements on accuracy and macro F-score measures in the OHSUMED data set achieve 28.3% and 37.0%, respectively. These performance enhancements demonstrate the effectiveness of the proposed models and solutions.
Zhiyang He, Ji Wu 0002, Ping Lv
Intell. Data Anal.2
2016 Hidden Softmax Sequence Model for Dialogue Structure Analysis
abstract
We propose a new unsupervised learning model, hidden softmax sequence model (HSSM), based on Boltzmann machine for dialogue structure analysis.The model employs three types of units in the hidden layer to discovery dialogue latent structures: softmax units which represent latent states of utterances; binary units which represent latent topics specified by dialogues; and a binary unit that represents the global general topic shared across the whole dialogue corpus.In addition, the model contains extra connections between adjacent hidden softmax units to formulate the dependency between latent states.Two different kinds of real world dialogue corpora, Twitter-Post and AirTicketBooking, are utilized for extensive comparing experiments, and the results illustrate that the proposed model outperforms sate-ofthe-art popular approaches.
Zhiyang He, Xien Liu, Ping Lv, Ji Wu 0002
ACL (1)4
2016 Improving the Probabilistic Framework for Representing Dialogue Systems with User Response Model
Miao Li 0003, Ji Wu 0002
INTERSPEECH3
2016 Target-Based State and Tracking Algorithm for Spoken Dialogue System
Miao Li 0003, Zhiyang He, Ji Wu 0002
INTERSPEECH3
2016 Objective Evaluation Methods for Chinese Text-To-Speech Systems
Ji Wu 0002, Sam Lai, Wenhui Lei, Carsten Isert
INTERSPEECH3
2016 The MSIIP system for dialog state tracking challenge 5
abstract
We present our work in Dialog State Tracking Challenge 5, the main task of which is to track dialog state on human-human conversations cross language. Firstly a probabilistic enhanced framework is used to represent sub-dialog, which consists of three parts, the input model for extracting features, the enhanced model for updating dialog state and the output model to give the tracking frame. Meanwhile, parallel language systems are proposed to overcome inaccuracy caused by machine translation for cross language testing. We also introduce a new iterative alignment method extended from our work in DSTC4. Furthermore, a slot-based score averaging method is introduced to build an ensemble by combining different trackers. Results of our DSTC5 system show that our method significantly improves tracking performance compared with baseline method.
Miao Li 0003, Ji Wu 0002
SLT3
2016 Entity disambiguation to Wikipedia using collective ranking
Ji Wu 0002, Dingding Wang 0001, Tao Li 0001
Inf. Process. Manag.2
2015 Rapid adaptation for deep neural networks through multi-task learning
abstract
We propose a novel approach to addressing the adaptation effectiveness issue in parameter adaptation for deep neural network (DNN) based acoustic models for automatic speech recognition by adding one or more small auxiliary output layers modeling broad acoustic units, such as mono-phones or tied-state (often called senone) clusters. In scenarios with a limited amount of available adaptation data, most senones are usually rarely seen or not observed, and consequently the ability to model them in a new condition is often not fully exploited. With the original senone classification task as the primary task, and adding auxiliary mono-phone/senone-cluster classification as the secondary tasks, multi-task learning (MTL) is employed to adapt the DNN parameters. With the proposed MTL adaptation framework, we improve the learning ability of the original DNN structure, then enlarge the coverage of the acoustic space to deal with the unseen senone problem, and thus enhance the discrimination power of the adapted DNN models. Experimental results on the 20,000-word open vocabulary WSJ task demonstrate that the proposed framework consistently outperforms the conventional linear hidden layer adaptation schemes without MTL by providing 5.4% relative reduction in word error rate (WERR) with only 1 single adaptation utterance, and 10.7% WERR with 40 adaptation utterances against the un-adapted DNN models.
Zhen Huang 0001, Jinyu Li 0001, Sabato Marco Siniscalchi, I-Fan Chen, Ji Wu 0002, Chin-Hui Lee 0001
INTERSPEECH5
2015 An entropy minimization framework for goal-driven dialogue management
Ji Wu 0002, Miao Li 0003, Chin-Hui Lee 0001
INTERSPEECH1
2015 A Probabilistic Framework for Representing Dialog Systems and Entropy-Based Dialog Management Through Dynamic Stochastic State Evolution
abstract
In this paper, we present a probabilistic framework for goal-driven spoken dialog systems. A new dynamic stochastic state (DS-state) is then defined to characterize the goal set of a dialog state at different stages of the dialog process. Furthermore, an entropy minimization dialog management (EMDM) strategy is also proposed to combine with the DS-states to facilitate a robust and efficient solution in reaching a user's goals. A song-on-demand task, with a total of 38 117 songs and 12 attributes corresponding to each song, is used to test the performance of the proposed approach. In an ideal simulation, assuming no errors, the EMDM strategy is the most efficient goal-seeking method among all tested approaches, returning the correct song within 3.3 dialog turns on average. Furthermore, in a practical scenario, with top five candidates to handle the unavoidable automatic speech recognition (ASR) and natural language understanding (NLU) errors, the results show that only 61.7% of the dialog goals can be successfully obtained in 6.23 dialog turns on average when random questions are asked by the system, whereas if the proposed DS-states are updated with the top five candidates from the SLU output using the proposed EMDM strategy executed at every DS-state, then a 86.7% dialog success rate can be accomplished effectively within 5.17 dialog turns on average. We also demonstrate that entropy-based DM strategies are more efficient than non-entropy based DM. Moreover, using the goal set distributions in EMDM, the results are better than those without them, such as in sate-of-the-art database summary DM.
Ji Wu 0002, Miao Li 0003, Chin-Hui Lee 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2014 Audio retrieval based on perceptual similarity
abstract
Given a short query audio clip, the goal of audio retrieval is to automatically fetch all similar clips from a given audio database. Different from traditional audio similarity which is mainly based on priori knowledge of objective reality, this paper proposes to use a more subjective method to m
Ji Wu 0002, Dingding Wang 0001, Tao Li 0001
CollaborateCom2
2014 Subword scheme for keyword search
abstract
Keyword search (KWS) is an important application of spoken language technology. The technique of Large Vocabulary Continuous Speech Recognition (LVCSR) is playing an important role in KWS system. However, for a language with large vocabulary and relatively insufficient text corpus, the vocabulary size keeps going up very quickly with the increasing amount of text, as we observed in Tamil. This brings difficulty in training a reliable language model, which may undermine KWS performance. Subword unit has been successfully employed in KWS system to handle out-of-vocabulary (OOV) problem. Inspired by this, we propose a novel subword scheme from the perspective of pronunciation to alleviate the large vocabulary problem. We find that the subword-based system outperforms our best word-based system on Tamil conversational telephone speech. The experiment of system combination shows that, over the best word-based system, a single subword-based system contains more complementary information than the total of that of the other three word-based systems.
Ji Wu 0002
SLT3
2014 Label correlation mixture model for multi-label text categorization
abstract
Multi-label text categorization is more difficult but practical than the conventional binary or multi-class text categorization. This paper propose a novel probabilistic generative model, label correlation mixture model (LCMM), to depict the multiple labeled documents, which can be used for multi-label text categorization. In LCMM, labels and topics have the one-to-one correspondences. LCMM consists of two parts: label correlation model and multi-label conditioned document model. The former one formulates the generating process of labels and the dependencies between the labels are taken into account. We also propose an efficient algorithm for calculating the probability of generating an arbitrary subset of labels. Multi-label conditioned document model can be regarded as a supervised label mixture model, in which the labels for a document are known. To evaluate LCMM, multi-label text categorization experiments on three standard text data sets are performed. The experimental results demonstrate the effectiveness of LCMM, comparing to other reported methods.
Zhiyang He, Ji Wu 0002, Ping Lv
SLT2
2013 Denoising deep neural networks based voice activity detection
abstract
Recently, the deep-belief-networks (DBN) based voice activity detection (VAD) has been proposed. It is powerful in fusing the advantages of multiple features, and achieves the state-of-the-art performance. However, the deep layers of the DBN-based VAD do not show an apparent superiority to the shallower layers. In this paper, we propose a denoising-deep-neural-network (DDNN) based VAD to address the aforementioned problem. Specifically, we pre-train a deep neural network in a special unsupervised denoising greedy layer-wise mode, and then fine-tune the whole network in a supervised way by the common back-propagation algorithm. In the pre-training phase, we take the noisy speech signals as the visible layer and try to extract a new feature that minimizes the reconstruction cross-entropy loss between the noisy speech signals and its corresponding clean speech signals. Experimental results show that the proposed DDNN-based VAD not only outperforms the DBN-based VAD but also shows an apparent performance improvement of the deep layers over shallower layers.
Xiao-Lei Zhang 0001, Ji Wu 0002
ICASSP2
2013 Weight optimization and layered clustering-based ECOC
abstract
Error correcting output code (ECOC) is a general framework of solving a multiclass classification problem via a binary-class classifier ensemble. In this paper, we propose a new heuristic coding method, named weight optimization and layered clustering-based ECOC (WOLC-ECOC). It iterates the following two steps until the training risk converges. The first step employs the layered clustering-based approach [1]. The approach can construct multiple different strong binary-class classifiers on a given binary-class problem, so that the heuristic training process will not be blocked by some difficult binary-class problems. The second step is the weight optimization technique [2]. It guarantees the non-increasing of the heuristic training process whenever we add new classifiers to the ECOC ensemble. Experimental results on several benchmark sets demonstrate that WOLC-ECOC is more effective than 15 referenced coding-decoding ECOC pairs.
Xiao-Lei Zhang 0001, Ji Wu 0002
ICASSP2
2013 Deep Belief Networks Based Voice Activity Detection
abstract
Fusing the advantages of multiple acoustic features is important for the robustness of voice activity detection (VAD). Recently, the machine-learning-based VADs have shown a superiority to traditional VADs on multiple feature fusion tasks. However, existing machine-learning-based VADs only utilize shallow models, which cannot explore the underlying manifold of the features. In this paper, we propose to fuse multiple features via a deep model, called deep belief network (DBN). DBN is a powerful hierarchical generative model for feature extraction. It can describe highly variant functions and discover the manifold of the features. We take the multiple serially-concatenated features as the input layer of DBN, and then extract a new feature by transferring these features through multiple nonlinear hidden layers. Finally, we predict the class of the new feature by a linear classifier. We further analyze that even a single-hidden-layer-based belief network is as powerful as the state-of-the-art models in the machine-learning-based VADs. In our empirical comparison, ten common features are used for performance analysis. Extensive experimental results on the AURORA2 corpus show that the DBN-based VAD not only outperforms eleven referenced VADs, but also can meet the real-time detection demand of VAD. The results also show that the DBN-based VAD can fuse the advantages of multiple features effectively.
Xiao-Lei Zhang 0001, Ji Wu 0002
IEEE Trans. Speech Audio Process.2
2012 Optimized weighted decoding for error-correcting output codes
abstract
A common method to solve a multiclass classification problem is to reduce the problem to a serial binary classification problems and combine them via Error-Correcting Output Codes (ECOC). The ECOC contains three parts: coding design, decoding algorithm, and base dichotomizer. Recently, the Loss-Weighted (LW) decoding algorithm (Escalera et al., PAMI2010), which introduces a weight matrix to the Loss-Based (LB) decoding (Allwein et al., JMLR2001), achieves improved performance over traditional decoding methods. However, the weight matrix is assigned empirically. In this paper, we present a theoretical global optimization method for the weight matrix, so as to achieve the minimal training risk. Although the experimental results on real-world image, audio and text classification tasks show that the proposed decoding method only leads to slightly better performances than others in the case of discrete outputs of the dichotomizers, the proposed method provides a new screen on the decoding methods of the ECOC.
Xiao-Lei Zhang 0001, Ji Wu 0002, Ping Lv
ICASSP2
2012 Linearithmic Time Sparse and Convex Maximum Margin Clustering
abstract
Recently, a new clustering method called maximum margin clustering (MMC) was proposed and has shown promising performances. It was originally formulated as a difficult nonconvex integer problem. To make the MMC problem practical, the researchers either relaxed the original MMC problem to inefficient convex optimization problems or reformulated it to nonconvex optimization problems, which sacrifice the convexity for efficiency. However, no approaches can both hold the convexity and be efficient. In this paper, a new linearithmic time sparse and convex MMC algorithm, called support-vector-regression-based MMC (SVR-MMC), is proposed. Generally, it first uses the SVR as the core of the MMC. Then, it is relaxed as a convex optimization problem, which is iteratively solved by the cutting-plane algorithm. Each cutting-plane subproblem is further decomposed to a serial supervised SVR problem by a new global extended-level method (GELM). Finally, each supervised SVR problem is solved in a linear time complexity by a new sparse-kernel SVR (SKSVR) algorithm. We further extend the SVR-MMC algorithm to the multiple-kernel clustering (MKC) problem and the multiclass MMC (M3C) problem, which are denoted as SVR-MKC and SVR-M3C, respectively. One key point of the algorithms is the utilization of the SVR. It can prevent the MMC and its extensions meeting an integer matrix programming problem. Another key point is the new SKSVR. It provides a linear time interface to the nonlinear kernel scenarios, so that the SVR-MMC and its extensions can keep a linearthmic time complexity in nonlinear kernel scenarios. Our experimental results on various real-world data sets demonstrate the effectiveness and the efficiency of the SVR-MMC and its two extensions. Moreover, the unsupervised application of the SVR-MKC to the voice activity detection (VAD) shows that the SVR-MKC can achieve good performances that are close to its supervised counterpart, meet the real-time demand of the VAD, and need no labeling for model training.
Xiao-Lei Zhang 0001, Ji Wu 0002
IEEE Trans. Syst. Man Cybern. Part B2
2011 An Active Learning Approach to Task Adaptation
Ji Wu 0002, Zhiyang He, Ping Lv
INTERSPEECH1
2011 Maximum Margin Clustering Based Statistical VAD With Multiple Observation Compound Feature
abstract
In this letter, we propose a new robust feature and an unsupervised learning approach for statistical voice activity detection (VAD). Maximum margin clustering (MMC), as an unsupervised classifier, can improve the robustness of support vector machine (SVM) based VAD while requiring no data labeling for model training. In the MMC framework, the multiple observation compound feature (MO-CF) is proposed to improve accuracy. MO-CF is composed of two subfeatures—multiple observation signal-to-noise ratio (MO-SNR) and multiple observation maximum probability (MO-MP). The contributions of the two subfeatures are balanced by a factor which is chosen to yield the largest area under the ROC curve (AUC) of the performance. The proposed approach obtains improved performance over seven commonly used VAD techniques in the experiments covering various noisy scenarios with low SNRs.
Ji Wu 0002, Xiao-Lei Zhang 0001
IEEE Signal Process. Lett.1
2011 Efficient Multiple Kernel Support Vector Machine Based Voice Activity Detection
abstract
In this letter, we propose a multiple kernel support vector machine (MK-SVM) method for multiple feature based VAD. To make the MK-SVM based VAD practical, we adapt the multiple kernel learning (MKL) thought to an efficient cutting-plane structural SVM solver. We further discuss the performances of the MK-SVM with two different optimization objectives, in terms of minimum classification errors (MCE) and improvement of receiver operating characteristic (ROC) curves. Our experimental results show that the proposed method not only leads to better global performances by taking the advantages of multiple features but also has a low computational complexity.
Ji Wu 0002, Xiao-Lei Zhang 0001
IEEE Signal Process. Lett.1
2010 A new VAD framework using statistical model and human knowledge based empirical rule
Ji Wu 0002, Xiao-Lei Zhang 0001
INTERSPEECH1
2009 Automatic punctuation generation for speech
abstract
Automatic generation of punctuation is an essential feature for many speech-to-text transcription tasks. This paper describes a maximum a-posteriori (MAP) approach for inserting punctuation marks into raw word sequences obtained from automatic speech recognition (ASR). The system consists of an ¿acoustic model¿ (AM) for prosodic features (actually pause duration) and a ¿language model¿ (LM) for text-only features. The LM combines three components: an MLP-based trigger-word model and a forward and a backward trigram punctuation predictor. The separation into acoustic and language model allows to learn these models on different corpora, especially allowing the LM to be trained on large amounts of data (text) for which no acoustic information is available. We find that the trigger-word LM is very useful, and further improvement can be achieved when combining both prosodic and lexical information. We achieve an F-measure of 81.0% and 56.5% for voicemails and podcasts, respectively, on reference transcripts, and 69.6% for voicemails on ASR transcripts.
Wenzhu Shen, Roger Peng Yu, Frank Seide, Ji Wu 0002
ASRU4
2009 Improvements on minimum covariance based Spatial correlation Transformation
abstract
In order to take advantage of the correlation information among different acoustic units in speech recognition, a novel approach named Minimum Covariance based Spatial Correlation Transformation was proposed in [8], which achieves satisfactory performance. However, there are two issues of this approach which can still be improved, 1) the estimation of the transformation matrix; 2) the construction of the history data. In this paper, a new algorithm of estimating the transformation matrix and a new strategy of constructing history supervector are proposed. Experimental results show that the improved approach achieves better performance than the original one.
Tengrong Su, Ji Wu 0002, Zuoying Wang
ICASSP2
2008 Spatial correlation transformation based on minimum covariance
abstract
In speech recognition, acoustic units are highly related. Different from some adaptation methods, such as Reference Speaker Weighting (RSW) and Eigenvoice, the correlation between different acoustic units in the feature space, which is called Spatial Correlation, focuses on the correlation information among different acoustic units of the same speaker. In this paper, a novel scheme using spatial correlation is proposed. In speech recognition system, with the spatial correlation information, the refined acoustic models are trained, and the transformation matrices are determined based on Minimum Covariance criteria. Experiments of this new algorithm show a significant improvement on speaker independent recognition systems.
Tengrong Su, Ji Wu 0002, Zuoying Wang
ICASSP2
2006 Pronunciation variation modeling for Mandarin with accent
Ji Wu 0002, Zuoying Wang
INTERSPEECH2
2004 A robust understanding model for spoken dialogues
Ji Wu 0002, Zuoying Wang
INTERSPEECH2
2003 Fuzzy clustering and Bayesian information criterion based threshold estimation for robust voice activity detection
abstract
In previous voice activity detection (VAD) approaches that use threshold, consistent accuracy cannot be achieved since the mean-value based and the histogram based threshold estimation algorithms are not robust. They strongly depend on the percentage of voice and background noise in the estimate interval. In this paper, fuzzy clustering and Bayesian information criterion are proposed to estimate the thresholds for VAD. Compared to previous algorithms, the new algorithm is more robust and heuristic-rules-free. It is insensitive to the estimated interval, and can maintain fast tracking speed of environment change when combined with online update. Experiment shows it works very well with energy features in both stationary and non-stationary environments.
Ji Wu 0002, Zuoying Wang, Dajin Lu
ICASSP (1)2
2003 A Chinese spoken dialogue system for train information
abstract
In this paper a Chinese spoken dialogue system developed for train information retrieval is presented. After a brief description of the system architecture and the individual modules, a dialogue manager, which integrates user plan inference with the topic tree model, is proposed. Also dialogue strategies based on this mechanism, including consistent information sharing across multiple topics, reliable user response expectation and proper system prompt design, are presented and explained in detail. Experiments show that sentence meaning understanding error rate decreased by 23.5% with the guide of user plan inference. Preliminary subjective evaluation shows that the users are interested and willing to talk with the system although there's still much to be improved.
Ji Wu 0002, Zuoying Wang
SMC2
2002 Robust Noisy Speech Recognition with Adaptive Frequency Bank Selection
abstract
With the development of automatic speech recognition technology, the robustness problem of speech recognition systems is becoming more and more important. This paper addresses the problem of speech recognition in an additive background noise environment. Since the frequency energy of different types of noise focuses on different frequency banks, the effects of additive noise on each frequency bank are different. The seriously obscured frequency banks have little word signal information left, and are harmful for subsequence speech processing. Wu and Lin (2000) applied the frequency bank selection theory to robust word boundary detection in a noisy environment, and obtained good detection results. In this paper, this theory is extended to noisy speech recognition. Unlike the standard MFCC which uses all frequency banks for cepstral coefficients, we only use the frequency banks that are slightly corrupted and discard the seriously obscured ones. Cepstral coefficients are calculated only on the selected frequency banks. Moreover, an acoustic model is also adapted to match the modification of the acoustic feature. Experiments on continuous digital speech recognition show that the proposed algorithm leads to better performance than spectral subtraction and cepstral mean normalization at low SNRs.
Ji Wu 0002, Zuoying Wang, Dajin Lu
ICMI2
2000 Gaussian similarity analysis and its application in speaker adaptation
Ji Wu 0002, Zuoying Wang
INTERSPEECH1