Wenxin Hou

dblp:270/4628 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
10since 2021 · last 2024
0000-0001-9297-6964ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 2 since 2021
YearPublicationVenuePosition
2024 The Good, The Bad, and Why: Unveiling Emotions in Generative AI
abstract
Emotion significantly impacts our daily behaviors and interactions. While recent generative AI models, such as large language models, have shown impressive performance in various tasks, it remains unclear whether they truly comprehend emotions and why. This paper aims to address this gap by incorporating psychological theories to gain a holistic understanding of emotions in generative AI models. Specifically, we propose three approaches: 1) EmotionPrompt to enhance AI model performance, 2) EmotionAttack to impair AI model performance, and 3) EmotionDecode to explain the effects of emotional stimuli, both benign and malignant. Through extensive experiments involving language and multi-modal models on semantic understanding, logical reasoning, and generation tasks, we demonstrate that both textual and visual EmotionPrompt can boost the performance of AI models while EmotionAttack can hinder it. More importantly, EmotionDecode reveals that AI models can comprehend emotional stimuli akin to the mechanism of dopamine in the human brain. Our work heralds a novel avenue for exploring psychology to enhance our understanding of generative AI models, thus boosting the research and development of human-AI collaboration and mitigating potential risks.
Jindong Wang 0001, Yixuan Zhang 0001, Kaijie Zhu, Wenxin Hou, Jianxun Lian, Qiang Yang 0001, Xing Xie 0001
ICML6
2024 Boosting Cross-Domain Speech Recognition With Self-Supervision
abstract
The cross-domain performance of automatic speech recognition (ASR) could be severely hampered due to the mismatch between training and testing distributions. Since the target domain usually lacks labeled data, and domain shifts exist at acoustic and linguistic levels, it is challenging to perform unsupervised domain adaptation (UDA) for ASR. Previous work has shown that self-supervised learning (SSL) or pseudo-labeling (PL) is effective in UDA by exploiting the self-supervisions of unlabeled data. However, these self-supervisions also face performance degradation in mismatched domain distributions, which previous work fails to address. This work presents a systematic UDA framework to fully utilize the unlabeled data with self-supervision in the pre-training and fine-tuning paradigm. On the one hand, we apply continued pre-training and data replay techniques to mitigate the domain mismatch of the SSL pre-trained model. On the other hand, we propose a domain-adaptive fine-tuning approach based on the PL technique with three unique modifications: Firstly, we design a dual-branch PL method to decrease the sensitivity to the erroneous pseudo-labels; Secondly, we devise an uncertainty-aware confidence filtering strategy to improve pseudo-label correctness; Thirdly, we introduce a two-step PL approach to incorporate target domain linguistic knowledge, thus generating more accurate target domain pseudo-labels. Experimental results on various cross-domain scenarios demonstrate that the proposed approach effectively boosts the cross-domain performance and significantly outperforms previous approaches.
Han Zhu 0004, Gaofeng Cheng, Jindong Wang 0001, Wenxin Hou, Pengyuan Zhang, Yonghong Yan 0002
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 FreeMatch: Self-adaptive Thresholding for Semi-supervised Learning
Yidong Wang 0003, Hao Chen 0102, Qiang Heng, Wenxin Hou, Zhen Wu 0002, Jindong Wang 0001, Marios Savvides, Takahiro Shinozaki, Bhiksha Raj, Bernt Schiele, Xing Xie 0001
ICLR4
2022 Margin Calibration for Long-Tailed Visual Recognition
Yidong Wang 0003, Wenxin Hou, Zhen Wu 0002, Jindong Wang 0001, Takahiro Shinozaki
ACML3
2022 Exploiting Unlabeled Data for Target-Oriented Opinion Words Extraction
abstract
Target-oriented Opinion Words Extraction (TOWE) is a fine-grained sentiment analysis task that aims to extract the corresponding opinion words of a given opinion target from the sentence. Recently, deep learning approaches have made remarkable progress on this task. Nevertheless, the TOWE task still suffers from the scarcity of training data due to the expensive data annotation process. Limited labeled data increase the risk of distribution shift between test data and training data. In this paper, we propose exploiting massive unlabeled data to reduce the risk by increasing the exposure of the model to varying distribution shifts. Specifically, we propose a novel Multi-Grained Consistency Regularization (MGCR) method to make use of unlabeled data and design two filters specifically for TOWE to filter noisy data at different granularity. Extensive experimental results on four TOWE benchmark datasets indicate the superiority of MGCR compared with current state-of-the-art methods. The in-depth analysis also demonstrates the effectiveness of the different-granularity filters.
Yidong Wang 0003, Hao Wu 0059, Ao Liu 0008, Wenxin Hou, Zhen Wu 0002, Jindong Wang 0001, Takahiro Shinozaki, Manabu Okumura, Yue Zhang 0004
COLING4
2022 USB: A Unified Semi-supervised Learning Benchmark for Classification
abstract
Semi-supervised learning (SSL) improves model generalization by leveraging massive unlabeled data to augment limited labeled samples. However, currently, popular SSL evaluation protocols are often constrained to computer vision (CV) tasks. In addition, previous work typically trains deep neural networks from scratch, which is time-consuming and environmentally unfriendly. To address the above issues, we construct a Unified SSL Benchmark (USB) for classification by selecting 15 diverse, challenging, and comprehensive tasks from CV, natural language processing (NLP), and audio processing (Audio), on which we systematically evaluate the dominant SSL methods, and also open-source a modular and extensible codebase for fair evaluation of these SSL methods. We further provide the pre-trained versions of the state-of-the-art neural models for CV tasks to make the cost affordable for further tuning. USB enables the evaluation of a single SSL algorithm on more tasks from multiple domains but with less cost. Specifically, on a single NVIDIA V100, only 39 GPU days are required to evaluate FixMatch on 15 tasks in USB while 335 GPU days (279 GPU days on 4 CV datasets except for ImageNet) are needed on 5 CV tasks with TorchSSL.
Yidong Wang 0003, Hao Chen 0102, Wang Sun, Ran Tao 0013, Wenxin Hou, Linyi Yang, Zhi Zhou 0007, Lan-Zhe Guo, Heli Qi, Zhen Wu 0002, Yufeng Li 0008, Satoshi Nakamura 0001, Wei Ye 0004, Marios Savvides, Bhiksha Raj, Takahiro Shinozaki, Bernt Schiele, Jindong Wang 0001, Xing Xie 0001, Yue Zhang 0004
NeurIPS6
2022 Exploiting Adapters for Cross-Lingual Low-Resource Speech Recognition
abstract
Cross-lingual speech adaptation aims to solve the problem of leveraging multiple rich-resource languages to build models for a low-resource target language. Since the low-resource language has limited training data, speech recognition models can easily overfit. Adapter is a versatile module that can be plugged into Transformer for parameter-efficient learning. In this paper, we propose to use adapters for parameter-efficient cross-lingual speech adaptation. Based on our previous MetaAdapter that implicitly leverages adapters, we propose a novel algorithm called SimAdapter for explicitly learning knowledge from adapters. Our algorithms can be easily integrated into the Transformer structure. MetaAdapter leverages meta-learning to transfer the general knowledge from training data to the test language. SimAdapter aims to learn the similarities between the source and target languages during fine-tuning using the adapters. We conduct extensive experiments on five-low-resource languages in the Common Voice dataset. Results demonstrate that MetaAdapter and SimAdapter can reduce WER by 2.98% and 2.55% with only 2.5% and 15.5% of trainable parameters compared to the strong full-model fine-tuning baseline. Moreover, we show that these two novel algorithms can be integrated for better performance with up to 3.55% relative WER reduction.
Wenxin Hou, Han Zhu 0004, Yidong Wang 0003, Jindong Wang 0001, Tao Qin 0001, Renjun Xu, Takahiro Shinozaki
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Meta-Adapter: Efficient Cross-Lingual Adaptation With Meta-Learning
abstract
Transfer learning from a multilingual model has shown favorable results on low-resource automatic speech recognition (ASR). However, full-model fine-tuning generates a separate model for every target language and is not suitable for deploying and maintaining in production. The key challenge lies in how to efficiently extend the pre-trained model with fewer parameters. In this paper, we propose to combine the adapter module with meta-learning algorithms to achieve high recognition performance under low-resource settings and improve the parameter-efficiency of the model. Extensive experiments show that our methods can achieve comparable or even superior recognition rates than the state-of-the-art baselines on low-resource languages, especially under very-low-resource conditions, with a significantly smaller model profile.
Wenxin Hou, Yidong Wang 0003, Shengzhou Gao, Takahiro Shinozaki
ICASSP1
2021 Cross-Domain Speech Recognition with Unsupervised Character-Level Distribution Matching
abstract
End-to-end automatic speech recognition (ASR) can achieve promising performance with large-scale training data. However, it is known that domain mismatch between training and testing data often leads to a degradation of recognition accuracy. In this work, we focus on the unsupervised domain adaptation for ASR and propose CMatch, a Character-level distribution matching method to perform fine-grained adaptation between each character in two domains. First, to obtain labels for the features belonging to each character, we achieve frame-level label assignment using the Connectionist Temporal Classification (CTC) pseudo labels. Then, we match the character-level distributions using Maximum Mean Discrepancy. We train our algorithm using the self-training technique. Experiments on the Libri-Adapt dataset show that our proposed approach achieves 14.39% and 16.50% relative Word Error Rate (WER) reduction on both cross-device and cross-environment ASR. We also comprehensively analyze the different strategies for frame-level label assignment and Transformer adaptations.
Wenxin Hou, Jindong Wang 0001, Xu Tan 0003, Tao Qin 0001, Takahiro Shinozaki
Interspeech1
2021 FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling
abstract
The recently proposed FixMatch achieved state-of-the-art results on most semi-supervised learning (SSL) benchmarks. However, like other modern SSL algorithms, FixMatch uses a pre-defined constant threshold for all classes to select unlabeled data that contribute to the training, thus failing to consider different learning status and learning difficulties of different classes. To address this issue, we propose Curriculum Pseudo Labeling (CPL), a curriculum learning approach to leverage unlabeled data according to the model's learning status. The core of CPL is to flexibly adjust thresholds for different classes at each time step to let pass informative unlabeled data and their pseudo labels. CPL does not introduce additional parameters or computations (forward or backward propagation). We apply CPL to FixMatch and call our improved algorithm FlexMatch. FlexMatch achieves state-of-the-art performance on a variety of SSL benchmarks, with especially strong performances when the labeled data are extremely limited or when the task is challenging. For example, FlexMatch achieves 13.96% and 18.96% error rate reduction over FixMatch on CIFAR-100 and STL-10 datasets respectively, when there are only 4 labels per class. CPL also significantly boosts the convergence speed, e.g., FlexMatch can use only 1/5 training time of FixMatch to achieve even better performance. Furthermore, we show that CPL can be easily adapted to other SSL algorithms and remarkably improve their performances. We open-source our code at https://github.com/TorchSSL/TorchSSL.
Yidong Wang 0003, Wenxin Hou, Hao Wu 0059, Jindong Wang 0001, Manabu Okumura, Takahiro Shinozaki
NeurIPS3
2020 Spoken Language Acquisition Based on Reinforcement Learning and Word Unit Segmentation
abstract
The process of spoken-language acquisition has been one of the topics of greatest interest to linguists for decades. By uti-lizing modern machine learning techniques, we simulated this process on computers, which helps to understand it and develop new possibilities of applying this concept on intelligent robots, among other things. This paper proposes a new framework for simulating spoken-language acquisition by combining reinforcement learning and unsupervised learning methods. Our experiments also show that a spoken language can be acquired considerably faster by identifying potential word segments from collected ambient sounds in an unsupervised manner.
Shengzhou Gao, Wenxin Hou, Tomohiro Tanaka, Takahiro Shinozaki
ICASSP2
2020 Large-Scale End-to-End Multilingual Speech Recognition and Language Identification with Multi-Task Learning
Wenxin Hou, Bairong Zhuang, Longfei Yang, Jiatong Shi, Takahiro Shinozaki
INTERSPEECH1
2020 Sound-Image Grounding Based Focusing Mechanism for Efficient Automatic Spoken Language Acquisition
Mingxin Zhang 0008, Tomohiro Tanaka, Wenxin Hou, Shengzhou Gao, Takahiro Shinozaki
INTERSPEECH3