Huaming Wang

dblp:120/5267 · DBLP profile ↗
← Back
26ranked-venue papers
1as first author
19since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 1 first-author · 16 since 2021Artificial intelligence and machine learning · 15 · 9 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 Training Audio Captioning Models without Audio
abstract
Automated Audio Captioning (AAC) is the task of generating natural language descriptions given an audio stream. A typical AAC system requires manually curated training data of audio segments and corresponding text caption annotations. The creation of these audio-caption pairs is costly, resulting in general data scarcity for the task. In this work, we address this major limitation and propose an approach to train AAC systems using only text. Our approach leverages the multimodal space of contrastively trained audio-text models, such as CLAP. During training, a decoder generates captions conditioned on the pretrained CLAP text encoder. During inference, the text encoder is replaced with the pretrained CLAP audio encoder. To bridge the modality gap between text and audio embeddings, we propose the use of noise injection or a learnable adapter, during training. We find that the proposed text-only framework performs competitively with stateof-the-art models trained with paired audio, showing that efficient text-to-audio transfer is possible. Finally, we showcase both stylized audio captioning and caption enrichment while training without audio or human-created text captions.
Soham Deshmukh, Benjamin Elizalde, Dimitra Emmanouilidou, Bhiksha Raj, Rita Singh, Huaming Wang
ICASSP6
2024 Prompting Audios Using Acoustic Properties for Emotion Representation
abstract
Emotions lie on a continuum, but current models treat emotions as a finite valued discrete variable. This representation does not capture the diversity in the expression of emotion. To better represent emotions we propose the use of natural language descriptions (or prompts). In this work, we address the challenge of automatically generating these prompts and training a model to better learn emotion representations from audio and prompt pairs. We use acoustic properties that are correlated to emotion like pitch, intensity, speech rate, and articulation rate to automatically generate prompts i.e. ‘acoustic prompts’. We use a contrastive learning objective to map speech to their respective acoustic prompts. We evaluate our model on Emotion Audio Retrieval and Speech Emotion Recognition. Our results show that the acoustic prompts significantly improve the model’s performance in EAR, in various Precision@K metrics. In SER, we observe a 3.8% relative accuracy improvement on the Ravdess dataset.
Hira Dhamyal, Benjamin Elizalde, Soham Deshmukh, Huaming Wang, Bhiksha Raj, Rita Singh
ICASSP4
2024 Natural Language Supervision For General-Purpose Audio Representations
abstract
Audio-Language models jointly learn multimodal text and audio representations that enable Zero-Shot inference. Models rely on the encoders to create powerful representations of the input and generalize to multiple tasks ranging from sounds, music, and speech. Although models have achieved remarkable performance, there is still a gap with task-specific models. In this paper, we propose a Contrastive Language-Audio Pretraining model that is pretrained with a diverse collection of 4.6M audio-text pairs employing two innovative encoders for Zero-Shot inference. To learn audio representations, we trained an audio encoder on 22 audio tasks, instead of the standard training of sound event classification. To learn language representations, we trained an autoregressive decoder-only model instead of the standard encoder-only models. Then, the audio and language representations are brought into a joint multimodal space using Contrastive Learning. We used our encoders to improve the downstream performance by a large margin. We extensively evaluated the generalization of our representations on 26 downstream tasks, the largest in the literature. Our model achieves state of the art results in several tasks outperforming 4 different models and leading the way towards general-purpose audio representations. Code is on GitHub1.
Benjamin Elizalde, Soham Deshmukh, Huaming Wang
ICASSP3
2024 PAM: Prompting Audio-Language Models for Audio Quality Assessment
Soham Deshmukh, Dareen Alharthi, Benjamin Elizalde, Hannes Gamper, Mahmoud Al Ismail, Rita Singh, Bhiksha Raj, Huaming Wang
INTERSPEECH8
2024 NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant Meeting Transcription
Alon Vinnikov, Amir Ivry, Aviv Hurvitz, Igor Abramovski, Sharon Koubi, Ilya Gurvich, Shai Pe'er, Benjamin Elizalde, Naoyuki Kanda, Xiaofei Wang 0009, Shalev Shaer, Stav Yagev, Yossi Asher, Sunit Sivasankaran, Yifan Gong 0001, Huaming Wang, Eyal Krupka
INTERSPEECH18
2024 Both real-valued and binary multi-feature fusion histograms for 3D local shape representation
Linbo Hao, Huaming Wang
Vis. Comput.5
2023 CLAP Learning Audio Concepts from Natural Language Supervision
abstract
Mainstream machine listening models are trained to learn audio concepts under the paradigm of one class label to many recordings focusing on one task. Learning under such restricted supervision limits the flexibility of models because they require labeled audio for training and can only predict the predefined categories. Instead, we propose to learn audio concepts from natural language supervision. We call our approach Contrastive Language-Audio Pretraining (CLAP), which connects language and audio by using two encoders and a contrastive learning objective, bringing audio and text descriptions into a joint multimodal space. We trained CLAP with 128k audio and text pairs and evaluated it on 16 downstream tasks across 7 domains, such as classification of sound events, scenes, music, and speech. CLAP establishes state-of-the-art (SoTA) in Zero-Shot performance. Also, we evaluated CLAP’s audio encoder in a supervised learning setup and achieved SoTA in 5 tasks. The Zero-Shot capability removes the need of training with class labeled audio, enables flexible class prediction at inference time, and generalizes well in multiple downstream tasks. Code is available at: https://github.com/microsoft/CLAP.
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, Huaming Wang
ICASSP4
2023 Real-Time Audio-Visual End-To-End Speech Enhancement
abstract
Audio-visual speech enhancement (AV-SE) methods utilize auxiliary visual cues to enhance speakers’ voices. Therefore, technically they should be able to outperform the audio-only speech enhancement (SE) methods. However, there are few works in the literature on an AV-SE system that can work in real time on a CPU. In this paper, we propose a low-latency real-time audio-visual end-to-end enhancement (AV-E3Net) model based on the recently proposed end-to-end enhancement network (E3Net). Our main contribution includes two aspects: 1) We employ a dense connection module to solve the performance degradation caused by the deep model structure. This module significantly improves the model’s performance on the AV-SE task. 2) We propose a multi-stage gating-and-summation (GS) fusion module to merge audio and visual cues. Our results show that the proposed model provides better perceptual quality and intelligibility than the baseline E3net model with a negligible computational cost increase.
Zirun Zhu, Hemin Yang, Sefik Emre Eskimez, Huaming Wang
ICASSP6
2023 General GAN-generated Image Detection by Data Augmentation in Fingerprint Domain
abstract
In this work, we investigate improving the generalizability of GAN-generated image detectors by performing data augmentation in the fingerprint domain. Specifically, we first separate the fingerprints and contents of the GAN-generated images using an autoencoder based GAN fingerprint extractor, followed by random perturbations of the fingerprints. Then the original fingerprints are substituted with the perturbed fingerprints and added to the original contents, to produce images that are visually invariant but with distinct fingerprints. The perturbed images can successfully imitate images generated by different GANs to improve the generalization of the detectors, which is demonstrated by the spectra visualization. To our knowledge, we are the first to conduct data augmentation in the fingerprint domain. Our work explores a novel prospect that is distinct from previous works on spatial and frequency domains augmentation. Extensive cross-GAN experiments demonstrate the effectiveness of our method compared to the state-of-the-art methods in detecting fake images generated by unknown GANs.
Huaming Wang, Jianwei Fei, Yunshu Dai, Lingyun Leng, Zhihua Xia
ICME1
2023 Audio Retrieval with WavText5K and CLAP Training
Soham Deshmukh, Benjamin Elizalde, Huaming Wang
INTERSPEECH3
2023 Real-Time Joint Personalized Speech Enhancement and Acoustic Echo Cancellation
Sefik Emre Eskimez, Takuya Yoshioka, Alex Ju, Tanel Pärnamaa, Huaming Wang
INTERSPEECH6
2023 Pengi: An Audio Language Model for Audio Tasks
abstract
In the domain of audio processing, Transfer Learning has facilitated the rise of Self-Supervised Learning and Zero-Shot Learning techniques. These approaches have led to the development of versatile models capable of tackling a wide array of tasks, while delivering state-of-the-art performance. However, current models inherently lack the capacity to produce the requisite language for open-ended tasks, such as Audio Captioning or Audio Question Answering. We introduce Pengi, a novel Audio Language Model that leverages Transfer Learning by framing all audio tasks as text-generation tasks. It takes as input, an audio recording, and text, and generates free-form text as output. The input audio is represented as a sequence of continuous embeddings by an audio encoder. A text encoder does the same for the corresponding text input. Both sequences are combined as a prefix to prompt a pre-trained frozen language model. The unified architecture of Pengi enables open-ended tasks and close-ended tasks without any additional fine-tuning or task-specific extensions. When evaluated on 21 downstream tasks, our approach yields state-of-the-art performance in several of them. Our results show that connecting language models with audio models is a major step towards general-purpose audio understanding.
Soham Deshmukh, Benjamin Elizalde, Rita Singh, Huaming Wang
NeurIPS4
2023 Rotational Voxels Statistics Histogram for both real-valued and binary feature representations of 3D local shape
Linbo Hao, Xuefeng Yang, Wentao Yi, Huaming Wang
J. Vis. Commun. Image Represent.6
2022 Attentional Local Contrastive Learning for Face Forgery Detection
Yunshu Dai, Jianwei Fei, Huaming Wang, Zhihua Xia
ICANN (1)3
2022 Personalized speech enhancement: new models and Comprehensive evaluation
abstract
Personalized speech enhancement (PSE) models utilize additional cues, such as speaker embeddings like d-vectors, to remove background noise and interfering speech in real-time and thus improve the speech quality of online video conferencing systems for various acoustic scenarios. In this work, we propose two neural networks for PSE that achieve superior performance to the previously proposed VoiceFilter. In addition, we create test sets that capture a variety of scenarios that users can encounter during video conferencing. Furthermore, we propose a new metric to measure the target speaker over-suppression (TSOS) problem, which was not sufficiently investigated before despite its critical importance in deployment. Besides, we propose multi-task training with a speech recognition back-end. Our results show that the proposed models can yield better speech recognition accuracy, speech intelligibility, and perceptual quality than the baseline models, and the multi-task training can alleviate the TSOS issue in addition to improving the speech recognition accuracy.
Sefik Emre Eskimez, Takuya Yoshioka, Huaming Wang, Xiaofei Wang 0009, Zhuo Chen 0006, Xuedong Huang 0001
ICASSP3
2022 One Model to Enhance Them All: Array Geometry Agnostic Multi-Channel Personalized Speech Enhancement
abstract
With the recent surge of video conferencing tools usage, providing high-quality speech signals and accurate captions have become essential to conduct day-to-day business or connect with friends and families. Single-channel personalized speech enhancement (PSE) methods show promising results compared with the unconditional speech enhancement (SE) methods in these scenarios due to their ability to remove interfering speech in addition to the environmental noise. In this work, we leverage spatial information afforded by microphone arrays to improve such systems’ performance further. We investigate the relative importance of speaker embeddings and spatial features. Moreover, we propose a new causal array-geometry-agnostic multi-channel PSE model, which can generate a high-quality enhanced signal from arbitrary microphone geometry. Experimental results show that the proposed geometry agnostic model outperforms the model trained on a specific microphone array geometry in both speech quality and automatic speech recognition accuracy. We also demonstrate the effectiveness of the proposed approach for unseen array geometries.
Hassan Taherian, Sefik Emre Eskimez, Takuya Yoshioka, Huaming Wang, Zhuo Chen 0006, Xuedong Huang 0001
ICASSP4
2022 Fast Real-time Personalized Speech Enhancement: End-to-End Enhancement Network (E3Net) and Knowledge Distillation
abstract
This paper investigates how to improve the runtime speed of personalized speech enhancement (PSE) networks while maintaining the model quality.Our approach includes two aspects: architecture and knowledge distillation (KD).We propose an end-to-end enhancement (E3Net) model architecture, which is 3× faster than a baseline STFT-based model.Besides, we use KD techniques to develop compressed student models without significantly degrading quality.In addition, we investigate using noisy data without reference clean signals for training the student models, where we combine KD with multi-task learning (MTL) using an automatic speech recognition (ASR) loss.Our results show that E3Net provides better speech and transcription quality with a lower target speaker over-suppression (TSOS) rate than the baseline model.Furthermore, we show that the KD methods can yield student models that are 2 -4× faster than the teacher and provides reasonable quality.Combining KD and MTL improves the ASR and TSOS metrics without degrading the speech quality.
Manthan Thakker, Sefik Emre Eskimez, Takuya Yoshioka, Huaming Wang
INTERSPEECH4
2022 Geometric feature statistics histogram for both real-valued and binary feature representations of 3D local shape
Linbo Hao, Huaming Wang
Image Vis. Comput.2
2021 Human Listening and Live Captioning: Multi-Task Training for Speech Enhancement
abstract
With the surge of online meetings, it has become more critical than ever to provide high-quality speech audio and live captioning under various noise conditions.However, most monaural speech enhancement (SE) models introduce processing artifacts and thus degrade the performance of downstream tasks, including automatic speech recognition (ASR).This paper proposes a multi-task training framework to make the SE models unharmful to ASR.Because most ASR training samples do not have corresponding clean signal references, we alternately perform two model update steps called SE-step and ASR-step.The SEstep uses clean and noisy signal pairs and a signal-based loss function.The ASR-step applies a pre-trained ASR model to training signals enhanced with the SE model.A cross-entropy loss between the ASR output and reference transcriptions is calculated to update the SE model parameters.Experimental results with realistic large-scale settings using ASR models trained on 75,000-hour data show that the proposed framework improves the word error rate for the SE output by 11.82% with little compromise in the SE quality.Performance analysis is also carried out by changing the ASR model, the data used for the ASR-step, and the schedule of the two update steps.
Sefik Emre Eskimez, Xiaofei Wang 0009, Hemin Yang, Zirun Zhu, Zhuo Chen 0006, Huaming Wang, Takuya Yoshioka
Interspeech7
2020 Knowledge Distance Measure for the Multigranularity Rough Approximations of a Fuzzy Concept
abstract
Different rough approximation spaces could be induced for an information system by its different attribute subsets, thus the multigranularity rough approximations of a fuzzy concept could be developed. Research on the uncertainty in multi-granulation spaces becomes a basic issue of uncertainty measure. If the uncertainty measure is not accurate enough, two different rough approximation spaces of a fuzzy concept may have the same uncertainty, and the difference between them for describing a fuzzy concept cannot be reflected. In this case, attribute reduction, granularity selection, and multigranularity measure cannot be conducted effectively. Therefore, establishing an uncertainty measure model with strong distinguishing ability in multi-granulation spaces is a key issue in uncertainty knowledge processing. In this paper, this problem will be solved in the view of knowledge distance. First, a fuzzy knowledge distance measure (FKD) based on the Earth Mover's distance is introduced. Even if two rough approximation spaces possess the same uncertainty when describing a fuzzy concept, they can be discriminated by FKD. Then, by studying the change rules of the FKD in a hierarchical quotient space structure, it is found that the FKD between any two rough approximation spaces in an HQSS is equal to the difference between their granularity measure or information measure. Furthermore, in order to show the applicability of the FKD, the FKD is used in granularity selection, attribute reduct, and multigranularity measure. The experimental results show that the FKD-based attribute significance function has a more powerful ability to obtain shorter reduct and it is more robustness, which show the effectiveness of the FKD.
Jie Yang 0052, Guoyin Wang 0001, Qinghua Zhang 0001, Huaming Wang
IEEE Trans. Fuzzy Syst.4
2019 Advances in Online Audio-Visual Meeting Transcription
abstract
This paper describes a system that generates speaker-annotated transcripts of meetings by using a microphone array and a 360-degree camera. The hallmark of the system is its ability to handle overlapped speech, which has been an unsolved problem in realistic settings for over a decade. We show that this problem can be addressed by using a continuous speech separation approach. In addition, we describe an online audio-visual speaker diarization method that leverages face tracking and identification, sound source localization, speaker identification, and, if available, prior speaker information for robustness to various real world challenges. All components are integrated in a meeting transcription framework called SRD, which stands for “separate, recognize, and diarize”. Experimental results using recordings of natural meetings involving up to 11 attendees are reported. The continuous speech separation improves a word error rate (WER) by 16.1% compared with a highly tuned beamformer. When a complete list of meeting attendees is available, the discrepancy between WER and speaker-attributed WER is only 1.0%, indicating accurate word-to-speaker association. This increases marginally to 1.6% when 50% of the attendees are unknown to the system.
Takuya Yoshioka, Yan Huang 0028, Aviv Hurvitz, Sharon Koubi, Eyal Krupka, Ido Leichter, Changliang Liu, Partha Parthasarathy, Alon Vinnikov, Lingfeng Wu, Igor Abramovski, Wayne Xiong, Huaming Wang, Jun Zhang 0066, Yong Zhao 0008, Tianyan Zhou, Cem Aksoylar, Zhuo Chen 0006, Moshe David, Dimitrios Dimitriadis, Yifan Gong 0001, Ilya Gurvich, Xuedong Huang 0001
ASRU15
2018 Attention-Aware Compositional Network for Person Re-Identification
abstract
Person re-identification (ReID) is to identify pedestrians observed from different camera views based on visual appearance. It is a challenging task due to large pose variations, complex background clutters and severe occlusions. Recently, human pose estimation by predicting joint locations was largely improved in accuracy. It is reasonable to use pose estimation results for handling pose variations and background clutters, and such attempts have obtained great improvement in ReID performance. However, we argue that the pose information was not well utilized and hasn't yet been fully exploited for person ReID. In this work, we introduce a novel framework called Attention-Aware Compositional Network (AACN) for person ReID. AACN consists of two main components: Pose-guided Part Attention (PPA) and Attention-aware Feature Composition (AFC). PPA is learned and applied to mask out undesirable background features in pedestrian feature maps. Furthermore, pose-guided visibility scores are estimated for body parts to deal with part occlusion in the proposed AFC module. Extensive experiments with ablation analysis show the effectiveness of our method, and state-of-the-art results are achieved on several public datasets, including Market-1501, CUHK03, CUHK01, SenseReID, CUHK03-NP and DukeMTMC-reID.
Rui Zhao 0001, Feng Zhu 0006, Huaming Wang, Wanli Ouyang
CVPR4
2017 Cracking the cocktail party problem by multi-beam deep attractor network
abstract
While recent progresses in neural network approaches to singlechannel speech separation, or more generally the cocktail party problem, achieved significant improvement, their performance for complex mixtures is still not satisfactory. In this work, we propose a novel multi-channel framework for multi-talker separation. In the proposed model, an input multi-channel mixture signal is firstly converted to a set of beamformed signals using fixed beam patterns. For this beamforming, we propose to use differential beamformers as they are more suitable for speech separation. Then each beamformed signal is fed into a single-channel anchored deep attractor network to generate separated signals. And the final separation is acquired by post selecting the separating output for each beams. To evaluate the proposed system, we create a challenging dataset comprising mixtures of 2, 3 or 4 speakers. Our results show that the proposed system largely improves the state of the art in speech separation, achieving 11.5 dB, 11.76 dB and 11.02 dB average signal-to-distortion ratio improvement for 4, 3 and 2 overlapped speaker mixtures, which is comparable to the performance of a minimum variance distortionless response beamformer that uses oracle location, source, and noise information. We also run speech recognition with a clean trained acoustic model on the separated speech, achieving relative word error rate (WER) reduction of 45.76%, 59.40% and 62.80% on fully overlapped speech of 4, 3 and 2 speakers, respectively. With a far talk acoustic model, the WER is further reduced.
Zhuo Chen 0006, Jinyu Li 0001, Takuya Yoshioka, Huaming Wang, Yifan Gong 0001
ASRU5
2015 Large-Scale Question Answering with Joint Embedding and Proof Tree Decoding
abstract
Question answering (QA) over a large-scale knowledge base (KB) such as Freebase is an important natural language processing application. There are linguistically oriented semantic parsing techniques and machine learning motivated statistical methods. Both of these approaches face a key challenge on how to handle diverse ways natural questions can be expressed about predicates and entities in the KB. This paper is to investigate how to combine these two approaches. We frame the problem from a proof-theoretic perspective, and formulate it as a proof tree search problem that seamlessly unifies semantic parsing, logic reasoning, and answer ranking. We combine our word entity joint embedding learned from web-scale data with other surface-form features to further boost accuracy improvements. Our real-time system on the Freebase QA task achieved a very high F1 score (47.2) on the standard Stanford WebQuestions benchmark test data.
Shengquan Yan, Huaming Wang, Xuedong Huang 0001
CIKM3
2014 An introduction to computational networks and the computational network toolkit (invited talk)
Dong Yu 0001, Adam Eversole, Michael L. Seltzer, Kaisheng Yao, Brian Guenter, Oleksii Kuchaiev, Frank Seide, Huaming Wang, Jasha Droppo, Zhiheng Huang, Geoffrey Zweig, Christopher J. Rossbach, Jon Currey
INTERSPEECH8
1996 Performance analysis of a new electroplating line configuration with a linear induction motor based material mover
abstract
A high-performance linear induction motor (LIM) can be used to increase material transfer speeds by an order of magnitude over conventional equipment. In this paper, we develop queueing network models for performance analysis of LIM based material mover in a new configuration for electroplating lines.
Ramakrishna Desiraju, Huaming Wang, Jerry L. Sanders
IEEE Trans. Robotics Autom.2