EDBT 2026 Demo / reviewers in the wild / expert
Ho-Hsiang Wu
dblp:00/8324
· DBLP profile ↗
13ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-1102-074XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Few-Shot Training-Free Anomaly Sound Detection
Ho-Hsiang Wu, Abinaya Kumar, Luca Bondi, Shabnam Ghaffarzadegan, Juan Pablo Bello |
INTERSPEECH | 1 |
| 2024 | Multi-Modal Continual Pre-Training For Audio EncodersabstractSeveral approaches have been proposed to pre-train an audio encoder to learn fundamental audio knowledge. These training frameworks range from supervised learning to self-supervised learning with a contrastive objective under multi-modal supervision. However, these approaches are constrained to a single pretext task, preventing their adaptability to multi-modal interactions beyond the modalities provided in training data. Continual learning (CL), in the meantime, allows machine learning systems to incrementally learn a new task while preserving the previously acquired knowledge, making the system more knowledgeable over time. The existing CL approaches are limited to learning downstream tasks such as classification. In this work, we propose to combine CL methods with several audio encoder pre-training methods. The audio encoders, when pre-trained continually over a sequence of multi-modal tasks, namely audiovisual and audio-text, exhibit improved performance across various downstream tasks compared to their non-continual learning counterparts, due to knowledge accumulation. The audio encoders are also capable of performing cross-modal tasks of all learned modalities. Gyuhak Kim, Ho-Hsiang Wu, Luca Bondi, Bing Liu 0001 |
ICASSP | 2 |
| 2024 | CLAP4Emo: ChatGPT-Assisted Speech Emotion Retrieval with Natural Language SupervisionabstractSpeech emotion retrieval is an important technique for large-scale and high-quality data collection. Conventional approach using ensemble of classification models might limit the retrieved emotion diversity and/or underperform in out-of-domain acoustic conditions. Natural language is diverse and agnostic to specific acoustic concepts, embedding a huge potential for developing language-based speech emotion retrieval system. In this paper we introduce CLAP4Emo, a novel framework to retrieve emotional speech via natural language prompts based on contrastive language-audio pretraining. To compensate for the absence of training captions in existing public datasets, we propose a systematic framework that applies ChatGPT to generate emotion captions. The experimental results demonstrate that our method can effectively improve the retrieved sample diversity while maintaining high precision across five benchmark datasets. By leveraging large language models, we establish a connection between audio and language for emotion description, culminating in an intuitive and interactive retrieval system. We release the generated emotion captions at: https://github.com/boschresearch/soundsee-emo-caps Shabnam Ghaffarzadegan, Luca Bondi, Abinaya Kumar, Samarjit Das, Ho-Hsiang Wu |
ICASSP | 6 |
| 2024 | Learning Audio Concepts from Counterfactual Natural LanguageabstractConventional audio classification relied on predefined classes, lacking the ability to learn from free-form text. Recent methods unlock learning joint audio-text embeddings from raw audio-text pairs describing audio in natural language. Despite recent advancements, there is little exploration of systematic methods to train models for recognizing sound events and sources in alternative scenarios, such as distinguishing fireworks from gunshots at outdoor events in similar situations. This study introduces causal reasoning and counterfactual analysis in the audio domain. We use counterfactual instances and include them in our model across different aspects. Our model considers acoustic characteristics and sound source information from human-annotated reference texts. To validate the effectiveness of our model, we conducted pre-training utilizing multiple audio captioning datasets. We then evaluate with several common downstream tasks, demonstrating the merits of the proposed method as one of the first works leveraging counterfactual information in audio domain. Specifically, the top-1 accuracy in open-ended language-based audio retrieval task increased by more than 43%. Ali Vosoughi, Luca Bondi, Ho-Hsiang Wu, Chenliang Xu |
ICASSP | 3 |
| 2024 | MOSAIC: Learning Unified Multi-Sensory Object Property Representations for Robot Learning via Interactive PerceptionabstractA holistic understanding of object properties across diverse sensory modalities (e.g., visual, audio, and haptic) is essential for tasks ranging from object categorization to complex manipulation. Drawing inspiration from cognitive science studies that emphasize the significance of multi-sensory integration in human perception, we introduce MOSAIC (Multimodal Object property learning with Self-Attention and Interactive Comprehension), a novel framework designed to facilitate the learning of unified multi-sensory object property representations. While it is undeniable that visual information plays a prominent role, we acknowledge that many fundamental object properties extend beyond the visual domain to encompass attributes like texture, mass distribution, or sounds, which significantly influence how we interact with objects. In MOSAIC, we leverage this profound insight by distilling knowledge from multimodal foundation models and aligning these representations not only across vision but also haptic and auditory sensory modalities. Through extensive experiments on a dataset where a humanoid robot interacts with 100 objects across 10 exploratory behaviors, we demonstrate the versatility of MOSAIC in two task families: object categorization and object-fetching tasks. Our results underscore the efficacy of MOSAIC's unified representations, showing competitive performance in category recognition through a simple linear probe setup and excelling in the fetch object task under zero-shot transfer conditions. This work pioneers the application of sensory grounding in foundation models for robotics, promising a significant leap in multi-sensory perception capabilities for autonomous systems. We have released the code, datasets, and additional results: https://github.com/gtatiya/MOSAIC. Gyan Tatiya, Jonathan Francis, Ho-Hsiang Wu, Yonatan Bisk, Jivko Sinapov |
ICRA | 3 |
| 2024 | Sound of Traffic: A Dataset for Acoustic Traffic Identification and Counting
Shabnam Ghaffarzadegan, Luca Bondi, Abinaya Kumar, Ho-Hsiang Wu, Hans-Georg Horst, Samarjit Das |
INTERSPEECH | 5 |
| 2023 | Audio-Text Models Do Not Yet Leverage Natural LanguageabstractMulti-modal contrastive learning techniques in the audio-text domain have quickly become a highly active area of research. Most works are evaluated with standard audio retrieval and classification benchmarks assuming that (i) these models are capable of leveraging the rich information contained in natural language, and (ii) current benchmarks are able to capture the nuances of such information. In this work, we show that state-of-the-art audio-text models do not yet really understand natural language, especially contextual concepts such as sequential or concurrent ordering of sound events. Our results suggest that existing benchmarks are not sufficient to assess these models' capabilities to match complex contexts from the audio and text modalities. We propose a Transformer-based architecture and show that, unlike prior work, it is capable of modeling the sequential relationship between sound events in the text and audio, given appropriate benchmark data. We advocate for the collection or generation of additional, diverse, data to allow future research to fully leverage natural language for audio-text modeling. Ho-Hsiang Wu, Oriol Nieto, Juan Pablo Bello, Justin Salamon |
ICASSP | 1 |
| 2023 | Active Learning for Abnormal Lung Sound Data Curation and Detection in Asthma
Shabnam Ghaffarzadegan, Luca Bondi, Ho-Hsiang Wu, Sirajum Munir, Kelly J. Shields, Samarjit Das, Joseph Aracri |
INTERSPEECH | 3 |
| 2022 | Wav2CLIP: Learning Robust Audio Representations from ClipabstractWe propose Wav2CLIP, a robust audio representation learning method by distilling from Contrastive Language-Image Pre-training (CLIP). We systematically evaluate Wav2CLIP on a variety of audio tasks including classification, retrieval, and generation, and show that Wav2CLIP can outperform several publicly available pre-trained audio representation algorithms. Wav2CLIP projects audio into a shared embedding space with images and text, which enables multimodal applications such as zero-shot classification, and cross-modal retrieval. Furthermore, Wav2CLIP needs just ∼10% of the data to achieve competitive performance on downstream tasks compared with fully supervised models, and is more efficient to pre-train than competing methods as it does not require learning a visual model in concert with an auditory model. Finally, we demonstrate image generation from Wav2CLIP as qualitative assessment of the shared embedding space. Our code and model weights are open sourced and made available for further applications. Ho-Hsiang Wu, Prem Seetharaman, Juan Pablo Bello |
ICASSP | 1 |
| 2022 | How to Listen? Rethinking Visual Sound LocalizationabstractLocalizing visual sounds consists on locating the position of objects that emit sound within an image.It is a growing research area with potential applications in monitoring natural and urban environments, such as wildlife migration and urban traffic.Previous works are usually evaluated with datasets having mostly a single dominant visible object, and proposed models usually require the introduction of localization modules during training or dedicated sampling strategies, but it remains unclear how these design choices play a role in the adaptability of these methods in more challenging scenarios.In this work, we analyze various model choices for visual sound localization and discuss how their different components affect the model's performance, namely the encoders' architecture, the loss function and the localization strategy.Furthermore, we study the interaction between these decisions, the model performance, and the data, by digging into different evaluation datasets spanning different difficulties and characteristics, and discuss the implications of such decisions in the context of real-world applications.Our code and model weights are open-sourced and made available for further applications. Ho-Hsiang Wu, Magdalena Fuentes, Prem Seetharaman, Juan Pablo Bello |
INTERSPEECH | 1 |
| 2021 | Multi-Task Self-Supervised Pre-Training for Music ClassificationabstractDeep learning is very data hungry, and supervised learning especially requires massive labeled data to work well. Machine listening research often suffers from limited labeled data problem, as human annotations are costly to acquire, and annotations for audio are time consuming and less intuitive. Besides, models learned from labeled dataset often embed biases specific to that particular dataset. Therefore, unsupervised learning techniques become popular approaches in solving machine listening problems. Particularly, a self-supervised learning technique utilizing reconstructions of multiple hand-crafted audio features has shown promising results when it is applied to speech domain such as emotion recognition and automatic speech recognition (ASR). In this paper, we apply self-supervised and multi-task learning methods for pre-training music encoders, and explore various design choices including encoder architectures, weighting mechanisms to combine losses from multiple tasks, and worker selections of pretext tasks. We investigate how these design choices interact with various downstream music classification tasks. We find that using various music specific workers altogether with weighting mechanisms to balance the losses during pre-training helps improve and generalize to the downstream tasks. Ho-Hsiang Wu, Chieh-Chi Kao, Qingming Tang, Ming Sun 0007, Brian McFee, Juan Pablo Bello, Chao Wang 0018 |
ICASSP | 1 |
| 2019 | Look, Listen, and Learn More: Design Choices for Deep Audio EmbeddingsabstractA considerable challenge in applying deep learning to audio classification is the scarcity of labeled data. An increasingly popular solution is to learn deep audio embeddings from large audio collections and use them to train shallow classifiers using small labeled datasets. Look, Listen, and Learn (L3-Net) is an embedding trained through self-supervised learning of audio-visual correspondence in videos as opposed to other embeddings requiring labeled data. This framework has the potential to produce powerful out-of-the-box embeddings for downstream audio classification tasks, but has a number of unexplained design choices that may impact the embeddings’ behavior. In this paper we investigate how L3-Net design choices impact the performance of downstream audio classifiers trained with these embeddings. We show that audio-informed choices of input representation are important, and that using sufficient data for training the embedding is key. Surprisingly, we find that matching the content for training the embedding to the downstream task is not beneficial. Finally, we show that our best variant of the L3-Net embedding outperforms both the VGGish and SoundNet embeddings, while having fewer parameters and being trained on less data. Our implementation of the L3-Net embedding model as well as pre-trained models are made freely available online. Jason Cramer, Ho-Hsiang Wu, Justin Salamon, Juan Pablo Bello |
ICASSP | 2 |
| 2009 | Towards a Class-Based Representation of Perceptual Tempo for Music RetrievalabstractTempo is a common criterion by which humans describe and categorize music, and this has spawned a large amount of research in the field of automatic tempo estimation. Most tempo estimation systems focus mainly on detecting the temporal repetition and periodicity present within a signal, and represent tempo as a count of beats-per-minute (BPM). However, in real-world music retrieval applications such as music navigation and playlist generation, a rough perceptual representation of tempo may be more appropriate than a BPM representation. In this paper, the problem of tempo estimation is presented as a statistical classification problem. Four perceptual tempo classes are defined which correspond to rough semantic terms that average users may use to describe tempo. Statistical models of each class are built using low-level audio features. Experimental results show that the perceptual tempo class representation outperforms several conventional BPM-based tempo estimation systems when applied to the tasks of music navigation and playlist generation. Ching-Wei Chen, Kyogu Lee, Ho-Hsiang Wu |
ICMLA | 3 |