VLDB 2026 Research / reviewers in the wild / expert
Yuanbo Hou
dblp:225/4791
· DBLP profile ↗
19ranked-venue papers
15as first author
17since 2021 · last 2026
0000-0001-8469-5740ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 13 first-author · 13 since 2021Artificial intelligence and machine learning · 10 · 7 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Soundscape Captioning Using Sound Affective Quality Network and Large Language ModelabstractWe live in a rich and varied acoustic world, which is experienced by individuals or communities as asoundscape. Computational auditory scene analysis, disentangling acoustic scenes by detecting and classifying events, focuses on objective attributes of sounds, such as their category and temporal characteristics, ignoring their effects on people, such as the emotions they evoke within a context. To fill this gap, we propose the affective soundscape captioning (ASSC) task, which enables automated soundscape analysis, thus avoiding labour-intensive subjective ratings and surveys in conventional methods. With soundscape captioning, context-aware descriptions are generated for soundscape by capturing the acoustic scenes (ASs), audio events (AEs) information, and the corresponding human affective qualities (AQs). To this end, we propose an automatic soundscape captioner (SoundSCaper) system composed of an acoustic model, i.e. SoundAQnet, and a large language model (LLM). SoundAQnet simultaneously models multi-scale information about ASs, AEs, and perceived AQs, while the LLM describes the soundscape with captions by parsing the information captured with SoundAQnet. SoundSCaper is assessed by two juries of 32 people. In expert evaluation, the average score of SoundSCaper-generated captions is slightly lower than that of two soundscape experts on the evaluation set D1 and the external mixed dataset D2, but not statistically significant. In layperson evaluation, SoundSCaper outperforms soundscape experts in several metrics on datasets D1 and D2. In addition to human evaluation, compared to other automated audio captioning (AAC) systems with and without LLM, SoundSCaper performs better on the ASSC task in several natural language processing (NLP) based metrics. Overall, SoundSCaper performs well in human subjective evaluation and various objective captioning metrics, and the generated captions are comparable to those annotated by soundscape experts. The model, source code, LLM scripts, human assessment data, instructions, and evaluation statistics are all publicly available. Yuanbo Hou, Qiaoqiao Ren, Wenwu Wang 0001, Jian Kang 0002, Tony Belpaeme, Dick Botteldooren |
IEEE Trans. Multim. | 1 |
| 2025 | Sound-Based Recognition of Touch Gestures and Emotions for Enhanced Human-Robot InteractionabstractEmotion recognition and touch gesture decoding are crucial for advancing human-robot interaction (HRI), especially in social environments where emotional cues and tactile perception play important roles. However, many humanoid robots, such as Pepper, Nao, and Furhat, lack full-body tactile skin, limiting their ability to engage in touch-based emotional and gesture interactions. In addition, vision-based emotion recognition methods usually face strict GDPR compliance challenges due to the need to collect personal facial data. To address these limitations and avoid privacy issues, this paper studies the potential of using the sounds produced by touching during HRI to recognise tactile gestures and classify emotions along the arousal and valence dimensions. Using a dataset of tactile gestures and emotional interactions from 28 participants with the humanoid robot Pepper, we design an audio-only lightweight touch gesture and emotion recognition model with only 0.24M parameters, 0.94MB model size, and 0.7G FLOPs. Experimental results show that the proposed model effectively recognises the arousal and valence states of different emotions, as well as various tactile gestures, when the input audio length varies. The proposed model is of low-latency and achieves similar results as well-known pretrained audio neural networks (PANNs), but with much smaller FLOPs, number of parameters, and model size. Yuanbo Hou, Qiaoqiao Ren, Wenwu Wang 0001, Dick Botteldooren |
ICASSP | 1 |
| 2025 | SCAN: Selective Contrastive Learning Against Noisy Data for Acoustic Anomaly DetectionabstractAcoustic Anomaly Detection (AAD) has gained significant attention for the detection of suspicious activities or faults. Contrastive learning-based unsupervised AAD has outperformed traditional models on academic datasets, however, its model training is predominantly based on datasets containing only normal samples. In real industrial settings, a dataset of normal samples can still be corrupted by abnormal samples. Handling such noisy data is a crucial challenge, yet it remains largely unsolved. To address this issue, this paper proposes a Selective Contrastive learning framework Against Noisy data (SCAN) to mitigate the adverse effects of training the AAD model with anomaly-corrupted data. Specifically, SCAN progressively constructs confidence sample pairs based on the Mahalanobis distance, which is derived from the geometric median. These selected pairs are then integrated into the contrastive learning framework to enhance representation learning and model robustness. Extensive experiments under varying levels of label noise (i.e., the proportion of mislabeled abnormal samples in training data) demonstrate that SCAN outperforms state-of-the-art (SOTA) AAD methods on the real-world industrial datasets DCASE2022 and DCASE2024 Task2. Zhaoyi Liu 0003, Yuanbo Hou, Wenwu Wang 0001, Sam Michiels, Danny Hughes 0001 |
IEEE Signal Process. Lett. | 2 |
| 2024 | Boosting Adversarial Transferability across Model Genus by Deformation-Constrained WarpingabstractAdversarial examples generated by a surrogate model typically exhibit limited transferability to unknown target systems. To address this problem, many transferability enhancement approaches (e.g., input transformation and model augmentation) have been proposed. However, they show poor performances in attacking systems having different model genera from the surrogate model. In this paper, we propose a novel and generic attacking strategy, called Deformation-Constrained Warping Attack (DeCoWA), that can be effectively applied to cross model genus attack. Specifically, DeCoWA firstly augments input examples via an elastic deformation, namely Deformation-Constrained Warping (DeCoW), to obtain rich local details of the augmented input. To avoid severe distortion of global semantics led by random deformation, DeCoW further constrains the strength and direction of the warping transformation by a novel adaptive control strategy. Extensive experiments demonstrate that the transferable examples crafted by our DeCoWA on CNN surrogates can significantly hinder the performance of Transformers (and vice versa) on various tasks, including image classification, video action recognition, and audio recognition. Code is made available at https://github.com/LinQinLiang/DeCoWA. Qinliang Lin, Zenghao Niu, Xilin He, Weicheng Xie 0001, Yuanbo Hou, LinLin Shen, Siyang Song |
AAAI | 6 |
| 2024 | Multi-Level Graph Learning For Audio Event Classification And Human-Perceived Annoyance Rating PredictionabstractWHO’s report on environmental noise estimates that 22 M people suffer from chronic annoyance related to noise caused by audio events (AEs) from various sources. Annoyance may lead to health issues and adverse effects on metabolic and cognitive systems. In cities, monitoring noise levels does not provide insights into noticeable AEs, let alone their relations to annoyance. To create annoyance-related monitoring, this paper proposes a graph-based model to identify AEs in a sound-scape, and explore relations between diverse AEs and human-perceived annoyance rating (AR). Specifically, this paper proposes a lightweight multi-level graph learning (MLGL) based on local and global semantic graphs to simultaneously perform audio event classification (AEC) and human annoyance rating prediction (ARP). Experiments show that: 1) MLGL with 4.1 M parameters improves AEC and ARP results by using semantic node information in local and global context-aware graphs; 2) MLGL captures relations between coarse-and fine-grained AEs and AR well; 3) Statistical analysis of MLGL results shows that some AEs from different sources significantly correlate with AR, which is consistent with previous research on human perception of these sound sources. Yuanbo Hou, Qiaoqiao Ren, Siyang Song, Wenwu Wang 0001, Dick Botteldooren |
ICASSP | 1 |
| 2024 | Cooperative Scene-Event Modelling for Acoustic Scene ClassificationabstractAcoustic scene classification (ASC) can be helpful for creating context awareness for intelligent robots. Humans naturally use the relations between acoustic scenes (AS) and audio events (AE) to understand and recognize their surrounding environments. However, in most previous works, ASC and audio event classification (AEC) are treated as independent tasks, with a focus primarily on audio features shared between scenes and events, but not their implicit relations. To address this limitation, we propose a cooperative scene-event modelling (cSEM) framework to automatically model the intricate scene-event relation by an adaptive coupling matrix to improve ASC. Compared with other scene-event modelling frameworks, the proposed cSEM offers the following advantages. First, it reduces the confusion between similar scenes by aligning the information of coarsegrained AS and fine-grained AE in the latent space, and reducing the redundant information between the AS and AE embeddings. Second, it exploits the relation information between AS and AE to improve ASC, which is shown to be beneficial, even if the information of AE is derived from unverified pseudo-labels. Third, it uses a regression-based loss function for cooperative modelling of scene-event relations, which is shown to be more effective than classification-based loss functions. Instantiated from four models based on either Transformer or convolutional neural networks, cSEM is evaluated on real-life and synthetic datasets. Experiments show that cSEM-based models work well in reallife scene-event analysis, offering competitive results on ASC as compared with other multi-feature or multi-model ensemble methods.TheASCaccuracyachievedontheTUT2018,TAU2019, and JSSED datasets is 81.0%, 88.9% and 97.2%, respectively Yuanbo Hou, Bo Kang, Wenwu Wang 0001, Jian Kang 0002, Dick Botteldooren |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Gct: Gated Contextual Transformer for Sequential Audio TaggingabstractAudio tagging aims to assign predefined tags to audio clips to indicate the class information of audio events. Sequential audio tagging (SAT) means detecting both the class information of audio events, and the order in which they occur within the audio clip. Most existing methods for SAT are based on connectionist temporal classification (CTC). However, CTC cannot effectively capture event connections due to the conditional independence assumption between outputs at different times. The contextual Transformer (cTransformer) addresses this issue by exploiting contextual information in SAT. Nevertheless, cTransformer is also limited in exploiting contextual information as it only uses forward information in inference. This paper proposes a gated contextual Transformer (GCT) with forward-backward inference (FBI). In addition, a gated contextual multi-layer perceptron (GCMLP) block is proposed in GCT to improve the performance of cTransformer structurally. Experiments on the two real-life audio datasets with manually annotated sequential labels show that the proposed GCT with GCMLP and FBI performs better than the CTC-based methods and cTransformer. Yuanbo Hou, Wenwu Wang 0001, Dick Botteldooren |
ICASSP | 1 |
| 2023 | CLF-AIAD: A Contrastive Learning Framework for Acoustic Industrial Anomaly Detection
Zhaoyi Liu 0003, Yuanbo Hou, Álvaro López-Chilet, Sam Michiels, Dick Botteldooren, Jon Ander Gómez, Danny Hughes 0001 |
ICONIP (7) | 2 |
| 2023 | Joint Prediction of Audio Event and Annoyance Rating in an Urban Soundscape by Hierarchical Graph Representation LearningabstractSound events in daily life carry rich information about the objective world. The composition of these sounds affects the mood of people in a soundscape. Most previous approaches only focus on classifying and detecting audio events and scenes, but may ignore their perceptual quality that may impact humans' listening mood for the environment, e.g. annoyance. To this end, this paper proposes a novel hierarchical graph representation learning (HGRL) approach which links objective audio events (AE) with subjective annoyance ratings (AR) of the soundscape perceived by humans. The hierarchical graph consists of fine-grained event (fAE) embeddings with single-class event semantics, coarse-grained event (cAE) embeddings with multi-class event semantics, and AR embeddings. Experiments show the proposed HGRL successfully integrates AE with AR for AEC and ARP tasks, while coordinating the relations between cAE and fAE and further aligning the two different grains of AE information with the AR. Yuanbo Hou, Siyang Song, Qiaoqiao Ren, Weicheng Xie 0001, Jian Kang 0002, Wenwu Wang 0001, Dick Botteldooren |
INTERSPEECH | 1 |
| 2023 | Audio Event-Relational Graph Representation Learning for Acoustic Scene ClassificationabstractMost deep learning-based acoustic scene classification (ASC) approaches identify scenes based on acoustic features converted from audio clips containing mixed information entangled by polyphonic audio events (AEs). However, these approaches have difficulties in explaining what cues they use to identify scenes. This paper conducts the first study on disclosing the relationship between real-life acoustic scenes and semantic embeddings from the most relevant AEs. Specifically, we propose an event-relational graph representation learning (ERGL) framework for ASC to classify scenes, and simultaneously answer clearly and straightly which cues are used in classifying. In the event-relational graph, embeddings of each event are treated as nodes, while relationship cues derived from each pair of nodes are described by multi-dimensional edge features. Experiments on a real-life ASC dataset show that the proposed ERGL achieves competitive performance on ASC by learning embeddings of only a limited number of AEs. The results show the feasibility of recognizing diverse acoustic scenes based on the audio event-relational graph. Visualizations of graph representations learned by ERGL are available here(https://github.com/Yuanbo2020/ERGL). Yuanbo Hou, Siyang Song, Chuang Yu 0001, Wenwu Wang 0001, Dick Botteldooren |
IEEE Signal Process. Lett. | 1 |
| 2022 | Sensor characterization for supporting de-noising and auto-calibration of sensor dataabstractIn earlier work, a mobile sensor network had been deployed in passenger cars to measure road quality using sound and vibration sensors. A self-calibration and confounder removal (SCCR) algorithm eliminated the effect of the measurement device in sound and vibration measurements. It relied on the hypothesis that two measurements of the same object should give the same result. However, once the deep learning model has been trained, the model cannot easily include more devices, because it relies on a one-hot encoded device identification. Therefore, additional characterization of both sensor and vehicle is proposed here using a batch of observations (a bag) from the same sensors. A set-encoder is trained with a device classification head. This set-encoder compresses the input bag into a latent representation, irrespective of the input order. This latent vector, which characterizes the sensor, is injected into the SCCR framework to steer the calibration. New devices are characterized by the pre-trained set-encoder. This relies on the hypothesis that closely related sensors with a similar response should be close to each other in the latent space. Therefore, SCCR is now possible with new devices, previously unseen at the training time of SCCR. However, an increase in reconstruction error is observed for the unseen devices. Wout Van Hauwermeiren, Yuanbo Hou, Karlo Filipan, Dick Botteldooren |
IJCNN | 2 |
| 2022 | Relation-guided acoustic scene classification aided with event embeddingsabstractIn real life, acoustic scenes and audio events are naturally correlated. Humans instinctively rely on fine-grained audio events as well as the overall sound characteristics to distinguish diverse acoustic scenes. Yet, most previous approaches treat acoustic scene classification (ASC) and audio event classification (AEC) as two independent tasks. A few studies on scene and event joint classification either use synthetic audio datasets that hardly match the real world, or simply use the multi-task framework to perform two tasks at the same time. Neither of these two ways makes full use of the implicit and inherent relation between fine-grained events and coarse-grained scenes. To this end, this paper proposes a relation-guided ASC (RGASC) model to further exploit and coordinate the scene-event relation for the mutual benefit of scene and event recognition. The TUT Urban Acoustic Scenes 2018 dataset (TUT2018) is annotated with pseudo labels of events by a simple and efficient audiorelated pre-trained model PANN, which is one of the state-of-the-art AEC models. Then, a prior scene-event relation matrix is defined as the average probability of the presence of each event type in each scene class. Finally, the two-tower RGASC model is jointly trained on the real-life dataset TUT2018 for both scene and event classification. The following results are achieved. 1) RGASC effectively coordinates the true information of coarsegrained scenes and the pseudo information of fine-grained events. 2) The event embeddings learned from pseudo labels under the guidance of prior scene-event relations help reduce the confusion between similar acoustic scenes. 3) Compared with other (non-ensemble) methods, RGASC improves the scene classification accuracy on the real-life dataset. Yuanbo Hou, Bo Kang, Wout Van Hauwermeiren, Dick Botteldooren |
IJCNN | 1 |
| 2022 | Event-related data conditioning for acoustic event classificationabstractModels based on diverse attention mechanisms have recently shined in tasks related to acoustic event classification (AEC). Among them, self-attention is often used in audio-only tasks to help the model recognize different acoustic events. Self-attention relies on the similarity between time frames, and uses global information from the whole segment to highlight specific features within a frame. In real life, information related to acoustic events will attenuate over time, which means the information within some frames around the event deserves more attention than distant time global information that may be unrelated to the event. This paper shows that self-attention may over-enhance certain segments of audio representations, and smooth out the boundaries between events representations and background noises. Hence, this paper proposes an event-related data conditioning (EDC) for AEC. EDC directly works on spectrograms. The idea of EDC is to adaptively select the frame-related attention range based on acoustic features, and gather the event-related local information to represent the frame. Experiments show that: 1) compared with spectrogram-based data augmentation methods and trainable feature weighting and self-attention, EDC outperforms them in both the original-size mode and the augmented mode; 2) EDC effectively gathers event-related local information and enhances boundaries between events and backgrounds, improving the performance of AEC. Yuanbo Hou, Dick Botteldooren |
INTERSPEECH | 1 |
| 2022 | CT-SAT: Contextual Transformer for Sequential Audio TaggingabstractSequential audio event tagging can provide not only the type information of audio events, but also the order information between events and the number of events that occur in an audio clip. Most previous works on audio event sequence analysis rely on connectionist temporal classification (CTC). However, CTC's conditional independence assumption prevents it from effectively learning correlations between diverse audio events. This paper first introduces the Transformer into sequential audio tagging, since Transformers perform well in sequence-related tasks. To better utilize contextual information of audio event sequences, we draw on the idea of bidirectional recurrent neural networks, and propose a contextual Transformer (cTransformer) with a bidirectional decoder that could exploit the forward and backward information of event sequences. Experiments on the real-life polyphonic audio dataset show that, compared to CTC-based methods, the cTransformer can effectively combine the fine-grained acoustic representations from the encoder and coarse-grained audio event cues to exploit contextual information to successfully recognize and predict the audio event sequence in polyphonic audio clips. Yuanbo Hou, Zhaoyi Liu 0003, Bo Kang, Dick Botteldooren |
INTERSPEECH | 1 |
| 2022 | Audio-visual scene classification via contrastive event-object alignment and semantic-based fusionabstractPrevious works on scene classification are mainly based on audio or visual signals, while humans perceive the environmental scenes through multiple senses. Recent studies on audio-visual scene classification separately fine-tune the large-scale audio and image pre-trained models on the target dataset, then either fuse the intermediate representations of the audio model and the visual model, or fuse the coarse-grained decision of both models at the clip level. Such methods ignore the detailed audio events and visual objects in audio-visual scenes (AVS), while humans often identify a scene through both audio events and visual objects within, and the congruence between them. To exploit the fine-grained information of audio events and visual objects in AVS, and coordinate the implicit relationship between audio events and visual objects, this paper proposes a multi-branch model equipped with contrastive event-object alignment (CEOA) and semantic-based fusion (SF) for AVSC. CEOA aims to align the learned embeddings of audio events and visual objects by comparing the difference between audio-visual event-object pairs. Then, visual objects associated with certain audio events and vice versa are accentuated by cross-attention and undergo SF for semantic-level fusion. Experiments show that: 1) the proposed AVSC model equipped with CEOA and SF outperforms the results of audio-only and visual-only models, i.e., the audio-visual results are better than the results from a single modality. 2) CEOA aligns the embeddings of audio events and related visual objects on a fine-grained level, and the SF effectively integrates both; 3) Compared with other large-scale integrated systems, the proposed model shows competitive performance, even without using additional datasets and data augmentation tricks. Yuanbo Hou, Bo Kang, Dick Botteldooren |
MMSP | 1 |
| 2021 | Rule-Embedded Network for Audio-Visual Voice Activity Detection in Live Musical Video StreamsabstractDetecting anchor’s voice in live musical streams is an important preprocessing step for music and speech signal processing. Existing approaches to voice activity detection (VAD) primarily rely on audio, however, audio-based VAD is difficult to effectively focus on the target voice in noisy environments. This paper proposes a rule-embedded network to fuse the audio-visual (A-V) inputs for better detection of the target voice. The core role of the rule in the model is to coordinate the relation between the bi-modal information and use visual representations as a mask to filter out the information of non-target sound. Experiments show that: 1) with the help of cross-modal fusion using the proposed rule, the detection results of the A-V branch outperform that of the audio branch in the same model framework; 2) the performance of the bimodal A-V model far outperforms that of audio-only models, indicating that the incorporation of both audio and visual signals is highly beneficial for VAD. To attract more attention to the cross-modal music and audio signal processing, a new live musical video corpus with frame-level labels is introduced. Yuanbo Hou, Bilei Zhu, Zejun Ma 0001, Dick Botteldooren |
ICASSP | 1 |
| 2021 | Attention-Based Cross-Modal Fusion for Audio-Visual Voice Activity Detection in Musical Video StreamsabstractMany previous audio-visual voice-related works focus on speech, ignoring the singing voice in the growing number of musical video streams on the Internet. For processing diverse musical video data, voice activity detection is a necessary step. This paper attempts to detect the speech and singing voices of target performers in musical video streams using audio-visual information. To integrate information of audio and visual modalities, a multi-branch network is proposed to learn audio and image representations, and the representations are fused by attention based on semantic similarity to shape the acoustic representations through the probability of anchor vocalization. Experiments show the proposed audio-visual multi-branch network far outperforms the audio-only model in challenging acoustic environments, indicating the cross-modal information fusion based on semantic correlation is sensible and successful. Yuanbo Hou, Zhesong Yu, Xia Liang, Xingjian Du, Bilei Zhu, Zejun Ma 0001, Dick Botteldooren |
Interspeech | 1 |
| 2020 | Transfer Learning for Improving Singing-Voice Detection in Polyphonic Instrumental MusicabstractDetecting singing-voice in polyphonic instrumental music is critical to music information retrieval.To train a robust vocal detector, a large dataset marked with vocal or non-vocal label at frame-level is essential.However, frame-level labeling is time-consuming and labor expensive, resulting there is little well-labeled dataset available for singing-voice detection (S-VD).Hence, we propose a data augmentation method for S-VD by transfer learning.In this study, clean speech clips with voice activity endpoints and separate instrumental music clips are artificially added together to simulate polyphonic vocals to train a vocal /non-vocal detector.Due to the different articulation and phonation between speaking and singing, the vocal detector trained with the artificial dataset does not match well with the polyphonic music which is singing vocals together with the instrumental accompaniments.To reduce this mismatch, transfer learning is used to transfer the knowledge learned from the artificial speech-plus-music training set to a small but matched polyphonic dataset, i.e., singing vocals with accompaniments.By transferring the related knowledge to make up for the lack of well-labeled training data in S-VD, the proposed data augmentation method by transfer learning can improve S-VD performance with an F-score improvement from 89.5% to 93.2%. Yuanbo Hou, Frank K. Soong, Jian Luan 0001, Shengchen Li |
INTERSPEECH | 1 |
| 2019 | Sound Event Detection with Sequentially Labelled Data Based on Connectionist Temporal Classification and Unsupervised ClusteringabstractSound event detection (SED) methods typically rely on either strongly labelled data or weakly labelled data. As an alternative, sequentially labelled data (SLD) was proposed. In SLD, the events and the order of events in audio clips are known, without knowing the occurrence time of events. This paper proposes a connectionist temporal classification (CTC) based SED system that uses SLD instead of strongly labelled data, with a novel unsupervised clustering stage. Experiments on 41 classes of sound events show that the proposed two-stage method trained on SLD achieves performance comparable to the previous state-of-the-art SED system trained on strongly labelled data, and is far better than another state-of-the-art SED system trained on weakly labelled data, which indicates the effectiveness of the proposed two-stage method trained on SLD without any onset/offset time of sound events. Yuanbo Hou, Qiuqiang Kong, Shengchen Li, Mark D. Plumbley |
ICASSP | 1 |