EDBT 2026 Demo / reviewers in the wild / expert
Dick Botteldooren
dblp:84/6768
· DBLP profile ↗
42ranked-venue papers
3as first author
29since 2021 · last 2026
0000-0002-7756-7238ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 3 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Soundscape Captioning Using Sound Affective Quality Network and Large Language ModelabstractWe live in a rich and varied acoustic world, which is experienced by individuals or communities as asoundscape. Computational auditory scene analysis, disentangling acoustic scenes by detecting and classifying events, focuses on objective attributes of sounds, such as their category and temporal characteristics, ignoring their effects on people, such as the emotions they evoke within a context. To fill this gap, we propose the affective soundscape captioning (ASSC) task, which enables automated soundscape analysis, thus avoiding labour-intensive subjective ratings and surveys in conventional methods. With soundscape captioning, context-aware descriptions are generated for soundscape by capturing the acoustic scenes (ASs), audio events (AEs) information, and the corresponding human affective qualities (AQs). To this end, we propose an automatic soundscape captioner (SoundSCaper) system composed of an acoustic model, i.e. SoundAQnet, and a large language model (LLM). SoundAQnet simultaneously models multi-scale information about ASs, AEs, and perceived AQs, while the LLM describes the soundscape with captions by parsing the information captured with SoundAQnet. SoundSCaper is assessed by two juries of 32 people. In expert evaluation, the average score of SoundSCaper-generated captions is slightly lower than that of two soundscape experts on the evaluation set D1 and the external mixed dataset D2, but not statistically significant. In layperson evaluation, SoundSCaper outperforms soundscape experts in several metrics on datasets D1 and D2. In addition to human evaluation, compared to other automated audio captioning (AAC) systems with and without LLM, SoundSCaper performs better on the ASSC task in several natural language processing (NLP) based metrics. Overall, SoundSCaper performs well in human subjective evaluation and various objective captioning metrics, and the generated captions are comparable to those annotated by soundscape experts. The model, source code, LLM scripts, human assessment data, instructions, and evaluation statistics are all publicly available. Yuanbo Hou, Qiaoqiao Ren, Wenwu Wang 0001, Jian Kang 0002, Tony Belpaeme, Dick Botteldooren |
IEEE Trans. Multim. | 7 |
| 2025 | Sound-Based Recognition of Touch Gestures and Emotions for Enhanced Human-Robot InteractionabstractEmotion recognition and touch gesture decoding are crucial for advancing human-robot interaction (HRI), especially in social environments where emotional cues and tactile perception play important roles. However, many humanoid robots, such as Pepper, Nao, and Furhat, lack full-body tactile skin, limiting their ability to engage in touch-based emotional and gesture interactions. In addition, vision-based emotion recognition methods usually face strict GDPR compliance challenges due to the need to collect personal facial data. To address these limitations and avoid privacy issues, this paper studies the potential of using the sounds produced by touching during HRI to recognise tactile gestures and classify emotions along the arousal and valence dimensions. Using a dataset of tactile gestures and emotional interactions from 28 participants with the humanoid robot Pepper, we design an audio-only lightweight touch gesture and emotion recognition model with only 0.24M parameters, 0.94MB model size, and 0.7G FLOPs. Experimental results show that the proposed model effectively recognises the arousal and valence states of different emotions, as well as various tactile gestures, when the input audio length varies. The proposed model is of low-latency and achieves similar results as well-known pretrained audio neural networks (PANNs), but with much smaller FLOPs, number of parameters, and model size. Yuanbo Hou, Qiaoqiao Ren, Wenwu Wang 0001, Dick Botteldooren |
ICASSP | 4 |
| 2025 | A General Closed-loop Predictive Coding Framework for Auditory Working MemoryabstractAuditory working memory is essential for various daily activities, such as language acquisition and conversation. It involves the temporary storage and manipulation of information that is no longer present in the environment. While extensively studied in neuroscience and cognitive science, research on its modeling within neural networks remains limited. To address this gap, we propose a general framework based on a closed-loop predictive coding paradigm to perform short auditory signal memory tasks. The framework is evaluated on two widely used benchmark datasets for environmental sound and speech, demonstrating high semantic similarity across both datasets. Zhongju Yuan, Geraint A. Wiggins, Dick Botteldooren |
IJCNN | 3 |
| 2025 | BioOSS: A Bio-Inspired Oscillatory State System with Spatio-Temporal DynamicsabstractToday’s deep learning architectures are primarily based on perceptron models, which do not capture the oscillatory dynamics characteristic of biological neurons. Although oscillatory systems have recently gained attention for their closer resemblance to neural behavior, they still fall short of modeling the intricate spatio-temporal interactions observed in natural neural circuits. In this paper, we propose a bio-inspired oscillatory state system (BioOSS) designed to emulate the wave-like propagation dynamics critical to neural processing, particularly in the prefrontal cortex (PFC), where complex activity patterns emerge. BioOSS comprises two interacting populations of neurons: p neurons, which represent simplified membrane-potential-like units inspired by pyramidal cells in cortical columns, and o neurons, which govern propagation velocities and modulate the lateral spread of activity. Through local interactions, these neurons produce wave-like propagation patterns. The model incorporates trainable parameters for damping and propagation speed, enabling flexible adaptation to task-specific spatio-temporal structures. We evaluate BioOSS on both synthetic and real-world tasks, demonstrating superior performance and enhanced interpretability compared to alternative architectures. Zhongju Yuan, Geraint A. Wiggins, Dick Botteldooren |
NeurIPS | 3 |
| 2025 | Electroencephalography Decoding with Conditional Identification GeneratorabstractDecoding Electroencephalography (EEG) signals are extremely useful for advancing and understanding human-artificial intelligence (AI) interaction systems. Recent advancements in deep neural networks (DNNs) have demonstrated significant promise in this respect due to their ability to model complex nonlinear relationships. However, DNNs face persistent challenges in addressing the inter-person variability inherent in EEG signals, which limits their generalizability. To tackle this limitation, we propose a novel framework that integrates conditional identification information, leveraging the interaction between EEG signals and individual traits to enhance the model's internal representation and improve decoding accuracy. Building on this foundation, we further introduce a privacy-preserving conditional information generator - a generative model that derives embedding knowledge directly from raw EEG signals. This approach eliminates the need for personal identification via individual tests, ensuring both efficiency and privacy. Experimental evaluations conducted on WithMe dataset confirm that this framework outperforms baseline network architectures. Notably, our approach achieves substantial improvements in decoding accuracy for both familiar and unseen subjects, paving the way for efficient, robust, and privacy-conscious human-computer interface systems. Pengfei Sun 0003, Jorg De Winne, Malu Zhang, Paul Devos, Dick Botteldooren |
Int. J. Neural Syst. | 5 |
| 2025 | Towards parameter-free attentional spiking neural networks
Pengfei Sun 0003, Jibin Wu, Paul Devos, Dick Botteldooren |
Neural Networks | 4 |
| 2025 | Delayed knowledge transfer: Cross-modal knowledge transfer from delayed stimulus to EEG for continuous attention detection based on spike-represented EEG signals
Pengfei Sun 0003, Jorg De Winne, Malu Zhang, Paul Devos, Dick Botteldooren |
Neural Networks | 5 |
| 2025 | A Dynamic Systems Approach to Modeling Human-Machine Rhythm InteractionabstractRhythm is an inherent aspect of human behavior, present from infancy and embedded in cultural practices. At the core of rhythm perception lies meter anticipation, a spontaneous process in the human brain that typically occurs before actual beats. This anticipation can be framed as a time series prediction problem. From the perspective of human embodied system behavior, although many models have been developed for time series prediction, most prioritize accuracy over biological realism, contrasting with the natural imprecision of human internal clocks. Neuroscientific evidence, such as infants' natural meter synchronization, underscores the need for biologically plausible models. Therefore, we propose a neuron oscillator-based dynamic system that simulates human behavior during meter perception. The model introduces two tunable parameters for local and global adjustments, fine-tuning the oscillation combinations to emulate human-like rhythmic behavior. The experiments are conducted under three common scenarios encountered during human-machine interaction, demonstrating that the proposed model can exhibit human-like reactions. Additionally, experiments involving human-machine and interhuman interactions show that the model successfully replicates real-world rhythmic behavior, advancing toward more natural and synchronized human-machine rhythm interaction. Zhongju Yuan, Wannes Van Ransbeeck, Geraint A. Wiggins, Dick Botteldooren |
IEEE Trans. Cybern. | 4 |
| 2025 | Delayed Memory Unit: Modeling Temporal Dependency Through Delay GateabstractRecurrent neural networks (RNNs) are widely recognized for their proficiency in modeling temporal dependencies, making them highly prevalent in sequential data processing applications. Nevertheless, vanilla RNNs are confronted with the well-known issue of gradient vanishing and exploding, posing a significant challenge for learning and establishing long-range dependencies. Additionally, gated RNNs tend to be over-parameterized, resulting in poor computational efficiency and network generalization. To address these challenges, this article proposes a novel delayed memory unit (DMU). The DMU incorporates a delay line structure along with delay gates into vanilla RNN, thereby enhancing temporal interaction and facilitating temporal credit assignment. Specifically, the DMU is designed to directly distribute the input information to the optimal time instant in the future, rather than aggregating and redistributing it over time through intricate network dynamics. Our proposed DMU demonstrates superior temporal modeling capabilities across a broad range of sequential modeling tasks, utilizing considerably fewer parameters than other state-of-the-art gated RNN models in applications such as speech recognition, radar gesture recognition, ECG waveform segmentation, and permuted sequential (PS) image classification. Pengfei Sun 0003, Jibin Wu, Malu Zhang, Paul Devos, Dick Botteldooren |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Multi-Level Graph Learning For Audio Event Classification And Human-Perceived Annoyance Rating PredictionabstractWHO’s report on environmental noise estimates that 22 M people suffer from chronic annoyance related to noise caused by audio events (AEs) from various sources. Annoyance may lead to health issues and adverse effects on metabolic and cognitive systems. In cities, monitoring noise levels does not provide insights into noticeable AEs, let alone their relations to annoyance. To create annoyance-related monitoring, this paper proposes a graph-based model to identify AEs in a sound-scape, and explore relations between diverse AEs and human-perceived annoyance rating (AR). Specifically, this paper proposes a lightweight multi-level graph learning (MLGL) based on local and global semantic graphs to simultaneously perform audio event classification (AEC) and human annoyance rating prediction (ARP). Experiments show that: 1) MLGL with 4.1 M parameters improves AEC and ARP results by using semantic node information in local and global context-aware graphs; 2) MLGL captures relations between coarse-and fine-grained AEs and AR well; 3) Statistical analysis of MLGL results shows that some AEs from different sources significantly correlate with AR, which is consistent with previous research on human perception of these sound sources. Yuanbo Hou, Qiaoqiao Ren, Siyang Song, Wenwu Wang 0001, Dick Botteldooren |
ICASSP | 6 |
| 2024 | A novel Reservoir Architecture for Periodic Time Series PredictionabstractThis paper introduces a novel approach to predicting periodic time series using reservoir computing. The model is tailored to deliver precise forecasts of rhythms, a crucial aspect for tasks such as generating musical rhythm. Leveraging reservoir computing, our proposed method is ultimately oriented towards predicting human perception of rhythm. Our network accurately predicts rhythmic signals within the human frequency perception range. The model architecture incorporates primary and intermediate neurons tasked with capturing and transmitting rhythmic information. Two parameter matrices, denoted as c and k, regulate the reservoir’s overall dynamics. We propose a loss function to adapt c post-training and introduce a dynamic selection (DS) mechanism that adjusts k to focus on areas with outstanding contributions. Experimental results on a diverse test set showcase accurate predictions, further improved through real-time tuning of the reservoir via c and k. Comparative assessments highlight its superior performance compared to conventional models. Zhongju Yuan, Geraint A. Wiggins, Dick Botteldooren |
IJCNN | 3 |
| 2024 | Cell-Stitching for Analog Neuromorphic ComputingabstractNeuromorphic computing is an innovative paradigm aiming to unify storage and computation, thereby addressing the constraints imposed by the Von Neumann bottleneck. Within this domain, analog computing based on memristors as synaptic elements has emerged as a promising avenue, though achieving consistent accuracy has proven to be challenging, due to limited precision of memristors. One promising strategy involves harnessing multiple memristor cells to improve synaptic precision, yet a mere concatenation approach as published in the literature falls short of meeting this goal, as inherent variations in the memristor writing process would render errors of higher order bit cells to overshadow lower order bits cells. In response to this challenge, we present a novel’ cell splicing’ methodology designed to enhance accuracy with analog computation. Experimental simulations using the ImageNet dataset demonstrate that it achieves accuracy similar to the baseline, and markedly outperforms the rudimentary concatenation approach. Pengfei Sun 0003, Piew Yoong Chee, Dick Botteldooren |
TENCON | 4 |
| 2024 | Delay learning based on temporal coding in Spiking Neural Networks
Pengfei Sun 0003, Jibin Wu, Malu Zhang, Paul Devos, Dick Botteldooren |
Neural Networks | 5 |
| 2024 | Cooperative Scene-Event Modelling for Acoustic Scene ClassificationabstractAcoustic scene classification (ASC) can be helpful for creating context awareness for intelligent robots. Humans naturally use the relations between acoustic scenes (AS) and audio events (AE) to understand and recognize their surrounding environments. However, in most previous works, ASC and audio event classification (AEC) are treated as independent tasks, with a focus primarily on audio features shared between scenes and events, but not their implicit relations. To address this limitation, we propose a cooperative scene-event modelling (cSEM) framework to automatically model the intricate scene-event relation by an adaptive coupling matrix to improve ASC. Compared with other scene-event modelling frameworks, the proposed cSEM offers the following advantages. First, it reduces the confusion between similar scenes by aligning the information of coarsegrained AS and fine-grained AE in the latent space, and reducing the redundant information between the AS and AE embeddings. Second, it exploits the relation information between AS and AE to improve ASC, which is shown to be beneficial, even if the information of AE is derived from unverified pseudo-labels. Third, it uses a regression-based loss function for cooperative modelling of scene-event relations, which is shown to be more effective than classification-based loss functions. Instantiated from four models based on either Transformer or convolutional neural networks, cSEM is evaluated on real-life and synthetic datasets. Experiments show that cSEM-based models work well in reallife scene-event analysis, offering competitive results on ASC as compared with other multi-feature or multi-model ensemble methods.TheASCaccuracyachievedontheTUT2018,TAU2019, and JSSED datasets is 81.0%, 88.9% and 97.2%, respectively Yuanbo Hou, Bo Kang, Wenwu Wang 0001, Jian Kang 0002, Dick Botteldooren |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | Gct: Gated Contextual Transformer for Sequential Audio TaggingabstractAudio tagging aims to assign predefined tags to audio clips to indicate the class information of audio events. Sequential audio tagging (SAT) means detecting both the class information of audio events, and the order in which they occur within the audio clip. Most existing methods for SAT are based on connectionist temporal classification (CTC). However, CTC cannot effectively capture event connections due to the conditional independence assumption between outputs at different times. The contextual Transformer (cTransformer) addresses this issue by exploiting contextual information in SAT. Nevertheless, cTransformer is also limited in exploiting contextual information as it only uses forward information in inference. This paper proposes a gated contextual Transformer (GCT) with forward-backward inference (FBI). In addition, a gated contextual multi-layer perceptron (GCMLP) block is proposed in GCT to improve the performance of cTransformer structurally. Experiments on the two real-life audio datasets with manually annotated sequential labels show that the proposed GCT with GCMLP and FBI performs better than the CTC-based methods and cTransformer. Yuanbo Hou, Wenwu Wang 0001, Dick Botteldooren |
ICASSP | 4 |
| 2023 | Adaptive Axonal Delays in Feedforward Spiking Neural Networks for Accurate Spoken Word RecognitionabstractSpiking neural networks (SNN) are a promising research avenue for building accurate and efficient automatic speech recognition systems. Recent advances in audio-to-spike encoding and training algorithms enable SNN to be applied in practical tasks. Biologically-inspired SNN communicates using sparse asynchronous events. Therefore, spike-timing is critical to SNN performance. In this aspect, most works focus on training synaptic weights and few have considered delays in event transmission, namely axonal delay. In this work, we consider a learnable axonal delay capped at a maximum value, which can be adapted according to the axonal delay distribution in each network layer. We show that our proposed method achieves the best classification results reported on the SHD dataset (92.45%) and NTIDIGITS dataset (95.09%). Our work illustrates the potential of training axonal delays for tasks with complex temporal structures. Pengfei Sun 0003, Ehsan Eqlimi, Yansong Chua, Paul Devos, Dick Botteldooren |
ICASSP | 5 |
| 2023 | CLF-AIAD: A Contrastive Learning Framework for Acoustic Industrial Anomaly Detection
Zhaoyi Liu 0003, Yuanbo Hou, Álvaro López-Chilet, Sam Michiels, Dick Botteldooren, Jon Ander Gómez, Danny Hughes 0001 |
ICONIP (7) | 6 |
| 2023 | Joint Prediction of Audio Event and Annoyance Rating in an Urban Soundscape by Hierarchical Graph Representation LearningabstractSound events in daily life carry rich information about the objective world. The composition of these sounds affects the mood of people in a soundscape. Most previous approaches only focus on classifying and detecting audio events and scenes, but may ignore their perceptual quality that may impact humans' listening mood for the environment, e.g. annoyance. To this end, this paper proposes a novel hierarchical graph representation learning (HGRL) approach which links objective audio events (AE) with subjective annoyance ratings (AR) of the soundscape perceived by humans. The hierarchical graph consists of fine-grained event (fAE) embeddings with single-class event semantics, coarse-grained event (cAE) embeddings with multi-class event semantics, and AR embeddings. Experiments show the proposed HGRL successfully integrates AE with AR for AEC and ARP tasks, while coordinating the relations between cAE and fAE and further aligning the two different grains of AE information with the AR. Yuanbo Hou, Siyang Song, Qiaoqiao Ren, Weicheng Xie 0001, Jian Kang 0002, Wenwu Wang 0001, Dick Botteldooren |
INTERSPEECH | 9 |
| 2023 | Personalising augmented soundscapes for supporting persons with dementia
Toon De Pessemier, Kris Vanhecke, Pieter Thomas, Tara Vander Mynsbrugge, Stefaan Vercoutere, Dominique Van de Velde, Patricia De Vriendt, Wout Joseph, Luc Martens, Dick Botteldooren, Paul Devos |
Multim. Tools Appl. | 10 |
| 2023 | Audio Event-Relational Graph Representation Learning for Acoustic Scene ClassificationabstractMost deep learning-based acoustic scene classification (ASC) approaches identify scenes based on acoustic features converted from audio clips containing mixed information entangled by polyphonic audio events (AEs). However, these approaches have difficulties in explaining what cues they use to identify scenes. This paper conducts the first study on disclosing the relationship between real-life acoustic scenes and semantic embeddings from the most relevant AEs. Specifically, we propose an event-relational graph representation learning (ERGL) framework for ASC to classify scenes, and simultaneously answer clearly and straightly which cues are used in classifying. In the event-relational graph, embeddings of each event are treated as nodes, while relationship cues derived from each pair of nodes are described by multi-dimensional edge features. Experiments on a real-life ASC dataset show that the proposed ERGL achieves competitive performance on ASC by learning embeddings of only a limited number of AEs. The results show the feasibility of recognizing diverse acoustic scenes based on the audio event-relational graph. Visualizations of graph representations learned by ERGL are available here(https://github.com/Yuanbo2020/ERGL). Yuanbo Hou, Siyang Song, Chuang Yu 0001, Wenwu Wang 0001, Dick Botteldooren |
IEEE Signal Process. Lett. | 5 |
| 2022 | Axonal Delay as a Short-Term Memory for Feed Forward Deep Spiking Neural NetworksabstractThe information of spiking neural networks (SNNs) are propagated between the adjacent biological neuron by spikes, which provides a computing paradigm with the promise of simulating the human brain. Recent studies have found that the time delay of neurons plays an important role in the learning process. Therefore, configuring the precise timing of the spike is a promising direction for understanding and improving the transmission process of temporal information in SNNs. However, most of the existing learning methods for spiking neurons are focusing on the adjustment of synaptic weight, while very few research has been working on axonal delay. In this paper, we verify the effectiveness of integrating time delay into supervised learning and propose a module that modulates the axonal delay through short-term memory. To this end, a rectified axonal delay (RAD) module is integrated with the spiking model to align the spike timing and thus improve the characterization learning ability of temporal features. Experiments on three neuromorphic benchmark datasets : NMNIST, DVS Gesture and N-TIDIGITS18 show that the proposed method achieves the state-of-the-art performance while using the fewest parameters. Pengfei Sun 0003, Longwei Zhu, Dick Botteldooren |
ICASSP | 3 |
| 2022 | Sensor characterization for supporting de-noising and auto-calibration of sensor dataabstractIn earlier work, a mobile sensor network had been deployed in passenger cars to measure road quality using sound and vibration sensors. A self-calibration and confounder removal (SCCR) algorithm eliminated the effect of the measurement device in sound and vibration measurements. It relied on the hypothesis that two measurements of the same object should give the same result. However, once the deep learning model has been trained, the model cannot easily include more devices, because it relies on a one-hot encoded device identification. Therefore, additional characterization of both sensor and vehicle is proposed here using a batch of observations (a bag) from the same sensors. A set-encoder is trained with a device classification head. This set-encoder compresses the input bag into a latent representation, irrespective of the input order. This latent vector, which characterizes the sensor, is injected into the SCCR framework to steer the calibration. New devices are characterized by the pre-trained set-encoder. This relies on the hypothesis that closely related sensors with a similar response should be close to each other in the latent space. Therefore, SCCR is now possible with new devices, previously unseen at the training time of SCCR. However, an increase in reconstruction error is observed for the unseen devices. Wout Van Hauwermeiren, Yuanbo Hou, Karlo Filipan, Dick Botteldooren |
IJCNN | 4 |
| 2022 | Relation-guided acoustic scene classification aided with event embeddingsabstractIn real life, acoustic scenes and audio events are naturally correlated. Humans instinctively rely on fine-grained audio events as well as the overall sound characteristics to distinguish diverse acoustic scenes. Yet, most previous approaches treat acoustic scene classification (ASC) and audio event classification (AEC) as two independent tasks. A few studies on scene and event joint classification either use synthetic audio datasets that hardly match the real world, or simply use the multi-task framework to perform two tasks at the same time. Neither of these two ways makes full use of the implicit and inherent relation between fine-grained events and coarse-grained scenes. To this end, this paper proposes a relation-guided ASC (RGASC) model to further exploit and coordinate the scene-event relation for the mutual benefit of scene and event recognition. The TUT Urban Acoustic Scenes 2018 dataset (TUT2018) is annotated with pseudo labels of events by a simple and efficient audiorelated pre-trained model PANN, which is one of the state-of-the-art AEC models. Then, a prior scene-event relation matrix is defined as the average probability of the presence of each event type in each scene class. Finally, the two-tower RGASC model is jointly trained on the real-life dataset TUT2018 for both scene and event classification. The following results are achieved. 1) RGASC effectively coordinates the true information of coarsegrained scenes and the pseudo information of fine-grained events. 2) The event embeddings learned from pseudo labels under the guidance of prior scene-event relations help reduce the confusion between similar acoustic scenes. 3) Compared with other (non-ensemble) methods, RGASC improves the scene classification accuracy on the real-life dataset. Yuanbo Hou, Bo Kang, Wout Van Hauwermeiren, Dick Botteldooren |
IJCNN | 4 |
| 2022 | Event-related data conditioning for acoustic event classificationabstractModels based on diverse attention mechanisms have recently shined in tasks related to acoustic event classification (AEC). Among them, self-attention is often used in audio-only tasks to help the model recognize different acoustic events. Self-attention relies on the similarity between time frames, and uses global information from the whole segment to highlight specific features within a frame. In real life, information related to acoustic events will attenuate over time, which means the information within some frames around the event deserves more attention than distant time global information that may be unrelated to the event. This paper shows that self-attention may over-enhance certain segments of audio representations, and smooth out the boundaries between events representations and background noises. Hence, this paper proposes an event-related data conditioning (EDC) for AEC. EDC directly works on spectrograms. The idea of EDC is to adaptively select the frame-related attention range based on acoustic features, and gather the event-related local information to represent the frame. Experiments show that: 1) compared with spectrogram-based data augmentation methods and trainable feature weighting and self-attention, EDC outperforms them in both the original-size mode and the augmented mode; 2) EDC effectively gathers event-related local information and enhances boundaries between events and backgrounds, improving the performance of AEC. Yuanbo Hou, Dick Botteldooren |
INTERSPEECH | 2 |
| 2022 | CT-SAT: Contextual Transformer for Sequential Audio TaggingabstractSequential audio event tagging can provide not only the type information of audio events, but also the order information between events and the number of events that occur in an audio clip. Most previous works on audio event sequence analysis rely on connectionist temporal classification (CTC). However, CTC's conditional independence assumption prevents it from effectively learning correlations between diverse audio events. This paper first introduces the Transformer into sequential audio tagging, since Transformers perform well in sequence-related tasks. To better utilize contextual information of audio event sequences, we draw on the idea of bidirectional recurrent neural networks, and propose a contextual Transformer (cTransformer) with a bidirectional decoder that could exploit the forward and backward information of event sequences. Experiments on the real-life polyphonic audio dataset show that, compared to CTC-based methods, the cTransformer can effectively combine the fine-grained acoustic representations from the encoder and coarse-grained audio event cues to exploit contextual information to successfully recognize and predict the audio event sequence in polyphonic audio clips. Yuanbo Hou, Zhaoyi Liu 0003, Bo Kang, Dick Botteldooren |
INTERSPEECH | 5 |
| 2022 | Audio-visual scene classification via contrastive event-object alignment and semantic-based fusionabstractPrevious works on scene classification are mainly based on audio or visual signals, while humans perceive the environmental scenes through multiple senses. Recent studies on audio-visual scene classification separately fine-tune the large-scale audio and image pre-trained models on the target dataset, then either fuse the intermediate representations of the audio model and the visual model, or fuse the coarse-grained decision of both models at the clip level. Such methods ignore the detailed audio events and visual objects in audio-visual scenes (AVS), while humans often identify a scene through both audio events and visual objects within, and the congruence between them. To exploit the fine-grained information of audio events and visual objects in AVS, and coordinate the implicit relationship between audio events and visual objects, this paper proposes a multi-branch model equipped with contrastive event-object alignment (CEOA) and semantic-based fusion (SF) for AVSC. CEOA aims to align the learned embeddings of audio events and visual objects by comparing the difference between audio-visual event-object pairs. Then, visual objects associated with certain audio events and vice versa are accentuated by cross-attention and undergo SF for semantic-level fusion. Experiments show that: 1) the proposed AVSC model equipped with CEOA and SF outperforms the results of audio-only and visual-only models, i.e., the audio-visual results are better than the results from a single modality. 2) CEOA aligns the embeddings of audio events and related visual objects on a fine-grained level, and the SF effectively integrates both; 3) Compared with other large-scale integrated systems, the proposed model shows competitive performance, even without using additional datasets and data augmentation tricks. Yuanbo Hou, Bo Kang, Dick Botteldooren |
MMSP | 3 |
| 2022 | Map Matching and Lane Detection Based on Markovian Behavior, GIS, and IMU DataabstractThis paper presents a fast, memory-efficient, and worldwide map matching algorithm based on raw geographic coordinates and enriched open map data with support for trajectories on foot, by bike, and motorized vehicles. The proposed algorithm combines the Markovian behavior and the shortest path aspect while taking into account the type and direction of all road segments, information about one-way traffic, maximum allowed speed per road segment, and driving behavior. Furthermore, a self-adapting lane detection algorithm based solely on accelerometer readings is added on top of the map matching algorithm. An experimental validation consisting of 30 trajectories on foot, by bike, and by car, showed the efficiency and accuracy of the proposed algorithms, with an average F1-score and median error of 99.5% and 1.89 m for the map matching algorithm and an average F1-score of 86.7% for the lane detection algorithm, which resulted in the correctly estimated lane 93.0% of the time. Moreover, the proposed technique outperforms existing state of the art techniques with accuracy improvements up to 45.2%. Jens Trogh, Dick Botteldooren, Bert De Coensel, Luc Martens, Wout Joseph, David Plets |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Rule-Embedded Network for Audio-Visual Voice Activity Detection in Live Musical Video StreamsabstractDetecting anchor’s voice in live musical streams is an important preprocessing step for music and speech signal processing. Existing approaches to voice activity detection (VAD) primarily rely on audio, however, audio-based VAD is difficult to effectively focus on the target voice in noisy environments. This paper proposes a rule-embedded network to fuse the audio-visual (A-V) inputs for better detection of the target voice. The core role of the rule in the model is to coordinate the relation between the bi-modal information and use visual representations as a mask to filter out the information of non-target sound. Experiments show that: 1) with the help of cross-modal fusion using the proposed rule, the detection results of the A-V branch outperform that of the audio branch in the same model framework; 2) the performance of the bimodal A-V model far outperforms that of audio-only models, indicating that the incorporation of both audio and visual signals is highly beneficial for VAD. To attract more attention to the cross-modal music and audio signal processing, a new live musical video corpus with frame-level labels is introduced. Yuanbo Hou, Bilei Zhu, Zejun Ma 0001, Dick Botteldooren |
ICASSP | 5 |
| 2021 | Attention-Based Cross-Modal Fusion for Audio-Visual Voice Activity Detection in Musical Video StreamsabstractMany previous audio-visual voice-related works focus on speech, ignoring the singing voice in the growing number of musical video streams on the Internet. For processing diverse musical video data, voice activity detection is a necessary step. This paper attempts to detect the speech and singing voices of target performers in musical video streams using audio-visual information. To integrate information of audio and visual modalities, a multi-branch network is proposed to learn audio and image representations, and the representations are fused by attention based on semantic similarity to shape the acoustic representations through the probability of anchor vocalization. Experiments show the proposed audio-visual multi-branch network far outperforms the audio-only model in challenging acoustic environments, indicating the cross-modal information fusion based on semantic correlation is sensible and successful. Yuanbo Hou, Zhesong Yu, Xia Liang, Xingjian Du, Bilei Zhu, Zejun Ma 0001, Dick Botteldooren |
Interspeech | 7 |
| 2020 | Detection of road pavement quality using statistical clustering methods
Joachim David, Toon De Pessemier, Luc Dekoninck, Bert De Coensel, Wout Joseph, Dick Botteldooren, Luc Martens |
J. Intell. Inf. Syst. | 6 |
| 2014 | Long-term learning behavior in a recurrent neural network for sound recognitionabstractIn this paper, the long-term learning properties of an artificial neural network model, designed for sound recognition and computational auditory scene analysis in general, are investigated. The model is designed to run for long periods of time (weeks to months) on low-cost hardware, used in a noise monitoring network, and builds upon previous work by the same authors. It consists of three neural layers, connected to each other by feedforward and feedback excitatory connections. It is shown that the different mechanisms that drive auditory attention emerge naturally from the way in which neural activation and intra-layer inhibitory connections are implemented in the model. Training of the artificial neural network is done following the Hebb principle, dictating that "Cells that fire together, wire together", with some important modifications, compared to standard Hebbian learning. As the model is designed to be on-line for extended periods of time, also learning mechanisms need to be adapted to this. The learning needs to be strongly attention- and saliency-driven, in order not to waste available memory space for sounds that are of no interest to the human listener. The model also implements plasticity, in order to deal with new or changing input over time, without catastrophically forgetting what it already learned. On top of that, it is shown that also the implementation of short-term memory plays an important role in the long-term learning properties of the model. The above properties are investigated and demonstrated by training on real urban sound recordings. Michiel Boes, Damiano Oldoni, Bert De Coensel, Dick Botteldooren |
IJCNN | 4 |
| 2013 | A biologically inspired recurrent neural network for sound source recognition incorporating auditory attentionabstractIn this paper, a human-mimicking model for sound source recognition is presented. It consists of an artificial neural network with three neuron layers (input, middle and output) that are connected by feedback connections between the output and middle layer, on top of feedforward connections from the input to middle and middle to output layers. Learning is accomplished by the model following the Hebb principle, dictating that “cells that fire together, wire together”, with some important alterations, compared to standard Hebbian learning, in order to prevent the model from forgetting previously learned patterns, when learning new ones. In addition, short-term memory is introduced into the model in order to facilitate and guide learning of neuronal synapses (long-term memory). As auditory attention is an essential part of human auditory scene analysis (ASA), it is also indispensable in any computational model mimicking it, and it is shown that different auditory attention mechanism naturally emerge from the neuronal behaviour as implemented in the model described in this paper. The learning behavior of the model is further investigated in the context of an urban sonic environment, and the importance of short-term memory in this process is demonstrated. Finally, the effectiveness of the model is evaluated by comparing model output on presented sound recordings to a human expert listeners evaluation of the same fragments. Michiel Boes, Damiano Oldoni, Bert De Coensel, Dick Botteldooren |
IJCNN | 4 |
| 2012 | Attention-driven auditory stream segregation using a SOM coupled with an excitatory-inhibitory ANNabstractAuditory attention is an essential property of human hearing. It is responsible for the selection of information to be sent to working memory and as such to be perceived consciously, from the abundance of auditory information that is continuously entering the ears. Thus, auditory attention heavily influences human auditory perception and systems simulating human auditory scene analysis would benefit from an attention model. In this paper, a human-mimicking model of auditory attention is presented, aimed to be used in environmental sound monitoring. It relies on a Self-Organizing Map (SOM) for learning and classifying sounds. Coupled to this SOM, an excitatory-inhibitory artificial neural network (ANN), simulating the auditory cortex, is defined. The activation of these neurons is calculated based on an interplay of various excitatory and inhibitory inputs. The latter simulate auditory attention mechanisms in a human-inspired but simplified way, in order to keep the computational cost within bounds. The behavior of the model incorporating all of these mechanisms is investigated, and plausible results are obtained. Michiel Boes, Damiano Oldoni, Bert De Coensel, Dick Botteldooren |
IJCNN | 4 |
| 2010 | Context-dependent environmental sound monitoring using SOM coupled with LEGIONabstractEnvironmental sound measurement networks are increasingly applied for monitoring noise pollution in an urban context. Intelligent measurement nodes offer the opportunity to perform advanced analysis of environmental sound, but tradeoffs between cost and functionality still have to be made. When using a tiered architecture, local nodes with limited computing capabilities can be used to detect sound events of potential interest, which are then further analyzed by more powerful nodes. This paper presents a human-mimicking model for detecting rare and conspicuous sound events. Features encoding spectro-temporal irregularities are extracted from the sound, and a Self-Organizing Map (SOM) is used to identify co-occurring features, which most likely belong to a single sound object. Extensive training allows this map to be tuned to the typical sounds that are heard at the microphone location. A Locally Excitatory Globally Inhibitory Oscillator Network (LEGION) is used to group units of the SOM in order to construct distinct sound objects. Damiano Oldoni, Bert De Coensel, Michaël Rademaker, Bernard De Baets, Dick Botteldooren |
IJCNN | 5 |
| 2008 | Clustering outdoor soundscapes using fuzzy antsabstractA classification algorithm for environmental sound recordings or “soundscapes” is outlined. An ant clustering approach is proposed, in which the behavior of the ants is governed by fuzzy rules. These rules are optimized by a genetic algorithm specially designed in order to achieve the optimal set of homogeneous clusters. Soundscape similarity is expressed as fuzzy resemblance of the shape of the sound pressure level histogram, the frequency spectrum and the spectrum of temporal fluctuations. These represent the loudness, the spectral and the temporal content of the soundscapes. Compared to traditional clustering methods, the advantages of this approach are that no a priori information is needed, such as the desired number of clusters, and that a flexible set of soundscape measures can be used. The clustering algorithm was applied to a set of 1116 acoustic measurements in 16 urban parks of Stockholm. The resulting clusters were validated against visitor’s perceptual measurements of soundscape quality. Bert De Coensel, Dick Botteldooren, Kenny Debacq, Mats E. Nilsson, Birgitta Berglund |
IEEE Congress on Evolutionary Computation | 2 |
| 2008 | A model for long-term environmental sound detectionabstractKnowledge on primary processing of sound by the human auditory system has tremendously increased. This paper exploits the opportunities this creates for assessing the impact of (unwanted) environmental noise on quality of life of people. In particular the effect of auditory attention in a multisource context is focused on. The typical application envisaged here is characterized by very long term exposure (days) and multiple listeners (thousands) that need to be assessed. Therefore, the proposed model introduces many simplifications. The results obtained show that the approach is nevertheless capable of generating insight in the emergence of annoyance and the appraisal of open area soundscapes. Dick Botteldooren, Bert De Coensel |
IJCNN | 1 |
| 2006 | Fuzzy Integrals as a Tool for Obtaining an Indicator for Quality of LifeabstractMany governments and international organizations recognize the fact that classical economic indicators do not accurately reflect the quality of life in a country or region. Indicators that reflect the subjective evaluation of well-being or quality of life by inhabitants have to be constructed based on multiple criteria. Multi-criteria evaluation systems based on fuzzy integrals seem very well suited for this task as they tend to approximate overall quality assessment by inhabitants quite well. This paper discusses the choice of fuzzy integrals and analyses the suitability of the approach based on a survey with 2000 people. Dick Botteldooren, Andy Verkeyn, Bernard De Baets, Peter Lercher |
FUZZ-IEEE | 1 |
| 2004 | Fuzzy translation tool for linguistic termsabstractAn automatic translation tool for linguistic terms is built. The terms are represented by fuzzy sets and the translations are based on the similarity degree between those fuzzy sets. The tool is tested on 21 adverbs in 9 languages. The fuzzy sets are constructed with a probability based approach, based on data from an International study on the choice of appropriate terms to label a noise annoyance scale. The results are in agreement with common sense translations. A detailed sensitivity analysis shows that the procedure is stable for many operator choices. Andy Verkeyn, Dick Botteldooren |
FUZZ-IEEE | 2 |
| 2003 | Uncertainty in Noise Mapping: Comparing a Probabilistic and a Fuzzy Set Approach
Tom De Muer, Dick Botteldooren |
IFSA | 2 |
| 2003 | Sugeno Integrals for the Modelling of Noise Annoyance Aggregation
Andy Verkeyn, Dick Botteldooren, Bernard De Baets, Guy De Tré |
IFSA | 2 |
| 2002 | An iterative fuzzy model for cognitive processes involved in environment quality judgementabstractSocial surveys are commonly used to assess the quality of the living environment. Models allow quantifying the impact of future developments. Starting from basic knowledge on the cognitive process of human judgement a fuzzy model is developed, focussing on the aggregation of the impact of particular activities to a global judgement. Dick Botteldooren, Andy Verkeyn |
FUZZ-IEEE | 1 |
| 2002 | Towards language independent models based on survey dataabstractThe modeling of the impact of environmental pollution is often based on survey data. Language related issues are known to complicate the comparison of such models. This paper tackles this problem by isolating the model from the language of the survey data using fuzzy set theory. Andy Verkeyn, Dick Botteldooren |
FUZZ-IEEE | 2 |