Lin Wang 0009

dblp:17/6729-9 · DBLP profile ↗
← Back
19ranked-venue papers
11as first author
6since 2021 · last 2026
0000-0001-8095-9518ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 8 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 On the Consistency Between Subjective and Objective Evaluation for Speech Enhancement Under Low-SNR Drone Noise
abstract
Speech enhancement in the presence of strong drone noise at extremely low signal-to-noise ratios (SNRs) remains challenging, and conventional objective metrics may fail to accurately reflect perceived speech quality. In this paper, we evaluate several deep learning-based speech enhancement algorithms under severe drone-noise conditions using a diverse set of objective metrics, including signal fidelity, intelligibility, and perceptual quality measures. The results reveal substantial inconsistencies among metrics, leading to conflicting conclusions regarding algorithm performance. To address this issue, we conduct a controlled subjective listening experiment under extremely low-SNR conditions to assess the perceptual relevance of these metrics. The consistency analysis shows that learning-based perceptual metrics (e.g., NISQA) and intelligibility-oriented measures (e.g., ESTOI) achieve the highest agreement with subjective judgments, followed by SI-SDR, while traditional quality metrics such as PESQ exhibit weaker agreement. These findings highlight the limitations of conventional metrics in extremely low-SNR conditions and provide guidance for selecting perceptually meaningful evaluation criteria in drone audition.
Dmitrii Mukhutdinov, Ashish Alex, Andrea Cavallaro, Lin Wang 0009
IEEE Signal Process. Lett.4
2024 When Cohesion Lies in the Embedding Space: Embedding-Based Reference-Free Metrics for Topic Segmentation
abstract
In this paper we propose a new framework and new methods for the reference-free evaluation of topic segmentation systems directly in the embedding space. Specifically, we define a common framework for reference-free, embedding-based topic segmentation metrics, and show how this applies to an existing metric. We then define new metrics, based on a previously defined cohesion score, Average Relative Proximity. Using this approach, we show that Large Language Models (LLMs) yield features that, if used correctly, can strongly correlate with traditional topic segmentation metrics based on costly and rare human annotations, while outperforming existing reference-free metrics borrowed from clustering evaluation in most domains. We then show that smaller language models specifically fine-tuned for different sentence-level tasks can outperform LLMs several orders of magnitude larger. Via a thorough comparison of our metric’s performance across different datasets, we see that conversational data present the biggest challenge in this framework. Finally, we analyse the behaviour of our metrics in specific error cases, such as those of under-generation and moving of ground truth topic boundaries, and show that our metrics behave more consistently than other reference-free methods.
Iacopo Ghinassi, Lin Wang 0009, Chris Newell, Matthew Purver
LREC/COLING2
2024 Enhanced Speech Emotion Recognition Incorporating Speaker-Sensitive Interactions in Conversations
abstract
Accurately detecting emotions in conversation is a necessary yet challenging task due to the complexity of emotions and dynamics in dialogues. The emotional state of a speaker can be influenced by many different factors, such as interlocutor stimulus, dialogue scene, and topic. In this work, we propose a conversational speech emotion recognition method to deal with capturing attentive contextual dependency and speaker-sensitive interactions. First, we use a pretrained WavLM model to extract frame-based audio representation in individual utterances. Second, an attentive bi-directional gated recurrent unit (GRU) models contextual-sensitive information and explores listener dependency and speaker influence jointly in a simple, fast, parameter-efficient way. The experiments conducted on the standard conversational dataset MELD demonstrate the effectiveness of the proposed method when compared against state-of the-art methods.
Jiachen Luo, Huy Phan, Lin Wang 0009, Joshua D. Reiss
ICME3
2023 Multimodal Topic Segmentation of Podcast Shows with Pre-trained Neural Encoders
abstract
We present two multimodal models for topic segmentation of podcasts built on pre-trained neural text and audio embeddings. We show that results can be improved by combining different modalities; but also by combining different encoders from the same modality, especially general-purpose sentence embeddings with specifically fine-tuned ones. We also show that audio embeddings can be substituted with two simple features related to sentence duration and inter-sentential pauses with comparable results. Finally, we publicly release our two datasets, the first in our knowledge publicly and freely available multimodal datasets for topic segmentation.
Iacopo Ghinassi, Lin Wang 0009, Chris Newell, Matthew Purver
ICMR2
2023 Data augmentation for speech separation
abstract
Deep learning models have advanced the state of the art of monaural speech separation. However, the performance of a separation model considerably decreases when tested on unseen speakers and noisy conditions. Separation models trained with data augmentation generalize better to unseen conditions. In this paper, we conduct a comprehensive survey of data augmentation techniques and apply them to improve the generalization of time-domain speech separation models. The augmentation techniques include seven source-preserving approaches (Gaussian noise, Gain, Time masking, frequency masking, Short noise, Time stretch, and Pitch shift) and three non-source preserving approaches (Dynamix mixing, Mixup, and Cutmix). After hyperparameter search for each augmentation method, we test the generalization of the augmented model by cross-corpus testing on three datasets (LibriMix, TIMIT, and VCTK), and identify the best augmentation combination that enhances generalization. Experimental results indicate that a combination of several non-source preserving strategies (CutMix, Mixup, and Dynamic mixing) resulted in the best generalization performance. Finally, the augmentation combinations also improved the performance of the speech separation model even when fewer training data are available.
Ashish Alex, Lin Wang 0009, Paolo Gastaldo, Andrea Cavallaro
Speech Commun.2
2021 Mixup Augmentation for Generalizable Speech Separation
abstract
Deep learning has advanced the state of the art of single-channel speech separation. However, separation models may overfit the training data and generalization across datasets is still an open problem in real-world conditions with noise. In this paper we address the generalization problem with Mixup as data augmentation approach. Mixup creates new training examples from linear combinations of samples during mini-batch training. We propose four variations of Mixup and assess the improved generalization of a speech separation model, DPRNN, with cross-corpus evaluation on LibriMix, TIMIT and VCTK datasets. DPRNN allows efficient modelling of longer input sequences by splitting the learnt representation from input mixture segment into small chunks and performing intra and inter chunk operations iteratively. We show that training DPRNN with the proposed Data-only Mixup augmentation variation improves performance on an unseen dataset in noisy conditions when compared to the baseline SpecAugment augmented models, while having comparable performance on the source dataset.
Ashish Alex, Lin Wang 0009, Paolo Gastaldo, Andrea Cavallaro
MMSP2
2020 A Blind Source Separation Framework for Ego-Noise Reduction on Multi-Rotor Drones
abstract
Acoustic sensing from a multi-rotor drone is heavily degraded by the strong ego-noise produced by the rotating motors and propellers. To address this problem, we propose a blind source separation (BSS) framework that extracts a target sound from noisy multi-channel signals captured by a microphone array mounted on a drone. The proposed method addresses the challenging problem of permutation alignment, in extremely low signal-to-noise-ratio scenarios (e.g. SNR <; -15 dB), by performing clustering on the time activities of the separated signals across frequencies. Since initialization plays an important role to the success of clustering, we propose a pre-processing algorithm which uses time-frequency spatial filtering (TFS) to generate a reference to pre-align the permutation. The pre-alignment not only improves the performance of clustering and permutation alignment, but also solves the target-channel selection problem for BSS. The proposed method integrates the advantages of both TFS and BSS. Experimental results with real-recorded data show that the proposed method is capable of processing the audio stream continuously in a blockwise manner and also remarkably outperforms the state-of-the-art.
Lin Wang 0009, Andrea Cavallaro
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 Sound-based Transportation Mode Recognition with Smartphones
abstract
Smartphone-based identification of the mode of transportation of the user is important for context-aware services. We investigate the feasibility of recognizing the 8 most common modes of locomotion and transportation from the sound recorded by a smartphone carried by the user. We propose a convolutional neural network based recognition pipeline, which operates on the short-time Fourier transform (STFT) spectrogram of the sound in the log domain. Experiment with the Sussex-Huawei locomotion-transportation (SHL) dataset on 366 hours of data shows promising results where the proposed pipeline can recognize the activities Still, Walk, Run, Bike, Car, Bus, Train and Subway with a global accuracy of 86.6%, which is 23% higher than classical machine learning pipelines. It is shown that sound is particularly useful for distinguishing between various vehicle activities (e.g. Car vs Bus, Train vs Subway). This discriminablity is complementary to the widely used motion sensors, which are poor at distinguish between rail and road transport.
Lin Wang 0009, Daniel Roggen
ICASSP1
2019 Audio-visual sensing from a quadcopter: dataset and baselines for source localization and sound enhancement
abstract
We present an audio-visual dataset recorded outdoors from a quadcopter and discuss baseline results for multiple applications. The dataset includes a scenario for source localization and sound enhancement with up to two static sources, and a scenario for source localization and tracking with a moving sound source. These sensing tasks are made challenging by the strong and time-varying ego-noise generated by the rotating motors and propellers. The dataset was collected using a small circular array with 8 microphones and a camera mounted on the quadcopter. The camera view was used to facilitate the annotation of the sound-source positions and can also be used for multi-modal sensing tasks. We discuss the audio-visual calibration procedure that is needed to generate the annotation for the dataset, which we make available to the research community.
Lin Wang 0009, Ricardo Sanchez-Matilla, Andrea Cavallaro
IROS1
2018 Tracking a moving sound source from a multi-rotor drone
abstract
We propose a method to track from a multi-rotor drone a moving source, such as a human speaker or an emergency whistle, whose sound is mixed with the strong ego-noise generated by rotating motors and propellers. The proposed method is independent of the specific drone and does not need pre-training nor reference signals. We first employ a time-frequency spatial filter to estimate, on short audio segments, the direction of arrival of the moving source and then we track these noisy estimations with a particle filter. We quantitatively evaluate the results using a ground-truth trajectory of the sound source obtained with an on-board camera and compare the performance of the proposed method with baseline solutions.
Lin Wang 0009, Ricardo Sanchez-Matilla, Andrea Cavallaro
IROS1
2018 Pseudo-Determined Blind Source Separation for Ad-hoc Microphone Networks
abstract
We propose a pseudo-determined blind source separation framework that exploits the information from a large number of microphones in an ad-hoc network to extract and enhance sound sources in a reverberant scenario. After compensating for the time offsets and sampling rate mismatch between (asynchronous) signals, we interpret as a determined M × M mixture the over-determined M × N mixture, where M ) N is the number of microphones and N is the number of sources. Next, we propose a pseudodetermined mixture model that can apply an M × M independent component analysis (ICA) directly to the M-channel recordings. Moreover, we propose a reference-based permutation alignment scheme that aligns the permutation of the ICA outputs and classifies them into target channels, which contain the N sources, and nontarget channels, which contain reverberation residuals. Finally, using the signals from nontarget channels, we estimate in each target channel the power spectral density of the noise component that we suppress with a spectral postfilter. Interestingly, we also obtain late-reverberation suppression as byproduct. Experiments show that each processing block improves incrementally source separation and that the performance of the proposed pseudodetermined separation improves as the number of microphones increases.
Lin Wang 0009, Andrea Cavallaro
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 Time-frequency processing for sound source localization from a micro aerial vehicle
abstract
We address the problem of sound source localization with a microphone array mounted on a micro aerial vehicle (MAV). Due to the noise generated by motors and propellers, this scenario is characterized by extremely low signal-to-noise ratios (SNR). Based on the observation that the energy of MAV sound recordings is usually concentrated at isolated time-frequency bins, we propose a time-frequency processing framework to address this problem. We first estimate the direction of arrival of the sound at individual time-frequency bins. Then we formulate a set of spatially informed filters pointing at candidate directions in the search space. The output of the filtering tends to present high non-Gaussianity when the spatial filter is steered towards the target sound source. Finally, by measuring the non-Gaussianity of the spatial filtering outputs we build a spatial likelihood function from which we estimate the direction of the target sound. Experimental results with real-recorded MAV ego-noise show the superiority of the proposed method over the state of the art in performing source localization robustly.
Lin Wang 0009, Andrea Cavallaro
ICASSP1
2017 Multi-Modal Localization and Enhancement of Multiple Sound Sources from a Micro Aerial Vehicle
abstract
The ego-noise generated by the motors and propellers of a micro aerial vehicle (MAV) masks the environmental sounds and considerably degrades the quality of the on-board sound recording. Sound enhancement approaches generally require knowledge of the direction of arrival of the target sound sources, which are difficult to estimate due to the low signal-to-noise-ratio (SNR) caused by the ego-noise and the interferences between multiple sources. To address this problem, we propose a multi-modal analysis approach that jointly exploits audio and video to enhance the sounds of multiple targets captured from an MAV equipped with a microphone array and a video camera. We first address audio-visual calibration via camera resectioning, audio-visual temporal alignment and geometrical alignment to jointly use the features in the audio and video streams, which are independently generated. The spatial information from the video is used to assist sound enhancement by tracking multiple potential sound sources with a particle filter. Then we infer the directions of arrival of the target sources from the video tracking results and extract the sound from the desired direction with a time-frequency spatial filter, which suppresses the ego-noise by exploiting its time-frequency sparsity. Experimental demonstration results with real outdoor data verify the robustness of the proposed multi-modal approach for multiple speakers in extremely low-SNR scenarios.
Ricardo Sanchez-Matilla, Lin Wang 0009, Andrea Cavallaro
ACM Multimedia2
2016 Ear in the sky: Ego-noise reduction for auditory micro aerial vehicles
abstract
We investigate the spectral and spatial characteristics of the ego-noise of a multirotor micro aerial vehicle (MAV) using audio signals captured with multiple onboard microphones and derive a noise model that grounds the feasibility of microphone-array techniques for noise reduction. The spectral analysis suggests that the ego-noise consists of narrowband harmonic noise and broadband noise, whose spectra vary dynamically with the motor rotation speed. The spatial analysis suggests that the ego-noise of a P-rotor MAV can be modeled as P directional noises plus one diffuse noise. Moreover, because of the fixed positions of the microphones and motors, we can assume that the acoustic mixing network of the ego-noise is stationary. We validate the proposed noise model and the stationary mixing assumption by applying blind source separation to multi-channel recordings from both a static and a moving MAV and quantify the signal-to-noise ratio improvement. Moreover, we make all the audio recordings publicly available.
Lin Wang 0009, Andrea Cavallaro
AVSS1
2016 An Iterative Approach to Source Counting and Localization Using Two Distant Microphones
abstract
We propose a time difference of arrival (TDOA) estimation framework based on time-frequency inter-channel phase difference (IPD) to count and localize multiple acoustic sources in a reverberant environment using two distant microphones. The time-frequency (T-F) processing enables exploitation of the nonstationarity and sparsity of audio signals, increasing robustness to multiple sources and ambient noise. For inter-channel phase difference estimation, we use a cost function, which is equivalent to the generalized cross correlation with phase transform (GCC) algorithm and which is robust to spatial aliasing caused by large inter-microphone distances. To estimate the number of sources, we further propose an iterative contribution removal (ICR) algorithm to count and locate the sources using the peaks of the GCC function. In each iteration, we first use IPD to calculate the GCC function, whose highest peak is detected as the location of a sound source; then we detect the T-F bins that are associated with this source and remove them from the IPD set. The proposed ICR algorithm successfully solves the GCC peak ambiguities between multiple sources and multiple reverberant paths.
Lin Wang 0009, Tsz-Kin Hon, Joshua D. Reiss, Andrea Cavallaro
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 Over-Determined Source Separation and Localization Using Distributed Microphones
abstract
We propose an overdetermined source separation and localization method for a set of M microphones distributed around an unknown number, N <; M, of sources. We reformulate the overdetermined acoustic mixing procedure with a new determined mixing model and apply a determined M × M independent component analysis ('CA) in each frequency bin directly. The reformulated 'CA operates without knowing N and also leads to better separation in reverberant scenarios. To solve the challenging permutation ambiguity problem, we first employ a time activity-based clustering approach to cluster the separated frequency components into M channels. We then propose a remixing procedure to detect and merge channels from the same source. The detection is done by analyzing time and frequency activities, spectral likeliness, and spatial location. To estimate the spatial location, we propose a time-frequency masking-based steered response power algorithm. Simulated and real-data experiments in a very challenging reverberant scenario confirm the effectiveness of the proposed method in obtaining the number of sources, the separated signals, and the location and spatial likelihood of each source.
Lin Wang 0009, Joshua D. Reiss, Andrea Cavallaro
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Audio Fingerprinting for Multi-Device Self-Localization
abstract
We investigate the self-localization problem of an ad-hoc network of randomly distributed and independent devices in an open-space environment with low reverberation but heavy noise (e.g. smartphones recording videos of an outdoor event). Assuming a sufficient number of sound sources, we estimate the distance between a pair of devices from the extreme (minimum and maximum) time difference of arrivals (TDOAs) from the sources to the pair of devices without knowing the time offset. The obtained inter-device distances are then exploited to derive the geometrical configuration of the network. In particular, we propose a robust audio fingerprinting algorithm for noisy recordings and perform landmark matching to construct a histogram of the TDOAs of multiple sources. The extreme TDOAs can be estimated from this histogram. By using audio fingerprinting features, the proposed algorithm works robustly in very noisy environments. Experiments with free-field simulation and open-space recordings prove the effectiveness of the proposed algorithm.
Tsz-Kin Hon, Lin Wang 0009, Joshua D. Reiss, Andrea Cavallaro
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 A Novel Hierarchical Decomposition Vector Quantization Method for High-Order LPC Parameters
abstract
The paper investigates vector quantization coding of high-order (e.g., 20th-50th order) linear prediction coding (LPC) parameters, and proposes a novel hierarchical decomposition vector quantization method for a scalable speech coding framework with variable orders of LPC analysis. Instead of vector quantizing the whole group of LPC parameters in the linear spectral frequency (LSF) domain directly, the proposed method decomposes the high-order LPC model into several low-order (e.g., 10th-order) LPC models, and vector quantizes them in the LSF domain separately. For the decomposition, the high-order LPC model is converted into a group of reflection coefficients at first, and then the group is split into several subgroups and converted into multiple low-order LPC models. It is shown that the proposed method is naturally suitable for a scalable coding framework where the information of the decomposed low-order LPC models can be encoded into a multi-layered bitstream and can be combined in a progressive way to recover the high-order LPC information. Experiments in a scalable coding framework with variable LPC analysis orders (10-50) reveal that, compared to a direct vector quantization scheme, the proposed method can reduce the size of the codebook and the number of coding bits significantly, and can also efficiently reduce the computation cost.
Lin Wang 0009, Zhe Chen 0005, Fuliang Yin
IEEE ACM Trans. Audio Speech Lang. Process.1
2011 A Region-Growing Permutation Alignment Approach in Frequency-Domain Blind Source Separation of Speech Mixtures
abstract
The convolutive blind source separation (BSS) problem can be solved efficiently in the frequency domain, where instantaneous BSS is performed separately in each frequency bin. However, the permutation ambiguity in each frequency bin should be resolved so that the separated frequency components from the same source are grouped together. To solve the permutation problem, this paper presents a new alignment method based on an inter-frequency dependence measure: the powers of separated signals. Bin-wise permutation alignment is applied first across all frequency bins, using the correlation of separated signal powers; then the full frequency band is partitioned into small regions based on the bin-wise permutation alignment result. Finally, region-wise permutation alignment is performed in a region-growing manner. The region-wise permutation correction scheme minimizes the spreading of the misalignment at isolated frequency bins to others, hence to improve permutation alignment. Experiment results in simulated and real environments verify the effectiveness of the proposed method. Analysis demonstrates that the proposed frequency-domain BSS method is computationally efficient.
Lin Wang 0009, Heping Ding, Fuliang Yin
IEEE Trans. Speech Audio Process.1