EDBT 2026 Demo / reviewers in the wild / expert
Philip J. B. Jackson
dblp:08/4640 · also Philip Jackson 0003
· DBLP profile ↗
49ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0001-7933-5935ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 21 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Effective-Efficient Approach for Dense Multi-Label Action DetectionabstractAbstract Unlike the sparse label action detection task, where a single action occurs in each timestamp of a video, in a dense multi-label scenario, actions can overlap temporally. To address this challenging task, it is necessary to simultaneously learn (i) co-occurrence action relationships and (ii) temporal dependencies. Current methods model co-occurrence action relationships by explicitly embedding class relations into the transformer network architecture. However, these approaches are not computationally efficient, as the network needs to compute all possible pair action class relations. In this paper, we overcome this by introducing a novel framework trained through a novel learning paradigm that allows the network to benefit from explicitly modelling temporal co-occurrence action dependencies during training without imposing their computational overhead during inference. Furthermore, to model temporal information, recent approaches extract multi-scale temporal features through hierarchical transformer-based networks. However, the self-attention mechanism in transformers inherently loses temporal positional information. We argue that combining this with multiple sub-sampling processes in hierarchical designs can lead to further loss of positional information. Preserving this information is essential for accurate action detection. In this paper, we address this issue by proposing a novel transformer network that (a) employs a non-hierarchical structure when modelling different ranges of temporal dependencies and (b) embeds relative positional encoding in its transformer layers. We evaluate the performance of our proposed approach on two challenging dense multi-label benchmark datasets and show that our method improves the current state-of-the-art results by 1.1% and 0.6% per-frame mAP on the Charades and MultiTHUMOS datasets, respectively, achieving new state-of-the-art per-frame mAP results at 26.5% and 44.6%, respectively. We also performed extensive ablation studies to examine the impact of the different components of our proposed approach. Our code will be released upon paper publication Faegheh Sardari, Armin Mustafa, Philip J. B. Jackson, Adrian Hilton 0001 |
Int. J. Comput. Vis. | 3 |
| 2026 | Reverberation-Based Features for Sound Event Localization and Detection With Distance EstimationabstractSound event localization and detection (SELD) involves predicting active sound event classes over time while estimating their positions. The localization subtask in SELD is usually treated as a direction of arrival estimation problem, ignoring source distance. Only recently, SELD was extended to 3D by incorporating distance estimation, enabling the prediction of sound event positions in 3D space (3D SELD). However, existing methods lack input features specifically designed for distance estimation. We address this gap by introducing two novel reverberation-based feature formats: one using the direct-toreverberant ratio (DRR) and another leveraging signal autocorrelation to capture early reflections. We extensively evaluate and benchmark these features on the STARSS23 dataset, combining them with established SELD features for sound event detection (SED) and direction-of-arrival estimation (DOAE), and testing across different network architectures. Our proposed features, applicable to both FOA and MIC formats, achieve state-of-theart distance estimation, enhancing overall 3D SELD performance. Davide Berghi, Philip J. B. Jackson |
IEEE Signal Process. Lett. | 2 |
| 2025 | SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic SoundscapesabstractSelf-supervised pre-trained audio networks have seen widespread adoption in real-world systems, particularly in multi-modal large language models. These networks are often employed in a frozen state, under the assumption that the self-supervised pre-training has sufficiently equipped them to handle real-world audio. However, a critical question remains: how well do these models actually perform in real-world conditions, where audio is typically polyphonic and complex, involving multiple overlapping sound sources? Current audio self-supervised learning (SSL) methods are often benchmarked on datasets predominantly featuring monophonic audio, such as environmental sounds, and speech. As a result, the ability of SSL models to generalize to polyphonic audio, a common characteristic in natural scenarios, remains underexplored. This limitation raises concerns about the practical robustness of SSL models in more realistic audio settings. To address this gap, we introduce Self-Supervised Learning from Audio Mixtures (SSLAM), a novel direction in audio SSL research, designed to improve the model’s ability to learn from polyphonic data while maintaining strong performance on monophonic data. We thoroughly evaluate SSLAM on standard audio SSL benchmark datasets which are predominantly monophonic and conduct a comprehensive comparative analysis against state-of-the-art (SOTA) methods using a range of high-quality, publicly available polyphonic datasets. SSLAM not only improves model performance on polyphonic audio, but also maintains or exceeds performance on standard audio SSL benchmarks. Notably, it achieves up to a 3.9% improvement on the AudioSet-2M(AS-2M), reaching a mean average precision (mAP) of 50.2. For polyphonic datasets, SSLAM sets new SOTA in both linear evaluation and fine-tuning regimes with performance improvements of up to 9.1%(mAP). These results demonstrate SSLAM's effectiveness in both polyphonic and monophonic soundscapes, significantly enhancing the performance of audio SSL models. Code and pre-trained models are available at https://github.com/ta012/SSLAM. Tony Alex, Sara Atito Ali Ahmed, Armin Mustafa, Muhammad Awais 0001, Philip J. B. Jackson |
ICLR | 5 |
| 2024 | DTF-AT: Decoupled Time-Frequency Audio Transformer for Event ClassificationabstractConvolutional neural networks (CNNs) and Transformer-based networks have recently enjoyed significant attention for various audio classification and tagging tasks following their wide adoption in the computer vision domain. Despite the difference in information distribution between audio spectrograms and natural images, there has been limited exploration of effective information retrieval from spectrograms using domain-specific layers tailored for the audio domain. In this paper, we leverage the power of the Multi-Axis Vision Transformer (MaxViT) to create DTF-AT (Decoupled Time-Frequency Audio Transformer) that facilitates interactions across time, frequency, spatial, and channel dimensions. The proposed DTF-AT architecture is rigorously evaluated across diverse audio and speech classification tasks, consistently establishing new benchmarks for state-of-the-art (SOTA) performance. Notably, on the challenging AudioSet 2M classification task, our approach demonstrates a substantial improvement of 4.4% when the model is trained from scratch and 3.2% when the model is initialised from ImageNet-1K pretrained weights. In addition, we present comprehensive ablation studies to investigate the impact and efficacy of our proposed approach. The codebase and pretrained weights are available on https://github.com/ta012/DTFAT.git Tony Alex, Armin Mustafa, Muhammad Awais 0001, Philip J. B. Jackson |
AAAI | 5 |
| 2024 | CoLeaF: A Contrastive-Collaborative Learning Framework for Weakly Supervised Audio-Visual Video Parsing
Faegheh Sardari, Armin Mustafa, Philip J. B. Jackson, Adrian Hilton 0001 |
ECCV (11) | 3 |
| 2024 | Max-AST: Combining Convolution, Local and Global Self-Attentions for Audio Event ClassificationabstractIn the domain of audio transformer architectures, prior research has extensively investigated isotropic architectures that capture the global context through full self-attention and hierarchical architectures that progressively transition from local to global context utilising hierarchical structures with convolutions or window-based attention. However, the idea of imbuing each individual block with both local and global contexts, thereby creating a hybrid transformer block, remains relatively under-explored in the field.To facilitate this exploration, we introduce Multi Axis Audio Spectrogram Transformer (Max-AST), an adaptation of MaxViT to the audio domain. Our approach leverages convolution, local window-attention, and global grid-attention in all the transformer blocks. The proposed model excels in efficiency compared to prior methods and consistently outperforms state-of-the-art techniques, achieving significant gains of up to 2.6% on the AudioSet full set. Further, we performed detailed ablations to analyse the impact of each of these components on audio feature learning. The source code is available at https://github.com/ta012/MaxAST.git Tony Alex, Armin Mustafa, Muhammad Awais 0001, Philip J. B. Jackson |
ICASSP | 5 |
| 2024 | Fusion of Audio and Visual Embeddings for Sound Event Localization and DetectionabstractSound event localization and detection (SELD) combines two subtasks: sound event detection (SED) and direction of arrival (DOA) estimation. SELD is usually tackled as an audio-only problem, but visual information has been recently included. Few audio-visual (AV)-SELD works have been published and most employ vision via face/object bounding boxes, or human pose keypoints. In contrast, we explore the integration of audio and visual feature embeddings extracted with pre-trained deep networks. For the visual modality, we tested ResNet50 and Inflated 3D ConvNet (I3D). Our comparison of AV fusion methods includes the AV-Conformer and Cross-Modal Attentive Fusion (CMAF) model. Our best models outperform the DCASE 2023 Task3 audio-only and AV baselines by a wide margin on the development set of the STARSS23 dataset, making them competitive amongst state-of-the-art results of the AV challenge, without model ensembling, heavy data augmentation, or prediction post-processing. Such techniques and further pre-training could be applied as next steps to improve performance. Davide Berghi, Peipei Wu, Jinzheng Zhao, Wenwu Wang 0001, Philip J. B. Jackson |
ICASSP | 5 |
| 2024 | ForecasterFlexOBM: A Multi-View Audio-Visual Dataset for Flexible Object-Based Media ProductionabstractLeveraging machine learning techniques, in the context of object-based media production, could enable provision of personalized media experiences to diverse audiences. To fine-tune and evaluate techniques for personalization applications, as well as more broadly, datasets which bridge the gap between research and production are needed. We introduce and release such a dataset, themed around a UK weather forecast and shot against a blue-screen background, of three professional actors/presenters – one male and one female (English) and one female (British Sign Language). Scenes include both production and research-oriented examples, with a range of dialogues and actions. Capture techniques consisted of a synchronized 4K resolution 16-camera array, production-typical microphones plus professional audio mix, a 16-channel microphone array with collocated Grasshopper3 camera, and a photogrammetry array. We demonstrate applications relevant to virtual production and creation of personalized media including neural radiance fields, shadow casting, action/event detection, speaker source tracking and video captioning. Davide Berghi, Craig Cieciura, Farshad Einabadi, Maxine Glancy, Oliver C. Camilleri, Philip Foster, Asmar Nadeem, Faegheh Sardari, Jinzheng Zhao, Marco Volino, Armin Mustafa, Philip J. B. Jackson, Adrian Hilton 0001 |
ICME | 12 |
| 2024 | Leveraging Visual Supervision for Array-Based Active Speaker Detection and LocalizationabstractConventional audio-visual approaches for active speaker detection (ASD) typically rely on visually pre-extracted face tracks and the corresponding single-channel audio to find the speaker in a video. Therefore, they tend to fail every time the face of the speaker is not visible. We demonstrate that a simple audio convolutional recurrent neural network (CRNN) trained with spatial input features extracted from multichannel audio can perform simultaneous horizontal active speaker detection and localization (ASDL), independently of the visual modality. To address the time and cost of generating ground truth labels to train such a system, we propose a new self-supervised training pipeline that embraces a “student-teacher” learning approach. A conventional pre-trained active speaker detector is adopted as a “teacher” network to provide the position of the speakers as pseudo-labels. The multichannel audio “student” network is trained to generate the same results. At inference, the student network can generalize and locate also the occluded speakers that the teacher network is not able to detect visually, yielding considerable improvements in recall rate. Experiments on the TragicTalkers dataset show that an audio network trained with the proposed self-supervised learning approach can exceed the performance of the typical audio-visual methods and produce results competitive with the costly conventional supervised training. We demonstrate that improvements can be achieved when minimal manual supervision is introduced in the learning pipeline. Further gains may be sought with larger training sets and integrating vision with the multichannel audio system. Davide Berghi, Philip J. B. Jackson |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Producing Personalised Object-Based Audio-Visual Experiences: an Ethnographic StudyabstractDevelopments in object-based media and IP-based delivery offer an opportunity to create superior audience experiences through personalisation. Towards the aim of making personalised experiences regularly available across the breadth of audio-visual media, we conducted a study to understand how personalised experiences are being created. This consisted of interviews with producers of six representative case studies, followed by a thematic analysis. We describe the workflows and report on the producers’ experiences and obstacles faced. We found that the metadata models, enabling personalisation, were developed independently for each experience, restricting interoperability of personalisation affordances provided to users. Furthermore, the available tools were not effectively integrated into preferred workflows, substantially increasing role responsibilities and production time. To ameliorate these issues, we propose the development of a unifying metadata framework and novel production tools. These tools should be integrated into existing workflows; improve efficiency using AI; and enable producers to serve more diverse audiences. Craig Cieciura, Maxine Glancy, Philip J. B. Jackson |
IMX | 3 |
| 2021 | Visually Supervised Speaker Detection and Localization via Microphone ArrayabstractActive speaker detection (ASD) is a multi-modal task that aims to identify who, if anyone, is speaking from a set of candidates. Current audio-visual approaches for ASD typically rely on visually pre-extracted face tracks (sequences of consecutive face crops) and the respective monaural audio. However, their recall rate is often low as only the visible faces are included in the set of candidates. Monaural audio may successfully detect the presence of speech activity but fails in localizing the speaker due to the lack of spatial cues. Our solution extends the audio front-end using a microphone array. We train an audio convolutional neural network (CNN) in combination with beamforming techniques to regress the speaker’s horizontal position directly in the video frames. We propose to generate weak labels using a pre-trained active speaker detector on pre-extracted face tracks. Our pipeline embraces the "student-teacher" paradigm, where a trained "teacher" network is used to produce pseudo-labels visually. The "student" network is an audio network trained to generate the same results. At inference, the student network can independently localize the speaker in the visual frames directly from the audio input. Experimental results on newly collected data prove that our approach significantly outperforms a variety of other baselines as well as the teacher network itself. It results in an excellent speech activity detector too. Davide Berghi, Adrian Hilton 0001, Philip J. B. Jackson |
MMSP | 3 |
| 2021 | Acoustic Room Modelling Using 360 Stereo CamerasabstractIn this paper we propose a pipeline for estimating acoustic 3D room structure with geometry and attribute prediction using spherical 360$^{\circ }$cameras. Instead of setting microphone arrays with loudspeakers to measure acoustic parameters for specific rooms, a simple and practical single-shot capture of the scene using a stereo pair of 360 cameras can be used to simulate those acoustic parameters. We assume that the room and objects can be represented as cuboids aligned to the main axes of the room coordinate (Manhattan world). The scene is captured as a stereo pair using off-the-shelf consumer spherical 360 cameras. A cuboid-based 3D room geometry model is estimated by correspondence matching between captured images and semantic labelling using a convolutional neural network (SegNet). The estimated geometry is used to produce frequency-dependent acoustic predictions of the scene. This is, to our knowledge, the first attempt in the literature to use visual geometry estimation and object classification algorithms to predict acoustic properties. Results are compared to measurements through calculated reverberant spatial audio object parameters used for reverberation reproduction customized to the given loudspeaker set up. Hansung Kim 0001, Luca Remaggi, Sam Fowler, Philip J. B. Jackson, Adrian Hilton 0001 |
IEEE Trans. Multim. | 4 |
| 2019 | Robust Full-sphere Binaural Sound Source Localization Using Interaural and Spectral CuesabstractA binaural sound source localization method is proposed that uses interaural and spectral cues for localization of sound sources with any direction of arrival on the full-sphere. The method is designed to be robust to the presence of reverberation, additive noise and different types of sounds. The method uses the interaural phase difference (IPD) for lateral angle localization, then interaural and spectral cues for polar angle localization. The method applies different weighting to the interaural and spectral cues depending on the estimated lateral angle. In particular, only the spectral cues are used for sound sources near or on the median plane. Benjamin R. Hammond, Philip J. B. Jackson |
ICASSP | 2 |
| 2019 | Generalisation in Environmental Sound Classification: The 'Making Sense of Sounds' Data Set and ChallengeabstractHumans are able to identify a large number of environmental sounds and categorise them according to high-level semantic categories, e.g. urban sounds or music. They are also capable of generalising from past experience to new sounds when applying these categories. In this paper we report on the creation of a data set that is structured according to the top-level of a taxonomy derived from human judgements and the design of an associated machine learning challenge, in which strong generalisation abilities are required to be successful. We introduce a baseline classification system, a deep convolutional network, which showed strong performance with an average accuracy on the evaluation data of 80.8%. The result is discussed in the light of two alternative explanations: An unlikely accidental category bias in the sound recordings or a more plausible true acoustic grounding of the high-level categories. Christian Kroos, Oliver Bones, Yin Cao, Lara Harris, Philip J. B. Jackson, William J. Davies, Wenwu Wang 0001, Trevor J. Cox, Mark D. Plumbley |
ICASSP | 5 |
| 2019 | Single-Channel Signal Separation and Deconvolution with Generative Adversarial NetworksabstractSingle-channel signal separation and deconvolution aims to separate and deconvolve individual sources from a single-channel mixture. Single-channel signal separation and deconvolution is a challenging problem in which no prior knowledge of the mixing filters is available. Both individual sources and mixing filters need to be estimated. In addition, a mixture may contain non-stationary noise which is unseen in the training set. We propose a synthesizing-decomposition (S-D) approach to solve the single-channel separation and deconvolution problem. In synthesizing, a generative model for sources is built using a generative adversarial network (GAN). In decomposition, both mixing filters and sources are optimized to minimize the reconstruction error of the mixture. The proposed S-D approach achieves a peak-to-noise-ratio (PSNR) of 18.9 dB and 15.4 dB in image inpainting and completion, outperforming a baseline convolutional neural network PSNR of 15.3 dB and 12.2 dB, respectively and achieves a PSNR of 13.2 dB in source separation together with deconvolution, outperforming a convolutive non-negative matrix factorization (NMF) baseline of 10.1 dB. Qiuqiang Kong, Yong Xu 0004, Philip J. B. Jackson, Wenwu Wang 0001, Mark D. Plumbley |
IJCAI | 3 |
| 2019 | Immersive Spatial Audio Reproduction for VR/AR Using Room Acoustic Modelling from 360° ImagesabstractRecent progresses in Virtual Reality (VR) and Augmented Reality (AR) allow us to experience various VR/AR applications in our daily life. In order to maximise the immersiveness of user in VR/AR environments, a plausible spatial audio reproduction synchronised with visual information is essential. In this paper, we propose a simple and efficient system to estimate room acoustic for plausible reproducton of spatial audio using 360° cameras for VR/AR applications. A pair of 360° images is used for room geometry and acoustic property estimation. A simplified 3D geometric model of the scene is estimated by depth estimation from captured images and semantic labelling using a convolutional neural network (CNN). The real environment acoustics are characterised by frequency-dependent acoustic predictions of the scene. Spatially synchronised audio is reproduced based on the estimated geometric and acoustic properties in the scene. The reconstructed scenes are rendered with synthesised spatial audio as VR/AR content. The results of estimated room geometry and simulated spatial audio are evaluated against the actual measurements and audio calculated from ground-truth Room Impulse Responses (RIRs) recorded in the rooms. Hansung Kim 0001, Luca Hernaggi, Philip J. B. Jackson, Adrian Hilton 0001 |
VR | 3 |
| 2019 | A Speech Synthesis Approach for High Quality Speech Separation and GenerationabstractWe propose a new method for source separation by synthesizing the source from a speech mixture corrupted by various environmental noise. Unlike traditional source separation methods which estimate the source from the mixture as a replica of the original source (e.g. by solving an inverse problem), our proposed method is a synthesis-based approach which aims to generate a new signal (i.e. “fake” source) that sounds similar to the original source. The proposed system has an encoder-decoder topology, where the encoder predicts intermediate-level features from the mixture, i.e. Mel-spectrum of the target source, using a hybrid recurrent and hourglass network, while the decoder is a state-of-the-art WaveNet speech synthesis network conditioned on the Mel-spectrum, which directly generates time-domain samples of the sources. Both objective and subjective evaluations were performed on the synthesized sources, and show great advantages of our proposed method for high-quality speech source separation and generation. Qingju Liu, Philip J. B. Jackson, Wenwu Wang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2019 | Modeling the Comb Filter Effect and Interaural Coherence for Binaural Source SeparationabstractTypical methods for binaural source separation consider only the direct sound as the target signal in a mixture. However, in most scenarios, this assumption limits the source separation performance. It is well known that the early reflections interact with the direct sound, producing acoustic effects at the listening position, e.g. the so-called comb filter effect. In this article, we propose a novel source separation model, that utilizes both the direct sound and the first early reflection information to model the comb filter effect. This is done by observing the interaural phase difference obtained from the time-frequency representation of binaural mixtures. Furthermore, a method is proposed to model the interaural coherence of the signals. Including information related to the sound multipath propagation, the performance of the proposed separation method is improved with respect to the baselines that did not use such information, as illustrated by using binaural recordings made in four rooms, having different sizes and reverberation times. Luca Remaggi, Philip J. B. Jackson, Wenwu Wang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Robust Full-Sphere Binaural Sound Source LocalizationabstractWe propose a novel method for full-sphere binaural sound source localization that is designed to be robust to real world recording conditions. A mask is proposed that is designed to remove diffuse noise and early room reflections. The method makes use of the interaural phase difference (IPD) for lateral angle localization and spectral cues for polar angle localization. The method is tested using different HRTF datasets to generate the test data and training data. The method is also tested with the presence of additive noise and reverberation. The method outperforms the state of the art binaural localization methods for most testing conditions. Benjamin R. Hammond, Philip J. B. Jackson |
ICASSP | 2 |
| 2018 | Synthesis of Images by Two-Stage Generative Adversarial NetworksabstractIn this paper, we propose a divide-and-conquer approach using two generative adversarial networks (GANs) to explore how a machine can draw colorful pictures (bird) using a small amount of training data. In our work, we simulate the procedure of an artist drawing a picture, where one begins with drawing objects' contours and edges and then paints them different colors. We adopt two GAN models to process basic visual features including shape, texture and color. We use the first GAN model to generate object shape, and then paint the black and white image based on the knowledge learned using the second GAN model. We run our experiments on 600 color images. The experimental results show that the use of our approach can generate good quality synthetic images, comparable to real ones. Qiang Huang 0007, Philip J. B. Jackson, Mark D. Plumbley, Wenwu Wang 0001 |
ICASSP | 2 |
| 2018 | Iterative Deep Neural Networks for Speaker-Independent Binaural Blind Speech SeparationabstractIn this paper, we propose an iterative deep neural network (DNN)-based binaural source separation scheme, for recovering two concurrent speech signals in a room environment. Besides the commonly-used spectral features, the DNN also takes non-linearly wrapped binaural spatial features as input, which are refined iteratively using parameters estimated from the DNN output via a feedback loop. Different DNN structures have been tested, including a classic multilayer perception regression architecture as well as a new hybrid network with both convolutional and densely-connected layers. Objective evaluations in terms of PESQ and STOI showed consistent improvement over baseline methods using traditional binaural features, especially when the hybrid DNN architecture was employed. In addition, our proposed scheme is robust to mismatches between the training and testing data. Qingju Liu, Yong Xu 0004, Philip J. B. Jackson, Wenwu Wang 0001, Philip Coleman |
ICASSP | 3 |
| 2018 | Acoustic Reflector Localization and ClassificationabstractThe process of understanding acoustic properties of environments is important for several applications, such as spatial audio, augmented reality and source separation. In this paper, multichannel room impulse responses are recorded and transformed into their direction of arrival (DOA)-time domain, by employing a superdirective beamformer. This domain can be represented as a 2D image. Hence, a novel image processing method is proposed to analyze the DOA-time domain, and estimate the reflection times of arrival and DOAs. The main acoustically reflective objects are then localized. Recent studies in acoustic reflector localization usually assume the room to be free from furniture. Here, by analyzing the scattered reflections, an algorithm is also proposed to binary classify reflectors into room boundaries and interior furniture. Experiments were conducted in four rooms. The classification algorithm showed high quality performance, also improving the localization accuracy, for non-static listener scenarios. Luca Remaggi, Hansung Kim 0001, Philip J. B. Jackson, Filippo Maria Fazi, Adrian Hilton 0001 |
ICASSP | 3 |
| 2018 | An Audio-Visual System for Object-Based Audio: From Recording to ListeningabstractObject-based audio is an emerging representation for audio content, where content is represented in a reproduction-format-agnostic way and, thus, produced once for consumption on many different kinds of devices. This affords new opportunities for immersive, personalized, and interactive listening experiences. This paper introduces an end-to-end object-based spatial audio pipeline, from sound recording to listening. A high-level system architecture is proposed, which includes novel audio-visual interfaces to support object-based capture and listener-tracked rendering, and incorporates a proposed component for objectification, that is, recording content directly into an object-based form. Text-based and extensible metadata enable communication between the system components. An open architecture for object rendering is also proposed. The system's capabilities are evaluated in two parts. First, listener-tracked reproduction of metadata automatically estimated from two moving talkers is evaluated using an objective binaural localization model. Second, object-based scene capture with audio extracted using blind source separation (to remix between two talkers) and beamforming (to remix a recording of a jazz group) is evaluated with perceptually motivated objective and subjective experiments. These experiments demonstrate that the novel components of the system add capabilities beyond the state of the art. Finally, we discuss challenges and future perspectives for object-based audio workflows. Philip Coleman, Andreas Franck, Jon Francombe, Qingju Liu, Teófilo Emídio de Campos, Richard J. Hughes 0001, Dylan Menzies, Marcos F. Simón Gálvez, James Woodcock, Philip J. B. Jackson, Frank Melchior, Chris Pike, Filippo Maria Fazi, Trevor J. Cox, Adrian Hilton 0001 |
IEEE Trans. Multim. | 11 |
| 2018 | Multiple Speaker Tracking in Spatial Audio via PHD Filtering and Depth-Audio FusionabstractIn the object-based spatial audio system, positions of the audio objects (e.g., speakers/talkers or voices) presented in the sound scene are required as important metadata attributes for object acquisition and reproduction. Binaural microphones are often used as a physical device to mimic human hearing and to monitor and analyze the scene, including localization and tracking of multiple speakers. The binaural audio tracker, however, is usually prone to the errors caused by room reverberation and background noise. To address this limitation, we present a multimodal tracking method by fusing the binaural audio with depth information (from a depth sensor, e.g., Kinect). More specifically, the probability hypothesis density (PHD) filtering framework is first applied to the depth stream, and a novel clutter intensity model is proposed to improve the robustness of the PHD filter when an object is occluded either by other objects or due to the limited field of view of the depth sensor. To compensate misdetections in the depth stream, a novel gap filling technique is presented to map audio azimuths obtained from the binaural audio tracker to 3D positions, using speaker-dependent spatial constraints learned from the depth stream. With our proposed method, both the errors in the binaural tracker and the misdetections in the depth tracker can be significantly reduced. Real-room recordings are used to show the improved performance of the proposed method in removing outliers and reducing misdetections. Qingju Liu, Wenwu Wang 0001, Teófilo Emídio de Campos, Philip J. B. Jackson, Adrian Hilton 0001 |
IEEE Trans. Multim. | 4 |
| 2017 | 3D Room Geometry Reconstruction Using Audio-Visual SensorsabstractIn this paper we propose a cuboid-based air-tight indoor room geometry estimation method using combination of audio-visual sensors. Existing vision-based 3D reconstruction methods are not applicable for scenes with transparent or reflective objects such as windows and mirrors. In this work we fuse multi-modal sensory information to overcome the limitations of purely visual reconstruction for reconstruction of complex scenes including transparent and mirror surfaces. A full scene is captured by 360$^{\circ}$ cameras and acoustic room impulse responses (RIRs) recorded by a loudspeaker and compact microphone array. Depth information of the scene is recovered by stereo matching from the captured images and estimation of major acoustic reflector locations from the sound. The coordinate systems for audio-visual sensors are aligned into a unified reference frame and plane elements are reconstructed from audio-visual data. Finally cuboid proxies are fitted to the planes to generate a complete room model. Experimental results show that the proposed system generates complete representations of the room structures regardless of transparent windows, featureless walls and shiny surfaces. Hansung Kim 0001, Luca Remaggi, Philip J. B. Jackson, Filippo Maria Fazi, Adrian Hilton 0001 |
3DV | 3 |
| 2017 | Fast tagging of natural sounds using marginal co-regularizationabstractAutomatic and fast tagging of natural sounds in audio collections is a very challenging task due to wide acoustic variations, the large number of possible tags, the incomplete and ambiguous tags provided by different labellers. To handle these problems, we use a co-regularization approach to learn a pair of classifiers on sound and text. The first classifier maps low-level audio features to a true tag list. The second classifier maps actively corrupted tags to the true tags, reducing incorrect mappings caused by low-level acoustic variations in the first classifier, and to augment the tags with additional relevant tags. Training the classifiers is implemented using marginal co-regularization, pair of which draws the two classifiers into agreement by a joint optimization. We evaluate this approach on two sound datasets, Freefield1010 and Task4 of DCASE2016. The results obtained show that marginal co-regularization outperforms the baseline GMM in both efficiency and effectiveness. Qiang Huang 0007, Yong Xu 0004, Philip J. B. Jackson, Wenwu Wang 0001, Mark D. Plumbley |
ICASSP | 3 |
| 2017 | Speech reaction time measurements for the evaluation of audio-visual spatial coherenceabstractPerception of spatial coherence and boundaries for the fusion of spatially incoherent stimuli in the context of multimedia reproduction have mainly been studied with continuous judgment scales. As results vary greatly between different studies, reaction time measurements are proposed as an indirect measure of the perception of spatial incoherence. Word recognition tasks for spatially coherent and incoherent audio-visual stimuli were conducted. Results show that reaction time measurements are sensitive to audio-visual offsets and that they are affected at the smallest measured offset (5.1°). Hanne Stenzel, Philip J. B. Jackson, Jon Francombe |
QoMEX | 2 |
| 2017 | Acoustic Reflector Localization: Novel Image Source Reversion and Direct Localization MethodsabstractAcoustic reflector localization is an important issue in audio signal processing, with direct applications in spatial audio, scene reconstruction, and source separation. Several methods have recently been proposed to estimate the 3-D positions of acoustic reflectors given room impulse responses (RIRs). In this paper, we categorize these methods as “image-source reversion,” which localizes the image source before finding the reflector position, and “direct localization,” which localizes the reflector without intermediate steps. We present five new contributions. First, an onset detector, called the clustered dynamic programing projected phase-slope algorithm, is proposed to automatically extract the time of arrival for early reflections within the RIRs of a compact microphone array. Second, we propose an image-source reversion method that uses the RIRs from a single loudspeaker. It is constructed by combining an image source locator (the image source direction and range (ISDAR) algorithm), and a reflector locator (using the loudspeaker-image bisection (LIB) algorithm). Third, two variants of it, exploiting multiple loudspeakers, are proposed. Fourth, we present a direct localization method, the ellipsoid tangent sample consensus (ETSAC), exploiting ellipsoid properties to localize the reflector. Finally, systematic experiments on simulated and measured RIRs are presented, comparing the proposed methods with the state-of-the-art. ETSAC generates errors lower than the alternative methods compared through our datasets. Nevertheless, the ISDAR-LIB combination performs well and has a run time 200 times faster than ETSAC. Luca Remaggi, Philip J. B. Jackson, Philip Coleman, Wenwu Wang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Unsupervised Feature Learning Based on Deep Models for Environmental Audio TaggingabstractEnvironmental audio tagging aims to predict only the presence or absence of certain acoustic events in the interested acoustic scene. In this paper, we make contributions to audio tagging in two parts, respectively, acoustic modeling and feature learning. We propose to use a shrinking deep neural network (DNN) framework incorporating unsupervised feature learning to handle the multilabel classification task. For the acoustic modeling, a large set of contextual frames of the chunk are fed into the DNN to perform a multilabel classification for the expected tags, considering that only chunk (or utterance) level rather than frame-level labels are available. Dropout and background noise aware training are also adopted to improve the generalization capability of the DNNs. For the unsupervised feature learning, we propose to use a symmetric or asymmetric deep denoising auto-encoder (syDAE or asyDAE) to generate new data-driven features from the logarithmic Mel-filter banks features. The new features, which are smoothed against background noise and more compact with contextual information, can further improve the performance of the DNN baseline. Compared with the standard Gaussian mixture model baseline of the DCASE 2016 audio tagging challenge, our proposed method obtains a significant equal error rate (EER) reduction from 0.21 to 0.13 on the development set. The proposed asyDAE system can get a relative 6.7% EER reduction compared with the strong DNN baseline on the development set. Finally, the results also show that our approach obtains the state-of-the-art performance with 0.15 EER on the evaluation set of the DCASE 2016 audio tagging task while EER of the first prize of this challenge is 0.17. Yong Xu 0004, Qiang Huang 0007, Wenwu Wang 0001, Peter Foster, Siddharth Sigtia, Philip J. B. Jackson, Mark D. Plumbley |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2016 | Predicting Binaural Speech Intelligibility from Signals Estimated by a Blind Source Separation AlgorithmabstractState-of-the-art binaural objective intelligibility measures (OIMs) require individual source signals for making intelligibility predictions, limiting their usability in real-time online operations. This limitation may be addressed by a blind source separation (BSS) process, which is able to extract the underlying sources from a mixture. In this study, a speech source is presented with either a stationary noise masker or a fluctuating noise masker whose azimuth varies in a horizontal plane, at two speech-to-noise ratios (SNRs). Three binaural OIMs are used to predict speech intelligibility from the signals separated by a BSS algorithm. The model predictions are compared with listeners' word identification rate in a perceptual listening experiment. The results suggest that with SNR compensation to the BSS-separated speech signal, the OIMs can maintain their predictive power for individual maskers compared to their performance measured from the direct signals. It also reveals that the errors in SNR between the estimated signals are not the only factors that decrease the predictive accuracy of the OIMs with the separated signals. Artefacts or distortions on the estimated signals caused by the BSS algorithm may also be concerns. Qingju Liu, Philip J. B. Jackson, Wenwu Wang 0001 |
INTERSPEECH | 3 |
| 2015 | IVA algorithms using a multivariate Student's t source prior for speech source separation in real room environmentsabstractThe independent vector analysis (IVA) algorithm employs a multivariate source prior to retain the dependency between different frequency bins of each source and thereby avoids the permutation problem that is inherent to blind source separation (BSS). In this paper, a multivariate Student's t distribution is adopted as the source prior, which because of its heavy tail nature can better model the large amplitude information in the frequency bins. Therefore it can improve the separation performance and the convergence speed of the IVA and fast version of the IVA (FastIVA) algorithms as compared with the IVA algorithm based on another multivariate super Gaussian source prior. Separation performance with real binaural room impulse responses (BRIRs) is evaluated by detailed simulation studies when using the different source priors, and the experimental results confirm that the IVA and the FastIVA with the proposed multivariate Student's t source prior can consistently achieve improved and faster separation performance. Waqas Rafique, Syed M. Naqvi, Philip J. B. Jackson, Jonathon A. Chambers |
ICASSP | 3 |
| 2015 | A 3D model for room boundary estimationabstractEstimating the geometric properties of an indoor environment through acoustic room impulse responses (RIRs) is useful in various applications, e.g., source separation, simultaneous localization and mapping, and spatial audio. Previously, we developed an algorithm to estimate the reflector's position by exploiting ellipses as projection of 3D spaces. In this article, we present a model for full 3D reconstruction of environments. More specifically, the three components of the previous method, respectively, MUSIC for direction of arrival (DOA) estimation, numerical search adopted for reflector estimation and the Hough transform to refine the results, are extended for 3D spaces. A variation is also proposed using RANSAC instead of the numerical search and the Hough transform wich significantly reduces the run time. Both methods are tested on simulated and measured RIR data. The proposed methods perform better than the baseline, reducing the estimation error. Luca Remaggi, Philip J. B. Jackson, Wenwu Wang 0001, Jonathon A. Chambers |
ICASSP | 2 |
| 2014 | Joint Mixing Vector and Binaural Model Based Stereo Source SeparationabstractIn this paper the mixing vector (MV) in the statistical mixing model is compared to the binaural cues represented by interaural level and phase differences (ILD and IPD). It is shown that the MV distributions are quite distinct while binaural models overlap when the sources are close to each other. On the other hand, the binaural cues are more robust to high reverberation than MV models. According to this complementary behavior we introduce a new robust algorithm for stereo speech separation which considers both additive and convolutive noise signals to model the MV and binaural cues in parallel and estimate probabilistic time-frequency masks. The contribution of each cue to the final decision is also adjusted by weighting the log-likelihoods of the cues empirically. Furthermore, the permutation problem of the frequency domain blind source separation (BSS) is addressed by initializing the MVs based on binaural cues. Experiments are performed systematically on determined and underdetermined speech mixtures in five rooms with various acoustic properties including anechoic, highly reverberant, and spatially-diffuse noise conditions. The results in terms of signal-to-distortion-ratio (SDR) confirm the benefits of integrating the MV and binaural cues, as compared with two state-of-the-art baseline algorithms which only use MV or the binaural cues. Atiyeh Alinaghi, Philip J. B. Jackson, Qingju Liu, Wenwu Wang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Spatial and coherence cues based time-frequency masking for binaural reverberant speech separationabstractMost of the binaural source separation algorithms only consider the dissimilarities between the recorded mixtures such as interaural phase and level differences (IPD, ILD) to classify and assign the time-frequency (T-F) regions of the mixture spectrograms to each source. However, in this paper we show that the coherence between the left and right recordings can provide extra information to label the T-F units from the sources. This also reduces the effect of reverberation which contains random reflections from different directions showing low correlation between the sensors. Our algorithm assigns the T-F regions into original sources based on weighted combination of IPD, ILD, the mixing vector models and the estimated interaural coherence (IC) between the left and right recordings. The binaural room impulse responses measured in four rooms with various acoustic conditions have been used to evaluate the performance of the proposed method which shows an average improvement of more than 2.23 dB in signal-to-distortion ratio (SDR) in room D with T60= 0.89 s over the state-of-the-art algorithms. Atiyeh Alinaghi, Wenwu Wang 0001, Philip J. B. Jackson |
ICASSP | 3 |
| 2012 | Use of bimodal coherence to resolve the permutation problem in convolutive BSS
Qingju Liu, Wenwu Wang 0001, Philip J. B. Jackson |
Signal Process. | 3 |
| 2011 | Integrating binaural cues and blind source separation method for separating reverberant speech mixturesabstractThis paper presents a new method for reverberant speech separation, based on the combination of binaural cues and blind source separation (BSS) for the automatic classification of the time-frequency (T-F) units of the speech mixture spectrogram. The main idea is to model interaural phase difference, interaural level difference and frequency bin-wise mixing vectors by Gaussian mixture models for each source and then evaluate that model at each T-F point and assign the units with high probability to that source. The model parameters and the assigned regions are refined iteratively using the Expectation-Maximization (EM) algorithm. The proposed method also addresses the permutation problem of the frequency domain BSS by initializing the mixing vectors for each frequency channel. The EM algorithm starts with binaural cues and after a few iterations the estimated probabilistic mask is used to initialize and re-estimate the mixing vector model parameters. We performed experiments on speech mixtures, and showed an average of about 0.8 dB improvement in signal-to-distortion (SDR) over the binaural only baseline. Atiyeh Alinaghi, Wenwu Wang 0001, Philip J. B. Jackson |
ICASSP | 3 |
| 2010 | Bimodal coherence based scale ambiguity cancellation for target speech extraction and enhancementabstractWe present a novel method for extracting target speech from au-ditory mixtures using bimodal coherence, which is statistically characterised by a Gaussian mixture modal (GMM) in the off-line training process, using the robust features obtained from the audio-visual speech. We then adjust the ICA-separated spectral components using the bimodal coherence in the time-frequency domain, to mitigate the scale ambiguities in different frequency bins. We tested our algorithm on the XM2VTS database, and the results show the performance improvement with our pro-posed algorithm in terms of signal to interference ratio (SIR) measurements. Index Terms: speech extraction, bimodal coherence, audio-visual, Gaussian mixture model (GMM), independent compo- Qingju Liu, Wenwu Wang 0001, Philip J. B. Jackson |
INTERSPEECH | 3 |
| 2009 | Statistical identification of articulation constraints in the production of speech
Philip J. B. Jackson, Veena D. Singampalli |
Speech Commun. | 1 |
| 2008 | Parallel model combination and word recognition in soccer audioabstractAudio from broadcast soccer can be used for identifying highlights from the game. Audio cues derived from these sources provide valuable information about game events, as can the detection of key words used by the commentators. In this paper we interpret the feasibility of incorporating both commentator word recognition and information about the additive background noise in an HMM structure. A limited set of audio cues, which have been extracted from data collected from the 2006 FIFA World Cup, are used to create an extension to the Aurora-2 database. The new database is then tested with various PMC models and compared to the standard baseline, clean and multi-condition training methods. It is found that incorporating SNR and noise type information into the PMC process is beneficial to recognition performance. Jack H. Longton, Philip J. B. Jackson |
ICME | 2 |
| 2007 | Statistical identification of critical, dependent and redundant articulatorsabstractA compact, data-driven statistical model for identifying roles played by articulators in production of English phones using 1D and 2D articulatory data is presented. Articulators critical in production of each phone were identified and were used to predict the pdfs of dependent articulators based on the strength of articulatory correlations. The performance of the model is evaluated on MOCHA database using proposed and exhaustive search techniques and the results of synthesised trajectories presented. Index Terms: coarticulation, speech production, articulatory modeling, critical articulator Veena D. Singampalli, Philip J. B. Jackson |
INTERSPEECH | 2 |
| 2007 | Visual analysis of lip coarticulation in VCV utterancesabstractThis paper presents an investigation of the visual variation on the bilabial plosive consonant /p/ in three coarticulation contexts. The aim is to provide detailed ensemble analysis to assist coarticulation modelling in visual speech synthesis. The underlying dynamics of labeled visual speech units, represented as lip shape, from symmetric VCV utterances, is investigated. Variation in lip dynamics is quantitively and qualitatively analyzed. This analysis shows that there are statistically significant differences in both the lip shape and trajectory during coarticulation. Aseel Turkmani, Adrian Hilton 0001, Philip J. B. Jackson, James D. Edge |
INTERSPEECH | 3 |
| 2006 | Enhancement of harmonic content of speech based on a dynamic programming pitch tracking algorithmabstractFor pitch tracking of a single speaker, a common requirement is to find the optimal path through a set of voiced or voiceless pitch estimates over a sequence of time frames.Dynamic programming (DP) algorithms have been applied before to this problem.Here, the pitch candidates are provided by a multi-channel autocorrelation-based estimator, and DP is extended to pitch tracking of multiple concurrent speakers.We use the resulting pitch information to enhance harmonic content in noisy speech and to obtain separations of target from interfering speech. Mark R. Every, Philip J. B. Jackson |
INTERSPEECH | 2 |
| 2005 | Amplitude modulation of frication noise by voicing saturatesabstractThe two distinct sound sources comprising voiced frication, voicing and frication, interact. One effect is that the periodic source at the glottis modulates the amplitude of the frication source originating in the vocal tract above the constriction. Voicing strength and modulation depth for frication noise were measured for sustained English voiced fricatives using high-pass filtering, spectral analysis in the modulation (envelope) domain, and a variable pitch compensation procedure. Results show a positive relationship between strength of the glottal source and modulation depth at voicing strengths below 66 dB SPL, at which point the modulation index was approximately 0.5 and saturation occurred. The alveolar [z] was found to be more modulated than other fricatives. Jonathan Pincas, Philip J. B. Jackson |
INTERSPEECH | 2 |
| 2005 | A multiple-level linear/linear segmental HMM with a formant-based intermediate layer
Martin J. Russell, Philip J. B. Jackson |
Comput. Speech Lang. | 2 |
| 2003 | Covariation and weighting of harmonically decomposed streams for ASRabstractDecomposition of speech signals into simultaneous streams of periodic and aperiodic information has been successfully applied to speech analysis, enhancement, modification and recently recognition. This paper examines the effect of different weightings of the two streams in a conventional HMM system in digit recognition tests on the Aurora 2.0 database. Comparison of the results from using matched weights during training showed a small improvement of approximately 10% relative to unmatched ones, under clean test conditions. Principal component analysis of the covariation amongst the periodic and aperiodic features indicated that only 45 (51) of the 78 coefficients were required to account for 99% of the variance, for clean (multi-condition) training, which yielded an 18.4% (10.3%) absolute increase in accuracy with respect to the baseline. These findings provide further evidence of the potential for harmonically-decomposed streams to improve performance and substantially to enhance recognition accuracy in noise. Philip J. B. Jackson, David M. Moreno, Martin J. Russell, Javier Hernando |
INTERSPEECH | 1 |
| 2003 | The effect of an intermediate articulatory layer on the performance of a segmental HMMabstractWe present a novel multi-level HMM in which an intermedi-ate ‘articulatory ’ representation is included between the state and surface-acoustic levels. A potential difficulty with such a model is that advantages gained by the introduction of an ar-ticulatory layer might be compromised by limitations due to an insufficiently rich articulatory representation, or by com-promises made for mathematical or computational expediency. This paper decribes a simple model in which speech dynam-ics are modelled as linear trajectories in a formant-based ‘ar-ticulatory ’ layer, and the articulatory-to-acoustic mappings are linear. Phone classification results for TIMIT are presented for monophone and triphone systems with a phone-level syn-tax. The results demonstrate that provided the intermediate rep-resentation is sufficiently rich, or a sufficiently large number of phone-class-dependent articulatory-to-acoustic mapping are employed, classification performance is not compromised. 1. Martin J. Russell, Philip J. B. Jackson |
INTERSPEECH | 2 |
| 2002 | Models of speech dynamics in a segmental-HMM recognizer using intermediate linear representationsabstractA theoretical and experimental analysis of a simple multilevel segmental HMM is presented in which the relationship between symbolic (phonetic) and surface (acoustic) representations of speech is regulated by an intermediate (articulatory) layer, where speech dynamics are modeled using linear trajectories. Three formant-based parameterizations and measured articulatory positions are considered as intermediate representations, from the TIMIT and MOCHA corpora respectively. The articulatory-to-acoustic mapping was performed by between 1 and 49 linear transformations. Results of phone-classification experiments demonstrate that, by appropriate choice of intermediate parameterization and mappings, it is possible to achieve close to optimal performance. Philip J. B. Jackson, Martin J. Russell |
INTERSPEECH | 1 |
| 2001 | Pitch-scaled estimation of simultaneous voiced and turbulence-noise components in speechabstractAlmost all speech contains simultaneous contributions from more than one acoustic source within the speaker's vocal tract. In this paper, we propose a method-the pitch-scaled harmonic filter (PSHF)-which aims to separate the voiced and turbulence-noise components of the speech signal during phonation, based on a maximum likelihood approach. The PSHF outputs periodic and aperiodic components that are estimates of the respective contributions of the different types of acoustic source. It produces four reconstructed time series signals by decomposing the original speech signal, first, according to amplitude, and then according to power of the Fourier coefficients. Thus, one pair of periodic and aperiodic signals is optimized for subsequent time-series analysis, and another pair for spectral analysis. The performance of the PSHF algorithm is tested on synthetic signals, using three forms of disturbance (jitter, shimmer and additive noise), and the results were used to predict the performance on real speech. Processing recorded speech examples elicited latent features from the signals, demonstrating the PSHF's potential for analysis of mixed-source speech. Philip J. B. Jackson, Christine H. Shadle |
IEEE Trans. Speech Audio Process. | 1 |
| 2000 | Performance of the pitch-scaled harmonic filter and applications in speech analysisabstractThe pitch-scaled harmonic filter (PSHF) is a technique for decomposing speech signals into their voiced and unvoiced constituents. In this paper, we evaluate its ability to reconstruct the time series of the two components accurately using a variety of synthetic, speech-like signals, and discuss its performance. These results determine the degree of confidence that can be expected for real speech signals: typically, 5 dB improvement in the signal-to-noise ratio of the harmonic component and approximately 5 dB more than the initial harmonics-to-noise ratio (HNR) in the anharmonic component. A selection of the analysis opportunities that the decomposition offers is demonstrated on speech recordings, including dynamic HNR estimation and separate linear prediction analyses of the two components. These new capabilities provided by the PSHF can facilitate discovering previously hidden features and investigating interactions of unvoiced sources, such as frication, with voicing. Philip J. B. Jackson, Christine H. Shadle |
ICASSP | 1 |