EDBT 2026 Demo / reviewers in the wild / expert
Craig T. Jin
dblp:03/148
· DBLP profile ↗
57ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0003-4636-753XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 2 first-author · 10 since 2021Systems, architecture and hardware · 15 · 1 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 7 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TactDeform: Finger Pad Deformation Inspired Spatial Tactile Feedback for Virtual Geometry ExplorationabstractSpatial tactile feedback can enhance the realism of geometry exploration in virtual reality applications. Current vibrotactile approaches often face challenges with the spatial and temporal resolution needed to render different 3D geometries. Inspired by the natural deformation of finger pads when exploring 3D objects and surfaces, we propose TactDeform, a parametric approach to render spatio-temporal tactile patterns using a finger-worn electro-tactile interface. The system dynamically renders electro-tactile patterns based on both interaction contexts (approaching, contact, and sliding) and geometric contexts (geometric features and textures), emulating deformations that occur during real-world touch exploration. Results from a user study \rr{(N=24)} show that the proposed approach enabled high texture discrimination and geometric feature identification compared to a baseline. Informed by results from a free 3D-geometry exploration phase, we provide insights that can inform future tactile interface designs. Yihao Dong, Praneeth Bimsara Perera, Chin-Teng Lin, Craig T. Jin, Anusha Withana |
CHI | 4 |
| 2025 | TB-HSU: Hierarchical 3D Scene Understanding with Contextual AffordancesabstractThe concept of function and affordance is a critical aspect of 3D scene understanding and supports task-oriented objectives. In this work, we develop a model that learns to structure and vary functional affordance across a 3D hierarchical scene graph representing the spatial organization of a scene. The varying functional affordance is designed to integrate with the varying spatial context of the graph. More specifically, we develop an algorithm that learns to construct a 3D hierarchical scene graph (3DHSG) that captures the spatial organization of the scene. Starting from segmented object point clouds and object semantic labels, we develop a 3DHSG with a top node that identifies the room label, child nodes that define local spatial regions inside the room with region-specific affordances, and grand-child nodes indicating object locations and object-specific affordances. To support this work, we create a custom 3DHSG dataset that provides ground truth data for local spatial regions with region-specific affordances and also object-specific affordances for each object. We employ a Transformer Based Hierarchical Scene Understanding (TB-HSU) model to learn the 3DHSG. We use a multi-task learning framework that learns both room classification and learns to define spatial regions within the room with region-specific affordances. Our work improves on the performance of state-of-the-art baseline models and shows one approach for applying transformer models to 3D scene understanding and the generation of 3DHSGs that capture the spatial organization of a room. The code and dataset are publicly available. Wenting Xu, Viorela Ila, Luping Zhou, Craig T. Jin |
AAAI | 4 |
| 2025 | Constrained LDDMM for Dynamic Vocal Tract Morphing: Integrating Volumetric and Real-Time MRI
Tharinda Piyadasa, Joan Glaunès, Amelia Gully, Michael Proctor, Kirrie J. Ballard, Tünde Szalay, Naeim Sanaei, Sheryl Foster, David Waddington, Craig T. Jin |
INTERSPEECH | 10 |
| 2025 | Rhotic Articulation in Australian English: Insights from MRI
Michael Proctor, Tünde Szalay, Tharinda Piyadasa, Craig T. Jin, Naeim Sanaei, Amelia Gully, David Waddington, Sheryl Foster, Kirrie J. Ballard |
INTERSPEECH | 4 |
| 2025 | Lateral Channel Formation in Australian English /l/: Insights from Magnetic Resonance Imaging
Tünde Szalay, Michael Proctor, Amelia Gully, Tharinda Piyadasa, Craig T. Jin, David Waddington, Naeim Sanaei, Sheryl Foster, Kirrie J. Ballard |
INTERSPEECH | 5 |
| 2024 | Addressing Data Scarcity in Voice Disorder Detection with Self-Supervised ModelsabstractMachine learning (ML) has shown promising results in the field of voice disorder detection over the past decade. However, the diversity of recording conditions, audio content, languages, and the scarcity of examples for each of these combinations pose a challenge in building ML models that can reliably detect voice disorders. Recent advancements in Self-Supervised Learning (SSL) offer hope by leveraging large datasets to pretrain models and extract audio features with high resilience for downstream tasks.In this paper, we fairly exhaustively explore commonly used SSL model representations to assess their suitability for addressing the downstream task of voice disorder detection. Using a combination of Support Vector Machines (SVM) and feedforward Deep Neural Networks (DNN) we show: i) that the combination of vowels /a/,/i/, and /u/ perform better than individual vowels; ii) SSL-based features generalize well to out-of-domain databases, and iii) that while spectral features like MFCC perform equally well compared to SSL-based features when trained and tested on the same database, performances seems to deteriorate when training and testing across different databases. Rijul Gupta, Catherine J. Madill, Dhanshree R. Gunjawate, Duy Duong Nguyen, Craig T. Jin |
ICASSP | 5 |
| 2024 | Active Noise Control Over 3D Space with A Dynamic Noise SourceabstractSpatial Active noise control (ANC) systems are proposed to minimize the noise over a spatial region of interest around people’s heads by generating an anti-noise field with multiple microphones and loudspeakers. Recently, a realistic microphone geometry was designed to allow the system to monitor the residual noise field without restricting people’s head movement. However, this design, which utilizes the remote microphone technique using an observation filter (OF) for noise field modelling, has a drawback that the system is not robust against variations in the noise source location. In this paper, we address this drawback and present an improved spatial ANC system that overcomes this limitation. The proposed system enhances its adaptability by selectively employing the most suitable OF from a set of pre-modelled discriminator filter (DF), tailored to the acoustic environment of the system. We demonstrate that the proposed method maintains noise reduction performance over the region of interest with a dynamic noise source with changing position and signal through simulation and physical measurements. Huiyuan Sun, Craig T. Jin, Thushara D. Abhayapala, Prasanga N. Samarasinghe |
ICASSP | 2 |
| 2024 | From RIR to BRIR: A Sparse Recovery Beamforming Approach for Virtual Binaural Sound RenderingabstractThe creation of a spatial sound scene through binaural rendering draws increasing research interest, given the rising application of virtual reality and augmented reality. High-fidelity augmented reality audio requires accurate room acoustic simulation. Typically, this is achieved with binaural rendering of spatial sounds using recording of a Higher-Order Microphone (HOM) and Head-Related Transfer Functions (HRTFs) based simulations. However, HOM often fails to deliver full immersion due to the limited order of recording. In this paper, we introduce a Binaural Room Impulse Response (BRIR) calculation method based on HOM recordings of Room Impulse Responses (RIRs) to improve binaural sound rendering. The proposed method enables a more accurate analysis and rendering of the directive components in the limited-order spherical harmonic (SH) recording of RIRs with sparse recovery beamforming. Results with measured RIRs using a HOM indicate that the proposed method achieves a more accurate BRIR calculation than the conventional SH domain method. Huiyuan Sun, Howe Yuan Zhu, Minh T. D. Nguyen, Chin-Teng Lin, Craig T. Jin |
ICASSP | 6 |
| 2024 | Towards Speech Classification from Acoustic and Vocal Tract data in Real-time MRI
Yaoyao Yue, Michael Proctor, Luping Zhou, Rijul Gupta, Tharinda Piyadasa, Amelia Gully, Kirrie J. Ballard, Craig T. Jin |
INTERSPEECH | 8 |
| 2024 | S$^{3}$CA: A Sparse Strip Spectral Correlation AnalyzerabstractThe spectral correlation density (SCD) is widely used to characterize cyclostationary signals and the strip spectral correlation analyzer (SSCA) is commonly used to estimate the SCD. Although the SSCA utilizes the fast Fourier transform (FFT) for computational efficiency, its real-time implementation still poses challenges as large input sizes are often involved. In this work, we present a sparse strip spectral correlation analyzer (S3CA) based on the sparse fast Fourier transform (SFFT). The S3CA approach involves computing a sparse, downsampled channel-data product (CDP) which is then passed to a modified SFFT implementation to obtain the spectral density. For an input of length 2 million samples, the S3CA is 30× faster than the conventional SSCA. Carol Jingyi Li, Richard Rademacher, David Boland, Craig T. Jin, Chad M. Spooner, Philip H. W. Leong |
IEEE Signal Process. Lett. | 4 |
| 2024 | Adjustable Coherent-to-Diffuse Power Estimator for Binaural Speech Enhancement in Multi-Talker EnvironmentsabstractThe binaural coherence-to-diffuse power ratio (CDR) estimate in reverberant environments is essential in many speech enhancement algorithms applied within hear-through systems. In this work, we propose a parameterised and adjustable binaural CDR estimator whose formulation is based on a geometrical interpretation of the short-time complex coherence function between binaural microphone signals. Conventional CDR estimators often distort the natural spectro-temporal behaviour of the noise field by relying on theoretical coherence models of the desired signal and/or diffuse noise field. Our proposed CDR estimator relies only on the observed spatial coherence and better preserves the natural characteristics of a binaural noise field. We demonstrate that the proposed CDR estimator can be used effectively for binaural dereverberation and denoising of broadside speech in multi-talker and noisy acoustic conditions and that it often outperforms state-of-the-art coherence-based methods for dereverberation and denoising. Furthermore, the adjustable parameter enables one to minimise the frequency-dependent estimation error of the binaural system in different environments. Reza Ghanavi, Craig T. Jin |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | HRTF Interpolation Using a Spherical Neural Process Meta-LearnerabstractSeveral individualization methods have recently been proposed to estimate a subject's Head-Related Transfer Function (HRTF) using convenient input modalities such as anthropometric measurements or pinnae photographs. There exists a need for adaptively correcting the estimation error committed by such methods using a few data point samples from the subject's HRTF, acquired using acoustic measurements or perceptual feedback. To facilitate this, we introduce a Convolutional Conditional Neural Process meta-learner specialized in HRTF error interpolation. In particular, the model includes a Spherical Convolutional Neural Network component to accommodate the spherical geometry of HRTF data. It also exploits potential symmetries between the HRTF's left and right channels about the median plane. In this work, we evaluate the proposed model's performance purely on time-aligned spectrum interpolation grounds under a simplified setup where a generic population-mean HRTF forms the initial estimates prior to corrections instead of individualized ones. The trained model achieves up to 3 dB relative error reduction compared to state-of-the-art interpolation methods despite being trained using only 85 subjects. This improvement translates up to nearly a halving of the data point count required to achieve comparable accuracy, in particular from 50 to 28 points to reach an average of -20 dB relative error per interpolated feature. Moreover, we show that the trained model provides well-calibrated uncertainty estimates. Accordingly, such estimates could inform the sequential decision problem of acquiring as few correcting HRTF data points as needed to meet a desired level of HRTF individualization accuracy. Etienne Thuillier, Craig T. Jin, Vesa Välimäki |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Fixed-point FPGA Implementation of the FFT Accumulation Method for Real-time Cyclostationary AnalysisabstractThe spectral correlation density (SCD) is an important tool in cyclostationary signal detection and classification. Even using efficient techniques based on the fast Fourier transform (FFT), real-time implementations are challenging because of the high computational complexity. A key dimension for computational optimization lies in minimizing the wordlength employed. In this article, we analyze the relationship between wordlength and signal-to-quantization noise in fixed-point implementations of the SCD function. A canonical SCD estimation algorithm, the FFT accumulation method (FAM) using fixed-point arithmetic, is studied. We derive closed-form expressions for SQNR and compare them at wordlengths ranging from 14 to 26 bits. The differences between the calculated SQNR and bit-exact simulations are less than 1 dB. Furthermore, an HLS-based FPGA design is implemented on a Xilinx Zynq UltraScale+ XCZU28DR-2FFVG1517E RFSoC. Using less than 25% of the logic fabric on the device, it consumes 7.7 W total on-chip power and has a power efficiency of 12.4 GOPS/W, which is an order of magnitude improvement over an Nvidia Tesla K40 graphics processing unit (GPU) implementation. In terms of throughput, it achieves 50 MS/sec, which is a speedup of 1.6 over a recent optimized FPGA implementation. Carol Jingyi Li, Xiangwei Li, Binglei Lou, Craig T. Jin, David Boland, Philip H. W. Leong |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2021 | Sparse Recovery Beamforming and Upscaling in the Ray SpaceabstractWe have been exploring the integration of sparse recovery methods into the ray space transform over the past years and now demonstrate the potential and benefits of beamforming and upscaling signals in the integrated ray space and sparse recovery domain. A primary advantage of the ray space approach derives from its robust ability to integrate information from multiple arrays and viewpoints. Nonetheless, for a given viewpoint, the ray space technique requires a dense array that can be divided into sub-arrays enabling the plenacoustic approach to signal processing. In this work, we explore a method to upscale an array beyond the limits imposed by the inter-microphone distances associated with the array and the concomitant spatial aliasing. In other words, sparse recovery enables one to synthesize or interpolate signals corresponding to an array with a greater number of microphones with a smaller inter-microphone distance. A critical issue is whether or not this interpolative synthesis actually improves array signal processing. This work shows that upscaling signals in the integrated ray space and sparse recovery domain can improve both source localization and separation. Shiduo Yu, Craig T. Jin, Fabio Antonacci, Augusto Sarti |
ICASSP | 2 |
| 2019 | Improved Multipath Time Delay Estimation Using Cepstrum SubtractionabstractWhen a motor-powered vessel travels past a fixed hydrophone in a multipath environment, a Lloyd's mirror constructive/destructive interference pattern is observed in the output spectrogram. The power cepstrum detects the periodic structure of the Lloyd's mirror pattern by generating a sequence of pulses (rahmonics) located at the fundamental quefrency (periodic time) and its multiples. This sequence is referred to here as the `rahmonic component' of the power cepstrum. The fundamental quefrency, which is the reciprocal of the frequency difference between adjacent interference fringes, equates to the multipath time delay. The other component of the power cepstrum is the non-rahmonic (extraneous) component, which combines with the rahmonic component to form the (total) power cepstrum. A data processing technique, termed `cepstrum subtraction', is described. This technique suppresses the extraneous component of the power cepstrum, leaving the rahmonic component that contains the desired multipath time delay information. This technique is applied to real acoustic recordings of motor-vessel transits in a shallow water environment, where the broadband noise radiated by the vessel arrives at the hydrophone via a direct ray path and a time-delayed multipath. The results show that cepstrum subtraction improves multipath time delay estimation by a factor of two for the at-sea experiment. Eric L. Ferguson, Stefan B. Williams, Craig T. Jin |
ICASSP | 3 |
| 2018 | Sound Source Localization in a Multipath Environment Using Convolutional Neural NetworksabstractThe propagation of sound in a shallow water environment is characterized by boundary reflections from the sea surface and sea floor. These reflections result in multiple (indirect) sound propagation paths, which can degrade the performance of passive sound source localization methods. This paper proposes the use of convolutional neural networks (CNNs) for the localization of sources of broadband acoustic radiated noise (such as motor vessels) in shallow water multipath environments. It is shown that CNNs operating on cepstrogram and generalized cross-correlogram inputs are able to estimate more reliably the instantaneous range and bearing of transiting motor vessels when the source localization performance of conventional passive ranging methods is degraded. The ensuing improvement in source localization performance is demonstrated using real data collected during an at-sea experiment. Eric L. Ferguson, Stefan B. Williams, Craig T. Jin |
ICASSP | 3 |
| 2018 | Considerations Regarding Individualization of Head-Related Transfer FunctionsabstractThis paper provides some considerations regarding using individualized head-related transfer functions for rendering binaural spatial audio over headphones. It briefly considers the degree of benefit that individualization may provide. It then examines the degree of variation existing within the ear morphology across listeners within the Sydney-York Morphological and Recording of Ears (SYMARE) database using kernel principal component analysis and the large deformation diffeomorphic metric mapping framework. The degree of variation across listeners in the directivity patterns associated with head-related transfer functions is also analyzed as a function of frequency. The variation in ear morphology is related to the variation in the directivity patterns using simple linear regression. Craig T. Jin, Reza Zolfaghari, Xian Long, Arun Sebastian, Shayikh Hossain, Joan Glaunès, Anthony I. Tew, Muhammad Shahnawaz, Augusto Sarti |
ICASSP | 1 |
| 2017 | Convolutional neural networks for passive monitoring of a shallow water environment using a single sensorabstractA cost effective approach to remote monitoring of protected areas such as marine reserves and restricted naval waters is to use passive sonar to detect, classify, localize, and track marine vessel activity (including small boats and autonomous underwater vehicles). Cepstral analysis of underwater acoustic data enables the time delay between the direct path arrival and the first multipath arrival to be measured, which in turn enables estimation of the instantaneous range of the source (a small boat). However, this conventional method is limited to ranges where the Lloyd's mirror effect (interference pattern formed between the direct and first multipath arrivals) is discernible. This paper proposes the use of convolutional neural networks (CNNs) for the joint detection and ranging of broadband acoustic noise sources such as marine vessels in conjunction with a data augmentation approach for improving network performance in varied signal-to-noise ratio (SNR) situations. Performance is compared with a conventional passive sonar ranging method for monitoring marine vessel activity using real data from a single hydrophone mounted above the sea floor. It is shown that CNNs operating on cepstrum data are able to detect the presence and estimate the range of transiting vessels at greater distances than the conventional method. Eric L. Ferguson, Rishi Ramakrishnan, Stefan B. Williams, Craig T. Jin |
ICASSP | 4 |
| 2017 | Kernel principal component analysis of the ear morphologyabstractThis paper describes features in the ear shape that change across a population of ears and explores the corresponding changes in ear acoustics. The statistical analysis conducted over the space of ear shapes uses a kernel principal component analysis (KPCA). Further, it utilizes the framework of large deformation diffeomorphic metric mapping and the vector space that is constructed over the space of initial momentums, which describes the diffeomorphic transformations from the reference template ear shape. The population of ear shapes examined by the KPCA are 124 left and right ear shapes from the SYMARE database that were rigidly aligned to the template (population average) ear. In the work presented here we show the morphological variations captured by the first two kernel principal components, and also show the acoustic transfer functions of the ears which are computed using fast multipole boundary element method simulations. Reza Zolfaghari, Nicolas Epain, Craig T. Jin, Joan Glaunès, Anthony I. Tew |
ICASSP | 3 |
| 2017 | FPGA Implementations of Kernel Normalised Least Mean Squares ProcessorsabstractKernel adaptive filters (KAFs) are online machine learning algorithms which are amenable to highly efficient streaming implementations. They require only a single pass through the data and can act as universal approximators, i.e. approximate any continuous function with arbitrary accuracy. KAFs are members of a family of kernel methods which apply an implicit non-linear mapping of input data to a high dimensional feature space, permitting learning algorithms to be expressed entirely as inner products. Such an approach avoids explicit projection into the feature space, enabling computational efficiency. In this paper, we propose the first fully pipelined implementation of the kernel normalised least mean squares algorithm for regression. Independent training tasks necessary for hyperparameter optimisation fill pipeline stages, so no stall cycles to resolve dependencies are required. Together with other optimisations to reduce resource utilisation and latency, our core achieves 161 GFLOPS on a Virtex 7 XC7VX485T FPGA for a floating point implementation and 211 GOPS for fixed point. Our PCI Express based floating-point system implementation achieves 80% of the core’s speed, this being a speedup of 10× over an optimised implementation on a desktop processor and 2.66× over a GPU. Nicholas J. Fraser, Duncan J. M. Moss, Julian Faraone, Stephen Tridgell, Craig T. Jin, Philip H. W. Leong |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2016 | Random projections for scaling machine learning on FPGAsabstractRandom projections have recently emerged as a powerful technique for large scale dimensionality reduction in machine learning applications. Crucially, the projection can be obtained from sparse probability distributions, enabling hardware implementations with little overhead. In this paper, we describe a Field-Programmable Gate Array (FPGA) implementation alongside a kernel adaptive filter (KAF) that is capable of reducing computational resources by introducing a controlled error term, achieving higher modelling capacity for given hardware resources. Empirical results involving classification, regression and novelty detection show that a 40% net increase in available resources and improvements in prediction accuracy is achievable for projections which halve the input vector length, enabling us to scale-up hardware implementations of KAF learning algorithms by at least a factor of 2. An implementation on a FPGA-based network card allows novelty detection of an 8× 24-bit input vector with latency of 404 ns, this being a 26-fold reduction compared to an Intel Core i5-2400 processor. Sean Fox, Stephen Tridgell, Craig T. Jin, Philip H. W. Leong |
FPT | 3 |
| 2016 | Generating a morphable model of earsabstractThis paper describes the generation of a morphable model for external ear shapes. The aim for the morphable model is to characterize an ear shape using only a few parameters in order to assist the study of morphoacoustics. The model is derived from a statistical analysis of a population of 58 ears from the SYMARE database. It is based upon the framework of large deformation diffeomorphic metric mapping (LDDMM) and the vector space that is constructed over the space of initial momentums describing the diffeomorphic transformations. To develop a morphable model using the LDDMM framework, the initial momentums are analyzed using a kernel based principal component analysis. In this paper, we examine the ability of our morphable model to construct test ear shapes not included in the principal component analysis. Reza Zolfaghari, Nicolas Epain, Craig T. Jin, Joan Glaunès, Anthony I. Tew |
ICASSP | 3 |
| 2016 | Spherical Harmonic Signal Covariance and Sound Field DiffusenessabstractCharacterizing sound field diffuseness has many practical applications, from room acoustics analysis to speech enhancement and sound field reproduction. In this paper, we investigate how spherical microphone arrays (SMAs) can be used to characterize diffuseness. Due to their specific geometry, SMAs are particularly well suited for analyzing the spatial properties of sound fields. In particular, the signals recorded by an SMA can be analyzed in the spherical harmonic (SH) domain, which has special and desirable mathematical properties when it comes to analyzing diffuse sound fields. We present a new measure of diffuseness, the COMEDIE diffuseness estimate, which is based on the analysis of the SH signal covariance matrix. This algorithm is suited for the estimation of diffuseness arising either from the presence of multiple sources distributed around the SMA or from the presence of a diffuse noise background. As well, we introduce the concept of a diffuseness profile, which consists in measuring the diffuseness for several SH orders simultaneously. Experimental results indicate that diffuseness profiles better describe the properties of the sound field than a single diffuseness measurement. Nicolas Epain, Craig T. Jin |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | A fully pipelined kernel normalised least mean squares processor for accelerated parameter optimisationabstractKernel adaptive filters (KAFs) are online machine learning algorithms which are amenable to highly efficient streaming implementations. They require only a single pass through the data during training and can act as universal approximators, i.e. approximate any continuous function with arbitrary accuracy. KAFs are members of a family of kernel methods which apply an implicit nonlinear mapping of input data to a high dimensional feature space, permitting learning algorithms to be expressed entirely as inner products. Such an approach avoids explicit projection into the feature space, enabling computational efficiency. In this paper, we propose the first fully pipelined floating point implementation of the kernel normalised least mean squares algorithm for regression. Independent training tasks necessary for parameter optimisation fill L cycles of latency ensuring the pipeline does not stall. Together with other optimisations to reduce resource utilisation and latency, our core achieves 160 GFLOPS on a Virtex 7 XC7VX485T FPGA, and the PCI-based system implementation is 70× faster than an optimised software implementation on a desktop processor. Nicholas J. Fraser, Duncan J. M. Moss, Stephen Tridgell, Craig T. Jin, Philip H. W. Leong |
FPL | 5 |
| 2015 | Super-resolution acoustic imaging using sparse recovery with spatial primingabstractIn this paper, we propose a new strategy to obtain superresolution maps of the sound field recorded by a spherical microphone array. In recent works, we have demonstrated that sparse recovery (SR) algorithms based on the minimisation of the lpnorm with 0pnorm when p<;1 is that it is a non-convex optimisation problem, thus it is likely that the algorithm converges to a local minimum. In this paper we show that we can improve the convergence of our SR acoustic imaging methods by providing, to the SR solver, priming information relating to the spatial location of the sound sources. This information can be acquired with a pre-processing, coarse analysis using standard blind source separation or direction-of-arrival techniques. Simulation results indicate that this approach can provide accurate estimates of the positions of multiple, simultaneous sound sources in the presence of noise or reverberation and even in an under-determined situation. Tahereh Noohi, Nicolas Epain, Craig T. Jin |
ICASSP | 3 |
| 2014 | Large Deformation Diffeomorphic Metric Mapping and Fast-Multipole Boundary Element Method provide new insights for Binaural acousticsabstractThis paper describes how Large Deformation Diffeomorphic Metric Mapping (LDDMM) can be coupled with a Fast Multipole (FM) Boundary Element Method (BEM) to investigate the relationship between morphological changes in the head, torso, and outer ears and their acoustic filtering (described by Head Related Transfer Functions, HRTFs). The LDDMM technique provides the ability to study and implement morphological changes in ear, head and torso shapes. The FM-BEM technique provides numerical simulations of the acoustic properties of an individual's head, torso, and outer ears. This paper describes the first application of LDDMM to the study of the relationship between a listener's morphology and a listener's HRTFs. To demonstrate some of the new capabilities provided by the coupling of these powerful tools, we morph the shape of a listener's ear, while keeping the torso and head shape essentially constant, and show changes in the acoustics. We validate the methodological framework by mapping the complete morphology of one listener to a target listener and obtaining the target listener's HRTFs. This work utilizes the data provided by the Sydney York Morphological and Acoustic Recordings of Ears (SYMARE) database. Reza Zolfaghari, Nicolas Epain, Craig T. Jin, Joan Glaunès, Anthony I. Tew |
ICASSP | 3 |
| 2014 | Design, Optimization and Evaluation of a Dual-Radius Spherical Microphone ArrayabstractSpherical Microphone Arrays (SMAs) constitute a powerful tool for analyzing the spatial properties of sound fields. However, the performance of SMA-based signal processing algorithms ultimately depends on the physical characteristics of the array. In particular, the range of frequencies over which an SMA provide rich spatial information is conditioned by the size of the array, the angular position of the sensors and other factors. In this work, we investigate the design of SMAs offering a wider frequency range of operation than that offered by conventional designs. To achieve this goal, microphones are distributed both on and at a distance from the surface of a rigid spherical baffle. The contributions of the paper are as follows. First, we present a general framework for modeling SMAs whose sensors are located at different distances from the array center and calculating optimal filters for the decomposition of the sound field into spherical harmonic modes. Second, we present an optimization method to design multi-radius SMAs with an optimally wide frequency range of operation given the total number of sensors available and target spatial resolution. Lastly, based on the optimization results, we built a prototype dual-radius SMA with 64 microphones. We present measurement results for the prototype microphone array and compare these results with theory. Craig T. Jin, Nicolas Epain, Abhaya Parthy |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | Creating the Sydney York Morphological and Acoustic Recordings of Ears DatabaseabstractThis paper introduces the process for creating the Sydney York Morphological and Acoustic Recordings of Ears (SYMARE) database. The SYMARE database supports research exploring the relationship between the morphology of human outer ears and their acoustic filtering properties-a relationship that is viewed by many as holding the key to human spatial hearing and the future of 3D personal audio. The SYMARE database is comprised of acoustically measured head-related impulse responses for 61 listeners (48 male/13 female), multiple high-resolution surface mesh models (upper torso, head and ears) for these listeners obtained from magnetic resonance imaging (MRI) data, and the corresponding simulated HRIR data for these listeners generated using the Fast Multipole Boundary Element Method (FM-BEM). In this work, we compare acoustically measured HRIR data for 61 listeners with the listeners' corresponding simulated HRIR data generated using the FM-BEM. Craig T. Jin, Pierre Guillon 0002, Nicolas Epain, Reza Zolfaghari, André van Schaik, Anthony I. Tew, Carl Hetherington, Jonathan Thorpe |
IEEE Trans. Multim. | 1 |
| 2013 | Super-resolution sound field imaging with sub-space pre-processingabstractSpherical microphone arrays are a powerful tool for sound field analysis. In previous work, we have shown that sparse recovery can be used to arbitrarily increase the resolution of the sound field recorded by a spherical microphone array. Because these super-resolution techniques rely on the assumption that the sound field results from a few dominant plane waves, they are not robust to the presence of noise or reverberation. In this paper we propose a simple method to separate the sound field into a directional component and a diffuse component prior to applying sparse recovery techniques. Simulations show that this pre-processing could dramatically improve the results of sparse recovery in noisy or reverberant environments. Nicolas Epain, Craig T. Jin |
ICASSP | 2 |
| 2013 | Direction of arrival estimation for spherical microphone arrays by combination of independent component analysis and sparse recoveryabstractSpherical microphone arrays provide a powerful tool for examining source localization and direction of arrival (DOA) estimation in the spherical harmonic domain. In previous work, we have investigated applying instantaneous independent component analysis (ICA) or sparse recovery separately in the spherical harmonic domain for DOA estimation. These algorithms work reasonably well, but rely on different signal characteristics: namely statistical independence or the spatial distribution of sources. In this paper, we describe methods to combine the ICA and sparse recovery algorithms to improve DOA estimation. The simulation results indicate that combining ICA and sparse recovery leads to more robust DOA estimation. Tahereh Noohi, Nicolas Epain, Craig T. Jin |
ICASSP | 3 |
| 2013 | A super-resolution beamforming algorithm for spherical microphone arrays using a compressed sensing approachabstractIn this paper, we present a novel beamforming algorithm that is designed for spherical microphone arrays and formulated in the spherical harmonic domain. The proposed algorithm employs sparse recovery, a compressed sensing technique, and assumes the position of the source signals are unknown. A formal listening test was conducted to evaluate the performance of the proposed algorithm and the results indicate the effectiveness of the proposed algorithm. Ping Kun Tony Wu, Nicolas Epain, Craig T. Jin |
ICASSP | 3 |
| 2013 | Structural Similarity Analysis of Modulation for audio quality assessmentabstractThis paper proposes an improved structural similarity method for audio quality assessment, which is Structural Similarity Analysis of Modulation (SSAM). Different from original structural similarity index measure, we introduce the analysis of structural similarity of modulation together with Computational Auditory Signal-processing and Perception (CASP) model. Audio features from CASP are extended to three dimensions: time, frequency and modulation spectrum. The combined architecture of ITU-R BS.1387-1 and proposed SSAM is given in this paper. Our proposed estimation system not only shows highly correlated with subjective results but also overcomes the shortage of ITU-R BS.1387-1 that only suitable for small impaired audio. Craig T. Jin |
ICASSP | 3 |
| 2012 | A frequency-domain algorithm to upscale ambisonic sound scenesabstractIn this paper, a novel algorithm for upscaling ambisonic sound scenes in the frequency domain is presented. This algorithm makes use of compressed sensing techniques to calculate a set of upscaling filters. These filters are then used to increase the spherical harmonic order of a set of ambisonic sound signals to higher orders. Upscaled ambisonic sound scenes have a greater spatial resolution, which allows more loudspeakers to be used during the playback, resulting in a larger sweet spot and improved sound quality. A formal listening test was conducted to evaluate the perceptual quality of sound fields reproduced using this technique. Results show that the proposed algorithm significantly improves the perceptual fidelity of the sound field reproduction, in comparison to classical ambisonic methods. Andrew Wabnitz, Nicolas Epain, Craig T. Jin |
ICASSP | 3 |
| 2012 | A dereverberation algorithm for spherical microphone arrays using compressed sensing techniquesabstractIn this paper, we present a novel multichannel dereverberation algorithm that enhances a target signal in a reverberant environment. The proposed algorithm is designed for a spherical microphone array and formulated in the spherical harmonic domain. The algorithm employs sparse recovery, a compressed sensing technique, to estimate the position of the target signal and its early reflections. Room impulse responses are obtained according to the estimations and the MINT (the multiple-input/output inverse-filtering theorem) is used to calculate the inverse filters. The performance of the proposed method is evaluated using computer simulation and our results indicate the effectiveness of the proposed dereverberation algorithm. Ping Kun Tony Wu, Nicolas Epain, Craig T. Jin |
ICASSP | 3 |
| 2012 | Creating the Sydney York Morphological and Acoustic Recordings of Ears DatabaseabstractThis paper introduces the process for creating the Sydney York Morphological and Acoustic Recordings of Ears (SYMARE) database. The SYMARE database supports research exploring the relationship between the morphology of human outer ears and their acoustic filtering properties - a relationship that is viewed by many as holding the key to human spatial hearing and the future of 3D personal audio. The SYMARE database is comprised of acoustically measured head-related impulse responses for 60 listeners, multiple high-resolution surface mesh models (upper torso, head and ears) for these listeners obtained from magnetic resonance imaging (MRI) data, and the corresponding simulated HRIR data for these listeners generated using the Fast Multipole Boundary Element Method (FM-BEM). In this work, we compare acoustically measured HRIR data for ten listeners with the listeners' corresponding simulated HRIR data generated using the FM-BEM. Pierre Guillon 0002, Reza Zolfaghari, Nicolas Epain, André van Schaik, Craig T. Jin, Carl Hetherington, Jonathan Thorpe, Anthony I. Tew |
ICME | 5 |
| 2011 | Spiking neural network-based auto-associative memory using FPGA interconnect delaysabstractThis paper describes the design of an auto-associative memory based on a spiking neural network (SNN). The architecture is able to effectively utilize the massive interconnect resources available in FPGA architectures as a good match to the axons in biological neural networks. A complete implementation of the memory on a single FPGA is presented. The signal processing circuitry is composed from simple, parallel building blocks and the training logic is implemented using an on-chip soft processor. Chong H. Ang, Craig T. Jin, Philip H. W. Leong, André van Schaik |
FPT | 2 |
| 2011 | Time domain reconstruction of spatial sound fields using compressed sensingabstractA novel technique for time domain spatial sound reproduction using compressed sensing is presented. The presented technique is based on the application of compressed sensing theory, which is used to improve the accuracy of the reconstructed sound field. In addition, singular value decomposition is also applied, which acts to significantly reduce the size of the data set to process, thus making it efficient and realisable for real-time applications. Results are presented from the preliminary performance evaluation of the compressed sensing technique in comparison to the Higher Order Ambisonic reconstruction technique. Andrew Wabnitz, Nicolas Epain, André van Schaik, Craig T. Jin |
ICASSP | 4 |
| 2011 | A programmable axonal propagation delay circuit for time-delay spiking neural networksabstractWe present an implementation of a programmable axonal propagation delay circuit which uses one first-order log-domain low-pass filter. Delays may be programmed in the 5-50ms range. It is designed to be a building block for time-delay spiking neural networks. It consists of a leaky-integrate-and-fire core, a spike generator circuit, and a delay adaptation circuit. Runchun Wang, Craig T. Jin, Alistair Lee McEwan, André van Schaik |
ISCAS | 2 |
| 2011 | Blind Image Watermarking Using a Sample Projection ApproachabstractThis paper presents a robust image watermarking scheme based on a sample projection approach. While we consider the human visual system in our watermarking algorithm, we use the low-frequency components of image blocks for data hiding to obtain high robustness against attacks. We use four samples of the approximation coefficients of the image blocks to construct a line segment in the 2-D space. The slope of this line segment, which is invariant to the gain factor, is employed for watermarking purpose. We embed the watermarking code by projecting the line segment on some specific lines according to message bits. To design a maximum likelihood decoder, we compute the distribution of the slope of the embedding line segment for Gaussian samples. The performance of the proposed technique is analytically investigated and verified via several simulations. Experimental results confirm the validity of our model and its high robustness against common attacks in comparison with similar watermarking techniques that are invariant to the gain attack. Mohammad Ali Akhaee, Sayed Mohammad Ebrahim Sahraeian, Craig T. Jin |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2010 | Investigating the implications of outer hair cell connectivity using a silicon cochleaabstractIn this paper we present results from several implementations of silicon cochleae whose dynamics are governed by the Hopf equation. These silicon cochleae exhibit the majority of active, nonlinear characteristics of the biological cochlea such as large-signal compression, two-tone suppression, the creation of distortion products and so forth. Here we explore the coupling between resonant sections of the basilar membrane to investigate phenomena such as masking and the characteristic frequency response curve of the cochlea at a particular place along the basilar membrane. We see that the interaction of resonant sections can account for these phenomena and that we can use these observations to partially explain the connectivity of the afferent and efferent fibres to the outer hair cells. This work not only gives us valuable insight into the dynamical behaviour of the early auditory system but it also highlights the benefits of building circuits of these complex systems in order to produce models whose parameters can be tuned and whose outputs can be observed and measured in realtime. Tara J. Hamilton, Jonathan Tapson, Craig T. Jin, André van Schaik |
ISCAS | 3 |
| 2010 | A log-domain implementation of the Izhikevich neuron modelabstractWe present an implementation of the Izhikevich neuron model which uses two first-order log-domain low-pass filters and two translinear multipliers. The neuron consists of a leaky-integrate-and-fire core, a slow adaptive state variable and quadratic positive feedback. Simulation results show that this neuron can emulate different spiking behaviours observed in biological neurons. André van Schaik, Craig T. Jin, Alistair Lee McEwan, Tara J. Hamilton |
ISCAS | 2 |
| 2010 | A log-domain implementation of the Mihalas-Niebur neuron modelabstractWe present an electronic neuron that uses first-order log-domain low-pass filters to implement the Mihalas-Niebur model. The neuron consists of a leaky-integrate-and-fire core and building blocks to implement an adaptive threshold and spike induced currents. Simulation results show that this modular neuron can emulate different spiking behaviours observed in biological neurons. André van Schaik, Craig T. Jin, Alistair Lee McEwan, Tara J. Hamilton, Stefan Mihalas, Ernst Niebur |
ISCAS | 2 |
| 2009 | Acoustic holography with a concentric rigid and open spherical microphone arrayabstractWe present a new method and performance data related to volumetric acoustic intensity imaging using a spherical microphone array (SMA) consisting of a dual, concentric rigid and open SMA. The dual, concentric array was designed to improve the frequency range of a standard SMA. We apply standard techniques associated with interior spherical near-field acoustic holography (NAH) and, in particular, consider issues related to the optimal use of information from both arrays for NAH projection and the advantages that thus accrue from utilising a dual, concentric SMA. Abhaya Parthy, Craig T. Jin, André van Schaik |
ICASSP | 2 |
| 2009 | Sound localisation with a silicon cochlea pairabstractA neuromorphic sound localisation system is proposed. It employs two microphones and a pair of silicon cochleae with address event interface for front-end processing. This allows subsequent processing to be implemented with spike-based algorithms. The system is adaptive and supports online learning. Its localisation capability was tested with white noise and pure tone stimuli, with an average error of around 3° in the −45° to 45° range. André van Schaik, Craig T. Jin |
ICASSP | 3 |
| 2009 | A First-Order Nonhomogeneous Markov Model for the Response of Spiking Neurons Stimulated by Small Phase-Continuous SignalsabstractWe present a first-order nonhomogeneous Markov model for the interspike-interval density of a continuously stimulated spiking neuron. The model allows the conditional interspike-interval density and the stationary interspike-interval density to be expressed as products of two separate functions, one of which describes only the neuron characteristics and the other of which describes only the signal characteristics. The approximation shows particularly clearly that signal autocorrelations and cross-correlations arise as natural features of the interspike-interval density and are particularly clear for small signals and moderate noise. We show that this model simplifies the design of spiking neuron cross-correlation systems and describe a four-neuron mutual inhibition network that generates a cross-correlation output for two input signals. Jonathan Tapson, Craig T. Jin, André van Schaik, Ralph Etienne-Cummings |
Neural Comput. | 2 |
| 2008 | A 2-D silicon cochlea with an improved automatic quality factor control-loopabstractIn this paper we present a 2-D silicon cochlea which includes an automatic quality factor control (AQC) loop. This control-loop is an improved version of that presented in [1] where the control-loop imposes both a ceiling and a threshold level on the output amplitude of the basilar membrane (BM) resonators. In this improved version we include only a single set-point in our control-loop. This allows us to tune the BM resonators close to a Hopf bifurcation. We present test results from a fabricated integrated circuit which, when compared with biological data, demonstrates the feasibility of our active 2-D cochlea model. Tara J. Hamilton, Craig T. Jin, André van Schaik, Jonathan Tapson |
ISCAS | 2 |
| 2008 | Self-tuned regenerative amplification and the hopf bifurcationabstractRecent work in cochlear amplifier modeling has focused on systems which show the dynamics of a Hopf bifurcation. We show that these systems are examples of a generic amplifier topology, the self-tuned regenerative amplifier (STRA). The STRA is a feedback-stabilized regenerative amplifier that can be operated in a region of supercritical stability. The signatures of Hopf amplification, such as a cubic nonlinear small-signal response at resonance, are general features of the topology. The topology is shown to include a degenerate parametric amplifier, which may explain its low noise and insensitivity to input-feedback phase mismatch. Jonathan Tapson, Tara J. Hamilton, Craig T. Jin, André van Schaik |
ISCAS | 3 |
| 2008 | A two-neuron cross-correlation circuit with a wide and continuous range of time delayabstractWe describe a circuit of two spiking neurons which extracts mathematically accurate cross-correlations from the signal inputs. It differs from prior circuits such as coincidence detectors or enhanced motion detectors in that it does not require anapriorifixed delay between input signals to be selected. The output, in the form of a differential spike histogram, displays a mathematical cross-correlation in the conventional correlation vs. time form. Jonathan Tapson, Mark P. Vismer, Craig T. Jin, André van Schaik, Fopefolu O. Folowosele, Ralph Etienne-Cummings |
ISCAS | 3 |
| 2007 | A Basilar Membrane Resonator for an Active 2-D CochleaabstractIn this paper we present a basilar membrane resonator design for an active 2D cochlea. It incorporates some of the non-linear behaviour exhibited in the real cochlea by utilizing a quality factor control loop. This control loop varies the gain and the frequency selectivity of the resonator based on the amplitude of the input signal. Tara J. Hamilton, Craig T. Jin, André van Schaik |
ISCAS | 2 |
| 2006 | Distance Variation Function for Simulation of Near-Field Virtual Auditory SpaceabstractWe present a method for simulating a near-field sound source in virtual auditory space (VAS). The method scales individualised HRTFs, measured at a distance of 1m to arbitrary distances in the near-field. It uses a model of the acoustic scattering for a point-source on a rigid sphere to calculate a distance variation function (DVF) to apply to the HRTFs. A sound localisation experiment was conducted in VAS with three subjects to evaluate the acoustic spatial fidelity of this method. Results show that with the modified HRTFs directional localisation is generally maintained at different distances and there is reasonable correlation between the perceived distance and target distance for distances up to 50cm from the centre of the subject's head. Alan Kan, Craig T. Jin, André van Schaik |
ICASSP (5) | 2 |
| 2006 | Listening Through Different Ears in the Sydney Opera HouseabstractWe present a psychoacoustic experiment that explores the ability of various listeners to discriminate between the virtual auditory space (VAS) stimuli generated using different binaural impulse response functions recorded in the Sydney Opera House. The binaural head-related impulse response (HRIR) functions were recorded for a group of subjects sitting in the same seat, P34, using a log sine sweep sound source located at the centre of the stage. The VAS stimuli generated using these HRIRs consist mostly of a variety of musical excerpts, speech, and white noise. Experimental results using an ABX test procedure show that out of a total of 1350 trials, 10 subjects responded correctly in 1230 of the test trials, indicating a discrimination performance greater than 90%. We also present data indicating the types of perceptual cues that aid in binaural sound discrimination process. Angela Qian Li, Craig T. Jin, André van Schaik |
ICASSP (5) | 2 |
| 2006 | An analysis of matching in the Tau cell log-domain filterabstractUsing various layout techniques and circuit configurations, the effects of matching on a log-domain filter were analyzed. It is shown here that one of three possible Tau cell configurations to implement the same 2nd order low pass filter clearly outperforms the others. Furthermore, application of a common centroid layout technique has had no noticeable improvement on filter matching Tara J. Hamilton, Craig T. Jin, André van Schaik |
ISCAS | 2 |
| 2005 | An aVLSI Cricket Ear ModelabstractFemale crickets can locate males by phonotaxis to the mating song they produce. The behaviour and underlying physiology has been studied in some depth showing that the cricket auditory system solves this complex problem in a unique manner. We present an analogue very large scale integrated (aVLSI) circuit model of this process and show that results from testing the circuit agree with simulation and what is known from the behaviour and physiology of the cricket auditory system. The aVLSI circuitry is now being extended to use on a robot along with previously modelled neural circuitry to better understand the complete sensorimotor pathway. 1 In t r o d u c t i o n Understanding how insects carry out complex sensorimotor tasks can help in the design of simple sensory and robotic systems. Often insect sensors have evolved into intricate filters matched to extract highly specific data from the environment which solves a particular problem directly with little or no need for further processing [1]. Examples include head stabilisation in the fly, which uses vision amongst other senses to estimate self-rotation and thus to stabilise its head in flight, and phonotaxis in the cricket. Because of the narrowness of the cricket body (only a few millimetres), the Interaural Time Difference (ITD) for sounds arriving at the two sides of the head is very small (1020s). Even with the tympanal membranes (eardrums) located, as they are, on the forelegs of the cricket, the ITD only reaches about 40s, which is too low to detect directly from timings of neural spikes. Because the wavelength of the cricket calling song is significantly greater than the width of the cricket body the Interaural Intensity Difference (IID) is also very low. In the absence of ITD or IID information, the cricket uses phase to determine direction. This is possible because the male cricket produces an almost pure tone for its calling song. * + School of Electrical and Information Engineering, Institute of Perception, Action and Behaviour. Figure 1: The cricket auditory system. Four acoustic inputs channel sounds directly or through tracheal tubes onto two tympanal membranes. Sound from contralateral inputs has to pass a (double) central membrane (the medial septum), inducing a phase delay and reduction in gain. The sound transmission from the contralateral tympanum is very weak, making each eardrum effectively a 3 input system. The physics of the cricket auditory system is well understood [2]; the system (see Figure 1) uses a pair of sound receivers with four acoustic inputs, two on the forelegs, which are the external surfaces of the tympana, and two on the body, the prothoracic or acoustic spiracles [3]. The connecting tracheal tubes are such that interference occurs as sounds travel inside the cricket, producing a directional response at the tympana to frequencies near to that of the calling song. The amplitude of vibration of the tympana, and hence the firing rate of the auditory afferent neurons attached to them, vary as a sound source is moved around the cricket and the sounds from the different inputs move in and out of phase. The outputs of the two tympana match when the sound is straight ahead, and the inputs are bilaterally symmetric with respect to the sound source. However, when sound at the calling song frequency is off-centre the phase of signals on the closer side comes better into alignment, and the signal increases on that side, and conversely decreases on the other. It is that crossover of tympanal vibration amplitudes which allows the cricket to track a sound source (see Figure 6 for example). A simplified version of the auditory system using only two acoustic inputs was implemented in hardware [4], and a simple 8-neuron network was all that was required to then direct a robot to carry out phonotaxis towards a species-specific calling song [5]. A simple simulator was also created to model the behaviour of the auditory system of Figure 1 at different frequencies [6]. Data from Michelsen et al. [2] (Figures 5 and 6) were digitised, and used together with average and "typical" values from the paper to choose gains and delays for the simulation. Figure 2 shows the model of the internal auditory system of the cricket from sound arriving at the acoustic inputs through to transmission down auditory receptor fibres. The simulator implements this model up to the summing of the delayed inputs, as well as modelling the external sound transmission. Results from the simulator were used to check the directionality of the system at different frequencies, and to gain a better understanding of its response. It was impractical to check the effect of leg movements or of complex sounds in the simulator due to the necessity of simulating the sound production and transmission. An aVLSI chip was designed to implement the same model, both allowing more complex experiments, such as leg movements to be run, and experiments to be run in the real world. Figure 2: A model of the auditory system of the cricket, used to build the simulator and the aVLSI implementation (shown in boxes). These experiments with the simulator and the circuits are being published in [6] and the reader is referred to those papers for more details. In the present paper we present the details of the circuits used for the aVLSI implementation. André van Schaik, Richard E. Reeve, Craig T. Jin, Tara J. Hamilton |
NIPS | 3 |
| 2001 | Parameterized Module Generator for an FPGA-Based Electronic Cochlea
Monk-Ping Leong, Craig T. Jin, Philip H. W. Leong |
FCCM | 2 |
| 1999 | Neural System Model of Human Sound Localization
Craig T. Jin, Simon Carlile |
NIPS | 1 |
| 1999 | Spectral Cues in Human Sound Localization
Craig T. Jin, Anna Corderoy, Simon Carlile, André van Schaik |
NIPS | 1 |
| 1999 | Human Localisation of Band-Pass Filtered NoiseabstractIn this work we study the influence and relationship of five different acoustical cues to the human sound localisation process. These cues are: interaural time delay, interaural level difference, interaural spectrum, monaural spectrum, and band-edge spectral contrast. Of particular interest was the synthesis and integration of the different cues to produce a coherent and robust percept of spatial location. The relative weighting and role of the different cues was investigated using band-pass filtered white noise with a frequency range (in kHz) of: 0.3-5, 0.3-7, 0.3-10, 0.3-14, 3-8, 4-9, and 7-14. These stimuli provided varying amounts of spectral information and physiologically detectable temporal information, thus probing the localisation process under varying sound conditions. Three subjects with normal hearing in both ears have performed five trials of 76 test positions for each of these stimuli in an anechoic room. All subjects showed systematic mislocalisation on most of these stimuli. The location to which they are mislocalised varies among subjects but in a systematic manner related to the five different acoustical cues. These cues have been correlated with the subject's localisation responses on an individual basis with the results suggesting that the internal weighting of the spectral cues may vary with the sound condition. André van Schaik, Craig T. Jin, Simon Carlile |
Int. J. Neural Syst. | 2 |