Mao-shen Jia

dblp:68/8762 · also Maoshen Jia · DBLP profile ↗
← Back
35ranked-venue papers
3as first author
27since 2021 · last 2026
0000-0002-3452-3913ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 1 first-author · 16 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021
YearPublicationVenuePosition
2026 Global diversity-based entropy minimization for source-free unsupervised domain adaptation of speaker verification
Xianhong Chen, Zhuorui Li, Mao-shen Jia
Speech Commun.4
2026 MulSE: Integrating Dual-Path Modeling and Global Attention for Multi-Channel Speech Enhancement
Mao-shen Jia, Yonggang Hu
IEEE Signal Process. Lett.2
2026 Multi-Level Interaction for Emotion Recognition From Unaligned Speech and Text
abstract
In multimodal emotion recognition, the diversity and temporal unalignment of speech and text modalities pose significant challenges for effective fusion. To address this issue, Multi-level Interaction for Emotion Recognition from Unaligned Speech and Text (MIUST) is proposed. Inspired by the hierarchical and multi-level integration process of human emotion cognition, the MIUST framework is designed to be consisted of a unimodal emotion recognition module, a multi-level cross-modal fusion module incorporating both coarse-grained feature learning and fine-grained modal fusion (FGMF), and an emotion classification module that synthesizes information from all stages. This multi-level branched fusion architecture more closely mirrors the human emotion understanding process. In addition, the FGMF can achieve cross-granularity fusion of speech and text. It learns emotional correlations between individual speech frames and textual words, thereby eliminating the need for explicit temporal alignment between modalities. Experimental results on the IEMOCAP and MELD datasets demonstrate the effectiveness of MIUST, achieving 77.78% weighted accuracy (WA), 78.79% unweighted accuracy (UA), and 77.82% weighted F1-score (W-F1) on IEMOCAP, outperforming existing state-of-the-art methods. These results validate that MIUST effectively improves multimodal emotion recognition performance by leveraging multi-level and cross-modal feature interaction. Our code is available athttps://github.com/HANLM15/MIUST.
Lingmin Han, Xianhong Chen, Mao-shen Jia, Changchun Bao
IEEE Trans. Affect. Comput.3
2025 SS-BRPE: Self-Supervised Blind Room Parameter Estimation Using Attention Mechanisms
abstract
In recent years, dynamic parameterization of acoustic environments has garnered attention in audio processing. This focus includes room volume and reverberation time (RT60), which defines local acoustics independently of sound source and receiver orientation. Previous studies show that purely attention-based models can achieve advanced results in room parameter estimation. However, their success relies on supervised pretrainings that require a large amount of labeled true values for room parameters and complex training pipelines. In light of this, we propose a novel Self-Supervised Blind Room Parameter Estimation (SS-BRPE) system. This system combines a purely attention-based model with self-supervised learning to estimate room acoustic parameters, from single-channel noisy speech signals. By utilizing unlabeled audio data for pretraining, the proposed system significantly reduces dependencies on costly labeled datasets. Our model also incorporates dynamic feature augmentation during fine-tuning to enhance adaptability and generalizability. Experimental results demonstrate that the SS-BRPE system not only achieves more superior performance in estimating room parameters than state-of-the-art methods but also effectively maintains high accuracy under conditions with limited labeled data. Code available at https://github.com/bjut-chunxiwang/SS-BRPE.
Chunxi Wang, Mao-shen Jia, Meiran Li, Changchun Bao, Wenyu Jin 0003
ICASSP2
2025 Using Corrected ASR Projection to Improve AD Recognition Performance from Spontaneous Speech
abstract
Alzheimer's Disease patients often exhibit cognitive decline, with language impairment being a prominent biomarker. Spontaneous speech analysis provides a non-invasive screening approach for AD. Large language models, increasingly employed for textual feature extraction, show potential in early AD prediction. However, Automatic Speech Recognition transcription errors, stemming from language impairments in AD and Mild Cognitive Impairment patients, can lead to information loss during feature extraction. To mitigate this, we introduce the Corrected ASR Projecting, CAP model. During training, ASR-transcribed text is manually corrected one by one, and then textual features are extracted independently using BERT, Claude, GLM, and GPT-3. The CAP model is trained by aligning the ASR transcription feature space with the corrected ASR transcription feature space. Experiments on the NCMMSC 2021 dataset demonstrate that the CAP model improves classification performance for AD recognition, with the maximum accuracy improvement reaching 5.55%.
Yun Jin, Mao-shen Jia, Peng Song 0007
ICASSP5
2025 Speech Enhancement with Dual-path Multi-Channel Linear Prediction Filter and Multi-norm Beamforming
Chengyuan Qin, Wenmeng Xiong, Mao-shen Jia, Changchun Bao
INTERSPEECH4
2025 Power Spectral Density Estimation for Acoustic Source Separation Using A Spherical Microphone Array
Mao-shen Jia, Yonggang Hu
INTERSPEECH2
2025 Direct-path Relative Harmonic Coefficients Detection for Multi-source Direction-of-Arrival Estimation in Reverberant Environments
Mao-shen Jia, Yonggang Hu
INTERSPEECH2
2025 Enhanced Prediction of Intracranial Aneurysm Rupture Risk via Multimodal Fusion
abstract
ABSTRACT It is well known that subarachnoid haemorrhage caused by intracranial aneurysm rupture has a high fatality rate. Therefore, the prediction of rupture risk can help doctors make targeted diagnoses and treatments in advance. In this study, an image‐text‐based hierarchical prediction model for the rupture risk of intracranial aneurysms (IAIT) is proposed, which combines medical images with structured texts to improve the prediction accuracy. This model captures the detailed features in the images, explores the interaction of multimodal features, and achieves better performance. Specifically, ternary‐view partial attention (TPA) is introduced into the image encoder to improve the model's attention to small lesions. With two symmetric local paths and one global path, local features can be better extracted, and Kronecker product is used for full fusion of image‐text features. Experiments on a private dataset show that the proposed model substantially outperforms both unimodal and existing multimodal baselines. It achieves over 12% higher accuracy than the best unimodal text model and over 45% higher than the best unimodal image model. Moreover, the TPA module further improves classification performance, validating the model's effectiveness. Overall, this study demonstrates the potential of multimodal fusion for accurate and interpretable prediction of intracranial aneurysm rupture risk.
Xinfeng Zhang 0002, Wei Guo 0020, Xiangsheng Li, Mao-shen Jia
IET Image Process.8
2025 Hybrid dual-path network: Singing voice separation in the waveform domain by combining Conformer and Transformer architectures
Chunxi Wang, Mao-shen Jia, Meiran Li, Dingding Yao
Speech Commun.2
2025 Cross-corpus speech emotion recognition using semi-supervised domain adaptation network
Mao-shen Jia, Xuan Cao, Jiawei Ru, Xinfeng Zhang 0002
Speech Commun.2
2024 Attention Is All You Need For Blind Room Volume Estimation
abstract
In recent years, dynamic parameterization of acoustic environments has raised increasing attention in the field of audio processing. One of the key parameters that characterize the local room acoustics in isolation from orientation and directivity of sources and receivers is the geometric room volume. Convolutional neural networks (CNNs) have been widely selected as the main models for conducting blind room acoustic parameter estimation, which aims to learn a direct mapping from audio spectrograms to corresponding labels. With the recent trend of self-attention mechanisms, this paper introduces a purely attention-based model to blindly estimate room volumes based on single-channel noisy speech signals. We demonstrate the feasibility of eliminating the reliance on CNNs for this task and the proposed Transformer architecture takes Gammatone magnitude spectral coefficients and phase spectrograms as inputs. To enhance the model performance given the task-specific dataset, cross-modality transfer learning is also applied. Experimental results demonstrate that the proposed model outperforms traditional CNN models across a wide range of real-world acoustics spaces, especially with the help of a dedicated pretraining and data augmentation schemes.
Chunxi Wang, Mao-shen Jia, Meiran Li, Changchun Bao, Wenyu Jin 0003
ICASSP2
2024 Spatial Acoustic Enhancement Using Unbiased Relative Harmonic Coefficients
Mao-shen Jia, Yonggang Hu, Changchun Bao
INTERSPEECH2
2024 TIM-Net: A multi-label classification network for TCM tongue images fusing global-local features
abstract
Abstract Combining the extracted tongue features with other medical indicators can effectively judge the diseases of patients. The previous work usually only analyzes a certain feature of the tongue body and is unable to extract multiple features simultaneously. In this study, a multi‐label classification network named TIM‐Net is proposed, which integrates global and local features to achieve multi‐label intelligent diagnosis of Chinese medicine tongue images. First, a feature extraction network based on ResNet is proposed to capture the features of tongue images more sufficiently. Then, a multi‐label classification algorithm fusing global and local features is proposed, and targeted screening operations are carried out on the class‐related feature maps based on global confidence. In addition, a logical masking algorithm is proposed to ensure that the local features can only correct the feature labels they represent, and do not interfere with other feature labels. The classification accuracy is further improved by using local feature confidence and correcting the global classification results. Finally, the experimental results indicate that the classification accuracy of the tongue images is gradually improved through optimizing the feature extraction network and fusing local features, and it exceeds other state‐of‐the‐art multi‐label classification networks.
Xinfeng Zhang 0002, Haonan Bian, Mao-shen Jia
IET Image Process.5
2024 A semi-supervised segmentation network fusing pseudo-label with multi-level feature consistency correction for hard exudates
abstract
Abstract Timely detection of hard exudates in fundus images can effectively avoid the severity of the disease, but the labelling of small and numerous lesion areas requires a lot of labour costs. This paper proposes a semi‐supervised segmentation network, which integrates pseudo‐labels and multi‐level features consistency correction. It achieves accurate segmentation of hard exudates by making full use of a small amount of labelled data and a large amount of unlabelled data. The network effectively extracts features from the unlabelled data through knowledge transfer of the teacher‐student model, and incorporates a Transformer network for auxiliary training to promote the quality of transfer. In addition, three unsupervised losses are introduced to improve the performance: the perturbation loss improves the robustness of the model to noise by adding different noises to the same input; the multi‐level feature consistency correction loss ensures the consistency of features of the student model at different scales; and the pseudo‐labelling cross‐supervision loss utilizes the generated pseudo‐labels for supervision between CNN and Transformer. By comparing the segmentation results with different proportion of the labelled data, it has better segmentation performance compared to other methods. The proposed methods can totally increase dice by 16.56% and mean intersection over union (MIoU) by 25.11%.
Xinfeng Zhang 0002, Mao-shen Jia
IET Image Process.6
2024 A distortionless convolution beamformer design method based on the weighted minimum mean square error for joint dereverberation and denoising
Changchun Bao, Mao-shen Jia, Wenmeng Xiong
Speech Commun.3
2024 Three-Dimensional Room Transfer Function Parameterization Based on Multiple Concentric Planar Circular Arrays
abstract
This study proposes a three-dimensional room transfer function (RTF) parameterization method based on multiple concentric planar circular arrays, which exhibits robustness to variations in the positions of both the receiver and source. According to the harmonic solution to the wave equation, the RTFs between two spherical regions (sound source and receiver) in a room can be expressed as a weighted sum of spherical harmonics, whose weight coefficients serve as the RTF parameters, which can be estimated by placing multiple concentric planar circular arrays composed of monopole-source pairs (MSPs) and multiple concentric planar circular arrays composed of omnidirectional-microphone pairs (OMPs) in respective source and receiver regions. We use MSP arrays to generate required outgoing soundfields originating from a source region. We derive a method to use OMP arrays to estimate RTF parameters that are concealed within the captured soundfield, which can be employed to reconstruct the RTF from any point in the source region to any point in the receiver region. The accuracy of the RTF parameterization method is validated through simulation testing.
Mao-shen Jia, Changchun Bao
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 First-Order Relative Harmonic Coefficient-Based Time-Frequency Points Selection for Multi-Source DOA Estimation
abstract
As a research focus within the field of array signal processing, multi-source direction-of-arrival (DOA) estimation in enclosed environments has been paid much attention. Contaminated by reverberation, noise, and inter-source interference, DOA estimation become challenging. Hence it is essential to identify time-frequency (TF) points dominated by only one source to alleviate these issues. This paper proposes a TF point selection method for DOA estimation based on the first-order relative harmonic coefficient (RHC). This is first analyzed on the “point” level from two perspective, and we design an adaptive single-source dominant zone (SSDZ) detection method. Subsequently, the relationship between first- and zero-order RHC magnitudes of different types of TF points is explored, and we develop a simple but useful rule to further select TF points in the detected SSDZs. Finally, we adopt two-dimensional (2-D) kernel density estimation (KDE) and peak search to estimate the DOAs of sources after calculating the angles of the detected TF points. The effectiveness and robustness of the proposed method are verified and compared with the reference methods through experiments with both the simulated and real-world recordings.
Mao-shen Jia, Changchun Bao, Wenmeng Xiong
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Harmonic-Aware Frequency and Time Attention for Automatic Piano Transcription
abstract
Automatic music transcription (AMT) is to transcribe music audio into note symbol representations. Concurrent notes overlapping in the frequency and time domains still hinder the performance of polyphonic piano transcription in current studies. In this work, we develop an attention-based method for piano transcription, where we propose a harmonic-aware attention to capture the musical frequency structure, and a local time attention to model temporal dependencies. The harmonic-aware frequency attention not only emphasizes the relationship between the obvious harmonics, but also extracts the correlation in the residual non-harmonic component. The time attention mechanism is improved using the learnable attention range masks to model frame-wise short-term dependencies on different subtasks. Experiments on the MAESTRO dataset demonstrate that the proposed system achieves state-of-the-art transcription performance on both frame-wise and note-wise F1 metrics. Considering the influence of the piano pedals' dynamic behavior on note duration, a note duration modification method is also proposed. With a more accurate annotation of the offset on MAESTRO, the transcription performance is further improved.
Qi Wang 0134, Mingkuan Liu, Changchun Bao, Mao-shen Jia
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 Joint DOA Estimation and Dereverberation Based on Multi-Channel Linear Prediction Filtering and Azimuth Sparsity
abstract
Source localization in reverberant environments has been a prominent research topic in the past two decades. In this paper, instead of the commonly employed time-frequency (TF) bin based methods which rely on empirically selected threshold values, we leverage the microphone array signal model comprising an early reverberant component and a late reverberant component, to propose a novel method for the source localization problem in reverberant environments. Our proposed criterion involves the joint removal of the late reverberant component using the multi-channel linear prediction (MCLP) filter, while estimating the directions of arrival (DOAs) of the actual sources using the early component signals. By applying the azimuth sparsity constraint, the true DOA can be estimated with high resolution and free from the interference of the early reflections. To solve the proposed criterion, DOAs, source signals, and MCLP filter coefficients are estimated by alternative iterations. Additionally, we present a source localization criterion specifically designed for the single source scenario as a special case of the multiple sources scenario. Finally, a source number estimation method and a postprocessing procedure are discussed for searching the global solutions to our proposed criteria. Evaluations with both simulated and realistic data demonstrate the advantages of our proposed methods over the baseline methods.
Wenmeng Xiong, Changchun Bao, Mao-shen Jia, José Picheral
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Adaptive learning Unet-based adversarial network with CNN and transformer for segmentation of hard exudates in diabetes retinopathy
abstract
Abstract Accurate segmentation of hard exudates in early non‐proliferative diabetic retinopathy can assist physicians in taking appropriate treatment in a more targeted manner, in order to avoid more serious damage to vision caused by the deterioration of the disease in the later stages. Here, an Adaptive Learning Unet‐based adversarial network with Convolutional neural network and Transformer (CT‐ALUnet) is proposed for automatic segmentation of hard exudates, combining the excellent local modelling ability of Unet with the global attention mechanism of transformer. Firstly, multi‐scale features are extracted through a CNN dual‐branch encoder. Then, the information fusion of features at adjacent scale is realized and the fused features are selected adaptively to maintain the overall consistency of features by attention‐guided multi‐scale fusion blocks (AGMFB). After that, the high‐level encoded features are input to transformer blocks to extract global contexts. Finally, these features are fused layer‐by‐layer to achieve accurate segmentation of hard exudates. In addition, adversarial training is incorporated into the above segmentation model, which improves Dice scores and MIoU scores by 7.5% and 3%, respectively. Experiments demonstrate that CT‐ALUnet shows more reliable segmentation and stronger generalization ability than other SOTA methods, which lays a good foundation for computer‐assisted diagnosis and assessment of efficacy.
Xinfeng Zhang 0002, Mao-shen Jia
IET Image Process.4
2023 Multi-Source Localization Using Optimized Time-Frequency Representation and Sparsity Component Analysis
abstract
This paper aims to address the multi-source localization problem by exploiting the sparsity of the speech signal in the time-frequency domain, where the challenge mainly lies in extracting the sparse component. An optimized time-frequency representation and sparsity component analysis-based multi-source localization method is proposed to overcome this challenge. Firstly, extracting the sparse components relies on the accurate representation in the time-frequency domain. However, the energy leakage problem caused by linear time-frequency transformation limits the accuracy of sparse component extraction. To tackle this problem, inspired by empirical mode decomposition, the proposed method classifies all the points in the time-frequency domain into four categories based on their phase feature and mode characteristics. Each type of the point is modeled separately, and a point-by-point analysis is conducted to remove all the points affected by energy leakage. Then, based on the optimized time-frequency representation, the phase coherence criterion is used to detect the sparse component in the point level. Following that, guided by the mode consistency characteristic of sparse components, an extension scheme is proposed to recover the falsely removed sparse components. Finally, the detected sparse components are applied for the multiple source localization. The objective evaluation is performed in both simulation and actual recording environments, and the proposed method can achieve better localization accuracy compared to several existing methods.
Mao-shen Jia, Dingding Yao, Jing Wang 0037
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Multiple-Speech-Source DOA Estimation Based on Single-Source Cluster Detection
abstract
This study proposes multiple-speech-source direction-of-arrival (DOA) estimation based on the distribution characteristic of the time-frequency (TF) point dominated by a single-source component (i.e., single-source point, SSP). By exploring the TF distribution characteristics of SSPs, we found that most are distributed in clusters in the TF domain. Hence, the concept of a single-source cluster (SSC) is given, each composed of adjacent TF points from one dominant sound source. Considering that SSCs have different shapes and sizes, an SSC detection method is designed based on point-to-cluster expansion, which is the research focus of this paper. A two-dimensional Gaussian function is introduced to model the theoretical distribution of the DOAs of SSPs, and a cluster expansion rule is proposed based on hypothesis testing of the DOA of a source. Two-dimensional kernel density estimation and peak search are adopted to estimate the DOAs and the number of sources using the detected SSCs. Experimental results in both simulated and real environments show that the proposed method can achieve better DOA estimation performance than some current techniques.
Mao-shen Jia, Jing Wang 0037, Ruiyuan Cao
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Speech Enhancement With Robust Beamforming for Spatially Overlapped and Distributed Sources
abstract
Most of the existing Beamforming methods are based on the assumptions that the sources are all point sources and the angular separation between the direction of arrival (DOA) of the source and the interference is large enough to assure good performance. In this paper, we consider a tough scenario where the target source and the interference are simultaneously spatially distributed and overlapped. To improve the performance of Beamforming in this scenario, we propose two approaches: the first approach exploits the non-Gaussianity as well as the spectrogram sparsity of the output of the microphone array; the second approach exploits the generalized sparsity with overlapped groups of the Beampattern. The proposed criteria are solved by methods based on linearized preconditioned alternating direction method of multipliers (LPADMM) with high accuracy and high computational efficiency. Numerical simulations and real data experiments show the advantages of the proposed approaches compared to previously proposed Beamforming methods for signal enhancement.
Wenmeng Xiong, Changchun Bao, Mao-shen Jia, José Picheral
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 A Hierarchical Retrieval Method Based on Hash Table for Audio Fingerprinting
Mao-shen Jia, Xuan Cao
ICIC (1)2
2021 Person Re-identification Based on Hash
Xinfeng Zhang 0002, Bowen Ren, Mao-shen Jia
ICIC (1)5
2021 Multi-Source DOA Estimation in Reverberant Environments by Jointing Detection and Modeling of Time-Frequency Points
abstract
In this article, the direction of arrival (DOA) estimation of multiple speech sources in reverberant environments is investigated based on the recording of a soundfield microphone. First, the recordings are analyzed in the time-frequency (T-F) domain to detect both “points” (single T-F points) and “regions” (multiple, adjacent T-F points) corresponding to a single source with low reverberation (known as low-reverberant-single-source (LRSS) points). Then, a LRSS point detection algorithm is proposed based on a joint dominance measure and instantaneous single-source point (SSP) identification. Following this, initial DOA estimates obtained for the detected LRSS points are analyzed using a Gaussian Mixture Model (GMM) derived by the Expectation-Maximization (EM) algorithm to cluster components into sources or outliers using a rule-based method. Finally, the DOA of each actual source is obtained from the estimated source components. Experiments on both simulated data and data recorded in an actual acoustic chamber demonstrate that the proposed algorithm exhibits improved performance for the DOA estimation in reverberant environments when compared to several existing approaches.
Mao-shen Jia, Changchun Bao, Christian H. Ritz
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 Sound Field Reproduction in Reverberant Room Using the Alternating Direction Method of Multipliers Based Lasso and Regularized Least-Square
Mao-shen Jia, Changchun Bao, Qi Wang 0134
ICIC (1)2
2018 Optical Character Detection and Recognition for Image-Based in Natural Scene
Bochao Wang, Xinfeng Zhang 0002, Yiheng Cai, Mao-shen Jia
ICIC (3)4
2018 Separation of multiple speech sources by recovering sparse and non-sparse components from B-format microphone recordings
Mao-shen Jia, Jundai Sun, Changchun Bao, Christian H. Ritz
Speech Commun.1
2018 Design of a Planar First-Order Loudspeaker Array for Global Active Noise Control
abstract
This paper proposes a method to design a planar first-order loudspeaker array structure for global active noise control. Compared with the traditional spherical loudspeaker array, the planar array provides a practical design with flexible source locations. The planar array is capable of achieving global noise control, provided that the loudspeakers have general variable first-order responses in elevation. On x-y plane, we use spherical harmonics to analyze the required first-order loudspeakers consisting of monopole and tangential dipole components. By exploiting the properties of the associated Legendre functions and its derivative, we can divide the primary soundfield into even harmonics controlled by the monopole component, and odd harmonics controlled by the dipole component. Through the appropriate choice of radii of circles, we avoid the ill-conditioning problem of matrix inversion and derive a robust solution for loudspeaker weights to suppress the primary noise field. Besides, we use the closely-located monopole pairs, instead of the ideal general first-order loudspeakers, to design an alternative planar array for practical implementation. As an illustration, we use several simulation examples to validate the performance of the two proposed planar loudspeaker arrays.
Changchun Bao, Mao-shen Jia
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Encoding Multiple Audio Objects Using Intra-Object Sparsity
abstract
Preserving audio scenes in the form of audio objects has become common in recent years. Object-based audio techniques provide more flexibility for personalized rendering as well as a more accurate audio object trajectory. For encoding and transmitting multiple audio objects in a lossy manner, a new compression framework for multiple simultaneously occurring audio objects is presented in this work. The proposed encoding approach is based on the intra-object sparsity (approximate k-sparsity). After establishing a quantitative measure of approximate k-sparsity, statistical analysis is employed to validate the proposed intra-object sparsity of audio objects. By exploring this intra-object sparsity, multiple simultaneously occurring audio objects are compressed into a mono downmix signal with side information. This downmix signal can be further compressed by legacy audio codecs. Meanwhile, the side information is transmitted in a lossless manner. The objective and subjective evaluations revealed that the proposed compression framework achieved better perceptual quality compared to an existing technique where up to eight audio objects are considered. The subjective evaluations also confirmed that the proposed approach is able to achieve scalable transmission according to the bandwidth while preserving the perceptual quality of both the individual audio objects and the spatial audio scenes.
Mao-shen Jia, Changchun Bao, Xiguang Zheng, Christian H. Ritz
IEEE ACM Trans. Audio Speech Lang. Process.1
2014 The design of Ambisonic reproduction system based on dynamic gain parameters
abstract
This paper describes a design approach of Ambisonic reproduction system based on dynamic gain parameters (DGP). In the conventional approaches, the fixed gain parameters are often optimized to minimize the overall objective function for whole 360° sound stage. The proposed approach has an advantage that the gain parameters vary with angles of source objects. The problem of optimization tradeoff among different angles is overcome by DGP, which achieves an optimal solution in each position. Source localizations of the B-Format signals were estimated in frequency bands in order to match the corresponding gain parameters. For the synthesized signals, the process was simplified by the given spatial information. Using the head-related transfer function (HRTF) analysis, the proposed approach was found to be significantly better than reference approaches in interaural time difference (ITD) and interaural level difference (ILD).
Changchun Bao, Mao-shen Jia
ICASSP3
2010 High frequency reconstruction of audio signal based on chaotic prediction theory
abstract
The quality of audio signals that have been encoded with low-bit rate audio coding standards is degraded because the high frequency information has been removed. The quality of such audio signals can, however, be improved by reconstructing the high frequency information which was lost. In this paper the principles of audio signal production and the characteristics of the human hearing system have been used to develop a blind high frequency reconstruction method based on chaotic prediction theory. Performance evaluation with objective and subjective tests has shown that this method is, in most cases, more efficient than other blind high frequency reconstruction methods.
Yong-tao Sha, Changchun Bao, Mao-shen Jia, Xin Liu 0043
ICASSP3
2008 A 8.32 kb/s embedded wideband speech coding candidate for ITU-t EV-VBR standardization
Changchun Bao, Hai-ting Li, Ze-xin Liu, Mao-shen Jia
INTERSPEECH6