Xiaochen Wang 0001

dblp:19/30-1 · DBLP profile ↗
← Back
57ranked-venue papers
0as first author
28since 2021 · last 2026
0000-0003-4177-7580ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 38 · 12 since 2021Artificial intelligence and machine learning · 11 · 11 since 2021Systems, architecture and hardware · 3 · 1 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FlowSR: Efficient Speech Super-resolution via Adaptive Prior Flow Matching
Jiajun Yuan, Xiaochen Wang 0001
ICIC (14)2
2025 Efficient Extended Neighborhoods Dynamic Selection Re-Ranking for Person Re-Identification
abstract
Person re-identification (re-ID) is a challenging retrieval task that requires matching a person’s captured images across non-overlapping camera views, with re-ranking being a critical step for improving accuracy. The k-nearest neighbors relationship is commonly used to determine the rank results by selecting only k fixed pedestrian image values for distance calculations. This operation, however, generates additional distance errors due to changes in the appearance of pedestrians. This paper addresses the above issue by proposing a simple but effective Extended Neighborhood Dynamic Selection (ENDS), distance to optimize the performance of ReID reranking. The number of images selected is then distributed over an interval. A limit of upper and lower is placed on the selected number to ensure that it is neither too low nor too high. The automatic selection of adjacent images is achieved using this method. This distance is determined by combining ENDS distances with Jaccard distances. It is the core principle of this method that, instead of using fixed values, the choice of the number of images to be included in each neighborhood should be made automatically. It also allows the removal of images that are dissimilar in favour of those that are more representative. Experimental results demonstrate the novel method of ranking by using Market-1501 and DukeMTMC’s reID dataset. In this paper, we propose a method that increases the Market-1501 mAP/Rank1 by 29.4%/12.9% while DukeMTMC-reID reranking by 35%/20.7%.
Chao Wang 0084, Zhongyuan Wang 0001, Xiaochen Wang 0001, Ruimin Hu, Mithun Mukherjee 0001
SMC3
2025 Multimodal and multichannel speech separation using location-guided speech feature mapping network
Yulin Wu 0003, Xiaochen Wang 0001, Dengshi Li, Ruimin Hu
Neurocomputing2
2025 Optimal Illumination Distance Metrics for Person Re-Identification in Complex Lighting Conditions
abstract
Person re-identification is extensively applied in public security and surveillance. However, environmental factors like time and location often lead to varying lighting conditions in captured pedestrian images, significantly impacting identification accuracy. Current approaches mitigate this issue through lighting transformation techniques, aiming to normalize images to a standard lighting condition for consistent person re-identification results. Yet, these methods overlook the fact that different content may hold distinct identification values under diverse lighting conditions. To address this, we conducted an analysis on the identification distance between images of the same or different pedestrians under pre-defined lighting conditions. From this analysis, we introduce the concept of optimal lighting: a condition where the distance between image pairs is minimized compared to other lighting scenarios. We propose utilizing this optimal lighting distance in the image retrieval process for final ranking. Our study, validated on synthetic datasets Market-IA and Duke-IA, demonstrates that optimal lighting is independent of image texture information. Each image pair exhibits a unique optimal lighting, yet consistently shows a minimum distance value.
Chao Wang 0084, Zhongyuan Wang 0001, Ruimin Hu, Xiaochen Wang 0001, Wen Zhou 0029
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Optimal Illumination Distance Metrics for Person Re-identification
Chao Wang 0084, Zhongyuan Wang 0001, Ruimin Hu, Xiaochen Wang 0001, Wen Zhou 0029
PRICAI (4)4
2024 Adaptive subband partition encoding scheme for multiple audio objects using CNN and residual dense blocks mixture network
Yulin Wu 0003, Ruimin Hu, Xiaochen Wang 0001
Expert Syst. Appl.3
2024 Learning with noisy labels for robust fatigue detection
Ruimin Hu, Xiaojie Zhu, Dongliang Zhu 0001, Xiaochen Wang 0001
Knowl. Based Syst.5
2023 Multi-speaker Direction of Arrival Estimation Using Audio and Visual Modalities with Convolutional Neural Network
abstract
In reality, audible and visible sound sources are closely aligned, and they can help humans locate sources exactly. To exploit the complementarity between audio and visual data in multi-speaker direction of arrival (DoA) estimation, we propose a novel network consisting of 3D convolution neural networks (3D-CNNs) and 2D-CNNs mixture networks with residual dense blocks. It has two main advantages: 1) both input audio and visual features are low-level signal representation: the real and imaginary parts of STFT coefficients for the audio feature and pixel coordinates for the visual feature, which can allow the network to learn to extract the most informative high-level features. 2) 3D-CNNs with the residual dense block are used for audio and visual feature mapping along the time and frequency axis. The following 2D-CNNs are to ensemble the high-level features along the DoA axis. Experimental results demonstrate promising SSL performance.
Yulin Wu 0003, Ruimin Hu, Xiaochen Wang 0001
ICME3
2023 Perceptual Audio Object Coding Using Adaptive Subband Grouping with CNN and Residual Block
abstract
Spatial audio content is becoming increasingly popular and is regarded as a set of object signals with associated metadata. The object-based content representation is independent of loudspeaker layouts and provides high spatial resolution when reproduced on more loudspeakers. The audio quality of the traditional spatial audio object coding (SAOC) method has severe aliasing distortion, which impairs the immersive listening experience. In this study, we reduce aliasing distortion by perceptual adaptive subband grouping strategy and use the convolutional neural network (CNN) and residual block to build the side information compressing model. Both objective and subjective experiments on benchmark datasets with different bitrates show that the proposed method achieves favorable performance against state-of-the-art methods.
Yulin Wu 0003, Ruimin Hu, Xiaochen Wang 0001
ICME3
2023 Multi-scale modeling temporal hierarchical attention for sequential recommendation
Nana Huang, Ruimin Hu, Xiaochen Wang 0001
Inf. Sci.3
2023 Cross-platform sequential recommendation with sharing item-level relevance data
Nana Huang, Ruimin Hu, Xiaochen Wang 0001, Xinjian Huang
Inf. Sci.3
2023 Single-channel Multi-speakers Speech Separation Based on Isolated Speech Segments
Shanfa Ke, Zhongyuan Wang 0001, Ruimin Hu, Xiaochen Wang 0001
Neural Process. Lett.4
2023 Multi-speaker DoA Estimation Using Audio and Visual Modality
Yulin Wu 0003, Ruimin Hu, Xiaochen Wang 0001, Shanfa Ke
Neural Process. Lett.3
2023 A Dual Self-Attention mechanism for vehicle re-Identification
Wenqian Zhu, Zhongyuan Wang 0001, Xiaochen Wang 0001, Ruimin Hu, Huikai Liu, Chao Wang 0084, Dengshi Li
Pattern Recognit.3
2023 From Collective Attribute Association of Groups to Precise Attribute Association of Individuals
abstract
Obscured person re-identification (Re-ID) aims to match an obscured image with a complete image of the same person captured by other cameras. As a major challenge in person identification, occlusion severely affects the effectiveness of most traditional person Re-ID methods. To solve this problem, this study proposes a trajectory association method, which, as a pre-processing technique for person Re-ID, can narrow the search range and reduce the problem of degradation caused by mixing. We investigate the method of converting the fuzzy association between sets into the precise association between elements for M video objects and N phone objects (trajectory information) with fuzzy group association relationships at the crime scene. First, we decompose the M-N precise association problem and analyze the similarity of the video objects in the source point and on the trajectories. Then, we define high-similarity points, study their distribution characteristics in different trajectories, and find that there is a significant difference between the distribution of high-similarity points in correct and incorrect matching trajectories. We simplify the full-path association problem into a partial-path high-similarity point distribution difference problem, which effectively reduces the difficulty in accurate association relationship construction. The association experiments in simple and mixed scenarios as well as Re-ID experiments on the PRPW and Market1501 demonstrate the effectiveness of our method.
Yun Lan, Ruimin Hu, Xin Xu 0007, Dengshi Li, Chao Wang 0084, Xiaochen Wang 0001
IEEE Trans. Multim.6
2022 $(\alpha, \ \beta)$-AWCS: $(\alpha, \ \beta)$-Attributed Weighted Community Search on Bipartite Graphs
abstract
Community search on bipartite graphs aims to find a community closely associated with the query vertex for personalized recommendation, fraud detection, and team formation. During community search, considering both the structural closeness and the homogeneity of attributes of nodes is the key to improving the quality of the output community. Traditional work uses the$(\alpha,\beta)$-core model to guarantee structural cohesion of the nodes (i.e., degree of each upper vertex is at least$\alpha$and degree of each lower vertex is at least$\beta$). However, it ignores the attributes of nodes, resulting in an average attribute similarity of only about 0.17 for the node pairs in the output community. In this paper, a framework for$(\alpha,\ \beta)$-Attributed Weighted Community Search ($(\alpha,\ \beta)$-AWCS) was proposed. It output a connected subgraph of$G$containing the query vertex, which satisfies both structurally cohesive (i.e., ($(\alpha,\ \beta){-}$-core) and keyword cohesiveness (i.e., its vertices share common keywords). The framework includes a pruning strategy to strip vertices that do not contain query attributes, thus effectively reducing the search space, and two algorithms improve the attributes cohesiveness of the output community. One of the exact algorithms first obtains a subgraph of attribute cohesion and subsequently keeps the structure cohesive. The other approximate algorithm has higher robustness, which iteratively removes the vertex with the lowest attribute score until the structural cohesion cannot be maintained. We have conducted experiments on real datasets of different sizes. Experiments show that both algorithms can improve the attribute cohesiveness metric by more than 25% compared to the traditional method. Meanwhile, structural cohesion was appropriate.
Dengshi Li, Xiaocong Liang, Ruimin Hu, Xiaochen Wang 0001
IJCNN5
2022 High Parameter Frequency Resolution Encoding Scheme for Spatial Audio Objects Using Stacked Sparse Autoencoder
Yulin Wu 0003, Ruimin Hu, Xiaochen Wang 0001, Chenhao Hu, Shanfa Ke
Neural Process. Lett.3
2022 COVID-19 contact tracking by group activity trajectory recovery over camera networks
Chao Wang 0084, Xiaochen Wang 0001, Zhongyuan Wang 0001, Wenqian Zhu, Ruimin Hu
Pattern Recognit.2
2021 Spatial Audio Object Coding Based on Time-Frequency Shifting and Scheduling
abstract
Spatial audio object coding (SAOC) is an effective method to transmit multiple audio objects. It divides the full frequency band into 28 subbands and extracts spatial parameters for de-coding. In this way, objects can be encoded into a downmix signal with a few parameters. However, using the same parameters in one subband will cause frequency aliasing distortion, which severely impacts the listening experience. Existing studies to enhance SAOC cannot eliminate the aliasing distortion of all objects effectively. This paper describes a new structure to balance the bit-rate and decode quality based on time-frequency (TF) shifting and scheduling. In this structure, a TF shifting strategy (contains global shifting and local shifting) is proposed to reduce frequency aliasing distortion. Furthermore, a scheduling strategy is used to decide which part should be shifted according to the degree of aliasing. From the experiment results, the performance of the proposed method is better than SAOC and other enhanced methods.
Chenhao Hu, Ruimin Hu, Xiaochen Wang 0001, Yulin Wu 0003
ICME3
2021 Efficient Multi-Step Audio Object Coding with Limited Residual Information
abstract
Spatial audio object coding (SAOC) is an effective method to transmit multiple audio objects. Audio systems can provide personalized services under this framework. However, this method causes frequency aliasing distortion, which severely impacts the listening experience. The multi-step SAOC (MS-SAOC) scheme was proposed to enhance the sound quality of each audio object by using residual information. Compared with SAOC, the bit-rate increases three times due to the residual data of multiple objects. In this paper, an efficient multi-step residual coding method is proposed to reduce the residual bit-rate of MS-SAOC. A two-level filter is designed to remove redundant residual information, and the limited residual information can efficiently compensate for frequency aliasing distortion. From experiment results, the residual bit-rate is half of MS-SAOC, and the sound quality is maintained at the Good-Excellent level.
Chenhao Hu, Ruimin Hu, Xiaochen Wang 0001, Yulin Wu 0003, Wenke Liu
ICME3
2021 Low Bitrates Audio Object Coding Using Convolutional Auto-Encoder and Densenet Mixture Model
abstract
The efficient transmission of the audio objects can be achieved by spatial audio object coding (SAOC) method that conveys a mono downmix signal together with side information parameters that enable object reconstruction in the decoder. To allow the transmission of audio objects at low bitrates, we present a new audio coding method with convolutional auto-encoder (CAE) and dense convolutional network (DenseNet) mixture model, optimizing the compression of side information parameters of audio objects. It has two main advantages: 1) Different from the linear transform methods, CAE can dig the nonlinear relationship of side information parameters and can effectively reduce the dimension of side information parameters; 2) DenseNet is adding in the decoder to make full use of low dimensional features of side information parameters, which improves the audio quality at low bitrate. Experiments show that our method outperforms base-line methods permitting bitrates as low as 1 kbps per object.
Yulin Wu 0003, Ruimin Hu, Chenhao Hu, Shanfa Ke, Xiaochen Wang 0001
ICME6
2021 Stacked Sparse Autoencoder for Audio Object Coding
Yulin Wu 0003, Ruimin Hu, Xiaochen Wang 0001, Chenhao Hu
MMM (1)3
2021 Intelligent cloud computing platform for three-dimensional sound reproduction
abstract
Summary Three‐dimensional (3D) audio reproduction techniques reproduce realistic sound sources and spatial perception for listeners. However, it is a rather difficult work to configure all the parameters in the previous 3D reproduction systems since ordinary users hardly possess professional acoustic knowledge. In this article, we developed a sound source distance reproduction model and design an online parameter computing framework for users to generate reproduction parameters based on a cloud computing platform. The proposed model reproduces sound direction and distance cues under specific conditions. Subjective experiments are executed to assess the spatial reproduction performance in a real environment and objective experiments are performed in simulated scenarios. Both the subjective and objective experiments show that the proposed method improves the spatial perception of sound events. Since the reproduction model depends on environment and loudspeaker array configuration, a cloud parameter generation framework is developed for users to generate acoustic parameters for the reproduction system.
Maosheng Zhang, Ruimin Hu, Jiang Lin, Xiaochen Wang 0001
Concurr. Comput. Pract. Exp.4
2021 Audio object coding based on N-step residual compensating
Chenhao Hu, Xiaochen Wang 0001, Ruimin Hu, Yulin Wu 0003
Multim. Tools Appl.2
2021 Optimization of sound fields reproduction based Higher-Order Ambisonics (HOA) using the Generative Adversarial Network (GAN)
Lingkun Zhang, Xiaochen Wang 0001, Ruimin Hu, Dengshi Li, Weiping Tu
Multim. Tools Appl.2
2021 Estimation of spherical harmonic coefficients in sound field recording using feed-forward neural networks
Lingkun Zhang, Xiaochen Wang 0001, Ruimin Hu, Dengshi Li, Weiping Tu
Multim. Tools Appl.2
2021 Trajectory Association for Person Re-identification
Ruimin Hu, Wenxin Huang, Dengshi Li, Xiaochen Wang 0001, Chenhao Hu
Neural Process. Lett.5
2021 Intelligibility Enhancement Via Normal-to-Lombard Speech Conversion With Long Short-Term Memory Network and Bayesian Gaussian Mixture Model
abstract
Speech communications and interactions frequently occur in a variety of environments. Noise in the environment significantly degrades speech intelligibility when speaking and listening. Especially in the listening stage, even if the multimedia terminal outputs clean speech, it is still difficult for listeners to obtain information. Intelligibility enhancement (IENH) of speech is a technique for overcoming the environmental noise in the listening stage. It implements a perceptual enhancement of non-noisy speech. This study focuses on IENH via normal-to-Lombard speech conversion, inspired by a well known acoustic mechanism named the Lombard effect. Our method combines the long short-term memory (LSTM) network and Bayesian Gaussian mixture model (BGMM) to build a conversion architecture. Compared with baselines, it has three main advantages: 1) an LSTM network is used for spectral tilt mapping with fully considering short-term correlations and high-dimensional expression abilities; 2) the aperiodicity (AP) is mapped together with the fundamental frequency ($F_0$) by a BGMM, which considers their relevance constraints and the importance of APs; 3) the gender-dependent mapping is used for$F_0$and APs to consider distribution differences between genders. Experiments indicate that our method gets better performance in both objective and subjective tests.
Xiaochen Wang 0001, Ruimin Hu, Huyin Zhang, Shanfa Ke
IEEE Trans. Multim.2
2020 Speech Intelligibility Enhancement Using Non-Parallel Speaking Style Conversion With Stargan And Dynamic Range Compression
abstract
Speech intelligibility enhancement is a perceptual enhancement technique for clean speech reproduced in noisy environments. It is typically used in the listening stage of multimedia communications. In this study, we enhance speech intelligibility by speaking style conversion (SSC), which is a data-driven approach inspired by a vocal mechanism named Lombard effect. The proposed SSC method combines star generative adversarial network (StarGAN) based mapping and dynamic range compression (DRC). It has two main advantages: 1) different from gender-independent conversion in previous studies, StarGAN can separately learn speech features of different genders to provide a differential conversion among genders with a single model and non-parallel training data; 2) we design a multi-level enhancement strategy with the use of DRC in the StarGAN architecture, which improves the SSC performance in strong noise interference. Experiments show that our method outperforms baseline methods.
Ruimin Hu, Shanfa Ke, Xiaochen Wang 0001
ICME5
2020 Normal-To-Lombard Speech Conversion by LSTM Network and BGMM for Intelligibility Enhancement of Telephone Speech
abstract
Noise in the environment significantly decreases the speech intelligibility of telephone conversations. Despite clean speech output from the device, the listener is still hard to get information. This study focuses on intelligibility enhancement (IENH) of telephone speech in near-end background noise based on normal-to-Lombard speech conversion. The proposed approach uses long short-term memory (LSTM) and Bayesian Gaussian mixture model (BGMM) to build the speech mapping model. Compared with previous studies, we fully consider the short-term correlations of speech and implement feature mappings with higher dimensional features and more types of features. Evaluations indicate that the proposed approach has achieved better results in both objective and subjective evaluation.
Xiaochen Wang 0001, Ruimin Hu, Huyin Zhang, Shanfa Ke
ICME2
2020 HRTF Representation with Convolutional Auto-encoder
Wei Chen 0143, Ruimin Hu, Xiaochen Wang 0001, Dengshi Li
MMM (1)3
2020 Perceptual Localization of Virtual Sound Source Based on Loudspeaker Triplet
Duanzheng Guan, Dengshi Li, Xuebei Cai, Xiaochen Wang 0001, Ruimin Hu
MMM (2)4
2020 Multi-step Coding Structure of Spatial Audio Object Coding
Chenhao Hu, Ruimin Hu, Xiaochen Wang 0001, Tingzhao Wu, Dengshi Li
MMM (1)3
2020 HMM-Based Person Re-identification in Large-Scale Open Scenario
Ruimin Hu, Wenxin Huang, Xiaochen Wang 0001, Dengshi Li
MMM (1)4
2020 Loudspeaker triplet selection based on low distortion within head for multichannel conversion of smart 3D home theater
abstract
Summary In recent years, with the vigorous development of 3D film industry, the demand for 3D Smart Home Theater, based on Internet of Things (IoT), continues to grow. Theaters are populated with a large number of loudspeakers for more realistic 3D sound effects. However the number of loudspeakers in home is limited. Therefore, multichannel conversion is required to achieve theater 3D sound effects in home. Traditionally, the replaced loudspeaker signal of the original system is assigned to a “loudspeaker triplet” of the converted system. A large amount of subjective evaluations is necessary to judge the consistency of the replaced loudspeaker position with the “phantom source” positions reconstructed by loudspeaker triplets. In this study, after calculating the least‐squares errors of the reproduced sound field within a given region, we explore the constraint between the low distortion of the reproduced sound field and the loudspeaker triplet positions. Using this constraint rule, we present a new loudspeaker triplet selection criteria that can greatly reduce the number and time of subjective evaluations for selecting the optimal loudspeaker triplet. Simulation and subjective evaluation experiments indicate that the proposed selection method outperforms the traditional method, and that the proposed method can be successfully applied to multichannel conversion.
Dengshi Li, Ruimin Hu, Xiaochen Wang 0001, Weiping Tu
Concurr. Comput. Pract. Exp.3
2020 Single Channel multi-speaker speech Separation based on quantized ratio mask and residual network
Shanfa Ke, Ruimin Hu, Xiaochen Wang 0001, Tingzhao Wu, Zhongyuan Wang 0001
Multim. Tools Appl.3
2020 A mapping model of spectral tilt in normal-to-Lombard speech conversion for intelligibility enhancement
Ruimin Hu, Xiaochen Wang 0001
Multim. Tools Appl.4
2019 Multi-speakers Speech Separation Based on Modified Attractor Points Estimation and GMM Clustering
abstract
In this paper, a new attractor points estimation method for DANet algorithm used in single channel multi-speaker speech separation has been proposed. A prerequisite is that there must be separate segments of each source in the mixture. This condition is met in the actual situation because the source signal is not overlapping at any time. With this prerequisite, an isolated source segments extracted from the mixture is converted to the embedding space. With the embedding of isolated source segments, a more accurately attractor point for each source will be created, due to it does not contain components of other sources. In addition, a gaussian mixture model(GMM) clustering method instead of K-means clustering method were used at run time. The experiment demonstrated that the proposed method gets a better separation performance than state of the art method up to 1.04dB in SDR.
Shanfa Ke, Ruimin Hu, Tingzhao Wu, Xiaochen Wang 0001, Zhongyuan Wang 0001
ICME5
2019 Spectral Tilt Estimation for Speech Intelligibility Enhancement Using RNN Based on All-Pole Model
Ruimin Hu, Xiaochen Wang 0001
MMM (2)4
2019 A near-end listening enhancement system by RNN-based noise cancellation and speech modification
Ruimin Hu, Xiaochen Wang 0001
Multim. Tools Appl.3
2019 Audio object coding based on optimal parameter frequency resolution
Tingzhao Wu, Ruimin Hu, Xiaochen Wang 0001, Shanfa Ke
Multim. Tools Appl.3
2018 Individualization of Head Related Transfer Functions Based on Radial Basis Function Neural Network
abstract
Head Related Transfer Functions (HRTFs) contain sound localization cues and are commonly used in 3D audio reproduction. Due to HRTFs are closely related to anthropometric parameters (head, pinna, torso), which means HRTFs vary with each individual, how to obtain a set of suitable HRTFs for each individual remains to be solved. In this paper, we investigated the complex relationship between HRTFs and anthropometric parameters through Radial Basis Function neural network (RBF), and proposed a method of generating individualized HRTFs with listener's anthropometric parameters. Objective experiments show that the estimated HRTFs have good consistency with measured ones, and the spectral distortion values have an average reduction of 0.59 dB compared with other methods. Subjective listening tests show that using estimated HRTFs enable accurate auditory localization.
Lian Meng, Xiaochen Wang 0001, Wei Chen 0143, Chunling Ai, Ruimin Hu
ICME2
2017 Sound physical property matching between non central listening point and central listening point for NHK 22.2 system reproduction
abstract
NHK has proposed a famous 3D audio system: 22.2 multi-channel system, but its loudspeakers are too many and are troublesome to put in home. Ando and Wang has proposed two simplification methods to reduce its channel number, but only 3D sound field at the central listening point can be recovered well by NHK 22.2 system and its simplified systems, the listening experience at a non central listening point is worse than that at the central listening point. In real life, listeners may stay at arbitrary listening point: central or non central point. Conventional pressure matching and particle matching method could be used for non central zone sound field reproduction, but they have some theoretical shortcomings. To address these problems, this paper propose a universal non central listening point sound field reproduction method by matching sound physical property between a non central listening point and the central listening point. Subjective and objective experiments show the effectiveness of the proposed method.
Song Wang 0011, Ruimin Hu, Xiaochen Wang 0001, Yuhong Yang 0001, Weiping Tu
ICASSP4
2017 The Perceptual Lossless Quantization of Spatial Parameter for 3D Audio Signals
Xiaochen Wang 0001, Ruimin Hu, Dengshi Li
MMM (2)2
2017 Frame-Independent and Parallel Method for 3D Audio Real-Time Rendering on Mobile Devices
Yucheng Song, Xiaochen Wang 0001, Wei Chen 0143, Weiping Tu
MMM (2)2
2017 3D Sound Field Reproduction at Non Central Point for NHK 22.2 System
Song Wang 0011, Ruimin Hu, Xiaochen Wang 0001, Yuhong Yang 0001, Weiping Tu
MMM (1)4
2017 Binaural Sound Source Distance Reproduction Based on Distance Variation Function and Artificial Reverberation
Jiawang Xu, Xiaochen Wang 0001, Maosheng Zhang
MMM (2)2
2016 Multichannel reduction based on sound field within two ears
abstract
People hope to use a small number of loudspeakers to get the experience of the film 3D sound at home. Considering that people use two ears to listen, this paper provides a method which reproduce the sound field within the region of two ears. We develop the fundamental performance limits for the truncated spherical harmonic function expansions of the sound field within the region of ears. Based on this, the low distortion of reproduced sound field within two ears is maintained in the processing of reducing loudspeakers from Q to Q-1. The 22.2 multichannel sound system without two low-frequency effect channels can be simplified to 6 channels automatically and the total of loudspeaker arrangements is ten. The subjective evaluation of the proposed method is better than that of the previous multichannel reduction method with the decrease of the number of loudspeakers.
Dengshi Li, Ruimin Hu, Xiaochen Wang 0001, Guo Wu, Weiping Tu
ICME3
2016 Analysis and Comparison of Inter-Channel Level Difference and Interaural Level Difference
Tingzhao Wu, Ruimin Hu, Xiaochen Wang 0001, Shanfa Ke
MMM (1)4
2016 Adaptive Multichannel Reduction Using Convex Polyhedral Loudspeaker Array
Lingkun Zhang, Ruimin Hu, Dengshi Li, Xiaochen Wang 0001, Weiping Tu
MMM (1)4
2015 A down-mixing method for 22.2 multichannel system reproduction
abstract
This paper proposes a general multichannel system reproduction method. Firstly, relative to original multichannel system, a general global model is build up by guaranteeing sound pressure and the direction of particle velocity at the receiving point constant, and making the square error of particle velocity magnitude at the receiving point as little as possible. Then the model is equivalent to a least squares problems with non-negative constraints, it can be worked out by existing mature algorithms, and the global optimal solution of simplifying multichannel system are obtained. The proposed method can be used to simplify 22.2 multichannel system to 10.2 and 8.2 multichannel system, objective and subjective experimental results demonstrate that it performs better than traditional method.
Song Wang 0011, Ruimin Hu, Xiaochen Wang 0001, Yuhong Yang 0001, Weiping Tu
ICASSP4
2015 Spatial perception reproduction of sound events based on sound property coincidences
abstract
Sound pressure and particle velocity are used to reproduce sound signals in multichannel systems. The two sound properties were estimated step by step and particle velocity was scaled due to ill-conditioned equations in Ando's study. We explore a new system of equations to maintain both sound pressure and particle velocity. The weight equations are solved in a non-traditional way to figure out exact solutions. Based on the proposed method, the perception of the direction of a sound event and the distance to the listening point are both reproduced correctly in a three-dimension reproduction system. The comparison between the proposed method and Ando's method is outlined and the proposed method is more flexible and useful. Objective evaluation shows the wavefront in the proposed method is more accurate than Ando's method and subjective evaluation confirms that the proposed method improves the spatial perception of sound events.
Maosheng Zhang, Ruimin Hu, Xiaochen Wang 0001, Dengshi Li
ICME4
2015 Unequal error protection for S3AC coding based on expanding window fountain codes
abstract
This paper presents a coding scheme to improve the performance of three-dimensional (3D) audio. The scheme is designed by the idea of joint source channel coding (JSCC) and is implemented by Expanding Window Fountain (EWF) codes for the 3D audio bitstreams after source coding of spatial squeeze surround audio coding (S3AC). EWF is one of unequal error protection (UEP) LT codes and the proposed scheme is achieved by that method for the case of two levels. Different from other transmissions with equal error protection (EEP) for each part of bitstreams, when transmitting the two parts of the bitstreams with downmixed mono signals and spatial side information after S3AC coding, this approach provides more protection to the part of spatial side information and comparatively less protection to the downmixed mono signals, and results in the improvement of the performance of the reproduced 3D audio especially on the aspect of the spatial perception. Objective simulation experiment has shown the proposed UEP scheme achieves a better performance than the EEP scheme on the aspect of spatial perception, for the bits error rates (BER) of spatial parameters can decrease dramatically to a low value of about 10-4, but BERs of downmixed mono signals and the case of EEP just decrease slightly to about 10-3.
Liuyue Su, Ruimin Hu, Xiaochen Wang 0001
ISCC5
2014 A 3D audio coding technique based on extracting the distance parameter
abstract
This paper presents a compression technique to improve the quality of three-dimensional (3D) audio produced by multiple loudspeaker channels or by headphone. The approach is based on extracting the side information of spatial sound sources within the three-dimensional space when capturing the sound sources. Different from other compression technique, the distances of sound sources are included in the side information. The separated signals of different sound sources are downmixed into one mono or stereo audio signal with the side information. The resulting downmixed signal is then compressed with traditional audio coder, resulting in a better perceptual quality of 3D audio by adding the distance parameter in the side information, and maintaining a low bit rates comparable with directional audio coding (DirAC).
Ruimin Hu, Liuyue Su, Weiping Tu, Xiaochen Wang 0001, Yuhong Yang 0001, Shi Dong 0004, Song Wang 0011, Maosheng Zhang, Furong Lei, Shiqing Li
ICME5
2014 The Perceptual Characteristics of 3D Orientation
Ruimin Hu, Weiping Tu, Xiaochen Wang 0001
MMM (2)5
2014 Comments on "Algorithmic Aspects of Hardware/Software Partitioning: 1D Search Algorithms"
abstract
In this paper, the work inis analyzed. An error in its theoretical description part is pointed out and illustrated by a simple example. A modification suggestion is proposed to make the theoretical description of the workmore deliberate and thus being used appropriately.
Hao-Jun Quan, Tao Zhang 0025, Qiang Liu 0011, Jichang Guo, Xiaochen Wang 0001, Ruimin Hu
IEEE Trans. Computers5
2013 An expanded Mid/Side coding for 3D audio signal compression
abstract
Three dimensional (3D) audio technologies are booming with the success of 3D video technology. The sharply increased audio channels make its huge data unacceptable for transmitting bandwidth and storage media. This paper investigates the conventional Mid/Side (M/S) coding method, and expands it to a Three-channel Dependent M/S coding (3D-M/S) method. 3D-M/S perform sum and difference coding based on three channels instead of conventional two channels, and corresponding transform matrixes are presented. Furthermore, a framework is proposed to enable 3D-M/S compress any number of audio channels. Experiment shows proposed method obtains 25.4% objective quality improvement comparing with independent channel coding, and only increases 11.3% complexity comparing with the 29.9% of PCA method.
Shi Dong 0004, Ruimin Hu, Xiaochen Wang 0001, Weiping Tu
ICASSP3