Fuliang Yin

dblp:25/3186 · DBLP profile ↗
← Back
55ranked-venue papers
0as first author
15since 2021 · last 2026
0000-0003-2642-8489ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 since 2021Computer networks · 8 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 A Neural Codec for Bone-Conducted Speech With Joint Denoising and Bandwidth Extension
abstract
In Internet of Things (IoT) communication scenarios such as wearable and industrial systems operating in extremely noisy environments, Bone-Conducted Microphone (BCM) speech inevitably suffers from noise interference. To address this issue, a neural codec for bone-conducted speech with joint denoising and bandwidth extension is proposed in this paper. Speciffcally, an encoder–decoder architecture composed of the MP-SENet and the BigVGAN vocoder is built as the backbone. On this basis, a Forward Time-Frequency Transformer (FTM) is introduced to effectively capture both local and global time-frequency dependencies in the speech signal. To improve the capacity of speech signal modeling, a Backward Time-Frequency Transformer (BTM) is presented to model the information flow in reverse, complementing the forward modeling and enabling more comprehensive extraction of time-frequency features. Subsequently, KalmanNet is designed to enhance the ability to capture temporal dependencies and dynamic information, compensating for the transformer’s limited effectiveness in modeling continuous dynamics. Finally, a Dual-Path Magnitude-Phase (DPMP) module is proposed to feature an interactive dual-stream architecture where magnitude and phase streams communicate with each other to supplement high-frequency components of the speech signal. The proposed method enables simultaneous speech encoding and enhancement under different SNRs using a single network. Furthermore, it supports coding at different bitrates without requiring modifications to the network architecture or retraining. Simulation experiments demonstrate the feasibility of the proposed approach, highlighting its potential for robust and efficient speech transmission in IoT-enabled communication.
Zhe Chen 0005, Fuliang Yin
IEEE Internet Things J.3
2026 Large-Scale Distributed Acoustic SLAM Based on Multiblock Alternating Direction Method of Multipliers
abstract
In large-scale acoustic SLAM, the rapid growth of measurement data significantly increases the computational burden of estimating robot trajectories and sound source positions. To overcome this challenge, a distributed acoustic SLAM method based on multi-block Alternating Direction Method of Multipliers (ADMM) is proposed in this paper. Firstly, to eliminate the influence of robot motion accumulated errors on the localization performance, the factor graph is adopted to construct the global consistency framework for multi-robot trajectories and multi-source positions. Then, the global factor graph is segmented into multiple interrelated subgraphs to reduce the computational complexity of SLAM in large-scale scenes. To further enhance robustness, the Huber function is incorporated to handle anomalous observations, and the multi-block ADMM is adopted to enable distributed optimization across subgraphs. Finally, the joint localization of robot motion trajectories and sound sources is obtained by fusing the local solutions from each subgraph via the average consensus algorithm. The simulation and real-world experiments show that the proposed method can estimate effectively the robot motion trajectories and sound source positions with lower computational complexity, and be robust to anomalous observations in large-scale scenes.
Zhe Chen 0005, Fuliang Yin
IEEE Internet Things J.3
2026 Blind compressed image diffusion restoration based on content prior and dense residual connection driven transformer
Shuang Yue, Zhe Chen 0005, Fuliang Yin
J. Vis. Commun. Image Represent.3
2025 Robust Fusion of Bone and Air-Conducted Sensors for Speech Enhancement with Adaptive Temporal-Frequency Attention
abstract
The multi-modal speech enhancement method has improved performance due to the diverse sources of its input data, which includes low-distortion air-conducted (AC) signals and low-noise bone-conducted (BC) signals. In light of these considerations, a novel complex-domain deep learning-based BC and AC speech fusion enhancement method is proposed. Specifically, the BC and AC spectrum is fed into a pre-fusion module firstly. Then the preliminary fused feature is concatenated with the original signals and fed into encoder, followed by adaptive time-frequency attention modules, in which the self-attention mechanism is employed to reconstruct the encoded features in both the time and frequency axes. After that, two masks are generated by decoders and multiplied with the original BC and AC spectrum. Finally, the two signals are summed and converted to the time-domain. The disabled situation when only one sensor works is considered to improve the robustness by introducing a probability of failure in the process of feeding the data to train the model. Experiments demonstrate that the proposed method outperforms existing fusion approaches, especially in the aspect of robustness against one invalid input channel.
Zhenglong Liu, Zhe Chen 0005, Fuliang Yin
ICASSP3
2025 A Robust and Secure Audio Zero-Watermarking Scheme Based on Novel Chaotic Map and Dual-Tree Complex Wavelet Transform
abstract
Zero-watermarking is recognized as a pivotal technology for copyright protection. However, the existing audio zero-watermarking schemes are often limited to insufficient robustness and inadequate security. To address these issues, a robust and secure audio zero-watermarking scheme based on novel chaotic map and dual-tree complex wavelet transform (DT-CWT) is proposed. Specifically, a two-dimensional logistic map with improving chaos complexity (2D-LMICC) is designed. Chaotic performance analysis demonstrates that 2D-LMICC exhibits superior randomness compared to the existing chaotic systems, thus enhancing the security of watermarking system. Then, an audio zero-watermark construction method is proposed. The carrier audio signal is firstly performed by DT-CWT to obtain coefficients of the low and middle frequency bands with good robustness. After that, these coefficients are processed by singular value decomposition to extract the robust features, and modulated with the watermark to generate plaintext zero-watermark. Finally, the plaintext zero-watermark information is encrypted based on the proposed 2D-LMICC to ensure information security. Extensive experiments demonstrate that the proposed zero-watermarking scheme outperforms the state-of-the-art schemes in resisting desynchronization attacks. It reduces the average bit error rate (BER) by 9.97% and 6.45% under 20% time-scale modification (TSM) and 20% pitch-scale modification (PSM), respectively. Additionally, the ciphertext zero-watermarking exhibits the information entropy near the theoretical maximum of 1 and the adjacent element correlation close to theoretical minimum of 0, effectively ensuring zero-watermarking security.
Donghan Li, Zhe Chen 0005, Fuliang Yin
IEEE Internet Things J.3
2025 Robust Distributed MVDR Beamforming in Wireless Acoustic Sensor Networks With Subnetwork Selection
abstract
Wireless acoustic sensor networks (WASNs) based on the Internet of Things (IoT) have gradually showcased significant advantages in speech processing tasks. Since audio signals captured by acoustic sensors are inevitably contaminated by environmental noises, exploring effective speech enhancement techniques is essential in WASNs. To this end, a robust distributed minimum variance distortionless response (MVDR) filtering method for speech enhancement is proposed in this paper. Specifically, a novel signal model for distributed networks is established by jointly exploiting both interchannel and interframe correlations of speech. Then, a centralized MVDR filter for WASNs is presented by utilizing information from all nodes to achieve good speech enhancement performance. Next, the distributed MVDR optimization problem is proposed and converted into its dual form, and further solved in a distributed manner using the primal-dual method of multipliers. Finally, to reduce computational cost and power consumption, the most efficient subnetwork for distributed speech enhancement, which balances the input signal-to-noise ratio and transmission power, is selected via a data-driven comparison algorithm. The proposed method can effectively suppress interference noise and improve speech quality and is robust to network topology changes and sound source movement. In addition, it does not need to activate all nodes during consensus iterations, thereby mitigating the communication and computational load. Experimental results under different conditions validate the effectiveness of the proposed method.
Qingying Zhao, Zhe Chen 0005, Fuliang Yin
IEEE Internet Things J.3
2025 Dual-Stream Multiscale Attention Monocular Depth Estimation Network
abstract
Deep-learning-based methods have shown superior performance in monocular depth estimation tasks. However, the existing methods often overlook small-scale objects and vertical information while suffering from edge blurring and loss of low-texture information issues. To remedy the issue, a dual-stream multiscale attention network (DMA-Net) for monocular depth estimation is proposed, featuring an encoder-decoder pattern. Specifically, two scales of inputs are, respectively, fed into the pretrained ResNeXt-101 to extract diverse image features. Then, a multiscale attention feature fusion model is constructed, where self-attention dilated convolution blocks effectively capture multiscale global features with long-distance dependencies and feature fusion blocks promote information exchange between the two tributaries, further reinforcing features. Next, a guiding decoder is designed to refine the restored depth map by assistively integrating the outputs of each encoder layer, and exploit efficient channel attention network to recalibrate the meaningful information. Finally, vertical information extractor is utilized to capture vertical features for enhancing the restore ability of longitudinal depth details. Extensive experiments are conducted on the KITTI and NYU Depth V2 datasets, and the results show that the proposed DMA-Net outperforms all previous methods, achieving competitive results on the majority of the metrics.
Ying Zou 0015, Zhe Chen 0005, Fuliang Yin
IEEE Internet Things J.3
2025 High-Order Multi-Scale Attention and Vertical Discriminator Enhanced CLIP for Monocular Depth Estimation
abstract
Multimodal monocular depth estimation methods based on deep learning have achieved competitive performance in recent years. However, the existing Contrastive Language-Image Pre-training (CLIP)-based multimodal networks often suffer from incomplete fusion of two modalities and lack multi-scale contextual information. To remedy these issues, this paper proposes a high-order feature and attention-assisted CLIP model HoCLIP for monocular depth estimation. Specifically, with the CLIP model as the backbone, Matrix Power Normalization Covariance Pooling (MPN-COV) technique is employed for high-order statistical modeling to capture image features by the visual encoder. These features are then combined with learnable deep prompts before being fed into the text encoder, facilitating enhanced fusion of text and image and enabling the extraction of more intricate statistical information and spatial structure. Furthermore, the Efficient Multi-Scale Attention (EMA)-Decoder is utilized for the reconstruction of depth maps. This structure captures contextual information across different scales, establishes long-range dependencies between features, and meticulously preserves spatial position information. Finally, a vertical discriminator with embedded vertical attention is integrated into the model’s final stages to capture vertical features and refine depth map generation. The extensive experiments on the NYU Depth V2 and KITTI datasets are conducted, and the results show that the proposed method has a decisive improvement over the state-of-the-art multimodal methods and exhibits robust competitiveness across all metrics.
Ying Zou 0015, Zhe Chen 0005, Fuliang Yin
IEEE Trans. Circuits Syst. Video Technol.3
2023 Acoustic SLAM With Moving Sound Event Based on Auxiliary Microphone Arrays
abstract
Acoustic simultaneous localization and mapping (ASLAM) aim to map the positions of sound sources while passively localizing the microphone array embedded in the robot platform. In this paper, an ASLAM method with auxiliary microphone arrays based on dual interacting multiple models and unscented Kalman filter (D-IMM-UKF) is proposed for the single moving source scenario. Firstly, a dual-unscented Kalman filter is presented, which can simultaneously track the robot and the speaker. Then, the interacting multiple models are adopted for the different motion dynamics of a robot and a speaker in space. To avoid the underdetermined condition when only the acoustic information is available, a small number of static microphone arrays are employed. Finally, the moving robot’s and speaker’s positions are estimated by the D-IMM-UKF algorithm. It can obtain the trajectories of the robot’s and speaker’s movements smoothly with good tracking accuracy. Experimental results verify the effectiveness of the proposed method.
De Hu, Zhe Chen 0005, Fuliang Yin
IEEE Trans. Intell. Transp. Syst.3
2022 Deep mutual information multi-view representation for visual recognition
Xianfa Xu, Zhe Chen 0005, Fuliang Yin
Appl. Intell.3
2022 Lightweight multi-scale aggregated residual attention networks for image super-resolution
Shurong Pang, Zhe Chen 0005, Fuliang Yin
Multim. Tools Appl.3
2021 Monocular Depth Estimation With Multi-Scale Feature Fusion
abstract
Depth estimation from a single image is a crucial but challenging task for reconstructing 3D structures and inferring scene geometry. However, most existing methods fail to extract more detailed information and estimate the distant small-scale objects well. In this paper, we propose a monocular depth estimation based on multi-scale feature fusion. Specifically, to obtain input features of different scales, we first feed the input images of different scales to pre-trained residual networks with sharing weights. Then, an attention mechanism is used to learn the salient features at different scales, which can integrate detailed information at large scale feature maps and scene information at small scale feature maps. Furthermore, inspired by the dense atrous spatial pyramid pooling in semantic segmentation, we build a multi-scale feature fusion dense pyramid to further improve the ability of the feature extraction. Last, a scale-invariant error loss is used to predict depth maps in log space. We evaluate our method on several public benchmark datasets (including NYU Depth V2 and KITTI). The experiment results show that the proposed method obtains better performance than the existing methods and achieves state-of-the-art results.
Xianfa Xu, Zhe Chen 0005, Fuliang Yin
IEEE Signal Process. Lett.3
2021 Geometry Calibration for Acoustic Transceiver Networks Based on Network Newton Distributed Optimization
abstract
Geometry calibration for distributed acoustic sensor networks is becoming increasing popular in the signal processing community. In this paper, a distributed geometry calibration method based on network Newton distributed optimization is proposed for the acoustic transceiver networks where each node consists of a microphone array and a loudspeaker. After collecting the direction-of-arrival and time-difference-of-arrival measurements, a two-stage centralized cost function is formulated to estimate the geometrical configuration of networks, and the corresponding identifiability conditions are discussed. Next, to achieve the distributed calibration, a distributed cost function is established by splitting the centralized cost function into multiple local cost functions. Finally, the distributed geometry calibration is carried out by using network Newton distributed optimization. The proposed method can effectively estimate the geometry structure of acoustic transceiver networks in noisy and reverberant environments. Compared with the existing approaches, it implements the calibration process in a distributed manner, which requires only the local communication among nodes and does not need an external central processor. Experimental results show the validity of the proposed method.
De Hu, Zhe Chen 0005, Fuliang Yin
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Passive Geometry Calibration for Microphone Arrays Based on Distributed Damped Newton Optimization
abstract
Geometry calibration is an inherent challenge in distributed acoustic sensor networks. To mitigate this problem, a passive geometry calibration approach based on distributed damped Newton optimization is proposed. Specifically, a geometric cost function incorporating direction of arrivals (DoAs) and time difference of arrivals (TDoAs) is first formulated, and then its identifiability conditions are given. Next, to achieve a distributed geometry calibration, the cost function is split into multiple local cost functions that are assigned to every node. After that, a distributed damped Newton optimization is presented to retrieve the geometry of microphone nodes and synchronize the internal delay between each two neighboring nodes. Finally, computational complexity and transmission bandwidth requirements are further analyzed. Compared with the existing approaches, the proposed method estimates the geometry structure of microphone networks in a distributed manner. Moreover, it requires a small number of acoustic sources. Experimental results show the validity of the proposed method.
De Hu, Zhe Chen 0005, Fuliang Yin
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Multi-Scale Spatial Attention-Guided Monocular Depth Estimation With Semantic Enhancement
abstract
Depth estimation from single monocular image is a vital but challenging task in 3D vision and scene understanding. Previous unsupervised methods have yielded impressive results, but the predicted depth maps still have several disadvantages such as missing small objects and object edge blurring. To address these problems, a multi-scale spatial attention guided monocular depth estimation method with semantic enhancement is proposed. Specifically, we first construct a multi-scale spatial attention-guided block based on atrous spatial pyramid pooling and spatial attention. Then, the correlation between the left and right views is fully explored by mutual information to obtain a more robust feature representation. Finally, we design a double-path prediction network to simultaneously generate depth maps and semantic labels. The proposed multi-scale spatial attention-guided block can focus more on the objects, especially on small objects. Moreover, the additional semantic information also enables the objects edge in the predicted depth maps more sharper. We conduct comprehensive evaluations on public benchmark datasets, such as KITTI and Make3D. The experiment results well demonstrate the effectiveness of the proposed method and achieve better performance than other self-supervised methods.
Xianfa Xu, Zhe Chen 0005, Fuliang Yin
IEEE Trans. Image Process.3
2020 Distributed multiple speaker tracking based on time delay estimation in microphone array network
abstract
Multiple speaker tracking in distributed microphone array (DMA) network is a challenging task. A critical issue for multiple speaker scenarios is to distinguish the ambiguous observation and associate it to the corresponding speaker, especially under reverberant and noisy environments. To address the problem, a distributed multiple speaker tracking method based on time delay estimation in DMA is proposed in this study. Specifically, the time delay estimated by the generalised cross‐correlation function is treated as an observation. In order to distinguish the observation for each speaker, the possible time delays, refer to as candidates, are extracted based on data association technique. Considering the ambient influence, a time delay estimation strategy is designed to calculate the time delay for each speaker from the candidates. Finally, only the reliable time delays in DMA are propagated throughout the whole network by diffusion fusion algorithm and used for updating the speakers' state within the distributed Kalman filter framework. The proposed approach can track multiple speakers successfully in a non‐centralised manner under reverberant and noisy environments. Simulation results indicate that, compared with other methods, the proposed method can achieve a smaller root mean square error for multiple speaker tracking, especially in adverse conditions.
Zhe Chen 0005, Fuliang Yin
IET Signal Process.3
2020 Analytical Geometry Calibration for Acoustic Transceiver Arrays
abstract
There are many intelligent devices around us that have the ability to emit and receive audio signals, such as smartphones, smartwatches, and tablets. If they constitute an acoustic transceiver network, it can boost the performance of many audio processing tasks. The speaker localization and tracking algorithms using such a network require the prior information of node positions. To acquire this knowledge, we derive an analytical solution for geometry calibration of acoustic transceiver networks where each node consists of a microphone array and a loudspeaker. Specifically, a linear cost function for node orientations is first established using the delivered direction-of-arrival (DoA) measurements, and its analytical solution is derived. Next, based on the DoA and time-of-arrival (ToA) measurements, another cost function for node positions is formulated, which can also be solved in the closed form. Experimental results show that the proposed method can successfully estimate the geometry structure of acoustic transceiver networks in noisy and reverberant environments. In contrast to most state-of-the-art geometry calibration algorithms, which iteratively solve the calibration problem by optimization methods, the proposed method can decrease the computational complexity greatly.
De Hu, Zhe Chen 0005, Fuliang Yin
IEEE Signal Process. Lett.3
2020 Active Sampling Rate Calibration Method for Acoustic Sensor Networks
abstract
The sampling rate mismatches among nodes severely degrade the performances of distributed signal processing methods in acoustic sensor networks. An active sampling rate calibration method based on the multiple signal classification (MUSIC) algorithm is proposed to solve this problem. Specifically, a pre-defined sinusoidal signal is used as the calibration sound source, and the sampling rate mismatch problem is formulated based on the Hadamard product between output signals of pairwise nodes. The Hadamard product is then filtered through the pre-designed filters to obtain the target sinusoidal signals whose frequencies linearly correlate with the sampling rate mismatch. Finally, the sampling rate differences among nodes are estimated by the root-MUSIC algorithm and compensated by the sinc-interpolation. The proposed method can effectively calibrate the sampling rate mismatch even under severe noise and reverberation. Moreover, it has a strong tolerance for deviations among the microphone frequency responses. Experimental results reveal the validity of the proposed method.
Rui Wang 0046, Zhe Chen 0005, Fuliang Yin
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Multi-Pitch Estimation of Polyphonic Music Based on Pseudo Two-Dimensional Spectrum
abstract
Multi-pitch estimation is a fundamental and key problem in music information retrieval, but still remains challenging due to the intrinsic complexity of polyphonic music. To address this problem, a pseudo 2-D spectrum-based method is proposed in this article. The pseudo 2-D spectrum is first constructed to map the time domain signal into the 2-D frequency space, where the harmonic signal exhibits a typical 2-D pattern. Then, pitch estimation is carried out by cross-correlation between the pseudo 2-D spectrum and the fixed 2-D harmonic template. Finally, the pitches of adjacent frames are grouped into pitch contours, where the contours whose lengths are shorter than the minimum note length limitation are discarded. And the remained pitches are refined using the estimates of neighboring frames by removing probable errors and reconstructing estimates. The proposed method exploits the harmonic structure of pitched sounds in a two-dimensional frequency plane, can work in the case where some notes contain few harmonics, and the harmonic overlap proportions are reduced greatly in the harmony cases. The experimental results show that the proposed method achieves promising performance comparing with the state-of-the-art methods on the evaluation datasets, and outperforms the bispectrum-based method on both evaluation datasets.
Weiwei Zhang 0008, Zhe Chen 0005, Fuliang Yin
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 A supervised term ranking model for diversity enhanced biomedical information retrieval
abstract
BACKGROUND: The number of biomedical research articles have increased exponentially with the advancement of biomedicine in recent years. These articles have thus brought a great difficulty in obtaining the needed information of researchers. Information retrieval technologies seek to tackle the problem. However, information needs cannot be completely satisfied by directly introducing the existing information retrieval techniques. Therefore, biomedical information retrieval not only focuses on the relevance of search results, but also aims to promote the completeness of the results, which is referred as the diversity-oriented retrieval. RESULTS: We address the diversity-oriented biomedical retrieval task using a supervised term ranking model. The model is learned through a supervised query expansion process for term refinement. Based on the model, the most relevant and diversified terms are selected to enrich the original query. The expanded query is then fed into a second retrieval to improve the relevance and diversity of search results. To this end, we propose three diversity-oriented optimization strategies in our model, including the diversified term labeling strategy, the biomedical resource-based term features and a diversity-oriented group sampling learning method. Experimental results on TREC Genomics collections demonstrate the effectiveness of the proposed model in improving the relevance and the diversity of search results. CONCLUSIONS: The proposed three strategies jointly contribute to the improvement of biomedical retrieval performance. Our model yields more relevant and diversified results than the state-of-the-art baseline models. Moreover, our method provides a general framework for improving biomedical retrieval performance, and can be used as the basis for future work.
Bo Xu 0009, Hongfei Lin, Liang Yang 0003, Kan Xu, Yi-Jia Zhang 0001, Dongyu Zhang 0001, Jian Wang 0021, Yuan Lin 0001, Fuliang Yin
BMC Bioinform.10
2019 Water rippling shaped clustering strategy for efficient performance of software define wireless sensor networks
Syed Bilal Hussian Shah, Zhe Chen 0005, Fuliang Yin, Awais Ahmad 0001
Peer-to-Peer Netw. Appl.3
2019 DOA-Based Three-Dimensional Node Geometry Calibration in Acoustic Sensor Networks and Its Cramér-Rao Bound and Sensitivity Analysis
abstract
Acoustic sensor networks (ASNs) are widely applied in scenarios like teleconference, teaching, and theatre. ASNs can be used in tracking speakers, enhancing the speaker's speech and human-machine interactions, etc., but the geometric structure of the ASN has to be calibrated. ASN geometry calibration is a challenging task due to the irregular geometric structures of ASNs. A three-dimensional (3D) node geometry calibration approach based on direction of arrival (DOA) measurements and artificial bee colony (ABC) algorithm is proposed in this paper. The theoretical DOAs of sound sources relative to nodes are first derived based on 3D rotation matrices and translation vectors, and the corresponding measured DOAs are estimated by the time-difference-of-arrival. Then, the node geometry calibration problem is formulated as the minimization of a cost function measuring the mismatch between theoretical and measured DOAs, and such non-convex minimization is effectively solved by the ABC algorithm. Next, Cramér-Rao bound is presented to provide a theoretical lower bound for DOA-based node geometry calibration. Finally, the sensitivity of the proposed method to the sound source position error is discussed. The proposed method can calibrate node geometry positions successfully in both 2D plane and 3D space and requires no information transmission among nodes when the positions of few sound sources and the relative geometry of microphones in each node are known. Experimental results reveal the validity of the proposed node geometry calibration method.
Rui Wang 0046, Zhe Chen 0005, Fuliang Yin
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 Improve Diversity-oriented Biomedical Information Retrieval using Supervised Query Expansion
Bo Xu 0009, Hongfei Lin, Liang Yang 0003, Kan Xu, Yi-Jia Zhang 0001, Dongyu Zhang 0001, Jian Wang 0021, Yuan Lin 0001, Fuliang Yin
BIBM10
2018 Energy and interoperable aware routing for throughput optimization in clustered IoT-wireless sensor networks
Syed Bilal Hussain Shah, Zhe Chen 0005, Fuliang Yin, Niqash Ahmad
Future Gener. Comput. Syst.3
2018 Melody Extraction From Polyphonic Music Using Particle Filter and Dynamic Programming
abstract
Melody extraction from polyphonic music is one important but challenging task in the music information retrieval community. In this paper, a new melody extraction method based on the particle filter and dynamic programming is proposed. The constant-Q transform is first introduced for multiresolution spectral analysis of polyphonic music. Then, the melody extraction is modeled in the Bayesian filtering framework, and the particle filter is used to get a rough melody contour. Specially, the pitch transition probability of adjacent frames is approximated according to the statistical analysis based on one publicly available dataset, and the likelihood of frame-wise pitches is defined by considering pitch salience, spectral smoothness, and timbre similarity. After that, the preliminary melodic contour obtained by particle filter is smoothed to achieve the frame-wise pitch range limitation. Finally, the dynamic programming is used to accurately track the final melodic contour. The proposed method requires no prior information, and is suitable for both instrumental and vocal melodies. The experimental results show that the performances of the proposed method is robust among four publicly available datasets comparing with the state-of-the-art methods, and it achieves the highest averaged raw pitch accuracy and raw chroma accuracy performances with lower octave errors.
Weiwei Zhang 0008, Zhe Chen 0005, Fuliang Yin, Qiaoling Zhang
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 Joint detection and decoding for physical-layer network coding in power line channel
abstract
In this study, a joint detection and generalised sum‐product decoding algorithm for a two‐way relay physical‐layer network coding (PLNC) system in power line communication is proposed. Distinct from the existing works which mainly focused on separate or iterative PLNC detection and decoding, the proposed algorithm is a belief propagation‐based joint PLNC detection and decoding algorithm, which integrates the PLNC, symbol detector and channel decoder together. The proposed algorithm is described by a joint factor graph, in which the detection nodes are connected to the factor graph of a generalised sum‐product decoder. The PLNC demodulation is redesigned to make sure that the degree of each detection node in the joint factor graph is more than one. A low computational complexity solution of the proposed joint decoding algorithm based on the majority‐logic algorithm is also proposed. The simulation results reveal the validity of the proposed algorithms.
Zhe Chen 0005, Fuliang Yin
IET Commun.3
2017 Eradication of pilot contamination and zero forcing precoding in the multi-cell TDD massive MIMO systems
abstract
In massive multiple‐input multiple‐output (M‐MIMO) systems, the base station (BS) estimates the downlink and uplink channels using the uplink pilot training in conjunction with channel reciprocity property of time division duplex (TDD) operation. However, this channel estimation is contaminated due to the use of non‐orthogonal pilot sequences for the uplink training in neighbouring cells that results in inter‐cell interference. This study proposes an effective pilot contamination elimination scheme along with zero forcing precoding for multi‐cell TDD M‐MIMO systems. In the proposed scheme, each BS is assigned with a specific orthogonal variable spreading factor (OVSF) code row and a set of Zadoff–Chu (ZC) sequences are used for uplink pilot training. Then, the set of ZC sequences is multiplied element‐wise at each BS with its OVSF code row to generate the orthogonality among the pilot sequences across the neighbouring cells. The proposed method can promise uncontaminated channel estimation and interference free downlink transmission without requiring the knowledge of the channels’ second‐order statistics and without increasing the training overhead by a factor equal to the number of interfering cells. The simulation results authenticate the effectiveness and superiority of the proposed scheme over other schemes.
Sajjad Ali Memon, Zhe Chen 0005, Fuliang Yin
IET Commun.3
2017 Speaker Tracking Based on Distributed Particle Filter in Distributed Microphone Networks
abstract
A speaker tracking method based on a distributed particle filter (DPF) for distributed microphone networks is proposed in this paper. First, the generalized cross-correlation (GCC) function is estimated at each node. To cope with the spurious effects due to the noise or reverberation, multiple delays related to the largest local peaks of the GCC constitute the local observation. Next, based on an optimal fusion rule, a modified DPF is presented, and a modified multiple-hypothesis model is also developed as its likelihood function by incorporating the information of the GCC. Finally, the modified DPF is used to track a moving speaker with a distributed microphone network. The proposed method requires only local communication among neighboring nodes, and is robust against nodes failure. Simulation and real-world experimental results demonstrate the validity of the proposed method.
Qiaoling Zhang, Zhe Chen 0005, Fuliang Yin
IEEE Trans. Syst. Man Cybern. Syst.3
2016 Distributed Marginalized Auxiliary Particle Filter for Speaker Tracking in Distributed Microphone Networks
abstract
In this paper, a distributed marginalized auxiliary particle filter (DMAPF) is proposed for speaker tracking in distributed microphone networks. After marginalizing the state-space model, the speaker's velocity and position are estimated using the distributed Kalman filter and the distributed auxiliary particle filter (APF), respectively. To overcome the adverse effects of noise and reverberation, a time difference of arrival selection scheme is presented to construct the local observation vector, based on the generalized cross-correlation function of the microphone pair signals at each node. Next, the multiple-hypothesis model is used as the local likelihood function of the DMAPF. Finally, the DMAPF is employed to estimate the time-varying positions of a moving speaker. The proposed method combines the strengths of the marginalized particle filter, APF, and distributed estimation. It can track the speaker successfully in noisy and reverberant environments. Moreover, it requires only local communication among neighboring nodes, and is scalable for speaker tracking. Experimental results reveal the validity of the proposed speaker tracking method.
Qiaoling Zhang, Zhe Chen 0005, Fuliang Yin
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Distributed IMM-Unscented Kalman Filter for Speaker Tracking in Microphone Array Networks
abstract
In this paper, we first propose a distributed unscented Kalman filter (DUKF) to overcome the nonlinearity of measurement model in speaker tracking. Next, for the different motion dynamics of a speaker in the in-door environment, we introduce the interacting multiple model (IMM) algorithm and propose a distributed interacting multiple model-unscented Kalman filter (IMM-UKF) for estimating time-varying speaker's positions in a microphone array network. In the distributed IMM-UKF based speaker tracking method, the time difference of arrival (TDOA) of the speech signals received by a pair of microphones at each node is estimated by the generalized cross-correlation (GCC) method, then the distributed IMM-UKF is used to track a speaker whose position and speed significantly vary over time in a microphone array network. The proposed method can estimate speaker's positions globally in the network and obtain a smoothed trajectory of the speaker's movement robustly in noisy and reverberant environments, and it is scalable for speaker tracking. Simulation and real-world experiment results reveal the effectiveness of the proposed speaker tracking method.
Zhe Chen 0005, Fuliang Yin
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 A Novel Hierarchical Decomposition Vector Quantization Method for High-Order LPC Parameters
abstract
The paper investigates vector quantization coding of high-order (e.g., 20th-50th order) linear prediction coding (LPC) parameters, and proposes a novel hierarchical decomposition vector quantization method for a scalable speech coding framework with variable orders of LPC analysis. Instead of vector quantizing the whole group of LPC parameters in the linear spectral frequency (LSF) domain directly, the proposed method decomposes the high-order LPC model into several low-order (e.g., 10th-order) LPC models, and vector quantizes them in the LSF domain separately. For the decomposition, the high-order LPC model is converted into a group of reflection coefficients at first, and then the group is split into several subgroups and converted into multiple low-order LPC models. It is shown that the proposed method is naturally suitable for a scalable coding framework where the information of the decomposed low-order LPC models can be encoded into a multi-layered bitstream and can be combined in a progressive way to recover the high-order LPC information. Experiments in a scalable coding framework with variable LPC analysis orders (10-50) reveal that, compared to a direct vector quantization scheme, the proposed method can reduce the size of the codebook and the number of coding bits significantly, and can also efficiently reduce the computation cost.
Lin Wang 0009, Zhe Chen 0005, Fuliang Yin
IEEE ACM Trans. Audio Speech Lang. Process.3
2012 Mask Cut Optimization in Two-Dimensional Phase Unwrapping
abstract
This letter proposes an improved mask cut method for 2-D phase unwrapping. The method is designed to generate thin mask cuts to balance the residues. Network flow ideas are taken to build the strategy of mask cut generation. The predecessor index is recorded for every pixel during searching, and the path from the current residue to its last residue ancestor is traced by the predecessor index information. This path is taken as an optimal path from residue to residue. All the pixels on these paths compose the final mask cuts. The proposed method successfully prevents most of the unnecessary pixels from going into the mask cuts. Comparing with the original mask cut method, experimental results show that the proposed method generates thinner and more accurate mask cuts than that generated by the original method. A significant improvement on the unwrapping result is achieved by the proposed method.
Dapeng Gao, Fuliang Yin
IEEE Geosci. Remote. Sens. Lett.2
2011 A Region-Growing Permutation Alignment Approach in Frequency-Domain Blind Source Separation of Speech Mixtures
abstract
The convolutive blind source separation (BSS) problem can be solved efficiently in the frequency domain, where instantaneous BSS is performed separately in each frequency bin. However, the permutation ambiguity in each frequency bin should be resolved so that the separated frequency components from the same source are grouped together. To solve the permutation problem, this paper presents a new alignment method based on an inter-frequency dependence measure: the powers of separated signals. Bin-wise permutation alignment is applied first across all frequency bins, using the correlation of separated signal powers; then the full frequency band is partitioned into small regions based on the bin-wise permutation alignment result. Finally, region-wise permutation alignment is performed in a region-growing manner. The region-wise permutation correction scheme minimizes the spreading of the misalignment at isolated frequency bins to others, hence to improve permutation alignment. Experiment results in simulated and real environments verify the effectiveness of the proposed method. Analysis demonstrates that the proposed frequency-domain BSS method is computationally efficient.
Lin Wang 0009, Heping Ding, Fuliang Yin
IEEE Trans. Speech Audio Process.3
2010 Bandwidth expansion of speech based on wavelet transform modulus maxima vector mapping
abstract
A novel approach to speech bandwidth expansion based on wavelet transform modulus maxima vector mapping is proposed. By taking advantage of the similarity of the modulus maxima vectors between narrowband and wideband wavelet-analyzed signals a neural network mapping structure can be established to perform bandwidth expansion given only the narrowband version of speech. Since the proposed algorithm works on the time-domain waveforms it offers a flexibility of variable-length frame selection that facilitates low delay and potentially data-dependent speech segment processing to further improve the speech quality. Evaluations based on both objective and subjective measures show that the proposed bandwidth expansion approach results in highquality synthesized wideband speech with little perceivable distortion from the original wideband speech signals.
Zhe Chen 0005, You-Chi Cheng, Fuliang Yin
INTERSPEECH3
2009 Blind Source Separation Based on Cumulants With Time and Frequency Non-Properties
abstract
This paper presents new results on blind separation of instantaneously mixed independent sources based on high-order statistics together with their time and frequency non-properties (i.e., the non-stationarity and non-whiteness of sources). Separation criteria of mixtures are established on a set of cumulants at different time instants using the non-stationarity of sources and/or time-delayed cumulants using the non-whiteness of sources. It is shown that cumulants at different time instants and time-delayed cumulants can be used as criteria for blind source separation (BSS). Furthermore, it is proved that the cumulant-based separation criteria are directly related to the separability conditions. Batch-data and online learning rules are developed based on the joint diagonalization of symmetric fourth-order cumulant matrices, and the learning rules are further simplified to correlation-based BSS algorithms. In addition, an initialization strategy is proposed for improving the convergence of the learning rules. Simulation results are given to demonstrate the validity and performance of the algorithms.
Tiemin Mei, Fuliang Yin, Jun Wang 0002
IEEE Trans. Speech Audio Process.2
2008 A blind source separation-based method for multiple images encryption
Qiu-Hua Lin, Fuliang Yin, Tiemin Mei, Hualou Liang
Image Vis. Comput.2
2008 Blind source separation for convolutive mixtures based on the joint diagonalization of power spectral density matrices
Tiemin Mei, Alfred Mertins, Fuliang Yin, Jiangtao Xi, Joe F. Chicharo
Signal Process.3
2007 A Speech Enhancement Method in Subband
Fuliang Yin
ISNN (3)4
2007 A fast algorithm for one-unit ICA-R
Qiu-Hua Lin, Yong-Rui Zheng, Fuliang Yin, Hualou Liang, Vince D. Calhoun
Inf. Sci.3
2006 A Fast Decryption Algorithm for BSS-Based Image Encryption
Qiu-Hua Lin, Fuliang Yin, Hualou Liang
ISNN (2)2
2006 A Blind Source Separation Based Multi-bit Digital Audio Watermarking Scheme
Xiaoyan Ding, Chong Wang 0003, Fuliang Yin
ISNN (2)4
2006 A Robust VAD Method for Array Signals
Fuliang Yin
ISNN (2)3
2006 Blind Source Separation Based on Time-Domain Optimization of a Frequency-Domain Independence Criterion
abstract
A new technique for the blind separation of convolutive mixtures is proposed in this paper. Inspired by the works of Amari, Sabala , and Rahbar, we firstly start from the application of Kullback-Leibler divergence in frequency domain, and then we integrate Kullback-Leibler divergence over the whole frequency range of interest to yield a new objective function which turns out to be time-domain variable dependent. In other words, the objective function is derived in frequency domain which can be optimized with respect to time domain variables. The proposed technique has the advantages of frequency domain approaches and is suitable for very long mixing channels, but does not suffer from the local permutation problem as the separation is achieved in time-domain
Tiemin Mei, Jiangtao Xi, Fuliang Yin, Alfred Mertins, Joe F. Chicharo
IEEE Trans. Speech Audio Process.3
2005 A fast running Hartley transform algorithm and its application in adaptive signal enhancement
abstract
A fast recursive algorithm for computation of the running discrete Hartley transform (RDHT) is presented. This method is based on the relation between the running discrete Fourier transform (RDFT) and the RDHT. The number of operations for the proposed recursive algorithm is only 2/N (N=length of the transform) of the direct computation of the RDHT. It also provides substantial computational savings compared with the recursive RDFT algorithm. A transform-domain adaptive digital filter is implemented based on the presented algorithm. Simulation results of its implementation on an adaptive line enhancer are given to demonstrate the efficiency of the presented fast algorithm.
Zhiyue Lin, Fuliang Yin, Richard W. McCallum
ICASSP (4)2
2005 Blind Source Separation-Based Encryption of Images and Speeches
Qiu-Hua Lin, Fuliang Yin, Hualou Liang
ISNN (2)2
2005 A Digital Audio Watermarking Scheme Based on Blind Source Separation
Chong Wang 0003, Xiangping Cong, Fuliang Yin
ISNN (2)4
2005 A New Speech Enhancement Method for Adverse Noise Environment
Fuliang Yin
ISNN (2)4
2005 Joint Diagonalization of Power Spectral Density Matrices for Blind Source Separation of Convolutive Mixtures
Tiemin Mei, Jiangtao Xi, Fuliang Yin, Joe F. Chicharo
ISNN (2)3
2005 An Audio Watermarking Scheme with Neural Network
Chong Wang 0003, Xiangping Cong, Fuliang Yin
ISNN (2)4
2005 A Subband Adaptive Learning Algorithm for Microphone Array Based Speech Enhancement
Dongxia Wang 0005, Fuliang Yin
ISNN (2)2
2004 Speech Segregation Using Constrained ICA
Qiu-Hua Lin, Yong-Rui Zheng, Fuliang Yin, Hualou Liang
ISNN (1)3
2004 Frequency-Domain Separation Algorithms for Instantaneous Mixtures
Tiemin Mei, Fuliang Yin
ISNN (1)2
2004 Cumulant-Based Blind Separation of Convolutive Mixtures
Tiemin Mei, Fuliang Yin, Jiangtao Xi, Joe F. Chicharo
ISNN (1)2
2004 Neural Congestion Control Algorithm in ATM Networks with Multiple Node
Ruijun Zhu, Fuliang Yin, Tianshuang Qiu
ISNN (2)2
2004 Blind separation of convolutive mixtures by decorrelation
Tiemin Mei, Fuliang Yin
Signal Process.2