Jiangtao Xi

dblp:06/4004 · DBLP profile ↗
← Back
41ranked-venue papers
5as first author
13since 2021 · last 2027
0000-0002-5550-1975ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 5 first-author · 1 since 2021Artificial intelligence and machine learning · 11 · 6 since 2021Computer networks · 10 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2027 DAFR-Net: Lightweight deformable adaptive feature reconstruction network for tiny object detection in remote sensing images
Guangshuai Gao, Yunqi Shang, Jiangtao Xi
Expert Syst. Appl.6
2026 On the Ambiguity Functions of Delay-Doppler Domain Multicarrier Modulation (DDMC)
Jun Tong, Jinhong Yuan, Akram Shafie, Jiangtao Xi
ICC4
2026 PFI-Net: A parallel feature interaction network for infrared and visible target detection
Xiaoxia Wang 0002, Jiangtao Xi, Fengbao Yang, Yunjia Yang
Pattern Recognit.2
2026 Equivalent Sampled Delay-Doppler (ESDD) Channel Models for ODDM Over Highly-Spread Channels
abstract
Delay-Doppler (DD) domain modulation such as orthogonal DD division multiplexing (ODDM) has been recently explored for communications over doubly selective channels. This paper derives equivalent sampled (on-grid) DD (ESDD) channel models for the effective discrete-time channels in ODDM systems over highly-spread off-grid physical channels with delay and Doppler shifts that can exceed the subpulse spacing and subtone spacing, respectively, of the DD orthogonal pulse (DDOP). The derived ESDD models account for i) off-grid delay and Doppler shifts present in practical physical channels, ii) sample-wise pulse shaping at the transmitter, and iii) matched filtering and windowing at the receiver. We then investigate the supports of the ESDD models and their implications on the input-output (IO) relation of ODDM adopting more general pulses and windows over highly-spread physical channels. In particular, we show that the digital sequences of ODDMcouple finelywith on-grid discrete-time DD channels, which leads to compact IO relation. We also analyze the folding and aliasing of the ESDD channel and their influence on the fading effect and predictability of effective channels experienced by the DD-domain symbols of ODDM. Based on the results, we identify conditions under which the ESDD channels can be estimated directly from embedded DD-domain pilots in a single ODDM frame. Finally, numerical results are provided to further demonstrate the findings of this paper.
Jun Tong, Akram Shafie, Jinhong Yuan, Hai Lin 0001, Jiangtao Xi
IEEE Trans. Wirel. Commun.5
2025 Orthogonal Delay-Doppler Division Multiplexing (ODDM) Modulation Over Highly-Spread Channels
abstract
This paper examines orthogonal delay-Doppler (DD) division multiplexing (ODDM) systems over highly-spread channels characterized by delay and Doppler shifts which can exceed, respectively, the sub-pulse separation and sub-tone separation of the DD orthogonal pulse employed by ODDM. We first derive an equivalent sampled DD (ESDD) channel model with on-grid delay and Doppler shifts specified by the sampling interval and frame duration of the transmission scheme. Our model accounts for off-grid delay and Doppler present in practical channels, samplewise pulse shaping at the transmitter, and matched filtering and windowing at the receiver. By examining their interaction, we show that the time-domain sequences of ODDM implemented approximately (digitally) using discrete Fourier transform (DFT) and inverse DFT (IDFT) and the discrete-time on-grid DD channel are finely coupled, and this leads to compact input-output (IO) relation for ODDM over on-grid channels. We then leverage the ESDD model to describe the IO relation of ODDM adopting more general pulses and windows over highly-spread off-grid channels, which reveals the potential folding (aliasing) of the ESDD channel and its implication on the signal model. We finally present numerical results. It is observed that pulses with shorter duration and windows with lower sidelobes in their spectrum lead to sparser ESDD channels.
Jun Tong, Akram Shafie, Jinhong Yuan, Hai Lin 0001, Jiangtao Xi
ICC5
2025 Provable causal distributed two-time-scale temporal-difference learning with instrumental variables
Jiamei Feng, Qingtao Wu, Ruijuan Zheng, Junlong Zhu, Jiangtao Xi, Mingchuan Zhang
Expert Syst. Appl.6
2024 Cross-Domain Calibration and Boundary Denoising Network for Weakly Supervised Semantic Segmentation
Zhoufeng Liu, Bingrui Li, Shumin Ding, Jiangtao Xi, Chunlei Li 0002
ICPR (3)4
2024 Grant-Free MIMO-NOMA With Differential Modulation for Machine-Type Communications
abstract
This article considers a challenging scenario of machine-type communications, where we assume Internet of Things (IoT) devices send short packets sporadically to an access point (AP) and the devices are not synchronized in the packet level. High-transmission efficiency and low latency are concerned. Motivated by the great potential of multiple-input-multiple-output nonorthogonal multiple access (MIMO-NOMA) in massive access, we design a grant-free MIMO-NOMA scheme, and in particular differential modulation is used so that expensive channel estimation at the receiver (AP) can be bypassed. The receiver at AP needs to carry out active device detection and multidevice data detection. The active user detection is formulated as the estimation of the common support of sparse signals, and a message-passing-based sparse Bayesian learning (SBL) algorithm is designed to solve the problem. Due to the use of differential modulation, we investigate the problem of noncoherent multidevice data detection, and develop a message-passing-based Bayesian data detector, where the constraint of differential modulation is exploited to drastically improve the detection performance, compared to the conventional noncoherent detection scheme. Simulation results demonstrate the effectiveness of the proposed active device detector and noncoherent multidevice data detector.
Yuanyuan Zhang 0005, Zhengdao Yuan, Qinghua Guo 0001, Zhongyong Wang, Jiangtao Xi, Yanguang Yu, Yonghui Li 0001
IEEE Internet Things J.5
2024 Orthogonal Delay-Doppler Division Multiplexing (ODDM) Over General Physical Channels
abstract
This paper investigates the characteristics and performance of orthogonal delay-Doppler division multiplexing (ODDM) modulation over doubly selective physical channels with general delay and Doppler. Assuming that the implementation of the ODDM is based on IDFT/DFT and sample-wise pulse shaping/receiver filtering, we study the input-output (IO) relation for ODDM and characterize the equivalent sampled delay-Doppler (ESDD) domain channel in terms of the parameters of the physical channel and the DD plane orthogonal pulse (DDOP). The established IO relation can describe the patterns of inter-symbol-interference (ISI) and inter-carrier-interference (ICI) for more general delay and Doppler shifts of the physical channel. Based on the results, we also examine the influence of the transmitter configuration on the sparsity of the ESDD channel and on the resulting performance-complexity tradeoff of ODDM. We further introduce a pilot-assisted method of estimating the physical channel parameters by leveraging the derived IO relation and the root-MUSIC algorithm. We also present a low-complexity symbol detector for ODDM systems based on conjugate gradients (CG). Simulation results under various settings show that the error performance of ODDM based on the estimate of the physical channel approaches that with perfect channel state information (CSI) at a low-to-medium signal-to-noise ratio (SNR), but has an increased gap from the perfect CSI case when the SNR increases.
Jun Tong, Jinhong Yuan, Hai Lin 0001, Jiangtao Xi
IEEE Trans. Commun.4
2023 Provable distributed adaptive temporal-difference learning over time-varying networks
Junlong Zhu, Bing Li 0031, Lin Wang 0039, Mingchuan Zhang, Ling Xing 0001, Jiangtao Xi, Qingtao Wu
Expert Syst. Appl.6
2023 Triple critical feature capture network: A triple critical feature capture network for weakly supervised object detection
abstract
Abstract Weakly supervised object detection (WSOD) is becoming increasingly important for computer vision tasks, as it alleviates the burden of manual annotation. Most WSOD techniques rely on multiple instance learning (MIL), which tends to localise the discriminative parts of salient objects instead of the whole object. In addition, network training is often supervised using simple image‐level annotations, without including object quantities or location information. However, this can lead to ambiguous differentiation of object instances, both in terms of location and semantics. To address these issues, propose an end‐to‐end triple critical feature capture network (TCFCNet) for WSOD is proposed. Specifically, a multi‐task branch, which can perform fully supervised classification and regression task, was integrated with a PCL in an end‐to‐end network for refining object locations in an online method. A cyclic parametric dropblock module (CPDM) was then designed to help the detector focus on the contextual information by using cyclic masking techniques to maximise the removal of the discriminative components of an object instance to alleviate the part domination problem. Finally, a feature decoupling module (FDM) is proposed to further reduce the ambiguous distinction of object instances by adaptively constructing robust critical features that adapt to multi‐task branch for classification and regression tasks, which contains a feature enhancement module and task‐specific polarisation functions. Comprehensive experiments are carried out on the challenging Pascal VOC 2007 and VOC 2012 datasets. The proposed method achieves a 54.6% mAP and a 44.3% mAP on the Pascal VOC 2007 and VOC 2012 datasets respectively, showed that our method outperformed existing mainstream techniques by a considerable margin.
Zhoufeng Liu, Kaihua Wang, Chunlei Li 0002, Shunmin Ding, Jiangtao Xi
IET Comput. Vis.5
2023 Inductive Matrix Completion and Root-MUSIC-Based Channel Estimation for Intelligent Reflecting Surface (IRS)-Aided Hybrid MIMO Systems
abstract
This paper studies the estimation of cascaded channels in passive intelligent reflective surface (IRS)-aided multiple-input multiple-output (MIMO) systems employing hybrid precoders and combiners. We propose a low-complexity solution that estimates the channel parameters progressively. The angles of departure (AoDs) and angles of arrival (AoAs) at the transmitter and receiver, respectively, are first estimated using inductive matrix completion (IMC) followed by root-MUSIC-based super-resolution spectrum estimation. Forward-backward spatial smoothing (FBSS) is applied to address the coherence issue. Using the estimated AoAs and AoDs, the training precoders and combiners are then optimized and the angle differences between the AoAs and AoDs at the IRS are estimated using the least squares (LS) method followed by FBSS and the root-MUSIC algorithm. Finally, the composite path gains of the cascaded channel are estimated using on-grid sparse recovery with a small-size dictionary. The simulation results suggest that the proposed estimator can achieve improved channel parameter estimation performance with lower complexity as compared to several recently reported alternatives, thanks to the exploitation of the knowledge of the array responses and low-rankness of the channel using low-complexity algorithms at all the stages.
Khawaja Fahad Masood, Jun Tong, Jiangtao Xi, Jinhong Yuan, Yanguang Yu
IEEE Trans. Wirel. Commun.3
2022 Regularized Covariance Estimation for Polarization Radar Detection in Compound Gaussian Sea Clutter
abstract
This article investigates regularized estimation of Kronecker-structured covariance matrices (CMs) for polarization radar in sea clutter scenarios where the data are assumed to follow the complex elliptically symmetric (CES) distributions with a Kronecker-structured CM. To obtain a well-conditioned estimate of the CM, we add penalty terms of Kullback–Leibler divergence to the negative log-likelihood function of the associated complex angular Gaussian (CAG) distribution. This is shown to be equivalent to regularizing Tyler’s fixed-point equations by shrinkage. A sufficient condition that the solution exists is discussed. An iterative algorithm is applied to solve the resulting fixed-point iterations, and its convergence is proven. In order to solve the critical problem of tuning the shrinkage factors, we then introduce two methods by exploiting oracle approximating shrinkage (OAS) and cross-validation (CV). The proposed estimator, referred to as the robust shrinkage Kronecker estimator (RSKE), is shown to achieve better performance compared with several existing methods when the training samples are limited. Simulations are conducted for validating the RSKE and demonstrating its high performance by using the IPIX 1998 real sea data.
Lei Xie 0009, Zishu He, Jun Tong, Jun Li 0038, Jiangtao Xi
IEEE Trans. Geosci. Remote. Sens.6
2019 Spectrum Sensing Using Multiple Large Eigenvalues and Its Performance Analysis
abstract
Cognitive radio (CR) is a promising technology to address the challenge of spectrum scarcity due to the massive number of objects in the Internet of Things (IoT). Equipping IoT objects with CR capability can also alleviate interference situations and achieve seamless connectivity in IoT. This paper deals with CR spectrum sensing and proposes a new eigenvalue-based detector by exploiting the summation of multiple large eigenvalues of the covariance matrix of received signals. By analyzing the distribution of the sum of the dependent large eigenvalues, we derive an approximate but explicit expression for the theoretical performance of the proposed detector. The theoretical analysis of the proposed detector is validated and its superior performance is demonstrated with real world signals. It is shown that the proposed detector outperforms the existing eigenvalue-based detectors and is more robust against noise uncertainty.
Ming Jin 0001, Qinghua Guo 0001, Youming Li, Jiangtao Xi, Defeng Huang
IEEE Internet Things J.4
2019 Effective Energy Detection for IoT Systems Against Noise Uncertainty at Low SNR
abstract
This paper deals with spectrum sensing for cognitive radio-based Internet of Things (IoT) systems and their coexistence with Long Term Evolution (LTE) systems. Due to the sparsity of the covariance matrix of IoT/LTE signals, we reveal that the likelihood ratio test approximates to energy detection (ED) at low signal to noise ratio. However, the noise (power) uncertainty can degrade the performance of ED severely, especially when low-cost IoT devices are employed for spectrum sensing. To tackle this issue, we derive the relationship among noise power, total power, and autocorrelation coefficient of received signals, and propose an unbiased estimator of noise power without the knowledge of the presence/absence of IoT/LTE signals. We then design a new ED with multiple estimates of noise power from historical and current sensing data, and analyze its theoretical performance. Numerical results are provided to verify the theoretical results and demonstrate the superior performance of the proposed detector. It is shown that, by exploiting sufficient historical sensing data, the performance of the proposed ED can closely approach that of the ideal ED.
Junteng Yao, Ming Jin 0001, Qinghua Guo 0001, Yonghui Li 0001, Jiangtao Xi
IEEE Internet Things J.5
2019 Energy Efficiency of Massive MIMO Systems With Low-Resolution ADCs and Successive Interference Cancellation
abstract
This paper studies the influence of signal detection schemes on the energy efficiency (EE) of uplink multiple-input-multiple-output (MIMO) systems with low-resolution analog-to-digital converters (ADCs). Assuming equal transmission rates for all users, we derive the optimal power allocation and their analytical approximations for zero-forcing (ZF) and ZF successive interference cancellation (ZF-SIC) receivers. Both the cases with perfect channel state information (CSI) and with imperfect CSI are considered. The EE with different receivers is compared. The results indicate that for uplink massive MIMO systems with low-resolution ADCs, the radio-frequency circuit power consumption can be significant because a large number of antennas are required to compensate for the loss due to quantization errors while the number of base station antennas needed with the ZF-SIC receiver is significantly smaller than that with the ZF receiver. Meanwhile, the increase of power consumption of signal processing with ZF-SIC can be moderate, due to the fact that the receiver weights are reused in a coherent block and the dimensionality is reduced. Consequently, the ZF-SIC receiver is able to improve the overall EE for massive MIMO systems with practical ADCs. We also conduct an approximation analysis for a multi-cell scenario with the pilot contamination and inter-cell-interference considered.
Jun Tong, Qinghua Guo 0001, Jiangtao Xi, Yanguang Yu, Zhitao Xiao
IEEE Trans. Wirel. Commun.4
2018 Cross-Validated Bandwidth Selection for Precision Matrix Estimation
abstract
Inverse covariance matrix, a.k.a. precision matrix, has wide applications in signal processing and is often estimated from training samples. The quality of estimation can be poor when the sample support is low. Banding/tapering are effective regularization approaches for covariance and precision matrix estimation but the bandwidth must be properly chosen. This paper investigates the bandwidth selection problem for banding/tapering-based precision matrix estimation. Exploiting a regression analysis interpretation of the precision matrix, we design a data-driven cross-validation (CV) method for automatically tuning the bandwidth. The effectiveness of the proposed method is demonstrated by numerical examples under a quadratic loss.
Jun Tong, Jiangtao Xi, Yanguang Yu, Philip Ogunbona
ICASSP2
2018 Joint spare channel estimation and decoding for orthogonal frequency division multiplexing using combined message passing
abstract
In this study, the authors investigate the use of combined belief propagation (BP), mean field (MF) and expectation propagation (EP) message passing to achieve joint channel estimation and decoding (JCED) for orthogonal frequency division multiplexing where the channel sparsity is exploited, and a low‐complexity BP–MF–EP‐based JCED receiver is designed. Moreover, comparisons in message updating of state‐of‐the‐art message passing‐based JCED receivers are provided to illustrate the merits of the proposed one. In addition, message passing schedules are optimised to achieve better system performance. Simulation results verify the superiority of the proposed combined message passing receiver in terms of both bit‐error‐rate performance and convergence speed.
Zhengdao Yuan, Chuanzong Zhang, Zhongyong Wang, Qinghua Guo 0001, Jiangtao Xi
IET Commun.5
2018 Hatching eggs classification based on deep learning
Lei Geng, Tingyu Yan, Zhitao Xiao, Jiangtao Xi
Multim. Tools Appl.4
2018 Linear shrinkage estimation of covariance matrices using low-complexity cross-validation
Jun Tong, Rui Hu 0009, Jiangtao Xi, Zhitao Xiao, Qinghua Guo 0001, Yanguang Yu
Signal Process.3
2017 Recovering the absolute phase maps of three selected spatial-frequency fringes with multi-color channels
Yi Ding 0038, Jiangtao Xi, Yanguang Yu, Fuqin Deng
Neurocomputing2
2017 An Auxiliary Variable-Aided Hybrid Message Passing Approach to Joint Channel Estimation and Decoding for MIMO-OFDM
abstract
This letter deals with message passing receiver design for joint channel estimation and decoding in MIMO-OFDM with unknown noise variance. The conventional factor graph representation for the system involves observation factors, which are functions of a number of variables in the form of multiplication and summation. In this work, by introducing some auxiliary variables, we further break each of the observation factors into several factors, which enables the use of hybrid mean field (MF), belief propagation (BP), and expectation propagation (EP) message passing to tackle the observation factors. It turns out that our approach is much more efficient than the existing approaches, leading to remarkable performance improvement as shown by simulation results.
Zhengdao Yuan, Chuanzong Zhang, Zhongyong Wang, Qinghua Guo 0001, Jiangtao Xi
IEEE Signal Process. Lett.5
2016 Choosing the diagonal loading factor for linear signal estimation using cross validation
abstract
Linear signal estimation based on sample covariance matrices (SCMs) can perform poorly if the training data are limited and the SCMs are ill-conditioned. Diagonal loading (DL) may be used to improve robustness in the face of limited training data. This paper introduces two leave-one-out cross-validation schemes for choosing the DL factor. One scheme repeatedly splits the training data with respect to time, while the other repeatedly splits the out-of-training data with respect to space. We derive computationally efficient implementations and compare them with the oracle choice in terms of the mean squared error.
Jun Tong, Qinghua Guo 0001, Jiangtao Xi, Yanguang Yu, Peter J. Schreier
ICASSP3
2016 Encoding and communicating navigable speech soundfields
Xiguang Zheng, Christian H. Ritz, Jiangtao Xi
Multim. Tools Appl.3
2014 Condition Number-Constrained Matrix Approximation With Applications to Signal Estimation in Communication Systems
abstract
This letter introduces condition number-constrained approximation to matrices used for signal estimation and detection. Under a Frobenius norm criterion, the closed-form solution to the optimal approximation is derived, which can be found efficiently for arbitrary condition number constraints. The resulting approximation techniques are applied to the imperfectly estimated covariance and channel matrices used for estimating transmit signals in communication systems. With an appropriately chosen value of condition number, the robustness of the linear and decision-feedback estimators (DFE) against model mismatch can be significantly improved.
Jun Tong, Qinghua Guo 0001, Sheng Tong, Jiangtao Xi, Yanguang Yu
IEEE Signal Process. Lett.4
2013 A psychoacoustic-based analysis-by-synthesis scheme for jointly encoding multiple audio objects into independent mixtures
abstract
Perceptually accurate representation of audio objects obtained from multi-track audio signals is desired for applications such as interactive soundfield rendering and browsing. Presented in this work is a scalable psychoacoustic analysis-by-synthesis approach to extract the perceptually dominant time-frequency audio objects from a multi-track audio signal. The proposed compression framework exploits sparsity in the perceptual time-frequency domain where up to eight audio objects can be efficiently encoded using only two audio mixtures with side information representing the origin of the time-frequency instances in the mixture signals. The proposed approach, judged by both objective and subjective tests, results in superior audio quality compared to existing techniques when encoding more than 5 audio objects.
Xiguang Zheng, Christian H. Ritz, Jiangtao Xi
ICASSP3
2013 Multisource DOA estimation based on time-frequency sparsity and joint inter-sensor data ratio with single acoustic vector sensor
abstract
By exploring the time-frequency (TF) sparsity property of the speech, the inter-sensor data ratios (ISDRs) of single acoustic vector sensor (AVS) have been derived and investigated. Under noiseless condition, ISDRs have favorable properties, such as being independent of frequency, DOA related with single valuedness, and no constraints on near or far field conditions. With these observations, we further investigated the behavior of ISDRs under noisy conditions and proposed a so-called ISDR-DOA estimation algorithm, where high local SNR data extraction and bivariate kernel density estimation techniques have been adopted to cluster the ISDRs representing the DOA information. Compared with the traditional DOA estimation methods with a small microphone array, the proposed algorithm has the merits of smaller size, no spatial aliasing and less computational cost. Simulation studies show that the proposed method with a single AVS can estimate up to seven sources simultaneously with high accuracy when the SNR is larger than 15dB. In addition, the DOA estimation results based on recorded data further validates the proposed algorithm.
Yue Xian Zou, Christian H. Ritz, Muawiyath Shujau, Jiangtao Xi
ICASSP6
2013 Iterative Frequency Domain Equalization With Generalized Approximate Message Passing
abstract
An iterative frequency domain equalization approach for coded single-carrier block transmissions over frequency selective channels is developed by using the recently proposed generalized approximate message passing (GAMP) algorithm. Compared with the low-complexity iterative frequency domain linear minimum mean square error (FD-LMMSE) equalization, the proposed approach can achieve significant performance gain with slight complexity increase.
Qinghua Guo 0001, Defeng Huang, Sven Nordholm, Jiangtao Xi, Yanguang Yu
IEEE Signal Process. Lett.4
2013 Collaborative Blind Source Separation Using Location Informed Spatial Microphones
abstract
This letter presents a new Collaborative Blind Source Separation (CBSS) technique that uses a pair of location informed coincident microphone arrays to jointly separate simultaneous speech sources based on time-frequency source localization estimates from each microphone recording. While existing BSS approaches are based on localization estimates of sparse time-frequency components, the proposed approach can also recover non-sparse (overlapping) time-frequency components. The proposed method has been evaluated using up to three simultaneous speech sources under both anechoic and reverberant conditions. Results from objective and subjective measures of the perceptual quality of the separated speech show that the proposed approach significantly outperforms existing BSS approaches.
Xiguang Zheng, Christian H. Ritz, Jiangtao Xi
IEEE Signal Process. Lett.3
2013 Encoding Navigable Speech Sources: A Psychoacoustic-Based Analysis-by-Synthesis Approach
abstract
This paper presents a psychoacoustic-based analysis-by-synthesis approach for compressing navigable speech sources. The approach targets multi-party teleconferencing applications, where selective reproduction of individual speech sources is desired. Based on exploiting sparsity of speech in the perceptual time-frequency domain, multiple speech signals are encoded into one mono mixture signal, which can be further compressed using a standard speech codec. Using side information indicating the active speech source for each time frequency instant enables flexible decoding and reproduction. Objective results highlight the importance of considering perception when exploiting the sparse nature of speech in the time-frequency domain. Results show that this sparsity, as measured by the preserved energy level of perceptually important time-frequency components extracted from mixtures of speech signals, is similar in both anechoic and reverberant environments. The proposed approach is applied to a series of simulated and real reverberant speech recordings, where the resulting speech mixtures are compressed using a standard speech codec operating at 32 kbps. The perceptual quality, as judged both by objective and subjective evaluations, outperforms a simple sparsity approach that does not consider perception as well as the approach that encodes each source separately. While the perceptual quality of individual speech sources is maintained, subjective tests also confirm the approach maintains the perceptual quality of the spatialized speech scene.
Xiguang Zheng, Christian H. Ritz, Jiangtao Xi
IEEE Trans. Speech Audio Process.3
2012 Encoding navigable speech sources: An analysis by synthesis approach
abstract
This paper pressents an analysis-by-synthesis coding architecture for compressing navigable speech sources. The proposed coding scheme encodes multiple overlapped speech sources recorded, for example, during a multi-participant meeting or teleconference, into a mono or stereo mixture signal that can be compressed with an existing speech coder. The individual speech sources can be separated from the received compressed mixture, which allows the listener to determine the active sources and their spatial locations at the reproduction site. The approach was applied to the compression of a series of speech soundfields created from multiple clean speech sentences and real meeting recordings, where each sound-field contained four participants with up to three simultaneous speech sources. At a total bit rate of 48 kbps, the perceptual quality of each decoded speech source, as judged by subjective listening tests, was found to be significantly better than either a non-a-by-s approach or separate encoding of each source at the same overall total bit rate. Subjective listening tests also confirm that the quality of the spatialised speech scene is maintained as well.
Xiguang Zheng, Christian H. Ritz, Jiangtao Xi
ICASSP3
2008 Blind source separation for convolutive mixtures based on the joint diagonalization of power spectral density matrices
Tiemin Mei, Alfred Mertins, Fuliang Yin, Jiangtao Xi, Joe F. Chicharo
Signal Process.4
2006 Blind Source Separation Based on Time-Domain Optimization of a Frequency-Domain Independence Criterion
abstract
A new technique for the blind separation of convolutive mixtures is proposed in this paper. Inspired by the works of Amari, Sabala , and Rahbar, we firstly start from the application of Kullback-Leibler divergence in frequency domain, and then we integrate Kullback-Leibler divergence over the whole frequency range of interest to yield a new objective function which turns out to be time-domain variable dependent. In other words, the objective function is derived in frequency domain which can be optimized with respect to time domain variables. The proposed technique has the advantages of frequency domain approaches and is suitable for very long mixing channels, but does not suffer from the local permutation problem as the separation is achieved in time-domain
Tiemin Mei, Jiangtao Xi, Fuliang Yin, Alfred Mertins, Joe F. Chicharo
IEEE Trans. Speech Audio Process.2
2005 Joint Diagonalization of Power Spectral Density Matrices for Blind Source Separation of Convolutive Mixtures
Tiemin Mei, Jiangtao Xi, Fuliang Yin, Joe F. Chicharo
ISNN (2)2
2004 Blind source separation of nonstationary convolutively mixed signals in the subband domain
abstract
The paper proposes a new technique for blind source separation (BSS) in the subband domain using an extended lapped transform (ELT) decomposition for nonstationary, convolutively mixed signals. As identified by S. Araki et al. (see Proc. 4th Int. Symp. on Independent Component Analysis and Blind Signal Separation - ICA2003, p.499-504, 2003), the motivation for subband-based BSS is the drawback of frequency domain BSS when dealing with separating mixed speech signals over a few seconds resulting in few samples in individual frequency bins leading to poor separation performance. In the proposed approach, mixed signals are decomposed into subband components by an ELT and within each subband a time domain Newton BSS algorithm is employed based on the nonstationarity property of the input signals and the joint diagonalization of output correlation matrices with time varying second order statistics (SOS). This subband version is compared to a fullband version using the same BSS algorithm.
Iain Russell, Jiangtao Xi, Alfred Mertins, Joe F. Chicharo
ICASSP (5)2
2004 Cumulant-Based Blind Separation of Convolutive Mixtures
Tiemin Mei, Fuliang Yin, Jiangtao Xi, Joe F. Chicharo
ISNN (1)3
1998 Computing running discrete cosine/sine transforms based on the adaptive LMS algorithm
abstract
Least mean square (LMS)-based computation of block-based discrete orthogonal transforms has been extensively studied in literature. This paper establishes the relationship between the running discrete cosine transform II (DCT-II), discrete sine transform II (DST-II), and the adaptive LMS algorithm. From this analysis a new parallel structure for computing the running DCT-II and DST-II is proposed.
Jiangtao Xi, Joe F. Chicharo
IEEE Trans. Circuits Syst. Video Technol.1
1997 Blind separation and restoration of signals mixed in convolutive environment
abstract
This paper proposes new neural network approaches for separating and restoring signals mixed through FIR channels. Firstly, a set of maximal entropy based training rules are developed. Secondly, a new scheme for restoring the original signals is proposed for the 2/spl times/2 case. Computer simulation results for speech signals are presented to verify the proposed approaches.
Jiangtao Xi, James P. Reilly
ICASSP1
1997 A time-domain interpolation approach for DFT harmonic analysis
Jiangtao Xi, Joe F. Chicharo
Signal Process.1
1996 A new structure for the running discrete Hartley transform
Jiangtao Xi, Joe F. Chicharo
Signal Process.1
1995 A block gradient-based algorithm for adaptive IIR filtering
Jiangtao Xi, Joe F. Chicharo
Signal Process.1