Zhongfu Ye

dblp:38/839 · DBLP profile ↗
← Back
83ranked-venue papers
2as first author
29since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 61 · 2 first-author · 20 since 2021Artificial intelligence and machine learning · 20 · 10 since 2021Computer networks · 7 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2027 MLFF-DA: Multi-level feature fusion with dual attention for enhanced image captioning
Mohammad Alamgir Hossain, Md. Bipul Hossen, Md. Ibrahim Abdullah, Zhongfu Ye
Expert Syst. Appl.6
2026 Hierarchical Region-Context Attention for image captioning
Mohammad Alamgir Hossain, Zhongfu Ye, Md. Bipul Hossen, Md. Ibrahim Abdullah
Eng. Appl. Artif. Intell.2
2026 GTA: Geometric transform-attention network for enhanced spatial reasoning in image captioning
Mohammad Alamgir Hossain, Zhongfu Ye, Md. Bipul Hossen, Md. Ibrahim Abdullah
Neurocomputing2
2026 Geometric understanding in spatial contexts for image captioning
Mohammad Alamgir Hossain, Md. Bipul Hossen, Zhongfu Ye, Md. Ibrahim Abdullah
Signal Process. Image Commun.3
2025 Attribute guided fusion network for obtaining fine-grained image captions
Md. Bipul Hossen, Zhongfu Ye, Amr Abdussalam, Fazal-E. Wahab
Multim. Tools Appl.2
2025 Sparse channel estimation and passive beamforming with practical phase shift model for IRS-assisted OFDM systems
Shabih Ul Hassan, Zhongfu Ye, Jiancheng An 0001, Md. Bipul Hossen
Signal Process.2
2025 Machine learning-inspired hybrid precoding with low-resolution phase shifters for intelligent reflecting surface (IRS) massive MIMO systems with limited RF chains
Shabih Ul Hassan, Zhongfu Ye, Talha Mir, Usama Mir
Wirel. Networks2
2024 A Spatial Long-Term Iterative Mask Estimation Approach for Multi-Channel Speaker Diarization and Speech Recognition
abstract
Deep learning (DL)-based speaker diarization methods have proven powerful performance comparing to traditional clustering-based methods for multi-talker speech diarization and recognition in farfield scenes. However, most DL-based approaches cannot utilize the spatial information well due to the poor robustness to unknown array topology and acoustic scenario. In this paper, a spatial long-term iterative mask estimation (SLT-IME) method is proposed to improve the performance of speaker diarization in various real-world acoustic scenarios. First, the complex angular central gaussian mixture model (cACGMM) with diarization results as initial values is used to estimate the presence probability of each speaker at each time-frequency bin, namely speaker masks, in a long-term chunk. Then, the speaker masks are converted to speaker activities according to the threshold, which deliver the diarization information of which speaker is active and when. Finally, the estimated speaker activity can also serve as the initial input for the diarization system, resulting in improved ASR performance. Experimental results on the CHiME-7 three datasets (CHiME-6, DiPCo, Mixer 6) show proposed method can improve diarization and recognition systems performance simultaneously. It also plays a key role in the ensemble system that achieves the best performance in the main track of CHiME-7 DASR Challenge.
Yanhui Tu, Maokui He, Ruoyu Wang 0029, Shutong Niu, Lei Sun 0010, Zhongfu Ye, Jun Du 0002, Chin-Hui Lee 0001
ICASSP7
2024 Lightweight Transducer Based on Frame-Level Criterion
abstract
The transducer model trained based on sequence-level criterion requires a lot of memory due to the generation of the large probability matrix. We proposed a lightweight transducer model based on frame-level criterion, which uses the results of the CTC forced alignment algorithm to determine the label for each frame. Then the encoder output can be combined with the decoder output at the corresponding time, rather than adding each element output by the encoder to each element output by the decoder as in the transducer. This significantly reduces memory and computation requirements. To address the problem of imbalanced classification caused by excessive blanks in the label, we decouple the blank and non-blank probabilities and truncate the gradient of the blank classifier to the main network. Experiments on the AISHELL-1 demonstrate that this enables the lightweight transducer to achieve similar results to transducer. Additionally, we use richer information to predict the probability of blank, achieving superior results to transducer.
Genshun Wan, Mengzhi Wang, Tingzhi Mao, Zhongfu Ye
INTERSPEECH5
2024 Attribute-Driven Filtering: A new attributes predicting approach for fine-grained image captioning
Md. Bipul Hossen, Zhongfu Ye, Amr Abdussalam, Shabih Ul Hassan
Eng. Appl. Artif. Intell.2
2024 GVA: guided visual attention approach for automatic image caption generation
Md. Bipul Hossen, Zhongfu Ye, Amr Abdussalam, Md. Imran Hossain
Multim. Syst.2
2024 Compact deep neural networks for real-time speech enhancement on resource-limited devices
Fazal-E. Wahab, Zhongfu Ye, Nasir Saleem, Rizwan Ullah
Speech Commun.2
2023 Representation for action recognition with motion vector termed as: SDQIO
M. Shujah Islam, Khush Bakhat, Mansoor Iqbal, Rashid Khan, Zhongfu Ye, M. Mattah Islam
Expert Syst. Appl.5
2023 QISampling: An Effective Sampling Strategy for Event-Based Sign Language Recognition
abstract
Event cameras are innovative neuromorphic sensors that detect changes in brightness and generate a stream of events, thus providing high temporal resolution and dynamic range advantages. With its ability to perceive motion information, event data is well-suited for sign language recognition (SLR). Existing methods of event-based SLR rely on a uniform sampling strategy, which may result in redundant and indiscriminate information being captured when segments are randomly sampled. In this letter, we propose an effective sampling strategy called Quantity-inspired Sampling (QISampling) that takes advantage of the quantity feature of event distribution to sample key segments containing more discriminative and remarkable motion features from a sign. These segments are then converted into frame-like representations, which serve as input to subsequent Deep Neural Networks (DNNs). In addition, we introduce a synthetic event-based sign language dataset N-WLASL to the community. We apply our strategy to DNNs and conduct experiments on this synthetic dataset and another real-world dataset. The results demonstrate that our QISampling strategy could accelerate training and make event data provide superior performance for the SLR task.
Zhongfu Ye, Jin Wang 0023, Yueyi Zhang 0001
IEEE Signal Process. Lett.2
2023 Sparse Channel Estimation With Surface Clustering for IRS-Assisted OFDM Systems
abstract
Intelligent reflecting surface (IRS) is deemed as a potential technology for future communications due to its adaptive enhancement for the propagation environment. To achieve the passive beamforming gain of IRS, accurate channel state information (CSI) is essential but practically challenging since its massive passive reflecting elements have no transmitting/receiving capability. This paper presents a new channel estimation problem formulation for IRS-assisted orthogonal frequency division multiplexing (OFDM) systems, where the channel sparsity is exploited in the time-domain. Considering the surfaces are physically close to each other, we further utilize the common sparsity among the different sub-surfaces and automatically cluster them into several groups by introducing a Dirichlet process (DP)-based clustering model. Then, a DP-based variational Bayesian inference (VBI) framework is proposed to jointly estimate the channel and cluster the sub-surfaces, which is expected to significantly improve the channel estimation performance. Moreover, a novel decoupling trick is combined into the VBI framework to efficiently handle the coupling effect brought by the reflection coefficients, as well as facilitate the Bayesian inference. Simulation results verify the effectiveness of the proposed channel estimation scheme and show its significant performance improvement over various benchmark schemes.
Haoyang Dong, Lei Zhou 0012, Jisheng Dai, Zhongfu Ye
IEEE Trans. Commun.5
2023 NumCap: A Number-controlled Multi-caption Image Captioning Network
abstract
Image captioning is a promising task that attracted researchers in the last few years. Existing image captioning models are primarily trained to generate one caption per image. However, an image may contain rich contents, and one caption cannot express its full details. A better solution is to describe an image with multiple captions, with each caption focusing on a specific aspect of the image. In this regard, we introduce a new number-based image captioning model that describes an image with multiple sentences. An image is annotated with multiple ground-truth captions; thus, we assign an external number to each caption to distinguish its order. Given an image-number pair as input, we could achieve different captions for the same image under different numbers. First, a number is attached to the image features to form an image-number vector (INV). Then, this vector and the corresponding caption are embedded using the order-embedding approach. Afterward, the INV’s embedding is fed to a language model to generate the caption. To show the efficiency of the numbers incorporation strategy, we conduct extensive experiments using MS-COCO, Flickr30K, and Flickr8K datasets. The proposed model attains 24.1 in METEOR on MS-COCO. The achieved results demonstrate that our method is competitive with a range of state-of-the-art models and validate its ability to produce different descriptions under different given numbers.
Amr Abdussalam, Zhongfu Ye, Ammar Hawbani, Majjed Al-Qatf, Rashid Khan
ACM Trans. Multim. Comput. Commun. Appl.2
2022 Joint Optimization of the Module and Sign of the Spectral Real Part Based on CRN for Speech Denoising
Zilu Guo, Xu Xu 0003, Zhongfu Ye
INTERSPEECH3
2022 Joint Estimation of Direction-of-Arrival and Distance for Arrays with Directional Sensors based on Sparse Bayesian Learning
Feifei Xiong, Pengyu Wang 0010, Zhongfu Ye, Jinwei Feng
INTERSPEECH3
2022 Dual transform based joint learning single channel speech separation using generative joint dictionary learning
Md. Imran Hossain, Tarek Hasan Al Mahmud, Md. Bipul Hossen, Rashid Khan, Zhongfu Ye
Multim. Tools Appl.6
2022 Deep neural combinational model (DNCM): digital image descriptor for child's independent learning
Nuzhat Naqvi, M. Shujah Islam, Mansoor Iqbal, Zhongfu Ye
Multim. Tools Appl.6
2022 Single and two-person(s) pose estimation based on R-WAA
M. Shujah Islam, Khush Bakhat, Rashid Khan, M. Mattah Islam, Zhongfu Ye
Multim. Tools Appl.5
2022 Applied Human Action Recognition Network Based on SNSP Features
M. Shujah Islam, Khush Bakhat, Rashid Khan, Nuzhat Naqvi, M. Mattah Islam, Zhongfu Ye
Neural Process. Lett.6
2022 An off-grid wideband DOA estimation method with the variational Bayes expectation-maximization framework
Pengyu Wang 0010, Huichao Yang, Zhongfu Ye
Signal Process.3
2022 PFRNet: Dual-Branch Progressive Fusion Rectification Network for Monaural Speech Enhancement
abstract
In recent years, the transformer-based dual-branch magnitude and complex spectrum estimation framework achieves state-of-the-art performance for monaural speech enhancement. However, the insufficient utilization of the interactive information in the middle layers makes each branch lack the ability of compensation and rectification. To address this problem, this letter proposes a novel dual-branch progressive fusion rectification network (PFRNet) for monaural speech enhancement. PFRNet is an encoder-decoder-based dual-branch structure with interactive improved real & complex transformers. In PFRNet, the fusion rectification block is proposed to convert the implicit relationship of the two branches into a fusion feature by the frequency-domain mutual attention mechanism. The fusion feature provides a platform for the interaction in the middle layers. The interactive time-frequency improved real & complex transformer can make better use of the long-term dependencies in the time-frequency domain. Experimental results show that the proposed PFRNet outperforms most advanced dual-branch speech enhancement approaches and previous advanced systems in terms of speech quality and intelligibility.
Runxiang Yu, Zhongfu Ye
IEEE Signal Process. Lett.3
2021 Action recognition using interrelationships of 3D joints and frames based on angle sine relation and distance features using interrelationships
M. Shujah Islam, Khush Bakhat, Rashid Khan, Mansoor Iqbal, M. Mattah Islam, Zhongfu Ye
Appl. Intell.6
2021 A novel automatic image caption generation using bidirectional long-short term memory framework
Zhongfu Ye, Rashid Khan, Nuzhat Naqvi, M. Shujah Islam
Multim. Tools Appl.1
2021 Robust adaptive beamforming based on a method for steering vector estimation and interference covariance matrix reconstruction
Sicong Sun, Zhongfu Ye
Signal Process.2
2021 Accurate 3D hand pose estimation network utilizing joints information
Xiongquan Zhang, Shiliang Huang, Zhongfu Ye
Signal Process. Image Commun.3
2021 Robust Adaptive Beamforming via Covariance Matrix Reconstruction Under Colored Noise
abstract
Aimed at the performance degradation of the standard Capon beamformer (SCB) when the signal of interest (SOI) appearing in the training data under the colored noise, a novel interference-plus-noise covariance matrix (INCM) reconstruction method is proposed in this letter. The colored noise covariance matrix (CNCM) is estimated by the cross Capon power in the noise sector and the interference power is obtained via the quasi-orthogonality based on the conventional beamforming (CBF) between different sparse steering vectors (SVs) to project the sample covariance matrix for the INCM reconstruction. Then, the SV of the SOI is updated by the principal eigenvector of the reconstructed SOI covariance matrix. Simulation results show that the proposed method is robust against some mismatch errors under the colored noise.
Huichao Yang, Pengyu Wang 0010, Zhongfu Ye
IEEE Signal Process. Lett.3
2020 Speaker Adaptive Training for Speech Recognition Based on Attention-Over-Attention Mechanism
Genshun Wan, Qingran Wang, Jianqing Gao, Zhongfu Ye
INTERSPEECH5
2020 CAD: concatenated action descriptor for one and two person(s), using silhouette and silhouette's skeleton
abstract
This study introduces an action descriptor that has the ability to perform human action recognition efficiently for one and two person(s). The authors’ proposed descriptor computes information like motion, spatial–temporal, diversion with respect to the centroid, critical point and keypoint detection, whereas the existing approaches lack to address this information efficiently. Action descriptors are developed from signature‐based optical flow, signature‐based corner points and binary robust invariant scalable keypoints. These action descriptors are applied to silhouette and silhouette's skeleton frames. These aforementioned action descriptors lead to developing the concatenated action descriptor (CAD). In order to develop action descriptors, the reference video frame plays an important role. Weizmann (one person) and both clean and noise versions of SBU Kinect Interaction (two persons) datasets are used for the evaluation of their proposed descriptors. On the other hand, classifications are performed by using support vector machine. Experimental results demonstrate that CAD not only outperforms among the entire proposed descriptors, but also provides better performance as compared to state‐of‐the‐art approaches.
M. Shujah Islam, Mansoor Iqbal, Nuzhat Naqvi, Khush Bakhat, M. Mattah Islam, Zhongfu Ye
IET Image Process.7
2020 Image captions: global-local and joint signals attention model (GL-JSAM)
Nuzhat Naqvi, Zhongfu Ye
Multim. Tools Appl.2
2020 Robust sparse Bayesian learning for DOA estimation in impulsive noise environments
Xu Xu 0003, Zhongfu Ye, Jisheng Dai
Signal Process.3
2020 Robust adaptive beamforming via subspace for interference covariance matrix reconstruction
Xingyu Zhu 0002, Xu Xu 0003, Zhongfu Ye
Signal Process.3
2020 Online Speaker Adaptation Using Memory-Aware Networks for Speech Recognition
abstract
In our previous work, we introduced our attention-based speaker adaptation method, which has been proved to be an efficient online speaker adaptation method for real-time speech recognition. In this paper, we present a more complete framework of this method named memory-aware networks, which consists of the main network, the memory module, the attention module and the connection module. A gate mechanism and a multiple-connections strategy are presented to connect the memory with the main network in order to take full advantage of the memory. An auxiliary speaker classification task is provided to improve the accuracy of the attention module. The fixed-size ordinally forgetting encoding method is used together with average pooling to gather both short-term and long-term information. Furthermore, instead of only using traditional speaker embeddings such as i-vectors or d-vectors as the memory, we design a new form of memory called residual vectors, which can represent different pronunciation habits. Experiments on both the Switchboard and AISHELL-2 tasks show that our method can perform online speaker adaptation very well with no additional adaptation data and with only a relative 3% increase in decoding computation complexity. Under the cross-entropy criterion, our method achieves a relative word error rate reduction of 9.4% and 8.3% compared to that of the speaker-independent model on the Switchboard task and the AISHELL-2 task, respectively, and approximately 7.0% compared to that of the traditional d-vector-based speaker adaptation method.
Genshun Wan, Jun Du 0002, Zhongfu Ye
IEEE ACM Trans. Audio Speech Lang. Process.4
2019 Acoustic Model Ensembling Using Effective Data Augmentation for CHiME-5 Challenge
Li Chai 0002, Jun Du 0002, Diyuan Liu, Zhongfu Ye, Chin-Hui Lee 0001
INTERSPEECH5
2019 A deep learning approach for face recognition based on angularly discriminative features
Mansoor Iqbal, M. Shujah Islam Sameem, Nuzhat Naqvi, Zhongfu Ye
Pattern Recognit. Lett.5
2019 Sparse Bayesian learning for off-grid DOA estimation with Gaussian mixture priors when both circular and non-circular sources coexist
Xu Xu 0003, Zhongfu Ye, Tarek Hasan Al Mahmud, Jisheng Dai, Kashif Shabir
Signal Process.3
2019 Blind Joint 2-D DOA/Symbols Estimation for 3-D Millimeter Wave Massive MIMO Communication Systems
abstract
By using a large number of antenna (sensor) elements at the receivers, massive multi-input multi-output (MIMO) offers many benefits for 5G communication systems, such as a huge spectral efficiency gain, significant reduction of latency, and robustness to interference. However, to get these benefits of massive MIMO, accuracy of the channel state information obtained at the transmitter is required. This article proposes a approach for blind joint channel/symbols estimation in 3-D millimeter wave massive MIMO systems based on tensor factorization. More specifically, we suggest a direction-of-arrival (DOA)-based channel estimation method, which provides the best performance in terms of error bound for channel estimation. We show that the massive MIMO signals can be expressed as a third-order (3-D) tensor model, where the matrices of channel (2-D DOA) and symbols can be viewed as two independent factor matrices. Such a hybrid tensorial modeling enables a blind joint estimation of 2-D DOA/symbols. To learn the tensor model, we develop two least squares--based algorithms. The first one is delta bilinear alternating least squares (DBALS) algorithm that exploits the increment values between two iterations of the factor matrices to provide the initializations for such matrices. This avoids the slow convergence caused by random initializations for factor matrices found in the traditional least squares algorithms. The other one is Vandermonde constrained DBALS that takes into account the potential Vandermonde nature structure of the DOA matrix in the DBALS algorithm. This provides the estimation for the DOA matrix and gives a better uniqueness results for the use of tensor model. The performance of the proposed approach is illustrated by means of simulation results, and a comparison is made with the recent approaches. Besides a blind joint 2-D DOA/symbols estimation, our approach offers a better performance due to avoiding the random initializations and taking in the Vandermonde structure of DOA matrix.
Chung Buiquang, Zhongfu Ye
ACM Trans. Sens. Networks2
2018 Networks Effectively Utilizing 2D Spatial Information for Accurate 3D Hand Pose Estimation
abstract
In this work, we propose a new method for accurate 3D hand pose estimation from a single depth map using convolutional neural networks (CNN). Our method effectively makes use of 2D spatial information to improve the performance by two means. Firstly, we formulate 3D hand pose estimation as a two-task (2D joints detection and depth regression) problem so that we can directly utilize the ability of hourglass module on processing multi-scale information for estimating 2D joint coordinates. Secondly, 2D spatial information is used to help depth regression by introducing the spatial attention mechanism to our method. The experimental results demonstrate that our method achieves the state-of-the-art performance on ICVL hand posture dataset and a comparable performance with the state-of-the-arts on NYU dataset.
Baoen Liu, Shiliang Huang, Zhongfu Ye
ICIP3
2018 Supervised single-channel speech dereverberation and denoising using a two-stage model based sparse representation
Xu Xu 0003, Zhongfu Ye
Speech Commun.5
2018 A Novel Chinese Sign Language Recognition Method Based on Keyframe-Centered Clips
abstract
Isolated sign language recognition (SLR) is a long-standing research problem. The existing methods consider inclusively ambiguous data to represent a sign and ignore the fact that only scarce key information can represent the sign efficiently since most information are redundant. Furthermore, inclusion of redundant information may result in inefficiency and difficulty in modeling the long-term dependency for SLR. This letter delivers a novel sequence-to-sequence learning method based on keyframe centered clips (KCCs) for Chinese SLR. Different from conventional methods, only key information is considered to represent a sign significantly. The frames-to-word task is transformed into a KCCs-to-subwords task successfully, to allow for different attention in the input data. The empirical results of the proposed method outperform significantly the state-of-the-art SLR systems on our dataset containing 310 Chinese sign language words.
Shiliang Huang, Chensi Mao, Jinxu Tao, Zhongfu Ye
IEEE Signal Process. Lett.4
2017 Locally controlled as-rigid-as-possible deformation for 2D characters
abstract
Abstract Due to its practical use in animation and image editing, two‐dimensional shape deformation has received great attention during the past decades. The traditional paradigms, spreading local modifications over the whole shape, will cause a global deformation. For the simultaneous realization of the local control and preservation of the rigidity of the shape, a novel deformation method is proposed for point, skeleton, and cage handles. Our framework can be accomplished in terms of the minimization of the improved as‐rigid‐as‐possible energy with local and smooth penalties, named locally controlled as‐rigid‐as‐possible energy. The visually pleasing experiment results demonstrate the success of our method in two‐dimensional character deformation.
Zhongfu Ye
Comput. Animat. Virtual Worlds5
2017 Effective subset approach for SVMpath singularities
Jisheng Dai, Weichao Xu, Zhongfu Ye, Chunqi Chang
Pattern Recognit. Lett.3
2017 Data fusion over localized sensor networks for parallel waveform enhancement based on 3-D tensor representations
Renjie Tong, Zhongfu Ye
Signal Process.2
2017 Sequential DOA estimation method for multi-group coherent signals
Jinqiang Wei, Xu Xu 0003, Dawei Luo, Zhongfu Ye
Signal Process.4
2017 Supplementations to the Higher Order Subspace Algorithm for Suppression of Spatially Colored Noise
abstract
Recently, researchers have proposed to represent the observed multichannel speech data as a 3-D tensor and then directly reduce the noise level in the time domain. For example, a higher order subspace algorithm (HOSA) was proposed for the reduction of spatially white noise (i.e., the noise added to different sensors is mutually uncorrelated) and yielded an excellent performance. Nevertheless, for spatially colored noise (i.e., the noise added to different sensors is correlated or even perfectly coherent), HOSA suffers from serious performance degradation because the T-weighted noise covariance matrix is not strictly diagonal any more. In this letter, we present three efficient supplementations to HOSA to improve its performance on spatially colored noise. First, we directly estimate the weighted signal covariance matrices by speech subtraction because noise statistics are unknown and have to be estimated in the case. Second, we clear all the off-diagonal elements of the whitened T-weighted noise covariance matrix as they generally have negligible absolute values. Third, we process speech frames in an overlap-free manner to better meet the white noise assumption. In both simulations and real experiments, these supplementations can greatly improve the performance of HOSA on spatially colored noise while not affecting its performance on spatially white noise.
Renjie Tong, Zhongfu Ye
IEEE Signal Process. Lett.2
2016 Continuous sign language recognition using level building based on fast hidden Markov model
Jinxu Tao, Zhongfu Ye
Pattern Recognit. Lett.3
2016 Widely linear minimum dispersion beamforming for sub-Gaussian noncircular signals
Zhongfu Ye
Signal Process.4
2016 A rank-reduction based 2-D DOA estimation algorithm for three parallel uniform linear arrays
Xu Xu 0003, Yawar A. Sheikh, Zhongfu Ye
Signal Process.4
2016 Supervised single-channel speech enhancement using ratio mask with joint dictionary learning
Guangzhao Bao, Zhongfu Ye
Speech Commun.4
2016 Supervised Monaural Speech Enhancement Using Complementary Joint Sparse Representations
abstract
Sparse representation is one of the most well-known methods that are applied to monaural speech enhancement. In order to make full use of the relationships among speech, noise, and mixture in sparse representation for speech enhancement, this letter proposes a novel sparsity model that consists of a couple of joint sparse representations (JSRs). One JSR uses the mapping relationship between mixture and speech while the other uses that between mixture and noise. Both relationships are used to constrain the joint dictionary learning, which effectively solves the source confusion problem of traditional methods. Moreover, the latter JSR can be complementary to the former JSR, depending on the level of structure of the noise. Thus, we propose a Gini index based weighting parameter to take their complementary advantages. The experimental results show that the proposed method outperforms state-of-the-art methods using various objective measures.
You Luo, Guangzhao Bao, Yangfei Xu, Zhongfu Ye
IEEE Signal Process. Lett.4
2015 Robust widely linear beamformer based on a projection constraint
abstract
For noncircular signals, optimal widely linear (WL) minimum variance distortionless response (MVDR) beamformer has a powerful performance by exploiting the noncircularity of the received signals. Though, the noncircularity rate can be estimated by the steering vector (SV) of the signal of interest (SOI), the performance degrades as there exist errors in the SOI's SV. This paper introduces a new robust WL beamformer. In the proposed approach, the assumed extended steering vector (ESV) of the SOI is used to construct an interference-plus-noise subspace projection matrix, and the new ESV is estimated by maximizing the WL beamformer output power under a constraint that prevents the ESV from converging to the interference. The proposed algorithm only needs imprecise knowledge of the antenna array geometry and the SOI's angular sector. Simulations verify the effectiveness of the proposed algorithm.
Zhongfu Ye
ICASSP5
2015 Single-channel speech separation using sequential discriminative dictionary learning
Yangfei Xu, Guangzhao Bao, Xu Xu 0003, Zhongfu Ye
Signal Process.4
2015 A Higher Order Subspace Algorithm for Multichannel Speech Enhancement
abstract
In this letter, we propose a tensor factorization approach for multichannel speech enhancement, which is very successful even when the noise level is high. Specifically, we extend the well-known subspace approach to arbitrary orders and present the higher order subspace approach for multichannel speech enhancement. Unlike previous algorithms, the proposed approach constructs a third order tensor from the noisy data and then applies a tensor operation to reduce the noise. Through this it preserves the original data structure and makes full use of the spatial and temporal correlations in the multichannel data. The proposed approach adopts an iterative and step-wise procedure which usually converges in a few iterations. At each step a subspace filter sharing the same form with the conventional subspace approach is updated. Experiments show that it has achieved considerable performance on white Gaussian noise in terms of segmental signal-to-noise ratio improvement. Rapid convergence of the proposed approach is also reported.
Renjie Tong, Guangzhao Bao, Zhongfu Ye
IEEE Signal Process. Lett.3
2015 A Robust Time-Frequency Decomposition Model for Suppression of Mixed Gaussian-Impulse Noise in Audio Signals
abstract
In this paper, we propose a robust time-frequency decomposition (RTFD) model to restore audio signals degraded by sparse impulse noise mixed with small dense Gaussian noise. This kind of noise is very common especially in old-time recordings. The proposed RTFD model is based on the observation that these degraded audio signals mainly contain four parts, i.e., the quasi-periodic and voiced part, the aperiodic and transient part, the arbitrarily large impulse noise and the small dense Gaussian noise. Sparsity and local correlations of corresponding parts are exploited to solve the RTFD model. We also heuristically develop a discriminative orthogonal matching pursuit (DOMP) algorithm to more precisely estimate sparse representing vectors. Specifically, the DOMP algorithm divides the whole atom set into two subsets, i.e., the active subset and the passive subset. Atoms in two subsets are treated discriminatively since sparsity regularization terms are not equally weighted. Based on RTFD and DOMP, we have developed two algorithms, i.e., the fidelity-oriented algorithm and the articulation-oriented algorithm. The proposed algorithms achieve considerable performance on both synthetic and real noisy signals. Results show that the articulation-oriented algorithm using DOMP obviously outperforms other algorithms in heavier impulse noise situations.
Renjie Tong, Yingyue Zhou, Guangzhao Bao, Zhongfu Ye
IEEE ACM Trans. Audio Speech Lang. Process.5
2014 DOA estimation for noncircular signals in the presence of mutual coupling
Shenghong Cao, Dongyang Xu 0004, Xu Xu 0003, Zhongfu Ye
Signal Process.4
2014 DOA estimation for wideband signals based on sparse signal reconstruction using prolate spheroidal wave functions
Nan Hu 0001, Xu Xu 0003, Zhongfu Ye
Signal Process.3
2014 Robust widely linear beamforming based on spatial spectrum of noncircularity coefficient
Dongyang Xu 0004, Can Gong, Shenghong Cao, Xu Xu 0003, Zhongfu Ye
Signal Process.5
2014 Off-grid DOA estimation using array covariance matrix and block-sparse Bayesian learning
Zhongfu Ye, Xu Xu 0003, Nan Hu 0001
Signal Process.2
2014 Learning a Discriminative Dictionary for Single-Channel Speech Separation
abstract
This paper presents a novel dictionary learning (DL) method to improve the performance of sparsity based single-channel speech separation (SCSS). The conventional approaches regard the sub-dictionaries as independent units and learn sub-dictionaries separately in the short-time Fourier transform (STFT) domain using their corresponding training sets respectively. However, we take the relationship between the sub-dictionaries into account and optimize the sub-dictionaries jointly in the time domain. By satisfying a designed discrimination constraint, a structured dictionary, whose atoms have better correspondences to the speaker labels, is learned so that the sources can be recovered by the corresponding reconstruction after sparse coding. An algorithm, which consists of sparse coding stage and dictionary updating stage, is proposed to deal with this DL optimization problem. Two strategies, i.e., direct learning and adaptive learning, are presented to select the training sets which are used to learn the discriminative dictionary. Experimental results show that the proposed SCSS algorithms have superior performance compared with other tested approaches.
Guangzhao Bao, Yangfei Xu, Zhongfu Ye
IEEE ACM Trans. Audio Speech Lang. Process.3
2013 A restoration algorithm for images contaminated by mixed Gaussian plus random-valued impulse noise
Yingyue Zhou, Zhongfu Ye
J. Vis. Commun. Image Represent.2
2013 DOA estimation based on fourth-order cumulants in the presence of sensor gain-phase errors
Shenghong Cao, Zhongfu Ye, Nan Hu 0001, Xu Xu 0003
Signal Process.2
2013 A source enumeration method based on subspace orthogonality and bootstrap technique
Yunxia Zhang, Nan Hu 0001, Zhongfu Ye
Signal Process.3
2013 A Compressed Sensing Approach to Blind Separation of Speech Mixture Based on a Two-Layer Sparsity Model
abstract
This paper discusses underdetermined blind source separation (BSS) using a compressed sensing (CS) approach, which contains two stages. In the first stage we exploit a modified K-means method to estimate the unknown mixing matrix. The second stage is to separate the sources from the mixed signals using the estimated mixing matrix from the first stage. In the second stage a two-layer sparsity model is used. The two-layer sparsity model assumes that the low frequency components of speech signals are sparse on K-SVD dictionary and the high frequency components are sparse on discrete cosine transformation (DCT) dictionary. This model, taking advantage of two dictionaries, can produce effective separation performance even if the sources are not sparse in time-frequency (TF) domain.
Guangzhao Bao, Zhongfu Ye, Xu Xu 0003, Yingyue Zhou
IEEE Trans. Speech Audio Process.2
2012 A Khatri-Rao based method for DOA estimation in the presence of mutual coupling
abstract
A Khatri-Rao (KR) product based method for direction-of-arrival (DOA) estimation using uniform linear array (ULA) in the presence of mutual coupling is presented. Based on the fact that mutual coupling matrices of ULAs can be modeled as banded complex symmetric Toeplitz matrices, a cost function with the form of KR product is derived. An alternating minimization procedure is employed to estimate the DOAs of all the radiating signals as well as the mutual coupling coefficients and the power of signals. Simulation results that demonstrate the validity of the proposed method are included.
Shenghong Cao, Zhongfu Ye, Xu Xu 0003
ICASSP2
2012 The estimate for DOAs of signals using sparse recovery method
abstract
This paper presents a new direction-of-arrival (DOA) estimation method using the concept of sparse representation of an array cross-correlation vector (ACCV), in which DOA estimation is achieved by finding the sparse parameter vector according to an optimization criterion. Compared with other sparse recovery algorithms the proposed method achieves a higher resolution and has a less computational complexity. The performance of our method is demonstrated and analyzed through numerical simulations.
Dongyang Xu 0004, Nan Hu 0001, Zhongfu Ye, Ming Bao
ICASSP3
2012 Wideband DOA estimation from the sparse recovery perspective for the spatial-only modeling of array data
Nan Hu 0001, Dongyang Xu 0004, Xu Xu 0003, Zhongfu Ye
Signal Process.4
2012 A sparse recovery algorithm for DOA estimation using weighted subspace fitting
Nan Hu 0001, Zhongfu Ye, Dongyang Xu 0004, Shenghong Cao
Signal Process.2
2012 DOA Estimation Based on Sparse Signal Recovery Utilizing Weighted l1-Norm Penalty
abstract
In this letter, a new DOA estimation method based on sparse signal recovery is proposed. We utilize the Capon spectrum to design a weightedl1-norm penalty in order to further enforce the sparsity and approximate the originall0-norm. A theoretical guidance for choosing a proper regularization parameter is also presented according to the dual form of the original problem. Simulation results demonstrate the effectiveness and efficiency of the proposed method.
Xu Xu 0003, Xiaohan Wei, Zhongfu Ye
IEEE Signal Process. Lett.3
2012 Linear Precoder Optimization for MIMO Systems with Joint Power Constraints
abstract
This paper considers linear precoder optimization problems for multiple-input multiple-output (MIMO) systems. In addition to the conventionally used sum-power constraint, maximum eigenvalue constraint on the precoding matrix is also considered so as to account for power limitations imposed on each antenna by the linearity of its own power amplifier in practical implementations. A framework employing directional derivative is developed to obtain optimal precoder designs for different criteria including maximizing the information rate and minimizing the sum of mean-square error (MSE). It turns out that power allocations in such situations are piecewise linear in sum-power space. The piecewise linear property allows us to generate the entire path of solution through finding out a finite number of breakpoints. A Homotopy-type algorithm is then proposed to obtain the solution for an arbitrary sum-power constraint. The number of breakpoints to be determined in our exact piecewise linear solution is in fact only about two times of the number of transmit antennas, so that our method is super fast and outperforms existing approximate solutions in the literature in both effectiveness and efficiency. Simulated experiments are performed to verify our theoretical analysis.
Jisheng Dai, Chunqi Chang, Weichao Xu, Zhongfu Ye
IEEE Trans. Commun.4
2011 A novel wideband DOA estimator based on Khatri-Rao subspace approach
Dahang Feng, Ming Bao, Zhongfu Ye, Luyang Guan, Xiaodong Li 0002
Signal Process.3
2010 Autocalibration algorithm for mutual coupling of planar array
Zhongfu Ye
Signal Process.2
2010 An extended TOPS algorithm based on incoherent signal subspace method
Jisheng Dai, Zhongfu Ye
Signal Process.3
2010 An efficient DOA estimation method in multipath environment
Zhongfu Ye
Signal Process.2
2009 Optimal designs for linear MIMO transceivers using directional derivative
abstract
Optimal designs for minimising the combination of symbol estimation errors subject to lp-norm constraint are investigated. Instead of considering each constraint in a separate way, the authors develop a unifying framework to obtain the optimal solution by employing a directional derivative method. The sum or peak power constraint turns out to be the case of p=1 or p→∞. Simulation results demonstrate the effectiveness of the proposed algorithm. Moreover, based on directional derivative, the authors show that the minimisation of the determinant of the minimum mean-square error (MMSE) matrix and the maximisation of mutual information are equivalent criteria.
Jisheng Dai, Zhongfu Ye
IET Commun.2
2009 DOA estimation based on fourth-order cumulants with unknown mutual coupling
Zhongfu Ye
Signal Process.2
2009 An efficient greedy scheduler for zero-forcing dirty-paper coding
abstract
In this paper, an efficient greedy scheduler for zero-forcing dirty-paper coding (ZF-DPC), which can be incorporated in complex Householder QR factorization of the channel matrix, is proposed. The ratio of the complexity of the proposed scheduler to the complexity of the channel matrix factorization required by ZF-DPC is O(M-1), while such ratio for the original greedy scheduler is O(M), where M is the number of transmitters. Therefore, the new scheduler reduces the overhead of scheduling from being the bottleneck of ZF-DPC to being negligible.
Jisheng Dai, Chunqi Chang, Zhongfu Ye, Yeung Sam Hung
IEEE Trans. Commun.3
2008 A multipattern matching algorithm using sampling and bit index
abstract
Pattern matching is one of the basic problems in computer science. In this paper we propose a new multiple pattern matching algorithm. Unlike the well known Knuth-Morris-Pratt, Boyer-Moore, Karp-Rabin and their variants, our algorithm is derived from the ideas of sampling and bit index, sampling for efficiency and bit index for flexibility, as a result providing the simplest way to search for multiple patterns. Theoretical analysis and experimental results show that our algorithm is average-optimal with average complexity of O(n/m) for the search of patterns of length m in a text of length n. It provides a proper solution to such needs as matching long dispersed patterns and especially bit pattern matching (newly introduced in this paper) in data analysis of some private protocols' communication.
Zhongfu Ye
AICCSA2
2006 Salience Preserving Image Fusion with Dynamic Range Compression
abstract
Gradient conveys important salient features in images. Traditional fusion methods based on gradient generally treat gradients from multichannels as a multi-valued vector, and compute its global statistics under the assumption of identical distribution. However, different source channels may reflect different important salient features, and their gradients are basically non-identically distributed. This prevents existing methods from successful salience preservation. In this paper, we propose to fuse the gradients from multi-channels in the concept of saliency. We first measure the salience map of each channel's gradient, and then use their saliency to weight their contribution in computing the global statistics. Gradients with high saliency are properly highlighted in the target gradient, and thereby salient features in the sources are well preserved. Furthermore, we handle the dynamic range problem by applying range compression on the target gradient, and thereby halo effect is effectively reduced.
Qiong Yang, Xiaoou Tang, Zhongfu Ye
ICIP4
2006 The Study of Classification of Motor Imaginaries Based on Kurtosis of EEG
Xiaopei Wu, Zhongfu Ye
ICONIP (3)2
2006 Progressive cut
abstract
Recently, interactive image cutout technique becomes prevalent for image segmentation problem due to its easy-to-use nature. However, most existing stroke-based interactive object cutout system did not consider the user intention inherent in the user interaction process. Strokes in sequential steps are treated as a collection rather than a process, and only the color information of the additional stroke is used to update the color model in the graph cut framework. Accordingly, unexpected fluctuation effect may occur during the process of interactive object cutout. In fact, each step of user interaction reflects the user's evaluation of previous result and his/her intention. By analyzing the user's intention behind the interaction, we propose a progressive cut algorithm, which explicitly models the user's intention into a graph cut framework for the object cutout task. Three aspects of user intention are utilized: 1) the color of the stroke indicates the kind of change s/he expects, 2) the location of the stroke indicates the region of interest, 3) the relative position between the stroke and the previous result indicates the segmentation error. By incorporating such information into the cutout system, the new algorithm removes the unexpected fluctuation effect of existing stroke-based graph-cut methods, and thus provides the user a more controllable result with fewer strokes and faster visual feedback. Experiments and user study show the strength of progressive cut in accuracy, speed, controllability, and user experience.
Qiong Yang, Xiaoou Tang, Zhongfu Ye
ACM Multimedia5
2003 Blind separation of convolutive mixtures based on second order and third order statistics
abstract
This paper addresses the problem of blind separation of linear convolutive mixtures. We first reformulate the problem into a blind separation of linear instantaneous mixtures, and then a statistical approach is applied to solve the reformulated problem. From the statistics of the mixtures, two kinds of matrix pencils are constructed to estimate the mixing matrix. The original sources are then separated with the estimated mixing matrix. For the purpose of computational efficiency and robustness, in the matrix pencil, one matrix is constructed from the second order statistics, and the other is constructed from the third order statistics. The proposed novel methods do not require the exact knowledge of the channel order. Simulation results show that the methods are robust and have good performance.
Zhongfu Ye, Chunqi Chang, Francis H. Y. Chan
ICASSP (5)1