Pak-Chung Ching

dblp:69/6105 · also P. C. Ching · DBLP profile ↗
← Back
145ranked-venue papers
2as first author
8since 2021 · last 2024
0000-0002-4692-8707ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 97 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 46 · 1 since 2021Computer networks · 31 · 5 since 2021Systems, architecture and hardware · 3 · 1 first-authorTheory of computation · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 The Generalized Degrees-of-Freedom Region of the Two-User MIMO Broadcast Channel With Delayed CSIT
abstract
In this paper, we characterize the generalized degrees-of-freedom (GDoF) region of the two-user$(M,N_{1},N_{2})$multiple-input multiple-output (MIMO) broadcast channel with delayed channel state information at the transmitter (CSIT), where there are one transmitter with$M$antennas and two receivers with$N_{1}$and$N_{2}$antennas, respectively. Under delayed CSIT, different from the existing converse approaches in the multiple-input single-output (MISO) GDoF and MIMO degrees-of-freedom (DoF) models, we incorporate new components into traditional approaches for this MIMO GDoF converse. For the achievability, we generalize the existing MISO achievable scheme. Our result reveals how the channel strength and antenna configuration impact the GDoF region of the two-user MIMO broadcast channel with delayed CSIT. Furthermore, the extension of our converse to a GDoF outer region of the$K$-user MIMO broadcast channel with delayed CSIT is also provided.
Tong Zhang 0026, Shuai Wang 0004, Yinfei Xu, Rui Wang 0007, Pak-Chung Ching, H. Vincent Poor
IEEE Trans. Inf. Theory5
2023 Joint Beamforming Design for Integrated Passive Sensing and Communications in V2I Networks
abstract
This letter investigates joint beamforming and power optimization in a vehicle-to-infrastructure (V2I) system for integrated passive sensing and communications techniques. Specifically, the base station (BS) with multiple antennas delivers downlink data to multiple vehicles, respectively, and a sensing receiver simultaneously utilizes the downlink data signal to sense the motion of one vehicle. The sensing receiver is deployed separately. It collects the line-of-sight (LoS) signals from the BS and the scattered signals via the target vehicle, such that the distance and velocity of the target vehicle can be detected in a passive manner. As a result, the joint downlink beamforming and power optimization at the BS would be able to maximize the weighted summation of downlink data rates, subject to constraints on the signal-to-interference-plus-noise ratios (SINRs) of the two signals for passive sensing. In order to solve the above non-convex optimization problem, we first derive the optimal receiving beams for the data and sensing receivers given arbitrary transmission beam design, and then propose a low-complexity iterative algorithm to find a sub-optimal design of transmission beams based on the successive convex approximation (SCA) method. Simulations demonstrate the good performance, fast convergence and useful design insights.
Bojie Li, Tong Zhang 0026, Rui Wang 0007, Pak-Chung Ching
IEEE Signal Process. Lett.5
2022 Energy Efficiency for Proactive Eavesdropping in Cooperative Cognitive Radio Networks
abstract
This article investigates a distant proactive eavesdropping system in cooperative cognitive radio (CR) networks. Specifically, an amplify-and-forward (AF) full-duplex (FD) secondary transmitter assists to relay the received signal from suspicious users to legitimate monitor for wireless information surveillance. In return, the secondary transmitter is granted to share the spectrum belonging to the suspicious users for its own information transmission. To improve the eavesdropping, the transmitted secondary user’s (SU) signal can also be used as a jamming signal to moderate the data rate of the suspicious link. We consider two cases, i.e., nonnegligible processing delay (NNPD) and negligible processing delay (NPD) at the secondary transmitter. Our target is to maximize network energy efficiency (NEE) via jointly optimizing the AF relay matrix and precoding vector at the secondary transmitter, as well as the receiver combining vector at the monitor, subject to the maximum power constraint at the secondary transmitter and minimum data rate requirement of the SU. We also guarantee that the achievable data rate of the eavesdropping link should be no less than that of the suspicious link for efficient surveillance. Due to the nonconvexity of the formulated NEE maximization problem, we develop an efficient path-following algorithm and a robust alternating optimization (AO) method as solutions under perfect and imperfect channel state information (CSI) conditions, respectively. We also analyze the convergence and computational complexity of the proposed schemes. Numerical results are provided to validate the effectiveness of our proposed schemes.
Yao Ge 0001, Pak-Chung Ching
IEEE Internet Things J.2
2022 Enhancing Segment-Based Speech Emotion Recognition by Iterative Self-Learning
abstract
Despite the widespread utilization of deep neural networks (DNNs) for speech emotion recognition (SER), they are severely restricted due to the paucity of labeled data for training. Recently, segment-based approaches for SER have been evolving, which train backbone networks on shorter segments instead of whole utterances, and thus naturally augments training examples without additional resources. However, one core challenge remains for segment-based approaches: most emotional corpora do not provide ground-truth labels at the segment level. To supervisely train a segment-based emotion model on such datasets, the most common way assigns each segment the corresponding utterance’s emotion label. However, this practice typically introduces noisy (incorrect) labels as emotional information is not uniformly distributed across the whole utterance. On the other hand, DNNs have been shown to easily over-fit a dataset when being trained with noisy labels. To this end, this work proposes a simple and effective iterative self-learning (ISL) framework, which comprises a procedure to progressively correct segment-level labels in an iterative learning manner. The ISL method produces dynamically-generated and soft emotion labels, leading to significant performance improvements. Experiments on three well-known emotional corpora demonstrate noticeable gains using the proposed method.
Shuiyang Mao, Pak-Chung Ching, Tan Lee
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 OTFS Signaling for Uplink NOMA of Heterogeneous Mobility Users
abstract
We investigate a coded uplink non-orthogonal multiple access (NOMA) configuration in which groups of co-channel users are modulated in accordance with orthogonal time frequency space (OTFS). We take advantage of OTFS characteristics to achieve NOMA spectrum sharing in the delay-Doppler domain between stationary and mobile users. We develop an efficient iterative turbo receiver based on the principle of successive interference cancellation (SIC) to overcome the co-channel interference (CCI). We propose two turbo detector algorithms: orthogonal approximate message passing with linear minimum mean squared error (OAMP-LMMSE) and Gaussian approximate message passing with expectation propagation (GAMP-EP). The interactive OAMP-LMMSE detector and GAMP-EP detector are respectively assigned for the reception of the stationary and mobile users. We analyze the convergence performance of our proposed iterative SIC turbo receiver by utilizing a customized extrinsic information transfer (EXIT) chart and simplify the corresponding detector algorithms to further reduce receiver complexity. Our proposed iterative SIC turbo receiver demonstrates performance improvement over existing receivers and robustness against imperfect SIC process and channel state information uncertainty.
Yao Ge 0001, Qinwen Deng, Pak-Chung Ching, Zhi Ding 0001
IEEE Trans. Commun.3
2021 Caching Transient Content for IoT Sensing: Multi-Agent Soft Actor-Critic
abstract
Edge nodes (ENs) in Internet of Things commonly serve as gateways to cache sensing data while providing accessing services for data consumers. This paper considers multiple ENs that cache sensing data under the coordination of the cloud. Particularly, each EN can fetch content generated by sensors within its coverage, which can be uploaded to the cloud via fronthaul and then be delivered to other ENs beyond the communication range. However, sensing data are usually transient with time whereas frequent cache updates could lead to considerable energy consumption at sensors and fronthaul traffic loads. Therefore, we adopt Age of Information to evaluate data freshness and investigate intelligent caching policies to preserve data freshness while reducing cache update costs. Specifically, we model the cache update problem as a cooperative multi-agent Markov decision process with the goal of minimizing the long-term average weighted cost. To efficiently handle the exponentially large number of actions, we devise a novel reinforcement learning approach, which is a discrete multi-agent variant of soft actor-critic (SAC). Furthermore, we generalize the proposed approach into a decentralized control, where each EN can make decisions based on local observations only. Simulation results demonstrate the superior performance of the proposed SAC-based caching schemes.
Xiongwei Wu, Xiuhua Li 0001, Jun Li 0004, Pak-Chung Ching, Victor C. M. Leung, H. Vincent Poor
IEEE Trans. Commun.4
2021 Receiver Design for OTFS with a Fractionally Spaced Sampling Approach
abstract
The recent emergence of orthogonal time frequency space (OTFS) modulation as a novel PHY-layer mechanism is more suitable in high-mobility wireless communication scenarios than traditional orthogonal frequency division multiplexing (OFDM). Although multiple studies have analyzed OTFS performance using theoretical and ideal baseband pulseshapes, a challenging and open problem is the development of effective receivers for practical OTFS systems that must rely on non-ideal pulseshapes for transmission. This work focuses on the design of practical receivers for OTFS. We consider a fractionally spaced sampling (FSS) receiver in which the sampling rate is an integer multiple of the symbol rate. For rectangular pulses used in OTFS transmission, we derive a general channel input-output relationship of OTFS in delay-Doppler domain without the common reliance on impractical assumptions such as ideal bi-orthogonal pulses and on-the-grid delay/Doppler shifts. We propose two equalization algorithms: iterative combining message passing (ICMP) and turbo message passing (TMP) for symbol detection by exploiting delay-Doppler channel sparsity and the channel diversity gain via FSS. We analyze the convergence performance of TMP receiver and propose simplified message passing (MP) receivers to further reduce complexity. Our FSS receivers demonstrate stronger performance than traditional receivers and robustness to the imperfect channel state information knowledge.
Yao Ge 0001, Qinwen Deng, Pak-Chung Ching, Zhi Ding 0001
IEEE Trans. Wirel. Commun.3
2021 Multi-Agent Reinforcement Learning for Cooperative Coded Caching via Homotopy Optimization
abstract
Introducing cooperative coded caching into small cell networks is a promising approach to reducing traffic loads. By encoding content via maximum distance separable (MDS) codes, coded fragments can be collectively cached at small-cell base stations (SBSs) to enhance caching efficiency. However, content popularity is usually time-varying and unknown in practice. As a result, cached content is anticipated to be intelligently updated by taking into account limited caching storage and interactive impacts among SBSs. In response to these challenges, we propose a multi-agent deep reinforcement learning (DRL) framework to intelligently update cached content in dynamic environments. With the goal of minimizing long-term expected fronthaul traffic loads, we first model dynamic coded caching as a cooperative multi-agent Markov decision process. Owing to the use of MDS coding, the resulting decision-making falls into a class of constrained reinforcement learning problems with continuous decision variables. To deal with this difficulty, we custom-build a novel DRL algorithm by embedding homotopy optimization into a deep deterministic policy gradient formalism. Next, to empower the caching framework with an effective trade-off between complexity and performance, we propose centralized, and partially and fully decentralized caching controls by applying the derived DRL approach. Simulation results demonstrate the superior performance of the proposed multi-agent framework.
Xiongwei Wu, Jun Li 0004, Ming Xiao 0001, Pak-Chung Ching, H. Vincent Poor
IEEE Trans. Wirel. Commun.4
2020 Deep Reinforcement Learning for IoT Networks: Age of Information and Energy Cost Tradeoff
abstract
In most Internet of Things (IoT) networks, edge nodes are commonly used as to relays to cache sensing data generated by IoT sensors as well as provide communication services for data consumers. However, a critical issue of IoT sensing is that data are usually transient, which necessitates temporal updates of caching content items while frequent cache updates could lead to considerable energy cost and challenge the lifetime of IoT sensors. To address this issue, we adopt the Age of Information (AoI) to quantity data freshness and propose an online cache update scheme to obtain an effective tradeoff between the average AoI and energy cost. Specifically, we first develop a characterization of transmission energy consumption at IoT sensors by incorporating a successful transmission condition. Then, we model cache updating as a Markov decision process to minimize average weighted cost with judicious definitions of state, action, and reward. Since user preference towards content items is usually unknown and often temporally evolving, we therefore develop a deep reinforcement learning (DRL) algorithm to enable intelligent cache updates. Through trial-and-error explorations, an effective caching policy can be learned without requiring exact knowledge of content popularity. Simulation results demonstrate the superiority of the proposed framework.
Xiongwei Wu, Xiuhua Li 0001, Jun Li 0004, Pak-Chung Ching, H. Vincent Poor
GLOBECOM4
2020 Latency-Minimized Design of secure transmissions in UAV-Aided Communications
abstract
Unmanned aerial vehicles (UAVs) can be utilized as aerial base stations to provide communication service for remote mobile users due to their high mobility and flexible deployment. However, the line-of-sight (LoS) wireless links are vulnerable to be intercepted by the eavesdropper (Eve), which presents a major challenge for UAV-aided communications. In this paper, we propose a latency-minimized transmission scheme for satisfying legitimate users' (LUs') content requests securely against Eve. By leveraging physical-layer security (PLS) techniques, we formulate a transmission latency minimization problem by jointly optimizing the UAV trajectory and user association. The resulting problem is a mixed-integer nonlinear program (MINLP), which is known to be NP hard. Furthermore, the dimension of optimization variables is indeterminate, which again makes our problem very challenging. To efficiently address this, we utilize bisection to search for the minimum transmission delay and introduce a variational penalty method to address the associated subproblem via an inexact block coordinate descent approach. Moreover, we present a characterization for the optimal solution. Simulation results are provided to demonstrate the superior performance of the proposed design.
Xiongwei Wu, Qiang Li 0017, Yawei Lu, H. Vincent Poor, Victor C. M. Leung, Pak-Chung Ching
ICASSP6
2020 Advancing Multiple Instance Learning with Attention Modeling for Categorical Speech Emotion Recognition
abstract
Categorical speech emotion recognition is typically performed as a sequence-to-label problem, i.e., to determine the discrete emotion label of the input utterance as a whole. One of the main challenges in practice is that most of the existing emotion corpora do not give ground truth labels for each segment; instead, we only have labels for whole utterances. To extract segment-level emotional information from such weakly labeled emotion corpora, we propose using multiple instance learning (MIL) to learn segment embeddings in a weakly supervised manner. Also, for a sufficiently long utterance, not all of the segments contain relevant emotional information. In this regard, three attention-based neural network models are then applied to the learned segment embeddings to attend the most salient part of a speech utterance. Experiments on the CASIA corpus and the IEMOCAP database show better or highly competitive results than other state-of-the-art approaches.
Shuiyang Mao, Pak-Chung Ching, C.-C. Jay Kuo, Tan Lee
INTERSPEECH2
2020 Emotion Profile Refinery for Speech Emotion Classification
abstract
Human emotions are inherently ambiguous and impure.When designing systems to anticipate human emotions based on speech, the lack of emotional purity must be considered.However, most of the current methods for speech emotion classification rest on the consensus, e. g., one single hard label for an utterance.This labeling principle imposes challenges for system performance considering emotional impurity.In this paper, we recommend the use of emotional profiles (EPs), which provides a time series of segment-level soft labels to capture the subtle blends of emotional cues present across a specific speech utterance.We further propose the emotion profile refinery (EPR), an iterative procedure to update EPs.The EPR method produces soft, dynamically-generated, multiple probabilistic class labels during successive stages of refinement, which results in significant improvements in the model accuracy.Experiments on three well-known emotion corpora show noticeable gain using the proposed method.
Shuiyang Mao, Pak-Chung Ching, Tan Lee
INTERSPEECH2
2020 EigenEmo: Spectral Utterance Representation Using Dynamic Mode Decomposition for Speech Emotion Classification
abstract
Human emotional speech is, by its very nature, a variant signal.This results in dynamics intrinsic to automatic emotion classification based on speech.In this work, we explore a spectral decomposition method stemming from fluid-dynamics, known as Dynamic Mode Decomposition (DMD), to computationally represent and analyze the global utterance-level dynamics of emotional speech.Specifically, segment-level emotion-specific representations are first learned through an Emotion Distillation process.This forms a multi-dimensional signal of emotion flow for each utterance, called Emotion Profiles (EPs).The DMD algorithm is then applied to the resultant EPs to capture the eigenfrequencies, and hence the fundamental transition dynamics of the emotion flow.Evaluation experiments using the proposed approach, which we call EigenEmo, show promising results.Moreover, due to the positive combination of their complementary properties, concatenating the utterance representations generated by EigenEmo with simple EPs averaging yields noticeable gains.
Shuiyang Mao, Pak-Chung Ching, Tan Lee
INTERSPEECH2
2020 Joint Long-Term Cache Updating and Short-Term Content Delivery in Cloud-Based Small Cell Networks
abstract
Explosive growth of mobile data demand may impose a heavy traffic burden on fronthaul links of cloud-based small cell networks (C-SCNs), which deteriorates users' quality of service (QoS) and requires substantial power consumption. This paper proposes an efficient maximum distance separable (MDS) coded caching framework for a cache-enabled C-SCNs, aiming at reducing long-term power consumption while satisfying users' QoS requirements in short-term transmissions. To achieve this goal, the cache resource in small-cell base stations (SBSs) needs to be reasonably updated by taking into account users' content preferences, SBS collaboration, and characteristics of wireless links. Specifically, without assuming any prior knowledge of content popularity, we formulate a mixed timescale problem to jointly optimize cache updating, multicast beamformers in fronthaul and edge links, and SBS clustering. Nevertheless, this problem is anti-causal because an optimal cache updating policy depends on future content requests and channel state information. To handle it, by properly leveraging historical observations, we propose a two-stage updating scheme by using Frobenius-Norm penalty and inexact block coordinate descent method. Furthermore, we derive a learning-based design, which can obtain effective trade-off between accuracy and computational complexity. Simulation results demonstrate the effectiveness of the proposed two-stage framework.
Xiongwei Wu, Qiang Li 0017, Xiuhua Li 0001, Victor C. M. Leung, Pak-Chung Ching
IEEE Trans. Commun.5
2019 Revisiting Hidden Markov Models for Speech Emotion Recognition
abstract
Hidden Markov models (HMMs) have a long tradition in automatic speech recognition (ASR) due to their capability of capturing temporal dynamic characteristics of speech. For emotion recognition from speech, three HMM based architectures are investigated and compared throughout the current paper, namely, the Gaussian mixture model based HMMs (GMM-HMMs), the subspace based Gaussian mixture model based HMMs (SGMM-HMMs) and the hybrid deep neural network HMMs (DNN-HMMs). Extensive emotion recognition experiments are carried out on these three architectures on the CASIA corpus, the Emo-DB corpus and the IEMOCAP database, respectively, and results are compared with those of state-of-the-art approaches. These HMM based architectures prove capable of constituting an effective model for speech emotion recognition. Also, the modeling accuracy is further enhanced by incorporating various advanced techniques from the ASR area. In particular, among all of the architectures, the SGMM-HMMs achieve the best performance in most of the experiments.
Shuiyang Mao, Dehua Tao, Guangyan Zhang, Pak-Chung Ching, Tan Lee
ICASSP4
2019 Latency Driven Fronthaul Bandwidth Allocation and Cooperative Beamforming for Cache-enabled Cloud-based Small Cell Networks
abstract
This paper considers content delivery of the cache-enabled small cell networks (C-SCNs), where users with the same request form a multicast group and are served by a cluster of small-cell base stations (SBSs) under the coordination of the central processor. The performance of such a coordination is severely limited by the fronthaul link, which may be saturated and degrade quality of service (QoS). To improve user QoS, we propose a latency driven scheme by jointly optimizing fronthaul bandwidth allocation, multicast beamforming, and BS clustering. Accordingly, with min-max fairness among multicast groups, a latency minimization problem is formulated under the constraints of fronthaul bandwidth and transmission power. The resultant problem is a mixed-integer nonlinear program, which is NP-hard. To address such a complex problem, a quadratic penalty-based algorithm is proposed by using a reformulation of binary constraint. Meanwhile, we present the necessary condition for an optimal solution, which shows that fronthaul bandwidth allocation is inherently adaptive to cached contents and patterns of BS cooperation. Finally, simulation results demonstrate that the proposed scheme can effectively reduce latency under different caching strategies.
Xiongwei Wu, Xiuhua Li 0001, Qiang Li 0017, Victor C. M. Leung, Pak-Chung Ching
ICASSP5
2019 Secure MIMO Interference Channel with Confidential Messages and Delayed CSIT
abstract
Secure degrees-of-freedom (SDoF) for multiple-input multiple-output (MIMO) interference channel with confidential messages and delayed channel state information at transmitter (CSIT) remains unclear, even for single-input single-output (SISO) case. In this paper, we propose an achievable SDoF for MIMO interference channel with confidential messages and delayed CSIT by designing a multi-phase achievable scheme. For the proposed scheme, we first establish a multi-phase transmission procedure with undetermined phase durations. In the next, the security and decoding conditions on feasible phase durations are derived. Finally, an achievable SDoF maximization problem with respect to phase durations is solved under security and decoding conditions. Due to effective coordination of interference, the proposed achievable SDoF can be 20% greater than the SDoF of MIMO wiretap channel with delayed CSIT, which removes one transmitter from the MIMO interference channel.
Tong Zhang 0026, Pak-Chung Ching
ICASSP2
2019 Joint Long-Term Cache Allocation and Short-Term Content Delivery in Green Cloud Small Cell Networks
abstract
Recent years have witnessed an exponential growth of mobile data traffic, which may lead to a serious traffic burn on the wireless networks and considerable power consumption. Network densification and edge caching are effective approaches to addressing these challenges. In this study, we investigate joint long-term cache allocation and short-term content delivery in cloud small cell networks (C-SCNs), where multiple small-cell BSs (SBSs) are connected to the central processor via fronthaul and can store popular contents so as to reduce the duplicated transmissions in networks. Accordingly, a long-term power minimization problem is formulated by jointly optimizing multicast beamforming, BS clustering, and cache allocation under quality of service (QoS) and storage constraints. The resultant mixed timescale design problem is an anticausal problem because the optimal cache allocation depends on the future file requests. To handle it, a two-stage optimization scheme is proposed by utilizing historical knowledge of users' requests and channel state information. Specifically, the online content delivery design is tackled with a penalty-based approach, and the periodic cache updating is optimized with a distributed alternating method. Simulation results indicate that the proposed scheme significantly outperforms conventional schemes and performs extremely close to a genie-aided lower bound in the low caching region.
Xiongwei Wu, Qiang Li 0017, Xiuhua Li 0001, Victor C. M. Leung, Pak-Chung Ching
ICC5
2019 Deep Learning of Segment-Level Feature Representation with Multiple Instance Learning for Utterance-Level Speech Emotion Recognition
Shuiyang Mao, Pak-Chung Ching, Tan Lee
INTERSPEECH2
2019 Spectrum Sharing and Energy Cooperation in Wireless Powered Cognitive Radio Networks
abstract
In this work, we consider a cognitive radio system, where the primary user (PU) owns the spectrum but has scarce energy while the secondary users (SUs) have adequate energy but lack of spectrum. Thus, a spectrum sharing and energy cooperation scheme is proposed, where the SUs help transfer energy to the PU in the first phase, and in return, the PU allows the SUs to access the spectrum in the second phase. This is particularly beneficial when the PU is energy-limited wireless sensor node or internet of things and the transmitters of SUs are base stations or access points with sufficient energy supply. Without loss of generality, we aim to maximize the minimum data rate among all SUs by jointly optimizing the time- splitting factor between the two phases, the transmission power at the primary transmitter (PT) and the precoding vectors for secondary transmitters (STs) under the minimum data rate requirement of the PU, and the power constraint at each ST. We also guarantee the energy causality constraint at the PT, i.e., the total consumed energy should be no larger than the total available energy. To solve this non-convex problem, we propose an efficient iterative algorithm by applying the successive convex approximation (SCA) and further show that the proposed algorithm is guaranteed to converge. Simulation results are finally presented to show the effectiveness of our proposed scheme.
Yao Ge 0001, Pak-Chung Ching
VTC Fall2
2019 Joint Fronthaul Multicast and Cooperative Beamforming for Cache-Enabled Cloud-Based Small Cell Networks: An MDS Codes-Aided Approach
abstract
The performance of cloud-based small cell networks (C-SCNs) relies highly on a capacity-limited fronthaul, which degrade quality of service when it is saturated. Coded caching is a promising approach to addressing these challenges, as it provides abundant opportunities for fronthaul multicast and cooperative transmissions. This paper investigates cache-enabled C-SCNs, in which small-cell base stations (SBSs) are connected to the central processor via fronthaul, and can prefetch popular contents by applying maximum distance separable (MDS) codes. To fully capture the benefits of fronthaul multicast and cooperative transmissions, an MDS codes-aided transmission scheme is first proposed. We formulate the problem to minimize the content delivery latency by jointly optimizing fronthaul bandwidth allocation, SBS clustering, and beamforming. To efficiently solve the resulting nonlinear integer programming problem, we propose a penalty-based design by leveraging variational reformulations of binary constraints. To improve the solution of the penalty-based design, a greedy SBS clustering design is also developed. Furthermore, closed-form characterization of the optimal solution is obtained, through which the benefits of MDS codes can be quantified. The simulation results are given to demonstrate the significant benefits of the proposed MDS codes-aided transmission scheme.
Xiongwei Wu, Qiang Li 0017, Victor C. M. Leung, Pak-Chung Ching
IEEE Trans. Wirel. Commun.4
2018 Content Delivery Design for Cache-Aided Cloud Radio Access Network to Achieve Low Latency
abstract
In this paper, we examine the content delivery design for a cache-aided cloud radio access network (CA-CRAN), where users are served by multiple base stations (BSs) that are connected to cloud processor via fronthaul link. We propose a unified framework for cooperative delivery, which aims to minimize the total latency in the network. With fairness among users and physical-layer transmission, beamformers and content assignment are jointly optimized to fully exploit the benefits of caching. To address the resulting mixed binary nonconvex problem, a successive convex approximation (SCA)-based algorithm is derived with low complexity. Through simulations, the proposed design reduces latency significantly compared with existing work.
Xiongwei Wu, Pak-Chung Ching
ICASSP2
2018 Three-User Mimo Broadcast Channel with Delayed Csit: A Higher Achievable DoF
abstract
Degrees of freedom (DoF) of the three-user multiple-input multiple-output (MIMO) broadcast channel (BC) with delayed CSIT was derived for most antenna configurations except for the case of , where transmitter has M antennas and each receiver has N antennas. In this paper, for that problem, we propose an effective scheme for acquiring a higher achievable DoF than the value via existing methods. In the initial transmission phase, we transmit more data symbols than the amount that the receivers can instantaneously decode. Then, we generate auxiliary symbols for decoding the data symbols. Specifically, our scheme introduces an integrated design for the generation of auxiliary symbols. As a result, a higher achievable DoF, i.e., [12MN/(7M+2N)], can be achieved for specific antenna configurations, where .2N <; M <; 2.5N.
Tong Zhang 0026, Xiongwei Wu, Yinfei Xu, Yao Ge 0001, Pak-Chung Ching
ICASSP5
2018 An Effective Discriminative Learning Approach for Emotion-Specific Features Using Deep Neural Networks
Shuiyang Mao, Pak-Chung Ching
ICONIP (4)2
2017 Joint transmit beamforming optimization and uplink/downlink user selection in a full-duplex multi-user MIMO system
abstract
This paper considers practical deployment issues of a multi-user MIMO system with full-duplex (FD) base station and half-duplex (HD) user equipment. The aim is to select a set of uplink (UL) and downlink (DL) users at any instant that will provide a satisfactory performance in system resource allocation. Furthermore, it is also necessary to deal with the interference created by the UL users to the DL users, which limits communication quality. In this work, we consider implementing a joint processing beamforming algorithm that can provide effective UL/DL selection and achieve system utility maximization. Our results show that with 20 dB self-interference cancellation, FD system significantly outperforms HD system under proportional fairness utility.
Man-Wai Un, Wing-Kin Ma, Pak-Chung Ching
ICASSP3
2017 Interference alignment on MIMO X channel with synergistic CSIT
abstract
The achievable degree of freedom (DoF) boosting has been demonstrated on a single-input single-output (SISO) X channel by using outdated and instantaneous channel state information at transmitter (CSIT) synergistically, in contrast to that of using completely outdated CSIT. However, the means by which the DoF gain can be obtained in a multiple-input multiple-output (MIMO) system remains unclear. This paper proposes an interference alignment scheme with synergistic CSIT for MIMO X channel. We show that the achievable DoF is greater than the optimal DoF obtained with outdated CSIT, and that achievable DoF equals to the optimal value with outdated CSIT and transmitter cooperation.
Tong Zhang 0026, Pak-Chung Ching
ICASSP2
2017 Acoustic Assessment of Disordered Voice with Continuous Speech Based on Utterance-Level ASR Posterior Features
Yuanyuan Liu 0002, Tan Lee, Pak-Chung Ching, Thomas K. T. Law, Kathy Yuet-Sheung Lee
INTERSPEECH3
2016 Cost-Efficient Optimization of Base Station Densities for Multitier Heterogeneous Cellular Networks
abstract
Heterogeneous cellular networks have emerged as the primary solution for overwhelming traffic demands. This study addresses the architecture optimization of multitier open-access downlink heterogeneous cellular networks. A distinguishing feature of this work is that the transmission requirements of mobile users and the profit requirements of operators are all considered. Spatial throughput is maximized under practical constrains on the network deployment cost and traffic loads of individual tiers. By using stochastic geometry, the problems of optimizing the densities of base stations (BSs) in different tiers are shown to be linear-fractional programming that can be further transformed into linear programming by employing Charnes-Cooper transformation. The linear-fractional programming can then be solved efficiently. For a two-tier network, the optimal BS densities are derived in closed form by using the graphical method. Our results indicate that enlarging the densities of BSs may not always improve network throughput and that the optimal densities of each tier exist to satisfy the communication demands of mobile users and the profit requirements of operators. Numerical results demonstrate the significant throughput gain from the optimal deployment of multitier heterogeneous cellular networks.
Ran Cai, Wei Zhang 0001, Pak-Chung Ching
IEEE Trans. Wirel. Commun.3
2016 Time-Reversal Space-Time Codes in Asynchronous Two-Way Relay Networks
abstract
We consider an asynchronous two-way relay network in which a few distributed relays assist in the communication between two single-antenna terminals through analog network coding. The asynchronous transmission between relays and terminals causes symbol misalignments and results in diversity loss in space-time codes (STCs). We propose a family of zero-padded time-reversal STC that can achieve full diversity with low-complexity maximum likelihood (ML) decoding, given a bound for the path delay difference. With ML decoding, the proposed code is decomposed into several independent parts, which greatly facilitate the decoding process. For two-relay scenarios, three code designs based on Alamouti code, quasi-orthogonal space-time block code, and multigroup decodable STC are provided for each relay with one, two, and four antennas, respectively. For three-relay scenarios, one code design based on quasi-orthogonal space-time block code is provided for each relay with one antenna. Proof of full diversity is established, and the decoding complexity order is analyzed for all four designs. Simulations confirm the full diversity gain of all designs. The bit-error-rate performance in asynchronous scenarios is almost similar to that in synchronous scenarios.
Wei Zhang 0001, Pak-Chung Ching
IEEE Trans. Wirel. Commun.3
2015 Optimal base station densities for cost-efficient multi-tier heterogeneous cellular networks
abstract
Heterogeneous cellular networks have emerged to be the primary solution for overwhelming traffic demands. This study addresses the architecture optimization of multi-tier open access downlink heterogeneous cellular networks to achieve the maximum network spatial throughput under practical constraints on the network deployment cost and traffic loads of individual tiers. In particular, with the use of a stochastic geometry-based network model, the problem of optimizing the densities of base stations (BSs) in different tiers is shown to be a linear-fractional programming that can be further transformed into a linear programming by employing Charnes-Cooper transformation. Such programming can, in turn, be solved efficiently. Furthermore, for the special case of a two-tier network, the optimal BS densities are derived in closed form using the graphical method. Simulation results demonstrate significant throughput gain from the optimal deployment of multi-tier heterogeneous cellular networks.
Ran Cai, Wei Zhang 0001, Pak-Chung Ching
ICASSP3
2015 Time-reversal space-time codes in asynchronous two-way double-antenna relay networks
abstract
We consider an asynchronous two-way relay network, in which two double-antenna relays assist in the communication between two single-antenna terminals through analog network coding. The asynchronous transmission between relays and terminals causes symbol misalignments and results in diversity loss in space-time block code (STBC). We propose a zero-padded time-reversal quasi-orthogonal STBC that can achieve full diversity with low-complexity maximum likelihood (ML) decoding, given a bounded delay. With ML decoding, the proposed code is decomposed into several independent parts, which leads to single complex symbol decoding. Proof of full diversity is established, and the decoding complexity order is analyzed for the proposed design. Simulations confirm the full diversity gain. The bit error rate performance in asynchronous scenarios is almost the same as that in synchronous scenarios.
Wei Zhang 0001, Pak-Chung Ching
ICASSP3
2015 Modeling temporal dependency for robust estimation of LP model parameters in speech enhancement
Chun Hoy Wong, Tan Lee, Yu Ting Yeung, Pak-Chung Ching
INTERSPEECH4
2015 A Power Allocation Strategy for Multiple Poisson Spectrum-Sharing Networks
abstract
This paper develops a power allocation strategy for multiple networks of Poisson-distributed single-antenna nodes that share the available spectrum in a spectrum underlay scenario. This strategy aims to maximize the overall throughput obtained by sharing the spectrum while limiting the degradation of the successful transmission probability of each network. In its original form, this joint power allocation problem is difficult to solve. However, we demonstrate that the problem can be transformed into a convex optimization formulation, which can be efficiently solved. Furthermore, we obtain a quasi-closed-form solution that has a water-filling interpretation by analyzing the optimality conditions. Numerical results indicate that, when a spectrum-sharing scheme employs the proposed optimal strategy of power allocation, the throughput substantially improves over that obtained by exclusively allocating the spectrum to the primary network. Moreover, when the number of spectrum-sharing networks increases, the enhancement is significant, being up to the limit imposed by the maximum allowable degradation in the performance of each network.
Ran Cai, Jian-Kang Zhang 0002, Timothy N. Davidson, Wei Zhang 0001, Kon Max Wong, Pak-Chung Ching
IEEE Trans. Wirel. Commun.6
2014 Performance of partial zero-forcing beamforming in large random spectrum sharing networks
abstract
Mutual interference is the main bottleneck on the throughput of large random spectrum sharing networks. This work examines the extent to which the performance of such networks can be improved by employing multiple transmitting antennas, without degrading the average performance of individual users. By extending partial zero-forcing beamforming to spectrum sharing networks, the aggregate interference towards primary receivers is reduced, and the desired signals at both primary and secondary receivers are boosted. Considering randomly distributed users and spatially independent Rayleigh fading channels, this work provides upper and lower bounds on the maximum permissible density of secondary transmitters with respect to the numbers of primary and secondary transmitting antennas. The simulation results show that substantial increase in the density of secondary transmitters can be obtained while meeting the outage requirements of the spectrum sharing users.
Ran Cai, Wei Zhang 0001, Pak-Chung Ching, Timothy N. Davidson, Jian-Kang Zhang 0002
ICASSP3
2014 Achieving full-diversity and fast maximum likelihood decoding in asynchronous analog network coding
abstract
This study designs space-time codes in analog network coding for asynchronous two-way relay networks, where asynchronous transmissions can cause diversity loss. We propose a novel code, called zero-padded interleave reversal Alamouti code (ZP-IR AC) to achieve full diversity with fast maximum likelihood decoding. Specifically, to combat symbol misalignment caused by asynchronous transmissions, the two terminals insert zero padding when transmitting to relays. Thereafter, the second relay performs an interleave reversal procedure to retain full diversity at each terminal. A salient feature of ZP-IR AC is that it can be decoupled into several independent parts that facilitate fast maximum likelihood decoding. Simulations of ZP-IR AC show full diversity gain. The bit error rate performance of ZP-IR AC is comparable to that of the synchronized Alamouti code and outperforms those of some recent schemes.
Wei Zhang 0001, Soung Chang Liew, Pak-Chung Ching
ICASSP4
2014 Opportunistic user scheduling in MIMO cognitive radio networks
abstract
This paper studies multiuser diversity of uplink MIMO cognitive radio network and proposes a two-stage opportunistic user scheduling scheme. In the first stage, a cognitive beamforming design is proposed to ensure the interference caused by secondary signals is canceled or minimized on the spatial dimensions occupied by primary MIMO system. Then, some secondary users that cause minimal interference leakage at primary system are pre-selected as candidate users. In the second stage, some candidate users that produce maximum sum secondary rate are further selected for uplink scheduling. The proposed scheme enables the secondary link to take advantage of multiuser diversity while ensuring that the interference on primary link is within a certain threshold. Analytical results show that the sum rate of secondary uplink scales as Nslog logK for K secondary users and Nsantennas on secondary receiver for very large K.
Lu Yang 0001, Wei Zhang 0001, Nengheng Zheng, Pak-Chung Ching
ICASSP4
2014 Bounded Delay-Tolerant Space-Time Codes for Distributed Antenna Systems
abstract
Distributed space-time codes (STCs) can achieve cooperative diversity in distributed antenna systems. But the path delay difference may cause diversity loss. In this paper, a family of asynchronous STC is proposed to achieve full diversity, given that the path delay difference is within a tolerance bound. The proposed code structure allows the overall decoding process to be decomposed into several independent sub-processes, which can be tackled by low-complexity maximum-likelihood decoders. Two code designs based on the Alamouti code and the Golden code are given for a system with two distributed antennas. Moreover, two designs based on orthogonal STC and fast group decodable STC are introduced for a system with four distributed antennas, with each transmitter having two antennas. Full diversity gain is proved for all code designs and their associated decoding complexities are analyzed. Simulations of the proposed codes confirm our theoretical results on full diversity gain.
Wei Zhang 0001, Soung Chang Liew, Pak-Chung Ching
IEEE Trans. Wirel. Commun.4
2013 Power control for multiple spectrum-sharing networks under random geometric topologies
abstract
This paper develops a power control strategy for multiple spectrum-sharing networks of single antenna nodes in a spectrum underlay scenario. A distinguishing feature of the proposed strategy is that it requires only knowledge of the spatial distribution of the nodes, rather than instantaneous channel state information. The strategy seeks to maximize a weighted sum of the throughput of each network while guaranteeing specified successful transmission probabilities. In its native form, this joint power allocation problem is difficult to solve. However, we show that the problem can be transformed into a convex optimization formulation that can be efficiently solved using general purpose tools. Furthermore, we analyze the optimality conditions and obtain a quasi-closed form solution reminiscent of waterfilling. Numerical results demonstrate that spectrum sharing employing the proposed optimal power yields a substantial throughput gain over allocating the spectrum to a single network.
Ran Cai, Jian-Kang Zhang 0002, Timothy N. Davidson, Kon Max Wong, Pak-Chung Ching
ICASSP5
2013 Full-diversity distributed space-time codes with an efficient ML decoder for asynchronous cooperative communications
abstract
We consider an asynchronous cooperative communication system where two distributed transmitters are communicating with one destination. In this paper, a bounded delay-tolerant time interleave reversal Alamouti code (BDT-TIR AC) is proposed. BDT-TIR AC achieves full diversity, given that the path delay difference is within a tolerance limit. Furthermore, we design an efficient maximum-likelihood (ML) decoder. By employing a divide-and-conquer strategy, the original decoding problem is decoupled into several sub-problems in small size. In fact, these sub-problems contain either Alamouti code structure or only four symbols. By parallel decoding on the sub-problems, the proposed ML decoder exhibits low complexity advantage in large scale problems. Simulations of BDT-TIR AC confirm the full diversity gain. The bit error rate performance of BDT-TIR AC in the considered asynchronous scenarios is comparable with that of synchronized Alamouti code, and outperforms the latest delay-tolerant space-time code.
Wei Zhang 0001, Pak-Chung Ching
ICASSP3
2013 Hybrid spectrum sharing with imperfect sensing in fading channels
abstract
This paper considers the hybrid spectrum sharing paradigm where a cognitive radio system first performs spectrum sensing to identify primary users' (PU) status (idle/busy) and then adapts its transmit power according to sensing outcomes. To maximize ergodic throughput in fading environments, joint sensing and power allocation has to be considered. However, existing studies determine the optimal sensing time based on instantaneous channel state information (CSI) at each time slot, which imposes a stringent requirement in practice. In this paper, we obtained a statistical CSI-based optimal sensing time by exploring the ergodic rates of both overlay and underlay access in Rayleigh fading environments. Simulation results validate the derived analytic expressions, showing that significantly higher maximum throughput can be achieved by hybrid access compared with conventional overlay access.
Yalin Zhang 0003, Pak-Chung Ching, Qinyu Zhang 0001
ICASSP2
2013 Achieving full cooperative and frequency diversity in bit-interleaved coded two-way relay networks
abstract
This paper investigates channel coded transmission schemes for two-way relay networks (TWRNs) where two terminal nodes exchange information through a set of amplify-and-forward (AF) relays over frequency-selective fading channels. Specifically, we assume that the two terminal nodes employ bit-interleaved coded modulation (BICM) and orthogonal frequency division multiplexing (OFDM) for coded data transmission, while the relays employ distributed space-time coding (DSTC) for forwarding the received signals. Our main contribution lies in analyzing the achievable diversity order of such channel coded AF-TWRNs. We show that the two terminal nodes, by using a maximum-likelihood (ML) BICM-OFDM decoder, can harvest the full cooperative and frequency diversity. Simulation results are presented to verify our analytical results.
Tsung-Hui Chang, Jianhua Ge, Wing-Kin Ma, Pak-Chung Ching
WCNC5
2013 Cooperative Secure Beamforming for AF Relay Networks With Multiple Eavesdroppers
abstract
This letter studies cooperative secure beamforming for amplify-and-forward (AF) relay networks in the presence of multiple eavesdroppers. Under both total and individual relay power constraints, we propose two schemes, namely secrecy rate maximization (SRM) beamforming and null-space beamforming. In the first scheme, our design problem is based on SRM. Using a suboptimal, but convex, technique-semidefinite relaxation (SDR), we show that this problem can be handled by performing a one-dimensional search which involves solving a sequence of semidefinite programs (SDPs). To reduce the complexity, in the second scheme, we instead maximize the information rate at the destination while completely eliminating the information leakage to all eavesdroppers. We prove that this problem can be exactly solved by SDR with one SDP only. Simulation results demonstrate the performance gains of the two proposed designs.
Qiang Li 0017, Wing-Kin Ma, Jianhua Ge, Pak-Chung Ching
IEEE Signal Process. Lett.5
2013 Cooperative Spectrum Sharing Based on Two-Path Successive Relaying
abstract
A spectrum sharing scheme is proposed for overlaid wireless networks based on the decode-and-forward two-path successive relaying technique. In this scheme, two secondary transmitters are used to transmit secondary information alternately to their respective receivers while relaying the primary signal at the same time. The transmission of the primary system is continuous, and the secondary system can opportunistically help the primary transmission in exchange for the spectrum sharing. Superposition coding is used at the secondary transmitters, where the primary signal and secondary signal are linearly combined. Successive interference decoding and cancelation is performed by secondary users to extract their desired signals. For the primary system, joint decoding is performed by treating the secondary signal as noise at the receiver side. The optimal power allocation is determined to maximize the success probability of the secondary system without violating the outage performance of the primary system. Numerical and simulation results show that the proposed scheme can conservatively protect the primary performance while realizing the secondary transmission requirement.
Chao Zhai 0001, Wei Zhang 0001, Pak-Chung Ching
IEEE Trans. Commun.3
2012 Spectrum leasing based on bandwidth efficient relaying in cognitive radio networks
abstract
A spectrum leasing scheme is proposed for overlaid wireless networks based on the spectral efficient relaying. The potential secondary transmitters compete to assist the primary transmission using the one-path alternate or two-path successive relaying in exchange for the opportunity of spectrum access. The cognitive relaying mode is dynamically chosen based on the availability of secondary receiver for the cooperation. The secondary user that can support the primary transmission and has the highest secondary rate is selected as the best to join the spectrum leasing. The rate and outage performance of both systems are studied. Numerical results show that the primary performance is greatly improved compared with direct transmission, and the secondary requirement is also well satisfied.
Chao Zhai 0001, Wei Zhang 0001, Pak-Chung Ching
GLOBECOM3
2012 Spectrum sharing between random geometric networks
abstract
In this paper a spectrum sharing scheme is proposed to maximize the successful transmission probability of a single-hop cognitive network coexisting with a random primary network. Both cognitive and primary networks exhibit randomness in topologies and endure imperfect wireless channel conditions. For a given primary outage probability bound, the maximum secondary transmit power is determined and then the maximum transmission capacity of the cognitive user is derived. Numerical results show that the proposed spectrum sharing scheme indeed boosts the transmission capacity of the cognitive network dramatically whilst having little performance loss of the primary network.
Ran Cai, Wei Zhang 0001, Pak-Chung Ching
ICASSP3
2012 Spectrum sharing based on two-path successive relaying
abstract
A spectrum sharing scheme is proposed for overlaid wireless networks based on the two-path successive relaying technique. In this scheme, two secondary transmitters are used to alternately transmit secondary information to their respective receivers whilst relaying the primary signal at the same time. By using superposition coding at the secondary transmitters and successive decoding at the receivers, a diversity gain of 2 is obtained for the primary system without losing the bandwidth efficiency. Outage probabilities of the primary system and secondary system are analyzed. The optimal power allocation factor of the secondary system is determined to minimize the outage probability of the secondary system while maintaining the performance of the primary system. Numerical results show that the proposed scheme can indeed enhance the outage performance of the primary and secondary systems simultaneously.
Chao Zhai 0001, Wei Zhang 0001, Pak-Chung Ching
ICASSP3
2012 Noncoherent Bit-Interleaved Coded OSTBC-OFDM with Maximum Spatial-Frequency Diversity
abstract
The combination of bit-interleaved coded modulation (BICM), orthogonal space-time block coding (OSTBC) and orthogonal frequency division multiplexing (OFDM) has been shown recently to be able to achieve maximum spatial-frequency diversity in frequency selective multi-path fading channels, provided that perfect channel state information (CSI) is available to the receiver. In view of the fact that perfect CSI can be obtained only if a sufficient amount of resource is allocated for training or pilot data, this paper investigates pilot-efficient noncoherent decoding methods for the BICM-OSTBC-OFDM system. In particular, we propose a noncoherent maximum-likelihood (ML) decoder that uses only one OSTBC-OFDM block. This block-wise decoder is suitable for relatively fast fading channels whose coherence time may be as short as one OSTBC-OFDM block. Our focus is mainly on noncoherent diversity analysis. We study a class of carefully designed transmission schemes, called perfect channel identifiability (PCI) achieving schemes, and show that they can exhibit good diversity performance. Specifically, we present a worst-case diversity analysis framework to show that PCI-achieving schemes can achieve the maximum noncoherent spatial-frequency diversity of BICM-OSTBC-OFDM. The developments are further extended to a distributed BICM-OSTBC-OFDM scenario in cooperative relay networks. Simulation results are presented to confirm our theoretical claims and show that the proposed noncoherent schemes can exhibit near-coherent performance.
Tsung-Hui Chang, Wing-Kin Ma, Jianhua Ge, Chong-Yung Chi, Pak-Chung Ching
IEEE Trans. Wirel. Commun.6
2011 Single-symbol decodable distributed STBC for two-path successive relay networks
abstract
A fast-decodable distributed space-time block code (STBC) is proposed for a two-path successive relay network that can achieve both full diversity and full transmission rate. The proposed STBC employs a precoder to rotate the constellation of transmitted symbols at source. After decode-and-forward by relays, the rotated information symbols are decoded at destination. It is shown that single complex symbol decoding can be obtained at the destination. Simulation results demonstrate that the full diversity is offered by the newly proposed distributed STBC for the two-path successive relay network.
Long Shi 0001, Wei Zhang 0001, Pak-Chung Ching
ICASSP3
2011 Robust Speaker Recognition Using Denoised Vocal Source and Vocal Tract Features
abstract
To alleviate the problem of severe degradation of speaker recognition performance under noisy environments because of inadequate and inaccurate speaker-discriminative information, a method of robust feature estimation that can capture both vocal source- and vocal tract-related characteristics from noisy speech utterances is proposed. Spectral subtraction, a simple yet useful speech enhancement technique, is employed to remove the noise-specific components prior to the feature extraction process. It has been shown through analytical derivation, as well as by simulation results, that the proposed feature estimation method leads to robust recognition performance, especially at low signal-to-noise ratios. In the context of Gaussian mixture model-based speaker recognition with the presence of additive white Gaussian noise, the new approach produces consistent reduction of both identification error rate and equal error rate at signal-to-noise ratios ranging from 0 to 15 dB.
Ning Wang 0052, Pak-Chung Ching, Nengheng Zheng, Tan Lee
IEEE Trans. Speech Audio Process.2
2011 An Effective Distributed Space-Time Code for Two-Path Successive Relay Network
abstract
Two-path successive relaying has recently been proposed as a mechanism for recovering the multiplexing loss of conventional relay networks; however, the full diversity cannot be guaranteed with this method. In this paper, distributed space-time block coding (STBC) is proposed for a two-path successive decode-and-forward relay network that can achieve both full rate and full diversity when perfect decoding is enabled at the relays. For the case of imperfect decoding, selection relaying is proposed with the distributed STBC to recover the full diversity. For the practical scenario in which the inter-relay channel is subject to fading, the two-path successive relaying transmission may not always be successful; for such cases an adaptive relaying scheme is proposed that can select a proper transmission mode between two-phase relaying transmission and two-path successive relaying transmission based on the channel quality of the relay network. Analytical results show that the proposed distributed STBC can offer both full diversity and a large symbol rate gain.
Wei Zhang 0001, Wing-Kin Ma, Pak-Chung Ching, H. Vincent Poor
IEEE Trans. Commun.4
2010 Distributed Space-Time Coding for Two-Path Successive Relaying
abstract
A distributed space-time block code (STBC) based on a decode-and-forward (DF) protocol is proposed for two-path successive relaying. By making the two relay nodes transmit and listen in turn, the proposed distributed STBC can achieve not only full rate but also full diversity when the relay nodes can correctly decode the information symbols from the source node and when the inter-relay channels are sufficiently strong. We also analyze the situation when there are decoding errors at the relay nodes. We notice in this case full diversity will not be achieved. We then propose the application of selection relaying to the proposed distributed STBC to ensure that full rate and full diversity can still be obtained with decoding errors at the relay nodes.
Wei Zhang 0001, Wing-Kin Ma, Pak-Chung Ching
GLOBECOM4
2010 Cross-lingual speaker adaptation via Gaussian component mapping
Houwei Cao, Tan Lee, Pak-Chung Ching
INTERSPEECH3
2010 Exploitation of phase information for speaker recognition
Ning Wang 0052, Pak-Chung Ching, Tan Lee
INTERSPEECH2
2009 Full diversity under multiple carrier frequency offsets of a family of space-frequency codes
abstract
A cooperative system may have both timing errors and multiple carrier frequency offsets (CFOs). To combat timing errors, space-frequency (SF) coded OFDM systems have been recently proposed for cooperative communications to achieve both full cooperative and full multipath diversities without time synchronization requirement. In this paper, we study the effect of multiple CFOs from relay nodes on a family of rotation based SF codes. We find that they can still achieve full diversities under the condition that the absolute values of normalized CFOs are less than 0.5. We further show that this full diversity property still holds for a complexity-reduced two-stage zero forcing (ZF) aided maximum likelihood (ML) decoding method, for which a ZF method is used to equalize multiple CFOs before ML decoding.
Xiang-Gen Xia 0001, Wing-Kin Ma, Pak-Chung Ching
ICASSP4
2009 Effects of language mixing for automatic recognition of Cantonese-English code-mixing utterances
Houwei Cao, Pak-Chung Ching, Tan Lee
INTERSPEECH2
2009 Exploration of vocal excitation modulation features for speaker recognition
Ning Wang 0052, Pak-Chung Ching, Tan Lee
INTERSPEECH2
2008 A simple ICI mitigation method for a space-frequency coded cooperative communication system with multiple CFOS
abstract
A cooperative system may have both timing errors and carrier frequency offsets (CFOs) from relay nodes. To combat timing errors, space-frequency (SF) coded OFDM system has been recently proposed for cooperative communications to achieve both full cooperative and multipath diversities without the time synchronization requirement among relay nodes. In this paper, we consider the signal detection problem in an SF coded OFDM system for cooperative communications where multiple CFOs may occur from relay nodes. By exploiting the structure of the SF codes that are rotational based in this paper, we propose a novel inter-carrier interference (ICI) mitigation method called the multiple fast Fourier transform (M-FFT) method. Simulation results illustrate that for the same symbol error rate (SER) level, the M-FFT based detection methods require less computations than other classic approaches.
Pak-Chung Ching, Xiang-Gen Xia 0001
ICASSP2
2008 Signal Detection in a Cooperative Communication System with Multiple CFOs by Exploiting the Properties of Space-Frequency Codes
abstract
In cooperative communications, carrier frequency offsets (CFOs) between pairs of nodes may be different, making CFOs compensation difficult if not impossible. These multiple CFOs may drastically degrade the performance of a space-frequency (SF) coded cooperative system. In this paper, we consider the signal detection problem in SF coded cooperative communication systems with multiple CFOs, where the SF codes are rotational based and can achieve both the full cooperative and the full multipath diversities. By exploiting the structure of the SF codes, we propose two signal detection methods to deal with the multiple CFOs problem in the SF coded OFDM systems. They are the minimum mean-squared error filtering (MMSE-F) method, the two-stage simple frequency shift Q taps (FS-Q-T) method, both of which offer different tradeoffs between performance and computational complexity. Simulation results indicate that our proposed detection methods work well as long as the carrier frequency offsets between nodes are not unreasonably large.
Xiang-Gen Xia 0001, Pak-Chung Ching
ICC3
2008 Language modeling for speech recognition of spoken Cantonese
Yu Ting Yeung, Houwei Cao, Nengheng Zheng, Tan Lee, Pak-Chung Ching
INTERSPEECH5
2008 Distributed Space-Frequency Coding for Cooperative Diversity in Broadband Wireless Ad Hoc Networks
abstract
A distributed space-frequency (SF) coded cooperative technique is proposed for broadband wireless ad-hoc networks, where both the channels from the source node to relay nodes and from the relay nodes to the destination node are characterized by frequency-selective fading. Using an SF coding at the source node and a circular shift at each relay node, we propose a scheme that can achieve both cooperative and multi-path diversity. The pairwise error probability is analyzed and the result demonstrates that a diversity gain of min(ML1, ML2) can be achieved, where M is the number of the relay nodes, and L1and L2are the number of taps of the multipath fading channels from the source node to the relay nodes and from the relay nodes to the destination node, respectively. Furthermore, the proposed distributed SF coding can achieve full-rate transmission for any number of relay nodes. In particular, it does not need IFFT/FFT processing or decoding at each relay node, thus providing a low-complexity design of the relay nodes in broadband wireless ad-hoc networks.
Wei Zhang 0001, Yabo Li, Xiang-Gen Xia 0001, Pak-Chung Ching, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.4
2007 Distributed Space-Frequency Coding in Broadband Ad Hoc Networks
abstract
A distributed space-frequency (SF) coded cooperative technique is proposed for broadband wireless ad-hoc networks. By employing an SF coding at the source node and a circular shift at each relay node, the proposed cooperative diversity technique can exploit the maximum relay diversity and multipath diversity. Pairwise error probability analysis shows that the achievable diversity gain is min (ML1, ML2), where M is the number of the relay nodes, and L1and L2are the number of taps of the multipath fading channels from the source node to the relay nodes and from the relay nodes to the destination node, respectively.
Wei Zhang 0001, Yabo Li, Xiang-Gen Xia 0001, Pak-Chung Ching, Khaled Ben Letaief
ICASSP (3)4
2007 Blind Separation of Moving Speech Sources using Short-Time LOD Based ICA Method
abstract
This paper describes the application of an effective short-time ICA method for blind speech separation of a moving-speaker system. For the situation where time-varying mixture exists, adaptive ICA techniques encounter difficulties due to the collapse of sources independence assumption under short-time analysis. In this paper, we propose a method based on the short-time local optima distribution (LOD) of feasible separation region to alleviate such problems. Based on the characteristics of these distributions, information is obtained for avoiding local traps and approaching the desired global optimum of the de-mixing matrix. Simulation tests of the proposed method show its effectiveness for blind separation of moving speech sources.
Pak-Chung Ching
ICASSP (3)2
2007 Model-based speech separation with single-microphone input
Siu Wa Lee, Frank K. Soong, Pak-Chung Ching
INTERSPEECH3
2007 Signal Detection for Space-Frequency Coded Cooperative Communication System with Multiple Carrier Frequency Offsets
abstract
To combat synchronization errors in cooperative communication, paper (Li, 2006) applied the orthogonal frequency division multiplexing (OFDM) technique to achieve full cooperative diversity and full multipath diversity for asynchronous cooperative communications, where a kind of high rate space-frequency (SF) codes was used at relay nodes. However, OFDM is very sensitive to carrier frequency offset (CFO). For conventional communication system, only one CFO exists and its effect is easy to be compensated. But for cooperative communication, the CFO between each relay and destination may be very different. These multiple CFO make frequency compensation difficult if not impossible. To solve this problem, we propose a detection scheme for SF codes which can work much better than the conventional SISO zero-forcing (ZF) and minimum mean square error (MMSE) methods. Moreover, parallel interference cancellation (PIC) can be used to further improve the system performance. Simulation results show that the proposed scheme combined with PIC is capable to achieve almost the same performance as cases without CFO.
Xiang-Gen Xia 0001, Pak-Chung Ching
WCNC3
2007 Integration of Complementary Acoustic Features for Speaker Recognition
abstract
This letter describes a speaker verification system that uses complementary acoustic features derived from the vocal source excitation and the vocal tract system. A new feature set, named the wavelet octave coefficients of residues (WOCOR), is proposed to capture the spectro-temporal source excitation characteristics embedded in the linear predictive residual signal. WOCOR is used to supplement the conventional vocal tract-related features, in this case, the Mel-frequency cepstral coefficients (MFCC), for speaker verification. A novel confidence measure-based score fusion technique is applied to integrate WOCOR and MFCC. Speaker verification experiments are carried out on the NIST 2001 database. The equal error rate (EER) attained with the proposed method is 7.67%, in comparison to 9.30% of the conventional MFCC-based system
Nengheng Zheng, Tan Lee, Pak-Chung Ching
IEEE Signal Process. Lett.3
2007 High-Rate Full-Diversity Space-Time-Frequency Codes for Broadband MIMO Block-Fading Channels
abstract
A systematic design of high-rate full-diversity space-time-frequency (STF) codes is proposed for multiple-input multiple-output frequency-selective block-fading channels. It is shown that the proposed STF codes can achieve rate Mtand full-diversity MtMrMbL, i.e., the product of the number of transmit antennas Mt, receive antennas Mr, fading blocks Mb, and channel taps L. The proposed STF codes are constructed from a layered algebraic design, where each layer of algebraic coded symbols are parsed into different transmit antennas, orthogonal frequency-division multiplexing tones, and fading blocks without rate loss. Simulation results show that the proposed STF codes achieve higher diversity gain in block-fading channels than some typical space-frequency codes
Wei Zhang 0001, Xiang-Gen Xia 0001, Pak-Chung Ching
IEEE Trans. Commun.3
2007 Clustered pilot tones for carrier frequency offset estimation in OFDM systems
abstract
In this paper, we propose a new pilot tone placement scheme and a novel pilot sequence design to estimate the carrier frequency offset (CFO) in OFDM systems. Unlike the conventional approach in which isolated pilot tones are used, we cluster two pilot tones together as a group and these tone groups are equally spaced along the subcarrier axis. The pilot sequence of each group is carefully designed, in which the pilot symbol on the left side in each cluster is made antipodal to the one on the right. The performance in terms of the signal-to-interference ratio (SIR) can be significantly improved by the proposed scheme. Theoretical analysis and simulation results show that the clustered pilot tones can give a substantially lower variance of CFO estimate than that of the isolated pilot tones. For a given CFO, it demonstrates that the variance of the estimate can be reduced by 3 - 14 dB. The bit error rate (BER) is reduced by about 1 dB both in AWGN and multipath fading channels in the simulated cases
Wei Zhang 0001, Xiang-Gen Xia 0001, Pak-Chung Ching
IEEE Trans. Wirel. Commun.3
2007 Full-Diversity and Fast ML Decoding Properties of General Orthogonal Space-Time Block Codes for MIMO-OFDM Systems
abstract
In this letter we apply the general orthogonal space-time block codes (OSTBC) to MIMO-OFDM systems over frequency-selective fading channels and aim to exploit the potential multipath diversity. By replacing the scalar entry of an OSTBC matrix with the vector of repeated symbols, we obtain a new OSTBC which can achieve both spatial diversity and multipath diversity. Moreover, a fast maximum likelihood (ML) decoding is admitted. Simulation results show that the proposed OSTBC, for two transmit antennas, can obtain a higher diversity gain than the Alamouti code at the same ML decoding complexity
Wei Zhang 0001, Xiang-Gen Xia 0001, Pak-Chung Ching
IEEE Trans. Wirel. Commun.3
2006 Design of Orthogonal Space-Time Block Codes for MIMO-OFDM Systems with Full Diversity and Fast Ml Decoding
abstract
In this paper, a design of orthogonal space-time block codes (OSTBC) is proposed for MIMO-OFDM systems operating in frequency-selective fading channels. By stacking the repeated symbols as entries of an OSTBC matrix, the proposed OSTBC can achieve full diversity in frequency-selective fading channels. Moreover, the ML decoding can be performed separately on each symbol, i.e., single-symbol decoding. Simulation results show that for two transmit antennas the proposed OSTBC can achieve a higher diversity gain than the Alamouti code yet maintaining a single-symbol decoding simplicity.
Wei Zhang 0001, Xiang-Gen Xia 0001, Pak-Chung Ching
ICASSP (4)3
2006 An Iterative Trajectory Regeneration Algorithm for Separating Mixed Speech Sources
abstract
Harmonicity and continuity are two important perceptual cues for separating mixed speech sources. This paper focuses on the separation of two speech sources with a single-microphone input. An iterative, least-squares (LS) based trajectory regeneration algorithm is proposed to estimate the magnitude spectrum of each source. Time-derivatives of the spectrum, or the dynamic spectral information, is used as a constraint in solving the resultant weighted normal equations. Each estimated spectral trajectory, as a result, exhibits similar temporal variations as the original source. Asymptotically, we also prove that the regenerated trajectory yields the same time variations as the given dynamic information. When cascaded with our previously proposed harmonic filtering algorithm to separate mixed voiced signals, the new trajectory regeneration is shown to be very effective to reduce mean squared errors by 82.2% and 69.5%, relatively, with ideal and approximated dynamic information, respectively
Siu Wa Lee, Frank K. Soong, Pak-Chung Ching
ICASSP (1)3
2006 Extended Differential Unitary Space-Time Modulation: A Non-Coherent Scheme with Error Penalty Less Than 3DB
abstract
In this paper we propose an extended differential unitary space-time modulation (xDUSTM) scheme that can offer improved error performance over the differential unitary space-time modulation (DUSTM) scheme. DUSTM is well suited to rapidly time-varying unknown channels. It has a simple structure, but incurs an error performance penalty of about 3dB compared to its coherent counterpart. The xDUSTM scheme considers moderately fast time-varying channels, and is designated to exploit such a characteristic for performance improvement. In xDUSTM, a problem that needs to be addressed is the complexity of its non-coherent maximum-likelihood (ML) receiver. We show that by choosing the orthogonal space-time block code (OSTBC) designs, the ML problem can be reduced to a Boolean quadratic program for which highly effective algorithms are available. Simulation results illustrate that the error performance penalty in xDUSTM can be reduced to 1dB.
Wing-Kin Ma, Chong-Yung Chi, Pak-Chung Ching
ICASSP (4)3
2006 Constrained Least Squares Estimation for Position Tracking
abstract
This paper describes an effective method for mobile location tracking and velocity estimation in a network-based wireless localization system. We propose a constrained least squares estimation algorithm for real-time tracking of the location and dynamic motion of a mobile user using TOA measurements. The tracking problem is formulated in a state-space framework and the constraints on system states are considered explicitly. Simulation results show that the proposed tracking algorithm can improve the accuracy through eliminating spurious position estimates
Pak-Chung Ching
ICASSP (4)2
2006 Automatic speech recognition of Cantonese-English code-mixing utterances
Joyce Y. C. Chan, Pak-Chung Ching, Tan Lee, Houwei Cao
INTERSPEECH2
2006 Automatic emotion recognition of speech signal in Mandarin
Pak-Chung Ching, Fanrang Kong
INTERSPEECH2
2005 High-rate full-diversity space-time-frequency codes for MIMO multipath block-fading channels
abstract
In this paper, we propose a systematic design of rate-M/sub t/ full-diversity space-time-frequency (STF) code for MIMO frequency-selective block-fading channels. The full-diversity can be proved to be the product of number of transmit antennas M/sub t/, receiver antennas M/sub r/, fading blocks M/sub b/ and channel taps L. The performance of the proposed STF codes are simulated and compared with some existing SF codes. The full-diversity of STF is further validated from the simulated case.
Wei Zhang 0001, Xiang-Gen Xia 0001, Pak-Chung Ching
GLOBECOM3
2005 On implementing the blind ML receiver for orthogonal space-time block codes
abstract
We consider the problem of blind maximum-likelihood (ML) detection for the orthogonal space-time block code (OSTBC) scheme. Our previous work has shown that the problem can be simplified to a Boolean quadratic program (BQP). This sequel focuses on effective optimization methods for that BQP, which, from an optimization viewpoint, is still a computationally hard problem. First, we consider semidefinite relaxation (SDR), a high-precision BQP approximation algorithm with a computational cost that is polynomial in the problem size. We also propose a simple method that can significantly reduce the average complexity of the SDR technique. Second, we consider sphere decoding, an exact BQP solver that can be computationally expensive in the worst case, but generally incurs a reasonable average complexity particularly at high SNRs. Simulation results indicate that these two blind ML algorithms provide very similar bit error rate performance. Moreover, numerical studies show that SDR provides better complexity performance than sphere decoding in the worst-case sense, while sphere decoding provides better complexity performance in the average sense.
Wing-Kin Ma, Ba-Ngu Vo, Timothy N. Davidson, Pak-Chung Ching
ICASSP (3)4
2005 Development of a Cantonese-English code-mixing speech corpus
Joyce Y. C. Chan, Pak-Chung Ching, Tan Lee
INTERSPEECH2
2005 Embedded Cantonese TTS for multi-device access to web content
abstract
This paper describes the development of an embedded Cantonese text-to-speech synthesizer to enable multi-device access to Chinese Web content. Advancements in wireless communication is driving Web visitors from using desktop PCs to mobile handheld devices. Significant reduction in the form factors of the client devices tends to shift information delivery from the visual to the aural modality. This calls for synthesizers that can run on relatively stringent computational and storage resources of handheld devices. We report on the migration of our Cantonese synthesizer, CU VOCAL, from the desktop to the embedded platform. Migration preserves the support for speech synthesis markups (SSML), ensures code compatibility and lowers the storage requirements of the syllable inventory. Results from listening tests indicate no signification deterioration in synthesis quality of embedded CU VOCAL when compared to its desktop counterpart.
Tien Ying Fung, Yuk-Chi Li, Eddie Sio, Icarus Lee, Helen M. Meng, Pak-Chung Ching
INTERSPEECH6
2005 Harmonic filtering for joint estimation of pitch and voiced source with single-microphone input
Siu Wa Lee, Frank K. Soong, Pak-Chung Ching
INTERSPEECH3
2004 A design of high-rate space-frequency codes for MIMO-OFDM systems
abstract
In this paper, a design of high-rate space-frequency codes (SFC) is proposed for MIMO-OFDM systems. The proposed SFC is achieved by partitioning data symbols in one OFDM block into many blocks and then precoding each block with an unitary matrix. For M/sub t/ transmit antennas, it can achieve a symbol rate M/sub t/ and it is numerically shown that it has full diversity in some interesting cases.
Wei Zhang 0001, Xiang-Gen Xia 0001, Pak-Chung Ching
GLOBECOM3
2004 On the number of pilots for OFDM system in multipath fading channels
abstract
The orthogonal frequency division multiplexing (OFDM) system has been demonstrated to be effective in a frequency selective multipath fading environment. To compensate the deleterious effects of the fading channels, pilot symbol assisted modulation (PSAM) technique has been used to obtain the channel estimate. In order to have an accurate estimate, a high percentage of pilot symbols is usually needed and thus reducing the spectrum efficiency. Therefore, the number of pilot symbols is a tradeoff between channel estimation accuracy and bandwidth efficiency. In this paper, we propose a novel method of deciding the optimal number of pilots in OFDM transmission. The method is based on the derivation and analysis of bit error rate (BER) of a PSAM-OFDM system in multipath fading channel in terms of the pilot spacing for a given block size.
Wei Zhang 0001, Xiang-Gen Xia 0001, Pak-Chung Ching, Wing-Kin Ma
ICASSP (4)3
2004 Blind symbol identifiability of orthogonal space-time block codes
abstract
This paper addresses the blind symbol identifiability of the orthogonal space-time block code (OSTBC) scheme. That is, the conditions under which OSTBC symbols can be identified without ambiguity when channel state information is not available. In many space-time communication schemes, achieving unique blind symbol identification requires certain assumptions on the number of receiver antennas and the rank of the channel matrix. In this paper, we show that unique blind symbol identification of OSTBCs is possible for any number of receiver antennas and for any (nonzero) channel matrix. This attractive unique identifiability result is shown to be achieved by a class of OSTBCs that exhibit certain matrix non-rotational properties. Using these properties, we validate the identifiability of a number of commonly used OSTBCs.
Wing-Kin Ma, Pak-Chung Ching, Timothy N. Davidson, Ba-Ngu Vo
ICASSP (4)2
2004 Bilingual Chinese/English voice browsing based on a VoiceXML platform
abstract
We report on the development of English-Chinese bilingual speech applications on a VXML platform. VXML supports displayless voice browsing of Web content. We have developed the CU voice browser based on OpenVXI 2.0. We have also integrated the voice browser with the OpenSpeech Recognizer (for English speech recognition), CU RSBB (for Chinese speech recognition), Speechify (for English speech synthesis) and CU VOCAL (for Chinese speech synthesis) in order to support bilingual voice browsing. The CU voice browser includes an attribute for identifying the appropriate language for speech input/output, thereby invoking the appropriate speech engine for recognition/synthesis. We have developed two bilingual sample applications $CU weather and CU news. This paper provides the associated VXML documents that specify these bilingual dialogs for browsing weather and news information respectively.
Helen M. Meng, Yuk-Chi Li, Tien Ying Fung, Kon Fan Low, Ka-Fai Chow, Tin Hang Lo, Man Cheuk Ho, Pak-Chung Ching
ICASSP (3)8
2004 Reduced-rank blind adaptive frequency-shift filtering for signal extraction
abstract
In this paper, we first illustrate that a blind adaptive frequency-shift (BA-FRESH) filter can be represented as a generalized sidelobe canceler (GSC). Since the computational power of the BA-FRESH filter is quite high, a reduced-rank implementation is thus proposed and achieved by using the eigen-subspace method. To avoid under-representation, a rule for choosing the rank/dimension of the signal subspace is introduced by looking at the eigenvalue spread of the signal covariance matrix. The proposed PCA-based reduced-rank BA-FRESH filter not only has a lower computational complexity, but is also more efficient in signal extraction when compared with the conventional, CSP-based and Krylov subspace-based BA-FRESH filters. The performance of this new method in reducing the spectrally overlapped interference of BPSK signals has been examined rigorously.
Lai Yin Ngan, Shan Ouyang 0001, Pak-Chung Ching
ICASSP (2)3
2004 Using Haar transformed vocal source information for automatic speaker recognition
abstract
This paper attempts to investigate the effectiveness of incorporating vocal source information for enhancing automatic speaker recognition accuracy. We propose a new method to extract discriminative features from the linear prediction (LP) residual signal, which are closely related to the glottal excitation of individual speaker. A complementary parameter set in addition to the commonly used linear predictive cepstral coefficients (LPCC), called Haar octave coefficients of residue (HOCOR), is obtained by applying a Haar transform to the LP residue. This additional feature vector retains the spectro-temporal characteristics of the source excitation sequences that are related to the fundamental frequency, harmonics, as well as their phases. Experimental evaluation over the YOHO corpus demonstrates the high speaker discriminative power and high inter-speaker variability of HOCOR. Speaker recognition tests with both vocal tract feature (LPCC) and vocal source information (HOCOR) outperform the conventional methods of using LPCC only.
Nengheng Zheng, Pak-Chung Ching
ICASSP (1)2
2004 In-phase feature induction: an effective compensation technique for robust speech recognition
Siu Wa Lee, Pak-Chung Ching
INTERSPEECH2
2004 Time -frequency analysis of vocal source signal for speaker recognition
abstract
This paper investigates the importance of spectro-temporal characteristics of the source excitation signal for speaker recognition. We propose an effective feature extraction technique for obtaining essential time-frequency information from the linear prediction (LP) residual signal, which are closely related to the glottal excitation of individual speaker. With pitch synchro-nous analysis, wavelet transform is applied to every two pitch cycles of the LP residual signal to generate a new feature vector, called Wavelet Octave Coefficients of Residues (WOCOR), which provides additional speaker discriminative power to the commonly used linear predictive Cepstral coefficients (LPCC). Experimental evaluation over a Cantonese speaker recognition corpus demonstrates the effectiveness of WOCOR for speaker recognition. Recognition tests with WOCOR and LPCC outperforms the conventional methods of using Mel Frequency Cepstral Coefficients (MFCC). 1.
Nengheng Zheng, Pak-Chung Ching, Tan Lee
INTERSPEECH2
2004 Crosstalk resilient interference cancellation in microphone arrays using Capon beamforming
abstract
This paper studies a reference-assisted approach for interference canceling (IC) in microphone array systems. Conventionally, reference-assisted IC is based on the zero crosstalk assumption; i.e., when the desired source signal is absent in the reference microphones. In applications where crosstalk is inevitable, the conventional IC approach usually exhibits degraded performance due to cancellation of the desired signal. In this paper, we develop a crosstalk resilient IC method based on the Capon beamforming technique. The proposed beamformer deals with the uncertainty of crosstalk by applying a constraint on the worst-case crosstalk magnitude. The proposed beamformer not only performs IC, it also provides blind beamforming of the desired signal. We show that a blind beamformer based on the traditional minimum-mean-square-error (MMSE) IC method is a special case of the proposed beamformer. One key step of implementing the proposed Capon beamformer lies in solving a difficult nonconvex optimization problem, and we illustrate how the Capon optimal solution can be effectively approximated using the so-called semidefinite relaxation algorithm. Simulation results demonstrate that the proposed beamformer is more robust against crosstalk-induced signal cancellation than beamformers based on the MMSE-IC methods.
Wing-Kin Ma, Pak-Chung Ching, Ba-Ngu Vo
IEEE Trans. Speech Audio Process.2
2004 ISIS: an adaptive, trilingual conversational system with interleaving interaction and delegation dialogs
abstract
ISIS (Intelligent Speech for Information Systems) is a trilingual spoken dialog system (SDS) for the stocks domain. It handles two dialects of Chinese (Cantonese and Putonghua) as well as English---the predominant languages in our region. The system supports spoken language queries regarding stock market information and simulated personal portfolios. The conversational interface is augmented with a screen display that can capture mouse-clicks as well as textual input by typing or stylus-writing. Real-time information is retrieved directly from a dedicated Reuters satellite feed. ISIS provides a system test-bed for our work in multilingual speech recognition and generation, speaker authentication, language understanding and dialog modeling. This article reports on our new explorations within the context of ISIS, including: (i) adaptivity to knowledge scope expansion; (ii) asynchronous human-computer interaction by task delegation to software agents; (iii) multi-threaded online interaction and offline delegation dialogs with interruptions for task switching.
Helen M. Meng, Pak-Chung Ching, Shuk Fong Chan, Yee Fong Wong, Cheong Chat Chan
ACM Trans. Comput. Hum. Interact.2
2003 Blind maximum-likelihood decoding for orthogonal space-time block codes: a semidefinite relaxation approach
abstract
Orthogonal space-time block codes (OSTBCs) have attracted much attention because they provide an effective and simple scheme for fully utilizing the diversity gain in multi-antenna systems. We address the problem of decoding OSTBCs without channel state information. We place our emphasis on the blind maximum-likelihood (ML) method with the BPSK constellation, and show that blind ML decoding requires the solution of a computationally hard optimization problem. To overcome this computational difficulty, we propose using a high-precision and efficient approximation algorithm, called semidefinite relaxation (SDR), to implement blind ML decoding suboptimally. The resultant SDR-ML blind decoder is efficient in that its complexity is approximately cubic in the number of symbols processed, and is promising for its appealing theoretical worst-case approximation accuracy. Simulation results show that the bit error performance of the SDR-ML blind decoder is substantially better than that of several other blind decoders including the cyclic ML method and the subspace method.
Wing-Kin Ma, Pak-Chung Ching, Timothy N. Davidson, Xiang-Gen Xia 0001
GLOBECOM2
2003 Robust interference suppression and blind speech beamforming in room reverberant environments
abstract
In a microphone array system where references of interference are additionally available, the unwanted signals in the received signals can be removed using Widrow's interference-cancelling (IC) approach. However, in the presence of crosstalk, IC can result in severe cancellation of the desired speech signal. In this paper, we propose a crosstalk-resistant method for joint interference suppression and blind speech signal beamforming. The proposed method is based on the Capon blind estimation principle, and is implemented using a powerful approximation tool, namely semidefinite relaxation. Simulation results show that the proposed method yields improved mean squared error performance compared with the IC-based method.
Wing-Kin Ma, Pak-Chung Ching
ICASSP (5)2
2003 Recent enhancements in CU VOCAL for Chinese TTS-enabled applications
abstract
CU VOCAL is a Cantonese text-to-speech (TTS) engine. We use a syllable-based concatenative synthesis approach to generate intelligible and natural synthesized speech [1]. This paper describes several recent enhancements in CU VOCAL. First, we have augmented the syllable unit selection strategy with a positional feature. This feature specifies the relative location of a syllable in a sentence and serves to improve the quality of Cantonese tone realization. Second, we have developed the CU VOCAL SAPI engine, a version of the synthesizer that eases integration with applications using SAPI (Speech Application Programming Interface). We demonstrate the use of CU VOCAL SAPI in an electronic book (e-book) reader. Third, we have made an initial attempt to use the CU VOCAL SAPI engine in Web content authored with Speech Application Language Tags (SALT). The use of SALT tags can ease the task of invoking Cantonese TTS service on webpages.
Helen M. Meng, Yuk-Chi Li, Tien Ying Fung, Man Cheuk Ho, Chi-Kin Keung, Tin Hang Lo, Wai Kit Lo, Pak-Chung Ching
INTERSPEECH8
2003 Joint time delay and frequency estimation via state-space realization
abstract
By applying a two-dimensional parameter estimation method proposed by M. Viberg and P. Stoica (see Conf. Rec. 32nd Asilomar Conf. Signals, Systems, Computers, vol.2, p.735-9, 1998), we develop a subspace method for estimating the differential delay of a sinusoidal signal received at two separated sensors as well as the sinusoidal frequencies. Using state-space realization, the time delay and frequency estimates are obtained from the state transition and observation matrices. Performance evaluation via computer simulations is included to demonstrate the effectiveness of the proposed algorithm.
Yuntao Wu, Hing-Cheung So, Pak-Chung Ching
IEEE Signal Process. Lett.3
2003 Cross-language spoken document retrieval using HMM-based retrieval model with multi-scale fusion
abstract
Cross-language spoken document retrieval (CL-SDR) is the technology that facilitates automatic retrieval of relevant information from a collection of spoken documents in a language that is different from that used in the queries. Information sources that are in different languages can then be retrieved automatically with CL-SDR, and the number of searchable information sources will increase significantly. The HMM-based retrieval model is a probabilistic formulation for the retrieval problem. Extensions to this retrieval model can be made by taking advantage of its probabilistic nature. Specifically, we have incorporated the translation component to make it possible to perform cross-language information retrieval (CLIR). In addition, this HMM-based CLIR retrieval model is also extended for retrieval at subword scales.In this work the extended HMM-based retrieval model has been applied to an English-Mandarin CL-SDR task, which is to search the Mandarin spoken document collection with English queries at word and subword scales. Retrieval results obtained from these indexing scales are then fused for multi-scale CL-SDR. Experimental results demonstrate that improvement in CL-SDR retrieval performance can be achieved by fusion of word and subword scales.
Wai Kit Lo, Helen M. Meng, Pak-Chung Ching
ACM Trans. Asian Lang. Inf. Process.3
2002 Multiuser detection for asynchronous CDMA using block coordinate ascent and semi-definite relaxation
abstract
Maximum-likelihood (ML) multiuser detection provides attractive bit error rate performance, but it is computationally prohibitive to implement (except in certain restricted cases). Recently, it has been shown that ML detection for synchronous CDMA can be efficiently and accurately approximated using the semi-definite relaxation (SDR) method. In this work, we consider the application of SDR to the more general scenario of asynchronous CDMA. To make this application computationally feasible, we incorporate a block coordinate ascent (BCA) technique into our detector. Simulation studies show that the resulting BCA-SDR detector has significantly better BER performance than several typical suboptimal multiuser detectors.
Wing-Kin Ma, Timothy N. Davidson, Kon Max Wong, Pak-Chung Ching
ICASSP4
2002 Fast algorithm for adaptive estimation of principal and minor components
abstract
With the penalty function method, a novel information criterion has been developed for searching the desired eigen-components. The estimation of principal and minor eigen-components is governed by a penalty factor, and a unified algorithm is devised that is capable of fast estimation of both the principal and minor components. Simulation results show that the new algorithm has a fast convergence rate when compared with the gradient-based fixed step-size algorithm for extracting principal and minor components.
Shan Ouyang 0001, Pak-Chung Ching
ICASSP2
2002 Multi-scale and multi-model integration for improved performance in Chinese spoken document retrieval
abstract
This paper describes our attempt to combine the relative merits of different indexing units (scales) and different retrieval models to improve performance in Chinese spoken document retrieval. Our study includes indexing units from three scales: words, character bigrams and syllable bigrams. We also include two different retrieval models: the HMM-based model and the vector space model (VSM). Our retrieval task is based on the TDT-2 Mandarin collection- news text is used to retrieve relevant Mandarin audio. We experimented with different scales and retrieval models. The HMMbased model retrieves better at the word scale (mAP=0.566). For the VSM, better performance is obtained at the character bigram scale (mAP=0.562). We proceeded with a series of integration experiments where the ranked retrieval lists from different runs are combined by rank-based re-scoring. The best retrieval performance (mAP=0.591) is achieved when we integrate the HMMword and VSM-character configurations. These results suggest that retrieval based on different scales and different models capture different kinds of knowledge, which can be integrated to improve retrieval performance. 1.
Wai Kit Lo, Helen M. Meng, Pak-Chung Ching
INTERSPEECH3
2002 ISIS: a multi-modal, trilingual, distributed spoken dialog system developed with CORBA, java, XML and KQML
abstract
ISIS (Intelligent Speech for Information Systems) is a trilingual spoken dialog system in the stocks domain. It supports the three languages commonly used in Hong Kong (Cantonese, Putonghua and English), and serves as a test-bed for our research in various speech and language technologies. ISIS also features combined interaction and delegation dialogs, and automatic assimilation of newly listed stock names into the system’s knowledge base. This paper focuses on the architecture and multi-modality of ISIS. We use the CORBA middleware to implement a distributed system that is interoperable across platforms. We also describe the incorporation of KQML (Knowledge Query and Manipulation Language) software agents in ISIS to handle delegation dialogs. The latest enhancement supports multi-modal and mixed-modal input which suit the natural affordances of certain interactions in order to improve usability. Input modalities include speaking, typing or mouse-clicking. Output media include synthesized speech, text, tables and graphics. 1.
Helen M. Meng, Pak-Chung Ching, Yee Fong Wong, Cheong Chat Chan
INTERSPEECH2
2002 CU VOCAL: corpus-based syllable concatenation for Chinese speech synthesis across domains and dialects
abstract
This paper describes CU VOCAL, a Chinese text-to-speech synthesis system that adopts the approach of corpus-based syllable concatenation. We have demonstrated the applicability of the approach primarily for Cantonese, a major dialect of Chinese predominant in Hong Kong, South China and many overseas Chinese communities. This work extends our previous work as described in [1]. Our approach is able to synthesize speech from free-form text, and it can also be optimized for response generation in specific application domains. We have also demonstrated the portability of the approach to Putonghua, the official Chinese dialect, in a domain-optimized setting. Coarticulatory context is expressed in terms of distinctive features. Tonal context is also included. We conducted a series of listening tests using CU VOCAL, which gave favorable performance. 1.
Helen M. Meng, Chi-Kin Keung, Kai-Chung Siu, Tien Ying Fung, Pak-Chung Ching
INTERSPEECH5
2002 Spoken language resources for Cantonese speech processing
Tan Lee, Wai Kit Lo, Pak-Chung Ching, Helen M. Meng
Speech Commun.3
2002 Using tone information in Cantonese continuous speech recognition
abstract
In Chinese languages, tones carry important information at various linguistic levels. This research is based on the belief that tone information, if acquired accurately and utilized effectively, contributes to the automatic speech recognition of Chinese. In particular, we focus on the Cantonese dialect, which is spoken by tens of millions of people in Southern China and Hong Kong. Cantonese is well known for its complicated tone system, which makes automatic tone recognition very difficult. This article describes an effective approach to explicit tone recognition of Cantonese in continuously spoken utterances. Tone feature vectors are derived, on a short-time basis, to characterize the syllable-wide patterns of F0 (fundamental frequency) and energy movements. A moving-window normalization technique is proposed to reduce the tone-irrelevant fluctuation of F0 and energy features. Hidden Markov models are employed for context-dependent acoustic modeling of different tones. A tone recognition accuracy of 66.4% has been achieved in the speaker-independent case. The recognized tone patterns are then utilized to assist Cantonese large-vocabulary continuous speech recognition (LVCSR) via a lattice expansion approach. Experimental results show that reliable tone information helps to improve the overall performance of LVCSR.
Tan Lee, Wai H. Lau, Yiu Wing Wong, Pak-Chung Ching
ACM Trans. Asian Lang. Inf. Process.4
2001 Joint time delay and frequency estimation of multiple sinusoids
abstract
We devise a new subspace method for estimating the differential time delay of a signal received at two separated sensors as well as the frequencies of the source signal, assuming that it consists of multiple sinusoids. The time delay and frequency estimates are related to the eigenvalues and eigenvectors of a matrix obtained from the covariances of the received signals. The effectiveness of the proposed algorithm is demonstrated via computer simulations using sinusoidal signals as well as real speech data.
Guisheng Liao, Hing-Cheung So, Pak-Chung Ching
ICASSP3
2001 Compensation of amplifier nonlinearities on wavelet packet division multiplexing
abstract
Wavelet packet division multiplexing (WPDM) is a high-capacity, flexible and robust orthogonal multiplexing scheme in which the message signals are waveform-coded onto wavelet packet basis functions for transmission. However, WPDM suffers from severe performance degradation in the presence of high-power amplifier (HPA) nonlinearities. Data predistortion using the pth-order Volterra inverse is proposed to combat the amplifier nonlinearities in a WPDM system. A 5th-order Volterra inverse with truncated memory length is designed based on the Volterra series channel model. Computer simulations are presented to demonstrate the capability of the proposed technique in compensating amplifier nonlinearities even under system parameter discrepancies. Guidelines are also proposed for designing a wavelet filter which leads to better predistortion with the truncated Volterra inverse.
Kin-Fai To, Pak-Chung Ching, Kon Max Wong
ICASSP2
2001 Efficient quasi-maximum-likelihood multiuser detection by semi-definite relaxation
abstract
In multiuser detection, maximum-likelihood detection (MLD) is optimum in the sense of minimum error probability. Unfortunately, MLD involves a computationally difficult optimization problem for which there is no known polynomial-time solution (with respect to the number of users). In this paper, we develop an approximate maximum-likelihood (ML) detector using semi-definite (SD) relaxation for the case of anti-podal data transmission, SD relaxation is an accurate and efficient approximation algorithm for certain difficult optimization problems. In MLD, SD relaxation is efficient in that its complexity is O(K/sup 3.5/), where K stands for the number of users. Simulation results indicate that the SD relaxation ML detector has its bit error performance close to the true ML detector, even when the cross-correlations between users are strong or the near-far effect is significant.
Wing-Kin Ma, Timothy N. Davidson, Kon Max Wong, Zhi-Quan Luo, Pak-Chung Ching
ICC5
2001 A "self-decorrelating" technique to enhance blind space-time RAKE receivers with single-user-type DS-CDMA detectors
abstract
A novel "self-decorrelating" technique is proposed to enhance the multipath constructive summation capability and the interference rejection capability of a class of maximum-SINR blind space-time CDMA RAKE receivers designed to tackle the near-far problem. This proposed "self-decorrelating" technique appears to have effectively removed the near-far problem's error floor at high SNR; and the proposed technique can also significant decrease the error rate. Moreover, this "blind" space-time processing receiver architecture needs no prior knowledge nor explicit estimation of (1) the channel's multipath arrival angle or arrival delay or power profile, (2) the receiver's nominal or actual antenna array manifold, and (3) the other CDMA users' signature spreading codes. Preliminary simulations suggest very significant performance improvement realizable from the proposed technique. Though developed for direct-sequence CDMA, the proposed technique might be adopted for frequency-hopping CDMA.
Kainam Thomas Wong, Guisheng Liao, Shun Keung Cheung, Michael D. Zoltowski, Javier Ramos 0001, Pak-Chung Ching
ICC6
2001 ISIS: a learning system with combined interaction and delegation dialogs
abstract
This paper presents a progress update of our ISIS 1 trilingual spoken dialog system. As described in [8], this is a conversational system for the stocks domain, and supports interactions in the languages of our region – English and two dialects of Chinese (Mandarin and Cantonese). ISIS provides a system test-bed for our initial explorations with the CORBA architecture, and delegation to KQML (Knowledge Query and Manipulation Language) agents. CORBA offers the advantages of interoperability, scalability and location transparency in client/server systems development. Users can delegate tasks to software agents to help monitor information (e.g. a drop in the price of a pre-specified stock), and generate user alert messages. Our current work presents new research directions in the context of ISIS: (i) automatic incorporation of newly listed stocks into our system’s knowledge base; (ii) switching between on-line interaction and off-line delegation in a single dialog thread. We will also report on enhancements in the system’s architecture and features (e.g. automatic end-point detection). delegation in a single dialog thread. We will also report on enhancements in the system’s architecture and features (e.g. automatic end-point detection). 2. System Architecture Previous work in the development of software infrastructures for dialog systems include [11] 2 [2]. Over the past year, we have continued to explore the development of a spoken dialog system based on the CORBA architecture. This middleware resides in between the operating system and the application layer 1.
Helen M. Meng, Shuk Fong Chan, Yee Fong Wong, Cheong Chat Chan, Yiu Wing Wong, Tien Ying Fung, Wai Ching Tsui, Ke Chen 0001, Ting-Yao Wu, Tan Lee, Wing Nin Choi, Pak-Chung Ching, Huisheng Chi
INTERSPEECH14
2001 On computing Verdu's upper bound for a class of maximum-likelihood multiuser detection and sequence detection problems
abstract
The upper bound derived by Verdu (1986) is often used to evaluate the bit error performance of both the maximum-likelihood (ML) sequence detector for single-user systems and the ML multiuser detector for code-division multiple-access (CDMA) systems. This upper bound, which is based on the concept of indecomposable error vectors (IEVs), can be expensive to compute because in general the IEVs may only be obtained using an exhaustive search. We consider the identification of IEVs for a particular class of ML detection problems commonly encountered in communications. By exploiting the properties of the IEVs for this case, we develop an IEV generation algorithm which has a complexity substantially lower than that of the exhaustive search. We also show that for specific communication systems, such as duobinary signaling, the expressions of Verdu's upper bound can be considerably simplified.
Wing-Kin Ma, Kon Max Wong, Pak-Chung Ching
IEEE Trans. Inf. Theory3
2000 Acoustic modeling for Chinese speech recognition: a comparative study of Mandarin and Cantonese
abstract
This paper presents a comparative study on automatic speech recognition for two different Chinese dialects, namely Mandarin and Cantonese. It focuses on decision-tree based context-dependent acoustic modeling for large-vocabulary continuous speech recognition. Extensive phonological and phonetic knowledge are incorporated to design questions concerning the left and right context of sub-syllable units, namely INITIALs and FINALs. This results in a set of class-triphone models for each dialect. Syllable recognition accuracy of 81.7% and 75.5% are attained for Mandarin and Cantonese respectively. Such a performance gap is accountable by various linguistic and practical reasons, including: 1) phonological and phonetic discrepancies between the two dialects; 2) design of training databases; and 3) design of phonetic questions in decision-tree clustering.
Tan Lee, Yiu Wing Wong, Bo Xu 0002, Pak-Chung Ching, Taiyi Huang 0001
ICASSP5
2000 Maximum likelihood detection for multicarrier systems employing non-orthogonal pulse shapes
abstract
Investigation of detection schemes for non-orthogonal multicarrier modulation (MCM) is motivated by two reasons. Firstly, non-orthogonal MCM offers a higher degree of freedom in pulse-shaping design. Secondly, the problem of detecting orthogonal MCM under channel distortion can be viewed as a problem of detecting non-orthogonal MCM. In this work, the maximum likelihood detector (MLD) is considered for non-orthogonal multicarrier systems. In the absence of inter-block interference, it is shown that the MLD can be efficiently achieved by a Viterbi algorithm (VA). In contrast to using the VA for channel equalization, the proposed VA has its survivor metrics running in the""frequency domain". Incorporating this VA with an interference-canceling approach, we also develop a decision feedback MLD for the case of non-zero inter-block interference. Superior bit error performance of the MLDs is demonstrated by simulations.
Wing-Kin Ma, Pak-Chung Ching, Kon Max Wong
ICASSP2
2000 Lexical tree decoding with a class-based language model for Chinese speech recognition
Wing Nin Choi, Yiu Wing Wong, Tan Lee, Pak-Chung Ching
INTERSPEECH4
2000 Incorporating tone information into Cantonese large-vocabulary continuous speech recognition
Wai H. Lau, Tan Lee, Yiu Wing Wong, Pak-Chung Ching
INTERSPEECH4
2000 ISIS: A multilingual spoken dialog system developed with CORBA and KQML agents
abstract
ISIS, which abbreviates Intelligent Speech for Information Systems, is a trilingual spoken dialog system (SDS) for the financial domain. It handles two dialects of Chinese (Cantonese and Putonghua), as well as English the predominant languages in our region. The system supports spoken language queries regarding stock market information and simulated personal portfolios. Real-time information is retrieved directly from a dedicated Reuters satellite feed. ISIS provides a system test-bed for our work in multilingual speech recognition and generation, speaker authentication, language understanding and dialog modeling. Furthermore, ISIS supports our initial explorations in: (i) CORBA's interoperability and scalability for SDS development; in conjunction with (ii) asynchronous human-computer interaction by delegation to KQML software agents...
Helen M. Meng, Shuk Fong Chan, Yee Fong Wong, Tien Ying Fung, Wai Ching Tsui, Tin Hang Lo, Cheong Chat Chan, Ke Chen 0001, Ting-Yao Wu, Tan Lee, Wing Nin Choi, Yiu Wing Wong, Pak-Chung Ching, Huisheng Chi
INTERSPEECH15
2000 Multi-scale audio indexing for Chinese spoken document retrieval
Helen M. Meng, Wai Kit Lo, Yuk-Chi Li, Pak-Chung Ching
INTERSPEECH4
2000 Performance of wavelet packet-division multiplexing in impulsive and Gaussian noise
abstract
Wavelet packet-division multiplexing (WPDM) is a high-capacity, flexible, and robust multiple-signal transmission technique in which the message signals are waveform coded onto wavelet packet basis functions for transmission. We derive an expression for the probability of error for a WPDM scheme in the presence of both impulsive and Gaussian noise sources and demonstrate that WPDM can provide greater immunity to impulsive noise than both a time-division multiplexing scheme and an orthogonal frequency-division multiplexing scheme.
Kon Max Wong, Jiangfeng Wu, Timothy N. Davidson, Qu Jin, Pak-Chung Ching
IEEE Trans. Commun.5
1999 Two-dimensional multi-resolution analysis of speech signals and its application to speech recognition
abstract
This paper describes a novel approach of using multi-resolution analysis (MRA) for automatic speech recognition. Two-dimensional MRA is applied to the short-time log spectrum of speech signal to extract the slowly varying spectral envelope that contains the most important articulatory and phonetic information. After passing through a standard cepstral analysis process, the MRA features are used for speech recognition in the same way as conventional short-time features like MFCCs, PLPs, etc. Preliminary experiments on both clean connected speech and noisy telephone conversation speech show that the use of MRA cepstra results in a significant reduction in insertion error when compared with MFCCs.
Chun-Ping Chan, Yiu Wing Wong, Tan Lee, Pak-Chung Ching
ICASSP4
1999 A 1.7KBPS waveform interpolation speech coder using decomposition of pitch cycle waveform
Pak-Chung Ching
EUROSPEECH2
1999 Micro-prosodic control in cantonese text-to-speech synthesis
Tan Lee, Helen M. Meng, Wai H. Lau, Wai Kit Lo, Pak-Chung Ching
EUROSPEECH5
1999 Acoustic modeling and language modeling for cantonese LVCSR
abstract
This paper describes our recent work on the development of a large-vocabulary, speaker-independent continuous speech recognition system for Cantonese (a major Chinese dialect). Both acoustic modeling and language modeling are being addressed. For acoustic modeling, we focus on right-context-dependent sub-syllable units. Tying of HMM at model as well as state level is applied based on phonetic knowledge and the decision-tree approach. Statistical language model is built from large amount of newspaper text. The overall recognition accuracy for syllable and Chinese character are 81.83% and 68.94% respectively. Keywords: LVCSR, Cantonese speech recognition, acoustic modeling, language modeling 1 INTRODUCTION Cantonese is one of the major Chinese dialects spoken by tens of millions of people in Hong Kong, Southern China as well as many overseas Chinese communities. With the great advancement of computer and information technology, there is an ever-increasing demand of largevocabulary con...
Yiu Wing Wong, Ka-Fai Chow, Wai H. Lau, Wai Kit Lo, Tan Lee, Pak-Chung Ching
EUROSPEECH6
1999 Cantonese syllable recognition using neural networks
abstract
This work describes a novel neural network based speech recognition system for isolated Cantonese syllables. Since Cantonese is a monosyllabic and tonal language, the recognition system is composed of two major components, namely the tone recognizer and the base syllable recognizer. The tone recognizer adopts the architecture of multilayer perceptron in which each output neuron represents a particular tone. The base syllable recognizer consists of a large number of independently trained recurrent networks, each representing a designated Cantonese syllable. An integrated recognition algorithm is developed to give the ultimate recognition results based on N-best outputs of the two subrecognizers. To demonstrate the effectiveness of the proposed methods, a speaker-dependent recognition system has been built with the vocabulary expanding progressively from 10 syllables to 200 syllables. In the case of 200 syllables, a top-1 recognition accuracy of 81.8% has been attained whilst the top-3 accuracy is 95.28.
Tan Lee, Pak-Chung Ching
IEEE Trans. Speech Audio Process.2
1998 A hybrid approach to synthesize high quality Cantonese speech
abstract
Synthesizing high quality speech necessitates an intelligent modification algorithm to adjust the important prosodic features of the pre-stored speech units to meet the desired output requirements, such as smoothness, naturalness and pleasantness. The time domain pitch-synchronous overlap and add (TD-PSOLA) scheme is a simple but effective method of varying the pitch and time-scaling of speech and it can produce high quality synthetic output. However, when the prosodic pattern requires a drastic modification in the spectral content of the stored units, TD-PSOLA often generates speech with reverberant sound. This paper develops a new hybrid synthesis method based on TD-PSOLA and shape-invariant sinusoidal technique to alleviate the problem of reverberation. It is particularly useful for the generation of Cantonese speech, since it can cope with the rapidly changing of the pitch profile of Cantonese, which is a mono-syllabic and tonal language. The proposed method has been applied to construct a Cantonese synthesizer which is shown to be capable of producing high quality Cantonese speech without reverberation.
Chu Min, Pak-Chung Ching
ICASSP2
1998 Articulatory synthesis of formant targeted sounds with parameters derived from the inverse solution of speech production
abstract
A new approach to produce high fidelity speech sounds by applying both the inverse solution of speech production and the pitch-synchronous articulatory synthesis technique is presented. Given a formant trace target, the dynamic vocal-tract area function together with the time variant VT length are estimated using an inverse solution of speech production. The improved Kelly-Lochbaum (1962) filter of the synthesizer, with multi-rate system sampling and dynamic scattering wave adjustment, is employed to deal with the variable VT length and acoustic continuity. The synthesizer is controlled by the estimated VT area function. A distinguished feature of this method is that artificially specified formant traces can be precisely obtained. Experimental results show that the formant targets can be precisely matched by the synthetic sounds. A potential application of this method for text-to-speech conversion is discussed.
Zhenli Yu, Pak-Chung Ching
ICASSP2
1998 Isolated word recognition using modular recurrent neural networks
Tan Lee, Pak-Chung Ching, Lai-Wan Chan
Pattern Recognit.2
1997 Development of a large vocabulary speech database for Cantonese
abstract
This paper describes work on developing a large vocabulary speech database for Cantonese. As a major Chinese dialect, Cantonese is spoken by tens of millions of people in Southern China and Hong Kong. It is very different from Mandarin or Putonghua in phonology, phonetics, vocabulary and grammatical structure. A speech database specially designed for Cantonese is urgently needed for the design, implementation and performance evaluation of various speech recognition systems. The proposed database contains a large number of speech utterances which include isolated syllables, polysyllabic words and phonetically rich sentences. It covers most of the intra-syllable and inter-syllable acoustic variations.
Pak-Chung Ching, Ka-Fai Chow, Tan Lee, Alfred Ying Pang Ng, Lai-Wan Chan
ICASSP1
1997 A neural network based speech recognition system for isolated Cantonese syllables
abstract
This paper describes a novel design of neural network based speech recognition system for isolated Cantonese syllables. Since Cantonese is a monosyllabic and tonal language, the recognition system consists of a tone recognizer and a base syllable recognizer. The tone recognizer adopts the architecture of a multi-layer perceptron in which each output neuron represents a particular tone. The syllable recognizer contains a large number of independently trained recurrent networks, each representing a designated Cantonese syllable. Such a modular structure provides greater flexibility to expand the system vocabulary progressively by adding new syllable models. To demonstrate the effectiveness of the proposed method, a speaker-dependent recognition system has been built with the vocabulary growing from 40 syllables to 200 syllables. In the case of 200 syllables, a top-1 recognition accuracy of 81.8% has been attained and the top-3 accuracy is 95.2%.
Tan Lee, Pak-Chung Ching
ICASSP2
1997 Improvement of TDOA measurement using wavelet denoising with a novel thresholding technique
abstract
Wavelet denoising is applied in time delay estimation between signals received at two spatially separated sensors in the presence of noise. Prior to cross correlation, each of the sensor outputs is denoised according to a novel thresholding rule in order to increase the input signal-to-noise ratio. Unlike conventional generalized cross correlators (GCCs), it does not require spectral estimation of the source signal and the corrupting noises which may introduce large delay variance. It is proved that the delay estimate provided by the proposed method is globally convergent to the true value with a high probability. Computer simulations illustrate that the technique outperforms other GCCs for different SNRs when the sampling rate is sufficiently high.
Shi-Quan Wu, Hing-Cheung So, Pak-Chung Ching
ICASSP3
1997 Automatic recognition of continuous Cantonese speech with very large vocabulary
Alfred Ying Pang Ng, Lai-Wan Chan, Pak-Chung Ching
EUROSPEECH3
1997 Geometrically and acoustically optimized codebook for unique mapping from formants to vocal-tract shape
Zhenli Yu, Pak-Chung Ching
EUROSPEECH2
1996 Determination of vocal-tract shapes from formant frequencies based on perturbation theory and interpolation method
abstract
Band-limited Fourier expansion, containing both odd and even components, has proven to be useful for representing human vocal-tract (VT) area function in the inverse problem of speech production. The Fourier coefficients can be estimated from the pole/zero frequencies of the VT transfer function. However, the difficulties encounted in determining the VT length and extracting zero frequencies from speech signal make this approach almost impossible to implement. This paper proposes an interpolation method that allows zero frequencies and the VT length to be derived accurately, and from which the VT area function can be determined subsequently based on perturbation theory. A root-cell code book is used to facilitate dynamic vowel-to-vowel (VV) transition. A computationally efficient implementation of the dynamic process for VV transition is presented. Simulation results are given.
Zhenli Yu, Pak-Chung Ching
ICASSP2
1996 On improving discrimination capability of an RNN based recognizer
Tan Lee, Pak-Chung Ching
ICSLP2
1996 Phone-based speech synthesis with neural network and articulatory control
Wai Kit Lo, Pak-Chung Ching
ICSLP2
1996 Characterization of UHF radio propagation channel in curved tunnels
abstract
We propose an imperfect waveguide model to represent empty curved tunnels. We use the field configuration in an optical dielectric guide together with the surface impedance boundary condition for the walls to simplify the complexity of derivation of the propagation modal equations. The curvature of the tunnels is assumed to be gentle. Our results show that increased propagation loss is almost linearly proportional to frequency and inversely proportional to the radius of curvature.
Y. P. Zhang, Y. Hwang, Pak-Chung Ching
PIMRC3
1995 Recurrent neural networks for speech modeling and speech recognition
abstract
Describes a new method of utilizing recurrent neural networks (RNNs) for speech modeling and speech recognition. For each particular speech unit, a fully connected recurrent neural network is built such that the static and dynamic speech characteristics are represented simultaneously by a specific temporal pattern of neuron activation states. By using the temporal RNN output, an input utterance can be represented as a number of stationary speech segments, which may be related to the basic phonetic components of the speech unit. An efficient self-supervised training algorithm has been developed for the RNN speech model. The segmentation for input utterances and the statistical modeling for individual phonetic segments are performed interactively in this training process. Some experimental results are used to demonstrate how the proposed RNN speech model can be used effectively for automatic recognition of isolated speech utterances.
Tan Lee, Pak-Chung Ching, Lai-Wan Chan
ICASSP2
1995 An improvement to the explicit time delay estimator
abstract
The explicit time delay estimator (ETDE) provides an efficient way to estimate the time difference of arrival between signals received at two separated sensors. However, the algorithm is biased for finite filter length and the delay bias increases when the signal-to-noise ratio (SNR) or the number of filter taps decreases. In this paper, we add an adaptive gain control to the ETDE to decouple the effect of changes in the SNR during adaptation. As a result, a smaller delay variance and an unbiased delay estimate for a wide range of filter lengths can be attained. Computer simulations are presented to validate the theoretical derivations of the proposed estimator for static and linearly time-varying delays under both stationary and nonstationary signal/noise power environments.
Hing-Cheung So, Pak-Chung Ching, Yiu Tong Chan
ICASSP2
1995 An RNN based speech recognition system with discriminative training
abstract
In our previous work #1#, a novel method of utilizing a set of fully connected recurrent neural networks #RNNs# for speech modeling has been proposed. Despite the e#ectiveness of the RNN model in characterizing individual speech units, the system performs less satisfactorily for speech recognition due to poor discrimination between models. In this paper, an e#cient discriminative training procedure is developed for the RNN based recognition system. By using discriminative training, each RNN speech model is adjusted to reduce its distance from the designated speech unit while increase distances from the others. In addition, a duration-screening process is introduced to enhance the discriminating power of the recognition system. Speaker-dependent recognition experiments have been carried out for 1# 11 isolated Cantonese digits, 2# 58 very confusing Cantonese CV syllables, and 3# 20 English isolated words. The recognition rates attained are 90.9#, 86.7# and 93.5# respectively. I. Int...
Tan Lee, Pak-Chung Ching, Lai-Wan Chan
EUROSPEECH2
1995 Automatic recognition of Cantonese lexical tones in connected speech by multi-layer perceptron
Alfred Ying Pang Ng, Pak-Chung Ching, Lai-Wan Chan
EUROSPEECH2
1995 A Unified Approach to Split Structure Adaptive Filtering
Pak-Chung Ching, K. F. Wan
ISCAS1
1995 Blind Estimation Using Higher-Order Cumulants
W. K. Lai, Pak-Chung Ching
ISCAS2
1995 Split filter structures for LMS adaptive filtering
K. C. Ho 0001, Pak-Chung Ching
Signal Process.2
1995 Tone recognition of isolated Cantonese syllables
abstract
Tone identification is essential for the recognition of the Chinese language, specifically far Cantonese which is well known for being very rich in tones. The paper presents an efficient method for tone recognition of isolated Cantonese syllables. Suprasegmental feature parameters are extracted from the voiced portion of a monosyllabic utterance and a three-layer feedforward neural network is used to classify these feature vectors. Using a phonologically complete vocabulary of 234 distinct syllables, the recognition accuracy for single-speaker and multispeaker is given by 89.0% and 87.6% respectively.>
Tan Lee, Pak-Chung Ching, Lai-Wan Chan, Y. H. Cheng, Brian Kan-Wing Mak
IEEE Trans. Speech Audio Process.2
1994 On the optimality of convergence behaviour for transform-domain split-path adaptive filter
abstract
A fast LMS adaptive filtering algorithm in transform-domain is developed. The algorithm is applied to a structure which decomposed an adaptive FIR filter into a parallel connection of two subfilters, one with a symmetric property and the other with an antisymmetric property. A detail analysis on the optimality of the convergence behaviour for this transform-domain split-path adaptive filter is presented. Simulation results show that the proposed algorithm has superior convergence performance while the increase in computation is only modest.>
K. F. Wan, Pak-Chung Ching
ICASSP (3)2
1994 A Fast Convergence Median LMS Algorithm
abstract
In this paper, a split-path median LMS algorithm is proposed. By reducing the eigenvalue spread of the input correlation matrix, the algorithm not only improves the convergence performance but also enhances the tracking behaviour of abrupt signal edges. The simulation results for line enhancement are presented and comparisons are made with LMS and median LMS algorithms.>
K. F. Wan, Pak-Chung Ching
ISCAS2
1993 A novel constrained algorithm for delay estimation in the presence of multipath transmissions
Hing-Cheung So, Pak-Chung Ching, K. C. Ho 0001, Yiu Tong Chan
ICASSP (1)2
1991 Adaptive time delay estimation in noisy environments
abstract
A model to improve the performance of an adaptive tracker for delay estimation in noisy environments is proposed. The system consists of an adaptive filter and a time varying gain control. While the adaptive filter inserts an appropriate time shift to the incoming signal, the gain control will provide a proper scaling to the filter output respective to the additive noise. Both the filter coefficients and the variable gain are adjusted simultaneously by minimizing the output mean-square error according to Widrow's least mean square (LMS) algorithm. This arrangement can decouple the adaptation of time shift and noise statistics, which in turn will give rise to a better convergence characteristics of the delay. Simulation results are included to demonstrate its capability in tracking time shift accurately under noisy environments.>
K. C. Ho 0001, Pak-Chung Ching, Yiu Tong Chan
ICASSP2
1990 Convergence speed up in adaptive time delay estimation
abstract
An adaptive configuration is introduced for time delay estimation whereby two adaptive filters, one giving a time delay and another a time advance of equal amount, are adapted simultaneously to minimize the errors. This configuration gives a fourfold increase in convergence speed compared with the conventional single adaptive filter. Proofs for convergence and speed up are given, together with verification from simulation results.>
K. C. Ho 0001, Pak-Chung Ching, Yiu Tong Chan
ICASSP2
1989 Non-stationary time delay estimation with a multipath
abstract
A description is given of a constrained adaptive scheme for time-delay estimation in the presence of multipath propagation. Two configurations are proposed. In either case, the role of the adaptive filters is to provide appropriate time shifts to the input signals. For a filter to function as a pure time shifter, its coefficients must take on the values of the samples of a sinc function. Applying this constraint on the coefficients, which are adapted by the LMS algorithm, reduces the amount of computation and speeds up the convergence rate considerably. since the number of adaptive coefficients is typically of the order of 30 or more, convergence time would be excessive without constraints. The effectiveness of the scheme is demonstrated by simulation results, which show its ability to track time-varying parameters accurately.>
Yiu Tong Chan, Pak-Chung Ching
ICASSP2