VLDB 2026 Research / reviewers in the wild / expert
Linhao Dong
dblp:39/10800
· DBLP profile ↗
31ranked-venue papers
10as first author
7since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 3 since 2021Computer networks · 5 · 3 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASRabstractMulti-talker automatic speech recognition plays a crucial role in scenarios involving multi-party interactions, such as meetings and conversations. Due to its inherent complexity, this task has been receiving increasing attention. Notably, the serialized output training (SOT) stands out among various approaches because of its simplistic architecture and exceptional performance. However, the frequent speaker changes in token-level SOT (t-SOT) present challenges for the autoregressive decoder in effectively utilizing context to predict output sequences. To address this issue, we introduce a masked t-SOT label, which serves as the cornerstone of an auxiliary training loss. Additionally, we utilize a speaker similarity matrix to refine the self-attention mechanism of the decoder. This strategic adjustment enhances contextual relationships within the same speaker’s tokens while minimizing interactions between different speakers’ tokens. We denote our method as speaker-aware SOT (SA-SOT). Experiments on the Librispeech datasets demonstrate that our SA-SOT obtains a relative cpWER reduction ranging from 12.75% to 22.03% on the multi-talker test sets. Furthermore, with more extensive training, our method achieves an impressive cpWER of 3.41%, establishing a new state-of-the-art result on the LibrispeechMix dataset. Zhiyun Fan, Linhao Dong, Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001 |
ICASSP | 2 |
| 2023 | Language-specific Boundary Learning for Improving Mandarin-English Code-switching Speech Recognition
Zhiyun Fan, Linhao Dong, Chen Shen 0011, Zhenlin Liang, Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001 |
INTERSPEECH | 2 |
| 2022 | Optimizing the Post-disaster Resource Allocation with Q-Learning: Demonstration of 2021 China Flood
Linhao Dong, Yanbing Bai, Qingsong Xu 0001, Erick Mas |
DEXA (2) | 1 |
| 2022 | Improving End-to-End Contextual Speech Recognition with Fine-Grained Contextual Knowledge SelectionabstractNowadays, most methods for end-to-end contextual speech recognition bias the recognition process towards contextual knowledge. Since all-neural contextual biasing methods rely on phrase-level contextual modeling and attention-based relevance modeling, they may suffer from the confusion between similar context-specific phrases, which hurts predictions at the token level. In this work, we focus on mitigating confusion problems with fine-grained contextual knowledge selection (FineCoS). In FineCoS, we introduce fine-grained knowledge to reduce the uncertainty of token predictions. Specifically, we first apply phrase selection to narrow the range of phrase candidates, and then conduct token attention on the tokens in the selected phrase candidates. Moreover, we re-normalize the attention weights of most relevant phrases in inference to obtain more focused phrase-level contextual representations, and inject position information to help model better discriminate phrases or tokens. On LibriSpeech and an in-house 160,000-hour dataset, we explore the proposed methods based on an all-neural biasing method, collaborative decoding (ColDec). The proposed methods further bring at most 6.1% relative word error rate reduction on LibriSpeech and 16.4% relative character error rate reduction on the in-house dataset. Minglun Han, Linhao Dong, Zhenlin Liang, Zejun Ma 0001, Bo Xu 0002 |
ICASSP | 2 |
| 2022 | Token-level Speaker Change Detection Using Speaker Difference and Speech Content via Continuous Integrate-and-fire
Zhiyun Fan, Zhenlin Liang, Linhao Dong, Jun Zhang 0066, Zejun Ma 0001, Bo Xu 0002 |
INTERSPEECH | 3 |
| 2022 | Sequence-Level Speaker Change Detection With Difference-Based Continuous Integrate-and-FireabstractSpeaker change detection is an important task in multi-party interactions such as meetings and conversations. In this paper, we address the speaker change detection task from the perspective of sequence transduction. Specifically, we propose a novel encoder-decoder framework that directly converts the input feature sequence to the speaker identity sequence. The difference-based continuous integrate-and-fire mechanism is designed to support this framework. It detects speaker changes by integrating the speaker difference between the encoder outputs frame-by-frame and transfers encoder outputs to segment-level speaker embeddings according to the detected speaker changes. The whole framework is supervised by the speaker identity sequence, a weaker label than the precise speaker change points. The experiments on the AMI and DIHARD-I corpora show that our sequence-level method consistently outperforms a strong frame-level baseline that uses the precise speaker change labels. Zhiyun Fan, Linhao Dong, Zejun Ma 0001, Bo Xu 0002 |
IEEE Signal Process. Lett. | 2 |
| 2021 | Cif-Based Collaborative Decoding for End-to-End Contextual Speech RecognitionabstractEnd-to-end (E2E) models have achieved promising results on multiple speech recognition benchmarks, and shown the potential to become the mainstream. However, the unified structure and the E2E training hamper injecting context information into them for contextual biasing. Though contextual LAS (CLAS) gives an excellent all-neural solution, the degree of biasing to given contextual information is not explicitly controllable. In this paper, we focus on incorporating contextual information into the continuous integrate-and-fire (CIF) based model that supports contextual biasing in a more controllable fashion. Specifically, an extra context processing network is introduced to extract contextual embeddings, integrate acoustically relevant contextual information and decode the contextual output distribution, thus forming a collaborative decoding with the decoder of the CIF-based model. Evaluated on the named entity rich evaluation sets of HKUST/AISHELL-2, our method brings relative character error rate (CER) reduction of 8.83%/21.13% and relative named entity character error rate (NE-CER) reduction of 40.14%/51.50% when compared with a strong baseline. Besides, it keeps the performance on original evaluation set without degradation. Minglun Han, Linhao Dong, Bo Xu 0002 |
ICASSP | 2 |
| 2020 | CIF: Continuous Integrate-And-Fire for End-To-End Speech RecognitionabstractIn this paper, we propose a novel soft and monotonic alignment mechanism used for sequence transduction. It is inspired by the integrate-and-fire model in spiking neural networks and employed in the encoder-decoder framework consists of continuous functions, thus being named as: Continuous Integrate-and-Fire (CIF). Applied to the ASR task, CIF not only shows a concise calculation, but also supports online recognition and acoustic boundary positioning, thus suitable for various ASR scenarios. Several support strategies are also proposed to alleviate the unique problems of CIF-based model. With the joint action of these methods, the CIF-based model shows competitive performance. Notably, it achieves a word error rate (WER) of 2.86% on the test-clean of Librispeech and creates new state-of-the-art result on Mandarin telephone ASR benchmark. Linhao Dong, Bo Xu 0002 |
ICASSP | 1 |
| 2019 | Self-attention Aligner: A Latency-control End-to-end Model for ASR Using Self-attention Network and Chunk-hoppingabstractSelf-attention network, an attention-based feedforward neural network, has recently shown the potential to replace recurrent neural networks (RNNs) in a variety of NLP tasks. However, it is not clear if the self-attention network could be a good alternative of RNNs in automatic speech recognition (ASR), which processes the longer speech sequences and may have online recognition requirements. In this paper, we present a RNN-free end-to-end model: self-attention aligner (SAA), which applies the self-attention networks to a simplified recurrent neural aligner (RNA) framework. We also propose a chunk-hopping mechanism, which enables the SAA model to encode on segmented frame chunks one after another to support online recognition. Experiments on two Mandarin ASR datasets show the replacement of RNNs by the self-attention networks yields a 8.4%-10.2% relative character error rate (CER) reduction. In addition, the chunk-hopping mechanism allows the SAA to have only a 2.5% relative CER degradation with a 320ms latency. After jointly training with a self-attention network language model, our SAA model obtains further error rate reduction on multiple datasets. Especially, it achieves 24.12% CER on the Mandarin ASR benchmark (HKUST), exceeding the best end-to-end model by over 2% absolute CER. Linhao Dong, Feng Wang 0023, Bo Xu 0002 |
ICASSP | 1 |
| 2019 | Boosting Character-Based Chinese Speech Synthesis via Multi-Task Learning and Dictionary Tutoring
Yuxiang Zou, Linhao Dong, Bo Xu 0002 |
INTERSPEECH | 2 |
| 2018 | Speech-Transformer: A No-Recurrence Sequence-to-Sequence Model for Speech RecognitionabstractRecurrent sequence-to-sequence models using encoder-decoder architecture have made great progress in speech recognition task. However, they suffer from the drawback of slow training speed because the internal recurrence limits the training parallelization. In this paper, we present the Speech-Transformer, a no-recurrence sequence-to-sequence model entirely relies on attention mechanisms to learn the positional dependencies, which can be trained faster with more efficiency. We also propose a 2D-Attention mechanism, which can jointly attend to the time and frequency axes of the 2-dimensional speech inputs, thus providing more expressive representations for the Speech-Transformer. Evaluated on the Wall Street Journal (WSJ) speech recognition dataset, our best model achieves competitive word error rate (WER) of 10.9%, while the whole training process only takes 1.2 days on 1 GPU, significantly faster than the published results of recurrent sequence-to-sequence models. Linhao Dong, Bo Xu 0002 |
ICASSP | 1 |
| 2018 | A Comparison of Modeling Units in Sequence-to-Sequence Speech Recognition with the Transformer on Mandarin Chinese
Linhao Dong, Bo Xu 0002 |
ICONIP (5) | 2 |
| 2018 | Syllable-Based Acoustic Modeling with CTC for Multi-Scenarios Mandarin speech recognitionabstractWith the improvement of speech recognition, voice products are gradually applied to every scene of life. The existing approaches to handle various scenarios are often to build many different acoustic models using scenario-dependent data only, with each for a special scene. The obvious weakness of these approaches is that it seriously hampers the large-scale application and maintenance of voice products. To address this issue, acoustic modeling based on context-independent syllables optimized with CTC loss is presented for multiple scenarios of Mandarin speech recognition. On the one hand, context-independent modeling overcomes the shortcomings of context-dependent modeling overfitting a particular scene. Also, it sidesteps decision trees used in context-dependent modeling so that there is no need to consider the building of decision tree and whether to start training again in a real application. On the other hand, choosing longer-length syllable acoustic units can effectively preserve the co-articulation effect that context-dependent phone can model. Also, syllables in the Chinese language have its inherent advantages, as its number is fixed and it is trainable, effective generalization and better robustness. This paper also explores the differences between wideband and narrowband data caused by the front-end signal acquisition block, and proposes a unified training method based on the use of VGG in the bottom layer, and introduces layer normalization. The experimental results demonstrate that the proposed syllable-based CTC acoustic model for multiple scenarios can achieve more than 15% and 7% relatively improvement for mobile phone data and telephone data separately compare with scenarios-dependent modeling. Linhao Dong, Bo Xu 0002 |
IJCNN | 2 |
| 2018 | Extending Recurrent Neural Aligner for Streaming End-to-End Speech Recognition in MandarinabstractEnd-to-end models have been showing superiority in Automatic Speech Recognition (ASR).At the same time, the capacity of streaming recognition has become a growing requirement for end-to-end models.Following these trends, an encoder-decoder recurrent neural network called Recurrent Neural Aligner (RNA) has been freshly proposed and shown its competitiveness on two English ASR tasks.However, it is not clear if RNA can be further improved and applied to other spoken language.In this work, we explore the applicability of RNA in Mandarin Chinese and present four effective extensions: In the encoder, we redesign the temporal downsampling and introduce a powerful convolutional structure.In the decoder, we utilize a regularizer to smooth the output distribution and conduct joint training with a language model.On two Mandarin Chinese conversational telephone speech recognition (MTS) datasets, our Extended-RNA obtains promising performance.Particularly, it achieves 27.7% character error rate (CER), which is superior to current state-of-the-art result on the popular HKUST task. Linhao Dong, Wei Chen 0048, Bo Xu 0002 |
INTERSPEECH | 1 |
| 2018 | Syllable-Based Sequence-to-Sequence Speech Recognition with the Transformer in Mandarin ChineseabstractSequence-to-sequence attention-based models have recently shown very promising results on automatic speech recognition (ASR) tasks, which integrate an acoustic, pronunciation and language model into a single neural network.In these models, the Transformer, a new sequence-to-sequence attention-based model relying entirely on self-attention without using RNNs or convolutions, achieves a new single-model state-of-the-art BLEU on neural machine translation (NMT) tasks.Since the outstanding performance of the Transformer, we extend it to speech and concentrate on it as the basic architecture of sequence-to-sequence attention-based model on Mandarin Chinese ASR tasks.Furthermore, we investigate a comparison between syllable based model and context-independent phoneme (CI-phoneme) based model with the Transformer in Mandarin Chinese.Additionally, a greedy cascading decoder with the Transformer is proposed for mapping CI-phoneme sequences and syllable sequences into word sequences.Experiments on HKUST datasets demonstrate that syllable based model with the Transformer performs better than CI-phoneme based counterpart, and achieves a character error rate (CER) of 28.77%, which is competitive to the state-of-the-art CER of 28.0% by the joint CTC-attention based encoder-decoder network. Linhao Dong, Bo Xu 0002 |
INTERSPEECH | 2 |
| 2017 | A Joint Scheduling and Content Caching Scheme for Energy Harvesting Access Points with MulticastabstractIn this work, we investigate a system where users are served by an access point that is equipped with energy harvesting and caching mechanism. Focusing on the design of an efficient content delivery scheduling, we propose a joint scheduling and caching scheme. The scheduling problem is formulated as a Markov decision process and solved by an on-line learning algorithm. To deal with large state space, we apply the linear approximation method to the state-action value functions, which significantly reduces the memory space for storing the function values. In addition, the preference learning is incorporated to speed up the convergence when dealing with the requests from users that have obvious content preferences. Simulation results confirm that the proposed scheme outperforms the baseline scheme in terms of convergence and system throughput, especially when the personal preference is concentrated to one or two contents. Linhao Dong, Dusit Niyato, Dong In Kim 0001, Dinh Thai Hoang |
GLOBECOM | 1 |
| 2017 | Wireless Information and Power Transfer: Spectral Efficiency Optimization for Asymmetric Full-Duplex Relay SystemsabstractTo address the problem of unbalanced received signal-to-interference-and-noise ratio (SINR) at relay and destination nodes in wireless power transfer (WPT)- supported relay system, we propose a novel asymmetric full-duplex (FD) decode-and-forward (DF) WPT relay strategy, where the transmission time slots are not necessarily identical. By introducing asymmetric time slots, higher degree of freedom is obtained than the conventional symmetric WPT relay system. Furthermore, based on the asymmetric strategy, we develop a spectral efficiency (SE)- oriented resource allocation algorithm by jointly designing time slots, transmission power at source and relay. Simulation results show that the proposed asymmetric system demonstrates higher SE than the symmetric WPT-powered FD and the time- switching based FD relay systems. Besides, more energy can be harvested at the relay node by the proposed system benefiting from the enhanced degree of freedom, showing its applicability in WPT-powered relay systems. Zhongxiang Wei, Sumei Sun, Xu Zhu 0001, Yi Huang 0001, Linhao Dong, Dong In Kim 0001 |
VTC Spring | 5 |
| 2017 | Online Bayesian Learning for Remote-Sensing Imagery CompressionabstractThis work investigates a statistical technique for high performance remote-sensing imagery compression. By exploiting existing remote-sensing data sets, useful structural and texture prior information can be learned. The main methodologies are Bayesian dictionary learning and stochastic approximation. A Bayesian network simulating the generation mechanism of remote- sensing images is modelled. The whole compression scheme is established. And the corresponding inference algorithm using Gibbs sampling is given, where the inference is realized in an online way. The performance of the proposed compressing scheme is evaluated over a high-resolution remote-sensing image data set captured by TH-1 series satellites. Experiment results have shown that our compression scheme outperforms JPEG-2000 by 3dB on average with same bits-per-pixel performance, and that Bayesian learning can provide a dictionary with high expressiveness for remote-sensing images. In addition, with online learning skills our proposed compression scheme can scale up to very large-scale training data. Zizhuo Zhang, Shaoyang Li, Xiaoming Tao 0001, Linhao Dong, Jianhua Lu |
VTC Spring | 4 |
| 2017 | Bayesian Hyperspectral and Multispectral Image Fusions via Double Matrix FactorizationabstractThis paper focuses on fusing hyperspectral and multispectral images with an unknown arbitrary point spread function (PSF). Instead of obtaining the fused image based on the estimation of the PSF, a novel model is proposed without intervention of the PSF under Bayesian framework, in which the fused image is decomposed into double subspace-constrained matrix-factorization-based components and residuals. On the basis of the model, the fusion problem is cast as a minimum mean square error estimator of three factor matrices. Then, to approximate the posterior distribution of the unknowns efficiently, an estimation approach is developed based on variational Bayesian inference. Different from most previous works, the PSF is not required in the proposed model and is not pre-assumed to be spatially invariant. Hence, the proposed approach is not related to the estimation errors of the PSF and has potential computational benefits when extended to spatially variant imaging system. Moreover, model parameters in our approach are less dependent on the input data sets and most of them can be learned automatically without manual intervention. Exhaustive experiments on three data sets verify that our approach shows excellent performance and more robustness to the noise with acceptable computational complexity, compared with other state-of-the-art methods. Baihong Lin, Xiaoming Tao 0001, Mai Xu, Linhao Dong, Jianhua Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2016 | Variational Bayesian image fusion based on combined sparse representationsabstractHyper-spectral image fusion has been a hot topic in medical imaging and remote sensing. This paper proposes a Bayesian fusion model which combines the panchromatic (PAN) image and the low spatial resolution hyper-spectral (HS) image under the same framework. Sparsity constraint is introduced as double "spike-and-slab" priors, and anisotropic Gaussian noise is adopted for accuracy. To achieve reduction in computational complexity, we turn the anisotropic Gaussian distribution into isotropic one with modified linear transformation and propose a variational Bayesian expectation maximization (EM) algorithm to calculate the result. Experiment results show that our solution can achieve comparable performance in pan-sharpening to other state-of-art algorithms while largely reducing the computational complexity. Baihong Lin, Xiaoming Tao 0001, Shaoyang Li, Linhao Dong, Jianhua Lu |
ICASSP | 4 |
| 2016 | Variational EM approach for high resolution hyper-spectral imaging based on probabilistic matrix factorizationabstractHigh resolution hyper-spectral imaging works as a scheme to obtain images with high spatial and spectral resolutions by merging a low spatial resolution hyper-spectral image (HSI) with a high spatial resolution multi-spectral image (MSI). In this paper, we propose a novel method based on probabilistic matrix factorization under Bayesian framework: First, Gaussian priors, as observations' distributions, are given upon two HSI-MSI-pair-based images, in which two variances share the same hyper-parameter to ensure fair and effective constraints on two observations. Second, to avoid the manual tuning process and learn a better setting automatically, hyper-priors are adopted for all hyper-parameters. To that end, a variational expectation-maximization (EM) approach is devised to figure out the result expectation for its simplicity and effectiveness. Exhaustive experiments of two different cases prove that our algorithm outperforms many state-of-the-art methods. Baihong Lin, Xiaoming Tao 0001, Linhao Dong, Jianhua Lu |
ICIP | 3 |
| 2016 | Mouse calibration aided real-time gaze estimation based on boost Gaussian Bayesian learningabstractIn this paper, we propose a novel gaze estimation method to evaluate the attention span of users upon on-screen content via a single webcam. Our method is based on supervised descent method for eye region of interest (ROI) extraction. Then, boost Gaussian Bayesian regressors are applied to learn a robust mapping from the input eye ROI to gaze coordinates. To get enough training samples, we implant our scheme as a plug-in into web browsers for data collection from users without bothering. To improve accuracy, we also introduce mouse click to help train the regressors. Experiment results show that our method outperforms the existing method and can provide gaze estimation data for user behaviour analysis in real-time implementation. Nanyang Ye 0001, Xiaoming Tao 0001, Linhao Dong, Ning Ge 0001 |
ICIP | 3 |
| 2016 | The THU multi-view face database for videoconferences and baseline evaluations
Xiaoming Tao 0001, Linhao Dong, Yang Li 0005, Jianhua Lu |
Neurocomputing | 2 |
| 2016 | HEMS: Hierarchical Exemplar-Based Matching-Synthesis for Object-Aware Image ReconstructionabstractMotivated by the attention on salient objects, conventional region-of-interest (ROI)-based image coding approaches attempt to assign more bits to ROIs and fewer bits to other regions. Thus, the perceptual quality of salient object regions is improved by sacrificing the quality of non-ROI regions with unpleasant artifacts. To address this issue, we concentrate on the efficient compression of object-centered images by encoding salient objects and background features separately. To fully recover the object and background, we propose a hierarchical exemplar-based matching-synthesis (HEMS) approach to reconstruct the image from exemplars. In the proposed framework, once the salient object regions are encoded, only the quantized color features and local descriptors of the background are kept, achieving bit-rate reduction. To make it possible and practical to reconstruct background regions, the hierarchical framework is designed in three layers, including relevant image search, patch candidates matching, and distortion optimized image synthesis. In the hierarchical framework, firstly, image search from an external database returns relevant images, limiting the search space to a feasible number of patch candidates. Secondly, patches are matched by color features to select the appropriate candidates. Finally, the distortion optimized image synthesis further makes it possible to automatically choose the most suitable texture sample, and seamlessly reconstruct the image. Compared to the conventional ROI-based image coding schemes, the proposed approach can achieve better visual quality on both ROI and background regions. Yipeng Sun, Xiaoming Tao 0001, Yang Li 0005, Linhao Dong, Jianhua Lu |
IEEE Trans. Multim. | 4 |
| 2015 | Pilot sequence design for multi-cell distributed MIMO systems with large-scale CSIabstractWhen non-orthogonal pilots are used in multi-cell systems, channel estimation would be corrupted by inter-cell interference. To tackle this problem, the design of pilot sequences for multi-cell distributed multiple-input multiple-output (MIMO) systems is addressed in this paper. We explore this issue by introducing discriminatory treatment of different channel parameters. Generally, the large-scale channel state information (CSI) is predictable and could be regarded as priori information in pilot design, due to its slowly-varying characteristics. In particular, by assuming the large-scale CSI is known a priori, we derive a lower bound on the achievable sum rate in the downlink with a linear detector at the mobile terminal (MT), taking both intercell interference and channel estimation error into account. The problem of pilot sequence design is first formulated, of which the target is to maximize the lower bound of the achievable sum rate with a total pilot power constraint for each cell. Afterwards, we solve the problem by introducing the iterative concave-convex procedure (CCCP) with the demonstration of its convergence. Simulation results illustrate the validity and superiority of the proposed pilot design scheme. Wei Feng 0001, Linhao Dong, Ning Ge 0001 |
ICC | 3 |
| 2015 | The THU multi-view face database for videoconferencesabstractIn this paper, we present a face video database that contains 31,500 videos of 100 individual volunteers. The primary purpose of building this database is to serve as a standardized test video sequences for any research related to video-conferences. Each of the volunteers was filmed by 9 groups of synchronized webcams under 7 illumination conditions, and was requested to complete a series designated actions. Thus, face variations on lip shape, occlusion, illumination, pose, and expression are presented in each video clip. Compared to the existing databases, THU face database provides multi-view video sequences with strict temporal synchronization, enabling evaluations on gaze-correction methods. Besides, based on our database, three well-known methods were tested, demonstrating the numerical performances under different circumstances. Free samples of this database can be downloaded at www.facedbv.com. Linhao Dong, Xiaoming Tao 0001, Yang Li 0005, Jichuan Lu, Zizhuo Zhang, Jingwen Cheng, Jianhua Lu |
ICIP | 1 |
| 2015 | Efficient Multi-Cell Clustering for Coordinated Multi-Point Transmission with Blossom Tree AlgorithmabstractCoordinated multi-point(CoMP) transmission clustering schemes could provide significant gains of system performance, such as throughput and cell- edge user data rates. Due to limitations of the backhaul communication and signal processing capability of base stations(BSs), the intrinsic problem of CoMP is that the selection of which BSs shall cooperate as only a few of BSs can be grouped in a cluster. However, approximating the theoretical performance bound of this clustering problem in CoMP at present is seldom discussed due to its inherent combinatorial complexity. In this paper, a novel efficient multi-cell clustering scheme based on blossom tree algorithm is proposed for cellular networks, incorporating CoMP with two cells in each cluster. With blossom tree algorithm, the proposed scheme can find out the optimal clustering strategy and help the CoMP transmission reach its theoretical performance bound on data rate in real-time computing(milliseconds in MATLAB simulation for one clustering). The simulation results show that our proposed method outperforms the existing dynamic greedy method in terms of cell edge users' average achievable data rate. Besides, it can also maintain high performance when extended to larger clusters in that with 4-cell clustering, the proposed method can reach 23.8% higher data rates than dynamic greedy method. Nanyang Ye 0001, Linhao Dong, Xiaoming Tao 0001, Ning Ge 0001 |
VTC Fall | 2 |
| 2015 | Full-Duplex Versus Half-Duplex Amplify-and-Forward Relaying: Which is More Energy Efficient in 60-GHz Dual-Hop Indoor Wireless Systems?abstractWe provide a comprehensive energy efficiency (EE) analysis of the full-duplex (FD) and half-duplex (HD) amplify-and-forward (AF) relay-assisted 60-GHz dual-hop indoor wireless systems, aiming to answer the question of which relaying mode is greener (more energy efficient) and to address the issue of EE optimization. We develop an opportunistic relaying mode selection scheme, where FD relaying with one-stage self-interference cancellation (passive suppression) or two-stage self-interference cancellation (passive suppression + analog cancellation) or HD relaying is opportunistically selected, together with transmission power adaptation, to maximize the EE with given channel gains. A low-complexity joint mode selection and EE optimization algorithm are proposed. We show a counter-intuitive finding that with a relatively loose maximum transmission power constraint, FD relaying with two-stage self-interference cancellation is preferable to both FD relaying with one-stage self-interference cancellation and HD relaying, resulting in a higher optimized EE. A full range of power consumption sources is considered to rationalize our analysis. The effects of imperfect self-interference cancellation at relay, drain efficiency, and static circuit power on EE are investigated. Simulation results verify our theoretical analysis. Zhongxiang Wei, Xu Zhu 0001, Sumei Sun, Yi Huang 0001, Linhao Dong, Yufei Jiang |
IEEE J. Sel. Areas Commun. | 5 |
| 2013 | Optimal asymmetric resource allocation for dual-hop multi-relay LTE-Advanced systems in the downlinkabstractWe propose an optimal asymmetric resource allocation (ARA) algorithm for the decode-and-forward (DF) dualhop multi-relay Long Term Evolution Advanced (LTE-Advanced) system in the downlink, with orthogonal frequency division multiplexing (OFDM) transmission. Our work is different in that the time slots for the two hops via each of the relays are designed to be asymmetric, i.e., with K relays in the cell, a total of 2K time slots may be of different durations, which enhances the degree of freedom over the previous work. Also, a destination may be served by multiple relays at the same time to enhance the transmission diversity. Moreover, closed-form results for optimal allocation are derived, which requires only limited feedback information. Simulation results show that, thanks to the multi-relay, multi-user and time diversities, the proposed ARA scheme can provide a much better performance than the scheme with symmetric time allocation, as well as the scheme with asymmetric time allocation for a system composed of independent single-relay sub-systems, especially when the relays are close to the source. Linhao Dong, Xu Zhu 0001, Yi Huang 0001 |
ICC | 1 |
| 2013 | Power efficient 60 GHz wireless communication networks with relaysabstractIn this paper, we study the power consumption in relay networks of the 60 GHz wireless communication based on amplify-and-forward (AF) and decode-and-forward (DF) relaying strategies. We propose a total power consumption model including drive power, decoding power, and power consumption of power amplifier (PA). This model is formulated as a function of drive power, which gives an easy access to the system level optimisation. The optimal drive power that minimises the total power consumption while satisfying the performance requirement can be found by numerical searching method. The impact of relay's locations on the total power consumption is also investigated. We show that, with the same performance requirement, in the small source-relay separation case AF consumes less power than DF, while with larger separation, AF consumes significantly more power than DF. This is different from the common intuition that DF is always more power consuming than AF due to the extra decoding power consumption at relay, which is due to the fact that the large source-relay separation limits the effective destination signal-to-noise ratio (SNR) in AF, leading to more substantial decoding power consumption in the many more decoding iterations than DF. Linhao Dong, Sumei Sun, Xu Zhu 0001, Yeow-Khiang Chia |
PIMRC | 1 |
| 2011 | Optimal Asymmetric Resource Allocation for Multi-Relay Based LTE-Advanced SystemsabstractIn this paper, we propose a novel asymmetric resource allocation (RA) scheme for the decode-and-forward (DF) multi-relay based Long Term Evolution Advanced (LTE-Advanced) system, with the orthogonal frequency division multiplexing (OFDM) modulation. Given a fixed total transmit time duration, the time slot durations for the two hops via each of the K relays are designed to be asymmetric, which increases the degree of freedom for transmission significantly. We propose an optimal algorithm to maximise the system capacity, with joint asymmetric time allocation, power allocation, and subcarrier selection/allocation. Simulation results show that the proposed asymmetric RA scheme provides an enhanced performance over algorithms with symmetric time slot allocation, and it is also robust against the variation of relay nodes' locations. Linhao Dong, Xu Zhu 0001, Yi Huang 0001 |
GLOBECOM | 1 |