VLDB 2026 Research / reviewers in the wild / expert
Zhendong Peng
dblp:243/6846
· DBLP profile ↗
23ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0003-4844-9942ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 11 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 10 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Error-Aware Super-Resolution Channel Estimation for RIS-Aided Multi-User mmWave Systems
Zhendong Peng, Gui Zhou, Cunhua Pan, Maged Elkashlan, Cyril Leung |
ICC | 1 |
| 2026 | Channel Estimation for RIS-Aided MU-MIMO mmWave Systems With Direct Channel LinksabstractIn this paper, we propose a three-stage unified channel estimation strategy for reconfigurable intelligent surface (RIS)-aided multi-user (MU) multiple-input multiple-output (MIMO) millimeter wave (mmWave) systems with the existence of the direct channels, where the base station (BS), the users and the RIS are equipped with uniform planar array (UPA). The effectiveness of the developed three-stage strategy stems from the careful design of both the pilot signal sequence of the users and the vectors of RIS. Specifically, in Stage I, the cascaded channel components are eliminated by configuring the RIS phase shift vectors with a π difference to estimate the direct channels for all users. The orthogonal subspace projection is employed in Stage II to obtain equivalent signal matrices, enabling the estimation of angles of departure (AoDs) of the user-RIS channel for all users. In Stage III, we combine the signals of the time slots with the same pilots and project obtained measurement matrix to the orthogonal complement space of the component consisting of the portion of the direct channel, which removes the direct components and thus prevents error propagation from the direct channels to the cascaded channels. Then, we estimate the angles of arrival (AoAs) of the RIS-BS channel and remaining parameters of the cascaded channel for all users by exploiting the sparsity and correlation in the obtained equivalent matrices. Simulation results demonstrate that the proposed method yields better estimation performance than the existing methods. Taihao Zhang, Zhendong Peng, Cunhua Pan, Hong Ren, Jiangzhou Wang |
IEEE Trans. Commun. | 2 |
| 2026 | Hybrid Beamforming for RIS-Assisted Multiuser Fluid Antenna SystemsabstractRecent advances in reconfigurable antennas have led to the new concept of the fluid antenna system (FAS) for shape and position flexibility, as another degree of freedom for wireless communication enhancement. This paper explores the integration of a transmit FAS array for hybrid beamforming (HBF) into a reconfigurable intelligent surface (RIS)-assisted communication architecture for multiuser communications in the downlink, corresponding to the downlink RIS-assisted multiuser multiple-input single-output (MISO) FAS model (Tx RIS-assisted-MISO-FAS). By considering Rician channel fading, we formulate a sum-rate maximization optimization problem to alternately optimize the HBF matrix, the RIS phase-shift matrix, and the FAS position. Due to the strong coupling of multiple optimization variables, the multi-fractional summation in the sum-rate expression, the modulus-1 limitation of analog phase shifters and RIS, and the antenna position variables appearing in the exponent, this problem is highly non-convex, which is addressed through the block coordinate descent (BCD) framework in conjunction with semidefinite relaxation (SDR) and majorization-minimization (MM) methods. To reduce the computational complexity, we then propose a low-complexity grating-lobe (GL)-based telescopic-FA (TFA) system with multiple delicately deployed RISs under the sub-connected HBF architecture and the line-of-sight (LoS)-dominant channel condition, to allow closed-form solutions for the HBF and TFA position. Our simulation results illustrate that the former optimization scheme significantly enhances the achievable rate of the proposed system, while the GL-based TFA scheme also provides a considerable gain over conventional fixed-position antenna (FPA) systems, requiring statistical channel state information (CSI) only and with low computational complexity. Jiangong Chen, Yue Xiao 0001, Zhendong Peng, Jing Zhu 0004, Xia Lei 0001, Christos Masouros, Kai-Kit Wong |
IEEE Trans. Wirel. Commun. | 3 |
| 2026 | Novel Synchronization Scheme Based on Pilot Sharing in Cell-Free Massive MIMO Systems
Qihao Peng, Hong Ren, Zhendong Peng, Cunhua Pan, Maged Elkashlan, Dongming Wang 0002, Jiangzhou Wang, Xiaohu You 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2025 | Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
Xingchen Song, Brendan Fahy, Qiaochu Song, Zhendong Peng, Anshul Wadhawan, Denglin Jiang, Apurv Verma, Vinay Ramesh, Srivas Prasad, Michele Franceschini |
INTERSPEECH | 6 |
| 2025 | Radar Rainbow Beams for Wideband mmWave Communication: Beam Training and TrackingabstractWe propose a novel integrated sensing and communication (ISAC) scheme that leverages sensing to assist communication in light-of-sight (LoS) environments, ensuring fast initial access, seamless user tracking, and uninterrupted communication for millimeter wave (mmWave) wideband systems. True-time-delayers (TTDs) are utilized to generate frequency-dependent radar rainbow beams by controlling the beam squint effect. These beams cover users across the entire angular space simultaneously for fast beam training using just one orthogonal frequency-division multiplexing (OFDM) symbol. Three detection and estimation schemes are proposed based on radar rainbow beams for estimation of the users’ directions, distances, and velocities, which are then exploited for communication beamformer design. The first proposed scheme utilizes a single-antenna radar receiver and one set of rainbow beams, but may cause a Doppler ambiguity. To tackle this limitation, two additional schemes are introduced, utilizing two sets of rainbow beams and a multi-antenna receiver, respectively. Furthermore, the proposed detection and estimation schemes are extended to realize user tracking by choosing different subsets of OFDM subcarriers. This approach eliminates the need to switch phase shifters and TTDs, which is typically required for existing tracking schemes. Simulation results reveal the effectiveness of the proposed rainbow beam-based training and tracking methods for mobile users. Gui Zhou, Moritz Garkisch, Zhendong Peng, Cunhua Pan, Robert Schober |
IEEE J. Sel. Areas Commun. | 3 |
| 2024 | Hydraformer: One Encoder for All Subsampling RatesabstractIn automatic speech recognition, subsampling is essential for tackling diverse scenarios. However, the inadequacy of a single subsampling rate to address various real-world situations often necessitates training and deploying multiple models, consequently increasing associated costs. To address this issue, we propose HydraFormer, comprising HydraSub, a Conformer-based encoder, and a BiTransformer-based decoder. HydraSub encompasses multiple branches, each representing a distinct subsampling rate, allowing for the flexible selection of any branch during inference based on the specific use case. HydraFormer can efficiently manage different subsampling rates, significantly reducing training and deployment expenses. Experiments on AISHELL-1 and LibriSpeech datasets reveal that HydraFormer effectively adapts to various subsampling rates and languages while maintaining high recognition performance. Additionally, HydraFormer showcases exceptional stability, sustaining consistent performance under various initialization conditions, and exhibits robust transferability by learning from pretrained single subsampling rate automatic speech recognition models1. Yaoxun Xu, Xingchen Song, Zhiyong Wu 0001, Di Wu 0061, Zhendong Peng |
ICME | 5 |
| 2024 | Rainbow Beams for Wideband mmWave Radar: Beam TrainingabstractWe present a novel fast beam training method for fast moving targets in millimeter wave (mmWave) wideband radar systems. True-time-delayers (TTDs) are utilized to generate frequency-dependent radar rainbow beams using one orthogonal frequency-division multiplexing (OFDM) symbol, simultaneously covering targets located in the entire angular space for fast beam training. We first propose a scheme based on a single-antenna radar receiver. It can effectively detect and estimate different parameters of interest of targets, including their angles, distance related delays, and velocity related Doppler frequencies, but faces a Doppler ambiguity challenge. To tackle this limitation, we further introduce a scheme based on a multi-antenna receiver, which provides high-precision estimation performance. Simulation results reveal the effectiveness of the proposed rainbow beam-based training method for detecting and estimating mobile targets. Gui Zhou, Zhendong Peng, Cunhua Pan, Robert Schober |
WCNC | 2 |
| 2024 | Individual Channel Estimation for RIS-Aided Communication Systems - A General FrameworkabstractWe propose new pilot transmission protocols for acquiring channel state information (CSI) of individual reconfigurable intelligent surface (RIS) assisted channels. Our approach addresses the challenge of individual CSI acquisition when the RIS lacks sensing and signal processing capabilities. We use monostatic and bistatic full-duplex base stations (BSs) and exploit the reciprocity of the uplink and downlink channels to design channel estimation algorithms based on both unstructured and geometric channel models. Specifically, for unstructured channel models, we develop two different channel estimation algorithms that provide high accuracy and low pilot overhead, respectively, depending on the type of full-duplex BS used. Moreover, a unified estimation framework is proposed to determine the CSI based on geometric channel models for both types of full-duplex BSs. For the angle estimation required as part of the proposed framework, we further develop a high-precision algorithm based on atomic norm minimization (ANM) and a low-complexity algorithm based on orthogonal matching pursuit (OMP). Simulation results reveal that the proposed algorithms are superior to existing methods in terms of estimation accuracy, complexity, and pilot overhead. Gui Zhou, Zhendong Peng, Cunhua Pan, Robert Schober |
IEEE Trans. Wirel. Commun. | 2 |
| 2023 | Individual Channel Estimation for RIS-Aided mm Wave Communication SystemsabstractWe propose new pilot transmission protocols for acquiring channel state information (CSI) of individual reconfig-urable intelligent surface (RIS) assisted channels. Our approach addresses the challenge of individual CSI acquisition when the RIS lacks sensing and signal processing capabilities. We use monostatic and bistatic full-duplex (FD) base stations (BSs) and exploit the reciprocity of the uplink and downlink channels to design channel estimation algorithms. Specifically, a unified estimation framework is proposed to estimate the CSI based on geometric channel models for both types of FD BSs. To handle the angle estimation required as part of the proposed framework, we further investigate a high-accuracy algorithm based on atomic norm minimization (ANM) and a low-complexity algorithm based on orthogonal matching pursuit (OMP). Simulation results reveal that the proposed ANM based estimation scheme for bistatic BSs outperforms that for monostatic BSs, since the former requires a identical pilot overhead and achieves a similar estimation accuracy while having a much simpler hardware implementation. Gui Zhou, Cunhua Pan, Zhendong Peng, Robert Schober |
GLOBECOM | 3 |
| 2023 | LightGrad: Lightweight Diffusion Probabilistic Model for Text-to-SpeechabstractRecent advances in neural text-to-speech (TTS) models bring thousands of TTS applications into daily life, where models are deployed in cloud to provide services for customs. Among these models are diffusion probabilistic models (DPMs), which can be stably trained and are more parameter-efficient compared with other generative models. As transmitting data between customs and the cloud introduces high latency and the risk of exposing private data, deploying TTS models on edge devices is preferred. When implementing DPMs onto edge devices, there are two practical problems. First, current DPMs are not lightweight enough for resource-constrained devices. Second, DPMs require many denoising steps in inference, which increases latency. In this work, we present LightGrad, a lightweight DPM for TTS. LightGrad is equipped with a lightweight U-Net diffusion decoder and a training-free fast sampling technique, reducing both model parameters and inference latency. Streaming inference is also implemented in LightGrad to reduce latency further. Compared with Grad-TTS, LightGrad achieves 62.2% reduction in paramters, 65.7% reduction in latency, while preserving comparable speech quality on both Chinese Mandarin and English in 4 denoising steps1. Xingchen Song, Zhendong Peng, Fuping Pan, Zhiyong Wu 0001 |
ICASSP | 3 |
| 2023 | Fast-U2++: Fast and Accurate End-to-End Speech Recognition in Joint CTC/Attention FramesabstractRecently, the unified streaming and non-streaming two-pass (U2/U2++) end-to-end model for speech recognition has shown great performance in terms of streaming capability, accuracy and latency. In this paper, we present fast-U2++, an enhanced version of U2++ to further reduce partial latency. The core idea of fast-U2++ is to output partial results of the bottom layers in its encoder with a small chunk, while using a large chunk in the top layers of its encoder to compensate the performance degradation caused by the small chunk. More-over, we use knowledge distillation method to reduce the token emission latency. We present extensive experiments on Aishell-1 dataset. Experiments and ablation studies show that compared to U2++, fast-U2++ reduces model latency from 320ms to 80ms, and achieves a character error rate (CER) of 5.06% with a streaming setup. Chengdong Liang, Xiao-Lei Zhang 0001, Di Wu 0061, Shengqiang Li, Xingchen Song, Zhendong Peng, Fuping Pan |
ICASSP | 7 |
| 2023 | TrimTail: Low-Latency Streaming ASR with Simple But Effective Spectrogram-Level Length PenaltyabstractIn this paper, we present TrimTail, a simple but effective emission regularization method to improve the latency of streaming ASR models. The core idea of TrimTail is to apply length penalty (i.e., by trimming trailing frames, see Fig. 1-(b)) directly on the spectrogram of input utterances, which does not require any alignment. We demonstrate that TrimTail is computationally cheap and can be applied online and optimized with any training loss or any model architecture on any dataset without any extra effort by applying it on various end-to-end streaming ASR networks either trained with CTC loss [1] or Transducer loss [2]. We achieve 100 ~ 200ms latency reduction with equal or even better accuracy on both Aishell-1 and Librispeech. Moreover, by using TrimTail, we can achieve a 400ms algorithmic improvement of User Sensitive Delay (USD) with an accuracy loss of less than 0.2. Xingchen Song, Di Wu 0061, Zhiyong Wu 0001, Yuekai Zhang, Zhendong Peng, Wenpeng Li, Fuping Pan, Changbao Zhu |
ICASSP | 6 |
| 2023 | Two-Phase Channel Estimation for UPA-Type RIS-Aided Multi-User mmWave Systems with Reduced Pilot Overhead and Error PropagationabstractIn this paper, an efficient two-phase channel estimation scheme with reduced pilot overhead and error propagation is proposed for a uniform planar array (UPA)-type reconfigurable intelligent surface (RIS)-aided multi-user (MU) millimeter wave (mmWave) system. In Phase I, based on the carefully designed RIS phase shift matrix, all users jointly transmit the pilot signals to estimate the correlation factors between different propagation paths of the common RIS-base station (BS) channel, which facilitates a significant MU diversity gain. Then, in Phase II, with the constructed ambiguous RIS-BS channel composed of the correlation factors obtained in the previous phase, each user independently sends a few pilots to estimate their own ambiguous user-RIS channel so as to obtain the entire cascaded channel. Simulation results validate that the proposed algorithm outperforms the existing algorithms in terms of both pilot overhead and estimation accuracy, and that its estimation performance improves as the number of users increases. Zhendong Peng, Cunhua Pan, Gui Zhou, Hong Ren |
ICC | 1 |
| 2023 | ZeroPrompt: Streaming Acoustic Encoders are Zero-Shot Masked LMsabstractIn this paper, we present ZeroPrompt (Figure 1-(a)) and the corresponding Prompt-and-Refine strategy (Figure 3), two simple but effective training-free methods to decrease the Token Display Time (TDT) of streaming ASR models without any accuracy loss.The core idea of ZeroPrompt is to append zeroed content to each chunk during inference, which acts like a prompt to encourage the model to predict future tokens even before they were spoken.We argue that streaming acoustic encoders naturally have the modeling ability of Masked Language Models and our experiments demonstrate that ZeroPrompt is engineering cheap and can be applied to streaming acoustic encoders on any dataset without any accuracy loss.Specifically, compared with our baseline models, we achieve 350 ∼ 700ms reduction on First Token Display Time (TDT-F) and 100 ∼ 400ms reduction on Last Token Display Time (TDT-L), with theoretically and experimentally equal WER on both Aishell-1 and Librispeech datasets. Xingchen Song, Di Wu 0061, Zhendong Peng, Bo Dang 0004, Fuping Pan, Zhiyong Wu 0001 |
INTERSPEECH | 4 |
| 2023 | Branch-ECAPA-TDNN: A Parallel Branch Architecture to Capture Local and Global Features for Speaker Verification
Jiadi Yao, Chengdong Liang, Zhendong Peng, Xiao-Lei Zhang 0001 |
INTERSPEECH | 3 |
| 2022 | Channel Estimation for RIS-Aided mmWave MIMO System from 1-Sparse Recovery PerspectiveabstractIn this paper, we develop a two-phase based uplink channel estimation strategy with reduced pilot overhead for an reconfigurable intelligent surface (RIS)-aided millimeter wave (mmWave) multiple-input multiple-output (MIMO) communication system. Specifically, in Phase I, an OMP-based method is adopted to estimate the AoDs at the users. The remaining parameters including the common AoAs at the BS, the cascaded AoDs at the RIS, and the cascaded channel gains are estimated in Phase II. In particular, the estimation of cascaded AoDs and channel gains can be formulated as 1-sparse recovery problems by decomposing the estimation of a multi-antenna channel with$J$scatterers into estimating$J$single-scatterer channels for virtual single-antenna users. Finally, the theoretical number of pilots required for the proposed method are analyzed and the simulation results are presented to demonstrate the high channel estimation accuracy. Zhendong Peng, Gui Zhou, Cunhua Pan, Hong Ren |
GLOBECOM | 1 |
| 2022 | WENETSPEECH: A 10000+ Hours Multi-Domain Mandarin Corpus for Speech RecognitionabstractIn this paper, we present WenetSpeech, a multi-domain Mandarin corpus consisting of 10000+ hours high-quality labeled speech, 2400+ hours weakly labeled speech, and about 10000 hours unlabeled speech, with 22400+ hours in total. We collect the data from YouTube and Podcast, which covers a variety of speaking styles, scenarios, domains, topics and noisy conditions. An optical character recognition (OCR) method is introduced to generate the audio/text segmentation candidates for the YouTube data on the corresponding video subtitles, while a high-quality ASR transcription system is used to generate audio/text pair candidates for the Podcast data. Then we propose a novel end-to-end label error detection approach to further validate and filter the candidates. We also provide three manually labelled high-quality test sets along with WenetSpeech for evaluation – Dev for cross-validation purpose in training, Test_Net, collected from Internet for matched test, and Test_Meeting, recorded from real meetings for more challenging mismatched test. Baseline systems trained with WenetSpeech are provided for three popular speech recognition toolkits, namely Kaldi, ESPnet, and WeNet, and recognition results on the three test sets are also provided as benchmarks. To the best of our knowledge, WenetSpeech is the current largest open-source Mandarin speech corpus with transcriptions, which benefits research on production-level speech recognition. Hang Lv 0001, Qijie Shao, Chao Yang 0031, Lei Xie 0001, Hui Bu, Chenchen Zeng, Di Wu 0061, Zhendong Peng |
ICASSP | 12 |
| 2022 | WeNet 2.0: More Productive End-to-End Speech Recognition ToolkitabstractRecently, we made available WeNet [1], a production-oriented end-to-end speech recognition toolkit, which introduces a unified two-pass (U2) framework and a built-in runtime to address the streaming and non-streaming decoding modes in a single model.To further improve ASR performance and facilitate various production requirements, in this paper, we present WeNet 2.0 with four important updates.(1) We propose U2++, a unified two-pass framework with bidirectional attention decoders, which includes the future contextual information by a right-toleft attention decoder to improve the representative ability of the shared encoder and the performance during the rescoring stage.(2) We introduce an n-gram based language model and a WFSTbased decoder into WeNet 2.0, promoting the use of rich text data in production scenarios.(3) We design a unified contextual biasing framework, which leverages user-specific context (e.g., contact lists) to provide rapid adaptation ability for production and improves ASR accuracy in both with-LM and without-LM scenarios.(4) We design a unified IO to support large-scale data for effective model training.In summary, the brand-new WeNet 2.0 achieves up to 10% relative recognition performance improvement over the original WeNet on various corpora and makes available several important production-oriented features. Di Wu 0061, Zhendong Peng, Xingchen Song, Zhuoyuan Yao, Hang Lv 0001, Lei Xie 0001, Chao Yang 0031, Fuping Pan, Jianwei Niu 0002 |
INTERSPEECH | 3 |
| 2022 | Channel Estimation for RIS-Aided Multi-User mmWave Systems With Uniform Planar ArraysabstractIn this paper, we adopt a three-stage based uplink channel estimation protocol with reduced pilot overhead for an reconfigurable intelligent surface (RIS)-aided multi-user (MU) millimeter wave (mmWave) communication system, in which both the base station (BS) and the RIS are equipped with a uniform planar array (UPA). Specifically, in Stage I, the channel state information (CSI) of a typical user is estimated. To address the power leakage issue for the common angles-of-arrival (AoAs) estimation in this stage, we develop a low-complexity one-dimensional search method. In Stage II, a re-parameterized common BS-RIS channel is constructed with the estimated information from Stage I to estimate other users’ CSI. In Stage III, only the rapidly varying channel gains need to re-estimated. Furthermore, the proposed method can be extended to multi-antenna UPA-type users, by decomposing the estimation of a multi-antenna channel with$J$scatterers into estimating$J$single-scatterer channels for a virtual single-antenna user. An orthogonal matching pursuit (OMP)-based method is proposed to estimate the angles-of-departure (AoDs) at the users. Simulation results demonstrate that the proposed algorithm significantly achieves high channel estimation accuracy, which approaches the genie-aided upper bound in the high signal-to-noise ratio (SNR) regime. Zhendong Peng, Gui Zhou, Cunhua Pan, Hong Ren, A. Lee Swindlehurst, Petar Popovski, Gang Wu 0001 |
IEEE Trans. Commun. | 1 |
| 2021 | WeNet: Production Oriented Streaming and Non-Streaming End-to-End Speech Recognition ToolkitabstractIn this paper, we propose an open source speech recognition toolkit called WeNet, in which a new two-pass approach named U2 is implemented to unify streaming and non-streaming endto-end (E2E) speech recognition in a single model.The main motivation of WeNet is to close the gap between the research and deployment of E2E speech recognition models.WeNet provides an efficient way to ship automatic speech recognition (ASR) applications in real-world scenarios, which is the main difference and advantage to other open source E2E speech recognition toolkits.We develop a hybird connectionist temporal classification (CTC)/attention architecture with transformer or conformer as encoder and an attention decoder to rescore th CTC hypotheses.To achieve streaming and non-streaming in a unified model, we use a dynamic chunk-based attention strategy which allows the self-attention to focus on the right context with random length.Our experiments on the AISHELL-1 dataset show that our model achieves 5.03% relative character error rate (CER) reduction in non-streaming ASR compared to a standard non-streaming transformer.After model quantification, our model achieves reasonable RTF and latency at runtime.The toolkit is publicly available at https://github.com/mobvoi/wenet. Zhuoyuan Yao, Di Wu 0061, Fan Yu 0002, Chao Yang 0031, Zhendong Peng, Lei Xie 0001 |
Interspeech | 7 |
| 2020 | ABFL: An autoencoder based practical approach for software fault localization
Zhendong Peng, Xi Xiao 0001, Guangwu Hu, Arun Kumar Sangaiah, Mohammed Atiquzzaman, Shutao Xia |
Inf. Sci. | 1 |
| 2019 | Non-local Self-attention Structure for Function Approximation in Deep Reinforcement LearningabstractReinforcement learning is a framework to make sequential decisions. The combination with deep neural networks further improves the ability of this framework. Convolutional nerual networks make it possible to make sequential decisions based on raw pixels information directly and make reinforcement learning achieve satisfying performances in series of tasks. However, convolutional neural networks still have own limitations in representing geometric patterns and long-term dependencies that occur consistently in state inputs. To tackle with the limitation, we propose the self-attention architecture to augment the original network. It provides a better balance between ability to model long-range dependencies and computational efficiency. Experiments on Atari games illustrate that self-attention structure is significantly effective for function approximation in deep reinforcement learning. Xi Xiao 0001, Guangwu Hu, Yao Yao 0006, Dianyan Zhang, Zhendong Peng, Qing Li 0006, Shutao Xia |
ICASSP | 6 |