VLDB 2026 Research / reviewers in the wild / expert
Zhenzi Weng
dblp:282/0253
· DBLP profile ↗
8ranked-venue papers
8as first author
8since 2021 · last 2026
0000-0003-2773-2722ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generative Semantic Communications for Robust Speech-to-Text TranslationabstractIn this article, we propose a robust semantic communication system for speech transmission, named Ross-S2T, to execute the speech-to-text translation (S2TT) transmission efficiently. First, a deep semantic encoder is developed to directly convert speech in the source language to textual features associated with the target language, facilitating the end-to-end (E2E) semantic exchange to perform the S2TT task and reducing the amount of transmission data without performance degradation. To mitigate semantic impairments inherent in the corrupted speech, a novel generative adversarial network (GAN)-enabled deep semantic compensator is established to estimate the hidden semantic information within the speech and extract deep semantic features simultaneously, which enables robust semantic transmission for corrupted speech. Furthermore, a semantic probe-aided compensator is devised to enhance the semantic fidelity of recovered semantic features and improve the understandability of the target text. According to simulation results, the proposed Ross-S2T exhibits superior S2TT performance compared to conventional approaches and high robustness against semantic impairments. Zhenzi Weng, Zhijin Qin, Xiaoming Tao 0001 |
IEEE Trans. Wirel. Commun. | 1 |
| 2025 | Robust Semantic Communications for Speech TransmissionabstractIn this paper, we propose a robust semantic communication system for speech transmission, named Ross-S2T, by delivering the essential semantic information. Specifically, we consider the speech-to-text translation (S2TT) as the transmission goal. First, a new deep semantic encoder is developed to convert speech in the source language to textual features associated with the target language, facilitating the end-to-end semantic exchange to perform the S2TT task and reducing the transmission data without performance degradation. To mitigate semantic impairments inherent in the corrupted speech, a novel generative adversarial network (GAN)-enabled deep semantic compensator is established to estimate the lost semantic information within the speech and extract deep semantic features simultaneously, which enables robust semantic transmission for corrupted speech. Furthermore, a semantic probe-aided compensator is devised to enhance the semantic fidelity of recovered semantic features and improve the understandability of the target text. According to simulation results, the proposed Ross-S2T exhibits superior S2TT performance compared to conventional approaches and high robustness against semantic impairments. Zhenzi Weng, Zhijin Qin, Geoffrey Ye Li |
ICASSP | 1 |
| 2024 | Semantic MIMO Systems for Speech-to-Text TransmissionabstractSemantic communications have been utilized to execute numerous intelligent tasks by transmitting task-related semantic information instead of bits. In this article, we propose a semantic-aware speech-to-text transmission system for the single-user multiple-input multiple-output (MIMO) and multi-user MIMO communication scenarios, named SAC-ST. Particularly, a semantic communication system to serve the speech-to-text task at the receiver is first designed, which compresses the semantic information and generates the low-dimensional semantic features by leveraging the transformer module. In addition, a novel semantic-aware network is proposed to facilitate transmission with high semantic fidelity by identifying the critical semantic information and guaranteeing its accurate recovery. Furthermore, we extend the SAC-ST with a neural network-enabled channel estimation network to mitigate the dependence on accurate channel state information and validate the feasibility of SAC-ST in practical communication environments. Simulation results will show that the proposed SAC-ST outperforms the communication framework without the semantic-aware network for speech-to-text transmission over the MIMO channels in terms of the speech-to-text metrics, especially in the low signal-to-noise regime. Moreover, the SAC-ST with the developed channel estimation network is comparable to the SAC-ST with perfect channel state information. Zhenzi Weng, Zhijin Qin, Huiqiang Xie, Xiaoming Tao 0001, Khaled Ben Letaief |
IEEE Trans. Wirel. Commun. | 1 |
| 2023 | Task-Oriented Semantic Communications for Speech TransmissionabstractSemantic communications execute intelligent tasks at the receiver by only transmitting necessary information. In this paper, we introduce TOS-ST, a task-oriented semantic communication system for speech transmission, which efficiently serves the semantic tasks at the receiver, including speech-to-text translation and speech-to-speech translation. Particularly, TOS-ST condenses the input speech in the source language and extracts the task-related semantics features prior to transmission. At the receiver, these features are recovered and utilized by the neural network-based semantic preserver and machine translation module to generate the uncorrupted text in the target language. To perform the speech-to-speech translation task, the translated text passes through a sophisticated neural network to obtain speech in the target language. According to the simulation results, the TOS-ST outperforms conventional speech transmission systems and exhibits higher robustness against channel impairment. Zhenzi Weng, Zhijin Qin, Xiaoming Tao 0001 |
VTC Fall | 1 |
| 2023 | Deep Learning Enabled Semantic Communications With Speech Recognition and SynthesisabstractIn this paper, we develop a deep learning based semantic communication system for speech transmission, named DeepSC-ST. We take the speech recognition and speech synthesis as the transmission tasks of the communication system, respectively. First, the speech recognition-related semantic features are extracted for transmission by a joint semantic-channel encoder and the text is recovered at the receiver based on the received semantic features, which significantly reduces the required amount of data transmission without performance degradation. Then, we perform speech synthesis at the receiver, which dedicates to re-generate the speech signals by feeding the recognized text and the speaker information into a neural network module. To enable the DeepSC-ST adaptive to dynamic channel environments, we identify a robust model to cope with different channel conditions. According to the simulation results, the proposed DeepSC-ST significantly outperforms conventional communication systems and existing DL-enabled communication systems, especially in the low signal-to-noise ratio (SNR) regime. A software demonstration is further developed as a proof-of-concept of the DeepSC-ST. Zhenzi Weng, Zhijin Qin, Xiaoming Tao 0001, Chengkang Pan, Guangyi Liu 0001, Geoffrey Ye Li |
IEEE Trans. Wirel. Commun. | 1 |
| 2021 | Semantic Communications for Speech RecognitionabstractThe traditional communications transmit all the source date represented by bits, regardless of the content of source and the semantic information required by the receiver. However, in some applications, the receiver only needs part of the source data that represents critical semantic information, which prompts to transmit the application-related information, especially when bandwidth resources are limited. In this paper, we consider a semantic communication system for speech recognition by designing the transceiver as an end-to-end (E2E) system. Particularly, a deep learning (DL)-enabled semantic communication system, named DeepSC-SR, is developed to learn and extract text-related semantic features at the transmitter, which motivates the system to transmit much less than the source speech data without performance degradation. Moreover, in order to facilitate the proposed DeepSC-SR for dynamic channel environments, we investigate a robust model to cope with various channel environments without requiring retraining. The simulation results demonstrate that our proposed DeepSC-SR outperforms the traditional communication systems in terms of the speech recognition metrics, such as character-error-rate and word-error-rate, and is more robust to channel variations, especially in the low signal-to-noise (SNR) regime. Zhenzi Weng, Zhijin Qin, Geoffrey Ye Li |
GLOBECOM | 1 |
| 2021 | Semantic Communications for Speech SignalsabstractWe consider a semantic communication system for speech signals, named DeepSC-S. Motivated by the breakthroughs in deep learning (DL), we make an effort to recover the transmitted speech signals in the semantic communication systems, which minimizes the error at the semantic level rather than the bit level or symbol level as in the traditional communication systems. Particularly, based on an attention mechanism employing squeeze-and-excitation (SE) networks, we design the transceiver as an end-to-end (E2E) system, which learns and extracts the essential speech information. Furthermore, in order to facilitate the proposed DeepSC-S to work well on dynamic practical communication scenarios, we find a model yielding good performance when coping with various channel environments without retraining process. The simulation results demonstrate that our proposed DeepSC-S is more robust to channel variations and outperforms the traditional communication systems, especially in the low signal-to-noise (SNR) regime. Zhenzi Weng, Zhijin Qin, Geoffrey Ye Li |
ICC | 1 |
| 2021 | Semantic Communication Systems for Speech TransmissionabstractSemantic communications could improve the transmission efficiency significantly by exploring the semantic information. In this paper, we make an effort to recover the transmitted speech signals in the semantic communication systems, which minimizes the error at the semantic level rather than the bit or symbol level. Particularly, we design a deep learning (DL)-enabled semantic communication system for speech signals, named DeepSC-S. In order to improve the recovery accuracy of speech signals, especially for the essential information, DeepSC-S is developed based on an attention mechanism by utilizing a squeeze-and-excitation (SE) network. The motivation behind the attention mechanism is to identify the essential speech information by providing higher weights to them when training the neural network. Moreover, in order to facilitate the proposed DeepSC-S for dynamic channel environments, we find a general model to cope with various channel conditions without retraining. Furthermore, we investigate DeepSC-S in telephone systems as well as multimedia transmission systems to verify the model adaptation in practice. The simulation results demonstrate that our proposed DeepSC-S outperforms the traditional communications in both cases in terms of the speech signals metrics, such as signal-to-distortion ration and perceptual evaluation of speech distortion. Besides, DeepSC-S is more robust to channel variations, especially in the low signal-to-noise (SNR) regime. Zhenzi Weng, Zhijin Qin |
IEEE J. Sel. Areas Commun. | 1 |