VLDB 2026 Research / reviewers in the wild / expert
Shengshi Yao
dblp:285/4531
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0001-5463-8614ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Error-Resilient Semantic Communication for Speech Transmission Over Packet-Loss NetworksabstractReal-time speech communication over wireless networks remains challenging, as conventional channel protection mechanisms cannot effectively counter packet loss under stringent bandwidth and latency constraints. Semantic communication has emerged as a promising paradigm for enhancing the robustness of speech transmission by means of joint source channel coding (JSCC). However, its cross-layer design hinders practical deployment due to the incompatibility with existing digital communication systems. To address this, we perform JSCC over the network layer to combat packet loss and support real deployment. Inspired by the generative latent modeling, we propose Glaris, a generative latent-prior-based resilient speech semantic communication framework that performs resilient transform coding in the generative latent space. Generative latent priors enable high-quality packet loss concealment (PLC) at the receiver side, well-balancing semantic consistency and reconstruction fidelity. Additionally, an integrated error resilience mechanism is designed to mitigate the error propagation and improve the effectiveness of PLC. Compared with traditional packet-level forward error correction (FEC) strategies, our new method achieves enhanced robustness over dynamic wireless networks while reducing redundancy overhead significantly. Experimental results on the LibriSpeech dataset demonstrate that Glaris consistently outperforms existing error-resilient codecs, achieving JSCC-level robustness while maintaining seamless compatibility with existing systems, and it also strikes a favorable balance between transmission efficiency and error resilience. Zhuohang Han, Jincheng Dai, Shengshi Yao, Junyi Wang 0002, Yanlong Li 0001, Kai Niu 0001, Wenjun Xu 0001, Ping Zhang 0003 |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | Task-Scalable Image Semantic Communication via Conditional Affine Transforms and Pixel-Wise Quality ControlabstractDeep autoencoder-based joint source-channel coding (JSCC) has gained significant attention for end-to-end image semantic communication systems. However, existing methods typically optimize a uniform bandwidth-distortion trade-off over the entire image, potentially leading to the loss of crucial details and inconsistent content for tasks with diverse regions of interest. In this paper, we propose a flexible fine-grained bandwidth allocation method for deep JSCC that enables highly efficient, task-scalable image transmission across various semantic communication scenarios using a single codec. Our method optimizes the bandwidth-distortion trade-off by constraining image distortion through a 2D pixel-wise quality map. Guided by the pixel-wise quality map, we introduce a novel conditional affine transformation that generates dedicated semantic feature maps tailored to specific tasks. Additionally, we introduce a semantic guidance network to automatically generate task-aware quality maps via backpropagation without additional retraining. This approach leverages a pretrained variable-length neural JSCC codec and adjusts the transmission quality on a fine-grained level, eliminating the need to train separate models for different tasks. Experimental results demonstrate the effectiveness of our bandwidth allocation method, enhancing task-specific performance in various goal-oriented image communication scenarios without additional training. Shengshi Yao, Sixian Wang, Zhongwei Si, Zhenyu Liu 0002, Jincheng Dai |
WCNC | 2 |
| 2025 | SoundSpring: Loss-Resilient Audio Transceiver With Dual-Functional Masked Language ModelingabstractIn this paper, we propose “SoundSpring”, a cutting-edge error-resilient audio transceiver that marries the robustness benefits of joint source-channel coding (JSCC) while also being compatible with current digital communication systems. Unlike recent deep JSCC transceivers, which learn to directly map audio signals to analog channel-input symbols via neural networks, our SoundSpring adopts the layered architecture that delineates audio compression from digital coded transmission, but it sufficiently exploits the impressive in-context predictive capabilities of large language (foundation) models. Integrated with the casual-order mask learning strategy, our single model operates on the latent feature domain and serve dual-functionalities: as efficient audio compressors at the transmitter and as effective mechanisms for packet loss concealment at the receiver. By jointly optimizing towards both audio compression efficiency and transmission error resiliency, we show that mask-learned language models are indeed powerful contextual predictors, and our dual-functional compression and concealment framework offers fresh perspectives on the application of foundation language models in audio communication. Through extensive experimental evaluations, we establish that SoundSpring apparently outperforms contemporary audio transmission systems in terms of signal fidelity metrics and perceptual quality scores. These new findings not only advocate for the practical deployment of SoundSpring in learning-based audio communication systems but also inspire the development of future audio semantic transceivers. Shengshi Yao, Jincheng Dai, Xiaoqi Qin, Sixian Wang, Siye Wang, Kai Niu 0001, Ping Zhang 0003 |
IEEE J. Sel. Areas Commun. | 1 |
| 2023 | Wireless Deep Speech Semantic TransmissionabstractIn this paper, we propose a new class of high-efficiency semantic coded transmission methods to realize end-to-end speech transmission over wireless channels. We name the whole system as Deep Speech Semantic Transmission (DSST). Specifically, we introduce a nonlinear transform to map the speech source to semantic latent space and feed semantic features into source-channel encoder to generate the channel-input sequence. Guided by the variational modeling idea, we set an entropy model on the latent space to estimate the importance diversity among semantic feature embeddings. Accordingly, these semantic features of different importance can be reasonably allocated with different coding rates, which maximizes the system coding gain. Furthermore, we introduce a channel signal-to-noise ratio (SNR) adaptation mechanism such that a single model can be applied over various channel states. The end-to-end optimization of our model leads to a flexible rate-distortion (RD) tradeoff, supporting an adaptive rate wireless speech semantic transmission. Experimental results verify that our DSST system clearly outperforms current engineered speech transmission systems on both objective and subjective metrics. Compared with existing neural speech semantic transmission methods, our model saves up to 75% of channel bandwidth costs when achieving the same quality. Audio samples are available at https://ximoo123.github.io/DSST. Zixuan Xiao, Shengshi Yao, Jincheng Dai, Sixian Wang, Kai Niu 0001, Ping Zhang 0003 |
ICASSP | 2 |
| 2023 | Learned Image Transmission over MIMO Fading ChannelsabstractLearned image transmission (LIT) has shown promising progress in recent years to boost the end-to-end transmission performance in semantic communications. To further enhance the system efficiency, in this paper, we propose a novel LIT framework built on multiple-input multiple-output (MIMO) fading channels. In particular, the proposed framework supports concurrent transmission of multiple streams, which can maximize the multiplexing gain in end-to-end semantic communication systems. By jointly considering the entropy distribution of the image semantic features and the wireless MIMO channel states, we design a spatial multiplexing mechanism that can adaptively realize coding rate allocation and stream mapping. As a result, source content and channel environment will be seamlessly coupled, which maximizes the coding gain. Moreover, the proposed LIT model is versatile: a single model can support various transmission rates. The whole model is optimized under the constraint of transmission rate-distortion (RD) tradeoff. Experimental results verify that our scheme substantially increases the throughput of semantic communication systems, and outperforms traditional MIMO communication systems under realistic fading channels. Shengshi Yao, Sixian Wang, Jincheng Dai, Kai Niu 0001 |
PIMRC | 1 |
| 2023 | Variational Speech Waveform Compression to Catalyze Semantic CommunicationsabstractWe propose a novel neural waveform compression method to catalyze emerging speech semantic communications. By introducing nonlinear transform and variational modeling, we effectively capture the dependencies within speech frames and estimate the probabilistic distribution of the speech feature more accurately, giving rise to better compression performance. In particular, the speech signals are analyzed and synthesized by a pair of nonlinear transforms, yielding latent features. An entropy model with hyperprior is built to capture the probabilistic distribution of latent features, followed by quantization and entropy coding. The proposed waveform codec can be optimized flexibly towards arbitrary rate, and the other appealing feature is that it can be easily optimized for any differentiable loss function, including perceptual loss used in semantic communications. To further improve the speech quality, we incorporate residual coding to mitigate the degradation arising from quantization distortion at the latent space. Results indicate that achieving the same perceptual quality score, the proposed method saves up to 27% coding rate than widely used adaptive multi-rate wideband (AMR-WB) codec as well as emerging neural waveform coding methods. Shengshi Yao, Zixuan Xiao, Sixian Wang, Jincheng Dai, Kai Niu 0001, Ping Zhang 0003 |
WCNC | 1 |
| 2021 | A Novel Deep Learning Architecture for Wireless Image TransmissionabstractIn this paper, the problem of neural compression based image transmission over wireless channels is studied. Since all procedures are considered over wireless links, the quality of training is affected by wireless factors such as packet errors. In the considered model, compressed data given by the neural source encoder (NSE) are fed into an error-control channel encoder and modulated as discrete symbols sent over a memoryless channel. In the receiving end, the channel decoder and the neural source decoder (NSD) forms an iterative structure to reconstruct the original image. Since all neural compressed data are transmitted over wireless channels, the training of NSD is affected by wireless channel factors such as residual bit errors given by the channel decoder. Meanwhile, during outer-loop iterations, the NSD needs to match the variant of information reliability output by the channel decoder so as to build a global optimal receiver. To this end, a refiner neural network is first attached after the NSD to adjust its output as the format of a priori information sent into the channel decoder. Then, the extrinsic information transfer (EXIT) functions of channel decoder and NSD are derived. At each iteration, the reliability of messages sent into the NSD is explicitly predicted by using the EXIT chart. By this means, the NSD can be trained in a residual bit error aware manner, and we realize a joint learning and iterative decoding framework to ensure the quality of neural image transmission over realistic wireless channels. Sixian Wang, Jincheng Dai, Shengshi Yao, Kai Niu 0001, Ping Zhang 0003 |
GLOBECOM | 3 |