EDBT 2026 Demo / reviewers in the wild / expert
Sixian Wang
dblp:79/11218
· DBLP profile ↗
20ranked-venue papers
5as first author
19since 2021 · last 2025
0000-0002-0621-1285ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 12 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Neural Hamiltonian Deformation Fields for Dynamic Scene RenderingabstractRepresenting and rendering dynamic scenes with complex motions remains challenging in computer vision and graphics. Recent dynamic view synthesis methods achieve high-quality rendering but often produce physically implausible motions. We introduce NeHaD, a neural deformation field for dynamic Gaussian Splatting governed by Hamiltonian mechanics. Our key observation is that existing methods using MLPs to predict deformation fields introduce inevitable biases, resulting in unnatural dynamics. By incorporating physics priors, we achieve robust and realistic dynamic scene rendering. Hamiltonian mechanics provides an ideal framework for modeling Gaussian deformation fields due to their shared phase-space structure, where primitives evolve along energy-conserving trajectories. We employ Hamiltonian neural networks to implicitly learn underlying physical laws governing deformation. Meanwhile, we introduce Boltzmann equilibrium decomposition, an energy-aware mechanism that adaptively separates static and dynamic Gaussians based on their spatial-temporal energy states for flexible rendering. To handle real-world dissipation, we employ second-order symplectic integration and local rigidity regularization as physics-informed constraints for robust dynamics modeling. Additionally, we extend NeHaD to adaptive streaming through scale-aware mipmapping and progressive optimization. Extensive experiments demonstrate that NeHaD achieves physically plausible results with a rendering quality-efficiency trade-off. To our knowledge, this is the first exploration leveraging Hamiltonian mechanics for neural Gaussian deformation, enabling physically realistic dynamic scene rendering with streaming capabilities. Hai-Long Qin, Sixian Wang, Guo Lu, Jincheng Dai |
SIGGRAPH Asia | 2 |
| 2025 | MaskDSC: Resilient Digital Semantic Communication with Masked Transformer and Unequal Error ProtectionabstractWe propose “MaskDSC”, a novel system designed to facilitate robust visual data transmission over unreliable wireless channels. MaskDSC effectively balances compression efficiency and transmission resilience by leveraging contextual modeling within the semantic latent space, complemented by unequal error protection mechanism at the physical layer, ensuring compatibility with existing digital communication systems. The novelty of our approach lies in a dual-functional masked Transformer architecture that exploits causal-order contextual dependencies among visual tokens. This architecture not only enhances compression efficiency through improved contextual entropy modeling but also provides robust error concealment capabilities to address diverse transmission error patterns inherent in volatile wireless channels. Our experimental evaluations conducted on image datasets demonstrate that MaskDSC outperforms state-of-the-art transmission systems, especially in terms of efficiency and resilience under dynamic wireless channel conditions. Kailin Tan, Sixian Wang, Xiaoqi Qin, Zhenyu Liu 0002, Jincheng Dai |
WCNC | 3 |
| 2025 | Task-Scalable Image Semantic Communication via Conditional Affine Transforms and Pixel-Wise Quality ControlabstractDeep autoencoder-based joint source-channel coding (JSCC) has gained significant attention for end-to-end image semantic communication systems. However, existing methods typically optimize a uniform bandwidth-distortion trade-off over the entire image, potentially leading to the loss of crucial details and inconsistent content for tasks with diverse regions of interest. In this paper, we propose a flexible fine-grained bandwidth allocation method for deep JSCC that enables highly efficient, task-scalable image transmission across various semantic communication scenarios using a single codec. Our method optimizes the bandwidth-distortion trade-off by constraining image distortion through a 2D pixel-wise quality map. Guided by the pixel-wise quality map, we introduce a novel conditional affine transformation that generates dedicated semantic feature maps tailored to specific tasks. Additionally, we introduce a semantic guidance network to automatically generate task-aware quality maps via backpropagation without additional retraining. This approach leverages a pretrained variable-length neural JSCC codec and adjusts the transmission quality on a fine-grained level, eliminating the need to train separate models for different tasks. Experimental results demonstrate the effectiveness of our bandwidth allocation method, enhancing task-specific performance in various goal-oriented image communication scenarios without additional training. Shengshi Yao, Sixian Wang, Zhongwei Si, Zhenyu Liu 0002, Jincheng Dai |
WCNC | 3 |
| 2025 | DiffCom: Channel Received Signal Is a Natural Condition to Guide Diffusion Posterior SamplingabstractEnd-to-end visual communication systems typically optimize a trade-off between channel bandwidth costs and signal-level distortion metrics. However, under challenging physical conditions, this traditional coding and transmission paradigm often results in unrealistic reconstructions with perceptible blurring and aliasing artifacts, despite the inclusion of perceptual or adversarial losses for optimizing. This issue primarily stems from the receiver’s limited knowledge about the underlying data manifold and the use of deterministic decoding mechanisms. To address these limitations, this paper introducesDiffCom, a novel end-to-endgenerative communicationparadigm that utilizes off-the-shelf generative priors and probabilistic diffusion models for decoding, thereby improving perceptual quality without heavily relying on bandwidth costs and received signal quality. Unlike traditional systems that rely on deterministic decoders optimized solely for distortion metrics, ourDiffComleverages raw channel-received signal as a fine-grained condition to guide stochastic posterior sampling. Our approach ensures that reconstructions remain on the manifold of real data with a novel confirming constraint, enhancing the robustness and reliability of the generated outcomes. Furthermore,DiffComincorporates a blind posterior sampling technique to address scenarios with unknown forward transmission characteristics. Extensive experimental validations demonstrate thatDiffComnot only produces realistic reconstructions with details faithful to the original data but also achieves superior robustness against diverse wireless transmission degradations. Collectively, these advancements establishDiffComas a new benchmark in designing generative communication systems that offer enhanced robustness and generalization superiorities. Sixian Wang, Jincheng Dai, Kailin Tan, Xiaoqi Qin, Kai Niu 0001, Ping Zhang 0003 |
IEEE J. Sel. Areas Commun. | 1 |
| 2025 | SoundSpring: Loss-Resilient Audio Transceiver With Dual-Functional Masked Language ModelingabstractIn this paper, we propose “SoundSpring”, a cutting-edge error-resilient audio transceiver that marries the robustness benefits of joint source-channel coding (JSCC) while also being compatible with current digital communication systems. Unlike recent deep JSCC transceivers, which learn to directly map audio signals to analog channel-input symbols via neural networks, our SoundSpring adopts the layered architecture that delineates audio compression from digital coded transmission, but it sufficiently exploits the impressive in-context predictive capabilities of large language (foundation) models. Integrated with the casual-order mask learning strategy, our single model operates on the latent feature domain and serve dual-functionalities: as efficient audio compressors at the transmitter and as effective mechanisms for packet loss concealment at the receiver. By jointly optimizing towards both audio compression efficiency and transmission error resiliency, we show that mask-learned language models are indeed powerful contextual predictors, and our dual-functional compression and concealment framework offers fresh perspectives on the application of foundation language models in audio communication. Through extensive experimental evaluations, we establish that SoundSpring apparently outperforms contemporary audio transmission systems in terms of signal fidelity metrics and perceptual quality scores. These new findings not only advocate for the practical deployment of SoundSpring in learning-based audio communication systems but also inspire the development of future audio semantic transceivers. Shengshi Yao, Jincheng Dai, Xiaoqi Qin, Sixian Wang, Siye Wang, Kai Niu 0001, Ping Zhang 0003 |
IEEE J. Sel. Areas Commun. | 4 |
| 2025 | ResiComp: Loss-Resilient Image Compression via Dual-Functional Masked Visual Token ModelingabstractRecent advancements in neural image codecs (NICs) are of significant compression performance, but limited attention has been paid to their error resilience. These resulting NICs tend to be sensitive to packet losses, which are prevalent in real-time communications. In this paper, we investigate how to elevate the resilience ability of NICs to combat packet losses. We propose ResiComp, a pioneering neural image compression framework with feature-domain packet loss concealment (PLC). Motivated by the inherent consistency between generation and compression, we advocate merging the tasks of entropy modeling and PLC into a unified framework focused on latent space context modeling. To this end, we take inspiration from the impressive generative capabilities of large language models (LLMs), particularly the recent advances of masked visual token modeling (MVTM). In specific, ResiComp develops a bi-directional masked Transformer to model the contextual dependencies among latents with dual-functionality: 1) iteratively acts as a conditional entropy model to boost compression efficiency; 2) operates latent PLC to improve resilience. During training, we integrate MVTM to mirror the effects of packet loss, enabling a dual-functional Transformer to restore the masked latents by predicting their missing values and conditional probability mass functions. Our ResiComp jointly optimizes compression efficiency and loss resilience. Moreover, ResiComp provides flexible coding modes, allowing for explicitly adjusting the efficiency-resilience trade-off in response to varying Internet or wireless network conditions. Extensive experiments demonstrate that ResiComp can significantly enhance the NIC’s resilience against packet losses, while exhibits a worthy trade-off between compression efficiency and packet loss resilience. Additionally, packet-level simulations, conducted using diverse network models based on real traces, demonstrate that ResiComp exhibits much better robustness to fluctuating network conditions compared to redundancy-based approaches like VTM + FEC. Sixian Wang, Jincheng Dai, Xiaoqi Qin, Ke Yang 0006, Kai Niu 0001, Ping Zhang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Wireless Deep Speech Semantic TransmissionabstractIn this paper, we propose a new class of high-efficiency semantic coded transmission methods to realize end-to-end speech transmission over wireless channels. We name the whole system as Deep Speech Semantic Transmission (DSST). Specifically, we introduce a nonlinear transform to map the speech source to semantic latent space and feed semantic features into source-channel encoder to generate the channel-input sequence. Guided by the variational modeling idea, we set an entropy model on the latent space to estimate the importance diversity among semantic feature embeddings. Accordingly, these semantic features of different importance can be reasonably allocated with different coding rates, which maximizes the system coding gain. Furthermore, we introduce a channel signal-to-noise ratio (SNR) adaptation mechanism such that a single model can be applied over various channel states. The end-to-end optimization of our model leads to a flexible rate-distortion (RD) tradeoff, supporting an adaptive rate wireless speech semantic transmission. Experimental results verify that our DSST system clearly outperforms current engineered speech transmission systems on both objective and subjective metrics. Compared with existing neural speech semantic transmission methods, our model saves up to 75% of channel bandwidth costs when achieving the same quality. Audio samples are available at https://ximoo123.github.io/DSST. Zixuan Xiao, Shengshi Yao, Jincheng Dai, Sixian Wang, Kai Niu 0001, Ping Zhang 0003 |
ICASSP | 4 |
| 2023 | WITT: A Wireless Image Transmission Transformer for Semantic CommunicationsabstractIn this paper, we aim to redesign the vision Transformer (ViT) as a new backbone to realize semantic image transmission, termed wireless image transmission transformer (WITT). Previous works build upon convolutional neural networks (CNNs), which are inefficient in capturing global dependencies, resulting in degraded end-to-end transmission performance especially for high-resolution images. To tackle this, the proposed WITT employs Swin Transformers as a more capable backbone to extract long-range information. Different from ViTs in image classification tasks, WITT is highly optimized for image transmission while considering the effect of the wireless channel. Specifically, we propose a spatial modulation module to scale the latent representations according to channel state information, which enhances the ability of a single model to deal with various channel conditions. As a result, extensive experiments verify that our WITT attains better performance for different image resolutions, distortion metrics, and channel conditions. The code is available at https://github.com/KeYang8/WITT. Ke Yang 0006, Sixian Wang, Jincheng Dai, Kailin Tan, Kai Niu 0001, Ping Zhang 0003 |
ICASSP | 2 |
| 2023 | Learned Image Transmission Toward Machine-Type Semantic CommunicationsabstractHumans tend to focus on only a few regions of interest (ROI) rather than perceiving the entire scene. This insight is also useful for machine tasks. Built upon the properties of ROI, in this paper, we propose a learned image transmission framework toward machine tasks, which ensures both the image reconstruction quality and the task accuracy. The whole system is optimized under a tripartite RDA tradeoff across the channel bandwidth cost (rate, R), the signal reconstruction quality (distortion, D), and the machine task performance (accuracy, A). According to the image content complexity distribution and the specific task, we incorporate both the entropy model and the ROI map to guide the source-channel coding rate allocation. As a result, we obtain the system coding gain. During this process, we develop two types of real-time ROI generation methods, suitable for high and low bandwidth cost regions, respectively. Experimental results show that our approach vastly outperforms state-of-the-art engineered image transmission methods and emerging image transmission methods. Moreover, we conduct an extensive ablation study to demonstrate the importance of individual components in our method, by which we expect to facilitate future research on this novel approach for machine-type semantic communications. Kailin Tan, Jincheng Dai, Sixian Wang, Ke Yang 0006, Kai Niu 0001 |
PIMRC | 3 |
| 2023 | Learned Image Transmission over MIMO Fading ChannelsabstractLearned image transmission (LIT) has shown promising progress in recent years to boost the end-to-end transmission performance in semantic communications. To further enhance the system efficiency, in this paper, we propose a novel LIT framework built on multiple-input multiple-output (MIMO) fading channels. In particular, the proposed framework supports concurrent transmission of multiple streams, which can maximize the multiplexing gain in end-to-end semantic communication systems. By jointly considering the entropy distribution of the image semantic features and the wireless MIMO channel states, we design a spatial multiplexing mechanism that can adaptively realize coding rate allocation and stream mapping. As a result, source content and channel environment will be seamlessly coupled, which maximizes the coding gain. Moreover, the proposed LIT model is versatile: a single model can support various transmission rates. The whole model is optimized under the constraint of transmission rate-distortion (RD) tradeoff. Experimental results verify that our scheme substantially increases the throughput of semantic communication systems, and outperforms traditional MIMO communication systems under realistic fading channels. Shengshi Yao, Sixian Wang, Jincheng Dai, Kai Niu 0001 |
PIMRC | 2 |
| 2023 | Variational Speech Waveform Compression to Catalyze Semantic CommunicationsabstractWe propose a novel neural waveform compression method to catalyze emerging speech semantic communications. By introducing nonlinear transform and variational modeling, we effectively capture the dependencies within speech frames and estimate the probabilistic distribution of the speech feature more accurately, giving rise to better compression performance. In particular, the speech signals are analyzed and synthesized by a pair of nonlinear transforms, yielding latent features. An entropy model with hyperprior is built to capture the probabilistic distribution of latent features, followed by quantization and entropy coding. The proposed waveform codec can be optimized flexibly towards arbitrary rate, and the other appealing feature is that it can be easily optimized for any differentiable loss function, including perceptual loss used in semantic communications. To further improve the speech quality, we incorporate residual coding to mitigate the degradation arising from quantization distortion at the latent space. Results indicate that achieving the same perceptual quality score, the proposed method saves up to 27% coding rate than widely used adaptive multi-rate wideband (AMR-WB) codec as well as emerging neural waveform coding methods. Shengshi Yao, Zixuan Xiao, Sixian Wang, Jincheng Dai, Kai Niu 0001, Ping Zhang 0003 |
WCNC | 3 |
| 2023 | Learned Source and Channel Coding for Talking-Head Semantic TransmissionabstractHow to efficiently transmit a special video over wireless channels? While the established systems work by combining H.26x video coding and 5G LDPC channel coding, its end-to-end transmission efficiency is still far away from the extreme for video sources in a specific domain. In this paper, we seek to design a special semantic communication system tailored for transmitting video calling streams over the wireless channels. Inspired by the recent progress in talking-head animation, we propose a talking- head semantic transmission (THST) system, which can efficiently transmit motion keypoint representation as compact semantic information to drive the free-view talk-heading synthesis at the receiver. Since the motion semantic key points are correlated, our THST system learns a nonlinear analysis transform to map the key points across multiple frames into latent space, then transmits the latent hyper semantic representation to the receiver via deep joint source-channel coding. Our system incorporates a latent prior to estimate the importance diversity on the semantic key points, accordingly, we realize variable rate joint source-channel coding to obtain system level coding gain. Extensive experimental validation shows that our THST system outperforms engineered competing systems on benchmark datasets. Moreover, due to the system level joint source and channel design, our method provides much more robust performance over noisy channels with only 33% bandwidth cost versus the current talking-head compression combined with 5G LDPC coded transmission systems. Weijie Yue, Jincheng Dai, Sixian Wang, Zhongwei Si, Kai Niu 0001 |
WCNC | 3 |
| 2023 | Toward Adaptive Semantic Communications: Efficient Data Transmission via Online Learned Nonlinear Transform Source-Channel CodingabstractThe emerging field semantic communication is driving the research of end-to-end data transmission. By utilizing the powerful representation ability of deep learning models, learned data transmission schemes have exhibited superior performance than the established source and channel coding methods. While, so far, research efforts mainly concentrated on architecture and model improvements toward a static target domain. Despite their successes, such learned models are still suboptimal due to the limitations in model capacity and imperfect optimization and generalization, particularly when the testing data distribution or channel response is different from that adopted for model training, as is likely to be the case in real-world. To tackle this, in this paper, we propose a novel online learned joint source and channel coding approach that leverages the deep learning model’s overfitting property. Specifically, we update the off-the-shelf pre-trained models after deployment in a lightweight online fashion to adapt to the distribution shifts in source data and environment domain. We take the overfitting concept to the extreme, proposing a series of implementation-friendly methods to adapt the codec model or representations to an individual data or channel state instance, which can further lead to substantial gains in terms of the end-to-end rate-distortion performance. Accordingly, the streaming ingredients include both the semantic representations of source data and the online updated decoder model parameters. The system design is formulated as a joint optimization problem whose goal is to minimize the loss function, a tripartite trade-off among the data stream bandwidth cost, model stream bandwidth cost, and end-to-end distortion. The proposed methods enable the communication-efficient adaptation for all parameters in the network without sacrificing decoding speed. Extensive experiments, including user study, on continually changing target source data and wireless channel environments, demonstrate the effectiveness and efficiency of our approach, on which we outperform existing state-of-the-art engineered transmission scheme (VVC combined with 5G LDPC coded transmission). Jincheng Dai, Sixian Wang, Ke Yang 0006, Kailin Tan, Xiaoqi Qin, Zhongwei Si, Kai Niu 0001, Ping Zhang 0003 |
IEEE J. Sel. Areas Commun. | 2 |
| 2023 | Wireless Deep Video Semantic TransmissionabstractIn this paper, we design a new class of high-efficiency deep joint source-channel coding methods to achieve end-to-end video transmission over wireless channels. The proposed methods exploit nonlinear transform and conditional coding architecture to adaptively extract semantic features across video frames, and transmit semantic feature domain representations over wireless channels via deep joint source-channel coding. Our framework is collected under the name deep video semantic transmission (DVST). In particular, benefiting from the strong temporal prior provided by the feature domain context, the learned nonlinear transform function becomes temporally adaptive, resulting in a richer and more accurate entropy model guiding the transmission of current frame. Accordingly, a novel rate adaptive transmission mechanism is developed to customize deep joint source-channel coding for video sources. It learns to allocate the limited channel bandwidth within and among video frames to maximize the overall transmission performance. The whole DVST design is formulated as an optimization problem whose goal is to minimize the end-to-end transmission rate-distortion performance under perceptual quality metrics or machine vision task performance metrics. Across standard video source test sequences and various communication scenarios, experiments show that our DVST can generally surpass traditional wireless video coded transmission schemes. The proposed DVST framework can well support future semantic communications due to its video content-aware and machine vision task integration abilities. Sixian Wang, Jincheng Dai, Kai Niu 0001, Zhongwei Si, Chao Dong 0002, Xiaoqi Qin, Ping Zhang 0003 |
IEEE J. Sel. Areas Commun. | 1 |
| 2022 | Perceptual Learned Source-Channel Coding for High-Fidelity Image Semantic TransmissionabstractAs one novel approach to realize end-to-end wireless image semantic transmission, deep learning-based joint source-channel coding (deep JSCC) method is emerging in both deep learning and communication communities. However, current deep JSCC image transmission systems are typically optimized for traditional distortion metrics such as peak signal-to-noise ratio (PSNR) or multi-scale structural similarity (MS-SSIM). But for low transmission rates, due to the imperfect wireless channel, these distortion metrics lose significance as they favor pixel-wise preservation. To account for human visual perception in semantic communications, it is of great importance to develop new deep JSCC systems optimized beyond traditional PSNR and MS-SSIM metrics. In this paper, we introduce adversarial losses to optimize deep JSCC, which tends to preserve global semantic information and local texture. Our new deep JSCC architecture combines encoder, wireless channel, decoder/generator, and discriminator, which are jointly learned under both perceptual and adversarial losses. Our method yields human visually much more pleasing results than state-of-the-art engineered image coded transmission systems and traditional deep JSCC systems. A user study confirms that achieving perceptually similar end-to-end image transmission quality, the proposed method can save about 50% wireless channel bandwidth costs. Sixian Wang, Jincheng Dai, Zhongwei Si, Dekun Zhou, Kai Niu 0001 |
GLOBECOM | 2 |
| 2022 | Resolution-Adaptive Source-Channel Coding for End-to-End Wireless Image TransmissionabstractThe recent deep learning-based joint source-channel coding (deep JSCC) framework has shown superior performance on end-to-end wireless image transmission without suffering from the “cliff effect”. However, a fundamental limit of current deep JSCC schemes is that the unbalanced regional importance of the source image has not been explicitly taken into account. It evenly distributes the coding rate to every image patch leading to an evident degradation of the overall coding efficiency. To break this fundamental limit, we propose a novel end-to-end wireless image transmission scheme in this paper. Our scheme integrates the deep JSCC architecture and the quadtree-structured regional rate allocation strategy adopted in the HEVC standard, collected under the name “resolution-adaptive deep JSCC (RaDJSCC)”. Our new architecture perceives the content of the transmitted image and adaptively allocates more channel bandwidth to the complex pixel blocks. Results show that for high-resolution images, the proposed RaDJSCC transmission method generally outperforms the emerging analog transmission schemes using deep JSCC and the digital transmission schemes using classical separated source and channel coding, e.g., BPG + LDPC. Ke Yang 0006, Sixian Wang, Kailin Tan, Jincheng Dai, Dekun Zhou, Kai Niu 0001 |
GLOBECOM | 2 |
| 2022 | Distributed Image Transmission Using Deep Joint Source-Channel CodingabstractWe study the problem of deep joint source-channel coding (D-JSCC) for correlated image sources, where each source is transmitted through a noisy independent channel to the common receiver. In particular, we consider a pair of images captured by two cameras with probably overlapping fields of view transmitted over wireless channels and reconstructed in the center node. The challenging problem involves designing a practical code to utilize both source and channel correlations to improve transmission efficiency without additional transmission overhead. To tackle this, we need to consider the common information across two stereo images as well as the differences between two transmission channels. In this case, we propose a deep neural networks solution that includes lightweight edge encoders and a powerful center decoder. Besides, in the decoder, we propose a novel channel state information aware cross attention module to highlight the overlapping fields and leverage the relevance between two noisy feature maps. Our results show the impressive improvement of reconstruction quality in both links by exploiting the noisy representations of the other link. Moreover, the proposed scheme shows competitive results compared to the separated schemes with capacity-achieving channel codes. Sixian Wang, Ke Yang 0006, Jincheng Dai, Kai Niu 0001 |
ICASSP | 1 |
| 2022 | Nonlinear Transform Source-Channel Coding for Semantic CommunicationsabstractIn this paper, we propose a class of high-efficiency deep joint source-channel coding methods that can closely adapt to the source distribution under the nonlinear transform, it can be collected under the name nonlinear transform source-channel coding (NTSCC). In the considered model, the transmitter first learns a nonlinear analysis transform to map the source data into latent space, then transmits the latent representation to the receiver via deep joint source-channel coding. Our model incorporates the nonlinear transform as a strong prior to effectively extract the source semantic features and provide side information for source-channel coding. Unlike existing conventional deep joint source-channel coding methods, the proposed NTSCC essentially learns both the source latent representation and an entropy model as the prior on the latent representation. Accordingly, novel adaptive rate transmission and hyperprior-aided codec refinement mechanisms are developed to upgrade deep joint source-channel coding. The whole system design is formulated as an optimization problem whose goal is to minimize the end-to-end transmission rate-distortion performance under established perceptual quality metrics. Across test image sources with various resolutions, we find that the proposed NTSCC transmission method generally outperforms both the analog transmission using the standard deep joint source-channel coding and the classical separation-based digital transmission. Notably, the proposed NTSCC method can potentially support future semantic communications due to its content-aware ability and perceptual optimization goal. Jincheng Dai, Sixian Wang, Kailin Tan, Zhongwei Si, Xiaoqi Qin, Kai Niu 0001, Ping Zhang 0003 |
IEEE J. Sel. Areas Commun. | 2 |
| 2021 | A Novel Deep Learning Architecture for Wireless Image TransmissionabstractIn this paper, the problem of neural compression based image transmission over wireless channels is studied. Since all procedures are considered over wireless links, the quality of training is affected by wireless factors such as packet errors. In the considered model, compressed data given by the neural source encoder (NSE) are fed into an error-control channel encoder and modulated as discrete symbols sent over a memoryless channel. In the receiving end, the channel decoder and the neural source decoder (NSD) forms an iterative structure to reconstruct the original image. Since all neural compressed data are transmitted over wireless channels, the training of NSD is affected by wireless channel factors such as residual bit errors given by the channel decoder. Meanwhile, during outer-loop iterations, the NSD needs to match the variant of information reliability output by the channel decoder so as to build a global optimal receiver. To this end, a refiner neural network is first attached after the NSD to adjust its output as the format of a priori information sent into the channel decoder. Then, the extrinsic information transfer (EXIT) functions of channel decoder and NSD are derived. At each iteration, the reliability of messages sent into the NSD is explicitly predicted by using the EXIT chart. By this means, the NSD can be trained in a residual bit error aware manner, and we realize a joint learning and iterative decoding framework to ensure the quality of neural image transmission over realistic wireless channels. Sixian Wang, Jincheng Dai, Shengshi Yao, Kai Niu 0001, Ping Zhang 0003 |
GLOBECOM | 1 |
| 1992 | The extraction of the best SGLD texture features in the ultrasound B-scan images of cancered stomach coatsabstractSGLD (spatial gray level dependence) matrices are used to analyze the B-scan images of 23 samples of normal stomach coats and 14 samples of cancerous stomach coats. According to these matrices, the values of eight texture features of each sample image are computed. Two groups of conditional frequency distributions are obtained. On the basis of these distributions, the authors evaluated the quality, which reflects the error probability in discriminating between pattern classes of all the features. By comparing the measurements of the quality, the authors select from these features the most effective ones in discriminating between a normal stomach and a cancerous stomach. The evaluation methods include normal distribution hypothesis testing, and T testing. The result of the experiments indicates that the selected texture features can be applied to an automatic diagnosis system in the near future.> Mengyang Liao, Jiamei Qin, Sixian Wang |
CBMS | 4 |