VLDB 2026 Research / reviewers in the wild / expert
Kailin Tan
dblp:284/9662
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2025
0009-0009-5802-0753ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MaskDSC: Resilient Digital Semantic Communication with Masked Transformer and Unequal Error ProtectionabstractWe propose “MaskDSC”, a novel system designed to facilitate robust visual data transmission over unreliable wireless channels. MaskDSC effectively balances compression efficiency and transmission resilience by leveraging contextual modeling within the semantic latent space, complemented by unequal error protection mechanism at the physical layer, ensuring compatibility with existing digital communication systems. The novelty of our approach lies in a dual-functional masked Transformer architecture that exploits causal-order contextual dependencies among visual tokens. This architecture not only enhances compression efficiency through improved contextual entropy modeling but also provides robust error concealment capabilities to address diverse transmission error patterns inherent in volatile wireless channels. Our experimental evaluations conducted on image datasets demonstrate that MaskDSC outperforms state-of-the-art transmission systems, especially in terms of efficiency and resilience under dynamic wireless channel conditions. Kailin Tan, Sixian Wang, Xiaoqi Qin, Zhenyu Liu 0002, Jincheng Dai |
WCNC | 2 |
| 2025 | DiffCom: Channel Received Signal Is a Natural Condition to Guide Diffusion Posterior SamplingabstractEnd-to-end visual communication systems typically optimize a trade-off between channel bandwidth costs and signal-level distortion metrics. However, under challenging physical conditions, this traditional coding and transmission paradigm often results in unrealistic reconstructions with perceptible blurring and aliasing artifacts, despite the inclusion of perceptual or adversarial losses for optimizing. This issue primarily stems from the receiver’s limited knowledge about the underlying data manifold and the use of deterministic decoding mechanisms. To address these limitations, this paper introducesDiffCom, a novel end-to-endgenerative communicationparadigm that utilizes off-the-shelf generative priors and probabilistic diffusion models for decoding, thereby improving perceptual quality without heavily relying on bandwidth costs and received signal quality. Unlike traditional systems that rely on deterministic decoders optimized solely for distortion metrics, ourDiffComleverages raw channel-received signal as a fine-grained condition to guide stochastic posterior sampling. Our approach ensures that reconstructions remain on the manifold of real data with a novel confirming constraint, enhancing the robustness and reliability of the generated outcomes. Furthermore,DiffComincorporates a blind posterior sampling technique to address scenarios with unknown forward transmission characteristics. Extensive experimental validations demonstrate thatDiffComnot only produces realistic reconstructions with details faithful to the original data but also achieves superior robustness against diverse wireless transmission degradations. Collectively, these advancements establishDiffComas a new benchmark in designing generative communication systems that offer enhanced robustness and generalization superiorities. Sixian Wang, Jincheng Dai, Kailin Tan, Xiaoqi Qin, Kai Niu 0001, Ping Zhang 0003 |
IEEE J. Sel. Areas Commun. | 3 |
| 2023 | WITT: A Wireless Image Transmission Transformer for Semantic CommunicationsabstractIn this paper, we aim to redesign the vision Transformer (ViT) as a new backbone to realize semantic image transmission, termed wireless image transmission transformer (WITT). Previous works build upon convolutional neural networks (CNNs), which are inefficient in capturing global dependencies, resulting in degraded end-to-end transmission performance especially for high-resolution images. To tackle this, the proposed WITT employs Swin Transformers as a more capable backbone to extract long-range information. Different from ViTs in image classification tasks, WITT is highly optimized for image transmission while considering the effect of the wireless channel. Specifically, we propose a spatial modulation module to scale the latent representations according to channel state information, which enhances the ability of a single model to deal with various channel conditions. As a result, extensive experiments verify that our WITT attains better performance for different image resolutions, distortion metrics, and channel conditions. The code is available at https://github.com/KeYang8/WITT. Ke Yang 0006, Sixian Wang, Jincheng Dai, Kailin Tan, Kai Niu 0001, Ping Zhang 0003 |
ICASSP | 4 |
| 2023 | Learned Image Transmission Toward Machine-Type Semantic CommunicationsabstractHumans tend to focus on only a few regions of interest (ROI) rather than perceiving the entire scene. This insight is also useful for machine tasks. Built upon the properties of ROI, in this paper, we propose a learned image transmission framework toward machine tasks, which ensures both the image reconstruction quality and the task accuracy. The whole system is optimized under a tripartite RDA tradeoff across the channel bandwidth cost (rate, R), the signal reconstruction quality (distortion, D), and the machine task performance (accuracy, A). According to the image content complexity distribution and the specific task, we incorporate both the entropy model and the ROI map to guide the source-channel coding rate allocation. As a result, we obtain the system coding gain. During this process, we develop two types of real-time ROI generation methods, suitable for high and low bandwidth cost regions, respectively. Experimental results show that our approach vastly outperforms state-of-the-art engineered image transmission methods and emerging image transmission methods. Moreover, we conduct an extensive ablation study to demonstrate the importance of individual components in our method, by which we expect to facilitate future research on this novel approach for machine-type semantic communications. Kailin Tan, Jincheng Dai, Sixian Wang, Ke Yang 0006, Kai Niu 0001 |
PIMRC | 1 |
| 2023 | Toward Adaptive Semantic Communications: Efficient Data Transmission via Online Learned Nonlinear Transform Source-Channel CodingabstractThe emerging field semantic communication is driving the research of end-to-end data transmission. By utilizing the powerful representation ability of deep learning models, learned data transmission schemes have exhibited superior performance than the established source and channel coding methods. While, so far, research efforts mainly concentrated on architecture and model improvements toward a static target domain. Despite their successes, such learned models are still suboptimal due to the limitations in model capacity and imperfect optimization and generalization, particularly when the testing data distribution or channel response is different from that adopted for model training, as is likely to be the case in real-world. To tackle this, in this paper, we propose a novel online learned joint source and channel coding approach that leverages the deep learning model’s overfitting property. Specifically, we update the off-the-shelf pre-trained models after deployment in a lightweight online fashion to adapt to the distribution shifts in source data and environment domain. We take the overfitting concept to the extreme, proposing a series of implementation-friendly methods to adapt the codec model or representations to an individual data or channel state instance, which can further lead to substantial gains in terms of the end-to-end rate-distortion performance. Accordingly, the streaming ingredients include both the semantic representations of source data and the online updated decoder model parameters. The system design is formulated as a joint optimization problem whose goal is to minimize the loss function, a tripartite trade-off among the data stream bandwidth cost, model stream bandwidth cost, and end-to-end distortion. The proposed methods enable the communication-efficient adaptation for all parameters in the network without sacrificing decoding speed. Extensive experiments, including user study, on continually changing target source data and wireless channel environments, demonstrate the effectiveness and efficiency of our approach, on which we outperform existing state-of-the-art engineered transmission scheme (VVC combined with 5G LDPC coded transmission). Jincheng Dai, Sixian Wang, Ke Yang 0006, Kailin Tan, Xiaoqi Qin, Zhongwei Si, Kai Niu 0001, Ping Zhang 0003 |
IEEE J. Sel. Areas Commun. | 4 |
| 2022 | Resolution-Adaptive Source-Channel Coding for End-to-End Wireless Image TransmissionabstractThe recent deep learning-based joint source-channel coding (deep JSCC) framework has shown superior performance on end-to-end wireless image transmission without suffering from the “cliff effect”. However, a fundamental limit of current deep JSCC schemes is that the unbalanced regional importance of the source image has not been explicitly taken into account. It evenly distributes the coding rate to every image patch leading to an evident degradation of the overall coding efficiency. To break this fundamental limit, we propose a novel end-to-end wireless image transmission scheme in this paper. Our scheme integrates the deep JSCC architecture and the quadtree-structured regional rate allocation strategy adopted in the HEVC standard, collected under the name “resolution-adaptive deep JSCC (RaDJSCC)”. Our new architecture perceives the content of the transmitted image and adaptively allocates more channel bandwidth to the complex pixel blocks. Results show that for high-resolution images, the proposed RaDJSCC transmission method generally outperforms the emerging analog transmission schemes using deep JSCC and the digital transmission schemes using classical separated source and channel coding, e.g., BPG + LDPC. Ke Yang 0006, Sixian Wang, Kailin Tan, Jincheng Dai, Dekun Zhou, Kai Niu 0001 |
GLOBECOM | 3 |
| 2022 | Nonlinear Transform Source-Channel Coding for Semantic CommunicationsabstractIn this paper, we propose a class of high-efficiency deep joint source-channel coding methods that can closely adapt to the source distribution under the nonlinear transform, it can be collected under the name nonlinear transform source-channel coding (NTSCC). In the considered model, the transmitter first learns a nonlinear analysis transform to map the source data into latent space, then transmits the latent representation to the receiver via deep joint source-channel coding. Our model incorporates the nonlinear transform as a strong prior to effectively extract the source semantic features and provide side information for source-channel coding. Unlike existing conventional deep joint source-channel coding methods, the proposed NTSCC essentially learns both the source latent representation and an entropy model as the prior on the latent representation. Accordingly, novel adaptive rate transmission and hyperprior-aided codec refinement mechanisms are developed to upgrade deep joint source-channel coding. The whole system design is formulated as an optimization problem whose goal is to minimize the end-to-end transmission rate-distortion performance under established perceptual quality metrics. Across test image sources with various resolutions, we find that the proposed NTSCC transmission method generally outperforms both the analog transmission using the standard deep joint source-channel coding and the classical separation-based digital transmission. Notably, the proposed NTSCC method can potentially support future semantic communications due to its content-aware ability and perceptual optimization goal. Jincheng Dai, Sixian Wang, Kailin Tan, Zhongwei Si, Xiaoqi Qin, Kai Niu 0001, Ping Zhang 0003 |
IEEE J. Sel. Areas Commun. | 3 |
| 2021 | Neural Layered Min-Sum Decoding for Protograph LDPC CodesabstractIn this paper, layered min-sum (MS) iterative decoding is formulated as a customized neural network following the sequential scheduling of check node (CN) updates. By virtue of the lifting structure of protograph low-density parity-check (LDPC) codes, identical network parameters are shared among all derived edges originating from the same edge in the protograph, which makes the number of learn- able parameters manageable. The proposed neural layered MS decoder can support arbitrary codelengths consequently. Moreover, an iteration-wise greedy training method is proposed to tune the parameters such that it avoids the vanishing gradient problem and accelerates the decoding convergence. Jincheng Dai, Kailin Tan, Kai Niu 0001, Mingzhe Chen, H. Vincent Poor, Shuguang Cui |
ICASSP | 3 |
| 2021 | Learning to Decode Protograph LDPC CodesabstractThe recent development of deep learning methods provides a new approach to optimize the belief propagation (BP) decoding of linear codes.However, the limitation of existing works is that the scale of neural networks increases rapidly with the codelength, thus they can only support short to moderate codelengths.From the point view of practicality, we propose a high-performance neural min-sum (MS) decoding method that makes full use of the lifting structure of protograph low-density parity-check (LDPC) codes.By this means, the size of the parameter array of each layer in the neural decoder only equals the number of edge-types for arbitrary codelengths.In particular, for protograph LDPC codes, the proposed neural MS decoder is constructed in a special way such that identical parameters are shared by a bundle of edges derived from the same edge-type.To reduce the complexity and overcome the vanishing gradient problem in training the proposed neural MS decoder, an iteration-byiteration (i.e., layer-by-layer in neural networks) greedy training method is proposed.With this, the proposed neural MS decoder tends to be optimized with faster convergence, which is aligned with the early termination mechanism widely used in practice.To further enhance the generalization ability of the proposed neural MS decoder, a codelength/rate compatible training method is proposed, which randomly selects samples from a set of codes lifted from the same base code.As a theoretical performance evaluation tool, a trajectory-based extrinsic information transfer (T-EXIT) chart is developed for various decoders.Both T-EXIT and simulation results show that the optimized MS decoding can provide faster convergence and up to 1dB gain compared with the plain MS decoding and its variants with only slightly increased complexity.In addition, it can even outperform the sum-product algorithm for some short codes. Jincheng Dai, Kailin Tan, Zhongwei Si, Kai Niu 0001, Mingzhe Chen, H. Vincent Poor, Shuguang Cui |
IEEE J. Sel. Areas Commun. | 2 |