Zhicheng Bao

dblp:321/5374 · DBLP profile ↗
← Back
21ranked-venue papers
5as first author
21since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 12 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Semantic Knowledge Base Based Dual-mode Video Semantic Communication
Zhicheng Bao, Nan Ma 0014, Chen Dong 0001, Hao Chen 0013, Xiaodong Xu 0001, Ping Zhang 0003
ICC1
2026 Breaking Semantic-Aware Watermarks via LLM-Guided Coherence-Preserving Semantic Injection
abstract
Generative images have proliferated on Web platforms in social media and online copyright distribution scenarios, and semantic watermarking has increasingly been integrated into diffusion models to support reliable provenance tracking and forgery prevention for web content. Traditional noise-layer-based watermarking, however, remains vulnerable to inversion attacks that can recover embedded signals. To mitigate this, recent content-aware semantic watermarking schemes bind watermark signals to high-level image semantics, constraining local edits that would otherwise disrupt global coherence. Yet, large language models (LLMs) possess structured reasoning capabilities that enable targeted exploration of semantic spaces, allowing locally fine-grained but globally coherent semantic alterations that invalidate such bindings. To expose this overlooked vulnerability, we introduce a Coherence-Preserving Semantic Injection (CSI) attack that leverages LLM-guided semantic manipulation under embedding-space similarity constraints. This alignment enforces visual-semantic consistency while selectively perturbing watermark-relevant semantics, ultimately inducing detector misclassification. Extensive empirical results show that CSI consistently outperforms prevailing attack baselines against content-aware semantic watermarking, revealing a fundamental security weakness of current semantic watermark designs when confronted with LLM-driven semantic perturbations.
Xiaoyu Li 0001, Zhicheng Bao, Xiaoyan Feng, Jiaojiao Jiang 0001
WWW3
2026 GSCNet: A transformer-based granular style control network for artifact-free image style transfer
Zhicheng Bao, Haoran Fan, Xiaoyu Li 0001, Qipei Nong, Jiaojiao Jiang 0001
J. Vis. Commun. Image Represent.2
2026 Coverage-Enhanced Semantic Communication Systems for Cellular Networks
Yunlu Wang, Chen Dong 0001, Wannian An, Zhicheng Bao, Hongchao Jiang, Mengying Sun, Xiaodong Xu 0001
IEEE Trans. Commun.4
2025 An Optional 2D Feature Scene Text Recognition Network Based on Transformer
Mayire Ibrayim, Jianjun Kang, Zhicheng Bao
ICIG (3)3
2025 Multi-modal End-to-End Text Spotting Networks: Interactive Enhancements Between Visual and Semantic Features
Mayire Ibrayim, Yefei Qian, Zhicheng Bao
ICIG (3)3
2025 Arbitrary-Shaped Text Detection with Hierarchical Feature Refinement and Nonlinear-Enhanced Transformers
abstract
Detecting arbitrary-shaped scene text is a challenging task due to the irregularities in font, size, color, orientation, and shape, which often lead to detection errors. In this paper, we propose a novel boundary-learning-based, coarse-to-fine text detection network that combines the strengths of CNNs and Transformers. Specifically, we introduce two key components: the Hierarchical Feature Refinement Network (HFRNet) and the Nonlinear-Enhanced Transformer Module (NETM). HFRNet enhances multi-scale feature extraction, improving the detection of text at various scales through dynamic convolution kernel sampling and attention mechanisms. This enables better adaptation to spatial scale variations and geometric deformations. NETM, on the other hand, leverages multi-head self-attention and nonlinear feature mapping to improve the representation of complex sequential data, allowing for more accurate text boundary detection in a coarse-to-fine manner. Our method, integrating HFRNet and NETM, achieves state-of-the-art performance on benchmark scene text detection datasets, including Total-Text and CTW1500.
Zhicheng Bao, Mayire Ibrayim
IJCNN1
2025 Research on Video Semantic Transmission Technology with Dynamic GOP Segmentation and Scene Adaptation
abstract
Video semantic communication, as a cutting-edge field in the convergence of communication and artificial intelligence, is dedicated to improving the quality of video communication through the efficient transmission of semantic features. However, existing semantic systems face problems such as the complexity of shared feature extraction and cross-scene feature conflicts during multi-scene switching. To address this challenge, this paper proposes a video semantic transmission technique based on dynamic Group of Pictures (GOP) segmentation. Specifically, the performance advantages of dynamic GOP under multiple wireless channels are verified by designing a multiscale fusion transition detection algorithm and a dynamic GOP division strategy. The experimental results show that the proposed method can adapt to different content scenarios, significantly optimize the semantic feature extraction and video reconstruction process, and provide reliable technical support for video semantic transmission over complex communication links.
Zhicheng Bao, Chen Dong 0001, Xiaodong Xu 0001
PIMRC2
2025 Cross-attention multi-perspective fusion network based fake news censorship
Weishan Zhang, Zhicheng Bao, Zhenqi Wang
Neurocomputing3
2025 MDVSC - Efficient Wireless Model Division Video Semantic Communication
abstract
This article introduces a novel method for transmitting video data over noisy wireless channels with high efficiency and controllability. The method derivates from model division multiple access (MDMA) to extract common semantic features from video frames. It also uses deep joint source-channel coding (JSCC) as the main framework to establish communication links and deal with channel noise. An entropy-based semantic importance coding scheme is developed to adjust the data amount accurately and explicitly. We name our method as model division video semantic communication (MDVSC). The main steps of our approach are as follows: first, video frames are transformed into a latent space to reduce computational complexity and redistribute data. Then, common features and individual features are extracted, and semantic importance coding is applied to further eliminate redundant semantic information under the communication bandwidth constraint. We evaluate our method on standard video test sequences and compare it with traditional wireless video coding methods. The results show that MDVSC generally surpasses the conventional methods in terms of quality metrics and has the capability to control code length precisely. Moreover, additional experiments and ablation studies are conducted to demonstrate its potential for various tasks.
Zhicheng Bao, Haotai Liang, Chen Dong 0001, Xiaodong Xu 0001, Ping Zhang 0003
IEEE Internet Things J.1
2025 Semantic Similarity Score for Measuring Visual Similarity at Semantic Level
abstract
With the rapid development of Internet of Things (IoT) technology, more sensors are required to operate in complex channel scenarios and under limited communication resources. Semantic communication, as an emerging paradigm, extracts, transmits, and reconstructs information at the semantic level, offering advantages, such as high compression rates and strong noise resistance. These features are expected to find widespread application across various IoT scenarios. However, widely used image similarity evaluation metrics like peak signal-to-noise ratio and multiscale structural similarity index primarily focus on pixel or structural features, making it challenging to accurately measure the loss of semantic-level information during transmission. This limitation poses challenges for the performance evaluation of visual semantic communication systems and restricts the emergence of more novel and efficient systems. To address this issue, we propose a new semantic evaluation metric-semantic similarity score (SeSS). This metric is based on Scene Graph Generation and graph matching techniques, transforming image similarity scores into graph matching scores. By manually annotating thousands of image pairs, we fine-tuned the hyperparameters within SeSS to align it more closely with human semantic perception. The performance of SeSS has been tested across various image datasets and specific IoT visual tasks. Experimental results demonstrate the effectiveness of SeSS in measuring differences in semantic-level information between images, making it a valuable tool for evaluating visual semantic communication systems. This development is expected to encourage the emergence of more robust systems suited for diverse IoT scenarios. The code of SeSS is openly available onhttps://github.com/FSR3340/Semantic_Similarty_ScoreGitHub.
Senran Fan, Zhicheng Bao, Chen Dong 0001, Haotai Liang, Xiaodong Xu 0001, Ping Zhang 0003
IEEE Internet Things J.2
2025 Semantic-Importance-Aware Communication Over MIMO Fading Channels
abstract
Semantic communication, a promising paradigm for next-generation wireless systems, optimizes the representation of semantic information and its resilience to channel effects, outperforming traditional systems in low signal-to-noise ratio (SNR) environments. However, most existing frameworks focus on Single-Input Single-Output (SISO) channels which limits their use in multi-antenna systems. To address this gap, we propose Semantic Importance-Aware Communication (SIAC-MIMO), a system designed for Multiple-Input Multiple-Output (MIMO) fading channels. SIAC-MIMO integrates semantic symbol inequality with advanced channel-aware techniques. SIAC-MIMO prioritizes critical semantic symbols, adapts transmission to MIMO channel states, and employs Orthogonal Model Division Multiple Access (O-MDMA) for multi-user broadcasting to mitigate interference while enhancing scalability. A bilateral progressive training algorithm is introduced to align semantic allocation with channel eigenmodes. To evaluate the effectiveness of this system, a theoretical framework is developed to analyze semantic performance metrics, such as semantic information distortion and semantic outage probability. The experiments across 2W2 to 64W64 MIMO setups demonstrate SIAC-MIMO’s superiority, achieving 5–18% improvements in Mean Structural Similarity Index Measure (MS-SSIM) at low SNR in single-user scenarios and 16–23% improvements in multi-user MIMO setups compared to traditional source-channel separation schemes, highlighting the system’s potential for efficient and robust communication.
Haotai Liang, Chen Dong 0001, Wannian An, Zhicheng Bao, Xiaodong Xu 0001
IEEE Internet Things J.4
2025 Semantic-Importance-Aware Reordering-Enhanced Semantic Communication System With OFDM Transmission
abstract
As a novel communication paradigm, semantic communication (SemCom) can greatly improve communication efficiency, which has aroused extensive research by scholars worldwide. As one of the important aspects of digital communication nowadays, how to combine channel estimation with SemCom is an important research direction. In this article, based on orthogonal frequency-division multiplexing (OFDM) communication architecture, the semantic importance-aware reordering-enhanced SemCom system (SIARE-SC) is proposed, which utilizes the inequality of semantic symbols combined with channel estimation in OFDM systems to reduce the distortion caused by channel estimation interpolation error (CEIE) and further improve the signal recovery quality. To enhance the generalizability of the system, we extend the verification of the effectiveness of SIARE-SC in various scenarios with different sources, channels, and pilot patterns. Furthermore, the importance reordering method proposed in the SIARE-SC has good applicability and effectiveness, which can be used to be compatible with other SemCom systems and has a significant suppression effect on the peak-to-average power ratio (PAPR). Meanwhile, CEIE has been considered for the first time to be included in the analysis of SemCom distortion, and mathematically derive the performance expressions of SIARE-SC under different channel and pilot pattern scenarios from three perspectives, namely, channel bandwidth ratio (CBR), signal-to-noise ratio (SNR), and CEIE, to obtain the corresponding bound of performance. The proposed SIARE-SC is shown to significantly improve semantic performance in various scenarios by conducting a large number of experimental tests.
Chen Dong 0001, Haotai Liang, Weizhi Li, Zhicheng Bao, Xiaodong Xu 0001, Ping Zhang 0003
IEEE Internet Things J.5
2025 Brain-Like Cognition-Driven Model Factory for IIoT Fault Diagnosis by Combining LLMs With Small Models
abstract
Fault diagnosis is important for predictive maintenance in smart manufacturing, which involves intelligent human-machine interactions in order to make smart decisions for potential problems. Large language model (LLM) is promising in providing general artificial intelligence capabilities in this regard. However, LLM itself can not accurately analyze faults due to heterogeneous data from different Industrial Internet of Things (IIoT) devices in different processes during the complete production process. To accurately diagnose faults and facilitate human-machine interaction, this article proposes a brain-like cognition-driven model factory (BC-MF), using an LLM as a supervisor to adaptively generate personalized small-scale models according to the features of these heterogeneous data, where the vertical federated learning (VFL) idea is adopted. This BC-MF-based fault diagnosis approach includes a preliminary diagnosis phase and a precise diagnosis phase. The preliminary diagnosis is accomplished by prompting the LLM using a brain-like chain of thoughts (BLCoTs). A hypernetwork uses the preliminary diagnostic results and the feature maps trained by each node in the VFL to generate dedicated diagnostic small models and uses these models for final precise diagnostics. The LLM provides fault maintenance recommendations interactively according to the final diagnostic results. Comprehensive evaluations are conducted using four open IIoT datasets and one self-made dataset. It shows that the proposed BC-MF approach is significantly better than the existing approaches, in terms of model accuracy, comprehension of faults, and so on.
Yuru Liu, Weishan Zhang, Zhicheng Bao, Xudong Chai, Mu Gu, Fei-Yue Wang 0001
IEEE Internet Things J.3
2025 sDAC - Semantic Digital Analog Converter for Semantic Communications
abstract
In this paper, we propose a novel semantic digital analog converter (sDAC) for the compatibility between semantic and digital communications. Most of the current semantic communication systems rely primarily on analog modulation, limiting their integration with digital communication systems, which are more common in practice. In fact, traditional quantization methods are unsuitable for semantic communication because they do not account for semantic information within symbols. These factors block the wide application of the semantic communication. To address these challenges, sDAC is proposed. It is a simple yet efficient and generative module used to realize digital and analog bi-directional conversion. The entire process is independent of any specific semantic model, modulation methods, or channel conditions. In the experiment section, the performance of sDAC is tested across different semantic models, semantic tasks, modulation methods, channel conditions and quantization orders. Test results show that the proposed sDAC has great generative properties and channel robustness.
Zhicheng Bao, Haotai Liang, Chen Dong 0001, Xiaodong Xu 0001, Cheng Guo 0004, Hao Chen 0013, Ping Zhang 0003
IEEE Trans. Commun.1
2025 Adaptive Bitrate Video Semantic Increment Transmission System Based on Buffer and Semantic Importance
abstract
Significant progress has been made in researching video semantic communication technology and adaptive bitrate (ABR) algorithms. However, wireless network fluctuations challenge video semantic communication systems without ABR algorithms to achieve a satisfactory balance between high semantic recovery accuracy and efficient bandwidth utilization. This paper proposes an adaptive bitrate video semantic increment transmission system based on buffer and semantic importance to address this issue. Firstly, a buffer-based video semantic increment transmission system is designed to dynamically adjust the amount of video semantic data transmitted by the transmitter based on network fluctuations. Then, a novel Deep Learning and Reinforcement Learning based ABR algorithm (DR-ABR) is developed to determine the optimal video incremental ratio under the current network conditions. Furthermore, a semantic feature compression technology based on semantic importance is proposed to compress the video data according to the abovementioned ratio. Experimental results demonstrate that the proposed method outperforms traditional approaches in terms of video semantic transmission performance.
Zhicheng Bao, Haotai Liang, Chen Dong 0001, Xiaodong Xu 0001, Lin Li 0062
IEEE Trans. Netw. Serv. Manag.2
2024 Entropy-Based Importance Reordering method for Mitigating Distortion in Slow Fading Channels
abstract
Semantic communication, as a new research paradigm, has garnered widespread attention from academia and industry. One of the important aspects is the study of channel estimation, which can further improve the recovery of signals in communication systems. However, most existing studies on semantic communication have only considered the case of perfect channel estimation. In this paper, pilots-assisted channel estimation is considered, and a symbol reordering method named Entropy-Based Importance Reordering (EBIR) is proposed to mitigate the distortions induced by slow fading channels. The method distinguishes important and unimportant semantic sym-bols based on the entropy value obtained from the entropy model. Based on the characteristics of channel estimation in slow fading time-varying channels, important semantic symbols are reassigned to improve signal recovery further. The results show that the effectiveness and universality of EBIR are validated for different sources, channel bandwidth ratios (CBRs) and channel states.
Zhicheng Bao, Haotai Liang, Chen Dong 0001, Xiaodong Xu 0001
WCNC2
2024 A Relay System for Semantic Image Transmission Based on Shared Feature Extraction and Hyperprior Entropy Compression
abstract
Nowadays, the need for high-quality image reconstruction and restoration is more and more urgent. However, most image transmission systems may suffer from image quality degradation or transmission interruption in the face of interference such as channel noise and link fading. To solve this problem, a relay communication network for semantic image transmission based on shared feature extraction and hyperprior entropy compression (HEC) is proposed, where the shared feature extraction technology based on Pearson correlation is proposed to eliminate partial shared feature of extracted semantic latent feature. In addition, the HEC technology is used to resist the effect of channel noise and link fading and carried out respectively at the source node and the relay node. Experimental results demonstrate that compared with other recent research methods, the proposed system has lower transmission overhead and higher semantic image transmission performance. Particularly, under the same conditions, the multi-scale structural similarity (MS-SSIM) of this system is superior to the comparison method by approximately 0.2.
Wannian An, Zhicheng Bao, Haotai Liang, Chen Dong 0001, Xiaodong Xu 0001
IEEE Internet Things J.2
2023 CFSL: A Credible Federated Self-Learning Framework
abstract
Federated learning can collaboratively train AI models while protecting data privacy. In practical industry environment, non-independent and identically distributed (Non-IID) characteristics of data affect the effectiveness of federated learning. Personalized federated learning can help resolve this, but it cannot adapt to unknown data. In addition, practical applications also call for trusted training environment and remain stable when there are security threats. In this article, we propose a credible federated self-learning (CFSL), based on the idea of hypernetwork supported by blockchain to achieve secured, credible, personalized federated self-learning, especially, for unknown data in Non-IID environment. Extensive experiments on three Non-IID data sets demonstrate the capabilities on adaptive resilience for security attacks and on accuracy of recognizing unknown objects, with good performance at the same time. CFSL outperforms the existing personalized federated learning methods, with an increase in average accuracy by 4.11%.
Weishan Zhang, Zhicheng Bao, Yuru Liu, Liang Xu 0009, Qinghua Lu 0001, Huansheng Ning, Xiao Wang 0002, Su Yang 0001, Fei-Yue Wang 0001, Zengxiang Li
IEEE Internet Things J.2
2023 Feature-Contrastive Graph Federated Learning: Responsible AI in Graph Information Analysis
abstract
Federated learning enables multiple clients to learn a general model without sharing local data, and the federated learning system also improves information security and advances responsible artificial intelligence (AI). However, the data of different clients in the system are non-independently and identically distributed (IID), which results in weight divergence, especially for complex graph data extraction. This article proposes a novel feature-contrastive graph federated (FcgFed) learning approach to improve the robustness of the federated learning system in graph data. First, we design an architecture for FcgFed learning systems to analyze graph information. Furthermore, we present a graph federated learning method based on contrastive learning to alleviate the weight divergence in federated learning. The experiments in node classification and graph classification demonstrate that our method achieves better performance than model-contrastive federated learning (MOON) and federated average (FedAvg). We also test the adaptability of our method in image classification, and the results demonstrate that weight similarity evaluation works for other frameworks and tasks.
Xingjie Zeng, Zhicheng Bao, Leiming Chen, Xiao Wang 0002, Fei-Yue Wang 0001
IEEE Trans. Comput. Soc. Syst.3
2022 DTCN: Dynamic Temporal Convolution Network for Evaluating Dividing Coefficients of Water Well
abstract
Multi-source data fusion is widely utilized to enrich the dataset for artificial intelligence methods. However, it also suffers the limitation that samples should have same format and can not be applied to some tasks where input has a different dimension. In the case of dividing coefficients evaluation on water well, the number of injection layers is variable based on the general injection plan of a field. No exiting methods can learn the input with dynamic data format. In order to address this challenge, we propose an intelligent water injection splitting method based on a dynamic temporal convolution network. Specifically, two improvements are proposed: 1) We design a dynamic activation strategy to build data groups according to the number of injection layers and the geographical relationship of each layer. 2) We design a dynamic temporal convolution network to evaluate the dividing coefficient with logging and production data in time series. The input has the data with different injection layers which enrich the samples for model training. We evaluate the model with real-world data from an oil field. The experimental results show the effectiveness of the model. We also compare it with CNN whose input is in the same format, the experimental results show that multi-source information fusion improves the accuracy.
Zhicheng Bao, Xingjie Zeng, Dakuang Han, Weishan Zhang
CSCWD1