VLDB 2026 Research / reviewers in the wild / expert
Peiwen Jiang
dblp:232/4024
· DBLP profile ↗
19ranked-venue papers
12as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 14 · 9 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic Communications With World Models
Peiwen Jiang, Jiajia Guo 0001, Chao-Kai Wen, Shi Jin 0002, Jun Zhang 0004 |
IEEE Trans. Commun. | 1 |
| 2026 | Foundation Model-Aided Channel-Adaptive Video Semantic Communication and Prototype ValidationabstractThe increasing demand for services such as live streaming and virtual reality places significant pressure on wireless communication systems. Enhancing system performance or reducing bandwidth consumption is critical for delivering high-quality video experiences. Semantic communication, which focuses on the transmission of meaning, offers a promising solution. However, existing approaches are often limited to single scenarios, rely on simple channels, lack adaptability to dynamic wireless environments, and remain untested in practical air interfaces. To address these challenges, we propose a foundation model-aided universal video semantic communication framework designed for pixel-wise reconstruction across diverse scenarios. This framework enables the transmission of entire videos using joint source-channel coding (JSCC) based on optical flow estimation and leverages multiple-input multiple-output orthogonal frequency division multiplexing (MIMO-OFDM) for efficient semantic delivery in 3rd generation partnership project (3GPP) standard channels. In scenarios requiring full transmission for regions of interest and selective transmission for other areas, the framework employs a foundation model for segmentation, followed by JSCC and delivery. Furthermore, we introduce a channel condition number-adaptive semantic remapping method based on an attention mechanism to mitigate the effects of wireless fading. To validate our approach, we implement the framework on a testbed and develop two online demonstrations. Simulations and over-the-air experiments confirm significant improvements in video quality and substantial reductions in bandwidth overhead compared to existing methods. Jiarun Ding, Peiwen Jiang, Chao-Kai Wen, Xiao Li 0001, Shi Jin 0002 |
IEEE Trans. Wirel. Commun. | 2 |
| 2026 | Position-Aided Semantic Communication for Efficient Image Transmission: Design, Implementation, and Experimental Results
Peiwen Jiang, Chao-Kai Wen, Shi Jin 0002, Jun Zhang 0004 |
IEEE Trans. Wirel. Commun. | 1 |
| 2026 | Adaptive Semantic Speech Transmission for High-Speed ScenariosabstractThe fast time-varying channels in high-speed scenarios impact signal transmission between transceivers and pose challenges to both the accuracy and bandwidth utilization of communication systems. Semantic communication, known for its ability to significantly reduce transmission bandwidth and enhance communication reliability, is especially effective in extreme environments. However, current semantic communication systems lack a comprehensive physical layer design, which limits their ability to achieve optimal performance in rapidly changing conditions. In this paper, we propose an adaptive semantic speech recognition and cloning transmission system with a superimposed pilot (SwitchAC-SIP) tailored for high-speed scenarios to ensure high-quality speech transmission. The system converts speech signals into textual content and speaker timbre features at the transmitter, while a speech cloning model reconstructs the speech at the receiver with a timbre closely resembling the original speaker based on these features, thereby eliminating the need to retrain the speech generation model for different users, ensuring both transmission quality and efficiency. To address the impact of high-speed environments on channel estimation performance, we introduce a superimposed pilot (SIP) in the physical layer. This method superimposes pilots and data across the entire time-frequency grid with a specific power ratio, significantly mitigating the detrimental effects of high-speed conditions on semantic communication systems. Furthermore, to enhance system flexibility in dynamic scenarios, we design a channel-adaptive network that dynamically allocates bandwidth ratios for text and audio semantics based on real-time channel conditions. This adaptive approach prioritizes the protection of critical semantic features according to user requirements. Simulation results demonstrate the substantial improvements in transmission efficiency and accuracy achieved by the proposed system. Peiwen Jiang, Wenjin Wang 0001, Xingyu Zhou 0011, Jing Zhang 0031, Chao-Kai Wen, Shi Jin 0002 |
IEEE Trans. Wirel. Commun. | 2 |
| 2025 | TST: A Schema-Based Top-Down and Dynamic-Aware Agent of Text-to-Table TasksabstractAs a bridge between natural texts and information systems like structured storage, statistical analysis, retrieving, and recommendation, the text-to-table task has received widespread attention recently. Existing researches have gone through a paradigm shift from traditional bottom-up IE (Information Extraction) to top-down LLMs-based question answering with RAG (Retrieval-Augmented Generation). Furthermore, these methods mainly adopt end-to-end models or use multi-stage pipelines to extract text content based on static table structures. However, they neglect to deal with precise inner-document evidence extraction and dynamic information such as multiple entities and events, which can not be defined in static table head format and are very common in natural texts.To address this issue, we propose a two-stage dynamic content extraction agent framework called TST (Text-Schema-Table), which uses type recognition methods to extract context evidences with the conduction of domain schema sequentially. Based on the evidence, firstly we quantify the total instances of each dynamic object and then extract them with ordered numerical prompts. Through extensive comparisons with existing methods across different datasets, our extraction framework exhibits state-of-the-art (SOTA) performance. Our codes are available at https://github.com/jiangpw41/TST. Peiwen Jiang, Haitong Jiang, Ruhui Ma, Yvonne Jie Chen, Jinhua Cheng |
ACL (1) | 1 |
| 2025 | Semantic Satellite Communications Based on Generative Foundation ModelabstractSatellite communications can provide massive connections and seamless coverage, but they also face several challenges, such as rain attenuation, long propagation delays, and co-channel interference. To improve transmission efficiency and address severe scenarios, semantic communication has become a popular choice, particularly when equipped with foundation models (FMs). In this study, we introduce an FM-based semantic satellite communication framework, termed FMSAT. This framework leverages FM-based segmentation and reconstruction to significantly reduce bandwidth requirements and accurately recover semantic features under high noise and interference. Considering the high speed of satellites, an adaptive encoder-decoder is proposed to protect important features and avoid frequent retransmissions. Meanwhile, a well-received image can provide a reference for repairing damaged images under sudden attenuation. Since acknowledgment feedback is subject to long propagation delays when retransmission is unavoidable, a novel error detection method is proposed to roughly detect semantic errors at the regenerative satellite. With the proposed detectors at both the satellite and the gateway, the quality of the received images can be ensured. The simulation results demonstrate that the proposed method can significantly reduce bandwidth requirements, adapt to complex satellite scenarios, and protect semantic information with an acceptable transmission delay. Peiwen Jiang, Chao-Kai Wen, Xiao Li 0001, Shi Jin 0002, Geoffrey Ye Li |
IEEE J. Sel. Areas Commun. | 1 |
| 2025 | PD-CEViT: A Novel Pilot Pattern Design and Channel Estimation Network for OFDM SystemsabstractDeep learning has been widely applied to channel estimation (CE), yielding significant performance improvements. However, existing research primarily focuses on static channel scenarios, leading to substantial performance degradation in dynamic environments. Furthermore, the use of fixed pilot patterns fails to adequately capture channel dynamics, resulting in unnecessary pilot overhead. In this study, we propose a Vision Transformer-based joint pilot design (PD) and CE network (PD-CEViT) for orthogonal frequency division multiplexing (OFDM) systems. The PD module leverages maximum Doppler shift and delay spread information to determine pilot positions, effectively capturing channel variations in dynamic scenarios. To further improve CE accuracy and robustness across diverse environments, the coarse CE from the PD module is passed to a CE module that utilizes a Vision Transformer (ViT), forming the joint PD-CEViT structure. Additionally, we introduce a pilot number switch network, named SwitchPD-CEViT, which dynamically adjusts between different PD-CEViT configurations based on the current channel conditions. This strategy balances network performance and pilot overhead, accommodating varying pilot requirements across different scenarios. Simulation results demonstrate that our proposed structure more effectively tracks channel variations compared to fixed pilot patterns. Even under challenging conditions with large Doppler shifts and delay spreads, our method significantly outperforms traditional and deep learning approaches in terms of mean square error (MSE) performance. Moreover, the integration of channel information further enhances estimation performance and robustness. Meanwhile, SwitchPD-CEViT achieves superior CE performance with reduced pilot overhead by efficiently managing pilot utilization. Peiwen Jiang, Jing Zhang 0031, Wenjin Wang 0001, Chao-Kai Wen, Shi Jin 0002 |
IEEE Trans. Commun. | 2 |
| 2025 | Hybrid Matching Teacher Framework for Cross-Domain Visual Detection TransformerabstractObject detection is a critical component of autonomous vehicle perception systems. However, domain shifts between training environments and real-world scenarios often degrade detector performance. Cross-domain object detection aims to adapt detectors to unlabeled target domains utilizing only labeled source data. Recent popular cross-domain object detection methods employ the mean teacher framework, which uses pseudo-labels generated by the teacher model to guide training on unlabeled real-world data. Despite its effectiveness, continuous training with noisy pseudo-labels leads to abnormal performance degradation in the later stages of training. To address this issue, we propose a novel Hybrid Matching Teacher (HMT) framework for cross-domain visual detection transformers, which enhances cross-domain knowledge transfer across pseudo-label generation, filtering, and training processes. Specifically, we design a Feature Sparse Alignment (FSA) module to adapt DETR tokens and queries, generate domain-adaptive weights to initialize the teacher-student models, and mitigate the inherent initial source bias in the teacher model. Next, a Localization-aware Pseudo-label Filtering (LPF) module ensures high-quality pseudo-labels by considering the consistency between localization and classification tasks. Furthermore, to improve the efficiency of pseudo-label training, the Cross-view Hybrid Matching (CHM) module introduces an auxiliary matching branch to increase the number of positive queries that match with pseudo-labels. Extensive experiments demonstrate that our approach achieves state-of-the-art performance, outperforming previous benchmarks by 3.1%, 8.5%, and 4.4% in adverse weather, diverse scenes, and synthetic-to-real, respectively. Xiaowei Wang 0001, Jinhui Suo, Yang Li 0093, Ming Gao 0012, Peiwen Jiang, Pengwen Dai |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Generative Diffusion Models for High Dimensional Channel EstimationabstractAlong with the prosperity of generative artificial intelligence (AI), its potential for solving conventional challenges in wireless communications has also surfaced. Inspired by this trend, we investigate the application of the advanced diffusion models (DMs), a representative class of generative AI models, to high dimensional wireless channel estimation. By capturing the structure of multiple-input multiple-output (MIMO) wireless channels via a deep generative prior encoded by DMs, we develop a novel posterior inference method for channel reconstruction. We further adapt the proposed method to recover channel information from low-resolution quantized measurements. Additionally, to enhance the over-the-air viability, we integrate the DM with the unsupervised Stein’s unbiased risk estimator to enable learning from noisy observations and circumvent the requirements for ground truth channel data that is hardly available in practice. Results reveal that the proposed estimator achieves high-fidelity channel recovery while reducing estimation latency by a factor of 10 compared to state-of-the-art schemes, facilitating real-time implementation. Moreover, our method outperforms existing estimators while reducing the pilot overhead by half, showcasing its scalability to ultra-massive antenna arrays. Xingyu Zhou 0011, Le Liang, Jing Zhang 0031, Peiwen Jiang, Shi Jin 0002 |
IEEE Trans. Wirel. Commun. | 4 |
| 2024 | TKGT: Redefinition and A New Way of Text-to-Table Tasks Based on Real World Demands and Knowledge Graphs Augmented LLMsabstractThe task of text-to-table receives widespread attention, yet its importance and difficulty are underestimated.Existing works use simple datasets similar to table-to-text tasks and employ methods that ignore domain structures.As a bridge between raw text and statistical analysis, the text-to-table task often deals with complex semi-structured texts that refer to specific domain topics in the real world with entities and events, especially from those of social sciences.In this paper, we analyze the limitations of benchmark datasets and methods used in the text-to-table literature and redefine the textto-table task to improve its compatibility with long text-processing tasks.Based on this redefinition, we propose a new dataset called CPL (Chinese Private Lending), which consists of judgments from China and is derived from a real-world legal academic project.We further propose TKGT (Text-KG-Table), a two stages domain-aware pipeline, which firstly generates domain knowledge graphs (KGs) classes semiautomatically from raw text with the mixed information extraction (Mixed-IE) method, then adopts the hybrid retrieval augmented generation (Hybird-RAG) method to transform it to tables for downstream needs under the guidance of KGs classes.Experiment results show that TKGT achieves state-of-the-art (SOTA) performance on both traditional datasets and the CPL.Our data and main code are available at https://github.com/jiangpw41/TKGT. Peiwen Jiang, Xinbo Lin, Ruhui Ma, Yvonne Jie Chen, Jinhua Cheng |
EMNLP | 1 |
| 2024 | RIS-Enhanced Semantic Communications Adaptive to User RequirementsabstractSemantic communication, through the interpretation of the semantic meaning of transmitted data, effectively reduces the required bandwidth. However, current deep learning-based methods face limitations due to their reliance on joint source-channel coding and end-to-end training, hindering adaptability to new channels and user demands. In this study, we introduce the Reconfigurable Intelligent Surface-Semantic Communication (RIS-SC) framework as a solution. This framework dynamically allocates semantic content, leveraging varying degrees of RIS assistance to cater to the evolving needs of users. It takes into account factors such as user mobility and obstacles in the line of sight, enabling the RIS resource to preserve essential semantics even in challenging channel conditions. While this ensures the preservation of core semantics in difficult channel conditions, it may also lead to the loss of some non-essential semantic details under extreme conditions. To counteract this, we have incorporated a reconstruction method that deduces the missing semantic elements, thereby enhancing visual understanding. The RIS-SC framework stands out for its adaptability, ensuring optimal resource distribution for users under favorable conditions and maintaining visual clarity in challenging scenarios. Simulations validate the effectiveness and adaptability of our approach in diverse channel conditions and user demands. Peiwen Jiang, Chao-Kai Wen, Shi Jin 0002, Geoffrey Ye Li |
IEEE Trans. Commun. | 1 |
| 2024 | SLoB: Suboptimal Load Balancing Scheduling in Local Heterogeneous GPU Clusters for Large Language Model InferenceabstractLarge language models (LLMs) are becoming powerful engines for social productivity in the manufacturing lifecycle. Existing application-level LLMs inference services focus on large datacenter and small edge intelligence (EI) scenarios, adopting iteration-level batch schedulers to solve resource utilization and inference speed problems. However, these services are incompatible with the scene of medium-sized local heterogeneous graphics processing unit (GPU) clusters with specific patterns, whose scale is between the two aforementioned scenarios. This type of scene proposes tradeoff problems for inference resource and speed, as well as user satisfaction problems for the semisparse frequency of queries with streaming responses. We propose suboptimal load balancing (SLoB), a distributed LLMs inference service scheduler in medium-sized local heterogeneous GPU clusters. SLoB leverages a multilevel adapter to accommodate LLMs usage patterns of scenes and balance resource utilization with inference efficiency. For semisparse problems, it adopts a mixed-priority pipeline scheduler with the least-padding principle to improve users’ satisfaction, a metric considering the weights of different tokens in streaming responses. Based on the system prototype, our experiments under simulated workloads demonstrate that SLoB gains a maximum improvement of 29.4$\times$under the satisfaction metric compared with the traditional run-to-completion scheduling solution while improving by up to 3.0$\times$compared with the state-of-the-art (SOTA) solution Orca. Peiwen Jiang, Haoxin Wang 0005, Zinuo Cai, Lintao Gao, Weishan Zhang, Ruhui Ma, Xiaokang Zhou |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2024 | Progressive Critical Region Transfer for Cross-Domain Visual Object DetectionabstractWell-trained visual object detectors are generally confronted with a severe performance decline when deployed in a novel driving scenario due to the impact of domain shift. Despite excellent improvements in unsupervised domain adaptive object detection achieved by adversarial training, those approaches fail to capture the transfer core underlying the holistic scenes. To solve this problem, we propose a progressive critical region transfer framework for cross-domain visual object detection. Specifically, we exploit a potential foreground mining (PFM) module and a semantic-specific RoI aggregation (SRA) module to improve the robustness of the cross-domain detection framework. Upon the critical regions in the broad sense, the PFM module first highlights the foreground regions by reweighting the hierarchical feature maps in sequence, and then modifies location biases at the downstream position of the backbone network for more accurate upstream predictions. Deep into the critical regions in the narrow sense, the SRA module concentrates on establishing an appropriate matching between batch-wise RoIs and all semantic centers, and further strengthens the aggregation of cross-domain identical semantic with the complement of context references. Together these modules are obligated to transform the adaptation importance from the whole scope to the latent foreground areas, and afterward to the informative regions of interest along the detection pipeline. Experiments show that our progressive critical region transfer framework achieves a state-of-the-art performance in adverse weather, camera configuration, and complicated scene adaptation, which outperforms the baselines by 19.4%, 5.0%, and 6.1%, respectively. Xiaowei Wang 0001, Peiwen Jiang, Yang Li 0093, Manjiang Hu, Ming Gao 0012, Dongpu Cao, Rongjun Ding |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | RIS-Enhanced Semantic Image Transmission Based on Reinforcement LearningabstractSemantic communication can significantly reduce transmission payload by sending only semantic information related to the task. However, existing end-to-end trained semantic studies degrade under extreme channel environments, while reconfigurable intelligent surface (RIS) technology offers a potential solution for realizing channel customization. In this work, we propose a reconfigurable RIS-enhanced semantic communication framework called RIS-SC. This framework allows for customization of the channel environment based on the user's requirements for different semantic parts, rather than relying solely on the conventional bit error rate requirement. Using reinforcement learning, the RIS controller interacts with varying channels to meet the user's different requirements. The RIS controller adaptively protects important semantic parts by adjusting the channel conditions. Simulation results demonstrate that the proposed RIS-SC framework can adapt to different channel environments and improve task performance under varying requirements, such as vertical semantic or true image reconstruction. Peiwen Jiang, Chao-Kai Wen, Shi Jin 0002, Xiao Li 0001, Geoffrey Ye Li |
GLOBECOM | 1 |
| 2023 | CE-ViT: A Robust Channel Estimator Based on Vision Transformer for OFDM SystemsabstractDeep learning (DL) has been widely utilized for channel estimation and has resulted in significant performance improvements. However, most existing research only performs training and testing in relatively static scenarios, leading to a serious deterioration in dynamic scenarios. In this paper, we propose a robust channel estimator for orthogonal frequency-division multiplexing (OFDM) systems in dynamic scenarios called channel estimator Vision Transformer (CE-ViT) based on attention mechanism. We perform a patch embedding operation to process data in both the time and frequency domains, addressing the limitations of the attention mechanism in extracting 2D correlations. Additionally, we introduce tokens that reflect channel characteristics into the network to enhance the robustness. Experimental results show that CE-ViT outperforms the state-of-the-art DL-based methods. Moreover, the addition of tokens significantly improves the performance of CE-ViT in dynamic channel conditions. Jing Zhang 0031, Peiwen Jiang, Chao-Kai Wen, Shi Jin 0002 |
GLOBECOM | 3 |
| 2023 | Wireless Semantic Communications for Video ConferencingabstractVideo conferencing has become a popular mode of meeting despite consuming considerable communication resources. Conventional video compression causes resolution reduction under a limited bandwidth. Semantic video conferencing (SVC) maintains a high resolution by transmitting some keypoints to represent the motions because the background is almost static, and the speakers do not change often. However, the study on the influence of transmission errors on keypoints is limited. In this paper, an SVC network based on keypoint transmission is established, which dramatically reduces transmission resources while only losing detailed expressions. Transmission errors in SVC only lead to a changed expression, whereas those in the conventional methods directly destroy pixels. However, the conventional error detector, such as cyclic redundancy check, cannot reflect the degree of expression changes. To overcome this issue, an incremental redundancy hybrid automatic repeat-request framework for varying channels (SVC-HARQ) incorporating a novel semantic error detector is developed. SVC-HARQ has flexibility in bit consumption and achieves a good performance. In addition, SVC-channel state information (CSI) is designed for CSI feedback to allocate the keypoint transmission and enhance the performance dramatically. Simulation shows that the proposed wireless semantic communication system can remarkably improve transmission efficiency. Peiwen Jiang, Chao-Kai Wen, Shi Jin 0002, Geoffrey Ye Li |
IEEE J. Sel. Areas Commun. | 1 |
| 2022 | Deep Source-Channel Coding for Sentence Semantic Transmission With HARQabstractRecently, semantic communication has been brought to the forefront because deep learning (DL)-based methods, such as Transformer, have achieved great success in semantic extraction. Although semantic communication has been successfully applied in sentence transmission to reduce semantic errors, the existing architecture is usually fixed in terms of codeword length and inefficient and inflexible for varying sentence lengths. In this study, we exploit hybrid automatic repeat request (HARQ) to reduce the semantic transmission error further. We combine semantic coding (SC) with Reed-Solomon (RS) channel coding and HARQ (called SC-RS-HARQ). SC-RS-HARQ exploits the superiority of SC and the reliability of conventional methods successfully. Although SC-RS-HARQ can be easily applied in existing HARQ systems, we also develop an end-to-end architecture called SCHARQ to pursue enhanced performance. Numerical results demonstrate that SCHARQ significantly reduces the required number of bits for semantic sentence transmission and the sentence error rate. We also attempt to replace error detection from cyclic redundancy check to a similarity detection network called Sim32 to allow the receiver to reserve wrong sentences with similar semantic information and conserve transmission resources. Peiwen Jiang, Chao-Kai Wen, Shi Jin 0002, Geoffrey Ye Li |
IEEE Trans. Commun. | 1 |
| 2021 | Dual CNN-Based Channel Estimation for MIMO-OFDM SystemsabstractRecently, convolutional neural network (CNN)-based channel estimation (CE) for massive multiple-input multiple-output communication systems has achieved remarkable success. However, complexity even needs to be reduced, and robustness can even be improved. Meanwhile, existing methods do not accurately explain which channel features help the denoising of CNNs. In this paper, we first compare the strengths and weaknesses of CNN-based CE in different domains. When complexity is limited, the channel sparsity in the angle-delay domain improves denoising and robustness whereas large noise power and pilot contamination are handled well in the spatial-frequency domain. Thus, we develop a novel network, called dual CNN, to exploit the advantages in the two domains. Furthermore, we introduce an extra neural network, called HyperNet, which learns to detect scenario changes from the same input as the dual CNN. HyperNet updates several parameters adaptively and combines the existing dual CNNs to improve robustness. Experimental results show improved estimation performance for the time-varying scenarios. To further exploit the correlation in the time domain, a recurrent neural network framework is developed, and training strategies are provided to ensure robustness to the changing of temporal correlation. This design improves channel estimation performance but its complexity is still low. Peiwen Jiang, Chao-Kai Wen, Shi Jin 0002, Geoffrey Ye Li |
IEEE Trans. Commun. | 1 |
| 2021 | AI-Aided Online Adaptive OFDM Receiver: Design and Experimental ResultsabstractOrthogonal frequency division multiplexing (OFDM) has been widely applied in many wireless communi- cation systems. The artificial intelligence (AI)-aided OFDM receivers are currently brought to the forefront to replace and improve the traditional OFDM receivers. In this paper, we first compare two AI-aided OFDM receivers, namely, data-driven fully connected deep neural network and model-driven ComNet, through extensive simulation and real-time video transmission using a 5G rapid prototyping system for an over-the-air (OTA) test. We find a performance gap between the simulation and the OTA test caused by the discrepancy between the channel model for offline training and the real environment. We develop a novel online training system, which is called SwitchNet receiver, to address this issue. This receiver has a flexible and extendable architecture and can adapt to real channels by training only several parameters online. From the OTA test, the AI-aided OFDM receivers, especially the SwitchNet receiver, are robust to OTA environments and promising for future communication systems. At the end of this paper, we discuss potential challenges and future research inspired by our initial study in this paper. Peiwen Jiang, Xuanxuan Gao, Jing Zhang 0031, Chao-Kai Wen, Shi Jin 0002, Geoffrey Ye Li |
IEEE Trans. Wirel. Commun. | 1 |