Chengpei Tang

dblp:133/2795 · DBLP profile ↗
← Back
20ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0002-8139-6742ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Cost-Effective Communication: An Auction-based Method for Language Agent Interaction
abstract
Multi-agent systems (MAS) built on large language models (LLMs) often suffer from inefficient ''free-for-all'' communication, leading to exponential token costs and low signal-to-noise ratios that hinder their practical deployment. We challenge the notion that more communication is always beneficial, hypothesizing instead that the core issue is the absence of resource rationality. We argue that "free'' communication, by ignoring the principle of scarcity, inherently breeds inefficiency and unnecessary expenses. To address this, we introduce the Dynamic Auction-based Language Agent (DALA), a novel framework that treats communication bandwidth as a scarce and tradable resource. Specifically, our DALA regards inter-agent communication as a centralized auction, where agents learn to bid for the opportunity to speak based on the predicted value density of their messages. Thus, our DALA intrinsically encourages agents to produce concise, informative messages while filtering out low-value communication. Extensive and comprehensive experiments demonstrate that our economically-driven DALA achieves new state-of-the-art performance across seven challenging reasoning benchmarks, including 84.32% on MMLU and a 91.21% pass@1 rate on HumanEval. Note that this is accomplished with remarkable efficiency, i.e., our DALA uses only 6.25 million tokens, a fraction of the resources consumed by current state-of-the-art methods on GSM8K. Further analysis reveals that our DALA cultivates the emergent skill of strategic silence, effectively adapting its communication strategies from verbosity to silence in a dynamic manner via resource constraints.
Yijia Fan, Jusheng Zhang, Kaitong Cai, Chengpei Tang, Jian Wang 0100, Keze Wang
AAAI5
2026 HiVA: Self-organized Hierarchical Variable Agent via Goal-driven Semantic-Topological Evolution
abstract
Autonomous agents play a crucial role in advancing Artificial General Intelligence, enabling problem decomposition and tool orchestration through Large Language Models (LLMs). However, existing paradigms face a critical trade-off. On one hand, reusable fixed workflows require manual reconfiguration upon environmental changes; on the other hand, flexible reactive loops fail to distill reasoning progress into transferable structures. We introduce Hierarchical Variable Agent (HiVA), a novel framework modeling agentic workflows as self-organized graphs with the Semantic-Topological Evolution (STEV) algorithm, which optimizes hybrid semantic-topological spaces using textual gradients as discrete-domain surrogates for backpropagation. The iterative process comprises Multi-Armed Bandit-infused forward routing, diagnostic gradient generation from environmental feedback, and coordinated updates that co-evolve individual semantics and topology for collective optimization in unknown environments. Experiments on dialogue, coding, Long-context Q&A, mathematical, and agentic benchmarks demonstrate improvements of 5-10% in task accuracy and enhanced resource efficiency over existing baselines, establishing HiVA's effectiveness in autonomous task execution.
Jinzhou Tang, Jusheng Zhang, Qinhan Lv, Sidi Liu, Chengpei Tang, Keze Wang
AAAI6
2026 Top-Down Semantic Refinement for Image Captioning
abstract
Large Vision-Language Models (VLMs) face an inherent contradiction in image captioning: their powerful single-step generation capabilities often lead to a myopic decision-making process. This makes it difficult to maintain global narrative coherence while capturing rich details, a limitation that is particularly pronounced in tasks that require multi-step and complex scene description. To overcome this fundamental challenge, we redefine image captioning as a goal-oriented hierarchical refinement planning problem, and further propose a novel framework, named Top-Down Semantic Refinement (TDSR), which models the generation process as a Markov Decision Process (MDP). However, planning within the vast state space of a VLM presents a significant computational hurdle. Our core contribution, therefore, is the design of a highly efficient Monte Carlo Tree Search (MCTS) algorithm tailored for VLMs. By incorporating a visual-guided parallel expansion and a lightweight value network, our TDSR reduces the call frequency to the expensive VLM by an order of magnitude without sacrificing planning quality. Furthermore, an adaptive early stopping mechanism dynamically matches computational overhead to the image's complexity. Extensive experiments on multiple benchmarks, including DetailCaps, COMPOSITIONCAP, and POPE, demonstrate that our TDSR, as a plug-and-play module, can significantly enhance the performance of existing VLMs (e.g., LLaVA-1.5, Qwen2.5-VL) by achieving state-of-the-art or highly competitive results in fine-grained description, compositional generalization, and hallucination suppression.
Jusheng Zhang, Kaitong Cai, Jian Wang 0100, Chengpei Tang, Keze Wang
AAAI5
2026 MTFuzz: A Novel Efficacy Fuzzing Framework for Aerospace Monolithic Firmware
Shuai Wang 0012, Xi Xiao 0001, Guangwu Hu, Kehuan Zhang, Le Yu 0002, Chengpei Tang, Qing Li 0006, Qizhen Xu
DSN6
2026 A Hybrid Mamba-Transformer Approach With Time-Frequency Fusion Attention for Fall Detection
abstract
With the rapid increase of the aging population, fall detection, as a key technology of intelligent medical treatment, has attracted extensive attention. Compared with the poor comfort of wearable devices and the light sensitivity of visual methods, the wireless sensing scheme based on channel state information (CSI) shows unique advantages with its non-contact and privacy friendly. However, the existing CSI-based methods have the problem of insufficient long-range modeling ability, and fail to establish the deep semantic correlation between time-domain and frequency-domain in the fall process. Therefore, this paper proposes a fall detection method based on hybrid Mamba and Transformer, namely HMT-Fall. Different from the existing methods that simply parallel the time-domain network and the frequency-domain network, HMT-Fall creatively constructs a multi-domain modeling mechanism with clear functional division and collaborative design. Firstly, Mamba network is introduced to process the raw CSI sequences, and its selective state space mechanism is used to achieve efficient long-range timing modeling. Secondly, the short-time Fourier transform (STFT) spectrum of CSI is analyzed by Swin Transformer, and the frequency domain representation with both local details and global context is efficiently extracted through the shifted window self-attention mechanism. Furthermore, a bidirectional cross-attention fusion module is designed to achieve dynamic alignment and mutual constraint between time-domain features and frequency-domain features at the semantic level, rather than simple static splicing or weighting, so as to form a physically consistent and more discriminative joint representation. The experimental results on the self-built dataset HMT-HAR and public datasets show that the detection performance of HMT-Fall is significantly better than that of the existing representative methods, achieving over 99% accuracy, which verifies the effectiveness and superiority of the proposed method.
Jiong Liang, Yingping Wang, Shaolin Liao, Chengpei Tang
IEEE Internet Things J.6
2026 Learning Prompt Adapters for Forgetting-Free Continual Image Super-Resolution
abstract
Continual image super-resolution (CISR) aims to efficiently adapt a pre-trained model to a variety of tasks while retaining knowledge from previously learned tasks, minimizing the need for intensive independent training. The primary challenges include catastrophic forgetting due to varying data distributions and degradation types, along with the necessity for high adaptability. While prompt-based continual learning has proven effective in image classification, its direct application to super-resolution (SR) often fails to meet the demands for detailed pixel-level restoration and domain discrimination in low-level characteristics. To address these challenges, we propose Learning Prompt Adapters (LPA), which dynamically generates pixel-wise prompts through a combination of multi-granularity prompt bases and identities. By adaptively integrating these prompts into the Transformer architecture, we effectively improve the model's performance on fine-grained details in super-resolution tasks, as well as enhancing the model's adaptability to new tasks and preserving knowledge from previous ones. Through organizing the low-rank prompt bases with specific identities, we set up an effective solution to managing cross-task differences and enhancing prompt richness. Extensive experiments on benchmarks comprising the NYU, RealSR, DIV2K, REDS, and MANGA109 datasets with diverse degradation types demonstrate that LPA significantly outperforms existing continual learning methods. Codes of this paper are available at: https://github.com/dummerchen/LPA.
Chaowei Fang, Bolin Fu, De Cheng, Chengpei Tang, Guanbin Li
IEEE Trans. Image Process.4
2026 IMRadar: Bidirectional Velocity Mamba for Contactless Human Behavior Sensing
abstract
In recent years, intelligent human behavior sensing based on channel state information (CSI) has garnered significant attention from researchers, serving as a pivotal application of contactless health monitoring. However, the feature extraction networks used in existing perception schemes have significant limitations in terms of global context perception, computational complexity, and only consider features in one direction. To address these issues, this article proposes a novel bidirectional velocity Mamba (BVMamba) model and constructs an intelligent behavior sensing system, named IMRadar. The system first analyzes the velocity information that better characterizes the human motion state from CSI data, and uses the BVMamba model to extract global deep behavioral features from both forward and reverse directions. The BVMamba model includes forward velocity Mamba block (FVMamba), reverse velocity Mamba block (RVMamba), and bidirectional velocity feature fusion block (FUBlock), which can comprehensively capture the dynamic characteristics of complex behaviors. Experiments have shown that IMRadar exhibits excellent recognition performance on both publicly available datasets (ARIL, Widar) and self-built dataset (IM-HAR), with accuracy rates exceeding 98% for all datasets, providing an efficient and robust solution for non-contact behavior perception technology.
Jiong Liang, Shaolin Liao, Henry Soekmadji, Chengpei Tang
IEEE J. Biomed. Health Informatics6
2025 DrDiff: Dynamic Routing Diffusion with Hierarchical Attention for Breaking the Efficiency-Quality Trade-off
abstract
This paper introduces DrDiff, a novel framework for long-text generation that overcomes the efficiency-quality trade-off through three core technologies.First, we design a dynamic expert scheduling mechanism that intelligently allocates computational resources during the diffusion process based on text complexity, enabling more efficient handling of text generation tasks of varying difficulty.Second, we introduce a Hierarchical Sparse Attention (HSA) mechanism that adaptively adjusts attention patterns according to a variety of input lengths, reducing computational complexity from O(n 2 ) to O(n) while maintaining model performance.Finally, we propose a Semantic Anchor States (SAS) module that combines with DPM-solver++ to reduce diffusion steps, significantly improving generation speed.Comprehensive experiments on various long-text generation benchmarks demonstrate the superiority of our DrDiff over the existing SOTA methods.
Jusheng Zhang, Yijia Fan, Kaitong Cai, Zimeng Huang, Jian Wang 0100, Chengpei Tang, Keze Wang
EMNLP7
2025 A Vehicle Path Planning and Prediction Algorithm Based on Attention Mechanism for Complex Traffic Intersection Collaboration in Intelligent Transportation
abstract
In the development of smart cities, the transportation system plays a crucial role, with road congestion being particularly prominent under conditions of long-distance travel and high traffic volumes. This paper proposes a Vehicle Path Planning and Prediction Algorithm (VPPPA) based on an attention mechanism for complex traffic intersections collaboration in intelligent transportation systems. Our proposed algorithm is designed to plan and analyze traffic flow at urban road intersections. Firstly, attention mechanisms are used to balance the number of vehicles at different traffic intersections, with a particular focus on alleviating congestion at critical intersections during peak hours. Secondly, a Convolutional Neural Network (CNN) is employed to capture the spatial relationships between different road segments and intersections. Moreover, the Long Short-Term Memory and CNN (LSTM-CNN) architecture effectively captures the important temporal correlations in the traffic flow data. Thirdly, the spatiotemporal attention mechanism of vehicles captures the local spatial correlation characteristics between the target intersection and adjacent intersections along the traffic network. Finally, our proposed VPPPA model leverages the advantages of the LSTM-CNN architecture, enhancing learning efficiency during the training process and extracting valuable information. Experimental results show that the proposed VPPPA has significant advantages and greater efficiency in reducing average travel time and improving throughput across various intersections.
Yan Li 0124, Chengpei Tang
IEEE Trans. Intell. Transp. Syst.3
2024 Gesture Generation Via Diffusion Model with Attention Mechanism
abstract
Generating natural and semantically aligned gestures from speech remains a challenging task in human-computer interaction due to the intricate relationship between speech and gestures. While recent advances in learning-based methodologies have shown progress, they exhibit limitations like limited diversity and fidelity, as well as a mismatch between generated gestures and the semantic and emotional context, impacting efficacy in conveying information. To address these multifaceted challenges, this study introduces Gesture Diffusion Attention (GDA), an innovative approach for generating gestures from spoken language. Diverging from conventional methods, GDA incorporates a sophisticated denoising diffusion probability module, progressively transforming simplistic probability distributions into more complex ones, consequently yielding a repertoire of natural and diverse gestures. Furthermore, the utilization of pretrained fastText models for textual feature extraction, coupled with attention mechanisms, ensures that generated gestures align with speech in terms of semantic content and emotional nuances. To empirically validate the efficacy of the proposed approach, a series of rigorous objective experiments were conducted. The results demonstrate the exceptional performance of GDA in generating natural and diversified gestures that accurately and coherently convey the intended information, surpassing the benchmarks established by traditional methods. Code is released at https://github.com/LEELLL/GDA-icassp2024.
Qiyuan Ding, Chengpei Tang, Keze Wang
ICASSP4
2024 Routing User-Interest Markov Tree for Scalable Personalized Knowledge-Aware Recommendation
abstract
To facilitate more accurate and explainable recommendation, it is crucial to incorporate side information into user-item interactions. Recently, knowledge graph (KG) has attracted much attention in a variety of domains due to its fruitful facts and abundant relations. However, the expanding scale of real-world data graphs poses severe challenges. In general, most existing KG-based algorithms adopt exhaustively hop-by-hop enumeration strategy to search all the possible relational paths, this manner involves extremely high-cost computations and is not scalable with the increase of hop numbers. To overcome these difficulties, in this article, we propose an end-to-end framework Knowledge-tree-routed UseR-Interest Trajectories Network (KURIT-Net). KURIT-Net employs the user-interest Markov trees (UIMTs) to reconfigure a recommendation-based KG, striking a good balance for routing knowledge between short-distance and long-distance relations between entities. Each tree starts from the preferred items for a user and routes the association reasoning paths along the entities in the KG to provide a human-readable explanation for model prediction. KURIT-Net receives entity and relation trajectory embedding (RTE) and fully reflects potential interests of each user by summarizing all reasoning paths in a KG. Besides, we conduct extensive experiments on six public datasets, our KURIT-Net significantly outperforms state-of-the-art approaches and shows its interpretability in recommendation.
Yongsen Zheng, Pengxu Wei, Ziliang Chen 0001, Chengpei Tang, Liang Lin 0004
IEEE Trans. Neural Networks Learn. Syst.4
2024 Train Once, Locate Anytime for Anyone: Adversarial Learning-based Wireless Localization
abstract
Among numerous indoor localization systems, WiFi fingerprint-based localization has been one of the most attractive solutions, which is known to be free of extra infrastructure and specialized hardware. To push forward this approach for wide deployment, three crucial goals on high deployment ubiquity, high localization accuracy, and low maintenance cost are desirable. However, due to severe challenges about signal variation, device heterogeneity, and database degradation root in environmental dynamics, pioneer works usually make a trade-off among them. In this article, we propose iToLoc, a deep learning-based localization system that achieves all three goals simultaneously. Once trained, iToLoc will provide accurate localization service for everyone using different devices and under diverse network conditions, and automatically update itself to maintain reliable performance anytime. iToLoc is purely based on WiFi fingerprints without relying on specific infrastructures. The core components of iToLoc are a domain adversarial neural network and a co-training-based semi-supervised learning framework. Extensive experiments across 7 months with eight different devices demonstrate that iToLoc achieves remarkable performance with an accuracy of 1.92 m and >95% localization success rate. Even 7 months after the original fingerprint database was established, the rate still maintains >90%, which significantly outperforms previous works.
Danyang Li 0005, Jingao Xu, Zheng Yang 0002, Chengpei Tang
ACM Trans. Sens. Networks4
2022 WiFi-Based Cross-Domain Gesture Recognition via Modified Prototypical Networks
abstract
Numerous deep learning studies have achieved remarkable advances in WiFi-based human gesture recognition (HGR) using channel state information (CSI). However, since the CSI patterns of the same gesture change across domains (i.e., users, environments, locations, and orientations), recognition accuracy might degrade significantly when applying the trained model to new domains. To overcome this problem, we propose a WiFi-based cross-domain gesture recognition system (WiGr) which has a domain-transferable mapping to construct an embedding space where the representations of samples from the same class are clustered, and those from different classes are separated. The key insight of WiGr is using the similarity between the query sample representation and the class prototypes in the embedding space to perform the gesture classification, which can avoid the influence of the cross-domain CSI patterns change. Meanwhile, we present a dual-path prototypical network (Dual-Path PN) which consists of a deep feature extractor and a dual-path (i.e., Path-A and Path-B substructures) recognizer. The trained feature extractor can extract the gesture-related domain-independent features from CSI, namely, the domain-transferable mapping. In addition, WiGr implements the cross-domain HGR based on only a pair of WiFi devices without retraining in the new domain. We conduct comprehensive experiments on three data sets, one is built by ourselves and the others are public data sets. The evaluation suggests that WiGr achieves 86.8%–92.7% in-domain recognition accuracy and 83.5%–93% cross-domain accuracy under the four-shot condition.
Xie Zhang, Chengpei Tang, Qingqian Ni
IEEE Internet Things J.2
2022 Intelligent Bus Operation Optimization by Integrating Cases and Data Driven Based on Business Chain and Enhanced Quantum Genetic Algorithm
abstract
Intelligent public transport systems are a key direction for the study of intelligent transportation systems (ITSs), and they perform the following functions: location and tracking, aided navigation, dispatch and command, and dynamic information release; these systems also help travelers determine optimal routes. This paper mainly studies intelligent bus operation optimization by integrating cases and data from business chains, including the optimization of public transport vehicle scheduling according to the characteristics of vehicle scheduling, considers the interests of passengers and public transport companies, and adopts a true-value encoding method with the departure time as the variable goal optimization. In addition, this paper builds a model for the intelligent scheduling problem in public transport based on an enhanced quantum genetic algorithm (EQGA) to find the optimal timetable. This model sets the minimum waiting time cost of passengers and the maximum interests of public transport companies as the goals, restricts the departure interval and two adjacent intervals, and constrains the load factor of passengers. Moreover, based on the analysis of public transport travel characteristics and passenger flow data, this paper evaluates the efficiency of public transport operation, reasonably analyzes the utilization of public transport resources and the travel time of public transport passengers, and thoroughly studies the public transport operation system. This paper selects actual bus line data for empirical analysis, and the empirical results show that the proposed algorithm and model can meet the requirements of many aspects and provides good intelligence, applicability, and optimization.
Haifeng Lin, Chengpei Tang
IEEE Trans. Intell. Transp. Syst.2
2022 Analysis and Optimization of Urban Public Transport Lines Based on Multiobjective Adaptive Particle Swarm Optimization
abstract
Urban public transport is a very complex system, and with the development of urbanization, there are many new urban traffic characteristics. Making bus routes and scheduling strategies more efficient, scientific and accurate has a positive impact on the actual operation of public transport. To solve the urban public transport line design problem, this paper describes the implicit law of the characteristics of public transport travel from a deep perspective and analyzes the forms, influencing factors and existing problems of bus dispatching. By establishing a multiobjective public transport dispatching optimization model, starting from bus companies, passengers and government departments, public transportation operating costs comprehensively consider the interests of various parties and finally realize the optimization objective of minimizing fixed costs, fuel costs, carbon emission costs and time window penalty costs. The objective function is set reasonably, and the generation and optimization method of the initial line set in the public transport line design problem is improved; suitable constraint conditions and evaluation indicators are considered. This paper attempts to control the overall length of the bus line on the premise of fully meeting the travel needs of passengers. By solving the multiobjective problem on the same network and comparing different multiobjective optimization algorithms, the effectiveness of the method is evaluated. Additionally, an improved multiobjective adaptive particle swarm optimization (MOAPSO) is proposed, which has the characteristics of faster convergence, higher efficiency and low computational complexity. The simulation experimental results show that the proposed algorithm in this paper can obtain a better Pareto optimal solution set and can effectively solve the multiobjective model. The departure interval conforms to the passenger flow distribution, which can effectively reduce costs and improve the travel service quality of passengers on a large-scale network.
Haifeng Lin, Chengpei Tang
IEEE Trans. Intell. Transp. Syst.2
2021 Robust Human Activity Recognition System with Wi-Fi Using Handcraft Feature
abstract
WiFi-based Human activity recognition (HAR) system has the drawback of the new domain inadaptability. Numerous studies have proposed to solve this problem, but these methods have the limitations of needing the new domain data or fine-tuning the model. In this paper, we propose HARW, a cross-domain HAR system using Wi-Fi. Specifically, a novel domain-independent feature extraction algorithm is proposed based on the multiple signal classification algorithm, which extracts three physical factors (i.e. time of flight, change rate of path length, and angle of arrival) simultaneously to construct the TCA feature. Then, A two-stage model is proposed to recognize activities based on TCA. The experimental results show that HARW can increase the average accuracy rate by 9 % and the best accuracy can reach 60%, without new domain data and fine-tuning the model, outperforming the method that only uses CSI raw data. In addition. HARW adonts onlv a nair of Wi-Fi devices.
Chengpei Tang, Xie Zhang, Hele Yao
ISCC2
2021 WiFi-Based Multi-task Sensing
Xie Zhang, Chengpei Tang, Yasong An
MobiQuitous2
2018 A Non-Intrusive Multi-Parameter Fault Diagnosis System for Industrial Machineries
abstract
Induction motor, especially driving motor, is the critical component for various modern industrial machineries. Fault diagnosis of induction motor is therefore a necessary and crucial task to ensure the machinery health and prevent vital damages. Conventional diagnosis methods are mainly based on intrusive sensors to measure certain physical parameters. However, intrusive sensors are costly and hard to apply to update traditional machines. In this paper, we propose EMFD, an energy-image based non-intrusive multi-parameter fault diagnosis system. We design an EMFD sensing platform to monitor the electric circuit parameters. Then we build a fault model that describes the relationships between two major kinds of faults and the electric circuit parameters. Based on the model, we propose a novel fault diagnosis algorithm that exploits a sparse auto-encoder based deep neural network. Different from the existing single-parameter methods, EMFD takes advantage of multiple circuit parameters and achieves accurate and robust diagnosis even in dynamic operating environments. We implement and deploy the proposed system in a real-world factory. The evaluation results show that EMFD can achieve the diagnosis accuracy of 96 %.
Shanqing Wang, Chengpei Tang, Chancheng Zhou
ICPADS2
2017 Virtual grid margin optimization and energy balancing scheme for mobile sinks in wireless sensor networks
Chengpei Tang, Nian Yang
Multim. Tools Appl.1
2015 Using Statistical Image Model for JPEG Steganography: Uniform Embedding Revisited
abstract
Uniform embedding was first introduced in 2012 for non-side-informed JPEG steganography, and then extended to the side-informed JPEG steganography in 2014. The idea behind uniform embedding is that, by uniformly spreading the embedding modifications to the quantized discrete cosine transform (DCT) coefficients of all possible magnitudes, the average changes of the first-order and the second-order statistics can be possibly minimized, which leads to less statistical detectability. The purpose of this paper is to refine the uniform embedding by considering the relative changes of statistical model for digital images, aiming to make the embedding modifications to be proportional to the coefficient of variation. Such a new strategy can be regarded as generalized uniform embedding in substantial sense. Compared with the original uniform embedding distortion (UED), the proposed method uses all the DCT coefficients (including the DC, zero, and non-zero AC coefficients) as the cover elements. We call the corresponding distortion function uniform embedding revisited distortion (UERD), which incorporates the complexities of both the DCT block and the DCT mode of each DCT coefficient (i.e., selection channel), and can be directly derived from the DCT domain. The effectiveness of the proposed scheme is verified with the evidence obtained from the exhaustive experiments using a popular steganalyzer with rich models on the BOSSbase database. The proposed UERD gains a significant performance improvement in terms of secure embedding capacity when compared with the original UED, and rivals the current state-of-the-art with much reduced computational complexity.
Linjie Guo, Jiangqun Ni, Wenkang Su 0001, Chengpei Tang, Yun Q. Shi 0001
IEEE Trans. Inf. Forensics Secur.4