Xudong Zhang 0001

dblp:58/6205-1 · also Xu-Dong Zhang 0001 · DBLP profile ↗
← Back
56ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0002-6465-7437ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 10Computer networks · 9 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient Reinforcement Learning for Zero-Shot Coordination in Evolving Games
abstract
Zero-shot coordination(ZSC), a key challenge in multi-agent game theory, has become a hot topic in reinforcement learning (RL) research recently, especially in complex evolving games. It focuses on the generalization ability of agents, requiring them to coordinate well with collaborators from a diverse, potentially evolving, pool of partners that are not seen before without any fine-tuning. Population-based training, which approximates such an evolving partner pool, has been proven to provide good zero-shot coordination performance; nevertheless, existing methods are limited by computational resources, mainly focusing on optimizing diversity in small populations while neglecting the potential performance gains from scaling population size. To address this issue, this paper proposes the Scalable Population Training (ScaPT), an efficient RL training framework comprising two key components: a meta-agent that efficiently realizes a population by selectively sharing parameters across agents, and a mutual information regularizer that guarantees population diversity. To empirically validate the effectiveness of ScaPT, this paper evaluates it along with representational frameworks in Hanabi cooperative game and confirms its superiority.
Bingyu Hui, Lebin Yu, Quanming Yao, Yunpeng Qu, Xudong Zhang 0001, Jian Wang 0030
AAAI5
2026 Enhancing Open-Set RFF Recognition With cGAN: Generating Multiple Unknown Classes
abstract
Open set recognition (OSR) in radio frequency fingerprint (RFF) is critical for securing Internet of Things (IoT) systems, where previously unseen or malicious devices may attempt unauthorized access. A widely used approach treats all unknown devices as a single additional class and assumes that they will produce low confidence scores during inference. However, due to the inherently subtle and highly similar RFF features across devices, this assumption often fails, leading to high false acceptance rates. To address this challenge, we propose a novel framework, called Multiple Unknown Classes Generation (MUCG), which replaces the single-class modeling of unknowns with a more expressive structure that simulates multiple distinct unknown classes. MUCG employs a conditional generative adversarial network (cGAN) guided by ideal signal priors to produce diverse and realistic unknown samples. Furthermore, we introduce a soft label perturbation (SLP) strategy that blends label semantics using Feature-wise Linear Modulation (FiLM), encouraging the generator to embed richer feature variations. Experiments on three public IoT datasets demonstrate that MUCG consistently outperforms state-of-the-art (SOTA) methods in OSR tasks, achieving superior accuracy and robustness under varying signal conditions.
Haohao Sun, Cong Zou, Qiexiang Wang, Jian Wang 0030, Xudong Zhang 0001
IEEE Internet Things J.5
2025 Mixed Policy-Space Response Oracles
abstract
Finding the Nash equilibrium in large-scale zero-sum games has long been a challenging problem due to the vast and unknown policy space and the utility matrix. The Policy-Space Response Oracles (PSRO) framework, as a combination of conventional game analysis with deep reinforcement learning, iteratively constructs a restricted game by including a pure policy that is the best response to the equilibrium of the restricted game in the previous iteration. However, adding a pure policy at a time is inefficient, as there may be multiple policies that dominate the equilibrium of the current restricted game. In this regard, we propose Mixed Policy-Space Response Oracles (M-PSRO), which add a mixed policy rather than a pure policy to the current policy set. To obtain more effective mixed policy candidates, we adopt a parallelized training framework to promote policy diversity. Theoretical analysis shows that M-PSRO can converge to an approximate Nash equilibrium. We conduct extensive experiments across a wide range of complex games, and results show that the M-PSRO algorithm achieves state-of-the-art performance across all the games and can approach the Nash equilibrium more efficiently. Numerous ablation studies show that the improvement benefits from the usage of mixed policies in M-PSRO, which offer an even more flexible optimization space.
Feihong Yang, Jian Wang 0030, Chao Wang 0083, Xudong Zhang 0001
IJCNN5
2025 Model-Based RF Fingerprint Extraction Approach for Robust IoT Device Identification
abstract
Radio frequency fingerprint identification (RFFI) leverages signal distortions caused by hardware impairments to identify transmitters, thereby enhancing IoT security. However, current radio frequency fingerprints (RFFs) modelings typically focus on partial hardware impairments, risking incomplete RFF extraction and limited RFF understanding. This study aims to refine the modeling of RFFs and guide the development of robust and accurate RFFI approaches based on this model. Specifically, we propose a comprehensive time-domain signal distortion model based on hardware impairments in wireless transmission circuit components, revealing that RFFs can be categorized into two types: 1) fine-grained RFF and 2) coarse-grained RFF. The former encompass localized distortions, such as mismatches, intersymbol interference, and nonlinear distortions; the latter relate to global features, including frequency spurs, phase noise, and crystal oscillator frequency offset. Subsequently, we analyze the impact of interference on the RFF model and propose necessary methods to mitigate the interference. Combining the comprehensive analysis of the RFF model and interference, we summarize three primary characteristics of RFFs: 1) multiscale; 2) fixedness; and 3) ubiquity. These characteristics indicate that convolutional neural networks (CNNs) from the visual domain cannot be directly transferred or simply adapted in terms of input data shape for application in RFFI. Therefore, we propose an enhanced CNN architecture with grouped convolutions and channel fusion modules for effective RFF extraction. To demonstrate the generalizability of our approach, we conduct extensive experiments using three public IoT signal datasets. Experimental results demonstrate that our method exhibits excellent identification performance and robustness against interference across various environments.
Qiexiang Wang, Yazhou Sun, Zhongfang Wang, Longhui Wang, Jian Wang 0030, Xudong Zhang 0001
IEEE Internet Things J.6
2024 Robust Communicative Multi-Agent Reinforcement Learning with Active Defense
abstract
Communication in multi-agent reinforcement learning (MARL) has been proven to effectively promote cooperation among agents recently. Since communication in real-world scenarios is vulnerable to noises and adversarial attacks, it is crucial to develop robust communicative MARL technique. However, existing research in this domain has predominantly focused on passive defense strategies, where agents receive all messages equally, making it hard to balance performance and robustness. We propose an active defense strategy, where agents automatically reduce the impact of potentially harmful messages on the final decision. There are two challenges to implement this strategy, that are defining unreliable messages and adjusting the unreliable messages' impact on the final decision properly. To address them, we design an Active Defense Multi-Agent Communication framework (ADMAC), which estimates the reliability of received messages and adjusts their impact on the final decision accordingly with the help of a decomposable decision structure. The superiority of ADMAC over existing methods is validated by experiments in three communication-critical tasks under four types of attacks.
Lebin Yu, Yunbo Qiu, Quanming Yao, Yuan Shen 0001, Xudong Zhang 0001, Jian Wang 0030
AAAI5
2024 Sensing-aided CSI Feedback with Deep Learning for Massive MIMO Systems
abstract
For frequency division duplexing massive multiple-input multiple-output systems, downlink channel state information (CSI) is required to be compressed and fed back to the base station (BS) to support beamforming. Recently, deep learning (DL) has demonstrated overwhelming performance in CSI feed-back, wherein multimodal information is explored to further improve the performance. With the emergence of integrated sensing and communications, radar-equipped BSs exhibit the capability to sense the wireless environment and assist communication design. In this paper, we propose a sensing-aided DL-based CSI feedback method, in which the angle information of scatters in the communication channel is sensed by the BS and utilized to reduce feedback overhead. A novel two-stage feedback scheme with lightweight network structures is carefully designed to improve feedback performance. Experiments demonstrate that compared to previous methods without utilization of sensing information, our sensing-aided methods significantly enhance performance in low-bit scenarios with reduced computational complexity. The open-source codes are available at https://github.com/zhang-xd18/safb.
Xudong Zhang 0001, Zhilin Lu 0002, Jintao Wang 0001
ICC1
2024 A Time-Varying and Time-Invariant RF Fingerprint Extraction Approach for IoT Device Identification
abstract
Radio Frequency Fingerprinting (RFF)-based identification methods have the potential to enhance the security of the Internet of Things (IoT). However, conventional fingerprinting techniques based on standard sample rates face limitations related to noise and device scale. The utilization of high sample rate receivers offers a promising solution to mitigate these constraints. Nonetheless, the challenge lies in extracting RFFs from the collected ultra-long signals. Image-based methods, which accumulate signals in the time domain, reduce the difficulty of extracting RFFs from long signals but overlook the fine-grained RFFs in the time domain. To solve this problem, this paper proposes RFF modeling for long signals, emphasizing the importance of obtaining short-term time-varying RFFs and time-invariant RFFs. Combining an analysis of the inductive biases of convolutional neural networks, we propose a backbone network named GResNet, which is capable to effectively extract these two types of RFFs. An information fusion module is added to improve identification performance. Extensive experiments are conducted with 100 LoRa devices, demonstrating that our method outperforms existing RFFI techniques based on standard sample rate or high sample rate signals. Furthermore, our approach maintains robust performance within a wide range of SNRs.
Qiexiang Wang, Yazhou Sun, Longhui Wang, Jian Wang 0030, Xudong Zhang 0001
ICC5
2024 Open-Set RF Fingerprint Identification with Synthetic Feature Constraint
abstract
The rapid expansion of the Internet of Things (IoT) has heightened the necessity for device identity authentication to ensure security. Radio frequency fingerprint identification (RFFI) has emerged as a promising solution for this purpose, which leverage unique signal distortions caused by hardware impairments to authenticate device identities. However, most RFFI methods operate under a closed-set assumption and usually mistakenly identify unknown devices from the open set as known devices. In this paper, we propose a Synthetic Feature Constrain for open-set Recognition (SFCR) method to maintain classification performance on known devices and identify unknowns. Specifically, we modify the nonlinear characteristics of known devices based on the power amplifier nonlinearity model of radio frequency fingerprints (RFF) to synthesize signals for unknown devices. Furthermore, we propose a synthetic feature constraint to calibrate the position of synthetic devices in the feature space, such that they lie between the feature centers of the collective synthetic and originating known devices. As synthetic devices represent only a subset of the real unknown devices, we also introduce a calibration method for the prediction results. Experiments on a publicly available LoRa device dataset have validated the effectiveness of our approach. The code is released on github.com/wzyxwqx/SFCR.
Qiexiang Wang, Haohao Sun, Yazhou Sun, Zhongfang Wang, Jian Wang 0030, Xudong Zhang 0001
VTC Fall6
2024 Effective Multi-Agent Communication Under Limited Bandwidth
abstract
With the fast development of multi-agent reinforcement learning, communication among agents has become a new research hotspot for its significant role in promoting the cooperation of automated devices. However, in real-world scenarios, agents such as unmanned vehicles and robots are likely to suffer from communication resource constraints, making designing efficient communication protocols essential. In this paper, we propose to quantize messages and reduce discrete entropy to achieve effective multi-agent communication under bandwidth limits. Achieving this goal requires solving two challenges: The first one is that the gradients of discrete entropy remain zero except for several discontinuous points wherein the gradients are undefined, making it hard to reduce discrete entropy via gradient-based training. To overcome it, we design Surrogate Entropy Minimization (SEM) scheme and confirm its effectiveness theoretically. The second challenge is maximizing cooperation performance under a given bandwidth limit. We model it as a constrained optimization problem and design Soft Barrier Method (SBM). Our proposed scheme is evaluated alongside four other methods in six environment settings and five different bandwidth limits, and demonstrates outstanding performance. Specifically, it manages to reduce bandwidth consumption by up to 90% with little or no loss of cooperation performance.
Lebin Yu, Qiexiang Wang, Yunbo Qiu, Jian Wang 0030, Xudong Zhang 0001, Zhu Han 0001
IEEE Trans. Mob. Comput.5
2023 Promoting Cooperation in Multi-Agent Reinforcement Learning via Mutual Help
abstract
Multi-agent reinforcement learning (MARL) has achieved great progress in cooperative tasks in recent years. However, in the local reward scheme, where only local rewards for each agent are given without global rewards shared by all the agents, traditional MARL algorithms lack sufficient consideration of agents’ mutual influence. In cooperative tasks, agents’ mutual influence is especially important since agents are supposed to coordinate to achieve better performance. In this paper, we propose a novel algorithm Mutual-Help-based MARL (MH-MARL) to instruct agents to help each other in order to promote cooperation. MH-MARL utilizes an expected action module to generate expected other agents’ actions for each particular agent. Then, the expected actions are delivered to other agents for selective imitation during training. Experimental results show that MH-MARL improves the performance of MARL both in success rate and cumulative reward.
Yunbo Qiu, Lebin Yu, Jian Wang 0030, Xudong Zhang 0001
ICASSP5
2023 Low Entropy Communication in Multi-Agent Reinforcement Learning
abstract
Communication in multi-agent reinforcement learning has been drawing attention recently for its significant role in cooperation. However, multi-agent systems may suffer from limitations on communication resources and thus need efficient communication techniques in real-world scenarios. According to the Shannon-Hartley theorem, messages to be transmitted reliably in worse channels require lower entropy. Therefore, we aim to reduce message entropy in multi-agent communication. A fundamental challenge is that the gradients of entropy are either 0 or ∞, disabling gradient-based methods. To handle it, we propose a pseudo gradient descent scheme, which reduces entropy by adjusting the distributions of messages wisely. We conduct experiments on two base communication frameworks with six environment settings and find that our scheme can reduce message entropy by up to 90% with nearly no loss of cooperation performance.
Lebin Yu, Yunbo Qiu, Qiexiang Wang, Xudong Zhang 0001, Jian Wang 0030
ICC4
2023 Improving Sample Efficiency of Multiagent Reinforcement Learning With Nonexpert Policy for Flocking Control
abstract
Control algorithms of a multiagent system (MAS) have been applied to many Internet of Things devices, such as unmanned aerial vehicles and autonomous underwater vehicles. Flocking control is a crucial problem in MAS to enhance the safety and cooperativity of agents, which requires the agents to maintain the flock when navigating to a target position and avoiding collisions. In comparison with the traditional algorithms, methods based on multiagent reinforcement learning (MARL) can solve the problem of flocking control more flexibly and adapt to more complex environments. However, the MARL-based methods demand a huge number of interactions between agents and the environment, resulting in the problem of sample inefficiency. In this article, we propose nonexpert policy-aided MARL (NPA-MARL) to improve sample efficiency, which utilizes a fundamental MARL algorithm and a prior policy whose performance can be nonexpert. Before online MARL training, NPA-MARL generates demonstrations by the nonexpert policy to pretrain agents, while preventing overfitting demonstrations. During online training, NPA-MARL instructs agents to imitate the nonexpert policy if the nonexpert policy is better in agents’ recognition. We leverage NPA-MARL to solve the problem of flocking control. Experimental results show that NPA-MARL improves sample efficiency and policy performance in flocking control. Besides, NPA-MARL has the scalability of more agents and the flexibility of choice of the nonexpert policy and a fundamental MARL algorithm.
Yunbo Qiu, Lebin Yu, Jian Wang 0030, Yu Wang 0002, Xudong Zhang 0001
IEEE Internet Things J.6
2023 Dual-Timescale Resource Allocation for Collaborative Service Caching and Computation Offloading in IoT Systems
abstract
Edge computing has been envisioned as a key enabler to provide computation-intensive and delay-sensitive services in the future Internet of Things systems. By offloading the computational tasks to the edge server, both the service latency and energy consumption can be reduced. Since devices may request various types of computing services, caching appropriate services in the edge server to immediately provide computing resources can improve the quality of service. Nevertheless, it brings new challenges to jointly optimize the resource allocation, where the timeliness of caching and offloading operations are different. In this article, we first formulate the collaborative service caching and computation offloading as a dual-timescale resource allocation problem to minimize the costs of latency and energy consumption. Under this framework, a novel scheme based on hierarchical deep reinforcement learning is proposed to output collaborative caching and computing actions. Specifically, the proposed approach contains the service caching policy and the device computing policy with hierarchical action–value functions, which allows a flexible configuration of caching timescales. The simulation results demonstrate that the proposed policy outperforms the existing schemes on convergence performance and various parameters.
Yuan Shen 0001, Yu Wang 0002, Xudong Zhang 0001, Jian Wang 0030
IEEE Trans. Ind. Informatics4
2023 Hierarchical and Stable Multiagent Reinforcement Learning for Cooperative Navigation Control
abstract
We solve an important and challenging cooperative navigation control problem, Multiagent Navigation to Unassigned Multiple targets (MNUM) in unknown environments with minimal time and without collision. Conventional methods are based on multiagent path planning that requires building an environment map and expensive real-time path planning computations. In this article, we formulate MNUM as a stochastic game and devise a novel multiagent deep reinforcement learning (MADRL) algorithm to learn an end-to-end solution, which directly maps raw sensor data to control signals. Once learned, the policy can be deployed onto each agent, and thereby, the expensive online planning computations can be offloaded. However, to solve MNUM, traditional MADRL suffers from large policy solution space and nonstationary environment when agents make decisions independently and concurrently. Accordingly, we propose a hierarchical and stable MADRL algorithm. The hierarchical learning part introduces a two-layer policy model to reduce the solution space and uses an interlaced learning paradigm to learn two coupled policies. In the stable learning part, we propose to learn an extended action-value function that implicitly incorporates estimations of other agents' actions, based on which the environment's nonstationarity caused by other agents' changing policies can be alleviated. Extensive experiments demonstrate that our method can converge in a fast way and generate more efficient cooperative navigation policies than comparable methods.
Shuangqing Wei, Xudong Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 UAV-Enabled Covert Federated Learning
abstract
Integrating unmanned aerial vehicles (UAVs) with federated learning (FL) has been seen as a promising paradigm for dealing with the massive amounts of data generated by intelligent devices. Nevertheless, although FL has natural advantages in data security protection, eavesdroppers can also deduce the raw data according to the shared parameters. Existing works mainly focused on encrypting the content of uploaded parameters, but we believe that it can improve security further by hiding the presence of parameter updating. Therefore, in this paper, we conceive a UAV-enabled covert federated learning architecture, where the UAV is not only responsible for orchestrating the operation of FL but also for emitting artificial noise (AN) to interfere with the eavesdropping of unintended users. To strike a balance between the security level and the training cost (including time overhead and energy consumption), we propose a distributed proximal policy optimization-based strategy for the sake of jointly optimizing the trajectory and AN transmitting power of the UAV, the CPU frequency, the transmitting power and the bandwidth allocation of the participated devices, as well as the needed accuracy of the local model. Furthermore, a series of experiments have been conducted to validate the effectiveness of our proposed scheme.
Xiangwang Hou, Jingjing Wang 0001, Chunxiao Jiang, Xudong Zhang 0001, Yong Ren 0001, Mérouane Debbah
IEEE Trans. Wirel. Commun.4
2022 Supervised Off-Policy Ranking
abstract
Off-policy evaluation (OPE) is to evaluate a target policy with data generated by other policies. Most previous OPE methods focus on precisely estimating the true performance of a policy. We observe that in many applications, (1) the end goal of OPE is to compare two or multiple candidate policies and choose a good one, which is a much simpler task than precisely evaluating their true performance; and (2) there are usually multiple policies that have been deployed to serve users in real-world systems and thus the true performance of these policies can be known. Inspired by the two observations, in this work, we study a new problem, supervised off-policy ranking (SOPR), which aims to rank a set of target policies based on supervised learning by leveraging off-policy data and policies with known performance. We propose a method to solve SOPR, which learns a policy scoring model by minimizing a ranking loss of the training policies rather than estimating the precise policy performance. The scoring model in our method, a hierarchical Transformer based model, maps a set of state-action pairs to a score, where the state of each pair comes from the off-policy data and the action is taken by a target policy on the state in an offline manner. Extensive experiments on public datasets show that our method outperforms baseline methods in terms of rank correlation, regret value, and stability. Our code is publicly available at GitHub.
Tao Qin 0001, Xudong Zhang 0001, Houqiang Li, Tie-Yan Liu
ICML4
2022 Sub-optimal Policy Aided Multi-Agent Reinforcement Learning for Flocking Control
abstract
Flocking control is a challenging problem, where multiple agents, such as drones or vehicles, need to reach a target position while maintaining the flock and avoiding collisions with obstacles and collisions among agents in the environment. Multi-agent reinforcement learning has achieved promising performance in flocking control. However, methods based on traditional reinforcement learning require a considerable number of interactions between agents and the environment. This paper proposes a sub-optimal policy aided multiagent reinforcement learning algorithm (SPA-MARL) to boost sample efficiency. SPA-MARL directly leverages a prior policy that can be manually designed or solved with a non-learning method to aid agents in learning, where the performance of the policy can be sub-optimal. SPA-MARL recognizes the difference in performance between the sub-optimal policy and itself, and then imitates the sub-optimal policy if the suboptimal policy is better. We leverage SPA-MARL to solve the flocking control problem. A traditional control method based on artificial potential fields is used to generate a sub-optimal policy. Experiments demonstrate that SPA-MARL can speed up the training process and outperform both the MARL baseline and the used sub-optimal policy.
Yunbo Qiu, Jian Wang 0030, Xudong Zhang 0001
SMC4
2022 Sample-Efficient Multi-Agent Reinforcement Learning with Demonstrations for Flocking Control
abstract
Flocking control is a significant problem in multi-agent systems such as multi-agent unmanned aerial vehicles and multi-agent autonomous underwater vehicles, which enhances the cooperativity and safety of agents. In contrast to traditional methods, multi-agent reinforcement learning (MARL) solves the problem of flocking control more flexibly. However, methods based on MARL suffer from sample inefficiency, since they require a huge number of experiences to be collected from interactions between agents and the environment. We propose a novel method Pretraining with Demonstrations for MARL (PwD-MARL), which can utilize non-expert demonstrations collected in advance with traditional methods to pretrain agents. During the process of pretraining, agents learn policies from demonstrations by MARL and behavior cloning simultaneously, and are prevented from overfitting demonstrations. By pretraining with non-expert demonstrations, PwD-MARL improves sample efficiency in the process of online MARL with a warm start. Experiments show that PwD-MARL improves sample efficiency and policy performance in the problem of flocking control, even with bad or few demonstrations.
Yunbo Qiu, Yuzhu Zhan, Jian Wang 0030, Xudong Zhang 0001
VTC Fall5
2022 The Optimized Sparse Fourier Transform for Band-Limited Signal
abstract
Sparse fast Fourier transform (SFFT) achieves spectrum sensing with sublinear computational and sample complexity, which has raised widely attention in the signal processing community recently. However, SFFT ignores the structure characteristics of spectrum. In this paper, we optimize the SFFT algorithm for band-limited spectrum sensing. The optimized permutation theory is given and proven first. Then, based on the optimized permutation theory, the optimized sparse Fourier transform for band-limited signal (OB-SFT) is designed. OB-SFT is a universal algorithm with deterministic parameters, and the computational and sample complexity are less than SFFT. Finally, numerical simulations verify the effectiveness and advantages of OB-SFT.
Longhui Wang, Qiexiang Wang, Jian Wang 0030, Xudong Zhang 0001
VTC Fall4
2020 Stabilizing Multi-Agent Deep Reinforcement Learning by Implicitly Estimating Other Agents' Behaviors
abstract
Deep reinforcement learning (DRL) is able to learn control policies for many complicated tasks, but it's power has not been unleashed to handle multi-agent circumstances. Independent learning, where each agent treats others as part of the environment and learns its own policy without considering others' policies is a simple way to apply DRL to multi-agent tasks. However, since agents' policies change as learning proceeds, from the perspective of each agent, the environment is non-stationary, which makes conventional DRL methods inefficient. To cope with this challenge, we propose a novel approach where each agent uses an implicit estimate of others' actions to guide its own policy learning. We demonstrate that given the implicit estimate of others' actions, each agent can learn its policy in a relatively stationary environment. Extensive experiments show that our method significantly alleviates the non-stationarity and outperforms the state-of-the-art in terms of both convergence speed and policy performance.
Shuangqing Wei, Xudong Zhang 0001, Chao Wang 0083
ICASSP4
2020 Clipping Noise Estimation Based on Deep Complex Neural Network with Sparsity Constraint
abstract
Clipping noise estimation and cancellation are essential in orthogonal frequency division multiplexing (OFDM) systems when clipping is performed to reduce the peak-to-average power ratio (PAPR). Motivated by the richer representational capacity of complex numbers and the fact that communication is a complex-valued problem, a novel clipping noise estimation scheme based on deep complex neural network is proposed in this paper. Specifically, the clipping noise is determined by a deep complex network, namely clipping noise estimation network (CNE-Net), such that the mean square error (MSE) and the sparsity of the estimated clipping noise are jointly optimized. Besides, an ordering based zero-forcing scheme is utilized to further ensure the sparsity of the estimated clipping noise. Simulation results show that the proposed CNE-Net shows comparable performance with the conventional decision-aided reconstruction (DAR) scheme and can achieve better performance than the one-iteration DAR scheme when the clipping noise is not sparse enough. In summary, the CNE-Net has a good capability to estimate the clipping noise from noise-affected features.
Xudong Zhang 0001, Yu Zhang 0050, Xiaohua Chang, Changyong Pan
VTC Spring1
2020 Robust Clipping Noise Cancellation Based on Location-Aware Compressed Sensing
abstract
For OFDM systems, to cope with the remain problems of high complexity and low performance in conventional clipping noise cancellation methods, a robust scheme based on location-aware compressed sensing (CS) and phase correction is proposed in this paper. Based on CS theory, a simple and configurable selection criterion is utilized to choose reliable observations for the clipping noise reconstruction. The transceiver is redesigned to transmit both the data and the clipping location. With the aid of clipping location information and phase information from the receiver, the proposed scheme improves both the accuracy and computational complexity. Simulation results show that the proposed scheme achieves excellent performance even in low signal-to-noise ratio (SNR) environments. Besides, due to the low computational complexity and excellent adaptivity, the proposed scheme is more feasible in practical engineering applications than other CS-based methods.
Xudong Zhang 0001, Yu Zhang 0050, Xiaohua Chang, Changyong Pan
VTC Spring1
2020 Deep-Reinforcement-Learning-Based Autonomous UAV Navigation With Sparse Rewards
abstract
Unmanned aerial vehicles (UAVs) have the potential in delivering Internet-of-Things (IoT) services from a great height, creating an airborne domain of the IoT. In this article, we address the problem of autonomous UAV navigation in large-scale complex environments by formulating it as a Markov decision process with sparse rewards and propose an algorithm named deep reinforcement learning (RL) with nonexpert helpers (LwH). In contrast to prior RL-based methods that put huge efforts into reward shaping, we adopt the sparse reward scheme, i.e., a UAV will be rewarded if and only if it completes navigation tasks. Using the sparse reward scheme ensures that the solution is not biased toward potentially suboptimal directions. However, having no intermediate rewards hinders the agent from efficient learning since informative states are rarely encountered. To handle the challenge, we assume that a prior policy (nonexpert helper) that might be of poor performance is available to the learning agent. The prior policy plays the role of guiding the agent in exploring the state space by reshaping the behavior policy used for environmental interaction. It also assists the agent in achieving goals by setting dynamic learning objectives with increasing difficulty. To evaluate our proposed method, we construct a simulator for UAV navigation in large-scale complex environments and compare our algorithm with several baselines. Experimental results demonstrate that LwH significantly outperforms the state-of-the-art algorithms handling sparse rewards and yields impressive navigation policies comparable to those learned in the environment with dense rewards.
Chao Wang 0083, Jian Wang 0030, Jingjing Wang 0001, Xudong Zhang 0001
IEEE Internet Things J.4
2019 Efficient Multi-agent Cooperative Navigation in Unknown Environments with Interlaced Deep Reinforcement Learning
abstract
This work addresses a multi-agent cooperative navigation problem that multiple agents work together in an unknown environment in order to reach different targets without collision and minimize the maximum navigation time they spend. Typical reinforcement learning-based solutions directly model the cooperative navigation policy as a steering policy. However, when each agent does not know which target to head for, this method could prolong convergence time and reduce overall performance. To this end, we model the navigation policy as a combination of a dynamic target selection policy and a collision avoidance policy. Since these two policies are coupled, an interlaced deep reinforcement learning method is proposed to simultaneously learn them. Additionally, a reward function is directly derived from the optimization objective function instead of using a heuristic design method. Extensive experiments demonstrate that the proposed method can converge in a fast way and generate a more efficient navigation policy compared with the state-of-the-art.
Xudong Zhang 0001
ICASSP4
2019 A Modified Inception-ResNet Network with Discriminant Weighting Loss for Handwritten Chinese Character Recognition
abstract
Handwritten Chinese character recognition (HCCR) is a representative large character set pattern classification task. Recently, convolutional neural networks have provided promising solutions for this challenging task. This paper adopts the modified Inception-ResNet network for handwritten Chinese character recognition, and proposes a discriminant weighting method for cross-entropy loss calculation which focuses on recognition errors in the training stage. Sparse training technique is also incorporated. Under the specific condition of utilizing the testing mini-batch mean and variance for batch normalization, the proposed method achieves improved performance on the ICDAR-2013 offline handwritten Chinese character competition dataset.
Linhui Chen, Liangrui Peng, Changsong Liu, Xudong Zhang 0001
ICDAR5
2017 Automatic radar waveform recognition based on time-frequency analysis and convolutional neural network
abstract
In this paper, we apply the idea of deep learning to radar waveform recognition. Since the frequency variation with time is the most essential distinction among radar signals with different modulation types, we transform one-dimensional radar signals into time-frequency images (TFIs) using time-frequency analysis and design a convolutional neural network to recognize the frequency variation patterns exhibited in TFIs. Furthermore, we analyze the statistical characteristics of the noise in TFIs and introduce a naive approach to reduce its influence on the frequency variation patterns. Simulation results demonstrate the impressive recognition rate under very low SNR conditions and the strong generalization ability of our proposed recognition method.
Chao Wang 0083, Jian Wang 0030, Xudong Zhang 0001
ICASSP3
2017 Hyper-spectral image reconstruction based on SL0-SL0 minimization
abstract
This paper proposes a new prior image constrained compressive sampling (PICCS) method to reconstruct hyper-spectral images, namely SL0-SL0minimization-based hyper-spectral imaging (HSI). This is a band-by-band reconstruction method, which reconstructs each hyper-spectral band based on the previous one. This method utilizes not only the sparsity of each hyper-spectral band in certain bases but also the similarity between two consecutive bands. In addition, compared with the popular approaches which reconstruct all the hyper-spectral bands simultaneously, SL0-SL0minimization-based HSI reduce the requirements to computational ability and memory of receivers for that only one hyper-spectral band is reconstructed at each time. Compared with the exiting PICCS methods, which lose efficiency to reconstruct signals with large size, the SL0-SL0minimization method significantly speeds up the reconstruction procedure. Some simulations are provided to illustrate the effectiveness of the proposed method.
Xinyue Zhang 0011, Xudong Zhang 0001
ICME2
2017 Finite budget analysis of multi-armed bandit problems
Yingce Xia, Tao Qin 0001, Wenkui Ding, Haifang Li 0002, Xudong Zhang 0001, Nenghai Yu, Tie-Yan Liu
Neurocomputing5
2016 Underwater sonar target imaging via compressed sensing with M sequences
Huichen Yan, Jia Xu 0001, Xiang-Gen Xia 0001, Xudong Zhang 0001, Teng Long 0001
Sci. China Inf. Sci.4
2016 Road-Aided Doppler Ambiguity Resolver for SAR Ground Moving Target in the Image Domain
abstract
A new Doppler ambiguity resolver (DAR) is proposed for ground moving targets of synthetic aperture radar (SAR) in the image domain. Based on the range-Doppler imaging of a static scene, the moving target's response is analyzed in the image domain, and the target's response slope is jointly determined by three motion parameters, i.e., azimuth velocity, ambiguous range velocity, and Doppler ambiguity number. A new DAR utilizing these three parameters is then proposed via the following steps. First, the moving target is detected after ground clutter cancelation between dual-channel images. Second, the ambiguous range velocity is estimated, and the road and its slope are extracted from the SAR image to establish the relationship between the 2-D velocities. Third, the Doppler ambiguity number is determined based on the response slope, the road slope, and the estimated ambiguous range velocity. Compared with the existing DARs in the 2-D time domain, the proposed algorithm works better in the most common signal-to-clutter-noise ratio scenarios. Finally, the results of the numerical experiments are provided to demonstrate the effectiveness of the proposed method.
Zhirui Wang 0003, Jia Xu 0001, Zu-Zhen Huang, Xudong Zhang 0001, Xiang-Gen Xia 0001, Teng Long 0001
IEEE Geosci. Remote. Sens. Lett.4
2015 Budgeted Bandit Problems with Continuous Random Costs
Yingce Xia, Wenkui Ding, Xudong Zhang 0001, Nenghai Yu, Tao Qin 0001
ACML3
2015 Wideband underwater sonar imaging via compressed sensing with scaling effect compensation
Huichen Yan, Jia Xu 0001, Xiang-Gen Xia 0001, Feng Liu 0010, Shibao Peng, Xudong Zhang 0001, Teng Long 0001
Sci. China Inf. Sci.6
2015 Learning to Rank from Noisy Data
abstract
Learning to rank, which learns the ranking function from training data, has become an emerging research area in information retrieval and machine learning. Most existing work on learning to rank assumes that the training data is clean, which is not always true, however. The ambiguity of query intent, the lack of domain knowledge, and the vague definition of relevance levels all make it difficult for common annotators to give reliable relevance labels to some documents. As a result, the relevance labels in the training data of learning to rank usually contain noise. If we ignore this fact, the performance of learning-to-rank algorithms will be damaged. In this article, we propose considering the labeling noise in the process of learning to rank and using a two-step approach to extend existing algorithms to handle noisy training data. In the first step, we estimate the degree of labeling noise for a training document. To this end, we assume that the majority of the relevance labels in the training data are reliable and we use a graphical model to describe the generative process of a training query, the feature vectors of its associated documents, and the relevance labels of these documents. The parameters in the graphical model are learned by means of maximum likelihood estimation. Then the conditional probability of the relevance label given the feature vector of a document is computed. If the probability is large, we regard the degree of labeling noise for this document as small; otherwise, we regard the degree as large. In the second step, we extend existing learning-to-rank algorithms by incorporating the estimated degree of labeling noise into their loss functions. Specifically, we give larger weights to those training documents with smaller degrees of labeling noise and smaller weights to those with larger degrees of labeling noise. As examples, we demonstrate the extensions for McRank, RankSVM, RankBoost, and RankNet. Empirical results on benchmark datasets show that the proposed approach can effectively distinguish noisy documents from clean ones, and the extended learning-to-rank algorithms can achieve better performances than baselines.
Wenkui Ding, Xiubo Geng, Xudong Zhang 0001
ACM Trans. Intell. Syst. Technol.3
2014 MIMO Radar Algorithm Parallel Implementation Based on TMS320C6678
abstract
To improve the real-time processing performance of MIMO radar signal processing, based on the analysis of working principles of MIMO radar and task level parallelism in multi-core DSP, this paper proposes an approach of MIMO radar algorithm parallel implementation on TMS320C6678. We focus on this high performance eight-core digital signal processor (DSP) and describe how to implement the MIMO radar algorithm on it efficiently. Our experimental results show that the implementation achieves desired parallel results and the real-time processing capability of the algorithm is improved much.
Yantao Gao, Xudong Zhang 0001
DASC3
2013 Multi-Armed Bandit with Budget Constraint and Variable Costs
abstract
We study the multi-armed bandit problems with budget constraint and variable costs (MAB-BV). In this setting, pulling an arm will receive a random reward together with a random cost, and the objective of an algorithm is to pull a sequence of arms in order to maximize the expected total reward with the costs of pulling those arms complying with a budget constraint. This new setting models many Internet applications (e.g., ad exchange, sponsored search, and cloud computing) in a more accurate manner than previous settings where the pulling of arms is either costless or with a fixed cost. We propose two UCB based algorithms for the new setting. The first algorithm needs prior knowledge about the lower bound of the expected costs when computing the exploration term. The second algorithm eliminates this need by estimating the minimal expected costs from empirical observations, and therefore can be applied to more real-world applications where prior knowledge is not available. We prove that both algorithms have nice learning abilities, with regret bounds of O(ln B). Furthermore, we show that when applying our proposed algorithms to a previous setting with fixed costs (which can be regarded as our special case), one can improve the previously obtained regret bound. Our simulation results on real-time bidding in ad exchange verify the effectiveness of the algorithms and are consistent with our theoretical analysis.
Wenkui Ding, Tao Qin 0001, Xudong Zhang 0001, Tie-Yan Liu
AAAI3
2008 Global Ranking Using Continuous Conditional Random Fields
abstract
This paper studies global ranking problem by learning to rank methods. Conventional learning to rank methods are usually designed for `local ranking', in the sense that the ranking model is defined on a single object, for example, a document in information retrieval. For many applications, this is a very loose approximation. Relations always exist between objects and it is better to define the ranking model as a function on all the objects to be ranked (i.e., the relations are also included). This paper refers to the problem as global ranking and proposes employing a Continuous Conditional Random Fields (CRF) for conducting the learning task. The Continuous CRF model is defined as a conditional probability distribution over ranking scores of objects conditioned on the objects. It can naturally represent the content information of objects as well as the relation information between objects, necessary for global ranking. Taking two specific information retrieval tasks as examples, the paper shows how the Continuous CRF method can perform global ranking better than baselines.
Tao Qin 0001, Tie-Yan Liu, Xudong Zhang 0001, De-Sheng Wang, Hang Li 0001
NIPS3
2008 Learning to rank relational objects and its application to web search
abstract
Learning to rank is a new statistical learning technology on creating a ranking model for sorting objects. The technology has been successfully applied to web search, and is becoming one of the key machineries for building search engines. Existing approaches to learning to rank, however, did not consider the cases in which there exists relationship between the objects to be ranked, despite of the fact that such situations are very common in practice. For example, in web search, given a query certain relationships usually exist among the the retrieved documents, e.g., URL hierarchy, similarity, etc., and sometimes it is necessary to utilize the information in ranking of the documents. This paper addresses the issue and formulates it as a novel learning problem, referred to as, 'learning to rank relational objects'. In the new learning task, the ranking model is defined as a function of not only the contents (features) of objects but also the relations between objects. The paper further focuses on one setting of the learning problem in which the way of using relation information is predetermined. It formalizes the learning task as an optimization problem in the setting. The paper then proposes a new method to perform the optimization task, particularly an implementation based on SVM. Experimental results show that the proposed method outperforms the baseline methods for two ranking tasks (Pseudo Relevance Feedback and Topic Distillation) in web search, indicating that the proposed method can indeed make effective use of relation information and content information in ranking.
Tao Qin 0001, Tie-Yan Liu, Xudong Zhang 0001, De-Sheng Wang, Wen-Ying Xiong, Hang Li 0001
WWW3
2008 Query-level loss functions for information retrieval
Tao Qin 0001, Xudong Zhang 0001, Ming-Feng Tsai, De-Sheng Wang, Tie-Yan Liu, Hang Li 0001
Inf. Process. Manag.2
2008 Content analysis based smart macroblock rearrangement for error resilience in wireless video transmission
Jian Feng 0006, Kwok-Tung Lo, Xudong Zhang 0001
J. Vis. Commun. Image Represent.4
2008 An active feedback framework for image retrieval
Tao Qin 0001, Xudong Zhang 0001, Tie-Yan Liu, De-Sheng Wang, Wei-Ying Ma, HongJiang Zhang
Pattern Recognit. Lett.2
2007 Ranking with multiple hyperplanes
abstract
The central problem for many applications in Information Retrieval is ranking and learning to rank is considered as a promising approach for addressing the issue. Ranking SVM, for example, is a state-of-the-art method for learning to rank and has been empirically demonstrated to be effective. In this paper, we study the issue of learning to rank, particularly the approach of using SVM techniques to perform the task. We point out that although Ranking SVM is advantageous, it still has shortcomings. Ranking SVM employs a single hyperplane in the feature space as the model for ranking, which is too simple to tackle complex ranking problems. Furthermore, the training of Ranking SVM is also computationally costly. In this paper, we look at an alternative approach to Ranking SVM, which we call "Multiple Hyperplane Ranker" (MHR), and make comparisons between the two approaches. MHR takes the divide-and-conquer strategy. It employs multiple hyperplanes to rank instances and finally aggregates the ranking results given by the hyperplanes. MHR contains Ranking SVM as a special case, and MHR can overcome the shortcomings which Ranking SVM suffers from. Experimental results on two information retrieval datasets show that MHR can outperform Ranking SVM in ranking.
Tao Qin 0001, Xudong Zhang 0001, De-Sheng Wang, Tie-Yan Liu, Hang Li 0001
SIGIR2
2007 Topic distillation via sub-site retrieval
Tao Qin 0001, Tie-Yan Liu, Xudong Zhang 0001, De-Sheng Wang, Wei-Ying Ma
Inf. Process. Manag.3
2006 Fast Robust Eigen-Background Updating for Foreground Detection
abstract
A fast robust eigen-background update algorithm is proposed for foreground object detection. The update procedure involves no eigen decomposition, thus faster than former eigen-background based algorithms. Meanwhile, the algorithm can robustly maintain the desired background model, resistant to outlying objects.
Xudong Zhang 0001
ICIP3
2006 Level-Biased Statistics in the Hierarchical Structure of the Web
Tie-Yan Liu, Xudong Zhang 0001, Wei-Ying Ma
PAKDD3
2006 AggregateRank: bringing order to web sites
abstract
Since the website is one of the most important organizational structures of the Web, how to effectively rank websites has been essential to many Web applications, such as Web search and crawling. In order to get the ranks of websites, researchers used to describe the inter-connectivity among websites with a so-called HostGraph in which the nodes denote websites and the edges denote linkages between websites (if and only if there are hyperlinks from the pages in one website to the pages in the other, there will be an edge between these two websites), and then adopted the random walk model in the HostGraph. However, as pointed in this paper, the random walk over such a HostGraph is not reasonable because it is not in accordance with the browsing behavior of web surfers. Therefore, the derivate rank cannot represent the true probability of visiting the corresponding website.In this work, we mathematically proved that the probability of visiting a website by the random web surfer should be equal to the sum of the PageRank values of the pages inside that website. Nevertheless, since the number of web pages is much larger than that of websites, it is not feasible to base the calculation of the ranks of websites on the calculation of PageRank. To tackle this problem, we proposed a novel method named AggregateRank rooted in the theory of stochastic complement, which cannot only approximate the sum of PageRank accurately, but also have a lower computational complexity than PageRank. Both theoretical analysis and experimental evaluation show that AggregateRank is a better method for ranking websites than previous methods.
Tie-Yan Liu, Ying Bao, Zhiming Ma, Xudong Zhang 0001, Wei-Ying Ma
SIGIR6
2005 Level-Based Link Analysis
Tie-Yan Liu, Xudong Zhang 0001, Tao Qin 0001, Bin Gao 0001, Wei-Ying Ma
APWeb3
2005 Subspace Clustering and Label Propagation for Active Feedback in Image Retrieval
abstract
In recent years, relevance feedback has been studied extensively as a way to improve performance of content-based image retrieval (CBIR). However, since users are usually unwilling to provide many feedbacks, the insufficiency of the training samples limited the success of relevance feedback. To tackle this problem, we propose two coupled algorithms: (i) overlapped subspace clustering to select representative images for user’s feedback; and (ii) multi-subspace label propagation to include unlabeled data in the training process. As these two algorithms are both working on sub feature spaces of the image database, they can not only deal with the insufficient training samples but also well capture the user’s attention during the retrieval process. Experimental results on a large database of general-purposed images demonstrated the high effectiveness of our proposed algorithms.
Tao Qin 0001, Tie-Yan Liu, Xudong Zhang 0001, Wei-Ying Ma, HongJiang Zhang
MMM3
2005 A study of relevance propagation for web search
abstract
Different from traditional information retrieval, both content and structure are critical to the success of Web information retrieval. In recent years, many relevance propagation techniques have been proposed to propagate content information between web pages through web structure to improve the performance of web search. In this paper, we first propose a generic relevance propagation framework, and then provide a comparison study on the effectiveness and efficiency of various representative propagation models that can be derived from this generic framework. We come to many conclusions that are useful for selecting a propagation model in real-world search applications, including 1) sitemap-based propagation models outperform hyperlink-based models in sense of both effectiveness and efficiency, and 2) sitemap-based term propagation is easier to be integrated into real-world search engines because of its parallel offline implementation and acceptable complexity. Some other more detailed study results are also reported in the paper.
Tao Qin 0001, Tie-Yan Liu, Xudong Zhang 0001, Zheng Chen 0001, Wei-Ying Ma
SIGIR3
2004 Color image segmentation in visual prostheses
abstract
The area of visual prosthesis by electrically stimulating the neural tissue of the blind person is more and more attractive with the development of electrode array and biological technology. Processing the images captured by the camera is the first step in the whole project. We use a new algorithm to get the accurate edge of color images. The result is better than that obtained through some edge operators like Canny, Soble. Then, we use simple methods to segment the image into regions that are based on gray level images of objects we are most interested in. Finally, we integrate the results from the two steps above. These methods are simple and have low computational complexity that can be used in real time process and in the volume-limited chip. The result is good in visual prostheses and also adapt to pixelized vision in later steps.
Yilun Cao, Xudong Zhang 0001
ICIG2
2004 A new full-pixel and sub-pixel motion vector search algorithm for fast block-matching motion estimation in H.264
abstract
H.264 is a new recommendation for moving picture coding proposed by JVT. In this standard, full pixel and sub-pixel motion estimation are the most important parts. The unequal-arm adaptive rood pattern search (APRS-3) algorithm has shown a good performance on search speed-up. In this paper, we propose an improved algorithm called APRS-4 which applies an early termination technology on APRS-3 by using adaptive threshold. This early termination technology is also applied on sub-pixel motion estimation, and a new fast sub-pixel motion estimation algorithm is proposed which uses the motion vector information of the neighbor blocks. Experimental results show that both algorithms have achieved a good speed-up ratio. Compared with full-pixel motion estimation, our new algorithm could keep image quality well with almost the same bit rate.
Baochen Jiang, Xudong Zhang 0001
ICIG3
2004 A new cut detection algorithm with constant false-alarm ratio for video segmentation
Tie-Yan Liu, Kwok-Tung Lo, Xudong Zhang 0001, Jian Feng 0006
J. Vis. Commun. Image Represent.3
2004 Shot reconstruction degree: a novel criterion for key frame selection
Tie-Yan Liu, Xudong Zhang 0001, Jian Feng 0006, Kwok-Tung Lo
Pattern Recognit. Lett.2
2003 Image watermarking using tree-based spatial-frequency feature of wavelet transform
Xudong Zhang 0001, Jian Feng 0006, Kwok-Tung Lo
J. Vis. Commun. Image Represent.1
2003 Dynamic selection and effective compression of key frames for video abstraction
Xudong Zhang 0001, Tie-Yan Liu, Kwok-Tung Lo, Jian Feng 0006
Pattern Recognit. Lett.1
2003 Frame interpolation scheme using inertia motion prediction
Tie-Yan Liu, Kwok-Tung Lo, Jian Feng 0006, Xudong Zhang 0001
Signal Process. Image Commun.4
2002 Constant false-alarm ratio processing for video cut detection
abstract
In this paper, a video cut detection algorithm with constant false-alarm ratio (CFAR) is proposed. Applying the non-parameter based CFAR processing technique from radar signal detection, a theoretical threshold determination strategy for video cut detection is developed, which results in controllable precision and evaluative recall performances. Detailed deductions and simulation results show that this algorithm leads to very good detection performance as compared to the previous works.
Tie-Yan Liu, Xudong Zhang 0001, Linwei Shan, Yingning Peng
ICIP (1)2