Xiao Zhang 0004

dblp:49/4478-4 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0003-4927-5016ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 39% Deep learning architectures and training · 30% Trustworthy machine learning · 30%
Computer networks
1 paper
Network optimization and economics · 100%
Network and information security
2 papers
Network security · 70% Privacy and data protection · 30%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
offline reinforcement learning
0.912025
Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood · ICLR 2025
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.912025
Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood · ICLR 2025
Machine learning › Deep learning architectures and training
physics-informed neural network
0.912025
Score-based free-form architectures for high-dimensional Fokker-Planck equations · ICLR 2025
Computational science and engineering › scientific machine learning › physics-informed machine learning › physics-informed neural networks
partial differential equation solving
0.912025
Score-based free-form architectures for high-dimensional Fokker-Planck equations · ICLR 2025
Network optimization and economics › resource allocation
bandwidth allocation
0.712023
Differential Pricing Strategies for Bandwidth Allocation With LFA Resilience: A Stackelberg Game Approach · IEEE Trans. Inf. Forensics Secur. 2023
Network optimization and economics › pricing
differential pricing
0.712023
Differential Pricing Strategies for Bandwidth Allocation With LFA Resilience: A Stackelberg Game Approach · IEEE Trans. Inf. Forensics Secur. 2023
Network optimization and economics
pricing
0.712023
Differential Pricing Strategies for Bandwidth Allocation With LFA Resilience: A Stackelberg Game Approach · IEEE Trans. Inf. Forensics Secur. 2023
Network optimization and economics
resource allocation
0.712023
Differential Pricing Strategies for Bandwidth Allocation With LFA Resilience: A Stackelberg Game Approach · IEEE Trans. Inf. Forensics Secur. 2023
Network security › attack strategy › denial-of-service attack
DDoS attack
0.712023
Differential Pricing Strategies for Bandwidth Allocation With LFA Resilience: A Stackelberg Game Approach · IEEE Trans. Inf. Forensics Secur. 2023
Network security › attack strategy › denial-of-service attack
link flooding attack
0.712023
Differential Pricing Strategies for Bandwidth Allocation With LFA Resilience: A Stackelberg Game Approach · IEEE Trans. Inf. Forensics Secur. 2023
Privacy and data protection
privacy-preserving data sharing
0.612022
Labrador: towards fair and auditable data sharing in cloud computing with long-term privacy · Sci. China Inf. Sci. 2022
Machine learning › Reinforcement learning › dynamic programming
bellman operator
0.312025
Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood · ICLR 2025
Software maintenance and evolution
software trustworthiness
0.112012
Optimized statistical analysis of software trustworthiness attributes · Sci. China Inf. Sci. 2012

Methods — techniques the papers use, named apart from their topics

score matching · 1.7physics-informed neural networks · 1.7stackelberg game · 1.3game theory · 1.3smooth bellman operator · 0.9convex hull neighborhood · 0.9cryptography · 0.6statistical analysis · 0.1
YearPublicationVenuePosition
2025 Score-based free-form architectures for high-dimensional Fokker-Planck equations
abstract
Deep learning methods incorporate PDE residuals as the loss function for solving Fokker-Planck equations, and usually impose the proper normalization condition to avoid a trivial solution. However, soft constraints require careful balancing of multi-objective loss functions, and specific network architectures may limit representation capacity under hard constraints. In this paper, we propose a novel framework: Fokker-Planck neural network (FPNN) that adopts a score PDE loss to decouple the score learning and the density normalization into two stages. Our method allows free-form network architectures to model the unnormalized density and strictly satisfy normalization constraints by post-processing. We demonstrate the effectiveness on various high-dimensional steady-state Fokker-Planck (SFP) equations, achieving superior accuracy and over a 20$\times$ speedup compared to state-of-the-art methods. Without any labeled data, FPNNs achieve the mean absolute percentage error (MAPE) of 11.36%, 13.87% and 12.72% for 4D Ring, 6D Unimodal and 6D Multi-modal problems respectively, requiring only 256, 980, and 980 parameters. Experimental results highlights the potential as a universal fast solver for handling more than 20-dimensional SFP equations, with great gains in efficiency, accuracy, memory and computational resource usage.
Faguo Wu, Xiao Zhang 0004
ICLR3
2025 Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood
abstract
Offline Reinforcement Learning (RL) struggles with distributional shifts, leading to the $Q$-value overestimation for out-of-distribution (OOD) actions. Existing methods address this issue by imposing constraints; however, they often become overly conservative when evaluating OOD regions, which constrains the $Q$-function generalization. This over-constraint issue results in poor $Q$-value estimation and hinders policy improvement. In this paper, we introduce a novel approach to achieve better $Q$-value estimation by enhancing $Q$-function generalization in OOD regions within Convex Hull and its Neighborhood (CHN). Under the safety generalization guarantees of the CHN, we propose the Smooth Bellman Operator (SBO), which updates OOD $Q$-values by smoothing them with neighboring in-sample $Q$-values. We theoretically show that SBO approximates true $Q$-values for both in-sample and OOD actions within the CHN. Our practical algorithm, Smooth Q-function OOD Generalization (SQOG), empirically alleviates the over-constraint issue, achieving near-accurate $Q$-value estimation. On the D4RL benchmarks, SQOG outperforms existing state-of-the-art methods in both performance and computational efficiency. Code is available at <https://github.com/yqpqry/SQOG>.
Qingmao Yao, Zhichao Lei, Tianyuan Chen, Ziyue Yuan, Xuefan Chen, Jianxiang Liu, Faguo Wu, Xiao Zhang 0004
ICLR8
2025 PIONM: A Generalized Approach to Solving Density-Constrained Mean-Field Games Equilibrium under Modified Boundary Conditions
abstract
Neural network-based methods are effective for solving equilibria in Mean-Field Games (MFGs), particularly in high-dimensional settings. However, solving the coupled partial differential equations (PDEs) in MFGs limits their applicability since solving coupled PDEs is computationally expensive. Additionally, modifying boundary conditions, such as the initial state distribution or terminal value function, necessitates extensive retraining, reducing scalability. To address these challenges, we propose a generalized framework, PIONM (Physics-Informed Neural Operator NF-MKV Net), which leverages physics-informed neural operators to solve MFGs equations. PIONM utilizes neural operators to compute MFGs equilibria for arbitrary boundary conditions. The method encodes boundary conditions as input features and trains the model to align them with density evolution, modeled using discrete-time normalizing flows. Once trained, the algorithm efficiently computes the density distribution at any time step for modified boundary condition, ensuring efficient adaptation to different boundary conditions in MFGs equilibria. Unlike traditional MFGs methods constrained by fixed coefficients, PIONM efficiently computes equilibria under varying boundary conditions, including obstacles, diffusion coefficients, initial densities, and terminal functions. PIONM can adapt to modified conditions while preserving density distribution constraints, demonstrating superior scalability and generalization capabilities compared to existing methods.
Xiao Zhang 0004
IJCNN3
2025 Decentralized Reward Allocation Mechanism with Sybil Resilience: The Case of Stake Pools
abstract
Proof-of-Stake (PoS) is an energy-efficient consensus where proposer election based on validator stakes causes centralization, especially in stake pools. However, existing reward allocation mechanisms aim to address centralization but increase the risk of Sybil attacks. Therefore, designing a reward allocation mechanism for stake pools that reconciles decentralization and Sybil resilience remains a key challenge. In this paper, we propose Dream-SR, a reward allocation mechanism established with formal function properties. Dream-SR enhances decentralization by imposing reward constraints to limit the dominance of large stake pools. Sybil resilience is ensured by aligning reward allocation with both the economic incentives and the influence, effectively addressing utility-maximizing and influence-maximizing Sybil attacks. Furthermore, we leverage a multi-leader-multi-follower (MLMF) Stackelberg game model to capture the interactions between validators and users regarding commission pricing and stake delegation within stake pools. The game model is used to analyze the impact of the reward allocation mechanism on the strategies of participants and the system equilibrium. Compared with other relevant mechanisms, numerical results reveal that Dream-SR effectively improves decentralization and resilience to Sybil attacks at equilibrium.
Binfeng Song, Lijia Xie, Xiao Zhang 0004
IWQoS3
2025 Integrated Task Assignment and Trajectory Planning for a Massive Number of Agents Based on Bilayer-Coupled Mean Field Games
abstract
Aiming at the problem of integrated task assignment and trajectory planning of a massive number of agents in the scenario with different priority task nodes and multiple static obstacles, this paper proposes a general framework based on bilayer-coupled mean field games, which couples the minimum cost of trajectory planning of an agent in the task assignment process to achieve a reasonable, globally optimal, and targeted adjustable assignment result. In the proposed general framework, firstly, the multi-population mean field game is used to plan the optimal trajectory of an agent between each pair of priority adjacent task nodes, and the minimum costs are calculated. Then, based on the discrete time finite state space mean field game, a task assignment model in the discrete task space is constructed, and the minimum costs obtained in the trajectory planning are coupled into the model as a reference, the task assignment strategies are finally obtained. Moreover, we give a specific example of the proposed general framework and prove the existence of equilibrium solutions of two mean field games. The effectiveness of the proposed general framework is demonstrated through simulation experiments and results analysis.Note to Practitioners—In multi-agent decision-making and control, task assignment and trajectory planning are two fundamental problems that coexist in many scenarios. Examples include the collaborative exploration of multiple task areas by UAV swarm, and the lane selection and efficient driving of autonomous vehicles. There are many methods for dealing with the integrated task assignment and trajectory planning. However, they have difficulties in dealing with large-scale agent problems, mainly due to the significant increase in communication and computation costs as the number of agents increases. In response to this problem, based on the characteristic of mean field game that transforms the game between individuals into a game between an individual and the whole, this paper proposes a general framework of bilayer-coupled mean field games for the scenario with different priority task nodes and multiple static obstacles. The multi-population mean field game is used to plan the optimal trajectory of agents, and the discrete time finite state space mean field game is utilized for task assignment, in which the cost of trajectory planning between task nodes is considered. We propose a specific example model and theoretically prove the existence of an equilibrium solution of this model. The effectiveness of the general framework is verified by simulation experiments and results analysis.
Zijia Niu, Sanjin Huang, Xiao Zhang 0004, Langyu Qian
IEEE Trans Autom. Sci. Eng.5
2024 Uncertainty modified policy for multi-agent reinforcement learning
Jianxiang Liu, Faguo Wu, Xiao Zhang 0004, Guojian Wang
Appl. Intell.4
2024 Adaptive trajectory-constrained exploration strategy for deep reinforcement learning
Guojian Wang, Faguo Wu, Xiao Zhang 0004, Zhiming Zheng 0001
Knowl. Based Syst.3
2023 Learning Diverse Policies with Soft Self-Generated Guidance
abstract
Reinforcement learning (RL) with sparse and deceptive rewards is a significant challenge because nonzero rewards are rarely obtained, and hence, the gradient calculated by the agent can be stochastic and without valid information. Recent work demonstrates that using memory buffers of previous experiences can lead to a more efficient learning process. However, existing methods usually require these experiences to be successful and may overly exploit them, which can cause the agent to adopt suboptimal behaviors. This study develops an approach that exploits diverse past trajectories for faster and more efficient online RL, even if these trajectories are suboptimal or not highly rewarded. The proposed algorithm merges a policy improvement step with an additional policy exploration step by using offline demonstration data. The main contribution of this study is that by regarding diverse past trajectories as guidance, instead of imitating them, our method directs its policy to follow and expand past trajectories, while still being able to learn without rewards and gradually approach optimality. Furthermore, a novel diversity measurement is introduced to maintain the diversity of the team and regulate exploration. The proposed algorithm is evaluated on a series of discrete and continuous control tasks with sparse and deceptive rewards. In comparison with the existing RL methods, the experimental results indicate that our proposed algorithm is significantly better than the baseline methods in terms of diverse exploration and avoiding local optima.
Guojian Wang, Faguo Wu, Xiao Zhang 0004, Jianxiang Liu
Int. J. Intell. Syst.3
2023 Differential Pricing Strategies for Bandwidth Allocation With LFA Resilience: A Stackelberg Game Approach
abstract
Link flooding attacks (LFAs) have always been a security concern as the impact of volumetric attacks on transit links are increasingly severe. Capacity expansion, while being effective in combating LFAs, involves considerable deployment costs. Therefore, how to efficiently manage the link resource among spatio-temporal dynamic customers remains a challenge for Internet service providers (ISPs). In this paper, we study the differential pricing strategy for bandwidth allocation with LFA resilience by leveraging a multi-leader-multi-follower (MLMF) Stackelberg game approach. Based on network capabilities, we employ pricing approaches instead of empirical assignments to regulate the privileged channel allocation, economically facilitating the domain-level resource coordination. And we formulate the bandwidth pricing and allocation decision problem by using a Stackelberg game-theoretic approach to capture the interactions between providers and customers. Then, we analyze the Stackelberg game equilibrium of uniform pricing strategy and differential pricing strategy, where differential pricing is applied to adjust the prices that individual customers receive according to heterogeneous factors. Furthermore, we give the optimal solution derivation in the case of differential pricing. By comparing with other relevant pricing strategies, our numerical results show that the differential pricing strategy achieves congestion-free property and provides enough incentives for ISPs to deploy.
Lijia Xie, Xiao Zhang 0004
IEEE Trans. Inf. Forensics Secur.4
2022 Labrador: towards fair and auditable data sharing in cloud computing with long-term privacy
Xiaojie Guo 0004, Jin Li 0002, Zheli Liu, Yu Wei 0007, Xiao Zhang 0004, Changyu Dong
Sci. China Inf. Sci.5
2021 Fine-Grained Intra-domain Bandwidth Allocation Against DDoS Attack
Lijia Xie, Xiao Zhang 0004, Yiming Shi, Zhiming Zheng 0001
SecureComm (1)3
2020 Spatio-temporal heterogeneous bandwidth allocation mechanism against DDoS attack
Xiao Zhang 0004, Lijia Xie
J. Netw. Comput. Appl.1
2019 Multi-authority attribute-based encryption scheme with constant-size ciphertexts and user revocation
abstract
Summary Ciphertext‐policy attribute‐based encryption (CP‐ABE) is regarded as one of the most suitable technologies for data access control in cloud storage system. It gives data owners direct and flexible control on access policies. However, there still exists practicality concerns in CP‐ABE applications, for example, the key escrow problem, user revocability, and large ciphertext size. Considering these problems, we propose a multi‐authority attribute‐based encryption scheme with constant‐size ciphertexts and user revocation for threshold access policy in this paper. The security proof shows that the proposed scheme is selectively secure under the augmented multi‐sequence of exponents decisional Diffie‐Hellman assumption, and it also achieves forward security, backward security, and collusion‐resistance.
Xiao Zhang 0004, Faguo Wu
Concurr. Comput. Pract. Exp.1
2019 Lattice based signature with outsourced revocation for Multimedia Social Networks in cloud computing
Faguo Wu, Xiao Zhang 0004, Zhiming Zheng 0001
Multim. Tools Appl.3
2017 Efficient Bloom filter for network protocols using AES instruction set
abstract
The Internet continues to flourish, while an increasing number of network applications are found deploying Bloom filters. However, the heterogeneity of the Bloom filter realisations complicates the utilisation of relevant applications. Moreover, when applying Bloom filter to traffic that usually has a gigabit capacity, even insignificant delays will accumulate and restrict the effectiveness of the real‐time protocols. In this study, the authors present a Bloom filter construction that can be easily and consistently adopted at network nodes, with also considerable processing speed. Specifically, the authors show that AES‐based hashes are adequate to create Bloom filters correctly. Then they illustrate how AES new instructions (AES‐NI) can be leveraged to accelerate the Bloom filter realisation. According to the authors' experimental results, the proposed Bloom filter enables the best speed performance compared to the competing approaches.
Zhiming Zheng 0001, Xiao Zhang 0004
IET Commun.3
2017 An efficient image encryption algorithm based on a novel chaotic map
Chengqi Wang, Xiao Zhang 0004, Zhiming Zheng 0001
Multim. Tools Appl.2
2017 An improved biometrics based authentication scheme using extended chaotic maps for multimedia medicine information systems
Chengqi Wang, Xiao Zhang 0004, Zhiming Zheng 0001
Multim. Tools Appl.2
2012 Optimized statistical analysis of software trustworthiness attributes
Xiao Zhang 0004, Wei Li 0022, Zhiming Zheng 0001
Sci. China Inf. Sci.1