VLDB 2026 Research / reviewers in the wild / expert
Shangshang Wang
dblp:238/9234
· DBLP profile ↗
13ranked-venue papers
7as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 6 first-author · 10 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Neural Constrained Combinatorial BanditsabstractConstrained combinatorial contextual bandits have emerged as trending tools in intelligent systems and networks to model reward and cost signals under combinatorial decision-making. On one hand, both signals are complex functions of the context, e.g., in federated learning, training loss (negative reward) and energy consumption (cost) are nonlinear functions of edge devices’ system conditions (context). On the other hand, there are cumulative constraints on costs, e.g., the accumulated energy consumption should be budgeted by energy resources. Besides, real-time systems often require such constraints to be guaranteed anytime or in each round, e.g., ensuring anytime fairness for task assignment to maintain the credibility of crowdsourcing platforms for workers. This bandit setting presents significant challenges, including modeling complex rewards/costs, satisfying anytime cumulative constraints, and balancing exploration and exploitation. Therefore, we propose a primal-dual algorithm (Neural-PD) with neural network-based estimations for rewards/costs and virtual queue-based optimization for constraints. Besides, we provide theoretical guarantees regarding the behavior of neural network training within the primal-dual framework and the dynamic neural tangent kernel (NTK) of the neural networks during online learning. By integrating NTK theory and Lyapunov-drift techniques, we prove Neural-PD achieves a sharp regret bound and a zero constraint violation. We also show Neural-PD outperforms existing algorithms with extensive experiments on both synthetic and real-world datasets. Shangshang Wang, Simeng Bian, Xin Liu 0049, Ziyu Shao |
IEEE Trans. Netw. | 1 |
| 2024 | GNN-Aided Distributed GAN with Partially Observable Social GraphabstractThe proliferation of edge computing has facilitated the edge-based artificial intelligence-generated content (AIGC) for ubiquitous and distributed end devices. To exemplify, we focus on the distributed implementation of one established instance, generative adversarial network (GAN), yielding the distributed GAN task. Practically speaking, this task usually is impeded by concerns including the unknown latency (of processing and transmission), the fairness requirement induced by heterogeneous distributed data and the limited energy budget of end devices. Besides, an often neglected factor is how to exploit feedback from networked end devices among which social ties indicate the flow of shared information. In practice, such social ties are partially observable to lack of exact knowledge of users, e.g., resulted from scarce historical data and privacy issues. Under this setting, we propose an online algorithm via integration of 1) online learning aided by graph neural network (GNN), aiming to recover social ties with GNN-based edge prediction, for accelerated learning of uncertainty and 2) online control to adaptively guarantee the constraints. We theoretically show that it not only achieves a sub-linear regret with guaranteed energy consumption and fairness but also leads to a superior global GAN. We also conduct simulations to justify its outperformance over online baselines. Shangshang Wang, Ziyu Shao, Yang Yang 0001 |
WCNC | 2 |
| 2024 | Privacy-Preserving Edge Intelligence: A Perspective of Constrained BanditsabstractAdvanced edge systems have brought intelligence to networked end devices at the network edge. In such systems, privacy preservation has been an integral role since users' privacy may be violated via edge-device interaction given unsafe decision-making on information sharing. Therefore, we in this paper study privacy preservation for decision-making under bandit models. Particularly, a canonical bandit model features an agent that aims to maximize attainable rewards based on feedback from arm selection. However, upon application in edge systems, such feedback becomes more complex given 1) privacy concern and 2) non-negligible cost feedback. Confronting such concerns during decision-making, we study a privacy-preserving constrained bandit variant where we face the challenge of guaranteeing privacy preservation and within-budget cost while striving for high rewards. In this paper, we address the challenge with an integration of local differential privacy mechanism, online control, and online learning. Theoretically, we prove that our algorithm maintains adjustable privacy, adheres to cost constraints, and achieves a sub-linear regret (i.e., loss of reward). Numerically, we conduct simulations to demonstrate the outperformance of our algorithm over baselines. Shangshang Wang, Yinxu Tang, Ziyu Shao, Yang Yang 0001 |
WCNC | 2 |
| 2024 | Multi-Agent Systems for Collaborative Inference Based on Deep Policy Q-Inference Network
Shangshang Wang, Yuqin Jing, Kezhu Wang |
J. Grid Comput. | 1 |
| 2024 | Next-Word Prediction: A Perspective of Energy-Aware Distributed InferenceabstractThe pursuit of high-quality artificial intelligence generated contents (AIGC) with fast response has prompted the evolution of natural language processing (NLP) services, notably those enabled at the edge (i.e., edge NLP). For concreteness, we study distributed inference for next-word prediction which is a prevalent edge NLP service for mobile keyboards on user devices. Accordingly, we optimize coupled metrics,i.e., maximize prediction click-through rate (CTR) for improved quality-of-service (QoS), minimize user impatience for enhanced quality-of-experience (QoE), and keep energy consumption within budget for sustainability. Moreover, we consider the real-world setting where there is no prior knowledge of heterogeneous NLP models' prediction accuracy. Via an integration of online learning and online control, we propose a novel distributed inference algorithm for online next-word prediction with user impatience (DONUT) to estimate models' prediction accuracy and balance the trade-offs among coupled metrics. Our theoretical analysis reveals that DONUT achieves sub-linear regret (loss of CTR), ensures bounded user impatience, and maintains within-budget energy consumption. Through numerical simulations, we not only establish DONUT's superior performance over other baseline methods, but also demonstrate its adaptability to various settings. Shangshang Wang, Ziyu Shao, John C. S. Lui |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | Social-Aware Distributed Meta-Learning: A Perspective of Constrained Graphical BanditsabstractMeta-learning has earned its wide popularity to handle a family of similar tasks (e.g., classification of pets and wildlife) with elaborately trained meta-knowledge (e.g., shared network architecture and neural network parameter initialization). In this paper, we focus on the distributed training of meta-knowledge via server-device collaboration at the edge (i.e., distributed meta-learning). Notably, its practical implementation often runs into concerns like 1) time-varying unknown wireless dynamics (e.g., transmission latency); 2) device-side fair device involvement in distributed training; 3) server-side resource efficiency. To address such concerns, 1) we employ online learning to estimate the unknown dynamics and further exploit social ties among device users to accelerate online learning; 2) we utilize online control techniques to handle long-term fairness and resource constraints. By characterizing inter-user social ties as a social graph, we study distributed meta-learning from the perspective of constrained graphical bandits. Therefore, we propose a SoCial-awarE meta-kNowledge dispaTch (SCENT) algorithm by effectively integrating graphical bandit learning and online control. Besides a sublinear regret (i.e., loss of performance), SCENT also guarantees a well-trained meta-knowledge under within-budget resource consumption and fair device involvement. We conduct simulations to justify the outperformance of SCENT compared with baselines. Shangshang Wang, Simeng Bian, Yinxu Tang, Ziyu Shao |
ICC | 1 |
| 2023 | Green Dueling BanditsabstractThe dueling bandit model has been acknowledged as an efficient analytic tool for sequential decision-making problems with qualitative pairwise comparison. For example, the comparison of workers' completion quality of assigned tasks in crowdsourcing systems; user's ranking of recommended items in recommender systems. In dueling bandits, an agent uses pairwise comparisons of selected arms to balance the exploitation-exploration trade-off during the online learning of uncertainties. Despite the wide application of dueling bandits, their green implementation should also consider the non-neglectable energy costs for selecting arms, implying the green dueling bandit model. Particularly, it requires online control to optimize energy costs adaptively in the long run for sustainable system deployment. Therefore, we 1) employ online learning methods to learn the uncertainties via qualitative pairwise comparisons; 2) utilize online control techniques to guarantee a within-budget energy cost for the green real-world deployment. Accordingly, we propose a Green Dueling Bandit Learning (GDBL) algorithm to effectively integrate dueling bandit learning for the exploration-exploitation trade-off and online control for the optimization of energy costs. We prove that GDBL achieves a sublinear round-averaged regret while keeping the energy cost under budget. We conduct simulations to demonstrate the outperformance of GDBL over baselines. Shangshang Wang, Ziyu Shao |
ICC | 1 |
| 2023 | Online Learning-Based Beamforming for Rate-Splitting Multiple Access: A Constrained Bandit ApproachabstractRate-splitting multiple access (RSMA) has emerged as a potential non-orthogonal transmission strategy and powerful interference management scheme for 6G. Most of the existing works on RSMA beamforming design assume instantaneous or statistical channel state information (CSI) is available at the transmitter. Such an assumption however is impractical especially in massive multiple-input multiple-output (MIMO) due to the dynamic wireless environments and the challenges in channel estimation. In this work, we propose a novel beamforming design framework based on online learning and online control to adaptively learn the best precoding action for a RSMA-aided downlink massive MIMO without explicit CSI feedback. In particular, we first formulate the precoder selection problem that maximizes the ergodic sum-rate subject to a long-term transmit power constraint as a constrained combinatorial multi-armed bandit (CMAB) problem. Then we propose a precoder selection with bandit learning algorithm for RSMA (PBR). Our theoretical analysis shows that PBR achieves a sublinear regret bound with a long-term power constraint guarantee. Through experimental results, we not only verify our theoretical analysis but also demonstrate the outperformance of PBR in terms of sum-rate and power consumption compared with the conventional transmission schemes without using RSMA. Shangshang Wang, Jingye Wang, Yijie Mao, Ziyu Shao |
ICC | 1 |
| 2023 | Neural Constrained Combinatorial BanditsabstractConstrained combinatorial contextual bandits have emerged as trending tools in intelligent systems and networks to model reward and cost signals under combinatorial decision-making. On one hand, both signals are complex functions of the context, e.g., in federated learning, training loss (negative reward) and energy consumption (cost) are nonlinear functions of edge devices’ system conditions (context). On the other hand, there are cumulative constraints on costs, e.g., the accumulated energy consumption should be budgeted by energy resources. Besides, real-time systems often require such constraints to be guaranteed anytime or in each round, e.g., ensuring anytime fairness for task assignment to maintain the credibility of crowdsourcing platforms for workers. This setting imposes a challenge on how to simultaneously achieve reward maximization while subjecting to anytime cumulative constraints. To address such challenge, we propose a primal-dual algorithm (Neural-PD) whose primal component adopts multi-layer perceptrons to estimate reward and cost functions, and its dual component estimates the Lagrange multiplier with the virtual queue. By integrating neural tangent kernel theory and Lyapunov-drift techniques, we prove Neural-PD achieves a sharp regret bound and a zero constraint violation. We also show Neural-PD outperforms existing algorithms with extensive experiments on both synthetic and real-world datasets. Shangshang Wang, Simeng Bian, Xin Liu 0049, Ziyu Shao |
INFOCOM | 1 |
| 2022 | Social-Aware Edge Intelligence: A Constrained Graphical Bandit ApproachabstractThe flourished edge intelligence has motivated the execution of machine learning tasks at the network edge. In this paper, we focus on distributing training, one of the core tasks, that is carried out by an edge server of limited communication capacity and multiple end devices. In distributed training, the key issue for the edge server is how to dynamically select a proper subset of end devices to periodically participate in the training. Such a dynamic end device selection problem is hindered by concerns like 1) unknown system dynamics, e.g., transmission latencies; 2) limited energy resources on end devices; and 3) unbalanced and non-IID data distribution over end devices. Therefore, the core challenge lies in the coordination of online learning and online control to fulfill both efficient learning of unknown statistics and guarantees of within-budget energy consumption and fairness selection. To address the above challenge, we first characterize the social ties among users of end devices as a social graph and then formulate the dynamic end device selection problem from the perspective of constrained graphical bandits. Under the formulation, we propose GRIND to effectively integrate graphical bandit learning methods with Lyapunov-drift techniques. The theoretical superiority of GRIND is not only 1) the achieved sub-linear round-averaged regret with satisfied long-term constraints but also 2) the characterization of graph structure with the independence number. Extensive simulations also verify the effectiveness of GRIND in terms of both latency reduction and long-term constraint satisfaction. Simeng Bian, Shangshang Wang, Yinxu Tang, Ziyu Shao |
GLOBECOM | 2 |
| 2022 | Memory Enhanced Spatial-Temporal Graph Convolutional Autoencoder for Human-Related Video Anomaly Detection
Sibo Luo, Shangshang Wang, Yuan Wu 0004, Cheng Jin 0001 |
PRCV (3) | 2 |
| 2022 | Decentralized Multi-Agent Bandit Learning for Intelligent Internet of Things SystemsabstractIn intelligent Internet of Things systems, data-hungry services are empowered by data collection, which is jointly accomplished by edge servers and data-collecting sensors. In this paper, we aim to achieve efficient data collection, i.e., maximize data rates from sensors to servers while mitigating the impact of data heterogeneity for data collected from sensors. Considering geographically distributed servers and sensors, we study the problem from the perspective of multi-agent multi-armed bandits. The key ideas of our approach are to 1) establish associations between servers and sensors under unknown wireless dynamics (i.e., channel state information) and selection fraction constraints; 2) utilize shared information via pairwise communication between servers to mitigate biased observations for data rates. To this end, we propose a scheme that leverages online learning to reduce uncertainties in wireless dynamics and online control to mitigate the impact of data heterogeneity. Based on an effective integration of bandit learning methods under pairwise communication and Lyapunov optimization techniques, we present a novel Decentralized sErver-Sensor association scheme with Multi-Agent learning under pairwise communication (DESMA). Our theoretical analysis demonstrates that DESMA achieves a tunable trade-off between maximizing data rate and mitigating the impact of data heterogeneity. Qiuyu Leng, Shangshang Wang, Xi Huang 0001, Ziyu Shao, Yang Yang 0001 |
WCNC | 2 |
| 2019 | Research on dynamic optimal control strategy of distributed super capacitor energy storage system based on convolution neural networkabstractSummary Compared with other modes of transport, urban rail transit has significant advantages such as large capacity, punctual safety, energy saving, and environmental protection. For pure electric buses, the problem of slow start, large battery energy loss, short battery life, and insufficient recovery of braking energy is caused by the use of single energy only. The power battery and super capacitor are combined to form a compound power supply to solve the problem. In order to improve the feature extraction ability of convolution neural networks, a deep convolution neural network model based on continuous convolution is designed. The model uses small scale convolution kernel to extract local features more carefully, and increases the nonlinear expression ability of the model by means of two continuous coiling layers. Dropout technology is used to reduce the interdependence between neurons, and the model is optimized by suppressing network overfitting. If the battery and the supercapacitor are combined to form a dual energy source system, it can not only meet the instantaneous high power demand of the electric vehicle but also prolong the life of the battery. Increasing the mileage of electric vehicles has become the development direction of pure electric vehicles. Due to its fast charging and discharging speed, long cycle life, and friendly environment, supercapacitor has unique advantages and development prospects in solving the problem of urban rail transit regeneration failure, performance analysis, modeling, and parameter identification of ultracapacitor. The performance of ultracapacitor is analyzed, and an improved dynamic equivalent circuit model of ultracapacitor is established. Therefore, this paper studies the control strategy of dual energy source storage system and the driving range of the dual energy source pure electric vehicle, which has a broad application prospect and practical significance. Hongyun Chen, Shangshang Wang |
Concurr. Comput. Pract. Exp. | 3 |