VLDB 2026 Research / reviewers in the wild / expert
Yinxu Tang
dblp:252/7419
· DBLP profile ↗
13ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-4254-8782ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Does Your AI Agent Get You? A Personalizable Framework for Approximating Human Models from Argumentation-based Dialogue TracesabstractExplainable AI is increasingly employing argumentation methods to facilitate interactive explanations between AI agents and human users. While existing approaches typically rely on predetermined human user models, there remains a critical gap in dynamically learning and updating these models during interactions. In this paper, we present a framework that enables AI agents to adapt their understanding of human users through argumentation-based dialogues. Our approach, called Persona, draws on prospect theory and integrates a probability weighting function with a Bayesian belief update mechanism that refines a probability distribution over possible human models based on exchanged arguments. Through empirical evaluations with human users in an applied argumentation setting, we demonstrate that Persona effectively captures evolving human beliefs, facilitates personalized interactions, and outperforms state-of-the-art methods. Yinxu Tang, Stylianos Loukas Vasileiou, William Yeoh 0001 |
AAAI | 1 |
| 2025 | Model Reconciliation via Cost-Optimal Explanations in Probabilistic Logic ProgrammingabstractIn human-AI interaction, effective communication relies on aligning the AI agent’s model with the human user’s mental model, a process known as model reconciliation. However, existing model reconciliation approaches predominantly assume deterministic models, overlooking the fact that human knowledge is often uncertain or probabilistic.
To bridge this gap, we present a probabilistic model reconciliation framework that resolves inconsistencies in MPE outcome probabilities between an agent’s and a user’s models.
Our approach is built on probabilistic logic programming (PLP) using ProbLog, where explanations are generated as cost-optimal model updates that reconcile these probabilistic differences.
We develop two search algorithms -- a generic baseline and an optimized version.
The latter is guided by theoretical insights and further extended with greedy and weighted variants to enhance scalability and efficiency.
Our approach is validated through a user study on explanation types and computational experiments showing that the optimized version consistently outperforms the generic baseline. Yinxu Tang, Stylianos Loukas Vasileiou, Vincent Derkinderen, William Yeoh 0001 |
NeurIPS | 1 |
| 2024 | Privacy-Preserving Edge Intelligence: A Perspective of Constrained BanditsabstractAdvanced edge systems have brought intelligence to networked end devices at the network edge. In such systems, privacy preservation has been an integral role since users' privacy may be violated via edge-device interaction given unsafe decision-making on information sharing. Therefore, we in this paper study privacy preservation for decision-making under bandit models. Particularly, a canonical bandit model features an agent that aims to maximize attainable rewards based on feedback from arm selection. However, upon application in edge systems, such feedback becomes more complex given 1) privacy concern and 2) non-negligible cost feedback. Confronting such concerns during decision-making, we study a privacy-preserving constrained bandit variant where we face the challenge of guaranteeing privacy preservation and within-budget cost while striving for high rewards. In this paper, we address the challenge with an integration of local differential privacy mechanism, online control, and online learning. Theoretically, we prove that our algorithm maintains adjustable privacy, adheres to cost constraints, and achieves a sub-linear regret (i.e., loss of reward). Numerically, we conduct simulations to demonstrate the outperformance of our algorithm over baselines. Shangshang Wang, Yinxu Tang, Ziyu Shao, Yang Yang 0001 |
WCNC | 3 |
| 2024 | Green Edge Intelligence Scheme for Mobile Keyboard Emoji PredictionabstractEmoji prediction has been widely adopted in most mobile keyboards to improve the quality of user experience. Considering the resource constraints of smartphones, it is promising to deploy well-trained prediction models on edge servers, with which smartphones can carry out emoji prediction in an online fashion. However, a key issue in such a scenario lies in how the smartphone should select a subset of models to achieve high-accuracy and real-time emoji prediction with energy efficiency (a.k.a.themodel selectionproblem). Moreover, part of the system dynamics such as the prediction accuracy and the inference latency of each model are usually unknowna prioriin practice, further complicating the problem. In this paper, with an effective integration of history-aware online learning and online control, we propose the first green edge intelligence scheme to solve the model selection problem for mobile keyboard emoji prediction. Our theoretical analysis and simulation results verify the effectiveness of our proposed scheme in achieving a sub-linear round-averaged regret bound and energy efficiency with a high prediction accuracy and a low latency. Yinxu Tang, Jianfeng Hou, Xi Huang 0001, Ziyu Shao, Yang Yang 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | Social-Aware Distributed Meta-Learning: A Perspective of Constrained Graphical BanditsabstractMeta-learning has earned its wide popularity to handle a family of similar tasks (e.g., classification of pets and wildlife) with elaborately trained meta-knowledge (e.g., shared network architecture and neural network parameter initialization). In this paper, we focus on the distributed training of meta-knowledge via server-device collaboration at the edge (i.e., distributed meta-learning). Notably, its practical implementation often runs into concerns like 1) time-varying unknown wireless dynamics (e.g., transmission latency); 2) device-side fair device involvement in distributed training; 3) server-side resource efficiency. To address such concerns, 1) we employ online learning to estimate the unknown dynamics and further exploit social ties among device users to accelerate online learning; 2) we utilize online control techniques to handle long-term fairness and resource constraints. By characterizing inter-user social ties as a social graph, we study distributed meta-learning from the perspective of constrained graphical bandits. Therefore, we propose a SoCial-awarE meta-kNowledge dispaTch (SCENT) algorithm by effectively integrating graphical bandit learning and online control. Besides a sublinear regret (i.e., loss of performance), SCENT also guarantees a well-trained meta-knowledge under within-budget resource consumption and fair device involvement. We conduct simulations to justify the outperformance of SCENT compared with baselines. Shangshang Wang, Simeng Bian, Yinxu Tang, Ziyu Shao |
ICC | 3 |
| 2022 | Social-Aware Edge Intelligence: A Constrained Graphical Bandit ApproachabstractThe flourished edge intelligence has motivated the execution of machine learning tasks at the network edge. In this paper, we focus on distributing training, one of the core tasks, that is carried out by an edge server of limited communication capacity and multiple end devices. In distributed training, the key issue for the edge server is how to dynamically select a proper subset of end devices to periodically participate in the training. Such a dynamic end device selection problem is hindered by concerns like 1) unknown system dynamics, e.g., transmission latencies; 2) limited energy resources on end devices; and 3) unbalanced and non-IID data distribution over end devices. Therefore, the core challenge lies in the coordination of online learning and online control to fulfill both efficient learning of unknown statistics and guarantees of within-budget energy consumption and fairness selection. To address the above challenge, we first characterize the social ties among users of end devices as a social graph and then formulate the dynamic end device selection problem from the perspective of constrained graphical bandits. Under the formulation, we propose GRIND to effectively integrate graphical bandit learning methods with Lyapunov-drift techniques. The theoretical superiority of GRIND is not only 1) the achieved sub-linear round-averaged regret with satisfied long-term constraints but also 2) the characterization of graph structure with the independence number. Extensive simulations also verify the effectiveness of GRIND in terms of both latency reduction and long-term constraint satisfaction. Simeng Bian, Shangshang Wang, Yinxu Tang, Ziyu Shao |
GLOBECOM | 3 |
| 2022 | Learning-Aided Stable Matching for Switch-Controller Association in SDN SystemsabstractThe scheme design of switch-controller association is an essential problem for software-defined networking (SDN) systems. A natural idea is to address the problem from the perspective of stable matching, since each switch (controller) often prefers to be associated with those controllers (switches) of lower communication costs and control traffic overhead. However, in practice, such system dynamics are usually unknown a priori, making it a challenging open problem. In this paper, we study such a problem of stable matching between switches and controllers with unknown communication costs from the perspective of multi-agent multi-armed bandit (MAMAB) learning. By integrating stable matching with online learning, we propose an effective Learning-aided Switch-controller Stable Matching (LS2M) scheme. Our theoretical analysis shows that LS2M effectively achieves a switch-optimal stable matching with a sublinear regret bound over time slots. Moreover, we conduct numerical simulations to verify the outperformance of LS2M over various baseline schemes. Yinxu Tang, Xi Huang 0001, Ziyu Shao, Yang Yang 0001 |
ICC | 1 |
| 2021 | Green Edge Intelligence Scheme for Mobile Keyboard Emoji PredictionabstractEmoji prediction has been widely adopted in most mobile keyboards to improve the quality of user experience. Considering the energy limitations of smartphones, it is promising to consider deploying pre-trained prediction models on edge servers, with which smartphones can carry out emoji prediction in an online fashion. However, given a limited connection capacity, a key issue under such a scheme lies in how each smartphone should select a subset of models to achieve high-accuracy and real-time emoji prediction with energy efficiency (a.k.a. the model selection problem). Moreover, part of the system dynamics such as the accuracy and the latency of individual models are usually unknown a priori in practice, further complicating the problem. In this paper, with an effective integration of history-aware online learning and online control, we propose the first green edge intelligence scheme to solve the model selection problem for edge-assisted mobile keyboard emoji prediction. Our theoretical analysis and simulation results verify the effectiveness of our proposed scheme in achieving a sublinear regret bound and energy efficiency with high accuracy and low latency. Jianfeng Hou, Yinxu Tang, Xi Huang 0001, Ziyu Shao, Yang Yang 0001 |
ICC | 2 |
| 2021 | History-Aware Online Cache Placement in Fog-Assisted IoT Systems: An Integration of Learning and ControlabstractIn fog-assisted Internet-of-Things systems, it is a common practice to cache popular content at the network edge to achieve high quality of service. Due to uncertainties, in practice, such as unknown file popularities, the cache placement scheme design is still an open problem with unresolved challenges: 1) how to maintain time-averaged storage costs under budgets; 2) how to incorporate online learning to aid cache placement to minimize performance loss [also known as (a.k.a.) regret]; and 3) how to exploit offline historical information to further reduce regret. In this article, we formulate the cache placement problem with unknown file popularities as a constrained combinatorial multiarmed bandit problem. To solve the problem, we employ virtual queue techniques to manage time-averaged storage cost constraints, and adopt history-aware bandit learning methods to integrate offline historical information into the online learning procedure to handle the exploration–exploitation tradeoff. With an effective combination of online control and history-aware online learning, we devise a cache placement scheme with history-aware bandit learning calledCPHBL. Our theoretical analysis and simulations show that CPHBL achieves a sublinear time-averaged regret bound. Moreover, the simulation results verify CPHBL’s advantage over the deep reinforcement learning-based approach. Xin Gao 0019, Xi Huang 0001, Yinxu Tang, Ziyu Shao, Yang Yang 0001 |
IEEE Internet Things J. | 3 |
| 2021 | Joint Switch-Controller Association and Control Devolution for SDN Systems: An Integrated Online Perspective of Control and LearningabstractIn software-defined networking (SDN) systems, it is a common practice to adopt a multi-controller design and control devolution techniques to improve the performance of the control plane. However, in such systems the decision-making for joint switch-controller association and control devolution often involves various uncertainties, e.g., the temporal variations of controller accessibility, and computation and communication costs of switches. In practice, statistics of such uncertainties are unattainable and need to be learned in an online fashion, calling for an integrated design of learning and control. In this article, we formulate a stochastic network optimization problem that aims to minimize time-average system costs and ensure queue stability. By transforming the problem into a combinatorial multi-armed bandit problem with long-term stability constraints, we adopt bandit learning methods and optimal control techniques to handle the exploration-exploitation tradeoff and long-term stability constraints, respectively. Through an integrated design of online learning and online control, we propose an effective Learning-Aided Switch-Controller Association and Control Devolution (LASAC) scheme. Our theoretical analysis and simulation results show that LASAC achieves a tunable tradeoff between queue stability and system cost reduction with a sublinear time-averaged regret bound over a finite time horizon. Xi Huang 0001, Yinxu Tang, Ziyu Shao, Yang Yang 0001, Hong Xu 0001 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2020 | Proactive Cache Placement with Bandit Learning in Fog-Assisted IoT SystemsabstractIn fog-assisted IoT systems, it is a common practice to cache popular content at the network edge to achieve high quality of service. Due to various uncertainties such as unknown file popularities in practice, the design of effective cache placement scheme is still an open problem with two key challenges: 1) how to incorporate online learning into the cache placement process to minimize performance loss (a.k.a. regret), and 2) how to maintain caching costs under budgets in the long run. In this paper, we formulate the content cache placement problem with unknown file popularities as a combinatorial multi-armed bandit (CMAB) problem with long-term time-average constraints. We adopt bandit learning methods and virtual queue technique to deal with the exploration-exploitation tradeoff and long-term time-average constraints, respectively. With an effective integration of online learning and online control, we devise a learning-aided cache placement scheme called CPB (Cache Placement with Bandit Learning). Our theoretical analysis and simulation results show that CPB achieves a tunable sublinear regret over a finite time horizon and keeps caching costs within budgets in the long run. Xin Gao 0019, Xi Huang 0001, Yinxu Tang, Ziyu Shao, Yang Yang 0001 |
ICC | 3 |
| 2020 | Joint Switch-Controller Association and Control Devolution for SDN Systems: An Integration of Online Control and Online LearningabstractIn software-defined networking (SDN) systems, it is a common practice to adopt a multi-controller design and control devolution techniques to improve the performance of the control plane. However, in such systems the decision making for joint switch-controller association and control devolution often involves various uncertainties, e.g., the temporal variations of controller accessibility, and computation and communication costs of switches. In practice, statistics of such uncertainties are unattainable and need to be learned in an online fashion, calling for an integrated design of learning and control. In this paper, we formulate a stochastic network optimization problem that aims to minimize time-average system costs and ensure queue stability. By transforming the problem into a combinatorial multi-armed bandit problem with long-term stability constraints, we adopt bandit learning methods and optimal control techniques to handle the exploration-exploitation tradeoff and long-term stability constraints, respectively. Through an integrated design of online learning and online control, we propose an effective Learning-Aided Switch-Controller Association and Control Devolution (LASAC) scheme. Our theoretical analysis and simulation results show that LASAC achieves a tunable tradeoff between queue stability and system cost reduction with a sublinear regret bound over a finite time horizon. Xi Huang 0001, Yinxu Tang, Ziyu Shao, Yang Yang 0001, Hong Xu 0001 |
IWQoS | 2 |
| 2019 | Learning-Aided Online Task Offloading for UAVs-Aided IoT SystemsabstractEquipped with specific IoT on-board devices, un- manned aerial vehicles (UAVs) can be orchestrated to assist in particular value-added service delivery with improved quality-of- service. Typically, services are delegated in the unit of tasks to a designated leader UAV, while the leader UAV splits each task into sub- tasks and offloads them to part of its nearby UAVs, a.k.a. helper UAVs, for timely processing. Such a decision making pro- cess, often referred to as UAV task offloading, still remains open and challenging to design, due to various uncertainties therein, such as the resource availability and instant workloads on helper UAVs. However, existing solutions often assume the knowledge of system dynamics is fully available and conduct decision making in an offline manner, resulting in excessive control overheads and scalability issues. In this paper, we study the UAV task offloading problem in an online setting and formulate it as a multi-armed bandits (MAB) problem with time-varying resource constraints. Then we propose VR-LATOS, a learning- aided offloading scheme that learns the unknown statistics from feedback signals while making effective offloading decisions in an online fashion. Results from both theoretical analysis and simulations demonstrate that VR-LATOS outperforms state-of-the-art schemes. Junge Zhu, Xi Huang 0001, Yinxu Tang, Ziyu Shao |
VTC Fall | 3 |