VLDB 2026 Research / reviewers in the wild / expert
Juaren Steiger
dblp:284/1392
· DBLP profile ↗
9ranked-venue papers
7as first author
8since 2021 · last 2026
0009-0000-7067-5087ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 6 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the Robustness of Age for Learning-Based Wireless Scheduling in Unknown Environments
Juaren Steiger, Bin Li 0014 |
INFOCOM | 1 |
| 2026 | Low-Overhead Scheduling for Synchronization in Large-Scale Heterogeneous Digital Twin Systems
Zifan Zhou, Juaren Steiger, Yin Sun 0001, Bin Li 0014 |
INFOCOM | 2 |
| 2026 | Learning to Balance Utility and Delay in Bipartite Queueing Networks With Sample Path ConstraintsabstractBipartite queueing networks with unknown statistics, where jobs are routed to and queued at servers and yield job-type and server-dependent utilities upon completion, model a wide range of problems in communications and related research areas (e.g., call routing in call centers, task assignment in crowdsourcing, job dispatching to cloud servers). Additionally, many such problems have additional routing constraints, such as quality of service or budgeted server cost constraints. The utility maximization problem in a bipartite queueing network with unknown statistics and subject to additional routing constraints is a constrained bandit learning problem with delayed feedback that depends on the server queueing delay. In this paper, we propose an efficient algorithm that overcomes the technical shortcomings of the state-of-the-art and effectively balances utility regret and peak job completion delay, while achieving constant peak constraint violation for the additional sample path routing constraints. Empirically, our algorithm is shown to simultaneously achieve low regret, peak delay, and peak constraint violation compared to existing algorithms. Juaren Steiger, Bin Li 0014, Ning Lu 0001 |
IEEE Trans. Netw. | 1 |
| 2025 | Optimal Hybrid Feedback-Driven Learning for Wireless Interactive Panoramic Scene DeliveryabstractImmersive technologies, such as virtual and augmented reality, demand high framerate, low latency, and precise synchronization between real and virtual environments. To meet these requirements, an edge server typically needs to perform high-quality rendering, and must predict user head motion and transmit a portion of the rendered panoramic scene that is large enough to cover the user's viewport, yet small enough to satisfy bandwidth constraints. Each portion yields two feedback signals: prediction feedback, indicating whether the selected portion covers the actual viewport, and transmission feedback, indicating whether all data packets are successfully delivered. While prior work models this setting as a multi-armed bandit with two-level bandit feedback, it overlooks that prediction feedback can be retrospectively computed for all possible portions, thus providing full-information feedback. In this work, we introduce a new two-level feedback model that combines full-information feedback with bandit feedback, and we formulate the portion selection problem as an online learning task under this hybrid setting. We derive an instance-dependent regret lower bound for this new hybrid feedback setting, and we propose AdaPort, a hybrid learning algorithm that leverages both the full-information feedback and bandit feedback to improve learning efficiency. We then show that the instance-dependent regret upper bound for AdaPort matches the lower bound asymptotically, proving its asymptotic optimality. Simulations using synthetic data and real-world traces demonstrate that AdaPort consistently outperforms state-of-the-art baselines, validating the benefits of exploiting the hybrid feedback structure. Xiaoyi Wu, Juaren Steiger, Bin Li 0014, R. Srikant 0001 |
MobiHoc | 2 |
| 2025 | Learning to Wirelessly Deliver Consistent and High-Quality Interactive Panoramic ScenesabstractWireless interactive panoramic scene delivery imposes unique challenges compared to its wired or non-interactive counterparts. The wireless channel is throughput-constrained, which limits the ability to deliver large and high-quality panoramic images. On the other hand, the interactivity imposes a real-time constraint on the system and limits the use of a playback buffer. Also, wireless inputs are not delivered instantaneously like wired interrupt-based inputs, so the system must predict the user's head pose and the portion of the scene visible to them, called the viewport. This reveals a tradeoff: delivering a portion too small may not cover the viewport if the prediction error is too large, while delivering a portion too large may result in a failed wireless transmission. Likewise, delivering the portion at too high a quality may result in a failed transmission. Despite these challenges, we would like to guarantee an immersive experience for the user by delivering a high-quality and visually consistent panoramic scene. To that end, we aim to maximize the user's quality of experience, which we define as the combination of (1) cumulative quality, (2) long-term consistency, which quantifies the overall variance in perceived quality, and (3) short-term consistency, which quantifies abrupt quality changes. We formulate this problem as a risk-averse multi-armed bandit problem with reward-dependent switching costs, and develop a novel block-based UCB algorithm with opportunistic switching. We derive its theoretical regret upper bound, which matches results in prior work, and corroborate this result in trace-based simulations using a panoramic video streaming data trace. Juaren Steiger, Xiaoyi Wu, Bin Li 0014 |
WiOpt | 1 |
| 2024 | Backlogged Bandits: Cost-Effective Learning for Utility Maximization in Queueing NetworksabstractBipartite queueing networks with unknown statistics, where jobs are routed to and queued at servers and yield server-dependent utilities upon completion, model a wide range of problems in communications and related research areas (e.g., call routing in call centers, task assignment in crowdsourcing, job dispatching to cloud servers). The utility maximization problem in bipartite queueing networks with unknown statistics is a bandit learning problem where the delayed semi-bandit feedback depends on the server queueing delay. In this paper, we propose an efficient algorithm that overcomes the technical shortcomings of the state-of-the-art and achieves square root regret, queue length, and feedback delay. Our approach also accommodates additional constraints, such as quality of service, fairness, and budgeted cost constraints, with constant expected peak violation and zero expected violation after a fixed timeslot. Empirically, our algorithm’s regret is competitive with the state-of-the-art for some problem instances and outperforms it in others, with much lower delay and constraint violation. Juaren Steiger, Bin Li 0014, Ning Lu 0001 |
INFOCOM | 1 |
| 2023 | Constrained Bandit Learning with Switching Costs for Wireless NetworksabstractBandits with arm selection constraints and bandits with switching costs have both gained recent attention in wireless networking research. Pessimistic-optimistic algorithms, which combine bandit learning with virtual queues to track the constraints, are commonly employed in the former. Block-based algorithms, where switching is disallowed within a block, are commonly employed in the latter. While efficient algorithms have been developed for both problems, it remains challenging to guarantee low regret and constraint violation in a bandit problem that includes both arm selection constraints and switching costs due to the tight coupling between the two. Here, switching may be necessary to decrease the constraint violation but comes at the cost of increased switching regret. In this paper, we tackle the constrained bandits with switching costs problem, for which we design a block-based pessimistic-optimistic algorithm. We identify three timely wireless networking applications for this framework in edge computing, mobile crowdsensing, and wireless network selection. We also prove that our algorithm achieves sublinear regret and vanishing constraint violation and corroborate these results with synthetic simulations and extensive trace-based simulations in the wireless network selection setting. Juaren Steiger, Bin Li 0014, Bo Ji 0001, Ning Lu 0001 |
INFOCOM | 1 |
| 2022 | Learning from Delayed Semi-Bandit Feedback under Strong Fairness GuaranteesabstractMulti-armed bandit frameworks, including combinatorial semi-bandits and sleeping bandits, are commonly employed to model problems in communication networks and other engineering domains. In such problems, feedback to the learning agent is often delayed (e.g. communication delays in a wireless network or conversion delays in online advertising). Moreover, arms in a bandit problem often represent entities required to be treated fairly, i.e. the arms should be played at least a required fraction of the time. In contrast to the previously studied asymptotic fairness, many real-time systems require such fairness guarantees to hold even in the short-term (e.g. ensuring the credibility of information flows in an industrial Internet of Things (IoT) system). To that end, we develop the Learning with Delays under Fairness (LDF) algorithm to solve combinatorial semi-bandit problems with sleeping arms and delayed feedback, which we prove guarantees strong (short-term) fairness. While previous theoretical work on bandit problems with delayed feedback typically derive instance-dependent regret bounds, this approach proves to be challenging when simultaneously considering fairness. We instead derive a novel instance-independent regret bound in this setting which agrees with state-of-the-art bounds. We verify our theoretical results with extensive simulations using both synthetic and real-world datasets. Juaren Steiger, Bin Li 0014, Ning Lu 0001 |
INFOCOM | 1 |
| 2020 | Learning for Path Planning and Coverage Mapping in UAV-Assisted Emergency CommunicationsabstractWe consider a setting in which a rotary-wing unmanned aerial vehicle (UAV) acts as an aerial base station to provide emergency communication service to an area of unknown and inhomogeneous user distribution. The UAV has communication with a ground node deployed to the area, which acts as a charging station. We are interested in two important problems in this setting, namely the path planning and coverage mapping problems. In the path planning problem, the UAV must plan its path starting and ending at the charging station, visiting a series of waypoints over which it hovers to provide coverage to surrounding users. On the other hand, the coverage mapping problem focuses on learning the distribution of user coverage over the area. We highlight the importance of learning this distribution to collect valuable data in an emergency situation. We then propose an online algorithm that simultaneously solves the path planning and coverage mapping problems using a deep learning model. We highlight the interplay and conflicting goals of path planning and coverage mapping, but show through Monte Carlo simulation that, under the correct parameters, the algorithm is able to achieve success on both problems. Juaren Steiger, Ning Lu 0001, Sameh Sorour |
GLOBECOM | 1 |