VLDB 2026 Research / reviewers in the wild / expert
Ujwal Dinesha
dblp:290/7228
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-5708-0553ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DOPL: Direct Online Preference Learning for Restless Bandits with Preference FeedbackabstractRestless multi-armed bandits (RMAB) has been widely used to model constrained sequential decision making problems, where the state of each restless arm evolves according to a Markov chain and each state transition generates a scalar reward. However, the success of RMAB crucially relies on the availability and quality of reward signals. Unfortunately, specifying an exact reward function in practice can be challenging and even infeasible. In this paper, we introduce Pref-RMAB, a new RMAB model in the presence of preference signals, where the decision maker only observes pairwise preference feedback rather than scalar reward from the activated arms at each decision epoch. Preference feedback, however, arguably contains less information than the scalar reward, which makes Pref-RMAB seemingly more difficult. To address this challenge, we present a direct online preference learning (DOPL) algorithm for Pref-RMAB to efficiently explore the unknown environments, adaptively collect preference data in an online manner, and directly leverage the preference feedback for decision-makings. We prove that DOPL yields a sublinear regret. To our best knowledge, this is the first algorithm to ensure $\tilde{\mathcal{O}}(\sqrt{T\ln T})$ regret for RMAB with preference feedback. Experimental results further demonstrate the effectiveness of DOPL. Guojun Xiong, Ujwal Dinesha, Debajoy Mukherjee, Jian Li 0008, Srinivas Shakkottai |
ICLR | 2 |
| 2024 | A Multi-Agent View of Wireless Video Streaming with Delayed Client-FeedbackabstractWe study the optimal control of multiple video streams over a wireless downlink from a base-transceiver-station (BTS)/access point to N end-devices (EDs). The BTS sends video packets to each ED under a joint transmission energy constraint, the EDs choose when to play out the received packets, and the collective goal is to provide a high Quality-of-Experience (QoE) to the clients/end-users. All EDs send feedback about their states and actions to the BTS which reaches it after a fixed deterministic delay. We analyze this team problem with delayed feedback as a cooperative Multi-Agent Constrained Partially Observable Markov Decision Process (MA-C-POMDP).First, using a recently established strong duality result for MAC-POMDPs, the original problem is decomposed into N independent unconstrained transmitter-receiver (two-agent) problems— all sharing a Lagrange multiplier (that also needs to be optimized for optimal control). Thereafter, the common information (CI) approach and the formalism of approximate information states (AISs) are used to guide the design of a neural-network based architecture for learning-based multi-agent control in a single unconstrained transmitter-receiver problem. Finally, simulations on a single transmitter-receiver pair with a stylized QoE model are performed to highlight the advantage of delay-aware two-agent coordination over the transmitter choosing both transmission and play-out actions (perceiving the delayed state of the receiver as its current state). Nouman Khan, Ujwal Dinesha, Subrahmanyam Arunachalam, Dheeraj Narasimha, Vijay G. Subramanian, Srinivas Shakkottai |
INFOCOM | 2 |
| 2024 | Risk-Averse Fine-tuning of Large Language ModelsabstractWe consider the challenge of mitigating the generation of negative or toxic content by the Large Language Models (LLMs) in response to certain prompts. We propose integrating risk-averse principles into LLM fine-tuning to minimize the occurrence of harmful outputs, particularly rare but significant events. By optimizing the risk measure of Conditional Value at Risk (CVaR), our methodology trains LLMs to exhibit superior performance in avoiding toxic outputs while maintaining effectiveness in generative tasks. Empirical evaluations on sentiment modification and toxicity mitigation tasks demonstrate the efficacy of risk-averse reinforcement learning with human feedback (RLHF) in promoting a safer and more constructive online discourse environment. Sapana Chaudhary, Ujwal Dinesha, Dileep M. Kalathil, Srinivas Shakkottai |
NeurIPS | 2 |
| 2024 | EdgeRIC: Empowering Real-time Intelligent Optimization and Control in NextG Cellular Networks
Woo-Hyun Ko, Ushasi Ghosh, Ujwal Dinesha, Raini Wu, Srinivas Shakkottai, Dinesh Bharadia |
NSDI | 3 |
| 2023 | Demo: EdgeRIC: Delivering Realtime RAN IntelligenceabstractNextG cellular networks must support diverse applications, such as interactive media streaming or robot control that have strict requirements on throughput, latency and reliability. These requirements must be met via optimizing wireless resources by utilizing application layer information, such as media streaming stall counts or robot pose estimates, along with network information, such as channel qualities and backlogs. Woo-Hyun Ko, Ushasi Ghosh, Ujwal Dinesha, Raini Wu, Srinivas Shakkottai, Dinesh Bharadia |
SIGCOMM | 3 |
| 2022 | Realtime intelligent control for NextG cellular radio access networksabstractRAN Intelligent Control (RIC) has developed in parallel with Open Radio Access Networks (O-RAN) as a means of utilizing newly available interfaces. Focus has been largely on non-realtime (non-RT: > 1 sec) dealing with RAN management and offline training, and near-realtime (near-RT: 10 ms to 1 sec) dealing with UE load balancing and RAN configuration. We contend that the true power of RIC can be unleashed only with realtime (RT: < 100 μs) measurement, optimization, and control of RAN resources, corresponding to the cellular transmission time interval (TTI: 125 μs to 1 ms). Harish Kumar Dureppagari, Ujwal Dinesha, Raini Wu, Venkata Siva Santosh Ganji, Woo-Hyun Ko, Srinivas Shakkottai, Dinesh Bharadia |
MobiSys | 2 |