EDBT 2026 Demo / reviewers in the wild / expert
Na Li 0002
dblp:18/3173-2
· DBLP profile ↗
38ranked-venue papers
1as first author
25since 2021 · last 2026
0000-0001-9545-3050ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 1 first-author · 22 since 2021Systems, architecture and hardware · 10 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Computer networks · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LILAD: Learning In-context Lyapunov-stable Adaptive Dynamics ModelsabstractSystem identification in control theory aims to approximate dynamical systems from trajectory data. While neural networks have demonstrated strong predictive accuracy, they often fail to preserve critical physical properties such as stability and typically assume stationary dynamics, limiting their applicability under distribution shifts. Existing approaches generally address either stability or adaptability in isolation, lacking a unified framework that ensures both. We propose LILAD (Learning In-Context Lyapunov-stable Adaptive Dynamics), a novel framework for system identification that jointly guarantees adaptability and stability. LILAD simultaneously learns a dynamics model and a Lyapunov function through in-context learning (ICL), explicitly accounting for parametric uncertainty. Trained across a diverse set of tasks, LILAD produces a stability-aware, adaptive dynamics model alongside an adaptive Lyapunov certificate. At test time, both components adapt to a new system instance using a short trajectory prompt, which enables fast generalization. To rigorously ensure stability, LILAD also computes a state-dependent attenuator that enforces a sufficient decrease condition on the Lyapunov function for any state in the new system instance. This mechanism extends stability guarantees even under out-of-distribution and out-of-task scenarios. We evaluate LILAD on benchmark autonomous systems and demonstrate that it outperforms adaptive, robust, and non-adaptive baselines in predictive accuracy. Amit Jena, Na Li 0002, Le Xie 0001 |
AAAI | 2 |
| 2025 | Primal-Dual Spectral Representation for Off-policy EvaluationabstractOff-policy evaluation (OPE) is one of the most fundamental problems in reinforcement learning (RL) to estimate the expected long-term payoff of a given target policy with \emph{only} experiences from another behavior policy that is potentially unknown. The distribution correction estimation (DICE) family of estimators have advanced the state of the art in OPE by breaking the \emph{curse of horizon}. However, the major bottleneck of applying DICE estimators lies in the difficulty of solving the saddle-point optimization involved, especially with neural network implementations. In this paper, we tackle this challenge by establishing a \emph{linear representation} of value function and stationary distribution correction ratio, \emph{i.e.}, primal and dual variables in the DICE framework, using the spectral decomposition of the transition operator. Such primal-dual representation not only bypasses the non-convex non-concave optimization in vanilla DICE, therefore enabling an computational efficient algorithm, but also paves the way for more efficient utilization of historical data. We highlight that our algorithm, \textbf{SpectralDICE}, is the first to leverage the linear representation of primal-dual variables that is both computation and sample efficient, the performance of which is supported by a rigorous theoretical sample complexity guarantee and a thorough empirical evaluation on various benchmarks. Tianyi Chen 0002, Na Li 0002, Kai Wang 0040, Bo Dai 0001 |
AISTATS | 3 |
| 2025 | Scalable spectral representations for multiagent reinforcement learning in network MDPsabstractNetwork Markov Decision Processes (MDPs), which are the de-facto model for multi-agent control, pose a significant challenge to efficient learning caused by the exponential growth of the global state-action space with the number of agents. In this work, utilizing the exponential decay property of network dynamics, we first derive scalable spectral local representations for multiagent reinforcement learning in network MDPs, which induces a network linear subspace for the local $Q$-function of each agent. Building on these local spectral representations, we design a scalable algorithmic framework for multiagent reinforcement learning in continuous state-action network MDPs, and provide end-to-end guarantees for the convergence of our algorithm. Empirically, we validate the effectiveness of our scalable representation-based approach on two benchmark problems, and demonstrate the advantages of our approach over generic function approximation approaches to representing the local $Q$-functions. Zhaolin Ren, Runyu Zhang 0001, Bo Dai 0001, Na Li 0002 |
AISTATS | 4 |
| 2025 | Efficient Online Reinforcement Learning for Diffusion PolicyabstractDiffusion policies have achieved superior performance in imitation learning and offline reinforcement learning (RL) due to their rich expressiveness. However, the conventional diffusion training procedure requires samples from target distribution, which is impossible in online RL since we cannot sample from the optimal policy. Backpropagating policy gradient through the diffusion process incurs huge computational costs and instability, thus being expensive and not scalable. To enable efficient training of diffusion policies in online RL, we generalize the conventional denoising score matching by reweighting the loss function. The resulting Reweighted Score Matching (RSM) preserves the optimal solution and low computational cost of denoising score matching, while eliminating the need to sample from the target distribution and allowing learning to optimize value functions. We introduce two tractable reweighted loss functions to solve two commonly used policy optimization problems, policy mirror descent and max-entropy policy, resulting in two practical algorithms named Diffusion Policy Mirror Descent (DPMD) and Soft Diffusion Actor-Critic (SDAC). We conducted comprehensive comparisons on MuJoCo benchmarks. The empirical results show that the proposed algorithms outperform recent diffusion-policy online RLs on most tasks, and the DPMD improves more than 120% over soft actor-critic on Humanoid and Ant. Haitong Ma, Tianyi Chen 0002, Kai Wang 0040, Na Li 0002, Bo Dai 0001 |
ICML | 4 |
| 2025 | Offline Imitation Learning upon Arbitrary Demonstrations by Pre-Training Dynamics RepresentationsabstractLimited data has become a major bottleneck in scaling up offline imitation learning (IL). In this paper, we propose enhancing IL performance under limited expert data by introducing a pre-training stage that learns dynamics representations, derived from factorizations of the transition dynamics. We first theoretically justify that the optimal decision variable of offline IL lies in the representation space, significantly reducing the parameters to learn in the downstream IL. Moreover, the dynamics representations can be learned from arbitrary data collected with the same dynamics, allowing the reuse of massive non-expert data and mitigating the limited data issues. We present a tractable loss function inspired by noise contrastive estimation to learn the dynamics representations at the pre-training stage. Experiments on MuJoCo demonstrate that our proposed algorithm can mimic expert policies with as few as a single trajectory. Experiments on real quadrupeds show that we can leverage pre-trained dynamics representations from simulator data to learn to walk from a few real-world demonstrations. Haitong Ma, Bo Dai 0001, Zhaolin Ren, Yebin Wang, Na Li 0002 |
IROS | 5 |
| 2025 | RODS: Robust Optimization Inspired Diffusion Sampling for Detecting and Reducing Hallucination in Generative ModelsabstractDiffusion models have achieved state-of-the-art performance in generative modeling, yet their sampling procedures remain vulnerable to hallucinations—often stemming from inaccuracies in score approximation. In this work, we reinterpret diffusion sampling through the lens of optimization and introduce RODS (Robust Optimization–inspired Diffusion Sampler), a novel method that detects and corrects high-risk sampling steps using geometric cues from the loss landscape. RODS enforces smoother sampling trajectories and \textit{adaptively} adjusts perturbations, reducing hallucinations without retraining and at minimal additional inference cost. Experiments on AFHQv2, FFHQ, and 11k-hands demonstrate that RODS maintains comparable image quality and preserves generation diversity. More importantly, it improves both sampling fidelity and robustness, detecting over 70\% of hallucinated samples and correcting more than 25\%, all while avoiding the introduction of new artifacts. We release our code at https://github.com/Yiqi-Verna-Tian/RODS. Yiqi Tian, Pengfei Jin, Mingze Yuan, Na Li 0002, Quanzheng Li |
NeurIPS | 4 |
| 2025 | Constrained Optimization From a Control Perspective via Feedback LinearizationabstractTools from control and dynamical systems have proven valuable for analyzing and developing optimization methods. In this paper, we establish rigorous theoretical foundations for using feedback linearization—a well-established nonlinear control technique—to solve constrained optimization problems. For equality-constrained optimization, we establish global convergence rates to first-order Karush-Kuhn-Tucker (KKT) points and uncover the close connection between the FL method and the Sequential Quadratic Programming (SQP) algorithm. Building on this relationship, we extend the FL approach to handle inequality-constrained problems. Furthermore, we introduce a momentum-accelerated feedback linearization algorithm and provide a rigorous convergence guarantee. Runyu Zhang 0001, Arvind Raghunathan, Jeff S. Shamma, Na Li 0002 |
NeurIPS | 4 |
| 2025 | Enhancing gaze estimation accuracy in wearable eye-tracking devices using neural networks
Jerome Charton, Na Li 0002, Quanzheng Li |
Neural Comput. Appl. | 3 |
| 2024 | Soft Robust MDPs and Risk-Sensitive MDPs: Equivalence, Policy Gradient, and Sample ComplexityabstractRobust Markov Decision Processes (MDPs) and risk-sensitive MDPs are both powerful tools for making decisions in the presence of uncertainties. Previous efforts have aimed to establish their connections, revealing equivalences in specific formulations. This paper introduces a new formulation for risk-sensitive MDPs, which assesses risk in a slightly different manner compared to the classical Markov risk measure [Ruszczy ́nski 2010], and establishes its equivalence with a class of soft robust MDP (RMDP) problems, including the standard RMDP as a special case. Leveraging this equivalence, we further derive the policy gradient theorem for both problems, proving gradient domination and global convergence of the exact policy gradient method under the tabular setting with direct parameterization. This forms a sharp contrast to the Markov risk measure, known to be potentially non-gradient-dominant [Huang et al. 2021]. We also propose a sample-based offline learning algorithm, namely the robust fitted-Z iteration (RFZI), for a specific soft RMDP problem with a KL-divergence regularization term (or equivalently the risk-sensitive MDP with an entropy risk measure). We showcase its streamlined
design and less stringent assumptions due to the equivalence and analyze its sample complexity. Runyu Zhang 0001, Na Li 0002 |
ICLR | 3 |
| 2024 | Skill Transfer and Discovery for Sim-to-Real Learning: A Representation-Based ViewpointabstractWe study sim-to-real skill transfer and discovery in the context of robotics control using representation learning. We draw inspiration from spectral decomposition of Markov decision processes. The spectral decomposition brings about representation that can linearly represent the state-action value function induced by any policies, thus can be regarded as skills. The skill representations are transferable across arbitrary tasks with the same transition dynamics. Moreover, to handle the sim-to-real gap in the dynamics, we propose a skill discovery algorithm that learns new skills caused by the sim-to-real gap from real-world data. We promote the discovery of new skills by enforcing orthogonal constraints between the skills to learn and the skills from simulators, and then synthesize the policy using the enlarged skill sets. We demonstrate our methodology by transferring quadrotor controllers from simulators to Crazyflie 2.1 quadrotors. We show that we can learn the skill representations from a single simulator task and transfer these to multiple different real-world tasks including hovering, taking off, landing and trajectory tracking. Our skill discovery approach helps narrow the sim-to-real gap and improve the real-world controller performance by up to 30.2%. Haitong Ma, Zhaolin Ren, Bo Dai 0001, Na Li 0002 |
IROS | 4 |
| 2024 | Enhancing Preference-based Linear Bandits via Human Response TimeabstractInteractive preference learning systems infer human preferences by presenting queries as pairs of options and collecting binary choices. Although binary choices are simple and widely used, they provide limited information about preference strength. To address this, we leverage human response times, which are inversely related to preference strength, as an additional signal. We propose a computationally efficient method that combines choices and response times to estimate human utility functions, grounded in the EZ diffusion model from psychology. Theoretical and empirical analyses show that for queries with strong preferences, response times complement choices by providing extra information about preference strength, leading to significantly improved utility estimation. We incorporate this estimator into preference-based linear bandits for fixed-budget best-arm identification. Simulations on three real-world datasets demonstrate that using response times significantly accelerates preference learning compared to choice-only approaches. Additional materials, such as code, slides, and talk video, are available at https://shenlirobot.github.io/pages/NeurIPS24.html. Shen Li 0003, Zhaolin Ren, Claire Liang, Na Li 0002, Julie A. Shah |
NeurIPS | 5 |
| 2024 | Guest Editorial Special Issue on Reinforcement Learning-Based Control: Data-Efficient and Resilient MethodsabstractAs an important branch of machine learning, reinforcement learning (RL) has proved its efficiency in many emerging applications in science and engineering. A remarkable advantage of RL is that it enables agents to maximize their cumulative rewards through online exploration and interactions with unknown (or partially unknown) and uncertain environments, which is regarded as a variant of data-driven adaptive optimal control methods. However, the successful implementation of RL-based control systems usually relies on a good quantity of online data due to its data-driven nature. Therefore, it is imperative to develop data-efficient RL methods for control systems to reduce the required number of interactions with the external environment. Moreover, network-aware issues, such as cyberattacks, dropout packet and communication latency, and actuator and sensor faults, are challenging conundrums that threaten the safety, security, stability, and reliability of network control systems. Consequently, it is significant to develop safe and resilient RL mechanisms. Weinan Gao, Na Li 0002, Kyriakos G. Vamvoudakis, F. Richard Yu, Zhong-Ping Jiang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Learning to Optimize with Stochastic Dominance ConstraintsabstractIn real-world decision-making, uncertainty is important yet difficult to handle. Stochastic dominance provides a theoretically sound approach to comparing uncertain quantities, but optimization with stochastic dominance constraints is often computationally expensive, which limits practical applicability. In this paper, we develop a simple yet efficient approach for the problem, Light Stochastic Dominance Solver (light-SD), by leveraging properties of the Lagrangian. We recast the inner optimization in the Lagrangian as a learning problem for surrogate approximation, which bypasses the intractability and leads to tractable updates or even closed-form solutions for gradient calculations. We prove convergence of the algorithm and test it empirically. The proposed light-SD demonstrates superior performance on several representative problems ranging from finance to supply chain management. Hanjun Dai, Yuan Xue 0001, Niao He, Na Li 0002, Dale Schuurmans, Bo Dai 0001 |
AISTATS | 5 |
| 2023 | Latent Variable Representation for Reinforcement Learning
Tongzheng Ren, Chenjun Xiao, Tianjun Zhang, Na Li 0002, Zhaoran Wang 0001, Sujay Sanghavi, Dale Schuurmans, Bo Dai 0001 |
ICLR | 4 |
| 2023 | FedDAR: Federated Domain-Aware Representation Learning
Aoxiao Zhong, Zhaolin Ren, Na Li 0002, Quanzheng Li |
ICLR | 4 |
| 2023 | Escaping saddle points in zeroth-order optimization: the power of two-point estimatorsabstractTwo-point zeroth order methods are important in many applications of zeroth-order optimization arising in robotics, wind farms, power systems, online optimization, and adversarial robustness to black-box attacks in deep neural networks, where the problem can be high-dimensional and/or time-varying. Furthermore, such problems may be nonconvex and contain saddle points. While existing works have shown that zeroth-order methods utilizing $\Omega(d)$ function valuations per iteration (with $d$ denoting the problem dimension) can escape saddle points efficiently, it remains an open question if zeroth-order methods based on two-point estimators can escape saddle points. In this paper, we show that by adding an appropriate isotropic perturbation at each iteration, a zeroth-order algorithm based on $2m$ (for any $1 \leq m \leq d$) function evaluations per iteration can not only find $\epsilon$-second order stationary points polynomially fast, but do so using only $\tilde{O}(\frac{d}{m\epsilon^{2}\bar{\psi}})$ function evaluations, where $\bar{\psi} \geq \tilde{\Omega}(\sqrt{\epsilon})$ is a parameter capturing the extent to which the function of interest exhibits the strict saddle property. Zhaolin Ren, Yujie Tang 0002, Na Li 0002 |
ICML | 3 |
| 2023 | Gaussian Max-Value Entropy Search for Multi-Agent Bayesian OptimizationabstractWe study the multi-agent Bayesian optimization (BO) problem, where multiple agents maximize a black-box function via iterative queries. We focus on Entropy Search (ES), a sample-efficient BO algorithm that selects queries to maximize the mutual information about the maximum of the black-box function. One of the main challenges of ES is that calculating the mutual information requires computationallycostly approximation techniques. For multi-agent BO problems, the computational cost of ES is exponential in the number of agents. To address this challenge, we propose the Gaussian Max-value Entropy Search, a multi-agent BO algorithm with favorable sample and computational efficiency. The key to our idea is to use a normal distribution to approximate the function maximum and calculate its mutual information accordingly. The resulting approximation allows queries to be cast as the solution of a closed-form optimization problem which, in turn, can be solved via a modified gradient ascent algorithm and scaled to a large number of agents. We demonstrate the effectiveness of Gaussian max-value Entropy Search through numerical experiments on standard test functions and real-robot experiments on the source seeking problem. Results show that the proposed algorithm outperforms the multi-agent BO baselines in the numerical experiments and can stably seek the source with a limited number of noisy observations on real robots. Haitong Ma, Flávio P. Calmon, Na Li 0002 |
IROS | 5 |
| 2023 | Distributed Information-Based Source SeekingabstractIn this article, we design an information-based multirobot source seeking algorithm where a group of mobile sensors localizes and moves close to a single source using only local range-based measurements. In the algorithm, the mobile sensors perform source identification/localization to estimate the source location; meanwhile, they move to new locations to maximize the Fisher information about the source contained in the sensor measurements. In doing so, they improve the source location estimate and move closer to the source. Our algorithm is superior in convergence speed compared with traditional field climbing algorithms, is flexible in the measurement model and the choice of information metric, and is robust to measurement model errors. Moreover, we provide a fully distributed version of our algorithm, where each sensor decides its own actions and only shares information with its neighbors through a sparse communication network. We perform extensive simulation experiments to test our algorithms on large-scale systems and implement physical experiments on small ground vehicles with light sensors, demonstrating success in seeking a light source. Victor Qin, Yujie Tang 0002, Na Li 0002 |
IEEE Trans. Robotics | 4 |
| 2022 | Improve Single-Point Zeroth-Order Optimization Using High-Pass and Low-Pass FiltersabstractSingle-point zeroth-order optimization (SZO) is useful in solving online black-box optimization and control problems in time-varying environments, as it queries the function value only once at each time step. However, the vanilla SZO method is known to suffer from a large estimation variance and slow convergence, which seriously limits its practical application. In this work, we borrow the idea of high-pass and low-pass filters from extremum seeking control (continuous-time version of SZO) and develop a novel SZO method called HLF-SZO by integrating these filters. It turns out that the high-pass filter coincides with the residual feedback method, and the low-pass filter can be interpreted as the momentum method. As a result, the proposed HLF-SZO achieves a much smaller variance and much faster convergence than the vanilla SZO method, and empirically outperforms the residual-feedback SZO method, which are verified via extensive numerical experiments. Xin Chen 0036, Yujie Tang 0002, Na Li 0002 |
ICML | 3 |
| 2022 | Policy Optimization for Markov Games: Unified Framework and Faster ConvergenceabstractThis paper studies policy optimization algorithms for multi-agent reinforcement learning. We begin by proposing an algorithm framework for two-player zero-sum Markov Games in the full-information setting, where each iteration consists of a policy update step at each state using a certain matrix game algorithm, and a value update step with a certain learning rate. This framework unifies many existing and new policy optimization algorithms. We show that the \emph{state-wise average policy} of this algorithm converges to an approximate Nash equilibrium (NE) of the game, as long as the matrix game algorithms achieve low weighted regret at each state, with respect to weights determined by the speed of the value updates. Next, we show that this framework instantiated with the Optimistic Follow-The-Regularized-Leader (OFTRL) algorithm at each state (and smooth value updates) can find an $\mathcal{\widetilde{O}}(T^{-5/6})$ approximate NE in $T$ iterations, and a similar algorithm with slightly modified value update rule achieves a faster $\mathcal{\widetilde{O}}(T^{-1})$ convergence rate. These improve over the current best $\mathcal{\widetilde{O}}(T^{-1/2})$ rate of symmetric policy optimization type algorithms. We also extend this algorithm to multi-player general-sum Markov Games and show an $\mathcal{\widetilde{O}}(T^{-3/4})$ convergence rate to Coarse Correlated Equilibria (CCE). Finally, we provide a numerical example to verify our theory and investigate the importance of smooth value updates, and find that using ''eager'' value updates instead (equivalent to the independent natural policy gradient algorithm) may significantly slow down the convergence, even on a simple game with $H=2$ layers. Runyu Zhang 0001, Huan Wang 0016, Caiming Xiong, Na Li 0002, Yu Bai 0017 |
NeurIPS | 5 |
| 2022 | On the Global Convergence Rates of Decentralized Softmax Gradient Play in Markov Potential GamesabstractSoftmax policy gradient is a popular algorithm for policy optimization in single-agent reinforcement learning, particularly since projection is not needed for each gradient update. However, in multi-agent systems, the lack of central coordination introduces significant additional difficulties in the convergence analysis. Even for a stochastic game with identical interest, there can be multiple Nash Equilibria (NEs), which disables proof techniques that rely on the existence of a unique global optimum. Moreover, the softmax parameterization introduces non-NE policies with zero gradient, making it difficult for gradient-based algorithms in seeking NEs. In this paper, we study the finite time convergence of decentralized softmax gradient play in a special form of game, Markov Potential Games (MPGs), which includes the identical interest game as a special case. We investigate both gradient play and natural gradient play, with and without $\log$-barrier regularization. The established convergence rates for the unregularized cases contain a trajectory dependent constant that can be \emph{arbitrarily large}, whereas the $\log$-barrier regularization overcomes this drawback, with the cost of slightly worse dependence on other factors such as the action set size. An empirical study on an identical interest matrix game confirms the theoretical findings. Runyu Zhang 0001, Jincheng Mei, Bo Dai 0001, Dale Schuurmans, Na Li 0002 |
NeurIPS | 5 |
| 2021 | Online Optimal Control with Affine ConstraintsabstractThis paper considers online optimal control with affine constraints on the states and actions under linear dynamics with bounded random disturbances. The system dynamics and constraints are assumed to be known and time invariant but the convex stage cost functions change adversarially. To solve this problem, we propose Online Gradient Descent with Buffer Zones (OGD-BZ). Theoretically, we show that OGD-BZ with proper parameters can guarantee the system to satisfy all the constraints despite any admissible disturbances. Further, we investigate the policy regret of OGD-BZ, which compares OGD-BZ's performance with the performance of the optimal linear policy in hindsight. We show that OGD-BZ can achieve a policy regret upper bound that is square root of the horizon length multiplied by some logarithmic terms of the horizon length under proper algorithm parameters. Yingying Li 0005, Subhro Das, Na Li 0002 |
AAAI | 3 |
| 2021 | Assisted Learning: Cooperative AI with AutonomyabstractThe rapid development in data collecting devices and computation platforms produces an emerging number of agents, each equipped with a unique data modality over a particular population of subjects. While an agent’s predictive performance may be enhanced by transmitting others’ data to it, this is often unrealistic due to intractable transmission costs and security concerns. In this paper, we propose a method named ASCII for an agent to improve its classification performance through assistance from other agents, without sharing proprietary data and model information. The main idea is to iteratively interchange an ignorance value between 0 and 1 for each collated sample among agents, where the value represents the urgency of further assistance needed. The method is naturally suitable for privacy-aware, transmission-economical, and decentralized learning scenarios. The method is also general as it allows the agents to use arbitrary classifiers such as logistic regression, ensemble tree, and neural network, and they may be heterogeneous among agents. We demonstrate the proposed method with extensive experimental studies. Jiaying Zhou, Xun Xian, Na Li 0002, Jie Ding 0002 |
ICASSP | 3 |
| 2021 | Federated Learning over Wireless Networks: A Band-limited Coordinated Descent ApproachabstractWe consider a many-to-one wireless architecture for federated learning at the network edge, where multiple edge devices collaboratively train a model using local data. The unreliable nature of wireless connectivity, together with constraints in computing resources at edge devices, dictates that the local updates at edge devices should be carefully crafted and compressed to match the wireless communication resources available and should work in concert with the receiver. Thus motivated, we propose SGD-based bandlimited coordinate descent algorithms for such settings. Specifically, for the wireless edge employing over-the-air computing, a common subset of k-coordinates of the gradient updates across edge devices are selected by the receiver in each iteration, and then transmitted simultaneously over k sub-carriers, each experiencing time-varying channel conditions. We characterize the impact of communication error and compression, in terms of the resulting gradient bias and mean squared error, on the convergence of the proposed algorithms. We then study learning-driven communication error minimization via joint optimization of power allocation and learning rates. Our findings reveal that optimal power allocation across different sub-carriers should take into account both the gradient values and channel conditions, thus generalizing the widely used water-filling policy. We also develop sub-optimal distributed solutions amenable to implementation. Junshan Zhang, Na Li 0002, Mehmet Dedeoglu |
INFOCOM | 2 |
| 2021 | Source Seeking by Dynamic Source Location EstimationabstractThis paper focuses on the problem of multi-robot source-seeking, where a group of mobile sensors localizes and moves close to a single source using only local measurements. Drawing inspiration from the optimal sensor placement research, we develop an algorithm that estimates the source location while approaches the source following gradient descent steps on a loss function defined on the Fisher information. We show that exploiting Fisher information gives a higher chance of obtaining an accurate source location estimate and naturally leads the sensors to the source. Our numerical experiments demonstrate the advantages of our algorithm, including faster convergence to the source than other algorithms, flexibility in the choice of the loss function, and robustness to measurement modeling errors. Moreover, the performance improves as the number of sensors increases, showing the advantage of using multi-robots in our source-seeking algorithm. We also implement physical experiments to test the algorithm on small ground vehicles with light sensors, demonstrating success in seeking a moving light source. Victor Qin, Yujie Tang 0002, Na Li 0002 |
IROS | 4 |
| 2020 | Soft Sensing Shirt for Shoulder Kinematics EstimationabstractSoft strain sensors have been explored as an unobtrusive approach for wearable motion tracking. However, accurate tracking of multi degree-of-freedom (DOF) noncyclic joint movements remains a challenge. This paper presents a soft sensing shirt for tracking shoulder kinematics of both cyclic and random arm movements in 3 DOFs: adduction/abduction, horizontal flexion/extension, and internal/external rotation. The sensing shirt consists of 8 textile-based capacitive strain sensors sewn around the shoulder joint that communicate to a customized readout electronics board through sewn micro-coaxial cables. An optimized sensor design includes passive shielding and demonstrates high linearity and low hysteresis, making it suitable for wearable motion tracking. In a study with a single human subject, we evaluated the tracking capability of the integrated shirt in comparison with a ground truth optical motion capture system. An ensemble-based regression algorithm was implemented in post-processing to estimate joint angles and angular velocities from the strain sensor data. Results demonstrated root mean square errors (RMSEs) less than 4.5° for joint angle estimation and normalized root mean square errors (NRMSEs) less than 4% for joint velocity estimation. Furthermore, we applied a recursive feature elimination (RFE)-based sensor selection analysis to down select the number of sensors for future shirt designs. This sensor selection analysis found that 5 sensors out of 8 were sufficient to generate comparable accuracies. Yichu Jin, Christina M. Glover, Haedo Cho, Oluwaseun A. Araromi, Moritz A. Graule, Na Li 0002, Robert J. Wood, Conor J. Walsh |
ICRA | 6 |
| 2020 | Leveraging Predictions in Smoothed Online Convex Optimization via Gradient-based AlgorithmsabstractWe consider online convex optimization with time-varying stage costs and additional switching costs. Since the switching costs introduce coupling across all stages, multi-step-ahead (long-term) predictions are incorporated to improve the online performance. However, longer-term predictions tend to suffer from lower quality. Thus, a critical question is: how to reduce the impact of long-term prediction errors on the online performance? To address this question, we introduce a gradient-based online algorithm, Receding Horizon Inexact Gradient (RHIG), and analyze its performance by dynamic regrets in terms of the temporal variation of the environment and the prediction errors. RHIG only considers at most $W$-step-ahead predictions to avoid being misled by worse predictions in the longer term. The optimal choice of $W$ suggested by our regret bounds depends on the tradeoff between the variation of the environment and the prediction accuracy. Additionally, we apply RHIG to a well-established stochastic prediction error model and provide expected regret and concentration bounds under correlated prediction errors. Lastly, we numerically test the performance of RHIG on quadrotor tracking problems. Yingying Li 0005, Na Li 0002 |
NeurIPS | 2 |
| 2020 | Scalable Multi-Agent Reinforcement Learning for Networked Systems with Average RewardabstractIt has long been recognized that multi-agent reinforcement learning (MARL) faces significant scalability issues due to the fact that the size of the state and action spaces are exponentially large in the number of agents. In this paper, we identify a rich class of networked MARL problems where the model exhibits a local dependence structure that allows it to be solved in a scalable manner. Specifically, we propose a Scalable Actor-Critic (SAC) method that can learn a near optimal localized policy for optimizing the average reward with complexity scaling with the state-action space size of local neighborhoods, as opposed to the entire network. Our result centers around identifying and exploiting an exponential decay property that ensures the effect of agents on each other decays exponentially fast in their graph distance. Guannan Qu, Yiheng Lin 0001, Adam Wierman, Na Li 0002 |
NeurIPS | 4 |
| 2019 | Online Optimal Control with Linear Dynamics and Predictions: Algorithms and Regret AnalysisabstractThis paper studies the online optimal control problem with time-varying convex stage costs for a time-invariant linear dynamical system, where a finite lookahead window of accurate predictions of the stage costs are available at each time. We design online algorithms, Receding Horizon Gradient-based Control (RHGC), that utilize the predictions through finite steps of gradient computations. We study the algorithm performance measured by dynamic regret: the online performance minus the optimal performance in hindsight. It is shown that the dynamic regret of RHGC decays exponentially with the size of the lookahead window. In addition, we provide a fundamental limit of the dynamic regret for any online algorithms by considering linear quadratic tracking problems. The regret upper bound of one RHGC method almost reaches the fundamental limit, demonstrating the effectiveness of the algorithm. Finally, we numerically test our algorithms for both linear and nonlinear systems to show the effectiveness and generality of our RHGC. Yingying Li 0005, Xin Chen 0036, Na Li 0002 |
NeurIPS | 3 |
| 2017 | Fast Decentralized Power Capping for Server ClustersabstractPower capping is a mechanism to ensure that the power consumption of clusters does not exceed the provisioned resources. A fast power capping method allows for a safe over-subscription of the rated power distribution devices, provides equipment protection, and enables large clusters to participate in demand-response programs. However, current methods have a slow response time with a large actuation latency when applied across a large number of servers as they rely on hierarchical management systems. We propose a fast decentralized power capping (DPC) technique that reduces the actuation latency by localizing power management at each server. The DPC method is based on a maximum throughput optimization formulation that takes into account the workloads priorities as well as the capacity of circuit breakers. Therefore, DPC significantly improves the cluster performance compared to alternative heuristics. We implement the proposed decentralized power management scheme on a real computing cluster. Compared to state-of-the-art hierarchical methods, DPC reduces the actuation latency by 72% up to 86% depending on the cluster size. In addition, DPC improves the system throughput performance by 16%, while using only 0.02% of the available network bandwidth. We describe how to minimize the overhead of each local DPC agent to a negligible amount. We also quantify the traffic and fault resilience of our decentralized power capping approach. Masoud Badiei, Xin Zhan, Na Li 0002, Sherief Reda |
HPCA | 4 |
| 2017 | Stochastic Primal-Dual Method on Riemannian Manifolds of Bounded Sectional CurvatureabstractWe study a stochastic primal-dual method for optimizing functions on elliptic (sub)manifolds- Riemannian (sub)manifolds with a positive bounded sectional curvature. In particular, we establish a convergence rate for geodesically convex functions that is related to the lower bound on the sectional curvature. The convergence analysis we present is based on Toponogov's comparison theorem, where geodesic triangles on the elliptic manifolds and a sphere are compared. We numerically demonstrate the performance of the proposed stochastic primal-dual algorithm on the sphere for non-negative principle component analysis (PCA), and on the Lie group SO(3) for the anchored localization from partial noisy measurements of relative rotations. In both applications, the proposed algorithm scales gracefully to high dimensions. Masoud Badiei, Na Li 0002 |
ICMLA | 2 |
| 2016 | DiBA: Distributed Power Budget Allocation for Large-Scale Computing ClustersabstractPower management has become a central issue inlarge-scale computing clusters where a considerable amount ofenergy is consumed and a large operational cost is incurredannually. Traditional power management techniques have a centralizeddesign that creates challenges for scalability of computingclusters. In this work, we develop a framework for distributedpower budget allocation that maximizes the utility of computingnodes subject to a total power budget constraint. To eliminate the role of central coordinator in the primaldualtechnique, we propose a distributed power budget allocationalgorithm (DiBA) which maximizes the combined performanceof a cluster subject to a power budget constraint in a distributedfashion. Specifically, DiBA is a consensus-based algorithm inwhich each server determines its optimal power consumptionlocally by communicating its state with neighbors (connectednodes) in a cluster. We characterize a synchronous primal-dualtechnique to obtain a benchmark for comparison with thedistributed algorithm that we propose. We demonstrate numericallythat DiBA is a scalable algorithm that outperforms theconventional primal-dual method on large scale clusters in termsof convergence time. Further, DiBA eliminates the communicationbottleneck in the primal-dual method. We thoroughly evaluatethe characteristics of DiBA through simulations of large-scaleclusters. Furthermore, we provide results from a proof-of-conceptimplementation on a real experimental cluster. Masoud Badiei, Xin Zhan, Sherief Reda, Na Li 0002 |
CCGrid | 5 |
| 2016 | Asynchronous local voltage control in power distribution networksabstractHigh penetration of distributed energy resources presents significant challenges and provides emerging opportunities for voltage regulation in power distribution systems. Advanced power-electronics technology makes it possible to control the reactive power output from these resources, in order to maintain a desirable voltage profile. This paper develops a local control framework to account for limits on reactive power resources using the gradient projection optimization method. Requiring only local voltage measurements, the proposed design does not suffer from the stability issues of (de-)centralized approaches caused by communication delays and noises. Our local voltage design is shown to be robust to potential asynchronous control updates among distributed resources in a "plug-and-play" distribution network. Na Li 0002 |
ICASSP | 2 |
| 2015 | TELLab: An Experiential Learning Tool for PsychologyabstractIn this paper, we discuss current practices and challenges of teaching psychology experiments. We review experiential learning and analogical learning pedagogies, which have informed the design of TELLab, an online platform for supporting effective experiential learning of psychology concepts. Na Li 0002, Krzysztof Z. Gajos, Ken Nakayama, Ryan Enos |
L@S | 1 |
| 2015 | Using and Designing Platforms for In Vivo Educational ExperimentsabstractIn contrast to typical laboratory experiments, the everyday use of online educational resources by large populations and the prevalence of software infrastructure for A/B testing leads us to consider how platforms can embed in vivo experiments that do not merely support research, but ensure practical improvements to their educational components. Examples are presented of randomized experimental comparisons conducted by subsets of the authors in three widely used online educational platforms -- Khan Academy, edX, and ASSISTments. We suggest design principles for platform technology to support randomized experiments that lead to practical improvements -- enabling Iterative Improvement and Collaborative Work -- and explain the benefit of their implementation by WPI co-authors in the ASSISTments platform. Joseph Jay Williams, Korinn S. Ostrow, Xiaolu Xiong, Elena L. Glassman, Juho Kim 0001, Samuel G. Maldonado, Na Li 0002, Justin Reich, Neil T. Heffernan |
L@S | 7 |
| 2015 | On the Interaction Between Load Balancing and Speed ScalingabstractSpeed scaling has been widely adopted in computer and communication systems, in particular, to reduce energy consumption. An important question is how speed scaling interacts with other resource allocation mechanisms such as scheduling and routing. In this paper, we study the interaction of speed scaling with load balancing. We characterize the equilibrium resulting from the load balancing and speed scaling interaction, and introduce two optimal load-balancing designs, in terms of traditional performance metric and cost-aware (in particular, energy-aware) performance metric, respectively. Especially, we characterize the load-balancing–speed-scaling equilibrium with respect to the optimal load-balancing schemes in processor-sharing systems. Our results show that the degree of inefficiency at the equilibrium is mostly bounded by the heterogeneity of the system, but independent of the number of servers. These results provide insights in understanding the interaction of load balancing with speed scaling and guiding new designs. Lijun Chen 0001, Na Li 0002 |
IEEE J. Sel. Areas Commun. | 2 |
| 2014 | Optimal Residential Demand Response in Distribution NetworksabstractDemand response (DR) enables customers to adjust their electricity usage to balance supply and demand. Most previous works on DR consider the supply-demand matching in an abstract way without taking into account the underlying power distribution network and the associated power flow and system operational constraints. As a result, the schemes proposed by those works may end up with electricity consumption/shedding decisions that violate those constraints and thus are not feasible. In this paper, we study residential DR with consideration of the power distribution network and the associated constraints. We formulate residential DR as an optimal power flow problem and propose a distributed scheme where the load service entity and the households interactively communicate to compute an optimal demand schedule. To complement our theoretical results, we also simulate an IEEE test distribution system. The simulation results demonstrate two interesting effects of DR. One is the location effect, meaning that the households far away from the feeder tend to reduce more demands in DR. The other is the rebound effect, meaning that DR may create a new peak after the DR event ends if the DR parameters are not chosen carefully. The two effects suggest certain rules we should follow when designing a DR program. Na Li 0002, Xiaorong Xie, Chi-Cheng Peter Chu, Rajit Gadh |
IEEE J. Sel. Areas Commun. | 2 |
| 2013 | Exact convex relaxation for optimal power flow in distribution networksabstractThe optimal power flow (OPF) problem seeks to control the power generation/consumption to minimize the generation cost, and is becoming important for distribution networks. OPF is nonconvex and a second-order cone programming (SOCP) relaxation has been proposed to solve it. We prove that after a "small" modification to OPF, the SOCP relaxation is exact under a "mild" condition. Empirical studies demonstrate that the modification to OPF is "small" and that the "mild" condition holds for all test networks, including the IEEE 13-bus test network and practical networks with high penetration of distributed generation. Lingwen Gan, Na Li 0002, Steven H. Low, Ufuk Topcu |
SIGMETRICS | 2 |