Lisha Chen

dblp:123/6690 · DBLP profile ↗
← Back
23ranked-venue papers
12as first author
13since 2021 · last 2025
0000-0001-8980-5275ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 11 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 4 since 2021Systems, architecture and hardware · 5 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Efficient First-Order Optimization on the Pareto Set for Multi-Objective Learning under Preference Guidance
abstract
Multi-objective learning under user-specified preference is common in real-world problems such as multi-lingual speech recognition under fairness. In this work, we frame such a problem as a semivectorial bilevel optimization problem, whose goal is to optimize a pre-defined preference function, subject to the constraint that the model parameters are weakly Pareto optimal. To solve this problem, we convert the multi-objective constraints to a single-objective constraint through a merit function with an easy-to-evaluate gradient, and then, we use a penalty-based reformulation of the bilevel optimization problem. We theoretically establish the properties of the merit function, and the relations of solutions for the penalty reformulation and the constrained formulation. Then we propose algorithms to solve the reformulated single-level problem, and establish its convergence guarantees. We test the method on various synthetic and real-world problems. The results demonstrate the effectiveness of the proposed method in finding preference-guided optimal solutions to the multi-objective problem.
Lisha Chen, Quan Xiao, Ellen H. Fukuda
ICML1
2025 Beyond Value Functions: Single-Loop Bilevel Optimization under Flatness Conditions
abstract
Bilevel optimization, a hierarchical optimization paradigm, has gained significant attention in a wide range of practical applications, notably in the fine-tuning of generative models. However, due to the nested problem structure, most existing algorithms require either the Hessian vector calculation or the nested loop updates, which are computationally inefficient in large language model (LLM) fine-tuning. In this paper, building upon the fully first-order penalty-based approach, we propose an efficient value function-free (\textsf{PBGD-Free}) algorithm that eliminates the loop of solving the lower-level problem and admits fully single-loop updates. Inspired by the landscape analysis of representation learning-based LLM fine-tuning problem, we propose a relaxed flatness condition for the upper-level function and prove the convergence of the proposed value-function-free algorithm. We test the performance of the proposed algorithm in various applications and demonstrate its superior computational efficiency over the state-of-the-art bilevel methods.
Liuyuan Jiang, Quan Xiao, Lisha Chen
NeurIPS3
2025 Objective Soups: Multilingual Multi-Task Modeling for Speech Processing
abstract
The need for training multilingual multi-task speech processing (MSP) models that perform both automatic speech recognition and speech-to-text translation is increasingly evident. However, a significant challenge arises from the conflicts among multiple objectives when using a single model. Multi-objective optimization can address this challenge by facilitating the optimization of multiple conflicting objectives and aligning the gradient updates in a common descent direction. While multi-objective optimization helps avoid conflicting gradient updates, a critical issue is that when there are many objectives, such as in MSP, it is often {\em difficult to find} a common descent direction. This leads to an important question: Is it more effective to separate highly conflicting objectives into different optimization levels or to keep them in a single level? To address this question, this paper investigates three multi-objective MSP formulations, which we refer to as \textbf{objective soup recipes}. These formulations apply multi-objective optimization at different optimization levels to mitigate potential conflicts among all objectives. To keep computation and memory overhead low, we incorporate a lightweight layer‑selection strategy that detects the most conflicting layers and uses only their gradients when computing the conflict‑avoidance direction. We conduct an extensive investigation using the CoVoST v2 dataset for combined multilingual ASR and ST tasks, along with the LibriSpeech and AISHELL-1 datasets for multilingual ASR, to identify highly conflicting objectives and determine the most effective training recipe among the three proposed multi-objective optimization algorithms.
A F M Saif, Lisha Chen, Songtao Lu, Brian Kingsbury, Tianyi Chen 0002
NeurIPS2
2024 Variance Reduction Can Improve Trade-Off in Multi-Objective Learning
abstract
Many machine learning problems today have multiple objective functions, which are often tackled by the multi-objective learning (MOL) framework. Albeit many encouraging results are obtained by MOL algorithms, a recent theoretical study [1] revealed that these gradient-based MOL methods (e.g., MGDA, CAGrad) all reflect an inherent trade-off between optimization convergence speeds and conflict-avoidance abilities. To this end, we develop an improved stochastic variance-reduced multi-objective gradient correction method for MOL, achieving the ${\mathcal{O}}\left({{\varepsilon ^{ - 1.5}}}\right)$ sample complexity. In addition, our proposed method simultaneously improves the theoretical guarantees for conflict avoidance and convergence rate compared to prior stochastic gradient-based MOL methods in the non-convex setting. We further validate the effectiveness of the proposed method empirically using popular multi-task learning (MTL) benchmarks.
Heshan Devaka Fernando, Lisha Chen, Songtao Lu, Miao Liu 0001, Subhajit Chaudhury, Keerthiram Murugesan, Gaowen Liu, Meng Wang 0003, Tianyi Chen 0002
ICASSP2
2024 M2ASR: Multilingual Multi-task Automatic Speech Recognition via Multi-objective Optimization
A F M Saif, Lisha Chen, Songtao Lu, Brian Kingsbury, Tianyi Chen 0002
INTERSPEECH2
2024 FERERO: A Flexible Framework for Preference-Guided Multi-Objective Learning
abstract
Finding specific preference-guided Pareto solutions that represent different trade-offs among multiple objectives is critical yet challenging in multi-objective problems. Existing methods are restrictive in preference definitions and/or their theoretical guarantees. In this work, we introduce a Flexible framEwork for pREfeRence-guided multi-Objective learning (**FERERO**) by casting it as a constrained vector optimization problem. Specifically, two types of preferences are incorporated into this formulation -- the *relative preference* defined by the partial ordering induced by a polyhedral cone, and the *absolute preference* defined by constraints that are linear functions of the objectives. To solve this problem, convergent algorithms are developed with both single-loop and stochastic variants. Notably, this is the *first single-loop primal algorithm* for constrained optimization to our knowledge. The proposed algorithms adaptively adjust to both constraint and objective values, eliminating the need to solve different subproblems at different stages of constraint satisfaction. Experiments on multiple benchmarks demonstrate the proposed method is very competitive in finding preference-guided optimal solutions. Code is available at https://github.com/lisha-chen/FERERO/.
Lisha Chen, A F M Saif, Yanning Shen, Tianyi Chen 0002
NeurIPS1
2024 Three-Way Trade-Off in Multi-Objective Learning: Optimization, Generalization and Conflict-Avoidance
abstract
Multi-objective learning (MOL) often arises in machine learning problems when there are multiple data modalities or tasks. One critical challenge in MOL is the potential conflict among different objectives during the optimization process. Recent works have developed various dynamic weighting algorithms for MOL, where the central idea is to find an update direction that avoids conflicts among objectives. Albeit its appealing intuition, empirical studies show that dynamic weighting methods may not outperform static ones. To understand this theory-practice gap, we focus on a stochastic variant of MGDA, the Multi-objective gradient with Double sampling (MoDo), and study the generalization performance and its interplay with optimization through the lens of algorithmic stability in the framework of statistical learning theory. We find that the key rationale behind MGDA—updating along conflict-avoidant direction—may hinder dynamic weighting algorithms from achieving the optimal $O(1/\sqrt{n})$ population risk, where $n$ is the number of training samples. We further demonstrate the impact of dynamic weights on the three-way trade-off among optimization, generalization, and conflict avoidance unique in MOL. We showcase the generality of our theoretical framework by analyzing other algorithms under the framework. Experiments on various multi-task learning benchmarks are performed to demonstrate the practical applicability. Code is available at https://github.com/heshandevaka/Trade-Off-MOL.
Lisha Chen, Heshan Devaka Fernando, Yiming Ying, Tianyi Chen 0002
J. Mach. Learn. Res.1
2023 A Nested Ensemble Method to Bilevel Machine Learning
abstract
Modern machine learning problems, such as hyperparameter optimization, meta learning, and adversarial training, adopt a bilevel learning formulation. Such problems involve a nested relation between inner- and outer-level problems, which often have suboptimal solutions with poor generalization ability. To address this issue, this paper proposes an ensemble method tailored to bilevel learning. Our method finds a nested ensemble of inner and outer parameters that improve generalization. We instantiate our general results with meta learning. We show theoretically and empirically that the diversity and the smoother loss landscape of the proposed ensemble methods lead to improved generalization over the state-of-the-art method.
Lisha Chen, Momin Abbas, Tianyi Chen 0002
ICASSP1
2023 Three-Way Trade-Off in Multi-Objective Learning: Optimization, Generalization and Conflict-Avoidance
abstract
Multi-objective learning (MOL) often arises in emerging machine learning problems when multiple learning criteria or tasks need to be addressed. Recent works have developed various _dynamic weighting_ algorithms for MOL, including MGDA and its variants, whose central idea is to find an update direction that _avoids conflicts_ among objectives. Albeit its appealing intuition, empirical studies show that dynamic weighting methods may not always outperform static alternatives. To bridge this gap between theory and practice, we focus on a new variant of stochastic MGDA - the Multi-objective gradient with Double sampling (MoDo) algorithm and study its generalization performance and the interplay with optimization through the lens of algorithm stability. We find that the rationale behind MGDA -- updating along conflict-avoidant direction - may \emph{impede} dynamic weighting algorithms from achieving the optimal ${\cal O}(1/\sqrt{n})$ population risk, where $n$ is the number of training samples. We further highlight the variability of dynamic weights and their impact on the three-way trade-off among optimization, generalization, and conflict avoidance that is unique in MOL. Code is available at https://github.com/heshandevaka/Trade-Off-MOL.
Lisha Chen, Heshan Devaka Fernando, Yiming Ying, Tianyi Chen 0002
NeurIPS1
2022 Is Bayesian Model-Agnostic Meta Learning Better than Model-Agnostic Meta Learning, Provably?
abstract
Meta learning aims at learning a model that can quickly adapt to unseen tasks. Widely used meta learning methods include model agnostic meta learning (MAML), implicit MAML, Bayesian MAML. Thanks to its ability of modeling uncertainty, Bayesian MAML often has advantageous empirical performance. However, the theoretical understanding of Bayesian MAML is still limited, especially on questions such as if and when Bayesian MAML has provably better performance than MAML. In this paper, we aim to provide theoretical justifications for Bayesian MAML’s advantageous performance by comparing the meta test risks of MAML and Bayesian MAML. In the meta linear regression, under both the distribution agnostic and linear centroid cases, we have established that Bayesian MAML indeed has provably lower meta test risks than MAML. We verify our theoretical results through experiments, the code of which is available at https://github.com/lishachen/Bayesian-MAML-vs-MAML.
Lisha Chen
AISTATS1
2022 Sharp-MAML: Sharpness-Aware Model-Agnostic Meta Learning
abstract
Model-agnostic meta learning (MAML) is currently one of the dominating approaches for few-shot meta-learning. Albeit its effectiveness, the optimization of MAML can be challenging due to the innate bilevel problem structure. Specifically, the loss landscape of MAML is much more complex with possibly more saddle points and local minimizers than its empirical risk minimization counterpart. To address this challenge, we leverage the recently invented sharpness-aware minimization and develop a sharpness-aware MAML approach that we term Sharp-MAML. We empirically demonstrate that Sharp-MAML and its computation-efficient variant can outperform the plain-vanilla MAML baseline (e.g., +3% accuracy on Mini-Imagenet). We complement the empirical study with the convergence rate analysis and the generalization bound of Sharp-MAML. To the best of our knowledge, this is the first empirical and theoretical study on sharpness-aware minimization in the context of bilevel learning.
Momin Abbas, Quan Xiao, Lisha Chen
ICML3
2022 Understanding Benign Overfitting in Gradient-Based Meta Learning
abstract
Meta learning has demonstrated tremendous success in few-shot learning with limited supervised data. In those settings, the meta model is usually overparameterized. While the conventional statistical learning theory suggests that overparameterized models tend to overfit, empirical evidence reveals that overparameterized meta learning methods still work well -- a phenomenon often called ``benign overfitting.'' To understand this phenomenon, we focus on the meta learning settings with a challenging bilevel structure that we term the gradient-based meta learning, and analyze its generalization performance under an overparameterized meta linear regression model. While our analysis uses the relatively tractable linear models, our theory contributes to understanding the delicate interplay among data heterogeneity, model adaptation and benign overfitting in gradient-based meta learning tasks. We corroborate our theoretical claims through numerical simulations.
Lisha Chen, Songtao Lu
NeurIPS1
2021 Uncertain Graph Neural Networks for Facial Action Unit Detection
abstract
Capturing the dependencies among different facial action units (AU) is extremely important for the AU detection task. Many studies have employed graph-based deep learning methods to exploit the dependencies among AUs. However, the dependencies among AUs in real world data are often noisy and the uncertainty is essential to be taken into consideration. Rather than employing a deterministic mode, we propose an uncertain graph neural network (UGN) to learn the probabilistic mask that simultaneously captures both the individual dependencies among AUs and the uncertainties. Further, we propose an adaptive weighted loss function based on the epistemic uncertainties to adaptively vary the weights of the training samples during the training process to account for unbalanced data distributions among AUs. We also provide an insightful analysis on how the uncertainties are related to the performance of AU detection. Extensive experiments, conducted on two benchmark datasets, i.e., BP4D and DISFA, demonstrate our method achieves the state-of-the-art performance.
Tengfei Song, Lisha Chen, Wenming Zheng
AAAI2
2019 Face Alignment With Kernel Density Deep Neural Network
abstract
Deep neural networks achieve good performance in many computer vision problems such as face alignment. However, when the testing image is challenging due to low resolution, occlusion or adversarial attacks, the accuracy of a deep neural network suffers greatly. Therefore, it is important to quantify the uncertainty in its predictions. A probabilistic neural network with Gaussian distribution over the target is typically used to quantify uncertainty for regression problems. However, in real-world problems especially computer vision tasks, the Gaussian assumption is too strong. To model more general distributions, such as multi-modal or asymmetric distributions, we propose to develop a kernel density deep neural network. Specifically, for face alignment, we adapt state-of-the-art hourglass neural network into a probabilistic neural network framework with landmark probability map as its output. The model is trained by maximizing the conditional log likelihood. To exploit the output probability map, we extend the model to multi-stage so that the logits map from the previous stage can feed into the next stage to progressively improve the landmark detection accuracy. Extensive experiments on benchmark datasets against state-of-the-art unconstrained deep learning method demonstrate that the proposed kernel density network achieves comparable or superior performance in terms of prediction accuracy. It further provides aleatoric uncertainty estimation in predictions.
Lisha Chen, Hui Su
ICCV1
2019 Embodied Conversational AI Agents in a Multi-modal Multi-agent Competitive Dialogue
abstract
In a setting where two AI agents embodied as animated humanoid avatars are engaged in a conversation with one human and each other, we see two challenges. One, determination by the AI agents about which one of them is being addressed. Two, determination by the AI agents if they may/could/should speak at the end of a turn. In this work we bring these two challenges together and explore the participation of AI agents in multi-party conversations. Particularly, we show two embodied AI shopkeeper agents who sell similar items aiming to get the business of a user by competing with each other on the price. In this scenario, we solve the first challenge by using headpose (estimated by deep learning techniques) to determine who the user is talking to. For the second challenge we use deontic logic to model rules of a negotiation conversation.
Rahul R. Divekar, Xiangyang Mou, Lisha Chen, Maíra Gatti de Bayser, Melina Alberio Guerra, Hui Su
IJCAI3
2019 You Talkin' to Me? A Practical Attention-Aware Embodied Agent
Rahul R. Divekar, Jeffrey O. Kephart, Xiangyang Mou, Lisha Chen, Hui Su
INTERACT (3)4
2019 Deep Structured Prediction for Facial Landmark Detection
abstract
Existing deep learning based facial landmark detection methods have achieved excellent performance. These methods, however, do not explicitly embed the structural dependencies among landmark points. They hence cannot preserve the geometric relationships between landmark points or generalize well to challenging conditions or unseen data. This paper proposes a method for deep structured facial landmark detection based on combining a deep Convolutional Network with a Conditional Random Field. We demonstrate its superior performance to existing state-of-the-art techniques in facial landmark detection, especially a better generalization ability on challenging datasets that include large pose and occlusion.
Lisha Chen, Hui Su
NeurIPS1
2017 A novel ZCC-MPC strategy with the input voltage observation for the induction motor drive system fed by an indirect matrix converter
abstract
In order to reduce the high cost of sampling circuits, simplify commutation procedure and improve the conversion efficiency of the indirect matrix converter (IMC), a zero current commutation model predictive control (ZCC-MPC) strategy with the input voltage observation is proposed in this paper. The input voltage observer is designed by using the model of the input LC filter, hence, three sampling and conditioning circuits for input phase voltages have been cancelled. The switching state combination with fixed duty cycle for the inverter stage is adopted in each sampling period, which ensures the zero dc-bus current during the whole commutation procedure in the rectifier stage. A simple two-step zero current commutation (ZCC) method has been adopted instead of complex four-step commutation. Simulation and experimental results show that sinusoidal input/output currents, good steady state/dynamic performance and unity input power factor have been achieved based on the proposed strategy. As well, the hardware cost has been lowered and the commutation reliability has been enhanced.
Yang Mei, Lisha Chen
IECON2
2014 Real-time damping estimation for variable impedance actuators
abstract
Recently-developed variable damping mechanisms have been exploited as a complement to compliant actuators. While accurate knowledge and control of generated damping is essential for achieving the desired performance, no physical sensor measuring the damping exists. This work introduces a novel non-model-based approach for the estimation of time-variant damping for variable impedance actuation systems. The approach is based only on torque and position/velocity measurements; without the knowledge of system's inputs, to ensure the estimation of both intentional and unintentional changes. Hence, a recursive least square estimator, modified for achieving a proper convergence for the estimation of time-variant parameters, is exploited. Experiments on a variable physical damping actuator are also presented to validate the performance of proposed approach.
Navvab Kashiri, Matteo Laffranchi, Jinoh Lee, Nikolaos G. Tsagarakis, Lisha Chen, Darwin G. Caldwell
ICRA5
2013 Optimal control for maximizing velocity of the CompAct™ compliant actuator
abstract
The CompAct™ actuator features a clutch mechanism placed in parallel with its passive series elastic transmission element and can therefore benefit from the advantages of both series elastic actuators (SEA) and rigid actuators. The actuator is capable of effectively managing the storage and release of the potential energy of the compliant element by the appropriate control of the clutch subsystem. Controlling the timing of the energy storage/release in the elastic element is exploited for improving motion control in this research. This paper analyses how this class of actuation systems can be used to maximize the link velocity of the joint. The dynamic model of the joint is derived and an optimal control strategy is proposed to identify optimal input reference profiles for the actuator (motor position/velocity and clutch activation timing) which permit the link velocity maximization. The effect of compliance of the joint on the performance of the system is studied and the optimal stiffness is analyzed.
Lisha Chen, Manolo Garabini, Matteo Laffranchi, Navvab Kashiri, Nikolaos G. Tsagarakis, Antonio Bicchi, Darwin G. Caldwell
ICRA1
2013 Link position control of a compliant actuator with unknown transmission friction torque
abstract
This paper proposes a control strategy for a compliant actuator, the CompAct™ actuator, which is equipped with semi active friction dampers in its transmission system. Both the transmission flexibility and the nonlinearity of the friction based damping torque makes the control of this actuator not a trivial task. This paper studies model of the presented actuator and the control problem of accurate link position tracking based on sliding mode approach that considers the friction torque as an uncertainty. Stability analysis and simulations highlight the effectiveness of the proposed controller in compensating for the deflections and unknown friction torque of the actuator. The performance of the controller is also validated by experiment results that demonstrate the tracking performance of the CompAct™ actuator achieved by the presented control strategy.
Lisha Chen, Matteo Laffranchi, Jinoh Lee, Navvab Kashiri, Nikolaos G. Tsagarakis, Darwin G. Caldwell
IROS1
2013 Stress functions for nonlinear dimension reduction, proximity analysis, and graph drawing
Lisha Chen, Andreas Buja
J. Mach. Learn. Res.1
2012 The role of physical damping in compliant actuation systems
abstract
Recently, compliance has been considered as one of the key physical properties that a robot should incorporate to be able to physically interact with humans and uncertain environments. Apart from the improved ability of interaction, mechanical robustness and higher safety-related performances, compliance introduces underdamped oscillatory modes and reduces the mechanical natural frequency of the plant to be controlled making its control much more complex than that of conventional stiff actuators. To overcome these drawbacks, some recent works focus on the incorporation of physical damping within compliant actuators. This work presents an analysis for the quantitative evaluation of the effects of physical damping in compliant robotic joints to demonstrate the improvements (dynamic performance, stability, controllability, tracking precision and energy efficiency) which can be gained by incorporating physical damping in such flexible transmission systems. Simulation and experimental results validate that these benefits can effectively be achieved on an existing compliant actuator prototype with variable physical damping.
Matteo Laffranchi, Lisha Chen, Nikolaos G. Tsagarakis, Darwin G. Caldwell
IROS2