EDBT 2026 Demo / reviewers in the wild / expert
Qianxiao Li
dblp:172/0930
· DBLP profile ↗
33ranked-venue papers
5as first author
26since 2021 · last 2025
0000-0002-3903-3737ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 5 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Autocorrelation Matters: Understanding the Role of Initialization Schemes for State Space ModelsabstractCurrent methods for initializing state space model (SSM) parameters primarily rely on the HiPPO framework \citep{gu2023how}, which is based on online function approximation with the SSM kernel basis.
However, the HiPPO framework does not explicitly account for the effects of the temporal structures of input sequences on the optimization of SSMs.
In this paper, we take a further step to investigate the roles of SSM initialization schemes by considering the autocorrelation of input sequences.
Specifically, we: (1) rigorously characterize the dependency of the SSM timescale on sequence length based on sequence autocorrelation; (2) find that with a proper timescale, allowing a zero real part for the eigenvalues of the SSM state matrix mitigates the curse of memory while still maintaining stability at initialization; (3) show that the imaginary part of the eigenvalues of the SSM state matrix determines the conditioning of SSM optimization problems, and uncover an approximation-estimation tradeoff when training SSMs with a specific class of target functions. Fusheng Liu, Qianxiao Li |
ICLR | 2 |
| 2025 | BP-Modified Local Loss for Efficient Training of Deep Neural NetworksabstractThe training of large models is memory-constrained, one direction to relieve
this is training using local loss, like GIM, LoCo, and Forward-Forward
algorithms. However, the local loss methods often face the issue of slow or
non-convergence. In this paper, we propose a novel BP-modified local loss
method that uses the true Backward Propagation (BP) gradient to modify the
local loss gradient to improve the performance of local loss training. We
use the stochastic modified equation to analyze our method and show that
modified offset decreases the bias between the BP gradient and local loss
gradient, but introduces additional variance, which results in a
bias-variance balance. Numerical experiments on full-tuning and LoKr tuning
on the ResNet-50 model and LoRA tuning on the ViT-b16 model on CIFAR-100
datasets show 20.5\% test top-1 accuracy improvement for the Forward-Forward
algorithm, 18.6\% improvement for LoCo algorithm and achieve only on average
7.7\% of test accuracy loss compared to the BP algorithm, with up to 75\%
memory savings. Lianhai Ren, Qianxiao Li |
ICLR | 2 |
| 2025 | Continuity-Preserving Convolutional Autoencoders for Learning Continuous Latent Dynamical Models from ImagesabstractContinuous dynamical systems are cornerstones of many scientific and engineering disciplines.
While machine learning offers powerful tools to model these systems from trajectory data, challenges arise when these trajectories are captured as images, resulting in pixel-level observations that are discrete in nature.
Consequently, a naive application of a convolutional autoencoder can result in latent coordinates that are discontinuous in time.
To resolve this, we propose continuity-preserving convolutional autoencoders (CpAEs) to learn continuous latent states and their corresponding continuous latent dynamical models from discrete image frames.
We present a mathematical formulation for learning dynamics from image frames, which illustrates issues with previous approaches and motivates our methodology based on promoting the continuity of convolution filters, thereby preserving the continuity of the latent states.
This approach enables CpAEs to produce latent states that evolve continuously with the underlying dynamics, leading to more accurate latent dynamical models.
Extensive experiments across various scenarios demonstrate the effectiveness of CpAEs. Aiqing Zhu, Yuting Pan, Qianxiao Li |
ICLR | 3 |
| 2025 | From Weight-Based to State-Based Fine-Tuning: Further Memory Reduction on LoRA with Parallel ControlabstractThe LoRA method has achieved notable success in reducing GPU memory usage by applying low-rank updates to weight matrices. Yet, one simple question remains: can we push this reduction even further? Furthermore, is it possible to achieve this while improving performance and reducing computation time? Answering these questions requires moving beyond the conventional weight-centric approach. In this paper, we present a state-based fine-tuning framework that shifts the focus from weight adaptation to optimizing forward states, with LoRA acting as a special example. Specifically, state-based tuning introduces parameterized perturbations to the states within the computational graph, allowing us to control states across an entire residual block. A key advantage of this approach is the potential to avoid storing large intermediate states in models like transformers. Empirical results across multiple architectures—including ViT, RoBERTa, LLaMA2-7B, and LLaMA3-8B—show that our method further reduces memory consumption and computation time while simultaneously improving performance. Moreover, as a result of memory reduction, we explore the feasibility to train 7B/8B models on consumer-level GPUs like Nvidia 3090, without model quantization. The code is available at an anonymous GitHub repository Lianhai Ren, Jingpu Cheng, Qianxiao Li |
ICML | 4 |
| 2025 | A unified framework for establishing the universal approximation of transformer-type architecturesabstractWe investigate the universal approximation property (UAP) of transformer-type architectures, providing a unified theoretical framework that extends prior results on residual networks to models incorporating attention mechanisms. Our work identifies token distinguishability as a fundamental requirement for UAP and introduces a general sufficient condition that applies to a broad class of architectures. Leveraging an analyticity assumption on the attention layer, we can significantly simplify the verification of this condition, providing a non-constructive approach in establishing UAP for such architectures. We demonstrate the applicability of our framework by proving UAP for transformers with various attention mechanisms, including kernel-based and sparse ones. The corollaries of our results either generalize prior works or establish UAP for architectures not previously covered. Furthermore, our framework offers a principled foundation for designing novel transformer architectures with inherent UAP guarantees, including those with specific functional symmetries. We propose examples to illustrate these insights. Jingpu Cheng, Ting Lin, Zuowei Shen, Qianxiao Li |
NeurIPS | 4 |
| 2024 | An Optimal Control View of LoRA and Binary Controller Design for Vision Transformers
Chi Zhang 0123, Jingpu Cheng, Qianxiao Li |
ECCV (53) | 3 |
| 2024 | Inverse Approximation Theory for Nonlinear Recurrent Neural NetworksabstractWe prove an inverse approximation theorem for the approximation of nonlinear sequence-to-sequence relationships using recurrent neural networks (RNNs). This is a so-called Bernstein-type result in approximation theory, which deduces properties of a target function under the assumption that it can be effectively approximated by a hypothesis space. In particular, we show that nonlinear sequence relationships that can be stably approximated by nonlinear RNNs must have an exponential decaying memory structure - a notion that can be made precise. This extends the previously identified curse of memory in linear RNNs into the general nonlinear setting, and quantifies the essential limitations of the RNN architecture for learning sequential relationships with long-term memory. Based on the analysis, we propose a principled reparameterization method to overcome the limitations. Our theoretical results are confirmed by numerical experiments. Zhong Li 0004, Qianxiao Li |
ICLR | 3 |
| 2024 | Parameter-Efficient Fine-Tuning with ControlsabstractIn contrast to the prevailing interpretation of Low-Rank Adaptation (LoRA) as a means of simulating weight changes in model adaptation, this paper introduces an alternative perspective by framing it as a control process. Specifically, we conceptualize lightweight matrices in LoRA as control modules tasked with perturbing the original, complex, yet frozen blocks on downstream tasks. Building upon this new understanding, we conduct a thorough analysis on the controllability of these modules, where we identify and establish sufficient conditions that facilitate their effective integration into downstream controls. Moreover, the control modules are redesigned by incorporating nonlinearities through a parameter-free attention mechanism. This modification allows for the intermingling of tokens within the controllers, enhancing the adaptability and performance of the system. Empirical findings substantiate that, without introducing any additional parameters, this approach surpasses the existing LoRA algorithms across all assessed datasets and rank configurations. Chi Zhang 0123, Jingpu Cheng, Yanyu Xu 0001, Qianxiao Li |
ICML | 4 |
| 2024 | Accelerating Legacy Numerical Solvers by Non-intrusive Gradient-based Meta-solvingabstractScientific computing is an essential tool for scientific discovery and engineering design, and its computational cost is always a main concern in practice. To accelerate scientific computing, it is a promising approach to use machine learning (especially meta-learning) techniques for selecting hyperparameters of traditional numerical methods. There have been numerous proposals to this direction, but many of them require automatic-differentiable numerical methods. However, in reality, many practical applications still depend on well-established but non-automatic-differentiable legacy codes, which prevents practitioners from applying the state-of-the-art research to their own problems. To resolve this problem, we propose a non-intrusive methodology with a novel gradient estimation technique to combine machine learning and legacy numerical codes without any modification. We theoretically and numerically show the advantage of the proposed method over other baselines and present applications of accelerating established non-automatic-differentiable numerical solvers implemented in PETSc, a widely used open-source numerical software library. Sohei Arisaka, Qianxiao Li |
ICML | 2 |
| 2024 | From Generalization Analysis to Optimization Designs for State Space ModelsabstractA State Space Model (SSM) is a foundation model in time series analysis, which has recently been shown as an alternative to transformers in sequence modeling. In this paper, we theoretically study the generalization of SSMs and propose improvements to training algorithms based on the generalization results. Specifically, we give a data-dependent generalization bound for SSMs, showing an interplay between the SSM parameters and the temporal dependencies of the training sequences. Leveraging the generalization bound, we (1) set up a scaling rule for model initialization based on the proposed generalization measure, which significantly improves the robustness of the output value scales on SSMs to different temporal patterns in the sequence data; (2) introduce a new regularization method for training SSMs to enhance the generalization performance. Numerical results are conducted to validate our results. Fusheng Liu, Qianxiao Li |
ICML | 2 |
| 2024 | StableSSM: Alleviating the Curse of Memory in State-space Models through Stable ReparameterizationabstractIn this paper, we investigate the long-term memory learning capabilities of state-space models (SSMs) from the perspective of parameterization. We prove that state-space models without any reparameterization exhibit a memory limitation similar to that of traditional RNNs: the target relationships that can be stably approximated by state-space models must have an exponential decaying memory. Our analysis identifies this ``curse of memory'' as a result of the recurrent weights converging to a stability boundary, suggesting that a reparameterization technique can be effective. To this end, we introduce a class of reparameterization techniques for SSMs that effectively lift its memory limitations. Besides improving approximation capabilities, we further illustrate that a principled choice of reparameterization scheme can also enhance optimization stability. We validate our findings using synthetic datasets, language models and image classifications. Qianxiao Li |
ICML | 2 |
| 2024 | Learning Macroscopic Dynamics from Partial Microscopic ObservationsabstractMacroscopic observables of a system are of keen interest in real applications such as the design of novel materials. Current methods rely on microscopic trajectory simulations, where the forces on all microscopic coordinates need to be computed or measured. However, this can be computationally prohibitive for realistic systems. In this paper, we propose a method to learn macroscopic dynamics requiring only force computations on a subset of the microscopic coordinates. Our method relies on a sparsity assumption: the force on each microscopic coordinate relies only on a small number of other coordinates. The main idea of our approach is to map the training procedure on the macroscopic coordinates back to the microscopic coordinates, on which partial force computations can be used as stochastic estimation to update model parameters. We provide a theoretical justification of this under suitable conditions. We demonstrate the accuracy, force computation efficiency, and robustness of our method on learning macroscopic closure models from a variety of microscopic systems, including those modeled by partial differential equations or molecular dynamics simulations. Mengyi Chen, Qianxiao Li |
NeurIPS | 2 |
| 2024 | Approximation Rate of the Transformer Architecture for Sequence ModelingabstractThe Transformer architecture is widely applied in sequence modeling applications, yet the theoretical understanding of its working principles remains limited. In this work, we investigate the approximation rate for single-layer Transformers with one head. We consider general non-linear relationships and identify a novel notion of complexity measures to establish an explicit Jackson-type approximation rate estimate for the Transformer. This rate reveals the structural properties of the Transformer and suggests the types of sequential relationships it is best suited for approximating. In particular, the results on approximation rates enable us to concretely analyze the differences between the Transformer and classical sequence modeling methods, such as recurrent neural networks. Qianxiao Li |
NeurIPS | 2 |
| 2024 | Deep Neural Network Approximation of Invariant Functions through Dynamical SystemsabstractWe study the approximation of functions which are invariant with respect to certain permutations of the input indices using flow maps of dynamical systems. Such invariant functions include the much studied translation-invariant ones involving image tasks, but also encompasses many permutation-invariant functions that find emerging applications in science and engineering. We prove sufficient conditions for universal approximation of these functions by a controlled dynamical system, which can be viewed as a general abstraction of deep residual networks with symmetry constraints. These results not only imply the universal approximation for a variety of commonly employed neural network architectures for symmetric function approximation, but also guide the design of architectures with approximation guarantees for applications involving new symmetry requirements. Qianxiao Li, Ting Lin, Zuowei Shen |
J. Mach. Learn. Res. | 1 |
| 2023 | Principled Acceleration of Iterative Numerical Methods Using Machine LearningabstractIterative methods are ubiquitous in large-scale scientific computing applications, and a number of approaches based on meta-learning have been recently proposed to accelerate them. However, a systematic study of these approaches and how they differ from meta-learning is lacking. In this paper, we propose a framework to analyze such learning-based acceleration approaches, where one can immediately identify a departure from classical meta-learning. We theoretically show that this departure may lead to arbitrary deterioration of model performance, and at the same time, we identify a methodology to ameliorate it by modifying the loss objective, leading to a novel training method for learning-based acceleration of iterative algorithms. We demonstrate the significant advantage and versatility of the proposed approach through various numerical applications. Sohei Arisaka, Qianxiao Li |
ICML | 2 |
| 2023 | An Annealing Mechanism for Adversarial Training AccelerationabstractDespite the empirical success in various domains, it has been revealed that deep neural networks are vulnerable to maliciously perturbed input data that can dramatically degrade their performance. These are known as adversarial attacks. To counter adversarial attacks, adversarial training formulated as a form of robust optimization has been demonstrated to be effective. However, conducting adversarial training brings much computational overhead compared with standard training. In order to reduce the computational cost, we propose an annealing mechanism, annealing mechanism for adversarial training acceleration (Amata), to reduce the overhead associated with adversarial training. The proposed Amata is provably convergent, well-motivated from the lens of optimal control theory, and can be combined with existing acceleration methods to further enhance performance. It is demonstrated that, on standard datasets, Amata can achieve similar or better robustness with around 1/3-1/2 the computational time compared with traditional methods. In addition, Amata can be incorporated into other adversarial training acceleration algorithms (e.g., YOPO, Free, Fast, and ATTA), which leads to a further reduction in computational time on large-scale problems. Nanyang Ye 0001, Qianxiao Li, Xiaoyun Zhou 0001, Zhanxing Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | On the approximation properties of recurrent encoder-decoder architectures
Zhong Li 0004, Qianxiao Li |
ICLR | 3 |
| 2022 | Unraveling Model-Agnostic Meta-Learning via The Adaptation Learning Rate
Yingtian Zou, Fusheng Liu, Qianxiao Li |
ICLR | 3 |
| 2022 | Self-Healing Robust Neural Networks via Closed-Loop ControlabstractDespite the wide applications of neural networks, there have been increasing concerns about their vulnerability issue. While numerous attack and defense techniques have been developed, this work investigates the robustness issue from a new angle: can we design a self-healing neural network that can automatically detect and fix the vulnerability issue by itself? A typical self-healing mechanism is the immune system of a human body. This biology-inspired idea has been used in many engineering designs but has rarely been investigated in deep learning. This paper considers the post-training self-healing of a neural network, and proposes a closed-loop control formulation to automatically detect and fix the errors caused by various attacks or perturbations. We provide a margin-based analysis to explain how this formulation can improve the robustness of a classifier. To speed up the inference, we convert the optimal control problem to Pontryagon's Maximum Principle and solve it via the method of successive approximation. Lastly, we present an error estimation of the proposed framework for neural networks with nonlinear activation functions. We validate the performance of several network architectures against various perturbations. Since the self-healing method does not need a-priori information about data perturbations or attacks, it can handle a broad class of unforeseen perturbations. Zhoutong Chen, Qianxiao Li |
J. Mach. Learn. Res. | 2 |
| 2022 | Approximation and Optimization Theory for Linear Continuous-Time Recurrent Neural NetworksabstractWe perform a systematic study of the approximation properties and optimization dynamics of recurrent neural networks (RNNs) when applied to learn input-output relationships in temporal data. We consider the simple but representative setting of using continuous-time linear RNNs to learn from data generated by linear relationships. On the approximation side, we prove a direct and an inverse approximation theorem of linear functionals using RNNs, which reveal the intricate connections between memory structures in the target and the corresponding approximation efficiency. In particular, we show that temporal relationships can be effectively approximated by RNNs if and only if the former possesses sufficient memory decay. On the optimization front, we perform detailed analysis of the optimization dynamics, including a precise understanding of the difficulty that may arise in learning relationships with long-term memory. The term “curse of memory” is coined to describe the uncovered phenomena, akin to the “curse of dimension” that plagues high-dimensional function approximation. These results form a relatively complete picture of the interaction of memory and recurrent structures in the linear dynamical setting. Zhong Li 0004, Jiequn Han, Weinan E, Qianxiao Li |
J. Mach. Learn. Res. | 4 |
| 2021 | Amata: An Annealing Mechanism for Adversarial Training AccelerationabstractDespite the empirical success in various domains, it has been revealed that deep neural networks are vulnerable to maliciously perturbed input data that much degrade their performance. This is known as adversarial attacks. To counter adversarial attacks, adversarial training formulated as a form of robust optimization has been demonstrated to be effective. However, conducting adversarial training brings much computational overhead compared with standard training. In order to reduce the computational cost, we propose an annealing mechanism, Amata, to reduce the overhead associated with adversarial training. The proposed Amata is provably convergent, well-motivated from the lens of optimal control theory and can be combined with existing acceleration methods to further enhance performance. It is demonstrated that on standard datasets, Amata can achieve similar or better robustness with around 1/3 to 1/2 the computational time compared with traditional methods. In addition, Amata can be incorporated into other adversarial training acceleration algorithms (e.g. YOPO, Free, Fast, and ATTA), which leads to further reduction in computational time on large-scale problems. Nanyang Ye 0001, Qianxiao Li, Xiaoyun Zhou 0001, Zhanxing Zhu |
AAAI | 2 |
| 2021 | Adversarial Invariant LearningabstractThough machine learning algorithms are able to achieve pattern recognition from the correlation between data and labels, the presence of spurious features in the data decreases the robustness of these learned relationships with respect to varied testing environments. This is known as out-of-distribution (OoD) generalization problem. Recently, invariant risk minimization (IRM) attempts to tackle this issue by penalizing predictions based on the unstable spurious features in the data collected from different environments. However, similar to domain adaptation or domain generalization, a prevalent non-trivial limitation in these works is that the environment information is assigned by human specialists, i.e. a priori, or determined heuristically. However, an inappropriate group partitioning can dramatically deteriorate the OoD generalization and this process is expensive and time-consuming. To deal with this issue, we propose a novel theoretically principled min-max framework to iteratively construct a worst-case splitting, i.e. creating the most challenging environment splittings for the backbone learning paradigm (e.g. IRM) to learn the robust feature representation. We also design a differentiable training strategy to facilitate the feasible gradient- based computation. Numerical experiments show that our algorithmic framework has achieved superior and stable performance in various datasets, such as Colored MNIST and Punctuated Stanford sentiment treebank (SST). Furthermore, we also find our algorithm to be robust even to a strong data poisoning attack. To the best of our knowledge, this is one of the first to adopt differentiable environment splitting method to enable stable predictions across environments without environment index information, which achieves the state-of-the-art performance on datasets with strong spurious correlation, such as Colored MNIST. Nanyang Ye 0001, Jingxuan Tang, Huayu Deng, Xiaoyun Zhou 0001, Qianxiao Li, Zhenguo Li, Guang-Zhong Yang, Zhanxing Zhu |
CVPR | 5 |
| 2021 | Towards Robust Neural Networks via Close-loop Control
Zhoutong Chen, Qianxiao Li |
ICLR | 2 |
| 2021 | On the Curse of Memory in Recurrent Neural Networks: Approximation and Optimization Analysis
Zhong Li 0004, Jiequn Han, Weinan E, Qianxiao Li |
ICLR | 4 |
| 2021 | Approximation Theory of Convolutional Architectures for Time Series ModellingabstractWe study the approximation properties of convolutional architectures applied to time series modelling, which can be formulated mathematically as a functional approximation problem. In the recurrent setting, recent results reveal an intricate connection between approximation efficiency and memory structures in the data generation process. In this paper, we derive parallel results for convolutional architectures, with WaveNet being a prime example. Our results reveal that in this new setting, approximation efficiency is not only characterised by memory, but also additional fine structures in the target relationship. This leads to a novel definition of spectrum-based regularity that measures the complexity of temporal relationships under the convolutional approximation scheme. These analyses provide a foundation to understand the differences between architectural choices for time series modelling and can give theoretically grounded guidance for practical applications. Zhong Li 0004, Qianxiao Li |
ICML | 3 |
| 2021 | Distributed optimization for degenerate loss functions arising from over-parameterization
Chi Zhang 0123, Qianxiao Li |
Artif. Intell. | 2 |
| 2019 | A Quantitative Analysis of the Effect of Batch Normalization on Gradient DescentabstractDespite its empirical success and recent theoretical progress, there generally lacks a quantitative analysis of the effect of batch normalization (BN) on the convergence and stability of gradient descent. In this paper, we provide such an analysis on the simple problem of ordinary least squares (OLS), where the precise dynamical properties of gradient descent (GD) is completely known, thus allowing us to isolate and compare the additional effects of BN. More precisely, we show that unlike GD, gradient descent with BN (BNGD) converges for arbitrary learning rates for the weights, and the convergence remains linear under mild conditions. Moreover, we quantify two different sources of acceleration of BNGD over GD – one due to over-parameterization which improves the effective condition number and another due having a large range of learning rates giving rise to fast descent. These phenomena set BNGD apart from GD and could account for much of its robustness properties. These findings are confirmed quantitatively by numerical experiments, which further show that many of the uncovered properties of BNGD in OLS are also observed qualitatively in more complex supervised learning problems. Yongqiang Cai, Qianxiao Li, Zuowei Shen |
ICML | 2 |
| 2019 | Decentralized Optimization with Edge SamplingabstractIn this paper, we propose a decentralized distributed algorithm with stochastic communication among nodes, building on a sampling method called "edge sampling''. Such a sampling algorithm allows us to avoid the heavy peer-to-peer communication cost when combining neighboring weights on dense networks while still maintains a comparable convergence rate. In particular, we quantitatively analyze its theoretical convergence properties, as well as the optimal sampling rate over the underlying network. When compared with previous methods, our solution is shown to be unbiased, communication-efficient and suffers from lower sampling variances. These theoretical findings are validated by both numerical experiments on the mixing rates of Markov Chains and distributed machine learning problems. Chi Zhang 0123, Qianxiao Li, Peilin Zhao |
IJCAI | 2 |
| 2019 | Stochastic Modified Equations and Dynamics of Stochastic Gradient Algorithms I: Mathematical FoundationsabstractWe develop the mathematical foundations of the stochastic modified equations (SME) framework for analyzing the dynamics of stochastic gradient algorithms, where the latter is approximated by a class of stochastic differential equations with small noise parameters. We prove that this approximation can be understood mathematically as an weak approximation, which leads to a number of precise and useful results on the approximations of stochastic gradient descent (SGD), momentum SGD and stochastic Nesterov's accelerated gradient method in the general setting of stochastic objectives. We also demonstrate through explicit calculations that this continuous-time approach can uncover important analytical insights into the stochastic gradient algorithms under consideration that may not be easy to obtain in a purely discrete-time setting. Qianxiao Li, Cheng Tai, Weinan E |
J. Mach. Learn. Res. | 1 |
| 2018 | An Optimal Control Approach to Deep Learning and Applications to Discrete-Weight Neural NetworksabstractDeep learning is formulated as a discrete-time optimal control problem. This allows one to characterize necessary conditions for optimality and develop training algorithms that do not rely on gradients with respect to the trainable parameters. In particular, we introduce the discrete-time method of successive approximations (MSA), which is based on the Pontryagin’s maximum principle, for training neural networks. A rigorous error estimate for the discrete MSA is obtained, which sheds light on its dynamics and the means to stabilize the algorithm. The developed methods are applied to train, in a rather principled way, neural networks with weights that are constrained to take values in a discrete set. We obtain competitive performance and interestingly, very sparse weights in the case of ternary networks, which may be useful in model deployment in low-memory devices. Qianxiao Li, Shuji Hao |
ICML | 1 |
| 2018 | Turn-by-turn Intelligent Manoeuvring of Driverless Taxis: A Recursive Value Model Enhanced by Reinforcement LearningabstractWe develop a protocol for constructing intelligent manoeuvring of driverless taxis - especially for empty taxis - with the aim of optimising the efficiency of the taxi system (e.g. minimising the average waiting time of the commuters). The manoeuvring is formulated as a stochastic policy matrix giving optimal choices on which way to go for taxis at any locations, based on the historical patterns of commuter origin and destination distribution. A recursive value (RV) model is formulated for turn-by-turn navigation. The policy matrix is solved by a combination of the RV model and reinforcement learning, together with a comprehensive agent-based microscopic simulation platform we have developed. We show that even for very large road networks, very efficient manoeuvring algorithms can be readily found with reinforcement learning, especially with our RV model as the initial state. Such algorithms are indispensable for the fleet of driverless taxis, and can also give helpful recommendations to human drivers operating taxis. The general framework of our protocol allows the construction of optimal manoeuvring algorithms for a wide range of road networks in a systematic way. Qianxiao Li |
Intelligent Vehicles Symposium | 2 |
| 2017 | Stochastic Modified Equations and Adaptive Stochastic Gradient AlgorithmsabstractWe develop the method of stochastic modified equations (SME), in which stochastic gradient algorithms are approximated in the weak sense by continuous-time stochastic differential equations. We exploit the continuous formulation together with optimal control theory to derive novel adaptive hyper-parameter adjustment policies. Our algorithms have competitive performance with the added benefit of being robust to varying models and datasets. This provides a general methodology for the analysis and design of stochastic gradient algorithms. Qianxiao Li, Cheng Tai, Weinan E |
ICML | 1 |
| 2017 | Maximum Principle Based Algorithms for Deep Learning
Qianxiao Li, Cheng Tai, Weinan E |
J. Mach. Learn. Res. | 1 |