Xingjian Li 0005

dblp:266/7377 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
0000-0001-7654-6833ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Deep learning architectures and training · 33% Generative modeling · 24% Motion planning and robot control · 21%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Motion planning and robot control › robot control
optimal control
0.912025
Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency · NeurIPS 2025
Machine learning › Trustworthy machine learning › robustness › robust learning
robust generalization
0.912025
Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency · NeurIPS 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency · NeurIPS 2025
Machine learning › Generative modeling › normalizing flow
continuous normalizing flow
0.512021
OT-Flow: Fast and Accurate Continuous Normalizing Flows via Optimal Transport · AAAI 2021
Machine learning › Generative modeling
normalizing flow
0.512021
OT-Flow: Fast and Accurate Continuous Normalizing Flows via Optimal Transport · AAAI 2021
Machine learning › Deep learning architectures and training › regularization
optimal transport regularization
0.512021
OT-Flow: Fast and Accurate Continuous Normalizing Flows via Optimal Transport · AAAI 2021

Methods — techniques the papers use, named apart from their topics

optimal control theory · 0.9continuous-time formulation · 0.9optimal transport · 0.5neural ordinary differential equation · 0.5
YearPublicationVenuePosition
2025 Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency
abstract
We study Transformers through the perspective of optimal control theory, using tools from continuous-time formulations to derive actionable insights into training and architecture design. This framework improves the performance of existing Transformer models while providing desirable theoretical guarantees, including generalization and robustness. Our framework is designed to be plug-and-play, enabling seamless integration with established Transformer models and requiring only slight changes to the implementation. We conduct seven extensive experiments on tasks motivated by text generation, sentiment analysis, image classification, and point cloud classification. Experimental results show that the framework improves the test performance of the baselines, while being more parameter-efficient. On character-level text generation with nanoGPT, our framework achieves a 46\% reduction in final test loss while using 42\% fewer parameters. On GPT-2, our framework achieves a 9.3\% reduction in final test loss, demonstrating scalability to larger models. To the best of our knowledge, this is the first work that applies optimal control theory to both the training and architecture of Transformers. It offers a new foundation for systematic, theory-driven improvements and moves beyond costly trial-and-error approaches.
Kelvin Kan, Xingjian Li 0005, Benjamin J. Zhang, Tuhin Sahai, Stanley J. Osher, Markos A. Katsoulakis
NeurIPS2
2021 OT-Flow: Fast and Accurate Continuous Normalizing Flows via Optimal Transport
abstract
A normalizing flow is an invertible mapping between an arbitrary probability distribution and a standard normal distribution; it can be used for density estimation and statistical inference. Computing the flow follows the change of variables formula and thus requires invertibility of the mapping and an efficient way to compute the determinant of its Jacobian. To satisfy these requirements, normalizing flows typically consist of carefully chosen components. Continuous normalizing flows (CNFs) are mappings obtained by solving a neural ordinary differential equation (ODE). The neural ODE's dynamics can be chosen almost arbitrarily while ensuring invertibility. Moreover, the log-determinant of the flow's Jacobian can be obtained by integrating the trace of the dynamics' Jacobian along the flow. Our proposed OT-Flow approach tackles two critical computational challenges that limit a more widespread use of CNFs. First, OT-Flow leverages optimal transport (OT) theory to regularize the CNF and enforce straight trajectories that are easier to integrate. Second, OT-Flow features exact trace computation with time complexity equal to trace estimators used in existing CNFs. On five high-dimensional density estimation and generative modeling tasks, OT-Flow performs competitively to state-of-the-art CNFs while on average requiring one-fourth of the number of weights with an 8x speedup in training time and 24x speedup in inference.
Derek Onken, Samy Wu Fung, Xingjian Li 0005, Lars Ruthotto
AAAI3