Yanxi Chen 0001

dblp:40/5750-1 · DBLP profile ↗
← Back
6ranked-venue papers
6as first author
4since 2021 · last 2025
0000-0003-0610-8103ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorTheory of computation · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 38% Efficient and distributed learning · 35% Probabilistic and Bayesian machine learning · 9%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
test-time scaling
0.912025
Provable Scaling Laws for the Test-Time Compute of Large Language Models · NeurIPS 2025
Machine learning › Efficient and distributed learning › distributed training › hybrid parallel training
3d parallelism
0.812024
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism · ICML 2024
Machine learning › Efficient and distributed learning › adaptive computation
early exit
0.812024
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism · ICML 2024
Machine learning › Efficient and distributed learning
inference acceleration
0.812024
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism · ICML 2024
Natural language and speech › Language models and text generation
large language model training
0.812024
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism · ICML 2024
Parallel and multicore computing
pipeline parallelism
0.812024
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.612022
Learning Mixtures of Linear Dynamical Systems · ICML 2022
Machine learning › Learning theory
sample complexity
0.612022
Learning Mixtures of Linear Dynamical Systems · ICML 2022
Machine learning › Representation and self-supervised learning
low-rank models
0.512021
Learning Mixtures of Low-Rank Models · IEEE Trans. Inf. Theory 2021
Machine learning › Optimization for machine learning
non-convex optimization
0.112021
Learning Mixtures of Low-Rank Models · IEEE Trans. Inf. Theory 2021

Methods — techniques the papers use, named apart from their topics

pipeline parallelism · 1.5backpropagation · 1.5KV caching · 1.5league-style evaluation · 0.9knockout tournament · 0.9spectral learning · 0.6method of moments · 0.6meta-algorithm · 0.5matrix sensing · 0.5
YearPublicationVenuePosition
2025 Provable Scaling Laws for the Test-Time Compute of Large Language Models
abstract
We propose two simple, principled and practical algorithms that enjoy provable scaling laws for the test-time compute of large language models (LLMs). The first one is a two-stage knockout-style algorithm: given an input problem, it first generates multiple candidate solutions, and then aggregate them via a knockout tournament for the final output. Assuming that the LLM can generate a correct solution with non-zero probability and do better than a random guess in comparing a pair of correct and incorrect solutions, we prove theoretically that the failure probability of this algorithm decays to zero exponentially or by a power law (depending on the specific way of scaling) as its test-time compute grows. The second one is a two-stage league-style algorithm, where each candidate is evaluated by its average win rate against multiple opponents, rather than eliminated upon loss to a single opponent. Under analogous but more robust assumptions, we prove that its failure probability also decays to zero exponentially with more test-time compute. Both algorithms require a black-box LLM and nothing else (e.g., no verifier or reward model) for a minimalistic implementation, which makes them appealing for practical applications and easy to adapt for different tasks. Through extensive experiments with diverse models and datasets, we validate the proposed theories and demonstrate the outstanding scaling properties of both algorithms.
Yanxi Chen 0001, Xuchen Pan, Yaliang Li, Bolin Ding, Jingren Zhou 0001
NeurIPS1
2024 EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism
abstract
We present EE-LLM, a framework for large-scale training and inference of early-exit large language models (LLMs). While recent works have shown preliminary evidence for the efficacy of early exiting in accelerating LLM inference, EE-LLM makes a foundational step towards scaling up early-exit LLMs by supporting their training and inference with massive 3D parallelism. Built upon Megatron-LM, EE-LLM implements a variety of algorithmic innovations and performance optimizations tailored to early exiting, including a lightweight method that facilitates backpropagation for the early-exit training objective with pipeline parallelism, techniques of leveraging idle resources in the original pipeline schedule for computation related to early-exit layers, and two approaches of early-exit inference that are compatible with KV caching for autoregressive generation. Our analytical and empirical study shows that EE-LLM achieves great training efficiency with negligible computational overhead compared to standard LLM training, as well as outstanding inference speedup without compromising output quality. To facilitate further research and adoption, we release EE-LLM at https://github.com/pan-x-c/EE-LLM.
Yanxi Chen 0001, Xuchen Pan, Yaliang Li, Bolin Ding, Jingren Zhou 0001
ICML1
2022 Learning Mixtures of Linear Dynamical Systems
abstract
We study the problem of learning a mixture of multiple linear dynamical systems (LDSs) from unlabeled short sample trajectories, each generated by one of the LDS models. Despite the wide applicability of mixture models for time-series data, learning algorithms that come with end-to-end performance guarantees are largely absent from existing literature. There are multiple sources of technical challenges, including but not limited to (1) the presence of latent variables (i.e. the unknown labels of trajectories); (2) the possibility that the sample trajectories might have lengths much smaller than the dimension $d$ of the LDS models; and (3) the complicated temporal dependence inherent to time-series data. To tackle these challenges, we develop a two-stage meta-algorithm, which is guaranteed to efficiently recover each ground-truth LDS model up to error $\tilde{O}(\sqrt{d/T})$, where $T$ is the total sample size. We validate our theoretical studies with numerical experiments, confirming the efficacy of the proposed algorithm.
Yanxi Chen 0001, H. Vincent Poor
ICML1
2021 Learning Mixtures of Low-Rank Models
abstract
We study the problem of learning mixtures of low-rank models, i.e. reconstructing multiple low-rank matrices from unlabelled linear measurements of each. This problem enriches two widely studied settings - low-rank matrix sensing and mixed linear regression - by bringing latent variables (i.e. unknown labels) and structural priors (i.e. low-rank structures) into consideration. To cope with the non-convexity issues arising from unlabelled heterogeneous data and low-complexity structure, we develop a three-stage meta-algorithm that is guaranteed to recover the unknown matrices with near-optimal sample and computational complexities under Gaussian designs. In addition, the proposed algorithm is provably stable against random noise. We complement the theoretical studies with empirical evidence that confirms the efficacy of our algorithm.
Yanxi Chen 0001, Cong Ma 0001, H. Vincent Poor, Yuxin Chen 0002
IEEE Trans. Inf. Theory1
2018 Change-Point Detection of Gaussian Graph Signals with Partial Information
abstract
In a change-point detection problem, a sequence of signals switches from one distribution to another at an unknown time step, and the goal is to quickly and reliably detect this change. By providing new insight into signal processing and data analysis, graph signal processing promises various applications including image processing and sensor network analysis, and becomes an emerging field of research. In this work, we formulate the problem of change-point detection on graph. Under the reasonable assumption of normality, we propose a CUSUM-based algorithm for change-point detection with an arbitrary, unknown and perhaps time-varying mean shift after the change-point. We further propose a decentralized, distributed algorithm, which requires no fusion center, to reduce computational complexity, as well as costs and delays of communication. Numerical results on both synthetic and real-world data demonstrate that our algorithms are efficient and accurate.
Yanxi Chen 0001, Xianghui Mao, Dan Ling, Yuantao Gu
ICASSP1
2018 Active Orthogonal Matching Pursuit for Sparse Subspace Clustering
abstract
Sparse subspace clustering (SSC) is a state-of-the-art method for clustering high-dimensional data points lying in a union of low-dimensional subspaces. However, while ℓ1optimization-based SSC algorithms suffer from high computational complexity, other variants of SSC, such as orthogonal-matching-pursuit-based SSC (OMP-SSC), lose clustering accuracy in pursuit of improving time efficiency. In this letter, we propose a novel active OMP-SSC, which improves clustering accuracy of OMP-SSC by adaptively updating data points and randomly dropping data points in the OMP process, while still enjoying the low computational complexity of greedy pursuit algorithms. We provide heuristic analysis of our approach and explain how these two active steps achieve a better tradeoff between connectivity and separation. Numerical results on both synthetic data and real-world data validate our analyses and show the advantages of the proposed active algorithm.
Yanxi Chen 0001, Gen Li 0005, Yuantao Gu
IEEE Signal Process. Lett.1