Yinuo Ren

dblp:281/6971 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
8since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Generative modeling · 62% Learning theory · 9% Deep learning architectures and training · 9%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model
discrete diffusion model
1.722025
Fast Solvers for Discrete Diffusion Models: Theory and Applications of High-Order Algorithms · NeurIPS 2025
How Discrete and Continuous Diffusion Meet: Comprehensive Analysis of Discrete Diffusion Models via a Stochastic Integral Framework · ICLR 2025
Machine learning › Generative modeling
diffusion model
1.622025
How Discrete and Continuous Diffusion Meet: Comprehensive Analysis of Discrete Diffusion Models via a Stochastic Integral Framework · ICLR 2025
Accelerating Diffusion Models with Parallel Sampling: Inference at Sub-Linear Time Complexity · NeurIPS 2024
Machine learning › Generative modeling › diffusion model
diffusion model inference
0.912025
Fast Solvers for Discrete Diffusion Models: Theory and Applications of High-Order Algorithms · NeurIPS 2025
Machine learning › Optimization for machine learning
convergence analysis
0.812024
Accelerating Diffusion Models with Parallel Sampling: Inference at Sub-Linear Time Complexity · NeurIPS 2024
Machine learning › Deep learning architectures and training
neural network estimator
0.812024
Statistical Spatially Inhomogeneous Diffusion Inference · AAAI 2024
Machine learning › Learning theory › statistical estimation
nonparametric estimation
0.812024
Statistical Spatially Inhomogeneous Diffusion Inference · AAAI 2024
Machine learning › Generative modeling › diffusion model
parallel sampling
0.812024
Accelerating Diffusion Models with Parallel Sampling: Inference at Sub-Linear Time Complexity · NeurIPS 2024
Machine learning › Reinforcement learning
sample efficiency
0.812024
Accelerating Diffusion Models with Parallel Sampling: Inference at Sub-Linear Time Complexity · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

tau-leaping · 2.6levy-type stochastic integrals · 1.7girsanov theorem · 1.7neural network · 1.5minimax optimal rate analysis · 1.5theta-trapezoidal method · 0.9high-order numerical scheme · 0.9probability flow ODE · 0.8picard iteration · 0.8girsanov's theorem · 0.8
YearPublicationVenuePosition
2025 How Discrete and Continuous Diffusion Meet: Comprehensive Analysis of Discrete Diffusion Models via a Stochastic Integral Framework
abstract
Discrete diffusion models have gained increasing attention for their ability to model complex distributions with tractable sampling and inference. However, the error analysis for discrete diffusion models remains less well-understood. In this work, we propose a comprehensive framework for the error analysis of discrete diffusion models based on Lévy-type stochastic integrals. By generalizing the Poisson random measure to that with a time-independent and state-dependent intensity, we rigorously establish a stochastic integral formulation of discrete diffusion models and provide the corresponding change of measure theorems that are intriguingly analogous to Itô integrals and Girsanov's theorem for their continuous counterparts. Our framework unifies and strengthens the current theoretical results on discrete diffusion models and obtains the first error bound for the $\tau$-leaping scheme in KL divergence. With error sources clearly identified, our analysis gives new insight into the mathematical properties of discrete diffusion models and offers guidance for the design of efficient and accurate algorithms for real-world discrete diffusion model applications.
Yinuo Ren, Haoxuan Chen, Grant M. Rotskoff, Lexing Ying
ICLR1
2025 Fast Solvers for Discrete Diffusion Models: Theory and Applications of High-Order Algorithms
abstract
Discrete diffusion models have emerged as a powerful generative modeling framework for discrete data with successful applications spanning from text generation to image synthesis. However, their deployment faces challenges due to the high dimensionality of the state space, necessitating the development of efficient inference algorithms. Current inference approaches mainly fall into two categories: exact simulation and approximate methods such as $\tau$-leaping. While exact methods suffer from unpredictable inference time and redundant function evaluations, $\tau$-leaping is limited by its first-order accuracy. In this work, we advance the latter category by tailoring the first extension of high-order numerical inference schemes to discrete diffusion models, enabling larger step sizes while reducing error. We rigorously analyze the proposed schemes and establish the second-order accuracy of the $\theta$-Trapezoidal method in KL divergence. Empirical evaluations on GSM8K-level math-reasoning, GPT-2-level text, and ImageNet-level image generation tasks demonstrate that our method achieves superior sample quality compared to existing approaches under equivalent computational constraints, with consistent performance gains across models ranging from 200M to 8B. Our code is available at https://github.com/yuchen-zhu-zyc/DiscreteFastSolver
Yinuo Ren, Haoxuan Chen, Grant M. Rotskoff, Molei Tao, Lexing Ying
NeurIPS1
2025 COS-DPO: Conditioned One-Shot Multi-Objective Fine-Tuning Framework
abstract
In LLM alignment and many other ML applications, one often faces the *Multi-Objective Fine-Tuning* (MOFT) problem, *i.e.*, fine-tuning an existing model with datasets labeled w.r.t. different objectives simultaneously. To address the challenge, we propose a *Conditioned One-Shot* fine-tuning framework (COS-DPO) that extends the Direct Preference Optimization technique, originally developed for efficient LLM alignment with preference data, to accommodate the MOFT settings. By direct conditioning on the weight across auxiliary objectives, our Weight-COS-DPO method enjoys an efficient one-shot training process for profiling the Pareto front and is capable of achieving comprehensive trade-off solutions even in the post-training stage. Based on our theoretical findings on the linear transformation properties of the loss function, we further propose the Temperature-COS-DPO method that augments the temperature parameter to the model input, enhancing the flexibility of post-training control over the trade-offs between the main and auxiliary objectives. We demonstrate the effectiveness and efficiency of the COS-DPO framework through its applications to various tasks, including the Learning-to-Rank (LTR) and LLM alignment tasks, highlighting its viability for large-scale ML deployments.
Yinuo Ren, Tesi Xiao, Michael Shavlovsky, Lexing Ying, Holakou Rahmanian
UAI1
2025 Optimal starting point for time series forecasting
Yinuo Ren, Guangyao Cao, Feng Li 0028, Haobo Qi
Expert Syst. Appl.2
2024 Statistical Spatially Inhomogeneous Diffusion Inference
abstract
Inferring a diffusion equation from discretely observed measurements is a statistical challenge of significant importance in a variety of fields, from single-molecule tracking in biophysical systems to modeling financial instruments. Assuming that the underlying dynamical process obeys a d-dimensional stochastic differential equation of the form dx_t = b(x_t)dt + \Sigma(x_t)dw_t, we propose neural network-based estimators of both the drift b and the spatially-inhomogeneous diffusion tensor D = \Sigma\Sigma^T/2 and provide statistical convergence guarantees when b and D are s-Hölder continuous. Notably, our bound aligns with the minimax optimal rate N^{-\frac{2s}{2s+d}} for nonparametric function estimation even in the presence of correlation within observational data, which necessitates careful handling when establishing fast-rate generalization bounds. Our theoretical results are bolstered by numerical experiments demonstrating accurate inference of spatially-inhomogeneous diffusion tensors.
Yinuo Ren, Yiping Lu 0001, Lexing Ying, Grant M. Rotskoff
AAAI1
2024 Understanding the Generalization Benefits of Late Learning Rate Decay
abstract
Why do neural networks trained with large learning rates for longer time often lead to better generalization? In this paper, we delve into this question by examining the relation between training and testing loss in neural networks. Through visualization of these losses, we note that the training trajectory with a large learning rate navigates through the minima manifold of the training loss, finally nearing the neighborhood of the testing loss minimum. Motivated by these findings, we introduce a nonlinear model whose loss landscapes mirror those observed for real neural networks. Upon investigating the training process using SGD on our model, we demonstrate that an extended phase with a large learning rate steers our model towards the minimum norm solution of the training loss, which may achieve near-optimal generalization, thereby affirming the empirically observed benefits of late learning rate decay.
Yinuo Ren, Chao Ma 0012, Lexing Ying
AISTATS1
2024 Multi-objective Optimization via Wasserstein-Fisher-Rao Gradient Flow
abstract
Multi-objective optimization (MOO) aims to optimize multiple, possibly conflicting objectives with widespread applications. We introduce a novel interacting particle method for MOO inspired by molecular dynamics simulations. Our approach combines overdamped Langevin and birth-death dynamics, incorporating a “dominance potential” to steer particles toward global Pareto optimality. In contrast to previous methods, our method is able to relocate dominated particles, making it particularly adept at managing Pareto fronts of complicated geometries. Our method is also theoretically grounded as a Wasserstein-Fisher-Rao gradient flow with convergence guarantees. Extensive experiments confirm that our approach outperforms state-of-the-art methods on challenging synthetic and real-world datasets.
Yinuo Ren, Tesi Xiao, Tanmay Gangwani, Anshuka Rangi, Holakou Rahmanian, Lexing Ying, Subhajit Sanyal
AISTATS1
2024 Accelerating Diffusion Models with Parallel Sampling: Inference at Sub-Linear Time Complexity
abstract
Diffusion models have become a leading method for generative modeling of both image and scientific data. As these models are costly to train and \emph{evaluate}, reducing the inference cost for diffusion models remains a major goal. Inspired by the recent empirical success in accelerating diffusion models via the parallel sampling technique~\cite{shih2024parallel}, we propose to divide the sampling process into $\mathcal{O}(1)$ blocks with parallelizable Picard iterations within each block. Rigorous theoretical analysis reveals that our algorithm achieves $\widetilde{\mathcal{O}}(\mathrm{poly} \log d)$ overall time complexity, marking \emph{the first implementation with provable sub-linear complexity w.r.t. the data dimension $d$}. Our analysis is based on a generalized version of Girsanov's theorem and is compatible with both the SDE and probability flow ODE implementations. Our results shed light on the potential of fast and efficient sampling of high-dimensional data on fast-evolving modern large-memory GPU clusters.
Haoxuan Chen, Yinuo Ren, Lexing Ying, Grant M. Rotskoff
NeurIPS2