Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Siqiao Mu

dblp:360/1210 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 41% Trustworthy machine learning · 32% Optimization for machine learning · 27%
Network and information security
1 paper
Privacy and data protection · 50% Security and privacy of machine learning · 50%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › machine unlearning
certified unlearning
0.912025
Rewind-to-Delete: Certified Machine Unlearning for Nonconvex Functions · NeurIPS 2025
Machine learning › Trustworthy machine learning
machine unlearning
0.912025
Rewind-to-Delete: Certified Machine Unlearning for Nonconvex Functions · NeurIPS 2025
Security and privacy of machine learning › machine unlearning
certified unlearning
0.912025
Rewind-to-Delete: Certified Machine Unlearning for Nonconvex Functions · NeurIPS 2025
Privacy and data protection
differential privacy
0.912025
Rewind-to-Delete: Certified Machine Unlearning for Nonconvex Functions · NeurIPS 2025
Machine learning › Reinforcement learning
actor-critic methods
0.812024
On the Second-Order Convergence of Biased Policy Gradient Algorithms · ICML 2024
Machine learning › Optimization for machine learning
convergence analysis
0.812024
On the Second-Order Convergence of Biased Policy Gradient Algorithms · ICML 2024
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.812024
On the Second-Order Convergence of Biased Policy Gradient Algorithms · ICML 2024
Machine learning › Optimization for machine learning › non-convex optimization
second-order stationary point
0.812024
On the Second-Order Convergence of Biased Policy Gradient Algorithms · ICML 2024
Machine learning › Reinforcement learning
value-based reinforcement learning
0.812024
On the Second-Order Convergence of Biased Policy Gradient Algorithms · ICML 2024

Methods — techniques the papers use, named apart from their topics

gradient descent · 1.7differential privacy · 1.7saddle point analysis · 0.8monte carlo sampling · 0.8biased gradient estimation · 0.8
YearPublicationVenuePosition
2025 Rewind-to-Delete: Certified Machine Unlearning for Nonconvex Functions
abstract
Machine unlearning algorithms aim to efficiently remove data from a model without retraining it from scratch, in order to remove corrupted or outdated data or respect a user's "right to be forgotten." Certified machine unlearning is a strong theoretical guarantee based on differential privacy that quantifies the extent to which an algorithm erases data from the model weights. In contrast to existing works in certified unlearning for convex or strongly convex loss functions, or nonconvex objectives with limiting assumptions, we propose the first, first-order, black-box (i.e., can be applied to models pretrained with vanilla gradient descent) algorithm for unlearning on general nonconvex loss functions, which unlearns by ``rewinding" to an earlier step during the learning process before performing gradient descent on the loss function of the retained data points. We prove $(\epsilon, \delta)$ certified unlearning and performance guarantees that establish the privacy-utility-complexity tradeoff of our algorithm, and we prove generalization guarantees for nonconvex functions that satisfy the Polyak-Lojasiewicz inequality. Finally, we demonstrate the superior performance of our algorithm compared to existing methods, within a new experimental framework that more accurately reflects unlearning user data in practice.
Siqiao Mu, Diego Klabjan
NeurIPS1
2024 On the Second-Order Convergence of Biased Policy Gradient Algorithms
abstract
Since the objective functions of reinforcement learning problems are typically highly nonconvex, it is desirable that policy gradient, the most popular algorithm, escapes saddle points and arrives at second-order stationary points. Existing results only consider vanilla policy gradient algorithms with unbiased gradient estimators, but practical implementations under the infinite-horizon discounted reward setting are biased due to finite-horizon sampling. Moreover, actor-critic methods, whose second-order convergence has not yet been established, are also biased due to the critic approximation of the value function. We provide a novel second-order analysis of biased policy gradient methods, including the vanilla gradient estimator computed from Monte-Carlo sampling of trajectories as well as the double-loop actor-critic algorithm, where in the inner loop the critic improves the approximation of the value function via TD(0) learning. Separately, we also establish the convergence of TD(0) on Markov chains irrespective of initial state distribution.
Siqiao Mu, Diego Klabjan
ICML1