Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zhaoxian Wu

dblp:255/6466 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0001-6724-6134ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 31% Deep learning architectures and training · 24% Learning theory · 24%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 88% Hardware accelerators and domain-specific architectures · 12%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › processing-in-memory › computing-in-memory
analog in-memory computing
1.622025
Analog In-memory Training on General Non-ideal Resistive Elements: The Impact of Response Functions · NeurIPS 2025
Towards Exact Gradient-based Training on Analog In-memory Computing · NeurIPS 2024
Machine learning › Efficient and distributed learning › distributed training
distributed stochastic gradient descent
0.912025
On the Trade-Off Between Flatness and Optimization in Distributed Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Learning theory › generalization
generalization analysis
0.912025
On the Trade-Off Between Flatness and Optimization in Distributed Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Deep learning architectures and training › training optimization
gradient-based training
0.912025
Analog In-memory Training on General Non-ideal Resistive Elements: The Impact of Response Functions · NeurIPS 2025
Machine learning › Optimization for machine learning
stochastic gradient descent
0.812024
Towards Exact Gradient-based Training on Analog In-memory Computing · NeurIPS 2024
Machine learning › Efficient and distributed learning › energy-efficient learning
energy-efficient training
0.312025
Analog In-memory Training on General Non-ideal Resistive Elements: The Impact of Response Functions · NeurIPS 2025
Hardware accelerators and domain-specific architectures
analog computing accelerator
0.212024
Towards Exact Gradient-based Training on Analog In-memory Computing · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

stochastic gradient descent · 2.4response function analysis · 1.7residual learning · 1.7tiki-taka · 1.5convergence analysis · 1.5diffusion strategy · 0.9consensus strategy · 0.9bilevel optimization · 0.9bi-level optimization · 0.9
YearPublicationVenuePosition
2025 Analog In-memory Training on General Non-ideal Resistive Elements: The Impact of Response Functions
abstract
As the economic and environmental costs of training and deploying large vision or language models increase dramatically, analog in-memory computing (AIMC) emerges as a promising energy-efficient solution. However, the training perspective, especially its training dynamic, is underexplored. In AIMC hardware, the trainable weights are represented by the conductance of resistive elements and updated using consecutive electrical pulses. While the conductance changes by a constant in response to each pulse, in reality, the change is scaled by asymmetric and non-linear response functions, leading to a non-ideal training dynamic. This paper provides a theoretical foundation for gradient-based training on AIMC hardware with non-ideal response functions. We demonstrate that asymmetric response functions negatively impact Analog SGD by imposing an implicit penalty on the objective. To overcome the issue, we propose residual learning algorithm, which provably converges exactly to a critical point by solving a bilevel optimization problem. We show that the proposed method can be extended to deal with other hardware imperfections like limited response granularity. As far as we know, it is the first paper to investigate the impact of a class of generic non-ideal response functions. The conclusion is supported by simulations validating our theoretical insights.
Zhaoxian Wu, Quan Xiao, Tayfun Gokmen, Omobayode Fagbohungbe, Tianyi Chen 0002
NeurIPS1
2025 On the Trade-Off Between Flatness and Optimization in Distributed Learning
abstract
This paper proposes a theoretical framework to evaluate and compare the performance of stochastic gradient algorithms for distributed learning in relation to their behavior around local minima in nonconvex environments. Previous works have noticed that convergence toward flat local minima tend to enhance the generalization ability of learning algorithms. This work discovers three interesting results. First, it shows that decentralized learning strategies are able to escape faster away from local minima and favor convergence toward flatter minima relative to the centralized solution. Second, in decentralized methods, the consensus strategy has a worse excess-risk performance than diffusion, giving it a better chance of escaping from local minima and favoring flatter minima. Third, and importantly, the ultimate classification accuracy is not solely dependent on the flatness of the local minimum but also on how well a learning algorithm can approach that minimum. In other words, the classification accuracy is a function of both flatness and optimization performance. In this regard, since diffusion has a lower excess-risk than consensus, when both algorithms are trained starting from random initial points, diffusion enhances the classification accuracy. The paper examines the interplay between the two measures of flatness and optimization error closely. One important conclusion is that decentralized strategies deliver in general enhanced classification accuracy because they strike a more favorable balance between flatness and optimization performance compared to the centralized solution.
Zhaoxian Wu, Kun Yuan 0001, Ali H. Sayed
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 On the Convergence of Single-Timescale Multi-Sequence Stochastic Approximation Without Fixed Point Smoothness
abstract
Stochastic approximation (SA) that involves multiple coupled sequences has diverse applications, including but not limited to bilevel optimization, meta learning and reinforcement learning. Unfortunately, the existing multi-timescale analysis of multiple-sequence SA (MSSA) implies a slow convergence rate, whereas the single-timescale analysis relies on assuming smoothness of fixed points. In this paper, we present tighter single-timescale analysis for MSSA, without assuming smoothness of fixed points. Our theoretical results demonstrate that, when all involved operators are strongly monotone, MSSA converges at a rate of $\tilde {\mathcal{O}}\left( {{K^{ - 1}}} \right)$, where K is the total number of iterations. Under a weaker assumption that all involved operators are strongly monotone except for$O\left( {{K^{ - \frac{1}{2}}}} \right)$ the main one, MSSA converges at a rate of . These theoretical results align with those established in single-sequence SA (SSSA). Applying these theoretical results to bilevel optimization offers relaxed assumptions and/or simpler algorithms with performance guarantees, as validated by numerical experiments.
Zhaoxian Wu, Qing Ling 0001
ICASSP2
2024 Towards Exact Gradient-based Training on Analog In-memory Computing
abstract
Given the high economic and environmental costs of using large vision or language models, analog in-memory accelerators present a promising solution for energy-efficient AI. While inference on analog accelerators has been studied recently, the training perspective is underexplored. Recent studies have shown that the "workhorse" of digital AI training - stochastic gradient descent (SGD) algorithm converges inexactly when applied to model training on non-ideal devices. This paper puts forth a theoretical foundation for gradient-based training on analog devices. We begin by characterizing the non-convergent issue of SGD, which is caused by the asymmetric updates on the analog devices. We then provide a lower bound of the asymptotic error to show that there is a fundamental performance limit of SGD-based analog training rather than an artifact of our analysis. To address this issue, we study a heuristic analog algorithm called Tiki-Taka that has recently exhibited superior empirical performance compared to SGD. We rigorously show its ability to converge to a critical point exactly and hence eliminate the asymptotic error. The simulations verify the correctness of the analyses.
Zhaoxian Wu, Tayfun Gokmen, Malte J. Rasch, Tianyi Chen 0002
NeurIPS1
2023 Distributed Online Learning With Adversarial Participants In An Adversarial Environment
abstract
This paper studies distributed online learning under Byzantine attacks. The performance of an online learning algorithm is characterized by (adversarial) regret, and a sublinear bound is preferred. But we prove that, even with a class of state-of-the-art robust aggregation rules, in an adversarial environment and with Byzantine participants, distributed online gradient descent can only achieve a linear adversarial regret bound, which is tight. This is the inevitable consequence of Byzantine attacks, even though we can control the constant of the linear adversarial regret to a reasonable level. Interestingly, when the environment is not fully adversarial so that the losses of the honest participants are i.i.d. (independent and identically distributed), we show that sublinear stochastic regret, in contrast to the aforementioned adversarial regret, is possible. We develop a Byzantine-robust distributed online gradient descent algorithm with momentum to attain such a sublinear stochastic regret bound.
Xingrong Dong, Zhaoxian Wu, Qing Ling 0001, Zhi Tian
ICASSP2
2022 A Byzantine-Resilient Dual Subgradient Method for Vertical Federated Learning
abstract
Federated learning (FL) raises new challenges on security risks, especially when the FL system involves Byzantine clients that send corrupted or adversarial messages to the central server for deteriorating the training paradigm. While there is an extensive research on robust algorithms for horizontal or data-partitioned FL problems, the exploration in Byzantine-resilient vertical or feature-partitioned FL is quite limited. In this paper, we provide a problem formulation of vertical FL in the presence of Byzantine attacks, and propose a Byzantine-resilient dual subgradient method. Convergence analysis is established, and the influence of the Byzantine clients is also clarified. Numerical experiments show the proposed algorithm is robust to various Byzantine attacks on vertical FL.
Kun Yuan 0001, Zhaoxian Wu, Qing Ling 0001
ICASSP2
2022 Byzantine-robust variance-reduced federated learning over distributed non-i.i.d. data
Zhaoxian Wu, Qing Ling 0001, Tianyi Chen 0002
Inf. Sci.2
2022 Communication-Censored Distributed Stochastic Gradient Descent
abstract
This article develops a communication-efficient algorithm to solve the stochastic optimization problem defined over a distributed network, aiming at reducing the burdensome communication in applications, such as distributed machine learning. Different from the existing works based on quantization and sparsification, we introduce a communication-censoring technique to reduce the transmissions of variables, which leads to our communication-censored distributed stochastic gradient descent (CSGD) algorithm. Specifically, in CSGD, the latest minibatch stochastic gradient at a worker will be transmitted to the server if and only if it is sufficiently informative. When the latest gradient is not available, the stale one will be reused at the server. To implement this communication-censoring strategy, the batch size is increasing in order to alleviate the effect of stochastic gradient noise. Theoretically, CSGD enjoys the same order of convergence rate as that of SGD but effectively reduces communication. Numerical experiments demonstrate the sizable communication saving of CSGD.
Zhaoxian Wu, Tianyi Chen 0002, Liping Li 0004, Qing Ling 0001
IEEE Trans. Neural Networks Learn. Syst.2
2021 Byzantine-Resilient Decentralized TD Learning with Linear Function Approximation
abstract
This paper considers the policy evaluation problem in reinforcement learning with agents of a decentralized and directed network. The focus is on decentralized temporal-difference (TD) learning with linear function approximation in the presence of unreliable or even malicious agents, termed as Byzantine agents. In order to evaluate the quality of a fixed policy in a common environment, agents usually run decentralized TD(λ) collaboratively. However, when some Byzantine agents behave adversarially, decentralized TD(λ) is unable to learn an accurate linear approximation for the true value function. We propose a trimmed-mean based decentralized TD(λ) algorithm to perform policy evaluation in this setting. We establish the finite-time convergence rate, as well as the asymptotic learning error that depends on the number of Byzantine agents. Numerical experiments corroborate the robustness of the proposed algorithm.
Zhaoxian Wu, Tianyi Chen 0002, Qing Ling 0001
ICASSP1
2020 Resilient to Byzantine Attacks Finite-Sum Optimization Over Networks
abstract
This contribution deals with distributed finite-sum optimization for learning over networks in the presence of malicious Byzantine attacks. To cope with such attacks, resilient approaches so far combine stochastic gradient descent (SGD) with different robust aggregation rules. However, the sizeable SGD-induced gradient noise makes it challenging to distinguish malicious messages sent by the Byzantine attackers from noisy stochastic gradients sent by the friendly workers. This motivates gradient noise reduction as a means of robustifying SGD in the presence of Byzantine attacks. To this end, the present work puts forth a Byzantine attack resilient distributed (Byrd-) SAGA approach for learning tasks involving finite-sum optimization over networks. Rather than the mean employed by distributed SAGA, the novel Byrd-SAGA relies on the geometric median to aggregate the corrected stochastic gradients sent by the workers. When less than half of the workers are Byzantine attackers, the robustness of geometric median to outliers enables Byrd-SAGA to achieve provable linear convergence to a neighborhood of the optimal solution, where the size of neighborhood is determined by the number of Byzantine workers. Numerical tests demonstrate the robustness of Byrd-SAGA to various Byzantine attacks, as well as the merits of Byrd-SAGA over Byzantine-resilient SGD.
Zhaoxian Wu, Qing Ling 0001, Tianyi Chen 0002, Georgios B. Giannakis
ICASSP1