VLDB 2026 Research / reviewers in the wild / expert
Kun Yuan 0001
dblp:74/4607-1
· DBLP profile ↗
46ranked-venue papers
12as first author
33since 2021 · last 2026
0000-0001-8394-8187ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 7 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gradient Normalization Enables Communication-Efficient Distributed Learning Under Initialization Data HeterogeneityabstractCommunication-efficient distributed learning has achieved remarkable progress in training large-scale deep neural networks across numerous clients. However, in heterogeneous environments-where local data distributions vary significantly-the empirical performance of many such algorithms degrades sharply. Moreover, existing theoretical analyses often depend on overly restrictive assumptions, such as bounded data heterogeneity, which may not be valid even for a single client. In this work, we introduce a general gradient normalization strategy that can be seamlessly integrated into a wide range of distributed learning algorithms, including compressed distributed stochastic gradient descent, federated averaging, and asynchronous variants. Our theoretical analysis demonstrates that this normalization technique effectively mitigates the negative impact of data heterogeneity, allowing these algorithms to achieve linear speedup rates with only requiring the boundedness of initialization data heterogeneity. Extensive numerical experiments further confirm the practical effectiveness and theoretical guarantees of our approach. Tao Sun 0005, Baihao Wu, Xinwang Liu 0002, Kun Yuan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Model-Free Test Time Adaptation for Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection is essential for the reliability of ML models. Most existing methods for OOD detection learn a fixed decision criterion from a given in-distribution dataset and apply it universally to decide if a data point is OOD. Recent work Fang et al. (2022) shows that given only in-distribution data, it is impossible to reliably detect OOD data without extra assumptions. Motivated by the theoretical result and recent exploration of test-time adaptation methods, we propose a Non-Parametric Test Time Adaptation framework for Out-Of-Distribution Detection (AdaODD). Unlike conventional methods, AdaODD utilizes online test samples for model adaptation during testing, enhancing adaptability to changing data distributions. The framework incorporates detected OOD instances into decision-making, reducing false positive rates, particularly when ID and OOD distributions overlap significantly. We demonstrate the effectiveness of AdaODD through comprehensive experiments on multiple OOD detection benchmarks, extensive empirical studies show that AdaODD significantly improves the performance of OOD detection over state-of-the-art methods. Specifically, AdaODD reduces the false positive rate (FPR95) by 23.23% on the CIFAR-10 benchmarks and 38% on the ImageNet-1 k benchmarks compared to the advanced methods. Lastly, we theoretically verify the effectiveness of AdaODD. Yifan Zhang 0004, Xue Wang 0010, Tian Zhou 0004, Kun Yuan 0001, Zhang Zhang 0001, Liang Wang 0001, Rong Jin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Greedy Low-Rank Gradient Compression Provably Converges for Distributed LearningabstractDistributed optimization is pivotal for large-scale signal processing and machine learning, yet communication overhead remains a major bottleneck. Low-rank gradient com-pression, in which the transmitted gradients are approximated by low-rank matrices to reduce communication, offers a promising remedy. Existing methods typically adopt either randomized or greedy compression strategies: randomized approaches project gradients onto randomly chosen subspaces, introducing high variance and degrading empirical performance; greedy methods select the most informative subspaces, achieving strong empirical results but lacking convergence guarantees. To address this gap, we propose GreedyLore-the first Greedy Low-Rank gradient compression algorithm for distributedlearning with rigorous convergence guarantees. GreedyLore incorporates error feedback to correct the bias introduced by greedy compression and introduces a semi-lazy subspace update that ensures the com-pression operator remains contractive throughout all iterations. With these techniques, we prove that GreedyLore achieves a convergence rate of$\mathcal{O}(\sigma/\sqrt{NT}^{-}+1/T)$under standard optimizers such as Adam-marking the first linear speedup convergence rate for low-rank gradient compression. Extensive experiments are conducted to validate our theoretical findings. Chuyan Chen, Pengrui Li, Weichen Jia, Yanjie Dong 0003, Kun Yuan 0001 |
CloudCom | 6 |
| 2025 | An All-Reduce Compatible Top-$K$ Compressor for Communication-Efficient Distributed LearningabstractCommunication remains a central bottleneck in large-scale distributed machine learning, and gradient sparsi-fication has emerged as a promising strategy to alleviate this challenge. However, existing gradient compressors face notable limitations: Rand-K discards structural information and per-forms poorly in practice, while Top-$K$preserves informative entries but loses the contraction property and requires costly All-Gather operations. In this paper, we propose ARC-Top-K, an All-Reduce-Compatible Top-K compressor that aligns sparsity patterns across nodes using a lightweight sketch of the gradient, enabling index-free All-Reduce while preserving globally significant information. ARC-Top-$K$is provably con-tractive and, when combined with momentum error feedback (EF21M), achieves linear speedup and sharper convergence rates than the original EF21M under standard assumptions. Empirically, ARC-Top-K matches the accuracy of Top-K while reducing wall-clock training time by up to 60.7%, offering an efficient and scalable solution that combines the robustness of Rand-K with the strong performance of Top-$K$. Chuyan Chen, Zhangxin Li, Yanjie Dong 0003, Kun Yuan 0001 |
CloudCom | 6 |
| 2025 | Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank StructuresabstractParameter-efficient fine-tuning (PEFT) significantly reduces memory costs when adapting large language models (LLMs) for downstream applications. However, traditional first-order (FO) fine-tuning algorithms incur substantial memory overhead due to the need to store activation values for back-propagation during gradient computation, particularly in long-context fine-tuning tasks. Zeroth-order (ZO) algorithms offer a promising alternative by approximating gradients using finite differences of function values, thus eliminating the need for activation storage. Nevertheless, existing ZO methods struggle to capture the low-rank gradient structure common in LLM fine-tuning, leading to suboptimal performance. This paper proposes a low-rank ZO gradient estimator and introduces a novel **lo**w-rank **ZO** algorithm (LOZO) that effectively captures this structure in LLMs. We provide convergence guarantees for LOZO by framing it as a subspace optimization method. Additionally, its low-rank nature enables LOZO to integrate with momentum techniques while incurring negligible extra memory costs. Extensive experiments across various model sizes and downstream tasks demonstrate that LOZO and its momentum-based variant outperform existing ZO methods and closely approach the performance of FO algorithms. Yiming Chen 0003, Kun Yuan 0001, Zaiwen Wen |
ICLR | 4 |
| 2025 | Subspace Optimization for Large Language Models with Convergence GuaranteesabstractSubspace optimization algorithms, such as GaLore (Zhao et al., 2024), have gained attention for pre-training and fine-tuning large language models (LLMs) due to their memory efficiency. However, their convergence guarantees remain unclear, particularly in stochastic settings. In this paper, we reveal that GaLore does not always converge to the optimal solution and provide an explicit counterexample to support this finding. We further explore the conditions under which GaLore achieves convergence, showing that it does so when either (i) a sufficiently large mini-batch size is used or (ii) the gradient noise is isotropic. More significantly, we introduce GoLore (Gradient random Low-rank projection), a novel variant of GaLore that provably converges in typical stochastic settings, even with standard batch sizes. Our convergence analysis extends naturally to other subspace optimization algorithms. Finally, we empirically validate our theoretical results and thoroughly test the proposed mechanisms. Codes are available at https://github.com/pkumelon/Golore. Pengrui Li, Yipeng Hu, Chuyan Chen, Kun Yuan 0001 |
ICML | 5 |
| 2025 | MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant OptimizationabstractAs distributed optimization scales to meet the demands of Large Language Model (LLM) training, hardware failures become increasingly non-negligible. Existing fault-tolerant training methods often introduce significant computational or memory overhead, demanding additional resources. To address this challenge, we propose **Me**mory- and **C**omputation- **e**fficient **F**ault-tolerant **O**ptimization (**MeCeFO**), a novel algorithm that ensures robust training with minimal overhead. When a computing node fails, MeCeFO seamlessly transfers its training task to a neighboring node while employing memory- and computation-efficient algorithmic optimizations to minimize the extra workload imposed on the neighboring node handling both tasks. MeCeFO leverages three key algorithmic designs: (i) Skip-connection, which drops the multi-head attention (MHA) module during backpropagation for memory- and computation-efficient approximation; (ii) Recomputation, which reduces activation memory in feedforward networks (FFNs); and (iii) Low-rank gradient approximation, enabling efficient estimation of FFN weight matrix gradients. Theoretically, MeCeFO matches the convergence rate of conventional distributed training, with a rate of $\mathcal{O}(1/\sqrt{nT})$, where $n$ is the data parallelism size and $T$ is the number of iterations. Empirically, MeCeFO maintains robust performance under high failure rates, incurring only a 4.18\% drop in throughput, demonstrating $5.0\times$ to $6.7\times$ greater resilience than previous SOTA approaches. Rizhen Hu, Mou Sun, Binhang Yuan, Kun Yuan 0001 |
NeurIPS | 6 |
| 2025 | MISA: Memory-Efficient LLMs Optimization with Module-wise Importance SamplingabstractThe substantial memory demands of pre-training and fine-tuning large language models (LLMs) require memory-efficient optimization algorithms. One promising approach is layer-wise optimization, which treats each transformer block as a single layer and optimizes it sequentially, while freezing the other layers to save optimizer states and activations. Although effective, these methods ignore the varying importance of the modules within each layer, leading to suboptimal performance. Moreover, layer-wise sampling provides only limited memory savings, as at least one full layer must remain active during optimization. To overcome these limitations, we propose **M**odule-wise **I**mportance **SA**mpling (**MISA**), a novel method that divides each layer into smaller modules and assigns importance scores to each module.
MISA uses a weighted random sampling mechanism to activate modules, provably reducing
gradient variance compared to layer-wise sampling.
Additionally, we establish an $\mathcal{O}(1/\sqrt{K})$ convergence rate under non-convex and stochastic conditions, where $K$ is the total number of training steps, and provide a detailed memory analysis showcasing MISA's superiority over existing baseline methods. Experiments on diverse learning tasks validate the effectiveness of MISA. Renjia Deng, Xue Wang 0010, Kun Yuan 0001 |
NeurIPS | 6 |
| 2025 | Revisiting Gradient Normalization and Clipping for Nonconvex SGD under Heavy-Tailed Noise: Necessity, Sufficiency, and AccelerationabstractGradient clipping has long been considered essential for ensuring the convergence of Stochastic Gradient Descent (SGD) in the presence of heavy-tailed gradient noise. In this paper, we revisit this belief and explore whether gradient normalization can serve as an effective alternative or complement. We prove that, under individual smoothness assumptions, gradient normalization alone is sufficient to guarantee convergence of the nonconvex SGD. Moreover, when combined with clipping, it yields far better rates of convergence under more challenging noise distributions. We provide a unifying theory describing normalization-only, clipping-only, and combined approaches. Moving forward, we investigate existing variance-reduced algorithms, establishing that, in such a setting, normalization alone is sufficient for convergence. Finally, we present an accelerated variant that under second-order smoothness improves convergence. Our results provide theoretical insights and practical guidance for using normalization and clipping in nonconvex optimization with heavy-tailed noise. Tao Sun 0005, Xinwang Liu 0002, Kun Yuan 0001 |
J. Mach. Learn. Res. | 3 |
| 2025 | Decentralized Bilevel Optimization: A Perspective from Transient Iteration ComplexityabstractStochastic bilevel optimization (SBO) is becoming increasingly essential in machine learning due to its versatility in handling nested structures. To address large-scale SBO, decentralized approaches have emerged as effective paradigms in which nodes communicate with immediate neighbors without a central server, thereby improving communication efficiency and enhancing algorithmic robustness. However, most decentralized SBO algorithms focus solely on asymptotic convergence rates, overlooking transient iteration complexity-the number of iterations required before asymptotic rates dominate, which results in limited understanding of the influence of network topology, data heterogeneity, and the nested bilevel algorithmic structures. To address this issue, this paper introduces D-SOBA, a Decentralized Stochastic One-loop Bilevel Algorithm framework. D-SOBA comprises two variants: D-SOBA-SO, which incorporates second-order Hessian and Jacobian matrices, and D-SOBA-FO, which relies entirely on first-order gradients. We provide a comprehensive non-asymptotic convergence analysis and establish the transient iteration complexity of D-SOBA. This provides the first theoretical understanding of how network topology, data heterogeneity, and nested bilevel structures influence decentralized SBO. Extensive experimental results demonstrate the efficiency and theoretical advantages of D-SOBA. Boao Kong, Shuchen Zhu, Songtao Lu, Xinmeng Huang, Kun Yuan 0001 |
J. Mach. Learn. Res. | 5 |
| 2025 | Optimal Complexity in Byzantine-Robust Distributed Stochastic Optimization with Data HeterogeneityabstractIn this paper, we establish tight lower bounds for Byzantine-robust distributed first-order stochastic methods in both strongly convex and non-convex stochastic optimization. We reveal that when the distributed nodes have heterogeneous data, the convergence error comprises two components: a non-vanishing Byzantine error and a vanishing optimization error. We establish the lower bounds on the Byzantine error and on the minimum number of queries to a stochastic gradient oracle for achieving an arbitrarily small optimization error. Nevertheless, we also identify significant discrepancies between our established lower bounds and the existing upper bounds. To fill this gap, we leverage the techniques of Nesterov's acceleration and variance reduction to develop novel Byzantine-robust distributed stochastic optimization methods that provably match these lower bounds, up to at most logarithmic factors, implying that our established lower bounds are tight. Qiankun Shi, Kun Yuan 0001, Qing Ling 0001 |
J. Mach. Learn. Res. | 3 |
| 2025 | On the Trade-Off Between Flatness and Optimization in Distributed LearningabstractThis paper proposes a theoretical framework to evaluate and compare the performance of stochastic gradient algorithms for distributed learning in relation to their behavior around local minima in nonconvex environments. Previous works have noticed that convergence toward flat local minima tend to enhance the generalization ability of learning algorithms. This work discovers three interesting results. First, it shows that decentralized learning strategies are able to escape faster away from local minima and favor convergence toward flatter minima relative to the centralized solution. Second, in decentralized methods, the consensus strategy has a worse excess-risk performance than diffusion, giving it a better chance of escaping from local minima and favoring flatter minima. Third, and importantly, the ultimate classification accuracy is not solely dependent on the flatness of the local minimum but also on how well a learning algorithm can approach that minimum. In other words, the classification accuracy is a function of both flatness and optimization performance. In this regard, since diffusion has a lower excess-risk than consensus, when both algorithms are trained starting from random initial points, diffusion enhances the classification accuracy. The paper examines the interplay between the two measures of flatness and optimization error closely. One important conclusion is that decentralized strategies deliver in general enhanced classification accuracy because they strike a more favorable balance between flatness and optimization performance compared to the centralized solution. Zhaoxian Wu, Kun Yuan 0001, Ali H. Sayed |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | BEVHeight++: Toward Robust Visual Centric 3D Object DetectionabstractWhile most recent autonomous driving system focuses on developing perception methods on ego-vehicle sensors, people tend to overlook an alternative approach to leverage intelligent roadside cameras to extend the perception ability beyond the visual range. We discover that the state-of-the-art vision-centric detection methods perform poorly on roadside cameras. This is because these methods mainly focus on recovering the depth regarding the camera center, where the depth difference between the car and the ground quickly shrinks while the distance increases. In this paper, we propose a simple yet effective approach, dubbed BEVHeight++, to address this issue. In essence, we regress the height to the ground to achieve a distance-agnostic formulation to ease the optimization process of camera-only perception methods. By incorporating both height and depth encoding techniques, we achieve a more accurate and robust projection from 2D to BEV spaces. On popular 3D detection benchmarks of roadside cameras, our method surpasses all previous vision-centric methods by a significant margin. In terms of the ego-vehicle scenario, BEVHeight++ surpasses depth-only methods with increases of +2.8% NDS and +1.7% mAP on the nuScenes test set, and even higher gains of +9.3% NDS and +8.8% mAP on the nuScenes-C benchmark with object-level distortion. Consistent and substantial performance improvements are achieved across the KITTI, KITTI-360, and Waymo datasets as well. Lei Yang 0060, Jun Li 0082, Kun Yuan 0001, Li Wang 0092, Yi Huang 0038, Xinyu Zhang 0001, Kaicheng Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Asynchronous Diffusion Learning with Agent Subsampling and Local UpdatesabstractIn this work, we examine a network of agents operating asynchronously, aiming to discover an ideal global model that suits individual local datasets. Our assumption is that each agent independently chooses when to participate throughout the algorithm and the specific subset of its neighbourhood with which it will cooperate at any given moment. When an agent chooses to take part, it undergoes multiple local updates before conveying its outcomes to the sub-sampled neighbourhood. Under this setup, we prove that the resulting asynchronous diffusion strategy is stable in the mean-square error sense and provide performance guarantees specifically for the federated learning setting. We illustrate the findings with numerical simulations. Elsa Rizk, Kun Yuan 0001, Ali H. Sayed |
ICASSP | 2 |
| 2024 | Momentum Benefits Non-iid Federated Learning Simply and ProvablyabstractFederated learning is a powerful paradigm for large-scale machine learning, but it
faces significant challenges due to unreliable network connections, slow commu-
nication, and substantial data heterogeneity across clients. FedAvg and SCAFFOLD are two prominent algorithms to address these challenges. In particular,
FedAvg employs multiple local updates before communicating with a central
server, while SCAFFOLD maintains a control variable on each client to compen-
sate for “client drift” in its local updates. Various methods have been proposed
to enhance the convergence of these two algorithms, but they either make imprac-
tical adjustments to algorithmic structure, or rely on the assumption of bounded
data heterogeneity. This paper explores the utilization of momentum to enhance
the performance of FedAvg and SCAFFOLD. When all clients participate in the
training process, we demonstrate that incorporating momentum allows FedAvg
to converge without relying on the assumption of bounded data heterogeneity even
using a constant local learning rate. This is novel and fairly suprising as existing
analyses for FedAvg require bounded data heterogeneity even with diminishing
local learning rates. In partial client participation, we show that momentum en-
ables SCAFFOLD to converge provably faster without imposing any additional
assumptions. Furthermore, we use momentum to develop new variance-reduced
extensions of FedAvg and SCAFFOLD, which exhibit state-of-the-art conver-
gence rates. Our experimental results support all theoretical findings. Xinmeng Huang, Pengfei Wu 0006, Kun Yuan 0001 |
ICLR | 4 |
| 2024 | Distributed Bilevel Optimization with Communication CompressionabstractStochastic bilevel optimization tackles challenges involving nested optimization structures. Its fast-growing scale nowadays necessitates efficient distributed algorithms. In conventional distributed bilevel methods, each worker must transmit full-dimensional stochastic gradients to the server every iteration, leading to significant communication overhead and thus hindering efficiency and scalability. To resolve this issue, we introduce the first family of distributed bilevel algorithms with communication compression. The primary challenge in algorithmic development is mitigating bias in hypergradient estimation caused by the nested structure. We first propose C-SOBA, a simple yet effective approach with unbiased compression and provable linear speedup convergence. However, it relies on strong assumptions on bounded gradients. To address this limitation, we explore the use of moving average, error feedback, and multi-step compression in bilevel optimization, resulting in a series of advanced algorithms with relaxed assumptions and improved convergence properties. Numerical experiments show that our compressed bilevel algorithms can achieve $10\times$ reduction in communication overhead without severe performance degradation. Jie Hu 0022, Xinmeng Huang, Songtao Lu, Kun Yuan 0001 |
ICML | 6 |
| 2024 | SPARKLE: A Unified Single-Loop Primal-Dual Framework for Decentralized Bilevel OptimizationabstractThis paper studies decentralized bilevel optimization, in which multiple agents collaborate to solve problems involving nested optimization structures with neighborhood communications. Most existing literature primarily utilizes gradient tracking to mitigate the influence of data heterogeneity, without exploring other well-known heterogeneity-correction techniques such as EXTRA or Exact Diffusion. Additionally, these studies often employ identical decentralized strategies for both upper- and lower-level problems, neglecting to leverage distinct mechanisms across different levels. To address these limitations, this paper proposes SPARKLE, a unified single-loop primal-dual algorithm framework for decentralized bilevel optimization. SPARKLE offers the flexibility to incorporate various heterogeneity-correction strategies into the algorithm. Moreover, SPARKLE allows for different strategies to solve upper- and lower-level problems. We present a unified convergence analysis for SPARKLE, applicable to all its variants, with state-of-the-art convergence rates compared to existing decentralized bilevel algorithms. Our results further reveal that EXTRA and Exact Diffusion are more suitable for decentralized bilevel optimization, and using mixed strategies in bilevel algorithms brings more benefits than relying solely on gradient tracking. Shuchen Zhu, Boao Kong, Songtao Lu, Xinmeng Huang, Kun Yuan 0001 |
NeurIPS | 5 |
| 2024 | An enhanced gradient-tracking bound for distributed online stochastic convex optimization
Sulaiman A. Alghunaim, Kun Yuan 0001 |
Signal Process. | 2 |
| 2023 | BEVHeight: A Robust Framework for Vision-based Roadside 3D Object DetectionabstractWhile most recent autonomous driving system focuses on developing perception methods on ego-vehicle sensors, people tend to overlook an alternative approach to leverage intelligent roadside cameras to extend the perception ability beyond the visual range. We discover that the state-of-the-art vision-centric bird's eye view detection methods have inferior performances on roadside cameras. This is because these methods mainly focus on recovering the depth regarding the camera center, where the depth difference between the car and the ground quickly shrinks while the distance increases. In this paper, we propose a simple yet effective approach, dubbed BEVHeight, to address this issue. In essence, instead of predicting the pixel-wise depth, we regress the height to the ground to achieve a distance-agnostic formulation to ease the optimization process of camera-only perception methods. On popular 3D detection benchmarks of roadside cameras, our method surpasses all previous vision-centric methods by a significant margin. The code is available at https://github.com/ADLab-AutoDrive/BEVHeight. Lei Yang 0060, Kaicheng Yu, Jun Li 0082, Kun Yuan 0001, Li Wang 0092, Xinyu Zhang 0001 |
CVPR | 5 |
| 2023 | DSGD-CECA: Decentralized SGD with Communication-Optimal Exact Consensus AlgorithmabstractDecentralized Stochastic Gradient Descent (SGD) is an emerging neural network training approach that enables multiple agents to train a model collaboratively and simultaneously. Rather than using a central parameter server to collect gradients from all the agents, each agent keeps a copy of the model parameters and communicates with a small number of other agents to exchange model updates. Their communication, governed by the communication topology and gossip weight matrices, facilitates the exchange of model updates. The state-of-the-art approach uses the dynamic one-peer exponential-2 topology, achieving faster training times and improved scalability than the ring, grid, torus, and hypercube topologies. However, this approach requires a power-of-2 number of agents, which is impractical at scale. In this paper, we remove this restriction and propose Decentralized SGD with Communication-optimal Exact Consensus Algorithm (DSGD-CECA), which works for any number of agents while still achieving state-of-the-art properties. In particular, DSGD-CECA incurs a unit per-iteration communication overhead and an $\tilde{O}(n^3)$ transient iteration complexity. Our proof is based on newly discovered properties of gossip weight matrices and a novel approach to combine them with DSGD’s convergence analysis. Numerical experiments show the efficiency of DSGD-CECA. Lisang Ding, Kexin Jin, Bicheng Ying, Kun Yuan 0001, Wotao Yin |
ICML | 4 |
| 2023 | AdaNPC: Exploring Non-Parametric Classifier for Test-Time AdaptationabstractMany recent machine learning tasks focus to develop models that can generalize to unseen distributions. Domain generalization (DG) has become one of the key topics in various fields. Several literatures show that DG can be arbitrarily hard without exploiting target domain information. To address this issue, test-time adaptive (TTA) methods are proposed. Existing TTA methods require offline target data or extra sophisticated optimization procedures during the inference stage. In this work, we adopt Non-Parametric Classifier to perform the test-time Adaptation (AdaNPC). In particular, we construct a memory that contains the feature and label pairs from training domains. During inference, given a test instance, AdaNPC first recalls $k$ closed samples from the memory to vote for the prediction, and then the test feature and predicted label are added to the memory. In this way, the sample distribution in the memory can be gradually changed from the training distribution towards the test distribution with very little extra computation cost. We theoretically justify the rationality behind the proposed method. Besides, we test our model on extensive numerical experiments. AdaNPC significantly outperforms competitive baselines on various DG benchmarks. In particular, when the adaptation target is a series of domains, the adaptation accuracy of AdaNPC is $50$% higher than advanced TTA methods. Yifan Zhang 0004, Xue Wang 0010, Kexin Jin, Kun Yuan 0001, Zhang Zhang 0001, Liang Wang 0001, Rong Jin 0001, Tieniu Tan |
ICML | 4 |
| 2023 | Unbiased Compression Saves Communication in Distributed Optimization: When and How Much?abstractCommunication compression is a common technique in distributed optimization
that can alleviate communication overhead by transmitting compressed gradients
and model parameters. However, compression can introduce information distortion,
which slows down convergence and incurs more communication rounds to achieve
desired solutions. Given the trade-off between lower per-round communication
costs and additional rounds of communication, it is unclear whether communication
compression reduces the total communication cost.
This paper explores the conditions under which unbiased compression, a widely
used form of compression, can reduce the total communication cost, as well as the
extent to which it can do so. To this end, we present the first theoretical formulation
for characterizing the total communication cost in distributed optimization with
unbiased compressors. We demonstrate that unbiased compression alone does not
necessarily save the total communication cost, but this outcome can be achieved
if the compressors used by all workers are further assumed independent. We
establish lower bounds on the communication rounds required by algorithms using
independent unbiased compressors to minimize smooth convex functions and
show that these lower bounds are tight by refining the analysis for ADIANA.
Our results reveal that using independent unbiased compression can reduce the
total communication cost by a factor of up to $\Theta(\sqrt{\min\\{n,\kappa\\}})$ when all local
smoothness constants are constrained by a common upper bound, where $n$ is the
number of workers and $\kappa$ is the condition number of the functions being minimized.
These theoretical findings are supported by experimental results. Xinmeng Huang, Kun Yuan 0001 |
NeurIPS | 3 |
| 2023 | Removing Data Heterogeneity Influence Enhances Network Topology Dependence of Decentralized SGDabstractWe consider decentralized stochastic optimization problems, where a network of $n$ nodes cooperates to find a minimizer of the globally-averaged cost. A widely studied decentralized algorithm for this problem is the decentralized SGD (D-SGD), in which each node averages only with its neighbors. D-SGD is efficient in single-iteration communication, but it is very sensitive to the network topology. For smooth objective functions, the transient stage (which measures the number of iterations the algorithm has to experience before achieving the linear speedup stage) of D-SGD is on the order of ${O}(n/(1-\beta)^2)$ and $O(n^3/(1-\beta)^4)$ for strongly and generally convex cost functions, respectively, where $1-\beta \in (0,1)$ is a topology-dependent quantity that approaches $0$ for a large and sparse network. Hence, D-SGD suffers from slow convergence for large and sparse networks. In this work, we revisit the convergence property of the D$^2$/Exact-Diffusion algorithm. By eliminating the influence of data heterogeneity between nodes, D$^2$/Exact-diffusion is shown to have an enhanced transient stage that is on the order of $\tilde{O}(n/(1-\beta))$ and $O(n^3/(1-\beta)^2)$ for strongly and generally convex cost functions (where $\tilde{O}(\cdot)$ hides all logarithm factors), respectively. Moreover, when D$^2$/Exact-Diffusion is implemented with both gradient accumulation and multi-round gossip communications, its transient stage can be further improved to $\tilde{O}(1/(1-\beta)^{\frac{1}{2}})$ and $\tilde{O}(n/(1-\beta))$ for strongly and generally convex cost functions, respectively. To our knowledge, these established results for D$^2$/Exact-Diffusion have the best, i.e., weakest) dependence on network topology compared to existing decentralized algorithms. Numerical simulations are conducted to validate our theories. Kun Yuan 0001, Sulaiman A. Alghunaim, Xinmeng Huang |
J. Mach. Learn. Res. | 1 |
| 2022 | CHEX: CHannel EXploration for CNN Model CompressionabstractChannel pruning has been broadly recognized as an effective technique to reduce the computation and memory cost of deep convolutional neural networks. However, conventional pruning methods have limitations in that: they are restricted to pruning process only, and they require a fully pre-trained large model. Such limitations may lead to sub-optimal model quality as well as excessive memory and training cost. In this paper, we propose a novel Channel Exploration methodology, dubbed as CHEX, to rectify these problems. As opposed to pruning-only strategy, we propose to repeatedly prune and regrow the channels throughout the training process, which reduces the risk of pruning important channels prematurely. More exactly: From intra-Layer's aspect, we tackle the channel pruning problem via a well-known column subset selection (CSS) formulation. From inter-Layer's aspect, our regrowing stages open a path for dynamically re-allocating the number of channels across all the layers under a global channel sparsity constraint. In addition, all the exploration process is done in a single training from scratch without the need of a pre-trained large model. Experimental results demonstrate that CHEX can effectively reduce the FLOPs of diverse CNN architectures on a variety of computer vision tasks, including image classification, object detection, instance segmentation, and 3D vision. For example, our compressed ResNet-50 model on ImageNet dataset achieves 76% top-l accuracy with only 25% FLOPs of the original ResNet-50 model, outperforming previous state-of-the-art channel pruning methods. The checkpoints and code are available at here. Zejiang Hou, Minghai Qin, Fei Sun 0002, Kun Yuan 0001, Yi Xu 0008, Yen-Kuang Chen, Rong Jin 0001, Yuan Xie 0001, Sun-Yuan Kung |
CVPR | 5 |
| 2022 | A Byzantine-Resilient Dual Subgradient Method for Vertical Federated LearningabstractFederated learning (FL) raises new challenges on security risks, especially when the FL system involves Byzantine clients that send corrupted or adversarial messages to the central server for deteriorating the training paradigm. While there is an extensive research on robust algorithms for horizontal or data-partitioned FL problems, the exploration in Byzantine-resilient vertical or feature-partitioned FL is quite limited. In this paper, we provide a problem formulation of vertical FL in the presence of Byzantine attacks, and propose a Byzantine-resilient dual subgradient method. Convergence analysis is established, and the influence of the Byzantine clients is also clarified. Numerical experiments show the proposed algorithm is robust to various Byzantine attacks on vertical FL. Kun Yuan 0001, Zhaoxian Wu, Qing Ling 0001 |
ICASSP | 1 |
| 2022 | Effective Model Sparsification by Scheduled Grow-and-Prune Methods
Minghai Qin, Fei Sun 0002, Zejiang Hou, Kun Yuan 0001, Yi Xu 0008, Yanzhi Wang 0001, Yen-Kuang Chen, Rong Jin 0001, Yuan Xie 0001 |
ICLR | 5 |
| 2022 | Lower Bounds and Nearly Optimal Algorithms in Distributed Learning with Communication CompressionabstractRecent advances in distributed optimization and learning have shown that communication compression is one of the most effective means of reducing communication. While there have been many results for convergence rates with compressed communication, a lower bound is still missing.Analyses of algorithms with communication compression have identified two abstract properties that guarantee convergence: the unbiased property or the contractive property. They can be applied either unidirectionally (compressing messages from worker to server) or bidirectionally. In the smooth and non-convex stochastic regime, this paper establishes a lower bound for distributed algorithms whether using unbiased or contractive compressors in unidirection or bidirection. To close the gap between this lower bound and the best existing upper bound, we further propose an algorithm, NEOLITHIC, that almost reaches our lower bound (except for a logarithm factor) under mild conditions. Our results also show that using contractive compressors in bidirection can yield iterative methods that converge as fast as those using unbiased compressors unidirectionally. We report experimental results that validate our findings. Xinmeng Huang, Yiming Chen 0003, Wotao Yin, Kun Yuan 0001 |
NeurIPS | 4 |
| 2022 | Communication-Efficient Topologies for Decentralized Learning with $O(1)$ Consensus RateabstractDecentralized optimization is an emerging paradigm in distributed learning in which agents achieve network-wide solutions by peer-to-peer communication without the central server. Since communication tends to be slower than computation, when each agent communicates with only a few neighboring agents per iteration, they can complete iterations faster than with more agents or a central server. However, the total number of iterations to reach a network-wide solution is affected by the speed at which the information of the agents is ``mixed'' by communication. We found that popular communication topologies either have large degrees (such as stars and complete graphs) or are ineffective at mixing information (such as rings and grids). To address this problem, we propose a new family of topologies, EquiTopo, which has an (almost) constant degree and network-size-independent consensus rate which is used to measure the mixing efficiency.In the proposed family, EquiStatic has a degree of $\Theta(\ln(n))$, where $n$ is the network size, and a series of time-varying one-peer topologies, EquiDyn, has a constant degree of 1. We generate EquiDyn through a certain random sampling procedure. Both of them achieve $n$-independent consensus rate. We apply them to decentralized SGD and decentralized gradient tracking and obtain faster communication and better convergence, both theoretically and empirically. Our code is implemented through BlueFog and available at https://github.com/kexinjinnn/EquiTopo. Zhuoqing Song, Kexin Jin, Lei Shi 0010, Ming Yan 0006, Wotao Yin, Kun Yuan 0001 |
NeurIPS | 7 |
| 2022 | Revisiting Optimal Convergence Rate for Smooth and Non-convex Stochastic Decentralized OptimizationabstractWhile numerous effective decentralized algorithms have been proposed with theoretical guarantees and empirical successes, the performance limits in decentralized optimization, especially the influence of network topology and its associated weight matrix on the optimal convergence rate, have not been fully understood. While Lu and Sa have recently provided an optimal rate for non-convex stochastic decentralized optimization using weight matrices associated with linear graphs, the optimal rate with general weight matrices remains unclear. This paper revisits non-convex stochastic decentralized optimization and establishes an optimal convergence rate with general weight matrices. In addition, we also establish the first optimal rate when non-convex loss functions further satisfy the Polyak-Lojasiewicz (PL) condition. Following existing lines of analysis in literature cannot achieve these results. Instead, we leverage the Ring-Lattice graph to admit general weight matrices while maintaining the optimal relation between the graph diameter and weight matrix connectivity. Lastly, we develop a new decentralized algorithm to attain the above two optimal rates up to logarithm factors. Kun Yuan 0001, Xinmeng Huang, Yiming Chen 0003, Yingya Zhang |
NeurIPS | 1 |
| 2021 | DecentLaM: Decentralized Momentum SGD for Large-batch Deep TrainingabstractThe scale of deep learning nowadays calls for efficient distributed training algorithms. Decentralized momentum SGD (DmSGD), in which each node averages only with its neighbors, is more communication efficient than vanilla Parallel momentum SGD that incurs global average across all computing nodes. On the other hand, the large-batch training has been demonstrated critical to achieve runtime speedup. This motivates us to investigate how DmSGD performs in the large-batch scenario.In this work, we find the momentum term can amplify the inconsistency bias in DmSGD. Such bias becomes more evident as batch-size grows large and hence results in severe performance degradation. We next propose DecentLaM, a novel decentralized large-batch momentum SGD to remove the momentum-incurred bias. The convergence rate for both strongly convex and non-convex scenarios is established. Our theoretical results justify the superiority of DecentLaM to DmSGD especially in the large-batch scenario. Experimental results on a a variety of computer vision tasks and models show that DecentLaM promises both efficient and high-quality training. Kun Yuan 0001, Yiming Chen 0003, Xinmeng Huang, Yingya Zhang, Wotao Yin |
ICCV | 1 |
| 2021 | Accelerating Gossip SGD with Periodic Global AveragingabstractCommunication overhead hinders the scalability of large-scale distributed training. Gossip SGD, where each node averages only with its neighbors, is more communication-efficient than the prevalent parallel SGD. However, its convergence rate is reversely proportional to quantity $1-\beta$ which measures the network connectivity. On large and sparse networks where $1-\beta \to 0$, Gossip SGD requires more iterations to converge, which offsets against its communication benefit. This paper introduces Gossip-PGA, which adds Periodic Global Averaging to accelerate Gossip SGD. Its transient stage, i.e., the iterations required to reach asymptotic linear speedup stage, improves from $\Omega(\beta^4 n^3/(1-\beta)^4)$ to $\Omega(\beta^4 n^3 H^4)$ for non-convex problems. The influence of network topology in Gossip-PGA can be controlled by the averaging period $H$. Its transient-stage complexity is also superior to local SGD which has order $\Omega(n^3 H^4)$. Empirical results of large-scale training on image classification (ResNet50) and language modeling (BERT) validate our theoretical findings. Yiming Chen 0003, Kun Yuan 0001, Yingya Zhang, Wotao Yin |
ICML | 2 |
| 2021 | An Improved Analysis and Rates for Variance Reduction under Without-replacement Sampling OrdersabstractWhen applying a stochastic algorithm, one must choose an order to draw samples. The practical choices are without-replacement sampling orders, which are empirically faster and more cache-friendly than uniform-iid-sampling but often have inferior theoretical guarantees. Without-replacement sampling is well understood only for SGD without variance reduction. In this paper, we will improve the convergence analysis and rates of variance reduction under without-replacement sampling orders for composite finite-sum minimization.Our results are in two-folds. First, we develop a damped variant of Finito called Prox-DFinito and establish its convergence rates with random reshuffling, cyclic sampling, and shuffling-once, under both generally and strongly convex scenarios. These rates match full-batch gradient descent and are state-of-the-art compared to the existing results for without-replacement sampling with variance-reduction. Second, our analysis can gauge how the cyclic order will influence the rate of cyclic sampling and, thus, allows us to derive the optimal fixed ordering. In the highly data-heterogeneous scenario, Prox-DFinito with optimal cyclic sampling can attain a sample-size-independent convergence rate, which, to our knowledge, is the first result that can match with uniform-iid-sampling with variance reduction. We also propose a practical method to discover the optimal cyclic ordering numerically. Xinmeng Huang, Kun Yuan 0001, Xianghui Mao, Wotao Yin |
NeurIPS | 2 |
| 2021 | Exponential Graph is Provably Efficient for Decentralized Deep TrainingabstractDecentralized SGD is an emerging training method for deep learning known for its much less (thus faster) communication per iteration, which relaxes the averaging step in parallel SGD to inexact averaging. The less exact the averaging is, however, the more the total iterations the training needs to take. Therefore, the key to making decentralized SGD efficient is to realize nearly-exact averaging using little communication. This requires a skillful choice of communication topology, which is an under-studied topic in decentralized optimization.In this paper, we study so-called exponential graphs where every node is connected to $O(\log(n))$ neighbors and $n$ is the total number of nodes. This work proves such graphs can lead to both fast communication and effective averaging simultaneously. We also discover that a sequence of $\log(n)$ one-peer exponential graphs, in which each node communicates to one single neighbor per iteration, can together achieve exact averaging. This favorable property enables one-peer exponential graph to average as effective as its static counterpart but communicates more efficiently. We apply these exponential graphs in decentralized (momentum) SGD to obtain the state-of-the-art balance between per-iteration communication and iteration complexity among all commonly-used topologies. Experimental results on a variety of tasks and models demonstrate that decentralized (momentum) SGD over exponential graphs promises both fast and high-quality training. Our code is implemented through BlueFog and available at https://github.com/Bluefog-Lib/NeurIPS2021-Exponential-Graph. Bicheng Ying, Kun Yuan 0001, Yiming Chen 0003, Hanbin Hu, Wotao Yin |
NeurIPS | 2 |
| 2019 | COVER: A Cluster-based Variance Reduced Method for Online LearningabstractIn this paper, we develop a stochastic-gradient learning algorithm for situations involving streaming data that arise from an underlying clustered structure. In such settings, the variance of gradient noise can be decomposed into the in-cluster variance σin2plus the between-cluster variance σbet2. We develop a cluster-based online variancereduced method (COVER) to eliminate σbet2and improve the MSD performance of stochastic-gradient descent (SGD) to the order of O(σin2). We establish the convergence property of COVER and derive a tight closed-form mean-square deviation (MSD) performance expression. Our simulations illustrate the improved performance of COVER in terms of steady-state performance. Kun Yuan 0001, Bicheng Ying, Ali H. Sayed |
ICASSP | 1 |
| 2019 | A Linearly Convergent Proximal Gradient Algorithm for Decentralized OptimizationabstractDecentralized optimization is a powerful paradigm that finds applications in engineering and learning design. This work studies decentralized composite optimization problems with non-smooth regularization terms. Most existing gradient-based proximal decentralized methods are known to converge to the optimal solution with sublinear rates, and it remains unclear whether this family of methods can achieve global linear convergence. To tackle this problem, this work assumes the non-smooth regularization term is common across all networked agents, which is the case for many machine learning problems. Under this condition, we design a proximal gradient decentralized algorithm whose fixed point coincides with the desired minimizer. We then provide a concise proof that establishes its linear convergence. In the absence of the non-smooth term, our analysis technique covers the well known EXTRA algorithm and provides useful bounds on the convergence rate and step-size. Sulaiman A. Alghunaim, Kun Yuan 0001, Ali H. Sayed |
NeurIPS | 2 |
| 2018 | Convergence of Variance-Reduced Learning Under Random ReshufflingabstractSeveral useful variance-reduced stochastic gradient algorithms, such as SVRG, SAGA, Finito, and SAG, have been proposed to minimize empirical risks with linear convergence properties to the exact minimizers. The existing convergence results assume uniform data sampling with replacement. However, it has been observed that random reshuffling can deliver superior performance and, yet, no formal proofs or guarantees of exact convergence exist for variance-reduced algorithms under random reshuffling. This paper makes two contributions. First, it resolves this open issue and provides the first theoretical guarantee of linear convergence under random reshuffling for SAGA; the argument is also adaptable to other variance-reduced algorithms. Second, under random reshuffling, the paper proposes a new amortized variance-reduced gradient (AVRG) algorithm with constant storage requirements compared to SAGA and with balanced gradient computations compared to SVRG. AVRG is also shown analytically to converge linearly. Bicheng Ying, Kun Yuan 0001, Ali H. Sayed |
ICASSP | 2 |
| 2016 | On the influence of momentum acceleration on online learningabstractThis paper examines the convergence rate and mean-square-error performance of momentum stochastic gradient methods in the constant step-size and slow adaptation regime. The results establish that momentum methods are equivalent to the standard stochastic gradient method with a re-scaled (larger) step-size value. The equivalence result is established for all time instants and not only in steady-state. The analysis is carried out for general risk functions, and is not limited to quadratic risks. One notable conclusion is that the well-known benefits of momentum constructions for deterministic optimization problems do not necessarily carry over to the stochastic setting when gradient noise is present and continuous adaptation is necessary. The analysis suggests a method to enhance performance in the stochastic setting by tuning the momentum parameter over time. Kun Yuan 0001, Bicheng Ying, Ali H. Sayed |
ICASSP | 1 |
| 2016 | Cooperative tracking for nonlinear multi-agent systems with hybrid time-delayed protocol
Jian-Qiang Hu, Jinde Cao, Kun Yuan 0001, Tasawar Hayat |
Neurocomputing | 3 |
| 2016 | On the Influence of Momentum Acceleration on Online LearningabstractThe article examines in some detail the convergence rate and mean-square-error performance of momentum stochastic gradient methods in the constant step-size and slow adaptation regime. The results establish that momentum methods are equivalent to the standard stochastic gradient method with a re-scaled (larger) step-size value. The size of the re-scaling is determined by the value of the momentum parameter. The equivalence result is established for all time instants and not only in steady-state. The analysis is carried out for general strongly convex and smooth risk functions, and is not limited to quadratic risks. One notable conclusion is that the well-known benefits of momentum constructions for deterministic optimization problems do not necessarily carry over to the adaptive online setting when small constant step-sizes are used to enable continuous adaptation and learning in the presence of persistent gradient noise. From simulations, the equivalence between momentum and standard stochastic gradient methods is also observed for non-differentiable and non-convex problems. Kun Yuan 0001, Bicheng Ying, Ali H. Sayed |
J. Mach. Learn. Res. | 1 |
| 2015 | Communication-Efficient Decentralized Event Monitoring in Wireless Sensor NetworksabstractIn this paper, we consider monitoring multiple events in a sensing field using a large-scale wireless sensor network (WSN). The goal is to develop communication-efficient algorithms that are scalable to the network size. Exploiting the sparse nature of the events, we formulate the event monitoring task as an `1 regularized nonnegative least squares problem where the optimization variable is a sparse vector representing the locations and magnitudes of events. Traditionally the problem can be reformulated by letting each sensor hold a local copy of the event vector and imposing consensus constraints on the local copies, and solved by decentralized algorithms such as the alternating direction method of multipliers (ADMM). This technique requires each sensor to exchange their estimates of the entire sparse vector and hence leads to high communication cost. Motivated by the observation that an event usually has limited influence range, we develop two communication-efficient decentralized algorithms, one is the partial consensus algorithm and the other is the Jacobi approach. In the partial consensus algorithm that is based on the ADMM, each sensor is responsible for recovering those events relevant to itself, and hence only consent with neighboring nodes on a part of the sparse vector. This strategy greatly reduces the amount of information exchanged among sensors. The Jacobi approach addresses the case that each sensor cares about the event occurring at its own position. Jacobi-like iterates are shown to be much faster than other algorithms, and incur minimal communication cost per iteration. Simulation results validate the effectiveness of the proposed algorithms and demonstrate the importance of proper modelling in designing communication-efficient decentralized algorithms. Kun Yuan 0001, Qing Ling 0001, Zhi Tian |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | Linearly convergent decentralized consensus optimization with the alternating direction method of multipliersabstractIn the decentralized consensus optimization problem, a network of agents minimizes the summation of their local objective functions on a common set of variables, allowing only information exchange among neighbors. The alternating direction method of multipliers (ADMM) has been shown to be a powerful tool for solving the problem with empirically fast convergence. This paper establishes the linear convergence rate of the ADMM in decentralized consensus optimization. The theoretical convergence rate is a function of the network topology, properties of the local objective functions, and the algorithm parameter. This result not only gives a performance guarantee for the ADMM but also provides a guideline to accelerate its convergence rate for the decentralized consensus optimization problems. Wei Shi 0010, Qing Ling 0001, Kun Yuan 0001, Gang Wu 0011, Wotao Yin |
ICASSP | 3 |
| 2008 | Robust Stabilization of the Distributed Parameter System With Time Delay via Fuzzy ControlabstractIn this paper, stabilization of the distributed parameter system (DPS) with time delay is studied using Galerkin's method and fuzzy control. With the help of Galerkin's method, the dynamics of DPS with time delay can be first converted into a group of low-order functional ordinary differential equations, which will be used for design of the robust fuzzy controller. The fuzzy controller designed can guarantee exponential stability of the closed-loop DPS. Some sufficient conditions are derived for the stabilization together with the linear matrix inequality design approach. The effectiveness of the proposed control design methodology is demonstrated in numerical simulations. Kun Yuan 0001, Han-Xiong Li, Jinde Cao |
IEEE Trans. Fuzzy Syst. | 1 |
| 2006 | Exponential stability and periodic solutions of fuzzy cellular neural networks with time-varying delays
Kun Yuan 0001, Jinde Cao, Jianming Deng |
Neurocomputing | 1 |
| 2006 | Global Asymptotical Stability of Recurrent Neural Networks With Multiple Discrete Delays and Distributed DelaysabstractBy employing the Lyapunov-Krasovskii functional and linear matrix inequality (LMI) approach, the problem of global asymptotical stability is studied for recurrent neural networks with both discrete time-varying delays and distributed time-varying delays. Some sufficient conditions are given for checking the global asymptotical stability of recurrent neural networks with mixed time-varying delay. The proposed LMI result is computationally efficient as it can be solved numerically using standard commercial software. Two examples are given to show the usefulness of the results. Jinde Cao, Kun Yuan 0001, Han-Xiong Li |
IEEE Trans. Neural Networks | 2 |
| 2006 | Robust Stability of Switched Cohen-Grossberg Neural Networks With Mixed Time-Varying DelaysabstractBy combining Cohen-Grossberg neural networks with an arbitrary switching rule, the mathematical model of a class of switched Cohen-Grossberg neural networks with mixed time-varying delays is established. Moreover, robust stability for such switched Cohen-Grossberg neural networks is analyzed based on a Lyapunov approach and linear matrix inequality (LMI) technique. Simple sufficient conditions are given to guarantee the switched Cohen-Grossberg neural networks to be globally asymptotically stable for all admissible parametric uncertainties. The proposed LMI-based results are computationally efficient as they can be solved numerically using standard commercial software. An example is given to illustrate the usefulness of the results. Kun Yuan 0001, Jinde Cao, Han-Xiong Li |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2004 | Global Exponential Stability of Cohen-Grossberg Neural Networks with Multiple Time-Varying Delays
Kun Yuan 0001, Jinde Cao |
ISNN (1) | 1 |