EDBT 2026 Demo / reviewers in the wild / expert
Sifan Yang
dblp:251/2905
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Optimization for machine learning · 76% Learning paradigms · 16% Kernel, tree and ensemble methods · 7% | |
| Theoretical computer science
3 papers |
Mathematical optimization · 100% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning
variance reduction |
2.4 | 3 | 2025 | Revisiting Stochastic Multi-Level Compositional Optimization · IEEE Trans. Pattern Anal. Mach. Intell. 2025 Efficient Sign-Based Optimization: Accelerating Convergence via Variance Reduction · NeurIPS 2024 Adaptive Variance Reduction for Stochastic Optimization under Weaker Assumptions · NeurIPS 2024 |
Mathematical optimization
online optimization |
1.7 | 2 | 2025 | Smoothed Online Convex Optimization with Delayed Feedback · IJCAI 2025 Online Nonsubmodular Optimization with Delayed Feedback in the Bandit Setting · AAAI 2025 |
Machine learning › Optimization for machine learning › optimization › compositional optimization
stochastic compositional optimization |
1.6 | 2 | 2025 | Revisiting Stochastic Multi-Level Compositional Optimization · IEEE Trans. Pattern Anal. Mach. Intell. 2025 Adaptive Variance Reduction for Stochastic Optimization under Weaker Assumptions · NeurIPS 2024 |
Machine learning › Optimization for machine learning
stochastic optimization |
1.5 | 2 | 2024 | Efficient Sign-Based Optimization: Accelerating Convergence via Variance Reduction · NeurIPS 2024 Adaptive Variance Reduction for Stochastic Optimization under Weaker Assumptions · NeurIPS 2024 |
Machine learning › Optimization for machine learning › adaptive optimization
adaptive subgradient method |
0.9 | 1 | 2025 | Dimension-Free Adaptive Subgradient Methods with Frequent Directions · ICML 2025 |
Machine learning › Learning paradigms › continual learning › catastrophic forgetting
catastrophic forgetting mitigation |
0.9 | 1 | 2025 | Learning without Isolation: Pathway Protection for Continual Learning · ICML 2025 |
Machine learning › Learning paradigms
continual learning |
0.9 | 1 | 2025 | Learning without Isolation: Pathway Protection for Continual Learning · ICML 2025 |
Mathematical optimization › online optimization
bandit optimization |
0.9 | 1 | 2025 | Online Nonsubmodular Optimization with Delayed Feedback in the Bandit Setting · AAAI 2025 |
Mathematical optimization › online optimization
delayed feedback |
0.9 | 1 | 2025 | Online Nonsubmodular Optimization with Delayed Feedback in the Bandit Setting · AAAI 2025 |
Mathematical optimization › online optimization › online convex optimization
smoothed online convex optimization |
0.9 | 1 | 2025 | Smoothed Online Convex Optimization with Delayed Feedback · IJCAI 2025 |
Mathematical optimization
submodular optimization |
0.9 | 1 | 2025 | Online Nonsubmodular Optimization with Delayed Feedback in the Bandit Setting · AAAI 2025 |
Machine learning › Optimization for machine learning
distributed optimization |
0.8 | 1 | 2024 | Efficient Sign-Based Optimization: Accelerating Convergence via Variance Reduction · NeurIPS 2024 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning
majority voting |
0.8 | 1 | 2024 | Efficient Sign-Based Optimization: Accelerating Convergence via Variance Reduction · NeurIPS 2024 |
Mathematical optimization › stochastic optimization
compositional optimization |
0.8 | 1 | 2024 | Projection-Free Variance Reduction Methods for Stochastic Constrained Multi-Level Compositional Optimization · ICML 2024 |
Mathematical optimization › continuous optimization › convex optimization › first-order methods
projection-free optimization |
0.8 | 1 | 2024 | Projection-Free Variance Reduction Methods for Stochastic Constrained Multi-Level Compositional Optimization · ICML 2024 |
Mathematical optimization
stochastic optimization |
0.8 | 1 | 2024 | Projection-Free Variance Reduction Methods for Stochastic Constrained Multi-Level Compositional Optimization · ICML 2024 |
Mathematical optimization › stochastic optimization
variance reduction |
0.8 | 1 | 2024 | Projection-Free Variance Reduction Methods for Stochastic Constrained Multi-Level Compositional Optimization · ICML 2024 |
Mathematical optimization
frank-wolfe algorithm |
0.2 | 1 | 2024 | Projection-Free Variance Reduction Methods for Stochastic Constrained Multi-Level Compositional Optimization · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
online gradient descent · 1.7variance reduction · 1.6adaptive learning rate · 1.6stage-wise optimization · 0.9primal-dual framework · 0.9one-point gradient estimator · 0.9model fusion · 0.9meta-expert framework · 0.9graph matching · 0.9frequent directions · 0.9fast frequent directions · 0.9stage-wise adaptation · 0.8signSGD · 0.8gradient mapping · 0.8STORM · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Online Nonsubmodular Optimization with Delayed Feedback in the Bandit SettingabstractWe investigate the online nonsubmodular optimization with delayed feedback in the bandit setting, where the loss function is α-weakly DR-submodular and β-weakly DR-supermodular. Previous work has established an (α,β)-regret bound of O(nd^⅓T^⅔), where n is the dimensionality and d is the maximum delay. However, its regret bound relies on the maximum delay and is thus sensitive to irregular delays. Additionally, it couples the effects of delays and bandit feedback as its bound is the product of the delay term and the O(nT^⅔) regret bound in the bandit setting without delayed feedback. In this paper, we develop two algorithms to address these limitations, respectively. Firstly, we propose a novel method, namely DBGD-NF, which employs the one-point gradient estimator and utilizes all the available estimated gradients in each round to update the decision. It achieves a better O(nd̅^⅓T^⅔) regret bound, which is relevant to the average delay d̅ = 1/T ∑ₜ₌₁ᵀ dₜ Sifan Yang, Yuanyu Wan |
AAAI | 1 |
| 2025 | Learning without Isolation: Pathway Protection for Continual LearningabstractDeep networks are prone to catastrophic forgetting during sequential task learning, i.e., losing the knowledge about old tasks upon learning new tasks. To this end, continual learning (CL) has emerged, whose existing methods focus mostly on regulating or protecting the parameters associated with the previous tasks. However, parameter protection is often impractical, since the size of parameters for storing the old-task knowledge increases linearly with the number of tasks, otherwise it is hard to preserve the parameters related to the old-task knowledge. In this work, we bring a dual opinion from neuroscience and physics to CL: in the whole networks, the pathways matter more than the parameters when concerning the knowledge acquired from the old tasks. Following this opinion, we propose a novel CL framework, learning without isolation (LwI), where model fusion is formulated as graph matching and the pathways occupied by the old tasks are protected without being isolated. Thanks to the sparsity of activation channels in a deep network, LwI can adaptively allocate available pathways for a new task, realizing pathway protection and addressing catastrophic forgetting in a parameter-effcient manner. Experiments on popular benchmark datasets demonstrate the superiority of the proposed LwI. Zhikang Chen, Abudukelimu Wuerkaixi, Sen Cui, Haoxuan Li 0001, Jingfeng Zhang, Bo Han 0003, Gang Niu 0001, Houfang Liu, Yi Yang 0039, Sifan Yang, Changshui Zhang |
ICML | 11 |
| 2025 | Dimension-Free Adaptive Subgradient Methods with Frequent DirectionsabstractIn this paper, we investigate the acceleration of adaptive subgradient methods through frequent directions (FD), a widely-used matrix sketching technique. The state-of-the-art regret bound exhibits a _linear_ dependence on the dimensionality $d$, leading to unsatisfactory guarantees for high-dimensional problems. Additionally, it suffers from an $O(\tau^2 d)$ time complexity per round, which scales quadratically with the sketching size $\tau$. To overcome these issues, we first propose an algorithm named FTSL, achieving a tighter regret bound that is independent of the dimensionality. The key idea is to integrate FD with adaptive subgradient methods under _the primal-dual framework_ and add the cumulative discarded information of FD back. To reduce its time complexity, we further utilize fast FD to expedite FTSL, yielding a better complexity of $O(\tau d)$ while maintaining the same regret bound. Moreover, to mitigate the computational cost for optimization problems involving matrix variables (e.g., training neural networks), we adapt FD to Shampoo, a popular optimization algorithm that accounts for the structure of decision, and give a novel analysis under _the primal-dual framework_. Our proposed method obtains an improved dimension-free regret bound. Experimental results have verified the efficiency and effectiveness of our approaches. Sifan Yang, Yuanyu Wan, Peijia Li, Yibo Wang 0005, Zhewei Wei, Lijun Zhang 0005 |
ICML | 1 |
| 2025 | Smoothed Online Convex Optimization with Delayed FeedbackabstractSmoothed online convex optimization (SOCO), in which the online player incurs both a hitting cost and a switching cost for changing its decisions, has garnered significant attention in recent years. While existing studies typically assume that the gradient information is revealed immediately, such an assumption may not hold in some real-world applications. To overcome this limitation, we investigate SOCO with delayed feedback, and develop two online algorithms that can minimize the dynamic regret with switching cost. Firstly, we extend Mild-OGD, an existing algorithm that adopts the meta-expert framework for online convex optimization with delayed feedback, to account for switching cost. Specifically, we analyze the switching cost in the expert-algorithm of Mild-OGD, and then modify its meta-algorithm to incorporate this cost when assigning the weight to each expert. We demonstrate that our proposed method, Smelt-DOGD can achieve an O(√(dT(P_T+1))) dynamic regret bound with switching cost, where d is the maximum delay and P_T is the path-length. Secondly, we develop an efficient variant to reduce the number of projections per round from O(log T) to 1, yet maintaining the same theoretical guarantee. The key idea is to construct a new surrogate loss defined over a simpler domain for expert-algorithms so that these experts do not need to perform the complex projection operations in each round. Finally, we conduct experiments to validate the effectiveness and efficiency of our algorithms. Sifan Yang, Yuanyu Wan |
IJCAI | 1 |
| 2025 | Revisiting Stochastic Multi-Level Compositional OptimizationabstractThis paper explores stochastic multi-level compositional optimization, where the objective function is a composition of multiple smooth functions. Traditional methods for solving this problem suffer from either sub-optimal sample complexities or require huge batch sizes. To address these limitations, we introduce the Stochastic Multi-level Variance Reduction (SMVR) method. In the expectation case, our SMVR method attains the optimal sample complexity of $\mathcal {O}(1/\epsilon ^{3})$O(1/ε3) to find an $\epsilon$ε-stationary point for non-convex objectives. When the function satisfies convexity or the Polyak-Łojasiewicz (PL) condition, we propose a stage-wise SMVR variant. This variant improves the sample complexity to $\mathcal {O}(1/\epsilon ^{2})$O(1/ε2) for convex functions and $\mathcal {O}(1/(\mu \epsilon ))$O(1/(με)) for functions meeting the $\mu$μ-PL condition or $\mu$μ-strong convexity. These complexities match the lower bounds not only in terms of $\epsilon$ε but also in terms of $\mu$μ (for PL or strongly convex functions), without relying on large batch sizes in each iteration. Furthermore, in the finite-sum case, we develop the SMVR-FS algorithm, which can achieve a complexity of $\mathcal {O}(\sqrt{n}/\epsilon ^{2})$O(n/ε2) for non-convex objectives, $\mathcal {O}(\sqrt{n}/\epsilon \log (1/\epsilon ))$O(n/εlog(1/ε)) for convex functions and $\mathcal {O}(\sqrt{n}/\mu \log (1/\epsilon ))$O(n/μlog(1/ε)) for objectives satisfying the $\mu$μ-PL condition, where $n$n denotes the number of functions in each level. To make use of adaptive learning rates, we propose the Adaptive SMVR method, which maintains the same complexities while demonstrating faster convergence in practice. Wei Jiang 0029, Sifan Yang, Yibo Wang 0005, Tianbao Yang, Lijun Zhang 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Projection-Free Variance Reduction Methods for Stochastic Constrained Multi-Level Compositional OptimizationabstractThis paper investigates projection-free algorithms for stochastic constrained multi-level optimization. In this context, the objective function is a nested composition of several smooth functions, and the decision set is closed and convex. Existing projection-free algorithms for solving this problem suffer from two limitations: 1) they solely focus on the gradient mapping criterion and fail to match the optimal sample complexities in unconstrained settings; 2) their analysis is exclusively applicable to non-convex functions, without considering convex and strongly convex objectives. To address these issues, we introduce novel projection-free variance reduction algorithms and analyze their complexities under different criteria. For gradient mapping, our complexities improve existing results and match the optimal rates for unconstrained problems. For the widely-used Frank-Wolfe gap criterion, we provide theoretical guarantees that align with those for single-level problems. Additionally, by using a stage-wise adaptation, we further obtain complexities for convex and strongly convex functions. Finally, numerical experiments on different tasks demonstrate the effectiveness of our methods. Wei Jiang 0029, Sifan Yang, Yibo Wang 0005, Yuanyu Wan, Lijun Zhang 0005 |
ICML | 2 |
| 2024 | Adaptive Variance Reduction for Stochastic Optimization under Weaker AssumptionsabstractThis paper explores adaptive variance reduction methods for stochastic optimization based on the STORM technique. Existing adaptive extensions of STORM rely on strong assumptions like bounded gradients and bounded function values, or suffer an additional $\mathcal{O}(\log T)$ term in the convergence rate. To address these limitations, we introduce a novel adaptive STORM method that achieves an optimal convergence rate of $\mathcal{O}(T^{-1/3})$ for non-convex functions with our newly designed learning rate strategy. Compared with existing approaches, our method requires weaker assumptions and attains the optimal convergence rate without the additional $\mathcal{O}(\log T)$ term. We also extend the proposed technique to stochastic compositional optimization, obtaining the same optimal rate of $\mathcal{O}(T^{-1/3})$. Furthermore, we investigate the non-convex finite-sum problem and develop another innovative adaptive variance reduction method that achieves an optimal convergence rate of $\mathcal{O}(n^{1/4} T^{-1/2} )$, where $n$ represents the number of component functions. Numerical experiments across various tasks validate the effectiveness of our method. Wei Jiang 0029, Sifan Yang, Yibo Wang 0005, Lijun Zhang 0005 |
NeurIPS | 2 |
| 2024 | Efficient Sign-Based Optimization: Accelerating Convergence via Variance ReductionabstractSign stochastic gradient descent (signSGD) is a communication-efficient method that transmits only the sign of stochastic gradients for parameter updating. Existing literature has demonstrated that signSGD can achieve a convergence rate of $\mathcal{O}(d^{1/2}T^{-1/4})$, where $d$ represents the dimension and $T$ is the iteration number. In this paper, we improve this convergence rate to $\mathcal{O}(d^{1/2}T^{-1/3})$ by introducing the Sign-based Stochastic Variance Reduction (SSVR) method, which employs variance reduction estimators to track gradients and leverages their signs to update. For finite-sum problems, our method can be further enhanced to achieve a convergence rate of $\mathcal{O}(m^{1/4}d^{1/2}T^{-1/2})$, where $m$ denotes the number of component functions. Furthermore, we investigate the heterogeneous majority vote in distributed settings and introduce two novel algorithms that attain improved convergence rates of $\mathcal{O}(d^{1/2}T^{-1/2} + dn^{-1/2})$ and $\mathcal{O}(d^{1/4}T^{-1/4})$ respectively, outperforming the previous results of $\mathcal{O}(dT^{-1/4} + dn^{-1/2})$ and $\mathcal{O}(d^{3/8}T^{-1/8})$, where $n$ represents the number of nodes. Numerical experiments across different tasks validate the effectiveness of our proposed methods. Wei Jiang 0029, Sifan Yang, Lijun Zhang 0005 |
NeurIPS | 2 |
| 2022 | SADG-Net: Sparse Adaptive Dynamic Guidance Network for Depth CompletionabstractExisting depth completion methods with standard CNN and fixed guidance information often produce invalid value diffusion and mismatch between the guidance and the depth. To address this issue, we propose a SADG-Net for depth completion. Specifically, a sparse adaptive module is designed to infer the initial dense depth map and its confidence, as well as the affinity between pixels. Then we develop a dynamic guidance spatial propagation network to refine the initial depth map and dynamically update the guidance information with the inferred depth. In contrast to previous algorithms, our method effectively optimizes the processing for sparse depth and significantly alleviates the error accumulation issues in spatial propagation. Extensive experiments demonstrate that our model improves upon the state-of-the-art performance on NYUv2 and BIDCD datasets. Guodong Zhang 0004, Chenchen Feng, Sifan Yang, Wenming Yang, Guijin Wang |
ICME | 4 |
| 2021 | Non-Local Aggregation for RGB-D Semantic SegmentationabstractExploiting both RGB (2D appearance) and Depth (3D geometry) information can improve the performance of semantic segmentation. However, due to the inherent difference between the RGB and Depth information, it remains a challenging problem in how to integrate RGB-D features effectively. In this letter, to address this issue, we propose a Non-local Aggregation Network (NANet), with a well-designed Multi-modality Non-local Aggregation Module (MNAM), to better exploit the non-local context of RGB-D features at multi-stage. Compared with most existing RGB-D semantic segmentation schemes, which only exploit local RGB-D features, the MNAM enables the aggregation of non-local RGB-D information along both spatial and channel dimensions. The proposed NANet achieves comparable performances with state-of-the-art methods on popular RGB-D benchmarks, NYUDv2 and SUN-RGBD. Guodong Zhang 0004, Jing-Hao Xue, Pengwei Xie, Sifan Yang, Guijin Wang |
IEEE Signal Process. Lett. | 4 |