EDBT 2026 Demo / reviewers in the wild / expert
Jian Li 0040
dblp:33/5448-40
· DBLP profile ↗
29ranked-venue papers
16as first author
24since 2021 · last 2026
0000-0003-4977-1802ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 15 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Redundancy to Precision: Constructing a Compressed Tree-Based Experiential Case Base for LLMs
Mucun Xie, Yijia Xu, Jian Li 0040 |
ICCBR | 3 |
| 2026 | Stability and generalization of differentially private minimax problems
Yilin Kang 0002, Jian Li 0040, Yong Liu 0018, Weiping Wang 0005 |
Neurocomputing | 2 |
| 2026 | FedNK-RF: Federated Kernel Learning With Heterogeneous Data and Optimal RatesabstractFederated learning (FL) has become a mainstream decentralized learning paradigm due to its privacy-preserving features. However, the heterogeneity of data in FL can reduce predictive accuracy and complicate the analysis of the generalization properties of FL methods. In this article, we propose efficient federated kernel learning (FedK) algorithms and study their generalization properties. We first devise FedK with random features (FedK-RF), which acquires global information through sharing RF of local data subsets, enhancing predictive capability while protecting privacy. We then propose federated Nyström approximation with RF (FedNK-RF) that reduces errors resulted from RF. Furthermore, using integral operator theory, we derive the excess risk bounds with minimax optimal rates, which illustrate the impacts from data heterogeneity and shared information. Finally, we conduct several experiments that demonstrate the superiority of the proposed FedNK-RF. Xuning Zhang, Jian Li 0040, Rong Yin 0001, Weiping Wang 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Improving Mathematical Reasoning Abilities of Small Language Models via Key-Point-Driven DistillationabstractLarge Language Models (LLMs) excel in mathematical reasoning due to their extensive parameters and training data, but their high computational demands hinder deployment. Distilling LLM reasoning into Smaller Language Models (SLMs, ≤ 1B parameters) is a potential solution, yet these models often struggle with calculation and semantic errors. Previous work introduced Program-of-Thought Distillation (PoTD) to reduce calculation mistakes. To tackle semantic errors, we propose Key-Point-Driven Mathematical Reasoning Distillation (KPDD), which improves SLM reasoning by splitting the problem-solving process into Key Points Extraction and Step-by-Step Solution. KPDD includes KPDD-CoT, generating Chain-of-Thought rationales, and KPDD-PoT, producing Program-of-Thought rationales. Experiments show KPDD-CoT enhances reasoning capabilities, while KPDD-PoT achieves state-of-the-art performance in mathematical tasks, effectively reducing misunderstanding errors and promoting the deployment of efficient, capable SLMs. Xunyu Zhu, Jian Li 0040, Rong Yin 0001, Can Ma, Weiping Wang 0005 |
IJCNN | 2 |
| 2025 | Improving Mathematical Reasoning Capabilities of Small Language Models via Feedback-Driven DistillationabstractLarge Language Models (LLMs) demonstrate exceptional reasoning capabilities, often achieving state-of-the-art performance in various tasks. However, their substantial computational and memory demands, due to billions of parameters, hinder deployment in resource-constrained environments. A promising solution is knowledge distillation, where LLMs transfer reasoning capabilities to Small Language Models (SLMs, ≤ 1B parameters), enabling wider deployment on low-resource devices. Existing methods primarily focus on generating high-quality reasoning rationales for distillation datasets but often neglect the critical role of data quantity and quality. To address these challenges, we propose a Feedback-Driven Distillation (FDD) framework to enhance SLMs’ mathematical reasoning capabilities. In the initialization stage, a distillation dataset is constructed by prompting LLMs to pair mathematical problems with corresponding reasoning rationales. We classify problems into easy and hard categories based on SLM performance. For easy problems, LLMs generate more complex variations, while for hard problems, new questions of similar complexity are synthesized. In addition, we propose a multi-round distillation paradigm to iteratively enrich the distillation datasets, thereby progressively improving the mathematical reasoning abilities of SLMs. Experimental results demonstrate that our method can make SLMs achieve SOTA mathematical reasoning performance. Xunyu Zhu, Jian Li 0040, Rong Yin 0001, Can Ma, Weiping Wang 0005 |
IJCNN | 2 |
| 2025 | Optimal Convergence for Agnostic Kernel Learning With Random FeaturesabstractOwing to their solid theoretical guarantees and flexible learning framework, random features (RFs) methods have drawn increasing attention in the field of nonparametric statistical learning. However, existing studies on RFs assume that the target function lies exactly in the associated kernel space, which may not hold true in practical applications. In this article, we investigate the effectiveness of RFs in an agnostic setting that the target regression may be out of the kernel space and prove that they can still achieve capacity-dependent statistical optimality. To achieve this, we provide a finer grained estimate for the capacity of the hypothesis space, and conduct a refined analysis of error terms after a concise error decomposition. Our results show that RF with uniform sampling can guarantee optimality in half of the agnostic situations, while RF with data-dependent sampling can achieve optimal rates in the entire agnostic setting. This finding suggests that using data-dependent sampling not only reduces the number of RFs but also improves their applicability in agnostic settings. Finally, we compare the performance of RFs with different sampling strategies on several real-world datasets. The experimental results provide supports for our theoretical findings. Jian Li 0040, Yong Liu 0018, Weiping Wang 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | High-Dimensional Analysis for Generalized Nonlinear Regression: From Asymptotics to AlgorithmabstractOverparameterization often leads to benign overfitting, where deep neural networks can be trained to overfit the training data but still generalize well on unseen data. However, it lacks a generalized asymptotic framework for nonlinear regressions and connections to conventional complexity notions. In this paper, we propose a generalized high-dimensional analysis for nonlinear regression models, including various nonlinear feature mapping methods and subsampling. Specifically, we first provide an implicit regularization parameter and asymptotic equivalents related to a classical complexity notion, i.e., effective dimension. We then present a high-dimensional analysis for nonlinear ridge regression and extend it to ridgeless regression in the under-parameterized and over-parameterized regimes, respectively. We find that the limiting risks decrease with the effective dimension. Motivated by these theoretical findings, we propose an algorithm, namely RFRed, to improve generalization ability. Finally, we validate our theoretical findings and the proposed algorithm through several experiments. Jian Li 0040, Yong Liu 0018, Weiping Wang 0005 |
AAAI | 1 |
| 2024 | FedNS: A Fast Sketching Newton-Type Algorithm for Federated LearningabstractRecent Newton-type federated learning algorithms have demonstrated linear convergence with respect to the communication rounds. However, communicating Hessian matrices is often unfeasible due to their quadratic communication complexity. In this paper, we introduce a novel approach to tackle this issue while still achieving fast convergence rates. Our proposed method, named as Federated Newton Sketch methods (FedNS), approximates the centralized Newton's method by communicating the sketched square-root Hessian instead of the exact Hessian. To enhance communication efficiency, we reduce the sketch size to match the effective dimension of the Hessian matrix. We provide convergence analysis based on statistical learning for the federated Newton sketch approaches. Specifically, our approaches reach super-linear convergence rates w.r.t. the communication rounds for the first time. We validate the effectiveness of our algorithms through various experiments, which coincide with our theoretical findings. Jian Li 0040, Yong Liu 0018, Weiping Wang 0005 |
AAAI | 1 |
| 2024 | Towards sharper excess risk bounds for differentially private pairwise learning
Yilin Kang 0002, Jian Li 0040, Yong Liu 0018, Weiping Wang 0005 |
Neurocomputing | 2 |
| 2024 | Distilling mathematical reasoning capabilities into Small Language Models
Xunyu Zhu, Jian Li 0040, Yong Liu 0018, Can Ma, Weiping Wang 0005 |
Neural Networks | 2 |
| 2024 | Optimal Rates for Agnostic Distributed LearningabstractThe existing optimal rates for distributed kernel ridge regression (DKRR) often rely on a strict assumption, assuming that the true concept belongs to the hypothesis space. However, agnostic distributed learning is more common in practice, where the target regression may lie outside the kernel space. In this paper, we refine the excess risk bounds for DKRR and demonstrate that DKRR still achieve capacity-dependent optimal rates in the agnostic setting. Our theoretical findings indicate that the condition on the number of partitions not only influences computational efficiency but also impacts the range of situations where optimal rates are applicable. To relax the strict condition on the number of partitions, we first derive a sharper estimate for the difference between empirical and expected covariance operators. We then leverage additional unlabeled examples to reduce the label-independent error terms, further extending the optimal rates to more situations in the agnostic setting. In addition to the generalization error bounds in expectation, we also present refined excess risk bounds in high probability, where the optimal rates can also pertain to the agnostic setting. Finally, through both theoretical and empirical comparisons with related work, we demonstrate that our findings provide higher statistical applicability and computational advantages. Jian Li 0040, Yong Liu 0018, Weiping Wang 0005 |
IEEE Trans. Inf. Theory | 1 |
| 2024 | Non-IID Federated Learning With Sharper Risk BoundabstractIn federated learning (FL), the not independently or identically distributed (non-IID) data partitioning impairs the performance of the global model, which is a severe problem to be solved. Despite the extensive literature related to the algorithmic novelties and optimization analysis of FL, there has been relatively little theoretical research devoted to studying the generalization performance of non-IID FL. The generalization research of non-IID FL still lack effective tools and analytical approach. In this article, we propose weighted local Rademacher complexity to pertinently analyze the generalization properties of non-IID FL and derive a sharper excess risk bound based on weighted local Rademacher complexity, where the convergence rate is much faster than the existing bounds. Based on the theoretical results, we present a general framework federated averaging with local rademacher complexity (FedALRC) to lower the excess risk without additional communication costs compared to some famous methods, such as FedAvg. Through extensive experiments, we show that FedALRC outperforms FedAvg, FedProx and FedNova, and those experimental results coincide with our theoretical findings. Bojian Wei, Jian Li 0040, Yong Liu 0018, Weiping Wang 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Optimal Convergence Rates for Agnostic Nyström Kernel LearningabstractNyström low-rank approximation has shown great potential in processing large-scale kernel matrix and neural networks. However, there lacks a unified analysis for Nyström approximation, and the asymptotical minimax optimality for Nyström methods usually require a strict condition, assuming that the target regression lies exactly in the hypothesis space. In this paper, to tackle these problems, we provide a refined generalization analysis for Nyström approximation in the agnostic setting, where the target regression may be out of the hypothesis space. Specifically, we show Nyström approximation can still achieve the capacity-dependent optimal rates in the agnostic setting. To this end, we first prove the capacity-dependent optimal guarantees of Nyström approximation with the standard uniform sampling, which covers both loss functions and applies to some agnostic settings. Then, using data-dependent sampling, for example, leverage scores sampling, we derive the capacity-dependent optimal rates that apply to the whole range of the agnostic setting. To our best knowledge, the capacity-dependent optimality for the whole range of the agnostic setting is first achieved and novel in Nyström approximation. Jian Li 0040, Yong Liu 0018, Weiping Wang 0005 |
ICML | 1 |
| 2023 | Towards Sharp Analysis for Distributed Learning with Random FeaturesabstractIn recent studies, the generalization properties for distributed learning and random features assumed the existence of the target concept over the hypothesis space. However, this strict condition is not applicable to the more common non-attainable case. In this paper, using refined proof techniques, we first extend the optimal rates for distributed learning with random features to the non-attainable case. Then, we reduce the number of required random features via data-dependent generating strategy, and improve the allowed number of partitions with additional unlabeled data. Theoretical analysis shows these techniques remarkably reduce computational cost while preserving the optimal generalization accuracy under standard assumptions. Finally, we conduct several experiments on both simulated and real-world datasets, and the empirical results validate our theoretical findings. Jian Li 0040, Yong Liu 0018 |
IJCAI | 1 |
| 2023 | Towards Sharper Risk Bounds for Agnostic Multi-objective LearningabstractMany real-world machine learning tasks have mul-tiple objectives, such as multi-object detection and product recommendation, which can not be optimized directly through a single objective function. Fortunately, multi-objective learning can be used to solve this problem efficiently by some vector-valued algorithms. Recently, researchers find that the performance of multi-objective learning will be impaired when the mixture weights are unknown, where a fixed algorithm is difficult to select the optimal model in the hypothesis space. Thus, agnostic multi-objective learning has been proposed, which provides an effective approach to solve the problem of simultaneously optimizing multiple objectives with unknown mixture weights. In this way, a proper model will be selected because the agnostic multi-objective learning can improve the worst case of the hypothesis space. However, the current generalization error bounds for agnostic multi-objective learning can not converge faster than$\mathcal{O}(1/\sqrt{n})$, which limits the generalization guarantee. In this paper, we provide a sharper excess risk bound for agnostic multi-objective learning with convergence rate of$\mathcal{O}(1/n)$, which is much faster than the existing results and matches the best theoretical results of centralized learning. Based on our theory, we then propose a novel algorithm to improve the generalization performance of agnostic multi-objective learning. Bojian Wei, Jian Li 0040, Weiping Wang 0005 |
IJCNN | 2 |
| 2023 | Optimal Convergence Rates for Distributed Nystroem ApproximationabstractThe distributed kernel ridge regression (DKRR) has shown great potential in processing complicated tasks. However, DKRR only made use of the local samples that failed to capture the global characteristics. Besides, the existing optimal learning guarantees were provided in expectation and only pertain to the attainable case that the target regression lies exactly in the kernel space. In this paper, we propose distributed learning with globally-shared Nystroem centers (DNystroem), which utilizes global information across the local clients. We also study the statistical properties of DNystroem in expectation and in probability, respectively, and obtain several state-of-the-art results with the minimax optimal learning rates. Note that, the optimal convergence rates for DNystroem pertain to the non-attainable case, while the statistical results allow more partitions and require fewer Nystroem centers. Finally, we conduct experiments on several real-world datasets to validate the effectiveness of the proposed algorithm, and the empirical results coincide with our theoretical findings. Jian Li 0040, Yong Liu 0018, Weiping Wang 0005 |
J. Mach. Learn. Res. | 1 |
| 2023 | Improving Differentiable Architecture Search via self-distillation
Xunyu Zhu, Jian Li 0040, Yong Liu 0018, Weiping Wang 0005 |
Neural Networks | 2 |
| 2023 | Semi-supervised vector-valued learning: Improved bounds and algorithms
Jian Li 0040, Yong Liu 0018, Weiping Wang 0005 |
Pattern Recognit. | 1 |
| 2022 | Sharper Utility Bounds for Differentially Private Models: Smooth and Non-smoothabstractIn this paper, by introducing Generalized Bernstein condition, we propose the first O(√p over n∈ ) high probability excess population risk bound for differentially private algorithms under the assumptions G-Lipschitz, L-smooth, and Polyak-Łojasiewicz condition, based on gradient perturbation method. If we replace the properties G-Lipschitz and L-smooth by α-Hölder smoothness (which can be used in non-smooth setting), the high probability bound comes to O(n-α over 1+2α) w.r.t n, which cannot achieve O (1/n) when α ∈(0,1]. To solve this problem, we propose a variant of gradient perturbation method, max1,g -Normalized Gradient Perturbation (m-NGP). We further show that by normalization, the high probability excess population risk bound under assumptions α-Hölder smooth and Polyak-Łojasiewicz condition can achieve O (√p over n∈), which is the first O (1/n) high probability excess population risk bound w.r.t n for differentially private algorithms under non-smooth conditions. Moreover, experimental results show that m-NGP improves the performance of the differentially private model over real datasets. Yilin Kang 0002, Yong Liu 0018, Jian Li 0040, Weiping Wang 0005 |
CIKM | 3 |
| 2022 | Ridgeless Regression with Random FeaturesabstractRecent theoretical studies illustrated that kernel ridgeless regression can guarantee good generalization ability without an explicit regularization. In this paper, we investigate the statistical properties of ridgeless regression with random features and stochastic gradient descent. We explore the effect of factors in the stochastic gradient and random features, respectively. Specifically, random features error exhibits the double-descent curve. Motivated by the theoretical findings, we propose a tunable kernel algorithm that optimizes the spectral density of kernel during training. Our work bridges the interpolation theory and practical algorithm. Jian Li 0040, Yong Liu 0018 |
IJCAI | 1 |
| 2022 | Non-IID Distributed Learning with Optimal Mixture Weights
Jian Li 0040, Bojian Wei, Yong Liu 0018, Weiping Wang 0005 |
ECML/PKDD (4) | 1 |
| 2022 | Convolutional spectral kernel learning with generalization guarantees
Jian Li 0040, Yong Liu 0018, Weiping Wang 0005 |
Artif. Intell. | 1 |
| 2021 | Operation-level Progressive Differentiable Architecture SearchabstractDifferentiable Neural Architecture Search (DARTS) is becoming more and more popular among Neural Architecture Search (NAS) methods because of its high search efficiency and low compute cost. However, the stability of DARTS is very inferior, especially skip connections aggregation that leads to performance collapse. Though existing methods leverage Hessian eigenvalues to alleviate skip connections aggregation, they make DARTS unable to explore architectures with better performance. In the paper, we propose operation-level progressive differentiable neural architecture search (OPP-DARTS) to avoid skip connections aggregation and explore better architectures simultaneously. We first divide the search process into several stages during the search phase and increase candidate operations into the search space progressively at the beginning of each stage. It can effectively alleviate the unfair competition between operations during the search phase of DARTS by offsetting the inherent unfair advantage of the skip connection over other operations. Besides, to keep the competition between operations relatively fair and select the operation from the candidate operations set that makes training loss of the supernet largest. The experiment results indicate that our method is effective and efficient. Our method’s performance on CIFAR-10 is superior to the architecture found by standard DARTS, and the transferability of our method also surpasses standard DARTS. We further demonstrate the robustness of our method on three simple search spaces, i.e., S2, S3, S4, and the results show us that our method is more robust than standard DARTS. Our code is available at https://github.com/zxunyu/OPP-DARTS. Xunyu Zhu, Jian Li 0040, Yong Liu 0018, Weiping Wang 0005 |
ICDM | 2 |
| 2021 | Federated Learning for Non-IID Data: From Theory to Algorithm
Bojian Wei, Jian Li 0040, Yong Liu 0018, Weiping Wang 0005 |
PRICAI (1) | 2 |
| 2020 | Automated Spectral Kernel LearningabstractThe generalization performance of kernel methods is largely determined by the kernel, but spectral representations of stationary kernels are both input-independent and output-independent, which limits their applications on complicated tasks. In this paper, we propose an efficient learning framework that incorporates the process of finding suitable kernels and model training. Using non-stationary spectral kernels and backpropagation w.r.t. the objective, we obtain favorable spectral representations that depends on both inputs and outputs. Further, based on Rademacher complexity, we derive data-dependent generalization error bounds, where we investigate the effect of those factors and introduce regularization terms to improve the performance. Extensive experimental results validate the effectiveness of the proposed algorithm and coincide with our theoretical findings. Jian Li 0040, Yong Liu 0018, Weiping Wang 0005 |
AAAI | 1 |
| 2019 | Multi-Class Learning using Unlabeled Samples: Theory and AlgorithmabstractIn this paper, we investigate the generalization performance of multi-class classification, for which we obtain a shaper error bound by using the notion of local Rademacher complexity and additional unlabeled samples, substantially improving the state-of-the-art bounds in existing multi-class learning methods. The statistical learning motivates us to devise an efficient multi-class learning framework with the local Rademacher complexity and Laplacian regularization. Coinciding with the theoretical analysis, experimental results demonstrate that the stated approach achieves better performance. Jian Li 0040, Yong Liu 0018, Rong Yin 0001, Weiping Wang 0005 |
IJCAI | 1 |
| 2019 | Approximate Manifold Regularization: Scalable Algorithm and Generalization AnalysisabstractGraph-based semi-supervised learning is one of the most popular and successful semi-supervised learning approaches. Unfortunately, it suffers from high time and space complexity, at least quadratic with the number of training samples. In this paper, we propose an efficient graph-based semi-supervised algorithm with a sound theoretical guarantee. The proposed method combines Nystrom subsampling and preconditioned conjugate gradient descent, substantially improving computational efficiency and reducing memory requirements. Extensive empirical results reveal that our method achieves the state-of-the-art performance in a short time even with limited computing resources. Jian Li 0040, Yong Liu 0018, Rong Yin 0001, Weiping Wang 0005 |
IJCAI | 1 |
| 2018 | Multi-Class Learning: From Theory to AlgorithmabstractIn this paper, we study the generalization performance of multi-class classification and obtain a shaper data-dependent generalization error bound with fast convergence rate, substantially improving the state-of-art bounds in the existing data-dependent generalization analysis. The theoretical analysis motivates us to devise two effective multi-class kernel learning algorithms with statistical guarantees. Experimental results show that our proposed methods can significantly outperform the existing multi-class classification methods. Jian Li 0040, Yong Liu 0018, Rong Yin 0001, Hua Zhang 0008, Lizhong Ding 0001, Weiping Wang 0005 |
NeurIPS | 1 |
| 2017 | Efficient Kernel Selection via Spectral AnalysisabstractKernel selection is a fundamental problem of kernel methods. Existing measures for kernel selection either provide less theoretical guarantee or have high computational complexity. In this paper, we propose a novel kernel selection criterion based on a newly defined spectral measure of a kernel matrix, with sound theoretical foundation and high computational efficiency. We first show that the spectral measure can be used to derive generalization bounds for some kernel-based algorithms. By minimizing the derived generalization bounds, we propose the kernel selection criterion with spectral measure. Moreover, we demonstrate that the popular minimum graph cut and maximum mean discrepancy are two special cases of the proposed criterion. Experimental results on lots of data sets show that our proposed criterion can not only give the comparable results as the state-of-the-art criterion, but also significantly improve the efficiency. Jian Li 0040, Yong Liu 0018, Hailun Lin, Yinliang Yue, Weiping Wang 0005 |
IJCAI | 1 |