Yingjie Wang 0007

dblp:33/6297-7 · also YingJie Wang 0007 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0002-1054-4197ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021
YearPublicationVenuePosition
2026 Causal-guided strength differential independence sample weighting for out-of-distribution generalization
Haoran Yu 0005, Weifeng Liu 0001, Yingjie Wang 0007, Baodi Liu, Dapeng Tao, Honglong Chen
Pattern Recognit.3
2025 Error Analysis Affected by Heavy-Tailed Gradients for Non-Convex Pairwise Stochastic Gradient Descent
abstract
In recent years, there have been a growing number of works studying the generalization properties of stochastic gradient descent (SGD) from the perspective of algorithmic stability. However, few of them devote to simultaneously studying the generalization and optimization for the non-convex setting, especially pairwise SGD with heavy-tailed gradient noise. This paper considers the impact of the heavy-tailed gradient noise obeying sub-Weibull distribution on the stability-based learning guarantees for non-convex pairwise SGD by investigating its generalization and optimization jointly. Specifically, based on two novel pairwise uniform model stability tools, we firstly bound the generalization error of pairwise SGD in the general non-convex setting after bridging the quantitative relationships between stability and generalization error. Then, we further consider the practical heavy-tailed sub-Weibull gradient noise condition to establish a refined generalization bound without the bounded gradient condition. Finally, sharper error bounds for generalization and optimization are built by introducing the gradient dominance condition. Comparing these results reveals that sub-Weibull gradient noise brings some positive dependencies on the heavy-tailed strength for generalization and optimization. Furthermore, we extend our analysis to the corresponding pairwise minibatch SGD and derive the first stability-based near-optimal generalization and optimization bounds which are consistent with many empirical observations.
Hong Chen 0004, Bin Gu 0001, Yingjie Wang 0007, Weifu Li
AAAI5
2025 IW-ViT: Independence-Driven Weighting Vision Transformer for out-of-distribution generalization
Weifeng Liu 0001, Haoran Yu 0005, Yingjie Wang 0007, Baodi Liu, Dapeng Tao, Honglong Chen
Pattern Recognit.3
2025 Generalization Bounds of Deep Neural Networks With τ-Mixing Samples
abstract
Deep neural networks (DNNs) have shown an astonishing ability to unlock the complicated relationships among the inputs and their responses. Along with empirical successes, some approximation analysis of DNNs has also been provided to understand their generalization performance. However, the existing analysis depends heavily on the independently identically distribution (i.i.d.) assumption of observations, which may be too ideal and often violated in real-world applications. To relax the i.i.d. assumption, this article develops the covering number-based concentration estimation to establish generalization bounds of DNNs with $\tau $ -mixing samples, where the dependency between samples is much general including $\alpha $ -mixing process as a special case. By assigning a specific parameter value to the $\tau $ -mixing process, our results are consistent with the existing convergence analysis under the i.i.d. case. Experiments on simulated data validate the theoretical findings.
Yaohui Chen 0002, Weifu Li, Yingjie Wang 0007, Bin Gu 0001, Feng Zheng 0001, Hong Chen 0004
IEEE Trans. Neural Networks Learn. Syst.4
2025 Sparse Additive Machine With the Correntropy-Induced Loss
abstract
Sparse additive machines (SAMs) have shown competitive performance on variable selection and classification in high-dimensional data due to their representation flexibility and interpretability. However, the existing methods often employ the unbounded or nonsmooth functions as the surrogates of 0-1 classification loss, which may encounter the degraded performance for data with outliers. To alleviate this problem, we propose a robust classification method, named SAM with the correntropy-induced loss (CSAM), by integrating the correntropy-induced loss (C-loss), the data-dependent hypothesis space, and the weighted -norm regularizer ( ) into additive machines. In theory, the generalization error bound is estimated via a novel error decomposition and the concentration estimation techniques, which shows that the convergence rate can be achieved under proper parameter conditions. In addition, the theoretical guarantee on variable selection consistency is analyzed. Experimental evaluations on both synthetic and real-world datasets consistently validate the effectiveness and robustness of the proposed approach.
Peipei Yuan, Xinge You, Hong Chen 0004, Yingjie Wang 0007, Qinmu Peng, Bin Zou 0002
IEEE Trans. Neural Networks Learn. Syst.4
2024 Towards Theoretical Understandings of Self-Consuming Generative Models
abstract
This paper tackles the emerging challenge of training generative models within a self-consuming loop, wherein successive generations of models are recursively trained on mixtures of real and synthetic data from previous generations. We construct a theoretical framework to rigorously evaluate how this training procedure impacts the data distributions learned by future models, including parametric and non-parametric models. Specifically, we derive bounds on the total variation (TV) distance between the synthetic data distributions produced by future models and the original real data distribution under various mixed training scenarios for diffusion models with a one-hidden-layer neural network score function. Our analysis demonstrates that this distance can be effectively controlled under the condition that mixed training dataset sizes or proportions of real data are large enough. Interestingly, we further unveil a phase transition induced by expanding synthetic data amounts, proving theoretically that while the TV distance exhibits an initial ascent, it declines beyond a threshold point. Finally, we present results for kernel density estimation, delivering nuanced insights such as the impact of mixed data training on error propagation.
Shi Fu, Sen Zhang 0006, Yingjie Wang 0007, Xinmei Tian 0001, Dacheng Tao
ICML3
2024 Discriminative Representation-Based Classifier for Few-Shot Remote Sensing Classification
Tianhao Yuan, Weifeng Liu 0001, Yingjie Wang 0007, Baodi Liu
PRCV (13)3
2023 Tilted Sparse Additive Models
abstract
Additive models have been burgeoning in data analysis due to their flexible representation and desirable interpretability. However, most existing approaches are constructed under empirical risk minimization (ERM), and thus perform poorly in situations where average performance is not a suitable criterion for the problems of interest, e.g., data with complex non-Gaussian noise, imbalanced labels or both of them. In this paper, a novel class of sparse additive models is proposed under tilted empirical risk minimization (TERM), which addresses the deficiencies in ERM by imposing tilted impact on individual losses, and is flexibly capable of achieving a variety of learning objectives, e.g., variable selection, robust estimation, imbalanced classification and multiobjective learning. On the theoretical side, a learning theory analysis which is centered around the generalization bound and function approximation error bound (under some specific data distributions) is conducted rigorously. On the practical side, an accelerated optimization algorithm is designed by integrating Prox-SVRG and random Fourier acceleration technique. The empirical assessments verify the competitive performance of our approach on both synthetic and real data.
Yingjie Wang 0007, Hong Chen 0004, Weifeng Liu 0001, Fengxiang He, Tieliang Gong, Youcheng Fu, Dacheng Tao
ICML1
2023 A Stable Vision Transformer for Out-of-Distribution Generalization
Haoran Yu 0005, Baodi Liu, Yingjie Wang 0007, Kai Zhang 0029, Dapeng Tao, Weifeng Liu 0001
PRCV (8)3
2023 Robust variable structure discovery based on tilted empirical risk minimization
Yingjie Wang 0007, Liangxuan Zhu, Hong Chen 0004, Lingjuan Wu
Appl. Intell.2
2022 Error-Based Knockoffs Inference for Controlled Feature Selection
abstract
Recently, the scheme of model-X knockoffs was proposed as a promising solution to address controlled feature selection under high-dimensional finite-sample settings. However, the procedure of model-X knockoffs depends heavily on the coefficient-based feature importance and only concerns the control of false discovery rate (FDR). To further improve its adaptivity and flexibility, in this paper, we propose an error-based knockoff inference method by integrating the knockoff features, the error-based feature importance statistics, and the stepdown procedure together. The proposed inference procedure does not require specifying a regression model and can handle feature selection with theoretical guarantees on controlling false discovery proportion (FDP), FDR, or k-familywise error rate (k-FWER). Empirical evaluations demonstrate the competitive performance of our approach on both simulated and real data.
Xuebin Zhao, Hong Chen 0004, Yingjie Wang 0007, Weifu Li, Tieliang Gong, Yulong Wang 0002, Feng Zheng 0001
AAAI3
2022 Huber Additive Models for Non-stationary Time Series Analysis
Yingjie Wang 0007, Xianrui Zhong, Fengxiang He, Hong Chen 0004, Dacheng Tao
ICLR1
2022 Distribution-dependent feature selection for deep neural networks
Xuebin Zhao, Weifu Li, Hong Chen 0004, Yingjie Wang 0007, Vijay John
Appl. Intell.4
2021 Distributed Ranking with Communications: Approximation Analysis and Applications
Hong Chen 0004, Yingjie Wang 0007, Yulong Wang 0002, Feng Zheng 0001
AAAI2
2021 Sparse additive machine with pinball loss
Yingjie Wang 0007, Hong Chen 0004, Tianjiao Yuan
Neurocomputing1
2021 Sparse Modal Additive Model
abstract
Sparse additive models have been successfully applied to high-dimensional data analysis due to the flexibility and interpretability of their representation. However, the existing methods are often formulated using the least-squares loss with learning the conditional mean, which is sensitive to data with the non-Gaussian noises, e.g., skewed noise, heavy-tailed noise, and outliers. To tackle this problem, we propose a new robust regression method, called as sparse modal additive model (SpMAM), by integrating the modal regression metric, the data-dependent hypothesis space, and the weightedlq,1-norm regularizer (q ≥ 1) into the additive models. Specifically, the modal regression metric assures the model robustness to complex noises via learning the conditional mode, the data-dependent hypothesis space offers the model adaptivity via sample-based presentation, and thelq,1-norm regularizer addresses the algorithmic interpretability via sparse variable selection. In theory, the proposed SpMAM enjoys statistical guarantees on asymptotic consistency for regression estimation and variable selection simultaneously. Experimental results on both synthetic and real-world benchmark data sets validate the effectiveness and robustness of the proposed model.
Hong Chen 0004, Yingjie Wang 0007, Feng Zheng 0001, Cheng Deng 0002, Heng Huang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2020 Multi-task Additive Models for Robust Estimation and Automatic Structure Discovery
abstract
Additive models have attracted much attention for high-dimensional regression estimation and variable selection. However, the existing models are usually limited to the single-task learning framework under the mean squared error (MSE) criterion, where the utilization of variable structure depends heavily on priori knowledge among variables. For high-dimensional observations in real environment, e.g., Coronal Mass Ejections (CMEs) data, the learning performance of previous methods may be degraded seriously due to the complex non-Gaussian noise and the insufficiency of prior knowledge on variable structure. To tackle this problem, we propose a new class of additive models, called Multi-task Additive Models (MAM), by integrating the mode-induced metric, the structure-based regularizer, and additive hypothesis spaces into a bilevel optimization framework. Our approach does not require any priori knowledge of variable structure and suits for high-dimensional data with complex noise, e.g., skewed noise, heavy-tailed noise, and outliers. A smooth iterative optimization algorithm with convergence guarantees is provided to implement MAM efficiently. Experiments on simulations and the CMEs analysis demonstrate the competitive performance of our approach for robust estimation and automatic structure discovery.
Yingjie Wang 0007, Hong Chen 0004, Feng Zheng 0001, Chen Xu 0007, Tieliang Gong
NeurIPS1