VLDB 2026 Research / reviewers in the wild / expert
Yiyuan She
dblp:42/4448
· DBLP profile ↗
14ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0002-5110-3179ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 since 2021Theory of computation · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorComputer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mind the jumps: A scalable robust local Gaussian process for multidimensional response surfaces with discontinuitiesabstractModeling response surfaces with abrupt jumps and discontinuities remains a major challenge across scientific and engineering domains. Although Gaussian process models excel at capturing smooth nonlinear relationships, their stationarity assumptions limit their ability to adapt to sudden input–output variations. Existing nonstationary extensions, particularly those based on domain partitioning, often struggle with boundary inconsistencies, sensitivity to outliers, and scalability issues in higher-dimensional settings, leading to reduced predictive accuracy and unreliable parameter estimation. To address the challenges posed by data heterogeneities and high dimensions, this paper proposes the Robust Local Gaussian Process (RLGP) model, a novel framework that integrates adaptive nearest-neighbor selection with a sparsity-driven robustification mechanism. Unlike existing methods, RLGP leverages an optimization-based mean-shift robustification after a multivariate perspective transformation combined with local neighborhood modeling to mitigate the influence of outliers. This approach enhances predictive accuracy near discontinuities while improving resistance to data heterogeneity. Comprehensive evaluations on real-world datasets show that RLGP consistently delivers high predictive accuracy and maintains competitive computational efficiency, especially in scenarios with sharp transitions and complex response structures. Scalability tests further confirm RLGP’s stability and reliability in higher-dimensional settings, where other methods falter. These outcomes establish RLGP as an effective and practical solution for modeling nonstationary and discontinuous response surfaces, applicable across a wide range of real-world scenarios. Isaac Adjetey, Yiyuan She |
Neurocomputing | 2 |
| 2024 | Nonuniform Sampling Pattern Design for Compressed Spectrum Sensing in Mobile Cognitive Radio NetworksabstractCompressed spectrum sensing (CSS) plays a pivotal role in dynamic spectrum access within mobile cognitive radio networks by offering reduced power consumption and lower hardware costs. The multicoset sampler, a well-known implementation for periodic nonuniform sampling, has been widely studied and is considered a promising architecture for realizing CSS. This paper focuses on the design of the multicoset sampling pattern, aiming at enhancing the isometry property of the sensing matrix. Unlike previous studies which assume a noise-free setup, our work considers the problem in a real-world environment with noise. First, we propose a deterministic algorithm for sampling pattern generation, particularly for specific hardware setup parameters. This algorithm offers strict mutual-coherence control in the multicoset sensing matrix. To address more general hardware configurations, we propose two optimization algorithms. One of them searches for nearly optimal sampling patterns through a random search strategy, while the other employs a greedy pursuit strategy to find a local optimizer. Furthermore, we propose an algorithm to iteratively optimize the sampling pattern between consecutive spectrum sensing windows by minimizing a restricted version of mutual coherence. The excellent performance of our proposed algorithms has been demonstrated through numerical experiments and has been verified on a self-developed hardware platform. Zihang Song, Yiyuan She, Jian Yang 0021, Jinbo Peng, Yue Gao 0001, Rahim Tafazolli |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | Slow Kill for Big Data LearningabstractBig-data applications often involve a vast number of observations and features, creating new challenges for variable selection and parameter estimation. This paper presents a novel technique called “slow kill,” which utilizes nonconvex constrained optimization, adaptive$\ell _{2}$-shrinkage, and increasing learning rates. The fact that the problem size can decrease during the slow kill iterations makes it particularly effective for large-scale variable screening. The interaction between statistics and optimization provides valuable insights into controlling quantiles, stepsize, and shrinkage parameters in order to relax the regularity conditions required to achieve the desired level of statistical accuracy. Experimental results on real and synthetic data show that slow kill outperforms state-of-the-art algorithms in various situations while being computationally efficient for large-scale data. Yiyuan She, Jiahui Shen, Adrian Barbu |
IEEE Trans. Inf. Theory | 1 |
| 2021 | Network Pruning via Annealing and Direct Sparsity ControlabstractArtificial neural networks (ANNs) especially deep convolutional neural networks are very popular these days and have been proved to successfully offer quite reliable solutions to many vision problems. However, the use of deep neural networks is widely impeded by their intensive computational and memory cost. In this paper, we propose a novel efficient network pruning framework that is suitable for both non-structured and structured channel-level pruning. Our proposed method tightens a sparsity constraint by gradually removing network parameters or filter channels based on a criterion and a schedule. The attractive fact that the network size keeps dropping throughout the iterations makes it suitable for the pruning of any untrained or pre-trained network. Because our method uses a$L_{0}$constraint instead of the$L_{1}$penalty, it does not introduce any bias in the training parameters or filter channels. Furthermore, the$L_{0}$constraint makes it easy to directly specify the desired sparsity level during the network pruning process. Finally, experimental validation on extensive synthetic and real vision datasets show that the proposed method obtains better or competitive performance compared to other states of art network pruning methods. Yangzi Guo, Yiyuan She, Adrian Barbu |
IJCNN | 2 |
| 2018 | Reinforced Robust Principal Component PursuitabstractHigh-dimensional data present in the real world is often corrupted by noise and gross outliers. Principal component analysis (PCA) fails to learn the true low-dimensional subspace in such cases. This is the reason why robust versions of PCA, which put a penalty on arbitrarily large outlying entries, are preferred to perform dimension reduction. In this paper, we argue that it is necessary to study the presence of outliers not only in the observed data matrix but also in the orthogonal complement subspace of the authentic principal subspace. In fact, the latter can seriously skew the estimation of the principal components. A reinforced robustification of principal component pursuit is designed in order to cater to the problem of finding out both types of outliers and eliminate their influence on the final subspace estimation. Simulation results under different design situations clearly show the superiority of our proposed method as compared with other popular implementations of robust PCA. This paper also showcases possible applications of our method in critically tough scenarios of face recognition and video background subtraction. Along with approximating a usable low-dimensional subspace from real-world data sets, the technique can capture semantically meaningful outliers. Pratik Prabhanjan Brahma, Yiyuan She, Jiade Li, Dapeng Oliver Wu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Feature Selection with Annealing for Computer Vision and Big Data LearningabstractMany computer vision and medical imaging problems are faced with learning from large-scale datasets, with millions of observations and features. In this paper we propose a novel efficient learning scheme that tightens a sparsity constraint by gradually removing variables based on a criterion and a schedule. The attractive fact that the problem size keeps dropping throughout the iterations makes it particularly suitable for big data learning. Our approach applies generically to the optimization of any differentiable loss function, and finds applications in regression, classification and ranking. The resultant algorithms build variable screening into estimation and are extremely simple to implement. We provide theoretical guarantees of convergence and selection consistency. In addition, one dimensional piecewise linear response functions are used to account for nonlinearity and a second order prior is imposed on these functions to avoid overfitting. Experiments on real and synthetic data show that the proposed method compares very well with other state of the art methods in regression, classification and ranking while being computationally very efficient and scalable. Adrian Barbu, Yiyuan She, Liangjing Ding, Gary Gramajo |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Why Deep Learning Works: A Manifold Disentanglement PerspectiveabstractDeep hierarchical representations of the data have been found out to provide better informative features for several machine learning applications. In addition, multilayer neural networks surprisingly tend to achieve better performance when they are subject to an unsupervised pretraining. The booming of deep learning motivates researchers to identify the factors that contribute to its success. One possible reason identified is the flattening of manifold-shaped data in higher layers of neural networks. However, it is not clear how to measure the flattening of such manifold-shaped data and what amount of flattening a deep neural network can achieve. For the first time, this paper provides quantitative evidence to validate the flattening hypothesis. To achieve this, we propose a few quantities for measuring manifold entanglement under certain assumptions and conduct experiments with both synthetic and real-world data. Our experimental results validate the proposition and lead to new insights on deep learning. Pratik Prabhanjan Brahma, Dapeng Oliver Wu, Yiyuan She |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Manifold-Regularized Selectable Factor Extraction for Semi-supervised Image Classification
Chao Zhang 0001, Fangyun Wei, Hongyang Zhang 0001, Yiyuan She |
BMVC | 5 |
| 2014 | The Group Square-Root Lasso: Theoretical Properties and Fast AlgorithmsabstractWe introduce and study the group square-root lasso (GSRL) method for estimation in high dimensional sparse regression models with group structure. The new estimator minimizes the square root of the residual sum of squares plus a penalty term proportional to the sum of the Euclidean norms of groups of the regression parameter vector. The net advantage of the method over the existing group lasso-type procedures consists in the form of the proportionality factor used in the penalty term, which for GSRL is independent of the variance of the error terms. This is of crucial importance in models with more parameters than the sample size, when estimating the variance of the noise becomes as difficult as the original problem. We show that the GSRL estimator adapts to the unknown sparsity of the regression vector, and has the same optimal estimation and prediction accuracy as the GL estimators, under the same minimal conditions on the model. This extends the results recently established for the square-root lasso, for sparse regression without group structure. Moreover, as a new type of result for square-root lasso methods, with or without groups, we study correct pattern recovery, and show that it can be achieved under conditions similar to those needed by the lasso or group-lasso-type methods, but with a simplified tuning strategy. We implement our method via a new algorithm, with proved convergence properties, which, unlike existing methods, scales well with the dimension of the problem. Our simulation studies support strongly our theoretical findings. Florentina Bunea, Johannes Lederer, Yiyuan She |
IEEE Trans. Inf. Theory | 3 |
| 2013 | A comparison of typical ℓp minimization algorithms
Qin Lyu, Zhouchen Lin, Yiyuan She, Chao Zhang 0001 |
Neurocomputing | 3 |
| 2013 | Stationary-sparse causality network learning
Yuejia He, Yiyuan She, Dapeng Oliver Wu |
J. Mach. Learn. Res. | 2 |
| 2010 | Approximating Higher-Order Distances Using Random Projections
Ping Li 0001, Michael W. Mahoney, Yiyuan She |
UAI | 3 |
| 2009 | Resolving deconvolution ambiguity in gene alternative splicingabstractBACKGROUND: For many gene structures it is impossible to resolve intensity data uniquely to establish abundances of splice variants. This was empirically noted by Wang et al. in which it was called a "degeneracy problem". The ambiguity results from an ill-posed problem where additional information is needed in order to obtain an unique answer in splice variant deconvolution. RESULTS: In this paper, we analyze the situations under which the problem occurs and perform a rigorous mathematical study which gives necessary and sufficient conditions on how many and what type of constraints are needed to resolve all ambiguity. This analysis is generally applicable to matrix models of splice variants. We explore the proposal that probe sequence information may provide sufficient additional constraints to resolve real-world instances. However, probe behavior cannot be predicted with sufficient accuracy by any existing probe sequence model, and so we present a Bayesian framework for estimating variant abundances by incorporating the prediction uncertainty from the micro-model of probe responsiveness into the macro-model of probe intensities. CONCLUSION: The matrix analysis of constraints provides a tool for detecting real-world instances in which additional constraints may be necessary to resolve splice variants. While purely mathematical constraints can be stated without error, real-world constraints may themselves be poorly resolved. Our Bayesian framework provides a generic solution to the problem of uniquely estimating transcript abundances given additional constraints that themselves may be uncertain, such as regression fit to probe sequence models. We demonstrate the efficacy of it by extensive simulations as well as various biological data. Yiyuan She, Earl Hubbell, Hui Wang 0007 |
BMC Bioinform. | 1 |
| 2004 | Block TERM factorization of block matrices
Yiyuan She, Pengwei Hao |
Sci. China Ser. F Inf. Sci. | 1 |