Yuesheng Xu

dblp:68/5920 · DBLP profile ↗
← Back
33ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 4 first-author · 7 since 2021Theory of computation · 8 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hypothesis spaces for deep learning
Rui Wang 0149, Yuesheng Xu, Mingsong Yan
Neural Networks2
2026 Computed Quantitative Planar Imaging for Targeted Alpha Therapy: Model-Based Sparse Reconstruction Validated With a Novel 225Ac Epoxy Phantom
abstract
Targeted Alpha Therapy (TAT), using alpha-emitting radionuclides (AER) such as225Ac, shows promise for the treatment of advanced and refractory cancers. Currently, TAT is prescribed on the basis of activity (e.g., MBq, kBq/kg), with no account taken of individual biodistribution or kinetics. The delivery of patient-specific treatment, based on absorbed dose criteria, requires in-vivo imaging of the AER biodistribution, a challenging scenario due to the scarcity of imageable photons. To address this, we present a novel computed quantitative planar (CQP) imaging method that reconstructs a coronal projection of the 3D AER distribution from anterior/posterior scintigraphy coregistered with CT. The model is regularized using maximum a posteriori estimation with sparse ℓ1 tight-framelet transforms and solved via a convergence-guaranteed fixed-point proximity algorithm. To experimentally evaluate our approach, we built a modular slab phantom containing a known distribution of225Ac vitrified in epoxy. CQP reconstruction was characterized by significantly reduced bias and noise, improved spatial resolution, and better signal-to-noise ratios, compared to geometric mean methods. The CQP approach is clinically implementable with conventional SPECT/CT systems, without need for hardware additions or modifications, and can assist dosimetry workflows, especially where 3D SPECT/PET is impractical.
Charles Ross Schmidtlein, Andrzej Król, Howard C. Gifford, Joseph O'Donoghue, Lisa Bodei, Yuesheng Xu
IEEE Trans. Medical Imaging7
2024 Addressing Spectral Bias of Deep Neural Networks by Multi-Grade Deep Learning
abstract
Deep neural networks (DNNs) have showcased their remarkable precision in approximating smooth functions. However, they suffer from the {\it spectral bias}, wherein DNNs typically exhibit a tendency to prioritize the learning of lower-frequency components of a function, struggling to effectively capture its high-frequency features. This paper is to address this issue. Notice that a function having only low frequency components may be well-represented by a shallow neural network (SNN), a network having only a few layers. By observing that composition of low frequency functions can effectively approximate a high-frequency function, we propose to learn a function containing high-frequency components by composing several SNNs, each of which learns certain low-frequency information from the given data. We implement the proposed idea by exploiting the multi-grade deep learning (MGDL) model, a recently introduced model that trains a DNN incrementally, grade by grade, a current grade learning from the residue of the previous grade only an SNN (with trainable parameters) composed with the SNNs (with fixed parameters) trained in the preceding grades as features. We apply MGDL to synthetic, manifold, colored images, and MNIST datasets, all characterized by presence of high-frequency features. Our study reveals that MGDL excels at representing functions containing high-frequency information. Specifically, the neural networks learned in each grade adeptly capture some low-frequency information, allowing their compositions with SNNs learned in the previous grades effectively representing the high-frequency features. Our experimental results underscore the efficacy of MGDL in addressing the spectral bias inherent in DNNs. By leveraging MGDL, we offer insights into overcoming spectral bias limitation of DNNs, thereby enhancing the performance and applicability of deep learning models in tasks requiring the representation of high-frequency information. This study confirms that the proposed method offers a promising solution to address the spectral bias of DNNs. The code is available on GitHub: \href{https://github.com/Ronglong-Fang/AddressingSpectralBiasviaMGDL}{\texttt{Addressing Spectral Bias via MGDL}}.
Ronglong Fang, Yuesheng Xu
NeurIPS2
2024 Convergence of deep ReLU networks
Yuesheng Xu, Haizhang Zhang
Neurocomputing1
2024 A duality approach to regularized learning problems in Banach spaces
Rui Wang 0096, Yuesheng Xu
J. Complex.3
2024 Sparse Representer Theorems for Learning in Reproducing Kernel Banach Spaces
abstract
Sparsity of a learning solution is a desirable feature in machine learning. Certain reproducing kernel Banach spaces (RKBSs) are appropriate hypothesis spaces for sparse learning methods. The goal of this paper is to understand what kind of RKBSs can promote sparsity for learning solutions. We consider two typical learning models in an RKBS: the minimum norm interpolation (MNI) problem and the regularization problem. We first establish an explicit representer theorem for solutions of these problems, which represents the extreme points of the solution set by a linear combination of the extreme points of the subdifferential set, of the norm function, which is data-dependent. We then propose sufficient conditions on the RKBS that can transform the explicit representation of the solutions to a sparse kernel representation having fewer terms than the number of the observed data. Under the proposed sufficient conditions, we investigate the role of the regularization parameter on sparsity of the regularized solutions. We further show that two specific RKBSs, the sequence space $\ell_1(\mathbb{N})$ and the measure space, can have sparse representer theorems for both MNI and regularization models.
Rui Wang 0096, Yuesheng Xu, Mingsong Yan
J. Mach. Learn. Res.2
2024 Uniform Convergence of Deep Neural Networks With Lipschitz Continuous Activation Functions and Variable Widths
abstract
We consider deep neural networks (DNNs) with a Lipschitz continuous activation function and with weight matrices of variable widths. We establish a uniform convergence analysis framework in which sufficient conditions on weight matrices and bias vectors together with the Lipschitz constant are provided to ensure uniform convergence of DNNs to a meaningful function as the number of their layers tends to infinity. In the framework, special results on uniform convergence of DNNs with a fixed width, bounded widths and unbounded widths are presented. In particular, as convolutional neural networks are special DNNs with weight matrices of increasing widths, we put forward conditions on the mask sequence which lead to uniform convergence of the resulting convolutional neural networks. The Lipschitz continuity assumption on the activation functions allows us to include in our theory most of commonly used activation functions in applications.
Yuesheng Xu, Haizhang Zhang
IEEE Trans. Inf. Theory1
2022 Convergence of deep convolutional neural networks
Yuesheng Xu, Haizhang Zhang
Neural Networks1
2022 A Fast Convergent Ordered-Subsets Algorithm With Subiteration-Dependent Preconditioners for PET Image Reconstruction
abstract
We investigated the imaging performance of a fast convergent ordered-subsets algorithm with subiteration-dependent preconditioners (SDPs) for positron emission tomography (PET) image reconstruction. In particular, we considered the use of SDP with the block sequential regularized expectation maximization (BSREM) approach with the relative difference prior (RDP) regularizer due to its prior clinical adaptation by vendors. Because the RDP regularization promotes smoothness in the reconstructed image, the directions of the gradients in smooth areas more accurately point toward the objective function's minimizer than those in variable areas. Motivated by this observation, two SDPs have been designed to increase iteration step-sizes in the smooth areas and reduce iteration step-sizes in the variable areas relative to a conventional expectation maximization preconditioner. The momentum technique used for convergence acceleration can be viewed as a special case of SDP. We have proved the global convergence of SDP-BSREM algorithms by assuming certain characteristics of the preconditioner. By means of numerical experiments using both simulated and clinical PET data, we have shown that the SDP-BSREM algorithms substantially improve the convergence rate, as compared to conventional BSREM and a vendor's implementation as Q.Clear. Specifically, SDP-BSREM algorithms converge 35%-50% faster in reaching the same objective function value than conventional BSREM and commercial Q.Clear algorithms. Moreover, we showed in phantoms with hot, cold and background regions that the SDP-BSREM algorithms approached the values of a highly converged reference image faster than conventional BSREM and commercial Q.Clear algorithms.
Charles Ross Schmidtlein, Andrzej Król, Si Li 0005, Yizun Lin, Sangtae Ahn, Charles W. Stearns, Yuesheng Xu
IEEE Trans. Medical Imaging8
2021 Certification and Trade-off of Multiple Fairness Criteria in Graph-based Spam Detection
abstract
Spamming reviews are prevalent in review systems to manipulate seller reputation and mislead customers. patterns to achieve state-of-the-art detection accuracy. The detection can influence a large number of real-world entities and it is ethical to treat different groups of entities as equally as possible. However, due to skewed distributions of the graphs, GNN can fail to meet diverse fairness criteria designed for different parties. We formulate linear systems of the input features and the adjacency matrix of the review graphs for the certification of multiple fairness criteria. When the criteria are competing, we relax the certification and design a multi-objective optimization (MOO) algorithm to explore multiple efficient trade-offs, so that no objective can be improved without harming another objective. We prove that the algorithm converges to a Pareto efficient solution using duality and the implicit function theorem. Since there can be exponentially many trade-offs of the criteria, we propose a data-driven stochastic search algorithm to approximate Pareto fronts consisting of multiple efficient trade-offs. Experimentally, we show that the algorithms converge to solutions that dominate baselines based on fairness regularization and adversarial training.
Kai Burkholder, Kenny Kwock, Yuesheng Xu, Sihong Xie
CIKM3
2021 Regularization in a functional reproducing kernel Hilbert space
Rui Wang 0096, Yuesheng Xu
J. Complex.2
2021 Representer Theorems in Banach Spaces: Minimum Norm Interpolation, Regularized Learning and Semi-Discrete Inverse Problems
abstract
Learning a function from a finite number of sampled data points (measurements) is a fundamental problem in science and engineering. This is often formulated as a minimum norm interpolation (MNI) problem, a regularized learning problem or, in general, a semi-discrete inverse problem (SDIP), in either Hilbert spaces or Banach spaces. The goal of this paper is to systematically study solutions of these problems in Banach spaces. We aim at obtaining explicit representer theorems for their solutions, on which convenient solution methods can then be developed. For the MNI problem, the explicit representer theorems enable us to express the infimum in terms of the norm of the linear combination of the interpolation functionals. For the purpose of developing efficient computational algorithms, we establish the fixed-point equation formulation of solutions of these problems. We reveal that unlike in a Hilbert space, in general, solutions of these problems in a Banach space may not be able to be reduced to truly finite dimensional problems (with certain infinite dimensional components hidden). We demonstrate how this obstacle can be removed, reducing the original problem to a truly finite dimensional one, in the special case when the Banach space is $\ell_1(\mathbb{N})$.
Rui Wang 0096, Yuesheng Xu
J. Mach. Learn. Res.2
2021 Synthesis of Mammogram From Digital Breast Tomosynthesis Using Deep Convolutional Neural Network With Gradient Guided cGANs
abstract
Synthetic digital mammography (SDM), a 2D image generated from digital breast tomosynthesis (DBT), is used as a potential substitute for full-field digital mammography (FFDM) in clinic to reduce the radiation dose for breast cancer screening. Previous studies exploited projection geometry and fused projection data and DBT volume, with different post-processing techniques applied on re-projection data which may generate different image appearance compared to FFDM. To alleviate this issue, one possible solution to generate an SDM image is using a learning-based method to model the transformation from the DBT volume to the FFDM image using current DBT/FFDM combo images. In this study, we proposed to use a deep convolutional neural network (DCNN) to learn the transformation to generate SDM using current DBT/FFDM combo images. Gradient guided conditional generative adversarial networks (GGGAN) objective function was designed to preserve subtle MCs and the perceptual loss was exploited to improve the performance of the proposed DCNN on perceptual quality. We used various image quality criteria for evaluation, including preserving masses and MCs which are important in mammogram. Experiment results demonstrated progressive performance improvement of network using different objective functions in terms of those image quality criteria. The methodology we exploited in the SDM generation task to analyze and progressively improve image quality by designing objective functions may be helpful to other image generation tasks.
Gongfa Jiang, Jun Wei 0002, Yuesheng Xu, Jiefang Wu, Genggeng Qin, Yao Lu 0007
IEEE Trans. Medical Imaging3
2020 Multiplicative Noise Removal: Nonlocal Low-Rank Model and Its Proximal Alternating Reweighted Minimization Algorithm
abstract
The goal of this paper is to develop a novel numerical method for efficient multiplicative noise removal. The nonlocal self-similarity of natural images implies that the matrices formed by their nonlocal similar patches are low-rank. By exploiting this low-rank prior with application to multiplicative noise removal, we propose a nonlocal low-rank model for this task and develop a proximal alternating reweighted minimization (PARM) algorithm to solve the optimization problem resulting from the model. Specifically, we utilize a generalized nonconvex surrogate of the rank function to regularize the patch matrices and develop a new nonlocal low-rank model, which is a nonconvex nonsmooth optimization problem having a patchwise data fidelity and a generalized nonlocal low-rank regularization term. To solve this optimization problem, we propose the PARM algorithm, which has a proximal alternating scheme with a reweighted approximation of its subproblem. A theoretical analysis of the proposed PARM algorithm is conducted to guarantee its global convergence to a critical point. Numerical experiments demonstrate that the proposed method for multiplicative noise removal significantly outperforms existing methods, such as the benchmark SAR-BM3D method, in terms of the visual quality of the denoised images, and of the peak-signal-to-noise ratio (PSNR) and the structural similarity index measure (SSIM) values.
Jian Lu 0002, Lixin Shen, Chen Xu 0004, Yuesheng Xu
SIAM J. Imaging Sci.5
2019 Synthesize Mammogram from Digital Breast Tomosynthesis with Gradient Guided cGANs
Gongfa Jiang, Yao Lu 0007, Jun Wei 0002, Yuesheng Xu
MICCAI (6)4
2019 A Higher-Order Polynomial Method for SPECT Reconstruction
abstract
Existing single-photon emission computed tomography (SPECT) reconstruction methods are mostly based on discrete models that may be viewed as piecewise constant approximations of a continuous data acquisition process. Due to low accuracy order of piecewise constant approximations, a traditional discrete model introduces irreducible model errors which are a bottleneck of the quality improvement of reconstructed images in clinical applications. To overcome this drawback, we develop a higher-order polynomial method for SPECT reconstruction. Specifically, we represent the data acquisition of SPECT imaging by using an integral equation model, approximate the solution of the underlying integral equation by higher-order piecewise polynomials leading to a new discrete system and introduce two novel regularizers for the system, by exploring the a priori knowledge of the radiotracer distribution, suitable for the approximation. The proposed higher-order polynomial method outperforms significantly the cutting edge reconstruction method based on a traditional discrete model in terms of model error reduction, noise suppression, and artifact reduction. In particular, the coefficient of variation of images reconstructed by the piecewise linear polynomial method is reduced by a factor of 10 in comparison to that of a traditional discrete model-based method.
Ying Jiang 0002, Si Li 0005, Yuesheng Xu
IEEE Trans. Medical Imaging3
2019 A Krasnoselskii-Mann Algorithm With an Improved EM Preconditioner for PET Image Reconstruction
abstract
This paper presents a preconditioned Krasnoselskii-Mann (KM) algorithm with an improved EM preconditioner (IEM-PKMA) for higher-order total variation (HOTV) regularized positron emission tomography (PET) image reconstruction. The PET reconstruction problem can be formulated as a three-term convex optimization model consisting of the Kullback-Leibler (KL) fidelity term, a nonsmooth penalty term, and a nonnegative constraint term which is also nonsmooth. We develop an efficient KM algorithm for solving this optimization problem based on a fixed-point characterization of its solution, with a preconditioner and a momentum technique for accelerating convergence. By combining the EM precondtioner, a thresholding, and a good inexpensive estimate of the solution, we propose an improved EM preconditioner that can not only accelerate convergence but also avoid the reconstructed image being "stuck at zero." Numerical results in this paper show that the proposed IEM-PKMA outperforms existing state-of-the-art algorithms including, the optimization transfer descent algorithm and the preconditioned L-BFGS-B algorithm for the differentiable smoothed anisotropic total variation regularized model, the preconditioned alternating projection algorithm, and the alternating direction method of multipliers for the nondifferentiable HOTV regularized model. Encouraging initial experiments using clinical data are presented.
Yizun Lin, Charles Ross Schmidtlein, Qia Li, Si Li 0005, Yuesheng Xu
IEEE Trans. Medical Imaging5
2017 H-BLAST: a fast protein sequence alignment toolkit on heterogeneous computers with GPUs
abstract
Motivation: The sequence alignment is a fundamental problem in bioinformatics. BLAST is a routinely used tool for this purpose with over 118 000 citations in the past two decades. As the size of bio-sequence databases grows exponentially, the computational speed of alignment softwares must be improved. Results: We develop the heterogeneous BLAST (H-BLAST), a fast parallel search tool for a heterogeneous computer that couples CPUs and GPUs, to accelerate BLASTX and BLASTP-basic tools of NCBI-BLAST. H-BLAST employs a locally decoupled seed-extension algorithm for better performance on GPUs, and offers a performance tuning mechanism for better efficiency among various CPUs and GPUs combinations. H-BLAST produces identical alignment results as NCBI-BLAST and its computational speed is much faster than that of NCBI-BLAST. Speedups achieved by H-BLAST over sequential NCBI-BLASTP (resp. NCBI-BLASTX) range mostly from 4 to 10 (resp. 5 to 7.2). With 2 CPU threads and 2 GPUs, H-BLAST can be faster than 16-threaded NCBI-BLASTX. Furthermore, H-BLAST is 1.5-4 times faster than GPU-BLAST. Availability and Implementation: https://github.com/Yeyke/H-BLAST.git. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Weicai Ye, Yongdong Zhang 0002, Yuesheng Xu
Bioinform.4
2014 Orthogonal polynomial expansions on sparse grids
Yanzhao Cao, Ying Jiang 0002, Yuesheng Xu
J. Complex.3
2012 Refinement of Operator-valued Reproducing Kernels
Haizhang Zhang, Yuesheng Xu
J. Mach. Learn. Res.2
2011 B-spline quasi-interpolation on sparse grids
Ying Jiang 0002, Yuesheng Xu
J. Complex.2
2010 Fast discrete algorithms for sparse Fourier expansions of high dimensional functions
Ying Jiang 0002, Yuesheng Xu
J. Complex.2
2010 Approximation of high-dimensional kernel matrices by multilevel circulant matrices
Guohui Song, Yuesheng Xu
J. Complex.2
2009 Reproducing kernel Banach spaces for machine learning
abstract
Reproducing kernel Hilbert space (RKHS) methods have become powerful tools in machine learning. However, their kernels, which measure similarity of inputs, are required to be symmetric, constraining certain applications in practice. Furthermore, the celebrated representer theorem only applies to regularizers induced by the norm of an RKHS. To remove these limitations, we introduce the notion of reproducing kernel Banach spaces (RKBS) for pairs of reflexive Banach spaces of functions by making use of semi-inner-products and the duality mapping. As applications, we develop the framework of RKBS standard learning schemes including minimal norm interpolation, regularization network, and support vector machines. In particular, existence, uniqueness and representer theorems are established.
Haizhang Zhang, Yuesheng Xu, Jun Zhang 0009
IJCNN2
2009 Optimal learning of bandlimited functions from localized sampling
Charles A. Micchelli, Yuesheng Xu, Haizhang Zhang
J. Complex.2
2009 Refinement of Reproducing Kernels
Yuesheng Xu, Haizhang Zhang
J. Mach. Learn. Res.1
2009 Reproducing Kernel Banach Spaces for Machine Learning
Haizhang Zhang, Yuesheng Xu, Jun Zhang 0009
J. Mach. Learn. Res.2
2008 Initial Experiences with the BEC Parallel Programming Environment
abstract
Bundle-exchange-compute (BEC) is a new virtual shared memory parallel programming environment for distributed-memory machines. Different from and complementary to other global address space (GAS) programming model research efforts, BEC has built-in efficient support for unstructured applications that inherently require high-volume random fine-grained communication, such as parallel graph algorithms, sparse-matrices, and large-scale physics simulations. In BEC, the global view of shared data structures enables ease of algorithm design and programming; and for good application performance, fine-grained (random) accesses to shared data are automatically and dynamically bundled together for coarse-grained message-passing. BEC frees the users from explicit management of data distribution, locality, and communication. Therefore, BEC is much easier to program than MPI, while achieving comparable application performance. This paper presents some initial BEC applications, which show that simple BEC programs can match very complex and highly optimized MPI codes.
Michael A. Heroux, Zhaofang Wen, Yuesheng Xu
ISPDC4
2007 Refinable Kernels
Yuesheng Xu, Haizhang Zhang
J. Mach. Learn. Res.1
2006 Universal Kernels
abstract
In this paper we investigate conditions on the features of a continuous kernel so that it may approximate an arbitrary continuous target function uniformly on any compact subset of the input space. A number of concrete examples are given of kernels with this universal approximating property.
Charles A. Micchelli, Yuesheng Xu, Haizhang Zhang
J. Mach. Learn. Res.2
2005 Deterministic convergence of an online gradient method for BP neural networks
abstract
Online gradient methods are widely used for training feedforward neural networks. We prove in this paper a convergence theorem for an online gradient method with variable step size for backward propagation (BP) neural networks with a hidden layer. Unlike most of the convergence results that are of probabilistic and nonmonotone nature, the convergence result that we establish here has a deterministic and monotone nature.
Wei Wu 0010, Guorui Feng, Zhengxue Li, Yuesheng Xu
IEEE Trans. Neural Networks4
2004 Recent Developments on Convergence of Online Gradient Methods for Neural Network Training
Wei Wu 0010, Zhengxue Li, Guorui Feng, Naimin Zhang, Dong Nan, Zhiqiong Shao, Jie Yang 0007, Yuesheng Xu
ISNN (1)9
1995 Degree reduction of Bézier curves by uniform approximation with endpoint interpolation
Przemyslaw Bogacki, Stanley E. Weinstein, Yuesheng Xu
Comput. Aided Des.3