Jack Xin

dblp:54/454 · DBLP profile ↗
← Back
32ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0002-6438-8476ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 5 since 2021Artificial intelligence and machine learning · 12 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Global Well-posedness and Convergence Analysis of Score-based Generative Models via Sharp Lipschitz Estimates
abstract
We establish global well-posedness and convergence of the score-based generative models (SGM) under minimal general assumptions of initial data for score estimation. For the smooth case, we start from a Lipschitz bound of the score function with optimal time length. The optimality is validated by an example whose Lipschitz constant of scores is bounded at initial but blows up in finite time. This necessitates the separation of time scales in conventional bounds for non-log-concave distributions. In contrast, our follow up analysis only relies on a local Lipschitz condition and is valid globally in time. This leads to the convergence of numerical scheme without time separation. For the non-smooth case, we show that the optimal Lipschitz bound is $O(1/t)$ in the point-wise sense for distributions supported on a compact, smooth and low-dimensional manifold with boundary.
Connor Mooney, Zhongjian Wang, Jack Xin
ICLR3
2024 Rethinking the Benefits of Steerable Features in 3D Equivariant Graph Neural Networks
abstract
Theoretical and empirical comparisons have been made to assess the expressive power and performance of invariant and equivariant GNNs. However, there is currently no theoretical result comparing the expressive power of $k$-hop invariant GNNs and equivariant GNNs. Additionally, little is understood about whether the performance of equivariant GNNs, employing steerable features up to type-$L$, increases as $L$ grows -- especially when the feature dimension is held constant. In this study, we introduce a key lemma that allows us to analyze steerable features by examining their corresponding invariant features. The lemma facilitates us in understanding the limitations of $k$-hop invariant GNNs, which fail to capture the global geometric structure due to the loss of geometric information between local structures. Furthermore, we investigate the invariant features associated with different types of steerable features and demonstrate that the expressiveness of steerable features is primarily determined by their dimension -- independent of their irreducible decomposition. This suggests that when the feature dimension is constant, increasing $L$ does not lead to essentially improved performance in equivariant GNNs employing steerable features up to type-$L$. We substantiate our theoretical insights with numerical evidence.
Shih-Hsin Wang, Yung-Chang Hsu, Justin M. Baker, Andrea L. Bertozzi, Jack Xin, Bao Wang 0001
ICLR5
2023 Weighted Anisotropic-Isotropic Total Variation for Poisson Denoising
abstract
Poisson noise commonly occurs in images captured by photon-limited imaging systems such as in astronomy and medicine. As the distribution of Poisson noise depends on the pixel intensity value, noise levels vary from pixels to pixels. Hence, denoising a Poisson-corrupted image while preserving important details can be challenging. In this paper, we propose a Poisson denoising model by incorporating the weighted anisotropic–isotropic total variation (AITV) as a regularization. We then develop an alternating direction method of multipliers with a combination of a proximal operator for an efficient implementation. Lastly, numerical experiments demonstrate that our algorithm outperforms other Poisson denoising methods in terms of image quality and computational efficiency.
Kevin Bui, Yifei Lou, Fredrick Park, Jack Xin
ICIP4
2022 Glassoformer: A Query-Sparse Transformer for Post-Fault Power Grid Voltage Prediction
abstract
We propose GLassoformer, a novel and efficient transformer architecture leveraging group Lasso regularization to reduce the number of queries of the standard self-attention mechanism. Due to the sparsified queries, GLassoformer is more computationally efficient than the standard transformers. On the power grid post-fault voltage prediction task, GLasso-former shows remarkably better prediction than many existing benchmark algorithms in terms of accuracy and stability.
Yunling Zheng, Carson Hu, Guang Lin 0001, Meng Yue 0001, Bao Wang 0001, Jack Xin
ICASSP6
2022 An Integrated Recurrent Neural Network and Regression Model with Spatial and Climatic Couplings for Vector-borne Disease Dynamics
abstract
We developed an integrated recurrent neural network and nonlinear regression spatio-temporal model for vector-borne disease evolution. We take into account climate data and seasonality as external factors that correlate with disease transmitting insects (e.g. flies), also spill-over infections from neighboring regions surrounding a region of interest. The climate data is encoded to the model through a quadratic embedding scheme motivated by recommendation systems. The neighboring regions' influence is modeled by a long short-term memory neural network. The integrated model is trained by stochastic gradient descent and tested on leish-maniasis data in Sri Lanka from 2013-2018 where infection outbreaks occurred. Our model outperformed ARIMA models across a number of regions with high infections, and an associated ablation study renders support to our modeling hypothesis and ideas.
Jack Xin, Guofa Zhou
ICPRAM2
2021 A Spatial-temporal Graph based Hybrid Infectious Disease Model with Application to COVID-19
Yunling Zheng, Jack Xin, Guofa Zhou
ICPRAM3
2021 A Weighted Difference of Anisotropic and Isotropic Total Variation for Relaxed Mumford-Shah Color and Multiphase Image Segmentation
abstract
In a class of piecewise-constant image segmentation models, we propose to incorporate a weighted difference of anisotropic and isotropic total variation (AITV) to regularize the partition boundaries in an image. In particular, we replace the total variation regularization in the Chan--Vese segmentation model and a fuzzy region competition model by the proposed AITV. To deal with the nonconvex nature of AITV, we apply the difference-of-convex algorithm (DCA), in which the subproblems can be minimized by the primal-dual hybrid gradient method with linesearch. The convergence of the DCA scheme is analyzed. In addition, a generalization to color image segmentation is discussed. In the numerical experiments, we compare the proposed models with the classic convex approaches and the two-stage segmentation methods (smoothing and then thresholding) on various images, showing that our models are effective in image segmentation and robust with respect to impulsive noises.
Kevin Bui, Fredrick Park, Yifei Lou, Jack Xin
SIAM J. Imaging Sci.4
2020 AutoShuffleNet: Learning Permutation Matrices via an Exact Lipschitz Continuous Penalty in Deep Convolutional Neural Networks
abstract
ShuffleNet is a state-of-the-art light weight convolutional neural network architecture. Its basic operations include group, channel-wise convolution and channel shuffling. However, channel shuffling is manually designed on empirical grounds. Mathematically, shuffling is a multiplication by a permutation matrix. In this paper, we propose to automate channel shuffling by learning permutation matrices in network training. We introduce an exact Lipschitz continuous non-convex penalty so that it can be incorporated in the stochastic gradient descent to approximate permutation at high precision. Exact permutations are obtained by simple rounding at the end of training and are used in inference. The resulting network, referred to as AutoShuffleNet, achieved improved classification accuracies on data from CIFAR-10, CIFAR-100 and ImageNet while preserving the inference costs of ShuffleNet. In addition, we found experimentally that the standard convex relaxation of permutation matrices into stochastic matrices leads to poor performance. We prove theoretically the exactness (error bounds) in recovering permutation matrices when our penalty function is zero (very small). We present examples of permutation optimization through graph matching and two-layer neural network models where the loss functions are calculated in closed analytical form. In the examples, convex relaxation failed to capture permutations whereas our penalty succeeded.
Jiancheng Lyu, Shuai Zhang 0009, Yingyong Qi, Jack Xin
KDD4
2019 Understanding Straight-Through Estimator in Training Activation Quantized Neural Nets
Penghang Yin, Jiancheng Lyu, Shuai Zhang 0009, Stanley J. Osher, Yingyong Qi, Jack Xin
ICLR (Poster)6
2018 BinaryRelax: A Relaxation Approach for Training Deep Neural Networks with Quantized Weights
abstract
We propose BinaryRelax, a simple two-phase algorithm, for training deep neural networks with quantized weights. The set constraint that characterizes the quantization of weights is not imposed until the late stage of training, and a sequence of pseudo quantized weights is maintained. Specifically, we relax the hard constraint into a continuous regularizer via a Moreau envelope, which turns out to be the squared Euclidean distance to the set of quantized weights. The pseudo quantized weights are obtained by linearly interpolating between the float weights and their quantizations. A continuation strategy is adopted to push the weights toward the quantized state by gradually increasing the regularization parameter. In the second phase, an exact quantization scheme with a small learning rate is invoked to guarantee fully quantized weights. We test BinaryRelax on the benchmark CIFAR and ImageNet color image datasets to demonstrate the superiority of the relaxed quantization approach and the improved accuracy over the state-of-the-art training methods. Finally, we prove the convergence of BinaryRelax under an approximate orthogonality condition.
Penghang Yin, Shuai Zhang 0009, Jiancheng Lyu, Stanley J. Osher, Yingyong Qi, Jack Xin
SIAM J. Imaging Sci.6
2016 A weighted difference of anisotropic and isotropic total variation for relaxed Mumford-Shah image segmentation
abstract
We propose to incorporate a weighted difference of anisotropic and isotropic total variation (TV) norms into a relaxed formulation of the two phase Mumford-Shah (MS) model for image segmentation. We show results exceeding those obtained by the MS model when using the standard TV norm to regularize partition boundaries. In particular, examples illustrating the qualitative differences between the proposed model and the standard MS one are shown. A fast numerical method is introduced to minimize the proposed model utilizing the difference-of-convex algorithm (DCA) and the primal dual hybrid gradient (PDHG) method.
Fredrick Park, Yifei Lou, Jack Xin
ICIP3
2015 Parallelization of a color-entropy preprocessed Chan-Vese model for face contour detection on multi-core CPU and GPU
Xiaohua Shi, Fredrick Park, Jack Xin, Yingyong Qi
Parallel Comput.4
2015 A Weighted Difference of Anisotropic and Isotropic Total Variation Model for Image Processing
abstract
We propose a weighted difference of anisotropic and isotropic total variation (TV) as a regularization for image processing tasks, based on the well-known TV model and natural image statistics. Due to the form of our model, it is natural to compute via a difference of convex algorithm (DCA). We draw its connection to the Bregman iteration for convex problems and prove that the iteration generated from our algorithm converges to a stationary point with the objective function values decreasing monotonically. A stopping strategy based on the stable oscillatory pattern of the iteration error from the ground truth is introduced. In numerical experiments on image denoising, image deblurring, and magnetic resonance imaging (MRI) reconstruction, our method improves on the classical TV model consistently and is on par with representative state-of-the-art methods.
Yifei Lou, Tieyong Zeng, Stanley J. Osher, Jack Xin
SIAM J. Imaging Sci.4
2014 Partially Blind Deblurring of Barcode from Out-of-Focus Blur
abstract
This paper addresses the nonstationary out-of-focus (OOF) blur removal in the application of barcode reconstruction. We propose a partially blind deblurring method when partial knowledge of the clean barcode is available. In particular, we consider an image formation model based on geometrical optics, which involves the point-spread function (PSF) for the OOF blur. With the known information, we can estimate a low-dimensional representation of the PSF using the Levenberg--Marquardt algorithm. Once the PSF is obtained, the deblurred image is computed by solving a quadratic program. We find that imposing a [0,1] box constraint is often good enough to enforce binary signal. Experiments on real data demonstrate that the forward model is physically realistic and our partially blind deblurring method can yield good reconstructions.
Yifei Lou, Ernie Esser, Hongkai Zhao, Jack Xin
SIAM J. Imaging Sci.4
2014 A sparse semi-blind source identification method and its application to Raman spectroscopy for explosives detection
Yuanchang Sun, Jack Xin
Signal Process.2
2013 A randomly perturbed infomax algorithm for blind source separation
abstract
We present a novel modification to the well-known infomax algorithm of blind source separation. Under natural gradient descent, the infomax algorithm converges to a stationary point of a limiting ordinary differential equation. However, due to the presence of saddle points or local minima of the corresponding likelihood function, the algorithm may be trapped around these “bad” stationary points for a long time, especially if the initial data are near them. To speed up convergence, we propose to add a sequence of random perturbations to the infomax algorithm to “shake” the iterating sequence so that it is “captured” by a path descending to a more stable stationary point. We analyze the convergence of the randomly perturbed algorithm, and illustrate its fast convergence through numerical examples on blind demixing of stochastic signals. The examples have analytical structures so that saddle points or local minima of the likelihood functions are explicit.
Jack Xin
ICASSP2
2013 A Method for Finding Structured Sparse Solutions to Nonnegative Least Squares Problems with Applications
abstract
Unmixing problems in many areas such as hyperspectral imaging and differential optical absorption spectroscopy (DOAS) often require finding sparse nonnegative linear combinations of dictionary elements that match observed data. We show how aspects of these problems, such as misalignment of DOAS references and uncertainty in hyperspectral endmembers, can be modeled by expanding the dictionary with grouped elements and imposing a structured sparsity assumption that the combinations within each group should be sparse or even 1-sparse. If the dictionary is highly coherent, it is difficult to obtain good solutions using convex or greedy methods, such as nonnegative least squares (NNLS) or orthogonal matching pursuit. We use penalties related to the Hoyer measure, which is the ratio of the $l_1$ and $l_2$ norms, as sparsity penalties to be added to the objective in NNLS-type models. For solving the resulting nonconvex models, we propose a scaled gradient projection algorithm that requires solving a sequence of strongly convex quadratic programs. We discuss its close connections to convex splitting methods and difference of convex programming. We also present promising numerical results for DOAS analysis and hyperspectral unmixing problems.
Ernie Esser, Yifei Lou, Jack Xin
SIAM J. Imaging Sci.3
2012 A Triple-Microphone Real-Time Speech Enhancement Algorithm Based on Approximate Array Analytical Solutions
Meng Yu 0003, Ryan Ritch, Jack Xin
INTERSPEECH3
2012 Exploring Off Time Nature for Speech Enhancement
Meng Yu 0003, Jack Xin
INTERSPEECH2
2012 Nonnegative Sparse Blind Source Separation for NMR Spectroscopy by Data Clustering, Model Reduction, and 1 Minimization
abstract
Motivated by applications in nuclear magnetic resonance (NMR) spectroscopy, we introduce a novel blind source separation (BSS) approach to treat nonnegative and correlated data. We consider the (over)-determined case where $n$ sources are to be separated from $m$ linear mixtures ($m\geq n$). Among the $n$ source signals, there are $n-1$ partially overlapping (Po) sources and one positive everywhere (Pe) source. This condition is applicable for many real-world signals such as NMR spectra of urine and blood serum for metabolic fingerprinting and disease diagnosis. The geometric properties of the mixture matrix and the sparseness structure of the source signals (in a transformed domain) are crucial to the identification of the mixing matrix and the sources. The method first identifies the mixing coefficients of the Pe source by exploiting geometry in data clustering. Then subsequent elimination of variables leads to a sub-BSS problem of the Po sources solvable by the minimal cone method and related linear programming. The last step is based on solving a convex $\ell_1$ minimization problem to extract the Pe source signals. Numerical results on NMR spectra show satisfactory performance of the method.
Yuanchang Sun, Jack Xin
SIAM J. Imaging Sci.2
2012 Multi-Channel l1 Regularized Convex Speech Enhancement Model and Fast Computation by the Split Bregman Method
abstract
A convex speech enhancement (CSE) method is presented based on convex optimization and pause detection of the speech sources. Channel spatial difference is identified for enhancing each speech source individually while suppressing other interfering sources. Sparse unmixing filters indicating channel spatial differences are sought byl1norm regularization and the split Bregman method. A subdivided split Bregman method is developed for efficiently solving the problem in severely reverberant environments. The speech pause detection is based on a binary mask source separation method. The CSE method is evaluated objectively and subjectively, and found to outperform a list of existing blind speech separation approaches on both synthetic and room recorded speech mixtures in terms of the overall computational speed and separation quality.
Meng Yu 0003, Wenye Ma, Jack Xin, Stanley J. Osher
IEEE Trans. Speech Audio Process.3
2012 A Convex Model for Nonnegative Matrix Factorization and Dimensionality Reduction on Physical Space
abstract
A collaborative convex framework for factoring a data matrix X into a nonnegative product AS , with a sparse coefficient matrix S, is proposed. We restrict the columns of the dictionary matrix A to coincide with certain columns of the data matrix X, thereby guaranteeing a physically meaningful dictionary and dimensionality reduction. We use l(1, ∞) regularization to select the dictionary from the data and show that this leads to an exact convex relaxation of l(0) in the case of distinct noise-free data. We also show how to relax the restriction-to- X constraint by initializing an alternating minimization approach with the solution of the convex model, obtaining a dictionary close to but not necessarily in X. We focus on applications of the proposed framework to hyperspectral endmember and abundance identification and also show an application to blind source separation of nuclear magnetic resonance data.
Ernie Esser, Michael Möller 0001, Stanley J. Osher, Guillermo Sapiro, Jack Xin
IEEE Trans. Image Process.5
2011 A Recursive Sparse Blind Source Separation Method for Nonnegative and Correlated Data in NMR Spectroscopy
Yuanchang Sun, Jack Xin
CAIP (2)2
2011 Content Adaptive Image Matching by Color-Entropy Segmentation and Inpainting
Yuanchang Sun, Jack Xin
CAIP (2)2
2011 Modeling Category Identification Using Sparse Instance Representation
Shunan Zhang, Michael D. Lee 0001, Meng Yu 0003, Jack Xin
CogSci4
2011 Postprocessing and sparse blind source separation of positive and partially overlapped data
Yuanchang Sun, C. Ridge, Federico del Rio-Portilla, Athan J. Shaka, Jack Xin
Signal Process.5
2010 Reducing musical noise in blind source separation by time-domain sparse filters and split bregman method
abstract
Musical noise often arises in the outputs of time-frequency binary mask based blind source separation approaches. Postprocessing is desired to enhance the separation quality. An efficient musical noise reduction method by time-domain sparse filters is presented using convex optimization. The sparse filters are sought by l1 regularization and the split Bregman method. The proposed musical noise reduction method is evaluated by both synthetic and room recorded speech and music data, and found to outperform existing musical noise reduction methods in terms of the objective and subjective measures. Index Terms: Musical noise, time-frequency mask, timedomain sparse filters, split Bregman method.
Wenye Ma, Meng Yu 0003, Jack Xin, Stanley J. Osher
INTERSPEECH3
2010 Convexity and fast speech extraction by split bregman method
abstract
A fast speech extraction (FSE) method is presented using convex optimization made possible by pause detection of the speech sources. Sparse unmixing filters are sought by l1 regularization and the split Bregman method. A subdivided split Bregman method is developed for efficiently estimating long reverberations in real room recordings. The speech pause detection is based on a binary mask source separation method. The FSE method is evaluated and found to outperform existing blind speech separation approaches on both synthetic and room recorded data in terms of the overall computational speed and separation quality. Index Terms: convexity, sparse filters, split Bregman method, fast blind speech extraction.
Meng Yu 0003, Wenye Ma, Jack Xin, Stanley J. Osher
INTERSPEECH3
2010 A Critical Quantity for Noise Attenuation in Feedback Systems
abstract
Feedback modules, which appear ubiquitously in biological regulations, are often subject to disturbances from the input, leading to fluctuations in the output. Thus, the question becomes how a feedback system can produce a faithful response with a noisy input. We employed multiple time scale analysis, Fluctuation Dissipation Theorem, linear stability, and numerical simulations to investigate a module with one positive feedback loop driven by an external stimulus, and we obtained a critical quantity in noise attenuation, termed as "signed activation time". We then studied the signed activation time for a system of two positive feedback loops, a system of one positive feedback loop and one negative feedback loop, and six other existing biological models consisting of multiple components along with positive and negative feedback loops. An inverse relationship is found between the noise amplification rate and the signed activation time, defined as the difference between the deactivation and activation time scales of the noise-free system, normalized by the frequency of noises presented in the input. Thus, the combination of fast activation and slow deactivation provides the best noise attenuation, and it can be attained in a single positive feedback loop system. An additional positive feedback loop often leads to a marked decrease in activation time, decrease or slight increase of deactivation time and allows larger kinetic rate variations for slow deactivation and fast activation. On the other hand, a negative feedback loop may increase the activation and deactivation times. The negative relationship between the noise amplification rate and the signed activation time also holds for the six other biological models with multiple components and feedback loops. This principle may be applicable to other feedback systems.
Jack Xin, Qing Nie
PLoS Comput. Biol.2
2008 A dynamic algorithm for blind separation of convolutive sound mixtures
Jie Liu 0003, Jack Xin, Yingyong Qi
Neurocomputing2
2006 Signal processing of acoustic signals in the time domain with an active nonlinear nonlocal cochlear model
Michael Drew Lamar, Jack Xin, Yingyong Qi
Signal Process.2
2000 A perception and PDE based nonlinear transformation for processing spoken words
abstract
Speech signals are often produced or received in the presence of noise, which is known to degrade the performance of a speech recognition system. In this paper, a perception- and PDE-based nonlinear transformation was developed to process spoken words in noisy environment. Our goal is to distinguish essential speech features and suppress noise so that the processed words are better recognized by a computer software. The nonlinear transformation was made on the spectrogram (short-term Fourier spectra) of speech signals, which reveals the signal energy distribution in time and frequency. The transformation reduces noise through time adaptation (reducing temporally slowly varying portions of spectra) and enhances spectral peaks (formants) by evolving a focusing quadratic fourth-order PDE. Short-term spectra of speech signals were initially divided into three (low, mid and high) frequency bands based on the critical bandwidth of human audition. An algorithm was developed to trace the upper and lower intensity envelopes of signal in each band. The difference between the upper and lower envelopes reflects the signal-to-noise (SNR) ratio of each band. Constant, low SNR signals in each band were adaptively decreased to reduce noise. Then evolution of the focusing PDE was used to enhance the spectral peaks, and further reduce noise interference. Numerical results on noisy spoken words indicated that the transformed spectral pattern of the spoken words was insensitive to noise for SNR ranging from 0 to 20 dB (decibel). The spectral distances between noisy words and original words decreased after the transformation. A numerical experiment was performed on 11 spoken words at SNR D 5 dB. A noisy word is recognized numerically by computing the closest L 2 spectral distance from the clean template. The experiment reached a recognition rate as high as 100%. Analyses on the properties of the transformation are provided. © 2001 Elsevier Science
Yingyong Qi, Jack Xin
INTERSPEECH2